System
A system that collects user data to estimate psychological state and generate personalized playlists, addressing the limitations of conventional music distribution by providing mood-matched music and real-time updates.
Patent Information
- Application Number
- JP2024123916
- Authority / Receiving Office
- JP · JP
- Patent Type
- Applications
- Current Assignee / Owner
- Filing Date
- 2024-07-30
- Publication Date
- 2026-02-12
AI Technical Summary
Conventional music distribution services struggle to provide music that matches a user's mood or psychological state, as they primarily rely on taste and preference without considering the user's environment, time of day, or daily activities, and lack flexibility in responding to mood changes during playback.
A system that collects user location, behavioral, and music preference data to estimate the current psychological state, generates a personalized playlist, and allows real-time updates based on user requests.
Provides music experiences that accurately match the user's mood and daily situation, offering flexibility and responsiveness to user preferences.
Smart Images

Figure 2026022399000001_ABST
Abstract
Description
[Technical Field]
[0001] The technology of the present disclosure relates to a system. [Background technology]
[0002] Patent document 1 discloses a persona chatbot control method performed by at least one processor, the method including the steps of receiving a user utterance, adding the user utterance to a prompt including an instruction sentence related to a description of the chatbot character, encoding the prompt, and inputting the encoded prompt into a language model to generate a chatbot utterance in response to the user utterance. [Prior art documents] [Patent documents]
[0003] [Patent Document 1] Japanese Patent Publication No. 2022-180282 Summary of the Invention [Problem to be solved by the invention]
[0004] Conventional music distribution services primarily provide recommendations based on a user's tastes and preferences, making it difficult to provide music that matches the user's mood or psychological state. Taking into account the user's environment, time of day, and daily activity history would enable more appropriate music selection, but no system for doing so has yet been developed. There is also a need for a flexible mechanism that can easily reflect a user's request if their mood changes while they are listening to a playlist. The present invention aims to solve these problems and provide a system that provides a music experience that matches the user's current psychological state and daily situation. [Means for solving the problem]
[0005] The present invention provides a means for collecting a user's current location information, behavioral data, and music preference data. It also includes a means for integrating the collected data and estimating the user's current psychological state. It also provides a means for generating an optimal music playlist based on the estimated psychological state and music preference data, and a means for transmitting the generated playlist to the user's device. It also provides a means for receiving a request from the user, reevaluating and updating the playlist, and a means for transmitting the updated playlist to the user's device. This allows the user to enjoy a music experience that perfectly matches their current mood.
[0006] "Device" means an electronic device owned by a User that collects and transmits location information, behavioral data, and music preference data.
[0007] "Location information" means data indicating a geographical location obtained by a terminal's GPS function or other location measurement function.
[0008] "Behavioral data" is data that indicates the user's daily behavior history, and includes, for example, payment history, schedule, activities, and the like.
[0009] "Music preference data" is data that indicates the user's music playback history and music evaluations.
[0010] "Mental state" refers to the user's current mood and emotional state estimated based on collected data.
[0011] A "playlist" is a list of songs selected based on specific criteria and sent to a user's terminal.
[0012] A "request" is a request made by a user for more suitable music for a currently playing playlist.
[0013] "Server" is a computer system that receives data sent from terminals, integrates and analyzes it, and generates and updates playlists. [Brief explanation of the drawings]
[0014] [Figure 1] 1 is a conceptual diagram showing an example of the configuration of a data processing system according to a first embodiment. [Figure 2] 1 is a conceptual diagram showing an example of main functions of a data processing device and a smart device according to a first embodiment. [Figure 3] FIG. 10 is a conceptual diagram showing an example of the configuration of a data processing system according to a second embodiment. [Figure 4] FIG. 10 is a conceptual diagram showing an example of main functions of a data processing device and smart glasses according to a second embodiment. [Figure 5] FIG. 10 is a conceptual diagram showing an example of the configuration of a data processing system according to a third embodiment. [Figure 6] FIG. 11 is a conceptual diagram showing an example of main functions of a data processing device and a headset-type terminal according to a third embodiment. [Figure 7] FIG. 10 is a conceptual diagram showing an example of the configuration of a data processing system according to a fourth embodiment. [Figure 8] FIG. 10 is a conceptual diagram showing an example of main functions of a data processing device and a robot according to a fourth embodiment. [Figure 9] 1 shows an emotion map onto which multiple emotions are mapped. [Figure 10] 1 shows an emotion map onto which multiple emotions are mapped. [Figure 11] FIG. 3 is a sequence diagram showing a processing flow of the data processing system according to the first embodiment. [Figure 12] FIG. 10 is a sequence diagram showing the flow of processing in the data processing system in Application Example 1. [Figure 13] FIG. 10 is a sequence diagram showing the flow of processing in the data processing system according to the second embodiment when an emotion engine is combined. [Figure 14]FIG. 10 is a sequence diagram showing the flow of processing in the data processing system in Application Example 2 when an emotion engine is combined. DETAILED DESCRIPTION OF THE INVENTION
[0015] An example of an embodiment of a system according to the technology of the present disclosure will be described below with reference to the accompanying drawings.
[0016] First, the terms used in the following description will be explained.
[0017] In the following embodiments, a coded processor (hereinafter simply referred to as a "processor") may be a single arithmetic device or a combination of multiple arithmetic devices. Furthermore, a processor may be a single type of arithmetic device or a combination of multiple types of arithmetic devices. Examples of arithmetic devices include a CPU (Central Processing Unit), a GPU (Graphics Processing Unit), a GPGPU (General-Purpose computing on Graphics Processing Units), and an APU (Accelerated Processing Unit).
[0018] In the following embodiments, a coded RAM (Random Access Memory) is a memory in which information is temporarily stored and is used as a working memory by a processor.
[0019] In the following embodiments, the coded storage is one or more non-volatile storage devices that store various programs, various parameters, etc. Examples of non-volatile storage devices include flash memory (SSD (Solid State Drive)), magnetic disks (e.g., hard disks), and magnetic tapes.
[0020] In the following embodiments, a communication I / F (Interface) with a symbol is an interface including a communication processor, an antenna, etc. The communication I / F controls communication between multiple computers. Examples of communication standards applied to the communication I / F include wireless communication standards including 5G (5th Generation Mobile Communication System), Wi-Fi (registered trademark), Bluetooth (registered trademark), etc.
[0021] In the following embodiments, "A and / or B" is synonymous with "at least one of A and B." In other words, "A and / or B" means that it may be only A, only B, or a combination of A and B. Furthermore, in this specification, the same concept as "A and / or B" is also applied when three or more things are expressed connected by "and / or."
[0022] [First embodiment]
[0023] FIG. 1 shows an example of the configuration of a data processing system 10 according to the first embodiment.
[0024] 1, a data processing system 10 includes a data processing device 12 and a smart device 14. An example of the data processing device 12 is a server.
[0025] The data processing device 12 includes a computer 22, a database 24, and a communication I / F 26. The computer 22 is an example of a "computer" according to the technology of the present disclosure. The computer 22 includes a processor 28, a RAM 30, and a storage 32. The processor 28, the RAM 30, and the storage 32 are connected to a bus 34. The database 24 and the communication I / F 26 are also connected to the bus 34. The communication I / F 26 is connected to a network 54. Examples of the network 54 include a WAN (Wide Area Network) and / or a LAN (Local Area Network).
[0026] The smart device 14 includes a computer 36, a reception device 38, an output device 40, a camera 42, and a communication I / F 44. The computer 36 includes a processor 46, a RAM 48, and a storage 50. The processor 46, the RAM 48, and the storage 50 are connected to a bus 52. The reception device 38, the output device 40, and the camera 42 are also connected to the bus 52.
[0027] The reception device 38 includes a touch panel 38A, a microphone 38B, and the like, and receives user input. The touch panel 38A detects contact with an indicator (for example, a pen or a finger) to receive user input by the touch of the indicator. The microphone 38B detects the user's voice to receive user input by voice. The control unit 46A transmits data indicating the user input received by the touch panel 38A and the microphone 38B to the data processing device 12. In the data processing device 12, the specific processing unit 290 acquires the data indicating the user input.
[0028] The output device 40 includes a display 40A and a speaker 40B, and presents data to the user 20 by outputting the data in a form of expression that the user 20 can perceive (for example, audio and / or text). The display 40A displays visible information such as text and images in accordance with instructions from the processor 46. The speaker 40B outputs audio in accordance with instructions from the processor 46. The camera 42 is a compact digital camera equipped with an optical system including a lens, aperture, and shutter, and an imaging element such as a CMOS (Complementary Metal-Oxide-Semiconductor) image sensor or a CCD (Charge Coupled Device) image sensor.
[0029] The communication I / F 44 is connected to a network 54. The communication I / Fs 44 and 26 control the exchange of various information between the processor 46 and the processor 28 via the network 54.
[0030] FIG. 2 shows an example of the main functions of the data processing device 12 and the smart device 14.
[0031] 2, in the data processing device 12, a specific process is performed by the processor 28. A specific processing program 56 is stored in the storage 32. The specific processing program 56 is an example of a "program" according to the technology of the present disclosure. The processor 28 reads the specific processing program 56 from the storage 32 and executes the read specific processing program 56 on the RAM 30. The specific process is realized by the processor 28 operating as a specific processing unit 290 in accordance with the specific processing program 56 executed on the RAM 30.
[0032] The storage 32 stores a data generation model 58 and an emotion identification model 59. The data generation model 58 and the emotion identification model 59 are used by the identification processing unit 290.
[0033] In the smart device 14, the processor 46 performs the reception output process. The storage 50 stores a reception output program 60. The reception output program 60 is used in conjunction with the specific processing program 56 by the data processing system 10. The processor 46 reads the reception output program 60 from the storage 50 and executes the read reception output program 60 on the RAM 48. The reception output process is realized by the processor 46 operating as the control unit 46A in accordance with the reception output program 60 executed on the RAM 48.
[0034] Next, a description will be given of the specific processing performed by the specific processing unit 290 of the data processing device 12. In the following description, the data processing device 12 will be referred to as a "server" and the smart device 14 will be referred to as a "terminal."
[0035] To implement the present invention, a user terminal, a server, and a communication system that links them together are required. A specific embodiment of this system will be described below.
[0036] Data Collection Phase
[0037] The device collects the user's current location information. Using the GPS function, it periodically records the user's current location and sends it to the server. The device also collects the user's daily behavioral data. This behavioral data includes payment history, calendar schedule information, and fitness tracker data. The device also collects the user's music preference data. Specifically, this includes the history of music that has been played in the past and the ratings that the user has given to songs. This data is also sent to the server.
[0038] Mood estimation phase
[0039] The server receives location information, behavioral data, and music preference data sent from the device and integrates them to create a user profile. Based on this profile, the server runs an algorithm to estimate the user's current psychological state. For example, say you work out at the gym at 7:00 AM and then have coffee at a cafe at 9:00 AM. From these actions, the server estimates your psychological state as "refreshed and active."
[0040] Playlist Generation Phase
[0041] The server selects the most suitable songs based on the estimated psychological state and music preference data. It then generates a playlist using the selected songs. The generated playlist includes song titles, artist names, playback links, and other information. This playlist is then sent to the device.
[0042] Request handling phase
[0043] A user can make requests for a playlist that is currently being played, such as "more upbeat songs" or "more calming songs." This request is sent to the server via the device. The server reevaluates the playlist based on the received request and adds new songs. In this way, the playlist can respond to the user's detailed requests. The updated playlist is then sent back to the device, where the user can listen to it.
[0044] Specific examples
[0045] Consider a scenario where a user goes to the gym at 7am and then has coffee at a cafe at 9am.
[0046] 1. The device records these activities as location and behavioral data and sends them to a server.
[0047] 2. The server receives this data and assumes the user is "refreshed and active."
[0048] 3. The server generates a playlist containing up-tempo pop songs that fit this state of mind and sends it to the device.
[0049] 4. The user starts listening to the playlist and requests a lighter song.
[0050] 5. The server receives this request, updates the playlist to include more upbeat songs, and sends it back to the device.
[0051] 6. The user can enjoy the updated playlist.
[0052] In this way, it is possible to provide music that best suits the user's mood.
[0053] The processing flow will be explained below.
[0054] Step 1:
[0055] The device uses the GPS function to obtain geographical data to obtain the user's current location information, and then sends the obtained location information to the server.
[0056] Step 2:
[0057] The device collects data about the user's daily activities, such as payment history, calendar schedule, and fitness tracker data, and transmits this data to a server.
[0058] Step 3:
[0059] The device collects the user's music preference data, specifically the history of songs played in the past and the user's ratings of songs, and transmits this data to the server.
[0060] Step 4:
[0061] The server combines location, behavioral, and music preference data to create a user profile that comprehensively describes the user's current situation.
[0062] Step 5:
[0063] The server estimates the user's current state of mind based on the user profile. For example, if the user went to the gym in the morning and then spent time at a cafe, the server estimates that the user feels refreshed and active.
[0064] Step 6:
[0065] The server selects the most suitable songs based on the estimated psychological state and music preference data, and generates a playlist based on the selected songs.
[0066] Step 7:
[0067] The server sends the generated playlist to the device, which includes detailed information about the selected songs (song title, artist name, playback link, etc.).
[0068] Step 8:
[0069] Users can make requests for songs on a playlist, such as "more upbeat songs" or "more calm songs," and these requests are sent to the server via the terminal.
[0070] Step 9:
[0071] The server receives the request from the user, re-evaluates the playlist according to the request, reselects songs based on the new criteria, and updates the playlist.
[0072] Step 10:
[0073] The server then sends the updated playlist back to the terminal, where the user can play the updated playlist and enjoy the new songs.
[0074] Through the above processing steps, a musical experience that matches the user's current mental state and specific requests can be provided.
[0075] Example 1
[0076] Next, a description will be given of Example 1. In the following description, the data processing device 12 will be referred to as a "server" and the smart device 14 will be referred to as a "terminal."
[0077] In modern society, users are surrounded by vast amounts of information and data. Providing music that matches a user's psychological state and mood is crucial for improving individual satisfaction. However, conventional music recommendation systems struggle to comprehensively analyze a user's current location, behavioral data, and musical preference data to provide optimal music in real time. They also lack the ability to flexibly update playlists in response to user requests. Given these circumstances, there is a need for a system that can provide music that matches a user's current psychological state and quickly respond to user requests.
[0078] The specific processing by the specific processing unit 290 of the data processing device 12 in the first embodiment is realized by the following means.
[0079] In this invention, the server
[0080] means for collecting current location information from a user's information processing device;
[0081] means for collecting behavioral data from a user's information processing device;
[0082] means for collecting music preference data from a user's information processing device;
[0083] This allows the system to integrate collected data to estimate the user's current psychological state using a generative AI model, generate an optimal music playlist based on the estimation results, and reevaluate and update the playlist according to user requests.
[0084] An "information processing device" is a device that collects, processes, stores, and communicates data, and examples of such devices include smartphones, tablets, and wearable devices.
[0085] "Current location information" is data indicating the geographic location of the user, and is information obtained using GPS or other location detection technology.
[0086] "Behavioral data" is information about a user's daily activities, including payment history, calendar schedule information, fitness tracker data, and the like.
[0087] "Music preference data" is information relating to the user's tastes and ratings of music, and includes the history of songs played in the past and rating data.
[0088] A "generative AI model" is a model of an algorithm or program that uses artificial intelligence technology to analyze and generate data, and is used to estimate a user's psychological state.
[0089] "Mental state" is an indicator of the user's current mental and emotional state, such as feeling refreshed, active, or relaxed.
[0090] A "playlist" is a list of songs to be played in a specific order, and is generated based on the user's psychological state and musical preference data.
[0091] A "request" refers to a request or wish made by a user regarding a playlist that is being played, and specifically includes requests such as "a more upbeat song" or "a more calming song."
[0092] The present invention relates to a system that provides an optimal music playlist based on the user's psychological state and also responds to user requests. Implementing this system requires a user's information processing device, a server, and a communication system that links them.
[0093] Data Collection Phase
[0094] The device collects the user's current location information. This location information is obtained using GPS and updated periodically. Examples of devices like smartphones and wearable devices. The device also collects the user's daily behavioral data. This behavioral data includes payment history (collected via NFC or QR code readers), calendar schedule information (obtained from the Google Calendar API or Apple Calendar API), and fitness tracker data (obtained from the Apple HealthKit or Google Fit API). The device also collects the user's music preference data. Specifically, it collects playback history (using the Spotify API or Apple Music API) and song rating data. The collected data is sent from the device to a server.
[0095] Data analysis phase
[0096] The server receives location information, behavioral data, and music preference data sent from the device and integrates them to create a user profile. The profile reflects the user's behavioral patterns and music preferences. Based on this profile, the server uses a generative AI model to estimate the user's current psychological state. For example, from data such as "The user worked out at the gym at 7am and then had coffee at a cafe at 9am," the server generates a prompt that estimates the user's psychological state as "Refreshed and Active." The following prompt is a concrete example:
[0097] "A user works out at the gym at 7am and then has coffee at a cafe at 9am. Generate a playlist with appropriate uptempo pop songs to infer that the user is refreshed and active."
[0098] Playlist Generation Phase
[0099] The server selects the most suitable songs based on the estimated psychological state and music preference data. Songs are searched from a database to select those that suit the user's psychological state. For example, for a "refreshed and active" psychological state, up-tempo pop songs are selected. A playlist is generated using the selected songs, including information such as song title, artist name, and playback link. The generated playlist is sent from the server to the device.
[0100] Request handling phase
[0101] Users can make requests for songs, such as "more upbeat songs" or "more calming songs," to the playlist being played. These requests are sent to the server via the device. The server reevaluates the playlist based on the received request and adds new songs. The updated playlist is then sent back to the device, allowing the user to enjoy the new playlist.
[0102] In this way, the system not only provides music that is optimal for the user's psychological state, but also has the feature of being able to quickly respond to user requests. The above is an embodiment of the present invention.
[0103] The flow of the identification process in the first embodiment will be described with reference to FIG.
[0104] Step 1: Start collecting data
[0105] The device starts collecting data when the user starts using the system. The input is the user's action, and the output is the start of the data collection process. For example, this process can be started by pressing a "start data collection" button on the app or device.
[0106] Step 2: Obtaining location information
[0107] The device uses the GPS function to obtain the user's current location information. The input is location information from the GPS sensor, and the output is to temporarily store this information within the device. Specifically, the device obtains location information once every minute and records it as latitude and longitude.
[0108] Step 3: Behavioral data collection
[0109] The device collects user behavioral data. The input is data from various sensors and APIs, and the output is collected behavioral data. For example, payment history is obtained via an NFC sensor or QR code reader, and calendar schedule information is obtained from the Google Calendar API or Apple Calendar API. Fitness tracker data is also collected using Apple HealthKit or Google Fit API.
[0110] Step 4: Collecting Music Preference Data
[0111] The device collects the user's music preference data. The input is data from streaming service APIs (Spotify API and Apple Music API), and the output is the collection of music preference data. Specifically, it collects information on playback history and user actions such as "like" and "skip."
[0112] Step 5: Send data
[0113] The device sends the collected location information, behavioral data, and music preference data to a server. The input is various data stored on the device, and the output is sending this data to the server. This communication is carried out via the Internet.
[0114] Step 6: Create a user profile
[0115] The server creates a user profile based on the received data. The input is the data received by the server, and the output is the creation of a user profile. Specifically, the server organizes the data into categories and uses statistical methods and machine learning models to analyze the user's behavioral patterns and musical preferences.
[0116] Step 7: Estimate mental state
[0117] The server uses a generative AI model to estimate the user's psychological state. The input is the user profile and behavioral data collected at that time, and the output is the estimated psychological state. Specifically, a prompt sentence is input into the generative AI model, and the psychological state is estimated from the response. For example, it may estimate that "the user worked out at the gym at 7 a.m. and had coffee at a cafe at 9 a.m., so they are refreshed and active."
[0118] Step 8: Selecting songs and creating a playlist
[0119] The server selects the most suitable songs based on the estimated psychological state and music preference data. The input is the psychological state and music preference data, and the output is a generated playlist. Specifically, it searches for suitable songs from a database and creates a playlist. For example, it selects an up-tempo pop song that is suitable for the "active" psychological state.
[0120] Step 9: Send Playlist
[0121] The server sends the generated playlist to the terminal. The input is the generated playlist, and the output is the playlist arriving at the terminal. This communication is performed via the Internet, and real-time updates are possible.
[0122] Step 10: Respond to requests
[0123] A user can make requests to a playlist that is currently playing, such as "more upbeat songs" or "more calming songs." The input is the user's request, and the output is a new playlist that has been re-evaluated based on that request. Specifically, the server receives the request, re-evaluates the playlist, and adds new songs. This updated playlist is then sent back to the device, where the user can listen to it.
[0124] The above are the processing steps of the program of this system and its specific operations.
[0125] (Application example 1)
[0126] Next, a description will be given of Application Example 1. In the following description, the data processing device 12 will be referred to as a "server" and the smart device 14 will be referred to as a "terminal."
[0127] In recent years, the spread of self-driving vehicles has created a demand for improved in-car entertainment experiences. However, conventional in-car entertainment systems have faced the challenge of providing music and content that is tailored to the user's current psychological state and mood. There is a demand for improving in-car comfort and satisfaction by providing music that matches the user's mood.
[0128] The specific processing by the specific processing unit 290 of the data processing device 12 in the application example 1 is realized by the following means.
[0129] In this invention, the server includes means for collecting current location information from the user's terminal, means for collecting behavioral data from the user's terminal, and means for collecting music preference data from the user's terminal, thereby making it possible to provide a music playlist that is optimal for the user's mood in the autonomous vehicle.
[0130] A "user's terminal" is a mobile information terminal such as a smartphone or tablet owned by a user.
[0131] "Current location information" means a user's real-time geographic location data obtained using location measurement technology such as GPS.
[0132] "Behavioral data" refers to information about a user's behavior, such as fitness tracker data or schedule information.
[0133] "Music preference data" is data relating to the user's musical preferences, such as the history of songs that the user has played in the past and their ratings of those songs.
[0134] "Mental state" is information that indicates the user's current mood and emotional state.
[0135] A "playlist" is a list of multiple songs provided to a user.
[0136] A "request" is a request or instruction given by a user to a playlist.
[0137] An "autonomous vehicle" is a vehicle that operates autonomously and requires minimal user intervention.
[0138] An "entertainment system" is a system that provides entertainment content such as music and video in a vehicle.
[0139] To implement the present invention, a user terminal, a server, and a communication system that links them together are required. A specific embodiment of this system will be described below.
[0140] 1. Data Collection Phase
[0141] The device collects the user's current location information. Using the GPS function, it periodically records the user's current location and sends it to the server. The device also collects the user's daily behavioral data. This behavioral data includes payment history, calendar schedule information, and fitness tracker data. The device also collects the user's music preference data. Specifically, this includes the history of music that has been played in the past and the ratings that the user has given to songs. This data is also sent to the server.
[0142] 2. Mood estimation phase
[0143] The server receives location information, behavioral data, and music preference data sent from the device and integrates them to create a user profile. Based on this profile, the server runs an algorithm to estimate the user's current psychological state. For example, say you work out at the gym at 7:00 AM and then have coffee at a cafe at 9:00 AM. From these actions, the server estimates your psychological state as "refreshed and active."
[0144] 3. Playlist generation phase
[0145] The server selects the most suitable songs based on the estimated psychological state and music preference data. It then generates a playlist using the selected songs. The generated playlist includes song titles, artist names, playback links, and other information. This playlist is then sent to the device.
[0146] 4. Request Response Phase
[0147] A user can make requests for a playlist that is currently being played, such as "more upbeat songs" or "more calming songs." This request is sent to the server via the device. The server reevaluates the playlist based on the received request and adds new songs. In this way, the playlist can respond to the user's detailed requests. The updated playlist is then sent back to the device, where the user can listen to it.
[0148] Specific examples
[0149] Consider a scenario where a user goes to the gym at 7am and then has coffee at a cafe at 9am.
[0150] 1. The device records these activities as location and behavioral data and sends them to a server.
[0151] 2. The server receives this data and assumes the user is "refreshed and active."
[0152] 3. The server generates a playlist containing up-tempo music that matches this psychological state (e.g., "Uptown Funk") and sends it to the terminal.
[0153] 4. The user starts listening to the playlist and requests a lighter song.
[0154] 5. The server receives this request, updates the playlist to include more upbeat songs, and sends it back to the device.
[0155] 6. The user can enjoy the updated playlist.
[0156] Example of input prompt for generative AI model
[0157] "Generate a music playlist that best suits the user's mood based on the following information:
[0158] Current location: Coordinates (35.6895, 139.6917)
[0159] Behavioral data: 1 hour training at 6am, commuting in the car at 9am
[0160] Music preference data: List of songs you've played in the past (e.g. Happy - Williams, Uptown Funk - Mars)
[0161] Choose the best songs to match the user's mood and provide a playlist."
[0162] The flow of the specific processing in the application example 1 will be described with reference to FIG.
[0163] Step 1:
[0164] The device collects the user's current location information. Specifically, it periodically obtains the user's geographical location data using the GPS function. In addition to this, it also collects the user's daily behavior data (e.g., fitness tracker data and calendar schedule information). It also collects the user's music preference data (such as the history and ratings of music played in the past). This data is then sent to the server.
[0165] Input: User location information, behavioral data, music preference data
[0166] Output: Consolidated data packet sent to the server
[0167] Step 2:
[0168] The server receives location information, behavioral data, and music preference data sent from the device and integrates them to create a user profile. Specifically, it uses a data analysis algorithm to integrate this diverse information and generate a user profile.
[0169] Input: location information, behavioral data, music preference data
[0170] Output: Unified user profile
[0171] Step 3:
[0172] The server runs an algorithm to estimate the user's current psychological state based on the generated user profile. Specifically, it uses a machine learning model to analyze the user's behavioral patterns and music preferences, and estimates the user's current psychological state (e.g., "refreshed and active").
[0173] Input: User profile
[0174] Output: Estimated mental state
[0175] Step 4:
[0176] The server selects the most suitable songs based on the estimated psychological state and music preference data. Specifically, it searches a music preference database to create a list of songs that suit the user's current mood. The server then generates a playlist using the selected songs.
[0177] Input: Psychological state, music preference data
[0178] Output: Generated playlist
[0179] Step 5:
[0180] The server then sends the generated playlist to the user's device. Specifically, the playlist contains song titles, artist names, playback links, etc.
[0181] Input: Generated playlist
[0182] Output: Playlist sent to user device
[0183] Step 6:
[0184] The user issues a request (e.g., "more upbeat songs," "more calm songs") to the playlist being played. This request is sent to the server via the device. The server reevaluates the playlist based on the received request and adds new songs.
[0185] Input: A request from the user
[0186] Output: Updated playlist
[0187] Step 7:
[0188] The server then sends the updated playlist back to the terminal, allowing the user to listen to the latest music playlist.
[0189] Input: Updated playlist
[0190] Output: The updated playlist sent to the user's device
[0191] (Example of input prompt for generative AI model)
[0192] "Generate a music playlist that best suits the user's mood based on the following information:
[0193] Current location: Coordinates (35.6895, 139.6917)
[0194] Behavioral data: 1 hour training at 6am, commuting in the car at 9am
[0195] Music preference data: List of songs you've played in the past (e.g. Happy - Williams, Uptown Funk - Mars)
[0196] Choose the best songs to match the user's mood and provide a playlist."
[0197] Furthermore, an emotion engine that estimates the user's emotion may be combined. That is, the identification processing unit 290 may estimate the user's emotion using the emotion identification model 59 and perform identification processing using the user's emotion.
[0198] To implement the present invention, a user terminal, a server, an emotion engine, and a communication system that links these components are required. A specific embodiment of this system will be described below.
[0199] Data Collection Phase
[0200] The device collects the user's current location information. Using the GPS function, it periodically records the user's current location and sends it to the server. The device also collects the user's daily behavioral data. This behavioral data includes payment history, calendar schedule information, and fitness tracker data. The device also collects the user's music preference data. Specifically, this includes the history of music that has been played in the past and the ratings that the user has given to songs. This data is also sent to the server.
[0201] Emotion Recognition Phase
[0202] The device recognizes the user's emotions through an emotion engine. The emotion engine analyzes the user's voice, facial expressions, and biometric information (heart rate, body temperature, etc.) obtained from input devices. The emotion data recognized by the emotion engine is sent to the server along with other collected data.
[0203] Mood estimation phase
[0204] The server receives location information, behavioral data, music preference data, and emotional data from the device and integrates them to create a user profile. Based on this profile, the server runs an algorithm to estimate the user's current psychological state. For example, suppose a user exercises at the gym in the morning and then has coffee at a cafe. Based on these actions and emotions, the server estimates the user's psychological state as "refreshed and active."
[0205] Playlist Generation Phase
[0206] The server selects the most suitable songs based on the estimated psychological state and music preference data. It then generates a playlist using the selected songs. The generated playlist includes song titles, artist names, playback links, and other information. This playlist is then sent to the device.
[0207] Request handling phase
[0208] A user can make requests for a playlist that is currently being played, such as "more upbeat songs" or "more calming songs." This request is sent to the server via the device. The server reevaluates the playlist based on the received request and adds new songs. In this way, the playlist can respond to the user's detailed requests. The updated playlist is then sent back to the device, where the user can listen to it.
[0209] Specific examples
[0210] Consider a scenario where a user goes to the gym at 7am and then has coffee at a cafe at 9am.
[0211] 1. The device records these activities as location and behavioral data and sends them to a server.
[0212] 2. The device uses an emotion engine to recognize the user's facial expressions, voice, and heart rate, and estimates that the user is "refreshed and active." This emotion data is also sent to the server.
[0213] 3. The server aggregates this data and assumes that the user is "refreshed and active."
[0214] 4. The server generates a playlist containing up-tempo pop songs that fit this state of mind and sends it to the device.
[0215] 5. The user starts listening to the playlist and requests a lighter song.
[0216] 6. The server receives this request, updates the playlist to include more upbeat songs, and sends it back to the device.
[0217] 7. The user can enjoy the updated playlist.
[0218] In this way, by utilizing comprehensive data, including the user's emotional state, it is possible to provide a more advanced and personalized music experience.
[0219] The processing flow will be explained below.
[0220] Step 1:
[0221] The device uses the GPS function to obtain geographical data to obtain the user's current location information, and then sends the obtained location information to the server.
[0222] Step 2:
[0223] The device collects data about the user's daily activities, such as payment history, calendar schedule, and fitness tracker data, and transmits this data to a server.
[0224] Step 3:
[0225] The device collects the user's music preference data, specifically the history of songs played in the past and the user's ratings of songs, and transmits this data to the server.
[0226] Step 4:
[0227] The device uses an emotion engine to recognize the user's emotions. The emotion engine analyzes the user's voice, facial expressions, and biometric information (heart rate, body temperature, etc.) obtained from the input device. The resulting emotion data is sent to the server.
[0228] Step 5:
[0229] The server receives location information, behavioral data, music preference data, and emotional data and integrates them to create a user profile that comprehensively describes the user's current situation.
[0230] Step 6:
[0231] The server estimates the user's current state of mind based on the user profile. For example, if the user works out at the gym in the morning and then spends time at a cafe, the server estimates that the user feels refreshed and active.
[0232] Step 7:
[0233] The server selects the most suitable songs based on the estimated psychological state and music preference data, and generates a playlist based on the selected songs. This playlist includes song titles, artist names, playback links, etc.
[0234] Step 8:
[0235] The server transmits the generated playlist to the terminal, and the playlist data is stored on the user's terminal.
[0236] Step 9:
[0237] The user sends a request to the playlist being played, such as "more upbeat songs" or "more calm songs." This request is sent from the terminal to the server.
[0238] Step 10:
[0239] The server re-evaluates the playlist based on the user's request, adds new songs that meet the criteria, and sends the updated playlist back to the device.
[0240] Step 11:
[0241] The user receives the updated playlist and begins playback, allowing them to enjoy music that better suits their mood.
[0242] Through this series of processing steps, a highly personalized music experience can be provided based on the user's current emotions and behavior.
[0243] Example 2
[0244] Next, a description will be given of Example 2. In the following description, the data processing device 12 will be referred to as a "server" and the smart device 14 will be referred to as a "terminal."
[0245] Conventional music recommendation systems often generate playlists based solely on a user's music preference data, without taking into account the user's real-time emotions or behavioral status. As a result, they are unable to suggest songs that fit the user's current psychological state, making it difficult to increase user satisfaction. To solve this problem, a music recommendation system that also takes into account the user's real-time emotions and behavioral data is needed.
[0246] The specific processing by the specific processing unit 290 of the data processing device 12 in the second embodiment is realized by the following means.
[0247] In this invention, the server includes: means for collecting current location information from the user's device; means for collecting behavioral data from the user's device; means for collecting music preference data from the user's device; means having an emotion analysis engine for collecting emotion data from the device; means for estimating the user's current psychological state by integrating the collected location information, behavioral data, music preference data, and emotion data; means for generating an optimal music playlist based on the estimated psychological state and music preference data; means for transmitting the generated playlist to the user's device; means for receiving a request from the user, reevaluating and updating the playlist, and means for transmitting the updated playlist to the user's device. This enables personalized music recommendations that take into account the user's real-time emotions and behavioral status.
[0248] "User's terminal" refers to electronic devices such as mobile terminals and wearable devices used by users.
[0249] "Location information" refers to data regarding the user's current location determined using GPS functions, etc.
[0250] "Behavioral data" refers to data that shows a user's daily activities, including payment history, calendar schedule information, and data from fitness trackers.
[0251] "Music preference data" refers to data that indicates a user's musical preferences, such as a history of songs that the user has played in the past and ratings of songs.
[0252] "Emotion data" refers to data obtained by an emotion analysis engine from a user's facial expressions, voice, and biometric information (e.g., heart rate and body temperature).
[0253] An "emotion analysis engine" refers to a software or hardware system for analyzing a user's emotional state.
[0254] "Mood" refers to information that indicates a user's current emotional or mental state.
[0255] A "playlist" refers to a collection of songs selected based on specific criteria.
[0256] A "request" refers to information indicating a specific instruction or request from a user.
[0257] A "server" refers to a computer system that communicates with user terminals via a network and processes and stores data.
[0258] The present invention relates to a music recommendation system that takes into account real-time user emotion and behavior data. Specific embodiments of the system are described below.
[0259] To implement this system, a user terminal, a server, an emotion analysis engine, and a communication system that links these are required. The following hardware and software are used as the main components.
[0260] 1. User devices: Mobile devices such as smartphones and wearable devices
[0261] 2. Server: High-performance computer system
[0262] 3. Sentiment analysis engine: Speech recognition software (e.g., speech recognition API, face recognition API)
[0263] Data collection
[0264] The device first collects the user's current location information using its GPS function. This data is periodically recorded and sent to a server. It also uses a fitness tracker to obtain exercise data (e.g., heart rate, number of steps), and collects behavioral data from the device's calendar and payment history. Furthermore, it collects the user's music preference data, such as the history of songs played and their ratings, and sends this data to the server.
[0265] Emotion analysis
[0266] The device uses a built-in camera and microphone to transmit the user's facial expressions and voice to an emotion analysis engine. The emotion analysis engine analyzes this data and detects the user's emotional state (e.g., joy, sadness). For example, if the device's camera captures the user's smile and the microphone detects a joyful voice tone, this information is sent to the server as emotion data.
[0267] Data integration and psychological state estimation
[0268] The server receives location information, behavioral data, music preference data, and emotional data from the device. It integrates these data to create a user profile and uses machine learning algorithms to estimate the user's current psychological state. For example, if a user exercises at the gym in the morning and then has coffee at a cafe, the server estimates the user's psychological state as "refreshed and active."
[0269] Playlist Generation
[0270] The server generates an optimal music playlist based on the estimated psychological state and the user's music preference data. The playlist includes song titles, artist names, and play links. The server then sends the playlist to the user's device, allowing the user to play songs based on the playlist.
[0271] Handling the request
[0272] Users can make requests for songs such as "more upbeat" or "more calming" to a playlist that is currently playing. This request is sent to the server via the device, and the server reevaluates the playlist based on the request and adds new songs. The updated playlist is then sent back to the device, where the user can enjoy it.
[0273] Specific examples
[0274] Consider a scenario where a user goes to the gym at 7am and then has coffee at a cafe at 9am. In this scenario, the following happens:
[0275] 1. The device activates the GPS function and records that the user is at the gym.
[0276] 2. Your device receives your heart rate and exercise data from your fitness tracker and records the information.
[0277] 3. The terminal obtains the payment history at the cafe and records that information.
[0278] 4. The device uses its built-in camera and microphone to transmit the user's facial expressions of delight and tone of voice to the emotion analysis engine.
[0279] 5. The server receives and integrates the location information, behavioral data, music preference data, and emotional data.
[0280] 6. The server uses a machine learning algorithm to estimate the user's psychological state as "refreshed and active."
[0281] 7. The server generates a playlist containing up-tempo pop songs based on the estimated psychological state and sends it to the device.
[0282] 8. While listening to a playlist, the user requests "something more upbeat."
[0283] 9. The server receives the request, updates the playlist to include more upbeat songs, and sends it back to the device.
[0284] 10. Users can enjoy the updated playlist.
[0285] Example prompt for a generative AI model:
[0286] "You have behavioral data that shows that a user works out at the gym at 7 AM and drinks coffee at a cafe at 9 AM. Furthermore, analysis by a sentiment analysis engine suggests that the user is in a psychological state of 'refreshed and active.' Based on this data, please generate a playlist to recommend to the user."
[0287] In this way, it is possible to provide an advanced music recommendation system that also takes into account the user's real-time emotions and behavior.
[0288] The flow of the identification process in the second embodiment will be described with reference to FIG.
[0289] Step 1: Data collection phase
[0290] The device collects the user's current location information. Input includes location information obtained from the device's GPS function. This data is recorded with a timestamp and sent to the server. For example, if the user is at the gym, the location coordinates and time are recorded. The device also collects exercise data from a fitness tracker. Input includes heart rate and step count data. This data is also recorded and sent to the server. The device also collects the user's daily behavior data (payment history, calendar schedule information) and music preference data (past playback history and ratings) and sends them to the server. Input is obtained from the smartphone, payment system, calendar app, and music app.
[0291] Step 2: Emotion Recognition Phase
[0292] The device uses a built-in camera and microphone to send the user's facial expressions and voice to an emotion analysis engine. Inputs include facial expression data captured by the camera, voice data picked up by the microphone, and biometric data (heart rate, body temperature, etc.). The emotion analysis engine analyzes this data and determines the user's emotional state. Emotional data such as "happiness" or "sadness" is generated as output and sent to the server. Specifically, the device analyzes the user's smile, which indicates joy, and the tone of their voice, which indicates happiness, and records this as emotion data.
[0293] Step 3: Data integration phase
[0294] The server receives location information, behavioral data, music preference data, and emotional data sent from the device. These multiple data sets serve as input. The server integrates them and performs the necessary processing to create a user profile. This profile is then processed using machine learning algorithms to generate an output that analyzes the interrelationships between each piece of data. For example, if a user works out at the gym in the morning and then has coffee at a cafe, this information can be used to estimate a psychological state of "refreshed and active."
[0295] Step 4: Mental state estimation phase
[0296] The server runs a machine learning algorithm on the integrated data to estimate the user's current state of mind. The input is the user profile generated in the data integration phase. The algorithm analyzes this data and estimates the user's state of mind (e.g., "refreshed and active") as output. Specifically, the server analyzes the correlations between each data set based on time of day, behavioral patterns, and emotional changes.
[0297] Step 5: Playlist generation phase
[0298] The server selects optimal songs based on the estimated psychological state and music preference data. The inputs are the estimated psychological state data and music preference data. Based on these data, the server filters the optimal songs from the music database and generates a playlist. The output is a playlist containing song titles, artist names, play links, etc. Specifically, up-tempo pop songs are selected for users in a "refreshed and active" psychological state. This playlist is then sent from the server to the device.
[0299] Step 6: Request handling phase
[0300] A user can make requests to a playlist that is currently being played. The input is a request sent by the user through the terminal (e.g., "More upbeat songs," "More calm songs"). The server receives this request and re-evaluates the playlist based on the input request. The server adds newly selected songs to the playlist, and an updated playlist is generated as output. Specifically, the server adds songs from its song database that fit the request, and sends the updated playlist back to the terminal.
[0301] (Application example 2)
[0302] Next, a description will be given of Application Example 2. In the following description, the data processing device 12 will be referred to as a "server" and the smart device 14 will be referred to as a "terminal."
[0303] Conventional music recommendation systems are limited to generating simple playlists based on a user's musical preferences, and have difficulty providing a personalized music experience that takes into account the user's emotions and psychological state. This has resulted in problems such as not being able to provide music that matches the user's current mood or emotion, leading to a decrease in satisfaction. Furthermore, the system lacks the ability to flexibly update playlists in response to user requests. A system that can solve these issues is needed.
[0304] The specific processing by the specific processing unit 290 of the data processing device 12 in the application example 2 is realized by the following means.
[0305] In this invention, the server includes means for recognizing a user's emotions by performing voice analysis and facial expression analysis, means for transmitting data based on the user's emotions to the server and estimating the user's mood, and means for generating an optimal music playlist based on the estimated mood, thereby enabling a more personalized music experience according to the user's emotions and psychological state.
[0306] A "user's terminal" is an electronic device that a user uses on a daily basis, and includes smartphones, tablets, smartwatches, etc.
[0307] "Current location information" refers to geographical location data collected from the user's device and is obtained using GPS.
[0308] "Behavioral data" is information about a user's daily activities, including schedules, payment history, fitness tracker data, etc.
[0309] "Music preference data" is information such as the user's favorite music genres and artists, playback history, and ratings.
[0310] "Emotion" indicates the emotional state the user is feeling at that time, and is based on analyzed biometric information, voice, facial expressions, etc.
[0311] "Means for recognizing emotions" refers to technologies and algorithms that analyze a user's voice, facial expressions, and biometric information in order to identify the user's emotional state.
[0312] The "means for estimating mood" is a method for estimating a user's current psychological state using a computer algorithm based on collected user behavioral data, location information, and emotional data.
[0313] A "playlist" is a list of music for a specific purpose or theme, including song titles, artist names, and playback links.
[0314] The "means for receiving a request" is an interface through which the system receives requests or wishes from the user, such as voice input or touch input.
[0315] The "means for reevaluating and updating a playlist" refers to a means for reviewing a playlist that has already been generated and adding or changing new songs, etc., based on a request received from a user.
[0316] A "generative AI model" is an artificial intelligence model used to generate new music playlists based on collected data, and includes machine learning algorithms.
[0317] A "prompt sentence" is a text sentence that serves as input to a generative AI model to perform a specific task.
[0318] To implement this invention, a server, a terminal, an emotion engine, and a communication system that links them together are required. A specific embodiment of this system will be described below.
[0319] Data Collection Phase
[0320] The device collects the user's current location information. Using its GPS function, it periodically records the user's current location and sends it to the server. The device also collects the user's daily behavioral data. This behavioral data includes payment history, calendar schedule information, and fitness tracker data. The device also collects the user's music preference data, including the history of music played in the past and the ratings the user has given to songs. This data is also sent to the server.
[0321] Emotion Recognition Phase
[0322] The device recognizes the user's emotions through an emotion engine. The emotion engine analyzes the user's voice, facial expressions, and biometric information (heart rate, body temperature, etc.) obtained from input devices. The emotion data recognized by the emotion engine is sent to the server along with other collected data.
[0323] Mood estimation phase
[0324] The server receives location information, behavioral data, music preference data, and emotional data from the device and integrates them to create a user profile. Based on this profile, the server runs an algorithm to estimate the user's current psychological state. For example, if a user exercises at the gym in the morning and then has coffee at a cafe, these actions and emotions can be used to estimate the user's psychological state as "refreshed and active."
[0325] Playlist Generation Phase
[0326] The server selects the most suitable songs based on the estimated psychological state and music preference data. The generated playlist includes song titles, artist names, and playback links. This playlist is then sent to the device and provided to the user.
[0327] Request handling phase
[0328] Users can make requests for songs, such as "more upbeat" or "more calming," to a playlist that is currently playing. These requests are sent to the server via the device. The server reevaluates the playlist based on the received request and adds new songs. The updated playlist is then sent back to the device, where the user can listen to it.
[0329] Specific examples
[0330] Consider a scenario where a user goes to the gym at 7am and then has coffee at a cafe at 9am.
[0331] 1. The device records these activities as location and behavioral data and sends them to a server.
[0332] 2. The device uses an emotion engine to recognize the user's facial expressions, voice, and heart rate, and estimates that the user is "refreshed and active." This emotion data is also sent to the server.
[0333] 3. The server aggregates this data and assumes that the user is "refreshed and active."
[0334] 4. The server generates a playlist containing up-tempo pop songs that fit this state of mind and sends it to the device.
[0335] 5. The user starts listening to the playlist and requests a lighter song.
[0336] 6. The server receives this request, updates the playlist to include more upbeat songs, and sends it back to the device.
[0337] 7. The user can enjoy the updated playlist.
[0338] Prompt sentences to input to the generative AI model
[0339] Use the following prompt:
[0340] Based on the user's daily behavior data, such as working out at the gym at 7am and then drinking coffee at a cafe at 9am, estimate his current psychological state as "refreshed and active." His musical preferences are pop and electronic. Generate a playlist that is appropriate for this state.
[0341] This allows for a more personalized music experience based on the user's emotions and behavior.
[0342] The flow of the specific processing in the application example 2 will be described with reference to FIG.
[0343] Step 1:
[0344] The device collects the user's current location information. Using the GPS function, it periodically obtains the user's latitude and longitude and sends that location information to the server. The input is the device's GPS data, and the output is the location information data sent to the server.
[0345] Step 2:
[0346] The device collects the user's daily behavioral data, including payment history, calendar schedule information, and fitness tracker data, and sends it to a server. The input is this behavioral data, and the output is the integrated data sent to the server.
[0347] Step 3:
[0348] The device collects the user's music preference data, including the history of songs played and rating information, and sends it to the server. The input is the music playback history and rating information, and the output is the music preference data sent to the server.
[0349] Step 4:
[0350] The device recognizes the user's emotions through an emotion engine. The emotion engine analyzes the user's voice, facial expression, and biometric information (heart rate, body temperature, etc.) to recognize the user's emotional state and output it as data. This data is sent to the server. The input at this time is voice, facial expression, biometric information, etc., and the output is recognized emotional data.
[0351] Step 5:
[0352] The server receives location information, behavioral data, music preference data, and emotional data sent from the device and integrates them to create a user profile. The server's algorithm analyzes this data and estimates the user's current psychological state. The input is the integrated user data, and the output is an estimated psychological state.
[0353] Step 6:
[0354] The server generates an optimal music playlist based on the estimated psychological state and music preference data. Using a generative AI model, the server selects songs that best fit the user's psychological state and creates a playlist. The input is the psychological state and music preference data, and the output is the generated playlist.
[0355] Step 7:
[0356] The server sends the generated playlist to the user's device and provides it to the user. The playlist includes song titles, artist names, playback links, etc. The input is the generated playlist, and the output is the playlist sent to the device.
[0357] Step 8:
[0358] A user can make a request for a playlist that is currently being played. For example, a request for "more upbeat songs" is sent to the server via the terminal. The input is the user's request, and the output is the request data sent to the server.
[0359] Step 9:
[0360] The server reevaluates the playlist based on the received request and adds new songs. It uses a generative AI model to generate a new playlist and sends it to the device. The input is the user request and existing playlist data, and the output is the updated playlist.
[0361] Step 10:
[0362] The terminal plays the updated playlist sent from the server and provides it to the user, allowing the user to enjoy a music experience tailored to their requests. The input is the updated playlist data, and the output is the songs to be played.
[0363] The specific processing unit 290 transmits the result of the specific processing to the smart device 14. In the smart device 14, the control unit 46A causes the output device 40 to output the result of the specific processing. The microphone 38B acquires audio indicating a user input regarding the result of the specific processing. The control unit 46A transmits audio data indicating the user input acquired by the microphone 38B to the data processing device 12. In the data processing device 12, the specific processing unit 290 acquires the audio data.
[0364] The data generation model 58 is a so-called generative AI (Artificial Intelligence). An example of the data generation model 58 is ChatGPT (Internet Search<URL: https: / / openai.com / blog / chatgpt> ), Gemini (Internet search <url: https: gemini.google.com ?hl="ja">) and other generation AIs. The data generation model 58 is obtained by performing deep learning on a neural network. A prompt including an instruction is input to the data generation model 58, and inference data such as voice data indicating voice, text data indicating text, and image data indicating an image is also input. The data generation model 58 performs inference on the input inference data in accordance with the instruction indicated by the prompt, and outputs the inference result in a data format such as voice data and text data. Here, inference refers to, for example, analysis, classification, prediction, and / or summarization.
[0365] In the above embodiment, an example in which the specific process is performed by the data processing device 12 has been given, but the technology of the present disclosure is not limited to this, and the specific process may be performed by the smart device 14.
[0366] [Second embodiment]
[0367] FIG. 3 shows an example of the configuration of a data processing system 210 according to the second embodiment.
[0368] 3, the data processing system 210 includes the data processing device 12 and smart glasses 214. An example of the data processing device 12 is a server.
[0369] The data processing device 12 includes a computer 22, a database 24, and a communication I / F 26. The computer 22 is an example of a "computer" according to the technology of the present disclosure. The computer 22 includes a processor 28, a RAM 30, and a storage 32. The processor 28, the RAM 30, and the storage 32 are connected to a bus 34. The database 24 and the communication I / F 26 are also connected to the bus 34. The communication I / F 26 is connected to a network 54. Examples of the network 54 include a WAN (Wide Area Network) and / or a LAN (Local Area Network).
[0370] The smart glasses 214 include a computer 36, a microphone 238, a speaker 240, a camera 42, and a communication I / F 44. The computer 36 includes a processor 46, a RAM 48, and a storage 50. The processor 46, the RAM 48, and the storage 50 are connected to a bus 52. The microphone 238, the speaker 240, and the camera 42 are also connected to the bus 52.
[0371] The microphone 238 receives instructions and the like from the user 20 by receiving voice uttered by the user 20. The microphone 238 captures the voice uttered by the user 20, converts the captured voice into audio data, and outputs it to the processor 46. The speaker 240 outputs audio in accordance with instructions from the processor 46.
[0372] Camera 42 is a small digital camera equipped with an optical system including a lens, aperture, and shutter, and an imaging element such as a CMOS (Complementary Metal-Oxide-Semiconductor) image sensor or a CCD (Charge Coupled Device) image sensor, and captures images of the surroundings of user 20 (for example, an imaging range defined by an angle of view equivalent to the field of vision of a typical healthy person).
[0373] The communication I / F 44 is connected to a network 54. The communication I / Fs 44 and 26 are responsible for the exchange of various information between the processor 46 and the processor 28 via the network 54. The exchange of various information between the processor 46 and the processor 28 using the communication I / Fs 44 and 26 is carried out in a secure state.
[0374] Fig. 4 shows an example of the main functions of the data processing device 12 and the smart glasses 214. As shown in Fig. 4, in the data processing device 12, a specific process is performed by the processor 28. A specific process program 56 is stored in the storage 32.
[0375] The specific processing program 56 is an example of a "program" according to the technology of the present disclosure. The processor 28 reads the specific processing program 56 from the storage 32 and executes the read specific processing program 56 on the RAM 30. The specific processing is realized by the processor 28 operating as a specific processing unit 290 in accordance with the specific processing program 56 executed on the RAM 30.
[0376] The storage 32 stores a data generation model 58 and an emotion identification model 59. The data generation model 58 and the emotion identification model 59 are used by the identification processing unit 290.
[0377] In the smart glasses 214, the reception output process is performed by the processor 46. A reception output program 60 is stored in the storage 50. The processor 46 reads the reception output program 60 from the storage 50 and executes the read reception output program 60 on the RAM 48. The reception output process is realized by the processor 46 operating as the control unit 46A in accordance with the reception output program 60 executed on the RAM 48.
[0378] Next, a description will be given of the identification process performed by the identification processing unit 290 of the data processing device 12. In the following description, the data processing device 12 will be referred to as the "server" and the smart glasses 214 will be referred to as the "terminal."
[0379] To implement the present invention, a user terminal, a server, and a communication system that links them together are required. A specific embodiment of this system will be described below.
[0380] Data Collection Phase
[0381] The device collects the user's current location information. Using the GPS function, it periodically records the user's current location and sends it to the server. The device also collects the user's daily behavioral data. This behavioral data includes payment history, calendar schedule information, and fitness tracker data. The device also collects the user's music preference data. Specifically, this includes the history of music that has been played in the past and the ratings that the user has given to songs. This data is also sent to the server.
[0382] Mood estimation phase
[0383] The server receives location information, behavioral data, and music preference data sent from the device and integrates them to create a user profile. Based on this profile, the server runs an algorithm to estimate the user's current psychological state. For example, say you work out at the gym at 7:00 AM and then have coffee at a cafe at 9:00 AM. From these actions, the server estimates your psychological state as "refreshed and active."
[0384] Playlist Generation Phase
[0385] The server selects the most suitable songs based on the estimated psychological state and music preference data. It then generates a playlist using the selected songs. The generated playlist includes song titles, artist names, playback links, and other information. This playlist is then sent to the device.
[0386] Request handling phase
[0387] A user can make requests for a playlist that is currently being played, such as "more upbeat songs" or "more calming songs." This request is sent to the server via the device. The server reevaluates the playlist based on the received request and adds new songs. In this way, the playlist can respond to the user's detailed requests. The updated playlist is then sent back to the device, where the user can listen to it.
[0388] Specific examples
[0389] Consider a scenario where a user goes to the gym at 7am and then has coffee at a cafe at 9am.
[0390] 1. The device records these activities as location and behavioral data and sends them to a server.
[0391] 2. The server receives this data and assumes the user is "refreshed and active."
[0392] 3. The server generates a playlist containing up-tempo pop songs that fit this state of mind and sends it to the device.
[0393] 4. The user starts listening to the playlist and requests a lighter song.
[0394] 5. The server receives this request, updates the playlist to include more upbeat songs, and sends it back to the device.
[0395] 6. The user can enjoy the updated playlist.
[0396] In this way, it is possible to provide music that best suits the user's mood.
[0397] The processing flow will be explained below.
[0398] Step 1:
[0399] The device uses the GPS function to obtain geographical data to obtain the user's current location information, and then sends the obtained location information to the server.
[0400] Step 2:
[0401] The device collects data about the user's daily activities, such as payment history, calendar schedule, and fitness tracker data, and transmits this data to a server.
[0402] Step 3:
[0403] The device collects the user's music preference data, specifically the history of songs played in the past and the user's ratings of songs, and transmits this data to the server.
[0404] Step 4:
[0405] The server combines location, behavioral, and music preference data to create a user profile that comprehensively describes the user's current situation.
[0406] Step 5:
[0407] The server estimates the user's current state of mind based on the user profile. For example, if the user went to the gym in the morning and then spent time at a cafe, the server estimates that the user feels refreshed and active.
[0408] Step 6:
[0409] The server selects the most suitable songs based on the estimated psychological state and music preference data, and generates a playlist based on the selected songs.
[0410] Step 7:
[0411] The server sends the generated playlist to the device, which includes detailed information about the selected songs (song title, artist name, playback link, etc.).
[0412] Step 8:
[0413] Users can make requests for songs on a playlist, such as "more upbeat songs" or "more calm songs," and these requests are sent to the server via the terminal.
[0414] Step 9:
[0415] The server receives the request from the user, re-evaluates the playlist according to the request, reselects songs based on the new criteria, and updates the playlist.
[0416] Step 10:
[0417] The server then sends the updated playlist back to the terminal, where the user can play the updated playlist and enjoy the new songs.
[0418] Through the above processing steps, a musical experience that matches the user's current mental state and specific requests can be provided.
[0419] Example 1
[0420] Next, a description will be given of Example 1. In the following description, the data processing device 12 will be referred to as a "server" and the smart glasses 214 will be referred to as a "terminal."
[0421] In modern society, users are surrounded by vast amounts of information and data. Providing music that matches a user's psychological state and mood is crucial for improving individual satisfaction. However, conventional music recommendation systems struggle to comprehensively analyze a user's current location, behavioral data, and musical preference data to provide optimal music in real time. They also lack the ability to flexibly update playlists in response to user requests. Given these circumstances, there is a need for a system that can provide music that matches a user's current psychological state and quickly respond to user requests.
[0422] The specific processing by the specific processing unit 290 of the data processing device 12 in the first embodiment is realized by the following means.
[0423] In this invention, the server
[0424] means for collecting current location information from a user's information processing device;
[0425] means for collecting behavioral data from a user's information processing device;
[0426] means for collecting music preference data from a user's information processing device;
[0427] This allows the system to integrate collected data to estimate the user's current psychological state using a generative AI model, generate an optimal music playlist based on the estimation results, and reevaluate and update the playlist according to user requests.
[0428] An "information processing device" is a device that collects, processes, stores, and communicates data, and examples of such devices include smartphones, tablets, and wearable devices.
[0429] "Current location information" is data indicating the geographic location of the user, and is information obtained using GPS or other location detection technology.
[0430] "Behavioral data" is information about a user's daily activities, including payment history, calendar schedule information, fitness tracker data, and the like.
[0431] "Music preference data" is information relating to the user's tastes and ratings of music, and includes the history of songs played in the past and rating data.
[0432] A "generative AI model" is a model of an algorithm or program that uses artificial intelligence technology to analyze and generate data, and is used to estimate a user's psychological state.
[0433] "Mental state" is an indicator of the user's current mental and emotional state, such as feeling refreshed, active, or relaxed.
[0434] A "playlist" is a list of songs to be played in a specific order, and is generated based on the user's psychological state and musical preference data.
[0435] A "request" refers to a request or wish made by a user regarding a playlist that is being played, and specifically includes requests such as "a more upbeat song" or "a more calming song."
[0436] The present invention relates to a system that provides an optimal music playlist based on the user's psychological state and also responds to user requests. Implementing this system requires a user's information processing device, a server, and a communication system that links them.
[0437] Data Collection Phase
[0438] The device collects the user's current location information. This location information is obtained using GPS and updated periodically. Examples of devices like smartphones and wearable devices. The device also collects the user's daily behavioral data. This behavioral data includes payment history (collected via NFC or QR code readers), calendar schedule information (obtained from the Google Calendar API or Apple Calendar API), and fitness tracker data (obtained from the Apple HealthKit or Google Fit API). The device also collects the user's music preference data. Specifically, it collects playback history (using the Spotify API or Apple Music API) and song rating data. The collected data is sent from the device to a server.
[0439] Data analysis phase
[0440] The server receives location information, behavioral data, and music preference data sent from the device and integrates them to create a user profile. The profile reflects the user's behavioral patterns and music preferences. Based on this profile, the server uses a generative AI model to estimate the user's current psychological state. For example, from data such as "The user worked out at the gym at 7am and then had coffee at a cafe at 9am," the server generates a prompt that estimates the user's psychological state as "Refreshed and Active." The following prompt is a concrete example:
[0441] "A user works out at the gym at 7am and then has coffee at a cafe at 9am. Generate a playlist with appropriate uptempo pop songs to infer that the user is refreshed and active."
[0442] Playlist Generation Phase
[0443] The server selects the most suitable songs based on the estimated psychological state and music preference data. Songs are searched from a database to select those that suit the user's psychological state. For example, for a "refreshed and active" psychological state, up-tempo pop songs are selected. A playlist is generated using the selected songs, including information such as song title, artist name, and playback link. The generated playlist is sent from the server to the device.
[0444] Request handling phase
[0445] Users can make requests for songs, such as "more upbeat songs" or "more calming songs," to the playlist being played. These requests are sent to the server via the device. The server reevaluates the playlist based on the received request and adds new songs. The updated playlist is then sent back to the device, allowing the user to enjoy the new playlist.
[0446] In this way, the system not only provides music that is optimal for the user's psychological state, but also has the feature of being able to quickly respond to user requests. The above is an embodiment of the present invention.
[0447] The flow of the identification process in the first embodiment will be described with reference to FIG.
[0448] Step 1: Start collecting data
[0449] The device starts collecting data when the user starts using the system. The input is the user's action, and the output is the start of the data collection process. For example, this process can be started by pressing a "start data collection" button on the app or device.
[0450] Step 2: Obtaining location information
[0451] The device uses the GPS function to obtain the user's current location information. The input is location information from the GPS sensor, and the output is to temporarily store this information within the device. Specifically, the device obtains location information once every minute and records it as latitude and longitude.
[0452] Step 3: Behavioral data collection
[0453] The device collects user behavioral data. The input is data from various sensors and APIs, and the output is collected behavioral data. For example, payment history is obtained via an NFC sensor or QR code reader, and calendar schedule information is obtained from the Google Calendar API or Apple Calendar API. Fitness tracker data is also collected using Apple HealthKit or Google Fit API.
[0454] Step 4: Collecting Music Preference Data
[0455] The device collects the user's music preference data. The input is data from streaming service APIs (Spotify API and Apple Music API), and the output is the collection of music preference data. Specifically, it collects information on playback history and user actions such as "like" and "skip."
[0456] Step 5: Send data
[0457] The device sends the collected location information, behavioral data, and music preference data to a server. The input is various data stored on the device, and the output is sending this data to the server. This communication is carried out via the Internet.
[0458] Step 6: Create a user profile
[0459] The server creates a user profile based on the received data. The input is the data received by the server, and the output is the creation of a user profile. Specifically, the server organizes the data into categories and uses statistical methods and machine learning models to analyze the user's behavioral patterns and musical preferences.
[0460] Step 7: Estimate mental state
[0461] The server uses a generative AI model to estimate the user's psychological state. The input is the user profile and behavioral data collected at that time, and the output is the estimated psychological state. Specifically, a prompt sentence is input into the generative AI model, and the psychological state is estimated from the response. For example, it may estimate that "the user worked out at the gym at 7 a.m. and had coffee at a cafe at 9 a.m., so they are refreshed and active."
[0462] Step 8: Selecting songs and creating a playlist
[0463] The server selects the most suitable songs based on the estimated psychological state and music preference data. The input is the psychological state and music preference data, and the output is a generated playlist. Specifically, it searches for suitable songs from a database and creates a playlist. For example, it selects an up-tempo pop song that is suitable for the "active" psychological state.
[0464] Step 9: Send Playlist
[0465] The server sends the generated playlist to the terminal. The input is the generated playlist, and the output is the playlist arriving at the terminal. This communication is performed via the Internet, and real-time updates are possible.
[0466] Step 10: Respond to requests
[0467] A user can make requests to a playlist that is currently playing, such as "more upbeat songs" or "more calming songs." The input is the user's request, and the output is a new playlist that has been re-evaluated based on that request. Specifically, the server receives the request, re-evaluates the playlist, and adds new songs. This updated playlist is then sent back to the device, where the user can listen to it.
[0468] The above are the processing steps of the program of this system and its specific operations.
[0469] (Application example 1)
[0470] Next, a description will be given of Application Example 1. In the following description, the data processing device 12 will be referred to as a "server" and the smart glasses 214 will be referred to as a "terminal."
[0471] In recent years, the spread of self-driving vehicles has created a demand for improved in-car entertainment experiences. However, conventional in-car entertainment systems have faced the challenge of providing music and content that is tailored to the user's current psychological state and mood. There is a demand for improving in-car comfort and satisfaction by providing music that matches the user's mood.
[0472] The specific processing by the specific processing unit 290 of the data processing device 12 in the application example 1 is realized by the following means.
[0473] In this invention, the server includes means for collecting current location information from the user's terminal, means for collecting behavioral data from the user's terminal, and means for collecting music preference data from the user's terminal, thereby making it possible to provide a music playlist that is optimal for the user's mood in the autonomous vehicle.
[0474] A "user's terminal" is a mobile information terminal such as a smartphone or tablet owned by a user.
[0475] "Current location information" means a user's real-time geographic location data obtained using location measurement technology such as GPS.
[0476] "Behavioral data" refers to information about a user's behavior, such as fitness tracker data or schedule information.
[0477] "Music preference data" is data relating to the user's musical preferences, such as the history of songs that the user has played in the past and their ratings of those songs.
[0478] "Mental state" is information that indicates the user's current mood and emotional state.
[0479] A "playlist" is a list of multiple songs provided to a user.
[0480] A "request" is a request or instruction given by a user to a playlist.
[0481] An "autonomous vehicle" is a vehicle that operates autonomously and requires minimal user intervention.
[0482] An "entertainment system" is a system that provides entertainment content such as music and video in a vehicle.
[0483] To implement the present invention, a user terminal, a server, and a communication system that links them together are required. A specific embodiment of this system will be described below.
[0484] 1. Data Collection Phase
[0485] The device collects the user's current location information. Using the GPS function, it periodically records the user's current location and sends it to the server. The device also collects the user's daily behavioral data. This behavioral data includes payment history, calendar schedule information, and fitness tracker data. The device also collects the user's music preference data. Specifically, this includes the history of music that has been played in the past and the ratings that the user has given to songs. This data is also sent to the server.
[0486] 2. Mood estimation phase
[0487] The server receives location information, behavioral data, and music preference data sent from the device and integrates them to create a user profile. Based on this profile, the server runs an algorithm to estimate the user's current psychological state. For example, say you work out at the gym at 7:00 AM and then have coffee at a cafe at 9:00 AM. From these actions, the server estimates your psychological state as "refreshed and active."
[0488] 3. Playlist generation phase
[0489] The server selects the most suitable songs based on the estimated psychological state and music preference data. It then generates a playlist using the selected songs. The generated playlist includes song titles, artist names, playback links, and other information. This playlist is then sent to the device.
[0490] 4. Request Response Phase
[0491] A user can make requests for a playlist that is currently being played, such as "more upbeat songs" or "more calming songs." This request is sent to the server via the device. The server reevaluates the playlist based on the received request and adds new songs. In this way, the playlist can respond to the user's detailed requests. The updated playlist is then sent back to the device, where the user can listen to it.
[0492] Specific examples
[0493] Consider a scenario where a user goes to the gym at 7am and then has coffee at a cafe at 9am.
[0494] 1. The device records these activities as location and behavioral data and sends them to a server.
[0495] 2. The server receives this data and assumes the user is "refreshed and active."
[0496] 3. The server generates a playlist containing up-tempo music that matches this psychological state (e.g., "Uptown Funk") and sends it to the terminal.
[0497] 4. The user starts listening to the playlist and requests a lighter song.
[0498] 5. The server receives this request, updates the playlist to include more upbeat songs, and sends it back to the device.
[0499] 6. The user can enjoy the updated playlist.
[0500] Example of input prompt for generative AI model
[0501] "Generate a music playlist that best suits the user's mood based on the following information:
[0502] Current location: Coordinates (35.6895, 139.6917)
[0503] Behavioral data: 1 hour training at 6am, commuting in the car at 9am
[0504] Music preference data: List of songs you've played in the past (e.g. Happy - Williams, Uptown Funk - Mars)
[0505] Choose the best songs to match the user's mood and provide a playlist."
[0506] The flow of the specific processing in the application example 1 will be described with reference to FIG.
[0507] Step 1:
[0508] The device collects the user's current location information. Specifically, it periodically obtains the user's geographical location data using the GPS function. In addition to this, it also collects the user's daily behavior data (e.g., fitness tracker data and calendar schedule information). It also collects the user's music preference data (such as the history and ratings of music played in the past). This data is then sent to the server.
[0509] Input: User location information, behavioral data, music preference data
[0510] Output: Consolidated data packet sent to the server
[0511] Step 2:
[0512] The server receives location information, behavioral data, and music preference data sent from the device and integrates them to create a user profile. Specifically, it uses a data analysis algorithm to integrate this diverse information and generate a user profile.
[0513] Input: location information, behavioral data, music preference data
[0514] Output: Unified user profile
[0515] Step 3:
[0516] The server runs an algorithm to estimate the user's current psychological state based on the generated user profile. Specifically, it uses a machine learning model to analyze the user's behavioral patterns and music preferences, and estimates the user's current psychological state (e.g., "refreshed and active").
[0517] Input: User profile
[0518] Output: Estimated mental state
[0519] Step 4:
[0520] The server selects the most suitable songs based on the estimated psychological state and music preference data. Specifically, it searches a music preference database to create a list of songs that suit the user's current mood. The server then generates a playlist using the selected songs.
[0521] Input: Psychological state, music preference data
[0522] Output: Generated playlist
[0523] Step 5:
[0524] The server then sends the generated playlist to the user's device. Specifically, the playlist contains song titles, artist names, playback links, etc.
[0525] Input: Generated playlist
[0526] Output: Playlist sent to user device
[0527] Step 6:
[0528] The user issues a request (e.g., "more upbeat songs," "more calm songs") to the playlist being played. This request is sent to the server via the device. The server reevaluates the playlist based on the received request and adds new songs.
[0529] Input: A request from the user
[0530] Output: Updated playlist
[0531] Step 7:
[0532] The server then sends the updated playlist back to the terminal, allowing the user to listen to the latest music playlist.
[0533] Input: Updated playlist
[0534] Output: The updated playlist sent to the user's device
[0535] (Example of input prompt for generative AI model)
[0536] "Generate a music playlist that best suits the user's mood based on the following information:
[0537] Current location: Coordinates (35.6895, 139.6917)
[0538] Behavioral data: 1 hour training at 6am, commuting in the car at 9am
[0539] Music preference data: List of songs you've played in the past (e.g. Happy - Williams, Uptown Funk - Mars)
[0540] Choose the best songs to match the user's mood and provide a playlist."
[0541] Furthermore, an emotion engine that estimates the user's emotion may be further combined. That is, the identification processing unit 290 may estimate the user's emotion using the emotion identification model 59, and perform identification processing using the user's emotion.
[0542] To implement the present invention, a user terminal, a server, an emotion engine, and a communication system that links these components are required. A specific embodiment of this system will be described below.
[0543] Data Collection Phase
[0544] The device collects the user's current location information. Using the GPS function, it periodically records the user's current location and sends it to the server. The device also collects the user's daily behavioral data. This behavioral data includes payment history, calendar schedule information, and fitness tracker data. The device also collects the user's music preference data. Specifically, this includes the history of music that has been played in the past and the ratings that the user has given to songs. This data is also sent to the server.
[0545] Emotion Recognition Phase
[0546] The device recognizes the user's emotions through an emotion engine. The emotion engine analyzes the user's voice, facial expressions, and biometric information (heart rate, body temperature, etc.) obtained from input devices. The emotion data recognized by the emotion engine is sent to the server along with other collected data.
[0547] Mood estimation phase
[0548] The server receives location information, behavioral data, music preference data, and emotional data from the device and integrates them to create a user profile. Based on this profile, the server runs an algorithm to estimate the user's current psychological state. For example, suppose a user exercises at the gym in the morning and then has coffee at a cafe. Based on these actions and emotions, the server estimates the user's psychological state as "refreshed and active."
[0549] Playlist Generation Phase
[0550] The server selects the most suitable songs based on the estimated psychological state and music preference data. It then generates a playlist using the selected songs. The generated playlist includes song titles, artist names, playback links, and other information. This playlist is then sent to the device.
[0551] Request handling phase
[0552] A user can make requests for a playlist that is currently being played, such as "more upbeat songs" or "more calming songs." This request is sent to the server via the device. The server reevaluates the playlist based on the received request and adds new songs. In this way, the playlist can respond to the user's detailed requests. The updated playlist is then sent back to the device, where the user can listen to it.
[0553] Specific examples
[0554] Consider a scenario where a user goes to the gym at 7am and then has coffee at a cafe at 9am.
[0555] 1. The device records these activities as location and behavioral data and sends them to a server.
[0556] 2. The device uses an emotion engine to recognize the user's facial expressions, voice, and heart rate, and estimates that the user is "refreshed and active." This emotion data is also sent to the server.
[0557] 3. The server aggregates this data and assumes that the user is "refreshed and active."
[0558] 4. The server generates a playlist containing up-tempo pop songs that fit this state of mind and sends it to the device.
[0559] 5. The user starts listening to the playlist and requests a lighter song.
[0560] 6. The server receives this request, updates the playlist to include more upbeat songs, and sends it back to the device.
[0561] 7. The user can enjoy the updated playlist.
[0562] In this way, by utilizing comprehensive data, including the user's emotional state, it is possible to provide a more advanced and personalized music experience.
[0563] The processing flow will be explained below.
[0564] Step 1:
[0565] The device uses the GPS function to obtain geographical data to obtain the user's current location information, and then sends the obtained location information to the server.
[0566] Step 2:
[0567] The device collects data about the user's daily activities, such as payment history, calendar schedule, and fitness tracker data, and transmits this data to a server.
[0568] Step 3:
[0569] The device collects the user's music preference data, specifically the history of songs played in the past and the user's ratings of songs, and transmits this data to the server.
[0570] Step 4:
[0571] The device uses an emotion engine to recognize the user's emotions. The emotion engine analyzes the user's voice, facial expressions, and biometric information (heart rate, body temperature, etc.) obtained from the input device. The resulting emotion data is sent to the server.
[0572] Step 5:
[0573] The server receives location information, behavioral data, music preference data, and emotional data and integrates them to create a user profile that comprehensively describes the user's current situation.
[0574] Step 6:
[0575] The server estimates the user's current state of mind based on the user profile. For example, if the user works out at the gym in the morning and then spends time at a cafe, the server estimates that the user feels refreshed and active.
[0576] Step 7:
[0577] The server selects the most suitable songs based on the estimated psychological state and music preference data, and generates a playlist based on the selected songs. This playlist includes song titles, artist names, playback links, etc.
[0578] Step 8:
[0579] The server transmits the generated playlist to the terminal, and the playlist data is stored on the user's terminal.
[0580] Step 9:
[0581] The user sends a request to the playlist being played, such as "more upbeat songs" or "more calm songs." This request is sent from the terminal to the server.
[0582] Step 10:
[0583] The server re-evaluates the playlist based on the user's request, adds new songs that meet the criteria, and sends the updated playlist back to the device.
[0584] Step 11:
[0585] The user receives the updated playlist and begins playback, allowing them to enjoy music that better suits their mood.
[0586] Through this series of processing steps, a highly personalized music experience can be provided based on the user's current emotions and behavior.
[0587] Example 2
[0588] Next, a description will be given of Example 2. In the following description, the data processing device 12 will be referred to as a "server" and the smart glasses 214 will be referred to as a "terminal."
[0589] Conventional music recommendation systems often generate playlists based solely on a user's music preference data, without taking into account the user's real-time emotions or behavioral status. As a result, they are unable to suggest songs that fit the user's current psychological state, making it difficult to increase user satisfaction. To solve this problem, a music recommendation system that also takes into account the user's real-time emotions and behavioral data is needed.
[0590] The specific processing by the specific processing unit 290 of the data processing device 12 in the second embodiment is realized by the following means.
[0591] In this invention, the server includes: means for collecting current location information from the user's device; means for collecting behavioral data from the user's device; means for collecting music preference data from the user's device; means having an emotion analysis engine for collecting emotion data from the device; means for estimating the user's current psychological state by integrating the collected location information, behavioral data, music preference data, and emotion data; means for generating an optimal music playlist based on the estimated psychological state and music preference data; means for transmitting the generated playlist to the user's device; means for receiving a request from the user, reevaluating and updating the playlist, and means for transmitting the updated playlist to the user's device. This enables personalized music recommendations that take into account the user's real-time emotions and behavioral status.
[0592] "User's terminal" refers to electronic devices such as mobile terminals and wearable devices used by users.
[0593] "Location information" refers to data regarding the user's current location determined using GPS functions, etc.
[0594] "Behavioral data" refers to data that shows a user's daily activities, including payment history, calendar schedule information, and data from fitness trackers.
[0595] "Music preference data" refers to data that indicates a user's musical preferences, such as a history of songs that the user has played in the past and ratings of songs.
[0596] "Emotion data" refers to data obtained by an emotion analysis engine from a user's facial expressions, voice, and biometric information (e.g., heart rate and body temperature).
[0597] An "emotion analysis engine" refers to a software or hardware system for analyzing a user's emotional state.
[0598] "Mood" refers to information that indicates a user's current emotional or mental state.
[0599] A "playlist" refers to a collection of songs selected based on specific criteria.
[0600] A "request" refers to information indicating a specific instruction or request from a user.
[0601] A "server" refers to a computer system that communicates with user terminals via a network and processes and stores data.
[0602] The present invention relates to a music recommendation system that takes into account real-time user emotion and behavior data. Specific embodiments of the system are described below.
[0603] To implement this system, a user terminal, a server, an emotion analysis engine, and a communication system that links these are required. The following hardware and software are used as the main components.
[0604] 1. User devices: Mobile devices such as smartphones and wearable devices
[0605] 2. Server: High-performance computer system
[0606] 3. Sentiment analysis engine: Speech recognition software (e.g., speech recognition API, face recognition API)
[0607] Data collection
[0608] The device first collects the user's current location information using its GPS function. This data is periodically recorded and sent to a server. It also uses a fitness tracker to obtain exercise data (e.g., heart rate, number of steps), and collects behavioral data from the device's calendar and payment history. Furthermore, it collects the user's music preference data, such as the history of songs played and their ratings, and sends this data to the server.
[0609] Emotion analysis
[0610] The device uses a built-in camera and microphone to transmit the user's facial expressions and voice to an emotion analysis engine. The emotion analysis engine analyzes this data and detects the user's emotional state (e.g., joy, sadness). For example, if the device's camera captures the user's smile and the microphone detects a joyful voice tone, this information is sent to the server as emotion data.
[0611] Data integration and psychological state estimation
[0612] The server receives location information, behavioral data, music preference data, and emotional data from the device. It integrates these data to create a user profile and uses machine learning algorithms to estimate the user's current psychological state. For example, if a user exercises at the gym in the morning and then has coffee at a cafe, the server estimates the user's psychological state as "refreshed and active."
[0613] Playlist Generation
[0614] The server generates an optimal music playlist based on the estimated psychological state and the user's music preference data. The playlist includes song titles, artist names, and play links. The server then sends the playlist to the user's device, allowing the user to play songs based on the playlist.
[0615] Handling the request
[0616] Users can make requests for songs such as "more upbeat" or "more calming" to a playlist that is currently playing. This request is sent to the server via the device, and the server reevaluates the playlist based on the request and adds new songs. The updated playlist is then sent back to the device, where the user can enjoy it.
[0617] Specific examples
[0618] Consider a scenario where a user goes to the gym at 7am and then has coffee at a cafe at 9am. In this scenario, the following happens:
[0619] 1. The device activates the GPS function and records that the user is at the gym.
[0620] 2. Your device receives your heart rate and exercise data from your fitness tracker and records the information.
[0621] 3. The terminal obtains the payment history at the cafe and records that information.
[0622] 4. The device uses its built-in camera and microphone to transmit the user's facial expressions of delight and tone of voice to the emotion analysis engine.
[0623] 5. The server receives and integrates the location information, behavioral data, music preference data, and emotional data.
[0624] 6. The server uses a machine learning algorithm to estimate the user's psychological state as "refreshed and active."
[0625] 7. The server generates a playlist containing up-tempo pop songs based on the estimated psychological state and sends it to the device.
[0626] 8. While listening to a playlist, the user requests "something more upbeat."
[0627] 9. The server receives the request, updates the playlist to include more upbeat songs, and sends it back to the device.
[0628] 10. Users can enjoy the updated playlist.
[0629] Example prompt for a generative AI model:
[0630] "You have behavioral data that shows that a user works out at the gym at 7 AM and drinks coffee at a cafe at 9 AM. Furthermore, analysis by a sentiment analysis engine suggests that the user is in a psychological state of 'refreshed and active.' Based on this data, please generate a playlist to recommend to the user."
[0631] In this way, it is possible to provide an advanced music recommendation system that also takes into account the user's real-time emotions and behavior.
[0632] The flow of the identification process in the second embodiment will be described with reference to FIG.
[0633] Step 1: Data collection phase
[0634] The device collects the user's current location information. Input includes location information obtained from the device's GPS function. This data is recorded with a timestamp and sent to the server. For example, if the user is at the gym, the location coordinates and time are recorded. The device also collects exercise data from a fitness tracker. Input includes heart rate and step count data. This data is also recorded and sent to the server. The device also collects the user's daily behavior data (payment history, calendar schedule information) and music preference data (past playback history and ratings) and sends them to the server. Input is obtained from the smartphone, payment system, calendar app, and music app.
[0635] Step 2: Emotion Recognition Phase
[0636] The device uses a built-in camera and microphone to send the user's facial expressions and voice to an emotion analysis engine. Inputs include facial expression data captured by the camera, voice data picked up by the microphone, and biometric data (heart rate, body temperature, etc.). The emotion analysis engine analyzes this data and determines the user's emotional state. Emotional data such as "happiness" or "sadness" is generated as output and sent to the server. Specifically, the device analyzes the user's smile, which indicates joy, and the tone of their voice, which indicates happiness, and records this as emotion data.
[0637] Step 3: Data integration phase
[0638] The server receives location information, behavioral data, music preference data, and emotional data sent from the device. These multiple data sets serve as input. The server integrates them and performs the necessary processing to create a user profile. This profile is then processed using machine learning algorithms to generate an output that analyzes the interrelationships between each piece of data. For example, if a user works out at the gym in the morning and then has coffee at a cafe, this information can be used to estimate a psychological state of "refreshed and active."
[0639] Step 4: Mental state estimation phase
[0640] The server runs a machine learning algorithm on the integrated data to estimate the user's current state of mind. The input is the user profile generated in the data integration phase. The algorithm analyzes this data and estimates the user's state of mind (e.g., "refreshed and active") as output. Specifically, the server analyzes the correlations between each data set based on time of day, behavioral patterns, and emotional changes.
[0641] Step 5: Playlist generation phase
[0642] The server selects optimal songs based on the estimated psychological state and music preference data. The inputs are the estimated psychological state data and music preference data. Based on these data, the server filters the optimal songs from the music database and generates a playlist. The output is a playlist containing song titles, artist names, play links, etc. Specifically, up-tempo pop songs are selected for users in a "refreshed and active" psychological state. This playlist is then sent from the server to the device.
[0643] Step 6: Request handling phase
[0644] A user can make requests to a playlist that is currently being played. The input is a request sent by the user through the terminal (e.g., "More upbeat songs," "More calm songs"). The server receives this request and re-evaluates the playlist based on the input request. The server adds newly selected songs to the playlist, and an updated playlist is generated as output. Specifically, the server adds songs from its song database that fit the request, and sends the updated playlist back to the terminal.
[0645] (Application example 2)
[0646] Next, a description will be given of Application Example 2. In the following description, the data processing device 12 will be referred to as a "server" and the smart glasses 214 will be referred to as a "terminal."
[0647] Conventional music recommendation systems are limited to generating simple playlists based on a user's musical preferences, and have difficulty providing a personalized music experience that takes into account the user's emotions and psychological state. This has resulted in problems such as not being able to provide music that matches the user's current mood or emotion, leading to a decrease in satisfaction. Furthermore, the system lacks the ability to flexibly update playlists in response to user requests. A system that can solve these issues is needed.
[0648] The specific processing by the specific processing unit 290 of the data processing device 12 in the application example 2 is realized by the following means.
[0649] In this invention, the server includes means for recognizing a user's emotions by performing voice analysis and facial expression analysis, means for transmitting data based on the user's emotions to the server and estimating the user's mood, and means for generating an optimal music playlist based on the estimated mood, thereby enabling a more personalized music experience according to the user's emotions and psychological state.
[0650] A "user's terminal" is an electronic device that a user uses on a daily basis, and includes smartphones, tablets, smartwatches, etc.
[0651] "Current location information" refers to geographical location data collected from the user's device and is obtained using GPS.
[0652] "Behavioral data" is information about a user's daily activities, including schedules, payment history, fitness tracker data, etc.
[0653] "Music preference data" is information such as the user's favorite music genres and artists, playback history, and ratings.
[0654] "Emotion" indicates the emotional state the user is feeling at that time, and is based on analyzed biometric information, voice, facial expressions, etc.
[0655] "Means for recognizing emotions" refers to technologies and algorithms that analyze a user's voice, facial expressions, and biometric information in order to identify the user's emotional state.
[0656] The "means for estimating mood" is a method for estimating a user's current psychological state using a computer algorithm based on collected user behavioral data, location information, and emotional data.
[0657] A "playlist" is a list of music for a specific purpose or theme, including song titles, artist names, and playback links.
[0658] The "means for receiving a request" is an interface through which the system receives requests or wishes from the user, such as voice input or touch input.
[0659] The "means for reevaluating and updating a playlist" refers to a means for reviewing a playlist that has already been generated and adding or changing new songs, etc., based on a request received from a user.
[0660] A "generative AI model" is an artificial intelligence model used to generate new music playlists based on collected data, and includes machine learning algorithms.
[0661] A "prompt sentence" is a text sentence that serves as input to a generative AI model to perform a specific task.
[0662] To implement this invention, a server, a terminal, an emotion engine, and a communication system that links them together are required. A specific embodiment of this system will be described below.
[0663] Data Collection Phase
[0664] The device collects the user's current location information. Using its GPS function, it periodically records the user's current location and sends it to the server. The device also collects the user's daily behavioral data. This behavioral data includes payment history, calendar schedule information, and fitness tracker data. The device also collects the user's music preference data, including the history of music played in the past and the ratings the user has given to songs. This data is also sent to the server.
[0665] Emotion Recognition Phase
[0666] The device recognizes the user's emotions through an emotion engine. The emotion engine analyzes the user's voice, facial expressions, and biometric information (heart rate, body temperature, etc.) obtained from input devices. The emotion data recognized by the emotion engine is sent to the server along with other collected data.
[0667] Mood estimation phase
[0668] The server receives location information, behavioral data, music preference data, and emotional data from the device and integrates them to create a user profile. Based on this profile, the server runs an algorithm to estimate the user's current psychological state. For example, if a user exercises at the gym in the morning and then has coffee at a cafe, these actions and emotions can be used to estimate the user's psychological state as "refreshed and active."
[0669] Playlist Generation Phase
[0670] The server selects the most suitable songs based on the estimated psychological state and music preference data. The generated playlist includes song titles, artist names, and playback links. This playlist is then sent to the device and provided to the user.
[0671] Request handling phase
[0672] Users can make requests for songs, such as "more upbeat" or "more calming," to a playlist that is currently playing. These requests are sent to the server via the device. The server reevaluates the playlist based on the received request and adds new songs. The updated playlist is then sent back to the device, where the user can listen to it.
[0673] Specific examples
[0674] Consider a scenario where a user goes to the gym at 7am and then has coffee at a cafe at 9am.
[0675] 1. The device records these activities as location and behavioral data and sends them to a server.
[0676] 2. The device uses an emotion engine to recognize the user's facial expressions, voice, and heart rate, and estimates that the user is "refreshed and active." This emotion data is also sent to the server.
[0677] 3. The server aggregates this data and assumes that the user is "refreshed and active."
[0678] 4. The server generates a playlist containing up-tempo pop songs that fit this state of mind and sends it to the device.
[0679] 5. The user starts listening to the playlist and requests a lighter song.
[0680] 6. The server receives this request, updates the playlist to include more upbeat songs, and sends it back to the device.
[0681] 7. The user can enjoy the updated playlist.
[0682] Prompt sentences to input to the generative AI model
[0683] Use the following prompt:
[0684] Based on the user's daily behavior data, such as working out at the gym at 7am and then drinking coffee at a cafe at 9am, estimate his current psychological state as "refreshed and active." His musical preferences are pop and electronic. Generate a playlist that is appropriate for this state.
[0685] This allows for a more personalized music experience based on the user's emotions and behavior.
[0686] The flow of the specific processing in the application example 2 will be described with reference to FIG.
[0687] Step 1:
[0688] The device collects the user's current location information. Using the GPS function, it periodically obtains the user's latitude and longitude and sends that location information to the server. The input is the device's GPS data, and the output is the location information data sent to the server.
[0689] Step 2:
[0690] The device collects the user's daily behavioral data, including payment history, calendar schedule information, and fitness tracker data, and sends it to a server. The input is this behavioral data, and the output is the integrated data sent to the server.
[0691] Step 3:
[0692] The device collects the user's music preference data, including the history of songs played and rating information, and sends it to the server. The input is the music playback history and rating information, and the output is the music preference data sent to the server.
[0693] Step 4:
[0694] The device recognizes the user's emotions through an emotion engine. The emotion engine analyzes the user's voice, facial expression, and biometric information (heart rate, body temperature, etc.) to recognize the user's emotional state and output it as data. This data is sent to the server. The input at this time is voice, facial expression, biometric information, etc., and the output is recognized emotional data.
[0695] Step 5:
[0696] The server receives location information, behavioral data, music preference data, and emotional data sent from the device and integrates them to create a user profile. The server's algorithm analyzes this data and estimates the user's current psychological state. The input is the integrated user data, and the output is an estimated psychological state.
[0697] Step 6:
[0698] The server generates an optimal music playlist based on the estimated psychological state and music preference data. Using a generative AI model, the server selects songs that best fit the user's psychological state and creates a playlist. The input is the psychological state and music preference data, and the output is the generated playlist.
[0699] Step 7:
[0700] The server sends the generated playlist to the user's device and provides it to the user. The playlist includes song titles, artist names, playback links, etc. The input is the generated playlist, and the output is the playlist sent to the device.
[0701] Step 8:
[0702] A user can make a request for a playlist that is currently being played. For example, a request for "more upbeat songs" is sent to the server via the terminal. The input is the user's request, and the output is the request data sent to the server.
[0703] Step 9:
[0704] The server reevaluates the playlist based on the received request and adds new songs. It uses a generative AI model to generate a new playlist and sends it to the device. The input is the user request and existing playlist data, and the output is the updated playlist.
[0705] Step 10:
[0706] The terminal plays the updated playlist sent from the server and provides it to the user, allowing the user to enjoy a music experience tailored to their requests. The input is the updated playlist data, and the output is the songs to be played.
[0707] The specific processing unit 290 transmits the result of the specific processing to the smart glasses 214. In the smart glasses 214, the control unit 46A causes the speaker 240 to output the result of the specific processing. The microphone 238 acquires audio indicating a user input regarding the result of the specific processing. The control unit 46A transmits audio data indicating the user input acquired by the microphone 238 to the data processing device 12. In the data processing device 12, the specific processing unit 290 acquires the audio data.
[0708] The data generation model 58 is a so-called generative AI (Artificial Intelligence). An example of the data generation model 58 is ChatGPT (Internet Search<URL: https: / / openai.com / blog / chatgpt> ), Gemini (Internet search <url: https: gemini.google.com ?hl="ja">) and other generation AIs. The data generation model 58 is obtained by performing deep learning on a neural network. A prompt including an instruction is input to the data generation model 58, and inference data such as voice data indicating voice, text data indicating text, and image data indicating an image is also input. The data generation model 58 performs inference on the input inference data in accordance with the instruction indicated by the prompt, and outputs the inference result in a data format such as voice data and text data. Here, inference refers to, for example, analysis, classification, prediction, and / or summarization.
[0709] In the above embodiment, an example in which the specific processing is performed by the data processing device 12 has been given, but the technology of the present disclosure is not limited to this, and the specific processing may be performed by the smart glasses 214.
[0710] [Third embodiment]
[0711] FIG. 5 shows an example of the configuration of a data processing system 310 according to the third embodiment.
[0712] 5, the data processing system 310 includes the data processing device 12 and a headset type terminal 314. An example of the data processing device 12 is a server.
[0713] The data processing device 12 includes a computer 22, a database 24, and a communication I / F 26. The computer 22 is an example of a "computer" according to the technology of the present disclosure. The computer 22 includes a processor 28, a RAM 30, and a storage 32. The processor 28, the RAM 30, and the storage 32 are connected to a bus 34. The database 24 and the communication I / F 26 are also connected to the bus 34. The communication I / F 26 is connected to a network 54. Examples of the network 54 include a WAN (Wide Area Network) and / or a LAN (Local Area Network).
[0714] The headset type terminal 314 includes a computer 36, a microphone 238, a speaker 240, a camera 42, a communication I / F 44, and a display 343. The computer 36 includes a processor 46, a RAM 48, and a storage 50. The processor 46, the RAM 48, and the storage 50 are connected to a bus 52. The microphone 238, the speaker 240, the camera 42, and the display 343 are also connected to the bus 52.
[0715] The microphone 238 receives instructions and the like from the user 20 by receiving voice uttered by the user 20. The microphone 238 captures the voice uttered by the user 20, converts the captured voice into audio data, and outputs it to the processor 46. The speaker 240 outputs audio in accordance with instructions from the processor 46.
[0716] Camera 42 is a small digital camera equipped with an optical system including a lens, aperture, and shutter, and an imaging element such as a CMOS (Complementary Metal-Oxide-Semiconductor) image sensor or a CCD (Charge Coupled Device) image sensor, and captures images of the surroundings of user 20 (for example, an imaging range defined by an angle of view equivalent to the field of vision of a typical healthy person).
[0717] The communication I / F 44 is connected to a network 54. The communication I / Fs 44 and 26 are responsible for the exchange of various information between the processor 46 and the processor 28 via the network 54. The exchange of various information between the processor 46 and the processor 28 using the communication I / Fs 44 and 26 is carried out in a secure state.
[0718] Fig. 6 shows an example of the main functions of the data processing device 12 and the headset type terminal 314. As shown in Fig. 6, in the data processing device 12, a specific process is performed by the processor 28. A specific process program 56 is stored in the storage 32.
[0719] The specific processing program 56 is an example of a "program" according to the technology of the present disclosure. The processor 28 reads the specific processing program 56 from the storage 32 and executes the read specific processing program 56 on the RAM 30. The specific processing is realized by the processor 28 operating as a specific processing unit 290 in accordance with the specific processing program 56 executed on the RAM 30.
[0720] The storage 32 stores a data generation model 58 and an emotion identification model 59. The data generation model 58 and the emotion identification model 59 are used by the identification processing unit 290.
[0721] In the headset type terminal 314, a reception output process is performed by the processor 46. A reception output program 60 is stored in the storage 50. The processor 46 reads the reception output program 60 from the storage 50 and executes the read reception output program 60 on the RAM 48. The reception output process is realized by the processor 46 operating as the control unit 46A in accordance with the reception output program 60 executed on the RAM 48.
[0722] Next, a description will be given of the identification process performed by the identification processing unit 290 of the data processing device 12. In the following description, the data processing device 12 will be referred to as the "server" and the headset type terminal 314 will be referred to as the "terminal."
[0723] To implement the present invention, a user terminal, a server, and a communication system that links them together are required. A specific embodiment of this system will be described below.
[0724] Data Collection Phase
[0725] The device collects the user's current location information. Using the GPS function, it periodically records the user's current location and sends it to the server. The device also collects the user's daily behavioral data. This behavioral data includes payment history, calendar schedule information, and fitness tracker data. The device also collects the user's music preference data. Specifically, this includes the history of music that has been played in the past and the ratings that the user has given to songs. This data is also sent to the server.
[0726] Mood estimation phase
[0727] The server receives location information, behavioral data, and music preference data sent from the device and integrates them to create a user profile. Based on this profile, the server runs an algorithm to estimate the user's current psychological state. For example, say you work out at the gym at 7:00 AM and then have coffee at a cafe at 9:00 AM. From these actions, the server estimates your psychological state as "refreshed and active."
[0728] Playlist Generation Phase
[0729] The server selects the most suitable songs based on the estimated psychological state and music preference data. It then generates a playlist using the selected songs. The generated playlist includes song titles, artist names, playback links, and other information. This playlist is then sent to the device.
[0730] Request handling phase
[0731] A user can make requests for a playlist that is currently being played, such as "more upbeat songs" or "more calming songs." This request is sent to the server via the device. The server reevaluates the playlist based on the received request and adds new songs. In this way, the playlist can respond to the user's detailed requests. The updated playlist is then sent back to the device, where the user can listen to it.
[0732] Specific examples
[0733] Consider a scenario where a user goes to the gym at 7am and then has coffee at a cafe at 9am.
[0734] 1. The device records these activities as location and behavioral data and sends them to a server.
[0735] 2. The server receives this data and assumes the user is "refreshed and active."
[0736] 3. The server generates a playlist containing up-tempo pop songs that fit this state of mind and sends it to the device.
[0737] 4. The user starts listening to the playlist and requests a lighter song.
[0738] 5. The server receives this request, updates the playlist to include more upbeat songs, and sends it back to the device.
[0739] 6. The user can enjoy the updated playlist.
[0740] In this way, it is possible to provide music that best suits the user's mood.
[0741] The processing flow will be explained below.
[0742] Step 1:
[0743] The device uses the GPS function to obtain geographical data to obtain the user's current location information, and then sends the obtained location information to the server.
[0744] Step 2:
[0745] The device collects data about the user's daily activities, such as payment history, calendar schedule, and fitness tracker data, and transmits this data to a server.
[0746] Step 3:
[0747] The device collects the user's music preference data, specifically the history of songs played in the past and the user's ratings of songs, and transmits this data to the server.
[0748] Step 4:
[0749] The server combines location, behavioral, and music preference data to create a user profile that comprehensively describes the user's current situation.
[0750] Step 5:
[0751] The server estimates the user's current state of mind based on the user profile. For example, if the user went to the gym in the morning and then spent time at a cafe, the server estimates that the user feels refreshed and active.
[0752] Step 6:
[0753] The server selects the most suitable songs based on the estimated psychological state and music preference data, and generates a playlist based on the selected songs.
[0754] Step 7:
[0755] The server sends the generated playlist to the device, which includes detailed information about the selected songs (song title, artist name, playback link, etc.).
[0756] Step 8:
[0757] Users can make requests for songs on a playlist, such as "more upbeat songs" or "more calm songs," and these requests are sent to the server via the terminal.
[0758] Step 9:
[0759] The server receives the request from the user, re-evaluates the playlist according to the request, reselects songs based on the new criteria, and updates the playlist.
[0760] Step 10:
[0761] The server then sends the updated playlist back to the terminal, where the user can play the updated playlist and enjoy the new songs.
[0762] Through the above processing steps, a musical experience that matches the user's current mental state and specific requests can be provided.
[0763] Example 1
[0764] Next, a description will be given of Example 1. In the following description, the data processing device 12 will be referred to as a "server" and the headset type terminal 314 will be referred to as a "terminal."
[0765] In modern society, users are surrounded by vast amounts of information and data. Providing music that matches a user's psychological state and mood is crucial for improving individual satisfaction. However, conventional music recommendation systems struggle to comprehensively analyze a user's current location, behavioral data, and musical preference data to provide optimal music in real time. They also lack the ability to flexibly update playlists in response to user requests. Given these circumstances, there is a need for a system that can provide music that matches a user's current psychological state and quickly respond to user requests.
[0766] The specific processing by the specific processing unit 290 of the data processing device 12 in the first embodiment is realized by the following means.
[0767] In this invention, the server
[0768] means for collecting current location information from a user's information processing device;
[0769] means for collecting behavioral data from a user's information processing device;
[0770] means for collecting music preference data from a user's information processing device;
[0771] This allows the system to integrate collected data to estimate the user's current psychological state using a generative AI model, generate an optimal music playlist based on the estimation results, and reevaluate and update the playlist according to user requests.
[0772] An "information processing device" is a device that collects, processes, stores, and communicates data, and examples of such devices include smartphones, tablets, and wearable devices.
[0773] "Current location information" is data indicating the geographic location of the user, and is information obtained using GPS or other location detection technology.
[0774] "Behavioral data" is information about a user's daily activities, including payment history, calendar schedule information, fitness tracker data, and the like.
[0775] "Music preference data" is information relating to the user's tastes and ratings of music, and includes the history of songs played in the past and rating data.
[0776] A "generative AI model" is a model of an algorithm or program that uses artificial intelligence technology to analyze and generate data, and is used to estimate a user's psychological state.
[0777] "Mental state" is an indicator of the user's current mental and emotional state, such as feeling refreshed, active, or relaxed.
[0778] A "playlist" is a list of songs to be played in a specific order, and is generated based on the user's psychological state and musical preference data.
[0779] A "request" refers to a request or wish made by a user regarding a playlist that is being played, and specifically includes requests such as "a more upbeat song" or "a more calming song."
[0780] The present invention relates to a system that provides an optimal music playlist based on the user's psychological state and also responds to user requests. Implementing this system requires a user's information processing device, a server, and a communication system that links them.
[0781] Data Collection Phase
[0782] The device collects the user's current location information. This location information is obtained using GPS and updated periodically. Examples of devices like smartphones and wearable devices. The device also collects the user's daily behavioral data. This behavioral data includes payment history (collected via NFC or QR code readers), calendar schedule information (obtained from the Google Calendar API or Apple Calendar API), and fitness tracker data (obtained from the Apple HealthKit or Google Fit API). The device also collects the user's music preference data. Specifically, it collects playback history (using the Spotify API or Apple Music API) and song rating data. The collected data is sent from the device to a server.
[0783] Data analysis phase
[0784] The server receives location information, behavioral data, and music preference data sent from the device and integrates them to create a user profile. The profile reflects the user's behavioral patterns and music preferences. Based on this profile, the server uses a generative AI model to estimate the user's current psychological state. For example, from data such as "The user worked out at the gym at 7am and then had coffee at a cafe at 9am," the server generates a prompt that estimates the user's psychological state as "Refreshed and Active." The following prompt is a concrete example:
[0785] "A user works out at the gym at 7am and then has coffee at a cafe at 9am. Generate a playlist with appropriate uptempo pop songs to infer that the user is refreshed and active."
[0786] Playlist Generation Phase
[0787] The server selects the most suitable songs based on the estimated psychological state and music preference data. Songs are searched from a database to select those that suit the user's psychological state. For example, for a "refreshed and active" psychological state, up-tempo pop songs are selected. A playlist is generated using the selected songs, including information such as song title, artist name, and playback link. The generated playlist is sent from the server to the device.
[0788] Request handling phase
[0789] Users can make requests for songs, such as "more upbeat songs" or "more calming songs," to the playlist being played. These requests are sent to the server via the device. The server reevaluates the playlist based on the received request and adds new songs. The updated playlist is then sent back to the device, allowing the user to enjoy the new playlist.
[0790] In this way, the system not only provides music that is optimal for the user's psychological state, but also has the feature of being able to quickly respond to user requests. The above is an embodiment of the present invention.
[0791] The flow of the identification process in the first embodiment will be described with reference to FIG.
[0792] Step 1: Start collecting data
[0793] The device starts collecting data when the user starts using the system. The input is the user's action, and the output is the start of the data collection process. For example, this process can be started by pressing a "start data collection" button on the app or device.
[0794] Step 2: Obtaining location information
[0795] The device uses the GPS function to obtain the user's current location information. The input is location information from the GPS sensor, and the output is to temporarily store this information within the device. Specifically, the device obtains location information once every minute and records it as latitude and longitude.
[0796] Step 3: Behavioral data collection
[0797] The device collects user behavioral data. The input is data from various sensors and APIs, and the output is collected behavioral data. For example, payment history is obtained via an NFC sensor or QR code reader, and calendar schedule information is obtained from the Google Calendar API or Apple Calendar API. Fitness tracker data is also collected using Apple HealthKit or Google Fit API.
[0798] Step 4: Collecting Music Preference Data
[0799] The device collects the user's music preference data. The input is data from streaming service APIs (Spotify API and Apple Music API), and the output is the collection of music preference data. Specifically, it collects information on playback history and user actions such as "like" and "skip."
[0800] Step 5: Send data
[0801] The device sends the collected location information, behavioral data, and music preference data to a server. The input is various data stored on the device, and the output is sending this data to the server. This communication is carried out via the Internet.
[0802] Step 6: Create a user profile
[0803] The server creates a user profile based on the received data. The input is the data received by the server, and the output is the creation of a user profile. Specifically, the server organizes the data into categories and uses statistical methods and machine learning models to analyze the user's behavioral patterns and musical preferences.
[0804] Step 7: Estimate mental state
[0805] The server uses a generative AI model to estimate the user's psychological state. The input is the user profile and behavioral data collected at that time, and the output is the estimated psychological state. Specifically, a prompt sentence is input into the generative AI model, and the psychological state is estimated from the response. For example, it may estimate that "the user worked out at the gym at 7 a.m. and had coffee at a cafe at 9 a.m., so they are refreshed and active."
[0806] Step 8: Selecting songs and creating a playlist
[0807] The server selects the most suitable songs based on the estimated psychological state and music preference data. The input is the psychological state and music preference data, and the output is a generated playlist. Specifically, it searches for suitable songs from a database and creates a playlist. For example, it selects an up-tempo pop song that is suitable for the "active" psychological state.
[0808] Step 9: Send Playlist
[0809] The server sends the generated playlist to the terminal. The input is the generated playlist, and the output is the playlist arriving at the terminal. This communication is performed via the Internet, and real-time updates are possible.
[0810] Step 10: Respond to requests
[0811] A user can make requests to a playlist that is currently playing, such as "more upbeat songs" or "more calming songs." The input is the user's request, and the output is a new playlist that has been re-evaluated based on that request. Specifically, the server receives the request, re-evaluates the playlist, and adds new songs. This updated playlist is then sent back to the device, where the user can listen to it.
[0812] The above are the processing steps of the program of this system and its specific operations.
[0813] (Application example 1)
[0814] Next, a description will be given of Application Example 1. In the following description, the data processing device 12 will be referred to as a "server" and the headset type terminal 314 will be referred to as a "terminal."
[0815] In recent years, the spread of self-driving vehicles has created a demand for improved in-car entertainment experiences. However, conventional in-car entertainment systems have faced the challenge of providing music and content that is tailored to the user's current psychological state and mood. There is a demand for improving in-car comfort and satisfaction by providing music that matches the user's mood.
[0816] The specific processing by the specific processing unit 290 of the data processing device 12 in the application example 1 is realized by the following means.
[0817] In this invention, the server includes means for collecting current location information from the user's terminal, means for collecting behavioral data from the user's terminal, and means for collecting music preference data from the user's terminal, thereby making it possible to provide a music playlist that is optimal for the user's mood in the autonomous vehicle.
[0818] A "user's terminal" is a mobile information terminal such as a smartphone or tablet owned by a user.
[0819] "Current location information" means a user's real-time geographic location data obtained using location measurement technology such as GPS.
[0820] "Behavioral data" refers to information about a user's behavior, such as fitness tracker data or schedule information.
[0821] "Music preference data" is data relating to the user's musical preferences, such as the history of songs that the user has played in the past and their ratings of those songs.
[0822] "Mental state" is information that indicates the user's current mood and emotional state.
[0823] A "playlist" is a list of multiple songs provided to a user.
[0824] A "request" is a request or instruction given by a user to a playlist.
[0825] An "autonomous vehicle" is a vehicle that operates autonomously and requires minimal user intervention.
[0826] An "entertainment system" is a system that provides entertainment content such as music and video in a vehicle.
[0827] To implement the present invention, a user terminal, a server, and a communication system that links them together are required. A specific embodiment of this system will be described below.
[0828] 1. Data Collection Phase
[0829] The device collects the user's current location information. Using the GPS function, it periodically records the user's current location and sends it to the server. The device also collects the user's daily behavioral data. This behavioral data includes payment history, calendar schedule information, and fitness tracker data. The device also collects the user's music preference data. Specifically, this includes the history of music that has been played in the past and the ratings that the user has given to songs. This data is also sent to the server.
[0830] 2. Mood estimation phase
[0831] The server receives location information, behavioral data, and music preference data sent from the device and integrates them to create a user profile. Based on this profile, the server runs an algorithm to estimate the user's current psychological state. For example, say you work out at the gym at 7:00 AM and then have coffee at a cafe at 9:00 AM. From these actions, the server estimates your psychological state as "refreshed and active."
[0832] 3. Playlist generation phase
[0833] The server selects the most suitable songs based on the estimated psychological state and music preference data. It then generates a playlist using the selected songs. The generated playlist includes song titles, artist names, playback links, and other information. This playlist is then sent to the device.
[0834] 4. Request Response Phase
[0835] A user can make requests for a playlist that is currently being played, such as "more upbeat songs" or "more calming songs." This request is sent to the server via the device. The server reevaluates the playlist based on the received request and adds new songs. In this way, the playlist can respond to the user's detailed requests. The updated playlist is then sent back to the device, where the user can listen to it.
[0836] Specific examples
[0837] Consider a scenario where a user goes to the gym at 7am and then has coffee at a cafe at 9am.
[0838] 1. The device records these activities as location and behavioral data and sends them to a server.
[0839] 2. The server receives this data and assumes the user is "refreshed and active."
[0840] 3. The server generates a playlist containing up-tempo music that matches this psychological state (e.g., "Uptown Funk") and sends it to the terminal.
[0841] 4. The user starts listening to the playlist and requests a lighter song.
[0842] 5. The server receives this request, updates the playlist to include more upbeat songs, and sends it back to the device.
[0843] 6. The user can enjoy the updated playlist.
[0844] Example of input prompt for generative AI model
[0845] "Generate a music playlist that best suits the user's mood based on the following information:
[0846] Current location: Coordinates (35.6895, 139.6917)
[0847] Behavioral data: 1 hour training at 6am, commuting in the car at 9am
[0848] Music preference data: List of songs you've played in the past (e.g. Happy - Williams, Uptown Funk - Mars)
[0849] Choose the best songs to match the user's mood and provide a playlist."
[0850] The flow of the specific processing in the application example 1 will be described with reference to FIG.
[0851] Step 1:
[0852] The device collects the user's current location information. Specifically, it periodically obtains the user's geographical location data using the GPS function. In addition to this, it also collects the user's daily behavior data (e.g., fitness tracker data and calendar schedule information). It also collects the user's music preference data (such as the history and ratings of music played in the past). This data is then sent to the server.
[0853] Input: User location information, behavioral data, music preference data
[0854] Output: Consolidated data packet sent to the server
[0855] Step 2:
[0856] The server receives location information, behavioral data, and music preference data sent from the device and integrates them to create a user profile. Specifically, it uses a data analysis algorithm to integrate this diverse information and generate a user profile.
[0857] Input: location information, behavioral data, music preference data
[0858] Output: Unified user profile
[0859] Step 3:
[0860] The server runs an algorithm to estimate the user's current psychological state based on the generated user profile. Specifically, it uses a machine learning model to analyze the user's behavioral patterns and music preferences, and estimates the user's current psychological state (e.g., "refreshed and active").
[0861] Input: User profile
[0862] Output: Estimated mental state
[0863] Step 4:
[0864] The server selects the most suitable songs based on the estimated psychological state and music preference data. Specifically, it searches a music preference database to create a list of songs that suit the user's current mood. The server then generates a playlist using the selected songs.
[0865] Input: Psychological state, music preference data
[0866] Output: Generated playlist
[0867] Step 5:
[0868] The server then sends the generated playlist to the user's device. Specifically, the playlist contains song titles, artist names, playback links, etc.
[0869] Input: Generated playlist
[0870] Output: Playlist sent to user device
[0871] Step 6:
[0872] The user issues a request (e.g., "more upbeat songs," "more calm songs") to the playlist being played. This request is sent to the server via the device. The server reevaluates the playlist based on the received request and adds new songs.
[0873] Input: A request from the user
[0874] Output: Updated playlist
[0875] Step 7:
[0876] The server then sends the updated playlist back to the terminal, allowing the user to listen to the latest music playlist.
[0877] Input: Updated playlist
[0878] Output: The updated playlist sent to the user's device
[0879] (Example of input prompt for generative AI model)
[0880] "Generate a music playlist that best suits the user's mood based on the following information:
[0881] Current location: Coordinates (35.6895, 139.6917)
[0882] Behavioral data: 1 hour training at 6am, commuting in the car at 9am
[0883] Music preference data: List of songs you've played in the past (e.g. Happy - Williams, Uptown Funk - Mars)
[0884] Choose the best songs to match the user's mood and provide a playlist."
[0885] Furthermore, an emotion engine that estimates the user's emotion may be further combined. That is, the identification processing unit 290 may estimate the user's emotion using the emotion identification model 59, and perform identification processing using the user's emotion.
[0886] To implement the present invention, a user terminal, a server, an emotion engine, and a communication system that links these components are required. A specific embodiment of this system will be described below.
[0887] Data Collection Phase
[0888] The device collects the user's current location information. Using the GPS function, it periodically records the user's current location and sends it to the server. The device also collects the user's daily behavioral data. This behavioral data includes payment history, calendar schedule information, and fitness tracker data. The device also collects the user's music preference data. Specifically, this includes the history of music that has been played in the past and the ratings that the user has given to songs. This data is also sent to the server.
[0889] Emotion Recognition Phase
[0890] The device recognizes the user's emotions through an emotion engine. The emotion engine analyzes the user's voice, facial expressions, and biometric information (heart rate, body temperature, etc.) obtained from input devices. The emotion data recognized by the emotion engine is sent to the server along with other collected data.
[0891] Mood estimation phase
[0892] The server receives location information, behavioral data, music preference data, and emotional data from the device and integrates them to create a user profile. Based on this profile, the server runs an algorithm to estimate the user's current psychological state. For example, suppose a user exercises at the gym in the morning and then has coffee at a cafe. Based on these actions and emotions, the server estimates the user's psychological state as "refreshed and active."
[0893] Playlist Generation Phase
[0894] The server selects the most suitable songs based on the estimated psychological state and music preference data. It then generates a playlist using the selected songs. The generated playlist includes song titles, artist names, playback links, and other information. This playlist is then sent to the device.
[0895] Request handling phase
[0896] A user can make requests for a playlist that is currently being played, such as "more upbeat songs" or "more calming songs." This request is sent to the server via the device. The server reevaluates the playlist based on the received request and adds new songs. In this way, the playlist can respond to the user's detailed requests. The updated playlist is then sent back to the device, where the user can listen to it.
[0897] Specific examples
[0898] Consider a scenario where a user goes to the gym at 7am and then has coffee at a cafe at 9am.
[0899] 1. The device records these activities as location and behavioral data and sends them to a server.
[0900] 2. The device uses an emotion engine to recognize the user's facial expressions, voice, and heart rate, and estimates that the user is "refreshed and active." This emotion data is also sent to the server.
[0901] 3. The server aggregates this data and assumes that the user is "refreshed and active."
[0902] 4. The server generates a playlist containing up-tempo pop songs that fit this state of mind and sends it to the device.
[0903] 5. The user starts listening to the playlist and requests a lighter song.
[0904] 6. The server receives this request, updates the playlist to include more upbeat songs, and sends it back to the device.
[0905] 7. The user can enjoy the updated playlist.
[0906] In this way, by utilizing comprehensive data, including the user's emotional state, it is possible to provide a more advanced and personalized music experience.
[0907] The processing flow will be explained below.
[0908] Step 1:
[0909] The device uses the GPS function to obtain geographical data to obtain the user's current location information, and then sends the obtained location information to the server.
[0910] Step 2:
[0911] The device collects data about the user's daily activities, such as payment history, calendar schedule, and fitness tracker data, and transmits this data to a server.
[0912] Step 3:
[0913] The device collects the user's music preference data, specifically the history of songs played in the past and the user's ratings of songs, and transmits this data to the server.
[0914] Step 4:
[0915] The device uses an emotion engine to recognize the user's emotions. The emotion engine analyzes the user's voice, facial expressions, and biometric information (heart rate, body temperature, etc.) obtained from the input device. The resulting emotion data is sent to the server.
[0916] Step 5:
[0917] The server receives location information, behavioral data, music preference data, and emotional data and integrates them to create a user profile that comprehensively describes the user's current situation.
[0918] Step 6:
[0919] The server estimates the user's current state of mind based on the user profile. For example, if the user works out at the gym in the morning and then spends time at a cafe, the server estimates that the user feels refreshed and active.
[0920] Step 7:
[0921] The server selects the most suitable songs based on the estimated psychological state and music preference data, and generates a playlist based on the selected songs. This playlist includes song titles, artist names, playback links, etc.
[0922] Step 8:
[0923] The server transmits the generated playlist to the terminal, and the playlist data is stored on the user's terminal.
[0924] Step 9:
[0925] The user sends a request to the playlist being played, such as "more upbeat songs" or "more calm songs." This request is sent from the terminal to the server.
[0926] Step 10:
[0927] The server re-evaluates the playlist based on the user's request, adds new songs that meet the criteria, and sends the updated playlist back to the device.
[0928] Step 11:
[0929] The user receives the updated playlist and begins playback, allowing them to enjoy music that better suits their mood.
[0930] Through this series of processing steps, a highly personalized music experience can be provided based on the user's current emotions and behavior.
[0931] Example 2
[0932] Next, a description will be given of Example 2. In the following description, the data processing device 12 will be referred to as a "server" and the headset type terminal 314 will be referred to as a "terminal."
[0933] Conventional music recommendation systems often generate playlists based solely on a user's music preference data, without taking into account the user's real-time emotions or behavioral status. As a result, they are unable to suggest songs that fit the user's current psychological state, making it difficult to increase user satisfaction. To solve this problem, a music recommendation system that also takes into account the user's real-time emotions and behavioral data is needed.
[0934] The specific processing by the specific processing unit 290 of the data processing device 12 in the second embodiment is realized by the following means.
[0935] In this invention, the server includes: means for collecting current location information from the user's device; means for collecting behavioral data from the user's device; means for collecting music preference data from the user's device; means having an emotion analysis engine for collecting emotion data from the device; means for estimating the user's current psychological state by integrating the collected location information, behavioral data, music preference data, and emotion data; means for generating an optimal music playlist based on the estimated psychological state and music preference data; means for transmitting the generated playlist to the user's device; means for receiving a request from the user, reevaluating and updating the playlist, and means for transmitting the updated playlist to the user's device. This enables personalized music recommendations that take into account the user's real-time emotions and behavioral status.
[0936] "User's terminal" refers to electronic devices such as mobile terminals and wearable devices used by users.
[0937] "Location information" refers to data regarding the user's current location determined using GPS functions, etc.
[0938] "Behavioral data" refers to data that shows a user's daily activities, including payment history, calendar schedule information, and data from fitness trackers.
[0939] "Music preference data" refers to data that indicates a user's musical preferences, such as a history of songs that the user has played in the past and ratings of songs.
[0940] "Emotion data" refers to data obtained by an emotion analysis engine from a user's facial expressions, voice, and biometric information (e.g., heart rate and body temperature).
[0941] An "emotion analysis engine" refers to a software or hardware system for analyzing a user's emotional state.
[0942] "Mood" refers to information that indicates a user's current emotional or mental state.
[0943] A "playlist" refers to a collection of songs selected based on specific criteria.
[0944] A "request" refers to information indicating a specific instruction or request from a user.
[0945] A "server" refers to a computer system that communicates with user terminals via a network and processes and stores data.
[0946] The present invention relates to a music recommendation system that takes into account real-time user emotion and behavior data. Specific embodiments of the system are described below.
[0947] To implement this system, a user terminal, a server, an emotion analysis engine, and a communication system that links these are required. The following hardware and software are used as the main components.
[0948] 1. User devices: Mobile devices such as smartphones and wearable devices
[0949] 2. Server: High-performance computer system
[0950] 3. Sentiment analysis engine: Speech recognition software (e.g., speech recognition API, face recognition API)
[0951] Data collection
[0952] The device first collects the user's current location information using its GPS function. This data is periodically recorded and sent to a server. It also uses a fitness tracker to obtain exercise data (e.g., heart rate, number of steps), and collects behavioral data from the device's calendar and payment history. Furthermore, it collects the user's music preference data, such as the history of songs played and their ratings, and sends this data to the server.
[0953] Emotion analysis
[0954] The device uses a built-in camera and microphone to transmit the user's facial expressions and voice to an emotion analysis engine. The emotion analysis engine analyzes this data and detects the user's emotional state (e.g., joy, sadness). For example, if the device's camera captures the user's smile and the microphone detects a joyful voice tone, this information is sent to the server as emotion data.
[0955] Data integration and psychological state estimation
[0956] The server receives location information, behavioral data, music preference data, and emotional data from the device. It integrates these data to create a user profile and uses machine learning algorithms to estimate the user's current psychological state. For example, if a user exercises at the gym in the morning and then has coffee at a cafe, the server estimates the user's psychological state as "refreshed and active."
[0957] Playlist Generation
[0958] The server generates an optimal music playlist based on the estimated psychological state and the user's music preference data. The playlist includes song titles, artist names, and play links. The server then sends the playlist to the user's device, allowing the user to play songs based on the playlist.
[0959] Handling the request
[0960] Users can make requests for songs such as "more upbeat" or "more calming" to a playlist that is currently playing. This request is sent to the server via the device, and the server reevaluates the playlist based on the request and adds new songs. The updated playlist is then sent back to the device, where the user can enjoy it.
[0961] Specific examples
[0962] Consider a scenario where a user goes to the gym at 7am and then has coffee at a cafe at 9am. In this scenario, the following happens:
[0963] 1. The device activates the GPS function and records that the user is at the gym.
[0964] 2. Your device receives your heart rate and exercise data from your fitness tracker and records the information.
[0965] 3. The terminal obtains the payment history at the cafe and records that information.
[0966] 4. The device uses its built-in camera and microphone to transmit the user's facial expressions of delight and tone of voice to the emotion analysis engine.
[0967] 5. The server receives and integrates the location information, behavioral data, music preference data, and emotional data.
[0968] 6. The server uses a machine learning algorithm to estimate the user's psychological state as "refreshed and active."
[0969] 7. The server generates a playlist containing up-tempo pop songs based on the estimated psychological state and sends it to the device.
[0970] 8. While listening to a playlist, the user requests "something more upbeat."
[0971] 9. The server receives the request, updates the playlist to include more upbeat songs, and sends it back to the device.
[0972] 10. Users can enjoy the updated playlist.
[0973] Example prompt for a generative AI model:
[0974] "You have behavioral data that shows that a user works out at the gym at 7 AM and drinks coffee at a cafe at 9 AM. Furthermore, analysis by a sentiment analysis engine suggests that the user is in a psychological state of 'refreshed and active.' Based on this data, please generate a playlist to recommend to the user."
[0975] In this way, it is possible to provide an advanced music recommendation system that also takes into account the user's real-time emotions and behavior.
[0976] The flow of the identification process in the second embodiment will be described with reference to FIG.
[0977] Step 1: Data collection phase
[0978] The device collects the user's current location information. Input includes location information obtained from the device's GPS function. This data is recorded with a timestamp and sent to the server. For example, if the user is at the gym, the location coordinates and time are recorded. The device also collects exercise data from a fitness tracker. Input includes heart rate and step count data. This data is also recorded and sent to the server. The device also collects the user's daily behavior data (payment history, calendar schedule information) and music preference data (past playback history and ratings) and sends them to the server. Input is obtained from the smartphone, payment system, calendar app, and music app.
[0979] Step 2: Emotion Recognition Phase
[0980] The device uses a built-in camera and microphone to send the user's facial expressions and voice to an emotion analysis engine. Inputs include facial expression data captured by the camera, voice data picked up by the microphone, and biometric data (heart rate, body temperature, etc.). The emotion analysis engine analyzes this data and determines the user's emotional state. Emotional data such as "happiness" or "sadness" is generated as output and sent to the server. Specifically, the device analyzes the user's smile, which indicates joy, and the tone of their voice, which indicates happiness, and records this as emotion data.
[0981] Step 3: Data integration phase
[0982] The server receives location information, behavioral data, music preference data, and emotional data sent from the device. These multiple data sets serve as input. The server integrates them and performs the necessary processing to create a user profile. This profile is then processed using machine learning algorithms to generate an output that analyzes the interrelationships between each piece of data. For example, if a user works out at the gym in the morning and then has coffee at a cafe, this information can be used to estimate a psychological state of "refreshed and active."
[0983] Step 4: Mental state estimation phase
[0984] The server runs a machine learning algorithm on the integrated data to estimate the user's current state of mind. The input is the user profile generated in the data integration phase. The algorithm analyzes this data and estimates the user's state of mind (e.g., "refreshed and active") as output. Specifically, the server analyzes the correlations between each data set based on time of day, behavioral patterns, and emotional changes.
[0985] Step 5: Playlist generation phase
[0986] The server selects optimal songs based on the estimated psychological state and music preference data. The inputs are the estimated psychological state data and music preference data. Based on these data, the server filters the optimal songs from the music database and generates a playlist. The output is a playlist containing song titles, artist names, play links, etc. Specifically, up-tempo pop songs are selected for users in a "refreshed and active" psychological state. This playlist is then sent from the server to the device.
[0987] Step 6: Request handling phase
[0988] A user can make requests to a playlist that is currently being played. The input is a request sent by the user through the terminal (e.g., "More upbeat songs," "More calm songs"). The server receives this request and re-evaluates the playlist based on the input request. The server adds newly selected songs to the playlist, and an updated playlist is generated as output. Specifically, the server adds songs from its song database that fit the request, and sends the updated playlist back to the terminal.
[0989] (Application example 2)
[0990] Next, a description will be given of Application Example 2. In the following description, the data processing device 12 will be referred to as a "server" and the headset type terminal 314 will be referred to as a "terminal."
[0991] Conventional music recommendation systems are limited to generating simple playlists based on a user's musical preferences, and have difficulty providing a personalized music experience that takes into account the user's emotions and psychological state. This has resulted in problems such as not being able to provide music that matches the user's current mood or emotion, leading to a decrease in satisfaction. Furthermore, the system lacks the ability to flexibly update playlists in response to user requests. A system that can solve these issues is needed.
[0992] The specific processing by the specific processing unit 290 of the data processing device 12 in the application example 2 is realized by the following means.
[0993] In this invention, the server includes means for recognizing a user's emotions by performing voice analysis and facial expression analysis, means for transmitting data based on the user's emotions to the server and estimating the user's mood, and means for generating an optimal music playlist based on the estimated mood, thereby enabling a more personalized music experience according to the user's emotions and psychological state.
[0994] A "user's terminal" is an electronic device that a user uses on a daily basis, and includes smartphones, tablets, smartwatches, etc.
[0995] "Current location information" refers to geographical location data collected from the user's device and is obtained using GPS.
[0996] "Behavioral data" is information about a user's daily activities, including schedules, payment history, fitness tracker data, etc.
[0997] "Music preference data" is information such as the user's favorite music genres and artists, playback history, and ratings.
[0998] "Emotion" indicates the emotional state the user is feeling at that time, and is based on analyzed biometric information, voice, facial expressions, etc.
[0999] "Means for recognizing emotions" refers to technologies and algorithms that analyze a user's voice, facial expressions, and biometric information in order to identify the user's emotional state.
[1000] The "means for estimating mood" is a method for estimating a user's current psychological state using a computer algorithm based on collected user behavioral data, location information, and emotional data.
[1001] A "playlist" is a list of music for a specific purpose or theme, including song titles, artist names, and playback links.
[1002] The "means for receiving a request" is an interface through which the system receives requests or wishes from the user, such as voice input or touch input.
[1003] The "means for reevaluating and updating a playlist" refers to a means for reviewing a playlist that has already been generated and adding or changing new songs, etc., based on a request received from a user.
[1004] A "generative AI model" is an artificial intelligence model used to generate new music playlists based on collected data, and includes machine learning algorithms.
[1005] A "prompt sentence" is a text sentence that serves as input to a generative AI model to perform a specific task.
[1006] To implement this invention, a server, a terminal, an emotion engine, and a communication system that links them together are required. A specific embodiment of this system will be described below.
[1007] Data Collection Phase
[1008] The device collects the user's current location information. Using its GPS function, it periodically records the user's current location and sends it to the server. The device also collects the user's daily behavioral data. This behavioral data includes payment history, calendar schedule information, and fitness tracker data. The device also collects the user's music preference data, including the history of music played in the past and the ratings the user has given to songs. This data is also sent to the server.
[1009] Emotion Recognition Phase
[1010] The device recognizes the user's emotions through an emotion engine. The emotion engine analyzes the user's voice, facial expressions, and biometric information (heart rate, body temperature, etc.) obtained from input devices. The emotion data recognized by the emotion engine is sent to the server along with other collected data.
[1011] Mood estimation phase
[1012] The server receives location information, behavioral data, music preference data, and emotional data from the device and integrates them to create a user profile. Based on this profile, the server runs an algorithm to estimate the user's current psychological state. For example, if a user exercises at the gym in the morning and then has coffee at a cafe, these actions and emotions can be used to estimate the user's psychological state as "refreshed and active."
[1013] Playlist Generation Phase
[1014] The server selects the most suitable songs based on the estimated psychological state and music preference data. The generated playlist includes song titles, artist names, and playback links. This playlist is then sent to the device and provided to the user.
[1015] Request handling phase
[1016] Users can make requests for songs, such as "more upbeat" or "more calming," to a playlist that is currently playing. These requests are sent to the server via the device. The server reevaluates the playlist based on the received request and adds new songs. The updated playlist is then sent back to the device, where the user can listen to it.
[1017] Specific examples
[1018] Consider a scenario where a user goes to the gym at 7am and then has coffee at a cafe at 9am.
[1019] 1. The device records these activities as location and behavioral data and sends them to a server.
[1020] 2. The device uses an emotion engine to recognize the user's facial expressions, voice, and heart rate, and estimates that the user is "refreshed and active." This emotion data is also sent to the server.
[1021] 3. The server aggregates this data and assumes that the user is "refreshed and active."
[1022] 4. The server generates a playlist containing up-tempo pop songs that fit this state of mind and sends it to the device.
[1023] 5. The user starts listening to the playlist and requests a lighter song.
[1024] 6. The server receives this request, updates the playlist to include more upbeat songs, and sends it back to the device.
[1025] 7. The user can enjoy the updated playlist.
[1026] Prompt sentences to input to the generative AI model
[1027] Use the following prompt:
[1028] Based on the user's daily behavior data, such as working out at the gym at 7am and then drinking coffee at a cafe at 9am, estimate his current psychological state as "refreshed and active." His musical preferences are pop and electronic. Generate a playlist that is appropriate for this state.
[1029] This allows for a more personalized music experience based on the user's emotions and behavior.
[1030] The flow of the specific processing in the application example 2 will be described with reference to FIG.
[1031] Step 1:
[1032] The device collects the user's current location information. Using the GPS function, it periodically obtains the user's latitude and longitude and sends that location information to the server. The input is the device's GPS data, and the output is the location information data sent to the server.
[1033] Step 2:
[1034] The device collects the user's daily behavioral data, including payment history, calendar schedule information, and fitness tracker data, and sends it to a server. The input is this behavioral data, and the output is the integrated data sent to the server.
[1035] Step 3:
[1036] The device collects the user's music preference data, including the history of songs played and rating information, and sends it to the server. The input is the music playback history and rating information, and the output is the music preference data sent to the server.
[1037] Step 4:
[1038] The device recognizes the user's emotions through an emotion engine. The emotion engine analyzes the user's voice, facial expression, and biometric information (heart rate, body temperature, etc.) to recognize the user's emotional state and output it as data. This data is sent to the server. The input at this time is voice, facial expression, biometric information, etc., and the output is recognized emotional data.
[1039] Step 5:
[1040] The server receives location information, behavioral data, music preference data, and emotional data sent from the device and integrates them to create a user profile. The server's algorithm analyzes this data and estimates the user's current psychological state. The input is the integrated user data, and the output is an estimated psychological state.
[1041] Step 6:
[1042] The server generates an optimal music playlist based on the estimated psychological state and music preference data. Using a generative AI model, the server selects songs that best fit the user's psychological state and creates a playlist. The input is the psychological state and music preference data, and the output is the generated playlist.
[1043] Step 7:
[1044] The server sends the generated playlist to the user's device and provides it to the user. The playlist includes song titles, artist names, playback links, etc. The input is the generated playlist, and the output is the playlist sent to the device.
[1045] Step 8:
[1046] A user can make a request for a playlist that is currently being played. For example, a request for "more upbeat songs" is sent to the server via the terminal. The input is the user's request, and the output is the request data sent to the server.
[1047] Step 9:
[1048] The server reevaluates the playlist based on the received request and adds new songs. It uses a generative AI model to generate a new playlist and sends it to the device. The input is the user request and existing playlist data, and the output is the updated playlist.
[1049] Step 10:
[1050] The terminal plays the updated playlist sent from the server and provides it to the user, allowing the user to enjoy a music experience tailored to their requests. The input is the updated playlist data, and the output is the songs to be played.
[1051] The specific processing unit 290 transmits the result of the specific processing to the headset type terminal 314. In the headset type terminal 314, the control unit 46A causes the speaker 240 and the display 343 to output the result of the specific processing. The microphone 238 acquires audio indicating a user input regarding the result of the specific processing. The control unit 46A transmits audio data indicating the user input acquired by the microphone 238 to the data processing device 12. In the data processing device 12, the specific processing unit 290 acquires the audio data.
[1052] The data generation model 58 is a so-called generative AI (Artificial Intelligence). An example of the data generation model 58 is ChatGPT (Internet Search<URL: https: / / openai.com / blog / chatgpt> ), Gemini (Internet search <url: https: gemini.google.com ?hl="ja">) and other generation AIs. The data generation model 58 is obtained by performing deep learning on a neural network. A prompt including an instruction is input to the data generation model 58, and inference data such as voice data indicating voice, text data indicating text, and image data indicating an image is also input. The data generation model 58 performs inference on the input inference data in accordance with the instruction indicated by the prompt, and outputs the inference result in a data format such as voice data and text data. Here, inference refers to, for example, analysis, classification, prediction, and / or summarization.
[1053] In the above embodiment, an example was given in which the specific processing is performed by the data processing device 12, but the technology of the present disclosure is not limited to this, and the specific processing may be performed by the headset type terminal 314.
[1054] [Fourth embodiment]
[1055] FIG. 7 shows an example of the configuration of a data processing system 410 according to the fourth embodiment.
[1056] 7, a data processing system 410 includes a data processing device 12 and a robot 414. An example of the data processing device 12 is a server.
[1057] The data processing device 12 includes a computer 22, a database 24, and a communication I / F 26. The computer 22 is an example of a "computer" according to the technology of the present disclosure. The computer 22 includes a processor 28, a RAM 30, and a storage 32. The processor 28, the RAM 30, and the storage 32 are connected to a bus 34. The database 24 and the communication I / F 26 are also connected to the bus 34. The communication I / F 26 is connected to a network 54. Examples of the network 54 include a WAN (Wide Area Network) and / or a LAN (Local Area Network).
[1058] The robot 414 includes a computer 36, a microphone 238, a speaker 240, a camera 42, a communication I / F 44, and a control target 443. The computer 36 includes a processor 46, a RAM 48, and a storage 50. The processor 46, the RAM 48, and the storage 50 are connected to a bus 52. The microphone 238, the speaker 240, the camera 42, and the control target 443 are also connected to the bus 52.
[1059] The microphone 238 receives instructions and the like from the user 20 by receiving voice uttered by the user 20. The microphone 238 captures the voice uttered by the user 20, converts the captured voice into audio data, and outputs it to the processor 46. The speaker 240 outputs audio in accordance with instructions from the processor 46.
[1060] Camera 42 is a small digital camera equipped with an optical system including a lens, aperture, and shutter, and an imaging element such as a CMOS (Complementary Metal-Oxide-Semiconductor) image sensor or a CCD (Charge Coupled Device) image sensor, and captures images of the surroundings of user 20 (for example, an imaging range defined by an angle of view equivalent to the field of vision of a typical healthy person).
[1061] The communication I / F 44 is connected to a network 54. The communication I / Fs 44 and 26 are responsible for the exchange of various information between the processor 46 and the processor 28 via the network 54. The exchange of various information between the processor 46 and the processor 28 using the communication I / Fs 44 and 26 is carried out in a secure state.
[1062] The control object 443 includes a display device, LEDs in the eyes, and motors for driving the arms, hands, and feet. The posture and gestures of the robot 414 are controlled by controlling the motors of the arms, hands, and feet. Some of the emotions of the robot 414 can be expressed by controlling these motors. In addition, the facial expressions of the robot 414 can also be expressed by controlling the light emission state of the LEDs in the eyes of the robot 414.
[1063] Fig. 8 shows an example of the main functions of the data processing device 12 and the robot 414. As shown in Fig. 8, in the data processing device 12, a specific process is performed by the processor 28. A specific process program 56 is stored in the storage 32.
[1064] The specific processing program 56 is an example of a "program" according to the technology of the present disclosure. The processor 28 reads the specific processing program 56 from the storage 32 and executes the read specific processing program 56 on the RAM 30. The specific processing is realized by the processor 28 operating as a specific processing unit 290 in accordance with the specific processing program 56 executed on the RAM 30.
[1065] The storage 32 stores a data generation model 58 and an emotion identification model 59. The data generation model 58 and the emotion identification model 59 are used by the identification processing unit 290.
[1066] In the robot 414, the processor 46 performs the reception output process. A reception output program 60 is stored in the storage 50. The processor 46 reads the reception output program 60 from the storage 50 and executes the read reception output program 60 on the RAM 48. The reception output process is realized by the processor 46 operating as the control unit 46A in accordance with the reception output program 60 executed on the RAM 48.
[1067] Next, a description will be given of the specific processing performed by the specific processing unit 290 of the data processing device 12. In the following description, the data processing device 12 will be referred to as a "server" and the robot 414 will be referred to as a "terminal."
[1068] To implement the present invention, a user terminal, a server, and a communication system that links them together are required. A specific embodiment of this system will be described below.
[1069] Data Collection Phase
[1070] The device collects the user's current location information. Using the GPS function, it periodically records the user's current location and sends it to the server. The device also collects the user's daily behavioral data. This behavioral data includes payment history, calendar schedule information, and fitness tracker data. The device also collects the user's music preference data. Specifically, this includes the history of music that has been played in the past and the ratings that the user has given to songs. This data is also sent to the server.
[1071] Mood estimation phase
[1072] The server receives location information, behavioral data, and music preference data sent from the device and integrates them to create a user profile. Based on this profile, the server runs an algorithm to estimate the user's current psychological state. For example, say you work out at the gym at 7:00 AM and then have coffee at a cafe at 9:00 AM. From these actions, the server estimates your psychological state as "refreshed and active."
[1073] Playlist Generation Phase
[1074] The server selects the most suitable songs based on the estimated psychological state and music preference data. It then generates a playlist using the selected songs. The generated playlist includes song titles, artist names, playback links, and other information. This playlist is then sent to the device.
[1075] Request handling phase
[1076] A user can make requests for a playlist that is currently being played, such as "more upbeat songs" or "more calming songs." This request is sent to the server via the device. The server reevaluates the playlist based on the received request and adds new songs. In this way, the playlist can respond to the user's detailed requests. The updated playlist is then sent back to the device, where the user can listen to it.
[1077] Specific examples
[1078] Consider a scenario where a user goes to the gym at 7am and then has coffee at a cafe at 9am.
[1079] 1. The device records these activities as location and behavioral data and sends them to a server.
[1080] 2. The server receives this data and assumes the user is "refreshed and active."
[1081] 3. The server generates a playlist containing up-tempo pop songs that fit this state of mind and sends it to the device.
[1082] 4. The user starts listening to the playlist and requests a lighter song.
[1083] 5. The server receives this request, updates the playlist to include more upbeat songs, and sends it back to the device.
[1084] 6. The user can enjoy the updated playlist.
[1085] In this way, it is possible to provide music that best suits the user's mood.
[1086] The processing flow will be explained below.
[1087] Step 1:
[1088] The device uses the GPS function to obtain geographical data to obtain the user's current location information, and then sends the obtained location information to the server.
[1089] Step 2:
[1090] The device collects data about the user's daily activities, such as payment history, calendar schedule, and fitness tracker data, and transmits this data to a server.
[1091] Step 3:
[1092] The device collects the user's music preference data, specifically the history of songs played in the past and the user's ratings of songs, and transmits this data to the server.
[1093] Step 4:
[1094] The server combines location, behavioral, and music preference data to create a user profile that comprehensively describes the user's current situation.
[1095] Step 5:
[1096] The server estimates the user's current state of mind based on the user profile. For example, if the user went to the gym in the morning and then spent time at a cafe, the server estimates that the user feels refreshed and active.
[1097] Step 6:
[1098] The server selects the most suitable songs based on the estimated psychological state and music preference data, and generates a playlist based on the selected songs.
[1099] Step 7:
[1100] The server sends the generated playlist to the device, which includes detailed information about the selected songs (song title, artist name, playback link, etc.).
[1101] Step 8:
[1102] Users can make requests for songs on a playlist, such as "more upbeat songs" or "more calm songs," and these requests are sent to the server via the terminal.
[1103] Step 9:
[1104] The server receives the request from the user, re-evaluates the playlist according to the request, reselects songs based on the new criteria, and updates the playlist.
[1105] Step 10:
[1106] The server then sends the updated playlist back to the terminal, where the user can play the updated playlist and enjoy the new songs.
[1107] Through the above processing steps, a musical experience that matches the user's current mental state and specific requests can be provided.
[1108] Example 1
[1109] Next, a description will be given of Example 1. In the following description, the data processing device 12 will be referred to as a "server" and the robot 414 will be referred to as a "terminal."
[1110] In modern society, users are surrounded by vast amounts of information and data. Providing music that matches a user's psychological state and mood is crucial for improving individual satisfaction. However, conventional music recommendation systems struggle to comprehensively analyze a user's current location, behavioral data, and musical preference data to provide optimal music in real time. They also lack the ability to flexibly update playlists in response to user requests. Given these circumstances, there is a need for a system that can provide music that matches a user's current psychological state and quickly respond to user requests.
[1111] The specific processing by the specific processing unit 290 of the data processing device 12 in the first embodiment is realized by the following means.
[1112] In this invention, the server
[1113] means for collecting current location information from a user's information processing device;
[1114] means for collecting behavioral data from a user's information processing device;
[1115] means for collecting music preference data from a user's information processing device;
[1116] This allows the system to integrate collected data to estimate the user's current psychological state using a generative AI model, generate an optimal music playlist based on the estimation results, and reevaluate and update the playlist according to user requests.
[1117] An "information processing device" is a device that collects, processes, stores, and communicates data, and examples of such devices include smartphones, tablets, and wearable devices.
[1118] "Current location information" is data indicating the geographic location of the user, and is information obtained using GPS or other location detection technology.
[1119] "Behavioral data" is information about a user's daily activities, including payment history, calendar schedule information, fitness tracker data, and the like.
[1120] "Music preference data" is information relating to the user's tastes and ratings of music, and includes the history of songs played in the past and rating data.
[1121] A "generative AI model" is a model of an algorithm or program that uses artificial intelligence technology to analyze and generate data, and is used to estimate a user's psychological state.
[1122] "Mental state" is an indicator of the user's current mental and emotional state, such as feeling refreshed, active, or relaxed.
[1123] A "playlist" is a list of songs to be played in a specific order, and is generated based on the user's psychological state and musical preference data.
[1124] A "request" refers to a request or wish made by a user regarding a playlist that is being played, and specifically includes requests such as "a more upbeat song" or "a more calming song."
[1125] The present invention relates to a system that provides an optimal music playlist based on the user's psychological state and also responds to user requests. Implementing this system requires a user's information processing device, a server, and a communication system that links them.
[1126] Data Collection Phase
[1127] The device collects the user's current location information. This location information is obtained using GPS and updated periodically. Examples of devices like smartphones and wearable devices. The device also collects the user's daily behavioral data. This behavioral data includes payment history (collected via NFC or QR code readers), calendar schedule information (obtained from the Google Calendar API or Apple Calendar API), and fitness tracker data (obtained from the Apple HealthKit or Google Fit API). The device also collects the user's music preference data. Specifically, it collects playback history (using the Spotify API or Apple Music API) and song rating data. The collected data is sent from the device to a server.
[1128] Data analysis phase
[1129] The server receives location information, behavioral data, and music preference data sent from the device and integrates them to create a user profile. The profile reflects the user's behavioral patterns and music preferences. Based on this profile, the server uses a generative AI model to estimate the user's current psychological state. For example, from data such as "The user worked out at the gym at 7am and then had coffee at a cafe at 9am," the server generates a prompt that estimates the user's psychological state as "Refreshed and Active." The following prompt is a concrete example:
[1130] "A user works out at the gym at 7am and then has coffee at a cafe at 9am. Generate a playlist with appropriate uptempo pop songs to infer that the user is refreshed and active."
[1131] Playlist Generation Phase
[1132] The server selects the most suitable songs based on the estimated psychological state and music preference data. Songs are searched from a database to select those that suit the user's psychological state. For example, for a "refreshed and active" psychological state, up-tempo pop songs are selected. A playlist is generated using the selected songs, including information such as song title, artist name, and playback link. The generated playlist is sent from the server to the device.
[1133] Request handling phase
[1134] Users can make requests for songs, such as "more upbeat songs" or "more calming songs," to the playlist being played. These requests are sent to the server via the device. The server reevaluates the playlist based on the received request and adds new songs. The updated playlist is then sent back to the device, allowing the user to enjoy the new playlist.
[1135] In this way, the system not only provides music that is optimal for the user's psychological state, but also has the feature of being able to quickly respond to user requests. The above is an embodiment of the present invention.
[1136] The flow of the identification process in the first embodiment will be described with reference to FIG.
[1137] Step 1: Start collecting data
[1138] The device starts collecting data when the user starts using the system. The input is the user's action, and the output is the start of the data collection process. For example, this process can be started by pressing a "start data collection" button on the app or device.
[1139] Step 2: Obtaining location information
[1140] The device uses the GPS function to obtain the user's current location information. The input is location information from the GPS sensor, and the output is to temporarily store this information within the device. Specifically, the device obtains location information once every minute and records it as latitude and longitude.
[1141] Step 3: Behavioral data collection
[1142] The device collects user behavioral data. The input is data from various sensors and APIs, and the output is collected behavioral data. For example, payment history is obtained via an NFC sensor or QR code reader, and calendar schedule information is obtained from the Google Calendar API or Apple Calendar API. Fitness tracker data is also collected using Apple HealthKit or Google Fit API.
[1143] Step 4: Collecting Music Preference Data
[1144] The device collects the user's music preference data. The input is data from streaming service APIs (Spotify API and Apple Music API), and the output is the collection of music preference data. Specifically, it collects information on playback history and user actions such as "like" and "skip."
[1145] Step 5: Send data
[1146] The device sends the collected location information, behavioral data, and music preference data to a server. The input is various data stored on the device, and the output is sending this data to the server. This communication is carried out via the Internet.
[1147] Step 6: Create a user profile
[1148] The server creates a user profile based on the received data. The input is the data received by the server, and the output is the creation of a user profile. Specifically, the server organizes the data into categories and uses statistical methods and machine learning models to analyze the user's behavioral patterns and musical preferences.
[1149] Step 7: Estimate mental state
[1150] The server uses a generative AI model to estimate the user's psychological state. The input is the user profile and behavioral data collected at that time, and the output is the estimated psychological state. Specifically, a prompt sentence is input into the generative AI model, and the psychological state is estimated from the response. For example, it may estimate that "the user worked out at the gym at 7 a.m. and had coffee at a cafe at 9 a.m., so they are refreshed and active."
[1151] Step 8: Selecting songs and creating a playlist
[1152] The server selects the most suitable songs based on the estimated psychological state and music preference data. The input is the psychological state and music preference data, and the output is a generated playlist. Specifically, it searches for suitable songs from a database and creates a playlist. For example, it selects an up-tempo pop song that is suitable for the "active" psychological state.
[1153] Step 9: Send Playlist
[1154] The server sends the generated playlist to the terminal. The input is the generated playlist, and the output is the playlist arriving at the terminal. This communication is performed via the Internet, and real-time updates are possible.
[1155] Step 10: Respond to requests
[1156] A user can make requests to a playlist that is currently playing, such as "more upbeat songs" or "more calming songs." The input is the user's request, and the output is a new playlist that has been re-evaluated based on that request. Specifically, the server receives the request, re-evaluates the playlist, and adds new songs. This updated playlist is then sent back to the device, where the user can listen to it.
[1157] The above are the processing steps of the program of this system and its specific operations.
[1158] (Application example 1)
[1159] Next, a description will be given of Application Example 1. In the following description, the data processing device 12 will be referred to as a "server" and the robot 414 will be referred to as a "terminal."
[1160] In recent years, the spread of self-driving vehicles has created a demand for improved in-car entertainment experiences. However, conventional in-car entertainment systems have faced the challenge of providing music and content that is tailored to the user's current psychological state and mood. There is a demand for improving in-car comfort and satisfaction by providing music that matches the user's mood.
[1161] The specific processing by the specific processing unit 290 of the data processing device 12 in the application example 1 is realized by the following means.
[1162] In this invention, the server includes means for collecting current location information from the user's terminal, means for collecting behavioral data from the user's terminal, and means for collecting music preference data from the user's terminal, thereby making it possible to provide a music playlist that is optimal for the user's mood in the autonomous vehicle.
[1163] A "user's terminal" is a mobile information terminal such as a smartphone or tablet owned by a user.
[1164] "Current location information" means a user's real-time geographic location data obtained using location measurement technology such as GPS.
[1165] "Behavioral data" refers to information about a user's behavior, such as fitness tracker data or schedule information.
[1166] "Music preference data" is data relating to the user's musical preferences, such as the history of songs that the user has played in the past and their ratings of those songs.
[1167] "Mental state" is information that indicates the user's current mood and emotional state.
[1168] A "playlist" is a list of multiple songs provided to a user.
[1169] A "request" is a request or instruction given by a user to a playlist.
[1170] An "autonomous vehicle" is a vehicle that operates autonomously and requires minimal user intervention.
[1171] An "entertainment system" is a system that provides entertainment content such as music and video in a vehicle.
[1172] To implement the present invention, a user terminal, a server, and a communication system that links them together are required. A specific embodiment of this system will be described below.
[1173] 1. Data Collection Phase
[1174] The device collects the user's current location information. Using the GPS function, it periodically records the user's current location and sends it to the server. The device also collects the user's daily behavioral data. This behavioral data includes payment history, calendar schedule information, and fitness tracker data. The device also collects the user's music preference data. Specifically, this includes the history of music that has been played in the past and the ratings that the user has given to songs. This data is also sent to the server.
[1175] 2. Mood estimation phase
[1176] The server receives location information, behavioral data, and music preference data sent from the device and integrates them to create a user profile. Based on this profile, the server runs an algorithm to estimate the user's current psychological state. For example, say you work out at the gym at 7:00 AM and then have coffee at a cafe at 9:00 AM. From these actions, the server estimates your psychological state as "refreshed and active."
[1177] 3. Playlist generation phase
[1178] The server selects the most suitable songs based on the estimated psychological state and music preference data. It then generates a playlist using the selected songs. The generated playlist includes song titles, artist names, playback links, and other information. This playlist is then sent to the device.
[1179] 4. Request Response Phase
[1180] A user can make requests for a playlist that is currently being played, such as "more upbeat songs" or "more calming songs." This request is sent to the server via the device. The server reevaluates the playlist based on the received request and adds new songs. In this way, the playlist can respond to the user's detailed requests. The updated playlist is then sent back to the device, where the user can listen to it.
[1181] Specific examples
[1182] Consider a scenario where a user goes to the gym at 7am and then has coffee at a cafe at 9am.
[1183] 1. The device records these activities as location and behavioral data and sends them to a server.
[1184] 2. The server receives this data and assumes the user is "refreshed and active."
[1185] 3. The server generates a playlist containing up-tempo music that matches this psychological state (e.g., "Uptown Funk") and sends it to the terminal.
[1186] 4. The user starts listening to the playlist and requests a lighter song.
[1187] 5. The server receives this request, updates the playlist to include more upbeat songs, and sends it back to the device.
[1188] 6. The user can enjoy the updated playlist.
[1189] Example of input prompt for generative AI model
[1190] "Generate a music playlist that best suits the user's mood based on the following information:
[1191] Current location: Coordinates (35.6895, 139.6917)
[1192] Behavioral data: 1 hour training at 6am, commuting in the car at 9am
[1193] Music preference data: List of songs you've played in the past (e.g. Happy - Williams, Uptown Funk - Mars)
[1194] Choose the best songs to match the user's mood and provide a playlist."
[1195] The flow of the specific processing in the application example 1 will be described with reference to FIG.
[1196] Step 1:
[1197] The device collects the user's current location information. Specifically, it periodically obtains the user's geographical location data using the GPS function. In addition to this, it also collects the user's daily behavior data (e.g., fitness tracker data and calendar schedule information). It also collects the user's music preference data (such as the history and ratings of music played in the past). This data is then sent to the server.
[1198] Input: User location information, behavioral data, music preference data
[1199] Output: Consolidated data packet sent to the server
[1200] Step 2:
[1201] The server receives location information, behavioral data, and music preference data sent from the device and integrates them to create a user profile. Specifically, it uses a data analysis algorithm to integrate this diverse information and generate a user profile.
[1202] Input: location information, behavioral data, music preference data
[1203] Output: Unified user profile
[1204] Step 3:
[1205] The server runs an algorithm to estimate the user's current psychological state based on the generated user profile. Specifically, it uses a machine learning model to analyze the user's behavioral patterns and music preferences, and estimates the user's current psychological state (e.g., "refreshed and active").
[1206] Input: User profile
[1207] Output: Estimated mental state
[1208] Step 4:
[1209] The server selects the most suitable songs based on the estimated psychological state and music preference data. Specifically, it searches a music preference database to create a list of songs that suit the user's current mood. The server then generates a playlist using the selected songs.
[1210] Input: Psychological state, music preference data
[1211] Output: Generated playlist
[1212] Step 5:
[1213] The server then sends the generated playlist to the user's device. Specifically, the playlist contains song titles, artist names, playback links, etc.
[1214] Input: Generated playlist
[1215] Output: Playlist sent to user device
[1216] Step 6:
[1217] The user issues a request (e.g., "more upbeat songs," "more calm songs") to the playlist being played. This request is sent to the server via the device. The server reevaluates the playlist based on the received request and adds new songs.
[1218] Input: A request from the user
[1219] Output: Updated playlist
[1220] Step 7:
[1221] The server then sends the updated playlist back to the terminal, allowing the user to listen to the latest music playlist.
[1222] Input: Updated playlist
[1223] Output: The updated playlist sent to the user's device
[1224] (Example of input prompt for generative AI model)
[1225] "Generate a music playlist that best suits the user's mood based on the following information:
[1226] Current location: Coordinates (35.6895, 139.6917)
[1227] Behavioral data: 1 hour training at 6am, commuting in the car at 9am
[1228] Music preference data: List of songs you've played in the past (e.g. Happy - Williams, Uptown Funk - Mars)
[1229] Choose the best songs to match the user's mood and provide a playlist."
[1230] Furthermore, an emotion engine that estimates the user's emotion may be further combined. That is, the identification processing unit 290 may estimate the user's emotion using the emotion identification model 59, and perform identification processing using the user's emotion.
[1231] To implement the present invention, a user terminal, a server, an emotion engine, and a communication system that links these components are required. A specific embodiment of this system will be described below.
[1232] Data Collection Phase
[1233] The device collects the user's current location information. Using the GPS function, it periodically records the user's current location and sends it to the server. The device also collects the user's daily behavioral data. This behavioral data includes payment history, calendar schedule information, and fitness tracker data. The device also collects the user's music preference data. Specifically, this includes the history of music that has been played in the past and the ratings that the user has given to songs. This data is also sent to the server.
[1234] Emotion Recognition Phase
[1235] The device recognizes the user's emotions through an emotion engine. The emotion engine analyzes the user's voice, facial expressions, and biometric information (heart rate, body temperature, etc.) obtained from input devices. The emotion data recognized by the emotion engine is sent to the server along with other collected data.
[1236] Mood estimation phase
[1237] The server receives location information, behavioral data, music preference data, and emotional data from the device and integrates them to create a user profile. Based on this profile, the server runs an algorithm to estimate the user's current psychological state. For example, suppose a user exercises at the gym in the morning and then has coffee at a cafe. Based on these actions and emotions, the server estimates the user's psychological state as "refreshed and active."
[1238] Playlist Generation Phase
[1239] The server selects the most suitable songs based on the estimated psychological state and music preference data. It then generates a playlist using the selected songs. The generated playlist includes song titles, artist names, playback links, and other information. This playlist is then sent to the device.
[1240] Request handling phase
[1241] A user can make requests for a playlist that is currently being played, such as "more upbeat songs" or "more calming songs." This request is sent to the server via the device. The server reevaluates the playlist based on the received request and adds new songs. In this way, the playlist can respond to the user's detailed requests. The updated playlist is then sent back to the device, where the user can listen to it.
[1242] Specific examples
[1243] Consider a scenario where a user goes to the gym at 7am and then has coffee at a cafe at 9am.
[1244] 1. The device records these activities as location and behavioral data and sends them to a server.
[1245] 2. The device uses an emotion engine to recognize the user's facial expressions, voice, and heart rate, and estimates that the user is "refreshed and active." This emotion data is also sent to the server.
[1246] 3. The server aggregates this data and assumes that the user is "refreshed and active."
[1247] 4. The server generates a playlist containing up-tempo pop songs that fit this state of mind and sends it to the device.
[1248] 5. The user starts listening to the playlist and requests a lighter song.
[1249] 6. The server receives this request, updates the playlist to include more upbeat songs, and sends it back to the device.
[1250] 7. The user can enjoy the updated playlist.
[1251] In this way, by utilizing comprehensive data, including the user's emotional state, it is possible to provide a more advanced and personalized music experience.
[1252] The processing flow will be explained below.
[1253] Step 1:
[1254] The device uses the GPS function to obtain geographical data to obtain the user's current location information, and then sends the obtained location information to the server.
[1255] Step 2:
[1256] The device collects data about the user's daily activities, such as payment history, calendar schedule, and fitness tracker data, and transmits this data to a server.
[1257] Step 3:
[1258] The device collects the user's music preference data, specifically the history of songs played in the past and the user's ratings of songs, and transmits this data to the server.
[1259] Step 4:
[1260] The device uses an emotion engine to recognize the user's emotions. The emotion engine analyzes the user's voice, facial expressions, and biometric information (heart rate, body temperature, etc.) obtained from the input device. The resulting emotion data is sent to the server.
[1261] Step 5:
[1262] The server receives location information, behavioral data, music preference data, and emotional data and integrates them to create a user profile that comprehensively describes the user's current situation.
[1263] Step 6:
[1264] The server estimates the user's current state of mind based on the user profile. For example, if the user works out at the gym in the morning and then spends time at a cafe, the server estimates that the user feels refreshed and active.
[1265] Step 7:
[1266] The server selects the most suitable songs based on the estimated psychological state and music preference data, and generates a playlist based on the selected songs. This playlist includes song titles, artist names, playback links, etc.
[1267] Step 8:
[1268] The server transmits the generated playlist to the terminal, and the playlist data is stored on the user's terminal.
[1269] Step 9:
[1270] The user sends a request to the playlist being played, such as "more upbeat songs" or "more calm songs." This request is sent from the terminal to the server.
[1271] Step 10:
[1272] The server re-evaluates the playlist based on the user's request, adds new songs that meet the criteria, and sends the updated playlist back to the device.
[1273] Step 11:
[1274] The user receives the updated playlist and begins playback, allowing them to enjoy music that better suits their mood.
[1275] Through this series of processing steps, a highly personalized music experience can be provided based on the user's current emotions and behavior.
[1276] Example 2
[1277] Next, a description will be given of Example 2. In the following description, the data processing device 12 will be referred to as a "server" and the robot 414 will be referred to as a "terminal."
[1278] Conventional music recommendation systems often generate playlists based solely on a user's music preference data, without taking into account the user's real-time emotions or behavioral status. As a result, they are unable to suggest songs that fit the user's current psychological state, making it difficult to increase user satisfaction. To solve this problem, a music recommendation system that also takes into account the user's real-time emotions and behavioral data is needed.
[1279] The specific processing by the specific processing unit 290 of the data processing device 12 in the second embodiment is realized by the following means.
[1280] In this invention, the server includes: means for collecting current location information from the user's device; means for collecting behavioral data from the user's device; means for collecting music preference data from the user's device; means having an emotion analysis engine for collecting emotion data from the device; means for estimating the user's current psychological state by integrating the collected location information, behavioral data, music preference data, and emotion data; means for generating an optimal music playlist based on the estimated psychological state and music preference data; means for transmitting the generated playlist to the user's device; means for receiving a request from the user, reevaluating and updating the playlist, and means for transmitting the updated playlist to the user's device. This enables personalized music recommendations that take into account the user's real-time emotions and behavioral status.
[1281] "User's terminal" refers to electronic devices such as mobile terminals and wearable devices used by users.
[1282] "Location information" refers to data regarding the user's current location determined using GPS functions, etc.
[1283] "Behavioral data" refers to data that shows a user's daily activities, including payment history, calendar schedule information, and data from fitness trackers.
[1284] "Music preference data" refers to data that indicates a user's musical preferences, such as a history of songs that the user has played in the past and ratings of songs.
[1285] "Emotion data" refers to data obtained by an emotion analysis engine from a user's facial expressions, voice, and biometric information (e.g., heart rate and body temperature).
[1286] An "emotion analysis engine" refers to a software or hardware system for analyzing a user's emotional state.
[1287] "Mood" refers to information that indicates a user's current emotional or mental state.
[1288] A "playlist" refers to a collection of songs selected based on specific criteria.
[1289] A "request" refers to information indicating a specific instruction or request from a user.
[1290] A "server" refers to a computer system that communicates with user terminals via a network and processes and stores data.
[1291] The present invention relates to a music recommendation system that takes into account real-time user emotion and behavior data. Specific embodiments of the system are described below.
[1292] To implement this system, a user terminal, a server, an emotion analysis engine, and a communication system that links these are required. The following hardware and software are used as the main components.
[1293] 1. User devices: Mobile devices such as smartphones and wearable devices
[1294] 2. Server: High-performance computer system
[1295] 3. Sentiment analysis engine: Speech recognition software (e.g., speech recognition API, face recognition API)
[1296] Data collection
[1297] The device first collects the user's current location information using its GPS function. This data is periodically recorded and sent to a server. It also uses a fitness tracker to obtain exercise data (e.g., heart rate, number of steps), and collects behavioral data from the device's calendar and payment history. Furthermore, it collects the user's music preference data, such as the history of songs played and their ratings, and sends this data to the server.
[1298] Emotion analysis
[1299] The device uses a built-in camera and microphone to transmit the user's facial expressions and voice to an emotion analysis engine. The emotion analysis engine analyzes this data and detects the user's emotional state (e.g., joy, sadness). For example, if the device's camera captures the user's smile and the microphone detects a joyful voice tone, this information is sent to the server as emotion data.
[1300] Data integration and psychological state estimation
[1301] The server receives location information, behavioral data, music preference data, and emotional data from the device. It integrates these data to create a user profile and uses machine learning algorithms to estimate the user's current psychological state. For example, if a user exercises at the gym in the morning and then has coffee at a cafe, the server estimates the user's psychological state as "refreshed and active."
[1302] Playlist Generation
[1303] The server generates an optimal music playlist based on the estimated psychological state and the user's music preference data. The playlist includes song titles, artist names, and play links. The server then sends the playlist to the user's device, allowing the user to play songs based on the playlist.
[1304] Handling the request
[1305] Users can make requests for songs such as "more upbeat" or "more calming" to a playlist that is currently playing. This request is sent to the server via the device, and the server reevaluates the playlist based on the request and adds new songs. The updated playlist is then sent back to the device, where the user can enjoy it.
[1306] Specific examples
[1307] Consider a scenario where a user goes to the gym at 7am and then has coffee at a cafe at 9am. In this scenario, the following happens:
[1308] 1. The device activates the GPS function and records that the user is at the gym.
[1309] 2. Your device receives your heart rate and exercise data from your fitness tracker and records the information.
[1310] 3. The terminal obtains the payment history at the cafe and records that information.
[1311] 4. The device uses its built-in camera and microphone to transmit the user's facial expressions of delight and tone of voice to the emotion analysis engine.
[1312] 5. The server receives and integrates the location information, behavioral data, music preference data, and emotional data.
[1313] 6. The server uses a machine learning algorithm to estimate the user's psychological state as "refreshed and active."
[1314] 7. The server generates a playlist containing up-tempo pop songs based on the estimated psychological state and sends it to the device.
[1315] 8. While listening to a playlist, the user requests "something more upbeat."
[1316] 9. The server receives the request, updates the playlist to include more upbeat songs, and sends it back to the device.
[1317] 10. Users can enjoy the updated playlist.
[1318] Example prompt for a generative AI model:
[1319] "You have behavioral data that shows that a user works out at the gym at 7 AM and drinks coffee at a cafe at 9 AM. Furthermore, analysis by a sentiment analysis engine suggests that the user is in a psychological state of 'refreshed and active.' Based on this data, please generate a playlist to recommend to the user."
[1320] In this way, it is possible to provide an advanced music recommendation system that also takes into account the user's real-time emotions and behavior.
[1321] The flow of the identification process in the second embodiment will be described with reference to FIG.
[1322] Step 1: Data collection phase
[1323] The device collects the user's current location information. Input includes location information obtained from the device's GPS function. This data is recorded with a timestamp and sent to the server. For example, if the user is at the gym, the location coordinates and time are recorded. The device also collects exercise data from a fitness tracker. Input includes heart rate and step count data. This data is also recorded and sent to the server. The device also collects the user's daily behavior data (payment history, calendar schedule information) and music preference data (past playback history and ratings) and sends them to the server. Input is obtained from the smartphone, payment system, calendar app, and music app.
[1324] Step 2: Emotion Recognition Phase
[1325] The device uses a built-in camera and microphone to send the user's facial expressions and voice to an emotion analysis engine. Inputs include facial expression data captured by the camera, voice data picked up by the microphone, and biometric data (heart rate, body temperature, etc.). The emotion analysis engine analyzes this data and determines the user's emotional state. Emotional data such as "happiness" or "sadness" is generated as output and sent to the server. Specifically, the device analyzes the user's smile, which indicates joy, and the tone of their voice, which indicates happiness, and records this as emotion data.
[1326] Step 3: Data integration phase
[1327] The server receives location information, behavioral data, music preference data, and emotional data sent from the device. These multiple data sets serve as input. The server integrates them and performs the necessary processing to create a user profile. This profile is then processed using machine learning algorithms to generate an output that analyzes the interrelationships between each piece of data. For example, if a user works out at the gym in the morning and then has coffee at a cafe, this information can be used to estimate a psychological state of "refreshed and active."
[1328] Step 4: Mental state estimation phase
[1329] The server runs a machine learning algorithm on the integrated data to estimate the user's current state of mind. The input is the user profile generated in the data integration phase. The algorithm analyzes this data and estimates the user's state of mind (e.g., "refreshed and active") as output. Specifically, the server analyzes the correlations between each data set based on time of day, behavioral patterns, and emotional changes.
[1330] Step 5: Playlist generation phase
[1331] The server selects optimal songs based on the estimated psychological state and music preference data. The inputs are the estimated psychological state data and music preference data. Based on these data, the server filters the optimal songs from the music database and generates a playlist. The output is a playlist containing song titles, artist names, play links, etc. Specifically, up-tempo pop songs are selected for users in a "refreshed and active" psychological state. This playlist is then sent from the server to the device.
[1332] Step 6: Request handling phase
[1333] A user can make requests to a playlist that is currently being played. The input is a request sent by the user through the terminal (e.g., "More upbeat songs," "More calm songs"). The server receives this request and re-evaluates the playlist based on the input request. The server adds newly selected songs to the playlist, and an updated playlist is generated as output. Specifically, the server adds songs from its song database that fit the request, and sends the updated playlist back to the terminal.
[1334] (Application example 2)
[1335] Next, a description will be given of Application Example 2. In the following description, the data processing device 12 will be referred to as a "server" and the robot 414 will be referred to as a "terminal."
[1336] Conventional music recommendation systems are limited to generating simple playlists based on a user's musical preferences, and have difficulty providing a personalized music experience that takes into account the user's emotions and psychological state. This has resulted in problems such as not being able to provide music that matches the user's current mood or emotion, leading to a decrease in satisfaction. Furthermore, the system lacks the ability to flexibly update playlists in response to user requests. A system that can solve these issues is needed.
[1337] The specific processing by the specific processing unit 290 of the data processing device 12 in the application example 2 is realized by the following means.
[1338] In this invention, the server includes means for recognizing a user's emotions by performing voice analysis and facial expression analysis, means for transmitting data based on the user's emotions to the server and estimating the user's mood, and means for generating an optimal music playlist based on the estimated mood, thereby enabling a more personalized music experience according to the user's emotions and psychological state.
[1339] A "user's terminal" is an electronic device that a user uses on a daily basis, and includes smartphones, tablets, smartwatches, etc.
[1340] "Current location information" refers to geographical location data collected from the user's device and is obtained using GPS.
[1341] "Behavioral data" is information about a user's daily activities, including schedules, payment history, fitness tracker data, etc.
[1342] "Music preference data" is information such as the user's favorite music genres and artists, playback history, and ratings.
[1343] "Emotion" indicates the emotional state the user is feeling at that time, and is based on analyzed biometric information, voice, facial expressions, etc.
[1344] "Means for recognizing emotions" refers to technologies and algorithms that analyze a user's voice, facial expressions, and biometric information in order to identify the user's emotional state.
[1345] The "means for estimating mood" is a method for estimating a user's current psychological state using a computer algorithm based on collected user behavioral data, location information, and emotional data.
[1346] A "playlist" is a list of music for a specific purpose or theme, including song titles, artist names, and playback links.
[1347] The "means for receiving a request" is an interface through which the system receives requests or wishes from the user, such as voice input or touch input.
[1348] The "means for reevaluating and updating a playlist" refers to a means for reviewing a playlist that has already been generated and adding or changing new songs, etc., based on a request received from a user.
[1349] A "generative AI model" is an artificial intelligence model used to generate new music playlists based on collected data, and includes machine learning algorithms.
[1350] A "prompt sentence" is a text sentence that serves as input to a generative AI model to perform a specific task.
[1351] To implement this invention, a server, a terminal, an emotion engine, and a communication system that links them together are required. A specific embodiment of this system will be described below.
[1352] Data Collection Phase
[1353] The device collects the user's current location information. Using its GPS function, it periodically records the user's current location and sends it to the server. The device also collects the user's daily behavioral data. This behavioral data includes payment history, calendar schedule information, and fitness tracker data. The device also collects the user's music preference data, including the history of music played in the past and the ratings the user has given to songs. This data is also sent to the server.
[1354] Emotion Recognition Phase
[1355] The device recognizes the user's emotions through an emotion engine. The emotion engine analyzes the user's voice, facial expressions, and biometric information (heart rate, body temperature, etc.) obtained from input devices. The emotion data recognized by the emotion engine is sent to the server along with other collected data.
[1356] Mood estimation phase
[1357] The server receives location information, behavioral data, music preference data, and emotional data from the device and integrates them to create a user profile. Based on this profile, the server runs an algorithm to estimate the user's current psychological state. For example, if a user exercises at the gym in the morning and then has coffee at a cafe, these actions and emotions can be used to estimate the user's psychological state as "refreshed and active."
[1358] Playlist Generation Phase
[1359] The server selects the most suitable songs based on the estimated psychological state and music preference data. The generated playlist includes song titles, artist names, and playback links. This playlist is then sent to the device and provided to the user.
[1360] Request handling phase
[1361] Users can make requests for songs, such as "more upbeat" or "more calming," to a playlist that is currently playing. These requests are sent to the server via the device. The server reevaluates the playlist based on the received request and adds new songs. The updated playlist is then sent back to the device, where the user can listen to it.
[1362] Specific examples
[1363] Consider a scenario where a user goes to the gym at 7am and then has coffee at a cafe at 9am.
[1364] 1. The device records these activities as location and behavioral data and sends them to a server.
[1365] 2. The device uses an emotion engine to recognize the user's facial expressions, voice, and heart rate, and estimates that the user is "refreshed and active." This emotion data is also sent to the server.
[1366] 3. The server aggregates this data and assumes that the user is "refreshed and active."
[1367] 4. The server generates a playlist containing up-tempo pop songs that fit this state of mind and sends it to the device.
[1368] 5. The user starts listening to the playlist and requests a lighter song.
[1369] 6. The server receives this request, updates the playlist to include more upbeat songs, and sends it back to the device.
[1370] 7. The user can enjoy the updated playlist.
[1371] Prompt sentences to input to the generative AI model
[1372] Use the following prompt:
[1373] Based on the user's daily behavior data, such as working out at the gym at 7am and then drinking coffee at a cafe at 9am, estimate his current psychological state as "refreshed and active." His musical preferences are pop and electronic. Generate a playlist that is appropriate for this state.
[1374] This allows for a more personalized music experience based on the user's emotions and behavior.
[1375] The flow of the specific processing in the application example 2 will be described with reference to FIG.
[1376] Step 1:
[1377] The device collects the user's current location information. Using the GPS function, it periodically obtains the user's latitude and longitude and sends that location information to the server. The input is the device's GPS data, and the output is the location information data sent to the server.
[1378] Step 2:
[1379] The device collects the user's daily behavioral data, including payment history, calendar schedule information, and fitness tracker data, and sends it to a server. The input is this behavioral data, and the output is the integrated data sent to the server.
[1380] Step 3:
[1381] The device collects the user's music preference data, including the history of songs played and rating information, and sends it to the server. The input is the music playback history and rating information, and the output is the music preference data sent to the server.
[1382] Step 4:
[1383] The device recognizes the user's emotions through an emotion engine. The emotion engine analyzes the user's voice, facial expression, and biometric information (heart rate, body temperature, etc.) to recognize the user's emotional state and output it as data. This data is sent to the server. The input at this time is voice, facial expression, biometric information, etc., and the output is recognized emotional data.
[1384] Step 5:
[1385] The server receives location information, behavioral data, music preference data, and emotional data sent from the device and integrates them to create a user profile. The server's algorithm analyzes this data and estimates the user's current psychological state. The input is the integrated user data, and the output is an estimated psychological state.
[1386] Step 6:
[1387] The server generates an optimal music playlist based on the estimated psychological state and music preference data. Using a generative AI model, the server selects songs that best fit the user's psychological state and creates a playlist. The input is the psychological state and music preference data, and the output is the generated playlist.
[1388] Step 7:
[1389] The server sends the generated playlist to the user's device and provides it to the user. The playlist includes song titles, artist names, playback links, etc. The input is the generated playlist, and the output is the playlist sent to the device.
[1390] Step 8:
[1391] A user can make a request for a playlist that is currently being played. For example, a request for "more upbeat songs" is sent to the server via the terminal. The input is the user's request, and the output is the request data sent to the server.
[1392] Step 9:
[1393] The server reevaluates the playlist based on the received request and adds new songs. It uses a generative AI model to generate a new playlist and sends it to the device. The input is the user request and existing playlist data, and the output is the updated playlist.
[1394] Step 10:
[1395] The terminal plays the updated playlist sent from the server and provides it to the user, allowing the user to enjoy a music experience tailored to their requests. The input is the updated playlist data, and the output is the songs to be played.
[1396] The specific processing unit 290 transmits the result of the specific processing to the robot 414. In the robot 414, the control unit 46A causes the speaker 240 and the control target 443 to output the result of the specific processing. The microphone 238 acquires voice indicating a user input regarding the result of the specific processing. The control unit 46A transmits voice data indicating the user input acquired by the microphone 238 to the data processing device 12. In the data processing device 12, the specific processing unit 290 acquires the voice data.
[1397] The data generation model 58 is a so-called generative AI (Artificial Intelligence). An example of the data generation model 58 is ChatGPT (Internet Search<URL: https: / / openai.com / blog / chatgpt> ), Gemini (Internet search <url: https: gemini.google.com ?hl="ja">) and other generation AIs. The data generation model 58 is obtained by performing deep learning on a neural network. A prompt including an instruction is input to the data generation model 58, and inference data such as voice data indicating voice, text data indicating text, and image data indicating an image is also input. The data generation model 58 performs inference on the input inference data in accordance with the instruction indicated by the prompt, and outputs the inference result in a data format such as voice data and text data. Here, inference refers to, for example, analysis, classification, prediction, and / or summarization.
[1398] In the above embodiment, an example was given in which the specific processing is performed by the data processing device 12, but the technology of the present disclosure is not limited to this, and the specific processing may be performed by the robot 414.
[1399] The emotion identification model 59 as an emotion engine may determine the user's emotion according to a specific mapping. Specifically, the emotion identification model 59 may determine the user's emotion according to an emotion map (see FIG. 9), which is a specific mapping. Similarly, the emotion identification model 59 may determine the robot's emotion, and the identification processing unit 290 may perform identification processing using the robot's emotion.
[1400] FIG. 9 is a diagram illustrating an emotion map 400 on which multiple emotions are mapped. In the emotion map 400, emotions are arranged in concentric circles radiating from the center. Emotions closer to the center of the concentric circles are more primitive. Emotions representing states and actions arising from a state of mind are arranged on the outer edges of the concentric circles. The concept of emotion includes both affect and mental states. Emotions generally generated from reactions occurring in the brain are arranged on the left side of the concentric circles. Emotions generally induced by situational judgment are arranged on the right side of the concentric circles. Emotions generally generated from reactions occurring in the brain and induced by situational judgment are arranged on the upper and lower sides of the concentric circles. Furthermore, the emotion of "pleasure" is arranged on the upper side of the concentric circles, and the emotion of "discomfort" is arranged on the lower side. In this way, in the emotion map 400, multiple emotions are mapped based on the structure by which emotions are generated, and emotions that tend to occur simultaneously are mapped close to each other.
[1401] These emotions are distributed in the 3 o'clock direction on emotion map 400, and typically fluctuate between relief and anxiety. In the right half of emotion map 400, situational awareness dominates over internal sensations, resulting in a sense of calm.
[1402] The inside of emotion map 400 represents what is going on in the mind, and the outside of emotion map 400 represents behavior, so the further you go outside emotion map 400, the more visible the emotions become (the more they are expressed in behavior).
[1403] Human emotions are based on various balances, such as posture and blood sugar levels. When these balances deviate from the ideal, a state of discomfort is indicated, and when they approach the ideal, a state of pleasure is indicated. Emotions can also be created for robots, automobiles, and motorcycles, based on various balances, such as posture and remaining battery life. When these balances deviate from the ideal, a state of discomfort is indicated, and when they approach the ideal, a state of pleasure is indicated. An emotion map can be generated, for example, based on Dr. Mitsuyoshi's emotion map (Research on Voice Emotion Recognition and Emotional Brain Physiological Signal Analysis Systems, Tokushima University, Doctoral Dissertation: https: / / ci.nii.ac.jp / naid / 500000375379). The left half of the emotion map lists emotions belonging to the "reaction" domain, where sensation is dominant. The right half of the emotion map lists emotions belonging to the "situation" domain, where situational awareness is dominant.
[1404] The emotion map defines two emotions that promote learning. One is a negative emotion on the situation side, around the middle of "repentance" or "reflection." In other words, this occurs when the robot experiences negative emotions such as "I never want to feel this way again" or "I don't want to be scolded again." The other is a positive emotion on the response side, around "desire." In other words, this occurs when the robot experiences positive feelings such as "I want more" or "I want to know more."
[1405] The emotion identification model 59 inputs user input into a pre-trained neural network, obtains emotion values indicating each emotion shown in the emotion map 400, and determines the user's emotion. This neural network is pre-trained based on multiple pieces of training data that are combinations of user input and emotion values indicating each emotion shown in the emotion map 400. Furthermore, this neural network is trained so that emotions that are located close to each other have similar values, as in the emotion map 900 shown in FIG. 10. FIG. 10 shows an example in which multiple emotions, "relieved," "calm," and "reassuring," have similar emotion values.
[1406] The system according to the present disclosure has been described above mainly with respect to the functions of the data processing device 12, but the system according to the present disclosure is not necessarily implemented on a server. The system according to the present disclosure may be implemented as a general information processing system. The present disclosure may be implemented, for example, as a software program running on a personal computer or an application running on a smartphone, etc. The method according to the present disclosure may be provided to users in the form of SaaS (Software as a Service).
[1407] In the above embodiment, an example was given in which the specific processing is performed by one computer 22, but the technology of the present disclosure is not limited to this, and the specific processing may be distributed and performed by a plurality of computers including the computer 22. For example, the data generation model 58 may be provided in an external device of the data processing device 12, and data may be generated in the external device in accordance with input data.
[1408] In the above embodiment, an example in which the specific processing program 56 is stored in the storage 32 has been described, but the technology of the present disclosure is not limited to this. For example, the specific processing program 56 may be stored in a portable, computer-readable, non-transitory storage medium such as a USB (Universal Serial Bus) memory. The specific processing program 56 stored in the non-transitory storage medium is installed in the computer 22 of the data processing device 12. The processor 28 executes the specific processing in accordance with the specific processing program 56.
[1409] Alternatively, the specific processing program 56 may be stored in a storage device such as a server connected to the data processing device 12 via the network 54, and the specific processing program 56 may be downloaded and installed on the computer 22 in response to a request from the data processing device 12.
[1410] It is not necessary to store all of the specific processing program 56 in a storage device such as a server connected to the data processing device 12 via the network 54, or to store all of the specific processing program 56 in the storage 32; only a portion of the specific processing program 56 may be stored.
[1411] The hardware resource for executing a specific process can be any of the following processors: An example of a processor is a CPU, which is a general-purpose processor that functions as a hardware resource for executing a specific process by executing software, i.e., a program. Another example of a processor is a dedicated electrical circuit, such as an FPGA (Field-Programmable Gate Array), a PLD (Programmable Logic Device), or an ASIC (Application Specific Integrated Circuit), which is a processor with a circuit configuration designed specifically for executing a specific process. Each processor has built-in or connected memory, and each processor uses the memory to execute the specific process.
[1412] The hardware resource that executes the specific processing may be configured with one of these various processors, or may be configured with a combination of two or more processors of the same or different types (for example, a combination of multiple FPGAs, or a combination of a CPU and an FPGA). Also, the hardware resource that executes the specific processing may be a single processor.
[1413] As an example of a system configured with a single processor, first, one processor is configured by combining one or more CPUs and software, and this processor functions as a hardware resource that executes a specific process. Second, there is a system that uses a processor that realizes the functions of an entire system including multiple hardware resources that execute a specific process on a single IC chip, as typified by SoC (System-on-a-chip). In this way, a specific process is realized using one or more of the above-mentioned various processors as hardware resources.
[1414] Furthermore, the hardware structure of these various processors can be, more specifically, an electric circuit that combines circuit elements such as semiconductor devices. The specific processing described above is merely an example. Therefore, it goes without saying that unnecessary steps may be deleted, new steps may be added, or the processing order may be rearranged, without departing from the spirit of the invention.
[1415] The above-described description and illustrations are a detailed explanation of the parts related to the technology of the present disclosure and are merely an example of the technology of the present disclosure. For example, the above description of the configuration, functions, actions, and effects is an explanation of an example of the configuration, functions, actions, and effects of the parts related to the technology of the present disclosure. Therefore, it goes without saying that unnecessary parts may be deleted, new elements may be added, or replacements may be made to the above-described description and illustrations within the scope of the gist of the technology of the present disclosure. Furthermore, to avoid confusion and facilitate understanding of the parts related to the technology of the present disclosure, the above-described description and illustrations omit explanations of common technical knowledge that do not require particular explanation to enable the implementation of the technology of the present disclosure.
[1416] All publications, patent applications, and technical standards mentioned in this specification are herein incorporated by reference to the same extent as if each individual publication, patent application, or technical standard was specifically and individually indicated to be incorporated by reference.
[1417] The following is further disclosed regarding the above embodiment.
[1418] (Claim 1)
[1419] means for collecting current location information from a user's device;
[1420] A means for collecting behavioral data from a user's device;
[1421] means for collecting music preference data from a user's device;
[1422] A means for estimating the user's current psychological state by integrating the collected location information, behavioral data, and music preference data;
[1423] means for generating an optimal music playlist based on the estimated psychological state and music preference data;
[1424] means for transmitting the generated playlist to a user's terminal;
[1425] means for receiving requests from a user to reevaluate and update the playlist;
[1426] means for transmitting the updated playlist to the user's terminal;
[1427] A system including:
[1428] (Claim 2)
[1429] 2. The system according to claim 1, wherein information about the location of the user is analyzed based on the user's current location information.
[1430] (Claim 3)
[1431] 10. The system of claim 1, further comprising analyzing payment history and schedule data included in the user's behavioral data.
[1432] "Example 1"
[1433] (Claim 1)
[1434] means for collecting current location information from a user's information processing device;
[1435] means for collecting behavioral data from a user's information processing device;
[1436] means for collecting music preference data from a user's information processing device;
[1437] using a generative AI model that integrates collected location, behavioral, and music preference data to estimate the user's current psychological state; and
[1438] means for generating an optimal music playlist based on the estimated psychological state and music preference data;
[1439] means for transmitting the generated playlist to a user's information processing device;
[1440] means for receiving requests from a user to reevaluate and update the playlist;
[1441] means for transmitting the updated playlist to the user's information processing device;
[1442] A system including:
[1443] (Claim 2)
[1444] 2. The system according to claim 1, wherein information about the location of the user is analyzed based on the user's current location information.
[1445] (Claim 3)
[1446] 2. The system according to claim 1, wherein the system analyzes financial transaction history and schedule information contained in the user's behavioral data.
[1447] "Application Example 1"
[1448] (Claim 1)
[1449] means for collecting current location information from a user's device;
[1450] A means for collecting behavioral data from a user's device;
[1451] means for collecting music preference data from a user's device;
[1452] A means for estimating the user's current psychological state by integrating the collected location information, behavioral data, and music preference data;
[1453] means for generating an optimal music playlist based on the estimated psychological state and music preference data;
[1454] means for transmitting the generated playlist to a user's terminal;
[1455] means for receiving requests from a user to reevaluate and update the playlist;
[1456] means for transmitting the updated playlist to the user's terminal;
[1457] A system including means for providing a music playlist that best suits a user's mood in an autonomous vehicle.
[1458] (Claim 2)
[1459] 2. The system according to claim 1, wherein information about the location of the user is analyzed based on the user's current location information.
[1460] (Claim 3)
[1461] 10. The system of claim 1, further comprising analyzing payment history and schedule data included in the user's behavioral data.
[1462] "Example 2: Combining Emotion Engines"
[1463] (Claim 1)
[1464] means for collecting current location information from a user's device;
[1465] A means for collecting behavioral data from a user's device;
[1466] means for collecting music preference data from a user's device;
[1467] means having an emotion analysis engine for collecting emotion data from a terminal;
[1468] a means for estimating a user's current psychological state by integrating the collected location information, behavioral data, music preference data, and emotional data;
[1469] means for generating an optimal music playlist based on the estimated psychological state and music preference data;
[1470] means for transmitting the generated playlist to a user's terminal;
[1471] means for receiving requests from a user to reevaluate and update the playlist;
[1472] means for transmitting the updated playlist to the user's terminal;
[1473] A system including:
[1474] (Claim 2)
[1475] 2. The system according to claim 1, wherein information about the location of the user is analyzed based on the user's current location information.
[1476] (Claim 3)
[1477] 10. The system of claim 1, further comprising analyzing payment history and schedule data included in the user's behavioral data.
[1478] "Application example 2 when combining emotion engines"
[1479] (Claim 1)
[1480] means for collecting current location information from a user's device;
[1481] A means for collecting behavioral data from a user's device;
[1482] means for collecting music preference data from a user's device;
[1483] A means for estimating the user's current psychological state by integrating the collected location information, behavioral data, and music preference data;
[1484] means for generating an optimal music playlist based on the estimated psychological state and music preference data;
[1485] means for transmitting the generated playlist to a user's terminal;
[1486] means for receiving requests from a user to reevaluate and update the playlist;
[1487] means for transmitting the updated playlist to the user's terminal;
[1488] A means for recognizing a user's emotions by performing voice analysis and facial expression analysis;
[1489] means for transmitting data based on the user's emotions to a server and estimating the user's mood at the server;
[1490] means for generating an optimal music playlist based on the estimated mood;
[1491] means for transmitting the generated playlist to a user's terminal and updating it upon request;
[1492] A system including:
[1493] (Claim 2)
[1494] 2. The system according to claim 1, wherein information about the location of the user is analyzed based on the user's current location information.
[1495] (Claim 3)
[1496] 10. The system of claim 1, further comprising analyzing payment history and schedule data included in the user's behavioral data. [Explanation of symbols]
[1497] 10, 210, 310, 410 Data Processing Systems 12 Data Processing Device 14 Smart Devices 214 Smart Glasses 314 Headset-type terminal 414 Robot< / url:> < / url:> < / url:> < / url:>
Claims
1. means for collecting current location information from a user's device; A means for collecting behavioral data from a user's device; means for collecting music preference data from a user's device; A means for estimating the user's current psychological state by integrating the collected location information, behavioral data, and music preference data; means for generating an optimal music playlist based on the estimated psychological state and music preference data; means for transmitting the generated playlist to a user's terminal; means for receiving requests from a user to reevaluate and update the playlist; means for transmitting the updated playlist to the user's terminal; A system including:
2. 2. The system according to claim 1, wherein information about the location of the user is analyzed based on information about the user's current location.
3. The system of claim 1 , further comprising analyzing payment history and schedule data included in the user behavior data.
Citation Information
Patent Citations
Persona chatbot control method and system
JP2022180282A