system
The system addresses the challenge of unsuitable music selection by using user location, time, and emotional state data to generate personalized music experiences, enhancing user engagement and enjoyment.
Patent Information
- Authority / Receiving Office
- JP · JP
- Patent Type
- Applications
- Current Assignee / Owner
- SOFTBANK GROUP CORP
- Filing Date
- 2024-10-10
- Publication Date
- 2026-04-22
AI Technical Summary
Existing music streaming systems fail to provide personalized music experiences that consider a user's location, time, and emotional state, leading to monotonous and unsuitable music selection.
A system that collects and analyzes user location information, location information, music listening history, and time information, and generates new music recommendations, and generates new music based on region and time, and provides users with a personalized music experience.
The system provides users with music tailored to their current location, time, and emotional state, enhancing the music experience by eliminating repetitive selections and improving user engagement.
Smart Images

Figure 2026068422000001_ABST
Abstract
Description
Technical Field
[0001] The technology of the present disclosure relates to a system.
Background Art
[0002] Patent Document 1 discloses a method for controlling a persona chatbot, which is performed by at least one processor and includes steps of receiving a user utterance, adding the user utterance to a prompt including an instruction sentence related to an explanation of a chatbot character, encoding the prompt, and inputting the encoded prompt into a language model to generate a chatbot utterance in response to the user utterance.
Prior Art Documents
Patent Documents
[0003]
Patent Document 1
Summary of the Invention
Problems to be Solved by the Invention
[0004] There are problems such as the routinization of music selection in music streaming and the selection of music that is not suitable for the individual location, time, and situation of the user. As a result, in the scene where music is provided, the user may not be able to enjoy the maximum music experience. Furthermore, it is difficult to select music suitable for the user's current emotion and location in real time, and there is a concern that the user experience will decline.
Means for Solving the Problems
[0005] This invention provides a system that selects the most suitable music for a user by collecting and analyzing the user's location information, music listening history, and regional music rankings, and further combining this with current time information. Furthermore, by generating new music based on region and time, it can provide users with an even more personalized music experience. This enables music playback tailored to the user's situation and eliminates the monotony of repetitive music selection.
[0006] "User" refers to an individual or legal entity that uses a music streaming service.
[0007] "Location information" refers to data that represents the user's current geographical location, including latitude and longitude.
[0008] "Music listening history" is a record of what kind of music and artists a user has played in the past.
[0009] "Regional music rankings" are data that ranks popular songs and artists within a specific geographical area.
[0010] "Analysis" is the process of examining collected data in detail to derive specific patterns or trends.
[0011] "New music" refers to songs or music content generated by the system to suit the user's current context.
[0012] "Current time information" refers to data indicating the date, time, and time zone at the time the user is playing music.
[0013] A "personalized music experience" refers to providing music that is customized to each user's preferences, circumstances, and location. [Brief explanation of the drawing]
[0014] [Figure 1]It is a conceptual diagram showing an example of the configuration of a data processing system according to the first embodiment. [Figure 2] It is a conceptual diagram showing an example of the main functions of a data processing device and a smart device according to the first embodiment. [Figure 3] It is a conceptual diagram showing an example of the configuration of a data processing system according to the second embodiment. [Figure 4] It is a conceptual diagram showing an example of the main functions of a data processing device and smart glasses according to the second embodiment. [Figure 5] It is a conceptual diagram showing an example of the configuration of a data processing system according to the third embodiment. [Figure 6] It is a conceptual diagram showing an example of the main functions of a data processing device and a headset-type terminal according to the third embodiment. [Figure 7] It is a conceptual diagram showing an example of the configuration of a data processing system according to the fourth embodiment. [Figure 8] It is a conceptual diagram showing an example of the main functions of a data processing device and a robot according to the fourth embodiment. [Figure 9] It shows an emotion map to which a plurality of emotions are mapped. [Figure 10] It shows an emotion map to which a plurality of emotions are mapped. [Figure 11] It is a sequence diagram showing the processing flow of the data processing system in Example 1. [Figure 12] It is a sequence diagram showing the processing flow of the data processing system in Application Example 1. [Figure 13] It is a sequence diagram showing the processing flow of the data processing system in Example 2 when an emotion engine is combined. [Figure 14] It is a sequence diagram showing the processing flow of the data processing system in Application Example 2 when an emotion engine is combined.
MODE FOR CARRYING OUT THE INVENTION
[0015] Hereinafter, an example of an embodiment of a system according to the technology of the present disclosure will be described with reference to the accompanying drawings.
[0016] First, the terms used in the following description will be explained.
[0017] In the following embodiments, the labeled processor (hereinafter simply referred to as "processor") may be a single arithmetic unit or a combination of multiple arithmetic units. Also, the processor may be one type of arithmetic unit or a combination of multiple types of arithmetic units. Examples of arithmetic units include a CPU (Central Processing Unit), a GPU (Graphics Processing Unit), a GPGPU (General-Purpose computing on Graphics Processing Units), an APU (Accelerated Processing Unit), and the like.
[0018] In the following embodiments, the labeled RAM (Random Access Memory) is a memory in which information is temporarily stored and is used as a work memory by the processor.
[0019] In the following embodiments, the labeled storage is one or more non-volatile storage devices that store various programs and various parameters, etc. Examples of non-volatile storage devices include flash memory (SSD (Solid State Drive)), magnetic disks (e.g., hard disks), or magnetic tapes, etc.
[0020] In the following embodiments, the labeled communication I / F (Interface) is an interface including a communication processor and an antenna, etc. The communication I / F controls communication between multiple computers. Examples of communication standards applied to the communication I / F include wireless communication standards including 5G (5th Generation Mobile Communication System), Wi-Fi (registered trademark), or Bluetooth (registered trademark), etc.
[0021] In the following embodiments, "A and / or B" is synonymous with "at least one of A and B." That is, "A and / or B" means that it may be A alone, or B alone, or a combination of A and B. Furthermore, in this specification, the same concept as "A and / or B" applies when expressing three or more things linked by "and / or."
[0022] [First Embodiment]
[0023] Figure 1 shows an example of the configuration of the data processing system 10 according to the first embodiment.
[0024] As shown in Figure 1, the data processing system 10 includes a data processing device 12 and a smart device 14. An example of the data processing device 12 is a server.
[0025] The data processing device 12 comprises a computer 22, a database 24, and a communication interface 26. The computer 22 is an example of a "computer" related to the technology of this disclosure. The computer 22 comprises a processor 28, RAM 30, and storage 32. The processor 28, RAM 30, and storage 32 are connected to a bus 34. The database 24 and the communication interface 26 are also connected to the bus 34. The communication interface 26 is connected to a network 54. An example of the network 54 is a WAN (Wide Area Network) and / or a LAN (Local Area Network).
[0026] The smart device 14 comprises a computer 36, a reception device 38, an output device 40, a camera 42, and a communication interface 44. The computer 36 comprises a processor 46, RAM 48, and storage 50. The processor 46, RAM 48, and storage 50 are connected to a bus 52. The reception device 38, output device 40, and camera 42 are also connected to the bus 52.
[0027] The reception device 38 is equipped with a touch panel 38A and a microphone 38B, etc., and receives user input. The touch panel 38A receives user input by detecting contact with an object (e.g., a pen or finger). The microphone 38B receives user input by detecting the user's voice. The control unit 46A transmits data indicating the user input received by the touch panel 38A and microphone 38B to the data processing device 12. In the data processing device 12, the specific processing unit 290 acquires the data indicating the user input.
[0028] The output device 40 includes a display 40A and a speaker 40B, and presents data to the user 20 by outputting the data in a form perceptible to the user 20 (e.g., audio and / or text). The display 40A displays visible information such as text and images according to instructions from the processor 46. The speaker 40B outputs audio according to instructions from the processor 46. The camera 42 is a small digital camera equipped with an optical system such as a lens, aperture, and shutter, and an image sensor such as a CMOS (Complementary Metal-Oxide-Semiconductor) image sensor or a CCD (Charge Coupled Device) image sensor.
[0029] Communication interface 44 is connected to network 54. Communication interfaces 44 and 26 are responsible for the exchange of various types of information between processor 46 and processor 28 via network 54.
[0030] Figure 2 shows an example of the main functions of the data processing device 12 and the smart device 14.
[0031] As shown in Figure 2, in the data processing device 12, specific processing is performed by the processor 28. A specific processing program 56 is stored in the storage 32. The specific processing program 56 is an example of a "program" related to the technology of this disclosure. The processor 28 reads the specific processing program 56 from the storage 32 and executes the read specific processing program 56 on the RAM 30. The specific processing is realized by the processor 28 operating as a specific processing unit 290 according to the specific processing program 56 executed on the RAM 30.
[0032] The storage 32 stores the data generation model 58 and the emotion identification model 59. The data generation model 58 and the emotion identification model 59 are used by the identification processing unit 290.
[0033] In the smart device 14, the processor 46 performs the reception output processing. The storage 50 stores the reception output program 60. The reception output program 60 is used in conjunction with a specific processing program 56 by the data processing system 10. The processor 46 reads the reception output program 60 from the storage 50 and executes the read reception output program 60 on the RAM 48. The reception output processing is realized by the processor 46 operating as a control unit 46A according to the reception output program 60 executed on the RAM 48.
[0034] Next, the specific processing performed by the specific processing unit 290 of the data processing device 12 will be described. In the following description, the data processing device 12 will be referred to as the "server" and the smart device 14 as the "terminal".
[0035] The system of the present invention collects and analyzes various data in order to provide users with a more suitable music experience. Specific embodiments are described below.
[0036] First, the device uses its GPS function to obtain the user's location information. This allows it to send precise latitude and longitude to the server. Additionally, it retrieves music listening history from the streaming application on the device and sends it to the server as well.
[0037] The server retrieves regional music ranking data from external sources based on location information and listening history received from the device. This allows it to identify currently popular songs in a specific area. Furthermore, the server analyzes user preferences and past trends based on this data.
[0038] Next, the server considers the user's current time information and performs the music selection process. This is made possible by using an AI algorithm. Specifically, it integrates the user's past listening habits, current location information, and time of day data to select or generate the most suitable music. The generated music and playlists are created live to match the user's situation at that moment.
[0039] The server then sends recommended music data and generated songs to the device. The device receives this and displays it to the user to support music playback. The user can select a playlist from the presented options and play it.
[0040] The following describes the processing flow.
[0041] Step 1:
[0042] The device uses its GPS function to obtain the user's location information. The obtained location information is sent to the server as latitude and longitude data.
[0043] Step 2:
[0044] The device collects the user's music listening history from music streaming applications and sends this data to a server.
[0045] Step 3:
[0046] The server identifies the user's area based on the received location information. It then retrieves regional music ranking data related to this area from an external source.
[0047] Step 4:
[0048] The server analyzes the received music listening history to understand the user's preferences and past playback trends.
[0049] Step 5:
[0050] The server considers music trends based on the current time of day and integrates the acquired data to select songs and playlists suitable for the user. An AI algorithm carries out this selection process.
[0051] Step 6:
[0052] The server uses AI to generate new music tailored to the user's current situation.
[0053] Step 7:
[0054] The server sends the selected playlist and generated music to the device.
[0055] Step 8:
[0056] The device displays the received music information on the user interface, making it easy for users to access.
[0057] Step 9:
[0058] The user selects a playlist from the suggestions and plays the music. The feedback received during this process is sent back to the server to further refine future music recommendations.
[0059] (Example 1)
[0060] Next, we will describe Example 1. In the following description, the data processing device 12 will be referred to as the "server," and the smart device 14 will be referred to as the "terminal."
[0061] Traditional systems have limitations in providing users with a more personalized and context-appropriate music experience. It's necessary to generate appropriate music in real time using AI technology, in addition to user location information, music listening history, and local music rankings. Furthermore, there's a demand for a richer music experience by appropriately selecting music according to the time of day.
[0062] The identification process performed by the identification processing unit 290 of the data processing device 12 in Example 1 is realized by the following means.
[0063] In this invention, the server includes means for acquiring the user's location information, means for acquiring the user's music listening history, and means for generating music using an AI algorithm. This makes it possible to provide personalized music in real time based on the user's location information, past listening history, and current time of day.
[0064] "Means for obtaining user location information" refers to a function that uses mobile devices or other devices to obtain geographical data indicating the user's current location.
[0065] "Means for obtaining a user's music listening history" refers to a function that collects historical information about music that a user has played in the past from a digital device.
[0066] "Methods for obtaining location-based music rankings by region" refers to a function that obtains data from an external source to rank popular songs in a specific region.
[0067] "A means of analyzing and selecting music suitable for the user" refers to a function that selects songs that match the user's preferences and activities based on acquired data.
[0068] "Methods for generating music using AI algorithms" refers to a function that uses artificial intelligence technology to automatically create music tailored to the user's situation.
[0069] "Means of providing selected or generated music to users" refers to functions for delivering selected or generated music to users in real time.
[0070] This invention is a system that provides users with a personalized music experience. This system aggregates the user's location information, music listening history, and time information, and uses AI technology to select and generate music in real time. In specific embodiments, it primarily operates as follows:
[0071] First, the device uses a GPS-equipped device to obtain the user's location information. This location information is sent to the server as latitude and longitude data. The device also collects the user's music listening history from digital streaming applications and sends this to the server as well. Furthermore, the device uses the system time to obtain the current time information.
[0072] The server then receives this data and retrieves regional music rankings from external music databases and APIs. This uses publicly available digital ranking data and APIs. For example, APIs from commercially available music information services are used. The server analyzes the collected information using an AI algorithm to select or generate the most suitable music for the user. A generative AI model is used for this AI algorithm. The prompt given to the model is an instruction such as, "Based on this user's current location and time, please recommend relaxing music."
[0073] The server sends selected or generated music data to the terminal, which receives it and displays it visually to the user. The user then plays music based on the presented playlist. For example, if the user is in a park in the afternoon, the terminal sends its location and time data, and the server selects recommended songs with a relaxing effect.
[0074] In this way, a music experience optimized for the user's situation is provided. This system makes it easy for users to enjoy music that suits their lifestyle and environment.
[0075] The flow of the specific processing in Example 1 will be explained using Figure 11.
[0076] Step 1:
[0077] The terminal obtains location information using the GPS function built into the user's device. This location information is provided in the form of latitude and longitude. The input is the device's current location, and based on this, specific geographical information is generated and sent to the server.
[0078] Step 2:
[0079] The device retrieves the user's music listening history from the streaming application. This history includes information such as the title of the played song, the artist's name, and the number of plays. This data is entered as a log on the device, organized, and then output to the server.
[0080] Step 3:
[0081] The server receives location information and music listening history from the terminal and retrieves regional music ranking data from external music databases and APIs. In this process, regional information is used as input, ranking data is searched based on it, and the retrieved information is imported into the server.
[0082] Step 4:
[0083] The server integrates the acquired information and uses an AI algorithm to select or generate music suitable for the user. At this stage, a generative AI model is used, and the prompt "Please recommend relaxing music based on this user's current location and time" is input, and the optimal music data is output.
[0084] Step 5:
[0085] The server sends selected or generated music data to the terminal. This transmission uses digital communication, with music data generated by an AI algorithm as input and the data sent to the terminal as output.
[0086] Step 6:
[0087] The terminal displays the received music data to the user and prepares for playback. The input is music data sent from the server, which is displayed on the terminal's user interface. The user then starts playing specific songs through the presented playlist.
[0088] This series of processing steps enables a system that provides users with the most suitable music in real time.
[0089] (Application Example 1)
[0090] Next, we will explain Application Example 1. In the following explanation, the data processing device 12 will be referred to as the "server," and the smart device 14 will be referred to as the "terminal."
[0091] While modern content distribution services allow users to easily enjoy a wide variety of music, automatically providing the optimal music for each individual user remains challenging. In particular, customizing music in real time based on a user's location and time of day presents a significant technical challenge. Therefore, there is a need for a system that dynamically generates and delivers music playlists, taking into account the user's location, time, and music listening history.
[0092] The specific processing performed by the specific processing unit 290 of the data processing device 12 in Application Example 1 is realized by the following means.
[0093] In this invention, the server includes means for acquiring the user's location information, means for acquiring the user's music listening history, and means for dynamically generating a music playlist considering the user's movement. This makes it possible to provide a customized music experience for each user.
[0094] A "user" is an individual who uses the system to experience music.
[0095] "Means of acquiring location information" refers to technical elements used to determine the user's current location, and usually refers to GPS data, etc.
[0096] "Means for obtaining music listening history" refers to the technical elements used to record and utilize information about songs and artists that users have listened to in the past.
[0097] "Means of obtaining regional music rankings" refers to the technical elements used to obtain data indicating the popularity of music in a specific geographical area.
[0098] "Methods for analyzing and selecting music suitable for the user" refer to algorithms and technical elements that analyze acquired data and select the music that best suits the user.
[0099] "Means of providing music to users" refers to the technical elements for delivering selected music to users in a playable format.
[0100] "Means for dynamically generating music playlists while considering movement circumstances" refers to the technical elements for creating a suitable music playlist each time, depending on the user's movement situation.
[0101] This music delivery system requires specific hardware and software configurations to provide users with a dynamic music experience. First, it accurately obtains the user's location information using the smartphone's GPS function. This sends the current location data to the server. The music streaming application also records the user's past music listening history and provides this information to the server. Based on the received location information and music listening history, the server accesses a regional music ranking database and retrieves data on popular songs.
[0102] Furthermore, the server uses AI algorithms, such as TENSORFLOW® and PyTorch, to analyze user preferences. It takes into account past listening trends, local music rankings, and time-of-day data to dynamically generate music playlists in real time. This process utilizes a generative AI model. The generated music playlists are then sent to the user's smartphone and provided in a playable format.
[0103] For example, if a user uses this system while in a moving car, their current location can be communicated, and popular songs in the area they are traveling in can be suggested. This feature helps users discover new music while simultaneously providing a music experience tailored to each user's environment.
[0104] Examples of prompt messages are as follows:
[0105] "I'm a user currently in New York who has listened to a lot of jazz in the past. Please suggest some music suitable for a morning walk."
[0106] In this way, the system can provide users with the optimal music experience in real time.
[0107] The flow of a specific process in Application Example 1 will be explained using Figure 12.
[0108] Step 1:
[0109] The device acquires the user's location information. Using the smartphone's GPS sensor, it obtains the current latitude and longitude and sends this data to the server. The input is location data obtained from GPS, and the output is location information sent to the server.
[0110] Step 2:
[0111] The device retrieves the user's music listening history from the music streaming application. It collects listening history data via the application API and sends it to the server. The input is the listening history, and the output is the listening history data sent to the server.
[0112] Step 3:
[0113] The server retrieves regional music rankings from an existing database based on the received location information. It uses the location information to search for the corresponding regional data and retrieves the ranking data. The input is location information, and the output is regional music ranking data.
[0114] Step 4:
[0115] The server uses an AI model to analyze music listening history, regional music ranking data, and time information. This analysis helps understand the user's musical preferences and tendencies. The inputs are listening history, ranking data, and time information, and the output is the analyzed pattern of the user's preferences.
[0116] Step 5:
[0117] The server uses a generative AI model to generate music playlists tailored to the user's preferences in real time. It integrates past preferences with the current context and uses prompts to determine the optimal set of songs for the AI model. A dynamically generated playlist is obtained at this stage.
[0118] Step 6:
[0119] The server sends the generated music playlist to the terminal. The terminal then provides this to the user in a playable format. The input is the generated playlist, and the output is delivered to the user as playable music data.
[0120] Furthermore, an emotion engine that estimates the user's emotions may be incorporated. That is, the identification processing unit 290 may use the emotion identification model 59 to estimate the user's emotions and perform identification processing using the user's emotions.
[0121] This invention is a system for providing a personalized music experience based on the user's emotional state. This system incorporates an emotion engine that can analyze the user's emotions in real time and select or generate music accordingly.
[0122] Specifically, the device detects the user's emotional state from their facial expressions and voice tone through sensors such as the camera and microphone. This data is analyzed by an emotion engine to identify the user's current emotional state. The acquired emotional data, along with location information and music listening history, is sent to the server.
[0123] Based on this emotional information, the server comprehensively analyzes regional music rankings and past listening data. Furthermore, it uses an AI algorithm to select music best suited to the user's emotions and, in some cases, generate new music tailored to their emotional state. This generation process takes into account the style, tempo, and genre that match the user's current emotions.
[0124] All of this information is transmitted from the server to the terminal, which then displays the selected music content on its user interface. Users can play music with simple operations and enjoy music that matches their mood at the time. For example, when a user is feeling stressed, they can choose relaxing music to promote endorphin release.
[0125] This means that music functions not only as mere entertainment, but also as an effective means of supporting the psychological state of its users.
[0126] The following describes the processing flow.
[0127] Step 1:
[0128] The device uses a camera and microphone to capture the user's facial expressions and voiceprints in real time. This collects data that indicates the user's emotional state.
[0129] Step 2:
[0130] The device inputs the collected emotional data into an emotion engine to analyze the user's emotions. The analysis results identify the emotions the user is currently experiencing (joy, sadness, stress, etc.).
[0131] Step 3:
[0132] The device sends analyzed sentiment data, user location information, and viewing history to the server. This allows the server to understand the user's comprehensive context.
[0133] Step 4:
[0134] The server uses received sentiment data and location information to refer to regional music rankings. It then analyzes user preferences in detail, along with related music data and listening history.
[0135] Step 5:
[0136] The server uses analysis results and AI algorithms to select music that best suits the user's specific emotional state. If necessary, it generates new music tailored to that emotional state.
[0137] Step 6:
[0138] The server sends selected or generated music data to the terminal, preparing to provide the user with the optimal music experience.
[0139] Step 7:
[0140] The device displays the received music data on the user interface, allowing users to easily play it.
[0141] Step 8:
[0142] Users can enjoy music that suits their mood by selecting and playing playlists and songs from the provided options. This song selection is sent back to the server as feedback to improve the accuracy of future AI analysis.
[0143] (Example 2)
[0144] Next, we will describe Example 2. In the following description, the data processing device 12 will be referred to as the "server" and the smart device 14 as the "terminal".
[0145] Conventional music delivery systems only recommend music based on simple location information and listening history, making it difficult to provide a personalized music experience that takes into account the user's complex emotional state. Therefore, there is a need to develop a system that allows users to easily enjoy music that matches their mood at any given time.
[0146] The identification process performed by the identification processing unit 290 of the data processing device 12 in Example 2 is realized by the following means.
[0147] In this invention, the server includes means for acquiring data from an image acquisition device and an audio acquisition device to detect the user's emotions, means for analyzing the acquired emotion data to identify the user's emotional state, and means for selecting or generating music suitable for the user using the identified emotional state, location information, and music listening history. This makes it possible to provide music optimized for the user's emotional state.
[0148] An "image acquisition device" is a device used to acquire a user's facial expression information, and includes equipment such as a camera or other video sensor.
[0149] A "voice acquisition device" is a device used to acquire the tone and volume of a user's voice, and includes equipment such as a microphone or other voice sensors.
[0150] "Emotional data" refers to data indicating the user's emotions, acquired through image acquisition devices and voice acquisition devices, and includes analysis results based on facial expression information and voice characteristics.
[0151] "Emotional state" refers to the user's psychological state identified through the analysis of emotional data, and indicates specific emotions such as joy, sadness, anger, and relaxation.
[0152] "Location information" refers to information about the geographical location where a user is located, and is data obtained using technologies such as GPS.
[0153] "Music listening history" refers to a record of music that a user has played in the past, including detailed information such as song title, artist, genre, and playback date and time.
[0154] A "music ranking" is a list that shows the popularity or priority of music based on a specific region or set criteria.
[0155] "Means for selecting or generating music" refers to technical methods or processes that select the most suitable music for a user, or create new music as needed, based on their emotional state, location information, and music listening history.
[0156] This system is a technology designed to provide a personalized music experience based on the user's emotional state. It primarily utilizes an image acquisition device, an audio acquisition device, an emotion engine, and a music generation and selection server.
[0157] The device is equipped with a camera for image acquisition and a microphone for audio acquisition, capturing the user's facial expressions and voice in real time. The collected facial expression data and voice tone are instantly transmitted to the server.
[0158] The server analyzes the received data using an emotion engine to determine the user's current emotional state. For example, if a user is smiling and speaking in a calm voice, they are judged to be happy and relaxed. Based on this analysis, the AI algorithm selects the most suitable music by referring to the identified emotional state, location information, and music listening history.
[0159] Furthermore, it can generate new music as needed. In this case, the server utilizes AI technology to create music that best suits the selected emotional state, considering the melody, tempo, and genre. This makes it possible to provide a musical experience that perfectly matches the user's mood at that moment.
[0160] All selected and generated music is sent from the server to the terminal and displayed on the terminal's user interface. Users can play this music with simple operations and enjoy a personalized musical experience. For example, if a user wants to relax at the end of a busy day, the system will select and provide piano music at a calm tempo.
[0161] As an example of a prompt, by providing the AI model with the input, "The user is currently feeling stressed. Please suggest music that is best suited to relieve stress," it is possible to suggest music that is best suited to the user.
[0162] The flow of the specific processing in Example 2 will be explained using Figure 13.
[0163] Step 1:
[0164] The device uses a camera (image acquisition device) and a microphone (audio acquisition device) to collect the user's facial expressions and voice. These sensors detect the user's facial movements, changes in expression, and voice tone and volume. The input for this step is the user's real-time facial and voice data, and the output is the collected raw emotion data. This data is collected as foundational information necessary for the next processing step.
[0165] Step 2:
[0166] The device sends the collected emotional data to a server, where it is analyzed by an emotion engine. This analysis process uses facial recognition technology and voice characteristic analysis to identify the emotions the user is expressing. The input is data about the user's facial expressions and voice, and the output is data representing the user's emotional state. For example, if the data shows a smile and a calm voice, the user's emotional state is identified as "relaxed."
[0167] Step 3:
[0168] The server uses identified emotional states to select or generate music by referencing the user's location information and music listening history. The input consists of identified emotional states, location information, and listening history, while the output is music optimized for the user. Specifically, an AI algorithm is used to select music with a tempo and genre that matches the user's current emotions. If necessary, a generative AI model is used to generate new music.
[0169] Step 4:
[0170] The server sends the selected or generated music to the terminal. The input here is the music selected or generated in step 3, and the output is the playlist displayed on the terminal's user interface. The terminal displays this music on its screen and configures the interface to make it easy for the user to select the music.
[0171] Step 5:
[0172] The user plays the music displayed on the device with simple controls. The input in this step is the selected music, and the output is the user starting the music playback and the resulting musical experience. For example, a user can further enhance their relaxed mood by selecting and playing specific relaxation music from a music playlist.
[0173] (Application Example 2)
[0174] Next, we will explain application example 2. In the following explanation, the data processing device 12 will be referred to as a "server" and the smart device 14 as a "terminal".
[0175] In modern society, entertainment that resonates with users' emotions is crucial for maintaining their psychological well-being and comfortable lives. However, existing music streaming services only offer generic music recommendations and services based on playback history, without considering the user's immediate emotional state. Therefore, there is a need for personalized services that can accurately respond to users' real-time emotions.
[0176] The specific processing performed by the specific processing unit 290 of the data processing device 12 in Application Example 2 is realized by the following means.
[0177] In this invention, the server includes an emotion engine that detects the user's emotional state, means for analyzing the user's biometric and voice information to acquire emotional data, means for analyzing the acquired emotional data, location information, music listening history, and regional music rankings to select the optimal music, and means for generating the selected music or providing it from existing music. This makes it possible to provide an optimal music experience that responds to the user's real-time emotional state.
[0178] "User" refers to an individual who uses the system to enjoy a music experience.
[0179] "Emotional state" refers to the user's psychological and emotional condition, and includes information inferred from facial expressions and tone of voice.
[0180] An "emotion engine" refers to an algorithm and system that analyzes a user's facial expressions and voice data to detect their emotional state in real time.
[0181] "Biometric information" refers to information such as the user's facial expressions and movements, which are acquired through cameras and sensors.
[0182] "Voice information" refers to information about changes in the user's voice and tone, captured by the microphone.
[0183] "Emotional data" refers to data that indicates the user's specific emotional state, obtained as a result of analysis by the emotion engine.
[0184] "Location information" refers to information that indicates the user's current geographical location, and is obtained using technologies such as GPS.
[0185] "Music listening history" refers to a record of songs and albums that a user has played so far.
[0186] "Regional music rankings" refer to information that shows the ranking of popular music within a specific region.
[0187] "Optimal music" refers to music that is considered most effective for the user, selected based on the user's emotional state, location information, and listening history.
[0188] "Means of generating music" refers to the process of composing new music using AI and algorithms, and providing a musical experience that matches the user's characteristics.
[0189] "Means of providing existing music" refers to the process of selecting the most suitable songs from existing tracks and delivering them to the user.
[0190] The system that realizes this invention mainly consists of a server, a terminal, and an emotion engine. The server analyzes the user's emotional state in real time using biometric and voice information transmitted from the user. This requires a terminal such as a smartphone equipped with a camera and microphone. The terminal analyzes the monitored facial expressions using OpenCV and analyzes the tone of voice using Google® Cloud Speech-to-Text. The results of this analysis are sent from the terminal to the server as emotion data.
[0191] The server analyzes emotional data through an emotion engine to select the most suitable music for the user. In addition to emotional data, it uses an AI algorithm (TensorFlow or PyTorch) to comprehensively analyze the user's location information, music listening history, and regional music rankings. Based on the results of this analysis and the Spotify API, it optimally selects generated or existing music and sends it to the device. This allows users to enjoy streaming music tailored to their emotional state.
[0192] For example, if analysis reveals that a user is experiencing stress, the server selects relaxing music and sends it to the device to help the user alleviate tension. This entire process is based on prompts such as: "If the user's facial expression indicates fatigue, recommend relaxing music."
[0193] This allows users to gain a musical experience that helps them better manage their emotions in their daily lives.
[0194] The flow of a specific process in Application Example 2 will be explained using Figure 14.
[0195] Step 1:
[0196] The device activates its camera and microphone to capture the user's current facial expressions and voice tone. Biometric information (facial expressions) and audio information (voice) are obtained as input. The output is the extracted biometric and audio data. This prepares the device for collecting real-time user emotion information.
[0197] Step 2:
[0198] The device analyzes captured video data using OpenCV and derives emotional states from the user's facial expressions. The input is biometric information, and the output is emotional data based on facial expressions. Specifically, it detects features such as eye and mouth movements to determine what emotions the user is experiencing.
[0199] Step 3:
[0200] The device uses Google Cloud Speech-to-Text to analyze voice data and determine emotional states from voice tone. The input is voice information, and the output is emotion data based on the voice. The user's mood is estimated by determining how excited or calm the voice is.
[0201] Step 4:
[0202] The device sends the emotion data obtained in steps 2 and 3 to the server. Location information and music listening history are also sent at this time. Inputs include emotion data, location information, and music listening history, while output is the transmission of data to the server.
[0203] Step 5:
[0204] The server analyzes received emotion data, location information, and music listening history using an AI algorithm (TensorFlow or PyTorch) to select the most suitable music for the user. The input is the transmitted data, and the output is the selected music playlist.
[0205] Step 6:
[0206] The server retrieves the selected music via the Spotify API and streams it to the device. The input is the selected music information, and the output is music playback on the device. This allows users to instantly enjoy music that matches their mood.
[0207] The specific processing unit 290 transmits the result of the specific processing to the smart device 14. In the smart device 14, the control unit 46A causes the output device 40 to output the result of the specific processing. The microphone 38B acquires audio indicating user input for the result of the specific processing. The control unit 46A transmits the audio data indicating user input acquired by the microphone 38B to the data processing device 12. In the data processing device 12, the specific processing unit 290 acquires the audio data.
[0208] Data generation model 58 is a so-called generative AI (Artificial Intelligence). An example of data generation model 58 is ChatGPT (registered trademark) (Internet search).<URL: https: / / openai.com / blog / chatgpt> ), Gemini (registered trademark) (Internet search) <url: https: gemini.google.com ?hl="ja">Examples of generative AI include the following. The data generation model 58 is obtained by performing deep learning on a neural network. The data generation model 58 is input with prompts containing instructions, and with inference data such as audio data representing speech, text data representing text, and image data representing images. The data generation model 58 infers from the input inference data according to the instructions indicated by the prompts, and outputs the inference results in data formats such as audio data and text data. Here, inference refers to, for example, analysis, classification, prediction, and / or summarization.
[0209] In the above embodiment, an example was given in which specific processing is performed by the data processing device 12, but the technology of this disclosure is not limited thereto, and the specific processing may also be performed by the smart device 14.
[0210] [Second Embodiment]
[0211] Figure 3 shows an example of the configuration of the data processing system 210 according to the second embodiment.
[0212] As shown in Figure 3, the data processing system 210 includes a data processing device 12 and smart glasses 214. An example of the data processing device 12 is a server.
[0213] The data processing device 12 comprises a computer 22, a database 24, and a communication interface 26. The computer 22 is an example of a "computer" related to the technology of this disclosure. The computer 22 comprises a processor 28, RAM 30, and storage 32. The processor 28, RAM 30, and storage 32 are connected to a bus 34. The database 24 and the communication interface 26 are also connected to the bus 34. The communication interface 26 is connected to a network 54. An example of the network 54 is a WAN (Wide Area Network) and / or a LAN (Local Area Network).
[0214] The smart glasses 214 include a computer 36, a microphone 238, a speaker 240, a camera 42, and a communication interface 44. The computer 36 includes a processor 46, RAM 48, and storage 50. The processor 46, RAM 48, and storage 50 are connected to a bus 52. The microphone 238, speaker 240, and camera 42 are also connected to the bus 52.
[0215] The microphone 238 receives voice signals from the user 20 and receives instructions from the user 20. The microphone 238 captures the voice signals from the user 20, converts the captured voice into audio data, and outputs it to the processor 46. The speaker 240 outputs audio according to the instructions from the processor 46.
[0216] Camera 42 is a small digital camera equipped with an optical system including a lens, aperture, and shutter, and an image sensor such as a CMOS (Complementary Metal-Oxide-Semiconductor) image sensor or a CCD (Charge Coupled Device) image sensor, and captures images of the area around the user 20 (for example, an imaging range defined by a field of view equivalent to the width of a typical healthy person's field of vision).
[0217] Communication interface 44 is connected to network 54. Communication interfaces 44 and 26 are responsible for the exchange of various information between processor 46 and processor 28 via network 54. The exchange of various information between processor 46 and processor 28 using communication interfaces 44 and 26 is performed in a secure manner.
[0218] Figure 4 shows an example of the main functions of the data processing device 12 and the smart glasses 214. As shown in Figure 4, the data processing device 12 performs specific processing using the processor 28. The storage 32 stores the specific processing program 56.
[0219] The specific processing program 56 is an example of a "program" relating to the technology of this disclosure. The processor 28 reads the specific processing program 56 from the storage 32 and executes the read specific processing program 56 on the RAM 30. The specific processing is realized by the processor 28 operating as a specific processing unit 290 in accordance with the specific processing program 56 executed on the RAM 30.
[0220] The storage 32 stores the data generation model 58 and the emotion identification model 59. The data generation model 58 and the emotion identification model 59 are used by the identification processing unit 290.
[0221] In the smart glasses 214, the processor 46 performs the reception output processing. The storage 50 stores the reception output program 60. The processor 46 reads the reception output program 60 from the storage 50 and executes the read reception output program 60 on the RAM 48. The reception output processing is realized by the processor 46 operating as a control unit 46A according to the reception output program 60 executed on the RAM 48.
[0222] Next, the identification processing performed by the identification processing unit 290 of the data processing device 12 will be described. In the following description, the data processing device 12 will be referred to as the "server" and the smart glasses 214 will be referred to as the "terminal".
[0223] The system of the present invention collects and analyzes various data in order to provide users with a more suitable music experience. Specific embodiments are described below.
[0224] First, the device uses its GPS function to obtain the user's location information. This allows it to send precise latitude and longitude to the server. Additionally, it retrieves music listening history from the streaming application on the device and sends it to the server as well.
[0225] The server retrieves regional music ranking data from external sources based on location information and listening history received from the device. This allows it to identify currently popular songs in a specific area. Furthermore, the server analyzes user preferences and past trends based on this data.
[0226] Next, the server considers the user's current time information and performs the music selection process. This is made possible by using an AI algorithm. Specifically, it integrates the user's past listening habits, current location information, and time of day data to select or generate the most suitable music. The generated music and playlists are created live to match the user's situation at that moment.
[0227] The server then sends recommended music data and generated songs to the device. The device receives this and displays it to the user to support music playback. The user can select a playlist from the presented options and play it.
[0228] The following describes the processing flow.
[0229] Step 1:
[0230] The device uses its GPS function to obtain the user's location information. The obtained location information is sent to the server as latitude and longitude data.
[0231] Step 2:
[0232] The device collects the user's music listening history from music streaming applications and sends this data to a server.
[0233] Step 3:
[0234] The server identifies the user's area based on the received location information. It then retrieves regional music ranking data related to this area from an external source.
[0235] Step 4:
[0236] The server analyzes the received music listening history to understand the user's preferences and past playback trends.
[0237] Step 5:
[0238] The server considers music trends based on the current time of day and integrates the acquired data to select songs and playlists suitable for the user. An AI algorithm carries out this selection process.
[0239] Step 6:
[0240] The server uses AI to generate new music tailored to the user's current situation.
[0241] Step 7:
[0242] The server sends the selected playlist and generated music to the device.
[0243] Step 8:
[0244] The device displays the received music information on the user interface, making it easy for users to access.
[0245] Step 9:
[0246] The user selects a playlist from the suggestions and plays the music. The feedback received during this process is sent back to the server to further refine future music recommendations.
[0247] (Example 1)
[0248] Next, we will describe Example 1. In the following description, the data processing device 12 will be referred to as the "server," and the smart glasses 214 will be referred to as the "terminal."
[0249] Traditional systems have limitations in providing users with a more personalized and context-appropriate music experience. It's necessary to generate appropriate music in real time using AI technology, in addition to user location information, music listening history, and local music rankings. Furthermore, there's a demand for a richer music experience by appropriately selecting music according to the time of day.
[0250] The identification process performed by the identification processing unit 290 of the data processing device 12 in Example 1 is realized by the following means.
[0251] In this invention, the server includes means for acquiring the user's location information, means for acquiring the user's music listening history, and means for generating music using an AI algorithm. This makes it possible to provide personalized music in real time based on the user's location information, past listening history, and current time of day.
[0252] "Means for obtaining user location information" refers to a function that uses mobile devices or other devices to obtain geographical data indicating the user's current location.
[0253] "Means for obtaining a user's music listening history" refers to a function that collects historical information about music that a user has played in the past from a digital device.
[0254] "Methods for obtaining location-based music rankings by region" refers to a function that obtains data from an external source to rank popular songs in a specific region.
[0255] "A means of analyzing and selecting music suitable for the user" refers to a function that selects songs that match the user's preferences and activities based on acquired data.
[0256] "Methods for generating music using AI algorithms" refers to a function that uses artificial intelligence technology to automatically create music tailored to the user's situation.
[0257] "Means of providing selected or generated music to users" refers to functions for delivering selected or generated music to users in real time.
[0258] This invention is a system that provides users with a personalized music experience. This system aggregates the user's location information, music listening history, and time information, and uses AI technology to select and generate music in real time. In specific embodiments, it primarily operates as follows:
[0259] First, the device uses a GPS-equipped device to obtain the user's location information. This location information is sent to the server as latitude and longitude data. The device also collects the user's music listening history from digital streaming applications and sends this to the server as well. Furthermore, the device uses the system time to obtain the current time information.
[0260] The server then receives this data and retrieves regional music rankings from external music databases and APIs. This uses publicly available digital ranking data and APIs. For example, APIs from commercially available music information services are used. The server analyzes the collected information using an AI algorithm to select or generate the most suitable music for the user. A generative AI model is used for this AI algorithm. The prompt given to the model is an instruction such as, "Based on this user's current location and time, please recommend relaxing music."
[0261] The server sends selected or generated music data to the terminal, which receives it and displays it visually to the user. The user then plays music based on the presented playlist. For example, if the user is in a park in the afternoon, the terminal sends its location and time data, and the server selects recommended songs with a relaxing effect.
[0262] In this way, a music experience optimized for the user's situation is provided. This system makes it easy for users to enjoy music that suits their lifestyle and environment.
[0263] The flow of the specific processing in Example 1 will be explained using Figure 11.
[0264] Step 1:
[0265] The terminal obtains location information using the GPS function built into the user's device. This location information is provided in the form of latitude and longitude. The input is the device's current location, and based on this, specific geographical information is generated and sent to the server.
[0266] Step 2:
[0267] The device retrieves the user's music listening history from the streaming application. This history includes information such as the title of the played song, the artist's name, and the number of plays. This data is entered as a log on the device, organized, and then output to the server.
[0268] Step 3:
[0269] The server receives location information and music listening history from the terminal and retrieves regional music ranking data from external music databases and APIs. In this process, regional information is used as input, ranking data is searched based on it, and the retrieved information is imported into the server.
[0270] Step 4:
[0271] The server integrates the acquired information and uses an AI algorithm to select or generate music suitable for the user. At this stage, a generative AI model is used, and the prompt "Please recommend relaxing music based on this user's current location and time" is input, and the optimal music data is output.
[0272] Step 5:
[0273] The server sends selected or generated music data to the terminal. This transmission uses digital communication, with music data generated by an AI algorithm as input and the data sent to the terminal as output.
[0274] Step 6:
[0275] The terminal displays the received music data to the user and prepares for playback. The input is music data sent from the server, which is displayed on the terminal's user interface. The user then starts playing specific songs through the presented playlist.
[0276] This series of processing steps enables a system that provides users with the most suitable music in real time.
[0277] (Application Example 1)
[0278] Next, we will explain Application Example 1. In the following explanation, the data processing device 12 will be referred to as the "server," and the smart glasses 214 will be referred to as the "terminal."
[0279] While modern content distribution services allow users to easily enjoy a wide variety of music, automatically providing the optimal music for each individual user remains challenging. In particular, customizing music in real time based on a user's location and time of day presents a significant technical challenge. Therefore, there is a need for a system that dynamically generates and delivers music playlists, taking into account the user's location, time, and music listening history.
[0280] The specific processing performed by the specific processing unit 290 of the data processing device 12 in Application Example 1 is realized by the following means.
[0281] In this invention, the server includes means for acquiring the location information of the user, means for acquiring the music listening history of the user, and means for dynamically generating a music playlist in consideration of the user's movement status. Thereby, it becomes possible to provide a customized music experience for each user.
[0282] The "user" is an individual who receives a music experience using the system.
[0283] The "means for acquiring location information" is a technical element used to identify the current location of the user, usually referring to GPS data, etc.
[0284] The "means for acquiring music listening history" is a technical element for recording and using the music and artist information that the user has listened to in the past.
[0285] The "means for acquiring music rankings for each region" is a technical element for obtaining data indicating the popularity of music in a specific geographical area.
[0286] The "means for analyzing and selecting music suitable for the user" is an algorithm or technical element for analyzing the acquired data and selecting the music that best suits the user.
[0287] The "means for providing music to the user" is a technical element for distributing the selected music to the user in a playable form.
[0288] The "means for dynamically generating a music playlist in consideration of the movement status" is a technical element for creating a suitable music playlist each time according to the situation where the user is moving.
[0289] This music delivery system requires specific hardware and software configurations to provide users with a dynamic music experience. First, it accurately obtains the user's location information using the smartphone's GPS function. This sends the current location data to the server. The music streaming application also records the user's past music listening history and provides this information to the server. Based on the received location information and music listening history, the server accesses a regional music ranking database and retrieves data on popular songs.
[0290] Furthermore, the server uses AI algorithms, such as TensorFlow or PyTorch, to analyze user preferences. It takes into account past listening trends, local music rankings, and time-of-day data to dynamically generate music playlists in real time. This process utilizes a generative AI model. The generated music playlists are then sent to the user's smartphone and provided in a playable format.
[0291] For example, if a user uses this system while in a moving car, their current location can be communicated, and popular songs in the area they are traveling in can be suggested. This feature helps users discover new music while simultaneously providing a music experience tailored to each user's environment.
[0292] Examples of prompt messages are as follows:
[0293] "I'm a user currently in New York who has listened to a lot of jazz in the past. Please suggest some music suitable for a morning walk."
[0294] In this way, the system can provide users with the optimal music experience in real time.
[0295] The flow of a specific process in Application Example 1 will be explained using Figure 12.
[0296] Step 1:
[0297] The device acquires the user's location information. Using the smartphone's GPS sensor, it obtains the current latitude and longitude and sends this data to the server. The input is location data obtained from GPS, and the output is location information sent to the server.
[0298] Step 2:
[0299] The device retrieves the user's music listening history from the music streaming application. It collects listening history data via the application API and sends it to the server. The input is the listening history, and the output is the listening history data sent to the server.
[0300] Step 3:
[0301] The server retrieves regional music rankings from an existing database based on the received location information. It uses the location information to search for the corresponding regional data and retrieves the ranking data. The input is location information, and the output is regional music ranking data.
[0302] Step 4:
[0303] The server uses an AI model to analyze music listening history, regional music ranking data, and time information. This analysis helps understand the user's musical preferences and tendencies. The inputs are listening history, ranking data, and time information, and the output is the analyzed pattern of the user's preferences.
[0304] Step 5:
[0305] The server uses a generative AI model to generate music playlists tailored to the user's preferences in real time. It integrates past preferences with the current context and uses prompts to determine the optimal set of songs for the AI model. A dynamically generated playlist is obtained at this stage.
[0306] Step 6:
[0307] The server transmits the generated music playlist to the terminal. The terminal provides this to the user in a playable form. The input is the generated playlist, and the output is delivered to the user as playable music data.
[0308] Furthermore, an emotion engine for estimating the user's emotions may be combined. That is, the specific processing unit 290 may estimate the user's emotions using the emotion identification model 59 and perform specific processing using the user's emotions.
[0309] The present invention is a system for providing a personalized music experience based on the user's emotional state. This system incorporates an emotion engine and can analyze the user's emotions in real time and select or generate appropriate music accordingly.
[0310] Specifically, the terminal detects the emotional state from the user's facial expressions and voice tones via sensors such as cameras and microphones. This data is analyzed by the emotion engine to identify the current emotional state of the user. The acquired emotion data is transmitted to the server together with the location information and music viewing history.
[0311] Based on this emotion information, the server comprehensively analyzes the music rankings for each region and past viewing data. Furthermore, using an AI algorithm, it selects music optimal for the user's emotions and, in some cases, generates new music according to the emotional state. In this generation process, the melody, tempo, and genre suitable for the user's current emotions are considered.
[0312] All this information is transmitted from the server to the terminal, and the terminal presents these selected music contents on the user interface. The user can play the music with a simple operation and enjoy the music suitable for the emotion at that time. As a specific example, when the user is feeling stressed, they can select music with a relaxing effect to promote the endorphin-inducing effect.
[0313] This means that music functions not only as mere entertainment, but also as an effective means of supporting the psychological state of its users.
[0314] The following describes the processing flow.
[0315] Step 1:
[0316] The device uses a camera and microphone to capture the user's facial expressions and voiceprints in real time. This collects data that indicates the user's emotional state.
[0317] Step 2:
[0318] The device inputs the collected emotional data into an emotion engine to analyze the user's emotions. The analysis results identify the emotions the user is currently experiencing (joy, sadness, stress, etc.).
[0319] Step 3:
[0320] The device sends analyzed sentiment data, user location information, and viewing history to the server. This allows the server to understand the user's comprehensive context.
[0321] Step 4:
[0322] The server uses received sentiment data and location information to refer to regional music rankings. It then analyzes user preferences in detail, along with related music data and listening history.
[0323] Step 5:
[0324] The server uses analysis results and AI algorithms to select music that best suits the user's specific emotional state. If necessary, it generates new music tailored to that emotional state.
[0325] Step 6:
[0326] The server sends selected or generated music data to the terminal, preparing to provide the user with the optimal music experience.
[0327] Step 7:
[0328] The device displays the received music data on the user interface, allowing users to easily play it.
[0329] Step 8:
[0330] Users can enjoy music that suits their mood by selecting and playing playlists and songs from the provided options. This song selection is sent back to the server as feedback to improve the accuracy of future AI analysis.
[0331] (Example 2)
[0332] Next, we will describe Example 2. In the following description, the data processing device 12 will be referred to as the "server" and the smart glasses 214 will be referred to as the "terminal".
[0333] Conventional music delivery systems only recommend music based on simple location information and listening history, making it difficult to provide a personalized music experience that takes into account the user's complex emotional state. Therefore, there is a need to develop a system that allows users to easily enjoy music that matches their mood at any given time.
[0334] The identification process performed by the identification processing unit 290 of the data processing device 12 in Example 2 is realized by the following means.
[0335] In this invention, the server includes means for acquiring data from an image acquisition device and an audio acquisition device to detect the user's emotions, means for analyzing the acquired emotion data to identify the user's emotional state, and means for selecting or generating music suitable for the user using the identified emotional state, location information, and music listening history. This makes it possible to provide music optimized for the user's emotional state.
[0336] An "image acquisition device" is a device used to acquire a user's facial expression information, and includes equipment such as a camera or other video sensor.
[0337] A "voice acquisition device" is a device used to acquire the tone and volume of a user's voice, and includes equipment such as a microphone or other voice sensors.
[0338] "Emotional data" refers to data indicating the user's emotions, acquired through image acquisition devices and voice acquisition devices, and includes analysis results based on facial expression information and voice characteristics.
[0339] "Emotional state" refers to the user's psychological state identified through the analysis of emotional data, and indicates specific emotions such as joy, sadness, anger, and relaxation.
[0340] "Location information" refers to information about the geographical location where a user is located, and is data obtained using technologies such as GPS.
[0341] "Music listening history" refers to a record of music that a user has played in the past, including detailed information such as song title, artist, genre, and playback date and time.
[0342] A "music ranking" is a list that shows the popularity or priority of music based on a specific region or set criteria.
[0343] "Means for selecting or generating music" refers to technical methods or processes that select the most suitable music for a user, or create new music as needed, based on their emotional state, location information, and music listening history.
[0344] This system is a technology designed to provide a personalized music experience based on the user's emotional state. It primarily utilizes an image acquisition device, an audio acquisition device, an emotion engine, and a music generation and selection server.
[0345] The device is equipped with a camera for image acquisition and a microphone for audio acquisition, capturing the user's facial expressions and voice in real time. The collected facial expression data and voice tone are instantly transmitted to the server.
[0346] The server analyzes the received data using an emotion engine to determine the user's current emotional state. For example, if a user is smiling and speaking in a calm voice, they are judged to be happy and relaxed. Based on this analysis, the AI algorithm selects the most suitable music by referring to the identified emotional state, location information, and music listening history.
[0347] Furthermore, it can generate new music as needed. In this case, the server utilizes AI technology to create music that best suits the selected emotional state, considering the melody, tempo, and genre. This makes it possible to provide a musical experience that perfectly matches the user's mood at that moment.
[0348] All selected and generated music is sent from the server to the terminal and displayed on the terminal's user interface. Users can play this music with simple operations and enjoy a personalized musical experience. For example, if a user wants to relax at the end of a busy day, the system will select and provide piano music at a calm tempo.
[0349] As an example of a prompt, by providing the AI model with the input, "The user is currently feeling stressed. Please suggest music that is best suited to relieve stress," it is possible to suggest music that is best suited to the user.
[0350] The flow of the specific processing in Example 2 will be explained using Figure 13.
[0351] Step 1:
[0352] The device uses a camera (image acquisition device) and a microphone (audio acquisition device) to collect the user's facial expressions and voice. These sensors detect the user's facial movements, changes in expression, and voice tone and volume. The input for this step is the user's real-time facial and voice data, and the output is the collected raw emotion data. This data is collected as foundational information necessary for the next processing step.
[0353] Step 2:
[0354] The device sends the collected emotional data to a server, where it is analyzed by an emotion engine. This analysis process uses facial recognition technology and voice characteristic analysis to identify the emotions the user is expressing. The input is data about the user's facial expressions and voice, and the output is data representing the user's emotional state. For example, if the data shows a smile and a calm voice, the user's emotional state is identified as "relaxed."
[0355] Step 3:
[0356] The server uses identified emotional states to select or generate music by referencing the user's location information and music listening history. The input consists of identified emotional states, location information, and listening history, while the output is music optimized for the user. Specifically, an AI algorithm is used to select music with a tempo and genre that matches the user's current emotions. If necessary, a generative AI model is used to generate new music.
[0357] Step 4:
[0358] The server sends the selected or generated music to the terminal. The input here is the music selected or generated in step 3, and the output is the playlist displayed on the terminal's user interface. The terminal displays this music on its screen and configures the interface to make it easy for the user to select the music.
[0359] Step 5:
[0360] The user plays the music displayed on the device with simple controls. The input in this step is the selected music, and the output is the user starting the music playback and the resulting musical experience. For example, a user can further enhance their relaxed mood by selecting and playing specific relaxation music from a music playlist.
[0361] (Application Example 2)
[0362] Next, we will explain application example 2. In the following explanation, the data processing device 12 will be referred to as the "server," and the smart glasses 214 will be referred to as the "terminal."
[0363] In modern society, entertainment that resonates with users' emotions is crucial for maintaining their psychological well-being and comfortable lives. However, existing music streaming services only offer generic music recommendations and services based on playback history, without considering the user's immediate emotional state. Therefore, there is a need for personalized services that can accurately respond to users' real-time emotions.
[0364] The specific processing performed by the specific processing unit 290 of the data processing device 12 in Application Example 2 is realized by the following means.
[0365] In this invention, the server includes an emotion engine that detects the user's emotional state, means for analyzing the user's biometric and voice information to acquire emotional data, means for analyzing the acquired emotional data, location information, music listening history, and regional music rankings to select the optimal music, and means for generating the selected music or providing it from existing music. This makes it possible to provide an optimal music experience that responds to the user's real-time emotional state.
[0366] "User" refers to an individual who uses the system to enjoy a music experience.
[0367] "Emotional state" refers to the user's psychological and emotional condition, and includes information inferred from facial expressions and tone of voice.
[0368] An "emotion engine" refers to an algorithm and system that analyzes a user's facial expressions and voice data to detect their emotional state in real time.
[0369] "Biometric information" refers to information such as the user's facial expressions and movements, which are acquired through cameras and sensors.
[0370] "Voice information" refers to information about changes in the user's voice and tone, captured by the microphone.
[0371] "Emotional data" refers to data that indicates the user's specific emotional state, obtained as a result of analysis by the emotion engine.
[0372] "Location information" refers to information that indicates the user's current geographical location, and is obtained using technologies such as GPS.
[0373] "Music listening history" refers to a record of songs and albums that a user has played so far.
[0374] "Regional music rankings" refer to information that shows the ranking of popular music within a specific region.
[0375] "Optimal music" refers to music that is considered most effective for the user, selected based on the user's emotional state, location information, and listening history.
[0376] "Means of generating music" refers to the process of composing new music using AI and algorithms, and providing a musical experience that matches the user's characteristics.
[0377] "Means of providing existing music" refers to the process of selecting the most suitable songs from existing tracks and delivering them to the user.
[0378] The system that realizes this invention mainly consists of a server, a terminal, and an emotion engine. The server analyzes the user's emotional state in real time using biometric and voice information transmitted from the user. This requires a terminal such as a smartphone equipped with a camera and microphone. The terminal analyzes monitored facial expressions using OpenCV and analyzes voice tone using Google Cloud Speech-to-Text. The results of this analysis are sent from the terminal to the server as emotion data.
[0379] The server analyzes emotional data through an emotion engine to select the most suitable music for the user. In addition to emotional data, it uses an AI algorithm (TensorFlow or PyTorch) to comprehensively analyze the user's location information, music listening history, and regional music rankings. Based on the results of this analysis and the Spotify API, it optimally selects generated or existing music and sends it to the device. This allows users to enjoy streaming music tailored to their emotional state.
[0380] For example, if analysis reveals that a user is experiencing stress, the server selects relaxing music and sends it to the device to help the user alleviate tension. This entire process is based on prompts such as: "If the user's facial expression indicates fatigue, recommend relaxing music."
[0381] This allows users to gain a musical experience that helps them better manage their emotions in their daily lives.
[0382] The flow of a specific process in Application Example 2 will be explained using Figure 14.
[0383] Step 1:
[0384] The device activates its camera and microphone to capture the user's current facial expressions and voice tone. Biometric information (facial expressions) and audio information (voice) are obtained as input. The output is the extracted biometric and audio data. This prepares the device for collecting real-time user emotion information.
[0385] Step 2:
[0386] The device analyzes captured video data using OpenCV and derives emotional states from the user's facial expressions. The input is biometric information, and the output is emotional data based on facial expressions. Specifically, it detects features such as eye and mouth movements to determine what emotions the user is experiencing.
[0387] Step 3:
[0388] The device uses Google Cloud Speech-to-Text to analyze voice data and determine emotional states from voice tone. The input is voice information, and the output is emotion data based on the voice. The user's mood is estimated by determining how excited or calm the voice is.
[0389] Step 4:
[0390] The device sends the emotion data obtained in steps 2 and 3 to the server. Location information and music listening history are also sent at this time. Inputs include emotion data, location information, and music listening history, while output is the transmission of data to the server.
[0391] Step 5:
[0392] The server analyzes received emotion data, location information, and music listening history using an AI algorithm (TensorFlow or PyTorch) to select the most suitable music for the user. The input is the transmitted data, and the output is the selected music playlist.
[0393] Step 6:
[0394] The server retrieves the selected music via the Spotify API and streams it to the device. The input is the selected music information, and the output is music playback on the device. This allows users to instantly enjoy music that matches their mood.
[0395] The specific processing unit 290 transmits the result of the specific processing to the smart glasses 214. In the smart glasses 214, the control unit 46A causes the speaker 240 to output the result of the specific processing. The microphone 238 acquires audio indicating user input for the result of the specific processing. The control unit 46A transmits the audio data indicating user input acquired by the microphone 238 to the data processing unit 12. In the data processing unit 12, the specific processing unit 290 acquires the audio data.
[0396] Data generation model 58 is a type of so-called generative AI (Artificial Intelligence). One example of data generation model 58 is ChatGPT (Internet search<URL: https: / / openai.com / blog / chatgpt> ), Gemini (Internet search) <url: https: gemini.google.com ?hl="ja">Examples of generative AI include the following. The data generation model 58 is obtained by performing deep learning on a neural network. The data generation model 58 is input with prompts containing instructions, and with inference data such as audio data representing speech, text data representing text, and image data representing images. The data generation model 58 infers from the input inference data according to the instructions indicated by the prompts, and outputs the inference results in data formats such as audio data and text data. Here, inference refers to, for example, analysis, classification, prediction, and / or summarization.
[0397] In the above embodiment, an example was given in which specific processing is performed by the data processing device 12, but the technology of this disclosure is not limited thereto, and the specific processing may also be performed by the smart glasses 214.
[0398] [Third Embodiment]
[0399] Figure 5 shows an example of the configuration of the data processing system 310 according to the third embodiment.
[0400] As shown in Figure 5, the data processing system 310 includes a data processing device 12 and a headset terminal 314. An example of the data processing device 12 is a server.
[0401] The data processing device 12 comprises a computer 22, a database 24, and a communication interface 26. The computer 22 is an example of a "computer" related to the technology of this disclosure. The computer 22 comprises a processor 28, RAM 30, and storage 32. The processor 28, RAM 30, and storage 32 are connected to a bus 34. The database 24 and the communication interface 26 are also connected to the bus 34. The communication interface 26 is connected to a network 54. An example of the network 54 is a WAN (Wide Area Network) and / or a LAN (Local Area Network).
[0402] The headset terminal 314 includes a computer 36, a microphone 238, a speaker 240, a camera 42, a communication interface 44, and a display 343. The computer 36 includes a processor 46, RAM 48, and storage 50. The processor 46, RAM 48, and storage 50 are connected to a bus 52. The microphone 238, speaker 240, camera 42, and display 343 are also connected to the bus 52.
[0403] The microphone 238 receives voice signals from the user 20 and receives instructions from the user 20. The microphone 238 captures the voice signals from the user 20, converts the captured voice into audio data, and outputs it to the processor 46. The speaker 240 outputs audio according to the instructions from the processor 46.
[0404] Camera 42 is a small digital camera equipped with an optical system including a lens, aperture, and shutter, and an image sensor such as a CMOS (Complementary Metal-Oxide-Semiconductor) image sensor or a CCD (Charge Coupled Device) image sensor, and captures images of the area around the user 20 (for example, an imaging range defined by a field of view equivalent to the width of a typical healthy person's field of vision).
[0405] Communication interface 44 is connected to network 54. Communication interfaces 44 and 26 are responsible for the exchange of various information between processor 46 and processor 28 via network 54. The exchange of various information between processor 46 and processor 28 using communication interfaces 44 and 26 is performed in a secure manner.
[0406] Figure 6 shows an example of the main functions of the data processing device 12 and the headset terminal 314. As shown in Figure 6, the data processing device 12 performs specific processing using the processor 28. The storage 32 stores the specific processing program 56.
[0407] The specific processing program 56 is an example of a "program" relating to the technology of this disclosure. The processor 28 reads the specific processing program 56 from the storage 32 and executes the read specific processing program 56 on the RAM 30. The specific processing is realized by the processor 28 operating as a specific processing unit 290 in accordance with the specific processing program 56 executed on the RAM 30.
[0408] The storage 32 stores the data generation model 58 and the emotion identification model 59. The data generation model 58 and the emotion identification model 59 are used by the identification processing unit 290.
[0409] In the headset terminal 314, the processor 46 performs the reception output processing. The storage 50 stores the reception output program 60. The processor 46 reads the reception output program 60 from the storage 50 and executes the read reception output program 60 on the RAM 48. The reception output processing is realized by the processor 46 operating as a control unit 46A according to the reception output program 60 executed on the RAM 48.
[0410] Next, the specific processing performed by the specific processing unit 290 of the data processing device 12 will be described. In the following description, the data processing device 12 will be referred to as the "server" and the headset terminal 314 will be referred to as the "terminal".
[0411] The system of the present invention collects and analyzes various data in order to provide users with a more suitable music experience. Specific embodiments are described below.
[0412] First, the device uses its GPS function to obtain the user's location information. This allows it to send precise latitude and longitude to the server. Additionally, it retrieves music listening history from the streaming application on the device and sends it to the server as well.
[0413] The server retrieves regional music ranking data from external sources based on location information and listening history received from the device. This allows it to identify currently popular songs in a specific area. Furthermore, the server analyzes user preferences and past trends based on this data.
[0414] Next, the server considers the user's current time information and performs the music selection process. This is made possible by using an AI algorithm. Specifically, it integrates the user's past listening habits, current location information, and time of day data to select or generate the most suitable music. The generated music and playlists are created live to match the user's situation at that moment.
[0415] The server then sends recommended music data and generated songs to the device. The device receives this and displays it to the user to support music playback. The user can select a playlist from the presented options and play it.
[0416] The following describes the processing flow.
[0417] Step 1:
[0418] The device uses its GPS function to obtain the user's location information. The obtained location information is sent to the server as latitude and longitude data.
[0419] Step 2:
[0420] The device collects the user's music listening history from music streaming applications and sends this data to a server.
[0421] Step 3:
[0422] The server identifies the user's area based on the received location information. It then retrieves regional music ranking data related to this area from an external source.
[0423] Step 4:
[0424] The server analyzes the received music listening history to understand the user's preferences and past playback trends.
[0425] Step 5:
[0426] The server considers music trends based on the current time of day and integrates the acquired data to select songs and playlists suitable for the user. An AI algorithm carries out this selection process.
[0427] Step 6:
[0428] The server uses AI to generate new music tailored to the user's current situation.
[0429] Step 7:
[0430] The server sends the selected playlist and generated music to the device.
[0431] Step 8:
[0432] The device displays the received music information on the user interface, making it easy for users to access.
[0433] Step 9:
[0434] The user selects a playlist from the suggestions and plays the music. The feedback received during this process is sent back to the server to further refine future music recommendations.
[0435] (Example 1)
[0436] Next, we will describe Example 1. In the following description, the data processing device 12 will be referred to as the "server," and the headset-type terminal 314 will be referred to as the "terminal."
[0437] Traditional systems have limitations in providing users with a more personalized and context-appropriate music experience. It's necessary to generate appropriate music in real time using AI technology, in addition to user location information, music listening history, and local music rankings. Furthermore, there's a demand for a richer music experience by appropriately selecting music according to the time of day.
[0438] The identification process performed by the identification processing unit 290 of the data processing device 12 in Example 1 is realized by the following means.
[0439] In this invention, the server includes means for acquiring the user's location information, means for acquiring the user's music listening history, and means for generating music using an AI algorithm. This makes it possible to provide personalized music in real time based on the user's location information, past listening history, and current time of day.
[0440] "Means for obtaining user location information" refers to a function that uses mobile devices or other devices to obtain geographical data indicating the user's current location.
[0441] "Means for obtaining a user's music listening history" refers to a function that collects historical information about music that a user has played in the past from a digital device.
[0442] "Methods for obtaining location-based music rankings by region" refers to a function that obtains data from an external source to rank popular songs in a specific region.
[0443] "A means of analyzing and selecting music suitable for the user" refers to a function that selects songs that match the user's preferences and activities based on acquired data.
[0444] "Methods for generating music using AI algorithms" refers to a function that uses artificial intelligence technology to automatically create music tailored to the user's situation.
[0445] "Means of providing selected or generated music to users" refers to functions for delivering selected or generated music to users in real time.
[0446] This invention is a system that provides users with a personalized music experience. This system aggregates the user's location information, music listening history, and time information, and uses AI technology to select and generate music in real time. In specific embodiments, it primarily operates as follows:
[0447] First, the device uses a GPS-equipped device to obtain the user's location information. This location information is sent to the server as latitude and longitude data. The device also collects the user's music listening history from digital streaming applications and sends this to the server as well. Furthermore, the device uses the system time to obtain the current time information.
[0448] The server then receives this data and retrieves regional music rankings from external music databases and APIs. This uses publicly available digital ranking data and APIs. For example, APIs from commercially available music information services are used. The server analyzes the collected information using an AI algorithm to select or generate the most suitable music for the user. A generative AI model is used for this AI algorithm. The prompt given to the model is an instruction such as, "Based on this user's current location and time, please recommend relaxing music."
[0449] The server sends selected or generated music data to the terminal, which receives it and displays it visually to the user. The user then plays music based on the presented playlist. For example, if the user is in a park in the afternoon, the terminal sends its location and time data, and the server selects recommended songs with a relaxing effect.
[0450] In this way, a music experience optimized for the user's situation is provided. This system makes it easy for users to enjoy music that suits their lifestyle and environment.
[0451] The flow of the specific processing in Example 1 will be explained using Figure 11.
[0452] Step 1:
[0453] The terminal obtains location information using the GPS function built into the user's device. This location information is provided in the form of latitude and longitude. The input is the device's current location, and based on this, specific geographical information is generated and sent to the server.
[0454] Step 2:
[0455] The device retrieves the user's music listening history from the streaming application. This history includes information such as the title of the played song, the artist's name, and the number of plays. This data is entered as a log on the device, organized, and then output to the server.
[0456] Step 3:
[0457] The server receives location information and music listening history from the terminal and retrieves regional music ranking data from external music databases and APIs. In this process, regional information is used as input, ranking data is searched based on it, and the retrieved information is imported into the server.
[0458] Step 4:
[0459] The server integrates the acquired information and uses an AI algorithm to select or generate music suitable for the user. At this stage, a generative AI model is used, and the prompt "Please recommend relaxing music based on this user's current location and time" is input, and the optimal music data is output.
[0460] Step 5:
[0461] The server sends selected or generated music data to the terminal. This transmission uses digital communication, with music data generated by an AI algorithm as input and the data sent to the terminal as output.
[0462] Step 6:
[0463] The terminal displays the received music data to the user and prepares for playback. The input is music data sent from the server, which is displayed on the terminal's user interface. The user then starts playing specific songs through the presented playlist.
[0464] This series of processing steps enables a system that provides users with the most suitable music in real time.
[0465] (Application Example 1)
[0466] Next, we will explain Application Example 1. In the following explanation, the data processing device 12 will be referred to as the "server," and the headset-type terminal 314 will be referred to as the "terminal."
[0467] While modern content distribution services allow users to easily enjoy a wide variety of music, automatically providing the optimal music for each individual user remains challenging. In particular, customizing music in real time based on a user's location and time of day presents a significant technical challenge. Therefore, there is a need for a system that dynamically generates and delivers music playlists, taking into account the user's location, time, and music listening history.
[0468] The specific processing performed by the specific processing unit 290 of the data processing device 12 in Application Example 1 is realized by the following means.
[0469] In this invention, the server includes means for acquiring the user's location information, means for acquiring the user's music listening history, and means for dynamically generating a music playlist considering the user's movement. This makes it possible to provide a customized music experience for each user.
[0470] A "user" is an individual who uses the system to experience music.
[0471] "Means of acquiring location information" refers to technical elements used to determine the user's current location, and usually refers to GPS data, etc.
[0472] "Means for obtaining music listening history" refers to the technical elements used to record and utilize information about songs and artists that users have listened to in the past.
[0473] "Means of obtaining regional music rankings" refers to the technical elements used to obtain data indicating the popularity of music in a specific geographical area.
[0474] "Methods for analyzing and selecting music suitable for the user" refer to algorithms and technical elements that analyze acquired data and select the music that best suits the user.
[0475] "Means of providing music to users" refers to the technical elements for delivering selected music to users in a playable format.
[0476] "Means for dynamically generating music playlists while considering movement circumstances" refers to the technical elements for creating a suitable music playlist each time, depending on the user's movement situation.
[0477] This music delivery system requires specific hardware and software configurations to provide users with a dynamic music experience. First, it accurately obtains the user's location information using the smartphone's GPS function. This sends the current location data to the server. The music streaming application also records the user's past music listening history and provides this information to the server. Based on the received location information and music listening history, the server accesses a regional music ranking database and retrieves data on popular songs.
[0478] Furthermore, the server uses AI algorithms, such as TensorFlow or PyTorch, to analyze user preferences. It takes into account past listening trends, local music rankings, and time-of-day data to dynamically generate music playlists in real time. This process utilizes a generative AI model. The generated music playlists are then sent to the user's smartphone and provided in a playable format.
[0479] For example, if a user uses this system while in a moving car, their current location can be communicated, and popular songs in the area they are traveling in can be suggested. This feature helps users discover new music while simultaneously providing a music experience tailored to each user's environment.
[0480] Examples of prompt messages are as follows:
[0481] "I'm a user currently in New York who has listened to a lot of jazz in the past. Please suggest some music suitable for a morning walk."
[0482] In this way, the system can provide users with the optimal music experience in real time.
[0483] The flow of a specific process in Application Example 1 will be explained using Figure 12.
[0484] Step 1:
[0485] The device acquires the user's location information. Using the smartphone's GPS sensor, it obtains the current latitude and longitude and sends this data to the server. The input is location data obtained from GPS, and the output is location information sent to the server.
[0486] Step 2:
[0487] The device retrieves the user's music listening history from the music streaming application. It collects listening history data via the application API and sends it to the server. The input is the listening history, and the output is the listening history data sent to the server.
[0488] Step 3:
[0489] The server retrieves regional music rankings from an existing database based on the received location information. It uses the location information to search for the corresponding regional data and retrieves the ranking data. The input is location information, and the output is regional music ranking data.
[0490] Step 4:
[0491] The server uses an AI model to analyze music listening history, regional music ranking data, and time information. This analysis helps understand the user's musical preferences and tendencies. The inputs are listening history, ranking data, and time information, and the output is the analyzed pattern of the user's preferences.
[0492] Step 5:
[0493] The server uses a generative AI model to generate music playlists tailored to the user's preferences in real time. It integrates past preferences with the current context and uses prompts to determine the optimal set of songs for the AI model. A dynamically generated playlist is obtained at this stage.
[0494] Step 6:
[0495] The server sends the generated music playlist to the terminal. The terminal then provides this to the user in a playable format. The input is the generated playlist, and the output is delivered to the user as playable music data.
[0496] Furthermore, an emotion engine that estimates the user's emotions may be incorporated. That is, the identification processing unit 290 may use the emotion identification model 59 to estimate the user's emotions and perform identification processing using the user's emotions.
[0497] This invention is a system for providing a personalized music experience based on the user's emotional state. This system incorporates an emotion engine that can analyze the user's emotions in real time and select or generate music accordingly.
[0498] Specifically, the device detects the user's emotional state from their facial expressions and voice tone through sensors such as the camera and microphone. This data is analyzed by an emotion engine to identify the user's current emotional state. The acquired emotional data, along with location information and music listening history, is sent to the server.
[0499] Based on this emotional information, the server comprehensively analyzes regional music rankings and past listening data. Furthermore, it uses an AI algorithm to select music best suited to the user's emotions and, in some cases, generate new music tailored to their emotional state. This generation process takes into account the style, tempo, and genre that match the user's current emotions.
[0500] All of this information is transmitted from the server to the terminal, which then displays the selected music content on its user interface. Users can play music with simple operations and enjoy music that matches their mood at the time. For example, when a user is feeling stressed, they can choose relaxing music to promote endorphin release.
[0501] This means that music functions not only as mere entertainment, but also as an effective means of supporting the psychological state of its users.
[0502] The following describes the processing flow.
[0503] Step 1:
[0504] The device uses a camera and microphone to capture the user's facial expressions and voiceprints in real time. This collects data that indicates the user's emotional state.
[0505] Step 2:
[0506] The device inputs the collected emotional data into an emotion engine to analyze the user's emotions. The analysis results identify the emotions the user is currently experiencing (joy, sadness, stress, etc.).
[0507] Step 3:
[0508] The device sends analyzed sentiment data, user location information, and viewing history to the server. This allows the server to understand the user's comprehensive context.
[0509] Step 4:
[0510] The server uses received sentiment data and location information to refer to regional music rankings. It then analyzes user preferences in detail, along with related music data and listening history.
[0511] Step 5:
[0512] The server uses analysis results and AI algorithms to select music that best suits the user's specific emotional state. If necessary, it generates new music tailored to that emotional state.
[0513] Step 6:
[0514] The server sends selected or generated music data to the terminal, preparing to provide the user with the optimal music experience.
[0515] Step 7:
[0516] The device displays the received music data on the user interface, allowing users to easily play it.
[0517] Step 8:
[0518] Users can enjoy music that suits their mood by selecting and playing playlists and songs from the provided options. This song selection is sent back to the server as feedback to improve the accuracy of future AI analysis.
[0519] (Example 2)
[0520] Next, we will describe Example 2. In the following description, the data processing device 12 will be referred to as the "server," and the headset-type terminal 314 will be referred to as the "terminal."
[0521] Conventional music delivery systems only recommend music based on simple location information and listening history, making it difficult to provide a personalized music experience that takes into account the user's complex emotional state. Therefore, there is a need to develop a system that allows users to easily enjoy music that matches their mood at any given time.
[0522] The identification process performed by the identification processing unit 290 of the data processing device 12 in Example 2 is realized by the following means.
[0523] In this invention, the server includes means for acquiring data from an image acquisition device and an audio acquisition device to detect the user's emotions, means for analyzing the acquired emotion data to identify the user's emotional state, and means for selecting or generating music suitable for the user using the identified emotional state, location information, and music listening history. This makes it possible to provide music optimized for the user's emotional state.
[0524] An "image acquisition device" is a device used to acquire a user's facial expression information, and includes equipment such as a camera or other video sensor.
[0525] A "voice acquisition device" is a device used to acquire the tone and volume of a user's voice, and includes equipment such as a microphone or other voice sensors.
[0526] "Emotional data" refers to data indicating the user's emotions, acquired through image acquisition devices and voice acquisition devices, and includes analysis results based on facial expression information and voice characteristics.
[0527] "Emotional state" refers to the user's psychological state identified through the analysis of emotional data, and indicates specific emotions such as joy, sadness, anger, and relaxation.
[0528] "Location information" refers to information about the geographical location where a user is located, and is data obtained using technologies such as GPS.
[0529] "Music listening history" refers to a record of music that a user has played in the past, including detailed information such as song title, artist, genre, and playback date and time.
[0530] A "music ranking" is a list that shows the popularity or priority of music based on a specific region or set criteria.
[0531] "Means for selecting or generating music" refers to technical methods or processes that select the most suitable music for a user, or create new music as needed, based on their emotional state, location information, and music listening history.
[0532] This system is a technology designed to provide a personalized music experience based on the user's emotional state. It primarily utilizes an image acquisition device, an audio acquisition device, an emotion engine, and a music generation and selection server.
[0533] The device is equipped with a camera for image acquisition and a microphone for audio acquisition, capturing the user's facial expressions and voice in real time. The collected facial expression data and voice tone are instantly transmitted to the server.
[0534] The server analyzes the received data using an emotion engine to determine the user's current emotional state. For example, if a user is smiling and speaking in a calm voice, they are judged to be happy and relaxed. Based on this analysis, the AI algorithm selects the most suitable music by referring to the identified emotional state, location information, and music listening history.
[0535] Furthermore, it can generate new music as needed. In this case, the server utilizes AI technology to create music that best suits the selected emotional state, considering the melody, tempo, and genre. This makes it possible to provide a musical experience that perfectly matches the user's mood at that moment.
[0536] All selected and generated music is sent from the server to the terminal and displayed on the terminal's user interface. Users can play this music with simple operations and enjoy a personalized musical experience. For example, if a user wants to relax at the end of a busy day, the system will select and provide piano music at a calm tempo.
[0537] As an example of a prompt, by providing the AI model with the input, "The user is currently feeling stressed. Please suggest music that is best suited to relieve stress," it is possible to suggest music that is best suited to the user.
[0538] The flow of the specific processing in Example 2 will be explained using Figure 13.
[0539] Step 1:
[0540] The device uses a camera (image acquisition device) and a microphone (audio acquisition device) to collect the user's facial expressions and voice. These sensors detect the user's facial movements, changes in expression, and voice tone and volume. The input for this step is the user's real-time facial and voice data, and the output is the collected raw emotion data. This data is collected as foundational information necessary for the next processing step.
[0541] Step 2:
[0542] The device sends the collected emotional data to a server, where it is analyzed by an emotion engine. This analysis process uses facial recognition technology and voice characteristic analysis to identify the emotions the user is expressing. The input is data about the user's facial expressions and voice, and the output is data representing the user's emotional state. For example, if the data shows a smile and a calm voice, the user's emotional state is identified as "relaxed."
[0543] Step 3:
[0544] The server uses identified emotional states to select or generate music by referencing the user's location information and music listening history. The input consists of identified emotional states, location information, and listening history, while the output is music optimized for the user. Specifically, an AI algorithm is used to select music with a tempo and genre that matches the user's current emotions. If necessary, a generative AI model is used to generate new music.
[0545] Step 4:
[0546] The server sends the selected or generated music to the terminal. The input here is the music selected or generated in step 3, and the output is the playlist displayed on the terminal's user interface. The terminal displays this music on its screen and configures the interface to make it easy for the user to select the music.
[0547] Step 5:
[0548] The user plays the music displayed on the device with simple controls. The input in this step is the selected music, and the output is the user starting the music playback and the resulting musical experience. For example, a user can further enhance their relaxed mood by selecting and playing specific relaxation music from a music playlist.
[0549] (Application Example 2)
[0550] Next, we will explain application example 2. In the following explanation, the data processing device 12 will be referred to as the "server," and the headset-type terminal 314 will be referred to as the "terminal."
[0551] In modern society, entertainment that resonates with users' emotions is crucial for maintaining their psychological well-being and comfortable lives. However, existing music streaming services only offer generic music recommendations and services based on playback history, without considering the user's immediate emotional state. Therefore, there is a need for personalized services that can accurately respond to users' real-time emotions.
[0552] The specific processing performed by the specific processing unit 290 of the data processing device 12 in Application Example 2 is realized by the following means.
[0553] In this invention, the server includes an emotion engine that detects the user's emotional state, means for analyzing the user's biometric and voice information to acquire emotional data, means for analyzing the acquired emotional data, location information, music listening history, and regional music rankings to select the optimal music, and means for generating the selected music or providing it from existing music. This makes it possible to provide an optimal music experience that responds to the user's real-time emotional state.
[0554] "User" refers to an individual who uses the system to enjoy a music experience.
[0555] "Emotional state" refers to the user's psychological and emotional condition, and includes information inferred from facial expressions and tone of voice.
[0556] An "emotion engine" refers to an algorithm and system that analyzes a user's facial expressions and voice data to detect their emotional state in real time.
[0557] "Biometric information" refers to information such as the user's facial expressions and movements, which are acquired through cameras and sensors.
[0558] "Voice information" refers to information about changes in the user's voice and tone, captured by the microphone.
[0559] "Emotional data" refers to data that indicates the user's specific emotional state, obtained as a result of analysis by the emotion engine.
[0560] "Location information" refers to information that indicates the user's current geographical location, and is obtained using technologies such as GPS.
[0561] "Music listening history" refers to a record of songs and albums that a user has played so far.
[0562] "Regional music rankings" refer to information that shows the ranking of popular music within a specific region.
[0563] "Optimal music" refers to music that is considered most effective for the user, selected based on the user's emotional state, location information, and listening history.
[0564] "Means of generating music" refers to the process of composing new music using AI and algorithms, and providing a musical experience that matches the user's characteristics.
[0565] "Means of providing existing music" refers to the process of selecting the most suitable songs from existing tracks and delivering them to the user.
[0566] The system that realizes this invention mainly consists of a server, a terminal, and an emotion engine. The server analyzes the user's emotional state in real time using biometric and voice information transmitted from the user. This requires a terminal such as a smartphone equipped with a camera and microphone. The terminal analyzes monitored facial expressions using OpenCV and analyzes voice tone using Google Cloud Speech-to-Text. The results of this analysis are sent from the terminal to the server as emotion data.
[0567] The server analyzes emotional data through an emotion engine to select the most suitable music for the user. In addition to emotional data, it uses an AI algorithm (TensorFlow or PyTorch) to comprehensively analyze the user's location information, music listening history, and regional music rankings. Based on the results of this analysis and the Spotify API, it optimally selects generated or existing music and sends it to the device. This allows users to enjoy streaming music tailored to their emotional state.
[0568] For example, if analysis reveals that a user is experiencing stress, the server selects relaxing music and sends it to the device to help the user alleviate tension. This entire process is based on prompts such as: "If the user's facial expression indicates fatigue, recommend relaxing music."
[0569] This allows users to gain a musical experience that helps them better manage their emotions in their daily lives.
[0570] The flow of a specific process in Application Example 2 will be explained using Figure 14.
[0571] Step 1:
[0572] The device activates its camera and microphone to capture the user's current facial expressions and voice tone. Biometric information (facial expressions) and audio information (voice) are obtained as input. The output is the extracted biometric and audio data. This prepares the device for collecting real-time user emotion information.
[0573] Step 2:
[0574] The device analyzes captured video data using OpenCV and derives emotional states from the user's facial expressions. The input is biometric information, and the output is emotional data based on facial expressions. Specifically, it detects features such as eye and mouth movements to determine what emotions the user is experiencing.
[0575] Step 3:
[0576] The device uses Google Cloud Speech-to-Text to analyze voice data and determine emotional states from voice tone. The input is voice information, and the output is emotion data based on the voice. The user's mood is estimated by determining how excited or calm the voice is.
[0577] Step 4:
[0578] The device sends the emotion data obtained in steps 2 and 3 to the server. Location information and music listening history are also sent at this time. Inputs include emotion data, location information, and music listening history, while output is the transmission of data to the server.
[0579] Step 5:
[0580] The server analyzes received emotion data, location information, and music listening history using an AI algorithm (TensorFlow or PyTorch) to select the most suitable music for the user. The input is the transmitted data, and the output is the selected music playlist.
[0581] Step 6:
[0582] The server retrieves the selected music via the Spotify API and streams it to the device. The input is the selected music information, and the output is music playback on the device. This allows users to instantly enjoy music that matches their mood.
[0583] The specific processing unit 290 transmits the result of the specific processing to the headset terminal 314. In the headset terminal 314, the control unit 46A causes the speaker 240 and display 343 to output the result of the specific processing. The microphone 238 acquires audio indicating user input for the result of the specific processing. The control unit 46A transmits the audio data indicating user input acquired by the microphone 238 to the data processing unit 12. In the data processing unit 12, the specific processing unit 290 acquires the audio data.
[0584] Data generation model 58 is a type of so-called generative AI (Artificial Intelligence). One example of data generation model 58 is ChatGPT (Internet search<URL: https: / / openai.com / blog / chatgpt> ), Gemini (Internet search) <url: https: gemini.google.com ?hl="ja">Examples of generative AI include the following. The data generation model 58 is obtained by performing deep learning on a neural network. The data generation model 58 is input with prompts containing instructions, and with inference data such as audio data representing speech, text data representing text, and image data representing images. The data generation model 58 infers from the input inference data according to the instructions indicated by the prompts, and outputs the inference results in data formats such as audio data and text data. Here, inference refers to, for example, analysis, classification, prediction, and / or summarization.
[0585] In the above embodiment, an example was given in which specific processing is performed by the data processing device 12, but the technology of this disclosure is not limited thereto, and specific processing may also be performed by the headset terminal 314.
[0586] [Fourth Embodiment]
[0587] Figure 7 shows an example of the configuration of the data processing system 410 according to the fourth embodiment.
[0588] As shown in Figure 7, the data processing system 410 includes a data processing device 12 and a robot 414. An example of the data processing device 12 is a server.
[0589] The data processing device 12 comprises a computer 22, a database 24, and a communication interface 26. The computer 22 is an example of a "computer" related to the technology of this disclosure. The computer 22 comprises a processor 28, RAM 30, and storage 32. The processor 28, RAM 30, and storage 32 are connected to a bus 34. The database 24 and the communication interface 26 are also connected to the bus 34. The communication interface 26 is connected to a network 54. An example of the network 54 is a WAN (Wide Area Network) and / or a LAN (Local Area Network).
[0590] The robot 414 includes a computer 36, a microphone 238, a speaker 240, a camera 42, a communication interface 44, and a controlled object 443. The computer 36 includes a processor 46, RAM 48, and storage 50. The processor 46, RAM 48, and storage 50 are connected to a bus 52. The microphone 238, speaker 240, camera 42, and controlled object 443 are also connected to the bus 52.
[0591] The microphone 238 receives voice signals from the user 20 and receives instructions from the user 20. The microphone 238 captures the voice signals from the user 20, converts the captured voice into audio data, and outputs it to the processor 46. The speaker 240 outputs audio according to the instructions from the processor 46.
[0592] Camera 42 is a small digital camera equipped with an optical system including a lens, aperture, and shutter, and an image sensor such as a CMOS (Complementary Metal-Oxide-Semiconductor) image sensor or a CCD (Charge Coupled Device) image sensor, and captures images of the area around the user 20 (for example, an imaging range defined by a field of view equivalent to the width of a typical healthy person's field of vision).
[0593] Communication interface 44 is connected to network 54. Communication interfaces 44 and 26 are responsible for the exchange of various information between processor 46 and processor 28 via network 54. The exchange of various information between processor 46 and processor 28 using communication interfaces 44 and 26 is performed in a secure manner.
[0594] The controlled object 443 includes a display device, LEDs in the eyes, and motors that drive the arms, hands, and feet. The posture and gestures of the robot 414 are controlled by controlling the motors of the arms, hands, and feet. Some of the robot 414's emotions can be expressed by controlling these motors. Furthermore, the robot 414's facial expressions can also be expressed by controlling the illumination state of the LEDs in its eyes.
[0595] Figure 8 shows an example of the main functions of the data processing device 12 and the robot 414. As shown in Figure 8, the data processing device 12 performs specific processing using the processor 28. The storage 32 stores the specific processing program 56.
[0596] The specific processing program 56 is an example of a "program" relating to the technology of this disclosure. The processor 28 reads the specific processing program 56 from the storage 32 and executes the read specific processing program 56 on the RAM 30. The specific processing is realized by the processor 28 operating as a specific processing unit 290 in accordance with the specific processing program 56 executed on the RAM 30.
[0597] The storage 32 stores the data generation model 58 and the emotion identification model 59. The data generation model 58 and the emotion identification model 59 are used by the identification processing unit 290.
[0598] In robot 414, the processor 46 performs the reception output processing. The storage 50 stores the reception output program 60. The processor 46 reads the reception output program 60 from the storage 50 and executes the read reception output program 60 on the RAM 48. The reception output processing is realized by the processor 46 operating as a control unit 46A according to the reception output program 60 executed on the RAM 48.
[0599] Next, the specific processing performed by the specific processing unit 290 of the data processing device 12 will be described. In the following description, the data processing device 12 will be referred to as the "server" and the robot 414 as the "terminal".
[0600] The system of the present invention collects and analyzes various data in order to provide users with a more suitable music experience. Specific embodiments are described below.
[0601] First, the device uses its GPS function to obtain the user's location information. This allows it to send precise latitude and longitude to the server. Additionally, it retrieves music listening history from the streaming application on the device and sends it to the server as well.
[0602] The server retrieves regional music ranking data from external sources based on location information and listening history received from the device. This allows it to identify currently popular songs in a specific area. Furthermore, the server analyzes user preferences and past trends based on this data.
[0603] Next, the server considers the user's current time information and performs the music selection process. This is made possible by using an AI algorithm. Specifically, it integrates the user's past listening habits, current location information, and time of day data to select or generate the most suitable music. The generated music and playlists are created live to match the user's situation at that moment.
[0604] The server then sends recommended music data and generated songs to the device. The device receives this and displays it to the user to support music playback. The user can select a playlist from the presented options and play it.
[0605] The following describes the processing flow.
[0606] Step 1:
[0607] The device uses its GPS function to obtain the user's location information. The obtained location information is sent to the server as latitude and longitude data.
[0608] Step 2:
[0609] The device collects the user's music listening history from music streaming applications and sends this data to a server.
[0610] Step 3:
[0611] The server identifies the user's area based on the received location information. It then retrieves regional music ranking data related to this area from an external source.
[0612] Step 4:
[0613] The server analyzes the received music listening history to understand the user's preferences and past playback trends.
[0614] Step 5:
[0615] The server considers music trends based on the current time of day and integrates the acquired data to select songs and playlists suitable for the user. An AI algorithm carries out this selection process.
[0616] Step 6:
[0617] The server uses AI to generate new music tailored to the user's current situation.
[0618] Step 7:
[0619] The server sends the selected playlist and generated music to the device.
[0620] Step 8:
[0621] The device displays the received music information on the user interface, making it easy for users to access.
[0622] Step 9:
[0623] The user selects a playlist from the suggestions and plays the music. The feedback received during this process is sent back to the server to further refine future music recommendations.
[0624] (Example 1)
[0625] Next, we will describe Example 1. In the following description, the data processing device 12 will be referred to as the "server" and the robot 414 as the "terminal".
[0626] Traditional systems have limitations in providing users with a more personalized and context-appropriate music experience. It's necessary to generate appropriate music in real time using AI technology, in addition to user location information, music listening history, and local music rankings. Furthermore, there's a demand for a richer music experience by appropriately selecting music according to the time of day.
[0627] The identification process performed by the identification processing unit 290 of the data processing device 12 in Example 1 is realized by the following means.
[0628] In this invention, the server includes means for acquiring the user's location information, means for acquiring the user's music listening history, and means for generating music using an AI algorithm. This makes it possible to provide personalized music in real time based on the user's location information, past listening history, and current time of day.
[0629] "Means for obtaining user location information" refers to a function that uses mobile devices or other devices to obtain geographical data indicating the user's current location.
[0630] "Means for obtaining a user's music listening history" refers to a function that collects historical information about music that a user has played in the past from a digital device.
[0631] "Methods for obtaining location-based music rankings by region" refers to a function that obtains data from an external source to rank popular songs in a specific region.
[0632] "A means of analyzing and selecting music suitable for the user" refers to a function that selects songs that match the user's preferences and activities based on acquired data.
[0633] "Methods for generating music using AI algorithms" refers to a function that uses artificial intelligence technology to automatically create music tailored to the user's situation.
[0634] "Means of providing selected or generated music to users" refers to functions for delivering selected or generated music to users in real time.
[0635] This invention is a system that provides users with a personalized music experience. This system aggregates the user's location information, music listening history, and time information, and uses AI technology to select and generate music in real time. In specific embodiments, it primarily operates as follows:
[0636] First, the device uses a GPS-equipped device to obtain the user's location information. This location information is sent to the server as latitude and longitude data. The device also collects the user's music listening history from digital streaming applications and sends this to the server as well. Furthermore, the device uses the system time to obtain the current time information.
[0637] The server then receives this data and retrieves regional music rankings from external music databases and APIs. This uses publicly available digital ranking data and APIs. For example, APIs from commercially available music information services are used. The server analyzes the collected information using an AI algorithm to select or generate the most suitable music for the user. A generative AI model is used for this AI algorithm. The prompt given to the model is an instruction such as, "Based on this user's current location and time, please recommend relaxing music."
[0638] The server sends selected or generated music data to the terminal, which receives it and displays it visually to the user. The user then plays music based on the presented playlist. For example, if the user is in a park in the afternoon, the terminal sends its location and time data, and the server selects recommended songs with a relaxing effect.
[0639] In this way, a music experience optimized for the user's situation is provided. This system makes it easy for users to enjoy music that suits their lifestyle and environment.
[0640] The flow of the specific processing in Example 1 will be explained using Figure 11.
[0641] Step 1:
[0642] The terminal obtains location information using the GPS function built into the user's device. This location information is provided in the form of latitude and longitude. The input is the device's current location, and based on this, specific geographical information is generated and sent to the server.
[0643] Step 2:
[0644] The device retrieves the user's music listening history from the streaming application. This history includes information such as the title of the played song, the artist's name, and the number of plays. This data is entered as a log on the device, organized, and then output to the server.
[0645] Step 3:
[0646] The server receives location information and music listening history from the terminal and retrieves regional music ranking data from external music databases and APIs. In this process, regional information is used as input, ranking data is searched based on it, and the retrieved information is imported into the server.
[0647] Step 4:
[0648] The server integrates the acquired information and uses an AI algorithm to select or generate music suitable for the user. At this stage, a generative AI model is used, and the prompt "Please recommend relaxing music based on this user's current location and time" is input, and the optimal music data is output.
[0649] Step 5:
[0650] The server sends selected or generated music data to the terminal. This transmission uses digital communication, with music data generated by an AI algorithm as input and the data sent to the terminal as output.
[0651] Step 6:
[0652] The terminal displays the received music data to the user and prepares for playback. The input is music data sent from the server, which is displayed on the terminal's user interface. The user then starts playing specific songs through the presented playlist.
[0653] This series of processing steps enables a system that provides users with the most suitable music in real time.
[0654] (Application Example 1)
[0655] Next, we will explain Application Example 1. In the following explanation, the data processing device 12 will be referred to as the "server" and the robot 414 as the "terminal".
[0656] While modern content distribution services allow users to easily enjoy a wide variety of music, automatically providing the optimal music for each individual user remains challenging. In particular, customizing music in real time based on a user's location and time of day presents a significant technical challenge. Therefore, there is a need for a system that dynamically generates and delivers music playlists, taking into account the user's location, time, and music listening history.
[0657] The specific processing performed by the specific processing unit 290 of the data processing device 12 in Application Example 1 is realized by the following means.
[0658] In this invention, the server includes means for acquiring the user's location information, means for acquiring the user's music listening history, and means for dynamically generating a music playlist considering the user's movement. This makes it possible to provide a customized music experience for each user.
[0659] A "user" is an individual who uses the system to experience music.
[0660] "Means of acquiring location information" refers to technical elements used to determine the user's current location, and usually refers to GPS data, etc.
[0661] "Means for obtaining music listening history" refers to the technical elements used to record and utilize information about songs and artists that users have listened to in the past.
[0662] "Means of obtaining regional music rankings" refers to the technical elements used to obtain data indicating the popularity of music in a specific geographical area.
[0663] "Methods for analyzing and selecting music suitable for the user" refer to algorithms and technical elements that analyze acquired data and select the music that best suits the user.
[0664] "Means of providing music to users" refers to the technical elements for delivering selected music to users in a playable format.
[0665] "Means for dynamically generating music playlists while considering movement circumstances" refers to the technical elements for creating a suitable music playlist each time, depending on the user's movement situation.
[0666] This music delivery system requires specific hardware and software configurations to provide users with a dynamic music experience. First, it accurately obtains the user's location information using the smartphone's GPS function. This sends the current location data to the server. The music streaming application also records the user's past music listening history and provides this information to the server. Based on the received location information and music listening history, the server accesses a regional music ranking database and retrieves data on popular songs.
[0667] Furthermore, the server uses AI algorithms, such as TensorFlow or PyTorch, to analyze user preferences. It takes into account past listening trends, local music rankings, and time-of-day data to dynamically generate music playlists in real time. This process utilizes a generative AI model. The generated music playlists are then sent to the user's smartphone and provided in a playable format.
[0668] For example, if a user uses this system while in a moving car, their current location can be communicated, and popular songs in the area they are traveling in can be suggested. This feature helps users discover new music while simultaneously providing a music experience tailored to each user's environment.
[0669] Examples of prompt messages are as follows:
[0670] "I'm a user currently in New York who has listened to a lot of jazz in the past. Please suggest some music suitable for a morning walk."
[0671] In this way, the system can provide users with the optimal music experience in real time.
[0672] The flow of a specific process in Application Example 1 will be explained using Figure 12.
[0673] Step 1:
[0674] The device acquires the user's location information. Using the smartphone's GPS sensor, it obtains the current latitude and longitude and sends this data to the server. The input is location data obtained from GPS, and the output is location information sent to the server.
[0675] Step 2:
[0676] The device retrieves the user's music listening history from the music streaming application. It collects listening history data via the application API and sends it to the server. The input is the listening history, and the output is the listening history data sent to the server.
[0677] Step 3:
[0678] The server retrieves regional music rankings from an existing database based on the received location information. It uses the location information to search for the corresponding regional data and retrieves the ranking data. The input is location information, and the output is regional music ranking data.
[0679] Step 4:
[0680] The server uses an AI model to analyze music listening history, regional music ranking data, and time information. This analysis helps understand the user's musical preferences and tendencies. The inputs are listening history, ranking data, and time information, and the output is the analyzed pattern of the user's preferences.
[0681] Step 5:
[0682] The server uses a generative AI model to generate music playlists tailored to the user's preferences in real time. It integrates past preferences with the current context and uses prompts to determine the optimal set of songs for the AI model. A dynamically generated playlist is obtained at this stage.
[0683] Step 6:
[0684] The server sends the generated music playlist to the terminal. The terminal then provides this to the user in a playable format. The input is the generated playlist, and the output is delivered to the user as playable music data.
[0685] Furthermore, an emotion engine that estimates the user's emotions may be incorporated. That is, the identification processing unit 290 may use the emotion identification model 59 to estimate the user's emotions and perform identification processing using the user's emotions.
[0686] This invention is a system for providing a personalized music experience based on the user's emotional state. This system incorporates an emotion engine that can analyze the user's emotions in real time and select or generate music accordingly.
[0687] Specifically, the device detects the user's emotional state from their facial expressions and voice tone through sensors such as the camera and microphone. This data is analyzed by an emotion engine to identify the user's current emotional state. The acquired emotional data, along with location information and music listening history, is sent to the server.
[0688] Based on this emotional information, the server comprehensively analyzes regional music rankings and past listening data. Furthermore, it uses an AI algorithm to select music best suited to the user's emotions and, in some cases, generate new music tailored to their emotional state. This generation process takes into account the style, tempo, and genre that match the user's current emotions.
[0689] All of this information is transmitted from the server to the terminal, which then displays the selected music content on its user interface. Users can play music with simple operations and enjoy music that matches their mood at the time. For example, when a user is feeling stressed, they can choose relaxing music to promote endorphin release.
[0690] This means that music functions not only as mere entertainment, but also as an effective means of supporting the psychological state of its users.
[0691] The following describes the processing flow.
[0692] Step 1:
[0693] The device uses a camera and microphone to capture the user's facial expressions and voiceprints in real time. This collects data that indicates the user's emotional state.
[0694] Step 2:
[0695] The device inputs the collected emotional data into an emotion engine to analyze the user's emotions. The analysis results identify the emotions the user is currently experiencing (joy, sadness, stress, etc.).
[0696] Step 3:
[0697] The device sends analyzed sentiment data, user location information, and viewing history to the server. This allows the server to understand the user's comprehensive context.
[0698] Step 4:
[0699] The server uses received sentiment data and location information to refer to regional music rankings. It then analyzes user preferences in detail, along with related music data and listening history.
[0700] Step 5:
[0701] The server uses analysis results and AI algorithms to select music that best suits the user's specific emotional state. If necessary, it generates new music tailored to that emotional state.
[0702] Step 6:
[0703] The server sends selected or generated music data to the terminal, preparing to provide the user with the optimal music experience.
[0704] Step 7:
[0705] The device displays the received music data on the user interface, allowing users to easily play it.
[0706] Step 8:
[0707] Users can enjoy music that suits their mood by selecting and playing playlists and songs from the provided options. This song selection is sent back to the server as feedback to improve the accuracy of future AI analysis.
[0708] (Example 2)
[0709] Next, we will describe Example 2. In the following description, the data processing device 12 will be referred to as the "server" and the robot 414 as the "terminal".
[0710] Conventional music delivery systems only recommend music based on simple location information and listening history, making it difficult to provide a personalized music experience that takes into account the user's complex emotional state. Therefore, there is a need to develop a system that allows users to easily enjoy music that matches their mood at any given time.
[0711] The identification process performed by the identification processing unit 290 of the data processing device 12 in Example 2 is realized by the following means.
[0712] In this invention, the server includes means for acquiring data from an image acquisition device and an audio acquisition device to detect the user's emotions, means for analyzing the acquired emotion data to identify the user's emotional state, and means for selecting or generating music suitable for the user using the identified emotional state, location information, and music listening history. This makes it possible to provide music optimized for the user's emotional state.
[0713] An "image acquisition device" is a device used to acquire a user's facial expression information, and includes equipment such as a camera or other video sensor.
[0714] A "voice acquisition device" is a device used to acquire the tone and volume of a user's voice, and includes equipment such as a microphone or other voice sensors.
[0715] "Emotional data" refers to data indicating the user's emotions, acquired through image acquisition devices and voice acquisition devices, and includes analysis results based on facial expression information and voice characteristics.
[0716] "Emotional state" refers to the user's psychological state identified through the analysis of emotional data, and indicates specific emotions such as joy, sadness, anger, and relaxation.
[0717] "Location information" refers to information about the geographical location where a user is located, and is data obtained using technologies such as GPS.
[0718] "Music listening history" refers to a record of music that a user has played in the past, including detailed information such as song title, artist, genre, and playback date and time.
[0719] A "music ranking" is a list that shows the popularity or priority of music based on a specific region or set criteria.
[0720] "Means for selecting or generating music" refers to technical methods or processes that select the most suitable music for a user, or create new music as needed, based on their emotional state, location information, and music listening history.
[0721] This system is a technology designed to provide a personalized music experience based on the user's emotional state. It primarily utilizes an image acquisition device, an audio acquisition device, an emotion engine, and a music generation and selection server.
[0722] The device is equipped with a camera for image acquisition and a microphone for audio acquisition, capturing the user's facial expressions and voice in real time. The collected facial expression data and voice tone are instantly transmitted to the server.
[0723] The server analyzes the received data using an emotion engine to determine the user's current emotional state. For example, if a user is smiling and speaking in a calm voice, they are judged to be happy and relaxed. Based on this analysis, the AI algorithm selects the most suitable music by referring to the identified emotional state, location information, and music listening history.
[0724] Furthermore, it can generate new music as needed. In this case, the server utilizes AI technology to create music that best suits the selected emotional state, considering the melody, tempo, and genre. This makes it possible to provide a musical experience that perfectly matches the user's mood at that moment.
[0725] All selected and generated music is sent from the server to the terminal and displayed on the terminal's user interface. Users can play this music with simple operations and enjoy a personalized musical experience. For example, if a user wants to relax at the end of a busy day, the system will select and provide piano music at a calm tempo.
[0726] As an example of a prompt, by providing the AI model with the input, "The user is currently feeling stressed. Please suggest music that is best suited to relieve stress," it is possible to suggest music that is best suited to the user.
[0727] The flow of the specific processing in Example 2 will be explained using Figure 13.
[0728] Step 1:
[0729] The device uses a camera (image acquisition device) and a microphone (audio acquisition device) to collect the user's facial expressions and voice. These sensors detect the user's facial movements, changes in expression, and voice tone and volume. The input for this step is the user's real-time facial and voice data, and the output is the collected raw emotion data. This data is collected as foundational information necessary for the next processing step.
[0730] Step 2:
[0731] The device sends the collected emotional data to a server, where it is analyzed by an emotion engine. This analysis process uses facial recognition technology and voice characteristic analysis to identify the emotions the user is expressing. The input is data about the user's facial expressions and voice, and the output is data representing the user's emotional state. For example, if the data shows a smile and a calm voice, the user's emotional state is identified as "relaxed."
[0732] Step 3:
[0733] The server uses identified emotional states to select or generate music by referencing the user's location information and music listening history. The input consists of identified emotional states, location information, and listening history, while the output is music optimized for the user. Specifically, an AI algorithm is used to select music with a tempo and genre that matches the user's current emotions. If necessary, a generative AI model is used to generate new music.
[0734] Step 4:
[0735] The server sends the selected or generated music to the terminal. The input here is the music selected or generated in step 3, and the output is the playlist displayed on the terminal's user interface. The terminal displays this music on its screen and configures the interface to make it easy for the user to select the music.
[0736] Step 5:
[0737] The user plays the music displayed on the device with simple controls. The input in this step is the selected music, and the output is the user starting the music playback and the resulting musical experience. For example, a user can further enhance their relaxed mood by selecting and playing specific relaxation music from a music playlist.
[0738] (Application Example 2)
[0739] Next, we will explain application example 2. In the following explanation, the data processing device 12 will be referred to as the "server" and the robot 414 as the "terminal".
[0740] In modern society, entertainment that resonates with users' emotions is crucial for maintaining their psychological well-being and comfortable lives. However, existing music streaming services only offer generic music recommendations and services based on playback history, without considering the user's immediate emotional state. Therefore, there is a need for personalized services that can accurately respond to users' real-time emotions.
[0741] The specific processing performed by the specific processing unit 290 of the data processing device 12 in Application Example 2 is realized by the following means.
[0742] In this invention, the server includes an emotion engine that detects the user's emotional state, means for analyzing the user's biometric and voice information to acquire emotional data, means for analyzing the acquired emotional data, location information, music listening history, and regional music rankings to select the optimal music, and means for generating the selected music or providing it from existing music. This makes it possible to provide an optimal music experience that responds to the user's real-time emotional state.
[0743] "User" refers to an individual who uses the system to enjoy a music experience.
[0744] "Emotional state" refers to the user's psychological and emotional condition, and includes information inferred from facial expressions and tone of voice.
[0745] An "emotion engine" refers to an algorithm and system that analyzes a user's facial expressions and voice data to detect their emotional state in real time.
[0746] "Biometric information" refers to information such as the user's facial expressions and movements, which are acquired through cameras and sensors.
[0747] "Voice information" refers to information about changes in the user's voice and tone, captured by the microphone.
[0748] "Emotional data" refers to data that indicates the user's specific emotional state, obtained as a result of analysis by the emotion engine.
[0749] "Location information" refers to information that indicates the user's current geographical location, and is obtained using technologies such as GPS.
[0750] "Music listening history" refers to a record of songs and albums that a user has played so far.
[0751] "Regional music rankings" refer to information that shows the ranking of popular music within a specific region.
[0752] "Optimal music" refers to music that is considered most effective for the user, selected based on the user's emotional state, location information, and listening history.
[0753] "Means of generating music" refers to the process of composing new music using AI and algorithms, and providing a musical experience that matches the user's characteristics.
[0754] "Means of providing existing music" refers to the process of selecting the most suitable songs from existing tracks and delivering them to the user.
[0755] The system that realizes this invention mainly consists of a server, a terminal, and an emotion engine. The server analyzes the user's emotional state in real time using biometric and voice information transmitted from the user. This requires a terminal such as a smartphone equipped with a camera and microphone. The terminal analyzes monitored facial expressions using OpenCV and analyzes voice tone using Google Cloud Speech-to-Text. The results of this analysis are sent from the terminal to the server as emotion data.
[0756] The server analyzes emotional data through an emotion engine to select the most suitable music for the user. In addition to emotional data, it uses an AI algorithm (TensorFlow or PyTorch) to comprehensively analyze the user's location information, music listening history, and regional music rankings. Based on the results of this analysis and the Spotify API, it optimally selects generated or existing music and sends it to the device. This allows users to enjoy streaming music tailored to their emotional state.
[0757] For example, if analysis reveals that a user is experiencing stress, the server selects relaxing music and sends it to the device to help the user alleviate tension. This entire process is based on prompts such as: "If the user's facial expression indicates fatigue, recommend relaxing music."
[0758] This allows users to gain a musical experience that helps them better manage their emotions in their daily lives.
[0759] The flow of a specific process in Application Example 2 will be explained using Figure 14.
[0760] Step 1:
[0761] The device activates its camera and microphone to capture the user's current facial expressions and voice tone. Biometric information (facial expressions) and audio information (voice) are obtained as input. The output is the extracted biometric and audio data. This prepares the device for collecting real-time user emotion information.
[0762] Step 2:
[0763] The device analyzes captured video data using OpenCV and derives emotional states from the user's facial expressions. The input is biometric information, and the output is emotional data based on facial expressions. Specifically, it detects features such as eye and mouth movements to determine what emotions the user is experiencing.
[0764] Step 3:
[0765] The device uses Google Cloud Speech-to-Text to analyze voice data and determine emotional states from voice tone. The input is voice information, and the output is emotion data based on the voice. The user's mood is estimated by determining how excited or calm the voice is.
[0766] Step 4:
[0767] The device sends the emotion data obtained in steps 2 and 3 to the server. Location information and music listening history are also sent at this time. Inputs include emotion data, location information, and music listening history, while output is the transmission of data to the server.
[0768] Step 5:
[0769] The server analyzes received emotion data, location information, and music listening history using an AI algorithm (TensorFlow or PyTorch) to select the most suitable music for the user. The input is the transmitted data, and the output is the selected music playlist.
[0770] Step 6:
[0771] The server retrieves the selected music via the Spotify API and streams it to the device. The input is the selected music information, and the output is music playback on the device. This allows users to instantly enjoy music that matches their mood.
[0772] The specific processing unit 290 transmits the result of the specific processing to the robot 414. In the robot 414, the control unit 46A causes the speaker 240 and the controlled object 443 to output the result of the specific processing. The microphone 238 acquires audio indicating user input for the result of the specific processing. The control unit 46A transmits the audio data indicating user input acquired by the microphone 238 to the data processing unit 12. In the data processing unit 12, the specific processing unit 290 acquires the audio data.
[0773] Data generation model 58 is a type of so-called generative AI (Artificial Intelligence). One example of data generation model 58 is ChatGPT (Internet search<URL: https: / / openai.com / blog / chatgpt> ), Gemini (Internet search) <url: https: gemini.google.com ?hl="ja">Examples of generative AI include the following. The data generation model 58 is obtained by performing deep learning on a neural network. The data generation model 58 is input with prompts containing instructions, and with inference data such as audio data representing speech, text data representing text, and image data representing images. The data generation model 58 infers from the input inference data according to the instructions indicated by the prompts, and outputs the inference results in data formats such as audio data and text data. Here, inference refers to, for example, analysis, classification, prediction, and / or summarization.
[0774] In the above embodiment, an example was given in which the specific processing is performed by the data processing device 12, but the technology of this disclosure is not limited thereto, and the specific processing may also be performed by the robot 414.
[0775] Furthermore, the emotion identification model 59, acting as an emotion engine, may determine the user's emotion according to a specific mapping. Specifically, the emotion identification model 59 may determine the user's emotion according to a specific mapping, which is an emotion map (see Figure 9). Similarly, the emotion identification model 59 may also determine the robot's emotion, and the identification processing unit 290 may perform identification processing using the robot's emotion.
[0776] Figure 9 shows an emotion map 400 in which multiple emotions are mapped. In the emotion map 400, emotions are arranged in concentric circles radiating from the center. The closer to the center of the concentric circles, the more primitive the emotions are located. Further out of the concentric circles, emotions representing states and actions arising from mental states are located. Emotion is a concept that includes feelings and mental states. On the left side of the concentric circles, emotions that are generally generated from reactions occurring in the brain are located. On the right side of the concentric circles, emotions that are generally induced by situational judgment are located. Above and below the concentric circles, emotions that are generally generated from reactions occurring in the brain and induced by situational judgment are located. In addition, the emotion of "pleasure" is located on the upper side of the concentric circles, and the emotion of "displeasure" is located on the lower side. Thus, in the emotion map 400, multiple emotions are mapped based on the structure in which emotions arise, and emotions that are likely to occur simultaneously are mapped close together.
[0777] These emotions are distributed at the 3 o'clock position on the Emotion Map 400, and usually fluctuate between feelings of security and anxiety. In the right half of the Emotion Map 400, situational awareness takes precedence over internal feelings, resulting in a calm impression.
[0778] The inside of the Emotion Map 400 represents inner thoughts, while the outside represents actions. Therefore, the further you go from the outside of the Emotion Map 400, the more visible (expressed in actions) your emotions become.
[0779] Here, human emotions are based on various balances, such as posture and blood sugar levels. When these balances deviate from the ideal, it results in discomfort, and when they approach the ideal, it results in pleasure. Similarly, in robots, cars, motorcycles, etc., emotions can be created based on various balances, such as posture and battery level. When these balances deviate from the ideal, it results in discomfort, and when they approach the ideal, it results in pleasure. The emotion map can be generated, for example, based on Dr. Mitsuyoshi's emotion map (Research on a system for analyzing brain physiological signals of speech emotion recognition and emotion, Tokushima University, doctoral dissertation: https: / / ci.nii.ac.jp / naid / 500000375379). The left half of the emotion map contains emotions belonging to a region called "response," where sensation is dominant. The right half of the emotion map contains emotions belonging to a region called "situation," where situational awareness is dominant.
[0780] The emotion map defines two emotions that promote learning. One is the emotion around the middle of the negative "repentance" and "reflection" on the situation side. In other words, it is when the robot experiences negative emotions such as "I never want to feel this way again" or "I don't want to be scolded again." The other is the emotion around the positive "desire" on the reaction side. In other words, it is when the robot has positive feelings such as "I want more" or "I want to know more."
[0781] The emotion identification model 59 inputs user input into a pre-trained neural network, obtains emotion values representing each emotion shown in the emotion map 400, and determines the user's emotion. This neural network is pre-trained based on multiple training data sets, which are combinations of user input and emotion values representing each emotion shown in the emotion map 400. Furthermore, this neural network is trained so that emotions located close together have similar values, as shown in the emotion map 900 in Figure 10. Figure 10 shows an example where multiple emotions such as "reassured," "calm," and "confident" have similar emotion values.
[0782] The above description primarily focuses on the functions of the data processing device 12 in relation to this disclosure. However, the system related to this disclosure is not necessarily implemented on a server. The system related to this disclosure may be implemented as a general information processing system. This disclosure may be implemented, for example, as a software program that runs on a personal computer or as an application that runs on a smartphone. The method related to this disclosure may be provided to users in SaaS (Software as a Service) format.
[0783] In the above embodiment, an example was given in which a specific process is performed by a single computer 22. However, the technology of this disclosure is not limited thereto, and a distributed processing of the specific process may be performed by multiple computers, including computer 22. For example, a data generation model 58 may be provided in an external device of the data processing device 12, and the external device may generate data according to the input data.
[0784] In the above embodiment, an example was given in which the specific processing program 56 is stored in the storage 32, but the technology of this disclosure is not limited thereto. For example, the specific processing program 56 may be stored in a portable, computer-readable, non-temporary storage medium such as a USB (Universal Serial Bus) memory. The specific processing program 56 stored in the non-temporary storage medium is installed in the computer 22 of the data processing device 12. The processor 28 executes specific processing according to the specific processing program 56.
[0785] Alternatively, the specific processing program 56 may be stored in a storage device such as a server connected to the data processing device 12 via the network 54, and the specific processing program 56 may be downloaded and installed on the computer 22 in response to a request from the data processing device 12.
[0786] Furthermore, it is not necessary to store the entirety of the specific processing program 56 in a storage device such as a server connected to the data processing device 12 via the network 54, or to store the entirety of the specific processing program 56 in the storage 32; it is acceptable to store only a portion of the specific processing program 56.
[0787] The following types of processors can be used as hardware resources to perform specific processing. Examples of processors include a CPU, a general-purpose processor that functions as a hardware resource to perform specific processing by executing software, i.e., a program. Other examples of processors include dedicated electrical circuits, such as FPGAs (Field-Programmable Gate Arrays), PLDs (Programmable Logic Devices), or ASICs (Application Specific Integrated Circuits), which have circuit configurations specifically designed to perform specific processing. All of these processors have built-in or connected memory, and all of them perform specific processing by using memory.
[0788] The hardware resource that performs a specific process may consist of one of these various processors, or it may consist of a combination of two or more processors of the same or different types (for example, a combination of multiple FPGAs, or a combination of a CPU and an FPGA). Alternatively, the hardware resource that performs a specific process may consist of a single processor.
[0789] Examples of configurations using a single processor include, firstly, a configuration in which one or more CPUs and software are combined to form a single processor, and this processor functions as a hardware resource that performs a specific process. Secondly, there is a configuration using a processor that realizes the functions of the entire system, including multiple hardware resources that perform a specific process, on a single IC chip, as exemplified by SoCs (System-on-a-chip). In this way, a specific process is realized using one or more of the above types of processors as hardware resources.
[0790] Furthermore, the hardware structure of these various processors can more specifically utilize electrical circuits that combine circuit elements such as semiconductor devices. Also, the specific processing described above is merely an example. Therefore, it goes without saying that unnecessary steps can be deleted, new steps added, or the processing order rearranged, as long as it does not deviate from the main purpose.
[0791] The descriptions and illustrations presented above are detailed explanations of the technical aspects of this disclosure and are merely examples of the technical aspects. For example, the above descriptions of the structure, function, operation, and effect are examples of the structure, function, operation, and effect of the technical aspects of this disclosure. Therefore, it goes without saying that you may delete unnecessary parts, add new elements, or replace elements in the descriptions and illustrations presented above, as long as you do not deviate from the essence of the technical aspects of this disclosure. Furthermore, in order to avoid confusion and facilitate understanding of the technical aspects of this disclosure, explanations of common technical knowledge and the like that do not require special explanation to enable the implementation of the technical aspects of this disclosure have been omitted from the descriptions and illustrations presented above.
[0792] All documents, patent applications, and technical standards described herein are incorporated by reference to the same extent as if each individual document, patent application, and technical standard were specifically and individually noted as being incorporated by reference.
[0793] The following is further disclosed regarding the embodiments described above.
[0794] (Claim 1)
[0795] A means of obtaining the user's location information,
[0796] A means of obtaining a user's music listening history,
[0797] A method for obtaining regional music rankings based on location information,
[0798] A method for selecting music suitable for the user by analyzing acquired location information, music listening history, and local music rankings,
[0799] Means of providing selected music to users,
[0800] A system that includes this.
[0801] (Claim 2)
[0802] The system according to claim 1, further comprising means for obtaining the user's current time information and selecting music based on that time information.
[0803] (Claim 3)
[0804] The system according to claim 1, further comprising means for generating new music based on region and time.
[0805] "Example 1"
[0806] (Claim 1)
[0807] A means of obtaining the user's location information,
[0808] A means of obtaining a user's music listening history,
[0809] A method for obtaining regional music rankings based on location information,
[0810] A method for selecting music suitable for the user by analyzing acquired location information, music listening history, and local music rankings,
[0811] A means of generating music using an AI algorithm,
[0812] Means for providing selected or generated music to users,
[0813] A system that includes this.
[0814] (Claim 2)
[0815] The system according to claim 1, further comprising means for obtaining the user's current time information and selecting music based on that time information.
[0816] (Claim 3)
[0817] The system according to claim 1, further comprising means for generating new music based on region and time.
[0818] "Application Example 1"
[0819] (Claim 1)
[0820] A means of obtaining the user's location information,
[0821] A means of obtaining a user's music listening history,
[0822] A method for obtaining regional music rankings based on location information,
[0823] A method for selecting music suitable for the user by analyzing acquired location information, music listening history, and local music rankings,
[0824] Means of providing selected music to users,
[0825] A means of dynamically generating music playlists that take into account the user's movement patterns,
[0826] A system that includes this.
[0827] (Claim 2)
[0828] The system according to claim 1, further comprising means for obtaining the user's current time information and selecting music based on that time information.
[0829] (Claim 3)
[0830] The system according to claim 1, further comprising means for generating new music based on region and time.
[0831] "Example 2 of combining an emotion engine"
[0832] (Claim 1)
[0833] A means for acquiring data from an image acquisition device and an audio acquisition device in order to detect the user's emotions,
[0834] A means of identifying the user's emotional state by analyzing acquired emotional data,
[0835] A means for selecting or generating music suitable for a user using identified emotional states, location information, and music listening history,
[0836] Means for providing selected or generated music to users,
[0837] A system that includes this.
[0838] (Claim 2)
[0839] The system according to claim 1, further comprising means for obtaining the user's current time information and selecting music based on that time information in addition to the user's emotional state.
[0840] (Claim 3)
[0841] The system according to claim 1, further comprising means for generating new music based on the emotional state of the user.
[0842] "Application example 2 when combining with an emotional engine"
[0843] (Claim 1)
[0844] A system equipped with an emotion engine that detects the user's emotional state, and means for acquiring emotional data by analyzing the user's biometric and voice information,
[0845] A method for selecting the optimal music by analyzing acquired emotional data, location information, music listening history, and regional music rankings,
[0846] A means of generating selected music or providing it from existing music,
[0847] A system that includes this.
[0848] (Claim 2)
[0849] The system according to claim 1, further comprising means for selecting or generating music appropriate to the user's emotional state based on the user's current time information.
[0850] (Claim 3)
[0851] The system according to claim 1, further comprising means for generating new music based on region, time, and the emotional state of the user. [Explanation of Symbols]
[0852] 10, 210, 310, 410 Data Processing Systems 12 Data Processing Devices 14 Smart Devices 214 Smart Glasses 314 Headset-type terminal 414 Robots< / url:> < / url:> < / url:> < / url:>
Claims
1. A means of obtaining the user's location information, A means of obtaining a user's music listening history, A method for obtaining regional music rankings based on location information, A method for selecting music suitable for the user by analyzing acquired location information, music listening history, and local music rankings, Means of providing selected music to users, A system that includes this.
2. The system according to claim 1, further comprising means for obtaining the user's current time information and selecting music based on that time information.
3. The system according to claim 1, further comprising means for generating new music based on region and time.
Citation Information
Patent Citations
Persona chatbot control method and system
JP2022180282A