System

A system that collects and analyzes users' music history, usage patterns, and emotional states to generate personalized playlists, continuously learning from feedback, addresses the limitations of conventional systems by offering tailored music experiences.

JP2026022489APending Publication Date: 2026-02-12SOFTBANK GROUP CORP
View PDF 1 Cites 0 Cited by

Patent Information

Application Number
JP2024124006
Authority / Receiving Office
JP · JP
Patent Type
Applications
Current Assignee / Owner
Filing Date
2024-07-30
Publication Date
2026-02-12

AI Technical Summary

Technical Problem

Conventional music recommendation systems fail to provide personalized music experiences that consider users' specific usage situations and emotions, lacking the ability to efficiently learn from user feedback and adapt recommendations accordingly.

Method used

A system that collects and analyzes users' playback history, usage time periods, activity patterns, and emotional states, generates personalized playlists, and continuously learns from user feedback to improve music recommendations.

Benefits of technology

Provides a highly personalized music experience that meets users' needs by optimizing playlists based on their preferences and behavioral patterns, enhancing user satisfaction.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure 2026022489000001_ABST
    Figure 2026022489000001_ABST
Patent Text Reader

Abstract

A system is provided.SOLUTION: A system comprising: means for collecting a reproduction history, a use time zone, an activity pattern, and an emotional situation of a user; means for analyzing the collected data and extracting a music preference and a behavior pattern of the user; means for generating a playlist based on an analysis result; means for providing the generated playlist to a user terminal; and means for collecting feedback from the user and performing learning of the system based on the feedback.SELECTED DRAWING: Figure 1
Need to check novelty before this filing date? Find Prior Art

Description

[Technical Field]

[0001] The technology of the present disclosure relates to a system. [Background technology]

[0002] Patent document 1 discloses a persona chatbot control method performed by at least one processor, the method including the steps of receiving a user utterance, adding the user utterance to a prompt including an instruction sentence related to a description of the chatbot character, encoding the prompt, and inputting the encoded prompt into a language model to generate a chatbot utterance in response to the user utterance. [Prior art documents] [Patent documents]

[0003] [Patent Document 1] Japanese Patent Publication No. 2022-180282 Summary of the Invention [Problem to be solved by the invention]

[0004] The streaming music market is expanding rapidly, but an increasing number of users are seeking personalized music experiences that meet their diverse needs. Individuals with busy daily lives and those who want to relax or concentrate through music face challenges, particularly in finding music that suits them and discovering new music. Conventional methods have struggled to efficiently meet these needs, necessitating the need for new methods to improve user satisfaction. [Means for solving the problem]

[0005] To solve the above-mentioned problems, we propose the following means as claimed in the claims. Specifically, we provide a system including: means for collecting a user's playback history, usage time periods, activity patterns, and emotional state; means for analyzing the collected data and extracting the user's music preferences and behavioral patterns; means for generating a playlist based on the analysis results; means for providing the generated playlist to a user terminal; and means for collecting user feedback and training the system based on the feedback. Furthermore, we can also use means for collecting the user's emotional state in tag form and reflecting it in playlist generation, or for analyzing the user's activity patterns using a clustering algorithm and using it in playlist generation. This allows us to provide a personalized music experience that meets the user's needs and improve user satisfaction.

[0006] "User" refers to an individual who uses the music streaming service.

[0007] "Playback history" refers to data relating to songs that a user has played in the past and the dates and times they were played.

[0008] "Time period of use" refers to information about music playing activities within a specific time range.

[0009] "Activity pattern" refers to information about the user's behavior when playing music (for example, while commuting, playing sports, etc.).

[0010] The "emotional state" refers to information that expresses, in tag form or other form, the emotion (for example, happy, sad, relaxed, etc.) that the user feels when listening to music.

[0011] "Feedback" refers to information regarding ratings and actions (e.g., skip, complete, like, dislike, etc.) provided by users in response to playback of a song.

[0012] "Analysis" refers to the processing of data and the execution of algorithms to process collected data and identify users' musical preferences and behavioral patterns.

[0013] "Playlist" means a list of songs provided to a user, which is a collection of songs organized according to certain criteria.

[0014] "Clustering algorithm" refers to a machine learning technique that groups data based on common characteristics.

[0015] "Personalization" refers to customizing a system based on individual user preferences and behavior. [Brief explanation of the drawings]

[0016] [Figure 1] 1 is a conceptual diagram showing an example of the configuration of a data processing system according to a first embodiment. [Figure 2] 1 is a conceptual diagram showing an example of main functions of a data processing device and a smart device according to a first embodiment. [Figure 3] FIG. 10 is a conceptual diagram showing an example of the configuration of a data processing system according to a second embodiment. [Figure 4] FIG. 10 is a conceptual diagram showing an example of main functions of a data processing device and smart glasses according to a second embodiment. [Figure 5] FIG. 10 is a conceptual diagram showing an example of the configuration of a data processing system according to a third embodiment. [Figure 6] FIG. 11 is a conceptual diagram showing an example of main functions of a data processing device and a headset-type terminal according to a third embodiment. [Figure 7] FIG. 10 is a conceptual diagram showing an example of the configuration of a data processing system according to a fourth embodiment. [Figure 8] FIG. 10 is a conceptual diagram showing an example of main functions of a data processing device and a robot according to a fourth embodiment. [Figure 9] 1 shows an emotion map onto which multiple emotions are mapped. [Figure 10]1 shows an emotion map onto which multiple emotions are mapped. [Figure 11] FIG. 3 is a sequence diagram showing a processing flow of the data processing system according to the first embodiment. [Figure 12] FIG. 10 is a sequence diagram showing the flow of processing in the data processing system in Application Example 1. [Figure 13] FIG. 10 is a sequence diagram showing the flow of processing in the data processing system according to the second embodiment when an emotion engine is combined. [Figure 14] FIG. 10 is a sequence diagram showing the flow of processing in the data processing system in Application Example 2 when an emotion engine is combined. DETAILED DESCRIPTION OF THE INVENTION

[0017] An example of an embodiment of a system according to the technology of the present disclosure will be described below with reference to the accompanying drawings.

[0018] First, the terms used in the following description will be explained.

[0019] In the following embodiments, a coded processor (hereinafter simply referred to as a "processor") may be a single arithmetic device or a combination of multiple arithmetic devices. Furthermore, a processor may be a single type of arithmetic device or a combination of multiple types of arithmetic devices. Examples of arithmetic devices include a CPU (Central Processing Unit), a GPU (Graphics Processing Unit), a GPGPU (General-Purpose computing on Graphics Processing Units), and an APU (Accelerated Processing Unit).

[0020] In the following embodiments, a coded RAM (Random Access Memory) is a memory in which information is temporarily stored and is used as a working memory by a processor.

[0021] In the following embodiments, the coded storage is one or more non-volatile storage devices that store various programs, various parameters, etc. Examples of non-volatile storage devices include flash memory (SSD (Solid State Drive)), magnetic disks (e.g., hard disks), and magnetic tapes.

[0022] In the following embodiments, a communication I / F (Interface) with a symbol is an interface including a communication processor, an antenna, etc. The communication I / F controls communication between multiple computers. Examples of communication standards applied to the communication I / F include wireless communication standards including 5G (5th Generation Mobile Communication System), Wi-Fi (registered trademark), Bluetooth (registered trademark), etc.

[0023] In the following embodiments, "A and / or B" is synonymous with "at least one of A and B." In other words, "A and / or B" means that it may be only A, only B, or a combination of A and B. Furthermore, in this specification, the same concept as "A and / or B" is also applied when three or more things are expressed connected by "and / or."

[0024] [First embodiment]

[0025] FIG. 1 shows an example of the configuration of a data processing system 10 according to the first embodiment.

[0026] 1, a data processing system 10 includes a data processing device 12 and a smart device 14. An example of the data processing device 12 is a server.

[0027] The data processing device 12 includes a computer 22, a database 24, and a communication I / F 26. The computer 22 is an example of a "computer" according to the technology of the present disclosure. The computer 22 includes a processor 28, a RAM 30, and a storage 32. The processor 28, the RAM 30, and the storage 32 are connected to a bus 34. The database 24 and the communication I / F 26 are also connected to the bus 34. The communication I / F 26 is connected to a network 54. Examples of the network 54 include a WAN (Wide Area Network) and / or a LAN (Local Area Network).

[0028] The smart device 14 includes a computer 36, a reception device 38, an output device 40, a camera 42, and a communication I / F 44. The computer 36 includes a processor 46, a RAM 48, and a storage 50. The processor 46, the RAM 48, and the storage 50 are connected to a bus 52. The reception device 38, the output device 40, and the camera 42 are also connected to the bus 52.

[0029] The reception device 38 includes a touch panel 38A, a microphone 38B, and the like, and receives user input. The touch panel 38A detects contact with an indicator (for example, a pen or a finger) to receive user input by the touch of the indicator. The microphone 38B detects the user's voice to receive user input by voice. The control unit 46A transmits data indicating the user input received by the touch panel 38A and the microphone 38B to the data processing device 12. In the data processing device 12, the specific processing unit 290 acquires the data indicating the user input.

[0030] The output device 40 includes a display 40A and a speaker 40B, and presents data to the user 20 by outputting the data in a form of expression that the user 20 can perceive (for example, audio and / or text). The display 40A displays visible information such as text and images in accordance with instructions from the processor 46. The speaker 40B outputs audio in accordance with instructions from the processor 46. The camera 42 is a compact digital camera equipped with an optical system including a lens, aperture, and shutter, and an imaging element such as a CMOS (Complementary Metal-Oxide-Semiconductor) image sensor or a CCD (Charge Coupled Device) image sensor.

[0031] The communication I / F 44 is connected to a network 54. The communication I / Fs 44 and 26 control the exchange of various information between the processor 46 and the processor 28 via the network 54.

[0032] FIG. 2 shows an example of the main functions of the data processing device 12 and the smart device 14.

[0033] 2, in the data processing device 12, a specific process is performed by the processor 28. A specific processing program 56 is stored in the storage 32. The specific processing program 56 is an example of a "program" according to the technology of the present disclosure. The processor 28 reads the specific processing program 56 from the storage 32 and executes the read specific processing program 56 on the RAM 30. The specific process is realized by the processor 28 operating as a specific processing unit 290 in accordance with the specific processing program 56 executed on the RAM 30.

[0034] The storage 32 stores a data generation model 58 and an emotion identification model 59. The data generation model 58 and the emotion identification model 59 are used by the identification processing unit 290.

[0035] In the smart device 14, the processor 46 performs the reception output process. The storage 50 stores a reception output program 60. The reception output program 60 is used in conjunction with the specific processing program 56 by the data processing system 10. The processor 46 reads the reception output program 60 from the storage 50 and executes the read reception output program 60 on the RAM 48. The reception output process is realized by the processor 46 operating as the control unit 46A in accordance with the reception output program 60 executed on the RAM 48.

[0036] Next, a description will be given of the specific processing performed by the specific processing unit 290 of the data processing device 12. In the following description, the data processing device 12 will be referred to as a "server" and the smart device 14 will be referred to as a "terminal."

[0037] This invention relates to a system that personalizes and optimizes a user's music experience. Specifically, it is a method for generating and providing playlists based on the user's preferences and behavioral patterns by collecting and analyzing the user's playback history, usage time period, activity patterns, and emotional state. Furthermore, by collecting user feedback and allowing the system to continuously learn, it provides a more accurate and personalized music experience.

[0038] User Data Collection

[0039] A user starts a music application and plays music in the course of daily life.

[0040] The device collects the user's music history, the start and end times of the music, whether the music was skipped, the user's current activity status (e.g., commuting, exercising), and emotion tags, which may be entered by the user in the form of "fun" or "want to relax."

[0041] The terminal converts this data into a specified format and transmits it to the server using a secure communication method.

[0042] Data analysis and learning

[0043] The server stores the received user data in a database.

[0044] The server analyzes the stored data and extracts features to identify users' musical preferences and behavioral patterns. This analysis uses machine learning algorithms. For example, a clustering algorithm can group users' activity patterns.

[0045] The server performs statistical analysis based on a large amount of user data and also learns general trends.

[0046] Playlist Generation

[0047] The server then generates the optimal playlist for a specific user based on the analysis results, which means selecting songs of different genres and moods according to the user's time of day, activity patterns, and emotional tags.

[0048] The server packages the list of songs included in the generated playlist and its associated data (song ID, playback URL, etc.) and sends it to the terminal.

[0049] Providing playlists

[0050] The terminal displays the received playlist on the application's user interface for easy access by the user.

[0051] The user selects a song from the playlist and plays it. When a playback command is given, the terminal streams the song over the Internet.

[0052] Gather feedback and continue learning

[0053] The device collects the songs the user has played, whether they have completed or skipped, and their ratings and feedback (e.g., "favorite marks") for specific songs.

[0054] The terminal transmits this feedback information to the server.

[0055] The server updates the user's preferences based on the received feedback and then reruns the analysis and learning process, which is then reflected in the next playlist generation.

[0056] As a concrete example, if a user sets the emotion tag "busy" during their morning commute, the server will select energetic, fast-paced music for that time period based on past data and learning results. The generated playlist is provided to the device and played by the user. On the other hand, when it's time to relax in the evening, the server will generate a playlist containing calming music and provide it to the user via the device. In this way, it is possible to provide a music experience that is optimal for the user's situation and emotion.

[0057] The processing flow will be explained below.

[0058] Step 1:

[0059] A user launches a music application and starts playing a song. At the same time, the user sets an activity state (e.g., "commuting" or "playing sports") and an emotion tag (e.g., "having fun" or "wanting to relax").

[0060] Step 2:

[0061] The device automatically collects information about the song the user is playing (playback start time, play end time, song ID, etc.), as well as the user's activity status and emotion tag. The collected data is temporarily stored on the device.

[0062] Step 3:

[0063] The device formats the collected user data into a specific format, encrypts it, and then sends it to a server using a secure protocol such as HTTPS.

[0064] Step 4:

[0065] The server stores the received user data in a database, and then starts a process to analyze the stored data. This analysis uses machine learning algorithms.

[0066] Step 5:

[0067] The server extracts features from the analyzed data to identify users' musical preferences and behavioral patterns, for example by using a clustering algorithm to group users based on their activity patterns and emotions.

[0068] Step 6:

[0069] The server selects music suitable for a specific time period or activity level and generates an optimal playlist for each user, for example, energetic music for the commute and quieter music for relaxation.

[0070] Step 7:

[0071] The server packages the generated playlist information (song ID, playback URL, etc.) and sends it to the device. This communication also uses a secure protocol.

[0072] Step 8:

[0073] The device displays the received playlist in the user interface of the music application, allowing the user to select and play it.

[0074] Step 9:

[0075] The user selects a song from the presented playlist and starts playing it. The device streams the song over the Internet.

[0076] Step 10:

[0077] The device collects the user's playback behavior (skip songs, complete playback, rating, etc.) and records it as feedback.

[0078] Step 11:

[0079] The terminal transmits the collected feedback information to the server.

[0080] Step 12:

[0081] The server analyzes the received feedback and uses it as data to update the user's preferences, which allows for more accurate personalization when generating the next playlist.

[0082] Example 1

[0083] Next, a description will be given of Example 1. In the following description, the data processing device 12 will be referred to as a "server" and the smart device 14 will be referred to as a "terminal."

[0084] Conventional music recommendation systems have difficulty generating playlists that fully consider a user's specific usage situation and emotions. Furthermore, they lack a system that can properly learn from user feedback and reflect it in future recommendations. As a result, users are unable to obtain the music experience that best suits their preferences and situation.

[0085] The specific processing by the specific processing unit 290 of the data processing device 12 in the first embodiment is realized by the following means.

[0086] In this invention, the server includes: means for collecting a user's playback history, usage time periods, activity patterns, and emotional state; means for formatting the collected data into a specified format and transmitting it to the server using a secure communication method; means for saving the received user data, analyzing it using a machine learning algorithm, and extracting the user's music preferences and behavioral patterns; means for generating an optimal playlist for a specific user based on the analysis results; means for packaging a list of songs included in the generated playlist and their associated data and transmitting it to a user terminal; means for displaying the received playlist on a user interface for easy user access; and means for collecting user feedback and re-executing the analysis and learning process based on the feedback, thereby enabling a more personalized music experience based on the user's usage history and emotional tags.

[0087] "Playback history" is information about songs that the user has played in the past, and includes song IDs, titles, artist names, and the like.

[0088] The "usage time period" is information about the date and time when the user used the music application, including the playback start time and end time.

[0089] The "activity pattern" refers to the activity status of the user while playing music, and indicates a specific operating state such as commuting, exercising, or relaxing.

[0090] The "emotional state" refers to the emotional or mood state of the user when listening to music, and is expressed in the form of emotion tags such as "fun" or "want to relax."

[0091] A "secure communication method" is a communication protocol that ensures safety when sending and receiving data, and a specific example is HTTPS.

[0092] "Machine learning algorithms" are algorithms that analyze large amounts of data and extract patterns and rules, and include clustering algorithms.

[0093] A "playlist" is a collection of multiple songs, and is a format for providing a collection of songs that meets the user's playback needs.

[0094] A "user interface" is a display screen that allows a user to directly interact with a system or application, and includes a song list, operation buttons, and the like.

[0095] "Feedback" refers to information provided by users to the system, such as playback history, song ratings and comments, which is reflected in the system's learning process.

[0096] The "analysis and learning process" is a series of steps that the system takes to learn the user's preferences and behavioral patterns based on data collected from the user, and then generate a new playlist.

[0097] This invention relates to a system that personalizes and optimizes a user's music experience. This system collects and analyzes a user's playback history, usage time, activity patterns, and emotional state, and generates and provides playlists based on the user's preferences and behavioral patterns. Furthermore, by collecting user feedback and continuously learning, the system provides an even more accurate and personalized music experience.

[0098] First, a user launches a music application and plays a song in their daily life. For example, a user presses the play button on the application to play a song. Here, the device collects the history of songs played by the user, the start and end times of playback, whether songs were skipped, the user's current activity state (commuting, exercising, etc.), and emotion tags (e.g., happy, want to relax). These emotion tags are input by the user.

[0099] The device then formats the collected data into a specified format and sends it to the server using HTTPS, a secure communication method. At this time, the data is converted into JSON format for efficient data processing.

[0100] The server stores the received user data in a database. For example, the server inserts data into the database using SQL queries. The server then analyzes the stored data using machine learning algorithms to identify users' musical preferences and behavioral patterns. This analysis uses the Python library scikit-learn, which runs a clustering algorithm to extract specific patterns.

[0101] Next, the server generates a playlist optimized for the specific user based on the analysis results. This means selecting songs of different genres and moods depending on the user's time of day, activity patterns, and emotional tags. The generated playlist is converted into JSON format and sent back to the device using the HTTPS protocol.

[0102] The device displays the received playlist on the application's user interface for easy access by the user. For example, the user can click a song in the playlist and press the play button to start streaming the song. Specifically, the device connects to the streaming server and streams the music using the playback URL.

[0103] Additionally, when a user provides feedback such as which songs they have completed playing, which songs they have skipped, or their ratings for specific songs, the device collects this feedback information and sends it to the server. The server then updates the user's preferences based on the received feedback and reruns the analysis and learning process, allowing the information to be reflected in the next playlist generation.

[0104] As a concrete example, if a user sets the emotion tag "busy" during their morning commute, the server will select energetic, fast-paced music for that time period based on past data and learning results. The generated playlist is provided to the device and played by the user. On the other hand, when it's time to relax in the evening, the server will generate a playlist containing calming music and provide it to the user via the device. In this way, it is possible to provide a music experience that is optimal for the user's situation and emotion.

[0105] Examples of prompt statements

[0106] "Please explain in detail the data analysis process that goes into generating an energizing playlist for your morning commute."

[0107] "How do I update my personalized music playlist based on user feedback?"

[0108] "Please explain in detail how you would generate an evening playlist based on the emotion tag 'I want to relax'."

[0109] The flow of the identification process in the first embodiment will be described with reference to FIG.

[0110] Step 1: Collect user data

[0111] In everyday life, users launch music applications and play music.

[0112] Specifically, the user presses the play button on the application on a device such as a smartphone or computer.

[0113] The device collects the user's history of songs played, the start and end times of playback, whether songs were skipped, the user's current activity status (commuting, exercising, etc.), and emotion tags (e.g., fun, want to relax).

[0114] The input is various data when the user operates the playlist.

[0115] As part of the data processing, the terminal records these data in its internal memory.

[0116] The output is formatted data, ready to be sent to the next step.

[0117] Step 2: Sending data

[0118] The terminal formats the collected data into the specified format.

[0119] The input is raw, unformatted data.

[0120] To process the data, the terminal converts the data into JSON format.

[0121] The output is data in JSON format.

[0122] The terminal transmits data to the server using a secure communication method (e.g., HTTPS).

[0123] Specifically, the terminal sends JSON format data to the server as a POST request using the HTTPS protocol.

[0124] Step 3: Data storage and analysis

[0125] The server stores the received user data in a database.

[0126] The input is JSON format data sent from the terminal.

[0127] For data processing, the server receives this data and inserts it into a database using an SQL query.

[0128] The output is user data stored in a database.

[0129] The server analyzes the stored data using machine learning algorithms (e.g., clustering algorithms).

[0130] The input is user data stored in a database.

[0131] For data calculation, the server applies machine learning models using Python libraries (e.g., scikit-learn) to identify users' musical preferences and behavioral patterns.

[0132] The output is the user's characteristics and pattern information as the analysis results.

[0133] Step 4: Generate a playlist

[0134] The server generates an optimal playlist for a particular user based on the analysis results.

[0135] The input is the user's characteristics and pattern information as the analysis result.

[0136] To process the data, the server narrows down the conditions based on the time of use, activity patterns, and emotion tags, and selects appropriate songs from the music database.

[0137] The output is the generated playlist.

[0138] The server packages the list of songs included in the generated playlist and its associated data (song ID, playback URL, etc.) and sends it to the terminal.

[0139] Specifically, the server converts the playlist into JSON format and sends a POST request to the terminal using the HTTPS protocol.

[0140] Step 5: Serve the playlist

[0141] The terminal displays the received playlist on the user interface of the application, making it easily accessible to the user.

[0142] The input is the playlist data in JSON format sent from the server.

[0143] For data processing, the device parses the JSON data and dynamically displays a song list in the UI component.

[0144] The output is the displayed playlist.

[0145] The user selects a song from the presented playlist and plays it.

[0146] Specifically, the user clicks on a song in the playlist and presses the play button.

[0147] The device streams the music over the Internet.

[0148] The input is the URL to play the song.

[0149] For data processing, the terminal connects to a streaming server and plays music.

[0150] Step 6: Gather feedback and continue learning

[0151] The device collects songs that the user has completed playing, songs that the user has skipped, and ratings and feedback for specific songs.

[0152] The input is the feedback information provided by the user.

[0153] As for data processing, the terminal tracks the user's operations and stores them in data storage.

[0154] The output is the collected feedback data.

[0155] The terminal transmits this feedback information to the server.

[0156] Specifically, the terminal formats the feedback data into JSON format and sends a POST request to the server using the HTTPS protocol.

[0157] The server updates the user's preferences based on the received feedback and then reruns the analysis and learning process.

[0158] The input is the feedback data sent from the terminal.

[0159] As the data is processed, the server stores it in a database and retrains the machine learning model.

[0160] The output is updated analysis results and data that will be reflected in the next playlist generation.

[0161] (Application example 1)

[0162] Next, a description will be given of Application Example 1. In the following description, the data processing device 12 will be referred to as a "server" and the smart device 14 will be referred to as a "terminal."

[0163] Conventionally, methods for optimizing the efficiency and accuracy of factory robots have relied solely on motion algorithms and machine learning. However, it is believed that even more efficient work can be achieved by providing an appropriate environment based on the robot's activity status and work content. The present invention aims to improve the efficiency and accuracy of work by playing personalized music while the robot is working.

[0164] The specific processing by the specific processing unit 290 of the data processing device 12 in the application example 1 is realized by the following means.

[0165] In this invention, the server includes means for collecting a user's playback history, usage time periods, activity patterns, and emotional state, means for analyzing the collected data and extracting the user's preferences and behavioral patterns, means for generating a playlist based on the analysis results, means for providing the generated playlist to the terminal, means for collecting the terminal's work history and activity level and generating an optimal music experience based on the data, means for collecting feedback and conducting system training, and means for generating applications for use in various facilities. This makes it possible to play optimal music according to the robot's activity status and work content, effectively improving work efficiency and accuracy.

[0166] "Playback history" is a record of music and content that a user has played in the past.

[0167] A "time period of use" refers to a specific time range during which a user uses music or content.

[0168] An "activity pattern" is a series of patterns of actions or states of a user or a terminal.

[0169] "Emotional state" refers to the user's current emotional or mood state.

[0170] "Analysis" is the process of examining collected data in detail to identify trends and characteristics.

[0171] A "playlist" refers to a list of songs or content to be played.

[0172] A "terminal" refers to an electronic device used by a user, and includes, for example, a smartphone or a robot.

[0173] "Work history" is a record of work performed by a terminal or robot.

[0174] "Activity level" is a measure that indicates the activity and intensity of the terminal or robot's movements.

[0175] "Feedback" refers to responses or opinions from users or systems.

[0176] A "clustering algorithm" is a machine learning technique for classifying data into several groups.

[0177] "System learning" is the process by which a system improves itself based on collected data and feedback.

[0178] An "application" is software or a program that provides a specific function.

[0179] This invention is a system that automatically plays appropriate music to improve the work efficiency and accuracy of robots working in factories. The system collects the user's playback history, usage time, activity patterns, and emotional state, and generates and provides a playlist based on the analysis results. It also has a function that collects the device's work history and activity level to generate optimal music.

[0180] Generating a Program

[0181] The system program includes the following main functions:

[0182] 1. Data Collection:

[0183] Sensors are used to collect records of tasks performed by the device, work time, and activity levels.

[0184] The user's emotional state and usage time period are also collected (for example, music playback history while working).

[0185] 2. Data Analysis:

[0186] The collected data is stored and analyzed on the server using Python.

[0187] The robot's activity patterns are classified using scikit-learn's clustering algorithm (KMeans).

[0188] Features are extracted from the data and music appropriate for the robot's task is selected.

[0189] 3. Playlist generation:

[0190] Generate the best playlist for a specific task based on the analysis results.

[0191] The playlist includes energetic and fast-paced songs, as well as calm and rhythmic songs.

[0192] 4. Playlist provided by:

[0193] The generated playlist is sent to the terminal and a playback instruction is issued.

[0194] 5. Feedback and learning:

[0195] The device collects the history of songs played and feedback from users and sends it to the server.

[0196] The system will self-learn based on the feedback and reflect this in the next playlist generation.

[0197] Processing Details

[0198] Hardware:

[0199] The robot is equipped with a head-mounted display (HMD) and built-in speakers.

[0200] Sensors for data collection (e.g., activity trackers and temperature sensors).

[0201] software:

[0202] Data analysis and learning program using Python.

[0203] Clustering algorithm using the scikit-learn library.

[0204] Data storage using a database (e.g., MongoDB).

[0205] Specific examples

[0206] For example, if a robot performing assembly work is engaged in the task "assembly," it will play lively, fast-paced music to improve the efficiency of that work. In this case, the server selects appropriate music based on the data collected so far, creates a playlist, and sends it to the device. It also receives feedback and selects the best music for the next task.

[0207] Prompt Sentence Examples

[0208] You can instruct the generative AI model to collect data and generate a playlist using prompts like the following:

[0209] Create a program for the data collection part. Collect the task ID, work time, and activity level while the robot is working, and save them in a data frame format.

[0210]

[0211] Please create a program for generating the playlist. Create an optimal playlist based on the robot's task ID and return it in list format.

[0212]

[0213] Write a program that provides the generated playlist. Design a system that plays all the songs in the playlist in order.

[0214] In this way, the present invention can optimize the efficiency and accuracy of the robot's work and provide an effective musical experience.

[0215] The flow of the specific processing in the application example 1 will be described with reference to FIG.

[0216] Step 1:

[0217] The device collects the user's or robot's playback history, usage time period, activity patterns, and emotional state from sensors and stores them in a data frame format. The input is sensor data, and the output is a formatted data frame. The device formats the collected data into a specified format and transmits it to the server using a secure communication method.

[0218] Step 2:

[0219] The server stores the data received from the terminal in a database. The input is the data sent from the terminal, and the output is the data stored in the database. The server performs a write operation to the database and checks whether the data is saved correctly.

[0220] Step 3:

[0221] The server analyzes the stored data and classifies the activity patterns of the user or robot using a clustering algorithm (e.g., KMeans in scikit-learn). The input is the data in the database, and the output is the classification results for each cluster. Data preprocessing and feature extraction are also performed during the analysis process.

[0222] Step 4:

[0223] The server generates a playlist optimized for a specific task based on the analysis results. The input is the clustering results and user preference data, and the output is the generated playlist. Specifically, it selects appropriate songs based on the data belonging to a specific cluster.

[0224] Step 5:

[0225] The server provides the generated playlist to the terminal. The input is the generated playlist, and the output is the playlist data sent to the terminal. The server packages and sends the list of songs included in the playlist and their associated data (song IDs, playback URLs, etc.).

[0226] Step 6:

[0227] The terminal displays the received playlist on the user interface, making it playable by the user or robot. The input is the playlist data received from the server, and the output is a playable playlist on the user interface. The terminal starts streaming the music over the Internet.

[0228] Step 7:

[0229] The device collects data during playback (e.g., playback completion, skip, rating) and feedback from the user. The input is user operations during playback and feedback data, and the output is the collected feedback data. The device securely transmits the feedback data to the server.

[0230] Step 8:

[0231] The server analyzes the received feedback data and continues to train the system. The input is the collected feedback data, and the output is an updated learning model for the system. The server reflects the feedback and uses it to generate the next playlist.

[0232] Through these processing steps, a system that optimizes the user's work efficiency can be effectively realized.

[0233] Furthermore, an emotion engine that estimates the user's emotion may be combined. That is, the identification processing unit 290 may estimate the user's emotion using the emotion identification model 59 and perform identification processing using the user's emotion.

[0234] This invention is a system for personalizing the music experience, and by combining it with an emotion engine that recognizes the user's emotions, it provides a more accurate music playlist. Specifically, it collects the user's playback history, usage time period, activity patterns, and emotional state, and then uses the emotion engine to recognize and analyze the user's real-time emotions, and generates and provides a playlist based on that data.

[0235] User Data Collection

[0236] A user launches a music application and starts playing a song. At the same time, the user sets an activity state (e.g., "commuting" or "playing sports") and an emotion tag (e.g., "having fun" or "wanting to relax").

[0237] Additionally, the device uses a built-in emotion engine to recognize real-time emotions by analyzing the user's voice, facial expressions, gestures, etc. This emotion recognition data may be stored in the form of tags.

[0238] The data collected by the device is temporarily stored and then transmitted to a server using a secure communication protocol.

[0239] Data analysis and learning

[0240] The server stores the received data in a database, which includes the user's playback history, usage time period, activity patterns, emotion tags, and real-time emotion recognition data.

[0241] The server analyzes the stored data using machine learning algorithms to extract features that can help identify users' musical preferences and behavioral patterns.

[0242] The server also performs statistical analysis of multiple user data and learns general trends, enabling more accurate personalization.

[0243] Playlist Generation

[0244] The server then generates a playlist optimized for a specific user based on the analyzed data, including songs that reflect the user's time of day, activity patterns, emotion tags, and real-time emotions recognized by the emotion engine.

[0245] The server packages the information of the generated playlist (music ID, playback URL, etc.) and sends it to the terminal.

[0246] Providing playlists

[0247] The terminal displays the received playlists in the user interface of the music application for easy access by the user.

[0248] The user selects a song from the playlist and plays it. When a playback command is given, the terminal streams the song over the Internet.

[0249] Gather feedback and continue learning

[0250] The device collects the user's playback behavior (skip songs, complete playback, rating, etc.) and their emotions regarding those actions (including real-time emotion recognition) and records them as feedback.

[0251] The terminal transmits this feedback information to the server.

[0252] The server analyzes the received feedback and uses it to update the user's preferences, which then reflects the feedback when generating the next playlist, resulting in more accurate personalization.

[0253] As a concrete example, if a user feels "busy" while commuting on the train in the morning, the emotion engine will recognize this emotion in real time from the user's facial expressions and tone of voice. Based on past data and learning results, the server will select energetic, fast-paced music for that time period and generate a playlist. This playlist will be provided to the user via their device. Similarly, if the user wants to relax in the evening, the server will generate a playlist containing calming music, which will also be provided to the user via their device. This system enables a music experience optimized for the user's situation and emotions.

[0254] The processing flow will be explained below.

[0255] Step 1:

[0256] The user launches a music application and starts playing a song, while manually inputting an activity state (e.g., "commuting" or "playing sports") and an emotion tag (e.g., "having fun" or "wanting to relax").

[0257] Step 2:

[0258] The device automatically collects information about the song being played by the user (playback start time, play end time, song ID, etc.), and uses a built-in emotion engine to analyze the user's voice, facial expressions, and gestures, thereby recognizing the user's emotions in real time.

[0259] Step 3:

[0260] The device formats the collected data (playback history, usage time, activity patterns, manually entered emotion tags, and real-time emotion recognition data) into a specific format, encrypts it, and then sends it to a server using a secure communication protocol (e.g., HTTPS).

[0261] Step 4:

[0262] The server stores the received data in a database, which is then analyzed using machine learning algorithms to extract features that identify the user's musical preferences and behavioral patterns.

[0263] Step 5:

[0264] The server uses the listening history and the real-time emotions recognized by the emotion engine to generate an optimal playlist that suits the user's mood and activity at that moment.

[0265] Step 6:

[0266] The server then packages the generated playlist information (song ID, playback URL, etc.) and sends it to the device. This communication is also carried out using a secure protocol.

[0267] Step 7:

[0268] The terminal displays the received playlist on the user interface of the music application, allowing the user to access it directly.

[0269] Step 8:

[0270] The user selects a song from the playlist and plays it. When the device receives the instruction to play, it starts streaming the song over the Internet.

[0271] Step 9:

[0272] The device collects the user's playback behavior (skip songs, complete playback, rating specific songs, etc.) and real-time emotional data, and then transmits this feedback information back to the server.

[0273] Step 10:

[0274] The server analyzes the received feedback and updates the user's preferences. The updated information is reflected in the next playlist generation, and the system uses this information for further learning.

[0275] Example 2

[0276] Next, a description will be given of Example 2. In the following description, the data processing device 12 will be referred to as a "server" and the smart device 14 will be referred to as a "terminal."

[0277] Current music playlist generation systems generate playlists based on a user's playback history, time of day, and basic activity patterns. However, it is difficult to reflect the user's momentary emotional state or real-time preferences. This makes it difficult for users to consistently enjoy the optimal music experience. Furthermore, there are limitations to methods for effectively utilizing feedback to continuously improve the system. To address these issues, a system that can recognize and analyze a user's real-time emotions and provide highly accurate personalized playlists is needed.

[0278] The specific processing by the specific processing unit 290 of the data processing device 12 in the second embodiment is realized by the following means.

[0279] In this invention, the server includes means for collecting a user's playback history, usage time periods, activity patterns, and emotional states; means for analyzing the collected data and extracting the user's music preferences and behavioral patterns; means for generating a playlist based on the analysis results; means for providing the generated playlist to the user terminal; means for analyzing the user's voice, facial expressions, and gestures to recognize and record real-time emotions; and means for collecting feedback from users and training the system based on the feedback.

[0280] This allows for the generation of an optimal playlist that reflects the user's real-time emotions, making it possible to continue providing a music experience that matches the user's preferences.

[0281] "User" means an individual who utilizes the System to play music and obtain a personalized music experience.

[0282] "Playback history" refers to a record of songs that a user has played in the past.

[0283] A "time period of use" refers to a segment of time during which a user uses a music application.

[0284] An "activity pattern" refers to the behavior or situation in which a user is playing music (e.g., commuting, playing sports).

[0285] "Emotional state" refers to the emotion or mood of the user when playing music.

[0286] "Means of collection" refers to the methods and devices used to acquire and record the necessary data (playback history, time of use, activity patterns, emotional state, etc.).

[0287] "Means for analysis" refers to methods or devices that use machine learning algorithms to analyze collected data and extract meaningful information.

[0288] "Means for generating a playlist" refers to a method or device that creates a list of songs based on the analysis results and provides it to the user.

[0289] "Means for providing" refers to a method or device for delivering the generated playlist to a user's terminal and making it accessible.

[0290] An "emotion engine" is software or algorithms that analyze voice, facial expressions, and gestures to recognize and record a user's real-time emotions.

[0291] "Feedback collection means" refers to a method or device that captures and records a user's playback behavior and emotional responses.

[0292] "Means for system training" refers to methods and devices that analyze collected feedback and improve the personalization accuracy of the entire system.

[0293] The present invention is a system for personalizing a user's music experience, and in particular, provides a music playlist that reflects the user's real-time emotions by combining an emotion engine. This system is realized through the following elements and processes.

[0294] 1. Collection of User Data

[0295] The user launches the music application and selects an activity status such as "commuting" or "playing sports" or an emotion tag such as "fun" or "want to relax" within the application when starting use.

[0296] The device collects this information and uses an emotion engine (e.g., OpenFace or Microsoft Azure's Emotion API) that analyzes voice, facial expressions, and gestures to obtain real-time emotional data about the user.

[0297] 2. Data transmission and storage

[0298] The device temporarily stores the user's playback history, usage time period, activity patterns, emotion tags, and real-time emotion recognition data, and transmits them to the server using a secure communication protocol (e.g., HTTPS).

[0299] 3. Data analysis and learning

[0300] The server stores the transmitted data in a structured database, which includes each user's playback history, usage time periods, activity patterns, emotion tags, and real-time emotion recognition data.

[0301] The server analyzes the stored data using machine learning algorithms (e.g., TensorFlow or PyTorch) to extract features that identify the user's musical preferences and behavioral patterns.

[0302] The server also performs statistical analysis of multiple user data and learns general trends to improve the accuracy of personalization across the system.

[0303] 4. Generate a playlist

[0304] The server then generates a playlist optimized for a specific user based on the analysis results, for example, selecting energetic music for a user commuting in the morning and calming music for a user wanting to relax in the evening.

[0305] The server packages the generated playlist (music IDs, playback URLs, etc.) and sends it to the terminal.

[0306] 5. Providing playlists

[0307] The device displays the received playlists in the music application's user interface for easy access by the user, with the playlists visually categorized.

[0308] The user selects a song from the playlist and plays it. When a playback command is given, the terminal streams the song over the Internet.

[0309] 6. Gather feedback and continue learning

[0310] The device collects and records the user's playback behavior (e.g., skipping songs, completing playback, rating, etc.) and emotional responses.

[0311] The terminal transmits this feedback information to the server.

[0312] The server analyzes the received feedback and uses it to update the user's preferences, so that the feedback is reflected the next time a playlist is generated, resulting in more accurate personalization.

[0313] Examples of concrete examples and prompts

[0314] For example, if a user feels "busy" while commuting on the train in the morning, the emotion engine recognizes this emotion in real time from the user's facial expressions and tone of voice. Based on past data and learning results, the server selects energetic, fast-paced music for that time period and generates a playlist. This playlist is provided to the user via their device. On the other hand, when the user wants to relax in the evening, the server generates a playlist containing calming music, which is also provided to the user via their device.

[0315] An example of a prompt to input to a generative AI model is as follows:

[0316] "I'm on the train to work this morning and I'm feeling busy. Make me a playlist of energetic music."

[0317] "Generate the perfect playlist for a relaxing evening."

[0318] This will provide the user with an optimal music experience that suits their real-time emotions and situation.

[0319] The flow of the identification process in the second embodiment will be described with reference to FIG.

[0320] Step 1: Collect user data

[0321] input:

[0322] User play history

[0323] User usage time period

[0324] User activity patterns (e.g., commuting, playing sports)

[0325] User's emotional state (e.g., "I'm having fun" or "I want to relax")

[0326] process:

[0327] The user starts the music application and selects the playback history and various tags when starting use.

[0328] The device obtains real-time emotional data using an emotion engine (e.g., OpenFace or Microsoft Azure's Emotion API) that analyzes the user's voice, facial expressions, and gestures.

[0329] output:

[0330] User playback history data

[0331] Current time zone data

[0332] Activity pattern data

[0333] Real-time emotion data

[0334] Specific behavior:

[0335] The user operates a music application and inputs settings that make them feel "busy" during their commute, and the device uses its camera and microphone to provide data to the emotion engine, generating real-time emotion data.

[0336] Step 2: Send and store data

[0337] input:

[0338] User playback history data, usage time zone data, activity pattern data, and real-time emotion data collected in Step 1

[0339] process:

[0340] The terminal temporarily stores the collected data.

[0341] The saved data is sent to the server using a secure communication protocol (e.g. HTTPS).

[0342] output:

[0343] Data sent to the server

[0344] Specific behavior:

[0345] The device saves the data in local storage and then transmits it to the server using HTTPS.

[0346] Step 3: Analyze and train the data

[0347] input:

[0348] Data sent in step 2

[0349] process:

[0350] The server stores the submitted data in a structured database.

[0351] Machine learning algorithms (e.g., TensorFlow and PyTorch) are used to analyze users' musical preferences and behavioral patterns and extract specific features.

[0352] The server performs statistical analysis of multiple user data to learn general trends.

[0353] output:

[0354] User music taste characteristic data

[0355] General Trend Data

[0356] Specific behavior:

[0357] The server accumulates data in a database, analyzes it using a machine learning model, extracts features based on the user's preferences and behavioral patterns, and builds a model based on these.

[0358] Step 4: Generate a playlist

[0359] input:

[0360] The characteristic data of the user's music preferences and general tendency data extracted in step 3

[0361] process:

[0362] Based on the analysis results, the server generates a playlist that is optimal for the specific user.

[0363] The generated playlist information (song ID, playback URL, etc.) is packaged and sent to the terminal.

[0364] output:

[0365] Generated playlist data

[0366] Specific behavior:

[0367] The server combines each user's characteristic data with general tendency data to generate an optimal playlist, which is then packaged and sent to the terminal.

[0368] Step 5: Serve the playlist

[0369] input:

[0370] The playlist data generated in step 4

[0371] process:

[0372] The terminal displays the received playlist on the user interface of the music application.

[0373] The user selects and plays songs from the playlist.

[0374] The device streams music over the internet.

[0375] output:

[0376] Displayed playlist

[0377] Streaming songs

[0378] Specific behavior:

[0379] The playlist received by the device is displayed on the app, the user performs operations to play the music, and internet streaming is performed when playback begins.

[0380] Step 6: Gather feedback and continue learning

[0381] input:

[0382] User playback behavior data (e.g., skipping songs, completing playback, rating, etc.)

[0383] Real-time emotion data

[0384] process:

[0385] The terminal collects the user's playing behavior and emotional responses and records them as feedback.

[0386] The recorded feedback information is sent to the server.

[0387] The server analyzes the feedback and uses it to update the user's preferences.

[0388] output:

[0389] Updated user preference data

[0390] Improving the system's learning model

[0391] Specific behavior:

[0392] The device collects real-time playback behavior and emotional responses through an emotion engine and sends the information to a server, which analyzes it as learning material to improve the system's personalization accuracy.

[0393] Through this series of processes, an optimal music playlist that reflects the user's real-time emotions is generated and provided to the user, significantly improving the user's music experience.

[0394] (Application example 2)

[0395] Next, a description will be given of Application Example 2. In the following description, the data processing device 12 will be referred to as a "server" and the smart device 14 will be referred to as a "terminal."

[0396] Conventional music playlist generation systems generate playlists based only on a user's playback history, usage time, and activity patterns, which means they lack the ability to personalize playlists in response to real-time emotional changes. Furthermore, it is difficult to provide a music experience that takes into account emotions in a virtual store. This can lead to situations where the user is feeling stressed, such as not playing the right music, which can detract from the shopping experience.

[0397] The specific processing by the specific processing unit 290 of the data processing device 12 in the application example 2 is realized by the following means.

[0398] In this invention, the server includes means for collecting a user's playback history, usage time period, activity pattern, and emotional state, means for analyzing the user's facial expressions and voice to recognize emotions in real time, and means for analyzing the collected data and real-time emotion recognition data to extract the user's music preferences and behavioral patterns. This makes it possible to generate a playlist based on real-time emotion recognition, and to provide optimal music that takes into account the user's emotions, particularly in a shopping experience in a virtual store.

[0399] The "playback history" is a history of songs and audio content that have been played by the user in the past.

[0400] The "usage time period" refers to the time or time period during which a user uses a music playlist.

[0401] An "activity pattern" refers to a user's daily activity or behavior pattern, such as a situation such as "commuting" or "exercising."

[0402] An "emotional state" is the emotion or mood a user is experiencing at a particular time.

[0403] An "emotion engine" is a technology that analyzes a user's facial expressions and voice to recognize emotions in real time.

[0404] A "user terminal" is a device used by a user, such as a smartphone, smart glasses, or head-mounted display.

[0405] "Feedback" refers to the user's playback behavior (skip songs, complete playback, rating, etc.) and their feelings about it.

[0406] A "playlist" is a list of songs that a user has selected to play consecutively.

[0407] A "virtual store" is a shopping environment provided in a virtual space where users can browse and purchase items from the store via digital devices.

[0408] "Real-time emotion recognition data" is data obtained by analyzing the user's current emotions in real time.

[0409] This invention is a system for personalizing the music experience in a virtual store using real-time emotion recognition data. The aim is to improve the user experience by generating and providing an optimal music playlist according to the user's emotional state while shopping in the virtual store through smart glasses.

[0410] 1. System configuration and programs

[0411] The system includes the following components:

[0412] Smart glasses equipped with an emotion recognition engine

[0413] Server that collects and analyzes data

[0414] A server that generates and provides playlists

[0415] The emotion recognition engine is integrated into the smart glasses and uses OpenCV and deep learning models to analyze the user's facial expressions and voice. Server technologies using frameworks such as Flask and Django are used for data collection and analysis.

[0416] 2. Processing Details

[0417] The emotion recognition engine captures the user's facial expressions through the camera and analyzes the data in real time. The analyzed emotional data, along with the user's playback history, usage time, and activity patterns, are sent to the server. The server then uses machine learning algorithms based on this data to extract the user's musical preferences and behavioral patterns.

[0418] The server then generates a playlist optimized for the specific user based on the emotion recognition data and the extracted music preference data. This playlist consists of songs that suit the user's current emotions. The generated playlist is then sent to the smart glasses via the Internet.

[0419] 3. Specific examples

[0420] For example, if a user feels stressed while shopping in a virtual store, the smart glasses' emotion recognition engine will recognize that emotion in real time. The emotion recognition data will be sent to the server, which will then generate a playlist containing music that promotes relaxation. This playlist will then be sent to the smart glasses, reducing the user's stress and providing a comfortable shopping experience.

[0421] An example of a prompt to input to a generative AI model would be:

[0422] Analyze in real time whether the user is currently feeling stressed and generate an appropriate relaxing music playlist based on that emotion. Example: If the user is feeling stressed, create a list of songs that are particularly effective for relaxing and play them through smart glasses.

[0423] In this way, the system realized by the present invention provides a music playlist that corresponds to the user's emotional state, thereby improving the shopping experience in a virtual store.

[0424] The flow of the specific processing in the application example 2 will be described with reference to FIG.

[0425] Step 1:

[0426] Smart glasses capture the user's facial expressions and voice.

[0427] Input: User's face image and voice.

[0428] Output: Captured face image data and audio data.

[0429] Specific operation: Uses the camera and microphone of the smart glasses to collect the user's face and voice in real time.

[0430] Step 2:

[0431] The emotion recognition engine analyzes the captured facial images and voice to recognize the user's emotions.

[0432] Input: Captured facial image data and audio data.

[0433] Output: Real-time emotion recognition data (e.g. "stressed", "relaxed", etc.).

[0434] How it works: It uses OpenCV and deep learning models to perform facial expression analysis and speech analysis, such as analyzing facial muscle movements and tone of voice, to create appropriate emotion tags.

[0435] Step 3:

[0436] The smart glasses send emotion recognition data, playback history, usage time, and activity patterns to a server.

[0437] Input: Real-time emotion recognition data, playback history, time of day usage, activity patterns.

[0438] Output: The aggregate data sent to the server.

[0439] Specific operation: Collected data is sent to the server using a secure protocol (e.g., HTTPS).

[0440] Step 4:

[0441] The server analyzes the received data and extracts the user's musical preferences and behavioral patterns.

[0442] Input: Playback history, time of day, activity patterns, real-time emotion recognition data.

[0443] Output: Music preference data and behavioral pattern data as analysis results.

[0444] Specific operation: Using machine learning algorithms (e.g., clustering algorithms and recommendation systems), analyze users' musical preferences and behavioral patterns.

[0445] Step 5:

[0446] The server generates a playlist based on the analysis results and real-time emotion recognition data.

[0447] Input: Music preference data, behavioral pattern data, real-time emotion recognition data.

[0448] Output: The generated playlist.

[0449] Specific operation: The analysis results are combined with emotional data to generate a list of songs that are optimal for the user, selecting songs that correspond to a specific emotional state (e.g., stress).

[0450] Step 6:

[0451] The server provides the generated playlist to the smart glasses.

[0452] Input: The generated playlist.

[0453] Output: Playlist data sent to the smart glasses.

[0454] Specific operation: The playlist data is packaged and sent to the smart glasses using a secure communication protocol.

[0455] Step 7:

[0456] The smart glasses play the received playlist.

[0457] Input: Playlist data sent to smart glasses.

[0458] Output: The song that is playing.

[0459] Specific operation: The received playlist is displayed in the user interface, and when the user selects it, the song is streamed over the Internet.

[0460] Step 8:

[0461] Collect user feedback and send it to the server.

[0462] Input: User playback behavior data (e.g., song skip, playback completion, song rating, etc.).

[0463] Output: Feedback data sent to the server.

[0464] Specific operation: Records various playback actions of users in real time and periodically sends the information to the server.

[0465] Step 9:

[0466] The server analyzes the feedback and trains the system.

[0467] Input: Feedback data.

[0468] Output: Updated music preference data and behavioral pattern data.

[0469] Specific operation: The received feedback data is analyzed using machine learning algorithms and updated to improve the system's personalization accuracy.

[0470] Example prompts to be input to the generative AI model:

[0471] Analyze in real time whether the user is currently feeling stressed and generate an appropriate relaxing music playlist based on that emotion. Example: If the user is feeling stressed, create a list of songs that are particularly effective for relaxing and play them through smart glasses.

[0472] The specific processing unit 290 transmits the result of the specific processing to the smart device 14. In the smart device 14, the control unit 46A causes the output device 40 to output the result of the specific processing. The microphone 38B acquires audio indicating a user input regarding the result of the specific processing. The control unit 46A transmits audio data indicating the user input acquired by the microphone 38B to the data processing device 12. In the data processing device 12, the specific processing unit 290 acquires the audio data.

[0473] The data generation model 58 is a so-called generative AI (Artificial Intelligence). An example of the data generation model 58 is ChatGPT (Internet Search<URL: https: / / openai.com / blog / chatgpt> ), Gemini (Internet search <url: https: gemini.google.com ?hl="ja">) and other generation AIs. The data generation model 58 is obtained by performing deep learning on a neural network. A prompt including an instruction is input to the data generation model 58, and inference data such as voice data indicating voice, text data indicating text, and image data indicating an image is also input. The data generation model 58 performs inference on the input inference data in accordance with the instruction indicated by the prompt, and outputs the inference result in a data format such as voice data and text data. Here, inference refers to, for example, analysis, classification, prediction, and / or summarization.

[0474] In the above embodiment, an example in which the specific process is performed by the data processing device 12 has been given, but the technology of the present disclosure is not limited to this, and the specific process may be performed by the smart device 14.

[0475] [Second embodiment]

[0476] FIG. 3 shows an example of the configuration of a data processing system 210 according to the second embodiment.

[0477] 3, the data processing system 210 includes the data processing device 12 and smart glasses 214. An example of the data processing device 12 is a server.

[0478] The data processing device 12 includes a computer 22, a database 24, and a communication I / F 26. The computer 22 is an example of a "computer" according to the technology of the present disclosure. The computer 22 includes a processor 28, a RAM 30, and a storage 32. The processor 28, the RAM 30, and the storage 32 are connected to a bus 34. The database 24 and the communication I / F 26 are also connected to the bus 34. The communication I / F 26 is connected to a network 54. Examples of the network 54 include a WAN (Wide Area Network) and / or a LAN (Local Area Network).

[0479] The smart glasses 214 include a computer 36, a microphone 238, a speaker 240, a camera 42, and a communication I / F 44. The computer 36 includes a processor 46, a RAM 48, and a storage 50. The processor 46, the RAM 48, and the storage 50 are connected to a bus 52. The microphone 238, the speaker 240, and the camera 42 are also connected to the bus 52.

[0480] The microphone 238 receives instructions and the like from the user 20 by receiving voice uttered by the user 20. The microphone 238 captures the voice uttered by the user 20, converts the captured voice into audio data, and outputs it to the processor 46. The speaker 240 outputs audio in accordance with instructions from the processor 46.

[0481] Camera 42 is a small digital camera equipped with an optical system including a lens, aperture, and shutter, and an imaging element such as a CMOS (Complementary Metal-Oxide-Semiconductor) image sensor or a CCD (Charge Coupled Device) image sensor, and captures images of the surroundings of user 20 (for example, an imaging range defined by an angle of view equivalent to the field of vision of a typical healthy person).

[0482] The communication I / F 44 is connected to a network 54. The communication I / Fs 44 and 26 are responsible for the exchange of various information between the processor 46 and the processor 28 via the network 54. The exchange of various information between the processor 46 and the processor 28 using the communication I / Fs 44 and 26 is carried out in a secure state.

[0483] Fig. 4 shows an example of the main functions of the data processing device 12 and the smart glasses 214. As shown in Fig. 4, in the data processing device 12, a specific process is performed by the processor 28. A specific process program 56 is stored in the storage 32.

[0484] The specific processing program 56 is an example of a "program" according to the technology of the present disclosure. The processor 28 reads the specific processing program 56 from the storage 32 and executes the read specific processing program 56 on the RAM 30. The specific processing is realized by the processor 28 operating as a specific processing unit 290 in accordance with the specific processing program 56 executed on the RAM 30.

[0485] The storage 32 stores a data generation model 58 and an emotion identification model 59. The data generation model 58 and the emotion identification model 59 are used by the identification processing unit 290.

[0486] In the smart glasses 214, the reception output process is performed by the processor 46. A reception output program 60 is stored in the storage 50. The processor 46 reads the reception output program 60 from the storage 50 and executes the read reception output program 60 on the RAM 48. The reception output process is realized by the processor 46 operating as the control unit 46A in accordance with the reception output program 60 executed on the RAM 48.

[0487] Next, a description will be given of the identification process performed by the identification processing unit 290 of the data processing device 12. In the following description, the data processing device 12 will be referred to as the "server" and the smart glasses 214 will be referred to as the "terminal."

[0488] This invention relates to a system that personalizes and optimizes a user's music experience. Specifically, it is a method for generating and providing playlists based on the user's preferences and behavioral patterns by collecting and analyzing the user's playback history, usage time period, activity patterns, and emotional state. Furthermore, by collecting user feedback and allowing the system to continuously learn, it provides a more accurate and personalized music experience.

[0489] User Data Collection

[0490] A user starts a music application and plays music in the course of daily life.

[0491] The device collects the user's music history, the start and end times of the music, whether the music was skipped, the user's current activity status (e.g., commuting, exercising), and emotion tags, which may be entered by the user in the form of "fun" or "want to relax."

[0492] The terminal converts this data into a specified format and transmits it to the server using a secure communication method.

[0493] Data analysis and learning

[0494] The server stores the received user data in a database.

[0495] The server analyzes the stored data and extracts features to identify users' musical preferences and behavioral patterns. This analysis uses machine learning algorithms. For example, a clustering algorithm can group users' activity patterns.

[0496] The server performs statistical analysis based on a large amount of user data and also learns general trends.

[0497] Playlist Generation

[0498] The server then generates the optimal playlist for a specific user based on the analysis results, which means selecting songs of different genres and moods according to the user's time of day, activity patterns, and emotional tags.

[0499] The server packages the list of songs included in the generated playlist and its associated data (song ID, playback URL, etc.) and sends it to the terminal.

[0500] Providing playlists

[0501] The terminal displays the received playlist on the application's user interface for easy access by the user.

[0502] The user selects a song from the playlist and plays it. When a playback command is given, the terminal streams the song over the Internet.

[0503] Gather feedback and continue learning

[0504] The device collects the songs the user has played, whether they have completed or skipped, and their ratings and feedback (e.g., "favorite marks") for specific songs.

[0505] The terminal transmits this feedback information to the server.

[0506] The server updates the user's preferences based on the received feedback and then reruns the analysis and learning process, which is then reflected in the next playlist generation.

[0507] As a concrete example, if a user sets the emotion tag "busy" during their morning commute, the server will select energetic, fast-paced music for that time period based on past data and learning results. The generated playlist is provided to the device and played by the user. On the other hand, when it's time to relax in the evening, the server will generate a playlist containing calming music and provide it to the user via the device. In this way, it is possible to provide a music experience that is optimal for the user's situation and emotion.

[0508] The processing flow will be explained below.

[0509] Step 1:

[0510] A user launches a music application and starts playing a song. At the same time, the user sets an activity state (e.g., "commuting" or "playing sports") and an emotion tag (e.g., "having fun" or "wanting to relax").

[0511] Step 2:

[0512] The device automatically collects information about the song the user is playing (playback start time, play end time, song ID, etc.), as well as the user's activity status and emotion tag. The collected data is temporarily stored on the device.

[0513] Step 3:

[0514] The device formats the collected user data into a specific format, encrypts it, and then sends it to a server using a secure protocol such as HTTPS.

[0515] Step 4:

[0516] The server stores the received user data in a database, and then starts a process to analyze the stored data. This analysis uses machine learning algorithms.

[0517] Step 5:

[0518] The server extracts features from the analyzed data to identify users' musical preferences and behavioral patterns, for example by using a clustering algorithm to group users based on their activity patterns and emotions.

[0519] Step 6:

[0520] The server selects music suitable for a specific time period or activity level and generates an optimal playlist for each user, for example, energetic music for the commute and quieter music for relaxation.

[0521] Step 7:

[0522] The server packages the generated playlist information (song ID, playback URL, etc.) and sends it to the device. This communication also uses a secure protocol.

[0523] Step 8:

[0524] The device displays the received playlist in the user interface of the music application, allowing the user to select and play it.

[0525] Step 9:

[0526] The user selects a song from the presented playlist and starts playing it. The device streams the song over the Internet.

[0527] Step 10:

[0528] The device collects the user's playback behavior (skip songs, complete playback, rating, etc.) and records it as feedback.

[0529] Step 11:

[0530] The terminal transmits the collected feedback information to the server.

[0531] Step 12:

[0532] The server analyzes the received feedback and uses it as data to update the user's preferences, which allows for more accurate personalization when generating the next playlist.

[0533] Example 1

[0534] Next, a description will be given of Example 1. In the following description, the data processing device 12 will be referred to as a "server" and the smart glasses 214 will be referred to as a "terminal."

[0535] Conventional music recommendation systems have difficulty generating playlists that fully consider a user's specific usage situation and emotions. Furthermore, they lack a system that can properly learn from user feedback and reflect it in future recommendations. As a result, users are unable to obtain the music experience that best suits their preferences and situation.

[0536] The specific processing by the specific processing unit 290 of the data processing device 12 in the first embodiment is realized by the following means.

[0537] In this invention, the server includes: means for collecting a user's playback history, usage time periods, activity patterns, and emotional state; means for formatting the collected data into a specified format and transmitting it to the server using a secure communication method; means for saving the received user data, analyzing it using a machine learning algorithm, and extracting the user's music preferences and behavioral patterns; means for generating an optimal playlist for a specific user based on the analysis results; means for packaging a list of songs included in the generated playlist and their associated data and transmitting it to a user terminal; means for displaying the received playlist on a user interface for easy user access; and means for collecting user feedback and re-executing the analysis and learning process based on the feedback, thereby enabling a more personalized music experience based on the user's usage history and emotional tags.

[0538] "Playback history" is information about songs that the user has played in the past, and includes song IDs, titles, artist names, and the like.

[0539] The "usage time period" is information about the date and time when the user used the music application, including the playback start time and end time.

[0540] The "activity pattern" refers to the activity status of the user while playing music, and indicates a specific operating state such as commuting, exercising, or relaxing.

[0541] The "emotional state" refers to the emotional or mood state of the user when listening to music, and is expressed in the form of emotion tags such as "fun" or "want to relax."

[0542] A "secure communication method" is a communication protocol that ensures safety when sending and receiving data, and a specific example is HTTPS.

[0543] "Machine learning algorithms" are algorithms that analyze large amounts of data and extract patterns and rules, and include clustering algorithms.

[0544] A "playlist" is a collection of multiple songs, and is a format for providing a collection of songs that meets the user's playback needs.

[0545] A "user interface" is a display screen that allows a user to directly interact with a system or application, and includes a song list, operation buttons, and the like.

[0546] "Feedback" refers to information provided by users to the system, such as playback history, song ratings and comments, which is reflected in the system's learning process.

[0547] The "analysis and learning process" is a series of steps that the system takes to learn the user's preferences and behavioral patterns based on data collected from the user, and then generate a new playlist.

[0548] This invention relates to a system that personalizes and optimizes a user's music experience. This system collects and analyzes a user's playback history, usage time, activity patterns, and emotional state, and generates and provides playlists based on the user's preferences and behavioral patterns. Furthermore, by collecting user feedback and continuously learning, the system provides an even more accurate and personalized music experience.

[0549] First, a user launches a music application and plays a song in their daily life. For example, a user presses the play button on the application to play a song. Here, the device collects the history of songs played by the user, the start and end times of playback, whether songs were skipped, the user's current activity state (commuting, exercising, etc.), and emotion tags (e.g., happy, want to relax). These emotion tags are input by the user.

[0550] The device then formats the collected data into a specified format and sends it to the server using HTTPS, a secure communication method. At this time, the data is converted into JSON format for efficient data processing.

[0551] The server stores the received user data in a database. For example, the server inserts data into the database using SQL queries. The server then analyzes the stored data using machine learning algorithms to identify users' musical preferences and behavioral patterns. This analysis uses the Python library scikit-learn, which runs a clustering algorithm to extract specific patterns.

[0552] Next, the server generates a playlist optimized for the specific user based on the analysis results. This means selecting songs of different genres and moods depending on the user's time of day, activity patterns, and emotional tags. The generated playlist is converted into JSON format and sent back to the device using the HTTPS protocol.

[0553] The device displays the received playlist on the application's user interface for easy access by the user. For example, the user can click a song in the playlist and press the play button to start streaming the song. Specifically, the device connects to the streaming server and streams the music using the playback URL.

[0554] Additionally, when a user provides feedback such as which songs they have completed playing, which songs they have skipped, or their ratings for specific songs, the device collects this feedback information and sends it to the server. The server then updates the user's preferences based on the received feedback and reruns the analysis and learning process, allowing the information to be reflected in the next playlist generation.

[0555] As a concrete example, if a user sets the emotion tag "busy" during their morning commute, the server will select energetic, fast-paced music for that time period based on past data and learning results. The generated playlist is provided to the device and played by the user. On the other hand, when it's time to relax in the evening, the server will generate a playlist containing calming music and provide it to the user via the device. In this way, it is possible to provide a music experience that is optimal for the user's situation and emotion.

[0556] Examples of prompt statements

[0557] "Please explain in detail the data analysis process that goes into generating an energizing playlist for your morning commute."

[0558] "How do I update my personalized music playlist based on user feedback?"

[0559] "Please explain in detail how you would generate an evening playlist based on the emotion tag 'I want to relax'."

[0560] The flow of the identification process in the first embodiment will be described with reference to FIG.

[0561] Step 1: Collect user data

[0562] In everyday life, users launch music applications and play music.

[0563] Specifically, the user presses the play button on the application on a device such as a smartphone or computer.

[0564] The device collects the user's history of songs played, the start and end times of playback, whether songs were skipped, the user's current activity status (commuting, exercising, etc.), and emotion tags (e.g., fun, want to relax).

[0565] The input is various data when the user operates the playlist.

[0566] As part of the data processing, the terminal records these data in its internal memory.

[0567] The output is formatted data, ready to be sent to the next step.

[0568] Step 2: Sending data

[0569] The terminal formats the collected data into the specified format.

[0570] The input is raw, unformatted data.

[0571] To process the data, the terminal converts the data into JSON format.

[0572] The output is data in JSON format.

[0573] The terminal transmits data to the server using a secure communication method (e.g., HTTPS).

[0574] Specifically, the terminal sends JSON format data to the server as a POST request using the HTTPS protocol.

[0575] Step 3: Data storage and analysis

[0576] The server stores the received user data in a database.

[0577] The input is JSON format data sent from the terminal.

[0578] For data processing, the server receives this data and inserts it into a database using an SQL query.

[0579] The output is user data stored in a database.

[0580] The server analyzes the stored data using machine learning algorithms (e.g., clustering algorithms).

[0581] The input is user data stored in a database.

[0582] For data calculation, the server applies machine learning models using Python libraries (e.g., scikit-learn) to identify users' musical preferences and behavioral patterns.

[0583] The output is the user's characteristics and pattern information as the analysis results.

[0584] Step 4: Generate a playlist

[0585] The server generates an optimal playlist for a particular user based on the analysis results.

[0586] The input is the user's characteristics and pattern information as the analysis result.

[0587] To process the data, the server narrows down the conditions based on the time of use, activity patterns, and emotion tags, and selects appropriate songs from the music database.

[0588] The output is the generated playlist.

[0589] The server packages the list of songs included in the generated playlist and its associated data (song ID, playback URL, etc.) and sends it to the terminal.

[0590] Specifically, the server converts the playlist into JSON format and sends a POST request to the terminal using the HTTPS protocol.

[0591] Step 5: Serve the playlist

[0592] The terminal displays the received playlist on the user interface of the application, making it easily accessible to the user.

[0593] The input is the playlist data in JSON format sent from the server.

[0594] For data processing, the device parses the JSON data and dynamically displays a song list in the UI component.

[0595] The output is the displayed playlist.

[0596] The user selects a song from the presented playlist and plays it.

[0597] Specifically, the user clicks on a song in the playlist and presses the play button.

[0598] The device streams the music over the Internet.

[0599] The input is the URL to play the song.

[0600] For data processing, the terminal connects to a streaming server and plays music.

[0601] Step 6: Gather feedback and continue learning

[0602] The device collects songs that the user has completed playing, songs that the user has skipped, and ratings and feedback for specific songs.

[0603] The input is the feedback information provided by the user.

[0604] As for data processing, the terminal tracks the user's operations and stores them in data storage.

[0605] The output is the collected feedback data.

[0606] The terminal transmits this feedback information to the server.

[0607] Specifically, the terminal formats the feedback data into JSON format and sends a POST request to the server using the HTTPS protocol.

[0608] The server updates the user's preferences based on the received feedback and then reruns the analysis and learning process.

[0609] The input is the feedback data sent from the terminal.

[0610] As the data is processed, the server stores it in a database and retrains the machine learning model.

[0611] The output is updated analysis results and data that will be reflected in the next playlist generation.

[0612] (Application example 1)

[0613] Next, a description will be given of Application Example 1. In the following description, the data processing device 12 will be referred to as a "server" and the smart glasses 214 will be referred to as a "terminal."

[0614] Conventionally, methods for optimizing the efficiency and accuracy of factory robots have relied solely on motion algorithms and machine learning. However, it is believed that even more efficient work can be achieved by providing an appropriate environment based on the robot's activity status and work content. The present invention aims to improve the efficiency and accuracy of work by playing personalized music while the robot is working.

[0615] The specific processing by the specific processing unit 290 of the data processing device 12 in the application example 1 is realized by the following means.

[0616] In this invention, the server includes means for collecting a user's playback history, usage time periods, activity patterns, and emotional state, means for analyzing the collected data and extracting the user's preferences and behavioral patterns, means for generating a playlist based on the analysis results, means for providing the generated playlist to the terminal, means for collecting the terminal's work history and activity level and generating an optimal music experience based on the data, means for collecting feedback and conducting system training, and means for generating applications for use in various facilities. This makes it possible to play optimal music according to the robot's activity status and work content, effectively improving work efficiency and accuracy.

[0617] "Playback history" is a record of music and content that a user has played in the past.

[0618] A "time period of use" refers to a specific time range during which a user uses music or content.

[0619] An "activity pattern" is a series of patterns of actions or states of a user or a terminal.

[0620] "Emotional state" refers to the user's current emotional or mood state.

[0621] "Analysis" is the process of examining collected data in detail to identify trends and characteristics.

[0622] A "playlist" refers to a list of songs or content to be played.

[0623] A "terminal" refers to an electronic device used by a user, and includes, for example, a smartphone or a robot.

[0624] "Work history" is a record of work performed by a terminal or robot.

[0625] "Activity level" is a measure that indicates the activity and intensity of the terminal or robot's movements.

[0626] "Feedback" refers to responses or opinions from users or systems.

[0627] A "clustering algorithm" is a machine learning technique for classifying data into several groups.

[0628] "System learning" is the process by which a system improves itself based on collected data and feedback.

[0629] An "application" is software or a program that provides a specific function.

[0630] This invention is a system that automatically plays appropriate music to improve the work efficiency and accuracy of robots working in factories. The system collects the user's playback history, usage time, activity patterns, and emotional state, and generates and provides a playlist based on the analysis results. It also has a function that collects the device's work history and activity level to generate optimal music.

[0631] Generating a Program

[0632] The system program includes the following main functions:

[0633] 1. Data Collection:

[0634] Sensors are used to collect records of tasks performed by the device, work time, and activity levels.

[0635] The user's emotional state and usage time period are also collected (for example, music playback history while working).

[0636] 2. Data Analysis:

[0637] The collected data is stored and analyzed on the server using Python.

[0638] The robot's activity patterns are classified using scikit-learn's clustering algorithm (KMeans).

[0639] Features are extracted from the data and music appropriate for the robot's task is selected.

[0640] 3. Playlist generation:

[0641] Generate the best playlist for a specific task based on the analysis results.

[0642] The playlist includes energetic and fast-paced songs, as well as calm and rhythmic songs.

[0643] 4. Playlist provided by:

[0644] The generated playlist is sent to the terminal and a playback instruction is issued.

[0645] 5. Feedback and learning:

[0646] The device collects the history of songs played and feedback from users and sends it to the server.

[0647] The system will self-learn based on the feedback and reflect this in the next playlist generation.

[0648] Processing Details

[0649] Hardware:

[0650] The robot is equipped with a head-mounted display (HMD) and built-in speakers.

[0651] Sensors for data collection (e.g., activity trackers and temperature sensors).

[0652] software:

[0653] Data analysis and learning program using Python.

[0654] Clustering algorithm using the scikit-learn library.

[0655] Data storage using a database (e.g., MongoDB).

[0656] Specific examples

[0657] For example, if a robot performing assembly work is engaged in the task "assembly," it will play lively, fast-paced music to improve the efficiency of that work. In this case, the server selects appropriate music based on the data collected so far, creates a playlist, and sends it to the device. It also receives feedback and selects the best music for the next task.

[0658] Prompt Sentence Examples

[0659] You can instruct the generative AI model to collect data and generate a playlist using prompts like the following:

[0660] Create a program for the data collection part. Collect the task ID, work time, and activity level while the robot is working, and save them in a data frame format.

[0661]

[0662] Please create a program for generating the playlist. Create an optimal playlist based on the robot's task ID and return it in list format.

[0663]

[0664] Write a program that provides the generated playlist. Design a system that plays all the songs in the playlist in order.

[0665] In this way, the present invention can optimize the efficiency and accuracy of the robot's work and provide an effective musical experience.

[0666] The flow of the specific processing in the application example 1 will be described with reference to FIG.

[0667] Step 1:

[0668] The device collects the user's or robot's playback history, usage time period, activity patterns, and emotional state from sensors and stores them in a data frame format. The input is sensor data, and the output is a formatted data frame. The device formats the collected data into a specified format and transmits it to the server using a secure communication method.

[0669] Step 2:

[0670] The server stores the data received from the terminal in a database. The input is the data sent from the terminal, and the output is the data stored in the database. The server performs a write operation to the database and checks whether the data is saved correctly.

[0671] Step 3:

[0672] The server analyzes the stored data and classifies the activity patterns of the user or robot using a clustering algorithm (e.g., KMeans in scikit-learn). The input is the data in the database, and the output is the classification results for each cluster. Data preprocessing and feature extraction are also performed during the analysis process.

[0673] Step 4:

[0674] The server generates a playlist optimized for a specific task based on the analysis results. The input is the clustering results and user preference data, and the output is the generated playlist. Specifically, it selects appropriate songs based on the data belonging to a specific cluster.

[0675] Step 5:

[0676] The server provides the generated playlist to the terminal. The input is the generated playlist, and the output is the playlist data sent to the terminal. The server packages and sends the list of songs included in the playlist and their associated data (song IDs, playback URLs, etc.).

[0677] Step 6:

[0678] The terminal displays the received playlist on the user interface, making it playable by the user or robot. The input is the playlist data received from the server, and the output is a playable playlist on the user interface. The terminal starts streaming the music over the Internet.

[0679] Step 7:

[0680] The device collects data during playback (e.g., playback completion, skip, rating) and feedback from the user. The input is user operations during playback and feedback data, and the output is the collected feedback data. The device securely transmits the feedback data to the server.

[0681] Step 8:

[0682] The server analyzes the received feedback data and continues to train the system. The input is the collected feedback data, and the output is an updated learning model for the system. The server reflects the feedback and uses it to generate the next playlist.

[0683] Through these processing steps, a system that optimizes the user's work efficiency can be effectively realized.

[0684] Furthermore, an emotion engine that estimates the user's emotion may be further combined. That is, the identification processing unit 290 may estimate the user's emotion using the emotion identification model 59, and perform identification processing using the user's emotion.

[0685] This invention is a system for personalizing the music experience, and by combining it with an emotion engine that recognizes the user's emotions, it provides a more accurate music playlist. Specifically, it collects the user's playback history, usage time period, activity patterns, and emotional state, and then uses the emotion engine to recognize and analyze the user's real-time emotions, and generates and provides a playlist based on that data.

[0686] User Data Collection

[0687] A user launches a music application and starts playing a song. At the same time, the user sets an activity state (e.g., "commuting" or "playing sports") and an emotion tag (e.g., "having fun" or "wanting to relax").

[0688] Additionally, the device uses a built-in emotion engine to recognize real-time emotions by analyzing the user's voice, facial expressions, gestures, etc. This emotion recognition data may be stored in the form of tags.

[0689] The data collected by the device is temporarily stored and then transmitted to a server using a secure communication protocol.

[0690] Data analysis and learning

[0691] The server stores the received data in a database, which includes the user's playback history, usage time period, activity patterns, emotion tags, and real-time emotion recognition data.

[0692] The server analyzes the stored data using machine learning algorithms to extract features that can help identify users' musical preferences and behavioral patterns.

[0693] The server also performs statistical analysis of multiple user data and learns general trends, enabling more accurate personalization.

[0694] Playlist Generation

[0695] The server then generates a playlist optimized for a specific user based on the analyzed data, including songs that reflect the user's time of day, activity patterns, emotion tags, and real-time emotions recognized by the emotion engine.

[0696] The server packages the information of the generated playlist (music ID, playback URL, etc.) and sends it to the terminal.

[0697] Providing playlists

[0698] The terminal displays the received playlists in the user interface of the music application for easy access by the user.

[0699] The user selects a song from the playlist and plays it. When a playback command is given, the terminal streams the song over the Internet.

[0700] Gather feedback and continue learning

[0701] The device collects the user's playback behavior (skip songs, complete playback, rating, etc.) and their emotions regarding those actions (including real-time emotion recognition) and records them as feedback.

[0702] The terminal transmits this feedback information to the server.

[0703] The server analyzes the received feedback and uses it to update the user's preferences, which then reflects the feedback when generating the next playlist, resulting in more accurate personalization.

[0704] As a concrete example, if a user feels "busy" while commuting on the train in the morning, the emotion engine will recognize this emotion in real time from the user's facial expressions and tone of voice. Based on past data and learning results, the server will select energetic, fast-paced music for that time period and generate a playlist. This playlist will be provided to the user via their device. Similarly, if the user wants to relax in the evening, the server will generate a playlist containing calming music, which will also be provided to the user via their device. This system enables a music experience optimized for the user's situation and emotions.

[0705] The processing flow will be explained below.

[0706] Step 1:

[0707] The user launches a music application and starts playing a song, while manually inputting an activity state (e.g., "commuting" or "playing sports") and an emotion tag (e.g., "having fun" or "wanting to relax").

[0708] Step 2:

[0709] The device automatically collects information about the song being played by the user (playback start time, play end time, song ID, etc.), and uses a built-in emotion engine to analyze the user's voice, facial expressions, and gestures, thereby recognizing the user's emotions in real time.

[0710] Step 3:

[0711] The device formats the collected data (playback history, usage time, activity patterns, manually entered emotion tags, and real-time emotion recognition data) into a specific format, encrypts it, and then sends it to a server using a secure communication protocol (e.g., HTTPS).

[0712] Step 4:

[0713] The server stores the received data in a database, which is then analyzed using machine learning algorithms to extract features that identify the user's musical preferences and behavioral patterns.

[0714] Step 5:

[0715] The server uses the listening history and the real-time emotions recognized by the emotion engine to generate an optimal playlist that suits the user's mood and activity at that moment.

[0716] Step 6:

[0717] The server then packages the generated playlist information (song ID, playback URL, etc.) and sends it to the device. This communication is also carried out using a secure protocol.

[0718] Step 7:

[0719] The terminal displays the received playlist on the user interface of the music application, allowing the user to access it directly.

[0720] Step 8:

[0721] The user selects a song from the playlist and plays it. When the device receives the instruction to play, it starts streaming the song over the Internet.

[0722] Step 9:

[0723] The device collects the user's playback behavior (skip songs, complete playback, rating specific songs, etc.) and real-time emotional data, and then transmits this feedback information back to the server.

[0724] Step 10:

[0725] The server analyzes the received feedback and updates the user's preferences. The updated information is reflected in the next playlist generation, and the system uses this information for further learning.

[0726] Example 2

[0727] Next, a description will be given of Example 2. In the following description, the data processing device 12 will be referred to as a "server" and the smart glasses 214 will be referred to as a "terminal."

[0728] Current music playlist generation systems generate playlists based on a user's playback history, time of day, and basic activity patterns. However, it is difficult to reflect the user's momentary emotional state or real-time preferences. This makes it difficult for users to consistently enjoy the optimal music experience. Furthermore, there are limitations to methods for effectively utilizing feedback to continuously improve the system. To address these issues, a system that can recognize and analyze a user's real-time emotions and provide highly accurate personalized playlists is needed.

[0729] The specific processing by the specific processing unit 290 of the data processing device 12 in the second embodiment is realized by the following means.

[0730] In this invention, the server includes means for collecting a user's playback history, usage time periods, activity patterns, and emotional states; means for analyzing the collected data and extracting the user's music preferences and behavioral patterns; means for generating a playlist based on the analysis results; means for providing the generated playlist to the user terminal; means for analyzing the user's voice, facial expressions, and gestures to recognize and record real-time emotions; and means for collecting feedback from users and training the system based on the feedback.

[0731] This allows for the generation of an optimal playlist that reflects the user's real-time emotions, making it possible to continue providing a music experience that matches the user's preferences.

[0732] "User" means an individual who utilizes the System to play music and obtain a personalized music experience.

[0733] "Playback history" refers to a record of songs that a user has played in the past.

[0734] A "time period of use" refers to a segment of time during which a user uses a music application.

[0735] An "activity pattern" refers to the behavior or situation in which a user is playing music (e.g., commuting, playing sports).

[0736] "Emotional state" refers to the emotion or mood of the user when playing music.

[0737] "Means of collection" refers to the methods and devices used to acquire and record the necessary data (playback history, time of use, activity patterns, emotional state, etc.).

[0738] "Means for analysis" refers to methods or devices that use machine learning algorithms to analyze collected data and extract meaningful information.

[0739] "Means for generating a playlist" refers to a method or device that creates a list of songs based on the analysis results and provides it to the user.

[0740] "Means for providing" refers to a method or device for delivering the generated playlist to a user's terminal and making it accessible.

[0741] An "emotion engine" is software or algorithms that analyze voice, facial expressions, and gestures to recognize and record a user's real-time emotions.

[0742] "Feedback collection means" refers to a method or device that captures and records a user's playback behavior and emotional responses.

[0743] "Means for system training" refers to methods and devices that analyze collected feedback and improve the personalization accuracy of the entire system.

[0744] The present invention is a system for personalizing a user's music experience, and in particular, provides a music playlist that reflects the user's real-time emotions by combining an emotion engine. This system is realized through the following elements and processes.

[0745] 1. Collection of User Data

[0746] The user launches the music application and selects an activity status such as "commuting" or "playing sports" or an emotion tag such as "fun" or "want to relax" within the application when starting use.

[0747] The device collects this information and uses an emotion engine (e.g., OpenFace or Microsoft Azure's Emotion API) that analyzes voice, facial expressions, and gestures to obtain real-time emotional data about the user.

[0748] 2. Data transmission and storage

[0749] The device temporarily stores the user's playback history, usage time period, activity patterns, emotion tags, and real-time emotion recognition data, and transmits them to the server using a secure communication protocol (e.g., HTTPS).

[0750] 3. Data analysis and learning

[0751] The server stores the transmitted data in a structured database, which includes each user's playback history, usage time periods, activity patterns, emotion tags, and real-time emotion recognition data.

[0752] The server analyzes the stored data using machine learning algorithms (e.g., TensorFlow or PyTorch) to extract features that identify the user's musical preferences and behavioral patterns.

[0753] The server also performs statistical analysis of multiple user data and learns general trends to improve the accuracy of personalization across the system.

[0754] 4. Generate a playlist

[0755] The server then generates a playlist optimized for a specific user based on the analysis results, for example, selecting energetic music for a user commuting in the morning and calming music for a user wanting to relax in the evening.

[0756] The server packages the generated playlist (music IDs, playback URLs, etc.) and sends it to the terminal.

[0757] 5. Providing playlists

[0758] The device displays the received playlists in the music application's user interface for easy access by the user, with the playlists visually categorized.

[0759] The user selects a song from the playlist and plays it. When a playback command is given, the terminal streams the song over the Internet.

[0760] 6. Gather feedback and continue learning

[0761] The device collects and records the user's playback behavior (e.g., skipping songs, completing playback, rating, etc.) and emotional responses.

[0762] The terminal transmits this feedback information to the server.

[0763] The server analyzes the received feedback and uses it to update the user's preferences, so that the feedback is reflected the next time a playlist is generated, resulting in more accurate personalization.

[0764] Examples of concrete examples and prompts

[0765] For example, if a user feels "busy" while commuting on the train in the morning, the emotion engine recognizes this emotion in real time from the user's facial expressions and tone of voice. Based on past data and learning results, the server selects energetic, fast-paced music for that time period and generates a playlist. This playlist is provided to the user via their device. On the other hand, when the user wants to relax in the evening, the server generates a playlist containing calming music, which is also provided to the user via their device.

[0766] An example of a prompt to input to a generative AI model is as follows:

[0767] "I'm on the train to work this morning and I'm feeling busy. Make me a playlist of energetic music."

[0768] "Generate the perfect playlist for a relaxing evening."

[0769] This will provide the user with an optimal music experience that suits their real-time emotions and situation.

[0770] The flow of the identification process in the second embodiment will be described with reference to FIG.

[0771] Step 1: Collect user data

[0772] input:

[0773] User play history

[0774] User usage time period

[0775] User activity patterns (e.g., commuting, playing sports)

[0776] User's emotional state (e.g., "I'm having fun" or "I want to relax")

[0777] process:

[0778] The user starts the music application and selects the playback history and various tags when starting use.

[0779] The device obtains real-time emotional data using an emotion engine (e.g., OpenFace or Microsoft Azure's Emotion API) that analyzes the user's voice, facial expressions, and gestures.

[0780] output:

[0781] User playback history data

[0782] Current time zone data

[0783] Activity pattern data

[0784] Real-time emotion data

[0785] Specific behavior:

[0786] The user operates a music application and inputs settings that make them feel "busy" during their commute, and the device uses its camera and microphone to provide data to the emotion engine, generating real-time emotion data.

[0787] Step 2: Send and store data

[0788] input:

[0789] User playback history data, usage time zone data, activity pattern data, and real-time emotion data collected in Step 1

[0790] process:

[0791] The terminal temporarily stores the collected data.

[0792] The saved data is sent to the server using a secure communication protocol (e.g. HTTPS).

[0793] output:

[0794] Data sent to the server

[0795] Specific behavior:

[0796] The device saves the data in local storage and then transmits it to the server using HTTPS.

[0797] Step 3: Analyze and train the data

[0798] input:

[0799] Data sent in step 2

[0800] process:

[0801] The server stores the submitted data in a structured database.

[0802] Machine learning algorithms (e.g., TensorFlow and PyTorch) are used to analyze users' musical preferences and behavioral patterns and extract specific features.

[0803] The server performs statistical analysis of multiple user data to learn general trends.

[0804] output:

[0805] User music taste characteristic data

[0806] General Trend Data

[0807] Specific behavior:

[0808] The server accumulates data in a database, analyzes it using a machine learning model, extracts features based on the user's preferences and behavioral patterns, and builds a model based on these.

[0809] Step 4: Generate a playlist

[0810] input:

[0811] The characteristic data of the user's music preferences and general tendency data extracted in step 3

[0812] process:

[0813] Based on the analysis results, the server generates a playlist that is optimal for the specific user.

[0814] The generated playlist information (song ID, playback URL, etc.) is packaged and sent to the terminal.

[0815] output:

[0816] Generated playlist data

[0817] Specific behavior:

[0818] The server combines each user's characteristic data with general tendency data to generate an optimal playlist, which is then packaged and sent to the terminal.

[0819] Step 5: Serve the playlist

[0820] input:

[0821] The playlist data generated in step 4

[0822] process:

[0823] The terminal displays the received playlist on the user interface of the music application.

[0824] The user selects and plays songs from the playlist.

[0825] The device streams music over the internet.

[0826] output:

[0827] Displayed playlist

[0828] Streaming songs

[0829] Specific behavior:

[0830] The playlist received by the device is displayed on the app, the user performs operations to play the music, and internet streaming is performed when playback begins.

[0831] Step 6: Gather feedback and continue learning

[0832] input:

[0833] User playback behavior data (e.g., skipping songs, completing playback, rating, etc.)

[0834] Real-time emotion data

[0835] process:

[0836] The terminal collects the user's playing behavior and emotional responses and records them as feedback.

[0837] The recorded feedback information is sent to the server.

[0838] The server analyzes the feedback and uses it to update the user's preferences.

[0839] output:

[0840] Updated user preference data

[0841] Improving the system's learning model

[0842] Specific behavior:

[0843] The device collects real-time playback behavior and emotional responses through an emotion engine and sends the information to a server, which analyzes it as learning material to improve the system's personalization accuracy.

[0844] Through this series of processes, an optimal music playlist that reflects the user's real-time emotions is generated and provided to the user, significantly improving the user's music experience.

[0845] (Application example 2)

[0846] Next, a description will be given of Application Example 2. In the following description, the data processing device 12 will be referred to as a "server" and the smart glasses 214 will be referred to as a "terminal."

[0847] Conventional music playlist generation systems generate playlists based only on a user's playback history, usage time, and activity patterns, which means they lack the ability to personalize playlists in response to real-time emotional changes. Furthermore, it is difficult to provide a music experience that takes into account emotions in a virtual store. This can lead to situations where the user is feeling stressed, such as not playing the right music, which can detract from the shopping experience.

[0848] The specific processing by the specific processing unit 290 of the data processing device 12 in the application example 2 is realized by the following means.

[0849] In this invention, the server includes means for collecting a user's playback history, usage time period, activity pattern, and emotional state, means for analyzing the user's facial expressions and voice to recognize emotions in real time, and means for analyzing the collected data and real-time emotion recognition data to extract the user's music preferences and behavioral patterns. This makes it possible to generate a playlist based on real-time emotion recognition, and to provide optimal music that takes into account the user's emotions, particularly in a shopping experience in a virtual store.

[0850] The "playback history" is a history of songs and audio content that have been played by the user in the past.

[0851] The "usage time period" refers to the time or time period during which a user uses a music playlist.

[0852] An "activity pattern" refers to a user's daily activity or behavior pattern, such as a situation such as "commuting" or "exercising."

[0853] An "emotional state" is the emotion or mood a user is experiencing at a particular time.

[0854] An "emotion engine" is a technology that analyzes a user's facial expressions and voice to recognize emotions in real time.

[0855] A "user terminal" is a device used by a user, such as a smartphone, smart glasses, or head-mounted display.

[0856] "Feedback" refers to the user's playback behavior (skip songs, complete playback, rating, etc.) and their feelings about it.

[0857] A "playlist" is a list of songs that a user has selected to play consecutively.

[0858] A "virtual store" is a shopping environment provided in a virtual space where users can browse and purchase items from the store via digital devices.

[0859] "Real-time emotion recognition data" is data obtained by analyzing the user's current emotions in real time.

[0860] This invention is a system for personalizing the music experience in a virtual store using real-time emotion recognition data. The aim is to improve the user experience by generating and providing an optimal music playlist according to the user's emotional state while shopping in the virtual store through smart glasses.

[0861] 1. System configuration and programs

[0862] The system includes the following components:

[0863] Smart glasses equipped with an emotion recognition engine

[0864] Server that collects and analyzes data

[0865] A server that generates and provides playlists

[0866] The emotion recognition engine is integrated into the smart glasses and uses OpenCV and deep learning models to analyze the user's facial expressions and voice. Server technologies using frameworks such as Flask and Django are used for data collection and analysis.

[0867] 2. Processing Details

[0868] The emotion recognition engine captures the user's facial expressions through the camera and analyzes the data in real time. The analyzed emotional data, along with the user's playback history, usage time, and activity patterns, are sent to the server. The server then uses machine learning algorithms based on this data to extract the user's musical preferences and behavioral patterns.

[0869] The server then generates a playlist optimized for the specific user based on the emotion recognition data and the extracted music preference data. This playlist consists of songs that suit the user's current emotions. The generated playlist is then sent to the smart glasses via the Internet.

[0870] 3. Specific examples

[0871] For example, if a user feels stressed while shopping in a virtual store, the smart glasses' emotion recognition engine will recognize that emotion in real time. The emotion recognition data will be sent to the server, which will then generate a playlist containing music that promotes relaxation. This playlist will then be sent to the smart glasses, reducing the user's stress and providing a comfortable shopping experience.

[0872] An example of a prompt to input to a generative AI model would be:

[0873] Analyze in real time whether the user is currently feeling stressed and generate an appropriate relaxing music playlist based on that emotion. Example: If the user is feeling stressed, create a list of songs that are particularly effective for relaxing and play them through smart glasses.

[0874] In this way, the system realized by the present invention provides a music playlist that corresponds to the user's emotional state, thereby improving the shopping experience in a virtual store.

[0875] The flow of the specific processing in the application example 2 will be described with reference to FIG.

[0876] Step 1:

[0877] Smart glasses capture the user's facial expressions and voice.

[0878] Input: User's face image and voice.

[0879] Output: Captured face image data and audio data.

[0880] Specific operation: Uses the camera and microphone of the smart glasses to collect the user's face and voice in real time.

[0881] Step 2:

[0882] The emotion recognition engine analyzes the captured facial images and voice to recognize the user's emotions.

[0883] Input: Captured facial image data and audio data.

[0884] Output: Real-time emotion recognition data (e.g. "stressed", "relaxed", etc.).

[0885] How it works: It uses OpenCV and deep learning models to perform facial expression analysis and speech analysis, such as analyzing facial muscle movements and tone of voice, to create appropriate emotion tags.

[0886] Step 3:

[0887] The smart glasses send emotion recognition data, playback history, usage time, and activity patterns to a server.

[0888] Input: Real-time emotion recognition data, playback history, time of day usage, activity patterns.

[0889] Output: The aggregate data sent to the server.

[0890] Specific operation: Collected data is sent to the server using a secure protocol (e.g., HTTPS).

[0891] Step 4:

[0892] The server analyzes the received data and extracts the user's musical preferences and behavioral patterns.

[0893] Input: Playback history, time of day, activity patterns, real-time emotion recognition data.

[0894] Output: Music preference data and behavioral pattern data as analysis results.

[0895] Specific operation: Using machine learning algorithms (e.g., clustering algorithms and recommendation systems), analyze users' musical preferences and behavioral patterns.

[0896] Step 5:

[0897] The server generates a playlist based on the analysis results and real-time emotion recognition data.

[0898] Input: Music preference data, behavioral pattern data, real-time emotion recognition data.

[0899] Output: The generated playlist.

[0900] Specific operation: The analysis results are combined with emotional data to generate a list of songs that are optimal for the user, selecting songs that correspond to a specific emotional state (e.g., stress).

[0901] Step 6:

[0902] The server provides the generated playlist to the smart glasses.

[0903] Input: The generated playlist.

[0904] Output: Playlist data sent to the smart glasses.

[0905] Specific operation: The playlist data is packaged and sent to the smart glasses using a secure communication protocol.

[0906] Step 7:

[0907] The smart glasses play the received playlist.

[0908] Input: Playlist data sent to smart glasses.

[0909] Output: The song that is playing.

[0910] Specific operation: The received playlist is displayed in the user interface, and when the user selects it, the song is streamed over the Internet.

[0911] Step 8:

[0912] Collect user feedback and send it to the server.

[0913] Input: User playback behavior data (e.g., song skip, playback completion, song rating, etc.).

[0914] Output: Feedback data sent to the server.

[0915] Specific operation: Records various playback actions of users in real time and periodically sends the information to the server.

[0916] Step 9:

[0917] The server analyzes the feedback and trains the system.

[0918] Input: Feedback data.

[0919] Output: Updated music preference data and behavioral pattern data.

[0920] Specific operation: The received feedback data is analyzed using machine learning algorithms and updated to improve the system's personalization accuracy.

[0921] Example prompts to be input to the generative AI model:

[0922] Analyze in real time whether the user is currently feeling stressed and generate an appropriate relaxing music playlist based on that emotion. Example: If the user is feeling stressed, create a list of songs that are particularly effective for relaxing and play them through smart glasses.

[0923] The specific processing unit 290 transmits the result of the specific processing to the smart glasses 214. In the smart glasses 214, the control unit 46A causes the speaker 240 to output the result of the specific processing. The microphone 238 acquires audio indicating a user input regarding the result of the specific processing. The control unit 46A transmits audio data indicating the user input acquired by the microphone 238 to the data processing device 12. In the data processing device 12, the specific processing unit 290 acquires the audio data.

[0924] The data generation model 58 is a so-called generative AI (Artificial Intelligence). An example of the data generation model 58 is ChatGPT (Internet Search<URL: https: / / openai.com / blog / chatgpt> ), Gemini (Internet search <url: https: gemini.google.com ?hl="ja">) and other generation AIs. The data generation model 58 is obtained by performing deep learning on a neural network. A prompt including an instruction is input to the data generation model 58, and inference data such as voice data indicating voice, text data indicating text, and image data indicating an image is also input. The data generation model 58 performs inference on the input inference data in accordance with the instruction indicated by the prompt, and outputs the inference result in a data format such as voice data and text data. Here, inference refers to, for example, analysis, classification, prediction, and / or summarization.

[0925] In the above embodiment, an example in which the specific processing is performed by the data processing device 12 has been given, but the technology of the present disclosure is not limited to this, and the specific processing may be performed by the smart glasses 214.

[0926] [Third embodiment]

[0927] FIG. 5 shows an example of the configuration of a data processing system 310 according to the third embodiment.

[0928] 5, the data processing system 310 includes the data processing device 12 and a headset type terminal 314. An example of the data processing device 12 is a server.

[0929] The data processing device 12 includes a computer 22, a database 24, and a communication I / F 26. The computer 22 is an example of a "computer" according to the technology of the present disclosure. The computer 22 includes a processor 28, a RAM 30, and a storage 32. The processor 28, the RAM 30, and the storage 32 are connected to a bus 34. The database 24 and the communication I / F 26 are also connected to the bus 34. The communication I / F 26 is connected to a network 54. Examples of the network 54 include a WAN (Wide Area Network) and / or a LAN (Local Area Network).

[0930] The headset type terminal 314 includes a computer 36, a microphone 238, a speaker 240, a camera 42, a communication I / F 44, and a display 343. The computer 36 includes a processor 46, a RAM 48, and a storage 50. The processor 46, the RAM 48, and the storage 50 are connected to a bus 52. The microphone 238, the speaker 240, the camera 42, and the display 343 are also connected to the bus 52.

[0931] The microphone 238 receives instructions and the like from the user 20 by receiving voice uttered by the user 20. The microphone 238 captures the voice uttered by the user 20, converts the captured voice into audio data, and outputs it to the processor 46. The speaker 240 outputs audio in accordance with instructions from the processor 46.

[0932] Camera 42 is a small digital camera equipped with an optical system including a lens, aperture, and shutter, and an imaging element such as a CMOS (Complementary Metal-Oxide-Semiconductor) image sensor or a CCD (Charge Coupled Device) image sensor, and captures images of the surroundings of user 20 (for example, an imaging range defined by an angle of view equivalent to the field of vision of a typical healthy person).

[0933] The communication I / F 44 is connected to a network 54. The communication I / Fs 44 and 26 are responsible for the exchange of various information between the processor 46 and the processor 28 via the network 54. The exchange of various information between the processor 46 and the processor 28 using the communication I / Fs 44 and 26 is carried out in a secure state.

[0934] Fig. 6 shows an example of the main functions of the data processing device 12 and the headset type terminal 314. As shown in Fig. 6, in the data processing device 12, a specific process is performed by the processor 28. A specific process program 56 is stored in the storage 32.

[0935] The specific processing program 56 is an example of a "program" according to the technology of the present disclosure. The processor 28 reads the specific processing program 56 from the storage 32 and executes the read specific processing program 56 on the RAM 30. The specific processing is realized by the processor 28 operating as a specific processing unit 290 in accordance with the specific processing program 56 executed on the RAM 30.

[0936] The storage 32 stores a data generation model 58 and an emotion identification model 59. The data generation model 58 and the emotion identification model 59 are used by the identification processing unit 290.

[0937] In the headset type terminal 314, a reception output process is performed by the processor 46. A reception output program 60 is stored in the storage 50. The processor 46 reads the reception output program 60 from the storage 50 and executes the read reception output program 60 on the RAM 48. The reception output process is realized by the processor 46 operating as the control unit 46A in accordance with the reception output program 60 executed on the RAM 48.

[0938] Next, a description will be given of the identification process performed by the identification processing unit 290 of the data processing device 12. In the following description, the data processing device 12 will be referred to as the "server" and the headset type terminal 314 will be referred to as the "terminal."

[0939] This invention relates to a system that personalizes and optimizes a user's music experience. Specifically, it is a method for generating and providing playlists based on the user's preferences and behavioral patterns by collecting and analyzing the user's playback history, usage time period, activity patterns, and emotional state. Furthermore, by collecting user feedback and allowing the system to continuously learn, it provides a more accurate and personalized music experience.

[0940] User Data Collection

[0941] A user starts a music application and plays music in the course of daily life.

[0942] The device collects the user's music history, the start and end times of the music, whether the music was skipped, the user's current activity status (e.g., commuting, exercising), and emotion tags, which may be entered by the user in the form of "fun" or "want to relax."

[0943] The terminal converts this data into a specified format and transmits it to the server using a secure communication method.

[0944] Data analysis and learning

[0945] The server stores the received user data in a database.

[0946] The server analyzes the stored data and extracts features to identify users' musical preferences and behavioral patterns. This analysis uses machine learning algorithms. For example, a clustering algorithm can group users' activity patterns.

[0947] The server performs statistical analysis based on a large amount of user data and also learns general trends.

[0948] Playlist Generation

[0949] The server then generates the optimal playlist for a specific user based on the analysis results, which means selecting songs of different genres and moods according to the user's time of day, activity patterns, and emotional tags.

[0950] The server packages the list of songs included in the generated playlist and its associated data (song ID, playback URL, etc.) and sends it to the terminal.

[0951] Providing playlists

[0952] The terminal displays the received playlist on the application's user interface for easy access by the user.

[0953] The user selects a song from the playlist and plays it. When a playback command is given, the terminal streams the song over the Internet.

[0954] Gather feedback and continue learning

[0955] The device collects the songs the user has played, whether they have completed or skipped, and their ratings and feedback (e.g., "favorite marks") for specific songs.

[0956] The terminal transmits this feedback information to the server.

[0957] The server updates the user's preferences based on the received feedback and then reruns the analysis and learning process, which is then reflected in the next playlist generation.

[0958] As a concrete example, if a user sets the emotion tag "busy" during their morning commute, the server will select energetic, fast-paced music for that time period based on past data and learning results. The generated playlist is provided to the device and played by the user. On the other hand, when it's time to relax in the evening, the server will generate a playlist containing calming music and provide it to the user via the device. In this way, it is possible to provide a music experience that is optimal for the user's situation and emotion.

[0959] The processing flow will be explained below.

[0960] Step 1:

[0961] A user launches a music application and starts playing a song. At the same time, the user sets an activity state (e.g., "commuting" or "playing sports") and an emotion tag (e.g., "having fun" or "wanting to relax").

[0962] Step 2:

[0963] The device automatically collects information about the song the user is playing (playback start time, play end time, song ID, etc.), as well as the user's activity status and emotion tag. The collected data is temporarily stored on the device.

[0964] Step 3:

[0965] The device formats the collected user data into a specific format, encrypts it, and then sends it to a server using a secure protocol such as HTTPS.

[0966] Step 4:

[0967] The server stores the received user data in a database, and then starts a process to analyze the stored data. This analysis uses machine learning algorithms.

[0968] Step 5:

[0969] The server extracts features from the analyzed data to identify users' musical preferences and behavioral patterns, for example by using a clustering algorithm to group users based on their activity patterns and emotions.

[0970] Step 6:

[0971] The server selects music suitable for a specific time period or activity level and generates an optimal playlist for each user, for example, energetic music for the commute and quieter music for relaxation.

[0972] Step 7:

[0973] The server packages the generated playlist information (song ID, playback URL, etc.) and sends it to the device. This communication also uses a secure protocol.

[0974] Step 8:

[0975] The device displays the received playlist in the user interface of the music application, allowing the user to select and play it.

[0976] Step 9:

[0977] The user selects a song from the presented playlist and starts playing it. The device streams the song over the Internet.

[0978] Step 10:

[0979] The device collects the user's playback behavior (skip songs, complete playback, rating, etc.) and records it as feedback.

[0980] Step 11:

[0981] The terminal transmits the collected feedback information to the server.

[0982] Step 12:

[0983] The server analyzes the received feedback and uses it as data to update the user's preferences, which allows for more accurate personalization when generating the next playlist.

[0984] Example 1

[0985] Next, a description will be given of Example 1. In the following description, the data processing device 12 will be referred to as a "server" and the headset type terminal 314 will be referred to as a "terminal."

[0986] Conventional music recommendation systems have difficulty generating playlists that fully consider a user's specific usage situation and emotions. Furthermore, they lack a system that can properly learn from user feedback and reflect it in future recommendations. As a result, users are unable to obtain the music experience that best suits their preferences and situation.

[0987] The specific processing by the specific processing unit 290 of the data processing device 12 in the first embodiment is realized by the following means.

[0988] In this invention, the server includes: means for collecting a user's playback history, usage time periods, activity patterns, and emotional state; means for formatting the collected data into a specified format and transmitting it to the server using a secure communication method; means for saving the received user data, analyzing it using a machine learning algorithm, and extracting the user's music preferences and behavioral patterns; means for generating an optimal playlist for a specific user based on the analysis results; means for packaging a list of songs included in the generated playlist and their associated data and transmitting it to a user terminal; means for displaying the received playlist on a user interface for easy user access; and means for collecting user feedback and re-executing the analysis and learning process based on the feedback, thereby enabling a more personalized music experience based on the user's usage history and emotional tags.

[0989] "Playback history" is information about songs that the user has played in the past, and includes song IDs, titles, artist names, and the like.

[0990] The "usage time period" is information about the date and time when the user used the music application, including the playback start time and end time.

[0991] The "activity pattern" refers to the activity status of the user while playing music, and indicates a specific operating state such as commuting, exercising, or relaxing.

[0992] The "emotional state" refers to the emotional or mood state of the user when listening to music, and is expressed in the form of emotion tags such as "fun" or "want to relax."

[0993] A "secure communication method" is a communication protocol that ensures safety when sending and receiving data, and a specific example is HTTPS.

[0994] "Machine learning algorithms" are algorithms that analyze large amounts of data and extract patterns and rules, and include clustering algorithms.

[0995] A "playlist" is a collection of multiple songs, and is a format for providing a collection of songs that meets the user's playback needs.

[0996] A "user interface" is a display screen that allows a user to directly interact with a system or application, and includes a song list, operation buttons, and the like.

[0997] "Feedback" refers to information provided by users to the system, such as playback history, song ratings and comments, which is reflected in the system's learning process.

[0998] The "analysis and learning process" is a series of steps that the system takes to learn the user's preferences and behavioral patterns based on data collected from the user, and then generate a new playlist.

[0999] This invention relates to a system that personalizes and optimizes a user's music experience. This system collects and analyzes a user's playback history, usage time, activity patterns, and emotional state, and generates and provides playlists based on the user's preferences and behavioral patterns. Furthermore, by collecting user feedback and continuously learning, the system provides an even more accurate and personalized music experience.

[1000] First, a user launches a music application and plays a song in their daily life. For example, a user presses the play button on the application to play a song. Here, the device collects the history of songs played by the user, the start and end times of playback, whether songs were skipped, the user's current activity state (commuting, exercising, etc.), and emotion tags (e.g., happy, want to relax). These emotion tags are input by the user.

[1001] The device then formats the collected data into a specified format and sends it to the server using HTTPS, a secure communication method. At this time, the data is converted into JSON format for efficient data processing.

[1002] The server stores the received user data in a database. For example, the server inserts data into the database using SQL queries. The server then analyzes the stored data using machine learning algorithms to identify users' musical preferences and behavioral patterns. This analysis uses the Python library scikit-learn, which runs a clustering algorithm to extract specific patterns.

[1003] Next, the server generates a playlist optimized for the specific user based on the analysis results. This means selecting songs of different genres and moods depending on the user's time of day, activity patterns, and emotional tags. The generated playlist is converted into JSON format and sent back to the device using the HTTPS protocol.

[1004] The device displays the received playlist on the application's user interface for easy access by the user. For example, the user can click a song in the playlist and press the play button to start streaming the song. Specifically, the device connects to the streaming server and streams the music using the playback URL.

[1005] Additionally, when a user provides feedback such as which songs they have completed playing, which songs they have skipped, or their ratings for specific songs, the device collects this feedback information and sends it to the server. The server then updates the user's preferences based on the received feedback and reruns the analysis and learning process, allowing the information to be reflected in the next playlist generation.

[1006] As a concrete example, if a user sets the emotion tag "busy" during their morning commute, the server will select energetic, fast-paced music for that time period based on past data and learning results. The generated playlist is provided to the device and played by the user. On the other hand, when it's time to relax in the evening, the server will generate a playlist containing calming music and provide it to the user via the device. In this way, it is possible to provide a music experience that is optimal for the user's situation and emotion.

[1007] Examples of prompt statements

[1008] "Please explain in detail the data analysis process that goes into generating an energizing playlist for your morning commute."

[1009] "How do I update my personalized music playlist based on user feedback?"

[1010] "Please explain in detail how you would generate an evening playlist based on the emotion tag 'I want to relax'."

[1011] The flow of the identification process in the first embodiment will be described with reference to FIG.

[1012] Step 1: Collect user data

[1013] In everyday life, users launch music applications and play music.

[1014] Specifically, the user presses the play button on the application on a device such as a smartphone or computer.

[1015] The device collects the user's history of songs played, the start and end times of playback, whether songs were skipped, the user's current activity status (commuting, exercising, etc.), and emotion tags (e.g., fun, want to relax).

[1016] The input is various data when the user operates the playlist.

[1017] As part of the data processing, the terminal records these data in its internal memory.

[1018] The output is formatted data, ready to be sent to the next step.

[1019] Step 2: Sending data

[1020] The terminal formats the collected data into the specified format.

[1021] The input is raw, unformatted data.

[1022] To process the data, the terminal converts the data into JSON format.

[1023] The output is data in JSON format.

[1024] The terminal transmits data to the server using a secure communication method (e.g., HTTPS).

[1025] Specifically, the terminal sends JSON format data to the server as a POST request using the HTTPS protocol.

[1026] Step 3: Data storage and analysis

[1027] The server stores the received user data in a database.

[1028] The input is JSON format data sent from the terminal.

[1029] For data processing, the server receives this data and inserts it into a database using an SQL query.

[1030] The output is user data stored in a database.

[1031] The server analyzes the stored data using machine learning algorithms (e.g., clustering algorithms).

[1032] The input is user data stored in a database.

[1033] For data calculation, the server applies machine learning models using Python libraries (e.g., scikit-learn) to identify users' musical preferences and behavioral patterns.

[1034] The output is the user's characteristics and pattern information as the analysis results.

[1035] Step 4: Generate a playlist

[1036] The server generates an optimal playlist for a particular user based on the analysis results.

[1037] The input is the user's characteristics and pattern information as the analysis result.

[1038] To process the data, the server narrows down the conditions based on the time of use, activity patterns, and emotion tags, and selects appropriate songs from the music database.

[1039] The output is the generated playlist.

[1040] The server packages the list of songs included in the generated playlist and its associated data (song ID, playback URL, etc.) and sends it to the terminal.

[1041] Specifically, the server converts the playlist into JSON format and sends a POST request to the terminal using the HTTPS protocol.

[1042] Step 5: Serve the playlist

[1043] The terminal displays the received playlist on the user interface of the application, making it easily accessible to the user.

[1044] The input is the playlist data in JSON format sent from the server.

[1045] For data processing, the device parses the JSON data and dynamically displays a song list in the UI component.

[1046] The output is the displayed playlist.

[1047] The user selects a song from the presented playlist and plays it.

[1048] Specifically, the user clicks on a song in the playlist and presses the play button.

[1049] The device streams the music over the Internet.

[1050] The input is the URL to play the song.

[1051] For data processing, the terminal connects to a streaming server and plays music.

[1052] Step 6: Gather feedback and continue learning

[1053] The device collects songs that the user has completed playing, songs that the user has skipped, and ratings and feedback for specific songs.

[1054] The input is the feedback information provided by the user.

[1055] As for data processing, the terminal tracks the user's operations and stores them in data storage.

[1056] The output is the collected feedback data.

[1057] The terminal transmits this feedback information to the server.

[1058] Specifically, the terminal formats the feedback data into JSON format and sends a POST request to the server using the HTTPS protocol.

[1059] The server updates the user's preferences based on the received feedback and then reruns the analysis and learning process.

[1060] The input is the feedback data sent from the terminal.

[1061] As the data is processed, the server stores it in a database and retrains the machine learning model.

[1062] The output is updated analysis results and data that will be reflected in the next playlist generation.

[1063] (Application example 1)

[1064] Next, a description will be given of Application Example 1. In the following description, the data processing device 12 will be referred to as a "server" and the headset type terminal 314 will be referred to as a "terminal."

[1065] Conventionally, methods for optimizing the efficiency and accuracy of factory robots have relied solely on motion algorithms and machine learning. However, it is believed that even more efficient work can be achieved by providing an appropriate environment based on the robot's activity status and work content. The present invention aims to improve the efficiency and accuracy of work by playing personalized music while the robot is working.

[1066] The specific processing by the specific processing unit 290 of the data processing device 12 in the application example 1 is realized by the following means.

[1067] In this invention, the server includes means for collecting a user's playback history, usage time periods, activity patterns, and emotional state, means for analyzing the collected data and extracting the user's preferences and behavioral patterns, means for generating a playlist based on the analysis results, means for providing the generated playlist to the terminal, means for collecting the terminal's work history and activity level and generating an optimal music experience based on the data, means for collecting feedback and conducting system training, and means for generating applications for use in various facilities. This makes it possible to play optimal music according to the robot's activity status and work content, effectively improving work efficiency and accuracy.

[1068] "Playback history" is a record of music and content that a user has played in the past.

[1069] A "time period of use" refers to a specific time range during which a user uses music or content.

[1070] An "activity pattern" is a series of patterns of actions or states of a user or a terminal.

[1071] "Emotional state" refers to the user's current emotional or mood state.

[1072] "Analysis" is the process of examining collected data in detail to identify trends and characteristics.

[1073] A "playlist" refers to a list of songs or content to be played.

[1074] A "terminal" refers to an electronic device used by a user, and includes, for example, a smartphone or a robot.

[1075] "Work history" is a record of work performed by a terminal or robot.

[1076] "Activity level" is a measure that indicates the activity and intensity of the terminal or robot's movements.

[1077] "Feedback" refers to responses or opinions from users or systems.

[1078] A "clustering algorithm" is a machine learning technique for classifying data into several groups.

[1079] "System learning" is the process by which a system improves itself based on collected data and feedback.

[1080] An "application" is software or a program that provides a specific function.

[1081] This invention is a system that automatically plays appropriate music to improve the work efficiency and accuracy of robots working in factories. The system collects the user's playback history, usage time, activity patterns, and emotional state, and generates and provides a playlist based on the analysis results. It also has a function that collects the device's work history and activity level to generate optimal music.

[1082] Generating a Program

[1083] The system program includes the following main functions:

[1084] 1. Data Collection:

[1085] Sensors are used to collect records of tasks performed by the device, work time, and activity levels.

[1086] The user's emotional state and usage time period are also collected (for example, music playback history while working).

[1087] 2. Data Analysis:

[1088] The collected data is stored and analyzed on the server using Python.

[1089] The robot's activity patterns are classified using scikit-learn's clustering algorithm (KMeans).

[1090] Features are extracted from the data and music appropriate for the robot's task is selected.

[1091] 3. Playlist generation:

[1092] Generate the best playlist for a specific task based on the analysis results.

[1093] The playlist includes energetic and fast-paced songs, as well as calm and rhythmic songs.

[1094] 4. Playlist provided by:

[1095] The generated playlist is sent to the terminal and a playback instruction is issued.

[1096] 5. Feedback and learning:

[1097] The device collects the history of songs played and feedback from users and sends it to the server.

[1098] The system will self-learn based on the feedback and reflect this in the next playlist generation.

[1099] Processing Details

[1100] Hardware:

[1101] The robot is equipped with a head-mounted display (HMD) and built-in speakers.

[1102] Sensors for data collection (e.g., activity trackers and temperature sensors).

[1103] software:

[1104] Data analysis and learning program using Python.

[1105] Clustering algorithm using the scikit-learn library.

[1106] Data storage using a database (e.g., MongoDB).

[1107] Specific examples

[1108] For example, if a robot performing assembly work is engaged in the task "assembly," it will play lively, fast-paced music to improve the efficiency of that work. In this case, the server selects appropriate music based on the data collected so far, creates a playlist, and sends it to the device. It also receives feedback and selects the best music for the next task.

[1109] Prompt Sentence Examples

[1110] You can instruct the generative AI model to collect data and generate a playlist using prompts like the following:

[1111] Create a program for the data collection part. Collect the task ID, work time, and activity level while the robot is working, and save them in a data frame format.

[1112]

[1113] Please create a program for generating the playlist. Create an optimal playlist based on the robot's task ID and return it in list format.

[1114]

[1115] Write a program that provides the generated playlist. Design a system that plays all the songs in the playlist in order.

[1116] In this way, the present invention can optimize the efficiency and accuracy of the robot's work and provide an effective musical experience.

[1117] The flow of the specific processing in the application example 1 will be described with reference to FIG.

[1118] Step 1:

[1119] The device collects the user's or robot's playback history, usage time period, activity patterns, and emotional state from sensors and stores them in a data frame format. The input is sensor data, and the output is a formatted data frame. The device formats the collected data into a specified format and transmits it to the server using a secure communication method.

[1120] Step 2:

[1121] The server stores the data received from the terminal in a database. The input is the data sent from the terminal, and the output is the data stored in the database. The server performs a write operation to the database and checks whether the data is saved correctly.

[1122] Step 3:

[1123] The server analyzes the stored data and classifies the activity patterns of the user or robot using a clustering algorithm (e.g., KMeans in scikit-learn). The input is the data in the database, and the output is the classification results for each cluster. Data preprocessing and feature extraction are also performed during the analysis process.

[1124] Step 4:

[1125] The server generates a playlist optimized for a specific task based on the analysis results. The input is the clustering results and user preference data, and the output is the generated playlist. Specifically, it selects appropriate songs based on the data belonging to a specific cluster.

[1126] Step 5:

[1127] The server provides the generated playlist to the terminal. The input is the generated playlist, and the output is the playlist data sent to the terminal. The server packages and sends the list of songs included in the playlist and their associated data (song IDs, playback URLs, etc.).

[1128] Step 6:

[1129] The terminal displays the received playlist on the user interface, making it playable by the user or robot. The input is the playlist data received from the server, and the output is a playable playlist on the user interface. The terminal starts streaming the music over the Internet.

[1130] Step 7:

[1131] The device collects data during playback (e.g., playback completion, skip, rating) and feedback from the user. The input is user operations during playback and feedback data, and the output is the collected feedback data. The device securely transmits the feedback data to the server.

[1132] Step 8:

[1133] The server analyzes the received feedback data and continues to train the system. The input is the collected feedback data, and the output is an updated learning model for the system. The server reflects the feedback and uses it to generate the next playlist.

[1134] Through these processing steps, a system that optimizes the user's work efficiency can be effectively realized.

[1135] Furthermore, an emotion engine that estimates the user's emotion may be further combined. That is, the identification processing unit 290 may estimate the user's emotion using the emotion identification model 59, and perform identification processing using the user's emotion.

[1136] This invention is a system for personalizing the music experience, and by combining it with an emotion engine that recognizes the user's emotions, it provides a more accurate music playlist. Specifically, it collects the user's playback history, usage time period, activity patterns, and emotional state, and then uses the emotion engine to recognize and analyze the user's real-time emotions, and generates and provides a playlist based on that data.

[1137] User Data Collection

[1138] A user launches a music application and starts playing a song. At the same time, the user sets an activity state (e.g., "commuting" or "playing sports") and an emotion tag (e.g., "having fun" or "wanting to relax").

[1139] Additionally, the device uses a built-in emotion engine to recognize real-time emotions by analyzing the user's voice, facial expressions, gestures, etc. This emotion recognition data may be stored in the form of tags.

[1140] The data collected by the device is temporarily stored and then transmitted to a server using a secure communication protocol.

[1141] Data analysis and learning

[1142] The server stores the received data in a database, which includes the user's playback history, usage time period, activity patterns, emotion tags, and real-time emotion recognition data.

[1143] The server analyzes the stored data using machine learning algorithms to extract features that can help identify users' musical preferences and behavioral patterns.

[1144] The server also performs statistical analysis of multiple user data and learns general trends, enabling more accurate personalization.

[1145] Playlist Generation

[1146] The server then generates a playlist optimized for a specific user based on the analyzed data, including songs that reflect the user's time of day, activity patterns, emotion tags, and real-time emotions recognized by the emotion engine.

[1147] The server packages the information of the generated playlist (music ID, playback URL, etc.) and sends it to the terminal.

[1148] Providing playlists

[1149] The terminal displays the received playlists in the user interface of the music application for easy access by the user.

[1150] The user selects a song from the playlist and plays it. When a playback command is given, the terminal streams the song over the Internet.

[1151] Gather feedback and continue learning

[1152] The device collects the user's playback behavior (skip songs, complete playback, rating, etc.) and their emotions regarding those actions (including real-time emotion recognition) and records them as feedback.

[1153] The terminal transmits this feedback information to the server.

[1154] The server analyzes the received feedback and uses it to update the user's preferences, which then reflects the feedback when generating the next playlist, resulting in more accurate personalization.

[1155] As a concrete example, if a user feels "busy" while commuting on the train in the morning, the emotion engine will recognize this emotion in real time from the user's facial expressions and tone of voice. Based on past data and learning results, the server will select energetic, fast-paced music for that time period and generate a playlist. This playlist will be provided to the user via their device. Similarly, if the user wants to relax in the evening, the server will generate a playlist containing calming music, which will also be provided to the user via their device. This system enables a music experience optimized for the user's situation and emotions.

[1156] The processing flow will be explained below.

[1157] Step 1:

[1158] The user launches a music application and starts playing a song, while manually inputting an activity state (e.g., "commuting" or "playing sports") and an emotion tag (e.g., "having fun" or "wanting to relax").

[1159] Step 2:

[1160] The device automatically collects information about the song being played by the user (playback start time, play end time, song ID, etc.), and uses a built-in emotion engine to analyze the user's voice, facial expressions, and gestures, thereby recognizing the user's emotions in real time.

[1161] Step 3:

[1162] The device formats the collected data (playback history, usage time, activity patterns, manually entered emotion tags, and real-time emotion recognition data) into a specific format, encrypts it, and then sends it to a server using a secure communication protocol (e.g., HTTPS).

[1163] Step 4:

[1164] The server stores the received data in a database, which is then analyzed using machine learning algorithms to extract features that identify the user's musical preferences and behavioral patterns.

[1165] Step 5:

[1166] The server uses the listening history and the real-time emotions recognized by the emotion engine to generate an optimal playlist that suits the user's mood and activity at that moment.

[1167] Step 6:

[1168] The server then packages the generated playlist information (song ID, playback URL, etc.) and sends it to the device. This communication is also carried out using a secure protocol.

[1169] Step 7:

[1170] The terminal displays the received playlist on the user interface of the music application, allowing the user to access it directly.

[1171] Step 8:

[1172] The user selects a song from the playlist and plays it. When the device receives the instruction to play, it starts streaming the song over the Internet.

[1173] Step 9:

[1174] The device collects the user's playback behavior (skip songs, complete playback, rating specific songs, etc.) and real-time emotional data, and then transmits this feedback information back to the server.

[1175] Step 10:

[1176] The server analyzes the received feedback and updates the user's preferences. The updated information is reflected in the next playlist generation, and the system uses this information for further learning.

[1177] Example 2

[1178] Next, a description will be given of Example 2. In the following description, the data processing device 12 will be referred to as a "server" and the headset type terminal 314 will be referred to as a "terminal."

[1179] Current music playlist generation systems generate playlists based on a user's playback history, time of day, and basic activity patterns. However, it is difficult to reflect the user's momentary emotional state or real-time preferences. This makes it difficult for users to consistently enjoy the optimal music experience. Furthermore, there are limitations to methods for effectively utilizing feedback to continuously improve the system. To address these issues, a system that can recognize and analyze a user's real-time emotions and provide highly accurate personalized playlists is needed.

[1180] The specific processing by the specific processing unit 290 of the data processing device 12 in the second embodiment is realized by the following means.

[1181] In this invention, the server includes means for collecting a user's playback history, usage time periods, activity patterns, and emotional states; means for analyzing the collected data and extracting the user's music preferences and behavioral patterns; means for generating a playlist based on the analysis results; means for providing the generated playlist to the user terminal; means for analyzing the user's voice, facial expressions, and gestures to recognize and record real-time emotions; and means for collecting feedback from users and training the system based on the feedback.

[1182] This allows for the generation of an optimal playlist that reflects the user's real-time emotions, making it possible to continue providing a music experience that matches the user's preferences.

[1183] "User" means an individual who utilizes the System to play music and obtain a personalized music experience.

[1184] "Playback history" refers to a record of songs that a user has played in the past.

[1185] A "time period of use" refers to a segment of time during which a user uses a music application.

[1186] An "activity pattern" refers to the behavior or situation in which a user is playing music (e.g., commuting, playing sports).

[1187] "Emotional state" refers to the emotion or mood of the user when playing music.

[1188] "Means of collection" refers to the methods and devices used to acquire and record the necessary data (playback history, time of use, activity patterns, emotional state, etc.).

[1189] "Means for analysis" refers to methods or devices that use machine learning algorithms to analyze collected data and extract meaningful information.

[1190] "Means for generating a playlist" refers to a method or device that creates a list of songs based on the analysis results and provides it to the user.

[1191] "Means for providing" refers to a method or device for delivering the generated playlist to a user's terminal and making it accessible.

[1192] An "emotion engine" is software or algorithms that analyze voice, facial expressions, and gestures to recognize and record a user's real-time emotions.

[1193] "Feedback collection means" refers to a method or device that captures and records a user's playback behavior and emotional responses.

[1194] "Means for system training" refers to methods and devices that analyze collected feedback and improve the personalization accuracy of the entire system.

[1195] The present invention is a system for personalizing a user's music experience, and in particular, provides a music playlist that reflects the user's real-time emotions by combining an emotion engine. This system is realized through the following elements and processes.

[1196] 1. Collection of User Data

[1197] The user launches the music application and selects an activity status such as "commuting" or "playing sports" or an emotion tag such as "fun" or "want to relax" within the application when starting use.

[1198] The device collects this information and uses an emotion engine (e.g., OpenFace or Microsoft Azure's Emotion API) that analyzes voice, facial expressions, and gestures to obtain real-time emotional data about the user.

[1199] 2. Data transmission and storage

[1200] The device temporarily stores the user's playback history, usage time period, activity patterns, emotion tags, and real-time emotion recognition data, and transmits them to the server using a secure communication protocol (e.g., HTTPS).

[1201] 3. Data analysis and learning

[1202] The server stores the transmitted data in a structured database, which includes each user's playback history, usage time periods, activity patterns, emotion tags, and real-time emotion recognition data.

[1203] The server analyzes the stored data using machine learning algorithms (e.g., TensorFlow or PyTorch) to extract features that identify the user's musical preferences and behavioral patterns.

[1204] The server also performs statistical analysis of multiple user data and learns general trends to improve the accuracy of personalization across the system.

[1205] 4. Generate a playlist

[1206] The server then generates a playlist optimized for a specific user based on the analysis results, for example, selecting energetic music for a user commuting in the morning and calming music for a user wanting to relax in the evening.

[1207] The server packages the generated playlist (music IDs, playback URLs, etc.) and sends it to the terminal.

[1208] 5. Providing playlists

[1209] The device displays the received playlists in the music application's user interface for easy access by the user, with the playlists visually categorized.

[1210] The user selects a song from the playlist and plays it. When a playback command is given, the terminal streams the song over the Internet.

[1211] 6. Gather feedback and continue learning

[1212] The device collects and records the user's playback behavior (e.g., skipping songs, completing playback, rating, etc.) and emotional responses.

[1213] The terminal transmits this feedback information to the server.

[1214] The server analyzes the received feedback and uses it to update the user's preferences, so that the feedback is reflected the next time a playlist is generated, resulting in more accurate personalization.

[1215] Examples of concrete examples and prompts

[1216] For example, if a user feels "busy" while commuting on the train in the morning, the emotion engine recognizes this emotion in real time from the user's facial expressions and tone of voice. Based on past data and learning results, the server selects energetic, fast-paced music for that time period and generates a playlist. This playlist is provided to the user via their device. On the other hand, when the user wants to relax in the evening, the server generates a playlist containing calming music, which is also provided to the user via their device.

[1217] An example of a prompt to input to a generative AI model is as follows:

[1218] "I'm on the train to work this morning and I'm feeling busy. Make me a playlist of energetic music."

[1219] "Generate the perfect playlist for a relaxing evening."

[1220] This will provide the user with an optimal music experience that suits their real-time emotions and situation.

[1221] The flow of the identification process in the second embodiment will be described with reference to FIG.

[1222] Step 1: Collect user data

[1223] input:

[1224] User play history

[1225] User usage time period

[1226] User activity patterns (e.g., commuting, playing sports)

[1227] User's emotional state (e.g., "I'm having fun" or "I want to relax")

[1228] process:

[1229] The user starts the music application and selects the playback history and various tags when starting use.

[1230] The device obtains real-time emotional data using an emotion engine (e.g., OpenFace or Microsoft Azure's Emotion API) that analyzes the user's voice, facial expressions, and gestures.

[1231] output:

[1232] User playback history data

[1233] Current time zone data

[1234] Activity pattern data

[1235] Real-time emotion data

[1236] Specific behavior:

[1237] The user operates a music application and inputs settings that make them feel "busy" during their commute, and the device uses its camera and microphone to provide data to the emotion engine, generating real-time emotion data.

[1238] Step 2: Send and store data

[1239] input:

[1240] User playback history data, usage time zone data, activity pattern data, and real-time emotion data collected in Step 1

[1241] process:

[1242] The terminal temporarily stores the collected data.

[1243] The saved data is sent to the server using a secure communication protocol (e.g. HTTPS).

[1244] output:

[1245] Data sent to the server

[1246] Specific behavior:

[1247] The device saves the data in local storage and then transmits it to the server using HTTPS.

[1248] Step 3: Analyze and train the data

[1249] input:

[1250] Data sent in step 2

[1251] process:

[1252] The server stores the submitted data in a structured database.

[1253] Machine learning algorithms (e.g., TensorFlow and PyTorch) are used to analyze users' musical preferences and behavioral patterns and extract specific features.

[1254] The server performs statistical analysis of multiple user data to learn general trends.

[1255] output:

[1256] User music taste characteristic data

[1257] General Trend Data

[1258] Specific behavior:

[1259] The server accumulates data in a database, analyzes it using a machine learning model, extracts features based on the user's preferences and behavioral patterns, and builds a model based on these.

[1260] Step 4: Generate a playlist

[1261] input:

[1262] The characteristic data of the user's music preferences and general tendency data extracted in step 3

[1263] process:

[1264] Based on the analysis results, the server generates a playlist that is optimal for the specific user.

[1265] The generated playlist information (song ID, playback URL, etc.) is packaged and sent to the terminal.

[1266] output:

[1267] Generated playlist data

[1268] Specific behavior:

[1269] The server combines each user's characteristic data with general tendency data to generate an optimal playlist, which is then packaged and sent to the terminal.

[1270] Step 5: Serve the playlist

[1271] input:

[1272] The playlist data generated in step 4

[1273] process:

[1274] The terminal displays the received playlist on the user interface of the music application.

[1275] The user selects and plays songs from the playlist.

[1276] The device streams music over the internet.

[1277] output:

[1278] Displayed playlist

[1279] Streaming songs

[1280] Specific behavior:

[1281] The playlist received by the device is displayed on the app, the user performs operations to play the music, and internet streaming is performed when playback begins.

[1282] Step 6: Gather feedback and continue learning

[1283] input:

[1284] User playback behavior data (e.g., skipping songs, completing playback, rating, etc.)

[1285] Real-time emotion data

[1286] process:

[1287] The terminal collects the user's playing behavior and emotional responses and records them as feedback.

[1288] The recorded feedback information is sent to the server.

[1289] The server analyzes the feedback and uses it to update the user's preferences.

[1290] output:

[1291] Updated user preference data

[1292] Improving the system's learning model

[1293] Specific behavior:

[1294] The device collects real-time playback behavior and emotional responses through an emotion engine and sends the information to a server, which analyzes it as learning material to improve the system's personalization accuracy.

[1295] Through this series of processes, an optimal music playlist that reflects the user's real-time emotions is generated and provided to the user, significantly improving the user's music experience.

[1296] (Application example 2)

[1297] Next, a description will be given of Application Example 2. In the following description, the data processing device 12 will be referred to as a "server" and the headset type terminal 314 will be referred to as a "terminal."

[1298] Conventional music playlist generation systems generate playlists based only on a user's playback history, usage time, and activity patterns, which means they lack the ability to personalize playlists in response to real-time emotional changes. Furthermore, it is difficult to provide a music experience that takes into account emotions in a virtual store. This can lead to situations where the user is feeling stressed, such as not playing the right music, which can detract from the shopping experience.

[1299] The specific processing by the specific processing unit 290 of the data processing device 12 in the application example 2 is realized by the following means.

[1300] In this invention, the server includes means for collecting a user's playback history, usage time period, activity pattern, and emotional state, means for analyzing the user's facial expressions and voice to recognize emotions in real time, and means for analyzing the collected data and real-time emotion recognition data to extract the user's music preferences and behavioral patterns. This makes it possible to generate a playlist based on real-time emotion recognition, and to provide optimal music that takes into account the user's emotions, particularly in a shopping experience in a virtual store.

[1301] The "playback history" is a history of songs and audio content that have been played by the user in the past.

[1302] The "usage time period" refers to the time or time period during which a user uses a music playlist.

[1303] An "activity pattern" refers to a user's daily activity or behavior pattern, such as a situation such as "commuting" or "exercising."

[1304] An "emotional state" is the emotion or mood a user is experiencing at a particular time.

[1305] An "emotion engine" is a technology that analyzes a user's facial expressions and voice to recognize emotions in real time.

[1306] A "user terminal" is a device used by a user, such as a smartphone, smart glasses, or head-mounted display.

[1307] "Feedback" refers to the user's playback behavior (skip songs, complete playback, rating, etc.) and their feelings about it.

[1308] A "playlist" is a list of songs that a user has selected to play consecutively.

[1309] A "virtual store" is a shopping environment provided in a virtual space where users can browse and purchase items from the store via digital devices.

[1310] "Real-time emotion recognition data" is data obtained by analyzing the user's current emotions in real time.

[1311] This invention is a system for personalizing the music experience in a virtual store using real-time emotion recognition data. The aim is to improve the user experience by generating and providing an optimal music playlist according to the user's emotional state while shopping in the virtual store through smart glasses.

[1312] 1. System configuration and programs

[1313] The system includes the following components:

[1314] Smart glasses equipped with an emotion recognition engine

[1315] Server that collects and analyzes data

[1316] A server that generates and provides playlists

[1317] The emotion recognition engine is integrated into the smart glasses and uses OpenCV and deep learning models to analyze the user's facial expressions and voice. Server technologies using frameworks such as Flask and Django are used for data collection and analysis.

[1318] 2. Processing Details

[1319] The emotion recognition engine captures the user's facial expressions through the camera and analyzes the data in real time. The analyzed emotional data, along with the user's playback history, usage time, and activity patterns, are sent to the server. The server then uses machine learning algorithms based on this data to extract the user's musical preferences and behavioral patterns.

[1320] The server then generates a playlist optimized for the specific user based on the emotion recognition data and the extracted music preference data. This playlist consists of songs that suit the user's current emotions. The generated playlist is then sent to the smart glasses via the Internet.

[1321] 3. Specific examples

[1322] For example, if a user feels stressed while shopping in a virtual store, the smart glasses' emotion recognition engine will recognize that emotion in real time. The emotion recognition data will be sent to the server, which will then generate a playlist containing music that promotes relaxation. This playlist will then be sent to the smart glasses, reducing the user's stress and providing a comfortable shopping experience.

[1323] An example of a prompt to input to a generative AI model would be:

[1324] Analyze in real time whether the user is currently feeling stressed and generate an appropriate relaxing music playlist based on that emotion. Example: If the user is feeling stressed, create a list of songs that are particularly effective for relaxing and play them through smart glasses.

[1325] In this way, the system realized by the present invention provides a music playlist that corresponds to the user's emotional state, thereby improving the shopping experience in a virtual store.

[1326] The flow of the specific processing in the application example 2 will be described with reference to FIG.

[1327] Step 1:

[1328] Smart glasses capture the user's facial expressions and voice.

[1329] Input: User's face image and voice.

[1330] Output: Captured face image data and audio data.

[1331] Specific operation: Uses the camera and microphone of the smart glasses to collect the user's face and voice in real time.

[1332] Step 2:

[1333] The emotion recognition engine analyzes the captured facial images and voice to recognize the user's emotions.

[1334] Input: Captured facial image data and audio data.

[1335] Output: Real-time emotion recognition data (e.g. "stressed", "relaxed", etc.).

[1336] How it works: It uses OpenCV and deep learning models to perform facial expression analysis and speech analysis, such as analyzing facial muscle movements and tone of voice, to create appropriate emotion tags.

[1337] Step 3:

[1338] The smart glasses send emotion recognition data, playback history, usage time, and activity patterns to a server.

[1339] Input: Real-time emotion recognition data, playback history, time of day usage, activity patterns.

[1340] Output: The aggregate data sent to the server.

[1341] Specific operation: Collected data is sent to the server using a secure protocol (e.g., HTTPS).

[1342] Step 4:

[1343] The server analyzes the received data and extracts the user's musical preferences and behavioral patterns.

[1344] Input: Playback history, time of day, activity patterns, real-time emotion recognition data.

[1345] Output: Music preference data and behavioral pattern data as analysis results.

[1346] Specific operation: Using machine learning algorithms (e.g., clustering algorithms and recommendation systems), analyze users' musical preferences and behavioral patterns.

[1347] Step 5:

[1348] The server generates a playlist based on the analysis results and real-time emotion recognition data.

[1349] Input: Music preference data, behavioral pattern data, real-time emotion recognition data.

[1350] Output: The generated playlist.

[1351] Specific operation: The analysis results are combined with emotional data to generate a list of songs that are optimal for the user, selecting songs that correspond to a specific emotional state (e.g., stress).

[1352] Step 6:

[1353] The server provides the generated playlist to the smart glasses.

[1354] Input: The generated playlist.

[1355] Output: Playlist data sent to the smart glasses.

[1356] Specific operation: The playlist data is packaged and sent to the smart glasses using a secure communication protocol.

[1357] Step 7:

[1358] The smart glasses play the received playlist.

[1359] Input: Playlist data sent to smart glasses.

[1360] Output: The song that is playing.

[1361] Specific operation: The received playlist is displayed in the user interface, and when the user selects it, the song is streamed over the Internet.

[1362] Step 8:

[1363] Collect user feedback and send it to the server.

[1364] Input: User playback behavior data (e.g., song skip, playback completion, song rating, etc.).

[1365] Output: Feedback data sent to the server.

[1366] Specific operation: Records various playback actions of users in real time and periodically sends the information to the server.

[1367] Step 9:

[1368] The server analyzes the feedback and trains the system.

[1369] Input: Feedback data.

[1370] Output: Updated music preference data and behavioral pattern data.

[1371] Specific operation: The received feedback data is analyzed using machine learning algorithms and updated to improve the system's personalization accuracy.

[1372] Example prompts to be input to the generative AI model:

[1373] Analyze in real time whether the user is currently feeling stressed and generate an appropriate relaxing music playlist based on that emotion. Example: If the user is feeling stressed, create a list of songs that are particularly effective for relaxing and play them through smart glasses.

[1374] The specific processing unit 290 transmits the result of the specific processing to the headset type terminal 314. In the headset type terminal 314, the control unit 46A causes the speaker 240 and the display 343 to output the result of the specific processing. The microphone 238 acquires audio indicating a user input regarding the result of the specific processing. The control unit 46A transmits audio data indicating the user input acquired by the microphone 238 to the data processing device 12. In the data processing device 12, the specific processing unit 290 acquires the audio data.

[1375] The data generation model 58 is a so-called generative AI (Artificial Intelligence). An example of the data generation model 58 is ChatGPT (Internet Search<URL: https: / / openai.com / blog / chatgpt> ), Gemini (Internet search <url: https: gemini.google.com ?hl="ja">) and other generation AIs. The data generation model 58 is obtained by performing deep learning on a neural network. A prompt including an instruction is input to the data generation model 58, and inference data such as voice data indicating voice, text data indicating text, and image data indicating an image is also input. The data generation model 58 performs inference on the input inference data in accordance with the instruction indicated by the prompt, and outputs the inference result in a data format such as voice data and text data. Here, inference refers to, for example, analysis, classification, prediction, and / or summarization.

[1376] In the above embodiment, an example was given in which the specific processing is performed by the data processing device 12, but the technology of the present disclosure is not limited to this, and the specific processing may be performed by the headset type terminal 314.

[1377] [Fourth embodiment]

[1378] FIG. 7 shows an example of the configuration of a data processing system 410 according to the fourth embodiment.

[1379] 7, a data processing system 410 includes a data processing device 12 and a robot 414. An example of the data processing device 12 is a server.

[1380] The data processing device 12 includes a computer 22, a database 24, and a communication I / F 26. The computer 22 is an example of a "computer" according to the technology of the present disclosure. The computer 22 includes a processor 28, a RAM 30, and a storage 32. The processor 28, the RAM 30, and the storage 32 are connected to a bus 34. The database 24 and the communication I / F 26 are also connected to the bus 34. The communication I / F 26 is connected to a network 54. Examples of the network 54 include a WAN (Wide Area Network) and / or a LAN (Local Area Network).

[1381] The robot 414 includes a computer 36, a microphone 238, a speaker 240, a camera 42, a communication I / F 44, and a control target 443. The computer 36 includes a processor 46, a RAM 48, and a storage 50. The processor 46, the RAM 48, and the storage 50 are connected to a bus 52. The microphone 238, the speaker 240, the camera 42, and the control target 443 are also connected to the bus 52.

[1382] The microphone 238 receives instructions and the like from the user 20 by receiving voice uttered by the user 20. The microphone 238 captures the voice uttered by the user 20, converts the captured voice into audio data, and outputs it to the processor 46. The speaker 240 outputs audio in accordance with instructions from the processor 46.

[1383] Camera 42 is a small digital camera equipped with an optical system including a lens, aperture, and shutter, and an imaging element such as a CMOS (Complementary Metal-Oxide-Semiconductor) image sensor or a CCD (Charge Coupled Device) image sensor, and captures images of the surroundings of user 20 (for example, an imaging range defined by an angle of view equivalent to the field of vision of a typical healthy person).

[1384] The communication I / F 44 is connected to a network 54. The communication I / Fs 44 and 26 are responsible for the exchange of various information between the processor 46 and the processor 28 via the network 54. The exchange of various information between the processor 46 and the processor 28 using the communication I / Fs 44 and 26 is carried out in a secure state.

[1385] The control object 443 includes a display device, LEDs in the eyes, and motors for driving the arms, hands, and feet. The posture and gestures of the robot 414 are controlled by controlling the motors of the arms, hands, and feet. Some of the emotions of the robot 414 can be expressed by controlling these motors. In addition, the facial expressions of the robot 414 can also be expressed by controlling the light emission state of the LEDs in the eyes of the robot 414.

[1386] Fig. 8 shows an example of the main functions of the data processing device 12 and the robot 414. As shown in Fig. 8, in the data processing device 12, a specific process is performed by the processor 28. A specific process program 56 is stored in the storage 32.

[1387] The specific processing program 56 is an example of a "program" according to the technology of the present disclosure. The processor 28 reads the specific processing program 56 from the storage 32 and executes the read specific processing program 56 on the RAM 30. The specific processing is realized by the processor 28 operating as a specific processing unit 290 in accordance with the specific processing program 56 executed on the RAM 30.

[1388] The storage 32 stores a data generation model 58 and an emotion identification model 59. The data generation model 58 and the emotion identification model 59 are used by the identification processing unit 290.

[1389] In the robot 414, the processor 46 performs the reception output process. A reception output program 60 is stored in the storage 50. The processor 46 reads the reception output program 60 from the storage 50 and executes the read reception output program 60 on the RAM 48. The reception output process is realized by the processor 46 operating as the control unit 46A in accordance with the reception output program 60 executed on the RAM 48.

[1390] Next, a description will be given of the specific processing performed by the specific processing unit 290 of the data processing device 12. In the following description, the data processing device 12 will be referred to as a "server" and the robot 414 will be referred to as a "terminal."

[1391] This invention relates to a system that personalizes and optimizes a user's music experience. Specifically, it is a method for generating and providing playlists based on the user's preferences and behavioral patterns by collecting and analyzing the user's playback history, usage time period, activity patterns, and emotional state. Furthermore, by collecting user feedback and allowing the system to continuously learn, it provides a more accurate and personalized music experience.

[1392] User Data Collection

[1393] A user starts a music application and plays music in the course of daily life.

[1394] The device collects the user's music history, the start and end times of the music, whether the music was skipped, the user's current activity status (e.g., commuting, exercising), and emotion tags, which may be entered by the user in the form of "fun" or "want to relax."

[1395] The terminal converts this data into a specified format and transmits it to the server using a secure communication method.

[1396] Data analysis and learning

[1397] The server stores the received user data in a database.

[1398] The server analyzes the stored data and extracts features to identify users' musical preferences and behavioral patterns. This analysis uses machine learning algorithms. For example, a clustering algorithm can group users' activity patterns.

[1399] The server performs statistical analysis based on a large amount of user data and also learns general trends.

[1400] Playlist Generation

[1401] The server then generates the optimal playlist for a specific user based on the analysis results, which means selecting songs of different genres and moods according to the user's time of day, activity patterns, and emotional tags.

[1402] The server packages the list of songs included in the generated playlist and its associated data (song ID, playback URL, etc.) and sends it to the terminal.

[1403] Providing playlists

[1404] The terminal displays the received playlist on the application's user interface for easy access by the user.

[1405] The user selects a song from the playlist and plays it. When a playback command is given, the terminal streams the song over the Internet.

[1406] Gather feedback and continue learning

[1407] The device collects the songs the user has played, whether they have completed or skipped, and their ratings and feedback (e.g., "favorite marks") for specific songs.

[1408] The terminal transmits this feedback information to the server.

[1409] The server updates the user's preferences based on the received feedback and then reruns the analysis and learning process, which is then reflected in the next playlist generation.

[1410] As a concrete example, if a user sets the emotion tag "busy" during their morning commute, the server will select energetic, fast-paced music for that time period based on past data and learning results. The generated playlist is provided to the device and played by the user. On the other hand, when it's time to relax in the evening, the server will generate a playlist containing calming music and provide it to the user via the device. In this way, it is possible to provide a music experience that is optimal for the user's situation and emotion.

[1411] The processing flow will be explained below.

[1412] Step 1:

[1413] A user launches a music application and starts playing a song. At the same time, the user sets an activity state (e.g., "commuting" or "playing sports") and an emotion tag (e.g., "having fun" or "wanting to relax").

[1414] Step 2:

[1415] The device automatically collects information about the song the user is playing (playback start time, play end time, song ID, etc.), as well as the user's activity status and emotion tag. The collected data is temporarily stored on the device.

[1416] Step 3:

[1417] The device formats the collected user data into a specific format, encrypts it, and then sends it to a server using a secure protocol such as HTTPS.

[1418] Step 4:

[1419] The server stores the received user data in a database, and then starts a process to analyze the stored data. This analysis uses machine learning algorithms.

[1420] Step 5:

[1421] The server extracts features from the analyzed data to identify users' musical preferences and behavioral patterns, for example by using a clustering algorithm to group users based on their activity patterns and emotions.

[1422] Step 6:

[1423] The server selects music suitable for a specific time period or activity level and generates an optimal playlist for each user, for example, energetic music for the commute and quieter music for relaxation.

[1424] Step 7:

[1425] The server packages the generated playlist information (song ID, playback URL, etc.) and sends it to the device. This communication also uses a secure protocol.

[1426] Step 8:

[1427] The device displays the received playlist in the user interface of the music application, allowing the user to select and play it.

[1428] Step 9:

[1429] The user selects a song from the presented playlist and starts playing it. The device streams the song over the Internet.

[1430] Step 10:

[1431] The device collects the user's playback behavior (skip songs, complete playback, rating, etc.) and records it as feedback.

[1432] Step 11:

[1433] The terminal transmits the collected feedback information to the server.

[1434] Step 12:

[1435] The server analyzes the received feedback and uses it as data to update the user's preferences, which allows for more accurate personalization when generating the next playlist.

[1436] Example 1

[1437] Next, a description will be given of Example 1. In the following description, the data processing device 12 will be referred to as a "server" and the robot 414 will be referred to as a "terminal."

[1438] Conventional music recommendation systems have difficulty generating playlists that fully consider a user's specific usage situation and emotions. Furthermore, they lack a system that can properly learn from user feedback and reflect it in future recommendations. As a result, users are unable to obtain the music experience that best suits their preferences and situation.

[1439] The specific processing by the specific processing unit 290 of the data processing device 12 in the first embodiment is realized by the following means.

[1440] In this invention, the server includes: means for collecting a user's playback history, usage time periods, activity patterns, and emotional state; means for formatting the collected data into a specified format and transmitting it to the server using a secure communication method; means for saving the received user data, analyzing it using a machine learning algorithm, and extracting the user's music preferences and behavioral patterns; means for generating an optimal playlist for a specific user based on the analysis results; means for packaging a list of songs included in the generated playlist and their associated data and transmitting it to a user terminal; means for displaying the received playlist on a user interface for easy user access; and means for collecting user feedback and re-executing the analysis and learning process based on the feedback, thereby enabling a more personalized music experience based on the user's usage history and emotional tags.

[1441] "Playback history" is information about songs that the user has played in the past, and includes song IDs, titles, artist names, and the like.

[1442] The "usage time period" is information about the date and time when the user used the music application, including the playback start time and end time.

[1443] The "activity pattern" refers to the activity status of the user while playing music, and indicates a specific operating state such as commuting, exercising, or relaxing.

[1444] The "emotional state" refers to the emotional or mood state of the user when listening to music, and is expressed in the form of emotion tags such as "fun" or "want to relax."

[1445] A "secure communication method" is a communication protocol that ensures safety when sending and receiving data, and a specific example is HTTPS.

[1446] "Machine learning algorithms" are algorithms that analyze large amounts of data and extract patterns and rules, and include clustering algorithms.

[1447] A "playlist" is a collection of multiple songs, and is a format for providing a collection of songs that meets the user's playback needs.

[1448] A "user interface" is a display screen that allows a user to directly interact with a system or application, and includes a song list, operation buttons, and the like.

[1449] "Feedback" refers to information provided by users to the system, such as playback history, song ratings and comments, which is reflected in the system's learning process.

[1450] The "analysis and learning process" is a series of steps that the system takes to learn the user's preferences and behavioral patterns based on data collected from the user, and then generate a new playlist.

[1451] This invention relates to a system that personalizes and optimizes a user's music experience. This system collects and analyzes a user's playback history, usage time, activity patterns, and emotional state, and generates and provides playlists based on the user's preferences and behavioral patterns. Furthermore, by collecting user feedback and continuously learning, the system provides an even more accurate and personalized music experience.

[1452] First, a user launches a music application and plays a song in their daily life. For example, a user presses the play button on the application to play a song. Here, the device collects the history of songs played by the user, the start and end times of playback, whether songs were skipped, the user's current activity state (commuting, exercising, etc.), and emotion tags (e.g., happy, want to relax). These emotion tags are input by the user.

[1453] The device then formats the collected data into a specified format and sends it to the server using HTTPS, a secure communication method. At this time, the data is converted into JSON format for efficient data processing.

[1454] The server stores the received user data in a database. For example, the server inserts data into the database using SQL queries. The server then analyzes the stored data using machine learning algorithms to identify users' musical preferences and behavioral patterns. This analysis uses the Python library scikit-learn, which runs a clustering algorithm to extract specific patterns.

[1455] Next, the server generates a playlist optimized for the specific user based on the analysis results. This means selecting songs of different genres and moods depending on the user's time of day, activity patterns, and emotional tags. The generated playlist is converted into JSON format and sent back to the device using the HTTPS protocol.

[1456] The device displays the received playlist on the application's user interface for easy access by the user. For example, the user can click a song in the playlist and press the play button to start streaming the song. Specifically, the device connects to the streaming server and streams the music using the playback URL.

[1457] Additionally, when a user provides feedback such as which songs they have completed playing, which songs they have skipped, or their ratings for specific songs, the device collects this feedback information and sends it to the server. The server then updates the user's preferences based on the received feedback and reruns the analysis and learning process, allowing the information to be reflected in the next playlist generation.

[1458] As a concrete example, if a user sets the emotion tag "busy" during their morning commute, the server will select energetic, fast-paced music for that time period based on past data and learning results. The generated playlist is provided to the device and played by the user. On the other hand, when it's time to relax in the evening, the server will generate a playlist containing calming music and provide it to the user via the device. In this way, it is possible to provide a music experience that is optimal for the user's situation and emotion.

[1459] Examples of prompt statements

[1460] "Please explain in detail the data analysis process that goes into generating an energizing playlist for your morning commute."

[1461] "How do I update my personalized music playlist based on user feedback?"

[1462] "Please explain in detail how you would generate an evening playlist based on the emotion tag 'I want to relax'."

[1463] The flow of the identification process in the first embodiment will be described with reference to FIG.

[1464] Step 1: Collect user data

[1465] In everyday life, users launch music applications and play music.

[1466] Specifically, the user presses the play button on the application on a device such as a smartphone or computer.

[1467] The device collects the user's history of songs played, the start and end times of playback, whether songs were skipped, the user's current activity status (commuting, exercising, etc.), and emotion tags (e.g., fun, want to relax).

[1468] The input is various data when the user operates the playlist.

[1469] As part of the data processing, the terminal records these data in its internal memory.

[1470] The output is formatted data, ready to be sent to the next step.

[1471] Step 2: Sending data

[1472] The terminal formats the collected data into the specified format.

[1473] The input is raw, unformatted data.

[1474] To process the data, the terminal converts the data into JSON format.

[1475] The output is data in JSON format.

[1476] The terminal transmits data to the server using a secure communication method (e.g., HTTPS).

[1477] Specifically, the terminal sends JSON format data to the server as a POST request using the HTTPS protocol.

[1478] Step 3: Data storage and analysis

[1479] The server stores the received user data in a database.

[1480] The input is JSON format data sent from the terminal.

[1481] For data processing, the server receives this data and inserts it into a database using an SQL query.

[1482] The output is user data stored in a database.

[1483] The server analyzes the stored data using machine learning algorithms (e.g., clustering algorithms).

[1484] The input is user data stored in a database.

[1485] For data calculation, the server applies machine learning models using Python libraries (e.g., scikit-learn) to identify users' musical preferences and behavioral patterns.

[1486] The output is the user's characteristics and pattern information as the analysis results.

[1487] Step 4: Generate a playlist

[1488] The server generates an optimal playlist for a particular user based on the analysis results.

[1489] The input is the user's characteristics and pattern information as the analysis result.

[1490] To process the data, the server narrows down the conditions based on the time of use, activity patterns, and emotion tags, and selects appropriate songs from the music database.

[1491] The output is the generated playlist.

[1492] The server packages the list of songs included in the generated playlist and its associated data (song ID, playback URL, etc.) and sends it to the terminal.

[1493] Specifically, the server converts the playlist into JSON format and sends a POST request to the terminal using the HTTPS protocol.

[1494] Step 5: Serve the playlist

[1495] The terminal displays the received playlist on the user interface of the application, making it easily accessible to the user.

[1496] The input is the playlist data in JSON format sent from the server.

[1497] For data processing, the device parses the JSON data and dynamically displays a song list in the UI component.

[1498] The output is the displayed playlist.

[1499] The user selects a song from the presented playlist and plays it.

[1500] Specifically, the user clicks on a song in the playlist and presses the play button.

[1501] The device streams the music over the Internet.

[1502] The input is the URL to play the song.

[1503] For data processing, the terminal connects to a streaming server and plays music.

[1504] Step 6: Gather feedback and continue learning

[1505] The device collects songs that the user has completed playing, songs that the user has skipped, and ratings and feedback for specific songs.

[1506] The input is the feedback information provided by the user.

[1507] As for data processing, the terminal tracks the user's operations and stores them in data storage.

[1508] The output is the collected feedback data.

[1509] The terminal transmits this feedback information to the server.

[1510] Specifically, the terminal formats the feedback data into JSON format and sends a POST request to the server using the HTTPS protocol.

[1511] The server updates the user's preferences based on the received feedback and then reruns the analysis and learning process.

[1512] The input is the feedback data sent from the terminal.

[1513] As the data is processed, the server stores it in a database and retrains the machine learning model.

[1514] The output is updated analysis results and data that will be reflected in the next playlist generation.

[1515] (Application example 1)

[1516] Next, a description will be given of Application Example 1. In the following description, the data processing device 12 will be referred to as a "server" and the robot 414 will be referred to as a "terminal."

[1517] Conventionally, methods for optimizing the efficiency and accuracy of factory robots have relied solely on motion algorithms and machine learning. However, it is believed that even more efficient work can be achieved by providing an appropriate environment based on the robot's activity status and work content. The present invention aims to improve the efficiency and accuracy of work by playing personalized music while the robot is working.

[1518] The specific processing by the specific processing unit 290 of the data processing device 12 in the application example 1 is realized by the following means.

[1519] In this invention, the server includes means for collecting a user's playback history, usage time periods, activity patterns, and emotional state, means for analyzing the collected data and extracting the user's preferences and behavioral patterns, means for generating a playlist based on the analysis results, means for providing the generated playlist to the terminal, means for collecting the terminal's work history and activity level and generating an optimal music experience based on the data, means for collecting feedback and conducting system training, and means for generating applications for use in various facilities. This makes it possible to play optimal music according to the robot's activity status and work content, effectively improving work efficiency and accuracy.

[1520] "Playback history" is a record of music and content that a user has played in the past.

[1521] A "time period of use" refers to a specific time range during which a user uses music or content.

[1522] An "activity pattern" is a series of patterns of actions or states of a user or a terminal.

[1523] "Emotional state" refers to the user's current emotional or mood state.

[1524] "Analysis" is the process of examining collected data in detail to identify trends and characteristics.

[1525] A "playlist" refers to a list of songs or content to be played.

[1526] A "terminal" refers to an electronic device used by a user, and includes, for example, a smartphone or a robot.

[1527] "Work history" is a record of work performed by a terminal or robot.

[1528] "Activity level" is a measure that indicates the activity and intensity of the terminal or robot's movements.

[1529] "Feedback" refers to responses or opinions from users or systems.

[1530] A "clustering algorithm" is a machine learning technique for classifying data into several groups.

[1531] "System learning" is the process by which a system improves itself based on collected data and feedback.

[1532] An "application" is software or a program that provides a specific function.

[1533] This invention is a system that automatically plays appropriate music to improve the work efficiency and accuracy of robots working in factories. The system collects the user's playback history, usage time, activity patterns, and emotional state, and generates and provides a playlist based on the analysis results. It also has a function that collects the device's work history and activity level to generate optimal music.

[1534] Generating a Program

[1535] The system program includes the following main functions:

[1536] 1. Data Collection:

[1537] Sensors are used to collect records of tasks performed by the device, work time, and activity levels.

[1538] The user's emotional state and usage time period are also collected (for example, music playback history while working).

[1539] 2. Data Analysis:

[1540] The collected data is stored and analyzed on the server using Python.

[1541] The robot's activity patterns are classified using scikit-learn's clustering algorithm (KMeans).

[1542] Features are extracted from the data and music appropriate for the robot's task is selected.

[1543] 3. Playlist generation:

[1544] Generate the best playlist for a specific task based on the analysis results.

[1545] The playlist includes energetic and fast-paced songs, as well as calm and rhythmic songs.

[1546] 4. Playlist provided by:

[1547] The generated playlist is sent to the terminal and a playback instruction is issued.

[1548] 5. Feedback and learning:

[1549] The device collects the history of songs played and feedback from users and sends it to the server.

[1550] The system will self-learn based on the feedback and reflect this in the next playlist generation.

[1551] Processing Details

[1552] Hardware:

[1553] The robot is equipped with a head-mounted display (HMD) and built-in speakers.

[1554] Sensors for data collection (e.g., activity trackers and temperature sensors).

[1555] software:

[1556] Data analysis and learning program using Python.

[1557] Clustering algorithm using the scikit-learn library.

[1558] Data storage using a database (e.g., MongoDB).

[1559] Specific examples

[1560] For example, if a robot performing assembly work is engaged in the task "assembly," it will play lively, fast-paced music to improve the efficiency of that work. In this case, the server selects appropriate music based on the data collected so far, creates a playlist, and sends it to the device. It also receives feedback and selects the best music for the next task.

[1561] Prompt Sentence Examples

[1562] You can instruct the generative AI model to collect data and generate a playlist using prompts like the following:

[1563] Create a program for the data collection part. Collect the task ID, work time, and activity level while the robot is working, and save them in a data frame format.

[1564]

[1565] Please create a program for generating the playlist. Create an optimal playlist based on the robot's task ID and return it in list format.

[1566]

[1567] Write a program that provides the generated playlist. Design a system that plays all the songs in the playlist in order.

[1568] In this way, the present invention can optimize the efficiency and accuracy of the robot's work and provide an effective musical experience.

[1569] The flow of the specific processing in the application example 1 will be described with reference to FIG.

[1570] Step 1:

[1571] The device collects the user's or robot's playback history, usage time period, activity patterns, and emotional state from sensors and stores them in a data frame format. The input is sensor data, and the output is a formatted data frame. The device formats the collected data into a specified format and transmits it to the server using a secure communication method.

[1572] Step 2:

[1573] The server stores the data received from the terminal in a database. The input is the data sent from the terminal, and the output is the data stored in the database. The server performs a write operation to the database and checks whether the data is saved correctly.

[1574] Step 3:

[1575] The server analyzes the stored data and classifies the activity patterns of the user or robot using a clustering algorithm (e.g., KMeans in scikit-learn). The input is the data in the database, and the output is the classification results for each cluster. Data preprocessing and feature extraction are also performed during the analysis process.

[1576] Step 4:

[1577] The server generates a playlist optimized for a specific task based on the analysis results. The input is the clustering results and user preference data, and the output is the generated playlist. Specifically, it selects appropriate songs based on the data belonging to a specific cluster.

[1578] Step 5:

[1579] The server provides the generated playlist to the terminal. The input is the generated playlist, and the output is the playlist data sent to the terminal. The server packages and sends the list of songs included in the playlist and their associated data (song IDs, playback URLs, etc.).

[1580] Step 6:

[1581] The terminal displays the received playlist on the user interface, making it playable by the user or robot. The input is the playlist data received from the server, and the output is a playable playlist on the user interface. The terminal starts streaming the music over the Internet.

[1582] Step 7:

[1583] The device collects data during playback (e.g., playback completion, skip, rating) and feedback from the user. The input is user operations during playback and feedback data, and the output is the collected feedback data. The device securely transmits the feedback data to the server.

[1584] Step 8:

[1585] The server analyzes the received feedback data and continues to train the system. The input is the collected feedback data, and the output is an updated learning model for the system. The server reflects the feedback and uses it to generate the next playlist.

[1586] Through these processing steps, a system that optimizes the user's work efficiency can be effectively realized.

[1587] Furthermore, an emotion engine that estimates the user's emotion may be further combined. That is, the identification processing unit 290 may estimate the user's emotion using the emotion identification model 59, and perform identification processing using the user's emotion.

[1588] This invention is a system for personalizing the music experience, and by combining it with an emotion engine that recognizes the user's emotions, it provides a more accurate music playlist. Specifically, it collects the user's playback history, usage time period, activity patterns, and emotional state, and then uses the emotion engine to recognize and analyze the user's real-time emotions, and generates and provides a playlist based on that data.

[1589] User Data Collection

[1590] A user launches a music application and starts playing a song. At the same time, the user sets an activity state (e.g., "commuting" or "playing sports") and an emotion tag (e.g., "having fun" or "wanting to relax").

[1591] Additionally, the device uses a built-in emotion engine to recognize real-time emotions by analyzing the user's voice, facial expressions, gestures, etc. This emotion recognition data may be stored in the form of tags.

[1592] The data collected by the device is temporarily stored and then transmitted to a server using a secure communication protocol.

[1593] Data analysis and learning

[1594] The server stores the received data in a database, which includes the user's playback history, usage time period, activity patterns, emotion tags, and real-time emotion recognition data.

[1595] The server analyzes the stored data using machine learning algorithms to extract features that can help identify users' musical preferences and behavioral patterns.

[1596] The server also performs statistical analysis of multiple user data and learns general trends, enabling more accurate personalization.

[1597] Playlist Generation

[1598] The server then generates a playlist optimized for a specific user based on the analyzed data, including songs that reflect the user's time of day, activity patterns, emotion tags, and real-time emotions recognized by the emotion engine.

[1599] The server packages the information of the generated playlist (music ID, playback URL, etc.) and sends it to the terminal.

[1600] Providing playlists

[1601] The terminal displays the received playlists in the user interface of the music application for easy access by the user.

[1602] The user selects a song from the playlist and plays it. When a playback command is given, the terminal streams the song over the Internet.

[1603] Gather feedback and continue learning

[1604] The device collects the user's playback behavior (skip songs, complete playback, rating, etc.) and their emotions regarding those actions (including real-time emotion recognition) and records them as feedback.

[1605] The terminal transmits this feedback information to the server.

[1606] The server analyzes the received feedback and uses it to update the user's preferences, which then reflects the feedback when generating the next playlist, resulting in more accurate personalization.

[1607] As a concrete example, if a user feels "busy" while commuting on the train in the morning, the emotion engine will recognize this emotion in real time from the user's facial expressions and tone of voice. Based on past data and learning results, the server will select energetic, fast-paced music for that time period and generate a playlist. This playlist will be provided to the user via their device. Similarly, if the user wants to relax in the evening, the server will generate a playlist containing calming music, which will also be provided to the user via their device. This system enables a music experience optimized for the user's situation and emotions.

[1608] The processing flow will be explained below.

[1609] Step 1:

[1610] The user launches a music application and starts playing a song, while manually inputting an activity state (e.g., "commuting" or "playing sports") and an emotion tag (e.g., "having fun" or "wanting to relax").

[1611] Step 2:

[1612] The device automatically collects information about the song being played by the user (playback start time, play end time, song ID, etc.), and uses a built-in emotion engine to analyze the user's voice, facial expressions, and gestures, thereby recognizing the user's emotions in real time.

[1613] Step 3:

[1614] The device formats the collected data (playback history, usage time, activity patterns, manually entered emotion tags, and real-time emotion recognition data) into a specific format, encrypts it, and then sends it to a server using a secure communication protocol (e.g., HTTPS).

[1615] Step 4:

[1616] The server stores the received data in a database, which is then analyzed using machine learning algorithms to extract features that identify the user's musical preferences and behavioral patterns.

[1617] Step 5:

[1618] The server uses the listening history and the real-time emotions recognized by the emotion engine to generate an optimal playlist that suits the user's mood and activity at that moment.

[1619] Step 6:

[1620] The server then packages the generated playlist information (song ID, playback URL, etc.) and sends it to the device. This communication is also carried out using a secure protocol.

[1621] Step 7:

[1622] The terminal displays the received playlist on the user interface of the music application, allowing the user to access it directly.

[1623] Step 8:

[1624] The user selects a song from the playlist and plays it. When the device receives the instruction to play, it starts streaming the song over the Internet.

[1625] Step 9:

[1626] The device collects the user's playback behavior (skip songs, complete playback, rating specific songs, etc.) and real-time emotional data, and then transmits this feedback information back to the server.

[1627] Step 10:

[1628] The server analyzes the received feedback and updates the user's preferences. The updated information is reflected in the next playlist generation, and the system uses this information for further learning.

[1629] Example 2

[1630] Next, a description will be given of Example 2. In the following description, the data processing device 12 will be referred to as a "server" and the robot 414 will be referred to as a "terminal."

[1631] Current music playlist generation systems generate playlists based on a user's playback history, time of day, and basic activity patterns. However, it is difficult to reflect the user's momentary emotional state or real-time preferences. This makes it difficult for users to consistently enjoy the optimal music experience. Furthermore, there are limitations to methods for effectively utilizing feedback to continuously improve the system. To address these issues, a system that can recognize and analyze a user's real-time emotions and provide highly accurate personalized playlists is needed.

[1632] The specific processing by the specific processing unit 290 of the data processing device 12 in the second embodiment is realized by the following means.

[1633] In this invention, the server includes means for collecting a user's playback history, usage time periods, activity patterns, and emotional states; means for analyzing the collected data and extracting the user's music preferences and behavioral patterns; means for generating a playlist based on the analysis results; means for providing the generated playlist to the user terminal; means for analyzing the user's voice, facial expressions, and gestures to recognize and record real-time emotions; and means for collecting feedback from users and training the system based on the feedback.

[1634] This allows for the generation of an optimal playlist that reflects the user's real-time emotions, making it possible to continue providing a music experience that matches the user's preferences.

[1635] "User" means an individual who utilizes the System to play music and obtain a personalized music experience.

[1636] "Playback history" refers to a record of songs that a user has played in the past.

[1637] A "time period of use" refers to a segment of time during which a user uses a music application.

[1638] An "activity pattern" refers to the behavior or situation in which a user is playing music (e.g., commuting, playing sports).

[1639] "Emotional state" refers to the emotion or mood of the user when playing music.

[1640] "Means of collection" refers to the methods and devices used to acquire and record the necessary data (playback history, time of use, activity patterns, emotional state, etc.).

[1641] "Means for analysis" refers to methods or devices that use machine learning algorithms to analyze collected data and extract meaningful information.

[1642] "Means for generating a playlist" refers to a method or device that creates a list of songs based on the analysis results and provides it to the user.

[1643] "Means for providing" refers to a method or device for delivering the generated playlist to a user's terminal and making it accessible.

[1644] An "emotion engine" is software or algorithms that analyze voice, facial expressions, and gestures to recognize and record a user's real-time emotions.

[1645] "Feedback collection means" refers to a method or device that captures and records a user's playback behavior and emotional responses.

[1646] "Means for system training" refers to methods and devices that analyze collected feedback and improve the personalization accuracy of the entire system.

[1647] The present invention is a system for personalizing a user's music experience, and in particular, provides a music playlist that reflects the user's real-time emotions by combining an emotion engine. This system is realized through the following elements and processes.

[1648] 1. Collection of User Data

[1649] The user launches the music application and selects an activity status such as "commuting" or "playing sports" or an emotion tag such as "fun" or "want to relax" within the application when starting use.

[1650] The device collects this information and uses an emotion engine (e.g., OpenFace or Microsoft Azure's Emotion API) that analyzes voice, facial expressions, and gestures to obtain real-time emotional data about the user.

[1651] 2. Data transmission and storage

[1652] The device temporarily stores the user's playback history, usage time period, activity patterns, emotion tags, and real-time emotion recognition data, and transmits them to the server using a secure communication protocol (e.g., HTTPS).

[1653] 3. Data analysis and learning

[1654] The server stores the transmitted data in a structured database, which includes each user's playback history, usage time periods, activity patterns, emotion tags, and real-time emotion recognition data.

[1655] The server analyzes the stored data using machine learning algorithms (e.g., TensorFlow or PyTorch) to extract features that identify the user's musical preferences and behavioral patterns.

[1656] The server also performs statistical analysis of multiple user data and learns general trends to improve the accuracy of personalization across the system.

[1657] 4. Generate a playlist

[1658] The server then generates a playlist optimized for a specific user based on the analysis results, for example, selecting energetic music for a user commuting in the morning and calming music for a user wanting to relax in the evening.

[1659] The server packages the generated playlist (music IDs, playback URLs, etc.) and sends it to the terminal.

[1660] 5. Providing playlists

[1661] The device displays the received playlists in the music application's user interface for easy access by the user, with the playlists visually categorized.

[1662] The user selects a song from the playlist and plays it. When a playback command is given, the terminal streams the song over the Internet.

[1663] 6. Gather feedback and continue learning

[1664] The device collects and records the user's playback behavior (e.g., skipping songs, completing playback, rating, etc.) and emotional responses.

[1665] The terminal transmits this feedback information to the server.

[1666] The server analyzes the received feedback and uses it to update the user's preferences, so that the feedback is reflected the next time a playlist is generated, resulting in more accurate personalization.

[1667] Examples of concrete examples and prompts

[1668] For example, if a user feels "busy" while commuting on the train in the morning, the emotion engine recognizes this emotion in real time from the user's facial expressions and tone of voice. Based on past data and learning results, the server selects energetic, fast-paced music for that time period and generates a playlist. This playlist is provided to the user via their device. On the other hand, when the user wants to relax in the evening, the server generates a playlist containing calming music, which is also provided to the user via their device.

[1669] An example of a prompt to input to a generative AI model is as follows:

[1670] "I'm on the train to work this morning and I'm feeling busy. Make me a playlist of energetic music."

[1671] "Generate the perfect playlist for a relaxing evening."

[1672] This will provide the user with an optimal music experience that suits their real-time emotions and situation.

[1673] The flow of the identification process in the second embodiment will be described with reference to FIG.

[1674] Step 1: Collect user data

[1675] input:

[1676] User play history

[1677] User usage time period

[1678] User activity patterns (e.g., commuting, playing sports)

[1679] User's emotional state (e.g., "I'm having fun" or "I want to relax")

[1680] process:

[1681] The user starts the music application and selects the playback history and various tags when starting use.

[1682] The device obtains real-time emotional data using an emotion engine (e.g., OpenFace or Microsoft Azure's Emotion API) that analyzes the user's voice, facial expressions, and gestures.

[1683] output:

[1684] User playback history data

[1685] Current time zone data

[1686] Activity pattern data

[1687] Real-time emotion data

[1688] Specific behavior:

[1689] The user operates a music application and inputs settings that make them feel "busy" during their commute, and the device uses its camera and microphone to provide data to the emotion engine, generating real-time emotion data.

[1690] Step 2: Send and store data

[1691] input:

[1692] User playback history data, usage time zone data, activity pattern data, and real-time emotion data collected in Step 1

[1693] process:

[1694] The terminal temporarily stores the collected data.

[1695] The saved data is sent to the server using a secure communication protocol (e.g. HTTPS).

[1696] output:

[1697] Data sent to the server

[1698] Specific behavior:

[1699] The device saves the data in local storage and then transmits it to the server using HTTPS.

[1700] Step 3: Analyze and train the data

[1701] input:

[1702] Data sent in step 2

[1703] process:

[1704] The server stores the submitted data in a structured database.

[1705] Machine learning algorithms (e.g., TensorFlow and PyTorch) are used to analyze users' musical preferences and behavioral patterns and extract specific features.

[1706] The server performs statistical analysis of multiple user data to learn general trends.

[1707] output:

[1708] User music taste characteristic data

[1709] General Trend Data

[1710] Specific behavior:

[1711] The server accumulates data in a database, analyzes it using a machine learning model, extracts features based on the user's preferences and behavioral patterns, and builds a model based on these.

[1712] Step 4: Generate a playlist

[1713] input:

[1714] The characteristic data of the user's music preferences and general tendency data extracted in step 3

[1715] process:

[1716] Based on the analysis results, the server generates a playlist that is optimal for the specific user.

[1717] The generated playlist information (song ID, playback URL, etc.) is packaged and sent to the terminal.

[1718] output:

[1719] Generated playlist data

[1720] Specific behavior:

[1721] The server combines each user's characteristic data with general tendency data to generate an optimal playlist, which is then packaged and sent to the terminal.

[1722] Step 5: Serve the playlist

[1723] input:

[1724] The playlist data generated in step 4

[1725] process:

[1726] The terminal displays the received playlist on the user interface of the music application.

[1727] The user selects and plays songs from the playlist.

[1728] The device streams music over the internet.

[1729] output:

[1730] Displayed playlist

[1731] Streaming songs

[1732] Specific behavior:

[1733] The playlist received by the device is displayed on the app, the user performs operations to play the music, and internet streaming is performed when playback begins.

[1734] Step 6: Gather feedback and continue learning

[1735] input:

[1736] User playback behavior data (e.g., skipping songs, completing playback, rating, etc.)

[1737] Real-time emotion data

[1738] process:

[1739] The terminal collects the user's playing behavior and emotional responses and records them as feedback.

[1740] The recorded feedback information is sent to the server.

[1741] The server analyzes the feedback and uses it to update the user's preferences.

[1742] output:

[1743] Updated user preference data

[1744] Improving the system's learning model

[1745] Specific behavior:

[1746] The device collects real-time playback behavior and emotional responses through an emotion engine and sends the information to a server, which analyzes it as learning material to improve the system's personalization accuracy.

[1747] Through this series of processes, an optimal music playlist that reflects the user's real-time emotions is generated and provided to the user, significantly improving the user's music experience.

[1748] (Application example 2)

[1749] Next, a description will be given of Application Example 2. In the following description, the data processing device 12 will be referred to as a "server" and the robot 414 will be referred to as a "terminal."

[1750] Conventional music playlist generation systems generate playlists based only on a user's playback history, usage time, and activity patterns, which means they lack the ability to personalize playlists in response to real-time emotional changes. Furthermore, it is difficult to provide a music experience that takes into account emotions in a virtual store. This can lead to situations where the user is feeling stressed, such as not playing the right music, which can detract from the shopping experience.

[1751] The specific processing by the specific processing unit 290 of the data processing device 12 in the application example 2 is realized by the following means.

[1752] In this invention, the server includes means for collecting a user's playback history, usage time period, activity pattern, and emotional state, means for analyzing the user's facial expressions and voice to recognize emotions in real time, and means for analyzing the collected data and real-time emotion recognition data to extract the user's music preferences and behavioral patterns. This makes it possible to generate a playlist based on real-time emotion recognition, and to provide optimal music that takes into account the user's emotions, particularly in a shopping experience in a virtual store.

[1753] The "playback history" is a history of songs and audio content that have been played by the user in the past.

[1754] The "usage time period" refers to the time or time period during which a user uses a music playlist.

[1755] An "activity pattern" refers to a user's daily activity or behavior pattern, such as a situation such as "commuting" or "exercising."

[1756] An "emotional state" is the emotion or mood a user is experiencing at a particular time.

[1757] An "emotion engine" is a technology that analyzes a user's facial expressions and voice to recognize emotions in real time.

[1758] A "user terminal" is a device used by a user, such as a smartphone, smart glasses, or head-mounted display.

[1759] "Feedback" refers to the user's playback behavior (skip songs, complete playback, rating, etc.) and their feelings about it.

[1760] A "playlist" is a list of songs that a user has selected to play consecutively.

[1761] A "virtual store" is a shopping environment provided in a virtual space where users can browse and purchase items from the store via digital devices.

[1762] "Real-time emotion recognition data" is data obtained by analyzing the user's current emotions in real time.

[1763] This invention is a system for personalizing the music experience in a virtual store using real-time emotion recognition data. The aim is to improve the user experience by generating and providing an optimal music playlist according to the user's emotional state while shopping in the virtual store through smart glasses.

[1764] 1. System configuration and programs

[1765] The system includes the following components:

[1766] Smart glasses equipped with an emotion recognition engine

[1767] Server that collects and analyzes data

[1768] A server that generates and provides playlists

[1769] The emotion recognition engine is integrated into the smart glasses and uses OpenCV and deep learning models to analyze the user's facial expressions and voice. Server technologies using frameworks such as Flask and Django are used for data collection and analysis.

[1770] 2. Processing Details

[1771] The emotion recognition engine captures the user's facial expressions through the camera and analyzes the data in real time. The analyzed emotional data, along with the user's playback history, usage time, and activity patterns, are sent to the server. The server then uses machine learning algorithms based on this data to extract the user's musical preferences and behavioral patterns.

[1772] The server then generates a playlist optimized for the specific user based on the emotion recognition data and the extracted music preference data. This playlist consists of songs that suit the user's current emotions. The generated playlist is then sent to the smart glasses via the Internet.

[1773] 3. Specific examples

[1774] For example, if a user feels stressed while shopping in a virtual store, the smart glasses' emotion recognition engine will recognize that emotion in real time. The emotion recognition data will be sent to the server, which will then generate a playlist containing music that promotes relaxation. This playlist will then be sent to the smart glasses, reducing the user's stress and providing a comfortable shopping experience.

[1775] An example of a prompt to input to a generative AI model would be:

[1776] Analyze in real time whether the user is currently feeling stressed and generate an appropriate relaxing music playlist based on that emotion. Example: If the user is feeling stressed, create a list of songs that are particularly effective for relaxing and play them through smart glasses.

[1777] In this way, the system realized by the present invention provides a music playlist that corresponds to the user's emotional state, thereby improving the shopping experience in a virtual store.

[1778] The flow of the specific processing in the application example 2 will be described with reference to FIG.

[1779] Step 1:

[1780] Smart glasses capture the user's facial expressions and voice.

[1781] Input: User's face image and voice.

[1782] Output: Captured face image data and audio data.

[1783] Specific operation: Uses the camera and microphone of the smart glasses to collect the user's face and voice in real time.

[1784] Step 2:

[1785] The emotion recognition engine analyzes the captured facial images and voice to recognize the user's emotions.

[1786] Input: Captured facial image data and audio data.

[1787] Output: Real-time emotion recognition data (e.g. "stressed", "relaxed", etc.).

[1788] How it works: It uses OpenCV and deep learning models to perform facial expression analysis and speech analysis, such as analyzing facial muscle movements and tone of voice, to create appropriate emotion tags.

[1789] Step 3:

[1790] The smart glasses send emotion recognition data, playback history, usage time, and activity patterns to a server.

[1791] Input: Real-time emotion recognition data, playback history, time of day usage, activity patterns.

[1792] Output: The aggregate data sent to the server.

[1793] Specific operation: Collected data is sent to the server using a secure protocol (e.g., HTTPS).

[1794] Step 4:

[1795] The server analyzes the received data and extracts the user's musical preferences and behavioral patterns.

[1796] Input: Playback history, time of day, activity patterns, real-time emotion recognition data.

[1797] Output: Music preference data and behavioral pattern data as analysis results.

[1798] Specific operation: Using machine learning algorithms (e.g., clustering algorithms and recommendation systems), analyze users' musical preferences and behavioral patterns.

[1799] Step 5:

[1800] The server generates a playlist based on the analysis results and real-time emotion recognition data.

[1801] Input: Music preference data, behavioral pattern data, real-time emotion recognition data.

[1802] Output: The generated playlist.

[1803] Specific operation: The analysis results are combined with emotional data to generate a list of songs that are optimal for the user, selecting songs that correspond to a specific emotional state (e.g., stress).

[1804] Step 6:

[1805] The server provides the generated playlist to the smart glasses.

[1806] Input: The generated playlist.

[1807] Output: Playlist data sent to the smart glasses.

[1808] Specific operation: The playlist data is packaged and sent to the smart glasses using a secure communication protocol.

[1809] Step 7:

[1810] The smart glasses play the received playlist.

[1811] Input: Playlist data sent to smart glasses.

[1812] Output: The song that is playing.

[1813] Specific operation: The received playlist is displayed in the user interface, and when the user selects it, the song is streamed over the Internet.

[1814] Step 8:

[1815] Collect user feedback and send it to the server.

[1816] Input: User playback behavior data (e.g., song skip, playback completion, song rating, etc.).

[1817] Output: Feedback data sent to the server.

[1818] Specific operation: Records various playback actions of users in real time and periodically sends the information to the server.

[1819] Step 9:

[1820] The server analyzes the feedback and trains the system.

[1821] Input: Feedback data.

[1822] Output: Updated music preference data and behavioral pattern data.

[1823] Specific operation: The received feedback data is analyzed using machine learning algorithms and updated to improve the system's personalization accuracy.

[1824] Example prompts to be input to the generative AI model:

[1825] Analyze in real time whether the user is currently feeling stressed and generate an appropriate relaxing music playlist based on that emotion. Example: If the user is feeling stressed, create a list of songs that are particularly effective for relaxing and play them through smart glasses.

[1826] The specific processing unit 290 transmits the result of the specific processing to the robot 414. In the robot 414, the control unit 46A causes the speaker 240 and the control target 443 to output the result of the specific processing. The microphone 238 acquires voice indicating a user input regarding the result of the specific processing. The control unit 46A transmits voice data indicating the user input acquired by the microphone 238 to the data processing device 12. In the data processing device 12, the specific processing unit 290 acquires the voice data.

[1827] The data generation model 58 is a so-called generative AI (Artificial Intelligence). An example of the data generation model 58 is ChatGPT (Internet Search<URL: https: / / openai.com / blog / chatgpt> ), Gemini (Internet search <url: https: gemini.google.com ?hl="ja">) and other generation AIs. The data generation model 58 is obtained by performing deep learning on a neural network. A prompt including an instruction is input to the data generation model 58, and inference data such as voice data indicating voice, text data indicating text, and image data indicating an image is also input. The data generation model 58 performs inference on the input inference data in accordance with the instruction indicated by the prompt, and outputs the inference result in a data format such as voice data and text data. Here, inference refers to, for example, analysis, classification, prediction, and / or summarization.

[1828] In the above embodiment, an example was given in which the specific processing is performed by the data processing device 12, but the technology of the present disclosure is not limited to this, and the specific processing may be performed by the robot 414.

[1829] The emotion identification model 59 as an emotion engine may determine the user's emotion according to a specific mapping. Specifically, the emotion identification model 59 may determine the user's emotion according to an emotion map (see FIG. 9), which is a specific mapping. Similarly, the emotion identification model 59 may determine the robot's emotion, and the identification processing unit 290 may perform identification processing using the robot's emotion.

[1830] FIG. 9 is a diagram illustrating an emotion map 400 on which multiple emotions are mapped. In the emotion map 400, emotions are arranged in concentric circles radiating from the center. Emotions closer to the center of the concentric circles are more primitive. Emotions representing states and actions arising from a state of mind are arranged on the outer edges of the concentric circles. The concept of emotion includes both affect and mental states. Emotions generally generated from reactions occurring in the brain are arranged on the left side of the concentric circles. Emotions generally induced by situational judgment are arranged on the right side of the concentric circles. Emotions generally generated from reactions occurring in the brain and induced by situational judgment are arranged on the upper and lower sides of the concentric circles. Furthermore, the emotion of "pleasure" is arranged on the upper side of the concentric circles, and the emotion of "discomfort" is arranged on the lower side. In this way, in the emotion map 400, multiple emotions are mapped based on the structure by which emotions are generated, and emotions that tend to occur simultaneously are mapped close to each other.

[1831] These emotions are distributed in the 3 o'clock direction on emotion map 400, and typically fluctuate between relief and anxiety. In the right half of emotion map 400, situational awareness dominates over internal sensations, resulting in a sense of calm.

[1832] The inside of emotion map 400 represents what is going on in the mind, and the outside of emotion map 400 represents behavior, so the further you go outside emotion map 400, the more visible the emotions become (the more they are expressed in behavior).

[1833] Human emotions are based on various balances, such as posture and blood sugar levels. When these balances deviate from the ideal, a state of discomfort is indicated, and when they approach the ideal, a state of pleasure is indicated. Emotions can also be created for robots, automobiles, and motorcycles, based on various balances, such as posture and remaining battery life. When these balances deviate from the ideal, a state of discomfort is indicated, and when they approach the ideal, a state of pleasure is indicated. An emotion map can be generated, for example, based on Dr. Mitsuyoshi's emotion map (Research on Voice Emotion Recognition and Emotional Brain Physiological Signal Analysis Systems, Tokushima University, Doctoral Dissertation: https: / / ci.nii.ac.jp / naid / 500000375379). The left half of the emotion map lists emotions belonging to the "reaction" domain, where sensation is dominant. The right half of the emotion map lists emotions belonging to the "situation" domain, where situational awareness is dominant.

[1834] The emotion map defines two emotions that promote learning. One is a negative emotion on the situation side, around the middle of "repentance" or "reflection." In other words, this occurs when the robot experiences negative emotions such as "I never want to feel this way again" or "I don't want to be scolded again." The other is a positive emotion on the response side, around "desire." In other words, this occurs when the robot experiences positive feelings such as "I want more" or "I want to know more."

[1835] The emotion identification model 59 inputs user input into a pre-trained neural network, obtains emotion values ​​indicating each emotion shown in the emotion map 400, and determines the user's emotion. This neural network is pre-trained based on multiple pieces of training data that are combinations of user input and emotion values ​​indicating each emotion shown in the emotion map 400. Furthermore, this neural network is trained so that emotions that are located close to each other have similar values, as in the emotion map 900 shown in FIG. 10. FIG. 10 shows an example in which multiple emotions, "relieved," "calm," and "reassuring," have similar emotion values.

[1836] The system according to the present disclosure has been described above mainly with respect to the functions of the data processing device 12, but the system according to the present disclosure is not necessarily implemented on a server. The system according to the present disclosure may be implemented as a general information processing system. The present disclosure may be implemented, for example, as a software program running on a personal computer or an application running on a smartphone, etc. The method according to the present disclosure may be provided to users in the form of SaaS (Software as a Service).

[1837] In the above embodiment, an example was given in which the specific processing is performed by one computer 22, but the technology of the present disclosure is not limited to this, and the specific processing may be distributed and performed by a plurality of computers including the computer 22. For example, the data generation model 58 may be provided in an external device of the data processing device 12, and data may be generated in the external device in accordance with input data.

[1838] In the above embodiment, an example in which the specific processing program 56 is stored in the storage 32 has been described, but the technology of the present disclosure is not limited to this. For example, the specific processing program 56 may be stored in a portable, computer-readable, non-transitory storage medium such as a USB (Universal Serial Bus) memory. The specific processing program 56 stored in the non-transitory storage medium is installed in the computer 22 of the data processing device 12. The processor 28 executes the specific processing in accordance with the specific processing program 56.

[1839] Alternatively, the specific processing program 56 may be stored in a storage device such as a server connected to the data processing device 12 via the network 54, and the specific processing program 56 may be downloaded and installed on the computer 22 in response to a request from the data processing device 12.

[1840] It is not necessary to store all of the specific processing program 56 in a storage device such as a server connected to the data processing device 12 via the network 54, or to store all of the specific processing program 56 in the storage 32; only a portion of the specific processing program 56 may be stored.

[1841] The hardware resource for executing a specific process can be any of the following processors: An example of a processor is a CPU, which is a general-purpose processor that functions as a hardware resource for executing a specific process by executing software, i.e., a program. Another example of a processor is a dedicated electrical circuit, such as an FPGA (Field-Programmable Gate Array), a PLD (Programmable Logic Device), or an ASIC (Application Specific Integrated Circuit), which is a processor with a circuit configuration designed specifically for executing a specific process. Each processor has built-in or connected memory, and each processor uses the memory to execute the specific process.

[1842] The hardware resource that executes the specific processing may be configured with one of these various processors, or may be configured with a combination of two or more processors of the same or different types (for example, a combination of multiple FPGAs, or a combination of a CPU and an FPGA). Also, the hardware resource that executes the specific processing may be a single processor.

[1843] As an example of a system configured with a single processor, first, one processor is configured by combining one or more CPUs and software, and this processor functions as a hardware resource that executes a specific process. Second, there is a system that uses a processor that realizes the functions of an entire system including multiple hardware resources that execute a specific process on a single IC chip, as typified by SoC (System-on-a-chip). In this way, a specific process is realized using one or more of the above-mentioned various processors as hardware resources.

[1844] Furthermore, the hardware structure of these various processors can be, more specifically, an electric circuit that combines circuit elements such as semiconductor devices. The specific processing described above is merely an example. Therefore, it goes without saying that unnecessary steps may be deleted, new steps may be added, or the processing order may be rearranged, without departing from the spirit of the invention.

[1845] The above-described description and illustrations are a detailed explanation of the parts related to the technology of the present disclosure and are merely an example of the technology of the present disclosure. For example, the above description of the configuration, functions, actions, and effects is an explanation of an example of the configuration, functions, actions, and effects of the parts related to the technology of the present disclosure. Therefore, it goes without saying that unnecessary parts may be deleted, new elements may be added, or replacements may be made to the above-described description and illustrations within the scope of the gist of the technology of the present disclosure. Furthermore, to avoid confusion and facilitate understanding of the parts related to the technology of the present disclosure, the above-described description and illustrations omit explanations of common technical knowledge that do not require particular explanation to enable the implementation of the technology of the present disclosure.

[1846] All publications, patent applications, and technical standards mentioned in this specification are herein incorporated by reference to the same extent as if each individual publication, patent application, or technical standard was specifically and individually indicated to be incorporated by reference.

[1847] The following is further disclosed regarding the above embodiment.

[1848] (Claim 1)

[1849] means for collecting a user's playback history, usage time period, activity pattern, and emotional state;

[1850] A means for analyzing the collected data and extracting the user's musical preferences and behavioral patterns;

[1851] means for generating a playlist based on the analysis results;

[1852] means for providing the generated playlist to a user terminal;

[1853] a means of collecting user feedback and training the system based on the feedback;

[1854] A system that includes...

[1855] (Claim 2)

[1856] 2. The system according to claim 1, wherein the emotional state of the user is collected in the form of tags and reflected in playlist generation.

[1857] (Claim 3)

[1858] 2. The system according to claim 1, wherein the user's activity patterns are analyzed using a clustering algorithm and utilized in generating the playlist.

[1859] "Example 1"

[1860] (Claim 1)

[1861] means for collecting a user's playback history, usage time period, activity pattern, and emotional state;

[1862] A means for converting the collected data into a specified format and transmitting the data to a server using a secure communication method;

[1863] means for storing the received user data and analyzing it using machine learning algorithms to extract user musical preferences and behavioral patterns;

[1864] A means for generating an optimal playlist for a specific user based on the analysis results;

[1865] a means for packaging a list of songs included in the generated playlist and related data and transmitting the packaged data to a user terminal;

[1866] means for displaying the received playlist in a user interface for easy user access;

[1867] A means of collecting user feedback and rerunning the analysis and learning process based on that feedback; and

[1868] A system including:

[1869] (Claim 2)

[1870] 2. The system according to claim 1, wherein the emotional state of the user is collected in the form of tags and reflected in playlist generation.

[1871] (Claim 3)

[1872] 2. The system according to claim 1, wherein the user's activity patterns are analyzed using a clustering algorithm and utilized in generating the playlist.

[1873] "Application Example 1"

[1874] (Claim 1)

[1875] means for collecting a user's playback history, usage time period, activity pattern, and emotional state;

[1876] A means for analyzing the collected data and extracting user preferences and behavioral patterns;

[1877] means for generating a playlist based on the analysis results;

[1878] a means for providing the generated playlist to a terminal;

[1879] means for collecting device activity history and activity levels and generating an optimal music experience based on that data;

[1880] a means of gathering feedback and training the system;

[1881] means for generating applications for use within various facilities;

[1882] A system including:

[1883] (Claim 2)

[1884] 2. The system according to claim 1, wherein the emotional state of the user is collected in the form of tags and reflected in the generated playlist.

[1885] (Claim 3)

[1886] 2. The system of claim 1, wherein the activity patterns of the terminal are analyzed using a clustering algorithm and used to generate the playlist.

[1887] "Example 2: Combining Emotion Engines"

[1888] (Claim 1)

[1889] means for collecting a user's playback history, usage time period, activity pattern, and emotional state;

[1890] A means for analyzing the collected data and extracting the user's musical preferences and behavioral patterns;

[1891] means for generating a playlist based on the analysis results;

[1892] means for providing the generated playlist to a user terminal;

[1893] A means of recognizing and recording real-time emotions by analyzing the user's voice, facial expressions, and gestures;

[1894] a means of collecting user feedback and training the system based on the feedback;

[1895] A system including:

[1896] (Claim 2)

[1897] 2. The system according to claim 1, wherein the emotional state of the user is collected in the form of tags and reflected in playlist generation.

[1898] (Claim 3)

[1899] 2. The system according to claim 1, wherein the user's activity patterns are analyzed using a clustering algorithm and utilized in generating the playlist.

[1900] "Application example 2 when combining emotion engines"

[1901] (Claim 1)

[1902] means for collecting a user's playback history, usage time period, activity pattern, and emotional state;

[1903] A means of recognizing emotions in real time by analyzing the user's facial expressions and voice,

[1904] A means for analyzing the collected data and real-time emotion recognition data to extract the user's musical preferences and behavioral patterns;

[1905] means for generating a playlist based on the analysis results and real-time emotion recognition data;

[1906] means for providing the generated playlist to a user terminal;

[1907] means for generating a music playlist based on the user's emotional state and providing an emotionally sensitive music experience within the virtual store;

[1908] a means of collecting user feedback and training the system based on the feedback;

[1909] A system including:

[1910] (Claim 2)

[1911] 10. The system of claim 1, wherein real-time emotion recognition data during a user's activity is used to generate a playlist.

[1912] (Claim 3)

[1913] 2. The system according to claim 1, wherein the user terminal is equipped with an emotion engine that analyzes the user's facial expressions and voice. [Explanation of symbols]

[1914] 10, 210, 310, 410 Data Processing Systems 12 Data Processing Device 14 Smart Devices 214 Smart Glasses 314 Headset-type terminal 414 Robot< / url:> < / url:> < / url:> < / url:>

Claims

1. means for collecting a user's playback history, usage time period, activity pattern, and emotional state; A means for analyzing the collected data and extracting the user's musical preferences and behavioral patterns; means for generating a playlist based on the analysis results; means for providing the generated playlist to a user terminal; a means of collecting user feedback and training the system based on the feedback; A system including:

2. 2. The system according to claim 1, wherein the emotional state of the user is collected in the form of tags and reflected in playlist generation.

3. 2. The system according to claim 1, wherein the user's activity patterns are analyzed using a clustering algorithm and utilized in generating the playlist.

Citation Information

Patent Citations

  • Persona chatbot control method and system

    JP2022180282A