System

The system addresses the challenge of providing personalized music information by using AI and emotional engines to analyze user preferences and emotions, resulting in a more engaging and satisfying music experience for users.

JP2025071051APending Publication Date: 2025-05-02SOFTBANK GROUP CORP
View PDF 1 Cites 0 Cited by

Patent Information

Application Number
JP2024182561
Authority / Receiving Office
JP · JP
Patent Type
Applications
Current Assignee / Owner
Priority Date
2023-10-20
Filing Date
2024-10-18
Publication Date
2025-05-02

AI Technical Summary

Technical Problem

Existing music information systems struggle to provide personalized and accurate music information to users, particularly seniors, who want to deepen their understanding and enjoyment of their favorite music.

Method used

A system that combines music information acquisition, user preference and question analysis, and personalized information provision using AI models and emotional engines, allowing users to access music information through displays and audio outputs.

Benefits of technology

Enables users to enrich their music experience by providing individually customized music information based on their preferences and past memories, improving user satisfaction and engagement.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure 2025071051000001_ABST
    Figure 2025071051000001_ABST
Patent Text Reader

Abstract

To provide a system.SOLUTION: A system includes: means for acquiring music information; means for analyzing emotion of a user by using an emotion recognition engine; means for providing information on music and artists related to a question or emotion of the user; and means for adjusting timing at which music is selected and information is provided on the basis of a result of analyzing the emotion of the user.SELECTED DRAWING: Figure 1
Need to check novelty before this filing date? Find Prior Art

Description

[Technical field]

[0001] The technology of the present disclosure relates to a system. [Background technology]

[0002] Patent document 1 discloses a persona chatbot control method performed by at least one processor, the method including the steps of receiving a user utterance, adding the user utterance to a prompt including a description and related instruction sentence regarding the chatbot character, encoding the prompt, and inputting the encoded prompt into a language model to generate a chatbot utterance in response to the user utterance. [Prior art documents] [Patent documents]

[0003] [Patent Document 1] JP 2022-180282 A Summary of the Invention [Problem to be solved by the invention]

[0004] Enriching seniors' musical experiences and helping them gain a deeper understanding of the music they love. [Means for solving the problem]

[0005] By providing a system that combines a means for acquiring music information, a means for analyzing a user's preferences and questions, and a means for providing information on music and artists related to the user's preferences and questions, we provide information that is individually customized based on the senior's musical preferences and past memories. By providing a deeper understanding of the music that the user likes, we can enrich the user's musical experience and rekindle their knowledge and passion for music. In addition, by providing information through access to the Internet and music information databases and through displays and audio output, users can learn about music easily and enjoyably.

[0006] "Means for acquiring music information" refers to means for accessing the Internet or a music information database. "Means for analyzing user's preferences and questions" refers to means for analyzing information provided by the user and selecting appropriate music and artist information based on the analysis results. "Means for providing information on music and artists related to the user's preferences and questions" refers to means for providing information on music and artists related to the user's preferences and questions through a display or audio output. [Brief description of the drawings]

[0007] [Figure 1] 1 is a conceptual diagram showing an example of the configuration of a data processing system according to a first embodiment. [Diagram 2] 1 is a conceptual diagram showing an example of main functions of a data processing device and a smart device according to a first embodiment. FIG. [Diagram 3] FIG. 11 is a conceptual diagram showing an example of the configuration of a data processing system according to a second embodiment. [Figure 4] FIG. 11 is a conceptual diagram showing an example of main functions of a data processing device and smart glasses according to a second embodiment. [Diagram 5] FIG. 13 is a conceptual diagram showing an example of the configuration of a data processing system according to a third embodiment. [Figure 6] FIG. 11 is a conceptual diagram showing an example of main functions of a data processing device and a headset-type terminal according to a third embodiment. [Figure 7]FIG. 13 is a conceptual diagram showing an example of the configuration of a data processing system according to a fourth embodiment. [Figure 8] FIG. 13 is a conceptual diagram showing an example of main functions of a data processing device and a robot according to a fourth embodiment. [Figure 9] 1 shows an emotion map onto which multiple emotions are mapped. [Figure 10] 1 shows an emotion map onto which multiple emotions are mapped. [Figure 11] 4 is a sequence diagram showing a process flow of the data processing system according to the first embodiment. FIG. [Figure 12] 11 is a sequence diagram showing a process flow of the data processing system in application example 1. FIG. [Figure 13] FIG. 11 is a sequence diagram showing the flow of processing of the data processing system according to the second embodiment when an emotion engine is combined. [Figure 14] FIG. 11 is a sequence diagram showing the flow of processing in the data processing system in application example 2 when combined with an emotion engine. DETAILED DESCRIPTION OF THE PREFERRED EMBODIMENTS

[0008] Hereinafter, an example of an embodiment of a system according to the technology of the present disclosure will be described with reference to the accompanying drawings.

[0009] First, the terms used in the following description will be explained.

[0010] In the following embodiments, a signed processor (hereinafter simply referred to as a "processor") may be one arithmetic device or a combination of multiple arithmetic devices. Furthermore, a processor may be one type of arithmetic device or a combination of multiple types of arithmetic devices. Examples of arithmetic devices include a CPU (Central Processing Unit), a GPU (Graphics Processing Unit), a GPGPU (General-Purpose computing on Graphics Processing Units), an APU (Accelerated Processing Unit), etc.

[0011] In the following embodiments, a signed RAM (Random Access Memory) is a memory in which information is temporarily stored and is used as a working memory by a processor.

[0012] In the following embodiments, the coded storage is one or more non-volatile storage devices that store various programs, various parameters, etc. Examples of non-volatile storage devices include flash memory (SSD (Solid State Drive)), magnetic disks (e.g., hard disks), and magnetic tapes.

[0013] In the following embodiments, a communication I / F (Interface) with a code is an interface including a communication processor and an antenna. The communication I / F controls communication between multiple computers. An example of a communication standard applied to the communication I / F is a wireless communication standard including 5G (5th Generation Mobile Communication System), Wi-Fi (registered trademark), Bluetooth (registered trademark), etc.

[0014] In the following embodiments, "A and / or B" is synonymous with "at least one of A and B." In other words, "A and / or B" means that it may be only A, only B, or a combination of A and B. In addition, in this specification, the same idea as "A and / or B" is also applied when three or more things are expressed by connecting them with "and / or."

[0015] [First embodiment] FIG. 1 shows an example of the configuration of a data processing system 10 according to the first embodiment.

[0016] 1, a data processing system 10 includes a data processing device 12 and a smart device 14. An example of the data processing device 12 is a server.

[0017] The data processing device 12 includes a computer 22, a database 24, and a communication I / F 26. The computer 22 is an example of a "computer" according to the technology of the present disclosure. The computer 22 includes a processor 28, a RAM 30, and a storage 32. The processor 28, the RAM 30, and the storage 32 are connected to a bus 34. The database 24 and the communication I / F 26 are also connected to the bus 34. The communication I / F 26 is connected to a network 54. Examples of the network 54 include a wide area network (WAN) and / or a local area network (LAN).

[0018] The smart device 14 includes a computer 36, a reception device 38, an output device 40, a camera 42, and a communication I / F 44. The computer 36 includes a processor 46, a RAM 48, and a storage 50. The processor 46, the RAM 48, and the storage 50 are connected to a bus 52. The reception device 38, the output device 40, and the camera 42 are also connected to the bus 52.

[0019] The reception device 38 includes a touch panel 38A and a microphone 38B, and receives user input. The touch panel 38A detects contact with a pointer (e.g., a pen or a finger) to receive user input by the touch of the pointer. The microphone 38B detects the user's voice to receive user input by voice. The control unit 46A transmits data indicating the user input received by the touch panel 38A and the microphone 38B to the data processing device 12. In the data processing device 12, the specific processing unit 290 acquires the data indicating the user input.

[0020] The output device 40 includes a display 40A and a speaker 40B, and presents data to the user 20 by outputting the data in a form of expression that the user 20 can perceive (e.g., voice and / or text). The display 40A displays visible information such as text and images according to instructions from the processor 46. The speaker 40B outputs voice according to instructions from the processor 46. The camera 42 is a compact digital camera equipped with an optical system including a lens, an aperture, and a shutter, and an imaging element such as a Complementary Metal-Oxide-Semiconductor (CMOS) image sensor or a Charge Coupled Device (CCD) image sensor.

[0021] The communication I / F 44 is connected to a network 54. The communication I / Fs 44 and 26 are responsible for transmitting and receiving various types of information between the processor 46 and the processor 28 via the network 54.

[0022] FIG. 2 shows an example of main functions of the data processing device 12 and the smart device 14.

[0023] As shown in Fig. 2, in the data processing device 12, a specific process is performed by the processor 28. A specific process program 56 is stored in the storage 32. The specific process program 56 is an example of a "program" according to the technology of the present disclosure. The processor 28 reads the specific process program 56 from the storage 32, and executes the read specific process program 56 on the RAM 30. The specific process is realized by the processor 28 operating as a specific processing unit 290 in accordance with the specific process program 56 executed on the RAM 30.

[0024] The storage 32 stores a data generation model 58 and an emotion identification model 59. The data generation model 58 and the emotion identification model 59 are used by the identification processing unit 290.

[0025] In the smart device 14, the processor 46 performs the reception output processing. The storage 50 stores a reception output program 60. The reception output program 60 is used together with the specific processing program 56 by the data processing system 10. The processor 46 reads out the reception output program 60 from the storage 50, and executes the read reception output program 60 on the RAM 48. The reception output processing is realized by the processor 46 operating as the control unit 46A in accordance with the reception output program 60 executed on the RAM 48.

[0026] Next, a description will be given of the identification process performed by the identification processing unit 290 of the data processing device 12. In the following description, the data processing device 12 will be referred to as a "server" and the smart device 14 will be referred to as a "terminal."

[0027] An embodiment for implementing the disclosed technology comprises the following elements. 1. How to get music information: The server obtains information about music history and artists by accessing the Internet and a music information database. 2. Means of analyzing user preferences and questions: - The server analyzes the information provided by the user and selects appropriate music and artists based on the user's preferences and questions. The server may also select appropriate music and artists based on the user's age, gender, and usual music listening habits. 3. Means of providing information about music and artists related to the user's preferences and questions: The server transmits the acquired information to the terminal, and the terminal provides information on music and artists related to the user's preferences and questions through a display and / or audio output. As a specific example, an embodiment will be described in which a user asks the question "Tell me about rock music from the 60's." 1. How to get music information: - The server accesses the Internet and music information databases to obtain information about rock music from the 1960s, such as representative bands and artists, popular songs, and the social context of the time. 2. Means of analyzing user preferences and questions: - The server parses the user's question, "What is rock music from the 60's?" and identifies information relevant to the question. 3. Means of providing information about music and artists related to the user's preferences and questions: The server transmits the acquired information to the terminal, and the terminal displays the characteristics, representative artists, popular songs, etc. of rock music in the 60s on the display. The terminal may also provide the characteristics, representative artists, popular songs, etc. of rock music in the 60s to the user through audio output.

[0028] The process flow will be explained below.

[0029] Step 1: Receive user input such as questions and preferences. - The user communicates questions and preferences to the system through the terminal. - The terminal sends the user's input to the server. Step 2: Search and retrieve information based on your questions and preferences. - The server receives and parses the user input. - The server accesses the Internet and music information databases to search for music and artist information related to the user's question and preferences. - The server organizes the obtained information and extracts the necessary information. Step 3: Provide information. - The server sends the acquired information to the terminal. - The device displays information on a display and conveys information through audio output. - Users can enjoy learning from the information provided.

[0030] Example 1

[0031] Next, a description will be given of Example 1. In the following description, the data processing device 12 is referred to as a "server" and the smart device 14 is referred to as a "terminal."

[0032] Conventional music information systems have difficulty in quickly and accurately responding to the diverse tastes and questions of users. In addition, few systems can provide a consistent process for efficiently collecting, analyzing, and presenting the information users want. As a result, the user experience is often poor and dissatisfied. In addition, when users want detailed information related to a specific era or genre, they have to search manually or refer to multiple sources, which is time-consuming and laborious.

[0033] The specific processing by the specific processing unit 290 of the data processing device 12 in the first embodiment is realized by the following means.

[0034] In this invention, the server includes a means for accessing the Internet or a database to acquire music information, a means for analyzing information and questions provided by the user to identify appropriate music and artists, and a means for transmitting information on the identified music and artists to the terminal and providing it to the user through a display or audio output, thereby making it possible to quickly and accurately collect, analyze, and provide the music information desired by the user.

[0035] The "Internet" is a collection of globally connected computer networks and an infrastructure for the exchange of information and communication.

[0036] A "database" is a systematically organized collection of information that can be efficiently accessed, searched, and managed through a specific application or query.

[0037] "User" refers to an individual or group that uses this system to obtain music information.

[0038] A "terminal" is a device from which a user receives information, and includes a personal computer, a smartphone, a tablet, etc.

[0039] "Display" refers to a screen or monitor for visually displaying information.

[0040] "Audio output" refers to technologies and devices that provide information to a user as audio.

[0041] "Music information" refers to all data and knowledge about music, such as music history, artist information, song details, and genre characteristics.

[0042] "Analysis" refers to the process of examining given information or data in detail to understand its structure and meaning.

[0043] "Acquisition" refers to the act of collecting information and making it available in a usable form.

[0044] "Providing" refers to the act of supplying information to a user and making it available for use.

[0045] A "network" refers to an infrastructure that connects multiple computers and devices and enables the exchange of information.

[0046] The present invention relates to a system for quickly and accurately providing music information desired by a user. Specific embodiments will be described below.

[0047] How to get music information The server uses the Internet and a database to collect information about music, utilizing existing web services such as Google® search API and music database API. For example, if a user asks the question "Tell me about rock music from the 60s," the server first uses the Google search API to gather the necessary information from the Internet. It also uses the music database API to obtain information about popular artists and songs from the 60s. This data is stored in the server's internal database and used for subsequent processing.

[0048] A means of analyzing user preferences and questions The server uses natural language processing techniques to analyze the questions and information entered by the user, specifically using high-performance generative AI models such as the ChatGPT model from OpenAI. If a user types in "Tell me about rock music from the 60s," the server will analyze the content and extract related keywords such as "60s," "rock music," etc. The results of this analysis will become the basis for the server to collect appropriate information and provide it to the user.

[0049] Here is an example of a prompt for parsing: Prompt: "A user has asked the question 'Rock music from the 60s.' Please analyze the information related to this question."

[0050] A way to provide users with music and artist information relevant to their preferences and questions In order to provide the collected and analyzed information to the user, the server first converts the data into an appropriate format, and then transmits the data to the user's terminal. The terminal displays the information received from the server on the display and outputs voice as necessary. Specifically, the terminal can use a Text-to-Speech (TTS) engine to play back text information as voice. For example, if a user asks "Tell me about rock music from the 60's," the server might provide text like this: "In 60's rock music, Singer A, Singer B, etc. are representative artists. Also, the XX Music Festival in 1969 is very famous." The device can both display this text information and play it aloud using a TTS engine.

[0051] Below are some examples of informational prompts: Prompt: "Provide the user with the information you have retrieved based on the user's question. Provide text on the display and speech output."

[0052] The flow of the identification process in the first embodiment will be described with reference to FIG.

[0053] Step 1: The server receives input from the user. Input: A user question (e.g., "What is rock music from the 60's?") Specific operation: The user enters a question into the terminal, which then sends this information to the server. Output: The server receives the user's question.

[0054] Step 2: The server uses a generative AI model to analyze the user's question. Input: A user question (e.g., "What is rock music from the 60's?") How it works: The server uses natural language processing techniques, such as OpenAI's ChatGPT model, to analyze the question, understand the intent of the question, and extract relevant keywords ("60s", "rock music"). Output: A list of keywords resulting from the analysis ("60s", "rock music").

[0055] Step 3: The server accesses the Internet and databases to retrieve relevant information. Input: Analyzed keyword list ("60s", "Rock music") How it works: The server uses the Google Search API to execute queries such as "60s rock music history" and collects the resulting webpage data. It also uses a music database API to retrieve information about popular artists and songs from the 60s. Output: Music information data (e.g. representative artists, popular songs, social background, etc.).

[0056] Step 4: The server organizes the retrieved information and converts it into a format for presentation to the user. Input: Retrieved music information data Specific operation: The server organizes the text data and organizes it into a format that is easy for the user to understand. For example, it may classify the acquired information into categories such as "famous artists," "popular songs," and "XX Music Festival in 1969." Output: Formatted information (e.g. "60s rock music is represented by artists Singer A, Singer B, etc. Also, the XX music festival in 1969 is very famous.").

[0057] Step 5: The server sends the formatted information to the user's terminal. Input: Formatted information Specific operation: The server sends information using the network address of the terminal. Output: The terminal receives the formatted information.

[0058] Step 6: The terminal displays the received information on a display and outputs audio as necessary. Input: Formatted information What it does: The device displays text information on a display and can also provide information aloud using a Text-to-Speech (TTS) engine. Output: The user receives information visually and audibly.

[0059] (Application example 1) Next, a description will be given of Application Example 1. In the following description, the data processing device 12 will be referred to as a "server" and the smart device 14 will be referred to as a "terminal."

[0060] Conventional music recommendation systems have limited functionality in providing appropriate music information based on users' questions and preferences. In particular, responding to specific user questions and diverse music genres is a challenge, and effective solutions to improve user experience are required. Furthermore, advanced functions for quickly and accurately analyzing user-entered questions and providing relevant music information are lacking.

[0061] The specific processing by the specific processing unit 290 of the data processing device 12 in the application example 1 is realized by the following means.

[0062] In this invention, the server includes a means for acquiring music information, a means for analyzing a user's preferences and questions, a means for providing music-related information related to the user's preferences and questions, a means for analyzing the user's questions using a generative AI model, and a means for generating a prompt sentence, which makes it possible to quickly and accurately analyze a specific question from a user and provide appropriate music information.

[0063] "Means for acquiring music information" refers to means for accessing the Internet or a music database to gather music-related information.

[0064] The "means for analyzing user preferences and questions" refers to a means for analyzing information provided by a user to understand the contents of the preferences and questions.

[0065] The "means for providing music-related information related to the user's preferences and questions" refers to a means for providing music-related information that the user is interested in based on collected information through a display device or audio output device.

[0066] "Means for analyzing a user's question using a generative AI model" refers to means for analyzing the content of a user's question using an AI model that employs natural language processing technology and identifying related information.

[0067] A "means for generating a prompt sentence" is a means for automatically creating text to generate an appropriate response to a user's question.

[0068] DETAILED DESCRIPTION OF THE PREFERRED EMBODIMENTS The present invention is configured to provide appropriate music-related information in response to specific user queries in a music recommendation system.

[0069] First, a user uses a smartphone application. When the user inputs a question, the question is sent to the application server. The server first accesses the Internet or a music database using a music information acquisition means to collect related music information.

[0070] The server then analyzes the user's question using a generative AI model, including the Transformer model, which is based on natural language processing techniques, that analyzes the text entered by the user and identifies relevant music genres and artists based on the content.

[0071] After obtaining the parsed result of the question, the server generates a prompt sentence, which is used to formulate an appropriate response to the user's question.

[0072] Finally, the server collects music information and provides the information to the user through a display device and audio output device using a means for providing music-related information related to the user's preferences and questions. As a specific example, information such as characteristics of music genres, representative artists, and popular songs can be displayed on the display and the same information can be reproduced by audio.

[0073] As a specific operational procedure of the embodiment, consider a case where a user asks the question, "Tell me about rock music from the 60's." In response to this question, the system operates as follows.

[0074] 1. The server accesses the Internet and a music database to retrieve information about 60's rock music. 2. Use a generative AI model to parse the user question and identify information related to 60s rock music. 3. Generate prompts and formulate appropriate responses to the user's questions. 4. Provide the collected information through a display or audio output device.

[0075] The following are some examples of prompt sentences: "Representative artists of 60s rock music include bands and popular singers. Representative songs include many hits from that time. For example, if you ask, 'What are the characteristics of rock music from the 60s and what are the famous bands?' we can provide you with this information."

[0076] In this way, the system of the present invention is able to quickly and accurately analyze specific questions from users and provide appropriate music information.

[0077] The flow of the specific process in the application example 1 will be described with reference to FIG.

[0078] Step 1: A user uses a smartphone application to input a specific question (e.g., "Tell me about rock music from the 60s"). The input text data is sent by the application to a server. The input here is the user's question text, and the output is the text data sent to the server.

[0079] Step 2: The server accesses the Internet and music databases to acquire music information. This is done using an acquisition means. In this step, the input is a search query based on the user's question, and the output is the acquired music-related information (e.g., data about rock music in the 60s). The server collects the required information from APIs and databases on the Internet and temporarily stores the information.

[0080] Step 3: The server uses a generative AI model to parse the user's question. During this parsing step, the user's input text is analyzed using natural language processing techniques to identify relevant music genres and artists. The input is the user's question text, and the output is the parsed theme or category (60's rock music in this example). The server uses the results of the parsing to extract information about a specific music genre.

[0081] Step 4: The server generates a prompt sentence in response to the user's question. In this step, an appropriate response text is created based on the analysis results. The input is the analysis results and collected music information, and the output is a completed response sentence to be provided to the user. The specific generation operation is a text generation process using a generative AI model.

[0082] Step 5: The server transmits the collected and generated information to the terminal using a means for providing music information. The terminal provides this information to the user through a display device or an audio output device. The input is the information transmitted from the server, and the output is a visual or audio presentation to the user. Specifically, text information is displayed on a display or audio information is played from a speaker.

[0083] Step 6: The user can then review the information provided and ask new questions or enter additional requests. At this step, the user's input is again sent to the server and the process repeats: the input is the user's new question or request, and the output is the updated information provided process.

[0084] Furthermore, an emotion engine that estimates the emotion of the user may be combined. That is, the identification processing unit 290 may estimate the emotion of the user using the emotion identification model 59, and perform identification processing using the emotion of the user.

[0085] This embodiment is composed of the following elements:

[0086] 1. Incorporating an emotion engine: - The system will incorporate an emotion engine to recognize the user's emotions. - The emotion engine recognizes emotions from the user's voice, facial expressions, text, etc., and customizes music and artist suggestions based on the recognition results. 2. Providing a music experience that responds to the user's emotions: - The system analyzes the user's emotional state through an emotion engine. - Based on the analysis results, the system will adjust music selection and the timing of information provided to provide the user with a more personalized music experience. As a specific example, an embodiment will be described in which a user makes a request to the system, "Tell me what song to listen to when I'm feeling sad." 1. Incorporating an emotion engine: - The system incorporates an emotion engine to recognize emotions from the user's voice, facial expressions, text, etc. - Analyze the user's tone of voice, changes in facial expressions, and the context of the text entered to understand the user's emotional state. 2. Providing a music experience that responds to the user's emotions: - The system analyzes the user's emotional state through an emotion engine. If a user requests "What songs do you listen to when you feel sad?" and the analyzed emotional state of the user is "sad", then the emotion engine will recognize the user's sad mood. - The system will select songs and artists that fit a sad mood and suggest them on the display or through audio output. - The emotion engine also monitors changes in the user's emotional state and adjusts music selection and information provision timing as needed.

[0087] The process flow will be explained below.

[0088] Step 1: Recognize the user's emotions. - The device collects information such as the user's voice, facial expressions, and text entered by the user. - The device transmits the collected information to the server. Step 2: Sentiment analysis by the sentiment engine. - The server inputs the received information into the emotion engine. - The emotion engine analyzes voice tone, facial expression changes, input text context, etc. to recognize the user's emotional state. Step 3: Select and suggest music according to emotions. - The server selects appropriate music based on the user's emotional state recognized by the emotion engine. - The server sends information about the selected music and artists to the device. Step 4: Deliver the music experience. - The device will display information about the received music and artists on the display and / or provide suggestions through audio output. - Users can express their emotions and feel comforted by the music provided.

[0089] Example 2 Next, a description will be given of Example 2. In the following description, the data processing device 12 is referred to as a "server" and the smart device 14 is referred to as a "terminal."

[0090] Conventional music provision systems provide music based on the user's preferences and questions, but they cannot take the user's emotions into account. This makes it difficult to provide music that is optimal for the user's current emotional state. In addition, because it is not possible to dynamically suggest music according to emotions, there is a problem in that it is not possible to increase user satisfaction.

[0091] The specific processing by the specific processing unit 290 of the data processing device 12 in the second embodiment is realized by the following means.

[0092] In this invention, the server includes a means for recognizing a user's emotion, a means for acquiring music information suitable for the user based on the recognized emotion, and a means for providing the acquired music information to the user, thereby making it possible to dynamically suggest music optimal to the user's emotional state.

[0093] "Means for recognizing the user's emotions" refers to technology that analyzes data such as the user's voice, facial expressions, and writing to determine the user's current emotional state.

[0094] "Means for acquiring music information appropriate for the user based on recognized emotions" refers to a technology that utilizes the results of emotion recognition to acquire information on music and artists that match the user's current emotional state from a music information database, etc.

[0095] The "means for providing the acquired music information to the user" refers to a technique for conveying the acquired music and artist information to the user through a display or audio output.

[0096] "Means for accessing the Internet or a music information database" refers to technology for connecting to a music information database on the Internet or locally or on the cloud, and searching for and acquiring music information.

[0097] "Means for providing information through a display or audio output" refers to devices or technologies for conveying acquired information to a user visually or audibly.

[0098] In order to implement the present invention, the following elements must be specifically configured and operated.

[0099] 1. Incorporating an emotion engine The server is equipped with an emotion engine to recognize the user's emotions. The emotion engine uses general emotion recognition technology such as "Microsoft (registered trademark) Azure (registered trademark) emotion recognition service." A user uses the device's microphone and camera to provide voice input, text input, facial expression data, and the like.

[0100] 2. User Emotion Recognition When a user speaks into the terminal, the terminal transmits the voice data to the server, and if necessary, facial expression data captured by the camera is also transmitted to the server. The server inputs the received data into the emotion engine to recognize the emotion. For example, if a user requests, "What song do you listen to when you feel sad?", the emotion engine recognizes the emotion "sad."

[0101] 3. Acquiring music information according to user emotions Based on the recognized emotion, the server accesses a music information database or a streaming service to obtain music information suitable for the user. Specifically, the server uses a general music information acquisition means such as a music database API. The server searches for "songs that fit a sad mood" and retrieves the results.

[0102] 4. Providing music information After the music information is obtained, the server transmits the information to the user's terminal. The terminal displays the received music information on the display and, if necessary, suggests the information to the user through audio output.

[0103] For example, when a user says to the terminal, "Please tell me what song you listen to when you feel sad," the specific operation is as follows.

[0104] 1. The user provides voice input. 2. The device sends the voice data to the server. 3. The server's emotion engine recognizes the emotion "sad." 4. The server uses a music database API to search for "songs that fit a sad mood." 5. The server sends the acquired music information to the terminal. 6. The device will show the results on the screen and also provide audio suggestions.

[0105] Examples of prompt statements "Generate a list of music suitable for when the user is sad." "Describe how you would adjust your musical suggestions if your emotional state changes."

[0106] This invention makes it possible to dynamically suggest music that is optimal for the emotional state of a user. Specifically, by using an emotion engine to recognize the user's emotion and providing optimal music information based on that emotion, it is possible to increase the user's satisfaction.

[0107] The flow of the identification process in the second embodiment will be described with reference to FIG.

[0108] Step 1: The user performs voice input. The user speaks into the device's microphone and says, "Please tell me what songs you listen to when you're feeling sad." This audio data is captured by the terminal. Input: User's voice data Output: Captured audio data

[0109] Step 2: The terminal transmits the voice data to the server. The device uses a network connection to transmit the captured audio data to a server. Input: Captured audio data Output: Audio data sent to the server

[0110] Step 3: The server inputs the voice data into an emotion engine to recognize emotions. The server passes the received voice data to emotion recognition software (for example, Microsoft Azure's emotion recognition service). The emotion engine analyzes the tone and content of the voice and recognizes the user's emotion as "sad." Input: Transmitted audio data Output: The recognized emotion (e.g. "sad")

[0111] Step 4: The server obtains music information based on the emotion. Based on the emotion recognition result, the server calls the music database API and sends a request to obtain related music information. Specifically, a query is sent to the API to search for "songs that fit a sad mood." Input: A recognized emotion (e.g. "sad") Output: Retrieved music information (song list)

[0112] Step 5: The server transmits the acquired music information to the terminal. The server sends the retrieved music information back to the device, including song titles and artist names in a list format. Input: Acquired music information (song list) Output: Music information sent to the device

[0113] Step 6: The device will display music information and provide audio suggestions. The terminal displays the received song list on the user's display. Additionally, it uses text-to-speech functionality to provide audio guidance to the user, such as, "Here's a list of songs that suit a sad mood." Input: Music information sent to the device Output: Display and audio suggestions

[0114] Step 7: The server continuously monitors the user's emotional state.

[0115] The server periodically receives new voice and facial expression data from the terminal and re-evaluates the user's emotions using the emotion engine. If the user's emotional state changes, new music information is obtained and re-suggestions are made as appropriate. Input: Voice and facial expression data sent periodically Output: New emotion recognition results and re-suggested music information

[0116] These processing steps allow the system to dynamically provide music selections that best suit the user's emotions.

[0117] (Application example 2) Next, a description will be given of Application Example 2. In the following description, the data processing device 12 will be referred to as a "server" and the smart device 14 will be referred to as a "terminal."

[0118] Conventional music distribution systems have had difficulty providing a personalized music experience that accurately reflects the user's feelings and emotions. Although there are systems that suggest music based on the user's preferences and questions, they lack the functionality to respond to the user's real-time emotional changes, and a deeper level of personalization is required.

[0119] The specific processing by the specific processing unit 290 of the data processing device 12 in the application example 2 is realized by the following means.

[0120] In this invention, the server includes a means for acquiring music information, a means for analyzing the user's emotions using an emotion recognition engine, a means for providing information on music and artists related to the user's preferences and emotions, and a means for adjusting the timing of music selection and information provision, thereby making it possible to provide a personalized music experience according to the user's emotional state.

[0121] The "means for acquiring music information" refers to a means for acquiring music information from a music database, the Internet, or the like.

[0122] An "emotion recognition engine" is software or hardware that analyzes a user's voice, facial expressions, text, etc., and recognizes the user's emotions.

[0123] "Means for providing information on music and artists related to the user's preferences and emotions" refers to means for providing information on music and artists based on the user's emotional state and preferences, and includes a display and audio output.

[0124] The "means for adjusting the timing of music selection and information provision" refers to a means for selecting music and providing information at optimal timing according to the user's emotional state.

[0125] The "means for monitoring changes in the user's emotional state" refers to a means for monitoring the user's emotional state in real time and adjusting the music experience in response to those changes.

[0126] In order to put the present invention into practice, it is first necessary to construct a system that includes an emotion recognition engine, a means for acquiring music information, and a means for selecting music and providing information. This system operates as follows.

[0127] The server has a means of accessing the Internet and a music information database to obtain music information. This means allows the server to always obtain the latest music information from the database. For example, the server can use the API of a well-known music streaming service.

[0128] An emotion recognition engine is implemented on the device and analyzes the user's voice, facial expression, text, etc. This makes it possible to recognize the user's emotions in real time. For example, a voice analysis API or facial expression recognition API is used for the emotion recognition engine. If the user inputs "I'm very sad today," the emotion recognition engine will recognize the emotion "sad."

[0129] Based on the recognized emotion, the server uses the music streaming service API to obtain information on music and artists that are suitable for the user's emotion. For example, it obtains a playlist that matches the "sad" mood. It then provides the information to the user via the device's display or audio output. For example, it may display "Recommended songs: ●● - ●●" on the display or notify the user by voice, "Here are some recommended songs."

[0130] In addition, the system monitors changes in the user's emotional state and adjusts music selection and information provision timing as necessary. Even if the user's emotion changes from "sad" to "happy," the system will detect this and suggest music appropriate to the new emotion. This allows the user to always enjoy a music experience that matches their emotion.

[0131] Specifically, the following prompt sentences are input to the generative AI model to perform emotion recognition: If a user types "I feel so sad today", an emotion recognition API will return the emotion "sad". Write a program that uses a music streaming service API to get a playlist that matches the "sad" mood and serves it to the user.

[0132] The system can provide a personalized music experience based on the user's emotional state, achieving a deeper level of personalization.

[0133] The flow of the specific process in the application example 2 will be described with reference to FIG.

[0134] Step 1: User emotion input The user inputs their mood or emotion into the device. For example, they can say, "I feel very sad today," or input it as text. This input becomes data for the next analysis.

[0135] Step 2: Emotion recognition The device uses the emotion recognition API to analyze the user's input data. Specifically, in the case of voice input, voice analysis is performed, and in the case of text input, sentence analysis is performed. Through this analysis, the user's emotion is recognized as "sad."

[0136] Step 3: Emotion-based music selection Based on the recognized emotion, the server uses the music streaming service API to search for music that matches the user's emotion. For example, if the emotion is "sad", it retrieves a playlist that matches the "sad" mood. At this stage, the retrieved data includes the song title, artist name, etc.

[0137] Step 4: Providing music information The music information obtained from the server is provided to the user via the device's display or audio output. Specifically, the device may show "Recommended songs: XX - XX" on the display or notify the user by voice, "Here are some recommended songs."

[0138] Step 5: Monitoring user emotional changes The system uses the device's emotion recognition engine to monitor the user's emotional state in real time, for example, re-recognizing emotions each time the user enters new voice or text input and determining whether the emotional state has changed.

[0139] Step 6: Timing adjustment of music provision If the user's emotion is recognized as having changed, the server again uses the music streaming service API to obtain music based on the new emotion (i.e., the changed emotion) and provides it to the user. For example, if the emotion changes from "sad" to "happy," the server obtains and provides a playlist that matches the "happy" mood.

[0140] This allows users to always enjoy a music experience that matches their emotions.

[0141] The specific processing unit 290 transmits the result of the specific processing to the smart device 14. In the smart device 14, the control unit 46A causes the output device 40 to output the result of the specific processing. The microphone 38B acquires a voice indicating a user input for the result of the specific processing. The control unit 46A transmits the voice data indicating the user input acquired by the microphone 38B to the data processing device 12. In the data processing device 12, the specific processing unit 290 acquires the voice data.

[0142] The data generation model 58 is a so-called generative AI (Artificial Intelligence). An example of the data generation model 58 is ChatGPT (registered trademark) (Internet search<URL: https: / / openai.com / blog / chatgpt> ), Gemini (registered trademark) (Internet search <url: https: gemini.google.com ?hl="ja">) and other generation AIs. The data generation model 58 is obtained by making a neural network perform deep learning. A prompt including an instruction is input to the data generation model 58, and inference data such as voice data indicating a voice, text data indicating a text, and image data indicating an image is input. The data generation model 58 performs inference on the input inference data according to the instruction indicated by the prompt, and outputs the inference result in a data format such as voice data and text data. Here, inference refers to, for example, analysis, classification, prediction, and / or summarization.

[0143] In the above embodiment, an example was given in which the specific process was performed by the data processing device 12, but the technology of the present disclosure is not limited to this, and the specific process may be performed by the smart device 14.

[0144] [Second embodiment] FIG. 3 shows an example of the configuration of a data processing system 210 according to the second embodiment.

[0145] 3, the data processing system 210 includes the data processing device 12 and smart glasses 214. An example of the data processing device 12 is a server.

[0146] The data processing device 12 includes a computer 22, a database 24, and a communication I / F 26. The computer 22 is an example of a "computer" according to the technology of the present disclosure. The computer 22 includes a processor 28, a RAM 30, and a storage 32. The processor 28, the RAM 30, and the storage 32 are connected to a bus 34. The database 24 and the communication I / F 26 are also connected to the bus 34. The communication I / F 26 is connected to a network 54. Examples of the network 54 include a wide area network (WAN) and / or a local area network (LAN).

[0147] The smart glasses 214 include a computer 36, a microphone 238, a speaker 240, a camera 42, and a communication I / F 44. The computer 36 includes a processor 46, a RAM 48, and a storage 50. The processor 46, the RAM 48, and the storage 50 are connected to a bus 52. The microphone 238, the speaker 240, and the camera 42 are also connected to the bus 52.

[0148] The microphone 238 receives instructions and the like from the user 20 by receiving voice uttered by the user 20. The microphone 238 captures the voice uttered by the user 20, converts the captured voice into audio data, and outputs it to the processor 46. The speaker 240 outputs the voice according to instructions from the processor 46.

[0149] Camera 42 is a small digital camera equipped with an optical system including a lens, an aperture, and a shutter, and an imaging element such as a CMOS (Complementary Metal-Oxide-Semiconductor) image sensor or a CCD (Charge Coupled Device) image sensor, and captures images of the surroundings of user 20 (e.g., an imaging range defined by an angle of view equivalent to the width of the field of vision of an average healthy person).

[0150] The communication I / F 44 is connected to a network 54. The communication I / Fs 44 and 26 are responsible for transmitting and receiving various types of information between the processor 46 and the processor 28 via the network 54. The transmission and reception of various types of information between the processor 46 and the processor 28 using the communication I / Fs 44 and 26 is performed in a secure state.

[0151] Fig. 4 shows an example of main functions of the data processing device 12 and the smart glasses 214. As shown in Fig. 4, in the data processing device 12, a specific process is performed by the processor 28. A specific process program 56 is stored in the storage 32.

[0152] The specific processing program 56 is an example of a "program" according to the technology of the present disclosure. The processor 28 reads the specific processing program 56 from the storage 32, and executes the read specific processing program 56 on the RAM 30. The specific processing is realized by the processor 28 operating as a specific processing unit 290 in accordance with the specific processing program 56 executed on the RAM 30.

[0153] The storage 32 stores a data generation model 58 and an emotion identification model 59. The data generation model 58 and the emotion identification model 59 are used by the identification processing unit 290.

[0154] In the smart glasses 214, the reception output process is performed by the processor 46. A reception output program 60 is stored in the storage 50. The processor 46 reads the reception output program 60 from the storage 50, and executes the read reception output program 60 on the RAM 48. The reception output process is realized by the processor 46 operating as the control unit 46A in accordance with the reception output program 60 executed on the RAM 48.

[0155] Next, a description will be given of the identification process performed by the identification processing unit 290 of the data processing device 12. In the following description, the data processing device 12 will be referred to as a "server" and the smart glasses 214 will be referred to as a "terminal".

[0156] An embodiment for implementing the disclosed technology comprises the following elements. 1. How to get music information: The server obtains information about music history and artists by accessing the Internet and a music information database. 2. Means of analyzing user preferences and questions: - The server analyzes the information provided by the user and selects appropriate music and artists based on the user's preferences and questions. The server may also select appropriate music and artists based on the user's age, gender, and usual music listening habits. 3. Means of providing information about music and artists related to the user's preferences and questions: The server transmits the acquired information to the terminal, and the terminal provides information on music and artists related to the user's preferences and questions through a display and / or audio output. As a specific example, an embodiment will be described in which a user asks the question "Tell me about rock music from the 60's." 1. How to get music information: - The server accesses the Internet and music information databases to obtain information about rock music from the 1960s, such as representative bands and artists, popular songs, and the social context of the time. 2. Means of analyzing user preferences and questions: - The server parses the user's question, "What is rock music from the 60's?" and identifies information relevant to the question. 3. Means of providing information about music and artists related to the user's preferences and questions: The server transmits the acquired information to the terminal, and the terminal displays the characteristics, representative artists, popular songs, etc. of rock music in the 60s on the display. The terminal may also provide the characteristics, representative artists, popular songs, etc. of rock music in the 60s to the user through audio output.

[0157] The process flow will be explained below. Step 1: Receive user input such as questions and preferences. - The user communicates questions and preferences to the system through the terminal. - The terminal sends the user's input to the server. Step 2: Search and retrieve information based on your questions and preferences. - The server receives and parses the user input. - The server accesses the Internet and music information databases to search for music and artist information related to the user's question and preferences. - The server organizes the obtained information and extracts the necessary information. Step 3: Provide information. - The server sends the acquired information to the terminal. - The device displays information on a display and conveys information through audio output. - Users can enjoy learning from the information provided.

[0158] Example 1 Next, a description will be given of Example 1. In the following description, the data processing device 12 is referred to as a "server" and the smart glasses 214 are referred to as a "terminal".

[0159] Conventional music information systems have difficulty in quickly and accurately responding to the diverse tastes and questions of users. In addition, few systems can provide a consistent process for efficiently collecting, analyzing, and presenting the information users want. As a result, the user experience is often poor and dissatisfied. In addition, when users want detailed information related to a specific era or genre, they have to search manually or refer to multiple sources, which is time-consuming and laborious.

[0160] The specific processing by the specific processing unit 290 of the data processing device 12 in the first embodiment is realized by the following means.

[0161] In this invention, the server includes a means for accessing the Internet or a database to acquire music information, a means for analyzing information and questions provided by the user to identify appropriate music and artists, and a means for transmitting information on the identified music and artists to the terminal and providing it to the user through a display or audio output, thereby making it possible to quickly and accurately collect, analyze, and provide the music information desired by the user.

[0162] The "Internet" is a collection of globally connected computer networks and an infrastructure for the exchange of information and communication.

[0163] A "database" is a systematically organized collection of information that can be efficiently accessed, searched, and managed through a specific application or query.

[0164] "User" refers to an individual or group that uses this system to obtain music information.

[0165] A "terminal" is a device from which a user receives information, and includes a personal computer, a smartphone, a tablet, etc.

[0166] "Display" refers to a screen or monitor for visually displaying information.

[0167] "Audio output" refers to technologies and devices that provide information to a user as audio.

[0168] "Music information" refers to all data and knowledge about music, such as music history, artist information, song details, and genre characteristics.

[0169] "Analysis" refers to the process of examining given information or data in detail to understand its structure and meaning.

[0170] "Acquisition" refers to the act of collecting information and making it available in a usable form.

[0171] "Providing" refers to the act of supplying information to a user and making it available for use.

[0172] A "network" refers to an infrastructure that connects multiple computers and devices and enables the exchange of information.

[0173] The present invention relates to a system for quickly and accurately providing music information desired by a user. Specific embodiments will be described below.

[0174] How to get music information The server uses the Internet and databases to collect information about music, utilizing existing web services such as Google Search API and music database API. For example, if a user asks the question "Tell me about rock music from the 60s," the server first uses the Google search API to gather the necessary information from the Internet. It also uses the music database API to obtain information about popular artists and songs from the 60s. This data is stored in the server's internal database and used for subsequent processing.

[0175] A means of analyzing user preferences and questions The server uses natural language processing techniques to analyze the questions and information entered by users, specifically using high-performance generative AI models such as OpenAI's ChatGPT model. If a user types in "Tell me about rock music from the 60s," the server will analyze the content and extract related keywords such as "60s," "rock music," etc. The results of this analysis will become the basis for the server to collect appropriate information and provide it to the user.

[0176] Here is an example of a prompt for parsing: Prompt: "A user has asked the question 'Rock music from the 60s.' Please analyze the information related to this question."

[0177] A way to provide users with music and artist information relevant to their preferences and questions In order to provide the collected and analyzed information to the user, the server first converts the data into an appropriate format, and then transmits the data to the user's terminal. The terminal displays the information received from the server on the display and outputs voice as necessary. Specifically, the terminal can use a Text-to-Speech (TTS) engine to play back text information as voice. For example, if a user asks "Tell me about rock music from the 60's," the server might provide text like this: "In 60's rock music, Singer A, Singer B, etc. are representative artists. Also, the XX Music Festival in 1969 is very famous." The device can both display this text information and play it aloud using a TTS engine.

[0178] Below are some examples of informational prompts: Prompt: "Provide the user with the information you have retrieved based on the user's question. Provide text on the display and speech output."

[0179] The flow of the identification process in the first embodiment will be described with reference to FIG.

[0180] Step 1: The server receives input from the user. Input: A user question (e.g., "What is rock music from the 60's?") Specific operation: The user enters a question into the terminal, which then sends this information to the server. Output: The server receives the user's question.

[0181] Step 2: The server uses a generative AI model to analyze the user's question. Input: A user question (e.g., "What is rock music from the 60's?") How it works: The server uses natural language processing techniques, such as OpenAI's ChatGPT model, to analyze the question, understand the intent of the question, and extract relevant keywords ("60s", "rock music"). Output: A list of keywords resulting from the analysis ("60s", "rock music").

[0182] Step 3: The server accesses the Internet and databases to retrieve relevant information. Input: Analyzed keyword list ("60s", "Rock music") How it works: The server uses the Google Search API to execute queries such as "60s rock music history" and collects the resulting webpage data. It also uses a music database API to retrieve information about popular artists and songs from the 60s. Output: Music information data (e.g. representative artists, popular songs, social background, etc.).

[0183] Step 4: The server organizes the retrieved information and converts it into a format for presentation to the user. Input: Retrieved music information data Specific operation: The server organizes the text data and organizes it into a format that is easy for the user to understand. For example, it may classify the acquired information into categories such as "famous artists," "popular songs," and "XX Music Festival in 1969." Output: Formatted information (e.g. "60s rock music is represented by artists Singer A, Singer B, etc. Also, the XX music festival in 1969 is very famous.").

[0184] Step 5: The server sends the formatted information to the user's terminal. Input: Formatted information Specific operation: The server sends information using the network address of the terminal. Output: The terminal receives the formatted information.

[0185] Step 6: The terminal displays the received information on a display and outputs audio as necessary. Input: Formatted information What it does: The device displays text information on a display and can also provide information aloud using a Text-to-Speech (TTS) engine. Output: The user receives information visually and audibly.

[0186] (Application example 1) Next, a description will be given of Application Example 1. In the following description, the data processing device 12 will be referred to as a "server" and the smart glasses 214 will be referred to as a "terminal".

[0187] Conventional music recommendation systems have limited functionality in providing appropriate music information based on users' questions and preferences. In particular, responding to specific user questions and diverse music genres is a challenge, and effective solutions to improve user experience are required. Furthermore, advanced functions for quickly and accurately analyzing user-entered questions and providing relevant music information are lacking.

[0188] The specific processing by the specific processing unit 290 of the data processing device 12 in the application example 1 is realized by the following means.

[0189] In this invention, the server includes a means for acquiring music information, a means for analyzing a user's preferences and questions, a means for providing music-related information related to the user's preferences and questions, a means for analyzing the user's questions using a generative AI model, and a means for generating a prompt sentence, which makes it possible to quickly and accurately analyze a specific question from a user and provide appropriate music information.

[0190] "Means for acquiring music information" refers to means for accessing the Internet or a music database to gather music-related information.

[0191] The "means for analyzing user preferences and questions" refers to a means for analyzing information provided by a user to understand the contents of the preferences and questions.

[0192] The "means for providing music-related information related to the user's preferences and questions" refers to a means for providing music-related information that the user is interested in based on collected information through a display device or audio output device.

[0193] "Means for analyzing a user's question using a generative AI model" refers to means for analyzing the content of a user's question using an AI model that employs natural language processing technology and identifying related information.

[0194] A "means for generating a prompt sentence" is a means for automatically creating text to generate an appropriate response to a user's question.

[0195] DETAILED DESCRIPTION OF THE PREFERRED EMBODIMENTS The present invention is configured to provide appropriate music-related information in response to specific user queries in a music recommendation system.

[0196] First, a user uses a smartphone application. When the user inputs a question, the question is sent to the application server. The server first accesses the Internet or a music database using a music information acquisition means to collect related music information.

[0197] The server then analyzes the user's question using a generative AI model, including the Transformer model, which is based on natural language processing techniques, that analyzes the text entered by the user and identifies relevant music genres and artists based on the content.

[0198] After obtaining the parsed result of the question, the server generates a prompt sentence, which is used to formulate an appropriate response to the user's question.

[0199] Finally, the server collects music information and provides the information to the user through a display device and audio output device using a means for providing music-related information related to the user's preferences and questions. As a specific example, information such as characteristics of music genres, representative artists, and popular songs can be displayed on the display and the same information can be reproduced by audio.

[0200] As a specific operational procedure of the embodiment, consider a case where a user asks the question, "Tell me about rock music from the 60's." In response to this question, the system operates as follows.

[0201] 1. The server accesses the Internet and a music database to retrieve information about 60's rock music. 2. Use a generative AI model to parse the user question and identify information related to 60s rock music. 3. Generate prompts and formulate appropriate responses to the user's questions. 4. Provide the collected information through a display or audio output device.

[0202] The following are some examples of prompt sentences: "Representative artists of 60s rock music include bands and popular singers. Representative songs include many hits from that time. For example, if you ask, 'What are the characteristics of rock music from the 60s and what are the famous bands?' we can provide you with this information."

[0203] In this way, the system of the present invention is able to quickly and accurately analyze specific questions from users and provide appropriate music information.

[0204] The flow of the specific process in the application example 1 will be described with reference to FIG.

[0205] Step 1: A user uses a smartphone application to input a specific question (e.g., "Tell me about rock music from the 60s"). The input text data is sent by the application to a server. The input here is the user's question text, and the output is the text data sent to the server.

[0206] Step 2: The server accesses the Internet and music databases to acquire music information. This is done using an acquisition means. In this step, the input is a search query based on the user's question, and the output is the acquired music-related information (e.g., data about rock music in the 60s). The server collects the required information from APIs and databases on the Internet and temporarily stores the information.

[0207] Step 3: The server uses a generative AI model to parse the user's question. During this parsing step, the user's input text is analyzed using natural language processing techniques to identify relevant music genres and artists. The input is the user's question text, and the output is the parsed theme or category (60's rock music in this example). The server uses the results of the parsing to extract information about a specific music genre.

[0208] Step 4: The server generates a prompt sentence in response to the user's question. In this step, an appropriate response text is created based on the analysis results. The input is the analysis results and collected music information, and the output is a completed response sentence to be provided to the user. The specific generation operation is a text generation process using a generative AI model.

[0209] Step 5: The server transmits the collected and generated information to the terminal using a means for providing music information. The terminal provides this information to the user through a display device or an audio output device. The input is the information transmitted from the server, and the output is a visual or audio presentation to the user. Specifically, text information is displayed on a display or audio information is played from a speaker.

[0210] Step 6: The user can then review the information provided and ask new questions or enter additional requests. At this step, the user's input is again sent to the server and the process repeats: the input is the user's new question or request, and the output is the updated information provided process.

[0211] In addition, an emotion engine that estimates the user's emotion may be further combined. That is, the identification processing unit 290 may estimate the user's emotion using the emotion identification model 59, and perform identification processing using the user's emotion.

[0212] This embodiment is composed of the following elements: 1. Incorporating an emotion engine: - The system will incorporate an emotion engine to recognize the user's emotions. - The emotion engine recognizes emotions from the user's voice, facial expressions, text, etc., and customizes music and artist suggestions based on the recognition results. 2. Providing a music experience that responds to the user's emotions: - The system analyzes the user's emotional state through an emotion engine. - Based on the analysis results, the system will adjust music selection and the timing of information provided to provide the user with a more personalized music experience. As a specific example, an embodiment will be described in which a user makes a request to the system, "Tell me what song to listen to when I'm feeling sad." 1. Incorporating an emotion engine: - The system incorporates an emotion engine to recognize emotions from the user's voice, facial expressions, text, etc. - Analyze the user's tone of voice, changes in facial expressions, and the context of the text entered to understand the user's emotional state. 2. Providing a music experience that responds to the user's emotions: - The system analyzes the user's emotional state through an emotion engine. If a user requests "What songs do you listen to when you feel sad?" and the analyzed emotional state of the user is "sad", then the emotion engine will recognize the user's sad mood. - The system will select songs and artists that fit a sad mood and suggest them on the display or through audio output. - The emotion engine also monitors changes in the user's emotional state and adjusts music selection and information provision timing as needed.

[0213] The process flow will be explained below.

[0214] Step 1: Recognize the user's emotions. - The device collects information such as the user's voice, facial expressions, and text entered by the user. - The device transmits the collected information to the server. Step 2: Sentiment analysis by the sentiment engine. - The server inputs the received information into the emotion engine. - The emotion engine analyzes voice tone, facial expression changes, input text context, etc. to recognize the user's emotional state. Step 3: Select and suggest music according to emotions. - The server selects appropriate music based on the user's emotional state recognized by the emotion engine. - The server sends information about the selected music and artists to the device. Step 4: Deliver the music experience. - The device will display information about the received music and artists on the display and / or provide suggestions through audio output. - Users can express their emotions and feel comforted by the music provided.

[0215] Example 2 Next, a description will be given of Example 2. In the following description, the data processing device 12 is referred to as a "server" and the smart glasses 214 are referred to as a "terminal".

[0216] Conventional music provision systems provide music based on the user's preferences and questions, but they cannot take the user's emotions into account. This makes it difficult to provide music that is optimal for the user's current emotional state. In addition, because it is not possible to dynamically suggest music according to emotions, there is a problem in that it is not possible to increase user satisfaction.

[0217] The specific processing by the specific processing unit 290 of the data processing device 12 in the second embodiment is realized by the following means.

[0218] In this invention, the server includes a means for recognizing a user's emotion, a means for acquiring music information suitable for the user based on the recognized emotion, and a means for providing the acquired music information to the user, thereby making it possible to dynamically suggest music optimal to the user's emotional state.

[0219] "Means for recognizing the user's emotions" refers to technology that analyzes data such as the user's voice, facial expressions, and writing to determine the user's current emotional state.

[0220] "Means for acquiring music information appropriate for the user based on recognized emotions" refers to a technology that utilizes the results of emotion recognition to acquire information on music and artists that match the user's current emotional state from a music information database, etc.

[0221] The "means for providing the acquired music information to the user" refers to a technique for conveying the acquired music and artist information to the user through a display or audio output.

[0222] "Means for accessing the Internet or a music information database" refers to technology for connecting to a music information database on the Internet or locally or on the cloud, and searching for and acquiring music information.

[0223] "Means for providing information through a display or audio output" refers to devices or technologies for conveying acquired information to a user visually or audibly.

[0224] In order to implement the present invention, the following elements must be specifically configured and operated.

[0225] 1. Incorporating an emotion engine The server is equipped with an emotion engine to recognize the user's emotions. The emotion engine uses general emotion recognition technologies such as the Microsoft Azure emotion recognition service. A user uses the device's microphone and camera to provide voice input, text input, facial expression data, and the like.

[0226] 2. User Emotion Recognition When a user speaks into the terminal, the terminal transmits the voice data to the server, and if necessary, facial expression data captured by the camera is also transmitted to the server. The server inputs the received data into the emotion engine to recognize the emotion. For example, if a user requests, "What song do you listen to when you feel sad?", the emotion engine recognizes the emotion "sad."

[0227] 3. Acquiring music information according to user emotions Based on the recognized emotion, the server accesses a music information database or a streaming service to obtain music information suitable for the user. Specifically, the server uses a general music information acquisition means such as a music database API. The server searches for "songs that fit a sad mood" and retrieves the results.

[0228] 4. Providing music information After the music information is obtained, the server transmits the information to the user's terminal. The terminal displays the received music information on the display and, if necessary, suggests the information to the user through audio output.

[0229] For example, when a user says to the terminal, "Please tell me what song you listen to when you feel sad," the specific operation is as follows. 1. The user provides voice input. 2. The device sends the voice data to the server. 3. The server's emotion engine recognizes the emotion "sad." 4. The server uses a music database API to search for "songs that fit a sad mood." 5. The server sends the acquired music information to the terminal. 6. The device will show the results on the screen and also provide audio suggestions.

[0230] Examples of prompt statements "Generate a list of music suitable for when the user is sad." "Describe how you would adjust your musical suggestions if your emotional state changes."

[0231] This invention makes it possible to dynamically suggest music that is optimal for the emotional state of a user. Specifically, by using an emotion engine to recognize the user's emotion and providing optimal music information based on that emotion, it is possible to increase the user's satisfaction.

[0232] The flow of the identification process in the second embodiment will be described with reference to FIG.

[0233] Step 1: The user performs voice input. The user speaks into the device's microphone and says, "Please tell me what songs you listen to when you're feeling sad." This audio data is captured by the terminal. Input: User's voice data Output: Captured audio data

[0234] Step 2: The terminal transmits the voice data to the server. The device uses a network connection to transmit the captured audio data to a server. Input: Captured audio data Output: Audio data sent to the server

[0235] Step 3: The server inputs the voice data into an emotion engine to recognize emotions. The server passes the received voice data to emotion recognition software (for example, Microsoft Azure's emotion recognition service). The emotion engine analyzes the tone and content of the voice and recognizes the user's emotion as "sad." Input: Transmitted audio data Output: The recognized emotion (e.g. "sad")

[0236] Step 4: The server obtains music information based on the emotion. Based on the emotion recognition result, the server calls the music database API and sends a request to obtain related music information. Specifically, a query is sent to the API to search for "songs that fit a sad mood." Input: A recognized emotion (e.g. "sad") Output: Retrieved music information (song list)

[0237] Step 5: The server transmits the acquired music information to the terminal. The server sends the retrieved music information back to the device, including song titles and artist names in a list format. Input: Acquired music information (song list) Output: Music information sent to the device

[0238] Step 6: The device will display music information and provide audio suggestions. The terminal displays the received song list on the user's display. Additionally, it uses text-to-speech functionality to provide audio guidance to the user, such as, "Here's a list of songs that suit a sad mood." Input: Music information sent to the device Output: Display and audio suggestions

[0239] Step 7: The server continuously monitors the user's emotional state. The server periodically receives new voice and facial expression data from the terminal and re-evaluates the user's emotions using the emotion engine. If the user's emotional state changes, new music information is obtained and re-suggestions are made as appropriate. Input: Voice and facial expression data sent periodically Output: New emotion recognition results and re-suggested music information

[0240] These processing steps allow the system to dynamically provide music selections that best suit the user's emotions.

[0241] (Application example 2) Next, a description will be given of Application Example 2. In the following description, the data processing device 12 will be referred to as a "server" and the smart glasses 214 will be referred to as a "terminal".

[0242] Conventional music distribution systems have had difficulty providing a personalized music experience that accurately reflects the user's feelings and emotions. Although there are systems that suggest music based on the user's preferences and questions, they lack the functionality to respond to the user's real-time emotional changes, and a deeper level of personalization is required.

[0243] The specific processing by the specific processing unit 290 of the data processing device 12 in the application example 2 is realized by the following means.

[0244] In this invention, the server includes a means for acquiring music information, a means for analyzing the user's emotions using an emotion recognition engine, a means for providing information on music and artists related to the user's preferences and emotions, and a means for adjusting the timing of music selection and information provision, thereby making it possible to provide a personalized music experience according to the user's emotional state.

[0245] The "means for acquiring music information" refers to a means for acquiring music information from a music database, the Internet, or the like.

[0246] An "emotion recognition engine" is software or hardware that analyzes a user's voice, facial expressions, text, etc., and recognizes the user's emotions.

[0247] "Means for providing information on music and artists related to the user's preferences and emotions" refers to means for providing information on music and artists based on the user's emotional state and preferences, and includes a display and audio output.

[0248] The "means for adjusting the timing of music selection and information provision" refers to a means for selecting music and providing information at optimal timing according to the user's emotional state.

[0249] The "means for monitoring changes in the user's emotional state" refers to a means for monitoring the user's emotional state in real time and adjusting the music experience in response to those changes.

[0250] In order to put the present invention into practice, it is first necessary to construct a system that includes an emotion recognition engine, a means for acquiring music information, and a means for selecting music and providing information. This system operates as follows.

[0251] The server has a means of accessing the Internet and a music information database to obtain music information. This means allows the server to always obtain the latest music information from the database. For example, the server can use the API of a well-known music streaming service.

[0252] An emotion recognition engine is implemented on the device and analyzes the user's voice, facial expression, text, etc. This makes it possible to recognize the user's emotions in real time. For example, a voice analysis API or facial expression recognition API is used for the emotion recognition engine. If the user inputs "I'm very sad today," the emotion recognition engine will recognize the emotion "sad."

[0253] Based on the recognized emotion, the server uses the music streaming service API to obtain information on music and artists that are suitable for the user's emotion. For example, it obtains a playlist that matches the "sad" mood. It then provides the information to the user via the device's display or audio output. For example, it may display "Recommended songs: ●● - ●●" on the display or notify the user by voice, "Here are some recommended songs."

[0254] In addition, the system monitors changes in the user's emotional state and adjusts music selection and information provision timing as necessary. Even if the user's emotion changes from "sad" to "happy," the system will detect this and suggest music appropriate to the new emotion. This allows the user to always enjoy a music experience that matches their emotion.

[0255] Specifically, the following prompt sentences are input to the generative AI model to perform emotion recognition: If a user types "I feel so sad today", an emotion recognition API will return the emotion "sad". Write a program that uses a music streaming service API to get a playlist that matches the "sad" mood and serves it to the user.

[0256] The system can provide a personalized music experience based on the user's emotional state, achieving a deeper level of personalization.

[0257] The flow of the specific process in the application example 2 will be described with reference to FIG.

[0258] Step 1: User emotion input The user inputs their mood or emotion into the device. For example, they can say, "I feel very sad today," or input it as text. This input becomes data for the next analysis.

[0259] Step 2: Emotion recognition The device uses the emotion recognition API to analyze the user's input data. Specifically, in the case of voice input, voice analysis is performed, and in the case of text input, sentence analysis is performed. Through this analysis, the user's emotion is recognized as "sad."

[0260] Step 3: Emotion-based music selection Based on the recognized emotion, the server uses the music streaming service API to search for music that matches the user's emotion. For example, if the emotion is "sad", it retrieves a playlist that matches the "sad" mood. At this stage, the retrieved data includes the song title, artist name, etc.

[0261] Step 4: Providing music information The music information obtained from the server is provided to the user via the device's display or audio output. Specifically, the device may show "Recommended songs: XX - XX" on the display or notify the user by voice, "Here are some recommended songs."

[0262] Step 5: Monitoring user emotional changes The system uses the device's emotion recognition engine to monitor the user's emotional state in real time, for example, re-recognizing emotions each time the user enters new voice or text input and determining whether the emotional state has changed.

[0263] Step 6: Timing adjustment of music provision If the user's emotion is recognized as having changed, the server again uses the music streaming service API to obtain music based on the new emotion (i.e., the changed emotion) and provides it to the user. For example, if the emotion changes from "sad" to "happy," the server obtains and provides a playlist that matches the "happy" mood.

[0264] This allows users to always enjoy a music experience that matches their emotions.

[0265] The specific processing unit 290 transmits the result of the specific processing to the smart glasses 214. In the smart glasses 214, the control unit 46A causes the speaker 240 to output the result of the specific processing. The microphone 238 acquires a voice indicating a user input for the result of the specific processing. The control unit 46A transmits the voice data indicating the user input acquired by the microphone 238 to the data processing device 12. In the data processing device 12, the specific processing unit 290 acquires the voice data.

[0266] The data generation model 58 is a so-called generative AI (Artificial Intelligence). An example of the data generation model 58 is ChatGPT (Internet search<URL: https: / / openai.com / blog / chatgpt> ), Gemini (Internet search <url: https: gemini.google.com ?hl="ja">) and other generation AIs. The data generation model 58 is obtained by making a neural network perform deep learning. A prompt including an instruction is input to the data generation model 58, and inference data such as voice data indicating a voice, text data indicating a text, and image data indicating an image is input. The data generation model 58 performs inference on the input inference data according to the instruction indicated by the prompt, and outputs the inference result in a data format such as voice data and text data. Here, inference refers to, for example, analysis, classification, prediction, and / or summarization.

[0267] In the above embodiment, an example was given in which the specific processing is performed by the data processing device 12, but the technology of the present disclosure is not limited to this, and the specific processing may be performed by the smart glasses 214.

[0268] [Third embodiment] FIG. 5 shows an example of the configuration of a data processing system 310 according to the third embodiment.

[0269] 5, the data processing system 310 includes the data processing device 12 and a headset type terminal 314. An example of the data processing device 12 is a server.

[0270] The data processing device 12 includes a computer 22, a database 24, and a communication I / F 26. The computer 22 is an example of a "computer" according to the technology of the present disclosure. The computer 22 includes a processor 28, a RAM 30, and a storage 32. The processor 28, the RAM 30, and the storage 32 are connected to a bus 34. The database 24 and the communication I / F 26 are also connected to the bus 34. The communication I / F 26 is connected to a network 54. Examples of the network 54 include a wide area network (WAN) and / or a local area network (LAN).

[0271] The headset type terminal 314 includes a computer 36, a microphone 238, a speaker 240, a camera 42, a communication I / F 44, and a display 343. The computer 36 includes a processor 46, a RAM 48, and a storage 50. The processor 46, the RAM 48, and the storage 50 are connected to a bus 52. The microphone 238, the speaker 240, the camera 42, and the display 343 are also connected to the bus 52.

[0272] The microphone 238 receives instructions and the like from the user 20 by receiving voice uttered by the user 20. The microphone 238 captures the voice uttered by the user 20, converts the captured voice into audio data, and outputs it to the processor 46. The speaker 240 outputs the voice according to instructions from the processor 46.

[0273] Camera 42 is a small digital camera equipped with an optical system including a lens, an aperture, and a shutter, and an imaging element such as a CMOS (Complementary Metal-Oxide-Semiconductor) image sensor or a CCD (Charge Coupled Device) image sensor, and captures images of the surroundings of user 20 (e.g., an imaging range defined by an angle of view equivalent to the width of the field of vision of an average healthy person).

[0274] The communication I / F 44 is connected to a network 54. The communication I / Fs 44 and 26 are responsible for transmitting and receiving various types of information between the processor 46 and the processor 28 via the network 54. The transmission and reception of various types of information between the processor 46 and the processor 28 using the communication I / Fs 44 and 26 is performed in a secure state.

[0275] Fig. 6 shows an example of main functions of the data processing device 12 and the headset type terminal 314. As shown in Fig. 6, in the data processing device 12, a specific process is performed by the processor 28. A specific process program 56 is stored in the storage 32.

[0276] The specific processing program 56 is an example of a "program" according to the technology of the present disclosure. The processor 28 reads the specific processing program 56 from the storage 32, and executes the read specific processing program 56 on the RAM 30. The specific processing is realized by the processor 28 operating as a specific processing unit 290 in accordance with the specific processing program 56 executed on the RAM 30.

[0277] The storage 32 stores a data generation model 58 and an emotion identification model 59. The data generation model 58 and the emotion identification model 59 are used by the identification processing unit 290.

[0278] In the headset type terminal 314, the reception output process is performed by the processor 46. A reception output program 60 is stored in the storage 50. The processor 46 reads the reception output program 60 from the storage 50, and executes the read reception output program 60 on the RAM 48. The reception output process is realized by the processor 46 operating as the control unit 46A in accordance with the reception output program 60 executed on the RAM 48.

[0279] Next, a description will be given of the identification process performed by the identification processing unit 290 of the data processing device 12. In the following description, the data processing device 12 will be referred to as the "server", and the headset type terminal 314 will be referred to as the "terminal".

[0280] An embodiment for implementing the disclosed technology comprises the following elements. 1. How to get music information: The server obtains information about music history and artists by accessing the Internet and a music information database. 2. Means of analyzing user preferences and questions: - The server analyzes the information provided by the user and selects appropriate music and artists based on the user's preferences and questions. The server may also select appropriate music and artists based on the user's age, gender, and usual music listening habits. 3. Means of providing information about music and artists related to the user's preferences and questions: The server transmits the acquired information to the terminal, and the terminal provides information on music and artists related to the user's preferences and questions through a display and / or audio output. As a specific example, an embodiment will be described in which a user asks the question "Tell me about rock music from the 60's." 1. How to get music information: - The server accesses the Internet and music information databases to obtain information about rock music from the 1960s, such as representative bands and artists, popular songs, and the social context of the time. 2. Means of analyzing user preferences and questions: - The server parses the user's question, "What is rock music from the 60's?" and identifies information relevant to the question. 3. Means of providing information about music and artists related to the user's preferences and questions: The server transmits the acquired information to the terminal, and the terminal displays the characteristics, representative artists, popular songs, etc. of rock music in the 60s on the display. The terminal may also provide the characteristics, representative artists, popular songs, etc. of rock music in the 60s to the user through audio output.

[0281] The process flow will be explained below.

[0282] Step 1: Receive user input such as questions and preferences. - The user communicates questions and preferences to the system through the terminal. - The terminal sends the user's input to the server. Step 2: Search and retrieve information based on your questions and preferences. - The server receives and parses the user input. - The server accesses the Internet and music information databases to search for music and artist information related to the user's question and preferences. - The server organizes the obtained information and extracts the necessary information. Step 3: Provide information. - The server sends the acquired information to the terminal. - The device displays information on a display and conveys information through audio output. - Users can enjoy learning from the information provided.

[0283] Example 1 Next, a description will be given of Example 1. In the following description, the data processing device 12 will be referred to as a "server" and the headset type terminal 314 will be referred to as a "terminal".

[0284] Conventional music information systems have difficulty in quickly and accurately responding to the diverse tastes and questions of users. In addition, few systems can provide a consistent process for efficiently collecting, analyzing, and presenting the information users want. As a result, the user experience is often poor and dissatisfied. In addition, when users want detailed information related to a specific era or genre, they have to search manually or refer to multiple sources, which is time-consuming and laborious.

[0285] The specific processing by the specific processing unit 290 of the data processing device 12 in the first embodiment is realized by the following means.

[0286] In this invention, the server includes a means for accessing the Internet or a database to acquire music information, a means for analyzing information and questions provided by the user to identify appropriate music and artists, and a means for transmitting information on the identified music and artists to the terminal and providing it to the user through a display or audio output, thereby making it possible to quickly and accurately collect, analyze, and provide the music information desired by the user.

[0287] The "Internet" is a collection of globally connected computer networks and an infrastructure for the exchange of information and communication.

[0288] A "database" is a systematically organized collection of information that can be efficiently accessed, searched, and managed through a specific application or query.

[0289] "User" refers to an individual or group that uses this system to obtain music information.

[0290] A "terminal" is a device from which a user receives information, and includes a personal computer, a smartphone, a tablet, etc.

[0291] "Display" refers to a screen or monitor for visually displaying information.

[0292] "Audio output" refers to technologies and devices that provide information to a user as audio.

[0293] "Music information" refers to all data and knowledge about music, such as music history, artist information, song details, and genre characteristics.

[0294] "Analysis" refers to the process of examining given information or data in detail to understand its structure and meaning.

[0295] "Acquisition" refers to the act of collecting information and making it available in a usable form.

[0296] "Providing" refers to the act of supplying information to a user and making it available for use.

[0297] A "network" refers to an infrastructure that connects multiple computers and devices and enables the exchange of information.

[0298] The present invention relates to a system for quickly and accurately providing music information desired by a user. Specific embodiments will be described below.

[0299] How to get music information The server uses the Internet and databases to collect information about music, utilizing existing web services such as Google Search API and music database API. For example, if a user asks the question "Tell me about rock music from the 60s," the server first uses the Google search API to gather the necessary information from the Internet. It also uses the music database API to obtain information about popular artists and songs from the 60s. This data is stored in the server's internal database and used for subsequent processing.

[0300] A means of analyzing user preferences and questions The server uses natural language processing techniques to analyze the questions and information entered by users, specifically using high-performance generative AI models such as OpenAI's ChatGPT model. If a user types in "Tell me about rock music from the 60s," the server will analyze the content and extract related keywords such as "60s," "rock music," etc. The results of this analysis will become the basis for the server to collect appropriate information and provide it to the user. Here is an example of a prompt for parsing: Prompt: "A user has asked the question 'Rock music from the 60s.' Please analyze the information related to this question."

[0301] A way to provide users with music and artist information relevant to their preferences and questions In order to provide the collected and analyzed information to the user, the server first converts the data into an appropriate format, and then transmits the data to the user's terminal. The terminal displays the information received from the server on the display and outputs voice as necessary. Specifically, the terminal can use a Text-to-Speech (TTS) engine to play back text information as voice. For example, if a user asks "Tell me about rock music from the 60's," the server might provide text like this: "In 60's rock music, Singer A, Singer B, etc. are representative artists. Also, the XX Music Festival in 1969 is very famous." The device can both display this text information and play it aloud using a TTS engine. Below are some examples of informational prompts: Prompt: "Provide the user with the information you have retrieved based on the user's question. Provide text on the display and speech output."

[0302] The flow of the identification process in the first embodiment will be described with reference to FIG.

[0303] Step 1: The server receives input from the user. Input: A user question (e.g., "What is rock music from the 60's?") Specific operation: The user enters a question into the terminal, which then sends this information to the server. Output: The server receives the user's question.

[0304] Step 2: The server uses a generative AI model to analyze the user's question. Input: A user question (e.g., "What is rock music from the 60's?") How it works: The server uses natural language processing techniques, such as OpenAI's ChatGPT model, to analyze the question, understand the intent of the question, and extract relevant keywords ("60s", "rock music"). Output: A list of keywords resulting from the analysis ("60s", "rock music").

[0305] Step 3: The server accesses the Internet and databases to retrieve relevant information. Input: Analyzed keyword list ("60s", "Rock music") How it works: The server uses the Google Search API to execute queries such as "60s rock music history" and collects the resulting webpage data. It also uses a music database API to retrieve information about popular artists and songs from the 60s. Output: Music information data (e.g. representative artists, popular songs, social background, etc.).

[0306] Step 4: The server organizes the retrieved information and converts it into a format for presentation to the user. Input: Retrieved music information data Specific operation: The server organizes the text data and organizes it into a format that is easy for the user to understand. For example, it may classify the acquired information into categories such as "famous artists," "popular songs," and "XX Music Festival in 1969." Output: Formatted information (e.g. "60s rock music is represented by artists Singer A, Singer B, etc. Also, the XX music festival in 1969 is very famous.").

[0307] Step 5: The server sends the formatted information to the user's terminal. Input: Formatted information Specific operation: The server sends information using the network address of the terminal. Output: The terminal receives the formatted information.

[0308] Step 6: The terminal displays the received information on a display and outputs audio as necessary. Input: Formatted information What it does: The device displays text information on a display and can also provide information aloud using a Text-to-Speech (TTS) engine. Output: The user receives information visually and audibly.

[0309] (Application example 1) Next, a description will be given of Application Example 1. In the following description, the data processing device 12 will be referred to as a "server" and the headset type terminal 314 will be referred to as a "terminal."

[0310] Conventional music recommendation systems have limited functionality in providing appropriate music information based on users' questions and preferences. In particular, responding to specific user questions and diverse music genres is a challenge, and effective solutions to improve user experience are required. Furthermore, advanced functions for quickly and accurately analyzing user-entered questions and providing relevant music information are lacking.

[0311] The specific processing by the specific processing unit 290 of the data processing device 12 in the application example 1 is realized by the following means.

[0312] In this invention, the server includes a means for acquiring music information, a means for analyzing a user's preferences and questions, a means for providing music-related information related to the user's preferences and questions, a means for analyzing the user's questions using a generative AI model, and a means for generating a prompt sentence, which makes it possible to quickly and accurately analyze a specific question from a user and provide appropriate music information.

[0313] "Means for acquiring music information" refers to means for accessing the Internet or a music database to gather music-related information.

[0314] The "means for analyzing user preferences and questions" refers to a means for analyzing information provided by a user to understand the contents of the preferences and questions.

[0315] The "means for providing music-related information related to the user's preferences and questions" refers to a means for providing music-related information that the user is interested in based on collected information through a display device or audio output device.

[0316] "Means for analyzing a user's question using a generative AI model" refers to means for analyzing the content of a user's question using an AI model that employs natural language processing technology and identifying related information.

[0317] A "means for generating a prompt sentence" is a means for automatically creating text to generate an appropriate response to a user's question.

[0318] DETAILED DESCRIPTION OF THE PREFERRED EMBODIMENTS The present invention is configured to provide appropriate music-related information in response to specific user queries in a music recommendation system.

[0319] First, a user uses a smartphone application. When the user inputs a question, the question is sent to the application server. The server first accesses the Internet or a music database using a music information acquisition means to collect related music information.

[0320] The server then analyzes the user's question using a generative AI model, including the Transformer model, which is based on natural language processing techniques, that analyzes the text entered by the user and identifies relevant music genres and artists based on the content.

[0321] After obtaining the parsed result of the question, the server generates a prompt sentence, which is used to formulate an appropriate response to the user's question.

[0322] Finally, the server collects music information and provides the information to the user through a display device and audio output device using a means for providing music-related information related to the user's preferences and questions. As a specific example, information such as characteristics of music genres, representative artists, and popular songs can be displayed on the display and the same information can be reproduced by audio.

[0323] As a specific operational procedure of the embodiment, consider a case where a user asks the question, "Tell me about rock music from the 60's." In response to this question, the system operates as follows.

[0324] 1. The server accesses the Internet and a music database to retrieve information about 60's rock music. 2. Use a generative AI model to parse the user question and identify information related to 60s rock music. 3. Generate prompts and formulate appropriate responses to the user's questions. 4. Provide the collected information through a display or audio output device.

[0325] The following are some examples of prompt sentences: "Representative artists of 60s rock music include bands and popular singers. Representative songs include many hits from that time. For example, if you ask, 'What are the characteristics of rock music from the 60s and what are the famous bands?' we can provide you with this information."

[0326] In this way, the system of the present invention is able to quickly and accurately analyze specific questions from users and provide appropriate music information.

[0327] The flow of the specific process in the application example 1 will be described with reference to FIG.

[0328] Step 1: A user uses a smartphone application to input a specific question (e.g., "Tell me about rock music from the 60s"). The input text data is sent by the application to a server. The input here is the user's question text, and the output is the text data sent to the server.

[0329] Step 2: The server accesses the Internet and music databases to acquire music information. This is done using an acquisition means. In this step, the input is a search query based on the user's question, and the output is the acquired music-related information (e.g., data about rock music in the 60s). The server collects the required information from APIs and databases on the Internet and temporarily stores the information.

[0330] Step 3: The server uses a generative AI model to parse the user's question. During this parsing step, the user's input text is analyzed using natural language processing techniques to identify relevant music genres and artists. The input is the user's question text, and the output is the parsed theme or category (60's rock music in this example). The server uses the results of the parsing to extract information about a specific music genre.

[0331] Step 4: The server generates a prompt sentence in response to the user's question. In this step, an appropriate response text is created based on the analysis results. The input is the analysis results and collected music information, and the output is a completed response sentence to be provided to the user. The specific generation operation is a text generation process using a generative AI model.

[0332] Step 5: The server transmits the collected and generated information to the terminal using a means for providing music information. The terminal provides this information to the user through a display device or an audio output device. The input is the information transmitted from the server, and the output is a visual or audio presentation to the user. Specifically, text information is displayed on a display or audio information is played from a speaker.

[0333] Step 6: The user can then review the information provided and ask new questions or enter additional requests. At this step, the user's input is again sent to the server and the process repeats: the input is the user's new question or request, and the output is the updated information provided process.

[0334] In addition, an emotion engine that estimates the user's emotion may be further combined. That is, the identification processing unit 290 may estimate the user's emotion using the emotion identification model 59, and perform identification processing using the user's emotion.

[0335] This embodiment is composed of the following elements: 1. Incorporating an emotion engine: - The system will incorporate an emotion engine to recognize the user's emotions. - The emotion engine recognizes emotions from the user's voice, facial expressions, text, etc., and customizes music and artist suggestions based on the recognition results. 2. Providing a music experience that responds to the user's emotions: - The system analyzes the user's emotional state through an emotion engine. - Based on the analysis results, the system will adjust music selection and the timing of information provided to provide the user with a more personalized music experience. As a specific example, an embodiment will be described in which a user makes a request to the system, "Tell me what song to listen to when I'm feeling sad." 1. Incorporating an emotion engine: - The system incorporates an emotion engine to recognize emotions from the user's voice, facial expressions, text, etc. - Analyze the user's tone of voice, changes in facial expressions, and the context of the text entered to understand the user's emotional state. 2. Providing a music experience that responds to the user's emotions: - The system analyzes the user's emotional state through an emotion engine. If a user requests "What songs do you listen to when you feel sad?" and the analyzed emotional state of the user is "sad", then the emotion engine will recognize the user's sad mood. - The system will select songs and artists that fit a sad mood and suggest them on the display or through audio output. - The emotion engine also monitors changes in the user's emotional state and adjusts music selection and information provision timing as needed.

[0336] The process flow will be explained below.

[0337] Step 1: Recognize the user's emotions. - The device collects information such as the user's voice, facial expressions, and text entered by the user. - The device transmits the collected information to the server. Step 2: Sentiment analysis by the sentiment engine. - The server inputs the received information into the emotion engine. - The emotion engine analyzes voice tone, facial expression changes, input text context, etc. to recognize the user's emotional state. Step 3: Select and suggest music according to emotions. - The server selects appropriate music based on the user's emotional state recognized by the emotion engine. - The server sends information about the selected music and artists to the device. Step 4: Deliver the music experience. - The device will display information about the received music and artists on the display and / or provide suggestions through audio output. - Users can express their emotions and feel comforted by the music provided.

[0338] Example 2

[0339] Next, a description will be given of Example 2. In the following description, the data processing device 12 will be referred to as a "server" and the headset type terminal 314 will be referred to as a "terminal".

[0340] Conventional music provision systems provide music based on the user's preferences and questions, but they cannot take the user's emotions into account. This makes it difficult to provide music that is optimal for the user's current emotional state. In addition, because it is not possible to dynamically suggest music according to emotions, there is a problem in that it is not possible to increase user satisfaction.

[0341] The specific processing by the specific processing unit 290 of the data processing device 12 in the second embodiment is realized by the following means.

[0342] In this invention, the server includes a means for recognizing a user's emotion, a means for acquiring music information suitable for the user based on the recognized emotion, and a means for providing the acquired music information to the user, thereby making it possible to dynamically suggest music optimal to the user's emotional state.

[0343] "Means for recognizing the user's emotions" refers to technology that analyzes data such as the user's voice, facial expressions, and writing to determine the user's current emotional state.

[0344] "Means for acquiring music information appropriate for the user based on recognized emotions" refers to a technology that utilizes the results of emotion recognition to acquire information on music and artists that match the user's current emotional state from a music information database, etc.

[0345] The "means for providing the acquired music information to the user" refers to a technique for conveying the acquired music and artist information to the user through a display or audio output.

[0346] "Means for accessing the Internet or a music information database" refers to technology for connecting to a music information database on the Internet or locally or on the cloud, and searching for and acquiring music information.

[0347] "Means for providing information through a display or audio output" refers to devices or technologies for conveying acquired information to a user visually or audibly.

[0348] In order to implement the present invention, the following elements must be specifically configured and operated.

[0349] 1. Incorporating an emotion engine The server is equipped with an emotion engine to recognize the user's emotions. The emotion engine uses general emotion recognition technologies such as the Microsoft Azure emotion recognition service. A user uses the device's microphone and camera to provide voice input, text input, facial expression data, and the like.

[0350] 2. User Emotion Recognition When a user speaks into the terminal, the terminal transmits the voice data to the server, and if necessary, facial expression data captured by the camera is also transmitted to the server. The server inputs the received data into the emotion engine to recognize the emotion. For example, if a user requests, "What song do you listen to when you feel sad?", the emotion engine recognizes the emotion "sad."

[0351] 3. Acquiring music information according to user emotions Based on the recognized emotion, the server accesses a music information database or a streaming service to obtain music information suitable for the user. Specifically, the server uses a general music information acquisition means such as a music database API. The server searches for "songs that fit a sad mood" and retrieves the results.

[0352] 4. Providing music information After the music information is obtained, the server transmits the information to the user's terminal. The terminal displays the received music information on the display and, if necessary, suggests the information to the user through audio output.

[0353] For example, when a user says to the terminal, "Please tell me what song you listen to when you feel sad," the specific operation is as follows.

[0354] 1. The user provides voice input. 2. The device sends the voice data to the server. 3. The server's emotion engine recognizes the emotion "sad." 4. The server uses a music database API to search for "songs that fit a sad mood." 5. The server sends the acquired music information to the terminal. 6. The device will show the results on the screen and also provide audio suggestions.

[0355] Examples of prompt statements "Generate a list of music suitable for when the user is sad." "Describe how you would adjust your musical suggestions if your emotional state changes."

[0356] This invention makes it possible to dynamically suggest music that is optimal for the emotional state of a user. Specifically, by using an emotion engine to recognize the user's emotion and providing optimal music information based on that emotion, it is possible to increase the user's satisfaction.

[0357] The flow of the identification process in the second embodiment will be described with reference to FIG.

[0358] Step 1: The user performs voice input. The user speaks into the device's microphone and says, "Please tell me what songs you listen to when you're feeling sad." This audio data is captured by the terminal. Input: User's voice data Output: Captured audio data

[0359] Step 2: The terminal transmits the voice data to the server. The device uses a network connection to transmit the captured audio data to a server. Input: Captured audio data Output: Audio data sent to the server

[0360] Step 3: The server inputs the voice data into an emotion engine to recognize emotions. The server passes the received voice data to emotion recognition software (for example, Microsoft Azure's emotion recognition service). The emotion engine analyzes the tone and content of the voice and recognizes the user's emotion as "sad." Input: Transmitted audio data Output: The recognized emotion (e.g. "sad")

[0361] Step 4: The server obtains music information based on the emotion. Based on the emotion recognition result, the server calls the music database API and sends a request to obtain related music information. Specifically, a query is sent to the API to search for "songs that fit a sad mood." Input: A recognized emotion (e.g. "sad") Output: Retrieved music information (song list)

[0362] Step 5: The server transmits the acquired music information to the terminal. The server sends the retrieved music information back to the device, including song titles and artist names in a list format. Input: Acquired music information (song list) Output: Music information sent to the device

[0363] Step 6: The device will display music information and provide audio suggestions. The terminal displays the received song list on the user's display. Additionally, it uses text-to-speech functionality to provide audio guidance to the user, such as, "Here's a list of songs that suit a sad mood." Input: Music information sent to the device Output: Display and audio suggestions

[0364] Step 7: The server continuously monitors the user's emotional state. The server periodically receives new voice and facial expression data from the terminal and re-evaluates the user's emotions using the emotion engine. If the user's emotional state changes, new music information is obtained and re-suggestions are made as appropriate. Input: Voice and facial expression data sent periodically Output: New emotion recognition results and re-suggested music information

[0365] These processing steps allow the system to dynamically provide music selections that best suit the user's emotions.

[0366] (Application example 2) Next, a description will be given of Application Example 2. In the following description, the data processing device 12 will be referred to as a "server", and the headset type terminal 314 will be referred to as a "terminal".

[0367] Conventional music distribution systems have had difficulty providing a personalized music experience that accurately reflects the user's feelings and emotions. Although there are systems that suggest music based on the user's preferences and questions, they lack the functionality to respond to the user's real-time emotional changes, and a deeper level of personalization is required.

[0368] The specific processing by the specific processing unit 290 of the data processing device 12 in the application example 2 is realized by the following means.

[0369] In this invention, the server includes a means for acquiring music information, a means for analyzing the user's emotions using an emotion recognition engine, a means for providing information on music and artists related to the user's preferences and emotions, and a means for adjusting the timing of music selection and information provision, thereby making it possible to provide a personalized music experience according to the user's emotional state.

[0370] The "means for acquiring music information" refers to a means for acquiring music information from a music database, the Internet, or the like.

[0371] An "emotion recognition engine" is software or hardware that analyzes a user's voice, facial expressions, text, etc., and recognizes the user's emotions.

[0372] "Means for providing information on music and artists related to the user's preferences and emotions" refers to means for providing information on music and artists based on the user's emotional state and preferences, and includes a display and audio output.

[0373] The "means for adjusting the timing of music selection and information provision" refers to a means for selecting music and providing information at optimal timing according to the user's emotional state.

[0374] The "means for monitoring changes in the user's emotional state" refers to a means for monitoring the user's emotional state in real time and adjusting the music experience in response to those changes.

[0375] In order to put the present invention into practice, it is first necessary to construct a system that includes an emotion recognition engine, a means for acquiring music information, and a means for selecting music and providing information. This system operates as follows.

[0376] The server has a means of accessing the Internet and a music information database to obtain music information. This means allows the server to always obtain the latest music information from the database. For example, the server can use the API of a well-known music streaming service.

[0377] An emotion recognition engine is implemented on the device and analyzes the user's voice, facial expression, text, etc. This makes it possible to recognize the user's emotions in real time. For example, a voice analysis API or facial expression recognition API is used for the emotion recognition engine. If the user inputs "I'm very sad today," the emotion recognition engine will recognize the emotion "sad."

[0378] Based on the recognized emotion, the server uses the music streaming service API to obtain information on music and artists that are suitable for the user's emotion. For example, it obtains a playlist that matches the "sad" mood. It then provides the information to the user via the device's display or audio output. For example, it may display "Recommended songs: ●● - ●●" on the display or notify the user by voice, "Here are some recommended songs."

[0379] In addition, the system monitors changes in the user's emotional state and adjusts music selection and information provision timing as necessary. Even if the user's emotion changes from "sad" to "happy," the system will detect this and suggest music appropriate to the new emotion. This allows the user to always enjoy a music experience that matches their emotion.

[0380] Specifically, the following prompt sentences are input to the generative AI model to perform emotion recognition: If a user types "I feel so sad today", an emotion recognition API will return the emotion "sad". Write a program that uses a music streaming service API to get a playlist that matches the "sad" mood and serves it to the user.

[0381] The system can provide a personalized music experience based on the user's emotional state, achieving a deeper level of personalization.

[0382] The flow of the specific process in the application example 2 will be described with reference to FIG.

[0383] Step 1: User emotion input The user inputs their mood or emotion into the device. For example, they can say, "I feel very sad today," or input it as text. This input becomes data for the next analysis.

[0384] Step 2: Emotion recognition The device uses the emotion recognition API to analyze the user's input data. Specifically, in the case of voice input, voice analysis is performed, and in the case of text input, sentence analysis is performed. Through this analysis, the user's emotion is recognized as "sad."

[0385] Step 3: Emotion-Based Music Selection Based on the recognized emotion, the server uses the music streaming service API to search for music that matches the user's emotion. For example, if the emotion is "sad", it retrieves a playlist that matches the "sad" mood. At this stage, the retrieved data includes the song title, artist name, etc.

[0386] Step 4: Providing music information The music information obtained from the server is provided to the user via the device's display or audio output. Specifically, the device may show "Recommended songs: XX - XX" on the display or notify the user by voice, "Here are some recommended songs."

[0387] Step 5: Monitoring user emotional changes The system uses the device's emotion recognition engine to monitor the user's emotional state in real time, for example, re-recognizing emotions each time the user enters new voice or text input and determining whether the emotional state has changed.

[0388] Step 6: Timing adjustment of music provision If the user's emotion is recognized as having changed, the server again uses the music streaming service API to obtain music based on the new emotion (i.e., the changed emotion) and provides it to the user. For example, if the emotion changes from "sad" to "happy," the server obtains and provides a playlist that matches the "happy" mood.

[0389] This allows users to always enjoy a music experience that matches their emotions.

[0390] The specific processing unit 290 transmits the result of the specific processing to the headset type terminal 314. In the headset type terminal 314, the control unit 46A causes the speaker 240 and the display 343 to output the result of the specific processing. The microphone 238 acquires voice indicating a user input for the result of the specific processing. The control unit 46A transmits voice data indicating the user input acquired by the microphone 238 to the data processing device 12. In the data processing device 12, the specific processing unit 290 acquires the voice data.

[0391] The data generation model 58 is a so-called generative AI (Artificial Intelligence). An example of the data generation model 58 is ChatGPT (Internet search<URL: https: / / openai.com / blog / chatgpt> ), Gemini (Internet search <url: https: gemini.google.com ?hl="ja">) and other generation AIs. The data generation model 58 is obtained by making a neural network perform deep learning. A prompt including an instruction is input to the data generation model 58, and inference data such as voice data indicating a voice, text data indicating a text, and image data indicating an image is input. The data generation model 58 performs inference on the input inference data according to the instruction indicated by the prompt, and outputs the inference result in a data format such as voice data and text data. Here, inference refers to, for example, analysis, classification, prediction, and / or summarization.

[0392] In the above embodiment, an example was given in which the specific processing is performed by the data processing device 12, but the technology of the present disclosure is not limited to this, and the specific processing may be performed by the headset type terminal 314.

[0393] [Fourth embodiment] FIG. 7 shows an example of the configuration of a data processing system 410 according to the fourth embodiment.

[0394] 7, a data processing system 410 includes a data processing device 12 and a robot 414. An example of the data processing device 12 is a server.

[0395] The data processing device 12 includes a computer 22, a database 24, and a communication I / F 26. The computer 22 is an example of a "computer" according to the technology of the present disclosure. The computer 22 includes a processor 28, a RAM 30, and a storage 32. The processor 28, the RAM 30, and the storage 32 are connected to a bus 34. The database 24 and the communication I / F 26 are also connected to the bus 34. The communication I / F 26 is connected to a network 54. Examples of the network 54 include a wide area network (WAN) and / or a local area network (LAN).

[0396] The robot 414 includes a computer 36, a microphone 238, a speaker 240, a camera 42, a communication I / F 44, and a control target 443. The computer 36 includes a processor 46, a RAM 48, and a storage 50. The processor 46, the RAM 48, and the storage 50 are connected to a bus 52. In addition, the microphone 238, the speaker 240, the camera 42, and the control target 443 are also connected to the bus 52.

[0397] The microphone 238 receives instructions and the like from the user 20 by receiving voice uttered by the user 20. The microphone 238 captures the voice uttered by the user 20, converts the captured voice into audio data, and outputs it to the processor 46. The speaker 240 outputs the voice according to instructions from the processor 46.

[0398] Camera 42 is a small digital camera equipped with an optical system including a lens, an aperture, and a shutter, and an imaging element such as a CMOS (Complementary Metal-Oxide-Semiconductor) image sensor or a CCD (Charge Coupled Device) image sensor, and captures images of the surroundings of user 20 (e.g., an imaging range defined by an angle of view equivalent to the width of the field of vision of an average healthy person).

[0399] The communication I / F 44 is connected to a network 54. The communication I / Fs 44 and 26 are responsible for transmitting and receiving various types of information between the processor 46 and the processor 28 via the network 54. The transmission and reception of various types of information between the processor 46 and the processor 28 using the communication I / Fs 44 and 26 is performed in a secure state.

[0400] The control target 443 includes a display device, LEDs in the eyes, and motors for driving the arms, hands, legs, etc. The posture and behavior of the robot 414 are controlled by controlling the motors of the arms, hands, legs, etc. Some of the emotions of the robot 414 can be expressed by controlling these motors. In addition, the facial expressions of the robot 414 can also be expressed by controlling the light emission state of the LEDs in the eyes of the robot 414.

[0401] Fig. 8 shows an example of main functions of the data processing device 12 and the robot 414. As shown in Fig. 8, in the data processing device 12, a specific process is performed by the processor 28. A specific process program 56 is stored in the storage 32.

[0402] The specific processing program 56 is an example of a "program" according to the technology of the present disclosure. The processor 28 reads the specific processing program 56 from the storage 32, and executes the read specific processing program 56 on the RAM 30. The specific processing is realized by the processor 28 operating as a specific processing unit 290 in accordance with the specific processing program 56 executed on the RAM 30.

[0403] The storage 32 stores a data generation model 58 and an emotion identification model 59. The data generation model 58 and the emotion identification model 59 are used by the identification processing unit 290.

[0404] In the robot 414, the reception output process is performed by the processor 46. A reception output program 60 is stored in the storage 50. The processor 46 reads the reception output program 60 from the storage 50, and executes the read reception output program 60 on the RAM 48. The reception output process is realized by the processor 46 operating as the control unit 46A in accordance with the reception output program 60 executed on the RAM 48.

[0405] Next, a description will be given of the specific processing by the specific processing unit 290 of the data processing device 12. In the following description, the data processing device 12 will be referred to as a "server" and the robot 414 will be referred to as a "terminal".

[0406] An embodiment for implementing the disclosed technology comprises the following elements. 1. How to get music information: The server obtains information about music history and artists by accessing the Internet and a music information database. 2. Means of analyzing user preferences and questions: - The server analyzes the information provided by the user and selects appropriate music and artists based on the user's preferences and questions. The server may also select appropriate music and artists based on the user's age, gender, and usual music listening habits. 3. Means of providing information about music and artists related to the user's preferences and questions: The server transmits the acquired information to the terminal, and the terminal provides information on music and artists related to the user's preferences and questions through a display and / or audio output. As a specific example, an embodiment will be described in which a user asks the question "Tell me about rock music from the 60's." 1. How to get music information: - The server accesses the Internet and music information databases to obtain information about rock music from the 1960s, such as representative bands and artists, popular songs, and the social context of the time. 2. Means of analyzing user preferences and questions: - The server parses the user's question, "What is rock music from the 60's?" and identifies information relevant to the question. 3. Means of providing information about music and artists related to the user's preferences and questions: The server transmits the acquired information to the terminal, and the terminal displays the characteristics, representative artists, popular songs, etc. of rock music in the 60s on the display. The terminal may also provide the characteristics, representative artists, popular songs, etc. of rock music in the 60s to the user through audio output.

[0407] The process flow will be explained below.

[0408] Step 1: Receive user input such as questions and preferences. - The user communicates questions and preferences to the system through the terminal. - The terminal sends the user's input to the server. Step 2: Search and retrieve information based on your questions and preferences. - The server receives and parses the user input. - The server accesses the Internet and music information databases to search for music and artist information related to the user's question and preferences. - The server organizes the obtained information and extracts the necessary information. Step 3: Provide information. - The server sends the acquired information to the terminal. - The device displays information on a display and conveys information through audio output. - Users can enjoy learning from the information provided.

[0409] Example 1 Next, a description will be given of Example 1. In the following description, the data processing device 12 is referred to as a "server" and the robot 414 is referred to as a "terminal."

[0410] Conventional music information systems have difficulty in quickly and accurately responding to the diverse tastes and questions of users. In addition, few systems can provide a consistent process for efficiently collecting, analyzing, and presenting the information users want. As a result, the user experience is often poor and dissatisfied. In addition, when users want detailed information related to a specific era or genre, they have to search manually or refer to multiple sources, which is time-consuming and laborious.

[0411] The specific processing by the specific processing unit 290 of the data processing device 12 in the first embodiment is realized by the following means.

[0412] In this invention, the server includes a means for accessing the Internet or a database to acquire music information, a means for analyzing information and questions provided by the user to identify appropriate music and artists, and a means for transmitting information on the identified music and artists to the terminal and providing it to the user through a display or audio output, thereby making it possible to quickly and accurately collect, analyze, and provide the music information desired by the user.

[0413] The "Internet" is a collection of globally connected computer networks and an infrastructure for the exchange of information and communication.

[0414] A "database" is a systematically organized collection of information that can be efficiently accessed, searched, and managed through a specific application or query.

[0415] "User" refers to an individual or group that uses this system to obtain music information.

[0416] A "terminal" is a device from which a user receives information, and includes a personal computer, a smartphone, a tablet, etc.

[0417] "Display" refers to a screen or monitor for visually displaying information.

[0418] "Audio output" refers to technologies and devices that provide information to a user as audio.

[0419] "Music information" refers to all data and knowledge about music, such as music history, artist information, song details, and genre characteristics.

[0420] "Analysis" refers to the process of examining given information or data in detail to understand its structure and meaning.

[0421] "Acquisition" refers to the act of collecting information and making it available in a usable form.

[0422] "Providing" refers to the act of supplying information to a user and making it available for use.

[0423] A "network" refers to an infrastructure that connects multiple computers and devices and enables the exchange of information.

[0424] The present invention relates to a system for quickly and accurately providing music information desired by a user. Specific embodiments will be described below.

[0425] How to get music information The server uses the Internet and databases to collect information about music, utilizing existing web services such as Google Search API and music database API. For example, if a user asks the question "Tell me about rock music from the 60s," the server first uses the Google search API to gather the necessary information from the Internet. It also uses the music database API to obtain information about popular artists and songs from the 60s. This data is stored in the server's internal database and used for subsequent processing.

[0426] A means of analyzing user preferences and questions The server uses natural language processing techniques to analyze the questions and information entered by users, specifically using high-performance generative AI models such as OpenAI's ChatGPT model. If a user types in "Tell me about rock music from the 60s," the server will analyze the content and extract related keywords such as "60s," "rock music," etc. The results of this analysis will become the basis for the server to collect appropriate information and provide it to the user.

[0427] Here is an example of a prompt for parsing: Prompt: "A user has asked the question 'Rock music from the 60s.' Please analyze the information related to this question."

[0428] A way to provide users with music and artist information relevant to their preferences and questions In order to provide the collected and analyzed information to the user, the server first converts the data into an appropriate format, and then transmits the data to the user's terminal. The terminal displays the information received from the server on the display and outputs voice as necessary. Specifically, the terminal can use a Text-to-Speech (TTS) engine to play back text information as voice. For example, if a user asks "Tell me about rock music from the 60's," the server might provide text like this: "In 60's rock music, Singer A, Singer B, etc. are representative artists. Also, the XX Music Festival in 1969 is very famous." The device can both display this text information and play it aloud using a TTS engine.

[0429] Below are some examples of informational prompts: Prompt: "Provide the user with the information you have retrieved based on the user's question. Provide text on the display and speech output."

[0430] The flow of the identification process in the first embodiment will be described with reference to FIG.

[0431] Step 1: The server receives input from the user. Input: A user question (e.g., "What is rock music from the 60's?") Specific operation: The user enters a question into the terminal, which then sends this information to the server. Output: The server receives the user's question.

[0432] Step 2: The server uses a generative AI model to analyze the user's question. Input: A user question (e.g., "What is rock music from the 60's?") How it works: The server uses natural language processing techniques, such as OpenAI's ChatGPT model, to analyze the question, understand the intent of the question, and extract relevant keywords ("60s", "rock music"). Output: A list of keywords resulting from the analysis ("60s", "rock music").

[0433] Step 3: The server accesses the Internet and databases to retrieve relevant information. Input: Analyzed keyword list ("60s", "Rock music") How it works: The server uses the Google Search API to execute queries such as "60s rock music history" and collects the resulting webpage data. It also uses a music database API to retrieve information about popular artists and songs from the 60s. Output: Music information data (e.g. representative artists, popular songs, social background, etc.).

[0434] Step 4: The server organizes the retrieved information and converts it into a format for presentation to the user. Input: Retrieved music information data Specific operation: The server organizes the text data and composes it in a format that is easy for the user to understand. For example, it may classify the acquired information into categories such as "famous artists," "popular songs," and "XX Music Festival in 1969." Output: Formatted information (e.g. "60s rock music is best known for Singer A, Singer B, etc. Also, the XX music festival in 1969 is very famous.").

[0435] Step 5: The server sends the formatted information to the user's terminal. Input: Formatted information Specific operation: The server sends information using the network address of the terminal. Output: The terminal receives the formatted information.

[0436] Step 6: The terminal displays the received information on a display and outputs audio as necessary. Input: Formatted information What it does: The device displays text information on a display and can also provide information aloud using a Text-to-Speech (TTS) engine. Output: The user receives information visually and audibly.

[0437] (Application example 1) Next, a description will be given of Application Example 1. In the following description, the data processing device 12 will be referred to as a "server" and the robot 414 will be referred to as a "terminal."

[0438] Conventional music recommendation systems have limited functionality in providing appropriate music information based on users' questions and preferences. In particular, responding to specific user questions and diverse music genres is a challenge, and effective solutions to improve user experience are required. Furthermore, advanced functions for quickly and accurately analyzing user-entered questions and providing relevant music information are lacking.

[0439] The specific processing by the specific processing unit 290 of the data processing device 12 in the application example 1 is realized by the following means.

[0440] In this invention, the server includes a means for acquiring music information, a means for analyzing a user's preferences and questions, a means for providing music-related information related to the user's preferences and questions, a means for analyzing the user's questions using a generative AI model, and a means for generating a prompt sentence, which makes it possible to quickly and accurately analyze a specific question from a user and provide appropriate music information.

[0441] "Means for acquiring music information" refers to means for accessing the Internet or a music database to gather music-related information.

[0442] The "means for analyzing user preferences and questions" refers to a means for analyzing information provided by a user to understand the contents of the preferences and questions.

[0443] The "means for providing music-related information related to the user's preferences and questions" refers to a means for providing music-related information that the user is interested in based on collected information through a display device or audio output device.

[0444] "Means for analyzing a user's question using a generative AI model" refers to means for analyzing the content of a user's question using an AI model that employs natural language processing technology and identifying related information.

[0445] A "means for generating a prompt sentence" is a means for automatically creating text to generate an appropriate response to a user's question.

[0446] DETAILED DESCRIPTION OF THE PREFERRED EMBODIMENTS The present invention is configured to provide appropriate music-related information in response to specific user queries in a music recommendation system.

[0447] First, a user uses a smartphone application. When the user inputs a question, the question is sent to the application server. The server first accesses the Internet or a music database using a music information acquisition means to collect related music information.

[0448] The server then analyzes the user's question using a generative AI model, including the Transformer model, which is based on natural language processing techniques, that analyzes the text entered by the user and identifies relevant music genres and artists based on the content.

[0449] After obtaining the parsed result of the question, the server generates a prompt sentence, which is used to formulate an appropriate response to the user's question.

[0450] Finally, the server collects music information and provides the information to the user through a display device and audio output device using a means for providing music-related information related to the user's preferences and questions. As a specific example, information such as characteristics of music genres, representative artists, and popular songs can be displayed on the display and the same information can be reproduced by audio.

[0451] As a specific operational procedure of the embodiment, consider a case where a user asks the question, "Tell me about rock music from the 60's." In response to this question, the system operates as follows.

[0452] 1. The server accesses the Internet and a music database to retrieve information about 60's rock music. 2. Use a generative AI model to parse the user question and identify information related to 60s rock music. 3. Generate prompts and formulate appropriate responses to the user's questions. 4. Provide the collected information through a display or audio output device.

[0453] The following are some examples of prompt sentences: "Representative artists of 60s rock music include bands and popular singers. Representative songs include many hits from that time. For example, if you ask, 'What are the characteristics of rock music from the 60s and what are the famous bands?' we can provide you with this information."

[0454] In this way, the system of the present invention is able to quickly and accurately analyze specific questions from users and provide appropriate music information.

[0455] The flow of the specific process in the application example 1 will be described with reference to FIG.

[0456] Step 1: A user uses a smartphone application to input a specific question (e.g., "Tell me about rock music from the 60s"). The input text data is sent by the application to a server. The input here is the user's question text, and the output is the text data sent to the server.

[0457] Step 2: The server accesses the Internet and music databases to acquire music information. This is done using an acquisition means. In this step, the input is a search query based on the user's question, and the output is the acquired music-related information (e.g., data about rock music in the 60s). The server collects the required information from APIs and databases on the Internet and temporarily stores the information.

[0458] Step 3: The server uses a generative AI model to parse the user's question. During this parsing step, the user's input text is analyzed using natural language processing techniques to identify relevant music genres and artists. The input is the user's question text, and the output is the parsed theme or category (60's rock music in this example). The server uses the results of the parsing to extract information about a specific music genre.

[0459] Step 4: The server generates a prompt sentence in response to the user's question. In this step, an appropriate response text is created based on the analysis results. The input is the analysis results and collected music information, and the output is a completed response sentence to be provided to the user. The specific generation operation is a text generation process using a generative AI model.

[0460] Step 5: The server transmits the collected and generated information to the terminal using a means for providing music information. The terminal provides this information to the user through a display device or an audio output device. The input is the information transmitted from the server, and the output is a visual or audio presentation to the user. Specifically, text information is displayed on a display or audio information is played from a speaker.

[0461] Step 6: The user can then review the information provided and ask new questions or enter additional requests. At this step, the user's input is again sent to the server and the process repeats: the input is the user's new question or request, and the output is the updated information provided process.

[0462] In addition, an emotion engine that estimates the user's emotion may be further combined. That is, the identification processing unit 290 may estimate the user's emotion using the emotion identification model 59, and perform identification processing using the user's emotion.

[0463] This embodiment is composed of the following elements: 1. Incorporating an emotion engine: - The system will incorporate an emotion engine to recognize the user's emotions. - The emotion engine recognizes emotions from the user's voice, facial expressions, text, etc., and customizes music and artist suggestions based on the recognition results. 2. Providing a music experience that responds to the user's emotions: - The system analyzes the user's emotional state through an emotion engine. - Based on the analysis results, the system will adjust music selection and the timing of information provided to provide the user with a more personalized music experience. As a specific example, an embodiment will be described in which a user makes a request to the system, "Tell me what song to listen to when I'm feeling sad." 1. Incorporating an emotion engine: - The system incorporates an emotion engine to recognize emotions from the user's voice, facial expressions, text, etc. - Analyze the user's tone of voice, changes in facial expressions, and the context of the text entered to understand the user's emotional state. 2. Providing a music experience that responds to the user's emotions: - The system analyzes the user's emotional state through an emotion engine. If a user requests "What songs do you listen to when you feel sad?" and the analyzed emotional state of the user is "sad", then the emotion engine will recognize the user's sad mood. - The system will select songs and artists that fit a sad mood and suggest them on the display or through audio output. - The emotion engine also monitors changes in the user's emotional state and adjusts music selection and information provision timing as needed.

[0464] The process flow will be explained below.

[0465] Step 1: Recognize the user's emotions. - The device collects information such as the user's voice, facial expressions, and text entered by the user. - The device transmits the collected information to the server. Step 2: Sentiment analysis by the sentiment engine. - The server inputs the received information into the emotion engine. - The emotion engine analyzes voice tone, facial expression changes, input text context, etc. to recognize the user's emotional state. Step 3: Select and suggest music according to emotions. - The server selects appropriate music based on the user's emotional state recognized by the emotion engine. - The server sends information about the selected music and artists to the device. Step 4: Deliver the music experience. - The device will display information about the received music and artists on the display and / or provide suggestions through audio output. - Users can express their emotions and feel comforted by the music provided.

[0466] Example 2 Next, a description will be given of Example 2. In the following description, the data processing device 12 is referred to as a "server" and the robot 414 is referred to as a "terminal."

[0467] Conventional music provision systems provide music based on the user's preferences and questions, but they cannot take the user's emotions into account. This makes it difficult to provide music that is optimal for the user's current emotional state. In addition, because it is not possible to dynamically suggest music according to emotions, there is a problem in that it is not possible to increase user satisfaction.

[0468] The specific processing by the specific processing unit 290 of the data processing device 12 in the second embodiment is realized by the following means.

[0469] In this invention, the server includes a means for recognizing a user's emotion, a means for acquiring music information suitable for the user based on the recognized emotion, and a means for providing the acquired music information to the user, thereby making it possible to dynamically suggest music optimal to the user's emotional state.

[0470] "Means for recognizing the user's emotions" refers to technology that analyzes data such as the user's voice, facial expressions, and writing to determine the user's current emotional state.

[0471] "Means for acquiring music information appropriate for the user based on recognized emotions" refers to a technology that utilizes the results of emotion recognition to acquire information on music and artists that match the user's current emotional state from a music information database, etc.

[0472] The "means for providing the acquired music information to the user" refers to a technique for conveying the acquired music and artist information to the user through a display or audio output.

[0473] "Means for accessing the Internet or a music information database" refers to technology for connecting to a music information database on the Internet or locally or on the cloud, and searching for and acquiring music information.

[0474] "Means for providing information through a display or audio output" refers to devices or technologies for conveying acquired information to a user visually or audibly.

[0475] In order to implement the present invention, the following elements must be specifically configured and operated.

[0476] 1. Incorporating an emotion engine The server is equipped with an emotion engine to recognize the user's emotions. The emotion engine uses general emotion recognition technologies such as the Microsoft Azure emotion recognition service. A user uses the device's microphone and camera to provide voice input, text input, facial expression data, and the like.

[0477] 2. User Emotion Recognition When a user speaks into the terminal, the terminal transmits the voice data to the server, and if necessary, facial expression data captured by the camera is also transmitted to the server. The server inputs the received data into the emotion engine to recognize the emotion. For example, if a user requests, "What song do you listen to when you feel sad?", the emotion engine recognizes the emotion "sad."

[0478] 3. Acquiring music information according to user emotions Based on the recognized emotion, the server accesses a music information database or a streaming service to obtain music information suitable for the user. Specifically, the server uses a general music information acquisition means such as a music database API. The server searches for "songs that fit a sad mood" and retrieves the results.

[0479] 4. Providing music information After the music information is obtained, the server transmits the information to the user's terminal. The terminal displays the received music information on the display and, if necessary, suggests the information to the user through audio output.

[0480] For example, when a user says to the terminal, "Please tell me what song you listen to when you feel sad," the specific operation is as follows.

[0481] 1. The user provides voice input. 2. The device sends the voice data to the server. 3. The server's emotion engine recognizes the emotion "sad." 4. The server uses a music database API to search for "songs that fit a sad mood." 5. The server sends the acquired music information to the terminal. 6. The device will show the results on the screen and also provide audio suggestions.

[0482] Examples of prompt statements "Generate a list of music suitable for when the user is sad." "Describe how you would adjust your musical suggestions if your emotional state changes."

[0483] This invention makes it possible to dynamically suggest music that is optimal for the emotional state of a user. Specifically, by using an emotion engine to recognize the user's emotion and providing optimal music information based on that emotion, it is possible to increase the user's satisfaction.

[0484] The flow of the identification process in the second embodiment will be described with reference to FIG.

[0485] Step 1: The user performs voice input. The user speaks into the device's microphone and says, "Please tell me what songs you listen to when you're feeling sad." This audio data is captured by the terminal. Input: User's voice data Output: Captured audio data

[0486] Step 2: The terminal transmits the voice data to the server. The device uses a network connection to transmit the captured audio data to a server. Input: Captured audio data Output: Audio data sent to the server

[0487] Step 3: The server inputs the voice data into an emotion engine to recognize emotions. The server passes the received voice data to emotion recognition software (for example, Microsoft Azure's emotion recognition service). The emotion engine analyzes the tone and content of the voice and recognizes the user's emotion as "sad." Input: Transmitted audio data Output: The recognized emotion (e.g. "sad")

[0488] Step 4: The server obtains music information based on the emotion. Based on the emotion recognition result, the server calls the music database API and sends a request to obtain related music information. Specifically, a query is sent to the API to search for "songs that fit a sad mood." Input: A recognized emotion (e.g. "sad") Output: Retrieved music information (song list)

[0489] Step 5: The server transmits the acquired music information to the terminal. The server sends the retrieved music information back to the device, including song titles and artist names in a list format. Input: Acquired music information (song list) Output: Music information sent to the device

[0490] Step 6: The device will display music information and provide audio suggestions. The terminal displays the received song list on the user's display. Additionally, it uses text-to-speech functionality to provide audio guidance to the user, such as, "Here's a list of songs that suit a sad mood." Input: Music information sent to the device Output: Display and audio suggestions

[0491] Step 7: The server continuously monitors the user's emotional state. The server periodically receives new voice and facial expression data from the terminal and re-evaluates the user's emotions using the emotion engine. If the user's emotional state changes, new music information is obtained and re-suggestions are made as appropriate. Input: Voice and facial expression data sent periodically Output: New emotion recognition results and re-suggested music information

[0492] These processing steps allow the system to dynamically provide music selections that best suit the user's emotions.

[0493] (Application example 2) Next, a description will be given of Application Example 2. In the following description, the data processing device 12 will be referred to as a "server" and the robot 414 will be referred to as a "terminal".

[0494] Conventional music distribution systems have had difficulty providing a personalized music experience that accurately reflects the user's feelings and emotions. Although there are systems that suggest music based on the user's preferences and questions, they lack the functionality to respond to the user's real-time emotional changes, and a deeper level of personalization is required.

[0495] The specific processing by the specific processing unit 290 of the data processing device 12 in the application example 2 is realized by the following means.

[0496] In this invention, the server includes a means for acquiring music information, a means for analyzing the user's emotions using an emotion recognition engine, a means for providing information on music and artists related to the user's preferences and emotions, and a means for adjusting the timing of music selection and information provision, thereby making it possible to provide a personalized music experience according to the user's emotional state.

[0497] The "means for acquiring music information" refers to a means for acquiring music information from a music database, the Internet, or the like.

[0498] An "emotion recognition engine" is software or hardware that analyzes a user's voice, facial expressions, text, etc., and recognizes the user's emotions.

[0499] "Means for providing information on music and artists related to the user's preferences and emotions" refers to means for providing information on music and artists based on the user's emotional state and preferences, and includes a display and audio output.

[0500] The "means for adjusting the timing of music selection and information provision" refers to a means for selecting music and providing information at optimal timing according to the user's emotional state.

[0501] The "means for monitoring changes in the user's emotional state" refers to a means for monitoring the user's emotional state in real time and adjusting the music experience in response to those changes.

[0502] In order to put the present invention into practice, it is first necessary to construct a system that includes an emotion recognition engine, a means for acquiring music information, and a means for selecting music and providing information. This system operates as follows.

[0503] The server has a means of accessing the Internet and a music information database to obtain music information. This means allows the server to always obtain the latest music information from the database. For example, the server can use the API of a well-known music streaming service.

[0504] An emotion recognition engine is implemented on the device and analyzes the user's voice, facial expression, text, etc. This makes it possible to recognize the user's emotions in real time. For example, a voice analysis API or facial expression recognition API is used for the emotion recognition engine. If the user inputs "I'm very sad today," the emotion recognition engine will recognize the emotion "sad."

[0505] Based on the recognized emotion, the server uses the music streaming service API to obtain information on music and artists that are suitable for the user's emotion. For example, it obtains a playlist that matches the "sad" mood. It then provides that information to the user via the device's display or audio output. For example, it may show "Recommended songs: ●● - ●●" on the display or notify the user by voice, "Here are some recommended songs."

[0506] In addition, the system monitors changes in the user's emotional state and adjusts music selection and information provision timing as necessary. Even if the user's emotion changes from "sad" to "happy," the system will detect this and suggest music appropriate to the new emotion. This allows the user to always enjoy a music experience that matches their emotion.

[0507] Specifically, the following prompt sentences are input to the generative AI model to perform emotion recognition: If a user types "I feel so sad today", an emotion recognition API will return the emotion "sad". Write a program that uses a music streaming service API to get a playlist that matches the "sad" mood and serves it to the user.

[0508] The system can provide a personalized music experience based on the user's emotional state, achieving a deeper level of personalization.

[0509] The flow of the specific process in the application example 2 will be described with reference to FIG.

[0510] Step 1: User emotion input The user inputs their mood or emotion into the device. For example, they can say, "I feel very sad today," or input it as text. This input becomes data for the next analysis.

[0511] Step 2: Emotion recognition The device uses the emotion recognition API to analyze the user's input data. Specifically, in the case of voice input, voice analysis is performed, and in the case of text input, sentence analysis is performed. Through this analysis, the user's emotion is recognized as "sad."

[0512] Step 3: Emotion-based music selection Based on the recognized emotion, the server uses the music streaming service API to search for music that matches the user's emotion. For example, if the emotion is "sad", it retrieves a playlist that matches the "sad" mood. At this stage, the retrieved data includes the song title, artist name, etc.

[0513] Step 4: Providing music information The music information obtained from the server is provided to the user via the device's display or audio output. Specifically, the device may show "Recommended songs: XX - XX" on the display or notify the user by voice, "Here are some recommended songs."

[0514] Step 5: Monitoring user emotional changes The system uses the device's emotion recognition engine to monitor the user's emotional state in real time, for example, re-recognizing emotions each time the user enters new voice or text input and determining whether the emotional state has changed.

[0515] Step 6: Timing adjustment of music provision If the user's emotion is recognized as having changed, the server again uses the music streaming service API to obtain music based on the new emotion (i.e., the changed emotion) and provides it to the user. For example, if the emotion changes from "sad" to "happy," the server obtains and provides a playlist that matches the "happy" mood.

[0516] This allows users to always enjoy a music experience that matches their emotions.

[0517] The specific processing unit 290 transmits the result of the specific processing to the robot 414. In the robot 414, the control unit 46A causes the speaker 240 and the control target 443 to output the result of the specific processing. The microphone 238 acquires a voice indicating a user input for the result of the specific processing. The control unit 46A transmits voice data indicating the user input acquired by the microphone 238 to the data processing device 12. In the data processing device 12, the specific processing unit 290 acquires the voice data.

[0518] The data generation model 58 is a so-called generative AI (Artificial Intelligence). An example of the data generation model 58 is ChatGPT (Internet search<URL: https: / / openai.com / blog / chatgpt> ), Gemini (Internet search <url: https: gemini.google.com ?hl="ja">) and other generation AIs. The data generation model 58 is obtained by making a neural network perform deep learning. A prompt including an instruction is input to the data generation model 58, and inference data such as voice data indicating a voice, text data indicating a text, and image data indicating an image is input. The data generation model 58 performs inference on the input inference data according to the instruction indicated by the prompt, and outputs the inference result in a data format such as voice data and text data. Here, inference refers to, for example, analysis, classification, prediction, and / or summarization.

[0519] In the above embodiment, an example was given in which the specific process was performed by the data processing device 12, but the technology of the present disclosure is not limited to this, and the specific process may be performed by the robot 414.

[0520] The emotion identification model 59 as an emotion engine may determine the emotion of the user according to a specific mapping. Specifically, the emotion identification model 59 may determine the emotion of the user according to an emotion map (see FIG. 9), which is a specific mapping. Similarly, the emotion identification model 59 may determine the emotion of the robot, and the identification processing unit 290 may perform identification processing using the emotion of the robot.

[0521] FIG. 9 is a diagram showing an emotion map 400 on which multiple emotions are mapped. In the emotion map 400, emotions are arranged in concentric circles radiating from the center. The closer to the center of the concentric circles, the more primitive emotions are arranged. The more outside the concentric circles, the more emotions that represent states and actions that arise from a state of mind are arranged. Emotions are a concept that includes emotions and mental states. On the left side of the concentric circles, emotions that are generally generated from reactions that occur in the brain are arranged. On the right side of the concentric circles, emotions that are generally induced by situational judgment are arranged. On the upper and lower sides of the concentric circles, emotions that are generally generated from reactions that occur in the brain and are induced by situational judgment are arranged. In addition, on the upper side of the concentric circles, emotions of "pleasure" are arranged, and on the lower side, emotions of "discomfort" are arranged. In this way, in the emotion map 400, multiple emotions are mapped based on the structure in which emotions are generated, and emotions that tend to occur simultaneously are mapped close to each other.

[0522] These emotions are distributed in the 3 o'clock direction of emotion map 400 and usually fluctuate between relief and anxiety. In the right half of emotion map 400, situational awareness dominates over internal sensations, resulting in a sense of calm.

[0523] The inside of emotion map 400 represents what is going on inside one's mind, and the outside of emotion map 400 represents behavior, so the further out on emotion map 400 you go, the more visible the emotions become (the more they are expressed in behavior).

[0524] Here, human emotions are based on various balances such as posture and blood sugar level, and when these balances are far from the ideal, it indicates an unpleasant state, and when they are close to the ideal, it indicates a pleasant state. Emotions can also be created for robots, cars, motorcycles, etc., based on various balances such as posture and battery level, so that when these balances are far from the ideal, it indicates an unpleasant state, and when they are close to the ideal, it indicates a pleasant state. The emotion map may be generated, for example, based on the emotion map of Dr. Mitsuyoshi (Research on speech emotion recognition and emotion brain physiological signal analysis system, Tokushima University, doctoral dissertation: https: / / ci.nii.ac.jp / naid / 500000375379). On the left half of the emotion map, emotions belonging to an area called "reaction" where sensation is dominant are lined up. On the right half of the emotion map, emotions belonging to an area called "situation" where situation recognition is dominant are lined up.

[0525] The emotion map defines two emotions that promote learning. The first is the negative emotion around the middle of "repentance" or "remorse" on the situation side. In other words, this is when the robot experiences negative emotions such as "I never want to feel this way again" or "I don't want to be scolded again." The other is the positive emotion around "desire" on the response side. In other words, this is when the robot has positive feelings such as "I want more" or "I want to know more."

[0526] The emotion identification model 59 inputs the user input to a pre-trained neural network, obtains emotion values ​​indicating each emotion shown in the emotion map 400, and determines the emotion of the user. This neural network is pre-trained based on multiple learning data that are combinations of the user input and emotion values ​​indicating each emotion shown in the emotion map 400. Furthermore, this neural network is trained so that emotions that are located close to each other have similar values, as in the emotion map 900 shown in Fig. 10. Fig. 10 shows an example in which multiple emotions, "relief," "calm," and "encouraging," have similar emotion values.

[0527] Although the system according to the present disclosure has been described above mainly with respect to the functions of the data processing device 12, the system according to the present disclosure is not necessarily implemented in a server. The system according to the present disclosure may be implemented as a general information processing system. The present disclosure may be implemented, for example, as a software program that runs on a personal computer, or an application that runs on a smartphone or the like. The method according to the present disclosure may be provided to a user in the form of SaaS (Software as a Service).

[0528] In the above embodiment, an example is given in which the specific processing is performed by one computer 22, but the technology of the present disclosure is not limited to this, and the specific processing may be distributed and performed by a plurality of computers including the computer 22. For example, the data generation model 58 may be provided in an external device of the data processing device 12, and the external device may generate data according to input data.

[0529] In the above embodiment, an example in which the specific processing program 56 is stored in the storage 32 has been described, but the technology of the present disclosure is not limited to this. For example, the specific processing program 56 may be stored in a portable computer-readable non-transitory storage medium such as a Universal Serial Bus (USB) memory. The specific processing program 56 stored in the non-transitory storage medium is installed in the computer 22 of the data processing device 12. The processor 28 executes the specific processing in accordance with the specific processing program 56.

[0530] In addition, the specific processing program 56 may be stored in a storage device such as a server connected to the data processing device 12 via the network 54, and the specific processing program 56 may be downloaded and installed on the computer 22 upon request from the data processing device 12.

[0531] It is not necessary to store all of the specific processing program 56 in a storage device such as a server connected to the data processing device 12 via the network 54, or to store all of the specific processing program 56 in the storage 32; only a portion of the specific processing program 56 may be stored.

[0532] As the hardware resource for executing the specific process, various processors as shown below can be used. An example of the processor is a CPU, which is a general-purpose processor that functions as a hardware resource for executing the specific process by executing software, i.e., a program. Another example of the processor is a dedicated electric circuit, which is a processor having a circuit configuration designed exclusively for executing the specific process, such as a Field-Programmable Gate Array (FPGA), a Programmable Logic Device (PLD), or an Application Specific Integrated Circuit (ASIC). Each processor has a built-in or connected memory, and each processor executes the specific process by using the memory.

[0533] The hardware resource that executes the specific process may be one of these various processors, or may be a combination of two or more processors of the same or different types (for example, a combination of multiple FPGAs, or a combination of a CPU and an FPGA). Also, the hardware resource that executes the specific process may be a single processor.

[0534] As an example of a configuration using one processor, first, one processor is configured by combining one or more CPUs and software, and this processor functions as a hardware resource that executes a specific process. Second, there is a configuration in which a processor is used that realizes the functions of the entire system including multiple hardware resources that execute a specific process on a single IC chip, as typified by SoC (System-on-a-chip). In this way, a specific process is realized using one or more of the above-mentioned various processors as hardware resources.

[0535] Furthermore, more specifically, the hardware structure of these various processors can be an electric circuit that combines circuit elements such as semiconductor elements. The specific processes described above are merely examples. It goes without saying that unnecessary steps may be deleted, new steps may be added, or the order of processes may be changed without departing from the spirit of the invention.

[0536] The above description and illustrations are detailed descriptions of the parts related to the technology of the present disclosure, and are merely an example of the technology of the present disclosure. For example, the above description of the configuration, function, action, and effect is an example of the configuration, function, action, and effect of the parts related to the technology of the present disclosure. Therefore, it goes without saying that unnecessary parts may be deleted, new elements may be added, or replacements may be made to the above description and illustrations, within the scope of the gist of the technology of the present disclosure. In addition, in order to avoid confusion and to facilitate understanding of the parts related to the technology of the present disclosure, the above description and illustrations omit explanations of technical common sense that do not require explanation in order to enable the implementation of the technology of the present disclosure.

[0537] All publications, patent applications, and standards mentioned in this specification are herein incorporated by reference to the same extent as if each individual publication, patent application, and standard was specifically and individually indicated to be incorporated by reference.

[0538] The following is further disclosed regarding the above embodiment.

[0539] (Appendix 1) A means for acquiring music information; a means for analyzing user preferences and questions; a means for providing music and artist information related to a user's preferences and queries; A system including:

[0540] (Appendix 2) 2. The system according to claim 1, wherein the means for acquiring music information is a means for accessing the Internet or a music information database.

[0541] (Appendix 3) The system according to claim 1, wherein the means for providing information about music or artists related to the user's preferences or questions is a means for providing information through a display or audio output.

[0542] (Appendix 4) 2. The system of claim 1, further comprising an emotion engine for recognizing user emotions.

[0543] (Appendix 5) The system described in Appendix 4 is characterized in that the emotion engine recognizes emotions from the user's voice, facial expressions, text, etc., and customizes music and artist suggestions based on the recognition results.

[0544] (Appendix 6) 5. The system of claim 4, wherein the emotion engine analyzes the user's emotional state and adjusts music selection and timing of information provision to provide a more personalized music experience.

[0545] "Example 1" (Claim 1) A means of accessing the Internet or a database to obtain music information; A means for analyzing user-supplied information and questions to identify appropriate music and artists; A means for transmitting information about the identified music or artist to a terminal and providing the information to a user through a display or audio output; A system including:

[0546] (Claim 2) 2. The system according to claim 1, wherein the means for acquiring music information is a means for accessing a network or a database.

[0547] (Claim 3) 2. The system according to claim 1, wherein the means for providing information about music or artists related to the user's preferences or questions is means for providing information through a display device or a voice conversion device.

[0548] "Application example 1" (Claim 1) A means for acquiring music information; a means for analyzing user preferences and questions; means for providing music-related information relevant to a user's preferences or queries; A means for analyzing a user's question using a generative AI model; and A means for generating a prompt sentence; A system including:

[0549] (Claim 2) 2. The system according to claim 1, wherein the means for acquiring music information is a means for accessing the Internet or a music database.

[0550] (Claim 3) 2. The system according to claim 1, wherein the means for providing music-related information related to the user's preferences or queries comprises means for providing information through a display device or an audio output device.

[0551] "Example 2 of combining emotion engines" (Claim 1) A means for recognizing a user's emotion; A means for obtaining music information suitable for a user based on the recognized emotion; A means for providing the acquired music information to a user; A system including:

[0552] (Claim 2) 2. The system according to claim 1, further comprising a means for accessing the Internet and a music information database.

[0553] (Claim 3) 2. The system of claim 1, further comprising a means for providing information through a display and / or audio output.

[0554] "Application example 2 when combining emotion engines" (Claim 1) A means for acquiring music information; A means for analyzing a user's emotions using an emotion recognition engine; A means for providing information on music and artists related to the user's tastes and emotions; A means to adjust the selection of music and the timing of information provision; A system including:

[0555] (Claim 2) 2. The system according to claim 1, wherein the means for acquiring music information is a means for accessing the Internet or a music information database.

[0556] (Claim 3) 2. The system according to claim 1, wherein the means for analyzing the user's emotions using an emotion recognition engine is a means for analyzing the user's voice, facial expression, and sentences.

[0557] (Claim 4) 2. The system according to claim 1, wherein the means for providing information about music and artists related to the user's preferences and emotions is means for providing information through a display and / or audio output.

[0558] (Claim 5) 2. The system of claim 1, wherein the means for adjusting the music selection and the timing of the information provision comprises means for monitoring changes in the user's emotional state. [Explanation of symbols]

[0559] 10, 210, 310, 410 Data Processing Systems 12 Data Processing Device 14 Smart Devices 214 Smart Glasses 314 Headset-type terminal 414 Robot< / url:> < / url:> < / url:> < / url:>

Claims

1. A means for acquiring music information; A means for analyzing a user's emotions using an emotion recognition engine; means for providing music and artist information relevant to a user's question or emotion; a means for adjusting a timing of selecting music and providing information based on the analysis result of the user's emotion; A system including:

2. The system of claim 1, characterized in that the means for providing information on music and artists related to the user's question or emotion is a means for providing the information on the music and artists based on a prompt sentence generated in response to the question and a generative AI.

3. 2. The system of claim 1, wherein the means for adjusting the timing of the music selection and the information presentation comprises means for monitoring changes in the user's emotional state.

4. The system according to claim 3, characterized in that the means for adjusting the timing of the music selection and information provision is means for providing music that matches the user's new emotions when it is recognized that the user's emotions have changed.

Citation Information

Patent Citations

  • Persona chatbot control method and system

    JP2022180282A