system
The system addresses the limitation of conventional communication devices by generating background music based on user conversation and musical preferences, enhancing emotional engagement and satisfaction.
Patent Information
- Authority / Receiving Office
- JP · JP
- Patent Type
- Applications
- Current Assignee / Owner
- SOFTBANK GROUP CORP
- Filing Date
- 2024-10-11
- Publication Date
- 2026-04-23
AI Technical Summary
Conventional dialogue experiences in communication devices lack the ability to effectively convey user emotions and intentions through background music, leading to a less engaging and satisfying communication experience.
A system that analyzes user conversation history and musical preferences to generate and play background music in real-time, using AI models to improve accuracy based on user feedback.
Enhances communication experiences by providing background music tailored to user emotions and conversation contexts, creating a richer and more memorable interaction.
Smart Images

Figure 2026069082000001_ABST
Abstract
Description
Technical Field
[0005] , ,
[0001] The technology of the present disclosure relates to a system.
Background Art
[0002] Patent Document 1 discloses a method for controlling a persona chatbot, which is performed by at least one processor, and includes steps of receiving a user utterance, adding the user utterance to a prompt including an instruction sentence related to an explanation of a chatbot character, encoding the prompt, and inputting the encoded prompt into a language model to generate a chatbot utterance in response to the user utterance.
Prior Art Documents
Patent Documents
[0003]
Patent Document 1
Summary of the Invention
Problems to be Solved by the Invention
[0004] In the conventional dialogue experience in communication devices, it is only composed of text and voice, and the means for effectively conveying the feelings and intentions of users are limited. In addition, due to the lack of appropriate background music during the dialogue, it is difficult to carry out the conversation in a relaxed atmosphere, and there may be an opportunity to miss improving the quality of communication. As a result, users cannot enjoy a music experience that suits their mood and situation, and dissatisfaction may occur.
Means for Solving the Problems
[0005] This invention utilizes the conversation history and music listening history of users of communication devices, and analyzes this history data to identify the characteristics of the conversation and the user's musical preferences. It then provides a means for generating background music based on the identified conversation characteristics and musical preferences, and transmitting the generated background music to the communication device for real-time playback. Furthermore, the accuracy of the generation means is continuously improved by optimizing the generation means based on user evaluation information and training an artificial intelligence model. In this way, the system can efficiently convey the user's emotions and intentions through music without disrupting the atmosphere of the conversation, providing a richer communication experience.
[0006] "Communication equipment" refers to devices used by users for communication, and includes electronic devices such as smartphones and tablets.
[0007] "User" refers to an individual who uses communication devices to communicate.
[0008] "Conversation history" refers to the record of text messages and voice messages sent by users through communication devices.
[0009] "Music listening history" refers to a record of music tracks that a user has previously played using a communication device.
[0010] "Analysis" is the process of extracting information from conversation history and music listening history to reveal patterns and characteristics.
[0011] "Background music" refers to music that plays during a user's conversation and serves to enhance the atmosphere of the conversation.
[0012] "Generation" refers to the process of creating new background music based on the analyzed information.
[0013] "Transmission" refers to the process of delivering the generated background music to communication devices.
[0014] "Playback" refers to the state where a user listens to music by outputting background music through a communication device.
[0015] "Evaluation information" refers to the feedback and accompanying data provided by a user to the system.
[0016] "Optimization" is a process of improving the performance of the system based on evaluation information.
[0017] "Artificial intelligence model" is an algorithm that learns based on data and executes complex generation tasks.
Brief Description of Drawings
[0018] [Figure 1] It is a conceptual diagram showing an example of the configuration of a data processing system according to the first embodiment. [Figure 2] It is a conceptual diagram showing an example of the main functions of a data processing device and a smart device according to the first embodiment. [Figure 3] It is a conceptual diagram showing an example of the configuration of a data processing system according to the second embodiment. [Figure 4] It is a conceptual diagram showing an example of the main functions of a data processing device and smart glasses according to the second embodiment. [Figure 5] It is a conceptual diagram showing an example of the configuration of a data processing system according to the third embodiment. [Figure 6] It is a conceptual diagram showing an example of the main functions of a data processing device and a headset-type terminal according to the third embodiment. [Figure 7] It is a conceptual diagram showing an example of the configuration of a data processing system according to the fourth embodiment. [Figure 8] It is a conceptual diagram showing an example of the main functions of a data processing device and a robot according to the fourth embodiment. [Figure 9] It shows an emotion map to which a plurality of emotions are mapped. [Figure 10] It shows an emotion map to which a plurality of emotions are mapped. [Figure 11]It is a sequence diagram showing the processing flow of the data processing system in Embodiment 1. [Figure 12] It is a sequence diagram showing the processing flow of the data processing system in Application Example 1. [Figure 13] It is a sequence diagram showing the processing flow of the data processing system in Embodiment 2 when the emotion engine is combined. [Figure 14] It is a sequence diagram showing the processing flow of the data processing system in Application Example 2 when the emotion engine is combined.
Mode for Carrying Out the Invention
[0019] Hereinafter, an example of an embodiment of the system according to the technology of the present disclosure will be described with reference to the accompanying drawings.
[0020] First, the terms used in the following description will be explained.
[0021] In the following embodiments, the numbered processor (hereinafter simply referred to as "processor") may be one arithmetic unit or a combination of multiple arithmetic units. Also, the processor may be one type of arithmetic unit or a combination of multiple types of arithmetic units. Examples of arithmetic units include CPU (Central Processing Unit), GPU (Graphics Processing Unit), GPGPU (General-Purpose computing on Graphics Processing Units), APU (Accelerated Processing Unit), etc.
[0022] In the following embodiments, the numbered RAM (Random Access Memory) is a memory in which information is temporarily stored and is used as a work memory by the processor.
[0023] In the following embodiments, the signed storage is one or more non-volatile storage devices that store various programs and various parameters. Examples of non-volatile storage devices include flash memory (SSD (Solid State Drive)), magnetic disks (e.g., hard disks), or magnetic tapes.
[0024] In the following embodiments, the signed communication interface (I / F) is an interface that includes a communication processor and an antenna, etc. The communication interface manages communication between multiple computers. Examples of communication standards applicable to the communication interface include wireless communication standards such as 5G (5th Generation Mobile Communication System), Wi-Fi (registered trademark), or Bluetooth (registered trademark).
[0025] In the following embodiments, "A and / or B" is synonymous with "at least one of A and B." That is, "A and / or B" means that it may be A alone, or B alone, or a combination of A and B. Furthermore, in this specification, the same concept as "A and / or B" applies when expressing three or more things linked by "and / or."
[0026] [First Embodiment]
[0027] Figure 1 shows an example of the configuration of the data processing system 10 according to the first embodiment.
[0028] As shown in Figure 1, the data processing system 10 includes a data processing device 12 and a smart device 14. An example of the data processing device 12 is a server.
[0029] The data processing device 12 comprises a computer 22, a database 24, and a communication interface 26. The computer 22 is an example of a "computer" related to the technology of this disclosure. The computer 22 comprises a processor 28, RAM 30, and storage 32. The processor 28, RAM 30, and storage 32 are connected to a bus 34. The database 24 and the communication interface 26 are also connected to the bus 34. The communication interface 26 is connected to a network 54. An example of the network 54 is a WAN (Wide Area Network) and / or a LAN (Local Area Network).
[0030] The smart device 14 comprises a computer 36, a reception device 38, an output device 40, a camera 42, and a communication interface 44. The computer 36 comprises a processor 46, RAM 48, and storage 50. The processor 46, RAM 48, and storage 50 are connected to a bus 52. The reception device 38, output device 40, and camera 42 are also connected to the bus 52.
[0031] The reception device 38 is equipped with a touch panel 38A and a microphone 38B, etc., and receives user input. The touch panel 38A receives user input by detecting contact with an object (e.g., a pen or finger). The microphone 38B receives user input by detecting the user's voice. The control unit 46A transmits data indicating the user input received by the touch panel 38A and microphone 38B to the data processing device 12. In the data processing device 12, the specific processing unit 290 acquires the data indicating the user input.
[0032] The output device 40 includes a display 40A and a speaker 40B, and presents data to the user 20 by outputting the data in a form perceptible to the user 20 (e.g., audio and / or text). The display 40A displays visible information such as text and images according to instructions from the processor 46. The speaker 40B outputs audio according to instructions from the processor 46. The camera 42 is a small digital camera equipped with an optical system such as a lens, aperture, and shutter, and an image sensor such as a CMOS (Complementary Metal-Oxide-Semiconductor) image sensor or a CCD (Charge Coupled Device) image sensor.
[0033] Communication interface 44 is connected to network 54. Communication interfaces 44 and 26 are responsible for the exchange of various types of information between processor 46 and processor 28 via network 54.
[0034] Figure 2 shows an example of the main functions of the data processing device 12 and the smart device 14.
[0035] As shown in Figure 2, in the data processing device 12, specific processing is performed by the processor 28. A specific processing program 56 is stored in the storage 32. The specific processing program 56 is an example of a "program" related to the technology of this disclosure. The processor 28 reads the specific processing program 56 from the storage 32 and executes the read specific processing program 56 on the RAM 30. The specific processing is realized by the processor 28 operating as a specific processing unit 290 according to the specific processing program 56 executed on the RAM 30.
[0036] The storage 32 stores the data generation model 58 and the emotion identification model 59. The data generation model 58 and the emotion identification model 59 are used by the identification processing unit 290.
[0037] In the smart device 14, the processor 46 performs the reception output processing. The storage 50 stores the reception output program 60. The reception output program 60 is used in conjunction with a specific processing program 56 by the data processing system 10. The processor 46 reads the reception output program 60 from the storage 50 and executes the read reception output program 60 on the RAM 48. The reception output processing is realized by the processor 46 operating as a control unit 46A according to the reception output program 60 executed on the RAM 48.
[0038] Next, the specific processing performed by the specific processing unit 290 of the data processing device 12 will be described. In the following description, the data processing device 12 will be referred to as the "server" and the smart device 14 as the "terminal".
[0039] This invention is a system that complements the communication experience using communication devices with music. This system provides users with background music that is adapted to their emotions and circumstances through the interaction between the server, terminal, and user.
[0040] The server first securely retrieves the user's conversation history and analyzes the conversation content. Natural language processing techniques are used for the analysis to identify keywords and emotional tones. In parallel, the server retrieves the user's music listening history to understand their preferred music genres and artists. Based on both sets of data, the server uses an AI model to generate background music that best suits the conversation context and the user's musical preferences.
[0041] The generated background music is transmitted to the user's device in real time. The device automatically plays the received music, allowing the user to hear appropriate background music during conversations. This playback enhances the user's communication experience, making it more emotional and memorable.
[0042] As a concrete example, consider a scenario where a user is having a fun conversation with a friend about an event. In this case, the server extracts keywords such as "event" and "fun" from the conversation and identifies upbeat pop music that the user has enjoyed listening to in the past from their listening history. The background music generated based on this information is played on the device, adding a fun atmosphere to the conversation and allowing the user to enjoy the conversation even more.
[0043] Furthermore, based on evaluation information obtained from users, the server continuously trains its AI model to improve the accuracy of generation. In this way, the present invention can provide music tailored to individual conversational experiences, creating new experiential value for users.
[0044] The following describes the processing flow.
[0045] Step 1:
[0046] With the user's consent, the server retrieves the user's conversation history and music listening history from the database. This involves secure data transfer using the latest security protocols.
[0047] Step 2:
[0048] The server analyzes the acquired conversation history using natural language processing (NLP) techniques. This analysis identifies keywords and basic emotional tones from the conversation content. For example, it classifies emotions such as "happy," "sad," and "busy."
[0049] Step 3:
[0050] The server analyzes the user's music listening history to identify their musical preferences. This includes a process of examining play counts, genres, and artist tendencies.
[0051] Step 4:
[0052] The server uses an AI model to generate optimal background music based on the characteristics of the conversation and the user's musical preferences. In this process, musical elements corresponding to the conversational context (e.g., formal, casual) are combined.
[0053] Step 5:
[0054] The generated background music is configured to be streamed from the server to the user's device. The device immediately plays the received music, seamlessly playing it in the background of the user's conversation.
[0055] Step 6:
[0056] Users can provide feedback on the background music being played. For example, they can rate whether they like the music or not using a score. They can also record actions taken during playback, such as adjusting the volume or skipping tracks.
[0057] Step 7:
[0058] The device sends this feedback and actions back to the server, which is used to train the AI model. The server uses this feedback to update the BGM generation algorithm and improve its accuracy.
[0059] This series of steps allows users to receive a music experience optimized for their individual conversations.
[0060] (Example 1)
[0061] Next, we will describe Example 1. In the following description, the data processing device 12 will be referred to as the "server," and the smart device 14 will be referred to as the "terminal."
[0062] In modern communication devices, there is a growing demand for richer and more emotionally engaging communication experiences by automatically providing background music that is appropriate to the user's emotions and preferences. However, conventional technologies struggle to generate and provide background music in real time based on the content of the conversation and the user's musical preferences. This results in users not being able to fully experience music that is appropriate to the situation.
[0063] The identification process performed by the identification processing unit 290 of the data processing device 12 in Example 1 is realized by the following means.
[0064] In this invention, the server includes a function to acquire a record of the communication device user's conversation, a function to acquire the user's music history, and a function to analyze the conversation record and music history to identify the characteristics of the conversation and the user's musical preferences. This makes it possible to generate optimal background music based on the characteristics and preferences of the conversation and deliver it to the user in real time.
[0065] A "communication device" is an electronic device used to send and receive voice and data, and is a terminal that supports user communication.
[0066] A "dialogue record" is data that digitally stores the content of conversations and messages conducted through communication devices.
[0067] "Music history" refers to data on songs and music that a user has listened to in the past, and is information that indicates the user's musical preferences.
[0068] A "generative model" is an algorithm based on machine learning or artificial intelligence that automatically creates new data, especially background music, based on acquired data.
[0069] A "control statement" is a text that represents commands or instructions used as input to a generative model, and is data used to guide the generative process.
[0070] An "intelligent model" is a system that uses machine learning algorithms to continuously learn based on user feedback, thereby improving the accuracy and performance of the model.
[0071] This invention relates to a system for enhancing the dialogue experience in a communication device through background music. This system has the function of providing music appropriate to the user's emotions and the context of the conversation through the interaction between the server, terminal, and user.
[0072] The server first acquires the content of conversations taking place on the communication device and analyzes it using natural language processing techniques. Possible software used for this would include natural language processing libraries such as "spaCy" and "NLTK." The server uses these libraries to extract keywords and emotional tones from the dialogue and also retrieves the user's music history from a database. This database contains information about songs and genres the user has previously played. Based on this data, the server creates control statements to input to the generative AI model.
[0073] The generative AI model uses technologies such as "OpenAI® GPT" to generate music data based on prompts created by the server. An example of such a prompt is "Emotional tone: Happy, Music genre: Pop." The generated music data is sent to the device via streaming.
[0074] The device receives music data sent from the server and plays it in the background. This playback uses the media player application installed on the device. This allows the user to naturally listen to music appropriate to the situation while having a conversation.
[0075] For example, if a user is having a conversation with a friend about a fun event, the server analyzes keywords such as "event" and "fun" and generates music based on pop music the user has liked in the past. This music is played on the device, adding an element of fun to the conversation and improving the user's experience.
[0076] Furthermore, users can provide feedback on the music played. The server collects this feedback and continuously trains its generation AI model to achieve more accurate music generation. With this configuration, the system can enrich the user's interactive experience with individually customized music.
[0077] The flow of the specific processing in Example 1 will be explained using Figure 11.
[0078] Step 1:
[0079] The server acquires dialogue data from the user's communication device. Input includes real-time ongoing conversation audio and text data. The server first converts this dialogue data into a digital format and securely acquires it using a secure communication protocol.
[0080] Step 2:
[0081] The server analyzes the acquired dialogue data using natural language processing (NLP) techniques. Input includes digital audio and text data. The server uses an NLP library to extract keywords and emotional tones from the text. This step results in outputting emotional tones such as "event" and "fun."
[0082] Step 3:
[0083] The server retrieves the user's music history from the database. Input includes the user ID and past playback history. The server uses SQL queries to retrieve listening data for songs related to the user. The output is a list of music genres and artists the user has previously enjoyed listening to.
[0084] Step 4:
[0085] The server creates prompts for the generative AI model based on the analysis results and music history. The input includes the output data from steps 2 and 3. The server constructs emotional tone and musical preferences as prompts and prepares them for input into the generative AI model. The output is the constructed prompts.
[0086] Step 5:
[0087] The server generates background music using a generative AI model. The input is a generated prompt message. Based on this, the generative AI model adjusts the parameters of the music data and generates new music. The output is the generated music data.
[0088] Step 6:
[0089] The server streams the generated music data to the user's device. The input includes music data from the generation AI model. The server converts the data into a compressed format and sends it to the device in real time. The output is the streaming music data.
[0090] Step 7:
[0091] The device decodes the received music data and plays it in the background. The input includes music data from the server. The device launches a media player and plays the music, providing the user with a real-time auditory experience. The output is the music playing from the device.
[0092] Step 8:
[0093] Users provide feedback on the played music via their device. Input includes user ratings regarding the music's suitability. Users input their ratings using an interface on their device. Output is the rating data sent to the server.
[0094] Step 9:
[0095] The server trains a generative AI model using user evaluation data. The input includes evaluation information. The server utilizes the collected data to provide feedback to the AI model, aiming to improve the accuracy of music generation. The output is the improved AI model.
[0096] (Application Example 1)
[0097] Next, we will explain Application Example 1. In the following explanation, the data processing device 12 will be referred to as the "server," and the smart device 14 will be referred to as the "terminal."
[0098] There is a need for ways to provide customers with a more comfortable and emotionally resonant experience within stores. However, current technology lacks the means to analyze customer emotions and conversation content in real time and automatically provide music that is appropriate for that, making it difficult to improve the customer experience.
[0099] The specific processing performed by the specific processing unit 290 of the data processing device 12 in Application Example 1 is realized by the following means.
[0100] In this invention, the server includes means for acquiring user voice data, means for analyzing the voice data to identify topics and emotions, and means for acquiring user music preference data. This makes it possible to generate background sounds in real time that correspond to the customer's emotions and conversation content, and to play them back in the store.
[0101] "User voice data" refers to audio information of conversations between customers and employees within the store.
[0102] "Methods for identifying topics and emotions" refer to technologies that analyze audio data to identify what customers are saying and the tone of their emotions.
[0103] "Music preference data" refers to information about music genres and artists that a user has liked in the past.
[0104] "Means of generating background sounds" refers to technology that automatically creates music to enhance the customer experience based on identified topics, emotions, and music preference data.
[0105] An "output device" is a device used to physically reproduce the generated background sound, and includes the sound system within the store.
[0106] This invention relates to a system that acquires audio data of conversations between customers and employees in a store and generates and plays background music in real time based on that data. The server receives the audio data acquired using the microphone of smart glasses and analyzes it using natural language processing (NLP) software (e.g., NLTK, SpaCy). This analysis identifies the customer's topics of conversation and emotions. The server further refers to the customer's music preference data and generates appropriate background music using an AI model (e.g., GPT-4®) based on this information.
[0107] The generated music is transmitted to the in-store speaker system, which acts as the output device, via a music streaming service API (e.g., Spotify API). This system allows customers to listen to background music that matches their emotions and conversation, resulting in a more comfortable and satisfying experience.
[0108] As a concrete example, if a customer in a general store is discussing gift ideas and the tone of their conversation is identified as bright and positive, the server can generate upbeat jazz music to liven up the store's atmosphere. An example of a prompt might be, "The customer is in a cheerful mood while choosing a gift. Please recommend music that would be perfect for this." In this way, the entire store environment can harmonize with the customer's needs, enabling the provision of high-value-added services.
[0109] The flow of a specific process in Application Example 1 will be explained using Figure 12.
[0110] Step 1:
[0111] The server receives audio data acquired by the device through the smart glasses' microphone. The received audio data includes the content of conversations between customers and employees. The input is audio data, and the output is an audio file in digital format for analysis preparation.
[0112] Step 2:
[0113] The server analyzes the received audio data using natural language processing libraries (e.g., NLTK, SpaCy). This analysis identifies the topic and sentiment. This process includes converting the audio data to text and performing sentiment analysis and keyword extraction. The input is a digital audio file, and the output is text-based topic and sentiment data.
[0114] Step 3:
[0115] The server retrieves customer music preference data from a database. This data is based on the customer's past music listening history, favorite genres, and artists. The input is the customer ID or profile information, and the output is music preference data.
[0116] Step 4:
[0117] The server generates appropriate background music using an AI model (e.g., GPT-4) based on identified topics, emotions, and music preference data. The AI model sets the context for music generation using prompts and generates the corresponding music clips. The input is topic, emotion data, and music preference data, and the output is the generated music file.
[0118] Step 5:
[0119] The server sends the generated music to the store's speaker system via a music streaming service API (e.g., Spotify API). This transmission allows background music to play in real time. The input is the generated music file, and the output is the playback of background music within the store.
[0120] Step 6:
[0121] Users listen to background music while making purchases, leading to increased satisfaction. User feedback is collected and used to improve the accuracy of subsequent music generation models. The input is user feedback data, and the output is updated music preference data and training data for the AI model.
[0122] Furthermore, an emotion engine that estimates the user's emotions may be incorporated. That is, the identification processing unit 290 may use the emotion identification model 59 to estimate the user's emotions and perform identification processing using the user's emotions.
[0123] This invention provides a system that further enhances the communication experience of users using communication devices through emotion recognition. In this system, a server, a terminal, and an emotion engine that analyzes the user's conversation history work together. The server first acquires the user's conversation history and performs analysis using natural language processing technology. Through this analysis, in addition to the characteristics of the conversation, the emotion engine recognizes the user's emotions.
[0124] The emotion engine classifies specific emotions such as joy, sadness, and anger based on the user's words and expressions. This emotion information is analyzed by the server along with the user's music listening history to more precisely identify the user's musical preferences. Based on this data, the server uses an AI model to generate background music appropriate to the emotion.
[0125] The generated background music is adjusted to match the user's emotions and sent to the device in real time. The device plays this music during the conversation, creating a space that complements the user's emotions. This feature enriches the user's emotional conversational experience and enables deeper communication.
[0126] As a concrete example, when a user's hard-worked project succeeds, the server uses an emotion engine to identify the emotion of "joy" from the conversation. It then generates upbeat, cheerful background music that emphasizes this emotion and plays it on the user's device. This allows the user to feel and experience that moment more vividly.
[0127] Furthermore, this system improves the accuracy of background music generation by continuously training its emotion engine and AI model based on user feedback. This allows it to provide a music experience that best suits the user's characteristics over time.
[0128] The following describes the processing flow.
[0129] Step 1:
[0130] The server collects conversation history data from communication devices based on the user's consent. This data is encrypted and securely transferred to the server while protecting privacy.
[0131] Step 2:
[0132] The server analyzes the collected conversation history using natural language processing algorithms. The purpose of the analysis is to extract keywords that appear in the conversation and to understand the overall atmosphere of the conversation.
[0133] Step 3:
[0134] The server uses an emotion engine to further examine the analysis results and recognize the user's emotional state (joy, sadness, tension, etc.). The emotion engine determines emotions based on contextual nuances and the frequency of use of specific words.
[0135] Step 4:
[0136] The server retrieves the user's music listening history data and analyzes the user's preferred music genres and styles. This process refers to the frequency of music playback and the trends in the playlists the user has selected.
[0137] Step 5:
[0138] The server uses an AI model based on emotions recognized from the conversation and analyzed musical preferences to generate background music optimized for the user's current emotional state.
[0139] Step 6:
[0140] The server streams the generated background music to the user's device in real time. The device plays this stream in the background, providing a pleasant musical experience that matches the user's conversation.
[0141] Step 7:
[0142] Users can send feedback on the background music being played to the server via their device. This feedback includes opinions on the music's rating and emotional coherence.
[0143] Step 8:
[0144] Based on the collected feedback, the server trains its emotion engine and AI model to further improve the accuracy of background music generation. This continuous learning process makes the music experience more personalized for the user.
[0145] (Example 2)
[0146] Next, we will describe Example 2. In the following description, the data processing device 12 will be referred to as the "server" and the smart device 14 as the "terminal".
[0147] In the modern era, while communication via devices is increasing, digital communication has limitations in conveying emotions, making it difficult to build the same emotional connection as in face-to-face conversations. Therefore, there is a need to provide music experiences that respond to users' emotions and improve the quality of communication.
[0148] The identification process performed by the identification processing unit 290 of the data processing device 12 in Example 2 is realized by the following means.
[0149] In this invention, the server includes means for acquiring the user's conversation history of a communication terminal, means for acquiring the user's music playback history, and means for analyzing the conversation history using a processing method and identifying the user's emotions using an emotion engine. This makes it possible to generate appropriate background music based on the user's emotions and provide it in real time.
[0150] A "communication terminal" is an electronic device used by users to send and receive voice and data.
[0151] "Dialogue history" refers to a record of conversations that users have had through their communication devices.
[0152] "Music playback history" is a record of the music a user has listened to in the past, and is used to analyze their preferences.
[0153] "Processing methods" is a general term for the techniques and methods used to analyze data and extract information.
[0154] An "emotion engine" is a program that analyzes dialogue content to identify the user's emotional state.
[0155] An "artificial intelligence model" is a collection of algorithms that utilize technologies such as machine learning to achieve specific capabilities.
[0156] "Background music" refers to music generated to complement the user's emotional state and enhance their experience.
[0157] "Real-time" refers to a situation where processing is performed instantly with virtually no delay.
[0158] This invention is a system that enhances the user's conversational experience using a communication terminal, adding depth to digital communication by providing background music that corresponds to the emotional state. The following describes a specific form for its implementation.
[0159] The server first retrieves the user's dialogue history from the communication terminal. This is done using speech recognition technology or a text message transfer protocol. Next, it analyzes the retrieved dialogue history using natural language processing technology and identifies the user's emotions using an emotion engine. This emotion engine is a machine learning model that uses a variety of dialogue patterns registered in a database to classify emotions.
[0160] In parallel, the server retrieves and analyzes the user's music playback history. This reveals what kind of music the user preferred to listen to in the past, and what emotional states they were in. Based on the identified emotional and musical preference data, the server creates new background music using a generative AI model. This AI model incorporates a music generation algorithm that utilizes deep learning technology, and is given information such as "the user is feeling happy" as an input prompt.
[0161] The generated music is transmitted from the server to the user's communication terminal in real time. The terminal plays this music in the background during the conversation, allowing the user to experience the emotions of the moment more deeply.
[0162] For example, if a user says, "I passed my exam today," the server analyzes the conversation history to identify the emotion of "joy." Based on this, upbeat, cheerful music is generated and immediately played on the device. An example of this prompt message might be, "Generate cheerful music that expresses the user's joy."
[0163] In this way, the system can provide a music experience tailored to the user's emotions, thereby improving the quality of communication.
[0164] The flow of the specific processing in Example 2 will be explained using Figure 13.
[0165] Step 1:
[0166] The server retrieves the user's conversation history from the communication terminal. It receives text and audio data sent from the terminal as input and saves it as conversation history. Specifically, it reads the text chat log or performs speech recognition to convert audio to text. This process saves the conversation content to the database in a specific format.
[0167] Step 2:
[0168] The server uses natural language processing (NLP) techniques to analyze the acquired dialogue history. The input is the stored dialogue history, and the server processes this data to identify emotions. Specifically, it uses an NLP library to perform morphological analysis and contextual understanding, extracting keywords and emotional phrases from the conversation as output. This extracts the characteristics of the conversation and provides an indicator of the user's emotions.
[0169] Step 3:
[0170] The server uses an emotion engine to identify the user's emotions from the analysis results. The input is keywords and emotion indicators extracted in the previous step, and the output is a specific emotion category (e.g., joy, sadness, anger). Specifically, it applies an emotion classification algorithm to calculate which emotion the dialogue content corresponds to. In this step, the emotion engine can train its model using past data to improve its accuracy.
[0171] Step 4:
[0172] The server retrieves the user's music playback history and analyzes it in conjunction with emotional information. It references the user's playback history database as input and creates a profile of the user's musical preferences as output. Specifically, it queries the playback history from the database, analyzing song genres, tempos, and emotional data from past listening sessions to identify the user's unique musical tastes.
[0173] Step 5:
[0174] The server uses a generation AI model to generate background music based on identified emotions and musical preferences. The input is emotion category and music profile information, and the output is a music file generated by the AI model. Specifically, the operation involves inputting the prompt "The user's emotion is 'joy,' so please generate bright and cheerful music" into the AI model and executing the music generation algorithm.
[0175] Step 6:
[0176] The server transmits the generated music to the communication terminal in real time, and the terminal plays it. The input is the generated music file, and the output is the music playback on the terminal. Specifically, the music file is sent to the terminal in audio stream format, and the terminal plays this data using an audio output device. Through this process, the user listens to music that resonates with their emotions during a conversation, improving the quality of the conversation.
[0177] (Application Example 2)
[0178] Next, we will explain application example 2. In the following explanation, the data processing device 12 will be referred to as a "server" and the smart device 14 as a "terminal".
[0179] Conventional communication devices have a problem in that they simply play music without considering the user's emotional state, making it difficult to provide a truly personalized music experience. Furthermore, there was a lack of means to instantly recognize emotions from the user's facial expressions and voice and generate and provide appropriate background music. Therefore, it was difficult to provide a music experience that would enhance the user's mental state.
[0180] The specific processing performed by the specific processing unit 290 of the data processing device 12 in Application Example 2 is realized by the following means.
[0181] In this invention, the server includes means for acquiring audio and visual data of a user of a communication device, means for analyzing the user's musical characteristics and identifying their emotional state, means for adjusting musical characteristics based on the emotional state obtained from the audio and visual data, means for generating background music using a generation AI model based on the adjusted musical characteristics, and means for transmitting and playing the generated background music to the communication device in real time. This makes it possible to provide background music that matches the user's emotional state and enrich the user's life experience.
[0182] "Communication equipment" refers to electronic devices used to send and receive voice and data, and which users use to communicate with others.
[0183] "Voice data" refers to information obtained from the user's voice, recorded in digital format for use in emotional analysis.
[0184] "Visual data" refers to information obtained from images and videos acquired using cameras and sensors, and is used to analyze the user's facial expressions and movements.
[0185] "Music characteristics" refer to information that indicates a user's musical preferences and tendencies, and are used for song selection and background music generation.
[0186] "Emotional state" refers to the state of a user's current emotions and is identified through the analysis of audio and visual data.
[0187] "Musical characteristics" refer to attributes of music such as tempo, melody, rhythm, and timbre, which are adjusted to suit the user's emotional state.
[0188] A "generative AI model" is a computer program that uses artificial intelligence and has the ability to generate new music based on input data.
[0189] "Background music" refers to music provided to complement or enhance the user's experience, and is generated to match the user's emotional state.
[0190] This invention is a system for providing a music experience tailored to the user's emotional state. The system comprises communication equipment, a server, and AI-based data analysis capabilities.
[0191] The server acquires the user's voice and visual data through communication devices such as smartphones and smart glasses. This utilizes microphones and cameras built into the devices.
[0192] The acquired data is analyzed using natural language processing technologies such as Google® Cloud Natural Language API and Microsoft® Azure® Emotion API, as well as image recognition technologies, and classified as the user's emotional state. For example, it is determined whether the user's voice tone and facial expressions indicate emotions such as "happiness" or "anger."
[0193] Next, this emotional information is used to adjust the musical characteristics. Specifically, this involves selecting songs that match the user's emotions from music services such as Spotify and Apple Music, or generating new music through a generative AI model.
[0194] The generated background music is transmitted in real time to the user's communication device and played on the device. The user can experience a space where emotions are complemented through the music.
[0195] For example, if a user uses the device while feeling stressed, the server will determine that the user is in a state where relaxation is needed, and will select, generate, and play calming classical music.
[0196] Examples of prompts to input into a generative AI model:
[0197] "The user's current emotional state is 'relaxed'. Please choose calming music."
[0198] This system allows users to receive an optimal musical experience tailored to their individual emotional state, enabling them to experience emotional richness in their daily lives.
[0199] The flow of a specific process in Application Example 2 will be explained using Figure 14.
[0200] Step 1:
[0201] The device acquires the user's voice and visual data. This involves using the device's built-in microphone and camera to collect digital data of voice and facial expressions in real time. This data is stored for a certain period for subsequent analysis.
[0202] Step 2:
[0203] The server receives the collected audio and visual data and prepares it for data analysis. The inputs are audio and images, and the output is digital data converted from this data into an analyzable format. At this stage, preprocessing such as noise reduction and image correction is performed.
[0204] Step 3:
[0205] The server uses pre-processed digital data and leverages the Google Cloud Natural Language API and Microsoft Azure Emotion API to analyze the user's emotional state. The analysis results in the output of numerical values or labels indicating the emotional state (e.g., happy, sad, angry).
[0206] Step 4:
[0207] The server adjusts musical characteristics based on the analyzed emotional state. The input is emotional data, and the output is the corresponding musical genre, tempo, and other characteristics. This identifies a musical style that suits the user's emotions.
[0208] Step 5:
[0209] The server uses the results from the previous step to generate background music using a generative AI model. The input for this step is musical characteristics, and the output is music data generated by the AI. The following prompt is used for the generative AI model: "The user's current emotional state is 'relaxed'. Please select calming music."
[0210] Step 6:
[0211] The server sends the generated music data to the terminal. The input is the generated music data, and the output is a notification confirming transmission. The terminal receives and stores this data.
[0212] Step 7:
[0213] The device plays the received music for the user. A music playback application automatically launches on the device, providing a musical experience tailored to the user's current emotional state. The input is music data, and the output is the played sound as an audio signal.
[0214] The specific processing unit 290 transmits the result of the specific processing to the smart device 14. In the smart device 14, the control unit 46A causes the output device 40 to output the result of the specific processing. The microphone 38B acquires audio indicating user input for the result of the specific processing. The control unit 46A transmits the audio data indicating user input acquired by the microphone 38B to the data processing device 12. In the data processing device 12, the specific processing unit 290 acquires the audio data.
[0215] Data generation model 58 is a so-called generative AI (Artificial Intelligence). An example of data generation model 58 is ChatGPT (registered trademark) (Internet search).<URL: https: / / openai.com / blog / chatgpt> ), Gemini (registered trademark) (Internet search) <url: https: gemini.google.com ?hl="ja">Examples of generative AI include the following. The data generation model 58 is obtained by performing deep learning on a neural network. The data generation model 58 is input with prompts containing instructions, and with inference data such as audio data representing speech, text data representing text, and image data representing images. The data generation model 58 infers from the input inference data according to the instructions indicated by the prompts, and outputs the inference results in data formats such as audio data and text data. Here, inference refers to, for example, analysis, classification, prediction, and / or summarization.
[0216] In the above embodiment, an example was given in which specific processing is performed by the data processing device 12, but the technology of this disclosure is not limited thereto, and the specific processing may also be performed by the smart device 14.
[0217] [Second Embodiment]
[0218] Figure 3 shows an example of the configuration of the data processing system 210 according to the second embodiment.
[0219] As shown in Figure 3, the data processing system 210 includes a data processing device 12 and smart glasses 214. An example of the data processing device 12 is a server.
[0220] The data processing device 12 comprises a computer 22, a database 24, and a communication interface 26. The computer 22 is an example of a "computer" related to the technology of this disclosure. The computer 22 comprises a processor 28, RAM 30, and storage 32. The processor 28, RAM 30, and storage 32 are connected to a bus 34. The database 24 and the communication interface 26 are also connected to the bus 34. The communication interface 26 is connected to a network 54. An example of the network 54 is a WAN (Wide Area Network) and / or a LAN (Local Area Network).
[0221] The smart glasses 214 include a computer 36, a microphone 238, a speaker 240, a camera 42, and a communication interface 44. The computer 36 includes a processor 46, RAM 48, and storage 50. The processor 46, RAM 48, and storage 50 are connected to a bus 52. The microphone 238, speaker 240, and camera 42 are also connected to the bus 52.
[0222] The microphone 238 receives voice signals from the user 20 and receives instructions from the user 20. The microphone 238 captures the voice signals from the user 20, converts the captured voice into audio data, and outputs it to the processor 46. The speaker 240 outputs audio according to the instructions from the processor 46.
[0223] Camera 42 is a small digital camera equipped with an optical system including a lens, aperture, and shutter, and an image sensor such as a CMOS (Complementary Metal-Oxide-Semiconductor) image sensor or a CCD (Charge Coupled Device) image sensor, and captures images of the area around the user 20 (for example, an imaging range defined by a field of view equivalent to the width of a typical healthy person's field of vision).
[0224] Communication interface 44 is connected to network 54. Communication interfaces 44 and 26 are responsible for the exchange of various information between processor 46 and processor 28 via network 54. The exchange of various information between processor 46 and processor 28 using communication interfaces 44 and 26 is performed in a secure manner.
[0225] Figure 4 shows an example of the main functions of the data processing device 12 and the smart glasses 214. As shown in Figure 4, the data processing device 12 performs specific processing using the processor 28. The storage 32 stores the specific processing program 56.
[0226] The specific processing program 56 is an example of a "program" relating to the technology of this disclosure. The processor 28 reads the specific processing program 56 from the storage 32 and executes the read specific processing program 56 on the RAM 30. The specific processing is realized by the processor 28 operating as a specific processing unit 290 in accordance with the specific processing program 56 executed on the RAM 30.
[0227] The storage 32 stores the data generation model 58 and the emotion identification model 59. The data generation model 58 and the emotion identification model 59 are used by the identification processing unit 290.
[0228] In the smart glasses 214, the processor 46 performs the reception output processing. The storage 50 stores the reception output program 60. The processor 46 reads the reception output program 60 from the storage 50 and executes the read reception output program 60 on the RAM 48. The reception output processing is realized by the processor 46 operating as a control unit 46A according to the reception output program 60 executed on the RAM 48.
[0229] Next, the identification processing performed by the identification processing unit 290 of the data processing device 12 will be described. In the following description, the data processing device 12 will be referred to as the "server" and the smart glasses 214 will be referred to as the "terminal".
[0230] This invention is a system that complements the communication experience using communication devices with music. This system provides users with background music that is adapted to their emotions and circumstances through the interaction between the server, terminal, and user.
[0231] The server first securely retrieves the user's conversation history and analyzes the conversation content. Natural language processing techniques are used for the analysis to identify keywords and emotional tones. In parallel, the server retrieves the user's music listening history to understand their preferred music genres and artists. Based on both sets of data, the server uses an AI model to generate background music that best suits the conversation context and the user's musical preferences.
[0232] The generated background music is transmitted to the user's device in real time. The device automatically plays the received music, allowing the user to hear appropriate background music during conversations. This playback enhances the user's communication experience, making it more emotional and memorable.
[0233] As a concrete example, consider a scenario where a user is having a fun conversation with a friend about an event. In this case, the server extracts keywords such as "event" and "fun" from the conversation and identifies upbeat pop music that the user has enjoyed listening to in the past from their listening history. The background music generated based on this information is played on the device, adding a fun atmosphere to the conversation and allowing the user to enjoy the conversation even more.
[0234] Furthermore, based on evaluation information obtained from users, the server continuously trains its AI model to improve the accuracy of generation. In this way, the present invention can provide music tailored to individual conversational experiences, creating new experiential value for users.
[0235] The following describes the processing flow.
[0236] Step 1:
[0237] With the user's consent, the server retrieves the user's conversation history and music listening history from the database. This involves secure data transfer using the latest security protocols.
[0238] Step 2:
[0239] The server analyzes the acquired conversation history using natural language processing (NLP) techniques. This analysis identifies keywords and basic emotional tones from the conversation content. For example, it classifies emotions such as "happy," "sad," and "busy."
[0240] Step 3:
[0241] The server analyzes the user's music listening history to identify their musical preferences. This includes a process of examining play counts, genres, and artist tendencies.
[0242] Step 4:
[0243] The server uses an AI model to generate optimal background music based on the characteristics of the conversation and the user's musical preferences. In this process, musical elements corresponding to the conversational context (e.g., formal, casual) are combined.
[0244] Step 5:
[0245] The generated background music is configured to be streamed from the server to the user's device. The device immediately plays the received music, seamlessly playing it in the background of the user's conversation.
[0246] Step 6:
[0247] Users can provide feedback on the background music being played. For example, they can rate whether they like the music or not using a score. They can also record actions taken during playback, such as adjusting the volume or skipping tracks.
[0248] Step 7:
[0249] The device sends this feedback and actions back to the server, which is used to train the AI model. The server uses this feedback to update the BGM generation algorithm and improve its accuracy.
[0250] This series of steps allows users to receive a music experience optimized for their individual conversations.
[0251] (Example 1)
[0252] Next, we will describe Example 1. In the following description, the data processing device 12 will be referred to as the "server," and the smart glasses 214 will be referred to as the "terminal."
[0253] In modern communication devices, there is a growing demand for richer and more emotionally engaging communication experiences by automatically providing background music that is appropriate to the user's emotions and preferences. However, conventional technologies struggle to generate and provide background music in real time based on the content of the conversation and the user's musical preferences. This results in users not being able to fully experience music that is appropriate to the situation.
[0254] The identification process performed by the identification processing unit 290 of the data processing device 12 in Example 1 is realized by the following means.
[0255] In this invention, the server includes a function to acquire a record of the communication device user's conversation, a function to acquire the user's music history, and a function to analyze the conversation record and music history to identify the characteristics of the conversation and the user's musical preferences. This makes it possible to generate optimal background music based on the characteristics and preferences of the conversation and deliver it to the user in real time.
[0256] A "communication device" is an electronic device used to send and receive voice and data, and is a terminal that supports user communication.
[0257] A "dialogue record" is data that digitally stores the content of conversations and messages conducted through communication devices.
[0258] "Music history" refers to data on songs and music that a user has listened to in the past, and is information that indicates the user's musical preferences.
[0259] A "generative model" is an algorithm based on machine learning or artificial intelligence that automatically creates new data, especially background music, based on acquired data.
[0260] A "control statement" is a text that represents commands or instructions used as input to a generative model, and is data used to guide the generative process.
[0261] An "intelligent model" is a system that uses machine learning algorithms to continuously learn based on user feedback, thereby improving the accuracy and performance of the model.
[0262] This invention relates to a system for enhancing the dialogue experience in a communication device through background music. This system has the function of providing music appropriate to the user's emotions and the context of the conversation through the interaction between the server, terminal, and user.
[0263] The server first acquires the content of conversations taking place on the communication device and analyzes it using natural language processing techniques. Possible software used for this would include natural language processing libraries such as "spaCy" and "NLTK." The server uses these libraries to extract keywords and emotional tones from the dialogue and also retrieves the user's music history from a database. This database contains information about songs and genres the user has previously played. Based on this data, the server creates control statements to input to the generative AI model.
[0264] The generative AI model uses technologies such as "OpenAI GPT" to generate music data based on prompts created by the server. An example of such a prompt is "Emotional tone: Happy, Music genre: Pop." The generated music data is sent to the device via streaming.
[0265] The device receives music data sent from the server and plays it in the background. This playback uses the media player application installed on the device. This allows the user to naturally listen to music appropriate to the situation while having a conversation.
[0266] For example, if a user is having a conversation with a friend about a fun event, the server analyzes keywords such as "event" and "fun" and generates music based on pop music the user has liked in the past. This music is played on the device, adding an element of fun to the conversation and improving the user's experience.
[0267] Furthermore, users can provide feedback on the music played. The server collects this feedback and continuously trains its generation AI model to achieve more accurate music generation. With this configuration, the system can enrich the user's interactive experience with individually customized music.
[0268] The flow of the specific processing in Example 1 will be explained using Figure 11.
[0269] Step 1:
[0270] The server acquires dialogue data from the user's communication device. Input includes real-time ongoing conversation audio and text data. The server first converts this dialogue data into a digital format and securely acquires it using a secure communication protocol.
[0271] Step 2:
[0272] The server analyzes the acquired dialogue data using natural language processing (NLP) techniques. Input includes digital audio and text data. The server uses an NLP library to extract keywords and emotional tones from the text. This step results in outputting emotional tones such as "event" and "fun."
[0273] Step 3:
[0274] The server retrieves the user's music history from the database. Input includes the user ID and past playback history. The server uses SQL queries to retrieve listening data for songs related to the user. The output is a list of music genres and artists the user has previously enjoyed listening to.
[0275] Step 4:
[0276] The server creates prompts for the generative AI model based on the analysis results and music history. The input includes the output data from steps 2 and 3. The server constructs emotional tone and musical preferences as prompts and prepares them for input into the generative AI model. The output is the constructed prompts.
[0277] Step 5:
[0278] The server generates background music using a generative AI model. The input is a generated prompt message. Based on this, the generative AI model adjusts the parameters of the music data and generates new music. The output is the generated music data.
[0279] Step 6:
[0280] The server streams the generated music data to the user's device. The input includes music data from the generation AI model. The server converts the data into a compressed format and sends it to the device in real time. The output is the streaming music data.
[0281] Step 7:
[0282] The device decodes the received music data and plays it in the background. The input includes music data from the server. The device launches a media player and plays the music, providing the user with a real-time auditory experience. The output is the music playing from the device.
[0283] Step 8:
[0284] The user provides feedback on the played music via the terminal. The input includes the user's evaluation information regarding the suitability of the music. The user inputs the evaluation using the interface on the terminal. The output is the evaluation data sent to the server.
[0285] Step 9:
[0286] The server uses the evaluation data from the user to train the AI model. The input includes the evaluation information. The server utilizes the collected data, provides feedback to the AI model, and aims to improve the accuracy of music generation. The output is the improved AI model.
[0287] (Application Example 1)
[0288] Next, Application Example 1 will be described. In the following description, the data processing device 12 is referred to as the "server", and the smart glasses 214 are referred to as the "terminal".
[0289] In the store, there is a demand for a method to enable customers to have a more comfortable and empathy-based experience. However, in the current technology, since there is a lack of means to analyze the customer's emotions and conversation content in real time and automatically provide suitable music, it is difficult to improve the customer experience.
[0290] The specific processing by the specific processing unit 290 of the data processing device 12 in Application Example 1 is realized by the following means.
[0291] In this invention, the server includes means for acquiring the user's voice data, means for analyzing the voice data to identify topics and emotions, and means for acquiring the user's music preference data. As a result, it becomes possible to generate background music in real time according to the customer's emotions and conversation content and play it in the store.
[0292] "User voice data" refers to audio information of conversations between customers and employees within the store.
[0293] "Methods for identifying topics and emotions" refer to technologies that analyze audio data to identify what customers are saying and the tone of their emotions.
[0294] "Music preference data" refers to information about music genres and artists that a user has liked in the past.
[0295] "Means of generating background sounds" refers to technology that automatically creates music to enhance the customer experience based on identified topics, emotions, and music preference data.
[0296] An "output device" is a device used to physically reproduce the generated background sound, and includes the sound system within the store.
[0297] This invention relates to a system that acquires audio data of conversations between customers and employees in a store and generates and plays background music in real time based on that data. The server receives the audio data acquired using the microphone of smart glasses and analyzes it using natural language processing (NLP) software (e.g., NLTK, SpaCy). This analysis identifies the customer's topics of conversation and emotions. The server further refers to the customer's music preference data and generates appropriate background music using an AI model (e.g., GPT-4) based on this information.
[0298] The generated music is transmitted to the in-store speaker system, which acts as the output device, via a music streaming service API (e.g., Spotify API). This system allows customers to listen to background music that matches their emotions and conversation, resulting in a more comfortable and satisfying experience.
[0299] As a specific example, when a customer is consulting about gift ideas at a grocery store, it is determined that the tone of the conversation is bright and positive. In this case, the server can generate bright jazz music to enliven the atmosphere in the store. An example of a prompt sentence would be in the form of "The customer is selecting a gift in a buoyant mood. Please recommend the most suitable music for this." In this way, the entire store environment can be harmonized with the customer's needs, and a high-value-added service can be provided.
[0300] The flow of the specific processing in Application Example 1 will be described using FIG. 12.
[0301] Step 1:
[0302] The server receives the voice data acquired by the terminal through the microphone of the smart glasses. The received voice data includes the conversation content between the customer and the employee. The input is voice data, and the output is a digital-formatted voice file for analysis preparation.
[0303] Step 2:
[0304] The server analyzes the received voice data using a natural language processing library (e.g., NLTK, SpaCy). Through the analysis, the topic and emotion are identified. This process includes steps of converting the voice data into text and performing sentiment analysis and keyword extraction. The input is a digital-formatted voice file, and the output is the topic and emotion data identified in text form.
[0305] Step 3:
[0306] The server obtains the customer's music preference data from the database. This data is based on the customer's past music viewing history, favorite genres, and artists. The input is the customer ID or profile information, and the output is the music preference data.
[0307] Step 4:
[0308] The server generates appropriate background music using an AI model (e.g., GPT-4) based on identified topics, emotions, and music preference data. The AI model sets the context for music generation using prompts and generates the corresponding music clips. The input is topic, emotion data, and music preference data, and the output is the generated music file.
[0309] Step 5:
[0310] The server sends the generated music to the store's speaker system via a music streaming service API (e.g., Spotify API). This transmission allows background music to play in real time. The input is the generated music file, and the output is the playback of background music within the store.
[0311] Step 6:
[0312] Users listen to background music while making purchases, leading to increased satisfaction. User feedback is collected and used to improve the accuracy of subsequent music generation models. The input is user feedback data, and the output is updated music preference data and training data for the AI model.
[0313] Furthermore, an emotion engine that estimates the user's emotions may be incorporated. That is, the identification processing unit 290 may use the emotion identification model 59 to estimate the user's emotions and perform identification processing using the user's emotions.
[0314] This invention provides a system that further enhances the communication experience of users using communication devices through emotion recognition. In this system, a server, a terminal, and an emotion engine that analyzes the user's conversation history work together. The server first acquires the user's conversation history and performs analysis using natural language processing technology. Through this analysis, in addition to the characteristics of the conversation, the emotion engine recognizes the user's emotions.
[0315] The emotion engine classifies specific emotions such as joy, sadness, and anger based on the user's words and expressions. This emotion information is analyzed by the server along with the user's music listening history to more precisely identify the user's musical preferences. Based on this data, the server uses an AI model to generate background music appropriate to the emotion.
[0316] The generated background music is adjusted to match the user's emotions and sent to the device in real time. The device plays this music during the conversation, creating a space that complements the user's emotions. This feature enriches the user's emotional conversational experience and enables deeper communication.
[0317] As a concrete example, when a user's hard-worked project succeeds, the server uses an emotion engine to identify the emotion of "joy" from the conversation. It then generates upbeat, cheerful background music that emphasizes this emotion and plays it on the user's device. This allows the user to feel and experience that moment more vividly.
[0318] Furthermore, this system improves the accuracy of background music generation by continuously training its emotion engine and AI model based on user feedback. This allows it to provide a music experience that best suits the user's characteristics over time.
[0319] The following describes the processing flow.
[0320] Step 1:
[0321] The server collects conversation history data from communication devices based on the user's consent. This data is encrypted and securely transferred to the server while protecting privacy.
[0322] Step 2:
[0323] The server analyzes the collected conversation history using natural language processing algorithms. The purpose of the analysis is to extract keywords that appear in the conversation and to understand the overall atmosphere of the conversation.
[0324] Step 3:
[0325] The server uses an emotion engine to further examine the analysis results and recognize the user's emotional state (joy, sadness, tension, etc.). The emotion engine determines emotions based on contextual nuances and the frequency of use of specific words.
[0326] Step 4:
[0327] The server retrieves the user's music listening history data and analyzes the user's preferred music genres and styles. This process refers to the frequency of music playback and the trends in the playlists the user has selected.
[0328] Step 5:
[0329] The server uses an AI model based on emotions recognized from the conversation and analyzed musical preferences to generate background music optimized for the user's current emotional state.
[0330] Step 6:
[0331] The server streams the generated background music to the user's device in real time. The device plays this stream in the background, providing a pleasant musical experience that matches the user's conversation.
[0332] Step 7:
[0333] Users can send feedback on the background music being played to the server via their device. This feedback includes opinions on the music's rating and emotional coherence.
[0334] Step 8:
[0335] Based on the collected feedback, the server trains its emotion engine and AI model to further improve the accuracy of background music generation. This continuous learning process makes the music experience more personalized for the user.
[0336] (Example 2)
[0337] Next, we will describe Example 2. In the following description, the data processing device 12 will be referred to as the "server" and the smart glasses 214 will be referred to as the "terminal".
[0338] In the modern era, while communication via devices is increasing, digital communication has limitations in conveying emotions, making it difficult to build the same emotional connection as in face-to-face conversations. Therefore, there is a need to provide music experiences that respond to users' emotions and improve the quality of communication.
[0339] The identification process performed by the identification processing unit 290 of the data processing device 12 in Example 2 is realized by the following means.
[0340] In this invention, the server includes means for acquiring the user's conversation history of a communication terminal, means for acquiring the user's music playback history, and means for analyzing the conversation history using a processing method and identifying the user's emotions using an emotion engine. This makes it possible to generate appropriate background music based on the user's emotions and provide it in real time.
[0341] A "communication terminal" is an electronic device used by users to send and receive voice and data.
[0342] "Dialogue history" refers to a record of conversations that users have had through their communication devices.
[0343] "Music playback history" is a record of the music a user has listened to in the past, and is used to analyze their preferences.
[0344] "Processing methods" is a general term for the techniques and methods used to analyze data and extract information.
[0345] An "emotion engine" is a program that analyzes dialogue content to identify the user's emotional state.
[0346] An "artificial intelligence model" is a collection of algorithms that utilize technologies such as machine learning to achieve specific capabilities.
[0347] "Background music" refers to music generated to complement the user's emotional state and enhance their experience.
[0348] "Real-time" refers to a situation where processing is performed instantly with virtually no delay.
[0349] This invention is a system that enhances the user's conversational experience using a communication terminal, adding depth to digital communication by providing background music that corresponds to the emotional state. The following describes a specific form for its implementation.
[0350] The server first retrieves the user's dialogue history from the communication terminal. This is done using speech recognition technology or a text message transfer protocol. Next, it analyzes the retrieved dialogue history using natural language processing technology and identifies the user's emotions using an emotion engine. This emotion engine is a machine learning model that uses a variety of dialogue patterns registered in a database to classify emotions.
[0351] In parallel, the server retrieves and analyzes the user's music playback history. This reveals what kind of music the user preferred to listen to in the past, and what emotional states they were in. Based on the identified emotional and musical preference data, the server creates new background music using a generative AI model. This AI model incorporates a music generation algorithm that utilizes deep learning technology, and is given information such as "the user is feeling happy" as an input prompt.
[0352] The generated music is transmitted from the server to the user's communication terminal in real time. The terminal plays this music in the background during the conversation, allowing the user to experience the emotions of the moment more deeply.
[0353] For example, if a user says, "I passed my exam today," the server analyzes the conversation history to identify the emotion of "joy." Based on this, upbeat, cheerful music is generated and immediately played on the device. An example of this prompt message might be, "Generate cheerful music that expresses the user's joy."
[0354] In this way, the system can provide a music experience tailored to the user's emotions, thereby improving the quality of communication.
[0355] The flow of the specific processing in Example 2 will be explained using Figure 13.
[0356] Step 1:
[0357] The server retrieves the user's conversation history from the communication terminal. It receives text and audio data sent from the terminal as input and saves it as conversation history. Specifically, it reads the text chat log or performs speech recognition to convert audio to text. This process saves the conversation content to the database in a specific format.
[0358] Step 2:
[0359] The server uses natural language processing (NLP) techniques to analyze the acquired dialogue history. The input is the stored dialogue history, and the server processes this data to identify emotions. Specifically, it uses an NLP library to perform morphological analysis and contextual understanding, extracting keywords and emotional phrases from the conversation as output. This extracts the characteristics of the conversation and provides an indicator of the user's emotions.
[0360] Step 3:
[0361] The server uses an emotion engine to identify the user's emotions from the analysis results. The input is keywords and emotion indicators extracted in the previous step, and the output is a specific emotion category (e.g., joy, sadness, anger). Specifically, it applies an emotion classification algorithm to calculate which emotion the dialogue content corresponds to. In this step, the emotion engine can train its model using past data to improve its accuracy.
[0362] Step 4:
[0363] The server retrieves the user's music playback history and analyzes it in conjunction with emotional information. It references the user's playback history database as input and creates a profile of the user's musical preferences as output. Specifically, it queries the playback history from the database, analyzing song genres, tempos, and emotional data from past listening sessions to identify the user's unique musical tastes.
[0364] Step 5:
[0365] The server uses a generation AI model to generate background music based on identified emotions and musical preferences. The input is emotion category and music profile information, and the output is a music file generated by the AI model. Specifically, the operation involves inputting the prompt "The user's emotion is 'joy,' so please generate bright and cheerful music" into the AI model and executing the music generation algorithm.
[0366] Step 6:
[0367] The server transmits the generated music to the communication terminal in real time, and the terminal plays it. The input is the generated music file, and the output is the music playback on the terminal. Specifically, the music file is sent to the terminal in audio stream format, and the terminal plays this data using an audio output device. Through this process, the user listens to music that resonates with their emotions during a conversation, improving the quality of the conversation.
[0368] (Application Example 2)
[0369] Next, we will explain application example 2. In the following explanation, the data processing device 12 will be referred to as the "server," and the smart glasses 214 will be referred to as the "terminal."
[0370] Conventional communication devices have a problem in that they simply play music without considering the user's emotional state, making it difficult to provide a truly personalized music experience. Furthermore, there was a lack of means to instantly recognize emotions from the user's facial expressions and voice and generate and provide appropriate background music. Therefore, it was difficult to provide a music experience that would enhance the user's mental state.
[0371] The specific processing performed by the specific processing unit 290 of the data processing device 12 in Application Example 2 is realized by the following means.
[0372] In this invention, the server includes means for acquiring audio and visual data of a user of a communication device, means for analyzing the user's musical characteristics and identifying their emotional state, means for adjusting musical characteristics based on the emotional state obtained from the audio and visual data, means for generating background music using a generation AI model based on the adjusted musical characteristics, and means for transmitting and playing the generated background music to the communication device in real time. This makes it possible to provide background music that matches the user's emotional state and enrich the user's life experience.
[0373] "Communication equipment" refers to electronic devices used to send and receive voice and data, and which users use to communicate with others.
[0374] "Voice data" refers to information obtained from the user's voice, recorded in digital format for use in emotional analysis.
[0375] "Visual data" refers to information obtained from images and videos acquired using cameras and sensors, and is used to analyze the user's facial expressions and movements.
[0376] "Music characteristics" refer to information that indicates a user's musical preferences and tendencies, and are used for song selection and background music generation.
[0377] "Emotional state" refers to the state of a user's current emotions and is identified through the analysis of audio and visual data.
[0378] "Musical characteristics" refer to attributes of music such as tempo, melody, rhythm, and timbre, which are adjusted to suit the user's emotional state.
[0379] A "generative AI model" is a computer program that uses artificial intelligence and has the ability to generate new music based on input data.
[0380] "Background music" refers to music provided to complement or enhance the user's experience, and is generated to match the user's emotional state.
[0381] This invention is a system for providing a music experience tailored to the user's emotional state. The system comprises communication equipment, a server, and AI-based data analysis capabilities.
[0382] The server acquires the user's voice and visual data through communication devices such as smartphones and smart glasses. This utilizes microphones and cameras built into the devices.
[0383] The acquired data is analyzed using natural language processing technologies such as the Google Cloud Natural Language API and the Microsoft Azure Emotion API, as well as image recognition technologies, and classified as the user's emotional state. For example, it is determined whether the user's voice tone and facial expressions indicate emotions such as "happiness" or "anger."
[0384] Next, this emotional information is used to adjust the musical characteristics. Specifically, this involves selecting songs that match the user's emotions from music services such as Spotify and Apple Music, or generating new music through a generative AI model.
[0385] The generated background music is transmitted in real time to the user's communication device and played on the device. The user can experience a space where emotions are complemented through the music.
[0386] For example, if a user uses the device while feeling stressed, the server will determine that the user "needs relaxation" and select, generate, and play calming classical music.
[0387] Examples of prompts to input into a generative AI model:
[0388] "The user's current emotional state is 'relaxed'. Please choose calming music."
[0389] This system allows users to receive an optimal musical experience tailored to their individual emotional state, enabling them to experience emotional richness in their daily lives.
[0390] The flow of a specific process in Application Example 2 will be explained using Figure 14.
[0391] Step 1:
[0392] The device acquires the user's voice and visual data. This involves using the device's built-in microphone and camera to collect digital data of voice and facial expressions in real time. This data is stored for a certain period for subsequent analysis.
[0393] Step 2:
[0394] The server receives the collected audio and visual data and prepares it for data analysis. The inputs are audio and images, and the output is digital data converted from this data into an analyzable format. At this stage, preprocessing such as noise reduction and image correction is performed.
[0395] Step 3:
[0396] The server uses pre-processed digital data and leverages the Google Cloud Natural Language API and Microsoft Azure Emotion API to analyze the user's emotional state. The analysis results in the output of numerical values or labels indicating the emotional state (e.g., happy, sad, angry).
[0397] Step 4:
[0398] The server adjusts musical characteristics based on the analyzed emotional state. The input is emotional data, and the output is the corresponding musical genre, tempo, and other characteristics. This identifies a musical style that suits the user's emotions.
[0399] Step 5:
[0400] The server uses the results from the previous step to generate background music using a generative AI model. The input for this step is musical characteristics, and the output is music data generated by the AI. The following prompt is used for the generative AI model: "The user's current emotional state is 'relaxed'. Please select calming music."
[0401] Step 6:
[0402] The server sends the generated music data to the terminal. The input is the generated music data, and the output is a notification confirming transmission. The terminal receives and stores this data.
[0403] Step 7:
[0404] The device plays the received music for the user. A music playback application automatically launches on the device, providing a musical experience tailored to the user's current emotional state. The input is music data, and the output is the played sound as an audio signal.
[0405] The specific processing unit 290 transmits the result of the specific processing to the smart glasses 214. In the smart glasses 214, the control unit 46A causes the speaker 240 to output the result of the specific processing. The microphone 238 acquires audio indicating user input for the result of the specific processing. The control unit 46A transmits the audio data indicating user input acquired by the microphone 238 to the data processing unit 12. In the data processing unit 12, the specific processing unit 290 acquires the audio data.
[0406] Data generation model 58 is a type of so-called generative AI (Artificial Intelligence). One example of data generation model 58 is ChatGPT (Internet search<URL: https: / / openai.com / blog / chatgpt> ), Gemini (Internet search) <url: https: gemini.google.com ?hl="ja">Examples of generative AI include the following. The data generation model 58 is obtained by performing deep learning on a neural network. The data generation model 58 is input with prompts containing instructions, and with inference data such as audio data representing speech, text data representing text, and image data representing images. The data generation model 58 infers from the input inference data according to the instructions indicated by the prompts, and outputs the inference results in data formats such as audio data and text data. Here, inference refers to, for example, analysis, classification, prediction, and / or summarization.
[0407] In the above embodiment, an example was given in which specific processing is performed by the data processing device 12, but the technology of this disclosure is not limited thereto, and the specific processing may also be performed by the smart glasses 214.
[0408] [Third Embodiment]
[0409] Figure 5 shows an example of the configuration of the data processing system 310 according to the third embodiment.
[0410] As shown in Figure 5, the data processing system 310 includes a data processing device 12 and a headset terminal 314. An example of the data processing device 12 is a server.
[0411] The data processing device 12 comprises a computer 22, a database 24, and a communication interface 26. The computer 22 is an example of a "computer" related to the technology of this disclosure. The computer 22 comprises a processor 28, RAM 30, and storage 32. The processor 28, RAM 30, and storage 32 are connected to a bus 34. The database 24 and the communication interface 26 are also connected to the bus 34. The communication interface 26 is connected to a network 54. An example of the network 54 is a WAN (Wide Area Network) and / or a LAN (Local Area Network).
[0412] The headset terminal 314 includes a computer 36, a microphone 238, a speaker 240, a camera 42, a communication interface 44, and a display 343. The computer 36 includes a processor 46, RAM 48, and storage 50. The processor 46, RAM 48, and storage 50 are connected to a bus 52. The microphone 238, speaker 240, camera 42, and display 343 are also connected to the bus 52.
[0413] The microphone 238 receives voice signals from the user 20 and receives instructions from the user 20. The microphone 238 captures the voice signals from the user 20, converts the captured voice into audio data, and outputs it to the processor 46. The speaker 240 outputs audio according to the instructions from the processor 46.
[0414] Camera 42 is a small digital camera equipped with an optical system including a lens, aperture, and shutter, and an image sensor such as a CMOS (Complementary Metal-Oxide-Semiconductor) image sensor or a CCD (Charge Coupled Device) image sensor, and captures images of the area around the user 20 (for example, an imaging range defined by a field of view equivalent to the width of a typical healthy person's field of vision).
[0415] Communication interface 44 is connected to network 54. Communication interfaces 44 and 26 are responsible for the exchange of various information between processor 46 and processor 28 via network 54. The exchange of various information between processor 46 and processor 28 using communication interfaces 44 and 26 is performed in a secure manner.
[0416] Figure 6 shows an example of the main functions of the data processing device 12 and the headset terminal 314. As shown in Figure 6, the data processing device 12 performs specific processing using the processor 28. The storage 32 stores the specific processing program 56.
[0417] The specific processing program 56 is an example of a "program" relating to the technology of this disclosure. The processor 28 reads the specific processing program 56 from the storage 32 and executes the read specific processing program 56 on the RAM 30. The specific processing is realized by the processor 28 operating as a specific processing unit 290 in accordance with the specific processing program 56 executed on the RAM 30.
[0418] The storage 32 stores the data generation model 58 and the emotion identification model 59. The data generation model 58 and the emotion identification model 59 are used by the identification processing unit 290.
[0419] In the headset terminal 314, the processor 46 performs the reception output processing. The storage 50 stores the reception output program 60. The processor 46 reads the reception output program 60 from the storage 50 and executes the read reception output program 60 on the RAM 48. The reception output processing is realized by the processor 46 operating as a control unit 46A according to the reception output program 60 executed on the RAM 48.
[0420] Next, the specific processing performed by the specific processing unit 290 of the data processing device 12 will be described. In the following description, the data processing device 12 will be referred to as the "server" and the headset terminal 314 will be referred to as the "terminal".
[0421] This invention is a system that complements the communication experience using communication devices with music. This system provides users with background music that is adapted to their emotions and circumstances through the interaction between the server, terminal, and user.
[0422] The server first securely retrieves the user's conversation history and analyzes the conversation content. Natural language processing techniques are used for the analysis to identify keywords and emotional tones. In parallel, the server retrieves the user's music listening history to understand their preferred music genres and artists. Based on both sets of data, the server uses an AI model to generate background music that best suits the conversation context and the user's musical preferences.
[0423] The generated background music is transmitted to the user's device in real time. The device automatically plays the received music, allowing the user to hear appropriate background music during conversations. This playback enhances the user's communication experience, making it more emotional and memorable.
[0424] As a concrete example, consider a scenario where a user is having a fun conversation with a friend about an event. In this case, the server extracts keywords such as "event" and "fun" from the conversation and identifies upbeat pop music that the user has enjoyed listening to in the past from their listening history. The background music generated based on this information is played on the device, adding a fun atmosphere to the conversation and allowing the user to enjoy the conversation even more.
[0425] Furthermore, based on evaluation information obtained from users, the server continuously trains its AI model to improve the accuracy of generation. In this way, the present invention can provide music tailored to individual conversational experiences, creating new experiential value for users.
[0426] The following describes the processing flow.
[0427] Step 1:
[0428] With the user's consent, the server retrieves the user's conversation history and music listening history from the database. This involves secure data transfer using the latest security protocols.
[0429] Step 2:
[0430] The server analyzes the acquired conversation history using natural language processing (NLP) techniques. This analysis identifies keywords and basic emotional tones from the conversation content. For example, it classifies emotions such as "happy," "sad," and "busy."
[0431] Step 3:
[0432] The server analyzes the user's music listening history to identify their musical preferences. This includes a process of examining play counts, genres, and artist tendencies.
[0433] Step 4:
[0434] The server uses an AI model to generate optimal background music based on the characteristics of the conversation and the user's musical preferences. In this process, musical elements corresponding to the conversational context (e.g., formal, casual) are combined.
[0435] Step 5:
[0436] The generated background music is configured to be streamed from the server to the user's device. The device immediately plays the received music, seamlessly playing it in the background of the user's conversation.
[0437] Step 6:
[0438] Users can provide feedback on the background music being played. For example, they can rate whether they like the music or not using a score. They can also record actions taken during playback, such as adjusting the volume or skipping tracks.
[0439] Step 7:
[0440] The device sends this feedback and actions back to the server, which is used to train the AI model. The server uses this feedback to update the BGM generation algorithm and improve its accuracy.
[0441] This series of steps allows users to receive a music experience optimized for their individual conversations.
[0442] (Example 1)
[0443] Next, we will describe Example 1. In the following description, the data processing device 12 will be referred to as the "server," and the headset-type terminal 314 will be referred to as the "terminal."
[0444] In modern communication devices, there is a growing demand for richer and more emotionally engaging communication experiences by automatically providing background music that is appropriate to the user's emotions and preferences. However, conventional technologies struggle to generate and provide background music in real time based on the content of the conversation and the user's musical preferences. This results in users not being able to fully experience music that is appropriate to the situation.
[0445] The identification process performed by the identification processing unit 290 of the data processing device 12 in Example 1 is realized by the following means.
[0446] In this invention, the server includes a function to acquire a record of the communication device user's conversation, a function to acquire the user's music history, and a function to analyze the conversation record and music history to identify the characteristics of the conversation and the user's musical preferences. This makes it possible to generate optimal background music based on the characteristics and preferences of the conversation and deliver it to the user in real time.
[0447] A "communication device" is an electronic device used to send and receive voice and data, and is a terminal that supports user communication.
[0448] A "dialogue record" is data that digitally stores the content of conversations and messages conducted through communication devices.
[0449] "Music history" refers to data on songs and music that a user has listened to in the past, and is information that indicates the user's musical preferences.
[0450] A "generative model" is an algorithm based on machine learning or artificial intelligence that automatically creates new data, especially background music, based on acquired data.
[0451] A "control statement" is a text that represents commands or instructions used as input to a generative model, and is data used to guide the generative process.
[0452] An "intelligent model" is a system that uses machine learning algorithms to continuously learn based on user feedback, thereby improving the accuracy and performance of the model.
[0453] This invention relates to a system for enhancing the dialogue experience in a communication device with background music. This system has the function of providing music appropriate to the user's emotions and the context of the conversation through the interaction between the server, terminal, and user.
[0454] The server first acquires the content of conversations taking place on the communication device and analyzes it using natural language processing techniques. Possible software used for this would include natural language processing libraries such as "spaCy" and "NLTK." The server uses these libraries to extract keywords and emotional tones from the dialogue and also retrieves the user's music history from a database. This database contains information about songs and genres the user has previously played. Based on this data, the server creates control statements to input to the generative AI model.
[0455] The generative AI model uses technologies such as "OpenAI GPT" to generate music data based on prompts created by the server. An example of such a prompt is "Emotional tone: Happy, Music genre: Pop." The generated music data is sent to the device via streaming.
[0456] The device receives music data sent from the server and plays it in the background. This playback uses the media player application installed on the device. This allows the user to naturally listen to music appropriate to the situation while having a conversation.
[0457] For example, if a user is having a conversation with a friend about a fun event, the server analyzes keywords such as "event" and "fun" and generates music based on pop music the user has liked in the past. This music is played on the device, adding an element of fun to the conversation and improving the user's experience.
[0458] Furthermore, users can provide feedback on the music played. The server collects this feedback and continuously trains its generation AI model to achieve more accurate music generation. With this configuration, the system can enrich the user's interactive experience with individually customized music.
[0459] The flow of the specific processing in Example 1 will be explained using Figure 11.
[0460] Step 1:
[0461] The server acquires dialogue data from the user's communication device. Input includes real-time ongoing conversation audio and text data. The server first converts this dialogue data into a digital format and securely acquires it using a secure communication protocol.
[0462] Step 2:
[0463] The server analyzes the acquired dialogue data using natural language processing (NLP) techniques. Input includes digital audio and text data. The server uses an NLP library to extract keywords and emotional tones from the text. This step results in outputting emotional tones such as "event" and "fun."
[0464] Step 3:
[0465] The server retrieves the user's music history from the database. Input includes the user ID and past playback history. The server uses SQL queries to retrieve listening data for songs related to the user. The output is a list of music genres and artists the user has previously enjoyed listening to.
[0466] Step 4:
[0467] The server creates prompts for the generative AI model based on the analysis results and music history. The input includes the output data from steps 2 and 3. The server constructs emotional tone and musical preferences as prompts and prepares them for input into the generative AI model. The output is the constructed prompts.
[0468] Step 5:
[0469] The server generates background music using a generative AI model. The input is a generated prompt message. Based on this, the generative AI model adjusts the parameters of the music data and generates new music. The output is the generated music data.
[0470] Step 6:
[0471] The server streams the generated music data to the user's device. The input includes music data from the generation AI model. The server converts the data into a compressed format and sends it to the device in real time. The output is the streaming music data.
[0472] Step 7:
[0473] The device decodes the received music data and plays it in the background. The input includes music data from the server. The device launches a media player and plays the music, providing the user with a real-time auditory experience. The output is the music playing from the device.
[0474] Step 8:
[0475] Users provide feedback on the played music via their device. Input includes user ratings regarding the music's suitability. Users input their ratings using an interface on their device. Output is the rating data sent to the server.
[0476] Step 9:
[0477] The server trains a generative AI model using user evaluation data. The input includes evaluation information. The server utilizes the collected data to provide feedback to the AI model, aiming to improve the accuracy of music generation. The output is the improved AI model.
[0478] (Application Example 1)
[0479] Next, we will explain Application Example 1. In the following explanation, the data processing device 12 will be referred to as the "server," and the headset-type terminal 314 will be referred to as the "terminal."
[0480] There is a need for ways to provide customers with a more comfortable and emotionally resonant experience within stores. However, current technology lacks the means to analyze customer emotions and conversation content in real time and automatically provide music that is appropriate for that, making it difficult to improve the customer experience.
[0481] The specific processing performed by the specific processing unit 290 of the data processing device 12 in Application Example 1 is realized by the following means.
[0482] In this invention, the server includes means for acquiring user voice data, means for analyzing the voice data to identify topics and emotions, and means for acquiring user music preference data. This makes it possible to generate background sounds in real time that correspond to the customer's emotions and conversation content, and to play them back in the store.
[0483] "User voice data" refers to audio information of conversations between customers and employees within the store.
[0484] "Methods for identifying topics and emotions" refer to technologies that analyze audio data to identify what customers are saying and the tone of their emotions.
[0485] "Music preference data" refers to information about music genres and artists that a user has liked in the past.
[0486] "Means of generating background sounds" refers to technology that automatically creates music to enhance the customer experience based on identified topics, emotions, and music preference data.
[0487] An "output device" is a device used to physically reproduce the generated background sound, and includes the sound system within the store.
[0488] This invention relates to a system that acquires audio data of conversations between customers and employees in a store and generates and plays background music in real time based on that data. The server receives the audio data acquired using the microphone of smart glasses and analyzes it using natural language processing (NLP) software (e.g., NLTK, SpaCy). This analysis identifies the customer's topics of conversation and emotions. The server further refers to the customer's music preference data and generates appropriate background music using an AI model (e.g., GPT-4) based on this information.
[0489] The generated music is transmitted to the in-store speaker system, which acts as the output device, via a music streaming service API (e.g., Spotify API). This system allows customers to listen to background music that matches their emotions and conversation, resulting in a more comfortable and satisfying experience.
[0490] As a concrete example, if a customer in a general store is discussing gift ideas and the tone of their conversation is identified as bright and positive, the server can generate upbeat jazz music to liven up the store's atmosphere. An example of a prompt might be, "The customer is in a cheerful mood while choosing a gift. Please recommend music that would be perfect for this." In this way, the entire store environment can harmonize with the customer's needs, enabling the provision of high-value-added services.
[0491] The flow of a specific process in Application Example 1 will be explained using Figure 12.
[0492] Step 1:
[0493] The server receives audio data acquired by the device through the smart glasses' microphone. The received audio data includes the content of conversations between customers and employees. The input is audio data, and the output is an audio file in digital format for analysis preparation.
[0494] Step 2:
[0495] The server analyzes the received audio data using natural language processing libraries (e.g., NLTK, SpaCy). This analysis identifies the topic and sentiment. This process includes converting the audio data to text and performing sentiment analysis and keyword extraction. The input is a digital audio file, and the output is text-based topic and sentiment data.
[0496] Step 3:
[0497] The server retrieves customer music preference data from a database. This data is based on the customer's past music listening history, favorite genres, and artists. The input is the customer ID or profile information, and the output is music preference data.
[0498] Step 4:
[0499] The server generates appropriate background music using an AI model (e.g., GPT-4) based on identified topics, emotions, and music preference data. The AI model sets the context for music generation using prompts and generates the corresponding music clips. The input is topic, emotion data, and music preference data, and the output is the generated music file.
[0500] Step 5:
[0501] The server sends the generated music to the store's speaker system via a music streaming service API (e.g., Spotify API). This transmission allows background music to play in real time. The input is the generated music file, and the output is the playback of background music within the store.
[0502] Step 6:
[0503] Users listen to background music while making purchases, leading to increased satisfaction. User feedback is collected and used to improve the accuracy of subsequent music generation models. The input is user feedback data, and the output is updated music preference data and training data for the AI model.
[0504] Furthermore, an emotion engine that estimates the user's emotions may be incorporated. That is, the identification processing unit 290 may use the emotion identification model 59 to estimate the user's emotions and perform identification processing using the user's emotions.
[0505] This invention provides a system that further enhances the communication experience of users using communication devices through emotion recognition. In this system, a server, a terminal, and an emotion engine that analyzes the user's conversation history work together. The server first acquires the user's conversation history and performs analysis using natural language processing technology. Through this analysis, in addition to the characteristics of the conversation, the emotion engine recognizes the user's emotions.
[0506] The emotion engine classifies specific emotions such as joy, sadness, and anger based on the user's words and expressions. This emotion information is analyzed by the server along with the user's music listening history to more precisely identify the user's musical preferences. Based on this data, the server uses an AI model to generate background music appropriate to the emotion.
[0507] The generated background music is adjusted to match the user's emotions and sent to the device in real time. The device plays this music during the conversation, creating a space that complements the user's emotions. This feature enriches the user's emotional conversational experience and enables deeper communication.
[0508] As a concrete example, when a user's hard-worked project succeeds, the server uses an emotion engine to identify the emotion of "joy" from the conversation. It then generates upbeat, cheerful background music that emphasizes this emotion and plays it on the user's device. This allows the user to feel and experience that moment more vividly.
[0509] Furthermore, this system improves the accuracy of background music generation by continuously training its emotion engine and AI model based on user feedback. This allows it to provide a music experience that best suits the user's characteristics over time.
[0510] The following describes the processing flow.
[0511] Step 1:
[0512] The server collects conversation history data from communication devices based on the user's consent. This data is encrypted and securely transferred to the server while protecting privacy.
[0513] Step 2:
[0514] The server analyzes the collected conversation history using natural language processing algorithms. The purpose of the analysis is to extract keywords that appear in the conversation and to understand the overall atmosphere of the conversation.
[0515] Step 3:
[0516] The server uses an emotion engine to further examine the analysis results and recognize the user's emotional state (joy, sadness, tension, etc.). The emotion engine determines emotions based on contextual nuances and the frequency of use of specific words.
[0517] Step 4:
[0518] The server retrieves the user's music listening history data and analyzes the user's preferred music genres and styles. This process refers to the frequency of music playback and the trends in the playlists the user has selected.
[0519] Step 5:
[0520] The server uses an AI model based on emotions recognized from the conversation and analyzed musical preferences to generate background music optimized for the user's current emotional state.
[0521] Step 6:
[0522] The server streams the generated background music to the user's device in real time. The device plays this stream in the background, providing a pleasant musical experience that matches the user's conversation.
[0523] Step 7:
[0524] Users can send feedback on the background music being played to the server via their device. This feedback includes opinions on the music's rating and emotional coherence.
[0525] Step 8:
[0526] Based on the collected feedback, the server trains its emotion engine and AI model to further improve the accuracy of background music generation. This continuous learning process makes the music experience more personalized for the user.
[0527] (Example 2)
[0528] Next, we will describe Example 2. In the following description, the data processing device 12 will be referred to as the "server," and the headset-type terminal 314 will be referred to as the "terminal."
[0529] In the modern era, while communication via devices is increasing, digital communication has limitations in conveying emotions, making it difficult to build the same emotional connection as in face-to-face conversations. Therefore, there is a need to provide music experiences that respond to users' emotions and improve the quality of communication.
[0530] The identification process performed by the identification processing unit 290 of the data processing device 12 in Example 2 is realized by the following means.
[0531] In this invention, the server includes means for acquiring the user's conversation history of a communication terminal, means for acquiring the user's music playback history, and means for analyzing the conversation history using a processing method and identifying the user's emotions using an emotion engine. This makes it possible to generate appropriate background music based on the user's emotions and provide it in real time.
[0532] A "communication terminal" is an electronic device used by users to send and receive voice and data.
[0533] "Dialogue history" refers to a record of conversations that users have had through their communication devices.
[0534] "Music playback history" is a record of the music a user has listened to in the past, and is used to analyze their preferences.
[0535] "Processing methods" is a general term for the techniques and methods used to analyze data and extract information.
[0536] An "emotion engine" is a program that analyzes dialogue content to identify the user's emotional state.
[0537] An "artificial intelligence model" is a collection of algorithms that utilize technologies such as machine learning to achieve specific capabilities.
[0538] "Background music" refers to music generated to complement the user's emotional state and enhance their experience.
[0539] "Real-time" refers to a situation where processing is performed instantly with virtually no delay.
[0540] This invention is a system that enhances the user's conversational experience using a communication terminal, adding depth to digital communication by providing background music that corresponds to the emotional state. The following describes a specific form for its implementation.
[0541] The server first retrieves the user's dialogue history from the communication terminal. This is done using speech recognition technology or a text message transfer protocol. Next, it analyzes the retrieved dialogue history using natural language processing technology and identifies the user's emotions using an emotion engine. This emotion engine is a machine learning model that uses a variety of dialogue patterns registered in a database to classify emotions.
[0542] In parallel, the server retrieves and analyzes the user's music playback history. This reveals what kind of music the user preferred to listen to in the past, and what emotional states they were in. Based on the identified emotional and musical preference data, the server creates new background music using a generative AI model. This AI model incorporates a music generation algorithm that utilizes deep learning technology, and is given information such as "the user is feeling happy" as an input prompt.
[0543] The generated music is transmitted from the server to the user's communication terminal in real time. The terminal plays this music in the background during the conversation, allowing the user to experience the emotions of the moment more deeply.
[0544] For example, if a user says, "I passed my exam today," the server analyzes the conversation history to identify the emotion of "joy." Based on this, upbeat, cheerful music is generated and immediately played on the device. An example of this prompt message might be, "Generate cheerful music that expresses the user's joy."
[0545] In this way, the system can provide a music experience tailored to the user's emotions, thereby improving the quality of communication.
[0546] The flow of the specific processing in Example 2 will be explained using Figure 13.
[0547] Step 1:
[0548] The server retrieves the user's conversation history from the communication terminal. It receives text and audio data sent from the terminal as input and saves it as conversation history. Specifically, it reads the text chat log or performs speech recognition to convert audio to text. This process saves the conversation content to the database in a specific format.
[0549] Step 2:
[0550] The server uses natural language processing (NLP) techniques to analyze the acquired dialogue history. The input is the stored dialogue history, and the server processes this data to identify emotions. Specifically, it uses an NLP library to perform morphological analysis and contextual understanding, extracting keywords and emotional phrases from the conversation as output. This extracts the characteristics of the conversation and provides an indicator of the user's emotions.
[0551] Step 3:
[0552] The server uses an emotion engine to identify the user's emotions from the analysis results. The input is keywords and emotion indicators extracted in the previous step, and the output is a specific emotion category (e.g., joy, sadness, anger). Specifically, it applies an emotion classification algorithm to calculate which emotion the dialogue content corresponds to. In this step, the emotion engine can train its model using past data to improve its accuracy.
[0553] Step 4:
[0554] The server retrieves the user's music playback history and analyzes it in conjunction with emotional information. It references the user's playback history database as input and creates a profile of the user's musical preferences as output. Specifically, it queries the playback history from the database, analyzing song genres, tempos, and emotional data from past listening sessions to identify the user's unique musical tastes.
[0555] Step 5:
[0556] The server uses a generation AI model to generate background music based on identified emotions and musical preferences. The input is emotion category and music profile information, and the output is a music file generated by the AI model. Specifically, the operation involves inputting the prompt "The user's emotion is 'joy,' so please generate bright and cheerful music" into the AI model and executing the music generation algorithm.
[0557] Step 6:
[0558] The server transmits the generated music to the communication terminal in real time, and the terminal plays it. The input is the generated music file, and the output is the music playback on the terminal. Specifically, the music file is sent to the terminal in audio stream format, and the terminal plays this data using an audio output device. Through this process, the user listens to music that resonates with their emotions during a conversation, improving the quality of the conversation.
[0559] (Application Example 2)
[0560] Next, we will explain application example 2. In the following explanation, the data processing device 12 will be referred to as the "server," and the headset-type terminal 314 will be referred to as the "terminal."
[0561] Conventional communication devices have a problem in that they simply play music without considering the user's emotional state, making it difficult to provide a truly personalized music experience. Furthermore, there was a lack of means to instantly recognize emotions from the user's facial expressions and voice and generate and provide appropriate background music. Therefore, it was difficult to provide a music experience that would enhance the user's mental state.
[0562] The specific processing performed by the specific processing unit 290 of the data processing device 12 in Application Example 2 is realized by the following means.
[0563] In this invention, the server includes means for acquiring audio and visual data of a user of a communication device, means for analyzing the user's musical characteristics and identifying their emotional state, means for adjusting musical characteristics based on the emotional state obtained from the audio and visual data, means for generating background music using a generation AI model based on the adjusted musical characteristics, and means for transmitting and playing the generated background music to the communication device in real time. This makes it possible to provide background music that matches the user's emotional state and enrich the user's life experience.
[0564] "Communication equipment" refers to electronic devices used to send and receive voice and data, and which users use to communicate with others.
[0565] "Voice data" refers to information obtained from the user's voice, recorded in digital format for use in emotional analysis.
[0566] "Visual data" refers to information obtained from images and videos acquired using cameras and sensors, and is used to analyze the user's facial expressions and movements.
[0567] "Music characteristics" refer to information that indicates a user's musical preferences and tendencies, and are used for song selection and background music generation.
[0568] "Emotional state" refers to the state of a user's current emotions and is identified through the analysis of audio and visual data.
[0569] "Musical characteristics" refer to attributes of music such as tempo, melody, rhythm, and timbre, which are adjusted to suit the user's emotional state.
[0570] A "generative AI model" is a computer program that uses artificial intelligence and has the ability to generate new music based on input data.
[0571] "Background music" refers to music provided to complement or enhance the user's experience, and is generated to match the user's emotional state.
[0572] This invention is a system for providing a music experience tailored to the user's emotional state. The system comprises communication equipment, a server, and AI-based data analysis capabilities.
[0573] The server acquires the user's voice and visual data through communication devices such as smartphones and smart glasses. This utilizes microphones and cameras built into the devices.
[0574] The acquired data is analyzed using natural language processing technologies such as the Google Cloud Natural Language API and the Microsoft Azure Emotion API, as well as image recognition technologies, and classified as the user's emotional state. For example, it is determined whether the user's voice tone and facial expressions indicate emotions such as "happiness" or "anger."
[0575] Next, this emotional information is used to adjust the musical characteristics. Specifically, this involves selecting songs that match the user's emotions from music services such as Spotify and Apple Music, or generating new music through a generative AI model.
[0576] The generated background music is transmitted in real time to the user's communication device and played on the device. The user can experience a space where emotions are complemented through the music.
[0577] For example, if a user uses the device while feeling stressed, the server will determine that the user "needs relaxation" and select, generate, and play calming classical music.
[0578] Examples of prompts to input into a generative AI model:
[0579] "The user's current emotional state is 'relaxed'. Please choose calming music."
[0580] This system allows users to receive an optimal musical experience tailored to their individual emotional state, enabling them to experience emotional richness in their daily lives.
[0581] The flow of a specific process in Application Example 2 will be explained using Figure 14.
[0582] Step 1:
[0583] The device acquires the user's voice and visual data. This involves using the device's built-in microphone and camera to collect digital data of voice and facial expressions in real time. This data is stored for a certain period for subsequent analysis.
[0584] Step 2:
[0585] The server receives the collected audio and visual data and prepares it for data analysis. The inputs are audio and images, and the output is digital data converted from this data into an analyzable format. At this stage, preprocessing such as noise reduction and image correction is performed.
[0586] Step 3:
[0587] The server uses pre-processed digital data and leverages the Google Cloud Natural Language API and Microsoft Azure Emotion API to analyze the user's emotional state. The analysis results in the output of numerical values or labels indicating the emotional state (e.g., happy, sad, angry).
[0588] Step 4:
[0589] The server adjusts musical characteristics based on the analyzed emotional state. The input is emotional data, and the output is the corresponding musical genre, tempo, and other characteristics. This identifies a musical style that suits the user's emotions.
[0590] Step 5:
[0591] The server uses the results from the previous step to generate background music using a generative AI model. The input for this step is musical characteristics, and the output is music data generated by the AI. The following prompt is used for the generative AI model: "The user's current emotional state is 'relaxed'. Please select calming music."
[0592] Step 6:
[0593] The server sends the generated music data to the terminal. The input is the generated music data, and the output is a notification confirming transmission. The terminal receives and stores this data.
[0594] Step 7:
[0595] The device plays the received music for the user. A music playback application automatically launches on the device, providing a musical experience tailored to the user's current emotional state. The input is music data, and the output is the played sound as an audio signal.
[0596] The specific processing unit 290 transmits the result of the specific processing to the headset terminal 314. In the headset terminal 314, the control unit 46A causes the speaker 240 and display 343 to output the result of the specific processing. The microphone 238 acquires audio indicating user input for the result of the specific processing. The control unit 46A transmits the audio data indicating user input acquired by the microphone 238 to the data processing unit 12. In the data processing unit 12, the specific processing unit 290 acquires the audio data.
[0597] Data generation model 58 is a type of so-called generative AI (Artificial Intelligence). One example of data generation model 58 is ChatGPT (Internet search<URL: https: / / openai.com / blog / chatgpt> ), Gemini (Internet search) <url: https: gemini.google.com ?hl="ja">Examples of generative AI include the following. The data generation model 58 is obtained by performing deep learning on a neural network. The data generation model 58 is input with prompts containing instructions, and with inference data such as audio data representing speech, text data representing text, and image data representing images. The data generation model 58 infers from the input inference data according to the instructions indicated by the prompts, and outputs the inference results in data formats such as audio data and text data. Here, inference refers to, for example, analysis, classification, prediction, and / or summarization.
[0598] In the above embodiment, an example was given in which specific processing is performed by the data processing device 12, but the technology of this disclosure is not limited thereto, and specific processing may also be performed by the headset terminal 314.
[0599] [Fourth Embodiment]
[0600] Figure 7 shows an example of the configuration of the data processing system 410 according to the fourth embodiment.
[0601] As shown in Figure 7, the data processing system 410 includes a data processing device 12 and a robot 414. An example of the data processing device 12 is a server.
[0602] The data processing device 12 comprises a computer 22, a database 24, and a communication interface 26. The computer 22 is an example of a "computer" related to the technology of this disclosure. The computer 22 comprises a processor 28, RAM 30, and storage 32. The processor 28, RAM 30, and storage 32 are connected to a bus 34. The database 24 and the communication interface 26 are also connected to the bus 34. The communication interface 26 is connected to a network 54. An example of the network 54 is a WAN (Wide Area Network) and / or a LAN (Local Area Network).
[0603] The robot 414 includes a computer 36, a microphone 238, a speaker 240, a camera 42, a communication interface 44, and a controlled object 443. The computer 36 includes a processor 46, RAM 48, and storage 50. The processor 46, RAM 48, and storage 50 are connected to a bus 52. The microphone 238, speaker 240, camera 42, and controlled object 443 are also connected to the bus 52.
[0604] The microphone 238 receives voice signals from the user 20 and receives instructions from the user 20. The microphone 238 captures the voice signals from the user 20, converts the captured voice into audio data, and outputs it to the processor 46. The speaker 240 outputs audio according to the instructions from the processor 46.
[0605] Camera 42 is a small digital camera equipped with an optical system including a lens, aperture, and shutter, and an image sensor such as a CMOS (Complementary Metal-Oxide-Semiconductor) image sensor or a CCD (Charge Coupled Device) image sensor, and captures images of the area around the user 20 (for example, an imaging range defined by a field of view equivalent to the width of a typical healthy person's field of vision).
[0606] Communication interface 44 is connected to network 54. Communication interfaces 44 and 26 are responsible for the exchange of various information between processor 46 and processor 28 via network 54. The exchange of various information between processor 46 and processor 28 using communication interfaces 44 and 26 is performed in a secure manner.
[0607] The controlled object 443 includes a display device, LEDs in the eyes, and motors that drive the arms, hands, and feet. The posture and gestures of the robot 414 are controlled by controlling the motors of the arms, hands, and feet. Some of the robot 414's emotions can be expressed by controlling these motors. Furthermore, the robot 414's facial expressions can also be expressed by controlling the illumination state of the LEDs in its eyes.
[0608] Figure 8 shows an example of the main functions of the data processing device 12 and the robot 414. As shown in Figure 8, the data processing device 12 performs specific processing using the processor 28. The storage 32 stores the specific processing program 56.
[0609] The specific processing program 56 is an example of a "program" relating to the technology of this disclosure. The processor 28 reads the specific processing program 56 from the storage 32 and executes the read specific processing program 56 on the RAM 30. The specific processing is realized by the processor 28 operating as a specific processing unit 290 in accordance with the specific processing program 56 executed on the RAM 30.
[0610] The storage 32 stores the data generation model 58 and the emotion identification model 59. The data generation model 58 and the emotion identification model 59 are used by the identification processing unit 290.
[0611] In robot 414, the processor 46 performs the reception output processing. The storage 50 stores the reception output program 60. The processor 46 reads the reception output program 60 from the storage 50 and executes the read reception output program 60 on the RAM 48. The reception output processing is realized by the processor 46 operating as a control unit 46A according to the reception output program 60 executed on the RAM 48.
[0612] Next, the specific processing performed by the specific processing unit 290 of the data processing device 12 will be described. In the following description, the data processing device 12 will be referred to as the "server" and the robot 414 as the "terminal".
[0613] This invention is a system that complements the communication experience using communication devices with music. This system provides users with background music that is adapted to their emotions and circumstances through the interaction between the server, terminal, and user.
[0614] The server first securely retrieves the user's conversation history and analyzes the conversation content. Natural language processing techniques are used for the analysis to identify keywords and emotional tones. In parallel, the server retrieves the user's music listening history to understand their preferred music genres and artists. Based on both sets of data, the server uses an AI model to generate background music that best suits the conversation context and the user's musical preferences.
[0615] The generated background music is transmitted to the user's device in real time. The device automatically plays the received music, allowing the user to hear appropriate background music during conversations. This playback enhances the user's communication experience, making it more emotional and memorable.
[0616] As a concrete example, consider a scenario where a user is having a fun conversation with a friend about an event. In this case, the server extracts keywords such as "event" and "fun" from the conversation and identifies upbeat pop music that the user has enjoyed listening to in the past from their listening history. The background music generated based on this information is played on the device, adding a fun atmosphere to the conversation and allowing the user to enjoy the conversation even more.
[0617] Furthermore, based on evaluation information obtained from users, the server continuously trains its AI model to improve the accuracy of generation. In this way, the present invention can provide music tailored to individual conversational experiences, creating new experiential value for users.
[0618] The following describes the processing flow.
[0619] Step 1:
[0620] With the user's consent, the server retrieves the user's conversation history and music listening history from the database. This involves secure data transfer using the latest security protocols.
[0621] Step 2:
[0622] The server analyzes the acquired conversation history using natural language processing (NLP) techniques. This analysis identifies keywords and basic emotional tones from the conversation content. For example, it classifies emotions such as "happy," "sad," and "busy."
[0623] Step 3:
[0624] The server analyzes the user's music listening history to identify their musical preferences. This includes a process of examining play counts, genres, and artist tendencies.
[0625] Step 4:
[0626] The server uses an AI model to generate optimal background music based on the characteristics of the conversation and the user's musical preferences. In this process, musical elements corresponding to the conversational context (e.g., formal, casual) are combined.
[0627] Step 5:
[0628] The generated background music is configured to be streamed from the server to the user's device. The device immediately plays the received music, seamlessly playing it in the background of the user's conversation.
[0629] Step 6:
[0630] Users can provide feedback on the background music being played. For example, they can rate whether they like the music or not using a score. They can also record actions taken during playback, such as adjusting the volume or skipping tracks.
[0631] Step 7:
[0632] The device sends this feedback and actions back to the server, which is used to train the AI model. The server uses this feedback to update the BGM generation algorithm and improve its accuracy.
[0633] This series of steps allows users to receive a music experience optimized for their individual conversations.
[0634] (Example 1)
[0635] Next, we will describe Example 1. In the following description, the data processing device 12 will be referred to as the "server" and the robot 414 as the "terminal".
[0636] In modern communication devices, there is a growing demand for richer and more emotionally engaging communication experiences by automatically providing background music that is appropriate to the user's emotions and preferences. However, conventional technologies struggle to generate and provide background music in real time based on the content of the conversation and the user's musical preferences. This results in users not being able to fully experience music that is appropriate to the situation.
[0637] The identification process performed by the identification processing unit 290 of the data processing device 12 in Example 1 is realized by the following means.
[0638] In this invention, the server includes a function to acquire a record of the communication device user's conversation, a function to acquire the user's music history, and a function to analyze the conversation record and music history to identify the characteristics of the conversation and the user's musical preferences. This makes it possible to generate optimal background music based on the characteristics and preferences of the conversation and deliver it to the user in real time.
[0639] A "communication device" is an electronic device used to send and receive voice and data, and is a terminal that supports user communication.
[0640] A "dialogue record" is data that digitally stores the content of conversations and messages conducted through communication devices.
[0641] "Music history" refers to data on songs and music that a user has listened to in the past, and is information that indicates the user's musical preferences.
[0642] A "generative model" is an algorithm based on machine learning or artificial intelligence that automatically creates new data, especially background music, based on acquired data.
[0643] A "control statement" is a text that represents commands or instructions used as input to a generative model, and is data used to guide the generative process.
[0644] An "intelligent model" is a system that uses machine learning algorithms to continuously learn based on user feedback, thereby improving the accuracy and performance of the model.
[0645] This invention relates to a system for enhancing the dialogue experience in a communication device through background music. This system has the function of providing music appropriate to the user's emotions and the context of the conversation through the interaction between the server, terminal, and user.
[0646] The server first acquires the content of conversations taking place on the communication device and analyzes it using natural language processing techniques. Possible software used for this would include natural language processing libraries such as "spaCy" and "NLTK." The server uses these libraries to extract keywords and emotional tones from the dialogue and also retrieves the user's music history from a database. This database contains information about songs and genres the user has previously played. Based on this data, the server creates control statements to input to the generative AI model.
[0647] The generative AI model uses technologies such as "OpenAI GPT" to generate music data based on prompts created by the server. An example of such a prompt is "Emotional tone: Happy, Music genre: Pop." The generated music data is sent to the device via streaming.
[0648] The device receives music data sent from the server and plays it in the background. This playback uses the media player application installed on the device. This allows the user to naturally listen to music appropriate to the situation while having a conversation.
[0649] For example, if a user is having a conversation with a friend about a fun event, the server analyzes keywords such as "event" and "fun" and generates music based on pop music the user has liked in the past. This music is played on the device, adding an element of fun to the conversation and improving the user's experience.
[0650] Furthermore, users can provide feedback on the music played. The server collects this feedback and continuously trains its generation AI model to achieve more accurate music generation. With this configuration, the system can enrich the user's interactive experience with individually customized music.
[0651] The flow of the specific processing in Example 1 will be explained using Figure 11.
[0652] Step 1:
[0653] The server acquires dialogue data from the user's communication device. Input includes real-time ongoing conversation audio and text data. The server first converts this dialogue data into a digital format and securely acquires it using a secure communication protocol.
[0654] Step 2:
[0655] The server analyzes the acquired dialogue data using natural language processing (NLP) techniques. Input includes digital audio and text data. The server uses an NLP library to extract keywords and emotional tones from the text. This step results in outputting emotional tones such as "event" and "fun."
[0656] Step 3:
[0657] The server retrieves the user's music history from the database. Input includes the user ID and past playback history. The server uses SQL queries to retrieve listening data for songs related to the user. The output is a list of music genres and artists the user has previously enjoyed listening to.
[0658] Step 4:
[0659] The server creates prompts for the generative AI model based on the analysis results and music history. The input includes the output data from steps 2 and 3. The server constructs emotional tone and musical preferences as prompts and prepares them for input into the generative AI model. The output is the constructed prompts.
[0660] Step 5:
[0661] The server generates background music using a generative AI model. The input is a generated prompt message. Based on this, the generative AI model adjusts the parameters of the music data and generates new music. The output is the generated music data.
[0662] Step 6:
[0663] The server streams the generated music data to the user's device. The input includes music data from the generation AI model. The server converts the data into a compressed format and sends it to the device in real time. The output is the streaming music data.
[0664] Step 7:
[0665] The device decodes the received music data and plays it in the background. The input includes music data from the server. The device launches a media player and plays the music, providing the user with a real-time auditory experience. The output is the music playing from the device.
[0666] Step 8:
[0667] Users provide feedback on the played music via their device. Input includes user ratings regarding the music's suitability. Users input their ratings using an interface on their device. Output is the rating data sent to the server.
[0668] Step 9:
[0669] The server trains a generative AI model using user evaluation data. The input includes evaluation information. The server utilizes the collected data to provide feedback to the AI model, aiming to improve the accuracy of music generation. The output is the improved AI model.
[0670] (Application Example 1)
[0671] Next, we will explain Application Example 1. In the following explanation, the data processing device 12 will be referred to as the "server" and the robot 414 as the "terminal".
[0672] There is a need for ways to provide customers with a more comfortable and emotionally resonant experience within stores. However, current technology lacks the means to analyze customer emotions and conversation content in real time and automatically provide music that is appropriate for that, making it difficult to improve the customer experience.
[0673] The specific processing performed by the specific processing unit 290 of the data processing device 12 in Application Example 1 is realized by the following means.
[0674] In this invention, the server includes means for acquiring user voice data, means for analyzing the voice data to identify topics and emotions, and means for acquiring user music preference data. This makes it possible to generate background sounds in real time that correspond to the customer's emotions and conversation content, and to play them back in the store.
[0675] "User voice data" refers to audio information of conversations between customers and employees within the store.
[0676] "Methods for identifying topics and emotions" refer to technologies that analyze audio data to identify what customers are saying and the tone of their emotions.
[0677] "Music preference data" refers to information about music genres and artists that a user has liked in the past.
[0678] "Means of generating background sounds" refers to technology that automatically creates music to enhance the customer experience based on identified topics, emotions, and music preference data.
[0679] An "output device" is a device used to physically reproduce the generated background sound, and includes the sound system within the store.
[0680] This invention relates to a system that acquires audio data of conversations between customers and employees in a store and generates and plays background music in real time based on that data. The server receives the audio data acquired using the microphone of smart glasses and analyzes it using natural language processing (NLP) software (e.g., NLTK, SpaCy). This analysis identifies the customer's topics of conversation and emotions. The server further refers to the customer's music preference data and generates appropriate background music using an AI model (e.g., GPT-4) based on this information.
[0681] The generated music is transmitted to the in-store speaker system, which acts as the output device, via a music streaming service API (e.g., Spotify API). This system allows customers to listen to background music that matches their emotions and conversation, resulting in a more comfortable and satisfying experience.
[0682] As a concrete example, if a customer in a general store is discussing gift ideas and the tone of their conversation is identified as bright and positive, the server can generate upbeat jazz music to liven up the store's atmosphere. An example of a prompt might be, "The customer is in a cheerful mood while choosing a gift. Please recommend music that would be perfect for this." In this way, the entire store environment can harmonize with the customer's needs, enabling the provision of high-value-added services.
[0683] The flow of a specific process in Application Example 1 will be explained using Figure 12.
[0684] Step 1:
[0685] The server receives audio data acquired by the device through the smart glasses' microphone. The received audio data includes the content of conversations between customers and employees. The input is audio data, and the output is an audio file in digital format for analysis preparation.
[0686] Step 2:
[0687] The server analyzes the received audio data using natural language processing libraries (e.g., NLTK, SpaCy). This analysis identifies the topic and sentiment. This process includes converting the audio data to text and performing sentiment analysis and keyword extraction. The input is a digital audio file, and the output is text-based topic and sentiment data.
[0688] Step 3:
[0689] The server retrieves customer music preference data from a database. This data is based on the customer's past music listening history, favorite genres, and artists. The input is the customer ID or profile information, and the output is music preference data.
[0690] Step 4:
[0691] The server generates appropriate background music using an AI model (e.g., GPT-4) based on identified topics, emotions, and music preference data. The AI model sets the context for music generation using prompts and generates the corresponding music clips. The input is topic, emotion data, and music preference data, and the output is the generated music file.
[0692] Step 5:
[0693] The server sends the generated music to the store's speaker system via a music streaming service API (e.g., Spotify API). This transmission allows background music to play in real time. The input is the generated music file, and the output is the playback of background music within the store.
[0694] Step 6:
[0695] Users listen to background music while making purchases, leading to increased satisfaction. User feedback is collected and used to improve the accuracy of subsequent music generation models. The input is user feedback data, and the output is updated music preference data and training data for the AI model.
[0696] Furthermore, an emotion engine that estimates the user's emotions may be incorporated. That is, the identification processing unit 290 may use the emotion identification model 59 to estimate the user's emotions and perform identification processing using the user's emotions.
[0697] This invention provides a system that further enhances the communication experience of users using communication devices through emotion recognition. In this system, a server, a terminal, and an emotion engine that analyzes the user's conversation history work together. The server first acquires the user's conversation history and performs analysis using natural language processing technology. Through this analysis, in addition to the characteristics of the conversation, the emotion engine recognizes the user's emotions.
[0698] The emotion engine classifies specific emotions such as joy, sadness, and anger based on the user's words and expressions. This emotion information is analyzed by the server along with the user's music listening history to more precisely identify the user's musical preferences. Based on this data, the server uses an AI model to generate background music appropriate to the emotion.
[0699] The generated background music is adjusted to match the user's emotions and sent to the device in real time. The device plays this music during the conversation, creating a space that complements the user's emotions. This feature enriches the user's emotional conversational experience and enables deeper communication.
[0700] As a concrete example, when a user's hard-worked project succeeds, the server uses an emotion engine to identify the emotion of "joy" from the conversation. It then generates upbeat, cheerful background music that emphasizes this emotion and plays it on the user's device. This allows the user to feel and experience that moment more vividly.
[0701] Furthermore, this system improves the accuracy of background music generation by continuously training its emotion engine and AI model based on user feedback. This allows it to provide a music experience that best suits the user's characteristics over time.
[0702] The following describes the processing flow.
[0703] Step 1:
[0704] The server collects conversation history data from communication devices based on the user's consent. This data is encrypted and securely transferred to the server while protecting privacy.
[0705] Step 2:
[0706] The server analyzes the collected conversation history using natural language processing algorithms. The purpose of the analysis is to extract keywords that appear in the conversation and to understand the overall atmosphere of the conversation.
[0707] Step 3:
[0708] The server uses an emotion engine to further examine the analysis results and recognize the user's emotional state (joy, sadness, tension, etc.). The emotion engine determines emotions based on contextual nuances and the frequency of use of specific words.
[0709] Step 4:
[0710] The server retrieves the user's music listening history data and analyzes the user's preferred music genres and styles. This process refers to the frequency of music playback and the trends in the playlists the user has selected.
[0711] Step 5:
[0712] The server uses an AI model based on emotions recognized from the conversation and analyzed musical preferences to generate background music optimized for the user's current emotional state.
[0713] Step 6:
[0714] The server streams the generated background music to the user's device in real time. The device plays this stream in the background, providing a pleasant musical experience that matches the user's conversation.
[0715] Step 7:
[0716] Users can send feedback on the background music being played to the server via their device. This feedback includes opinions on the music's rating and emotional coherence.
[0717] Step 8:
[0718] Based on the collected feedback, the server trains its emotion engine and AI model to further improve the accuracy of background music generation. This continuous learning process makes the music experience more personalized for the user.
[0719] (Example 2)
[0720] Next, we will describe Example 2. In the following description, the data processing device 12 will be referred to as the "server" and the robot 414 as the "terminal".
[0721] In the modern era, while communication via devices is increasing, digital communication has limitations in conveying emotions, making it difficult to build the same emotional connection as in face-to-face conversations. Therefore, there is a need to provide music experiences that respond to users' emotions and improve the quality of communication.
[0722] The identification process performed by the identification processing unit 290 of the data processing device 12 in Example 2 is realized by the following means.
[0723] In this invention, the server includes means for acquiring the user's conversation history of a communication terminal, means for acquiring the user's music playback history, and means for analyzing the conversation history using a processing method and identifying the user's emotions using an emotion engine. This makes it possible to generate appropriate background music based on the user's emotions and provide it in real time.
[0724] A "communication terminal" is an electronic device used by users to send and receive voice and data.
[0725] "Dialogue history" refers to a record of conversations that users have had through their communication devices.
[0726] "Music playback history" is a record of the music a user has listened to in the past, and is used to analyze their preferences.
[0727] "Processing methods" is a general term for the techniques and methods used to analyze data and extract information.
[0728] An "emotion engine" is a program that analyzes dialogue content to identify the user's emotional state.
[0729] An "artificial intelligence model" is a collection of algorithms that utilize technologies such as machine learning to achieve specific capabilities.
[0730] "Background music" refers to music generated to complement the user's emotional state and enhance their experience.
[0731] "Real-time" refers to a situation where processing is performed instantly with virtually no delay.
[0732] This invention is a system that enhances the user's conversational experience using a communication terminal, adding depth to digital communication by providing background music that corresponds to the emotional state. The following describes a specific form for its implementation.
[0733] The server first retrieves the user's dialogue history from the communication terminal. This is done using speech recognition technology or a text message transfer protocol. Next, it analyzes the retrieved dialogue history using natural language processing technology and identifies the user's emotions using an emotion engine. This emotion engine is a machine learning model that uses a variety of dialogue patterns registered in a database to classify emotions.
[0734] In parallel, the server retrieves and analyzes the user's music playback history. This reveals what kind of music the user preferred to listen to in the past, and what emotional states they were in. Based on the identified emotional and musical preference data, the server creates new background music using a generative AI model. This AI model incorporates a music generation algorithm that utilizes deep learning technology, and is given information such as "the user is feeling happy" as an input prompt.
[0735] The generated music is transmitted from the server to the user's communication terminal in real time. The terminal plays this music in the background during the conversation, allowing the user to experience the emotions of the moment more deeply.
[0736] For example, if a user says, "I passed my exam today," the server analyzes the conversation history to identify the emotion of "joy." Based on this, upbeat, cheerful music is generated and immediately played on the device. An example of this prompt message might be, "Generate cheerful music that expresses the user's joy."
[0737] In this way, the system can provide a music experience tailored to the user's emotions, thereby improving the quality of communication.
[0738] The flow of the specific processing in Example 2 will be explained using Figure 13.
[0739] Step 1:
[0740] The server retrieves the user's conversation history from the communication terminal. It receives text and audio data sent from the terminal as input and saves it as conversation history. Specifically, it reads the text chat log or performs speech recognition to convert audio to text. This process saves the conversation content to the database in a specific format.
[0741] Step 2:
[0742] The server uses natural language processing (NLP) techniques to analyze the acquired dialogue history. The input is the stored dialogue history, and the server processes this data to identify emotions. Specifically, it uses an NLP library to perform morphological analysis and contextual understanding, extracting keywords and emotional phrases from the conversation as output. This extracts the characteristics of the conversation and provides an indicator of the user's emotions.
[0743] Step 3:
[0744] The server uses an emotion engine to identify the user's emotions from the analysis results. The input is keywords and emotion indicators extracted in the previous step, and the output is a specific emotion category (e.g., joy, sadness, anger). Specifically, it applies an emotion classification algorithm to calculate which emotion the dialogue content corresponds to. In this step, the emotion engine can train its model using past data to improve its accuracy.
[0745] Step 4:
[0746] The server retrieves the user's music playback history and analyzes it in conjunction with emotional information. It references the user's playback history database as input and creates a profile of the user's musical preferences as output. Specifically, it queries the playback history from the database, analyzing song genres, tempos, and emotional data from past listening sessions to identify the user's unique musical tastes.
[0747] Step 5:
[0748] The server uses a generation AI model to generate background music based on identified emotions and musical preferences. The input is emotion category and music profile information, and the output is a music file generated by the AI model. Specifically, the operation involves inputting the prompt "The user's emotion is 'joy,' so please generate bright and cheerful music" into the AI model and executing the music generation algorithm.
[0749] Step 6:
[0750] The server transmits the generated music to the communication terminal in real time, and the terminal plays it. The input is the generated music file, and the output is the music playback on the terminal. Specifically, the music file is sent to the terminal in audio stream format, and the terminal plays this data using an audio output device. Through this process, the user listens to music that resonates with their emotions during a conversation, improving the quality of the conversation.
[0751] (Application Example 2)
[0752] Next, we will explain application example 2. In the following explanation, the data processing device 12 will be referred to as the "server" and the robot 414 as the "terminal".
[0753] Conventional communication devices have a problem in that they simply play music without considering the user's emotional state, making it difficult to provide a truly personalized music experience. Furthermore, there was a lack of means to instantly recognize emotions from the user's facial expressions and voice and generate and provide appropriate background music. Therefore, it was difficult to provide a music experience that would enhance the user's mental state.
[0754] The specific processing performed by the specific processing unit 290 of the data processing device 12 in Application Example 2 is realized by the following means.
[0755] In this invention, the server includes means for acquiring audio and visual data of a user of a communication device, means for analyzing the user's musical characteristics and identifying their emotional state, means for adjusting musical characteristics based on the emotional state obtained from the audio and visual data, means for generating background music using a generation AI model based on the adjusted musical characteristics, and means for transmitting and playing the generated background music to the communication device in real time. This makes it possible to provide background music that matches the user's emotional state and enrich the user's life experience.
[0756] "Communication equipment" refers to electronic devices used to send and receive voice and data, and which users use to communicate with others.
[0757] "Voice data" refers to information obtained from the user's voice, recorded in digital format for use in emotional analysis.
[0758] "Visual data" refers to information obtained from images and videos acquired using cameras and sensors, and is used to analyze the user's facial expressions and movements.
[0759] "Music characteristics" refer to information that indicates a user's musical preferences and tendencies, and are used for song selection and background music generation.
[0760] "Emotional state" refers to the state of a user's current emotions and is identified through the analysis of audio and visual data.
[0761] "Musical characteristics" refer to attributes of music such as tempo, melody, rhythm, and timbre, which are adjusted to suit the user's emotional state.
[0762] A "generative AI model" is a computer program that uses artificial intelligence and has the ability to generate new music based on input data.
[0763] "Background music" refers to music provided to complement or enhance the user's experience, and is generated to match the user's emotional state.
[0764] This invention is a system for providing a music experience tailored to the user's emotional state. The system comprises communication equipment, a server, and AI-based data analysis capabilities.
[0765] The server acquires the user's voice and visual data through communication devices such as smartphones and smart glasses. This utilizes microphones and cameras built into the devices.
[0766] The acquired data is analyzed using natural language processing technologies such as the Google Cloud Natural Language API and the Microsoft Azure Emotion API, as well as image recognition technologies, and classified as the user's emotional state. For example, it is determined whether the user's voice tone and facial expressions indicate emotions such as "happiness" or "anger."
[0767] Next, this emotional information is used to adjust the musical characteristics. Specifically, this involves selecting songs that match the user's emotions from music services such as Spotify and Apple Music, or generating new music through a generative AI model.
[0768] The generated background music is transmitted in real time to the user's communication device and played on the device. The user can experience a space where emotions are complemented through the music.
[0769] For example, if a user uses the device while feeling stressed, the server will determine that the user "needs relaxation" and select, generate, and play calming classical music.
[0770] Examples of prompts to input into a generative AI model:
[0771] "The user's current emotional state is 'relaxed'. Please choose calming music."
[0772] This system allows users to receive an optimal musical experience tailored to their individual emotional state, enabling them to experience emotional richness in their daily lives.
[0773] The flow of a specific process in Application Example 2 will be explained using Figure 14.
[0774] Step 1:
[0775] The device acquires the user's voice and visual data. This involves using the device's built-in microphone and camera to collect digital data of voice and facial expressions in real time. This data is stored for a certain period for subsequent analysis.
[0776] Step 2:
[0777] The server receives the collected audio and visual data and prepares it for data analysis. The inputs are audio and images, and the output is digital data converted from this data into an analyzable format. At this stage, preprocessing such as noise reduction and image correction is performed.
[0778] Step 3:
[0779] The server uses pre-processed digital data and leverages the Google Cloud Natural Language API and Microsoft Azure Emotion API to analyze the user's emotional state. The analysis results in the output of numerical values or labels indicating the emotional state (e.g., happy, sad, angry).
[0780] Step 4:
[0781] The server adjusts musical characteristics based on the analyzed emotional state. The input is emotional data, and the output is the corresponding musical genre, tempo, and other characteristics. This identifies a musical style that suits the user's emotions.
[0782] Step 5:
[0783] The server uses the results from the previous step to generate background music using a generative AI model. The input for this step is musical characteristics, and the output is music data generated by the AI. The following prompt is used for the generative AI model: "The user's current emotional state is 'relaxed'. Please select calming music."
[0784] Step 6:
[0785] The server sends the generated music data to the terminal. The input is the generated music data, and the output is a notification confirming transmission. The terminal receives and stores this data.
[0786] Step 7:
[0787] The device plays the received music for the user. A music playback application automatically launches on the device, providing a musical experience tailored to the user's current emotional state. The input is music data, and the output is the played sound as an audio signal.
[0788] The specific processing unit 290 transmits the result of the specific processing to the robot 414. In the robot 414, the control unit 46A causes the speaker 240 and the controlled object 443 to output the result of the specific processing. The microphone 238 acquires audio indicating user input for the result of the specific processing. The control unit 46A transmits the audio data indicating user input acquired by the microphone 238 to the data processing unit 12. In the data processing unit 12, the specific processing unit 290 acquires the audio data.
[0789] Data generation model 58 is a type of so-called generative AI (Artificial Intelligence). One example of data generation model 58 is ChatGPT (Internet search<URL: https: / / openai.com / blog / chatgpt> ), Gemini (Internet search) <url: https: gemini.google.com ?hl="ja">Examples of generative AI include the following. The data generation model 58 is obtained by performing deep learning on a neural network. The data generation model 58 is input with prompts containing instructions, and with inference data such as audio data representing speech, text data representing text, and image data representing images. The data generation model 58 infers from the input inference data according to the instructions indicated by the prompts, and outputs the inference results in data formats such as audio data and text data. Here, inference refers to, for example, analysis, classification, prediction, and / or summarization.
[0790] In the above embodiment, an example was given in which the specific processing is performed by the data processing device 12, but the technology of this disclosure is not limited thereto, and the specific processing may also be performed by the robot 414.
[0791] Furthermore, the emotion identification model 59, acting as an emotion engine, may determine the user's emotion according to a specific mapping. Specifically, the emotion identification model 59 may determine the user's emotion according to a specific mapping, which is an emotion map (see Figure 9). Similarly, the emotion identification model 59 may also determine the robot's emotion, and the identification processing unit 290 may perform identification processing using the robot's emotion.
[0792] Figure 9 shows an emotion map 400 in which multiple emotions are mapped. In the emotion map 400, emotions are arranged in concentric circles radiating from the center. The closer to the center of the concentric circles, the more primitive the emotions are located. Further out of the concentric circles, emotions representing states and actions arising from mental states are located. Emotion is a concept that includes feelings and mental states. On the left side of the concentric circles, emotions that are generally generated from reactions occurring in the brain are located. On the right side of the concentric circles, emotions that are generally induced by situational judgment are located. Above and below the concentric circles, emotions that are generally generated from reactions occurring in the brain and induced by situational judgment are located. In addition, the emotion of "pleasure" is located on the upper side of the concentric circles, and the emotion of "displeasure" is located on the lower side. Thus, in the emotion map 400, multiple emotions are mapped based on the structure in which emotions arise, and emotions that are likely to occur simultaneously are mapped close together.
[0793] These emotions are distributed at the 3 o'clock position on the Emotion Map 400, and usually fluctuate between feelings of security and anxiety. In the right half of the Emotion Map 400, situational awareness takes precedence over internal feelings, resulting in a calm impression.
[0794] The inside of the Emotion Map 400 represents inner thoughts, while the outside represents actions. Therefore, the further you go from the outside of the Emotion Map 400, the more visible (expressed in actions) your emotions become.
[0795] Here, human emotions are based on various balances, such as posture and blood sugar levels. When these balances deviate from the ideal, it results in discomfort, and when they approach the ideal, it results in pleasure. Similarly, in robots, cars, motorcycles, etc., emotions can be created based on various balances, such as posture and battery level. When these balances deviate from the ideal, it results in discomfort, and when they approach the ideal, it results in pleasure. The emotion map can be generated, for example, based on Dr. Mitsuyoshi's emotion map (Research on a system for analyzing brain physiological signals of speech emotion recognition and emotion, Tokushima University, doctoral dissertation: https: / / ci.nii.ac.jp / naid / 500000375379). The left half of the emotion map contains emotions belonging to a region called "response," where sensation is dominant. The right half of the emotion map contains emotions belonging to a region called "situation," where situational awareness is dominant.
[0796] The emotion map defines two emotions that promote learning. One is the emotion around the middle of the negative "repentance" and "reflection" on the situation side. In other words, it is when the robot experiences negative emotions such as "I never want to feel this way again" or "I don't want to be scolded again." The other is the emotion around the positive "desire" on the reaction side. In other words, it is when the robot has positive feelings such as "I want more" or "I want to know more."
[0797] The emotion identification model 59 inputs user input into a pre-trained neural network, obtains emotion values representing each emotion shown in the emotion map 400, and determines the user's emotion. This neural network is pre-trained based on multiple training data sets, which are combinations of user input and emotion values representing each emotion shown in the emotion map 400. Furthermore, this neural network is trained so that emotions located close together have similar values, as shown in the emotion map 900 in Figure 10. Figure 10 shows an example where multiple emotions such as "reassured," "calm," and "confident" have similar emotion values.
[0798] The above description primarily focuses on the functions of the data processing device 12 in relation to this disclosure. However, the system related to this disclosure is not necessarily implemented on a server. The system related to this disclosure may be implemented as a general information processing system. This disclosure may be implemented, for example, as a software program that runs on a personal computer or as an application that runs on a smartphone. The method related to this disclosure may be provided to users in SaaS (Software as a Service) format.
[0799] In the above embodiment, an example was given in which a specific process is performed by a single computer 22. However, the technology of this disclosure is not limited thereto, and a distributed processing of the specific process may be performed by multiple computers, including computer 22. For example, a data generation model 58 may be provided in an external device of the data processing device 12, and the external device may generate data according to the input data.
[0800] In the above embodiment, an example was given in which the specific processing program 56 is stored in the storage 32, but the technology of this disclosure is not limited thereto. For example, the specific processing program 56 may be stored in a portable, computer-readable, non-temporary storage medium such as a USB (Universal Serial Bus) memory. The specific processing program 56 stored in the non-temporary storage medium is installed in the computer 22 of the data processing device 12. The processor 28 executes specific processing according to the specific processing program 56.
[0801] Alternatively, the specific processing program 56 may be stored in a storage device such as a server connected to the data processing device 12 via the network 54, and the specific processing program 56 may be downloaded and installed on the computer 22 in response to a request from the data processing device 12.
[0802] Furthermore, it is not necessary to store the entirety of the specific processing program 56 in a storage device such as a server connected to the data processing device 12 via the network 54, or to store the entirety of the specific processing program 56 in the storage 32; it is acceptable to store only a portion of the specific processing program 56.
[0803] The following types of processors can be used as hardware resources to perform specific processing. Examples of processors include a CPU, a general-purpose processor that functions as a hardware resource to perform specific processing by executing software, i.e., a program. Other examples of processors include dedicated electrical circuits, such as FPGAs (Field-Programmable Gate Arrays), PLDs (Programmable Logic Devices), or ASICs (Application Specific Integrated Circuits), which have circuit configurations specifically designed to perform specific processing. All of these processors have built-in or connected memory, and all of them perform specific processing by using memory.
[0804] The hardware resource that performs a specific process may consist of one of these various processors, or it may consist of a combination of two or more processors of the same or different types (for example, a combination of multiple FPGAs, or a combination of a CPU and an FPGA). Alternatively, the hardware resource that performs a specific process may consist of a single processor.
[0805] Examples of configurations using a single processor include, firstly, a configuration in which one or more CPUs and software are combined to form a single processor, and this processor functions as a hardware resource that performs a specific process. Secondly, there is a configuration using a processor that realizes the functions of the entire system, including multiple hardware resources that perform a specific process, on a single IC chip, as exemplified by SoCs (System-on-a-chip). In this way, a specific process is realized using one or more of the above types of processors as hardware resources.
[0806] Furthermore, the hardware structure of these various processors can more specifically utilize electrical circuits that combine circuit elements such as semiconductor devices. Also, the specific processing described above is merely an example. Therefore, it goes without saying that unnecessary steps can be deleted, new steps added, or the processing order rearranged, as long as it does not deviate from the main purpose.
[0807] The descriptions and illustrations presented above are detailed explanations of the technical aspects of this disclosure and are merely examples of the technical aspects. For example, the above descriptions of the structure, function, operation, and effect are examples of the structure, function, operation, and effect of the technical aspects of this disclosure. Therefore, it goes without saying that you may delete unnecessary parts, add new elements, or replace elements in the descriptions and illustrations presented above, as long as you do not deviate from the essence of the technical aspects of this disclosure. Furthermore, in order to avoid confusion and facilitate understanding of the technical aspects of this disclosure, explanations of common technical knowledge and the like that do not require special explanation to enable the implementation of the technical aspects of this disclosure have been omitted from the descriptions and illustrations presented above.
[0808] All documents, patent applications, and technical standards described herein are incorporated by reference to the same extent as if each individual document, patent application, and technical standard were specifically and individually noted as being incorporated by reference.
[0809] The following is further disclosed regarding the embodiments described above.
[0810] (Claim 1)
[0811] A means of obtaining the conversation history of a user of a communication device,
[0812] A means of obtaining a user's music listening history,
[0813] A means for analyzing the aforementioned conversation history and the aforementioned music listening history to identify the characteristics of the conversation and the user's musical preferences,
[0814] Means for generating background music based on the identified characteristics of the conversation and musical preferences,
[0815] A means for transmitting and playing the generated background music to the communication device,
[0816] A system that includes this.
[0817] (Claim 2)
[0818] The system according to claim 1, further comprising means for optimizing the generation means based on evaluation information from the user.
[0819] (Claim 3)
[0820] The system according to claim 1, further comprising means for training an artificial intelligence model based on the evaluation information and improving the accuracy of the generation means.
[0821] "Example 1"
[0822] (Claim 1)
[0823] A function to acquire a record of the conversations of users of the communication device,
[0824] A function to retrieve the user's music history,
[0825] The function analyzes the aforementioned dialogue records and music history to identify the characteristics of the dialogue and the user's musical preferences.
[0826] Based on the identified dialogue characteristics and musical preferences, a function is provided to create background music using a generative model.
[0827] The function includes distributing the created background music to the communication device and playing it back.
[0828] A function for generating control statements to be input to the aforementioned generation model,
[0829] A system that includes this.
[0830] (Claim 2)
[0831] The system according to claim 1, which includes a function to improve the creation function based on evaluation information from the user.
[0832] (Claim 3)
[0833] The system according to claim 1, which includes a function to learn an intelligent model based on the evaluation information and improve the accuracy of the creation function.
[0834] "Application Example 1"
[0835] (Claim 1)
[0836] Means for acquiring user voice data,
[0837] A means for analyzing the aforementioned audio data to identify topics and emotions,
[0838] A means of obtaining user music preference data,
[0839] A means for generating background sounds based on the aforementioned topic, emotion, and music preference data,
[0840] A means for transmitting the generated background sound to an output device and playing it back,
[0841] A system that includes this.
[0842] (Claim 2)
[0843] The system according to claim 1, comprising means for adjusting the generating means based on the means for obtaining.
[0844] (Claim 3)
[0845] The system according to claim 1, further comprising means for learning an intelligent model based on the adjustment information and improving the accuracy of the generation means.
[0846] "Example 2 of combining an emotion engine"
[0847] (Claim 1)
[0848] A means of obtaining the conversation history of a communication terminal user,
[0849] A means of obtaining a user's music playback history,
[0850] A means for analyzing the aforementioned dialogue history using a processing method and identifying the user's emotions using an emotion engine,
[0851] A means for clarifying the user's musical preferences based on the aforementioned emotional information and the aforementioned music playback history,
[0852] A means of generating background music appropriate to a specific emotion using an artificial intelligence model,
[0853] A means for transmitting and playing the generated background music to the communication terminal in real time,
[0854] A system that includes this.
[0855] (Claim 2)
[0856] The system according to claim 1, which optimizes the generation means and the emotion engine based on the response information from the user.
[0857] (Claim 3)
[0858] The system according to claim 1, wherein the intelligent model is adjusted using the aforementioned reaction information, and the accuracy of the generation means is improved.
[0859] "Application example 2 when combining with an emotional engine"
[0860] (Claim 1)
[0861] Means for acquiring voice data and visual data of users of communication devices,
[0862] A means of analyzing the user's musical characteristics and identifying their emotional state,
[0863] Means for adjusting musical characteristics based on emotional states obtained from the aforementioned audio and visual data,
[0864] A means for generating background music using a generative AI model based on the adjusted musical characteristics,
[0865] A means for transmitting and playing the generated background music to the communication device in real time,
[0866] A system that includes this.
[0867] (Claim 2)
[0868] The system according to claim 1, which optimizes the means for generating based on evaluation data from the aforementioned users.
[0869] (Claim 3)
[0870] The system according to claim 1, which trains an artificial intelligence model based on the evaluation data and improves the accuracy of the generating means. [Explanation of Symbols]
[0871] 10, 210, 310, 410 Data Processing Systems 12 Data Processing Devices 14 Smart Devices 214 Smart Glasses 314 Headset-type terminal 414 Robots< / url:> < / url:> < / url:> < / url:>
Claims
1. A means of obtaining the conversation history of a user of a communication device, A means of obtaining a user's music listening history, A means for analyzing the aforementioned conversation history and the aforementioned music listening history to identify the characteristics of the conversation and the user's musical preferences, Means for generating background music based on the identified characteristics of the conversation and musical preferences, A means for transmitting and playing the generated background music to the communication device, A system that includes this.
2. The system according to claim 1, further comprising means for optimizing the generation means based on evaluation information from the user.
3. The system according to claim 1, further comprising means for training an artificial intelligence model based on the evaluation information and improving the accuracy of the generation means.
Citation Information
Patent Citations
Persona chatbot control method and system
JP2022180282A