system

A system that analyzes biometric data to generate personalized music in real time addresses the limitations of existing music services by providing emotional state-responsive music for improved psychological well-being.

JP2026070110APending Publication Date: 2026-04-27SOFTBANK GROUP CORP
View PDF 1 Cites 0 Cited by

Patent Information

Authority / Receiving Office
JP · JP
Patent Type
Applications
Current Assignee / Owner
SOFTBANK GROUP CORP
Filing Date
2024-10-15
Publication Date
2026-04-27

AI Technical Summary

Technical Problem

Current music services fail to respond to a user's real-time emotional state, limiting their effectiveness in improving psychological well-being.

Method used

A system that acquires biometric information in real time, analyzes the user's emotional state, and generates personalized music tailored to their emotional state using generative AI models.

Benefits of technology

Provides a real-time, personalized music experience that supports psychological well-being by addressing the user's emotional needs.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure 2026070110000001_ABST
    Figure 2026070110000001_ABST
Patent Text Reader

Abstract

We provide the system. [Solution] A means of collecting information to acquire the user's biometric information, An emotion analysis means for analyzing the user's emotional state based on the aforementioned biometric information, A music generation means that generates music data according to the aforementioned emotional state, Music provision means for providing the generated music data to a user device, A system that includes this.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The technology of the present disclosure relates to a system.

Background Art

[0002] Patent Document 1 discloses a persona chatbot control method performed by at least one processor, including steps of receiving a user utterance, adding the user utterance to a prompt including an instruction sentence related to an explanation of a chatbot character, encoding the prompt, and inputting the encoded prompt into a language model to generate a chatbot utterance as a response to the user utterance.

Prior Art Documents

Patent Documents

[0003]

Patent Document 1

Summary of the Invention

Problems to be Solved by the Invention

[0004] In modern society, the increase in stress and the disruption of the life rhythm have a great impact on an individual's mental health. Therefore, means for improving these problems and providing appropriate support according to an individual's emotional state are required. However, current music services are generally one-way and cannot respond to a user's real-time emotional state, so there is a problem that the user's psychological well-being cannot be effectively improved.

Means for Solving the Problems

[0005] To address these challenges, the present invention provides a system that acquires a user's biometric information in real time and analyzes the user's emotional state based on that information. Specifically, it performs emotional analysis based on the acquired biometric and voice information, generates appropriate music data according to the analysis results, and provides it to the user's device. This makes it possible to provide personalized music that matches the user's emotional state in real time, thereby supporting their psychological well-being.

[0006] "Information gathering means" refers to a device or function for acquiring a user's biometric information.

[0007] "Emotional analysis means" refers to algorithms and devices that determine a user's emotional state based on acquired biometric information.

[0008] "Music generation means" refers to algorithms and devices for creating appropriate music data according to the analyzed emotional state.

[0009] "Music delivery means" refers to the means of delivering generated music data to users, that is, devices or functions for streaming or data transfer.

[0010] "Biometric information" refers to data about the user's body, such as heart rate and body temperature.

[0011] "Voice data" refers to information about the user's voice, including acoustic characteristics used for emotion analysis.

[0012] A "generative opposing network" is a deep learning technique used for data generation, and it is a neural network used to generate data that closely resembles reality.

[0013] A "recurrent neural network" is a type of neural network that excels at processing time-series data and is used to learn recurring patterns in the generation of music data. [Brief explanation of the drawing]

[0014] [Figure 1] This is a conceptual diagram showing an example of the configuration of a data processing system according to the first embodiment. [Figure 2] This is a conceptual diagram showing an example of the essential functions of a data processing device and a smart device according to the first embodiment. [Figure 3] This is a conceptual diagram showing an example of the configuration of a data processing system according to the second embodiment. [Figure 4] This is a conceptual diagram showing an example of the main functions of a data processing device and smart glasses according to the second embodiment. [Figure 5] This is a conceptual diagram showing an example of the configuration of a data processing system according to the third embodiment. [Figure 6] This is a conceptual diagram showing an example of the main functions of a data processing device and a headset-type terminal according to the third embodiment. [Figure 7] This is a conceptual diagram showing an example of the configuration of a data processing system according to the fourth embodiment. [Figure 8] This is a conceptual diagram showing an example of the main functions of a data processing device and a robot according to the fourth embodiment. [Figure 9] This shows an emotion map where multiple emotions are mapped. [Figure 10] This shows an emotion map where multiple emotions are mapped. [Figure 11] This is a sequence diagram showing the processing flow of the data processing system in Example 1. [Figure 12] This is a sequence diagram showing the processing flow of the data processing system in Application Example 1. [Figure 13] This is a sequence diagram showing the processing flow of the data processing system in Example 2, which incorporates an emotion engine. [Figure 14] This is a sequence diagram showing the processing flow of the data processing system in Application Example 2, which combines an emotion engine. [Modes for carrying out the invention]

[0015] Hereinafter, an example of an embodiment of a system according to the technology of the present disclosure will be described with reference to the accompanying drawings.

[0016] First, the terms used in the following description will be explained.

[0017] In the following embodiments, the numbered processor (hereinafter simply referred to as "processor") may be a single arithmetic unit or a combination of multiple arithmetic units. Also, the processor may be a single type of arithmetic unit or a combination of multiple types of arithmetic units. Examples of arithmetic units include a CPU (Central Processing Unit), a GPU (Graphics Processing Unit), a GPGPU (General-Purpose computing on Graphics Processing Units), an APU (Accelerated Processing Unit), and the like.

[0018] In the following embodiments, the numbered RAM (Random Access Memory) is a memory in which information is temporarily stored and is used as a work memory by the processor.

[0019] In the following embodiments, the numbered storage is one or more non-volatile storage devices that store various programs and various parameters, etc. Examples of non-volatile storage devices include flash memory (SSD (Solid State Drive)), magnetic disks (e.g., hard disks), or magnetic tapes, etc.

[0020] In the following embodiments, the signed communication interface (I / F) is an interface that includes a communication processor and an antenna, etc. The communication interface manages communication between multiple computers. Examples of communication standards applicable to the communication interface include wireless communication standards such as 5G (5th Generation Mobile Communication System), Wi-Fi (registered trademark), or Bluetooth (registered trademark).

[0021] In the following embodiments, "A and / or B" is synonymous with "at least one of A and B." That is, "A and / or B" means that it may be A alone, or B alone, or a combination of A and B. Furthermore, in this specification, the same concept as "A and / or B" applies when expressing three or more things linked by "and / or."

[0022] [First Embodiment]

[0023] Figure 1 shows an example of the configuration of the data processing system 10 according to the first embodiment.

[0024] As shown in Figure 1, the data processing system 10 includes a data processing device 12 and a smart device 14. An example of the data processing device 12 is a server.

[0025] The data processing device 12 comprises a computer 22, a database 24, and a communication interface 26. The computer 22 is an example of a "computer" related to the technology of this disclosure. The computer 22 comprises a processor 28, RAM 30, and storage 32. The processor 28, RAM 30, and storage 32 are connected to a bus 34. The database 24 and the communication interface 26 are also connected to the bus 34. The communication interface 26 is connected to a network 54. An example of the network 54 is a WAN (Wide Area Network) and / or a LAN (Local Area Network).

[0026] The smart device 14 comprises a computer 36, a reception device 38, an output device 40, a camera 42, and a communication interface 44. The computer 36 comprises a processor 46, RAM 48, and storage 50. The processor 46, RAM 48, and storage 50 are connected to a bus 52. The reception device 38, output device 40, and camera 42 are also connected to the bus 52.

[0027] The reception device 38 is equipped with a touch panel 38A and a microphone 38B, etc., and receives user input. The touch panel 38A receives user input by detecting contact with an object (e.g., a pen or finger). The microphone 38B receives user input by detecting the user's voice. The control unit 46A transmits data indicating the user input received by the touch panel 38A and microphone 38B to the data processing device 12. In the data processing device 12, the specific processing unit 290 acquires the data indicating the user input.

[0028] The output device 40 includes a display 40A and a speaker 40B, and presents data to the user 20 by outputting the data in a form perceptible to the user 20 (e.g., audio and / or text). The display 40A displays visible information such as text and images according to instructions from the processor 46. The speaker 40B outputs audio according to instructions from the processor 46. The camera 42 is a small digital camera equipped with an optical system such as a lens, aperture, and shutter, and an image sensor such as a CMOS (Complementary Metal-Oxide-Semiconductor) image sensor or a CCD (Charge Coupled Device) image sensor.

[0029] Communication interface 44 is connected to network 54. Communication interfaces 44 and 26 are responsible for the exchange of various types of information between processor 46 and processor 28 via network 54.

[0030] Figure 2 shows an example of the main functions of the data processing device 12 and the smart device 14.

[0031] As shown in Figure 2, in the data processing device 12, a specific processing is performed by the processor 28. A specific processing program 56 is stored in the storage 32. The specific processing program 56 is an example of a "program" related to the technology of this disclosure. The processor 28 reads the specific processing program 56 from the storage 32 and executes the read specific processing program 56 on the RAM 30. The specific processing is realized by the processor 28 operating as a specific processing unit 290 according to the specific processing program 56 executed on the RAM 30.

[0032] The storage 32 stores the data generation model 58 and the emotion identification model 59. The data generation model 58 and the emotion identification model 59 are used by the identification processing unit 290.

[0033] In the smart device 14, the processor 46 performs the reception output processing. The storage 50 stores the reception output program 60. The reception output program 60 is used in conjunction with a specific processing program 56 by the data processing system 10. The processor 46 reads the reception output program 60 from the storage 50 and executes the read reception output program 60 on the RAM 48. The reception output processing is realized by the processor 46 operating as a control unit 46A according to the reception output program 60 executed on the RAM 48.

[0034] Next, the specific processing performed by the specific processing unit 290 of the data processing device 12 will be described. In the following description, the data processing device 12 will be referred to as the "server" and the smart device 14 as the "terminal".

[0035] The system of the present invention supports the user's psychological health by acquiring the user's biometric information and voice data in real time, analyzing the user's emotional state based on this data, and generating and providing appropriate music.

[0036] In this system, the device first collects data from smartwatches and monitoring cameras. Specifically, it acquires heart rate and voice data and sends this information to a server. When using this system, users can naturally provide data from devices they wear daily without any physical burden.

[0037] The server uses an emotion analysis model to identify the user's current emotional state based on the received biometric and voice data. This emotion analysis is performed by comprehensively analyzing acoustic characteristics such as voice tone and pitch, as well as physiological data such as changes in heart rate and body temperature.

[0038] After the emotional state is identified, the server uses music generation tools to create music appropriate to the user's emotions. This process utilizes prompts to take into account the user's musical preferences and cultural background, allowing for the provision of music that is highly relatable to the user.

[0039] Ultimately, the music generated from the server is streamed to the user's device, allowing them to listen to the music in real time and use it to relax, concentrate, or boost their motivation. In this way, the system provides a real-time musical experience tailored to the user's psychological needs.

[0040] For example, if a user feels stressed while working, the system can determine their stress level from an increase in heart rate and a change in voice tone, generate relaxing music, and immediately begin playing it. In this way, this invention contributes to improving the user's quality of life.

[0041] The following describes the processing flow.

[0042] Step 1:

[0043] The device transmits biometric and audio data acquired from smartwatches and monitoring cameras to the server in real time. The data is transmitted encrypted to protect user privacy.

[0044] Step 2:

[0045] The server inputs the received biometric and voice data into the data analysis module and cleanses the data. In this process, missing values ​​are imputed and outliers are removed, preparing clean data for analysis.

[0046] Step 3:

[0047] The server uses cleansed data to apply acoustic feature extraction algorithms, extracting features such as tone and pitch from the audio data. It also analyzes changes in heart rate and body temperature from biometric information and uses this data to run an emotion analysis model.

[0048] Step 4:

[0049] The server updates user profiles based on sentiment scores obtained from sentiment analysis models. These profiles take into account each user's characteristics and historical data, and also reflect long-term trend analysis.

[0050] Step 5:

[0051] The server determines an appropriate music style that corresponds to the user's current emotional state. Then, a music generation algorithm (such as GAN or RNN) is used to generate an original music track based on the selected style.

[0052] Step 6:

[0053] The server sends the generated music data to the user's device, allowing the user to play and enjoy it. If user feedback is received, this data is also saved and used to improve the application in the future.

[0054] (Example 1)

[0055] Next, we will describe Example 1. In the following description, the data processing device 12 will be referred to as the "server," and the smart device 14 will be referred to as the "terminal."

[0056] In modern society, psychological stress and anxiety are problems faced by many people, and there is a need to provide means to appropriately alleviate them. However, conventional methods have the challenge of not being able to identify the emotional state of individual users in real time and provide relaxation that is appropriate to that state.

[0057] The identification process performed by the identification processing unit 290 of the data processing device 12 in Example 1 is realized by the following means.

[0058] In this invention, the server includes data collection means for acquiring the user's biometric indicators and voice information, analysis means for analyzing the biometric indicators and voice information to identify the user's emotional state, and generation means for generating music that reflects the user's preferences and cultural background in relation to the emotional state. This makes it possible to provide the user with a personalized music experience in real time, thereby reducing psychological stress and anxiety.

[0059] A "user" is an individual who uses the system and provides biometric information and voice data.

[0060] "Biometric indicators" refer to data that numerically represents a user's physical condition, such as heart rate and body temperature.

[0061] "Audio information" refers to data that includes acoustic characteristics such as the tone and pitch of the user's voice.

[0062] "Data collection means" refers to a device or method for detecting and acquiring a user's biometric indicators and voice information.

[0063] "Analysis means" refers to a device or method that analyzes acquired biometric indicators and voice information to identify the user's emotional state.

[0064] "Emotional state" refers to information that indicates the user's mental state, and includes specific psychological conditions such as stress, stability, and happiness.

[0065] "Generating means" refers to a device or method for creating music in relation to the user's emotional state.

[0066] "Preferences" refer to information about the music genres and styles that individual users particularly enjoy.

[0067] "Cultural background" refers to information about a user's expectations and preferences regarding music, based on their lifestyle and social and regional characteristics.

[0068] This system supports users' psychological well-being by acquiring biometric and voice information from users in real time, analyzing their emotional state based on that data, and generating and providing appropriate music.

[0069] The device collects biometric and audio information using smartwatches and monitoring cameras that users use daily. This involves using hardware such as heart rate sensors and microphones to acquire the user's physical and acoustic data. This data is transmitted to a server using communication protocols such as Bluetooth and Wi-Fi.

[0070] The server is equipped with a generative AI model for analyzing received biometric and voice information. This model comprehensively analyzes acoustic features such as voice tone and pitch, as well as physiological data such as heart rate and body temperature, to identify the user's emotional state. The generative AI model can also consider the user's musical preferences and cultural background using prompts. For example, it can use prompts such as, "Your current emotional state is stress; please provide relaxing jazz music."

[0071] After identifying the user's emotional state, the server uses a generative AI model to automatically generate music appropriate to the user's emotions. The generated music is streamed to the user's device and provided in real time. This allows the user to relax, concentrate, or improve their motivation. This system aims to improve the user's quality of life by providing a music experience optimized for each individual user.

[0072] The flow of the specific processing in Example 1 will be explained using Figure 11.

[0073] Step 1:

[0074] The device uses the user's smartwatch or monitoring camera to collect biometric data (e.g., heart rate) and audio information (e.g., voice tone) in real time. This data is acquired by the device as input and prepared for the next processing step. Specifically, the heart rate sensor periodically records the pulse, and the microphone records the user's voice.

[0075] Step 2:

[0076] The device transmits collected biometric and voice information to a server using network technologies such as Bluetooth or Wi-Fi. The output of this step is composed of transmission data and prepared for analysis on the server. Specifically, data packets are transferred to the server via network communication.

[0077] Step 3:

[0078] The server analyzes the emotional state using a generative AI model based on the received data. Inputs include heart rate variability and voice tone patterns, and the output identifies the user's current emotional state (e.g., stress, relaxation). Specifically, the AI ​​model adapts the data to prompt statements to reveal the emotional state.

[0079] Step 4:

[0080] The server uses music generation methods based on the user's emotional state to create music suitable for the user. The generation AI model processes prompts such as "Please generate music for relaxation," and music data is constructed as output. Specifically, the AI ​​algorithm creates music while considering the user's musical preferences and cultural background.

[0081] Step 5:

[0082] The server then streams the generated music data back to the terminal. The input for this step is the generated music data, and the output is provided in a format playable on the user's device. Specifically, the music data stream is transmitted to the terminal via the network, and playback takes place in real time.

[0083] In this way, the system aims to provide users with a music experience tailored to their individual needs and to reduce psychological stress.

[0084] (Application Example 1)

[0085] Next, we will explain Application Example 1. In the following explanation, the data processing device 12 will be referred to as the "server," and the smart device 14 will be referred to as the "terminal."

[0086] In production environments, workers' mental stress and decreased concentration significantly impact work efficiency. This increases the risk of reduced quality and accidents, making it crucial to understand workers' psychological states in real time and take appropriate action. Conventional methods have struggled to quickly and accurately assess workers' emotional states and provide appropriate relaxation measures.

[0087] The specific processing performed by the specific processing unit 290 of the data processing device 12 in Application Example 1 is realized by the following means.

[0088] In this invention, the server includes information gathering means for acquiring the user's biometric information, emotion analysis means for analyzing the user's emotional state based on the biometric information, and music generation means for generating music data according to the emotional state. This makes it possible to automatically provide music suitable for the work environment while monitoring the psychological state of the worker.

[0089] "Information gathering means" refers to a device or method for acquiring a user's biometric information, such as voice data or heart rate data.

[0090] "Emotional analysis means" refers to a process or device for analyzing and identifying a user's emotional state based on acquired biometric information.

[0091] "Music generation means" refers to a technology or device for generating optimal music data according to an analyzed emotional state.

[0092] "Music delivery means" refers to a method or device for outputting generated music data to a user or working environment.

[0093] "Work environment" refers to the physical or virtual workplace where music data is output by the music delivery means.

[0094] The system of this invention is implemented via a terminal and server installed in the manufacturing environment. The terminal is connected to information collection devices such as a smartwatch and a camera, which acquire the worker's biometric information in real time. Specifically, the terminal collects the worker's heart rate and voice data and transmits the data to the server via the network.

[0095] The server runs analysis software using programming languages ​​such as Python and R, which implements an emotion analysis model. Pandas and NumPy are used for data processing, including cleaning and preprocessing. The server inputs received biometric information into the emotion analysis model and analyzes the user's emotional state using machine learning frameworks such as TENSORFLOW® and PyTorch.

[0096] Based on the analyzed emotional state, the server utilizes a generative AI model to generate music. Open-source music generation tools and APIs are used to create customized music tailored to the user's psychological state. The generated music data is streamed to the terminal through the work environment's speaker system. This allows workers to listen to stress-reducing music in real time.

[0097] As a concrete example, when a worker feels stressed, the system generates relaxing music based on an increase in heart rate and a change in voice tone, and distributes it to the workspace. An example of a prompt message to the generation AI model would be, "Generate music that will help employees relax, and distribute that music through the speaker system."

[0098] The flow of a specific process in Application Example 1 will be explained using Figure 12.

[0099] Step 1:

[0100] The device acquires biometric information from a smartwatch and a monitoring camera. It receives user heart rate and voice data as input. This data is collected via a sensor interface and converted to JSON format for portability.

[0101] Step 2:

[0102] The device transmits collected biometric information to the server. Input data includes heart rate, voice pitch, and tone. The data is transferred to the server via a REST API, and the server then incorporates it into its data analysis system.

[0103] Step 3:

[0104] The server analyzes the received biometric information. The input data includes heart rate and voice data, and data cleaning and normalization are performed using Pandas and NumPy. This prepares the input for the sentiment analysis model. A machine learning framework (e.g., TensorFlow) is used for sentiment analysis, and the analysis results identify the user's current emotional state.

[0105] Step 4:

[0106] The server generates music data based on the analyzed emotional state. It receives emotional state data as input and uses a generative AI model (e.g., a music generation API) to generate customized music. The output includes music data.

[0107] Step 5:

[0108] The server streams the generated music data to terminals in the work environment. The terminals output the music through a speaker system. Employees can listen to this music to relax or improve their concentration. The input is the generated music data, and the output is music playback from the speakers.

[0109] Furthermore, an emotion engine that estimates the user's emotions may be incorporated. That is, the identification processing unit 290 may use the emotion identification model 59 to estimate the user's emotions and perform identification processing using the user's emotions.

[0110] This invention relates to a system that recognizes emotions based on a user's biometric information and voice data, and provides music that corresponds to the user's emotional state. This system achieves highly accurate emotion analysis by incorporating an emotion engine, and generates music using the results of the emotion analysis.

[0111] The device acquires biometric information such as heart rate, body temperature, and voice data in real time from smartwatches and monitoring cameras, and transmits it to a server. Users can use this system naturally in their daily lives without requiring any special operation.

[0112] The server receives biometric information and voice data transmitted from the terminal and inputs this data into the emotion engine. The emotion engine uses advanced machine learning algorithms to analyze features such as tone, pitch, and tempo of the voice data to precisely identify the user's emotions. It can also recognize patterns of emotional fluctuations by referring to past emotional state history and predict current emotions.

[0113] Based on these results, the server generates music data that corresponds to the user's current and predicted emotional state. Modern algorithms such as generative opposing networks and recurrent neural networks are used for music generation to create music tracks that match the user's preferences and emotions.

[0114] The generated music is streamed to the user's device, allowing the user to play the music through their mobile device or smart speaker for relaxation or a change of pace. For example, if a user is feeling stressed at work, the system will provide music with a relaxing tone to alleviate stress. In this way, the present invention realizes an emotion-responsive music experience tailored to individual users, contributing to the promotion of the user's psychological well-being.

[0115] The following describes the processing flow.

[0116] Step 1:

[0117] The device acquires the user's biometric information and voice data in real time from smartwatches and monitoring cameras. This data includes heart rate, body temperature, and voice tone.

[0118] Step 2:

[0119] The device encrypts the acquired data and sends it to the server using a secure communication protocol. This protects the user's privacy.

[0120] Step 3:

[0121] The server inputs the received data into the emotion engine. The emotion engine analyzes the tone, pitch, and tempo of the audio data and runs a machine learning model to identify the user's emotions.

[0122] Step 4:

[0123] The server updates the user profile using emotion identification results obtained from the emotion engine. This profile is used to store past emotional state information and understand the user's patterns over the long term.

[0124] Step 5:

[0125] The server uses the latest music generation algorithms to generate music data based on emotion analysis. By employing generative opposing networks and recurrent neural networks, it generates music tracks that match individual emotions.

[0126] Step 6:

[0127] The server streams the generated music data to the user's device. Users can then enjoy the music in real time via their mobile devices or smart speakers.

[0128] Step 7:

[0129] Users can relax and change their mood through music suited to their emotions. They can also provide feedback as needed to help improve the system.

[0130] (Example 2)

[0131] Next, we will describe Example 2. In the following description, the data processing device 12 will be referred to as the "server" and the smart device 14 as the "terminal".

[0132] In modern society, there is a need to manage and appropriately respond to people's emotions and stress, but conventional systems have difficulty providing music that is appropriate to the user's physiological and psychological state. Therefore, users have to select music that suits their emotions themselves, which presents a challenge in terms of efficient emotional management.

[0133] The identification process performed by the identification processing unit 290 of the data processing device 12 in Example 2 is realized by the following means.

[0134] In this invention, the server includes data collection means for acquiring the user's biometric data, emotion analysis means for analyzing the user's emotional state based on the biometric data and voice data, emotion prediction means for predicting emotional fluctuation patterns considering the emotional state and past emotional history, and music generation means for generating music data based on the predicted emotional state. This allows the user to passively receive music appropriate to their state, making it easier to manage their emotions on a daily basis.

[0135] "Biometric data" refers to measurable information about a user's body, such as heart rate and body temperature, which indicate the user's physiological state.

[0136] "Data collection means" refers to devices and technologies that can acquire biometric data or voice data from users, such as devices that use sensors or microphones.

[0137] "Emotional analysis means" refers to algorithms and technologies that analyze acquired biometric and voice data to identify the user's emotional state.

[0138] "Emotion prediction methods" refer to technologies and models that predict a user's emotional fluctuation patterns based on current and past emotional data.

[0139] "Music generation means" refers to a technology for creating music data suitable for the user according to their analyzed emotional state, and is a function that generates music tracks using a generation AI model.

[0140] A "generative AI model" is an artificial intelligence model that has the ability to learn from data and generate new data, and specifically includes generative opposing networks and recurrent neural networks.

[0141] This invention is a system for providing a music experience tailored to the physiological and psychological characteristics of users. This system collects the user's biometric and vocal data in real time, analyzes the user's emotional state based on that data, and generates music data according to the analysis results.

[0142] The device uses the user's smartwatch or monitoring camera to acquire physiological data such as heart rate and body temperature, as well as voice data. This allows for the accurate and continuous collection of the user's biometric data. The device uses a secure communication protocol to safely transmit this data to the server.

[0143] The server receives the transmitted data and uses sentiment analysis tools to accurately identify the user's emotional state. This analysis utilizes machine learning algorithms and tools to analyze features such as tone, pitch, and tempo of the voice. Furthermore, sentiment prediction tools allow the server to refer to past emotional history and predict not only the current emotional state but also its fluctuation patterns.

[0144] The server then uses a generative AI model to generate music data tailored to the user's emotions. Specifically, it uses generative opposing networks (GANs) and recurrent neural networks (RNNs) to create music tracks that the user will like. This generated music is then delivered to the user's device via streaming. The user can play this music on their mobile device or smart speaker to help them relax or change their mood.

[0145] For example, when a user feels anxious in a large crowd, this system analyzes voice and biometric data to generate anxiety-relieving music. By listening to this music, the user can reduce stress and regain calmness.

[0146] An example of a prompt message is: "If the emotional state obtained from the user's biometric information is anxious, generate a music track to alleviate anxiety." This format allows you to instruct the generation AI model in this way.

[0147] In this way, the present invention provides a music experience tailored to the user's emotional state, realizing a system that contributes to maintaining and improving psychological health.

[0148] The flow of the specific processing in Example 2 will be explained using Figure 13.

[0149] Step 1:

[0150] The device acquires biometric information and voice data through the user's smartwatch or monitoring camera. Specifically, the smartwatch's sensors measure heart rate and body temperature, and voice data is recorded by the camera's microphone. This input data is sent to the server as physiological indicators and voice characteristics.

[0151] Step 2:

[0152] The server receives biometric information and voice data transmitted from the terminal. Input data includes heart rate, body temperature, voice tone, pitch, and tempo. The server preprocesses the data, removing noise to make it suitable for analysis. This processed data is then passed to the emotion analysis system.

[0153] Step 3:

[0154] The server analyzes pre-processed data using emotion analysis tools. The audio data is analyzed for changes in tone and rhythm, and combined with heart rate and body temperature to determine the user's emotional state. The user's emotional state is output as a result of this calculation.

[0155] Step 4:

[0156] The server uses emotion prediction tools in addition to emotion analysis tools to refer to the user's past emotion data and predict how their current emotional state will change. By utilizing past records, it calculates patterns of temporal changes in emotions and outputs predicted emotion values.

[0157] Step 5:

[0158] The server generates music data using a generative AI model based on the obtained emotional state and predicted values. Specifically, it uses generative opposing networks (GANs) and recurrent neural networks (RNNs) to create music tracks that are optimal for the user's emotions. Personalized music is output based on the input emotional information and the user's musical preferences.

[0159] Step 6:

[0160] The generated music data is streamed from the server to the user's device. The user plays the music on their smartphone or smart speaker and enjoys its relaxing effects. This process involves the user receiving the music and taking actual actions to maintain their psychological well-being and comfort.

[0161] (Application Example 2)

[0162] Next, we will explain application example 2. In the following explanation, the data processing device 12 will be referred to as a "server" and the smart device 14 as a "terminal".

[0163] Conventional music delivery systems struggle to select music that matches a user's emotional state, making it difficult to provide a musical experience that responds to the instantaneous emotional changes of individual users. Furthermore, existing services lack sufficient mechanisms to analyze biometric information and voice data in real time and generate and provide appropriate music. Therefore, there is a need for technology that efficiently delivers music tailored to a user's emotional state and improves satisfaction.

[0164] The specific processing performed by the specific processing unit 290 of the data processing device 12 in Application Example 2 is realized by the following means.

[0165] In this invention, the server includes information acquisition means for acquiring the user's biometric information, emotion analysis means for analyzing the user's emotional state based on the biometric information and voice data, and music creation means for generating music data according to the emotional state. This makes it possible to automatically generate and provide music that is tailored to the emotional state of each individual user.

[0166] "Information acquisition means" refers to a device or mechanism that has the function of acquiring a user's biometric information or voice data.

[0167] "Emotional analysis means" refers to a device or algorithm that has the function of analyzing and identifying a user's emotional state based on acquired biometric information and voice data.

[0168] "Music creation means" refers to a device or algorithm that has the function of generating music data in accordance with the analyzed emotional state.

[0169] "Music transmission means" refers to a device or structure that has the function of transmitting and providing generated music data to a user.

[0170] To implement this invention, it is desirable to use a smartwatch or smartphone as a terminal for acquiring the user's biometric information in real time. These terminals function as "information acquisition means" for acquiring biometric information such as heart rate and body temperature, and acquire voice data using a microphone built into the terminal.

[0171] The acquired biometric information and voice data are transmitted to a server via the network. After receiving this data, the server uses an "emotion analysis tool" that implements a machine learning algorithm to analyze the user's emotional state. To do this, the server uses advanced machine learning software to analyze the tone, pitch, tempo, and other aspects of the voice data.

[0172] Based on the analysis results, the server uses a "music creation method" to generate music data that corresponds to the user's emotions. At this time, algorithms such as generative opposition learning models and recurrent neural networks are utilized to generate music that fits the user's emotions.

[0173] The generated music data is transmitted to the user's device in real time. The device has music playback capabilities, and the user can play the generated music through their smartphone or smart speaker to relax or uplift their emotions.

[0174] For example, if a user wants to relax, the system will automatically generate relaxing music and provide it to the user's device. Another example prompt is, "If the user's heart rate is elevated and a desire to be energized is detected, suggest a music track that generates uplifting music." The server processes this prompt, generating and providing music that matches the user's mood at that moment.

[0175] The flow of a specific process in Application Example 2 will be explained using Figure 14.

[0176] Step 1:

[0177] The terminal acquires the user's biometric information (heart rate, body temperature, etc.) via a device such as a smartwatch. This input data is measured in real time. Based on these results, the terminal packages the data and similarly acquires voice data, preparing to send it to the server.

[0178] Step 2:

[0179] The device transmits biometric information and voice data to the server. The input at this time is the biometric information acquired by the device and the recorded voice data. The server receives this data, appropriately decodes it, and formats it into the format necessary for analysis.

[0180] Step 3:

[0181] The server inputs the received biometric and voice data into a machine learning model for emotion analysis. In this step, the user's emotional state is identified by analyzing features such as voice tone, pitch, and tempo, and the analysis results are output. The results include the main emotions the user is feeling and their intensity.

[0182] Step 4:

[0183] The server generates music data using a generative AI model based on the emotion analysis results. This algorithm utilizes generative opposing networks and recurrent neural networks to create music that corresponds to the analyzed emotions. At this point, the input is the emotion analysis results, and the output is the generated music data.

[0184] Step 5:

[0185] The server streams the generated music data to the terminal in real time. The terminal receives this music and prepares it for playback by the user. The input is the generated music data, and the output is the music track ready for the user to listen to.

[0186] Step 6:

[0187] Users listen to music streamed through their devices. Based on user feedback and usage patterns, the server collects further data and continuously improves the system accordingly. This feedback mechanism enables specific actions to enhance the individual user experience.

[0188] The specific processing unit 290 transmits the result of the specific processing to the smart device 14. In the smart device 14, the control unit 46A causes the output device 40 to output the result of the specific processing. The microphone 38B acquires audio indicating user input for the result of the specific processing. The control unit 46A transmits the audio data indicating user input acquired by the microphone 38B to the data processing device 12. In the data processing device 12, the specific processing unit 290 acquires the audio data.

[0189] Data generation model 58 is a so-called generative AI (Artificial Intelligence). An example of data generation model 58 is ChatGPT (registered trademark) (Internet search).<URL: https: / / openai.com / blog / chatgpt> ), Gemini (registered trademark) (Internet search) <url: https: gemini.google.com ?hl="ja">Examples of generative AI include the following. The data generation model 58 is obtained by performing deep learning on a neural network. The data generation model 58 is input with prompts containing instructions, and with inference data such as audio data representing speech, text data representing text, and image data representing images. The data generation model 58 infers from the input inference data according to the instructions indicated by the prompts, and outputs the inference results in data formats such as audio data and text data. Here, inference refers to, for example, analysis, classification, prediction, and / or summarization.

[0190] In the above embodiment, an example was given in which specific processing is performed by the data processing device 12, but the technology of this disclosure is not limited thereto, and the specific processing may also be performed by the smart device 14.

[0191] [Second Embodiment]

[0192] Figure 3 shows an example of the configuration of the data processing system 210 according to the second embodiment.

[0193] As shown in Figure 3, the data processing system 210 includes a data processing device 12 and smart glasses 214. An example of the data processing device 12 is a server.

[0194] The data processing device 12 comprises a computer 22, a database 24, and a communication interface 26. The computer 22 is an example of a "computer" related to the technology of this disclosure. The computer 22 comprises a processor 28, RAM 30, and storage 32. The processor 28, RAM 30, and storage 32 are connected to a bus 34. The database 24 and the communication interface 26 are also connected to the bus 34. The communication interface 26 is connected to a network 54. An example of the network 54 is a WAN (Wide Area Network) and / or a LAN (Local Area Network).

[0195] The smart glasses 214 include a computer 36, a microphone 238, a speaker 240, a camera 42, and a communication interface 44. The computer 36 includes a processor 46, RAM 48, and storage 50. The processor 46, RAM 48, and storage 50 are connected to a bus 52. The microphone 238, speaker 240, and camera 42 are also connected to the bus 52.

[0196] The microphone 238 receives voice signals from the user 20 and receives instructions from the user 20. The microphone 238 captures the voice signals from the user 20, converts the captured voice into audio data, and outputs it to the processor 46. The speaker 240 outputs audio according to the instructions from the processor 46.

[0197] Camera 42 is a small digital camera equipped with an optical system including a lens, aperture, and shutter, and an image sensor such as a CMOS (Complementary Metal-Oxide-Semiconductor) image sensor or a CCD (Charge Coupled Device) image sensor, and captures images of the area around the user 20 (for example, an imaging range defined by a field of view equivalent to the width of a typical healthy person's field of vision).

[0198] Communication interface 44 is connected to network 54. Communication interfaces 44 and 26 are responsible for the exchange of various information between processor 46 and processor 28 via network 54. The exchange of various information between processor 46 and processor 28 using communication interfaces 44 and 26 is performed in a secure manner.

[0199] Figure 4 shows an example of the main functions of the data processing device 12 and the smart glasses 214. As shown in Figure 4, the data processing device 12 performs specific processing using the processor 28. The storage 32 stores the specific processing program 56.

[0200] The specific processing program 56 is an example of a "program" relating to the technology of this disclosure. The processor 28 reads the specific processing program 56 from the storage 32 and executes the read specific processing program 56 on the RAM 30. The specific processing is realized by the processor 28 operating as a specific processing unit 290 in accordance with the specific processing program 56 executed on the RAM 30.

[0201] The storage 32 stores the data generation model 58 and the emotion identification model 59. The data generation model 58 and the emotion identification model 59 are used by the identification processing unit 290.

[0202] In the smart glasses 214, the processor 46 performs the reception output processing. The storage 50 stores the reception output program 60. The processor 46 reads the reception output program 60 from the storage 50 and executes the read reception output program 60 on the RAM 48. The reception output processing is realized by the processor 46 operating as a control unit 46A according to the reception output program 60 executed on the RAM 48.

[0203] Next, the identification processing performed by the identification processing unit 290 of the data processing device 12 will be described. In the following description, the data processing device 12 will be referred to as the "server" and the smart glasses 214 will be referred to as the "terminal".

[0204] The system of the present invention supports the user's psychological health by acquiring the user's biometric information and voice data in real time, analyzing the user's emotional state based on this data, and generating and providing appropriate music.

[0205] In this system, the device first collects data from smartwatches and monitoring cameras. Specifically, it acquires heart rate and voice data and sends this information to a server. When using this system, users can naturally provide data from devices they wear daily without any physical burden.

[0206] The server uses an emotion analysis model to identify the user's current emotional state based on the received biometric and voice data. This emotion analysis is performed by comprehensively analyzing acoustic characteristics such as voice tone and pitch, as well as physiological data such as changes in heart rate and body temperature.

[0207] After the emotional state is identified, the server uses music generation tools to create music appropriate to the user's emotions. This process utilizes prompts to take into account the user's musical preferences and cultural background, allowing for the provision of music that is highly relatable to the user.

[0208] Ultimately, the music generated from the server is streamed to the user's device, allowing them to listen to the music in real time and use it to relax, concentrate, or boost their motivation. In this way, the system provides a real-time musical experience tailored to the user's psychological needs.

[0209] For example, if a user feels stressed while working, the system can determine their stress level from an increase in heart rate and a change in voice tone, generate relaxing music, and immediately begin playing it. In this way, this invention contributes to improving the user's quality of life.

[0210] The following describes the processing flow.

[0211] Step 1:

[0212] The device transmits biometric and audio data acquired from smartwatches and monitoring cameras to the server in real time. The data is transmitted encrypted to protect user privacy.

[0213] Step 2:

[0214] The server inputs the received biometric and voice data into the data analysis module and cleanses the data. In this process, missing values ​​are imputed and outliers are removed, preparing clean data for analysis.

[0215] Step 3:

[0216] The server uses cleansed data to apply acoustic feature extraction algorithms, extracting features such as tone and pitch from the audio data. It also analyzes changes in heart rate and body temperature from biometric information and uses this data to run an emotion analysis model.

[0217] Step 4:

[0218] The server updates user profiles based on sentiment scores obtained from sentiment analysis models. These profiles take into account each user's characteristics and historical data, and also reflect long-term trend analysis.

[0219] Step 5:

[0220] The server determines an appropriate music style that corresponds to the user's current emotional state. Then, a music generation algorithm (such as GAN or RNN) is used to generate an original music track based on the selected style.

[0221] Step 6:

[0222] The server sends the generated music data to the user's device, allowing the user to play and enjoy it. If user feedback is received, this data is also saved and used to improve the application in the future.

[0223] (Example 1)

[0224] Next, we will describe Example 1. In the following description, the data processing device 12 will be referred to as the "server," and the smart glasses 214 will be referred to as the "terminal."

[0225] In modern society, psychological stress and anxiety are problems faced by many people, and there is a need to provide means to appropriately alleviate them. However, conventional methods have the challenge of not being able to identify the emotional state of individual users in real time and provide relaxation that is appropriate to that state.

[0226] The identification process performed by the identification processing unit 290 of the data processing device 12 in Example 1 is realized by the following means.

[0227] In this invention, the server includes data collection means for acquiring the user's biometric indicators and voice information, analysis means for analyzing the biometric indicators and voice information to identify the user's emotional state, and generation means for generating music that reflects the user's preferences and cultural background in relation to the emotional state. This makes it possible to provide the user with a personalized music experience in real time, thereby reducing psychological stress and anxiety.

[0228] A "user" is an individual who uses the system and provides biometric information and voice data.

[0229] "Biometric indicators" refer to data that numerically represents a user's physical condition, such as heart rate and body temperature.

[0230] "Audio information" refers to data that includes acoustic characteristics such as the tone and pitch of the user's voice.

[0231] "Data collection means" refers to a device or method for detecting and acquiring a user's biometric indicators and voice information.

[0232] "Analysis means" refers to a device or method that analyzes acquired biometric indicators and voice information to identify the user's emotional state.

[0233] "Emotional state" refers to information that indicates the user's mental state, and includes specific psychological conditions such as stress, stability, and happiness.

[0234] "Generating means" refers to a device or method for creating music in relation to the user's emotional state.

[0235] "Preferences" refer to information about the music genres and styles that individual users particularly enjoy.

[0236] "Cultural background" refers to information about a user's expectations and preferences regarding music, based on their lifestyle and social and regional characteristics.

[0237] This system supports users' psychological well-being by acquiring biometric and voice information from users in real time, analyzing their emotional state based on that data, and generating and providing appropriate music.

[0238] The device collects biometric and audio information using smartwatches and monitoring cameras that users use daily. This involves using hardware such as heart rate sensors and microphones to acquire the user's physical and acoustic data. This data is transmitted to a server using communication protocols such as Bluetooth and Wi-Fi.

[0239] The server is equipped with a generative AI model for analyzing received biometric and voice information. This model comprehensively analyzes acoustic features such as voice tone and pitch, as well as physiological data such as heart rate and body temperature, to identify the user's emotional state. The generative AI model can also consider the user's musical preferences and cultural background using prompts. For example, it can use prompts such as, "Your current emotional state is stress; please provide relaxing jazz music."

[0240] After identifying the user's emotional state, the server uses a generative AI model to automatically generate music appropriate to the user's emotions. The generated music is streamed to the user's device and provided in real time. This allows the user to relax, concentrate, or improve their motivation. This system aims to improve the user's quality of life by providing a music experience optimized for each individual user.

[0241] The flow of the specific processing in Example 1 will be explained using Figure 11.

[0242] Step 1:

[0243] The device uses the user's smartwatch or monitoring camera to collect biometric data (e.g., heart rate) and audio information (e.g., voice tone) in real time. This data is acquired by the device as input and prepared for the next processing step. Specifically, the heart rate sensor periodically records the pulse, and the microphone records the user's voice.

[0244] Step 2:

[0245] The device transmits collected biometric and voice information to a server using network technologies such as Bluetooth or Wi-Fi. The output of this step is composed of transmission data and prepared for analysis on the server. Specifically, data packets are transferred to the server via network communication.

[0246] Step 3:

[0247] The server analyzes the emotional state using a generative AI model based on the received data. Inputs include heart rate variability and voice tone patterns, and the output identifies the user's current emotional state (e.g., stress, relaxation). Specifically, the AI ​​model adapts the data to prompt statements to reveal the emotional state.

[0248] Step 4:

[0249] The server uses music generation methods based on the user's emotional state to create music suitable for the user. The generation AI model processes prompts such as "Please generate music for relaxation," and music data is constructed as output. Specifically, the AI ​​algorithm creates music while considering the user's musical preferences and cultural background.

[0250] Step 5:

[0251] The server then streams the generated music data back to the terminal. The input for this step is the generated music data, and the output is provided in a format playable on the user's device. Specifically, the music data stream is transmitted to the terminal via the network, and playback takes place in real time.

[0252] In this way, the system aims to provide users with a music experience tailored to their individual needs and to reduce psychological stress.

[0253] (Application Example 1)

[0254] Next, we will explain Application Example 1. In the following explanation, the data processing device 12 will be referred to as the "server," and the smart glasses 214 will be referred to as the "terminal."

[0255] In production environments, workers' mental stress and decreased concentration significantly impact work efficiency. This increases the risk of reduced quality and accidents, making it crucial to understand workers' psychological states in real time and take appropriate action. Conventional methods have struggled to quickly and accurately assess workers' emotional states and provide appropriate relaxation measures.

[0256] The specific processing performed by the specific processing unit 290 of the data processing device 12 in Application Example 1 is realized by the following means.

[0257] In this invention, the server includes information gathering means for acquiring the user's biometric information, emotion analysis means for analyzing the user's emotional state based on the biometric information, and music generation means for generating music data according to the emotional state. This makes it possible to automatically provide music suitable for the work environment while monitoring the psychological state of the worker.

[0258] "Information gathering means" refers to a device or method for acquiring a user's biometric information, such as voice data or heart rate data.

[0259] "Emotional analysis means" refers to a process or device for analyzing and identifying a user's emotional state based on acquired biometric information.

[0260] "Music generation means" refers to a technology or device for generating optimal music data according to an analyzed emotional state.

[0261] "Music delivery means" refers to a method or device for outputting generated music data to a user or working environment.

[0262] "Work environment" refers to the physical or virtual workplace where music data is output by the music delivery means.

[0263] The system of this invention is implemented via a terminal and server installed in the manufacturing environment. The terminal is connected to information collection devices such as a smartwatch and a camera, which acquire the worker's biometric information in real time. Specifically, the terminal collects the worker's heart rate and voice data and transmits the data to the server via the network.

[0264] The server runs analysis software using programming languages ​​such as Python and R, which implements an emotion analysis model. Pandas and NumPy are used for data processing, including cleaning and preprocessing. The server inputs received biometric information into the emotion analysis model and analyzes the user's emotional state using machine learning frameworks such as TensorFlow and PyTorch.

[0265] Based on the analyzed emotional state, the server utilizes a generative AI model to generate music. Open-source music generation tools and APIs are used to create customized music tailored to the user's psychological state. The generated music data is streamed to the terminal through the work environment's speaker system. This allows workers to listen to stress-reducing music in real time.

[0266] As a concrete example, when a worker feels stressed, the system generates relaxing music based on an increase in heart rate and a change in voice tone, and distributes it to the workspace. An example of a prompt message to the generation AI model would be, "Generate music that will help employees relax, and distribute that music through the speaker system."

[0267] The flow of a specific process in Application Example 1 will be explained using Figure 12.

[0268] Step 1:

[0269] The device acquires biometric information from a smartwatch and a monitoring camera. It receives user heart rate and voice data as input. This data is collected via a sensor interface and converted to JSON format for portability.

[0270] Step 2:

[0271] The device transmits collected biometric information to the server. Input data includes heart rate, voice pitch, and tone. The data is transferred to the server via a REST API, and the server then incorporates it into its data analysis system.

[0272] Step 3:

[0273] The server analyzes the received biometric information. The input data includes heart rate and voice data, and data cleaning and normalization are performed using Pandas and NumPy. This prepares the input for the sentiment analysis model. A machine learning framework (e.g., TensorFlow) is used for sentiment analysis, and the analysis results identify the user's current emotional state.

[0274] Step 4:

[0275] The server generates music data based on the analyzed emotional state. It receives emotional state data as input and uses a generative AI model (e.g., a music generation API) to generate customized music. The output includes music data.

[0276] Step 5:

[0277] The server streams the generated music data to terminals in the work environment. The terminals output the music through a speaker system. Employees can listen to this music to relax or improve their concentration. The input is the generated music data, and the output is music playback from the speakers.

[0278] Furthermore, an emotion engine that estimates the user's emotions may be incorporated. That is, the identification processing unit 290 may use the emotion identification model 59 to estimate the user's emotions and perform identification processing using the user's emotions.

[0279] The present invention is a system that recognizes emotions based on a user's biometric information and voice data, and provides music according to the user's emotional state. By incorporating an emotion engine, this system achieves highly accurate emotion analysis and uses the emotion analysis results to generate music.

[0280] The terminal acquires biometric information such as heart rate, body temperature, and voice data in real time from a smartwatch or a monitoring camera, and transmits it to the server. The user can naturally use this system in daily life without the need for special operations.

[0281] The server receives the biometric information and voice data transmitted from the terminal, and inputs these data into the emotion engine. Using advanced machine learning algorithms, the emotion engine analyzes features such as the tone, pitch, and tempo of the voice data to precisely identify the user's emotions. It is also possible to recognize the emotional fluctuation pattern and predict the current emotion by referring to the past emotional state history.

[0282] Based on this result, the server generates music data according to the user's current and predicted emotional states. For music generation, the latest algorithms such as generative adversarial networks and recurrent neural networks are used to create music tracks suitable for the user's preferences and emotions.

[0283] The generated music is delivered to the user's terminal via streaming, and the user can play the music through a mobile device or a smart speaker to achieve relaxation and mood change. For example, when the user feels stressed during work, the system provides relaxing-tone music to relieve stress. Thus, the present invention realizes an emotion-responsive music experience suitable for individual users and contributes to the promotion of the user's mental health.

[0284] The following describes the processing flow.

[0285] Step 1:

[0286] The terminal acquires the user's biometric information and voice data in real time from a smartwatch or a surveillance camera. This data includes heart rate, body temperature, and voice tone, etc.

[0287] Step 2:

[0288] The terminal encrypts the acquired data and sends it to the server using a secure communication protocol. This protects the user's privacy.

[0289] Step 3:

[0290] The server inputs the received data into the emotion engine. The emotion engine analyzes the tone, pitch, and tempo of the voice data and executes a machine learning model for identifying the user's emotion.

[0291] Step 4: [[ID=​​​​​​​​​​​​​​​​​​​​​​​

[0298] Users can relax and change their mood through music suited to their emotions. They can also provide feedback as needed to help improve the system.

[0299] (Example 2)

[0300] Next, we will describe Example 2. In the following description, the data processing device 12 will be referred to as the "server" and the smart glasses 214 will be referred to as the "terminal".

[0301] In modern society, there is a need to manage and appropriately respond to people's emotions and stress, but conventional systems have difficulty providing music that is appropriate to the user's physiological and psychological state. Therefore, users have to select music that suits their emotions themselves, which presents a challenge in terms of efficient emotional management.

[0302] The identification process performed by the identification processing unit 290 of the data processing device 12 in Example 2 is realized by the following means.

[0303] In this invention, the server includes data collection means for acquiring the user's biometric data, emotion analysis means for analyzing the user's emotional state based on the biometric data and voice data, emotion prediction means for predicting emotional fluctuation patterns considering the emotional state and past emotional history, and music generation means for generating music data based on the predicted emotional state. This allows the user to passively receive music appropriate to their state, making it easier to manage their emotions on a daily basis.

[0304] "Biometric data" refers to measurable information about a user's body, such as heart rate and body temperature, which indicate the user's physiological state.

[0305] "Data collection means" refers to devices and technologies that can acquire biometric data or voice data from users, such as devices that use sensors or microphones.

[0306] The "emotion analysis means" is an algorithm or technology for analyzing the acquired biological data and voice data to identify the user's emotional state.

[0307] The "emotion prediction means" is a technology or model for predicting the variation pattern of the user's emotions based on the current and past emotion data.

[0308] The "music generation means" is a technology for creating music data suitable for the user according to the analyzed emotional state, and is a function for generating music tracks by utilizing a generative AI model.

[0309] The "generative AI model" is a model of artificial intelligence that has the ability to learn from data and generate new data, and specifically includes a generative adversarial network and a recurrent neural network.

[0310] The present invention is a system for providing a music experience tailored to the physiological and psychological characteristics of the user. This system collects the user's biological data and voice data in real time, analyzes the user's emotional state based on these data, and plays a role in generating music data according to the analysis results.

[0311] The terminal uses the user's smartwatch and monitoring camera to acquire physiological data such as heart rate and body temperature, and voice data. Thereby, the user's biological data can be accurately and continuously collected. The terminal uses a secure communication protocol to securely transmit this data to the server.

[0312] The server receives the transmitted data and accurately identifies the user's emotional state using the emotion analysis means. For this analysis, machine learning algorithms and tools for analyzing features such as the tone, pitch, and tempo of the voice are used. Also, by means of the emotion prediction means, it is possible to refer to the past emotion history and predict not only the current emotional state but also its variation pattern.

[0313] The server then uses a generative AI model to generate music data tailored to the user's emotions. Specifically, it uses generative opposing networks (GANs) and recurrent neural networks (RNNs) to create music tracks that the user will like. This generated music is then delivered to the user's device via streaming. The user can play this music on their mobile device or smart speaker to help them relax or change their mood.

[0314] For example, when a user feels anxious in a large crowd, this system analyzes voice and biometric data to generate anxiety-relieving music. By listening to this music, the user can reduce stress and regain calmness.

[0315] An example of a prompt message is: "If the emotional state obtained from the user's biometric information is anxious, generate a music track to alleviate anxiety." This format allows you to instruct the generation AI model in this way.

[0316] In this way, the present invention provides a music experience tailored to the user's emotional state, realizing a system that contributes to maintaining and improving psychological health.

[0317] The flow of the specific processing in Example 2 will be explained using Figure 13.

[0318] Step 1:

[0319] The device acquires biometric information and voice data through the user's smartwatch or monitoring camera. Specifically, the smartwatch's sensors measure heart rate and body temperature, and voice data is recorded by the camera's microphone. This input data is sent to the server as physiological indicators and voice characteristics.

[0320] Step 2:

[0321] The server receives biometric information and voice data transmitted from the terminal. Input data includes heart rate, body temperature, voice tone, pitch, and tempo. The server preprocesses the data, removing noise to make it suitable for analysis. This processed data is then passed to the emotion analysis system.

[0322] Step 3:

[0323] The server analyzes pre-processed data using emotion analysis tools. The audio data is analyzed for changes in tone and rhythm, and combined with heart rate and body temperature to determine the user's emotional state. The user's emotional state is output as a result of this calculation.

[0324] Step 4:

[0325] The server uses emotion prediction tools in addition to emotion analysis tools to refer to the user's past emotion data and predict how their current emotional state will change. By utilizing past records, it calculates patterns of temporal changes in emotions and outputs predicted emotion values.

[0326] Step 5:

[0327] The server generates music data using a generative AI model based on the obtained emotional state and predicted values. Specifically, it uses generative opposing networks (GANs) and recurrent neural networks (RNNs) to create music tracks that are optimal for the user's emotions. Personalized music is output based on the input emotional information and the user's musical preferences.

[0328] Step 6:

[0329] The generated music data is streamed from the server to the user's device. The user plays the music on their smartphone or smart speaker and enjoys its relaxing effects. This process involves the user receiving the music and taking actual actions to maintain their psychological well-being and comfort.

[0330] (Application Example 2)

[0331] Next, we will explain application example 2. In the following explanation, the data processing device 12 will be referred to as the "server," and the smart glasses 214 will be referred to as the "terminal."

[0332] Conventional music delivery systems struggle to select music that matches a user's emotional state, making it difficult to provide a musical experience that responds to the instantaneous emotional changes of individual users. Furthermore, existing services lack sufficient mechanisms to analyze biometric information and voice data in real time and generate and provide appropriate music. Therefore, there is a need for technology that efficiently delivers music tailored to a user's emotional state and improves satisfaction.

[0333] The specific processing performed by the specific processing unit 290 of the data processing device 12 in Application Example 2 is realized by the following means.

[0334] In this invention, the server includes information acquisition means for acquiring the user's biometric information, emotion analysis means for analyzing the user's emotional state based on the biometric information and voice data, and music creation means for generating music data according to the emotional state. This makes it possible to automatically generate and provide music that is tailored to the emotional state of each individual user.

[0335] "Information acquisition means" refers to a device or mechanism that has the function of acquiring a user's biometric information or voice data.

[0336] "Emotional analysis means" refers to a device or algorithm that has the function of analyzing and identifying a user's emotional state based on acquired biometric information and voice data.

[0337] "Music creation means" refers to a device or algorithm that has the function of generating music data in accordance with the analyzed emotional state.

[0338] "Music transmission means" refers to a device or structure that has the function of transmitting and providing generated music data to a user.

[0339] To implement this invention, it is desirable to use a smartwatch or smartphone as a terminal for acquiring the user's biometric information in real time. These terminals function as "information acquisition means" for acquiring biometric information such as heart rate and body temperature, and acquire voice data using a microphone built into the terminal.

[0340] The acquired biometric information and voice data are transmitted to a server via the network. After receiving this data, the server uses an "emotion analysis tool" that implements a machine learning algorithm to analyze the user's emotional state. To do this, the server uses advanced machine learning software to analyze the tone, pitch, tempo, and other aspects of the voice data.

[0341] Based on the analysis results, the server uses a "music creation method" to generate music data that corresponds to the user's emotions. At this time, algorithms such as generative opposition learning models and recurrent neural networks are utilized to generate music that fits the user's emotions.

[0342] The generated music data is transmitted to the user's device in real time. The device has music playback capabilities, and the user can play the generated music through their smartphone or smart speaker to relax or uplift their emotions.

[0343] For example, if a user wants to relax, the system will automatically generate relaxing music and provide it to the user's device. Another example prompt is, "If the user's heart rate is elevated and a desire to be energized is detected, suggest a music track that generates uplifting music." The server processes this prompt, generating and providing music that matches the user's mood at that moment.

[0344] The flow of a specific process in Application Example 2 will be explained using Figure 14.

[0345] Step 1:

[0346] The terminal acquires the user's biometric information (heart rate, body temperature, etc.) via a device such as a smartwatch. This input data is measured in real time. Based on these results, the terminal packages the data and similarly acquires voice data, preparing to send it to the server.

[0347] Step 2:

[0348] The device transmits biometric information and voice data to the server. The input at this time is the biometric information acquired by the device and the recorded voice data. The server receives this data, appropriately decodes it, and formats it into the format necessary for analysis.

[0349] Step 3:

[0350] The server inputs the received biometric and voice data into a machine learning model for emotion analysis. In this step, the user's emotional state is identified by analyzing features such as voice tone, pitch, and tempo, and the analysis results are output. The results include the main emotions the user is feeling and their intensity.

[0351] Step 4:

[0352] The server generates music data using a generative AI model based on the emotion analysis results. This algorithm utilizes generative opposing networks and recurrent neural networks to create music that corresponds to the analyzed emotions. At this point, the input is the emotion analysis results, and the output is the generated music data.

[0353] Step 5:

[0354] The server streams the generated music data to the terminal in real time. The terminal receives this music and prepares it for playback by the user. The input is the generated music data, and the output is the music track ready for the user to listen to.

[0355] Step 6:

[0356] Users listen to music streamed through their devices. Based on user feedback and usage patterns, the server collects further data and continuously improves the system accordingly. This feedback mechanism enables specific actions to enhance the individual user experience.

[0357] The specific processing unit 290 transmits the result of the specific processing to the smart glasses 214. In the smart glasses 214, the control unit 46A causes the speaker 240 to output the result of the specific processing. The microphone 238 acquires audio indicating user input for the result of the specific processing. The control unit 46A transmits the audio data indicating user input acquired by the microphone 238 to the data processing unit 12. In the data processing unit 12, the specific processing unit 290 acquires the audio data.

[0358] Data generation model 58 is a type of so-called generative AI (Artificial Intelligence). One example of data generation model 58 is ChatGPT (Internet search<URL: https: / / openai.com / blog / chatgpt> ), Gemini (Internet search) <url: https: gemini.google.com ?hl="ja">Examples of generative AI include the following. The data generation model 58 is obtained by performing deep learning on a neural network. The data generation model 58 is input with prompts containing instructions, and with inference data such as audio data representing speech, text data representing text, and image data representing images. The data generation model 58 infers from the input inference data according to the instructions indicated by the prompts, and outputs the inference results in data formats such as audio data and text data. Here, inference refers to, for example, analysis, classification, prediction, and / or summarization.

[0359] In the above embodiment, an example was given in which specific processing is performed by the data processing device 12, but the technology of this disclosure is not limited thereto, and the specific processing may also be performed by the smart glasses 214.

[0360] [Third Embodiment]

[0361] Figure 5 shows an example of the configuration of the data processing system 310 according to the third embodiment.

[0362] As shown in Figure 5, the data processing system 310 includes a data processing device 12 and a headset terminal 314. An example of the data processing device 12 is a server.

[0363] The data processing device 12 comprises a computer 22, a database 24, and a communication interface 26. The computer 22 is an example of a "computer" related to the technology of this disclosure. The computer 22 comprises a processor 28, RAM 30, and storage 32. The processor 28, RAM 30, and storage 32 are connected to a bus 34. The database 24 and the communication interface 26 are also connected to the bus 34. The communication interface 26 is connected to a network 54. An example of the network 54 is a WAN (Wide Area Network) and / or a LAN (Local Area Network).

[0364] The headset terminal 314 includes a computer 36, a microphone 238, a speaker 240, a camera 42, a communication interface 44, and a display 343. The computer 36 includes a processor 46, RAM 48, and storage 50. The processor 46, RAM 48, and storage 50 are connected to a bus 52. The microphone 238, speaker 240, camera 42, and display 343 are also connected to the bus 52.

[0365] The microphone 238 receives voice signals from the user 20 and receives instructions from the user 20. The microphone 238 captures the voice signals from the user 20, converts the captured voice into audio data, and outputs it to the processor 46. The speaker 240 outputs audio according to the instructions from the processor 46.

[0366] Camera 42 is a small digital camera equipped with an optical system including a lens, aperture, and shutter, and an image sensor such as a CMOS (Complementary Metal-Oxide-Semiconductor) image sensor or a CCD (Charge Coupled Device) image sensor, and captures images of the area around the user 20 (for example, an imaging range defined by a field of view equivalent to the width of a typical healthy person's field of vision).

[0367] Communication interface 44 is connected to network 54. Communication interfaces 44 and 26 are responsible for the exchange of various information between processor 46 and processor 28 via network 54. The exchange of various information between processor 46 and processor 28 using communication interfaces 44 and 26 is performed in a secure manner.

[0368] Figure 6 shows an example of the main functions of the data processing device 12 and the headset terminal 314. As shown in Figure 6, the data processing device 12 performs specific processing using the processor 28. The storage 32 stores the specific processing program 56.

[0369] The specific processing program 56 is an example of a "program" relating to the technology of this disclosure. The processor 28 reads the specific processing program 56 from the storage 32 and executes the read specific processing program 56 on the RAM 30. The specific processing is realized by the processor 28 operating as a specific processing unit 290 in accordance with the specific processing program 56 executed on the RAM 30.

[0370] The storage 32 stores the data generation model 58 and the emotion identification model 59. The data generation model 58 and the emotion identification model 59 are used by the identification processing unit 290.

[0371] In the headset terminal 314, the processor 46 performs the reception output processing. The storage 50 stores the reception output program 60. The processor 46 reads the reception output program 60 from the storage 50 and executes the read reception output program 60 on the RAM 48. The reception output processing is realized by the processor 46 operating as a control unit 46A according to the reception output program 60 executed on the RAM 48.

[0372] Next, the specific processing performed by the specific processing unit 290 of the data processing device 12 will be described. In the following description, the data processing device 12 will be referred to as the "server" and the headset terminal 314 will be referred to as the "terminal".

[0373] The system of the present invention supports the user's psychological health by acquiring the user's biometric information and voice data in real time, analyzing the user's emotional state based on this data, and generating and providing appropriate music.

[0374] In this system, the device first collects data from smartwatches and monitoring cameras. Specifically, it acquires heart rate and voice data and sends this information to a server. When using this system, users can naturally provide data from devices they wear daily without any physical burden.

[0375] The server uses an emotion analysis model to identify the user's current emotional state based on the received biometric and voice data. This emotion analysis is performed by comprehensively analyzing acoustic characteristics such as voice tone and pitch, as well as physiological data such as changes in heart rate and body temperature.

[0376] After the emotional state is identified, the server uses music generation tools to create music appropriate to the user's emotions. This process utilizes prompts to take into account the user's musical preferences and cultural background, allowing for the provision of music that is highly relatable to the user.

[0377] Ultimately, the music generated from the server is streamed to the user's device, allowing them to listen to the music in real time and use it to relax, concentrate, or boost their motivation. In this way, the system provides a real-time musical experience tailored to the user's psychological needs.

[0378] For example, if a user feels stressed while working, the system can determine their stress level from an increase in heart rate and a change in voice tone, generate relaxing music, and immediately begin playing it. In this way, this invention contributes to improving the user's quality of life.

[0379] The following describes the processing flow.

[0380] Step 1:

[0381] The device transmits biometric and audio data acquired from smartwatches and monitoring cameras to the server in real time. The data is transmitted encrypted to protect user privacy.

[0382] Step 2:

[0383] The server inputs the received biometric and voice data into the data analysis module and cleanses the data. In this process, missing values ​​are imputed and outliers are removed, preparing clean data for analysis.

[0384] Step 3:

[0385] The server uses cleansed data to apply acoustic feature extraction algorithms, extracting features such as tone and pitch from the audio data. It also analyzes changes in heart rate and body temperature from biometric information and uses this data to run an emotion analysis model.

[0386] Step 4:

[0387] The server updates user profiles based on sentiment scores obtained from sentiment analysis models. These profiles take into account each user's characteristics and historical data, and also reflect long-term trend analysis.

[0388] Step 5:

[0389] The server determines an appropriate music style that corresponds to the user's current emotional state. Then, a music generation algorithm (such as GAN or RNN) is used to generate an original music track based on the selected style.

[0390] Step 6:

[0391] The server sends the generated music data to the user's device, allowing the user to play and enjoy it. If user feedback is received, this data is also saved and used to improve the application in the future.

[0392] (Example 1)

[0393] Next, we will describe Example 1. In the following description, the data processing device 12 will be referred to as the "server," and the headset-type terminal 314 will be referred to as the "terminal."

[0394] In modern society, psychological stress and anxiety are problems faced by many people, and there is a need to provide means to appropriately alleviate them. However, conventional methods have the challenge of not being able to identify the emotional state of individual users in real time and provide relaxation that is appropriate to that state.

[0395] The identification process performed by the identification processing unit 290 of the data processing device 12 in Example 1 is realized by the following means.

[0396] In this invention, the server includes data collection means for acquiring the user's biometric indicators and voice information, analysis means for analyzing the biometric indicators and voice information to identify the user's emotional state, and generation means for generating music that reflects the user's preferences and cultural background in relation to the emotional state. This makes it possible to provide the user with a personalized music experience in real time, thereby reducing psychological stress and anxiety.

[0397] A "user" is an individual who uses the system and provides biometric information and voice data.

[0398] "Biometric indicators" refer to data that numerically represents a user's physical condition, such as heart rate and body temperature.

[0399] "Audio information" refers to data that includes acoustic characteristics such as the tone and pitch of the user's voice.

[0400] "Data collection means" refers to a device or method for detecting and acquiring a user's biometric indicators and voice information.

[0401] "Analysis means" refers to a device or method that analyzes acquired biometric indicators and voice information to identify the user's emotional state.

[0402] "Emotional state" refers to information that indicates the user's mental state, and includes specific psychological conditions such as stress, stability, and happiness.

[0403] "Generating means" refers to a device or method for creating music in relation to the user's emotional state.

[0404] "Preferences" refer to information about the music genres and styles that individual users particularly enjoy.

[0405] "Cultural background" refers to information about a user's expectations and preferences regarding music, based on their lifestyle and social and regional characteristics.

[0406] This system supports users' psychological well-being by acquiring biometric and voice information from users in real time, analyzing their emotional state based on that data, and generating and providing appropriate music.

[0407] The device collects biometric and audio information using smartwatches and monitoring cameras that users use daily. This involves using hardware such as heart rate sensors and microphones to acquire the user's physical and acoustic data. This data is transmitted to a server using communication protocols such as Bluetooth and Wi-Fi.

[0408] The server is equipped with a generative AI model for analyzing received biometric and voice information. This model comprehensively analyzes acoustic features such as voice tone and pitch, as well as physiological data such as heart rate and body temperature, to identify the user's emotional state. The generative AI model can also consider the user's musical preferences and cultural background using prompts. For example, it can use prompts such as, "Your current emotional state is stress; please provide relaxing jazz music."

[0409] After identifying the user's emotional state, the server uses a generative AI model to automatically generate music appropriate to the user's emotions. The generated music is streamed to the user's device and provided in real time. This allows the user to relax, concentrate, or improve their motivation. This system aims to improve the user's quality of life by providing a music experience optimized for each individual user.

[0410] The flow of the specific processing in Example 1 will be explained using Figure 11.

[0411] Step 1:

[0412] The device uses the user's smartwatch or monitoring camera to collect biometric data (e.g., heart rate) and audio information (e.g., voice tone) in real time. This data is acquired by the device as input and prepared for the next processing step. Specifically, the heart rate sensor periodically records the pulse, and the microphone records the user's voice.

[0413] Step 2:

[0414] The device transmits collected biometric and voice information to a server using network technologies such as Bluetooth or Wi-Fi. The output of this step is composed of transmission data and prepared for analysis on the server. Specifically, data packets are transferred to the server via network communication.

[0415] Step 3:

[0416] The server analyzes the emotional state using a generative AI model based on the received data. Inputs include heart rate variability and voice tone patterns, and the output identifies the user's current emotional state (e.g., stress, relaxation). Specifically, the AI ​​model adapts the data to prompt statements to reveal the emotional state.

[0417] Step 4:

[0418] The server uses music generation methods based on the user's emotional state to create music suitable for the user. The generation AI model processes prompts such as "Please generate music for relaxation," and music data is constructed as output. Specifically, the AI ​​algorithm creates music while considering the user's musical preferences and cultural background.

[0419] Step 5:

[0420] The server then streams the generated music data back to the terminal. The input for this step is the generated music data, and the output is provided in a format playable on the user's device. Specifically, the music data stream is transmitted to the terminal via the network, and playback takes place in real time.

[0421] In this way, the system aims to provide users with a music experience tailored to their individual needs and to reduce psychological stress.

[0422] (Application Example 1)

[0423] Next, we will explain Application Example 1. In the following explanation, the data processing device 12 will be referred to as the "server," and the headset-type terminal 314 will be referred to as the "terminal."

[0424] In production environments, workers' mental stress and decreased concentration significantly impact work efficiency. This increases the risk of reduced quality and accidents, making it crucial to understand workers' psychological states in real time and take appropriate action. Conventional methods have struggled to quickly and accurately assess workers' emotional states and provide appropriate relaxation measures.

[0425] The specific processing performed by the specific processing unit 290 of the data processing device 12 in Application Example 1 is realized by the following means.

[0426] In this invention, the server includes information gathering means for acquiring the user's biometric information, emotion analysis means for analyzing the user's emotional state based on the biometric information, and music generation means for generating music data according to the emotional state. This makes it possible to automatically provide music suitable for the work environment while monitoring the psychological state of the worker.

[0427] "Information gathering means" refers to a device or method for acquiring a user's biometric information, such as voice data or heart rate data.

[0428] "Emotional analysis means" refers to a process or device for analyzing and identifying a user's emotional state based on acquired biometric information.

[0429] "Music generation means" refers to a technology or device for generating optimal music data according to an analyzed emotional state.

[0430] "Music delivery means" refers to a method or device for outputting generated music data to a user or working environment.

[0431] "Work environment" refers to the physical or virtual workplace where music data is output by the music delivery means.

[0432] The system of this invention is implemented via a terminal and server installed in the manufacturing environment. The terminal is connected to information collection devices such as a smartwatch and a camera, which acquire the worker's biometric information in real time. Specifically, the terminal collects the worker's heart rate and voice data and transmits the data to the server via the network.

[0433] The server runs analysis software using programming languages ​​such as Python and R, which implements an emotion analysis model. Pandas and NumPy are used for data processing, including cleaning and preprocessing. The server inputs received biometric information into the emotion analysis model and analyzes the user's emotional state using machine learning frameworks such as TensorFlow and PyTorch.

[0434] Based on the analyzed emotional state, the server utilizes a generative AI model to generate music. Open-source music generation tools and APIs are used to create customized music tailored to the user's psychological state. The generated music data is streamed to the terminal through the work environment's speaker system. This allows workers to listen to stress-reducing music in real time.

[0435] As a concrete example, when a worker feels stressed, the system generates relaxing music based on an increase in heart rate and a change in voice tone, and distributes it to the workspace. An example of a prompt message to the generation AI model would be, "Generate music that will help employees relax, and distribute that music through the speaker system."

[0436] The flow of a specific process in Application Example 1 will be explained using Figure 12.

[0437] Step 1:

[0438] The device acquires biometric information from a smartwatch and a monitoring camera. It receives user heart rate and voice data as input. This data is collected via a sensor interface and converted to JSON format for portability.

[0439] Step 2:

[0440] The device transmits collected biometric information to the server. Input data includes heart rate, voice pitch, and tone. The data is transferred to the server via a REST API, and the server then incorporates it into its data analysis system.

[0441] Step 3:

[0442] The server analyzes the received biometric information. The input data includes heart rate and voice data, and data cleaning and normalization are performed using Pandas and NumPy. This prepares the input for the sentiment analysis model. A machine learning framework (e.g., TensorFlow) is used for sentiment analysis, and the analysis results identify the user's current emotional state.

[0443] Step 4:

[0444] The server generates music data based on the analyzed emotional state. It receives emotional state data as input and uses a generative AI model (e.g., a music generation API) to generate customized music. The output includes music data.

[0445] Step 5:

[0446] The server streams the generated music data to terminals in the work environment. The terminals output the music through a speaker system. Employees can listen to this music to relax or improve their concentration. The input is the generated music data, and the output is music playback from the speakers.

[0447] Furthermore, an emotion engine that estimates the user's emotions may be incorporated. That is, the identification processing unit 290 may use the emotion identification model 59 to estimate the user's emotions and perform identification processing using the user's emotions.

[0448] This invention relates to a system that recognizes emotions based on a user's biometric information and voice data, and provides music that corresponds to the user's emotional state. This system achieves highly accurate emotion analysis by incorporating an emotion engine, and generates music using the results of the emotion analysis.

[0449] The device acquires biometric information such as heart rate, body temperature, and voice data in real time from smartwatches and monitoring cameras, and transmits it to a server. Users can use this system naturally in their daily lives without requiring any special operation.

[0450] The server receives biometric information and voice data transmitted from the terminal and inputs this data into the emotion engine. The emotion engine uses advanced machine learning algorithms to analyze features such as tone, pitch, and tempo of the voice data to precisely identify the user's emotions. It can also recognize patterns of emotional fluctuations by referring to past emotional state history and predict current emotions.

[0451] Based on these results, the server generates music data that corresponds to the user's current and predicted emotional state. Modern algorithms such as generative opposing networks and recurrent neural networks are used for music generation to create music tracks that match the user's preferences and emotions.

[0452] The generated music is streamed to the user's device, allowing the user to play the music through their mobile device or smart speaker for relaxation or a change of pace. For example, if a user is feeling stressed at work, the system will provide music with a relaxing tone to alleviate stress. In this way, the present invention realizes an emotion-responsive music experience tailored to individual users, contributing to the promotion of the user's psychological well-being.

[0453] The following describes the processing flow.

[0454] Step 1:

[0455] The device acquires the user's biometric information and voice data in real time from smartwatches and monitoring cameras. This data includes heart rate, body temperature, and voice tone.

[0456] Step 2:

[0457] The device encrypts the acquired data and sends it to the server using a secure communication protocol. This protects the user's privacy.

[0458] Step 3:

[0459] The server inputs the received data into the emotion engine. The emotion engine analyzes the tone, pitch, and tempo of the audio data and runs a machine learning model to identify the user's emotions.

[0460] Step 4:

[0461] The server updates the user profile using emotion identification results obtained from the emotion engine. This profile is used to store past emotional state information and understand the user's patterns over the long term.

[0462] Step 5:

[0463] The server uses the latest music generation algorithms to generate music data based on emotion analysis. By employing generative opposing networks and recurrent neural networks, it generates music tracks that match individual emotions.

[0464] Step 6:

[0465] The server streams the generated music data to the user's device. Users can then enjoy the music in real time via their mobile devices or smart speakers.

[0466] Step 7:

[0467] Users can relax and change their mood through music suited to their emotions. They can also provide feedback as needed to help improve the system.

[0468] (Example 2)

[0469] Next, we will describe Example 2. In the following description, the data processing device 12 will be referred to as the "server," and the headset-type terminal 314 will be referred to as the "terminal."

[0470] In modern society, there is a need to manage and appropriately respond to people's emotions and stress, but conventional systems have difficulty providing music that is appropriate to the user's physiological and psychological state. Therefore, users have to select music that suits their emotions themselves, which presents a challenge in terms of efficient emotional management.

[0471] The identification process performed by the identification processing unit 290 of the data processing device 12 in Example 2 is realized by the following means.

[0472] In this invention, the server includes data collection means for acquiring the user's biometric data, emotion analysis means for analyzing the user's emotional state based on the biometric data and voice data, emotion prediction means for predicting emotional fluctuation patterns considering the emotional state and past emotional history, and music generation means for generating music data based on the predicted emotional state. This allows the user to passively receive music appropriate to their state, making it easier to manage their emotions on a daily basis.

[0473] "Biometric data" refers to measurable information about a user's body, such as heart rate and body temperature, which indicate the user's physiological state.

[0474] "Data collection means" refers to devices and technologies that can acquire biometric data or voice data from users, such as devices that use sensors or microphones.

[0475] "Emotional analysis means" refers to algorithms and technologies that analyze acquired biometric and voice data to identify the user's emotional state.

[0476] "Emotion prediction methods" refer to technologies and models that predict a user's emotional fluctuation patterns based on current and past emotional data.

[0477] "Music generation means" refers to a technology for creating music data suitable for the user according to their analyzed emotional state, and is a function that generates music tracks using a generation AI model.

[0478] A "generative AI model" is an artificial intelligence model that has the ability to learn from data and generate new data, and specifically includes generative opposing networks and recurrent neural networks.

[0479] This invention is a system for providing a music experience tailored to the physiological and psychological characteristics of users. This system collects the user's biometric and vocal data in real time, analyzes the user's emotional state based on that data, and generates music data according to the analysis results.

[0480] The device uses the user's smartwatch or monitoring camera to acquire physiological data such as heart rate and body temperature, as well as voice data. This allows for the accurate and continuous collection of the user's biometric data. The device uses a secure communication protocol to safely transmit this data to the server.

[0481] The server receives the transmitted data and uses sentiment analysis tools to accurately identify the user's emotional state. This analysis utilizes machine learning algorithms and tools to analyze features such as tone, pitch, and tempo of the voice. Furthermore, sentiment prediction tools allow the server to refer to past emotional history and predict not only the current emotional state but also its fluctuation patterns.

[0482] The server then uses a generative AI model to generate music data tailored to the user's emotions. Specifically, it uses generative opposing networks (GANs) and recurrent neural networks (RNNs) to create music tracks that the user will like. This generated music is then delivered to the user's device via streaming. The user can play this music on their mobile device or smart speaker to help them relax or change their mood.

[0483] For example, when a user feels anxious in a large crowd, this system analyzes voice and biometric data to generate anxiety-relieving music. By listening to this music, the user can reduce stress and regain calmness.

[0484] An example of a prompt message is: "If the emotional state obtained from the user's biometric information is anxious, generate a music track to alleviate anxiety." This format allows you to instruct the generation AI model in this way.

[0485] In this way, the present invention provides a music experience tailored to the user's emotional state, realizing a system that contributes to maintaining and improving psychological health.

[0486] The flow of the specific processing in Example 2 will be explained using Figure 13.

[0487] Step 1:

[0488] The device acquires biometric information and voice data through the user's smartwatch or monitoring camera. Specifically, the smartwatch's sensors measure heart rate and body temperature, and voice data is recorded by the camera's microphone. This input data is sent to the server as physiological indicators and voice characteristics.

[0489] Step 2:

[0490] The server receives biometric information and voice data transmitted from the terminal. Input data includes heart rate, body temperature, voice tone, pitch, and tempo. The server preprocesses the data, removing noise to make it suitable for analysis. This processed data is then passed to the emotion analysis system.

[0491] Step 3:

[0492] The server analyzes pre-processed data using emotion analysis tools. The audio data is analyzed for changes in tone and rhythm, and combined with heart rate and body temperature to determine the user's emotional state. The user's emotional state is output as a result of this calculation.

[0493] Step 4:

[0494] The server uses emotion prediction tools in addition to emotion analysis tools to refer to the user's past emotion data and predict how their current emotional state will change. By utilizing past records, it calculates patterns of temporal changes in emotions and outputs predicted emotion values.

[0495] Step 5:

[0496] The server generates music data using a generative AI model based on the obtained emotional state and predicted values. Specifically, it uses generative opposing networks (GANs) and recurrent neural networks (RNNs) to create music tracks that are optimal for the user's emotions. Personalized music is output based on the input emotional information and the user's musical preferences.

[0497] Step 6:

[0498] The generated music data is streamed from the server to the user's device. The user plays the music on their smartphone or smart speaker and enjoys its relaxing effects. This process involves the user receiving the music and taking actual actions to maintain their psychological well-being and comfort.

[0499] (Application Example 2)

[0500] Next, we will explain application example 2. In the following explanation, the data processing device 12 will be referred to as the "server," and the headset-type terminal 314 will be referred to as the "terminal."

[0501] Conventional music delivery systems struggle to select music that matches a user's emotional state, making it difficult to provide a musical experience that responds to the instantaneous emotional changes of individual users. Furthermore, existing services lack sufficient mechanisms to analyze biometric information and voice data in real time and generate and provide appropriate music. Therefore, there is a need for technology that efficiently delivers music tailored to a user's emotional state and improves satisfaction.

[0502] The specific processing performed by the specific processing unit 290 of the data processing device 12 in Application Example 2 is realized by the following means.

[0503] In this invention, the server includes information acquisition means for acquiring the user's biometric information, emotion analysis means for analyzing the user's emotional state based on the biometric information and voice data, and music creation means for generating music data according to the emotional state. This makes it possible to automatically generate and provide music that is tailored to the emotional state of each individual user.

[0504] "Information acquisition means" refers to a device or mechanism that has the function of acquiring a user's biometric information or voice data.

[0505] "Emotional analysis means" refers to a device or algorithm that has the function of analyzing and identifying a user's emotional state based on acquired biometric information and voice data.

[0506] "Music creation means" refers to a device or algorithm that has the function of generating music data in accordance with the analyzed emotional state.

[0507] "Music transmission means" refers to a device or structure that has the function of transmitting and providing generated music data to a user.

[0508] To implement this invention, it is desirable to use a smartwatch or smartphone as a terminal for acquiring the user's biometric information in real time. These terminals function as "information acquisition means" for acquiring biometric information such as heart rate and body temperature, and acquire voice data using a microphone built into the terminal.

[0509] The acquired biometric information and voice data are transmitted to a server via the network. After receiving this data, the server uses an "emotion analysis tool" that implements a machine learning algorithm to analyze the user's emotional state. To do this, the server uses advanced machine learning software to analyze the tone, pitch, tempo, and other aspects of the voice data.

[0510] Based on the analysis results, the server uses a "music creation method" to generate music data that corresponds to the user's emotions. At this time, algorithms such as generative opposition learning models and recurrent neural networks are utilized to generate music that fits the user's emotions.

[0511] The generated music data is transmitted to the user's device in real time. The device has music playback capabilities, and the user can play the generated music through their smartphone or smart speaker to relax or uplift their emotions.

[0512] For example, if a user wants to relax, the system will automatically generate relaxing music and provide it to the user's device. Another example prompt is, "If the user's heart rate is elevated and a desire to be energized is detected, suggest a music track that generates uplifting music." The server processes this prompt, generating and providing music that matches the user's mood at that moment.

[0513] The flow of a specific process in Application Example 2 will be explained using Figure 14.

[0514] Step 1:

[0515] The terminal acquires the user's biometric information (heart rate, body temperature, etc.) via a device such as a smartwatch. This input data is measured in real time. Based on these results, the terminal packages the data and similarly acquires voice data, preparing to send it to the server.

[0516] Step 2:

[0517] The device transmits biometric information and voice data to the server. The input at this time is the biometric information acquired by the device and the recorded voice data. The server receives this data, appropriately decodes it, and formats it into the format necessary for analysis.

[0518] Step 3:

[0519] The server inputs the received biometric and voice data into a machine learning model for emotion analysis. In this step, the user's emotional state is identified by analyzing features such as voice tone, pitch, and tempo, and the analysis results are output. The results include the main emotions the user is feeling and their intensity.

[0520] Step 4:

[0521] The server generates music data using a generative AI model based on the emotion analysis results. This algorithm utilizes generative opposing networks and recurrent neural networks to create music that corresponds to the analyzed emotions. At this point, the input is the emotion analysis results, and the output is the generated music data.

[0522] Step 5:

[0523] The server streams the generated music data to the terminal in real time. The terminal receives this music and prepares it for playback by the user. The input is the generated music data, and the output is the music track ready for the user to listen to.

[0524] Step 6:

[0525] Users listen to music streamed through their devices. Based on user feedback and usage patterns, the server collects further data and continuously improves the system accordingly. This feedback mechanism enables specific actions to enhance the individual user experience.

[0526] The specific processing unit 290 transmits the result of the specific processing to the headset terminal 314. In the headset terminal 314, the control unit 46A causes the speaker 240 and display 343 to output the result of the specific processing. The microphone 238 acquires audio indicating user input for the result of the specific processing. The control unit 46A transmits the audio data indicating user input acquired by the microphone 238 to the data processing unit 12. In the data processing unit 12, the specific processing unit 290 acquires the audio data.

[0527] Data generation model 58 is a type of so-called generative AI (Artificial Intelligence). One example of data generation model 58 is ChatGPT (Internet search<URL: https: / / openai.com / blog / chatgpt> ), Gemini (Internet search) <url: https: gemini.google.com ?hl="ja">Examples of generative AI include the following. The data generation model 58 is obtained by performing deep learning on a neural network. The data generation model 58 is input with prompts containing instructions, and with inference data such as audio data representing speech, text data representing text, and image data representing images. The data generation model 58 infers from the input inference data according to the instructions indicated by the prompts, and outputs the inference results in data formats such as audio data and text data. Here, inference refers to, for example, analysis, classification, prediction, and / or summarization.

[0528] In the above embodiment, an example was given in which specific processing is performed by the data processing device 12, but the technology of this disclosure is not limited thereto, and specific processing may also be performed by the headset terminal 314.

[0529] [Fourth Embodiment]

[0530] Figure 7 shows an example of the configuration of the data processing system 410 according to the fourth embodiment.

[0531] As shown in Figure 7, the data processing system 410 includes a data processing device 12 and a robot 414. An example of the data processing device 12 is a server.

[0532] The data processing device 12 comprises a computer 22, a database 24, and a communication interface 26. The computer 22 is an example of a "computer" related to the technology of this disclosure. The computer 22 comprises a processor 28, RAM 30, and storage 32. The processor 28, RAM 30, and storage 32 are connected to a bus 34. The database 24 and the communication interface 26 are also connected to the bus 34. The communication interface 26 is connected to a network 54. An example of the network 54 is a WAN (Wide Area Network) and / or a LAN (Local Area Network).

[0533] The robot 414 includes a computer 36, a microphone 238, a speaker 240, a camera 42, a communication interface 44, and a controlled object 443. The computer 36 includes a processor 46, RAM 48, and storage 50. The processor 46, RAM 48, and storage 50 are connected to a bus 52. The microphone 238, speaker 240, camera 42, and controlled object 443 are also connected to the bus 52.

[0534] The microphone 238 receives voice signals from the user 20 and receives instructions from the user 20. The microphone 238 captures the voice signals from the user 20, converts the captured voice into audio data, and outputs it to the processor 46. The speaker 240 outputs audio according to the instructions from the processor 46.

[0535] Camera 42 is a small digital camera equipped with an optical system including a lens, aperture, and shutter, and an image sensor such as a CMOS (Complementary Metal-Oxide-Semiconductor) image sensor or a CCD (Charge Coupled Device) image sensor, and captures images of the area around the user 20 (for example, an imaging range defined by a field of view equivalent to the width of a typical healthy person's field of vision).

[0536] Communication interface 44 is connected to network 54. Communication interfaces 44 and 26 are responsible for the exchange of various information between processor 46 and processor 28 via network 54. The exchange of various information between processor 46 and processor 28 using communication interfaces 44 and 26 is performed in a secure manner.

[0537] The controlled object 443 includes a display device, LEDs in the eyes, and motors that drive the arms, hands, and feet. The posture and gestures of the robot 414 are controlled by controlling the motors of the arms, hands, and feet. Some of the robot 414's emotions can be expressed by controlling these motors. Furthermore, the robot 414's facial expressions can also be expressed by controlling the illumination state of the LEDs in its eyes.

[0538] Figure 8 shows an example of the main functions of the data processing device 12 and the robot 414. As shown in Figure 8, the data processing device 12 performs specific processing using the processor 28. The storage 32 stores the specific processing program 56.

[0539] The specific processing program 56 is an example of a "program" relating to the technology of this disclosure. The processor 28 reads the specific processing program 56 from the storage 32 and executes the read specific processing program 56 on the RAM 30. The specific processing is realized by the processor 28 operating as a specific processing unit 290 in accordance with the specific processing program 56 executed on the RAM 30.

[0540] The storage 32 stores the data generation model 58 and the emotion identification model 59. The data generation model 58 and the emotion identification model 59 are used by the identification processing unit 290.

[0541] In robot 414, the processor 46 performs the reception output processing. The storage 50 stores the reception output program 60. The processor 46 reads the reception output program 60 from the storage 50 and executes the read reception output program 60 on the RAM 48. The reception output processing is realized by the processor 46 operating as a control unit 46A according to the reception output program 60 executed on the RAM 48.

[0542] Next, the specific processing performed by the specific processing unit 290 of the data processing device 12 will be described. In the following description, the data processing device 12 will be referred to as the "server" and the robot 414 as the "terminal".

[0543] The system of the present invention supports the user's psychological health by acquiring the user's biometric information and voice data in real time, analyzing the user's emotional state based on this data, and generating and providing appropriate music.

[0544] In this system, the device first collects data from smartwatches and monitoring cameras. Specifically, it acquires heart rate and voice data and sends this information to a server. When using this system, users can naturally provide data from devices they wear daily without any physical burden.

[0545] The server uses an emotion analysis model to identify the user's current emotional state based on the received biometric and voice data. This emotion analysis is performed by comprehensively analyzing acoustic characteristics such as voice tone and pitch, as well as physiological data such as changes in heart rate and body temperature.

[0546] After the emotional state is identified, the server uses music generation tools to create music appropriate to the user's emotions. This process utilizes prompts to take into account the user's musical preferences and cultural background, allowing for the provision of music that is highly relatable to the user.

[0547] Ultimately, the music generated from the server is streamed to the user's device, allowing them to listen to the music in real time and use it to relax, concentrate, or boost their motivation. In this way, the system provides a real-time musical experience tailored to the user's psychological needs.

[0548] For example, if a user feels stressed while working, the system can determine their stress level from an increase in heart rate and a change in voice tone, generate relaxing music, and immediately begin playing it. In this way, this invention contributes to improving the user's quality of life.

[0549] The following describes the processing flow.

[0550] Step 1:

[0551] The device transmits biometric and audio data acquired from smartwatches and monitoring cameras to the server in real time. The data is transmitted encrypted to protect user privacy.

[0552] Step 2:

[0553] The server inputs the received biometric and voice data into the data analysis module and cleanses the data. In this process, missing values ​​are imputed and outliers are removed, preparing clean data for analysis.

[0554] Step 3:

[0555] The server uses cleansed data to apply acoustic feature extraction algorithms, extracting features such as tone and pitch from the audio data. It also analyzes changes in heart rate and body temperature from biometric information and uses this data to run an emotion analysis model.

[0556] Step 4:

[0557] The server updates user profiles based on sentiment scores obtained from sentiment analysis models. These profiles take into account each user's characteristics and historical data, and also reflect long-term trend analysis.

[0558] Step 5:

[0559] The server determines an appropriate music style that corresponds to the user's current emotional state. Then, a music generation algorithm (such as GAN or RNN) is used to generate an original music track based on the selected style.

[0560] Step 6:

[0561] The server sends the generated music data to the user's device, allowing the user to play and enjoy it. If user feedback is received, this data is also saved and used to improve the application in the future.

[0562] (Example 1)

[0563] Next, we will describe Example 1. In the following description, the data processing device 12 will be referred to as the "server" and the robot 414 as the "terminal".

[0564] In modern society, psychological stress and anxiety are problems faced by many people, and there is a need to provide means to appropriately alleviate them. However, conventional methods have the challenge of not being able to identify the emotional state of individual users in real time and provide relaxation that is appropriate to that state.

[0565] The identification process performed by the identification processing unit 290 of the data processing device 12 in Example 1 is realized by the following means.

[0566] In this invention, the server includes data collection means for acquiring the user's biometric indicators and voice information, analysis means for analyzing the biometric indicators and voice information to identify the user's emotional state, and generation means for generating music that reflects the user's preferences and cultural background in relation to the emotional state. This makes it possible to provide the user with a personalized music experience in real time, thereby reducing psychological stress and anxiety.

[0567] A "user" is an individual who uses the system and provides biometric information and voice data.

[0568] "Biometric indicators" refer to data that numerically represents a user's physical condition, such as heart rate and body temperature.

[0569] "Audio information" refers to data that includes acoustic characteristics such as the tone and pitch of the user's voice.

[0570] "Data collection means" refers to a device or method for detecting and acquiring a user's biometric indicators and voice information.

[0571] "Analysis means" refers to a device or method that analyzes acquired biometric indicators and voice information to identify the user's emotional state.

[0572] "Emotional state" refers to information that indicates the user's mental state, and includes specific psychological conditions such as stress, stability, and happiness.

[0573] "Generating means" refers to a device or method for creating music in relation to the user's emotional state.

[0574] "Preferences" refer to information about the music genres and styles that individual users particularly enjoy.

[0575] "Cultural background" refers to information about a user's expectations and preferences regarding music, based on their lifestyle and social and regional characteristics.

[0576] This system supports users' psychological well-being by acquiring biometric and voice information from users in real time, analyzing their emotional state based on that data, and generating and providing appropriate music.

[0577] The device collects biometric and audio information using smartwatches and monitoring cameras that users use daily. This involves using hardware such as heart rate sensors and microphones to acquire the user's physical and acoustic data. This data is transmitted to a server using communication protocols such as Bluetooth and Wi-Fi.

[0578] The server is equipped with a generative AI model for analyzing received biometric and voice information. This model comprehensively analyzes acoustic features such as voice tone and pitch, as well as physiological data such as heart rate and body temperature, to identify the user's emotional state. The generative AI model can also consider the user's musical preferences and cultural background using prompts. For example, it can use prompts such as, "Your current emotional state is stress; please provide relaxing jazz music."

[0579] After identifying the user's emotional state, the server uses a generative AI model to automatically generate music appropriate to the user's emotions. The generated music is streamed to the user's device and provided in real time. This allows the user to relax, concentrate, or improve their motivation. This system aims to improve the user's quality of life by providing a music experience optimized for each individual user.

[0580] The flow of the specific processing in Example 1 will be explained using Figure 11.

[0581] Step 1:

[0582] The device uses the user's smartwatch or monitoring camera to collect biometric data (e.g., heart rate) and audio information (e.g., voice tone) in real time. This data is acquired by the device as input and prepared for the next processing step. Specifically, the heart rate sensor periodically records the pulse, and the microphone records the user's voice.

[0583] Step 2:

[0584] The device transmits collected biometric and voice information to a server using network technologies such as Bluetooth or Wi-Fi. The output of this step is composed of transmission data and prepared for analysis on the server. Specifically, data packets are transferred to the server via network communication.

[0585] Step 3:

[0586] The server analyzes the emotional state using a generative AI model based on the received data. Inputs include heart rate variability and voice tone patterns, and the output identifies the user's current emotional state (e.g., stress, relaxation). Specifically, the AI ​​model adapts the data to prompt statements to reveal the emotional state.

[0587] Step 4:

[0588] The server uses music generation methods based on the user's emotional state to create music suitable for the user. The generation AI model processes prompts such as "Please generate music for relaxation," and music data is constructed as output. Specifically, the AI ​​algorithm creates music while considering the user's musical preferences and cultural background.

[0589] Step 5:

[0590] The server then streams the generated music data back to the terminal. The input for this step is the generated music data, and the output is provided in a format playable on the user's device. Specifically, the music data stream is transmitted to the terminal via the network, and playback takes place in real time.

[0591] In this way, the system aims to provide users with a music experience tailored to their individual needs and to reduce psychological stress.

[0592] (Application Example 1)

[0593] Next, we will explain Application Example 1. In the following explanation, the data processing device 12 will be referred to as the "server" and the robot 414 as the "terminal".

[0594] In production environments, workers' mental stress and decreased concentration significantly impact work efficiency. This increases the risk of reduced quality and accidents, making it crucial to understand workers' psychological states in real time and take appropriate action. Conventional methods have struggled to quickly and accurately assess workers' emotional states and provide appropriate relaxation measures.

[0595] The specific processing performed by the specific processing unit 290 of the data processing device 12 in Application Example 1 is realized by the following means.

[0596] In this invention, the server includes information gathering means for acquiring the user's biometric information, emotion analysis means for analyzing the user's emotional state based on the biometric information, and music generation means for generating music data according to the emotional state. This makes it possible to automatically provide music suitable for the work environment while monitoring the psychological state of the worker.

[0597] "Information gathering means" refers to a device or method for acquiring a user's biometric information, such as voice data or heart rate data.

[0598] "Emotional analysis means" refers to a process or device for analyzing and identifying a user's emotional state based on acquired biometric information.

[0599] "Music generation means" refers to a technology or device for generating optimal music data according to an analyzed emotional state.

[0600] "Music delivery means" refers to a method or device for outputting generated music data to a user or working environment.

[0601] "Work environment" refers to the physical or virtual workplace where music data is output by the music delivery means.

[0602] The system of this invention is implemented via a terminal and server installed in the manufacturing environment. The terminal is connected to information collection devices such as a smartwatch and a camera, which acquire the worker's biometric information in real time. Specifically, the terminal collects the worker's heart rate and voice data and transmits the data to the server via the network.

[0603] The server runs analysis software using programming languages ​​such as Python and R, which implements an emotion analysis model. Pandas and NumPy are used for data processing, including cleaning and preprocessing. The server inputs received biometric information into the emotion analysis model and analyzes the user's emotional state using machine learning frameworks such as TensorFlow and PyTorch.

[0604] Based on the analyzed emotional state, the server utilizes a generative AI model to generate music. Open-source music generation tools and APIs are used to create customized music tailored to the user's psychological state. The generated music data is streamed to the terminal through the work environment's speaker system. This allows workers to listen to stress-reducing music in real time.

[0605] As a concrete example, when a worker feels stressed, the system generates relaxing music based on an increase in heart rate and a change in voice tone, and distributes it to the workspace. An example of a prompt message to the generation AI model would be, "Generate music that will help employees relax, and distribute that music through the speaker system."

[0606] The flow of a specific process in Application Example 1 will be explained using Figure 12.

[0607] Step 1:

[0608] The device acquires biometric information from a smartwatch and a monitoring camera. It receives user heart rate and voice data as input. This data is collected via a sensor interface and converted to JSON format for portability.

[0609] Step 2:

[0610] The device transmits collected biometric information to the server. Input data includes heart rate, voice pitch, and tone. The data is transferred to the server via a REST API, and the server then incorporates it into its data analysis system.

[0611] Step 3:

[0612] The server analyzes the received biometric information. The input data includes heart rate and voice data, and data cleaning and normalization are performed using Pandas and NumPy. This prepares the input for the sentiment analysis model. A machine learning framework (e.g., TensorFlow) is used for sentiment analysis, and the analysis results identify the user's current emotional state.

[0613] Step 4:

[0614] The server generates music data based on the analyzed emotional state. It receives emotional state data as input and uses a generative AI model (e.g., a music generation API) to generate customized music. The output includes music data.

[0615] Step 5:

[0616] The server streams the generated music data to terminals in the work environment. The terminals output the music through a speaker system. Employees can listen to this music to relax or improve their concentration. The input is the generated music data, and the output is music playback from the speakers.

[0617] Furthermore, an emotion engine that estimates the user's emotions may be incorporated. That is, the identification processing unit 290 may use the emotion identification model 59 to estimate the user's emotions and perform identification processing using the user's emotions.

[0618] This invention relates to a system that recognizes emotions based on a user's biometric information and voice data, and provides music that corresponds to the user's emotional state. This system achieves highly accurate emotion analysis by incorporating an emotion engine, and generates music using the results of the emotion analysis.

[0619] The device acquires biometric information such as heart rate, body temperature, and voice data in real time from smartwatches and monitoring cameras, and transmits it to a server. Users can use this system naturally in their daily lives without requiring any special operation.

[0620] The server receives biometric information and voice data transmitted from the terminal and inputs this data into the emotion engine. The emotion engine uses advanced machine learning algorithms to analyze features such as tone, pitch, and tempo of the voice data to precisely identify the user's emotions. It can also recognize patterns of emotional fluctuations by referring to past emotional state history and predict current emotions.

[0621] Based on these results, the server generates music data that corresponds to the user's current and predicted emotional state. Modern algorithms such as generative opposing networks and recurrent neural networks are used for music generation to create music tracks that match the user's preferences and emotions.

[0622] The generated music is streamed to the user's device, allowing the user to play the music through their mobile device or smart speaker for relaxation or a change of pace. For example, if a user is feeling stressed at work, the system will provide music with a relaxing tone to alleviate stress. In this way, the present invention realizes an emotion-responsive music experience tailored to individual users, contributing to the promotion of the user's psychological well-being.

[0623] The following describes the processing flow.

[0624] Step 1:

[0625] The device acquires the user's biometric information and voice data in real time from smartwatches and monitoring cameras. This data includes heart rate, body temperature, and voice tone.

[0626] Step 2:

[0627] The device encrypts the acquired data and sends it to the server using a secure communication protocol. This protects the user's privacy.

[0628] Step 3:

[0629] The server inputs the received data into the emotion engine. The emotion engine analyzes the tone, pitch, and tempo of the audio data and runs a machine learning model to identify the user's emotions.

[0630] Step 4:

[0631] The server updates the user profile using emotion identification results obtained from the emotion engine. This profile is used to store past emotional state information and understand the user's patterns over the long term.

[0632] Step 5:

[0633] The server uses the latest music generation algorithms to generate music data based on emotion analysis. By employing generative opposing networks and recurrent neural networks, it generates music tracks that match individual emotions.

[0634] Step 6:

[0635] The server streams the generated music data to the user's device. Users can then enjoy the music in real time via their mobile devices or smart speakers.

[0636] Step 7:

[0637] Users can relax and change their mood through music suited to their emotions. They can also provide feedback as needed to help improve the system.

[0638] (Example 2)

[0639] Next, we will describe Example 2. In the following description, the data processing device 12 will be referred to as the "server" and the robot 414 as the "terminal".

[0640] In modern society, there is a need to manage and appropriately respond to people's emotions and stress, but conventional systems have difficulty providing music that is appropriate to the user's physiological and psychological state. Therefore, users have to select music that suits their emotions themselves, which presents a challenge in terms of efficient emotional management.

[0641] The identification process performed by the identification processing unit 290 of the data processing device 12 in Example 2 is realized by the following means.

[0642] In this invention, the server includes data collection means for acquiring the user's biometric data, emotion analysis means for analyzing the user's emotional state based on the biometric data and voice data, emotion prediction means for predicting emotional fluctuation patterns considering the emotional state and past emotional history, and music generation means for generating music data based on the predicted emotional state. This allows the user to passively receive music appropriate to their state, making it easier to manage their emotions on a daily basis.

[0643] "Biometric data" refers to measurable information about a user's body, such as heart rate and body temperature, which indicate the user's physiological state.

[0644] "Data collection means" refers to devices and technologies that can acquire biometric data or voice data from users, such as devices that use sensors or microphones.

[0645] "Emotional analysis means" refers to algorithms and technologies that analyze acquired biometric and voice data to identify the user's emotional state.

[0646] "Emotion prediction methods" refer to technologies and models that predict a user's emotional fluctuation patterns based on current and past emotional data.

[0647] "Music generation means" refers to a technology for creating music data suitable for the user according to their analyzed emotional state, and is a function that generates music tracks using a generation AI model.

[0648] A "generative AI model" is an artificial intelligence model that has the ability to learn from data and generate new data, and specifically includes generative opposing networks and recurrent neural networks.

[0649] This invention is a system for providing a music experience tailored to the physiological and psychological characteristics of users. This system collects the user's biometric and vocal data in real time, analyzes the user's emotional state based on that data, and generates music data according to the analysis results.

[0650] The device uses the user's smartwatch or monitoring camera to acquire physiological data such as heart rate and body temperature, as well as voice data. This allows for the accurate and continuous collection of the user's biometric data. The device uses a secure communication protocol to safely transmit this data to the server.

[0651] The server receives the transmitted data and uses sentiment analysis tools to accurately identify the user's emotional state. This analysis utilizes machine learning algorithms and tools to analyze features such as tone, pitch, and tempo of the voice. Furthermore, sentiment prediction tools allow the server to refer to past emotional history and predict not only the current emotional state but also its fluctuation patterns.

[0652] The server then uses a generative AI model to generate music data tailored to the user's emotions. Specifically, it uses generative opposing networks (GANs) and recurrent neural networks (RNNs) to create music tracks that the user will like. This generated music is then delivered to the user's device via streaming. The user can play this music on their mobile device or smart speaker to help them relax or change their mood.

[0653] For example, when a user feels anxious in a large crowd, this system analyzes voice and biometric data to generate anxiety-relieving music. By listening to this music, the user can reduce stress and regain calmness.

[0654] An example of a prompt message is: "If the emotional state obtained from the user's biometric information is anxious, generate a music track to alleviate anxiety." This format allows you to instruct the generation AI model in this way.

[0655] In this way, the present invention provides a music experience tailored to the user's emotional state, realizing a system that contributes to maintaining and improving psychological health.

[0656] The flow of the specific processing in Example 2 will be explained using Figure 13.

[0657] Step 1:

[0658] The device acquires biometric information and voice data through the user's smartwatch or monitoring camera. Specifically, the smartwatch's sensors measure heart rate and body temperature, and voice data is recorded by the camera's microphone. This input data is sent to the server as physiological indicators and voice characteristics.

[0659] Step 2:

[0660] The server receives biometric information and voice data transmitted from the terminal. Input data includes heart rate, body temperature, voice tone, pitch, and tempo. The server preprocesses the data, removing noise to make it suitable for analysis. This processed data is then passed to the emotion analysis system.

[0661] Step 3:

[0662] The server analyzes pre-processed data using emotion analysis tools. The audio data is analyzed for changes in tone and rhythm, and combined with heart rate and body temperature to determine the user's emotional state. The user's emotional state is output as a result of this calculation.

[0663] Step 4:

[0664] The server uses emotion prediction tools in addition to emotion analysis tools to refer to the user's past emotion data and predict how their current emotional state will change. By utilizing past records, it calculates patterns of temporal changes in emotions and outputs predicted emotion values.

[0665] Step 5:

[0666] The server generates music data using a generative AI model based on the obtained emotional state and predicted values. Specifically, it uses generative opposing networks (GANs) and recurrent neural networks (RNNs) to create music tracks that are optimal for the user's emotions. Personalized music is output based on the input emotional information and the user's musical preferences.

[0667] Step 6:

[0668] The generated music data is streamed from the server to the user's device. The user plays the music on their smartphone or smart speaker and enjoys its relaxing effects. This process involves the user receiving the music and taking actual actions to maintain their psychological well-being and comfort.

[0669] (Application Example 2)

[0670] Next, we will explain application example 2. In the following explanation, the data processing device 12 will be referred to as the "server" and the robot 414 as the "terminal".

[0671] Conventional music delivery systems struggle to select music that matches a user's emotional state, making it difficult to provide a musical experience that responds to the instantaneous emotional changes of individual users. Furthermore, existing services lack sufficient mechanisms to analyze biometric information and voice data in real time and generate and provide appropriate music. Therefore, there is a need for technology that efficiently delivers music tailored to a user's emotional state and improves satisfaction.

[0672] The specific processing performed by the specific processing unit 290 of the data processing device 12 in Application Example 2 is realized by the following means.

[0673] In this invention, the server includes information acquisition means for acquiring the user's biometric information, emotion analysis means for analyzing the user's emotional state based on the biometric information and voice data, and music creation means for generating music data according to the emotional state. This makes it possible to automatically generate and provide music that is tailored to the emotional state of each individual user.

[0674] "Information acquisition means" refers to a device or mechanism that has the function of acquiring a user's biometric information or voice data.

[0675] "Emotional analysis means" refers to a device or algorithm that has the function of analyzing and identifying a user's emotional state based on acquired biometric information and voice data.

[0676] "Music creation means" refers to a device or algorithm that has the function of generating music data in accordance with the analyzed emotional state.

[0677] "Music transmission means" refers to a device or structure that has the function of transmitting and providing generated music data to a user.

[0678] To implement this invention, it is desirable to use a smartwatch or smartphone as a terminal for acquiring the user's biometric information in real time. These terminals function as "information acquisition means" for acquiring biometric information such as heart rate and body temperature, and acquire voice data using a microphone built into the terminal.

[0679] The acquired biometric information and voice data are transmitted to a server via the network. After receiving this data, the server uses an "emotion analysis tool" that implements a machine learning algorithm to analyze the user's emotional state. To do this, the server uses advanced machine learning software to analyze the tone, pitch, tempo, and other aspects of the voice data.

[0680] Based on the analysis results, the server uses a "music creation method" to generate music data that corresponds to the user's emotions. At this time, algorithms such as generative opposition learning models and recurrent neural networks are utilized to generate music that fits the user's emotions.

[0681] The generated music data is transmitted to the user's device in real time. The device has music playback capabilities, and the user can play the generated music through their smartphone or smart speaker to relax or uplift their emotions.

[0682] For example, if a user wants to relax, the system will automatically generate relaxing music and provide it to the user's device. Another example prompt is, "If the user's heart rate is elevated and a desire to be energized is detected, suggest a music track that generates uplifting music." The server processes this prompt, generating and providing music that matches the user's mood at that moment.

[0683] The flow of a specific process in Application Example 2 will be explained using Figure 14.

[0684] Step 1:

[0685] The terminal acquires the user's biometric information (heart rate, body temperature, etc.) via a device such as a smartwatch. This input data is measured in real time. Based on these results, the terminal packages the data and similarly acquires voice data, preparing to send it to the server.

[0686] Step 2:

[0687] The device transmits biometric information and voice data to the server. The input at this time is the biometric information acquired by the device and the recorded voice data. The server receives this data, appropriately decodes it, and formats it into the format necessary for analysis.

[0688] Step 3:

[0689] The server inputs the received biometric and voice data into a machine learning model for emotion analysis. In this step, the user's emotional state is identified by analyzing features such as voice tone, pitch, and tempo, and the analysis results are output. The results include the main emotions the user is feeling and their intensity.

[0690] Step 4:

[0691] The server generates music data using a generative AI model based on the emotion analysis results. This algorithm utilizes generative opposing networks and recurrent neural networks to create music that corresponds to the analyzed emotions. At this point, the input is the emotion analysis results, and the output is the generated music data.

[0692] Step 5:

[0693] The server streams the generated music data to the terminal in real time. The terminal receives this music and prepares it for playback by the user. The input is the generated music data, and the output is the music track ready for the user to listen to.

[0694] Step 6:

[0695] Users listen to music streamed through their devices. Based on user feedback and usage patterns, the server collects further data and continuously improves the system accordingly. This feedback mechanism enables specific actions to enhance the individual user experience.

[0696] The specific processing unit 290 transmits the result of the specific processing to the robot 414. In the robot 414, the control unit 46A causes the speaker 240 and the controlled object 443 to output the result of the specific processing. The microphone 238 acquires audio indicating user input for the result of the specific processing. The control unit 46A transmits the audio data indicating user input acquired by the microphone 238 to the data processing unit 12. In the data processing unit 12, the specific processing unit 290 acquires the audio data.

[0697] Data generation model 58 is a type of so-called generative AI (Artificial Intelligence). One example of data generation model 58 is ChatGPT (Internet search<URL: https: / / openai.com / blog / chatgpt> ), Gemini (Internet search) <url: https: gemini.google.com ?hl="ja">Examples of generative AI include the following. The data generation model 58 is obtained by performing deep learning on a neural network. The data generation model 58 is input with prompts containing instructions, and with inference data such as audio data representing speech, text data representing text, and image data representing images. The data generation model 58 infers from the input inference data according to the instructions indicated by the prompts, and outputs the inference results in data formats such as audio data and text data. Here, inference refers to, for example, analysis, classification, prediction, and / or summarization.

[0698] In the above embodiment, an example was given in which specific processing is performed by the data processing device 12, but the technology of this disclosure is not limited thereto, and the specific processing may also be performed by the robot 414.

[0699] Furthermore, the emotion identification model 59, acting as an emotion engine, may determine the user's emotion according to a specific mapping. Specifically, the emotion identification model 59 may determine the user's emotion according to a specific mapping, which is an emotion map (see Figure 9). Similarly, the emotion identification model 59 may also determine the robot's emotion, and the identification processing unit 290 may perform identification processing using the robot's emotion.

[0700] Figure 9 shows an emotion map 400 in which multiple emotions are mapped. In the emotion map 400, emotions are arranged in concentric circles radiating from the center. The closer to the center of the concentric circles, the more primitive the emotions are located. Further out of the concentric circles, emotions representing states and actions arising from mental states are located. Emotion is a concept that includes feelings and mental states. On the left side of the concentric circles, emotions that are generally generated from reactions occurring in the brain are located. On the right side of the concentric circles, emotions that are generally induced by situational judgment are located. Above and below the concentric circles, emotions that are generally generated from reactions occurring in the brain and induced by situational judgment are located. In addition, the emotion of "pleasure" is located on the upper side of the concentric circles, and the emotion of "displeasure" is located on the lower side. Thus, in the emotion map 400, multiple emotions are mapped based on the structure in which emotions arise, and emotions that are likely to occur simultaneously are mapped close together.

[0701] These emotions are distributed at the 3 o'clock position on the Emotion Map 400, and usually fluctuate between feelings of security and anxiety. In the right half of the Emotion Map 400, situational awareness takes precedence over internal feelings, resulting in a calm impression.

[0702] The inside of the Emotion Map 400 represents inner thoughts, while the outside represents actions. Therefore, the further you go from the outside of the Emotion Map 400, the more visible (expressed in actions) your emotions become.

[0703] Here, human emotions are based on various balances, such as posture and blood sugar levels. When these balances deviate from the ideal, it results in discomfort, and when they approach the ideal, it results in pleasure. Similarly, in robots, cars, motorcycles, etc., emotions can be created based on various balances, such as posture and battery level. When these balances deviate from the ideal, it results in discomfort, and when they approach the ideal, it results in pleasure. The emotion map can be generated, for example, based on Dr. Mitsuyoshi's emotion map (Research on a system for analyzing brain physiological signals of speech emotion recognition and emotion, Tokushima University, doctoral dissertation: https: / / ci.nii.ac.jp / naid / 500000375379). The left half of the emotion map contains emotions belonging to a region called "response," where sensation is dominant. The right half of the emotion map contains emotions belonging to a region called "situation," where situational awareness is dominant.

[0704] The emotion map defines two emotions that promote learning. One is the emotion around the middle of the negative "repentance" and "reflection" on the situation side. In other words, it is when the robot experiences negative emotions such as "I never want to feel this way again" or "I don't want to be scolded again." The other is the emotion around the positive "desire" on the reaction side. In other words, it is when the robot has positive feelings such as "I want more" or "I want to know more."

[0705] The emotion identification model 59 inputs user input into a pre-trained neural network, obtains emotion values ​​representing each emotion shown in the emotion map 400, and determines the user's emotion. This neural network is pre-trained based on multiple training data sets, which are combinations of user input and emotion values ​​representing each emotion shown in the emotion map 400. Furthermore, this neural network is trained so that emotions located close together have similar values, as shown in the emotion map 900 in Figure 10. Figure 10 shows an example where multiple emotions such as "reassured," "calm," and "confident" have similar emotion values.

[0706] The above description primarily focuses on the functions of the data processing device 12 in relation to this disclosure. However, the system related to this disclosure is not necessarily implemented on a server. The system related to this disclosure may be implemented as a general information processing system. This disclosure may be implemented, for example, as a software program that runs on a personal computer or as an application that runs on a smartphone. The method related to this disclosure may be provided to users in SaaS (Software as a Service) format.

[0707] In the above embodiment, an example was given in which a specific process is performed by a single computer 22. However, the technology of this disclosure is not limited thereto, and a distributed processing of the specific process may be performed by multiple computers, including computer 22. For example, a data generation model 58 may be provided in an external device of the data processing device 12, and the external device may generate data according to the input data.

[0708] In the above embodiment, an example was given in which the specific processing program 56 is stored in the storage 32, but the technology of this disclosure is not limited thereto. For example, the specific processing program 56 may be stored in a portable, computer-readable, non-temporary storage medium such as a USB (Universal Serial Bus) memory. The specific processing program 56 stored in the non-temporary storage medium is installed in the computer 22 of the data processing device 12. The processor 28 executes specific processing according to the specific processing program 56.

[0709] Alternatively, the specific processing program 56 may be stored in a storage device such as a server connected to the data processing device 12 via the network 54, and the specific processing program 56 may be downloaded and installed on the computer 22 in response to a request from the data processing device 12.

[0710] Furthermore, it is not necessary to store the entirety of the specific processing program 56 in a storage device such as a server connected to the data processing device 12 via the network 54, or to store the entirety of the specific processing program 56 in the storage 32; it is acceptable to store only a portion of the specific processing program 56.

[0711] The following types of processors can be used as hardware resources to perform specific processing. Examples of processors include a CPU, a general-purpose processor that functions as a hardware resource to perform specific processing by executing software, i.e., a program. Other examples of processors include dedicated electrical circuits, such as FPGAs (Field-Programmable Gate Arrays), PLDs (Programmable Logic Devices), or ASICs (Application Specific Integrated Circuits), which have circuit configurations specifically designed to perform specific processing. All of these processors have built-in or connected memory, and all of them perform specific processing by using memory.

[0712] The hardware resource that performs a specific process may consist of one of these various processors, or it may consist of a combination of two or more processors of the same or different types (for example, a combination of multiple FPGAs, or a combination of a CPU and an FPGA). Alternatively, the hardware resource that performs a specific process may consist of a single processor.

[0713] Examples of configurations using a single processor include, firstly, a configuration in which one or more CPUs and software are combined to form a single processor, and this processor functions as a hardware resource that performs a specific process. Secondly, there is a configuration using a processor that realizes the functions of the entire system, including multiple hardware resources that perform a specific process, on a single IC chip, as exemplified by SoCs (System-on-a-chip). In this way, a specific process is realized using one or more of the above types of processors as hardware resources.

[0714] Furthermore, the hardware structure of these various processors can more specifically utilize electrical circuits that combine circuit elements such as semiconductor devices. Also, the specific processing described above is merely an example. Therefore, it goes without saying that unnecessary steps can be deleted, new steps added, or the processing order rearranged, as long as it does not deviate from the main purpose.

[0715] The descriptions and illustrations presented above are detailed explanations of the technical aspects of this disclosure and are merely examples of the technical aspects. For example, the above descriptions of the structure, function, operation, and effect are examples of the structure, function, operation, and effect of the technical aspects of this disclosure. Therefore, it goes without saying that you may delete unnecessary parts, add new elements, or replace elements in the descriptions and illustrations presented above, as long as you do not deviate from the essence of the technical aspects of this disclosure. Furthermore, in order to avoid confusion and facilitate understanding of the technical aspects of this disclosure, explanations of common technical knowledge and the like that do not require special explanation to enable the implementation of the technical aspects of this disclosure have been omitted from the descriptions and illustrations presented above.

[0716] All documents, patent applications, and technical standards described herein are incorporated by reference to the same extent as if each individual document, patent application, and technical standard were specifically and individually noted to be incorporated by reference.

[0717] The following is further disclosed regarding the embodiments described above.

[0718] (Claim 1)

[0719] A means of collecting information to acquire the user's biometric information,

[0720] An emotion analysis means for analyzing the user's emotional state based on the aforementioned biometric information,

[0721] A music generation means that generates music data according to the aforementioned emotional state,

[0722] Music provision means for providing the generated music data to a user device,

[0723] A system that includes this.

[0724] (Claim 2)

[0725] The system according to claim 1, characterized in that the information gathering means acquires voice data and heart rate data.

[0726] (Claim 3)

[0727] The system according to claim 1, characterized in that the music generation means generates music data using a generative opposing network or a recurrent neural network.

[0728] "Example 1"

[0729] (Claim 1)

[0730] A data collection means for acquiring user biometric indicators and voice information,

[0731] An analysis means for analyzing the aforementioned biometric indicators and voice information to identify the user's emotional state,

[0732] A generation means that generates music that reflects the user's preferences and cultural background in relation to the aforementioned emotional state,

[0733] A means for distributing the generated music to the user's device and providing a real-time music experience,

[0734] A system that includes this.

[0735] (Claim 2)

[0736] The system according to claim 1, characterized in that the data acquisition means acquires voice information and physiological indicators such as heart rate.

[0737] (Claim 3)

[0738] The system according to claim 1, characterized in that the generation means uses a generation AI model to generate music via prompt sentences.

[0739] "Application Example 1"

[0740] (Claim 1)

[0741] A means of collecting information to acquire the user's biometric information,

[0742] An emotion analysis means for analyzing the user's emotional state based on the aforementioned biometric information,

[0743] A music generation means that generates music data according to the aforementioned emotional state,

[0744] A music provision means that outputs the generated music data in the working environment,

[0745] A system that includes this.

[0746] (Claim 2)

[0747] The system according to claim 1, characterized in that the information gathering means acquires voice data and heart rate data, thereby monitoring the worker's condition.

[0748] (Claim 3)

[0749] The system according to claim 1, characterized in that the music generation means generates music data using a generation model.

[0750] "Example 2 of combining an emotion engine"

[0751] (Claim 1)

[0752] A data collection method for acquiring user biometric data,

[0753] An emotion analysis means for analyzing the user's emotional state based on the aforementioned biometric data and voice data,

[0754] An emotion prediction means that predicts the pattern of emotional fluctuations considering the aforementioned emotional state and past emotional history,

[0755] Music generation means that generates music data based on the predicted emotional state,

[0756] Music provision means for providing the generated music data to a user device,

[0757] A system that includes this.

[0758] (Claim 2)

[0759] The system according to claim 1, characterized in that the data acquisition means acquires audio data including tone, pitch, and tempo characteristics, and heart rate data.

[0760] (Claim 3)

[0761] The system according to claim 1, characterized in that the music generation means generates music data using a generative opposing network or a recurrent neural network as a generative AI model.

[0762] "Application example 2 when combining with an emotional engine"

[0763] (Claim 1)

[0764] A means for acquiring information to obtain the user's biometric information,

[0765] An emotion analysis means for analyzing the user's emotional state based on the aforementioned biometric information and voice data,

[0766] A music creation means that generates music data according to the aforementioned emotional state,

[0767] Music transmission means for transmitting the generated music data to a user device,

[0768] A system that includes this.

[0769] (Claim 2)

[0770] The system according to claim 1, characterized in that the information acquisition means acquires voice data and heart rate data.

[0771] (Claim 3)

[0772] The system according to claim 1, characterized in that the music creation means generates music data using a generative opposition learning method or a recurrent neural network. [Explanation of symbols]

[0773] 10, 210, 310, 410 Data Processing Systems 12 Data Processing Devices 14 Smart Devices 214 Smart Glasses 314 Headset-type terminal 414 Robots< / url:> < / url:> < / url:> < / url:>

Claims

1. A means of collecting information to acquire the user's biometric information, An emotion analysis means for analyzing the user's emotional state based on the aforementioned biometric information, A music generation means that generates music data according to the aforementioned emotional state, Music provision means for providing the generated music data to a user device, A system that includes this.

2. The system according to claim 1, characterized in that the information gathering means acquires voice data and heart rate data.

3. The system according to claim 1, characterized in that the music generation means generates music data using a generative opposing network or a recurrent neural network.

Citation Information

Patent Citations

  • Persona chatbot control method and system

    JP2022180282A