system

A biometric-driven system provides personalized audio data based on heart rate and mental state evaluation, addressing the limitations of conventional music therapy by offering adaptable and effective stress management and concentration improvement.

JP2026069026APending Publication Date: 2026-04-23SOFTBANK GROUP CORP
View PDF 1 Cites 0 Cited by

Patent Information

Authority / Receiving Office
JP · JP
Patent Type
Applications
Current Assignee / Owner
SOFTBANK GROUP CORP
Filing Date
2024-10-11
Publication Date
2026-04-23

AI Technical Summary

Technical Problem

Existing methods for stress management and concentration improvement, such as music therapy, lack customization based on individual biological information and are not highly versatile.

Method used

A system that collects biometric information, such as heart rate, to evaluate a user's mental state and provides tailored audio data using generative AI, incorporating sports psychology and music therapy principles.

Benefits of technology

Enables personalized audio experiences that dynamically adapt to individual mental states, enhancing stress management and concentration.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure 2026069026000001_ABST
    Figure 2026069026000001_ABST
Patent Text Reader

Abstract

We provide the system. [Solution] A device that acquires the user's biometric information, A means of evaluating the user's mental state based on acquired biometric information, A means for generating or selecting acoustic data suitable for the user based on the evaluation results, A system including means for providing generated or selected acoustic data to a user.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The technology of the present disclosure relates to a system.

Background Art

[0002] Patent Document 1 discloses a method for controlling a persona chatbot, which is performed by at least one processor, including steps of receiving a user utterance, adding the user utterance to a prompt including an instruction sentence related to an explanation of a character of the chatbot, encoding the prompt, and inputting the encoded prompt into a language model to generate a chatbot utterance in response to the user utterance.

Prior Art Documents

Patent Documents

[0003]

Patent Document 1

Summary of the Invention

Problems to be Solved by the Invention

[0004] In modern society, there are various situations where stress management and concentration improvement are required. In particular, sports players, students, businesspersons, etc. need a method to effectively balance their physical and mental states. However, the means to receive mental support suitable for individual situations are limited, and it is difficult to provide a highly versatile solution. Also, existing music therapy and mental care methods have problems that it is difficult to customize based on individual biological information and the effects are limited.

Means for Solving the Problems

[0005] This invention uses a device to collect the user's biometric information, acquiring data such as heart rate in real time. Based on the acquired biometric information, it evaluates the user's mental state and provides the user with generated or selected audio data based on this evaluation. This makes it possible to provide music tailored to each user's individual condition by utilizing information from various mental care, sports psychology, and music therapy fields. Furthermore, this system is compatible with diverse biometric information sources and functions as a highly versatile mental care means.

[0006] "Biometric information" refers to data that indicates the user's physical condition, and includes heart rate, blood pressure, pulse rate, etc.

[0007] "Mental state" refers to the user's psychological and emotional state, and may include situations such as anxiety, tension, and relaxation.

[0008] "Audio data" refers to digital data that contains information about physical sound, and can include music, sound effects, and rhythmic patterns.

[0009] "Generation" refers to the process of creating new data or information based on specific conditions or algorithms.

[0010] "Selection" is the act of choosing appropriate data or information from existing data based on specific criteria. [Brief explanation of the drawing]

[0011] [Figure 1] This is a conceptual diagram showing an example of the configuration of a data processing system according to the first embodiment. [Figure 2] This is a conceptual diagram showing an example of the essential functions of a data processing device and a smart device according to the first embodiment. [Figure 3] This is a conceptual diagram showing an example of the configuration of a data processing system according to the second embodiment. [Figure 4] This is a conceptual diagram showing an example of the main functions of a data processing device and smart glasses according to the second embodiment. [Figure 5] This is a conceptual diagram showing an example of the configuration of a data processing system according to the third embodiment. [Figure 6] This is a conceptual diagram showing an example of the main functions of a data processing device and a headset-type terminal according to the third embodiment. [Figure 7] This is a conceptual diagram showing an example of the configuration of a data processing system according to the fourth embodiment. [Figure 8] This is a conceptual diagram showing an example of the main functions of a data processing device and a robot according to the fourth embodiment. [Figure 9] This shows an emotion map where multiple emotions are mapped. [Figure 10] This shows an emotion map where multiple emotions are mapped. [Figure 11] This is a sequence diagram showing the processing flow of the data processing system in Example 1. [Figure 12] This is a sequence diagram showing the processing flow of the data processing system in Application Example 1. [Figure 13] This is a sequence diagram showing the processing flow of the data processing system in Example 2, which incorporates an emotion engine. [Figure 14] This is a sequence diagram showing the processing flow of the data processing system in Application Example 2, which combines an emotion engine. [Modes for carrying out the invention]

[0012] Hereinafter, an example of an embodiment of the system relating to the technology of this disclosure will be described with reference to the attached drawings.

[0013] First, let's explain the terminology used in the following explanation.

[0014] In the following embodiments, the labeled processor (hereinafter simply referred to as "processor") may be a single arithmetic unit or a combination of multiple arithmetic units. Also, the processor may be a single type of arithmetic unit or a combination of multiple types of arithmetic units. Examples of arithmetic units include a CPU (Central Processing Unit), a GPU (Graphics Processing Unit), a GPGPU (General-Purpose computing on Graphics Processing Units), an APU (Accelerated Processing Unit), and the like.

[0015] In the following embodiments, the labeled RAM (Random Access Memory) is a memory in which information is temporarily stored and is used as a work memory by the processor.

[0016] In the following embodiments, the labeled storage is one or more non-volatile storage devices that store various programs, various parameters, and the like. Examples of non-volatile storage devices include flash memory (SSD (Solid State Drive)), magnetic disks (e.g., hard disks), or magnetic tapes, and the like.

[0017] In the following embodiments, the labeled communication I / F (Interface) is an interface including a communication processor, an antenna, and the like. The communication I / F controls communication between multiple computers. Examples of communication standards applied to the communication I / F include wireless communication standards including 5G (5th Generation Mobile Communication System), Wi-Fi (registered trademark), or Bluetooth (registered trademark).

[0018] In the following embodiments, "A and / or B" is synonymous with "at least one of A and B." That is, "A and / or B" means that it may be A alone, or B alone, or a combination of A and B. Furthermore, in this specification, the same concept as "A and / or B" applies when expressing three or more things linked by "and / or."

[0019] [First Embodiment]

[0020] Figure 1 shows an example of the configuration of the data processing system 10 according to the first embodiment.

[0021] As shown in Figure 1, the data processing system 10 includes a data processing device 12 and a smart device 14. An example of the data processing device 12 is a server.

[0022] The data processing device 12 comprises a computer 22, a database 24, and a communication interface 26. The computer 22 is an example of a "computer" related to the technology of this disclosure. The computer 22 comprises a processor 28, RAM 30, and storage 32. The processor 28, RAM 30, and storage 32 are connected to a bus 34. The database 24 and the communication interface 26 are also connected to the bus 34. The communication interface 26 is connected to a network 54. An example of the network 54 is a WAN (Wide Area Network) and / or a LAN (Local Area Network).

[0023] The smart device 14 comprises a computer 36, a reception device 38, an output device 40, a camera 42, and a communication interface 44. The computer 36 comprises a processor 46, RAM 48, and storage 50. The processor 46, RAM 48, and storage 50 are connected to a bus 52. The reception device 38, output device 40, and camera 42 are also connected to the bus 52.

[0024] The reception device 38 is equipped with a touch panel 38A and a microphone 38B, etc., and receives user input. The touch panel 38A receives user input by detecting contact with an object (e.g., a pen or finger). The microphone 38B receives user input by detecting the user's voice. The control unit 46A transmits data indicating the user input received by the touch panel 38A and microphone 38B to the data processing device 12. In the data processing device 12, the specific processing unit 290 acquires the data indicating the user input.

[0025] The output device 40 includes a display 40A and a speaker 40B, and presents data to the user 20 by outputting the data in a form perceptible to the user 20 (e.g., audio and / or text). The display 40A displays visible information such as text and images according to instructions from the processor 46. The speaker 40B outputs audio according to instructions from the processor 46. The camera 42 is a small digital camera equipped with an optical system such as a lens, aperture, and shutter, and an image sensor such as a CMOS (Complementary Metal-Oxide-Semiconductor) image sensor or a CCD (Charge Coupled Device) image sensor.

[0026] Communication interface 44 is connected to network 54. Communication interfaces 44 and 26 are responsible for the exchange of various types of information between processor 46 and processor 28 via network 54.

[0027] Figure 2 shows an example of the main functions of the data processing device 12 and the smart device 14.

[0028] As shown in Figure 2, in the data processing device 12, a specific processing is performed by the processor 28. A specific processing program 56 is stored in the storage 32. The specific processing program 56 is an example of a "program" related to the technology of this disclosure. The processor 28 reads the specific processing program 56 from the storage 32 and executes the read specific processing program 56 on the RAM 30. The specific processing is realized by the processor 28 operating as a specific processing unit 290 according to the specific processing program 56 executed on the RAM 30.

[0029] The storage 32 stores the data generation model 58 and the emotion identification model 59. The data generation model 58 and the emotion identification model 59 are used by the identification processing unit 290.

[0030] In the smart device 14, the processor 46 performs the reception output processing. The storage 50 stores the reception output program 60. The reception output program 60 is used in conjunction with a specific processing program 56 by the data processing system 10. The processor 46 reads the reception output program 60 from the storage 50 and executes the read reception output program 60 on the RAM 48. The reception output processing is realized by the processor 46 operating as a control unit 46A according to the reception output program 60 executed on the RAM 48.

[0031] Next, the specific processing performed by the specific processing unit 290 of the data processing device 12 will be described. In the following description, the data processing device 12 will be referred to as the "server" and the smart device 14 as the "terminal".

[0032] This invention is a system that uses a smart device to acquire a user's biometric information and, based on that information, provides acoustic data optimized for the user's mental state. The details are described below.

[0033] The server receives biometric information transmitted from the terminal in real time. This information primarily includes heart rate and is used as data indicating the user's current physical and mental state. The received data is analyzed to assess the user's situation, whether relaxation is needed, or whether concentration is required.

[0034] The terminal typically functions as a smartphone or tablet and is responsible for acquiring biometric information from portable devices such as smartwatches. The terminal transmits the acquired data to a server via the internet. Upon receiving processed results or acoustic data from the server, the terminal becomes an interface for providing that data to the user.

[0035] This system allows users to receive dynamically generated and selected audio data based on their heart rate and mental state. This audio data comes in various formats, including relaxing music and rhythmic patterns designed to enhance concentration. Users can operate the application to listen to the corresponding audio data in real time.

[0036] For example, if concentration is needed during class or work, the device will determine that the heart rate is constant or indicates a state of tension. If the analysis indicates the user needs improved concentration, the server selects music to enhance concentration and sends it to the device. The audio data is played through the device to support the user's concentration. This allows the user to receive support tailored to their mental state.

[0037] The following describes the processing flow.

[0038] Step 1:

[0039] The user launches the smartphone app and connects it to the smartwatch. The app prepares to periodically collect heart rate data from the smartwatch via Bluetooth or other means.

[0040] Step 2:

[0041] The device acquires biometric information, such as heart rate, from the smartwatch. This data is updated at regular intervals, reflecting current information in real time.

[0042] Step 3:

[0043] The device organizes the acquired biometric information into data packets and sends them to the server via the internet. The communication is encrypted to ensure secure transmission.

[0044] Step 4:

[0045] The server analyzes biometric information received from the terminal. Based on the heart rate data, the server evaluates the user's current state and determines what kind of acoustic data is appropriate.

[0046] Step 5:

[0047] Based on the analysis results, the server selects or generates appropriate acoustic data. Specifically, it uses AI to choose music that is most suitable for the user's condition, referencing information on mental care and music therapy.

[0048] Step 6:

[0049] The server compresses the selected or generated audio data and sends it to the terminal in streaming format.

[0050] Step 7:

[0051] The device decompresses the audio data received from the server and begins playback in real time. Users can listen to the played audio data via a smartphone app.

[0052] Step 8:

[0053] Users can provide feedback on audio data through the app. This feedback is sent to the server and used to improve future data selection and generation algorithms.

[0054] (Example 1)

[0055] Next, we will describe Example 1. In the following description, the data processing device 12 will be referred to as the "server," and the smart device 14 will be referred to as the "terminal."

[0056] In modern society, individual users routinely experience psychological problems such as stress and lack of concentration. Under these circumstances, there are limited means of dynamically providing acoustic information tailored to each user's psychological state. Furthermore, conventional solutions lack sufficient technology to individually incorporate user feedback. Therefore, there is a need for a more adaptable and reliable system that can deliver an optimized acoustic experience for each individual.

[0057] The identification process performed by the identification processing unit 290 of the data processing device 12 in Example 1 is realized by the following means.

[0058] In this invention, the server includes means for acquiring the user's biometric parameters, means for evaluating the user's psychological state based on the acquired biometric parameters, and means for generating or selecting optimal acoustic information for the user based on the evaluation results. This makes it possible to provide appropriate acoustic information that corresponds to the user's unique psychological state.

[0059] "Biometric parameters" refer to information that indicates the user's physical condition, and include data such as heart rate, skin temperature, and blood pressure.

[0060] "Psychological state" refers to the emotional or mental condition of the user, and represents mental conditions such as stress levels and concentration levels.

[0061] "Acoustic information" refers to sound wave data provided to users, specifically a collection of sounds that have a psychological effect, such as music, nature sounds, and white noise.

[0062] A "generative AI model" is a form of artificial intelligence that uses learning algorithms to analyze data and generate new data.

[0063] "Feedback" refers to the evaluations and opinions that users provide regarding the acoustic information they receive, and is used as response information to improve and optimize the system.

[0064] This invention is a system for providing individually optimized acoustic information based on the user's biological parameters. The details of this system are described below.

[0065] Server operation

[0066] The server centrally receives biometric parameters from users transmitted from multiple terminals. These biometric parameters include heart rate, skin temperature, and blood pressure. The server analyzes this data in real time to estimate the user's psychological state. Specifically, it uses a generative AI model to analyze fluctuations in biometric parameters and evaluate stress levels and concentration levels. Based on this analysis, the server selects appropriate acoustic information from a digital library for the user and sends it to the terminal. The software used includes data mining tools and machine learning algorithms.

[0067] Terminal operation

[0068] The device typically functions as a smartphone or tablet, receiving biometric parameters acquired by smartwatches and other wearable devices via Bluetooth or other wireless communication methods. The device has the capability to transmit this data to a server over the internet. It also functions as an interface providing the user with analysis results and acoustic information received from the server. The application on the device features a user interface that allows the user to easily play acoustic information and perform operations such as volume adjustment and playback / pause.

[0069] User experience

[0070] Users can receive individually generated and selected acoustic information based on their heart rate and mental state via their smart devices. This includes relaxing music and rhythmic patterns that improve concentration. Users can experience the acoustics in real time using the provided application and send feedback to the server through the application. This feedback will be used to improve future acoustic information provision using a generative AI model.

[0071] Specific example

[0072] For example, when a user is doing work that requires concentration, the device can determine their current level of tension from heart rate data. The server analyzes this data, selects music that will help the user concentrate, and sends it to the device. The audio information is played through the device to support the user's concentration.

[0073] Example of a prompt

[0074] "What music would you recommend for when you want to relax?"

[0075] "I want to concentrate on my work, what kind of music would be good?"

[0076] This system allows users to dynamically receive sound support tailored to their different psychological states, thereby improving their quality of daily life.

[0077] The flow of the specific processing in Example 1 will be explained using Figure 11.

[0078] Step 1:

[0079] The terminal acquires biometric parameters from the wearable device worn by the user. Specifically, it uses Bluetooth communication to receive data such as heart rate and skin temperature. The input is biometric data from the wearable device, and the output is formatted data ready to be sent to the server.

[0080] Step 2:

[0081] The device transmits acquired biometric parameters to the server. Specifically, it uploads data to the server in real time using a secure communication protocol (e.g., HTTPS). The input is sensor data, and the output is the data sent to the server.

[0082] Step 3:

[0083] The server analyzes the received biometric parameters. Specifically, it uses a generative AI model to analyze heart rate fluctuations, thereby evaluating the user's psychological state. It generates analysis results including stress levels and concentration levels. The input is biometric data sent from the terminal, and the output is the analyzed psychological state evaluation result.

[0084] Step 4:

[0085] The server selects the most suitable acoustic information for the user based on the analysis results. It queries a digital library to identify music and rhythms that match the user's psychological state. The input is the evaluation results, and the output is the selected acoustic information.

[0086] Step 5:

[0087] The server transmits the selected acoustic information to the terminal. The acoustic information is converted into a data format and sent to the terminal via the internet. The input is the selected acoustic information, and the output is the data transmitted to the terminal.

[0088] Step 6:

[0089] The terminal provides the user with audio information received from the server. Specifically, it decodes the audio data and plays it back through speakers or headphones. The input is the audio information from the server, and the output is the played sound.

[0090] Step 7:

[0091] Users send feedback on acoustic information via their devices. They input subjective evaluations through an interface within the app and send them to the server. The input is feedback data, and the output is the evaluation information sent to the server.

[0092] Step 8:

[0093] The server updates the generated AI model based on user feedback. By analyzing the feedback and adjusting the model's algorithm, it improves the accuracy of future acoustic information provision. The input is the feedback information, and the output is the updated AI model.

[0094] (Application Example 1)

[0095] Next, we will explain Application Example 1. In the following explanation, the data processing device 12 will be referred to as the "server," and the smart device 14 will be referred to as the "terminal."

[0096] Traditionally, methods for improving the customer experience within stores have been limited, making it difficult to provide an optimal environment tailored to each customer's individual mental state. In particular, dynamically adjusting the sound environment according to the customer's mental state to enhance purchasing intent or create a relaxed atmosphere has been challenging to achieve.

[0097] The specific processing performed by the specific processing unit 290 of the data processing device 12 in Application Example 1 is realized by the following means.

[0098] In this invention, the server includes means for acquiring the user's biometric data, means for determining the user's psychological state based on the acquired biometric data, and means for generating or selecting acoustic information suitable for the user based on the determination result. This makes it possible to provide an optimal acoustic environment in a store that corresponds to the customer's mental state, thereby improving the customer experience.

[0099] "Users" refer to individual customers or consumers who provide biometric data and receive acoustic information tailored to their mental state.

[0100] "Biometric data" refers to information that indicates a user's physical condition, such as their heart rate, and is used to assess their psychological state.

[0101] "Psychological state" refers to the user's current mental health and emotional state, including states such as relaxation and concentration.

[0102] "Acoustic information" refers to sound data, such as music and sound effects, that are selected or generated based on a person's mental state.

[0103] "Means for generating or selecting" refers to processes or devices for creating or selecting appropriate acoustic information based on the results of a psychological state assessment.

[0104] "Means for delivering acoustic information to a store environment" refers to a system for providing selected acoustic information to customers within a physical or virtual store space.

[0105] This invention is a system that improves the customer experience by utilizing the user's biometric data to provide optimal acoustic information in the store environment.

[0106] The server acquires biometric data, such as heart rate, from the user's smartwatch or equivalent portable device. This data is transmitted to the server via the internet, where a machine learning model written in Python is used to determine the user's psychological state in real time. This machine learning model is capable of distinguishing between states of relaxation and concentration.

[0107] Based on the determined psychological state, the server selects or generates the most suitable audio information for the user. This audio information may be selected using the Spotify API or similar music streaming platforms. The selected audio information is then distributed via Wi-Fi or Bluetooth to speaker systems installed within the store. This ensures that soothing music plays when the user wants to relax, and music with an appropriate tempo plays when they need to concentrate.

[0108] As a concrete example, if a customer visiting a shopping mall is determined to be stressed due to an elevated heart rate, the server will select healing music and distribute it throughout the store to promote relaxation. An example of a prompt message in this case would be: "Heart rate data: 85, Mental state: Needs relaxation, Music type: Pleasant melody."

[0109] The flow of a specific process in Application Example 1 will be explained using Figure 12.

[0110] Step 1:

[0111] The device acquires heart rate data from the user's smartwatch. This data is obtained directly from the device's built-in sensors and forms the basis for processing as the user's biometric information. This data is then sent to a server for subsequent analysis processes.

[0112] Step 2:

[0113] The server receives heart rate data transmitted from the terminal. The server uses this input data to perform data analysis using a machine learning model built in Python. Specifically, the model processes the heart rate data as a feedback loop to determine the user's psychological state (e.g., stressed or relaxed).

[0114] Step 3:

[0115] The server selects appropriate audio information based on the results of a machine learning model. It references the psychological state data received as input, searches for suitable music from its music library using the Spotify API, and selects the appropriate track. The selected audio information is then sent to the subsequent distribution process.

[0116] Step 4:

[0117] The server distributes selected audio information to the store's speaker system via Wi-Fi or Bluetooth. This causes music tailored to the customer's mental state to begin playing in the store environment. Because the audio information is played in real time, the customer's experience is in line with their biological state.

[0118] Step 5:

[0119] Users can enjoy shopping in a comfortable environment tailored to their mental state by listening to the music playing in the store. An example prompt message, "Heart rate data: 85, Mental state: Need to relax, Music selection type: Pleasant melody," serves as the trigger to initiate this process.

[0120] Furthermore, an emotion engine that estimates the user's emotions may be incorporated. That is, the identification processing unit 290 may use the emotion identification model 59 to estimate the user's emotions and perform identification processing using the user's emotions.

[0121] This invention is a system that combines the user's biometric information with an emotion engine to more accurately evaluate the user's mental state and provide appropriate acoustic data. In this embodiment, a smart device and a communication terminal work together to analyze the user's emotions in real time using data such as the user's heart rate, voice, and facial expressions.

[0122] The server receives biometric and emotional data transmitted from the terminal and integrates this information to evaluate the user's current mental state. The emotion engine uses machine learning models to analyze the user's voice tone and facial movements to recognize emotions such as joy, sadness, and surprise. This enables a comprehensive evaluation of mental state that includes not only biometric information but also emotional elements.

[0123] The device provides a user-friendly interface. Specifically, it uses a smartphone or PC to send the user's voice and facial expression data to an emotion engine, which then sends the results to a server. Furthermore, the device receives appropriate audio data from the server and plays it back to the user in real time.

[0124] Users can play music through the application and provide feedback on the music's impact on their emotions. This feedback information is used to improve the quality of future music generation and selection processes. For example, if the server determines that a user is experiencing stress, it will select relaxing music based on the emotion engine's recognition results and provide it to the user via the device.

[0125] This system goes beyond conventional music delivery based on biometric information, making it possible to provide highly accurate care for a variety of mental states stemming from emotions.

[0126] The following describes the processing flow.

[0127] Step 1:

[0128] The user launches the smartphone app and connects it to the smartwatch and smart device. The app prepares to collect data such as the user's heart rate, voice, and facial expressions.

[0129] Step 2:

[0130] The device acquires heart rate data in real time from the smartwatch, and also collects the user's voice and facial expression data via a smartphone or webcam. This allows for the accumulation of comprehensive biometric information.

[0131] Step 3:

[0132] The device combines collected heart rate data with voice and facial expression data to generate data packets, which are then transmitted to a server via the internet. Communication is conducted using a secure method.

[0133] Step 4:

[0134] The server analyzes the received biometric and emotional data. It evaluates changes in heart rate, voice tone, and facial expressions to identify the user's current mental state.

[0135] Step 5:

[0136] The emotion engine operates on the server and uses machine learning models to recognize the user's emotions from their voice and facial expressions. Based on the recognition results, it determines whether the user is happy, sad, or otherwise.

[0137] Step 6:

[0138] The server generates or selects appropriate acoustic data based on the analysis results and emotion recognition results. The acoustic data is prepared to be effective for relaxation and energy enhancement.

[0139] Step 7:

[0140] The server sends the generated or selected audio data to the terminal. The terminal receives this data and plays music for the user.

[0141] Step 8:

[0142] Users can provide feedback on the effects of the music while listening to it. This feedback is sent from the device to the server and used to improve the performance of the emotion engine and the accuracy of the sound data.

[0143] (Example 2)

[0144] Next, we will describe Example 2. In the following description, the data processing device 12 will be referred to as the "server" and the smart device 14 as the "terminal".

[0145] A comprehensive assessment of a user's mental state is needed, taking into account not only their biometric information but also emotional information such as voice and facial expressions. However, conventional systems have the challenge of being unable to integrate and analyze this diverse data and provide accurate acoustic data in real time.

[0146] The identification process performed by the identification processing unit 290 of the data processing device 12 in Example 2 is realized by the following means.

[0147] In this invention, the server includes means for integrating and evaluating the user's biometric information, voice, and facial expression data; means for selecting appropriate acoustic data based on the evaluated mental state; and means for providing the selected acoustic data in real time. This enables a more accurate and comprehensive evaluation of the mental state and the provision of detailed acoustic data that corresponds to the user's emotional state.

[0148] "User" refers to an individual who uses this system to perform biometric and emotional analysis.

[0149] "Biometric information" refers to data that indicates an individual's physical or physiological state, such as heart rate, voice, and facial expressions.

[0150] "Mental state" refers to the user's psychological and emotional state, including emotions such as joy, sadness, and surprise.

[0151] "Audio data" refers to digital data related to the user's voice, including their tone and tempo.

[0152] "Facial expression data" refers to digital data related to facial movements and changes, and is used to infer the user's emotions.

[0153] "Audio data" refers to digital data related to sound, such as music and sound effects, intended to influence the user's emotions and mental state.

[0154] "Integration" refers to the process of combining and analyzing biometric information and emotional data to evaluate the overall mental state.

[0155] "Real-time" refers to a situation where latency is minimized, allowing users to receive data services instantly.

[0156] This invention is a system that integrates the user's biometric information and emotional data to more accurately analyze the user's mental state and provide acoustic data based on that analysis. This system operates through the coordinated efforts of three entities: a server, a terminal, and a user.

[0157] The server receives biometric information (heart rate, voice, facial expression data, etc.) transmitted from the terminal and analyzes this data using an emotion engine. The emotion engine utilizes machine learning models to recognize emotions based on voice tone and facial movements. This allows for a comprehensive evaluation of emotional and physiological data, enabling an assessment of the user's current mental state.

[0158] On the other hand, the terminal provides an interface with the user. Using a smart device or personal computer, it sends the user's voice and facial expression data to the emotion engine and then sends the results to the server. It also has the function of receiving appropriate audio data sent from the server and playing it back to the user in real time.

[0159] Users provide biometric data through their smart devices and listen to the generated acoustic data. Furthermore, they can provide feedback on how the acoustic data affects their emotions. This feedback helps improve the accuracy of the next acoustic data generation and selection process. For example, if a user is feeling stressed, the system can select and play relaxing music in real time.

[0160] An example of a prompt might be, "Generate a music selection algorithm to provide relaxing music to users experiencing stress." In this way, the system can assess the user's mental state in real time and provide an appropriate sound experience based on that assessment.

[0161] The flow of the specific processing in Example 2 will be explained using Figure 13.

[0162] Step 1:

[0163] The user acquires biometric information such as heart rate, voice, and facial expressions using a smart device. The user uses a smartphone or wearable device to measure biometric data in real time through sensors on the device. The input to this process is the user's heart rate, voice, and facial expression data, and the output is transmitted to the terminal as biometric information.

[0164] Step 2:

[0165] The device transmits the acquired biometric information to the server. Specifically, the device transfers the information to the server as data packets via an internet connection. The input for this step is the biometric information collected from the user, and the output is the digital data sent to the server.

[0166] Step 3:

[0167] The server analyzes the received biometric information, along with the user's voice and facial expression data, using an emotion engine. On the server, a machine learning model processes the input data, analyzing voice tone and facial expressions to recognize the emotional state. In this analysis process, the input is digital data received from the terminal, and the output is the analyzed emotional information.

[0168] Step 4:

[0169] The server evaluates the user's current mental state based on the analyzed emotional information. This process integrates biometric and emotional information to comprehensively assess the mental state using a computer algorithm. The input in this step is the analyzed emotional information, and the output is the evaluated mental state data.

[0170] Step 5:

[0171] The server selects appropriate sound data based on the assessed mental state. For example, if the user is assessed as feeling stressed, it selects music with a relaxing effect. The input is mental state data, and the output is the selected sound data.

[0172] Step 6:

[0173] The selected audio data is sent from the server to the terminal. The server encodes the data and sends it in a format that the terminal can receive. The input for this step is the audio data, and the output is the audio file sent to the terminal.

[0174] Step 7:

[0175] The terminal plays the received audio data in real time and provides it to the user. The terminal plays music using an audio player, which the user can listen to. The input to this process is the received audio file, and the output is the audio playback as a user experience.

[0176] Step 8:

[0177] Users provide feedback on the impact of the provided audio data on their emotions. This feedback is sent to the server via the application and used to improve the audio data selection process. The input for this step is the user's experience with the audio data, and the output is the feedback information sent through the application.

[0178] (Application Example 2)

[0179] Next, we will explain application example 2. In the following explanation, the data processing device 12 will be referred to as a "server" and the smart device 14 as a "terminal".

[0180] Traditional methods have struggled to accurately provide acoustic information based on the user's mental state. In particular, evaluations based solely on simple biometric data may fail to adequately reflect the user's emotions through music. As a result, there are limitations to enhancing the user's emotional satisfaction.

[0181] The specific processing performed by the specific processing unit 290 of the data processing device 12 in Application Example 2 is realized by the following means.

[0182] In this invention, the server includes means for acquiring the user's biometric data, means for evaluating the user's mental state based on the acquired biometric data, voice tone, and facial movements, means for using a communication device to generate or select acoustic information suitable for the user based on the evaluation results and provide it to the user, means for analyzing voice tone and facial movements using an emotion analysis engine, and means for selecting acoustic information from a database based on the analysis results and providing it to the user in real time. This makes it possible to accurately evaluate the user's emotions and provide acoustic information suitable for those emotions.

[0183] "Biometric data" refers to individual physical and physiological information such as a user's heart rate pattern, voice tone, and facial movements.

[0184] "Mental state" refers to the emotional and psychological state or tendencies of the user, and includes emotions such as joy, sadness, and surprise.

[0185] "Acoustic information" refers to data of sounds and music that are suitable for the user's mental state.

[0186] "Communication equipment" refers to the hardware and software infrastructure used to send and receive data between servers and terminals and to provide acoustic information to users in real time.

[0187] An "emotion analysis engine" refers to a software component that uses machine learning models to analyze the tone of a user's voice and facial movements to recognize their emotions.

[0188] A "database" is a system for storing and managing audio information provided to users.

[0189] The system for realizing this invention uses a terminal equipped with various sensors and a server with advanced analytical capabilities. Specifically, a terminal such as a smartphone collects the user's biometric data and transmits that data to the server via the internet. The user's biometric data includes heart rate patterns, voice tone, and facial movements.

[0190] The server runs an emotion analysis engine that utilizes machine learning models based on biometric data. This emotion analysis engine uses libraries such as TENSORFLOW® for processing audio and images. Based on the analysis results, the user's mental state is evaluated, and the most appropriate acoustic information is selected from the database. In this process, a cloud database such as Firebase is used to manage and provide the acoustic data.

[0191] If the system determines that a user needs relaxation, it might select a piece of music, such as a classical piano sonata, and stream it in real time. This is expected to promote mental relaxation in the user.

[0192] As an example of a prompt to a generative AI model, the question, "What genre of music is suitable when a user needs to relax?" is used to optimize the system's suggestions. In this way, the user is provided with the most personalized music experience possible.

[0193] The flow of a specific process in Application Example 2 will be explained using Figure 14.

[0194] Step 1:

[0195] The device collects biometric data such as the user's heart rate pattern, voice tone, and facial movements using various sensors. This data is temporarily stored within the device and then transmitted to a server via the internet. The input is the user's biometric data, and the output is the generation of data packets to be sent to the server.

[0196] Step 2:

[0197] The server receives biometric data from the terminal and processes it as input data for running a machine learning model. Specifically, it uses TensorFlow to analyze voice tone and facial movements to evaluate the user's mental state. This process outputs mental states such as "relaxed," "lively," and "sad."

[0198] Step 3:

[0199] The server selects the most suitable audio information from a Firebase database based on the mental state assessment results. The selection process involves rule-based data processing, such as matching the assessment results and selecting "classical music" for a relaxed state. The input is the mental state assessment results, and the output is the selected audio information.

[0200] Step 4:

[0201] The server transmits the selected audio information to the terminal. The audio information is converted into real-time streaming data and prepared to be provided to the user. The input is the selected audio information, and the output is streaming data that can be played on the user's terminal.

[0202] Step 5:

[0203] Users enjoy audio information provided through their devices. Actual playback takes place through the smartphone's audio function, allowing users to obtain a musical experience suited to their own mental state. The input is streaming data, and the output is the musical experience mediated through the user's hearing.

[0204] The specific processing unit 290 transmits the result of the specific processing to the smart device 14. In the smart device 14, the control unit 46A causes the output device 40 to output the result of the specific processing. The microphone 38B acquires audio indicating user input for the result of the specific processing. The control unit 46A transmits the audio data indicating user input acquired by the microphone 38B to the data processing device 12. In the data processing device 12, the specific processing unit 290 acquires the audio data.

[0205] Data generation model 58 is a so-called generative AI (Artificial Intelligence). An example of data generation model 58 is ChatGPT (registered trademark) (Internet search).<URL: https: / / openai.com / blog / chatgpt> ), Gemini (registered trademark) (Internet search) <url: https: gemini.google.com ?hl="ja">Examples of generative AI include the following. The data generation model 58 is obtained by performing deep learning on a neural network. The data generation model 58 is input with prompts containing instructions, and with inference data such as audio data representing speech, text data representing text, and image data representing images. The data generation model 58 infers from the input inference data according to the instructions indicated by the prompts, and outputs the inference results in data formats such as audio data and text data. Here, inference refers to, for example, analysis, classification, prediction, and / or summarization.

[0206] In the above embodiment, an example was given in which specific processing is performed by the data processing device 12, but the technology of this disclosure is not limited thereto, and the specific processing may also be performed by the smart device 14.

[0207] [Second Embodiment]

[0208] Figure 3 shows an example of the configuration of the data processing system 210 according to the second embodiment.

[0209] As shown in Figure 3, the data processing system 210 includes a data processing device 12 and smart glasses 214. An example of the data processing device 12 is a server.

[0210] The data processing device 12 comprises a computer 22, a database 24, and a communication interface 26. The computer 22 is an example of a "computer" related to the technology of this disclosure. The computer 22 comprises a processor 28, RAM 30, and storage 32. The processor 28, RAM 30, and storage 32 are connected to a bus 34. The database 24 and the communication interface 26 are also connected to the bus 34. The communication interface 26 is connected to a network 54. An example of the network 54 is a WAN (Wide Area Network) and / or a LAN (Local Area Network).

[0211] The smart glasses 214 include a computer 36, a microphone 238, a speaker 240, a camera 42, and a communication interface 44. The computer 36 includes a processor 46, RAM 48, and storage 50. The processor 46, RAM 48, and storage 50 are connected to a bus 52. The microphone 238, speaker 240, and camera 42 are also connected to the bus 52.

[0212] The microphone 238 receives voice signals from the user 20 and receives instructions from the user 20. The microphone 238 captures the voice signals from the user 20, converts the captured voice into audio data, and outputs it to the processor 46. The speaker 240 outputs audio according to the instructions from the processor 46.

[0213] Camera 42 is a small digital camera equipped with an optical system including a lens, aperture, and shutter, and an image sensor such as a CMOS (Complementary Metal-Oxide-Semiconductor) image sensor or a CCD (Charge Coupled Device) image sensor, and captures images of the area around the user 20 (for example, an imaging range defined by a field of view equivalent to the width of a typical healthy person's field of vision).

[0214] Communication interface 44 is connected to network 54. Communication interfaces 44 and 26 are responsible for the exchange of various information between processor 46 and processor 28 via network 54. The exchange of various information between processor 46 and processor 28 using communication interfaces 44 and 26 is performed in a secure manner.

[0215] Figure 4 shows an example of the main functions of the data processing device 12 and the smart glasses 214. As shown in Figure 4, the data processing device 12 performs specific processing using the processor 28. The storage 32 stores the specific processing program 56.

[0216] The specific processing program 56 is an example of a "program" relating to the technology of this disclosure. The processor 28 reads the specific processing program 56 from the storage 32 and executes the read specific processing program 56 on the RAM 30. The specific processing is realized by the processor 28 operating as a specific processing unit 290 in accordance with the specific processing program 56 executed on the RAM 30.

[0217] The storage 32 stores the data generation model 58 and the emotion identification model 59. The data generation model 58 and the emotion identification model 59 are used by the identification processing unit 290.

[0218] In the smart glasses 214, the processor 46 performs the reception output processing. The storage 50 stores the reception output program 60. The processor 46 reads the reception output program 60 from the storage 50 and executes the read reception output program 60 on the RAM 48. The reception output processing is realized by the processor 46 operating as a control unit 46A according to the reception output program 60 executed on the RAM 48.

[0219] Next, the identification processing performed by the identification processing unit 290 of the data processing device 12 will be described. In the following description, the data processing device 12 will be referred to as the "server" and the smart glasses 214 will be referred to as the "terminal".

[0220] This invention is a system that uses a smart device to acquire a user's biometric information and, based on that information, provides acoustic data optimized for the user's mental state. The details are described below.

[0221] The server receives biometric information transmitted from the terminal in real time. This information primarily includes heart rate and is used as data indicating the user's current physical and mental state. The received data is analyzed to assess the user's situation, whether relaxation is needed, or whether concentration is required.

[0222] The terminal typically functions as a smartphone or tablet and is responsible for acquiring biometric information from portable devices such as smartwatches. The terminal transmits the acquired data to a server via the internet. Upon receiving processed results or acoustic data from the server, the terminal becomes an interface for providing that data to the user.

[0223] This system allows users to receive dynamically generated and selected audio data based on their heart rate and mental state. This audio data comes in various formats, including relaxing music and rhythmic patterns designed to enhance concentration. Users can operate the application to listen to the corresponding audio data in real time.

[0224] For example, if concentration is needed during class or work, the device will determine that the heart rate is constant or indicates a state of tension. If the analysis indicates the user needs improved concentration, the server selects music to enhance concentration and sends it to the device. The audio data is played through the device to support the user's concentration. This allows the user to receive support tailored to their mental state.

[0225] The following describes the processing flow.

[0226] Step 1:

[0227] The user launches the smartphone app and connects it to the smartwatch. The app prepares to periodically collect heart rate data from the smartwatch via Bluetooth or other means.

[0228] Step 2:

[0229] The device acquires biometric information, such as heart rate, from the smartwatch. This data is updated at regular intervals, reflecting current information in real time.

[0230] Step 3:

[0231] The device organizes the acquired biometric information into data packets and sends them to the server via the internet. The communication is encrypted to ensure secure transmission.

[0232] Step 4:

[0233] The server analyzes biometric information received from the terminal. Based on the heart rate data, the server evaluates the user's current state and determines what kind of acoustic data is appropriate.

[0234] Step 5:

[0235] Based on the analysis results, the server selects or generates appropriate acoustic data. Specifically, it uses AI to choose music that is most suitable for the user's condition, referencing information on mental care and music therapy.

[0236] Step 6:

[0237] The server compresses the selected or generated audio data and sends it to the terminal in streaming format.

[0238] Step 7:

[0239] The device decompresses the audio data received from the server and begins playback in real time. Users can listen to the played audio data via a smartphone app.

[0240] Step 8:

[0241] Users can provide feedback on audio data through the app. This feedback is sent to the server and used to improve future data selection and generation algorithms.

[0242] (Example 1)

[0243] Next, we will describe Example 1. In the following description, the data processing device 12 will be referred to as the "server," and the smart glasses 214 will be referred to as the "terminal."

[0244] In modern society, individual users routinely experience psychological problems such as stress and lack of concentration. Under these circumstances, there are limited means of dynamically providing acoustic information tailored to each user's psychological state. Furthermore, conventional solutions lack sufficient technology to individually incorporate user feedback. Therefore, there is a need for a more adaptable and reliable system that can deliver an optimized acoustic experience for each individual.

[0245] The identification process performed by the identification processing unit 290 of the data processing device 12 in Example 1 is realized by the following means.

[0246] In this invention, the server includes means for acquiring the user's biometric parameters, means for evaluating the user's psychological state based on the acquired biometric parameters, and means for generating or selecting optimal acoustic information for the user based on the evaluation results. This makes it possible to provide appropriate acoustic information that corresponds to the user's unique psychological state.

[0247] "Biometric parameters" refer to information that indicates the user's physical condition, and include data such as heart rate, skin temperature, and blood pressure.

[0248] "Psychological state" refers to the emotional or mental condition of the user, and represents mental conditions such as stress levels and concentration levels.

[0249] "Acoustic information" refers to sound wave data provided to users, specifically a collection of sounds that have a psychological effect, such as music, nature sounds, and white noise.

[0250] A "generative AI model" is a form of artificial intelligence that uses learning algorithms to analyze data and generate new data.

[0251] "Feedback" refers to the evaluations and opinions that users provide regarding the acoustic information they receive, and is used as response information to improve and optimize the system.

[0252] This invention is a system for providing individually optimized acoustic information based on the user's biological parameters. The details of this system are described below.

[0253] Server operation

[0254] The server centrally receives biometric parameters from users transmitted from multiple terminals. These biometric parameters include heart rate, skin temperature, and blood pressure. The server analyzes this data in real time to estimate the user's psychological state. Specifically, it uses a generative AI model to analyze fluctuations in biometric parameters and evaluate stress levels and concentration levels. Based on this analysis, the server selects appropriate acoustic information from a digital library for the user and sends it to the terminal. The software used includes data mining tools and machine learning algorithms.

[0255] Terminal operation

[0256] The device typically functions as a smartphone or tablet, receiving biometric parameters acquired by smartwatches and other wearable devices via Bluetooth or other wireless communication methods. The device has the capability to transmit this data to a server over the internet. It also functions as an interface providing the user with analysis results and acoustic information received from the server. The application on the device features a user interface that allows the user to easily play acoustic information and perform operations such as volume adjustment and playback / pause.

[0257] User experience

[0258] Users can receive individually generated and selected acoustic information based on their heart rate and mental state via their smart devices. This includes relaxing music and rhythmic patterns that improve concentration. Users can experience the acoustics in real time using the provided application and send feedback to the server through the application. This feedback will be used to improve future acoustic information provision using a generative AI model.

[0259] Specific example

[0260] For example, when a user is doing work that requires concentration, the device can determine their current level of tension from heart rate data. The server analyzes this data, selects music that will help the user concentrate, and sends it to the device. The audio information is played through the device to support the user's concentration.

[0261] Example of a prompt

[0262] "What music would you recommend for when you want to relax?"

[0263] "I want to concentrate on my work, what kind of music would be good?"

[0264] This system allows users to dynamically receive sound support tailored to their different psychological states, thereby improving their quality of daily life.

[0265] The flow of the specific processing in Example 1 will be explained using Figure 11.

[0266] Step 1:

[0267] The terminal acquires biometric parameters from the wearable device worn by the user. Specifically, it uses Bluetooth communication to receive data such as heart rate and skin temperature. The input is biometric data from the wearable device, and the output is formatted data ready to be sent to the server.

[0268] Step 2:

[0269] The device transmits acquired biometric parameters to the server. Specifically, it uploads data to the server in real time using a secure communication protocol (e.g., HTTPS). The input is sensor data, and the output is the data sent to the server.

[0270] Step 3:

[0271] The server analyzes the received biometric parameters. Specifically, it uses a generative AI model to analyze heart rate fluctuations, thereby evaluating the user's psychological state. It generates analysis results including stress levels and concentration levels. The input is biometric data sent from the terminal, and the output is the analyzed psychological state evaluation result.

[0272] Step 4:

[0273] The server selects the most suitable acoustic information for the user based on the analysis results. It queries a digital library to identify music and rhythms that match the user's psychological state. The input is the evaluation results, and the output is the selected acoustic information.

[0274] Step 5:

[0275] The server transmits the selected acoustic information to the terminal. The acoustic information is converted into a data format and sent to the terminal via the internet. The input is the selected acoustic information, and the output is the data transmitted to the terminal.

[0276] Step 6:

[0277] The terminal provides the user with the acoustic information received from the server. As a specific operation, it decodes the acoustic data and performs an operation of playing it through a speaker or headphones. The input is the acoustic information from the server, and the output is the played sound.

[0278] Step 7:

[0279] The user transmits feedback on the acoustic information via the terminal. Through the interface on the app, the user inputs subjective evaluations and sends them to the server. The input is the feedback data, and the output is the evaluation information sent to the server.

[0280] Step 8:

[0281] The server updates the AI model generated based on the user's feedback. By analyzing the feedback and adjusting the model's algorithm, the accuracy of providing acoustic information in subsequent times is improved. The input is the feedback information, and the output is the updated AI model.

[0282] (Application Example 1)

[0283] Next, Application Example 1 will be described. In the following description, the data processing device 12 is referred to as the "server", and the smart glasses 214 are referred to as the "terminal".

[0284] Conventionally, the methods for improving the customer experience in a store have been limited, and it has been difficult to provide an optimal environment according to the individual mental states of customers. In particular, although it has been required to dynamically adjust the acoustic environment according to the mental state of customers to enhance the willingness to purchase or create a relaxing atmosphere, its realization has been difficult.

[0285] The specific processing by the specific processing unit 290 of the data processing device 12 in Application Example 1 is realized by the following respective means.

[0286] In this invention, the server includes means for acquiring the user's biometric data, means for determining the user's mental state based on the acquired biometric data, and means for generating or selecting acoustic information suitable for the user based on the determination result. Thereby, it becomes possible to provide an optimal acoustic environment according to the mental state of customers in the store and improve the customer experience.

[0287] The "user" refers to individual customers or consumers who provide biometric data and receive acoustic information according to their mental state.

[0288] The "biometric data" is information indicating the physical state of the user, such as the user's heart rate, and is data used for determining the mental state.

[0289] The "mental state" refers to the mental health and emotional state of the user at the current time, and includes states such as relaxation and concentration.

[0290] The "acoustic information" is data of sounds such as music and sound effects, which is selected or generated based on the mental state.

[0291] The "means for generating or selecting" is a process or device for creating or selecting appropriate acoustic information based on the determination result of the mental state.

[0292] The "means for distributing acoustic information in the store environment" is a system for providing the selected acoustic information to customers in the physical or virtual store space.

[0293] This invention is a system that utilizes the user's biometric data and provides optimal acoustic information in the store environment to improve the customer experience.

[0294] The server acquires biometric data, such as heart rate, from the user's smartwatch or equivalent portable device. This data is transmitted to the server via the internet, where a machine learning model written in Python is used to determine the user's psychological state in real time. This machine learning model is capable of distinguishing between states of relaxation and concentration.

[0295] Based on the determined psychological state, the server selects or generates the most suitable audio information for the user. This audio information may be selected using the Spotify API or similar music streaming platforms. The selected audio information is then distributed via Wi-Fi or Bluetooth to speaker systems installed within the store. This ensures that soothing music plays when the user wants to relax, and music with an appropriate tempo plays when they need to concentrate.

[0296] As a concrete example, if a customer visiting a shopping mall is determined to be stressed due to an elevated heart rate, the server will select healing music and distribute it throughout the store to promote relaxation. An example of a prompt message in this case would be: "Heart rate data: 85, Mental state: Needs relaxation, Music type: Pleasant melody."

[0297] The flow of a specific process in Application Example 1 will be explained using Figure 12.

[0298] Step 1:

[0299] The device acquires heart rate data from the user's smartwatch. This data is obtained directly from the device's built-in sensors and forms the basis for processing as the user's biometric information. This data is then sent to a server for subsequent analysis processes.

[0300] Step 2:

[0301] The server receives the heart rate data transmitted from the terminal. Using this input data, the server performs data analysis in a machine learning model built with Python. Specifically, the model processes the heart rate data as a feedback loop to determine the user's mental state (e.g., stress state or relaxation state).

[0302] Step 3:

[0303] Based on the determination result of the machine learning model, the server selects appropriate acoustic information. Referring to the data of the mental state received as input, the server searches for and selects suitable music from the music library using an API such as the Spotify API. The selected acoustic information is sent to subsequent distribution processing.

[0304] Step 4:

[0305] The server distributes the selected acoustic information to the speaker system in the store via WiFi or Bluetooth. As a result, music suitable for the user's mental state starts to play in the store environment. Since the acoustic information is played in real time, the user's experience will be in line with their physiological state.

[0306] Step 5:

[0307] The user can enjoy shopping in a comfortable environment according to their mental state by listening to the music playing in the store. An example of a prompt sentence, "Heart rate data: 85, Mental state: Need relaxation, Song selection type: Comfortable melody," presented as an example triggers this series of processes.

[0308] Furthermore, an emotion engine for estimating the user's emotions may be combined. That is, the specific processing unit 290 may estimate the user's emotions using the emotion recognition model 59 and perform specific processing using the user's emotions.

[0309] This invention is a system that combines the user's biometric information with an emotion engine to more accurately evaluate the user's mental state and provide appropriate acoustic data. In this embodiment, a smart device and a communication terminal work together to analyze the user's emotions in real time using data such as the user's heart rate, voice, and facial expressions.

[0310] The server receives biometric and emotional data transmitted from the terminal and integrates this information to evaluate the user's current mental state. The emotion engine uses machine learning models to analyze the user's voice tone and facial movements to recognize emotions such as joy, sadness, and surprise. This enables a comprehensive evaluation of mental state that includes not only biometric information but also emotional elements.

[0311] The device provides a user-friendly interface. Specifically, it uses a smartphone or PC to send the user's voice and facial expression data to an emotion engine, which then sends the results to a server. Furthermore, the device receives appropriate audio data from the server and plays it back to the user in real time.

[0312] Users can play music through the application and provide feedback on the music's impact on their emotions. This feedback information is used to improve the quality of future music generation and selection processes. For example, if the server determines that a user is experiencing stress, it will select relaxing music based on the emotion engine's recognition results and provide it to the user via the device.

[0313] This system goes beyond conventional music delivery based on biometric information, making it possible to provide highly accurate care for a variety of mental states stemming from emotions.

[0314] The following describes the processing flow.

[0315] Step 1:

[0316] The user launches the smartphone app and connects it to the smartwatch and smart device. The app prepares to collect data such as the user's heart rate, voice, and facial expressions.

[0317] Step 2:

[0318] The device acquires heart rate data in real time from the smartwatch, and also collects the user's voice and facial expression data via a smartphone or webcam. This allows for the accumulation of comprehensive biometric information.

[0319] Step 3:

[0320] The device combines collected heart rate data with voice and facial expression data to generate data packets, which are then transmitted to a server via the internet. Communication is conducted using a secure method.

[0321] Step 4:

[0322] The server analyzes the received biometric and emotional data. It evaluates changes in heart rate, voice tone, and facial expressions to identify the user's current mental state.

[0323] Step 5:

[0324] The emotion engine operates on the server and uses machine learning models to recognize the user's emotions from their voice and facial expressions. Based on the recognition results, it determines whether the user is happy, sad, or otherwise.

[0325] Step 6:

[0326] The server generates or selects appropriate acoustic data based on the analysis results and emotion recognition results. The acoustic data is prepared to be effective for relaxation and energy enhancement.

[0327] Step 7:

[0328] The server sends the generated or selected audio data to the terminal. The terminal receives this data and plays music for the user.

[0329] Step 8:

[0330] Users can provide feedback on the effects of the music while listening to it. This feedback is sent from the device to the server and used to improve the performance of the emotion engine and the accuracy of the sound data.

[0331] (Example 2)

[0332] Next, we will describe Example 2. In the following description, the data processing device 12 will be referred to as the "server" and the smart glasses 214 will be referred to as the "terminal".

[0333] A comprehensive assessment of a user's mental state is needed, taking into account not only their biometric information but also emotional information such as voice and facial expressions. However, conventional systems have the challenge of being unable to integrate and analyze this diverse data and provide accurate acoustic data in real time.

[0334] The identification process performed by the identification processing unit 290 of the data processing device 12 in Example 2 is realized by the following means.

[0335] In this invention, the server includes means for integrating and evaluating the user's biometric information, voice, and facial expression data; means for selecting appropriate acoustic data based on the evaluated mental state; and means for providing the selected acoustic data in real time. This enables a more accurate and comprehensive evaluation of the mental state and the provision of detailed acoustic data that corresponds to the user's emotional state.

[0336] "User" refers to an individual who uses this system to perform biometric and emotional analysis.

[0337] "Biometric information" refers to data that indicates an individual's physical or physiological state, such as heart rate, voice, and facial expressions.

[0338] "Mental state" refers to the user's psychological and emotional state, including emotions such as joy, sadness, and surprise.

[0339] "Audio data" refers to digital data related to the user's voice, including their tone and tempo.

[0340] "Facial expression data" refers to digital data related to facial movements and changes, and is used to infer the user's emotions.

[0341] "Audio data" refers to digital data related to sound, such as music and sound effects, intended to influence the user's emotions and mental state.

[0342] "Integration" refers to the process of combining and analyzing biometric information and emotional data to evaluate the overall mental state.

[0343] "Real-time" refers to a situation where latency is minimized, allowing users to receive data services instantly.

[0344] This invention is a system that integrates the user's biometric information and emotional data to more accurately analyze the user's mental state and provide acoustic data based on that analysis. This system operates through the coordinated efforts of three entities: a server, a terminal, and a user.

[0345] The server receives biometric information (heart rate, voice, facial expression data, etc.) transmitted from the terminal and analyzes this data using an emotion engine. The emotion engine utilizes machine learning models to recognize emotions based on voice tone and facial movements. This allows for a comprehensive evaluation of emotional and physiological data, enabling an assessment of the user's current mental state.

[0346] On the other hand, the terminal provides an interface with the user. Using a smart device or personal computer, it sends the user's voice and facial expression data to the emotion engine and then sends the results to the server. It also has the function of receiving appropriate audio data sent from the server and playing it back to the user in real time.

[0347] Users provide biometric data through their smart devices and listen to the generated acoustic data. Furthermore, they can provide feedback on how the acoustic data affects their emotions. This feedback helps improve the accuracy of the next acoustic data generation and selection process. For example, if a user is feeling stressed, the system can select and play relaxing music in real time.

[0348] An example of a prompt might be, "Generate a music selection algorithm to provide relaxing music to users experiencing stress." In this way, the system can assess the user's mental state in real time and provide an appropriate sound experience based on that assessment.

[0349] The flow of the specific processing in Example 2 will be explained using Figure 13.

[0350] Step 1:

[0351] The user acquires biometric information such as heart rate, voice, and facial expressions using a smart device. The user uses a smartphone or wearable device to measure biometric data in real time through sensors on the device. The input to this process is the user's heart rate, voice, and facial expression data, and the output is transmitted to the terminal as biometric information.

[0352] Step 2:

[0353] The device transmits the acquired biometric information to the server. Specifically, the device transfers the information to the server as data packets via an internet connection. The input for this step is the biometric information collected from the user, and the output is the digital data sent to the server.

[0354] Step 3:

[0355] The server analyzes the received biometric information, along with the user's voice and facial expression data, using an emotion engine. On the server, a machine learning model processes the input data, analyzing voice tone and facial expressions to recognize the emotional state. In this analysis process, the input is digital data received from the terminal, and the output is the analyzed emotional information.

[0356] Step 4:

[0357] The server evaluates the user's current mental state based on the analyzed emotional information. This process integrates biometric and emotional information to comprehensively assess the mental state using a computer algorithm. The input in this step is the analyzed emotional information, and the output is the evaluated mental state data.

[0358] Step 5:

[0359] The server selects appropriate sound data based on the assessed mental state. For example, if the user is assessed as feeling stressed, it selects music with a relaxing effect. The input is mental state data, and the output is the selected sound data.

[0360] Step 6:

[0361] The selected audio data is sent from the server to the terminal. The server encodes the data and sends it in a format that the terminal can receive. The input for this step is the audio data, and the output is the audio file sent to the terminal.

[0362] Step 7:

[0363] The terminal plays the received audio data in real time and provides it to the user. The terminal plays music using an audio player, which the user can listen to. The input to this process is the received audio file, and the output is the audio playback as a user experience.

[0364] Step 8:

[0365] Users provide feedback on the impact of the provided audio data on their emotions. This feedback is sent to the server via the application and used to improve the audio data selection process. The input for this step is the user's experience with the audio data, and the output is the feedback information sent through the application.

[0366] (Application Example 2)

[0367] Next, we will explain application example 2. In the following explanation, the data processing device 12 will be referred to as the "server," and the smart glasses 214 will be referred to as the "terminal."

[0368] Traditional methods have struggled to accurately provide acoustic information based on the user's mental state. In particular, evaluations based solely on simple biometric data may fail to adequately reflect the user's emotions through music. As a result, there are limitations to enhancing the user's emotional satisfaction.

[0369] The specific processing performed by the specific processing unit 290 of the data processing device 12 in Application Example 2 is realized by the following means.

[0370] In this invention, the server includes means for acquiring the user's biometric data, means for evaluating the user's mental state based on the acquired biometric data, voice tone, and facial movements, means for using a communication device to generate or select acoustic information suitable for the user based on the evaluation results and provide it to the user, means for analyzing voice tone and facial movements using an emotion analysis engine, and means for selecting acoustic information from a database based on the analysis results and providing it to the user in real time. This makes it possible to accurately evaluate the user's emotions and provide acoustic information suitable for those emotions.

[0371] "Biometric data" refers to individual physical and physiological information such as a user's heart rate pattern, voice tone, and facial movements.

[0372] "Mental state" refers to the emotional and psychological state or tendencies of the user, and includes emotions such as joy, sadness, and surprise.

[0373] "Acoustic information" refers to data of sounds and music that are suitable for the user's mental state.

[0374] "Communication equipment" refers to the hardware and software infrastructure used to send and receive data between servers and terminals and to provide acoustic information to users in real time.

[0375] An "emotion analysis engine" refers to a software component that uses machine learning models to analyze the tone of a user's voice and facial movements to recognize their emotions.

[0376] A "database" is a system for storing and managing audio information provided to users.

[0377] The system for realizing this invention uses a terminal equipped with various sensors and a server with advanced analytical capabilities. Specifically, a terminal such as a smartphone collects the user's biometric data and transmits that data to the server via the internet. The user's biometric data includes heart rate patterns, voice tone, and facial movements.

[0378] The server runs an emotion analysis engine that utilizes machine learning models based on biometric data. This emotion analysis engine uses libraries such as TensorFlow for processing audio and images. Based on the analysis results, the user's mental state is evaluated, and the acoustic information best suited to that state is selected from the database. In this process, a cloud database such as Firebase is used to manage and provide the acoustic data.

[0379] If the system determines that a user needs relaxation, it might select a piece of music, such as a classical piano sonata, and stream it in real time. This is expected to promote mental relaxation in the user.

[0380] As an example of a prompt to a generative AI model, the question, "What genre of music is suitable when a user needs to relax?" is used to optimize the system's suggestions. In this way, the user is provided with the most personalized music experience possible.

[0381] The flow of a specific process in Application Example 2 will be explained using Figure 14.

[0382] Step 1:

[0383] The device collects biometric data such as the user's heart rate pattern, voice tone, and facial movements using various sensors. This data is temporarily stored within the device and then transmitted to a server via the internet. The input is the user's biometric data, and the output is the generation of data packets to be sent to the server.

[0384] Step 2:

[0385] The server receives biometric data from the terminal and processes it as input data for running a machine learning model. Specifically, it uses TensorFlow to analyze voice tone and facial movements to evaluate the user's mental state. This process outputs mental states such as "relaxed," "lively," and "sad."

[0386] Step 3:

[0387] The server selects the most suitable audio information from a Firebase database based on the mental state assessment results. The selection process involves rule-based data processing, such as matching the assessment results and selecting "classical music" for a relaxed state. The input is the mental state assessment results, and the output is the selected audio information.

[0388] Step 4:

[0389] The server transmits the selected audio information to the terminal. The audio information is converted into real-time streaming data and prepared to be provided to the user. The input is the selected audio information, and the output is streaming data that can be played on the user's terminal.

[0390] Step 5:

[0391] Users enjoy audio information provided through their devices. Actual playback takes place through the smartphone's audio function, allowing users to obtain a musical experience suited to their own mental state. The input is streaming data, and the output is the musical experience mediated through the user's hearing.

[0392] The specific processing unit 290 transmits the result of the specific processing to the smart glasses 214. In the smart glasses 214, the control unit 46A causes the speaker 240 to output the result of the specific processing. The microphone 238 acquires audio indicating user input for the result of the specific processing. The control unit 46A transmits the audio data indicating user input acquired by the microphone 238 to the data processing unit 12. In the data processing unit 12, the specific processing unit 290 acquires the audio data.

[0393] Data generation model 58 is a type of so-called generative AI (Artificial Intelligence). One example of data generation model 58 is ChatGPT (Internet Search).<URL: https: / / openai.com / blog / chatgpt> ), Gemini (Internet search) <url: https: gemini.google.com ?hl="ja">Examples of generative AI include the following. The data generation model 58 is obtained by performing deep learning on a neural network. The data generation model 58 is input with prompts containing instructions, and with inference data such as audio data representing speech, text data representing text, and image data representing images. The data generation model 58 infers from the input inference data according to the instructions indicated by the prompts, and outputs the inference results in data formats such as audio data and text data. Here, inference refers to, for example, analysis, classification, prediction, and / or summarization.

[0394] In the above embodiment, an example was given in which specific processing is performed by the data processing device 12, but the technology of this disclosure is not limited thereto, and the specific processing may also be performed by the smart glasses 214.

[0395] [Third Embodiment]

[0396] Figure 5 shows an example of the configuration of the data processing system 310 according to the third embodiment.

[0397] As shown in Figure 5, the data processing system 310 includes a data processing device 12 and a headset terminal 314. An example of the data processing device 12 is a server.

[0398] The data processing device 12 comprises a computer 22, a database 24, and a communication interface 26. The computer 22 is an example of a "computer" related to the technology of this disclosure. The computer 22 comprises a processor 28, RAM 30, and storage 32. The processor 28, RAM 30, and storage 32 are connected to a bus 34. The database 24 and the communication interface 26 are also connected to the bus 34. The communication interface 26 is connected to a network 54. An example of the network 54 is a WAN (Wide Area Network) and / or a LAN (Local Area Network).

[0399] The headset terminal 314 includes a computer 36, a microphone 238, a speaker 240, a camera 42, a communication interface 44, and a display 343. The computer 36 includes a processor 46, RAM 48, and storage 50. The processor 46, RAM 48, and storage 50 are connected to a bus 52. The microphone 238, speaker 240, camera 42, and display 343 are also connected to the bus 52.

[0400] The microphone 238 receives voice signals from the user 20 and receives instructions from the user 20. The microphone 238 captures the voice signals from the user 20, converts the captured voice into audio data, and outputs it to the processor 46. The speaker 240 outputs audio according to the instructions from the processor 46.

[0401] Camera 42 is a small digital camera equipped with an optical system including a lens, aperture, and shutter, and an image sensor such as a CMOS (Complementary Metal-Oxide-Semiconductor) image sensor or a CCD (Charge Coupled Device) image sensor, and captures images of the area around the user 20 (for example, an imaging range defined by a field of view equivalent to the width of a typical healthy person's field of vision).

[0402] Communication interface 44 is connected to network 54. Communication interfaces 44 and 26 are responsible for the exchange of various information between processor 46 and processor 28 via network 54. The exchange of various information between processor 46 and processor 28 using communication interfaces 44 and 26 is performed in a secure manner.

[0403] Figure 6 shows an example of the main functions of the data processing device 12 and the headset terminal 314. As shown in Figure 6, the data processing device 12 performs specific processing using the processor 28. The storage 32 stores the specific processing program 56.

[0404] The specific processing program 56 is an example of a "program" relating to the technology of this disclosure. The processor 28 reads the specific processing program 56 from the storage 32 and executes the read specific processing program 56 on the RAM 30. The specific processing is realized by the processor 28 operating as a specific processing unit 290 in accordance with the specific processing program 56 executed on the RAM 30.

[0405] The storage 32 stores the data generation model 58 and the emotion identification model 59. The data generation model 58 and the emotion identification model 59 are used by the identification processing unit 290.

[0406] In the headset terminal 314, the processor 46 performs the reception output processing. The storage 50 stores the reception output program 60. The processor 46 reads the reception output program 60 from the storage 50 and executes the read reception output program 60 on the RAM 48. The reception output processing is realized by the processor 46 operating as a control unit 46A according to the reception output program 60 executed on the RAM 48.

[0407] Next, the specific processing performed by the specific processing unit 290 of the data processing device 12 will be described. In the following description, the data processing device 12 will be referred to as the "server" and the headset terminal 314 will be referred to as the "terminal".

[0408] This invention is a system that uses a smart device to acquire a user's biometric information and, based on that information, provides acoustic data optimized for the user's mental state. The details are described below.

[0409] The server receives biometric information transmitted from the terminal in real time. This information primarily includes heart rate and is used as data indicating the user's current physical and mental state. The received data is analyzed to assess the user's situation, whether relaxation is needed, or whether concentration is required.

[0410] The terminal typically functions as a smartphone or tablet and is responsible for acquiring biometric information from portable devices such as smartwatches. The terminal transmits the acquired data to a server via the internet. Upon receiving processed results or acoustic data from the server, the terminal becomes an interface for providing that data to the user.

[0411] This system allows users to receive dynamically generated and selected audio data based on their heart rate and mental state. This audio data comes in various formats, including relaxing music and rhythmic patterns designed to enhance concentration. Users can operate the application to listen to the corresponding audio data in real time.

[0412] For example, if concentration is needed during class or work, the device will determine that the heart rate is constant or indicates a state of tension. If the analysis indicates the user needs improved concentration, the server selects music to enhance concentration and sends it to the device. The audio data is played through the device to support the user's concentration. This allows the user to receive support tailored to their mental state.

[0413] The following describes the processing flow.

[0414] Step 1:

[0415] The user launches the smartphone app and connects it to the smartwatch. The app prepares to periodically collect heart rate data from the smartwatch via Bluetooth or other means.

[0416] Step 2:

[0417] The device acquires biometric information, such as heart rate, from the smartwatch. This data is updated at regular intervals, reflecting current information in real time.

[0418] Step 3:

[0419] The device organizes the acquired biometric information into data packets and sends them to the server via the internet. The communication is encrypted to ensure secure transmission.

[0420] Step 4:

[0421] The server analyzes biometric information received from the terminal. Based on the heart rate data, the server evaluates the user's current state and determines what kind of acoustic data is appropriate.

[0422] Step 5:

[0423] Based on the analysis results, the server selects or generates appropriate acoustic data. Specifically, it uses AI to choose music that is most suitable for the user's condition, referencing information on mental care and music therapy.

[0424] Step 6:

[0425] The server compresses the selected or generated audio data and sends it to the terminal in streaming format.

[0426] Step 7:

[0427] The device decompresses the audio data received from the server and begins playback in real time. Users can listen to the played audio data via a smartphone app.

[0428] Step 8:

[0429] Users can provide feedback on audio data through the app. This feedback is sent to the server and used to improve future data selection and generation algorithms.

[0430] (Example 1)

[0431] Next, we will describe Example 1. In the following description, the data processing device 12 will be referred to as the "server," and the headset-type terminal 314 will be referred to as the "terminal."

[0432] In modern society, individual users routinely experience psychological problems such as stress and lack of concentration. Under these circumstances, there are limited means of dynamically providing acoustic information tailored to each user's psychological state. Furthermore, conventional solutions lack sufficient technology to individually incorporate user feedback. Therefore, there is a need for a more adaptable and reliable system that can deliver an optimized acoustic experience for each individual.

[0433] The identification process performed by the identification processing unit 290 of the data processing device 12 in Example 1 is realized by the following means.

[0434] In this invention, the server includes means for acquiring the user's biometric parameters, means for evaluating the user's psychological state based on the acquired biometric parameters, and means for generating or selecting optimal acoustic information for the user based on the evaluation results. This makes it possible to provide appropriate acoustic information that corresponds to the user's unique psychological state.

[0435] "Biometric parameters" refer to information that indicates the user's physical condition, and include data such as heart rate, skin temperature, and blood pressure.

[0436] "Psychological state" refers to the emotional or mental condition of the user, and represents mental conditions such as stress levels and concentration levels.

[0437] "Acoustic information" refers to sound wave data provided to users, specifically a collection of sounds that have a psychological effect, such as music, nature sounds, and white noise.

[0438] A "generative AI model" is a form of artificial intelligence that uses learning algorithms to analyze data and generate new data.

[0439] "Feedback" refers to the evaluations and opinions that users provide regarding the acoustic information they receive, and is used as response information to improve and optimize the system.

[0440] This invention is a system for providing individually optimized acoustic information based on the user's biological parameters. The details of this system are described below.

[0441] Server operation

[0442] The server centrally receives biometric parameters from users transmitted from multiple terminals. These biometric parameters include heart rate, skin temperature, and blood pressure. The server analyzes this data in real time to estimate the user's psychological state. Specifically, it uses a generative AI model to analyze fluctuations in biometric parameters and evaluate stress levels and concentration levels. Based on this analysis, the server selects appropriate acoustic information from a digital library for the user and sends it to the terminal. The software used includes data mining tools and machine learning algorithms.

[0443] Terminal operation

[0444] The device typically functions as a smartphone or tablet, receiving biometric parameters acquired by smartwatches and other wearable devices via Bluetooth or other wireless communication methods. The device has the capability to transmit this data to a server over the internet. It also functions as an interface providing the user with analysis results and acoustic information received from the server. The application on the device features a user interface that allows the user to easily play acoustic information and perform operations such as volume adjustment and playback / pause.

[0445] User experience

[0446] Users can receive individually generated and selected acoustic information based on their heart rate and mental state via their smart devices. This includes relaxing music and rhythmic patterns that improve concentration. Users can experience the acoustics in real time using the provided application and send feedback to the server through the application. This feedback will be used to improve future acoustic information provision using a generative AI model.

[0447] Specific example

[0448] For example, when a user is doing work that requires concentration, the device can determine their current level of tension from heart rate data. The server analyzes this data, selects music that will help the user concentrate, and sends it to the device. The audio information is played through the device to support the user's concentration.

[0449] Example of a prompt

[0450] "What music would you recommend for when you want to relax?"

[0451] "I want to concentrate on my work, what kind of music would be good?"

[0452] This system allows users to dynamically receive sound support tailored to their different psychological states, thereby improving their quality of daily life.

[0453] The flow of the specific processing in Example 1 will be explained using Figure 11.

[0454] Step 1:

[0455] The terminal acquires biometric parameters from the wearable device worn by the user. Specifically, it uses Bluetooth communication to receive data such as heart rate and skin temperature. The input is biometric data from the wearable device, and the output is formatted data ready to be sent to the server.

[0456] Step 2:

[0457] The device transmits acquired biometric parameters to the server. Specifically, it uploads data to the server in real time using a secure communication protocol (e.g., HTTPS). The input is sensor data, and the output is the data sent to the server.

[0458] Step 3:

[0459] The server analyzes the received biometric parameters. Specifically, it uses a generative AI model to analyze heart rate fluctuations, thereby evaluating the user's psychological state. It generates analysis results including stress levels and concentration levels. The input is biometric data sent from the terminal, and the output is the analyzed psychological state evaluation result.

[0460] Step 4:

[0461] The server selects the most suitable acoustic information for the user based on the analysis results. It queries a digital library to identify music and rhythms that match the user's psychological state. The input is the evaluation results, and the output is the selected acoustic information.

[0462] Step 5:

[0463] The server transmits the selected acoustic information to the terminal. The acoustic information is converted into a data format and sent to the terminal via the internet. The input is the selected acoustic information, and the output is the data transmitted to the terminal.

[0464] Step 6:

[0465] The terminal provides the user with audio information received from the server. Specifically, it decodes the audio data and plays it back through speakers or headphones. The input is the audio information from the server, and the output is the played sound.

[0466] Step 7:

[0467] Users send feedback on acoustic information via their devices. They input subjective evaluations through an interface within the app and send them to the server. The input is feedback data, and the output is the evaluation information sent to the server.

[0468] Step 8:

[0469] The server updates the generated AI model based on user feedback. By analyzing the feedback and adjusting the model's algorithm, it improves the accuracy of future acoustic information provision. The input is the feedback information, and the output is the updated AI model.

[0470] (Application Example 1)

[0471] Next, we will explain Application Example 1. In the following explanation, the data processing device 12 will be referred to as the "server," and the headset-type terminal 314 will be referred to as the "terminal."

[0472] Traditionally, methods for improving the customer experience within stores have been limited, making it difficult to provide an optimal environment tailored to each customer's individual mental state. In particular, dynamically adjusting the sound environment according to the customer's mental state to enhance purchasing intent or create a relaxed atmosphere has been challenging to achieve.

[0473] The specific processing performed by the specific processing unit 290 of the data processing device 12 in Application Example 1 is realized by the following means.

[0474] In this invention, the server includes means for acquiring the user's biometric data, means for determining the user's psychological state based on the acquired biometric data, and means for generating or selecting acoustic information suitable for the user based on the determination result. This makes it possible to provide an optimal acoustic environment in a store that corresponds to the customer's mental state, thereby improving the customer experience.

[0475] "Users" refer to individual customers or consumers who provide biometric data and receive acoustic information tailored to their mental state.

[0476] "Biometric data" refers to information that indicates a user's physical condition, such as their heart rate, and is used to assess their psychological state.

[0477] "Psychological state" refers to the user's current mental health and emotional state, including states such as relaxation and concentration.

[0478] "Acoustic information" refers to sound data, such as music and sound effects, that are selected or generated based on a person's mental state.

[0479] "Means for generating or selecting" refers to processes or devices for creating or selecting appropriate acoustic information based on the results of a psychological state assessment.

[0480] "Means for delivering acoustic information to a store environment" refers to a system for providing selected acoustic information to customers within a physical or virtual store space.

[0481] This invention is a system that improves the customer experience by utilizing the user's biometric data to provide optimal acoustic information in the store environment.

[0482] The server acquires biometric data, such as heart rate, from the user's smartwatch or equivalent portable device. This data is transmitted to the server via the internet, where a machine learning model written in Python is used to determine the user's psychological state in real time. This machine learning model is capable of distinguishing between states of relaxation and concentration.

[0483] Based on the determined psychological state, the server selects or generates the most suitable audio information for the user. This audio information may be selected using the Spotify API or similar music streaming platforms. The selected audio information is then distributed via Wi-Fi or Bluetooth to speaker systems installed within the store. This ensures that soothing music plays when the user wants to relax, and music with an appropriate tempo plays when they need to concentrate.

[0484] As a concrete example, if a customer visiting a shopping mall is determined to be stressed due to an elevated heart rate, the server will select healing music and distribute it throughout the store to promote relaxation. An example of a prompt message in this case would be: "Heart rate data: 85, Mental state: Needs relaxation, Music type: Pleasant melody."

[0485] The flow of a specific process in Application Example 1 will be explained using Figure 12.

[0486] Step 1:

[0487] The device acquires heart rate data from the user's smartwatch. This data is obtained directly from the device's built-in sensors and forms the basis for processing as the user's biometric information. This data is then sent to a server for subsequent analysis processes.

[0488] Step 2:

[0489] The server receives heart rate data transmitted from the terminal. The server uses this input data to perform data analysis using a machine learning model built in Python. Specifically, the model processes the heart rate data as a feedback loop to determine the user's psychological state (e.g., stressed or relaxed).

[0490] Step 3:

[0491] The server selects appropriate audio information based on the results of a machine learning model. It references the psychological state data received as input, searches for suitable music from its music library using the Spotify API, and selects the appropriate track. The selected audio information is then sent to the subsequent distribution process.

[0492] Step 4:

[0493] The server distributes selected audio information to the store's speaker system via Wi-Fi or Bluetooth. This causes music tailored to the customer's mental state to begin playing in the store environment. Because the audio information is played in real time, the customer's experience is in line with their biological state.

[0494] Step 5:

[0495] Users can enjoy shopping in a comfortable environment tailored to their mental state by listening to the music playing in the store. An example prompt message, "Heart rate data: 85, Mental state: Need to relax, Music selection type: Pleasant melody," serves as the trigger to initiate this process.

[0496] Furthermore, an emotion engine that estimates the user's emotions may be incorporated. That is, the identification processing unit 290 may use the emotion identification model 59 to estimate the user's emotions and perform identification processing using the user's emotions.

[0497] This invention is a system that combines the user's biometric information with an emotion engine to more accurately evaluate the user's mental state and provide appropriate acoustic data. In this embodiment, a smart device and a communication terminal work together to analyze the user's emotions in real time using data such as the user's heart rate, voice, and facial expressions.

[0498] The server receives biometric and emotional data transmitted from the terminal and integrates this information to evaluate the user's current mental state. The emotion engine uses machine learning models to analyze the user's voice tone and facial movements to recognize emotions such as joy, sadness, and surprise. This enables a comprehensive evaluation of mental state that includes not only biometric information but also emotional elements.

[0499] The device provides a user-friendly interface. Specifically, it uses a smartphone or PC to send the user's voice and facial expression data to an emotion engine, which then sends the results to a server. Furthermore, the device receives appropriate audio data from the server and plays it back to the user in real time.

[0500] Users can play music through the application and provide feedback on the music's impact on their emotions. This feedback information is used to improve the quality of future music generation and selection processes. For example, if the server determines that a user is experiencing stress, it will select relaxing music based on the emotion engine's recognition results and provide it to the user via the device.

[0501] This system goes beyond conventional music delivery based on biometric information, making it possible to provide highly accurate care for a variety of mental states stemming from emotions.

[0502] The following describes the processing flow.

[0503] Step 1:

[0504] The user launches the smartphone app and connects it to the smartwatch and smart device. The app prepares to collect data such as the user's heart rate, voice, and facial expressions.

[0505] Step 2:

[0506] The device acquires heart rate data in real time from the smartwatch, and also collects the user's voice and facial expression data via a smartphone or webcam. This allows for the accumulation of comprehensive biometric information.

[0507] Step 3:

[0508] The device combines collected heart rate data with voice and facial expression data to generate data packets, which are then transmitted to a server via the internet. Communication is conducted using a secure method.

[0509] Step 4:

[0510] The server analyzes the received biometric and emotional data. It evaluates changes in heart rate, voice tone, and facial expressions to identify the user's current mental state.

[0511] Step 5:

[0512] The emotion engine operates on the server and uses machine learning models to recognize the user's emotions from their voice and facial expressions. Based on the recognition results, it determines whether the user is happy, sad, or otherwise.

[0513] Step 6:

[0514] The server generates or selects appropriate acoustic data based on the analysis results and emotion recognition results. The acoustic data is prepared to be effective for relaxation and energy enhancement.

[0515] Step 7:

[0516] The server sends the generated or selected audio data to the terminal. The terminal receives this data and plays music for the user.

[0517] Step 8:

[0518] Users can provide feedback on the effects of the music while listening to it. This feedback is sent from the device to the server and used to improve the performance of the emotion engine and the accuracy of the sound data.

[0519] (Example 2)

[0520] Next, we will describe Example 2. In the following description, the data processing device 12 will be referred to as the "server," and the headset-type terminal 314 will be referred to as the "terminal."

[0521] A comprehensive assessment of a user's mental state is needed, taking into account not only their biometric information but also emotional information such as voice and facial expressions. However, conventional systems have the challenge of being unable to integrate and analyze this diverse data and provide accurate acoustic data in real time.

[0522] The identification process performed by the identification processing unit 290 of the data processing device 12 in Example 2 is realized by the following means.

[0523] In this invention, the server includes means for integrating and evaluating the user's biometric information, voice, and facial expression data; means for selecting appropriate acoustic data based on the evaluated mental state; and means for providing the selected acoustic data in real time. This enables a more accurate and comprehensive evaluation of the mental state and the provision of detailed acoustic data that corresponds to the user's emotional state.

[0524] "User" refers to an individual who uses this system to perform biometric and emotional analysis.

[0525] "Biometric information" refers to data that indicates an individual's physical or physiological state, such as heart rate, voice, and facial expressions.

[0526] "Mental state" refers to the user's psychological and emotional state, including emotions such as joy, sadness, and surprise.

[0527] "Audio data" refers to digital data related to the user's voice, including their tone and tempo.

[0528] "Facial expression data" refers to digital data related to facial movements and changes, and is used to infer the user's emotions.

[0529] "Audio data" refers to digital data related to sound, such as music and sound effects, intended to influence the user's emotions and mental state.

[0530] "Integration" refers to the process of combining and analyzing biometric information and emotional data to evaluate the overall mental state.

[0531] "Real-time" refers to a situation where latency is minimized, allowing users to receive data services instantly.

[0532] This invention is a system that integrates the user's biometric information and emotional data to more accurately analyze the user's mental state and provide acoustic data based on that analysis. This system operates through the coordinated efforts of three entities: a server, a terminal, and a user.

[0533] The server receives biometric information (heart rate, voice, facial expression data, etc.) transmitted from the terminal and analyzes this data using an emotion engine. The emotion engine utilizes machine learning models to recognize emotions based on voice tone and facial movements. This allows for a comprehensive evaluation of emotional and physiological data, enabling an assessment of the user's current mental state.

[0534] On the other hand, the terminal provides an interface with the user. Using a smart device or personal computer, it sends the user's voice and facial expression data to the emotion engine and then sends the results to the server. It also has the function of receiving appropriate audio data sent from the server and playing it back to the user in real time.

[0535] Users provide biometric data through their smart devices and listen to the generated acoustic data. Furthermore, they can provide feedback on how the acoustic data affects their emotions. This feedback helps improve the accuracy of the next acoustic data generation and selection process. For example, if a user is feeling stressed, the system can select and play relaxing music in real time.

[0536] An example of a prompt might be, "Generate a music selection algorithm to provide relaxing music to users experiencing stress." In this way, the system can assess the user's mental state in real time and provide an appropriate sound experience based on that assessment.

[0537] The flow of the specific processing in Example 2 will be explained using Figure 13.

[0538] Step 1:

[0539] The user acquires biometric information such as heart rate, voice, and facial expressions using a smart device. The user uses a smartphone or wearable device to measure biometric data in real time through sensors on the device. The input to this process is the user's heart rate, voice, and facial expression data, and the output is transmitted to the terminal as biometric information.

[0540] Step 2:

[0541] The device transmits the acquired biometric information to the server. Specifically, the device transfers the information to the server as data packets via an internet connection. The input for this step is the biometric information collected from the user, and the output is the digital data sent to the server.

[0542] Step 3:

[0543] The server analyzes the received biometric information, along with the user's voice and facial expression data, using an emotion engine. On the server, a machine learning model processes the input data, analyzing voice tone and facial expressions to recognize the emotional state. In this analysis process, the input is digital data received from the terminal, and the output is the analyzed emotional information.

[0544] Step 4:

[0545] The server evaluates the user's current mental state based on the analyzed emotional information. This process integrates biometric and emotional information to comprehensively assess the mental state using a computer algorithm. The input in this step is the analyzed emotional information, and the output is the evaluated mental state data.

[0546] Step 5:

[0547] The server selects appropriate sound data based on the assessed mental state. For example, if the user is assessed as feeling stressed, it selects music with a relaxing effect. The input is mental state data, and the output is the selected sound data.

[0548] Step 6:

[0549] The selected audio data is sent from the server to the terminal. The server encodes the data and sends it in a format that the terminal can receive. The input for this step is the audio data, and the output is the audio file sent to the terminal.

[0550] Step 7:

[0551] The terminal plays the received audio data in real time and provides it to the user. The terminal plays music using an audio player, which the user can listen to. The input to this process is the received audio file, and the output is the audio playback as a user experience.

[0552] Step 8:

[0553] Users provide feedback on the impact of the provided audio data on their emotions. This feedback is sent to the server via the application and used to improve the audio data selection process. The input for this step is the user's experience with the audio data, and the output is the feedback information sent through the application.

[0554] (Application Example 2)

[0555] Next, we will explain application example 2. In the following explanation, the data processing device 12 will be referred to as the "server," and the headset-type terminal 314 will be referred to as the "terminal."

[0556] Traditional methods have struggled to accurately provide acoustic information based on the user's mental state. In particular, evaluations based solely on simple biometric data may fail to adequately reflect the user's emotions through music. As a result, there are limitations to enhancing the user's emotional satisfaction.

[0557] The specific processing performed by the specific processing unit 290 of the data processing device 12 in Application Example 2 is realized by the following means.

[0558] In this invention, the server includes means for acquiring the user's biometric data, means for evaluating the user's mental state based on the acquired biometric data, voice tone, and facial movements, means for using a communication device to generate or select acoustic information suitable for the user based on the evaluation results and provide it to the user, means for analyzing voice tone and facial movements using an emotion analysis engine, and means for selecting acoustic information from a database based on the analysis results and providing it to the user in real time. This makes it possible to accurately evaluate the user's emotions and provide acoustic information suitable for those emotions.

[0559] "Biometric data" refers to individual physical and physiological information such as a user's heart rate pattern, voice tone, and facial movements.

[0560] "Mental state" refers to the emotional and psychological state or tendencies of the user, and includes emotions such as joy, sadness, and surprise.

[0561] "Acoustic information" refers to data of sounds and music that are suitable for the user's mental state.

[0562] "Communication equipment" refers to the hardware and software infrastructure used to send and receive data between servers and terminals and to provide acoustic information to users in real time.

[0563] An "emotion analysis engine" refers to a software component that uses machine learning models to analyze the tone of a user's voice and facial movements to recognize their emotions.

[0564] A "database" is a system for storing and managing audio information provided to users.

[0565] The system for realizing this invention uses a terminal equipped with various sensors and a server with advanced analytical capabilities. Specifically, a terminal such as a smartphone collects the user's biometric data and transmits that data to the server via the internet. The user's biometric data includes heart rate patterns, voice tone, and facial movements.

[0566] The server runs an emotion analysis engine that utilizes machine learning models based on biometric data. This emotion analysis engine uses libraries such as TensorFlow for processing audio and images. Based on the analysis results, the user's mental state is evaluated, and the acoustic information best suited to that state is selected from the database. In this process, a cloud database such as Firebase is used to manage and provide the acoustic data.

[0567] If the system determines that a user needs relaxation, it might select a piece of music, such as a classical piano sonata, and stream it in real time. This is expected to promote mental relaxation in the user.

[0568] As an example of a prompt to a generative AI model, the question, "What genre of music is suitable when a user needs to relax?" is used to optimize the system's suggestions. In this way, the user is provided with the most personalized music experience possible.

[0569] The flow of a specific process in Application Example 2 will be explained using Figure 14.

[0570] Step 1:

[0571] The device collects biometric data such as the user's heart rate pattern, voice tone, and facial movements using various sensors. This data is temporarily stored within the device and then transmitted to a server via the internet. The input is the user's biometric data, and the output is the generation of data packets to be sent to the server.

[0572] Step 2:

[0573] The server receives biometric data from the terminal and processes it as input data for running a machine learning model. Specifically, it uses TensorFlow to analyze voice tone and facial movements to evaluate the user's mental state. This process outputs mental states such as "relaxed," "lively," and "sad."

[0574] Step 3:

[0575] The server selects the most suitable audio information from a Firebase database based on the mental state assessment results. The selection process involves rule-based data processing, such as matching the assessment results and selecting "classical music" for a relaxed state. The input is the mental state assessment results, and the output is the selected audio information.

[0576] Step 4:

[0577] The server transmits the selected audio information to the terminal. The audio information is converted into real-time streaming data and prepared to be provided to the user. The input is the selected audio information, and the output is streaming data that can be played on the user's terminal.

[0578] Step 5:

[0579] Users enjoy audio information provided through their devices. Actual playback takes place through the smartphone's audio function, allowing users to obtain a musical experience suited to their own mental state. The input is streaming data, and the output is the musical experience mediated through the user's hearing.

[0580] The specific processing unit 290 transmits the result of the specific processing to the headset terminal 314. In the headset terminal 314, the control unit 46A causes the speaker 240 and display 343 to output the result of the specific processing. The microphone 238 acquires audio indicating user input for the result of the specific processing. The control unit 46A transmits the audio data indicating user input acquired by the microphone 238 to the data processing unit 12. In the data processing unit 12, the specific processing unit 290 acquires the audio data.

[0581] Data generation model 58 is a type of so-called generative AI (Artificial Intelligence). One example of data generation model 58 is ChatGPT (Internet Search).<URL: https: / / openai.com / blog / chatgpt> ), Gemini (Internet search) <url: https: gemini.google.com ?hl="ja">Examples of generative AI include the following. The data generation model 58 is obtained by performing deep learning on a neural network. The data generation model 58 is input with prompts containing instructions, and with inference data such as audio data representing speech, text data representing text, and image data representing images. The data generation model 58 infers from the input inference data according to the instructions indicated by the prompts, and outputs the inference results in data formats such as audio data and text data. Here, inference refers to, for example, analysis, classification, prediction, and / or summarization.

[0582] In the above embodiment, an example was given in which specific processing is performed by the data processing device 12, but the technology of this disclosure is not limited thereto, and specific processing may also be performed by the headset terminal 314.

[0583] [Fourth Embodiment]

[0584] Figure 7 shows an example of the configuration of the data processing system 410 according to the fourth embodiment.

[0585] As shown in Figure 7, the data processing system 410 includes a data processing device 12 and a robot 414. An example of the data processing device 12 is a server.

[0586] The data processing device 12 comprises a computer 22, a database 24, and a communication interface 26. The computer 22 is an example of a "computer" related to the technology of this disclosure. The computer 22 comprises a processor 28, RAM 30, and storage 32. The processor 28, RAM 30, and storage 32 are connected to a bus 34. The database 24 and the communication interface 26 are also connected to the bus 34. The communication interface 26 is connected to a network 54. An example of the network 54 is a WAN (Wide Area Network) and / or a LAN (Local Area Network).

[0587] The robot 414 includes a computer 36, a microphone 238, a speaker 240, a camera 42, a communication interface 44, and a controlled object 443. The computer 36 includes a processor 46, RAM 48, and storage 50. The processor 46, RAM 48, and storage 50 are connected to a bus 52. The microphone 238, speaker 240, camera 42, and controlled object 443 are also connected to the bus 52.

[0588] The microphone 238 receives voice signals from the user 20 and receives instructions from the user 20. The microphone 238 captures the voice signals from the user 20, converts the captured voice into audio data, and outputs it to the processor 46. The speaker 240 outputs audio according to the instructions from the processor 46.

[0589] Camera 42 is a small digital camera equipped with an optical system including a lens, aperture, and shutter, and an image sensor such as a CMOS (Complementary Metal-Oxide-Semiconductor) image sensor or a CCD (Charge Coupled Device) image sensor, and captures images of the area around the user 20 (for example, an imaging range defined by a field of view equivalent to the width of a typical healthy person's field of vision).

[0590] Communication interface 44 is connected to network 54. Communication interfaces 44 and 26 are responsible for the exchange of various information between processor 46 and processor 28 via network 54. The exchange of various information between processor 46 and processor 28 using communication interfaces 44 and 26 is performed in a secure manner.

[0591] The controlled object 443 includes a display device, LEDs in the eyes, and motors that drive the arms, hands, and feet. The posture and gestures of the robot 414 are controlled by controlling the motors of the arms, hands, and feet. Some of the robot 414's emotions can be expressed by controlling these motors. Furthermore, the robot 414's facial expressions can also be expressed by controlling the illumination state of the LEDs in its eyes.

[0592] Figure 8 shows an example of the main functions of the data processing device 12 and the robot 414. As shown in Figure 8, the data processing device 12 performs specific processing using the processor 28. The storage 32 stores the specific processing program 56.

[0593] The specific processing program 56 is an example of a "program" relating to the technology of this disclosure. The processor 28 reads the specific processing program 56 from the storage 32 and executes the read specific processing program 56 on the RAM 30. The specific processing is realized by the processor 28 operating as a specific processing unit 290 in accordance with the specific processing program 56 executed on the RAM 30.

[0594] The storage 32 stores the data generation model 58 and the emotion identification model 59. The data generation model 58 and the emotion identification model 59 are used by the identification processing unit 290.

[0595] In robot 414, the processor 46 performs the reception output processing. The storage 50 stores the reception output program 60. The processor 46 reads the reception output program 60 from the storage 50 and executes the read reception output program 60 on the RAM 48. The reception output processing is realized by the processor 46 operating as a control unit 46A according to the reception output program 60 executed on the RAM 48.

[0596] Next, the specific processing performed by the specific processing unit 290 of the data processing device 12 will be described. In the following description, the data processing device 12 will be referred to as the "server" and the robot 414 as the "terminal".

[0597] This invention is a system that uses a smart device to acquire a user's biometric information and, based on that information, provides acoustic data optimized for the user's mental state. The details are described below.

[0598] The server receives biometric information transmitted from the terminal in real time. This information primarily includes heart rate and is used as data indicating the user's current physical and mental state. The received data is analyzed to assess the user's situation, whether relaxation is needed, or whether concentration is required.

[0599] The terminal typically functions as a smartphone or tablet and is responsible for acquiring biometric information from portable devices such as smartwatches. The terminal transmits the acquired data to a server via the internet. Upon receiving processed results or acoustic data from the server, the terminal becomes an interface for providing that data to the user.

[0600] This system allows users to receive dynamically generated and selected audio data based on their heart rate and mental state. This audio data comes in various formats, including relaxing music and rhythmic patterns designed to enhance concentration. Users can operate the application to listen to the corresponding audio data in real time.

[0601] For example, if concentration is needed during class or work, the device will determine that the heart rate is constant or indicates a state of tension. If the analysis indicates the user needs improved concentration, the server selects music to enhance concentration and sends it to the device. The audio data is played through the device to support the user's concentration. This allows the user to receive support tailored to their mental state.

[0602] The following describes the processing flow.

[0603] Step 1:

[0604] The user launches the smartphone app and connects it to the smartwatch. The app prepares to periodically collect heart rate data from the smartwatch via Bluetooth or other means.

[0605] Step 2:

[0606] The device acquires biometric information, such as heart rate, from the smartwatch. This data is updated at regular intervals, reflecting current information in real time.

[0607] Step 3:

[0608] The device organizes the acquired biometric information into data packets and sends them to the server via the internet. The communication is encrypted to ensure secure transmission.

[0609] Step 4:

[0610] The server analyzes biometric information received from the terminal. Based on the heart rate data, the server evaluates the user's current state and determines what kind of acoustic data is appropriate.

[0611] Step 5:

[0612] Based on the analysis results, the server selects or generates appropriate acoustic data. Specifically, it uses AI to choose music that is most suitable for the user's condition, referencing information on mental care and music therapy.

[0613] Step 6:

[0614] The server compresses the selected or generated audio data and sends it to the terminal in streaming format.

[0615] Step 7:

[0616] The device decompresses the audio data received from the server and begins playback in real time. Users can listen to the played audio data via a smartphone app.

[0617] Step 8:

[0618] Users can provide feedback on audio data through the app. This feedback is sent to the server and used to improve future data selection and generation algorithms.

[0619] (Example 1)

[0620] Next, we will describe Example 1. In the following description, the data processing device 12 will be referred to as the "server" and the robot 414 as the "terminal".

[0621] In modern society, individual users routinely experience psychological problems such as stress and lack of concentration. Under these circumstances, there are limited means of dynamically providing acoustic information tailored to each user's psychological state. Furthermore, conventional solutions lack sufficient technology to individually incorporate user feedback. Therefore, there is a need for a more adaptable and reliable system that can deliver an optimized acoustic experience for each individual.

[0622] The identification process performed by the identification processing unit 290 of the data processing device 12 in Example 1 is realized by the following means.

[0623] In this invention, the server includes means for acquiring the user's biometric parameters, means for evaluating the user's psychological state based on the acquired biometric parameters, and means for generating or selecting optimal acoustic information for the user based on the evaluation results. This makes it possible to provide appropriate acoustic information that corresponds to the user's unique psychological state.

[0624] "Biometric parameters" refer to information that indicates the user's physical condition, and include data such as heart rate, skin temperature, and blood pressure.

[0625] "Psychological state" refers to the emotional or mental condition of the user, and represents mental conditions such as stress levels and concentration levels.

[0626] "Acoustic information" refers to sound wave data provided to users, specifically a collection of sounds that have a psychological effect, such as music, nature sounds, and white noise.

[0627] A "generative AI model" is a form of artificial intelligence that uses learning algorithms to analyze data and generate new data.

[0628] "Feedback" refers to the evaluations and opinions that users provide regarding the acoustic information they receive, and is used as response information to improve and optimize the system.

[0629] This invention is a system for providing individually optimized acoustic information based on the user's biological parameters. The details of this system are described below.

[0630] Server operation

[0631] The server centrally receives biometric parameters from users transmitted from multiple terminals. These biometric parameters include heart rate, skin temperature, and blood pressure. The server analyzes this data in real time to estimate the user's psychological state. Specifically, it uses a generative AI model to analyze fluctuations in biometric parameters and evaluate stress levels and concentration levels. Based on this analysis, the server selects appropriate acoustic information from a digital library for the user and sends it to the terminal. The software used includes data mining tools and machine learning algorithms.

[0632] Terminal operation

[0633] The device typically functions as a smartphone or tablet, receiving biometric parameters acquired by smartwatches and other wearable devices via Bluetooth or other wireless communication methods. The device has the capability to transmit this data to a server over the internet. It also functions as an interface providing the user with analysis results and acoustic information received from the server. The application on the device features a user interface that allows the user to easily play acoustic information and perform operations such as volume adjustment and playback / pause.

[0634] User experience

[0635] Users can receive individually generated and selected acoustic information based on their heart rate and mental state via their smart devices. This includes relaxing music and rhythmic patterns that improve concentration. Users can experience the acoustics in real time using the provided application and send feedback to the server through the application. This feedback will be used to improve future acoustic information provision using a generative AI model.

[0636] Specific example

[0637] For example, when a user is doing work that requires concentration, the device can determine their current level of tension from heart rate data. The server analyzes this data, selects music that will help the user concentrate, and sends it to the device. The audio information is played through the device to support the user's concentration.

[0638] Example of a prompt

[0639] "What music would you recommend for when you want to relax?"

[0640] "I want to concentrate on my work, what kind of music would be good?"

[0641] This system allows users to dynamically receive sound support tailored to their different psychological states, thereby improving their quality of daily life.

[0642] The flow of the specific processing in Example 1 will be explained using Figure 11.

[0643] Step 1:

[0644] The terminal acquires biometric parameters from the wearable device worn by the user. Specifically, it uses Bluetooth communication to receive data such as heart rate and skin temperature. The input is biometric data from the wearable device, and the output is formatted data ready to be sent to the server.

[0645] Step 2:

[0646] The device transmits acquired biometric parameters to the server. Specifically, it uploads data to the server in real time using a secure communication protocol (e.g., HTTPS). The input is sensor data, and the output is the data sent to the server.

[0647] Step 3:

[0648] The server analyzes the received biometric parameters. Specifically, it uses a generative AI model to analyze heart rate fluctuations, thereby evaluating the user's psychological state. It generates analysis results including stress levels and concentration levels. The input is biometric data sent from the terminal, and the output is the analyzed psychological state evaluation result.

[0649] Step 4:

[0650] The server selects the most suitable acoustic information for the user based on the analysis results. It queries a digital library to identify music and rhythms that match the user's psychological state. The input is the evaluation results, and the output is the selected acoustic information.

[0651] Step 5:

[0652] The server transmits the selected acoustic information to the terminal. The acoustic information is converted into a data format and sent to the terminal via the internet. The input is the selected acoustic information, and the output is the data transmitted to the terminal.

[0653] Step 6:

[0654] The terminal provides the user with audio information received from the server. Specifically, it decodes the audio data and plays it back through speakers or headphones. The input is the audio information from the server, and the output is the played sound.

[0655] Step 7:

[0656] Users send feedback on acoustic information via their devices. They input subjective evaluations through an interface within the app and send them to the server. The input is feedback data, and the output is the evaluation information sent to the server.

[0657] Step 8:

[0658] The server updates the generated AI model based on user feedback. By analyzing the feedback and adjusting the model's algorithm, it improves the accuracy of future acoustic information provision. The input is the feedback information, and the output is the updated AI model.

[0659] (Application Example 1)

[0660] Next, we will explain Application Example 1. In the following explanation, the data processing device 12 will be referred to as the "server" and the robot 414 as the "terminal".

[0661] Traditionally, methods for improving the customer experience within stores have been limited, making it difficult to provide an optimal environment tailored to each customer's individual mental state. In particular, dynamically adjusting the sound environment according to the customer's mental state to enhance purchasing intent or create a relaxed atmosphere has been challenging to achieve.

[0662] The specific processing performed by the specific processing unit 290 of the data processing device 12 in Application Example 1 is realized by the following means.

[0663] In this invention, the server includes means for acquiring the user's biometric data, means for determining the user's psychological state based on the acquired biometric data, and means for generating or selecting acoustic information suitable for the user based on the determination result. This makes it possible to provide an optimal acoustic environment in a store that corresponds to the customer's mental state, thereby improving the customer experience.

[0664] "Users" refer to individual customers or consumers who provide biometric data and receive acoustic information tailored to their mental state.

[0665] "Biometric data" refers to information that indicates a user's physical condition, such as their heart rate, and is used to assess their psychological state.

[0666] "Psychological state" refers to the user's current mental health and emotional state, including states such as relaxation and concentration.

[0667] "Acoustic information" refers to sound data, such as music and sound effects, that are selected or generated based on a person's mental state.

[0668] "Means for generating or selecting" refers to processes or devices for creating or selecting appropriate acoustic information based on the results of a psychological state assessment.

[0669] "Means for delivering acoustic information to a store environment" refers to a system for providing selected acoustic information to customers within a physical or virtual store space.

[0670] This invention is a system that improves the customer experience by utilizing the user's biometric data to provide optimal acoustic information in the store environment.

[0671] The server acquires biometric data, such as heart rate, from the user's smartwatch or equivalent portable device. This data is transmitted to the server via the internet, where a machine learning model written in Python is used to determine the user's psychological state in real time. This machine learning model is capable of distinguishing between states of relaxation and concentration.

[0672] Based on the determined psychological state, the server selects or generates the most suitable audio information for the user. This audio information may be selected using the Spotify API or similar music streaming platforms. The selected audio information is then distributed via Wi-Fi or Bluetooth to speaker systems installed within the store. This ensures that soothing music plays when the user wants to relax, and music with an appropriate tempo plays when they need to concentrate.

[0673] As a concrete example, if a customer visiting a shopping mall is determined to be stressed due to an elevated heart rate, the server will select healing music and distribute it throughout the store to promote relaxation. An example of a prompt message in this case would be: "Heart rate data: 85, Mental state: Needs relaxation, Music type: Pleasant melody."

[0674] The flow of a specific process in Application Example 1 will be explained using Figure 12.

[0675] Step 1:

[0676] The device acquires heart rate data from the user's smartwatch. This data is obtained directly from the device's built-in sensors and forms the basis for processing as the user's biometric information. This data is then sent to a server for subsequent analysis processes.

[0677] Step 2:

[0678] The server receives heart rate data transmitted from the terminal. The server uses this input data to perform data analysis using a machine learning model built in Python. Specifically, the model processes the heart rate data as a feedback loop to determine the user's psychological state (e.g., stressed or relaxed).

[0679] Step 3:

[0680] The server selects appropriate audio information based on the results of a machine learning model. It references the psychological state data received as input, searches for suitable music from its music library using the Spotify API, and selects the appropriate track. The selected audio information is then sent to the subsequent distribution process.

[0681] Step 4:

[0682] The server distributes selected audio information to the store's speaker system via Wi-Fi or Bluetooth. This causes music tailored to the customer's mental state to begin playing in the store environment. Because the audio information is played in real time, the customer's experience is in line with their biological state.

[0683] Step 5:

[0684] Users can enjoy shopping in a comfortable environment tailored to their mental state by listening to the music playing in the store. An example prompt message, "Heart rate data: 85, Mental state: Need to relax, Music selection type: Pleasant melody," serves as the trigger to initiate this process.

[0685] Furthermore, an emotion engine that estimates the user's emotions may be incorporated. That is, the identification processing unit 290 may use the emotion identification model 59 to estimate the user's emotions and perform identification processing using the user's emotions.

[0686] This invention is a system that combines the user's biometric information with an emotion engine to more accurately evaluate the user's mental state and provide appropriate acoustic data. In this embodiment, a smart device and a communication terminal work together to analyze the user's emotions in real time using data such as the user's heart rate, voice, and facial expressions.

[0687] The server receives biometric and emotional data transmitted from the terminal and integrates this information to evaluate the user's current mental state. The emotion engine uses machine learning models to analyze the user's voice tone and facial movements to recognize emotions such as joy, sadness, and surprise. This enables a comprehensive evaluation of mental state that includes not only biometric information but also emotional elements.

[0688] The device provides a user-friendly interface. Specifically, it uses a smartphone or PC to send the user's voice and facial expression data to an emotion engine, which then sends the results to a server. Furthermore, the device receives appropriate audio data from the server and plays it back to the user in real time.

[0689] Users can play music through the application and provide feedback on the music's impact on their emotions. This feedback information is used to improve the quality of future music generation and selection processes. For example, if the server determines that a user is experiencing stress, it will select relaxing music based on the emotion engine's recognition results and provide it to the user via the device.

[0690] This system goes beyond conventional music delivery based on biometric information, making it possible to provide highly accurate care for a variety of mental states stemming from emotions.

[0691] The following describes the processing flow.

[0692] Step 1:

[0693] The user launches the smartphone app and connects it to the smartwatch and smart device. The app prepares to collect data such as the user's heart rate, voice, and facial expressions.

[0694] Step 2:

[0695] The device acquires heart rate data in real time from the smartwatch, and also collects the user's voice and facial expression data via a smartphone or webcam. This allows for the accumulation of comprehensive biometric information.

[0696] Step 3:

[0697] The device combines collected heart rate data with voice and facial expression data to generate data packets, which are then transmitted to a server via the internet. Communication is conducted using a secure method.

[0698] Step 4:

[0699] The server analyzes the received biometric and emotional data. It evaluates changes in heart rate, voice tone, and facial expressions to identify the user's current mental state.

[0700] Step 5:

[0701] The emotion engine operates on the server and uses machine learning models to recognize the user's emotions from their voice and facial expressions. Based on the recognition results, it determines whether the user is happy, sad, or otherwise.

[0702] Step 6:

[0703] The server generates or selects appropriate acoustic data based on the analysis results and emotion recognition results. The acoustic data is prepared to be effective for relaxation and energy enhancement.

[0704] Step 7:

[0705] The server sends the generated or selected audio data to the terminal. The terminal receives this data and plays music for the user.

[0706] Step 8:

[0707] Users can provide feedback on the effects of the music while listening to it. This feedback is sent from the device to the server and used to improve the performance of the emotion engine and the accuracy of the sound data.

[0708] (Example 2)

[0709] Next, we will describe Example 2. In the following description, the data processing device 12 will be referred to as the "server" and the robot 414 as the "terminal".

[0710] A comprehensive assessment of a user's mental state is needed, taking into account not only their biometric information but also emotional information such as voice and facial expressions. However, conventional systems have the challenge of being unable to integrate and analyze this diverse data and provide accurate acoustic data in real time.

[0711] The identification process performed by the identification processing unit 290 of the data processing device 12 in Example 2 is realized by the following means.

[0712] In this invention, the server includes means for integrating and evaluating the user's biometric information, voice, and facial expression data; means for selecting appropriate acoustic data based on the evaluated mental state; and means for providing the selected acoustic data in real time. This enables a more accurate and comprehensive evaluation of the mental state and the provision of detailed acoustic data that corresponds to the user's emotional state.

[0713] "User" refers to an individual who uses this system to perform biometric and emotional analysis.

[0714] "Biometric information" refers to data that indicates an individual's physical or physiological state, such as heart rate, voice, and facial expressions.

[0715] "Mental state" refers to the user's psychological and emotional state, including emotions such as joy, sadness, and surprise.

[0716] "Audio data" refers to digital data related to the user's voice, including their tone and tempo.

[0717] "Facial expression data" refers to digital data related to facial movements and changes, and is used to infer the user's emotions.

[0718] "Audio data" refers to digital data related to sound, such as music and sound effects, intended to influence the user's emotions and mental state.

[0719] "Integration" refers to the process of combining and analyzing biometric information and emotional data to evaluate the overall mental state.

[0720] "Real-time" refers to a situation where latency is minimized, allowing users to receive data services instantly.

[0721] This invention is a system that integrates the user's biometric information and emotional data to more accurately analyze the user's mental state and provide acoustic data based on that analysis. This system operates through the coordinated efforts of three entities: a server, a terminal, and a user.

[0722] The server receives biometric information (heart rate, voice, facial expression data, etc.) transmitted from the terminal and analyzes this data using an emotion engine. The emotion engine utilizes machine learning models to recognize emotions based on voice tone and facial movements. This allows for a comprehensive evaluation of emotional and physiological data, enabling an assessment of the user's current mental state.

[0723] On the other hand, the terminal provides an interface with the user. Using a smart device or personal computer, it sends the user's voice and facial expression data to the emotion engine and then sends the results to the server. It also has the function of receiving appropriate audio data sent from the server and playing it back to the user in real time.

[0724] Users provide biometric data through their smart devices and listen to the generated acoustic data. Furthermore, they can provide feedback on how the acoustic data affects their emotions. This feedback helps improve the accuracy of the next acoustic data generation and selection process. For example, if a user is feeling stressed, the system can select and play relaxing music in real time.

[0725] An example of a prompt might be, "Generate a music selection algorithm to provide relaxing music to users experiencing stress." In this way, the system can assess the user's mental state in real time and provide an appropriate sound experience based on that assessment.

[0726] The flow of the specific processing in Example 2 will be explained using Figure 13.

[0727] Step 1:

[0728] The user acquires biometric information such as heart rate, voice, and facial expressions using a smart device. The user uses a smartphone or wearable device to measure biometric data in real time through sensors on the device. The input to this process is the user's heart rate, voice, and facial expression data, and the output is transmitted to the terminal as biometric information.

[0729] Step 2:

[0730] The device transmits the acquired biometric information to the server. Specifically, the device transfers the information to the server as data packets via an internet connection. The input for this step is the biometric information collected from the user, and the output is the digital data sent to the server.

[0731] Step 3:

[0732] The server analyzes the received biometric information, along with the user's voice and facial expression data, using an emotion engine. On the server, a machine learning model processes the input data, analyzing voice tone and facial expressions to recognize the emotional state. In this analysis process, the input is digital data received from the terminal, and the output is the analyzed emotional information.

[0733] Step 4:

[0734] The server evaluates the user's current mental state based on the analyzed emotional information. This process integrates biometric and emotional information to comprehensively assess the mental state using a computer algorithm. The input in this step is the analyzed emotional information, and the output is the evaluated mental state data.

[0735] Step 5:

[0736] The server selects appropriate sound data based on the assessed mental state. For example, if the user is assessed as feeling stressed, it selects music with a relaxing effect. The input is mental state data, and the output is the selected sound data.

[0737] Step 6:

[0738] The selected audio data is sent from the server to the terminal. The server encodes the data and sends it in a format that the terminal can receive. The input for this step is the audio data, and the output is the audio file sent to the terminal.

[0739] Step 7:

[0740] The terminal plays the received audio data in real time and provides it to the user. The terminal plays music using an audio player, which the user can listen to. The input to this process is the received audio file, and the output is the audio playback as a user experience.

[0741] Step 8:

[0742] Users provide feedback on the impact of the provided audio data on their emotions. This feedback is sent to the server via the application and used to improve the audio data selection process. The input for this step is the user's experience with the audio data, and the output is the feedback information sent through the application.

[0743] (Application Example 2)

[0744] Next, we will explain application example 2. In the following explanation, the data processing device 12 will be referred to as the "server" and the robot 414 as the "terminal".

[0745] Traditional methods have struggled to accurately provide acoustic information based on the user's mental state. In particular, evaluations based solely on simple biometric data may fail to adequately reflect the user's emotions through music. As a result, there are limitations to enhancing the user's emotional satisfaction.

[0746] The specific processing performed by the specific processing unit 290 of the data processing device 12 in Application Example 2 is realized by the following means.

[0747] In this invention, the server includes means for acquiring the user's biometric data, means for evaluating the user's mental state based on the acquired biometric data, voice tone, and facial movements, means for using a communication device to generate or select acoustic information suitable for the user based on the evaluation results and provide it to the user, means for analyzing voice tone and facial movements using an emotion analysis engine, and means for selecting acoustic information from a database based on the analysis results and providing it to the user in real time. This makes it possible to accurately evaluate the user's emotions and provide acoustic information suitable for those emotions.

[0748] "Biometric data" refers to individual physical and physiological information such as a user's heart rate pattern, voice tone, and facial movements.

[0749] "Mental state" refers to the emotional and psychological state or tendencies of the user, and includes emotions such as joy, sadness, and surprise.

[0750] "Acoustic information" refers to data of sounds and music that are suitable for the user's mental state.

[0751] "Communication equipment" refers to the hardware and software infrastructure used to send and receive data between servers and terminals and to provide acoustic information to users in real time.

[0752] An "emotion analysis engine" refers to a software component that uses machine learning models to analyze the tone of a user's voice and facial movements to recognize their emotions.

[0753] A "database" is a system for storing and managing audio information provided to users.

[0754] The system for realizing this invention uses a terminal equipped with various sensors and a server with advanced analytical capabilities. Specifically, a terminal such as a smartphone collects the user's biometric data and transmits that data to the server via the internet. The user's biometric data includes heart rate patterns, voice tone, and facial movements.

[0755] The server runs an emotion analysis engine that utilizes machine learning models based on biometric data. This emotion analysis engine uses libraries such as TensorFlow for processing audio and images. Based on the analysis results, the user's mental state is evaluated, and the acoustic information best suited to that state is selected from the database. In this process, a cloud database such as Firebase is used to manage and provide the acoustic data.

[0756] If the system determines that a user needs relaxation, it might select a piece of music, such as a classical piano sonata, and stream it in real time. This is expected to promote mental relaxation in the user.

[0757] As an example of a prompt to a generative AI model, the question, "What genre of music is suitable when a user needs to relax?" is used to optimize the system's suggestions. In this way, the user is provided with the most personalized music experience possible.

[0758] The flow of a specific process in Application Example 2 will be explained using Figure 14.

[0759] Step 1:

[0760] The device collects biometric data such as the user's heart rate pattern, voice tone, and facial movements using various sensors. This data is temporarily stored within the device and then transmitted to a server via the internet. The input is the user's biometric data, and the output is the generation of data packets to be sent to the server.

[0761] Step 2:

[0762] The server receives biometric data from the terminal and processes it as input data for running a machine learning model. Specifically, it uses TensorFlow to analyze voice tone and facial movements to evaluate the user's mental state. This process outputs mental states such as "relaxed," "lively," and "sad."

[0763] Step 3:

[0764] The server selects the most suitable audio information from a Firebase database based on the mental state assessment results. The selection process involves rule-based data processing, such as matching the assessment results and selecting "classical music" for a relaxed state. The input is the mental state assessment results, and the output is the selected audio information.

[0765] Step 4:

[0766] The server transmits the selected audio information to the terminal. The audio information is converted into real-time streaming data and prepared to be provided to the user. The input is the selected audio information, and the output is streaming data that can be played on the user's terminal.

[0767] Step 5:

[0768] Users enjoy audio information provided through their devices. Actual playback takes place through the smartphone's audio function, allowing users to obtain a musical experience suited to their own mental state. The input is streaming data, and the output is the musical experience mediated through the user's hearing.

[0769] The specific processing unit 290 transmits the result of the specific processing to the robot 414. In the robot 414, the control unit 46A causes the speaker 240 and the controlled object 443 to output the result of the specific processing. The microphone 238 acquires audio indicating user input for the result of the specific processing. The control unit 46A transmits the audio data indicating user input acquired by the microphone 238 to the data processing unit 12. In the data processing unit 12, the specific processing unit 290 acquires the audio data.

[0770] Data generation model 58 is a type of so-called generative AI (Artificial Intelligence). One example of data generation model 58 is ChatGPT (Internet Search).<URL: https: / / openai.com / blog / chatgpt> ), Gemini (Internet search) <url: https: gemini.google.com ?hl="ja">Examples of generative AI include the following. The data generation model 58 is obtained by performing deep learning on a neural network. The data generation model 58 is input with prompts containing instructions, and with inference data such as audio data representing speech, text data representing text, and image data representing images. The data generation model 58 infers from the input inference data according to the instructions indicated by the prompts, and outputs the inference results in data formats such as audio data and text data. Here, inference refers to, for example, analysis, classification, prediction, and / or summarization.

[0771] In the above embodiment, an example was given in which specific processing is performed by the data processing device 12, but the technology of this disclosure is not limited thereto, and the specific processing may also be performed by the robot 414.

[0772] Furthermore, the emotion identification model 59, acting as an emotion engine, may determine the user's emotion according to a specific mapping. Specifically, the emotion identification model 59 may determine the user's emotion according to a specific mapping, which is an emotion map (see Figure 9). Similarly, the emotion identification model 59 may also determine the robot's emotion, and the identification processing unit 290 may perform identification processing using the robot's emotion.

[0773] Figure 9 shows an emotion map 400 in which multiple emotions are mapped. In the emotion map 400, emotions are arranged in concentric circles radiating from the center. The closer to the center of the concentric circles, the more primitive the emotions are located. Further out of the concentric circles, emotions representing states and actions arising from mental states are located. Emotion is a concept that includes feelings and mental states. On the left side of the concentric circles, emotions that are generally generated from reactions occurring in the brain are located. On the right side of the concentric circles, emotions that are generally induced by situational judgment are located. Above and below the concentric circles, emotions that are generally generated from reactions occurring in the brain and induced by situational judgment are located. In addition, the emotion of "pleasure" is located on the upper side of the concentric circles, and the emotion of "displeasure" is located on the lower side. Thus, in the emotion map 400, multiple emotions are mapped based on the structure in which emotions arise, and emotions that are likely to occur simultaneously are mapped close together.

[0774] These emotions are distributed at the 3 o'clock position on the Emotion Map 400, and usually fluctuate between feelings of security and anxiety. In the right half of the Emotion Map 400, situational awareness takes precedence over internal feelings, resulting in a calm impression.

[0775] The inside of the Emotion Map 400 represents inner thoughts, while the outside represents actions. Therefore, the further you go from the outside of the Emotion Map 400, the more visible (expressed in actions) your emotions become.

[0776] Here, human emotions are based on various balances, such as posture and blood sugar levels. When these balances deviate from the ideal, it results in discomfort, and when they approach the ideal, it results in pleasure. Similarly, in robots, cars, motorcycles, etc., emotions can be created based on various balances, such as posture and battery level. When these balances deviate from the ideal, it results in discomfort, and when they approach the ideal, it results in pleasure. The emotion map can be generated, for example, based on Dr. Mitsuyoshi's emotion map (Research on a system for analyzing brain physiological signals of speech emotion recognition and emotion, Tokushima University, doctoral dissertation: https: / / ci.nii.ac.jp / naid / 500000375379). The left half of the emotion map contains emotions belonging to a region called "response," where sensation is dominant. The right half of the emotion map contains emotions belonging to a region called "situation," where situational awareness is dominant.

[0777] The emotion map defines two emotions that promote learning. One is the emotion around the middle of the negative "repentance" and "reflection" on the situation side. In other words, it is when the robot experiences negative emotions such as "I never want to feel this way again" or "I don't want to be scolded again." The other is the emotion around the positive "desire" on the reaction side. In other words, it is when the robot has positive feelings such as "I want more" or "I want to know more."

[0778] The emotion identification model 59 inputs user input into a pre-trained neural network, obtains emotion values ​​representing each emotion shown in the emotion map 400, and determines the user's emotion. This neural network is pre-trained based on multiple training data sets, which are combinations of user input and emotion values ​​representing each emotion shown in the emotion map 400. Furthermore, this neural network is trained so that emotions located close together have similar values, as shown in the emotion map 900 in Figure 10. Figure 10 shows an example where multiple emotions such as "reassured," "calm," and "confident" have similar emotion values.

[0779] The above description primarily focuses on the functions of the data processing device 12 in relation to this disclosure. However, the system related to this disclosure is not necessarily implemented on a server. The system related to this disclosure may be implemented as a general information processing system. This disclosure may be implemented, for example, as a software program that runs on a personal computer or as an application that runs on a smartphone. The method related to this disclosure may be provided to users in SaaS (Software as a Service) format.

[0780] In the above embodiment, an example was given in which a specific process is performed by a single computer 22. However, the technology of this disclosure is not limited thereto, and a distributed processing of the specific process may be performed by multiple computers, including computer 22. For example, a data generation model 58 may be provided in an external device of the data processing device 12, and the external device may generate data according to the input data.

[0781] In the above embodiment, an example was given in which the specific processing program 56 is stored in the storage 32, but the technology of this disclosure is not limited thereto. For example, the specific processing program 56 may be stored in a portable, computer-readable, non-temporary storage medium such as a USB (Universal Serial Bus) memory. The specific processing program 56 stored in the non-temporary storage medium is installed in the computer 22 of the data processing device 12. The processor 28 executes specific processing according to the specific processing program 56.

[0782] Alternatively, the specific processing program 56 may be stored in a storage device such as a server connected to the data processing device 12 via the network 54, and the specific processing program 56 may be downloaded and installed on the computer 22 in response to a request from the data processing device 12.

[0783] Furthermore, it is not necessary to store the entirety of the specific processing program 56 in a storage device such as a server connected to the data processing device 12 via the network 54, or to store the entirety of the specific processing program 56 in the storage 32; it is acceptable to store only a portion of the specific processing program 56.

[0784] The following types of processors can be used as hardware resources to perform specific processing. Examples of processors include a CPU, a general-purpose processor that functions as a hardware resource to perform specific processing by executing software, i.e., a program. Other examples of processors include dedicated electrical circuits, such as FPGAs (Field-Programmable Gate Arrays), PLDs (Programmable Logic Devices), or ASICs (Application Specific Integrated Circuits), which have circuit configurations specifically designed to perform specific processing. All of these processors have built-in or connected memory, and all of them perform specific processing by using memory.

[0785] The hardware resource that performs a specific process may consist of one of these various processors, or it may consist of a combination of two or more processors of the same or different types (for example, a combination of multiple FPGAs, or a combination of a CPU and an FPGA). Alternatively, the hardware resource that performs a specific process may consist of a single processor.

[0786] Examples of configurations using a single processor include, firstly, a configuration in which one or more CPUs and software are combined to form a single processor, and this processor functions as a hardware resource that performs a specific process. Secondly, there is a configuration using a processor that realizes the functions of the entire system, including multiple hardware resources that perform a specific process, on a single IC chip, as exemplified by SoCs (System-on-a-chip). In this way, a specific process is realized using one or more of the above types of processors as hardware resources.

[0787] Furthermore, the hardware structure of these various processors can more specifically utilize electrical circuits that combine circuit elements such as semiconductor devices. Also, the specific processing described above is merely an example. Therefore, it goes without saying that unnecessary steps can be deleted, new steps added, or the processing order rearranged, as long as it does not deviate from the main purpose.

[0788] The descriptions and illustrations presented above are detailed explanations of the technical aspects of this disclosure and are merely examples of the technical aspects. For example, the above descriptions of the structure, function, operation, and effect are examples of the structure, function, operation, and effect of the technical aspects of this disclosure. Therefore, it goes without saying that you may delete unnecessary parts, add new elements, or replace elements in the descriptions and illustrations presented above, as long as you do not deviate from the essence of the technical aspects of this disclosure. Furthermore, in order to avoid confusion and facilitate understanding of the technical aspects of this disclosure, explanations of common technical knowledge and the like that do not require special explanation to enable the implementation of the technical aspects of this disclosure have been omitted from the descriptions and illustrations presented above.

[0789] All documents, patent applications, and technical standards described herein are incorporated by reference to the same extent as if each individual document, patent application, and technical standard were specifically and individually noted as being incorporated by reference.

[0790] The following is further disclosed regarding the embodiments described above.

[0791] (Claim 1)

[0792] A device that acquires the user's biometric information,

[0793] A means of evaluating the user's mental state based on acquired biometric information,

[0794] A means for generating or selecting acoustic data suitable for the user based on the evaluation results,

[0795] A system including means for providing generated or selected acoustic data to a user.

[0796] (Claim 2)

[0797] The system according to claim 1, which generates or selects acoustic data by comparing evaluation results with information related to mental care, sports psychology, and music therapy.

[0798] (Claim 3)

[0799] The system according to claim 1, which uses heart rate data as biometric information.

[0800] "Example 1"

[0801] (Claim 1)

[0802] A means of acquiring the user's biometric parameters,

[0803] A means of evaluating the user's psychological state based on acquired biometric parameters,

[0804] A means for generating or selecting optimal acoustic information for the user based on the evaluation results,

[0805] Means for providing generated or selected acoustic information to the user,

[0806] A means of obtaining user feedback on the provided audio information,

[0807] A system that includes means for updating the generated AI model based on acquired feedback.

[0808] (Claim 2)

[0809] The system according to claim 1, which generates or selects acoustic information by comparing evaluation results with data related to psychotherapy, exercise psychology, and sound therapy.

[0810] (Claim 3)

[0811] The system according to claim 1, which uses heart rate information as a biological parameter.

[0812] "Application Example 1"

[0813] (Claim 1)

[0814] A device that acquires the user's biometric data,

[0815] A means for determining the user's psychological state based on acquired biometric data,

[0816] A means for generating or selecting acoustic information suitable for the user based on the judgment result,

[0817] Means for providing generated or selected acoustic information to the user,

[0818] A system that includes means for delivering audio information to a store environment.

[0819] (Claim 2)

[0820] The system according to claim 1, which generates or selects acoustic information by comparing the judgment result with information on mental care, exercise psychology, and sound therapy.

[0821] (Claim 3)

[0822] The system according to claim 1, which uses heart rate information as biometric data.

[0823] "Example 2 of combining an emotion engine"

[0824] (Claim 1)

[0825] A device that acquires the user's biometric information,

[0826] A means of evaluating the user's mental state by integrating acquired biometric information and voice and facial expression data,

[0827] A means of selecting and transmitting audio data suitable for the user to the terminal based on the assessed mental state,

[0828] A system that includes means for providing selected audio data to users in real time via a terminal.

[0829] (Claim 2)

[0830] The system according to claim 1, which selects acoustic data by comparing the evaluation results with information on mental care, sports psychology, and music therapy.

[0831] (Claim 3)

[0832] The system according to claim 1, which utilizes heart rate data and voice tone and facial expression data, including emotion analysis, as biometric information.

[0833] "Application example 2 of combining emotional engines"

[0834] (Claim 1)

[0835] Means for acquiring users' biometric data,

[0836] A method for evaluating the user's mental state based on acquired biometric data, voice tone, and facial movements,

[0837] A means of using a communication device to generate or select acoustic information suitable for the user based on the evaluation results and provide it to the user,

[0838] A means of analyzing voice tone and facial movements using an emotion analysis engine,

[0839] Based on the analysis results, a means of selecting acoustic information from the database and providing it to the user in real time,

[0840] ...

[0841] A system that includes this.

[0842] (Claim 2)

[0843] The system according to claim 1, which includes the step of selecting an evaluation result from a library of acoustic information and playing it back by a streaming means.

[0844] (Claim 3)

[0845] The system according to claim 1, which utilizes heart rate pattern data as biological information. [Explanation of Symbols]

[0846] 10, 210, 310, 410 Data Processing Systems 12 Data Processing Devices 14 Smart Devices 214 Smart Glasses 314 Headset-type terminal 414 Robots< / url:> < / url:> < / url:> < / url:>

Claims

1. A device that acquires the user's biometric information, A means of evaluating the user's mental state based on acquired biometric information, A means for generating or selecting acoustic data suitable for the user based on the evaluation results, A system including means for providing generated or selected acoustic data to a user.

2. The system according to claim 1, which generates or selects acoustic data by comparing evaluation results with information related to mental care, sports psychology, and music therapy.

3. The system according to claim 1, which uses heart rate data as biometric information.

Citation Information

Patent Citations

  • Persona chatbot control method and system

    JP2022180282A