system

The system addresses the challenge of accurately recognizing non-verbal emotions by acquiring and analyzing audio and video data, visualizing results, and refining its analysis through user feedback, enhancing communication effectiveness.

JP2026069054APending Publication Date: 2026-04-23SOFTBANK GROUP CORP
View PDF 1 Cites 0 Cited by

Patent Information

Authority / Receiving Office
JP · JP
Patent Type
Applications
Current Assignee / Owner
SOFTBANK GROUP CORP
Filing Date
2024-10-11
Publication Date
2026-04-23

AI Technical Summary

Technical Problem

Existing systems struggle to accurately grasp non-verbal emotions in communication environments, making it difficult to take appropriate actions based on nuanced emotional cues in business and educational settings.

Method used

A system that acquires audio and video data via a communication device, performs sentiment analysis, visualizes the results, and optimizes learning through user feedback to enhance emotional understanding.

Benefits of technology

Enables accurate recognition of non-verbal sentiment information, supporting better communication by continuously improving analysis accuracy over time.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure 2026069054000001_ABST
    Figure 2026069054000001_ABST
Patent Text Reader

Abstract

We provide the system. [Solution] Means for acquiring audio data and video data via a communication device, Processing means for performing emotion analysis based on the aforementioned audio data and video data, A means of visualizing the analysis results and presenting them to the user using an output device, A means of optimizing learning by receiving feedback information and updating the sentiment analysis model, A system that includes this.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The technology of the present disclosure relates to a system.

Background Art

[0002] Patent Document 1 discloses a persona chatbot control method performed by at least one processor, including steps of receiving a user utterance, adding the user utterance to a prompt including an instruction sentence related to an explanation of a chatbot character, encoding the prompt, and inputting the encoded prompt into a language model to generate a chatbot utterance as a response to the user utterance.

Prior Art Documents

Patent Documents

[0003]

Patent Document 1

Summary of the Invention

Problems to be Solved by the Invention

[0004] In a modern communication environment, it is difficult to accurately grasp what kind of emotions the conversation partner is feeling. For this reason, there is a problem that it is difficult to take appropriate actions based on non - verbal information in business and educational settings. In particular, it is required to accurately recognize nuances and emotions that are difficult to convey only through speech and take prompt actions accordingly.

Means for Solving the Problems

[0005] To solve this problem, the present invention provides a system that includes means for acquiring audio data and video data via a communication device. This system is equipped with processing means for performing sentiment analysis using the acquired audio data and video data. Furthermore, it includes an output device for visualizing the analysis results and presenting them to the user. In addition, by providing means for optimizing learning by receiving feedback information from the user and updating the sentiment analysis model, it is possible to accurately grasp nonverbal sentiment information and support better communication.

[0006] A "communication device" refers to a device that has the function of acquiring voice and video data and sending and receiving data over a network.

[0007] "Audio data" refers to acoustic signals acquired by sensor devices such as microphones and that can be processed in digital format.

[0008] "Video data" refers to data that represents visual information captured by video acquisition devices such as cameras in a digital format.

[0009] "Emotional analysis" refers to the process of inferring and identifying a subject's emotional state based on audio and video data.

[0010] "Processing means" refers to hardware and software components that analyze data within a system and perform desired actions.

[0011] "Visualization" refers to techniques for displaying digital data in a way that is easy for humans to understand.

[0012] An "output device" refers to a device used to display analysis results, either physically or digitally.

[0013] "Feedback information" refers to information provided by users to the system, including evaluations of analysis and operation, and requests for changes.

[0014] "Optimizing learning" refers to the process by which the system makes adjustments to improve its operation and analysis accuracy based on the feedback it receives. [Brief explanation of the drawing]

[0015] [Figure 1] This is a conceptual diagram showing an example of the configuration of a data processing system according to the first embodiment. [Figure 2] This is a conceptual diagram showing an example of the essential functions of a data processing device and a smart device according to the first embodiment. [Figure 3] This is a conceptual diagram showing an example of the configuration of a data processing system according to the second embodiment. [Figure 4] This is a conceptual diagram showing an example of the main functions of a data processing device and smart glasses according to the second embodiment. [Figure 5] This is a conceptual diagram showing an example of the configuration of a data processing system according to the third embodiment. [Figure 6] This is a conceptual diagram showing an example of the main functions of a data processing device and a headset-type terminal according to the third embodiment. [Figure 7] This is a conceptual diagram showing an example of the configuration of a data processing system according to the fourth embodiment. [Figure 8] This is a conceptual diagram showing an example of the main functions of a data processing device and a robot according to the fourth embodiment. [Figure 9] This shows an emotion map where multiple emotions are mapped. [Figure 10] This shows an emotion map where multiple emotions are mapped. [Figure 11] This is a sequence diagram showing the processing flow of the data processing system in Example 1. [Figure 12] This is a sequence diagram showing the processing flow of the data processing system in Application Example 1. [Figure 13] This is a sequence diagram showing the processing flow of the data processing system in Example 2, which incorporates an emotion engine. [Figure 14]It is a sequence diagram showing the processing flow of a data processing system in Application Example 2 when a sentiment engine is combined.

Embodiment for Carrying Out the Invention

[0016] Hereinafter, an example of an embodiment of a system according to the technology of the present disclosure will be described with reference to the accompanying drawings.

[0017] First, the terms used in the following description will be explained.

[0018] In the following embodiments, a numbered processor (hereinafter simply referred to as "processor") may be a single arithmetic unit or a combination of multiple arithmetic units. Also, the processor may be a single type of arithmetic unit or a combination of multiple types of arithmetic units. Examples of arithmetic units include a CPU (Central Processing Unit), a GPU (Graphics Processing Unit), a GPGPU (General-Purpose computing on Graphics Processing Units), an APU (Accelerated Processing Unit), and the like.

[0019] In the following embodiments, a numbered RAM (Random Access Memory) is a memory in which information is temporarily stored and is used as a work memory by the processor.

[0020] In the following embodiments, a numbered storage is one or more non-volatile storage devices that store various programs and various parameters, etc. Examples of non-volatile storage devices include flash memory (SSD (Solid State Drive)), magnetic disks (e.g., hard disks), or magnetic tapes, and the like.

[0021] In the following embodiments, the signed communication interface (I / F) is an interface that includes a communication processor and an antenna, etc. The communication interface manages communication between multiple computers. Examples of communication standards applicable to the communication interface include wireless communication standards such as 5G (5th Generation Mobile Communication System), Wi-Fi (registered trademark), or Bluetooth (registered trademark).

[0022] In the following embodiments, "A and / or B" is synonymous with "at least one of A and B." That is, "A and / or B" means that it may be A alone, or B alone, or a combination of A and B. Furthermore, in this specification, the same concept as "A and / or B" applies when expressing three or more things linked by "and / or."

[0023] [First Embodiment]

[0024] Figure 1 shows an example of the configuration of the data processing system 10 according to the first embodiment.

[0025] As shown in Figure 1, the data processing system 10 includes a data processing device 12 and a smart device 14. An example of the data processing device 12 is a server.

[0026] The data processing device 12 comprises a computer 22, a database 24, and a communication interface 26. The computer 22 is an example of a "computer" related to the technology of this disclosure. The computer 22 comprises a processor 28, RAM 30, and storage 32. The processor 28, RAM 30, and storage 32 are connected to a bus 34. The database 24 and the communication interface 26 are also connected to the bus 34. The communication interface 26 is connected to a network 54. An example of the network 54 is a WAN (Wide Area Network) and / or a LAN (Local Area Network).

[0027] The smart device 14 comprises a computer 36, a reception device 38, an output device 40, a camera 42, and a communication interface 44. The computer 36 comprises a processor 46, RAM 48, and storage 50. The processor 46, RAM 48, and storage 50 are connected to a bus 52. The reception device 38, output device 40, and camera 42 are also connected to the bus 52.

[0028] The reception device 38 is equipped with a touch panel 38A and a microphone 38B, etc., and receives user input. The touch panel 38A receives user input by detecting contact with an object (e.g., a pen or finger). The microphone 38B receives user input by detecting the user's voice. The control unit 46A transmits data indicating the user input received by the touch panel 38A and microphone 38B to the data processing device 12. In the data processing device 12, the specific processing unit 290 acquires the data indicating the user input.

[0029] The output device 40 includes a display 40A and a speaker 40B, and presents data to the user 20 by outputting the data in a form perceptible to the user 20 (e.g., audio and / or text). The display 40A displays visible information such as text and images according to instructions from the processor 46. The speaker 40B outputs audio according to instructions from the processor 46. The camera 42 is a small digital camera equipped with an optical system such as a lens, aperture, and shutter, and an image sensor such as a CMOS (Complementary Metal-Oxide-Semiconductor) image sensor or a CCD (Charge Coupled Device) image sensor.

[0030] Communication interface 44 is connected to network 54. Communication interfaces 44 and 26 are responsible for the exchange of various types of information between processor 46 and processor 28 via network 54.

[0031] Figure 2 shows an example of the main functions of the data processing device 12 and the smart device 14.

[0032] As shown in Figure 2, in the data processing device 12, specific processing is performed by the processor 28. A specific processing program 56 is stored in the storage 32. The specific processing program 56 is an example of a "program" related to the technology of this disclosure. The processor 28 reads the specific processing program 56 from the storage 32 and executes the read specific processing program 56 on the RAM 30. The specific processing is realized by the processor 28 operating as a specific processing unit 290 according to the specific processing program 56 executed on the RAM 30.

[0033] The storage 32 stores the data generation model 58 and the emotion identification model 59. The data generation model 58 and the emotion identification model 59 are used by the identification processing unit 290.

[0034] In the smart device 14, the processor 46 performs the reception output processing. The storage 50 stores the reception output program 60. The reception output program 60 is used in conjunction with a specific processing program 56 by the data processing system 10. The processor 46 reads the reception output program 60 from the storage 50 and executes the read reception output program 60 on the RAM 48. The reception output processing is realized by the processor 46 operating as a control unit 46A according to the reception output program 60 executed on the RAM 48.

[0035] Next, the specific processing performed by the specific processing unit 290 of the data processing device 12 will be described. In the following description, the data processing device 12 will be referred to as the "server" and the smart device 14 as the "terminal".

[0036] The system according to the present invention has the function of acquiring voice and video data using a specific communication device and performing emotion analysis. To achieve this, a server, a terminal, and a user must cooperate in operation.

[0037] First, the device acquires audio data from the user's surroundings using a microphone and simultaneously acquires video data using a camera. This enables real-time data collection. The collected data is sent to a server via a security protocol, and the server receives the data. Next, the server activates an advanced algorithm for sentiment analysis based on the collected audio and video data. This algorithm utilizes natural language processing and computer vision technologies to analyze the emotional state.

[0038] Specifically, the server extracts factors such as voice tone, pitch, and speed from audio data and uses this information to infer emotional states. It also analyzes facial expressions, eye movements, and cheek redness from video data to obtain additional emotional information. By integrating these results, the server then performs the final emotion identification.

[0039] The server then visualizes the analysis results and returns them to the terminal as data. The terminal intuitively displays this data on a dashboard. For example, it can simultaneously monitor the emotions of multiple participants during a meeting and inform the user of each participant's emotional state. The user reviews the dashboard and adjusts their approach to the conversation as needed.

[0040] Furthermore, the server receives feedback information provided by users and uses it to improve the analysis model. This allows the system to improve its accuracy over time, enabling more reliable sentiment analysis. This feedback loop is a crucial element in continuously optimizing the system's performance.

[0041] The following describes the processing flow.

[0042] Step 1:

[0043] The device uses a microphone to collect audio data, including ambient sounds, and simultaneously captures the user's video data with a camera. This allows for the collection of emotion-related information, such as specific utterances and changes in facial expressions.

[0044] Step 2:

[0045] The device sends collected audio and video data to the server via a secure protocol. During transmission, the data is encrypted to protect privacy.

[0046] Step 3:

[0047] To analyze the data received by the server in real time, natural language processing algorithms are applied to the audio data to extract speech features. These extracted features include speed, emphasis, and intonation.

[0048] Step 4:

[0049] The server uses computer vision algorithms to analyze video data and identify emotion-related features from the user's facial expressions and movements. For example, it can infer emotions such as joy, anger, sadness, or happiness based on smiles and eyebrow furrowing.

[0050] Step 5:

[0051] The server integrates the audio and video analysis results and uses statistical models and machine learning algorithms to comprehensively identify the user's emotional state. It then outputs a specific emotional category (e.g., joy, anxiety, surprise).

[0052] Step 6:

[0053] The server processes the sentiment analysis results as visual data and generates a dataset for the dashboard. This includes visual elements such as color coding and icons.

[0054] Step 7:

[0055] The device acquires visual data and displays it in real time on the user's dashboard. The user can observe this and take appropriate action based on the situation.

[0056] Step 8:

[0057] When a user provides feedback based on the system's analysis results, the terminal receives this information and sends it to the server. This feedback is used for the continuous training of the analysis model, contributing to improving the system's accuracy.

[0058] (Example 1)

[0059] Next, we will describe Example 1. In the following description, the data processing device 12 will be referred to as the "server," and the smart device 14 will be referred to as the "terminal."

[0060] When analyzing human emotions from audio and video data, there is a challenge in accurately identifying emotions using analysis based on a single source of information. In addition, there is a problem in that the use of feedback to effectively present analysis results to users and improve the accuracy of the system is insufficient.

[0061] The identification process performed by the identification processing unit 290 of the data processing device 12 in Example 1 is realized by the following means.

[0062] In this invention, the server includes acquisition means for acquiring audio and video signals, analysis means for analyzing emotional states based on the audio and video signals using natural language processing and image analysis techniques, and integration means for integrating the analysis results from the analysis means and identifying emotions. This enables accurate emotion identification utilizing multimodal information and improved accuracy based on user feedback.

[0063] "Audio signals" are data obtained by converting sound vibrations acquired from individuals or their environment into a format that can be electrically transmitted.

[0064] A "video signal" is data obtained by converting visual information acquired by a visual device such as a camera into a format that can be transmitted electronically.

[0065] "Natural language processing" is a technology that uses computers to process and analyze human language, and is used to recognize emotions and intentions from speech and text.

[0066] "Image analysis technology" refers to computational techniques that extract features from video data to understand the state of objects and people.

[0067] "Emotional state" refers to the emotions an individual experiences at a particular moment, and can be inferred from their tone of voice and facial expressions.

[0068] "Integration methods" refer to processes and techniques for consolidating analytical results obtained from multiple sources into a single conclusion.

[0069] An "analytical model" is a set of algorithms and computational procedures used to analyze data, which are learned to recognize specific patterns or meanings.

[0070] "Feedback" refers to evaluations and information provided by users to improve the performance of a system.

[0071] "Improving accuracy" refers to increasing the performance of a system, specifically how closely its output matches real-world conditions.

[0072] This invention relates to a system that identifies emotional states by analyzing audio and video signals and provides the user with the analysis results. To implement the invention, a microphone for acquiring audio signals, a camera for acquiring video signals, a server for data processing, and a terminal for presenting these to the user are required.

[0073] terminal

[0074] The device acquires audio signals using microphones placed around the user and video signals using a camera. This data is transmitted to a server in real time. In particular, the device performs acoustic processing to remove ambient noise and allow the user to focus on the conversation. The camera also uses a face detection algorithm to facilitate analysis of the user's facial expressions.

[0075] server

[0076] The server receives audio and video signals transmitted from the terminal. Audio signals are analyzed using natural language processing techniques to extract features such as tone, pitch, and speed. Video signals are analyzed using image analysis techniques to extract facial movements and changes in expression. These processes utilize algorithms powered by generative AI models to estimate emotional states. Based on the analysis results, the server identifies each user's emotions and converts them into a visualizeable data format.

[0077] User

[0078] Users visually review the analysis results sent back from the server via a dashboard on their device. This allows them to understand the emotional state of other participants during the meeting. Users provide feedback to the server via their device, and this feedback is used to improve the analysis model.

[0079] For example, if a participant is feeling anxious during an online meeting, their name and emotional state will be displayed on the device's dashboard. A prompt such as, "Analyze the emotional states of current meeting participants in real time and display any participants who appear anxious," allows for a quick response.

[0080] The flow of the specific processing in Example 1 will be explained using Figure 11.

[0081] Step 1:

[0082] The device acquires audio signals through microphones placed around the user and collects video signals using a camera. The inputs are ambient audio and visual data. The audio signal undergoes acoustic processing to remove unwanted ambient noise. The video signal uses a face detection algorithm to pinpoint the location of faces. The output of this processing is clean audio data and video data with faces detected.

[0083] Step 2:

[0084] The terminal transmits the processed audio and video data to the server via a security protocol. Here, the data is encrypted before transmission to ensure privacy and security. The input is the output data from the previous step, and the output is encrypted audio and video data.

[0085] Step 3:

[0086] The server receives the transmitted data and applies natural language processing techniques to the audio signal. Specifically, it extracts tone, pitch, and speed, and uses these features to estimate emotion. The input is encrypted audio data, and the output is an emotion estimation score.

[0087] Step 4:

[0088] The server analyzes video data and detects features such as facial expressions and movements. Using advanced computer vision technology, it quantifies various facial expressions and estimates possible emotional states. The input is encrypted video data, and the output is an emotion estimation score based on facial expressions.

[0089] Step 5:

[0090] The server integrates estimated scores obtained from audio and video and performs final emotion identification using a generative AI model. The resulting overall emotion score is stored in a database and converted into a data format for visualization. The input is the estimated score for each emotion, and the output is a visualized emotion identification result.

[0091] Step 6:

[0092] The server sends emotion recognition data to the terminal, which displays it on a dashboard. Users refer to the information displayed on the dashboard to adjust their approach to meetings and conversations. The input is a visualized emotion recognition result, and the output is an intuitive interface display presented to the user.

[0093] Step 7:

[0094] The user inputs feedback on the system's sentiment analysis results into a terminal. The server receives the feedback and uses it to improve the analysis model in the generated AI model. The input is the user's feedback, and the output is the updated analysis model.

[0095] (Application Example 1)

[0096] Next, we will explain Application Example 1. In the following explanation, the data processing device 12 will be referred to as the "server," and the smart device 14 will be referred to as the "terminal."

[0097] In modern customer service, understanding individual customer emotions in real time and improving service quality is becoming increasingly important. However, traditional methods make it difficult to accurately interpret emotions from customers' facial expressions and voices, hindering service improvement. To solve this problem, a system is needed that utilizes acoustic and visual information to analyze individual customer emotions in real time and respond quickly.

[0098] The specific processing performed by the specific processing unit 290 of the data processing device 12 in Application Example 1 is realized by the following means.

[0099] In this invention, the server includes means for acquiring acoustic and video information via communication equipment, computation means for analyzing emotions based on the acoustic and video information, and means for visualizing the analysis results and providing them to the individual using a display device. This enables service staff to immediately grasp the customer's emotions and provide appropriate service.

[0100] "Communication equipment" refers to devices used to acquire and transmit or receive audio and video information externally.

[0101] "Acoustic information" refers to data related to speech and tone, including characteristics such as the tone and speed of the speaker's voice.

[0102] "Visual information" refers to data related to images or videos that is used to analyze facial expressions and movements.

[0103] "Emotion" refers to an expression of an individual's emotional state or psychological reaction.

[0104] The "computation means" is a function that analyzes the acquired acoustic and visual information and performs processing to identify emotions.

[0105] A "display device" refers to a screen device or glasses-type information device that visually presents analyzed emotions to an individual.

[0106] "Visualization" is the process of representing data visually and providing it in an easily understandable format.

[0107] This invention aims to realize a system that utilizes communication equipment to collect acoustic and visual information and perform emotional analysis in real time. The user wears glasses-type information devices to acquire acoustic and visual information obtained from interactions with customers. These devices use built-in microphones and cameras to collect ambient sounds and customer facial expressions in real time. The collected data is transmitted to a server via Bluetooth or Wi-Fi.

[0108] The server performs emotion analysis using TENSORFLOW® and OpenCV based on acquired acoustic and video information. From the acoustic information, it analyzes the tone, speed, and pitch of speech, while from the video information, it evaluates facial features and movements. The resulting emotion data is then comprehensively evaluated using computer vision and natural language processing technologies.

[0109] The analyzed emotional data is visualized in real time on the user's glasses-type information device display. As a result, the user can quickly grasp the customer's emotional state and provide appropriate service. Furthermore, feedback information is sent to the server, and the accuracy is continuously improved by refining the emotional analysis algorithm. This feedback is obtained through manual input and automatic detection, contributing to the improvement of the system's performance.

[0110] As a concrete example, when a store clerk is introducing a product to a customer, if anxiety is detected from the customer's facial expression and voice, the display on the glasses will show instructions such as, "The customer is feeling anxious. Please provide reassuring information." This information is then analyzed by a generative AI model using prompts such as, "Analyze how customer A feels about the new product. Audio and video samples are available."

[0111] Through the above, it is possible to improve customer satisfaction in customer service operations.

[0112] The flow of a specific process in Application Example 1 will be explained using Figure 12.

[0113] Step 1:

[0114] The device uses a microphone and camera to acquire ambient acoustic and visual information. The input here is real-time collected audio and image data. The device temporarily stores this data and prepares it for transmission to a server via Bluetooth or Wi-Fi.

[0115] Step 2:

[0116] The server receives audio and video information transmitted from the terminal. Based on this input data, it performs data preprocessing. Audio data is subjected to noise reduction filtering and processed to make the audio signal clearer. Video data is filtered to extract features important for image recognition. This results in a clean dataset ready for analysis.

[0117] Step 3:

[0118] The server performs emotion analysis using TensorFlow and OpenCV with pre-processed data. Clean audio and video datasets are used as input. Tone, pitch, and speed are detected from the audio data, and facial features and movements are detected from the video data, and data calculations are performed. The analyzed emotion data is output, clarifying the user's emotional state.

[0119] Step 4:

[0120] The analysis results are visualized by the server and sent back to the terminal in real time. The input is analyzed emotional data, and visualized information is generated based on it. The user can check the analysis results on the display of a glasses-type information device. This display shows icons and text messages indicating the customer's emotional state, and prompts the user to take action.

[0121] Step 5:

[0122] User feedback is sent to the server via the terminal. Based on the feedback received as input, the server updates its emotion analysis algorithm to improve its accuracy. The processing of feedback information is supported by specific generative AI models and prompt statements. This allows the system to continuously provide more sophisticated analysis results.

[0123] Furthermore, an emotion engine that estimates the user's emotions may be incorporated. That is, the identification processing unit 290 may use the emotion identification model 59 to estimate the user's emotions and perform identification processing using the user's emotions.

[0124] This invention realizes a system that incorporates an emotion engine for analyzing audio and video data and recognizing the user's emotions. The system consists of a server, a terminal, and the emotion engine, which work together to analyze the user's emotional state in real time.

[0125] First, the device records the user's voice using a microphone device and simultaneously acquires video data using a camera device. This audio and video data is then transmitted to the server, timestamped as it goes. The server processes this data using an emotion engine.

[0126] The emotion engine extracts features such as voice intonation and speed by applying natural language processing techniques to audio data. Furthermore, it uses computer vision technology to analyze changes in facial expressions and eye movements from video data. Based on this information, the emotion engine estimates multiple emotional states of the user.

[0127] As a concrete example of its application, consider communication in a confined space during a meeting. The terminal records the voice and facial expressions of meeting participants and immediately transmits them to the server. The emotion engine evaluates when participants are feeling joy or anxiety and visualizes this as an analysis result. The server displays the analysis results on the terminal's dashboard and presents them to the user.

[0128] Users can refer to this dashboard to adjust the meeting's progress and facilitate smoother dialogue. Furthermore, incorporating user feedback into the emotion engine optimizes model training and further improves prediction accuracy. By repeating this process, the system can evolve continuously over the long term, increasing its practical usability.

[0129] The following describes the processing flow.

[0130] Step 1:

[0131] The device uses a microphone and camera to collect the user's voice and video data. Voice data includes the speaker's tone, intonation, and speed. Video data includes visual information such as facial expressions and eye movements.

[0132] Step 2:

[0133] The device transmits collected audio and video data to the server in real time. The data is sent in a stream format to minimize latency.

[0134] Step 3:

[0135] The server receives the transmitted data and analyzes it using an emotion engine. For audio data, natural language processing techniques are applied to extract voice features and identify parameters that provide clues to emotion.

[0136] Step 4:

[0137] The server analyzes the video data using computer vision technology to detect changes in facial expressions and gaze. This reveals nonverbal emotional elements derived from the video.

[0138] Step 5:

[0139] The server's emotion engine integrates audio and video data and utilizes multiple emotion models to infer the overall emotional state. For example, it may detect "interest" and "anxiety" simultaneously.

[0140] Step 6:

[0141] The server visualizes the analysis results and generates an intuitive dashboard dataset. This dataset includes color-coded graphs and sentiment icons.

[0142] Step 7:

[0143] The device receives the dataset and displays it on the user's dashboard. This allows the user to understand their emotional state in real time.

[0144] Step 8:

[0145] Users can adjust their responses or change the direction of discussions based on the information on the dashboard. Furthermore, the device sends feedback received from users to the server, which is then incorporated into the emotion engine's machine learning model.

[0146] Step 9:

[0147] The server analyzes the feedback information and updates the emotion engine to improve analysis accuracy. Through this process, the system accumulates insights over time and becomes able to recognize user emotions more accurately.

[0148] (Example 2)

[0149] Next, we will describe Example 2. In the following description, the data processing device 12 will be referred to as the "server" and the smart device 14 as the "terminal".

[0150] In recent years, with the advancement of information and communication technologies using voice and video, accurately understanding users' emotions in real time has become crucial. However, conventional systems often analyze voice and video information separately, which limits the accuracy of emotional state recognition. Furthermore, there is a lack of mechanisms to effectively utilize user feedback to improve the accuracy of emotion analysis models. Therefore, there is a need for a system that can analyze emotions with high accuracy and continuously optimize learning.

[0151] The identification process performed by the identification processing unit 290 of the data processing device 12 in Example 2 is realized by the following means.

[0152] In this invention, the server includes means for collecting user acoustic and video information via communication means, means for applying natural language processing techniques to extract features based on the acoustic information, and means for applying computer vision techniques to analyze facial expressions and gaze based on the video information. This enables integrated analysis of audio and video data, allowing for more accurate estimation of the user's emotional state. Furthermore, by visualizing the analysis results and presenting them to the user, feedback can be received to optimize the learning of the emotion analysis model.

[0153] "Communication means" refers to a device or mechanism for collecting audio and video information from a user and transferring it to a server.

[0154] "Acoustic information" refers to data obtained from the user's voice and is used to estimate their emotional state.

[0155] "Visual information" refers to data that records changes in the user's face and facial expressions, and is used to estimate their visual emotional state.

[0156] "Natural language processing technology" refers to techniques for converting acoustic information into text and extracting features such as intonation and speed.

[0157] "Computer vision technology" refers to techniques that analyze video information to detect facial expressions and eye movements.

[0158] An "emotion analysis model" is an algorithm or data model that estimates a user's emotional state based on acoustic and visual information.

[0159] "Visualization" is the process of displaying analysis results in the form of graphs, charts, and other formats to allow users to understand them intuitively.

[0160] "Feedback" refers to opinions and evaluations obtained from users, which are used to improve the accuracy of the model and optimize its learning process.

[0161] The system for implementing this invention mainly consists of a terminal, a server, and an emotion analysis engine. This system analyzes the user's emotional state using acoustic and visual information.

[0162] First, the device uses a microphone device and a camera device to simultaneously collect the user's voice and video. The hardware used includes high-precision microphone devices (e.g., typical audio recording devices) and high-resolution cameras (e.g., typical video recording devices). The audio and video information is time-stamped and transmitted to the server in real-time or near real-time.

[0163] The server performs advanced processing on the received data. Natural language processing techniques are used for acoustic information, and common acoustic analysis software such as Google® Cloud Speech-to-Text API is used to convert speech into text data, further extracting features such as intonation and speed. This analysis provides a means to estimate the user's emotional state.

[0164] Furthermore, the server applies computer vision technology to the video information, using common video analysis software such as OpenCV and TensorFlow to analyze the user's facial expressions and eye movements. This allows for the acquisition of emotional information that can be gleaned from the video.

[0165] The server integrates features obtained from audio and video and estimates the user's emotional state based on an emotion analysis model. The estimation results are visualized and displayed on the device's dashboard in the form of graphs and charts. This allows the user to intuitively understand the analysis results.

[0166] Furthermore, users contribute to optimizing the sentiment analysis model by providing feedback. This feedback information is fed into a continuous learning process, making it possible to improve prediction accuracy.

[0167] A concrete example is the ability to monitor the emotional state of each participant in a meeting in real time and adjust the meeting's progress as needed. A possible prompt in this scenario might be, "Please tell me how to assess participants' emotional states in real time and ensure the meeting runs smoothly." This system would enable better communication and decision-making.

[0168] The flow of the specific processing in Example 2 will be explained using Figure 13.

[0169] Step 1:

[0170] The device records the user's voice using a microphone device and acquires video using a camera device. The input is real-time audio and video. This data is timestamped and synchronized. The output is timestamped audio and video data, ready to be sent to the server immediately.

[0171] Step 2:

[0172] The terminal transmits the collected data to the server via a high-speed communication network. The input is the time-stamped audio and video data collected in step 1. The output is the audio and video data received by the server. This process enables real-time analysis of the data.

[0173] Step 3:

[0174] The server processes the audio data using acoustic analysis software such as the Google Cloud Speech-to-Text API. The input is the acoustic information sent to the server. Specifically, the audio is converted into text, and features related to intonation and speed are extracted. The output of this process is a feature vector for estimating the user's emotional state.

[0175] Step 4:

[0176] The server processes video data using tools such as OpenCV and TensorFlow. The input is video information sent to the server. Specifically, it analyzes facial expressions and eye movements from the video and extracts feature information obtained from this analysis. The output is a dataset containing this aggregated feature information.

[0177] Step 5:

[0178] The server integrates features obtained from audio and video and estimates the user's emotional state using an emotion analysis model. The input is the output from steps 3 and 4. The emotion analysis model utilizes a generative AI model to calculate the probability of various emotions. The output is the estimated result indicating the user's emotional state.

[0179] Step 6:

[0180] The server visualizes the estimation results and displays them on the terminal's dashboard in the form of a graph or chart. The input is the estimated emotional state obtained in step 5. The output is the analysis results provided in a visually understandable format for the user. This information serves as foundational data for the user to intuitively understand the analysis results and provide appropriate feedback.

[0181] Step 7:

[0182] The user reviews the analysis results presented on the dashboard and provides feedback to the system. The input is the analysis results presented to the user. The output is feedback data to improve the accuracy of the sentiment analysis model. This facilitates the continuous learning and optimization of the model.

[0183] (Application Example 2)

[0184] Next, we will explain application example 2. In the following explanation, the data processing device 12 will be referred to as a "server" and the smart device 14 as a "terminal".

[0185] In recent years, brick-and-mortar stores have been required to understand customer satisfaction in real time and provide services accordingly. However, traditional methods have the problem of not being able to immediately perceive customers' emotions. There is a need to provide an effective means for employees to accurately and quickly evaluate customers' facial expressions and voices and directly link that evaluation to service improvement.

[0186] The specific processing performed by the specific processing unit 290 of the data processing device 12 in Application Example 2 is realized by the following means.

[0187] In this invention, the server includes means for acquiring audio and video data via a communication device, processing means for performing sentiment analysis based on the audio and video data, and means for visualizing the analysis results and presenting them to the user using an output device. This enables real-time analysis of customer sentiment, allowing employees to immediately provide services based on that analysis.

[0188] A "communication device" is a device used to acquire audio and video data and transfer it to a server.

[0189] "Voice data" refers to digital information about human speech collected using voice acquisition devices such as microphones.

[0190] "Video data" refers to digital data based on visual information, acquired through video acquisition devices such as cameras.

[0191] "Emotional analysis" is a process that estimates a user's emotional state using audio and video data.

[0192] "Processing means" refers to means consisting of software or hardware for performing sentiment analysis and visualizing the data.

[0193] An "output device" is a device such as a display that visually presents the analysis results from the server to the user.

[0194] "A means of analyzing customer emotions in real time from facial expressions and voice and presenting them to workers" refers to a technology and means for instantly evaluating the emotional state of customers in physical stores and providing feedback on the results to workers.

[0195] "Feedback information" refers to evaluations and comments from users regarding sentiment analysis, and is data used to improve the analysis model.

[0196] In implementing this invention, a system will be constructed to perform sentiment analysis in customer-employee interactions in physical stores. This will be achieved using a server, terminals, and a sentiment analysis engine.

[0197] The server receives audio and video data from terminals via a communication device. Terminals, such as smart glasses or smartphones, have built-in microphones and cameras to capture customer audio and video in real time. After receiving the data, the server processes it using an emotion analysis engine. For audio data, natural language processing techniques are applied to analyze properties such as intonation and speed, and for video data, computer vision techniques are used to capture changes in facial expressions and gaze.

[0198] The data is first processed by a Python-based program. Audio data has its features extracted using the Librosa library, and video data is analyzed using OpenCV and a TensorFlow machine learning model. The analysis results are then visually displayed using an output device, for example, on the display of smart glasses worn by employees.

[0199] As a concrete example of this system, when a customer enters a store, the staff can instantly analyze whether the customer is smiling when they greet them with "Welcome." If a positive emotion is detected from the customer's expression, a message such as "Satisfied" will appear on the staff member's glasses, allowing for immediate improvement in the quality of service.

[0200] An example of a prompt message would be, "Please tell me the steps to develop an application that analyzes customer facial expressions and voice in real time and provides emotional feedback on a smart device."

[0201] The flow of a specific process in Application Example 2 will be explained using Figure 14.

[0202] Step 1:

[0203] The device uses smart glasses or a smartphone to record audio data with a microphone and capture video data with a camera. This provides instant digital information about the customer's voice tone and facial expressions.

[0204] Step 2:

[0205] The terminal transfers the acquired audio and video data to the server in real time. During this process, the data is time-stamped and synchronized. The input data transmitted consists of recorded audio and video files.

[0206] Step 3:

[0207] The server analyzes the received audio data using the Librosa library. This process extracts features such as intonation and speed from the speech and generates feature data based on these features. The audio waveform and spectral information are processed, and numerical data suggesting emotional state is output.

[0208] Step 4:

[0209] The server analyzes video data using the OpenCV library. Specifically, it recognizes facial expressions and eye movements, and converts this visual information into numerical data. It detects facial landmarks from video frames and identifies features that indicate emotion. The output is a parameter set indicating emotion.

[0210] Step 5:

[0211] The emotion analysis engine on the server integrates audio feature data and video parameter sets, and performs analysis using a machine learning model. TensorFlow is used to estimate the customer's emotional state through the integrated data. As a result of this integration, an emotion label and its confidence level are output.

[0212] Step 6:

[0213] The user (store staff) visually checks the analysis results received from the server on the display of smart glasses. This allows them to understand the customer's emotional state in real time. The output displays a specific emotional state (e.g., "satisfied," "dissatisfied").

[0214] The specific processing unit 290 transmits the result of the specific processing to the smart device 14. In the smart device 14, the control unit 46A causes the output device 40 to output the result of the specific processing. The microphone 38B acquires audio indicating user input for the result of the specific processing. The control unit 46A transmits the audio data indicating user input acquired by the microphone 38B to the data processing device 12. In the data processing device 12, the specific processing unit 290 acquires the audio data.

[0215] Data generation model 58 is a so-called generative AI (Artificial Intelligence). An example of data generation model 58 is ChatGPT (registered trademark) (Internet search).<URL: https: / / openai.com / blog / chatgpt> ), Gemini (registered trademark) (Internet search) <url: https: gemini.google.com ?hl="ja">Examples of generative AI include the following. The data generation model 58 is obtained by performing deep learning on a neural network. The data generation model 58 is input with prompts containing instructions, and with inference data such as audio data representing speech, text data representing text, and image data representing images. The data generation model 58 infers from the input inference data according to the instructions indicated by the prompts, and outputs the inference results in data formats such as audio data and text data. Here, inference refers to, for example, analysis, classification, prediction, and / or summarization.

[0216] In the above embodiment, an example was given in which specific processing is performed by the data processing device 12, but the technology of this disclosure is not limited thereto, and the specific processing may also be performed by the smart device 14.

[0217] [Second Embodiment]

[0218] Figure 3 shows an example of the configuration of the data processing system 210 according to the second embodiment.

[0219] As shown in Figure 3, the data processing system 210 includes a data processing device 12 and smart glasses 214. An example of the data processing device 12 is a server.

[0220] The data processing device 12 comprises a computer 22, a database 24, and a communication interface 26. The computer 22 is an example of a "computer" related to the technology of this disclosure. The computer 22 comprises a processor 28, RAM 30, and storage 32. The processor 28, RAM 30, and storage 32 are connected to a bus 34. The database 24 and the communication interface 26 are also connected to the bus 34. The communication interface 26 is connected to a network 54. An example of the network 54 is a WAN (Wide Area Network) and / or a LAN (Local Area Network).

[0221] The smart glasses 214 include a computer 36, a microphone 238, a speaker 240, a camera 42, and a communication interface 44. The computer 36 includes a processor 46, RAM 48, and storage 50. The processor 46, RAM 48, and storage 50 are connected to a bus 52. The microphone 238, speaker 240, and camera 42 are also connected to the bus 52.

[0222] The microphone 238 receives voice signals from the user 20 and receives instructions from the user 20. The microphone 238 captures the voice signals from the user 20, converts the captured voice into audio data, and outputs it to the processor 46. The speaker 240 outputs audio according to the instructions from the processor 46.

[0223] Camera 42 is a small digital camera equipped with an optical system including a lens, aperture, and shutter, and an image sensor such as a CMOS (Complementary Metal-Oxide-Semiconductor) image sensor or a CCD (Charge Coupled Device) image sensor, and captures images of the area around the user 20 (for example, an imaging range defined by a field of view equivalent to the width of a typical healthy person's field of vision).

[0224] Communication interface 44 is connected to network 54. Communication interfaces 44 and 26 are responsible for the exchange of various information between processor 46 and processor 28 via network 54. The exchange of various information between processor 46 and processor 28 using communication interfaces 44 and 26 is performed in a secure manner.

[0225] Figure 4 shows an example of the main functions of the data processing device 12 and the smart glasses 214. As shown in Figure 4, the data processing device 12 performs specific processing using the processor 28. The storage 32 stores the specific processing program 56.

[0226] The specific processing program 56 is an example of a "program" relating to the technology of this disclosure. The processor 28 reads the specific processing program 56 from the storage 32 and executes the read specific processing program 56 on the RAM 30. The specific processing is realized by the processor 28 operating as a specific processing unit 290 in accordance with the specific processing program 56 executed on the RAM 30.

[0227] The storage 32 stores the data generation model 58 and the emotion identification model 59. The data generation model 58 and the emotion identification model 59 are used by the identification processing unit 290.

[0228] In the smart glasses 214, the processor 46 performs the reception output processing. The storage 50 stores the reception output program 60. The processor 46 reads the reception output program 60 from the storage 50 and executes the read reception output program 60 on the RAM 48. The reception output processing is realized by the processor 46 operating as a control unit 46A according to the reception output program 60 executed on the RAM 48.

[0229] Next, the identification processing performed by the identification processing unit 290 of the data processing device 12 will be described. In the following description, the data processing device 12 will be referred to as the "server" and the smart glasses 214 will be referred to as the "terminal".

[0230] The system according to the present invention has the function of acquiring voice and video data using a specific communication device and performing emotion analysis. To achieve this, a server, a terminal, and a user must cooperate in operation.

[0231] First, the device acquires audio data from the user's surroundings using a microphone and simultaneously acquires video data using a camera. This enables real-time data collection. The collected data is sent to a server via a security protocol, and the server receives the data. Next, the server activates an advanced algorithm for sentiment analysis based on the collected audio and video data. This algorithm utilizes natural language processing and computer vision technologies to analyze the emotional state.

[0232] Specifically, the server extracts factors such as voice tone, pitch, and speed from audio data and uses this information to infer emotional states. It also analyzes facial expressions, eye movements, and cheek redness from video data to obtain additional emotional information. By integrating these results, the server then performs the final emotion identification.

[0233] The server then visualizes the analysis results and returns them to the terminal as data. The terminal intuitively displays this data on a dashboard. For example, it can simultaneously monitor the emotions of multiple participants during a meeting and inform the user of each participant's emotional state. The user reviews the dashboard and adjusts their approach to the conversation as needed.

[0234] Furthermore, the server receives feedback information provided by users and uses it to improve the analysis model. This allows the system to improve its accuracy over time, enabling more reliable sentiment analysis. This feedback loop is a crucial element in continuously optimizing the system's performance.

[0235] The following describes the processing flow.

[0236] Step 1:

[0237] The device uses a microphone to collect audio data, including ambient sounds, and simultaneously captures the user's video data with a camera. This allows for the collection of emotion-related information, such as specific utterances and changes in facial expressions.

[0238] Step 2:

[0239] The device sends collected audio and video data to the server via a secure protocol. During transmission, the data is encrypted to protect privacy.

[0240] Step 3:

[0241] To analyze the data received by the server in real time, natural language processing algorithms are applied to the audio data to extract speech features. These extracted features include speed, emphasis, and intonation.

[0242] Step 4:

[0243] The server uses computer vision algorithms to analyze video data and identify emotion-related features from the user's facial expressions and movements. For example, it can infer emotions such as joy, anger, sadness, or happiness based on smiles and eyebrow furrowing.

[0244] Step 5:

[0245] The server integrates the audio and video analysis results and uses statistical models and machine learning algorithms to comprehensively identify the user's emotional state. It then outputs a specific emotional category (e.g., joy, anxiety, surprise).

[0246] Step 6:

[0247] The server processes the sentiment analysis results as visual data and generates a dataset for the dashboard. This includes visual elements such as color coding and icons.

[0248] Step 7:

[0249] The device acquires visual data and displays it in real time on the user's dashboard. The user can observe this and take appropriate action based on the situation.

[0250] Step 8:

[0251] When a user provides feedback based on the system's analysis results, the terminal receives this information and sends it to the server. This feedback is used for the continuous training of the analysis model, contributing to improving the system's accuracy.

[0252] (Example 1)

[0253] Next, we will describe Example 1. In the following description, the data processing device 12 will be referred to as the "server," and the smart glasses 214 will be referred to as the "terminal."

[0254] When analyzing human emotions from audio and video data, there is a challenge in accurately identifying emotions using analysis based on a single source of information. In addition, there is a problem in that the use of feedback to effectively present analysis results to users and improve the accuracy of the system is insufficient.

[0255] The identification process performed by the identification processing unit 290 of the data processing device 12 in Example 1 is realized by the following means.

[0256] In this invention, the server includes acquisition means for acquiring audio and video signals, analysis means for analyzing emotional states based on the audio and video signals using natural language processing and image analysis techniques, and integration means for integrating the analysis results from the analysis means and identifying emotions. This enables accurate emotion identification utilizing multimodal information and improved accuracy based on user feedback.

[0257] "Audio signals" are data obtained by converting sound vibrations acquired from individuals or their environment into a format that can be electrically transmitted.

[0258] A "video signal" is data obtained by converting visual information acquired by a visual device such as a camera into a format that can be transmitted electronically.

[0259] "Natural language processing" is a technology that uses computers to process and analyze human language, and is used to recognize emotions and intentions from speech and text.

[0260] "Image analysis technology" refers to computational techniques that extract features from video data to understand the state of objects and people.

[0261] "Emotional state" refers to the emotions an individual experiences at a particular moment, and can be inferred from their tone of voice and facial expressions.

[0262] "Integration methods" refer to processes and techniques for consolidating analytical results obtained from multiple sources into a single conclusion.

[0263] An "analytical model" is a set of algorithms and computational procedures used to analyze data, which are learned to recognize specific patterns or meanings.

[0264] "Feedback" refers to evaluations and information provided by users to improve the performance of a system.

[0265] "Improving accuracy" refers to increasing the performance of a system, specifically how closely its output matches real-world conditions.

[0266] This invention relates to a system that identifies emotional states by analyzing audio and video signals and provides the user with the analysis results. To implement the invention, a microphone for acquiring audio signals, a camera for acquiring video signals, a server for data processing, and a terminal for presenting these to the user are required.

[0267] terminal

[0268] The device acquires audio signals using microphones placed around the user and video signals using a camera. This data is transmitted to a server in real time. In particular, the device performs acoustic processing to remove ambient noise and allow the user to focus on the conversation. The camera also uses a face detection algorithm to facilitate analysis of the user's facial expressions.

[0269] server

[0270] The server receives audio and video signals transmitted from the terminal. Audio signals are analyzed using natural language processing techniques to extract features such as tone, pitch, and speed. Video signals are analyzed using image analysis techniques to extract facial movements and changes in expression. These processes utilize algorithms powered by generative AI models to estimate emotional states. Based on the analysis results, the server identifies each user's emotions and converts them into a visualizeable data format.

[0271] User

[0272] Users visually review the analysis results sent back from the server via a dashboard on their device. This allows them to understand the emotional state of other participants during the meeting. Users provide feedback to the server via their device, and this feedback is used to improve the analysis model.

[0273] For example, if a participant is feeling anxious during an online meeting, their name and emotional state will be displayed on the device's dashboard. A prompt such as, "Analyze the emotional states of current meeting participants in real time and display any participants who appear anxious," allows for a quick response.

[0274] The flow of the specific processing in Example 1 will be explained using Figure 11.

[0275] Step 1:

[0276] The terminal acquires an audio signal through a microphone installed around the user and collects a video signal using a camera. The inputs are the surrounding audio and visual data. The audio signal undergoes acoustic processing to remove unnecessary ambient noise. The video signal uses a face detection algorithm to identify the position of the face. The output of this process is clean audio data and video data with detected faces.

[0277] Step 2:

[0278] The terminal transmits the processed audio and video data to the server via a security protocol. Here, the data is encrypted before transmission to ensure privacy and security. The input is the output data from the previous step, and the output is encrypted audio data and video data.

[0279] Step 3:

[0280] The server receives the transmitted data and applies natural language processing technology to the audio signal. Specifically, it extracts tone, pitch, and speed, and performs emotion estimation based on these features. The input is the encrypted audio data, and the output is the emotion estimation score.

[0281] Step 4:

[0282] The server analyzes the video data and detects features such as facial expressions and movements. It utilizes computer vision technology to quantify various expressions and infer possible emotional states. The input is the encrypted video data, and the output is the emotion estimation score based on expressions.

[0283] Step 5:

[0284] The server integrates the estimation scores obtained from audio and video, and performs final emotion recognition using a generated AI model. The resulting comprehensive emotion score is saved in a database and converted into a visualization data format. The input is the estimation score of each emotion, and the output is the emotion recognition result that can be visualized.

[0285] Step 6:

[0286] The server sends the data of the emotion recognition result to the terminal, and the terminal displays it on the dashboard. The user refers to the information displayed on the dashboard to adjust the approach of the meeting or interaction. The input is the visualizable emotion recognition result, and the output is the intuitive interface display presented to the user.

[0287] Step 7:

[0288] The user inputs feedback on the emotion analysis result of the system into the terminal. The server receives the feedback and uses it to improve the analysis model in the generated AI model. The input is the feedback from the user, and the output is the updated analysis model.

[0289] (Application Example 1)

[0290] Next, Application Example 1 will be described. In the following description, the data processing device 12 is referred to as the "server", and the smart glasses 214 are referred to as the "terminal".

[0291] In modern customer service operations, it is becoming increasingly important to grasp the emotions of individual customers in real time and improve the quality of service. However, with conventional methods, it is difficult to appropriately read emotions from customers' expressions and voices, which has hindered service improvement. To solve this problem, a system that utilizes acoustic information and video information to analyze the emotions of individual customers in real time and can respond quickly is necessary.

[0292] The specific processing by the specific processing unit 290 of the data processing device 12 in Application Example 1 is realized by the following means.

[0293] In this invention, the server includes means for acquiring acoustic and video information via communication equipment, computation means for analyzing emotions based on the acoustic and video information, and means for visualizing the analysis results and providing them to the individual using a display device. This enables service staff to immediately grasp the customer's emotions and provide appropriate service.

[0294] "Communication equipment" refers to devices used to acquire and transmit or receive audio and video information externally.

[0295] "Acoustic information" refers to data related to speech and tone, including characteristics such as the tone and speed of the speaker's voice.

[0296] "Visual information" refers to data related to images or videos that is used to analyze facial expressions and movements.

[0297] "Emotion" refers to an expression of an individual's emotional state or psychological reaction.

[0298] The "computation means" is a function that analyzes the acquired acoustic and visual information and performs processing to identify emotions.

[0299] A "display device" refers to a screen device or glasses-type information device that visually presents analyzed emotions to an individual.

[0300] "Visualization" is the process of representing data visually and providing it in an easily understandable format.

[0301] This invention aims to realize a system that utilizes communication equipment to collect acoustic and visual information and perform emotional analysis in real time. The user wears glasses-type information devices to acquire acoustic and visual information obtained from interactions with customers. These devices use built-in microphones and cameras to collect ambient sounds and customer facial expressions in real time. The collected data is transmitted to a server via Bluetooth or Wi-Fi.

[0302] Based on the acquired acoustic information and video information, the server performs emotion analysis using TensorFlow and OpenCV. From the acoustic information, the tone, speed, and pitch of the voice are analyzed, and from the video information, the facial features and movements are evaluated. The emotion data obtained as a result of the analysis is comprehensively judged by applying computer vision technology and natural language processing technology.

[0303] The analyzed emotion data is visualized in real time on the display of the user's glasses-type information device. As a result, the user can quickly grasp the emotional state of the customer and implement appropriate service responses. Furthermore, by sending feedback information to the server and continuously improving the emotion analysis algorithm, the accuracy is improved. This feedback is obtained through manual input or automatic detection and contributes to the improvement of the system's performance.

[0304] As a specific example, when a store clerk is introducing a product to a certain customer and uneasiness is perceived from the customer's expression and voice, an instruction such as "The customer is feeling uneasy. Please provide reassuring information" is displayed on the glasses display. This information is analyzed in the generative AI model using a prompt sentence such as "Analyze what kind of emotions customer A has towards the new product. There are audio samples and video samples."

[0305] Through the above, it is possible to improve customer satisfaction in customer service operations.

[0306] The flow of the specific process in Application Example 1 will be described using FIG. 12.

[0307] Step 1:

[0308] The device uses a microphone and camera to acquire ambient acoustic and visual information. The input here is real-time collected audio and image data. The device temporarily stores this data and prepares it for transmission to a server via Bluetooth or Wi-Fi.

[0309] Step 2:

[0310] The server receives audio and video information transmitted from the terminal. Based on this input data, it performs data preprocessing. Audio data is subjected to noise reduction filtering and processed to make the audio signal clearer. Video data is filtered to extract features important for image recognition. This results in a clean dataset ready for analysis.

[0311] Step 3:

[0312] The server performs emotion analysis using TensorFlow and OpenCV with pre-processed data. Clean audio and video datasets are used as input. Tone, pitch, and speed are detected from the audio data, and facial features and movements are detected from the video data, and data calculations are performed. The analyzed emotion data is output, clarifying the user's emotional state.

[0313] Step 4:

[0314] The analysis results are visualized by the server and sent back to the terminal in real time. The input is analyzed emotional data, and visualized information is generated based on it. The user can check the analysis results on the display of a glasses-type information device. This display shows icons and text messages indicating the customer's emotional state, and prompts the user to take action.

[0315] Step 5:

[0316] User feedback is sent to the server via the terminal. Based on the feedback received as input, the server updates its emotion analysis algorithm to improve its accuracy. The processing of feedback information is supported by specific generative AI models and prompt statements. This allows the system to continuously provide more sophisticated analysis results.

[0317] Furthermore, an emotion engine that estimates the user's emotions may be incorporated. That is, the identification processing unit 290 may use the emotion identification model 59 to estimate the user's emotions and perform identification processing using the user's emotions.

[0318] This invention realizes a system that incorporates an emotion engine for analyzing audio and video data and recognizing the user's emotions. The system consists of a server, a terminal, and the emotion engine, which work together to analyze the user's emotional state in real time.

[0319] First, the device records the user's voice using a microphone device and simultaneously acquires video data using a camera device. This audio and video data is then transmitted to the server, timestamped as it goes. The server processes this data using an emotion engine.

[0320] The emotion engine extracts features such as voice intonation and speed by applying natural language processing techniques to audio data. Furthermore, it uses computer vision technology to analyze changes in facial expressions and eye movements from video data. Based on this information, the emotion engine estimates multiple emotional states of the user.

[0321] As a concrete example of its application, consider communication in a confined space during a meeting. The terminal records the voice and facial expressions of meeting participants and immediately transmits them to the server. The emotion engine evaluates when participants are feeling joy or anxiety and visualizes this as an analysis result. The server displays the analysis results on the terminal's dashboard and presents them to the user.

[0322] Users can refer to this dashboard to adjust the meeting's progress and facilitate smoother dialogue. Furthermore, incorporating user feedback into the emotion engine optimizes model training and further improves prediction accuracy. By repeating this process, the system can evolve continuously over the long term, increasing its practical usability.

[0323] The following describes the processing flow.

[0324] Step 1:

[0325] The device uses a microphone and camera to collect the user's voice and video data. Voice data includes the speaker's tone, intonation, and speed. Video data includes visual information such as facial expressions and eye movements.

[0326] Step 2:

[0327] The device transmits collected audio and video data to the server in real time. The data is sent in a stream format to minimize latency.

[0328] Step 3:

[0329] The server receives the transmitted data and analyzes it using an emotion engine. For audio data, natural language processing techniques are applied to extract voice features and identify parameters that provide clues to emotion.

[0330] Step 4:

[0331] The server analyzes the video data using computer vision technology to detect changes in facial expressions and gaze. This reveals nonverbal emotional elements derived from the video.

[0332] Step 5:

[0333] The server's emotion engine integrates audio and video data and utilizes multiple emotion models to infer the overall emotional state. For example, it may detect "interest" and "anxiety" simultaneously.

[0334] Step 6:

[0335] The server visualizes the analysis results and generates an intuitive dashboard dataset. This dataset includes color-coded graphs and sentiment icons.

[0336] Step 7:

[0337] The device receives the dataset and displays it on the user's dashboard. This allows the user to understand their emotional state in real time.

[0338] Step 8:

[0339] Users can adjust their responses or change the direction of discussions based on the information on the dashboard. Furthermore, the device sends feedback received from users to the server, which is then incorporated into the emotion engine's machine learning model.

[0340] Step 9:

[0341] The server analyzes the feedback information and updates the emotion engine to improve analysis accuracy. Through this process, the system accumulates insights over time and becomes able to recognize user emotions more accurately.

[0342] (Example 2)

[0343] Next, we will describe Example 2. In the following description, the data processing device 12 will be referred to as the "server" and the smart glasses 214 will be referred to as the "terminal".

[0344] In recent years, with the advancement of information and communication technologies using voice and video, accurately understanding users' emotions in real time has become crucial. However, conventional systems often analyze voice and video information separately, which limits the accuracy of emotional state recognition. Furthermore, there is a lack of mechanisms to effectively utilize user feedback to improve the accuracy of emotion analysis models. Therefore, there is a need for a system that can analyze emotions with high accuracy and continuously optimize learning.

[0345] The identification process performed by the identification processing unit 290 of the data processing device 12 in Example 2 is realized by the following means.

[0346] In this invention, the server includes means for collecting user acoustic and video information via communication means, means for applying natural language processing techniques to extract features based on the acoustic information, and means for applying computer vision techniques to analyze facial expressions and gaze based on the video information. This enables integrated analysis of audio and video data, allowing for more accurate estimation of the user's emotional state. Furthermore, by visualizing the analysis results and presenting them to the user, feedback can be received to optimize the learning of the emotion analysis model.

[0347] "Communication means" refers to a device or mechanism for collecting audio and video information from a user and transferring it to a server.

[0348] "Acoustic information" refers to data obtained from the user's voice and is used to estimate their emotional state.

[0349] "Visual information" refers to data that records changes in the user's face and facial expressions, and is used to estimate their visual emotional state.

[0350] "Natural language processing technology" refers to techniques for converting acoustic information into text and extracting features such as intonation and speed.

[0351] "Computer vision technology" refers to techniques that analyze video information to detect facial expressions and eye movements.

[0352] An "emotion analysis model" is an algorithm or data model that estimates a user's emotional state based on acoustic and visual information.

[0353] "Visualization" is the process of displaying analysis results in the form of graphs, charts, and other formats to allow users to understand them intuitively.

[0354] "Feedback" refers to opinions and evaluations obtained from users, which are used to improve the accuracy of the model and optimize its learning process.

[0355] The system for implementing this invention mainly consists of a terminal, a server, and an emotion analysis engine. This system analyzes the user's emotional state using acoustic and visual information.

[0356] First, the device uses a microphone device and a camera device to simultaneously collect the user's voice and video. The hardware used includes high-precision microphone devices (e.g., typical audio recording devices) and high-resolution cameras (e.g., typical video recording devices). The audio and video information is time-stamped and transmitted to the server in real-time or near real-time.

[0357] The server performs advanced processing on the received data. Natural language processing techniques are used for acoustic information, and common acoustic analysis software such as the Google Cloud Speech-to-Text API is used to convert speech into text data, further extracting features such as intonation and speed. This analysis provides a means to estimate the user's emotional state.

[0358] Furthermore, the server applies computer vision technology to the video information, using common video analysis software such as OpenCV and TensorFlow to analyze the user's facial expressions and eye movements. This allows for the acquisition of emotional information that can be gleaned from the video.

[0359] The server integrates features obtained from audio and video and estimates the user's emotional state based on an emotion analysis model. The estimation results are visualized and displayed on the device's dashboard in the form of graphs and charts. This allows the user to intuitively understand the analysis results.

[0360] Furthermore, users contribute to optimizing the sentiment analysis model by providing feedback. This feedback information is fed into a continuous learning process, making it possible to improve prediction accuracy.

[0361] A concrete example is the ability to monitor the emotional state of each participant in a meeting in real time and adjust the meeting's progress as needed. A possible prompt in this scenario might be, "Please tell me how to assess participants' emotional states in real time and ensure the meeting runs smoothly." This system would enable better communication and decision-making.

[0362] The flow of the specific processing in Example 2 will be explained using Figure 13.

[0363] Step 1:

[0364] The device records the user's voice using a microphone device and acquires video using a camera device. The input is real-time audio and video. This data is timestamped and synchronized. The output is timestamped audio and video data, ready to be sent to the server immediately.

[0365] Step 2:

[0366] The terminal transmits the collected data to the server via a high-speed communication network. The input is the time-stamped audio and video data collected in step 1. The output is the audio and video data received by the server. This process enables real-time analysis of the data.

[0367] Step 3:

[0368] The server processes the audio data using acoustic analysis software such as the Google Cloud Speech-to-Text API. The input is the acoustic information sent to the server. Specifically, the audio is converted into text, and features related to intonation and speed are extracted. The output of this process is a feature vector for estimating the user's emotional state.

[0369] Step 4:

[0370] The server processes video data using tools such as OpenCV and TensorFlow. The input is video information sent to the server. Specifically, it analyzes facial expressions and eye movements from the video and extracts feature information obtained from this analysis. The output is a dataset containing this aggregated feature information.

[0371] Step 5:

[0372] The server integrates features obtained from audio and video and estimates the user's emotional state using an emotion analysis model. The input is the output from steps 3 and 4. The emotion analysis model utilizes a generative AI model to calculate the probability of various emotions. The output is the estimated result indicating the user's emotional state.

[0373] Step 6:

[0374] The server visualizes the estimation results and displays them on the terminal's dashboard in the form of a graph or chart. The input is the estimated emotional state obtained in step 5. The output is the analysis results provided in a visually understandable format for the user. This information serves as foundational data for the user to intuitively understand the analysis results and provide appropriate feedback.

[0375] Step 7:

[0376] The user reviews the analysis results presented on the dashboard and provides feedback to the system. The input is the analysis results presented to the user. The output is feedback data to improve the accuracy of the sentiment analysis model. This facilitates the continuous learning and optimization of the model.

[0377] (Application Example 2)

[0378] Next, we will explain application example 2. In the following explanation, the data processing device 12 will be referred to as the "server," and the smart glasses 214 will be referred to as the "terminal."

[0379] In recent years, brick-and-mortar stores have been required to understand customer satisfaction in real time and provide services accordingly. However, traditional methods have the problem of not being able to immediately perceive customers' emotions. There is a need to provide an effective means for employees to accurately and quickly evaluate customers' facial expressions and voices and directly link that evaluation to service improvement.

[0380] The specific processing performed by the specific processing unit 290 of the data processing device 12 in Application Example 2 is realized by the following means.

[0381] In this invention, the server includes means for acquiring audio and video data via a communication device, processing means for performing sentiment analysis based on the audio and video data, and means for visualizing the analysis results and presenting them to the user using an output device. This enables real-time analysis of customer sentiment, allowing employees to immediately provide services based on that analysis.

[0382] A "communication device" is a device used to acquire audio and video data and transfer it to a server.

[0383] "Voice data" refers to digital information about human speech collected using voice acquisition devices such as microphones.

[0384] "Video data" refers to digital data based on visual information, acquired through video acquisition devices such as cameras.

[0385] "Emotional analysis" is a process that estimates a user's emotional state using audio and video data.

[0386] "Processing means" refers to means consisting of software or hardware for performing sentiment analysis and visualizing the data.

[0387] An "output device" is a device such as a display that visually presents the analysis results from the server to the user.

[0388] "A means of analyzing customer emotions in real time from facial expressions and voice and presenting them to workers" refers to a technology and means for instantly evaluating the emotional state of customers in physical stores and providing feedback on the results to workers.

[0389] "Feedback information" refers to evaluations and comments from users regarding sentiment analysis, and is data used to improve the analysis model.

[0390] In implementing this invention, a system will be constructed to perform sentiment analysis in customer-employee interactions in physical stores. This will be achieved using a server, terminals, and a sentiment analysis engine.

[0391] The server receives audio and video data from terminals via a communication device. Terminals, such as smart glasses or smartphones, have built-in microphones and cameras to capture customer audio and video in real time. After receiving the data, the server processes it using an emotion analysis engine. For audio data, natural language processing techniques are applied to analyze properties such as intonation and speed, and for video data, computer vision techniques are used to capture changes in facial expressions and gaze.

[0392] The data is first processed by a Python-based program. Audio data has its features extracted using the Librosa library, and video data is analyzed using OpenCV and a TensorFlow machine learning model. The analysis results are then visually displayed using an output device, for example, on the display of smart glasses worn by employees.

[0393] As a concrete example of this system, when a customer enters a store, the staff can instantly analyze whether the customer is smiling when they greet them with "Welcome." If a positive emotion is detected from the customer's expression, a message such as "Satisfied" will appear on the staff member's glasses, allowing for immediate improvement in the quality of service.

[0394] An example of a prompt message would be, "Please tell me the steps to develop an application that analyzes customer facial expressions and voice in real time and provides emotional feedback on a smart device."

[0395] The flow of a specific process in Application Example 2 will be explained using Figure 14.

[0396] Step 1:

[0397] The device uses smart glasses or a smartphone to record audio data with a microphone and capture video data with a camera. This provides instant digital information about the customer's voice tone and facial expressions.

[0398] Step 2:

[0399] The terminal transfers the acquired audio and video data to the server in real time. During this process, the data is time-stamped and synchronized. The input data transmitted consists of recorded audio and video files.

[0400] Step 3:

[0401] The server analyzes the received audio data using the Librosa library. This process extracts features such as intonation and speed from the speech and generates feature data based on these features. The audio waveform and spectral information are processed, and numerical data suggesting emotional state is output.

[0402] Step 4:

[0403] The server analyzes video data using the OpenCV library. Specifically, it recognizes facial expressions and eye movements, and converts this visual information into numerical data. It detects facial landmarks from video frames and identifies features that indicate emotion. The output is a parameter set indicating emotion.

[0404] Step 5:

[0405] The emotion analysis engine on the server integrates audio feature data and video parameter sets, and performs analysis using a machine learning model. TensorFlow is used to estimate the customer's emotional state through the integrated data. As a result of this integration, an emotion label and its confidence level are output.

[0406] Step 6:

[0407] The user (store staff) visually checks the analysis results received from the server on the display of smart glasses. This allows them to understand the customer's emotional state in real time. The output displays a specific emotional state (e.g., "satisfied," "dissatisfied").

[0408] The specific processing unit 290 transmits the result of the specific processing to the smart glasses 214. In the smart glasses 214, the control unit 46A causes the speaker 240 to output the result of the specific processing. The microphone 238 acquires audio indicating user input for the result of the specific processing. The control unit 46A transmits the audio data indicating user input acquired by the microphone 238 to the data processing unit 12. In the data processing unit 12, the specific processing unit 290 acquires the audio data.

[0409] Data generation model 58 is a type of so-called generative AI (Artificial Intelligence). One example of data generation model 58 is ChatGPT (Internet search<URL: https: / / openai.com / blog / chatgpt> ), Gemini (Internet search) <url: https: gemini.google.com ?hl="ja">Examples of generative AI include the following. The data generation model 58 is obtained by performing deep learning on a neural network. The data generation model 58 is input with prompts containing instructions, and with inference data such as audio data representing speech, text data representing text, and image data representing images. The data generation model 58 infers from the input inference data according to the instructions indicated by the prompts, and outputs the inference results in data formats such as audio data and text data. Here, inference refers to, for example, analysis, classification, prediction, and / or summarization.

[0410] In the above embodiment, an example was given in which specific processing is performed by the data processing device 12, but the technology of this disclosure is not limited thereto, and the specific processing may also be performed by the smart glasses 214.

[0411] [Third Embodiment]

[0412] Figure 5 shows an example of the configuration of the data processing system 310 according to the third embodiment.

[0413] As shown in Figure 5, the data processing system 310 includes a data processing device 12 and a headset terminal 314. An example of the data processing device 12 is a server.

[0414] The data processing device 12 comprises a computer 22, a database 24, and a communication interface 26. The computer 22 is an example of a "computer" related to the technology of this disclosure. The computer 22 comprises a processor 28, RAM 30, and storage 32. The processor 28, RAM 30, and storage 32 are connected to a bus 34. The database 24 and the communication interface 26 are also connected to the bus 34. The communication interface 26 is connected to a network 54. An example of the network 54 is a WAN (Wide Area Network) and / or a LAN (Local Area Network).

[0415] The headset terminal 314 includes a computer 36, a microphone 238, a speaker 240, a camera 42, a communication interface 44, and a display 343. The computer 36 includes a processor 46, RAM 48, and storage 50. The processor 46, RAM 48, and storage 50 are connected to a bus 52. The microphone 238, speaker 240, camera 42, and display 343 are also connected to the bus 52.

[0416] The microphone 238 receives voice signals from the user 20 and receives instructions from the user 20. The microphone 238 captures the voice signals from the user 20, converts the captured voice into audio data, and outputs it to the processor 46. The speaker 240 outputs audio according to the instructions from the processor 46.

[0417] Camera 42 is a small digital camera equipped with an optical system including a lens, aperture, and shutter, and an image sensor such as a CMOS (Complementary Metal-Oxide-Semiconductor) image sensor or a CCD (Charge Coupled Device) image sensor, and captures images of the area around the user 20 (for example, an imaging range defined by a field of view equivalent to the width of a typical healthy person's field of vision).

[0418] Communication interface 44 is connected to network 54. Communication interfaces 44 and 26 are responsible for the exchange of various information between processor 46 and processor 28 via network 54. The exchange of various information between processor 46 and processor 28 using communication interfaces 44 and 26 is performed in a secure manner.

[0419] Figure 6 shows an example of the main functions of the data processing device 12 and the headset terminal 314. As shown in Figure 6, the data processing device 12 performs specific processing using the processor 28. The storage 32 stores the specific processing program 56.

[0420] The specific processing program 56 is an example of a "program" relating to the technology of this disclosure. The processor 28 reads the specific processing program 56 from the storage 32 and executes the read specific processing program 56 on the RAM 30. The specific processing is realized by the processor 28 operating as a specific processing unit 290 in accordance with the specific processing program 56 executed on the RAM 30.

[0421] The storage 32 stores the data generation model 58 and the emotion identification model 59. The data generation model 58 and the emotion identification model 59 are used by the identification processing unit 290.

[0422] In the headset terminal 314, the processor 46 performs the reception output processing. The storage 50 stores the reception output program 60. The processor 46 reads the reception output program 60 from the storage 50 and executes the read reception output program 60 on the RAM 48. The reception output processing is realized by the processor 46 operating as a control unit 46A according to the reception output program 60 executed on the RAM 48.

[0423] Next, the specific processing performed by the specific processing unit 290 of the data processing device 12 will be described. In the following description, the data processing device 12 will be referred to as the "server" and the headset terminal 314 will be referred to as the "terminal".

[0424] The system according to the present invention has the function of acquiring voice and video data using a specific communication device and performing emotion analysis. To achieve this, a server, a terminal, and a user must cooperate in operation.

[0425] First, the device acquires audio data from the user's surroundings using a microphone and simultaneously acquires video data using a camera. This enables real-time data collection. The collected data is sent to a server via a security protocol, and the server receives the data. Next, the server activates an advanced algorithm for sentiment analysis based on the collected audio and video data. This algorithm utilizes natural language processing and computer vision technologies to analyze the emotional state.

[0426] Specifically, the server extracts factors such as voice tone, pitch, and speed from audio data and uses this information to infer emotional states. It also analyzes facial expressions, eye movements, and cheek redness from video data to obtain additional emotional information. By integrating these results, the server then performs the final emotion identification.

[0427] The server then visualizes the analysis results and returns them to the terminal as data. The terminal intuitively displays this data on a dashboard. For example, it can simultaneously monitor the emotions of multiple participants during a meeting and inform the user of each participant's emotional state. The user reviews the dashboard and adjusts their approach to the conversation as needed.

[0428] Furthermore, the server receives feedback information provided by users and uses it to improve the analysis model. This allows the system to improve its accuracy over time, enabling more reliable sentiment analysis. This feedback loop is a crucial element in continuously optimizing the system's performance.

[0429] The following describes the processing flow.

[0430] Step 1:

[0431] The device uses a microphone to collect audio data, including ambient sounds, and simultaneously captures the user's video data with a camera. This allows for the collection of emotion-related information, such as specific utterances and changes in facial expressions.

[0432] Step 2:

[0433] The device sends collected audio and video data to the server via a secure protocol. During transmission, the data is encrypted to protect privacy.

[0434] Step 3:

[0435] To analyze the data received by the server in real time, natural language processing algorithms are applied to the audio data to extract speech features. These extracted features include speed, emphasis, and intonation.

[0436] Step 4:

[0437] The server uses computer vision algorithms to analyze video data and identify emotion-related features from the user's facial expressions and movements. For example, it can infer emotions such as joy, anger, sadness, or happiness based on smiles and eyebrow furrowing.

[0438] Step 5:

[0439] The server integrates the audio and video analysis results and uses statistical models and machine learning algorithms to comprehensively identify the user's emotional state. It then outputs a specific emotional category (e.g., joy, anxiety, surprise).

[0440] Step 6:

[0441] The server processes the sentiment analysis results as visual data and generates a dataset for the dashboard. This includes visual elements such as color coding and icons.

[0442] Step 7:

[0443] The device acquires visual data and displays it in real time on the user's dashboard. The user can observe this and take appropriate action based on the situation.

[0444] Step 8:

[0445] When a user provides feedback based on the system's analysis results, the terminal receives this information and sends it to the server. This feedback is used for the continuous training of the analysis model, contributing to improving the system's accuracy.

[0446] (Example 1)

[0447] Next, we will describe Example 1. In the following description, the data processing device 12 will be referred to as the "server," and the headset-type terminal 314 will be referred to as the "terminal."

[0448] When analyzing human emotions from audio and video data, there is a challenge in accurately identifying emotions using analysis based on a single source of information. In addition, there is a problem in that the use of feedback to effectively present analysis results to users and improve the accuracy of the system is insufficient.

[0449] The identification process performed by the identification processing unit 290 of the data processing device 12 in Example 1 is realized by the following means.

[0450] In this invention, the server includes acquisition means for acquiring audio and video signals, analysis means for analyzing emotional states based on the audio and video signals using natural language processing and image analysis techniques, and integration means for integrating the analysis results from the analysis means and identifying emotions. This enables accurate emotion identification utilizing multimodal information and improved accuracy based on user feedback.

[0451] "Audio signals" are data obtained by converting sound vibrations acquired from individuals or their environment into a format that can be electrically transmitted.

[0452] A "video signal" is data obtained by converting visual information acquired by a visual device such as a camera into a format that can be transmitted electronically.

[0453] "Natural language processing" is a technology that uses computers to process and analyze human language, and is used to recognize emotions and intentions from speech and text.

[0454] "Image analysis technology" refers to computational techniques that extract features from video data to understand the state of objects and people.

[0455] "Emotional state" refers to the emotions an individual experiences at a particular moment, and can be inferred from their tone of voice and facial expressions.

[0456] "Integration methods" refer to processes and techniques for consolidating analytical results obtained from multiple sources into a single conclusion.

[0457] An "analytical model" is a set of algorithms and computational procedures used to analyze data, which are learned to recognize specific patterns or meanings.

[0458] "Feedback" refers to evaluations and information provided by users to improve the performance of a system.

[0459] "Improving accuracy" refers to increasing the performance of a system, specifically how closely its output matches real-world conditions.

[0460] This invention relates to a system that identifies emotional states by analyzing audio and video signals and provides the user with the analysis results. To implement the invention, a microphone for acquiring audio signals, a camera for acquiring video signals, a server for data processing, and a terminal for presenting these to the user are required.

[0461] terminal

[0462] The device acquires audio signals using microphones placed around the user and video signals using a camera. This data is transmitted to a server in real time. In particular, the device performs acoustic processing to remove ambient noise and allow the user to focus on the conversation. The camera also uses a face detection algorithm to facilitate analysis of the user's facial expressions.

[0463] server

[0464] The server receives audio and video signals transmitted from the terminal. Audio signals are analyzed using natural language processing techniques to extract features such as tone, pitch, and speed. Video signals are analyzed using image analysis techniques to extract facial movements and changes in expression. These processes utilize algorithms powered by generative AI models to estimate emotional states. Based on the analysis results, the server identifies each user's emotions and converts them into a visualizeable data format.

[0465] User

[0466] Users visually review the analysis results sent back from the server via a dashboard on their device. This allows them to understand the emotional state of other participants during the meeting. Users provide feedback to the server via their device, and this feedback is used to improve the analysis model.

[0467] For example, if a participant is feeling anxious during an online meeting, their name and emotional state will be displayed on the device's dashboard. A prompt such as, "Analyze the emotional states of current meeting participants in real time and display any participants who appear anxious," allows for a quick response.

[0468] The flow of the specific processing in Example 1 will be explained using Figure 11.

[0469] Step 1:

[0470] The device acquires audio signals through microphones placed around the user and collects video signals using a camera. The inputs are ambient audio and visual data. The audio signal undergoes acoustic processing to remove unwanted ambient noise. The video signal uses a face detection algorithm to pinpoint the location of faces. The output of this processing is clean audio data and video data with faces detected.

[0471] Step 2:

[0472] The terminal transmits the processed audio and video data to the server via a security protocol. Here, the data is encrypted before transmission to ensure privacy and security. The input is the output data from the previous step, and the output is encrypted audio and video data.

[0473] Step 3:

[0474] The server receives the transmitted data and applies natural language processing techniques to the audio signal. Specifically, it extracts tone, pitch, and speed, and uses these features to estimate emotion. The input is encrypted audio data, and the output is an emotion estimation score.

[0475] Step 4:

[0476] The server analyzes video data and detects features such as facial expressions and movements. Using advanced computer vision technology, it quantifies various facial expressions and estimates possible emotional states. The input is encrypted video data, and the output is an emotion estimation score based on facial expressions.

[0477] Step 5:

[0478] The server integrates estimated scores obtained from audio and video and performs final emotion identification using a generative AI model. The resulting overall emotion score is stored in a database and converted into a data format for visualization. The input is the estimated score for each emotion, and the output is a visualized emotion identification result.

[0479] Step 6:

[0480] The server sends emotion recognition data to the terminal, which displays it on a dashboard. Users refer to the information displayed on the dashboard to adjust their approach to meetings and conversations. The input is a visualized emotion recognition result, and the output is an intuitive interface display presented to the user.

[0481] Step 7:

[0482] The user inputs feedback on the system's sentiment analysis results into a terminal. The server receives the feedback and uses it to improve the analysis model in the generated AI model. The input is the user's feedback, and the output is the updated analysis model.

[0483] (Application Example 1)

[0484] Next, we will explain Application Example 1. In the following explanation, the data processing device 12 will be referred to as the "server," and the headset-type terminal 314 will be referred to as the "terminal."

[0485] In modern customer service, understanding individual customer emotions in real time and improving service quality is becoming increasingly important. However, traditional methods make it difficult to accurately interpret emotions from customers' facial expressions and voices, hindering service improvement. To solve this problem, a system is needed that utilizes acoustic and visual information to analyze individual customer emotions in real time and respond quickly.

[0486] The specific processing performed by the specific processing unit 290 of the data processing device 12 in Application Example 1 is realized by the following means.

[0487] In this invention, the server includes means for acquiring acoustic and video information via communication equipment, computation means for analyzing emotions based on the acoustic and video information, and means for visualizing the analysis results and providing them to the individual using a display device. This enables service staff to immediately grasp the customer's emotions and provide appropriate service.

[0488] "Communication equipment" refers to devices used to acquire and transmit or receive audio and video information externally.

[0489] "Acoustic information" refers to data related to speech and tone, including characteristics such as the tone and speed of the speaker's voice.

[0490] "Visual information" refers to data related to images or videos that is used to analyze facial expressions and movements.

[0491] "Emotion" refers to an expression of an individual's emotional state or psychological reaction.

[0492] The "computation means" is a function that analyzes the acquired acoustic and visual information and performs processing to identify emotions.

[0493] A "display device" refers to a screen device or glasses-type information device that visually presents analyzed emotions to an individual.

[0494] "Visualization" is the process of representing data visually and providing it in an easily understandable format.

[0495] This invention aims to realize a system that utilizes communication equipment to collect acoustic and visual information and perform emotional analysis in real time. The user wears glasses-type information devices to acquire acoustic and visual information obtained from interactions with customers. These devices use built-in microphones and cameras to collect ambient sounds and customer facial expressions in real time. The collected data is transmitted to a server via Bluetooth or Wi-Fi.

[0496] The server performs emotion analysis using TensorFlow and OpenCV based on acquired acoustic and video information. It analyzes speech tone, speed, and pitch from the acoustic information, and evaluates facial features and movements from the video information. The resulting emotion data is then comprehensively evaluated using computer vision and natural language processing technologies.

[0497] The analyzed emotional data is visualized in real time on the user's glasses-type information device display. As a result, the user can quickly grasp the customer's emotional state and provide appropriate service. Furthermore, feedback information is sent to the server, and the accuracy is continuously improved by refining the emotional analysis algorithm. This feedback is obtained through manual input and automatic detection, contributing to the improvement of the system's performance.

[0498] As a concrete example, when a store clerk is introducing a product to a customer, if anxiety is detected from the customer's facial expression and voice, the display on the glasses will show instructions such as, "The customer is feeling anxious. Please provide reassuring information." This information is then analyzed by a generative AI model using prompts such as, "Analyze how customer A feels about the new product. Audio and video samples are available."

[0499] Through the above, it is possible to improve customer satisfaction in customer service operations.

[0500] The flow of a specific process in Application Example 1 will be explained using Figure 12.

[0501] Step 1:

[0502] The device uses a microphone and camera to acquire ambient acoustic and visual information. The input here is real-time collected audio and image data. The device temporarily stores this data and prepares it for transmission to a server via Bluetooth or Wi-Fi.

[0503] Step 2:

[0504] The server receives audio and video information transmitted from the terminal. Based on this input data, it performs data preprocessing. Audio data is subjected to noise reduction filtering and processed to make the audio signal clearer. Video data is filtered to extract features important for image recognition. This results in a clean dataset ready for analysis.

[0505] Step 3:

[0506] The server performs emotion analysis using TensorFlow and OpenCV with pre-processed data. Clean audio and video datasets are used as input. Tone, pitch, and speed are detected from the audio data, and facial features and movements are detected from the video data, and data calculations are performed. The analyzed emotion data is output, clarifying the user's emotional state.

[0507] Step 4:

[0508] The analysis results are visualized by the server and sent back to the terminal in real time. The input is analyzed emotional data, and visualized information is generated based on it. The user can check the analysis results on the display of a glasses-type information device. This display shows icons and text messages indicating the customer's emotional state, and prompts the user to take action.

[0509] Step 5:

[0510] User feedback is sent to the server via the terminal. Based on the feedback received as input, the server updates its emotion analysis algorithm to improve its accuracy. The processing of feedback information is supported by specific generative AI models and prompt statements. This allows the system to continuously provide more sophisticated analysis results.

[0511] Furthermore, an emotion engine that estimates the user's emotions may be incorporated. That is, the identification processing unit 290 may use the emotion identification model 59 to estimate the user's emotions and perform identification processing using the user's emotions.

[0512] This invention realizes a system that incorporates an emotion engine for analyzing audio and video data and recognizing the user's emotions. The system consists of a server, a terminal, and the emotion engine, which work together to analyze the user's emotional state in real time.

[0513] First, the device records the user's voice using a microphone device and simultaneously acquires video data using a camera device. This audio and video data is then transmitted to the server, timestamped as it goes. The server processes this data using an emotion engine.

[0514] The emotion engine extracts features such as voice intonation and speed by applying natural language processing techniques to audio data. Furthermore, it uses computer vision technology to analyze changes in facial expressions and eye movements from video data. Based on this information, the emotion engine estimates multiple emotional states of the user.

[0515] As a concrete example of its application, consider communication in a confined space during a meeting. The terminal records the voice and facial expressions of meeting participants and immediately transmits them to the server. The emotion engine evaluates when participants are feeling joy or anxiety and visualizes this as an analysis result. The server displays the analysis results on the terminal's dashboard and presents them to the user.

[0516] Users can refer to this dashboard to adjust the meeting's progress and facilitate smoother dialogue. Furthermore, incorporating user feedback into the emotion engine optimizes model training and further improves prediction accuracy. By repeating this process, the system can evolve continuously over the long term, increasing its practical usability.

[0517] The following describes the processing flow.

[0518] Step 1:

[0519] The device uses a microphone and camera to collect the user's voice and video data. Voice data includes the speaker's tone, intonation, and speed. Video data includes visual information such as facial expressions and eye movements.

[0520] Step 2:

[0521] The device transmits collected audio and video data to the server in real time. The data is sent in a stream format to minimize latency.

[0522] Step 3:

[0523] The server receives the transmitted data and analyzes it using an emotion engine. For audio data, natural language processing techniques are applied to extract voice features and identify parameters that provide clues to emotion.

[0524] Step 4:

[0525] The server analyzes the video data using computer vision technology to detect changes in facial expressions and gaze. This reveals nonverbal emotional elements derived from the video.

[0526] Step 5:

[0527] The server's emotion engine integrates audio and video data and utilizes multiple emotion models to infer the overall emotional state. For example, it may detect "interest" and "anxiety" simultaneously.

[0528] Step 6:

[0529] The server visualizes the analysis results and generates an intuitive dashboard dataset. This dataset includes color-coded graphs and sentiment icons.

[0530] Step 7:

[0531] The device receives the dataset and displays it on the user's dashboard. This allows the user to understand their emotional state in real time.

[0532] Step 8:

[0533] Users can adjust their responses or change the direction of discussions based on the information on the dashboard. Furthermore, the device sends feedback received from users to the server, which is then incorporated into the emotion engine's machine learning model.

[0534] Step 9:

[0535] The server analyzes the feedback information and updates the emotion engine to improve analysis accuracy. Through this process, the system accumulates insights over time and becomes able to recognize user emotions more accurately.

[0536] (Example 2)

[0537] Next, we will describe Example 2. In the following description, the data processing device 12 will be referred to as the "server," and the headset-type terminal 314 will be referred to as the "terminal."

[0538] In recent years, with the advancement of information and communication technologies using voice and video, accurately understanding users' emotions in real time has become crucial. However, conventional systems often analyze voice and video information separately, which limits the accuracy of emotional state recognition. Furthermore, there is a lack of mechanisms to effectively utilize user feedback to improve the accuracy of emotion analysis models. Therefore, there is a need for a system that can analyze emotions with high accuracy and continuously optimize learning.

[0539] The identification process performed by the identification processing unit 290 of the data processing device 12 in Example 2 is realized by the following means.

[0540] In this invention, the server includes means for collecting user acoustic and video information via communication means, means for applying natural language processing techniques to extract features based on the acoustic information, and means for applying computer vision techniques to analyze facial expressions and gaze based on the video information. This enables integrated analysis of audio and video data, allowing for more accurate estimation of the user's emotional state. Furthermore, by visualizing the analysis results and presenting them to the user, feedback can be received to optimize the learning of the emotion analysis model.

[0541] "Communication means" refers to a device or mechanism for collecting audio and video information from a user and transferring it to a server.

[0542] "Acoustic information" refers to data obtained from the user's voice and is used to estimate their emotional state.

[0543] "Visual information" refers to data that records changes in the user's face and facial expressions, and is used to estimate their visual emotional state.

[0544] "Natural language processing technology" refers to techniques for converting acoustic information into text and extracting features such as intonation and speed.

[0545] "Computer vision technology" refers to techniques that analyze video information to detect facial expressions and eye movements.

[0546] An "emotion analysis model" is an algorithm or data model that estimates a user's emotional state based on acoustic and visual information.

[0547] "Visualization" is the process of displaying analysis results in the form of graphs, charts, and other formats to allow users to understand them intuitively.

[0548] "Feedback" refers to opinions and evaluations obtained from users, which are used to improve the accuracy of the model and optimize its learning process.

[0549] The system for implementing this invention mainly consists of a terminal, a server, and an emotion analysis engine. This system analyzes the user's emotional state using acoustic and visual information.

[0550] First, the device uses a microphone device and a camera device to simultaneously collect the user's voice and video. The hardware used includes high-precision microphone devices (e.g., typical audio recording devices) and high-resolution cameras (e.g., typical video recording devices). The audio and video information is time-stamped and transmitted to the server in real-time or near real-time.

[0551] The server performs advanced processing on the received data. Natural language processing techniques are used for acoustic information, and common acoustic analysis software such as the Google Cloud Speech-to-Text API is used to convert speech into text data, further extracting features such as intonation and speed. This analysis provides a means to estimate the user's emotional state.

[0552] Furthermore, the server applies computer vision technology to the video information, using common video analysis software such as OpenCV and TensorFlow to analyze the user's facial expressions and eye movements. This allows for the acquisition of emotional information that can be gleaned from the video.

[0553] The server integrates features obtained from audio and video and estimates the user's emotional state based on an emotion analysis model. The estimation results are visualized and displayed on the device's dashboard in the form of graphs and charts. This allows the user to intuitively understand the analysis results.

[0554] Furthermore, users contribute to optimizing the sentiment analysis model by providing feedback. This feedback information is fed into a continuous learning process, making it possible to improve prediction accuracy.

[0555] A concrete example is the ability to monitor the emotional state of each participant in a meeting in real time and adjust the meeting's progress as needed. A possible prompt in this scenario might be, "Please tell me how to assess participants' emotional states in real time and ensure the meeting runs smoothly." This system would enable better communication and decision-making.

[0556] The flow of the specific processing in Example 2 will be explained using Figure 13.

[0557] Step 1:

[0558] The device records the user's voice using a microphone device and acquires video using a camera device. The input is real-time audio and video. This data is timestamped and synchronized. The output is timestamped audio and video data, ready to be sent to the server immediately.

[0559] Step 2:

[0560] The terminal transmits the collected data to the server via a high-speed communication network. The input is the time-stamped audio and video data collected in step 1. The output is the audio and video data received by the server. This process enables real-time analysis of the data.

[0561] Step 3:

[0562] The server processes the audio data using acoustic analysis software such as the Google Cloud Speech-to-Text API. The input is the acoustic information sent to the server. Specifically, the audio is converted into text, and features related to intonation and speed are extracted. The output of this process is a feature vector for estimating the user's emotional state.

[0563] Step 4:

[0564] The server processes video data using tools such as OpenCV and TensorFlow. The input is video information sent to the server. Specifically, it analyzes facial expressions and eye movements from the video and extracts feature information obtained from this analysis. The output is a dataset containing this aggregated feature information.

[0565] Step 5:

[0566] The server integrates features obtained from audio and video and estimates the user's emotional state using an emotion analysis model. The input is the output from steps 3 and 4. The emotion analysis model utilizes a generative AI model to calculate the probability of various emotions. The output is the estimated result indicating the user's emotional state.

[0567] Step 6:

[0568] The server visualizes the estimation results and displays them on the terminal's dashboard in the form of a graph or chart. The input is the estimated emotional state obtained in step 5. The output is the analysis results provided in a visually understandable format for the user. This information serves as foundational data for the user to intuitively understand the analysis results and provide appropriate feedback.

[0569] Step 7:

[0570] The user reviews the analysis results presented on the dashboard and provides feedback to the system. The input is the analysis results presented to the user. The output is feedback data to improve the accuracy of the sentiment analysis model. This facilitates the continuous learning and optimization of the model.

[0571] (Application Example 2)

[0572] Next, we will explain application example 2. In the following explanation, the data processing device 12 will be referred to as the "server," and the headset-type terminal 314 will be referred to as the "terminal."

[0573] In recent years, brick-and-mortar stores have been required to understand customer satisfaction in real time and provide services accordingly. However, traditional methods have the problem of not being able to immediately perceive customers' emotions. There is a need to provide an effective means for employees to accurately and quickly evaluate customers' facial expressions and voices and directly link that evaluation to service improvement.

[0574] The specific processing performed by the specific processing unit 290 of the data processing device 12 in Application Example 2 is realized by the following means.

[0575] In this invention, the server includes means for acquiring audio and video data via a communication device, processing means for performing sentiment analysis based on the audio and video data, and means for visualizing the analysis results and presenting them to the user using an output device. This enables real-time analysis of customer sentiment, allowing employees to immediately provide services based on that analysis.

[0576] A "communication device" is a device used to acquire audio and video data and transfer it to a server.

[0577] "Voice data" refers to digital information about human speech collected using voice acquisition devices such as microphones.

[0578] "Video data" refers to digital data based on visual information, acquired through video acquisition devices such as cameras.

[0579] "Emotional analysis" is a process that estimates a user's emotional state using audio and video data.

[0580] "Processing means" refers to means consisting of software or hardware for performing sentiment analysis and visualizing the data.

[0581] An "output device" is a device such as a display that visually presents the analysis results from the server to the user.

[0582] "A means of analyzing customer emotions in real time from facial expressions and voice and presenting them to workers" refers to a technology and means for instantly evaluating the emotional state of customers in physical stores and providing feedback on the results to workers.

[0583] "Feedback information" refers to evaluations and comments from users regarding sentiment analysis, and is data used to improve the analysis model.

[0584] In implementing this invention, a system will be constructed to perform sentiment analysis in customer-employee interactions in physical stores. This will be achieved using a server, terminals, and a sentiment analysis engine.

[0585] The server receives audio and video data from terminals via a communication device. Terminals, such as smart glasses or smartphones, have built-in microphones and cameras to capture customer audio and video in real time. After receiving the data, the server processes it using an emotion analysis engine. For audio data, natural language processing techniques are applied to analyze properties such as intonation and speed, and for video data, computer vision techniques are used to capture changes in facial expressions and gaze.

[0586] The data is first processed by a Python-based program. Audio data has its features extracted using the Librosa library, and video data is analyzed using OpenCV and a TensorFlow machine learning model. The analysis results are then visually displayed using an output device, for example, on the display of smart glasses worn by employees.

[0587] As a concrete example of this system, when a customer enters a store, the staff can instantly analyze whether the customer is smiling when they greet them with "Welcome." If a positive emotion is detected from the customer's expression, a message such as "Satisfied" will appear on the staff member's glasses, allowing for immediate improvement in the quality of service.

[0588] An example of a prompt message would be, "Please tell me the steps to develop an application that analyzes customer facial expressions and voice in real time and provides emotional feedback on a smart device."

[0589] The flow of a specific process in Application Example 2 will be explained using Figure 14.

[0590] Step 1:

[0591] The device uses smart glasses or a smartphone to record audio data with a microphone and capture video data with a camera. This provides instant digital information about the customer's voice tone and facial expressions.

[0592] Step 2:

[0593] The terminal transfers the acquired audio and video data to the server in real time. During this process, the data is time-stamped and synchronized. The input data transmitted consists of recorded audio and video files.

[0594] Step 3:

[0595] The server analyzes the received audio data using the Librosa library. This process extracts features such as intonation and speed from the speech and generates feature data based on these features. The audio waveform and spectral information are processed, and numerical data suggesting emotional state is output.

[0596] Step 4:

[0597] The server analyzes video data using the OpenCV library. Specifically, it recognizes facial expressions and eye movements, and converts this visual information into numerical data. It detects facial landmarks from video frames and identifies features that indicate emotion. The output is a parameter set indicating emotion.

[0598] Step 5:

[0599] The emotion analysis engine on the server integrates audio feature data and video parameter sets, and performs analysis using a machine learning model. TensorFlow is used to estimate the customer's emotional state through the integrated data. As a result of this integration, an emotion label and its confidence level are output.

[0600] Step 6:

[0601] The user (store staff) visually checks the analysis results received from the server on the display of smart glasses. This allows them to understand the customer's emotional state in real time. The output displays a specific emotional state (e.g., "satisfied," "dissatisfied").

[0602] The specific processing unit 290 transmits the result of the specific processing to the headset terminal 314. In the headset terminal 314, the control unit 46A causes the speaker 240 and display 343 to output the result of the specific processing. The microphone 238 acquires audio indicating user input for the result of the specific processing. The control unit 46A transmits the audio data indicating user input acquired by the microphone 238 to the data processing unit 12. In the data processing unit 12, the specific processing unit 290 acquires the audio data.

[0603] Data generation model 58 is a type of so-called generative AI (Artificial Intelligence). One example of data generation model 58 is ChatGPT (Internet search<URL: https: / / openai.com / blog / chatgpt> ), Gemini (Internet search) <url: https: gemini.google.com ?hl="ja">Examples of generative AI include the following. The data generation model 58 is obtained by performing deep learning on a neural network. The data generation model 58 is input with prompts containing instructions, and with inference data such as audio data representing speech, text data representing text, and image data representing images. The data generation model 58 infers from the input inference data according to the instructions indicated by the prompts, and outputs the inference results in data formats such as audio data and text data. Here, inference refers to, for example, analysis, classification, prediction, and / or summarization.

[0604] In the above embodiment, an example was given in which specific processing is performed by the data processing device 12, but the technology of this disclosure is not limited thereto, and specific processing may also be performed by the headset terminal 314.

[0605] [Fourth Embodiment]

[0606] Figure 7 shows an example of the configuration of the data processing system 410 according to the fourth embodiment.

[0607] As shown in Figure 7, the data processing system 410 includes a data processing device 12 and a robot 414. An example of the data processing device 12 is a server.

[0608] The data processing device 12 comprises a computer 22, a database 24, and a communication interface 26. The computer 22 is an example of a "computer" related to the technology of this disclosure. The computer 22 comprises a processor 28, RAM 30, and storage 32. The processor 28, RAM 30, and storage 32 are connected to a bus 34. The database 24 and the communication interface 26 are also connected to the bus 34. The communication interface 26 is connected to a network 54. An example of the network 54 is a WAN (Wide Area Network) and / or a LAN (Local Area Network).

[0609] The robot 414 includes a computer 36, a microphone 238, a speaker 240, a camera 42, a communication interface 44, and a controlled object 443. The computer 36 includes a processor 46, RAM 48, and storage 50. The processor 46, RAM 48, and storage 50 are connected to a bus 52. The microphone 238, speaker 240, camera 42, and controlled object 443 are also connected to the bus 52.

[0610] The microphone 238 receives voice signals from the user 20 and receives instructions from the user 20. The microphone 238 captures the voice signals from the user 20, converts the captured voice into audio data, and outputs it to the processor 46. The speaker 240 outputs audio according to the instructions from the processor 46.

[0611] Camera 42 is a small digital camera equipped with an optical system including a lens, aperture, and shutter, and an image sensor such as a CMOS (Complementary Metal-Oxide-Semiconductor) image sensor or a CCD (Charge Coupled Device) image sensor, and captures images of the area around the user 20 (for example, an imaging range defined by a field of view equivalent to the width of a typical healthy person's field of vision).

[0612] Communication interface 44 is connected to network 54. Communication interfaces 44 and 26 are responsible for the exchange of various information between processor 46 and processor 28 via network 54. The exchange of various information between processor 46 and processor 28 using communication interfaces 44 and 26 is performed in a secure manner.

[0613] The controlled object 443 includes a display device, LEDs in the eyes, and motors that drive the arms, hands, and feet. The posture and gestures of the robot 414 are controlled by controlling the motors of the arms, hands, and feet. Some of the robot 414's emotions can be expressed by controlling these motors. Furthermore, the robot 414's facial expressions can also be expressed by controlling the illumination state of the LEDs in its eyes.

[0614] Figure 8 shows an example of the main functions of the data processing device 12 and the robot 414. As shown in Figure 8, the data processing device 12 performs specific processing using the processor 28. The storage 32 stores the specific processing program 56.

[0615] The specific processing program 56 is an example of a "program" relating to the technology of this disclosure. The processor 28 reads the specific processing program 56 from the storage 32 and executes the read specific processing program 56 on the RAM 30. The specific processing is realized by the processor 28 operating as a specific processing unit 290 in accordance with the specific processing program 56 executed on the RAM 30.

[0616] The storage 32 stores the data generation model 58 and the emotion identification model 59. The data generation model 58 and the emotion identification model 59 are used by the identification processing unit 290.

[0617] In robot 414, the processor 46 performs the reception output processing. The storage 50 stores the reception output program 60. The processor 46 reads the reception output program 60 from the storage 50 and executes the read reception output program 60 on the RAM 48. The reception output processing is realized by the processor 46 operating as a control unit 46A according to the reception output program 60 executed on the RAM 48.

[0618] Next, the specific processing performed by the specific processing unit 290 of the data processing device 12 will be described. In the following description, the data processing device 12 will be referred to as the "server" and the robot 414 as the "terminal".

[0619] The system according to the present invention has the function of acquiring voice and video data using a specific communication device and performing emotion analysis. To achieve this, a server, a terminal, and a user must cooperate in operation.

[0620] First, the device acquires audio data from the user's surroundings using a microphone and simultaneously acquires video data using a camera. This enables real-time data collection. The collected data is sent to a server via a security protocol, and the server receives the data. Next, the server activates an advanced algorithm for sentiment analysis based on the collected audio and video data. This algorithm utilizes natural language processing and computer vision technologies to analyze the emotional state.

[0621] Specifically, the server extracts factors such as voice tone, pitch, and speed from audio data and uses this information to infer emotional states. It also analyzes facial expressions, eye movements, and cheek redness from video data to obtain additional emotional information. By integrating these results, the server then performs the final emotion identification.

[0622] The server then visualizes the analysis results and returns them to the terminal as data. The terminal intuitively displays this data on a dashboard. For example, it can simultaneously monitor the emotions of multiple participants during a meeting and inform the user of each participant's emotional state. The user reviews the dashboard and adjusts their approach to the conversation as needed.

[0623] Furthermore, the server receives feedback information provided by users and uses it to improve the analysis model. This allows the system to improve its accuracy over time, enabling more reliable sentiment analysis. This feedback loop is a crucial element in continuously optimizing the system's performance.

[0624] The following describes the processing flow.

[0625] Step 1:

[0626] The device uses a microphone to collect audio data, including ambient sounds, and simultaneously captures the user's video data with a camera. This allows for the collection of emotion-related information, such as specific utterances and changes in facial expressions.

[0627] Step 2:

[0628] The device sends collected audio and video data to the server via a secure protocol. During transmission, the data is encrypted to protect privacy.

[0629] Step 3:

[0630] To analyze the data received by the server in real time, natural language processing algorithms are applied to the audio data to extract speech features. These extracted features include speed, emphasis, and intonation.

[0631] Step 4:

[0632] The server uses computer vision algorithms to analyze video data and identify emotion-related features from the user's facial expressions and movements. For example, it can infer emotions such as joy, anger, sadness, or happiness based on smiles and eyebrow furrowing.

[0633] Step 5:

[0634] The server integrates the audio and video analysis results and uses statistical models and machine learning algorithms to comprehensively identify the user's emotional state. It then outputs a specific emotional category (e.g., joy, anxiety, surprise).

[0635] Step 6:

[0636] The server processes the sentiment analysis results as visual data and generates a dataset for the dashboard. This includes visual elements such as color coding and icons.

[0637] Step 7:

[0638] The device acquires visual data and displays it in real time on the user's dashboard. The user can observe this and take appropriate action based on the situation.

[0639] Step 8:

[0640] When a user provides feedback based on the system's analysis results, the terminal receives this information and sends it to the server. This feedback is used for the continuous training of the analysis model, contributing to improving the system's accuracy.

[0641] (Example 1)

[0642] Next, we will describe Example 1. In the following description, the data processing device 12 will be referred to as the "server" and the robot 414 as the "terminal".

[0643] When analyzing human emotions from audio and video data, there is a challenge in accurately identifying emotions using analysis based on a single source of information. In addition, there is a problem in that the use of feedback to effectively present analysis results to users and improve the accuracy of the system is insufficient.

[0644] The identification process performed by the identification processing unit 290 of the data processing device 12 in Example 1 is realized by the following means.

[0645] In this invention, the server includes acquisition means for acquiring audio and video signals, analysis means for analyzing emotional states based on the audio and video signals using natural language processing and image analysis techniques, and integration means for integrating the analysis results from the analysis means and identifying emotions. This enables accurate emotion identification utilizing multimodal information and improved accuracy based on user feedback.

[0646] "Audio signals" are data obtained by converting sound vibrations acquired from individuals or their environment into a format that can be electrically transmitted.

[0647] A "video signal" is data obtained by converting visual information acquired by a visual device such as a camera into a format that can be transmitted electronically.

[0648] "Natural language processing" is a technology that uses computers to process and analyze human language, and is used to recognize emotions and intentions from speech and text.

[0649] "Image analysis technology" refers to computational techniques that extract features from video data to understand the state of objects and people.

[0650] "Emotional state" refers to the emotions an individual experiences at a particular moment, and can be inferred from their tone of voice and facial expressions.

[0651] "Integration methods" refer to processes and techniques for consolidating analytical results obtained from multiple sources into a single conclusion.

[0652] An "analytical model" is a set of algorithms and computational procedures used to analyze data, which are learned to recognize specific patterns or meanings.

[0653] "Feedback" refers to evaluations and information provided by users to improve the performance of a system.

[0654] "Improving accuracy" refers to increasing the performance of a system, specifically how closely its output matches real-world conditions.

[0655] This invention relates to a system that identifies emotional states by analyzing audio and video signals and provides the user with the analysis results. To implement the invention, a microphone for acquiring audio signals, a camera for acquiring video signals, a server for data processing, and a terminal for presenting these to the user are required.

[0656] terminal

[0657] The device acquires audio signals using microphones placed around the user and video signals using a camera. This data is transmitted to a server in real time. In particular, the device performs acoustic processing to remove ambient noise and allow the user to focus on the conversation. The camera also uses a face detection algorithm to facilitate analysis of the user's facial expressions.

[0658] server

[0659] The server receives audio and video signals transmitted from the terminal. Audio signals are analyzed using natural language processing techniques to extract features such as tone, pitch, and speed. Video signals are analyzed using image analysis techniques to extract facial movements and changes in expression. These processes utilize algorithms powered by generative AI models to estimate emotional states. Based on the analysis results, the server identifies each user's emotions and converts them into a visualizeable data format.

[0660] User

[0661] Users visually review the analysis results sent back from the server via a dashboard on their device. This allows them to understand the emotional state of other participants during the meeting. Users provide feedback to the server via their device, and this feedback is used to improve the analysis model.

[0662] For example, if a participant is feeling anxious during an online meeting, their name and emotional state will be displayed on the device's dashboard. A prompt such as, "Analyze the emotional states of current meeting participants in real time and display any participants who appear anxious," allows for a quick response.

[0663] The flow of the specific processing in Example 1 will be explained using Figure 11.

[0664] Step 1:

[0665] The device acquires audio signals through microphones placed around the user and collects video signals using a camera. The inputs are ambient audio and visual data. The audio signal undergoes acoustic processing to remove unwanted ambient noise. The video signal uses a face detection algorithm to pinpoint the location of faces. The output of this processing is clean audio data and video data with faces detected.

[0666] Step 2:

[0667] The terminal transmits the processed audio and video data to the server via a security protocol. Here, the data is encrypted before transmission to ensure privacy and security. The input is the output data from the previous step, and the output is encrypted audio and video data.

[0668] Step 3:

[0669] The server receives the transmitted data and applies natural language processing techniques to the audio signal. Specifically, it extracts tone, pitch, and speed, and uses these features to estimate emotion. The input is encrypted audio data, and the output is an emotion estimation score.

[0670] Step 4:

[0671] The server analyzes video data and detects features such as facial expressions and movements. Using advanced computer vision technology, it quantifies various facial expressions and estimates possible emotional states. The input is encrypted video data, and the output is an emotion estimation score based on facial expressions.

[0672] Step 5:

[0673] The server integrates estimated scores obtained from audio and video and performs final emotion identification using a generative AI model. The resulting overall emotion score is stored in a database and converted into a data format for visualization. The input is the estimated score for each emotion, and the output is a visualized emotion identification result.

[0674] Step 6:

[0675] The server sends emotion recognition data to the terminal, which displays it on a dashboard. Users refer to the information displayed on the dashboard to adjust their approach to meetings and conversations. The input is a visualized emotion recognition result, and the output is an intuitive interface display presented to the user.

[0676] Step 7:

[0677] The user inputs feedback on the system's sentiment analysis results into a terminal. The server receives the feedback and uses it to improve the analysis model in the generated AI model. The input is the user's feedback, and the output is the updated analysis model.

[0678] (Application Example 1)

[0679] Next, we will explain Application Example 1. In the following explanation, the data processing device 12 will be referred to as the "server" and the robot 414 as the "terminal".

[0680] In modern customer service, understanding individual customer emotions in real time and improving service quality is becoming increasingly important. However, traditional methods make it difficult to accurately interpret emotions from customers' facial expressions and voices, hindering service improvement. To solve this problem, a system is needed that utilizes acoustic and visual information to analyze individual customer emotions in real time and respond quickly.

[0681] The specific processing performed by the specific processing unit 290 of the data processing device 12 in Application Example 1 is realized by the following means.

[0682] In this invention, the server includes means for acquiring acoustic and video information via communication equipment, computation means for analyzing emotions based on the acoustic and video information, and means for visualizing the analysis results and providing them to the individual using a display device. This enables service staff to immediately grasp the customer's emotions and provide appropriate service.

[0683] "Communication equipment" refers to devices used to acquire and transmit or receive audio and video information externally.

[0684] "Acoustic information" refers to data related to speech and tone, including characteristics such as the tone and speed of the speaker's voice.

[0685] "Visual information" refers to data related to images or videos that is used to analyze facial expressions and movements.

[0686] "Emotion" refers to an expression of an individual's emotional state or psychological reaction.

[0687] The "computation means" is a function that analyzes the acquired acoustic and visual information and performs processing to identify emotions.

[0688] A "display device" refers to a screen device or glasses-type information device that visually presents analyzed emotions to an individual.

[0689] "Visualization" is the process of representing data visually and providing it in an easily understandable format.

[0690] This invention aims to realize a system that utilizes communication equipment to collect acoustic and visual information and perform emotional analysis in real time. The user wears glasses-type information devices to acquire acoustic and visual information obtained from interactions with customers. These devices use built-in microphones and cameras to collect ambient sounds and customer facial expressions in real time. The collected data is transmitted to a server via Bluetooth or Wi-Fi.

[0691] The server performs emotion analysis using TensorFlow and OpenCV based on acquired acoustic and video information. It analyzes speech tone, speed, and pitch from the acoustic information, and evaluates facial features and movements from the video information. The resulting emotion data is then comprehensively evaluated using computer vision and natural language processing technologies.

[0692] The analyzed emotional data is visualized in real time on the user's glasses-type information device display. As a result, the user can quickly grasp the customer's emotional state and provide appropriate service. Furthermore, feedback information is sent to the server, and the accuracy is continuously improved by refining the emotional analysis algorithm. This feedback is obtained through manual input and automatic detection, contributing to the improvement of the system's performance.

[0693] As a concrete example, when a store clerk is introducing a product to a customer, if anxiety is detected from the customer's facial expression and voice, the display on the glasses will show instructions such as, "The customer is feeling anxious. Please provide reassuring information." This information is then analyzed by a generative AI model using prompts such as, "Analyze how customer A feels about the new product. Audio and video samples are available."

[0694] Through the above, it is possible to improve customer satisfaction in customer service operations.

[0695] The flow of a specific process in Application Example 1 will be explained using Figure 12.

[0696] Step 1:

[0697] The device uses a microphone and camera to acquire ambient acoustic and visual information. The input here is real-time collected audio and image data. The device temporarily stores this data and prepares it for transmission to a server via Bluetooth or Wi-Fi.

[0698] Step 2:

[0699] The server receives audio and video information transmitted from the terminal. Based on this input data, it performs data preprocessing. Audio data is subjected to noise reduction filtering and processed to make the audio signal clearer. Video data is filtered to extract features important for image recognition. This results in a clean dataset ready for analysis.

[0700] Step 3:

[0701] The server performs emotion analysis using TensorFlow and OpenCV with pre-processed data. Clean audio and video datasets are used as input. Tone, pitch, and speed are detected from the audio data, and facial features and movements are detected from the video data, and data calculations are performed. The analyzed emotion data is output, clarifying the user's emotional state.

[0702] Step 4:

[0703] The analysis results are visualized by the server and sent back to the terminal in real time. The input is analyzed emotional data, and visualized information is generated based on it. The user can check the analysis results on the display of a glasses-type information device. This display shows icons and text messages indicating the customer's emotional state, and prompts the user to take action.

[0704] Step 5:

[0705] User feedback is sent to the server via the terminal. Based on the feedback received as input, the server updates its emotion analysis algorithm to improve its accuracy. The processing of feedback information is supported by specific generative AI models and prompt statements. This allows the system to continuously provide more sophisticated analysis results.

[0706] Furthermore, an emotion engine that estimates the user's emotions may be incorporated. That is, the identification processing unit 290 may use the emotion identification model 59 to estimate the user's emotions and perform identification processing using the user's emotions.

[0707] This invention realizes a system that incorporates an emotion engine for analyzing audio and video data and recognizing the user's emotions. The system consists of a server, a terminal, and the emotion engine, which work together to analyze the user's emotional state in real time.

[0708] First, the device records the user's voice using a microphone device and simultaneously acquires video data using a camera device. This audio and video data is then transmitted to the server, timestamped as it goes. The server processes this data using an emotion engine.

[0709] The emotion engine extracts features such as voice intonation and speed by applying natural language processing techniques to audio data. Furthermore, it uses computer vision technology to analyze changes in facial expressions and eye movements from video data. Based on this information, the emotion engine estimates multiple emotional states of the user.

[0710] As a concrete example of its application, consider communication in a confined space during a meeting. The terminal records the voice and facial expressions of meeting participants and immediately transmits them to the server. The emotion engine evaluates when participants are feeling joy or anxiety and visualizes this as an analysis result. The server displays the analysis results on the terminal's dashboard and presents them to the user.

[0711] Users can refer to this dashboard to adjust the meeting's progress and facilitate smoother dialogue. Furthermore, incorporating user feedback into the emotion engine optimizes model training and further improves prediction accuracy. By repeating this process, the system can evolve continuously over the long term, increasing its practical usability.

[0712] The following describes the processing flow.

[0713] Step 1:

[0714] The device uses a microphone and camera to collect the user's voice and video data. Voice data includes the speaker's tone, intonation, and speed. Video data includes visual information such as facial expressions and eye movements.

[0715] Step 2:

[0716] The device transmits collected audio and video data to the server in real time. The data is sent in a stream format to minimize latency.

[0717] Step 3:

[0718] The server receives the transmitted data and analyzes it using an emotion engine. For audio data, natural language processing techniques are applied to extract voice features and identify parameters that provide clues to emotion.

[0719] Step 4:

[0720] The server analyzes the video data using computer vision technology to detect changes in facial expressions and gaze. This reveals nonverbal emotional elements derived from the video.

[0721] Step 5:

[0722] The server's emotion engine integrates audio and video data and utilizes multiple emotion models to infer the overall emotional state. For example, it may detect "interest" and "anxiety" simultaneously.

[0723] Step 6:

[0724] The server visualizes the analysis results and generates an intuitive dashboard dataset. This dataset includes color-coded graphs and sentiment icons.

[0725] Step 7:

[0726] The device receives the dataset and displays it on the user's dashboard. This allows the user to understand their emotional state in real time.

[0727] Step 8:

[0728] Users can adjust their responses or change the direction of discussions based on the information on the dashboard. Furthermore, the device sends feedback received from users to the server, which is then incorporated into the emotion engine's machine learning model.

[0729] Step 9:

[0730] The server analyzes the feedback information and updates the emotion engine to improve analysis accuracy. Through this process, the system accumulates insights over time and becomes able to recognize user emotions more accurately.

[0731] (Example 2)

[0732] Next, we will describe Example 2. In the following description, the data processing device 12 will be referred to as the "server" and the robot 414 as the "terminal".

[0733] In recent years, with the advancement of information and communication technologies using voice and video, accurately understanding users' emotions in real time has become crucial. However, conventional systems often analyze voice and video information separately, which limits the accuracy of emotional state recognition. Furthermore, there is a lack of mechanisms to effectively utilize user feedback to improve the accuracy of emotion analysis models. Therefore, there is a need for a system that can analyze emotions with high accuracy and continuously optimize learning.

[0734] The identification process performed by the identification processing unit 290 of the data processing device 12 in Example 2 is realized by the following means.

[0735] In this invention, the server includes means for collecting user acoustic and video information via communication means, means for applying natural language processing techniques to extract features based on the acoustic information, and means for applying computer vision techniques to analyze facial expressions and gaze based on the video information. This enables integrated analysis of audio and video data, allowing for more accurate estimation of the user's emotional state. Furthermore, by visualizing the analysis results and presenting them to the user, feedback can be received to optimize the learning of the emotion analysis model.

[0736] "Communication means" refers to a device or mechanism for collecting audio and video information from a user and transferring it to a server.

[0737] "Acoustic information" refers to data obtained from the user's voice and is used to estimate their emotional state.

[0738] "Visual information" refers to data that records changes in the user's face and facial expressions, and is used to estimate their visual emotional state.

[0739] "Natural language processing technology" refers to techniques for converting acoustic information into text and extracting features such as intonation and speed.

[0740] "Computer vision technology" refers to techniques that analyze video information to detect facial expressions and eye movements.

[0741] An "emotion analysis model" is an algorithm or data model that estimates a user's emotional state based on acoustic and visual information.

[0742] "Visualization" is the process of displaying analysis results in the form of graphs, charts, and other formats to allow users to understand them intuitively.

[0743] "Feedback" refers to opinions and evaluations obtained from users, which are used to improve the accuracy of the model and optimize its learning process.

[0744] The system for implementing this invention mainly consists of a terminal, a server, and an emotion analysis engine. This system analyzes the user's emotional state using acoustic and visual information.

[0745] First, the device uses a microphone device and a camera device to simultaneously collect the user's voice and video. The hardware used includes high-precision microphone devices (e.g., typical audio recording devices) and high-resolution cameras (e.g., typical video recording devices). The audio and video information is time-stamped and transmitted to the server in real-time or near real-time.

[0746] The server performs advanced processing on the received data. Natural language processing techniques are used for acoustic information, and common acoustic analysis software such as the Google Cloud Speech-to-Text API is used to convert speech into text data, further extracting features such as intonation and speed. This analysis provides a means to estimate the user's emotional state.

[0747] Furthermore, the server applies computer vision technology to the video information, using common video analysis software such as OpenCV and TensorFlow to analyze the user's facial expressions and eye movements. This allows for the acquisition of emotional information that can be gleaned from the video.

[0748] The server integrates features obtained from audio and video and estimates the user's emotional state based on an emotion analysis model. The estimation results are visualized and displayed on the device's dashboard in the form of graphs and charts. This allows the user to intuitively understand the analysis results.

[0749] Furthermore, users contribute to optimizing the sentiment analysis model by providing feedback. This feedback information is fed into a continuous learning process, making it possible to improve prediction accuracy.

[0750] A concrete example is the ability to monitor the emotional state of each participant in a meeting in real time and adjust the meeting's progress as needed. A possible prompt in this scenario might be, "Please tell me how to assess participants' emotional states in real time and ensure the meeting runs smoothly." This system would enable better communication and decision-making.

[0751] The flow of the specific processing in Example 2 will be explained using Figure 13.

[0752] Step 1:

[0753] The device records the user's voice using a microphone device and acquires video using a camera device. The input is real-time audio and video. This data is timestamped and synchronized. The output is timestamped audio and video data, ready to be sent to the server immediately.

[0754] Step 2:

[0755] The terminal transmits the collected data to the server via a high-speed communication network. The input is the time-stamped audio and video data collected in step 1. The output is the audio and video data received by the server. This process enables real-time analysis of the data.

[0756] Step 3:

[0757] The server processes the audio data using acoustic analysis software such as the Google Cloud Speech-to-Text API. The input is the acoustic information sent to the server. Specifically, the audio is converted into text, and features related to intonation and speed are extracted. The output of this process is a feature vector for estimating the user's emotional state.

[0758] Step 4:

[0759] The server processes video data using tools such as OpenCV and TensorFlow. The input is video information sent to the server. Specifically, it analyzes facial expressions and eye movements from the video and extracts feature information obtained from this analysis. The output is a dataset containing this aggregated feature information.

[0760] Step 5:

[0761] The server integrates features obtained from audio and video and estimates the user's emotional state using an emotion analysis model. The input is the output from steps 3 and 4. The emotion analysis model utilizes a generative AI model to calculate the probability of various emotions. The output is the estimated result indicating the user's emotional state.

[0762] Step 6:

[0763] The server visualizes the estimation results and displays them on the terminal's dashboard in the form of a graph or chart. The input is the estimated emotional state obtained in step 5. The output is the analysis results provided in a visually understandable format for the user. This information serves as foundational data for the user to intuitively understand the analysis results and provide appropriate feedback.

[0764] Step 7:

[0765] The user reviews the analysis results presented on the dashboard and provides feedback to the system. The input is the analysis results presented to the user. The output is feedback data to improve the accuracy of the sentiment analysis model. This facilitates the continuous learning and optimization of the model.

[0766] (Application Example 2)

[0767] Next, we will explain application example 2. In the following explanation, the data processing device 12 will be referred to as the "server" and the robot 414 as the "terminal".

[0768] In recent years, brick-and-mortar stores have been required to understand customer satisfaction in real time and provide services accordingly. However, traditional methods have the problem of not being able to immediately perceive customers' emotions. There is a need to provide an effective means for employees to accurately and quickly evaluate customers' facial expressions and voices and directly link that evaluation to service improvement.

[0769] The specific processing performed by the specific processing unit 290 of the data processing device 12 in Application Example 2 is realized by the following means.

[0770] In this invention, the server includes means for acquiring audio and video data via a communication device, processing means for performing sentiment analysis based on the audio and video data, and means for visualizing the analysis results and presenting them to the user using an output device. This enables real-time analysis of customer sentiment, allowing employees to immediately provide services based on that analysis.

[0771] A "communication device" is a device used to acquire audio and video data and transfer it to a server.

[0772] "Voice data" refers to digital information about human speech collected using voice acquisition devices such as microphones.

[0773] "Video data" refers to digital data based on visual information, acquired through video acquisition devices such as cameras.

[0774] "Emotional analysis" is a process that estimates a user's emotional state using audio and video data.

[0775] "Processing means" refers to means consisting of software or hardware for performing sentiment analysis and visualizing the data.

[0776] An "output device" is a device such as a display that visually presents the analysis results from the server to the user.

[0777] "A means of analyzing customer emotions in real time from facial expressions and voice and presenting them to workers" refers to a technology and means for instantly evaluating the emotional state of customers in physical stores and providing feedback on the results to workers.

[0778] "Feedback information" refers to evaluations and comments from users regarding sentiment analysis, and is data used to improve the analysis model.

[0779] In implementing this invention, a system will be constructed to perform sentiment analysis in customer-employee interactions in physical stores. This will be achieved using a server, terminals, and a sentiment analysis engine.

[0780] The server receives audio and video data from terminals via a communication device. Terminals, such as smart glasses or smartphones, have built-in microphones and cameras to capture customer audio and video in real time. After receiving the data, the server processes it using an emotion analysis engine. For audio data, natural language processing techniques are applied to analyze properties such as intonation and speed, and for video data, computer vision techniques are used to capture changes in facial expressions and gaze.

[0781] The data is first processed by a Python-based program. Audio data has its features extracted using the Librosa library, and video data is analyzed using OpenCV and a TensorFlow machine learning model. The analysis results are then visually displayed using an output device, for example, on the display of smart glasses worn by employees.

[0782] As a concrete example of this system, when a customer enters a store, the staff can instantly analyze whether the customer is smiling when they greet them with "Welcome." If a positive emotion is detected from the customer's expression, a message such as "Satisfied" will appear on the staff member's glasses, allowing for immediate improvement in the quality of service.

[0783] An example of a prompt message would be, "Please tell me the steps to develop an application that analyzes customer facial expressions and voice in real time and provides emotional feedback on a smart device."

[0784] The flow of a specific process in Application Example 2 will be explained using Figure 14.

[0785] Step 1:

[0786] The device uses smart glasses or a smartphone to record audio data with a microphone and capture video data with a camera. This provides instant digital information about the customer's voice tone and facial expressions.

[0787] Step 2:

[0788] The terminal transfers the acquired audio and video data to the server in real time. During this process, the data is time-stamped and synchronized. The input data transmitted consists of recorded audio and video files.

[0789] Step 3:

[0790] The server analyzes the received audio data using the Librosa library. This process extracts features such as intonation and speed from the speech and generates feature data based on these features. The audio waveform and spectral information are processed, and numerical data suggesting emotional state is output.

[0791] Step 4:

[0792] The server analyzes video data using the OpenCV library. Specifically, it recognizes facial expressions and eye movements, and converts this visual information into numerical data. It detects facial landmarks from video frames and identifies features that indicate emotion. The output is a parameter set indicating emotion.

[0793] Step 5:

[0794] The emotion analysis engine on the server integrates audio feature data and video parameter sets, and performs analysis using a machine learning model. TensorFlow is used to estimate the customer's emotional state through the integrated data. As a result of this integration, an emotion label and its confidence level are output.

[0795] Step 6:

[0796] The user (store staff) visually checks the analysis results received from the server on the display of smart glasses. This allows them to understand the customer's emotional state in real time. The output displays a specific emotional state (e.g., "satisfied," "dissatisfied").

[0797] The specific processing unit 290 transmits the result of the specific processing to the robot 414. In the robot 414, the control unit 46A causes the speaker 240 and the controlled object 443 to output the result of the specific processing. The microphone 238 acquires audio indicating user input for the result of the specific processing. The control unit 46A transmits the audio data indicating user input acquired by the microphone 238 to the data processing unit 12. In the data processing unit 12, the specific processing unit 290 acquires the audio data.

[0798] Data generation model 58 is a type of so-called generative AI (Artificial Intelligence). One example of data generation model 58 is ChatGPT (Internet search<URL: https: / / openai.com / blog / chatgpt> ), Gemini (Internet search) <url: https: gemini.google.com ?hl="ja">Examples of generative AI include the following. The data generation model 58 is obtained by performing deep learning on a neural network. The data generation model 58 is input with prompts containing instructions, and with inference data such as audio data representing speech, text data representing text, and image data representing images. The data generation model 58 infers from the input inference data according to the instructions indicated by the prompts, and outputs the inference results in data formats such as audio data and text data. Here, inference refers to, for example, analysis, classification, prediction, and / or summarization.

[0799] In the above embodiment, an example was given in which the specific processing is performed by the data processing device 12, but the technology of this disclosure is not limited thereto, and the specific processing may also be performed by the robot 414.

[0800] Furthermore, the emotion identification model 59, acting as an emotion engine, may determine the user's emotion according to a specific mapping. Specifically, the emotion identification model 59 may determine the user's emotion according to a specific mapping, which is an emotion map (see Figure 9). Similarly, the emotion identification model 59 may also determine the robot's emotion, and the identification processing unit 290 may perform identification processing using the robot's emotion.

[0801] Figure 9 shows an emotion map 400 in which multiple emotions are mapped. In the emotion map 400, emotions are arranged in concentric circles radiating from the center. The closer to the center of the concentric circles, the more primitive the emotions are located. Further out of the concentric circles, emotions representing states and actions arising from mental states are located. Emotion is a concept that includes feelings and mental states. On the left side of the concentric circles, emotions that are generally generated from reactions occurring in the brain are located. On the right side of the concentric circles, emotions that are generally induced by situational judgment are located. Above and below the concentric circles, emotions that are generally generated from reactions occurring in the brain and induced by situational judgment are located. In addition, the emotion of "pleasure" is located on the upper side of the concentric circles, and the emotion of "displeasure" is located on the lower side. Thus, in the emotion map 400, multiple emotions are mapped based on the structure in which emotions arise, and emotions that are likely to occur simultaneously are mapped close together.

[0802] These emotions are distributed at the 3 o'clock position on the Emotion Map 400, and usually fluctuate between feelings of security and anxiety. In the right half of the Emotion Map 400, situational awareness takes precedence over internal feelings, resulting in a calm impression.

[0803] The inside of the Emotion Map 400 represents inner thoughts, while the outside represents actions. Therefore, the further you go from the outside of the Emotion Map 400, the more visible (expressed in actions) your emotions become.

[0804] Here, human emotions are based on various balances, such as posture and blood sugar levels. When these balances deviate from the ideal, it results in discomfort, and when they approach the ideal, it results in pleasure. Similarly, in robots, cars, motorcycles, etc., emotions can be created based on various balances, such as posture and battery level. When these balances deviate from the ideal, it results in discomfort, and when they approach the ideal, it results in pleasure. The emotion map can be generated, for example, based on Dr. Mitsuyoshi's emotion map (Research on a system for analyzing brain physiological signals of speech emotion recognition and emotion, Tokushima University, doctoral dissertation: https: / / ci.nii.ac.jp / naid / 500000375379). The left half of the emotion map contains emotions belonging to a region called "response," where sensation is dominant. The right half of the emotion map contains emotions belonging to a region called "situation," where situational awareness is dominant.

[0805] The emotion map defines two emotions that promote learning. One is the emotion around the middle of the negative "repentance" and "reflection" on the situation side. In other words, it is when the robot experiences negative emotions such as "I never want to feel this way again" or "I don't want to be scolded again." The other is the emotion around the positive "desire" on the reaction side. In other words, it is when the robot has positive feelings such as "I want more" or "I want to know more."

[0806] The emotion identification model 59 inputs user input into a pre-trained neural network, obtains emotion values ​​representing each emotion shown in the emotion map 400, and determines the user's emotion. This neural network is pre-trained based on multiple training data sets, which are combinations of user input and emotion values ​​representing each emotion shown in the emotion map 400. Furthermore, this neural network is trained so that emotions located close together have similar values, as shown in the emotion map 900 in Figure 10. Figure 10 shows an example where multiple emotions such as "reassured," "calm," and "confident" have similar emotion values.

[0807] The above description primarily focuses on the functions of the data processing device 12 in relation to this disclosure. However, the system related to this disclosure is not necessarily implemented on a server. The system related to this disclosure may be implemented as a general information processing system. This disclosure may be implemented, for example, as a software program that runs on a personal computer or as an application that runs on a smartphone. The method related to this disclosure may be provided to users in SaaS (Software as a Service) format.

[0808] In the above embodiment, an example was given in which a specific process is performed by a single computer 22. However, the technology of this disclosure is not limited thereto, and a distributed processing of the specific process may be performed by multiple computers, including computer 22. For example, a data generation model 58 may be provided in an external device of the data processing device 12, and the external device may generate data according to the input data.

[0809] In the above embodiment, an example was given in which the specific processing program 56 is stored in the storage 32, but the technology of this disclosure is not limited thereto. For example, the specific processing program 56 may be stored in a portable, computer-readable, non-temporary storage medium such as a USB (Universal Serial Bus) memory. The specific processing program 56 stored in the non-temporary storage medium is installed in the computer 22 of the data processing device 12. The processor 28 executes specific processing according to the specific processing program 56.

[0810] Alternatively, the specific processing program 56 may be stored in a storage device such as a server connected to the data processing device 12 via the network 54, and the specific processing program 56 may be downloaded and installed on the computer 22 in response to a request from the data processing device 12.

[0811] Furthermore, it is not necessary to store the entirety of the specific processing program 56 in a storage device such as a server connected to the data processing device 12 via the network 54, or to store the entirety of the specific processing program 56 in the storage 32; it is acceptable to store only a portion of the specific processing program 56.

[0812] The following types of processors can be used as hardware resources to perform specific processing. Examples of processors include a CPU, a general-purpose processor that functions as a hardware resource to perform specific processing by executing software, i.e., a program. Other examples of processors include dedicated electrical circuits, such as FPGAs (Field-Programmable Gate Arrays), PLDs (Programmable Logic Devices), or ASICs (Application Specific Integrated Circuits), which have circuit configurations specifically designed to perform specific processing. All of these processors have built-in or connected memory, and all of them perform specific processing by using memory.

[0813] The hardware resource that performs a specific process may consist of one of these various processors, or it may consist of a combination of two or more processors of the same or different types (for example, a combination of multiple FPGAs, or a combination of a CPU and an FPGA). Alternatively, the hardware resource that performs a specific process may consist of a single processor.

[0814] Examples of configurations using a single processor include, firstly, a configuration in which one or more CPUs and software are combined to form a single processor, and this processor functions as a hardware resource that performs a specific process. Secondly, there is a configuration using a processor that realizes the functions of the entire system, including multiple hardware resources that perform a specific process, on a single IC chip, as exemplified by SoCs (System-on-a-chip). In this way, a specific process is realized using one or more of the above types of processors as hardware resources.

[0815] Furthermore, the hardware structure of these various processors can more specifically utilize electrical circuits that combine circuit elements such as semiconductor devices. Also, the specific processing described above is merely an example. Therefore, it goes without saying that unnecessary steps can be deleted, new steps added, or the processing order rearranged, as long as it does not deviate from the main purpose.

[0816] The descriptions and illustrations presented above are detailed explanations of the technical aspects of this disclosure and are merely examples of the technical aspects. For example, the above descriptions of the structure, function, operation, and effect are examples of the structure, function, operation, and effect of the technical aspects of this disclosure. Therefore, it goes without saying that you may delete unnecessary parts, add new elements, or replace elements in the descriptions and illustrations presented above, as long as you do not deviate from the essence of the technical aspects of this disclosure. Furthermore, in order to avoid confusion and facilitate understanding of the technical aspects of this disclosure, explanations of common technical knowledge and the like that do not require special explanation to enable the implementation of the technical aspects of this disclosure have been omitted from the descriptions and illustrations presented above.

[0817] All documents, patent applications, and technical standards described herein are incorporated by reference to the same extent as if each individual document, patent application, and technical standard were specifically and individually noted as being incorporated by reference.

[0818] The following is further disclosed regarding the embodiments described above.

[0819] (Claim 1)

[0820] Means for acquiring audio data and video data via a communication device,

[0821] Processing means for performing emotion analysis based on the aforementioned audio data and video data,

[0822] A means of visualizing the analysis results and presenting them to the user using an output device,

[0823] A means of optimizing learning by receiving feedback information and updating the sentiment analysis model,

[0824] A system that includes this.

[0825] (Claim 2)

[0826] The system according to claim 1, comprising processing means for integrating multiple analytical models for analyzing multimodal data and enabling accurate identification of emotions.

[0827] (Claim 3)

[0828] The system according to claim 1, further comprising a feedback processing means for receiving feedback from a user and reflecting it in the emotion analysis model.

[0829] "Example 1"

[0830] (Claim 1)

[0831] means for acquiring audio signals and video signals,

[0832] An analysis means for analyzing emotional states using natural language processing and image analysis techniques based on the aforementioned audio and video signals,

[0833] The analysis results from the aforementioned analysis means are integrated, and an integration means for identifying emotions is provided,

[0834] A means for presenting visualized analysis results to the user via an interface,

[0835] An optimization method that improves the accuracy of sentiment analysis by receiving evaluation information from users and updating the analysis model,

[0836] A system that includes this.

[0837] (Claim 2)

[0838] The system according to claim 1, which integrates multiple technologies to analyze multimodal information and enables more precise emotion classification.

[0839] (Claim 3)

[0840] The system according to claim 1, which receives evaluation information from users as feedback and applies it to the analysis model to improve performance.

[0841] "Application Example 1"

[0842] (Claim 1)

[0843] Means for acquiring audio and video information via communication equipment,

[0844] A computation means for analyzing emotions based on the aforementioned acoustic and visual information,

[0845] A means of visualizing the analysis results and providing them to an individual using a display device,

[0846] A means of optimizing learning by receiving feedback information and updating the emotion analysis algorithm,

[0847] A means of supporting coping strategies by using glasses-type information devices to present an individual's emotional state in real time,

[0848] A system that includes this.

[0849] (Claim 2)

[0850] The system according to claim 1, comprising processing means that integrate multiple computational algorithms for analyzing multimodal data and enable accurate identification of emotions.

[0851] (Claim 3)

[0852] The system according to claim 1, comprising a feedback processing means for receiving feedback from an individual and reflecting it in the emotion analysis algorithm.

[0853] "Example 2 of combining an emotion engine"

[0854] (Claim 1)

[0855] A means for collecting user audio and video information via communication means,

[0856] A means for applying natural language processing techniques based on the aforementioned acoustic information to extract features,

[0857] A means for applying computer vision technology based on the aforementioned video information to analyze facial expressions and gaze,

[0858] A means for estimating an emotional state by combining features obtained from the aforementioned acoustic and visual information,

[0859] A means for visualizing the analysis results and presenting them to a person using a display device,

[0860] A means for optimizing learning by receiving feedback information and updating the emotion analysis model,

[0861] A system that includes this.

[0862] (Claim 2)

[0863] The system according to claim 1, comprising processing means that integrate multiple analytical techniques for analyzing multimodal data and enable accurate identification of emotions.

[0864] (Claim 3)

[0865] The system according to claim 1, further comprising a feedback processing means for receiving feedback from a person and reflecting it in the emotion analysis model.

[0866] "Application example 2 when combining with an emotional engine"

[0867] (Claim 1)

[0868] Means for acquiring audio data and video data via a communication device,

[0869] Processing means for performing emotion analysis based on the aforementioned audio data and video data,

[0870] A means of visualizing the analysis results and presenting them to the user using an output device,

[0871] A method for analyzing customer emotions in real time from their facial expressions and voice and presenting them to workers,

[0872] A means of optimizing learning by receiving feedback information and updating the sentiment analysis model,

[0873] A system that includes this.

[0874] (Claim 2)

[0875] The system according to claim 1, comprising processing means for integrating multiple analytical models for analyzing multimodal data and enabling accurate identification of emotions.

[0876] (Claim 3)

[0877] The system according to claim 1, further comprising a feedback processing means for receiving feedback from a user and reflecting it in the emotion analysis model. [Explanation of Symbols]

[0878] 10, 210, 310, 410 Data Processing Systems 12 Data Processing Devices 14 Smart Devices 214 Smart Glasses 314 Headset-type terminal 414 Robots< / url:> < / url:> < / url:> < / url:>

Claims

1. Means for acquiring audio data and video data via a communication device, Processing means for performing emotion analysis based on the aforementioned audio data and video data, A means of visualizing the analysis results and presenting them to the user using an output device, A means of optimizing learning by receiving feedback information and updating the sentiment analysis model, A system that includes this.

2. The system according to claim 1, comprising processing means that integrate multiple analytical models for analyzing multimodal data and enable accurate identification of emotions.

3. The system according to claim 1, further comprising a feedback processing means for receiving feedback from a user and reflecting it in the emotion analysis model.

Citation Information

Patent Citations

  • Persona chatbot control method and system

    JP2022180282A