System
A system that converts and analyzes voice data from meetings to identify emotional states and risk levels, facilitating timely interventions for improved employee health and productivity.
Patent Information
- Application Number
- JP2024116379
- Authority / Receiving Office
- JP · JP
- Patent Type
- Applications
- Current Assignee / Owner
- Filing Date
- 2024-07-19
- Publication Date
- 2026-01-29
AI Technical Summary
Existing methods fail to efficiently and systematically detect employee emotional and health risks during meetings, making it difficult to address declines in work performance and mental health issues promptly.
A system that acquires voice data, converts it to text, analyzes the content and emotional state, calculates risk levels, and sends notifications when thresholds are exceeded, enabling real-time monitoring and management of employee health.
Enables continuous tracking of employee work progress and emotional states, allowing for timely intervention to improve productivity and health management within organizations.
Smart Images

Figure 2026014905000001_ABST
Abstract
Description
[Technical Field]
[0001] The technology of the present disclosure relates to a system. [Background technology]
[0002] Patent document 1 discloses a persona chatbot control method performed by at least one processor, the method including the steps of receiving a user utterance, adding the user utterance to a prompt including an instruction sentence related to a description of the chatbot character, encoding the prompt, and inputting the encoded prompt into a language model to generate a chatbot utterance in response to the user utterance. [Prior art documents] [Patent documents]
[0003] [Patent Document 1] Japanese Patent Publication No. 2022-180282 Summary of the Invention [Problem to be solved by the invention]
[0004] As work style reform and mental health become increasingly important, it is extremely important to understand employees' work situations and mental and physical conditions in real time. However, in typical one-on-one or team meetings, it is difficult to efficiently obtain and analyze this information without direct observation by managers. Furthermore, there are no effective and systematic methods for detecting employee emotional and health risks in a timely manner and taking early action. This makes it difficult to detect declines in employee work performance and mental health issues early, negatively impacting the health and productivity of the entire organization. [Means for solving the problem]
[0005] To solve the above problems, a system is provided that includes a means for acquiring voice data, a means for converting voice data into text data, a means for analyzing the text data and extracting the content of the conversation, a means for identifying an emotional state from the text data, a means for calculating and visualizing a risk level based on the identified emotional state, and a means for sending a notification when the risk level exceeds a predetermined threshold.This system analyzes meeting voice data in real time, making it possible to constantly keep track of employees' work progress, consultation details, and emotional state.Furthermore, by calculating the risk level based on the emotional state and responding quickly when a high risk is detected, it is possible to manage employee health and improve productivity throughout the organization.
[0006] "Audio data" refers to digital or analog signal data for recording and transmitting audio.
[0007] "Text data" is data in the form of a character string converted from voice data, and is used for analyzing and processing the content of a conversation.
[0008] "Analysis" is the process of examining and breaking down data to understand its components and extract specific information.
[0009] "Conversational content" refers to the specific matters, events, and intended messages discussed during a meeting or conversation.
[0010] "Emotional state" refers to the speaker's mental state or emotion that is detected based on speech and text analysis.
[0011] "Risk level" is the degree of mental and physical risk calculated based on the identified emotional state.
[0012] "Visualization" is the process of displaying analyzed and calculated data in the form of graphs, charts, etc., to make it easier to understand.
[0013] A "notification" is an alert or message sent to interested parties when a particular event or condition occurs. [Brief explanation of the drawings]
[0014] [Figure 1] 1 is a conceptual diagram showing an example of the configuration of a data processing system according to a first embodiment. [Figure 2] 1 is a conceptual diagram showing an example of main functions of a data processing device and a smart device according to a first embodiment. [Figure 3] FIG. 10 is a conceptual diagram showing an example of the configuration of a data processing system according to a second embodiment. [Figure 4] FIG. 10 is a conceptual diagram showing an example of main functions of a data processing device and smart glasses according to a second embodiment. [Figure 5] FIG. 10 is a conceptual diagram showing an example of the configuration of a data processing system according to a third embodiment. [Figure 6] FIG. 11 is a conceptual diagram showing an example of main functions of a data processing device and a headset-type terminal according to a third embodiment. [Figure 7] FIG. 10 is a conceptual diagram showing an example of the configuration of a data processing system according to a fourth embodiment. [Figure 8] FIG. 10 is a conceptual diagram showing an example of main functions of a data processing device and a robot according to a fourth embodiment. [Figure 9] 1 shows an emotion map onto which multiple emotions are mapped. [Figure 10] 1 shows an emotion map onto which multiple emotions are mapped. [Figure 11] FIG. 3 is a sequence diagram showing a processing flow of the data processing system according to the first embodiment. [Figure 12] FIG. 10 is a sequence diagram showing the flow of processing in the data processing system in Application Example 1. [Figure 13] FIG. 10 is a sequence diagram showing the flow of processing in the data processing system according to the second embodiment when an emotion engine is combined. [Figure 14]FIG. 10 is a sequence diagram showing the flow of processing in the data processing system in Application Example 2 when an emotion engine is combined. DETAILED DESCRIPTION OF THE INVENTION
[0015] An example of an embodiment of a system according to the technology of the present disclosure will be described below with reference to the accompanying drawings.
[0016] First, the terms used in the following description will be explained.
[0017] In the following embodiments, a coded processor (hereinafter simply referred to as a "processor") may be a single arithmetic device or a combination of multiple arithmetic devices. Furthermore, a processor may be a single type of arithmetic device or a combination of multiple types of arithmetic devices. Examples of arithmetic devices include a CPU (Central Processing Unit), a GPU (Graphics Processing Unit), a GPGPU (General-Purpose computing on Graphics Processing Units), and an APU (Accelerated Processing Unit).
[0018] In the following embodiments, a coded RAM (Random Access Memory) is a memory in which information is temporarily stored and is used as a working memory by a processor.
[0019] In the following embodiments, the coded storage is one or more non-volatile storage devices that store various programs, various parameters, etc. Examples of non-volatile storage devices include flash memory (SSD (Solid State Drive)), magnetic disks (e.g., hard disks), and magnetic tapes.
[0020] In the following embodiments, a communication I / F (Interface) with a symbol is an interface including a communication processor, an antenna, etc. The communication I / F controls communication between multiple computers. Examples of communication standards applied to the communication I / F include wireless communication standards including 5G (5th Generation Mobile Communication System), Wi-Fi (registered trademark), Bluetooth (registered trademark), etc.
[0021] In the following embodiments, "A and / or B" is synonymous with "at least one of A and B." In other words, "A and / or B" means that it may be only A, only B, or a combination of A and B. Furthermore, in this specification, the same concept as "A and / or B" is also applied when three or more things are expressed connected by "and / or."
[0022] [First embodiment]
[0023] FIG. 1 shows an example of the configuration of a data processing system 10 according to the first embodiment.
[0024] 1, a data processing system 10 includes a data processing device 12 and a smart device 14. An example of the data processing device 12 is a server.
[0025] The data processing device 12 includes a computer 22, a database 24, and a communication I / F 26. The computer 22 is an example of a "computer" according to the technology of the present disclosure. The computer 22 includes a processor 28, a RAM 30, and a storage 32. The processor 28, the RAM 30, and the storage 32 are connected to a bus 34. The database 24 and the communication I / F 26 are also connected to the bus 34. The communication I / F 26 is connected to a network 54. Examples of the network 54 include a WAN (Wide Area Network) and / or a LAN (Local Area Network).
[0026] The smart device 14 includes a computer 36, a reception device 38, an output device 40, a camera 42, and a communication I / F 44. The computer 36 includes a processor 46, a RAM 48, and a storage 50. The processor 46, the RAM 48, and the storage 50 are connected to a bus 52. The reception device 38, the output device 40, and the camera 42 are also connected to the bus 52.
[0027] The reception device 38 includes a touch panel 38A, a microphone 38B, and the like, and receives user input. The touch panel 38A detects contact with an indicator (for example, a pen or a finger) to receive user input by the touch of the indicator. The microphone 38B detects the user's voice to receive user input by voice. The control unit 46A transmits data indicating the user input received by the touch panel 38A and the microphone 38B to the data processing device 12. In the data processing device 12, the specific processing unit 290 acquires the data indicating the user input.
[0028] The output device 40 includes a display 40A and a speaker 40B, and presents data to the user 20 by outputting the data in a form of expression that the user 20 can perceive (for example, audio and / or text). The display 40A displays visible information such as text and images in accordance with instructions from the processor 46. The speaker 40B outputs audio in accordance with instructions from the processor 46. The camera 42 is a compact digital camera equipped with an optical system including a lens, aperture, and shutter, and an imaging element such as a CMOS (Complementary Metal-Oxide-Semiconductor) image sensor or a CCD (Charge Coupled Device) image sensor.
[0029] The communication I / F 44 is connected to a network 54. The communication I / Fs 44 and 26 control the exchange of various information between the processor 46 and the processor 28 via the network 54.
[0030] FIG. 2 shows an example of the main functions of the data processing device 12 and the smart device 14.
[0031] 2, in the data processing device 12, a specific process is performed by the processor 28. A specific processing program 56 is stored in the storage 32. The specific processing program 56 is an example of a "program" according to the technology of the present disclosure. The processor 28 reads the specific processing program 56 from the storage 32 and executes the read specific processing program 56 on the RAM 30. The specific process is realized by the processor 28 operating as a specific processing unit 290 in accordance with the specific processing program 56 executed on the RAM 30.
[0032] The storage 32 stores a data generation model 58 and an emotion identification model 59. The data generation model 58 and the emotion identification model 59 are used by the identification processing unit 290.
[0033] In the smart device 14, the processor 46 performs the reception output process. The storage 50 stores a reception output program 60. The reception output program 60 is used in conjunction with the specific processing program 56 by the data processing system 10. The processor 46 reads the reception output program 60 from the storage 50 and executes the read reception output program 60 on the RAM 48. The reception output process is realized by the processor 46 operating as the control unit 46A in accordance with the reception output program 60 executed on the RAM 48.
[0034] Next, a description will be given of the specific processing performed by the specific processing unit 290 of the data processing device 12. In the following description, the data processing device 12 will be referred to as a "server" and the smart device 14 will be referred to as a "terminal."
[0035] The present invention relates to a system that starts by acquiring voice data, converts the voice data into text data, analyzes the text data to extract the content of the conversation and the emotional state, and calculates and visualizes the risk level. Below, we will explain in detail each component of this system and its operation.
[0036] System Embodiments
[0037] Acquiring audio data
[0038] When the device detects the start of a meeting, it starts capturing audio data in real time through the microphone, and sends the captured audio data to the server in a specified format.
[0039] Saving audio data
[0040] The server receives the voice data sent from the device and stores it in a designated database, which is used in the subsequent analysis process.
[0041] Speech-to-Text (STT)
[0042] The server calls the speech-to-text engine to convert the stored voice data into text data. For example, a speech saying, "The project is going well, but I've been feeling tired lately" is converted into text. This converted text data is temporarily stored and passed to the analysis engine.
[0043] Conversation content analysis
[0044] The server inputs the text data into a natural language processing (NLP) module and analyzes the content of the conversation. During this process, keywords related to the progress of work and the content of the consultation (e.g., "progress," "problem," "fatigue") are extracted. The extracted keywords and content are stored in a database and used for subsequent processing and report creation.
[0045] sentiment analysis
[0046] The server inputs the text data into a sentiment analysis module to analyze the emotions and tone of the conversation. During this process, emotional states such as "happiness," "sadness," "anger," and "fatigue" are identified. The emotional state is compared with past data, and the difference from the normal state is calculated. If an anomaly is detected, the information is stored in a database.
[0047] Risk Assessment
[0048] The server calculates mental and physical risk levels based on the results of emotion analysis. Using a big database, it compares current risk levels with past data and calculates the current risk level. The calculated risk level is visualized, and the impact and trends are displayed in graphs and dashboard format.
[0049] Feedback and Notifications
[0050] The server generates reports based on the results of sentiment analysis and risk assessment. If a risk signal is detected, a notification is sent to the employee and their manager. This notification can be sent via email or a chat app, for example. Periodic reports are also generated and distributed to managers.
[0051] Reviewing feedback
[0052] Users can check the feedback and reports sent from the server, which can then be used to take measures regarding employees' work status and health status, for example, suggesting that fatigued employees take time off.
[0053] Specific examples
[0054] scenario:
[0055] A progress meeting is held for a project, and a project member says, "I've been feeling tired lately."
[0056] How it works:
[0057] 1. The device collects the audio data of the meeting in real time and sends it to the server.
[0058] 2. The server receives the voice data and converts it into text data.
[0059] 3. The server analyzes the text data "The project is going well, but I've been feeling tired lately" and extracts the keyword "fatigue."
[0060] 4. The server uses the emotion analysis module to identify the "feeling of fatigue."
[0061] 5. The server calculates the "mental fatigue level" and determines that the risk is high.
[0062] 6. The server sends a "high fatigue risk" notification to the manager and employee.
[0063] 7. The user reviews the notification and takes action such as following up with the employee or suggesting time off.
[0064] The above is a specific embodiment of the present invention. This system can effectively manage the health of employees and improve productivity throughout the organization.
[0065] The processing flow will be explained below.
[0066] Step 1:
[0067] A user logs in to the system and enters information such as the date and time of the meeting, participants, purpose, etc. For example, information such as "October 5th, 10:00, progress meeting, participants: Mr. A, Mr. B" is registered in the system.
[0068] Step 2:
[0069] When the device detects the start of a meeting, it starts collecting audio data in real time through the microphone, and sends the collected audio data to the server in a specified format.
[0070] Step 3:
[0071] The server receives the voice data sent from the device and stores it in a database, which is used in the subsequent analysis process.
[0072] Step 4:
[0073] The server inputs the stored voice data into a Speech-to-Text (STT) engine and converts the voice into text data. For example, a speech saying, "The project is going well, but I've been feeling tired lately" is converted into text. This text data is temporarily stored.
[0074] Step 5:
[0075] The server inputs the text data into a natural language processing (NLP) module and analyzes the content of the conversation. During this process, keywords related to the progress of work and the content of the consultation (e.g., "progress," "problem," "fatigue") are extracted. The extracted keywords and content are stored in a database and used for subsequent processing and report creation.
[0076] Step 6:
[0077] The server inputs the text data into a sentiment analysis module to analyze the emotions and tone of the conversation. During this process, emotional states such as "happiness," "sadness," "anger," and "fatigue" are identified. The emotional state is compared with past data, and the difference from the normal state is calculated. If an anomaly is detected, the information is stored in a database.
[0078] Step 7:
[0079] The server calculates the mental and physical risk level based on the results of the emotion analysis. It uses a big database to compare it with past data and calculate the current risk level. The results are visualized, and the impact and trends are displayed in graphs and dashboard format.
[0080] Step 8:
[0081] The server generates reports based on the results of sentiment analysis and risk assessment. If a risk signal is detected, a notification is sent to the employee and their supervisor so that they can take action. This notification can be sent via email or a chat app, for example. Periodic reports are also generated and distributed to supervisors.
[0082] Step 9:
[0083] The user checks the feedback and reports sent from the server, which allows the user to take measures regarding the work situation and health status of employees. For example, the user can suggest that a fatigued employee take time off.
[0084] Example 1
[0085] Next, a description will be given of Example 1. In the following description, the data processing device 12 will be referred to as a "server" and the smart device 14 will be referred to as a "terminal."
[0086] There is a need to monitor employees' mental and physical health conditions in real time to prevent declines in work efficiency and productivity. However, conventional methods have made it difficult to properly analyze the content of comments and emotional states made during meetings, assess risks, and take prompt measures. The present invention aims to solve these problems and provide a system for effectively managing employees' health conditions.
[0087] The specific processing by the specific processing unit 290 of the data processing device 12 in the first embodiment is realized by the following means.
[0088] In this invention, the server includes means for acquiring voice data, means for converting the voice data into text data, means for analyzing the text data and extracting the content of the conversation, means for identifying an emotional state from the text data, means for calculating and visualizing a risk level based on the identified emotional state, means for sending a notification when the risk level exceeds a predetermined threshold, means for a terminal to detect the start of a meeting and send the voice data to the server in real time, means for the server to store the voice data in a database and convert the voice data into text data using a speech-to-text engine, means for the server to extract keywords related to work progress and consultation content from the text data using a natural language processing module, means for the server to analyze the emotional state using a sentiment analysis module, means for the server to calculate a risk level using a big data analysis framework and visualize the risk level using a data visualization tool, means for the server to generate a report and send a notification to the employee and their supervisor using a predetermined notification means, and means for a user to check the notification from the server and take measures. This enables real-time monitoring of employees' mental and physical health conditions and rapid visualization and notification of risks.
[0089] "Audio data" refers to data in which voice or acoustic information is recorded in digital format.
[0090] "Text data" refers to data obtained by converting voice data into character information.
[0091] A "natural language processing (NLP) module" is software and algorithms used to analyze text data and understand its content and meaning.
[0092] "Sentiment Analysis Module" means software and algorithms for identifying emotional states from text data.
[0093] "Risk level" is a calculated degree of risk based on the identified emotional state and other data.
[0094] "Visualization" is the presentation of data in a visual format such as a graph, chart, or dashboard.
[0095] "Notification" means sending a warning or information to the user when a specific condition is met.
[0096] A "terminal" is a device that acquires voice data and transmits it to a server, and includes devices such as personal computers and smartphones.
[0097] A "server" is a computer system for storing and processing audio data.
[0098] A "database" is a system that systematically stores and manages information.
[0099] A "Speech-to-Text engine" is software for converting voice data into text data.
[0100] "Work progress status" is information that indicates the progress of projects and tasks.
[0101] A "big data analytics framework" is software and a platform for analyzing large amounts of data and extracting meaningful information.
[0102] A "data visualization tool" is a tool for visually displaying data, and includes tools such as Grafana and Tableau.
[0103] A "report" is a document that summarizes the results of analysis and evaluation.
[0104] "User" means a person who uses the System to review feedback and notifications and take appropriate action.
[0105] "Detecting the start of a meeting" means that the system automatically recognizes the start of a meeting.
[0106] System Overview
[0107] This invention is a system that captures audio data during meetings in real time, analyzes it to extract the content of the conversation and the emotional state, and calculates and visualizes the mental and physical risk levels. This system consists of three main components: the terminal, the server, and the user. The following describes the details of each component and their operation.
[0108] Acquiring audio data
[0109] When the device detects the start of a meeting, it starts capturing audio data using the built-in microphone or an external microphone. This detection can be done using conferencing software such as Zoom or Microsoft Teams. The audio data is sent to the server in real time in a specified format (WAV, MP3, etc.). This transmission uses the HTTPS protocol.
[0110] Saving audio data
[0111] The server receives the voice data sent from the device and stores it in a database (e.g., MySQL or PostgreSQL). During this storage process, the voice data is properly indexed and prepared for subsequent analysis.
[0112] Speech-to-Text (STT)
[0113] The server calls the Google Speech-to-Text API or IBM Watson's Speech-to-Text service to convert the stored audio data into text data. This process involves splitting the audio data into short clips and converting each into text. This converted text data is temporarily stored in a database.
[0114] Conversation content analysis
[0115] The server inputs the text data into a natural language processing (NLP) module to analyze the content of the conversation. Specifically, it uses libraries such as SpaCy and NLTK to split sentences, tag parts of speech, and recognize named entities. During this process, it extracts keywords related to the progress of work and the content of the consultation. The extracted keywords are stored in a database and used for subsequent processing.
[0116] sentiment analysis
[0117] The server analyzes the emotional state from the text data using an emotion analysis module (e.g., TextBlob or Hugging Face Transformers). Here, emotional states such as "joy," "sadness," "anger," and "fatigue" are identified, and their intensity and frequency are calculated. If an abnormality in the emotion is detected, the information is stored in a database.
[0118] Risk Assessment
[0119] The server calculates the mental and physical risk level based on the results of the sentiment analysis. This risk level is calculated using a big data analysis framework (e.g., Apache Hadoop or Spark). The risk level is calculated through comparison with past data and visualized using a data visualization tool (e.g., Grafana or Tableau).
[0120] Feedback and Notifications
[0121] Based on the analysis results and risk assessment, the server generates regular reports and specific notifications, which are sent via email, Slack, Microsoft Teams, etc. The notifications include details of the risk level and recommended countermeasures.
[0122] Reviewing feedback
[0123] Users can review the feedback and reports sent from the server and take appropriate measures for the mental and physical health of employees. For example, if a high-risk employee is notified, the user can directly follow up with the employee and take measures such as suggesting time off if necessary.
[0124] Specific examples
[0125] As a specific example, consider a situation where a progress meeting is held for a project and a project member says, "I've been feeling tired lately."
[0126] 1. The device collects the audio data of the meeting in real time and sends it to the server.
[0127] 2. The server receives the audio data and converts it to text using the Google Speech-to-Text API.
[0128] 3. The server analyzes the text data "The project is going well, but I've been feeling tired lately" using a natural language processing module (e.g., SpaCy) and extracts the keyword "fatigue."
[0129] 4. The server uses a sentiment analysis module (e.g., TextBlob) to identify "fatigue."
[0130] 5. The server uses a big data analysis framework (e.g., Apache Hadoop) to calculate the "level of mental fatigue" and determine the risk as high.
[0131] 6. The server sends a "High Fatigue Risk" notification to the manager and employee using the specified notification method (e.g., email or Slack).
[0132] 7. The user reviews the notification and takes action such as following up with the employee or suggesting time off.
[0133] Prompt Sentence Examples
[0134] At the next project meeting, we will collect the comments of the members, convert them into text data, and analyze them. Please propose a method to visualize the emotional state and risk level and send a notification to the manager.
[0135] The flow of the identification process in the first embodiment will be described with reference to FIG.
[0136] Step 1:
[0137] The device detects the start of a meeting and begins capturing audio data. As input, it recognizes the meeting session using conferencing software such as Zoom or Microsoft Teams. As output, it converts the audio data captured from the microphone into WAV or MP3 format.
[0138] Specific behavior:
[0139] Activate the device's microphone when the meeting start event is triggered.
[0140] Audio data is collected in real time while being temporarily stored in a buffer.
[0141] Step 2:
[0142] The device sends the collected audio data to the server using the HTTPS protocol. The input is the audio data acquired in step 1. The output is an audio data file that is uploaded to the server.
[0143] Specific behavior:
[0144] The audio data is divided into packets of a fixed size.
[0145] The split data packets are sent to the server via HTTPS.
[0146] Step 3:
[0147] The server stores the received voice data in a database.,The input is the voice data sent from the terminal.,The output is the voice data stored in the database.
[0148] Specific behavior:
[0149] Connect to the database and select the table for storing the audio data.
[0150] The received audio data is properly indexed and written to a database.
[0151] Step 4:
[0152] The server converts the stored voice data into text data using a speech-to-text engine. The input is the voice data stored in the database. The output is the converted text data.
[0153] Specific behavior:
[0154] Split the audio data into short clips.
[0155] It calls the Google Speech-to-Text API or IBM Watson's Speech-to-Text service to convert audio clips individually into text.
[0156] Step 5:
[0157] The server inputs the text data into a natural language processing (NLP) module to analyze the content of the conversation. The input is the text data obtained in step 4. The output is extracted keywords and contextual information.
[0158] Specific behavior:
[0159] Use libraries such as SpaCy and NLTK to perform sentence segmentation, part-of-speech tagging, and named entity recognition.
[0160] Keywords such as "fatigue" and "project" are extracted and stored in a database.
[0161] Step 6:
[0162] The server inputs the text data into the sentiment analysis module to analyze the emotion and tone. The input is the text data obtained in step 5. The output is an analysis of the emotional state.
[0163] Specific behavior:
[0164] Classify emotional states using TextBlob and Hugging Face Transformers.
[0165] Identify emotions such as "joy," "sadness," "anger," and "fatigue" and calculate their intensity.
[0166] Step 7:
[0167] The server calculates the mental and physical risk levels based on the results of the sentiment analysis. The input is the sentiment analysis result obtained in step 6. The output is the calculated risk level.
[0168] Specific behavior:
[0169] The risk level is calculated using a big data analysis framework (e.g., Apache Hadoop or Spark).
[0170] The current risk level is evaluated by comparing it with past data and the results are stored in a database.
[0171] Step 8:
[0172] The server visualizes the calculated risk level using a data visualization tool. The input is the risk level calculated in step 7. The output is a graph or dashboard of the visualized risk level.
[0173] Specific behavior:
[0174] Use Grafana or Tableau to display the risk level in graphs and charts.
[0175] Configure the dashboard so users can understand the real-time risk situation.
[0176] Step 9:
[0177] The server sends notifications via email, Slack, etc. when the risk level exceeds a predetermined threshold. The input is the visualized risk level data. The output is a risk notification message that is generated and sent.
[0178] Specific behavior:
[0179] The risk level data is checked periodically and if a threshold is exceeded, a notification generation process is initiated.
[0180] Send messages in the appropriate notification format (email or chat message) for each user.
[0181] Step 10:
[0182] The user reviews the notifications and reports sent by the server and takes appropriate action. As input, there are notifications and reports sent by the server. As output, actions are taken, such as employee follow-ups and leave suggestions.
[0183] Specific behavior:
[0184] If you receive a notification, review its contents to understand the details of the risk level.
[0185] If necessary, communicate directly with employees and take measures such as suggesting time off if there is a high risk of fatigue.
[0186] Through the above steps, this system can optimize employee health management and help improve work efficiency.
[0187] (Application example 1)
[0188] Next, a description will be given of Application Example 1. In the following description, the data processing device 12 will be referred to as a "server" and the smart device 14 will be referred to as a "terminal."
[0189] In security-related meetings, it is necessary to grasp important information and risk factors in real time without missing them and respond quickly. However, there is a lack of effective means to analyze the content of conversations and emotional states during meetings and instantly assess risks, so it is necessary to improve the responsiveness and accuracy of security management.
[0190] The specific processing by the specific processing unit 290 of the data processing device 12 in the application example 1 is realized by the following means.
[0191] In this invention, the server includes means for acquiring voice data, means for converting the voice data into text data, means for analyzing the text data and extracting the content of the conversation, means for identifying an emotional state from the text data, means for calculating and visualizing a risk level based on the identified emotional state, means for sending a notification when the risk level exceeds a predetermined threshold, means for acquiring voice data of a security conference in real time, identifying risk-related keywords, and evaluating the risk level, and means for notifying a security administrator based on the evaluation results. This makes it possible to quickly and accurately evaluate risks during a conference and take appropriate measures in a timely manner.
[0192] "Audio Data" means digitized audio information collected through a sound capture device such as a microphone.
[0193] "Text data" is data that has been analyzed and converted into textual information and is expressed in natural language.
[0194] "Emotional state" is a classification of a speaker's emotion or mental state identified from text data, and includes, for example, joy, sadness, anger, fatigue, etc.
[0195] "Risk level" is a numerical value or rating score that indicates the likelihood of potential danger or problem, calculated based on the identified emotional state.
[0196] "Security Administrator" means a person or team with a specialized role responsible for security-related monitoring, analysis, and response.
[0197] "Real time" refers to processing that responds quickly in time, with acquisition and analysis occurring almost simultaneously.
[0198] A "notification" is a message or warning sent from the system to a user (in this case, a security administrator) that contains important information or risk assessment results.
[0199] "Keywords" are important words or phrases extracted from the meeting content that serve as indicators of specific risks or issues.
[0200] "Visualization" refers to the visual representation of data and evaluation results, including displaying them in graph or dashboard format.
[0201] "Evaluation results" are specific results and numerical values derived from the analysis and evaluation process, and are information about the level of risk and emotional state that is communicated to managers.
[0202] In order to implement the present invention, the following hardware and software configuration is required.
[0203] Hardware and Software
[0204] 1. Smartphone: Used to capture voice data in real time.
[0205] 2. Microphone: Used to capture accurate voice data.
[0206] 3. Server: The device responsible for the main data processing and storage.
[0207] 4. Speech-to-Text engine: Software that converts voice data, such as Google Cloud Speech-to-Text, into text data.
[0208] 5. Natural Language Processing (NLP) module: Software such as spaCy that analyzes text data and extracts keywords.
[0209] 6. Sentiment Analysis Module: Software that identifies emotional states from text such as VADER.
[0210] 7. Database: A system that stores data, such as PostgreSQL.
[0211] 8. Visualization tools: Tools for visualizing risk levels, such as D3.js.
[0212] 9. Notification system: A system that sends notifications, such as Firebase Cloud Messaging.
[0213] Specific operation of the system
[0214] When the device detects the start of a meeting, it starts capturing audio data through the microphone, and transmits the captured audio data to the server in a specified format.
[0215] The server receives the voice data sent from the device and stores it in a designated database, which is then used in the subsequent analysis process.
[0216] The server calls the speech-to-text engine to convert the stored voice data into text data, which is then temporarily stored and passed to the analysis engine.
[0217] The server inputs the text data into a natural language processing module to analyze the content of the conversation. During this process, risk-related keywords are extracted. The extracted keywords and content are stored in a database for subsequent processing and report generation.
[0218] The server inputs the text data into a sentiment analysis module to analyze the emotions and tone of the conversation. During this process, emotional states such as "joy," "sadness," "anger," and "fatigue" are identified. A risk level is calculated by comparing with past data, and if an abnormality is detected, the information is stored in a database.
[0219] The server calculates the risk level based on the results of sentiment analysis. Using a big database, it compares the current risk level with past data and calculates the current risk level. The calculated risk level is visualized, and the impact and trends are displayed in graphs and dashboard format.
[0220] The server generates reports based on the results of sentiment analysis and risk assessment. If a risk signal is detected, a notification is sent to the security administrator. This notification is sent via email or chat app. Periodic reports are also generated and distributed to the administrator.
[0221] The user (security administrator) checks the feedback and reports sent from the server and takes measures. For example, if a particular risk is determined to be high, detailed investigation or emergency countermeasures can be implemented.
[0222] Examples and prompts
[0223] Examples:
[0224] If a member of a security meeting says, "There has been an increase in external attacks recently," this voice data is captured in real time, converted into text, and risk-related keywords such as "attack" and "external" are extracted. Sentiment analysis then identifies negative sentiment and evaluates it as a high risk. The results are visualized and a notification is sent to the security administrator.
[0225] Prompt statement:
[0226] 1. Speech to text:
[0227] Please tell me how to convert Japanese audio data into text using Google Cloud Speech-to-Text.
[0228] 2. Conversation content analysis:
[0229] Can you please give me the Python code to extract nouns and verbs from text using spaCy's Japanese model?
[0230] 3. Sentiment analysis:
[0231] How do I use VADER to evaluate the negative sentiment score of text?
[0232] 4. Risk Assessment and Visualization:
[0233] Can you please provide some Python code to visualize the risk scores in a pie chart using Matplotlib?
[0234] 5. Notification sending:
[0235] Can you please provide me some Python code to send a notification to a specific topic using Firebase Cloud Messaging?
[0236] By implementing the system in this way, security-related risks in meetings can be assessed quickly and accurately, and appropriate countermeasures can be taken.
[0237] The flow of the specific processing in the application example 1 will be described with reference to FIG.
[0238] Step 1:
[0239] When the device detects the start of a meeting, it starts capturing audio data through the microphone. The captured audio data is sent to the server in real time. The input is an analog audio signal captured from the microphone, and the output is digitized audio data. The device converts the signal into a digital format and sends it to the server.
[0240] Step 2:
[0241] The server receives the voice data sent from the device and stores it in a database. This process prepares the voice data for use in subsequent analysis processes. The input is digitized voice data, and the output is voice data stored in the database. Specifically, the server registers the data as an entry in the database and assigns a timestamp.
[0242] Step 3:
[0243] The server calls the Speech-to-Text engine to convert the stored voice data into text data. The converted text data is temporarily stored and passed to the analysis engine. The input is digitized voice data and the output is text data. The server sends the voice data to the Speech-to-Text engine and temporarily stores the obtained text data.
[0244] Step 4:
[0245] The server inputs the text data into a natural language processing (NLP) module and analyzes the content of the conversation. During this process, risk-related keywords are extracted. The extracted keywords and content are stored in a database and used for subsequent processing and report creation. The input is text data, and the output is a list of extracted keywords. The server analyzes the text data, extracts important keywords, and stores them.
[0246] Step 5:
[0247] The server inputs the text data into the sentiment analysis module to analyze the emotions and tones in the conversation. During this process, emotional states such as "joy," "sadness," "anger," and "fatigue" are identified. The input is text data, and the output is the classification result of the emotional state. The server analyzes the text data, identifies the emotional state, and stores it.
[0248] Step 6:
[0249] The server calculates the risk level based on the results of the emotion analysis. Using a big database, it compares it with past data to calculate the current risk level. The input is the classification result of the emotional state, and the output is a numerical value of the risk level. The server performs a comparison operation, calculates the risk level, and creates data for visualization.
[0250] Step 7:
[0251] The server generates a report based on the results of sentiment analysis and risk assessment. If a risk signal is detected, it sends a notification to the security administrator. This notification is sent via email or chat app. The input is the risk level number and sentiment classification result, and the output is a specific report and notification message. The server generates a report based on this data and sends it to the administrator using the notification system.
[0252] Step 8:
[0253] The user (security administrator) checks the feedback and reports sent from the server and takes appropriate measures. For example, if a particular risk is determined to be high, detailed investigations and emergency countermeasures can be implemented. The input is the reports and notifications from the server, and the output is the administrator's countermeasure actions. The user takes specific response steps based on the contents of the report.
[0254] In this way, the system can quickly and accurately assess security-related risks in meetings and take appropriate action in a timely manner.
[0255] Furthermore, an emotion engine that estimates the user's emotion may be combined. That is, the identification processing unit 290 may estimate the user's emotion using the emotion identification model 59 and perform identification processing using the user's emotion.
[0256] The present invention relates to a system that starts by acquiring voice data, converts the voice data into text data, analyzes the text data to extract the content of the conversation and the emotional state, and calculates and visualizes the risk level. Furthermore, the present invention uses an emotion engine to recognize the user's emotions, and based on this, it is possible to more precisely evaluate the risk level. Below, each component of this system and its operation are specifically described.
[0257] System Embodiments
[0258] Acquiring audio data
[0259] The user enters the meeting details into the system and instructs the system to start the meeting. When the device detects the start of the meeting, it starts capturing audio data in real time through the microphone. The captured audio data is sent to the server in a specified format.
[0260] Saving audio data
[0261] The server receives the voice data sent from the device and stores it in a database, which is used in the subsequent analysis process.
[0262] Speech-to-Text (STT)
[0263] The server inputs the stored voice data into a Speech-to-Text (STT) engine and converts the voice into text data. For example, a speech saying, "The project is going well, but I've been feeling tired lately" is converted into text. This text data is temporarily stored.
[0264] Conversation content analysis
[0265] The server inputs the text data into a natural language processing (NLP) module and analyzes the content of the conversation. During this process, keywords related to the progress of work and the content of the consultation (e.g., "progress," "problem," "fatigue") are extracted. The extracted keywords and content are stored in a database and used for subsequent processing and report creation.
[0266] sentiment analysis
[0267] The server inputs the text data into an emotion engine, which analyzes the emotions and tone of the conversation. During this process, emotional states such as "happiness," "sadness," "anger," and "fatigue" are identified. The emotional state is compared with past data, and the difference from the usual state is calculated. The emotion engine also uses both voice and text data to identify the emotional state in detail. Furthermore, the emotion engine analyzes the user's facial expression data to identify the emotional state. This multi-layered emotion analysis enables more accurate risk calculations.
[0268] Risk Assessment
[0269] The server calculates the mental and physical risk level based on the results of the emotion analysis. It uses a big database to compare it with past data and calculate the current risk level. The results are visualized, and the impact and trends are displayed in graphs and dashboards.
[0270] Feedback and Notifications
[0271] The server generates reports based on the results of sentiment analysis and risk assessment. If a risk signal is detected, a notification is sent to the employee and their manager. This notification can be sent via email or a chat app, for example. Periodic reports are also generated and distributed to managers.
[0272] Reviewing feedback
[0273] Users can check the feedback and reports sent from the server, which can then be used to take measures regarding employees' work status and health status, such as suggesting that fatigued employees take time off.
[0274] Specific examples
[0275] scenario:
[0276] During a progress meeting for a project, one of the members said, "I've been feeling tired lately." Furthermore, it was confirmed that the member had a tired expression throughout the meeting.
[0277] How it works:
[0278] 1. The device collects the audio data of the meeting in real time and sends it to the server.
[0279] 2. The server receives the voice data and converts it into text data.
[0280] 3. The server analyzes the text data "The project is going well, but I've been feeling tired lately" and extracts the keyword "fatigue."
[0281] 4. The server uses the emotion engine to identify the "feeling of fatigue" and then analyzes the user's facial expression data to confirm that they are expressing fatigue.
[0282] 5. The server calculates the "mental fatigue level" and determines that the risk is high.
[0283] 6. The server sends a "high fatigue risk" notification to the manager and employee.
[0284] 7. The user reviews the notification and takes action, such as following up with the employee or suggesting time off.
[0285] In this way, the system of the present invention can perform multi-layered emotion analysis and effectively realize employee health management and productivity improvement throughout the organization.
[0286] The processing flow will be explained below.
[0287] Step 1:
[0288] A user logs in to the system and enters the date and time of the meeting, the participants, the purpose, etc. For example, the following information is registered in the system: "October 5th, 10:00, progress meeting, participants: Mr. A, Mr. B."
[0289] Step 2:
[0290] When the device detects the start of a meeting, it starts collecting audio data in real time through the microphone, and the collected audio data is sent to the server in a specified format.
[0291] Step 3:
[0292] The server receives the voice data sent from the device and stores it in a database, which is used in subsequent analysis processes.
[0293] Step 4:
[0294] The server inputs the stored voice data into a Speech-to-Text (STT) engine and converts the voice into text data. For example, a speech saying, "The project is going well, but I've been feeling tired lately" is converted into text. This converted text data is temporarily stored.
[0295] Step 5:
[0296] The server inputs the text data into a natural language processing (NLP) module and analyzes the content of the conversation. During this process, keywords related to the progress of work and the content of the consultation (e.g., "progress," "problems," "fatigue," etc.) are extracted. The extracted keywords and content are stored in a database.
[0297] Step 6:
[0298] The server inputs the text data into the emotion engine, which analyzes the emotions and tone of the conversation. During this process, emotional states such as "happiness," "sadness," "anger," and "fatigue" are identified. The emotion engine uses both the voice and text data to identify the emotional state in detail. The emotion engine also analyzes the user's facial expression data to complement the emotional state.
[0299] Step 7:
[0300] The server calculates the mental and physical risk level based on the results of the emotion analysis. It uses a big database to compare it with past data and calculate the current risk level. The results are visualized, and the impact and trends are displayed in graphs and dashboard format.
[0301] Step 8:
[0302] The server generates reports based on the results of sentiment analysis and risk assessment. If a risk signal is detected, a notification is sent to the employee and their manager. This notification can be sent via email or a chat app, for example. Periodic reports are also generated and distributed to managers.
[0303] Step 9:
[0304] Users can check the feedback and reports sent from the server, which can then be used to take measures regarding employees' work status and health status, such as suggesting that fatigued employees take time off.
[0305] Specific examples
[0306] scenario:
[0307] During a progress meeting for a project, a project member says, "I've been feeling tired lately." Furthermore, that member has a tired expression throughout the meeting.
[0308] How it works:
[0309] 1. The device collects the audio data of the meeting in real time and sends it to the server.
[0310] 2. The server receives the voice data and converts it into text data.
[0311] 3. The server analyzes the text data "The project is going well, but I've been feeling tired lately" and extracts the keyword "fatigue."
[0312] 4. The server uses the emotion engine to identify the "feeling of fatigue" and then analyzes the user's facial expression data to confirm the tiredness.
[0313] 5. The server calculates the "mental fatigue level" and determines that the risk is high.
[0314] 6. The server sends a "high fatigue risk" notification to the manager and employee.
[0315] 7. The user reviews the notification and takes action, such as following up with the employee or suggesting time off.
[0316] In this way, the system of the present invention can effectively manage employee health and improve productivity throughout the organization by performing multi-layered emotion analysis.
[0317] Example 2
[0318] Next, a description will be given of Example 2. In the following description, the data processing device 12 will be referred to as a "server" and the smart device 14 will be referred to as a "terminal."
[0319] There is a problem in that it is difficult to grasp the work status and health status of employees in real time and respond quickly and accurately. In particular, it is difficult with conventional systems to precisely evaluate emotional states and risk levels and provide appropriate feedback and notifications. This poses a challenge in terms of improving work efficiency and effectively managing employee health.
[0320] The specific processing by the specific processing unit 290 of the data processing device 12 in the second embodiment is realized by the following means.
[0321] In this invention, the server includes means for acquiring voice data, means for converting voice data into text data, means for identifying an emotional state from the text data and voice data, means for calculating and visualizing a risk level by comparing it with past data, and means for sending a notification when the risk level exceeds a predetermined threshold. This makes it possible to precisely evaluate an employee's emotions and risk level and provide feedback and notifications at the appropriate time.
[0322] "Voice data" refers to digital data of voice or sound collected by a voice input device such as a microphone.
[0323] "Text data" is digital data that has been converted from voice data into character information.
[0324] "Emotional state" refers to the psychological and emotional state analyzed from the user's statements, voice data, facial expression data, etc.
[0325] "Risk level" is the degree of mental and physical risk calculated in comparison with emotional state and past data.
[0326] "Visualization" refers to visually displaying data such as risk level and emotional state in the form of graphs, dashboards, etc.
[0327] "Notification" refers to the sending of alerts or information to users or their superiors by the system based on specific conditions.
[0328] A "meeting" refers to a conference or meeting that takes place at a specific date and time, the details of which are entered into the system.
[0329] An "NLP (natural language processing) module" is a program module that analyzes text data and extracts conversation content and keywords.
[0330] An "emotion engine" is a software engine that analyzes voice data, text data, and facial expression data to identify emotional states.
[0331] A "big database" is a large-scale database for storing and analyzing large amounts of past data.
[0332] The system starts by acquiring voice data, converting it into text data, analyzing it to extract the content of the conversation and the emotional state, and then calculating and visualizing the risk level. It also uses an emotion engine to recognize the user's emotions and precisely evaluates the risk level based on that.
[0333] System Embodiments
[0334] Acquiring audio data
[0335] The user enters meeting details (date and time, participant list, purpose, etc.) through the system's user interface and presses the Start Meeting button to issue instructions. When the device detects the start of the meeting, it begins capturing audio data in real time via the microphone. The captured audio data is sent to the server in an appropriate format (e.g., WAV, MP3).
[0336] Saving audio data
[0337] The server receives the voice data sent from the terminal and stores it in a database, which is used in the subsequent analysis process.
[0338] Speech-to-text (STT)
[0339] The server inputs the stored voice data into a Speech-to-Text (STT) engine and converts the voice into text data. For example, a voice saying, "The project is going well, but I've been feeling tired lately" is converted into text data. This text data is temporarily stored.
[0340] Conversation content analysis
[0341] The server inputs the text data into a natural language processing (NLP) module and analyzes the content of the conversation. During this process, keywords related to the progress of work and the content of the consultation (e.g., "progress," "problem," "fatigue") are extracted. The extracted keywords and content are stored in a database and used for subsequent processing and report creation.
[0342] sentiment analysis
[0343] The server inputs text and voice data into an emotion engine to analyze the emotions and tone of the conversation. During this process, emotional states such as "happiness," "sadness," "anger," and "fatigue" are identified. The emotional state is compared with past data, and the difference from the usual state is calculated. Furthermore, in addition to voice and text data, facial expression data is used to identify the emotional state in more detail. This multi-layered emotion analysis enables more accurate risk calculations.
[0344] Risk Assessment
[0345] The server calculates the mental and physical risk level based on the results of the emotion analysis. It uses a big database to compare it with past data and calculate the current risk level. The results are visualized, and the impact and trends are displayed in graphs and dashboards.
[0346] Generate feedback and notifications
[0347] The server generates reports based on the results of sentiment analysis and risk assessment. If a risk signal is detected, a notification is sent to the employee and their manager. This notification can be sent via email or a chat app, for example. Periodic reports are also generated and distributed to managers.
[0348] Review feedback and implement measures
[0349] Users can check the feedback and reports sent from the server, which can then be used to take measures regarding employees' work status and health status, such as suggesting that fatigued employees take time off.
[0350] Specific examples
[0351] scenario:
[0352] During a progress meeting for a project, one of the members said, "I've been feeling tired lately." Furthermore, it was confirmed that the member had a tired expression throughout the meeting.
[0353] How it works:
[0354] 1. The device collects the audio data of the meeting in real time and sends it to the server.
[0355] 2. The server receives the voice data and converts it into text data.
[0356] 3. The server analyzes the text data "The project is going well, but I've been feeling tired lately" and extracts the keyword "fatigue."
[0357] 4. The server uses the emotion engine to identify the "feeling of fatigue" and then analyzes the user's facial expression data to confirm that they are expressing fatigue.
[0358] 5. The server calculates the "mental fatigue level" and determines that the risk is high.
[0359] 6. The server sends a "high fatigue risk" notification to the manager and employee.
[0360] 7. The user reviews the notification and takes action, such as following up with the employee or suggesting time off.
[0361] Prompt Sentence Examples
[0362] "Please explain how a system can collect and analyze everyone's comments during a status meeting, and then analyze their emotional state. Based on specific keywords and emotional tone, it can generate risk notifications."
[0363] In this way, the system of the present invention can perform multi-layered emotion analysis, effectively managing employee health and improving productivity throughout the organization.
[0364] The flow of the identification process in the second embodiment will be described with reference to FIG.
[0365] Step 1:
[0366] The user inputs detailed information about the meeting through the system's user interface and presses the start button to give instructions.
[0367] Input: Meeting details (date and time, participant list, purpose, etc.)
[0368] Output: Trigger to start a meeting
[0369] Specific operation: The user clicks the "Start Meeting" button using the system's web application, and the entered information is sent to the terminal.
[0370] Step 2:
[0371] When the terminal detects the start of a meeting, it starts capturing audio data in real time via the microphone.
[0372] Input: Meeting start trigger
[0373] Output: Real-time audio data
[0374] Specific behavior: The device's microphone will be turned on, a notification will appear saying "Audio collection during the meeting will begin," and audio data will begin to be collected.
[0375] Step 3:
[0376] The device sends the acquired audio data to the server in a specified format (e.g., WAV, MP3).
[0377] Input: Real-time audio data
[0378] Output: Audio data sent to the server in a given format
[0379] Specific operation: The device compresses the collected audio data and sends it to the server in the specified format.
[0380] Step 4:
[0381] The server receives the voice data sent from the terminal and stores it in a predetermined database.
[0382] Input: Audio data sent in a given format
[0383] Output: Audio data stored in a database
[0384] Specific operation: The server logs "Audio data received" and saves the file in the specified database.
[0385] Step 5:
[0386] The server inputs the stored voice data into a Speech-to-Text (STT) engine and converts the voice into text data.
[0387] Input: Audio data stored in a database
[0388] Output: Converted text data
[0389] Specific operation: The STT engine will start, and after a few seconds, the message "Speech to text conversion complete" will be displayed.
[0390] Step 6:
[0391] The server inputs the text data into a natural language processing (NLP) module to analyze the content of the conversation.
[0392] Input: Converted text data
[0393] Output: Extracted keywords and conversation content
[0394] Specific operation: The NLP module will start analysis and the message "Keyword extraction completed" will be displayed.
[0395] Step 7:
[0396] The server inputs the text and voice data into an emotion engine to analyze the emotion and tone of the conversation.
[0397] Input: Converted text data, audio data
[0398] Output: Identified emotional state
[0399] Specific behavior: The sentiment engine will be activated and the results will be displayed on the dashboard as "Sentiment analysis completed."
[0400] Step 8:
[0401] The server calculates the mental and physical risk levels based on the results of the emotion analysis.
[0402] Input: Identified emotional state
[0403] Output: Calculated risk level
[0404] Specific behavior: The risk assessment module will perform the analysis, the message "Risk assessment completed" will be displayed, and the results will be displayed in graphical form on the dashboard.
[0405] Step 9:
[0406] The server sends a notification if the risk exceeds a predetermined threshold.
[0407] Input: Calculated risk level
[0408] Output: Notification sent
[0409] Specific behavior: If the "high risk of fatigue" is determined, a notification will be created and the message "Notification sent" will be displayed.
[0410] Step 10:
[0411] Users check the feedback and reports sent from the server and take action.
[0412] Input: Notifications sent, reports
[0413] Output: Measures taken
[0414] Specific actions: The user receives a notification, checks the details on the dashboard, and clicks the "Suggest time off to employee" button to take action.
[0415] (Application example 2)
[0416] Next, a description will be given of Application Example 2. In the following description, the data processing device 12 will be referred to as a "server" and the smart device 14 will be referred to as a "terminal."
[0417] Modern security services require rapid and accurate detection of potential risks posed by employees and visitors. However, previous technologies often only analyzed voice data, resulting in incomplete understanding of emotional states and making it difficult to accurately assess risk levels. Furthermore, the lack of an effective system for simultaneously performing real-time emotion analysis and risk assessment makes it difficult to respond quickly on-site. Therefore, the present invention aims to provide a system that uses multi-layered emotion analysis to integrate both voice data and facial expression data, enabling more accurate and rapid risk detection and assessment.
[0418] The specific processing by the specific processing unit 290 of the data processing device 12 in the application example 2 is realized by the following means.
[0419] In this invention, the server includes means for acquiring voice data, means for converting voice data into text data, means for analyzing the text data and extracting the content of the conversation, means for identifying an emotional state from the text data, means for calculating and visualizing a risk level based on the identified emotional state, means for sending a notification when the risk level exceeds a predetermined threshold, means for identifying an emotional state using both voice data and facial expression data, and a device (such as smart glasses) for analyzing the emotional state in real time. This enables integrated analysis of voice data and facial expression data, enabling more accurate risk assessment and rapid risk detection in real time.
[0420] "Audio data" is a digital representation of sound captured through an audio input device such as a microphone.
[0421] "Text data" is digital information that expresses voice data as characters.
[0422] The "content of the conversation" is information that indicates the topic and context extracted from the acquired text data.
[0423] "Emotional state" refers to the speaker's psychological state, which is determined by analyzing voice data, text data, and facial expression data.
[0424] "Risk level" is an assessment of potential danger calculated based on emotional state and conversation content.
[0425] "Visualization" is the process of displaying data in a graphical format (e.g., graphs, dashboards, etc.).
[0426] "Means for sending notifications" refers to a function that sends an alert via email, chat app, etc. when the risk level exceeds a specified threshold.
[0427] A "server" is a computer system that processes and stores voice and text data and analyzes them as needed.
[0428] "Apparatus for real-time analysis" refers to a device (e.g., smart glasses) that can acquire voice data and facial expression data in real time and perform analysis and processing immediately.
[0429] "Smart glasses" are eyeglass-type wearable devices that have built-in microphones and cameras and can acquire and analyze voice data and facial expression data.
[0430] overview
[0431] The present invention relates to a system that uses voice data and facial expression data to analyze emotional states and assess risk levels, particularly for use in security and surveillance applications, and is designed to enable security guards wearing smart glasses to more effectively detect risks.
[0432] System Configuration
[0433] The server operates to perform the following main functions:
[0434] 1. Acquiring audio data: Using the microphone built into the smart glasses, surrounding audio is collected in real time.
[0435] 2. Storage of voice data: The collected voice data is sent to a server and stored in a database.
[0436] 3. Speech-to-Text (STT): The stored voice data is converted into text data using a Speech-to-Text (STT) engine.
[0437] 4. Conversation content analysis: Text data is input into a natural language processing (NLP) module to extract the conversation content and important keywords.
[0438] 5. Sentiment analysis: Text data and facial expression data are input into the emotion engine to analyze the emotional state.
[0439] 6. Risk assessment: Based on the results of sentiment analysis, the risk level is calculated and visualized.
[0440] 7. Feedback and Notification: Send notifications to users and supervisors when risks are detected that exceed predetermined thresholds.
[0441] Specific examples
[0442] A security guard at a department store is patrolling while wearing smart glasses. The guard collects a visitor's statement, "I'm frustrated because there have been a lot of strange things happening at the shopping mall recently." Based on this statement, the system extracts keywords such as "strange" and "frustrated," and also detects anger from the visitor's facial expression. Combining this information, the system assesses the person as being at high risk and displays a notification on the guard's HUD indicating that the person is at high risk.
[0443] Hardware and software used
[0444] Hardware: Smart glasses with built-in microphone and camera
[0445] Software: Speech Recognition Library (speech-to-text conversion), NLP Module (natural language processing), Emotion Engine (sentiment analysis), Risk Assessment Module (risk assessment)
[0446] Prompt Sentence Examples
[0447] We want to design a system that analyzes the emotional state of a user from their voice data and facial expression data, and evaluates the risk level in real time. This system has the following functions:
[0448] 1. Acquiring audio data and converting it to text
[0449] 2. Content analysis of text data
[0450] 3. Emotional state analysis (based on voice and facial expression data)
[0451] 4. Calculating and visualizing risk levels
[0452] Generate pseudocode for a program that combines these elements.
[0453] The flow of the specific processing in the application example 2 will be described with reference to FIG.
[0454] Step 1:
[0455] Acquiring audio data
[0456] The server captures voice data in real time through the microphone of the smart glasses. This data is sent to the server as voice input. The input is the voice data uttered by the user, and the output is the voice file sent to the server.
[0457] Step 2:
[0458] Saving audio data
[0459] The server stores the captured audio data in a database, which is used in the subsequent analysis process. The input is the audio file sent to the server, and the output is the audio data stored in the database.
[0460] Step 3:
[0461] Speech-to-text (STT)
[0462] The server inputs the stored voice data into a Speech-to-Text (STT) engine and converts the voice into text data. The input is the voice data stored in the database, and the output is text data. This allows the voice data to be treated as text information.
[0463] Step 4:
[0464] Conversation content analysis
[0465] The server inputs the text data into a natural language processing (NLP) module and analyzes the content of the conversation. During this process, important keywords (e.g., "progress," "problems," and "fatigue") are extracted. The input is text data, and the output is a list of extracted keywords.
[0466] Step 5:
[0467] sentiment analysis
[0468] The server inputs text data and facial expression data captured by the smart glasses' built-in camera into the emotion engine to analyze the emotional state. During this process, emotions such as "happiness," "sadness," "anger," and "fatigue" are identified. The input is text data and facial expression data, and the output is the identified emotional state.
[0469] Step 6:
[0470] Risk Assessment
[0471] The server calculates the risk level based on the emotional state and the results of the conversation content analysis. It uses a big database to compare it with past data and evaluate the current risk level. The input is the emotional state and a list of keywords, and the output is the calculated risk level.
[0472] Step 7:
[0473] Feedback and Notifications
[0474] The server sends a notification to the user and supervisor if the calculated risk level exceeds a predetermined threshold. This notification is sent via email or a chat app. The input is the risk level and the output is the notification message.
[0475] Step 8:
[0476] Display and confirmation
[0477] The user checks the results of the risk level and emotional state on the head-up display (HUD) of the smart glasses. The input is a notification message, and the output is the information displayed on the HUD. This allows the user to understand the situation in real time and respond quickly.
[0478] The specific processing unit 290 transmits the result of the specific processing to the smart device 14. In the smart device 14, the control unit 46A causes the output device 40 to output the result of the specific processing. The microphone 38B acquires audio indicating a user input regarding the result of the specific processing. The control unit 46A transmits audio data indicating the user input acquired by the microphone 38B to the data processing device 12. In the data processing device 12, the specific processing unit 290 acquires the audio data.
[0479] The data generation model 58 is a so-called generative AI (Artificial Intelligence). An example of the data generation model 58 is ChatGPT (Internet Search<URL: https: / / openai.com / blog / chatgpt> ), Gemini (Internet search <url: https: gemini.google.com ?hl="ja">) and other generation AIs. The data generation model 58 is obtained by performing deep learning on a neural network. A prompt including an instruction is input to the data generation model 58, and inference data such as voice data indicating voice, text data indicating text, and image data indicating an image is also input. The data generation model 58 performs inference on the input inference data in accordance with the instruction indicated by the prompt, and outputs the inference result in a data format such as voice data and text data. Here, inference refers to, for example, analysis, classification, prediction, and / or summarization.
[0480] In the above embodiment, an example in which the specific process is performed by the data processing device 12 has been given, but the technology of the present disclosure is not limited to this, and the specific process may be performed by the smart device 14.
[0481] [Second embodiment]
[0482] FIG. 3 shows an example of the configuration of a data processing system 210 according to the second embodiment.
[0483] 3, the data processing system 210 includes the data processing device 12 and smart glasses 214. An example of the data processing device 12 is a server.
[0484] The data processing device 12 includes a computer 22, a database 24, and a communication I / F 26. The computer 22 is an example of a "computer" according to the technology of the present disclosure. The computer 22 includes a processor 28, a RAM 30, and a storage 32. The processor 28, the RAM 30, and the storage 32 are connected to a bus 34. The database 24 and the communication I / F 26 are also connected to the bus 34. The communication I / F 26 is connected to a network 54. Examples of the network 54 include a WAN (Wide Area Network) and / or a LAN (Local Area Network).
[0485] The smart glasses 214 include a computer 36, a microphone 238, a speaker 240, a camera 42, and a communication I / F 44. The computer 36 includes a processor 46, a RAM 48, and a storage 50. The processor 46, the RAM 48, and the storage 50 are connected to a bus 52. The microphone 238, the speaker 240, and the camera 42 are also connected to the bus 52.
[0486] The microphone 238 receives instructions and the like from the user 20 by receiving voice uttered by the user 20. The microphone 238 captures the voice uttered by the user 20, converts the captured voice into audio data, and outputs it to the processor 46. The speaker 240 outputs audio in accordance with instructions from the processor 46.
[0487] Camera 42 is a small digital camera equipped with an optical system including a lens, aperture, and shutter, and an imaging element such as a CMOS (Complementary Metal-Oxide-Semiconductor) image sensor or a CCD (Charge Coupled Device) image sensor, and captures images of the surroundings of user 20 (for example, an imaging range defined by an angle of view equivalent to the field of vision of a typical healthy person).
[0488] The communication I / F 44 is connected to a network 54. The communication I / Fs 44 and 26 are responsible for the exchange of various information between the processor 46 and the processor 28 via the network 54. The exchange of various information between the processor 46 and the processor 28 using the communication I / Fs 44 and 26 is carried out in a secure state.
[0489] Fig. 4 shows an example of the main functions of the data processing device 12 and the smart glasses 214. As shown in Fig. 4, in the data processing device 12, a specific process is performed by the processor 28. A specific process program 56 is stored in the storage 32.
[0490] The specific processing program 56 is an example of a "program" according to the technology of the present disclosure. The processor 28 reads the specific processing program 56 from the storage 32 and executes the read specific processing program 56 on the RAM 30. The specific processing is realized by the processor 28 operating as a specific processing unit 290 in accordance with the specific processing program 56 executed on the RAM 30.
[0491] The storage 32 stores a data generation model 58 and an emotion identification model 59. The data generation model 58 and the emotion identification model 59 are used by the identification processing unit 290.
[0492] In the smart glasses 214, the reception output process is performed by the processor 46. A reception output program 60 is stored in the storage 50. The processor 46 reads the reception output program 60 from the storage 50 and executes the read reception output program 60 on the RAM 48. The reception output process is realized by the processor 46 operating as the control unit 46A in accordance with the reception output program 60 executed on the RAM 48.
[0493] Next, a description will be given of the identification process performed by the identification processing unit 290 of the data processing device 12. In the following description, the data processing device 12 will be referred to as the "server" and the smart glasses 214 will be referred to as the "terminal."
[0494] The present invention relates to a system that starts by acquiring voice data, converts the voice data into text data, analyzes the text data to extract the content of the conversation and the emotional state, and calculates and visualizes the risk level. Below, we will explain in detail each component of this system and its operation.
[0495] System Embodiments
[0496] Acquiring audio data
[0497] When the device detects the start of a meeting, it starts capturing audio data in real time through the microphone, and sends the captured audio data to the server in a specified format.
[0498] Saving audio data
[0499] The server receives the voice data sent from the device and stores it in a designated database, which is used in the subsequent analysis process.
[0500] Speech-to-Text (STT)
[0501] The server calls the speech-to-text engine to convert the stored voice data into text data. For example, a speech saying, "The project is going well, but I've been feeling tired lately" is converted into text. This converted text data is temporarily stored and passed to the analysis engine.
[0502] Conversation content analysis
[0503] The server inputs the text data into a natural language processing (NLP) module and analyzes the content of the conversation. During this process, keywords related to the progress of work and the content of the consultation (e.g., "progress," "problem," "fatigue") are extracted. The extracted keywords and content are stored in a database and used for subsequent processing and report creation.
[0504] sentiment analysis
[0505] The server inputs the text data into a sentiment analysis module to analyze the emotions and tone of the conversation. During this process, emotional states such as "happiness," "sadness," "anger," and "fatigue" are identified. The emotional state is compared with past data, and the difference from the normal state is calculated. If an anomaly is detected, the information is stored in a database.
[0506] Risk Assessment
[0507] The server calculates mental and physical risk levels based on the results of emotion analysis. Using a big database, it compares current risk levels with past data and calculates the current risk level. The calculated risk level is visualized, and the impact and trends are displayed in graphs and dashboard format.
[0508] Feedback and Notifications
[0509] The server generates reports based on the results of sentiment analysis and risk assessment. If a risk signal is detected, a notification is sent to the employee and their manager. This notification can be sent via email or a chat app, for example. Periodic reports are also generated and distributed to managers.
[0510] Reviewing feedback
[0511] Users can check the feedback and reports sent from the server, which can then be used to take measures regarding employees' work status and health status, for example, suggesting that fatigued employees take time off.
[0512] Specific examples
[0513] scenario:
[0514] A progress meeting is held for a project, and a project member says, "I've been feeling tired lately."
[0515] How it works:
[0516] 1. The device collects the audio data of the meeting in real time and sends it to the server.
[0517] 2. The server receives the voice data and converts it into text data.
[0518] 3. The server analyzes the text data "The project is going well, but I've been feeling tired lately" and extracts the keyword "fatigue."
[0519] 4. The server uses the emotion analysis module to identify the "feeling of fatigue."
[0520] 5. The server calculates the "mental fatigue level" and determines that the risk is high.
[0521] 6. The server sends a "high fatigue risk" notification to the manager and employee.
[0522] 7. The user reviews the notification and takes action such as following up with the employee or suggesting time off.
[0523] The above is a specific embodiment of the present invention. This system can effectively manage the health of employees and improve productivity throughout the organization.
[0524] The processing flow will be explained below.
[0525] Step 1:
[0526] A user logs in to the system and enters information such as the date and time of the meeting, participants, purpose, etc. For example, information such as "October 5th, 10:00, progress meeting, participants: Mr. A, Mr. B" is registered in the system.
[0527] Step 2:
[0528] When the device detects the start of a meeting, it starts collecting audio data in real time through the microphone, and sends the collected audio data to the server in a specified format.
[0529] Step 3:
[0530] The server receives the voice data sent from the device and stores it in a database, which is used in the subsequent analysis process.
[0531] Step 4:
[0532] The server inputs the stored voice data into a Speech-to-Text (STT) engine and converts the voice into text data. For example, a speech saying, "The project is going well, but I've been feeling tired lately" is converted into text. This text data is temporarily stored.
[0533] Step 5:
[0534] The server inputs the text data into a natural language processing (NLP) module and analyzes the content of the conversation. During this process, keywords related to the progress of work and the content of the consultation (e.g., "progress," "problem," "fatigue") are extracted. The extracted keywords and content are stored in a database and used for subsequent processing and report creation.
[0535] Step 6:
[0536] The server inputs the text data into a sentiment analysis module to analyze the emotions and tone of the conversation. During this process, emotional states such as "happiness," "sadness," "anger," and "fatigue" are identified. The emotional state is compared with past data, and the difference from the normal state is calculated. If an anomaly is detected, the information is stored in a database.
[0537] Step 7:
[0538] The server calculates the mental and physical risk level based on the results of the emotion analysis. It uses a big database to compare it with past data and calculate the current risk level. The results are visualized, and the impact and trends are displayed in graphs and dashboard format.
[0539] Step 8:
[0540] The server generates reports based on the results of sentiment analysis and risk assessment. If a risk signal is detected, a notification is sent to the employee and their supervisor so that they can take action. This notification can be sent via email or a chat app, for example. Periodic reports are also generated and distributed to supervisors.
[0541] Step 9:
[0542] The user checks the feedback and reports sent from the server, which allows the user to take measures regarding the work situation and health status of employees. For example, the user can suggest that a fatigued employee take time off.
[0543] Example 1
[0544] Next, a description will be given of Example 1. In the following description, the data processing device 12 will be referred to as a "server" and the smart glasses 214 will be referred to as a "terminal."
[0545] There is a need to monitor employees' mental and physical health conditions in real time to prevent declines in work efficiency and productivity. However, conventional methods have made it difficult to properly analyze the content of comments and emotional states made during meetings, assess risks, and take prompt measures. The present invention aims to solve these problems and provide a system for effectively managing employees' health conditions.
[0546] The specific processing by the specific processing unit 290 of the data processing device 12 in the first embodiment is realized by the following means.
[0547] In this invention, the server includes means for acquiring voice data, means for converting the voice data into text data, means for analyzing the text data and extracting the content of the conversation, means for identifying an emotional state from the text data, means for calculating and visualizing a risk level based on the identified emotional state, means for sending a notification when the risk level exceeds a predetermined threshold, means for a terminal to detect the start of a meeting and send the voice data to the server in real time, means for the server to store the voice data in a database and convert the voice data into text data using a speech-to-text engine, means for the server to extract keywords related to work progress and consultation content from the text data using a natural language processing module, means for the server to analyze the emotional state using a sentiment analysis module, means for the server to calculate a risk level using a big data analysis framework and visualize the risk level using a data visualization tool, means for the server to generate a report and send a notification to the employee and their supervisor using a predetermined notification means, and means for a user to check the notification from the server and take measures. This enables real-time monitoring of employees' mental and physical health conditions and rapid visualization and notification of risks.
[0548] "Audio data" refers to data in which voice or acoustic information is recorded in digital format.
[0549] "Text data" refers to data obtained by converting voice data into character information.
[0550] A "natural language processing (NLP) module" is software and algorithms used to analyze text data and understand its content and meaning.
[0551] "Sentiment Analysis Module" means software and algorithms for identifying emotional states from text data.
[0552] "Risk level" is a calculated degree of risk based on the identified emotional state and other data.
[0553] "Visualization" is the presentation of data in a visual format such as a graph, chart, or dashboard.
[0554] "Notification" means sending a warning or information to the user when a specific condition is met.
[0555] A "terminal" is a device that acquires voice data and transmits it to a server, and includes devices such as personal computers and smartphones.
[0556] A "server" is a computer system for storing and processing audio data.
[0557] A "database" is a system that systematically stores and manages information.
[0558] A "Speech-to-Text engine" is software for converting voice data into text data.
[0559] "Work progress status" is information that indicates the progress of projects and tasks.
[0560] A "big data analytics framework" is software and a platform for analyzing large amounts of data and extracting meaningful information.
[0561] A "data visualization tool" is a tool for visually displaying data, and includes tools such as Grafana and Tableau.
[0562] A "report" is a document that summarizes the results of analysis and evaluation.
[0563] "User" means a person who uses the System to review feedback and notifications and take appropriate action.
[0564] "Detecting the start of a meeting" means that the system automatically recognizes the start of a meeting.
[0565] System Overview
[0566] This invention is a system that captures audio data during meetings in real time, analyzes it to extract the content of the conversation and the emotional state, and calculates and visualizes the mental and physical risk levels. This system consists of three main components: the terminal, the server, and the user. The following describes the details of each component and their operation.
[0567] Acquiring audio data
[0568] When the device detects the start of a meeting, it starts capturing audio data using the built-in microphone or an external microphone. This detection can be done using conferencing software such as Zoom or Microsoft Teams. The audio data is sent to the server in real time in a specified format (WAV, MP3, etc.). This transmission uses the HTTPS protocol.
[0569] Saving audio data
[0570] The server receives the voice data sent from the device and stores it in a database (e.g., MySQL or PostgreSQL). During this storage process, the voice data is properly indexed and prepared for subsequent analysis.
[0571] Speech-to-Text (STT)
[0572] The server calls the Google Speech-to-Text API or IBM Watson's Speech-to-Text service to convert the stored audio data into text data. This process involves splitting the audio data into short clips and converting each into text. This converted text data is temporarily stored in a database.
[0573] Conversation content analysis
[0574] The server inputs the text data into a natural language processing (NLP) module to analyze the content of the conversation. Specifically, it uses libraries such as SpaCy and NLTK to split sentences, tag parts of speech, and recognize named entities. During this process, it extracts keywords related to the progress of work and the content of the consultation. The extracted keywords are stored in a database and used for subsequent processing.
[0575] sentiment analysis
[0576] The server analyzes the emotional state from the text data using an emotion analysis module (e.g., TextBlob or Hugging Face Transformers). Here, emotional states such as "joy," "sadness," "anger," and "fatigue" are identified, and their intensity and frequency are calculated. If an abnormality in the emotion is detected, the information is stored in a database.
[0577] Risk Assessment
[0578] The server calculates the mental and physical risk level based on the results of the sentiment analysis. This risk level is calculated using a big data analysis framework (e.g., Apache Hadoop or Spark). The risk level is calculated through comparison with past data and visualized using a data visualization tool (e.g., Grafana or Tableau).
[0579] Feedback and Notifications
[0580] Based on the analysis results and risk assessment, the server generates regular reports and specific notifications, which are sent via email, Slack, Microsoft Teams, etc. The notifications include details of the risk level and recommended countermeasures.
[0581] Reviewing feedback
[0582] Users can review the feedback and reports sent from the server and take appropriate measures for the mental and physical health of employees. For example, if a high-risk employee is notified, the user can directly follow up with the employee and take measures such as suggesting time off if necessary.
[0583] Specific examples
[0584] As a specific example, consider a situation where a progress meeting is held for a project and a project member says, "I've been feeling tired lately."
[0585] 1. The device collects the audio data of the meeting in real time and sends it to the server.
[0586] 2. The server receives the audio data and converts it to text using the Google Speech-to-Text API.
[0587] 3. The server analyzes the text data "The project is going well, but I've been feeling tired lately" using a natural language processing module (e.g., SpaCy) and extracts the keyword "fatigue."
[0588] 4. The server uses a sentiment analysis module (e.g., TextBlob) to identify "fatigue."
[0589] 5. The server uses a big data analysis framework (e.g., Apache Hadoop) to calculate the "level of mental fatigue" and determine the risk as high.
[0590] 6. The server sends a "High Fatigue Risk" notification to the manager and employee using the specified notification method (e.g., email or Slack).
[0591] 7. The user reviews the notification and takes action such as following up with the employee or suggesting time off.
[0592] Prompt Sentence Examples
[0593] At the next project meeting, we will collect the comments of the members, convert them into text data, and analyze them. Please propose a method to visualize the emotional state and risk level and send a notification to the manager.
[0594] The flow of the identification process in the first embodiment will be described with reference to FIG.
[0595] Step 1:
[0596] The device detects the start of a meeting and begins capturing audio data. As input, it recognizes the meeting session using conferencing software such as Zoom or Microsoft Teams. As output, it converts the audio data captured from the microphone into WAV or MP3 format.
[0597] Specific behavior:
[0598] Activate the device's microphone when the meeting start event is triggered.
[0599] Audio data is collected in real time while being temporarily stored in a buffer.
[0600] Step 2:
[0601] The device sends the collected audio data to the server using the HTTPS protocol. The input is the audio data acquired in step 1. The output is an audio data file that is uploaded to the server.
[0602] Specific behavior:
[0603] The audio data is divided into packets of a fixed size.
[0604] The split data packets are sent to the server via HTTPS.
[0605] Step 3:
[0606] The server stores the received voice data in a database.,The input is the voice data sent from the terminal.,The output is the voice data stored in the database.
[0607] Specific behavior:
[0608] Connect to the database and select the table for storing the audio data.
[0609] The received audio data is properly indexed and written to a database.
[0610] Step 4:
[0611] The server converts the stored voice data into text data using a speech-to-text engine. The input is the voice data stored in the database. The output is the converted text data.
[0612] Specific behavior:
[0613] Split the audio data into short clips.
[0614] It calls the Google Speech-to-Text API or IBM Watson's Speech-to-Text service to convert audio clips individually into text.
[0615] Step 5:
[0616] The server inputs the text data into a natural language processing (NLP) module to analyze the content of the conversation. The input is the text data obtained in step 4. The output is extracted keywords and contextual information.
[0617] Specific behavior:
[0618] Use libraries such as SpaCy and NLTK to perform sentence segmentation, part-of-speech tagging, and named entity recognition.
[0619] Keywords such as "fatigue" and "project" are extracted and stored in a database.
[0620] Step 6:
[0621] The server inputs the text data into the sentiment analysis module to analyze the emotion and tone. The input is the text data obtained in step 5. The output is an analysis of the emotional state.
[0622] Specific behavior:
[0623] Classify emotional states using TextBlob and Hugging Face Transformers.
[0624] Identify emotions such as "joy," "sadness," "anger," and "fatigue" and calculate their intensity.
[0625] Step 7:
[0626] The server calculates the mental and physical risk levels based on the results of the sentiment analysis. The input is the sentiment analysis result obtained in step 6. The output is the calculated risk level.
[0627] Specific behavior:
[0628] The risk level is calculated using a big data analysis framework (e.g., Apache Hadoop or Spark).
[0629] The current risk level is evaluated by comparing it with past data and the results are stored in a database.
[0630] Step 8:
[0631] The server visualizes the calculated risk level using a data visualization tool. The input is the risk level calculated in step 7. The output is a graph or dashboard of the visualized risk level.
[0632] Specific behavior:
[0633] Use Grafana or Tableau to display the risk level in graphs and charts.
[0634] Configure the dashboard so users can understand the real-time risk situation.
[0635] Step 9:
[0636] The server sends notifications via email, Slack, etc. when the risk level exceeds a predetermined threshold. The input is the visualized risk level data. The output is a risk notification message that is generated and sent.
[0637] Specific behavior:
[0638] The risk level data is checked periodically and if a threshold is exceeded, a notification generation process is initiated.
[0639] Send messages in the appropriate notification format (email or chat message) for each user.
[0640] Step 10:
[0641] The user reviews the notifications and reports sent by the server and takes appropriate action. As input, there are notifications and reports sent by the server. As output, actions are taken, such as employee follow-ups and leave suggestions.
[0642] Specific behavior:
[0643] If you receive a notification, review its contents to understand the details of the risk level.
[0644] If necessary, communicate directly with employees and take measures such as suggesting time off if there is a high risk of fatigue.
[0645] Through the above steps, this system can optimize employee health management and help improve work efficiency.
[0646] (Application example 1)
[0647] Next, a description will be given of Application Example 1. In the following description, the data processing device 12 will be referred to as a "server" and the smart glasses 214 will be referred to as a "terminal."
[0648] In security-related meetings, it is necessary to grasp important information and risk factors in real time without missing them and respond quickly. However, there is a lack of effective means to analyze the content of conversations and emotional states during meetings and instantly assess risks, so it is necessary to improve the responsiveness and accuracy of security management.
[0649] The specific processing by the specific processing unit 290 of the data processing device 12 in the application example 1 is realized by the following means.
[0650] In this invention, the server includes means for acquiring voice data, means for converting the voice data into text data, means for analyzing the text data and extracting the content of the conversation, means for identifying an emotional state from the text data, means for calculating and visualizing a risk level based on the identified emotional state, means for sending a notification when the risk level exceeds a predetermined threshold, means for acquiring voice data of a security conference in real time, identifying risk-related keywords, and evaluating the risk level, and means for notifying a security administrator based on the evaluation results. This makes it possible to quickly and accurately evaluate risks during a conference and take appropriate measures in a timely manner.
[0651] "Audio Data" means digitized audio information collected through a sound capture device such as a microphone.
[0652] "Text data" is data that has been analyzed and converted into textual information and is expressed in natural language.
[0653] "Emotional state" is a classification of a speaker's emotion or mental state identified from text data, and includes, for example, joy, sadness, anger, fatigue, etc.
[0654] "Risk level" is a numerical value or rating score that indicates the likelihood of potential danger or problem, calculated based on the identified emotional state.
[0655] "Security Administrator" means a person or team with a specialized role responsible for security-related monitoring, analysis, and response.
[0656] "Real time" refers to processing that responds quickly in time, with acquisition and analysis occurring almost simultaneously.
[0657] A "notification" is a message or warning sent from the system to a user (in this case, a security administrator) that contains important information or risk assessment results.
[0658] "Keywords" are important words or phrases extracted from the meeting content that serve as indicators of specific risks or issues.
[0659] "Visualization" refers to the visual representation of data and evaluation results, including displaying them in graph or dashboard format.
[0660] "Evaluation results" are specific results and numerical values derived from the analysis and evaluation process, and are information about the level of risk and emotional state that is communicated to managers.
[0661] In order to implement the present invention, the following hardware and software configuration is required.
[0662] Hardware and Software
[0663] 1. Smartphone: Used to capture voice data in real time.
[0664] 2. Microphone: Used to capture accurate voice data.
[0665] 3. Server: The device responsible for the main data processing and storage.
[0666] 4. Speech-to-Text engine: Software that converts voice data, such as Google Cloud Speech-to-Text, into text data.
[0667] 5. Natural Language Processing (NLP) module: Software such as spaCy that analyzes text data and extracts keywords.
[0668] 6. Sentiment Analysis Module: Software that identifies emotional states from text such as VADER.
[0669] 7. Database: A system that stores data, such as PostgreSQL.
[0670] 8. Visualization tools: Tools for visualizing risk levels, such as D3.js.
[0671] 9. Notification system: A system that sends notifications, such as Firebase Cloud Messaging.
[0672] Specific operation of the system
[0673] When the device detects the start of a meeting, it starts capturing audio data through the microphone, and transmits the captured audio data to the server in a specified format.
[0674] The server receives the voice data sent from the device and stores it in a designated database, which is then used in the subsequent analysis process.
[0675] The server calls the speech-to-text engine to convert the stored voice data into text data, which is then temporarily stored and passed to the analysis engine.
[0676] The server inputs the text data into a natural language processing module to analyze the content of the conversation. During this process, risk-related keywords are extracted. The extracted keywords and content are stored in a database for subsequent processing and report generation.
[0677] The server inputs the text data into a sentiment analysis module to analyze the emotions and tone of the conversation. During this process, emotional states such as "joy," "sadness," "anger," and "fatigue" are identified. A risk level is calculated by comparing with past data, and if an abnormality is detected, the information is stored in a database.
[0678] The server calculates the risk level based on the results of sentiment analysis. Using a big database, it compares the current risk level with past data and calculates the current risk level. The calculated risk level is visualized, and the impact and trends are displayed in graphs and dashboard format.
[0679] The server generates reports based on the results of sentiment analysis and risk assessment. If a risk signal is detected, a notification is sent to the security administrator. This notification is sent via email or chat app. Periodic reports are also generated and distributed to the administrator.
[0680] The user (security administrator) checks the feedback and reports sent from the server and takes measures. For example, if a particular risk is determined to be high, detailed investigation or emergency countermeasures can be implemented.
[0681] Examples and prompts
[0682] Examples:
[0683] If a member of a security meeting says, "There has been an increase in external attacks recently," this voice data is captured in real time, converted into text, and risk-related keywords such as "attack" and "external" are extracted. Sentiment analysis then identifies negative sentiment and evaluates it as a high risk. The results are visualized and a notification is sent to the security administrator.
[0684] Prompt statement:
[0685] 1. Speech to text:
[0686] Please tell me how to convert Japanese audio data into text using Google Cloud Speech-to-Text.
[0687] 2. Conversation content analysis:
[0688] Can you please give me the Python code to extract nouns and verbs from text using spaCy's Japanese model?
[0689] 3. Sentiment analysis:
[0690] How do I use VADER to evaluate the negative sentiment score of text?
[0691] 4. Risk Assessment and Visualization:
[0692] Can you please provide some Python code to visualize the risk scores in a pie chart using Matplotlib?
[0693] 5. Notification sending:
[0694] Can you please provide me some Python code to send a notification to a specific topic using Firebase Cloud Messaging?
[0695] By implementing the system in this way, security-related risks in meetings can be assessed quickly and accurately, and appropriate countermeasures can be taken.
[0696] The flow of the specific processing in the application example 1 will be described with reference to FIG.
[0697] Step 1:
[0698] When the device detects the start of a meeting, it starts capturing audio data through the microphone. The captured audio data is sent to the server in real time. The input is an analog audio signal captured from the microphone, and the output is digitized audio data. The device converts the signal into a digital format and sends it to the server.
[0699] Step 2:
[0700] The server receives the voice data sent from the device and stores it in a database. This process prepares the voice data for use in subsequent analysis processes. The input is digitized voice data, and the output is voice data stored in the database. Specifically, the server registers the data as an entry in the database and assigns a timestamp.
[0701] Step 3:
[0702] The server calls the Speech-to-Text engine to convert the stored voice data into text data. The converted text data is temporarily stored and passed to the analysis engine. The input is digitized voice data and the output is text data. The server sends the voice data to the Speech-to-Text engine and temporarily stores the obtained text data.
[0703] Step 4:
[0704] The server inputs the text data into a natural language processing (NLP) module and analyzes the content of the conversation. During this process, risk-related keywords are extracted. The extracted keywords and content are stored in a database and used for subsequent processing and report creation. The input is text data, and the output is a list of extracted keywords. The server analyzes the text data, extracts important keywords, and stores them.
[0705] Step 5:
[0706] The server inputs the text data into the sentiment analysis module to analyze the emotions and tones in the conversation. During this process, emotional states such as "joy," "sadness," "anger," and "fatigue" are identified. The input is text data, and the output is the classification result of the emotional state. The server analyzes the text data, identifies the emotional state, and stores it.
[0707] Step 6:
[0708] The server calculates the risk level based on the results of the emotion analysis. Using a big database, it compares it with past data to calculate the current risk level. The input is the classification result of the emotional state, and the output is a numerical value of the risk level. The server performs a comparison operation, calculates the risk level, and creates data for visualization.
[0709] Step 7:
[0710] The server generates a report based on the results of sentiment analysis and risk assessment. If a risk signal is detected, it sends a notification to the security administrator. This notification is sent via email or chat app. The input is the risk level number and sentiment classification result, and the output is a specific report and notification message. The server generates a report based on this data and sends it to the administrator using the notification system.
[0711] Step 8:
[0712] The user (security administrator) checks the feedback and reports sent from the server and takes appropriate measures. For example, if a particular risk is determined to be high, detailed investigations and emergency countermeasures can be implemented. The input is the reports and notifications from the server, and the output is the administrator's countermeasure actions. The user takes specific response steps based on the contents of the report.
[0713] In this way, the system can quickly and accurately assess security-related risks in meetings and take appropriate action in a timely manner.
[0714] Furthermore, an emotion engine that estimates the user's emotion may be further combined. That is, the identification processing unit 290 may estimate the user's emotion using the emotion identification model 59, and perform identification processing using the user's emotion.
[0715] The present invention relates to a system that starts by acquiring voice data, converts the voice data into text data, analyzes the text data to extract the content of the conversation and the emotional state, and calculates and visualizes the risk level. Furthermore, the present invention uses an emotion engine to recognize the user's emotions, and based on this, it is possible to more precisely evaluate the risk level. Below, each component of this system and its operation are specifically described.
[0716] System Embodiments
[0717] Acquiring audio data
[0718] The user enters the meeting details into the system and instructs the system to start the meeting. When the device detects the start of the meeting, it starts capturing audio data in real time through the microphone. The captured audio data is sent to the server in a specified format.
[0719] Saving audio data
[0720] The server receives the voice data sent from the device and stores it in a database, which is used in the subsequent analysis process.
[0721] Speech-to-Text (STT)
[0722] The server inputs the stored voice data into a Speech-to-Text (STT) engine and converts the voice into text data. For example, a speech saying, "The project is going well, but I've been feeling tired lately" is converted into text. This text data is temporarily stored.
[0723] Conversation content analysis
[0724] The server inputs the text data into a natural language processing (NLP) module and analyzes the content of the conversation. During this process, keywords related to the progress of work and the content of the consultation (e.g., "progress," "problem," "fatigue") are extracted. The extracted keywords and content are stored in a database and used for subsequent processing and report creation.
[0725] sentiment analysis
[0726] The server inputs the text data into an emotion engine, which analyzes the emotions and tone of the conversation. During this process, emotional states such as "happiness," "sadness," "anger," and "fatigue" are identified. The emotional state is compared with past data, and the difference from the usual state is calculated. The emotion engine also uses both voice and text data to identify the emotional state in detail. Furthermore, the emotion engine analyzes the user's facial expression data to identify the emotional state. This multi-layered emotion analysis enables more accurate risk calculations.
[0727] Risk Assessment
[0728] The server calculates the mental and physical risk level based on the results of the emotion analysis. It uses a big database to compare it with past data and calculate the current risk level. The results are visualized, and the impact and trends are displayed in graphs and dashboards.
[0729] Feedback and Notifications
[0730] The server generates reports based on the results of sentiment analysis and risk assessment. If a risk signal is detected, a notification is sent to the employee and their manager. This notification can be sent via email or a chat app, for example. Periodic reports are also generated and distributed to managers.
[0731] Reviewing feedback
[0732] Users can check the feedback and reports sent from the server, which can then be used to take measures regarding employees' work status and health status, such as suggesting that fatigued employees take time off.
[0733] Specific examples
[0734] scenario:
[0735] During a progress meeting for a project, one of the members said, "I've been feeling tired lately." Furthermore, it was confirmed that the member had a tired expression throughout the meeting.
[0736] How it works:
[0737] 1. The device collects the audio data of the meeting in real time and sends it to the server.
[0738] 2. The server receives the voice data and converts it into text data.
[0739] 3. The server analyzes the text data "The project is going well, but I've been feeling tired lately" and extracts the keyword "fatigue."
[0740] 4. The server uses the emotion engine to identify the "feeling of fatigue" and then analyzes the user's facial expression data to confirm that they are expressing fatigue.
[0741] 5. The server calculates the "mental fatigue level" and determines that the risk is high.
[0742] 6. The server sends a "high fatigue risk" notification to the manager and employee.
[0743] 7. The user reviews the notification and takes action, such as following up with the employee or suggesting time off.
[0744] In this way, the system of the present invention can perform multi-layered emotion analysis and effectively realize employee health management and productivity improvement throughout the organization.
[0745] The processing flow will be explained below.
[0746] Step 1:
[0747] A user logs in to the system and enters the date and time of the meeting, the participants, the purpose, etc. For example, the following information is registered in the system: "October 5th, 10:00, progress meeting, participants: Mr. A, Mr. B."
[0748] Step 2:
[0749] When the device detects the start of a meeting, it starts collecting audio data in real time through the microphone, and the collected audio data is sent to the server in a specified format.
[0750] Step 3:
[0751] The server receives the voice data sent from the device and stores it in a database, which is used in subsequent analysis processes.
[0752] Step 4:
[0753] The server inputs the stored voice data into a Speech-to-Text (STT) engine and converts the voice into text data. For example, a speech saying, "The project is going well, but I've been feeling tired lately" is converted into text. This converted text data is temporarily stored.
[0754] Step 5:
[0755] The server inputs the text data into a natural language processing (NLP) module and analyzes the content of the conversation. During this process, keywords related to the progress of work and the content of the consultation (e.g., "progress," "problems," "fatigue," etc.) are extracted. The extracted keywords and content are stored in a database.
[0756] Step 6:
[0757] The server inputs the text data into the emotion engine, which analyzes the emotions and tone of the conversation. During this process, emotional states such as "happiness," "sadness," "anger," and "fatigue" are identified. The emotion engine uses both the voice and text data to identify the emotional state in detail. The emotion engine also analyzes the user's facial expression data to complement the emotional state.
[0758] Step 7:
[0759] The server calculates the mental and physical risk level based on the results of the emotion analysis. It uses a big database to compare it with past data and calculate the current risk level. The results are visualized, and the impact and trends are displayed in graphs and dashboard format.
[0760] Step 8:
[0761] The server generates reports based on the results of sentiment analysis and risk assessment. If a risk signal is detected, a notification is sent to the employee and their manager. This notification can be sent via email or a chat app, for example. Periodic reports are also generated and distributed to managers.
[0762] Step 9:
[0763] Users can check the feedback and reports sent from the server, which can then be used to take measures regarding employees' work status and health status, such as suggesting that fatigued employees take time off.
[0764] Specific examples
[0765] scenario:
[0766] During a progress meeting for a project, a project member says, "I've been feeling tired lately." Furthermore, that member has a tired expression throughout the meeting.
[0767] How it works:
[0768] 1. The device collects the audio data of the meeting in real time and sends it to the server.
[0769] 2. The server receives the voice data and converts it into text data.
[0770] 3. The server analyzes the text data "The project is going well, but I've been feeling tired lately" and extracts the keyword "fatigue."
[0771] 4. The server uses the emotion engine to identify the "feeling of fatigue" and then analyzes the user's facial expression data to confirm the tiredness.
[0772] 5. The server calculates the "mental fatigue level" and determines that the risk is high.
[0773] 6. The server sends a "high fatigue risk" notification to the manager and employee.
[0774] 7. The user reviews the notification and takes action, such as following up with the employee or suggesting time off.
[0775] In this way, the system of the present invention can effectively manage employee health and improve productivity throughout the organization by performing multi-layered emotion analysis.
[0776] Example 2
[0777] Next, a description will be given of Example 2. In the following description, the data processing device 12 will be referred to as a "server" and the smart glasses 214 will be referred to as a "terminal."
[0778] There is a problem in that it is difficult to grasp the work status and health status of employees in real time and respond quickly and accurately. In particular, it is difficult with conventional systems to precisely evaluate emotional states and risk levels and provide appropriate feedback and notifications. This poses a challenge in terms of improving work efficiency and effectively managing employee health.
[0779] The specific processing by the specific processing unit 290 of the data processing device 12 in the second embodiment is realized by the following means.
[0780] In this invention, the server includes means for acquiring voice data, means for converting voice data into text data, means for identifying an emotional state from the text data and voice data, means for calculating and visualizing a risk level by comparing it with past data, and means for sending a notification when the risk level exceeds a predetermined threshold. This makes it possible to precisely evaluate an employee's emotions and risk level and provide feedback and notifications at the appropriate time.
[0781] "Voice data" refers to digital data of voice or sound collected by a voice input device such as a microphone.
[0782] "Text data" is digital data that has been converted from voice data into character information.
[0783] "Emotional state" refers to the psychological and emotional state analyzed from the user's statements, voice data, facial expression data, etc.
[0784] "Risk level" is the degree of mental and physical risk calculated in comparison with emotional state and past data.
[0785] "Visualization" refers to visually displaying data such as risk level and emotional state in the form of graphs, dashboards, etc.
[0786] "Notification" refers to the sending of alerts or information to users or their superiors by the system based on specific conditions.
[0787] A "meeting" refers to a conference or meeting that takes place at a specific date and time, the details of which are entered into the system.
[0788] An "NLP (natural language processing) module" is a program module that analyzes text data and extracts conversation content and keywords.
[0789] An "emotion engine" is a software engine that analyzes voice data, text data, and facial expression data to identify emotional states.
[0790] A "big database" is a large-scale database for storing and analyzing large amounts of past data.
[0791] The system starts by acquiring voice data, converting it into text data, analyzing it to extract the content of the conversation and the emotional state, and then calculating and visualizing the risk level. It also uses an emotion engine to recognize the user's emotions and precisely evaluates the risk level based on that.
[0792] System Embodiments
[0793] Acquiring audio data
[0794] The user enters meeting details (date and time, participant list, purpose, etc.) through the system's user interface and presses the Start Meeting button to issue instructions. When the device detects the start of the meeting, it begins capturing audio data in real time via the microphone. The captured audio data is sent to the server in an appropriate format (e.g., WAV, MP3).
[0795] Saving audio data
[0796] The server receives the voice data sent from the terminal and stores it in a database, which is used in the subsequent analysis process.
[0797] Speech-to-text (STT)
[0798] The server inputs the stored voice data into a Speech-to-Text (STT) engine and converts the voice into text data. For example, a voice saying, "The project is going well, but I've been feeling tired lately" is converted into text data. This text data is temporarily stored.
[0799] Conversation content analysis
[0800] The server inputs the text data into a natural language processing (NLP) module and analyzes the content of the conversation. During this process, keywords related to the progress of work and the content of the consultation (e.g., "progress," "problem," "fatigue") are extracted. The extracted keywords and content are stored in a database and used for subsequent processing and report creation.
[0801] sentiment analysis
[0802] The server inputs text and voice data into an emotion engine to analyze the emotions and tone of the conversation. During this process, emotional states such as "happiness," "sadness," "anger," and "fatigue" are identified. The emotional state is compared with past data, and the difference from the usual state is calculated. Furthermore, in addition to voice and text data, facial expression data is used to identify the emotional state in more detail. This multi-layered emotion analysis enables more accurate risk calculations.
[0803] Risk Assessment
[0804] The server calculates the mental and physical risk level based on the results of the emotion analysis. It uses a big database to compare it with past data and calculate the current risk level. The results are visualized, and the impact and trends are displayed in graphs and dashboards.
[0805] Generate feedback and notifications
[0806] The server generates reports based on the results of sentiment analysis and risk assessment. If a risk signal is detected, a notification is sent to the employee and their manager. This notification can be sent via email or a chat app, for example. Periodic reports are also generated and distributed to managers.
[0807] Review feedback and implement measures
[0808] Users can check the feedback and reports sent from the server, which can then be used to take measures regarding employees' work status and health status, such as suggesting that fatigued employees take time off.
[0809] Specific examples
[0810] scenario:
[0811] During a progress meeting for a project, one of the members said, "I've been feeling tired lately." Furthermore, it was confirmed that the member had a tired expression throughout the meeting.
[0812] How it works:
[0813] 1. The device collects the audio data of the meeting in real time and sends it to the server.
[0814] 2. The server receives the voice data and converts it into text data.
[0815] 3. The server analyzes the text data "The project is going well, but I've been feeling tired lately" and extracts the keyword "fatigue."
[0816] 4. The server uses the emotion engine to identify the "feeling of fatigue" and then analyzes the user's facial expression data to confirm that they are expressing fatigue.
[0817] 5. The server calculates the "mental fatigue level" and determines that the risk is high.
[0818] 6. The server sends a "high fatigue risk" notification to the manager and employee.
[0819] 7. The user reviews the notification and takes action, such as following up with the employee or suggesting time off.
[0820] Prompt Sentence Examples
[0821] "Please explain how a system can collect and analyze everyone's comments during a status meeting, and then analyze their emotional state. Based on specific keywords and emotional tone, it can generate risk notifications."
[0822] In this way, the system of the present invention can perform multi-layered emotion analysis, effectively managing employee health and improving productivity throughout the organization.
[0823] The flow of the identification process in the second embodiment will be described with reference to FIG.
[0824] Step 1:
[0825] The user inputs detailed information about the meeting through the system's user interface and presses the start button to give instructions.
[0826] Input: Meeting details (date and time, participant list, purpose, etc.)
[0827] Output: Trigger to start a meeting
[0828] Specific operation: The user clicks the "Start Meeting" button using the system's web application, and the entered information is sent to the terminal.
[0829] Step 2:
[0830] When the terminal detects the start of a meeting, it starts capturing audio data in real time via the microphone.
[0831] Input: Meeting start trigger
[0832] Output: Real-time audio data
[0833] Specific behavior: The device's microphone will be turned on, a notification will appear saying "Audio collection during the meeting will begin," and audio data will begin to be collected.
[0834] Step 3:
[0835] The device sends the acquired audio data to the server in a specified format (e.g., WAV, MP3).
[0836] Input: Real-time audio data
[0837] Output: Audio data sent to the server in a given format
[0838] Specific operation: The device compresses the collected audio data and sends it to the server in the specified format.
[0839] Step 4:
[0840] The server receives the voice data sent from the terminal and stores it in a predetermined database.
[0841] Input: Audio data sent in a given format
[0842] Output: Audio data stored in a database
[0843] Specific operation: The server logs "Audio data received" and saves the file in the specified database.
[0844] Step 5:
[0845] The server inputs the stored voice data into a Speech-to-Text (STT) engine and converts the voice into text data.
[0846] Input: Audio data stored in a database
[0847] Output: Converted text data
[0848] Specific operation: The STT engine will start, and after a few seconds, the message "Speech to text conversion complete" will be displayed.
[0849] Step 6:
[0850] The server inputs the text data into a natural language processing (NLP) module to analyze the content of the conversation.
[0851] Input: Converted text data
[0852] Output: Extracted keywords and conversation content
[0853] Specific operation: The NLP module will start analysis and the message "Keyword extraction completed" will be displayed.
[0854] Step 7:
[0855] The server inputs the text and voice data into an emotion engine to analyze the emotion and tone of the conversation.
[0856] Input: Converted text data, audio data
[0857] Output: Identified emotional state
[0858] Specific behavior: The sentiment engine will be activated and the results will be displayed on the dashboard as "Sentiment analysis completed."
[0859] Step 8:
[0860] The server calculates the mental and physical risk levels based on the results of the emotion analysis.
[0861] Input: Identified emotional state
[0862] Output: Calculated risk level
[0863] Specific behavior: The risk assessment module will perform the analysis, the message "Risk assessment completed" will be displayed, and the results will be displayed in graphical form on the dashboard.
[0864] Step 9:
[0865] The server sends a notification if the risk exceeds a predetermined threshold.
[0866] Input: Calculated risk level
[0867] Output: Notification sent
[0868] Specific behavior: If the "high risk of fatigue" is determined, a notification will be created and the message "Notification sent" will be displayed.
[0869] Step 10:
[0870] Users check the feedback and reports sent from the server and take action.
[0871] Input: Notifications sent, reports
[0872] Output: Measures taken
[0873] Specific actions: The user receives a notification, checks the details on the dashboard, and clicks the "Suggest time off to employee" button to take action.
[0874] (Application example 2)
[0875] Next, a description will be given of Application Example 2. In the following description, the data processing device 12 will be referred to as a "server" and the smart glasses 214 will be referred to as a "terminal."
[0876] Modern security services require rapid and accurate detection of potential risks posed by employees and visitors. However, previous technologies often only analyzed voice data, resulting in incomplete understanding of emotional states and making it difficult to accurately assess risk levels. Furthermore, the lack of an effective system for simultaneously performing real-time emotion analysis and risk assessment makes it difficult to respond quickly on-site. Therefore, the present invention aims to provide a system that uses multi-layered emotion analysis to integrate both voice data and facial expression data, enabling more accurate and rapid risk detection and assessment.
[0877] The specific processing by the specific processing unit 290 of the data processing device 12 in the application example 2 is realized by the following means.
[0878] In this invention, the server includes means for acquiring voice data, means for converting voice data into text data, means for analyzing the text data and extracting the content of the conversation, means for identifying an emotional state from the text data, means for calculating and visualizing a risk level based on the identified emotional state, means for sending a notification when the risk level exceeds a predetermined threshold, means for identifying an emotional state using both voice data and facial expression data, and a device (such as smart glasses) for analyzing the emotional state in real time. This enables integrated analysis of voice data and facial expression data, enabling more accurate risk assessment and rapid risk detection in real time.
[0879] "Audio data" is a digital representation of sound captured through an audio input device such as a microphone.
[0880] "Text data" is digital information that expresses voice data as characters.
[0881] The "content of the conversation" is information that indicates the topic and context extracted from the acquired text data.
[0882] "Emotional state" refers to the speaker's psychological state, which is determined by analyzing voice data, text data, and facial expression data.
[0883] "Risk level" is an assessment of potential danger calculated based on emotional state and conversation content.
[0884] "Visualization" is the process of displaying data in a graphical format (e.g., graphs, dashboards, etc.).
[0885] "Means for sending notifications" refers to a function that sends an alert via email, chat app, etc. when the risk level exceeds a specified threshold.
[0886] A "server" is a computer system that processes and stores voice and text data and analyzes them as needed.
[0887] "Apparatus for real-time analysis" refers to a device (e.g., smart glasses) that can acquire voice data and facial expression data in real time and perform analysis and processing immediately.
[0888] "Smart glasses" are eyeglass-type wearable devices that have built-in microphones and cameras and can acquire and analyze voice data and facial expression data.
[0889] overview
[0890] The present invention relates to a system that uses voice data and facial expression data to analyze emotional states and assess risk levels, particularly for use in security and surveillance applications, and is designed to enable security guards wearing smart glasses to more effectively detect risks.
[0891] System Configuration
[0892] The server operates to perform the following main functions:
[0893] 1. Acquiring audio data: Using the microphone built into the smart glasses, surrounding audio is collected in real time.
[0894] 2. Storage of voice data: The collected voice data is sent to a server and stored in a database.
[0895] 3. Speech-to-Text (STT): The stored voice data is converted into text data using a Speech-to-Text (STT) engine.
[0896] 4. Conversation content analysis: Text data is input into a natural language processing (NLP) module to extract the conversation content and important keywords.
[0897] 5. Sentiment analysis: Text data and facial expression data are input into the emotion engine to analyze the emotional state.
[0898] 6. Risk assessment: Based on the results of sentiment analysis, the risk level is calculated and visualized.
[0899] 7. Feedback and Notification: Send notifications to users and supervisors when risks are detected that exceed predetermined thresholds.
[0900] Specific examples
[0901] A security guard at a department store is patrolling while wearing smart glasses. The guard collects a visitor's statement, "I'm frustrated because there have been a lot of strange things happening at the shopping mall recently." Based on this statement, the system extracts keywords such as "strange" and "frustrated," and also detects anger from the visitor's facial expression. Combining this information, the system assesses the person as being at high risk and displays a notification on the guard's HUD indicating that the person is at high risk.
[0902] Hardware and software used
[0903] Hardware: Smart glasses with built-in microphone and camera
[0904] Software: Speech Recognition Library (speech-to-text conversion), NLP Module (natural language processing), Emotion Engine (sentiment analysis), Risk Assessment Module (risk assessment)
[0905] Prompt Sentence Examples
[0906] We want to design a system that analyzes the emotional state of a user from their voice data and facial expression data, and evaluates the risk level in real time. This system has the following functions:
[0907] 1. Acquiring audio data and converting it to text
[0908] 2. Content analysis of text data
[0909] 3. Emotional state analysis (based on voice and facial expression data)
[0910] 4. Calculating and visualizing risk levels
[0911] Generate pseudocode for a program that combines these elements.
[0912] The flow of the specific processing in the application example 2 will be described with reference to FIG.
[0913] Step 1:
[0914] Acquiring audio data
[0915] The server captures voice data in real time through the microphone of the smart glasses. This data is sent to the server as voice input. The input is the voice data uttered by the user, and the output is the voice file sent to the server.
[0916] Step 2:
[0917] Saving audio data
[0918] The server stores the captured audio data in a database, which is used in the subsequent analysis process. The input is the audio file sent to the server, and the output is the audio data stored in the database.
[0919] Step 3:
[0920] Speech-to-text (STT)
[0921] The server inputs the stored voice data into a Speech-to-Text (STT) engine and converts the voice into text data. The input is the voice data stored in the database, and the output is text data. This allows the voice data to be treated as text information.
[0922] Step 4:
[0923] Conversation content analysis
[0924] The server inputs the text data into a natural language processing (NLP) module and analyzes the content of the conversation. During this process, important keywords (e.g., "progress," "problems," and "fatigue") are extracted. The input is text data, and the output is a list of extracted keywords.
[0925] Step 5:
[0926] sentiment analysis
[0927] The server inputs text data and facial expression data captured by the smart glasses' built-in camera into the emotion engine to analyze the emotional state. During this process, emotions such as "happiness," "sadness," "anger," and "fatigue" are identified. The input is text data and facial expression data, and the output is the identified emotional state.
[0928] Step 6:
[0929] Risk Assessment
[0930] The server calculates the risk level based on the emotional state and the results of the conversation content analysis. It uses a big database to compare it with past data and evaluate the current risk level. The input is the emotional state and a list of keywords, and the output is the calculated risk level.
[0931] Step 7:
[0932] Feedback and Notifications
[0933] The server sends a notification to the user and supervisor if the calculated risk level exceeds a predetermined threshold. This notification is sent via email or a chat app. The input is the risk level and the output is the notification message.
[0934] Step 8:
[0935] Display and confirmation
[0936] The user checks the results of the risk level and emotional state on the head-up display (HUD) of the smart glasses. The input is a notification message, and the output is the information displayed on the HUD. This allows the user to understand the situation in real time and respond quickly.
[0937] The specific processing unit 290 transmits the result of the specific processing to the smart glasses 214. In the smart glasses 214, the control unit 46A causes the speaker 240 to output the result of the specific processing. The microphone 238 acquires audio indicating a user input regarding the result of the specific processing. The control unit 46A transmits audio data indicating the user input acquired by the microphone 238 to the data processing device 12. In the data processing device 12, the specific processing unit 290 acquires the audio data.
[0938] The data generation model 58 is a so-called generative AI (Artificial Intelligence). An example of the data generation model 58 is ChatGPT (Internet Search<URL: https: / / openai.com / blog / chatgpt> ), Gemini (Internet search <url: https: gemini.google.com ?hl="ja">) and other generation AIs. The data generation model 58 is obtained by performing deep learning on a neural network. A prompt including an instruction is input to the data generation model 58, and inference data such as voice data indicating voice, text data indicating text, and image data indicating an image is also input. The data generation model 58 performs inference on the input inference data in accordance with the instruction indicated by the prompt, and outputs the inference result in a data format such as voice data and text data. Here, inference refers to, for example, analysis, classification, prediction, and / or summarization.
[0939] In the above embodiment, an example in which the specific processing is performed by the data processing device 12 has been given, but the technology of the present disclosure is not limited to this, and the specific processing may be performed by the smart glasses 214.
[0940] [Third embodiment]
[0941] FIG. 5 shows an example of the configuration of a data processing system 310 according to the third embodiment.
[0942] 5, the data processing system 310 includes the data processing device 12 and a headset type terminal 314. An example of the data processing device 12 is a server.
[0943] The data processing device 12 includes a computer 22, a database 24, and a communication I / F 26. The computer 22 is an example of a "computer" according to the technology of the present disclosure. The computer 22 includes a processor 28, a RAM 30, and a storage 32. The processor 28, the RAM 30, and the storage 32 are connected to a bus 34. The database 24 and the communication I / F 26 are also connected to the bus 34. The communication I / F 26 is connected to a network 54. Examples of the network 54 include a WAN (Wide Area Network) and / or a LAN (Local Area Network).
[0944] The headset type terminal 314 includes a computer 36, a microphone 238, a speaker 240, a camera 42, a communication I / F 44, and a display 343. The computer 36 includes a processor 46, a RAM 48, and a storage 50. The processor 46, the RAM 48, and the storage 50 are connected to a bus 52. The microphone 238, the speaker 240, the camera 42, and the display 343 are also connected to the bus 52.
[0945] The microphone 238 receives instructions and the like from the user 20 by receiving voice uttered by the user 20. The microphone 238 captures the voice uttered by the user 20, converts the captured voice into audio data, and outputs it to the processor 46. The speaker 240 outputs audio in accordance with instructions from the processor 46.
[0946] Camera 42 is a small digital camera equipped with an optical system including a lens, aperture, and shutter, and an imaging element such as a CMOS (Complementary Metal-Oxide-Semiconductor) image sensor or a CCD (Charge Coupled Device) image sensor, and captures images of the surroundings of user 20 (for example, an imaging range defined by an angle of view equivalent to the field of vision of a typical healthy person).
[0947] The communication I / F 44 is connected to a network 54. The communication I / Fs 44 and 26 are responsible for the exchange of various information between the processor 46 and the processor 28 via the network 54. The exchange of various information between the processor 46 and the processor 28 using the communication I / Fs 44 and 26 is carried out in a secure state.
[0948] Fig. 6 shows an example of the main functions of the data processing device 12 and the headset type terminal 314. As shown in Fig. 6, in the data processing device 12, a specific process is performed by the processor 28. A specific process program 56 is stored in the storage 32.
[0949] The specific processing program 56 is an example of a "program" according to the technology of the present disclosure. The processor 28 reads the specific processing program 56 from the storage 32 and executes the read specific processing program 56 on the RAM 30. The specific processing is realized by the processor 28 operating as a specific processing unit 290 in accordance with the specific processing program 56 executed on the RAM 30.
[0950] The storage 32 stores a data generation model 58 and an emotion identification model 59. The data generation model 58 and the emotion identification model 59 are used by the identification processing unit 290.
[0951] In the headset type terminal 314, a reception output process is performed by the processor 46. A reception output program 60 is stored in the storage 50. The processor 46 reads the reception output program 60 from the storage 50 and executes the read reception output program 60 on the RAM 48. The reception output process is realized by the processor 46 operating as the control unit 46A in accordance with the reception output program 60 executed on the RAM 48.
[0952] Next, a description will be given of the identification process performed by the identification processing unit 290 of the data processing device 12. In the following description, the data processing device 12 will be referred to as the "server" and the headset type terminal 314 will be referred to as the "terminal."
[0953] The present invention relates to a system that starts by acquiring voice data, converts the voice data into text data, analyzes the text data to extract the content of the conversation and the emotional state, and calculates and visualizes the risk level. Below, we will explain in detail each component of this system and its operation.
[0954] System Embodiments
[0955] Acquiring audio data
[0956] When the device detects the start of a meeting, it starts capturing audio data in real time through the microphone, and sends the captured audio data to the server in a specified format.
[0957] Saving audio data
[0958] The server receives the voice data sent from the device and stores it in a designated database, which is used in the subsequent analysis process.
[0959] Speech-to-Text (STT)
[0960] The server calls the speech-to-text engine to convert the stored voice data into text data. For example, a speech saying, "The project is going well, but I've been feeling tired lately" is converted into text. This converted text data is temporarily stored and passed to the analysis engine.
[0961] Conversation content analysis
[0962] The server inputs the text data into a natural language processing (NLP) module and analyzes the content of the conversation. During this process, keywords related to the progress of work and the content of the consultation (e.g., "progress," "problem," "fatigue") are extracted. The extracted keywords and content are stored in a database and used for subsequent processing and report creation.
[0963] sentiment analysis
[0964] The server inputs the text data into a sentiment analysis module to analyze the emotions and tone of the conversation. During this process, emotional states such as "happiness," "sadness," "anger," and "fatigue" are identified. The emotional state is compared with past data, and the difference from the normal state is calculated. If an anomaly is detected, the information is stored in a database.
[0965] Risk Assessment
[0966] The server calculates mental and physical risk levels based on the results of emotion analysis. Using a big database, it compares current risk levels with past data and calculates the current risk level. The calculated risk level is visualized, and the impact and trends are displayed in graphs and dashboard format.
[0967] Feedback and Notifications
[0968] The server generates reports based on the results of sentiment analysis and risk assessment. If a risk signal is detected, a notification is sent to the employee and their manager. This notification can be sent via email or a chat app, for example. Periodic reports are also generated and distributed to managers.
[0969] Reviewing feedback
[0970] Users can check the feedback and reports sent from the server, which can then be used to take measures regarding employees' work status and health status, for example, suggesting that fatigued employees take time off.
[0971] Specific examples
[0972] scenario:
[0973] A progress meeting is held for a project, and a project member says, "I've been feeling tired lately."
[0974] How it works:
[0975] 1. The device collects the audio data of the meeting in real time and sends it to the server.
[0976] 2. The server receives the voice data and converts it into text data.
[0977] 3. The server analyzes the text data "The project is going well, but I've been feeling tired lately" and extracts the keyword "fatigue."
[0978] 4. The server uses the emotion analysis module to identify the "feeling of fatigue."
[0979] 5. The server calculates the "mental fatigue level" and determines that the risk is high.
[0980] 6. The server sends a "high fatigue risk" notification to the manager and employee.
[0981] 7. The user reviews the notification and takes action such as following up with the employee or suggesting time off.
[0982] The above is a specific embodiment of the present invention. This system can effectively manage the health of employees and improve productivity throughout the organization.
[0983] The processing flow will be explained below.
[0984] Step 1:
[0985] A user logs in to the system and enters information such as the date and time of the meeting, participants, purpose, etc. For example, information such as "October 5th, 10:00, progress meeting, participants: Mr. A, Mr. B" is registered in the system.
[0986] Step 2:
[0987] When the device detects the start of a meeting, it starts collecting audio data in real time through the microphone, and sends the collected audio data to the server in a specified format.
[0988] Step 3:
[0989] The server receives the voice data sent from the device and stores it in a database, which is used in the subsequent analysis process.
[0990] Step 4:
[0991] The server inputs the stored voice data into a Speech-to-Text (STT) engine and converts the voice into text data. For example, a speech saying, "The project is going well, but I've been feeling tired lately" is converted into text. This text data is temporarily stored.
[0992] Step 5:
[0993] The server inputs the text data into a natural language processing (NLP) module and analyzes the content of the conversation. During this process, keywords related to the progress of work and the content of the consultation (e.g., "progress," "problem," "fatigue") are extracted. The extracted keywords and content are stored in a database and used for subsequent processing and report creation.
[0994] Step 6:
[0995] The server inputs the text data into a sentiment analysis module to analyze the emotions and tone of the conversation. During this process, emotional states such as "happiness," "sadness," "anger," and "fatigue" are identified. The emotional state is compared with past data, and the difference from the normal state is calculated. If an anomaly is detected, the information is stored in a database.
[0996] Step 7:
[0997] The server calculates the mental and physical risk level based on the results of the emotion analysis. It uses a big database to compare it with past data and calculate the current risk level. The results are visualized, and the impact and trends are displayed in graphs and dashboard format.
[0998] Step 8:
[0999] The server generates reports based on the results of sentiment analysis and risk assessment. If a risk signal is detected, a notification is sent to the employee and their supervisor so that they can take action. This notification can be sent via email or a chat app, for example. Periodic reports are also generated and distributed to supervisors.
[1000] Step 9:
[1001] The user checks the feedback and reports sent from the server, which allows the user to take measures regarding the work situation and health status of employees. For example, the user can suggest that a fatigued employee take time off.
[1002] Example 1
[1003] Next, a description will be given of Example 1. In the following description, the data processing device 12 will be referred to as a "server" and the headset type terminal 314 will be referred to as a "terminal."
[1004] There is a need to monitor employees' mental and physical health conditions in real time to prevent declines in work efficiency and productivity. However, conventional methods have made it difficult to properly analyze the content of comments and emotional states made during meetings, assess risks, and take prompt measures. The present invention aims to solve these problems and provide a system for effectively managing employees' health conditions.
[1005] The specific processing by the specific processing unit 290 of the data processing device 12 in the first embodiment is realized by the following means.
[1006] In this invention, the server includes means for acquiring voice data, means for converting the voice data into text data, means for analyzing the text data and extracting the content of the conversation, means for identifying an emotional state from the text data, means for calculating and visualizing a risk level based on the identified emotional state, means for sending a notification when the risk level exceeds a predetermined threshold, means for a terminal to detect the start of a meeting and send the voice data to the server in real time, means for the server to store the voice data in a database and convert the voice data into text data using a speech-to-text engine, means for the server to extract keywords related to work progress and consultation content from the text data using a natural language processing module, means for the server to analyze the emotional state using a sentiment analysis module, means for the server to calculate a risk level using a big data analysis framework and visualize the risk level using a data visualization tool, means for the server to generate a report and send a notification to the employee and their supervisor using a predetermined notification means, and means for a user to check the notification from the server and take measures. This enables real-time monitoring of employees' mental and physical health conditions and rapid visualization and notification of risks.
[1007] "Audio data" refers to data in which voice or acoustic information is recorded in digital format.
[1008] "Text data" refers to data obtained by converting voice data into character information.
[1009] A "natural language processing (NLP) module" is software and algorithms used to analyze text data and understand its content and meaning.
[1010] "Sentiment Analysis Module" means software and algorithms for identifying emotional states from text data.
[1011] "Risk level" is a calculated degree of risk based on the identified emotional state and other data.
[1012] "Visualization" is the presentation of data in a visual format such as a graph, chart, or dashboard.
[1013] "Notification" means sending a warning or information to the user when a specific condition is met.
[1014] A "terminal" is a device that acquires voice data and transmits it to a server, and includes devices such as personal computers and smartphones.
[1015] A "server" is a computer system for storing and processing audio data.
[1016] A "database" is a system that systematically stores and manages information.
[1017] A "Speech-to-Text engine" is software for converting voice data into text data.
[1018] "Work progress status" is information that indicates the progress of projects and tasks.
[1019] A "big data analytics framework" is software and a platform for analyzing large amounts of data and extracting meaningful information.
[1020] A "data visualization tool" is a tool for visually displaying data, and includes tools such as Grafana and Tableau.
[1021] A "report" is a document that summarizes the results of analysis and evaluation.
[1022] "User" means a person who uses the System to review feedback and notifications and take appropriate action.
[1023] "Detecting the start of a meeting" means that the system automatically recognizes the start of a meeting.
[1024] System Overview
[1025] This invention is a system that captures audio data during meetings in real time, analyzes it to extract the content of the conversation and the emotional state, and calculates and visualizes the mental and physical risk levels. This system consists of three main components: the terminal, the server, and the user. The following describes the details of each component and their operation.
[1026] Acquiring audio data
[1027] When the device detects the start of a meeting, it starts capturing audio data using the built-in microphone or an external microphone. This detection can be done using conferencing software such as Zoom or Microsoft Teams. The audio data is sent to the server in real time in a specified format (WAV, MP3, etc.). This transmission uses the HTTPS protocol.
[1028] Saving audio data
[1029] The server receives the voice data sent from the device and stores it in a database (e.g., MySQL or PostgreSQL). During this storage process, the voice data is properly indexed and prepared for subsequent analysis.
[1030] Speech-to-Text (STT)
[1031] The server calls the Google Speech-to-Text API or IBM Watson's Speech-to-Text service to convert the stored audio data into text data. This process involves splitting the audio data into short clips and converting each into text. This converted text data is temporarily stored in a database.
[1032] Conversation content analysis
[1033] The server inputs the text data into a natural language processing (NLP) module to analyze the content of the conversation. Specifically, it uses libraries such as SpaCy and NLTK to split sentences, tag parts of speech, and recognize named entities. During this process, it extracts keywords related to the progress of work and the content of the consultation. The extracted keywords are stored in a database and used for subsequent processing.
[1034] sentiment analysis
[1035] The server analyzes the emotional state from the text data using an emotion analysis module (e.g., TextBlob or Hugging Face Transformers). Here, emotional states such as "joy," "sadness," "anger," and "fatigue" are identified, and their intensity and frequency are calculated. If an abnormality in the emotion is detected, the information is stored in a database.
[1036] Risk Assessment
[1037] The server calculates the mental and physical risk level based on the results of the sentiment analysis. This risk level is calculated using a big data analysis framework (e.g., Apache Hadoop or Spark). The risk level is calculated through comparison with past data and visualized using a data visualization tool (e.g., Grafana or Tableau).
[1038] Feedback and Notifications
[1039] Based on the analysis results and risk assessment, the server generates regular reports and specific notifications, which are sent via email, Slack, Microsoft Teams, etc. The notifications include details of the risk level and recommended countermeasures.
[1040] Reviewing feedback
[1041] Users can review the feedback and reports sent from the server and take appropriate measures for the mental and physical health of employees. For example, if a high-risk employee is notified, the user can directly follow up with the employee and take measures such as suggesting time off if necessary.
[1042] Specific examples
[1043] As a specific example, consider a situation where a progress meeting is held for a project and a project member says, "I've been feeling tired lately."
[1044] 1. The device collects the audio data of the meeting in real time and sends it to the server.
[1045] 2. The server receives the audio data and converts it to text using the Google Speech-to-Text API.
[1046] 3. The server analyzes the text data "The project is going well, but I've been feeling tired lately" using a natural language processing module (e.g., SpaCy) and extracts the keyword "fatigue."
[1047] 4. The server uses a sentiment analysis module (e.g., TextBlob) to identify "fatigue."
[1048] 5. The server uses a big data analysis framework (e.g., Apache Hadoop) to calculate the "level of mental fatigue" and determine the risk as high.
[1049] 6. The server sends a "High Fatigue Risk" notification to the manager and employee using the specified notification method (e.g., email or Slack).
[1050] 7. The user reviews the notification and takes action such as following up with the employee or suggesting time off.
[1051] Prompt Sentence Examples
[1052] At the next project meeting, we will collect the comments of the members, convert them into text data, and analyze them. Please propose a method to visualize the emotional state and risk level and send a notification to the manager.
[1053] The flow of the identification process in the first embodiment will be described with reference to FIG.
[1054] Step 1:
[1055] The device detects the start of a meeting and begins capturing audio data. As input, it recognizes the meeting session using conferencing software such as Zoom or Microsoft Teams. As output, it converts the audio data captured from the microphone into WAV or MP3 format.
[1056] Specific behavior:
[1057] Activate the device's microphone when the meeting start event is triggered.
[1058] Audio data is collected in real time while being temporarily stored in a buffer.
[1059] Step 2:
[1060] The device sends the collected audio data to the server using the HTTPS protocol. The input is the audio data acquired in step 1. The output is an audio data file that is uploaded to the server.
[1061] Specific behavior:
[1062] The audio data is divided into packets of a fixed size.
[1063] The split data packets are sent to the server via HTTPS.
[1064] Step 3:
[1065] The server stores the received voice data in a database.,The input is the voice data sent from the terminal.,The output is the voice data stored in the database.
[1066] Specific behavior:
[1067] Connect to the database and select the table for storing the audio data.
[1068] The received audio data is properly indexed and written to a database.
[1069] Step 4:
[1070] The server converts the stored voice data into text data using a speech-to-text engine. The input is the voice data stored in the database. The output is the converted text data.
[1071] Specific behavior:
[1072] Split the audio data into short clips.
[1073] It calls the Google Speech-to-Text API or IBM Watson's Speech-to-Text service to convert audio clips individually into text.
[1074] Step 5:
[1075] The server inputs the text data into a natural language processing (NLP) module to analyze the content of the conversation. The input is the text data obtained in step 4. The output is extracted keywords and contextual information.
[1076] Specific behavior:
[1077] Use libraries such as SpaCy and NLTK to perform sentence segmentation, part-of-speech tagging, and named entity recognition.
[1078] Keywords such as "fatigue" and "project" are extracted and stored in a database.
[1079] Step 6:
[1080] The server inputs the text data into the sentiment analysis module to analyze the emotion and tone. The input is the text data obtained in step 5. The output is an analysis of the emotional state.
[1081] Specific behavior:
[1082] Classify emotional states using TextBlob and Hugging Face Transformers.
[1083] Identify emotions such as "joy," "sadness," "anger," and "fatigue" and calculate their intensity.
[1084] Step 7:
[1085] The server calculates the mental and physical risk levels based on the results of the sentiment analysis. The input is the sentiment analysis result obtained in step 6. The output is the calculated risk level.
[1086] Specific behavior:
[1087] The risk level is calculated using a big data analysis framework (e.g., Apache Hadoop or Spark).
[1088] The current risk level is evaluated by comparing it with past data and the results are stored in a database.
[1089] Step 8:
[1090] The server visualizes the calculated risk level using a data visualization tool. The input is the risk level calculated in step 7. The output is a graph or dashboard of the visualized risk level.
[1091] Specific behavior:
[1092] Use Grafana or Tableau to display the risk level in graphs and charts.
[1093] Configure the dashboard so users can understand the real-time risk situation.
[1094] Step 9:
[1095] The server sends notifications via email, Slack, etc. when the risk level exceeds a predetermined threshold. The input is the visualized risk level data. The output is a risk notification message that is generated and sent.
[1096] Specific behavior:
[1097] The risk level data is checked periodically and if a threshold is exceeded, a notification generation process is initiated.
[1098] Send messages in the appropriate notification format (email or chat message) for each user.
[1099] Step 10:
[1100] The user reviews the notifications and reports sent by the server and takes appropriate action. As input, there are notifications and reports sent by the server. As output, actions are taken, such as employee follow-ups and leave suggestions.
[1101] Specific behavior:
[1102] If you receive a notification, review its contents to understand the details of the risk level.
[1103] If necessary, communicate directly with employees and take measures such as suggesting time off if there is a high risk of fatigue.
[1104] Through the above steps, this system can optimize employee health management and help improve work efficiency.
[1105] (Application example 1)
[1106] Next, a description will be given of Application Example 1. In the following description, the data processing device 12 will be referred to as a "server" and the headset type terminal 314 will be referred to as a "terminal."
[1107] In security-related meetings, it is necessary to grasp important information and risk factors in real time without missing them and respond quickly. However, there is a lack of effective means to analyze the content of conversations and emotional states during meetings and instantly assess risks, so it is necessary to improve the responsiveness and accuracy of security management.
[1108] The specific processing by the specific processing unit 290 of the data processing device 12 in the application example 1 is realized by the following means.
[1109] In this invention, the server includes means for acquiring voice data, means for converting the voice data into text data, means for analyzing the text data and extracting the content of the conversation, means for identifying an emotional state from the text data, means for calculating and visualizing a risk level based on the identified emotional state, means for sending a notification when the risk level exceeds a predetermined threshold, means for acquiring voice data of a security conference in real time, identifying risk-related keywords, and evaluating the risk level, and means for notifying a security administrator based on the evaluation results. This makes it possible to quickly and accurately evaluate risks during a conference and take appropriate measures in a timely manner.
[1110] "Audio Data" means digitized audio information collected through a sound capture device such as a microphone.
[1111] "Text data" is data that has been analyzed and converted into textual information and is expressed in natural language.
[1112] "Emotional state" is a classification of a speaker's emotion or mental state identified from text data, and includes, for example, joy, sadness, anger, fatigue, etc.
[1113] "Risk level" is a numerical value or rating score that indicates the likelihood of potential danger or problem, calculated based on the identified emotional state.
[1114] "Security Administrator" means a person or team with a specialized role responsible for security-related monitoring, analysis, and response.
[1115] "Real time" refers to processing that responds quickly in time, with acquisition and analysis occurring almost simultaneously.
[1116] A "notification" is a message or warning sent from the system to a user (in this case, a security administrator) that contains important information or risk assessment results.
[1117] "Keywords" are important words or phrases extracted from the meeting content that serve as indicators of specific risks or issues.
[1118] "Visualization" refers to the visual representation of data and evaluation results, including displaying them in graph or dashboard format.
[1119] "Evaluation results" are specific results and numerical values derived from the analysis and evaluation process, and are information about the level of risk and emotional state that is communicated to managers.
[1120] In order to implement the present invention, the following hardware and software configuration is required.
[1121] Hardware and Software
[1122] 1. Smartphone: Used to capture voice data in real time.
[1123] 2. Microphone: Used to capture accurate voice data.
[1124] 3. Server: The device responsible for the main data processing and storage.
[1125] 4. Speech-to-Text engine: Software that converts voice data, such as Google Cloud Speech-to-Text, into text data.
[1126] 5. Natural Language Processing (NLP) module: Software such as spaCy that analyzes text data and extracts keywords.
[1127] 6. Sentiment Analysis Module: Software that identifies emotional states from text such as VADER.
[1128] 7. Database: A system that stores data, such as PostgreSQL.
[1129] 8. Visualization tools: Tools for visualizing risk levels, such as D3.js.
[1130] 9. Notification system: A system that sends notifications, such as Firebase Cloud Messaging.
[1131] Specific operation of the system
[1132] When the device detects the start of a meeting, it starts capturing audio data through the microphone, and transmits the captured audio data to the server in a specified format.
[1133] The server receives the voice data sent from the device and stores it in a designated database, which is then used in the subsequent analysis process.
[1134] The server calls the speech-to-text engine to convert the stored voice data into text data, which is then temporarily stored and passed to the analysis engine.
[1135] The server inputs the text data into a natural language processing module to analyze the content of the conversation. During this process, risk-related keywords are extracted. The extracted keywords and content are stored in a database for subsequent processing and report generation.
[1136] The server inputs the text data into a sentiment analysis module to analyze the emotions and tone of the conversation. During this process, emotional states such as "joy," "sadness," "anger," and "fatigue" are identified. A risk level is calculated by comparing with past data, and if an abnormality is detected, the information is stored in a database.
[1137] The server calculates the risk level based on the results of sentiment analysis. Using a big database, it compares the current risk level with past data and calculates the current risk level. The calculated risk level is visualized, and the impact and trends are displayed in graphs and dashboard format.
[1138] The server generates reports based on the results of sentiment analysis and risk assessment. If a risk signal is detected, a notification is sent to the security administrator. This notification is sent via email or chat app. Periodic reports are also generated and distributed to the administrator.
[1139] The user (security administrator) checks the feedback and reports sent from the server and takes measures. For example, if a particular risk is determined to be high, detailed investigation or emergency countermeasures can be implemented.
[1140] Examples and prompts
[1141] Examples:
[1142] If a member of a security meeting says, "There has been an increase in external attacks recently," this voice data is captured in real time, converted into text, and risk-related keywords such as "attack" and "external" are extracted. Sentiment analysis then identifies negative sentiment and evaluates it as a high risk. The results are visualized and a notification is sent to the security administrator.
[1143] Prompt statement:
[1144] 1. Speech to text:
[1145] Please tell me how to convert Japanese audio data into text using Google Cloud Speech-to-Text.
[1146] 2. Conversation content analysis:
[1147] Can you please give me the Python code to extract nouns and verbs from text using spaCy's Japanese model?
[1148] 3. Sentiment analysis:
[1149] How do I use VADER to evaluate the negative sentiment score of text?
[1150] 4. Risk Assessment and Visualization:
[1151] Can you please provide some Python code to visualize the risk scores in a pie chart using Matplotlib?
[1152] 5. Notification sending:
[1153] Can you please provide me some Python code to send a notification to a specific topic using Firebase Cloud Messaging?
[1154] By implementing the system in this way, security-related risks in meetings can be assessed quickly and accurately, and appropriate countermeasures can be taken.
[1155] The flow of the specific processing in the application example 1 will be described with reference to FIG.
[1156] Step 1:
[1157] When the device detects the start of a meeting, it starts capturing audio data through the microphone. The captured audio data is sent to the server in real time. The input is an analog audio signal captured from the microphone, and the output is digitized audio data. The device converts the signal into a digital format and sends it to the server.
[1158] Step 2:
[1159] The server receives the voice data sent from the device and stores it in a database. This process prepares the voice data for use in subsequent analysis processes. The input is digitized voice data, and the output is voice data stored in the database. Specifically, the server registers the data as an entry in the database and assigns a timestamp.
[1160] Step 3:
[1161] The server calls the Speech-to-Text engine to convert the stored voice data into text data. The converted text data is temporarily stored and passed to the analysis engine. The input is digitized voice data and the output is text data. The server sends the voice data to the Speech-to-Text engine and temporarily stores the obtained text data.
[1162] Step 4:
[1163] The server inputs the text data into a natural language processing (NLP) module and analyzes the content of the conversation. During this process, risk-related keywords are extracted. The extracted keywords and content are stored in a database and used for subsequent processing and report creation. The input is text data, and the output is a list of extracted keywords. The server analyzes the text data, extracts important keywords, and stores them.
[1164] Step 5:
[1165] The server inputs the text data into the sentiment analysis module to analyze the emotions and tones in the conversation. During this process, emotional states such as "joy," "sadness," "anger," and "fatigue" are identified. The input is text data, and the output is the classification result of the emotional state. The server analyzes the text data, identifies the emotional state, and stores it.
[1166] Step 6:
[1167] The server calculates the risk level based on the results of the emotion analysis. Using a big database, it compares it with past data to calculate the current risk level. The input is the classification result of the emotional state, and the output is a numerical value of the risk level. The server performs a comparison operation, calculates the risk level, and creates data for visualization.
[1168] Step 7:
[1169] The server generates a report based on the results of sentiment analysis and risk assessment. If a risk signal is detected, it sends a notification to the security administrator. This notification is sent via email or chat app. The input is the risk level number and sentiment classification result, and the output is a specific report and notification message. The server generates a report based on this data and sends it to the administrator using the notification system.
[1170] Step 8:
[1171] The user (security administrator) checks the feedback and reports sent from the server and takes appropriate measures. For example, if a particular risk is determined to be high, detailed investigations and emergency countermeasures can be implemented. The input is the reports and notifications from the server, and the output is the administrator's countermeasure actions. The user takes specific response steps based on the contents of the report.
[1172] In this way, the system can quickly and accurately assess security-related risks in meetings and take appropriate action in a timely manner.
[1173] Furthermore, an emotion engine that estimates the user's emotion may be further combined. That is, the identification processing unit 290 may estimate the user's emotion using the emotion identification model 59, and perform identification processing using the user's emotion.
[1174] The present invention relates to a system that starts by acquiring voice data, converts the voice data into text data, analyzes the text data to extract the content of the conversation and the emotional state, and calculates and visualizes the risk level. Furthermore, the present invention uses an emotion engine to recognize the user's emotions, and based on this, it is possible to more precisely evaluate the risk level. Below, each component of this system and its operation are specifically described.
[1175] System Embodiments
[1176] Acquiring audio data
[1177] The user enters the meeting details into the system and instructs the system to start the meeting. When the device detects the start of the meeting, it starts capturing audio data in real time through the microphone. The captured audio data is sent to the server in a specified format.
[1178] Saving audio data
[1179] The server receives the voice data sent from the device and stores it in a database, which is used in the subsequent analysis process.
[1180] Speech-to-Text (STT)
[1181] The server inputs the stored voice data into a Speech-to-Text (STT) engine and converts the voice into text data. For example, a speech saying, "The project is going well, but I've been feeling tired lately" is converted into text. This text data is temporarily stored.
[1182] Conversation content analysis
[1183] The server inputs the text data into a natural language processing (NLP) module and analyzes the content of the conversation. During this process, keywords related to the progress of work and the content of the consultation (e.g., "progress," "problem," "fatigue") are extracted. The extracted keywords and content are stored in a database and used for subsequent processing and report creation.
[1184] sentiment analysis
[1185] The server inputs the text data into an emotion engine, which analyzes the emotions and tone of the conversation. During this process, emotional states such as "happiness," "sadness," "anger," and "fatigue" are identified. The emotional state is compared with past data, and the difference from the usual state is calculated. The emotion engine also uses both voice and text data to identify the emotional state in detail. Furthermore, the emotion engine analyzes the user's facial expression data to identify the emotional state. This multi-layered emotion analysis enables more accurate risk calculations.
[1186] Risk Assessment
[1187] The server calculates the mental and physical risk level based on the results of the emotion analysis. It uses a big database to compare it with past data and calculate the current risk level. The results are visualized, and the impact and trends are displayed in graphs and dashboards.
[1188] Feedback and Notifications
[1189] The server generates reports based on the results of sentiment analysis and risk assessment. If a risk signal is detected, a notification is sent to the employee and their manager. This notification can be sent via email or a chat app, for example. Periodic reports are also generated and distributed to managers.
[1190] Reviewing feedback
[1191] Users can check the feedback and reports sent from the server, which can then be used to take measures regarding employees' work status and health status, such as suggesting that fatigued employees take time off.
[1192] Specific examples
[1193] scenario:
[1194] During a progress meeting for a project, one of the members said, "I've been feeling tired lately." Furthermore, it was confirmed that the member had a tired expression throughout the meeting.
[1195] How it works:
[1196] 1. The device collects the audio data of the meeting in real time and sends it to the server.
[1197] 2. The server receives the voice data and converts it into text data.
[1198] 3. The server analyzes the text data "The project is going well, but I've been feeling tired lately" and extracts the keyword "fatigue."
[1199] 4. The server uses the emotion engine to identify the "feeling of fatigue" and then analyzes the user's facial expression data to confirm that they are expressing fatigue.
[1200] 5. The server calculates the "mental fatigue level" and determines that the risk is high.
[1201] 6. The server sends a "high fatigue risk" notification to the manager and employee.
[1202] 7. The user reviews the notification and takes action, such as following up with the employee or suggesting time off.
[1203] In this way, the system of the present invention can perform multi-layered emotion analysis and effectively realize employee health management and productivity improvement throughout the organization.
[1204] The processing flow will be explained below.
[1205] Step 1:
[1206] A user logs in to the system and enters the date and time of the meeting, the participants, the purpose, etc. For example, the following information is registered in the system: "October 5th, 10:00, progress meeting, participants: Mr. A, Mr. B."
[1207] Step 2:
[1208] When the device detects the start of a meeting, it starts collecting audio data in real time through the microphone, and the collected audio data is sent to the server in a specified format.
[1209] Step 3:
[1210] The server receives the voice data sent from the device and stores it in a database, which is used in subsequent analysis processes.
[1211] Step 4:
[1212] The server inputs the stored voice data into a Speech-to-Text (STT) engine and converts the voice into text data. For example, a speech saying, "The project is going well, but I've been feeling tired lately" is converted into text. This converted text data is temporarily stored.
[1213] Step 5:
[1214] The server inputs the text data into a natural language processing (NLP) module and analyzes the content of the conversation. During this process, keywords related to the progress of work and the content of the consultation (e.g., "progress," "problems," "fatigue," etc.) are extracted. The extracted keywords and content are stored in a database.
[1215] Step 6:
[1216] The server inputs the text data into the emotion engine, which analyzes the emotions and tone of the conversation. During this process, emotional states such as "happiness," "sadness," "anger," and "fatigue" are identified. The emotion engine uses both the voice and text data to identify the emotional state in detail. The emotion engine also analyzes the user's facial expression data to complement the emotional state.
[1217] Step 7:
[1218] The server calculates the mental and physical risk level based on the results of the emotion analysis. It uses a big database to compare it with past data and calculate the current risk level. The results are visualized, and the impact and trends are displayed in graphs and dashboard format.
[1219] Step 8:
[1220] The server generates reports based on the results of sentiment analysis and risk assessment. If a risk signal is detected, a notification is sent to the employee and their manager. This notification can be sent via email or a chat app, for example. Periodic reports are also generated and distributed to managers.
[1221] Step 9:
[1222] Users can check the feedback and reports sent from the server, which can then be used to take measures regarding employees' work status and health status, such as suggesting that fatigued employees take time off.
[1223] Specific examples
[1224] scenario:
[1225] During a progress meeting for a project, a project member says, "I've been feeling tired lately." Furthermore, that member has a tired expression throughout the meeting.
[1226] How it works:
[1227] 1. The device collects the audio data of the meeting in real time and sends it to the server.
[1228] 2. The server receives the voice data and converts it into text data.
[1229] 3. The server analyzes the text data "The project is going well, but I've been feeling tired lately" and extracts the keyword "fatigue."
[1230] 4. The server uses the emotion engine to identify the "feeling of fatigue" and then analyzes the user's facial expression data to confirm the tiredness.
[1231] 5. The server calculates the "mental fatigue level" and determines that the risk is high.
[1232] 6. The server sends a "high fatigue risk" notification to the manager and employee.
[1233] 7. The user reviews the notification and takes action, such as following up with the employee or suggesting time off.
[1234] In this way, the system of the present invention can effectively manage employee health and improve productivity throughout the organization by performing multi-layered emotion analysis.
[1235] Example 2
[1236] Next, a description will be given of Example 2. In the following description, the data processing device 12 will be referred to as a "server" and the headset type terminal 314 will be referred to as a "terminal."
[1237] There is a problem in that it is difficult to grasp the work status and health status of employees in real time and respond quickly and accurately. In particular, it is difficult with conventional systems to precisely evaluate emotional states and risk levels and provide appropriate feedback and notifications. This poses a challenge in terms of improving work efficiency and effectively managing employee health.
[1238] The specific processing by the specific processing unit 290 of the data processing device 12 in the second embodiment is realized by the following means.
[1239] In this invention, the server includes means for acquiring voice data, means for converting voice data into text data, means for identifying an emotional state from the text data and voice data, means for calculating and visualizing a risk level by comparing it with past data, and means for sending a notification when the risk level exceeds a predetermined threshold. This makes it possible to precisely evaluate an employee's emotions and risk level and provide feedback and notifications at the appropriate time.
[1240] "Voice data" refers to digital data of voice or sound collected by a voice input device such as a microphone.
[1241] "Text data" is digital data that has been converted from voice data into character information.
[1242] "Emotional state" refers to the psychological and emotional state analyzed from the user's statements, voice data, facial expression data, etc.
[1243] "Risk level" is the degree of mental and physical risk calculated in comparison with emotional state and past data.
[1244] "Visualization" refers to visually displaying data such as risk level and emotional state in the form of graphs, dashboards, etc.
[1245] "Notification" refers to the sending of alerts or information to users or their superiors by the system based on specific conditions.
[1246] A "meeting" refers to a conference or meeting that takes place at a specific date and time, the details of which are entered into the system.
[1247] An "NLP (natural language processing) module" is a program module that analyzes text data and extracts conversation content and keywords.
[1248] An "emotion engine" is a software engine that analyzes voice data, text data, and facial expression data to identify emotional states.
[1249] A "big database" is a large-scale database for storing and analyzing large amounts of past data.
[1250] The system starts by acquiring voice data, converting it into text data, analyzing it to extract the content of the conversation and the emotional state, and then calculating and visualizing the risk level. It also uses an emotion engine to recognize the user's emotions and precisely evaluates the risk level based on that.
[1251] System Embodiments
[1252] Acquiring audio data
[1253] The user enters meeting details (date and time, participant list, purpose, etc.) through the system's user interface and presses the Start Meeting button to issue instructions. When the device detects the start of the meeting, it begins capturing audio data in real time via the microphone. The captured audio data is sent to the server in an appropriate format (e.g., WAV, MP3).
[1254] Saving audio data
[1255] The server receives the voice data sent from the terminal and stores it in a database, which is used in the subsequent analysis process.
[1256] Speech-to-text (STT)
[1257] The server inputs the stored voice data into a Speech-to-Text (STT) engine and converts the voice into text data. For example, a voice saying, "The project is going well, but I've been feeling tired lately" is converted into text data. This text data is temporarily stored.
[1258] Conversation content analysis
[1259] The server inputs the text data into a natural language processing (NLP) module and analyzes the content of the conversation. During this process, keywords related to the progress of work and the content of the consultation (e.g., "progress," "problem," "fatigue") are extracted. The extracted keywords and content are stored in a database and used for subsequent processing and report creation.
[1260] sentiment analysis
[1261] The server inputs text and voice data into an emotion engine to analyze the emotions and tone of the conversation. During this process, emotional states such as "happiness," "sadness," "anger," and "fatigue" are identified. The emotional state is compared with past data, and the difference from the usual state is calculated. Furthermore, in addition to voice and text data, facial expression data is used to identify the emotional state in more detail. This multi-layered emotion analysis enables more accurate risk calculations.
[1262] Risk Assessment
[1263] The server calculates the mental and physical risk level based on the results of the emotion analysis. It uses a big database to compare it with past data and calculate the current risk level. The results are visualized, and the impact and trends are displayed in graphs and dashboards.
[1264] Generate feedback and notifications
[1265] The server generates reports based on the results of sentiment analysis and risk assessment. If a risk signal is detected, a notification is sent to the employee and their manager. This notification can be sent via email or a chat app, for example. Periodic reports are also generated and distributed to managers.
[1266] Review feedback and implement measures
[1267] Users can check the feedback and reports sent from the server, which can then be used to take measures regarding employees' work status and health status, such as suggesting that fatigued employees take time off.
[1268] Specific examples
[1269] scenario:
[1270] During a progress meeting for a project, one of the members said, "I've been feeling tired lately." Furthermore, it was confirmed that the member had a tired expression throughout the meeting.
[1271] How it works:
[1272] 1. The device collects the audio data of the meeting in real time and sends it to the server.
[1273] 2. The server receives the voice data and converts it into text data.
[1274] 3. The server analyzes the text data "The project is going well, but I've been feeling tired lately" and extracts the keyword "fatigue."
[1275] 4. The server uses the emotion engine to identify the "feeling of fatigue" and then analyzes the user's facial expression data to confirm that they are expressing fatigue.
[1276] 5. The server calculates the "mental fatigue level" and determines that the risk is high.
[1277] 6. The server sends a "high fatigue risk" notification to the manager and employee.
[1278] 7. The user reviews the notification and takes action, such as following up with the employee or suggesting time off.
[1279] Prompt Sentence Examples
[1280] "Please explain how a system can collect and analyze everyone's comments during a status meeting, and then analyze their emotional state. Based on specific keywords and emotional tone, it can generate risk notifications."
[1281] In this way, the system of the present invention can perform multi-layered emotion analysis, effectively managing employee health and improving productivity throughout the organization.
[1282] The flow of the identification process in the second embodiment will be described with reference to FIG.
[1283] Step 1:
[1284] The user inputs detailed information about the meeting through the system's user interface and presses the start button to give instructions.
[1285] Input: Meeting details (date and time, participant list, purpose, etc.)
[1286] Output: Trigger to start a meeting
[1287] Specific operation: The user clicks the "Start Meeting" button using the system's web application, and the entered information is sent to the terminal.
[1288] Step 2:
[1289] When the terminal detects the start of a meeting, it starts capturing audio data in real time via the microphone.
[1290] Input: Meeting start trigger
[1291] Output: Real-time audio data
[1292] Specific behavior: The device's microphone will be turned on, a notification will appear saying "Audio collection during the meeting will begin," and audio data will begin to be collected.
[1293] Step 3:
[1294] The device sends the acquired audio data to the server in a specified format (e.g., WAV, MP3).
[1295] Input: Real-time audio data
[1296] Output: Audio data sent to the server in a given format
[1297] Specific operation: The device compresses the collected audio data and sends it to the server in the specified format.
[1298] Step 4:
[1299] The server receives the voice data sent from the terminal and stores it in a predetermined database.
[1300] Input: Audio data sent in a given format
[1301] Output: Audio data stored in a database
[1302] Specific operation: The server logs "Audio data received" and saves the file in the specified database.
[1303] Step 5:
[1304] The server inputs the stored voice data into a Speech-to-Text (STT) engine and converts the voice into text data.
[1305] Input: Audio data stored in a database
[1306] Output: Converted text data
[1307] Specific operation: The STT engine will start, and after a few seconds, the message "Speech to text conversion complete" will be displayed.
[1308] Step 6:
[1309] The server inputs the text data into a natural language processing (NLP) module to analyze the content of the conversation.
[1310] Input: Converted text data
[1311] Output: Extracted keywords and conversation content
[1312] Specific operation: The NLP module will start analysis and the message "Keyword extraction completed" will be displayed.
[1313] Step 7:
[1314] The server inputs the text and voice data into an emotion engine to analyze the emotion and tone of the conversation.
[1315] Input: Converted text data, audio data
[1316] Output: Identified emotional state
[1317] Specific behavior: The sentiment engine will be activated and the results will be displayed on the dashboard as "Sentiment analysis completed."
[1318] Step 8:
[1319] The server calculates the mental and physical risk levels based on the results of the emotion analysis.
[1320] Input: Identified emotional state
[1321] Output: Calculated risk level
[1322] Specific behavior: The risk assessment module will perform the analysis, the message "Risk assessment completed" will be displayed, and the results will be displayed in graphical form on the dashboard.
[1323] Step 9:
[1324] The server sends a notification if the risk exceeds a predetermined threshold.
[1325] Input: Calculated risk level
[1326] Output: Notification sent
[1327] Specific behavior: If the "high risk of fatigue" is determined, a notification will be created and the message "Notification sent" will be displayed.
[1328] Step 10:
[1329] Users check the feedback and reports sent from the server and take action.
[1330] Input: Notifications sent, reports
[1331] Output: Measures taken
[1332] Specific actions: The user receives a notification, checks the details on the dashboard, and clicks the "Suggest time off to employee" button to take action.
[1333] (Application example 2)
[1334] Next, a description will be given of Application Example 2. In the following description, the data processing device 12 will be referred to as a "server" and the headset type terminal 314 will be referred to as a "terminal."
[1335] Modern security services require rapid and accurate detection of potential risks posed by employees and visitors. However, previous technologies often only analyzed voice data, resulting in incomplete understanding of emotional states and making it difficult to accurately assess risk levels. Furthermore, the lack of an effective system for simultaneously performing real-time emotion analysis and risk assessment makes it difficult to respond quickly on-site. Therefore, the present invention aims to provide a system that uses multi-layered emotion analysis to integrate both voice data and facial expression data, enabling more accurate and rapid risk detection and assessment.
[1336] The specific processing by the specific processing unit 290 of the data processing device 12 in the application example 2 is realized by the following means.
[1337] In this invention, the server includes means for acquiring voice data, means for converting voice data into text data, means for analyzing the text data and extracting the content of the conversation, means for identifying an emotional state from the text data, means for calculating and visualizing a risk level based on the identified emotional state, means for sending a notification when the risk level exceeds a predetermined threshold, means for identifying an emotional state using both voice data and facial expression data, and a device (such as smart glasses) for analyzing the emotional state in real time. This enables integrated analysis of voice data and facial expression data, enabling more accurate risk assessment and rapid risk detection in real time.
[1338] "Audio data" is a digital representation of sound captured through an audio input device such as a microphone.
[1339] "Text data" is digital information that expresses voice data as characters.
[1340] The "content of the conversation" is information that indicates the topic and context extracted from the acquired text data.
[1341] "Emotional state" refers to the speaker's psychological state, which is determined by analyzing voice data, text data, and facial expression data.
[1342] "Risk level" is an assessment of potential danger calculated based on emotional state and conversation content.
[1343] "Visualization" is the process of displaying data in a graphical format (e.g., graphs, dashboards, etc.).
[1344] "Means for sending notifications" refers to a function that sends an alert via email, chat app, etc. when the risk level exceeds a specified threshold.
[1345] A "server" is a computer system that processes and stores voice and text data and analyzes them as needed.
[1346] "Apparatus for real-time analysis" refers to a device (e.g., smart glasses) that can acquire voice data and facial expression data in real time and perform analysis and processing immediately.
[1347] "Smart glasses" are eyeglass-type wearable devices that have built-in microphones and cameras and can acquire and analyze voice data and facial expression data.
[1348] overview
[1349] The present invention relates to a system that uses voice data and facial expression data to analyze emotional states and assess risk levels, particularly for use in security and surveillance applications, and is designed to enable security guards wearing smart glasses to more effectively detect risks.
[1350] System Configuration
[1351] The server operates to perform the following main functions:
[1352] 1. Acquiring audio data: Using the microphone built into the smart glasses, surrounding audio is collected in real time.
[1353] 2. Storage of voice data: The collected voice data is sent to a server and stored in a database.
[1354] 3. Speech-to-Text (STT): The stored voice data is converted into text data using a Speech-to-Text (STT) engine.
[1355] 4. Conversation content analysis: Text data is input into a natural language processing (NLP) module to extract the conversation content and important keywords.
[1356] 5. Sentiment analysis: Text data and facial expression data are input into the emotion engine to analyze the emotional state.
[1357] 6. Risk assessment: Based on the results of sentiment analysis, the risk level is calculated and visualized.
[1358] 7. Feedback and Notification: Send notifications to users and supervisors when risks are detected that exceed predetermined thresholds.
[1359] Specific examples
[1360] A security guard at a department store is patrolling while wearing smart glasses. The guard collects a visitor's statement, "I'm frustrated because there have been a lot of strange things happening at the shopping mall recently." Based on this statement, the system extracts keywords such as "strange" and "frustrated," and also detects anger from the visitor's facial expression. Combining this information, the system assesses the person as being at high risk and displays a notification on the guard's HUD indicating that the person is at high risk.
[1361] Hardware and software used
[1362] Hardware: Smart glasses with built-in microphone and camera
[1363] Software: Speech Recognition Library (speech-to-text conversion), NLP Module (natural language processing), Emotion Engine (sentiment analysis), Risk Assessment Module (risk assessment)
[1364] Prompt Sentence Examples
[1365] We want to design a system that analyzes the emotional state of a user from their voice data and facial expression data, and evaluates the risk level in real time. This system has the following functions:
[1366] 1. Acquiring audio data and converting it to text
[1367] 2. Content analysis of text data
[1368] 3. Emotional state analysis (based on voice and facial expression data)
[1369] 4. Calculating and visualizing risk levels
[1370] Generate pseudocode for a program that combines these elements.
[1371] The flow of the specific processing in the application example 2 will be described with reference to FIG.
[1372] Step 1:
[1373] Acquiring audio data
[1374] The server captures voice data in real time through the microphone of the smart glasses. This data is sent to the server as voice input. The input is the voice data uttered by the user, and the output is the voice file sent to the server.
[1375] Step 2:
[1376] Saving audio data
[1377] The server stores the captured audio data in a database, which is used in the subsequent analysis process. The input is the audio file sent to the server, and the output is the audio data stored in the database.
[1378] Step 3:
[1379] Speech-to-text (STT)
[1380] The server inputs the stored voice data into a Speech-to-Text (STT) engine and converts the voice into text data. The input is the voice data stored in the database, and the output is text data. This allows the voice data to be treated as text information.
[1381] Step 4:
[1382] Conversation content analysis
[1383] The server inputs the text data into a natural language processing (NLP) module and analyzes the content of the conversation. During this process, important keywords (e.g., "progress," "problems," and "fatigue") are extracted. The input is text data, and the output is a list of extracted keywords.
[1384] Step 5:
[1385] sentiment analysis
[1386] The server inputs text data and facial expression data captured by the smart glasses' built-in camera into the emotion engine to analyze the emotional state. During this process, emotions such as "happiness," "sadness," "anger," and "fatigue" are identified. The input is text data and facial expression data, and the output is the identified emotional state.
[1387] Step 6:
[1388] Risk Assessment
[1389] The server calculates the risk level based on the emotional state and the results of the conversation content analysis. It uses a big database to compare it with past data and evaluate the current risk level. The input is the emotional state and a list of keywords, and the output is the calculated risk level.
[1390] Step 7:
[1391] Feedback and Notifications
[1392] The server sends a notification to the user and supervisor if the calculated risk level exceeds a predetermined threshold. This notification is sent via email or a chat app. The input is the risk level and the output is the notification message.
[1393] Step 8:
[1394] Display and confirmation
[1395] The user checks the results of the risk level and emotional state on the head-up display (HUD) of the smart glasses. The input is a notification message, and the output is the information displayed on the HUD. This allows the user to understand the situation in real time and respond quickly.
[1396] The specific processing unit 290 transmits the result of the specific processing to the headset type terminal 314. In the headset type terminal 314, the control unit 46A causes the speaker 240 and the display 343 to output the result of the specific processing. The microphone 238 acquires audio indicating a user input regarding the result of the specific processing. The control unit 46A transmits audio data indicating the user input acquired by the microphone 238 to the data processing device 12. In the data processing device 12, the specific processing unit 290 acquires the audio data.
[1397] The data generation model 58 is a so-called generative AI (Artificial Intelligence). An example of the data generation model 58 is ChatGPT (Internet Search<URL: https: / / openai.com / blog / chatgpt> ), Gemini (Internet search <url: https: gemini.google.com ?hl="ja">) and other generation AIs. The data generation model 58 is obtained by performing deep learning on a neural network. A prompt including an instruction is input to the data generation model 58, and inference data such as voice data indicating voice, text data indicating text, and image data indicating an image is also input. The data generation model 58 performs inference on the input inference data in accordance with the instruction indicated by the prompt, and outputs the inference result in a data format such as voice data and text data. Here, inference refers to, for example, analysis, classification, prediction, and / or summarization.
[1398] In the above embodiment, an example was given in which the specific processing is performed by the data processing device 12, but the technology of the present disclosure is not limited to this, and the specific processing may be performed by the headset type terminal 314.
[1399] [Fourth embodiment]
[1400] FIG. 7 shows an example of the configuration of a data processing system 410 according to the fourth embodiment.
[1401] 7, a data processing system 410 includes a data processing device 12 and a robot 414. An example of the data processing device 12 is a server.
[1402] The data processing device 12 includes a computer 22, a database 24, and a communication I / F 26. The computer 22 is an example of a "computer" according to the technology of the present disclosure. The computer 22 includes a processor 28, a RAM 30, and a storage 32. The processor 28, the RAM 30, and the storage 32 are connected to a bus 34. The database 24 and the communication I / F 26 are also connected to the bus 34. The communication I / F 26 is connected to a network 54. Examples of the network 54 include a WAN (Wide Area Network) and / or a LAN (Local Area Network).
[1403] The robot 414 includes a computer 36, a microphone 238, a speaker 240, a camera 42, a communication I / F 44, and a control target 443. The computer 36 includes a processor 46, a RAM 48, and a storage 50. The processor 46, the RAM 48, and the storage 50 are connected to a bus 52. The microphone 238, the speaker 240, the camera 42, and the control target 443 are also connected to the bus 52.
[1404] The microphone 238 receives instructions and the like from the user 20 by receiving voice uttered by the user 20. The microphone 238 captures the voice uttered by the user 20, converts the captured voice into audio data, and outputs it to the processor 46. The speaker 240 outputs audio in accordance with instructions from the processor 46.
[1405] Camera 42 is a small digital camera equipped with an optical system including a lens, aperture, and shutter, and an imaging element such as a CMOS (Complementary Metal-Oxide-Semiconductor) image sensor or a CCD (Charge Coupled Device) image sensor, and captures images of the surroundings of user 20 (for example, an imaging range defined by an angle of view equivalent to the field of vision of a typical healthy person).
[1406] The communication I / F 44 is connected to a network 54. The communication I / Fs 44 and 26 are responsible for the exchange of various information between the processor 46 and the processor 28 via the network 54. The exchange of various information between the processor 46 and the processor 28 using the communication I / Fs 44 and 26 is carried out in a secure state.
[1407] The control object 443 includes a display device, LEDs in the eyes, and motors for driving the arms, hands, and feet. The posture and gestures of the robot 414 are controlled by controlling the motors of the arms, hands, and feet. Some of the emotions of the robot 414 can be expressed by controlling these motors. In addition, the facial expressions of the robot 414 can also be expressed by controlling the light emission state of the LEDs in the eyes of the robot 414.
[1408] Fig. 8 shows an example of the main functions of the data processing device 12 and the robot 414. As shown in Fig. 8, in the data processing device 12, a specific process is performed by the processor 28. A specific process program 56 is stored in the storage 32.
[1409] The specific processing program 56 is an example of a "program" according to the technology of the present disclosure. The processor 28 reads the specific processing program 56 from the storage 32 and executes the read specific processing program 56 on the RAM 30. The specific processing is realized by the processor 28 operating as a specific processing unit 290 in accordance with the specific processing program 56 executed on the RAM 30.
[1410] The storage 32 stores a data generation model 58 and an emotion identification model 59. The data generation model 58 and the emotion identification model 59 are used by the identification processing unit 290.
[1411] In the robot 414, the processor 46 performs the reception output process. A reception output program 60 is stored in the storage 50. The processor 46 reads the reception output program 60 from the storage 50 and executes the read reception output program 60 on the RAM 48. The reception output process is realized by the processor 46 operating as the control unit 46A in accordance with the reception output program 60 executed on the RAM 48.
[1412] Next, a description will be given of the specific processing performed by the specific processing unit 290 of the data processing device 12. In the following description, the data processing device 12 will be referred to as a "server" and the robot 414 will be referred to as a "terminal."
[1413] The present invention relates to a system that starts by acquiring voice data, converts the voice data into text data, analyzes the text data to extract the content of the conversation and the emotional state, and calculates and visualizes the risk level. Below, we will explain in detail each component of this system and its operation.
[1414] System Embodiments
[1415] Acquiring audio data
[1416] When the device detects the start of a meeting, it starts capturing audio data in real time through the microphone, and sends the captured audio data to the server in a specified format.
[1417] Saving audio data
[1418] The server receives the voice data sent from the device and stores it in a designated database, which is used in the subsequent analysis process.
[1419] Speech-to-Text (STT)
[1420] The server calls the speech-to-text engine to convert the stored voice data into text data. For example, a speech saying, "The project is going well, but I've been feeling tired lately" is converted into text. This converted text data is temporarily stored and passed to the analysis engine.
[1421] Conversation content analysis
[1422] The server inputs the text data into a natural language processing (NLP) module and analyzes the content of the conversation. During this process, keywords related to the progress of work and the content of the consultation (e.g., "progress," "problem," "fatigue") are extracted. The extracted keywords and content are stored in a database and used for subsequent processing and report creation.
[1423] sentiment analysis
[1424] The server inputs the text data into a sentiment analysis module to analyze the emotions and tone of the conversation. During this process, emotional states such as "happiness," "sadness," "anger," and "fatigue" are identified. The emotional state is compared with past data, and the difference from the normal state is calculated. If an anomaly is detected, the information is stored in a database.
[1425] Risk Assessment
[1426] The server calculates mental and physical risk levels based on the results of emotion analysis. Using a big database, it compares current risk levels with past data and calculates the current risk level. The calculated risk level is visualized, and the impact and trends are displayed in graphs and dashboard format.
[1427] Feedback and Notifications
[1428] The server generates reports based on the results of sentiment analysis and risk assessment. If a risk signal is detected, a notification is sent to the employee and their manager. This notification can be sent via email or a chat app, for example. Periodic reports are also generated and distributed to managers.
[1429] Reviewing feedback
[1430] Users can check the feedback and reports sent from the server, which can then be used to take measures regarding employees' work status and health status, for example, suggesting that fatigued employees take time off.
[1431] Specific examples
[1432] scenario:
[1433] A progress meeting is held for a project, and a project member says, "I've been feeling tired lately."
[1434] How it works:
[1435] 1. The device collects the audio data of the meeting in real time and sends it to the server.
[1436] 2. The server receives the voice data and converts it into text data.
[1437] 3. The server analyzes the text data "The project is going well, but I've been feeling tired lately" and extracts the keyword "fatigue."
[1438] 4. The server uses the emotion analysis module to identify the "feeling of fatigue."
[1439] 5. The server calculates the "mental fatigue level" and determines that the risk is high.
[1440] 6. The server sends a "high fatigue risk" notification to the manager and employee.
[1441] 7. The user reviews the notification and takes action such as following up with the employee or suggesting time off.
[1442] The above is a specific embodiment of the present invention. This system can effectively manage the health of employees and improve productivity throughout the organization.
[1443] The processing flow will be explained below.
[1444] Step 1:
[1445] A user logs in to the system and enters information such as the date and time of the meeting, participants, purpose, etc. For example, information such as "October 5th, 10:00, progress meeting, participants: Mr. A, Mr. B" is registered in the system.
[1446] Step 2:
[1447] When the device detects the start of a meeting, it starts collecting audio data in real time through the microphone, and sends the collected audio data to the server in a specified format.
[1448] Step 3:
[1449] The server receives the voice data sent from the device and stores it in a database, which is used in the subsequent analysis process.
[1450] Step 4:
[1451] The server inputs the stored voice data into a Speech-to-Text (STT) engine and converts the voice into text data. For example, a speech saying, "The project is going well, but I've been feeling tired lately" is converted into text. This text data is temporarily stored.
[1452] Step 5:
[1453] The server inputs the text data into a natural language processing (NLP) module and analyzes the content of the conversation. During this process, keywords related to the progress of work and the content of the consultation (e.g., "progress," "problem," "fatigue") are extracted. The extracted keywords and content are stored in a database and used for subsequent processing and report creation.
[1454] Step 6:
[1455] The server inputs the text data into a sentiment analysis module to analyze the emotions and tone of the conversation. During this process, emotional states such as "happiness," "sadness," "anger," and "fatigue" are identified. The emotional state is compared with past data, and the difference from the normal state is calculated. If an anomaly is detected, the information is stored in a database.
[1456] Step 7:
[1457] The server calculates the mental and physical risk level based on the results of the emotion analysis. It uses a big database to compare it with past data and calculate the current risk level. The results are visualized, and the impact and trends are displayed in graphs and dashboard format.
[1458] Step 8:
[1459] The server generates reports based on the results of sentiment analysis and risk assessment. If a risk signal is detected, a notification is sent to the employee and their supervisor so that they can take action. This notification can be sent via email or a chat app, for example. Periodic reports are also generated and distributed to supervisors.
[1460] Step 9:
[1461] The user checks the feedback and reports sent from the server, which allows the user to take measures regarding the work situation and health status of employees. For example, the user can suggest that a fatigued employee take time off.
[1462] Example 1
[1463] Next, a description will be given of Example 1. In the following description, the data processing device 12 will be referred to as a "server" and the robot 414 will be referred to as a "terminal."
[1464] There is a need to monitor employees' mental and physical health conditions in real time to prevent declines in work efficiency and productivity. However, conventional methods have made it difficult to properly analyze the content of comments and emotional states made during meetings, assess risks, and take prompt measures. The present invention aims to solve these problems and provide a system for effectively managing employees' health conditions.
[1465] The specific processing by the specific processing unit 290 of the data processing device 12 in the first embodiment is realized by the following means.
[1466] In this invention, the server includes means for acquiring voice data, means for converting the voice data into text data, means for analyzing the text data and extracting the content of the conversation, means for identifying an emotional state from the text data, means for calculating and visualizing a risk level based on the identified emotional state, means for sending a notification when the risk level exceeds a predetermined threshold, means for a terminal to detect the start of a meeting and send the voice data to the server in real time, means for the server to store the voice data in a database and convert the voice data into text data using a speech-to-text engine, means for the server to extract keywords related to work progress and consultation content from the text data using a natural language processing module, means for the server to analyze the emotional state using a sentiment analysis module, means for the server to calculate a risk level using a big data analysis framework and visualize the risk level using a data visualization tool, means for the server to generate a report and send a notification to the employee and their supervisor using a predetermined notification means, and means for a user to check the notification from the server and take measures. This enables real-time monitoring of employees' mental and physical health conditions and rapid visualization and notification of risks.
[1467] "Audio data" refers to data in which voice or acoustic information is recorded in digital format.
[1468] "Text data" refers to data obtained by converting voice data into character information.
[1469] A "natural language processing (NLP) module" is software and algorithms used to analyze text data and understand its content and meaning.
[1470] "Sentiment Analysis Module" means software and algorithms for identifying emotional states from text data.
[1471] "Risk level" is a calculated degree of risk based on the identified emotional state and other data.
[1472] "Visualization" is the presentation of data in a visual format such as a graph, chart, or dashboard.
[1473] "Notification" means sending a warning or information to the user when a specific condition is met.
[1474] A "terminal" is a device that acquires voice data and transmits it to a server, and includes devices such as personal computers and smartphones.
[1475] A "server" is a computer system for storing and processing audio data.
[1476] A "database" is a system that systematically stores and manages information.
[1477] A "Speech-to-Text engine" is software for converting voice data into text data.
[1478] "Work progress status" is information that indicates the progress of projects and tasks.
[1479] A "big data analytics framework" is software and a platform for analyzing large amounts of data and extracting meaningful information.
[1480] A "data visualization tool" is a tool for visually displaying data, and includes tools such as Grafana and Tableau.
[1481] A "report" is a document that summarizes the results of analysis and evaluation.
[1482] "User" means a person who uses the System to review feedback and notifications and take appropriate action.
[1483] "Detecting the start of a meeting" means that the system automatically recognizes the start of a meeting.
[1484] System Overview
[1485] This invention is a system that captures audio data during meetings in real time, analyzes it to extract the content of the conversation and the emotional state, and calculates and visualizes the mental and physical risk levels. This system consists of three main components: the terminal, the server, and the user. The following describes the details of each component and their operation.
[1486] Acquiring audio data
[1487] When the device detects the start of a meeting, it starts capturing audio data using the built-in microphone or an external microphone. This detection can be done using conferencing software such as Zoom or Microsoft Teams. The audio data is sent to the server in real time in a specified format (WAV, MP3, etc.). This transmission uses the HTTPS protocol.
[1488] Saving audio data
[1489] The server receives the voice data sent from the device and stores it in a database (e.g., MySQL or PostgreSQL). During this storage process, the voice data is properly indexed and prepared for subsequent analysis.
[1490] Speech-to-Text (STT)
[1491] The server calls the Google Speech-to-Text API or IBM Watson's Speech-to-Text service to convert the stored audio data into text data. This process involves splitting the audio data into short clips and converting each into text. This converted text data is temporarily stored in a database.
[1492] Conversation content analysis
[1493] The server inputs the text data into a natural language processing (NLP) module to analyze the content of the conversation. Specifically, it uses libraries such as SpaCy and NLTK to split sentences, tag parts of speech, and recognize named entities. During this process, it extracts keywords related to the progress of work and the content of the consultation. The extracted keywords are stored in a database and used for subsequent processing.
[1494] sentiment analysis
[1495] The server analyzes the emotional state from the text data using an emotion analysis module (e.g., TextBlob or Hugging Face Transformers). Here, emotional states such as "joy," "sadness," "anger," and "fatigue" are identified, and their intensity and frequency are calculated. If an abnormality in the emotion is detected, the information is stored in a database.
[1496] Risk Assessment
[1497] The server calculates the mental and physical risk level based on the results of the sentiment analysis. This risk level is calculated using a big data analysis framework (e.g., Apache Hadoop or Spark). The risk level is calculated through comparison with past data and visualized using a data visualization tool (e.g., Grafana or Tableau).
[1498] Feedback and Notifications
[1499] Based on the analysis results and risk assessment, the server generates regular reports and specific notifications, which are sent via email, Slack, Microsoft Teams, etc. The notifications include details of the risk level and recommended countermeasures.
[1500] Reviewing feedback
[1501] Users can review the feedback and reports sent from the server and take appropriate measures for the mental and physical health of employees. For example, if a high-risk employee is notified, the user can directly follow up with the employee and take measures such as suggesting time off if necessary.
[1502] Specific examples
[1503] As a specific example, consider a situation where a progress meeting is held for a project and a project member says, "I've been feeling tired lately."
[1504] 1. The device collects the audio data of the meeting in real time and sends it to the server.
[1505] 2. The server receives the audio data and converts it to text using the Google Speech-to-Text API.
[1506] 3. The server analyzes the text data "The project is going well, but I've been feeling tired lately" using a natural language processing module (e.g., SpaCy) and extracts the keyword "fatigue."
[1507] 4. The server uses a sentiment analysis module (e.g., TextBlob) to identify "fatigue."
[1508] 5. The server uses a big data analysis framework (e.g., Apache Hadoop) to calculate the "level of mental fatigue" and determine the risk as high.
[1509] 6. The server sends a "High Fatigue Risk" notification to the manager and employee using the specified notification method (e.g., email or Slack).
[1510] 7. The user reviews the notification and takes action such as following up with the employee or suggesting time off.
[1511] Prompt Sentence Examples
[1512] At the next project meeting, we will collect the comments of the members, convert them into text data, and analyze them. Please propose a method to visualize the emotional state and risk level and send a notification to the manager.
[1513] The flow of the identification process in the first embodiment will be described with reference to FIG.
[1514] Step 1:
[1515] The device detects the start of a meeting and begins capturing audio data. As input, it recognizes the meeting session using conferencing software such as Zoom or Microsoft Teams. As output, it converts the audio data captured from the microphone into WAV or MP3 format.
[1516] Specific behavior:
[1517] Activate the device's microphone when the meeting start event is triggered.
[1518] Audio data is collected in real time while being temporarily stored in a buffer.
[1519] Step 2:
[1520] The device sends the collected audio data to the server using the HTTPS protocol. The input is the audio data acquired in step 1. The output is an audio data file that is uploaded to the server.
[1521] Specific behavior:
[1522] The audio data is divided into packets of a fixed size.
[1523] The split data packets are sent to the server via HTTPS.
[1524] Step 3:
[1525] The server stores the received voice data in a database.,The input is the voice data sent from the terminal.,The output is the voice data stored in the database.
[1526] Specific behavior:
[1527] Connect to the database and select the table for storing the audio data.
[1528] The received audio data is properly indexed and written to a database.
[1529] Step 4:
[1530] The server converts the stored voice data into text data using a speech-to-text engine. The input is the voice data stored in the database. The output is the converted text data.
[1531] Specific behavior:
[1532] Split the audio data into short clips.
[1533] It calls the Google Speech-to-Text API or IBM Watson's Speech-to-Text service to convert audio clips individually into text.
[1534] Step 5:
[1535] The server inputs the text data into a natural language processing (NLP) module to analyze the content of the conversation. The input is the text data obtained in step 4. The output is extracted keywords and contextual information.
[1536] Specific behavior:
[1537] Use libraries such as SpaCy and NLTK to perform sentence segmentation, part-of-speech tagging, and named entity recognition.
[1538] Keywords such as "fatigue" and "project" are extracted and stored in a database.
[1539] Step 6:
[1540] The server inputs the text data into the sentiment analysis module to analyze the emotion and tone. The input is the text data obtained in step 5. The output is an analysis of the emotional state.
[1541] Specific behavior:
[1542] Classify emotional states using TextBlob and Hugging Face Transformers.
[1543] Identify emotions such as "joy," "sadness," "anger," and "fatigue" and calculate their intensity.
[1544] Step 7:
[1545] The server calculates the mental and physical risk levels based on the results of the sentiment analysis. The input is the sentiment analysis result obtained in step 6. The output is the calculated risk level.
[1546] Specific behavior:
[1547] The risk level is calculated using a big data analysis framework (e.g., Apache Hadoop or Spark).
[1548] The current risk level is evaluated by comparing it with past data and the results are stored in a database.
[1549] Step 8:
[1550] The server visualizes the calculated risk level using a data visualization tool. The input is the risk level calculated in step 7. The output is a graph or dashboard of the visualized risk level.
[1551] Specific behavior:
[1552] Use Grafana or Tableau to display the risk level in graphs and charts.
[1553] Configure the dashboard so users can understand the real-time risk situation.
[1554] Step 9:
[1555] The server sends notifications via email, Slack, etc. when the risk level exceeds a predetermined threshold. The input is the visualized risk level data. The output is a risk notification message that is generated and sent.
[1556] Specific behavior:
[1557] The risk level data is checked periodically and if a threshold is exceeded, a notification generation process is initiated.
[1558] Send messages in the appropriate notification format (email or chat message) for each user.
[1559] Step 10:
[1560] The user reviews the notifications and reports sent by the server and takes appropriate action. As input, there are notifications and reports sent by the server. As output, actions are taken, such as employee follow-ups and leave suggestions.
[1561] Specific behavior:
[1562] If you receive a notification, review its contents to understand the details of the risk level.
[1563] If necessary, communicate directly with employees and take measures such as suggesting time off if there is a high risk of fatigue.
[1564] Through the above steps, this system can optimize employee health management and help improve work efficiency.
[1565] (Application example 1)
[1566] Next, a description will be given of Application Example 1. In the following description, the data processing device 12 will be referred to as a "server" and the robot 414 will be referred to as a "terminal."
[1567] In security-related meetings, it is necessary to grasp important information and risk factors in real time without missing them and respond quickly. However, there is a lack of effective means to analyze the content of conversations and emotional states during meetings and instantly assess risks, so it is necessary to improve the responsiveness and accuracy of security management.
[1568] The specific processing by the specific processing unit 290 of the data processing device 12 in the application example 1 is realized by the following means.
[1569] In this invention, the server includes means for acquiring voice data, means for converting the voice data into text data, means for analyzing the text data and extracting the content of the conversation, means for identifying an emotional state from the text data, means for calculating and visualizing a risk level based on the identified emotional state, means for sending a notification when the risk level exceeds a predetermined threshold, means for acquiring voice data of a security conference in real time, identifying risk-related keywords, and evaluating the risk level, and means for notifying a security administrator based on the evaluation results. This makes it possible to quickly and accurately evaluate risks during a conference and take appropriate measures in a timely manner.
[1570] "Audio Data" means digitized audio information collected through a sound capture device such as a microphone.
[1571] "Text data" is data that has been analyzed and converted into textual information and is expressed in natural language.
[1572] "Emotional state" is a classification of a speaker's emotion or mental state identified from text data, and includes, for example, joy, sadness, anger, fatigue, etc.
[1573] "Risk level" is a numerical value or rating score that indicates the likelihood of potential danger or problem, calculated based on the identified emotional state.
[1574] "Security Administrator" means a person or team with a specialized role responsible for security-related monitoring, analysis, and response.
[1575] "Real time" refers to processing that responds quickly in time, with acquisition and analysis occurring almost simultaneously.
[1576] A "notification" is a message or warning sent from the system to a user (in this case, a security administrator) that contains important information or risk assessment results.
[1577] "Keywords" are important words or phrases extracted from the meeting content that serve as indicators of specific risks or issues.
[1578] "Visualization" refers to the visual representation of data and evaluation results, including displaying them in graph or dashboard format.
[1579] "Evaluation results" are specific results and numerical values derived from the analysis and evaluation process, and are information about the level of risk and emotional state that is communicated to managers.
[1580] In order to implement the present invention, the following hardware and software configuration is required.
[1581] Hardware and Software
[1582] 1. Smartphone: Used to capture voice data in real time.
[1583] 2. Microphone: Used to capture accurate voice data.
[1584] 3. Server: The device responsible for the main data processing and storage.
[1585] 4. Speech-to-Text engine: Software that converts voice data, such as Google Cloud Speech-to-Text, into text data.
[1586] 5. Natural Language Processing (NLP) module: Software such as spaCy that analyzes text data and extracts keywords.
[1587] 6. Sentiment Analysis Module: Software that identifies emotional states from text such as VADER.
[1588] 7. Database: A system that stores data, such as PostgreSQL.
[1589] 8. Visualization tools: Tools for visualizing risk levels, such as D3.js.
[1590] 9. Notification system: A system that sends notifications, such as Firebase Cloud Messaging.
[1591] Specific operation of the system
[1592] When the device detects the start of a meeting, it starts capturing audio data through the microphone, and transmits the captured audio data to the server in a specified format.
[1593] The server receives the voice data sent from the device and stores it in a designated database, which is then used in the subsequent analysis process.
[1594] The server calls the speech-to-text engine to convert the stored voice data into text data, which is then temporarily stored and passed to the analysis engine.
[1595] The server inputs the text data into a natural language processing module to analyze the content of the conversation. During this process, risk-related keywords are extracted. The extracted keywords and content are stored in a database for subsequent processing and report generation.
[1596] The server inputs the text data into a sentiment analysis module to analyze the emotions and tone of the conversation. During this process, emotional states such as "joy," "sadness," "anger," and "fatigue" are identified. A risk level is calculated by comparing with past data, and if an abnormality is detected, the information is stored in a database.
[1597] The server calculates the risk level based on the results of sentiment analysis. Using a big database, it compares the current risk level with past data and calculates the current risk level. The calculated risk level is visualized, and the impact and trends are displayed in graphs and dashboard format.
[1598] The server generates reports based on the results of sentiment analysis and risk assessment. If a risk signal is detected, a notification is sent to the security administrator. This notification is sent via email or chat app. Periodic reports are also generated and distributed to the administrator.
[1599] The user (security administrator) checks the feedback and reports sent from the server and takes measures. For example, if a particular risk is determined to be high, detailed investigation or emergency countermeasures can be implemented.
[1600] Examples and prompts
[1601] Examples:
[1602] If a member of a security meeting says, "There has been an increase in external attacks recently," this voice data is captured in real time, converted into text, and risk-related keywords such as "attack" and "external" are extracted. Sentiment analysis then identifies negative sentiment and evaluates it as a high risk. The results are visualized and a notification is sent to the security administrator.
[1603] Prompt statement:
[1604] 1. Speech to text:
[1605] Please tell me how to convert Japanese audio data into text using Google Cloud Speech-to-Text.
[1606] 2. Conversation content analysis:
[1607] Can you please give me the Python code to extract nouns and verbs from text using spaCy's Japanese model?
[1608] 3. Sentiment analysis:
[1609] How do I use VADER to evaluate the negative sentiment score of text?
[1610] 4. Risk Assessment and Visualization:
[1611] Can you please provide some Python code to visualize the risk scores in a pie chart using Matplotlib?
[1612] 5. Notification sending:
[1613] Can you please provide me some Python code to send a notification to a specific topic using Firebase Cloud Messaging?
[1614] By implementing the system in this way, security-related risks in meetings can be assessed quickly and accurately, and appropriate countermeasures can be taken.
[1615] The flow of the specific processing in the application example 1 will be described with reference to FIG.
[1616] Step 1:
[1617] When the device detects the start of a meeting, it starts capturing audio data through the microphone. The captured audio data is sent to the server in real time. The input is an analog audio signal captured from the microphone, and the output is digitized audio data. The device converts the signal into a digital format and sends it to the server.
[1618] Step 2:
[1619] The server receives the voice data sent from the device and stores it in a database. This process prepares the voice data for use in subsequent analysis processes. The input is digitized voice data, and the output is voice data stored in the database. Specifically, the server registers the data as an entry in the database and assigns a timestamp.
[1620] Step 3:
[1621] The server calls the Speech-to-Text engine to convert the stored voice data into text data. The converted text data is temporarily stored and passed to the analysis engine. The input is digitized voice data and the output is text data. The server sends the voice data to the Speech-to-Text engine and temporarily stores the obtained text data.
[1622] Step 4:
[1623] The server inputs the text data into a natural language processing (NLP) module and analyzes the content of the conversation. During this process, risk-related keywords are extracted. The extracted keywords and content are stored in a database and used for subsequent processing and report creation. The input is text data, and the output is a list of extracted keywords. The server analyzes the text data, extracts important keywords, and stores them.
[1624] Step 5:
[1625] The server inputs the text data into the sentiment analysis module to analyze the emotions and tones in the conversation. During this process, emotional states such as "joy," "sadness," "anger," and "fatigue" are identified. The input is text data, and the output is the classification result of the emotional state. The server analyzes the text data, identifies the emotional state, and stores it.
[1626] Step 6:
[1627] The server calculates the risk level based on the results of the emotion analysis. Using a big database, it compares it with past data to calculate the current risk level. The input is the classification result of the emotional state, and the output is a numerical value of the risk level. The server performs a comparison operation, calculates the risk level, and creates data for visualization.
[1628] Step 7:
[1629] The server generates a report based on the results of sentiment analysis and risk assessment. If a risk signal is detected, it sends a notification to the security administrator. This notification is sent via email or chat app. The input is the risk level number and sentiment classification result, and the output is a specific report and notification message. The server generates a report based on this data and sends it to the administrator using the notification system.
[1630] Step 8:
[1631] The user (security administrator) checks the feedback and reports sent from the server and takes appropriate measures. For example, if a particular risk is determined to be high, detailed investigations and emergency countermeasures can be implemented. The input is the reports and notifications from the server, and the output is the administrator's countermeasure actions. The user takes specific response steps based on the contents of the report.
[1632] In this way, the system can quickly and accurately assess security-related risks in meetings and take appropriate action in a timely manner.
[1633] Furthermore, an emotion engine that estimates the user's emotion may be further combined. That is, the identification processing unit 290 may estimate the user's emotion using the emotion identification model 59, and perform identification processing using the user's emotion.
[1634] The present invention relates to a system that starts by acquiring voice data, converts the voice data into text data, analyzes the text data to extract the content of the conversation and the emotional state, and calculates and visualizes the risk level. Furthermore, the present invention uses an emotion engine to recognize the user's emotions, and based on this, it is possible to more precisely evaluate the risk level. Below, each component of this system and its operation are specifically described.
[1635] System Embodiments
[1636] Acquiring audio data
[1637] The user enters the meeting details into the system and instructs the system to start the meeting. When the device detects the start of the meeting, it starts capturing audio data in real time through the microphone. The captured audio data is sent to the server in a specified format.
[1638] Saving audio data
[1639] The server receives the voice data sent from the device and stores it in a database, which is used in the subsequent analysis process.
[1640] Speech-to-Text (STT)
[1641] The server inputs the stored voice data into a Speech-to-Text (STT) engine and converts the voice into text data. For example, a speech saying, "The project is going well, but I've been feeling tired lately" is converted into text. This text data is temporarily stored.
[1642] Conversation content analysis
[1643] The server inputs the text data into a natural language processing (NLP) module and analyzes the content of the conversation. During this process, keywords related to the progress of work and the content of the consultation (e.g., "progress," "problem," "fatigue") are extracted. The extracted keywords and content are stored in a database and used for subsequent processing and report creation.
[1644] sentiment analysis
[1645] The server inputs the text data into an emotion engine, which analyzes the emotions and tone of the conversation. During this process, emotional states such as "happiness," "sadness," "anger," and "fatigue" are identified. The emotional state is compared with past data, and the difference from the usual state is calculated. The emotion engine also uses both voice and text data to identify the emotional state in detail. Furthermore, the emotion engine analyzes the user's facial expression data to identify the emotional state. This multi-layered emotion analysis enables more accurate risk calculations.
[1646] Risk Assessment
[1647] The server calculates the mental and physical risk level based on the results of the emotion analysis. It uses a big database to compare it with past data and calculate the current risk level. The results are visualized, and the impact and trends are displayed in graphs and dashboards.
[1648] Feedback and Notifications
[1649] The server generates reports based on the results of sentiment analysis and risk assessment. If a risk signal is detected, a notification is sent to the employee and their manager. This notification can be sent via email or a chat app, for example. Periodic reports are also generated and distributed to managers.
[1650] Reviewing feedback
[1651] Users can check the feedback and reports sent from the server, which can then be used to take measures regarding employees' work status and health status, such as suggesting that fatigued employees take time off.
[1652] Specific examples
[1653] scenario:
[1654] During a progress meeting for a project, one of the members said, "I've been feeling tired lately." Furthermore, it was confirmed that the member had a tired expression throughout the meeting.
[1655] How it works:
[1656] 1. The device collects the audio data of the meeting in real time and sends it to the server.
[1657] 2. The server receives the voice data and converts it into text data.
[1658] 3. The server analyzes the text data "The project is going well, but I've been feeling tired lately" and extracts the keyword "fatigue."
[1659] 4. The server uses the emotion engine to identify the "feeling of fatigue" and then analyzes the user's facial expression data to confirm that they are expressing fatigue.
[1660] 5. The server calculates the "mental fatigue level" and determines that the risk is high.
[1661] 6. The server sends a "high fatigue risk" notification to the manager and employee.
[1662] 7. The user reviews the notification and takes action, such as following up with the employee or suggesting time off.
[1663] In this way, the system of the present invention can perform multi-layered emotion analysis and effectively realize employee health management and productivity improvement throughout the organization.
[1664] The processing flow will be explained below.
[1665] Step 1:
[1666] A user logs in to the system and enters the date and time of the meeting, the participants, the purpose, etc. For example, the following information is registered in the system: "October 5th, 10:00, progress meeting, participants: Mr. A, Mr. B."
[1667] Step 2:
[1668] When the device detects the start of a meeting, it starts collecting audio data in real time through the microphone, and the collected audio data is sent to the server in a specified format.
[1669] Step 3:
[1670] The server receives the voice data sent from the device and stores it in a database, which is used in subsequent analysis processes.
[1671] Step 4:
[1672] The server inputs the stored voice data into a Speech-to-Text (STT) engine and converts the voice into text data. For example, a speech saying, "The project is going well, but I've been feeling tired lately" is converted into text. This converted text data is temporarily stored.
[1673] Step 5:
[1674] The server inputs the text data into a natural language processing (NLP) module and analyzes the content of the conversation. During this process, keywords related to the progress of work and the content of the consultation (e.g., "progress," "problems," "fatigue," etc.) are extracted. The extracted keywords and content are stored in a database.
[1675] Step 6:
[1676] The server inputs the text data into the emotion engine, which analyzes the emotions and tone of the conversation. During this process, emotional states such as "happiness," "sadness," "anger," and "fatigue" are identified. The emotion engine uses both the voice and text data to identify the emotional state in detail. The emotion engine also analyzes the user's facial expression data to complement the emotional state.
[1677] Step 7:
[1678] The server calculates the mental and physical risk level based on the results of the emotion analysis. It uses a big database to compare it with past data and calculate the current risk level. The results are visualized, and the impact and trends are displayed in graphs and dashboard format.
[1679] Step 8:
[1680] The server generates reports based on the results of sentiment analysis and risk assessment. If a risk signal is detected, a notification is sent to the employee and their manager. This notification can be sent via email or a chat app, for example. Periodic reports are also generated and distributed to managers.
[1681] Step 9:
[1682] Users can check the feedback and reports sent from the server, which can then be used to take measures regarding employees' work status and health status, such as suggesting that fatigued employees take time off.
[1683] Specific examples
[1684] scenario:
[1685] During a progress meeting for a project, a project member says, "I've been feeling tired lately." Furthermore, that member has a tired expression throughout the meeting.
[1686] How it works:
[1687] 1. The device collects the audio data of the meeting in real time and sends it to the server.
[1688] 2. The server receives the voice data and converts it into text data.
[1689] 3. The server analyzes the text data "The project is going well, but I've been feeling tired lately" and extracts the keyword "fatigue."
[1690] 4. The server uses the emotion engine to identify the "feeling of fatigue" and then analyzes the user's facial expression data to confirm the tiredness.
[1691] 5. The server calculates the "mental fatigue level" and determines that the risk is high.
[1692] 6. The server sends a "high fatigue risk" notification to the manager and employee.
[1693] 7. The user reviews the notification and takes action, such as following up with the employee or suggesting time off.
[1694] In this way, the system of the present invention can effectively manage employee health and improve productivity throughout the organization by performing multi-layered emotion analysis.
[1695] Example 2
[1696] Next, a description will be given of Example 2. In the following description, the data processing device 12 will be referred to as a "server" and the robot 414 will be referred to as a "terminal."
[1697] There is a problem in that it is difficult to grasp the work status and health status of employees in real time and respond quickly and accurately. In particular, it is difficult with conventional systems to precisely evaluate emotional states and risk levels and provide appropriate feedback and notifications. This poses a challenge in terms of improving work efficiency and effectively managing employee health.
[1698] The specific processing by the specific processing unit 290 of the data processing device 12 in the second embodiment is realized by the following means.
[1699] In this invention, the server includes means for acquiring voice data, means for converting voice data into text data, means for identifying an emotional state from the text data and voice data, means for calculating and visualizing a risk level by comparing it with past data, and means for sending a notification when the risk level exceeds a predetermined threshold. This makes it possible to precisely evaluate an employee's emotions and risk level and provide feedback and notifications at the appropriate time.
[1700] "Voice data" refers to digital data of voice or sound collected by a voice input device such as a microphone.
[1701] "Text data" is digital data that has been converted from voice data into character information.
[1702] "Emotional state" refers to the psychological and emotional state analyzed from the user's statements, voice data, facial expression data, etc.
[1703] "Risk level" is the degree of mental and physical risk calculated in comparison with emotional state and past data.
[1704] "Visualization" refers to visually displaying data such as risk level and emotional state in the form of graphs, dashboards, etc.
[1705] "Notification" refers to the sending of alerts or information to users or their superiors by the system based on specific conditions.
[1706] A "meeting" refers to a conference or meeting that takes place at a specific date and time, the details of which are entered into the system.
[1707] An "NLP (natural language processing) module" is a program module that analyzes text data and extracts conversation content and keywords.
[1708] An "emotion engine" is a software engine that analyzes voice data, text data, and facial expression data to identify emotional states.
[1709] A "big database" is a large-scale database for storing and analyzing large amounts of past data.
[1710] The system starts by acquiring voice data, converting it into text data, analyzing it to extract the content of the conversation and the emotional state, and then calculating and visualizing the risk level. It also uses an emotion engine to recognize the user's emotions and precisely evaluates the risk level based on that.
[1711] System Embodiments
[1712] Acquiring audio data
[1713] The user enters meeting details (date and time, participant list, purpose, etc.) through the system's user interface and presses the Start Meeting button to issue instructions. When the device detects the start of the meeting, it begins capturing audio data in real time via the microphone. The captured audio data is sent to the server in an appropriate format (e.g., WAV, MP3).
[1714] Saving audio data
[1715] The server receives the voice data sent from the terminal and stores it in a database, which is used in the subsequent analysis process.
[1716] Speech-to-text (STT)
[1717] The server inputs the stored voice data into a Speech-to-Text (STT) engine and converts the voice into text data. For example, a voice saying, "The project is going well, but I've been feeling tired lately" is converted into text data. This text data is temporarily stored.
[1718] Conversation content analysis
[1719] The server inputs the text data into a natural language processing (NLP) module and analyzes the content of the conversation. During this process, keywords related to the progress of work and the content of the consultation (e.g., "progress," "problem," "fatigue") are extracted. The extracted keywords and content are stored in a database and used for subsequent processing and report creation.
[1720] sentiment analysis
[1721] The server inputs text and voice data into an emotion engine to analyze the emotions and tone of the conversation. During this process, emotional states such as "happiness," "sadness," "anger," and "fatigue" are identified. The emotional state is compared with past data, and the difference from the usual state is calculated. Furthermore, in addition to voice and text data, facial expression data is used to identify the emotional state in more detail. This multi-layered emotion analysis enables more accurate risk calculations.
[1722] Risk Assessment
[1723] The server calculates the mental and physical risk level based on the results of the emotion analysis. It uses a big database to compare it with past data and calculate the current risk level. The results are visualized, and the impact and trends are displayed in graphs and dashboards.
[1724] Generate feedback and notifications
[1725] The server generates reports based on the results of sentiment analysis and risk assessment. If a risk signal is detected, a notification is sent to the employee and their manager. This notification can be sent via email or a chat app, for example. Periodic reports are also generated and distributed to managers.
[1726] Review feedback and implement measures
[1727] Users can check the feedback and reports sent from the server, which can then be used to take measures regarding employees' work status and health status, such as suggesting that fatigued employees take time off.
[1728] Specific examples
[1729] scenario:
[1730] During a progress meeting for a project, one of the members said, "I've been feeling tired lately." Furthermore, it was confirmed that the member had a tired expression throughout the meeting.
[1731] How it works:
[1732] 1. The device collects the audio data of the meeting in real time and sends it to the server.
[1733] 2. The server receives the voice data and converts it into text data.
[1734] 3. The server analyzes the text data "The project is going well, but I've been feeling tired lately" and extracts the keyword "fatigue."
[1735] 4. The server uses the emotion engine to identify the "feeling of fatigue" and then analyzes the user's facial expression data to confirm that they are expressing fatigue.
[1736] 5. The server calculates the "mental fatigue level" and determines that the risk is high.
[1737] 6. The server sends a "high fatigue risk" notification to the manager and employee.
[1738] 7. The user reviews the notification and takes action, such as following up with the employee or suggesting time off.
[1739] Prompt Sentence Examples
[1740] "Please explain how a system can collect and analyze everyone's comments during a status meeting, and then analyze their emotional state. Based on specific keywords and emotional tone, it can generate risk notifications."
[1741] In this way, the system of the present invention can perform multi-layered emotion analysis, effectively managing employee health and improving productivity throughout the organization.
[1742] The flow of the identification process in the second embodiment will be described with reference to FIG.
[1743] Step 1:
[1744] The user inputs detailed information about the meeting through the system's user interface and presses the start button to give instructions.
[1745] Input: Meeting details (date and time, participant list, purpose, etc.)
[1746] Output: Trigger to start a meeting
[1747] Specific operation: The user clicks the "Start Meeting" button using the system's web application, and the entered information is sent to the terminal.
[1748] Step 2:
[1749] When the terminal detects the start of a meeting, it starts capturing audio data in real time via the microphone.
[1750] Input: Meeting start trigger
[1751] Output: Real-time audio data
[1752] Specific behavior: The device's microphone will be turned on, a notification will appear saying "Audio collection during the meeting will begin," and audio data will begin to be collected.
[1753] Step 3:
[1754] The device sends the acquired audio data to the server in a specified format (e.g., WAV, MP3).
[1755] Input: Real-time audio data
[1756] Output: Audio data sent to the server in a given format
[1757] Specific operation: The device compresses the collected audio data and sends it to the server in the specified format.
[1758] Step 4:
[1759] The server receives the voice data sent from the terminal and stores it in a predetermined database.
[1760] Input: Audio data sent in a given format
[1761] Output: Audio data stored in a database
[1762] Specific operation: The server logs "Audio data received" and saves the file in the specified database.
[1763] Step 5:
[1764] The server inputs the stored voice data into a Speech-to-Text (STT) engine and converts the voice into text data.
[1765] Input: Audio data stored in a database
[1766] Output: Converted text data
[1767] Specific operation: The STT engine will start, and after a few seconds, the message "Speech to text conversion complete" will be displayed.
[1768] Step 6:
[1769] The server inputs the text data into a natural language processing (NLP) module to analyze the content of the conversation.
[1770] Input: Converted text data
[1771] Output: Extracted keywords and conversation content
[1772] Specific operation: The NLP module will start analysis and the message "Keyword extraction completed" will be displayed.
[1773] Step 7:
[1774] The server inputs the text and voice data into an emotion engine to analyze the emotion and tone of the conversation.
[1775] Input: Converted text data, audio data
[1776] Output: Identified emotional state
[1777] Specific behavior: The sentiment engine will be activated and the results will be displayed on the dashboard as "Sentiment analysis completed."
[1778] Step 8:
[1779] The server calculates the mental and physical risk levels based on the results of the emotion analysis.
[1780] Input: Identified emotional state
[1781] Output: Calculated risk level
[1782] Specific behavior: The risk assessment module will perform the analysis, the message "Risk assessment completed" will be displayed, and the results will be displayed in graphical form on the dashboard.
[1783] Step 9:
[1784] The server sends a notification if the risk exceeds a predetermined threshold.
[1785] Input: Calculated risk level
[1786] Output: Notification sent
[1787] Specific behavior: If the "high risk of fatigue" is determined, a notification will be created and the message "Notification sent" will be displayed.
[1788] Step 10:
[1789] Users check the feedback and reports sent from the server and take action.
[1790] Input: Notifications sent, reports
[1791] Output: Measures taken
[1792] Specific actions: The user receives a notification, checks the details on the dashboard, and clicks the "Suggest time off to employee" button to take action.
[1793] (Application example 2)
[1794] Next, a description will be given of Application Example 2. In the following description, the data processing device 12 will be referred to as a "server" and the robot 414 will be referred to as a "terminal."
[1795] Modern security services require rapid and accurate detection of potential risks posed by employees and visitors. However, previous technologies often only analyzed voice data, resulting in incomplete understanding of emotional states and making it difficult to accurately assess risk levels. Furthermore, the lack of an effective system for simultaneously performing real-time emotion analysis and risk assessment makes it difficult to respond quickly on-site. Therefore, the present invention aims to provide a system that uses multi-layered emotion analysis to integrate both voice data and facial expression data, enabling more accurate and rapid risk detection and assessment.
[1796] The specific processing by the specific processing unit 290 of the data processing device 12 in the application example 2 is realized by the following means.
[1797] In this invention, the server includes means for acquiring voice data, means for converting voice data into text data, means for analyzing the text data and extracting the content of the conversation, means for identifying an emotional state from the text data, means for calculating and visualizing a risk level based on the identified emotional state, means for sending a notification when the risk level exceeds a predetermined threshold, means for identifying an emotional state using both voice data and facial expression data, and a device (such as smart glasses) for analyzing the emotional state in real time. This enables integrated analysis of voice data and facial expression data, enabling more accurate risk assessment and rapid risk detection in real time.
[1798] "Audio data" is a digital representation of sound captured through an audio input device such as a microphone.
[1799] "Text data" is digital information that expresses voice data as characters.
[1800] The "content of the conversation" is information that indicates the topic and context extracted from the acquired text data.
[1801] "Emotional state" refers to the speaker's psychological state, which is determined by analyzing voice data, text data, and facial expression data.
[1802] "Risk level" is an assessment of potential danger calculated based on emotional state and conversation content.
[1803] "Visualization" is the process of displaying data in a graphical format (e.g., graphs, dashboards, etc.).
[1804] "Means for sending notifications" refers to a function that sends an alert via email, chat app, etc. when the risk level exceeds a specified threshold.
[1805] A "server" is a computer system that processes and stores voice and text data and analyzes them as needed.
[1806] "Apparatus for real-time analysis" refers to a device (e.g., smart glasses) that can acquire voice data and facial expression data in real time and perform analysis and processing immediately.
[1807] "Smart glasses" are eyeglass-type wearable devices that have built-in microphones and cameras and can acquire and analyze voice data and facial expression data.
[1808] overview
[1809] The present invention relates to a system that uses voice data and facial expression data to analyze emotional states and assess risk levels, particularly for use in security and surveillance applications, and is designed to enable security guards wearing smart glasses to more effectively detect risks.
[1810] System Configuration
[1811] The server operates to perform the following main functions:
[1812] 1. Acquiring audio data: Using the microphone built into the smart glasses, surrounding audio is collected in real time.
[1813] 2. Storage of voice data: The collected voice data is sent to a server and stored in a database.
[1814] 3. Speech-to-Text (STT): The stored voice data is converted into text data using a Speech-to-Text (STT) engine.
[1815] 4. Conversation content analysis: Text data is input into a natural language processing (NLP) module to extract the conversation content and important keywords.
[1816] 5. Sentiment analysis: Text data and facial expression data are input into the emotion engine to analyze the emotional state.
[1817] 6. Risk assessment: Based on the results of sentiment analysis, the risk level is calculated and visualized.
[1818] 7. Feedback and Notification: Send notifications to users and supervisors when risks are detected that exceed predetermined thresholds.
[1819] Specific examples
[1820] A security guard at a department store is patrolling while wearing smart glasses. The guard collects a visitor's statement, "I'm frustrated because there have been a lot of strange things happening at the shopping mall recently." Based on this statement, the system extracts keywords such as "strange" and "frustrated," and also detects anger from the visitor's facial expression. Combining this information, the system assesses the person as being at high risk and displays a notification on the guard's HUD indicating that the person is at high risk.
[1821] Hardware and software used
[1822] Hardware: Smart glasses with built-in microphone and camera
[1823] Software: Speech Recognition Library (speech-to-text conversion), NLP Module (natural language processing), Emotion Engine (sentiment analysis), Risk Assessment Module (risk assessment)
[1824] Prompt Sentence Examples
[1825] We want to design a system that analyzes the emotional state of a user from their voice data and facial expression data, and evaluates the risk level in real time. This system has the following functions:
[1826] 1. Acquiring audio data and converting it to text
[1827] 2. Content analysis of text data
[1828] 3. Emotional state analysis (based on voice and facial expression data)
[1829] 4. Calculating and visualizing risk levels
[1830] Generate pseudocode for a program that combines these elements.
[1831] The flow of the specific processing in the application example 2 will be described with reference to FIG.
[1832] Step 1:
[1833] Acquiring audio data
[1834] The server captures voice data in real time through the microphone of the smart glasses. This data is sent to the server as voice input. The input is the voice data uttered by the user, and the output is the voice file sent to the server.
[1835] Step 2:
[1836] Saving audio data
[1837] The server stores the captured audio data in a database, which is used in the subsequent analysis process. The input is the audio file sent to the server, and the output is the audio data stored in the database.
[1838] Step 3:
[1839] Speech-to-text (STT)
[1840] The server inputs the stored voice data into a Speech-to-Text (STT) engine and converts the voice into text data. The input is the voice data stored in the database, and the output is text data. This allows the voice data to be treated as text information.
[1841] Step 4:
[1842] Conversation content analysis
[1843] The server inputs the text data into a natural language processing (NLP) module and analyzes the content of the conversation. During this process, important keywords (e.g., "progress," "problems," and "fatigue") are extracted. The input is text data, and the output is a list of extracted keywords.
[1844] Step 5:
[1845] sentiment analysis
[1846] The server inputs text data and facial expression data captured by the smart glasses' built-in camera into the emotion engine to analyze the emotional state. During this process, emotions such as "happiness," "sadness," "anger," and "fatigue" are identified. The input is text data and facial expression data, and the output is the identified emotional state.
[1847] Step 6:
[1848] Risk Assessment
[1849] The server calculates the risk level based on the emotional state and the results of the conversation content analysis. It uses a big database to compare it with past data and evaluate the current risk level. The input is the emotional state and a list of keywords, and the output is the calculated risk level.
[1850] Step 7:
[1851] Feedback and Notifications
[1852] The server sends a notification to the user and supervisor if the calculated risk level exceeds a predetermined threshold. This notification is sent via email or a chat app. The input is the risk level and the output is the notification message.
[1853] Step 8:
[1854] Display and confirmation
[1855] The user checks the results of the risk level and emotional state on the head-up display (HUD) of the smart glasses. The input is a notification message, and the output is the information displayed on the HUD. This allows the user to understand the situation in real time and respond quickly.
[1856] The specific processing unit 290 transmits the result of the specific processing to the robot 414. In the robot 414, the control unit 46A causes the speaker 240 and the control target 443 to output the result of the specific processing. The microphone 238 acquires voice indicating a user input regarding the result of the specific processing. The control unit 46A transmits voice data indicating the user input acquired by the microphone 238 to the data processing device 12. In the data processing device 12, the specific processing unit 290 acquires the voice data.
[1857] The data generation model 58 is a so-called generative AI (Artificial Intelligence). An example of the data generation model 58 is ChatGPT (Internet Search<URL: https: / / openai.com / blog / chatgpt> ), Gemini (Internet search <url: https: gemini.google.com ?hl="ja">) and other generation AIs. The data generation model 58 is obtained by performing deep learning on a neural network. A prompt including an instruction is input to the data generation model 58, and inference data such as voice data indicating voice, text data indicating text, and image data indicating an image is also input. The data generation model 58 performs inference on the input inference data in accordance with the instruction indicated by the prompt, and outputs the inference result in a data format such as voice data and text data. Here, inference refers to, for example, analysis, classification, prediction, and / or summarization.
[1858] In the above embodiment, an example was given in which the specific processing is performed by the data processing device 12, but the technology of the present disclosure is not limited to this, and the specific processing may be performed by the robot 414.
[1859] The emotion identification model 59 as an emotion engine may determine the user's emotion according to a specific mapping. Specifically, the emotion identification model 59 may determine the user's emotion according to an emotion map (see FIG. 9), which is a specific mapping. Similarly, the emotion identification model 59 may determine the robot's emotion, and the identification processing unit 290 may perform identification processing using the robot's emotion.
[1860] FIG. 9 is a diagram illustrating an emotion map 400 on which multiple emotions are mapped. In the emotion map 400, emotions are arranged in concentric circles radiating from the center. Emotions closer to the center of the concentric circles are more primitive. Emotions representing states and actions arising from a state of mind are arranged on the outer edges of the concentric circles. The concept of emotion includes both affect and mental states. Emotions generally generated from reactions occurring in the brain are arranged on the left side of the concentric circles. Emotions generally induced by situational judgment are arranged on the right side of the concentric circles. Emotions generally generated from reactions occurring in the brain and induced by situational judgment are arranged on the upper and lower sides of the concentric circles. Furthermore, the emotion of "pleasure" is arranged on the upper side of the concentric circles, and the emotion of "discomfort" is arranged on the lower side. In this way, in the emotion map 400, multiple emotions are mapped based on the structure by which emotions are generated, and emotions that tend to occur simultaneously are mapped close to each other.
[1861] These emotions are distributed in the 3 o'clock direction on emotion map 400, and typically fluctuate between relief and anxiety. In the right half of emotion map 400, situational awareness dominates over internal sensations, resulting in a sense of calm.
[1862] The inside of emotion map 400 represents what is going on in the mind, and the outside of emotion map 400 represents behavior, so the further you go outside emotion map 400, the more visible the emotions become (the more they are expressed in behavior).
[1863] Human emotions are based on various balances, such as posture and blood sugar levels. When these balances deviate from the ideal, a state of discomfort is indicated, and when they approach the ideal, a state of pleasure is indicated. Emotions can also be created for robots, automobiles, and motorcycles, based on various balances, such as posture and remaining battery life. When these balances deviate from the ideal, a state of discomfort is indicated, and when they approach the ideal, a state of pleasure is indicated. An emotion map can be generated, for example, based on Dr. Mitsuyoshi's emotion map (Research on Voice Emotion Recognition and Emotional Brain Physiological Signal Analysis Systems, Tokushima University, Doctoral Dissertation: https: / / ci.nii.ac.jp / naid / 500000375379). The left half of the emotion map lists emotions belonging to the "reaction" domain, where sensation is dominant. The right half of the emotion map lists emotions belonging to the "situation" domain, where situational awareness is dominant.
[1864] The emotion map defines two emotions that promote learning. One is a negative emotion on the situation side, around the middle of "repentance" or "reflection." In other words, this occurs when the robot experiences negative emotions such as "I never want to feel this way again" or "I don't want to be scolded again." The other is a positive emotion on the response side, around "desire." In other words, this occurs when the robot experiences positive feelings such as "I want more" or "I want to know more."
[1865] The emotion identification model 59 inputs user input into a pre-trained neural network, obtains emotion values indicating each emotion shown in the emotion map 400, and determines the user's emotion. This neural network is pre-trained based on multiple pieces of training data that are combinations of user input and emotion values indicating each emotion shown in the emotion map 400. Furthermore, this neural network is trained so that emotions that are located close to each other have similar values, as in the emotion map 900 shown in FIG. 10. FIG. 10 shows an example in which multiple emotions, "relieved," "calm," and "reassuring," have similar emotion values.
[1866] The system according to the present disclosure has been described above mainly with respect to the functions of the data processing device 12, but the system according to the present disclosure is not necessarily implemented on a server. The system according to the present disclosure may be implemented as a general information processing system. The present disclosure may be implemented, for example, as a software program running on a personal computer or an application running on a smartphone, etc. The method according to the present disclosure may be provided to users in the form of SaaS (Software as a Service).
[1867] In the above embodiment, an example was given in which the specific processing is performed by one computer 22, but the technology of the present disclosure is not limited to this, and the specific processing may be distributed and performed by a plurality of computers including the computer 22. For example, the data generation model 58 may be provided in an external device of the data processing device 12, and data may be generated in the external device in accordance with input data.
[1868] In the above embodiment, an example in which the specific processing program 56 is stored in the storage 32 has been described, but the technology of the present disclosure is not limited to this. For example, the specific processing program 56 may be stored in a portable, computer-readable, non-transitory storage medium such as a USB (Universal Serial Bus) memory. The specific processing program 56 stored in the non-transitory storage medium is installed in the computer 22 of the data processing device 12. The processor 28 executes the specific processing in accordance with the specific processing program 56.
[1869] Alternatively, the specific processing program 56 may be stored in a storage device such as a server connected to the data processing device 12 via the network 54, and the specific processing program 56 may be downloaded and installed on the computer 22 in response to a request from the data processing device 12.
[1870] It is not necessary to store all of the specific processing program 56 in a storage device such as a server connected to the data processing device 12 via the network 54, or to store all of the specific processing program 56 in the storage 32; only a portion of the specific processing program 56 may be stored.
[1871] The hardware resource for executing a specific process can be any of the following processors: An example of a processor is a CPU, which is a general-purpose processor that functions as a hardware resource for executing a specific process by executing software, i.e., a program. Another example of a processor is a dedicated electrical circuit, such as an FPGA (Field-Programmable Gate Array), a PLD (Programmable Logic Device), or an ASIC (Application Specific Integrated Circuit), which is a processor with a circuit configuration designed specifically for executing a specific process. Each processor has built-in or connected memory, and each processor uses the memory to execute the specific process.
[1872] The hardware resource that executes the specific processing may be configured with one of these various processors, or may be configured with a combination of two or more processors of the same or different types (for example, a combination of multiple FPGAs, or a combination of a CPU and an FPGA). Also, the hardware resource that executes the specific processing may be a single processor.
[1873] As an example of a system configured with a single processor, first, one processor is configured by combining one or more CPUs and software, and this processor functions as a hardware resource that executes a specific process. Second, there is a system that uses a processor that realizes the functions of an entire system including multiple hardware resources that execute a specific process on a single IC chip, as typified by SoC (System-on-a-chip). In this way, a specific process is realized using one or more of the above-mentioned various processors as hardware resources.
[1874] Furthermore, the hardware structure of these various processors can be, more specifically, an electric circuit that combines circuit elements such as semiconductor devices. The specific processing described above is merely an example. Therefore, it goes without saying that unnecessary steps may be deleted, new steps may be added, or the processing order may be rearranged, without departing from the spirit of the invention.
[1875] The above-described description and illustrations are a detailed explanation of the parts related to the technology of the present disclosure and are merely an example of the technology of the present disclosure. For example, the above description of the configuration, functions, actions, and effects is an explanation of an example of the configuration, functions, actions, and effects of the parts related to the technology of the present disclosure. Therefore, it goes without saying that unnecessary parts may be deleted, new elements may be added, or replacements may be made to the above-described description and illustrations within the scope of the gist of the technology of the present disclosure. Furthermore, to avoid confusion and facilitate understanding of the parts related to the technology of the present disclosure, the above-described description and illustrations omit explanations of common technical knowledge that do not require particular explanation to enable the implementation of the technology of the present disclosure.
[1876] All publications, patent applications, and technical standards mentioned in this specification are herein incorporated by reference to the same extent as if each individual publication, patent application, or technical standard was specifically and individually indicated to be incorporated by reference.
[1877] The following is further disclosed regarding the above embodiment.
[1878] (Claim 1)
[1879] means for acquiring audio data;
[1880] means for converting the voice data into text data;
[1881] means for analyzing the text data and extracting the contents of the conversation;
[1882] means for identifying an emotional state from the text data;
[1883] A means for calculating and visualizing a risk level based on the identified emotional state;
[1884] The system includes means for sending a notification when the risk exceeds a predetermined threshold.
[1885] (Claim 2)
[1886] 10. The system of claim 1, wherein the system detects the start of a meeting and starts capturing audio data.
[1887] (Claim 3)
[1888] 2. The system according to claim 1, wherein the conversation content extraction means extracts keywords related to the progress of work and the content of consultation.
[1889] "Example 1"
[1890] (Claim 1)
[1891] means for acquiring audio data;
[1892] means for converting the voice data into text data;
[1893] means for analyzing the text data and extracting the contents of the conversation;
[1894] means for identifying an emotional state from the text data;
[1895] A means for calculating and visualizing a risk level based on the identified emotional state;
[1896] means for sending a notification when the risk level exceeds a predetermined threshold;
[1897] A means for the terminal to detect the start of a meeting and transmit audio data to a server in real time;
[1898] The server stores the voice data in a database and converts the voice data into text data using a speech-to-text engine;
[1899] A server uses a natural language processing module to extract keywords related to the progress of work and the content of consultation from the text data;
[1900] means for the server to analyze the emotional state using an emotion analysis module;
[1901] a means for the server to calculate the risk level using a big data analysis framework and visualize the risk level using a data visualization tool;
[1902] means for the server to generate a report and send notifications to the employee and his / her supervisor using a predetermined notification means;
[1903] A system that includes a means for users to check notifications from the server and take appropriate action.
[1904] (Claim 2)
[1905] 2. The system according to claim 1, wherein the terminal detects the start of a meeting and starts capturing audio data.
[1906] (Claim 3)
[1907] 2. The system according to claim 1, wherein the server analyzes the content of the conversation and extracts keywords relating to the progress of work and the content of the consultation.
[1908] "Application Example 1"
[1909] (Claim 1)
[1910] means for acquiring audio data;
[1911] means for converting the voice data into text data;
[1912] means for analyzing the text data and extracting the contents of the conversation;...
Claims
1. means for acquiring audio data; means for converting the voice data into text data; means for analyzing the text data and extracting the contents of the conversation; means for identifying an emotional state from the text data; A means for calculating and visualizing a risk level based on the identified emotional state; The system includes means for sending a notification when the risk exceeds a predetermined threshold.
2. 2. The system according to claim 1, wherein the system detects the start of a meeting and starts capturing audio data.
3. 2. The system according to claim 1, wherein said conversation content extraction means extracts keywords relating to the progress of work and the content of consultation.
Citation Information
Patent Citations
Persona chatbot control method and system
JP2022180282A