system
A system analyzing voice, facial expressions, and gestures provides real-time harassment risk evaluation and feedback to promote safe workplace communication.
Patent Information
- Authority / Receiving Office
- JP · JP
- Patent Type
- Applications
- Current Assignee / Owner
- SOFTBANK GROUP CORP
- Filing Date
- 2024-10-11
- Publication Date
- 2026-04-23
AI Technical Summary
The challenge of objectively judging harassment in a workplace environment is difficult due to subjective boundaries and the lack of real-time emotional state assessment, leading to potential inadvertent harassment during communication.
A system that simultaneously analyzes voice, facial expressions, and gestures to evaluate harassment risk by preprocessing data, analyzing emotional states, and providing real-time feedback to users.
The system effectively reduces the risk of harassment by clarifying emotional boundaries and promoting safe communication through timely adjustments based on emotional analysis.
Smart Images

Figure 2026069061000001_ABST
Abstract
Description
Technical Field
[0001] The technology of the present disclosure relates to a system.
Background Art
[0002] Patent Document 1 discloses a persona chatbot control method performed by at least one processor, the method including steps of receiving a user utterance, adding the user utterance to a prompt including an instruction sentence related to an explanation of a character of the chatbot, encoding the prompt, and inputting the encoded prompt into a language model to generate a chatbot utterance as a response to the user utterance.
Prior Art Documents
Patent Documents
[0003]
Patent Document 1
Summary of the Invention
Problems to be Solved by the Invention
[0004] Since the boundary of harassment depends on the victim's subjectivity, it is difficult to objectively judge. Further, in a workplace environment, there is a lack of means to immediately grasp the alternating emotional changes, and there is a risk of inadvertently causing harassment during communication. To solve this problem, it is required to analyze emotions in real time and explicitly evaluate the harassment risk.
Means for Solving the Problems
[0005] To solve the aforementioned problems, the present invention includes acquisition means for acquiring voice data, facial expression data, and gesture data, and preprocessing means for preprocessing and standardizing this data. Furthermore, it includes analysis means for analyzing the preprocessed data and evaluating emotional states, and evaluation means for evaluating harassment risk based on the analysis results. The evaluation results are provided to the user in real time by feedback means. This makes it possible to prevent the risk of harassment during communication and realize a safe workplace environment.
[0006] "Audio data" refers to data that records and stores sound waveforms in digital format, and includes human voice information.
[0007] "Facial expression data" refers to digital information obtained by capturing or recognizing a person's facial movements and expressions, and is used to analyze emotions and attitudes.
[0008] "Gesture data" refers to digital information obtained by capturing human body movements such as gestures and hand movements and converting them into an analyzable format.
[0009] "Acquisition means" refers to a mechanism that includes sensors and devices for collecting data such as voice, facial expressions, and gestures.
[0010] "Preprocessing means" refers to functions that perform processing such as filtering and standardization in order to prepare acquired data into an appropriate format for analysis.
[0011] An "analysis method" is a mechanism that uses a specific algorithm to evaluate emotions and states based on pre-processed data.
[0012] An "evaluation tool" is a function that uses information obtained through analysis to determine harassment risk according to specific criteria.
[0013] A "feedback mechanism" is an interface or device that notifies the user of evaluation results and prompts them to adjust their behavior or receive warnings. [Brief explanation of the drawing]
[0014] [Figure 1] This is a conceptual diagram showing an example of the configuration of a data processing system according to the first embodiment. [Figure 2] This is a conceptual diagram showing an example of the essential functions of a data processing device and a smart device according to the first embodiment. [Figure 3] This is a conceptual diagram showing an example of the configuration of a data processing system according to the second embodiment. [Figure 4] This is a conceptual diagram showing an example of the main functions of a data processing device and smart glasses according to the second embodiment. [Figure 5] This is a conceptual diagram showing an example of the configuration of a data processing system according to the third embodiment. [Figure 6] This is a conceptual diagram showing an example of the main functions of a data processing device and a headset-type terminal according to the third embodiment. [Figure 7] This is a conceptual diagram showing an example of the configuration of a data processing system according to the fourth embodiment. [Figure 8] This is a conceptual diagram showing an example of the main functions of a data processing device and a robot according to the fourth embodiment. [Figure 9] This shows an emotion map where multiple emotions are mapped. [Figure 10] This shows an emotion map where multiple emotions are mapped. [Figure 11] This is a sequence diagram showing the processing flow of the data processing system in Example 1. [Figure 12] This is a sequence diagram showing the processing flow of the data processing system in Application Example 1. [Figure 13] This is a sequence diagram showing the processing flow of the data processing system in Example 2, which incorporates an emotion engine. [Figure 14]It is a sequence diagram showing the processing flow of a data processing system in Application Example 2 when combined with an emotion engine.
Embodiments for Carrying Out the Invention
[0015] Hereinafter, an example of an embodiment of a system according to the technology of the present disclosure will be described with reference to the accompanying drawings.
[0016] First, the terms used in the following description will be explained.
[0017] In the following embodiments, a processor with a reference numeral (hereinafter simply referred to as "processor") may be a single arithmetic unit or a combination of multiple arithmetic units. Also, the processor may be a single type of arithmetic unit or a combination of multiple types of arithmetic units. Examples of arithmetic units include a CPU (Central Processing Unit), a GPU (Graphics Processing Unit), a GPGPU (General-Purpose computing on Graphics Processing Units), an APU (Accelerated Processing Unit), and the like.
[0018] In the following embodiments, a RAM (Random Access Memory) with a reference numeral is a memory in which information is temporarily stored and is used as a work memory by the processor.
[0019] In the following embodiments, a storage with a reference numeral is one or more non-volatile storage devices that store various programs and various parameters, etc. Examples of non-volatile storage devices include flash memory (SSD (Solid State Drive)), magnetic disks (e.g., hard disks), or magnetic tapes, and the like.
[0020] In the following embodiments, the signed communication interface (I / F) is an interface that includes a communication processor and an antenna, etc. The communication interface manages communication between multiple computers. Examples of communication standards applicable to the communication interface include wireless communication standards such as 5G (5th Generation Mobile Communication System), Wi-Fi (registered trademark), or Bluetooth (registered trademark).
[0021] In the following embodiments, "A and / or B" is synonymous with "at least one of A and B." That is, "A and / or B" means that it may be A alone, or B alone, or a combination of A and B. Furthermore, in this specification, the same concept as "A and / or B" applies when expressing three or more things linked by "and / or."
[0022] [First Embodiment]
[0023] Figure 1 shows an example of the configuration of the data processing system 10 according to the first embodiment.
[0024] As shown in Figure 1, the data processing system 10 includes a data processing device 12 and a smart device 14. An example of the data processing device 12 is a server.
[0025] The data processing device 12 comprises a computer 22, a database 24, and a communication interface 26. The computer 22 is an example of a "computer" related to the technology of this disclosure. The computer 22 comprises a processor 28, RAM 30, and storage 32. The processor 28, RAM 30, and storage 32 are connected to a bus 34. The database 24 and the communication interface 26 are also connected to the bus 34. The communication interface 26 is connected to a network 54. An example of the network 54 is a WAN (Wide Area Network) and / or a LAN (Local Area Network).
[0026] The smart device 14 comprises a computer 36, a reception device 38, an output device 40, a camera 42, and a communication interface 44. The computer 36 comprises a processor 46, RAM 48, and storage 50. The processor 46, RAM 48, and storage 50 are connected to a bus 52. The reception device 38, output device 40, and camera 42 are also connected to the bus 52.
[0027] The reception device 38 is equipped with a touch panel 38A and a microphone 38B, etc., and receives user input. The touch panel 38A receives user input by detecting contact with an object (e.g., a pen or finger). The microphone 38B receives user input by detecting the user's voice. The control unit 46A transmits data indicating the user input received by the touch panel 38A and microphone 38B to the data processing device 12. In the data processing device 12, the specific processing unit 290 acquires the data indicating the user input.
[0028] The output device 40 includes a display 40A and a speaker 40B, and presents data to the user 20 by outputting the data in a form perceptible to the user 20 (e.g., audio and / or text). The display 40A displays visible information such as text and images according to instructions from the processor 46. The speaker 40B outputs audio according to instructions from the processor 46. The camera 42 is a small digital camera equipped with an optical system such as a lens, aperture, and shutter, and an image sensor such as a CMOS (Complementary Metal-Oxide-Semiconductor) image sensor or a CCD (Charge Coupled Device) image sensor.
[0029] Communication interface 44 is connected to network 54. Communication interfaces 44 and 26 are responsible for the exchange of various types of information between processor 46 and processor 28 via network 54.
[0030] Figure 2 shows an example of the main functions of the data processing device 12 and the smart device 14.
[0031] As shown in Figure 2, in the data processing device 12, a specific processing is performed by the processor 28. A specific processing program 56 is stored in the storage 32. The specific processing program 56 is an example of a "program" related to the technology of this disclosure. The processor 28 reads the specific processing program 56 from the storage 32 and executes the read specific processing program 56 on the RAM 30. The specific processing is realized by the processor 28 operating as a specific processing unit 290 according to the specific processing program 56 executed on the RAM 30.
[0032] The storage 32 stores the data generation model 58 and the emotion identification model 59. The data generation model 58 and the emotion identification model 59 are used by the identification processing unit 290.
[0033] In the smart device 14, the processor 46 performs the reception output processing. The storage 50 stores the reception output program 60. The reception output program 60 is used in conjunction with a specific processing program 56 by the data processing system 10. The processor 46 reads the reception output program 60 from the storage 50 and executes the read reception output program 60 on the RAM 48. The reception output processing is realized by the processor 46 operating as a control unit 46A according to the reception output program 60 executed on the RAM 48.
[0034] Next, the specific processing performed by the specific processing unit 290 of the data processing device 12 will be described. In the following description, the data processing device 12 will be referred to as the "server" and the smart device 14 as the "terminal".
[0035] This invention provides a system that simultaneously analyzes voice, facial expressions, and gestures to reduce the risk of harassment in the workplace. This system includes three main components: a terminal, a server, and a user, each playing a specific role.
[0036] The device is responsible for collecting data from the user. Specifically, it uses a microphone and camera installed in the device to acquire the user's voice and video in real time. The acquired voice data is subjected to noise reduction within the device, and facial expressions and gestures are extracted from the video data using image recognition.
[0037] Preprocessed data is sent to a server via the network. The server analyzes the received data and uses various algorithms to analyze the user's emotional state. This analysis assesses the risk of harassment and generates a risk score. The evaluation system determines the risk based on specific thresholds and generates warnings as needed.
[0038] The generated feedback is then sent back to the device via the network. Upon receiving the feedback, the device immediately provides the user with a visual or auditory warning. Based on this warning, the user can adjust their statements and behavior in a timely manner.
[0039] As a concrete example, when a user speaks during a meeting, the device records their voice and facial expressions, and the server detects changes in emotion. Through analysis, if the tone of voice becomes higher or the facial expression becomes more serious, the server assesses a high harassment risk and provides the user with a warning through the device.
[0040] In this way, this system clarifies the boundaries of harassment in real time and provides an effective means to promote safe and healthy communication.
[0041] The following describes the processing flow.
[0042] Step 1:
[0043] The device automatically activates the microphone and camera when the user starts a conversation, and collects audio and video data in real time.
[0044] Step 2:
[0045] The device applies a noise reduction filter to the collected audio data to extract clear audio, and uses an image recognition algorithm to extract facial expressions and gestures from the video data.
[0046] Step 3:
[0047] The terminal sends pre-processed voice, facial expression, and gesture data to the server as secure data packets.
[0048] Step 4:
[0049] The server decodes the received data packets, uses a speech analysis algorithm to infer emotions from the tone and speed of the voice, and analyzes facial expression and gesture data using a deep learning model.
[0050] Step 5:
[0051] Based on the analysis results, the server performs an overall sentiment assessment and scores the harassment risk based on specific criteria.
[0052] Step 6:
[0053] If the server determines that the assessed risk score exceeds a threshold, it generates a warning message in real time and sends that feedback to the terminal.
[0054] Step 7:
[0055] The device receives feedback from the server and provides the user with warning messages visually or audibly. The user then adjusts their speech and gestures based on this feedback.
[0056] (Example 1)
[0057] Next, we will describe Example 1. In the following description, the data processing device 12 will be referred to as the "server," and the smart device 14 will be referred to as the "terminal."
[0058] Preventing the deterioration of interpersonal relationships and the occurrence of harassment in the workplace is extremely important. However, traditional methods have challenges in effectively managing risks, such as difficulty in understanding emotional states in real time and providing appropriate feedback.
[0059] The identification process performed by the identification processing unit 290 of the data processing device 12 in Example 1 is realized by the following means.
[0060] In this invention, the server includes acquisition means for simultaneously acquiring audio and video information, preprocessing means for preprocessing the acquired information using noise reduction and image understanding, and analysis means for analyzing the preprocessed information and using a generative artificial intelligence model to evaluate emotional states. This makes it possible to grasp the user's emotional state in real time, assess the risks in interpersonal relationships, and provide appropriate feedback.
[0061] "Voice information" refers to sound wave data that is electronically recorded by detecting the user's speech and conversation.
[0062] "Video information" refers to image data that visually records the user's posture, movements, and facial expressions.
[0063] "Acquisition means" refers to a mechanism that simultaneously detects audio and video information and collects it as digital data.
[0064] "Preprocessing means" refers to processing means that remove noise from acquired audio information and refine video information through image analysis.
[0065] "Noise reduction" is a technology that removes unwanted noise and other sounds from audio information, thereby improving signal clarity.
[0066] "Image understanding" is a technology that recognizes facial expressions and gestures from video information and extracts them as data.
[0067] A "generative artificial intelligence model" is a trained algorithm used to identify a user's emotional state based on acquired and pre-processed information.
[0068] "Analysis means" refers to a technical mechanism introduced to analyze pre-processed information and evaluate the user's emotional state.
[0069] "Feedback" refers to notifications and advice provided to users based on results obtained from analysis.
[0070] "Interpersonal risk" is an indicator of the possibility that dialogue in the workplace environment may cause discomfort or harm to the other party.
[0071] This invention is a system that simultaneously acquires and analyzes audio and video information and provides feedback to the user in order to reduce the risk of harassment in the workplace environment.
[0072] The device uses a high-performance microphone to acquire audio information and features noise cancellation to suppress ambient noise. It also has a high-resolution camera to collect video information, meticulously recording the user's face and body movements. This two types of information are transmitted to a server via the network.
[0073] The server applies noise reduction processing to audio information and uses image recognition technology (such as OpenCV) on video information to extract the user's facial expressions and gestures. The pre-processed data is then analyzed by a generative AI model to identify the user's emotional state. Based on the analysis results, the server calculates the risk of harassment and generates feedback as needed.
[0074] Upon receiving the results of their harassment risk assessment, users receive visual or auditory feedback through their device. This allows users to appropriately adjust their words and actions on the spot, thereby maintaining a safe and healthy work environment.
[0075] As a concrete example, when a user speaks during a meeting, the device records their voice and facial expressions, which are then analyzed by a server. If changes in emotion are detected, such as a rise in voice tone or a hostile expression, the server assesses the high risk of harassment and provides the user with a warning through the device.
[0076] An example of a prompt message is, "Assess the risk that the statements made during this meeting constitute harassment." This system supports good communication in the workplace by understanding emotional states in real time and conducting appropriate risk assessments.
[0077] The flow of the specific processing in Example 1 will be explained using Figure 11.
[0078] Step 1:
[0079] The device simultaneously acquires user audio and video information using a high-performance microphone and camera. Specifically, the microphone records audio as digital data, and the camera captures the user's face and body movements in real time. The input for this step is the user's real-time conversation and movements, and the output is audio and video data converted into digital format.
[0080] Step 2:
[0081] The device performs noise reduction processing on the acquired audio data. Specifically, it removes unwanted components such as ambient noise to output a clear audio signal. For video data, it uses image understanding technology to extract facial expressions and gestures. The input is the audio and video data acquired in step 1, and the output is the noise-reduced audio data and image data formatted for analysis.
[0082] Step 3:
[0083] The terminal transmits pre-processed audio and video data to the server over the network. Specifically, the data is securely transferred using a secure protocol (e.g., SSL / TLS). The input to this step is the output data from step 2, and the output is the data ready for analysis that has been passed to the server.
[0084] Step 4:
[0085] The server uses a generative AI model to analyze the received data. Specifically, it analyzes voice tone and emotion from audio data, and facial expressions and movements from video data. The input is the data sent in step 3, and the output is the analysis result representing the user's emotional state.
[0086] Step 5:
[0087] The server assesses the user's harassment risk based on the analysis results. Specifically, it compares the calculated emotional state with a specific threshold to calculate a risk score. The input for this step is the analysis results from step 4, and the output is a score indicating the harassment risk.
[0088] Step 6:
[0089] The server generates feedback for the user based on the risk assessment. Specifically, if the risk is high, it creates a warning message and prepares the feedback content. The input is the risk score from step 5, and the output is the generated feedback message.
[0090] Step 7:
[0091] The terminal receives feedback sent from the server and notifies the user visually or audibly. Specifically, it displays a pop-up on the screen or provides an audible alert. The input for this step is the feedback message from step 6, and the output is a warning or alert notification to the user.
[0092] (Application Example 1)
[0093] Next, we will explain Application Example 1. In the following explanation, the data processing device 12 will be referred to as the "server," and the smart device 14 will be referred to as the "terminal."
[0094] There is a need to provide a system that reduces the risk of harassment in the work environment and facilitates communication between workers and autonomous machines. This invention aims to solve the problem of providing a means to detect stress and communication problems during work in real time and to respond quickly.
[0095] The specific processing performed by the specific processing unit 290 of the data processing device 12 in Application Example 1 is realized by the following means.
[0096] In this invention, the server includes a collection means for acquiring voice data, facial expression data, and motion data; a preprocessing means for preprocessing and generalizing the acquired data; and an analysis means for analyzing the preprocessed data and determining the emotional state. This makes it possible to reduce the risk of harassment in the work environment and provide appropriate feedback to workers.
[0097] "Audio data" refers to information recorded in digital format for analysis, such as the subject's speech or other sound information.
[0098] "Facial expression data" refers to information used to acquire and analyze the facial movements and expressions that indicate specific emotions of a subject, as images or videos.
[0099] "Motion data" refers to information used to digitally record and analyze gestures, including the movements of the subject's arms and body.
[0100] "Collection means" refers to a device or mechanism for acquiring voice data, facial expression data, and motion data from a subject and transmitting them to a system.
[0101] "Preprocessing means" refers to the process of processing acquired data into a format that is easy to analyze, reducing noise, and generalizing the data.
[0102] "Analysis means" refers to methods and devices for determining the emotional state of a subject and evaluating the risk using pre-processed data.
[0103] "Evaluation methods" refer to the process of determining the risk of harassment based on the analysis results and generating feedback as needed.
[0104] "Notification means" refers to a device or method for communicating information, including risks, based on evaluation results, to the target person or relevant party.
[0105] "Transmission means" refers to the process of using autonomous machines to communicate evaluation results and feedback to workers or stakeholders.
[0106] This invention is a system designed to reduce the risk of harassment in the work environment and to facilitate communication between workers and autonomous machines. This system collects and analyzes voice data, facial expression data, and motion data in real time.
[0107] The terminal is equipped with a microphone and camera to capture worker voice and video. These devices are used to acquire audio and video from the work environment in real time. The acquired data undergoes noise reduction and image recognition within the terminal. Specific hardware includes high-performance cameras and directional microphones. Software libraries such as TENSORFLOW® and OpenCV are used for data preprocessing.
[0108] The server is responsible for analyzing the pre-processed data. Built on a cloud platform, the server utilizes Google® Cloud Speech-to-Text and Azure® Cognitive Services for speech recognition and sentiment analysis. The server determines the emotional state of workers and assesses harassment risk based on the analyzed audio and video data.
[0109] The evaluation results are transmitted to workers and stakeholders via notification systems. Specifically, appropriate feedback is provided through displays and speakers installed in the autonomous machines. For example, if a worker exhibits stressful behavior while working with an autonomous machine, the system will offer advice such as, "Try to calm your tone."
[0110] By utilizing generative AI models and prompt statements, feedback can become flexible and effective, tailored to the specific situation on site. An example of a prompt statement would be, "We want to design an algorithm to analyze employee stress levels in real time using audio and video data from factory work." A system configured in this way can effectively reduce the risk of harassment and enhance the safety of the work environment.
[0111] The flow of a specific process in Application Example 1 will be explained using Figure 12.
[0112] Step 1:
[0113] The terminal collects the voice and video of workers. It receives raw audio and video data acquired by a microphone and camera as input. Based on this, it performs preprocessing such as noise reduction and image recognition to remove noise from the audio and outputs standardized data in which facial expressions and movements are extracted. Specifically, the terminal runs an audio filtering algorithm and a facial recognition model in parallel.
[0114] Step 2:
[0115] The pre-processed data is sent from the terminal to the server. The server receives this as input and analyzes the voice and facial expression data. For data processing, Google Cloud Speech-to-Text is used to convert the voice to text, and emotion analysis is performed using Azure Cognitive Services. As a result, evaluation data indicating the worker's emotional state is output. Specifically, the server sequentially converts the voice to text and detects changes in emotion from the continuous data stream.
[0116] Step 3:
[0117] The server determines harassment risk based on evaluation data. It receives evaluation data indicating emotional state as input and calculates a risk score based on an algorithm. This process determines whether the risk exceeds a certain threshold. The output determines the content of the notification to the user. Specifically, the server compares the calculated risk score against multiple criteria.
[0118] Step 4:
[0119] Based on the evaluation results, the terminal sends feedback to the worker. It receives notifications from the server as input and provides feedback visually and audibly. The software used includes a speaker and display, outputting visual warning screens and audio notifications. Specifically, the terminal flexibly generates voice messages and displays related information on the screen.
[0120] Furthermore, an emotion engine that estimates the user's emotions may be incorporated. That is, the identification processing unit 290 may use the emotion identification model 59 to estimate the user's emotions and perform identification processing using the user's emotions.
[0121] This invention relates to a system that simultaneously acquires voice data, facial expression data, and gesture data, and analyzes them to evaluate the user's emotional state. By incorporating an emotion engine, this system can recognize the user's emotions with greater accuracy and assess harassment risk. The emotion engine has advanced functionality to perform integrated analysis of voice and facial expression data and identify the intensity and type of emotion.
[0122] The device continuously collects necessary data using its microphone and camera while the user performs daily tasks and conversations. The collected data undergoes basic noise reduction and data format conversion within the device to improve data quality. The pre-processed data is then transferred to a server via the network.
[0123] The server analyzes the received voice, facial expression, and gesture data. The analysis incorporates an emotion engine that comprehensively analyzes changes in voice tone, tempo, and facial expression. This allows for a more detailed assessment of emotional states, enabling a highly accurate determination of how the user is feeling. This emotional information is crucial in harassment risk assessment, and a risk score is generated accordingly.
[0124] Based on the results obtained through analysis and evaluation, the server generates warnings and advice as needed. For example, if the emotion engine identifies user stress, the server immediately sends that information to the terminal, providing the user with excellent feedback.
[0125] This feedback is provided to the user visually or audibly through their device. Based on the information provided, the user can adjust their own behavior and actions. As a result, the risk of harassment is reduced, and a more comfortable and secure work environment can be created. For example, during a team meeting, if the emotional engine detects tension or anxiety, the user is given a warning encouraging a softer approach, improving the quality of communication.
[0126] In this way, this system utilizes an emotion engine to provide an effective solution aimed at improving communication in the workplace and preventing harassment.
[0127] The following describes the processing flow.
[0128] Step 1:
[0129] The device activates its microphone and camera as soon as the user starts a conversation, collecting audio and video data.
[0130] Step 2:
[0131] The device removes noise from collected audio data and detects facial feature points from video data to preprocess initial data of facial expressions and gestures.
[0132] Step 3:
[0133] The terminal packets the processed voice, facial expression, and gesture data and sends it to the server via a secure protocol.
[0134] Step 4:
[0135] The server decodes the data received from the terminal and uses an emotion engine to accurately analyze the intensity and type of emotion from voice tone, facial expressions, and gestures.
[0136] Step 5:
[0137] The server uses the analyzed sentiment information to score the harassment risk and generates a warning when the result exceeds a certain threshold.
[0138] Step 6:
[0139] The server transmits the generated warnings and advice to the terminal, which then provides them to the user in real time, either visually or audibly.
[0140] Step 7:
[0141] Users can review the feedback they receive and adjust their attitudes and statements as needed to maintain better communication.
[0142] (Example 2)
[0143] Next, we will describe Example 2. In the following description, the data processing device 12 will be referred to as the "server" and the smart device 14 as the "terminal".
[0144] In today's workplace environment, improving the quality of communication and reducing the risk of harassment is crucial. However, traditional methods make it difficult to accurately assess an employee's emotional state and provide appropriate feedback based on that assessment. To address this challenge, it is necessary to efficiently acquire and analyze acoustic, visual, and motion signals.
[0145] The identification process performed by the identification processing unit 290 of the data processing device 12 in Example 2 is realized by the following means.
[0146] In this invention, the server includes means for acquiring acoustic signals, visual signals, and motion signals; means for preprocessing the acquired signals and converting them into a unified format; and means for analyzing the preprocessed signals and identifying emotional states. This makes it possible to evaluate the user's emotional state with high accuracy and provide rapid feedback if there is a risk.
[0147] An "acoustic signal" is a representation of sound or audio data as an electrical signal.
[0148] "Visual signals" are representations of visual information as electrical signals, and include images and videos.
[0149] "Motion signals" are electrical signals that capture human movements and gestures.
[0150] "Means of acquisition" refers to the devices and technologies used to collect data or signals.
[0151] "Preprocessing means" refers to devices or techniques that perform initial data processing to make the data easier to analyze.
[0152] "Converting to a unified format" means unifying data obtained in different formats or standards into a single standard format that is easier to analyze.
[0153] "Means of analysis" refers to devices or technologies for analyzing acquired data in detail and extracting specific information.
[0154] "Identifying emotional states" means determining the type and intensity of a user's emotions based on the results of data analysis.
[0155] "Risk assessment" is the process of evaluating the extent to which potential dangers or problems exist based on the analyzed information.
[0156] "Feedback" refers to the act of providing advice and information to users based on acquired data and analysis results.
[0157] This invention provides a system that uses acoustic signals, visual signals, and motion signals to analyze a user's emotional state with high accuracy and reduce the risk of harassment.
[0158] The device acquires acoustic and visual signals using a microphone and camera. Acoustic signals include the user's voice, while visual signals include the user's facial expressions and gestures. The acquired signals are preprocessed within the device to improve data quality through noise reduction and formatting standardization. This preprocessing prepares the data into a consistent format for analysis.
[0159] The server receives pre-processed signals, which are then analyzed by an emotion engine. This process involves analyzing the speaker's tone and tempo from the acoustic signals, and changes in facial expressions and movements from the visual signals. The emotion engine employs sophisticated algorithms to integrate these individual elements to identify the type and intensity of emotion. The specific hardware consists of standard server equipment, while the software incorporates signal processing and machine learning models.
[0160] The analysis results are used for risk assessment to evaluate the likelihood of harassment. If the risk exceeds the assessed threshold, the server generates a warning on the terminal and provides appropriate feedback to the user. This feedback is presented visually or audibly from the terminal, allowing the user to improve their behavior based on it.
[0161] As a concrete example, during a team meeting, if the emotion engine detects tension or anxiety in the user, the server provides feedback to encourage relaxation. This can facilitate smoother communication during the meeting and reduce stress.
[0162] Examples of prompt messages include the following:
[0163] "Propose a system that analyzes a user's emotions based on acoustic, visual, and motion signals, generates feedback when stress or anxiety is detected, and demonstrate how it can be used to improve workplace communication."
[0164] In this way, the present invention provides a practical solution that analyzes the emotional state of users in detail, improves the quality of communication in the workplace, and reduces the risk of harassment.
[0165] The flow of the specific processing in Example 2 will be explained using Figure 13.
[0166] Step 1:
[0167] The device acquires acoustic and visual signals using a microphone and camera. The user's voice and facial expressions are captured as input by the microphone and camera. Specifically, the microphone collects ambient sounds, and the camera continuously captures the user's face and its movements. These signals are sent to the device's data storage.
[0168] Step 2:
[0169] The terminal performs noise reduction and data format standardization on the acquired signals. The acoustic and visual signals obtained in step 1 are used as input. Specifically, noise filtering is performed to reduce background noise. After this, the audio and visual signals are converted to a common data format to ensure data consistency. The output is a pre-processed signal suitable for analysis.
[0170] Step 3:
[0171] The server receives pre-processed signals and performs analysis using an emotion engine. The input consists of standardized acoustic and visual signals transmitted from the terminal. Specifically, the emotion engine analyzes the tone and tempo of sounds, changes in facial expressions, and gestures, correlating these elements to identify the type and intensity of emotion. The output is the analyzed emotion information.
[0172] Step 4:
[0173] The server performs a risk assessment based on the analyzed emotional information. The emotional information obtained in step 3 is used as input. Specifically, it calculates a harassment risk score using the emotional information and compares this score against the evaluation criteria. The output is the risk score and the evaluation result.
[0174] Step 5:
[0175] The server generates feedback based on the evaluation results and sends it to the terminal. The risk score obtained in step 4 is used as input. Specifically, if the risk is high, it generates a warning message or advice and sends it to the terminal. The output is specific feedback to the user.
[0176] Step 6:
[0177] Users receive feedback through their device and adjust their behavior accordingly. Input includes visual or auditory feedback provided by the device. Specific actions include, for example, taking deep breaths or adjusting their work pace if they receive relaxing feedback. The output is improved behavior or psychological state.
[0178] (Application Example 2)
[0179] Next, we will explain application example 2. In the following explanation, the data processing device 12 will be referred to as a "server" and the smart device 14 as a "terminal".
[0180] While technologies exist that use voice, facial expressions, and movement information to evaluate human emotional states, many of them suffer from insufficient data acquisition and analysis, particularly in detecting negative emotions in communication and providing real-time feedback. Furthermore, risk assessment and countermeasures in response to emotional changes tend to lag behind, making it difficult to ensure healthy communication in workplaces and educational institutions.
[0181] The specific processing performed by the specific processing unit 290 of the data processing device 12 in Application Example 2 is realized by the following means.
[0182] In this invention, the server includes means for acquiring voice information, facial expression information, and motion information; means for preprocessing and standardizing the acquired information; and means for analyzing the preprocessed information and evaluating the emotional state. This makes it possible to provide real-time feedback and facilitate communication.
[0183] "Audio information" refers to data used to record sounds emitted by the user and analyze their tone and tempo.
[0184] "Facial expression information" refers to data used to capture changes in a user's face and analyze changes in their emotions.
[0185] "Motion information" refers to data used to capture and analyze the user's gestures and body movements.
[0186] "Means of acquisition" refers to devices and mechanisms for collecting voice information, facial expression information, and motion information in real time.
[0187] "Preprocessing methods" refer to processing techniques used to reduce noise and standardize collected data in order to improve its quality.
[0188] "Standardization" refers to the process of adjusting data into a unified format to make it easier to analyze.
[0189] "Means of analysis" refer to algorithms and software that comprehensively evaluate voice information, facial expression information, and movement information to identify emotional states.
[0190] "Means of evaluation" refers to a system that determines potential risks based on the analysis results.
[0191] "Means of provision" refers to an interface for providing feedback to users based on evaluation results and encouraging improvement.
[0192] A "means of providing real-time feedback" is a system that responds immediately to changes in emotions and provides information back to the user right away.
[0193] "Means of facilitating communication" refer to feedback mechanisms that facilitate interactions between users and reduce negative emotions.
[0194] The system for realizing this invention mainly consists of a server and a terminal. The terminal uses a microphone and camera to acquire voice information, facial expression information, and motion information in real time during the user's daily conversations and activities. The collected information undergoes initial noise reduction on the terminal, is converted into standardized data, and then transferred to the server.
[0195] The server analyzes the received data using an advanced emotion engine. This emotion engine incorporates data analysis libraries such as TensorFlow and OpenCV. The emotion engine comprehensively analyzes the tone and tempo of the voice, as well as changes in facial expressions and movements, to identify the intensity and type of the user's emotions. Subsequently, a harassment risk assessment is performed, enabling immediate, real-time feedback.
[0196] Feedback is provided to users visually or audibly through their devices. For example, if one participant shows signs of tension or anxiety during a team meeting, the server sends a notification to all participants such as, "Let's stay calm." In this way, it improves the quality of communication and contributes to promoting healthy dialogue and relationships in the workplace.
[0197] Furthermore, to realize this system, an example of a prompt using a generative AI model is: "Analyze the emotional state of the participants from the following call log, and generate appropriate feedback if negative emotions are strong." Based on the results obtained from this prompt, the system can construct more accurate feedback.
[0198] The flow of a specific process in Application Example 2 will be explained using Figure 14.
[0199] Step 1:
[0200] The device collects the user's voice information, facial expression information, and motion information. This data is acquired using a microphone and camera, respectively. Since the acquired data contains noise, noise reduction is performed. The input is raw voice information, facial expression information, and motion information, and the output is standardized data with the noise removed.
[0201] Step 2:
[0202] The terminal standardizes and formats the data, which has had noise reduced through preprocessing. This process converts the data into a format suitable for analysis on the server. The input is noise-reduced data, and the output is in a standardized data format.
[0203] Step 3:
[0204] The server receives standardized data sent from the terminal and analyzes it using an emotion engine. The emotion engine uses libraries such as TensorFlow and OpenCV to analyze voice tone, facial expression changes, and actions. The input is standardized data, and the output is the analysis result regarding the intensity and type of emotion.
[0205] Step 4:
[0206] The server assesses harassment risk based on the analysis results and generates feedback as needed. A generative AI model is used for the assessment, with a focus on detecting negative emotions. The input is the analysis results, and the output is a risk assessment score and feedback content.
[0207] Step 5:
[0208] The server sends the generated feedback to the terminal, notifying the user visually or audibly. The user can then adjust their communication style based on this. The input is the feedback content, and the output is the notification to the user and their response action.
[0209] Step 6:
[0210] Users can adjust their behavior based on feedback received from the server and take actions to improve the quality of their communication. This specifically includes softening their tone of voice or using more gentle facial expressions. The input is the feedback action, and the output is the improved communication style.
[0211] The specific processing unit 290 transmits the result of the specific processing to the smart device 14. In the smart device 14, the control unit 46A causes the output device 40 to output the result of the specific processing. The microphone 38B acquires audio indicating user input for the result of the specific processing. The control unit 46A transmits the audio data indicating user input acquired by the microphone 38B to the data processing device 12. In the data processing device 12, the specific processing unit 290 acquires the audio data.
[0212] Data generation model 58 is a so-called generative AI (Artificial Intelligence). An example of data generation model 58 is ChatGPT (registered trademark) (Internet search).<URL: https: / / openai.com / blog / chatgpt> ), Gemini (registered trademark) (Internet search) <url: https: gemini.google.com ?hl="ja">Examples of generative AI include the following. The data generation model 58 is obtained by performing deep learning on a neural network. The data generation model 58 is input with prompts containing instructions, and with inference data such as audio data representing speech, text data representing text, and image data representing images. The data generation model 58 infers from the input inference data according to the instructions indicated by the prompts, and outputs the inference results in data formats such as audio data and text data. Here, inference refers to, for example, analysis, classification, prediction, and / or summarization.
[0213] In the above embodiment, an example was given in which specific processing is performed by the data processing device 12, but the technology of this disclosure is not limited thereto, and the specific processing may also be performed by the smart device 14.
[0214] [Second Embodiment]
[0215] Figure 3 shows an example of the configuration of the data processing system 210 according to the second embodiment.
[0216] As shown in Figure 3, the data processing system 210 includes a data processing device 12 and smart glasses 214. An example of the data processing device 12 is a server.
[0217] The data processing device 12 comprises a computer 22, a database 24, and a communication interface 26. The computer 22 is an example of a "computer" related to the technology of this disclosure. The computer 22 comprises a processor 28, RAM 30, and storage 32. The processor 28, RAM 30, and storage 32 are connected to a bus 34. The database 24 and the communication interface 26 are also connected to the bus 34. The communication interface 26 is connected to a network 54. An example of the network 54 is a WAN (Wide Area Network) and / or a LAN (Local Area Network).
[0218] The smart glasses 214 include a computer 36, a microphone 238, a speaker 240, a camera 42, and a communication interface 44. The computer 36 includes a processor 46, RAM 48, and storage 50. The processor 46, RAM 48, and storage 50 are connected to a bus 52. The microphone 238, speaker 240, and camera 42 are also connected to the bus 52.
[0219] The microphone 238 receives voice signals from the user 20 and receives instructions from the user 20. The microphone 238 captures the voice signals from the user 20, converts the captured voice into audio data, and outputs it to the processor 46. The speaker 240 outputs audio according to the instructions from the processor 46.
[0220] Camera 42 is a small digital camera equipped with an optical system including a lens, aperture, and shutter, and an image sensor such as a CMOS (Complementary Metal-Oxide-Semiconductor) image sensor or a CCD (Charge Coupled Device) image sensor, and captures images of the area around the user 20 (for example, an imaging range defined by a field of view equivalent to the width of a typical healthy person's field of vision).
[0221] Communication interface 44 is connected to network 54. Communication interfaces 44 and 26 are responsible for the exchange of various information between processor 46 and processor 28 via network 54. The exchange of various information between processor 46 and processor 28 using communication interfaces 44 and 26 is performed in a secure manner.
[0222] Figure 4 shows an example of the main functions of the data processing device 12 and the smart glasses 214. As shown in Figure 4, the data processing device 12 performs specific processing using the processor 28. The storage 32 stores the specific processing program 56.
[0223] The specific processing program 56 is an example of a "program" relating to the technology of this disclosure. The processor 28 reads the specific processing program 56 from the storage 32 and executes the read specific processing program 56 on the RAM 30. The specific processing is realized by the processor 28 operating as a specific processing unit 290 in accordance with the specific processing program 56 executed on the RAM 30.
[0224] The storage 32 stores the data generation model 58 and the emotion identification model 59. The data generation model 58 and the emotion identification model 59 are used by the identification processing unit 290.
[0225] In the smart glasses 214, the processor 46 performs the reception output processing. The storage 50 stores the reception output program 60. The processor 46 reads the reception output program 60 from the storage 50 and executes the read reception output program 60 on the RAM 48. The reception output processing is realized by the processor 46 operating as a control unit 46A according to the reception output program 60 executed on the RAM 48.
[0226] Next, the identification processing performed by the identification processing unit 290 of the data processing device 12 will be described. In the following description, the data processing device 12 will be referred to as the "server" and the smart glasses 214 will be referred to as the "terminal".
[0227] This invention provides a system that simultaneously analyzes voice, facial expressions, and gestures to reduce the risk of harassment in the workplace. This system includes three main components: a terminal, a server, and a user, each playing a specific role.
[0228] The device is responsible for collecting data from the user. Specifically, it uses a microphone and camera installed in the device to acquire the user's voice and video in real time. The acquired voice data is subjected to noise reduction within the device, and facial expressions and gestures are extracted from the video data using image recognition.
[0229] Preprocessed data is sent to a server via the network. The server analyzes the received data and uses various algorithms to analyze the user's emotional state. This analysis assesses the risk of harassment and generates a risk score. The evaluation system determines the risk based on specific thresholds and generates warnings as needed.
[0230] The generated feedback is then sent back to the device via the network. Upon receiving the feedback, the device immediately provides the user with a visual or auditory warning. Based on this warning, the user can adjust their statements and behavior in a timely manner.
[0231] As a concrete example, when a user speaks during a meeting, the device records their voice and facial expressions, and the server detects changes in emotion. Through analysis, if the tone of voice becomes higher or the facial expression becomes more serious, the server assesses a high harassment risk and provides the user with a warning through the device.
[0232] In this way, this system clarifies the boundaries of harassment in real time and provides an effective means to promote safe and healthy communication.
[0233] The following describes the processing flow.
[0234] Step 1:
[0235] The device automatically activates the microphone and camera when the user starts a conversation, and collects audio and video data in real time.
[0236] Step 2:
[0237] The device applies a noise reduction filter to the collected audio data to extract clear audio, and uses an image recognition algorithm to extract facial expressions and gestures from the video data.
[0238] Step 3:
[0239] The terminal sends pre-processed voice, facial expression, and gesture data to the server as secure data packets.
[0240] Step 4:
[0241] The server decodes the received data packets, uses a speech analysis algorithm to infer emotions from the tone and speed of the voice, and analyzes facial expression and gesture data using a deep learning model.
[0242] Step 5:
[0243] Based on the analysis results, the server performs an overall sentiment assessment and scores the harassment risk based on specific criteria.
[0244] Step 6:
[0245] If the server determines that the assessed risk score exceeds a threshold, it generates a warning message in real time and sends that feedback to the terminal.
[0246] Step 7:
[0247] The device receives feedback from the server and provides the user with warning messages visually or audibly. The user then adjusts their speech and gestures based on this feedback.
[0248] (Example 1)
[0249] Next, we will describe Example 1. In the following description, the data processing device 12 will be referred to as the "server," and the smart glasses 214 will be referred to as the "terminal."
[0250] Preventing the deterioration of interpersonal relationships and the occurrence of harassment in the workplace is extremely important. However, traditional methods have challenges in effectively managing risks, such as difficulty in understanding emotional states in real time and providing appropriate feedback.
[0251] The identification process performed by the identification processing unit 290 of the data processing device 12 in Example 1 is realized by the following means.
[0252] In this invention, the server includes acquisition means for simultaneously acquiring audio and video information, preprocessing means for preprocessing the acquired information using noise reduction and image understanding, and analysis means for analyzing the preprocessed information and using a generative artificial intelligence model to evaluate emotional states. This makes it possible to grasp the user's emotional state in real time, assess the risks in interpersonal relationships, and provide appropriate feedback.
[0253] "Voice information" refers to sound wave data that is electronically recorded by detecting the user's speech and conversation.
[0254] "Video information" refers to image data that visually records the user's posture, movements, and facial expressions.
[0255] "Acquisition means" refers to a mechanism that simultaneously detects audio and video information and collects it as digital data.
[0256] "Preprocessing means" refers to processing means that remove noise from acquired audio information and refine video information through image analysis.
[0257] "Noise reduction" is a technology that removes unwanted noise and other sounds from audio information, thereby improving signal clarity.
[0258] "Image understanding" is a technology that recognizes facial expressions and gestures from video information and extracts them as data.
[0259] A "generative artificial intelligence model" is a trained algorithm used to identify a user's emotional state based on acquired and pre-processed information.
[0260] "Analysis means" refers to a technical mechanism introduced to analyze pre-processed information and evaluate the user's emotional state.
[0261] "Feedback" refers to notifications and advice provided to users based on results obtained from analysis.
[0262] "Interpersonal risk" is an indicator of the possibility that dialogue in the workplace environment may cause discomfort or harm to the other party.
[0263] This invention is a system that simultaneously acquires and analyzes audio and video information and provides feedback to the user in order to reduce the risk of harassment in the workplace environment.
[0264] The device uses a high-performance microphone to acquire audio information and features noise cancellation to suppress ambient noise. It also has a high-resolution camera to collect video information, meticulously recording the user's face and body movements. This two types of information are transmitted to a server via the network.
[0265] The server applies noise reduction processing to audio information and uses image recognition technology (such as OpenCV) on video information to extract the user's facial expressions and gestures. The pre-processed data is then analyzed by a generative AI model to identify the user's emotional state. Based on the analysis results, the server calculates the risk of harassment and generates feedback as needed.
[0266] Upon receiving the results of their harassment risk assessment, users receive visual or auditory feedback through their device. This allows users to appropriately adjust their words and actions on the spot, thereby maintaining a safe and healthy work environment.
[0267] As a concrete example, when a user speaks during a meeting, the device records their voice and facial expressions, which are then analyzed by a server. If changes in emotion are detected, such as a rise in voice tone or a hostile expression, the server assesses the high risk of harassment and provides the user with a warning through the device.
[0268] An example of a prompt message is, "Assess the risk that the statements made during this meeting constitute harassment." This system supports good communication in the workplace by understanding emotional states in real time and conducting appropriate risk assessments.
[0269] The flow of the specific processing in Example 1 will be explained using Figure 11.
[0270] Step 1:
[0271] The device simultaneously acquires user audio and video information using a high-performance microphone and camera. Specifically, the microphone records audio as digital data, and the camera captures the user's face and body movements in real time. The input for this step is the user's real-time conversation and movements, and the output is audio and video data converted into digital format.
[0272] Step 2:
[0273] The device performs noise reduction processing on the acquired audio data. Specifically, it removes unwanted components such as ambient noise to output a clear audio signal. For video data, it uses image understanding technology to extract facial expressions and gestures. The input is the audio and video data acquired in step 1, and the output is the noise-reduced audio data and image data formatted for analysis.
[0274] Step 3:
[0275] The terminal transmits pre-processed audio and video data to the server over the network. Specifically, the data is securely transferred using a secure protocol (e.g., SSL / TLS). The input to this step is the output data from step 2, and the output is the data ready for analysis that has been passed to the server.
[0276] Step 4:
[0277] The server uses a generative AI model to analyze the received data. Specifically, it analyzes voice tone and emotion from audio data, and facial expressions and movements from video data. The input is the data sent in step 3, and the output is the analysis result representing the user's emotional state.
[0278] Step 5:
[0279] The server assesses the user's harassment risk based on the analysis results. Specifically, it compares the calculated emotional state with a specific threshold to calculate a risk score. The input for this step is the analysis results from step 4, and the output is a score indicating the harassment risk.
[0280] Step 6:
[0281] The server generates feedback to the user based on risk assessment. As a specific operation, when the risk is high, it creates a warning message and prepares the feedback content. The input is the risk score in step 5, and the output is the generated feedback message.
[0282] Step 7:
[0283] The terminal receives the feedback sent from the server and notifies the user visually or aurally. Specifically, it displays a popup on the screen or prompts with voice. The input for this step is the feedback message in step 6, and the output is the notification of warning or alert to the user.
[0284] (Application Example 1)
[0285] Next, Application Example 1 will be described. In the following description, the data processing device 12 is referred to as the "server", and the smart glasses 214 are referred to as the "terminal".
[0286] There is a need to provide a system for reducing the risk of harassment in the work environment and for smoothing the communication between workers and autonomous machines. The present invention aims to solve the problem of providing means for detecting stress and communication problems during work in real time and responding promptly.
[0287] The specific processing by the specific processing unit 290 of the data processing device 12 in Application Example 1 is realized by the following respective means.
[0288] In this invention, the server includes a collection means for acquiring voice data, expression data, and motion data, a preprocessing means for preprocessing and generalizing the acquired data, and an analysis means for analyzing the preprocessed data and determining the emotional state. Thereby, it becomes possible to reduce the harassment risk in the work environment and provide appropriate feedback to the worker.
[0289] "Audio data" refers to information recorded in digital format for analysis, such as the subject's speech or other sound information.
[0290] "Facial expression data" refers to information used to acquire and analyze the facial movements and expressions that indicate specific emotions of a subject, as images or videos.
[0291] "Motion data" refers to information used to digitally record and analyze gestures, including the movements of the subject's arms and body.
[0292] "Collection means" refers to a device or mechanism for acquiring voice data, facial expression data, and motion data from a subject and transmitting them to a system.
[0293] "Preprocessing means" refers to the process of processing acquired data into a format that is easy to analyze, reducing noise, and generalizing the data.
[0294] "Analysis means" refers to methods and devices for determining the emotional state of a subject and evaluating the risk using pre-processed data.
[0295] "Evaluation methods" refer to the process of determining the risk of harassment based on the analysis results and generating feedback as needed.
[0296] "Notification means" refers to a device or method for communicating information, including risks, based on evaluation results, to the target person or relevant party.
[0297] "Transmission means" refers to the process of using autonomous machines to communicate evaluation results and feedback to workers or stakeholders.
[0298] This invention is a system designed to reduce the risk of harassment in the work environment and to facilitate communication between workers and autonomous machines. This system collects and analyzes voice data, facial expression data, and motion data in real time.
[0299] The terminal is equipped with a microphone and camera to capture the worker's voice and video. These devices are used to acquire audio and video from the work environment in real time. The acquired data undergoes noise reduction and image recognition within the terminal. Specific hardware components include high-performance cameras and directional microphones. Software libraries such as TensorFlow and OpenCV are used for data preprocessing.
[0300] The server is responsible for analyzing the pre-processed data. Built on a cloud platform, the server utilizes Google Cloud Speech-to-Text and Azure Cognitive Services for speech recognition and sentiment analysis. From the analyzed audio and video data, the server determines the emotional state of workers and assesses the risk of harassment.
[0301] The evaluation results are transmitted to workers and stakeholders via notification systems. Specifically, appropriate feedback is provided through displays and speakers installed in the autonomous machines. For example, if a worker exhibits stressful behavior while working with an autonomous machine, the system will offer advice such as, "Try to calm your tone."
[0302] By utilizing generative AI models and prompt statements, feedback can become flexible and effective, tailored to the specific situation on site. An example of a prompt statement would be, "We want to design an algorithm to analyze employee stress levels in real time using audio and video data from factory work." A system configured in this way can effectively reduce the risk of harassment and enhance the safety of the work environment.
[0303] The flow of a specific process in Application Example 1 will be explained using Figure 12.
[0304] Step 1:
[0305] The terminal collects the voices and videos of workers. As input, it receives the raw voice data and video data obtained by the microphone and camera. Based on this, it performs preprocessing such as noise reduction and image recognition, removes noise from the voice, and outputs standardized data extracted from facial expressions and movements. As a specific operation, the terminal executes a voice filtering algorithm and a face recognition model in parallel.
[0306] Step 2:
[0307] The preprocessed data is sent from the terminal to the server. The server receives this as input and analyzes the voice and expression data. As data processing, it uses Google Cloud Speech-to-Text to convert the voice into text and performs sentiment analysis with Azure Cognitive Services. As a result, it outputs evaluation data indicating the emotional state of the worker. As a specific operation, the server sequentially converts the voice into text and detects changes in emotion from the continuous data stream.
[0308] Step 3:
[0309] The server determines the harassment risk based on the evaluation data. As input, it receives the evaluation data indicating the emotional state and calculates a risk score based on an algorithm. Thereby, a process is performed to determine whether the risk exceeds a specific criterion. As output, the content of the notification to the user is determined. As a specific operation, the server compares the calculated risk score with multiple criteria.
[0310] Step 4:
[0311] Based on the evaluation result, the terminal sends feedback to the worker. As input, it receives the content of the notification from the server and provides feedback visually and auditorily. As the software used, speakers and displays are utilized to output a visual warning screen and voice notifications. As a specific operation, the terminal flexibly generates a voice message and displays relevant information on the display.
[0312] Furthermore, an emotion engine that estimates the user's emotions may be incorporated. That is, the identification processing unit 290 may use the emotion identification model 59 to estimate the user's emotions and perform identification processing using the user's emotions.
[0313] This invention relates to a system that simultaneously acquires voice data, facial expression data, and gesture data, and analyzes them to evaluate the user's emotional state. By incorporating an emotion engine, this system can recognize the user's emotions with greater accuracy and assess harassment risk. The emotion engine has advanced functionality to perform integrated analysis of voice and facial expression data and identify the intensity and type of emotion.
[0314] The device continuously collects necessary data using its microphone and camera while the user performs daily tasks and conversations. The collected data undergoes basic noise reduction and data format conversion within the device to improve data quality. The pre-processed data is then transferred to a server via the network.
[0315] The server analyzes the received voice, facial expression, and gesture data. The analysis incorporates an emotion engine that comprehensively analyzes changes in voice tone, tempo, and facial expression. This allows for a more detailed assessment of emotional states, enabling a highly accurate determination of how the user is feeling. This emotional information is crucial in harassment risk assessment, and a risk score is generated accordingly.
[0316] Based on the results obtained through analysis and evaluation, the server generates warnings and advice as needed. For example, if the emotion engine identifies user stress, the server immediately sends that information to the terminal, providing the user with excellent feedback.
[0317] This feedback is provided to the user visually or audibly through their device. Based on the information provided, the user can adjust their own behavior and actions. As a result, the risk of harassment is reduced, and a more comfortable and secure work environment can be created. For example, during a team meeting, if the emotional engine detects tension or anxiety, the user is given a warning encouraging a softer approach, improving the quality of communication.
[0318] In this way, this system utilizes an emotion engine to provide an effective solution aimed at improving communication in the workplace and preventing harassment.
[0319] The following describes the processing flow.
[0320] Step 1:
[0321] The device activates its microphone and camera as soon as the user starts a conversation, collecting audio and video data.
[0322] Step 2:
[0323] The device removes noise from collected audio data and detects facial feature points from video data to preprocess initial data of facial expressions and gestures.
[0324] Step 3:
[0325] The terminal packets the processed voice, facial expression, and gesture data and sends it to the server via a secure protocol.
[0326] Step 4:
[0327] The server decodes the data received from the terminal and uses an emotion engine to accurately analyze the intensity and type of emotion from voice tone, facial expressions, and gestures.
[0328] Step 5:
[0329] The server uses the analyzed sentiment information to score the harassment risk and generates a warning when the result exceeds a certain threshold.
[0330] Step 6:
[0331] The server transmits the generated warnings and advice to the terminal, which then provides them to the user in real time, either visually or audibly.
[0332] Step 7:
[0333] Users can review the feedback they receive and adjust their attitudes and statements as needed to maintain better communication.
[0334] (Example 2)
[0335] Next, we will describe Example 2. In the following description, the data processing device 12 will be referred to as the "server" and the smart glasses 214 will be referred to as the "terminal".
[0336] In today's workplace environment, improving the quality of communication and reducing the risk of harassment is crucial. However, traditional methods make it difficult to accurately assess an employee's emotional state and provide appropriate feedback based on that assessment. To address this challenge, it is necessary to efficiently acquire and analyze acoustic, visual, and motion signals.
[0337] The identification process performed by the identification processing unit 290 of the data processing device 12 in Example 2 is realized by the following means.
[0338] In this invention, the server includes means for acquiring acoustic signals, visual signals, and motion signals; means for preprocessing the acquired signals and converting them into a unified format; and means for analyzing the preprocessed signals and identifying emotional states. This makes it possible to evaluate the user's emotional state with high accuracy and provide rapid feedback if there is a risk.
[0339] An "acoustic signal" is a representation of sound or audio data as an electrical signal.
[0340] "Visual signals" are representations of visual information as electrical signals, and include images and videos.
[0341] "Motion signals" are electrical signals that capture human movements and gestures.
[0342] "Means of acquisition" refers to the devices and technologies used to collect data or signals.
[0343] "Preprocessing means" refers to devices or techniques that perform initial data processing to make the data easier to analyze.
[0344] "Converting to a unified format" means unifying data obtained in different formats or standards into a single standard format that is easier to analyze.
[0345] "Means of analysis" refers to devices or technologies for analyzing acquired data in detail and extracting specific information.
[0346] "Identifying emotional states" means determining the type and intensity of a user's emotions based on the results of data analysis.
[0347] "Risk assessment" is the process of evaluating the extent to which potential dangers or problems exist based on the analyzed information.
[0348] "Feedback" refers to the act of providing advice and information to users based on acquired data and analysis results.
[0349] This invention provides a system that uses acoustic signals, visual signals, and motion signals to analyze a user's emotional state with high accuracy and reduce the risk of harassment.
[0350] The device acquires acoustic and visual signals using a microphone and camera. Acoustic signals include the user's voice, while visual signals include the user's facial expressions and gestures. The acquired signals are preprocessed within the device to improve data quality through noise reduction and formatting standardization. This preprocessing prepares the data into a consistent format for analysis.
[0351] The server receives pre-processed signals, which are then analyzed by an emotion engine. This process involves analyzing the speaker's tone and tempo from the acoustic signals, and changes in facial expressions and movements from the visual signals. The emotion engine employs sophisticated algorithms to integrate these individual elements to identify the type and intensity of emotion. The specific hardware consists of standard server equipment, while the software incorporates signal processing and machine learning models.
[0352] The analysis results are used for risk assessment to evaluate the likelihood of harassment. If the risk exceeds the assessed threshold, the server generates a warning on the terminal and provides appropriate feedback to the user. This feedback is presented visually or audibly from the terminal, allowing the user to improve their behavior based on it.
[0353] As a concrete example, during a team meeting, if the emotion engine detects tension or anxiety in the user, the server provides feedback to encourage relaxation. This can facilitate smoother communication during the meeting and reduce stress.
[0354] Examples of prompt messages include the following:
[0355] "Propose a system that analyzes a user's emotions based on acoustic, visual, and motion signals, generates feedback when stress or anxiety is detected, and demonstrate how it can be used to improve workplace communication."
[0356] In this way, the present invention provides a practical solution that analyzes the emotional state of users in detail, improves the quality of communication in the workplace, and reduces the risk of harassment.
[0357] The flow of the specific processing in Example 2 will be explained using Figure 13.
[0358] Step 1:
[0359] The device acquires acoustic and visual signals using a microphone and camera. The user's voice and facial expressions are captured as input by the microphone and camera. Specifically, the microphone collects ambient sounds, and the camera continuously captures the user's face and its movements. These signals are sent to the device's data storage.
[0360] Step 2:
[0361] The terminal performs noise reduction and data format standardization on the acquired signals. The acoustic and visual signals obtained in step 1 are used as input. Specifically, noise filtering is performed to reduce background noise. After this, the audio and visual signals are converted to a common data format to ensure data consistency. The output is a pre-processed signal suitable for analysis.
[0362] Step 3:
[0363] The server receives pre-processed signals and performs analysis using an emotion engine. The input consists of standardized acoustic and visual signals transmitted from the terminal. Specifically, the emotion engine analyzes the tone and tempo of sounds, changes in facial expressions, and gestures, correlating these elements to identify the type and intensity of emotion. The output is the analyzed emotion information.
[0364] Step 4:
[0365] The server performs a risk assessment based on the analyzed emotional information. The emotional information obtained in step 3 is used as input. Specifically, it calculates a harassment risk score using the emotional information and compares this score against the evaluation criteria. The output is the risk score and the evaluation result.
[0366] Step 5:
[0367] The server generates feedback based on the evaluation results and sends it to the terminal. The risk score obtained in step 4 is used as input. Specifically, if the risk is high, it generates a warning message or advice and sends it to the terminal. The output is specific feedback to the user.
[0368] Step 6:
[0369] Users receive feedback through their device and adjust their behavior accordingly. Input includes visual or auditory feedback provided by the device. Specific actions include, for example, taking deep breaths or adjusting their work pace if they receive relaxing feedback. The output is improved behavior or psychological state.
[0370] (Application Example 2)
[0371] Next, we will explain application example 2. In the following explanation, the data processing device 12 will be referred to as the "server," and the smart glasses 214 will be referred to as the "terminal."
[0372] While technologies exist that use voice, facial expressions, and movement information to evaluate human emotional states, many of them suffer from insufficient data acquisition and analysis, particularly in detecting negative emotions in communication and providing real-time feedback. Furthermore, risk assessment and countermeasures in response to emotional changes tend to lag behind, making it difficult to ensure healthy communication in workplaces and educational institutions.
[0373] The specific processing performed by the specific processing unit 290 of the data processing device 12 in Application Example 2 is realized by the following means.
[0374] In this invention, the server includes means for acquiring voice information, facial expression information, and motion information; means for preprocessing and standardizing the acquired information; and means for analyzing the preprocessed information and evaluating the emotional state. This makes it possible to provide real-time feedback and facilitate communication.
[0375] "Audio information" refers to data used to record sounds emitted by the user and analyze their tone and tempo.
[0376] "Facial expression information" refers to data used to capture changes in a user's face and analyze changes in their emotions.
[0377] "Motion information" refers to data used to capture and analyze the user's gestures and body movements.
[0378] "Means of acquisition" refers to devices and mechanisms for collecting voice information, facial expression information, and motion information in real time.
[0379] "Preprocessing methods" refer to processing techniques used to reduce noise and standardize collected data in order to improve its quality.
[0380] "Standardization" refers to the process of adjusting data into a unified format to make it easier to analyze.
[0381] "Means of analysis" refer to algorithms and software that comprehensively evaluate voice information, facial expression information, and movement information to identify emotional states.
[0382] "Means of evaluation" refers to a system that determines potential risks based on the analysis results.
[0383] "Means of provision" refers to an interface for providing feedback to users based on evaluation results and encouraging improvement.
[0384] A "means of providing real-time feedback" is a system that responds immediately to changes in emotions and provides information back to the user right away.
[0385] "Means of facilitating communication" refer to feedback mechanisms that facilitate interactions between users and reduce negative emotions.
[0386] The system for realizing this invention mainly consists of a server and a terminal. The terminal uses a microphone and camera to acquire voice information, facial expression information, and motion information in real time during the user's daily conversations and activities. The collected information undergoes initial noise reduction on the terminal, is converted into standardized data, and then transferred to the server.
[0387] The server analyzes the received data using an advanced emotion engine. This emotion engine incorporates data analysis libraries such as TensorFlow and OpenCV. The emotion engine comprehensively analyzes the tone and tempo of the voice, as well as changes in facial expressions and movements, to identify the intensity and type of the user's emotions. Subsequently, a harassment risk assessment is performed, enabling immediate, real-time feedback.
[0388] Feedback is provided to users visually or audibly through their devices. For example, if one participant shows signs of tension or anxiety during a team meeting, the server sends a notification to all participants such as, "Let's stay calm." In this way, it improves the quality of communication and contributes to promoting healthy dialogue and relationships in the workplace.
[0389] Furthermore, to realize this system, an example of a prompt using a generative AI model is: "Analyze the emotional state of the participants from the following call log, and generate appropriate feedback if negative emotions are strong." Based on the results obtained from this prompt, the system can construct more accurate feedback.
[0390] The flow of a specific process in Application Example 2 will be explained using Figure 14.
[0391] Step 1:
[0392] The device collects the user's voice information, facial expression information, and motion information. This data is acquired using a microphone and camera, respectively. Since the acquired data contains noise, noise reduction is performed. The input is raw voice information, facial expression information, and motion information, and the output is standardized data with the noise removed.
[0393] Step 2:
[0394] The terminal standardizes and formats the data, which has had noise reduced through preprocessing. This process converts the data into a format suitable for analysis on the server. The input is noise-reduced data, and the output is in a standardized data format.
[0395] Step 3:
[0396] The server receives standardized data sent from the terminal and analyzes it using an emotion engine. The emotion engine uses libraries such as TensorFlow and OpenCV to analyze voice tone, facial expression changes, and actions. The input is standardized data, and the output is the analysis result regarding the intensity and type of emotion.
[0397] Step 4:
[0398] The server assesses harassment risk based on the analysis results and generates feedback as needed. A generative AI model is used for the assessment, with a focus on detecting negative emotions. The input is the analysis results, and the output is a risk assessment score and feedback content.
[0399] Step 5:
[0400] The server sends the generated feedback to the terminal, notifying the user visually or audibly. The user can then adjust their communication style based on this. The input is the feedback content, and the output is the notification to the user and their response action.
[0401] Step 6:
[0402] Users can adjust their behavior based on feedback received from the server and take actions to improve the quality of their communication. This specifically includes softening their tone of voice or using more gentle facial expressions. The input is the feedback action, and the output is the improved communication style.
[0403] The specific processing unit 290 transmits the result of the specific processing to the smart glasses 214. In the smart glasses 214, the control unit 46A causes the speaker 240 to output the result of the specific processing. The microphone 238 acquires audio indicating user input for the result of the specific processing. The control unit 46A transmits the audio data indicating user input acquired by the microphone 238 to the data processing unit 12. In the data processing unit 12, the specific processing unit 290 acquires the audio data.
[0404] Data generation model 58 is a type of so-called generative AI (Artificial Intelligence). One example of data generation model 58 is ChatGPT (Internet Search).<URL: https: / / openai.com / blog / chatgpt> ), Gemini (Internet search) <url: https: gemini.google.com ?hl="ja">Examples of generative AI include the following. The data generation model 58 is obtained by performing deep learning on a neural network. The data generation model 58 is input with prompts containing instructions, and with inference data such as audio data representing speech, text data representing text, and image data representing images. The data generation model 58 infers from the input inference data according to the instructions indicated by the prompts, and outputs the inference results in data formats such as audio data and text data. Here, inference refers to, for example, analysis, classification, prediction, and / or summarization.
[0405] In the above embodiment, an example was given in which specific processing is performed by the data processing device 12, but the technology of this disclosure is not limited thereto, and the specific processing may also be performed by the smart glasses 214.
[0406] [Third Embodiment]
[0407] Figure 5 shows an example of the configuration of the data processing system 310 according to the third embodiment.
[0408] As shown in Figure 5, the data processing system 310 includes a data processing device 12 and a headset terminal 314. An example of the data processing device 12 is a server.
[0409] The data processing device 12 comprises a computer 22, a database 24, and a communication interface 26. The computer 22 is an example of a "computer" related to the technology of this disclosure. The computer 22 comprises a processor 28, RAM 30, and storage 32. The processor 28, RAM 30, and storage 32 are connected to a bus 34. The database 24 and the communication interface 26 are also connected to the bus 34. The communication interface 26 is connected to a network 54. An example of the network 54 is a WAN (Wide Area Network) and / or a LAN (Local Area Network).
[0410] The headset terminal 314 includes a computer 36, a microphone 238, a speaker 240, a camera 42, a communication interface 44, and a display 343. The computer 36 includes a processor 46, RAM 48, and storage 50. The processor 46, RAM 48, and storage 50 are connected to a bus 52. The microphone 238, speaker 240, camera 42, and display 343 are also connected to the bus 52.
[0411] The microphone 238 receives voice signals from the user 20 and receives instructions from the user 20. The microphone 238 captures the voice signals from the user 20, converts the captured voice into audio data, and outputs it to the processor 46. The speaker 240 outputs audio according to the instructions from the processor 46.
[0412] Camera 42 is a small digital camera equipped with an optical system including a lens, aperture, and shutter, and an image sensor such as a CMOS (Complementary Metal-Oxide-Semiconductor) image sensor or a CCD (Charge Coupled Device) image sensor, and captures images of the area around the user 20 (for example, an imaging range defined by a field of view equivalent to the width of a typical healthy person's field of vision).
[0413] Communication interface 44 is connected to network 54. Communication interfaces 44 and 26 are responsible for the exchange of various information between processor 46 and processor 28 via network 54. The exchange of various information between processor 46 and processor 28 using communication interfaces 44 and 26 is performed in a secure manner.
[0414] Figure 6 shows an example of the main functions of the data processing device 12 and the headset terminal 314. As shown in Figure 6, the data processing device 12 performs specific processing using the processor 28. The storage 32 stores the specific processing program 56.
[0415] The specific processing program 56 is an example of a "program" relating to the technology of this disclosure. The processor 28 reads the specific processing program 56 from the storage 32 and executes the read specific processing program 56 on the RAM 30. The specific processing is realized by the processor 28 operating as a specific processing unit 290 in accordance with the specific processing program 56 executed on the RAM 30.
[0416] The storage 32 stores the data generation model 58 and the emotion identification model 59. The data generation model 58 and the emotion identification model 59 are used by the identification processing unit 290.
[0417] In the headset terminal 314, the processor 46 performs the reception output processing. The storage 50 stores the reception output program 60. The processor 46 reads the reception output program 60 from the storage 50 and executes the read reception output program 60 on the RAM 48. The reception output processing is realized by the processor 46 operating as a control unit 46A according to the reception output program 60 executed on the RAM 48.
[0418] Next, the specific processing performed by the specific processing unit 290 of the data processing device 12 will be described. In the following description, the data processing device 12 will be referred to as the "server" and the headset terminal 314 will be referred to as the "terminal".
[0419] This invention provides a system that simultaneously analyzes voice, facial expressions, and gestures to reduce the risk of harassment in the workplace. This system includes three main components: a terminal, a server, and a user, each playing a specific role.
[0420] The device is responsible for collecting data from the user. Specifically, it uses a microphone and camera installed in the device to acquire the user's voice and video in real time. The acquired voice data is subjected to noise reduction within the device, and facial expressions and gestures are extracted from the video data using image recognition.
[0421] Preprocessed data is sent to a server via the network. The server analyzes the received data and uses various algorithms to analyze the user's emotional state. This analysis assesses the risk of harassment and generates a risk score. The evaluation system determines the risk based on specific thresholds and generates warnings as needed.
[0422] The generated feedback is then sent back to the device via the network. Upon receiving the feedback, the device immediately provides the user with a visual or auditory warning. Based on this warning, the user can adjust their statements and behavior in a timely manner.
[0423] As a concrete example, when a user speaks during a meeting, the device records their voice and facial expressions, and the server detects changes in emotion. Through analysis, if the tone of voice becomes higher or the facial expression becomes more serious, the server assesses a high harassment risk and provides the user with a warning through the device.
[0424] In this way, this system clarifies the boundaries of harassment in real time and provides an effective means to promote safe and healthy communication.
[0425] The following describes the processing flow.
[0426] Step 1:
[0427] The device automatically activates the microphone and camera when the user starts a conversation, and collects audio and video data in real time.
[0428] Step 2:
[0429] The device applies a noise reduction filter to the collected audio data to extract clear audio, and uses an image recognition algorithm to extract facial expressions and gestures from the video data.
[0430] Step 3:
[0431] The terminal sends pre-processed voice, facial expression, and gesture data to the server as secure data packets.
[0432] Step 4:
[0433] The server decodes the received data packets, uses a speech analysis algorithm to infer emotions from the tone and speed of the voice, and analyzes facial expression and gesture data using a deep learning model.
[0434] Step 5:
[0435] Based on the analysis results, the server performs an overall sentiment assessment and scores the harassment risk based on specific criteria.
[0436] Step 6:
[0437] If the server determines that the assessed risk score exceeds a threshold, it generates a warning message in real time and sends that feedback to the terminal.
[0438] Step 7:
[0439] The device receives feedback from the server and provides the user with warning messages visually or audibly. The user then adjusts their speech and gestures based on this feedback.
[0440] (Example 1)
[0441] Next, we will describe Example 1. In the following description, the data processing device 12 will be referred to as the "server," and the headset-type terminal 314 will be referred to as the "terminal."
[0442] Preventing the deterioration of interpersonal relationships and the occurrence of harassment in the workplace is extremely important. However, traditional methods have challenges in effectively managing risks, such as difficulty in understanding emotional states in real time and providing appropriate feedback.
[0443] The identification process performed by the identification processing unit 290 of the data processing device 12 in Example 1 is realized by the following means.
[0444] In this invention, the server includes acquisition means for simultaneously acquiring audio and video information, preprocessing means for preprocessing the acquired information using noise reduction and image understanding, and analysis means for analyzing the preprocessed information and using a generative artificial intelligence model to evaluate emotional states. This makes it possible to grasp the user's emotional state in real time, assess the risks in interpersonal relationships, and provide appropriate feedback.
[0445] "Voice information" refers to sound wave data that is electronically recorded by detecting the user's speech and conversation.
[0446] "Video information" refers to image data that visually records the user's posture, movements, and facial expressions.
[0447] "Acquisition means" refers to a mechanism that simultaneously detects audio and video information and collects it as digital data.
[0448] "Preprocessing means" refers to processing means that remove noise from acquired audio information and refine video information through image analysis.
[0449] "Noise reduction" is a technology that removes unwanted noise and other sounds from audio information, thereby improving signal clarity.
[0450] "Image understanding" is a technology that recognizes facial expressions and gestures from video information and extracts them as data.
[0451] A "generative artificial intelligence model" is a trained algorithm used to identify a user's emotional state based on acquired and pre-processed information.
[0452] "Analysis means" refers to a technical mechanism introduced to analyze pre-processed information and evaluate the user's emotional state.
[0453] "Feedback" refers to notifications and advice provided to users based on results obtained from analysis.
[0454] "Interpersonal risk" is an indicator of the possibility that dialogue in the workplace environment may cause discomfort or harm to the other party.
[0455] This invention is a system that simultaneously acquires and analyzes audio and video information and provides feedback to the user in order to reduce the risk of harassment in the workplace environment.
[0456] The device uses a high-performance microphone to acquire audio information and features noise cancellation to suppress ambient noise. It also has a high-resolution camera to collect video information, meticulously recording the user's face and body movements. This two types of information are transmitted to a server via the network.
[0457] The server applies noise reduction processing to audio information and uses image recognition technology (such as OpenCV) on video information to extract the user's facial expressions and gestures. The pre-processed data is then analyzed by a generative AI model to identify the user's emotional state. Based on the analysis results, the server calculates the risk of harassment and generates feedback as needed.
[0458] Upon receiving the results of their harassment risk assessment, users receive visual or auditory feedback through their device. This allows users to appropriately adjust their words and actions on the spot, thereby maintaining a safe and healthy work environment.
[0459] As a concrete example, when a user speaks during a meeting, the device records their voice and facial expressions, which are then analyzed by a server. If changes in emotion are detected, such as a rise in voice tone or a hostile expression, the server assesses the high risk of harassment and provides the user with a warning through the device.
[0460] An example of a prompt message is, "Assess the risk that the statements made during this meeting constitute harassment." This system supports good communication in the workplace by understanding emotional states in real time and conducting appropriate risk assessments.
[0461] The flow of the specific processing in Example 1 will be explained using Figure 11.
[0462] Step 1:
[0463] The device simultaneously acquires user audio and video information using a high-performance microphone and camera. Specifically, the microphone records audio as digital data, and the camera captures the user's face and body movements in real time. The input for this step is the user's real-time conversation and movements, and the output is audio and video data converted into digital format.
[0464] Step 2:
[0465] The device performs noise reduction processing on the acquired audio data. Specifically, it removes unwanted components such as ambient noise to output a clear audio signal. For video data, it uses image understanding technology to extract facial expressions and gestures. The input is the audio and video data acquired in step 1, and the output is the noise-reduced audio data and image data formatted for analysis.
[0466] Step 3:
[0467] The terminal transmits pre-processed audio and video data to the server over the network. Specifically, the data is securely transferred using a secure protocol (e.g., SSL / TLS). The input to this step is the output data from step 2, and the output is the data ready for analysis that has been passed to the server.
[0468] Step 4:
[0469] The server uses a generative AI model to analyze the received data. Specifically, it analyzes voice tone and emotion from audio data, and facial expressions and movements from video data. The input is the data sent in step 3, and the output is the analysis result representing the user's emotional state.
[0470] Step 5:
[0471] The server assesses the user's harassment risk based on the analysis results. Specifically, it compares the calculated emotional state with a specific threshold to calculate a risk score. The input for this step is the analysis results from step 4, and the output is a score indicating the harassment risk.
[0472] Step 6:
[0473] The server generates feedback for the user based on the risk assessment. Specifically, if the risk is high, it creates a warning message and prepares the feedback content. The input is the risk score from step 5, and the output is the generated feedback message.
[0474] Step 7:
[0475] The terminal receives feedback sent from the server and notifies the user visually or audibly. Specifically, it displays a pop-up on the screen or provides an audible alert. The input for this step is the feedback message from step 6, and the output is a warning or alert notification to the user.
[0476] (Application Example 1)
[0477] Next, we will explain Application Example 1. In the following explanation, the data processing device 12 will be referred to as the "server," and the headset-type terminal 314 will be referred to as the "terminal."
[0478] There is a need to provide a system that reduces the risk of harassment in the work environment and facilitates communication between workers and autonomous machines. This invention aims to solve the problem of providing a means to detect stress and communication problems during work in real time and to respond quickly.
[0479] The specific processing performed by the specific processing unit 290 of the data processing device 12 in Application Example 1 is realized by the following means.
[0480] In this invention, the server includes a collection means for acquiring voice data, facial expression data, and motion data; a preprocessing means for preprocessing and generalizing the acquired data; and an analysis means for analyzing the preprocessed data and determining the emotional state. This makes it possible to reduce the risk of harassment in the work environment and provide appropriate feedback to workers.
[0481] "Audio data" refers to information recorded in digital format for analysis, such as the subject's speech or other sound information.
[0482] "Facial expression data" refers to information used to acquire and analyze the facial movements and expressions that indicate specific emotions of a subject, as images or videos.
[0483] "Motion data" refers to information used to digitally record and analyze gestures, including the movements of the subject's arms and body.
[0484] "Collection means" refers to a device or mechanism for acquiring voice data, facial expression data, and motion data from a subject and transmitting them to a system.
[0485] "Preprocessing means" refers to the process of processing acquired data into a format that is easy to analyze, reducing noise, and generalizing the data.
[0486] "Analysis means" refers to methods and devices for determining the emotional state of a subject and evaluating the risk using pre-processed data.
[0487] "Evaluation methods" refer to the process of determining the risk of harassment based on the analysis results and generating feedback as needed.
[0488] "Notification means" refers to a device or method for communicating information, including risks, based on evaluation results, to the target person or relevant party.
[0489] "Transmission means" refers to the process of using autonomous machines to communicate evaluation results and feedback to workers or stakeholders.
[0490] This invention is a system designed to reduce the risk of harassment in the work environment and to facilitate communication between workers and autonomous machines. This system collects and analyzes voice data, facial expression data, and motion data in real time.
[0491] The terminal is equipped with a microphone and camera to capture the worker's voice and video. These devices are used to acquire audio and video from the work environment in real time. The acquired data undergoes noise reduction and image recognition within the terminal. Specific hardware components include high-performance cameras and directional microphones. Software libraries such as TensorFlow and OpenCV are used for data preprocessing.
[0492] The server is responsible for analyzing the pre-processed data. Built on a cloud platform, the server utilizes Google Cloud Speech-to-Text and Azure Cognitive Services for speech recognition and sentiment analysis. From the analyzed audio and video data, the server determines the emotional state of workers and assesses the risk of harassment.
[0493] The evaluation results are transmitted to workers and stakeholders via notification systems. Specifically, appropriate feedback is provided through displays and speakers installed in the autonomous machines. For example, if a worker exhibits stressful behavior while working with an autonomous machine, the system will offer advice such as, "Try to calm your tone."
[0494] By utilizing generative AI models and prompt statements, feedback can become flexible and effective, tailored to the specific situation on site. An example of a prompt statement would be, "We want to design an algorithm to analyze employee stress levels in real time using audio and video data from factory work." A system configured in this way can effectively reduce the risk of harassment and enhance the safety of the work environment.
[0495] The flow of a specific process in Application Example 1 will be explained using Figure 12.
[0496] Step 1:
[0497] The terminal collects the voice and video of workers. It receives raw audio and video data acquired by a microphone and camera as input. Based on this, it performs preprocessing such as noise reduction and image recognition to remove noise from the audio and outputs standardized data in which facial expressions and movements are extracted. Specifically, the terminal runs an audio filtering algorithm and a facial recognition model in parallel.
[0498] Step 2:
[0499] The pre-processed data is sent from the terminal to the server. The server receives this as input and analyzes the voice and facial expression data. For data processing, Google Cloud Speech-to-Text is used to convert the voice to text, and emotion analysis is performed using Azure Cognitive Services. As a result, evaluation data indicating the worker's emotional state is output. Specifically, the server sequentially converts the voice to text and detects changes in emotion from the continuous data stream.
[0500] Step 3:
[0501] The server determines harassment risk based on evaluation data. It receives evaluation data indicating emotional state as input and calculates a risk score based on an algorithm. This process determines whether the risk exceeds a certain threshold. The output determines the content of the notification to the user. Specifically, the server compares the calculated risk score against multiple criteria.
[0502] Step 4:
[0503] Based on the evaluation results, the terminal sends feedback to the worker. It receives notifications from the server as input and provides feedback visually and audibly. The software used includes a speaker and display, outputting visual warning screens and audio notifications. Specifically, the terminal flexibly generates voice messages and displays related information on the screen.
[0504] Furthermore, an emotion engine that estimates the user's emotions may be incorporated. That is, the identification processing unit 290 may use the emotion identification model 59 to estimate the user's emotions and perform identification processing using the user's emotions.
[0505] This invention relates to a system that simultaneously acquires voice data, facial expression data, and gesture data, and analyzes them to evaluate the user's emotional state. By incorporating an emotion engine, this system can recognize the user's emotions with greater accuracy and assess harassment risk. The emotion engine has advanced functionality to perform integrated analysis of voice and facial expression data and identify the intensity and type of emotion.
[0506] The device continuously collects necessary data using its microphone and camera while the user performs daily tasks and conversations. The collected data undergoes basic noise reduction and data format conversion within the device to improve data quality. The pre-processed data is then transferred to a server via the network.
[0507] The server analyzes the received voice, facial expression, and gesture data. The analysis incorporates an emotion engine that comprehensively analyzes changes in voice tone, tempo, and facial expression. This allows for a more detailed assessment of emotional states, enabling a highly accurate determination of how the user is feeling. This emotional information is crucial in harassment risk assessment, and a risk score is generated accordingly.
[0508] Based on the results obtained through analysis and evaluation, the server generates warnings and advice as needed. For example, if the emotion engine identifies user stress, the server immediately sends that information to the terminal, providing the user with excellent feedback.
[0509] This feedback is provided to the user visually or audibly through their device. Based on the information provided, the user can adjust their own behavior and actions. As a result, the risk of harassment is reduced, and a more comfortable and secure work environment can be created. For example, during a team meeting, if the emotional engine detects tension or anxiety, the user is given a warning encouraging a softer approach, improving the quality of communication.
[0510] In this way, this system utilizes an emotion engine to provide an effective solution aimed at improving communication in the workplace and preventing harassment.
[0511] The following describes the processing flow.
[0512] Step 1:
[0513] The device activates its microphone and camera as soon as the user starts a conversation, collecting audio and video data.
[0514] Step 2:
[0515] The device removes noise from collected audio data and detects facial feature points from video data to preprocess initial data of facial expressions and gestures.
[0516] Step 3:
[0517] The terminal packets the processed voice, facial expression, and gesture data and sends it to the server via a secure protocol.
[0518] Step 4:
[0519] The server decodes the data received from the terminal and uses an emotion engine to accurately analyze the intensity and type of emotion from voice tone, facial expressions, and gestures.
[0520] Step 5:
[0521] The server uses the analyzed sentiment information to score the harassment risk and generates a warning when the result exceeds a certain threshold.
[0522] Step 6:
[0523] The server transmits the generated warnings and advice to the terminal, which then provides them to the user in real time, either visually or audibly.
[0524] Step 7:
[0525] Users can review the feedback they receive and adjust their attitudes and statements as needed to maintain better communication.
[0526] (Example 2)
[0527] Next, we will describe Example 2. In the following description, the data processing device 12 will be referred to as the "server," and the headset-type terminal 314 will be referred to as the "terminal."
[0528] In today's workplace environment, improving the quality of communication and reducing the risk of harassment is crucial. However, traditional methods make it difficult to accurately assess an employee's emotional state and provide appropriate feedback based on that assessment. To address this challenge, it is necessary to efficiently acquire and analyze acoustic, visual, and motion signals.
[0529] The identification process performed by the identification processing unit 290 of the data processing device 12 in Example 2 is realized by the following means.
[0530] In this invention, the server includes means for acquiring acoustic signals, visual signals, and motion signals; means for preprocessing the acquired signals and converting them into a unified format; and means for analyzing the preprocessed signals and identifying emotional states. This makes it possible to evaluate the user's emotional state with high accuracy and provide rapid feedback if there is a risk.
[0531] An "acoustic signal" is a representation of sound or audio data as an electrical signal.
[0532] "Visual signals" are representations of visual information as electrical signals, and include images and videos.
[0533] "Motion signals" are electrical signals that capture human movements and gestures.
[0534] "Means of acquisition" refers to the devices and technologies used to collect data or signals.
[0535] "Preprocessing means" refers to devices or techniques that perform initial data processing to make the data easier to analyze.
[0536] "Converting to a unified format" means unifying data obtained in different formats or standards into a single standard format that is easier to analyze.
[0537] "Means of analysis" refers to devices or technologies for analyzing acquired data in detail and extracting specific information.
[0538] "Identifying emotional states" means determining the type and intensity of a user's emotions based on the results of data analysis.
[0539] "Risk assessment" is the process of evaluating the extent to which potential dangers or problems exist based on the analyzed information.
[0540] "Feedback" refers to the act of providing advice and information to users based on acquired data and analysis results.
[0541] This invention provides a system that uses acoustic signals, visual signals, and motion signals to analyze a user's emotional state with high accuracy and reduce the risk of harassment.
[0542] The device acquires acoustic and visual signals using a microphone and camera. Acoustic signals include the user's voice, while visual signals include the user's facial expressions and gestures. The acquired signals are preprocessed within the device to improve data quality through noise reduction and formatting standardization. This preprocessing prepares the data into a consistent format for analysis.
[0543] The server receives pre-processed signals, which are then analyzed by an emotion engine. This process involves analyzing the speaker's tone and tempo from the acoustic signals, and changes in facial expressions and movements from the visual signals. The emotion engine employs sophisticated algorithms to integrate these individual elements to identify the type and intensity of emotion. The specific hardware consists of standard server equipment, while the software incorporates signal processing and machine learning models.
[0544] The analysis results are used for risk assessment to evaluate the likelihood of harassment. If the risk exceeds the assessed threshold, the server generates a warning on the terminal and provides appropriate feedback to the user. This feedback is presented visually or audibly from the terminal, allowing the user to improve their behavior based on it.
[0545] As a concrete example, during a team meeting, if the emotion engine detects tension or anxiety in the user, the server provides feedback to encourage relaxation. This can facilitate smoother communication during the meeting and reduce stress.
[0546] Examples of prompt messages include the following:
[0547] "Propose a system that analyzes a user's emotions based on acoustic, visual, and motion signals, generates feedback when stress or anxiety is detected, and demonstrate how it can be used to improve workplace communication."
[0548] In this way, the present invention provides a practical solution that analyzes the emotional state of users in detail, improves the quality of communication in the workplace, and reduces the risk of harassment.
[0549] The flow of the specific processing in Example 2 will be explained using Figure 13.
[0550] Step 1:
[0551] The device acquires acoustic and visual signals using a microphone and camera. The user's voice and facial expressions are captured as input by the microphone and camera. Specifically, the microphone collects ambient sounds, and the camera continuously captures the user's face and its movements. These signals are sent to the device's data storage.
[0552] Step 2:
[0553] The terminal performs noise reduction and data format standardization on the acquired signals. The acoustic and visual signals obtained in step 1 are used as input. Specifically, noise filtering is performed to reduce background noise. After this, the audio and visual signals are converted to a common data format to ensure data consistency. The output is a pre-processed signal suitable for analysis.
[0554] Step 3:
[0555] The server receives pre-processed signals and performs analysis using an emotion engine. The input consists of standardized acoustic and visual signals transmitted from the terminal. Specifically, the emotion engine analyzes the tone and tempo of sounds, changes in facial expressions, and gestures, correlating these elements to identify the type and intensity of emotion. The output is the analyzed emotion information.
[0556] Step 4:
[0557] The server performs a risk assessment based on the analyzed emotional information. The emotional information obtained in step 3 is used as input. Specifically, it calculates a harassment risk score using the emotional information and compares this score against the evaluation criteria. The output is the risk score and the evaluation result.
[0558] Step 5:
[0559] The server generates feedback based on the evaluation results and sends it to the terminal. The risk score obtained in step 4 is used as input. Specifically, if the risk is high, it generates a warning message or advice and sends it to the terminal. The output is specific feedback to the user.
[0560] Step 6:
[0561] Users receive feedback through their device and adjust their behavior accordingly. Input includes visual or auditory feedback provided by the device. Specific actions include, for example, taking deep breaths or adjusting their work pace if they receive relaxing feedback. The output is improved behavior or psychological state.
[0562] (Application Example 2)
[0563] Next, we will explain application example 2. In the following explanation, the data processing device 12 will be referred to as the "server," and the headset-type terminal 314 will be referred to as the "terminal."
[0564] While technologies exist that use voice, facial expressions, and movement information to evaluate human emotional states, many of them suffer from insufficient data acquisition and analysis, particularly in detecting negative emotions in communication and providing real-time feedback. Furthermore, risk assessment and countermeasures in response to emotional changes tend to lag behind, making it difficult to ensure healthy communication in workplaces and educational institutions.
[0565] The specific processing performed by the specific processing unit 290 of the data processing device 12 in Application Example 2 is realized by the following means.
[0566] In this invention, the server includes means for acquiring voice information, facial expression information, and motion information; means for preprocessing and standardizing the acquired information; and means for analyzing the preprocessed information and evaluating the emotional state. This makes it possible to provide real-time feedback and facilitate communication.
[0567] "Audio information" refers to data used to record sounds emitted by the user and analyze their tone and tempo.
[0568] "Facial expression information" refers to data used to capture changes in a user's face and analyze changes in their emotions.
[0569] "Motion information" refers to data used to capture and analyze the user's gestures and body movements.
[0570] "Means of acquisition" refers to devices and mechanisms for collecting voice information, facial expression information, and motion information in real time.
[0571] "Preprocessing methods" refer to processing techniques used to reduce noise and standardize collected data in order to improve its quality.
[0572] "Standardization" refers to the process of adjusting data into a unified format to make it easier to analyze.
[0573] "Means of analysis" refer to algorithms and software that comprehensively evaluate voice information, facial expression information, and movement information to identify emotional states.
[0574] "Means of evaluation" refers to a system that determines potential risks based on the analysis results.
[0575] "Means of provision" refers to an interface for providing feedback to users based on evaluation results and encouraging improvement.
[0576] A "means of providing real-time feedback" is a system that responds immediately to changes in emotions and provides information back to the user right away.
[0577] "Means of facilitating communication" refer to feedback mechanisms that facilitate interactions between users and reduce negative emotions.
[0578] The system for realizing this invention mainly consists of a server and a terminal. The terminal uses a microphone and camera to acquire voice information, facial expression information, and motion information in real time during the user's daily conversations and activities. The collected information undergoes initial noise reduction on the terminal, is converted into standardized data, and then transferred to the server.
[0579] The server analyzes the received data using an advanced emotion engine. This emotion engine incorporates data analysis libraries such as TensorFlow and OpenCV. The emotion engine comprehensively analyzes the tone and tempo of the voice, as well as changes in facial expressions and movements, to identify the intensity and type of the user's emotions. Subsequently, a harassment risk assessment is performed, enabling immediate, real-time feedback.
[0580] Feedback is provided to users visually or audibly through their devices. For example, if one participant shows signs of tension or anxiety during a team meeting, the server sends a notification to all participants such as, "Let's stay calm." In this way, it improves the quality of communication and contributes to promoting healthy dialogue and relationships in the workplace.
[0581] Furthermore, to realize this system, an example of a prompt using a generative AI model is: "Analyze the emotional state of the participants from the following call log, and generate appropriate feedback if negative emotions are strong." Based on the results obtained from this prompt, the system can construct more accurate feedback.
[0582] The flow of a specific process in Application Example 2 will be explained using Figure 14.
[0583] Step 1:
[0584] The device collects the user's voice information, facial expression information, and motion information. This data is acquired using a microphone and camera, respectively. Since the acquired data contains noise, noise reduction is performed. The input is raw voice information, facial expression information, and motion information, and the output is standardized data with the noise removed.
[0585] Step 2:
[0586] The terminal standardizes and formats the data, which has had noise reduced through preprocessing. This process converts the data into a format suitable for analysis on the server. The input is noise-reduced data, and the output is in a standardized data format.
[0587] Step 3:
[0588] The server receives standardized data sent from the terminal and analyzes it using an emotion engine. The emotion engine uses libraries such as TensorFlow and OpenCV to analyze voice tone, facial expression changes, and actions. The input is standardized data, and the output is the analysis result regarding the intensity and type of emotion.
[0589] Step 4:
[0590] The server assesses harassment risk based on the analysis results and generates feedback as needed. A generative AI model is used for the assessment, with a focus on detecting negative emotions. The input is the analysis results, and the output is a risk assessment score and feedback content.
[0591] Step 5:
[0592] The server sends the generated feedback to the terminal, notifying the user visually or audibly. The user can then adjust their communication style based on this. The input is the feedback content, and the output is the notification to the user and their response action.
[0593] Step 6:
[0594] Users can adjust their behavior based on feedback received from the server and take actions to improve the quality of their communication. This specifically includes softening their tone of voice or using more gentle facial expressions. The input is the feedback action, and the output is the improved communication style.
[0595] The specific processing unit 290 transmits the result of the specific processing to the headset terminal 314. In the headset terminal 314, the control unit 46A causes the speaker 240 and display 343 to output the result of the specific processing. The microphone 238 acquires audio indicating user input for the result of the specific processing. The control unit 46A transmits the audio data indicating user input acquired by the microphone 238 to the data processing unit 12. In the data processing unit 12, the specific processing unit 290 acquires the audio data.
[0596] Data generation model 58 is a type of so-called generative AI (Artificial Intelligence). One example of data generation model 58 is ChatGPT (Internet Search).<URL: https: / / openai.com / blog / chatgpt> ), Gemini (Internet search) <url: https: gemini.google.com ?hl="ja">Examples of generative AI include the following. The data generation model 58 is obtained by performing deep learning on a neural network. The data generation model 58 is input with prompts containing instructions, and with inference data such as audio data representing speech, text data representing text, and image data representing images. The data generation model 58 infers from the input inference data according to the instructions indicated by the prompts, and outputs the inference results in data formats such as audio data and text data. Here, inference refers to, for example, analysis, classification, prediction, and / or summarization.
[0597] In the above embodiment, an example was given in which specific processing is performed by the data processing device 12, but the technology of this disclosure is not limited thereto, and specific processing may also be performed by the headset terminal 314.
[0598] [Fourth Embodiment]
[0599] Figure 7 shows an example of the configuration of the data processing system 410 according to the fourth embodiment.
[0600] As shown in Figure 7, the data processing system 410 includes a data processing device 12 and a robot 414. An example of the data processing device 12 is a server.
[0601] The data processing device 12 comprises a computer 22, a database 24, and a communication interface 26. The computer 22 is an example of a "computer" related to the technology of this disclosure. The computer 22 comprises a processor 28, RAM 30, and storage 32. The processor 28, RAM 30, and storage 32 are connected to a bus 34. The database 24 and the communication interface 26 are also connected to the bus 34. The communication interface 26 is connected to a network 54. An example of the network 54 is a WAN (Wide Area Network) and / or a LAN (Local Area Network).
[0602] The robot 414 includes a computer 36, a microphone 238, a speaker 240, a camera 42, a communication interface 44, and a controlled object 443. The computer 36 includes a processor 46, RAM 48, and storage 50. The processor 46, RAM 48, and storage 50 are connected to a bus 52. The microphone 238, speaker 240, camera 42, and controlled object 443 are also connected to the bus 52.
[0603] The microphone 238 receives voice signals from the user 20 and receives instructions from the user 20. The microphone 238 captures the voice signals from the user 20, converts the captured voice into audio data, and outputs it to the processor 46. The speaker 240 outputs audio according to the instructions from the processor 46.
[0604] Camera 42 is a small digital camera equipped with an optical system including a lens, aperture, and shutter, and an image sensor such as a CMOS (Complementary Metal-Oxide-Semiconductor) image sensor or a CCD (Charge Coupled Device) image sensor, and captures images of the area around the user 20 (for example, an imaging range defined by a field of view equivalent to the width of a typical healthy person's field of vision).
[0605] Communication interface 44 is connected to network 54. Communication interfaces 44 and 26 are responsible for the exchange of various information between processor 46 and processor 28 via network 54. The exchange of various information between processor 46 and processor 28 using communication interfaces 44 and 26 is performed in a secure manner.
[0606] The controlled object 443 includes a display device, LEDs in the eyes, and motors that drive the arms, hands, and feet. The posture and gestures of the robot 414 are controlled by controlling the motors of the arms, hands, and feet. Some of the robot 414's emotions can be expressed by controlling these motors. Furthermore, the robot 414's facial expressions can also be expressed by controlling the illumination state of the LEDs in its eyes.
[0607] Figure 8 shows an example of the main functions of the data processing device 12 and the robot 414. As shown in Figure 8, the data processing device 12 performs specific processing using the processor 28. The storage 32 stores the specific processing program 56.
[0608] The specific processing program 56 is an example of a "program" relating to the technology of this disclosure. The processor 28 reads the specific processing program 56 from the storage 32 and executes the read specific processing program 56 on the RAM 30. The specific processing is realized by the processor 28 operating as a specific processing unit 290 in accordance with the specific processing program 56 executed on the RAM 30.
[0609] The storage 32 stores the data generation model 58 and the emotion identification model 59. The data generation model 58 and the emotion identification model 59 are used by the identification processing unit 290.
[0610] In robot 414, the processor 46 performs the reception output processing. The storage 50 stores the reception output program 60. The processor 46 reads the reception output program 60 from the storage 50 and executes the read reception output program 60 on the RAM 48. The reception output processing is realized by the processor 46 operating as a control unit 46A according to the reception output program 60 executed on the RAM 48.
[0611] Next, the specific processing performed by the specific processing unit 290 of the data processing device 12 will be described. In the following description, the data processing device 12 will be referred to as the "server" and the robot 414 as the "terminal".
[0612] This invention provides a system that simultaneously analyzes voice, facial expressions, and gestures to reduce the risk of harassment in the workplace. This system includes three main components: a terminal, a server, and a user, each playing a specific role.
[0613] The device is responsible for collecting data from the user. Specifically, it uses a microphone and camera installed in the device to acquire the user's voice and video in real time. The acquired voice data is subjected to noise reduction within the device, and facial expressions and gestures are extracted from the video data using image recognition.
[0614] Preprocessed data is sent to a server via the network. The server analyzes the received data and uses various algorithms to analyze the user's emotional state. This analysis assesses the risk of harassment and generates a risk score. The evaluation system determines the risk based on specific thresholds and generates warnings as needed.
[0615] The generated feedback is then sent back to the device via the network. Upon receiving the feedback, the device immediately provides the user with a visual or auditory warning. Based on this warning, the user can adjust their statements and behavior in a timely manner.
[0616] As a concrete example, when a user speaks during a meeting, the device records their voice and facial expressions, and the server detects changes in emotion. Through analysis, if the tone of voice becomes higher or the facial expression becomes more serious, the server assesses a high harassment risk and provides the user with a warning through the device.
[0617] In this way, this system clarifies the boundaries of harassment in real time and provides an effective means to promote safe and healthy communication.
[0618] The following describes the processing flow.
[0619] Step 1:
[0620] The device automatically activates the microphone and camera when the user starts a conversation, and collects audio and video data in real time.
[0621] Step 2:
[0622] The device applies a noise reduction filter to the collected audio data to extract clear audio, and uses an image recognition algorithm to extract facial expressions and gestures from the video data.
[0623] Step 3:
[0624] The terminal sends pre-processed voice, facial expression, and gesture data to the server as secure data packets.
[0625] Step 4:
[0626] The server decodes the received data packets, uses a speech analysis algorithm to infer emotions from the tone and speed of the voice, and analyzes facial expression and gesture data using a deep learning model.
[0627] Step 5:
[0628] Based on the analysis results, the server performs an overall sentiment assessment and scores the harassment risk based on specific criteria.
[0629] Step 6:
[0630] If the server determines that the assessed risk score exceeds a threshold, it generates a warning message in real time and sends that feedback to the terminal.
[0631] Step 7:
[0632] The device receives feedback from the server and provides the user with warning messages visually or audibly. The user then adjusts their speech and gestures based on this feedback.
[0633] (Example 1)
[0634] Next, we will describe Example 1. In the following description, the data processing device 12 will be referred to as the "server" and the robot 414 as the "terminal".
[0635] Preventing the deterioration of interpersonal relationships and the occurrence of harassment in the workplace is extremely important. However, traditional methods have challenges in effectively managing risks, such as difficulty in understanding emotional states in real time and providing appropriate feedback.
[0636] The identification process performed by the identification processing unit 290 of the data processing device 12 in Example 1 is realized by the following means.
[0637] In this invention, the server includes acquisition means for simultaneously acquiring audio and video information, preprocessing means for preprocessing the acquired information using noise reduction and image understanding, and analysis means for analyzing the preprocessed information and using a generative artificial intelligence model to evaluate emotional states. This makes it possible to grasp the user's emotional state in real time, assess the risks in interpersonal relationships, and provide appropriate feedback.
[0638] "Voice information" refers to sound wave data that is electronically recorded by detecting the user's speech and conversation.
[0639] "Video information" refers to image data that visually records the user's posture, movements, and facial expressions.
[0640] "Acquisition means" refers to a mechanism that simultaneously detects audio and video information and collects it as digital data.
[0641] "Preprocessing means" refers to processing means that remove noise from acquired audio information and refine video information through image analysis.
[0642] "Noise reduction" is a technology that removes unwanted noise and other sounds from audio information, thereby improving signal clarity.
[0643] "Image understanding" is a technology that recognizes facial expressions and gestures from video information and extracts them as data.
[0644] A "generative artificial intelligence model" is a trained algorithm used to identify a user's emotional state based on acquired and pre-processed information.
[0645] "Analysis means" refers to a technical mechanism introduced to analyze pre-processed information and evaluate the user's emotional state.
[0646] "Feedback" refers to notifications and advice provided to users based on results obtained from analysis.
[0647] "Interpersonal risk" is an indicator of the possibility that dialogue in the workplace environment may cause discomfort or harm to the other party.
[0648] This invention is a system that simultaneously acquires and analyzes audio and video information and provides feedback to the user in order to reduce the risk of harassment in the workplace environment.
[0649] The device uses a high-performance microphone to acquire audio information and features noise cancellation to suppress ambient noise. It also has a high-resolution camera to collect video information, meticulously recording the user's face and body movements. This two types of information are transmitted to a server via the network.
[0650] The server applies noise reduction processing to audio information and uses image recognition technology (such as OpenCV) on video information to extract the user's facial expressions and gestures. The pre-processed data is then analyzed by a generative AI model to identify the user's emotional state. Based on the analysis results, the server calculates the risk of harassment and generates feedback as needed.
[0651] Upon receiving the results of their harassment risk assessment, users receive visual or auditory feedback through their device. This allows users to appropriately adjust their words and actions on the spot, thereby maintaining a safe and healthy work environment.
[0652] As a concrete example, when a user speaks during a meeting, the device records their voice and facial expressions, which are then analyzed by a server. If changes in emotion are detected, such as a rise in voice tone or a hostile expression, the server assesses the high risk of harassment and provides the user with a warning through the device.
[0653] An example of a prompt message is, "Assess the risk that the statements made during this meeting constitute harassment." This system supports good communication in the workplace by understanding emotional states in real time and conducting appropriate risk assessments.
[0654] The flow of the specific processing in Example 1 will be explained using Figure 11.
[0655] Step 1:
[0656] The device simultaneously acquires user audio and video information using a high-performance microphone and camera. Specifically, the microphone records audio as digital data, and the camera captures the user's face and body movements in real time. The input for this step is the user's real-time conversation and movements, and the output is audio and video data converted into digital format.
[0657] Step 2:
[0658] The device performs noise reduction processing on the acquired audio data. Specifically, it removes unwanted components such as ambient noise to output a clear audio signal. For video data, it uses image understanding technology to extract facial expressions and gestures. The input is the audio and video data acquired in step 1, and the output is the noise-reduced audio data and image data formatted for analysis.
[0659] Step 3:
[0660] The terminal transmits pre-processed audio and video data to the server over the network. Specifically, the data is securely transferred using a secure protocol (e.g., SSL / TLS). The input to this step is the output data from step 2, and the output is the data ready for analysis that has been passed to the server.
[0661] Step 4:
[0662] The server uses a generative AI model to analyze the received data. Specifically, it analyzes voice tone and emotion from audio data, and facial expressions and movements from video data. The input is the data sent in step 3, and the output is the analysis result representing the user's emotional state.
[0663] Step 5:
[0664] The server assesses the user's harassment risk based on the analysis results. Specifically, it compares the calculated emotional state with a specific threshold to calculate a risk score. The input for this step is the analysis results from step 4, and the output is a score indicating the harassment risk.
[0665] Step 6:
[0666] The server generates feedback for the user based on the risk assessment. Specifically, if the risk is high, it creates a warning message and prepares the feedback content. The input is the risk score from step 5, and the output is the generated feedback message.
[0667] Step 7:
[0668] The terminal receives feedback sent from the server and notifies the user visually or audibly. Specifically, it displays a pop-up on the screen or provides an audible alert. The input for this step is the feedback message from step 6, and the output is a warning or alert notification to the user.
[0669] (Application Example 1)
[0670] Next, we will explain Application Example 1. In the following explanation, the data processing device 12 will be referred to as the "server" and the robot 414 as the "terminal".
[0671] There is a need to provide a system that reduces the risk of harassment in the work environment and facilitates communication between workers and autonomous machines. This invention aims to solve the problem of providing a means to detect stress and communication problems during work in real time and to respond quickly.
[0672] The specific processing performed by the specific processing unit 290 of the data processing device 12 in Application Example 1 is realized by the following means.
[0673] In this invention, the server includes a collection means for acquiring voice data, facial expression data, and motion data; a preprocessing means for preprocessing and generalizing the acquired data; and an analysis means for analyzing the preprocessed data and determining the emotional state. This makes it possible to reduce the risk of harassment in the work environment and provide appropriate feedback to workers.
[0674] "Audio data" refers to information recorded in digital format for analysis, such as the subject's speech or other sound information.
[0675] "Facial expression data" refers to information used to acquire and analyze the facial movements and expressions that indicate specific emotions of a subject, as images or videos.
[0676] "Motion data" refers to information used to digitally record and analyze gestures, including the movements of the subject's arms and body.
[0677] "Collection means" refers to a device or mechanism for acquiring voice data, facial expression data, and motion data from a subject and transmitting them to a system.
[0678] "Preprocessing means" refers to the process of processing acquired data into a format that is easy to analyze, reducing noise, and generalizing the data.
[0679] "Analysis means" refers to methods and devices for determining the emotional state of a subject and evaluating the risk using pre-processed data.
[0680] "Evaluation methods" refer to the process of determining the risk of harassment based on the analysis results and generating feedback as needed.
[0681] "Notification means" refers to a device or method for communicating information, including risks, based on evaluation results, to the target person or relevant party.
[0682] "Transmission means" refers to the process of using autonomous machines to communicate evaluation results and feedback to workers or stakeholders.
[0683] This invention is a system designed to reduce the risk of harassment in the work environment and to facilitate communication between workers and autonomous machines. This system collects and analyzes voice data, facial expression data, and motion data in real time.
[0684] The terminal is equipped with a microphone and camera to capture the worker's voice and video. These devices are used to acquire audio and video from the work environment in real time. The acquired data undergoes noise reduction and image recognition within the terminal. Specific hardware components include high-performance cameras and directional microphones. Software libraries such as TensorFlow and OpenCV are used for data preprocessing.
[0685] The server is responsible for analyzing the pre-processed data. Built on a cloud platform, the server utilizes Google Cloud Speech-to-Text and Azure Cognitive Services for speech recognition and sentiment analysis. From the analyzed audio and video data, the server determines the emotional state of workers and assesses the risk of harassment.
[0686] The evaluation results are transmitted to workers and stakeholders via notification systems. Specifically, appropriate feedback is provided through displays and speakers installed in the autonomous machines. For example, if a worker exhibits stressful behavior while working with an autonomous machine, the system will offer advice such as, "Try to calm your tone."
[0687] By utilizing generative AI models and prompt statements, feedback can become flexible and effective, tailored to the specific situation on site. An example of a prompt statement would be, "We want to design an algorithm to analyze employee stress levels in real time using audio and video data from factory work." A system configured in this way can effectively reduce the risk of harassment and enhance the safety of the work environment.
[0688] The flow of a specific process in Application Example 1 will be explained using Figure 12.
[0689] Step 1:
[0690] The terminal collects the voice and video of workers. It receives raw audio and video data acquired by a microphone and camera as input. Based on this, it performs preprocessing such as noise reduction and image recognition to remove noise from the audio and outputs standardized data in which facial expressions and movements are extracted. Specifically, the terminal runs an audio filtering algorithm and a facial recognition model in parallel.
[0691] Step 2:
[0692] The pre-processed data is sent from the terminal to the server. The server receives this as input and analyzes the voice and facial expression data. For data processing, Google Cloud Speech-to-Text is used to convert the voice to text, and emotion analysis is performed using Azure Cognitive Services. As a result, evaluation data indicating the worker's emotional state is output. Specifically, the server sequentially converts the voice to text and detects changes in emotion from the continuous data stream.
[0693] Step 3:
[0694] The server determines harassment risk based on evaluation data. It receives evaluation data indicating emotional state as input and calculates a risk score based on an algorithm. This process determines whether the risk exceeds a certain threshold. The output determines the content of the notification to the user. Specifically, the server compares the calculated risk score against multiple criteria.
[0695] Step 4:
[0696] Based on the evaluation results, the terminal sends feedback to the worker. It receives notifications from the server as input and provides feedback visually and audibly. The software used includes a speaker and display, outputting visual warning screens and audio notifications. Specifically, the terminal flexibly generates voice messages and displays related information on the screen.
[0697] Furthermore, an emotion engine that estimates the user's emotions may be incorporated. That is, the identification processing unit 290 may use the emotion identification model 59 to estimate the user's emotions and perform identification processing using the user's emotions.
[0698] This invention relates to a system that simultaneously acquires voice data, facial expression data, and gesture data, and analyzes them to evaluate the user's emotional state. By incorporating an emotion engine, this system can recognize the user's emotions with greater accuracy and assess harassment risk. The emotion engine has advanced functionality to perform integrated analysis of voice and facial expression data and identify the intensity and type of emotion.
[0699] The device continuously collects necessary data using its microphone and camera while the user performs daily tasks and conversations. The collected data undergoes basic noise reduction and data format conversion within the device to improve data quality. The pre-processed data is then transferred to a server via the network.
[0700] The server analyzes the received voice, facial expression, and gesture data. The analysis incorporates an emotion engine that comprehensively analyzes changes in voice tone, tempo, and facial expression. This allows for a more detailed assessment of emotional states, enabling a highly accurate determination of how the user is feeling. This emotional information is crucial in harassment risk assessment, and a risk score is generated accordingly.
[0701] Based on the results obtained through analysis and evaluation, the server generates warnings and advice as needed. For example, if the emotion engine identifies user stress, the server immediately sends that information to the terminal, providing the user with excellent feedback.
[0702] This feedback is provided to the user visually or audibly through their device. Based on the information provided, the user can adjust their own behavior and actions. As a result, the risk of harassment is reduced, and a more comfortable and secure work environment can be created. For example, during a team meeting, if the emotional engine detects tension or anxiety, the user is given a warning encouraging a softer approach, improving the quality of communication.
[0703] In this way, this system utilizes an emotion engine to provide an effective solution aimed at improving communication in the workplace and preventing harassment.
[0704] The following describes the processing flow.
[0705] Step 1:
[0706] The device activates its microphone and camera as soon as the user starts a conversation, collecting audio and video data.
[0707] Step 2:
[0708] The device removes noise from collected audio data and detects facial feature points from video data to preprocess initial data of facial expressions and gestures.
[0709] Step 3:
[0710] The terminal packets the processed voice, facial expression, and gesture data and sends it to the server via a secure protocol.
[0711] Step 4:
[0712] The server decodes the data received from the terminal and uses an emotion engine to accurately analyze the intensity and type of emotion from voice tone, facial expressions, and gestures.
[0713] Step 5:
[0714] The server uses the analyzed sentiment information to score the harassment risk and generates a warning when the result exceeds a certain threshold.
[0715] Step 6:
[0716] The server transmits the generated warnings and advice to the terminal, which then provides them to the user in real time, either visually or audibly.
[0717] Step 7:
[0718] Users can review the feedback they receive and adjust their attitudes and statements as needed to maintain better communication.
[0719] (Example 2)
[0720] Next, we will describe Example 2. In the following description, the data processing device 12 will be referred to as the "server" and the robot 414 as the "terminal".
[0721] In today's workplace environment, improving the quality of communication and reducing the risk of harassment is crucial. However, traditional methods make it difficult to accurately assess an employee's emotional state and provide appropriate feedback based on that assessment. To address this challenge, it is necessary to efficiently acquire and analyze acoustic, visual, and motion signals.
[0722] The identification process performed by the identification processing unit 290 of the data processing device 12 in Example 2 is realized by the following means.
[0723] In this invention, the server includes means for acquiring acoustic signals, visual signals, and motion signals; means for preprocessing the acquired signals and converting them into a unified format; and means for analyzing the preprocessed signals and identifying emotional states. This makes it possible to evaluate the user's emotional state with high accuracy and provide rapid feedback if there is a risk.
[0724] An "acoustic signal" is a representation of sound or audio data as an electrical signal.
[0725] "Visual signals" are representations of visual information as electrical signals, and include images and videos.
[0726] "Motion signals" are electrical signals that capture human movements and gestures.
[0727] "Means of acquisition" refers to the devices and technologies used to collect data or signals.
[0728] "Preprocessing means" refers to devices or techniques that perform initial data processing to make the data easier to analyze.
[0729] "Converting to a unified format" means unifying data obtained in different formats or standards into a single standard format that is easier to analyze.
[0730] "Means of analysis" refers to devices or technologies for analyzing acquired data in detail and extracting specific information.
[0731] "Identifying emotional states" means determining the type and intensity of a user's emotions based on the results of data analysis.
[0732] "Risk assessment" is the process of evaluating the extent to which potential dangers or problems exist based on the analyzed information.
[0733] "Feedback" refers to the act of providing advice and information to users based on acquired data and analysis results.
[0734] This invention provides a system that uses acoustic signals, visual signals, and motion signals to analyze a user's emotional state with high accuracy and reduce the risk of harassment.
[0735] The device acquires acoustic and visual signals using a microphone and camera. Acoustic signals include the user's voice, while visual signals include the user's facial expressions and gestures. The acquired signals are preprocessed within the device to improve data quality through noise reduction and formatting standardization. This preprocessing prepares the data into a consistent format for analysis.
[0736] The server receives pre-processed signals, which are then analyzed by an emotion engine. This process involves analyzing the speaker's tone and tempo from the acoustic signals, and changes in facial expressions and movements from the visual signals. The emotion engine employs sophisticated algorithms to integrate these individual elements to identify the type and intensity of emotion. The specific hardware consists of standard server equipment, while the software incorporates signal processing and machine learning models.
[0737] The analysis results are used for risk assessment to evaluate the likelihood of harassment. If the risk exceeds the assessed threshold, the server generates a warning on the terminal and provides appropriate feedback to the user. This feedback is presented visually or audibly from the terminal, allowing the user to improve their behavior based on it.
[0738] As a concrete example, during a team meeting, if the emotion engine detects tension or anxiety in the user, the server provides feedback to encourage relaxation. This can facilitate smoother communication during the meeting and reduce stress.
[0739] Examples of prompt messages include the following:
[0740] "Propose a system that analyzes a user's emotions based on acoustic, visual, and motion signals, generates feedback when stress or anxiety is detected, and demonstrate how it can be used to improve workplace communication."
[0741] In this way, the present invention provides a practical solution that analyzes the emotional state of users in detail, improves the quality of communication in the workplace, and reduces the risk of harassment.
[0742] The flow of the specific processing in Example 2 will be explained using Figure 13.
[0743] Step 1:
[0744] The device acquires acoustic and visual signals using a microphone and camera. The user's voice and facial expressions are captured as input by the microphone and camera. Specifically, the microphone collects ambient sounds, and the camera continuously captures the user's face and its movements. These signals are sent to the device's data storage.
[0745] Step 2:
[0746] The terminal performs noise reduction and data format standardization on the acquired signals. The acoustic and visual signals obtained in step 1 are used as input. Specifically, noise filtering is performed to reduce background noise. After this, the audio and visual signals are converted to a common data format to ensure data consistency. The output is a pre-processed signal suitable for analysis.
[0747] Step 3:
[0748] The server receives pre-processed signals and performs analysis using an emotion engine. The input consists of standardized acoustic and visual signals transmitted from the terminal. Specifically, the emotion engine analyzes the tone and tempo of sounds, changes in facial expressions, and gestures, correlating these elements to identify the type and intensity of emotion. The output is the analyzed emotion information.
[0749] Step 4:
[0750] The server performs a risk assessment based on the analyzed emotional information. The emotional information obtained in step 3 is used as input. Specifically, it calculates a harassment risk score using the emotional information and compares this score against the evaluation criteria. The output is the risk score and the evaluation result.
[0751] Step 5:
[0752] The server generates feedback based on the evaluation results and sends it to the terminal. The risk score obtained in step 4 is used as input. Specifically, if the risk is high, it generates a warning message or advice and sends it to the terminal. The output is specific feedback to the user.
[0753] Step 6:
[0754] Users receive feedback through their device and adjust their behavior accordingly. Input includes visual or auditory feedback provided by the device. Specific actions include, for example, taking deep breaths or adjusting their work pace if they receive relaxing feedback. The output is improved behavior or psychological state.
[0755] (Application Example 2)
[0756] Next, we will explain application example 2. In the following explanation, the data processing device 12 will be referred to as the "server" and the robot 414 as the "terminal".
[0757] While technologies exist that use voice, facial expressions, and movement information to evaluate human emotional states, many of them suffer from insufficient data acquisition and analysis, particularly in detecting negative emotions in communication and providing real-time feedback. Furthermore, risk assessment and countermeasures in response to emotional changes tend to lag behind, making it difficult to ensure healthy communication in workplaces and educational institutions.
[0758] The specific processing performed by the specific processing unit 290 of the data processing device 12 in Application Example 2 is realized by the following means.
[0759] In this invention, the server includes means for acquiring voice information, facial expression information, and motion information; means for preprocessing and standardizing the acquired information; and means for analyzing the preprocessed information and evaluating the emotional state. This makes it possible to provide real-time feedback and facilitate communication.
[0760] "Audio information" refers to data used to record sounds emitted by the user and analyze their tone and tempo.
[0761] "Facial expression information" refers to data used to capture changes in a user's face and analyze changes in their emotions.
[0762] "Motion information" refers to data used to capture and analyze the user's gestures and body movements.
[0763] "Means of acquisition" refers to devices and mechanisms for collecting voice information, facial expression information, and motion information in real time.
[0764] "Preprocessing methods" refer to processing techniques used to reduce noise and standardize collected data in order to improve its quality.
[0765] "Standardization" refers to the process of adjusting data into a unified format to make it easier to analyze.
[0766] "Means of analysis" refer to algorithms and software that comprehensively evaluate voice information, facial expression information, and movement information to identify emotional states.
[0767] "Means of evaluation" refers to a system that determines potential risks based on the analysis results.
[0768] "Means of provision" refers to an interface for providing feedback to users based on evaluation results and encouraging improvement.
[0769] A "means of providing real-time feedback" is a system that responds immediately to changes in emotions and provides information back to the user right away.
[0770] "Means of facilitating communication" refer to feedback mechanisms that facilitate interactions between users and reduce negative emotions.
[0771] The system for realizing this invention mainly consists of a server and a terminal. The terminal uses a microphone and camera to acquire voice information, facial expression information, and motion information in real time during the user's daily conversations and activities. The collected information undergoes initial noise reduction on the terminal, is converted into standardized data, and then transferred to the server.
[0772] The server analyzes the received data using an advanced emotion engine. This emotion engine incorporates data analysis libraries such as TensorFlow and OpenCV. The emotion engine comprehensively analyzes the tone and tempo of the voice, as well as changes in facial expressions and movements, to identify the intensity and type of the user's emotions. Subsequently, a harassment risk assessment is performed, enabling immediate, real-time feedback.
[0773] Feedback is provided to users visually or audibly through their devices. For example, if one participant shows signs of tension or anxiety during a team meeting, the server sends a notification to all participants such as, "Let's stay calm." In this way, it improves the quality of communication and contributes to promoting healthy dialogue and relationships in the workplace.
[0774] Furthermore, to realize this system, an example of a prompt using a generative AI model is: "Analyze the emotional state of the participants from the following call log, and generate appropriate feedback if negative emotions are strong." Based on the results obtained from this prompt, the system can construct more accurate feedback.
[0775] The flow of a specific process in Application Example 2 will be explained using Figure 14.
[0776] Step 1:
[0777] The device collects the user's voice information, facial expression information, and motion information. This data is acquired using a microphone and camera, respectively. Since the acquired data contains noise, noise reduction is performed. The input is raw voice information, facial expression information, and motion information, and the output is standardized data with the noise removed.
[0778] Step 2:
[0779] The terminal standardizes and formats the data, which has had noise reduced through preprocessing. This process converts the data into a format suitable for analysis on the server. The input is noise-reduced data, and the output is in a standardized data format.
[0780] Step 3:
[0781] The server receives standardized data sent from the terminal and analyzes it using an emotion engine. The emotion engine uses libraries such as TensorFlow and OpenCV to analyze voice tone, facial expression changes, and actions. The input is standardized data, and the output is the analysis result regarding the intensity and type of emotion.
[0782] Step 4:
[0783] The server assesses harassment risk based on the analysis results and generates feedback as needed. A generative AI model is used for the assessment, with a focus on detecting negative emotions. The input is the analysis results, and the output is a risk assessment score and feedback content.
[0784] Step 5:
[0785] The server sends the generated feedback to the terminal, notifying the user visually or audibly. The user can then adjust their communication style based on this. The input is the feedback content, and the output is the notification to the user and their response action.
[0786] Step 6:
[0787] Users can adjust their behavior based on feedback received from the server and take actions to improve the quality of their communication. This specifically includes softening their tone of voice or using more gentle facial expressions. The input is the feedback action, and the output is the improved communication style.
[0788] The specific processing unit 290 transmits the result of the specific processing to the robot 414. In the robot 414, the control unit 46A causes the speaker 240 and the controlled object 443 to output the result of the specific processing. The microphone 238 acquires audio indicating user input for the result of the specific processing. The control unit 46A transmits the audio data indicating user input acquired by the microphone 238 to the data processing unit 12. In the data processing unit 12, the specific processing unit 290 acquires the audio data.
[0789] Data generation model 58 is a type of so-called generative AI (Artificial Intelligence). One example of data generation model 58 is ChatGPT (Internet Search).<URL: https: / / openai.com / blog / chatgpt> ), Gemini (Internet search) <url: https: gemini.google.com ?hl="ja">Examples of generative AI include the following. The data generation model 58 is obtained by performing deep learning on a neural network. The data generation model 58 is input with prompts containing instructions, and with inference data such as audio data representing speech, text data representing text, and image data representing images. The data generation model 58 infers from the input inference data according to the instructions indicated by the prompts, and outputs the inference results in data formats such as audio data and text data. Here, inference refers to, for example, analysis, classification, prediction, and / or summarization.
[0790] In the above embodiment, an example was given in which specific processing is performed by the data processing device 12, but the technology of this disclosure is not limited thereto, and the specific processing may also be performed by the robot 414.
[0791] Furthermore, the emotion identification model 59, acting as an emotion engine, may determine the user's emotion according to a specific mapping. Specifically, the emotion identification model 59 may determine the user's emotion according to a specific mapping, which is an emotion map (see Figure 9). Similarly, the emotion identification model 59 may also determine the robot's emotion, and the identification processing unit 290 may perform identification processing using the robot's emotion.
[0792] Figure 9 shows an emotion map 400 in which multiple emotions are mapped. In the emotion map 400, emotions are arranged in concentric circles radiating from the center. The closer to the center of the concentric circles, the more primitive the emotions are located. Further out of the concentric circles, emotions representing states and actions arising from mental states are located. Emotion is a concept that includes feelings and mental states. On the left side of the concentric circles, emotions that are generally generated from reactions occurring in the brain are located. On the right side of the concentric circles, emotions that are generally induced by situational judgment are located. Above and below the concentric circles, emotions that are generally generated from reactions occurring in the brain and induced by situational judgment are located. In addition, the emotion of "pleasure" is located on the upper side of the concentric circles, and the emotion of "displeasure" is located on the lower side. Thus, in the emotion map 400, multiple emotions are mapped based on the structure in which emotions arise, and emotions that are likely to occur simultaneously are mapped close together.
[0793] These emotions are distributed at the 3 o'clock position on the Emotion Map 400, and usually fluctuate between feelings of security and anxiety. In the right half of the Emotion Map 400, situational awareness takes precedence over internal feelings, resulting in a calm impression.
[0794] The inside of the Emotion Map 400 represents inner thoughts, while the outside represents actions. Therefore, the further you go from the outside of the Emotion Map 400, the more visible (expressed in actions) your emotions become.
[0795] Here, human emotions are based on various balances, such as posture and blood sugar levels. When these balances deviate from the ideal, it results in discomfort, and when they approach the ideal, it results in pleasure. Similarly, in robots, cars, motorcycles, etc., emotions can be created based on various balances, such as posture and battery level. When these balances deviate from the ideal, it results in discomfort, and when they approach the ideal, it results in pleasure. The emotion map can be generated, for example, based on Dr. Mitsuyoshi's emotion map (Research on a system for analyzing brain physiological signals of speech emotion recognition and emotion, Tokushima University, doctoral dissertation: https: / / ci.nii.ac.jp / naid / 500000375379). The left half of the emotion map contains emotions belonging to a region called "response," where sensation is dominant. The right half of the emotion map contains emotions belonging to a region called "situation," where situational awareness is dominant.
[0796] The emotion map defines two emotions that promote learning. One is the emotion around the middle of the negative "repentance" and "reflection" on the situation side. In other words, it is when the robot experiences negative emotions such as "I never want to feel this way again" or "I don't want to be scolded again." The other is the emotion around the positive "desire" on the reaction side. In other words, it is when the robot has positive feelings such as "I want more" or "I want to know more."
[0797] The emotion identification model 59 inputs user input into a pre-trained neural network, obtains emotion values representing each emotion shown in the emotion map 400, and determines the user's emotion. This neural network is pre-trained based on multiple training data sets, which are combinations of user input and emotion values representing each emotion shown in the emotion map 400. Furthermore, this neural network is trained so that emotions located close together have similar values, as shown in the emotion map 900 in Figure 10. Figure 10 shows an example where multiple emotions such as "reassured," "calm," and "confident" have similar emotion values.
[0798] The above description primarily focuses on the functions of the data processing device 12 in relation to this disclosure. However, the system related to this disclosure is not necessarily implemented on a server. The system related to this disclosure may be implemented as a general information processing system. This disclosure may be implemented, for example, as a software program that runs on a personal computer or as an application that runs on a smartphone. The method related to this disclosure may be provided to users in SaaS (Software as a Service) format.
[0799] In the above embodiment, an example was given in which a specific process is performed by a single computer 22. However, the technology of this disclosure is not limited thereto, and a distributed processing of the specific process may be performed by multiple computers, including computer 22. For example, a data generation model 58 may be provided in an external device of the data processing device 12, and the external device may generate data according to the input data.
[0800] In the above embodiment, an example was given in which the specific processing program 56 is stored in the storage 32, but the technology of this disclosure is not limited thereto. For example, the specific processing program 56 may be stored in a portable, computer-readable, non-temporary storage medium such as a USB (Universal Serial Bus) memory. The specific processing program 56 stored in the non-temporary storage medium is installed in the computer 22 of the data processing device 12. The processor 28 executes specific processing according to the specific processing program 56.
[0801] Alternatively, the specific processing program 56 may be stored in a storage device such as a server connected to the data processing device 12 via the network 54, and the specific processing program 56 may be downloaded and installed on the computer 22 in response to a request from the data processing device 12.
[0802] Furthermore, it is not necessary to store the entirety of the specific processing program 56 in a storage device such as a server connected to the data processing device 12 via the network 54, or to store the entirety of the specific processing program 56 in the storage 32; it is acceptable to store only a portion of the specific processing program 56.
[0803] The following types of processors can be used as hardware resources to perform specific processing. Examples of processors include a CPU, a general-purpose processor that functions as a hardware resource to perform specific processing by executing software, i.e., a program. Other examples of processors include dedicated electrical circuits, such as FPGAs (Field-Programmable Gate Arrays), PLDs (Programmable Logic Devices), or ASICs (Application Specific Integrated Circuits), which have circuit configurations specifically designed to perform specific processing. All of these processors have built-in or connected memory, and all of them perform specific processing by using memory.
[0804] The hardware resource that performs a specific process may consist of one of these various processors, or it may consist of a combination of two or more processors of the same or different types (for example, a combination of multiple FPGAs, or a combination of a CPU and an FPGA). Alternatively, the hardware resource that performs a specific process may consist of a single processor.
[0805] Examples of configurations using a single processor include, firstly, a configuration in which one or more CPUs and software are combined to form a single processor, and this processor functions as a hardware resource that performs a specific process. Secondly, there is a configuration using a processor that realizes the functions of the entire system, including multiple hardware resources that perform a specific process, on a single IC chip, as exemplified by SoCs (System-on-a-chip). In this way, a specific process is realized using one or more of the above types of processors as hardware resources.
[0806] Furthermore, the hardware structure of these various processors can more specifically utilize electrical circuits that combine circuit elements such as semiconductor devices. Also, the specific processing described above is merely an example. Therefore, it goes without saying that unnecessary steps can be deleted, new steps added, or the processing order rearranged, as long as it does not deviate from the main purpose.
[0807] The descriptions and illustrations presented above are detailed explanations of the technical aspects of this disclosure and are merely examples of the technical aspects. For example, the above descriptions of the structure, function, operation, and effect are examples of the structure, function, operation, and effect of the technical aspects of this disclosure. Therefore, it goes without saying that you may delete unnecessary parts, add new elements, or replace elements in the descriptions and illustrations presented above, as long as you do not deviate from the essence of the technical aspects of this disclosure. Furthermore, in order to avoid confusion and facilitate understanding of the technical aspects of this disclosure, explanations of common technical knowledge and the like that do not require special explanation to enable the implementation of the technical aspects of this disclosure have been omitted from the descriptions and illustrations presented above.
[0808] All documents, patent applications, and technical standards described herein are incorporated by reference to the same extent as if each individual document, patent application, and technical standard were specifically and individually noted as being incorporated by reference.
[0809] The following is further disclosed regarding the embodiments described above.
[0810] (Claim 1)
[0811] A means of acquiring voice data, facial expression data, and gesture data,
[0812] A preprocessing means for preprocessing and standardizing the acquired data,
[0813] An analytical means for analyzing pre-processed data and evaluating emotional states,
[0814] An evaluation method for assessing harassment risk based on analysis results,
[0815] A feedback mechanism for providing evaluation results to the user,
[0816] A system that includes this.
[0817] (Claim 2)
[0818] The system according to claim 1, wherein an evaluation means generates a warning when it exceeds a specific threshold, and a feedback means notifies the user of this warning.
[0819] (Claim 3)
[0820] The system according to claim 1, wherein the acquisition means acquires audio and video data simultaneously, and the preprocessing means processes the data using noise reduction and image recognition.
[0821] "Example 1"
[0822] (Claim 1)
[0823] A means for simultaneously acquiring audio and video information,
[0824] A preprocessing means that preprocesses the acquired information and processes it using noise reduction and image understanding,
[0825] An analysis means that uses a generative artificial intelligence model to analyze preprocessed information and evaluate emotional states,
[0826] An evaluation method for assessing the risks in interpersonal relationships based on the analysis results,
[0827] A feedback mechanism that provides warnings or improvement instructions based on evaluation results and notifies the user,
[0828] A system that includes this.
[0829] (Claim 2)
[0830] The system according to claim 1, wherein if the evaluation means exceeds a set threshold value, the feedback means generates a warning and notifies the user.
[0831] (Claim 3)
[0832] The system according to claim 1, wherein the acquisition means simultaneously acquires audio and video information, and the preprocessing means processes the information using noise suppression and image analysis.
[0833] "Application Example 1"
[0834] (Claim 1)
[0835] A means for acquiring voice data, facial expression data, and motion data,
[0836] A preprocessing means for preprocessing and generalizing the acquired data,
[0837] An analysis means for analyzing pre-processed data and determining emotional state,
[0838] An evaluation method that assesses harassment risk based on analysis results and generates feedback for the work environment,
[0839] A means of providing feedback on evaluation results to workers,
[0840] A transmission means for sending evaluation results and feedback to workers via an autonomous machine,
[0841] A system that includes this.
[0842] (Claim 2)
[0843] The system according to claim 1, wherein an evaluation means generates a warning when it exceeds a certain standard, and a feedback means notifies the worker of this warning.
[0844] (Claim 3)
[0845] The system according to claim 1, wherein the acquisition means simultaneously acquires audio and video data, and the preprocessing means processes the data using noise reduction and image analysis.
[0846] "Example 2 of combining an emotion engine"
[0847] (Claim 1)
[0848] Means for acquiring acoustic signals, visual signals, and motion signals,
[0849] A means for preprocessing the acquired signal and converting it into a unified format,
[0850] A means for analyzing pre-processed signals and identifying emotional states,
[0851] A means of conducting a risk assessment based on the analysis results,
[0852] Means for providing evaluation results to users,
[0853] A system that includes this.
[0854] (Claim 2)
[0855] The system according to claim 1, further comprising a feedback means for generating a warning and notifying the user when the risk assessment means exceeds a defined standard.
[0856] (Claim 3)
[0857] The system according to claim 1, wherein the acquisition means simultaneously acquires an acoustic signal and a visual signal, and the preprocessing means optimizes the signal using signal processing technology.
[0858] "Application example 2 of combining emotional engines"
[0859] (Claim 1)
[0860] Means for acquiring voice information, facial expression information, and motion information,
[0861] A means of preprocessing and standardizing the acquired information,
[0862] A means of analyzing pre-processed information and evaluating emotional states,
[0863] A means of evaluating risk based on the analysis results,
[0864] Means of providing evaluation results to users,
[0865] A means of providing real-time feedback and facilitating communication,
[0866] A device that includes this.
[0867] (Claim 2)
[0868] The apparatus according to claim 1, wherein an evaluation means generates a notification when it exceeds a predetermined standard, and a providing means provides this notification to the user.
[0869] (Claim 3)
[0870] The apparatus according to claim 1, wherein the acquisition means acquires audio information and video information simultaneously, and the preprocessing means processes the information using noise reduction and recognition techniques. [Explanation of Symbols]
[0871] 10, 210, 310, 410 Data Processing Systems 12 Data Processing Devices 14 Smart Devices 214 Smart Glasses 314 Headset-type terminal 414 Robots< / url:> < / url:> < / url:> < / url:>
Claims
1. A means of acquiring voice data, facial expression data, and gesture data, A preprocessing means for preprocessing and standardizing the acquired data, An analytical means for analyzing pre-processed data and evaluating emotional states, An evaluation method for assessing harassment risk based on analysis results, A feedback mechanism for providing evaluation results to the user, A system that includes this.
2. The system according to claim 1, wherein an evaluation means generates a warning when it exceeds a specific threshold, and a feedback means notifies the user of this warning.
3. The system according to claim 1, wherein the acquisition means acquires audio and video data simultaneously, and the preprocessing means processes the data using noise reduction and image recognition.
Citation Information
Patent Citations
Persona chatbot control method and system
JP2022180282A