system

A system for detecting power harassment through speech recognition and sentiment analysis in workplace meetings provides immediate alerts, addressing the challenge of delayed detection and enhancing workplace safety and mental health.

JP2026070175APending Publication Date: 2026-04-27SOFTBANK GROUP CORP
View PDF 1 Cites 0 Cited by

Patent Information

Authority / Receiving Office
JP · JP
Patent Type
Applications
Current Assignee / Owner
SOFTBANK GROUP CORP
Filing Date
2024-10-15
Publication Date
2026-04-27

AI Technical Summary

Technical Problem

Power harassment in the workplace causes stress and decreased motivation among employees, and conventional detection methods rely on delayed victim reports, lacking effective early detection and countermeasures.

Method used

A system that acquires audio data from workplace meetings, converts it into text using speech recognition, performs sentiment analysis to detect signs of power harassment, and automatically generates warning information via email or messaging platforms, enabling immediate notification and response.

Benefits of technology

Enables rapid detection and prevention of power harassment, improving employee safety and workplace mental health by allowing immediate action on inappropriate remarks.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure 2026070175000001_ABST
    Figure 2026070175000001_ABST
Patent Text Reader

Abstract

We provide the system. [Solution] A voice input means for acquiring voice data, A speech recognition means that converts speech into text data based on acquired speech data, A sentiment analysis method that analyzes text data and extracts emotional information, An alert generation means that generates warning information when specific conditions are met based on emotional information, A notification means that notifies relevant parties of warning information using a pre-configured communication method, A system that includes this.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The technology of the present disclosure relates to a system.

Background Art

[0002] Patent Document 1 discloses a method for controlling a persona chatbot, which is performed by at least one processor and includes steps of receiving a user utterance, adding the user utterance to a prompt including an instruction sentence related to an explanation of a character of the chatbot, encoding the prompt, and inputting the encoded prompt into a language model to generate a chatbot utterance in response to the user utterance.

Prior Art Documents

Patent Documents

[0003]

Patent Document 1

Summary of the Invention

Problems to be Solved by the Invention

[0004] Power harassment in the workplace is a problem that causes stress and decreased motivation among employees, and ultimately has an adverse impact on the productivity of the entire company. Conventionally, the detection and countermeasures of power harassment rely on the reports of victims or the observations of third parties, and there is a risk that the damage will expand due to the delay in timing. In addition, although it is required to detect the signs of power harassment in advance and take prompt countermeasures, there is currently a problem that effective means to achieve this are lacking.

Means for Solving the Problems

[0005] This invention provides a system that acquires audio data from workplace meetings and other events in real time and converts the audio into text data using speech recognition technology. It performs sentiment analysis on the text data to detect signs of power harassment. If signs are detected, it generates warning information according to pre-set criteria and automatically notifies relevant parties via email or online messaging platforms. This enables a rapid response, leading to the early detection and prevention of power harassment. Furthermore, by providing control means to instruct the user to take specified actions based on the warning information, the system can enhance the effectiveness of preventing further harm and improving corrective measures.

[0006] "Voice input means" refers to devices or programs for acquiring voice, and their role is to collect voice data from meetings and conversations.

[0007] "Speech recognition means" refers to technologies and devices that analyze acquired speech data and convert the speech into corresponding text data.

[0008] "Emotional analysis tools" refer to algorithms and programs that analyze text data to extract the emotional state of the speaker.

[0009] An "alert generation method" refers to a means of creating warning information when specific conditions are met, based on information obtained through sentiment analysis.

[0010] "Notification means" refers to the technology or mechanism for sending generated warning information to relevant parties through communication means such as email or online messaging platforms.

[0011] "Control means" refers to a device or program that has the function of instructing the system or user to take appropriate action based on warning information. [Brief explanation of the drawing]

[0012] [Figure 1]This is a conceptual diagram showing an example of the configuration of a data processing system according to the first embodiment. [Figure 2] This is a conceptual diagram showing an example of the essential functions of a data processing device and a smart device according to the first embodiment. [Figure 3] This is a conceptual diagram showing an example of the configuration of a data processing system according to the second embodiment. [Figure 4] This is a conceptual diagram showing an example of the main functions of a data processing device and smart glasses according to the second embodiment. [Figure 5] This is a conceptual diagram showing an example of the configuration of a data processing system according to the third embodiment. [Figure 6] This is a conceptual diagram showing an example of the main functions of a data processing device and a headset-type terminal according to the third embodiment. [Figure 7] This is a conceptual diagram showing an example of the configuration of a data processing system according to the fourth embodiment. [Figure 8] This is a conceptual diagram showing an example of the main functions of a data processing device and a robot according to the fourth embodiment. [Figure 9] This shows an emotion map where multiple emotions are mapped. [Figure 10] This shows an emotion map where multiple emotions are mapped. [Figure 11] This is a sequence diagram showing the processing flow of the data processing system in Example 1. [Figure 12] This is a sequence diagram showing the processing flow of the data processing system in Application Example 1. [Figure 13] This is a sequence diagram showing the processing flow of the data processing system in Example 2, which incorporates an emotion engine. [Figure 14] This is a sequence diagram showing the processing flow of the data processing system in Application Example 2, which combines an emotion engine. [Modes for carrying out the invention]

[0013] Hereinafter, an example of an embodiment of the system relating to the technology of this disclosure will be described with reference to the attached drawings.

[0014] First, the terms used in the following description will be explained.

[0015] In the following embodiments, the numbered processor (hereinafter simply referred to as "processor") may be a single arithmetic unit or a combination of multiple arithmetic units. Also, the processor may be a single type of arithmetic unit or a combination of multiple types of arithmetic units. Examples of arithmetic units include a CPU (Central Processing Unit), a GPU (Graphics Processing Unit), a GPGPU (General-Purpose computing on Graphics Processing Units), an APU (Accelerated Processing Unit), and the like.

[0016] In the following embodiments, the numbered RAM (Random Access Memory) is a memory in which information is temporarily stored and is used as a work memory by the processor.

[0017] In the following embodiments, the numbered storage is one or more non-volatile storage devices that store various programs and various parameters, etc. Examples of non-volatile storage devices include flash memory (SSD (Solid State Drive)), magnetic disks (e.g., hard disks), or magnetic tapes, and the like.

[0018] In the following embodiments, the numbered communication I / F (Interface) is an interface including a communication processor and an antenna, etc. The communication I / F controls communication between multiple computers. Examples of communication standards applied to the communication I / F include wireless communication standards including 5G (5th Generation Mobile Communication System), Wi-Fi (registered trademark), or Bluetooth (registered trademark), and the like.

[0019] In the following embodiments, "A and / or B" is synonymous with "at least one of A and B." That is, "A and / or B" means that it may be A alone, or B alone, or a combination of A and B. Furthermore, in this specification, the same concept as "A and / or B" applies when expressing three or more things linked by "and / or."

[0020] [First Embodiment]

[0021] Figure 1 shows an example of the configuration of the data processing system 10 according to the first embodiment.

[0022] As shown in Figure 1, the data processing system 10 includes a data processing device 12 and a smart device 14. An example of the data processing device 12 is a server.

[0023] The data processing device 12 comprises a computer 22, a database 24, and a communication interface 26. The computer 22 is an example of a "computer" related to the technology of this disclosure. The computer 22 comprises a processor 28, RAM 30, and storage 32. The processor 28, RAM 30, and storage 32 are connected to a bus 34. The database 24 and the communication interface 26 are also connected to the bus 34. The communication interface 26 is connected to a network 54. An example of the network 54 is a WAN (Wide Area Network) and / or a LAN (Local Area Network).

[0024] The smart device 14 comprises a computer 36, a reception device 38, an output device 40, a camera 42, and a communication interface 44. The computer 36 comprises a processor 46, RAM 48, and storage 50. The processor 46, RAM 48, and storage 50 are connected to a bus 52. The reception device 38, output device 40, and camera 42 are also connected to the bus 52.

[0025] The reception device 38 is equipped with a touch panel 38A and a microphone 38B, etc., and receives user input. The touch panel 38A receives user input by detecting contact with an object (e.g., a pen or finger). The microphone 38B receives user input by detecting the user's voice. The control unit 46A transmits data indicating the user input received by the touch panel 38A and microphone 38B to the data processing device 12. In the data processing device 12, the specific processing unit 290 acquires the data indicating the user input.

[0026] The output device 40 includes a display 40A and a speaker 40B, and presents data to the user 20 by outputting the data in a form perceptible to the user 20 (e.g., audio and / or text). The display 40A displays visible information such as text and images according to instructions from the processor 46. The speaker 40B outputs audio according to instructions from the processor 46. The camera 42 is a small digital camera equipped with an optical system such as a lens, aperture, and shutter, and an image sensor such as a CMOS (Complementary Metal-Oxide-Semiconductor) image sensor or a CCD (Charge Coupled Device) image sensor.

[0027] Communication interface 44 is connected to network 54. Communication interfaces 44 and 26 are responsible for the exchange of various types of information between processor 46 and processor 28 via network 54.

[0028] Figure 2 shows an example of the main functions of the data processing device 12 and the smart device 14.

[0029] As shown in Figure 2, in the data processing device 12, a specific processing is performed by the processor 28. A specific processing program 56 is stored in the storage 32. The specific processing program 56 is an example of a "program" related to the technology of this disclosure. The processor 28 reads the specific processing program 56 from the storage 32 and executes the read specific processing program 56 on the RAM 30. The specific processing is realized by the processor 28 operating as a specific processing unit 290 according to the specific processing program 56 executed on the RAM 30.

[0030] The storage 32 stores the data generation model 58 and the emotion identification model 59. The data generation model 58 and the emotion identification model 59 are used by the identification processing unit 290.

[0031] In the smart device 14, the processor 46 performs the reception output processing. The storage 50 stores the reception output program 60. The reception output program 60 is used in conjunction with a specific processing program 56 by the data processing system 10. The processor 46 reads the reception output program 60 from the storage 50 and executes the read reception output program 60 on the RAM 48. The reception output processing is realized by the processor 46 operating as a control unit 46A according to the reception output program 60 executed on the RAM 48.

[0032] Next, the specific processing performed by the specific processing unit 290 of the data processing device 12 will be described. In the following description, the data processing device 12 will be referred to as the "server" and the smart device 14 as the "terminal".

[0033] This invention provides a system that monitors meetings and daily voice communications in the workplace and detects signs of power harassment in real time. This system mainly consists of voice input means, voice recognition means, emotion analysis means, alert generation means, and notification means.

[0034] First, a microphone device, acting as a terminal, collects audio data from the meeting. This data is transmitted to a server in real time. The server uses speech recognition to convert the received audio data into text data. This converted text data is further processed, and sentiment analysis is used to identify the emotional state of the speaker. In this system, emotions are classified as "positive," "negative," or "neutral" based on the words used and the tone of voice during the conversation.

[0035] Next, the server uses an alert generation mechanism based on the sentiment analysis results to generate warning information if a certain threshold is exceeded. Key indicators in this process include whether the content of the speech is aggressive or if the tone is inappropriate. The generated warning information is automatically communicated to relevant parties through notification mechanisms. For example, it may be sent to HR personnel or managers via email or online messaging platforms.

[0036] For example, suppose during a meeting, a superior says to a subordinate in an overbearing tone, "This result is completely disappointing, and it's entirely your fault." The server converts this audio into text, then uses sentiment analysis to determine that it is "aggressive" and "blame-shifting," and the alert generation mechanism determines that it has exceeded the threshold. In this case, the system sends an alert to the human resources department, enabling immediate action.

[0037] Thus, this system allows companies to quickly and effectively manage signs of harassment in order to ensure employee safety and maintain workplace mental health.

[0038] The following describes the processing flow.

[0039] Step 1:

[0040] The terminal collects audio data in real time through a microphone installed in the conference room. This audio data is immediately converted into a digital format and transmitted to a server via the network.

[0041] Step 2:

[0042] The server inputs the received audio data into the speech recognition service, which converts it from speech to text data. The speech recognition engine analyzes phonemes, intonation, and speaker characteristics to generate highly accurate text output.

[0043] Step 3:

[0044] The server passes the obtained text data to a sentiment analysis algorithm. Here, based on the content of the text, the selected words, and the context, the emotional nuance of the text is classified as "positive," "negative," "neutral," etc.

[0045] Step 4:

[0046] The server evaluates the results of sentiment analysis, sets thresholds, and makes a decision on whether to generate an alert. If the sentiment exceeds the set criteria, especially if it is determined to be aggressive or domineering, a warning message is generated.

[0047] Step 5:

[0048] The server processes the generated warning information and prepares to notify specific parties. The information includes the specific content of the statement, the analysis results, and a timestamp.

[0049] Step 6:

[0050] The server sends alert information to HR personnel and managers via notification methods such as email and online messaging platforms like Slack. This gives stakeholders an advantage in taking early action and resolving problems.

[0051] Step 7:

[0052] The user (HR department) reviews the received warning information and, if necessary, conducts a detailed investigation of the problem and interviews with those involved. They also plan and implement corrective measures and follow-up actions as needed.

[0053] (Example 1)

[0054] Next, we will describe Example 1. In the following description, the data processing device 12 will be referred to as the "server," and the smart device 14 will be referred to as the "terminal."

[0055] There is a challenge in early detection and appropriate response to potential signs of power harassment during workplace conversations and meetings. This challenge is particularly pronounced in situations where real-time monitoring and immediate notification are required.

[0056] The identification process performed by the identification processing unit 290 of the data processing device 12 in Example 1 is realized by the following means.

[0057] In this invention, the server includes a receiving device for recording audio signals, a conversion device for converting the recorded audio signals into text information, and an evaluation device for evaluating the text information and identifying emotional characteristics. This enables the real-time detection of signs of power harassment, prompt notification to those involved, and appropriate responses.

[0058] An "audio signal" is an electrical signal obtained by converting sound vibrations collected through a microphone or recording device into an electrical signal.

[0059] A "receiving device" is a hardware or software component for capturing audio signals.

[0060] A "conversion device" is a device that has the function of analyzing audio signals and converting them into text information.

[0061] "Textual information" refers to data obtained by converting audio signals into text format.

[0062] An "evaluation device" is a device that analyzes textual information and processes it to identify emotional characteristics.

[0063] "Emotional characteristics" refer to emotional traits classified based on the evaluation results of the content of statements.

[0064] A "signal generation device" is a device that has the function of generating a warning signal when certain conditions are met.

[0065] A "warning signal" is a signal generated to draw attention based on an assessment of emotional characteristics.

[0066] "Communication equipment" refers to any device or system used to transmit or receive information.

[0067] A "transmitting device" is a device that has the function of distributing the generated warning signal to a third party via a communication means.

[0068] A "command system" is a device that has the function of issuing instructions to prompt specific actions based on warning signals.

[0069] "Electronic communication" refers to the means of sending and receiving information using electronic methods.

[0070] "Network communication means" refers to means of sending and receiving information via a digital network.

[0071] This invention is a system that monitors voice communication during workplace conversations and meetings, and quickly identifies inappropriate remarks. This system operates through the cooperation of three entities: a terminal, a server, and a user.

[0072] First, the receiving device, acting as a terminal, collects audio signals using, for example, a microphone equipped with a highly sensitive acoustic sensor. The collected audio signals are recorded as digital data and immediately transmitted to the server.

[0073] Next, the server receives this audio signal and converts it into text information using a conversion device. This procedure can utilize speech recognition technologies such as Google® Cloud Speech-to-Text or Amazon Transcribe. The converted text information is then evaluated by an evaluation device to identify sentiment characteristics. In this process, a natural language processing model using the Hugging Face Transformers library is used to classify the text information into categories such as "positive," "negative," and "neutral."

[0074] After emotional characteristics are identified, the server uses a signal generator to produce warning signals as needed. For example, if aggressive remarks towards others are detected, the system will issue a warning signal based on predefined criteria. These criteria include rules that are triggered when a threshold is exceeded. The generated warning signals are quickly notified to users and administrators using communication devices. Specifically, immediate message transmissions are made via Slack or MICROSOFT® TEAMS®.

[0075] This system enables users to address workplace environment issues quickly and accurately.

[0076] As a concrete example, consider a situation where a superior aggressively tells a subordinate during a meeting, "This result is completely disappointing, and it's entirely your fault." In this case, the server converts the audio data into text, performs sentiment analysis to determine that the statement is "aggressive" and "blame-shifting," and immediately sends a warning to the HR department. An example of a prompt for this system would be: "We want to analyze signs of power harassment from audio communication during workplace meetings. We want to classify emotions from the audio data and issue real-time alerts if aggressive remarks are made. Please detail the necessary steps and technologies to use to achieve this process."

[0077] The flow of the specific processing in Example 1 will be explained using Figure 11.

[0078] Step 1:

[0079] The receiving device, acting as a terminal, operates and collects audio signals from ambient noise during the meeting. During this process, the terminal performs filtering to reduce noise within the meeting room, resulting in clear audio data. The input data is the analog audio signal picked up by the microphone device, while the output data is the signal converted to a digital audio format.

[0080] Step 2:

[0081] The terminal transmits the collected digital audio signal to the server in real time. This process requires a stable wireless or wired network connection and minimizes data delay using a certain amount of buffering. The input data is the digital audio signal acquired in the previous step, and the output data is the same audio signal received on the server side.

[0082] Step 3:

[0083] The server activates a speech recognition device to convert the received digital audio signal into text information. The speech recognition software uses an existing model (e.g., Google Cloud Speech-to-Text) and configures language profiles to improve recognition accuracy. The input data is an audio signal, and the output data is the text data converted from that audio signal.

[0084] Step 4:

[0085] The server passes the converted text data to an evaluation device to identify sentiment characteristics. This evaluation process uses a natural language processing model (e.g., BERT) to perform sentiment analysis on the text and classify it as "positive," "negative," or "neutral." The input data is textual information, and the output data is the classified sentiment characteristics.

[0086] Step 5:

[0087] The server processes data with emotional characteristics using a signal generator and generates a warning signal if it exceeds a pre-set threshold. For example, if negative emotions are pronounced, it determines the warning level and prepares the necessary alarm message. The input data is the result of the emotional characteristics evaluation, and the output data is the message used as the warning signal.

[0088] Step 6:

[0089] The server notifies users and relevant parties of the generated warning signal via communication devices. Email and online messaging platforms are used as notification methods, and multiple notification methods are used simultaneously when immediate action is required. The input data is the warning signal, and the output data is the notification message delivered to the user's terminal.

[0090] (Application Example 1)

[0091] Next, we will explain Application Example 1. In the following explanation, the data processing device 12 will be referred to as the "server," and the smart device 14 will be referred to as the "terminal."

[0092] In the workplace environment, particularly in manufacturing and other industrial settings, inappropriate communication among employees can negatively impact work efficiency and workplace atmosphere. This can potentially impair employees' mental health and work performance. Preventing and improving such problems is therefore crucial.

[0093] The specific processing performed by the specific processing unit 290 of the data processing device 12 in Application Example 1 is realized by the following means.

[0094] In this invention, the server includes an acoustic input means for acquiring acoustic information, an acoustic recognition means for converting speech into text information based on the acquired acoustic information, and an emotion analysis means for analyzing the text information and extracting emotion information. This enables real-time monitoring of communication in the work environment, immediate detection of problematic conversations, and appropriate responses.

[0095] "Acoustic information" refers to all data related to speech and sound waves, and forms the basis for speech recognition and analysis.

[0096] "Acoustic input means" refers to a device or part of a device that has the function of collecting acoustic information. Specifically, this includes sensors such as microphones.

[0097] "Acoustic recognition means" refers to a function that converts acquired acoustic information into text information, and is a technology for converting audio data into text data.

[0098] "Textual information" refers to the text data that results from the recognition and conversion of audio information.

[0099] "Emotional analysis tools" refer to functions that extract and analyze the speaker's emotional state from textual information. They perform emotional evaluations such as positive, negative, and neutral.

[0100] "Warning information" refers to information generated as a result of sentiment analysis to alert users when certain conditions are met.

[0101] "Notification means" refers to a device or method that has the function of transmitting generated warning information or other information to relevant parties.

[0102] "Mechanical equipment" refers to devices and facilities designed to operate automatically and perform specified functions. Factory robots are a concrete example.

[0103] "Information transmission means" refers to the technologies and methods used to transmit information to designated recipients.

[0104] The system implementing this invention is primarily intended for the detection and management of inappropriate communication in the workplace environment. Specific embodiments of this system are described below.

[0105] The server acquires acoustic information using microphones installed on factory robots. This acoustic information obtained through the acoustic input means is transmitted to the server. The acoustic information is converted into text information on the server using acoustic recognition means. In this step, a speech recognition API such as Google Cloud Speech-to-Text is used. This converts the speech into text data.

[0106] Next, the converted text information is analyzed through sentiment analysis tools. This process utilizes sentiment analysis libraries such as Microsoft Azure® Text Analytics. The server uses this sentiment analysis to extract the speaker's emotional state from the text information and performs sentiment evaluations such as positive, negative, or neutral.

[0107] Subsequently, based on the sentiment analysis results, the warning generation system generates warning information if specific conditions are met. This warning information is then communicated to relevant parties via email or digital messaging using the notification system.

[0108] For example, if a tense exchange occurs between a supervisor and a subordinate while machinery is operating in a factory, the audio is collected, and if sentiment analysis determines it to be "negative," a warning is generated. This warning information is then sent to the administrator to help with preventative measures.

[0109] An example of a prompt message that can be given to a generative AI model is, "Please analyze the sentiment of the following text and determine whether it is positive or negative: text."

[0110] Thus, through this invention, communication in the workplace environment can be monitored in real time, enabling early detection and response to problems.

[0111] The flow of a specific process in Application Example 1 will be explained using Figure 12.

[0112] Step 1:

[0113] Microphones equipped on factory robots collect acoustic information from the surroundings. Audio data is obtained as input and sent to a server for processing in real time.

[0114] Step 2:

[0115] The server converts the received audio data into text information using a speech recognition API such as Google Cloud Speech-to-Text. In this step, the audio information, which is the input data, is output as text data. The specific actions of recognizing the audio signal and converting it into text format are performed.

[0116] Step 3:

[0117] The converted text information is analyzed by sentiment analysis tools on the server. Using tools such as Microsoft Azure Text Analytics, sentiment information is extracted, and a sentiment rating (positive, negative, neutral, etc.) is output based on the input text. Specifically, the emotional value of each word and phrase is analyzed to determine the overall sentiment.

[0118] Step 4:

[0119] The server uses a warning generation mechanism based on the sentiment analysis results to generate warning information if the conditions are met. In this step, sentiment information is used as input, and warning information is output if it exceeds a threshold. Specific actions are taken to generate a warning when aggressive remarks or negative tones are detected.

[0120] Step 5:

[0121] The generated warning information is communicated to relevant parties via notification methods. The input is the warning information, and the output is the sending of emails or digital messages. Specifically, the warning notification is delivered to the administrator's mailbox or messaging platform.

[0122] Furthermore, an emotion engine that estimates the user's emotions may be incorporated. That is, the identification processing unit 290 may use the emotion identification model 59 to estimate the user's emotions and perform identification processing using the user's emotions.

[0123] This invention provides a system that incorporates an emotion engine to recognize user emotions when analyzing voice data. The purpose is to monitor workplace meetings and communications, detect signs of power harassment, and understand individual emotional states in real time.

[0124] As a terminal, a microphone device installed in the conference room collects participants' conversations in real time. The acquired audio data is immediately transmitted to the server. The server uses speech recognition to convert the audio data into text data. Next, sentiment analysis is performed on this text data to extract the speaker's emotional information.

[0125] Furthermore, the server incorporates an emotion engine that analyzes the user's voice tone, speed, and word choice to recognize their emotional state in real time. This information is recorded in a database and used for future analysis.

[0126] As a concrete example, suppose in a meeting, a superior tells a subordinate in a harsh tone, "Failure of this project is unacceptable." The server converts this audio into text and, through sentiment analysis and an emotion engine, determines that the statement is "aggressive" and that the speaker's emotion is "frustrated." Based on this, the alert generation mechanism determines that a threshold has been exceeded and generates a warning.

[0127] The generated warning information is automatically sent to HR personnel and managers via email or online messaging platforms. The user (HR department) uses this information to plan early intervention measures such as interviews and counseling.

[0128] Thus, by using an emotion engine, this system can analyze employees' emotions more precisely and manage workplace problems quickly and effectively. As a result, it can contribute to maintaining mental health in the workplace and improving the working environment.

[0129] The following describes the processing flow.

[0130] Step 1:

[0131] The terminal collects participants' conversations in real time through microphones installed in the conference room. The collected audio data is immediately converted into a digital signal and sent to the server.

[0132] Step 2:

[0133] The server processes the received audio data using speech recognition technology and converts the audio into text data. This conversion is achieved by the speech recognition engine analyzing the phonemes and intonation of the speech.

[0134] Step 3:

[0135] The server inputs the converted text data into a sentiment analysis system, which extracts the speaker's emotional information from the text. Sentiment analysis classifies the text into emotional categories such as "positive" or "negative" based on keywords and context.

[0136] Step 4:

[0137] The server, along with the results of sentiment analysis, uses an emotion engine to analyze the user's voice tone, speed, and word choice in real time, gaining a more detailed understanding of the emotions being expressed during speech. This analysis helps to grasp changes in emotions and their intensity.

[0138] Step 5:

[0139] The server sets thresholds based on the emotion engine and emotion analysis results, and performs evaluation using an alert generation mechanism. If aggressive or inappropriate emotions exceeding a specific threshold are detected during this evaluation, warning information is generated.

[0140] Step 6:

[0141] The server sends the generated warning information to the relevant parties via a notification system using pre-configured communication tools (e.g., email, online messaging platform). The notification includes the content of the message and the sentiment analysis results.

[0142] Step 7:

[0143] The user (HR department) reviews the submitted alert information and evaluates the identification of the problem and the response measures for the victim. If necessary, they develop and implement interviews and mental health support measures. These measures include providing feedback for improvement and implementing training programs.

[0144] (Example 2)

[0145] Next, we will describe Example 2. In the following description, the data processing device 12 will be referred to as the "server" and the smart device 14 as the "terminal".

[0146] In today's workplace, communication-related friction and mental health issues are on the rise. In particular, it is difficult to detect emotional triggers hidden in meetings and daily conversations early and to take appropriate action. Therefore, there is a need for a system that can monitor employees' emotional states in real time, alert them to potential problems early, and enable them to take countermeasures.

[0147] The identification process performed by the identification processing unit 290 of the data processing device 12 in Example 2 is realized by the following means.

[0148] In this invention, the server includes an input means for acquiring an audio signal, a recognition means for converting the acquired audio signal into text information, an analysis means for analyzing the text information and extracting emotion data, and an emotion recognition means for recognizing the emotional state in real time based on the characteristics of the voice and recording it in a database. This makes it possible to automatically detect emotional triggers hidden in communication and quickly notify relevant parties of alarm information.

[0149] An "audio signal" is data that converts air vibrations into electrical signals, and includes conversations and sounds.

[0150] "Input means" refers to devices or equipment used to acquire audio signals, such as microphones.

[0151] "Recognition means" refers to technologies and systems for converting audio signals into text information, and includes speech recognition software.

[0152] "Textual information" refers to text data generated based on audio signals, representing the content of the audio as written characters.

[0153] "Analysis means" refers to algorithms and programs used to analyze textual information and extract emotional data.

[0154] "Emotional data" refers to information extracted from textual information and audio characteristics, indicating the speaker's emotional state and type of emotion.

[0155] An "alert generation method" refers to a system or technology that automatically creates alarm information based on emotional data.

[0156] "Notification means" refers to methods and technologies for transmitting generated alarm information to relevant parties, and includes electronic communication means and messaging systems.

[0157] "Emotion recognition means" refers to systems or processes for understanding a speaker's emotions in real time from audio signals or text information.

[0158] This invention provides a system that detects potential emotional triggers early in workplace communication and prompts appropriate responses. This system is realized through the integrated involvement of terminals, servers, and users.

[0159] The device used will be a microphone device installed in the conference room or workplace. This microphone device can effectively capture the participants' voice signals and collect clear audio data using noise reduction technology. Specifically, it is desirable to use a microphone with noise-canceling capabilities.

[0160] The server receives the collected audio signals and converts them into text information using speech recognition software. For example, cloud-based speech recognition services can be used, such as Google Cloud Speech-to-Text. The converted text information is then analyzed using natural language processing techniques to extract sentiment data. The analysis employs algorithms utilizing generative AI models, such as BERT and GPT-3(registered trademark).

[0161] Based on emotion data, the server analyzes the characteristics of the voice in the audio, such as tone and speed, and evaluates the speaker's emotions in real time through emotion recognition means. This recognized emotion information is recorded in a database and stored for further analysis and future responses. The server also uses an alert generation means to automatically generate an alarm when the emotion data exceeds a certain threshold.

[0162] Alarm information is sent to users via electronic communication or online messaging systems through notification means. Examples of such means include email and online messaging platforms (such as Slack or Microsoft Teams).

[0163] Based on the received alert information, users such as HR departments and managers can quickly take appropriate countermeasures, such as interviews and counseling. As a concrete example, a prompt sentence to be input into the generating AI model could be something like, "Determine the speaker's emotions from the audio data spoken during the meeting, and propose countermeasures if an emotional trigger occurs."

[0164] Thus, by automating the collection, conversion, and analysis of voice data, the present invention enables smoother communication in the workplace and early problem resolution.

[0165] The flow of the specific processing in Example 2 will be explained using Figure 13.

[0166] Step 1:

[0167] The terminal uses a microphone device installed in the conference room to acquire participants' audio signals in real time. The microphone reduces ambient noise and ensures clear audio capture. The acquired audio signals are temporarily stored as digital data within the terminal.

[0168] Step 2:

[0169] The terminal sends the collected audio signal to the server. The server receives this audio signal using a secure communication protocol. The received digital audio signal is converted into text data using speech recognition software. This conversion analyzes the waveform data of the audio signal and replaces it with the corresponding string.

[0170] Step 3:

[0171] The server uses natural language processing techniques to extract sentiment data from the converted text data. A generative AI model is used to understand the context of the text and determine the speaker's emotions. Sentiment labels such as "joy" and "anger" are output from the input text data.

[0172] Step 4:

[0173] The server uses emotion recognition technology to analyze the characteristics of the audio signal, also evaluating the speaker's tone and speed. This evaluation combines supplementary information with text-based emotion data, allowing for real-time recognition of the emotional state. The analysis results are recorded in a database as the emotional state.

[0174] Step 5:

[0175] The server uses an alert generation mechanism based on emotional data and voice characteristics information to generate an alarm when a certain threshold is exceeded. For example, if a particular statement matches an aggressive tone, alarm data is generated. This alarm data is prepared for notification to relevant parties.

[0176] Step 6:

[0177] The notification method involves sending the generated alarm data to the user via email or online messaging system. The communication method used is selected by the system settings. Upon receiving this alarm notification, the Human Resources Department plans and prepares to implement appropriate countermeasures.

[0178] (Application Example 2)

[0179] Next, we will explain application example 2. In the following explanation, the data processing device 12 will be referred to as a "server" and the smart device 14 as a "terminal".

[0180] Preventing troubles and discords in commercial facilities and offices is crucial, but conventional security systems have difficulty providing early warnings based on emotional shifts. Therefore, there is a need to monitor emotional escalations that may occur between users in real time and respond appropriately.

[0181] The specific processing performed by the specific processing unit 290 of the data processing device 12 in Application Example 2 is realized by the following means.

[0182] In this invention, the server includes voice input means for acquiring voice data, voice recognition means for converting voice into text data based on the acquired voice data, and emotion analysis means for analyzing the text data and extracting emotion information. This makes it possible to identify signs of discomfort or trouble in real time based on emotion information and generate warning information.

[0183] "Voice input means" refers to a device or system that acquires voice data.

[0184] "Speech recognition means" refers to a technology or device that converts acquired speech data into text data.

[0185] "Emotional analysis means" refers to a method or device that analyzes text data and extracts emotional information from it.

[0186] An "alert generation means" is a mechanism or function that generates warning information when specific conditions are met based on emotional information.

[0187] "Notification means" refers to a system or function that transmits generated warning information to a specific recipient using communication means.

[0188] A "control means" is a means for executing specific instructions based on warning information.

[0189] This invention is for constructing a security system based on emotion analysis. First, the entire system includes a voice input means for acquiring voice data. The voice input means consists of a device that collects voices in commercial facilities or offices. This device can be a smartphone or a dedicated voice capture device.

[0190] The server acts as a speech recognition system. This system converts the acquired speech data into text data. The software used here is the speech_recognition library for Python. After this, the text data is analyzed by an emotion analysis system. This analysis system incorporates a specific algorithm for identifying emotions, and the emotion_recognition library is used for this purpose.

[0191] The server also functions as an alert generation mechanism. Based on the sentiment analysis, if the received sentiment information meets certain conditions, the system automatically generates a warning. This warning is generated when it detects discomfort or signs of trouble that exceed a pre-set threshold.

[0192] The generated warning information is sent to security guards or personnel via notification methods. These notification methods include email and the communication functions of online messaging platforms.

[0193] As a concrete example, consider a scenario where signs of an impending argument between customers in a shopping mall are detected. This system quickly performs sentiment analysis when the volume of voices or aggressive language increases. It then notifies security personnel that "caution is needed."

[0194] The following prompt statements can be used with the generative AI model.

[0195] "From this audio data, we have detected changes in emotion. In particular, we observed 'anger' or 'fear,' which strongly suggest the need for security. We recommend the following steps." This allows for a swift and appropriate response.

[0196] The flow of a specific process in Application Example 2 will be explained using Figure 14.

[0197] Step 1:

[0198] The terminal acquires audio from commercial facilities and offices using an audio input device. The audio data is collected in real time from a microphone device. The input consists of ambient sounds and people's conversations, and the output is an audio file in digital format.

[0199] Step 2:

[0200] The server processes the acquired audio data using speech recognition. In this step, the audio data is converted into text data. The input is digital audio data, and the output is text data. Specifically, speech recognition is performed using the Python speech_recognition library.

[0201] Step 3:

[0202] The server analyzes text data using sentiment analysis tools. The input is text data, and the output is sentiment information. At this stage, the emotion_recognition library is used to identify emotions from the wording and tone contained in the text. Specifically, it analyzes keywords and phrases in the text to determine a particular emotion (e.g., "anger," "sadness").

[0203] Step 4:

[0204] The server generates warning information using an alert generation mechanism based on emotional information. The input is emotional information obtained through emotional analysis, and the output is warning information. Specifically, if the detected emotion exceeds a pre-set threshold, a warning state is confirmed. The server generates an alarm based on this information.

[0205] Step 5:

[0206] The server sends the generated warning information to the relevant parties through notification channels. The input is the warning information, and the output is a notification message. In this process, warnings are sent to security guards and administrators using email or messaging platforms. Specifically, the warning content is distributed to pre-configured email addresses or chat groups.

[0207] The specific processing unit 290 transmits the result of the specific processing to the smart device 14. In the smart device 14, the control unit 46A causes the output device 40 to output the result of the specific processing. The microphone 38B acquires audio indicating user input for the result of the specific processing. The control unit 46A transmits the audio data indicating user input acquired by the microphone 38B to the data processing device 12. In the data processing device 12, the specific processing unit 290 acquires the audio data.

[0208] Data generation model 58 is a so-called generative AI (Artificial Intelligence). An example of data generation model 58 is ChatGPT (registered trademark) (Internet search).<URL: https: / / openai.com / blog / chatgpt> ), Gemini (registered trademark) (Internet search) <url: https: gemini.google.com ?hl="ja">Examples of generative AI include the following. The data generation model 58 is obtained by performing deep learning on a neural network. The data generation model 58 is input with prompts containing instructions, and with inference data such as audio data representing speech, text data representing text, and image data representing images. The data generation model 58 infers from the input inference data according to the instructions indicated by the prompts, and outputs the inference results in data formats such as audio data and text data. Here, inference refers to, for example, analysis, classification, prediction, and / or summarization.

[0209] In the above embodiment, an example was given in which specific processing is performed by the data processing device 12, but the technology of this disclosure is not limited thereto, and the specific processing may also be performed by the smart device 14.

[0210] [Second Embodiment]

[0211] Figure 3 shows an example of the configuration of the data processing system 210 according to the second embodiment.

[0212] As shown in Figure 3, the data processing system 210 includes a data processing device 12 and smart glasses 214. An example of the data processing device 12 is a server.

[0213] The data processing device 12 comprises a computer 22, a database 24, and a communication interface 26. The computer 22 is an example of a "computer" related to the technology of this disclosure. The computer 22 comprises a processor 28, RAM 30, and storage 32. The processor 28, RAM 30, and storage 32 are connected to a bus 34. The database 24 and the communication interface 26 are also connected to the bus 34. The communication interface 26 is connected to a network 54. An example of the network 54 is a WAN (Wide Area Network) and / or a LAN (Local Area Network).

[0214] The smart glasses 214 include a computer 36, a microphone 238, a speaker 240, a camera 42, and a communication interface 44. The computer 36 includes a processor 46, RAM 48, and storage 50. The processor 46, RAM 48, and storage 50 are connected to a bus 52. The microphone 238, speaker 240, and camera 42 are also connected to the bus 52.

[0215] The microphone 238 receives voice signals from the user 20 and receives instructions from the user 20. The microphone 238 captures the voice signals from the user 20, converts the captured voice into audio data, and outputs it to the processor 46. The speaker 240 outputs audio according to the instructions from the processor 46.

[0216] Camera 42 is a small digital camera equipped with an optical system including a lens, aperture, and shutter, and an image sensor such as a CMOS (Complementary Metal-Oxide-Semiconductor) image sensor or a CCD (Charge Coupled Device) image sensor, and captures images of the area around the user 20 (for example, an imaging range defined by a field of view equivalent to the width of a typical healthy person's field of vision).

[0217] Communication interface 44 is connected to network 54. Communication interfaces 44 and 26 are responsible for the exchange of various information between processor 46 and processor 28 via network 54. The exchange of various information between processor 46 and processor 28 using communication interfaces 44 and 26 is performed in a secure manner.

[0218] Figure 4 shows an example of the main functions of the data processing device 12 and the smart glasses 214. As shown in Figure 4, the data processing device 12 performs specific processing using the processor 28. The storage 32 stores the specific processing program 56.

[0219] The specific processing program 56 is an example of a "program" relating to the technology of this disclosure. The processor 28 reads the specific processing program 56 from the storage 32 and executes the read specific processing program 56 on the RAM 30. The specific processing is realized by the processor 28 operating as a specific processing unit 290 in accordance with the specific processing program 56 executed on the RAM 30.

[0220] The storage 32 stores the data generation model 58 and the emotion identification model 59. The data generation model 58 and the emotion identification model 59 are used by the identification processing unit 290.

[0221] In the smart glasses 214, the processor 46 performs the reception output processing. The storage 50 stores the reception output program 60. The processor 46 reads the reception output program 60 from the storage 50 and executes the read reception output program 60 on the RAM 48. The reception output processing is realized by the processor 46 operating as a control unit 46A according to the reception output program 60 executed on the RAM 48.

[0222] Next, the identification processing performed by the identification processing unit 290 of the data processing device 12 will be described. In the following description, the data processing device 12 will be referred to as the "server" and the smart glasses 214 will be referred to as the "terminal".

[0223] This invention provides a system that monitors meetings and daily voice communications in the workplace and detects signs of power harassment in real time. This system mainly consists of voice input means, voice recognition means, emotion analysis means, alert generation means, and notification means.

[0224] First, a microphone device, acting as a terminal, collects audio data from the meeting. This data is transmitted to a server in real time. The server uses speech recognition to convert the received audio data into text data. This converted text data is further processed, and sentiment analysis is used to identify the emotional state of the speaker. In this system, emotions are classified as "positive," "negative," or "neutral" based on the words used and the tone of voice during the conversation.

[0225] Next, the server uses an alert generation mechanism based on the sentiment analysis results to generate warning information if a certain threshold is exceeded. Key indicators in this process include whether the content of the speech is aggressive or if the tone is inappropriate. The generated warning information is automatically communicated to relevant parties through notification mechanisms. For example, it may be sent to HR personnel or managers via email or online messaging platforms.

[0226] For example, suppose during a meeting, a superior says to a subordinate in an overbearing tone, "This result is completely disappointing, and it's entirely your fault." The server converts this audio into text, then uses sentiment analysis to determine that it is "aggressive" and "blame-shifting," and the alert generation mechanism determines that it has exceeded the threshold. In this case, the system sends an alert to the human resources department, enabling immediate action.

[0227] Thus, this system allows companies to quickly and effectively manage signs of harassment in order to ensure employee safety and maintain workplace mental health.

[0228] The following describes the processing flow.

[0229] Step 1:

[0230] The terminal collects audio data in real time through a microphone installed in the conference room. This audio data is immediately converted into a digital format and transmitted to a server via the network.

[0231] Step 2:

[0232] The server inputs the received audio data into the speech recognition service, which converts it from speech to text data. The speech recognition engine analyzes phonemes, intonation, and speaker characteristics to generate highly accurate text output.

[0233] Step 3:

[0234] The server passes the obtained text data to a sentiment analysis algorithm. Here, based on the content of the text, the selected words, and the context, the emotional nuance of the text is classified as "positive," "negative," "neutral," etc.

[0235] Step 4:

[0236] The server evaluates the results of sentiment analysis, sets thresholds, and makes a decision on whether to generate an alert. If the sentiment exceeds the set criteria, especially if it is determined to be aggressive or domineering, a warning message is generated.

[0237] Step 5:

[0238] The server processes the generated warning information and prepares to notify specific parties. The information includes the specific content of the statement, the analysis results, and a timestamp.

[0239] Step 6:

[0240] The server sends alert information to HR personnel and managers via notification methods such as email and online messaging platforms like Slack. This gives stakeholders an advantage in taking early action and resolving problems.

[0241] Step 7:

[0242] The user (HR department) reviews the received warning information and, if necessary, conducts a detailed investigation of the problem and interviews with those involved. They also plan and implement corrective measures and follow-up actions as needed.

[0243] (Example 1)

[0244] Next, we will describe Example 1. In the following description, the data processing device 12 will be referred to as the "server," and the smart glasses 214 will be referred to as the "terminal."

[0245] There is a challenge in early detection and appropriate response to potential signs of power harassment during workplace conversations and meetings. This challenge is particularly pronounced in situations where real-time monitoring and immediate notification are required.

[0246] The identification process performed by the identification processing unit 290 of the data processing device 12 in Example 1 is realized by the following means.

[0247] In this invention, the server includes a receiving device for recording audio signals, a conversion device for converting the recorded audio signals into text information, and an evaluation device for evaluating the text information and identifying emotional characteristics. This enables the real-time detection of signs of power harassment, prompt notification to those involved, and appropriate responses.

[0248] An "audio signal" is an electrical signal obtained by converting sound vibrations collected through a microphone or recording device into an electrical signal.

[0249] A "receiving device" is a hardware or software component for capturing audio signals.

[0250] A "conversion device" is a device that has the function of analyzing audio signals and converting them into text information.

[0251] "Textual information" refers to data obtained by converting audio signals into text format.

[0252] An "evaluation device" is a device that analyzes textual information and processes it to identify emotional characteristics.

[0253] "Emotional characteristics" refer to emotional traits classified based on the evaluation results of the content of statements.

[0254] A "signal generation device" is a device that has the function of generating a warning signal when certain conditions are met.

[0255] A "warning signal" is a signal generated to draw attention based on an assessment of emotional characteristics.

[0256] "Communication equipment" refers to any device or system used to transmit or receive information.

[0257] A "transmitting device" is a device that has the function of distributing the generated warning signal to a third party via a communication means.

[0258] A "command system" is a device that has the function of issuing instructions to prompt specific actions based on warning signals.

[0259] "Electronic communication" refers to the means of sending and receiving information using electronic methods.

[0260] "Network communication means" refers to means of sending and receiving information via a digital network.

[0261] This invention is a system that monitors voice communication during workplace conversations and meetings, and quickly identifies inappropriate remarks. This system operates through the cooperation of three entities: a terminal, a server, and a user.

[0262] First, the receiving device, acting as a terminal, collects audio signals using, for example, a microphone equipped with a highly sensitive acoustic sensor. The collected audio signals are recorded as digital data and immediately transmitted to the server.

[0263] Next, the server receives this audio signal and converts it into text information using a conversion device. This procedure can utilize speech recognition technologies such as Google Cloud Speech-to-Text or Amazon Transcribe. The converted text information is then evaluated by an evaluation device to identify sentiment characteristics. In this process, a natural language processing model using the Hugging Face Transformers library is employed to classify the text information into categories such as "positive," "negative," and "neutral."

[0264] After emotional characteristics are identified, the server uses a signal generator to produce warning signals as needed. For example, if aggressive remarks towards others are detected, the system will issue a warning signal based on predefined criteria. These criteria include rules that are triggered when a threshold is exceeded. The generated warning signals are quickly notified to users and administrators using communication devices. Specifically, immediate messages are sent via Slack or Microsoft Teams.

[0265] This system enables users to address workplace environment issues quickly and accurately.

[0266] As a concrete example, consider a situation where a superior aggressively tells a subordinate during a meeting, "This result is completely disappointing, and it's entirely your fault." In this case, the server converts the audio data into text, performs sentiment analysis to determine that the statement is "aggressive" and "blame-shifting," and immediately sends a warning to the HR department. An example of a prompt for this system would be: "We want to analyze signs of power harassment from audio communication during workplace meetings. We want to classify emotions from the audio data and issue real-time alerts if aggressive remarks are made. Please detail the necessary steps and technologies to use to achieve this process."

[0267] The flow of the specific processing in Example 1 will be explained using Figure 11.

[0268] Step 1:

[0269] The receiving device, acting as a terminal, operates and collects audio signals from ambient noise during the meeting. During this process, the terminal performs filtering to reduce noise within the meeting room, resulting in clear audio data. The input data is the analog audio signal picked up by the microphone device, while the output data is the signal converted to a digital audio format.

[0270] Step 2:

[0271] The terminal transmits the collected digital audio signal to the server in real time. This process requires a stable wireless or wired network connection and minimizes data delay using a certain amount of buffering. The input data is the digital audio signal acquired in the previous step, and the output data is the same audio signal received on the server side.

[0272] Step 3:

[0273] The server activates a speech recognition device to convert the received digital audio signal into text information. The speech recognition software uses an existing model (e.g., Google Cloud Speech-to-Text) and configures language profiles to improve recognition accuracy. The input data is an audio signal, and the output data is the text data converted from that audio signal.

[0274] Step 4:

[0275] The server passes the converted text data to an evaluation device to identify sentiment characteristics. This evaluation process uses a natural language processing model (e.g., BERT) to perform sentiment analysis on the text and classify it as "positive," "negative," or "neutral." The input data is textual information, and the output data is the classified sentiment characteristics.

[0276] Step 5:

[0277] The server processes data with emotional characteristics by a signal generation device and generates a warning signal when it exceeds a pre-set standard. For example, if the negative emotion is significant, it determines the warning level and prepares the necessary warning message. The input data is the evaluation result of emotional characteristics, and the output data is the message as a warning signal.

[0278] Step 6:

[0279] The server notifies the user or relevant parties of the generated warning signal via a communication device. As notification means, email or an online messaging platform is used, and when immediate response is required, multiple notification means are used simultaneously. The input data is the warning signal, and the output data is the notification message delivered to the user terminal.

[0280] (Application Example 1)

[0281] Next, Application Example 1 will be described. In the following description, the data processing device 12 is referred to as the "server", and the smart glasses 214 are referred to as the "terminal".

[0282] In a workplace environment, especially at a business site such as manufacturing, inappropriate communication among employees may have an adverse impact on work efficiency and the workplace atmosphere. As a result, the mental health and work performance of employees may be impaired. It is required to prevent and improve such problems.

[0283] The specific processing by the specific processing unit 290 of the data processing device 12 in Application Example 1 is realized by the following respective means.

[0284] In this invention, the server includes an acoustic input means for acquiring acoustic information, an acoustic recognition means for converting voice into character information based on the acquired acoustic information, and an emotion analysis means for analyzing the character information and extracting emotion information. Thereby, the communication in the business environment can be monitored in real time, a problematic conversation can be immediately detected, and appropriate countermeasures can be taken.

[0285] "Acoustic information" refers to all data related to voice and sound waves, which is the basis for voice recognition and analysis.

[0286] "Acoustic input means" refers to a device or a part of a device that has the function of collecting acoustic information. Specifically, it includes sensors such as microphones.

[0287] "Acoustic recognition means" refers to the function of converting the acquired acoustic information into character information, which is a technology for converting voice data into text data.

[0288] "Character information" refers to the text data that is the result of the recognition and conversion of acoustic information.

[0289] "Emotion analysis means" refers to the function of extracting and analyzing the speaker's emotional state from character information. It performs emotion evaluations such as positive, negative, and neutral.

[0290] "Warning information" refers to the information for alerting generated when specific conditions are met as a result of emotion analysis.

[0291] "Notification means" refers to a device or method that has the function of transmitting the generated warning information and other information to relevant parties.

[0292] "Mechanical device" refers to instruments and equipment that operate automatically and are designed to perform specified functions. A factory robot is a specific example.

[0293] "Information transmission means" refers to technologies and methods for transmitting information to specified recipients.

[0294] The system for implementing the present invention is mainly aimed at detecting and managing inappropriate communication in the workplace environment. The specific embodiments of this system will be described below.

[0295] The server acquires acoustic information using microphones installed on factory robots. This acoustic information obtained through the acoustic input means is transmitted to the server. The acoustic information is converted into text information on the server using acoustic recognition means. In this step, a speech recognition API such as Google Cloud Speech-to-Text is used. This converts the speech into text data.

[0296] Next, the converted text information is analyzed through sentiment analysis tools. This process utilizes sentiment analysis libraries such as Microsoft Azure Text Analytics. The server uses this sentiment analysis to extract the speaker's emotional state from the text information and performs sentiment evaluations such as positive, negative, or neutral.

[0297] Subsequently, based on the sentiment analysis results, the warning generation system generates warning information if specific conditions are met. This warning information is then communicated to relevant parties via email or digital messaging using the notification system.

[0298] For example, if a tense exchange occurs between a supervisor and a subordinate while machinery is operating in a factory, the audio is collected, and if sentiment analysis determines it to be "negative," a warning is generated. This warning information is then sent to the administrator to help with preventative measures.

[0299] An example of a prompt message that can be given to a generative AI model is, "Please analyze the sentiment of the following text and determine whether it is positive or negative: text."

[0300] Thus, through this invention, communication in the workplace environment can be monitored in real time, enabling early detection and response to problems.

[0301] The flow of a specific process in Application Example 1 will be explained using Figure 12.

[0302] Step 1:

[0303] A microphone equipped on an industrial robot collects ambient acoustic information. Voice data is obtained as input, and this data is transmitted to a server for real-time processing.

[0304] Step 2:

[0305] The server converts the received voice data into character information using a speech recognition API such as Google Cloud Speech-to-Text. In this step, voice information as input data is output as character data. Specific operations are performed where the voice signal is recognized and converted into text format.

[0306] Step 3:

[0307] The converted character information is analyzed by sentiment analysis means on the server. Tools such as Microsoft Azure Text Analytics are used to extract sentiment information, and based on the input text, sentiment evaluations such as positive, negative, and neutral are output. Specifically, the sentiment value of each word and phrase is analyzed to determine the overall sentiment.

[0308] Step 4:

[0309] The server uses warning generation means from the results of sentiment analysis to generate warning information when conditions are met. In this step, sentiment information is used as input, and warning information is output when a threshold is exceeded. Specific operations are performed where a warning is generated when an aggressive statement or negative tone is detected.

[0310] Step 5:

[0311] The generated warning information is transmitted to relevant parties by notification means. The input is the warning information, and emails and digital messages are sent as output. As specific operations, warning notifications are delivered to the administrator's mailbox and messaging platform.

[0312] Furthermore, an emotion engine that estimates the user's emotions may be incorporated. That is, the identification processing unit 290 may use the emotion identification model 59 to estimate the user's emotions and perform identification processing using the user's emotions.

[0313] This invention provides a system that incorporates an emotion engine to recognize user emotions when analyzing voice data. The purpose is to monitor workplace meetings and communications, detect signs of power harassment, and understand individual emotional states in real time.

[0314] As a terminal, a microphone device installed in the conference room collects participants' conversations in real time. The acquired audio data is immediately transmitted to the server. The server uses speech recognition to convert the audio data into text data. Next, sentiment analysis is performed on this text data to extract the speaker's emotional information.

[0315] Furthermore, the server incorporates an emotion engine that analyzes the user's voice tone, speed, and word choice to recognize their emotional state in real time. This information is recorded in a database and used for future analysis.

[0316] As a concrete example, suppose in a meeting, a superior tells a subordinate in a harsh tone, "Failure of this project is unacceptable." The server converts this audio into text and, through sentiment analysis and an emotion engine, determines that the statement is "aggressive" and that the speaker's emotion is "frustrated." Based on this, the alert generation mechanism determines that a threshold has been exceeded and generates a warning.

[0317] The generated warning information is automatically sent to HR personnel and managers via email or online messaging platforms. The user (HR department) uses this information to plan early intervention measures such as interviews and counseling.

[0318] Thus, by using an emotion engine, this system can analyze employees' emotions more precisely and manage workplace problems quickly and effectively. As a result, it can contribute to maintaining mental health in the workplace and improving the working environment.

[0319] The following describes the processing flow.

[0320] Step 1:

[0321] The terminal collects participants' conversations in real time through microphones installed in the conference room. The collected audio data is immediately converted into a digital signal and sent to the server.

[0322] Step 2:

[0323] The server processes the received audio data using speech recognition technology and converts the audio into text data. This conversion is achieved by the speech recognition engine analyzing the phonemes and intonation of the speech.

[0324] Step 3:

[0325] The server inputs the converted text data into a sentiment analysis system, which extracts the speaker's emotional information from the text. Sentiment analysis classifies the text into emotional categories such as "positive" or "negative" based on keywords and context.

[0326] Step 4:

[0327] The server, along with the results of sentiment analysis, uses an emotion engine to analyze the user's voice tone, speed, and word choice in real time, gaining a more detailed understanding of the emotions being expressed during speech. This analysis helps to grasp changes in emotions and their intensity.

[0328] Step 5:

[0329] The server sets thresholds based on the emotion engine and emotion analysis results, and performs evaluation using an alert generation mechanism. If aggressive or inappropriate emotions exceeding a specific threshold are detected during this evaluation, warning information is generated.

[0330] Step 6:

[0331] The server sends the generated warning information to the relevant parties via a notification system using pre-configured communication tools (e.g., email, online messaging platform). The notification includes the content of the message and the sentiment analysis results.

[0332] Step 7:

[0333] The user (HR department) reviews the submitted alert information and evaluates the identification of the problem and the response measures for the victim. If necessary, they develop and implement interviews and mental health support measures. These measures include providing feedback for improvement and implementing training programs.

[0334] (Example 2)

[0335] Next, we will describe Example 2. In the following description, the data processing device 12 will be referred to as the "server" and the smart glasses 214 will be referred to as the "terminal".

[0336] In today's workplace, communication-related friction and mental health issues are on the rise. In particular, it is difficult to detect emotional triggers hidden in meetings and daily conversations early and to take appropriate action. Therefore, there is a need for a system that can monitor employees' emotional states in real time, alert them to potential problems early, and enable them to take countermeasures.

[0337] The identification process performed by the identification processing unit 290 of the data processing device 12 in Example 2 is realized by the following means.

[0338] In this invention, the server includes an input means for acquiring an audio signal, a recognition means for converting the acquired audio signal into text information, an analysis means for analyzing the text information and extracting emotion data, and an emotion recognition means for recognizing the emotional state in real time based on the characteristics of the voice and recording it in a database. This makes it possible to automatically detect emotional triggers hidden in communication and quickly notify relevant parties of alarm information.

[0339] An "audio signal" is data that converts air vibrations into electrical signals, and includes conversations and sounds.

[0340] "Input means" refers to devices or equipment used to acquire audio signals, such as microphones.

[0341] "Recognition means" refers to technologies and systems for converting audio signals into text information, and includes speech recognition software.

[0342] "Textual information" refers to text data generated based on audio signals, representing the content of the audio as written characters.

[0343] "Analysis means" refers to algorithms and programs used to analyze textual information and extract emotional data.

[0344] "Emotional data" refers to information extracted from textual information and audio characteristics, indicating the speaker's emotional state and type of emotion.

[0345] An "alert generation method" refers to a system or technology that automatically creates alarm information based on emotional data.

[0346] "Notification means" refers to methods and technologies for transmitting generated alarm information to relevant parties, and includes electronic communication means and messaging systems.

[0347] "Emotion recognition means" refers to systems or processes for understanding a speaker's emotions in real time from audio signals or text information.

[0348] This invention provides a system that detects potential emotional triggers early in workplace communication and prompts appropriate responses. This system is realized through the integrated involvement of terminals, servers, and users.

[0349] The device used will be a microphone device installed in the conference room or workplace. This microphone device can effectively capture the participants' voice signals and collect clear audio data using noise reduction technology. Specifically, it is desirable to use a microphone with noise-canceling capabilities.

[0350] The server receives the collected audio signals and converts them into text information using speech recognition software. For example, cloud-based speech recognition services can be used, such as Google Cloud Speech-to-Text. The converted text information is then analyzed using natural language processing techniques to extract sentiment data. The analysis employs algorithms utilizing generative AI models, such as BERT and GPT-3.

[0351] Based on emotion data, the server analyzes the characteristics of the voice in the audio, such as tone and speed, and evaluates the speaker's emotions in real time through emotion recognition means. This recognized emotion information is recorded in a database and stored for further analysis and future responses. The server also uses an alert generation means to automatically generate an alarm when the emotion data exceeds a certain threshold.

[0352] Alarm information is sent to users via electronic communication or online messaging systems through notification means. Examples of such means include email and online messaging platforms (such as Slack or Microsoft Teams).

[0353] Based on the received alert information, users such as HR departments and managers can quickly take appropriate countermeasures, such as interviews and counseling. As a concrete example, a prompt sentence to be input into the generating AI model could be something like, "Determine the speaker's emotions from the audio data spoken during the meeting, and propose countermeasures if an emotional trigger occurs."

[0354] Thus, by automating the collection, conversion, and analysis of voice data, the present invention enables smoother communication in the workplace and early problem resolution.

[0355] The flow of the specific processing in Example 2 will be explained using Figure 13.

[0356] Step 1:

[0357] The terminal uses a microphone device installed in the conference room to acquire participants' audio signals in real time. The microphone reduces ambient noise and ensures clear audio capture. The acquired audio signals are temporarily stored as digital data within the terminal.

[0358] Step 2:

[0359] The terminal sends the collected audio signal to the server. The server receives this audio signal using a secure communication protocol. The received digital audio signal is converted into text data using speech recognition software. This conversion analyzes the waveform data of the audio signal and replaces it with the corresponding string.

[0360] Step 3:

[0361] The server uses natural language processing techniques to extract sentiment data from the converted text data. A generative AI model is used to understand the context of the text and determine the speaker's emotions. Sentiment labels such as "joy" and "anger" are output from the input text data.

[0362] Step 4:

[0363] The server uses emotion recognition technology to analyze the characteristics of the audio signal, also evaluating the speaker's tone and speed. This evaluation combines supplementary information with text-based emotion data, allowing for real-time recognition of the emotional state. The analysis results are recorded in a database as the emotional state.

[0364] Step 5:

[0365] The server uses an alert generation mechanism based on emotional data and voice characteristics information to generate an alarm when a certain threshold is exceeded. For example, if a particular statement matches an aggressive tone, alarm data is generated. This alarm data is prepared for notification to relevant parties.

[0366] Step 6:

[0367] The notification method involves sending the generated alarm data to the user via email or online messaging system. The communication method used is selected by the system settings. Upon receiving this alarm notification, the Human Resources Department plans and prepares to implement appropriate countermeasures.

[0368] (Application Example 2)

[0369] Next, we will explain application example 2. In the following explanation, the data processing device 12 will be referred to as the "server," and the smart glasses 214 will be referred to as the "terminal."

[0370] Preventing troubles and discords in commercial facilities and offices is crucial, but conventional security systems have difficulty providing early warnings based on emotional shifts. Therefore, there is a need to monitor emotional escalations that may occur between users in real time and respond appropriately.

[0371] The specific processing performed by the specific processing unit 290 of the data processing device 12 in Application Example 2 is realized by the following means.

[0372] In this invention, the server includes voice input means for acquiring voice data, voice recognition means for converting voice into text data based on the acquired voice data, and emotion analysis means for analyzing the text data and extracting emotion information. This makes it possible to identify signs of discomfort or trouble in real time based on emotion information and generate warning information.

[0373] "Voice input means" refers to a device or system that acquires voice data.

[0374] "Speech recognition means" refers to a technology or device that converts acquired speech data into text data.

[0375] "Emotional analysis means" refers to a method or device that analyzes text data and extracts emotional information from it.

[0376] An "alert generation means" is a mechanism or function that generates warning information when specific conditions are met based on emotional information.

[0377] "Notification means" refers to a system or function that transmits generated warning information to a specific recipient using communication means.

[0378] A "control means" is a means for executing specific instructions based on warning information.

[0379] This invention is for constructing a security system based on emotion analysis. First, the entire system includes a voice input means for acquiring voice data. The voice input means consists of a device that collects voices in commercial facilities or offices. This device can be a smartphone or a dedicated voice capture device.

[0380] The server acts as a speech recognition system. This system converts the acquired speech data into text data. The software used here is the speech_recognition library for Python. After this, the text data is analyzed by an emotion analysis system. This analysis system incorporates a specific algorithm for identifying emotions, and the emotion_recognition library is used for this purpose.

[0381] The server also functions as an alert generation mechanism. Based on the sentiment analysis, if the received sentiment information meets certain conditions, the system automatically generates a warning. This warning is generated when it detects discomfort or signs of trouble that exceed a pre-set threshold.

[0382] The generated warning information is sent to security guards or personnel via notification methods. These notification methods include email and the communication functions of online messaging platforms.

[0383] As a concrete example, consider a scenario where signs of an impending argument between customers in a shopping mall are detected. This system quickly performs sentiment analysis when the volume of voices or aggressive language increases. It then notifies security personnel that "caution is needed."

[0384] The following prompt statements can be used with the generative AI model.

[0385] "From this audio data, we have detected changes in emotion. In particular, we observed 'anger' or 'fear,' which strongly suggest the need for security. We recommend the following steps." This allows for a swift and appropriate response.

[0386] The flow of a specific process in Application Example 2 will be explained using Figure 14.

[0387] Step 1:

[0388] The terminal acquires audio from commercial facilities and offices using an audio input device. The audio data is collected in real time from a microphone device. The input consists of ambient sounds and people's conversations, and the output is an audio file in digital format.

[0389] Step 2:

[0390] The server processes the acquired audio data using speech recognition. In this step, the audio data is converted into text data. The input is digital audio data, and the output is text data. Specifically, speech recognition is performed using the Python speech_recognition library.

[0391] Step 3:

[0392] The server analyzes text data using sentiment analysis tools. The input is text data, and the output is sentiment information. At this stage, the emotion_recognition library is used to identify emotions from the wording and tone contained in the text. Specifically, it analyzes keywords and phrases in the text to determine a particular emotion (e.g., "anger," "sadness").

[0393] Step 4:

[0394] The server generates warning information using an alert generation mechanism based on emotional information. The input is emotional information obtained through emotional analysis, and the output is warning information. Specifically, if the detected emotion exceeds a pre-set threshold, a warning state is confirmed. The server generates an alarm based on this information.

[0395] Step 5:

[0396] The server sends the generated warning information to the relevant parties through notification channels. The input is the warning information, and the output is a notification message. In this process, warnings are sent to security guards and administrators using email or messaging platforms. Specifically, the warning content is distributed to pre-configured email addresses or chat groups.

[0397] The specific processing unit 290 transmits the result of the specific processing to the smart glasses 214. In the smart glasses 214, the control unit 46A causes the speaker 240 to output the result of the specific processing. The microphone 238 acquires audio indicating user input for the result of the specific processing. The control unit 46A transmits the audio data indicating user input acquired by the microphone 238 to the data processing unit 12. In the data processing unit 12, the specific processing unit 290 acquires the audio data.

[0398] Data generation model 58 is a type of so-called generative AI (Artificial Intelligence). One example of data generation model 58 is ChatGPT (Internet search<URL: https: / / openai.com / blog / chatgpt> ), Gemini (Internet search) <url: https: gemini.google.com ?hl="ja">Examples of generative AI include the following. The data generation model 58 is obtained by performing deep learning on a neural network. The data generation model 58 is input with prompts containing instructions, and with inference data such as audio data representing speech, text data representing text, and image data representing images. The data generation model 58 infers from the input inference data according to the instructions indicated by the prompts, and outputs the inference results in data formats such as audio data and text data. Here, inference refers to, for example, analysis, classification, prediction, and / or summarization.

[0399] In the above embodiment, an example was given in which specific processing is performed by the data processing device 12, but the technology of this disclosure is not limited thereto, and the specific processing may also be performed by the smart glasses 214.

[0400] [Third Embodiment]

[0401] Figure 5 shows an example of the configuration of the data processing system 310 according to the third embodiment.

[0402] As shown in Figure 5, the data processing system 310 includes a data processing device 12 and a headset terminal 314. An example of the data processing device 12 is a server.

[0403] The data processing device 12 comprises a computer 22, a database 24, and a communication interface 26. The computer 22 is an example of a "computer" related to the technology of this disclosure. The computer 22 comprises a processor 28, RAM 30, and storage 32. The processor 28, RAM 30, and storage 32 are connected to a bus 34. The database 24 and the communication interface 26 are also connected to the bus 34. The communication interface 26 is connected to a network 54. An example of the network 54 is a WAN (Wide Area Network) and / or a LAN (Local Area Network).

[0404] The headset terminal 314 includes a computer 36, a microphone 238, a speaker 240, a camera 42, a communication interface 44, and a display 343. The computer 36 includes a processor 46, RAM 48, and storage 50. The processor 46, RAM 48, and storage 50 are connected to a bus 52. The microphone 238, speaker 240, camera 42, and display 343 are also connected to the bus 52.

[0405] The microphone 238 receives voice signals from the user 20 and receives instructions from the user 20. The microphone 238 captures the voice signals from the user 20, converts the captured voice into audio data, and outputs it to the processor 46. The speaker 240 outputs audio according to the instructions from the processor 46.

[0406] Camera 42 is a small digital camera equipped with an optical system including a lens, aperture, and shutter, and an image sensor such as a CMOS (Complementary Metal-Oxide-Semiconductor) image sensor or a CCD (Charge Coupled Device) image sensor, and captures images of the area around the user 20 (for example, an imaging range defined by a field of view equivalent to the width of a typical healthy person's field of vision).

[0407] Communication interface 44 is connected to network 54. Communication interfaces 44 and 26 are responsible for the exchange of various information between processor 46 and processor 28 via network 54. The exchange of various information between processor 46 and processor 28 using communication interfaces 44 and 26 is performed in a secure manner.

[0408] Figure 6 shows an example of the main functions of the data processing device 12 and the headset terminal 314. As shown in Figure 6, the data processing device 12 performs specific processing using the processor 28. The storage 32 stores the specific processing program 56.

[0409] The specific processing program 56 is an example of a "program" relating to the technology of this disclosure. The processor 28 reads the specific processing program 56 from the storage 32 and executes the read specific processing program 56 on the RAM 30. The specific processing is realized by the processor 28 operating as a specific processing unit 290 in accordance with the specific processing program 56 executed on the RAM 30.

[0410] The storage 32 stores the data generation model 58 and the emotion identification model 59. The data generation model 58 and the emotion identification model 59 are used by the identification processing unit 290.

[0411] In the headset terminal 314, the processor 46 performs the reception output processing. The storage 50 stores the reception output program 60. The processor 46 reads the reception output program 60 from the storage 50 and executes the read reception output program 60 on the RAM 48. The reception output processing is realized by the processor 46 operating as a control unit 46A according to the reception output program 60 executed on the RAM 48.

[0412] Next, the specific processing performed by the specific processing unit 290 of the data processing device 12 will be described. In the following description, the data processing device 12 will be referred to as the "server" and the headset terminal 314 will be referred to as the "terminal".

[0413] This invention provides a system that monitors meetings and daily voice communications in the workplace and detects signs of power harassment in real time. This system mainly consists of voice input means, voice recognition means, emotion analysis means, alert generation means, and notification means.

[0414] First, a microphone device, acting as a terminal, collects audio data from the meeting. This data is transmitted to a server in real time. The server uses speech recognition to convert the received audio data into text data. This converted text data is further processed, and sentiment analysis is used to identify the emotional state of the speaker. In this system, emotions are classified as "positive," "negative," or "neutral" based on the words used and the tone of voice during the conversation.

[0415] Next, the server uses an alert generation mechanism based on the sentiment analysis results to generate warning information if a certain threshold is exceeded. Key indicators in this process include whether the content of the speech is aggressive or if the tone is inappropriate. The generated warning information is automatically communicated to relevant parties through notification mechanisms. For example, it may be sent to HR personnel or managers via email or online messaging platforms.

[0416] For example, suppose during a meeting, a superior says to a subordinate in an overbearing tone, "This result is completely disappointing, and it's entirely your fault." The server converts this audio into text, then uses sentiment analysis to determine that it is "aggressive" and "blame-shifting," and the alert generation mechanism determines that it has exceeded the threshold. In this case, the system sends an alert to the human resources department, enabling immediate action.

[0417] Thus, this system allows companies to quickly and effectively manage signs of harassment in order to ensure employee safety and maintain workplace mental health.

[0418] The following describes the processing flow.

[0419] Step 1:

[0420] The terminal collects audio data in real time through a microphone installed in the conference room. This audio data is immediately converted into a digital format and transmitted to a server via the network.

[0421] Step 2:

[0422] The server inputs the received audio data into the speech recognition service, which converts it from speech to text data. The speech recognition engine analyzes phonemes, intonation, and speaker characteristics to generate highly accurate text output.

[0423] Step 3:

[0424] The server passes the obtained text data to a sentiment analysis algorithm. Here, based on the content of the text, the selected words, and the context, the emotional nuance of the text is classified as "positive," "negative," "neutral," etc.

[0425] Step 4:

[0426] The server evaluates the results of sentiment analysis, sets thresholds, and makes a decision on whether to generate an alert. If the sentiment exceeds the set criteria, especially if it is determined to be aggressive or domineering, a warning message is generated.

[0427] Step 5:

[0428] The server processes the generated warning information and prepares to notify specific parties. The information includes the specific content of the statement, the analysis results, and a timestamp.

[0429] Step 6:

[0430] The server sends alert information to HR personnel and managers via notification methods such as email and online messaging platforms like Slack. This gives stakeholders an advantage in taking early action and resolving problems.

[0431] Step 7:

[0432] The user (HR department) reviews the received warning information and, if necessary, conducts a detailed investigation of the problem and interviews with those involved. They also plan and implement corrective measures and follow-up actions as needed.

[0433] (Example 1)

[0434] Next, we will describe Example 1. In the following description, the data processing device 12 will be referred to as the "server," and the headset-type terminal 314 will be referred to as the "terminal."

[0435] There is a challenge in early detection and appropriate response to potential signs of power harassment during workplace conversations and meetings. This challenge is particularly pronounced in situations where real-time monitoring and immediate notification are required.

[0436] The identification process performed by the identification processing unit 290 of the data processing device 12 in Example 1 is realized by the following means.

[0437] In this invention, the server includes a receiving device for recording audio signals, a conversion device for converting the recorded audio signals into text information, and an evaluation device for evaluating the text information and identifying emotional characteristics. This enables the real-time detection of signs of power harassment, prompt notification to those involved, and appropriate responses.

[0438] An "audio signal" is an electrical signal obtained by converting sound vibrations collected through a microphone or recording device into an electrical signal.

[0439] A "receiving device" is a hardware or software component for capturing audio signals.

[0440] A "conversion device" is a device that has the function of analyzing audio signals and converting them into text information.

[0441] "Textual information" refers to data obtained by converting audio signals into text format.

[0442] An "evaluation device" is a device that analyzes textual information and processes it to identify emotional characteristics.

[0443] "Emotional characteristics" refer to emotional traits classified based on the evaluation results of the content of statements.

[0444] A "signal generation device" is a device that has the function of generating a warning signal when certain conditions are met.

[0445] A "warning signal" is a signal generated to draw attention based on an assessment of emotional characteristics.

[0446] "Communication equipment" refers to any device or system used to transmit or receive information.

[0447] A "transmitting device" is a device that has the function of distributing the generated warning signal to a third party via a communication means.

[0448] A "command system" is a device that has the function of issuing instructions to prompt specific actions based on warning signals.

[0449] "Electronic communication" refers to the means of sending and receiving information using electronic methods.

[0450] "Network communication means" refers to means of sending and receiving information via a digital network.

[0451] This invention is a system that monitors voice communication during workplace conversations and meetings, and quickly identifies inappropriate remarks. This system operates through the cooperation of three entities: a terminal, a server, and a user.

[0452] First, the receiving device, acting as a terminal, collects audio signals using, for example, a microphone equipped with a highly sensitive acoustic sensor. The collected audio signals are recorded as digital data and immediately transmitted to the server.

[0453] Next, the server receives this audio signal and converts it into text information using a conversion device. This procedure can utilize speech recognition technologies such as Google Cloud Speech-to-Text or Amazon Transcribe. The converted text information is then evaluated by an evaluation device to identify sentiment characteristics. In this process, a natural language processing model using the Hugging Face Transformers library is employed to classify the text information into categories such as "positive," "negative," and "neutral."

[0454] After emotional characteristics are identified, the server uses a signal generator to produce warning signals as needed. For example, if aggressive remarks towards others are detected, the system will issue a warning signal based on predefined criteria. These criteria include rules that are triggered when a threshold is exceeded. The generated warning signals are quickly notified to users and administrators using communication devices. Specifically, immediate messages are sent via Slack or Microsoft Teams.

[0455] This system enables users to address workplace environment issues quickly and accurately.

[0456] As a concrete example, consider a situation where a superior aggressively tells a subordinate during a meeting, "This result is completely disappointing, and it's entirely your fault." In this case, the server converts the audio data into text, performs sentiment analysis to determine that the statement is "aggressive" and "blame-shifting," and immediately sends a warning to the HR department. An example of a prompt for this system would be: "We want to analyze signs of power harassment from audio communication during workplace meetings. We want to classify emotions from the audio data and issue real-time alerts if aggressive remarks are made. Please detail the necessary steps and technologies to use to achieve this process."

[0457] The flow of the specific processing in Example 1 will be explained using Figure 11.

[0458] Step 1:

[0459] The receiving device, acting as a terminal, operates and collects audio signals from ambient noise during the meeting. During this process, the terminal performs filtering to reduce noise within the meeting room, resulting in clear audio data. The input data is the analog audio signal picked up by the microphone device, while the output data is the signal converted to a digital audio format.

[0460] Step 2:

[0461] The terminal transmits the collected digital audio signal to the server in real time. This process requires a stable wireless or wired network connection and minimizes data delay using a certain amount of buffering. The input data is the digital audio signal acquired in the previous step, and the output data is the same audio signal received on the server side.

[0462] Step 3:

[0463] The server activates a speech recognition device to convert the received digital audio signal into text information. The speech recognition software uses an existing model (e.g., Google Cloud Speech-to-Text) and configures language profiles to improve recognition accuracy. The input data is an audio signal, and the output data is the text data converted from that audio signal.

[0464] Step 4:

[0465] The server passes the converted text data to an evaluation device to identify sentiment characteristics. This evaluation process uses a natural language processing model (e.g., BERT) to perform sentiment analysis on the text and classify it as "positive," "negative," or "neutral." The input data is textual information, and the output data is the classified sentiment characteristics.

[0466] Step 5:

[0467] The server processes data with emotional characteristics using a signal generator and generates a warning signal if it exceeds a pre-set threshold. For example, if negative emotions are pronounced, it determines the warning level and prepares the necessary alarm message. The input data is the result of the emotional characteristics evaluation, and the output data is the message used as the warning signal.

[0468] Step 6:

[0469] The server notifies users and relevant parties of the generated warning signal via communication devices. Email and online messaging platforms are used as notification methods, and multiple notification methods are used simultaneously when immediate action is required. The input data is the warning signal, and the output data is the notification message delivered to the user's terminal.

[0470] (Application Example 1)

[0471] Next, we will explain Application Example 1. In the following explanation, the data processing device 12 will be referred to as the "server," and the headset-type terminal 314 will be referred to as the "terminal."

[0472] In the workplace environment, particularly in manufacturing and other industrial settings, inappropriate communication among employees can negatively impact work efficiency and workplace atmosphere. This can potentially impair employees' mental health and work performance. Preventing and improving such problems is therefore crucial.

[0473] The specific processing performed by the specific processing unit 290 of the data processing device 12 in Application Example 1 is realized by the following means.

[0474] In this invention, the server includes an acoustic input means for acquiring acoustic information, an acoustic recognition means for converting speech into text information based on the acquired acoustic information, and an emotion analysis means for analyzing the text information and extracting emotion information. This enables real-time monitoring of communication in the work environment, immediate detection of problematic conversations, and appropriate responses.

[0475] "Acoustic information" refers to all data related to speech and sound waves, and forms the basis for speech recognition and analysis.

[0476] "Acoustic input means" refers to a device or part of a device that has the function of collecting acoustic information. Specifically, this includes sensors such as microphones.

[0477] "Acoustic recognition means" refers to a function that converts acquired acoustic information into text information, and is a technology for converting audio data into text data.

[0478] "Textual information" refers to the text data that results from the recognition and conversion of audio information.

[0479] "Emotional analysis tools" refer to functions that extract and analyze the speaker's emotional state from textual information. They perform emotional evaluations such as positive, negative, and neutral.

[0480] "Warning information" refers to information generated as a result of sentiment analysis to alert users when certain conditions are met.

[0481] "Notification means" refers to a device or method that has the function of transmitting generated warning information or other information to relevant parties.

[0482] "Mechanical equipment" refers to devices and facilities designed to operate automatically and perform specified functions. Factory robots are a concrete example.

[0483] "Information transmission means" refers to the technologies and methods used to transmit information to designated recipients.

[0484] The system implementing this invention is primarily intended for the detection and management of inappropriate communication in the workplace environment. Specific embodiments of this system are described below.

[0485] The server acquires acoustic information using microphones installed on factory robots. This acoustic information obtained through the acoustic input means is transmitted to the server. The acoustic information is converted into text information on the server using acoustic recognition means. In this step, a speech recognition API such as Google Cloud Speech-to-Text is used. This converts the speech into text data.

[0486] Next, the converted text information is analyzed through sentiment analysis tools. This process utilizes sentiment analysis libraries such as Microsoft Azure Text Analytics. The server uses this sentiment analysis to extract the speaker's emotional state from the text information and performs sentiment evaluations such as positive, negative, or neutral.

[0487] Subsequently, based on the sentiment analysis results, the warning generation system generates warning information if specific conditions are met. This warning information is then communicated to relevant parties via email or digital messaging using the notification system.

[0488] For example, if a tense exchange occurs between a supervisor and a subordinate while machinery is operating in a factory, the audio is collected, and if sentiment analysis determines it to be "negative," a warning is generated. This warning information is then sent to the administrator to help with preventative measures.

[0489] An example of a prompt message that can be given to a generative AI model is, "Please analyze the sentiment of the following text and determine whether it is positive or negative: text."

[0490] Thus, through this invention, communication in the workplace environment can be monitored in real time, enabling early detection and response to problems.

[0491] The flow of a specific process in Application Example 1 will be explained using Figure 12.

[0492] Step 1:

[0493] Microphones equipped on factory robots collect acoustic information from the surroundings. Audio data is obtained as input and sent to a server for processing in real time.

[0494] Step 2:

[0495] The server converts the received audio data into text information using a speech recognition API such as Google Cloud Speech-to-Text. In this step, the audio information, which is the input data, is output as text data. The specific actions of recognizing the audio signal and converting it into text format are performed.

[0496] Step 3:

[0497] The converted text information is analyzed by sentiment analysis tools on the server. Using tools such as Microsoft Azure Text Analytics, sentiment information is extracted, and a sentiment rating (positive, negative, neutral, etc.) is output based on the input text. Specifically, the emotional value of each word and phrase is analyzed to determine the overall sentiment.

[0498] Step 4:

[0499] The server uses a warning generation mechanism based on the sentiment analysis results to generate warning information if the conditions are met. In this step, sentiment information is used as input, and warning information is output if it exceeds a threshold. Specific actions are taken to generate a warning when aggressive remarks or negative tones are detected.

[0500] Step 5:

[0501] The generated warning information is communicated to relevant parties via notification methods. The input is the warning information, and the output is the sending of emails or digital messages. Specifically, the warning notification is delivered to the administrator's mailbox or messaging platform.

[0502] Furthermore, an emotion engine that estimates the user's emotions may be incorporated. That is, the identification processing unit 290 may use the emotion identification model 59 to estimate the user's emotions and perform identification processing using the user's emotions.

[0503] This invention provides a system that incorporates an emotion engine to recognize user emotions when analyzing voice data. The purpose is to monitor workplace meetings and communications, detect signs of power harassment, and understand individual emotional states in real time.

[0504] As a terminal, a microphone device installed in the conference room collects participants' conversations in real time. The acquired audio data is immediately transmitted to the server. The server uses speech recognition to convert the audio data into text data. Next, sentiment analysis is performed on this text data to extract the speaker's emotional information.

[0505] Furthermore, the server incorporates an emotion engine that analyzes the user's voice tone, speed, and word choice to recognize their emotional state in real time. This information is recorded in a database and used for future analysis.

[0506] As a concrete example, suppose in a meeting, a superior tells a subordinate in a harsh tone, "Failure of this project is unacceptable." The server converts this audio into text and, through sentiment analysis and an emotion engine, determines that the statement is "aggressive" and that the speaker's emotion is "frustrated." Based on this, the alert generation mechanism determines that a threshold has been exceeded and generates a warning.

[0507] The generated warning information is automatically sent to HR personnel and managers via email or online messaging platforms. The user (HR department) uses this information to plan early intervention measures such as interviews and counseling.

[0508] Thus, by using an emotion engine, this system can analyze employees' emotions more precisely and manage workplace problems quickly and effectively. As a result, it can contribute to maintaining mental health in the workplace and improving the working environment.

[0509] The following describes the processing flow.

[0510] Step 1:

[0511] The terminal collects participants' conversations in real time through microphones installed in the conference room. The collected audio data is immediately converted into a digital signal and sent to the server.

[0512] Step 2:

[0513] The server processes the received audio data using speech recognition technology and converts the audio into text data. This conversion is achieved by the speech recognition engine analyzing the phonemes and intonation of the speech.

[0514] Step 3:

[0515] The server inputs the converted text data into a sentiment analysis system, which extracts the speaker's emotional information from the text. Sentiment analysis classifies the text into emotional categories such as "positive" or "negative" based on keywords and context.

[0516] Step 4:

[0517] The server, along with the results of sentiment analysis, uses an emotion engine to analyze the user's voice tone, speed, and word choice in real time, gaining a more detailed understanding of the emotions being expressed during speech. This analysis helps to grasp changes in emotions and their intensity.

[0518] Step 5:

[0519] The server sets thresholds based on the emotion engine and emotion analysis results, and performs evaluation using an alert generation mechanism. If aggressive or inappropriate emotions exceeding a specific threshold are detected during this evaluation, warning information is generated.

[0520] Step 6:

[0521] The server sends the generated warning information to the relevant parties via a notification system using pre-configured communication tools (e.g., email, online messaging platform). The notification includes the content of the message and the sentiment analysis results.

[0522] Step 7:

[0523] The user (HR department) reviews the submitted alert information and evaluates the identification of the problem and the response measures for the victim. If necessary, they develop and implement interviews and mental health support measures. These measures include providing feedback for improvement and implementing training programs.

[0524] (Example 2)

[0525] Next, we will describe Example 2. In the following description, the data processing device 12 will be referred to as the "server," and the headset-type terminal 314 will be referred to as the "terminal."

[0526] In today's workplace, communication-related friction and mental health issues are on the rise. In particular, it is difficult to detect emotional triggers hidden in meetings and daily conversations early and to take appropriate action. Therefore, there is a need for a system that can monitor employees' emotional states in real time, alert them to potential problems early, and enable them to take countermeasures.

[0527] The identification process performed by the identification processing unit 290 of the data processing device 12 in Example 2 is realized by the following means.

[0528] In this invention, the server includes an input means for acquiring an audio signal, a recognition means for converting the acquired audio signal into text information, an analysis means for analyzing the text information and extracting emotion data, and an emotion recognition means for recognizing the emotional state in real time based on the characteristics of the voice and recording it in a database. This makes it possible to automatically detect emotional triggers hidden in communication and quickly notify relevant parties of alarm information.

[0529] An "audio signal" is data that converts air vibrations into electrical signals, and includes conversations and sounds.

[0530] "Input means" refers to devices or equipment used to acquire audio signals, such as microphones.

[0531] "Recognition means" refers to technologies and systems for converting audio signals into text information, and includes speech recognition software.

[0532] "Textual information" refers to text data generated based on audio signals, representing the content of the audio as written characters.

[0533] "Analysis means" refers to algorithms and programs used to analyze textual information and extract emotional data.

[0534] "Emotional data" refers to information extracted from textual information and audio characteristics, indicating the speaker's emotional state and type of emotion.

[0535] An "alert generation method" refers to a system or technology that automatically creates alarm information based on emotional data.

[0536] "Notification means" refers to methods and technologies for transmitting generated alarm information to relevant parties, and includes electronic communication means and messaging systems.

[0537] "Emotion recognition means" refers to systems or processes for understanding a speaker's emotions in real time from audio signals or text information.

[0538] This invention provides a system that detects potential emotional triggers early in workplace communication and prompts appropriate responses. This system is realized through the integrated involvement of terminals, servers, and users.

[0539] The device used will be a microphone device installed in the conference room or workplace. This microphone device can effectively capture the participants' voice signals and collect clear audio data using noise reduction technology. Specifically, it is desirable to use a microphone with noise-canceling capabilities.

[0540] The server receives the collected audio signals and converts them into text information using speech recognition software. For example, cloud-based speech recognition services can be used, such as Google Cloud Speech-to-Text. The converted text information is then analyzed using natural language processing techniques to extract sentiment data. The analysis employs algorithms utilizing generative AI models, such as BERT and GPT-3.

[0541] Based on emotion data, the server analyzes the characteristics of the voice in the audio, such as tone and speed, and evaluates the speaker's emotions in real time through emotion recognition means. This recognized emotion information is recorded in a database and stored for further analysis and future responses. The server also uses an alert generation means to automatically generate an alarm when the emotion data exceeds a certain threshold.

[0542] Alarm information is sent to users via electronic communication or online messaging systems through notification means. Examples of such means include email and online messaging platforms (such as Slack or Microsoft Teams).

[0543] Based on the received alert information, users such as HR departments and managers can quickly take appropriate countermeasures, such as interviews and counseling. As a concrete example, a prompt sentence to be input into the generating AI model could be something like, "Determine the speaker's emotions from the audio data spoken during the meeting, and propose countermeasures if an emotional trigger occurs."

[0544] Thus, by automating the collection, conversion, and analysis of voice data, the present invention enables smoother communication in the workplace and early problem resolution.

[0545] The flow of the specific processing in Example 2 will be explained using Figure 13.

[0546] Step 1:

[0547] The terminal uses a microphone device installed in the conference room to acquire participants' audio signals in real time. The microphone reduces ambient noise and ensures clear audio capture. The acquired audio signals are temporarily stored as digital data within the terminal.

[0548] Step 2:

[0549] The terminal sends the collected audio signal to the server. The server receives this audio signal using a secure communication protocol. The received digital audio signal is converted into text data using speech recognition software. This conversion analyzes the waveform data of the audio signal and replaces it with the corresponding string.

[0550] Step 3:

[0551] The server uses natural language processing techniques to extract sentiment data from the converted text data. A generative AI model is used to understand the context of the text and determine the speaker's emotions. Sentiment labels such as "joy" and "anger" are output from the input text data.

[0552] Step 4:

[0553] The server uses emotion recognition technology to analyze the characteristics of the audio signal, also evaluating the speaker's tone and speed. This evaluation combines supplementary information with text-based emotion data, allowing for real-time recognition of the emotional state. The analysis results are recorded in a database as the emotional state.

[0554] Step 5:

[0555] The server uses an alert generation mechanism based on emotional data and voice characteristics information to generate an alarm when a certain threshold is exceeded. For example, if a particular statement matches an aggressive tone, alarm data is generated. This alarm data is prepared for notification to relevant parties.

[0556] Step 6:

[0557] The notification method involves sending the generated alarm data to the user via email or online messaging system. The communication method used is selected by the system settings. Upon receiving this alarm notification, the Human Resources Department plans and prepares to implement appropriate countermeasures.

[0558] (Application Example 2)

[0559] Next, we will explain application example 2. In the following explanation, the data processing device 12 will be referred to as the "server," and the headset-type terminal 314 will be referred to as the "terminal."

[0560] Preventing troubles and discords in commercial facilities and offices is crucial, but conventional security systems have difficulty providing early warnings based on emotional shifts. Therefore, there is a need to monitor emotional escalations that may occur between users in real time and respond appropriately.

[0561] The specific processing performed by the specific processing unit 290 of the data processing device 12 in Application Example 2 is realized by the following means.

[0562] In this invention, the server includes voice input means for acquiring voice data, voice recognition means for converting voice into text data based on the acquired voice data, and emotion analysis means for analyzing the text data and extracting emotion information. This makes it possible to identify signs of discomfort or trouble in real time based on emotion information and generate warning information.

[0563] "Voice input means" refers to a device or system that acquires voice data.

[0564] "Speech recognition means" refers to a technology or device that converts acquired speech data into text data.

[0565] "Emotional analysis means" refers to a method or device that analyzes text data and extracts emotional information from it.

[0566] An "alert generation means" is a mechanism or function that generates warning information when specific conditions are met based on emotional information.

[0567] "Notification means" refers to a system or function that transmits generated warning information to a specific recipient using communication means.

[0568] A "control means" is a means for executing specific instructions based on warning information.

[0569] This invention is for constructing a security system based on emotion analysis. First, the entire system includes a voice input means for acquiring voice data. The voice input means consists of a device that collects voices in commercial facilities or offices. This device can be a smartphone or a dedicated voice capture device.

[0570] The server acts as a speech recognition system. This system converts the acquired speech data into text data. The software used here is the speech_recognition library for Python. After this, the text data is analyzed by an emotion analysis system. This analysis system incorporates a specific algorithm for identifying emotions, and the emotion_recognition library is used for this purpose.

[0571] The server also functions as an alert generation mechanism. Based on the sentiment analysis, if the received sentiment information meets certain conditions, the system automatically generates a warning. This warning is generated when it detects discomfort or signs of trouble that exceed a pre-set threshold.

[0572] The generated warning information is sent to security guards or personnel via notification methods. These notification methods include email and the communication functions of online messaging platforms.

[0573] As a concrete example, consider a scenario where signs of an impending argument between customers in a shopping mall are detected. This system quickly performs sentiment analysis when the volume of voices or aggressive language increases. It then notifies security personnel that "caution is needed."

[0574] The following prompt statements can be used with the generative AI model.

[0575] "From this audio data, we have detected changes in emotion. In particular, we observed 'anger' or 'fear,' which strongly suggest the need for security. We recommend the following steps." This allows for a swift and appropriate response.

[0576] The flow of a specific process in Application Example 2 will be explained using Figure 14.

[0577] Step 1:

[0578] The terminal acquires audio from commercial facilities and offices using an audio input device. The audio data is collected in real time from a microphone device. The input consists of ambient sounds and people's conversations, and the output is an audio file in digital format.

[0579] Step 2:

[0580] The server processes the acquired audio data using speech recognition. In this step, the audio data is converted into text data. The input is digital audio data, and the output is text data. Specifically, speech recognition is performed using the Python speech_recognition library.

[0581] Step 3:

[0582] The server analyzes text data using sentiment analysis tools. The input is text data, and the output is sentiment information. At this stage, the emotion_recognition library is used to identify emotions from the wording and tone contained in the text. Specifically, it analyzes keywords and phrases in the text to determine a particular emotion (e.g., "anger," "sadness").

[0583] Step 4:

[0584] The server generates warning information using an alert generation mechanism based on emotional information. The input is emotional information obtained through emotional analysis, and the output is warning information. Specifically, if the detected emotion exceeds a pre-set threshold, a warning state is confirmed. The server generates an alarm based on this information.

[0585] Step 5:

[0586] The server sends the generated warning information to the relevant parties through notification channels. The input is the warning information, and the output is a notification message. In this process, warnings are sent to security guards and administrators using email or messaging platforms. Specifically, the warning content is distributed to pre-configured email addresses or chat groups.

[0587] The specific processing unit 290 transmits the result of the specific processing to the headset terminal 314. In the headset terminal 314, the control unit 46A causes the speaker 240 and display 343 to output the result of the specific processing. The microphone 238 acquires audio indicating user input for the result of the specific processing. The control unit 46A transmits the audio data indicating user input acquired by the microphone 238 to the data processing unit 12. In the data processing unit 12, the specific processing unit 290 acquires the audio data.

[0588] Data generation model 58 is a type of so-called generative AI (Artificial Intelligence). One example of data generation model 58 is ChatGPT (Internet search<URL: https: / / openai.com / blog / chatgpt> ), Gemini (Internet search) <url: https: gemini.google.com ?hl="ja">Examples of generative AI include the following. The data generation model 58 is obtained by performing deep learning on a neural network. The data generation model 58 is input with prompts containing instructions, and with inference data such as audio data representing speech, text data representing text, and image data representing images. The data generation model 58 infers from the input inference data according to the instructions indicated by the prompts, and outputs the inference results in data formats such as audio data and text data. Here, inference refers to, for example, analysis, classification, prediction, and / or summarization.

[0589] In the above embodiment, an example was given in which specific processing is performed by the data processing device 12, but the technology of this disclosure is not limited thereto, and specific processing may also be performed by the headset terminal 314.

[0590] [Fourth Embodiment]

[0591] Figure 7 shows an example of the configuration of the data processing system 410 according to the fourth embodiment.

[0592] As shown in Figure 7, the data processing system 410 includes a data processing device 12 and a robot 414. An example of the data processing device 12 is a server.

[0593] The data processing device 12 comprises a computer 22, a database 24, and a communication interface 26. The computer 22 is an example of a "computer" related to the technology of this disclosure. The computer 22 comprises a processor 28, RAM 30, and storage 32. The processor 28, RAM 30, and storage 32 are connected to a bus 34. The database 24 and the communication interface 26 are also connected to the bus 34. The communication interface 26 is connected to a network 54. An example of the network 54 is a WAN (Wide Area Network) and / or a LAN (Local Area Network).

[0594] The robot 414 includes a computer 36, a microphone 238, a speaker 240, a camera 42, a communication interface 44, and a controlled object 443. The computer 36 includes a processor 46, RAM 48, and storage 50. The processor 46, RAM 48, and storage 50 are connected to a bus 52. The microphone 238, speaker 240, camera 42, and controlled object 443 are also connected to the bus 52.

[0595] The microphone 238 receives voice signals from the user 20 and receives instructions from the user 20. The microphone 238 captures the voice signals from the user 20, converts the captured voice into audio data, and outputs it to the processor 46. The speaker 240 outputs audio according to the instructions from the processor 46.

[0596] Camera 42 is a small digital camera equipped with an optical system including a lens, aperture, and shutter, and an image sensor such as a CMOS (Complementary Metal-Oxide-Semiconductor) image sensor or a CCD (Charge Coupled Device) image sensor, and captures images of the area around the user 20 (for example, an imaging range defined by a field of view equivalent to the width of a typical healthy person's field of vision).

[0597] Communication interface 44 is connected to network 54. Communication interfaces 44 and 26 are responsible for the exchange of various information between processor 46 and processor 28 via network 54. The exchange of various information between processor 46 and processor 28 using communication interfaces 44 and 26 is performed in a secure manner.

[0598] The controlled object 443 includes a display device, LEDs in the eyes, and motors that drive the arms, hands, and feet. The posture and gestures of the robot 414 are controlled by controlling the motors of the arms, hands, and feet. Some of the robot 414's emotions can be expressed by controlling these motors. Furthermore, the robot 414's facial expressions can also be expressed by controlling the illumination state of the LEDs in its eyes.

[0599] Figure 8 shows an example of the main functions of the data processing device 12 and the robot 414. As shown in Figure 8, the data processing device 12 performs specific processing using the processor 28. The storage 32 stores the specific processing program 56.

[0600] The specific processing program 56 is an example of a "program" relating to the technology of this disclosure. The processor 28 reads the specific processing program 56 from the storage 32 and executes the read specific processing program 56 on the RAM 30. The specific processing is realized by the processor 28 operating as a specific processing unit 290 in accordance with the specific processing program 56 executed on the RAM 30.

[0601] The storage 32 stores the data generation model 58 and the emotion identification model 59. The data generation model 58 and the emotion identification model 59 are used by the identification processing unit 290.

[0602] In robot 414, the processor 46 performs the reception output processing. The storage 50 stores the reception output program 60. The processor 46 reads the reception output program 60 from the storage 50 and executes the read reception output program 60 on the RAM 48. The reception output processing is realized by the processor 46 operating as a control unit 46A according to the reception output program 60 executed on the RAM 48.

[0603] Next, the specific processing performed by the specific processing unit 290 of the data processing device 12 will be described. In the following description, the data processing device 12 will be referred to as the "server" and the robot 414 as the "terminal".

[0604] This invention provides a system that monitors meetings and daily voice communications in the workplace and detects signs of power harassment in real time. This system mainly consists of voice input means, voice recognition means, emotion analysis means, alert generation means, and notification means.

[0605] First, a microphone device, acting as a terminal, collects audio data from the meeting. This data is transmitted to a server in real time. The server uses speech recognition to convert the received audio data into text data. This converted text data is further processed, and sentiment analysis is used to identify the emotional state of the speaker. In this system, emotions are classified as "positive," "negative," or "neutral" based on the words used and the tone of voice during the conversation.

[0606] Next, the server uses an alert generation mechanism based on the sentiment analysis results to generate warning information if a certain threshold is exceeded. Key indicators in this process include whether the content of the speech is aggressive or if the tone is inappropriate. The generated warning information is automatically communicated to relevant parties through notification mechanisms. For example, it may be sent to HR personnel or managers via email or online messaging platforms.

[0607] For example, suppose during a meeting, a superior says to a subordinate in an overbearing tone, "This result is completely disappointing, and it's entirely your fault." The server converts this audio into text, then uses sentiment analysis to determine that it is "aggressive" and "blame-shifting," and the alert generation mechanism determines that it has exceeded the threshold. In this case, the system sends an alert to the human resources department, enabling immediate action.

[0608] Thus, this system allows companies to quickly and effectively manage signs of harassment in order to ensure employee safety and maintain workplace mental health.

[0609] The following describes the processing flow.

[0610] Step 1:

[0611] The terminal collects audio data in real time through a microphone installed in the conference room. This audio data is immediately converted into a digital format and transmitted to a server via the network.

[0612] Step 2:

[0613] The server inputs the received audio data into the speech recognition service, which converts it from speech to text data. The speech recognition engine analyzes phonemes, intonation, and speaker characteristics to generate highly accurate text output.

[0614] Step 3:

[0615] The server passes the obtained text data to a sentiment analysis algorithm. Here, based on the content of the text, the selected words, and the context, the emotional nuance of the text is classified as "positive," "negative," "neutral," etc.

[0616] Step 4:

[0617] The server evaluates the results of sentiment analysis, sets thresholds, and makes a decision on whether to generate an alert. If the sentiment exceeds the set criteria, especially if it is determined to be aggressive or domineering, a warning message is generated.

[0618] Step 5:

[0619] The server processes the generated warning information and prepares to notify specific parties. The information includes the specific content of the statement, the analysis results, and a timestamp.

[0620] Step 6:

[0621] The server sends alert information to HR personnel and managers via notification methods such as email and online messaging platforms like Slack. This gives stakeholders an advantage in taking early action and resolving problems.

[0622] Step 7:

[0623] The user (HR department) reviews the received warning information and, if necessary, conducts a detailed investigation of the problem and interviews with those involved. They also plan and implement corrective measures and follow-up actions as needed.

[0624] (Example 1)

[0625] Next, we will describe Example 1. In the following description, the data processing device 12 will be referred to as the "server" and the robot 414 as the "terminal".

[0626] There is a challenge in early detection and appropriate response to potential signs of power harassment during workplace conversations and meetings. This challenge is particularly pronounced in situations where real-time monitoring and immediate notification are required.

[0627] The identification process performed by the identification processing unit 290 of the data processing device 12 in Example 1 is realized by the following means.

[0628] In this invention, the server includes a receiving device for recording audio signals, a conversion device for converting the recorded audio signals into text information, and an evaluation device for evaluating the text information and identifying emotional characteristics. This enables the real-time detection of signs of power harassment, prompt notification to those involved, and appropriate responses.

[0629] An "audio signal" is an electrical signal obtained by converting sound vibrations collected through a microphone or recording device into an electrical signal.

[0630] A "receiving device" is a hardware or software component for capturing audio signals.

[0631] A "conversion device" is a device that has the function of analyzing audio signals and converting them into text information.

[0632] "Textual information" refers to data obtained by converting audio signals into text format.

[0633] An "evaluation device" is a device that analyzes textual information and processes it to identify emotional characteristics.

[0634] "Emotional characteristics" refer to emotional traits classified based on the evaluation results of the content of statements.

[0635] A "signal generation device" is a device that has the function of generating a warning signal when certain conditions are met.

[0636] A "warning signal" is a signal generated to draw attention based on an assessment of emotional characteristics.

[0637] "Communication equipment" refers to any device or system used to transmit or receive information.

[0638] A "transmitting device" is a device that has the function of distributing the generated warning signal to a third party via a communication means.

[0639] A "command system" is a device that has the function of issuing instructions to prompt specific actions based on warning signals.

[0640] "Electronic communication" refers to the means of sending and receiving information using electronic methods.

[0641] "Network communication means" refers to means of sending and receiving information via a digital network.

[0642] This invention is a system that monitors voice communication during workplace conversations and meetings, and quickly identifies inappropriate remarks. This system operates through the cooperation of three entities: a terminal, a server, and a user.

[0643] First, the receiving device, acting as a terminal, collects audio signals using, for example, a microphone equipped with a highly sensitive acoustic sensor. The collected audio signals are recorded as digital data and immediately transmitted to the server.

[0644] Next, the server receives this audio signal and converts it into text information using a conversion device. This procedure can utilize speech recognition technologies such as Google Cloud Speech-to-Text or Amazon Transcribe. The converted text information is then evaluated by an evaluation device to identify sentiment characteristics. In this process, a natural language processing model using the Hugging Face Transformers library is employed to classify the text information into categories such as "positive," "negative," and "neutral."

[0645] After emotional characteristics are identified, the server uses a signal generator to produce warning signals as needed. For example, if aggressive remarks towards others are detected, the system will issue a warning signal based on predefined criteria. These criteria include rules that are triggered when a threshold is exceeded. The generated warning signals are quickly notified to users and administrators using communication devices. Specifically, immediate messages are sent via Slack or Microsoft Teams.

[0646] This system enables users to address workplace environment issues quickly and accurately.

[0647] As a concrete example, consider a situation where a superior aggressively tells a subordinate during a meeting, "This result is completely disappointing, and it's entirely your fault." In this case, the server converts the audio data into text, performs sentiment analysis to determine that the statement is "aggressive" and "blame-shifting," and immediately sends a warning to the HR department. An example of a prompt for this system would be: "We want to analyze signs of power harassment from audio communication during workplace meetings. We want to classify emotions from the audio data and issue real-time alerts if aggressive remarks are made. Please detail the necessary steps and technologies to use to achieve this process."

[0648] The flow of the specific processing in Example 1 will be explained using Figure 11.

[0649] Step 1:

[0650] The receiving device, acting as a terminal, operates and collects audio signals from ambient noise during the meeting. During this process, the terminal performs filtering to reduce noise within the meeting room, resulting in clear audio data. The input data is the analog audio signal picked up by the microphone device, while the output data is the signal converted to a digital audio format.

[0651] Step 2:

[0652] The terminal transmits the collected digital audio signal to the server in real time. This process requires a stable wireless or wired network connection and minimizes data delay using a certain amount of buffering. The input data is the digital audio signal acquired in the previous step, and the output data is the same audio signal received on the server side.

[0653] Step 3:

[0654] The server activates a speech recognition device to convert the received digital audio signal into text information. The speech recognition software uses an existing model (e.g., Google Cloud Speech-to-Text) and configures language profiles to improve recognition accuracy. The input data is an audio signal, and the output data is the text data converted from that audio signal.

[0655] Step 4:

[0656] The server passes the converted text data to an evaluation device to identify sentiment characteristics. This evaluation process uses a natural language processing model (e.g., BERT) to perform sentiment analysis on the text and classify it as "positive," "negative," or "neutral." The input data is textual information, and the output data is the classified sentiment characteristics.

[0657] Step 5:

[0658] The server processes data with emotional characteristics using a signal generator and generates a warning signal if it exceeds a pre-set threshold. For example, if negative emotions are pronounced, it determines the warning level and prepares the necessary alarm message. The input data is the result of the emotional characteristics evaluation, and the output data is the message used as the warning signal.

[0659] Step 6:

[0660] The server notifies users and relevant parties of the generated warning signal via communication devices. Email and online messaging platforms are used as notification methods, and multiple notification methods are used simultaneously when immediate action is required. The input data is the warning signal, and the output data is the notification message delivered to the user's terminal.

[0661] (Application Example 1)

[0662] Next, we will explain Application Example 1. In the following explanation, the data processing device 12 will be referred to as the "server" and the robot 414 as the "terminal".

[0663] In the workplace environment, particularly in manufacturing and other industrial settings, inappropriate communication among employees can negatively impact work efficiency and workplace atmosphere. This can potentially impair employees' mental health and work performance. Preventing and improving such problems is therefore crucial.

[0664] The specific processing performed by the specific processing unit 290 of the data processing device 12 in Application Example 1 is realized by the following means.

[0665] In this invention, the server includes an acoustic input means for acquiring acoustic information, an acoustic recognition means for converting speech into text information based on the acquired acoustic information, and an emotion analysis means for analyzing the text information and extracting emotion information. This enables real-time monitoring of communication in the work environment, immediate detection of problematic conversations, and appropriate responses.

[0666] "Acoustic information" refers to all data related to speech and sound waves, and forms the basis for speech recognition and analysis.

[0667] "Acoustic input means" refers to a device or part of a device that has the function of collecting acoustic information. Specifically, this includes sensors such as microphones.

[0668] "Acoustic recognition means" refers to a function that converts acquired acoustic information into text information, and is a technology for converting audio data into text data.

[0669] "Textual information" refers to the text data that results from the recognition and conversion of audio information.

[0670] "Emotional analysis tools" refer to functions that extract and analyze the speaker's emotional state from textual information. They perform emotional evaluations such as positive, negative, and neutral.

[0671] "Warning information" refers to information generated as a result of sentiment analysis to alert users when certain conditions are met.

[0672] "Notification means" refers to a device or method that has the function of transmitting generated warning information or other information to relevant parties.

[0673] "Mechanical equipment" refers to devices and facilities designed to operate automatically and perform specified functions. Factory robots are a concrete example.

[0674] "Information transmission means" refers to the technologies and methods used to transmit information to designated recipients.

[0675] The system implementing this invention is primarily intended for the detection and management of inappropriate communication in the workplace environment. Specific embodiments of this system are described below.

[0676] The server acquires acoustic information using microphones installed on factory robots. This acoustic information obtained through the acoustic input means is transmitted to the server. The acoustic information is converted into text information on the server using acoustic recognition means. In this step, a speech recognition API such as Google Cloud Speech-to-Text is used. This converts the speech into text data.

[0677] Next, the converted text information is analyzed through sentiment analysis tools. This process utilizes sentiment analysis libraries such as Microsoft Azure Text Analytics. The server uses this sentiment analysis to extract the speaker's emotional state from the text information and performs sentiment evaluations such as positive, negative, or neutral.

[0678] Subsequently, based on the sentiment analysis results, the warning generation system generates warning information if specific conditions are met. This warning information is then communicated to relevant parties via email or digital messaging using the notification system.

[0679] For example, if a tense exchange occurs between a supervisor and a subordinate while machinery is operating in a factory, the audio is collected, and if sentiment analysis determines it to be "negative," a warning is generated. This warning information is then sent to the administrator to help with preventative measures.

[0680] An example of a prompt message that can be given to a generative AI model is, "Please analyze the sentiment of the following text and determine whether it is positive or negative: text."

[0681] Thus, through this invention, communication in the workplace environment can be monitored in real time, enabling early detection and response to problems.

[0682] The flow of a specific process in Application Example 1 will be explained using Figure 12.

[0683] Step 1:

[0684] Microphones equipped on factory robots collect acoustic information from the surroundings. Audio data is obtained as input and sent to a server for processing in real time.

[0685] Step 2:

[0686] The server converts the received audio data into text information using a speech recognition API such as Google Cloud Speech-to-Text. In this step, the audio information, which is the input data, is output as text data. The specific actions of recognizing the audio signal and converting it into text format are performed.

[0687] Step 3:

[0688] The converted text information is analyzed by sentiment analysis tools on the server. Using tools such as Microsoft Azure Text Analytics, sentiment information is extracted, and a sentiment rating (positive, negative, neutral, etc.) is output based on the input text. Specifically, the emotional value of each word and phrase is analyzed to determine the overall sentiment.

[0689] Step 4:

[0690] The server uses a warning generation mechanism based on the sentiment analysis results to generate warning information if the conditions are met. In this step, sentiment information is used as input, and warning information is output if it exceeds a threshold. Specific actions are taken to generate a warning when aggressive remarks or negative tones are detected.

[0691] Step 5:

[0692] The generated warning information is communicated to relevant parties via notification methods. The input is the warning information, and the output is the sending of emails or digital messages. Specifically, the warning notification is delivered to the administrator's mailbox or messaging platform.

[0693] Furthermore, an emotion engine that estimates the user's emotions may be incorporated. That is, the identification processing unit 290 may use the emotion identification model 59 to estimate the user's emotions and perform identification processing using the user's emotions.

[0694] This invention provides a system that incorporates an emotion engine to recognize user emotions when analyzing voice data. The purpose is to monitor workplace meetings and communications, detect signs of power harassment, and understand individual emotional states in real time.

[0695] As a terminal, a microphone device installed in the conference room collects participants' conversations in real time. The acquired audio data is immediately transmitted to the server. The server uses speech recognition to convert the audio data into text data. Next, sentiment analysis is performed on this text data to extract the speaker's emotional information.

[0696] Furthermore, the server incorporates an emotion engine that analyzes the user's voice tone, speed, and word choice to recognize their emotional state in real time. This information is recorded in a database and used for future analysis.

[0697] As a concrete example, suppose in a meeting, a superior tells a subordinate in a harsh tone, "Failure of this project is unacceptable." The server converts this audio into text and, through sentiment analysis and an emotion engine, determines that the statement is "aggressive" and that the speaker's emotion is "frustrated." Based on this, the alert generation mechanism determines that a threshold has been exceeded and generates a warning.

[0698] The generated warning information is automatically sent to HR personnel and managers via email or online messaging platforms. The user (HR department) uses this information to plan early intervention measures such as interviews and counseling.

[0699] Thus, by using an emotion engine, this system can analyze employees' emotions more precisely and manage workplace problems quickly and effectively. As a result, it can contribute to maintaining mental health in the workplace and improving the working environment.

[0700] The following describes the processing flow.

[0701] Step 1:

[0702] The terminal collects participants' conversations in real time through microphones installed in the conference room. The collected audio data is immediately converted into a digital signal and sent to the server.

[0703] Step 2:

[0704] The server processes the received audio data using speech recognition technology and converts the audio into text data. This conversion is achieved by the speech recognition engine analyzing the phonemes and intonation of the speech.

[0705] Step 3:

[0706] The server inputs the converted text data into a sentiment analysis system, which extracts the speaker's emotional information from the text. Sentiment analysis classifies the text into emotional categories such as "positive" or "negative" based on keywords and context.

[0707] Step 4:

[0708] The server, along with the results of sentiment analysis, uses an emotion engine to analyze the user's voice tone, speed, and word choice in real time, gaining a more detailed understanding of the emotions being expressed during speech. This analysis helps to grasp changes in emotions and their intensity.

[0709] Step 5:

[0710] The server sets thresholds based on the emotion engine and emotion analysis results, and performs evaluation using an alert generation mechanism. If aggressive or inappropriate emotions exceeding a specific threshold are detected during this evaluation, warning information is generated.

[0711] Step 6:

[0712] The server sends the generated warning information to the relevant parties via a notification system using pre-configured communication tools (e.g., email, online messaging platform). The notification includes the content of the message and the sentiment analysis results.

[0713] Step 7:

[0714] The user (HR department) reviews the submitted alert information and evaluates the identification of the problem and the response measures for the victim. If necessary, they develop and implement interviews and mental health support measures. These measures include providing feedback for improvement and implementing training programs.

[0715] (Example 2)

[0716] Next, we will describe Example 2. In the following description, the data processing device 12 will be referred to as the "server" and the robot 414 as the "terminal".

[0717] In today's workplace, communication-related friction and mental health issues are on the rise. In particular, it is difficult to detect emotional triggers hidden in meetings and daily conversations early and to take appropriate action. Therefore, there is a need for a system that can monitor employees' emotional states in real time, alert them to potential problems early, and enable them to take countermeasures.

[0718] The identification process performed by the identification processing unit 290 of the data processing device 12 in Example 2 is realized by the following means.

[0719] In this invention, the server includes an input means for acquiring an audio signal, a recognition means for converting the acquired audio signal into text information, an analysis means for analyzing the text information and extracting emotion data, and an emotion recognition means for recognizing the emotional state in real time based on the characteristics of the voice and recording it in a database. This makes it possible to automatically detect emotional triggers hidden in communication and quickly notify relevant parties of alarm information.

[0720] An "audio signal" is data that converts air vibrations into electrical signals, and includes conversations and sounds.

[0721] "Input means" refers to devices or equipment used to acquire audio signals, such as microphones.

[0722] "Recognition means" refers to technologies and systems for converting audio signals into text information, and includes speech recognition software.

[0723] "Textual information" refers to text data generated based on audio signals, representing the content of the audio as written characters.

[0724] "Analysis means" refers to algorithms and programs used to analyze textual information and extract emotional data.

[0725] "Emotional data" refers to information extracted from textual information and audio characteristics, indicating the speaker's emotional state and type of emotion.

[0726] An "alert generation method" refers to a system or technology that automatically creates alarm information based on emotional data.

[0727] "Notification means" refers to methods and technologies for transmitting generated alarm information to relevant parties, and includes electronic communication means and messaging systems.

[0728] "Emotion recognition means" refers to systems or processes for understanding a speaker's emotions in real time from audio signals or text information.

[0729] This invention provides a system that detects potential emotional triggers early in workplace communication and prompts appropriate responses. This system is realized through the integrated involvement of terminals, servers, and users.

[0730] The device used will be a microphone device installed in the conference room or workplace. This microphone device can effectively capture the participants' voice signals and collect clear audio data using noise reduction technology. Specifically, it is desirable to use a microphone with noise-canceling capabilities.

[0731] The server receives the collected audio signals and converts them into text information using speech recognition software. For example, cloud-based speech recognition services can be used, such as Google Cloud Speech-to-Text. The converted text information is then analyzed using natural language processing techniques to extract sentiment data. The analysis employs algorithms utilizing generative AI models, such as BERT and GPT-3.

[0732] Based on emotion data, the server analyzes the characteristics of the voice in the audio, such as tone and speed, and evaluates the speaker's emotions in real time through emotion recognition means. This recognized emotion information is recorded in a database and stored for further analysis and future responses. The server also uses an alert generation means to automatically generate an alarm when the emotion data exceeds a certain threshold.

[0733] Alarm information is sent to users via electronic communication or online messaging systems through notification means. Examples of such means include email and online messaging platforms (such as Slack or Microsoft Teams).

[0734] Based on the received alert information, users such as HR departments and managers can quickly take appropriate countermeasures, such as interviews and counseling. As a concrete example, a prompt sentence to be input into the generating AI model could be something like, "Determine the speaker's emotions from the audio data spoken during the meeting, and propose countermeasures if an emotional trigger occurs."

[0735] Thus, by automating the collection, conversion, and analysis of voice data, the present invention enables smoother communication in the workplace and early problem resolution.

[0736] The flow of the specific processing in Example 2 will be explained using Figure 13.

[0737] Step 1:

[0738] The terminal uses a microphone device installed in the conference room to acquire participants' audio signals in real time. The microphone reduces ambient noise and ensures clear audio capture. The acquired audio signals are temporarily stored as digital data within the terminal.

[0739] Step 2:

[0740] The terminal sends the collected audio signal to the server. The server receives this audio signal using a secure communication protocol. The received digital audio signal is converted into text data using speech recognition software. This conversion analyzes the waveform data of the audio signal and replaces it with the corresponding string.

[0741] Step 3:

[0742] The server uses natural language processing techniques to extract sentiment data from the converted text data. A generative AI model is used to understand the context of the text and determine the speaker's emotions. Sentiment labels such as "joy" and "anger" are output from the input text data.

[0743] Step 4:

[0744] The server uses emotion recognition technology to analyze the characteristics of the audio signal, also evaluating the speaker's tone and speed. This evaluation combines supplementary information with text-based emotion data, allowing for real-time recognition of the emotional state. The analysis results are recorded in a database as the emotional state.

[0745] Step 5:

[0746] The server uses an alert generation mechanism based on emotional data and voice characteristics information to generate an alarm when a certain threshold is exceeded. For example, if a particular statement matches an aggressive tone, alarm data is generated. This alarm data is prepared for notification to relevant parties.

[0747] Step 6:

[0748] The notification method involves sending the generated alarm data to the user via email or online messaging system. The communication method used is selected by the system settings. Upon receiving this alarm notification, the Human Resources Department plans and prepares to implement appropriate countermeasures.

[0749] (Application Example 2)

[0750] Next, we will explain application example 2. In the following explanation, the data processing device 12 will be referred to as the "server" and the robot 414 as the "terminal".

[0751] Preventing troubles and discords in commercial facilities and offices is crucial, but conventional security systems have difficulty providing early warnings based on emotional shifts. Therefore, there is a need to monitor emotional escalations that may occur between users in real time and respond appropriately.

[0752] The specific processing performed by the specific processing unit 290 of the data processing device 12 in Application Example 2 is realized by the following means.

[0753] In this invention, the server includes voice input means for acquiring voice data, voice recognition means for converting voice into text data based on the acquired voice data, and emotion analysis means for analyzing the text data and extracting emotion information. This makes it possible to identify signs of discomfort or trouble in real time based on emotion information and generate warning information.

[0754] "Voice input means" refers to a device or system that acquires voice data.

[0755] "Speech recognition means" refers to a technology or device that converts acquired speech data into text data.

[0756] "Emotional analysis means" refers to a method or device that analyzes text data and extracts emotional information from it.

[0757] An "alert generation means" is a mechanism or function that generates warning information when specific conditions are met based on emotional information.

[0758] "Notification means" refers to a system or function that transmits generated warning information to a specific recipient using communication means.

[0759] A "control means" is a means for executing specific instructions based on warning information.

[0760] This invention is for constructing a security system based on emotion analysis. First, the entire system includes a voice input means for acquiring voice data. The voice input means consists of a device that collects voices in commercial facilities or offices. This device can be a smartphone or a dedicated voice capture device.

[0761] The server acts as a speech recognition system. This system converts the acquired speech data into text data. The software used here is the speech_recognition library for Python. After this, the text data is analyzed by an emotion analysis system. This analysis system incorporates a specific algorithm for identifying emotions, and the emotion_recognition library is used for this purpose.

[0762] The server also functions as an alert generation mechanism. Based on the sentiment analysis, if the received sentiment information meets certain conditions, the system automatically generates a warning. This warning is generated when it detects discomfort or signs of trouble that exceed a pre-set threshold.

[0763] The generated warning information is sent to security guards or personnel via notification methods. These notification methods include email and the communication functions of online messaging platforms.

[0764] As a concrete example, consider a scenario where signs of an impending argument between customers in a shopping mall are detected. This system quickly performs sentiment analysis when the volume of voices or aggressive language increases. It then notifies security personnel that "caution is needed."

[0765] The following prompt statements can be used with the generative AI model.

[0766] "From this audio data, we have detected changes in emotion. In particular, we observed 'anger' or 'fear,' which strongly suggest the need for security. We recommend the following steps." This allows for a swift and appropriate response.

[0767] The flow of a specific process in Application Example 2 will be explained using Figure 14.

[0768] Step 1:

[0769] The terminal acquires audio from commercial facilities and offices using an audio input device. The audio data is collected in real time from a microphone device. The input consists of ambient sounds and people's conversations, and the output is an audio file in digital format.

[0770] Step 2:

[0771] The server processes the acquired audio data using speech recognition. In this step, the audio data is converted into text data. The input is digital audio data, and the output is text data. Specifically, speech recognition is performed using the Python speech_recognition library.

[0772] Step 3:

[0773] The server analyzes text data using sentiment analysis tools. The input is text data, and the output is sentiment information. At this stage, the emotion_recognition library is used to identify emotions from the wording and tone contained in the text. Specifically, it analyzes keywords and phrases in the text to determine a particular emotion (e.g., "anger," "sadness").

[0774] Step 4:

[0775] The server generates warning information using an alert generation mechanism based on emotional information. The input is emotional information obtained through emotional analysis, and the output is warning information. Specifically, if the detected emotion exceeds a pre-set threshold, a warning state is confirmed. The server generates an alarm based on this information.

[0776] Step 5:

[0777] The server sends the generated warning information to the relevant parties through notification channels. The input is the warning information, and the output is a notification message. In this process, warnings are sent to security guards and administrators using email or messaging platforms. Specifically, the warning content is distributed to pre-configured email addresses or chat groups.

[0778] The specific processing unit 290 transmits the result of the specific processing to the robot 414. In the robot 414, the control unit 46A causes the speaker 240 and the controlled object 443 to output the result of the specific processing. The microphone 238 acquires audio indicating user input for the result of the specific processing. The control unit 46A transmits the audio data indicating user input acquired by the microphone 238 to the data processing unit 12. In the data processing unit 12, the specific processing unit 290 acquires the audio data.

[0779] Data generation model 58 is a type of so-called generative AI (Artificial Intelligence). One example of data generation model 58 is ChatGPT (Internet search<URL: https: / / openai.com / blog / chatgpt> ), Gemini (Internet search) <url: https: gemini.google.com ?hl="ja">Examples of generative AI include the following. The data generation model 58 is obtained by performing deep learning on a neural network. The data generation model 58 is input with prompts containing instructions, and with inference data such as audio data representing speech, text data representing text, and image data representing images. The data generation model 58 infers from the input inference data according to the instructions indicated by the prompts, and outputs the inference results in data formats such as audio data and text data. Here, inference refers to, for example, analysis, classification, prediction, and / or summarization.

[0780] In the above embodiment, an example was given in which specific processing is performed by the data processing device 12, but the technology of this disclosure is not limited thereto, and the specific processing may also be performed by the robot 414.

[0781] Furthermore, the emotion identification model 59, acting as an emotion engine, may determine the user's emotion according to a specific mapping. Specifically, the emotion identification model 59 may determine the user's emotion according to a specific mapping, which is an emotion map (see Figure 9). Similarly, the emotion identification model 59 may also determine the robot's emotion, and the identification processing unit 290 may perform identification processing using the robot's emotion.

[0782] Figure 9 shows an emotion map 400 in which multiple emotions are mapped. In the emotion map 400, emotions are arranged in concentric circles radiating from the center. The closer to the center of the concentric circles, the more primitive the emotions are located. Further out of the concentric circles, emotions representing states and actions arising from mental states are located. Emotion is a concept that includes feelings and mental states. On the left side of the concentric circles, emotions that are generally generated from reactions occurring in the brain are located. On the right side of the concentric circles, emotions that are generally induced by situational judgment are located. Above and below the concentric circles, emotions that are generally generated from reactions occurring in the brain and induced by situational judgment are located. In addition, the emotion of "pleasure" is located on the upper side of the concentric circles, and the emotion of "displeasure" is located on the lower side. Thus, in the emotion map 400, multiple emotions are mapped based on the structure in which emotions arise, and emotions that are likely to occur simultaneously are mapped close together.

[0783] These emotions are distributed at the 3 o'clock position on the Emotion Map 400, and usually fluctuate between feelings of security and anxiety. In the right half of the Emotion Map 400, situational awareness takes precedence over internal feelings, resulting in a calm impression.

[0784] The inside of the Emotion Map 400 represents inner thoughts, while the outside represents actions. Therefore, the further you go from the outside of the Emotion Map 400, the more visible (expressed in actions) your emotions become.

[0785] Here, human emotions are based on various balances, such as posture and blood sugar levels. When these balances deviate from the ideal, it results in discomfort, and when they approach the ideal, it results in pleasure. Similarly, in robots, cars, motorcycles, etc., emotions can be created based on various balances, such as posture and battery level. When these balances deviate from the ideal, it results in discomfort, and when they approach the ideal, it results in pleasure. The emotion map can be generated, for example, based on Dr. Mitsuyoshi's emotion map (Research on a system for analyzing brain physiological signals of speech emotion recognition and emotion, Tokushima University, doctoral dissertation: https: / / ci.nii.ac.jp / naid / 500000375379). The left half of the emotion map contains emotions belonging to a region called "response," where sensation is dominant. The right half of the emotion map contains emotions belonging to a region called "situation," where situational awareness is dominant.

[0786] The emotion map defines two emotions that promote learning. One is the emotion around the middle of the negative "repentance" and "reflection" on the situation side. In other words, it is when the robot experiences negative emotions such as "I never want to feel this way again" or "I don't want to be scolded again." The other is the emotion around the positive "desire" on the reaction side. In other words, it is when the robot has positive feelings such as "I want more" or "I want to know more."

[0787] The emotion identification model 59 inputs user input into a pre-trained neural network, obtains emotion values ​​representing each emotion shown in the emotion map 400, and determines the user's emotion. This neural network is pre-trained based on multiple training data sets, which are combinations of user input and emotion values ​​representing each emotion shown in the emotion map 400. Furthermore, this neural network is trained so that emotions located close together have similar values, as shown in the emotion map 900 in Figure 10. Figure 10 shows an example where multiple emotions such as "reassured," "calm," and "confident" have similar emotion values.

[0788] The above description primarily focuses on the functions of the data processing device 12 in relation to this disclosure. However, the system related to this disclosure is not necessarily implemented on a server. The system related to this disclosure may be implemented as a general information processing system. This disclosure may be implemented, for example, as a software program that runs on a personal computer or as an application that runs on a smartphone. The method related to this disclosure may be provided to users in SaaS (Software as a Service) format.

[0789] In the above embodiment, an example was given in which a specific process is performed by a single computer 22. However, the technology of this disclosure is not limited thereto, and a distributed processing of the specific process may be performed by multiple computers, including computer 22. For example, a data generation model 58 may be provided in an external device of the data processing device 12, and the external device may generate data according to the input data.

[0790] In the above embodiment, an example was given in which the specific processing program 56 is stored in the storage 32, but the technology of this disclosure is not limited thereto. For example, the specific processing program 56 may be stored in a portable, computer-readable, non-temporary storage medium such as a USB (Universal Serial Bus) memory. The specific processing program 56 stored in the non-temporary storage medium is installed in the computer 22 of the data processing device 12. The processor 28 executes specific processing according to the specific processing program 56.

[0791] Alternatively, the specific processing program 56 may be stored in a storage device such as a server connected to the data processing device 12 via the network 54, and the specific processing program 56 may be downloaded and installed on the computer 22 in response to a request from the data processing device 12.

[0792] Furthermore, it is not necessary to store the entirety of the specific processing program 56 in a storage device such as a server connected to the data processing device 12 via the network 54, or to store the entirety of the specific processing program 56 in the storage 32; it is acceptable to store only a portion of the specific processing program 56.

[0793] The following types of processors can be used as hardware resources to perform specific processing. Examples of processors include a CPU, a general-purpose processor that functions as a hardware resource to perform specific processing by executing software, i.e., a program. Other examples of processors include dedicated electrical circuits, such as FPGAs (Field-Programmable Gate Arrays), PLDs (Programmable Logic Devices), or ASICs (Application Specific Integrated Circuits), which have circuit configurations specifically designed to perform specific processing. All of these processors have built-in or connected memory, and all of them perform specific processing by using memory.

[0794] The hardware resource that performs a specific process may consist of one of these various processors, or it may consist of a combination of two or more processors of the same or different types (for example, a combination of multiple FPGAs, or a combination of a CPU and an FPGA). Alternatively, the hardware resource that performs a specific process may consist of a single processor.

[0795] Examples of configurations using a single processor include, firstly, a configuration in which one or more CPUs and software are combined to form a single processor, and this processor functions as a hardware resource that performs a specific process. Secondly, there is a configuration using a processor that realizes the functions of the entire system, including multiple hardware resources that perform a specific process, on a single IC chip, as exemplified by SoCs (System-on-a-chip). In this way, a specific process is realized using one or more of the above types of processors as hardware resources.

[0796] Furthermore, the hardware structure of these various processors can more specifically utilize electrical circuits that combine circuit elements such as semiconductor devices. Also, the specific processing described above is merely an example. Therefore, it goes without saying that unnecessary steps can be deleted, new steps added, or the processing order rearranged, as long as it does not deviate from the main purpose.

[0797] The descriptions and illustrations presented above are detailed explanations of the technical aspects of this disclosure and are merely examples of the technical aspects. For example, the above descriptions of the structure, function, operation, and effect are examples of the structure, function, operation, and effect of the technical aspects of this disclosure. Therefore, it goes without saying that you may delete unnecessary parts, add new elements, or replace elements in the descriptions and illustrations presented above, as long as you do not deviate from the essence of the technical aspects of this disclosure. Furthermore, in order to avoid confusion and facilitate understanding of the technical aspects of this disclosure, explanations of common technical knowledge and the like that do not require special explanation to enable the implementation of the technical aspects of this disclosure have been omitted from the descriptions and illustrations presented above.

[0798] All documents, patent applications, and technical standards described herein are incorporated by reference to the same extent as if each individual document, patent application, and technical standard were specifically and individually noted to be incorporated by reference.

[0799] The following is further disclosed regarding the embodiments described above.

[0800] (Claim 1)

[0801] A voice input means for acquiring voice data,

[0802] A speech recognition means that converts speech into text data based on acquired speech data,

[0803] A sentiment analysis method that analyzes text data and extracts emotional information,

[0804] An alert generation means that generates warning information when specific conditions are met based on emotional information,

[0805] A notification means that notifies relevant parties of warning information using a pre-configured communication method,

[0806] A system that includes this.

[0807] (Claim 2)

[0808] The system according to claim 1, further comprising control means for instructing a specified action to be taken based on warning information.

[0809] (Claim 3)

[0810] The system according to claim 1, wherein the notification means transmits warning information using email or an online messaging platform.

[0811] "Example 1"

[0812] (Claim 1)

[0813] A receiving device for recording audio signals,

[0814] A conversion device that converts recorded audio signals into text information,

[0815] An evaluation device that evaluates textual information and identifies emotional characteristics,

[0816] A signal generating device that generates a warning signal when a predefined standard is exceeded based on emotional characteristics,

[0817] A transmitting device that sends a warning signal to a third party using a pre-configured communication device,

[0818] A system that includes this.

[0819] (Claim 2)

[0820] The system according to claim 1, further comprising a command device that instructs a specific response to be prompted based on a warning signal.

[0821] (Claim 3)

[0822] The system according to claim 1, wherein the transmitting device distributes a warning signal using electronic communication or network communication means.

[0823] "Application Example 1"

[0824] (Claim 1)

[0825] An acoustic input means for acquiring acoustic information,

[0826] Acoustic recognition means that converts speech into text information based on acquired acoustic information,

[0827] A means of sentiment analysis that analyzes textual information and extracts emotional information,

[0828] A warning generation means that generates warning information when specific conditions are met based on emotional information,

[0829] A notification means that notifies relevant parties of warning information using a pre-configured information transmission means,

[0830] A means of monitoring acoustic information in the work environment in real time using an acoustic information input device installed on a machine,

[0831] A system that includes this.

[0832] (Claim 2)

[0833] The system according to claim 1, further comprising control means for instructing a specified action to be taken based on warning information.

[0834] (Claim 3)

[0835] The system according to claim 1, wherein the notification means transmits warning information using electronic communication or digital messaging means.

[0836] "Example 2 of combining an emotion engine"

[0837] (Claim 1)

[0838] An input means for acquiring audio signals,

[0839] A recognition means that converts a signal into text information based on the acquired audio signal,

[0840] An analytical means for analyzing textual information and extracting emotional data,

[0841] An alert generation means that generates alarm information when specific conditions are met based on emotional data and voice characteristics,

[0842] A notification means that transmits alarm information to relevant parties using a pre-configured communication method,

[0843] A system that includes this.

[0844] (Claim 2)

[0845] The system according to claim 1, further comprising emotion recognition means for recognizing emotional states in real time based on voice characteristics and recording them in a database.

[0846] (Claim 3)

[0847] The system according to claim 1, wherein the notification means transmits alarm information using electronic communication or an online messaging method.

[0848] "Application example 2 when combining with an emotional engine"

[0849] (Claim 1)

[0850] A voice input means for acquiring voice data,

[0851] A speech recognition means that converts speech into text data based on acquired speech data,

[0852] A sentiment analysis method that analyzes text data and extracts emotional information,

[0853] An alert generation method that identifies signs of discomfort or trouble in real time based on emotional information and generates warning information when specific conditions are met,

[0854] A notification system that uses pre-configured communication methods to notify security guards or personnel of warning information,

[0855] A system that includes this.

[0856] (Claim 2)

[0857] The system according to claim 1, further comprising control means for instructing a designated defensive action to be taken based on warning information.

[0858] (Claim 3)

[0859] The system according to claim 1, wherein the notification means transmits warning information using electronic communication means. [Explanation of symbols]

[0860] 10, 210, 310, 410 Data Processing Systems 12 Data Processing Devices 14 Smart Devices 214 Smart Glasses 314 Headset-type terminal 414 Robots< / url:> < / url:> < / url:> < / url:>

Claims

1. A voice input means for acquiring voice data, A speech recognition means that converts speech into text data based on acquired speech data, A sentiment analysis method that analyzes text data and extracts emotional information, An alert generation means that generates warning information when specific conditions are met based on emotional information, A notification means that notifies relevant parties of warning information using a pre-configured communication method, A system that includes this.

2. The system according to claim 1, further comprising control means for instructing a specified action to be taken based on warning information.

3. The system according to claim 1, wherein the notification means transmits warning information using email or an online messaging platform.

Citation Information

Patent Citations

  • Persona chatbot control method and system

    JP2022180282A