system
An information processing device converts voice input to text, analyzes for harassment signs, and generates real-time warnings, effectively preventing unconscious harassment in communication settings.
Patent Information
- Authority / Receiving Office
- JP · JP
- Patent Type
- Applications
- Current Assignee / Owner
- SOFTBANK GROUP CORP
- Filing Date
- 2024-12-09
- Publication Date
- 2026-06-19
Smart Images

Figure 2026100721000001_ABST
Abstract
Description
Technical Field
[0001] The technology of the present disclosure relates to a system.
Background Art
[0002] Patent Document 1 discloses a method for controlling a persona chatbot performed by at least one processor, the method including steps of receiving a user utterance, adding the user utterance to a prompt including an instruction sentence related to an explanation of a character of the chatbot, encoding the prompt, and inputting the encoded prompt into a language model to generate a chatbot utterance as a response to the user utterance.
Prior Art Documents
Patent Documents
[0003]
Patent Document 1
Summary of the Invention
Problems to be Solved by the Invention
[0004] In modern society, it is known that harassment has a profound impact on an individual's mental health and workplace environment. However, since many harassments are carried out unconsciously, the perpetrator may not notice it and escalate it. There is a need for an effective means to detect harassment in advance and prevent it.
Means for Solving the Problems
[0005] To detect signs of harassment, an information processing device is used that acquires voice input and converts the voice data into text data. The converted text is analyzed to identify features and phrases suggestive of harassment. Based on this analysis, a warning is generated and notified to the user visually or audibly, thereby encouraging awareness of harassing behavior. Furthermore, by cross-referencing with a database of past cases, the accuracy of the warning is improved, effectively resolving the problem.
[0006] "Voice input" is the process of acquiring sound from an external environment through an input device such as a microphone.
[0007] An "information processing device" is an electronic device capable of acquiring, processing, analyzing, and outputting data.
[0008] "Audio data" refers to data in which audio information acquired through voice input is stored in digital format.
[0009] "Text data" refers to information obtained by converting audio data into written text.
[0010] "Conversion means" refers to a technology or device for converting audio data into text data.
[0011] "Analysis means" refers to a technology or device for processing text data, understanding its meaning and content, and detecting specific features or signs.
[0012] "Signs of harassment" are expressions or patterns, in terms of language and context, that may cause discomfort or psychological pressure to others.
[0013] A "warning" is a message or notification that alerts a user when signs of harassment are detected.
[0014] "Dissemination means" refers to technology or devices for generating warnings and notifying users.
[0015] "Visual notification" is a method of conveying information to the user using screen displays and lights.
[0016] "Auditory notification" is a method of conveying information to the user using voice and sound signals.
[0017] "Matching" is a process of comparing the analyzed data with past stored cases to determine similarity.
[0018] "Database" is a system that stores a collection of data organized according to a specific purpose.
Brief Description of the Drawings
[0019] [Figure 1] It is a conceptual diagram showing an example of the configuration of a data processing system according to the first embodiment. [Figure 2] It is a conceptual diagram showing an example of the main functions of a data processing device and a smart device according to the first embodiment. [Figure 3] It is a conceptual diagram showing an example of the configuration of a data processing system according to the second embodiment. [Figure 4] It is a conceptual diagram showing an example of the main functions of a data processing device and smart glasses according to the second embodiment. [Figure 5] It is a conceptual diagram showing an example of the configuration of a data processing system according to the third embodiment. [Figure 6] It is a conceptual diagram showing an example of the main functions of a data processing device and a headset-type terminal according to the third embodiment. [Figure 7] It is a conceptual diagram showing an example of the configuration of a data processing system according to the fourth embodiment. [Figure 8] It is a conceptual diagram showing an example of the main functions of a data processing device and a robot according to the fourth embodiment. [Figure 9] It shows an emotion map to which multiple emotions are mapped. [Figure 10] It shows an emotion map to which multiple emotions are mapped. [Figure 11] It is a sequence diagram showing the processing flow of the data processing system in Embodiment 1. [Figure 12] It is a sequence diagram showing the processing flow of the data processing system in Application Example 1. [Figure 13] It is a sequence diagram showing the processing flow of the data processing system in Embodiment 2 when the emotion engine is combined. [Figure 14] It is a sequence diagram showing the processing flow of the data processing system in Application Example 2 when the emotion engine is combined.
Mode for Carrying Out the Invention
[0020] Hereinafter, an example of an embodiment of a system according to the technology of the present disclosure will be described with reference to the accompanying drawings.
[0021] First, the terms used in the following description will be explained.
[0022] In the following embodiments, a numbered processor (hereinafter simply referred to as "processor") may be a single arithmetic unit or a combination of multiple arithmetic units. Also, the processor may be a single type of arithmetic unit or a combination of multiple types of arithmetic units. Examples of arithmetic units include CPU (Central Processing Unit), GPU (Graphics Processing Unit), GPGPU (General-Purpose computing on Graphics Processing Units), APU (Accelerated Processing Unit), etc.
[0023] In the following embodiments, a numbered RAM (Random Access Memory) is a memory in which information is temporarily stored and is used as a work memory by the processor.
[0024] In the following embodiments, the signed storage is one or more non-volatile storage devices that store various programs and various parameters. Examples of non-volatile storage devices include flash memory (SSD (Solid State Drive)), magnetic disks (e.g., hard disks), or magnetic tapes.
[0025] In the following embodiments, the signed communication interface (I / F) is an interface that includes a communication processor and an antenna, etc. The communication interface manages communication between multiple computers. Examples of communication standards applicable to the communication interface include wireless communication standards such as 5G (5th Generation Mobile Communication System), Wi-Fi (registered trademark), or Bluetooth (registered trademark).
[0026] In the following embodiments, "A and / or B" is synonymous with "at least one of A and B." That is, "A and / or B" means that it may be A alone, or B alone, or a combination of A and B. Furthermore, in this specification, the same concept as "A and / or B" applies when expressing three or more things linked by "and / or."
[0027] [First Embodiment]
[0028] Figure 1 shows an example of the configuration of the data processing system 10 according to the first embodiment.
[0029] As shown in Figure 1, the data processing system 10 includes a data processing device 12 and a smart device 14. An example of the data processing device 12 is a server.
[0030] The data processing device 12 comprises a computer 22, a database 24, and a communication interface 26. The computer 22 is an example of a "computer" related to the technology of this disclosure. The computer 22 comprises a processor 28, RAM 30, and storage 32. The processor 28, RAM 30, and storage 32 are connected to a bus 34. The database 24 and the communication interface 26 are also connected to the bus 34. The communication interface 26 is connected to a network 54. An example of the network 54 is a WAN (Wide Area Network) and / or a LAN (Local Area Network).
[0031] The smart device 14 comprises a computer 36, a reception device 38, an output device 40, a camera 42, and a communication interface 44. The computer 36 comprises a processor 46, RAM 48, and storage 50. The processor 46, RAM 48, and storage 50 are connected to a bus 52. The reception device 38, output device 40, and camera 42 are also connected to the bus 52.
[0032] The reception device 38 is equipped with a touch panel 38A and a microphone 38B, etc., and receives user input. The touch panel 38A receives user input by detecting contact with an object (e.g., a pen or finger). The microphone 38B receives user input by detecting the user's voice. The control unit 46A transmits data indicating the user input received by the touch panel 38A and microphone 38B to the data processing device 12. In the data processing device 12, the specific processing unit 290 acquires the data indicating the user input.
[0033] The output device 40 includes a display 40A and a speaker 40B, and presents data to the user 20 by outputting the data in a form perceptible to the user 20 (e.g., audio and / or text). The display 40A displays visible information such as text and images according to instructions from the processor 46. The speaker 40B outputs audio according to instructions from the processor 46. The camera 42 is a small digital camera equipped with an optical system such as a lens, aperture, and shutter, and an image sensor such as a CMOS (Complementary Metal-Oxide-Semiconductor) image sensor or a CCD (Charge Coupled Device) image sensor.
[0034] Communication interface 44 is connected to network 54. Communication interfaces 44 and 26 are responsible for the exchange of various types of information between processor 46 and processor 28 via network 54.
[0035] Figure 2 shows an example of the main functions of the data processing device 12 and the smart device 14.
[0036] As shown in Figure 2, in the data processing device 12, a specific processing is performed by the processor 28. A specific processing program 56 is stored in the storage 32. The specific processing program 56 is an example of a "program" related to the technology of this disclosure. The processor 28 reads the specific processing program 56 from the storage 32 and executes the read specific processing program 56 on the RAM 30. The specific processing is realized by the processor 28 operating as a specific processing unit 290 according to the specific processing program 56 executed on the RAM 30.
[0037] The storage 32 stores the data generation model 58 and the emotion identification model 59. The data generation model 58 and the emotion identification model 59 are used by the identification processing unit 290.
[0038] In the smart device 14, the processor 46 performs the reception output processing. The storage 50 stores the reception output program 60. The reception output program 60 is used in conjunction with a specific processing program 56 by the data processing system 10. The processor 46 reads the reception output program 60 from the storage 50 and executes the read reception output program 60 on the RAM 48. The reception output processing is realized by the processor 46 operating as a control unit 46A according to the reception output program 60 executed on the RAM 48.
[0039] Next, the specific processing performed by the specific processing unit 290 of the data processing device 12 will be described. In the following description, the data processing device 12 will be referred to as the "server" and the smart device 14 as the "terminal".
[0040] This invention provides a system that acquires voice input, detects signs of harassment, and notifies the user of a warning. Specific embodiments thereof are described below.
[0041] The device acquires audio during conversations via its microphone and collects audio data in real time. This input audio data is converted into text data using speech recognition technology built into the device. In this way, the content spoken is organized into written information.
[0042] Next, the server receives the converted text data and performs analysis using natural language processing (NLP) techniques. During the analysis process, it compares the text with a database of harassment cases collected in the past to determine if there are any phrases or patterns in the text that are considered harassment.
[0043] If the analysis results indicate signs of harassment, the server immediately generates a warning and sends it to the terminal. This warning includes an explanation of the specific phrase or context in question and is communicated to the user visually or audibly. This allows the user to receive real-time feedback and reflect on and improve their words and actions.
[0044] As a concrete example, consider a scenario in a workplace meeting. If a supervisor begins using harsh language towards a subordinate, the device immediately detects this and records it as text. The server analyzes this text to determine if it matches past case data. If it is determined that there is a risk of harassment, an alert is sent to the supervisor's device. This process can prevent harassment that may occur unintentionally.
[0045] This invention is expected to improve the quality of communication in the workplace and society as a whole, as harassment, whether conscious or unconscious, can be quickly recognized and dealt with appropriately.
[0046] The following describes the processing flow.
[0047] Step 1:
[0048] The device acquires audio from the conversation via the microphone. The audio data is temporarily stored in a buffer in real time.
[0049] Step 2:
[0050] The device uses speech recognition technology to convert the acquired speech data into text data. At this stage, noise reduction is performed to minimize background noise in the conversation.
[0051] Step 3:
[0052] The terminal cleans the converted text, removing unnecessary spaces and symbols. This text data is then divided into phrases or sentences.
[0053] Step 4:
[0054] The server receives text data and begins analysis using natural language processing (NLP) techniques. This analysis includes keyword research and sentiment analysis.
[0055] Step 5:
[0056] The server analyzes the text and compares it with a database of harassment cases. Here, it checks whether the statements in the text are similar to those in past harassment cases.
[0057] Step 6:
[0058] If the server determines there is a risk of harassment, it generates an alert and sends it to the terminal. The alert includes information about the problematic statement and its impact.
[0059] Step 7:
[0060] The device visually or audibly notifies the user of alerts it has received. This allows the user to re-evaluate and improve their statements.
[0061] (Example 1)
[0062] Next, we will describe Example 1. In the following description, the data processing device 12 will be referred to as the "server," and the smart device 14 will be referred to as the "terminal."
[0063] There is a need for a system that can quickly detect harassment that occurs unconsciously in the workplace and other communication settings, and provide advance warnings, thereby preventing conflicts and discomfort among those involved.
[0064] The identification process performed by the identification processing unit 290 of the data processing device 12 in Example 1 is realized by the following means.
[0065] In this invention, the server includes means for acquiring voice input, means for converting the acquired voice data into text information, and means for analyzing the converted text information to detect signs of harassment. This makes it possible to quickly recognize signs of harassment and issue a warning to the user.
[0066] "Voice input" is the process of acquiring conversations or voice information through devices such as microphones.
[0067] A "terminal" is an electronic device that receives voice input and performs the necessary processing.
[0068] "Textual information" refers to data obtained by converting audio data into text format.
[0069] "Means of conversion" refers to the technology or device used to convert audio data into text information.
[0070] "Means of analysis" refer to techniques and functions for examining textual information and identifying specific patterns or signs.
[0071] "Signs of harassment" refer to phrases or behavioral patterns that may cause discomfort or threat to others.
[0072] A "processing device" is a system that includes computers and servers for analyzing speech and text information.
[0073] "Means of generating warnings and notifying users" refers to a function that automatically creates a message based on detected signs of harassment and notifies the user visually or audibly.
[0074] A "database" is an information system that organizes past cases and data to facilitate searching and matching.
[0075] "Visual or auditory notification" refers to a method of conveying warning information to the user through screen displays or audio output.
[0076] This invention relates to a system that acquires voice input and detects signs of harassment. The system consists of a terminal, a server, and communication between the two.
[0077] The device acquires the speaker's voice through the microphone and collects it as audio data. For speech recognition, it is necessary to run speech recognition software suitable for the device. Examples include Google® Speech-to-Text API and IBM Watson® Speech to Text. This software is used to convert the collected audio data into text information.
[0078] The converted text information is sent to a server and analyzed using natural language processing techniques. This analysis includes extracting words and phrases from the text information and verifying whether they match patterns related to harassment registered in a database of past cases. Suitable software for this purpose on the server includes Python's NLTK and spaCy.
[0079] Based on the analysis results, the server immediately generates a warning if signs of harassment are detected. The warning message includes information about the problematic content of the statement and how to improve it, and is sent to the terminal. The terminal presents the received warning to the user through visual messages and audio notifications. This allows the user to receive real-time feedback and reflect on and improve their statements and actions.
[0080] As a concrete example, consider a workplace meeting scenario. If a supervisor uses harsh language towards a subordinate, the terminal instantly captures the words and converts them into text. The server analyzes this text, compares it against a past database, and assesses the harassment risk. A warning message is then generated and sent to the supervisor's terminal. This procedure allows for the prevention of harassment in advance.
[0081] Examples of input prompts for a generative AI model:
[0082] "Please explain the specific implementation procedures for a harassment detection system that uses audio data."
[0083] The flow of the specific processing in Example 1 will be explained using Figure 11.
[0084] Step 1:
[0085] The device uses a microphone to acquire voice input and collects it as audio data. The input is raw speech during a conversation, and the output is audio data in digital format. This data is temporarily stored within the device.
[0086] Step 2:
[0087] The terminal uses speech recognition software to convert speech data into text information. The input is the speech data obtained in the previous step, and the output is text information (text data). Specifically, the speech recognition software analyzes the speech waveform and generates the corresponding text information.
[0088] Step 3:
[0089] The server receives character information transmitted from the terminal. This character information is used as input and stored as preparation for further analysis. The output is character information stored on the server in a parseable format.
[0090] Step 4:
[0091] The server uses natural language processing techniques to analyze textual information and detect signs of harassment. The input is the textual information received in the previous step, and the output is the result of the harassment risk assessment. Specifically, the server compares the textual information with past cases stored in the database to determine whether or not there is a risk.
[0092] Step 5:
[0093] If the server detects signs of harassment, it generates a warning message. The input is the result of the risk assessment obtained in step 4, and the output is a warning message to notify the user. This warning includes a description of the specific problem phrase or context.
[0094] Step 6:
[0095] The terminal receives warning messages sent from the server and notifies the user visually or audibly. The input is the warning message from the server, and the output is a warning display or audio notification to the user. Specifically, the terminal provides feedback to the user by displaying a message on the screen or emitting an audio alert.
[0096] (Application Example 1)
[0097] Next, we will explain Application Example 1. In the following explanation, the data processing device 12 will be referred to as the "server," and the smart device 14 will be referred to as the "terminal."
[0098] There is a need to prevent harassment in internal communications by detecting conversations containing signs of harassment in real time and immediately sending warnings to those involved. However, current systems have difficulty with real-time detection and immediate feedback provision, and there is room for improvement.
[0099] The specific processing performed by the specific processing unit 290 of the data processing device 12 in Application Example 1 is realized by the following means.
[0100] In this invention, the server includes an information processing device for acquiring voice input, a conversion means for converting the acquired voice data into text data, an analysis means for analyzing the text data and detecting signs of harassment, and a means for notifying the user of the generated warning in real time and providing feedback to improve behavior. This makes it possible to detect signs of harassment in a conversation in real time and immediately issue a warning to the relevant parties.
[0101] An "information processing device for acquiring voice input" is a device that captures external voice signals, thereby enabling it to capture human speech as digital data.
[0102] A "conversion method for converting to text data" refers to a method that has the function of converting acquired audio data into text information using speech recognition technology, and plays a role in making audio information visible.
[0103] An "analytical means for detecting signs of harassment" is a means that has the ability to identify words and patterns that suggest harassment by analyzing text data and comparing it with past cases.
[0104] A "warning generation and transmission device" is a device that has the function of communicating warnings to relevant parties based on detected signs of harassment, thereby providing users with immediate feedback.
[0105] A "means of providing real-time notifications and feedback to improve behavior" refers to a means of immediately notifying users based on detected signs, providing specific improvement suggestions, and supporting user behavior.
[0106] The system to realize this application first involves the user acquiring audio using a device such as a smartphone or personal computer. The microphone in the device collects the audio of the conversation and digitizes the data. This digitized audio data is then converted into text data using speech recognition software such as the Google Speech-to-Text API.
[0107] The server receives this text data and analyzes it using natural language processing libraries such as Apache® OpenNLP. During the analysis process, the acquired text data is compared against a database of past harassment cases to check for the presence of specific phrases or patterns. If signs of harassment are detected, a generative AI model is used to generate a warning message for the user.
[0108] Warnings are sent to the device in real time. This allows users to receive immediate feedback to improve their words and actions. Notifications are sent to the device visually or audibly using communication tools such as the Twilio API.
[0109] As a concrete example, consider a scenario where a user is using the application during a workplace meeting. In this situation, if a supervisor makes a negative comment to a subordinate, the audio is immediately transcribed and analyzed. If signs of harassment are detected, a warning message such as "The current statement may be harassment" will appear on the supervisor's device, providing an opportunity for them to recognize areas for improvement.
[0110] Examples of prompts to input into a generative AI model:
[0111] "Identify any remarks in this conversation that could be considered harassment and issue a warning to the user."
[0112] The flow of a specific process in Application Example 1 will be explained using Figure 12.
[0113] Step 1:
[0114] The device acquires voice input. The user launches an application on a device such as a smartphone or personal computer and captures physical voice data through the microphone. The input is an analog audio signal and is output as digital audio data.
[0115] Step 2:
[0116] The device converts the audio data into text data. The acquired digital audio data is then converted into text format using the Google Speech-to-Text API. The input is digital audio data, and the output is natural language text data.
[0117] Step 3:
[0118] The server analyzes the text data. The server uses natural language processing libraries such as Apache OpenNLP to analyze the text data. The input is text data, and the server recognizes phrases and patterns that may indicate harassment. The output is harassment detection information as a result of the analysis.
[0119] Step 4:
[0120] The server generates warnings using a generative AI model. Based on the analyzed data, the generative AI model is used to generate warning content. The input is harassment detection information, and a warning message is output to notify the user.
[0121] Step 5:
[0122] The server sends a warning to the device. The generated warning message is sent to the target device using a communication service such as the Twilio API. The input is a warning message, and the output is a notification that the user receives the warning visually or audibly.
[0123] Step 6:
[0124] The user receives feedback. The user reviews the warnings displayed on their device and receives feedback on how to improve their behavior. The input is in the form of visual and auditory notifications, and the output is provided as feedback for behavioral improvement.
[0125] Furthermore, an emotion engine that estimates the user's emotions may be incorporated. That is, the identification processing unit 290 may use the emotion identification model 59 to estimate the user's emotions and perform identification processing using the user's emotions.
[0126] This invention provides a system that combines voice and emotion recognition functions to more precisely detect signs of harassment and notify the user of appropriate warnings.
[0127] First, the device acquires audio during a conversation via its microphone, collecting audio data in real time. This audio data is then converted into text data using the device's built-in speech recognition technology. This text data undergoes a series of preprocessing steps before being prepared for analysis.
[0128] A key element here is the inclusion of an emotion engine in the device. This emotion engine recognizes and analyzes the user's emotional state from voice and text data. This allows for more than just text analysis; it also takes into account the speaker's emotional tone and intent. The emotion engine determines which emotion the user is expressing—such as anger, sadness, or joy—and provides that emotional information to the analysis system.
[0129] Next, the server receives the text data and accompanying sentiment information and performs analysis using natural language processing (NLP) techniques. The analysis identifies phrases and patterns in the text that may be considered harassment and further compares them with a database of past cases. Using the sentiment information obtained from the sentiment engine, the server evaluates the severity of the detected signs and applies it to generate warnings.
[0130] If a warning is deemed necessary, the server generates a warning based on emotional information and sends it to the terminal. This warning is communicated to the user through visual and auditory notification methods. The user can use this real-time feedback to adjust their words and actions and improve their behavior to prevent harassment.
[0131] As a concrete example, consider a conversation between a supervisor and a subordinate in the workplace. If the supervisor makes a statement expressing frustration towards the subordinate, the terminal's emotion engine recognizes the emotion of frustration from the supervisor's voice and transmits that information to the server. The server takes this emotion information, adjusts the severity of the warning, and then sends feedback to the supervisor that includes specific areas for improvement. By combining emotion information in this way, more accurate and appropriate harassment prevention measures become possible.
[0132] The following describes the processing flow.
[0133] Step 1:
[0134] The device uses its microphone to capture audio during a conversation in real time. The captured audio data is clarified through noise filtering and temporarily stored in a buffer.
[0135] Step 2:
[0136] The device uses speech recognition technology to convert the audio data in the buffer into text data. This conversion uses an algorithm that turns words and phrases in the audio into characters.
[0137] Step 3:
[0138] The device's emotion engine analyzes the user's emotional state from voice data. The analysis considers factors such as voice tone, pitch, and speed to identify emotions like joy and anger.
[0139] Step 4:
[0140] Text data and sentiment information are sent from the terminal to the server. The server receives this data and prepares it for analysis.
[0141] Step 5:
[0142] The server uses natural language processing (NLP) techniques to analyze text data. This analysis includes detecting specific keywords, understanding contextual meaning, and identifying signs of harassment.
[0143] Step 6:
[0144] The server takes in emotional information provided by the emotion engine and generates warnings based on the analysis results. Emotional information is a crucial factor in determining the severity of the warning and how it should be presented.
[0145] Step 7:
[0146] The server sends a generated warning to the terminal. The warning is adjusted to take into account the user's emotional state and the content of their statements.
[0147] Step 8:
[0148] The device notifies the user of any warnings it has received. The notification is delivered visually or audibly, prompting the user to re-evaluate their words and actions.
[0149] (Example 2)
[0150] Next, we will describe Example 2. In the following description, the data processing device 12 will be referred to as the "server" and the smart device 14 as the "terminal".
[0151] In recent years, harassment has become a serious problem in the workplace and society, requiring immediate and appropriate responses. However, conventional systems are limited to text analysis and have the drawback of not being able to adequately consider the speaker's emotions and intentions. As a result, there is a risk that appropriate warnings may not be issued due to incorrect judgments, and problems may be overlooked. To solve this problem, a new system is needed that can perform more sophisticated analysis using voice and emotional information.
[0152] The identification process performed by the identification processing unit 290 of the data processing device 12 in Example 2 is realized by the following means.
[0153] In this invention, the server includes terminal means for acquiring voice input, voice recognition means for converting the acquired voice data into text data, and means for performing analysis using the converted text data in addition to emotional information recognized from the voice. This makes it possible to more accurately detect signs of harassment and to quickly notify the user of appropriate warnings.
[0154] "Voice input" is the process of acquiring acoustic information emitted by the user through a device such as a microphone.
[0155] A "terminal device" is an information processing device that acquires audio and processes and transmits data.
[0156] "Audio data" refers to information recorded in digital format from collected audio input.
[0157] "Speech recognition means" refers to technology that analyzes speech data and converts it into corresponding text data.
[0158] "Text data" refers to character information converted by speech recognition technology.
[0159] "Emotional information" refers to information extracted from audio and text that indicates the speaker's emotional state.
[0160] "Means of analysis" refers to the process of using text data and sentiment information to determine signs of harassment.
[0161] "Signs of harassment" are patterns or phrases that may indicate potential aggression or harassment in one's words or actions.
[0162] "Means of generating warnings" refers to the process of creating notifications that prompt users to take action based on detected signs of harassment.
[0163] "Notification output means" refers to technologies for conveying warnings to users visually or audibly.
[0164] This invention provides a system that utilizes speech recognition and sentiment analysis technology to detect signs of harassment and generate appropriate warnings. This enables users to maintain a healthier environment in their daily communication.
[0165] First, the user's device acquires conversations and voice input through a high-performance microphone. This device incorporates general-purpose voice processing software as a voice recognition technology. For example, a specific service as a "voice recognition means" employs technology to instantly convert voice data into text data. Noise reduction and text organization are performed at this stage to improve the accuracy of subsequent analysis.
[0166] Next, the emotion engine installed in the device analyzes the emotional information from this text and audio data. The emotion engine uses emotion analysis software to determine the emotional state from the tone of voice and the content of the text. This not only analyzes the words themselves, but also detects what emotions the speaker is feeling while speaking, and sends that information to the server.
[0167] The server receives the transmitted text data and sentiment information and performs analysis using natural language processing (NLP) techniques. Widely used NLP libraries are utilized for information processing for the analysis. Based on the analysis results, the server identifies signs of harassment and evaluates the severity of those signs by cross-referencing them with the sentiment information.
[0168] If necessary, the server generates and sends a warning prompting appropriate action to the terminal. The warning is communicated to the user visually or audibly, helping to create a comfortable and safe communication environment. A concrete example of this process is a conversation between a supervisor and an employee in the workplace. If the supervisor makes an irritated remark, the emotion engine detects that emotion, and the server provides appropriate feedback to help improve the supervisor's communication style.
[0169] Furthermore, a concrete example of a prompt using a generative AI model might be something like, "Please tell me how to generate appropriate feedback when a user becomes emotionally agitated." By combining emotional information with harassment detection in this way, more accurate responses become possible.
[0170] The flow of the specific processing in Example 2 will be explained using Figure 13.
[0171] Step 1:
[0172] The device acquires the user's voice through a high-performance microphone. The input is real-time audio data, which is converted into a digital format to prepare for subsequent processing. This digital audio data is then processed to remove background noise and format it for analysis.
[0173] Step 2:
[0174] The device uses speech recognition technology to convert audio data into text data. The input is pre-processed audio data, and the output is text data in string format. This conversion process analyzes speech patterns using acoustic and language models and generates the corresponding text.
[0175] Step 3:
[0176] The device generates emotion information from text data and speech characteristics. The input is the text data and speech tone information obtained in step 2, and the output is emotion information indicating the user's emotional state. The emotion engine analyzes the content of the text and speech characteristics such as emphasis and speed to identify the emotion the user is expressing.
[0177] Step 4:
[0178] The terminal sends the generated text data and sentiment information to the server. The input consists of text data and sentiment information, which are securely transmitted to the server using a network communication protocol. This step enables subsequent data analysis on the server.
[0179] Step 5:
[0180] The server receives the transmitted data and performs analysis using natural language processing techniques. The input consists of text data and sentiment information, and the output is an evaluation of potential signs of harassment. Here, the server extracts specific phrases and patterns from the text and combines them with sentiment information to assess their severity.
[0181] Step 6:
[0182] The server generates a warning based on the analysis results and sends it to the terminal. The input is the analysis results, and the output is the warning message that is sent to the user. This process creates a warning that highlights the areas that need improvement for the user.
[0183] Step 7:
[0184] The terminal notifies the user of any received warnings. The input is a warning message sent from the server, which is presented to the user using visual or auditory means. Based on this feedback, the user can review their words and actions and adjust their behavior accordingly.
[0185] (Application Example 2)
[0186] Next, we will explain application example 2. In the following explanation, the data processing device 12 will be referred to as the "server," and the smart device 14 will be referred to as the "terminal."
[0187] In various settings such as workplaces and public facilities, it is currently difficult to detect signs of harassment, and systems for preventing it are not sufficiently developed. Conventional systems mainly analyze only voice and text data, and because they cannot consider emotional nuances or context, they can make incorrect judgments. This invention aims to solve these problems and achieve more accurate harassment detection and notification.
[0188] The specific processing performed by the specific processing unit 290 of the data processing device 12 in Application Example 2 is realized by the following means.
[0189] In this invention, the server includes means for using a device to acquire voice input, means for converting the acquired voice data into text data, means for analyzing the text data and sentiment data extracted from the voice to detect signs of harassment, means for generating a warning based on the signs and sentiment data detected by the analysis means, and means for notifying the user of the warning via a visual interface. This enables advanced automated harassment detection and real-time notification combined with sentiment analysis.
[0190] "Voice input" refers to information data obtained through a device using human speech.
[0191] A "device" is an electronic device used to acquire and process audio and other data.
[0192] "Text data" refers to string data converted based on voice input.
[0193] "Emotional data" refers to information indicating emotional states, analyzed from audio and text data.
[0194] "Analysis methods" refer to processes and algorithms that detect specific patterns or emotions based on acquired data.
[0195] A "warning" is a message that notifies users of risks or situations requiring attention, based on observed data.
[0196] A "visual interface" is a means of displaying information to a user in a visual way.
[0197] To realize this invention, a system including a server and a terminal is required. The terminal is equipped with a microphone as a device for acquiring voice input and collects voice data in real time during conversations. Using speech recognition technology, this voice data is converted into text data, and further emotion data is extracted from the voice and text data using an emotion engine. The emotion engine implements algorithms for determining the tone of voice and the emotional state of the speaker.
[0198] The server receives this text and sentiment data and uses natural language processing (NLP) techniques to analyze it. The analysis includes procedures to identify potential signs of harassment within the text data and, if necessary, has the ability to cross-reference it with a database of past cases. Based on this, a warning is generated according to the detected signs and sentiment data. This warning is notified to the user via the terminal and presented using a visual interface.
[0199] The hardware and software used include Google Cloud Speech-to-Text for speech recognition, IBM Watson Tone Analyzer for sentiment analysis, and libraries such as spaCy and BERT for natural language processing.
[0200] As a concrete example, if a participant makes a statement that indicates strong stress towards another participant during a workplace discussion, the device will capture the audio, and the emotion engine will recognize the emotion of stress. The server will analyze this data, generate a warning, and provide feedback to the user through a visual interface.
[0201] An example of a prompt to input into a generative AI model is: "Design a system that analyzes voice and sentiment data to identify potential signs of harassment in workplace conversations and provides real-time notifications."
[0202] The flow of a specific process in Application Example 2 will be explained using Figure 14.
[0203] Step 1:
[0204] The device acquires voice input through the microphone. The acquired voice is transmitted to the server in real time as digital audio data. At this stage, the input is human speech, and the output is a digitized audio file.
[0205] Step 2:
[0206] The server uses a speech recognition engine (e.g., Google Cloud Speech-to-Text) to convert the received audio data into text data. The input here is audio data, and the speech recognition process generates text data as output.
[0207] Step 3:
[0208] The server extracts sentiment data from the generated text data using a sentiment engine (e.g., IBM Watson Tone Analyzer). The input is text data, and after the analysis process, sentiment data indicating the user's emotional state is output.
[0209] Step 4:
[0210] The server uses natural language processing (NLP) techniques (e.g., spaCy, BERT) to analyze text data and detect signs of harassment. The input consists of text data and sentiment data, and signs are identified based on the NLP algorithm. The output is data indicating the presence or absence of signs and their content.
[0211] Step 5:
[0212] The server generates warnings based on the analysis results, cross-referencing them with a database of past cases as needed. Here, the inputs are the analysis results and the database of past cases, and the output is a appropriately adjusted warning message.
[0213] Step 6:
[0214] The terminal notifies the user of generated warnings through a visual interface. It communicates problems to the user in real time by visually displaying warning data received from the server. The input is warning data, and the output is visual feedback to the user.
[0215] The specific processing unit 290 transmits the result of the specific processing to the smart device 14. In the smart device 14, the control unit 46A causes the output device 40 to output the result of the specific processing. The microphone 38B acquires audio indicating user input for the result of the specific processing. The control unit 46A transmits the audio data indicating user input acquired by the microphone 38B to the data processing device 12. In the data processing device 12, the specific processing unit 290 acquires the audio data.
[0216] Data generation model 58 is a so-called generative AI (Artificial Intelligence). An example of data generation model 58 is ChatGPT (registered trademark) (Internet search).<URL: https: / / openai.com / blog / chatgpt> ), Gemini (registered trademark) (Internet search) <url: https: gemini.google.com ?hl="ja">Examples of generative AI include the following. The data generation model 58 is obtained by performing deep learning on a neural network. The data generation model 58 is input with prompts containing instructions, and with inference data such as audio data representing speech, text data representing text, and image data representing images. The data generation model 58 infers from the input inference data according to the instructions indicated by the prompts, and outputs the inference results in data formats such as audio data and text data. Here, inference refers to, for example, analysis, classification, prediction, and / or summarization.
[0217] In the above embodiment, an example was given in which specific processing is performed by the data processing device 12, but the technology of this disclosure is not limited thereto, and the specific processing may also be performed by the smart device 14.
[0218] [Second Embodiment]
[0219] Figure 3 shows an example of the configuration of the data processing system 210 according to the second embodiment.
[0220] As shown in Figure 3, the data processing system 210 includes a data processing device 12 and smart glasses 214. An example of the data processing device 12 is a server.
[0221] The data processing device 12 comprises a computer 22, a database 24, and a communication interface 26. The computer 22 is an example of a "computer" related to the technology of this disclosure. The computer 22 comprises a processor 28, RAM 30, and storage 32. The processor 28, RAM 30, and storage 32 are connected to a bus 34. The database 24 and the communication interface 26 are also connected to the bus 34. The communication interface 26 is connected to a network 54. An example of the network 54 is a WAN (Wide Area Network) and / or a LAN (Local Area Network).
[0222] The smart glasses 214 include a computer 36, a microphone 238, a speaker 240, a camera 42, and a communication interface 44. The computer 36 includes a processor 46, RAM 48, and storage 50. The processor 46, RAM 48, and storage 50 are connected to a bus 52. The microphone 238, speaker 240, and camera 42 are also connected to the bus 52.
[0223] The microphone 238 receives voice signals from the user 20 and receives instructions from the user 20. The microphone 238 captures the voice signals from the user 20, converts the captured voice into audio data, and outputs it to the processor 46. The speaker 240 outputs audio according to the instructions from the processor 46.
[0224] Camera 42 is a small digital camera equipped with an optical system including a lens, aperture, and shutter, and an image sensor such as a CMOS (Complementary Metal-Oxide-Semiconductor) image sensor or a CCD (Charge Coupled Device) image sensor, and captures images of the area around the user 20 (for example, an imaging range defined by a field of view equivalent to the width of a typical healthy person's field of vision).
[0225] Communication interface 44 is connected to network 54. Communication interfaces 44 and 26 are responsible for the exchange of various information between processor 46 and processor 28 via network 54. The exchange of various information between processor 46 and processor 28 using communication interfaces 44 and 26 is performed in a secure manner.
[0226] Figure 4 shows an example of the main functions of the data processing device 12 and the smart glasses 214. As shown in Figure 4, the data processing device 12 performs specific processing using the processor 28. The storage 32 stores the specific processing program 56.
[0227] The specific processing program 56 is an example of a "program" relating to the technology of this disclosure. The processor 28 reads the specific processing program 56 from the storage 32 and executes the read specific processing program 56 on the RAM 30. The specific processing is realized by the processor 28 operating as a specific processing unit 290 in accordance with the specific processing program 56 executed on the RAM 30.
[0228] The storage 32 stores the data generation model 58 and the emotion identification model 59. The data generation model 58 and the emotion identification model 59 are used by the identification processing unit 290.
[0229] In the smart glasses 214, the processor 46 performs the reception output processing. The storage 50 stores the reception output program 60. The processor 46 reads the reception output program 60 from the storage 50 and executes the read reception output program 60 on the RAM 48. The reception output processing is realized by the processor 46 operating as a control unit 46A according to the reception output program 60 executed on the RAM 48.
[0230] Next, the identification processing performed by the identification processing unit 290 of the data processing device 12 will be described. In the following description, the data processing device 12 will be referred to as the "server" and the smart glasses 214 will be referred to as the "terminal".
[0231] This invention provides a system that acquires voice input, detects signs of harassment, and notifies the user of a warning. Specific embodiments thereof are described below.
[0232] The device acquires audio during conversations via its microphone and collects audio data in real time. This input audio data is converted into text data using speech recognition technology built into the device. In this way, the content spoken is organized into written information.
[0233] Next, the server receives the converted text data and performs analysis using natural language processing (NLP) techniques. During the analysis process, it compares the text with a database of harassment cases collected in the past to determine if there are any phrases or patterns in the text that are considered harassment.
[0234] If the analysis results indicate signs of harassment, the server immediately generates a warning and sends it to the terminal. This warning includes an explanation of the specific phrase or context in question and is communicated to the user visually or audibly. This allows the user to receive real-time feedback and reflect on and improve their words and actions.
[0235] As a concrete example, consider a scenario in a workplace meeting. If a supervisor begins using harsh language towards a subordinate, the device immediately detects this and records it as text. The server analyzes this text to determine if it matches past case data. If it is determined that there is a risk of harassment, an alert is sent to the supervisor's device. This process can prevent harassment that may occur unintentionally.
[0236] This invention is expected to improve the quality of communication in the workplace and society as a whole, as harassment, whether conscious or unconscious, can be quickly recognized and dealt with appropriately.
[0237] The following describes the processing flow.
[0238] Step 1:
[0239] The device acquires audio from the conversation via the microphone. The audio data is temporarily stored in a buffer in real time.
[0240] Step 2:
[0241] The device uses speech recognition technology to convert the acquired speech data into text data. At this stage, noise reduction is performed to minimize background noise in the conversation.
[0242] Step 3:
[0243] The terminal cleans the converted text, removing unnecessary spaces and symbols. This text data is then divided into phrases or sentences.
[0244] Step 4:
[0245] The server receives text data and begins analysis using natural language processing (NLP) techniques. This analysis includes keyword research and sentiment analysis.
[0246] Step 5:
[0247] The server analyzes the text and compares it with a database of harassment cases. Here, it checks whether the statements in the text are similar to those in past harassment cases.
[0248] Step 6:
[0249] If the server determines there is a risk of harassment, it generates an alert and sends it to the terminal. The alert includes information about the problematic statement and its impact.
[0250] Step 7:
[0251] The device visually or audibly notifies the user of alerts it has received. This allows the user to re-evaluate and improve their statements.
[0252] (Example 1)
[0253] Next, we will describe Example 1. In the following description, the data processing device 12 will be referred to as the "server," and the smart glasses 214 will be referred to as the "terminal."
[0254] There is a need for a system that can quickly detect harassment that occurs unconsciously in the workplace and other communication settings, and provide advance warnings, thereby preventing conflicts and discomfort among those involved.
[0255] The identification process performed by the identification processing unit 290 of the data processing device 12 in Example 1 is realized by the following means.
[0256] In this invention, the server includes means for acquiring voice input, means for converting the acquired voice data into text information, and means for analyzing the converted text information to detect signs of harassment. This makes it possible to quickly recognize signs of harassment and issue a warning to the user.
[0257] "Voice input" is the process of acquiring conversations or voice information through devices such as microphones.
[0258] A "terminal" is an electronic device that receives voice input and performs the necessary processing.
[0259] "Textual information" refers to data obtained by converting audio data into text format.
[0260] "Means of conversion" refers to the technology or device used to convert audio data into text information.
[0261] "Means of analysis" refer to techniques and functions for examining textual information and identifying specific patterns or signs.
[0262] "Signs of harassment" refer to phrases or behavioral patterns that may cause discomfort or threat to others.
[0263] A "processing device" is a system that includes computers and servers for analyzing speech and text information.
[0264] "Means of generating warnings and notifying users" refers to a function that automatically creates a message based on detected signs of harassment and notifies the user visually or audibly.
[0265] A "database" is an information system that organizes past cases and data to facilitate searching and matching.
[0266] "Visual or auditory notification" refers to a method of conveying warning information to the user through screen displays or audio output.
[0267] This invention relates to a system that acquires voice input and detects signs of harassment. The system consists of a terminal, a server, and communication between the two.
[0268] The device acquires the speaker's voice through the microphone and collects it as audio data. For speech recognition, the device needs to run appropriate speech recognition software. Examples include the Google Speech-to-Text API and IBM Watson Speech to Text. This software is then used to convert the collected audio data into text information.
[0269] The converted text information is sent to a server and analyzed using natural language processing techniques. This analysis includes extracting words and phrases from the text information and verifying whether they match patterns related to harassment registered in a database of past cases. Suitable software for this purpose on the server includes Python's NLTK and spaCy.
[0270] Based on the analysis results, the server immediately generates a warning if signs of harassment are detected. The warning message includes information about the problematic content of the statement and how to improve it, and is sent to the terminal. The terminal presents the received warning to the user through visual messages and audio notifications. This allows the user to receive real-time feedback and reflect on and improve their statements and actions.
[0271] As a concrete example, consider a workplace meeting scenario. If a supervisor uses harsh language towards a subordinate, the terminal instantly captures the words and converts them into text. The server analyzes this text, compares it against a past database, and assesses the harassment risk. A warning message is then generated and sent to the supervisor's terminal. This procedure allows for the prevention of harassment in advance.
[0272] Examples of input prompts for a generative AI model:
[0273] "Please explain the specific implementation procedures for a harassment detection system that uses audio data."
[0274] The flow of the specific processing in Example 1 will be explained using Figure 11.
[0275] Step 1:
[0276] The device uses a microphone to acquire voice input and collects it as audio data. The input is raw speech during a conversation, and the output is audio data in digital format. This data is temporarily stored within the device.
[0277] Step 2:
[0278] The terminal uses speech recognition software to convert speech data into text information. The input is the speech data obtained in the previous step, and the output is text information (text data). Specifically, the speech recognition software analyzes the speech waveform and generates the corresponding text information.
[0279] Step 3:
[0280] The server receives the character information sent from the terminal. This character information is adopted as input and stored as a preparation stage for further analysis. The output is the character information stored in the server in an analyzable form.
[0281] Step 4:
[0282] The server analyzes the character information using natural language processing technology to detect signs of harassment. The input is the character information received in the previous step, and the output is the risk assessment result of harassment. Specifically, the server compares the character information with past cases stored in the database to determine the presence or absence of risk.
[0283] Step 5:
[0284] When the server recognizes signs of harassment, it generates a warning message. The input is the result of the risk assessment obtained in Step 4, and the output is a warning message for notifying the user. This warning includes an explanation of specific problem phrases or contexts.
[0285] Step 6:
[0286] The terminal receives the warning message sent from the server and notifies the user visually or audibly. The input is the warning message from the server, and the output is a warning display or voice notification to the user. As a specific operation, the terminal provides feedback to the user by displaying the message on the screen or emitting a voice alert.
[0287] (Application Example 1)
[0288] Next, Application Example 1 will be described. In the following description, the data processing device 12 is referred to as the "server", and the smart glasses 214 are referred to as the "terminal".
[0289] There is a need to prevent harassment in internal communications by detecting conversations containing signs of harassment in real time and immediately sending warnings to those involved. However, current systems have difficulty with real-time detection and immediate feedback provision, and there is room for improvement.
[0290] The specific processing performed by the specific processing unit 290 of the data processing device 12 in Application Example 1 is realized by the following means.
[0291] In this invention, the server includes an information processing device for acquiring voice input, a conversion means for converting the acquired voice data into text data, an analysis means for analyzing the text data and detecting signs of harassment, and a means for notifying the user of the generated warning in real time and providing feedback to improve behavior. This makes it possible to detect signs of harassment in a conversation in real time and immediately issue a warning to the relevant parties.
[0292] An "information processing device for acquiring voice input" is a device that captures external voice signals, thereby enabling it to capture human speech as digital data.
[0293] A "conversion method for converting to text data" refers to a method that has the function of converting acquired audio data into text information using speech recognition technology, and plays a role in making audio information visible.
[0294] An "analytical means for detecting signs of harassment" is a means that has the ability to identify words and patterns that suggest harassment by analyzing text data and comparing it with past cases.
[0295] A "warning generation and transmission device" is a device that has the function of communicating warnings to relevant parties based on detected signs of harassment, thereby providing users with immediate feedback.
[0296] A "means of providing real-time notifications and feedback to improve behavior" refers to a means of immediately notifying users based on detected signs, providing specific improvement suggestions, and supporting user behavior.
[0297] The system to realize this application first involves the user acquiring audio using a device such as a smartphone or personal computer. The microphone in the device collects the audio of the conversation and digitizes the data. This digitized audio data is then converted into text data using speech recognition software such as the Google Speech-to-Text API.
[0298] The server receives this text data and analyzes it using natural language processing libraries such as Apache OpenNLP. During the analysis process, the acquired text data is compared against a database of past harassment cases to check for the presence of specific phrases or patterns. If signs of harassment are detected, a generative AI model is used to generate a warning message for the user.
[0299] Warnings are sent to the device in real time. This allows users to receive immediate feedback to improve their words and actions. Notifications are sent to the device visually or audibly using communication tools such as the Twilio API.
[0300] As a concrete example, consider a scenario where a user is using the application during a workplace meeting. In this situation, if a supervisor makes a negative comment to a subordinate, the audio is immediately transcribed and analyzed. If signs of harassment are detected, a warning message such as "The current statement may be harassment" will appear on the supervisor's device, providing an opportunity for them to recognize areas for improvement.
[0301] Examples of prompts to input into a generative AI model:
[0302] "Identify statements in this conversation that could be considered harassment and issue a warning to the user."
[0303] The flow of the specific process in Application Example 1 will be described using FIG. 12.
[0304] Step 1:
[0305] The terminal acquires voice input. The user starts an application on a device such as a smartphone or personal computer and captures physical voice data through the microphone. The input is voice as an analog signal and is output as digital voice data.
[0306] Step 2:
[0307] The terminal converts the voice data into text data. The acquired digital voice data is converted into text format using the Google Speech-to-Text API. The input is digital voice data and is output as text data in natural language.
[0308] Step 3:
[0309] The server analyzes the text data. The server uses a natural language processing library such as Apache OpenNLP to analyze the text data. The input is text data, and phrases or patterns that could be signs of harassment are recognized. The output is detection information of harassment as the analysis result.
[0310] Step 4:
[0311] The server creates a warning using the generative AI model. Based on the analyzed data, the generative AI model is utilized to generate the warning content. The input is the detection information of harassment, and a warning message for notifying the user is output.
[0312] Step 5:
[0313] The server sends a warning to the device. The generated warning message is sent to the target device using a communication service such as the Twilio API. The input is a warning message, and the output is a notification that the user receives the warning visually or audibly.
[0314] Step 6:
[0315] The user receives feedback. The user reviews the warnings displayed on their device and receives feedback on how to improve their behavior. The input is in the form of visual and auditory notifications, and the output is provided as feedback for behavioral improvement.
[0316] Furthermore, an emotion engine that estimates the user's emotions may be incorporated. That is, the identification processing unit 290 may use the emotion identification model 59 to estimate the user's emotions and perform identification processing using the user's emotions.
[0317] This invention provides a system that combines voice and emotion recognition functions to more precisely detect signs of harassment and notify the user of appropriate warnings.
[0318] First, the device acquires audio during a conversation via its microphone, collecting audio data in real time. This audio data is then converted into text data using the device's built-in speech recognition technology. This text data undergoes a series of preprocessing steps before being prepared for analysis.
[0319] A key element here is the inclusion of an emotion engine in the device. This emotion engine recognizes and analyzes the user's emotional state from voice and text data. This allows for more than just text analysis; it also takes into account the speaker's emotional tone and intent. The emotion engine determines which emotion the user is expressing—such as anger, sadness, or joy—and provides that emotional information to the analysis system.
[0320] Next, the server receives the text data and accompanying sentiment information and performs analysis using natural language processing (NLP) techniques. The analysis identifies phrases and patterns in the text that may be considered harassment and further compares them with a database of past cases. Using the sentiment information obtained from the sentiment engine, the server evaluates the severity of the detected signs and applies it to generate warnings.
[0321] If a warning is deemed necessary, the server generates a warning based on emotional information and sends it to the terminal. This warning is communicated to the user through visual and auditory notification methods. The user can use this real-time feedback to adjust their words and actions and improve their behavior to prevent harassment.
[0322] As a concrete example, consider a conversation between a supervisor and a subordinate in the workplace. If the supervisor makes a statement expressing frustration towards the subordinate, the terminal's emotion engine recognizes the emotion of frustration from the supervisor's voice and transmits that information to the server. The server takes this emotion information, adjusts the severity of the warning, and then sends feedback to the supervisor that includes specific areas for improvement. By combining emotion information in this way, more accurate and appropriate harassment prevention measures become possible.
[0323] The following describes the processing flow.
[0324] Step 1:
[0325] The device uses its microphone to capture audio during a conversation in real time. The captured audio data is clarified through noise filtering and temporarily stored in a buffer.
[0326] Step 2:
[0327] The device uses speech recognition technology to convert the audio data in the buffer into text data. This conversion uses an algorithm that turns words and phrases in the audio into characters.
[0328] Step 3:
[0329] The device's emotion engine analyzes the user's emotional state from voice data. The analysis considers factors such as voice tone, pitch, and speed to identify emotions like joy and anger.
[0330] Step 4:
[0331] Text data and sentiment information are sent from the terminal to the server. The server receives this data and prepares it for analysis.
[0332] Step 5:
[0333] The server uses natural language processing (NLP) techniques to analyze text data. This analysis includes detecting specific keywords, understanding contextual meaning, and identifying signs of harassment.
[0334] Step 6:
[0335] The server takes in emotional information provided by the emotion engine and generates warnings based on the analysis results. Emotional information is a crucial factor in determining the severity of the warning and how it should be presented.
[0336] Step 7:
[0337] The server sends a generated warning to the terminal. The warning is adjusted to take into account the user's emotional state and the content of their statements.
[0338] Step 8:
[0339] The device notifies the user of any warnings it has received. The notification is delivered visually or audibly, prompting the user to re-evaluate their words and actions.
[0340] (Example 2)
[0341] Next, we will describe Example 2. In the following description, the data processing device 12 will be referred to as the "server" and the smart glasses 214 will be referred to as the "terminal".
[0342] In recent years, harassment has become a serious problem in the workplace and society, requiring immediate and appropriate responses. However, conventional systems are limited to text analysis and have the drawback of not being able to adequately consider the speaker's emotions and intentions. As a result, there is a risk that appropriate warnings may not be issued due to incorrect judgments, and problems may be overlooked. To solve this problem, a new system is needed that can perform more sophisticated analysis using voice and emotional information.
[0343] The identification process performed by the identification processing unit 290 of the data processing device 12 in Example 2 is realized by the following means.
[0344] In this invention, the server includes terminal means for acquiring voice input, voice recognition means for converting the acquired voice data into text data, and means for performing analysis using the converted text data in addition to emotional information recognized from the voice. This makes it possible to more accurately detect signs of harassment and to quickly notify the user of appropriate warnings.
[0345] "Voice input" is the process of acquiring acoustic information emitted by the user through a device such as a microphone.
[0346] A "terminal device" is an information processing device that acquires audio and processes and transmits data.
[0347] "Audio data" refers to information recorded in digital format from collected audio input.
[0348] "Speech recognition means" refers to technology that analyzes speech data and converts it into corresponding text data.
[0349] "Text data" refers to character information converted by speech recognition technology.
[0350] "Emotional information" refers to information extracted from audio and text that indicates the speaker's emotional state.
[0351] "Means of analysis" refers to the process of using text data and sentiment information to determine signs of harassment.
[0352] "Signs of harassment" are patterns or phrases that may indicate potential aggression or harassment in one's words or actions.
[0353] "Means of generating warnings" refers to the process of creating notifications that prompt users to take action based on detected signs of harassment.
[0354] "Notification output means" refers to technologies for conveying warnings to users visually or audibly.
[0355] This invention provides a system that utilizes speech recognition and sentiment analysis technology to detect signs of harassment and generate appropriate warnings. This enables users to maintain a healthier environment in their daily communication.
[0356] First, the user's device acquires conversations and voice input through a high-performance microphone. This device incorporates general-purpose voice processing software as a voice recognition technology. For example, a specific service as a "voice recognition means" employs technology to instantly convert voice data into text data. Noise reduction and text organization are performed at this stage to improve the accuracy of subsequent analysis.
[0357] Next, the emotion engine installed in the device analyzes the emotional information from this text and audio data. The emotion engine uses emotion analysis software to determine the emotional state from the tone of voice and the content of the text. This not only analyzes the words themselves, but also detects what emotions the speaker is feeling while speaking, and sends that information to the server.
[0358] The server receives the transmitted text data and sentiment information and performs analysis using natural language processing (NLP) techniques. Widely used NLP libraries are utilized for information processing for the analysis. Based on the analysis results, the server identifies signs of harassment and evaluates the severity of those signs by cross-referencing them with the sentiment information.
[0359] If necessary, the server generates and sends a warning prompting appropriate action to the terminal. The warning is communicated to the user visually or audibly, helping to create a comfortable and safe communication environment. A concrete example of this process is a conversation between a supervisor and an employee in the workplace. If the supervisor makes an irritated remark, the emotion engine detects that emotion, and the server provides appropriate feedback to help improve the supervisor's communication style.
[0360] Furthermore, a concrete example of a prompt using a generative AI model might be something like, "Please tell me how to generate appropriate feedback when a user becomes emotionally agitated." By combining emotional information with harassment detection in this way, more accurate responses become possible.
[0361] The flow of the specific processing in Example 2 will be explained using Figure 13.
[0362] Step 1:
[0363] The device acquires the user's voice through a high-performance microphone. The input is real-time audio data, which is converted into a digital format to prepare for subsequent processing. This digital audio data is then processed to remove background noise and format it for analysis.
[0364] Step 2:
[0365] The device uses speech recognition technology to convert audio data into text data. The input is pre-processed audio data, and the output is text data in string format. This conversion process analyzes speech patterns using acoustic and language models and generates the corresponding text.
[0366] Step 3:
[0367] The device generates emotion information from text data and speech characteristics. The input is the text data and speech tone information obtained in step 2, and the output is emotion information indicating the user's emotional state. The emotion engine analyzes the content of the text and speech characteristics such as emphasis and speed to identify the emotion the user is expressing.
[0368] Step 4:
[0369] The terminal sends the generated text data and sentiment information to the server. The input consists of text data and sentiment information, which are securely transmitted to the server using a network communication protocol. This step enables subsequent data analysis on the server.
[0370] Step 5:
[0371] The server receives the transmitted data and performs analysis using natural language processing techniques. The input consists of text data and sentiment information, and the output is an evaluation of potential signs of harassment. Here, the server extracts specific phrases and patterns from the text and combines them with sentiment information to assess their severity.
[0372] Step 6:
[0373] The server generates a warning based on the analysis results and sends it to the terminal. The input is the analysis results, and the output is the warning message that is sent to the user. This process creates a warning that highlights the areas that need improvement for the user.
[0374] Step 7:
[0375] The terminal notifies the user of any received warnings. The input is a warning message sent from the server, which is presented to the user using visual or auditory means. Based on this feedback, the user can review their words and actions and adjust their behavior accordingly.
[0376] (Application Example 2)
[0377] Next, we will explain application example 2. In the following explanation, the data processing device 12 will be referred to as the "server," and the smart glasses 214 will be referred to as the "terminal."
[0378] In various settings such as workplaces and public facilities, it is currently difficult to detect signs of harassment, and systems for preventing it are not sufficiently developed. Conventional systems mainly analyze only voice and text data, and because they cannot consider emotional nuances or context, they can make incorrect judgments. This invention aims to solve these problems and achieve more accurate harassment detection and notification.
[0379] The specific processing performed by the specific processing unit 290 of the data processing device 12 in Application Example 2 is realized by the following means.
[0380] In this invention, the server includes means for using a device to acquire voice input, means for converting the acquired voice data into text data, means for analyzing the text data and sentiment data extracted from the voice to detect signs of harassment, means for generating a warning based on the signs and sentiment data detected by the analysis means, and means for notifying the user of the warning via a visual interface. This enables advanced automated harassment detection and real-time notification combined with sentiment analysis.
[0381] "Voice input" refers to information data obtained through a device using human speech.
[0382] A "device" is an electronic device used to acquire and process audio and other data.
[0383] "Text data" refers to string data converted based on voice input.
[0384] "Emotional data" refers to information indicating emotional states, analyzed from audio and text data.
[0385] "Analysis methods" refer to processes and algorithms that detect specific patterns or emotions based on acquired data.
[0386] A "warning" is a message that notifies users of risks or situations requiring attention, based on observed data.
[0387] A "visual interface" is a means of displaying information to a user in a visual way.
[0388] To realize this invention, a system including a server and a terminal is required. The terminal is equipped with a microphone as a device for acquiring voice input and collects voice data in real time during conversations. Using speech recognition technology, this voice data is converted into text data, and further emotion data is extracted from the voice and text data using an emotion engine. The emotion engine implements algorithms for determining the tone of voice and the emotional state of the speaker.
[0389] The server receives this text and sentiment data and uses natural language processing (NLP) techniques to analyze it. The analysis includes procedures to identify potential signs of harassment within the text data and, if necessary, has the ability to cross-reference it with a database of past cases. Based on this, a warning is generated according to the detected signs and sentiment data. This warning is notified to the user via the terminal and presented using a visual interface.
[0390] The hardware and software used include Google Cloud Speech-to-Text for speech recognition, IBM Watson Tone Analyzer for sentiment analysis, and libraries such as spaCy and BERT for natural language processing.
[0391] As a concrete example, if a participant makes a statement that indicates strong stress towards another participant during a workplace discussion, the device will capture the audio, and the emotion engine will recognize the emotion of stress. The server will analyze this data, generate a warning, and provide feedback to the user through a visual interface.
[0392] An example of a prompt to input into a generative AI model is: "Design a system that analyzes voice and sentiment data to identify potential signs of harassment in workplace conversations and provides real-time notifications."
[0393] The flow of a specific process in Application Example 2 will be explained using Figure 14.
[0394] Step 1:
[0395] The device acquires voice input through the microphone. The acquired voice is transmitted to the server in real time as digital audio data. At this stage, the input is human speech, and the output is a digitized audio file.
[0396] Step 2:
[0397] The server uses a speech recognition engine (e.g., Google Cloud Speech-to-Text) to convert the received audio data into text data. The input here is audio data, and the speech recognition process generates text data as output.
[0398] Step 3:
[0399] The server extracts sentiment data from the generated text data using a sentiment engine (e.g., IBM Watson Tone Analyzer). The input is text data, and after the analysis process, sentiment data indicating the user's emotional state is output.
[0400] Step 4:
[0401] The server uses natural language processing (NLP) techniques (e.g., spaCy, BERT) to analyze text data and detect signs of harassment. The input consists of text data and sentiment data, and signs are identified based on the NLP algorithm. The output is data indicating the presence or absence of signs and their content.
[0402] Step 5:
[0403] The server generates warnings based on the analysis results, cross-referencing them with a database of past cases as needed. Here, the inputs are the analysis results and the database of past cases, and the output is a appropriately adjusted warning message.
[0404] Step 6:
[0405] The terminal notifies the user of generated warnings through a visual interface. It communicates problems to the user in real time by visually displaying warning data received from the server. The input is warning data, and the output is visual feedback to the user.
[0406] The specific processing unit 290 transmits the result of the specific processing to the smart glasses 214. In the smart glasses 214, the control unit 46A causes the speaker 240 to output the result of the specific processing. The microphone 238 acquires audio indicating user input for the result of the specific processing. The control unit 46A transmits the audio data indicating user input acquired by the microphone 238 to the data processing unit 12. In the data processing unit 12, the specific processing unit 290 acquires the audio data.
[0407] Data generation model 58 is a type of so-called generative AI (Artificial Intelligence). An example of data generation model 58 is ChatGPT (Internet Search<URL: https: / / openai.com / blog / chatgpt> ), Gemini (Internet search) <url: https: gemini.google.com ?hl="ja">Examples of generative AI include the following. The data generation model 58 is obtained by performing deep learning on a neural network. The data generation model 58 is input with prompts containing instructions, and with inference data such as audio data representing speech, text data representing text, and image data representing images. The data generation model 58 infers from the input inference data according to the instructions indicated by the prompts, and outputs the inference results in data formats such as audio data and text data. Here, inference refers to, for example, analysis, classification, prediction, and / or summarization.
[0408] In the above embodiment, an example was given in which specific processing is performed by the data processing device 12, but the technology of this disclosure is not limited thereto, and the specific processing may also be performed by the smart glasses 214.
[0409] [Third Embodiment]
[0410] Figure 5 shows an example of the configuration of the data processing system 310 according to the third embodiment.
[0411] As shown in Figure 5, the data processing system 310 includes a data processing device 12 and a headset terminal 314. An example of the data processing device 12 is a server.
[0412] The data processing device 12 comprises a computer 22, a database 24, and a communication interface 26. The computer 22 is an example of a "computer" related to the technology of this disclosure. The computer 22 comprises a processor 28, RAM 30, and storage 32. The processor 28, RAM 30, and storage 32 are connected to a bus 34. The database 24 and the communication interface 26 are also connected to the bus 34. The communication interface 26 is connected to a network 54. An example of the network 54 is a WAN (Wide Area Network) and / or a LAN (Local Area Network).
[0413] The headset terminal 314 includes a computer 36, a microphone 238, a speaker 240, a camera 42, a communication interface 44, and a display 343. The computer 36 includes a processor 46, RAM 48, and storage 50. The processor 46, RAM 48, and storage 50 are connected to a bus 52. The microphone 238, speaker 240, camera 42, and display 343 are also connected to the bus 52.
[0414] The microphone 238 receives voice signals from the user 20 and receives instructions from the user 20. The microphone 238 captures the voice signals from the user 20, converts the captured voice into audio data, and outputs it to the processor 46. The speaker 240 outputs audio according to the instructions from the processor 46.
[0415] Camera 42 is a small digital camera equipped with an optical system including a lens, aperture, and shutter, and an image sensor such as a CMOS (Complementary Metal-Oxide-Semiconductor) image sensor or a CCD (Charge Coupled Device) image sensor, and captures images of the area around the user 20 (for example, an imaging range defined by a field of view equivalent to the width of a typical healthy person's field of vision).
[0416] Communication interface 44 is connected to network 54. Communication interfaces 44 and 26 are responsible for the exchange of various information between processor 46 and processor 28 via network 54. The exchange of various information between processor 46 and processor 28 using communication interfaces 44 and 26 is performed in a secure manner.
[0417] Figure 6 shows an example of the main functions of the data processing device 12 and the headset terminal 314. As shown in Figure 6, the data processing device 12 performs specific processing using the processor 28. The storage 32 stores the specific processing program 56.
[0418] The specific processing program 56 is an example of a "program" relating to the technology of this disclosure. The processor 28 reads the specific processing program 56 from the storage 32 and executes the read specific processing program 56 on the RAM 30. The specific processing is realized by the processor 28 operating as a specific processing unit 290 in accordance with the specific processing program 56 executed on the RAM 30.
[0419] The storage 32 stores the data generation model 58 and the emotion identification model 59. The data generation model 58 and the emotion identification model 59 are used by the identification processing unit 290.
[0420] In the headset terminal 314, the processor 46 performs the reception output processing. The storage 50 stores the reception output program 60. The processor 46 reads the reception output program 60 from the storage 50 and executes the read reception output program 60 on the RAM 48. The reception output processing is realized by the processor 46 operating as a control unit 46A according to the reception output program 60 executed on the RAM 48.
[0421] Next, the specific processing performed by the specific processing unit 290 of the data processing device 12 will be described. In the following description, the data processing device 12 will be referred to as the "server" and the headset terminal 314 will be referred to as the "terminal".
[0422] This invention provides a system that acquires voice input, detects signs of harassment, and notifies the user of a warning. Specific embodiments thereof are described below.
[0423] The device acquires audio during conversations via its microphone and collects audio data in real time. This input audio data is converted into text data using speech recognition technology built into the device. In this way, the content spoken is organized into written information.
[0424] Next, the server receives the converted text data and performs analysis using natural language processing (NLP) techniques. During the analysis process, it compares the text with a database of harassment cases collected in the past to determine if there are any phrases or patterns in the text that are considered harassment.
[0425] If the analysis results indicate signs of harassment, the server immediately generates a warning and sends it to the terminal. This warning includes an explanation of the specific phrase or context in question and is communicated to the user visually or audibly. This allows the user to receive real-time feedback and reflect on and improve their words and actions.
[0426] As a concrete example, consider a scenario in a workplace meeting. If a supervisor begins using harsh language towards a subordinate, the device immediately detects this and records it as text. The server analyzes this text to determine if it matches past case data. If it is determined that there is a risk of harassment, an alert is sent to the supervisor's device. This process can prevent harassment that may occur unintentionally.
[0427] This invention is expected to improve the quality of communication in the workplace and society as a whole, as harassment, whether conscious or unconscious, can be quickly recognized and dealt with appropriately.
[0428] The following describes the processing flow.
[0429] Step 1:
[0430] The device acquires audio from the conversation via the microphone. The audio data is temporarily stored in a buffer in real time.
[0431] Step 2:
[0432] The device uses speech recognition technology to convert the acquired speech data into text data. At this stage, noise reduction is performed to minimize background noise in the conversation.
[0433] Step 3:
[0434] The terminal cleans the converted text, removing unnecessary spaces and symbols. This text data is then divided into phrases or sentences.
[0435] Step 4:
[0436] The server receives text data and begins analysis using natural language processing (NLP) techniques. This analysis includes keyword research and sentiment analysis.
[0437] Step 5:
[0438] The server analyzes the text and compares it with a database of harassment cases. Here, it checks whether the statements in the text are similar to those in past harassment cases.
[0439] Step 6:
[0440] If the server determines there is a risk of harassment, it generates an alert and sends it to the terminal. The alert includes information about the problematic statement and its impact.
[0441] Step 7:
[0442] The device visually or audibly notifies the user of alerts it has received. This allows the user to re-evaluate and improve their statements.
[0443] (Example 1)
[0444] Next, we will describe Example 1. In the following description, the data processing device 12 will be referred to as the "server," and the headset-type terminal 314 will be referred to as the "terminal."
[0445] There is a need for a system that can quickly detect harassment that occurs unconsciously in the workplace and other communication settings, and provide advance warnings, thereby preventing conflicts and discomfort among those involved.
[0446] The identification process performed by the identification processing unit 290 of the data processing device 12 in Example 1 is realized by the following means.
[0447] In this invention, the server includes means for acquiring voice input, means for converting the acquired voice data into text information, and means for analyzing the converted text information to detect signs of harassment. This makes it possible to quickly recognize signs of harassment and issue a warning to the user.
[0448] "Voice input" is the process of acquiring conversations or voice information through devices such as microphones.
[0449] A "terminal" is an electronic device that receives voice input and performs the necessary processing.
[0450] "Textual information" refers to data obtained by converting audio data into text format.
[0451] "Means of conversion" refers to the technology or device used to convert audio data into text information.
[0452] "Means of analysis" refer to techniques and functions for examining textual information and identifying specific patterns or signs.
[0453] "Signs of harassment" refer to phrases or behavioral patterns that may cause discomfort or threat to others.
[0454] A "processing device" is a system that includes computers and servers for analyzing speech and text information.
[0455] "Means of generating warnings and notifying users" refers to a function that automatically creates a message based on detected signs of harassment and notifies the user visually or audibly.
[0456] A "database" is an information system that organizes past cases and data to facilitate searching and matching.
[0457] "Visual or auditory notification" refers to a method of conveying warning information to the user through screen displays or audio output.
[0458] This invention relates to a system that acquires voice input and detects signs of harassment. The system consists of a terminal, a server, and communication between the two.
[0459] The device acquires the speaker's voice through the microphone and collects it as audio data. For speech recognition, the device needs to run appropriate speech recognition software. Examples include the Google Speech-to-Text API and IBM Watson Speech to Text. This software is then used to convert the collected audio data into text information.
[0460] The converted text information is sent to a server and analyzed using natural language processing techniques. This analysis includes extracting words and phrases from the text information and verifying whether they match patterns related to harassment registered in a database of past cases. Suitable software for this purpose on the server includes Python's NLTK and spaCy.
[0461] Based on the analysis results, the server immediately generates a warning if signs of harassment are detected. The warning message includes information about the problematic content of the statement and how to improve it, and is sent to the terminal. The terminal presents the received warning to the user through visual messages and audio notifications. This allows the user to receive real-time feedback and reflect on and improve their statements and actions.
[0462] As a concrete example, consider a workplace meeting scenario. If a supervisor uses harsh language towards a subordinate, the terminal instantly captures the words and converts them into text. The server analyzes this text, compares it against a past database, and assesses the harassment risk. A warning message is then generated and sent to the supervisor's terminal. This procedure allows for the prevention of harassment in advance.
[0463] Examples of input prompts for a generative AI model:
[0464] "Please explain the specific implementation procedures for a harassment detection system that uses audio data."
[0465] The flow of the specific processing in Example 1 will be explained using Figure 11.
[0466] Step 1:
[0467] The device uses a microphone to acquire voice input and collects it as audio data. The input is raw speech during a conversation, and the output is audio data in digital format. This data is temporarily stored within the device.
[0468] Step 2:
[0469] The terminal uses speech recognition software to convert speech data into text information. The input is the speech data obtained in the previous step, and the output is text information (text data). Specifically, the speech recognition software analyzes the speech waveform and generates the corresponding text information.
[0470] Step 3:
[0471] The server receives character information transmitted from the terminal. This character information is used as input and stored as preparation for further analysis. The output is character information stored on the server in a parseable format.
[0472] Step 4:
[0473] The server uses natural language processing techniques to analyze textual information and detect signs of harassment. The input is the textual information received in the previous step, and the output is the result of the harassment risk assessment. Specifically, the server compares the textual information with past cases stored in the database to determine whether or not there is a risk.
[0474] Step 5:
[0475] If the server detects signs of harassment, it generates a warning message. The input is the result of the risk assessment obtained in step 4, and the output is a warning message to notify the user. This warning includes a description of the specific problem phrase or context.
[0476] Step 6:
[0477] The terminal receives warning messages sent from the server and notifies the user visually or audibly. The input is the warning message from the server, and the output is a warning display or audio notification to the user. Specifically, the terminal provides feedback to the user by displaying a message on the screen or emitting an audio alert.
[0478] (Application Example 1)
[0479] Next, we will explain Application Example 1. In the following explanation, the data processing device 12 will be referred to as the "server," and the headset-type terminal 314 will be referred to as the "terminal."
[0480] There is a need to prevent harassment in internal communications by detecting conversations containing signs of harassment in real time and immediately sending warnings to those involved. However, current systems have difficulty with real-time detection and immediate feedback provision, and there is room for improvement.
[0481] The specific processing performed by the specific processing unit 290 of the data processing device 12 in Application Example 1 is realized by the following means.
[0482] In this invention, the server includes an information processing device for acquiring voice input, a conversion means for converting the acquired voice data into text data, an analysis means for analyzing the text data and detecting signs of harassment, and a means for notifying the user of the generated warning in real time and providing feedback to improve behavior. This makes it possible to detect signs of harassment in a conversation in real time and immediately issue a warning to the relevant parties.
[0483] An "information processing device for acquiring voice input" is a device that captures external voice signals, thereby enabling it to capture human speech as digital data.
[0484] A "conversion method for converting to text data" refers to a method that has the function of converting acquired audio data into text information using speech recognition technology, and plays a role in making audio information visible.
[0485] An "analytical means for detecting signs of harassment" is a means that has the ability to identify words and patterns that suggest harassment by analyzing text data and comparing it with past cases.
[0486] A "warning generation and transmission device" is a device that has the function of communicating warnings to relevant parties based on detected signs of harassment, thereby providing users with immediate feedback.
[0487] A "means of providing real-time notifications and feedback to improve behavior" refers to a means of immediately notifying users based on detected signs, providing specific improvement suggestions, and supporting user behavior.
[0488] The system to realize this application first involves the user acquiring audio using a device such as a smartphone or personal computer. The microphone in the device collects the audio of the conversation and digitizes the data. This digitized audio data is then converted into text data using speech recognition software such as the Google Speech-to-Text API.
[0489] The server receives this text data and analyzes it using natural language processing libraries such as Apache OpenNLP. During the analysis process, the acquired text data is compared against a database of past harassment cases to check for the presence of specific phrases or patterns. If signs of harassment are detected, a generative AI model is used to generate a warning message for the user.
[0490] Warnings are sent to the device in real time. This allows users to receive immediate feedback to improve their words and actions. Notifications are sent to the device visually or audibly using communication tools such as the Twilio API.
[0491] As a concrete example, consider a scenario where a user is using the application during a workplace meeting. In this situation, if a supervisor makes a negative comment to a subordinate, the audio is immediately transcribed and analyzed. If signs of harassment are detected, a warning message such as "The current statement may be harassment" will appear on the supervisor's device, providing an opportunity for them to recognize areas for improvement.
[0492] Examples of prompts to input into a generative AI model:
[0493] "Identify any remarks in this conversation that could be considered harassment and issue a warning to the user."
[0494] The flow of a specific process in Application Example 1 will be explained using Figure 12.
[0495] Step 1:
[0496] The device acquires voice input. The user launches an application on a device such as a smartphone or personal computer and captures physical voice data through the microphone. The input is an analog audio signal and is output as digital audio data.
[0497] Step 2:
[0498] The device converts the audio data into text data. The acquired digital audio data is then converted into text format using the Google Speech-to-Text API. The input is digital audio data, and the output is natural language text data.
[0499] Step 3:
[0500] The server analyzes the text data. The server uses natural language processing libraries such as Apache OpenNLP to analyze the text data. The input is text data, and the server recognizes phrases and patterns that may indicate harassment. The output is harassment detection information as a result of the analysis.
[0501] Step 4:
[0502] The server generates warnings using a generative AI model. Based on the analyzed data, the generative AI model is used to generate warning content. The input is harassment detection information, and a warning message is output to notify the user.
[0503] Step 5:
[0504] The server sends a warning to the device. The generated warning message is sent to the target device using a communication service such as the Twilio API. The input is a warning message, and the output is a notification that the user receives the warning visually or audibly.
[0505] Step 6:
[0506] The user receives feedback. The user reviews the warnings displayed on their device and receives feedback on how to improve their behavior. The input is in the form of visual and auditory notifications, and the output is provided as feedback for behavioral improvement.
[0507] Furthermore, an emotion engine that estimates the user's emotions may be incorporated. That is, the identification processing unit 290 may use the emotion identification model 59 to estimate the user's emotions and perform identification processing using the user's emotions.
[0508] This invention provides a system that combines voice and emotion recognition functions to more precisely detect signs of harassment and notify the user of appropriate warnings.
[0509] First, the device acquires audio during a conversation via its microphone, collecting audio data in real time. This audio data is then converted into text data using the device's built-in speech recognition technology. This text data undergoes a series of preprocessing steps before being prepared for analysis.
[0510] A key element here is the inclusion of an emotion engine in the device. This emotion engine recognizes and analyzes the user's emotional state from voice and text data. This allows for more than just text analysis; it also takes into account the speaker's emotional tone and intent. The emotion engine determines which emotion the user is expressing—such as anger, sadness, or joy—and provides that emotional information to the analysis system.
[0511] Next, the server receives the text data and accompanying sentiment information and performs analysis using natural language processing (NLP) techniques. The analysis identifies phrases and patterns in the text that may be considered harassment and further compares them with a database of past cases. Using the sentiment information obtained from the sentiment engine, the server evaluates the severity of the detected signs and applies it to generate warnings.
[0512] If a warning is deemed necessary, the server generates a warning based on emotional information and sends it to the terminal. This warning is communicated to the user through visual and auditory notification methods. The user can use this real-time feedback to adjust their words and actions and improve their behavior to prevent harassment.
[0513] As a concrete example, consider a conversation between a supervisor and a subordinate in the workplace. If the supervisor makes a statement expressing frustration towards the subordinate, the terminal's emotion engine recognizes the emotion of frustration from the supervisor's voice and transmits that information to the server. The server takes this emotion information, adjusts the severity of the warning, and then sends feedback to the supervisor that includes specific areas for improvement. By combining emotion information in this way, more accurate and appropriate harassment prevention measures become possible.
[0514] The following describes the processing flow.
[0515] Step 1:
[0516] The device uses its microphone to capture audio during a conversation in real time. The captured audio data is clarified through noise filtering and temporarily stored in a buffer.
[0517] Step 2:
[0518] The device uses speech recognition technology to convert the audio data in the buffer into text data. This conversion uses an algorithm that turns words and phrases in the audio into characters.
[0519] Step 3:
[0520] The device's emotion engine analyzes the user's emotional state from voice data. The analysis considers factors such as voice tone, pitch, and speed to identify emotions like joy and anger.
[0521] Step 4:
[0522] Text data and sentiment information are sent from the terminal to the server. The server receives this data and prepares it for analysis.
[0523] Step 5:
[0524] The server uses natural language processing (NLP) techniques to analyze text data. This analysis includes detecting specific keywords, understanding contextual meaning, and identifying signs of harassment.
[0525] Step 6:
[0526] The server takes in emotional information provided by the emotion engine and generates warnings based on the analysis results. Emotional information is a crucial factor in determining the severity of the warning and how it should be presented.
[0527] Step 7:
[0528] The server sends a generated warning to the terminal. The warning is adjusted to take into account the user's emotional state and the content of their statements.
[0529] Step 8:
[0530] The device notifies the user of any warnings it has received. The notification is delivered visually or audibly, prompting the user to re-evaluate their words and actions.
[0531] (Example 2)
[0532] Next, we will describe Example 2. In the following description, the data processing device 12 will be referred to as the "server," and the headset-type terminal 314 will be referred to as the "terminal."
[0533] In recent years, harassment has become a serious problem in the workplace and society, requiring immediate and appropriate responses. However, conventional systems are limited to text analysis and have the drawback of not being able to adequately consider the speaker's emotions and intentions. As a result, there is a risk that appropriate warnings may not be issued due to incorrect judgments, and problems may be overlooked. To solve this problem, a new system is needed that can perform more sophisticated analysis using voice and emotional information.
[0534] The identification process performed by the identification processing unit 290 of the data processing device 12 in Example 2 is realized by the following means.
[0535] In this invention, the server includes terminal means for acquiring voice input, voice recognition means for converting the acquired voice data into text data, and means for performing analysis using the converted text data in addition to emotional information recognized from the voice. This makes it possible to more accurately detect signs of harassment and to quickly notify the user of appropriate warnings.
[0536] "Voice input" is the process of acquiring acoustic information emitted by the user through a device such as a microphone.
[0537] A "terminal device" is an information processing device that acquires audio and processes and transmits data.
[0538] "Audio data" refers to information recorded in digital format from collected audio input.
[0539] "Speech recognition means" refers to technology that analyzes speech data and converts it into corresponding text data.
[0540] "Text data" refers to character information converted by speech recognition technology.
[0541] "Emotional information" refers to information extracted from audio and text that indicates the speaker's emotional state.
[0542] "Means of analysis" refers to the process of using text data and sentiment information to determine signs of harassment.
[0543] "Signs of harassment" are patterns or phrases that may indicate potential aggression or harassment in one's words or actions.
[0544] "Means of generating warnings" refers to the process of creating notifications that prompt users to take action based on detected signs of harassment.
[0545] "Notification output means" refers to technologies for conveying warnings to users visually or audibly.
[0546] This invention provides a system that utilizes speech recognition and sentiment analysis technology to detect signs of harassment and generate appropriate warnings. This enables users to maintain a healthier environment in their daily communication.
[0547] First, the user's device acquires conversations and voice input through a high-performance microphone. This device incorporates general-purpose voice processing software as a voice recognition technology. For example, a specific service as a "voice recognition means" employs technology to instantly convert voice data into text data. Noise reduction and text organization are performed at this stage to improve the accuracy of subsequent analysis.
[0548] Next, the emotion engine installed in the device analyzes the emotional information from this text and audio data. The emotion engine uses emotion analysis software to determine the emotional state from the tone of voice and the content of the text. This not only analyzes the words themselves, but also detects what emotions the speaker is feeling while speaking, and sends that information to the server.
[0549] The server receives the transmitted text data and sentiment information and performs analysis using natural language processing (NLP) techniques. Widely used NLP libraries are utilized for information processing for the analysis. Based on the analysis results, the server identifies signs of harassment and evaluates the severity of those signs by cross-referencing them with the sentiment information.
[0550] If necessary, the server generates and sends a warning prompting appropriate action to the terminal. The warning is communicated to the user visually or audibly, helping to create a comfortable and safe communication environment. A concrete example of this process is a conversation between a supervisor and an employee in the workplace. If the supervisor makes an irritated remark, the emotion engine detects that emotion, and the server provides appropriate feedback to help improve the supervisor's communication style.
[0551] Furthermore, a concrete example of a prompt using a generative AI model might be something like, "Please tell me how to generate appropriate feedback when a user becomes emotionally agitated." By combining emotional information with harassment detection in this way, more accurate responses become possible.
[0552] The flow of the specific processing in Example 2 will be explained using Figure 13.
[0553] Step 1:
[0554] The device acquires the user's voice through a high-performance microphone. The input is real-time audio data, which is converted into a digital format to prepare for subsequent processing. This digital audio data is then processed to remove background noise and format it for analysis.
[0555] Step 2:
[0556] The device uses speech recognition technology to convert audio data into text data. The input is pre-processed audio data, and the output is text data in string format. This conversion process analyzes speech patterns using acoustic and language models and generates the corresponding text.
[0557] Step 3:
[0558] The device generates emotion information from text data and speech characteristics. The input is the text data and speech tone information obtained in step 2, and the output is emotion information indicating the user's emotional state. The emotion engine analyzes the content of the text and speech characteristics such as emphasis and speed to identify the emotion the user is expressing.
[0559] Step 4:
[0560] The terminal sends the generated text data and sentiment information to the server. The input consists of text data and sentiment information, which are securely transmitted to the server using a network communication protocol. This step enables subsequent data analysis on the server.
[0561] Step 5:
[0562] The server receives the transmitted data and performs analysis using natural language processing techniques. The input consists of text data and sentiment information, and the output is an evaluation of potential signs of harassment. Here, the server extracts specific phrases and patterns from the text and combines them with sentiment information to assess their severity.
[0563] Step 6:
[0564] The server generates a warning based on the analysis results and sends it to the terminal. The input is the analysis results, and the output is the warning message that is sent to the user. This process creates a warning that highlights the areas that need improvement for the user.
[0565] Step 7:
[0566] The terminal notifies the user of any received warnings. The input is a warning message sent from the server, which is presented to the user using visual or auditory means. Based on this feedback, the user can review their words and actions and adjust their behavior accordingly.
[0567] (Application Example 2)
[0568] Next, we will explain application example 2. In the following explanation, the data processing device 12 will be referred to as the "server," and the headset-type terminal 314 will be referred to as the "terminal."
[0569] In various settings such as workplaces and public facilities, it is currently difficult to detect signs of harassment, and systems for preventing it are not sufficiently developed. Conventional systems mainly analyze only voice and text data, and because they cannot consider emotional nuances or context, they can make incorrect judgments. This invention aims to solve these problems and achieve more accurate harassment detection and notification.
[0570] The specific processing performed by the specific processing unit 290 of the data processing device 12 in Application Example 2 is realized by the following means.
[0571] In this invention, the server includes means for using a device to acquire voice input, means for converting the acquired voice data into text data, means for analyzing the text data and sentiment data extracted from the voice to detect signs of harassment, means for generating a warning based on the signs and sentiment data detected by the analysis means, and means for notifying the user of the warning via a visual interface. This enables advanced automated harassment detection and real-time notification combined with sentiment analysis.
[0572] "Voice input" refers to information data obtained through a device using human speech.
[0573] A "device" is an electronic device used to acquire and process audio and other data.
[0574] "Text data" refers to string data converted based on voice input.
[0575] "Emotional data" refers to information indicating emotional states, analyzed from audio and text data.
[0576] "Analysis methods" refer to processes and algorithms that detect specific patterns or emotions based on acquired data.
[0577] A "warning" is a message that notifies users of risks or situations requiring attention, based on observed data.
[0578] A "visual interface" is a means of displaying information to a user in a visual way.
[0579] To realize this invention, a system including a server and a terminal is required. The terminal is equipped with a microphone as a device for acquiring voice input and collects voice data in real time during conversations. Using speech recognition technology, this voice data is converted into text data, and further emotion data is extracted from the voice and text data using an emotion engine. The emotion engine implements algorithms for determining the tone of voice and the emotional state of the speaker.
[0580] The server receives this text and sentiment data and uses natural language processing (NLP) techniques to analyze it. The analysis includes procedures to identify potential signs of harassment within the text data and, if necessary, has the ability to cross-reference it with a database of past cases. Based on this, a warning is generated according to the detected signs and sentiment data. This warning is notified to the user via the terminal and presented using a visual interface.
[0581] The hardware and software used include Google Cloud Speech-to-Text for speech recognition, IBM Watson Tone Analyzer for sentiment analysis, and libraries such as spaCy and BERT for natural language processing.
[0582] As a concrete example, if a participant makes a statement that indicates strong stress towards another participant during a workplace discussion, the device will capture the audio, and the emotion engine will recognize the emotion of stress. The server will analyze this data, generate a warning, and provide feedback to the user through a visual interface.
[0583] An example of a prompt to input into a generative AI model is: "Design a system that analyzes voice and sentiment data to identify potential signs of harassment in workplace conversations and provides real-time notifications."
[0584] The flow of a specific process in Application Example 2 will be explained using Figure 14.
[0585] Step 1:
[0586] The device acquires voice input through the microphone. The acquired voice is transmitted to the server in real time as digital audio data. At this stage, the input is human speech, and the output is a digitized audio file.
[0587] Step 2:
[0588] The server uses a speech recognition engine (e.g., Google Cloud Speech-to-Text) to convert the received audio data into text data. The input here is audio data, and the speech recognition process generates text data as output.
[0589] Step 3:
[0590] The server extracts sentiment data from the generated text data using a sentiment engine (e.g., IBM Watson Tone Analyzer). The input is text data, and after the analysis process, sentiment data indicating the user's emotional state is output.
[0591] Step 4:
[0592] The server uses natural language processing (NLP) techniques (e.g., spaCy, BERT) to analyze text data and detect signs of harassment. The input consists of text data and sentiment data, and signs are identified based on the NLP algorithm. The output is data indicating the presence or absence of signs and their content.
[0593] Step 5:
[0594] The server generates warnings based on the analysis results, cross-referencing them with a database of past cases as needed. Here, the inputs are the analysis results and the database of past cases, and the output is a appropriately adjusted warning message.
[0595] Step 6:
[0596] The terminal notifies the user of generated warnings through a visual interface. It communicates problems to the user in real time by visually displaying warning data received from the server. The input is warning data, and the output is visual feedback to the user.
[0597] The specific processing unit 290 transmits the result of the specific processing to the headset terminal 314. In the headset terminal 314, the control unit 46A causes the speaker 240 and display 343 to output the result of the specific processing. The microphone 238 acquires audio indicating user input for the result of the specific processing. The control unit 46A transmits the audio data indicating user input acquired by the microphone 238 to the data processing unit 12. In the data processing unit 12, the specific processing unit 290 acquires the audio data.
[0598] Data generation model 58 is a type of so-called generative AI (Artificial Intelligence). An example of data generation model 58 is ChatGPT (Internet Search<URL: https: / / openai.com / blog / chatgpt> ), Gemini (Internet search) <url: https: gemini.google.com ?hl="ja">Examples of generative AI include the following. The data generation model 58 is obtained by performing deep learning on a neural network. The data generation model 58 is input with prompts containing instructions, and with inference data such as audio data representing speech, text data representing text, and image data representing images. The data generation model 58 infers from the input inference data according to the instructions indicated by the prompts, and outputs the inference results in data formats such as audio data and text data. Here, inference refers to, for example, analysis, classification, prediction, and / or summarization.
[0599] In the above embodiment, an example was given in which specific processing is performed by the data processing device 12, but the technology of this disclosure is not limited thereto, and specific processing may also be performed by the headset terminal 314.
[0600] [Fourth Embodiment]
[0601] Figure 7 shows an example of the configuration of the data processing system 410 according to the fourth embodiment.
[0602] As shown in Figure 7, the data processing system 410 includes a data processing device 12 and a robot 414. An example of the data processing device 12 is a server.
[0603] The data processing device 12 comprises a computer 22, a database 24, and a communication interface 26. The computer 22 is an example of a "computer" related to the technology of this disclosure. The computer 22 comprises a processor 28, RAM 30, and storage 32. The processor 28, RAM 30, and storage 32 are connected to a bus 34. The database 24 and the communication interface 26 are also connected to the bus 34. The communication interface 26 is connected to a network 54. An example of the network 54 is a WAN (Wide Area Network) and / or a LAN (Local Area Network).
[0604] The robot 414 includes a computer 36, a microphone 238, a speaker 240, a camera 42, a communication interface 44, and a controlled object 443. The computer 36 includes a processor 46, RAM 48, and storage 50. The processor 46, RAM 48, and storage 50 are connected to a bus 52. The microphone 238, speaker 240, camera 42, and controlled object 443 are also connected to the bus 52.
[0605] The microphone 238 receives voice signals from the user 20 and receives instructions from the user 20. The microphone 238 captures the voice signals from the user 20, converts the captured voice into audio data, and outputs it to the processor 46. The speaker 240 outputs audio according to the instructions from the processor 46.
[0606] Camera 42 is a small digital camera equipped with an optical system including a lens, aperture, and shutter, and an image sensor such as a CMOS (Complementary Metal-Oxide-Semiconductor) image sensor or a CCD (Charge Coupled Device) image sensor, and captures images of the area around the user 20 (for example, an imaging range defined by a field of view equivalent to the width of a typical healthy person's field of vision).
[0607] Communication interface 44 is connected to network 54. Communication interfaces 44 and 26 are responsible for the exchange of various information between processor 46 and processor 28 via network 54. The exchange of various information between processor 46 and processor 28 using communication interfaces 44 and 26 is performed in a secure manner.
[0608] The controlled object 443 includes a display device, LEDs in the eyes, and motors that drive the arms, hands, and feet. The posture and gestures of the robot 414 are controlled by controlling the motors of the arms, hands, and feet. Some of the robot 414's emotions can be expressed by controlling these motors. Furthermore, the robot 414's facial expressions can also be expressed by controlling the illumination state of the LEDs in its eyes.
[0609] Figure 8 shows an example of the main functions of the data processing device 12 and the robot 414. As shown in Figure 8, the data processing device 12 performs specific processing using the processor 28. The storage 32 stores the specific processing program 56.
[0610] The specific processing program 56 is an example of a "program" relating to the technology of this disclosure. The processor 28 reads the specific processing program 56 from the storage 32 and executes the read specific processing program 56 on the RAM 30. The specific processing is realized by the processor 28 operating as a specific processing unit 290 in accordance with the specific processing program 56 executed on the RAM 30.
[0611] The storage 32 stores the data generation model 58 and the emotion identification model 59. The data generation model 58 and the emotion identification model 59 are used by the identification processing unit 290.
[0612] In robot 414, the processor 46 performs the reception output processing. The storage 50 stores the reception output program 60. The processor 46 reads the reception output program 60 from the storage 50 and executes the read reception output program 60 on the RAM 48. The reception output processing is realized by the processor 46 operating as a control unit 46A according to the reception output program 60 executed on the RAM 48.
[0613] Next, the specific processing performed by the specific processing unit 290 of the data processing device 12 will be described. In the following description, the data processing device 12 will be referred to as the "server" and the robot 414 as the "terminal".
[0614] This invention provides a system that acquires voice input, detects signs of harassment, and notifies the user of a warning. Specific embodiments thereof are described below.
[0615] The device acquires audio during conversations via its microphone and collects audio data in real time. This input audio data is converted into text data using speech recognition technology built into the device. In this way, the content spoken is organized into written information.
[0616] Next, the server receives the converted text data and performs analysis using natural language processing (NLP) techniques. During the analysis process, it compares the text with a database of harassment cases collected in the past to determine if there are any phrases or patterns in the text that are considered harassment.
[0617] If the analysis results indicate signs of harassment, the server immediately generates a warning and sends it to the terminal. This warning includes an explanation of the specific phrase or context in question and is communicated to the user visually or audibly. This allows the user to receive real-time feedback and reflect on and improve their words and actions.
[0618] As a concrete example, consider a scenario in a workplace meeting. If a supervisor begins using harsh language towards a subordinate, the device immediately detects this and records it as text. The server analyzes this text to determine if it matches past case data. If it is determined that there is a risk of harassment, an alert is sent to the supervisor's device. This process can prevent harassment that may occur unintentionally.
[0619] This invention is expected to improve the quality of communication in the workplace and society as a whole, as harassment, whether conscious or unconscious, can be quickly recognized and dealt with appropriately.
[0620] The following describes the processing flow.
[0621] Step 1:
[0622] The device acquires audio from the conversation via the microphone. The audio data is temporarily stored in a buffer in real time.
[0623] Step 2:
[0624] The device uses speech recognition technology to convert the acquired speech data into text data. At this stage, noise reduction is performed to minimize background noise in the conversation.
[0625] Step 3:
[0626] The terminal cleans the converted text, removing unnecessary spaces and symbols. This text data is then divided into phrases or sentences.
[0627] Step 4:
[0628] The server receives text data and begins analysis using natural language processing (NLP) techniques. This analysis includes keyword research and sentiment analysis.
[0629] Step 5:
[0630] The server analyzes the text and compares it with a database of harassment cases. Here, it checks whether the statements in the text are similar to those in past harassment cases.
[0631] Step 6:
[0632] If the server determines there is a risk of harassment, it generates an alert and sends it to the terminal. The alert includes information about the problematic statement and its impact.
[0633] Step 7:
[0634] The device visually or audibly notifies the user of alerts it has received. This allows the user to re-evaluate and improve their statements.
[0635] (Example 1)
[0636] Next, we will describe Example 1. In the following description, the data processing device 12 will be referred to as the "server" and the robot 414 as the "terminal".
[0637] There is a need for a system that can quickly detect harassment that occurs unconsciously in the workplace and other communication settings, and provide advance warnings, thereby preventing conflicts and discomfort among those involved.
[0638] The identification process performed by the identification processing unit 290 of the data processing device 12 in Example 1 is realized by the following means.
[0639] In this invention, the server includes means for acquiring voice input, means for converting the acquired voice data into text information, and means for analyzing the converted text information to detect signs of harassment. This makes it possible to quickly recognize signs of harassment and issue a warning to the user.
[0640] "Voice input" is the process of acquiring conversations or voice information through devices such as microphones.
[0641] A "terminal" is an electronic device that receives voice input and performs the necessary processing.
[0642] "Textual information" refers to data obtained by converting audio data into text format.
[0643] "Means of conversion" refers to the technology or device used to convert audio data into text information.
[0644] "Means of analysis" refer to techniques and functions for examining textual information and identifying specific patterns or signs.
[0645] "Signs of harassment" refer to phrases or behavioral patterns that may cause discomfort or threat to others.
[0646] A "processing device" is a system that includes computers and servers for analyzing speech and text information.
[0647] "Means of generating warnings and notifying users" refers to a function that automatically creates a message based on detected signs of harassment and notifies the user visually or audibly.
[0648] A "database" is an information system that organizes past cases and data to facilitate searching and matching.
[0649] "Visual or auditory notification" refers to a method of conveying warning information to the user through screen displays or audio output.
[0650] This invention relates to a system that acquires voice input and detects signs of harassment. The system consists of a terminal, a server, and communication between the two.
[0651] The device acquires the speaker's voice through the microphone and collects it as audio data. For speech recognition, the device needs to run appropriate speech recognition software. Examples include the Google Speech-to-Text API and IBM Watson Speech to Text. This software is then used to convert the collected audio data into text information.
[0652] The converted text information is sent to a server and analyzed using natural language processing techniques. This analysis includes extracting words and phrases from the text information and verifying whether they match patterns related to harassment registered in a database of past cases. Suitable software for this purpose on the server includes Python's NLTK and spaCy.
[0653] Based on the analysis results, the server immediately generates a warning if signs of harassment are detected. The warning message includes information about the problematic content of the statement and how to improve it, and is sent to the terminal. The terminal presents the received warning to the user through visual messages and audio notifications. This allows the user to receive real-time feedback and reflect on and improve their statements and actions.
[0654] As a concrete example, consider a workplace meeting scenario. If a supervisor uses harsh language towards a subordinate, the terminal instantly captures the words and converts them into text. The server analyzes this text, compares it against a past database, and assesses the harassment risk. A warning message is then generated and sent to the supervisor's terminal. This procedure allows for the prevention of harassment in advance.
[0655] Examples of input prompts for a generative AI model:
[0656] "Please explain the specific implementation procedures for a harassment detection system that uses audio data."
[0657] The flow of the specific processing in Example 1 will be explained using Figure 11.
[0658] Step 1:
[0659] The device uses a microphone to acquire voice input and collects it as audio data. The input is raw speech during a conversation, and the output is audio data in digital format. This data is temporarily stored within the device.
[0660] Step 2:
[0661] The terminal uses speech recognition software to convert speech data into text information. The input is the speech data obtained in the previous step, and the output is text information (text data). Specifically, the speech recognition software analyzes the speech waveform and generates the corresponding text information.
[0662] Step 3:
[0663] The server receives character information transmitted from the terminal. This character information is used as input and stored as preparation for further analysis. The output is character information stored on the server in a parseable format.
[0664] Step 4:
[0665] The server uses natural language processing techniques to analyze textual information and detect signs of harassment. The input is the textual information received in the previous step, and the output is the result of the harassment risk assessment. Specifically, the server compares the textual information with past cases stored in the database to determine whether or not there is a risk.
[0666] Step 5:
[0667] If the server detects signs of harassment, it generates a warning message. The input is the result of the risk assessment obtained in step 4, and the output is a warning message to notify the user. This warning includes a description of the specific problem phrase or context.
[0668] Step 6:
[0669] The terminal receives warning messages sent from the server and notifies the user visually or audibly. The input is the warning message from the server, and the output is a warning display or audio notification to the user. Specifically, the terminal provides feedback to the user by displaying a message on the screen or emitting an audio alert.
[0670] (Application Example 1)
[0671] Next, we will explain Application Example 1. In the following explanation, the data processing device 12 will be referred to as the "server" and the robot 414 as the "terminal".
[0672] There is a need to prevent harassment in internal communications by detecting conversations containing signs of harassment in real time and immediately sending warnings to those involved. However, current systems have difficulty with real-time detection and immediate feedback provision, and there is room for improvement.
[0673] The specific processing performed by the specific processing unit 290 of the data processing device 12 in Application Example 1 is realized by the following means.
[0674] In this invention, the server includes an information processing device for acquiring voice input, a conversion means for converting the acquired voice data into text data, an analysis means for analyzing the text data and detecting signs of harassment, and a means for notifying the user of the generated warning in real time and providing feedback to improve behavior. This makes it possible to detect signs of harassment in a conversation in real time and immediately issue a warning to the relevant parties.
[0675] An "information processing device for acquiring voice input" is a device that captures external voice signals, thereby enabling it to capture human speech as digital data.
[0676] A "conversion method for converting to text data" refers to a method that has the function of converting acquired audio data into text information using speech recognition technology, and plays a role in making audio information visible.
[0677] An "analytical means for detecting signs of harassment" is a means that has the ability to identify words and patterns that suggest harassment by analyzing text data and comparing it with past cases.
[0678] A "warning generation and transmission device" is a device that has the function of communicating warnings to relevant parties based on detected signs of harassment, thereby providing users with immediate feedback.
[0679] A "means of providing real-time notifications and feedback to improve behavior" refers to a means of immediately notifying users based on detected signs, providing specific improvement suggestions, and supporting user behavior.
[0680] The system to realize this application first involves the user acquiring audio using a device such as a smartphone or personal computer. The microphone in the device collects the audio of the conversation and digitizes the data. This digitized audio data is then converted into text data using speech recognition software such as the Google Speech-to-Text API.
[0681] The server receives this text data and analyzes it using natural language processing libraries such as Apache OpenNLP. During the analysis process, the acquired text data is compared against a database of past harassment cases to check for the presence of specific phrases or patterns. If signs of harassment are detected, a generative AI model is used to generate a warning message for the user.
[0682] Warnings are sent to the device in real time. This allows users to receive immediate feedback to improve their words and actions. Notifications are sent to the device visually or audibly using communication tools such as the Twilio API.
[0683] As a concrete example, consider a scenario where a user is using the application during a workplace meeting. In this situation, if a supervisor makes a negative comment to a subordinate, the audio is immediately transcribed and analyzed. If signs of harassment are detected, a warning message such as "The current statement may be harassment" will appear on the supervisor's device, providing an opportunity for them to recognize areas for improvement.
[0684] Examples of prompts to input into a generative AI model:
[0685] "Identify any remarks in this conversation that could be considered harassment and issue a warning to the user."
[0686] The flow of a specific process in Application Example 1 will be explained using Figure 12.
[0687] Step 1:
[0688] The device acquires voice input. The user launches an application on a device such as a smartphone or personal computer and captures physical voice data through the microphone. The input is an analog audio signal and is output as digital audio data.
[0689] Step 2:
[0690] The device converts the audio data into text data. The acquired digital audio data is then converted into text format using the Google Speech-to-Text API. The input is digital audio data, and the output is natural language text data.
[0691] Step 3:
[0692] The server analyzes the text data. The server uses natural language processing libraries such as Apache OpenNLP to analyze the text data. The input is text data, and the server recognizes phrases and patterns that may indicate harassment. The output is harassment detection information as a result of the analysis.
[0693] Step 4:
[0694] The server generates warnings using a generative AI model. Based on the analyzed data, the generative AI model is used to generate warning content. The input is harassment detection information, and a warning message is output to notify the user.
[0695] Step 5:
[0696] The server sends a warning to the device. The generated warning message is sent to the target device using a communication service such as the Twilio API. The input is a warning message, and the output is a notification that the user receives the warning visually or audibly.
[0697] Step 6:
[0698] The user receives feedback. The user reviews the warnings displayed on their device and receives feedback on how to improve their behavior. The input is in the form of visual and auditory notifications, and the output is provided as feedback for behavioral improvement.
[0699] Furthermore, an emotion engine that estimates the user's emotions may be incorporated. That is, the identification processing unit 290 may use the emotion identification model 59 to estimate the user's emotions and perform identification processing using the user's emotions.
[0700] This invention provides a system that combines voice and emotion recognition functions to more precisely detect signs of harassment and notify the user of appropriate warnings.
[0701] First, the device acquires audio during a conversation via its microphone, collecting audio data in real time. This audio data is then converted into text data using the device's built-in speech recognition technology. This text data undergoes a series of preprocessing steps before being prepared for analysis.
[0702] A key element here is the inclusion of an emotion engine in the device. This emotion engine recognizes and analyzes the user's emotional state from voice and text data. This allows for more than just text analysis; it also takes into account the speaker's emotional tone and intent. The emotion engine determines which emotion the user is expressing—such as anger, sadness, or joy—and provides that emotional information to the analysis system.
[0703] Next, the server receives the text data and accompanying sentiment information and performs analysis using natural language processing (NLP) techniques. The analysis identifies phrases and patterns in the text that may be considered harassment and further compares them with a database of past cases. Using the sentiment information obtained from the sentiment engine, the server evaluates the severity of the detected signs and applies it to generate warnings.
[0704] If a warning is deemed necessary, the server generates a warning based on emotional information and sends it to the terminal. This warning is communicated to the user through visual and auditory notification methods. The user can use this real-time feedback to adjust their words and actions and improve their behavior to prevent harassment.
[0705] As a concrete example, consider a conversation between a supervisor and a subordinate in the workplace. If the supervisor makes a statement expressing frustration towards the subordinate, the terminal's emotion engine recognizes the emotion of frustration from the supervisor's voice and transmits that information to the server. The server takes this emotion information, adjusts the severity of the warning, and then sends feedback to the supervisor that includes specific areas for improvement. By combining emotion information in this way, more accurate and appropriate harassment prevention measures become possible.
[0706] The following describes the processing flow.
[0707] Step 1:
[0708] The device uses its microphone to capture audio during a conversation in real time. The captured audio data is clarified through noise filtering and temporarily stored in a buffer.
[0709] Step 2:
[0710] The device uses speech recognition technology to convert the audio data in the buffer into text data. This conversion uses an algorithm that turns words and phrases in the audio into characters.
[0711] Step 3:
[0712] The device's emotion engine analyzes the user's emotional state from voice data. The analysis considers factors such as voice tone, pitch, and speed to identify emotions like joy and anger.
[0713] Step 4:
[0714] Text data and sentiment information are sent from the terminal to the server. The server receives this data and prepares it for analysis.
[0715] Step 5:
[0716] The server uses natural language processing (NLP) techniques to analyze text data. This analysis includes detecting specific keywords, understanding contextual meaning, and identifying signs of harassment.
[0717] Step 6:
[0718] The server takes in emotional information provided by the emotion engine and generates warnings based on the analysis results. Emotional information is a crucial factor in determining the severity of the warning and how it should be presented.
[0719] Step 7:
[0720] The server sends a generated warning to the terminal. The warning is adjusted to take into account the user's emotional state and the content of their statements.
[0721] Step 8:
[0722] The device notifies the user of any warnings it has received. The notification is delivered visually or audibly, prompting the user to re-evaluate their words and actions.
[0723] (Example 2)
[0724] Next, we will describe Example 2. In the following description, the data processing device 12 will be referred to as the "server" and the robot 414 as the "terminal".
[0725] In recent years, harassment has become a serious problem in the workplace and society, requiring immediate and appropriate responses. However, conventional systems are limited to text analysis and have the drawback of not being able to adequately consider the speaker's emotions and intentions. As a result, there is a risk that appropriate warnings may not be issued due to incorrect judgments, and problems may be overlooked. To solve this problem, a new system is needed that can perform more sophisticated analysis using voice and emotional information.
[0726] The identification process performed by the identification processing unit 290 of the data processing device 12 in Example 2 is realized by the following means.
[0727] In this invention, the server includes terminal means for acquiring voice input, voice recognition means for converting the acquired voice data into text data, and means for performing analysis using the converted text data in addition to emotional information recognized from the voice. This makes it possible to more accurately detect signs of harassment and to quickly notify the user of appropriate warnings.
[0728] "Voice input" is the process of acquiring acoustic information emitted by the user through a device such as a microphone.
[0729] A "terminal device" is an information processing device that acquires audio and processes and transmits data.
[0730] "Audio data" refers to information recorded in digital format from collected audio input.
[0731] "Speech recognition means" refers to technology that analyzes speech data and converts it into corresponding text data.
[0732] "Text data" refers to character information converted by speech recognition technology.
[0733] "Emotional information" refers to information extracted from audio and text that indicates the speaker's emotional state.
[0734] "Means of analysis" refers to the process of using text data and sentiment information to determine signs of harassment.
[0735] "Signs of harassment" are patterns or phrases that may indicate potential aggression or harassment in one's words or actions.
[0736] "Means of generating warnings" refers to the process of creating notifications that prompt users to take action based on detected signs of harassment.
[0737] "Notification output means" refers to technologies for conveying warnings to users visually or audibly.
[0738] This invention provides a system that utilizes speech recognition and sentiment analysis technology to detect signs of harassment and generate appropriate warnings. This enables users to maintain a healthier environment in their daily communication.
[0739] First, the user's device acquires conversations and voice input through a high-performance microphone. This device incorporates general-purpose voice processing software as a voice recognition technology. For example, a specific service as a "voice recognition means" employs technology to instantly convert voice data into text data. Noise reduction and text organization are performed at this stage to improve the accuracy of subsequent analysis.
[0740] Next, the emotion engine installed in the device analyzes the emotional information from this text and audio data. The emotion engine uses emotion analysis software to determine the emotional state from the tone of voice and the content of the text. This not only analyzes the words themselves, but also detects what emotions the speaker is feeling while speaking, and sends that information to the server.
[0741] The server receives the transmitted text data and sentiment information and performs analysis using natural language processing (NLP) techniques. Widely used NLP libraries are utilized for information processing for the analysis. Based on the analysis results, the server identifies signs of harassment and evaluates the severity of those signs by cross-referencing them with the sentiment information.
[0742] If necessary, the server generates and sends a warning prompting appropriate action to the terminal. The warning is communicated to the user visually or audibly, helping to create a comfortable and safe communication environment. A concrete example of this process is a conversation between a supervisor and an employee in the workplace. If the supervisor makes an irritated remark, the emotion engine detects that emotion, and the server provides appropriate feedback to help improve the supervisor's communication style.
[0743] Furthermore, a concrete example of a prompt using a generative AI model might be something like, "Please tell me how to generate appropriate feedback when a user becomes emotionally agitated." By combining emotional information with harassment detection in this way, more accurate responses become possible.
[0744] The flow of the specific processing in Example 2 will be explained using Figure 13.
[0745] Step 1:
[0746] The device acquires the user's voice through a high-performance microphone. The input is real-time audio data, which is converted into a digital format to prepare for subsequent processing. This digital audio data is then processed to remove background noise and format it for analysis.
[0747] Step 2:
[0748] The device uses speech recognition technology to convert audio data into text data. The input is pre-processed audio data, and the output is text data in string format. This conversion process analyzes speech patterns using acoustic and language models and generates the corresponding text.
[0749] Step 3:
[0750] The device generates emotion information from text data and speech characteristics. The input is the text data and speech tone information obtained in step 2, and the output is emotion information indicating the user's emotional state. The emotion engine analyzes the content of the text and speech characteristics such as emphasis and speed to identify the emotion the user is expressing.
[0751] Step 4:
[0752] The terminal sends the generated text data and sentiment information to the server. The input consists of text data and sentiment information, which are securely transmitted to the server using a network communication protocol. This step enables subsequent data analysis on the server.
[0753] Step 5:
[0754] The server receives the transmitted data and performs analysis using natural language processing techniques. The input consists of text data and sentiment information, and the output is an evaluation of potential signs of harassment. Here, the server extracts specific phrases and patterns from the text and combines them with sentiment information to assess their severity.
[0755] Step 6:
[0756] The server generates a warning based on the analysis results and sends it to the terminal. The input is the analysis results, and the output is the warning message that is sent to the user. This process creates a warning that highlights the areas that need improvement for the user.
[0757] Step 7:
[0758] The terminal notifies the user of any received warnings. The input is a warning message sent from the server, which is presented to the user using visual or auditory means. Based on this feedback, the user can review their words and actions and adjust their behavior accordingly.
[0759] (Application Example 2)
[0760] Next, we will explain application example 2. In the following explanation, the data processing device 12 will be referred to as the "server" and the robot 414 as the "terminal".
[0761] In various settings such as workplaces and public facilities, it is currently difficult to detect signs of harassment, and systems for preventing it are not sufficiently developed. Conventional systems mainly analyze only voice and text data, and because they cannot consider emotional nuances or context, they can make incorrect judgments. This invention aims to solve these problems and achieve more accurate harassment detection and notification.
[0762] The specific processing performed by the specific processing unit 290 of the data processing device 12 in Application Example 2 is realized by the following means.
[0763] In this invention, the server includes means for using a device to acquire voice input, means for converting the acquired voice data into text data, means for analyzing the text data and sentiment data extracted from the voice to detect signs of harassment, means for generating a warning based on the signs and sentiment data detected by the analysis means, and means for notifying the user of the warning via a visual interface. This enables advanced automated harassment detection and real-time notification combined with sentiment analysis.
[0764] "Voice input" refers to information data obtained through a device using human speech.
[0765] A "device" is an electronic device used to acquire and process audio and other data.
[0766] "Text data" refers to string data converted based on voice input.
[0767] "Emotional data" refers to information indicating emotional states, analyzed from audio and text data.
[0768] "Analysis methods" refer to processes and algorithms that detect specific patterns or emotions based on acquired data.
[0769] A "warning" is a message that notifies users of risks or situations requiring attention, based on observed data.
[0770] A "visual interface" is a means of displaying information to a user in a visual way.
[0771] To realize this invention, a system including a server and a terminal is required. The terminal is equipped with a microphone as a device for acquiring voice input and collects voice data in real time during conversations. Using speech recognition technology, this voice data is converted into text data, and further emotion data is extracted from the voice and text data using an emotion engine. The emotion engine implements algorithms for determining the tone of voice and the emotional state of the speaker.
[0772] The server receives this text and sentiment data and uses natural language processing (NLP) techniques to analyze it. The analysis includes procedures to identify potential signs of harassment within the text data and, if necessary, has the ability to cross-reference it with a database of past cases. Based on this, a warning is generated according to the detected signs and sentiment data. This warning is notified to the user via the terminal and presented using a visual interface.
[0773] The hardware and software used include Google Cloud Speech-to-Text for speech recognition, IBM Watson Tone Analyzer for sentiment analysis, and libraries such as spaCy and BERT for natural language processing.
[0774] As a concrete example, if a participant makes a statement that indicates strong stress towards another participant during a workplace discussion, the device will capture the audio, and the emotion engine will recognize the emotion of stress. The server will analyze this data, generate a warning, and provide feedback to the user through a visual interface.
[0775] An example of a prompt to input into a generative AI model is: "Design a system that analyzes voice and sentiment data to identify potential signs of harassment in workplace conversations and provides real-time notifications."
[0776] The flow of a specific process in Application Example 2 will be explained using Figure 14.
[0777] Step 1:
[0778] The device acquires voice input through the microphone. The acquired voice is transmitted to the server in real time as digital audio data. At this stage, the input is human speech, and the output is a digitized audio file.
[0779] Step 2:
[0780] The server uses a speech recognition engine (e.g., Google Cloud Speech-to-Text) to convert the received audio data into text data. The input here is audio data, and the speech recognition process generates text data as output.
[0781] Step 3:
[0782] The server extracts sentiment data from the generated text data using a sentiment engine (e.g., IBM Watson Tone Analyzer). The input is text data, and after the analysis process, sentiment data indicating the user's emotional state is output.
[0783] Step 4:
[0784] The server uses natural language processing (NLP) techniques (e.g., spaCy, BERT) to analyze text data and detect signs of harassment. The input consists of text data and sentiment data, and signs are identified based on the NLP algorithm. The output is data indicating the presence or absence of signs and their content.
[0785] Step 5:
[0786] The server generates warnings based on the analysis results, cross-referencing them with a database of past cases as needed. Here, the inputs are the analysis results and the database of past cases, and the output is a appropriately adjusted warning message.
[0787] Step 6:
[0788] The terminal notifies the user of generated warnings through a visual interface. It communicates problems to the user in real time by visually displaying warning data received from the server. The input is warning data, and the output is visual feedback to the user.
[0789] The specific processing unit 290 transmits the result of the specific processing to the robot 414. In the robot 414, the control unit 46A causes the speaker 240 and the controlled object 443 to output the result of the specific processing. The microphone 238 acquires audio indicating user input for the result of the specific processing. The control unit 46A transmits the audio data indicating user input acquired by the microphone 238 to the data processing unit 12. In the data processing unit 12, the specific processing unit 290 acquires the audio data.
[0790] Data generation model 58 is a type of so-called generative AI (Artificial Intelligence). An example of data generation model 58 is ChatGPT (Internet Search<URL: https: / / openai.com / blog / chatgpt> ), Gemini (Internet search) <url: https: gemini.google.com ?hl="ja">Examples of generative AI include the following. The data generation model 58 is obtained by performing deep learning on a neural network. The data generation model 58 is input with prompts containing instructions, and with inference data such as audio data representing speech, text data representing text, and image data representing images. The data generation model 58 infers from the input inference data according to the instructions indicated by the prompts, and outputs the inference results in data formats such as audio data and text data. Here, inference refers to, for example, analysis, classification, prediction, and / or summarization.
[0791] In the above embodiment, an example was given in which the specific processing is performed by the data processing device 12, but the technology of this disclosure is not limited thereto, and the specific processing may also be performed by the robot 414.
[0792] Furthermore, the emotion identification model 59, acting as an emotion engine, may determine the user's emotion according to a specific mapping. Specifically, the emotion identification model 59 may determine the user's emotion according to a specific mapping, which is an emotion map (see Figure 9). Similarly, the emotion identification model 59 may also determine the robot's emotion, and the identification processing unit 290 may perform identification processing using the robot's emotion.
[0793] Figure 9 shows an emotion map 400 in which multiple emotions are mapped. In the emotion map 400, emotions are arranged in concentric circles radiating from the center. The closer to the center of the concentric circles, the more primitive the emotions are located. Further out of the concentric circles, emotions representing states and actions arising from mental states are located. Emotion is a concept that includes feelings and mental states. On the left side of the concentric circles, emotions that are generally generated from reactions occurring in the brain are located. On the right side of the concentric circles, emotions that are generally induced by situational judgment are located. Above and below the concentric circles, emotions that are generally generated from reactions occurring in the brain and induced by situational judgment are located. In addition, the emotion of "pleasure" is located on the upper side of the concentric circles, and the emotion of "displeasure" is located on the lower side. Thus, in the emotion map 400, multiple emotions are mapped based on the structure in which emotions arise, and emotions that are likely to occur simultaneously are mapped close together.
[0794] These emotions are distributed at the 3 o'clock position on the Emotion Map 400, and usually fluctuate between feelings of security and anxiety. In the right half of the Emotion Map 400, situational awareness takes precedence over internal feelings, resulting in a calm impression.
[0795] The inside of the Emotion Map 400 represents inner thoughts, while the outside represents actions. Therefore, the further you go from the outside of the Emotion Map 400, the more visible (expressed in actions) your emotions become.
[0796] Here, human emotions are based on various balances, such as posture and blood sugar levels. When these balances deviate from the ideal, it results in discomfort, and when they approach the ideal, it results in pleasure. Similarly, in robots, cars, motorcycles, etc., emotions can be created based on various balances, such as posture and battery level. When these balances deviate from the ideal, it results in discomfort, and when they approach the ideal, it results in pleasure. The emotion map can be generated, for example, based on Dr. Mitsuyoshi's emotion map (Research on a system for analyzing brain physiological signals of speech emotion recognition and emotion, Tokushima University, doctoral dissertation: https: / / ci.nii.ac.jp / naid / 500000375379). The left half of the emotion map contains emotions belonging to a region called "response," where sensation is dominant. The right half of the emotion map contains emotions belonging to a region called "situation," where situational awareness is dominant.
[0797] The emotion map defines two emotions that promote learning. One is the emotion around the middle of the negative "repentance" and "reflection" on the situation side. In other words, it is when the robot experiences negative emotions such as "I never want to feel this way again" or "I don't want to be scolded again." The other is the emotion around the positive "desire" on the reaction side. In other words, it is when the robot has positive feelings such as "I want more" or "I want to know more."
[0798] The emotion identification model 59 inputs user input into a pre-trained neural network, obtains emotion values representing each emotion shown in the emotion map 400, and determines the user's emotion. This neural network is pre-trained based on multiple training data sets, which are combinations of user input and emotion values representing each emotion shown in the emotion map 400. Furthermore, this neural network is trained so that emotions located close together have similar values, as shown in the emotion map 900 in Figure 10. Figure 10 shows an example where multiple emotions such as "reassured," "calm," and "confident" have similar emotion values.
[0799] The above description primarily focuses on the functions of the data processing device 12 in relation to this disclosure. However, the system related to this disclosure is not necessarily implemented on a server. The system related to this disclosure may be implemented as a general information processing system. This disclosure may be implemented, for example, as a software program that runs on a personal computer or as an application that runs on a smartphone. The method related to this disclosure may be provided to users in SaaS (Software as a Service) format.
[0800] In the above embodiment, an example was given in which a specific process is performed by a single computer 22. However, the technology of this disclosure is not limited thereto, and a distributed processing of the specific process may be performed by multiple computers, including computer 22. For example, a data generation model 58 may be provided in an external device of the data processing device 12, and the external device may generate data according to the input data.
[0801] In the above embodiment, an example was given in which the specific processing program 56 is stored in the storage 32, but the technology of this disclosure is not limited thereto. For example, the specific processing program 56 may be stored in a portable, computer-readable, non-temporary storage medium such as a USB (Universal Serial Bus) memory. The specific processing program 56 stored in the non-temporary storage medium is installed in the computer 22 of the data processing device 12. The processor 28 executes specific processing according to the specific processing program 56.
[0802] Alternatively, the specific processing program 56 may be stored in a storage device such as a server connected to the data processing device 12 via the network 54, and the specific processing program 56 may be downloaded and installed on the computer 22 in response to a request from the data processing device 12.
[0803] Furthermore, it is not necessary to store the entirety of the specific processing program 56 in a storage device such as a server connected to the data processing device 12 via the network 54, or to store the entirety of the specific processing program 56 in the storage 32; it is acceptable to store only a portion of the specific processing program 56.
[0804] The following types of processors can be used as hardware resources to perform specific processing. Examples of processors include a CPU, a general-purpose processor that functions as a hardware resource to perform specific processing by executing software, i.e., a program. Other examples of processors include dedicated electrical circuits, such as FPGAs (Field-Programmable Gate Arrays), PLDs (Programmable Logic Devices), or ASICs (Application Specific Integrated Circuits), which have circuit configurations specifically designed to perform specific processing. All of these processors have built-in or connected memory, and all of them perform specific processing by using memory.
[0805] The hardware resource that performs a specific process may consist of one of these various processors, or it may consist of a combination of two or more processors of the same or different types (for example, a combination of multiple FPGAs, or a combination of a CPU and an FPGA). Alternatively, the hardware resource that performs a specific process may consist of a single processor.
[0806] Examples of configurations using a single processor include, firstly, a configuration in which one or more CPUs and software are combined to form a single processor, and this processor functions as a hardware resource that performs a specific process. Secondly, there is a configuration using a processor that realizes the functions of the entire system, including multiple hardware resources that perform a specific process, on a single IC chip, as exemplified by SoCs (System-on-a-chip). In this way, a specific process is realized using one or more of the above types of processors as hardware resources.
[0807] Furthermore, the hardware structure of these various processors can more specifically utilize electrical circuits that combine circuit elements such as semiconductor devices. Also, the specific processing described above is merely an example. Therefore, it goes without saying that unnecessary steps can be deleted, new steps added, or the processing order rearranged, as long as it does not deviate from the main purpose.
[0808] The descriptions and illustrations presented above are detailed explanations of the technical aspects of this disclosure and are merely examples of the technical aspects. For example, the above descriptions of the structure, function, operation, and effect are examples of the structure, function, operation, and effect of the technical aspects of this disclosure. Therefore, it goes without saying that you may delete unnecessary parts, add new elements, or replace elements in the descriptions and illustrations presented above, as long as you do not deviate from the essence of the technical aspects of this disclosure. Furthermore, in order to avoid confusion and facilitate understanding of the technical aspects of this disclosure, explanations of common technical knowledge and the like that do not require special explanation to enable the implementation of the technical aspects of this disclosure have been omitted from the descriptions and illustrations presented above.
[0809] All documents, patent applications, and technical standards described herein are incorporated by reference to the same extent as if each individual document, patent application, and technical standard were specifically and individually noted to be incorporated by reference.
[0810] The following is further disclosed regarding the embodiments described above.
[0811] (Claim 1)
[0812] An information processing device that acquires voice input,
[0813] A conversion means for converting acquired audio data into text data,
[0814] An analytical means for analyzing the text data and detecting signs of harassment,
[0815] A means for issuing a warning based on the signs detected by the analysis means,
[0816] A system that includes this.
[0817] (Claim 2)
[0818] The system according to claim 1, further comprising an information processing device that generates a warning by comparing the analyzed text data with past cases.
[0819] (Claim 3)
[0820] The system according to claim 1, further comprising means for visually or audibly notifying the user when a warning is issued.
[0821] "Example 1"
[0822] (Claim 1)
[0823] A device that acquires voice input,
[0824] A means of converting acquired audio data into text information,
[0825] A processing device that analyzes converted character information and detects signs of harassment,
[0826] A means for generating a warning based on the signs detected by the aforementioned processing device and notifying the user,
[0827] A system that includes this.
[0828] (Claim 2)
[0829] The system according to claim 1, which generates a warning by comparing the analyzed character information with a database of past cases.
[0830] (Claim 3)
[0831] The system according to claim 1, further comprising means for providing visual or auditory notification to the user when a warning is issued.
[0832] "Application Example 1"
[0833] (Claim 1)
[0834] An information processing device that acquires voice input,
[0835] A conversion means for converting acquired audio data into text data,
[0836] An analytical means for analyzing the text data and detecting signs of harassment,
[0837] A means for issuing a warning based on the signs detected by the analysis means,
[0838] A means of notifying users of generated warnings in real time and providing feedback to improve their behavior,
[0839] A system that includes this.
[0840] (Claim 2)
[0841] The system according to claim 1, further comprising an information processing device that compares the analyzed text data with past cases and provides notification based on technical analysis.
[0842] (Claim 3)
[0843] The system according to claim 1, further comprising a function that suggests specific improvements to the user when a warning is issued.
[0844] "Example 2 of combining an emotion engine"
[0845] (Claim 1)
[0846] A terminal means for acquiring voice input,
[0847] A speech recognition means that converts acquired audio data into text data,
[0848] A means of performing analysis using emotional information recognized from speech in addition to the converted text data,
[0849] A means for generating a warning based on signs and emotional information of harassment,
[0850] An output means that sends a warning to the terminal and notifies the user,
[0851] A system that includes this.
[0852] (Claim 2)
[0853] The system according to claim 1, further comprising information processing means for comparing analyzed text data and sentiment information with past cases and evaluating the severity of the warning to be generated.
[0854] (Claim 3)
[0855] The system according to claim 1, further comprising means for visually or audibly notifying the user so that the user can adjust their actions based on the warnings they have received.
[0856] "Application example 2 when combining with an emotional engine"
[0857] (Claim 1)
[0858] A device that acquires voice input,
[0859] A means of converting acquired audio data into text data,
[0860] A means for analyzing emotional data extracted from the text data and audio to detect signs of harassment,
[0861] Means for generating warnings based on signs and emotion data detected by analysis means,
[0862] A means of notifying the user of a warning via a visual interface,
[0863] A system that includes this.
[0864] (Claim 2)
[0865] The system according to claim 1, further comprising an information processing device that generates a warning by matching analyzed text data and sentiment data with case information.
[0866] (Claim 3)
[0867] The system according to claim 1, further comprising an interface that visually notifies the user, including additional information, when a warning is issued. [Explanation of Symbols]
[0868] 10, 210, 310, 410 Data Processing Systems 12 Data Processing Devices 14 Smart Devices 214 Smart Glasses 314 Headset-type terminal 414 Robots< / url:> < / url:> < / url:> < / url:>
Claims
1. An information processing device that acquires voice input, A conversion means for converting acquired audio data into text data, An analytical means for analyzing the text data and detecting signs of harassment, A means for issuing a warning based on the signs detected by the analysis means, A system that includes this.
2. The system according to claim 1, further comprising an information processing device that generates a warning by comparing the analyzed text data with past cases.
3. The system according to claim 1, further comprising means for visually or audibly notifying the user when a warning is issued.