system
The system addresses the challenge of unconscious negative remarks by detecting and converting them into positive expressions, improving communication skills through real-time feedback and reports.
Patent Information
- Authority / Receiving Office
- JP · JP
- Patent Type
- Applications
- Current Assignee / Owner
- SOFTBANK GROUP CORP
- Filing Date
- 2024-10-15
- Publication Date
- 2026-04-27
AI Technical Summary
In modern communication, individuals often unconsciously make negative remarks, leading to friction in relationships and hindering good communication, as they are unaware of these remarks and struggle to convert them into positive expressions.
A system comprising voice acquisition, speech recognition, natural language processing, notification, and report generation means to detect negative remarks in real-time, providing visual or haptic feedback and post-conversation reports to improve communication skills.
Enables users to recognize and convert negative expressions into positive ones, enhancing interpersonal relationships through immediate feedback and personalized improvement suggestions.
Smart Images

Figure 2026070243000001_ABST
Abstract
Description
Technical Field
[0001] The technology of the present disclosure relates to a system.
Background Art
[0002] Patent Document 1 discloses a method for controlling a persona chatbot, which is performed by at least one processor, the method including the steps of receiving a user utterance, adding the user utterance to a prompt including an instruction sentence related to an explanation of a character of the chatbot, encoding the prompt, and inputting the encoded prompt into a language model to generate a chatbot utterance in response to the user utterance.
Prior Art Documents
Patent Documents
[0003]
Patent Document 1
Summary of the Invention
Problems to be Solved by the Invention
[0004] In modern society, it is common to unconsciously make negative remarks in daily communication, which may cause friction in human relationships. In such situations, it is difficult for the speaker himself / herself to notice negative expressions and it is hard to consciously convert them into positive expressions. In addition, the conversation with others may stagnate and good communication may be hindered. The object of this invention is to detect and notify negative remarks in real time and support the speaker to easily convert them into positive expressions.
Means for Solving the Problems
[0005] To solve the above-mentioned problems, the present invention comprises a voice acquisition means for acquiring voice in real time, a voice recognition means for converting the acquired voice into text, a natural language processing means for analyzing the text data to detect negative remarks, a notification means for notifying the user visually or haptically based on the detected negative remarks, and a report generation means for providing the analysis results of negative remarks after the conversation. This allows the user to immediately become aware of negative expressions during a conversation and switch to positive expressions. Furthermore, by reviewing their own communication patterns through the report and receiving specific advice for improvement, it helps in building good interpersonal relationships.
[0006] A "voice acquisition means" is a device that has the function of capturing the user's conversation voice in real time and incorporating it into the system as a digital signal.
[0007] A "speech recognition means" is a device that has the function of analyzing acquired speech data and converting it into corresponding text data.
[0008] "Natural language processing means" refers to a device or system that uses algorithms or methods to analyze text data and identify specific expressions, including negative statements.
[0009] "Notification means" refers to a device or method that provides visual or tactile feedback to the user based on detected negative remarks.
[0010] A "report generation means" is a device or system that has the function of compiling the analysis results of negative comments, organizing the data necessary to provide users with information for improvement, and presenting it to them. [Brief explanation of the drawing]
[0011] [Figure 1] This is a conceptual diagram showing an example of the configuration of a data processing system according to the first embodiment. [Figure 2]This is a conceptual diagram showing an example of the essential functions of a data processing device and a smart device according to the first embodiment. [Figure 3] This is a conceptual diagram showing an example of the configuration of a data processing system according to the second embodiment. [Figure 4] This is a conceptual diagram showing an example of the main functions of a data processing device and smart glasses according to the second embodiment. [Figure 5] This is a conceptual diagram showing an example of the configuration of a data processing system according to the third embodiment. [Figure 6] This is a conceptual diagram showing an example of the main functions of a data processing device and a headset-type terminal according to the third embodiment. [Figure 7] This is a conceptual diagram showing an example of the configuration of a data processing system according to the fourth embodiment. [Figure 8] This is a conceptual diagram showing an example of the main functions of a data processing device and a robot according to the fourth embodiment. [Figure 9] This shows an emotion map where multiple emotions are mapped. [Figure 10] This shows an emotion map where multiple emotions are mapped. [Figure 11] This is a sequence diagram showing the processing flow of the data processing system in Example 1. [Figure 12] This is a sequence diagram showing the processing flow of the data processing system in Application Example 1. [Figure 13] This is a sequence diagram showing the processing flow of the data processing system in Example 2 when an emotion engine is combined. [Figure 14] This is a sequence diagram showing the processing flow of the data processing system in Application Example 2, which combines an emotion engine. [Modes for carrying out the invention]
[0012] Hereinafter, an example of an embodiment of the system relating to the technology of this disclosure will be described with reference to the attached drawings.
[0013] First, the terms used in the following description will be explained.
[0014] In the following embodiments, a processor with a reference number (hereinafter simply referred to as "processor") may be a single arithmetic unit or a combination of multiple arithmetic units. Further, the processor may be a single type of arithmetic unit or a combination of multiple types of arithmetic units. Examples of arithmetic units include a CPU (Central Processing Unit), a GPU (Graphics Processing Unit), a GPGPU (General-Purpose computing on Graphics Processing Units), an APU (Accelerated Processing Unit), and the like.
[0015] In the following embodiments, a RAM (Random Access Memory) with a reference number is a memory in which information is temporarily stored and is used as a work memory by the processor.
[0016] In the following embodiments, a storage with a reference number is one or more non-volatile storage devices that store various programs and various parameters, etc. Examples of non-volatile storage devices include flash memory (SSD (Solid State Drive)), magnetic disks (e.g., hard disks), or magnetic tapes, and the like.
[0017] In the following embodiments, a communication I / F (Interface) with a reference number is an interface including a communication processor and an antenna, etc. The communication I / F controls communication between multiple computers. Examples of communication standards applied to the communication I / F include wireless communication standards including 5G (5th Generation Mobile Communication System), Wi-Fi (registered trademark), or Bluetooth (registered trademark), and the like.
[0018] In the following embodiments, "A and / or B" is synonymous with "at least one of A and B." That is, "A and / or B" means that it may be A alone, or B alone, or a combination of A and B. Furthermore, in this specification, the same concept as "A and / or B" applies when expressing three or more things linked by "and / or."
[0019] [First Embodiment]
[0020] Figure 1 shows an example of the configuration of the data processing system 10 according to the first embodiment.
[0021] As shown in Figure 1, the data processing system 10 includes a data processing device 12 and a smart device 14. An example of the data processing device 12 is a server.
[0022] The data processing device 12 comprises a computer 22, a database 24, and a communication interface 26. The computer 22 is an example of a "computer" related to the technology of this disclosure. The computer 22 comprises a processor 28, RAM 30, and storage 32. The processor 28, RAM 30, and storage 32 are connected to a bus 34. The database 24 and the communication interface 26 are also connected to the bus 34. The communication interface 26 is connected to a network 54. An example of the network 54 is a WAN (Wide Area Network) and / or a LAN (Local Area Network).
[0023] The smart device 14 comprises a computer 36, a reception device 38, an output device 40, a camera 42, and a communication interface 44. The computer 36 comprises a processor 46, RAM 48, and storage 50. The processor 46, RAM 48, and storage 50 are connected to a bus 52. The reception device 38, output device 40, and camera 42 are also connected to the bus 52.
[0024] The reception device 38 is equipped with a touch panel 38A and a microphone 38B, etc., and receives user input. The touch panel 38A receives user input by detecting contact with an object (e.g., a pen or finger). The microphone 38B receives user input by detecting the user's voice. The control unit 46A transmits data indicating the user input received by the touch panel 38A and microphone 38B to the data processing device 12. In the data processing device 12, the specific processing unit 290 acquires the data indicating the user input.
[0025] The output device 40 includes a display 40A and a speaker 40B, and presents data to the user 20 by outputting the data in a form perceptible to the user 20 (e.g., audio and / or text). The display 40A displays visible information such as text and images according to instructions from the processor 46. The speaker 40B outputs audio according to instructions from the processor 46. The camera 42 is a small digital camera equipped with an optical system such as a lens, aperture, and shutter, and an image sensor such as a CMOS (Complementary Metal-Oxide-Semiconductor) image sensor or a CCD (Charge Coupled Device) image sensor.
[0026] Communication interface 44 is connected to network 54. Communication interfaces 44 and 26 are responsible for the exchange of various types of information between processor 46 and processor 28 via network 54.
[0027] Figure 2 shows an example of the main functions of the data processing device 12 and the smart device 14.
[0028] As shown in Figure 2, in the data processing device 12, a specific processing is performed by the processor 28. A specific processing program 56 is stored in the storage 32. The specific processing program 56 is an example of a "program" related to the technology of this disclosure. The processor 28 reads the specific processing program 56 from the storage 32 and executes the read specific processing program 56 on the RAM 30. The specific processing is realized by the processor 28 operating as a specific processing unit 290 according to the specific processing program 56 executed on the RAM 30.
[0029] The storage 32 stores the data generation model 58 and the emotion identification model 59. The data generation model 58 and the emotion identification model 59 are used by the identification processing unit 290.
[0030] In the smart device 14, the processor 46 performs the reception output processing. The storage 50 stores the reception output program 60. The reception output program 60 is used in conjunction with a specific processing program 56 by the data processing system 10. The processor 46 reads the reception output program 60 from the storage 50 and executes the read reception output program 60 on the RAM 48. The reception output processing is realized by the processor 46 operating as a control unit 46A according to the reception output program 60 executed on the RAM 48.
[0031] Next, the specific processing performed by the specific processing unit 290 of the data processing device 12 will be described. In the following description, the data processing device 12 will be referred to as the "server" and the smart device 14 as the "terminal".
[0032] In embodiments of the present invention, the system includes a bio-worn terminal and integrates various functions to support natural conversation for the user. A detailed description of its operation and specific examples are provided below.
[0033] In this system, the terminal functions as a voice acquisition device, capturing ambient sound in real time. The acquired sound is converted into text data by the speech recognition device within the terminal. Speech recognition is performed using pre-trained acoustic and language models, ensuring accurate digital conversion of spoken content.
[0034] The converted text is analyzed using natural language processing tools. This analysis evaluates whether negative remarks are included and scores them using keyword lists and contextual analysis. Since the terminal processes this in real time, it is possible to provide immediate feedback to the user.
[0035] If negative comments are detected, the device provides feedback to the user via notification. These notifications are delivered through devices such as glasses, rings, or watches, using subtle vibrations or flashing lights to discreetly signal the user. This allows the user to become aware of their comments without interrupting the conversation and, if necessary, switch to positive language.
[0036] Furthermore, once the conversation ends, the data collected by the device is compiled by a report generation system, and the information is provided to the user's smartphone. The report includes the percentage and timing of negative comments, as well as specific advice for improvement, which users can use to improve their everyday communication skills.
[0037] For example, if a user says, "I'm really tired today, and I don't want to do anything anymore," the device detects negative phrases such as "tired" and "don't want to do anything." Then, a glasses-type device vibrates subtly to inform the user of this fact. After the conversation, the user receives suggestions on how to improve their expression through a report displayed on a smartphone app. In this way, the system provides users with a way to continuously train their positive communication skills.
[0038] The following describes the processing flow.
[0039] Step 1:
[0040] The device uses its built-in microphone to capture audio from the user's surroundings in real time. This audio data is captured within the device as a digital signal.
[0041] Step 2:
[0042] The device uses a speech recognition module to convert acquired speech data into text data. Speech recognition utilizes pre-trained acoustic and language models.
[0043] Step 3:
[0044] The terminal analyzes the converted text data using a natural language processing module. This module detects negative expressions using keyword lists and sentiment analysis algorithms, and calculates a negative score for each statement.
[0045] Step 4:
[0046] The device calculates a negative score and determines whether it exceeds a certain threshold. If it exceeds the threshold, the information is passed to the notification module.
[0047] Step 5:
[0048] The device will notify the user. Depending on the notification method, a device shaped like glasses or a ring will use vibration or flashing LEDs to inform the user of the presence of negative comments.
[0049] Step 6:
[0050] After the conversation ends, the device compiles all the spoken data and negative scores, and then organizes the analysis results.
[0051] Step 7:
[0052] The device generates a report using a smartphone app based on the collected data. This report includes the number and timing of negative comments, as well as suggestions for improvement.
[0053] Step 8:
[0054] Users can view reports on their smartphones and use them to improve their communication skills. This allows them to recognize their own negative tendencies and consciously switch to more positive expressions.
[0055] (Example 1)
[0056] Next, we will describe Example 1. In the following description, the data processing device 12 will be referred to as the "server," and the smart device 14 will be referred to as the "terminal."
[0057] This invention aims to solve the problem of users unconsciously creating emotional biases in their speech during everyday conversations. Therefore, it seeks to encourage improvement in expression by immediately identifying negative expressions used by users during conversations and providing feedback based on those findings. Conventional methods have faced challenges in visualizing emotional balance during conversations and providing timely feedback.
[0058] The identification process performed by the identification processing unit 290 of the data processing device 12 in Example 1 is realized by the following means.
[0059] In this invention, the server includes means for acquiring audio in real time, means for converting the acquired audio into text data, and means for using natural language processing techniques to measure emotional tendencies. This allows users to instantly recognize negative patterns in their own speech during a conversation and receive appropriate feedback, thereby improving the quality of communication.
[0060] "Means for acquiring audio in real time" refers to technologies and devices for collecting ambient audio signals and using them for subsequent processing.
[0061] "Means of converting to text data" refers to technology or equipment that converts acquired audio signals into text information, thereby achieving the digitization of audio information.
[0062] "Means for detecting negative descriptions" refers to algorithms and technologies for identifying negative emotions and expressions within text, thereby performing sentiment analysis of speech.
[0063] "Means of providing immediate notification" refers to devices or technologies for quickly communicating information to users based on detected negative reviews.
[0064] "Means of generating reports" refers to systems and technologies for creating systematic reports based on analysis results and providing them to users.
[0065] "Natural language processing technology" refers to technologies for processing human language using computers, and includes grammatical analysis, sentiment analysis, and other text processing.
[0066] "The shape of eyeglasses, rings, or watches" refers to the shape of a device that may be worn by a user and is a wearable device with notification capabilities.
[0067] This invention provides an information processing device to support natural conversations by users. Primarily, it utilizes a bio-worn terminal and has the function of capturing and analyzing the user's speech in real time.
[0068] The device uses advanced microphone technology to capture audio in real time. This audio input is optimized using filtering technology to remove background noise. The captured audio data is converted into text data by speech recognition software. This process utilizes pre-trained acoustic and language models to improve accuracy.
[0069] The converted text data is analyzed by the server using natural language processing techniques. This analysis determines whether negative descriptions are detected. Sentiment analysis includes keyword searching and contextual evaluation, and incorporates algorithms to measure emotional tendencies. The analysis results are immediately provided as feedback, using subtle vibrations or LED displays as notification methods. This allows users to receive appropriate feedback during conversations.
[0070] As a concrete example of its use, if a user says, "I'm really tired today, and I don't want to do anything anymore," the device will detect phrases like "tired" and "don't want to do anything" as negative. The glasses-type device will then vibrate to inform the user of this fact, providing an immediate opportunity for them to consciously correct their statement.
[0071] Furthermore, once the conversation ends, the server analyzes the results and sends a detailed report to the user's smartphone. This report includes areas for improvement in speech and specific advice, allowing users to receive feedback to improve their everyday communication skills. An example of a prompt might be a question such as, "What aspects of today's speech could be improved?"
[0072] The flow of the specific processing in Example 1 will be explained using Figure 11.
[0073] Step 1:
[0074] The device acquires the sound surrounding the user in real time through the microphone. This audio signal becomes the input. The device applies filtering to remove background noise and outputs clear audio data. This process allows for more accurate capture of the user's speech.
[0075] Step 2:
[0076] The device's speech recognition software takes filtered audio data as input and converts it into text data. This involves calculations based on a generative AI model that utilizes acoustic and linguistic models, outputting the audio signal as text information. Specifically, this process includes analyzing each phoneme and mapping it to the most matching word.
[0077] Step 3:
[0078] The server uses natural language processing techniques to analyze the generated text data as input. This analysis aims to detect negative descriptions, processing the data using keyword lists and contextual analysis, and outputting the results as an evaluation score. At this stage, sentiment tendencies become clear.
[0079] Step 4:
[0080] Based on the analysis results score, the device will immediately notify the user as needed. The input is the evaluation score, and the output is feedback through subtle vibrations and LED displays. Specifically, the glasses-type device vibrates to inform the user that negative comments have been detected.
[0081] Step 5:
[0082] Once the conversation ends, the server generates a detailed report based on the collected data. The input for this process is past conversation data and its analysis results, while the output is a report that includes the percentage of negative comments and specific advice for improvement. This report is provided through the user's smartphone app, allowing the user to review its contents and use them to improve future communication.
[0083] (Application Example 1)
[0084] Next, we will explain Application Example 1. In the following explanation, the data processing device 12 will be referred to as the "server," and the smart device 14 will be referred to as the "terminal."
[0085] In many situations, it is crucial to proactively detect negative remarks and risky conversations and respond appropriately. However, existing technologies have struggled to detect risks in real time and provide immediate feedback to operators. As a result, misunderstandings and conflicts can arise during conversations. In addition, the limited feedback available after conversations has hindered long-term improvement of communication skills.
[0086] The specific processing performed by the specific processing unit 290 of the data processing device 12 in Application Example 1 is realized by the following means.
[0087] In this invention, the server includes an acoustic acquisition means for acquiring sound in real time, a speech recognition means for converting the acquired sound into text data, a natural language analysis means for analyzing the text data to detect negative statements, and a warning means for detecting risks predicted during the operator's statements in real time and providing corresponding notifications. This enables the operator to instantly grasp negative phrases and potential risks and take appropriate action. Furthermore, the use of the report generation means and suggestion generation means contributes to improving negative statements and enhancing the quality of communication.
[0088] An "acoustic acquisition means" is a device that captures ambient sounds in real time and converts them into a format that can be processed within the system.
[0089] "Speech recognition means" refers to a processing device for converting acquired acoustic data into text format.
[0090] A "natural language analysis tool" is a device that analyzes text data and has the function of detecting negative statements and important language patterns.
[0091] A "notification means" is a device that provides visual or tactile feedback to the operator based on detected information.
[0092] A "report generation means" is a device that compiles the results of the analysis of conversation data and creates documents or digital reports to provide to the operator.
[0093] A "warning device" is a device that detects potential risks during speech in real time and immediately issues a warning to the operator.
[0094] A "suggestion generation device" is a device that provides feedback to the operator, including improvement measures and suggestions, based on detected negative statements and risks.
[0095] This invention is a system that includes sound acquisition means, speech recognition means, natural language analysis means, notification means, report generation means, warning means, and suggestion generation means. A server acts as the central point, monitoring user statements in real time and detecting negative phrases and risks.
[0096] First, a microphone is used to capture ambient sound. This acoustic data is then converted into text format by a speech recognition system. This process utilizes speech recognition technologies such as Google® Cloud Speech-to-Text API.
[0097] Next, natural language processing tools analyze the text data to detect negative statements and risks in real time. This utilizes natural language processing technologies such as the Google NLP API.
[0098] If detected, a notification system activates, sending a subtle vibration or light to the user via glasses or accessories. This allows the user to receive feedback without interrupting the conversation.
[0099] After the conversation ends, a detailed report is generated by the report generation mechanism. This report includes improvement suggestions and is provided to the user through a smartphone application. Based on this report, the system, through the suggestion generation mechanism, makes specific suggestions to support the user in improving their communication skills.
[0100] For example, if a user says, "I'm really tired today," the system immediately detects the negative phrase "tired." The glasses-type device then vibrates slightly to draw the user's attention. A subsequent report offers suggestions on how to switch to more positive language.
[0101] Examples of prompts include questions such as, "What negative phrases were detected in the conversation?" or "What advice would you give to improve these negative elements?"
[0102] This system allows users to quickly identify risk factors during conversations and continuously improve their skills.
[0103] The flow of a specific process in Application Example 1 will be explained using Figure 12.
[0104] Step 1:
[0105] Acoustic Acquisition: The device uses a microphone to acquire ambient sounds in real time. The input is ambient sound, and the output is acoustic data. This acoustic data is converted to a digital format in preparation for processing.
[0106] Step 2:
[0107] Speech Recognition: The server uses speech recognition to convert acquired audio data into text data. The input is digital audio data, and the output is the corresponding text data. The Google Cloud Speech-to-Text API is used for speech recognition to accurately transcribe the content of the audio data into text.
[0108] Step 3:
[0109] Natural Language Processing: The server uses natural language analysis tools to analyze text data and detect negative statements. The input is text data, and the output is the analysis results and evaluation score. Using the Google NLP API, the text content is scanned, sentiment analysis is performed, and negative keywords are extracted.
[0110] Step 4:
[0111] Feedback Notification: The device uses notification methods to send a subtle vibration or light alert to the operator based on detected negative statements. The input is the analysis result, and the output is a physical notification. Feedback is provided discreetly by operating glasses-type devices or accessories.
[0112] Step 5:
[0113] Post-conversation report generation: The server uses a report generation mechanism to generate a report based on a detailed analysis of the conversation and sends it to the user's smartphone. The input is the analysis data of the entire conversation, and the output is a detailed report. This report includes suggestions for improvement.
[0114] Step 6:
[0115] Providing suggestions: The server provides users with suggestions for improving negative statements through a suggestion generation mechanism. The input is the content of the report, and the output is personalized suggestions. Users can utilize the provided advice in their daily communication.
[0116] Furthermore, an emotion engine that estimates the user's emotions may be incorporated. That is, the identification processing unit 290 may use the emotion identification model 59 to estimate the user's emotions and perform identification processing using the user's emotions.
[0117] This invention is a system that combines speech recognition technology, natural language processing technology, and an emotion engine to support improved user communication. This system is implemented as a bio-worn device that understands the user's natural conversational situation and provides real-time feedback.
[0118] In this system, the terminal first functions as a voice acquisition device, capturing surrounding sounds in real time. The acquired voice is converted into text data by a speech recognition device. This converted data is analyzed by a natural language processing device to evaluate whether it contains negative expressions.
[0119] The key is that an emotion engine is incorporated into this process. This emotion engine analyzes voice tone and contextual information of the utterance to recognize the user's emotional state. This allows for analysis that takes into account not only the words themselves, but also the psychological background behind them. For example, if someone says "I can't do this anymore" in a tired voice, it can determine whether that emotional state is depression or normal fatigue.
[0120] The system provides subtle notifications to the user in response to negative comments or emotional states. These notifications take the form of items such as glasses, rings, or watches, and provide feedback through vibration and light. The intensity and frequency of these notifications are adjusted based on data from the emotion engine, enabling more precise and personalized responses.
[0121] Furthermore, after a conversation, a detailed report based on the collected data is generated in the smartphone app. This report includes specific areas for improvement and suggestions for positive communication. Based on this information, users can effectively improve their communication skills. For example, if a user expresses emotional frustration, the system can analyze the tone and content and suggest alternative, positive expressions while preventing unnecessary intensity in notifications. In this way, it is an invention that comprehensively supports the user's emotions and the content of the conversation.
[0122] The following describes the processing flow.
[0123] Step 1:
[0124] The device uses its built-in microphone to capture audio from the user's surroundings in real time. This audio data is captured within the device as a digital signal.
[0125] Step 2:
[0126] The device converts the acquired audio data into text data using speech recognition technology. In this process, acoustic and language models are utilized to improve the accuracy of the speech-to-text conversion.
[0127] Step 3:
[0128] The device uses a natural language processing module to analyze the converted text data. This analysis includes keyword analysis and contextual analysis to assess whether negative statements are included.
[0129] Step 4:
[0130] The device uses an emotion engine to analyze voice tone, content, and context to recognize the user's emotional state. This information helps determine the user's psychological state when speaking.
[0131] Step 5:
[0132] The device determines whether a notification is necessary based on negative scores and emotional states. If necessary, it subtly notifies the user using vibration or light. The intensity and frequency of these notifications are adjusted according to the analysis results of the emotion engine.
[0133] Step 6:
[0134] After the conversation ends, the device collects all the data and generates a report based on the analysis of emotional state and content of speech. This report is sent to the user via a smartphone app.
[0135] Step 7:
[0136] The system censors reports provided by users via smartphone and uses the information gathered to improve their communication skills. The reports include negative aspects and emotional context of statements, as well as suggestions for improvement.
[0137] (Example 2)
[0138] Next, we will describe Example 2. In the following description, the data processing device 12 will be referred to as the "server" and the smart device 14 as the "terminal".
[0139] In modern society, it is crucial to accurately understand an individual's emotional state and improve negative communication. However, conventional technologies are limited to analyzing simple text data obtained from voice, lacking the means to recognize emotional changes in real time and respond appropriately. As a result, users have difficulty accurately understanding their own emotional state, posing a challenge to improving the quality of communication.
[0140] The identification process performed by the identification processing unit 290 of the data processing device 12 in Example 2 is realized by the following means.
[0141] In this invention, the server includes an input means for acquiring audio in real time, a conversion means for converting the acquired audio into text, and an analysis means for analyzing the text and evaluating the emotional state. This allows for direct analysis of the emotional state from the audio data, rapid understanding of emotional changes in real-time dialogue, and the provision of appropriate feedback and improvement measures.
[0142] "Input means for acquiring audio in real time" refers to a device or function of a device that instantly captures audio present in the user's surroundings and provides the audio data for subsequent processing.
[0143] "A conversion method for converting acquired audio into text" refers to a system or algorithm that handles the process of converting audio input into textual information, and uses speech recognition technology to transcribe spoken content into text.
[0144] "Analysis means for analyzing text and evaluating emotional state" refers to a technology or process for analyzing converted text data and identifying and evaluating the emotional nuances and state of the user.
[0145] "Presentation means for providing physical notifications to the user based on emotional state" refers to a device or function for providing visual or tactile stimuli to the user and transmitting feedback based on the results of emotion analysis.
[0146] "A means for generating reports that include improvement suggestions while considering the user's emotional state" refers to a system or process for creating reports that provide specific areas for improvement and suggestions to the user based on the results of emotional and communication analysis.
[0147] A "wearable information display device" refers to a device that can be worn directly by the user and provides visual or tactile information.
[0148] This invention is a system that highly analyzes the user's emotional state and communication content, and provides appropriate feedback and improvement suggestions. The system consists of means for voice acquisition, speech recognition, natural language processing, emotion analysis, feedback notification, and report generation.
[0149] The device is equipped with an input mechanism for acquiring audio in real time, which is installed in wearable information display devices such as glasses, rings, and watches. This makes it possible to instantly collect the user's ambient sounds and conversation content. The collected audio data is sent to a server and converted into text using speech recognition technology such as Google Cloud Speech-to-Text.
[0150] Once the audio is converted to text, the server uses a natural language processing system to analyze the text data and evaluate the emotional state, particularly negative expressions. This evaluation also includes the user's voice tone and contextual information, allowing the emotion engine to analyze the underlying psychological nuances.
[0151] Based on the analysis results, the device provides visual or tactile notifications to the user. For example, if the device determines that the user is feeling fatigued, it will provide a soft vibration or light notification to encourage relaxation.
[0152] Furthermore, after a conversation session, the server generates a detailed report based on the analysis data and provides it to the user via a smartphone app. This report includes specific areas for improvement based on sentiment analysis and suggestions for positive communication.
[0153] For example, if a user says, "I'm feeling irritated today and can't concentrate on work," the system will detect that negative emotion and, taking into account the tone of voice, make suggestions to minimize its impact.
[0154] An example of a prompt for a generative AI model is: "Analyze the user's statement 'I'm tired today, but I'm satisfied,' determine the emotional state, and suggest feedback." This prompt allows the model to analyze the user's statement and the emotions behind it, and generate appropriate feedback.
[0155] The flow of the specific processing in Example 2 will be explained using Figure 13.
[0156] Step 1:
[0157] The device collects audio in real time. When the user begins speaking, the device uses its built-in microphone to input ambient sounds and generates digital audio data. This data is immediately sent to the server, enabling real-time processing.
[0158] Step 2:
[0159] The server inputs the received audio data into speech recognition software, which converts it into text data. This process uses a speech recognition engine, specifically analyzing the words in the audio and outputting the corresponding text. The output obtained at this stage is text information that includes all the spoken content.
[0160] Step 3:
[0161] The server passes text data to a natural language processing system for sentiment analysis. It analyzes specific keywords and phrases from the input text and applies natural language processing algorithms to evaluate negative or positive emotional states. The output here is an evaluation of the user's emotional state.
[0162] Step 4:
[0163] The server determines what to notify the user based on the sentiment analysis results. Based on the analysis information, it generates appropriate feedback and prepares data to present it to the user, for example, through a slight vibration or visual sign. The output is the specific feedback message to be notified.
[0164] Step 5:
[0165] The device uses the generated feedback message to physically notify the user. It uses vibration motors, LEDs, etc., to convey feedback to the user. At this stage, the user receives appropriate signals based on the sentiment analysis results.
[0166] Step 6:
[0167] After all processing is complete, the server generates a detailed report based on the conversation content and sentiment analysis results. This report includes suggestions for improvement and positive feedback. The generated report is sent to the user's smartphone app.
[0168] Step 7:
[0169] Users can review reports on a smartphone app and consider actions to improve their communication skills based on the suggested improvements. This allows for practical improvement while reflecting on daily conversations.
[0170] (Application Example 2)
[0171] Next, we will explain application example 2. In the following explanation, the data processing device 12 will be referred to as a "server" and the smart device 14 as a "terminal".
[0172] With the advancement of technology, improving the user experience is becoming increasingly important in many industries. In particular, improving the quality of communication between staff and customers is a challenge in customer service at physical stores. However, it is difficult for staff to respond quickly to customers' emotions and negative comments using conventional methods. This invention aims to solve these problems and improve customer service at physical stores.
[0173] The specific processing performed by the specific processing unit 290 of the data processing device 12 in Application Example 2 is realized by the following means. In this invention, the server includes a voice acquisition means for acquiring voice in real time, a voice recognition means for converting the acquired voice into text data, and a natural language processing means for analyzing the text data to detect negative statements. This makes it possible to grasp the emotional state of customers in real time and notify staff of that information as needed.
[0174] "Voice acquisition means" refers to a means for acquiring ambient sounds in real time and using them for subsequent processing.
[0175] "Speech recognition means" refers to a means for converting acquired speech into text data.
[0176] "Natural language processing tools" are methods for analyzing text data to detect negative statements or specific emotional states.
[0177] "Notification means" refers to means of providing users with visual or tactile notifications based on detected negative statements and the user's emotional state.
[0178] A "record generation means" is a means for providing analysis results based on negative remarks and emotional evaluations after a conversation.
[0179] A "visual support device" is a device that helps users intuitively understand information by displaying it visually.
[0180] A "finger-worn device" is a device worn on the finger to provide physical or informational notifications.
[0181] A "wrist-worn device" is a device worn on the arm that provides physical or informational notifications.
[0182] "Positive dialogue methods" refer to methods and practical suggestions for users to engage in better communication.
[0183] The system for carrying out this invention consists of a process that acquires voice in real time, performs emotion analysis, and provides feedback to the user. Specifically, the terminal captures ambient sound using voice acquisition means. This voice is converted into text data using speech recognition means.
[0184] Text data is analyzed by natural language processing on the server to detect negative remarks and specific emotional states. The analyzed emotional data is used to determine the user's emotional state and the quality of communication. Based on the emotional state, the user is notified through visual support devices, finger-worn devices, and arm-worn devices. This allows the user to understand customer reactions in real time and take appropriate action.
[0185] Furthermore, after the conversation ends, the server uses a record generation mechanism to create a detailed report and suggests positive ways of interacting to the user. This report derives areas for improvement and recommendations through a generative AI model.
[0186] As a concrete example, this system could be used in physical stores to support staff in customer service. If negative remarks or feelings of depression are detected during a conversation with a customer, the staff will be immediately notified with information to address the issue. As a result, it becomes possible to provide higher quality customer support.
[0187] An example of a prompt is, "Create an app that analyzes customer conversation audio and provides staff with real-time emotional feedback."
[0188] The flow of a specific process in Application Example 2 will be explained using Figure 14.
[0189] Step 1:
[0190] The device uses voice acquisition means to capture ambient sounds in real time. The input is natural ambient sound, which is then captured as digital audio data using acoustic sensors and microphones. This process requires filtering out external noise and appropriately capturing the audio signal.
[0191] Step 2:
[0192] The audio data sent to the server is converted into text data by speech recognition technology. In this process, the input digital audio data is converted into string data using speech recognition software (e.g., Google Speech Recognition API). The resulting text data serves as the basis for communication analysis.
[0193] Step 3:
[0194] The server uses natural language processing (NLTK) to analyze text data and detect negative statements and specific emotional states. The input text data undergoes morphological analysis and sentiment scoring using natural language processing libraries (e.g., NLTK and SpaCy), and the presence or absence of negative expressions is output. This process also includes tone analysis by an emotion engine to understand the context.
[0195] Step 4:
[0196] The server provides feedback to the user via notification methods based on the analysis results. The input consists of analysis results of negative statements and emotions, and physical notification devices (e.g., visual support devices or vibration notification devices) are used for feedback. This output allows the user to understand the customer's emotional state in a timely manner and take appropriate action.
[0197] Step 5:
[0198] After the conversation ends, the server generates a detailed report using a record generation system and provides it to the user. Using a generation AI model, the report outputs information including areas for improvement and positive dialogue methods, based on the input analysis data. This report helps improve the user's communication skills.
[0199] The specific processing unit 290 transmits the result of the specific processing to the smart device 14. In the smart device 14, the control unit 46A causes the output device 40 to output the result of the specific processing. The microphone 38B acquires audio indicating user input for the result of the specific processing. The control unit 46A transmits the audio data indicating user input acquired by the microphone 38B to the data processing device 12. In the data processing device 12, the specific processing unit 290 acquires the audio data.
[0200] Data generation model 58 is a so-called generative AI (Artificial Intelligence). An example of data generation model 58 is ChatGPT (registered trademark) (Internet search).<URL: https: / / openai.com / blog / chatgpt> ), Gemini (registered trademark) (Internet search) <url: https: gemini.google.com ?hl="ja">Examples of generative AI include the following. The data generation model 58 is obtained by performing deep learning on a neural network. The data generation model 58 is input with prompts containing instructions, and with inference data such as audio data representing speech, text data representing text, and image data representing images. The data generation model 58 infers from the input inference data according to the instructions indicated by the prompts, and outputs the inference results in data formats such as audio data and text data. Here, inference refers to, for example, analysis, classification, prediction, and / or summarization.
[0201] In the above embodiment, an example was given in which specific processing is performed by the data processing device 12, but the technology of this disclosure is not limited thereto, and the specific processing may also be performed by the smart device 14.
[0202] [Second Embodiment]
[0203] Figure 3 shows an example of the configuration of the data processing system 210 according to the second embodiment.
[0204] As shown in Figure 3, the data processing system 210 includes a data processing device 12 and smart glasses 214. An example of the data processing device 12 is a server.
[0205] The data processing device 12 comprises a computer 22, a database 24, and a communication interface 26. The computer 22 is an example of a "computer" related to the technology of this disclosure. The computer 22 comprises a processor 28, RAM 30, and storage 32. The processor 28, RAM 30, and storage 32 are connected to a bus 34. The database 24 and the communication interface 26 are also connected to the bus 34. The communication interface 26 is connected to a network 54. An example of the network 54 is a WAN (Wide Area Network) and / or a LAN (Local Area Network).
[0206] The smart glasses 214 include a computer 36, a microphone 238, a speaker 240, a camera 42, and a communication interface 44. The computer 36 includes a processor 46, RAM 48, and storage 50. The processor 46, RAM 48, and storage 50 are connected to a bus 52. The microphone 238, speaker 240, and camera 42 are also connected to the bus 52.
[0207] The microphone 238 receives voice signals from the user 20 and receives instructions from the user 20. The microphone 238 captures the voice signals from the user 20, converts the captured voice into audio data, and outputs it to the processor 46. The speaker 240 outputs audio according to the instructions from the processor 46.
[0208] Camera 42 is a small digital camera equipped with an optical system including a lens, aperture, and shutter, and an image sensor such as a CMOS (Complementary Metal-Oxide-Semiconductor) image sensor or a CCD (Charge Coupled Device) image sensor, and captures images of the area around the user 20 (for example, an imaging range defined by a field of view equivalent to the width of a typical healthy person's field of vision).
[0209] Communication interface 44 is connected to network 54. Communication interfaces 44 and 26 are responsible for the exchange of various information between processor 46 and processor 28 via network 54. The exchange of various information between processor 46 and processor 28 using communication interfaces 44 and 26 is performed in a secure manner.
[0210] Figure 4 shows an example of the main functions of the data processing device 12 and the smart glasses 214. As shown in Figure 4, the data processing device 12 performs specific processing using the processor 28. The storage 32 stores the specific processing program 56.
[0211] The specific processing program 56 is an example of a "program" relating to the technology of this disclosure. The processor 28 reads the specific processing program 56 from the storage 32 and executes the read specific processing program 56 on the RAM 30. The specific processing is realized by the processor 28 operating as a specific processing unit 290 in accordance with the specific processing program 56 executed on the RAM 30.
[0212] The storage 32 stores the data generation model 58 and the emotion identification model 59. The data generation model 58 and the emotion identification model 59 are used by the identification processing unit 290.
[0213] In the smart glasses 214, the processor 46 performs the reception output processing. The storage 50 stores the reception output program 60. The processor 46 reads the reception output program 60 from the storage 50 and executes the read reception output program 60 on the RAM 48. The reception output processing is realized by the processor 46 operating as a control unit 46A according to the reception output program 60 executed on the RAM 48.
[0214] Next, the identification processing performed by the identification processing unit 290 of the data processing device 12 will be described. In the following description, the data processing device 12 will be referred to as the "server" and the smart glasses 214 will be referred to as the "terminal".
[0215] In embodiments of the present invention, the system includes a bio-worn terminal and integrates various functions to support natural conversation for the user. A detailed description of its operation and specific examples are provided below.
[0216] In this system, the terminal functions as a voice acquisition device, capturing ambient sound in real time. The acquired sound is converted into text data by the speech recognition device within the terminal. Speech recognition is performed using pre-trained acoustic and language models, ensuring accurate digital conversion of spoken content.
[0217] The converted text is analyzed using natural language processing tools. This analysis evaluates whether negative remarks are included and scores them using keyword lists and contextual analysis. Since the terminal processes this in real time, it is possible to provide immediate feedback to the user.
[0218] If negative comments are detected, the device provides feedback to the user via notification. These notifications are delivered through devices such as glasses, rings, or watches, using subtle vibrations or flashing lights to discreetly signal the user. This allows the user to become aware of their comments without interrupting the conversation and, if necessary, switch to positive language.
[0219] Furthermore, once the conversation ends, the data collected by the device is compiled by a report generation system, and the information is provided to the user's smartphone. The report includes the percentage and timing of negative comments, as well as specific advice for improvement, which users can use to improve their everyday communication skills.
[0220] For example, if a user says, "I'm really tired today, and I don't want to do anything anymore," the device detects negative phrases such as "tired" and "don't want to do anything." Then, a glasses-type device vibrates subtly to inform the user of this fact. After the conversation, the user receives suggestions on how to improve their expression through a report displayed on a smartphone app. In this way, the system provides users with a way to continuously train their positive communication skills.
[0221] The following describes the processing flow.
[0222] Step 1:
[0223] The device uses its built-in microphone to capture audio from the user's surroundings in real time. This audio data is captured within the device as a digital signal.
[0224] Step 2:
[0225] The device uses a speech recognition module to convert acquired speech data into text data. Speech recognition utilizes pre-trained acoustic and language models.
[0226] Step 3:
[0227] The terminal analyzes the converted text data using a natural language processing module. This module detects negative expressions using keyword lists and sentiment analysis algorithms, and calculates a negative score for each statement.
[0228] Step 4:
[0229] The device calculates a negative score and determines whether it exceeds a certain threshold. If it exceeds the threshold, the information is passed to the notification module.
[0230] Step 5:
[0231] The device will notify the user. Depending on the notification method, a device shaped like glasses or a ring will use vibration or flashing LEDs to inform the user of the presence of negative comments.
[0232] Step 6:
[0233] After the conversation ends, the device compiles all the spoken data and negative scores, and then organizes the analysis results.
[0234] Step 7:
[0235] The device generates a report using a smartphone app based on the collected data. This report includes the number and timing of negative comments, as well as suggestions for improvement.
[0236] Step 8:
[0237] Users can view reports on their smartphones and use them to improve their communication skills. This allows them to recognize their own negative tendencies and consciously switch to more positive expressions.
[0238] (Example 1)
[0239] Next, we will describe Example 1. In the following description, the data processing device 12 will be referred to as the "server," and the smart glasses 214 will be referred to as the "terminal."
[0240] This invention aims to solve the problem of users unconsciously creating emotional biases in their speech during everyday conversations. Therefore, it seeks to encourage improvement in expression by immediately identifying negative expressions used by users during conversations and providing feedback based on those findings. Conventional methods have faced challenges in visualizing emotional balance during conversations and providing timely feedback.
[0241] The identification process performed by the identification processing unit 290 of the data processing device 12 in Example 1 is realized by the following means.
[0242] In this invention, the server includes means for acquiring audio in real time, means for converting the acquired audio into text data, and means for using natural language processing techniques to measure emotional tendencies. This allows users to instantly recognize negative patterns in their own speech during a conversation and receive appropriate feedback, thereby improving the quality of communication.
[0243] "Means for acquiring audio in real time" refers to technologies and devices for collecting ambient audio signals and using them for subsequent processing.
[0244] "Means of converting to text data" refers to technology or equipment that converts acquired audio signals into text information, thereby achieving the digitization of audio information.
[0245] "Means for detecting negative descriptions" refers to algorithms and technologies for identifying negative emotions and expressions within text, thereby performing sentiment analysis of speech.
[0246] "Means of providing immediate notification" refers to devices or technologies for quickly communicating information to users based on detected negative reviews.
[0247] "Means of generating reports" refers to systems and technologies for creating systematic reports based on analysis results and providing them to users.
[0248] "Natural language processing technology" refers to technologies for processing human language using computers, and includes grammatical analysis, sentiment analysis, and other text processing.
[0249] "The shape of eyeglasses, rings, or watches" refers to the shape of a device that may be worn by a user and is a wearable device with notification capabilities.
[0250] This invention provides an information processing device to support natural conversations by users. Primarily, it utilizes a bio-worn terminal and has the function of capturing and analyzing the user's speech in real time.
[0251] The device uses advanced microphone technology to capture audio in real time. This audio input is optimized using filtering technology to remove background noise. The captured audio data is converted into text data by speech recognition software. This process utilizes pre-trained acoustic and language models to improve accuracy.
[0252] The converted text data is analyzed by the server using natural language processing techniques. This analysis determines whether negative descriptions are detected. Sentiment analysis includes keyword searching and contextual evaluation, and incorporates algorithms to measure emotional tendencies. The analysis results are immediately provided as feedback, using subtle vibrations or LED displays as notification methods. This allows users to receive appropriate feedback during conversations.
[0253] As a concrete example of its use, if a user says, "I'm really tired today, and I don't want to do anything anymore," the device will detect phrases like "tired" and "don't want to do anything" as negative. The glasses-type device will then vibrate to inform the user of this fact, providing an immediate opportunity for them to consciously correct their statement.
[0254] Furthermore, once the conversation ends, the server analyzes the results and sends a detailed report to the user's smartphone. This report includes areas for improvement in speech and specific advice, allowing users to receive feedback to improve their everyday communication skills. An example of a prompt might be a question such as, "What aspects of today's speech could be improved?"
[0255] The flow of the specific processing in Example 1 will be explained using Figure 11.
[0256] Step 1:
[0257] The device acquires the sound surrounding the user in real time through the microphone. This audio signal becomes the input. The device applies filtering to remove background noise and outputs clear audio data. This process allows for more accurate capture of the user's speech.
[0258] Step 2:
[0259] The device's speech recognition software takes filtered audio data as input and converts it into text data. This involves calculations based on a generative AI model that utilizes acoustic and linguistic models, outputting the audio signal as text information. Specifically, this process includes analyzing each phoneme and mapping it to the most matching word.
[0260] Step 3:
[0261] The server uses natural language processing techniques to analyze the generated text data as input. This analysis aims to detect negative descriptions, processing the data using keyword lists and contextual analysis, and outputting the results as an evaluation score. At this stage, sentiment tendencies become clear.
[0262] Step 4:
[0263] Based on the analysis results score, the device will immediately notify the user as needed. The input is the evaluation score, and the output is feedback through subtle vibrations and LED displays. Specifically, the glasses-type device vibrates to inform the user that negative comments have been detected.
[0264] Step 5:
[0265] Once the conversation ends, the server generates a detailed report based on the collected data. The input for this process is past conversation data and its analysis results, while the output is a report that includes the percentage of negative comments and specific advice for improvement. This report is provided through the user's smartphone app, allowing the user to review its contents and use them to improve future communication.
[0266] (Application Example 1)
[0267] Next, we will explain Application Example 1. In the following explanation, the data processing device 12 will be referred to as the "server," and the smart glasses 214 will be referred to as the "terminal."
[0268] In many situations, it is crucial to proactively detect negative remarks and risky conversations and respond appropriately. However, existing technologies have struggled to detect risks in real time and provide immediate feedback to operators. As a result, misunderstandings and conflicts can arise during conversations. In addition, the limited feedback available after conversations has hindered long-term improvement of communication skills.
[0269] The specific processing performed by the specific processing unit 290 of the data processing device 12 in Application Example 1 is realized by the following means.
[0270] In this invention, the server includes an acoustic acquisition means for acquiring sound in real time, a speech recognition means for converting the acquired sound into text data, a natural language analysis means for analyzing the text data to detect negative statements, and a warning means for detecting risks predicted during the operator's statements in real time and providing corresponding notifications. This enables the operator to instantly grasp negative phrases and potential risks and take appropriate action. Furthermore, the use of the report generation means and suggestion generation means contributes to improving negative statements and enhancing the quality of communication.
[0271] An "acoustic acquisition means" is a device that captures ambient sounds in real time and converts them into a format that can be processed within the system.
[0272] "Speech recognition means" refers to a processing device for converting acquired acoustic data into text format.
[0273] A "natural language analysis tool" is a device that analyzes text data and has the function of detecting negative statements and important language patterns.
[0274] A "notification means" is a device that provides visual or tactile feedback to the operator based on detected information.
[0275] A "report generation means" is a device that compiles the results of the analysis of conversation data and creates documents or digital reports to provide to the operator.
[0276] A "warning device" is a device that detects potential risks during speech in real time and immediately issues a warning to the operator.
[0277] A "suggestion generation device" is a device that provides feedback to the operator, including improvement measures and suggestions, based on detected negative statements and risks.
[0278] This invention is a system that includes sound acquisition means, speech recognition means, natural language analysis means, notification means, report generation means, warning means, and suggestion generation means. A server acts as the central point, monitoring user statements in real time and detecting negative phrases and risks.
[0279] First, a microphone is used to capture ambient sound. This acoustic data is then converted into text format by a speech recognition system. This process utilizes speech recognition technologies such as the Google Cloud Speech-to-Text API.
[0280] Next, natural language processing tools analyze the text data to detect negative statements and risks in real time. This utilizes natural language processing technologies such as the Google NLP API.
[0281] If detected, a notification system activates, sending a subtle vibration or light to the user via glasses or accessories. This allows the user to receive feedback without interrupting the conversation.
[0282] After the conversation ends, a detailed report is generated by the report generation mechanism. This report includes improvement suggestions and is provided to the user through a smartphone application. Based on this report, the system, through the suggestion generation mechanism, makes specific suggestions to support the user in improving their communication skills.
[0283] As a specific example, when a user says "I'm really tired today", the system immediately detects the negative phrase "tired". Then, the glasses-type device causes a slight vibration to prompt the user's conscious attention. In the subsequent report, suggestions on how to switch to positive expressions are made.
[0284] As examples of prompt sentences, questions such as "What negative phrases were detected in the current conversation?" and "What advice is there to improve these negative elements?" can be considered.
[0285] With this system, users can quickly grasp the risk elements during conversations and aim for continuous skill improvement.
[0286] The flow of the specific process in Application Example 1 will be described using FIG. 12.
[0287] Step 1:
[0288] Acoustic acquisition: The terminal uses a microphone to acquire ambient sound in real time. The input is environmental sound, and the output is acoustic data. Convert it to digital format as preparation for processing this acoustic data.
[0289] Step 2:
[0290] Speech recognition: The server uses speech recognition means to convert the acquired acoustic data into text data. The input is digital acoustic data, and the output is the corresponding text data. Use the Google Cloud Speech-to-Text API for speech recognition to precisely convert the content of the acoustic data into text.
[0291] Step 3:
[0292] Natural Language Processing: The server uses natural language analysis tools to analyze text data and detect negative statements. The input is text data, and the output is the analysis results and evaluation score. Using the Google NLP API, the text content is scanned, sentiment analysis is performed, and negative keywords are extracted.
[0293] Step 4:
[0294] Feedback Notification: The device uses notification methods to send a subtle vibration or light alert to the operator based on detected negative statements. The input is the analysis result, and the output is a physical notification. Feedback is provided discreetly by operating glasses-type devices or accessories.
[0295] Step 5:
[0296] Post-conversation report generation: The server uses a report generation mechanism to generate a report based on a detailed analysis of the conversation and sends it to the user's smartphone. The input is the analysis data of the entire conversation, and the output is a detailed report. This report includes suggestions for improvement.
[0297] Step 6:
[0298] Providing suggestions: The server provides users with suggestions for improving negative statements through a suggestion generation mechanism. The input is the content of the report, and the output is personalized suggestions. Users can utilize the provided advice in their daily communication.
[0299] Furthermore, an emotion engine that estimates the user's emotions may be incorporated. That is, the identification processing unit 290 may use the emotion identification model 59 to estimate the user's emotions and perform identification processing using the user's emotions.
[0300] This invention is a system that combines speech recognition technology, natural language processing technology, and an emotion engine to support improved user communication. This system is implemented as a bio-worn device that understands the user's natural conversational situation and provides real-time feedback.
[0301] In this system, the terminal first functions as a voice acquisition device, capturing surrounding sounds in real time. The acquired voice is converted into text data by a speech recognition device. This converted data is analyzed by a natural language processing device to evaluate whether it contains negative expressions.
[0302] The key is that an emotion engine is incorporated into this process. This emotion engine analyzes voice tone and contextual information of the utterance to recognize the user's emotional state. This allows for analysis that takes into account not only the words themselves, but also the psychological background behind them. For example, if someone says "I can't do this anymore" in a tired voice, it can determine whether that emotional state is depression or normal fatigue.
[0303] The system provides subtle notifications to the user in response to negative comments or emotional states. These notifications take the form of items such as glasses, rings, or watches, and provide feedback through vibration and light. The intensity and frequency of these notifications are adjusted based on data from the emotion engine, enabling more precise and personalized responses.
[0304] Furthermore, after a conversation, a detailed report based on the collected data is generated in the smartphone app. This report includes specific areas for improvement and suggestions for positive communication. Based on this information, users can effectively improve their communication skills. For example, if a user expresses emotional frustration, the system can analyze the tone and content and suggest alternative, positive expressions while preventing unnecessary intensity in notifications. In this way, it is an invention that comprehensively supports the user's emotions and the content of the conversation.
[0305] The processing flow will be described below.
[0306] Step 1:
[0307] The terminal uses the built-in microphone to acquire the surrounding sound of the user in real time. This audio data is captured into the terminal as a digital signal.
[0308] Step 2:
[0309] The terminal converts the acquired audio data into text data by means of speech recognition. At this time, an acoustic model and a language model are used to improve the accuracy of converting speech into text.
[0310] Step 3:
[0311] The terminal uses the natural language processing module to analyze the converted text data. In this analysis, a keyword list and context analysis are performed to evaluate whether negative statements are included.
[0312] Step 4:
[0313] The terminal uses the emotion engine to analyze the voice tone, speech content, and context, and recognize the user's emotional state. Based on this information, it is determined what kind of psychological state the user is speaking in.
[0314] Step 5:
[0315] The terminal determines whether a notification is necessary based on the negative score and emotional state. If necessary, a subtle notification is given to the user using vibration or light. This notification is adjusted in intensity and frequency according to the analysis results of the emotion engine.
[0316] Step 6:
[0317] After the conversation ends, the device collects all the data and generates a report based on the analysis of emotional state and content of speech. This report is sent to the user via a smartphone app.
[0318] Step 7:
[0319] The system censors reports provided by users via smartphone and uses the information gathered to improve their communication skills. The reports include negative aspects and emotional context of statements, as well as suggestions for improvement.
[0320] (Example 2)
[0321] Next, we will describe Example 2. In the following description, the data processing device 12 will be referred to as the "server" and the smart glasses 214 will be referred to as the "terminal".
[0322] In modern society, it is crucial to accurately understand an individual's emotional state and improve negative communication. However, conventional technologies are limited to analyzing simple text data obtained from voice, lacking the means to recognize emotional changes in real time and respond appropriately. As a result, users have difficulty accurately understanding their own emotional state, posing a challenge to improving the quality of communication.
[0323] The identification process performed by the identification processing unit 290 of the data processing device 12 in Example 2 is realized by the following means.
[0324] In this invention, the server includes an input means for acquiring audio in real time, a conversion means for converting the acquired audio into text, and an analysis means for analyzing the text and evaluating the emotional state. This allows for direct analysis of the emotional state from the audio data, rapid understanding of emotional changes in real-time dialogue, and the provision of appropriate feedback and improvement measures.
[0325] "Input means for acquiring audio in real time" refers to a device or function of a device that instantly captures audio present in the user's surroundings and provides the audio data for subsequent processing.
[0326] "A conversion method for converting acquired audio into text" refers to a system or algorithm that handles the process of converting audio input into textual information, and uses speech recognition technology to transcribe spoken content into text.
[0327] "Analysis means for analyzing text and evaluating emotional state" refers to a technology or process for analyzing converted text data and identifying and evaluating the emotional nuances and state of the user.
[0328] "Presentation means for providing physical notifications to the user based on emotional state" refers to a device or function for providing visual or tactile stimuli to the user and transmitting feedback based on the results of emotion analysis.
[0329] "A means for generating reports that include improvement suggestions while considering the user's emotional state" refers to a system or process for creating reports that provide specific areas for improvement and suggestions to the user based on the results of emotional and communication analysis.
[0330] A "wearable information display device" refers to a device that can be worn directly by the user and provides visual or tactile information.
[0331] This invention is a system that highly analyzes the user's emotional state and communication content, and provides appropriate feedback and improvement suggestions. The system consists of means for voice acquisition, speech recognition, natural language processing, emotion analysis, feedback notification, and report generation.
[0332] The device is equipped with an input mechanism for acquiring audio in real time, which is installed in wearable information display devices such as glasses, rings, and watches. This makes it possible to instantly collect the user's ambient sounds and conversation content. The collected audio data is sent to a server and converted into text using speech recognition technology such as Google Cloud Speech-to-Text.
[0333] Once the audio is converted to text, the server uses a natural language processing system to analyze the text data and evaluate the emotional state, particularly negative expressions. This evaluation also includes the user's voice tone and contextual information, allowing the emotion engine to analyze the underlying psychological nuances.
[0334] Based on the analysis results, the device provides visual or tactile notifications to the user. For example, if the device determines that the user is feeling fatigued, it will provide a soft vibration or light notification to encourage relaxation.
[0335] Furthermore, after a conversation session, the server generates a detailed report based on the analysis data and provides it to the user via a smartphone app. This report includes specific areas for improvement based on sentiment analysis and suggestions for positive communication.
[0336] For example, if a user says, "I'm feeling irritated today and can't concentrate on work," the system will detect that negative emotion and, taking into account the tone of voice, make suggestions to minimize its impact.
[0337] An example of a prompt for a generative AI model is: "Analyze the user's statement 'I'm tired today, but I'm satisfied,' determine the emotional state, and suggest feedback." This prompt allows the model to analyze the user's statement and the emotions behind it, and generate appropriate feedback.
[0338] The flow of the specific processing in Example 2 will be explained using Figure 13.
[0339] Step 1:
[0340] The device collects audio in real time. When the user begins speaking, the device uses its built-in microphone to input ambient sounds and generates digital audio data. This data is immediately sent to the server, enabling real-time processing.
[0341] Step 2:
[0342] The server inputs the received audio data into speech recognition software, which converts it into text data. This process uses a speech recognition engine, specifically analyzing the words in the audio and outputting the corresponding text. The output obtained at this stage is text information that includes all the spoken content.
[0343] Step 3:
[0344] The server passes text data to a natural language processing system for sentiment analysis. It analyzes specific keywords and phrases from the input text and applies natural language processing algorithms to evaluate negative or positive emotional states. The output here is an evaluation of the user's emotional state.
[0345] Step 4:
[0346] The server determines what to notify the user based on the sentiment analysis results. Based on the analysis information, it generates appropriate feedback and prepares data to present it to the user, for example, through a slight vibration or visual sign. The output is the specific feedback message to be notified.
[0347] Step 5:
[0348] The device uses the generated feedback message to physically notify the user. It uses vibration motors, LEDs, etc., to convey feedback to the user. At this stage, the user receives appropriate signals based on the sentiment analysis results.
[0349] Step 6:
[0350] After all processing is complete, the server generates a detailed report based on the conversation content and sentiment analysis results. This report includes suggestions for improvement and positive feedback. The generated report is sent to the user's smartphone app.
[0351] Step 7:
[0352] Users can review reports on a smartphone app and consider actions to improve their communication skills based on the suggested improvements. This allows for practical improvement while reflecting on daily conversations.
[0353] (Application Example 2)
[0354] Next, we will explain application example 2. In the following explanation, the data processing device 12 will be referred to as the "server," and the smart glasses 214 will be referred to as the "terminal."
[0355] With the advancement of technology, improving the user experience is becoming increasingly important in many industries. In particular, improving the quality of communication between staff and customers is a challenge in customer service at physical stores. However, it is difficult for staff to respond quickly to customers' emotions and negative comments using conventional methods. This invention aims to solve these problems and improve customer service at physical stores.
[0356] The specific processing performed by the specific processing unit 290 of the data processing device 12 in Application Example 2 is realized by the following means. In this invention, the server includes a voice acquisition means for acquiring voice in real time, a voice recognition means for converting the acquired voice into text data, and a natural language processing means for analyzing the text data to detect negative statements. This makes it possible to grasp the emotional state of customers in real time and notify staff of that information as needed.
[0357] "Voice acquisition means" refers to a means for acquiring ambient sounds in real time and using them for subsequent processing.
[0358] "Speech recognition means" refers to a means for converting acquired speech into text data.
[0359] "Natural language processing tools" are methods for analyzing text data to detect negative statements or specific emotional states.
[0360] "Notification means" refers to means of providing users with visual or tactile notifications based on detected negative statements and the user's emotional state.
[0361] A "record generation means" is a means for providing analysis results based on negative remarks and emotional evaluations after a conversation.
[0362] A "visual support device" is a device that helps users intuitively understand information by displaying it visually.
[0363] A "finger-worn device" is a device worn on the finger to provide physical or informational notifications.
[0364] A "wrist-worn device" is a device worn on the arm that provides physical or informational notifications.
[0365] "Positive dialogue methods" refer to methods and practical suggestions for users to engage in better communication.
[0366] The system for carrying out this invention consists of a process that acquires voice in real time, performs emotion analysis, and provides feedback to the user. Specifically, the terminal captures ambient sound using voice acquisition means. This voice is converted into text data using speech recognition means.
[0367] Text data is analyzed by natural language processing on the server to detect negative remarks and specific emotional states. The analyzed emotional data is used to determine the user's emotional state and the quality of communication. Based on the emotional state, the user is notified through visual support devices, finger-worn devices, and arm-worn devices. This allows the user to understand customer reactions in real time and take appropriate action.
[0368] Furthermore, after the conversation ends, the server uses a record generation mechanism to create a detailed report and suggests positive ways of interacting to the user. This report derives areas for improvement and recommendations through a generative AI model.
[0369] As a concrete example, this system could be used in physical stores to support staff in customer service. If negative remarks or feelings of depression are detected during a conversation with a customer, the staff will be immediately notified with information to address the issue. As a result, it becomes possible to provide higher quality customer support.
[0370] An example of a prompt is, "Create an app that analyzes customer conversation audio and provides staff with real-time emotional feedback."
[0371] The flow of a specific process in Application Example 2 will be explained using Figure 14.
[0372] Step 1:
[0373] The device uses voice acquisition means to capture ambient sounds in real time. The input is natural ambient sound, which is then captured as digital audio data using acoustic sensors and microphones. This process requires filtering out external noise and appropriately capturing the audio signal.
[0374] Step 2:
[0375] The audio data sent to the server is converted into text data by speech recognition technology. In this process, the input digital audio data is converted into string data using speech recognition software (e.g., Google Speech Recognition API). The resulting text data serves as the basis for communication analysis.
[0376] Step 3:
[0377] The server uses natural language processing (NLTK) to analyze text data and detect negative statements and specific emotional states. The input text data undergoes morphological analysis and sentiment scoring using natural language processing libraries (e.g., NLTK and SpaCy), and the presence or absence of negative expressions is output. This process also includes tone analysis by an emotion engine to understand the context.
[0378] Step 4:
[0379] The server provides feedback to the user via notification methods based on the analysis results. The input consists of analysis results of negative statements and emotions, and physical notification devices (e.g., visual support devices or vibration notification devices) are used for feedback. This output allows the user to understand the customer's emotional state in a timely manner and take appropriate action.
[0380] Step 5:
[0381] After the conversation ends, the server generates a detailed report using a record generation system and provides it to the user. Using a generation AI model, the report outputs information including areas for improvement and positive dialogue methods, based on the input analysis data. This report helps improve the user's communication skills.
[0382] The specific processing unit 290 transmits the result of the specific processing to the smart glasses 214. In the smart glasses 214, the control unit 46A causes the speaker 240 to output the result of the specific processing. The microphone 238 acquires audio indicating user input for the result of the specific processing. The control unit 46A transmits the audio data indicating user input acquired by the microphone 238 to the data processing unit 12. In the data processing unit 12, the specific processing unit 290 acquires the audio data.
[0383] Data generation model 58 is a type of so-called generative AI (Artificial Intelligence). One example of data generation model 58 is ChatGPT (Internet search<URL: https: / / openai.com / blog / chatgpt> ), Gemini (Internet search) <url: https: gemini.google.com ?hl="ja">Examples of generative AI include the following. The data generation model 58 is obtained by performing deep learning on a neural network. The data generation model 58 is input with prompts containing instructions, and with inference data such as audio data representing speech, text data representing text, and image data representing images. The data generation model 58 infers from the input inference data according to the instructions indicated by the prompts, and outputs the inference results in data formats such as audio data and text data. Here, inference refers to, for example, analysis, classification, prediction, and / or summarization.
[0384] In the above embodiment, an example was given in which specific processing is performed by the data processing device 12, but the technology of this disclosure is not limited thereto, and the specific processing may also be performed by the smart glasses 214.
[0385] [Third Embodiment]
[0386] Figure 5 shows an example of the configuration of the data processing system 310 according to the third embodiment.
[0387] As shown in Figure 5, the data processing system 310 includes a data processing device 12 and a headset terminal 314. An example of the data processing device 12 is a server.
[0388] The data processing device 12 comprises a computer 22, a database 24, and a communication interface 26. The computer 22 is an example of a "computer" related to the technology of this disclosure. The computer 22 comprises a processor 28, RAM 30, and storage 32. The processor 28, RAM 30, and storage 32 are connected to a bus 34. The database 24 and the communication interface 26 are also connected to the bus 34. The communication interface 26 is connected to a network 54. An example of the network 54 is a WAN (Wide Area Network) and / or a LAN (Local Area Network).
[0389] The headset terminal 314 includes a computer 36, a microphone 238, a speaker 240, a camera 42, a communication interface 44, and a display 343. The computer 36 includes a processor 46, RAM 48, and storage 50. The processor 46, RAM 48, and storage 50 are connected to a bus 52. The microphone 238, speaker 240, camera 42, and display 343 are also connected to the bus 52.
[0390] The microphone 238 receives voice signals from the user 20 and receives instructions from the user 20. The microphone 238 captures the voice signals from the user 20, converts the captured voice into audio data, and outputs it to the processor 46. The speaker 240 outputs audio according to the instructions from the processor 46.
[0391] Camera 42 is a small digital camera equipped with an optical system including a lens, aperture, and shutter, and an image sensor such as a CMOS (Complementary Metal-Oxide-Semiconductor) image sensor or a CCD (Charge Coupled Device) image sensor, and captures images of the area around the user 20 (for example, an imaging range defined by a field of view equivalent to the width of a typical healthy person's field of vision).
[0392] Communication interface 44 is connected to network 54. Communication interfaces 44 and 26 are responsible for the exchange of various information between processor 46 and processor 28 via network 54. The exchange of various information between processor 46 and processor 28 using communication interfaces 44 and 26 is performed in a secure manner.
[0393] Figure 6 shows an example of the main functions of the data processing device 12 and the headset terminal 314. As shown in Figure 6, the data processing device 12 performs specific processing using the processor 28. The storage 32 stores the specific processing program 56.
[0394] The specific processing program 56 is an example of a "program" relating to the technology of this disclosure. The processor 28 reads the specific processing program 56 from the storage 32 and executes the read specific processing program 56 on the RAM 30. The specific processing is realized by the processor 28 operating as a specific processing unit 290 in accordance with the specific processing program 56 executed on the RAM 30.
[0395] The storage 32 stores the data generation model 58 and the emotion identification model 59. The data generation model 58 and the emotion identification model 59 are used by the identification processing unit 290.
[0396] In the headset terminal 314, the processor 46 performs the reception output processing. The storage 50 stores the reception output program 60. The processor 46 reads the reception output program 60 from the storage 50 and executes the read reception output program 60 on the RAM 48. The reception output processing is realized by the processor 46 operating as a control unit 46A according to the reception output program 60 executed on the RAM 48.
[0397] Next, the specific processing performed by the specific processing unit 290 of the data processing device 12 will be described. In the following description, the data processing device 12 will be referred to as the "server" and the headset terminal 314 will be referred to as the "terminal".
[0398] In embodiments of the present invention, the system includes a bio-worn terminal and integrates various functions to support natural conversation for the user. A detailed description of its operation and specific examples are provided below.
[0399] In this system, the terminal functions as a voice acquisition device, capturing ambient sound in real time. The acquired sound is converted into text data by the speech recognition device within the terminal. Speech recognition is performed using pre-trained acoustic and language models, ensuring accurate digital conversion of spoken content.
[0400] The converted text is analyzed using natural language processing tools. This analysis evaluates whether negative remarks are included and scores them using keyword lists and contextual analysis. Since the terminal processes this in real time, it is possible to provide immediate feedback to the user.
[0401] If negative comments are detected, the device provides feedback to the user via notification. These notifications are delivered through devices such as glasses, rings, or watches, using subtle vibrations or flashing lights to discreetly signal the user. This allows the user to become aware of their comments without interrupting the conversation and, if necessary, switch to positive language.
[0402] Furthermore, once the conversation ends, the data collected by the device is compiled by a report generation system, and the information is provided to the user's smartphone. The report includes the percentage and timing of negative comments, as well as specific advice for improvement, which users can use to improve their everyday communication skills.
[0403] For example, if a user says, "I'm really tired today, and I don't want to do anything anymore," the device detects negative phrases such as "tired" and "don't want to do anything." Then, a glasses-type device vibrates subtly to inform the user of this fact. After the conversation, the user receives suggestions on how to improve their expression through a report displayed on a smartphone app. In this way, the system provides users with a way to continuously train their positive communication skills.
[0404] The following describes the processing flow.
[0405] Step 1:
[0406] The device uses its built-in microphone to capture audio from the user's surroundings in real time. This audio data is captured within the device as a digital signal.
[0407] Step 2:
[0408] The device uses a speech recognition module to convert acquired speech data into text data. Speech recognition utilizes pre-trained acoustic and language models.
[0409] Step 3:
[0410] The terminal analyzes the converted text data using a natural language processing module. This module detects negative expressions using keyword lists and sentiment analysis algorithms, and calculates a negative score for each statement.
[0411] Step 4:
[0412] The device calculates a negative score and determines whether it exceeds a certain threshold. If it exceeds the threshold, the information is passed to the notification module.
[0413] Step 5:
[0414] The device will notify the user. Depending on the notification method, a device shaped like glasses or a ring will use vibration or flashing LEDs to inform the user of the presence of negative comments.
[0415] Step 6:
[0416] After the conversation ends, the device compiles all the spoken data and negative scores, and then organizes the analysis results.
[0417] Step 7:
[0418] The device generates a report using a smartphone app based on the collected data. This report includes the number and timing of negative comments, as well as suggestions for improvement.
[0419] Step 8:
[0420] Users can view reports on their smartphones and use them to improve their communication skills. This allows them to recognize their own negative tendencies and consciously switch to more positive expressions.
[0421] (Example 1)
[0422] Next, we will describe Example 1. In the following description, the data processing device 12 will be referred to as the "server," and the headset-type terminal 314 will be referred to as the "terminal."
[0423] This invention aims to solve the problem of users unconsciously creating emotional biases in their speech during everyday conversations. Therefore, it seeks to encourage improvement in expression by immediately identifying negative expressions used by users during conversations and providing feedback based on those findings. Conventional methods have faced challenges in visualizing emotional balance during conversations and providing timely feedback.
[0424] The identification process performed by the identification processing unit 290 of the data processing device 12 in Example 1 is realized by the following means.
[0425] In this invention, the server includes means for acquiring audio in real time, means for converting the acquired audio into text data, and means for using natural language processing techniques to measure emotional tendencies. This allows users to instantly recognize negative patterns in their own speech during a conversation and receive appropriate feedback, thereby improving the quality of communication.
[0426] "Means for acquiring audio in real time" refers to technologies and devices for collecting ambient audio signals and using them for subsequent processing.
[0427] "Means of converting to text data" refers to technology or equipment that converts acquired audio signals into text information, thereby achieving the digitization of audio information.
[0428] "Means for detecting negative descriptions" refers to algorithms and technologies for identifying negative emotions and expressions within text, thereby performing sentiment analysis of speech.
[0429] "Means of providing immediate notification" refers to devices or technologies for quickly communicating information to users based on detected negative reviews.
[0430] "Means of generating reports" refers to systems and technologies for creating systematic reports based on analysis results and providing them to users.
[0431] "Natural language processing technology" refers to technologies for processing human language using computers, and includes grammatical analysis, sentiment analysis, and other text processing.
[0432] "The shape of eyeglasses, rings, or watches" refers to the shape of a device that may be worn by a user and is a wearable device with notification capabilities.
[0433] This invention provides an information processing device to support natural conversations by users. Primarily, it utilizes a bio-worn terminal and has the function of capturing and analyzing the user's speech in real time.
[0434] The device uses advanced microphone technology to capture audio in real time. This audio input is optimized using filtering technology to remove background noise. The captured audio data is converted into text data by speech recognition software. This process utilizes pre-trained acoustic and language models to improve accuracy.
[0435] The converted text data is analyzed by the server using natural language processing techniques. This analysis determines whether negative descriptions are detected. Sentiment analysis includes keyword searching and contextual evaluation, and incorporates algorithms to measure emotional tendencies. The analysis results are immediately provided as feedback, using subtle vibrations or LED displays as notification methods. This allows users to receive appropriate feedback during conversations.
[0436] As a concrete example of its use, if a user says, "I'm really tired today, and I don't want to do anything anymore," the device will detect phrases like "tired" and "don't want to do anything" as negative. The glasses-type device will then vibrate to inform the user of this fact, providing an immediate opportunity for them to consciously correct their statement.
[0437] Furthermore, once the conversation ends, the server analyzes the results and sends a detailed report to the user's smartphone. This report includes areas for improvement in speech and specific advice, allowing users to receive feedback to improve their everyday communication skills. An example of a prompt might be a question such as, "What aspects of today's speech could be improved?"
[0438] The flow of the specific processing in Example 1 will be explained using Figure 11.
[0439] Step 1:
[0440] The device acquires the sound surrounding the user in real time through the microphone. This audio signal becomes the input. The device applies filtering to remove background noise and outputs clear audio data. This process allows for more accurate capture of the user's speech.
[0441] Step 2:
[0442] The device's speech recognition software takes filtered audio data as input and converts it into text data. This involves calculations based on a generative AI model that utilizes acoustic and linguistic models, outputting the audio signal as text information. Specifically, this process includes analyzing each phoneme and mapping it to the most matching word.
[0443] Step 3:
[0444] The server uses natural language processing techniques to analyze the generated text data as input. This analysis aims to detect negative descriptions, processing the data using keyword lists and contextual analysis, and outputting the results as an evaluation score. At this stage, sentiment tendencies become clear.
[0445] Step 4:
[0446] Based on the analysis results score, the device will immediately notify the user as needed. The input is the evaluation score, and the output is feedback through subtle vibrations and LED displays. Specifically, the glasses-type device vibrates to inform the user that negative comments have been detected.
[0447] Step 5:
[0448] Once the conversation ends, the server generates a detailed report based on the collected data. The input for this process is past conversation data and its analysis results, while the output is a report that includes the percentage of negative comments and specific advice for improvement. This report is provided through the user's smartphone app, allowing the user to review its contents and use them to improve future communication.
[0449] (Application Example 1)
[0450] Next, we will explain Application Example 1. In the following explanation, the data processing device 12 will be referred to as the "server," and the headset-type terminal 314 will be referred to as the "terminal."
[0451] In many situations, it is crucial to proactively detect negative remarks and risky conversations and respond appropriately. However, existing technologies have struggled to detect risks in real time and provide immediate feedback to operators. As a result, misunderstandings and conflicts can arise during conversations. In addition, the limited feedback available after conversations has hindered long-term improvement of communication skills.
[0452] The specific processing performed by the specific processing unit 290 of the data processing device 12 in Application Example 1 is realized by the following means.
[0453] In this invention, the server includes an acoustic acquisition means for acquiring sound in real time, a speech recognition means for converting the acquired sound into text data, a natural language analysis means for analyzing the text data to detect negative statements, and a warning means for detecting risks predicted during the operator's statements in real time and providing corresponding notifications. This enables the operator to instantly grasp negative phrases and potential risks and take appropriate action. Furthermore, the use of the report generation means and suggestion generation means contributes to improving negative statements and enhancing the quality of communication.
[0454] An "acoustic acquisition means" is a device that captures ambient sounds in real time and converts them into a format that can be processed within the system.
[0455] "Speech recognition means" refers to a processing device for converting acquired acoustic data into text format.
[0456] A "natural language analysis tool" is a device that analyzes text data and has the function of detecting negative statements and important language patterns.
[0457] A "notification means" is a device that provides visual or tactile feedback to the operator based on detected information.
[0458] A "report generation means" is a device that compiles the results of the analysis of conversation data and creates documents or digital reports to provide to the operator.
[0459] A "warning device" is a device that detects potential risks during speech in real time and immediately issues a warning to the operator.
[0460] A "suggestion generation device" is a device that provides feedback to the operator, including improvement measures and suggestions, based on detected negative statements and risks.
[0461] This invention is a system that includes sound acquisition means, speech recognition means, natural language analysis means, notification means, report generation means, warning means, and suggestion generation means. A server acts as the central point, monitoring user statements in real time and detecting negative phrases and risks.
[0462] First, a microphone is used to capture ambient sound. This acoustic data is then converted into text format by a speech recognition system. This process utilizes speech recognition technologies such as the Google Cloud Speech-to-Text API.
[0463] Next, natural language processing tools analyze the text data to detect negative statements and risks in real time. This utilizes natural language processing technologies such as the Google NLP API.
[0464] If detected, a notification system activates, sending a subtle vibration or light to the user via glasses or accessories. This allows the user to receive feedback without interrupting the conversation.
[0465] After the conversation ends, a detailed report is generated by the report generation mechanism. This report includes improvement suggestions and is provided to the user through a smartphone application. Based on this report, the system, through the suggestion generation mechanism, makes specific suggestions to support the user in improving their communication skills.
[0466] For example, if a user says, "I'm really tired today," the system immediately detects the negative phrase "tired." The glasses-type device then vibrates slightly to draw the user's attention. A subsequent report offers suggestions on how to switch to more positive language.
[0467] Examples of prompts include questions such as, "What negative phrases were detected in the conversation?" or "What advice would you give to improve these negative elements?"
[0468] This system allows users to quickly identify risk factors during conversations and continuously improve their skills.
[0469] The flow of a specific process in Application Example 1 will be explained using Figure 12.
[0470] Step 1:
[0471] Acoustic Acquisition: The device uses a microphone to acquire ambient sounds in real time. The input is ambient sound, and the output is acoustic data. This acoustic data is converted to a digital format in preparation for processing.
[0472] Step 2:
[0473] Speech Recognition: The server uses speech recognition to convert acquired audio data into text data. The input is digital audio data, and the output is the corresponding text data. The Google Cloud Speech-to-Text API is used for speech recognition to accurately transcribe the content of the audio data into text.
[0474] Step 3:
[0475] Natural Language Processing: The server uses natural language analysis tools to analyze text data and detect negative statements. The input is text data, and the output is the analysis results and evaluation score. Using the Google NLP API, the text content is scanned, sentiment analysis is performed, and negative keywords are extracted.
[0476] Step 4:
[0477] Feedback Notification: The device uses notification methods to send a subtle vibration or light alert to the operator based on detected negative statements. The input is the analysis result, and the output is a physical notification. Feedback is provided discreetly by operating glasses-type devices or accessories.
[0478] Step 5:
[0479] Post-conversation report generation: The server uses a report generation mechanism to generate a report based on a detailed analysis of the conversation and sends it to the user's smartphone. The input is the analysis data of the entire conversation, and the output is a detailed report. This report includes suggestions for improvement.
[0480] Step 6:
[0481] Providing suggestions: The server provides users with suggestions for improving negative statements through a suggestion generation mechanism. The input is the content of the report, and the output is personalized suggestions. Users can utilize the provided advice in their daily communication.
[0482] Furthermore, an emotion engine that estimates the user's emotions may be incorporated. That is, the identification processing unit 290 may use the emotion identification model 59 to estimate the user's emotions and perform identification processing using the user's emotions.
[0483] This invention is a system that combines speech recognition technology, natural language processing technology, and an emotion engine to support improved user communication. This system is implemented as a bio-worn device that understands the user's natural conversational situation and provides real-time feedback.
[0484] In this system, the terminal first functions as a voice acquisition device, capturing surrounding sounds in real time. The acquired voice is converted into text data by a speech recognition device. This converted data is analyzed by a natural language processing device to evaluate whether it contains negative expressions.
[0485] The key is that an emotion engine is incorporated into this process. This emotion engine analyzes voice tone and contextual information of the utterance to recognize the user's emotional state. This allows for analysis that takes into account not only the words themselves, but also the psychological background behind them. For example, if someone says "I can't do this anymore" in a tired voice, it can determine whether that emotional state is depression or normal fatigue.
[0486] The system provides subtle notifications to the user in response to negative comments or emotional states. These notifications take the form of items such as glasses, rings, or watches, and provide feedback through vibration and light. The intensity and frequency of these notifications are adjusted based on data from the emotion engine, enabling more precise and personalized responses.
[0487] Furthermore, after a conversation, a detailed report based on the collected data is generated in the smartphone app. This report includes specific areas for improvement and suggestions for positive communication. Based on this information, users can effectively improve their communication skills. For example, if a user expresses emotional frustration, the system can analyze the tone and content and suggest alternative, positive expressions while preventing unnecessary intensity in notifications. In this way, it is an invention that comprehensively supports the user's emotions and the content of the conversation.
[0488] The following describes the processing flow.
[0489] Step 1:
[0490] The device uses its built-in microphone to capture audio from the user's surroundings in real time. This audio data is captured within the device as a digital signal.
[0491] Step 2:
[0492] The device converts the acquired audio data into text data using speech recognition technology. In this process, acoustic and language models are utilized to improve the accuracy of the speech-to-text conversion.
[0493] Step 3:
[0494] The device uses a natural language processing module to analyze the converted text data. This analysis includes keyword analysis and contextual analysis to assess whether negative statements are included.
[0495] Step 4:
[0496] The device uses an emotion engine to analyze voice tone, content, and context to recognize the user's emotional state. This information helps determine the user's psychological state when speaking.
[0497] Step 5:
[0498] The device determines whether a notification is necessary based on negative scores and emotional states. If necessary, it subtly notifies the user using vibration or light. The intensity and frequency of these notifications are adjusted according to the analysis results of the emotion engine.
[0499] Step 6:
[0500] After the conversation ends, the device collects all the data and generates a report based on the analysis of emotional state and content of speech. This report is sent to the user via a smartphone app.
[0501] Step 7:
[0502] The system censors reports provided by users via smartphone and uses the information gathered to improve their communication skills. The reports include negative aspects and emotional context of statements, as well as suggestions for improvement.
[0503] (Example 2)
[0504] Next, we will describe Example 2. In the following description, the data processing device 12 will be referred to as the "server," and the headset-type terminal 314 will be referred to as the "terminal."
[0505] In modern society, it is crucial to accurately understand an individual's emotional state and improve negative communication. However, conventional technologies are limited to analyzing simple text data obtained from voice, lacking the means to recognize emotional changes in real time and respond appropriately. As a result, users have difficulty accurately understanding their own emotional state, posing a challenge to improving the quality of communication.
[0506] The identification process performed by the identification processing unit 290 of the data processing device 12 in Example 2 is realized by the following means.
[0507] In this invention, the server includes an input means for acquiring audio in real time, a conversion means for converting the acquired audio into text, and an analysis means for analyzing the text and evaluating the emotional state. This allows for direct analysis of the emotional state from the audio data, rapid understanding of emotional changes in real-time dialogue, and the provision of appropriate feedback and improvement measures.
[0508] "Input means for acquiring audio in real time" refers to a device or function of a device that instantly captures audio present in the user's surroundings and provides the audio data for subsequent processing.
[0509] "A conversion method for converting acquired audio into text" refers to a system or algorithm that handles the process of converting audio input into textual information, and uses speech recognition technology to transcribe spoken content into text.
[0510] "Analysis means for analyzing text and evaluating emotional state" refers to a technology or process for analyzing converted text data and identifying and evaluating the emotional nuances and state of the user.
[0511] "Presentation means for providing physical notifications to the user based on emotional state" refers to a device or function for providing visual or tactile stimuli to the user and transmitting feedback based on the results of emotion analysis.
[0512] "A means for generating reports that include improvement suggestions while considering the user's emotional state" refers to a system or process for creating reports that provide specific areas for improvement and suggestions to the user based on the results of emotional and communication analysis.
[0513] A "wearable information display device" refers to a device that can be worn directly by the user and provides visual or tactile information.
[0514] This invention is a system that highly analyzes the user's emotional state and communication content, and provides appropriate feedback and improvement suggestions. The system consists of means for voice acquisition, speech recognition, natural language processing, emotion analysis, feedback notification, and report generation.
[0515] The device is equipped with an input mechanism for acquiring audio in real time, which is installed in wearable information display devices such as glasses, rings, and watches. This makes it possible to instantly collect the user's ambient sounds and conversation content. The collected audio data is sent to a server and converted into text using speech recognition technology such as Google Cloud Speech-to-Text.
[0516] Once the audio is converted to text, the server uses a natural language processing system to analyze the text data and evaluate the emotional state, particularly negative expressions. This evaluation also includes the user's voice tone and contextual information, allowing the emotion engine to analyze the underlying psychological nuances.
[0517] Based on the analysis results, the device provides visual or tactile notifications to the user. For example, if the device determines that the user is feeling fatigued, it will provide a soft vibration or light notification to encourage relaxation.
[0518] Furthermore, after a conversation session, the server generates a detailed report based on the analysis data and provides it to the user via a smartphone app. This report includes specific areas for improvement based on sentiment analysis and suggestions for positive communication.
[0519] For example, if a user says, "I'm feeling irritated today and can't concentrate on work," the system will detect that negative emotion and, taking into account the tone of voice, make suggestions to minimize its impact.
[0520] An example of a prompt for a generative AI model is: "Analyze the user's statement 'I'm tired today, but I'm satisfied,' determine the emotional state, and suggest feedback." This prompt allows the model to analyze the user's statement and the emotions behind it, and generate appropriate feedback.
[0521] The flow of the specific processing in Example 2 will be explained using Figure 13.
[0522] Step 1:
[0523] The device collects audio in real time. When the user begins speaking, the device uses its built-in microphone to input ambient sounds and generates digital audio data. This data is immediately sent to the server, enabling real-time processing.
[0524] Step 2:
[0525] The server inputs the received audio data into speech recognition software, which converts it into text data. This process uses a speech recognition engine, specifically analyzing the words in the audio and outputting the corresponding text. The output obtained at this stage is text information that includes all the spoken content.
[0526] Step 3:
[0527] The server passes text data to a natural language processing system for sentiment analysis. It analyzes specific keywords and phrases from the input text and applies natural language processing algorithms to evaluate negative or positive emotional states. The output here is an evaluation of the user's emotional state.
[0528] Step 4:
[0529] The server determines what to notify the user based on the sentiment analysis results. Based on the analysis information, it generates appropriate feedback and prepares data to present it to the user, for example, through a slight vibration or visual sign. The output is the specific feedback message to be notified.
[0530] Step 5:
[0531] The device uses the generated feedback message to physically notify the user. It uses vibration motors, LEDs, etc., to convey feedback to the user. At this stage, the user receives appropriate signals based on the sentiment analysis results.
[0532] Step 6:
[0533] After all processing is complete, the server generates a detailed report based on the conversation content and sentiment analysis results. This report includes suggestions for improvement and positive feedback. The generated report is sent to the user's smartphone app.
[0534] Step 7:
[0535] Users can review reports on a smartphone app and consider actions to improve their communication skills based on the suggested improvements. This allows for practical improvement while reflecting on daily conversations.
[0536] (Application Example 2)
[0537] Next, we will explain application example 2. In the following explanation, the data processing device 12 will be referred to as the "server," and the headset-type terminal 314 will be referred to as the "terminal."
[0538] With the advancement of technology, improving the user experience is becoming increasingly important in many industries. In particular, improving the quality of communication between staff and customers is a challenge in customer service at physical stores. However, it is difficult for staff to respond quickly to customers' emotions and negative comments using conventional methods. This invention aims to solve these problems and improve customer service at physical stores.
[0539] The specific processing performed by the specific processing unit 290 of the data processing device 12 in Application Example 2 is realized by the following means. In this invention, the server includes a voice acquisition means for acquiring voice in real time, a voice recognition means for converting the acquired voice into text data, and a natural language processing means for analyzing the text data to detect negative statements. This makes it possible to grasp the emotional state of customers in real time and notify staff of that information as needed.
[0540] "Voice acquisition means" refers to a means for acquiring ambient sounds in real time and using them for subsequent processing.
[0541] "Speech recognition means" refers to a means for converting acquired speech into text data.
[0542] "Natural language processing tools" are methods for analyzing text data to detect negative statements or specific emotional states.
[0543] "Notification means" refers to means of providing users with visual or tactile notifications based on detected negative statements and the user's emotional state.
[0544] A "record generation means" is a means for providing analysis results based on negative remarks and emotional evaluations after a conversation.
[0545] A "visual support device" is a device that helps users intuitively understand information by displaying it visually.
[0546] A "finger-worn device" is a device worn on the finger to provide physical or informational notifications.
[0547] A "wrist-worn device" is a device worn on the arm that provides physical or informational notifications.
[0548] "Positive dialogue methods" refer to methods and practical suggestions for users to engage in better communication.
[0549] The system for carrying out this invention consists of a process that acquires voice in real time, performs emotion analysis, and provides feedback to the user. Specifically, the terminal captures ambient sound using voice acquisition means. This voice is converted into text data using speech recognition means.
[0550] Text data is analyzed by natural language processing on the server to detect negative remarks and specific emotional states. The analyzed emotional data is used to determine the user's emotional state and the quality of communication. Based on the emotional state, the user is notified through visual support devices, finger-worn devices, and arm-worn devices. This allows the user to understand customer reactions in real time and take appropriate action.
[0551] Furthermore, after the conversation ends, the server uses a record generation mechanism to create a detailed report and suggests positive ways of interacting to the user. This report derives areas for improvement and recommendations through a generative AI model.
[0552] As a concrete example, this system could be used in physical stores to support staff in customer service. If negative remarks or feelings of depression are detected during a conversation with a customer, the staff will be immediately notified with information to address the issue. As a result, it becomes possible to provide higher quality customer support.
[0553] An example of a prompt is, "Create an app that analyzes customer conversation audio and provides staff with real-time emotional feedback."
[0554] The flow of a specific process in Application Example 2 will be explained using Figure 14.
[0555] Step 1:
[0556] The device uses voice acquisition means to capture ambient sounds in real time. The input is natural ambient sound, which is then captured as digital audio data using acoustic sensors and microphones. This process requires filtering out external noise and appropriately capturing the audio signal.
[0557] Step 2:
[0558] The audio data sent to the server is converted into text data by speech recognition technology. In this process, the input digital audio data is converted into string data using speech recognition software (e.g., Google Speech Recognition API). The resulting text data serves as the basis for communication analysis.
[0559] Step 3:
[0560] The server uses natural language processing (NLTK) to analyze text data and detect negative statements and specific emotional states. The input text data undergoes morphological analysis and sentiment scoring using natural language processing libraries (e.g., NLTK and SpaCy), and the presence or absence of negative expressions is output. This process also includes tone analysis by an emotion engine to understand the context.
[0561] Step 4:
[0562] The server provides feedback to the user via notification methods based on the analysis results. The input consists of analysis results of negative statements and emotions, and physical notification devices (e.g., visual support devices or vibration notification devices) are used for feedback. This output allows the user to understand the customer's emotional state in a timely manner and take appropriate action.
[0563] Step 5:
[0564] After the conversation ends, the server generates a detailed report using a record generation system and provides it to the user. Using a generation AI model, the report outputs information including areas for improvement and positive dialogue methods, based on the input analysis data. This report helps improve the user's communication skills.
[0565] The specific processing unit 290 transmits the result of the specific processing to the headset terminal 314. In the headset terminal 314, the control unit 46A causes the speaker 240 and display 343 to output the result of the specific processing. The microphone 238 acquires audio indicating user input for the result of the specific processing. The control unit 46A transmits the audio data indicating user input acquired by the microphone 238 to the data processing unit 12. In the data processing unit 12, the specific processing unit 290 acquires the audio data.
[0566] Data generation model 58 is a type of so-called generative AI (Artificial Intelligence). One example of data generation model 58 is ChatGPT (Internet search<URL: https: / / openai.com / blog / chatgpt> ), Gemini (Internet search) <url: https: gemini.google.com ?hl="ja">Examples of generative AI include the following. The data generation model 58 is obtained by performing deep learning on a neural network. The data generation model 58 is input with prompts containing instructions, and with inference data such as audio data representing speech, text data representing text, and image data representing images. The data generation model 58 infers from the input inference data according to the instructions indicated by the prompts, and outputs the inference results in data formats such as audio data and text data. Here, inference refers to, for example, analysis, classification, prediction, and / or summarization.
[0567] In the above embodiment, an example was given in which specific processing is performed by the data processing device 12, but the technology of this disclosure is not limited thereto, and specific processing may also be performed by the headset terminal 314.
[0568] [Fourth Embodiment]
[0569] Figure 7 shows an example of the configuration of the data processing system 410 according to the fourth embodiment.
[0570] As shown in Figure 7, the data processing system 410 includes a data processing device 12 and a robot 414. An example of the data processing device 12 is a server.
[0571] The data processing device 12 comprises a computer 22, a database 24, and a communication interface 26. The computer 22 is an example of a "computer" related to the technology of this disclosure. The computer 22 comprises a processor 28, RAM 30, and storage 32. The processor 28, RAM 30, and storage 32 are connected to a bus 34. The database 24 and the communication interface 26 are also connected to the bus 34. The communication interface 26 is connected to a network 54. An example of the network 54 is a WAN (Wide Area Network) and / or a LAN (Local Area Network).
[0572] The robot 414 includes a computer 36, a microphone 238, a speaker 240, a camera 42, a communication interface 44, and a controlled object 443. The computer 36 includes a processor 46, RAM 48, and storage 50. The processor 46, RAM 48, and storage 50 are connected to a bus 52. The microphone 238, speaker 240, camera 42, and controlled object 443 are also connected to the bus 52.
[0573] The microphone 238 receives voice signals from the user 20 and receives instructions from the user 20. The microphone 238 captures the voice signals from the user 20, converts the captured voice into audio data, and outputs it to the processor 46. The speaker 240 outputs audio according to the instructions from the processor 46.
[0574] Camera 42 is a small digital camera equipped with an optical system including a lens, aperture, and shutter, and an image sensor such as a CMOS (Complementary Metal-Oxide-Semiconductor) image sensor or a CCD (Charge Coupled Device) image sensor, and captures images of the area around the user 20 (for example, an imaging range defined by a field of view equivalent to the width of a typical healthy person's field of vision).
[0575] Communication interface 44 is connected to network 54. Communication interfaces 44 and 26 are responsible for the exchange of various information between processor 46 and processor 28 via network 54. The exchange of various information between processor 46 and processor 28 using communication interfaces 44 and 26 is performed in a secure manner.
[0576] The controlled object 443 includes a display device, LEDs in the eyes, and motors that drive the arms, hands, and feet. The posture and gestures of the robot 414 are controlled by controlling the motors of the arms, hands, and feet. Some of the robot 414's emotions can be expressed by controlling these motors. Furthermore, the robot 414's facial expressions can also be expressed by controlling the illumination state of the LEDs in its eyes.
[0577] Figure 8 shows an example of the main functions of the data processing device 12 and the robot 414. As shown in Figure 8, the data processing device 12 performs specific processing using the processor 28. The storage 32 stores the specific processing program 56.
[0578] The specific processing program 56 is an example of a "program" relating to the technology of this disclosure. The processor 28 reads the specific processing program 56 from the storage 32 and executes the read specific processing program 56 on the RAM 30. The specific processing is realized by the processor 28 operating as a specific processing unit 290 in accordance with the specific processing program 56 executed on the RAM 30.
[0579] The storage 32 stores the data generation model 58 and the emotion identification model 59. The data generation model 58 and the emotion identification model 59 are used by the identification processing unit 290.
[0580] In robot 414, the processor 46 performs the reception output processing. The storage 50 stores the reception output program 60. The processor 46 reads the reception output program 60 from the storage 50 and executes the read reception output program 60 on the RAM 48. The reception output processing is realized by the processor 46 operating as a control unit 46A according to the reception output program 60 executed on the RAM 48.
[0581] Next, the specific processing performed by the specific processing unit 290 of the data processing device 12 will be described. In the following description, the data processing device 12 will be referred to as the "server" and the robot 414 as the "terminal".
[0582] In embodiments of the present invention, the system includes a bio-worn terminal and integrates various functions to support natural conversation for the user. A detailed description of its operation and specific examples are provided below.
[0583] In this system, the terminal functions as a voice acquisition device, capturing ambient sound in real time. The acquired sound is converted into text data by the speech recognition device within the terminal. Speech recognition is performed using pre-trained acoustic and language models, ensuring accurate digital conversion of spoken content.
[0584] The converted text is analyzed using natural language processing tools. This analysis evaluates whether negative remarks are included and scores them using keyword lists and contextual analysis. Since the terminal processes this in real time, it is possible to provide immediate feedback to the user.
[0585] If negative comments are detected, the device provides feedback to the user via notification. These notifications are delivered through devices such as glasses, rings, or watches, using subtle vibrations or flashing lights to discreetly signal the user. This allows the user to become aware of their comments without interrupting the conversation and, if necessary, switch to positive language.
[0586] Furthermore, once the conversation ends, the data collected by the device is compiled by a report generation system, and the information is provided to the user's smartphone. The report includes the percentage and timing of negative comments, as well as specific advice for improvement, which users can use to improve their everyday communication skills.
[0587] For example, if a user says, "I'm really tired today, and I don't want to do anything anymore," the device detects negative phrases such as "tired" and "don't want to do anything." Then, a glasses-type device vibrates subtly to inform the user of this fact. After the conversation, the user receives suggestions on how to improve their expression through a report displayed on a smartphone app. In this way, the system provides users with a way to continuously train their positive communication skills.
[0588] The following describes the processing flow.
[0589] Step 1:
[0590] The device uses its built-in microphone to capture audio from the user's surroundings in real time. This audio data is captured within the device as a digital signal.
[0591] Step 2:
[0592] The device uses a speech recognition module to convert acquired speech data into text data. Speech recognition utilizes pre-trained acoustic and language models.
[0593] Step 3:
[0594] The terminal analyzes the converted text data using a natural language processing module. This module detects negative expressions using keyword lists and sentiment analysis algorithms, and calculates a negative score for each statement.
[0595] Step 4:
[0596] The device calculates a negative score and determines whether it exceeds a certain threshold. If it exceeds the threshold, the information is passed to the notification module.
[0597] Step 5:
[0598] The device will notify the user. Depending on the notification method, a device shaped like glasses or a ring will use vibration or flashing LEDs to inform the user of the presence of negative comments.
[0599] Step 6:
[0600] After the conversation ends, the device compiles all the spoken data and negative scores, and then organizes the analysis results.
[0601] Step 7:
[0602] The device generates a report using a smartphone app based on the collected data. This report includes the number and timing of negative comments, as well as suggestions for improvement.
[0603] Step 8:
[0604] Users can view reports on their smartphones and use them to improve their communication skills. This allows them to recognize their own negative tendencies and consciously switch to more positive expressions.
[0605] (Example 1)
[0606] Next, we will describe Example 1. In the following description, the data processing device 12 will be referred to as the "server" and the robot 414 as the "terminal".
[0607] This invention aims to solve the problem of users unconsciously creating emotional biases in their speech during everyday conversations. Therefore, it seeks to encourage improvement in expression by immediately identifying negative expressions used by users during conversations and providing feedback based on those findings. Conventional methods have faced challenges in visualizing emotional balance during conversations and providing timely feedback.
[0608] The identification process performed by the identification processing unit 290 of the data processing device 12 in Example 1 is realized by the following means.
[0609] In this invention, the server includes means for acquiring audio in real time, means for converting the acquired audio into text data, and means for using natural language processing techniques to measure emotional tendencies. This allows users to instantly recognize negative patterns in their own speech during a conversation and receive appropriate feedback, thereby improving the quality of communication.
[0610] "Means for acquiring audio in real time" refers to technologies and devices for collecting ambient audio signals and using them for subsequent processing.
[0611] "Means of converting to text data" refers to technology or equipment that converts acquired audio signals into text information, thereby achieving the digitization of audio information.
[0612] "Means for detecting negative descriptions" refers to algorithms and technologies for identifying negative emotions and expressions within text, thereby performing sentiment analysis of speech.
[0613] "Means of providing immediate notification" refers to devices or technologies for quickly communicating information to users based on detected negative reviews.
[0614] "Means of generating reports" refers to systems and technologies for creating systematic reports based on analysis results and providing them to users.
[0615] "Natural language processing technology" refers to technologies for processing human language using computers, and includes grammatical analysis, sentiment analysis, and other text processing.
[0616] "The shape of eyeglasses, rings, or watches" refers to the shape of a device that may be worn by a user and is a wearable device with notification capabilities.
[0617] This invention provides an information processing device to support natural conversations by users. Primarily, it utilizes a bio-worn terminal and has the function of capturing and analyzing the user's speech in real time.
[0618] The device uses advanced microphone technology to capture audio in real time. This audio input is optimized using filtering technology to remove background noise. The captured audio data is converted into text data by speech recognition software. This process utilizes pre-trained acoustic and language models to improve accuracy.
[0619] The converted text data is analyzed by the server using natural language processing techniques. This analysis determines whether negative descriptions are detected. Sentiment analysis includes keyword searching and contextual evaluation, and incorporates algorithms to measure emotional tendencies. The analysis results are immediately provided as feedback, using subtle vibrations or LED displays as notification methods. This allows users to receive appropriate feedback during conversations.
[0620] As a concrete example of its use, if a user says, "I'm really tired today, and I don't want to do anything anymore," the device will detect phrases like "tired" and "don't want to do anything" as negative. The glasses-type device will then vibrate to inform the user of this fact, providing an immediate opportunity for them to consciously correct their statement.
[0621] Furthermore, once the conversation ends, the server analyzes the results and sends a detailed report to the user's smartphone. This report includes areas for improvement in speech and specific advice, allowing users to receive feedback to improve their everyday communication skills. An example of a prompt might be a question such as, "What aspects of today's speech could be improved?"
[0622] The flow of the specific processing in Example 1 will be explained using Figure 11.
[0623] Step 1:
[0624] The device acquires the sound surrounding the user in real time through the microphone. This audio signal becomes the input. The device applies filtering to remove background noise and outputs clear audio data. This process allows for more accurate capture of the user's speech.
[0625] Step 2:
[0626] The device's speech recognition software takes filtered audio data as input and converts it into text data. This involves calculations based on a generative AI model that utilizes acoustic and linguistic models, outputting the audio signal as text information. Specifically, this process includes analyzing each phoneme and mapping it to the most matching word.
[0627] Step 3:
[0628] The server uses natural language processing techniques to analyze the generated text data as input. This analysis aims to detect negative descriptions, processing the data using keyword lists and contextual analysis, and outputting the results as an evaluation score. At this stage, sentiment tendencies become clear.
[0629] Step 4:
[0630] Based on the analysis results score, the device will immediately notify the user as needed. The input is the evaluation score, and the output is feedback through subtle vibrations and LED displays. Specifically, the glasses-type device vibrates to inform the user that negative comments have been detected.
[0631] Step 5:
[0632] Once the conversation ends, the server generates a detailed report based on the collected data. The input for this process is past conversation data and its analysis results, while the output is a report that includes the percentage of negative comments and specific advice for improvement. This report is provided through the user's smartphone app, allowing the user to review its contents and use them to improve future communication.
[0633] (Application Example 1)
[0634] Next, we will explain Application Example 1. In the following explanation, the data processing device 12 will be referred to as the "server" and the robot 414 as the "terminal".
[0635] In many situations, it is crucial to proactively detect negative remarks and risky conversations and respond appropriately. However, existing technologies have struggled to detect risks in real time and provide immediate feedback to operators. As a result, misunderstandings and conflicts can arise during conversations. In addition, the limited feedback available after conversations has hindered long-term improvement of communication skills.
[0636] The specific processing performed by the specific processing unit 290 of the data processing device 12 in Application Example 1 is realized by the following means.
[0637] In this invention, the server includes an acoustic acquisition means for acquiring sound in real time, a speech recognition means for converting the acquired sound into text data, a natural language analysis means for analyzing the text data to detect negative statements, and a warning means for detecting risks predicted during the operator's statements in real time and providing corresponding notifications. This enables the operator to instantly grasp negative phrases and potential risks and take appropriate action. Furthermore, the use of the report generation means and suggestion generation means contributes to improving negative statements and enhancing the quality of communication.
[0638] An "acoustic acquisition means" is a device that captures ambient sounds in real time and converts them into a format that can be processed within the system.
[0639] "Speech recognition means" refers to a processing device for converting acquired acoustic data into text format.
[0640] A "natural language analysis tool" is a device that analyzes text data and has the function of detecting negative statements and important language patterns.
[0641] A "notification means" is a device that provides visual or tactile feedback to the operator based on detected information.
[0642] A "report generation means" is a device that compiles the results of the analysis of conversation data and creates documents or digital reports to provide to the operator.
[0643] A "warning device" is a device that detects potential risks during speech in real time and immediately issues a warning to the operator.
[0644] A "suggestion generation device" is a device that provides feedback to the operator, including improvement measures and suggestions, based on detected negative statements and risks.
[0645] This invention is a system that includes sound acquisition means, speech recognition means, natural language analysis means, notification means, report generation means, warning means, and suggestion generation means. A server acts as the central point, monitoring user statements in real time and detecting negative phrases and risks.
[0646] First, a microphone is used to capture ambient sound. This acoustic data is then converted into text format by a speech recognition system. This process utilizes speech recognition technologies such as the Google Cloud Speech-to-Text API.
[0647] Next, natural language processing tools analyze the text data to detect negative statements and risks in real time. This utilizes natural language processing technologies such as the Google NLP API.
[0648] If detected, a notification system activates, sending a subtle vibration or light alert to the user through a glasses-type device or accessory. This allows the user to receive feedback without interrupting the conversation.
[0649] After the conversation ends, a detailed report is generated by the report generation mechanism. This report includes improvement suggestions and is provided to the user through a smartphone application. Based on this report, the system, through the suggestion generation mechanism, makes specific suggestions to support the user in improving their communication skills.
[0650] For example, if a user says, "I'm really tired today," the system immediately detects the negative phrase "tired." The glasses-type device then vibrates slightly to draw the user's attention. A subsequent report offers suggestions on how to switch to more positive language.
[0651] Examples of prompts include questions such as, "What negative phrases were detected in the conversation?" or "What advice would you give to improve these negative elements?"
[0652] This system allows users to quickly identify risk factors during conversations and continuously improve their skills.
[0653] The flow of a specific process in Application Example 1 will be explained using Figure 12.
[0654] Step 1:
[0655] Acoustic Acquisition: The device uses a microphone to acquire ambient sounds in real time. The input is ambient sound, and the output is acoustic data. This acoustic data is converted to a digital format in preparation for processing.
[0656] Step 2:
[0657] Speech Recognition: The server uses speech recognition to convert acquired audio data into text data. The input is digital audio data, and the output is the corresponding text data. The Google Cloud Speech-to-Text API is used for speech recognition to accurately transcribe the content of the audio data into text.
[0658] Step 3:
[0659] Natural Language Processing: The server uses natural language analysis tools to analyze text data and detect negative statements. The input is text data, and the output is the analysis results and evaluation score. Using the Google NLP API, the text content is scanned, sentiment analysis is performed, and negative keywords are extracted.
[0660] Step 4:
[0661] Feedback Notification: The device uses notification methods to send a subtle vibration or light alert to the operator based on detected negative statements. The input is the analysis result, and the output is a physical notification. Feedback is provided discreetly by operating glasses-type devices or accessories.
[0662] Step 5:
[0663] Post-conversation report generation: The server uses a report generation mechanism to generate a report based on a detailed analysis of the conversation and sends it to the user's smartphone. The input is the analysis data of the entire conversation, and the output is a detailed report. This report includes suggestions for improvement.
[0664] Step 6:
[0665] Providing suggestions: The server provides users with suggestions for improving negative statements through a suggestion generation mechanism. The input is the content of the report, and the output is personalized suggestions. Users can utilize the provided advice in their daily communication.
[0666] Furthermore, an emotion engine that estimates the user's emotions may be incorporated. That is, the identification processing unit 290 may use the emotion identification model 59 to estimate the user's emotions and perform identification processing using the user's emotions.
[0667] This invention is a system that combines speech recognition technology, natural language processing technology, and an emotion engine to support improved user communication. This system is implemented as a bio-worn device that understands the user's natural conversational situation and provides real-time feedback.
[0668] In this system, the terminal first functions as a voice acquisition device, capturing surrounding sounds in real time. The acquired voice is converted into text data by a speech recognition device. This converted data is analyzed by a natural language processing device to evaluate whether it contains negative expressions.
[0669] The key is that an emotion engine is incorporated into this process. This emotion engine analyzes voice tone and contextual information of the utterance to recognize the user's emotional state. This allows for analysis that takes into account not only the words themselves, but also the psychological background behind them. For example, if someone says "I can't do this anymore" in a tired voice, it can determine whether that emotional state is depression or normal fatigue.
[0670] The system provides subtle notifications to the user in response to negative comments or emotional states. These notifications take the form of items such as glasses, rings, or watches, and provide feedback through vibration and light. The intensity and frequency of these notifications are adjusted based on data from the emotion engine, enabling more precise and personalized responses.
[0671] Furthermore, after a conversation, a detailed report based on the collected data is generated in the smartphone app. This report includes specific areas for improvement and suggestions for positive communication. Based on this information, users can effectively improve their communication skills. For example, if a user expresses emotional frustration, the system can analyze the tone and content and suggest alternative, positive expressions while preventing unnecessary intensity in notifications. In this way, it is an invention that comprehensively supports the user's emotions and the content of the conversation.
[0672] The following describes the processing flow.
[0673] Step 1:
[0674] The device uses its built-in microphone to capture audio from the user's surroundings in real time. This audio data is captured within the device as a digital signal.
[0675] Step 2:
[0676] The device converts the acquired audio data into text data using speech recognition technology. In this process, acoustic and language models are utilized to improve the accuracy of the speech-to-text conversion.
[0677] Step 3:
[0678] The device uses a natural language processing module to analyze the converted text data. This analysis includes keyword analysis and contextual analysis to assess whether negative statements are included.
[0679] Step 4:
[0680] The device uses an emotion engine to analyze voice tone, content, and context to recognize the user's emotional state. This information helps determine the user's psychological state when speaking.
[0681] Step 5:
[0682] The device determines whether a notification is necessary based on negative scores and emotional states. If necessary, it subtly notifies the user using vibration or light. The intensity and frequency of these notifications are adjusted according to the analysis results of the emotion engine.
[0683] Step 6:
[0684] After the conversation ends, the device collects all the data and generates a report based on the analysis of emotional state and content of speech. This report is sent to the user via a smartphone app.
[0685] Step 7:
[0686] The system censors reports provided by users via smartphone and uses the information gathered to improve their communication skills. The reports include negative aspects and emotional context of statements, as well as suggestions for improvement.
[0687] (Example 2)
[0688] Next, we will describe Example 2. In the following description, the data processing device 12 will be referred to as the "server" and the robot 414 as the "terminal".
[0689] In modern society, it is crucial to accurately understand an individual's emotional state and improve negative communication. However, conventional technologies are limited to analyzing simple text data obtained from voice, lacking the means to recognize emotional changes in real time and respond appropriately. As a result, users have difficulty accurately understanding their own emotional state, posing a challenge to improving the quality of communication.
[0690] The identification process performed by the identification processing unit 290 of the data processing device 12 in Example 2 is realized by the following means.
[0691] In this invention, the server includes an input means for acquiring audio in real time, a conversion means for converting the acquired audio into text, and an analysis means for analyzing the text and evaluating the emotional state. This allows for direct analysis of the emotional state from the audio data, rapid understanding of emotional changes in real-time dialogue, and the provision of appropriate feedback and improvement measures.
[0692] "Input means for acquiring audio in real time" refers to a device or function of a device that instantly captures audio present in the user's surroundings and provides the audio data for subsequent processing.
[0693] "A conversion method for converting acquired audio into text" refers to a system or algorithm that handles the process of converting audio input into textual information, and uses speech recognition technology to transcribe spoken content into text.
[0694] "Analysis means for analyzing text and evaluating emotional state" refers to a technology or process for analyzing converted text data and identifying and evaluating the emotional nuances and state of the user.
[0695] "Presentation means for providing physical notifications to the user based on emotional state" refers to a device or function for providing visual or tactile stimuli to the user and transmitting feedback based on the results of emotion analysis.
[0696] "A means for generating reports that include improvement suggestions while considering the user's emotional state" refers to a system or process for creating reports that provide specific areas for improvement and suggestions to the user based on the results of emotional and communication analysis.
[0697] A "wearable information display device" refers to a device that can be worn directly by the user and provides visual or tactile information.
[0698] This invention is a system that highly analyzes the user's emotional state and communication content, and provides appropriate feedback and improvement suggestions. The system consists of means for voice acquisition, speech recognition, natural language processing, emotion analysis, feedback notification, and report generation.
[0699] The device is equipped with an input mechanism for acquiring audio in real time, which is installed in wearable information display devices such as glasses, rings, and watches. This makes it possible to instantly collect the user's ambient sounds and conversation content. The collected audio data is sent to a server and converted into text using speech recognition technology such as Google Cloud Speech-to-Text.
[0700] Once the audio is converted to text, the server uses a natural language processing system to analyze the text data and evaluate the emotional state, particularly negative expressions. This evaluation also includes the user's voice tone and contextual information, allowing the emotion engine to analyze the underlying psychological nuances.
[0701] Based on the analysis results, the device provides visual or tactile notifications to the user. For example, if the device determines that the user is feeling fatigued, it will provide a soft vibration or light notification to encourage relaxation.
[0702] Furthermore, after a conversation session, the server generates a detailed report based on the analysis data and provides it to the user via a smartphone app. This report includes specific areas for improvement based on sentiment analysis and suggestions for positive communication.
[0703] For example, if a user says, "I'm feeling irritated today and can't concentrate on work," the system will detect that negative emotion and, taking into account the tone of voice, make suggestions to minimize its impact.
[0704] An example of a prompt for a generative AI model is: "Analyze the user's statement 'I'm tired today, but I'm satisfied,' determine the emotional state, and suggest feedback." This prompt allows the model to analyze the user's statement and the emotions behind it, and generate appropriate feedback.
[0705] The flow of the specific processing in Example 2 will be explained using Figure 13.
[0706] Step 1:
[0707] The device collects audio in real time. When the user begins speaking, the device uses its built-in microphone to input ambient sounds and generates digital audio data. This data is immediately sent to the server, enabling real-time processing.
[0708] Step 2:
[0709] The server inputs the received audio data into speech recognition software, which converts it into text data. This process uses a speech recognition engine, specifically analyzing the words in the audio and outputting the corresponding text. The output obtained at this stage is text information that includes all the spoken content.
[0710] Step 3:
[0711] The server passes text data to a natural language processing system for sentiment analysis. It analyzes specific keywords and phrases from the input text and applies natural language processing algorithms to evaluate negative or positive emotional states. The output here is an evaluation of the user's emotional state.
[0712] Step 4:
[0713] The server determines what to notify the user based on the sentiment analysis results. Based on the analysis information, it generates appropriate feedback and prepares data to present it to the user, for example, through a slight vibration or visual sign. The output is the specific feedback message to be notified.
[0714] Step 5:
[0715] The device uses the generated feedback message to physically notify the user. It uses vibration motors, LEDs, etc., to convey feedback to the user. At this stage, the user receives appropriate signals based on the sentiment analysis results.
[0716] Step 6:
[0717] After all processing is complete, the server generates a detailed report based on the conversation content and sentiment analysis results. This report includes suggestions for improvement and positive feedback. The generated report is sent to the user's smartphone app.
[0718] Step 7:
[0719] Users can review reports on a smartphone app and consider actions to improve their communication skills based on the suggested improvements. This allows for practical improvement while reflecting on daily conversations.
[0720] (Application Example 2)
[0721] Next, we will explain application example 2. In the following explanation, the data processing device 12 will be referred to as the "server" and the robot 414 as the "terminal".
[0722] With the advancement of technology, improving the user experience is becoming increasingly important in many industries. In particular, improving the quality of communication between staff and customers is a challenge in customer service at physical stores. However, it is difficult for staff to respond quickly to customers' emotions and negative comments using conventional methods. This invention aims to solve these problems and improve customer service at physical stores.
[0723] The specific processing performed by the specific processing unit 290 of the data processing device 12 in Application Example 2 is realized by the following means. In this invention, the server includes a voice acquisition means for acquiring voice in real time, a voice recognition means for converting the acquired voice into text data, and a natural language processing means for analyzing the text data to detect negative statements. This makes it possible to grasp the emotional state of customers in real time and notify staff of that information as needed.
[0724] "Voice acquisition means" refers to a means for acquiring ambient sounds in real time and using them for subsequent processing.
[0725] "Speech recognition means" refers to a means for converting acquired speech into text data.
[0726] "Natural language processing tools" are methods for analyzing text data to detect negative statements or specific emotional states.
[0727] "Notification means" refers to means of providing users with visual or tactile notifications based on detected negative statements and the user's emotional state.
[0728] A "record generation means" is a means for providing analysis results based on negative remarks and emotional evaluations after a conversation.
[0729] A "visual support device" is a device that helps users intuitively understand information by displaying it visually.
[0730] A "finger-worn device" is a device worn on the finger to provide physical or informational notifications.
[0731] A "wrist-worn device" is a device worn on the arm that provides physical or informational notifications.
[0732] "Positive dialogue methods" refer to methods and practical suggestions for users to engage in better communication.
[0733] The system for carrying out this invention consists of a process that acquires voice in real time, performs emotion analysis, and provides feedback to the user. Specifically, the terminal captures ambient sound using voice acquisition means. This voice is converted into text data using speech recognition means.
[0734] Text data is analyzed by natural language processing on the server to detect negative remarks and specific emotional states. The analyzed emotional data is used to determine the user's emotional state and the quality of communication. Based on the emotional state, the user is notified through visual support devices, finger-worn devices, and arm-worn devices. This allows the user to understand customer reactions in real time and take appropriate action.
[0735] Furthermore, after the conversation ends, the server uses a record generation mechanism to create a detailed report and suggests positive ways of interacting to the user. This report derives areas for improvement and recommendations through a generative AI model.
[0736] As a concrete example, this system could be used in physical stores to support staff in customer service. If negative remarks or feelings of depression are detected during a conversation with a customer, the staff will be immediately notified with information to address the issue. As a result, it becomes possible to provide higher quality customer support.
[0737] An example of a prompt is, "Create an app that analyzes customer conversation audio and provides staff with real-time emotional feedback."
[0738] The flow of a specific process in Application Example 2 will be explained using Figure 14.
[0739] Step 1:
[0740] The device uses voice acquisition means to capture ambient sounds in real time. The input is natural ambient sound, which is then captured as digital audio data using acoustic sensors and microphones. This process requires filtering out external noise and appropriately capturing the audio signal.
[0741] Step 2:
[0742] The audio data sent to the server is converted into text data by speech recognition technology. In this process, the input digital audio data is converted into string data using speech recognition software (e.g., Google Speech Recognition API). The resulting text data serves as the basis for communication analysis.
[0743] Step 3:
[0744] The server uses natural language processing (NLTK) to analyze text data and detect negative statements and specific emotional states. The input text data undergoes morphological analysis and sentiment scoring using natural language processing libraries (e.g., NLTK and SpaCy), and the presence or absence of negative expressions is output. This process also includes tone analysis by an emotion engine to understand the context.
[0745] Step 4:
[0746] The server provides feedback to the user via notification methods based on the analysis results. The input consists of analysis results of negative statements and emotions, and physical notification devices (e.g., visual support devices or vibration notification devices) are used for feedback. This output allows the user to understand the customer's emotional state in a timely manner and take appropriate action.
[0747] Step 5:
[0748] After the conversation ends, the server generates a detailed report using a record generation system and provides it to the user. Using a generation AI model, the report outputs information including areas for improvement and positive dialogue methods, based on the input analysis data. This report helps improve the user's communication skills.
[0749] The specific processing unit 290 transmits the result of the specific processing to the robot 414. In the robot 414, the control unit 46A causes the speaker 240 and the controlled object 443 to output the result of the specific processing. The microphone 238 acquires audio indicating user input for the result of the specific processing. The control unit 46A transmits the audio data indicating user input acquired by the microphone 238 to the data processing unit 12. In the data processing unit 12, the specific processing unit 290 acquires the audio data.
[0750] Data generation model 58 is a type of so-called generative AI (Artificial Intelligence). One example of data generation model 58 is ChatGPT (Internet search<URL: https: / / openai.com / blog / chatgpt> ), Gemini (Internet search) <url: https: gemini.google.com ?hl="ja">Examples of generative AI include the following. The data generation model 58 is obtained by performing deep learning on a neural network. The data generation model 58 is input with prompts containing instructions, and with inference data such as audio data representing speech, text data representing text, and image data representing images. The data generation model 58 infers from the input inference data according to the instructions indicated by the prompts, and outputs the inference results in data formats such as audio data and text data. Here, inference refers to, for example, analysis, classification, prediction, and / or summarization.
[0751] In the above embodiment, an example was given in which specific processing is performed by the data processing device 12, but the technology of this disclosure is not limited thereto, and the specific processing may also be performed by the robot 414.
[0752] Furthermore, the emotion identification model 59, acting as an emotion engine, may determine the user's emotion according to a specific mapping. Specifically, the emotion identification model 59 may determine the user's emotion according to a specific mapping, which is an emotion map (see Figure 9). Similarly, the emotion identification model 59 may also determine the robot's emotion, and the identification processing unit 290 may perform identification processing using the robot's emotion.
[0753] Figure 9 shows an emotion map 400 in which multiple emotions are mapped. In the emotion map 400, emotions are arranged in concentric circles radiating from the center. The closer to the center of the concentric circles, the more primitive the emotions are located. Further out of the concentric circles, emotions representing states and actions arising from mental states are located. Emotion is a concept that includes feelings and mental states. On the left side of the concentric circles, emotions that are generally generated from reactions occurring in the brain are located. On the right side of the concentric circles, emotions that are generally induced by situational judgment are located. Above and below the concentric circles, emotions that are generally generated from reactions occurring in the brain and induced by situational judgment are located. In addition, the emotion of "pleasure" is located on the upper side of the concentric circles, and the emotion of "displeasure" is located on the lower side. Thus, in the emotion map 400, multiple emotions are mapped based on the structure in which emotions arise, and emotions that are likely to occur simultaneously are mapped close together.
[0754] These emotions are distributed at the 3 o'clock position on the Emotion Map 400, and usually fluctuate between feelings of security and anxiety. In the right half of the Emotion Map 400, situational awareness takes precedence over internal feelings, resulting in a calm impression.
[0755] The inside of the Emotion Map 400 represents inner thoughts, while the outside represents actions. Therefore, the further you go from the outside of the Emotion Map 400, the more visible (expressed in actions) your emotions become.
[0756] Here, human emotions are based on various balances, such as posture and blood sugar levels. When these balances deviate from the ideal, it results in discomfort, and when they approach the ideal, it results in pleasure. Similarly, in robots, cars, motorcycles, etc., emotions can be created based on various balances, such as posture and battery level. When these balances deviate from the ideal, it results in discomfort, and when they approach the ideal, it results in pleasure. The emotion map can be generated, for example, based on Dr. Mitsuyoshi's emotion map (Research on a system for analyzing brain physiological signals of speech emotion recognition and emotion, Tokushima University, doctoral dissertation: https: / / ci.nii.ac.jp / naid / 500000375379). The left half of the emotion map contains emotions belonging to a region called "response," where sensation is dominant. The right half of the emotion map contains emotions belonging to a region called "situation," where situational awareness is dominant.
[0757] The emotion map defines two emotions that promote learning. One is the emotion around the middle of the negative "repentance" and "reflection" on the situation side. In other words, it is when the robot experiences negative emotions such as "I never want to feel this way again" or "I don't want to be scolded again." The other is the emotion around the positive "desire" on the reaction side. In other words, it is when the robot has positive feelings such as "I want more" or "I want to know more."
[0758] The emotion identification model 59 inputs user input into a pre-trained neural network, obtains emotion values representing each emotion shown in the emotion map 400, and determines the user's emotion. This neural network is pre-trained based on multiple training data sets, which are combinations of user input and emotion values representing each emotion shown in the emotion map 400. Furthermore, this neural network is trained so that emotions located close together have similar values, as shown in the emotion map 900 in Figure 10. Figure 10 shows an example where multiple emotions such as "reassured," "calm," and "confident" have similar emotion values.
[0759] The above description primarily focuses on the functions of the data processing device 12 in relation to this disclosure. However, the system related to this disclosure is not necessarily implemented on a server. The system related to this disclosure may be implemented as a general information processing system. This disclosure may be implemented, for example, as a software program that runs on a personal computer or as an application that runs on a smartphone. The method related to this disclosure may be provided to users in SaaS (Software as a Service) format.
[0760] In the above embodiment, an example was given in which a specific process is performed by a single computer 22. However, the technology of this disclosure is not limited thereto, and a distributed processing of the specific process may be performed by multiple computers, including computer 22. For example, a data generation model 58 may be provided in an external device of the data processing device 12, and the external device may generate data according to the input data.
[0761] In the above embodiment, an example was given in which the specific processing program 56 is stored in the storage 32, but the technology of this disclosure is not limited thereto. For example, the specific processing program 56 may be stored in a portable, computer-readable, non-temporary storage medium such as a USB (Universal Serial Bus) memory. The specific processing program 56 stored in the non-temporary storage medium is installed in the computer 22 of the data processing device 12. The processor 28 executes specific processing according to the specific processing program 56.
[0762] Alternatively, the specific processing program 56 may be stored in a storage device such as a server connected to the data processing device 12 via the network 54, and the specific processing program 56 may be downloaded and installed on the computer 22 in response to a request from the data processing device 12.
[0763] Furthermore, it is not necessary to store the entirety of the specific processing program 56 in a storage device such as a server connected to the data processing device 12 via the network 54, or to store the entirety of the specific processing program 56 in the storage 32; it is acceptable to store only a portion of the specific processing program 56.
[0764] The following types of processors can be used as hardware resources to perform specific processing. Examples of processors include a CPU, a general-purpose processor that functions as a hardware resource to perform specific processing by executing software, i.e., a program. Other examples of processors include dedicated electrical circuits, such as FPGAs (Field-Programmable Gate Arrays), PLDs (Programmable Logic Devices), or ASICs (Application Specific Integrated Circuits), which have circuit configurations specifically designed to perform specific processing. All of these processors have built-in or connected memory, and all of them perform specific processing by using memory.
[0765] The hardware resource that performs a specific process may consist of one of these various processors, or it may consist of a combination of two or more processors of the same or different types (for example, a combination of multiple FPGAs, or a combination of a CPU and an FPGA). Alternatively, the hardware resource that performs a specific process may consist of a single processor.
[0766] Examples of configurations using a single processor include, firstly, a configuration in which one or more CPUs and software are combined to form a single processor, and this processor functions as a hardware resource that performs a specific process. Secondly, there is a configuration using a processor that realizes the functions of the entire system, including multiple hardware resources that perform a specific process, on a single IC chip, as exemplified by SoCs (System-on-a-chip). In this way, a specific process is realized using one or more of the above types of processors as hardware resources.
[0767] Furthermore, the hardware structure of these various processors can more specifically utilize electrical circuits that combine circuit elements such as semiconductor devices. Also, the specific processing described above is merely an example. Therefore, it goes without saying that unnecessary steps can be deleted, new steps added, or the processing order rearranged, as long as it does not deviate from the main purpose.
[0768] The descriptions and illustrations presented above are detailed explanations of the technical aspects of this disclosure and are merely examples of the technical aspects. For example, the above descriptions of the structure, function, operation, and effect are examples of the structure, function, operation, and effect of the technical aspects of this disclosure. Therefore, it goes without saying that you may delete unnecessary parts, add new elements, or replace elements in the descriptions and illustrations presented above, as long as you do not deviate from the essence of the technical aspects of this disclosure. Furthermore, in order to avoid confusion and facilitate understanding of the technical aspects of this disclosure, explanations of common technical knowledge and the like that do not require special explanation to enable the implementation of the technical aspects of this disclosure have been omitted from the descriptions and illustrations presented above.
[0769] All documents, patent applications, and technical standards described herein are incorporated by reference to the same extent as if each individual document, patent application, and technical standard were specifically and individually noted to be incorporated by reference.
[0770] The following is further disclosed regarding the embodiments described above.
[0771] (Claim 1)
[0772] A voice acquisition method that acquires voice in real time,
[0773] A speech recognition means that converts acquired audio into text data,
[0774] A natural language processing method that analyzes text data to detect negative statements,
[0775] A notification means that visually or tactilely notifies the user based on detected negative statements,
[0776] A report generation method that provides the results of an analysis of negative statements after a conversation,
[0777] A system that includes this.
[0778] (Claim 2)
[0779] The system according to claim 1, wherein the notification means has the shape of eyeglasses, a ring, or a wristwatch, and provides physical notification to the user.
[0780] (Claim 3)
[0781] The system according to claim 1, wherein the report generation means suggests to the user areas for improvement regarding negative statements.
[0782] "Example 1"
[0783] (Claim 1)
[0784] A means of acquiring audio in real time,
[0785] A means of converting acquired audio into text data,
[0786] A means of analyzing text data to detect negative descriptions,
[0787] A means of immediately notifying users based on detected negative reviews,
[0788] A means of generating a report that provides an analysis of negative descriptions after a conversation and suggests areas for improvement,
[0789] A method using natural language processing techniques to measure emotional tendencies,
[0790] Information processing device including
[0791] (Claim 2)
[0792] The information processing apparatus according to claim 1, wherein the notification means has the shape of eyeglasses, a ring, or a wristwatch, and provides tactile notification to the user.
[0793] (Claim 3)
[0794] The information processing apparatus according to claim 1, wherein the generated report includes specific suggestions for helping to improve the user's expression.
[0795] "Application Example 1"
[0796] (Claim 1)
[0797] A means of acquiring sound in real time,
[0798] A speech recognition means that converts acquired sound into text data,
[0799] A natural language analysis method that analyzes text data to detect negative statements,
[0800] A notification means that visually or tactilely notifies the operator based on the detected negative statement,
[0801] A report generation means that provides the results of analyzing negative statements after a conversation,
[0802] A warning system that detects foreseeable risks during the operator's speech in real time and issues corresponding notifications,
[0803] A proposal generation means that provides the operator with analysis results including suggestions for improvement based on predicted risks,
[0804] A system that includes this.
[0805] (Claim 2)
[0806] The system according to claim 1, wherein the notification means has the shape of a visual device, an ornament, or a wearable device, and provides physical notification to the operator.
[0807] (Claim 3)
[0808] The system according to claim 1, wherein the report generation means provides the operator with a continuous approach regarding areas for improvement in negative statements and anticipated risks.
[0809] "Example 2 of combining an emotion engine"
[0810] (Claim 1)
[0811] An input method for acquiring audio in real time,
[0812] A conversion method for converting acquired audio into text,
[0813] An analytical method for analyzing text and evaluating emotional state,
[0814] A means of providing physical notifications to the user based on their emotional state,
[0815] A generation means for creating a report that includes improvement suggestions, taking into account the user's emotional state,
[0816] A system that includes this.
[0817] (Claim 2)
[0818] The system according to claim 1, wherein the display means has the shape of a wearable information display device and provides notification to the user.
[0819] (Claim 3)
[0820] The system according to claim 1, wherein the generation means proposes improvement measures to the user based on the emotion analysis results.
[0821] "Application example 2 when combining with an emotional engine"
[0822] (Claim 1)
[0823] A voice acquisition method that acquires voice in real time,
[0824] A speech recognition means that converts acquired audio into text data,
[0825] A natural language processing method that analyzes text data to detect negative statements,
[0826] A notification system that provides visual or tactile notifications based on detected negative statements and the user's emotional state,
[0827] A record generation means that provides analysis results of negative statements and emotional evaluations after a conversation,
[0828] A system that includes this.
[0829] (Claim 2)
[0830] The system according to claim 1, wherein the notification means has the shape of a visual support device, a finger-worn device, or an arm-worn device, and provides physical or visual notification to the user.
[0831] (Claim 3)
[0832] The system according to claim 1, wherein the record generation means proposes to the user areas for improvement in negative statements and methods for positive dialogue. [Explanation of Symbols]
[0833] 10, 210, 310, 410 Data Processing Systems 12 Data Processing Devices 14 Smart Devices 214 Smart Glasses 314 Headset-type terminal 414 Robots< / url:> < / url:> < / url:> < / url:>
Claims
1. A voice acquisition method that acquires voice in real time, A speech recognition means that converts acquired audio into text data, A natural language processing method that analyzes text data to detect negative statements, A notification means that visually or tactilely notifies the user based on detected negative statements, A report generation method that provides the results of an analysis of negative statements after a conversation, A system that includes this.
2. The system according to claim 1, wherein the notification means has the shape of eyeglasses, a ring, or a wristwatch, and provides physical notification to the user.
3. The system according to claim 1, wherein the report generation means suggests to the user areas for improvement regarding negative statements.
Citation Information
Patent Citations
Persona chatbot control method and system
JP2022180282A