system

A real-time voice transcription and analysis system provides consistent, high-quality answers by recording, transcribing, and generating relevant information, addressing the variability in staff skills and enhancing customer satisfaction.

JP2026064641APending Publication Date: 2026-04-14SOFTBANK GROUP CORP
View PDF 1 Cites 0 Cited by

Patent Information

Authority / Receiving Office
JP · JP
Patent Type
Applications
Current Assignee / Owner
SOFTBANK GROUP CORP
Filing Date
2024-10-02
Publication Date
2026-04-14

AI Technical Summary

Technical Problem

Existing systems in call centers and sales settings face challenges in providing consistent, high-quality answers to inquiries due to variations in staff skills and knowledge, leading to inefficiencies and reduced customer satisfaction.

Method used

A system that records user voice in real-time, transcribes it into text, analyzes the data for important keywords and context, generates relevant answers using a server, and displays them on a terminal, enabling rapid and accurate information provision.

Benefits of technology

The system ensures consistent, high-quality answers and information delivery, reducing the burden on users and improving operational efficiency.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure 2026064641000001_ABST
    Figure 2026064641000001_ABST
Patent Text Reader

Abstract

We provide the system. [Solution] A means by which a terminal uses a communication method to record the user's voice and transcribe that voice into text in real time, A means by which the terminal sends the text data to the server, A means by which a server analyzes transcribed data, extracts important keywords and context, and generates answers and related information based on them, A means for the server to send the generated answers and related information to the terminal, The device displays answers and related information on the screen and notifies the user of important information via voice. A system that includes this.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The technology of the present disclosure relates to a system.

Background Art

[0002] Patent Document 1 discloses a method for controlling a persona chatbot, which is performed by at least one processor, including steps of receiving a user utterance, adding the user utterance to a prompt including an instruction sentence related to an explanation of a chatbot character, encoding the prompt, and inputting the encoded prompt into a language model to generate a chatbot utterance in response to the user utterance.

Prior Art Documents

Patent Documents

[0003]

Patent Document 1

Summary of the Invention

Problems to be Solved by the Invention

[0004] In recent years, in call centers and sales settings, there has been a demand for providing quick and accurate answers to inquiries and consultations regarding products and services. However, due to differences in the skills and knowledge of the staff, there is a problem that the quality and accuracy of the provided answers vary. There is a need for a system that solves this problem and provides consistent high-quality answers without relying on the skills of the staff. In addition, this system is required to be operated in real time and provide information efficiently while reducing the burden on the user.

Means for Solving the Problems

[0005] To solve the above problems, the present invention provides the following means. First, the terminal is provided with means for recording the user's voice using communication means and transcribing the voice in real time. Next, the terminal is provided with means for transmitting the transcribed data to a server. The server is provided with means for analyzing the received transcribed data, extracting important keywords and context, and generating answers and related information based on them. These generated answers and related information are then transmitted from the server back to the terminal. Finally, the terminal is provided with means for displaying the received answers and related information on a screen and notifying the user of important information by voice, thereby realizing rapid and accurate information provision.

[0006] A "terminal" is a device used by a user to record audio and communicate with a server.

[0007] "Recording audio" means acquiring the content of what the user says as digital data.

[0008] "Real-time" means that processing is performed instantly without delay.

[0009] "Transcription" is the process of converting audio data into text data.

[0010] "Communication methods" refer to the technologies and protocols used to transmit data to other devices or servers.

[0011] A "server" is a central processing unit that processes, analyzes, and generates data it receives.

[0012] "Analysis" is the process of extracting and understanding necessary information based on received text data.

[0013] "Keywords" refer to important words or phrases in the consultation.

[0014] "Context" refers to the relationships and background information between keywords in text data.

[0015] "Answer" refers to an appropriate response to the user's consultation generated by the server.

[0016] "Relevant information" refers to additional information and materials related to the user's consultation.

[0017] "Display on the screen" means visually presenting information on the display of the terminal.

[0018] "Notify by voice" means the terminal conveys information to the user by voice.

Brief Description of the Drawings

[0019] [Figure 1] It is a conceptual diagram showing an example of the configuration of a data processing system according to the first embodiment. [Figure 2] It is a conceptual diagram showing an example of the main functions of a data processing device and a smart device according to the first embodiment. [Figure 3] It is a conceptual diagram showing an example of the configuration of a data processing system according to the second embodiment. [Figure 4] It is a conceptual diagram showing an example of the main functions of a data processing device and smart glasses according to the second embodiment. [Figure 5] It is a conceptual diagram showing an example of the configuration of a data processing system according to the third embodiment. [Figure 6] It is a conceptual diagram showing an example of the main functions of a data processing device and a headset-type terminal according to the third embodiment. [Figure 7] It is a conceptual diagram showing an example of the configuration of a data processing system according to the fourth embodiment. [Figure 8] It is a conceptual diagram showing an example of the main functions of a data processing device and a robot according to the fourth embodiment. [Figure 9] It shows an emotion map to which a plurality of emotions are mapped. [Figure 10] It shows an emotion map to which a plurality of emotions are mapped. [Figure 11] It is a sequence diagram showing the processing flow of the data processing system in Example 1. [Figure 12] It is a sequence diagram showing the processing flow of the data processing system in Application Example 1. [Figure 13] It is a sequence diagram showing the processing flow of the data processing system in Example 2 when combined with an emotion engine. [Figure 14] It is a sequence diagram showing the processing flow of the data processing system in Application Example 2 when combined with an emotion engine.

Mode for Carrying Out the Invention

[0020] Hereinafter, an example of an embodiment of the system according to the technology of the present disclosure will be described with reference to the accompanying drawings.

[0021] First, the terms used in the following description will be explained.

[0022] In the following embodiments, a numbered processor (hereinafter simply referred to as "processor") may be a single arithmetic unit or a combination of multiple arithmetic units. Also, the processor may be a single type of arithmetic unit or a combination of multiple types of arithmetic units. Examples of arithmetic units include a CPU (Central Processing Unit), a GPU (Graphics Processing Unit), a GPGPU (General-Purpose computing on Graphics Processing Units), an APU (Accelerated Processing Unit), and the like.

[0023] In the following embodiments, a numbered RAM (Random Access Memory) is a memory in which information is temporarily stored and is used as a work memory by the processor.

[0024] In the following embodiments, the signed storage is one or more non-volatile storage devices that store various programs and various parameters. Examples of non-volatile storage devices include flash memory (SSD (Solid State Drive)), magnetic disks (e.g., hard disks), or magnetic tapes.

[0025] In the following embodiments, the signed communication interface (I / F) is an interface that includes a communication processor and an antenna, etc. The communication interface manages communication between multiple computers. Examples of communication standards applicable to the communication interface include wireless communication standards such as 5G (5th Generation Mobile Communication System), Wi-Fi (registered trademark), or Bluetooth (registered trademark).

[0026] In the following embodiments, "A and / or B" is synonymous with "at least one of A and B." That is, "A and / or B" means that it may be A alone, or B alone, or a combination of A and B. Furthermore, in this specification, the same concept as "A and / or B" applies when expressing three or more things linked by "and / or."

[0027] [First Embodiment]

[0028] Figure 1 shows an example of the configuration of the data processing system 10 according to the first embodiment.

[0029] As shown in Figure 1, the data processing system 10 includes a data processing device 12 and a smart device 14. An example of the data processing device 12 is a server.

[0030] The data processing device 12 comprises a computer 22, a database 24, and a communication interface 26. The computer 22 is an example of a "computer" related to the technology of this disclosure. The computer 22 comprises a processor 28, RAM 30, and storage 32. The processor 28, RAM 30, and storage 32 are connected to a bus 34. The database 24 and the communication interface 26 are also connected to the bus 34. The communication interface 26 is connected to a network 54. An example of the network 54 is a WAN (Wide Area Network) and / or a LAN (Local Area Network).

[0031] The smart device 14 comprises a computer 36, a reception device 38, an output device 40, a camera 42, and a communication interface 44. The computer 36 comprises a processor 46, RAM 48, and storage 50. The processor 46, RAM 48, and storage 50 are connected to a bus 52. The reception device 38, output device 40, and camera 42 are also connected to the bus 52.

[0032] The reception device 38 is equipped with a touch panel 38A and a microphone 38B, etc., and receives user input. The touch panel 38A receives user input by detecting contact with an object (e.g., a pen or finger). The microphone 38B receives user input by detecting the user's voice. The control unit 46A transmits data indicating the user input received by the touch panel 38A and microphone 38B to the data processing device 12. In the data processing device 12, the specific processing unit 290 acquires the data indicating the user input.

[0033] The output device 40 includes a display 40A and a speaker 40B, and presents data to the user 20 by outputting the data in a form perceptible to the user 20 (e.g., audio and / or text). The display 40A displays visible information such as text and images according to instructions from the processor 46. The speaker 40B outputs audio according to instructions from the processor 46. The camera 42 is a small digital camera equipped with an optical system such as a lens, aperture, and shutter, and an image sensor such as a CMOS (Complementary Metal-Oxide-Semiconductor) image sensor or a CCD (Charge Coupled Device) image sensor.

[0034] Communication interface 44 is connected to network 54. Communication interfaces 44 and 26 are responsible for the exchange of various types of information between processor 46 and processor 28 via network 54.

[0035] Figure 2 shows an example of the main functions of the data processing device 12 and the smart device 14.

[0036] As shown in Figure 2, in the data processing device 12, a specific processing is performed by the processor 28. A specific processing program 56 is stored in the storage 32. The specific processing program 56 is an example of a "program" related to the technology of this disclosure. The processor 28 reads the specific processing program 56 from the storage 32 and executes the read specific processing program 56 on the RAM 30. The specific processing is realized by the processor 28 operating as a specific processing unit 290 according to the specific processing program 56 executed on the RAM 30.

[0037] The storage 32 stores the data generation model 58 and the emotion identification model 59. The data generation model 58 and the emotion identification model 59 are used by the identification processing unit 290.

[0038] In the smart device 14, the processor 46 performs the reception output processing. The storage 50 stores the reception output program 60. The reception output program 60 is used in conjunction with a specific processing program 56 by the data processing system 10. The processor 46 reads the reception output program 60 from the storage 50 and executes the read reception output program 60 on the RAM 48. The reception output processing is realized by the processor 46 operating as a control unit 46A according to the reception output program 60 executed on the RAM 48.

[0039] Next, the specific processing performed by the specific processing unit 290 of the data processing device 12 will be described. In the following description, the data processing device 12 will be referred to as the "server" and the smart device 14 as the "terminal".

[0040] System Overview

[0041] This invention relates to a system that analyzes telephone and in-person consultations in real time and provides appropriate answers and related information. This system consists of a user's terminal, a cloud-based server, and communication means. The following describes each part of the system in detail.

[0042] Voice input and text conversion

[0043] 1. Users receive consultations via telephone or in person.

[0044] Users will use their own devices (smartphones, PCs, smart glasses, etc.) to provide consultation.

[0045] 2. The device records the user's consultation content as audio.

[0046] The device's application uses the microphone function to capture the audio of the consultation in real time. The audio data is temporarily stored in a buffer.

[0047] 3. The device uses a speech recognition engine to convert speech into text.

[0048] The device sends voice data to a speech recognition engine, which converts it into text data. The converted text is then stored within the application.

[0049] Text data transmission and analysis

[0050] 4. The device sends the transcribed consultation content to the server.

[0051] Text data is sent from the device to a server in the cloud using the HTTPS protocol.

[0052] 5. The server receives the transcribed data and parses it.

[0053] The server's natural language processing (NLP) engine analyzes the text data and extracts important keywords and context.

[0054] Solution generation and search

[0055] 6. The server uses generative AI to generate answers and related information.

[0056] The server generates answers and related information using a database of FAQs related to the problem and a pre-trained generative AI model.

[0057] 7. The server filters the generated answers and related information and selects the appropriate ones.

[0058] From the multiple generated answers, filter and select the most appropriate information.

[0059] Submit and display of answers

[0060] 8. The server sends the answer and related information to the terminal.

[0061] Encode the optimal answers, links, and materials and send them back to your device.

[0062] 9. The device notifies the user of the answer and related information.

[0063] The application on the device displays the received answers and information on the screen and notifies the user by voice as needed.

[0064] Specific example

[0065] For customer support

[0066] 1. A user contacts us because the product is not working.

[0067] 2. The device records the audio and transcribes the voice message, "The product isn't working, what should I do?"

[0068] 3. The device sends the converted data to the server.

[0069] 4. The server analyzes the text data and extracts the keywords "product" and "not working".

[0070] 5. The server refers to the relevant FAQ database and generates answers, reset instructions, and support video links.

[0071] 6. Filter the information generated by the server and send it to the terminal.

[0072] 7. The device will display instructions on how to reset the device and a link to a support video on the screen, and will announce a voice message saying, "Here's how to reset the device."

[0073] In the case of sales

[0074] 1. The user receives an inquiry about the details of the new product.

[0075] 2. The device records the audio and transcribes the phrase "Please tell me more about the new product" into text.

[0076] 3. The device sends the converted data to the server.

[0077] 4. The server analyzes the text data and extracts the keywords "new product" and "details".

[0078] 5. The server generates related sales materials, product brochures, and links to explanatory videos.

[0079] 6. Filter the information generated by the server and send it to the terminal.

[0080] 7. The device displays a link to a brochure and product features on the screen, and announces with a voice message, "Here are the details of the new product."

[0081] Thus, the system of the present invention enables consistent, high-quality answers and information provision, regardless of the user's skill level. It offers high practicality in the field and contributes to improved work efficiency.

[0082] The following describes the processing flow.

[0083] Step 1:

[0084] Users receive consultations via phone or in person.

[0085] Users will use their own devices (smartphones, PCs, smart glasses, etc.) to provide consultation.

[0086] Step 2:

[0087] The device records the user's conversation audio.

[0088] The device's application uses the microphone function to capture the audio of the consultation in real time. The audio data is temporarily stored in a buffer.

[0089] Step 3:

[0090] The device uses a speech recognition engine to convert speech into text.

[0091] The device sends voice data to a speech recognition engine, which converts it into text data. The converted text is then stored within the application.

[0092] Step 4:

[0093] The terminal sends the transcribed consultation content to the server.

[0094] Text data is sent from the device to a server in the cloud using the HTTPS protocol.

[0095] Step 5:

[0096] The server receives the transcribed data and parses it.

[0097] The server's natural language processing (NLP) engine analyzes the text data and extracts important keywords and context.

[0098] Step 6:

[0099] The server uses generative AI to generate answers and related information.

[0100] The server generates answers and related information using a database of FAQs related to the problem and a pre-trained generative AI model.

[0101] Step 7:

[0102] The server filters the generated answers and related information and selects the appropriate ones.

[0103] From the multiple generated answers, filter and select the most appropriate information.

[0104] Step 8:

[0105] The server sends the answer and related information to the terminal.

[0106] Encode the optimal answers, links, and materials and send them back to your device.

[0107] Step 9:

[0108] The device notifies the user of the answer and related information.

[0109] The application on the device displays the received answers and information on the screen and notifies the user by voice as needed.

[0110] (Example 1)

[0111] Next, we will describe Example 1. In the following description, the data processing device 12 will be referred to as the "server," and the smart device 14 will be referred to as the "terminal."

[0112] In modern call centers and customer support, there is a demand for providing quick and accurate answers to user inquiries. However, conventional systems often suffer from low accuracy in voice-to-text conversion, and answer generation is frequently done manually, leading to delays and inconsistent quality. This has resulted in decreased customer satisfaction and reduced operational efficiency. This invention aims to solve these problems, streamline the process from voice input to answer provision, and achieve consistent, high-quality information delivery.

[0113] The identification process performed by the identification processing unit 290 of the data processing device 12 in Example 1 is realized by the following means.

[0114] In this invention, the server includes: means for the terminal to record the user's voice using communication means and transcribe the voice in real time; means for the terminal to transmit the transcribed data to the server; means for the server to analyze the transcribed data, extract important keywords and context, and generate answers and related information based on them; means for the server to filter the generated answers and related information and transmit it to the terminal; and means for the terminal to display the answers and related information on the screen and notify the user of important information by voice. This automates the process from voice input to answer provision, enabling the rapid, consistent, and high-quality provision of information.

[0115] A "terminal" is an information processing device used by a user, which acquires and processes data, including voice input. Specifically, this refers to smartphones, personal computers, smart glasses, etc.

[0116] "Communication means" refers to technical means for sending and receiving data, including the internet, Wi-Fi, Bluetooth, etc.

[0117] "Recording audio" means capturing a user's speech as digital audio data using a microphone and storing it.

[0118] "Real-time transcription of speech" means instantly analyzing speech data and converting it into corresponding text data. This is typically done using a speech recognition engine.

[0119] "Transcribed data" refers to text data obtained after converting voice input into written information.

[0120] A "server" is a high-performance information processing device that operates on the cloud and has functions such as data analysis, storage, and transmission.

[0121] "Analysis" is the process of breaking down received data and performing operations to understand it.

[0122] "Important keywords and context" refer to words and phrases that are particularly meaningful from the input data, as well as the relationships between them.

[0123] "Answers and related information" refers to information that allows for appropriate responses to user questions and inquiries. This includes generated text, links, and documents.

[0124] "Generating" means creating new information based on input data.

[0125] "Filtering" is the process of selecting the most appropriate answer or piece of information from multiple generated responses.

[0126] "Displaying on the screen" means making information visible as text or images on the device's display.

[0127] "Notifying by voice" means conveying text data to the user as voice using speech synthesis.

[0128] Modes for carrying out the invention

[0129] System Overview

[0130] This invention relates to a system that analyzes telephone and in-person consultations in real time and provides appropriate answers and related information. This system consists of a user's terminal, a cloud-based server, and communication means. The following describes each component of the system in detail.

[0131] Voice input and text conversion

[0132] The device records audio when the user has a consultation in person or over the phone. The device uses an information processing device such as a smartphone, PC, or smart glasses. The application on this device uses the microphone function to acquire the user's voice in real time and temporarily stores the audio data in memory.

[0133] Next, the device uses a speech recognition engine to convert the speech into text. For example, it uses the Google® Cloud Speech-to-Text API to convert the acquired speech data into text data. This converted text data is then stored within the application.

[0134] Text data transmission and analysis

[0135] The terminal sends the transcribed consultation content to the server. The HTTPS protocol is used for this transmission, ensuring the secure transfer of data.

[0136] The server analyzes the received text data. A natural language processing (NLP) engine is used for the analysis. For example, the Google Cloud Natural Language API is used to extract important keywords and context from the text data.

[0137] Solution generation and search

[0138] The server uses a generative AI model to generate answers and related information. For example, a generative AI model such as OpenAI's GPT-3® is used. The server also refers to a FAQ database related to the problem to generate answers and related information.

[0139] The server filters the multiple generated answers and related information, selecting the most appropriate one. This ensures that the necessary information is provided efficiently.

[0140] Submit and display of answers

[0141] The selected answers and related information are sent from the server to the terminal. The HTTPS protocol is used for this transmission, ensuring the secure exchange of information.

[0142] The device notifies the user of the received answers and related information. The application displays this information on the screen and, if necessary, informs the user via voice.

[0143] Specific example

[0144] Customer support scenarios

[0145] 1. The user asks, "The product isn't working, what should I do?"

[0146] 2. The device records this audio and transcribes it into text.

[0147] 3. The device sends the converted data to the server.

[0148] 4. The server analyzes the text data and extracts the keywords "product" and "not working".

[0149] 5. The server generates answers using the relevant FAQ database and generative AI models.

[0150] 6. The server filters the generated answers and selects the most appropriate information.

[0151] 7. The server sends the selected answer to the terminal.

[0152] 8. The device will announce a voice message saying, "Here's how to reset," and display the reset instructions on the screen.

[0153] Sales scenario

[0154] 1. The user asks, "Please tell me more about the new product."

[0155] 2. The device records this audio and transcribes it into text.

[0156] 3. The device sends the converted data to the server.

[0157] 4. The server analyzes the text data and extracts the keywords "new product" and "details".

[0158] 5. The server generates related sales materials, product brochures, and links to explanatory videos.

[0159] 6. The server filters the generated information and selects the most appropriate information.

[0160] 7. The server sends the selected information to the terminal.

[0161] 8. The device will announce, "Here are the details of the new product," and display a link to the brochure on the screen.

[0162] The flow of the specific processing in Example 1 will be explained using Figure 11.

[0163] Step 1:

[0164] The user initiates a consultation via voice. The user launches a dedicated application on their device, such as a smartphone or PC.

[0165] Specific action: The user asks a question such as, "Please tell me more about the new product."

[0166] Input: User's voice.

[0167] Output: Audio data is input to the terminal.

[0168] Step 2:

[0169] The device records the user's voice. The device's application uses the microphone function to capture the user's speech in real time and stores the audio data in memory.

[0170] Specific operation: The application acquires audio data from the device's microphone and temporarily stores it in a buffer.

[0171] Input: User's voice data.

[0172] Output: Audio data stored in the buffer.

[0173] Step 3:

[0174] The device uses a speech recognition engine to convert speech into text. For example, it sends speech data to the Google Cloud Speech-to-Text API and converts it into text data.

[0175] Specific operation: The speech recognition engine analyzes the audio data and generates text data.

[0176] Input: Audio data.

[0177] Output: Text data.

[0178] Step 4:

[0179] The device sends the transcribed data to the server. The device uses the HTTPS protocol to securely send the text data to the server in the cloud.

[0180] Specific operation: Encoded text data is sent via the HTTPS protocol.

[0181] Input: Text data.

[0182] Output: Text data sent to the server.

[0183] Step 5:

[0184] The server analyzes the received text data. An NLP engine is used to extract important keywords and context from the text. This process includes, for example, the Google Cloud Natural Language API.

[0185] Specific operation: The server's NLP engine analyzes the text data and extracts keywords such as "new product" and "details."

[0186] Input: Text data.

[0187] Output: Analyzed keywords and contextual information.

[0188] Step 6:

[0189] The server generates answers and related information using a generative AI model. It uses a generative AI model (for example, OpenAI's GPT-3) to generate appropriate answers and information based on the consultation content.

[0190] Specific operation: Input prompts into the generative AI model and obtain the generated answer.

[0191] Input: Analyzed keywords and contextual information.

[0192] Output: Generated answers and related information.

[0193] Step 7:

[0194] The server filters the generated answers and related information. It selects the most appropriate answer from the multiple answers generated. Criteria for evaluating reliability and relevance are applied to this process.

[0195] Specific operation: The filtering algorithm selects the optimal solution.

[0196] Input: Multiple generated answers.

[0197] Output: Optimal answers and related information.

[0198] Step 8:

[0199] The server sends the answer and related information to the terminal. The selected answer is re-encoded and sent to the terminal using the HTTPS protocol.

[0200] Specific operation: Encoded information is sent via the HTTPS protocol.

[0201] Input: The best answer or related information.

[0202] Output: Answers and related information sent to the terminal.

[0203] Step 9:

[0204] The device notifies the user of the answer and related information. The device's application displays the received information on the screen and, if necessary, communicates it to the user verbally.

[0205] Specific action: The application displays information on the screen and announces with a voice message, "Here are the details of the new product."

[0206] Input: Answers and related information sent to the device.

[0207] Output: Information displayed or notified to the user.

[0208] (Application Example 1)

[0209] Next, we will explain Application Example 1. In the following explanation, the data processing device 12 will be referred to as the "server," and the smart device 14 will be referred to as the "terminal."

[0210] Conventional customer support systems have struggled to provide timely and appropriate answers to user inquiries. This often resulted in delays in response times and a decline in the quality of answers. Furthermore, limitations in communication methods and devices posed a challenge in terms of usability. This invention aims to solve these problems and enable rapid and high-quality customer support.

[0211] The specific processing performed by the specific processing unit 290 of the data processing device 12 in Application Example 1 is realized by the following means.

[0212] In this invention, the server includes: [means for the terminal to record the user's voice using communication means and transcribe the voice in real time; [means for the terminal to transmit the transcribed data to the server]; [means for the server to analyze the transcribed data, extract important keywords and context, and generate answers and related information based thereon]; [means for the server to transmit the generated answers and related information to the terminal]; [means for the terminal to display the answers and related information on a screen and notify the user of important information by voice]; and [a portable information processing device on which an application that accepts the user's voice input is executed]. This enables rapid and high-quality customer support by transcribing voice from the terminal in real time and generating and displaying appropriate answers.

[0213] A "terminal" is a computer device used by a user that records audio and transmits the transcribed data to a server.

[0214] "Communication methods" refer to the technologies and protocols used for sending and receiving data, and are the means used to exchange information between a server and a terminal.

[0215] "Recording audio" refers to collecting the user's spoken content as digital data using a microphone.

[0216] "Real-time transcription" refers to the process of instantly converting acquired audio data into text data.

[0217] "Text-based data" refers to information that has been converted into text format by a speech recognition engine.

[0218] A "server" is a computer system located in the cloud that has the function of analyzing digitized data and generating answers.

[0219] "Analyzing" refers to the process of understanding received data using a natural language processing engine and extracting important keywords and context.

[0220] "Extracting important keywords and context" refers to the process of taking meaningful words and vocabulary from text data and combining them based on their context.

[0221] "Generating related information" refers to constructing answers or additional information based on extracted keywords and context.

[0222] "Displaying on the screen" refers to providing answers or information in a visually apparent form on the device's display.

[0223] "Notifying the user via voice" refers to the process of communicating the generated answer to the user using speech synthesis technology.

[0224] A "portable information processing device" refers to an electronic device that is portable and has the functionality to accept voice input.

[0225] This invention relates to a system that transcribes a user's voice into text in real time and provides answers and related information based on that transcription. This system mainly consists of a terminal, communication means, and a server. Specific embodiments of this system are described below.

[0226] Hardware and software

[0227] The terminals used are portable information processing devices such as smartphones and smart glasses. These devices are equipped with microphones and speech recognition engines to record and transcribe user voice input in real time. Furthermore, applications running on the terminals also have the functionality to send the text data to a server.

[0228] The server resides in the cloud and uses a natural language processing (NLP) engine and generative AI models to generate answers and related information. The server analyzes the text data sent by the user and extracts important keywords and context.

[0229] The communication method used is an internet connection for sending and receiving data between the terminal and the server. The communication protocol used is HTTPS to ensure data security and efficient transfer.

[0230] Data processing and data calculation

[0231] The device uses a speech recognition engine (for example, Google Speech Recognition API) to convert speech into text data in real time. This text data is then sent to a server in the cloud via the application.

[0232] The server uses a natural language processing engine (such as SpaCy or NLTK) to analyze the text data and extract important keywords and context. Then, it uses a generative AI model (such as OpenAI's GPT-3) to generate the optimal answer and relevant information.

[0233] The generated answers are sent from the server to the terminal, where the answers and related information are displayed on the screen. Important information is also communicated to the user via voice using speech synthesis technology.

[0234] Specific example

[0235] For example, consider a customer support scenario where a user asks, "My product isn't working, what should I do?" The device records the user's voice and converts it into text data using a speech recognition engine. This text data is then sent to a server, where a natural language processing engine extracts the keywords "product" and "not working." A generative AI model then consults a relevant FAQ database to generate an answer that includes instructions on how to reset the product and links to support videos. Finally, the generated answer is sent back to the device, displayed on the screen, and announced via voice.

[0236] Example of a prompt

[0237] Prompt: "User Inquiry: [User Question]"

[0238] Example: "User inquiry: What should I do if the product doesn't work?"

[0239] Thus, the system of the present invention allows users to receive prompt and high-quality customer support. Furthermore, support staff can provide consistent and high-quality answers to user inquiries.

[0240] The flow of a specific process in Application Example 1 will be explained using Figure 12.

[0241] Step 1:

[0242] The user asks a question into the device (smartphone or smart glasses). The device's microphone records the audio. The input is the user's voice, and the output is the recorded audio data. Specifically, the microphone captures the user's speech and temporarily stores it as a buffer within the application.

[0243] Step 2:

[0244] The device sends recorded audio data to a speech recognition engine in real time, where it is converted into text data. The input is recorded audio data, and the output is transcribed text data. Specifically, the SpeechRecognition library is used to convert the audio data into text, which is then stored in a variable within the application.

[0245] Step 3:

[0246] The device sends characterized text data to a server in the cloud using the HTTPS protocol. The input is character data, and the output is the state as it has been sent to the server. Specifically, the Requests library is used to send the character data as a POST request to the server's API endpoint.

[0247] Step 4:

[0248] The server receives text data and analyzes it using a natural language processing engine (e.g., SpaCy). The input is text data, and the output is data with key keywords and context extracted. Specifically, the NLP engine on the cloud server analyzes the text data and identifies important words and phrases within the text.

[0249] Step 5:

[0250] The server uses a generative AI model (e.g., GPT-3) to generate answers and related information based on the analysis results. The input is keywords and contextual information from the analysis results, and the output is the generated answers and related information. Specifically, it generates a prompt sentence, inputs it into the generative AI model, and generates an appropriate answer.

[0251] Step 6:

[0252] The server sends the generated answer and related information back to the terminal using the HTTPS protocol. The input is the generated answer and related information, and the output is the state as it has been sent to the terminal. Specifically, the server encodes the appropriate answer and sends it in a format that the terminal can understand.

[0253] Step 7:

[0254] The device displays answers and related information on its screen and notifies the user of important information via voice. Input consists of answers and related information returned from the server, while output is visual and auditory notification to the user. Specifically, the application displays the answers on the device's display and uses speech synthesis technology to convey important points verbally.

[0255] By following these steps, the system will be able to provide real-time answers to user inquiries.

[0256] Furthermore, an emotion engine that estimates the user's emotions may be incorporated. That is, the identification processing unit 290 may use the emotion identification model 59 to estimate the user's emotions and perform identification processing using the user's emotions.

[0257] System Overview

[0258] This invention relates to a system that analyzes telephone and in-person consultations in real time and provides appropriate answers and related information. Furthermore, it is a system that achieves higher user satisfaction by analyzing the user's emotions and adjusting the content of the answers and information provided based on those emotions. This system consists of the user's terminal, a server on the cloud, an emotion engine, and communication means. The following describes each part of the system in detail.

[0259] Voice input and text conversion

[0260] 1. Users receive consultations via telephone or in person.

[0261] Users will use their own devices (smartphones, PCs, smart glasses, etc.) to provide consultation.

[0262] 2. The device records the user's consultation content as audio.

[0263] The device's application uses the microphone function to capture the audio of the consultation in real time. The audio data is temporarily stored in a buffer.

[0264] 3. The device uses a speech recognition engine to convert speech into text.

[0265] The device sends voice data to a speech recognition engine, which converts it into text data. The converted text is then stored within the application.

[0266] Emotion analysis

[0267] 4. The device sends the recorded audio data to the emotion engine, which analyzes the user's emotions.

[0268] The emotion engine analyzes the user's emotions (e.g., anger, anxiety, joy) from voice data and generates emotion data.

[0269] Text data transmission and analysis

[0270] 5. The device sends the transcribed consultation content and emotional data to the server.

[0271] Text data and sentiment data are sent from the device to a server in the cloud using the HTTPS protocol.

[0272] 6. The server receives the transcribed data and parses it.

[0273] The server's natural language processing (NLP) engine analyzes the text data and extracts important keywords and context.

[0274] 7. The server uses generative AI to generate answers and related information.

[0275] The server uses a FAQ database related to the question and a pre-trained generative AI model to generate answers and related information.

[0276] Adjustment and transmission of answers

[0277] 8. The server uses sentiment data to adjust and filter answers and related information.

[0278] The server filters and selects the most appropriate answers and related information considering the user's sentiment data received from the sentiment engine.

[0279] 9. The server sends answers and related information to the terminal.

[0280] Encode the optimal answers, links, and materials and return them to the terminal.

[0281] Display and notification of answers

[0282] 10. The terminal notifies the user of the answers and related information.

[0283] The application on the terminal displays the received answers and information on the screen and notifies the user audibly if necessary. Adjust the display content and tone of the audible notification based on the sentiment data.

[0284] Specific example

[0285] In the case of customer support

[0286] 1. The user receives a consultation that the product does not work.

[0287] 2. The terminal records the voice and converts the voice "The product doesn't work. What should I do?" into text.

[0288] 3. The device sends voice data to the emotion engine, which analyzes whether the user is feeling anxious.

[0289] 4. The device sends the transcribed data and sentiment data to the server.

[0290] 5. The server analyzes the text data and extracts the keywords "product" and "not working".

[0291] 6. The server refers to the relevant FAQ database and generates answers, reset instructions, and support video links.

[0292] 7. The server considers emotional data and selects a more polite and reassuring answer.

[0293] 8. The server sends the selected information to the terminal.

[0294] 9. The device displays instructions on how to reset the device and a link to a support video on the screen, and provides a voice notification in a gentle tone saying, "Here's how to reset the device."

[0295] In the case of sales

[0296] 1. The user receives an inquiry about the details of the new product.

[0297] 2. The device records the audio and transcribes the phrase "Please tell me more about the new product" into text.

[0298] 3. The device sends voice data to the emotion engine, which analyzes whether the user is interested.

[0299] 4. The device sends the transcribed data and sentiment data to the server.

[0300] 5. The server analyzes the text data and extracts the keywords "new product" and "details".

[0301] 6. The server generates sales materials related to it, product brochures, and links to explanatory videos.

[0302] 7. The server selects an interesting explanation considering the emotion data.

[0303] 8. The server sends the selected information to the terminal.

[0304] 9. The terminal displays the link to the brochure and the features of the product on the screen, and gives an audio notification in a lively tone saying "This is the detailed information of the new product".

[0305] In this way, the system of the present invention enables consistent high-quality answers and information provision based on the emotions of users, has high practicality in the field, and contributes to the improvement of business efficiency and user satisfaction.

[0306] The following describes the processing flow.

[0307] Step 1:

[0308] The user receives consultations by phone or in person.

[0309] The user uses their own terminal (such as a smartphone, PC, smart glasses, etc.) to handle the consultations.

[0310] Step 2:

[0311] The terminal records the voice of the user's consultation content.

[0312] The application on the terminal utilizes the microphone function to obtain the voice of the consultation content in real time. The voice data is temporarily stored in the buffer.

[0313] Step 3:

[0314] The terminal sends the voice data to the emotion engine to analyze the user's emotion.

[0315] The device sends voice data to an emotion engine (for example, a voice analysis API), which analyzes the user's emotions (anger, anxiety, joy, etc.) from the voice. The analysis results are stored as emotion data.

[0316] Step 4:

[0317] The device uses a speech recognition engine to convert speech into text.

[0318] The device sends the voice data to a speech recognition engine (e.g., a speech recognition API) and converts it into text data. The converted text is then stored within the application.

[0319] Step 5:

[0320] The device sends the transcribed consultation content and emotional data to the server.

[0321] Text data and sentiment data are sent from the device to a server in the cloud using the HTTPS protocol.

[0322] Step 6:

[0323] The server receives the transcribed data and parses it.

[0324] The server's natural language processing (NLP) engine analyzes the text data and extracts important keywords and context. The analysis results are saved as analysis data.

[0325] Step 7:

[0326] The server uses generative AI to generate answers and related information.

[0327] Based on the analysis data, the server generates answers and related information using a database of FAQs relevant to the problem and a pre-trained generative AI model. The generated information is temporarily stored.

[0328] Step 8:

[0329] The server uses sentiment data to adjust and filter answers and related information.

[0330] The server considers the user's emotional data received from the emotion engine and filters and selects the most appropriate answers and relevant information. The selected information is then prepared as data to be sent.

[0331] Step 9:

[0332] The server sends the answer and related information to the terminal.

[0333] The transmitted data is encoded and sent back to the device. This return transmission takes place via cloud communication.

[0334] Step 10:

[0335] The device notifies the user of the answer and related information.

[0336] The application on the device displays the received answers and information on the screen and notifies the user by voice as needed. Based on sentiment data, the displayed content and the tone of voice notifications are adjusted.

[0337] (Example 2)

[0338] Next, we will describe Example 2. In the following description, the data processing device 12 will be referred to as the "server" and the smart device 14 as the "terminal".

[0339] Conventional consultation systems only analyze data transcribed from user speech, failing to consider user emotions and making it difficult to provide appropriate and personalized answers. Furthermore, they cannot provide notifications in a tone adjusted to the user's emotions, thus failing to adequately improve user satisfaction. Therefore, there was a need for a more advanced consultation system that analyzes user emotions in real time and provides emotion-based answers.

[0340] The identification process performed by the identification processing unit 290 of the data processing device 12 in Example 2 is realized by the following means.

[0341] In this invention, the server includes means for the terminal to transmit text data and sentiment data to the server, means for the server to analyze the text data, extract important keywords and context, and adjust the answer and related information using the sentiment data, and means for the server to transmit the generated answer and related information to the terminal. This enables the provision of appropriate answers based on the user's emotions and notifications in an adjusted tone.

[0342] A "terminal" is an information processing device used by users to input voice and send and receive data.

[0343] "Communication methods" refer to network technologies used to send and receive data between terminals and servers.

[0344] "Recording audio" means acquiring and saving the user's speech as digital audio data.

[0345] "Real-time transcription" refers to the process of instantly converting recorded audio data into text data.

[0346] "Emotional data" refers to information that indicates the user's emotional state, analyzed from voice data.

[0347] A "server" is a remote computer system that analyzes data received from a terminal and generates appropriate answers or information.

[0348] "Analyzing text data" means processing the text data received by the server and extracting important keywords and context.

[0349] "Key keywords" are words or phrases that are particularly important for understanding the user's inquiry.

[0350] "Extracting context" means analyzing the relationships between words in order to understand the meaning and intent of the entire text.

[0351] "Generating answers and related information" means that the server automatically creates appropriate answers and reference information based on the user's inquiry.

[0352] A "generative AI model" is an artificial intelligence system that uses technologies such as natural language processing to automatically generate text from input data.

[0353] "Adjusted tone" refers to appropriately modifying the speech style and nuances of voice notifications based on the user's emotional data.

[0354] "Displaying on the screen" means visually presenting answers or information on the device's display.

[0355] "Notifying by voice" means presenting answers or information to the user audibly using speech synthesis technology.

[0356] This invention is a system that analyzes user inquiries in real time and provides appropriate answers and related information. This system also analyzes the user's emotions and adjusts the content of the answers and information provided based on the emotional data, thereby achieving higher user satisfaction.

[0357] The main components of the system consist of the user's terminal, a server in the cloud, an emotion engine, and communication methods. Details of each component are shown below.

[0358] terminal

[0359] This is an information processing device that allows users to input voice data and send and receive data. Specifically, this includes smartphones, PCs, and smart glasses. The device incorporates a voice recording function, a speech recognition engine to convert voice data into text, and communication means for sending and receiving the results of sentiment analysis.

[0360] server

[0361] This is a cloud-based computer system that analyzes text and sentiment data received from terminals. The server uses a natural language processing (NLP) engine to analyze text data and generates answers and related information using generative AI models (e.g., OpenAI GPT-3). It also plays a role in adjusting answers by taking sentiment data into consideration.

[0362] means of communication

[0363] This is a network technology for sending and receiving data between a terminal and a server. Specifically, it involves data transmission using the HTTPS protocol.

[0364] Examples

[0365] For customer support

[0366] 1. The user asks, "The product isn't working, what should I do?"

[0367] 2. The device records the audio and uses a speech recognition engine (e.g., Google Speech-to-Text API) to transcribe it. The generated text data will be "The product isn't working, what should I do?"

[0368] 3. The device sends voice data to an emotion engine (e.g., IBM Watson® Tone Analyzer) to obtain emotion data indicating anxiety.

[0369] 4. The device sends text data and sentiment data to the server.

[0370] 5. The server uses a natural language processing engine (e.g., BERT model) to analyze the text data and extract key keywords such as "product" and "not working".

[0371] 6. The server uses a generative AI model (e.g., OpenAI GPT-3) to generate "Product Reset Instructions" and "Support Video Links".

[0372] 7. Take into account the server's emotional data indicating anxiety, and adjust the response to be polite and reassuring.

[0373] 8. The server sends the adjusted answer to the terminal.

[0374] 9. The device displays "Here's how to reset it" on the screen and also provides a voice notification in a gentle tone.

[0375] Examples of prompt statements

[0376] For customer support

[0377] Prompt message:

[0378] User: "The product isn't working, what should I do?" Sentiment analysis result: Anxiety.

[0379] Generative AI models:

[0380] 1. How to reset the product

[0381] 2. Support video links

[0382] 3. Related answers from the FAQ

[0383] Taking these factors into consideration, please generate your answer in a gentle tone.

[0384] The above details the embodiment of the system of the present invention. This system provides high-quality answers and information based on the user's emotions, thereby improving work efficiency and user satisfaction.

[0385] The flow of the specific processing in Example 2 will be explained using Figure 13.

[0386] Step 1:

[0387] The user initiates a consultation.

[0388] Input: User's voice (e.g., "The product isn't working, what should I do?").

[0389] Specific action: The user speaks their question into the device's microphone.

[0390] Output: Raw audio data.

[0391] Step 2:

[0392] The device records audio.

[0393] Input: Raw audio data.

[0394] Specific operation: The device uses its microphone function to record the user's voice in real time and temporarily stores it in a buffer as audio data.

[0395] Output: Audio data stored in the buffer.

[0396] Step 3:

[0397] The device converts speech into text.

[0398] Input: Audio data stored in the buffer.

[0399] Specific operation: The device sends audio data to a speech recognition engine (e.g., Google Speech-to-Text API), and the audio data is converted into text data.

[0400] Output: Text data (Example: "The product isn't working, what should I do?").

[0401] Step 4:

[0402] The device performs emotion analysis.

[0403] Input: Audio data and text data.

[0404] Specific operation: The device sends voice data to an emotion engine (e.g., IBM Watson Tone Analyzer), which analyzes the emotion and obtains emotion data.

[0405] Output: Emotional data (e.g., anxiety).

[0406] Step 5:

[0407] The device sends data to the server.

[0408] Input: Text data and sentiment data.

[0409] Specific operation: The device sends text data and sentiment data to a server in the cloud using the HTTPS protocol.

[0410] Output: Text data and sentiment data sent to the server.

[0411] Step 6:

[0412] The server analyzes the text data.

[0413] Input: Character data sent to a server in the cloud.

[0414] Specific operation: The server uses a natural language processing (NLP) engine (e.g., BERT model) to analyze the text data and extract important keywords and context.

[0415] Output: Extracted keywords (e.g., "product", "not working").

[0416] Step 7:

[0417] The server generates the answers and related information.

[0418] Input: Extracted keywords.

[0419] Specific operation: The server uses a generative AI model (e.g., OpenAI GPT-3) to generate answers and related information.

[0420] Output: Generated answers and related information (e.g., "Product Reset Instructions," "Support Video Link").

[0421] Step 8:

[0422] The server adjusts the answers based on sentiment data.

[0423] Input: Sentimental data and generated answers or related information.

[0424] Specific operation: The server considers the sentiment data received from the sentiment engine and adjusts the response to have the most appropriate tone and content.

[0425] Output: Adjusted answer (e.g., a polite and reassuring answer).

[0426] Step 9:

[0427] The server sends the answer to the terminal.

[0428] Input: Adjusted answers and related information.

[0429] Specific operation: The server encodes the adjusted answer and related information and sends it to the terminal.

[0430] Output: The adjusted answers and related information sent to the terminal.

[0431] Step 10:

[0432] The device will display and notify you of the answer.

[0433] Input: Adjusted answers and related information sent from the server.

[0434] Specific actions: The device displays answers and information on the screen and provides voice notifications in a tone adjusted based on emotional data.

[0435] Output: The answers and information notified to the user.

[0436] (Application Example 2)

[0437] Next, we will explain application example 2. In the following explanation, the data processing device 12 will be referred to as a "server" and the smart device 14 as a "terminal".

[0438] Conventional consultation systems using speech recognition have difficulty providing answers and information that take into account the user's emotions, and in particular, providing appropriate response instructions in real time has been difficult in security service settings. As a result, the quality of responses by field staff has declined, hindering security operations that require rapid response. The present invention aims to solve these problems and provide a system that improves the quality of responses by field staff by analyzing the user's emotions and providing appropriate response instructions in real time.

[0439] The specific processing performed by the specific processing unit 290 of the data processing device 12 in Application Example 2 is realized by the following means.

[0440] In this invention, the server includes means for analyzing transcribed data, extracting important keywords and context, and generating answers and related information based on them; means for analyzing the user's emotions and adjusting the answers and related information based on the analysis results; and means for transmitting the generated answers and related information to the terminal. This makes it possible to provide appropriate response instructions and information based on the user's emotions.

[0441] A "terminal" is an electronic device operated by a user that performs functions such as recording voice, transcribing it into text, displaying sentiment analysis results, and sending notifications.

[0442] "Communication method" refers to technology that uses Internet protocols (e.g., HTTPS) to send and receive data between a terminal and a server.

[0443] A "speech recognition engine" is software or a system that converts speech data into text data.

[0444] An "emotion analysis engine" is software or a system that analyzes a user's emotions from voice data or text data and outputs the results as data.

[0445] A "server" is a computer system that resides in the cloud and performs tasks such as analyzing text data, utilizing sentiment analysis results, and generating and transmitting answers and related information.

[0446] A "natural language processing engine" is software or a system that analyzes text data, understands keywords and context, and generates appropriate answers and related information.

[0447] "Generative AI" is artificial intelligence that uses a pre-trained model to generate answers and information in natural language based on input data.

[0448] "Answers and related information" refers to the answers, related data, links, and materials provided in response to a user's query.

[0449] "Emotional data" refers to information generated by an emotion analysis engine that indicates the user's emotional state.

[0450] The system of the present invention analyzes voice data in real time and provides appropriate answers and relevant information based on the user's emotions. This system is intended to support on-site response in security services, and its main components include a terminal, a server, an emotion analysis engine, a speech recognition engine, a natural language processing engine, and a generative AI model.

[0451] Overall system configuration

[0452] 1. Terminal:

[0453] It is a user-operated electronic device that records voice, transcribes it into text, displays sentiment analysis results, and provides notifications.

[0454] Smartphones, tablets, and PCs are examples.

[0455] 2. Speech recognition engine:

[0456] Software that converts audio data into text data. This often utilizes APIs such as Google's Speech-to-Text API.

[0457] 3. Emotion Analysis Engine:

[0458] Software that analyzes a user's emotions from audio or text data and outputs the results as data. IBM Watson Tone Analyzer is used.

[0459] 4. Server:

[0460] It is a computer system that resides in the cloud and performs text data analysis, utilizes sentiment analysis results, and generates and transmits answers and related information.

[0461] The cloud server runs the natural language processing engines (NLTK, SpaCy) and the generative AI model (OpenAI GPT-4®).

[0462] 5. Means of communication:

[0463] The Internet Protocol (HTTPS) is used to send and receive data between the terminal and the server.

[0464] System Operation Overview

[0465] When a user makes a voice report at a security site, the device records the voice and converts it into text data in real time using a speech recognition engine. This text data is sent to an emotion analysis engine to identify the user's emotions. The emotion data and text data are sent to a cloud server, where a natural language processing engine analyzes the text data and extracts important keywords and context. Subsequently, a generative AI model generates appropriate answers and relevant information, and the answers are refined based on the emotion data. The optimized answers are sent to the device and displayed and notified to the user.

[0466] Specific example

[0467] For example, suppose a security officer reports, "An alarm has been triggered at the north entrance of the building. Please check the situation." This audio data is recorded on a terminal and converted into text data, "An alarm has been triggered at the north entrance of the building," using Google's Speech-to-Text API. This text data is then analyzed by IBM Watson Tone Analyzer, generating the sentiment data "urgent." The text data and sentiment data sent to the cloud server are analyzed by a natural language processing engine, and a generative AI model generates the optimal response based on the following prompt sentences.

[0468] Example of a prompt:

[0469] "Alarm: An alarm has been triggered at the north entrance of the building."

[0470] Emotion: Urgent

[0471] Possible steps: 1. Check the north entrance with the camera.

[0472] 2. Contact the nearest guard.

[0473] 3. Call the police if necessary.

[0474] Please generate the optimal solution.

[0475] Based on this, the generated answers and related information are sent to the device and notified to the user. The user can then quickly take appropriate action by following the instructions displayed on the device.

[0476] The flow of a specific process in Application Example 2 will be explained using Figure 14.

[0477] Step 1:

[0478] The device records audio reported by the user at the security site. This audio data is temporarily stored in a buffer on the device. The input is the user's voice, and the output is the recorded audio data.

[0479] Step 2:

[0480] The device sends recorded audio data in real time to a speech recognition engine (for example, Google's Speech-to-Text API). This engine converts the audio data into text data. Specifically, the audio data is sent to the API, and text data is returned. The input is the recorded audio data, and the output is the converted text data.

[0481] Step 3:

[0482] The terminal sends transcribed text data to an emotion analysis engine (e.g., IBM Watson Tone Analyzer) to analyze the user's emotions. The emotion analysis engine identifies emotions (such as urgency, anxiety, or joy) from the text data and generates emotion data. The input is transcribed text data, and the output is emotion data.

[0483] Step 4:

[0484] The device sends transcribed text data and sentiment data to a server in the cloud using the HTTPS protocol. Specifically, data packaged in JSON format or similar is sent. The input is text data and sentiment data, and the output is the data sent to the server.

[0485] Step 5:

[0486] The server analyzes the received text and sentiment data. It uses its natural language processing engine (e.g., NLTK or SpaCy) to extract keywords and context from the text data. The input is the text and sentiment data sent to the server, and the output is the extracted keyword and context data.

[0487] Step 6:

[0488] The server uses a generative AI model (e.g., OpenAI GPT-4) to generate answers and related information based on extracted keywords and contextual data. An example of a generated prompt is as follows:

[0489] "Alarm: An alarm has been triggered at the north entrance of the building. Emotion: Emergency Possible actions: 1. Check the north entrance with the camera 2. Contact the nearest guard 3. Call the police if necessary Please generate the best solution."

[0490] The input consists of keywords and contextual data, while the output consists of generated answers and related information.

[0491] Step 7:

[0492] The server adjusts the generated answers and related information based on sentiment data and selects the most appropriate one. This adjustment takes into account the user's current emotional state (e.g., urgent). The input is sentiment data and the generated answers and related information, and the output is the adjusted answers and related information.

[0493] Step 8:

[0494] The server sends the adjusted answers and related information to the terminal. The input is the adjusted answers and related information, and the output is the data sent to the terminal.

[0495] Step 9:

[0496] The device displays the adjusted answers and related information received from the server on its screen and, if necessary, informs the user via voice notification. Specifically, the answers are displayed on the device's screen, and necessary notifications are played through the speaker. The input is the answers and related information sent to the device, and the output is the answers and related information notified to the user.

[0497] The specific processing unit 290 transmits the result of the specific processing to the smart device 14. In the smart device 14, the control unit 46A causes the output device 40 to output the result of the specific processing. The microphone 38B acquires audio indicating user input for the result of the specific processing. The control unit 46A transmits the audio data indicating user input acquired by the microphone 38B to the data processing device 12. In the data processing device 12, the specific processing unit 290 acquires the audio data.

[0498] Data generation model 58 is a so-called generative AI (Artificial Intelligence). An example of data generation model 58 is ChatGPT (registered trademark) (Internet search).<URL: https: / / openai.com / blog / chatgpt> ), Gemini (registered trademark) (Internet search) <url: https: gemini.google.com ?hl="ja">Examples of generative AI include the following. The data generation model 58 is obtained by performing deep learning on a neural network. The data generation model 58 is input with prompts containing instructions, and with inference data such as audio data representing speech, text data representing text, and image data representing images. The data generation model 58 infers from the input inference data according to the instructions indicated by the prompts, and outputs the inference results in data formats such as audio data and text data. Here, inference refers to, for example, analysis, classification, prediction, and / or summarization.

[0499] In the above embodiment, an example was given in which specific processing is performed by the data processing device 12, but the technology of this disclosure is not limited thereto, and the specific processing may also be performed by the smart device 14.

[0500] [Second Embodiment]

[0501] Figure 3 shows an example of the configuration of the data processing system 210 according to the second embodiment.

[0502] As shown in Figure 3, the data processing system 210 includes a data processing device 12 and smart glasses 214. An example of the data processing device 12 is a server.

[0503] The data processing device 12 comprises a computer 22, a database 24, and a communication interface 26. The computer 22 is an example of a "computer" related to the technology of this disclosure. The computer 22 comprises a processor 28, RAM 30, and storage 32. The processor 28, RAM 30, and storage 32 are connected to a bus 34. The database 24 and the communication interface 26 are also connected to the bus 34. The communication interface 26 is connected to a network 54. An example of the network 54 is a WAN (Wide Area Network) and / or a LAN (Local Area Network).

[0504] The smart glasses 214 include a computer 36, a microphone 238, a speaker 240, a camera 42, and a communication interface 44. The computer 36 includes a processor 46, RAM 48, and storage 50. The processor 46, RAM 48, and storage 50 are connected to a bus 52. The microphone 238, speaker 240, and camera 42 are also connected to the bus 52.

[0505] The microphone 238 receives voice signals from the user 20 and receives instructions from the user 20. The microphone 238 captures the voice signals from the user 20, converts the captured voice into audio data, and outputs it to the processor 46. The speaker 240 outputs audio according to the instructions from the processor 46.

[0506] Camera 42 is a small digital camera equipped with an optical system including a lens, aperture, and shutter, and an image sensor such as a CMOS (Complementary Metal-Oxide-Semiconductor) image sensor or a CCD (Charge Coupled Device) image sensor, and captures images of the area around the user 20 (for example, an imaging range defined by a field of view equivalent to the width of a typical healthy person's field of vision).

[0507] Communication interface 44 is connected to network 54. Communication interfaces 44 and 26 are responsible for the exchange of various information between processor 46 and processor 28 via network 54. The exchange of various information between processor 46 and processor 28 using communication interfaces 44 and 26 is performed in a secure manner.

[0508] Figure 4 shows an example of the main functions of the data processing device 12 and the smart glasses 214. As shown in Figure 4, the data processing device 12 performs specific processing using the processor 28. The storage 32 stores the specific processing program 56.

[0509] The specific processing program 56 is an example of a "program" relating to the technology of this disclosure. The processor 28 reads the specific processing program 56 from the storage 32 and executes the read specific processing program 56 on the RAM 30. The specific processing is realized by the processor 28 operating as a specific processing unit 290 in accordance with the specific processing program 56 executed on the RAM 30.

[0510] The storage 32 stores the data generation model 58 and the emotion identification model 59. The data generation model 58 and the emotion identification model 59 are used by the identification processing unit 290.

[0511] In the smart glasses 214, the processor 46 performs the reception output processing. The storage 50 stores the reception output program 60. The processor 46 reads the reception output program 60 from the storage 50 and executes the read reception output program 60 on the RAM 48. The reception output processing is realized by the processor 46 operating as a control unit 46A according to the reception output program 60 executed on the RAM 48.

[0512] Next, the identification processing performed by the identification processing unit 290 of the data processing device 12 will be described. In the following description, the data processing device 12 will be referred to as the "server" and the smart glasses 214 will be referred to as the "terminal".

[0513] System Overview

[0514] This invention relates to a system that analyzes telephone and in-person consultations in real time and provides appropriate answers and related information. This system consists of a user's terminal, a cloud-based server, and communication means. The following describes each part of the system in detail.

[0515] Voice input and text conversion

[0516] 1. Users receive consultations via telephone or in person.

[0517] Users will use their own devices (smartphones, PCs, smart glasses, etc.) to provide consultation.

[0518] 2. The device records the user's consultation content as audio.

[0519] The device's application uses the microphone function to capture the audio of the consultation in real time. The audio data is temporarily stored in a buffer.

[0520] 3. The device uses a speech recognition engine to convert speech into text.

[0521] The device sends voice data to a speech recognition engine, which converts it into text data. The converted text is then stored within the application.

[0522] Text data transmission and analysis

[0523] 4. The device sends the transcribed consultation content to the server.

[0524] Text data is sent from the device to a server in the cloud using the HTTPS protocol.

[0525] 5. The server receives the transcribed data and parses it.

[0526] The server's natural language processing (NLP) engine analyzes the text data and extracts important keywords and context.

[0527] Solution generation and search

[0528] 6. The server uses generative AI to generate answers and related information.

[0529] The server generates answers and related information using a database of FAQs related to the problem and a pre-trained generative AI model.

[0530] 7. The server filters the generated answers and related information and selects the appropriate ones.

[0531] From the multiple generated answers, filter and select the most appropriate information.

[0532] Submit and display of answers

[0533] 8. The server sends the answer and related information to the terminal.

[0534] Encode the optimal answers, links, and materials and send them back to your device.

[0535] 9. The device notifies the user of the answer and related information.

[0536] The application on the device displays the received answers and information on the screen and notifies the user by voice as needed.

[0537] Specific example

[0538] For customer support

[0539] 1. A user contacts us because the product is not working.

[0540] 2. The device records the audio and transcribes the voice message, "The product isn't working, what should I do?"

[0541] 3. The device sends the converted data to the server.

[0542] 4. The server analyzes the text data and extracts the keywords "product" and "not working".

[0543] 5. The server refers to the relevant FAQ database and generates answers, reset instructions, and support video links.

[0544] 6. Filter the information generated by the server and send it to the terminal.

[0545] 7. The device will display instructions on how to reset the device and a link to a support video on the screen, and will announce a voice message saying, "Here's how to reset the device."

[0546] In the case of sales

[0547] 1. The user receives an inquiry about the details of the new product.

[0548] 2. The device records the audio and transcribes the phrase "Please tell me more about the new product" into text.

[0549] 3. The device sends the converted data to the server.

[0550] 4. The server analyzes the text data and extracts the keywords "new product" and "details".

[0551] 5. The server generates related sales materials, product brochures, and links to explanatory videos.

[0552] 6. Filter the information generated by the server and send it to the terminal.

[0553] 7. The device displays a link to a brochure and product features on the screen, and announces with a voice message, "Here are the details of the new product."

[0554] Thus, the system of the present invention enables consistent, high-quality answers and information provision, regardless of the user's skill level. It offers high practicality in the field and contributes to improved work efficiency.

[0555] The following describes the processing flow.

[0556] Step 1:

[0557] Users receive consultations via phone or in person.

[0558] Users will use their own devices (smartphones, PCs, smart glasses, etc.) to provide consultation.

[0559] Step 2:

[0560] The device records the user's conversation audio.

[0561] The device's application uses the microphone function to capture the audio of the consultation in real time. The audio data is temporarily stored in a buffer.

[0562] Step 3:

[0563] The device uses a speech recognition engine to convert speech into text.

[0564] The device sends voice data to a speech recognition engine, which converts it into text data. The converted text is then stored within the application.

[0565] Step 4:

[0566] The terminal sends the transcribed consultation content to the server.

[0567] Text data is sent from the device to a server in the cloud using the HTTPS protocol.

[0568] Step 5:

[0569] The server receives the transcribed data and parses it.

[0570] The server's natural language processing (NLP) engine analyzes the text data and extracts important keywords and context.

[0571] Step 6:

[0572] The server uses generative AI to generate answers and related information.

[0573] The server generates answers and related information using a database of FAQs related to the problem and a pre-trained generative AI model.

[0574] Step 7:

[0575] The server filters the generated answers and related information and selects the appropriate ones.

[0576] From the multiple generated answers, filter and select the most appropriate information.

[0577] Step 8:

[0578] The server sends the answer and related information to the terminal.

[0579] Encode the optimal answers, links, and materials and send them back to your device.

[0580] Step 9:

[0581] The device notifies the user of the answer and related information.

[0582] The application on the device displays the received answers and information on the screen and notifies the user by voice as needed.

[0583] (Example 1)

[0584] Next, we will describe Example 1. In the following description, the data processing device 12 will be referred to as the "server," and the smart glasses 214 will be referred to as the "terminal."

[0585] In modern call centers and customer support, there is a demand for providing quick and accurate answers to user inquiries. However, conventional systems often suffer from low accuracy in voice-to-text conversion, and answer generation is frequently done manually, leading to delays and inconsistent quality. This has resulted in decreased customer satisfaction and reduced operational efficiency. This invention aims to solve these problems, streamline the process from voice input to answer provision, and achieve consistent, high-quality information delivery.

[0586] The identification process performed by the identification processing unit 290 of the data processing device 12 in Example 1 is realized by the following means.

[0587] In this invention, the server includes: means for the terminal to record the user's voice using communication means and transcribe the voice in real time; means for the terminal to transmit the transcribed data to the server; means for the server to analyze the transcribed data, extract important keywords and context, and generate answers and related information based on them; means for the server to filter the generated answers and related information and transmit it to the terminal; and means for the terminal to display the answers and related information on the screen and notify the user of important information by voice. This automates the process from voice input to answer provision, enabling the rapid, consistent, and high-quality provision of information.

[0588] A "terminal" is an information processing device used by a user, which acquires and processes data, including voice input. Specifically, this refers to smartphones, personal computers, smart glasses, etc.

[0589] "Communication means" refers to technical means for sending and receiving data, including the internet, Wi-Fi, Bluetooth, etc.

[0590] "Recording audio" means capturing a user's speech as digital audio data using a microphone and storing it.

[0591] "Real-time transcription of speech" means instantly analyzing speech data and converting it into corresponding text data. This is typically done using a speech recognition engine.

[0592] "Transcribed data" refers to text data obtained after converting voice input into written information.

[0593] A "server" is a high-performance information processing device that operates on the cloud and has functions such as data analysis, storage, and transmission.

[0594] "Analysis" is the process of breaking down received data and performing operations to understand it.

[0595] "Important keywords and context" refer to words and phrases that are particularly meaningful from the input data, as well as the relationships between them.

[0596] "Answers and related information" refers to information that allows for appropriate responses to user questions and inquiries. This includes generated text, links, and documents.

[0597] "Generating" means creating new information based on input data.

[0598] "Filtering" is the process of selecting the most appropriate answer or piece of information from multiple generated responses.

[0599] "Displaying on the screen" means making information visible as text or images on the device's display.

[0600] "Notifying by voice" means conveying text data to the user as voice using speech synthesis.

[0601] Modes for carrying out the invention

[0602] System Overview

[0603] This invention relates to a system that analyzes telephone and in-person consultations in real time and provides appropriate answers and related information. This system consists of a user's terminal, a cloud-based server, and communication means. The following describes each component of the system in detail.

[0604] Voice input and text conversion

[0605] The device records audio when the user has a consultation in person or over the phone. The device uses an information processing device such as a smartphone, PC, or smart glasses. The application on this device uses the microphone function to acquire the user's voice in real time and temporarily stores the audio data in memory.

[0606] Next, the device uses a speech recognition engine to convert the speech into text. For example, it might use the Google Cloud Speech-to-Text API to convert the acquired speech data into text data. This converted text data is then stored within the application.

[0607] Text data transmission and analysis

[0608] The terminal sends the transcribed consultation content to the server. The HTTPS protocol is used for this transmission, ensuring the secure transfer of data.

[0609] The server analyzes the received text data. A natural language processing (NLP) engine is used for the analysis. For example, the Google Cloud Natural Language API is used to extract important keywords and context from the text data.

[0610] Solution generation and search

[0611] The server uses a generative AI model to generate answers and related information. For example, a generative AI model like OpenAI's GPT-3 is used. The server also refers to a FAQ database related to the problem to generate answers and related information.

[0612] The server filters the multiple generated answers and related information, selecting the most appropriate one. This ensures that the necessary information is provided efficiently.

[0613] Submit and display of answers

[0614] The selected answers and related information are sent from the server to the terminal. The HTTPS protocol is used for this transmission, ensuring the secure exchange of information.

[0615] The device notifies the user of the received answers and related information. The application displays this information on the screen and, if necessary, informs the user via voice.

[0616] Specific example

[0617] Customer support scenarios

[0618] 1. The user asks, "The product isn't working, what should I do?"

[0619] 2. The device records this audio and transcribes it into text.

[0620] 3. The device sends the converted data to the server.

[0621] 4. The server analyzes the text data and extracts the keywords "product" and "not working".

[0622] 5. The server generates answers using the relevant FAQ database and generative AI models.

[0623] 6. The server filters the generated answers and selects the most appropriate information.

[0624] 7. The server sends the selected answer to the terminal.

[0625] 8. The device will announce a voice message saying, "Here's how to reset," and display the reset instructions on the screen.

[0626] Sales scenario

[0627] 1. The user asks, "Please tell me more about the new product."

[0628] 2. The device records this audio and transcribes it into text.

[0629] 3. The device sends the converted data to the server.

[0630] 4. The server analyzes the text data and extracts the keywords "new product" and "details".

[0631] 5. The server generates related sales materials, product brochures, and links to explanatory videos.

[0632] 6. The server filters the generated information and selects the most appropriate information.

[0633] 7. The server sends the selected information to the terminal.

[0634] 8. The device will announce, "Here are the details of the new product," and display a link to the brochure on the screen.

[0635] The flow of the specific processing in Example 1 will be explained using Figure 11.

[0636] Step 1:

[0637] The user initiates a consultation via voice. The user launches a dedicated application on their device, such as a smartphone or PC.

[0638] Specific action: The user asks a question such as, "Please tell me more about the new product."

[0639] Input: User's voice.

[0640] Output: Audio data is input to the terminal.

[0641] Step 2:

[0642] The device records the user's voice. The device's application uses the microphone function to capture the user's speech in real time and stores the audio data in memory.

[0643] Specific operation: The application acquires audio data from the device's microphone and temporarily stores it in a buffer.

[0644] Input: User's voice data.

[0645] Output: Audio data stored in the buffer.

[0646] Step 3:

[0647] The device uses a speech recognition engine to convert speech into text. For example, it sends speech data to the Google Cloud Speech-to-Text API and converts it into text data.

[0648] Specific operation: The speech recognition engine analyzes the audio data and generates text data.

[0649] Input: Audio data.

[0650] Output: Text data.

[0651] Step 4:

[0652] The device sends the transcribed data to the server. The device uses the HTTPS protocol to securely send the text data to the server in the cloud.

[0653] Specific operation: Encoded text data is sent via the HTTPS protocol.

[0654] Input: Text data.

[0655] Output: Text data sent to the server.

[0656] Step 5:

[0657] The server analyzes the received text data. An NLP engine is used to extract important keywords and context from the text. This process includes, for example, the Google Cloud Natural Language API.

[0658] Specific operation: The server's NLP engine analyzes the text data and extracts keywords such as "new product" and "details."

[0659] Input: Text data.

[0660] Output: Analyzed keywords and contextual information.

[0661] Step 6:

[0662] The server generates answers and related information using a generative AI model. It uses a generative AI model (for example, OpenAI's GPT-3) to generate appropriate answers and information based on the consultation content.

[0663] Specific operation: Input prompts into the generative AI model and obtain the generated answer.

[0664] Input: Analyzed keywords and contextual information.

[0665] Output: Generated answers and related information.

[0666] Step 7:

[0667] The server filters the generated answers and related information. It selects the most appropriate answer from the multiple answers generated. Criteria for evaluating reliability and relevance are applied to this process.

[0668] Specific operation: The filtering algorithm selects the optimal solution.

[0669] Input: Multiple generated answers.

[0670] Output: Optimal answers and related information.

[0671] Step 8:

[0672] The server sends the answer and related information to the terminal. The selected answer is re-encoded and sent to the terminal using the HTTPS protocol.

[0673] Specific operation: Encoded information is sent via the HTTPS protocol.

[0674] Input: The best answer or related information.

[0675] Output: Answers and related information sent to the terminal.

[0676] Step 9:

[0677] The device notifies the user of the answer and related information. The device's application displays the received information on the screen and, if necessary, communicates it to the user verbally.

[0678] Specific action: The application displays information on the screen and announces with a voice message, "Here are the details of the new product."

[0679] Input: Answers and related information sent to the device.

[0680] Output: Information displayed or notified to the user.

[0681] (Application Example 1)

[0682] Next, we will explain Application Example 1. In the following explanation, the data processing device 12 will be referred to as the "server," and the smart glasses 214 will be referred to as the "terminal."

[0683] Conventional customer support systems have struggled to provide timely and appropriate answers to user inquiries. This often resulted in delays in response times and a decline in the quality of answers. Furthermore, limitations in communication methods and devices posed a challenge in terms of usability. This invention aims to solve these problems and enable rapid and high-quality customer support.

[0684] The specific processing performed by the specific processing unit 290 of the data processing device 12 in Application Example 1 is realized by the following means.

[0685] In this invention, the server includes: [means for the terminal to record the user's voice using communication means and transcribe the voice in real time; [means for the terminal to transmit the transcribed data to the server]; [means for the server to analyze the transcribed data, extract important keywords and context, and generate answers and related information based thereon]; [means for the server to transmit the generated answers and related information to the terminal]; [means for the terminal to display the answers and related information on a screen and notify the user of important information by voice]; and [a portable information processing device on which an application that accepts the user's voice input is executed]. This enables rapid and high-quality customer support by transcribing voice from the terminal in real time and generating and displaying appropriate answers.

[0686] A "terminal" is a computer device used by a user that records audio and transmits the transcribed data to a server.

[0687] "Communication methods" refer to the technologies and protocols used for sending and receiving data, and are the means used to exchange information between a server and a terminal.

[0688] "Recording audio" refers to collecting the user's spoken content as digital data using a microphone.

[0689] "Real-time transcription" refers to the process of instantly converting acquired audio data into text data.

[0690] "Text-based data" refers to information that has been converted into text format by a speech recognition engine.

[0691] A "server" is a computer system located in the cloud that has the function of analyzing digitized data and generating answers.

[0692] "Analyzing" refers to the process of understanding received data using a natural language processing engine and extracting important keywords and context.

[0693] "Extracting important keywords and context" refers to the process of taking meaningful words and vocabulary from text data and combining them based on their context.

[0694] "Generating related information" refers to constructing answers or additional information based on extracted keywords and context.

[0695] "Displaying on the screen" refers to providing answers or information in a visually apparent form on the device's display.

[0696] "Notifying the user via voice" refers to the process of communicating the generated answer to the user using speech synthesis technology.

[0697] A "portable information processing device" refers to an electronic device that is portable and has the functionality to accept voice input.

[0698] This invention relates to a system that transcribes a user's voice into text in real time and provides answers and related information based on that transcription. This system mainly consists of a terminal, communication means, and a server. Specific embodiments of this system are described below.

[0699] Hardware and software

[0700] The terminals used are portable information processing devices such as smartphones and smart glasses. These devices are equipped with microphones and speech recognition engines to record and transcribe user voice input in real time. Furthermore, applications running on the terminals also have the functionality to send the text data to a server.

[0701] The server resides in the cloud and uses a natural language processing (NLP) engine and generative AI models to generate answers and related information. The server analyzes the text data sent by the user and extracts important keywords and context.

[0702] The communication method used is an internet connection for sending and receiving data between the terminal and the server. The communication protocol used is HTTPS to ensure data security and efficient transfer.

[0703] Data processing and data calculation

[0704] The device uses a speech recognition engine (for example, Google Speech Recognition API) to convert speech into text data in real time. This text data is then sent to a server in the cloud via the application.

[0705] The server uses a natural language processing engine (such as SpaCy or NLTK) to analyze the text data and extract important keywords and context. Then, it uses a generative AI model (such as OpenAI's GPT-3) to generate the optimal answer and relevant information.

[0706] The generated answers are sent from the server to the terminal, where the answers and related information are displayed on the screen. Important information is also communicated to the user via voice using speech synthesis technology.

[0707] Specific example

[0708] For example, consider a customer support scenario where a user asks, "My product isn't working, what should I do?" The device records the user's voice and converts it into text data using a speech recognition engine. This text data is then sent to a server, where a natural language processing engine extracts the keywords "product" and "not working." A generative AI model then consults a relevant FAQ database to generate an answer that includes instructions on how to reset the product and links to support videos. Finally, the generated answer is sent back to the device, displayed on the screen, and announced via voice.

[0709] Example of a prompt

[0710] Prompt: "User Inquiry: [User Question]"

[0711] Example: "User inquiry: What should I do if the product doesn't work?"

[0712] Thus, the system of the present invention allows users to receive prompt and high-quality customer support. Furthermore, support staff can provide consistent and high-quality answers to user inquiries.

[0713] The flow of a specific process in Application Example 1 will be explained using Figure 12.

[0714] Step 1:

[0715] The user asks a question into the device (smartphone or smart glasses). The device's microphone records the audio. The input is the user's voice, and the output is the recorded audio data. Specifically, the microphone captures the user's speech and temporarily stores it as a buffer within the application.

[0716] Step 2:

[0717] The device sends recorded audio data to a speech recognition engine in real time, where it is converted into text data. The input is recorded audio data, and the output is transcribed text data. Specifically, the SpeechRecognition library is used to convert the audio data into text, which is then stored in a variable within the application.

[0718] Step 3:

[0719] The device sends characterized text data to a server in the cloud using the HTTPS protocol. The input is character data, and the output is the state as it has been sent to the server. Specifically, the Requests library is used to send the character data as a POST request to the server's API endpoint.

[0720] Step 4:

[0721] The server receives text data and analyzes it using a natural language processing engine (e.g., SpaCy). The input is text data, and the output is data with key keywords and context extracted. Specifically, the NLP engine on the cloud server analyzes the text data and identifies important words and phrases within the text.

[0722] Step 5:

[0723] The server uses a generative AI model (e.g., GPT-3) to generate answers and related information based on the analysis results. The input is keywords and contextual information from the analysis results, and the output is the generated answers and related information. Specifically, it generates a prompt sentence, inputs it into the generative AI model, and generates an appropriate answer.

[0724] Step 6:

[0725] The server sends the generated answer and related information back to the terminal using the HTTPS protocol. The input is the generated answer and related information, and the output is the state as it has been sent to the terminal. Specifically, the server encodes the appropriate answer and sends it in a format that the terminal can understand.

[0726] Step 7:

[0727] The device displays answers and related information on its screen and notifies the user of important information via voice. Input consists of answers and related information returned from the server, while output is visual and auditory notification to the user. Specifically, the application displays the answers on the device's display and uses speech synthesis technology to convey important points verbally.

[0728] By following these steps, the system will be able to provide real-time answers to user inquiries.

[0729] Furthermore, an emotion engine that estimates the user's emotions may be incorporated. That is, the identification processing unit 290 may use the emotion identification model 59 to estimate the user's emotions and perform identification processing using the user's emotions.

[0730] System Overview

[0731] This invention relates to a system that analyzes telephone and in-person consultations in real time and provides appropriate answers and related information. Furthermore, it is a system that achieves higher user satisfaction by analyzing the user's emotions and adjusting the content of the answers and information provided based on those emotions. This system consists of the user's terminal, a server on the cloud, an emotion engine, and communication means. The following describes each part of the system in detail.

[0732] Voice input and text conversion

[0733] 1. Users receive consultations via telephone or in person.

[0734] Users will use their own devices (smartphones, PCs, smart glasses, etc.) to provide consultation.

[0735] 2. The device records the user's consultation content as audio.

[0736] The device's application uses the microphone function to capture the audio of the consultation in real time. The audio data is temporarily stored in a buffer.

[0737] 3. The device uses a speech recognition engine to convert speech into text.

[0738] The device sends voice data to a speech recognition engine, which converts it into text data. The converted text is then stored within the application.

[0739] Emotion analysis

[0740] 4. The device sends the recorded audio data to the emotion engine, which analyzes the user's emotions.

[0741] The emotion engine analyzes the user's emotions (e.g., anger, anxiety, joy) from voice data and generates emotion data.

[0742] Text data transmission and analysis

[0743] 5. The device sends the transcribed consultation content and emotional data to the server.

[0744] Text data and sentiment data are sent from the device to a server in the cloud using the HTTPS protocol.

[0745] 6. The server receives the transcribed data and parses it.

[0746] The server's natural language processing (NLP) engine analyzes the text data and extracts important keywords and context.

[0747] 7. The server uses generative AI to generate answers and related information.

[0748] The server generates answers and related information using a database of FAQs related to the problem and a pre-trained generative AI model.

[0749] Adjust and submit your answer.

[0750] 8. The server uses sentiment data to adjust and filter answers and related information.

[0751] The server considers the user's sentiment data received from the sentiment engine to filter and select the most appropriate answers and relevant information.

[0752] 9. The server sends the answer and related information to the terminal.

[0753] Encode the optimal answers, links, and materials and send them back to your device.

[0754] Display and notification of answers

[0755] 10. The device notifies the user of the answer and related information.

[0756] The application on the device displays the received answers and information on the screen and notifies the user by voice as needed. Based on sentiment data, the displayed content and the tone of voice notifications are adjusted.

[0757] Specific example

[0758] For customer support

[0759] 1. A user contacts us because the product is not working.

[0760] 2. The device records the audio and transcribes the voice message, "The product isn't working, what should I do?"

[0761] 3. The device sends voice data to the emotion engine, which analyzes whether the user is feeling anxious.

[0762] 4. The device sends the transcribed data and sentiment data to the server.

[0763] 5. The server analyzes the text data and extracts the keywords "product" and "not working".

[0764] 6. The server refers to the relevant FAQ database and generates answers, reset instructions, and support video links.

[0765] 7. The server considers emotional data and selects a more polite and reassuring answer.

[0766] 8. The server sends the selected information to the terminal.

[0767] 9. The device displays instructions on how to reset the device and a link to a support video on the screen, and provides a voice notification in a gentle tone saying, "Here's how to reset the device."

[0768] In the case of sales

[0769] 1. The user receives an inquiry about the details of the new product.

[0770] 2. The device records the audio and transcribes the phrase "Please tell me more about the new product" into text.

[0771] 3. The device sends voice data to the emotion engine, which analyzes whether the user is interested.

[0772] 4. The device sends the transcribed data and sentiment data to the server.

[0773] 5. The server analyzes the text data and extracts the keywords "new product" and "details".

[0774] 6. The server generates related sales materials, product brochures, and links to explanatory videos.

[0775] 7. The server considers sentiment data and selects an explanation that will pique interest.

[0776] 8. The server sends the selected information to the terminal.

[0777] 9. The device displays a link to a brochure and product features on the screen, and provides a lively voice notification saying, "Here are the details of the new product."

[0778] Thus, the system of the present invention enables consistent, high-quality answers and information provision based on the user's emotions, is highly practical in the field, and contributes to improving work efficiency and user satisfaction.

[0779] The following describes the processing flow.

[0780] Step 1:

[0781] Users receive consultations via phone or in person.

[0782] Users will use their own devices (smartphones, PCs, smart glasses, etc.) to provide consultation.

[0783] Step 2:

[0784] The device records the user's conversation audio.

[0785] The device's application uses the microphone function to capture the audio of the consultation in real time. The audio data is temporarily stored in a buffer.

[0786] Step 3:

[0787] The device sends voice data to an emotion engine, which then analyzes the user's emotions.

[0788] The device sends voice data to an emotion engine (for example, a voice analysis API), which analyzes the user's emotions (anger, anxiety, joy, etc.) from the voice. The analysis results are stored as emotion data.

[0789] Step 4:

[0790] The device uses a speech recognition engine to convert speech into text.

[0791] The device sends the voice data to a speech recognition engine (e.g., a speech recognition API) and converts it into text data. The converted text is then stored within the application.

[0792] Step 5:

[0793] The device sends the transcribed consultation content and emotional data to the server.

[0794] Text data and sentiment data are sent from the device to a server in the cloud using the HTTPS protocol.

[0795] Step 6:

[0796] The server receives the transcribed data and parses it.

[0797] The server's natural language processing (NLP) engine analyzes the text data and extracts important keywords and context. The analysis results are saved as analysis data.

[0798] Step 7:

[0799] The server uses generative AI to generate answers and related information.

[0800] Based on the analysis data, the server generates answers and related information using a database of FAQs relevant to the problem and a pre-trained generative AI model. The generated information is temporarily stored.

[0801] Step 8:

[0802] The server uses sentiment data to adjust and filter answers and related information.

[0803] The server considers the user's emotional data received from the emotion engine and filters and selects the most appropriate answers and relevant information. The selected information is then prepared as data to be sent.

[0804] Step 9:

[0805] The server sends the answer and related information to the terminal.

[0806] The transmitted data is encoded and sent back to the device. This return transmission takes place via cloud communication.

[0807] Step 10:

[0808] The device notifies the user of the answer and related information.

[0809] The application on the device displays the received answers and information on the screen and notifies the user by voice as needed. Based on sentiment data, the displayed content and the tone of voice notifications are adjusted.

[0810] (Example 2)

[0811] Next, we will describe Example 2. In the following description, the data processing device 12 will be referred to as the "server" and the smart glasses 214 will be referred to as the "terminal".

[0812] Conventional consultation systems only analyze data transcribed from user speech, failing to consider user emotions and making it difficult to provide appropriate and personalized answers. Furthermore, they cannot provide notifications in a tone adjusted to the user's emotions, thus failing to adequately improve user satisfaction. Therefore, there was a need for a more advanced consultation system that analyzes user emotions in real time and provides emotion-based answers.

[0813] The identification process performed by the identification processing unit 290 of the data processing device 12 in Example 2 is realized by the following means.

[0814] In this invention, the server includes means for the terminal to transmit text data and sentiment data to the server, means for the server to analyze the text data, extract important keywords and context, and adjust the answer and related information using the sentiment data, and means for the server to transmit the generated answer and related information to the terminal. This enables the provision of appropriate answers based on the user's emotions and notifications in an adjusted tone.

[0815] A "terminal" is an information processing device used by users to input voice and send and receive data.

[0816] "Communication methods" refer to network technologies used to send and receive data between terminals and servers.

[0817] "Recording audio" means acquiring and saving the user's speech as digital audio data.

[0818] "Real-time transcription" refers to the process of instantly converting recorded audio data into text data.

[0819] "Emotional data" refers to information that indicates the user's emotional state, analyzed from voice data.

[0820] A "server" is a remote computer system that analyzes data received from a terminal and generates appropriate answers or information.

[0821] "Analyzing text data" means processing the text data received by the server and extracting important keywords and context.

[0822] "Key keywords" are words or phrases that are particularly important for understanding the user's inquiry.

[0823] "Extracting context" means analyzing the relationships between words in order to understand the meaning and intent of the entire text.

[0824] "Generating answers and related information" means that the server automatically creates appropriate answers and reference information based on the user's inquiry.

[0825] A "generative AI model" is an artificial intelligence system that uses technologies such as natural language processing to automatically generate text from input data.

[0826] "Adjusted tone" refers to appropriately modifying the speech style and nuances of voice notifications based on the user's emotional data.

[0827] "Displaying on the screen" means visually presenting answers or information on the device's display.

[0828] "Notifying by voice" means presenting answers or information to the user audibly using speech synthesis technology.

[0829] This invention is a system that analyzes user inquiries in real time and provides appropriate answers and related information. This system also analyzes the user's emotions and adjusts the content of the answers and information provided based on the emotional data, thereby achieving higher user satisfaction.

[0830] The main components of the system consist of the user's terminal, a server in the cloud, an emotion engine, and communication methods. Details of each component are shown below.

[0831] terminal

[0832] This is an information processing device that allows users to input voice data and send and receive data. Specifically, this includes smartphones, PCs, and smart glasses. The device incorporates a voice recording function, a speech recognition engine to convert voice data into text, and communication means for sending and receiving the results of sentiment analysis.

[0833] server

[0834] This is a cloud-based computer system that analyzes text and sentiment data received from terminals. The server uses a natural language processing (NLP) engine to analyze text data and generates answers and related information using generative AI models (e.g., OpenAI GPT-3). It also plays a role in adjusting answers by taking sentiment data into consideration.

[0835] means of communication

[0836] This is a network technology for sending and receiving data between a terminal and a server. Specifically, it involves data transmission using the HTTPS protocol.

[0837] Examples

[0838] For customer support

[0839] 1. The user asks, "The product isn't working, what should I do?"

[0840] 2. The device records the audio and uses a speech recognition engine (e.g., Google Speech-to-Text API) to transcribe it. The generated text data will be "The product isn't working, what should I do?"

[0841] 3. The device sends voice data to an emotion engine (e.g., IBM Watson Tone Analyzer) to obtain emotion data indicating anxiety.

[0842] 4. The device sends text data and sentiment data to the server.

[0843] 5. The server uses a natural language processing engine (e.g., BERT model) to analyze the text data and extract key keywords such as "product" and "not working".

[0844] 6. The server uses a generative AI model (e.g., OpenAI GPT-3) to generate "Product Reset Instructions" and "Support Video Links".

[0845] 7. Take into account the server's emotional data indicating anxiety, and adjust the response to be polite and reassuring.

[0846] 8. The server sends the adjusted answer to the terminal.

[0847] 9. The device displays "Here's how to reset it" on the screen and also provides a voice notification in a gentle tone.

[0848] Examples of prompt statements

[0849] For customer support

[0850] Prompt message:

[0851] User: "The product isn't working, what should I do?" Sentiment analysis result: Anxiety.

[0852] Generative AI models:

[0853] 1. How to reset the product

[0854] 2. Support video links

[0855] 3. Related answers from the FAQ

[0856] Taking these factors into consideration, please generate your answer in a gentle tone.

[0857] The above details the embodiment of the system of the present invention. This system provides high-quality answers and information based on the user's emotions, thereby improving work efficiency and user satisfaction.

[0858] The flow of the specific processing in Example 2 will be explained using Figure 13.

[0859] Step 1:

[0860] The user initiates a consultation.

[0861] Input: User's voice (e.g., "The product isn't working, what should I do?").

[0862] Specific action: The user speaks their question into the device's microphone.

[0863] Output: Raw audio data.

[0864] Step 2:

[0865] The device records audio.

[0866] Input: Raw audio data.

[0867] Specific operation: The device uses its microphone function to record the user's voice in real time and temporarily stores it in a buffer as audio data.

[0868] Output: Audio data stored in the buffer.

[0869] Step 3:

[0870] The device converts speech into text.

[0871] Input: Audio data stored in the buffer.

[0872] Specific operation: The device sends audio data to a speech recognition engine (e.g., Google Speech-to-Text API), and the audio data is converted into text data.

[0873] Output: Text data (Example: "The product isn't working, what should I do?").

[0874] Step 4:

[0875] The device performs emotion analysis.

[0876] Input: Audio data and text data.

[0877] Specific operation: The device sends voice data to an emotion engine (e.g., IBM Watson Tone Analyzer), which analyzes the emotion and obtains emotion data.

[0878] Output: Emotional data (e.g., anxiety).

[0879] Step 5:

[0880] The device sends data to the server.

[0881] Input: Text data and sentiment data.

[0882] Specific operation: The device sends text data and sentiment data to a server in the cloud using the HTTPS protocol.

[0883] Output: Text data and sentiment data sent to the server.

[0884] Step 6:

[0885] The server analyzes the text data.

[0886] Input: Character data sent to a server in the cloud.

[0887] Specific operation: The server uses a natural language processing (NLP) engine (e.g., BERT model) to analyze the text data and extract important keywords and context.

[0888] Output: Extracted keywords (e.g., "product", "not working").

[0889] Step 7:

[0890] The server generates the answers and related information.

[0891] Input: Extracted keywords.

[0892] Specific operation: The server uses a generative AI model (e.g., OpenAI GPT-3) to generate answers and related information.

[0893] Output: Generated answers and related information (e.g., "Product Reset Instructions," "Support Video Link").

[0894] Step 8:

[0895] The server adjusts the answers based on sentiment data.

[0896] Input: Sentimental data and generated answers or related information.

[0897] Specific operation: The server considers the sentiment data received from the sentiment engine and adjusts the response to have the most appropriate tone and content.

[0898] Output: Adjusted answer (e.g., a polite and reassuring answer).

[0899] Step 9:

[0900] The server sends the answer to the terminal.

[0901] Input: Adjusted answers and related information.

[0902] Specific operation: The server encodes the adjusted answer and related information and sends it to the terminal.

[0903] Output: The adjusted answers and related information sent to the terminal.

[0904] Step 10:

[0905] The device will display and notify you of the answer.

[0906] Input: Adjusted answers and related information sent from the server.

[0907] Specific actions: The device displays answers and information on the screen and provides voice notifications in a tone adjusted based on emotional data.

[0908] Output: The answers and information notified to the user.

[0909] (Application Example 2)

[0910] Next, we will explain application example 2. In the following explanation, the data processing device 12 will be referred to as the "server," and the smart glasses 214 will be referred to as the "terminal."

[0911] Conventional consultation systems using speech recognition have difficulty providing answers and information that take into account the user's emotions, and in particular, providing appropriate response instructions in real time has been difficult in security service settings. As a result, the quality of responses by field staff has declined, hindering security operations that require rapid response. The present invention aims to solve these problems and provide a system that improves the quality of responses by field staff by analyzing the user's emotions and providing appropriate response instructions in real time.

[0912] The specific processing performed by the specific processing unit 290 of the data processing device 12 in Application Example 2 is realized by the following means.

[0913] In this invention, the server includes means for analyzing transcribed data, extracting important keywords and context, and generating answers and related information based on them; means for analyzing the user's emotions and adjusting the answers and related information based on the analysis results; and means for transmitting the generated answers and related information to the terminal. This makes it possible to provide appropriate response instructions and information based on the user's emotions.

[0914] A "terminal" is an electronic device operated by a user that performs functions such as recording voice, transcribing it into text, displaying sentiment analysis results, and sending notifications.

[0915] "Communication method" refers to technology that uses Internet protocols (e.g., HTTPS) to send and receive data between a terminal and a server.

[0916] A "speech recognition engine" is software or a system that converts speech data into text data.

[0917] An "emotion analysis engine" is software or a system that analyzes a user's emotions from voice data or text data and outputs the results as data.

[0918] A "server" is a computer system that resides in the cloud and performs tasks such as analyzing text data, utilizing sentiment analysis results, and generating and transmitting answers and related information.

[0919] A "natural language processing engine" is software or a system that analyzes text data, understands keywords and context, and generates appropriate answers and related information.

[0920] "Generative AI" is artificial intelligence that uses a pre-trained model to generate answers and information in natural language based on input data.

[0921] "Answers and related information" refers to the answers, related data, links, and materials provided in response to a user's query.

[0922] "Emotional data" refers to information generated by an emotion analysis engine that indicates the user's emotional state.

[0923] The system of the present invention analyzes voice data in real time and provides appropriate answers and relevant information based on the user's emotions. This system is intended to support on-site response in security services, and its main components include a terminal, a server, an emotion analysis engine, a speech recognition engine, a natural language processing engine, and a generative AI model.

[0924] Overall system configuration

[0925] 1. Terminal:

[0926] It is a user-operated electronic device that records voice, transcribes it into text, displays sentiment analysis results, and provides notifications.

[0927] Smartphones, tablets, and PCs are examples.

[0928] 2. Speech recognition engine:

[0929] Software that converts audio data into text data. This often utilizes APIs such as Google's Speech-to-Text API.

[0930] 3. Emotion Analysis Engine:

[0931] Software that analyzes a user's emotions from audio or text data and outputs the results as data. IBM Watson Tone Analyzer is used.

[0932] 4. Server:

[0933] It is a computer system that resides in the cloud and performs text data analysis, utilizes sentiment analysis results, and generates and transmits answers and related information.

[0934] The cloud server runs the natural language processing engines (NLTK, SpaCy) and the generative AI model (OpenAI GPT-4).

[0935] 5. Means of communication:

[0936] The Internet Protocol (HTTPS) is used to send and receive data between the terminal and the server.

[0937] System Operation Overview

[0938] When a user makes a voice report at a security site, the device records the voice and converts it into text data in real time using a speech recognition engine. This text data is sent to an emotion analysis engine to identify the user's emotions. The emotion data and text data are sent to a cloud server, where a natural language processing engine analyzes the text data and extracts important keywords and context. Subsequently, a generative AI model generates appropriate answers and relevant information, and the answers are refined based on the emotion data. The optimized answers are sent to the device and displayed and notified to the user.

[0939] Specific example

[0940] For example, suppose a security officer reports, "An alarm has been triggered at the north entrance of the building. Please check the situation." This audio data is recorded on a terminal and converted into text data, "An alarm has been triggered at the north entrance of the building," using Google's Speech-to-Text API. This text data is then analyzed by IBM Watson Tone Analyzer, generating the sentiment data "urgent." The text data and sentiment data sent to the cloud server are analyzed by a natural language processing engine, and a generative AI model generates the optimal response based on the following prompt sentences.

[0941] Example of a prompt:

[0942] "Alarm: An alarm has been triggered at the north entrance of the building."

[0943] Emotion: Urgent

[0944] Possible steps: 1. Check the north entrance with the camera.

[0945] 2. Contact the nearest guard.

[0946] 3. Call the police if necessary.

[0947] Please generate the optimal solution.

[0948] Based on this, the generated answers and related information are sent to the device and notified to the user. The user can then quickly take appropriate action by following the instructions displayed on the device.

[0949] The flow of a specific process in Application Example 2 will be explained using Figure 14.

[0950] Step 1:

[0951] The device records audio reported by the user at the security site. This audio data is temporarily stored in a buffer on the device. The input is the user's voice, and the output is the recorded audio data.

[0952] Step 2:

[0953] The device sends recorded audio data in real time to a speech recognition engine (for example, Google's Speech-to-Text API). This engine converts the audio data into text data. Specifically, the audio data is sent to the API, and text data is returned. The input is the recorded audio data, and the output is the converted text data.

[0954] Step 3:

[0955] The terminal sends transcribed text data to an emotion analysis engine (e.g., IBM Watson Tone Analyzer) to analyze the user's emotions. The emotion analysis engine identifies emotions (such as urgency, anxiety, or joy) from the text data and generates emotion data. The input is transcribed text data, and the output is emotion data.

[0956] Step 4:

[0957] The device sends transcribed text data and sentiment data to a server in the cloud using the HTTPS protocol. Specifically, data packaged in JSON format or similar is sent. The input is text data and sentiment data, and the output is the data sent to the server.

[0958] Step 5:

[0959] The server analyzes the received text and sentiment data. It uses its natural language processing engine (e.g., NLTK or SpaCy) to extract keywords and context from the text data. The input is the text and sentiment data sent to the server, and the output is the extracted keyword and context data.

[0960] Step 6:

[0961] The server uses a generative AI model (e.g., OpenAI GPT-4) to generate answers and related information based on extracted keywords and contextual data. An example of a generated prompt is as follows:

[0962] "Alarm: An alarm has been triggered at the north entrance of the building. Emotion: Emergency Possible actions: 1. Check the north entrance with the camera 2. Contact the nearest guard 3. Call the police if necessary Please generate the best solution."

[0963] The input consists of keywords and contextual data, while the output consists of generated answers and related information.

[0964] Step 7:

[0965] The server adjusts the generated answers and related information based on sentiment data and selects the most appropriate one. This adjustment takes into account the user's current emotional state (e.g., urgent). The input is sentiment data and the generated answers and related information, and the output is the adjusted answers and related information.

[0966] Step 8:

[0967] The server sends the adjusted answers and related information to the terminal. The input is the adjusted answers and related information, and the output is the data sent to the terminal.

[0968] Step 9:

[0969] The device displays the adjusted answers and related information received from the server on its screen and, if necessary, informs the user via voice notification. Specifically, the answers are displayed on the device's screen, and necessary notifications are played through the speaker. The input is the answers and related information sent to the device, and the output is the answers and related information notified to the user.

[0970] The specific processing unit 290 transmits the result of the specific processing to the smart glasses 214. In the smart glasses 214, the control unit 46A causes the speaker 240 to output the result of the specific processing. The microphone 238 acquires audio indicating user input for the result of the specific processing. The control unit 46A transmits the audio data indicating user input acquired by the microphone 238 to the data processing unit 12. In the data processing unit 12, the specific processing unit 290 acquires the audio data.

[0971] Data generation model 58 is a type of so-called generative AI (Artificial Intelligence). One example of data generation model 58 is ChatGPT (Internet search<URL: https: / / openai.com / blog / chatgpt> ), Gemini (Internet search) <url: https: gemini.google.com ?hl="ja">Examples of generative AI include the following. The data generation model 58 is obtained by performing deep learning on a neural network. The data generation model 58 is input with prompts containing instructions, and with inference data such as audio data representing speech, text data representing text, and image data representing images. The data generation model 58 infers from the input inference data according to the instructions indicated by the prompts, and outputs the inference results in data formats such as audio data and text data. Here, inference refers to, for example, analysis, classification, prediction, and / or summarization.

[0972] In the above embodiment, an example was given in which specific processing is performed by the data processing device 12, but the technology of this disclosure is not limited thereto, and the specific processing may also be performed by the smart glasses 214.

[0973] [Third Embodiment]

[0974] Figure 5 shows an example of the configuration of the data processing system 310 according to the third embodiment.

[0975] As shown in Figure 5, the data processing system 310 includes a data processing device 12 and a headset terminal 314. An example of the data processing device 12 is a server.

[0976] The data processing device 12 comprises a computer 22, a database 24, and a communication interface 26. The computer 22 is an example of a "computer" related to the technology of this disclosure. The computer 22 comprises a processor 28, RAM 30, and storage 32. The processor 28, RAM 30, and storage 32 are connected to a bus 34. The database 24 and the communication interface 26 are also connected to the bus 34. The communication interface 26 is connected to a network 54. An example of the network 54 is a WAN (Wide Area Network) and / or a LAN (Local Area Network).

[0977] The headset terminal 314 includes a computer 36, a microphone 238, a speaker 240, a camera 42, a communication interface 44, and a display 343. The computer 36 includes a processor 46, RAM 48, and storage 50. The processor 46, RAM 48, and storage 50 are connected to a bus 52. The microphone 238, speaker 240, camera 42, and display 343 are also connected to the bus 52.

[0978] The microphone 238 receives voice signals from the user 20 and receives instructions from the user 20. The microphone 238 captures the voice signals from the user 20, converts the captured voice into audio data, and outputs it to the processor 46. The speaker 240 outputs audio according to the instructions from the processor 46.

[0979] Camera 42 is a small digital camera equipped with an optical system including a lens, aperture, and shutter, and an image sensor such as a CMOS (Complementary Metal-Oxide-Semiconductor) image sensor or a CCD (Charge Coupled Device) image sensor, and captures images of the area around the user 20 (for example, an imaging range defined by a field of view equivalent to the width of a typical healthy person's field of vision).

[0980] Communication interface 44 is connected to network 54. Communication interfaces 44 and 26 are responsible for the exchange of various information between processor 46 and processor 28 via network 54. The exchange of various information between processor 46 and processor 28 using communication interfaces 44 and 26 is performed in a secure manner.

[0981] Figure 6 shows an example of the main functions of the data processing device 12 and the headset terminal 314. As shown in Figure 6, the data processing device 12 performs specific processing using the processor 28. The storage 32 stores the specific processing program 56.

[0982] The specific processing program 56 is an example of a "program" relating to the technology of this disclosure. The processor 28 reads the specific processing program 56 from the storage 32 and executes the read specific processing program 56 on the RAM 30. The specific processing is realized by the processor 28 operating as a specific processing unit 290 in accordance with the specific processing program 56 executed on the RAM 30.

[0983] The storage 32 stores the data generation model 58 and the emotion identification model 59. The data generation model 58 and the emotion identification model 59 are used by the identification processing unit 290.

[0984] In the headset terminal 314, the processor 46 performs the reception output processing. The storage 50 stores the reception output program 60. The processor 46 reads the reception output program 60 from the storage 50 and executes the read reception output program 60 on the RAM 48. The reception output processing is realized by the processor 46 operating as a control unit 46A according to the reception output program 60 executed on the RAM 48.

[0985] Next, the specific processing performed by the specific processing unit 290 of the data processing device 12 will be described. In the following description, the data processing device 12 will be referred to as the "server" and the headset terminal 314 will be referred to as the "terminal".

[0986] System Overview

[0987] This invention relates to a system that analyzes telephone and in-person consultations in real time and provides appropriate answers and related information. This system consists of a user's terminal, a cloud-based server, and communication means. The following describes each part of the system in detail.

[0988] Voice input and text conversion

[0989] 1. Users receive consultations via telephone or in person.

[0990] Users will use their own devices (smartphones, PCs, smart glasses, etc.) to provide consultation.

[0991] 2. The device records the user's consultation content as audio.

[0992] The device's application uses the microphone function to capture the audio of the consultation in real time. The audio data is temporarily stored in a buffer.

[0993] 3. The device uses a speech recognition engine to convert speech into text.

[0994] The device sends voice data to a speech recognition engine, which converts it into text data. The converted text is then stored within the application.

[0995] Text data transmission and analysis

[0996] 4. The device sends the transcribed consultation content to the server.

[0997] Text data is sent from the device to a server in the cloud using the HTTPS protocol.

[0998] 5. The server receives the transcribed data and parses it.

[0999] The server's natural language processing (NLP) engine analyzes the text data and extracts important keywords and context.

[1000] Solution generation and search

[1001] 6. The server uses generative AI to generate answers and related information.

[1002] The server generates answers and related information using a database of FAQs related to the problem and a pre-trained generative AI model.

[1003] 7. The server filters the generated answers and related information and selects the appropriate ones.

[1004] From the multiple generated answers, filter and select the most appropriate information.

[1005] Submit and display of answers

[1006] 8. The server sends the answer and related information to the terminal.

[1007] Encode the optimal answers, links, and materials and send them back to your device.

[1008] 9. The device notifies the user of the answer and related information.

[1009] The application on the device displays the received answers and information on the screen and notifies the user by voice as needed.

[1010] Specific example

[1011] For customer support

[1012] 1. A user contacts us because the product is not working.

[1013] 2. The device records the audio and transcribes the voice message, "The product isn't working, what should I do?"

[1014] 3. The device sends the converted data to the server.

[1015] 4. The server analyzes the text data and extracts the keywords "product" and "not working".

[1016] 5. The server refers to the relevant FAQ database and generates answers, reset instructions, and support video links.

[1017] 6. Filter the information generated by the server and send it to the terminal.

[1018] 7. The device will display instructions on how to reset the device and a link to a support video on the screen, and will announce a voice message saying, "Here's how to reset the device."

[1019] In the case of sales

[1020] 1. The user receives an inquiry about the details of the new product.

[1021] 2. The device records the audio and transcribes the phrase "Please tell me more about the new product" into text.

[1022] 3. The device sends the converted data to the server.

[1023] 4. The server analyzes the text data and extracts the keywords "new product" and "details".

[1024] 5. The server generates related sales materials, product brochures, and links to explanatory videos.

[1025] 6. Filter the information generated by the server and send it to the terminal.

[1026] 7. The device displays a link to a brochure and product features on the screen, and announces with a voice message, "Here are the details of the new product."

[1027] Thus, the system of the present invention enables consistent, high-quality answers and information provision, regardless of the user's skill level. It offers high practicality in the field and contributes to improved work efficiency.

[1028] The following describes the processing flow.

[1029] Step 1:

[1030] Users receive consultations via phone or in person.

[1031] Users will use their own devices (smartphones, PCs, smart glasses, etc.) to provide consultation.

[1032] Step 2:

[1033] The device records the user's conversation audio.

[1034] The device's application uses the microphone function to capture the audio of the consultation in real time. The audio data is temporarily stored in a buffer.

[1035] Step 3:

[1036] The device uses a speech recognition engine to convert speech into text.

[1037] The device sends voice data to a speech recognition engine, which converts it into text data. The converted text is then stored within the application.

[1038] Step 4:

[1039] The terminal sends the transcribed consultation content to the server.

[1040] Text data is sent from the device to a server in the cloud using the HTTPS protocol.

[1041] Step 5:

[1042] The server receives the transcribed data and parses it.

[1043] The server's natural language processing (NLP) engine analyzes the text data and extracts important keywords and context.

[1044] Step 6:

[1045] The server uses generative AI to generate answers and related information.

[1046] The server generates answers and related information using a database of FAQs related to the problem and a pre-trained generative AI model.

[1047] Step 7:

[1048] The server filters the generated answers and related information and selects the appropriate ones.

[1049] From the multiple generated answers, filter and select the most appropriate information.

[1050] Step 8:

[1051] The server sends the answer and related information to the terminal.

[1052] Encode the optimal answers, links, and materials and send them back to your device.

[1053] Step 9:

[1054] The device notifies the user of the answer and related information.

[1055] The application on the device displays the received answers and information on the screen and notifies the user by voice as needed.

[1056] (Example 1)

[1057] Next, we will describe Example 1. In the following description, the data processing device 12 will be referred to as the "server," and the headset-type terminal 314 will be referred to as the "terminal."

[1058] In modern call centers and customer support, there is a demand for providing quick and accurate answers to user inquiries. However, conventional systems often suffer from low accuracy in voice-to-text conversion, and answer generation is frequently done manually, leading to delays and inconsistent quality. This has resulted in decreased customer satisfaction and reduced operational efficiency. This invention aims to solve these problems, streamline the process from voice input to answer provision, and achieve consistent, high-quality information delivery.

[1059] The identification process performed by the identification processing unit 290 of the data processing device 12 in Example 1 is realized by the following means.

[1060] In this invention, the server includes: means for the terminal to record the user's voice using communication means and transcribe the voice in real time; means for the terminal to transmit the transcribed data to the server; means for the server to analyze the transcribed data, extract important keywords and context, and generate answers and related information based on them; means for the server to filter the generated answers and related information and transmit it to the terminal; and means for the terminal to display the answers and related information on the screen and notify the user of important information by voice. This automates the process from voice input to answer provision, enabling the rapid, consistent, and high-quality provision of information.

[1061] A "terminal" is an information processing device used by a user, which acquires and processes data, including voice input. Specifically, this refers to smartphones, personal computers, smart glasses, etc.

[1062] "Communication means" refers to technical means for sending and receiving data, including the internet, Wi-Fi, Bluetooth, etc.

[1063] "Recording audio" means capturing a user's speech as digital audio data using a microphone and storing it.

[1064] "Real-time transcription of speech" means instantly analyzing speech data and converting it into corresponding text data. This is typically done using a speech recognition engine.

[1065] "Transcribed data" refers to text data obtained after converting voice input into written information.

[1066] A "server" is a high-performance information processing device that operates on the cloud and has functions such as data analysis, storage, and transmission.

[1067] "Analysis" is the process of breaking down received data and performing operations to understand it.

[1068] "Important keywords and context" refer to words and phrases that are particularly meaningful from the input data, as well as the relationships between them.

[1069] "Answers and related information" refers to information that allows for appropriate responses to user questions and inquiries. This includes generated text, links, and documents.

[1070] "Generating" means creating new information based on input data.

[1071] "Filtering" is the process of selecting the most appropriate answer or piece of information from multiple generated responses.

[1072] "Displaying on the screen" means making information visible as text or images on the device's display.

[1073] "Notifying by voice" means conveying text data to the user as voice using speech synthesis.

[1074] Modes for carrying out the invention

[1075] System Overview

[1076] This invention relates to a system that analyzes telephone and in-person consultations in real time and provides appropriate answers and related information. This system consists of a user's terminal, a cloud-based server, and communication means. The following describes each component of the system in detail.

[1077] Voice input and text conversion

[1078] The device records audio when the user has a consultation in person or over the phone. The device uses an information processing device such as a smartphone, PC, or smart glasses. The application on this device uses the microphone function to acquire the user's voice in real time and temporarily stores the audio data in memory.

[1079] Next, the device uses a speech recognition engine to convert the speech into text. For example, it might use the Google Cloud Speech-to-Text API to convert the acquired speech data into text data. This converted text data is then stored within the application.

[1080] Text data transmission and analysis

[1081] The terminal sends the transcribed consultation content to the server. The HTTPS protocol is used for this transmission, ensuring the secure transfer of data.

[1082] The server analyzes the received text data. A natural language processing (NLP) engine is used for the analysis. For example, the Google Cloud Natural Language API is used to extract important keywords and context from the text data.

[1083] Solution generation and search

[1084] The server uses a generative AI model to generate answers and related information. For example, a generative AI model like OpenAI's GPT-3 is used. The server also refers to a FAQ database related to the problem to generate answers and related information.

[1085] The server filters the multiple generated answers and related information, selecting the most appropriate one. This ensures that the necessary information is provided efficiently.

[1086] Submit and display of answers

[1087] The selected answers and related information are sent from the server to the terminal. The HTTPS protocol is used for this transmission, ensuring the secure exchange of information.

[1088] The device notifies the user of the received answers and related information. The application displays this information on the screen and, if necessary, informs the user via voice.

[1089] Specific example

[1090] Customer support scenarios

[1091] 1. The user asks, "The product isn't working, what should I do?"

[1092] 2. The device records this audio and transcribes it into text.

[1093] 3. The device sends the converted data to the server.

[1094] 4. The server analyzes the text data and extracts the keywords "product" and "not working".

[1095] 5. The server generates answers using the relevant FAQ database and generative AI models.

[1096] 6. The server filters the generated answers and selects the most appropriate information.

[1097] 7. The server sends the selected answer to the terminal.

[1098] 8. The device will announce a voice message saying, "Here's how to reset," and display the reset instructions on the screen.

[1099] Sales scenario

[1100] 1. The user asks, "Please tell me more about the new product."

[1101] 2. The device records this audio and transcribes it into text.

[1102] 3. The device sends the converted data to the server.

[1103] 4. The server analyzes the text data and extracts the keywords "new product" and "details".

[1104] 5. The server generates related sales materials, product brochures, and links to explanatory videos.

[1105] 6. The server filters the generated information and selects the most appropriate information.

[1106] 7. The server sends the selected information to the terminal.

[1107] 8. The device will announce, "Here are the details of the new product," and display a link to the brochure on the screen.

[1108] The flow of the specific processing in Example 1 will be explained using Figure 11.

[1109] Step 1:

[1110] The user initiates a consultation via voice. The user launches a dedicated application on their device, such as a smartphone or PC.

[1111] Specific action: The user asks a question such as, "Please tell me more about the new product."

[1112] Input: User's voice.

[1113] Output: Audio data is input to the terminal.

[1114] Step 2:

[1115] The device records the user's voice. The device's application uses the microphone function to capture the user's speech in real time and stores the audio data in memory.

[1116] Specific operation: The application acquires audio data from the device's microphone and temporarily stores it in a buffer.

[1117] Input: User's voice data.

[1118] Output: Audio data stored in the buffer.

[1119] Step 3:

[1120] The device uses a speech recognition engine to convert speech into text. For example, it sends speech data to the Google Cloud Speech-to-Text API and converts it into text data.

[1121] Specific operation: The speech recognition engine analyzes the audio data and generates text data.

[1122] Input: Audio data.

[1123] Output: Text data.

[1124] Step 4:

[1125] The device sends the transcribed data to the server. The device uses the HTTPS protocol to securely send the text data to the server in the cloud.

[1126] Specific operation: Encoded text data is sent via the HTTPS protocol.

[1127] Input: Text data.

[1128] Output: Text data sent to the server.

[1129] Step 5:

[1130] The server analyzes the received text data. An NLP engine is used to extract important keywords and context from the text. This process includes, for example, the Google Cloud Natural Language API.

[1131] Specific operation: The server's NLP engine analyzes the text data and extracts keywords such as "new product" and "details."

[1132] Input: Text data.

[1133] Output: Analyzed keywords and contextual information.

[1134] Step 6:

[1135] The server generates answers and related information using a generative AI model. It uses a generative AI model (for example, OpenAI's GPT-3) to generate appropriate answers and information based on the consultation content.

[1136] Specific operation: Input prompts into the generative AI model and obtain the generated answer.

[1137] Input: Analyzed keywords and contextual information.

[1138] Output: Generated answers and related information.

[1139] Step 7:

[1140] The server filters the generated answers and related information. It selects the most appropriate answer from the multiple answers generated. Criteria for evaluating reliability and relevance are applied to this process.

[1141] Specific operation: The filtering algorithm selects the optimal solution.

[1142] Input: Multiple generated answers.

[1143] Output: Optimal answers and related information.

[1144] Step 8:

[1145] The server sends the answer and related information to the terminal. The selected answer is re-encoded and sent to the terminal using the HTTPS protocol.

[1146] Specific operation: Encoded information is sent via the HTTPS protocol.

[1147] Input: The best answer or related information.

[1148] Output: Answers and related information sent to the terminal.

[1149] Step 9:

[1150] The device notifies the user of the answer and related information. The device's application displays the received information on the screen and, if necessary, communicates it to the user verbally.

[1151] Specific action: The application displays information on the screen and announces with a voice message, "Here are the details of the new product."

[1152] Input: Answers and related information sent to the device.

[1153] Output: Information displayed or notified to the user.

[1154] (Application Example 1)

[1155] Next, we will explain Application Example 1. In the following explanation, the data processing device 12 will be referred to as the "server," and the headset-type terminal 314 will be referred to as the "terminal."

[1156] Conventional customer support systems have struggled to provide timely and appropriate answers to user inquiries. This often resulted in delays in response times and a decline in the quality of answers. Furthermore, limitations in communication methods and devices posed a challenge in terms of usability. This invention aims to solve these problems and enable rapid and high-quality customer support.

[1157] The specific processing performed by the specific processing unit 290 of the data processing device 12 in Application Example 1 is realized by the following means.

[1158] In this invention, the server includes: [means for the terminal to record the user's voice using communication means and transcribe the voice in real time; [means for the terminal to transmit the transcribed data to the server]; [means for the server to analyze the transcribed data, extract important keywords and context, and generate answers and related information based thereon]; [means for the server to transmit the generated answers and related information to the terminal]; [means for the terminal to display the answers and related information on a screen and notify the user of important information by voice]; and [a portable information processing device on which an application that accepts the user's voice input is executed]. This enables rapid and high-quality customer support by transcribing voice from the terminal in real time and generating and displaying appropriate answers.

[1159] A "terminal" is a computer device used by a user that records audio and transmits the transcribed data to a server.

[1160] "Communication methods" refer to the technologies and protocols used for sending and receiving data, and are the means used to exchange information between a server and a terminal.

[1161] "Recording audio" refers to collecting the user's spoken content as digital data using a microphone.

[1162] "Real-time transcription" refers to the process of instantly converting acquired audio data into text data.

[1163] "Text-based data" refers to information that has been converted into text format by a speech recognition engine.

[1164] A "server" is a computer system located in the cloud that has the function of analyzing digitized data and generating answers.

[1165] "Analyzing" refers to the process of understanding received data using a natural language processing engine and extracting important keywords and context.

[1166] "Extracting important keywords and context" refers to the process of taking meaningful words and vocabulary from text data and combining them based on their context.

[1167] "Generating related information" refers to constructing answers or additional information based on extracted keywords and context.

[1168] "Displaying on the screen" refers to providing answers or information in a visually apparent form on the device's display.

[1169] "Notifying the user via voice" refers to the process of communicating the generated answer to the user using speech synthesis technology.

[1170] A "portable information processing device" refers to an electronic device that is portable and has the functionality to accept voice input.

[1171] This invention relates to a system that transcribes a user's voice into text in real time and provides answers and related information based on that transcription. This system mainly consists of a terminal, communication means, and a server. Specific embodiments of this system are described below.

[1172] Hardware and software

[1173] The terminals used are portable information processing devices such as smartphones and smart glasses. These devices are equipped with microphones and speech recognition engines to record and transcribe user voice input in real time. Furthermore, applications running on the terminals also have the functionality to send the text data to a server.

[1174] The server resides in the cloud and uses a natural language processing (NLP) engine and generative AI models to generate answers and related information. The server analyzes the text data sent by the user and extracts important keywords and context.

[1175] The communication method used is an internet connection for sending and receiving data between the terminal and the server. The communication protocol used is HTTPS to ensure data security and efficient transfer.

[1176] Data processing and data calculation

[1177] The device uses a speech recognition engine (for example, Google Speech Recognition API) to convert speech into text data in real time. This text data is then sent to a server in the cloud via the application.

[1178] The server uses a natural language processing engine (such as SpaCy or NLTK) to analyze the text data and extract important keywords and context. Then, it uses a generative AI model (such as OpenAI's GPT-3) to generate the optimal answer and relevant information.

[1179] The generated answers are sent from the server to the terminal, where the answers and related information are displayed on the screen. Important information is also communicated to the user via voice using speech synthesis technology.

[1180] Specific example

[1181] For example, consider a customer support scenario where a user asks, "My product isn't working, what should I do?" The device records the user's voice and converts it into text data using a speech recognition engine. This text data is then sent to a server, where a natural language processing engine extracts the keywords "product" and "not working." A generative AI model then consults a relevant FAQ database to generate an answer that includes instructions on how to reset the product and links to support videos. Finally, the generated answer is sent back to the device, displayed on the screen, and announced via voice.

[1182] Example of a prompt

[1183] Prompt: "User Inquiry: [User Question]"

[1184] Example: "User inquiry: What should I do if the product doesn't work?"

[1185] Thus, the system of the present invention allows users to receive prompt and high-quality customer support. Furthermore, support staff can provide consistent and high-quality answers to user inquiries.

[1186] The flow of a specific process in Application Example 1 will be explained using Figure 12.

[1187] Step 1:

[1188] The user asks a question into the device (smartphone or smart glasses). The device's microphone records the audio. The input is the user's voice, and the output is the recorded audio data. Specifically, the microphone captures the user's speech and temporarily stores it as a buffer within the application.

[1189] Step 2:

[1190] The device sends recorded audio data to a speech recognition engine in real time, where it is converted into text data. The input is recorded audio data, and the output is transcribed text data. Specifically, the SpeechRecognition library is used to convert the audio data into text, which is then stored in a variable within the application.

[1191] Step 3:

[1192] The device sends characterized text data to a server in the cloud using the HTTPS protocol. The input is character data, and the output is the state as it has been sent to the server. Specifically, the Requests library is used to send the character data as a POST request to the server's API endpoint.

[1193] Step 4:

[1194] The server receives text data and analyzes it using a natural language processing engine (e.g., SpaCy). The input is text data, and the output is data with key keywords and context extracted. Specifically, the NLP engine on the cloud server analyzes the text data and identifies important words and phrases within the text.

[1195] Step 5:

[1196] The server uses a generative AI model (e.g., GPT-3) to generate answers and related information based on the analysis results. The input is keywords and contextual information from the analysis results, and the output is the generated answers and related information. Specifically, it generates a prompt sentence, inputs it into the generative AI model, and generates an appropriate answer.

[1197] Step 6:

[1198] The server sends the generated answer and related information back to the terminal using the HTTPS protocol. The input is the generated answer and related information, and the output is the state as it has been sent to the terminal. Specifically, the server encodes the appropriate answer and sends it in a format that the terminal can understand.

[1199] Step 7:

[1200] The device displays answers and related information on its screen and notifies the user of important information via voice. Input consists of answers and related information returned from the server, while output is visual and auditory notification to the user. Specifically, the application displays the answers on the device's display and uses speech synthesis technology to convey important points verbally.

[1201] By following these steps, the system will be able to provide real-time answers to user inquiries.

[1202] Furthermore, an emotion engine that estimates the user's emotions may be incorporated. That is, the identification processing unit 290 may use the emotion identification model 59 to estimate the user's emotions and perform identification processing using the user's emotions.

[1203] System Overview

[1204] This invention relates to a system that analyzes telephone and in-person consultations in real time and provides appropriate answers and related information. Furthermore, it is a system that achieves higher user satisfaction by analyzing the user's emotions and adjusting the content of the answers and information provided based on those emotions. This system consists of the user's terminal, a server on the cloud, an emotion engine, and communication means. The following describes each part of the system in detail.

[1205] Voice input and text conversion

[1206] 1. Users receive consultations via telephone or in person.

[1207] Users will use their own devices (smartphones, PCs, smart glasses, etc.) to provide consultation.

[1208] 2. The device records the user's consultation content as audio.

[1209] The device's application uses the microphone function to capture the audio of the consultation in real time. The audio data is temporarily stored in a buffer.

[1210] 3. The device uses a speech recognition engine to convert speech into text.

[1211] The device sends voice data to a speech recognition engine, which converts it into text data. The converted text is then stored within the application.

[1212] Emotion analysis

[1213] 4. The device sends the recorded audio data to the emotion engine, which analyzes the user's emotions.

[1214] The emotion engine analyzes the user's emotions (e.g., anger, anxiety, joy) from voice data and generates emotion data.

[1215] Text data transmission and analysis

[1216] 5. The device sends the transcribed consultation content and emotional data to the server.

[1217] Text data and sentiment data are sent from the device to a server in the cloud using the HTTPS protocol.

[1218] 6. The server receives the transcribed data and parses it.

[1219] The server's natural language processing (NLP) engine analyzes the text data and extracts important keywords and context.

[1220] 7. The server uses generative AI to generate answers and related information.

[1221] The server generates answers and related information using a database of FAQs related to the problem and a pre-trained generative AI model.

[1222] Adjust and submit your answer.

[1223] 8. The server uses sentiment data to adjust and filter answers and related information.

[1224] The server considers the user's sentiment data received from the sentiment engine to filter and select the most appropriate answers and relevant information.

[1225] 9. The server sends the answer and related information to the terminal.

[1226] Encode the optimal answers, links, and materials and send them back to your device.

[1227] Display and notification of answers

[1228] 10. The device notifies the user of the answer and related information.

[1229] The application on the device displays the received answers and information on the screen and notifies the user by voice as needed. Based on sentiment data, the displayed content and the tone of voice notifications are adjusted.

[1230] Specific example

[1231] For customer support

[1232] 1. A user contacts us because the product is not working.

[1233] 2. The device records the audio and transcribes the voice message, "The product isn't working, what should I do?"

[1234] 3. The device sends voice data to the emotion engine, which analyzes whether the user is feeling anxious.

[1235] 4. The device sends the transcribed data and sentiment data to the server.

[1236] 5. The server analyzes the text data and extracts the keywords "product" and "not working".

[1237] 6. The server refers to the relevant FAQ database and generates answers, reset instructions, and support video links.

[1238] 7. The server considers emotional data and selects a more polite and reassuring answer.

[1239] 8. The server sends the selected information to the terminal.

[1240] 9. The device displays instructions on how to reset the device and a link to a support video on the screen, and provides a voice notification in a gentle tone saying, "Here's how to reset the device."

[1241] In the case of sales

[1242] 1. The user receives an inquiry about the details of the new product.

[1243] 2. The device records the audio and transcribes the phrase "Please tell me more about the new product" into text.

[1244] 3. The device sends voice data to the emotion engine, which analyzes whether the user is interested.

[1245] 4. The device sends the transcribed data and sentiment data to the server.

[1246] 5. The server analyzes the text data and extracts the keywords "new product" and "details".

[1247] 6. The server generates related sales materials, product brochures, and links to explanatory videos.

[1248] 7. The server considers sentiment data and selects an explanation that will pique interest.

[1249] 8. The server sends the selected information to the terminal.

[1250] 9. The device displays a link to a brochure and product features on the screen, and provides a lively voice notification saying, "Here are the details of the new product."

[1251] Thus, the system of the present invention enables consistent, high-quality answers and information provision based on the user's emotions, is highly practical in the field, and contributes to improving work efficiency and user satisfaction.

[1252] The following describes the processing flow.

[1253] Step 1:

[1254] Users receive consultations via phone or in person.

[1255] Users will use their own devices (smartphones, PCs, smart glasses, etc.) to provide consultation.

[1256] Step 2:

[1257] The device records the user's conversation audio.

[1258] The device's application uses the microphone function to capture the audio of the consultation in real time. The audio data is temporarily stored in a buffer.

[1259] Step 3:

[1260] The device sends voice data to an emotion engine, which then analyzes the user's emotions.

[1261] The device sends voice data to an emotion engine (for example, a voice analysis API), which analyzes the user's emotions (anger, anxiety, joy, etc.) from the voice. The analysis results are stored as emotion data.

[1262] Step 4:

[1263] The device uses a speech recognition engine to convert speech into text.

[1264] The device sends the voice data to a speech recognition engine (e.g., a speech recognition API) and converts it into text data. The converted text is then stored within the application.

[1265] Step 5:

[1266] The device sends the transcribed consultation content and emotional data to the server.

[1267] Text data and sentiment data are sent from the device to a server in the cloud using the HTTPS protocol.

[1268] Step 6:

[1269] The server receives the transcribed data and parses it.

[1270] The server's natural language processing (NLP) engine analyzes the text data and extracts important keywords and context. The analysis results are saved as analysis data.

[1271] Step 7:

[1272] The server uses generative AI to generate answers and related information.

[1273] Based on the analysis data, the server generates answers and related information using a database of FAQs relevant to the problem and a pre-trained generative AI model. The generated information is temporarily stored.

[1274] Step 8:

[1275] The server uses sentiment data to adjust and filter answers and related information.

[1276] The server considers the user's emotional data received from the emotion engine and filters and selects the most appropriate answers and relevant information. The selected information is then prepared as data to be sent.

[1277] Step 9:

[1278] The server sends the answer and related information to the terminal.

[1279] The transmitted data is encoded and sent back to the device. This return transmission takes place via cloud communication.

[1280] Step 10:

[1281] The device notifies the user of the answer and related information.

[1282] The application on the device displays the received answers and information on the screen and notifies the user by voice as needed. Based on sentiment data, the displayed content and the tone of voice notifications are adjusted.

[1283] (Example 2)

[1284] Next, we will describe Example 2. In the following description, the data processing device 12 will be referred to as the "server," and the headset-type terminal 314 will be referred to as the "terminal."

[1285] Conventional consultation systems only analyze data transcribed from user speech, failing to consider user emotions and making it difficult to provide appropriate and personalized answers. Furthermore, they cannot provide notifications in a tone adjusted to the user's emotions, thus failing to adequately improve user satisfaction. Therefore, there was a need for a more advanced consultation system that analyzes user emotions in real time and provides emotion-based answers.

[1286] The identification process performed by the identification processing unit 290 of the data processing device 12 in Example 2 is realized by the following means.

[1287] In this invention, the server includes means for the terminal to transmit text data and sentiment data to the server, means for the server to analyze the text data, extract important keywords and context, and adjust the answer and related information using the sentiment data, and means for the server to transmit the generated answer and related information to the terminal. This enables the provision of appropriate answers based on the user's emotions and notifications in an adjusted tone.

[1288] A "terminal" is an information processing device used by users to input voice and send and receive data.

[1289] "Communication methods" refer to network technologies used to send and receive data between terminals and servers.

[1290] "Recording audio" means acquiring and saving the user's speech as digital audio data.

[1291] "Real-time transcription" refers to the process of instantly converting recorded audio data into text data.

[1292] "Emotional data" refers to information that indicates the user's emotional state, analyzed from voice data.

[1293] A "server" is a remote computer system that analyzes data received from a terminal and generates appropriate answers or information.

[1294] "Analyzing text data" means processing the text data received by the server and extracting important keywords and context.

[1295] "Key keywords" are words or phrases that are particularly important for understanding the user's inquiry.

[1296] "Extracting context" means analyzing the relationships between words in order to understand the meaning and intent of the entire text.

[1297] "Generating answers and related information" means that the server automatically creates appropriate answers and reference information based on the user's inquiry.

[1298] A "generative AI model" is an artificial intelligence system that uses technologies such as natural language processing to automatically generate text from input data.

[1299] "Adjusted tone" refers to appropriately modifying the speech style and nuances of voice notifications based on the user's emotional data.

[1300] "Displaying on the screen" means visually presenting answers or information on the device's display.

[1301] "Notifying by voice" means presenting answers or information to the user audibly using speech synthesis technology.

[1302] This invention is a system that analyzes user inquiries in real time and provides appropriate answers and related information. This system also analyzes the user's emotions and adjusts the content of the answers and information provided based on the emotional data, thereby achieving higher user satisfaction.

[1303] The main components of the system consist of the user's terminal, a server in the cloud, an emotion engine, and communication methods. Details of each component are shown below.

[1304] terminal

[1305] This is an information processing device that allows users to input voice data and send and receive data. Specifically, this includes smartphones, PCs, and smart glasses. The device incorporates a voice recording function, a speech recognition engine to convert voice data into text, and communication means for sending and receiving the results of sentiment analysis.

[1306] server

[1307] This is a cloud-based computer system that analyzes text and sentiment data received from terminals. The server uses a natural language processing (NLP) engine to analyze text data and generates answers and related information using generative AI models (e.g., OpenAI GPT-3). It also plays a role in adjusting answers by taking sentiment data into consideration.

[1308] means of communication

[1309] This is a network technology for sending and receiving data between a terminal and a server. Specifically, it involves data transmission using the HTTPS protocol.

[1310] Examples

[1311] For customer support

[1312] 1. The user asks, "The product isn't working, what should I do?"

[1313] 2. The device records the audio and uses a speech recognition engine (e.g., Google Speech-to-Text API) to transcribe it. The generated text data will be "The product isn't working, what should I do?"

[1314] 3. The device sends voice data to an emotion engine (e.g., IBM Watson Tone Analyzer) to obtain emotion data indicating anxiety.

[1315] 4. The device sends text data and sentiment data to the server.

[1316] 5. The server uses a natural language processing engine (e.g., BERT model) to analyze the text data and extract key keywords such as "product" and "not working".

[1317] 6. The server uses a generative AI model (e.g., OpenAI GPT-3) to generate "Product Reset Instructions" and "Support Video Links".

[1318] 7. Take into account the server's emotional data indicating anxiety, and adjust the response to be polite and reassuring.

[1319] 8. The server sends the adjusted answer to the terminal.

[1320] 9. The device displays "Here's how to reset it" on the screen and also provides a voice notification in a gentle tone.

[1321] Examples of prompt statements

[1322] For customer support

[1323] Prompt message:

[1324] User: "The product isn't working, what should I do?" Sentiment analysis result: Anxiety.

[1325] Generative AI models:

[1326] 1. How to reset the product

[1327] 2. Support video links

[1328] 3. Related answers from the FAQ

[1329] Taking these factors into consideration, please generate your answer in a gentle tone.

[1330] The above details the embodiment of the system of the present invention. This system provides high-quality answers and information based on the user's emotions, thereby improving work efficiency and user satisfaction.

[1331] The flow of the specific processing in Example 2 will be explained using Figure 13.

[1332] Step 1:

[1333] The user initiates a consultation.

[1334] Input: User's voice (e.g., "The product isn't working, what should I do?").

[1335] Specific action: The user speaks their question into the device's microphone.

[1336] Output: Raw audio data.

[1337] Step 2:

[1338] The device records audio.

[1339] Input: Raw audio data.

[1340] Specific operation: The device uses its microphone function to record the user's voice in real time and temporarily stores it in a buffer as audio data.

[1341] Output: Audio data stored in the buffer.

[1342] Step 3:

[1343] The device converts speech into text.

[1344] Input: Audio data stored in the buffer.

[1345] Specific operation: The device sends audio data to a speech recognition engine (e.g., Google Speech-to-Text API), and the audio data is converted into text data.

[1346] Output: Text data (Example: "The product isn't working, what should I do?").

[1347] Step 4:

[1348] The device performs emotion analysis.

[1349] Input: Audio data and text data.

[1350] Specific operation: The device sends voice data to an emotion engine (e.g., IBM Watson Tone Analyzer), which analyzes the emotion and obtains emotion data.

[1351] Output: Emotional data (e.g., anxiety).

[1352] Step 5:

[1353] The device sends data to the server.

[1354] Input: Text data and sentiment data.

[1355] Specific operation: The device sends text data and sentiment data to a server in the cloud using the HTTPS protocol.

[1356] Output: Text data and sentiment data sent to the server.

[1357] Step 6:

[1358] The server analyzes the text data.

[1359] Input: Character data sent to a server in the cloud.

[1360] Specific operation: The server uses a natural language processing (NLP) engine (e.g., BERT model) to analyze the text data and extract important keywords and context.

[1361] Output: Extracted keywords (e.g., "product", "not working").

[1362] Step 7:

[1363] The server generates the answers and related information.

[1364] Input: Extracted keywords.

[1365] Specific operation: The server uses a generative AI model (e.g., OpenAI GPT-3) to generate answers and related information.

[1366] Output: Generated answers and related information (e.g., "Product Reset Instructions," "Support Video Link").

[1367] Step 8:

[1368] The server adjusts the answers based on sentiment data.

[1369] Input: Sentimental data and generated answers or related information.

[1370] Specific operation: The server considers the sentiment data received from the sentiment engine and adjusts the response to have the most appropriate tone and content.

[1371] Output: Adjusted answer (e.g., a polite and reassuring answer).

[1372] Step 9:

[1373] The server sends the answer to the terminal.

[1374] Input: Adjusted answers and related information.

[1375] Specific operation: The server encodes the adjusted answer and related information and sends it to the terminal.

[1376] Output: The adjusted answers and related information sent to the terminal.

[1377] Step 10:

[1378] The device will display and notify you of the answer.

[1379] Input: Adjusted answers and related information sent from the server.

[1380] Specific actions: The device displays answers and information on the screen and provides voice notifications in a tone adjusted based on emotional data.

[1381] Output: The answers and information notified to the user.

[1382] (Application Example 2)

[1383] Next, we will explain application example 2. In the following explanation, the data processing device 12 will be referred to as the "server," and the headset-type terminal 314 will be referred to as the "terminal."

[1384] Conventional consultation systems using speech recognition have difficulty providing answers and information that take into account the user's emotions, and in particular, providing appropriate response instructions in real time has been difficult in security service settings. As a result, the quality of responses by field staff has declined, hindering security operations that require rapid response. The present invention aims to solve these problems and provide a system that improves the quality of responses by field staff by analyzing the user's emotions and providing appropriate response instructions in real time.

[1385] The specific processing performed by the specific processing unit 290 of the data processing device 12 in Application Example 2 is realized by the following means.

[1386] In this invention, the server includes means for analyzing transcribed data, extracting important keywords and context, and generating answers and related information based on them; means for analyzing the user's emotions and adjusting the answers and related information based on the analysis results; and means for transmitting the generated answers and related information to the terminal. This makes it possible to provide appropriate response instructions and information based on the user's emotions.

[1387] A "terminal" is an electronic device operated by a user that performs functions such as recording voice, transcribing it into text, displaying sentiment analysis results, and sending notifications.

[1388] "Communication method" refers to technology that uses Internet protocols (e.g., HTTPS) to send and receive data between a terminal and a server.

[1389] A "speech recognition engine" is software or a system that converts speech data into text data.

[1390] An "emotion analysis engine" is software or a system that analyzes a user's emotions from voice data or text data and outputs the results as data.

[1391] A "server" is a computer system that resides in the cloud and performs tasks such as analyzing text data, utilizing sentiment analysis results, and generating and transmitting answers and related information.

[1392] A "natural language processing engine" is software or a system that analyzes text data, understands keywords and context, and generates appropriate answers and related information.

[1393] "Generative AI" is artificial intelligence that uses a pre-trained model to generate answers and information in natural language based on input data.

[1394] "Answers and related information" refers to the answers, related data, links, and materials provided in response to a user's query.

[1395] "Emotional data" refers to information generated by an emotion analysis engine that indicates the user's emotional state.

[1396] The system of the present invention analyzes voice data in real time and provides appropriate answers and relevant information based on the user's emotions. This system is intended to support on-site response in security services, and its main components include a terminal, a server, an emotion analysis engine, a speech recognition engine, a natural language processing engine, and a generative AI model.

[1397] Overall system configuration

[1398] 1. Terminal:

[1399] It is a user-operated electronic device that records voice, transcribes it into text, displays sentiment analysis results, and provides notifications.

[1400] Smartphones, tablets, and PCs are examples.

[1401] 2. Speech recognition engine:

[1402] Software that converts audio data into text data. This often utilizes APIs such as Google's Speech-to-Text API.

[1403] 3. Emotion Analysis Engine:

[1404] Software that analyzes a user's emotions from audio or text data and outputs the results as data. IBM Watson Tone Analyzer is used.

[1405] 4. Server:

[1406] It is a computer system that resides in the cloud and performs text data analysis, utilizes sentiment analysis results, and generates and transmits answers and related information.

[1407] The cloud server runs the natural language processing engines (NLTK, SpaCy) and the generative AI model (OpenAI GPT-4).

[1408] 5. Means of communication:

[1409] The Internet Protocol (HTTPS) is used to send and receive data between the terminal and the server.

[1410] System Operation Overview

[1411] When a user makes a voice report at a security site, the device records the voice and converts it into text data in real time using a speech recognition engine. This text data is sent to an emotion analysis engine to identify the user's emotions. The emotion data and text data are sent to a cloud server, where a natural language processing engine analyzes the text data and extracts important keywords and context. Subsequently, a generative AI model generates appropriate answers and relevant information, and the answers are refined based on the emotion data. The optimized answers are sent to the device and displayed and notified to the user.

[1412] Specific example

[1413] For example, suppose a security officer reports, "An alarm has been triggered at the north entrance of the building. Please check the situation." This audio data is recorded on a terminal and converted into text data, "An alarm has been triggered at the north entrance of the building," using Google's Speech-to-Text API. This text data is then analyzed by IBM Watson Tone Analyzer, generating the sentiment data "urgent." The text data and sentiment data sent to the cloud server are analyzed by a natural language processing engine, and a generative AI model generates the optimal response based on the following prompt sentences.

[1414] Example of a prompt:

[1415] "Alarm: An alarm has been triggered at the north entrance of the building."

[1416] Emotion: Urgent

[1417] Possible steps: 1. Check the north entrance with the camera.

[1418] 2. Contact the nearest guard.

[1419] 3. Call the police if necessary.

[1420] Please generate the optimal solution.

[1421] Based on this, the generated answers and related information are sent to the device and notified to the user. The user can then quickly take appropriate action by following the instructions displayed on the device.

[1422] The flow of a specific process in Application Example 2 will be explained using Figure 14.

[1423] Step 1:

[1424] The device records audio reported by the user at the security site. This audio data is temporarily stored in a buffer on the device. The input is the user's voice, and the output is the recorded audio data.

[1425] Step 2:

[1426] The device sends recorded audio data in real time to a speech recognition engine (for example, Google's Speech-to-Text API). This engine converts the audio data into text data. Specifically, the audio data is sent to the API, and text data is returned. The input is the recorded audio data, and the output is the converted text data.

[1427] Step 3:

[1428] The terminal sends transcribed text data to an emotion analysis engine (e.g., IBM Watson Tone Analyzer) to analyze the user's emotions. The emotion analysis engine identifies emotions (such as urgency, anxiety, or joy) from the text data and generates emotion data. The input is transcribed text data, and the output is emotion data.

[1429] Step 4:

[1430] The device sends transcribed text data and sentiment data to a server in the cloud using the HTTPS protocol. Specifically, data packaged in JSON format or similar is sent. The input is text data and sentiment data, and the output is the data sent to the server.

[1431] Step 5:

[1432] The server analyzes the received text and sentiment data. It uses its natural language processing engine (e.g., NLTK or SpaCy) to extract keywords and context from the text data. The input is the text and sentiment data sent to the server, and the output is the extracted keyword and context data.

[1433] Step 6:

[1434] The server uses a generative AI model (e.g., OpenAI GPT-4) to generate answers and related information based on extracted keywords and contextual data. An example of a generated prompt is as follows:

[1435] "Alarm: An alarm has been triggered at the north entrance of the building. Emotion: Emergency Possible actions: 1. Check the north entrance with the camera 2. Contact the nearest guard 3. Call the police if necessary Please generate the best solution."

[1436] The input consists of keywords and contextual data, while the output consists of generated answers and related information.

[1437] Step 7:

[1438] The server adjusts the generated answers and related information based on sentiment data and selects the most appropriate one. This adjustment takes into account the user's current emotional state (e.g., urgent). The input is sentiment data and the generated answers and related information, and the output is the adjusted answers and related information.

[1439] Step 8:

[1440] The server sends the adjusted answers and related information to the terminal. The input is the adjusted answers and related information, and the output is the data sent to the terminal.

[1441] Step 9:

[1442] The device displays the adjusted answers and related information received from the server on its screen and, if necessary, informs the user via voice notification. Specifically, the answers are displayed on the device's screen, and necessary notifications are played through the speaker. The input is the answers and related information sent to the device, and the output is the answers and related information notified to the user.

[1443] The specific processing unit 290 transmits the result of the specific processing to the headset terminal 314. In the headset terminal 314, the control unit 46A causes the speaker 240 and display 343 to output the result of the specific processing. The microphone 238 acquires audio indicating user input for the result of the specific processing. The control unit 46A transmits the audio data indicating user input acquired by the microphone 238 to the data processing unit 12. In the data processing unit 12, the specific processing unit 290 acquires the audio data.

[1444] Data generation model 58 is a type of so-called generative AI (Artificial Intelligence). One example of data generation model 58 is ChatGPT (Internet search<URL: https: / / openai.com / blog / chatgpt> ), Gemini (Internet search) <url: https: gemini.google.com ?hl="ja">Examples of generative AI include the following. The data generation model 58 is obtained by performing deep learning on a neural network. The data generation model 58 is input with prompts containing instructions, and with inference data such as audio data representing speech, text data representing text, and image data representing images. The data generation model 58 infers from the input inference data according to the instructions indicated by the prompts, and outputs the inference results in data formats such as audio data and text data. Here, inference refers to, for example, analysis, classification, prediction, and / or summarization.

[1445] In the above embodiment, an example was given in which specific processing is performed by the data processing device 12, but the technology of this disclosure is not limited thereto, and specific processing may also be performed by the headset terminal 314.

[1446] [Fourth Embodiment]

[1447] Figure 7 shows an example of the configuration of the data processing system 410 according to the fourth embodiment.

[1448] As shown in Figure 7, the data processing system 410 includes a data processing device 12 and a robot 414. An example of the data processing device 12 is a server.

[1449] The data processing device 12 comprises a computer 22, a database 24, and a communication interface 26. The computer 22 is an example of a "computer" related to the technology of this disclosure. The computer 22 comprises a processor 28, RAM 30, and storage 32. The processor 28, RAM 30, and storage 32 are connected to a bus 34. The database 24 and the communication interface 26 are also connected to the bus 34. The communication interface 26 is connected to a network 54. An example of the network 54 is a WAN (Wide Area Network) and / or a LAN (Local Area Network).

[1450] The robot 414 includes a computer 36, a microphone 238, a speaker 240, a camera 42, a communication interface 44, and a controlled object 443. The computer 36 includes a processor 46, RAM 48, and storage 50. The processor 46, RAM 48, and storage 50 are connected to a bus 52. The microphone 238, speaker 240, camera 42, and controlled object 443 are also connected to the bus 52.

[1451] The microphone 238 receives voice signals from the user 20 and receives instructions from the user 20. The microphone 238 captures the voice signals from the user 20, converts the captured voice into audio data, and outputs it to the processor 46. The speaker 240 outputs audio according to the instructions from the processor 46.

[1452] Camera 42 is a small digital camera equipped with an optical system including a lens, aperture, and shutter, and an image sensor such as a CMOS (Complementary Metal-Oxide-Semiconductor) image sensor or a CCD (Charge Coupled Device) image sensor, and captures images of the area around the user 20 (for example, an imaging range defined by a field of view equivalent to the width of a typical healthy person's field of vision).

[1453] Communication interface 44 is connected to network 54. Communication interfaces 44 and 26 are responsible for the exchange of various information between processor 46 and processor 28 via network 54. The exchange of various information between processor 46 and processor 28 using communication interfaces 44 and 26 is performed in a secure manner.

[1454] The controlled object 443 includes a display device, LEDs in the eyes, and motors that drive the arms, hands, and feet. The posture and gestures of the robot 414 are controlled by controlling the motors of the arms, hands, and feet. Some of the robot 414's emotions can be expressed by controlling these motors. Furthermore, the robot 414's facial expressions can also be expressed by controlling the illumination state of the LEDs in its eyes.

[1455] Figure 8 shows an example of the main functions of the data processing device 12 and the robot 414. As shown in Figure 8, the data processing device 12 performs specific processing using the processor 28. The storage 32 stores the specific processing program 56.

[1456] The specific processing program 56 is an example of a "program" relating to the technology of this disclosure. The processor 28 reads the specific processing program 56 from the storage 32 and executes the read specific processing program 56 on the RAM 30. The specific processing is realized by the processor 28 operating as a specific processing unit 290 in accordance with the specific processing program 56 executed on the RAM 30.

[1457] The storage 32 stores the data generation model 58 and the emotion identification model 59. The data generation model 58 and the emotion identification model 59 are used by the identification processing unit 290.

[1458] In robot 414, the processor 46 performs the reception output processing. The storage 50 stores the reception output program 60. The processor 46 reads the reception output program 60 from the storage 50 and executes the read reception output program 60 on the RAM 48. The reception output processing is realized by the processor 46 operating as a control unit 46A according to the reception output program 60 executed on the RAM 48.

[1459] Next, the specific processing performed by the specific processing unit 290 of the data processing device 12 will be described. In the following description, the data processing device 12 will be referred to as the "server" and the robot 414 as the "terminal".

[1460] System Overview

[1461] This invention relates to a system that analyzes telephone and in-person consultations in real time and provides appropriate answers and related information. This system consists of a user's terminal, a cloud-based server, and communication means. The following describes each part of the system in detail.

[1462] Voice input and text conversion

[1463] 1. Users receive consultations via telephone or in person.

[1464] Users will use their own devices (smartphones, PCs, smart glasses, etc.) to provide consultation.

[1465] 2. The device records the user's consultation content as audio.

[1466] The device's application uses the microphone function to capture the audio of the consultation in real time. The audio data is temporarily stored in a buffer.

[1467] 3. The device uses a speech recognition engine to convert speech into text.

[1468] The device sends voice data to a speech recognition engine, which converts it into text data. The converted text is then stored within the application.

[1469] Text data transmission and analysis

[1470] 4. The device sends the transcribed consultation content to the server.

[1471] Text data is sent from the device to a server in the cloud using the HTTPS protocol.

[1472] 5. The server receives the transcribed data and parses it.

[1473] The server's natural language processing (NLP) engine analyzes the text data and extracts important keywords and context.

[1474] Solution generation and search

[1475] 6. The server uses generative AI to generate answers and related information.

[1476] The server generates answers and related information using a database of FAQs related to the problem and a pre-trained generative AI model.

[1477] 7. The server filters the generated answers and related information and selects the appropriate ones.

[1478] From the multiple generated answers, filter and select the most appropriate information.

[1479] Submit and display of answers

[1480] 8. The server sends the answer and related information to the terminal.

[1481] Encode the optimal answers, links, and materials and send them back to your device.

[1482] 9. The device notifies the user of the answer and related information.

[1483] The application on the device displays the received answers and information on the screen and notifies the user by voice as needed.

[1484] Specific example

[1485] For customer support

[1486] 1. A user contacts us because the product is not working.

[1487] 2. The device records the audio and transcribes the voice message, "The product isn't working, what should I do?"

[1488] 3. The device sends the converted data to the server.

[1489] 4. The server analyzes the text data and extracts the keywords "product" and "not working".

[1490] 5. The server refers to the relevant FAQ database and generates answers, reset instructions, and support video links.

[1491] 6. Filter the information generated by the server and send it to the terminal.

[1492] 7. The device will display instructions on how to reset the device and a link to a support video on the screen, and will announce a voice message saying, "Here's how to reset the device."

[1493] In the case of sales

[1494] 1. The user receives an inquiry about the details of the new product.

[1495] 2. The device records the audio and transcribes the phrase "Please tell me more about the new product" into text.

[1496] 3. The device sends the converted data to the server.

[1497] 4. The server analyzes the text data and extracts the keywords "new product" and "details".

[1498] 5. The server generates related sales materials, product brochures, and links to explanatory videos.

[1499] 6. Filter the information generated by the server and send it to the terminal.

[1500] 7. The device displays a link to a brochure and product features on the screen, and announces with a voice message, "Here are the details of the new product."

[1501] Thus, the system of the present invention enables consistent, high-quality answers and information provision, regardless of the user's skill level. It offers high practicality in the field and contributes to improved work efficiency.

[1502] The following describes the processing flow.

[1503] Step 1:

[1504] Users receive consultations via phone or in person.

[1505] Users will use their own devices (smartphones, PCs, smart glasses, etc.) to provide consultation.

[1506] Step 2:

[1507] The device records the user's conversation audio.

[1508] The device's application uses the microphone function to capture the audio of the consultation in real time. The audio data is temporarily stored in a buffer.

[1509] Step 3:

[1510] The device uses a speech recognition engine to convert speech into text.

[1511] The device sends voice data to a speech recognition engine, which converts it into text data. The converted text is then stored within the application.

[1512] Step 4:

[1513] The terminal sends the transcribed consultation content to the server.

[1514] Text data is sent from the device to a server in the cloud using the HTTPS protocol.

[1515] Step 5:

[1516] The server receives the transcribed data and parses it.

[1517] The server's natural language processing (NLP) engine analyzes the text data and extracts important keywords and context.

[1518] Step 6:

[1519] The server uses generative AI to generate answers and related information.

[1520] The server generates answers and related information using a database of FAQs related to the problem and a pre-trained generative AI model.

[1521] Step 7:

[1522] The server filters the generated answers and related information and selects the appropriate ones.

[1523] From the multiple generated answers, filter and select the most appropriate information.

[1524] Step 8:

[1525] The server sends the answer and related information to the terminal.

[1526] Encode the optimal answers, links, and materials and send them back to your device.

[1527] Step 9:

[1528] The device notifies the user of the answer and related information.

[1529] The application on the device displays the received answers and information on the screen and notifies the user by voice as needed.

[1530] (Example 1)

[1531] Next, we will describe Example 1. In the following description, the data processing device 12 will be referred to as the "server" and the robot 414 as the "terminal".

[1532] In modern call centers and customer support, there is a demand for providing quick and accurate answers to user inquiries. However, conventional systems often suffer from low accuracy in voice-to-text conversion, and answer generation is frequently done manually, leading to delays and inconsistent quality. This has resulted in decreased customer satisfaction and reduced operational efficiency. This invention aims to solve these problems, streamline the process from voice input to answer provision, and achieve consistent, high-quality information delivery.

[1533] The identification process performed by the identification processing unit 290 of the data processing device 12 in Example 1 is realized by the following means.

[1534] In this invention, the server includes: means for the terminal to record the user's voice using communication means and transcribe the voice in real time; means for the terminal to transmit the transcribed data to the server; means for the server to analyze the transcribed data, extract important keywords and context, and generate answers and related information based on them; means for the server to filter the generated answers and related information and transmit it to the terminal; and means for the terminal to display the answers and related information on the screen and notify the user of important information by voice. This automates the process from voice input to answer provision, enabling the rapid, consistent, and high-quality provision of information.

[1535] A "terminal" is an information processing device used by a user, which acquires and processes data, including voice input. Specifically, this refers to smartphones, personal computers, smart glasses, etc.

[1536] "Communication means" refers to technical means for sending and receiving data, including the internet, Wi-Fi, Bluetooth, etc.

[1537] "Recording audio" means capturing a user's speech as digital audio data using a microphone and storing it.

[1538] "Real-time transcription of speech" means instantly analyzing speech data and converting it into corresponding text data. This is typically done using a speech recognition engine.

[1539] "Transcribed data" refers to text data obtained after converting voice input into written information.

[1540] A "server" is a high-performance information processing device that operates on the cloud and has functions such as data analysis, storage, and transmission.

[1541] "Analysis" is the process of breaking down received data and performing operations to understand it.

[1542] "Important keywords and context" refer to words and phrases that are particularly meaningful from the input data, as well as the relationships between them.

[1543] "Answers and related information" refers to information that allows for appropriate responses to user questions and inquiries. This includes generated text, links, and documents.

[1544] "Generating" means creating new information based on input data.

[1545] "Filtering" is the process of selecting the most appropriate answer or piece of information from multiple generated responses.

[1546] "Displaying on the screen" means making information visible as text or images on the device's display.

[1547] "Notifying by voice" means conveying text data to the user as voice using speech synthesis.

[1548] Modes for carrying out the invention

[1549] System Overview

[1550] This invention relates to a system that analyzes telephone and in-person consultations in real time and provides appropriate answers and related information. This system consists of a user's terminal, a cloud-based server, and communication means. The following describes each component of the system in detail.

[1551] Voice input and text conversion

[1552] The device records audio when the user has a consultation in person or over the phone. The device uses an information processing device such as a smartphone, PC, or smart glasses. The application on this device uses the microphone function to acquire the user's voice in real time and temporarily stores the audio data in memory.

[1553] Next, the device uses a speech recognition engine to convert the speech into text. For example, it might use the Google Cloud Speech-to-Text API to convert the acquired speech data into text data. This converted text data is then stored within the application.

[1554] Text data transmission and analysis

[1555] The terminal sends the transcribed consultation content to the server. The HTTPS protocol is used for this transmission, ensuring the secure transfer of data.

[1556] The server analyzes the received text data. A natural language processing (NLP) engine is used for the analysis. For example, the Google Cloud Natural Language API is used to extract important keywords and context from the text data.

[1557] Solution generation and search

[1558] The server uses a generative AI model to generate answers and related information. For example, a generative AI model like OpenAI's GPT-3 is used. The server also refers to a FAQ database related to the problem to generate answers and related information.

[1559] The server filters the multiple generated answers and related information, selecting the most appropriate one. This ensures that the necessary information is provided efficiently.

[1560] Submit and display of answers

[1561] The selected answers and related information are sent from the server to the terminal. The HTTPS protocol is used for this transmission, ensuring the secure exchange of information.

[1562] The device notifies the user of the received answers and related information. The application displays this information on the screen and, if necessary, informs the user via voice.

[1563] Specific example

[1564] Customer support scenarios

[1565] 1. The user asks, "The product isn't working, what should I do?"

[1566] 2. The device records this audio and transcribes it into text.

[1567] 3. The device sends the converted data to the server.

[1568] 4. The server analyzes the text data and extracts the keywords "product" and "not working".

[1569] 5. The server generates answers using the relevant FAQ database and generative AI models.

[1570] 6. The server filters the generated answers and selects the most appropriate information.

[1571] 7. The server sends the selected answer to the terminal.

[1572] 8. The device will announce a voice message saying, "Here's how to reset," and display the reset instructions on the screen.

[1573] Sales scenario

[1574] 1. The user asks, "Please tell me more about the new product."

[1575] 2. The device records this audio and transcribes it into text.

[1576] 3. The device sends the converted data to the server.

[1577] 4. The server analyzes the text data and extracts the keywords "new product" and "details".

[1578] 5. The server generates related sales materials, product brochures, and links to explanatory videos.

[1579] 6. The server filters the generated information and selects the most appropriate information.

[1580] 7. The server sends the selected information to the terminal.

[1581] 8. The device will announce, "Here are the details of the new product," and display a link to the brochure on the screen.

[1582] The flow of the specific processing in Example 1 will be explained using Figure 11.

[1583] Step 1:

[1584] The user initiates a consultation via voice. The user launches a dedicated application on their device, such as a smartphone or PC.

[1585] Specific action: The user asks a question such as, "Please tell me more about the new product."

[1586] Input: User's voice.

[1587] Output: Audio data is input to the terminal.

[1588] Step 2:

[1589] The device records the user's voice. The device's application uses the microphone function to capture the user's speech in real time and stores the audio data in memory.

[1590] Specific operation: The application acquires audio data from the device's microphone and temporarily stores it in a buffer.

[1591] Input: User's voice data.

[1592] Output: Audio data stored in the buffer.

[1593] Step 3:

[1594] The device uses a speech recognition engine to convert speech into text. For example, it sends speech data to the Google Cloud Speech-to-Text API and converts it into text data.

[1595] Specific operation: The speech recognition engine analyzes the audio data and generates text data.

[1596] Input: Audio data.

[1597] Output: Text data.

[1598] Step 4:

[1599] The device sends the transcribed data to the server. The device uses the HTTPS protocol to securely send the text data to the server in the cloud.

[1600] Specific operation: Encoded text data is sent via the HTTPS protocol.

[1601] Input: Text data.

[1602] Output: Text data sent to the server.

[1603] Step 5:

[1604] The server analyzes the received text data. An NLP engine is used to extract important keywords and context from the text. This process includes, for example, the Google Cloud Natural Language API.

[1605] Specific operation: The server's NLP engine analyzes the text data and extracts keywords such as "new product" and "details."

[1606] Input: Text data.

[1607] Output: Analyzed keywords and contextual information.

[1608] Step 6:

[1609] The server generates answers and related information using a generative AI model. It uses a generative AI model (for example, OpenAI's GPT-3) to generate appropriate answers and information based on the consultation content.

[1610] Specific operation: Input prompts into the generative AI model and obtain the generated answer.

[1611] Input: Analyzed keywords and contextual information.

[1612] Output: Generated answers and related information.

[1613] Step 7:

[1614] The server filters the generated answers and related information. It selects the most appropriate answer from the multiple answers generated. Criteria for evaluating reliability and relevance are applied to this process.

[1615] Specific operation: The filtering algorithm selects the optimal solution.

[1616] Input: Multiple generated answers.

[1617] Output: Optimal answers and related information.

[1618] Step 8:

[1619] The server sends the answer and related information to the terminal. The selected answer is re-encoded and sent to the terminal using the HTTPS protocol.

[1620] Specific operation: Encoded information is sent via the HTTPS protocol.

[1621] Input: The best answer or related information.

[1622] Output: Answers and related information sent to the terminal.

[1623] Step 9:

[1624] The device notifies the user of the answer and related information. The device's application displays the received information on the screen and, if necessary, communicates it to the user verbally.

[1625] Specific action: The application displays information on the screen and announces with a voice message, "Here are the details of the new product."

[1626] Input: Answers and related information sent to the device.

[1627] Output: Information displayed or notified to the user.

[1628] (Application Example 1)

[1629] Next, we will explain Application Example 1. In the following explanation, the data processing device 12 will be referred to as the "server" and the robot 414 as the "terminal".

[1630] Conventional customer support systems have struggled to provide timely and appropriate answers to user inquiries. This often resulted in delays in response times and a decline in the quality of answers. Furthermore, limitations in communication methods and devices posed a challenge in terms of usability. This invention aims to solve these problems and enable rapid and high-quality customer support.

[1631] The specific processing performed by the specific processing unit 290 of the data processing device 12 in Application Example 1 is realized by the following means.

[1632] In this invention, the server includes: [means for the terminal to record the user's voice using communication means and transcribe the voice in real time; [means for the terminal to transmit the transcribed data to the server]; [means for the server to analyze the transcribed data, extract important keywords and context, and generate answers and related information based thereon]; [means for the server to transmit the generated answers and related information to the terminal]; [means for the terminal to display the answers and related information on a screen and notify the user of important information by voice]; and [a portable information processing device on which an application that accepts the user's voice input is executed]. This enables rapid and high-quality customer support by transcribing voice from the terminal in real time and generating and displaying appropriate answers.

[1633] A "terminal" is a computer device used by a user that records audio and transmits the transcribed data to a server.

[1634] "Communication methods" refer to the technologies and protocols used for sending and receiving data, and are the means used to exchange information between a server and a terminal.

[1635] "Recording audio" refers to collecting the user's spoken content as digital data using a microphone.

[1636] "Real-time transcription" refers to the process of instantly converting acquired audio data into text data.

[1637] "Text-based data" refers to information that has been converted into text format by a speech recognition engine.

[1638] A "server" is a computer system located in the cloud that has the function of analyzing digitized data and generating answers.

[1639] "Analyzing" refers to the process of understanding received data using a natural language processing engine and extracting important keywords and context.

[1640] "Extracting important keywords and context" refers to the process of taking meaningful words and vocabulary from text data and combining them based on their context.

[1641] "Generating related information" refers to constructing answers or additional information based on extracted keywords and context.

[1642] "Displaying on the screen" refers to providing answers or information in a visually apparent form on the device's display.

[1643] "Notifying the user via voice" refers to the process of communicating the generated answer to the user using speech synthesis technology.

[1644] A "portable information processing device" refers to an electronic device that is portable and has the functionality to accept voice input.

[1645] This invention relates to a system that transcribes a user's voice into text in real time and provides answers and related information based on that transcription. This system mainly consists of a terminal, communication means, and a server. Specific embodiments of this system are described below.

[1646] Hardware and software

[1647] The terminals used are portable information processing devices such as smartphones and smart glasses. These devices are equipped with microphones and speech recognition engines to record and transcribe user voice input in real time. Furthermore, applications running on the terminals also have the functionality to send the text data to a server.

[1648] The server resides in the cloud and uses a natural language processing (NLP) engine and generative AI models to generate answers and related information. The server analyzes the text data sent by the user and extracts important keywords and context.

[1649] The communication method used is an internet connection for sending and receiving data between the terminal and the server. The communication protocol used is HTTPS to ensure data security and efficient transfer.

[1650] Data processing and data calculation

[1651] The device uses a speech recognition engine (for example, Google Speech Recognition API) to convert speech into text data in real time. This text data is then sent to a server in the cloud via the application.

[1652] The server uses a natural language processing engine (such as SpaCy or NLTK) to analyze the text data and extract important keywords and context. Then, it uses a generative AI model (such as OpenAI's GPT-3) to generate the optimal answer and relevant information.

[1653] The generated answers are sent from the server to the terminal, where the answers and related information are displayed on the screen. Important information is also communicated to the user via voice using speech synthesis technology.

[1654] Specific example

[1655] For example, consider a customer support scenario where a user asks, "My product isn't working, what should I do?" The device records the user's voice and converts it into text data using a speech recognition engine. This text data is then sent to a server, where a natural language processing engine extracts the keywords "product" and "not working." A generative AI model then consults a relevant FAQ database to generate an answer that includes instructions on how to reset the product and links to support videos. Finally, the generated answer is sent back to the device, displayed on the screen, and announced via voice.

[1656] Example of a prompt

[1657] Prompt: "User Inquiry: [User Question]"

[1658] Example: "User inquiry: What should I do if the product doesn't work?"

[1659] Thus, the system of the present invention allows users to receive prompt and high-quality customer support. Furthermore, support staff can provide consistent and high-quality answers to user inquiries.

[1660] The flow of a specific process in Application Example 1 will be explained using Figure 12.

[1661] Step 1:

[1662] The user asks a question into the device (smartphone or smart glasses). The device's microphone records the audio. The input is the user's voice, and the output is the recorded audio data. Specifically, the microphone captures the user's speech and temporarily stores it as a buffer within the application.

[1663] Step 2:

[1664] The device sends recorded audio data to a speech recognition engine in real time, where it is converted into text data. The input is recorded audio data, and the output is transcribed text data. Specifically, the SpeechRecognition library is used to convert the audio data into text, which is then stored in a variable within the application.

[1665] Step 3:

[1666] The device sends characterized text data to a server in the cloud using the HTTPS protocol. The input is character data, and the output is the state as it has been sent to the server. Specifically, the Requests library is used to send the character data as a POST request to the server's API endpoint.

[1667] Step 4:

[1668] The server receives text data and analyzes it using a natural language processing engine (e.g., SpaCy). The input is text data, and the output is data with key keywords and context extracted. Specifically, the NLP engine on the cloud server analyzes the text data and identifies important words and phrases within the text.

[1669] Step 5:

[1670] The server uses a generative AI model (e.g., GPT-3) to generate answers and related information based on the analysis results. The input is keywords and contextual information from the analysis results, and the output is the generated answers and related information. Specifically, it generates a prompt sentence, inputs it into the generative AI model, and generates an appropriate answer.

[1671] Step 6:

[1672] The server sends the generated answer and related information back to the terminal using the HTTPS protocol. The input is the generated answer and related information, and the output is the state as it has been sent to the terminal. Specifically, the server encodes the appropriate answer and sends it in a format that the terminal can understand.

[1673] Step 7:

[1674] The device displays answers and related information on its screen and notifies the user of important information via voice. Input consists of answers and related information returned from the server, while output is visual and auditory notification to the user. Specifically, the application displays the answers on the device's display and uses speech synthesis technology to convey important points verbally.

[1675] By following these steps, the system will be able to provide real-time answers to user inquiries.

[1676] Furthermore, an emotion engine that estimates the user's emotions may be incorporated. That is, the identification processing unit 290 may use the emotion identification model 59 to estimate the user's emotions and perform identification processing using the user's emotions.

[1677] System Overview

[1678] This invention relates to a system that analyzes telephone and in-person consultations in real time and provides appropriate answers and related information. Furthermore, it is a system that achieves higher user satisfaction by analyzing the user's emotions and adjusting the content of the answers and information provided based on those emotions. This system consists of the user's terminal, a server on the cloud, an emotion engine, and communication means. The following describes each part of the system in detail.

[1679] Voice input and text conversion

[1680] 1. Users receive consultations via telephone or in person.

[1681] Users will use their own devices (smartphones, PCs, smart glasses, etc.) to provide consultation.

[1682] 2. The device records the user's consultation content as audio.

[1683] The device's application uses the microphone function to capture the audio of the consultation in real time. The audio data is temporarily stored in a buffer.

[1684] 3. The device uses a speech recognition engine to convert speech into text.

[1685] The device sends voice data to a speech recognition engine, which converts it into text data. The converted text is then stored within the application.

[1686] Emotion analysis

[1687] 4. The device sends the recorded audio data to the emotion engine, which analyzes the user's emotions.

[1688] The emotion engine analyzes the user's emotions (e.g., anger, anxiety, joy) from voice data and generates emotion data.

[1689] Text data transmission and analysis

[1690] 5. The device sends the transcribed consultation content and emotional data to the server.

[1691] Text data and sentiment data are sent from the device to a server in the cloud using the HTTPS protocol.

[1692] 6. The server receives the transcribed data and parses it.

[1693] The server's natural language processing (NLP) engine analyzes the text data and extracts important keywords and context.

[1694] 7. The server uses generative AI to generate answers and related information.

[1695] The server generates answers and related information using a database of FAQs related to the problem and a pre-trained generative AI model.

[1696] Adjust and submit your answer.

[1697] 8. The server uses sentiment data to adjust and filter answers and related information.

[1698] The server considers the user's sentiment data received from the sentiment engine to filter and select the most appropriate answers and relevant information.

[1699] 9. The server sends the answer and related information to the terminal.

[1700] Encode the optimal answers, links, and materials and send them back to your device.

[1701] Display and notification of answers

[1702] 10. The device notifies the user of the answer and related information.

[1703] The application on the device displays the received answers and information on the screen and notifies the user by voice as needed. Based on sentiment data, the displayed content and the tone of voice notifications are adjusted.

[1704] Specific example

[1705] For customer support

[1706] 1. A user contacts us because the product is not working.

[1707] 2. The device records the audio and transcribes the voice message, "The product isn't working, what should I do?"

[1708] 3. The device sends voice data to the emotion engine, which analyzes whether the user is feeling anxious.

[1709] 4. The device sends the transcribed data and sentiment data to the server.

[1710] 5. The server analyzes the text data and extracts the keywords "product" and "not working".

[1711] 6. The server refers to the relevant FAQ database and generates answers, reset instructions, and support video links.

[1712] 7. The server considers emotional data and selects a more polite and reassuring answer.

[1713] 8. The server sends the selected information to the terminal.

[1714] 9. The device displays instructions on how to reset the device and a link to a support video on the screen, and provides a voice notification in a gentle tone saying, "Here's how to reset the device."

[1715] In the case of sales

[1716] 1. The user receives an inquiry about the details of the new product.

[1717] 2. The device records the audio and transcribes the phrase "Please tell me more about the new product" into text.

[1718] 3. The device sends voice data to the emotion engine, which analyzes whether the user is interested.

[1719] 4. The device sends the transcribed data and sentiment data to the server.

[1720] 5. The server analyzes the text data and extracts the keywords "new product" and "details".

[1721] 6. The server generates related sales materials, product brochures, and links to explanatory videos.

[1722] 7. The server considers sentiment data and selects an explanation that will pique interest.

[1723] 8. The server sends the selected information to the terminal.

[1724] 9. The device displays a link to a brochure and product features on the screen, and provides a lively voice notification saying, "Here are the details of the new product."

[1725] Thus, the system of the present invention enables consistent, high-quality answers and information provision based on the user's emotions, is highly practical in the field, and contributes to improving work efficiency and user satisfaction.

[1726] The following describes the processing flow.

[1727] Step 1:

[1728] Users receive consultations via phone or in person.

[1729] Users will use their own devices (smartphones, PCs, smart glasses, etc.) to provide consultation.

[1730] Step 2:

[1731] The device records the user's conversation audio.

[1732] The device's application uses the microphone function to capture the audio of the consultation in real time. The audio data is temporarily stored in a buffer.

[1733] Step 3:

[1734] The device sends voice data to an emotion engine, which then analyzes the user's emotions.

[1735] The device sends voice data to an emotion engine (for example, a voice analysis API), which analyzes the user's emotions (anger, anxiety, joy, etc.) from the voice. The analysis results are stored as emotion data.

[1736] Step 4:

[1737] The device uses a speech recognition engine to convert speech into text.

[1738] The device sends the voice data to a speech recognition engine (e.g., a speech recognition API) and converts it into text data. The converted text is then stored within the application.

[1739] Step 5:

[1740] The device sends the transcribed consultation content and emotional data to the server.

[1741] Text data and sentiment data are sent from the device to a server in the cloud using the HTTPS protocol.

[1742] Step 6:

[1743] The server receives the transcribed data and parses it.

[1744] The server's natural language processing (NLP) engine analyzes the text data and extracts important keywords and context. The analysis results are saved as analysis data.

[1745] Step 7:

[1746] The server uses generative AI to generate answers and related information.

[1747] Based on the analysis data, the server generates answers and related information using a database of FAQs relevant to the problem and a pre-trained generative AI model. The generated information is temporarily stored.

[1748] Step 8:

[1749] The server uses sentiment data to adjust and filter answers and related information.

[1750] The server considers the user's emotional data received from the emotion engine and filters and selects the most appropriate answers and relevant information. The selected information is then prepared as data to be sent.

[1751] Step 9:

[1752] The server sends the answer and related information to the terminal.

[1753] The transmitted data is encoded and sent back to the device. This return transmission takes place via cloud communication.

[1754] Step 10:

[1755] The device notifies the user of the answer and related information.

[1756] The application on the device displays the received answers and information on the screen and notifies the user by voice as needed. Based on sentiment data, the displayed content and the tone of voice notifications are adjusted.

[1757] (Example 2)

[1758] Next, we will describe Example 2. In the following description, the data processing device 12 will be referred to as the "server" and the robot 414 as the "terminal".

[1759] Conventional consultation systems only analyze data transcribed from user speech, failing to consider user emotions and making it difficult to provide appropriate and personalized answers. Furthermore, they cannot provide notifications in a tone adjusted to the user's emotions, thus failing to adequately improve user satisfaction. Therefore, there was a need for a more advanced consultation system that analyzes user emotions in real time and provides emotion-based answers.

[1760] The identification process performed by the identification processing unit 290 of the data processing device 12 in Example 2 is realized by the following means.

[1761] In this invention, the server includes means for the terminal to transmit text data and sentiment data to the server, means for the server to analyze the text data, extract important keywords and context, and adjust the answer and related information using the sentiment data, and means for the server to transmit the generated answer and related information to the terminal. This enables the provision of appropriate answers based on the user's emotions and notifications in an adjusted tone.

[1762] A "terminal" is an information processing device used by users to input voice and send and receive data.

[1763] "Communication methods" refer to network technologies used to send and receive data between terminals and servers.

[1764] "Recording audio" means acquiring and saving the user's speech as digital audio data.

[1765] "Real-time transcription" refers to the process of instantly converting recorded audio data into text data.

[1766] "Emotional data" refers to information that indicates the user's emotional state, analyzed from voice data.

[1767] A "server" is a remote computer system that analyzes data received from a terminal and generates appropriate answers or information.

[1768] "Analyzing text data" means processing the text data received by the server and extracting important keywords and context.

[1769] "Key keywords" are words or phrases that are particularly important for understanding the user's inquiry.

[1770] "Extracting context" means analyzing the relationships between words in order to understand the meaning and intent of the entire text.

[1771] "Generating answers and related information" means that the server automatically creates appropriate answers and reference information based on the user's inquiry.

[1772] A "generative AI model" is an artificial intelligence system that uses technologies such as natural language processing to automatically generate text from input data.

[1773] "Adjusted tone" refers to appropriately modifying the speech style and nuances of voice notifications based on the user's emotional data.

[1774] "Displaying on the screen" means visually presenting answers or information on the device's display.

[1775] "Notifying by voice" means presenting answers or information to the user audibly using speech synthesis technology.

[1776] This invention is a system that analyzes user inquiries in real time and provides appropriate answers and related information. This system also analyzes the user's emotions and adjusts the content of the answers and information provided based on the emotional data, thereby achieving higher user satisfaction.

[1777] The main components of the system consist of the user's terminal, a server in the cloud, an emotion engine, and communication methods. Details of each component are shown below.

[1778] terminal

[1779] This is an information processing device that allows users to input voice data and send and receive data. Specifically, this includes smartphones, PCs, and smart glasses. The device incorporates a voice recording function, a speech recognition engine to convert voice data into text, and communication means for sending and receiving the results of sentiment analysis.

[1780] server

[1781] This is a cloud-based computer system that analyzes text and sentiment data received from terminals. The server uses a natural language processing (NLP) engine to analyze text data and generates answers and related information using generative AI models (e.g., OpenAI GPT-3). It also plays a role in adjusting answers by taking sentiment data into consideration.

[1782] means of communication

[1783] This is a network technology for sending and receiving data between a terminal and a server. Specifically, it involves data transmission using the HTTPS protocol.

[1784] Examples

[1785] For customer support

[1786] 1. The user asks, "The product isn't working, what should I do?"

[1787] 2. The device records the audio and uses a speech recognition engine (e.g., Google Speech-to-Text API) to transcribe it. The generated text data will be "The product isn't working, what should I do?"

[1788] 3. The device sends voice data to an emotion engine (e.g., IBM Watson Tone Analyzer) to obtain emotion data indicating anxiety.

[1789] 4. The device sends text data and sentiment data to the server.

[1790] 5. The server uses a natural language processing engine (e.g., BERT model) to analyze the text data and extract key keywords such as "product" and "not working".

[1791] 6. The server uses a generative AI model (e.g., OpenAI GPT-3) to generate "Product Reset Instructions" and "Support Video Links".

[1792] 7. Take into account the server's emotional data indicating anxiety, and adjust the response to be polite and reassuring.

[1793] 8. The server sends the adjusted answer to the terminal.

[1794] 9. The device displays "Here's how to reset it" on the screen and also provides a voice notification in a gentle tone.

[1795] Examples of prompt statements

[1796] For customer support

[1797] Prompt message:

[1798] User: "The product isn't working, what should I do?" Sentiment analysis result: Anxiety.

[1799] Generative AI models:

[1800] 1. How to reset the product

[1801] 2. Support video links

[1802] 3. Related answers from the FAQ

[1803] Taking these factors into consideration, please generate your answer in a gentle tone.

[1804] The above details the embodiment of the system of the present invention. This system provides high-quality answers and information based on the user's emotions, thereby improving work efficiency and user satisfaction.

[1805] The flow of the specific processing in Example 2 will be explained using Figure 13.

[1806] Step 1:

[1807] The user initiates a consultation.

[1808] Input: User's voice (e.g., "The product isn't working, what should I do?").

[1809] Specific action: The user speaks their question into the device's microphone.

[1810] Output: Raw audio data.

[1811] Step 2:

[1812] The device records audio.

[1813] Input: Raw audio data.

[1814] Specific operation: The device uses its microphone function to record the user's voice in real time and temporarily stores it in a buffer as audio data.

[1815] Output: Audio data stored in the buffer.

[1816] Step 3:

[1817] The device converts speech into text.

[1818] Input: Audio data stored in the buffer.

[1819] Specific operation: The device sends audio data to a speech recognition engine (e.g., Google Speech-to-Text API), and the audio data is converted into text data.

[1820] Output: Text data (Example: "The product isn't working, what should I do?").

[1821] Step 4:

[1822] The device performs emotion analysis.

[1823] Input: Audio data and text data.

[1824] Specific operation: The device sends voice data to an emotion engine (e.g., IBM Watson Tone Analyzer), which analyzes the emotion and obtains emotion data.

[1825] Output: Emotional data (e.g., anxiety).

[1826] Step 5:

[1827] The device sends data to the server.

[1828] Input: Text data and sentiment data.

[1829] Specific operation: The device sends text data and sentiment data to a server in the cloud using the HTTPS protocol.

[1830] Output: Text data and sentiment data sent to the server.

[1831] Step 6:

[1832] The server analyzes the text data.

[1833] Input: Character data sent to a server in the cloud.

[1834] Specific operation: The server uses a natural language processing (NLP) engine (e.g., BERT model) to analyze the text data and extract important keywords and context.

[1835] Output: Extracted keywords (e.g., "product", "not working").

[1836] Step 7:

[1837] The server generates the answers and related information.

[1838] Input: Extracted keywords.

[1839] Specific operation: The server uses a generative AI model (e.g., OpenAI GPT-3) to generate answers and related information.

[1840] Output: Generated answers and related information (e.g., "Product Reset Instructions," "Support Video Link").

[1841] Step 8:

[1842] The server adjusts the answers based on sentiment data.

[1843] Input: Sentimental data and generated answers or related information.

[1844] Specific operation: The server considers the sentiment data received from the sentiment engine and adjusts the response to have the most appropriate tone and content.

[1845] Output: Adjusted answer (e.g., a polite and reassuring answer).

[1846] Step 9:

[1847] The server sends the answer to the terminal.

[1848] Input: Adjusted answers and related information.

[1849] Specific operation: The server encodes the adjusted answer and related information and sends it to the terminal.

[1850] Output: The adjusted answers and related information sent to the terminal.

[1851] Step 10:

[1852] The device will display and notify you of the answer.

[1853] Input: Adjusted answers and related information sent from the server.

[1854] Specific actions: The device displays answers and information on the screen and provides voice notifications in a tone adjusted based on emotional data.

[1855] Output: The answers and information notified to the user.

[1856] (Application Example 2)

[1857] Next, we will explain application example 2. In the following explanation, the data processing device 12 will be referred to as the "server" and the robot 414 as the "terminal".

[1858] Conventional consultation systems using speech recognition have difficulty providing answers and information that take into account the user's emotions, and in particular, providing appropriate response instructions in real time has been difficult in security service settings. As a result, the quality of responses by field staff has declined, hindering security operations that require rapid response. The present invention aims to solve these problems and provide a system that improves the quality of responses by field staff by analyzing the user's emotions and providing appropriate response instructions in real time.

[1859] The specific processing performed by the specific processing unit 290 of the data processing device 12 in Application Example 2 is realized by the following means.

[1860] In this invention, the server includes means for analyzing transcribed data, extracting important keywords and context, and generating answers and related information based on them; means for analyzing the user's emotions and adjusting the answers and related information based on the analysis results; and means for transmitting the generated answers and related information to the terminal. This makes it possible to provide appropriate response instructions and information based on the user's emotions.

[1861] A "terminal" is an electronic device operated by a user that performs functions such as recording voice, transcribing it into text, displaying sentiment analysis results, and sending notifications.

[1862] "Communication method" refers to technology that uses Internet protocols (e.g., HTTPS) to send and receive data between a terminal and a server.

[1863] A "speech recognition engine" is software or a system that converts speech data into text data.

[1864] An "emotion analysis engine" is software or a system that analyzes a user's emotions from voice data or text data and outputs the results as data.

[1865] A "server" is a computer system that resides in the cloud and performs tasks such as analyzing text data, utilizing sentiment analysis results, and generating and transmitting answers and related information.

[1866] A "natural language processing engine" is software or a system that analyzes text data, understands keywords and context, and generates appropriate answers and related information.

[1867] "Generative AI" is artificial intelligence that uses a pre-trained model to generate answers and information in natural language based on input data.

[1868] "Answers and related information" refers to the answers, related data, links, and materials provided in response to a user's query.

[1869] "Emotional data" refers to information generated by an emotion analysis engine that indicates the user's emotional state.

[1870] The system of the present invention analyzes voice data in real time and provides appropriate answers and relevant information based on the user's emotions. This system is intended to support on-site response in security services, and its main components include a terminal, a server, an emotion analysis engine, a speech recognition engine, a natural language processing engine, and a generative AI model.

[1871] Overall system configuration

[1872] 1. Terminal:

[1873] It is a user-operated electronic device that records voice, transcribes it into text, displays sentiment analysis results, and provides notifications.

[1874] Smartphones, tablets, and PCs are examples.

[1875] 2. Speech recognition engine:

[1876] Software that converts audio data into text data. This often utilizes APIs such as Google's Speech-to-Text API.

[1877] 3. Emotion Analysis Engine:

[1878] Software that analyzes a user's emotions from audio or text data and outputs the results as data. IBM Watson Tone Analyzer is used.

[1879] 4. Server:

[1880] It is a computer system that resides in the cloud and performs text data analysis, utilizes sentiment analysis results, and generates and transmits answers and related information.

[1881] The cloud server runs the natural language processing engines (NLTK, SpaCy) and the generative AI model (OpenAI GPT-4).

[1882] 5. Means of communication:

[1883] The Internet Protocol (HTTPS) is used to send and receive data between the terminal and the server.

[1884] System Operation Overview

[1885] When a user makes a voice report at a security site, the device records the voice and converts it into text data in real time using a speech recognition engine. This text data is sent to an emotion analysis engine to identify the user's emotions. The emotion data and text data are sent to a cloud server, where a natural language processing engine analyzes the text data and extracts important keywords and context. Subsequently, a generative AI model generates appropriate answers and relevant information, and the answers are refined based on the emotion data. The optimized answers are sent to the device and displayed and notified to the user.

[1886] Specific example

[1887] For example, suppose a security officer reports, "An alarm has been triggered at the north entrance of the building. Please check the situation." This audio data is recorded on a terminal and converted into text data, "An alarm has been triggered at the north entrance of the building," using Google's Speech-to-Text API. This text data is then analyzed by IBM Watson Tone Analyzer, generating the sentiment data "urgent." The text data and sentiment data sent to the cloud server are analyzed by a natural language processing engine, and a generative AI model generates the optimal response based on the following prompt sentences.

[1888] Example of a prompt:

[1889] "Alarm: An alarm has been triggered at the north entrance of the building."

[1890] Emotion: Urgent

[1891] Possible steps: 1. Check the north entrance with the camera.

[1892] 2. Contact the nearest guard.

[1893] 3. Call the police if necessary.

[1894] Please generate the optimal solution.

[1895] Based on this, the generated answers and related information are sent to the device and notified to the user. The user can then quickly take appropriate action by following the instructions displayed on the device.

[1896] The flow of a specific process in Application Example 2 will be explained using Figure 14.

[1897] Step 1:

[1898] The device records audio reported by the user at the security site. This audio data is temporarily stored in a buffer on the device. The input is the user's voice, and the output is the recorded audio data.

[1899] Step 2:

[1900] The device sends recorded audio data in real time to a speech recognition engine (for example, Google's Speech-to-Text API). This engine converts the audio data into text data. Specifically, the audio data is sent to the API, and text data is returned. The input is the recorded audio data, and the output is the converted text data.

[1901] Step 3:

[1902] The terminal sends transcribed text data to an emotion analysis engine (e.g., IBM Watson Tone Analyzer) to analyze the user's emotions. The emotion analysis engine identifies emotions (such as urgency, anxiety, or joy) from the text data and generates emotion data. The input is transcribed text data, and the output is emotion data.

[1903] Step 4:

[1904] The device sends transcribed text data and sentiment data to a server in the cloud using the HTTPS protocol. Specifically, data packaged in JSON format or similar is sent. The input is text data and sentiment data, and the output is the data sent to the server.

[1905] Step 5:

[1906] The server analyzes the received text and sentiment data. It uses its natural language processing engine (e.g., NLTK or SpaCy) to extract keywords and context from the text data. The input is the text and sentiment data sent to the server, and the output is the extracted keyword and context data.

[1907] Step 6:

[1908] The server uses a generative AI model (e.g., OpenAI GPT-4) to generate answers and related information based on extracted keywords and contextual data. An example of a generated prompt is as follows:

[1909] "Alarm: An alarm has been triggered at the north entrance of the building. Emotion: Emergency Possible actions: 1. Check the north entrance with the camera 2. Contact the nearest guard 3. Call the police if necessary Please generate the best solution."

[1910] The input consists of keywords and contextual data, while the output consists of generated answers and related information.

[1911] Step 7:

[1912] The server adjusts the generated answers and related information based on sentiment data and selects the most appropriate one. This adjustment takes into account the user's current emotional state (e.g., urgent). The input is sentiment data and the generated answers and related information, and the output is the adjusted answers and related information.

[1913] Step 8:

[1914] The server sends the adjusted answers and related information to the terminal. The input is the adjusted answers and related information, and the output is the data sent to the terminal.

[1915] Step 9:

[1916] The device displays the adjusted answers and related information received from the server on its screen and, if necessary, informs the user via voice notification. Specifically, the answers are displayed on the device's screen, and necessary notifications are played through the speaker. The input is the answers and related information sent to the device, and the output is the answers and related information notified to the user.

[1917] The specific processing unit 290 transmits the result of the specific processing to the robot 414. In the robot 414, the control unit 46A causes the speaker 240 and the controlled object 443 to output the result of the specific processing. The microphone 238 acquires audio indicating user input for the result of the specific processing. The control unit 46A transmits the audio data indicating user input acquired by the microphone 238 to the data processing unit 12. In the data processing unit 12, the specific processing unit 290 acquires the audio data.

[1918] Data generation model 58 is a type of so-called generative AI (Artificial Intelligence). One example of data generation model 58 is ChatGPT (Internet search<URL: https: / / openai.com / blog / chatgpt> ), Gemini (Internet search) <url: https: gemini.google.com ?hl="ja">Examples of generative AI include the following. The data generation model 58 is obtained by performing deep learning on a neural network. The data generation model 58 is input with prompts containing instructions, and with inference data such as audio data representing speech, text data representing text, and image data representing images. The data generation model 58 infers from the input inference data according to the instructions indicated by the prompts, and outputs the inference results in data formats such as audio data and text data. Here, inference refers to, for example, analysis, classification, prediction, and / or summarization.

[1919] In the above embodiment, an example was given in which specific processing is performed by the data processing device 12, but the technology of this disclosure is not limited thereto, and the specific processing may also be performed by the robot 414.

[1920] Furthermore, the emotion identification model 59, acting as an emotion engine, may determine the user's emotion according to a specific mapping. Specifically, the emotion identification model 59 may determine the user's emotion according to a specific mapping, which is an emotion map (see Figure 9). Similarly, the emotion identification model 59 may also determine the robot's emotion, and the identification processing unit 290 may perform identification processing using the robot's emotion.

[1921] Figure 9 shows an emotion map 400 in which multiple emotions are mapped. In the emotion map 400, emotions are arranged in concentric circles radiating from the center. The closer to the center of the concentric circles, the more primitive the emotions are located. Further out of the concentric circles, emotions representing states and actions arising from mental states are located. Emotion is a concept that includes feelings and mental states. On the left side of the concentric circles, emotions that are generally generated from reactions occurring in the brain are located. On the right side of the concentric circles, emotions that are generally induced by situational judgment are located. Above and below the concentric circles, emotions that are generally generated from reactions occurring in the brain and induced by situational judgment are located. In addition, the emotion of "pleasure" is located on the upper side of the concentric circles, and the emotion of "displeasure" is located on the lower side. Thus, in the emotion map 400, multiple emotions are mapped based on the structure in which emotions arise, and emotions that are likely to occur simultaneously are mapped close together.

[1922] These emotions are distributed at the 3 o'clock position on the Emotion Map 400, and usually fluctuate between feelings of security and anxiety. In the right half of the Emotion Map 400, situational awareness takes precedence over internal feelings, resulting in a calm impression.

[1923] The inside of the Emotion Map 400 represents inner thoughts, while the outside represents actions. Therefore, the further you go from the outside of the Emotion Map 400, the more visible (expressed in actions) your emotions become.

[1924] Here, human emotions are based on various balances, such as posture and blood sugar levels. When these balances deviate from the ideal, it results in discomfort, and when they approach the ideal, it results in pleasure. Similarly, in robots, cars, motorcycles, etc., emotions can be created based on various balances, such as posture and battery level. When these balances deviate from the ideal, it results in discomfort, and when they approach the ideal, it results in pleasure. The emotion map can be generated, for example, based on Dr. Mitsuyoshi's emotion map (Research on a system for analyzing brain physiological signals of speech emotion recognition and emotion, Tokushima University, doctoral dissertation: https: / / ci.nii.ac.jp / naid / 500000375379). The left half of the emotion map contains emotions belonging to a region called "response," where sensation is dominant. The right half of the emotion map contains emotions belonging to a region called "situation," where situational awareness is dominant.

[1925] The emotion map defines two emotions that promote learning. One is the emotion around the middle of the negative "repentance" and "reflection" on the situation side. In other words, it is when the robot experiences negative emotions such as "I never want to feel this way again" or "I don't want to be scolded again." The other is the emotion around the positive "desire" on the reaction side. In other words, it is when the robot has positive feelings such as "I want more" or "I want to know more."

[1926] The emotion identification model 59 inputs user input into a pre-trained neural network, obtains emotion values ​​representing each emotion shown in the emotion map 400, and determines the user's emotion. This neural network is pre-trained based on multiple training data sets, which are combinations of user input and emotion values ​​representing each emotion shown in the emotion map 400. Furthermore, this neural network is trained so that emotions located close together have similar values, as shown in the emotion map 900 in Figure 10. Figure 10 shows an example where multiple emotions such as "reassured," "calm," and "confident" have similar emotion values.

[1927] The above description primarily focuses on the functions of the data processing device 12 in relation to this disclosure. However, the system related to this disclosure is not necessarily implemented on a server. The system related to this disclosure may be implemented as a general information processing system. This disclosure may be implemented, for example, as a software program that runs on a personal computer or as an application that runs on a smartphone. The method related to this disclosure may be provided to users in SaaS (Software as a Service) format.

[1928] In the above embodiment, an example was given in which a specific process is performed by a single computer 22. However, the technology of this disclosure is not limited thereto, and a distributed processing of the specific process may be performed by multiple computers, including computer 22. For example, a data generation model 58 may be provided in an external device of the data processing device 12, and the external device may generate data according to the input data.

[1929] In the above embodiment, an example was given in which the specific processing program 56 is stored in the storage 32, but the technology of this disclosure is not limited thereto. For example, the specific processing program 56 may be stored in a portable, computer-readable, non-temporary storage medium such as a USB (Universal Serial Bus) memory. The specific processing program 56 stored in the non-temporary storage medium is installed in the computer 22 of the data processing device 12. The processor 28 executes specific processing according to the specific processing program 56.

[1930] Alternatively, the specific processing program 56 may be stored in a storage device such as a server connected to the data processing device 12 via the network 54, and the specific processing program 56 may be downloaded and installed on the computer 22 in response to a request from the data processing device 12.

[1931] Furthermore, it is not necessary to store the entirety of the specific processing program 56 in a storage device such as a server connected to the data processing device 12 via the network 54, or to store the entirety of the specific processing program 56 in the storage 32; it is acceptable to store only a portion of the specific processing program 56.

[1932] The following types of processors can be used as hardware resources to perform specific processing. Examples of processors include a CPU, a general-purpose processor that functions as a hardware resource to perform specific processing by executing software, i.e., a program. Other examples of processors include dedicated electrical circuits, such as FPGAs (Field-Programmable Gate Arrays), PLDs (Programmable Logic Devices), or ASICs (Application Specific Integrated Circuits), which have circuit configurations specifically designed to perform specific processing. All of these processors have built-in or connected memory, and all of them perform specific processing by using memory.

[1933] The hardware resource that performs a specific process may consist of one of these various processors, or it may consist of a combination of two or more processors of the same or different types (for example, a combination of multiple FPGAs, or a combination of a CPU and an FPGA). Alternatively, the hardware resource that performs a specific process may consist of a single processor.

[1934] Examples of configurations using a single processor include, firstly, a configuration in which one or more CPUs and software are combined to form a single processor, and this processor functions as a hardware resource that performs a specific process. Secondly, there is a configuration using a processor that realizes the functions of the entire system, including multiple hardware resources that perform a specific process, on a single IC chip, as exemplified by SoCs (System-on-a-chip). In this way, a specific process is realized using one or more of the above types of processors as hardware resources.

[1935] Furthermore, the hardware structure of these various processors can more specifically utilize electrical circuits that combine circuit elements such as semiconductor devices. Also, the specific processing described above is merely an example. Therefore, it goes without saying that unnecessary steps can be deleted, new steps added, or the processing order rearranged, as long as it does not deviate from the main purpose.

[1936] The descriptions and illustrations presented above are detailed explanations of the technical aspects of this disclosure and are merely examples of the technical aspects. For example, the above descriptions of the structure, function, operation, and effect are examples of the structure, function, operation, and effect of the technical aspects of this disclosure. Therefore, it goes without saying that you may delete unnecessary parts, add new elements, or replace elements in the descriptions and illustrations presented above, as long as you do not deviate from the essence of the technical aspects of this disclosure. Furthermore, in order to avoid confusion and facilitate understanding of the technical aspects of this disclosure, explanations of common technical knowledge and the like that do not require special explanation to enable the implementation of the technical aspects of this disclosure have been omitted from the descriptions and illustrations presented above.

[1937] All documents, patent applications, and technical standards described herein are incorporated by reference to the same extent as if each individual document, patent application, and technical standard were specifically and individually noted as being incorporated by reference.

[1938] The following is further disclosed regarding the embodiments described above.

[1939] (Claim 1)

[1940] [A means by which the terminal uses communication means to record the user's voice and transcribe that voice into text in real time,

[1941] [Means by which the terminal sends the converted data to the server,

[1942] [A means by which the server analyzes the transcribed data, extracts important keywords and context, and generates answers and related information based on them,

[1943] [Means for the server to send generated answers and related information to the terminal,

[1944] [A means by which the device displays answers and related information on the screen and notifies the user of important information by voice,

[1945] A system that includes this.

[1946] (Claim 2)

[1947] [The system according to claim 1, wherein the terminal transmits audio data of the user's consultation content to a speech recognition engine in real time to transcribe it into text.

[1948] (Claim 3)

[1949] [The system according to claim 1, wherein the server analyzes text data using a natural language processing engine located in the cloud and generates answers and related information using a generative AI.

[1950] "Example 1"

[1951] (Claim 1)

[1952] [A means by which the terminal uses communication means to record the user's voice and transcribe that voice into text in real time,

[1953] [Means by which the terminal sends the converted data to the server,

[1954] [A means by which the server analyzes the transcribed data, extracts important keywords and context, and generates answers and related information based on them,

[1955] [Methods for filtering the generated answers and related information on the server and sending them to the terminal,

[1956] [A means by which the device displays answers and related information on the screen and notifies the user of important information by voice,

[1957] A system that includes this.

[1958] (Claim 2)

[1959] [The system according to claim 1, wherein the terminal transmits audio data of the user's consultation content to a speech recognition engine in real time to transcribe it into text.

[1960] (Claim 3)

[1961] [The system according to claim 1, wherein the server analyzes text data using a natural language processing engine located in the cloud and generates answers and related information using a generative AI.

[1962] "Application Example 1"

[1963] (Claim 1)

[1964] [A means by which the terminal uses communication means to record the user's voice and transcribe that voice into text in real time,

[1965] [Means by which the terminal sends the converted data to the server,

[1966] [A means by which the server analyzes the transcribed data, extracts important keywords and context, and generates answers and related information based on them,

[1967] [Means for the server to send generated answers and related information to the terminal,

[1968] [A means by which the device displays answers and related information on the screen and notifies the user of important information by voice,

[1969] A system including a portable information processing device on which an application that accepts user voice input is run.

[1970] (Claim 2)

[1971] [The system according to claim 1, wherein the terminal transmits audio data of the user's consultation content to a speech recognition engine in real time to transcribe it into text.

[1972] (Claim 3)

[1973] [The system according to claim 1, wherein the server analyzes text data using a natural language processing engine located in the cloud and generates answers and related information using a generative AI.

[1974] "Example 2 of combining an emotion engine"

[1975] (Claim 1)

[1976] [A means by which the terminal uses communication means to record the user's voice and transcribe that voice into text in real time,

[1977] [Means by which the terminal transmits text data and sentiment data to the server,

[1978] [A means by which the server analyzes the transcribed data, extracts important keywords and context, and adjusts the answers and related information using sentiment data,

[1979] [Means for the server to send generated answers and related information to the terminal,

[1980] [A means by which the device displays answers and related information on the screen and provides voice notifications in a tone adjusted based on emotional data,

[1981] A system that includes this.

[1982] (Claim 2)

[1983] [The system according to claim 1, wherein the terminal transmits audio data of the user's consultation content to a speech recognition engine in real time to transcribe it into text.

[1984] (Claim 3)

[1985] [The system according to claim 1, wherein the server analyzes text data using a natural language processing engine located in the cloud, generates answers and related information using a generative AI, and adjusts the answers based on sentiment data.

[1986] "Application example 2 when combining with an emotional engine"

[1987] (Claim 1)

[1988] [A means by which the terminal uses communication means to record the user's voice and transcribe that voice into text in real time,

[1989] [Means by which the terminal sends the converted data to the server,

[1990] [A means by which the server analyzes the transcribed data, extracts important keywords and context, and generates answers and related information based on them,

[1991] [Means for the server to send generated answers and related information to the terminal,

[1992] [A means by which the device displays answers and related information on the screen and notifies the user of important information by voice,

[1993] [A means by which the server analyzes the user's emotions and adjusts the answers and related information based on the analysis results,

[1994] [Means for the device to adjust the tone and display content of notifications based on the results of sentiment analysis,

[1995] A system that includes this.

[1996] (Claim 2)

[1997] [The system according to claim 1, wherein the terminal transmits audio data of the user's consultation content to a speech recognition engine in real time to transcribe it into text.

[1998] (Claim 3)

[1999] [The system according to claim 1, wherein the server analyzes text data using a natural language processing engine located in the cloud and generates answers and related information using a generative AI. [Explanation of Symbols]

[2000] 10, 210, 310, 410 Data Processing Systems 12 Data Processing Devices 14 Smart Devices 214 Smart Glasses 314 Headset-type terminal 414 Robots< / url:> < / url:> < / url:> < / url:>

Claims

1. A means by which a terminal uses a communication method to record the user's voice and transcribe that voice into text in real time, A means by which the terminal sends the text data to the server, A means by which a server analyzes transcribed data, extracts important keywords and context, and generates answers and related information based on them, A means for the server to send the generated answers and related information to the terminal, The device displays answers and related information on the screen and notifies the user of important information via voice. A system that includes this.

2. The system according to claim 1, wherein the terminal transmits audio data of the user's consultation content to a speech recognition engine in real time to transcribe it into text.

3. The system according to claim 1, wherein the server analyzes text data using a natural language processing engine located in the cloud and generates answers and related information using a generative AI.

Citation Information

Patent Citations

  • Persona chatbot control method and system

    JP2022180282A