system
The system addresses the challenge of handling suspicious calls by switching to an AI automated response, saving and analyzing conversation content, and suggesting appropriate actions, enhancing response efficiency and accuracy.
Patent Information
- Authority / Receiving Office
- JP · JP
- Patent Type
- Patents
- Current Assignee / Owner
- SOFTBANK GROUP CORP
- Filing Date
- 2025-03-19
- Publication Date
- 2026-07-29
AI Technical Summary
Individuals receiving calls from potential fraudsters or malicious peddlers often struggle to respond calmly and determine the authenticity of the call, leading to difficulties in taking appropriate actions.
A system that switches calls to an AI automatic response, saving conversation content, summarizing the opponent's speech for falsifiability, and proposing the next action.
Enables users to respond calmly and efficiently by automating the call handling process, reducing user burden and improving response accuracy.
Smart Images

Figure 0007897370000001_ABST
Abstract
Description
Technical Field
[0001] The technology of the present disclosure relates to a system.
Background Art
[0002] Patent Document No. 1 discloses a method for controlling a persona chatbot, which is performed by at least one processor, including steps of receiving a user utterance, adding the user utterance to a prompt including an instruction sentence related to an explanation of a chatbot character, encoding the prompt, and inputting the encoded prompt into a language model to generate a chatbot utterance as a response to the user utterance.
Prior Art Documents
Patent Documents
[0003]
Patent Document 1
Summary of the Invention
Problems to be Solved by the Invention
[0004] When a call is received that seems to be from a fraudster or a malicious product peddler, the person receiving the call may not be able to respond calmly. Also, it is difficult to determine whether it is fraud, and it is difficult to take appropriate actions.
Means for Solving the Problems
[0005] The present invention provides means for switching a received call from a human to an AI automatic response. The AI automatic response saves the conversation content and summarizes the falsifiability of the opponent's speech. Furthermore, the AI automatic response proposes the next action. Thereby, the person receiving the call can calm down and execute the action of the AI.
Brief Description of the Drawings
[0006] [Figure 1] This is a conceptual diagram showing an example of the configuration of a data processing system according to the first embodiment. [Figure 2] This is a conceptual diagram showing an example of the essential functions of the data processing device and smart device according to the first embodiment. [Figure 3] This is a conceptual diagram showing an example of the configuration of a data processing system according to the second embodiment. [Figure 4] This is a conceptual diagram showing an example of the main functions of a data processing device and smart glasses according to the second embodiment. [Figure 5] This is a conceptual diagram showing an example of the configuration of a data processing system according to the third embodiment. [Figure 6] This is a conceptual diagram showing an example of the main functions of a data processing device and a headset-type terminal according to the third embodiment. [Figure 7] This is a conceptual diagram showing an example of the configuration of a data processing system according to the fourth embodiment. [Figure 8] This is a conceptual diagram showing an example of the main functions of a data processing device and a robot according to the fourth embodiment. [Figure 9] This shows an emotion map where multiple emotions are mapped. [Figure 10] This shows an emotion map where multiple emotions are mapped. [Figure 11] This is a sequence diagram showing the processing flow of the data processing system in Embodiment 1 of Example 1. [Figure 12] This is a sequence diagram showing the processing flow of the data processing system in Application Example 1 of Form Example 1. [Figure 13] This is a sequence diagram showing the processing flow of the data processing system in Example 2 of Embodiment 2. [Figure 14] This is a sequence diagram showing the processing flow of the data processing system in Application Example 2 of Form Example 2. [Figure 15] This is a sequence diagram showing the processing flow of the data processing system in Embodiment 3 of Example 3. [Figure 16]It is a sequence diagram showing the processing flow of the data processing system in Application Example 3 of Form Example 3. [Figure 17] It is a sequence diagram showing the processing flow of the data processing system in Example 1 of Form Example 1 when combined with an emotion engine. [Figure 18] It is a sequence diagram showing the processing flow of the data processing system in Application Example 1 of Form Example 1 when combined with an emotion engine. [Figure 19] It is a sequence diagram showing the processing flow of the data processing system in Example 2 of Form Example 2 when combined with an emotion engine. [Figure 20] It is a sequence diagram showing the processing flow of the data processing system in Application Example 2 of Form Example 2 when combined with an emotion engine. [Figure 21] It is a sequence diagram showing the processing flow of the data processing system in Example 3 of Form Example 3 when combined with an emotion engine. [Figure 22] It is a sequence diagram showing the processing flow of the data processing system in Application Example 3 of Form Example 3 when combined with an emotion engine. [Figure 23] It is a sequence diagram showing the processing flow of the data processing system in other embodiments.
Embodiments for Carrying Out the Invention
[0007] Hereinafter, an example of an embodiment of the system according to the technology of the present disclosure will be described with reference to the accompanying drawings.
[0008] First, the language used in the following description will be explained.
[0009] In the following embodiments, the labeled processor (hereinafter simply referred to as "processor") may be a single arithmetic unit or a combination of multiple arithmetic units. Also, the processor may be a single type of arithmetic unit or a combination of multiple types of arithmetic units. Examples of arithmetic units include a CPU (Central Processing Unit), a GPU (Graphics Processing Unit), a GPGPU (General-Purpose computing on Graphics Processing Units), an APU (Accelerated Processing Unit), or a TPU (TENSOR PROCESSING UNIT (registered trademark)), etc.
[0010] In the following embodiments, the labeled RAM (Random Access Memory) is a memory in which information is temporarily stored and is used as a work memory by the processor.
[0011] In the following embodiments, the labeled storage is one or more non-volatile storage devices that store various programs, various parameters, etc. Examples of non-volatile storage devices include flash memory (SSD (Solid State Drive)), magnetic disks (e.g., hard disks), or magnetic tapes, etc.
[0012] In the following embodiments, the labeled communication I / F (Interface) is an interface including a communication processor, an antenna, etc. The communication I / F controls communication between multiple computers. Examples of communication standards applied to the communication I / F include wireless communication standards including 5G (5th Generation Mobile Communication System), Wi-Fi (registered trademark), or Bluetooth (registered trademark), etc.
[0013] In the following embodiments, "A and / or B" is synonymous with "at least one of A and B." That is, "A and / or B" means that it may be A alone, or B alone, or a combination of A and B. Furthermore, in this specification, the same concept as "A and / or B" applies when expressing three or more things linked by "and / or."
[0014] [First Embodiment]
[0015] Figure 1 shows an example of the configuration of the data processing system 10 according to the first embodiment.
[0016] As shown in Figure 1, the data processing system 10 includes a data processing device 12 and a smart device 14. An example of the data processing device 12 is a server.
[0017] The data processing device 12 comprises a computer 22, a database 24, and a communication interface 26. The computer 22 is an example of a "computer" related to the technology of this disclosure. The computer 22 comprises a processor 28, RAM 30, and storage 32. The processor 28, RAM 30, and storage 32 are connected to a bus 34. The database 24 and the communication interface 26 are also connected to the bus 34. The communication interface 26 is connected to a network 54. An example of the network 54 is a WAN (Wide Area Network) and / or a LAN (Local Area Network).
[0018] The smart device 14 comprises a computer 36, a reception device 38, an output device 40, a camera 42, and a communication interface 44. The computer 36 comprises a processor 46, RAM 48, and storage 50. The processor 46, RAM 48, and storage 50 are connected to a bus 52. The reception device 38, output device 40, and camera 42 are also connected to the bus 52.
[0019] The reception device 38 is equipped with a touch panel 38A and a microphone 38B, etc., and receives user input. The touch panel 38A receives user input by detecting contact with an object (e.g., a pen or finger). The microphone 38B receives user input by detecting the user's voice. The control unit 46A transmits data indicating the user input received by the touch panel 38A and microphone 38B to the data processing device 12. In the data processing device 12, the specific processing unit 290 acquires the data indicating the user input.
[0020] The output device 40 includes a display 40A and a speaker 40B, and presents data to the user 20 by outputting the data in a form perceptible to the user 20 (e.g., audio and / or text). The display 40A displays visible information such as text and images according to instructions from the processor 46. The speaker 40B outputs audio according to instructions from the processor 46. The camera 42 is a small digital camera equipped with an optical system such as a lens, aperture, and shutter, and an image sensor such as a CMOS (Complementary Metal-Oxide-Semiconductor) image sensor or a CCD (Charge Coupled Device) image sensor.
[0021] Communication interface 44 is connected to network 54. Communication interfaces 44 and 26 are responsible for the exchange of various types of information between processor 46 and processor 28 via network 54.
[0022] Figure 2 shows an example of the main functions of the data processing device 12 and the smart device 14.
[0023] As shown in Figure 2, in the data processing device 12, a specific processing is performed by the processor 28. A specific processing program 56 is stored in the storage 32. The specific processing program 56 is an example of a "program" related to the technology of this disclosure. The processor 28 reads the specific processing program 56 from the storage 32 and executes the read specific processing program 56 on the RAM 30. The specific processing is realized by the processor 28 operating as a specific processing unit 290 according to the specific processing program 56 executed on the RAM 30.
[0024] The storage 32 stores the data generation model 58 and the emotion identification model 59. The data generation model 58 and the emotion identification model 59 are used by the identification processing unit 290.
[0025] In the smart device 14, the processor 46 performs the reception output processing. The storage 50 stores the reception output program 60. The reception output program 60 is used in conjunction with a specific processing program 56 by the data processing system 10. The processor 46 reads the reception output program 60 from the storage 50 and executes the read reception output program 60 on the RAM 48. The reception output processing is realized by the processor 46 operating as a control unit 46A according to the reception output program 60 executed on the RAM 48.
[0026] Next, the identification process performed by the identification processing unit 290 of the data processing device 12 will be described.
[0027] "Example of form 1"
[0028] One embodiment of the present invention provides a means for a human to switch to an AI automated response system when a phone call comes in. Specifically, when a phone call comes in, the person answering the call can switch to the AI automated response system by pressing a specific button.
[0029] "Example of form 2"
[0030] Furthermore, the AI automated response system saves the conversation content and summarizes the potential for perjury in the other party's statements. Specifically, the AI automated response system transcribes the other party's statements into text and stores that text in a database. It also analyzes the text of the other party's statements to determine whether they are potentially perjury.
[0031] "Example of form 3"
[0032] Furthermore, the AI automated response will suggest the next action. Specifically, based on the result of the assessment of the perjury potential of the other party's statement, the AI automated response will suggest what action should be taken next. For example, if there is a high probability that the other party's statement is perjury, the AI automated response will suggest reporting it to the police.
[0033] The following describes the processing flow for each example of the form.
[0034] "Example of form 1"
[0035] Step 1: When a phone call comes in, the person answering the phone presses a specific button.
[0036] Step 2: When a specific button is pressed, the system switches to AI automated response.
[0037] Step 3: The AI automated response system answers the call and starts the conversation.
[0038] "Example of form 2"
[0039] Step 1: The AI automated response system transcribes the conversation into text.
[0040] Step 2: Save the transcribed conversation content to the database.
[0041] Step 3: The AI automated response analyzes the saved text to determine whether the other party's statement is perjury.
[0042] "Example of form 3"
[0043] Step 1: The AI automated response system determines the next action based on its assessment of the perjury potential of the other party's statement.
[0044] Step 2: The AI automated response system suggests the next action and notifies the person who answered the call.
[0045] Step 3: The person who received the call takes action based on the AI automated response's suggestions. For example, they might call the police.
[0046] (Example 1)
[0047] Next, we will describe Example 1 of Form Example 1. In the following description, the data processing device 12 will be referred to as the "server," and the smart device 14 will be referred to as the "terminal."
[0048] Traditional communication systems required human intervention upon receiving a call, leading to challenges in response efficiency and accuracy. Furthermore, the lack of features to assess the reliability of the caller's statements and suggest subsequent actions placed a significant burden on users.
[0049] The identification process performed by the identification processing unit 290 of the data processing device 12 in Example 1 is realized by the following means.
[0050] In this invention, the server includes means for switching received communications from a human to an automated response system, means for the automated response system to record the content of the conversation and evaluate the reliability of the other party's statements, and means for the automated response system to propose the next action. This enables efficient communication responses and reduces the burden on the user.
[0051] "A means of switching received communications from a human to an automated response system" refers to a function in which, when a communication device receives an incoming call, the user performs a specific operation to transfer the response from a human to an automated response system.
[0052] "A means by which an automated response system records the content of a conversation and evaluates the reliability of the other party's statements" refers to a function in which an automated response system saves the conversation in progress, analyzes the content of the statements, and determines their reliability.
[0053] "Means by which an automated response system suggests the next action" refers to a function in which an automated response system suggests the next action the user should take based on the content of the conversation.
[0054] "Means for detecting a specific operation and activating an automatic response device" refers to a function that detects an operation performed by a user, such as pressing a specific button on a communication device, and activates the automatic response device.
[0055] "A means by which an automated response device can offer a greeting at the start of a conversation" refers to a function in which an automated response device plays a pre-set greeting message to the other party at the start of communication.
[0056] A description of embodiments for carrying out this invention will be given.
[0057] The server receives incoming signals from communication devices and detects when a user performs a specific action. Specifically, when a user presses the "AI Response" button on their device, the server activates an automated response system. This automated response system uses AI software such as Google® Dialogflow or Amazon Lex to record the conversation and evaluate the reliability of the other party's statements.
[0058] When the user presses a button on the terminal, the terminal sends that information to the server. Based on the received information, the server activates an automated response system and hands over the communication response to the AI. The automated response system plays a greeting message at the start of the conversation, such as, "Hello, this is the automated response system. How can I help you?"
[0059] As a concrete example, when a user receives a call, pressing the "AI Answer" button on their device activates the automated answering system, and the AI greets the caller. This process allows the user to leave the phone call to the AI.
[0060] An example of a prompt to input into a generating AI model is, "Please tell me the specific steps to switch to AI automatic answering when a phone call comes in." This prompt allows the AI to provide detailed information about how the system works and how to configure it.
[0061] The flow of the specific processing in Example 1 will be explained using Figure 11.
[0062] Step 1:
[0063] The terminal detects when a communication device receives an incoming signal. The input is the incoming signal from the communication device. The terminal sends this signal to the server, notifying it of the incoming call. The output is an incoming call notification sent to the server.
[0064] Step 2:
[0065] The user presses the "AI Response" button on the device. The input is the user's button press. The device detects this action and notifies the server that the button has been pressed. The output is a button press notification sent to the server.
[0066] Step 3:
[0067] The server receives a button press notification from the terminal. The input is the notification from the terminal. Based on this notification, the server generates an instruction to activate the automated response system. The output is an instruction to activate the automated response system.
[0068] Step 4:
[0069] The automated response system starts operating upon receiving a startup command from the server. Its input is the startup command from the server. At the start of the conversation, the automated response system plays a greeting message such as, "Hello, this is the automated response system. How can I help you?" Its output is a greeting message for the recipient.
[0070] Step 5:
[0071] The automated response system records the content of the conversation with the caller and evaluates the reliability of the caller's statements. The input is the caller's statements. The automated response system analyzes the statements and processes the data to evaluate their reliability. The output is the reliability evaluation result.
[0072] (Application Example 1)
[0073] Next, we will describe Application Example 1 of Form Example 1. In the following description, the data processing device 12 will be referred to as a "server," and the smart device 14 will be referred to as a "terminal."
[0074] In today's communication environment, suspicious communications and fraudulent phone calls are on the rise, requiring a swift and effective response. However, traditional methods require human recipients to handle all inquiries, placing a heavy burden on them. Furthermore, there is a risk of errors due to misjudgments. To address these challenges, the introduction of an automated response system utilizing artificial intelligence is necessary.
[0075] The specific processing performed by the specific processing unit 290 of the data processing device 12 in Application Example 1 is realized by the following means.
[0076] In this invention, the server includes means for switching received communications from human to artificial intelligence automated response, means for the artificial intelligence automated response to save the conversation content and summarize the perjury of the other party's statements, means for the artificial intelligence automated response to propose the next action, and means for the artificial intelligence automated response to respond to suspicious communications based on pre-set instructions. This enables a rapid and accurate response to suspicious communications.
[0077] "Received communications" refers to the transmission of information from external sources, such as phone calls and messages.
[0078] "Artificial intelligence automated response" refers to a system that uses artificial intelligence technology to automatically generate responses and send replies to the communication partner.
[0079] "Saving conversation content" refers to recording information exchanged during communication so that it can be referenced later.
[0080] "Summarizing the perjury" refers to evaluating the veracity of the other party's statements and identifying any questionable points.
[0081] "Suggesting the next course of action" means suggesting the appropriate course of action to take based on the content of the communication.
[0082] "Suspicious communications" refer to the transmission of information that differs from normal communications and may be fraudulent or illegal.
[0083] "Pre-set instructions" refer to guidelines for responses and actions that artificial intelligence should follow in specific situations.
[0084] The system for implementing this invention mainly consists of a server and a terminal. The server runs a program to switch received communications from human interaction to automated artificial intelligence responses. Specifically, when a communication is received, the server activates an automated artificial intelligence response based on instructions from the terminal and saves the conversation content. The saved data is used to evaluate the perjury potential of the other party's statements.
[0085] The server uses a generative AI model to propose the next action based on the content of the communication. This uses generative AI models such as OpenAI's GPT-3®. The server responds to suspicious communications based on pre-configured instructions. This allows users to respond quickly and accurately to suspicious communications.
[0086] As a concrete example, when a device receives a suspicious call, the user presses a specific button, and the server activates an AI automated response system that says, "This is the security service. How can I help you?" An example of the prompt might be, "Please provide an example of how to respond to a suspicious call. Please emphasize that this is a security service response."
[0087] The flow of a specific process in Application Example 1 will be explained using Figure 12.
[0088] Step 1:
[0089] The terminal receives a communication. If the user determines the communication is suspicious, they press a specific button on the terminal. This causes the terminal to send a signal to the server requesting the activation of the artificial intelligence automated response system. The input is the received communication, and the output is the activation request signal to the server.
[0090] Step 2:
[0091] The server receives an activation request signal from the terminal and activates an artificial intelligence automated response system. The server uses a generative AI model to generate a response based on the prompt. The input is the activation request signal and prompt from the terminal, and the output is the generated response. Specifically, the server inputs the prompt "Please provide an example of how to respond to a suspicious call. Emphasize that this is a security service response." into the generative AI model and generates a response.
[0092] Step 3:
[0093] The server sends the generated response to the terminal. The terminal then sends this response to the communication partner. The input is the generated response from the server, and the output is the response sent to the communication partner. Specifically, the terminal sends a response to the communication partner saying, "This is the security service. How can I help you?"
[0094] Step 4:
[0095] The server stores the content of the communication and evaluates the perjury potential of the other party's statements. The input is the content of the communication, and the output is the result of the perjury evaluation. Specifically, the server records the conversation content in a database and analyzes the reliability of the statements using natural language processing techniques.
[0096] Step 5:
[0097] Based on the results of the perjury assessment, the server proposes the next action. The input is the result of the perjury assessment, and the output is the proposed action. Specifically, the server generates a suggestion such as "This communication is suspicious. Further verification is required," and sends it to the terminal.
[0098] (Example 2)
[0099] Next, we will describe Example 2 of Form Example 2. In the following description, the data processing device 12 will be referred to as the "server" and the smart device 14 as the "terminal".
[0100] Conventional automated response systems have shortcomings, such as insufficient recording of conversation content and evaluation of the reliability of statements, making them unable to appropriately suggest the next course of action for the user. Furthermore, because fraud detection based on perjury assessment is not performed, users are vulnerable to fraudulent activity.
[0101] The identification process performed by the identification processing unit 290 of the data processing device 12 in Example 2 is realized by the following means.
[0102] In this invention, the server includes means for acquiring voice input, means for converting voice data into text, and means for storing the text data in a storage device. This enables accurate recording of conversation content and evaluation of the perjury potential of statements.
[0103] "Means for acquiring voice input" refers to a device or method for receiving a user's voice and processing it as digital data.
[0104] "Means for converting audio data to text" refers to a technology or device that analyzes an audio signal and converts it into a corresponding string of characters.
[0105] "Means for storing text data in a storage device" refers to a method or apparatus for recording converted text data in a database or other storage medium.
[0106] "Means for analyzing text data and determining the perjury of a statement" refers to a technology or device for processing text data and evaluating the reliability and truthfulness of its content.
[0107] "Means for outputting analysis results" refers to a method or apparatus for presenting the results of text data analysis to the user.
[0108] "Means for switching received communications from human to automated responses" refers to a method or device for transferring a human response to an automated response system upon receiving a communication.
[0109] "Means of suggesting the next action through automated response" refers to a technology or device that suggests the next action a user should take based on analysis results.
[0110] This invention is a system that determines the perjury of a statement by acquiring voice input, converting it into text data, and analyzing it. A specific embodiment of this system is shown below.
[0111] The user speaks aloud through the device's microphone. The device captures this audio as digital data and sends it to the server. The server uses speech recognition software to convert the audio data into text. High-precision text conversion is achieved by using speech recognition technologies such as the Google Cloud Speech-to-Text API.
[0112] The converted text data is stored on a storage device by the server. A database management system is used for storage, and metadata such as the message timestamp and user ID are also recorded.
[0113] Next, the server analyzes the text data using a natural language processing library. Specifically, it performs grammatical analysis using spaCy and then scrutinizes the content of the statements using OpenAI's GPT generative AI model. This allows it to determine the perjury potential of the statements and detect inconsistencies and unnatural points.
[0114] The analysis results are sent from the server to the terminal and presented to the user. The user can then review the results on the terminal screen and decide on their next course of action.
[0115] For example, if a user says, "It didn't rain yesterday," the system transcribes the statement into text and saves it to a database. Then, it refers to weather information and compares it with the actual weather to determine whether the statement is false.
[0116] An example of a prompt would be: "Assess the likelihood that the following statement is perjury: 'It didn't rain yesterday.' Please take actual weather data into consideration."
[0117] In this way, the system can automate the entire process from voice input to perjury detection, enabling it to provide users with quick and accurate information.
[0118] The flow of the specific processing in Example 2 will be explained using Figure 13.
[0119] Step 1:
[0120] The user speaks aloud through the device's microphone. The device captures this audio as digital data. The input is the user's voice, and the output is digital audio data. The device prompts the user to press the "Start Recording" button to begin voice input.
[0121] Step 2:
[0122] The device sends the captured audio data to the server. The server uses speech recognition software to convert the audio data into text. The input is digital audio data, and the output is text data. The server uses the Google Cloud Speech-to-Text API to convert the audio to text.
[0123] Step 3:
[0124] The server stores the converted text data in storage. The input is text data, and the output is text data stored in the database. The server executes SQL queries to insert the text data and associated metadata into the database.
[0125] Step 4:
[0126] The server analyzes stored text data to determine the perjury potential of a statement. The input is text data obtained from a database, and the output is the result of the perjury evaluation. The server performs grammatical analysis using the natural language processing library spaCy and scrutinizes the content of the statement using OpenAI's GPT generative AI model.
[0127] Step 5:
[0128] The server sends the analysis results to the terminal and presents them to the user. The input is the result of the perjury assessment, and the output is the result displayed on the terminal's screen. The terminal displays the analysis results on the screen, and the user confirms the results.
[0129] (Application Example 2)
[0130] Next, we will describe application example 2 of form example 2. In the following description, the data processing device 12 will be referred to as a "server" and the smart device 14 as a "terminal".
[0131] In today's communication environment, real-time evaluation of conversation reliability is crucial. However, conventional technologies have struggled to efficiently determine the perjury potential of statements and present this information to users immediately. Therefore, new methods are needed to improve conversation reliability.
[0132] The specific processing performed by the specific processing unit 290 of the data processing device 12 in Application Example 2 is realized by the following means.
[0133] In this invention, the server includes means for switching received communications from human to artificial intelligence automated response, means for the artificial intelligence automated response to save the conversation content and summarize the perjury potential of the other party's statements, and means for the artificial intelligence automated response to suggest the next action. This makes it possible to evaluate the reliability of the conversation in real time and present it to the user immediately.
[0134] "Communication" refers to the act or means of sending and receiving voice or data.
[0135] "Artificial intelligence automated response" refers to a system that uses artificial intelligence technology to automatically generate responses and engage in conversation.
[0136] "Conversation content" refers to the entirety of the statements and information exchanged during communication.
[0137] "Perjury potential" is an indicator of the degree to which a statement is likely to be false.
[0138] "Speech recognition technology" is a technology that converts speech into text.
[0139] "Natural language processing technology" is a technology used to analyze text data and understand its meaning and intent.
[0140] A "display device" is a device used to present information visually.
[0141] "User" refers to an individual or group that uses the system.
[0142] "Action" refers to the next step or action suggested by the system.
[0143] The system for carrying out this invention includes a server and a terminal. The server has the function of switching received communications from human to artificial intelligence automated responses. The terminal uses speech recognition technology to transcribe speech during communication into text in real time. Specifically, it uses the Google Cloud Speech-to-Text API to convert speech to text.
[0144] The server analyzes the transcribed speech using natural language processing techniques. This analysis utilizes natural language processing libraries such as spaCy and Transformers. The analysis results in an evaluation of the perjury potential of the speech. This evaluation uses a pre-trained generative AI model (e.g., OpenAI GPT-3).
[0145] The evaluation results are displayed in real time on the terminal's display device. This allows users to instantly verify the reliability of the conversation. For example, if someone says "This product is the best-selling in the market" during a meeting, the server will refer to past data and market information to evaluate the reliability of the statement.
[0146] An example of a prompt for a generative AI model might be: "Evaluate the likelihood that this statement is true. Market data is as follows." Using this prompt, the model determines the reliability of the statement and provides the result to the user.
[0147] The flow of a specific process in Application Example 2 will be explained using Figure 14.
[0148] Step 1:
[0149] The device acquires audio during communication via its microphone. The input is audio data, which is converted into text data using the Google Cloud Speech-to-Text API. The output is the transcribed speech.
[0150] Step 2:
[0151] The server receives text data sent from the terminal. The input is text data, which is then analyzed using a natural language processing library (e.g., spaCy, Transformers). The purpose of the analysis is to understand the content of the statement and extract information to assess its perjury potential. The output is the analysis result.
[0152] Step 3:
[0153] The server sends a prompt to a generative AI model (e.g., OpenAI GPT-3) based on the analysis results. The input consists of the analysis results and the prompt text. An example of the prompt text is "Evaluate the likelihood that this statement is true. The market data is as follows." The model evaluates the reliability of the statement and outputs the result.
[0154] Step 4:
[0155] The server receives evaluation results from the generated AI model and sends them to the terminal. The input is the evaluation result, and the output is reliability evaluation information to be presented to the user.
[0156] Step 5:
[0157] The terminal displays reliability evaluation information received from the server on its display device. The input is reliability evaluation information, and the output is information that the user can visually confirm. This allows the user to instantly verify the reliability of the conversation.
[0158] (Example 3)
[0159] Next, we will describe Embodiment 3 of Embodiment Example 3. In the following description, the data processing device 12 will be referred to as the "server" and the smart device 14 as the "terminal".
[0160] In today's communication environment, there is a need to quickly and accurately determine the veracity of what the other party says and to suggest appropriate actions. However, conventional systems have the challenge of having a complex process for evaluating the veracity of statements, making it difficult to clearly indicate the next action the user should take.
[0161] The identification process performed by the identification processing unit 290 of the data processing device 12 in Example 3 is realized by the following means.
[0162] In this invention, the server includes means for switching received communications from a human to an automated response system, means for the automated response system to store conversation information and evaluate the veracity of the other party's statements, and means for the automated response system to propose the next action based on the evaluation result. This makes it possible to quickly evaluate the veracity of statements and propose appropriate actions to the user.
[0163] "Received communications" refers to audio and data signals sent from external sources.
[0164] An "automatic response device" refers to a device that has the function of automatically responding to received communications.
[0165] "Conversation information" refers to the content of audio and text exchanged during communication.
[0166] "Evaluating the truthfulness of a statement or piece of information" refers to the process of judging the accuracy and reliability of such statements or information.
[0167] "Suggesting action" means indicating specific actions to take next based on the evaluation results.
[0168] A "communication terminal" refers to a device used by a user that has the function of performing communication.
[0169] "User" refers to a person who uses the system.
[0170] A description of embodiments for carrying out this invention will be given.
[0171] First, the user enters a prompt message using a communication terminal. This prompt message is used to determine the truthfulness of the other party's statement. As a concrete example, consider the prompt message, "Is this statement true?"
[0172] Next, the terminal sends the entered prompt message to the server. The server uses a generative AI model to analyze the received prompt message. This analysis utilizes natural language processing techniques, employing models such as "OpenAI GPT-3" and "Google BERT." The server uses these models to scrutinize the content of the utterance and evaluate its truthfulness.
[0173] Based on the evaluation results, the server suggests the next course of action. For example, if it is determined that the statement is highly likely to be perjury, the server will decide to "suggest reporting to the police."
[0174] The server then sends the suggested action to the terminal. The terminal displays the received suggestion to the user. The user can then use this suggestion to decide on their next action.
[0175] In this way, it becomes possible to quickly evaluate the veracity of statements and propose appropriate actions. The flow of the specific processing in Example 3 will be explained using Figure 15.
[0176] Step 1:
[0177] The user enters a prompt message using a communication terminal. For example, they might enter the prompt message, "Is this statement true?" This input serves as the basis for determining the truthfulness of the other party's statement.
[0178] Step 2:
[0179] The terminal sends the entered prompt message to the server. Here, the input is the prompt message, and the output is the data sent to the server. The terminal securely transmits the data using the HTTPS protocol.
[0180] Step 3:
[0181] The server inputs the received prompt text into a generating AI model. The input is the prompt text, and the output is the analysis result by the AI model. The server uses natural language processing models such as "OpenAI GPT-3" and "Google BERT" to analyze the content of the utterance and evaluate its truthfulness.
[0182] Step 4:
[0183] The server proposes the next course of action based on the analysis results. The input is the analysis results of the AI model, and the output is the proposed action. For example, if the server determines that there is a high probability that the statement is perjury, it will decide to take the action of "I suggest reporting this to the police."
[0184] Step 5:
[0185] The server sends the proposed action to the terminal. The input is the proposed action, and the output is the transmission of data to the terminal. The server securely transmits the data using the HTTPS protocol.
[0186] Step 6:
[0187] The terminal displays received suggestions to the user. The input is suggestions from the server, and the output is what is displayed to the user. The terminal displays the suggestions on the screen, allowing the user to decide on their next action.
[0188] (Application Example 3)
[0189] Next, we will describe application example 3 of form example 3. In the following description, the data processing device 12 will be referred to as a "server" and the smart device 14 as a "terminal".
[0190] In modern society, the spread of false information and fraudulent activities through communications is increasing, and there is a need for a swift and effective response. However, conventional systems have difficulty determining the falsehood of communications in real time and suggesting appropriate actions. As a result, the risk of users becoming involved in fraudulent activities is increasing.
[0191] The specific processing performed by the specific processing unit 290 of the data processing device 12 in Application Example 3 is realized by the following means.
[0192] In this invention, the server includes means for switching received communications from a human to an AI-powered automated response, means for the AI automated response to record the content of the conversation and evaluate the falsity of the other party's statements, and means for the AI automated response to propose the next action. This makes it possible to determine the falsity of the communication content in real time and propose appropriate actions to the user.
[0193] "Means for switching received communications from human response to AI-powered automated response" refers to technology that switches from a conventional human response to an automated response by artificial intelligence when communication begins.
[0194] "An artificial intelligence automated response system that records the content of conversations and evaluates the veracity of the other party's statements" refers to a technology in which artificial intelligence stores the content of communications, analyzes that content, and determines the truthfulness of the statements.
[0195] "An artificial intelligence automated response system that suggests the next action" refers to a technology that, based on analysis results, suggests the appropriate next action to the user.
[0196] "Means of converting dialogue into text information using speech recognition technology" refers to technology for converting spoken dialogue into text format.
[0197] "Methods for analyzing the possibility of falsehood using generative AI models" refers to techniques that utilize generative AI models to evaluate the falsehood of dialogue content.
[0198] "A means of displaying warnings based on analysis results and suggesting reporting to legal authorities if necessary" refers to a technology that issues warnings to users based on the results of a falsehood assessment and, if necessary, encourages them to report to legal authorities.
[0199] The system for implementing this invention mainly consists of a server and a terminal. The server has the function of switching received communications from human to AI-powered automated responses. The terminal converts the dialogue into text information using speech recognition technology. Specifically, it uses speech recognition software such as the Google Speech-to-Text API to convert speech data into text data.
[0200] The server uses a generative AI model to analyze the falsity of the converted text data. This analysis utilizes generative AI models such as OpenAI's GPT-4®. The generative AI model evaluates the likelihood that the content is false based on the input text.
[0201] Based on the analysis results, the server displays a warning on the terminal and, if necessary, suggests reporting to legal authorities. This allows users to understand the falsity of communications in real time and take appropriate action.
[0202] As a concrete example, consider a scenario where a user is using this system during a business meeting. When someone says, "This product is 100% safe," the device uses speech recognition technology to convert the statement into text, which is then analyzed by a server using a generated AI model. If the analysis determines that the statement is highly false, the device displays a warning to the user: "This statement requires caution. Please review the details or consider legal action if necessary."
[0203] Examples of prompt statements to input into a generative AI model include the following:
[0204] "Assess the likelihood that the following statement is perjury: 'This product is 100% safe.'"
[0205] The flow of the specific processing in Application Example 3 will be explained using Figure 16.
[0206] Step 1:
[0207] The device receives voice input from the user. This voice input is audio data of the user's conversation. The device converts this audio data into text data using the Google Speech-to-Text API. The converted text data is then output.
[0208] Step 2:
[0209] The server receives text data sent from the terminal. The server inputs this text data into a generative AI model (e.g., GPT-4). The generative AI model analyzes the content of the text data and evaluates the likelihood of the statement being false. As a result of the evaluation, a score indicating the likelihood of falsehood is output.
[0210] Step 3:
[0211] The server determines whether to display a warning to the user based on the evaluation results from the generated AI model. If the falsehood score exceeds a certain threshold, the server sends a warning message to the terminal. The warning message states that caution is needed regarding what is said.
[0212] Step 4:
[0213] The terminal receives a warning message sent from the server and displays it to the user. The user reviews the displayed warning message and considers reporting it to legal authorities if necessary. This allows the user to understand the falsity of the communication content in real time and take appropriate action.
[0214] Furthermore, an emotion engine that estimates the user's emotions may be incorporated. That is, the identification processing unit 290 may use the emotion identification model 59 to estimate the user's emotions and perform identification processing using the user's emotions.
[0215] "Example of form 1"
[0216] One embodiment of the present invention provides a system in which an AI automated response system includes an emotion engine that recognizes the user's emotions. This emotion engine analyzes the user's emotions from the tone of their voice, manner of speaking, word choice, etc. For example, if the user is feeling angry or anxious, the emotion engine feeds that information back to the AI automated response system, which adjusts its response accordingly.
[0217] "Example of form 2"
[0218] The system also provides an emotion engine that adjusts the AI's automated responses based on the user's emotions. Specifically, if the system determines that the user is angry, the AI's automated response will be adjusted to use more polite language and a calmer tone. This allows for more appropriate responses that take the user's emotions into consideration.
[0219] "Example of form 3"
[0220] Furthermore, the system provides an emotion engine that adjusts the AI's automated response suggestions for the next action based on the user's emotions. For example, if the system determines that the user is feeling anxious, the AI automated response will suggest actions to alleviate that anxiety. This enables more appropriate action suggestions that are tailored to the user's emotions.
[0221] The following describes the processing flow for each example of the form.
[0222] "Example of form 1"
[0223] Step 1: The user receives a suspicious phone call.
[0224] Step 2: The user switches the phone to AI automated answering.
[0225] Step 3: The emotion engine analyzes the user's emotions from their voice.
[0226] Step 4: The AI automated response system adjusts its response based on feedback from the emotion engine.
[0227] "Example of form 2"
[0228] Step 1: The user receives a suspicious phone call.
[0229] Step 2: The user switches the phone to AI automated answering.
[0230] Step 3: The emotion engine analyzes the user's emotions from their voice.
[0231] Step 4: The AI automated response system adjusts its response based on feedback from the emotion engine.
[0232] "Example of form 3"
[0233] Step 1: The user receives a suspicious phone call.
[0234] Step 2: The user switches the phone to AI automated answering.
[0235] Step 3: The emotion engine analyzes the user's emotions from their voice.
[0236] Step 4: The AI automated response system adjusts its next action suggestion based on feedback from the emotion engine.
[0237] (Example 1)
[0238] Next, we will describe Example 1 of Form Example 1. In the following description, the data processing device 12 will be referred to as the "server," and the smart device 14 will be referred to as the "terminal."
[0239] Traditional communication systems required recipients to respond manually, making it difficult to provide appropriate responses that reflected emotions. Furthermore, they lacked automated processes for saving conversation content and determining perjury, and also lacked features to suggest the next course of action.
[0240] The identification process performed by the identification processing unit 290 of the data processing device 12 in Example 1 is realized by the following means.
[0241] In this invention, the server includes means for switching received communications from a human to an AI-powered automated response, means for the AI automated response to acquire voice data and perform emotion analysis, and means for the AI automated response to adjust the response content based on the emotion analysis. This enables appropriate responses according to emotions without the recipient having to respond manually.
[0242] "Received communications" refers to audio information sent to the terminal from an external source, such as phone calls or voice messages.
[0243] "Artificial intelligence automated response" refers to a program or system that automatically generates and transmits a response to an incoming communication.
[0244] "Audio data" refers to data that represents audio information obtained from received communications in digital format.
[0245] "Emotional analysis" refers to the process of analyzing audio data to identify the speaker's emotional state.
[0246] "Adjusting the response content" refers to changing the content and tone of the response generated by the AI automated response system based on the results of sentiment analysis.
[0247] "Saving conversation content" refers to recording audio information exchanged during communication so that it can be referenced later.
[0248] "Summarizing the perjury" refers to analyzing the content of a conversation and organizing the information needed to evaluate the veracity of the statements.
[0249] "Suggesting the next action" refers to the AI automated response system suggesting the appropriate course of action to take based on the content and analysis results of the communication.
[0250] This invention provides a system that efficiently responds to incoming communications using artificial intelligence-based automated response. Specific embodiments of this system are described below.
[0251] When the server receives a communication, it first activates an AI-powered automated response system. This automated response system uses speech recognition software to acquire voice data and perform sentiment analysis. Specifically, it uses a "voice capture system" as the speech recognition software and a "sentiment analysis engine" for sentiment analysis. This analyzes the user's voice tone, speaking style, and word choice to identify their emotions.
[0252] The device switches to an AI-powered automated response system when the user presses the "AI Response Switch Button" upon receiving a communication. This action sends a signal to the server, initiating the AI automated response. The server adjusts the response based on the sentiment analysis results, providing the user with an appropriate response.
[0253] For example, when a user receives a call and presses the "AI response switch button," the server immediately acquires the voice data and performs sentiment analysis. If the analysis indicates that the user is feeling anxious, the AI automated response will respond in a gentle tone, saying, "Please don't worry. How can I help you?" This response is recorded on the server and used to improve the generated AI model.
[0254] An example of a prompt sentence to input into a generative AI model might be, "If a user is angry on the phone, how should the AI automated response system respond?" By using this prompt sentence, the generative AI model can suggest an appropriate response.
[0255] The flow of the specific processing in Example 1 will be explained using Figure 17.
[0256] Step 1:
[0257] The terminal detects incoming communications. When the user presses the "AI response switch button" upon receiving a communication, the terminal sends a signal to the server requesting the activation of the artificial intelligence automatic response. The input for this step is the incoming communication signal, and the output is the activation request signal to the server.
[0258] Step 2:
[0259] When the server receives a startup request from the terminal, it activates an artificial intelligence automated response system. The server uses speech recognition software to obtain voice data from the communication. The input for this step is the startup request signal from the terminal, and the output is the obtained voice data.
[0260] Step 3:
[0261] The server sends the acquired audio data to the emotion analysis engine. The emotion analysis engine analyzes the audio data and identifies the user's emotions. The input to this step is the audio data, and the output is the analyzed emotion information. Specifically, it analyzes the tone of voice, speaking style, and word choice.
[0262] Step 4:
[0263] The server adjusts the content of the AI-powered automated response based on emotional information from the emotion analysis engine. The server uses a generative AI model to generate an appropriate response that matches the user's emotions. The input for this step is the analyzed emotional information, and the output is the adjusted response.
[0264] Step 5:
[0265] The server sends the refined response to the user. The user receives the response from the AI automated response system. The input for this step is the refined response, and the output is the response to the user. Specifically, this involves using a gentle tone and reassuring language in the response.
[0266] (Application Example 1)
[0267] Next, we will describe Application Example 1 of Form Example 1. In the following description, the data processing device 12 will be referred to as a "server," and the smart device 14 will be referred to as a "terminal."
[0268] Conventional communication systems made it difficult for recipients to quickly assess the urgency of a call and respond appropriately. Furthermore, they were unable to accurately analyze the caller's emotions and adjust their responses accordingly, resulting in inadequate responses, particularly in emergencies where a rapid and appropriate response was required.
[0269] The specific processing performed by the specific processing unit 290 of the data processing device 12 in Application Example 1 is realized by the following means.
[0270] In this invention, the server includes means for switching the received communication from a human to an artificial intelligence automatic response, means for analyzing the emotions of the caller using an emotion analysis engine and determining the urgency, and means for automatically executing an appropriate response according to the determined urgency. This enables a prompt and appropriate response based on the emotions of the caller.
[0271] "Received communication" refers to the act of receiving information including voice such as a phone call or a voice message.
[0272] "Artificial intelligence automatic response" refers to an artificial intelligence system that automatically generates a response to the received communication and conducts a conversation.
[0273] "Save the conversation content" refers to the act of recording the content of a phone call or conversation so that it can be referred to later.
[0274] "Summarize the perjury" refers to the act of judging whether the other party's statement is true and organizing the result.
[0275] "Propose the next action" refers to suggesting the next action to be taken based on the current situation.
[0276] "Emotion analysis engine" refers to a technology for analyzing emotions from the voice and words of a caller and identifying those emotions.
[0277] "Determine the urgency" refers to the act of evaluating the importance and urgency of a situation and determining the priority of the response.
[0278] "Automatically execute an appropriate response" refers to automatically performing the necessary actions based on the determined urgency.
[0279] The system for implementing this invention operates in a network environment including a server and terminals. The server has a function to switch the received communication to an AI automatic response and uses an emotion analysis engine to analyze the voice of the caller. The emotion analysis engine analyzes the tone, speaking style, and word choice of the caller's voice to identify emotions. Thereby, the server determines the urgency and automatically executes an appropriate response.
[0280] The terminal provides an interface for transmitting voice to the server, and the user can switch to the AI automatic response by pressing a specific button. The server uses the speech_recognition library of Python as voice recognition software and a custom EmotionEngine for emotion analysis. Furthermore, the generated response is provided by the AIResponse module.
[0281] As a specific example, when the user is using a terminal of a security company, in the event of an emergency call, the user can switch to the AI automatic response by pressing a button on the terminal. The server analyzes the caller's statement "Help, someone has broken into my house", and the emotion analysis engine detects strong fear. Thereby, the server immediately issues an instruction to dispatch a security guard.
[0282] An example of a prompt sentence for the generation AI model is "Analyze the voice of the caller and identify the emotion. If the emotion is fear, a prompt response is required."
[0283] The flow of the specific process in Application Example 1 will be described using FIG. 18.
[0284] Step 1:
[0285] When the terminal detects that the user presses a button to switch the received communication to the AI automatic response, it transmits the voice data to the server. The input is the user's voice, and the output is the transmission of the voice data to the server. The terminal converts the voice into digital data and transmits it to the server via the network.
[0286] Step 2:
[0287] The server converts the received audio data into text using the speech_recognition library. The input is audio data, and the output is text data. The server applies a speech recognition algorithm to extract text from the audio.
[0288] Step 3:
[0289] The server inputs text data into an emotion analysis engine to identify the caller's emotions. The input is text data, and the output is emotion information. The emotion analysis engine analyzes the content and word choice of the text to determine the emotion.
[0290] Step 4:
[0291] The server determines the urgency based on emotional information and decides on the appropriate response. The input is emotional information, and the output is a response instruction. If the emotion is fear or anxiety, the server determines that a rapid response is necessary and generates a response instruction.
[0292] Step 5:
[0293] The server executes the generated response instructions and takes specific actions as needed, such as dispatching security personnel. The input is the response instructions, and the output is the actions taken. The server notifies the relevant systems and services and implements the response according to the instructions.
[0294] (Example 2)
[0295] Next, we will describe Example 2 of Form Example 2. In the following description, the data processing device 12 will be referred to as the "server" and the smart device 14 as the "terminal".
[0296] In a conventional automatic response system, only the conversation content with the user is recorded, and no adjustment of the response based on the falsity of the speech or the user's emotions is performed. Therefore, the user experience is poor, and misunderstandings and dissatisfaction may occur. In addition, since the function of proposing the next action is lacking, there is a problem that it is difficult for the user to take appropriate actions.
[0297] The specific processing by the specific processing unit 290 of the data processing apparatus 12 in the second embodiment is realized by the following means.
[0298] In this invention, the server includes means for switching the received communication from a human to an automatic response device, means for the automatic response device to record the conversation content and analyze the falsity of the speech of the other party, and means for the automatic response device to adjust the response content based on emotion analysis. Thereby, in the conversation with the user, it is possible to provide an appropriate response according to the emotion while considering the falsity of the speech, and further propose the next action.
[0299] The "received communication" refers to the voice and data sent from the outside, and is intended to be switched to an automatic response device.
[0300] The "automatic response device" refers to a device or system for automatically processing the conversation with the user, recording, analyzing, and responding.
[0301] "Recording the conversation content" refers to the act of saving the conversation and communication content with the user as data, which is used for later analysis and reference.
[0302] "Analyzing the falsity of the speech" refers to the process of analyzing the speech content of the user and judging whether the content is based on facts.
[0303] "Emotion analysis" refers to the technology of reading emotions from the user's speech and voice and identifying emotional states such as anger, joy, and sadness.
[0304] <"Adjusting response content" refers to changing the content and tone of the response generated by the automated response system according to the analyzed emotions and situation.
[0305] "Suggesting the next action" refers to showing the user the appropriate next steps or actions based on their current situation.
[0306] This invention is a system that efficiently manages user interaction using an automated response device and improves the user experience. Specific embodiments of this system are described below.
[0307] The server switches incoming communications from a human to an automated response system. This involves receiving voice data over the communication network and converting it to text using speech recognition software. Specifically, it uses a "speech recognition API" for speech recognition.
[0308] The terminal receives the user's voice input and sends it to the server. The server converts the received voice data into text using a "speech recognition API" and stores that text in a database. A "database management system" is used for this database.
[0309] The server analyzes the stored text using a natural language processing library to determine whether the statement is false. This allows the server to verify whether the user's statement is based on facts.
[0310] Furthermore, the device uses a "sentiment analysis engine" to analyze the user's emotions from their voice and text. The server receives the results of the sentiment analysis and determines what emotion the user is experiencing, such as anger, joy, or sadness.
[0311] The server adjusts the content of the AI automated response based on the results of the sentiment analysis. For example, if it is determined that the user is angry, the AI automated response will be set to respond in a calm tone, such as, "I'm sorry. Please let me know how we can improve."
[0312] For example, if a user asks, "Does this product really work?", the server transcribes the question into text and stores it in the database. It then compares this question with past data and product information to provide accurate information.
[0313] Examples of prompts include, "Transcribe the user's utterance into text and generate a prompt to determine its falsity," and "Create a prompt that adjusts the AI's response based on the user's sentiment."
[0314] This system makes it possible to manage user interactions in a more accurate and emotionally sensitive manner.
[0315] The flow of the specific processing in Example 2 will be explained using Figure 19.
[0316] Step 1:
[0317] The device receives voice input from the user. For example, the user might say, "Does this product really work?" The device then sends this voice data to the server. The input is the user's voice data, and the output is the transfer of the voice data to the server.
[0318] Step 2:
[0319] The server converts the received audio data into text using a speech recognition API. The speech recognition API analyzes the audio signal and generates corresponding text data. The input is audio data, and the output is text data.
[0320] Step 3:
[0321] The server stores the transcribed conversation content in a database. A database management system is used to record the text data in the appropriate format. The input is text data, and the output is storage in the database.
[0322] Step 4:
[0323] The server analyzes the stored text using a natural language processing library to determine the falsity of the statements. The natural language processing library analyzes the text data and evaluates its falsity by comparing it with known facts. The input is text data, and the output is the result of the falsity evaluation.
[0324] Step 5:
[0325] The device uses an emotion analysis engine to analyze the user's emotions from their voice or text. The emotion analysis engine identifies the user's emotional state from the input data. The input is voice or text data, and the output is the result of the emotion analysis.
[0326] Step 6:
[0327] The server adjusts the content of the AI-generated response based on the results of the sentiment analysis. The server changes the tone and content of the response according to the user's emotions to generate an appropriate response. The input is the result of the sentiment analysis, and the output is the adjusted response content.
[0328] Step 7:
[0329] The server sends a pre-arranged response to the terminal and provides the response to the user in voice or text. The terminal receives the response from the server and relays it to the user. The input is the pre-arranged response, and the output is the provision of the response to the user.
[0330] (Application Example 2)
[0331] Next, we will describe application example 2 of form example 2. In the following description, the data processing device 12 will be referred to as a "server" and the smart device 14 as a "terminal".
[0332] In modern communications, judging the reliability of what the other party says is crucial, but this is difficult with conventional systems. Furthermore, there is a need to provide appropriate responses that respond to the user's emotions, but effective means to achieve this are lacking. Additionally, there is a demand for a function that issues warnings based on the assessment of perjury, but the technology to implement this is still immature.
[0333] The specific processing performed by the specific processing unit 290 of the data processing device 12 in Application Example 2 is realized by the following means.
[0334] In this invention, the server includes means for switching received communications from human to machine learning automated responses, means for the machine learning automated response to save the conversation content and summarize the perjury potential of the other party's statements, and means for the machine learning automated response to adjust the response content based on sentiment analysis. This makes it possible to determine the reliability of statements in communications and to provide appropriate responses that correspond to the user's emotions.
[0335] "Communication" refers to the act or process of sending and receiving information, and includes forms such as voice, text, and data.
[0336] "Machine learning automated response" refers to a system that uses machine learning technology to automatically generate responses, providing appropriate answers based on user input.
[0337] "Saving conversation content" means recording information exchanged during communication and making it available for later reference.
[0338] "Summarizing the perjury" means evaluating whether the other party's statements are true or false, and then organizing and presenting the results.
[0339] "Emotional analysis" is a technology that infers emotions from a user's words and actions and analyzes their emotional state.
[0340] "Adjusting the response content" means changing the content and tone of the response according to the user's emotions and situation.
[0341] "Issuing a warning" means providing a message to draw attention when certain conditions are met.
[0342] "Suggesting an action" means indicating to the user what they should do next.
[0343] The system for implementing this invention mainly consists of a server and a terminal. The server has the function of switching received communications from human input to automated machine learning responses. The terminal is responsible for receiving input from the user and sending it to the server.
[0344] The server uses natural language processing libraries (e.g., spaCy, NLTK) to analyze the received conversation content and determine the perjury potential of the other party's statements. Furthermore, it uses sentiment analysis APIs (e.g., IBM Watson® Tone Analyzer) to analyze the user's emotions and adjust the response based on the results. This enables appropriate responses tailored to the user's emotions.
[0345] Furthermore, the server has a function to issue warnings based on its assessment of perjury, allowing it to alert users. This can improve the reliability of communications.
[0346] For example, if a user asks, "Is this transaction safe?", the server will refer to past data and evaluate the reliability of the statement. Furthermore, if the server determines that the user is feeling uneasy, it will respond in a calm tone, such as, "Please rest assured, this transaction is safe."
[0347] An example of a prompt for a generative AI model is: "Analyze the user's statements and determine their perjury potential. Also, analyze the user's emotions and generate an appropriate response."
[0348] The flow of a specific process in Application Example 2 will be explained using Figure 20.
[0349] Step 1:
[0350] The device receives voice input from the user. This voice input is converted into text data using a speech recognition API (e.g., Google Speech-to-Text). This text data then becomes the input for the next process.
[0351] Step 2:
[0352] The server analyzes the received text data using a natural language processing library (e.g., spaCy, NLTK). The analysis extracts the content of the statement and generates data for determining its perjury potential. This data then becomes the input for the next process.
[0353] Step 3:
[0354] The server performs a perjury assessment. Specifically, it compares the statement against past database data to evaluate its reliability. This evaluation result becomes the input for the next process.
[0355] Step 4:
[0356] The server uses a sentiment analysis API (e.g., IBM Watson Tone Analyzer) to analyze the user's emotions from text data. Based on the analysis results, the user's emotional state is identified. This emotional state becomes the input for the next process.
[0357] Step 5:
[0358] The server generates an appropriate response based on the perjury assessment results and emotional state. A generative AI model is used to create the response to the user. This response becomes the input for the next process.
[0359] Step 6:
[0360] The server issues a warning as needed based on the results of the perjury assessment. A warning message is generated and sent to the user.
[0361] Step 7:
[0362] The terminal displays responses and warning messages received from the server to the user. Based on this, the user can decide on their next action.
[0363] (Example 3)
[0364] Next, we will describe Embodiment 3 of Embodiment Example 3. In the following description, the data processing device 12 will be referred to as the "server" and the smart device 14 as the "terminal".
[0365] Conventional AI-powered automated response systems could determine the perjury potential of a statement, but they lacked the ability to offer appropriate action suggestions that took the user's emotional state into account. Furthermore, their action suggestions based on the perjury potential were limited, resulting in insufficient responses to alleviate user anxiety.
[0366] The identification process performed by the identification processing unit 290 of the data processing device 12 in Example 3 is realized by the following means.
[0367] In this invention, the server includes means for switching received communications from human to artificial intelligence automated response, means for the artificial intelligence automated response to save the conversation content and determine the perjury of the other party's statements, means for analyzing the user's emotional state using an emotion analysis engine, and means for proposing appropriate actions based on the results of the perjury determination and the emotion analysis. This enables flexible and appropriate action suggestions that respond to the user's emotions.
[0368] "Means for switching received communications from human to AI-powered automated responses" refers to a function that, after a communication has been received by a human, switches that response to an automated response by artificial intelligence.
[0369] "An artificial intelligence automated response system that saves conversation content and determines whether the other party's statements are perjury" refers to a function in which artificial intelligence records the content of a conversation, analyzes that content, and evaluates the truthfulness of the statements.
[0370] "Methods for analyzing a user's emotional state using an emotion analysis engine" refers to technologies that infer emotions from a user's statements and actions and analyze that state.
[0371] "A means of suggesting appropriate actions based on the results of perjury assessment and sentiment analysis" refers to a function that suggests the next course of action, taking into account the truthfulness of a statement and the user's emotional state.
[0372] A description of embodiments for carrying out this invention will be given.
[0373] The server generates a program to switch the communication received from the user to an AI-powered automated response. This program runs using server hardware equipped with an NVIDIA GPU and a generative AI model using TENSORFLOW®. The server receives the user's statements as text data and analyzes their content using natural language processing techniques. The analysis includes data processing to determine the perjury potential of the statements and data calculations using an emotion analysis engine.
[0374] For example, if a user asks "Does this product really work?" through their device, the server uses a generative AI model to analyze the statement and extract key keywords. The server then compares the statement against a past database to evaluate its credibility. Furthermore, if the sentiment analysis engine determines that the user is feeling anxious, the server suggests actions such as, "This information may not be reliable. We recommend consulting a professional for further information."
[0375] An example of a prompt to be input to the generating AI model is, "Analyze the user's statement, determine its perjury potential, and suggest an appropriate action." This prompt allows the server to generate an appropriate response to the user's statement. The flow of specific processing in Example 3 will be explained using Figure 21.
[0376] Step 1:
[0377] The user inputs their message into the system via a terminal. The entered text data is sent to the server. The server receives this text data and prepares it for the next analysis step.
[0378] Step 2:
[0379] The server analyzes the received text data using natural language processing techniques. Specifically, it uses a generative AI model to understand the content of the utterance and extract important keywords and context. This analysis clarifies the intent and subject of the utterance. The input is the user's utterance, and the output is the analyzed keywords and contextual information.
[0380] Step 3:
[0381] The server determines the perjury potential of a statement based on the analysis results. It compares the statement against a past database to evaluate its credibility. This process involves database searches and comparison operations. The input is the analyzed keywords and contextual information, and the output is the result of the perjury determination.
[0382] Step 4:
[0383] The server uses an emotion analysis engine to analyze the user's emotional state. It infers emotions from the user's statements and identifies emotions such as anxiety and anger. The input is the user's statements, and the output is the evaluation result of the emotional state.
[0384] Step 5:
[0385] The server suggests the next course of action based on the results of the perjury assessment and sentiment analysis. For example, if there is a high probability that the statement is perjury and the user is feeling anxious, the server might suggest, "This information may not be reliable. We recommend consulting a professional." The input is the result of the perjury assessment and the evaluation of the emotional state, and the output is the suggested action.
[0386] Step 6:
[0387] The server sends the generated suggestions to the terminal and presents them to the user. The user can then decide on their next action based on these suggestions. The input is the suggested action, and the output is the content presented to the user.
[0388] (Application Example 3)
[0389] Next, we will describe application example 3 of form example 3. In the following description, the data processing device 12 will be referred to as a "server" and the smart device 14 as a "terminal".
[0390] In modern society, fraudulent activities and the spread of false information through communications are on the rise, making it difficult for users to protect themselves from these threats. Furthermore, while appropriate responses that respond to users' emotions are required, conventional systems have not been able to adequately achieve this. Therefore, ensuring user safety and peace of mind remains a challenge.
[0391] The specific processing performed by the specific processing unit 290 of the data processing device 12 in Application Example 3 is realized by the following means.
[0392] In this invention, the server includes means for switching received communications from an information processing device to an automated response device, means for the automated response device to record conversation information and evaluate the falsity of the other party's statements, means for the automated response device to propose the next action, and means for an emotion analysis device to evaluate the user's emotions and adjust the proposed action. This makes it possible for the user to evaluate the falsity of the communication content and receive an appropriate response according to their emotions.
[0393] An "information processing device" is a device that receives communications and, if necessary, switches to an automatic response device.
[0394] An "automatic response device" is a device that records received conversation information and has the function of evaluating the likelihood of the other party's statements being false.
[0395] "Conversational information" refers to voice or text data exchanged through communication.
[0396] "Falsehood" is an evaluation criterion that indicates the possibility that a statement is not based on facts.
[0397] An "emotion analysis device" is a device that evaluates the user's emotions and adjusts the suggested actions accordingly.
[0398] A "means of suggesting actions" refers to a function that indicates to the user the next action they should take based on the evaluation results.
[0399] To implement this invention, a server plays a central role. The server is equipped with an information processing device that receives communications and switches the received communications to an automatic response device. The automatic response device records conversation information and evaluates the falsity of the other party's statements. The evaluation is performed using natural language processing techniques, specifically using software such as Python or TensorFlow.
[0400] Furthermore, an emotion analysis device evaluates the user's emotions and adjusts the suggested actions accordingly. Emotion analysis utilizes emotion analysis APIs such as IBM Watson. This enables the provision of appropriate action suggestions based on the user's emotional state.
[0401] For example, if a user receives a potentially fraudulent phone call, the server will suggest, "This call may be fraudulent. Do you want to report it to the police?" If the user is feeling uneasy, it will also suggest, "Do you want to send a notification to a trusted contact?"
[0402] Examples of prompts generated using AI models include, "Analyze the content of this message and assess the likelihood of perjury," and "Analyze the user's sentiment and suggest appropriate action."
[0403] The flow of the specific processing in Application Example 3 will be explained using Figure 22.
[0404] Step 1:
[0405] The server receives communications from the terminal using an information processing device. Input can include voice or text data. The server sends this data to an automated response system for recording as conversation information.
[0406] Step 2:
[0407] The automated response system analyzes recorded conversation information. The input is the audio or text data recorded in step 1. The server uses natural language processing techniques to evaluate the falsity of the statements. Specifically, it analyzes the data using Python or TensorFlow and outputs a falsity score.
[0408] Step 3:
[0409] The server suggests the next action based on the falsehood score. The input is the falsehood score obtained in step 2. If the score is high, the server will generate a suggestion such as, "This call may be perjury. Do you want to report it to the police?"
[0410] Step 4:
[0411] The emotion analysis device evaluates the user's emotions. The input is the voice or text data received in step 1. The server analyzes the emotions using an emotion analysis API such as IBM Watson and outputs the emotional state.
[0412] Step 5:
[0413] The server adjusts the suggested actions based on the emotional state. The input is the emotional state obtained in step 4. If the user is feeling anxious, the server will generate a suggestion such as, "Do you want to send a notification to a trusted contact?"
[0414] Step 6:
[0415] The user receives suggestions from the server and selects the appropriate action. The input consists of the suggestions generated in steps 3 and 5. The user decides on an action based on the suggestions and sends feedback to the server as needed.
[0416] (Other examples)
[0417] Next, other embodiments will be described. In the following description, the data processing device 12 will be referred to as a "server" and the smart device 14 as a "terminal".
[0418] In today's communication environment, there is a demand for quick and appropriate responses to received communications. However, conventional systems struggle to consider the reliability of communication content and the emotional state of the user, sometimes resulting in inappropriate responses. To solve this problem, an automated response system that analyzes communication content and considers the emotional state of the user is necessary.
[0419] The identification process performed by the identification processing unit 290 of the data processing device 12 in other embodiments is realized by the following means.
[0420] In this invention, the server includes means for switching received communications from human response to automated artificial intelligence response, means for acquiring voice data and converting it to text using speech recognition technology, and means for analyzing the texted data using natural language processing technology and generating prompts for determining perjury. This enables evaluation of the reliability of the communication content and appropriate responses that take into account the user's emotional state.
[0421] "Artificial intelligence automated response" refers to a system that processes received communications on behalf of humans, analyzes voice data, and generates appropriate responses.
[0422] "Speech recognition technology" is a technology that converts speech data into text data, and is used to make speech information into an analyzable format.
[0423] "Natural language processing technology" is a technique for analyzing text data to understand its meaning and context, and is used to evaluate the reliability of communication content.
[0424] A "generative AI model" is an artificial intelligence model that generates text based on input prompts and is used for perjury assessment and generating action suggestions.
[0425] A "prompt" is an input sentence used to instruct a generative AI model on a specific task, and it includes instructions for the model to produce an appropriate output.
[0426] An "emotion analysis engine" is a technology that analyzes a user's emotional state and adjusts the response based on the results.
[0427] This invention is a system that receives communications and takes appropriate action using artificial intelligence automated response. When the server receives a communication, it switches from human response to AI automated response based on pre-configured conditions. The voice data is converted into text data using the Google Cloud Speech-to-Text API. This conversion makes the voice information into a format that can be analyzed.
[0428] The server analyzes the converted text data using natural language processing techniques. Specifically, it uses OpenAI's GPT model and evaluates perjury by inputting prompt sentences into a generative AI model. An example of a prompt sentence is, "Please evaluate the reliability of this statement."
[0429] Based on the perjury assessment, the server generates a prompt to suggest the next action. For example, it might generate a prompt such as, "Based on this statement, please suggest the next action to take," and input it into the AI model. The model then generates specific suggestions to present to the user.
[0430] Furthermore, the server uses an emotion analysis engine (e.g., IBM Watson Tone Analyzer) to analyze the user's emotional state. Based on the analysis results, it generates prompts to adjust the suggested actions. For example, it might generate a prompt such as, "If the user is feeling anxious, how should the suggestions be adjusted?" and input this into a generation AI model.
[0431] The terminal displays suggestions sent from the server to the user. The user interface is designed using HTML / CSS to present the suggestions in a visually clear and easy-to-understand manner. The user reviews the suggestions and chooses to accept or reject them.
[0432] This system automates the entire process from receiving communications to providing suggestions to the user, utilizing a generative AI model to make sophisticated judgments and suggestions. By generating and analyzing specific prompt sentences, it can provide appropriate responses to users.
[0433] The flow of specific processing in other embodiments will be explained using Figure 23.
[0434] Step 1:
[0435] When the server receives a communication, it switches from human to AI-powered automated response based on pre-configured conditions. The input is the received communication, and the output is the switch to the AI-powered automated response. This switch prepares the server for automatic processing of the communication content.
[0436] Step 2:
[0437] The server converts received audio data into text data using the Google Cloud Speech-to-Text API. The input is audio data, and the output is text data. Speech recognition technology is used to convert the audio information into a format that can be analyzed.
[0438] Step 3:
[0439] The server analyzes the transcribed data using natural language processing techniques. Specifically, it uses OpenAI's GPT model to generate a prompt message: "Please evaluate the reliability of this statement." The input is text data, and the output is the result of a perjury assessment. By inputting the prompt into the generating AI model, the reliability of the statement is evaluated.
[0440] Step 4:
[0441] The server generates a prompt to suggest the next action based on the result of its perjury assessment. The input is the result of the perjury assessment, and the output is the action suggestion. For example, by generating a prompt such as "Please suggest the next action to take based on this statement" and inputting it into the generation AI model, a specific suggestion will be generated.
[0442] Step 5:
[0443] The server analyzes the user's emotional state using an emotion analysis engine. The input is text data, and the output is the emotion analysis result. Based on the analysis result, it generates prompt sentences to adjust the suggested actions. For example, it generates a prompt such as "How should the suggestions be adjusted if the user is feeling anxious?" and inputs it into the generation AI model.
[0444] Step 6:
[0445] The terminal displays suggestions sent from the server to the user. The input is the action suggestion, and the output is the display on the user interface. The user interface is designed using HTML / CSS to present the suggestion content in a visually clear and easy-to-understand manner.
[0446] Step 7:
[0447] The user reviews the proposal displayed on the device and chooses to accept or reject it. The input is the displayed proposal, and the output is the user's choice. The user's choice influences the next action.
[0448] The specific processing unit 290 transmits the result of the specific processing to the smart device 14. In the smart device 14, the control unit 46A causes the output device 40 to output the result of the specific processing. The microphone 38B acquires audio indicating user input for the result of the specific processing. The control unit 46A transmits the audio data indicating user input acquired by the microphone 38B to the data processing device 12. In the data processing device 12, the specific processing unit 290 acquires the audio data.
[0449] Data generation model 58 is a form of so-called generative AI (Artificial Intelligence). An example of data generation model 58 is ChatGPT (registered trademark) (Internet search).<URL: https: / / openai.com / blog / chatgpt> Examples of generative AI include the following. The data generation model 58 is obtained by performing deep learning on a neural network. The data generation model 58 is input with prompts containing instructions, and with inference data such as audio data representing speech, text data representing text, and image data representing images. The data generation model 58 infers from the input inference data according to the instructions indicated by the prompts, and outputs the inference results in data formats such as audio data and text data. Here, inference refers to, for example, analysis, classification, prediction, and / or summarization.
[0450] Other examples of generative AI include Gemini® (registered trademark) (Internet search). <url: https: gemini.google.com ?hl="ja">) are some examples.
[0451] In the above embodiment, an example was given in which specific processing is performed by the data processing device 12, but the technology of this disclosure is not limited thereto, and the specific processing may also be performed by the smart device 14.
[0452] [Second Embodiment]
[0453] Figure 3 shows an example of the configuration of the data processing system 210 according to the second embodiment.
[0454] As shown in Figure 3, the data processing system 210 includes a data processing device 12 and smart glasses 214. An example of the data processing device 12 is a server.
[0455] The data processing device 12 comprises a computer 22, a database 24, and a communication interface 26. The computer 22 is an example of a "computer" related to the technology of this disclosure. The computer 22 comprises a processor 28, RAM 30, and storage 32. The processor 28, RAM 30, and storage 32 are connected to a bus 34. The database 24 and the communication interface 26 are also connected to the bus 34. The communication interface 26 is connected to a network 54. An example of the network 54 is a WAN (Wide Area Network) and / or a LAN (Local Area Network).
[0456] The smart glasses 214 include a computer 36, a microphone 238, a speaker 240, a camera 42, and a communication interface 44. The computer 36 includes a processor 46, RAM 48, and storage 50. The processor 46, RAM 48, and storage 50 are connected to a bus 52. The microphone 238, speaker 240, and camera 42 are also connected to the bus 52.
[0457] The microphone 238 receives voice signals from the user 20 and receives instructions from the user 20. The microphone 238 captures the voice signals from the user 20, converts the captured voice into audio data, and outputs it to the processor 46. The speaker 240 outputs audio according to the instructions from the processor 46.
[0458] Camera 42 is a small digital camera equipped with an optical system including a lens, aperture, and shutter, and an image sensor such as a CMOS (Complementary Metal-Oxide-Semiconductor) image sensor or a CCD (Charge Coupled Device) image sensor, and captures images of the area around the user 20 (for example, an imaging range defined by a field of view equivalent to the width of a typical healthy person's field of vision).
[0459] Communication interface 44 is connected to network 54. Communication interfaces 44 and 26 are responsible for the exchange of various information between processor 46 and processor 28 via network 54. The exchange of various information between processor 46 and processor 28 using communication interfaces 44 and 26 is performed in a secure manner.
[0460] Figure 4 shows an example of the main functions of the data processing device 12 and the smart glasses 214. As shown in Figure 4, the data processing device 12 performs specific processing using the processor 28. The storage 32 stores the specific processing program 56.
[0461] The specific processing program 56 is an example of a "program" relating to the technology of this disclosure. The processor 28 reads the specific processing program 56 from the storage 32 and executes the read specific processing program 56 on the RAM 30. The specific processing is realized by the processor 28 operating as a specific processing unit 290 in accordance with the specific processing program 56 executed on the RAM 30.
[0462] The storage 32 stores the data generation model 58 and the emotion identification model 59. The data generation model 58 and the emotion identification model 59 are used by the identification processing unit 290.
[0463] In the smart glasses 214, the processor 46 performs the reception output processing. The storage 50 stores the reception output program 60. The processor 46 reads the reception output program 60 from the storage 50 and executes the read reception output program 60 on the RAM 48. The reception output processing is realized by the processor 46 operating as a control unit 46A according to the reception output program 60 executed on the RAM 48.
[0464] Next, the identification process performed by the identification processing unit 290 of the data processing device 12 will be described.
[0465] "Example of form 1"
[0466] One embodiment of the present invention provides a means for a human to switch to an AI automated response system when a phone call comes in. Specifically, when a phone call comes in, the person answering the call can switch to the AI automated response system by pressing a specific button.
[0467] "Example of form 2"
[0468] Furthermore, the AI automated response system saves the conversation content and summarizes the potential for perjury in the other party's statements. Specifically, the AI automated response system transcribes the other party's statements into text and stores that text in a database. It also analyzes the text of the other party's statements to determine whether they are potentially perjury.
[0469] "Example of form 3"
[0470] Furthermore, the AI automated response will suggest the next action. Specifically, based on the result of the assessment of the perjury potential of the other party's statement, the AI automated response will suggest what action should be taken next. For example, if there is a high probability that the other party's statement is perjury, the AI automated response will suggest reporting it to the police.
[0471] The following describes the processing flow for each example of the form.
[0472] "Example of form 1"
[0473] Step 1: When a phone call comes in, the person answering the phone presses a specific button.
[0474] Step 2: When a specific button is pressed, the system switches to AI automated response.
[0475] Step 3: The AI automated response system answers the call and starts the conversation.
[0476] "Example of form 2"
[0477] Step 1: The AI automated response system transcribes the conversation into text.
[0478] Step 2: Save the transcribed conversation content to the database.
[0479] Step 3: The AI automated response analyzes the saved text to determine whether the other party's statement is perjury.
[0480] "Example of form 3"
[0481] Step 1: The AI automated response system determines the next action based on its assessment of the perjury potential of the other party's statement.
[0482] Step 2: The AI automated response system suggests the next action and notifies the person who answered the call.
[0483] Step 3: The person who received the call takes action based on the AI automated response's suggestions. For example, they might call the police.
[0484] (Example 1)
[0485] Next, we will describe Example 1 of Form Example 1. In the following description, the data processing device 12 will be referred to as the "server" and the smart glasses 214 will be referred to as the "terminal".
[0486] Traditional communication systems required human intervention upon receiving a call, leading to challenges in response efficiency and accuracy. Furthermore, the lack of features to assess the reliability of the caller's statements and suggest subsequent actions placed a significant burden on users.
[0487] The identification process performed by the identification processing unit 290 of the data processing device 12 in Example 1 is realized by the following means.
[0488] In this invention, the server includes means for switching received communications from a human to an automated response system, means for the automated response system to record the content of the conversation and evaluate the reliability of the other party's statements, and means for the automated response system to propose the next action. This enables efficient communication responses and reduces the burden on the user.
[0489] "A means of switching received communications from a human to an automated response system" refers to a function in which, when a communication device receives an incoming call, the user performs a specific operation to transfer the response from a human to an automated response system.
[0490] "A means by which an automated response system records the content of a conversation and evaluates the reliability of the other party's statements" refers to a function in which an automated response system saves the conversation in progress, analyzes the content of the statements, and determines their reliability.
[0491] "Means by which an automated response system suggests the next action" refers to a function in which an automated response system suggests the next action the user should take based on the content of the conversation.
[0492] "Means for detecting a specific operation and activating an automatic response device" refers to a function that detects an operation performed by a user, such as pressing a specific button on a communication device, and activates the automatic response device.
[0493] "A means by which an automated response device can offer a greeting at the start of a conversation" refers to a function in which an automated response device plays a pre-set greeting message to the other party at the start of communication.
[0494] A description of embodiments for carrying out this invention will be given.
[0495] The server receives incoming signals from communication devices and detects when a user performs a specific action. Specifically, when a user presses the "AI Response" button on their device, the server activates an automated response system. This automated response system uses AI software such as Google Dialogflow or Amazon Lex to record the conversation and evaluate the reliability of the other party's statements.
[0496] When the user presses a button on the terminal, the terminal sends that information to the server. Based on the received information, the server activates an automated response system and hands over the communication response to the AI. The automated response system plays a greeting message at the start of the conversation, such as, "Hello, this is the automated response system. How can I help you?"
[0497] As a concrete example, when a user receives a call, pressing the "AI Answer" button on their device activates the automated answering system, and the AI greets the caller. This process allows the user to leave the phone call to the AI.
[0498] An example of a prompt to input into a generating AI model is, "Please tell me the specific steps to switch to AI automatic answering when a phone call comes in." This prompt allows the AI to provide detailed information about how the system works and how to configure it.
[0499] The flow of the specific processing in Example 1 will be explained using Figure 11.
[0500] Step 1:
[0501] The terminal detects when a communication device receives an incoming signal. The input is the incoming signal from the communication device. The terminal sends this signal to the server, notifying it of the incoming call. The output is an incoming call notification sent to the server.
[0502] Step 2:
[0503] The user presses the "AI Response" button on the device. The input is the user's button press. The device detects this action and notifies the server that the button has been pressed. The output is a button press notification sent to the server.
[0504] Step 3:
[0505] The server receives a button press notification from the terminal. The input is the notification from the terminal. Based on this notification, the server generates an instruction to activate the automated response system. The output is an instruction to activate the automated response system.
[0506] Step 4:
[0507] The automated response system starts operating upon receiving a startup command from the server. Its input is the startup command from the server. At the start of the conversation, the automated response system plays a greeting message such as, "Hello, this is the automated response system. How can I help you?" Its output is a greeting message for the recipient.
[0508] Step 5:
[0509] The automated response system records the content of the conversation with the caller and evaluates the reliability of the caller's statements. The input is the caller's statements. The automated response system analyzes the statements and processes the data to evaluate their reliability. The output is the reliability evaluation result.
[0510] (Application Example 1)
[0511] Next, we will describe Application Example 1 of Form Example 1. In the following description, the data processing device 12 will be referred to as the "server," and the smart glasses 214 will be referred to as the "terminal."
[0512] In today's communication environment, suspicious communications and fraudulent phone calls are on the rise, requiring a swift and effective response. However, traditional methods require human recipients to handle all inquiries, placing a heavy burden on them. Furthermore, there is a risk of errors due to misjudgments. To address these challenges, the introduction of an automated response system utilizing artificial intelligence is necessary.
[0513] The specific processing performed by the specific processing unit 290 of the data processing device 12 in Application Example 1 is realized by the following means.
[0514] In this invention, the server includes means for switching received communications from human to artificial intelligence automated response, means for the artificial intelligence automated response to save the conversation content and summarize the perjury of the other party's statements, means for the artificial intelligence automated response to propose the next action, and means for the artificial intelligence automated response to respond to suspicious communications based on pre-set instructions. This enables a rapid and accurate response to suspicious communications.
[0515] "Received communications" refers to the transmission of information from external sources, such as phone calls and messages.
[0516] "Artificial intelligence automated response" refers to a system that uses artificial intelligence technology to automatically generate responses and send replies to the communication partner.
[0517] "Saving conversation content" refers to recording information exchanged during communication so that it can be referenced later.
[0518] "Summarizing the perjury" refers to evaluating the veracity of the other party's statements and identifying any questionable points.
[0519] "Suggesting the next course of action" means suggesting the appropriate course of action to take based on the content of the communication.
[0520] "Suspicious communications" refer to the transmission of information that differs from normal communications and may be fraudulent or illegal.
[0521] "Pre-set instructions" refer to guidelines for responses and actions that artificial intelligence should follow in specific situations.
[0522] The system for implementing this invention mainly consists of a server and a terminal. The server runs a program to switch received communications from human interaction to automated artificial intelligence responses. Specifically, when a communication is received, the server activates an automated artificial intelligence response based on instructions from the terminal and saves the conversation content. The saved data is used to evaluate the perjury potential of the other party's statements.
[0523] The server uses a generative AI model to propose the next action based on the content of the communication. This uses generative AI models such as OpenAI's GPT-3. The server responds to suspicious communications based on pre-configured instructions. This allows users to respond quickly and accurately to suspicious communications.
[0524] As a concrete example, when a device receives a suspicious call, the user presses a specific button, and the server activates an AI automated response system that says, "This is the security service. How can I help you?" An example of the prompt might be, "Please provide an example of how to respond to a suspicious call. Please emphasize that this is a security service response."
[0525] The flow of a specific process in Application Example 1 will be explained using Figure 12.
[0526] Step 1:
[0527] The terminal receives a communication. If the user determines the communication is suspicious, they press a specific button on the terminal. This causes the terminal to send a signal to the server requesting the activation of the artificial intelligence automated response system. The input is the received communication, and the output is the activation request signal to the server.
[0528] Step 2:
[0529] The server receives an activation request signal from the terminal and activates an artificial intelligence automated response system. The server uses a generative AI model to generate a response based on the prompt. The input is the activation request signal and prompt from the terminal, and the output is the generated response. Specifically, the server inputs the prompt "Please provide an example of how to respond to a suspicious call. Emphasize that this is a security service response." into the generative AI model and generates a response.
[0530] Step 3:
[0531] The server sends the generated response to the terminal. The terminal then sends this response to the communication partner. The input is the generated response from the server, and the output is the response sent to the communication partner. Specifically, the terminal sends a response to the communication partner saying, "This is the security service. How can I help you?"
[0532] Step 4:
[0533] The server stores the content of the communication and evaluates the perjury potential of the other party's statements. The input is the content of the communication, and the output is the result of the perjury evaluation. Specifically, the server records the conversation content in a database and analyzes the reliability of the statements using natural language processing techniques.
[0534] Step 5:
[0535] Based on the results of the perjury assessment, the server proposes the next action. The input is the result of the perjury assessment, and the output is the proposed action. Specifically, the server generates a suggestion such as "This communication is suspicious. Further verification is required," and sends it to the terminal.
[0536] (Example 2)
[0537] Next, we will describe Example 2 of Form Example 2. In the following description, the data processing device 12 will be referred to as the "server" and the smart glasses 214 will be referred to as the "terminal".
[0538] Conventional automated response systems have shortcomings, such as insufficient recording of conversation content and evaluation of the reliability of statements, making them unable to appropriately suggest the next course of action for the user. Furthermore, because fraud detection based on perjury assessment is not performed, users are vulnerable to fraudulent activity.
[0539] The identification process performed by the identification processing unit 290 of the data processing device 12 in Example 2 is realized by the following means.
[0540] In this invention, the server includes means for acquiring voice input, means for converting voice data into text, and means for storing the text data in a storage device. This enables accurate recording of conversation content and evaluation of the perjury potential of statements.
[0541] "Means for acquiring voice input" refers to a device or method for receiving a user's voice and processing it as digital data.
[0542] "Means for converting audio data to text" refers to a technology or device that analyzes an audio signal and converts it into a corresponding string of characters.
[0543] "Means for storing text data in a storage device" refers to a method or apparatus for recording converted text data in a database or other storage medium.
[0544] "Means for analyzing text data and determining the perjury of a statement" refers to a technology or device for processing text data and evaluating the reliability and truthfulness of its content.
[0545] "Means for outputting analysis results" refers to a method or apparatus for presenting the results of text data analysis to the user.
[0546] "Means for switching received communications from human to automated responses" refers to a method or device for transferring a human response to an automated response system upon receiving a communication.
[0547] "Means of suggesting the next action through automated response" refers to a technology or device that suggests the next action a user should take based on analysis results.
[0548] This invention is a system that determines the perjury of a statement by acquiring voice input, converting it into text data, and analyzing it. A specific embodiment of this system is shown below.
[0549] The user speaks aloud through the device's microphone. The device captures this audio as digital data and sends it to the server. The server uses speech recognition software to convert the audio data into text. High-precision text conversion is achieved by using speech recognition technologies such as the Google Cloud Speech-to-Text API.
[0550] The converted text data is stored on a storage device by the server. A database management system is used for storage, and metadata such as the message timestamp and user ID are also recorded.
[0551] Next, the server analyzes the text data using a natural language processing library. Specifically, it performs grammatical analysis using spaCy and then scrutinizes the content of the statements using OpenAI's GPT generative AI model. This allows it to determine the perjury potential of the statements and detect inconsistencies and unnatural points.
[0552] The analysis results are sent from the server to the terminal and presented to the user. The user can then review the results on the terminal screen and decide on their next course of action.
[0553] For example, if a user says, "It didn't rain yesterday," the system transcribes the statement into text and saves it to a database. Then, it refers to weather information and compares it with the actual weather to determine whether the statement is false.
[0554] An example of a prompt would be: "Assess the likelihood that the following statement is perjury: 'It didn't rain yesterday.' Please take actual weather data into consideration."
[0555] In this way, the system can automate the entire process from voice input to perjury detection, enabling it to provide users with quick and accurate information.
[0556] The flow of the specific processing in Example 2 will be explained using Figure 13.
[0557] Step 1:
[0558] The user speaks aloud through the device's microphone. The device captures this audio as digital data. The input is the user's voice, and the output is digital audio data. The device prompts the user to press the "Start Recording" button to begin voice input.
[0559] Step 2:
[0560] The device sends the captured audio data to the server. The server uses speech recognition software to convert the audio data into text. The input is digital audio data, and the output is text data. The server uses the Google Cloud Speech-to-Text API to convert the audio to text.
[0561] Step 3:
[0562] The server stores the converted text data in storage. The input is text data, and the output is text data stored in the database. The server executes SQL queries to insert the text data and associated metadata into the database.
[0563] Step 4:
[0564] The server analyzes stored text data to determine the perjury potential of a statement. The input is text data obtained from a database, and the output is the result of the perjury evaluation. The server performs grammatical analysis using the natural language processing library spaCy and scrutinizes the content of the statement using OpenAI's GPT generative AI model.
[0565] Step 5:
[0566] The server sends the analysis results to the terminal and presents them to the user. The input is the result of the perjury assessment, and the output is the result displayed on the terminal's screen. The terminal displays the analysis results on the screen, and the user confirms the results.
[0567] (Application Example 2)
[0568] Next, we will describe application example 2 of form example 2. In the following description, the data processing device 12 will be referred to as the "server" and the smart glasses 214 as the "terminal".
[0569] In today's communication environment, real-time evaluation of conversation reliability is crucial. However, conventional technologies have struggled to efficiently determine the perjury potential of statements and present this information to users immediately. Therefore, new methods are needed to improve conversation reliability.
[0570] The specific processing performed by the specific processing unit 290 of the data processing device 12 in Application Example 2 is realized by the following means.
[0571] In this invention, the server includes means for switching received communications from human to artificial intelligence automated response, means for the artificial intelligence automated response to save the conversation content and summarize the perjury potential of the other party's statements, and means for the artificial intelligence automated response to suggest the next action. This makes it possible to evaluate the reliability of the conversation in real time and present it to the user immediately.
[0572] "Communication" refers to the act or means of sending and receiving voice or data.
[0573] "Artificial intelligence automated response" refers to a system that uses artificial intelligence technology to automatically generate responses and engage in conversation.
[0574] "Conversation content" refers to the entirety of the statements and information exchanged during communication.
[0575] "Perjury potential" is an indicator of the degree to which a statement is likely to be false.
[0576] "Speech recognition technology" is a technology that converts speech into text.
[0577] "Natural language processing technology" is a technology used to analyze text data and understand its meaning and intent.
[0578] A "display device" is a device used to present information visually.
[0579] "User" refers to an individual or group that uses the system.
[0580] "Action" refers to the next step or action suggested by the system.
[0581] The system for carrying out this invention includes a server and a terminal. The server has the function of switching received communications from human to artificial intelligence automated responses. The terminal uses speech recognition technology to transcribe speech during communication into text in real time. Specifically, it uses the Google Cloud Speech-to-Text API to convert speech to text.
[0582] The server analyzes the transcribed speech using natural language processing techniques. This analysis utilizes natural language processing libraries such as spaCy and Transformers. The analysis results in an evaluation of the perjury potential of the speech. This evaluation uses a pre-trained generative AI model (e.g., OpenAI GPT-3).
[0583] The evaluation results are displayed in real time on the terminal's display device. This allows users to instantly verify the reliability of the conversation. For example, if someone says "This product is the best-selling in the market" during a meeting, the server will refer to past data and market information to evaluate the reliability of the statement.
[0584] An example of a prompt for a generative AI model might be: "Evaluate the likelihood that this statement is true. Market data is as follows." Using this prompt, the model determines the reliability of the statement and provides the result to the user.
[0585] The flow of a specific process in Application Example 2 will be explained using Figure 14.
[0586] Step 1:
[0587] The device acquires audio during communication via its microphone. The input is audio data, which is converted into text data using the Google Cloud Speech-to-Text API. The output is the transcribed speech.
[0588] Step 2:
[0589] The server receives text data sent from the terminal. The input is text data, which is then analyzed using a natural language processing library (e.g., spaCy, Transformers). The purpose of the analysis is to understand the content of the statement and extract information to assess its perjury potential. The output is the analysis result.
[0590] Step 3:
[0591] The server sends a prompt to a generative AI model (e.g., OpenAI GPT-3) based on the analysis results. The input consists of the analysis results and the prompt text. An example of the prompt text is "Evaluate the likelihood that this statement is true. The market data is as follows." The model evaluates the reliability of the statement and outputs the result.
[0592] Step 4:
[0593] The server receives evaluation results from the generated AI model and sends them to the terminal. The input is the evaluation result, and the output is reliability evaluation information to be presented to the user.
[0594] Step 5:
[0595] The terminal displays reliability evaluation information received from the server on its display device. The input is reliability evaluation information, and the output is information that the user can visually confirm. This allows the user to instantly verify the reliability of the conversation.
[0596] (Example 3)
[0597] Next, we will describe Embodiment 3 of Embodiment Example 3. In the following description, the data processing device 12 will be referred to as the "server," and the smart glasses 214 will be referred to as the "terminal."
[0598] In today's communication environment, there is a need to quickly and accurately determine the veracity of what the other party says and to suggest appropriate actions. However, conventional systems have the challenge of having a complex process for evaluating the veracity of statements, making it difficult to clearly indicate the next action the user should take.
[0599] The identification process performed by the identification processing unit 290 of the data processing device 12 in Example 3 is realized by the following means.
[0600] In this invention, the server includes means for switching received communications from a human to an automated response system, means for the automated response system to store conversation information and evaluate the veracity of the other party's statements, and means for the automated response system to propose the next action based on the evaluation result. This makes it possible to quickly evaluate the veracity of statements and propose appropriate actions to the user.
[0601] "Received communications" refers to audio and data signals sent from external sources.
[0602] An "automatic response device" refers to a device that has the function of automatically responding to received communications.
[0603] "Conversation information" refers to the content of audio and text exchanged during communication.
[0604] "Evaluating the truthfulness of a statement or piece of information" refers to the process of judging the accuracy and reliability of such statements or information.
[0605] "Suggesting action" means indicating specific actions to take next based on the evaluation results.
[0606] A "communication terminal" refers to a device used by a user that has the function of performing communication.
[0607] "User" refers to a person who uses the system.
[0608] A description of embodiments for carrying out this invention will be given.
[0609] First, the user enters a prompt message using a communication terminal. This prompt message is used to determine the truthfulness of the other party's statement. As a concrete example, consider the prompt message, "Is this statement true?"
[0610] Next, the terminal sends the entered prompt message to the server. The server uses a generative AI model to analyze the received prompt message. This analysis utilizes natural language processing techniques, employing models such as "OpenAI GPT-3" and "Google BERT." The server uses these models to scrutinize the content of the utterance and evaluate its truthfulness.
[0611] Based on the evaluation results, the server suggests the next course of action. For example, if it is determined that the statement is highly likely to be perjury, the server will decide to "suggest reporting to the police."
[0612] The server then sends the suggested action to the terminal. The terminal displays the received suggestion to the user. The user can then use this suggestion to decide on their next action.
[0613] In this way, it becomes possible to quickly evaluate the veracity of statements and propose appropriate actions. The flow of the specific processing in Example 3 will be explained using Figure 15.
[0614] Step 1:
[0615] The user enters a prompt message using a communication terminal. For example, they might enter the prompt message, "Is this statement true?" This input serves as the basis for determining the truthfulness of the other party's statement.
[0616] Step 2:
[0617] The terminal sends the entered prompt message to the server. Here, the input is the prompt message, and the output is the data sent to the server. The terminal securely transmits the data using the HTTPS protocol.
[0618] Step 3:
[0619] The server inputs the received prompt text into a generating AI model. The input is the prompt text, and the output is the analysis result by the AI model. The server uses natural language processing models such as "OpenAI GPT-3" and "Google BERT" to analyze the content of the utterance and evaluate its truthfulness.
[0620] Step 4:
[0621] The server proposes the next course of action based on the analysis results. The input is the analysis results of the AI model, and the output is the proposed action. For example, if the server determines that there is a high probability that the statement is perjury, it will decide to take the action of "I suggest reporting this to the police."
[0622] Step 5:
[0623] The server sends the proposed action to the terminal. The input is the proposed action, and the output is the transmission of data to the terminal. The server securely transmits the data using the HTTPS protocol.
[0624] Step 6:
[0625] The terminal displays received suggestions to the user. The input is suggestions from the server, and the output is what is displayed to the user. The terminal displays the suggestions on the screen, allowing the user to decide on their next action.
[0626] (Application Example 3)
[0627] Next, we will describe application example 3 of form example 3. In the following description, the data processing device 12 will be referred to as the "server" and the smart glasses 214 as the "terminal".
[0628] In modern society, the spread of false information and fraudulent activities through communications is increasing, and there is a need for a swift and effective response. However, conventional systems have difficulty determining the falsehood of communications in real time and suggesting appropriate actions. As a result, the risk of users becoming involved in fraudulent activities is increasing.
[0629] The specific processing performed by the specific processing unit 290 of the data processing device 12 in Application Example 3 is realized by the following means.
[0630] In this invention, the server includes means for switching received communications from a human to an AI-powered automated response, means for the AI automated response to record the content of the conversation and evaluate the falsity of the other party's statements, and means for the AI automated response to propose the next action. This makes it possible to determine the falsity of the communication content in real time and propose appropriate actions to the user.
[0631] "Means for switching received communications from human response to AI-powered automated response" refers to technology that switches from a conventional human response to an automated response by artificial intelligence when communication begins.
[0632] "An artificial intelligence automated response system that records the content of conversations and evaluates the veracity of the other party's statements" refers to a technology in which artificial intelligence stores the content of communications, analyzes that content, and determines the truthfulness of the statements.
[0633] "An artificial intelligence automated response system that suggests the next action" refers to a technology that, based on analysis results, suggests the appropriate next action to the user.
[0634] "Means of converting dialogue into text information using speech recognition technology" refers to technology for converting spoken dialogue into text format.
[0635] "Methods for analyzing the possibility of falsehood using generative AI models" refers to techniques that utilize generative AI models to evaluate the falsehood of dialogue content.
[0636] "A means of displaying warnings based on analysis results and suggesting reporting to legal authorities if necessary" refers to a technology that issues warnings to users based on the results of a falsehood assessment and, if necessary, encourages them to report to legal authorities.
[0637] The system for implementing this invention mainly consists of a server and a terminal. The server has the function of switching received communications from human to AI-powered automated responses. The terminal converts the dialogue into text information using speech recognition technology. Specifically, it uses speech recognition software such as the Google Speech-to-Text API to convert speech data into text data.
[0638] The server uses a generative AI model to analyze the falsity of the converted text data. This analysis leverages generative AI models such as OpenAI's GPT-4. The generative AI model assesses the likelihood that the content is false based on the input text.
[0639] Based on the analysis results, the server displays a warning on the terminal and, if necessary, suggests reporting to legal authorities. This allows users to understand the falsity of communications in real time and take appropriate action.
[0640] As a concrete example, consider a scenario where a user is using this system during a business meeting. When someone says, "This product is 100% safe," the device uses speech recognition technology to convert the statement into text, which is then analyzed by a server using a generated AI model. If the analysis determines that the statement is highly false, the device displays a warning to the user: "This statement requires caution. Please review the details or consider legal action if necessary."
[0641] Examples of prompt statements to input into a generative AI model include the following:
[0642] "Assess the likelihood that the following statement is perjury: 'This product is 100% safe.'"
[0643] The flow of the specific processing in Application Example 3 will be explained using Figure 16.
[0644] Step 1:
[0645] The device receives voice input from the user. This voice input is audio data of the user's conversation. The device converts this audio data into text data using the Google Speech-to-Text API. The converted text data is then output.
[0646] Step 2:
[0647] The server receives text data sent from the terminal. The server inputs this text data into a generative AI model (e.g., GPT-4). The generative AI model analyzes the content of the text data and evaluates the likelihood of the statement being false. As a result of the evaluation, a score indicating the likelihood of falsehood is output.
[0648] Step 3:
[0649] The server determines whether to display a warning to the user based on the evaluation results from the generated AI model. If the falsehood score exceeds a certain threshold, the server sends a warning message to the terminal. The warning message states that caution is needed regarding what is said.
[0650] Step 4:
[0651] The terminal receives a warning message sent from the server and displays it to the user. The user reviews the displayed warning message and considers reporting it to legal authorities if necessary. This allows the user to understand the falsity of the communication content in real time and take appropriate action.
[0652] Furthermore, an emotion engine that estimates the user's emotions may be incorporated. That is, the identification processing unit 290 may use the emotion identification model 59 to estimate the user's emotions and perform identification processing using the user's emotions.
[0653] "Example of form 1"
[0654] One embodiment of the present invention provides a system in which an AI automated response system includes an emotion engine that recognizes the user's emotions. This emotion engine analyzes the user's emotions from the tone of their voice, manner of speaking, word choice, etc. For example, if the user is feeling angry or anxious, the emotion engine feeds that information back to the AI automated response system, which adjusts its response accordingly.
[0655] "Example of form 2"
[0656] The system also provides an emotion engine that adjusts the AI's automated responses based on the user's emotions. Specifically, if the system determines that the user is angry, the AI's automated response will be adjusted to use more polite language and a calmer tone. This allows for more appropriate responses that take the user's feelings into consideration.
[0657] "Example of form 3"
[0658] Furthermore, the system provides an emotion engine that adjusts the AI's automated response suggestions for the next action based on the user's emotions. For example, if the system determines that the user is feeling anxious, the AI automated response will suggest actions to alleviate that anxiety. This enables more appropriate action suggestions that are tailored to the user's emotions.
[0659] The following describes the processing flow for each example of the form.
[0660] "Example of form 1"
[0661] Step 1: The user receives a suspicious phone call.
[0662] Step 2: The user switches the phone to AI automated answering.
[0663] Step 3: The emotion engine analyzes the user's emotions from their voice.
[0664] Step 4: The AI automated response system adjusts its response based on feedback from the emotion engine.
[0665] "Example of form 2"
[0666] Step 1: The user receives a suspicious phone call.
[0667] Step 2: The user switches the phone to AI automated answering.
[0668] Step 3: The emotion engine analyzes the user's emotions from their voice.
[0669] Step 4: The AI automated response system adjusts its response based on feedback from the emotion engine.
[0670] "Example of form 3"
[0671] Step 1: The user receives a suspicious phone call.
[0672] Step 2: The user switches the phone to AI automated answering.
[0673] Step 3: The emotion engine analyzes the user's emotions from their voice.
[0674] Step 4: The AI automated response system adjusts its next action suggestion based on feedback from the emotion engine.
[0675] (Example 1)
[0676] Next, we will describe Example 1 of Form Example 1. In the following description, the data processing device 12 will be referred to as the "server" and the smart glasses 214 will be referred to as the "terminal".
[0677] Traditional communication systems required recipients to respond manually, making it difficult to provide appropriate responses that reflected emotions. Furthermore, they lacked automated processes for saving conversation content and determining perjury, and also lacked features to suggest the next course of action.
[0678] The identification process performed by the identification processing unit 290 of the data processing device 12 in Example 1 is realized by the following means.
[0679] In this invention, the server includes means for switching received communications from a human to an AI-powered automated response, means for the AI automated response to acquire voice data and perform emotion analysis, and means for the AI automated response to adjust the response content based on the emotion analysis. This enables appropriate responses according to emotions without the recipient having to respond manually.
[0680] "Received communications" refers to audio information sent to the terminal from an external source, such as phone calls or voice messages.
[0681] "Artificial intelligence automated response" refers to a program or system that automatically generates and transmits a response to an incoming communication.
[0682] "Audio data" refers to data that represents audio information obtained from received communications in digital format.
[0683] "Emotional analysis" refers to the process of analyzing audio data to identify the speaker's emotional state.
[0684] "Adjusting the response content" refers to changing the content and tone of the response generated by the AI automated response system based on the results of sentiment analysis.
[0685] "Saving conversation content" refers to recording audio information exchanged during communication so that it can be referenced later.
[0686] "Summarizing the perjury" refers to analyzing the content of a conversation and organizing the information needed to evaluate the veracity of the statements.
[0687] "Suggesting the next action" refers to the AI automated response system suggesting the appropriate course of action to take based on the content and analysis results of the communication.
[0688] This invention provides a system that efficiently responds to incoming communications using artificial intelligence-based automated response. Specific embodiments of this system are described below.
[0689] When the server receives a communication, it first activates an AI-powered automated response system. This automated response system uses speech recognition software to acquire voice data and perform sentiment analysis. Specifically, it uses a "voice capture system" as the speech recognition software and a "sentiment analysis engine" for sentiment analysis. This analyzes the user's voice tone, speaking style, and word choice to identify their emotions.
[0690] The device switches to an AI-powered automated response system when the user presses the "AI Response Switch Button" upon receiving a communication. This action sends a signal to the server, initiating the AI automated response. The server adjusts the response based on the sentiment analysis results, providing the user with an appropriate response.
[0691] For example, when a user receives a call and presses the "AI response switch button," the server immediately acquires the voice data and performs sentiment analysis. If the analysis indicates that the user is feeling anxious, the AI automated response will respond in a gentle tone, saying, "Please don't worry. How can I help you?" This response is recorded on the server and used to improve the generated AI model.
[0692] An example of a prompt sentence to input into a generative AI model might be, "If a user is angry on the phone, how should the AI automated response system respond?" By using this prompt sentence, the generative AI model can suggest an appropriate response.
[0693] The flow of the specific processing in Example 1 will be explained using Figure 17.
[0694] Step 1:
[0695] The terminal detects incoming communications. When the user presses the "AI response switch button" upon receiving a communication, the terminal sends a signal to the server requesting the activation of the artificial intelligence automatic response. The input for this step is the incoming communication signal, and the output is the activation request signal to the server.
[0696] Step 2:
[0697] When the server receives a startup request from the terminal, it activates an artificial intelligence automated response system. The server uses speech recognition software to obtain voice data from the communication. The input for this step is the startup request signal from the terminal, and the output is the obtained voice data.
[0698] Step 3:
[0699] The server sends the acquired audio data to the emotion analysis engine. The emotion analysis engine analyzes the audio data and identifies the user's emotions. The input to this step is the audio data, and the output is the analyzed emotion information. Specifically, it analyzes the tone of voice, speaking style, and word choice.
[0700] Step 4:
[0701] The server adjusts the content of the AI-powered automated response based on emotional information from the emotion analysis engine. The server uses a generative AI model to generate an appropriate response that matches the user's emotions. The input for this step is the analyzed emotional information, and the output is the adjusted response.
[0702] Step 5:
[0703] The server sends the refined response to the user. The user receives the response from the AI automated response system. The input for this step is the refined response, and the output is the response to the user. Specifically, this involves using a gentle tone and reassuring language in the response.
[0704] (Application Example 1)
[0705] Next, we will describe Application Example 1 of Form Example 1. In the following description, the data processing device 12 will be referred to as the "server," and the smart glasses 214 will be referred to as the "terminal."
[0706] Conventional communication systems made it difficult for recipients to quickly assess the urgency of a call and respond appropriately. Furthermore, they were unable to accurately analyze the caller's emotions and adjust their responses accordingly, resulting in inadequate responses, particularly in emergencies where a rapid and appropriate response was required.
[0707] The specific processing performed by the specific processing unit 290 of the data processing device 12 in Application Example 1 is realized by the following means.
[0708] In this invention, the server includes means for switching received communications from human response to automated artificial intelligence response, means for analyzing the caller's emotions using an emotion analysis engine and determining the urgency, and means for automatically executing an appropriate response according to the urgency. This enables a rapid and appropriate response based on the caller's emotions.
[0709] "Received communications" refers to the act of receiving information, including audio, such as phone calls and voice messages.
[0710] "Artificial intelligence automated response" refers to an artificial intelligence system that automatically generates responses to received communications and engages in dialogue.
[0711] "Saving conversation content" refers to the act of recording the content of a phone call or conversation so that it can be referenced later.
[0712] "Summarizing the perjury" is the act of determining whether the other party's statement is true or false and organizing the results of that determination.
[0713] "Suggesting the next course of action" means suggesting the next course of action to take based on the current situation.
[0714] An "emotion analysis engine" is a technology that analyzes emotions from a caller's voice and words to identify those emotions.
[0715] "Assessing the urgency" means evaluating the importance and urgency of a situation and determining the priority of responses.
[0716] "Automatically executing appropriate actions" means automatically taking the necessary actions based on the determined level of urgency.
[0717] The system for implementing this invention operates in a network environment including a server and terminals. The server has the function of switching received communications to an artificial intelligence-based automated response system and uses an emotion analysis engine to analyze the caller's voice. The emotion analysis engine analyzes the caller's tone of voice, speaking style, and word choice to identify emotions. Based on this, the server determines the urgency and automatically takes an appropriate action.
[0718] The terminal provides an interface for sending voice to the server, and the user can switch to an AI-powered automated response by pressing a specific button. The server uses the Python speech_recognition library as speech recognition software and a custom EmotionEngine for sentiment analysis. Furthermore, the generated response is provided by the AIResponse module.
[0719] For example, if a user is using a security company's terminal, they can switch to an AI-powered automated response system by pressing a button on the terminal when an emergency call comes in. The server analyzes the caller's statement, "Help me, someone is breaking into my house," and an emotion analysis engine detects a high level of fear. As a result, the server immediately issues an order to dispatch security personnel.
[0720] An example of a prompt message for a generative AI model is: "Analyze the caller's voice and identify their emotion. If the emotion is fear, immediate action is required."
[0721] The flow of a specific process in Application Example 1 will be explained using Figure 18.
[0722] Step 1:
[0723] When the user presses a button to switch the received communication to an AI-powered automated response, the terminal sends voice data to the server. The input is the user's voice, and the output is the transmission of voice data to the server. The terminal converts the voice into digital data and sends it to the server over the network.
[0724] Step 2:
[0725] The server converts the received audio data into text using the speech_recognition library. The input is audio data, and the output is text data. The server applies a speech recognition algorithm to extract text from the audio.
[0726] Step 3:
[0727] The server inputs text data into an emotion analysis engine to identify the caller's emotions. The input is text data, and the output is emotion information. The emotion analysis engine analyzes the content and word choice of the text to determine the emotion.
[0728] Step 4:
[0729] The server determines the urgency based on emotional information and decides on the appropriate response. The input is emotional information, and the output is a response instruction. If the emotion is fear or anxiety, the server determines that a rapid response is necessary and generates a response instruction.
[0730] Step 5:
[0731] The server executes the generated response instructions and takes specific actions as needed, such as dispatching security personnel. The input is the response instructions, and the output is the actions taken. The server notifies the relevant systems and services and implements the response according to the instructions.
[0732] (Example 2)
[0733] Next, we will describe Example 2 of Form Example 2. In the following description, the data processing device 12 will be referred to as the "server" and the smart glasses 214 will be referred to as the "terminal".
[0734] Conventional automated response systems simply record the content of conversations with users without adjusting responses based on the falsity of statements or the user's emotions. This results in a poor user experience and can lead to misunderstandings and dissatisfaction. Furthermore, the lack of features to suggest the next course of action makes it difficult for users to take appropriate action.
[0735] The identification process performed by the identification processing unit 290 of the data processing device 12 in Example 2 is realized by the following means.
[0736] In this invention, the server includes means for switching received communications from a human to an automated response system, means for the automated response system to record the content of the conversation and analyze the falsity of the other party's statements, and means for the automated response system to adjust the content of the response based on sentiment analysis. This makes it possible to provide appropriate responses that take into account the falsity of statements and the emotions of the user during conversations, and to further suggest the next action.
[0737] "Received communications" refers to audio and data sent from external sources, and the purpose is to switch these to an automated response system.
[0738] An "automated response device" refers to a device or system that automatically processes, records, analyzes, and responds to user interactions.
[0739] "Recording conversation content" refers to the act of saving conversations and communications with users as data, which can then be used for later analysis and reference.
[0740] "Analyzing the falsity of statements" refers to the process of analyzing the content of a user's statements and determining whether or not that content is based on facts.
[0741] "Emotion analysis" refers to a technology that reads emotions from a user's speech and voice, and identifies emotional states such as anger, joy, and sadness.
[0742] "Adjusting response content" refers to changing the content and tone of the response generated by the automated response system according to the analyzed emotions and situation.
[0743] "Suggesting the next action" refers to showing the user the appropriate next steps or actions based on their current situation.
[0744] This invention is a system that efficiently manages user interaction using an automated response device and improves the user experience. Specific embodiments of this system are described below.
[0745] The server switches incoming communications from a human to an automated response system. This involves receiving voice data over the communication network and converting it to text using speech recognition software. Specifically, it uses a "speech recognition API" for speech recognition.
[0746] The terminal receives the user's voice input and sends it to the server. The server converts the received voice data into text using a "speech recognition API" and stores that text in a database. A "database management system" is used for this database.
[0747] The server analyzes the stored text using a natural language processing library to determine whether the statement is false. This allows the server to verify whether the user's statement is based on facts.
[0748] Furthermore, the device uses a "sentiment analysis engine" to analyze the user's emotions from their voice and text. The server receives the results of the sentiment analysis and determines what emotion the user is experiencing, such as anger, joy, or sadness.
[0749] The server adjusts the content of the AI automated response based on the results of the sentiment analysis. For example, if it is determined that the user is angry, the AI automated response will be set to respond in a calm tone, such as, "I'm sorry. Please let me know how we can improve."
[0750] For example, if a user asks, "Does this product really work?", the server transcribes the question into text and stores it in the database. It then compares this question with past data and product information to provide accurate information.
[0751] Examples of prompts include, "Transcribe the user's utterance into text and generate a prompt to determine its falsity," and "Create a prompt that adjusts the AI's response based on the user's sentiment."
[0752] This system makes it possible to manage user interactions in a more accurate and emotionally sensitive manner.
[0753] The flow of the specific processing in Example 2 will be explained using Figure 19.
[0754] Step 1:
[0755] The device receives voice input from the user. For example, the user might say, "Does this product really work?" The device then sends this voice data to the server. The input is the user's voice data, and the output is the transfer of the voice data to the server.
[0756] Step 2:
[0757] The server converts the received audio data into text using a speech recognition API. The speech recognition API analyzes the audio signal and generates corresponding text data. The input is audio data, and the output is text data.
[0758] Step 3:
[0759] The server stores the transcribed conversation content in a database. A database management system is used to record the text data in the appropriate format. The input is text data, and the output is storage in the database.
[0760] Step 4:
[0761] The server analyzes the stored text using a natural language processing library to determine the falsity of the statements. The natural language processing library analyzes the text data and evaluates its falsity by comparing it with known facts. The input is text data, and the output is the result of the falsity evaluation.
[0762] Step 5:
[0763] The device uses an emotion analysis engine to analyze the user's emotions from their voice or text. The emotion analysis engine identifies the user's emotional state from the input data. The input is voice or text data, and the output is the result of the emotion analysis.
[0764] Step 6:
[0765] The server adjusts the content of the AI-generated response based on the results of the sentiment analysis. The server changes the tone and content of the response according to the user's emotions to generate an appropriate response. The input is the result of the sentiment analysis, and the output is the adjusted response content.
[0766] Step 7:
[0767] The server sends a pre-arranged response to the terminal and provides the response to the user in voice or text. The terminal receives the response from the server and relays it to the user. The input is the pre-arranged response, and the output is the provision of the response to the user.
[0768] (Application Example 2)
[0769] Next, we will describe application example 2 of form example 2. In the following description, the data processing device 12 will be referred to as the "server" and the smart glasses 214 as the "terminal".
[0770] In modern communications, judging the reliability of what the other party says is crucial, but this is difficult with conventional systems. Furthermore, there is a need to provide appropriate responses that respond to the user's emotions, but effective means to achieve this are lacking. Additionally, there is a demand for a function that issues warnings based on the assessment of perjury, but the technology to implement this is still immature.
[0771] The specific processing performed by the specific processing unit 290 of the data processing device 12 in Application Example 2 is realized by the following means.
[0772] In this invention, the server includes means for switching received communications from human to machine learning automated responses, means for the machine learning automated response to save the conversation content and summarize the perjury potential of the other party's statements, and means for the machine learning automated response to adjust the response content based on sentiment analysis. This makes it possible to determine the reliability of statements in communications and to provide appropriate responses that correspond to the user's emotions.
[0773] "Communication" refers to the act or process of sending and receiving information, and includes forms such as voice, text, and data.
[0774] "Machine learning automated response" refers to a system that uses machine learning technology to automatically generate responses, providing appropriate answers based on user input.
[0775] "Saving conversation content" means recording information exchanged during communication and making it available for later reference.
[0776] "Summarizing the perjury" means evaluating whether the other party's statements are true or false, and then organizing and presenting the results.
[0777] "Emotional analysis" is a technology that infers emotions from a user's words and actions and analyzes their emotional state.
[0778] "Adjusting the response content" means changing the content and tone of the response according to the user's emotions and situation.
[0779] "Issuing a warning" means providing a message to draw attention when certain conditions are met.
[0780] "Suggesting an action" means indicating to the user what they should do next.
[0781] The system for implementing this invention mainly consists of a server and a terminal. The server has the function of switching received communications from human input to automated machine learning responses. The terminal is responsible for receiving input from the user and sending it to the server.
[0782] The server uses natural language processing libraries (e.g., spaCy, NLTK) to analyze the received conversation content and determine the perjury potential of the other party's statements. Furthermore, it uses sentiment analysis APIs (e.g., IBM Watson Tone Analyzer) to analyze the user's emotions and adjust the response based on the results. This enables appropriate responses tailored to the user's emotions.
[0783] Furthermore, the server has a function to issue warnings based on its assessment of perjury, allowing it to alert users. This can improve the reliability of communications.
[0784] For example, if a user asks, "Is this transaction safe?", the server will refer to past data and evaluate the reliability of the statement. Furthermore, if the server determines that the user is feeling uneasy, it will respond in a calm tone, such as, "Please rest assured, this transaction is safe."
[0785] An example of a prompt for a generative AI model is: "Analyze the user's statements and determine their perjury potential. Also, analyze the user's emotions and generate an appropriate response."
[0786] The flow of a specific process in Application Example 2 will be explained using Figure 20.
[0787] Step 1:
[0788] The device receives voice input from the user. This voice input is converted into text data using a speech recognition API (e.g., Google Speech-to-Text). This text data then becomes the input for the next process.
[0789] Step 2:
[0790] The server analyzes the received text data using a natural language processing library (e.g., spaCy, NLTK). The analysis extracts the content of the statement and generates data for determining its perjury potential. This data then becomes the input for the next process.
[0791] Step 3:
[0792] The server performs a perjury assessment. Specifically, it compares the statement against past database data to evaluate its reliability. This evaluation result becomes the input for the next process.
[0793] Step 4:
[0794] The server uses a sentiment analysis API (e.g., IBM Watson Tone Analyzer) to analyze the user's emotions from text data. Based on the analysis results, the user's emotional state is identified. This emotional state becomes the input for the next process.
[0795] Step 5:
[0796] The server generates an appropriate response based on the perjury assessment results and emotional state. A generative AI model is used to create the response to the user. This response becomes the input for the next process.
[0797] Step 6:
[0798] The server issues a warning as needed based on the results of the perjury assessment. A warning message is generated and sent to the user.
[0799] Step 7:
[0800] The terminal displays responses and warning messages received from the server to the user. Based on this, the user can decide on their next action.
[0801] (Example 3)
[0802] Next, we will describe Embodiment 3 of Embodiment Example 3. In the following description, the data processing device 12 will be referred to as the "server," and the smart glasses 214 will be referred to as the "terminal."
[0803] Conventional AI-powered automated response systems could determine the perjury potential of a statement, but they lacked the ability to offer appropriate action suggestions that took the user's emotional state into account. Furthermore, their action suggestions based on the perjury potential were limited, resulting in insufficient responses to alleviate user anxiety.
[0804] The identification process performed by the identification processing unit 290 of the data processing device 12 in Example 3 is realized by the following means.
[0805] In this invention, the server includes means for switching received communications from human to artificial intelligence automated response, means for the artificial intelligence automated response to save the conversation content and determine the perjury of the other party's statements, means for analyzing the user's emotional state using an emotion analysis engine, and means for proposing appropriate actions based on the results of the perjury determination and the emotion analysis. This enables flexible and appropriate action suggestions that respond to the user's emotions.
[0806] "Means for switching received communications from human to AI-powered automated responses" refers to a function that, after a communication has been received by a human, switches that response to an automated response by artificial intelligence.
[0807] "An artificial intelligence automated response system that saves conversation content and determines whether the other party's statements are perjury" refers to a function in which artificial intelligence records the content of a conversation, analyzes that content, and evaluates the truthfulness of the statements.
[0808] "Methods for analyzing a user's emotional state using an emotion analysis engine" refers to technologies that infer emotions from a user's statements and actions and analyze that state.
[0809] "A means of suggesting appropriate actions based on the results of perjury assessment and sentiment analysis" refers to a function that suggests the next course of action, taking into account the truthfulness of a statement and the user's emotional state.
[0810] A description of embodiments for carrying out this invention will be given.
[0811] The server generates a program to switch the communication received from the user to an AI-powered automated response. This program runs using server hardware equipped with an NVIDIA GPU and a generative AI model using TensorFlow. The server receives the user's utterances as text data and analyzes their content using natural language processing techniques. The analysis includes data processing to determine the perjury potential of the utterances and data calculations using a sentiment analysis engine.
[0812] For example, if a user asks "Does this product really work?" through their device, the server uses a generative AI model to analyze the statement and extract key keywords. The server then compares the statement against a past database to evaluate its credibility. Furthermore, if the sentiment analysis engine determines that the user is feeling anxious, the server suggests actions such as, "This information may not be reliable. We recommend consulting a professional for further information."
[0813] An example of a prompt to be input to the generating AI model is, "Analyze the user's statement, determine its perjury potential, and suggest an appropriate action." This prompt allows the server to generate an appropriate response to the user's statement. The flow of specific processing in Example 3 will be explained using Figure 21.
[0814] Step 1:
[0815] The user inputs their message into the system via a terminal. The entered text data is sent to the server. The server receives this text data and prepares it for the next analysis step.
[0816] Step 2:
[0817] The server analyzes the received text data using natural language processing techniques. Specifically, it uses a generative AI model to understand the content of the utterance and extract important keywords and context. This analysis clarifies the intent and subject of the utterance. The input is the user's utterance, and the output is the analyzed keywords and contextual information.
[0818] Step 3:
[0819] The server determines the perjury potential of a statement based on the analysis results. It compares the statement against a past database to evaluate its credibility. This process involves database searches and comparison operations. The input is the analyzed keywords and contextual information, and the output is the result of the perjury determination.
[0820] Step 4:
[0821] The server uses an emotion analysis engine to analyze the user's emotional state. It infers emotions from the user's statements and identifies emotions such as anxiety and anger. The input is the user's statements, and the output is the evaluation result of the emotional state.
[0822] Step 5:
[0823] The server suggests the next course of action based on the results of the perjury assessment and sentiment analysis. For example, if there is a high probability that the statement is perjury and the user is feeling anxious, the server might suggest, "This information may not be reliable. We recommend consulting a professional." The input is the result of the perjury assessment and the evaluation of the emotional state, and the output is the suggested action.
[0824] Step 6:
[0825] The server sends the generated suggestions to the terminal and presents them to the user. The user can then decide on their next action based on these suggestions. The input is the suggested action, and the output is the content presented to the user.
[0826] (Application Example 3)
[0827] Next, we will describe application example 3 of form example 3. In the following description, the data processing device 12 will be referred to as the "server" and the smart glasses 214 as the "terminal".
[0828] In modern society, fraudulent activities and the spread of false information through communications are on the rise, making it difficult for users to protect themselves from these threats. Furthermore, while appropriate responses that respond to users' emotions are required, conventional systems have not been able to adequately achieve this. Therefore, ensuring user safety and peace of mind remains a challenge.
[0829] The specific processing performed by the specific processing unit 290 of the data processing device 12 in Application Example 3 is realized by the following means.
[0830] In this invention, the server includes means for switching received communications from an information processing device to an automated response device, means for the automated response device to record conversation information and evaluate the falsity of the other party's statements, means for the automated response device to propose the next action, and means for an emotion analysis device to evaluate the user's emotions and adjust the proposed action. This makes it possible for the user to evaluate the falsity of the communication content and receive an appropriate response according to their emotions.
[0831] An "information processing device" is a device that receives communications and, if necessary, switches to an automatic response device.
[0832] An "automatic response device" is a device that records received conversation information and has the function of evaluating the likelihood of the other party's statements being false.
[0833] "Conversational information" refers to voice or text data exchanged through communication.
[0834] "Falsehood" is an evaluation criterion that indicates the possibility that a statement is not based on facts.
[0835] An "emotion analysis device" is a device that evaluates the user's emotions and adjusts the suggested actions accordingly.
[0836] A "means of suggesting actions" refers to a function that indicates to the user the next action they should take based on the evaluation results.
[0837] To implement this invention, a server plays a central role. The server is equipped with an information processing device that receives communications and switches the received communications to an automatic response device. The automatic response device records conversation information and evaluates the falsity of the other party's statements. The evaluation is performed using natural language processing techniques, specifically using software such as Python or TensorFlow.
[0838] Furthermore, an emotion analysis device evaluates the user's emotions and adjusts the suggested actions accordingly. Emotion analysis utilizes emotion analysis APIs such as IBM Watson. This enables the provision of appropriate action suggestions based on the user's emotional state.
[0839] For example, if a user receives a potentially fraudulent phone call, the server will suggest, "This call may be fraudulent. Do you want to report it to the police?" If the user is feeling uneasy, it will also suggest, "Do you want to send a notification to a trusted contact?"
[0840] Examples of prompts generated using AI models include, "Analyze the content of this message and assess the likelihood of perjury," and "Analyze the user's sentiment and suggest appropriate action."
[0841] The flow of the specific processing in Application Example 3 will be explained using Figure 22.
[0842] Step 1:
[0843] The server receives communications from the terminal using an information processing device. Input can include voice or text data. The server sends this data to an automated response system for recording as conversation information.
[0844] Step 2:
[0845] The automated response system analyzes recorded conversation information. The input is the audio or text data recorded in step 1. The server uses natural language processing techniques to evaluate the falsity of the statements. Specifically, it analyzes the data using Python or TensorFlow and outputs a falsity score.
[0846] Step 3:
[0847] The server suggests the next action based on the falsehood score. The input is the falsehood score obtained in step 2. If the score is high, the server will generate a suggestion such as, "This call may be perjury. Do you want to report it to the police?"
[0848] Step 4:
[0849] The emotion analysis device evaluates the user's emotions. The input is the voice or text data received in step 1. The server analyzes the emotions using an emotion analysis API such as IBM Watson and outputs the emotional state.
[0850] Step 5:
[0851] The server adjusts the suggested actions based on the emotional state. The input is the emotional state obtained in step 4. If the user is feeling anxious, the server will generate a suggestion such as, "Do you want to send a notification to a trusted contact?"
[0852] Step 6:
[0853] The user receives suggestions from the server and selects the appropriate action. The input consists of the suggestions generated in steps 3 and 5. The user decides on an action based on the suggestions and sends feedback to the server as needed.
[0854] (Other examples)
[0855] Since this is the same as the specific processing described in the other embodiments of the first embodiment above, the explanation will be omitted.
[0856] The specific processing unit 290 transmits the result of the specific processing to the smart glasses 214. In the smart glasses 214, the control unit 46A causes the speaker 240 to output the result of the specific processing. The microphone 238 acquires audio indicating user input for the result of the specific processing. The control unit 46A transmits the audio data indicating user input acquired by the microphone 238 to the data processing unit 12. In the data processing unit 12, the specific processing unit 290 acquires the audio data.
[0857] The data generation model 58 is a form of so-called generative AI (Artificial Intelligence). One example of the data generation model 58 is ChatGPT (Internet Search).<URL: https: / / openai.com / blog / chatgpt> Examples of generative AI include the following. The data generation model 58 is obtained by performing deep learning on a neural network. The data generation model 58 is input with prompts containing instructions, and with inference data such as audio data representing speech, text data representing text, and image data representing images. The data generation model 58 infers from the input inference data according to the instructions indicated by the prompts, and outputs the inference results in data formats such as audio data and text data. Here, inference refers to, for example, analysis, classification, prediction, and / or summarization.
[0858] Other examples of generative AI include Gemini (Internet search <url: https: gemini.google.com ?hl="ja">) are some examples.
[0859] In the above embodiment, an example was given in which specific processing is performed by the data processing device 12, but the technology of this disclosure is not limited thereto, and the specific processing may also be performed by the smart glasses 214.
[0860] [Third Embodiment]
[0861] Figure 5 shows an example of the configuration of the data processing system 310 according to the third embodiment.
[0862] As shown in Figure 5, the data processing system 310 includes a data processing device 12 and a headset terminal 314. An example of the data processing device 12 is a server.
[0863] The data processing device 12 comprises a computer 22, a database 24, and a communication interface 26. The computer 22 is an example of a "computer" related to the technology of this disclosure. The computer 22 comprises a processor 28, RAM 30, and storage 32. The processor 28, RAM 30, and storage 32 are connected to a bus 34. The database 24 and the communication interface 26 are also connected to the bus 34. The communication interface 26 is connected to a network 54. An example of the network 54 is a WAN (Wide Area Network) and / or a LAN (Local Area Network).
[0864] The headset terminal 314 includes a computer 36, a microphone 238, a speaker 240, a camera 42, a communication interface 44, and a display 343. The computer 36 includes a processor 46, RAM 48, and storage 50. The processor 46, RAM 48, and storage 50 are connected to a bus 52. The microphone 238, speaker 240, camera 42, and display 343 are also connected to the bus 52.
[0865] The microphone 238 receives voice signals from the user 20 and receives instructions from the user 20. The microphone 238 captures the voice signals from the user 20, converts the captured voice into audio data, and outputs it to the processor 46. The speaker 240 outputs audio according to the instructions from the processor 46.
[0866] Camera 42 is a small digital camera equipped with an optical system including a lens, aperture, and shutter, and an image sensor such as a CMOS (Complementary Metal-Oxide-Semiconductor) image sensor or a CCD (Charge Coupled Device) image sensor, and captures images of the area around the user 20 (for example, an imaging range defined by a field of view equivalent to the width of a typical healthy person's field of vision).
[0867] Communication interface 44 is connected to network 54. Communication interfaces 44 and 26 are responsible for the exchange of various information between processor 46 and processor 28 via network 54. The exchange of various information between processor 46 and processor 28 using communication interfaces 44 and 26 is performed in a secure manner.
[0868] Figure 6 shows an example of the main functions of the data processing device 12 and the headset terminal 314. As shown in Figure 6, the data processing device 12 performs specific processing using the processor 28. The storage 32 stores the specific processing program 56.
[0869] The specific processing program 56 is an example of a "program" relating to the technology of this disclosure. The processor 28 reads the specific processing program 56 from the storage 32 and executes the read specific processing program 56 on the RAM 30. The specific processing is realized by the processor 28 operating as a specific processing unit 290 in accordance with the specific processing program 56 executed on the RAM 30.
[0870] The storage 32 stores the data generation model 58 and the emotion identification model 59. The data generation model 58 and the emotion identification model 59 are used by the identification processing unit 290.
[0871] In the headset terminal 314, the processor 46 performs the reception output processing. The storage 50 stores the reception output program 60. The processor 46 reads the reception output program 60 from the storage 50 and executes the read reception output program 60 on the RAM 48. The reception output processing is realized by the processor 46 operating as a control unit 46A according to the reception output program 60 executed on the RAM 48.
[0872] Next, the identification process performed by the identification processing unit 290 of the data processing device 12 will be described.
[0873] "Example of form 1"
[0874] One embodiment of the present invention provides a means for a human to switch to an AI automated response system when a phone call comes in. Specifically, when a phone call comes in, the person answering the call can switch to the AI automated response system by pressing a specific button.
[0875] "Example of form 2"
[0876] Furthermore, the AI automated response system saves the conversation content and summarizes the potential for perjury in the other party's statements. Specifically, the AI automated response system transcribes the other party's statements into text and stores that text in a database. It also analyzes the text of the other party's statements to determine whether they are potentially perjury.
[0877] "Example of form 3"
[0878] Furthermore, the AI automated response will suggest the next action. Specifically, based on the result of the assessment of the perjury potential of the other party's statement, the AI automated response will suggest what action should be taken next. For example, if there is a high probability that the other party's statement is perjury, the AI automated response will suggest reporting it to the police.
[0879] The following describes the processing flow for each example of the form.
[0880] "Example of form 1"
[0881] Step 1: When a phone call comes in, the person answering the phone presses a specific button.
[0882] Step 2: When a specific button is pressed, the system switches to AI automated response.
[0883] Step 3: The AI automated response system answers the call and starts the conversation.
[0884] "Example of form 2"
[0885] Step 1: The AI automated response system transcribes the conversation into text.
[0886] Step 2: Save the transcribed conversation content to the database.
[0887] Step 3: The AI automated response analyzes the saved text to determine whether the other party's statement is perjury.
[0888] "Example of form 3"
[0889] Step 1: The AI automated response system determines the next action based on its assessment of the perjury potential of the other party's statement.
[0890] Step 2: The AI automated response system suggests the next action and notifies the person who answered the call.
[0891] Step 3: The person who received the call takes action based on the AI automated response's suggestions. For example, they might call the police.
[0892] (Example 1)
[0893] Next, we will describe Embodiment 1 of Embodiment Example 1. In the following description, the data processing device 12 will be referred to as the "server," and the headset-type terminal 314 will be referred to as the "terminal."
[0894] Traditional communication systems required human intervention upon receiving a call, leading to challenges in response efficiency and accuracy. Furthermore, the lack of features to assess the reliability of the caller's statements and suggest subsequent actions placed a significant burden on users.
[0895] The identification process performed by the identification processing unit 290 of the data processing device 12 in Example 1 is realized by the following means.
[0896] In this invention, the server includes means for switching received communications from a human to an automated response system, means for the automated response system to record the content of the conversation and evaluate the reliability of the other party's statements, and means for the automated response system to propose the next action. This enables efficient communication responses and reduces the burden on the user.
[0897] "A means of switching received communications from a human to an automated response system" refers to a function in which, when a communication device receives an incoming call, the user performs a specific operation to transfer the response from a human to an automated response system.
[0898] "A means by which an automated response system records the content of a conversation and evaluates the reliability of the other party's statements" refers to a function in which an automated response system saves the conversation in progress, analyzes the content of the statements, and determines their reliability.
[0899] "Means by which an automated response system suggests the next action" refers to a function in which an automated response system suggests the next action the user should take based on the content of the conversation.
[0900] "Means for detecting a specific operation and activating an automatic response device" refers to a function that detects an operation performed by a user, such as pressing a specific button on a communication device, and activates the automatic response device.
[0901] "A means by which an automated response device can offer a greeting at the start of a conversation" refers to a function in which an automated response device plays a pre-set greeting message to the other party at the start of communication.
[0902] A description of embodiments for carrying out this invention will be given.
[0903] The server receives incoming signals from communication devices and detects when a user performs a specific action. Specifically, when a user presses the "AI Response" button on their device, the server activates an automated response system. This automated response system uses AI software such as Google Dialogflow or Amazon Lex to record the conversation and evaluate the reliability of the other party's statements.
[0904] When the user presses a button on the terminal, the terminal sends that information to the server. Based on the received information, the server activates an automated response system and hands over the communication response to the AI. The automated response system plays a greeting message at the start of the conversation, such as, "Hello, this is the automated response system. How can I help you?"
[0905] As a concrete example, when a user receives a call, pressing the "AI Answer" button on their device activates the automated answering system, and the AI greets the caller. This process allows the user to leave the phone call to the AI.
[0906] An example of a prompt to input into a generating AI model is, "Please tell me the specific steps to switch to AI automatic answering when a phone call comes in." This prompt allows the AI to provide detailed information about how the system works and how to configure it.
[0907] The flow of the specific processing in Example 1 will be explained using Figure 11.
[0908] Step 1:
[0909] The terminal detects when a communication device receives an incoming signal. The input is the incoming signal from the communication device. The terminal sends this signal to the server, notifying it of the incoming call. The output is an incoming call notification sent to the server.
[0910] Step 2:
[0911] The user presses the "AI Response" button on the device. The input is the user's button press. The device detects this action and notifies the server that the button has been pressed. The output is a button press notification sent to the server.
[0912] Step 3:
[0913] The server receives a button press notification from the terminal. The input is the notification from the terminal. Based on this notification, the server generates an instruction to activate the automated response system. The output is an instruction to activate the automated response system.
[0914] Step 4:
[0915] The automated response system starts operating upon receiving a startup command from the server. Its input is the startup command from the server. At the start of the conversation, the automated response system plays a greeting message such as, "Hello, this is the automated response system. How can I help you?" Its output is a greeting message for the recipient.
[0916] Step 5:
[0917] The automated response system records the content of the conversation with the caller and evaluates the reliability of the caller's statements. The input is the caller's statements. The automated response system analyzes the statements and processes the data to evaluate their reliability. The output is the reliability evaluation result.
[0918] (Application Example 1)
[0919] Next, we will describe Application Example 1 of Form Example 1. In the following description, the data processing device 12 will be referred to as the "server," and the headset-type terminal 314 will be referred to as the "terminal."
[0920] In today's communication environment, suspicious communications and fraudulent phone calls are on the rise, requiring a swift and effective response. However, traditional methods require human recipients to handle all inquiries, placing a heavy burden on them. Furthermore, there is a risk of errors due to misjudgments. To address these challenges, the introduction of an automated response system utilizing artificial intelligence is necessary.
[0921] The specific processing performed by the specific processing unit 290 of the data processing device 12 in Application Example 1 is realized by the following means.
[0922] In this invention, the server includes means for switching received communications from human to artificial intelligence automated response, means for the artificial intelligence automated response to save the conversation content and summarize the perjury of the other party's statements, means for the artificial intelligence automated response to propose the next action, and means for the artificial intelligence automated response to respond to suspicious communications based on pre-set instructions. This enables a rapid and accurate response to suspicious communications.
[0923] "Received communications" refers to the transmission of information from external sources, such as phone calls and messages.
[0924] "Artificial intelligence automated response" refers to a system that uses artificial intelligence technology to automatically generate responses and send replies to the communication partner.
[0925] "Saving conversation content" refers to recording information exchanged during communication so that it can be referenced later.
[0926] "Summarizing the perjury" refers to evaluating the veracity of the other party's statements and identifying any questionable points.
[0927] "Suggesting the next course of action" means suggesting the appropriate course of action to take based on the content of the communication.
[0928] "Suspicious communications" refer to the transmission of information that differs from normal communications and may be fraudulent or illegal.
[0929] "Pre-set instructions" refer to guidelines for responses and actions that artificial intelligence should follow in specific situations.
[0930] The system for implementing this invention mainly consists of a server and a terminal. The server runs a program to switch received communications from human interaction to automated artificial intelligence responses. Specifically, when a communication is received, the server activates an automated artificial intelligence response based on instructions from the terminal and saves the conversation content. The saved data is used to evaluate the perjury potential of the other party's statements.
[0931] The server uses a generative AI model to propose the next action based on the content of the communication. This uses generative AI models such as OpenAI's GPT-3. The server responds to suspicious communications based on pre-configured instructions. This allows users to respond quickly and accurately to suspicious communications.
[0932] As a concrete example, when a device receives a suspicious call, the user presses a specific button, and the server activates an AI automated response system that says, "This is the security service. How can I help you?" An example of the prompt might be, "Please provide an example of how to respond to a suspicious call. Please emphasize that this is a security service response."
[0933] The flow of a specific process in Application Example 1 will be explained using Figure 12.
[0934] Step 1:
[0935] The terminal receives a communication. If the user determines the communication is suspicious, they press a specific button on the terminal. This causes the terminal to send a signal to the server requesting the activation of the artificial intelligence automated response system. The input is the received communication, and the output is the activation request signal to the server.
[0936] Step 2:
[0937] The server receives an activation request signal from the terminal and activates an artificial intelligence automated response system. The server uses a generative AI model to generate a response based on the prompt. The input is the activation request signal and prompt from the terminal, and the output is the generated response. Specifically, the server inputs the prompt "Please provide an example of how to respond to a suspicious call. Emphasize that this is a security service response." into the generative AI model and generates a response.
[0938] Step 3:
[0939] The server sends the generated response to the terminal. The terminal then sends this response to the communication partner. The input is the generated response from the server, and the output is the response sent to the communication partner. Specifically, the terminal sends a response to the communication partner saying, "This is the security service. How can I help you?"
[0940] Step 4:
[0941] The server stores the content of the communication and evaluates the perjury potential of the other party's statements. The input is the content of the communication, and the output is the result of the perjury evaluation. Specifically, the server records the conversation content in a database and analyzes the reliability of the statements using natural language processing techniques.
[0942] Step 5:
[0943] Based on the results of the perjury assessment, the server proposes the next action. The input is the result of the perjury assessment, and the output is the proposed action. Specifically, the server generates a suggestion such as "This communication is suspicious. Further verification is required," and sends it to the terminal.
[0944] (Example 2)
[0945] Next, we will describe Example 2 of the morphological example. In the following description, the data processing device 12 will be referred to as the "server," and the headset-type terminal 314 will be referred to as the "terminal."
[0946] Conventional automated response systems have shortcomings, such as insufficient recording of conversation content and evaluation of the reliability of statements, making them unable to appropriately suggest the next course of action for the user. Furthermore, because fraud detection based on perjury assessment is not performed, users are vulnerable to fraudulent activity.
[0947] The identification process performed by the identification processing unit 290 of the data processing device 12 in Example 2 is realized by the following means.
[0948] In this invention, the server includes means for acquiring voice input, means for converting voice data into text, and means for storing the text data in a storage device. This enables accurate recording of conversation content and evaluation of the perjury potential of statements.
[0949] "Means for acquiring voice input" refers to a device or method for receiving a user's voice and processing it as digital data.
[0950] "Means for converting audio data to text" refers to a technology or device that analyzes an audio signal and converts it into a corresponding string of characters.
[0951] "Means for storing text data in a storage device" refers to a method or apparatus for recording converted text data in a database or other storage medium.
[0952] "Means for analyzing text data and determining the perjury of a statement" refers to a technology or device for processing text data and evaluating the reliability and truthfulness of its content.
[0953] "Means for outputting analysis results" refers to a method or apparatus for presenting the results of text data analysis to the user.
[0954] "Means for switching received communications from human to automated responses" refers to a method or device for transferring a human response to an automated response system upon receiving a communication.
[0955] "Means of suggesting the next action through automated response" refers to a technology or device that suggests the next action a user should take based on analysis results.
[0956] This invention is a system that determines the perjury of a statement by acquiring voice input, converting it into text data, and analyzing it. A specific embodiment of this system is shown below.
[0957] The user speaks aloud through the device's microphone. The device captures this audio as digital data and sends it to the server. The server uses speech recognition software to convert the audio data into text. High-precision text conversion is achieved by using speech recognition technologies such as the Google Cloud Speech-to-Text API.
[0958] The converted text data is stored on a storage device by the server. A database management system is used for storage, and metadata such as the message timestamp and user ID are also recorded.
[0959] Next, the server analyzes the text data using a natural language processing library. Specifically, it performs grammatical analysis using spaCy and then scrutinizes the content of the statements using OpenAI's GPT generative AI model. This allows it to determine the perjury potential of the statements and detect inconsistencies and unnatural points.
[0960] The analysis results are sent from the server to the terminal and presented to the user. The user can then review the results on the terminal screen and decide on their next course of action.
[0961] For example, if a user says, "It didn't rain yesterday," the system transcribes the statement into text and saves it to a database. Then, it refers to weather information and compares it with the actual weather to determine whether the statement is false.
[0962] An example of a prompt would be: "Assess the likelihood that the following statement is perjury: 'It didn't rain yesterday.' Please take actual weather data into consideration."
[0963] In this way, the system can automate the entire process from voice input to perjury detection, enabling it to provide users with quick and accurate information.
[0964] The flow of the specific processing in Example 2 will be explained using Figure 13.
[0965] Step 1:
[0966] The user speaks aloud through the device's microphone. The device captures this audio as digital data. The input is the user's voice, and the output is digital audio data. The device prompts the user to press the "Start Recording" button to begin voice input.
[0967] Step 2:
[0968] The device sends the captured audio data to the server. The server uses speech recognition software to convert the audio data into text. The input is digital audio data, and the output is text data. The server uses the Google Cloud Speech-to-Text API to convert the audio to text.
[0969] Step 3:
[0970] The server stores the converted text data in storage. The input is text data, and the output is text data stored in the database. The server executes SQL queries to insert the text data and associated metadata into the database.
[0971] Step 4:
[0972] The server analyzes stored text data to determine the perjury potential of a statement. The input is text data obtained from a database, and the output is the result of the perjury evaluation. The server performs grammatical analysis using the natural language processing library spaCy and scrutinizes the content of the statement using OpenAI's GPT generative AI model.
[0973] Step 5:
[0974] The server sends the analysis results to the terminal and presents them to the user. The input is the result of the perjury assessment, and the output is the result displayed on the terminal's screen. The terminal displays the analysis results on the screen, and the user confirms the results.
[0975] (Application Example 2)
[0976] Next, we will describe application example 2 of form example 2. In the following description, the data processing device 12 will be referred to as the "server," and the headset-type terminal 314 will be referred to as the "terminal."
[0977] In today's communication environment, real-time evaluation of conversation reliability is crucial. However, conventional technologies have struggled to efficiently determine the perjury potential of statements and present this information to users immediately. Therefore, new methods are needed to improve conversation reliability.
[0978] The specific processing performed by the specific processing unit 290 of the data processing device 12 in Application Example 2 is realized by the following means.
[0979] In this invention, the server includes means for switching received communications from human to artificial intelligence automated response, means for the artificial intelligence automated response to save the conversation content and summarize the perjury potential of the other party's statements, and means for the artificial intelligence automated response to suggest the next action. This makes it possible to evaluate the reliability of the conversation in real time and present it to the user immediately.
[0980] "Communication" refers to the act or means of sending and receiving voice or data.
[0981] "Artificial intelligence automated response" refers to a system that uses artificial intelligence technology to automatically generate responses and engage in conversation.
[0982] "Conversation content" refers to the entirety of the statements and information exchanged during communication.
[0983] "Perjury potential" is an indicator of the degree to which a statement is likely to be false.
[0984] "Speech recognition technology" is a technology that converts speech into text.
[0985] "Natural language processing technology" is a technology used to analyze text data and understand its meaning and intent.
[0986] A "display device" is a device used to present information visually.
[0987] "User" refers to an individual or group that uses the system.
[0988] "Action" refers to the next step or action suggested by the system.
[0989] The system for carrying out this invention includes a server and a terminal. The server has the function of switching received communications from human to artificial intelligence automated responses. The terminal uses speech recognition technology to transcribe speech during communication into text in real time. Specifically, it uses the Google Cloud Speech-to-Text API to convert speech to text.
[0990] The server analyzes the transcribed speech using natural language processing techniques. This analysis utilizes natural language processing libraries such as spaCy and Transformers. The analysis results in an evaluation of the perjury potential of the speech. This evaluation uses a pre-trained generative AI model (e.g., OpenAI GPT-3).
[0991] The evaluation results are displayed in real time on the terminal's display device. This allows users to instantly verify the reliability of the conversation. For example, if someone says "This product is the best-selling in the market" during a meeting, the server will refer to past data and market information to evaluate the reliability of the statement.
[0992] An example of a prompt for a generative AI model might be: "Evaluate the likelihood that this statement is true. Market data is as follows." Using this prompt, the model determines the reliability of the statement and provides the result to the user.
[0993] The flow of a specific process in Application Example 2 will be explained using Figure 14.
[0994] Step 1:
[0995] The device acquires audio during communication via its microphone. The input is audio data, which is converted into text data using the Google Cloud Speech-to-Text API. The output is the transcribed speech.
[0996] Step 2:
[0997] The server receives text data sent from the terminal. The input is text data, which is then analyzed using a natural language processing library (e.g., spaCy, Transformers). The purpose of the analysis is to understand the content of the statement and extract information to assess its perjury potential. The output is the analysis result.
[0998] Step 3:
[0999] The server sends a prompt to a generative AI model (e.g., OpenAI GPT-3) based on the analysis results. The input consists of the analysis results and the prompt text. An example of the prompt text is "Evaluate the likelihood that this statement is true. The market data is as follows." The model evaluates the reliability of the statement and outputs the result.
[1000] Step 4:
[1001] The server receives evaluation results from the generated AI model and sends them to the terminal. The input is the evaluation result, and the output is reliability evaluation information to be presented to the user.
[1002] Step 5:
[1003] The terminal displays reliability evaluation information received from the server on its display device. The input is reliability evaluation information, and the output is information that the user can visually confirm. This allows the user to instantly verify the reliability of the conversation.
[1004] (Example 3)
[1005] Next, we will describe Embodiment 3 of Embodiment Example 3. In the following description, the data processing device 12 will be referred to as the "server," and the headset-type terminal 314 will be referred to as the "terminal."
[1006] In today's communication environment, there is a need to quickly and accurately determine the veracity of what the other party says and to suggest appropriate actions. However, conventional systems have the challenge of having a complex process for evaluating the veracity of statements, making it difficult to clearly indicate the next action the user should take.
[1007] The identification process performed by the identification processing unit 290 of the data processing device 12 in Example 3 is realized by the following means.
[1008] In this invention, the server includes means for switching received communications from a human to an automated response system, means for the automated response system to store conversation information and evaluate the veracity of the other party's statements, and means for the automated response system to propose the next action based on the evaluation result. This makes it possible to quickly evaluate the veracity of statements and propose appropriate actions to the user.
[1009] "Received communications" refers to audio and data signals sent from external sources.
[1010] An "automatic response device" refers to a device that has the function of automatically responding to received communications.
[1011] "Conversation information" refers to the content of audio and text exchanged during communication.
[1012] "Evaluating the truthfulness of a statement or piece of information" refers to the process of judging the accuracy and reliability of such statements or information.
[1013] "Suggesting action" means indicating specific actions to take next based on the evaluation results.
[1014] A "communication terminal" refers to a device used by a user that has the function of performing communication.
[1015] "User" refers to a person who uses the system.
[1016] A description of embodiments for carrying out this invention will be given.
[1017] First, the user enters a prompt message using a communication terminal. This prompt message is used to determine the truthfulness of the other party's statement. As a concrete example, consider the prompt message, "Is this statement true?"
[1018] Next, the terminal sends the entered prompt message to the server. The server uses a generative AI model to analyze the received prompt message. This analysis utilizes natural language processing techniques, employing models such as "OpenAI GPT-3" and "Google BERT." The server uses these models to scrutinize the content of the utterance and evaluate its truthfulness.
[1019] Based on the evaluation results, the server suggests the next course of action. For example, if it is determined that the statement is highly likely to be perjury, the server will decide to "suggest reporting to the police."
[1020] The server then sends the suggested action to the terminal. The terminal displays the received suggestion to the user. The user can then use this suggestion to decide on their next action.
[1021] In this way, it becomes possible to quickly evaluate the veracity of statements and propose appropriate actions. The flow of the specific processing in Example 3 will be explained using Figure 15.
[1022] Step 1:
[1023] The user enters a prompt message using a communication terminal. For example, they might enter the prompt message, "Is this statement true?" This input serves as the basis for determining the truthfulness of the other party's statement.
[1024] Step 2:
[1025] The terminal sends the entered prompt message to the server. Here, the input is the prompt message, and the output is the data sent to the server. The terminal securely transmits the data using the HTTPS protocol.
[1026] Step 3:
[1027] The server inputs the received prompt text into a generating AI model. The input is the prompt text, and the output is the analysis result by the AI model. The server uses natural language processing models such as "OpenAI GPT-3" and "Google BERT" to analyze the content of the utterance and evaluate its truthfulness.
[1028] Step 4:
[1029] The server proposes the next course of action based on the analysis results. The input is the analysis results of the AI model, and the output is the proposed action. For example, if the server determines that there is a high probability that the statement is perjury, it will decide to take the action of "I suggest reporting this to the police."
[1030] Step 5:
[1031] The server sends the proposed action to the terminal. The input is the proposed action, and the output is the transmission of data to the terminal. The server securely transmits the data using the HTTPS protocol.
[1032] Step 6:
[1033] The terminal displays received suggestions to the user. The input is suggestions from the server, and the output is what is displayed to the user. The terminal displays the suggestions on the screen, allowing the user to decide on their next action.
[1034] (Application Example 3)
[1035] Next, we will describe application example 3 of form example 3. In the following description, the data processing device 12 will be referred to as the "server," and the headset-type terminal 314 will be referred to as the "terminal."
[1036] In modern society, the spread of false information and fraudulent activities through communications is increasing, and there is a need for a swift and effective response. However, conventional systems have difficulty determining the falsehood of communications in real time and suggesting appropriate actions. As a result, the risk of users becoming involved in fraudulent activities is increasing.
[1037] The specific processing performed by the specific processing unit 290 of the data processing device 12 in Application Example 3 is realized by the following means.
[1038] In this invention, the server includes means for switching received communications from a human to an AI-powered automated response, means for the AI automated response to record the content of the conversation and evaluate the falsity of the other party's statements, and means for the AI automated response to propose the next action. This makes it possible to determine the falsity of the communication content in real time and propose appropriate actions to the user.
[1039] "Means for switching received communications from human response to AI-powered automated response" refers to technology that switches from a conventional human response to an automated response by artificial intelligence when communication begins.
[1040] "An artificial intelligence automated response system that records the content of conversations and evaluates the veracity of the other party's statements" refers to a technology in which artificial intelligence stores the content of communications, analyzes that content, and determines the truthfulness of the statements.
[1041] "An artificial intelligence automated response system that suggests the next action" refers to a technology that, based on analysis results, suggests the appropriate next action to the user.
[1042] "Means of converting dialogue into text information using speech recognition technology" refers to technology for converting spoken dialogue into text format.
[1043] "Methods for analyzing the possibility of falsehood using generative AI models" refers to techniques that utilize generative AI models to evaluate the falsehood of dialogue content.
[1044] "A means of displaying warnings based on analysis results and suggesting reporting to legal authorities if necessary" refers to a technology that issues warnings to users based on the results of a falsehood assessment and, if necessary, encourages them to report to legal authorities.
[1045] The system for implementing this invention mainly consists of a server and a terminal. The server has the function of switching received communications from human to AI-powered automated responses. The terminal converts the dialogue into text information using speech recognition technology. Specifically, it uses speech recognition software such as the Google Speech-to-Text API to convert speech data into text data.
[1046] The server uses a generative AI model to analyze the falsity of the converted text data. This analysis leverages generative AI models such as OpenAI's GPT-4. The generative AI model assesses the likelihood that the content is false based on the input text.
[1047] Based on the analysis results, the server displays a warning on the terminal and, if necessary, suggests reporting to legal authorities. This allows users to understand the falsity of communications in real time and take appropriate action.
[1048] As a concrete example, consider a scenario where a user is using this system during a business meeting. When someone says, "This product is 100% safe," the device uses speech recognition technology to convert the statement into text, which is then analyzed by a server using a generated AI model. If the analysis determines that the statement is highly false, the device displays a warning to the user: "This statement requires caution. Please review the details or consider legal action if necessary."
[1049] Examples of prompt statements to input into a generative AI model include the following:
[1050] "Assess the likelihood that the following statement is perjury: 'This product is 100% safe.'"
[1051] The flow of the specific processing in Application Example 3 will be explained using Figure 16.
[1052] Step 1:
[1053] The device receives voice input from the user. This voice input is audio data of the user's conversation. The device converts this audio data into text data using the Google Speech-to-Text API. The converted text data is then output.
[1054] Step 2:
[1055] The server receives text data sent from the terminal. The server inputs this text data into a generative AI model (e.g., GPT-4). The generative AI model analyzes the content of the text data and evaluates the likelihood of the statement being false. As a result of the evaluation, a score indicating the likelihood of falsehood is output.
[1056] Step 3:
[1057] The server determines whether to display a warning to the user based on the evaluation results from the generated AI model. If the falsehood score exceeds a certain threshold, the server sends a warning message to the terminal. The warning message states that caution is needed regarding what is said.
[1058] Step 4:
[1059] The terminal receives a warning message sent from the server and displays it to the user. The user reviews the displayed warning message and considers reporting it to legal authorities if necessary. This allows the user to understand the falsity of the communication content in real time and take appropriate action.
[1060] Furthermore, an emotion engine that estimates the user's emotions may be incorporated. That is, the identification processing unit 290 may use the emotion identification model 59 to estimate the user's emotions and perform identification processing using the user's emotions.
[1061] "Example of form 1"
[1062] One embodiment of the present invention provides a system in which an AI automated response system includes an emotion engine that recognizes the user's emotions. This emotion engine analyzes the user's emotions from the tone of their voice, manner of speaking, word choice, etc. For example, if the user is feeling angry or anxious, the emotion engine feeds that information back to the AI automated response system, which adjusts its response accordingly.
[1063] "Example of form 2"
[1064] The system also provides an emotion engine that adjusts the AI's automated responses based on the user's emotions. Specifically, if the system determines that the user is angry, the AI's automated response will be adjusted to use more polite language and a calmer tone. This allows for more appropriate responses that take the user's emotions into consideration.
[1065] "Example of form 3"
[1066] Furthermore, the system provides an emotion engine that adjusts the AI's automated response suggestions for the next action based on the user's emotions. For example, if the system determines that the user is feeling anxious, the AI automated response will suggest actions to alleviate that anxiety. This enables more appropriate action suggestions that are tailored to the user's emotions.
[1067] The following describes the processing flow for each example of the form.
[1068] "Example of form 1"
[1069] Step 1: The user receives a suspicious phone call.
[1070] Step 2: The user switches the phone to AI automated answering.
[1071] Step 3: The emotion engine analyzes the user's emotions from their voice.
[1072] Step 4: The AI automated response system adjusts its response based on feedback from the emotion engine.
[1073] "Example of form 2"
[1074] Step 1: The user receives a suspicious phone call.
[1075] Step 2: The user switches the phone to AI automated answering.
[1076] Step 3: The emotion engine analyzes the user's emotions from their voice.
[1077] Step 4: The AI automated response system adjusts its response based on feedback from the emotion engine.
[1078] "Example of form 3"
[1079] Step 1: The user receives a suspicious phone call.
[1080] Step 2: The user switches the phone to AI automated answering.
[1081] Step 3: The emotion engine analyzes the user's emotions from their voice.
[1082] Step 4: The AI automated response system adjusts its next action suggestion based on feedback from the emotion engine.
[1083] (Example 1)
[1084] Next, we will describe Embodiment 1 of Embodiment Example 1. In the following description, the data processing device 12 will be referred to as the "server," and the headset-type terminal 314 will be referred to as the "terminal."
[1085] Traditional communication systems required recipients to respond manually, making it difficult to provide appropriate responses that reflected emotions. Furthermore, they lacked automated processes for saving conversation content and determining perjury, and also lacked features to suggest the next course of action.
[1086] The identification process performed by the identification processing unit 290 of the data processing device 12 in Example 1 is realized by the following means.
[1087] In this invention, the server includes means for switching received communications from a human to an AI-powered automated response, means for the AI automated response to acquire voice data and perform emotion analysis, and means for the AI automated response to adjust the response content based on the emotion analysis. This enables appropriate responses according to emotions without the recipient having to respond manually.
[1088] "Received communications" refers to audio information sent to the terminal from an external source, such as phone calls or voice messages.
[1089] "Artificial intelligence automated response" refers to a program or system that automatically generates and transmits a response to an incoming communication.
[1090] "Audio data" refers to data that represents audio information obtained from received communications in digital format.
[1091] "Emotional analysis" refers to the process of analyzing audio data to identify the speaker's emotional state.
[1092] "Adjusting the response content" refers to changing the content and tone of the response generated by the AI automated response system based on the results of sentiment analysis.
[1093] "Saving conversation content" refers to recording audio information exchanged during communication so that it can be referenced later.
[1094] "Summarizing the perjury" refers to analyzing the content of a conversation and organizing the information needed to evaluate the veracity of the statements.
[1095] "Suggesting the next action" refers to the AI automated response system suggesting the appropriate course of action to take based on the content and analysis results of the communication.
[1096] This invention provides a system that efficiently responds to incoming communications using artificial intelligence-based automated response. Specific embodiments of this system are described below.
[1097] When the server receives a communication, it first activates an AI-powered automated response system. This automated response system uses speech recognition software to acquire voice data and perform sentiment analysis. Specifically, it uses a "voice capture system" as the speech recognition software and a "sentiment analysis engine" for sentiment analysis. This analyzes the user's voice tone, speaking style, and word choice to identify their emotions.
[1098] The device switches to an AI-powered automated response system when the user presses the "AI Response Switch Button" upon receiving a communication. This action sends a signal to the server, initiating the AI automated response. The server adjusts the response based on the sentiment analysis results, providing the user with an appropriate response.
[1099] For example, when a user receives a call and presses the "AI response switch button," the server immediately acquires the voice data and performs sentiment analysis. If the analysis indicates that the user is feeling anxious, the AI automated response will respond in a gentle tone, saying, "Please don't worry. How can I help you?" This response is recorded on the server and used to improve the generated AI model.
[1100] An example of a prompt sentence to input into a generative AI model might be, "If a user is angry on the phone, how should the AI automated response system respond?" By using this prompt sentence, the generative AI model can suggest an appropriate response.
[1101] The flow of the specific processing in Example 1 will be explained using Figure 17.
[1102] Step 1:
[1103] The terminal detects incoming communications. When the user presses the "AI response switch button" upon receiving a communication, the terminal sends a signal to the server requesting the activation of the artificial intelligence automatic response. The input for this step is the incoming communication signal, and the output is the activation request signal to the server.
[1104] Step 2:
[1105] When the server receives a startup request from the terminal, it activates an artificial intelligence automated response system. The server uses speech recognition software to obtain voice data from the communication. The input for this step is the startup request signal from the terminal, and the output is the obtained voice data.
[1106] Step 3:
[1107] The server sends the acquired audio data to the emotion analysis engine. The emotion analysis engine analyzes the audio data and identifies the user's emotions. The input to this step is the audio data, and the output is the analyzed emotion information. Specifically, it analyzes the tone of voice, speaking style, and word choice.
[1108] Step 4:
[1109] The server adjusts the content of the AI-powered automated response based on emotional information from the emotion analysis engine. The server uses a generative AI model to generate an appropriate response that matches the user's emotions. The input for this step is the analyzed emotional information, and the output is the adjusted response.
[1110] Step 5:
[1111] The server sends the refined response to the user. The user receives the response from the AI automated response system. The input for this step is the refined response, and the output is the response to the user. Specifically, this involves using a gentle tone and reassuring language in the response.
[1112] (Application Example 1)
[1113] Next, we will describe Application Example 1 of Form Example 1. In the following description, the data processing device 12 will be referred to as the "server," and the headset-type terminal 314 will be referred to as the "terminal."
[1114] Conventional communication systems made it difficult for recipients to quickly assess the urgency of a call and respond appropriately. Furthermore, they were unable to accurately analyze the caller's emotions and adjust their responses accordingly, resulting in inadequate responses, particularly in emergencies where a rapid and appropriate response was required.
[1115] The specific processing performed by the specific processing unit 290 of the data processing device 12 in Application Example 1 is realized by the following means.
[1116] In this invention, the server includes means for switching received communications from human response to automated artificial intelligence response, means for analyzing the caller's emotions using an emotion analysis engine and determining the urgency, and means for automatically executing an appropriate response according to the urgency. This enables a rapid and appropriate response based on the caller's emotions.
[1117] "Received communications" refers to the act of receiving information, including audio, such as phone calls and voice messages.
[1118] "Artificial intelligence automated response" refers to an artificial intelligence system that automatically generates responses to received communications and engages in dialogue.
[1119] "Saving conversation content" refers to the act of recording the content of a phone call or conversation so that it can be referenced later.
[1120] "Summarizing the perjury" is the act of determining whether the other party's statement is true or false and organizing the results of that determination.
[1121] "Suggesting the next course of action" means suggesting the next course of action to take based on the current situation.
[1122] An "emotion analysis engine" is a technology that analyzes emotions from a caller's voice and words to identify those emotions.
[1123] "Assessing the urgency" means evaluating the importance and urgency of a situation and determining the priority of responses.
[1124] "Automatically executing appropriate actions" means automatically taking the necessary actions based on the determined level of urgency.
[1125] The system for implementing this invention operates in a network environment including a server and terminals. The server has the function of switching received communications to an artificial intelligence-based automated response system and uses an emotion analysis engine to analyze the caller's voice. The emotion analysis engine analyzes the caller's tone of voice, speaking style, and word choice to identify emotions. Based on this, the server determines the urgency and automatically takes an appropriate action.
[1126] The terminal provides an interface for sending voice to the server, and the user can switch to an AI-powered automated response by pressing a specific button. The server uses the Python speech_recognition library as speech recognition software and a custom EmotionEngine for sentiment analysis. Furthermore, the generated response is provided by the AIResponse module.
[1127] For example, if a user is using a security company's terminal, they can switch to an AI-powered automated response system by pressing a button on the terminal when an emergency call comes in. The server analyzes the caller's statement, "Help me, someone is breaking into my house," and an emotion analysis engine detects a high level of fear. As a result, the server immediately issues an order to dispatch security personnel.
[1128] An example of a prompt message for a generative AI model is: "Analyze the caller's voice and identify their emotion. If the emotion is fear, immediate action is required."
[1129] The flow of a specific process in Application Example 1 will be explained using Figure 18.
[1130] Step 1:
[1131] When the user presses a button to switch the received communication to an AI-powered automated response, the terminal sends voice data to the server. The input is the user's voice, and the output is the transmission of voice data to the server. The terminal converts the voice into digital data and sends it to the server over the network.
[1132] Step 2:
[1133] The server converts the received audio data into text using the speech_recognition library. The input is audio data, and the output is text data. The server applies a speech recognition algorithm to extract text from the audio.
[1134] Step 3:
[1135] The server inputs text data into an emotion analysis engine to identify the caller's emotions. The input is text data, and the output is emotion information. The emotion analysis engine analyzes the content and word choice of the text to determine the emotion.
[1136] Step 4:
[1137] The server determines the urgency based on emotional information and decides on the appropriate response. The input is emotional information, and the output is a response instruction. If the emotion is fear or anxiety, the server determines that a rapid response is necessary and generates a response instruction.
[1138] Step 5:
[1139] The server executes the generated response instructions and takes specific actions as needed, such as dispatching security personnel. The input is the response instructions, and the output is the actions taken. The server notifies the relevant systems and services and implements the response according to the instructions.
[1140] (Example 2)
[1141] Next, we will describe Example 2 of the morphological example. In the following description, the data processing device 12 will be referred to as the "server," and the headset-type terminal 314 will be referred to as the "terminal."
[1142] Conventional automated response systems simply record the content of conversations with users without adjusting responses based on the falsity of statements or the user's emotions. This results in a poor user experience and can lead to misunderstandings and dissatisfaction. Furthermore, the lack of features to suggest the next course of action makes it difficult for users to take appropriate action.
[1143] The identification process performed by the identification processing unit 290 of the data processing device 12 in Example 2 is realized by the following means.
[1144] In this invention, the server includes means for switching received communications from a human to an automated response system, means for the automated response system to record the content of the conversation and analyze the falsity of the other party's statements, and means for the automated response system to adjust the content of the response based on sentiment analysis. This makes it possible to provide appropriate responses that take into account the falsity of statements and the emotions of the user during conversations, and to further suggest the next action.
[1145] "Received communications" refers to audio and data sent from external sources, and the purpose is to switch these to an automated response system.
[1146] An "automated response device" refers to a device or system that automatically processes, records, analyzes, and responds to user interactions.
[1147] "Recording conversation content" refers to the act of saving conversations and communications with users as data, which can then be used for later analysis and reference.
[1148] "Analyzing the falsity of statements" refers to the process of analyzing the content of a user's statements and determining whether or not that content is based on facts.
[1149] "Emotion analysis" refers to a technology that reads emotions from a user's speech and voice, and identifies emotional states such as anger, joy, and sadness.
[1150] "Adjusting response content" refers to changing the content and tone of the response generated by the automated response system according to the analyzed emotions and situation.
[1151] "Suggesting the next action" refers to showing the user the appropriate next steps or actions based on their current situation.
[1152] This invention is a system that efficiently manages user interaction using an automated response device and improves the user experience. Specific embodiments of this system are described below.
[1153] The server switches incoming communications from a human to an automated response system. This involves receiving voice data over the communication network and converting it to text using speech recognition software. Specifically, it uses a "speech recognition API" for speech recognition.
[1154] The terminal receives the user's voice input and sends it to the server. The server converts the received voice data into text using a "speech recognition API" and stores that text in a database. A "database management system" is used for this database.
[1155] The server analyzes the stored text using a natural language processing library to determine whether the statement is false. This allows the server to verify whether the user's statement is based on facts.
[1156] Furthermore, the device uses a "sentiment analysis engine" to analyze the user's emotions from their voice and text. The server receives the results of the sentiment analysis and determines what emotion the user is experiencing, such as anger, joy, or sadness.
[1157] The server adjusts the content of the AI automated response based on the results of the sentiment analysis. For example, if it is determined that the user is angry, the AI automated response will be set to respond in a calm tone, such as, "I'm sorry. Please let me know how we can improve."
[1158] For example, if a user asks, "Does this product really work?", the server transcribes the question into text and stores it in the database. It then compares this question with past data and product information to provide accurate information.
[1159] Examples of prompts include, "Transcribe the user's utterance into text and generate a prompt to determine its falsity," and "Create a prompt that adjusts the AI's response based on the user's sentiment."
[1160] This system makes it possible to manage user interactions in a more accurate and emotionally sensitive manner.
[1161] The flow of the specific processing in Example 2 will be explained using Figure 19.
[1162] Step 1:
[1163] The device receives voice input from the user. For example, the user might say, "Does this product really work?" The device then sends this voice data to the server. The input is the user's voice data, and the output is the transfer of the voice data to the server.
[1164] Step 2:
[1165] The server converts the received audio data into text using a speech recognition API. The speech recognition API analyzes the audio signal and generates corresponding text data. The input is audio data, and the output is text data.
[1166] Step 3:
[1167] The server stores the transcribed conversation content in a database. A database management system is used to record the text data in the appropriate format. The input is text data, and the output is storage in the database.
[1168] Step 4:
[1169] The server analyzes the stored text using a natural language processing library to determine the falsity of the statements. The natural language processing library analyzes the text data and evaluates its falsity by comparing it with known facts. The input is text data, and the output is the result of the falsity evaluation.
[1170] Step 5:
[1171] The device uses an emotion analysis engine to analyze the user's emotions from their voice or text. The emotion analysis engine identifies the user's emotional state from the input data. The input is voice or text data, and the output is the result of the emotion analysis.
[1172] Step 6:
[1173] The server adjusts the content of the AI-generated response based on the results of the sentiment analysis. The server changes the tone and content of the response according to the user's emotions to generate an appropriate response. The input is the result of the sentiment analysis, and the output is the adjusted response content.
[1174] Step 7:
[1175] The server sends a pre-arranged response to the terminal and provides the response to the user in voice or text. The terminal receives the response from the server and relays it to the user. The input is the pre-arranged response, and the output is the provision of the response to the user.
[1176] (Application Example 2)
[1177] Next, we will describe application example 2 of form example 2. In the following description, the data processing device 12 will be referred to as the "server," and the headset-type terminal 314 will be referred to as the "terminal."
[1178] In modern communications, judging the reliability of what the other party says is crucial, but this is difficult with conventional systems. Furthermore, there is a need to provide appropriate responses that respond to the user's emotions, but effective means to achieve this are lacking. Additionally, there is a demand for a function that issues warnings based on the assessment of perjury, but the technology to implement this is still immature.
[1179] The specific processing performed by the specific processing unit 290 of the data processing device 12 in Application Example 2 is realized by the following means.
[1180] In this invention, the server includes means for switching received communications from human to machine learning automated responses, means for the machine learning automated response to save the conversation content and summarize the perjury potential of the other party's statements, and means for the machine learning automated response to adjust the response content based on sentiment analysis. This makes it possible to determine the reliability of statements in communications and to provide appropriate responses that correspond to the user's emotions.
[1181] "Communication" refers to the act or process of sending and receiving information, and includes forms such as voice, text, and data.
[1182] "Machine learning automated response" refers to a system that uses machine learning technology to automatically generate responses, providing appropriate answers based on user input.
[1183] "Saving conversation content" means recording information exchanged during communication and making it available for later reference.
[1184] "Summarizing the perjury" means evaluating whether the other party's statements are true or false, and then organizing and presenting the results.
[1185] "Emotional analysis" is a technology that infers emotions from a user's words and actions and analyzes their emotional state.
[1186] "Adjusting the response content" means changing the content and tone of the response according to the user's emotions and situation.
[1187] "Issuing a warning" means providing a message to draw attention when certain conditions are met.
[1188] "Suggesting an action" means indicating to the user what they should do next.
[1189] The system for implementing this invention mainly consists of a server and a terminal. The server has the function of switching received communications from human input to automated machine learning responses. The terminal is responsible for receiving input from the user and sending it to the server.
[1190] The server uses natural language processing libraries (e.g., spaCy, NLTK) to analyze the received conversation content and determine the perjury potential of the other party's statements. Furthermore, it uses sentiment analysis APIs (e.g., IBM Watson Tone Analyzer) to analyze the user's emotions and adjust the response based on the results. This enables appropriate responses tailored to the user's emotions.
[1191] Furthermore, the server has a function to issue warnings based on its assessment of perjury, allowing it to alert users. This can improve the reliability of communications.
[1192] For example, if a user asks, "Is this transaction safe?", the server will refer to past data and evaluate the reliability of the statement. Furthermore, if the server determines that the user is feeling uneasy, it will respond in a calm tone, such as, "Please rest assured, this transaction is safe."
[1193] An example of a prompt for a generative AI model is: "Analyze the user's statements and determine their perjury potential. Also, analyze the user's emotions and generate an appropriate response."
[1194] The flow of a specific process in Application Example 2 will be explained using Figure 20.
[1195] Step 1:
[1196] The device receives voice input from the user. This voice input is converted into text data using a speech recognition API (e.g., Google Speech-to-Text). This text data then becomes the input for the next process.
[1197] Step 2:
[1198] The server analyzes the received text data using a natural language processing library (e.g., spaCy, NLTK). The analysis extracts the content of the statement and generates data for determining its perjury potential. This data then becomes the input for the next process.
[1199] Step 3:
[1200] The server performs a perjury assessment. Specifically, it compares the statement against past database data to evaluate its reliability. This evaluation result becomes the input for the next process.
[1201] Step 4:
[1202] The server uses a sentiment analysis API (e.g., IBM Watson Tone Analyzer) to analyze the user's emotions from text data. Based on the analysis results, the user's emotional state is identified. This emotional state becomes the input for the next process.
[1203] Step 5:
[1204] The server generates an appropriate response based on the perjury assessment results and emotional state. A generative AI model is used to create the response to the user. This response becomes the input for the next process.
[1205] Step 6:
[1206] The server issues a warning as needed based on the results of the perjury assessment. A warning message is generated and sent to the user.
[1207] Step 7:
[1208] The terminal displays responses and warning messages received from the server to the user. Based on this, the user can decide on their next action.
[1209] (Example 3)
[1210] Next, we will describe Embodiment 3 of Embodiment Example 3. In the following description, the data processing device 12 will be referred to as the "server," and the headset-type terminal 314 will be referred to as the "terminal."
[1211] Conventional AI-powered automated response systems could determine the perjury potential of a statement, but they lacked the ability to offer appropriate action suggestions that took the user's emotional state into account. Furthermore, their action suggestions based on the perjury potential were limited, resulting in insufficient responses to alleviate user anxiety.
[1212] The identification process performed by the identification processing unit 290 of the data processing device 12 in Example 3 is realized by the following means.
[1213] In this invention, the server includes means for switching received communications from human to artificial intelligence automated response, means for the artificial intelligence automated response to save the conversation content and determine the perjury of the other party's statements, means for analyzing the user's emotional state using an emotion analysis engine, and means for proposing appropriate actions based on the results of the perjury determination and the emotion analysis. This enables flexible and appropriate action suggestions that respond to the user's emotions.
[1214] "Means for switching received communications from human to AI-powered automated responses" refers to a function that, after a communication has been received by a human, switches that response to an automated response by artificial intelligence.
[1215] "An artificial intelligence automated response system that saves conversation content and determines whether the other party's statements are perjury" refers to a function in which artificial intelligence records the content of a conversation, analyzes that content, and evaluates the truthfulness of the statements.
[1216] "Methods for analyzing a user's emotional state using an emotion analysis engine" refers to technologies that infer emotions from a user's statements and actions and analyze that state.
[1217] "A means of suggesting appropriate actions based on the results of perjury assessment and sentiment analysis" refers to a function that suggests the next course of action, taking into account the truthfulness of a statement and the user's emotional state.
[1218] A description of embodiments for carrying out this invention will be given.
[1219] The server generates a program to switch the communication received from the user to an AI-powered automated response. This program runs using server hardware equipped with an NVIDIA GPU and a generative AI model using TensorFlow. The server receives the user's utterances as text data and analyzes their content using natural language processing techniques. The analysis includes data processing to determine the perjury potential of the utterances and data calculations using a sentiment analysis engine.
[1220] For example, if a user asks "Does this product really work?" through their device, the server uses a generative AI model to analyze the statement and extract key keywords. The server then compares the statement against a past database to evaluate its credibility. Furthermore, if the sentiment analysis engine determines that the user is feeling anxious, the server suggests actions such as, "This information may not be reliable. We recommend consulting a professional for further information."
[1221] An example of a prompt to be input to the generating AI model is, "Analyze the user's statement, determine its perjury potential, and suggest an appropriate action." This prompt allows the server to generate an appropriate response to the user's statement. The flow of specific processing in Example 3 will be explained using Figure 21.
[1222] Step 1:
[1223] The user inputs their message into the system via a terminal. The entered text data is sent to the server. The server receives this text data and prepares it for the next analysis step.
[1224] Step 2:
[1225] The server analyzes the received text data using natural language processing techniques. Specifically, it uses a generative AI model to understand the content of the utterance and extract important keywords and context. This analysis clarifies the intent and subject of the utterance. The input is the user's utterance, and the output is the analyzed keywords and contextual information.
[1226] Step 3:
[1227] The server determines the perjury potential of a statement based on the analysis results. It compares the statement against a past database to evaluate its credibility. This process involves database searches and comparison operations. The input is the analyzed keywords and contextual information, and the output is the result of the perjury determination.
[1228] Step 4:
[1229] The server uses an emotion analysis engine to analyze the user's emotional state. It infers emotions from the user's statements and identifies emotions such as anxiety and anger. The input is the user's statements, and the output is the evaluation result of the emotional state.
[1230] Step 5:
[1231] The server suggests the next course of action based on the results of the perjury assessment and sentiment analysis. For example, if there is a high probability that the statement is perjury and the user is feeling anxious, the server might suggest, "This information may not be reliable. We recommend consulting a professional." The input is the result of the perjury assessment and the evaluation of the emotional state, and the output is the suggested action.
[1232] Step 6:
[1233] The server sends the generated suggestions to the terminal and presents them to the user. The user can then decide on their next action based on these suggestions. The input is the suggested action, and the output is the content presented to the user.
[1234] (Application Example 3)
[1235] Next, we will describe application example 3 of form example 3. In the following description, the data processing device 12 will be referred to as the "server," and the headset-type terminal 314 will be referred to as the "terminal."
[1236] In modern society, fraudulent activities and the spread of false information through communications are on the rise, making it difficult for users to protect themselves from these threats. Furthermore, while appropriate responses that respond to users' emotions are required, conventional systems have not been able to adequately achieve this. Therefore, ensuring user safety and peace of mind remains a challenge.
[1237] The specific processing performed by the specific processing unit 290 of the data processing device 12 in Application Example 3 is realized by the following means.
[1238] In this invention, the server includes means for switching received communications from an information processing device to an automated response device, means for the automated response device to record conversation information and evaluate the falsity of the other party's statements, means for the automated response device to propose the next action, and means for an emotion analysis device to evaluate the user's emotions and adjust the proposed action. This makes it possible for the user to evaluate the falsity of the communication content and receive an appropriate response according to their emotions.
[1239] An "information processing device" is a device that receives communications and, if necessary, switches to an automatic response device.
[1240] An "automatic response device" is a device that records received conversation information and has the function of evaluating the likelihood of the other party's statements being false.
[1241] "Conversational information" refers to voice or text data exchanged through communication.
[1242] "Falsehood" is an evaluation criterion that indicates the possibility that a statement is not based on facts.
[1243] An "emotion analysis device" is a device that evaluates the user's emotions and adjusts the suggested actions accordingly.
[1244] A "means of suggesting actions" refers to a function that indicates to the user the next action they should take based on the evaluation results.
[1245] To implement this invention, a server plays a central role. The server is equipped with an information processing device that receives communications and switches the received communications to an automatic response device. The automatic response device records conversation information and evaluates the falsity of the other party's statements. The evaluation is performed using natural language processing techniques, specifically using software such as Python or TensorFlow.
[1246] Furthermore, an emotion analysis device evaluates the user's emotions and adjusts the suggested actions accordingly. Emotion analysis utilizes emotion analysis APIs such as IBM Watson. This enables the provision of appropriate action suggestions based on the user's emotional state.
[1247] For example, if a user receives a potentially fraudulent phone call, the server will suggest, "This call may be fraudulent. Do you want to report it to the police?" If the user is feeling uneasy, it will also suggest, "Do you want to send a notification to a trusted contact?"
[1248] Examples of prompts generated using AI models include, "Analyze the content of this message and assess the likelihood of perjury," and "Analyze the user's sentiment and suggest appropriate action."
[1249] The flow of the specific processing in Application Example 3 will be explained using Figure 22.
[1250] Step 1:
[1251] The server receives communications from the terminal using an information processing device. Input can include voice or text data. The server sends this data to an automated response system for recording as conversation information.
[1252] Step 2:
[1253] The automated response system analyzes recorded conversation information. The input is the audio or text data recorded in step 1. The server uses natural language processing techniques to evaluate the falsity of the statements. Specifically, it analyzes the data using Python or TensorFlow and outputs a falsity score.
[1254] Step 3:
[1255] The server suggests the next action based on the falsehood score. The input is the falsehood score obtained in step 2. If the score is high, the server will generate a suggestion such as, "This call may be perjury. Do you want to report it to the police?"
[1256] Step 4:
[1257] The emotion analysis device evaluates the user's emotions. The input is the voice or text data received in step 1. The server analyzes the emotions using an emotion analysis API such as IBM Watson and outputs the emotional state.
[1258] Step 5:
[1259] The server adjusts the suggested actions based on the emotional state. The input is the emotional state obtained in step 4. If the user is feeling anxious, the server will generate a suggestion such as, "Do you want to send a notification to a trusted contact?"
[1260] Step 6:
[1261] The user receives suggestions from the server and selects the appropriate action. The input consists of the suggestions generated in steps 3 and 5. The user decides on an action based on the suggestions and sends feedback to the server as needed.
[1262] (Other examples)
[1263] Since this is the same as the specific processing described in the other embodiments of the first embodiment above, the explanation will be omitted.
[1264] The specific processing unit 290 transmits the result of the specific processing to the headset terminal 314. In the headset terminal 314, the control unit 46A causes the speaker 240 and display 343 to output the result of the specific processing. The microphone 238 acquires audio indicating user input for the result of the specific processing. The control unit 46A transmits the audio data indicating user input acquired by the microphone 238 to the data processing unit 12. In the data processing unit 12, the specific processing unit 290 acquires the audio data.
[1265] The data generation model 58 is a form of so-called generative AI (Artificial Intelligence). One example of the data generation model 58 is ChatGPT (Internet Search).<URL: https: / / openai.com / blog / chatgpt> Examples of generative AI include the following. The data generation model 58 is obtained by performing deep learning on a neural network. The data generation model 58 is input with prompts containing instructions, and with inference data such as audio data representing speech, text data representing text, and image data representing images. The data generation model 58 infers from the input inference data according to the instructions indicated by the prompts, and outputs the inference results in data formats such as audio data and text data. Here, inference refers to, for example, analysis, classification, prediction, and / or summarization.
[1266] Other examples of generative AI include Gemini (Internet search <url: https: gemini.google.com ?hl="ja">) are some examples.
[1267] In the above embodiment, an example was given in which specific processing is performed by the data processing device 12, but the technology of this disclosure is not limited thereto, and specific processing may also be performed by the headset terminal 314.
[1268] [Fourth Embodiment]
[1269] Figure 7 shows an example of the configuration of the data processing system 410 according to the fourth embodiment.
[1270] As shown in Figure 7, the data processing system 410 includes a data processing device 12 and a robot 414. An example of the data processing device 12 is a server.
[1271] The data processing device 12 comprises a computer 22, a database 24, and a communication interface 26. The computer 22 is an example of a "computer" related to the technology of this disclosure. The computer 22 comprises a processor 28, RAM 30, and storage 32. The processor 28, RAM 30, and storage 32 are connected to a bus 34. The database 24 and the communication interface 26 are also connected to the bus 34. The communication interface 26 is connected to a network 54. An example of the network 54 is a WAN (Wide Area Network) and / or a LAN (Local Area Network).
[1272] The robot 414 includes a computer 36, a microphone 238, a speaker 240, a camera 42, a communication interface 44, and a controlled object 443. The computer 36 includes a processor 46, RAM 48, and storage 50. The processor 46, RAM 48, and storage 50 are connected to a bus 52. The microphone 238, speaker 240, camera 42, and controlled object 443 are also connected to the bus 52.
[1273] The microphone 238 receives voice signals from the user 20 and receives instructions from the user 20. The microphone 238 captures the voice signals from the user 20, converts the captured voice into audio data, and outputs it to the processor 46. The speaker 240 outputs audio according to the instructions from the processor 46.
[1274] Camera 42 is a small digital camera equipped with an optical system including a lens, aperture, and shutter, and an image sensor such as a CMOS (Complementary Metal-Oxide-Semiconductor) image sensor or a CCD (Charge Coupled Device) image sensor, and captures images of the area around the user 20 (for example, an imaging range defined by a field of view equivalent to the width of a typical healthy person's field of vision).
[1275] Communication interface 44 is connected to network 54. Communication interfaces 44 and 26 are responsible for the exchange of various information between processor 46 and processor 28 via network 54. The exchange of various information between processor 46 and processor 28 using communication interfaces 44 and 26 is performed in a secure manner.
[1276] The controlled object 443 includes a display device, LEDs in the eyes, and motors that drive the arms, hands, and feet. The posture and gestures of the robot 414 are controlled by controlling the motors of the arms, hands, and feet. Some of the robot 414's emotions can be expressed by controlling these motors. Furthermore, the robot 414's facial expressions can also be expressed by controlling the illumination state of the LEDs in its eyes.
[1277] Figure 8 shows an example of the main functions of the data processing device 12 and the robot 414. As shown in Figure 8, the data processing device 12 performs specific processing using the processor 28. The storage 32 stores the specific processing program 56.
[1278] The specific processing program 56 is an example of a "program" relating to the technology of this disclosure. The processor 28 reads the specific processing program 56 from the storage 32 and executes the read specific processing program 56 on the RAM 30. The specific processing is realized by the processor 28 operating as a specific processing unit 290 in accordance with the specific processing program 56 executed on the RAM 30.
[1279] The storage 32 stores the data generation model 58 and the emotion identification model 59. The data generation model 58 and the emotion identification model 59 are used by the identification processing unit 290.
[1280] In robot 414, the processor 46 performs the reception output processing. The storage 50 stores the reception output program 60. The processor 46 reads the reception output program 60 from the storage 50 and executes the read reception output program 60 on the RAM 48. The reception output processing is realized by the processor 46 operating as a control unit 46A according to the reception output program 60 executed on the RAM 48.
[1281] Next, the identification process performed by the identification processing unit 290 of the data processing device 12 will be described.
[1282] "Example of form 1"
[1283] One embodiment of the present invention provides a means for a human to switch to an AI automated response system when a phone call comes in. Specifically, when a phone call comes in, the person answering the call can switch to the AI automated response system by pressing a specific button.
[1284] "Example of form 2"
[1285] Furthermore, the AI automated response system saves the conversation content and summarizes the potential for perjury in the other party's statements. Specifically, the AI automated response system transcribes the other party's statements into text and stores that text in a database. It also analyzes the text of the other party's statements to determine whether they are potentially perjury.
[1286] "Example of form 3"
[1287] Furthermore, the AI automated response will suggest the next action. Specifically, based on the result of the assessment of the perjury potential of the other party's statement, the AI automated response will suggest what action should be taken next. For example, if there is a high probability that the other party's statement is perjury, the AI automated response will suggest reporting it to the police.
[1288] The following describes the processing flow for each example of the form.
[1289] "Example of form 1"
[1290] Step 1: When a phone call comes in, the person answering the phone presses a specific button.
[1291] Step 2: When a specific button is pressed, the system switches to AI automated response.
[1292] Step 3: The AI automated response system answers the call and starts the conversation.
[1293] "Example of form 2"
[1294] Step 1: The AI automated response system transcribes the conversation into text.
[1295] Step 2: Save the transcribed conversation content to the database.
[1296] Step 3: The AI automated response analyzes the saved text to determine whether the other party's statement is perjury.
[1297] "Example of form 3"
[1298] Step 1: The AI automated response system determines the next action based on its assessment of the perjury potential of the other party's statement.
[1299] Step 2: The AI automated response system suggests the next action and notifies the person who answered the call.
[1300] Step 3: The person who received the call takes action based on the AI automated response's suggestions. For example, they might call the police.
[1301] (Example 1)
[1302] Next, we will describe Embodiment 1 of Example Form 1. In the following description, the data processing device 12 will be referred to as the "server" and the robot 414 as the "terminal".
[1303] Traditional communication systems required human intervention upon receiving a call, leading to challenges in response efficiency and accuracy. Furthermore, the lack of features to assess the reliability of the caller's statements and suggest subsequent actions placed a significant burden on users.
[1304] The identification process performed by the identification processing unit 290 of the data processing device 12 in Example 1 is realized by the following means.
[1305] In this invention, the server includes means for switching received communications from a human to an automated response system, means for the automated response system to record the content of the conversation and evaluate the reliability of the other party's statements, and means for the automated response system to propose the next action. This enables efficient communication responses and reduces the burden on the user.
[1306] "A means of switching received communications from a human to an automated response system" refers to a function in which, when a communication device receives an incoming call, the user performs a specific operation to transfer the response from a human to an automated response system.
[1307] "A means by which an automated response system records the content of a conversation and evaluates the reliability of the other party's statements" refers to a function in which an automated response system saves the conversation in progress, analyzes the content of the statements, and determines their reliability.
[1308] "Means by which an automated response system suggests the next action" refers to a function in which an automated response system suggests the next action the user should take based on the content of the conversation.
[1309] "Means for detecting a specific operation and activating an automatic response device" refers to a function that detects an operation performed by a user, such as pressing a specific button on a communication device, and activates the automatic response device.
[1310] "A means by which an automated response device can offer a greeting at the start of a conversation" refers to a function in which an automated response device plays a pre-set greeting message to the other party at the start of communication.
[1311] A description of embodiments for carrying out this invention will be given.
[1312] The server receives incoming signals from communication devices and detects when a user performs a specific action. Specifically, when a user presses the "AI Response" button on their device, the server activates an automated response system. This automated response system uses AI software such as Google Dialogflow or Amazon Lex to record the conversation and evaluate the reliability of the other party's statements.
[1313] When the user presses a button on the terminal, the terminal sends that information to the server. Based on the received information, the server activates an automated response system and hands over the communication response to the AI. The automated response system plays a greeting message at the start of the conversation, such as, "Hello, this is the automated response system. How can I help you?"
[1314] As a concrete example, when a user receives a call, pressing the "AI Answer" button on their device activates the automated answering system, and the AI greets the caller. This process allows the user to leave the phone call to the AI.
[1315] An example of a prompt to input into a generating AI model is, "Please tell me the specific steps to switch to AI automatic answering when a phone call comes in." This prompt allows the AI to provide detailed information about how the system works and how to configure it.
[1316] The flow of the specific processing in Example 1 will be explained using Figure 11.
[1317] Step 1:
[1318] The terminal detects when a communication device receives an incoming signal. The input is the incoming signal from the communication device. The terminal sends this signal to the server, notifying it of the incoming call. The output is an incoming call notification sent to the server.
[1319] Step 2:
[1320] The user presses the "AI Response" button on the device. The input is the user's button press. The device detects this action and notifies the server that the button has been pressed. The output is a button press notification sent to the server.
[1321] Step 3:
[1322] The server receives a button press notification from the terminal. The input is the notification from the terminal. Based on this notification, the server generates an instruction to activate the automated response system. The output is an instruction to activate the automated response system.
[1323] Step 4:
[1324] The automated response system starts operating upon receiving a startup command from the server. Its input is the startup command from the server. At the start of the conversation, the automated response system plays a greeting message such as, "Hello, this is the automated response system. How can I help you?" Its output is a greeting message for the recipient.
[1325] Step 5:
[1326] The automated response system records the content of the conversation with the caller and evaluates the reliability of the caller's statements. The input is the caller's statements. The automated response system analyzes the statements and processes the data to evaluate their reliability. The output is the reliability evaluation result.
[1327] (Application Example 1)
[1328] Next, we will describe Application Example 1 of Form Example 1. In the following description, the data processing device 12 will be referred to as the "server" and the robot 414 as the "terminal".
[1329] In today's communication environment, suspicious communications and fraudulent phone calls are on the rise, requiring a swift and effective response. However, traditional methods require human recipients to handle all inquiries, placing a heavy burden on them. Furthermore, there is a risk of errors due to misjudgments. To address these challenges, the introduction of an automated response system utilizing artificial intelligence is necessary.
[1330] The specific processing performed by the specific processing unit 290 of the data processing device 12 in Application Example 1 is realized by the following means.
[1331] In this invention, the server includes means for switching received communications from human to artificial intelligence automated response, means for the artificial intelligence automated response to save the conversation content and summarize the perjury of the other party's statements, means for the artificial intelligence automated response to propose the next action, and means for the artificial intelligence automated response to respond to suspicious communications based on pre-set instructions. This enables a rapid and accurate response to suspicious communications.
[1332] "Received communications" refers to the transmission of information from external sources, such as phone calls and messages.
[1333] "Artificial intelligence automated response" refers to a system that uses artificial intelligence technology to automatically generate responses and send replies to the communication partner.
[1334] "Saving conversation content" refers to recording information exchanged during communication so that it can be referenced later.
[1335] "Summarizing the perjury" refers to evaluating the veracity of the other party's statements and identifying any questionable points.
[1336] "Suggesting the next course of action" means suggesting the appropriate course of action to take based on the content of the communication.
[1337] "Suspicious communications" refer to the transmission of information that differs from normal communications and may be fraudulent or illegal.
[1338] "Pre-set instructions" refer to guidelines for responses and actions that artificial intelligence should follow in specific situations.
[1339] The system for implementing this invention mainly consists of a server and a terminal. The server runs a program to switch received communications from human interaction to automated artificial intelligence responses. Specifically, when a communication is received, the server activates an automated artificial intelligence response based on instructions from the terminal and saves the conversation content. The saved data is used to evaluate the perjury potential of the other party's statements.
[1340] The server uses a generative AI model to propose the next action based on the content of the communication. This uses generative AI models such as OpenAI's GPT-3. The server responds to suspicious communications based on pre-configured instructions. This allows users to respond quickly and accurately to suspicious communications.
[1341] As a concrete example, when a device receives a suspicious call, the user presses a specific button, and the server activates an AI automated response system that says, "This is the security service. How can I help you?" An example of the prompt might be, "Please provide an example of how to respond to a suspicious call. Please emphasize that this is a security service response."
[1342] The flow of a specific process in Application Example 1 will be explained using Figure 12.
[1343] Step 1:
[1344] The terminal receives a communication. If the user determines the communication is suspicious, they press a specific button on the terminal. This causes the terminal to send a signal to the server requesting the activation of the artificial intelligence automated response system. The input is the received communication, and the output is the activation request signal to the server.
[1345] Step 2:
[1346] The server receives an activation request signal from the terminal and activates an artificial intelligence automated response system. The server uses a generative AI model to generate a response based on the prompt. The input is the activation request signal and prompt from the terminal, and the output is the generated response. Specifically, the server inputs the prompt "Please provide an example of how to respond to a suspicious call. Emphasize that this is a security service response." into the generative AI model and generates a response.
[1347] Step 3:
[1348] The server sends the generated response to the terminal. The terminal then sends this response to the communication partner. The input is the generated response from the server, and the output is the response sent to the communication partner. Specifically, the terminal sends a response to the communication partner saying, "This is the security service. How can I help you?"
[1349] Step 4:
[1350] The server stores the content of the communication and evaluates the perjury potential of the other party's statements. The input is the content of the communication, and the output is the result of the perjury evaluation. Specifically, the server records the conversation content in a database and analyzes the reliability of the statements using natural language processing techniques.
[1351] Step 5:
[1352] Based on the results of the perjury assessment, the server proposes the next action. The input is the result of the perjury assessment, and the output is the proposed action. Specifically, the server generates a suggestion such as "This communication is suspicious. Further verification is required," and sends it to the terminal.
[1353] (Example 2)
[1354] Next, we will describe Example 2 of the morphological example. In the following description, the data processing device 12 will be referred to as the "server" and the robot 414 as the "terminal".
[1355] Conventional automated response systems have shortcomings, such as insufficient recording of conversation content and evaluation of the reliability of statements, making them unable to appropriately suggest the next course of action for the user. Furthermore, because fraud detection based on perjury assessment is not performed, users are vulnerable to fraudulent activity.
[1356] The identification process performed by the identification processing unit 290 of the data processing device 12 in Example 2 is realized by the following means.
[1357] In this invention, the server includes means for acquiring voice input, means for converting voice data into text, and means for storing the text data in a storage device. This enables accurate recording of conversation content and evaluation of the perjury potential of statements.
[1358] "Means for acquiring voice input" refers to a device or method for receiving a user's voice and processing it as digital data.
[1359] "Means for converting audio data to text" refers to a technology or device that analyzes an audio signal and converts it into a corresponding string of characters.
[1360] "Means for storing text data in a storage device" refers to a method or apparatus for recording converted text data in a database or other storage medium.
[1361] "Means for analyzing text data and determining the perjury of a statement" refers to a technology or device for processing text data and evaluating the reliability and truthfulness of its content.
[1362] "Means for outputting analysis results" refers to a method or apparatus for presenting the results of text data analysis to the user.
[1363] "Means for switching received communications from human to automated responses" refers to a method or device for transferring a human response to an automated response system upon receiving a communication.
[1364] "Means of suggesting the next action through automated response" refers to a technology or device that suggests the next action a user should take based on analysis results.
[1365] This invention is a system that determines the perjury of a statement by acquiring voice input, converting it into text data, and analyzing it. A specific embodiment of this system is shown below.
[1366] The user speaks aloud through the device's microphone. The device captures this audio as digital data and sends it to the server. The server uses speech recognition software to convert the audio data into text. High-precision text conversion is achieved by using speech recognition technologies such as the Google Cloud Speech-to-Text API.
[1367] The converted text data is stored on a storage device by the server. A database management system is used for storage, and metadata such as the message timestamp and user ID are also recorded.
[1368] Next, the server analyzes the text data using a natural language processing library. Specifically, it performs grammatical analysis using spaCy and then scrutinizes the content of the statements using OpenAI's GPT generative AI model. This allows it to determine the perjury potential of the statements and detect inconsistencies and unnatural points.
[1369] The analysis results are sent from the server to the terminal and presented to the user. The user can then review the results on the terminal screen and decide on their next course of action.
[1370] For example, if a user says, "It didn't rain yesterday," the system transcribes the statement into text and saves it to a database. Then, it refers to weather information and compares it with the actual weather to determine whether the statement is false.
[1371] An example of a prompt would be: "Assess the likelihood that the following statement is perjury: 'It didn't rain yesterday.' Please take actual weather data into consideration."
[1372] In this way, the system can automate the entire process from voice input to perjury detection, enabling it to provide users with quick and accurate information.
[1373] The flow of the specific processing in Example 2 will be explained using Figure 13.
[1374] Step 1:
[1375] The user speaks aloud through the device's microphone. The device captures this audio as digital data. The input is the user's voice, and the output is digital audio data. The device prompts the user to press the "Start Recording" button to begin voice input.
[1376] Step 2:
[1377] The device sends the captured audio data to the server. The server uses speech recognition software to convert the audio data into text. The input is digital audio data, and the output is text data. The server uses the Google Cloud Speech-to-Text API to convert the audio to text.
[1378] Step 3:
[1379] The server stores the converted text data in storage. The input is text data, and the output is text data stored in the database. The server executes SQL queries to insert the text data and associated metadata into the database.
[1380] Step 4:
[1381] The server analyzes stored text data to determine the perjury potential of a statement. The input is text data obtained from a database, and the output is the result of the perjury evaluation. The server performs grammatical analysis using the natural language processing library spaCy and scrutinizes the content of the statement using OpenAI's GPT generative AI model.
[1382] Step 5:
[1383] The server sends the analysis results to the terminal and presents them to the user. The input is the result of the perjury assessment, and the output is the result displayed on the terminal's screen. The terminal displays the analysis results on the screen, and the user confirms the results.
[1384] (Application Example 2)
[1385] Next, we will describe application example 2 of form example 2. In the following description, the data processing device 12 will be referred to as the "server" and the robot 414 as the "terminal".
[1386] In today's communication environment, real-time evaluation of conversation reliability is crucial. However, conventional technologies have struggled to efficiently determine the perjury potential of statements and present this information to users immediately. Therefore, new methods are needed to improve conversation reliability.
[1387] The specific processing performed by the specific processing unit 290 of the data processing device 12 in Application Example 2 is realized by the following means.
[1388] In this invention, the server includes means for switching received communications from human to artificial intelligence automated response, means for the artificial intelligence automated response to save the conversation content and summarize the perjury potential of the other party's statements, and means for the artificial intelligence automated response to suggest the next action. This makes it possible to evaluate the reliability of the conversation in real time and present it to the user immediately.
[1389] "Communication" refers to the act or means of sending and receiving voice or data.
[1390] "Artificial intelligence automated response" refers to a system that uses artificial intelligence technology to automatically generate responses and engage in conversation.
[1391] "Conversation content" refers to the entirety of the statements and information exchanged during communication.
[1392] "Perjury potential" is an indicator of the degree to which a statement is likely to be false.
[1393] "Speech recognition technology" is a technology that converts speech into text.
[1394] "Natural language processing technology" is a technology used to analyze text data and understand its meaning and intent.
[1395] A "display device" is a device used to present information visually.
[1396] "User" refers to an individual or group that uses the system.
[1397] "Action" refers to the next step or action suggested by the system.
[1398] The system for carrying out this invention includes a server and a terminal. The server has the function of switching received communications from human to artificial intelligence automated responses. The terminal uses speech recognition technology to transcribe speech during communication into text in real time. Specifically, it uses the Google Cloud Speech-to-Text API to convert speech to text.
[1399] The server analyzes the transcribed speech using natural language processing techniques. This analysis utilizes natural language processing libraries such as spaCy and Transformers. The analysis results in an evaluation of the perjury potential of the speech. This evaluation uses a pre-trained generative AI model (e.g., OpenAI GPT-3).
[1400] The evaluation results are displayed in real time on the terminal's display device. This allows users to instantly verify the reliability of the conversation. For example, if someone says "This product is the best-selling in the market" during a meeting, the server will refer to past data and market information to evaluate the reliability of the statement.
[1401] An example of a prompt for a generative AI model might be: "Evaluate the likelihood that this statement is true. Market data is as follows." Using this prompt, the model determines the reliability of the statement and provides the result to the user.
[1402] The flow of a specific process in Application Example 2 will be explained using Figure 14.
[1403] Step 1:
[1404] The device acquires audio during communication via its microphone. The input is audio data, which is converted into text data using the Google Cloud Speech-to-Text API. The output is the transcribed speech.
[1405] Step 2:
[1406] The server receives text data sent from the terminal. The input is text data, which is then analyzed using a natural language processing library (e.g., spaCy, Transformers). The purpose of the analysis is to understand the content of the statement and extract information to assess its perjury potential. The output is the analysis result.
[1407] Step 3:
[1408] The server sends a prompt to a generative AI model (e.g., OpenAI GPT-3) based on the analysis results. The input consists of the analysis results and the prompt text. An example of the prompt text is "Evaluate the likelihood that this statement is true. The market data is as follows." The model evaluates the reliability of the statement and outputs the result.
[1409] Step 4:
[1410] The server receives evaluation results from the generated AI model and sends them to the terminal. The input is the evaluation result, and the output is reliability evaluation information to be presented to the user.
[1411] Step 5:
[1412] The terminal displays reliability evaluation information received from the server on its display device. The input is reliability evaluation information, and the output is information that the user can visually confirm. This allows the user to instantly verify the reliability of the conversation.
[1413] (Example 3)
[1414] Next, we will describe Embodiment 3 of Example 3. In the following description, the data processing device 12 will be referred to as the "server" and the robot 414 as the "terminal".
[1415] In today's communication environment, there is a need to quickly and accurately determine the veracity of what the other party says and to suggest appropriate actions. However, conventional systems have the challenge of having a complex process for evaluating the veracity of statements, making it difficult to clearly indicate the next action the user should take.
[1416] The identification process performed by the identification processing unit 290 of the data processing device 12 in Example 3 is realized by the following means.
[1417] In this invention, the server includes means for switching received communications from a human to an automated response system, means for the automated response system to store conversation information and evaluate the veracity of the other party's statements, and means for the automated response system to propose the next action based on the evaluation result. This makes it possible to quickly evaluate the veracity of statements and propose appropriate actions to the user.
[1418] "Received communications" refers to audio and data signals sent from external sources.
[1419] An "automatic response device" refers to a device that has the function of automatically responding to received communications.
[1420] "Conversation information" refers to the content of audio and text exchanged during communication.
[1421] "Evaluating the truthfulness of a statement or piece of information" refers to the process of judging the accuracy and reliability of such statements or information.
[1422] "Suggesting action" means indicating specific actions to take next based on the evaluation results.
[1423] A "communication terminal" refers to a device used by a user that has the function of performing communication.
[1424] "User" refers to a person who uses the system.
[1425] A description of embodiments for carrying out this invention will be given.
[1426] First, the user enters a prompt message using a communication terminal. This prompt message is used to determine the truthfulness of the other party's statement. As a concrete example, consider the prompt message, "Is this statement true?"
[1427] Next, the terminal sends the entered prompt message to the server. The server uses a generative AI model to analyze the received prompt message. This analysis utilizes natural language processing techniques, employing models such as "OpenAI GPT-3" and "Google BERT." The server uses these models to scrutinize the content of the utterance and evaluate its truthfulness.
[1428] Based on the evaluation results, the server suggests the next course of action. For example, if it is determined that the statement is highly likely to be perjury, the server will decide to "suggest reporting to the police."
[1429] The server then sends the suggested action to the terminal. The terminal displays the received suggestion to the user. The user can then use this suggestion to decide on their next action.
[1430] In this way, it becomes possible to quickly evaluate the veracity of statements and propose appropriate actions. The flow of the specific processing in Example 3 will be explained using Figure 15.
[1431] Step 1:
[1432] The user enters a prompt message using a communication terminal. For example, they might enter the prompt message, "Is this statement true?" This input serves as the basis for determining the truthfulness of the other party's statement.
[1433] Step 2:
[1434] The terminal sends the entered prompt message to the server. Here, the input is the prompt message, and the output is the data sent to the server. The terminal securely transmits the data using the HTTPS protocol.
[1435] Step 3:
[1436] The server inputs the received prompt text into a generating AI model. The input is the prompt text, and the output is the analysis result by the AI model. The server uses natural language processing models such as "OpenAI GPT-3" and "Google BERT" to analyze the content of the utterance and evaluate its truthfulness.
[1437] Step 4:
[1438] The server proposes the next course of action based on the analysis results. The input is the analysis results of the AI model, and the output is the proposed action. For example, if the server determines that there is a high probability that the statement is perjury, it will decide to take the action of "I suggest reporting this to the police."
[1439] Step 5:
[1440] The server sends the proposed action to the terminal. The input is the proposed action, and the output is the transmission of data to the terminal. The server securely transmits the data using the HTTPS protocol.
[1441] Step 6:
[1442] The terminal displays received suggestions to the user. The input is suggestions from the server, and the output is what is displayed to the user. The terminal displays the suggestions on the screen, allowing the user to decide on their next action.
[1443] (Application Example 3)
[1444] Next, we will describe application example 3 of form example 3. In the following description, the data processing device 12 will be referred to as the "server" and the robot 414 as the "terminal".
[1445] In modern society, the spread of false information and fraudulent activities through communications is increasing, and there is a need for a swift and effective response. However, conventional systems have difficulty determining the falsehood of communications in real time and suggesting appropriate actions. As a result, the risk of users becoming involved in fraudulent activities is increasing.
[1446] The specific processing performed by...
Claims
1. Equipped with a processor, The aforementioned processor, Switching the received communication from human to AI automated response, Audio data is acquired and converted into text data using speech recognition technology. The text data is analyzed using natural language processing techniques. Based on the text data, a prompt is generated to determine whether the statements made by the caller are perjudicial, and based on the prompt, the perjudicial status is determined. Based on the results of the perjury determination, the following actions should be suggested to the user: The system analyzes the user's emotional state using an emotion analysis engine, generates prompts to adjust suggested actions based on the analysis results, and adjusts the actions based on the prompts. system.
2. The system according to claim 1, wherein the processor generates a prompt to propose the next action based on the result of the perjury determination, and proposes the action based on the prompt.