System

The system addresses the challenge of real-time fraud detection during phone calls by recording, converting audio to text, analyzing for fraud, and generating automated responses, effectively preventing fraud and safeguarding users.

JP2026036106APending Publication Date: 2026-03-05SOFTBANK GROUP CORP
View PDF 1 Cites 0 Cited by

Patent Information

Application Number
JP2024138621
Authority / Receiving Office
JP · JP
Patent Type
Applications
Current Assignee / Owner
Filing Date
2024-08-20
Publication Date
2026-03-05

AI Technical Summary

Technical Problem

Current systems are inadequate in detecting and preventing fraud during phone calls, particularly affecting elderly individuals, as they lack real-time fraud detection and response capabilities.

Method used

A system that automatically records calls, converts audio to text using generative AI, analyzes for fraud keywords, takes over conversations if fraud is suspected, generates automated responses, and stores records for later provision to authorities.

Benefits of technology

Enables real-time fraud detection and prevention, reducing the risk of users becoming victims by automatically responding to suspected fraud and providing evidence for authorities.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure 2026036106000001_ABST
    Figure 2026036106000001_ABST
Patent Text Reader

Abstract

A system is provided.SOLUTION: A system including means for automatically starting recording when a call is started, means for transmitting voice data of the call, means for converting the received voice data into a text by a generation AI, means for detecting a keyword related to fraud, means for automatically taking over a conversation by an AI when there is a suspicion of fraud, and means for storing and providing a record suspected of fraud.SELECTED DRAWING: Figure 1
Need to check novelty before this filing date? Find Prior Art

Description

[Technical Field]

[0001] The technology of the present disclosure relates to a system. [Background technology]

[0002] Patent document 1 discloses a persona chatbot control method performed by at least one processor, the method including the steps of receiving a user utterance, adding the user utterance to a prompt including an instruction sentence related to a description of the chatbot character, encoding the prompt, and inputting the encoded prompt into a language model to generate a chatbot utterance in response to the user utterance. [Prior art documents] [Patent documents]

[0003] [Patent Document 1] Japanese Patent Publication No. 2022-180282 Summary of the Invention [Problem to be solved by the invention]

[0004] Special frauds remain common, with the majority of victims being elderly. Current countermeasures are insufficient in many areas, making it difficult for elderly people to use the telephone safely. This invention aims to create an environment in which all users, including elderly people, can use the telephone safely by providing a system that immediately detects suspected fraud during a call and automatically blocks the conversation with the fraudster, thereby preventing fraud before it occurs. [Means for solving the problem]

[0005] This invention relates to a system that automatically starts recording when a call is initiated and converts the audio data of the call into text in real time using a generation AI. It also has a means for analyzing the converted conversation and detecting keywords related to fraud. It also has a means for the AI ​​to automatically take over the conversation if fraud is suspected, and includes a means for saving and providing records of suspected fraud. It also includes a means for the generation AI to generate and play back an automated response message, and a means for notifying the user if suspected fraud is detected. This prevents fraud before it occurs and provides an environment where telephones can be used safely.

[0006] "Call" means the activity of transmitting and receiving audio between two parties via telephone equipment.

[0007] "Recording" refers to the act of saving audio data during a call on a device.

[0008] "Generative AI" refers to a system that uses artificial intelligence technology to convert voice data into text data.

[0009] "Text conversion" refers to the process of analyzing audio data and expressing it as written information.

[0010] "Scam-related keywords" refer to specific words and phrases that are frequently used in scam conversations.

[0011] "AI taking over the conversation" refers to the act of artificial intelligence continuing the call on behalf of the user.

[0012] "Storing records" refers to the act of storing call data in a storage device such as a database.

[0013] "Means of providing" refers to the method of providing stored data to a third party organization.

[0014] "Automatic response AI" refers to a system that uses artificial intelligence to automatically generate appropriate responses and continue the call.

[0015] "Playback as audio" refers to the act of converting text data into audio and outputting it.

[0016] "Notifying the user" refers to the act of the system informing the user of specific information. [Brief explanation of the drawings]

[0017] [Figure 1] 1 is a conceptual diagram showing an example of the configuration of a data processing system according to a first embodiment. [Figure 2] 1 is a conceptual diagram showing an example of main functions of a data processing device and a smart device according to a first embodiment. [Figure 3] FIG. 10 is a conceptual diagram showing an example of the configuration of a data processing system according to a second embodiment. [Figure 4] FIG. 10 is a conceptual diagram showing an example of main functions of a data processing device and smart glasses according to a second embodiment. [Figure 5] FIG. 10 is a conceptual diagram showing an example of the configuration of a data processing system according to a third embodiment. [Figure 6] FIG. 11 is a conceptual diagram showing an example of main functions of a data processing device and a headset-type terminal according to a third embodiment. [Figure 7] FIG. 10 is a conceptual diagram showing an example of the configuration of a data processing system according to a fourth embodiment. [Figure 8] FIG. 10 is a conceptual diagram showing an example of main functions of a data processing device and a robot according to a fourth embodiment. [Figure 9] 1 shows an emotion map onto which multiple emotions are mapped. [Figure 10] 1 shows an emotion map onto which multiple emotions are mapped. [Figure 11] FIG. 3 is a sequence diagram showing a processing flow of the data processing system according to the first embodiment. [Figure 12] FIG. 10 is a sequence diagram showing the flow of processing in the data processing system in Application Example 1. [Figure 13]FIG. 10 is a sequence diagram showing the flow of processing in the data processing system according to the second embodiment when an emotion engine is combined. [Figure 14] FIG. 10 is a sequence diagram showing the flow of processing in the data processing system in Application Example 2 when an emotion engine is combined. DETAILED DESCRIPTION OF THE INVENTION

[0018] An example of an embodiment of a system according to the technology of the present disclosure will be described below with reference to the accompanying drawings.

[0019] First, the terms used in the following description will be explained.

[0020] In the following embodiments, a coded processor (hereinafter simply referred to as a "processor") may be a single arithmetic device or a combination of multiple arithmetic devices. Furthermore, a processor may be a single type of arithmetic device or a combination of multiple types of arithmetic devices. Examples of arithmetic devices include a CPU (Central Processing Unit), a GPU (Graphics Processing Unit), a GPGPU (General-Purpose computing on Graphics Processing Units), and an APU (Accelerated Processing Unit).

[0021] In the following embodiments, a coded RAM (Random Access Memory) is a memory in which information is temporarily stored and is used as a working memory by a processor.

[0022] In the following embodiments, the coded storage is one or more non-volatile storage devices that store various programs, various parameters, etc. Examples of non-volatile storage devices include flash memory (SSD (Solid State Drive)), magnetic disks (e.g., hard disks), and magnetic tapes.

[0023] In the following embodiments, a communication I / F (Interface) with a symbol is an interface including a communication processor, an antenna, etc. The communication I / F controls communication between multiple computers. Examples of communication standards applied to the communication I / F include wireless communication standards including 5G (5th Generation Mobile Communication System), Wi-Fi (registered trademark), Bluetooth (registered trademark), etc.

[0024] In the following embodiments, "A and / or B" is synonymous with "at least one of A and B." In other words, "A and / or B" means that it may be only A, only B, or a combination of A and B. Furthermore, in this specification, the same concept as "A and / or B" is also applied when three or more things are expressed connected by "and / or."

[0025] [First embodiment]

[0026] FIG. 1 shows an example of the configuration of a data processing system 10 according to the first embodiment.

[0027] 1, a data processing system 10 includes a data processing device 12 and a smart device 14. An example of the data processing device 12 is a server.

[0028] The data processing device 12 includes a computer 22, a database 24, and a communication I / F 26. The computer 22 is an example of a "computer" according to the technology of the present disclosure. The computer 22 includes a processor 28, a RAM 30, and a storage 32. The processor 28, the RAM 30, and the storage 32 are connected to a bus 34. The database 24 and the communication I / F 26 are also connected to the bus 34. The communication I / F 26 is connected to a network 54. Examples of the network 54 include a WAN (Wide Area Network) and / or a LAN (Local Area Network).

[0029] The smart device 14 includes a computer 36, a reception device 38, an output device 40, a camera 42, and a communication I / F 44. The computer 36 includes a processor 46, a RAM 48, and a storage 50. The processor 46, the RAM 48, and the storage 50 are connected to a bus 52. The reception device 38, the output device 40, and the camera 42 are also connected to the bus 52.

[0030] The reception device 38 includes a touch panel 38A, a microphone 38B, and the like, and receives user input. The touch panel 38A detects contact with an indicator (for example, a pen or a finger) to receive user input by the touch of the indicator. The microphone 38B detects the user's voice to receive user input by voice. The control unit 46A transmits data indicating the user input received by the touch panel 38A and the microphone 38B to the data processing device 12. In the data processing device 12, the specific processing unit 290 acquires the data indicating the user input.

[0031] The output device 40 includes a display 40A and a speaker 40B, and presents data to the user 20 by outputting the data in a form of expression that the user 20 can perceive (for example, audio and / or text). The display 40A displays visible information such as text and images in accordance with instructions from the processor 46. The speaker 40B outputs audio in accordance with instructions from the processor 46. The camera 42 is a compact digital camera equipped with an optical system including a lens, aperture, and shutter, and an imaging element such as a CMOS (Complementary Metal-Oxide-Semiconductor) image sensor or a CCD (Charge Coupled Device) image sensor.

[0032] The communication I / F 44 is connected to a network 54. The communication I / Fs 44 and 26 control the exchange of various information between the processor 46 and the processor 28 via the network 54.

[0033] FIG. 2 shows an example of the main functions of the data processing device 12 and the smart device 14.

[0034] 2, in the data processing device 12, a specific process is performed by the processor 28. A specific processing program 56 is stored in the storage 32. The specific processing program 56 is an example of a "program" according to the technology of the present disclosure. The processor 28 reads the specific processing program 56 from the storage 32 and executes the read specific processing program 56 on the RAM 30. The specific process is realized by the processor 28 operating as a specific processing unit 290 in accordance with the specific processing program 56 executed on the RAM 30.

[0035] The storage 32 stores a data generation model 58 and an emotion identification model 59. The data generation model 58 and the emotion identification model 59 are used by the identification processing unit 290.

[0036] In the smart device 14, the processor 46 performs the reception output process. The storage 50 stores a reception output program 60. The reception output program 60 is used in conjunction with the specific processing program 56 by the data processing system 10. The processor 46 reads the reception output program 60 from the storage 50 and executes the read reception output program 60 on the RAM 48. The reception output process is realized by the processor 46 operating as the control unit 46A in accordance with the reception output program 60 executed on the RAM 48.

[0037] Next, a description will be given of the specific processing performed by the specific processing unit 290 of the data processing device 12. In the following description, the data processing device 12 will be referred to as a "server" and the smart device 14 will be referred to as a "terminal."

[0038] This invention relates to a system for preventing special fraud during a call. Below, each element of this system and its processing will be explained in natural language.

[0039] Start recording a call

[0040] 1. Device:

[0041] It has a function that automatically starts recording when a call is started, so that audio data is recorded from the moment the user starts a call.

[0042] Sending audio data

[0043] 2. Terminal:

[0044] Recorded audio data is sent to the server in fixed buffer sizes, allowing the content of the call to be analyzed in real time.

[0045] Converting audio data to text

[0046] 3. Server:

[0047] The voice data sent from the device is received and converted into text in real time using a generation AI, which uses voice recognition technology to convert it into text information accurately and quickly.

[0048] Text data analysis

[0049] 4. Server:

[0050] The text of the conversation is analyzed, where fraud-related keywords and similar words are retrieved from a database, and an analysis algorithm is used to determine whether there is any suspicion of fraud.

[0051] What to do if fraud is suspected

[0052] 5. Server:

[0053] If fraud is suspected, the AI ​​automatically takes over the conversation, generating an appropriate response and sending it to the device.

[0054] 6. Terminal:

[0055] It responds on behalf of the user by playing back an automated response received from the server, which is designed to be unpleasant for scammers.

[0056] Record keeping and provision

[0057] 7. Server:

[0058] The data of calls suspected to be fraudulent will be saved and later provided to the police and other relevant authorities. This information will be used to help prevent fraud and to prevent future victims from becoming victims of fraud.

[0059] Specific examples

[0060] Here is a specific scenario:

[0061] 1. User:

[0062] Elderly user receives a call.

[0063] 2. Terminal:

[0064] When a call is initiated, the device automatically starts recording, and the recorded audio data is sent to the server in real time.

[0065] 3. Server:

[0066] The received voice data is converted into text using a generative AI, and keywords related to fraud, such as "depositing money" and "bank account," are analyzed.

[0067] If the server suspects fraud, it will activate an automated response AI that will generate a response such as "Please wait a moment, we're checking."

[0068] 4. Terminal:

[0069] Based on instructions from the server, an automated response is played and the conversation continues to be recorded.

[0070] 5. Server:

[0071] After the service is terminated, conversation data suspected of fraud will be saved and provided to the police on a regular basis.

[0072] This will help prevent victims of special fraud and allow all users, including the elderly, to use the telephone with peace of mind.

[0073] The processing flow will be explained below.

[0074] Step 1:

[0075] User: Start a call.

[0076] Device: Recording will start automatically when a call is initiated.

[0077] Step 2:

[0078] Terminal: Recorded audio data is sent to the server at fixed buffer size intervals.

[0079] Step 3:

[0080] Server: Receives voice data sent from the device and converts the voice into text using generative AI.

[0081] Step 4:

[0082] Server: Analyzes the text of the conversation, retrieves fraud-related keywords and similar words from a database, and determines whether there is any suspicion of fraud.

[0083] Step 5:

[0084] Server: If fraud is suspected, an automated response AI is activated to generate an appropriate response.

[0085] Terminal: Plays back the generated response aloud and acts on behalf of the user.

[0086] Step 6:

[0087] Device: Continue recording and continue sending conversation data to the server.

[0088] Step 7:

[0089] Server: Stores conversation data suspected of fraud in a database. This data will be provided to relevant agencies such as the police at a later date.

[0090] Step 8:

[0091] Server: If necessary, notify the user of suspected fraud.

[0092] This series of steps will protect users from special fraud by monitoring call content in real time and automatically responding if fraud is suspected.

[0093] Example 1

[0094] Next, a description will be given of Example 1. In the following description, the data processing device 12 will be referred to as a "server" and the smart device 14 will be referred to as a "terminal."

[0095] Special frauds are becoming more sophisticated every day, and many people, including the elderly, are falling victim to them. Effective measures to prevent these frauds are needed, but existing technology makes it difficult to detect fraud during a call in real time and take appropriate action. Furthermore, there is no system that can instantly recognize and respond to suspected fraudulent calls. Therefore, there is a need to provide a system that prevents fraud during a call and create an environment in which users can use the telephone with peace of mind.

[0096] The specific processing by the specific processing unit 290 of the data processing device 12 in the first embodiment is realized by the following means.

[0097] In this invention, the server includes a means for transmitting the voice data of the call at fixed buffer size intervals, a means for converting the received voice data into text in real time using a generation AI, a means for retrieving and detecting fraud-related keywords from a database, a means for the generation AI to automatically take over the conversation and generate and send an appropriate response if fraud is suspected, and a means for saving records of suspected fraud and providing them to relevant authorities. This makes it possible to detect special frauds in real time during a call and automatically provide an appropriate response, significantly reducing the risk of users becoming victims of fraud.

[0098] The "means for automatically starting recording when a call is started" refers to a device or software that automatically starts a recording function the moment the user starts a call and records the audio data of the call.

[0099] The "means for transmitting voice data of a call at fixed buffer size intervals" refers to a device or software for transmitting collected voice data to a server at fixed volume or time intervals.

[0100] "Means for converting received voice data into text in real time using a generation AI" refers to a device or software in which a server receives voice data sent from a terminal and converts the voice into text in real time using a generation AI.

[0101] "Means for retrieving and detecting fraud-related keywords from a database" refers to a device or software for searching a database for keywords that indicate signs of fraud contained in text-converted voice data and detecting them.

[0102] "Means for a generation AI to automatically take over a conversation when fraud is suspected, and generate and send an appropriate response" refers to a device or software that, when it is determined that fraud is suspected, automatically generates an appropriate countermeasure response from a generation AI and sends it on behalf of the user.

[0103] "Means for storing records of suspected fraud and providing them to relevant authorities" refers to devices or software that record the content of suspected fraudulent calls and the results of their analysis, and provide them to relevant authorities at a later date.

[0104] "Means for generating lines for automated responses using a generation AI and playing them audibly" refers to a device or software that converts the text of an automated response generated by a generation AI into audio and plays it back through a speaker or the like.

[0105] "Means for notifying the user when suspected fraud is detected" refers to a device or software that notifies the user of the information when the system determines that there is a suspicion of fraud.

[0106] This invention relates to a system for preventing special fraud during a call. Each element of this system and its processing will be specifically described below.

[0107] System configuration and processing

[0108] Call recording

[0109] When a user picks up the receiver or presses the call button to start a call, the device's recording function automatically starts and records the audio data of the call. Once the recorded audio data reaches a certain buffer size, it is sent from the device to the server. This is usually done using an HTTP POST request.

[0110] Converting audio data to text

[0111] The server receives the voice data sent from the device and converts it into text in real time using a generation AI (e.g., DeepSpeech, Google® Cloud Speech-to-Text, etc.) This generation AI inputs the voice data and outputs the corresponding text data.

[0112] Text data analysis

[0113] The text data is first stored in a database on the server, and then keywords related to fraud are retrieved from the database and detected. This analysis is performed using regular expressions and natural language processing tools (e.g., NLTK). Based on the keyword detection results, a judgment is made as to whether there is suspicion of fraud, and if so, the system proceeds to the next step.

[0114] Generate and send automatic responses

[0115] If fraud is suspected, the server uses a generative AI model (e.g., GPT-3 (registered trademark)) to generate an appropriate automated response. A specific example of the generative AI's prompt is, "This call is suspected to be fraudulent. Please play the next message." The generated automated response is sent to the device, which plays it aloud. This allows the server to respond appropriately to the fraudster on behalf of the user.

[0116] Record keeping and provision

[0117] Records of calls suspected of fraud are stored on a server. This information will be provided to relevant authorities, such as the police, at a later date. The stored data includes voice data, text data, detected keywords, and responses generated by the AI.

[0118] Specific examples

[0119] The following is a specific example:

[0120] The user is an elderly person and receives a phone call. The device automatically records the call and sends the recorded voice data to the server in real time. The server converts the voice data into text using a generative AI and detects fraud-related keywords such as "depositing money" and "bank account." If the server determines that there is suspicion of fraud, it uses the generative AI to generate an automated response such as "Please wait a moment, we're checking," and sends it to the device. The device then plays this automated response aloud to ensure the user's safety.

[0121] This system will prevent victims of special fraud before they occur, and will enable all users, including the elderly, to use the telephone with peace of mind.

[0122] Prompt Sentence Examples

[0123] 1. "This call is suspected to be fraudulent. Please play the following message."

[0124] 2. "Do not provide bank account information."

[0125] 3. "Always check with family and friends before sending money."

[0126] Using such prompts allows the generative AI to provide appropriate warnings and responses to scammers.

[0127] The flow of the identification process in the first embodiment will be described with reference to FIG.

[0128] Step 1:

[0129] The user answers or makes a call. The input is the action of lifting the telephone receiver or pressing the talk button, and the output is the start of the call signal, which starts recording the call.

[0130] Step 2:

[0131] The device detects that a call has started and automatically starts recording. The input is a call start signal, and the output is recording of audio data. The device's internal recording function is activated, and audio data is collected through the microphone.

[0132] Step 3:

[0133] The device collects recorded audio data in fixed buffer sizes and sends them to the server in bulk. The input is the audio data being recorded, and the output is the audio data buffer. When the buffer reaches a certain size, it is sent to the server via an HTTP POST request.

[0134] Step 4:

[0135] The server receives the audio data sent from the device. The input is the audio data sent via an HTTP POST request, and the output is the audio data stored on the server. The server temporarily stores this audio data and proceeds to the next step.

[0136] Step 5:

[0137] The server converts the received voice data into text in real time using a generation AI (e.g., DeepSpeech, Google Cloud Speech-to-Text). The input is voice data, and the output is text data. The generation AI analyzes the voice data and generates the corresponding text.

[0138] Step 6:

[0139] The server analyzes the data converted to text by the generative AI and detects fraud-related keywords. The input is text data, and the output is the keyword detection results. A natural language processing tool (e.g., NLTK) is used to search for fraud-related keywords.

[0140] Step 7:

[0141] The server determines whether fraud is suspected based on the keyword detection results. The input is the keyword detection results, and the output is a flag indicating whether fraud is suspected or not. A judgment algorithm (e.g., decision tree, SVM) is executed.

[0142] Step 8:

[0143] If the server determines that fraud is suspected, it uses a generative AI model (e.g., GPT-3) to generate an appropriate automated response. The input is a flag indicating suspected fraud, and the output is the automated response. The generative AI inputs a prompt sentence and generates a corresponding response.

[0144] Step 9:

[0145] The server sends the generated automatic response to the terminal. The input is the automatic response text, and the output is the data to be sent to the terminal. An HTTP POST request is used for sending.

[0146] Step 10:

[0147] The terminal plays back the received automated response lines aloud. The input is the data sent from the server, and the output is the audio data. The TTS (Text-to-Speech) function is used to convert the text data into audio, which is then played back through the speaker.

[0148] Step 11:

[0149] The server stores records of suspected fraudulent phone calls and periodically provides them to relevant authorities (e.g., police). The input is the call's audio data, text data, and analysis results, and the output is the stored recorded data and data provided to relevant authorities. The records are stored in a database and provided to relevant authorities via an API or interface.

[0150] (Application example 1)

[0151] Next, a description will be given of Application Example 1. In the following description, the data processing device 12 will be referred to as a "server" and the smart device 14 will be referred to as a "terminal."

[0152] In modern society, the number of fraud cases is increasing, and vulnerable people such as the elderly are often the targets. To solve this problem, a system that can identify fraudulent activities in real time and take prompt countermeasures is needed. However, current systems have difficulty monitoring the content of phone calls one by one, making it difficult to prevent fraudulent activities. Therefore, there is a need to provide a system that can detect fraudulent activities early and respond appropriately.

[0153] The specific processing by the specific processing unit 290 of the data processing device 12 in the application example 1 is realized by the following means.

[0154] In this invention, the server includes means for automatically starting recording when a call is initiated, means for transmitting the voice data of the call, means for converting the received voice data into text using a generating AI, means for detecting fraud-related keywords, means for the AI ​​to automatically take over the conversation if fraud is suspected, means for recording data in real time and providing it to relevant authorities, and means for saving and providing records of suspected fraud. This makes it possible to detect special frauds in real time and take appropriate action.

[0155] "Means for automatically starting recording when a call starts" refers to a function that automatically starts recording audio data as soon as a call starts.

[0156] "Means for transmitting voice data of a call" refers to a function for transmitting recorded voice data to a server in real time or at fixed buffer size intervals.

[0157] "Means of converting received voice data into text using generation AI" refers to the function of converting voice data received by the server into text data using generation AI technology.

[0158] "Means for detecting fraud-related keywords" refers to the function of searching for and detecting specific fraud-related keywords from text data.

[0159] "Means for AI to automatically take over the conversation if fraud is suspected" refers to a function in which AI automatically generates and responds to conversations when it is determined that fraud is suspected.

[0160] "Means of recording data in real time and providing it to relevant authorities" refers to the function of recording the contents of phone calls in real time and providing it to relevant authorities such as the police if necessary.

[0161] "Means for storing and providing records of suspected fraud" refers to the function of storing records of phone calls suspected of fraud and providing them to relevant authorities at a later date if necessary.

[0162] "Means for generating lines for automatic responses using generative AI and playing them back as audio" refers to the function of creating the content of automatic responses using generative AI technology and playing them back as audio.

[0163] "Means of notifying users when suspected fraud is detected" refers to a function that warns or notifies users when suspected fraud is detected.

[0164] This invention relates to a system for preventing special fraud during a call. Each element of this system and its processing will be explained below.

[0165] Terminal behavior:

[0166] The device has a means to automatically start recording when a call is initiated, and the recorded audio data is sent to the server in fixed buffer sizes, allowing the content of the call to be analyzed in real time.

[0167] Server behavior:

[0168] The server converts the received voice data into text using a generative AI model, "Google Cloud Speech-to-Text," which allows for accurate and fast conversion of voice to text.

[0169] The text data is then analyzed to detect fraud-related keywords, such as "bank transfer fraud," "depositing money," and "bank account," and an analysis algorithm is used to detect these.

[0170] If fraud is suspected, the server automatically activates AI to generate an appropriate response. This response is generated using a generative AI model such as OpenAI (registered trademark) GPT-4 (registered trademark). The generated response is sent to the device and played back as audio on the device.

[0171] Data storage and provision:

[0172] The server stores records of calls suspected of fraud and has the means to provide them to relevant authorities if necessary, so that police and other relevant authorities can obtain the necessary information at a later date.

[0173] Examples:

[0174] Here's a specific scenario:

[0175] 1. User:

[0176] An elderly user receives a suspicious phone call.

[0177] 2. Terminal:

[0178] Recording will start automatically as soon as the call is started.

[0179] The recorded audio data is sent to the server in real time.

[0180] 3. Server:

[0181] The received voice data is converted into text using the generative AI model "Google Cloud Speech-to-Text."

[0182] From the text data, keywords related to fraud, such as "depositing money" and "bank account," are detected.

[0183] If it determines that fraud is suspected, an automatic response message is generated using OpenAI GPT-4.

[0184] For example, generate a response like "Please wait, we're checking."

[0185] 4. Terminal:

[0186] The automatic response message sent from the server is played back and the conversation is taken over.

[0187] 5. Server:

[0188] Records of calls suspected of fraud will be saved and provided to the police and other relevant authorities as necessary.

[0189] Example prompt sentence:

[0190] System: "If fraud is detected, generate an appropriate response."

[0191] User: "I'm from the bank. Please give me your account number so I can deposit some money."

[0192] This will enable vulnerable people, such as the elderly, to make calls safely without falling victim to fraud.

[0193] The flow of the specific processing in the application example 1 will be described with reference to FIG.

[0194] Step 1:

[0195] When a user starts a call, the device automatically starts recording. The input is the call audio, and the output is real-time recorded audio data. Specifically, the device's microphone captures the audio and stores it in a buffer.

[0196] Step 2:

[0197] The device sends recorded audio data to the server in fixed buffer size increments. The input is the recorded audio data, and the output is the buffer data sent to the server. Specifically, the device sends the audio data stored in the buffer to the server in real time.

[0198] Step 3:

[0199] The server converts the received voice data into text using a generative AI model. The input is voice data, and the output is text data. Specifically, it uses "Google Cloud Speech-to-Text" to convert the voice data into text data.

[0200] Step 4:

[0201] The server analyzes the text data and detects fraud-related keywords. The input is text data, and the output is the results of fraud keyword detection. Specifically, it uses an algorithm to search for specific keywords (e.g., "transfer fraud" or "bank account") within the text data.

[0202] Step 5:

[0203] If the server determines that fraud is suspected, it activates a generative AI model to generate an appropriate automated response. The input is the fraud keyword detection results and the conversation text, and the output is the generated automated response text. Specific operations include using "OpenAI GPT-4" to generate a response based on the prompt text. Examples of prompt text include "If fraud is detected, please generate an appropriate response" and "This is from a bank. Please tell me your account number so I can deposit the money."

[0204] Step 6:

[0205] The device receives the automated response sent from the server and plays it back as audio. The input is the generated automated response text, and the output is the automated response in audio. Specifically, the AI ​​response is played back as audio using the device's speaker.

[0206] Step 7:

[0207] The server stores call records suspected of fraud and provides them to relevant authorities as necessary. The input is the call records and fraud detection results, and the output is the stored data and submitted data. Specifically, the server stores call records suspected of fraud in a database and provides them to relevant authorities such as the police.

[0208] This will create a system that can detect fraud in real time and respond appropriately.

[0209] Furthermore, an emotion engine that estimates the user's emotion may be combined. That is, the identification processing unit 290 may estimate the user's emotion using the emotion identification model 59 and perform identification processing using the user's emotion.

[0210] This invention combines an emotion engine with a system for preventing special fraud during phone calls. Below, each element of this system and its processing will be explained in natural language.

[0211] Start recording a call

[0212] 1. User: Start a call.

[0213] 2. Device: Automatically starts recording when a call is initiated. This means that audio data is recorded from the moment the user starts a call.

[0214] Sending audio data

[0215] 3. Terminal: Recorded audio data is sent to the server at fixed buffer size intervals.

[0216] Speech data conversion and sentiment analysis

[0217] 4. Server: Receives the voice data sent from the device, converts the voice to text using generative AI, and analyzes the user's emotions from the voice data using an emotion engine.

[0218] Analysis of text and emotion data

[0219] 5. Server: Analyzes the text of the conversation. Retrieves fraud-related keywords and similar words from the database, and uses an analysis algorithm to determine whether there is suspicion of fraud. Taking into account the emotional data analyzed by the emotion engine, the server strengthens detection of suspected fraud.

[0220] What to do if fraud is suspected

[0221] 6. Server: If fraud is suspected, the automated AI generates an appropriate response. The response is dynamically adjusted based on the user's emotional data obtained from the emotion engine.

[0222] 7. Device: Plays an automated response and responds on behalf of the user. During this time, the call continues to be recorded and the conversation data continues to be sent to the server.

[0223] Record keeping and provision

[0224] 8. Server: Conversation data suspected of fraud is stored in a database. The data will be provided to relevant agencies such as the police at a later date. The analysis results of the emotion engine will also be stored.

[0225] User Notification

[0226] 9. Server: Notify the user of suspected fraud and sentiment analysis results, if necessary.

[0227] Specific examples

[0228] Here is a specific scenario:

[0229] 1. User: An elderly user receives the call.

[0230] 2. Device: When a call is started, the device automatically starts recording. The recorded audio data is sent to the server in real time.

[0231] 3. Server: The received voice data is converted into text using a generative AI, and then the emotion engine analyzes the user's emotions. Fraud-related keywords such as "depositing money" and "bank account" are detected, and a suspicion of fraud is determined based on the analysis results and the user's emotional data.

[0232] 4. Server: If the server detects a suspected fraud, it will activate an automated AI response, generating a response such as "Please wait a moment, we're checking." The response is dynamically adjusted based on the results of sentiment analysis.

[0233] 5. Terminal: Based on instructions from the server, it plays an automated response and continues to record the conversation.

[0234] 6. Server: After the call ends, the server stores the conversation data and sentiment analysis results of any suspected fraud and periodically provides them to the police.

[0235] This series of steps protects users from specialized fraud and enables more advanced fraud detection and response through the emotion engine, providing an environment where users can use their phones with peace of mind.

[0236] The processing flow will be explained below.

[0237] Step 1:

[0238] User: Start a call.

[0239] Device: Automatically starts recording as soon as the call starts.

[0240] Step 2:

[0241] Terminal: Recorded audio data is sent to the server in real time at fixed buffer size intervals.

[0242] Step 3:

[0243] Server: Receives voice data sent from the device and converts it into text using a generation AI. The converted text data is temporarily stored.

[0244] Step 4:

[0245] Server: At the same time, the emotion engine is used to analyze the user's emotions from the voice data, and the analysis results are temporarily saved.

[0246] Step 5:

[0247] Server: The server compares the text of the conversation with a database of fraud-related keywords to determine whether there is any suspicion of fraud. It also considers the analysis results of the emotion engine to enhance detection of suspected fraud.

[0248] Step 6:

[0249] Server: If fraud is suspected, the automated AI responds by generating an appropriate response. The response is dynamically adjusted based on the user's emotional data obtained from the emotion engine.

[0250] Step 7:

[0251] Terminal: Plays back the automated response sent from the server and responds on behalf of the user. Records during the conversation and continues to send voice data to the server.

[0252] Step 8:

[0253] Server: After the call ends, the conversation data suspected of fraud and the results of sentiment analysis are stored in a database.

[0254] Step 9:

[0255] Server: Provides the relevant data to the police and other relevant authorities on a regular basis. If necessary, notifies the user of suspected fraud or the results of sentiment analysis.

[0256] This series of steps will protect users from specialized fraud, enable advanced fraud detection and response using an emotion engine, and provide recorded data to police to help prevent future crimes.

[0257] Example 2

[0258] Next, a description will be given of Example 2. In the following description, the data processing device 12 will be referred to as a "server" and the smart device 14 will be referred to as a "terminal."

[0259] Currently, special frauds over the phone are on the rise, and elderly people are particularly vulnerable to these crimes. Conventional countermeasures have limitations, and more effective systems are needed to detect suspected fraud in real time and protect users. It is also important to ensure that records of fraudulent conversations are stored and provided to the necessary authorities.

[0260] The specific processing by the specific processing unit 290 of the data processing device 12 in the second embodiment is realized by the following means.

[0261] In this invention, the server includes a means for converting the voice data of the call into text using a generation AI, a means for analyzing emotions from the converted text, and a means for detecting fraud-related keywords and determining suspicion of fraud based on the emotion analysis results. This makes it possible to analyze the content of the call in real time and strengthen suspicion of fraud based on the emotion data.

[0262] The "means for starting recording" is a device or program that has the function of automatically starting recording of voice data as soon as a call is started.

[0263] The "means for transmitting audio data" is a device or program for transmitting recorded audio data to a server in real time with a certain buffer size.

[0264] "Means for converting audio data into text using generative AI" refers to an artificial intelligence algorithm or software for automatically converting audio data into text data.

[0265] "Means for analyzing emotions" refers to software or algorithms that analyze the user's emotional state from textual data and obtain the results.

[0266] A "means for detecting fraud-related keywords" is a program or algorithm for extracting specific words or phrases related to fraudulent activities from text data.

[0267] "Means for determining suspicion of fraud by taking into account the results of sentiment analysis" refers to software or algorithms that have the function of comprehensively evaluating and determining the possibility of fraud based on the results of sentiment analysis data and keywords related to fraud.

[0268] "Means for automatic AI conversation takeover" refers to a program or device that allows artificial intelligence to automatically continue a conversation on behalf of a user when a suspected fraud is determined.

[0269] "Means for adjusting the content of the automated response based on the results of sentiment analysis" refers to software or algorithms for appropriately changing the content and tone of the automated response message based on the results of sentiment analysis.

[0270] The "means for playing the message aloud" refers to a device or software for converting the generated automated response message into aloud voice and playing it back to the other party on the call.

[0271] The "means of storing and providing records of suspected fraud" refers to a system for storing call records and analysis data of suspected fraud in a database and providing them to relevant authorities as necessary.

[0272] This invention relates to a system for preventing special fraud during phone calls, which operates by combining a generative AI and an emotion engine. This system is realized using the following specific hardware and software.

[0273] Hardware and Software Configuration

[0274] server

[0275] 1. Generative AI model (e.g., Google Cloud's Speech-to-Text API): Converts voice data into text data.

[0276] 2. Sentiment engine (e.g., Microsoft® Azure® Text Analytics API): Analyzes emotions from text data.

[0277] 3. Automated response AI (e.g., OpenAI’s GPT-3): Generates appropriate responses in cases of suspected fraud.

[0278] Terminal

[0279] 1. Calling app (e.g. Skype, Google Voice): An application for making calls.

[0280] 2. Recording function (e.g., iOS's Call Recorder app): A function that automatically records phone calls.

[0281] 3. Text-to-speech API (e.g., Amazon Polly): Speech the generated text response.

[0282] Overview of program processing

[0283] Start recording a call

[0284] When a user makes or receives a call using a calling app, the device automatically activates the recording function and records the audio data. This audio data is sent to the server in real time at fixed buffer size intervals.

[0285] Speech data conversion and sentiment analysis

[0286] The server uses a generative AI model to convert the voice data into text, which is then analyzed by an emotion engine to determine the user's emotional state. For example, emotions such as "happy," "sad," and "angry" are detected.

[0287] Determine suspected fraud

[0288] Based on the textual content of the conversation and the emotional data, the server detects keywords related to fraud. For example, it scans for words such as "money," "bank," and "account." At the same time, it also takes into account the results of the emotional analysis and makes a comprehensive judgment on the suspicion of fraud. For example, if the phrase "depositing money" matches the emotional data of "anxiety," it is determined that there is a high suspicion of fraud.

[0289] Auto-response generation and playback

[0290] If suspected fraud is detected, the server activates an automated AI to generate an appropriate response. This response is adjusted based on the results of sentiment analysis. The generated response is sent in text format to the device, which then uses a speech synthesis API to convert it into audio and play it back to the other party. For example, the response could be, "Please wait a moment, we're checking."

[0291] Record keeping and provision

[0292] Conversation data and sentiment analysis results for suspected fraud are stored in a secure database on a server. This data can be provided to relevant authorities as needed, such as the police, to enable further investigation and countermeasures.

[0293] User Notifications

[0294] If a suspected fraud is detected, the server will notify the user via email and / or SMS.

[0295] Examples of concrete examples and prompts

[0296] Specific examples

[0297] When an elderly user answers the phone and the call begins, the device automatically starts recording. This audio data is sent to a server in real time and converted into text by a generative AI model. If the phrase "depositing money" and the emotion "anxiety" are detected, the server activates an automated response AI and generates a response saying, "Please wait a moment, we're checking." This response is then converted into audio and played back to the caller.

[0298] Prompt Sentence Examples

[0299] The voice data is converted into text and analyzed, and if it contains keywords such as "money" or "bank," it is judged to be suspected of fraud. For example, if the phrase "depositing money" is included, it is considered highly suspected of fraud.

[0300] This effectively protects users from specialized fraud while on the phone, while the emotion engine enables more advanced fraud detection and response.

[0301] The flow of the identification process in the second embodiment will be described with reference to FIG.

[0302] Processing Steps

[0303] Step 1: Start recording the call

[0304] User: A user initiates or receives a call using a calling app.

[0305] Terminal: When a call is initiated, the recording function is automatically activated and audio data is recorded.

[0306] Input: Call Start Trigger

[0307] Output: Recorded audio data

[0308] Specific operation: When the device's calling app detects the start of a call, the recording function is activated and the call audio from that point on is saved as digital data.

[0309] Step 2: Sending audio data

[0310] Terminal: Recorded audio data is sent to the server in real time at fixed buffer size intervals.

[0311] Input: Recorded audio data

[0312] Output: Audio data sent to the server

[0313] What it does: The audio data is split into one-minute buffers and sent to the server using an HTTP POST request.

[0314] Step 3: Converting audio data to text

[0315] Server: Uses a generative AI model to convert voice data into text data.

[0316] Input: Audio data received from the device

[0317] Output: Text data

[0318] Specific operation: After receiving the voice data, the server calls Google Cloud's Speech-to-Text API to convert the voice to text.

[0319] Step 4: Sentiment Analysis

[0320] Server: Analyzes emotions from text data using an emotion engine.

[0321] Input: Text data

[0322] Output: Emotion analysis results

[0323] How it works: The text data is sent to Microsoft Azure's Text Analytics API, where it is analyzed to see if it contains emotions such as "happiness," "sadness," or "anger."

[0324] Step 5: Determine if there is any fraud

[0325] Server: Based on the text data and sentiment analysis results, detects fraud-related keywords and determines whether fraud is suspected.

[0326] Input: Text data, sentiment analysis results

[0327] Output: Suspected fraud determination result

[0328] How it works: The server scans the text data and compares it with a list of keywords (e.g., "money," "bank," "account"), while also taking into account the results of sentiment analysis to make a comprehensive judgment on the suspicion of fraud.

[0329] Step 6: Generate an autoresponder

[0330] Server: If fraud is suspected, an automated AI is used to generate an appropriate response.

[0331] Input: Suspected fraud judgment result

[0332] Output: Auto-response text

[0333] Specific operation: The server calls OpenAI's GPT-3 and generates a response message based on the situation, such as "Please wait a moment, checking."

[0334] Step 7: Play Auto Attendant Audio

[0335] Terminal: The text of the automated response is converted into speech using a speech synthesis API and played back to the other party.

[0336] Input: Auto-response text

[0337] Output: Voiced auto attendant

[0338] How it works: The device uses a speech synthesis API such as Amazon Polly to convert the generated text into speech and play it back to the other party in real time.

[0339] Step 8: Record keeping and provision

[0340] Server: Stores suspected fraudulent conversation data and sentiment analysis results in a database and provides them to relevant authorities as needed.

[0341] Input: Suspected fraud voice data, emotion analysis results

[0342] Output: Data stored, data provided

[0343] What it does: The server stores suspected fraudulent conversation data in a secure database and provides it to police and other relevant authorities as needed.

[0344] Step 9: User Notification

[0345] Server: Notifies the user if suspected fraud is detected.

[0346] Input: Suspected fraud judgment result

[0347] Output: User notification

[0348] What it does: The server notifies the user of suspected fraud via email address or mobile phone number. This notification is sent via email using an SMTP server or as a text message using an SMS API.

[0349] Through the above steps, the system can detect suspected special fraud in real time during a call and effectively protect users.

[0350] (Application example 2)

[0351] Next, a description will be given of Application Example 2. In the following description, the data processing device 12 will be referred to as a "server" and the smart device 14 will be referred to as a "terminal."

[0352] Conventional call recording systems could record phone conversations and, in some cases, detect keywords that could be suspected of fraud. However, they lacked the ability to analyze user sentiment and use that information to further improve the accuracy of fraud detection. As a result, even if fraud was suspected, it was not possible to evaluate whether the user actually perceived the risk, limiting the accuracy of fraud detection. Furthermore, they lacked the ability to provide appropriate automated responses, making it difficult to respond quickly in situations where there was a high risk of actually being scammed. The challenge is to solve these problems and provide a more advanced fraud prevention system.

[0353] The identification processing by the identification processing unit 290 of the data processing device 12 in Application Example 2 is realized by the following means. In this invention, the server includes means for transmitting voice data of the call, means for converting the received voice data into text using a generation AI, means for detecting fraud-related keywords, means for analyzing user emotions from the voice data using an emotion engine, means for dynamically adjusting automatic responses based on the emotion data, and means for saving and providing records of suspected fraud. This enables advanced fraud detection that takes into account user emotion analysis and the provision of appropriate and dynamically adjusted automatic responses.

[0354] A "call" is an act of exchanging voice information between two people using a telephone line.

[0355] "Recording" is the act of recording sound in digital or analog form.

[0356] "Generative AI" is a system that uses artificial intelligence technology to generate and convert voice and text.

[0357] "Text conversion" is the process of converting audio data into text data.

[0358] "Fraud Keywords" are words or phrases that are likely to be associated with fraudulent activity.

[0359] An "emotion engine" is a system that analyzes user emotions from voice and text data.

[0360] An "auto-response" is a message that the system automatically generates in response to a user.

[0361] "Dynamic adjustment" refers to changing the response in real time based on specific situations and conditions.

[0362] "Record keeping" is the act of storing data in a file or database.

[0363] "Notification" means the act of informing a user of specific information.

[0364] This invention is a system for preventing special fraud during phone calls, and aims to further improve accuracy by analyzing the user's emotions using an emotion engine.

[0365] The system is roughly configured as follows:

[0366] Hardware and Software Configuration

[0367] 1. Device:

[0368] A communication device such as a smartphone that has the ability to record voice data during calls and send it to a server.

[0369] 2. Server:

[0370] The server converts the received voice data into text using a generative AI model, detects fraud-related keywords, and analyzes the user's emotions using an emotion engine.

[0371] Program Overview

[0372] The system operates in the following steps:

[0373] Call Recording:

[0374] When a user starts a call, the device automatically starts recording, and the recorded audio data is sent to the server in real time.

[0375] Speech to text conversion and sentiment analysis:

[0376] The server converts the received voice data into text using a generative AI, and then uses an emotion engine to analyze emotions from the voice data. Specifically, the server uses the speech_recognition library and the transformers library.

[0377] Suspected fraud detection:

[0378] The server analyzes the converted text data to detect fraud-related keywords and, taking into account the results of the emotion engine analysis, determines whether or not there is any suspicion of fraud.

[0379] Generate an auto-response:

[0380] If fraud is suspected, the server will activate an automated AI response that generates a dynamically tailored response based on the results of sentiment analysis. This automated response is played back via voice and acts on behalf of the user.

[0381] Record-Keeping and Notification:

[0382] During and after the call, the server stores the records of suspected fraud and the results of the sentiment analysis. If necessary, the server periodically provides the data to the relevant authorities and notifies the user of the suspected fraud.

[0383] Specific examples

[0384] For example, if an elderly person receives a phone call and is told, "This is confirmation from the bank," the device automatically records the call and sends the audio data to a server. The server converts the audio data into text and analyzes the user's emotions using an emotion engine. If a fraud keyword such as "depositing money" is detected and the analysis finds that the elderly person is anxious, the server generates an automatic response such as "Are you okay? We're checking," and plays it back.

[0385] Prompt Sentence Examples

[0386] Prompt: Convert the speech data into text and detect whether it contains fraud-related keywords. Also, perform sentiment analysis and generate appropriate responses based on that.

[0387] In this way, it provides an advanced prevention system to prevent users from falling victim to fraud during calls.

[0388] The flow of the specific processing in the application example 2 will be described with reference to FIG.

[0389] Step 1:

[0390] Start recording a call

[0391] Input: The user receives a phone call.

[0392] How it works: When you start a call, your device automatically starts recording the call.

[0393] Output: The recorded audio data is generated.

[0394] Step 2:

[0395] Sending audio data

[0396] Input: Recorded audio data.

[0397] How it works: The device sends the recorded audio data to the server in real time.

[0398] Output: Audio data transferred to the server.

[0399] Step 3:

[0400] Converting audio data to text

[0401] Input: The audio data sent to the server.

[0402] How it works: The server uses a generative AI model (e.g., the speech_recognition library) to convert the audio data to text.

[0403] Output: Text data is generated.

[0404] Step 4:

[0405] Emotion analysis

[0406] Input: The text data generated in step 3.

[0407] How it works: The server uses an emotion engine (e.g., the sentiment-analysis model from the transformers library) to analyze user sentiment from text data.

[0408] Output: Emotion data is generated.

[0409] Step 5:

[0410] Fraudulent Keyword Detection

[0411] Input: Text data.

[0412] How it works: The server scans the text data using a pre-defined list of keywords (e.g., "deposit money," "bank account," etc.) to detect fraud-related keywords in the text data.

[0413] Output: The result whether fraud keywords were detected or not.

[0414] Step 6:

[0415] Determining suspected fraud

[0416] Input: Sentiment data and fraud keyword detection results.

[0417] Operation: The server makes a comprehensive judgment on whether there is suspicion of fraud based on the emotional data and the fraud keyword detection results.

[0418] Output: Judgment result on whether fraud is suspected.

[0419] Step 7:

[0420] Generate an auto-response

[0421] Input: Suspected fraud judgment results and sentiment data.

[0422] How it works: If fraud is suspected, the server uses a generative AI model to generate an automated response that dynamically adjusts based on sentiment data.

[0423] Output: A well-tuned auto-response.

[0424] Step 8:

[0425] Playing an Auto Attendant

[0426] Input: The generated auto-response.

[0427] How it works: The generated automated response is output as audio and played back to the user through the device.

[0428] Output: The audio response played to the user.

[0429] Step 9:

[0430] Record keeping and provision

[0431] Input: call content, fraud detection results, emotional data.

[0432] How it works: The server stores suspected fraudulent calls and their analysis results in a database, and provides them to relevant authorities as needed.

[0433] Output: Stored records and provided data.

[0434] Step 10:

[0435] User Notification

[0436] Input: Fraud determination result.

[0437] How it works: The server generates a warning message and sends it to the user via their device to inform them of the suspected fraud.

[0438] Output: The warning message sent to the user.

[0439] These steps will enable an advanced prevention system that analyzes call content in real time and protects users from the risk of special fraud.

[0440] The specific processing unit 290 transmits the result of the specific processing to the smart device 14. In the smart device 14, the control unit 46A causes the output device 40 to output the result of the specific processing. The microphone 38B acquires audio indicating a user input regarding the result of the specific processing. The control unit 46A transmits audio data indicating the user input acquired by the microphone 38B to the data processing device 12. In the data processing device 12, the specific processing unit 290 acquires the audio data.

[0441] The data generation model 58 is a so-called generative AI (Artificial Intelligence). An example of the data generation model 58 is ChatGPT (registered trademark) (Internet search engine).<URL: https: / / openai.com / blog / chatgpt> ), Gemini (registered trademark) (Internet search <url: https: gemini.google.com ?hl="ja">) and other generation AIs. The data generation model 58 is obtained by performing deep learning on a neural network. A prompt including an instruction is input to the data generation model 58, and inference data such as voice data indicating voice, text data indicating text, and image data indicating an image is also input. The data generation model 58 performs inference on the input inference data in accordance with the instruction indicated by the prompt, and outputs the inference result in a data format such as voice data and text data. Here, inference refers to, for example, analysis, classification, prediction, and / or summarization.

[0442] In the above embodiment, an example in which the specific process is performed by the data processing device 12 has been given, but the technology of the present disclosure is not limited to this, and the specific process may be performed by the smart device 14.

[0443] [Second embodiment]

[0444] FIG. 3 shows an example of the configuration of a data processing system 210 according to the second embodiment.

[0445] 3, the data processing system 210 includes the data processing device 12 and smart glasses 214. An example of the data processing device 12 is a server.

[0446] The data processing device 12 includes a computer 22, a database 24, and a communication I / F 26. The computer 22 is an example of a "computer" according to the technology of the present disclosure. The computer 22 includes a processor 28, a RAM 30, and a storage 32. The processor 28, the RAM 30, and the storage 32 are connected to a bus 34. The database 24 and the communication I / F 26 are also connected to the bus 34. The communication I / F 26 is connected to a network 54. Examples of the network 54 include a WAN (Wide Area Network) and / or a LAN (Local Area Network).

[0447] The smart glasses 214 include a computer 36, a microphone 238, a speaker 240, a camera 42, and a communication I / F 44. The computer 36 includes a processor 46, a RAM 48, and a storage 50. The processor 46, the RAM 48, and the storage 50 are connected to a bus 52. The microphone 238, the speaker 240, and the camera 42 are also connected to the bus 52.

[0448] The microphone 238 receives instructions and the like from the user 20 by receiving voice uttered by the user 20. The microphone 238 captures the voice uttered by the user 20, converts the captured voice into audio data, and outputs it to the processor 46. The speaker 240 outputs audio in accordance with instructions from the processor 46.

[0449] Camera 42 is a small digital camera equipped with an optical system including a lens, aperture, and shutter, and an imaging element such as a CMOS (Complementary Metal-Oxide-Semiconductor) image sensor or a CCD (Charge Coupled Device) image sensor, and captures images of the surroundings of user 20 (for example, an imaging range defined by an angle of view equivalent to the field of vision of a typical healthy person).

[0450] The communication I / F 44 is connected to a network 54. The communication I / Fs 44 and 26 are responsible for the exchange of various information between the processor 46 and the processor 28 via the network 54. The exchange of various information between the processor 46 and the processor 28 using the communication I / Fs 44 and 26 is carried out in a secure state.

[0451] Fig. 4 shows an example of the main functions of the data processing device 12 and the smart glasses 214. As shown in Fig. 4, in the data processing device 12, a specific process is performed by the processor 28. A specific process program 56 is stored in the storage 32.

[0452] The specific processing program 56 is an example of a "program" according to the technology of the present disclosure. The processor 28 reads the specific processing program 56 from the storage 32 and executes the read specific processing program 56 on the RAM 30. The specific processing is realized by the processor 28 operating as a specific processing unit 290 in accordance with the specific processing program 56 executed on the RAM 30.

[0453] The storage 32 stores a data generation model 58 and an emotion identification model 59. The data generation model 58 and the emotion identification model 59 are used by the identification processing unit 290.

[0454] In the smart glasses 214, the reception output process is performed by the processor 46. A reception output program 60 is stored in the storage 50. The processor 46 reads the reception output program 60 from the storage 50 and executes the read reception output program 60 on the RAM 48. The reception output process is realized by the processor 46 operating as the control unit 46A in accordance with the reception output program 60 executed on the RAM 48.

[0455] Next, a description will be given of the identification process performed by the identification processing unit 290 of the data processing device 12. In the following description, the data processing device 12 will be referred to as the "server" and the smart glasses 214 will be referred to as the "terminal."

[0456] This invention relates to a system for preventing special fraud during a call. Below, each element of this system and its processing will be explained in natural language.

[0457] Start recording a call

[0458] 1. Device:

[0459] It has a function that automatically starts recording when a call is started, so that audio data is recorded from the moment the user starts a call.

[0460] Sending audio data

[0461] 2. Terminal:

[0462] Recorded audio data is sent to the server in fixed buffer sizes, allowing the content of the call to be analyzed in real time.

[0463] Converting audio data to text

[0464] 3. Server:

[0465] The voice data sent from the device is received and converted into text in real time using a generation AI, which uses voice recognition technology to convert it into text information accurately and quickly.

[0466] Text data analysis

[0467] 4. Server:

[0468] The text of the conversation is analyzed, where fraud-related keywords and similar words are retrieved from a database, and an analysis algorithm is used to determine whether there is any suspicion of fraud.

[0469] What to do if fraud is suspected

[0470] 5. Server:

[0471] If fraud is suspected, the AI ​​automatically takes over the conversation, generating an appropriate response and sending it to the device.

[0472] 6. Terminal:

[0473] It responds on behalf of the user by playing back an automated response received from the server, which is designed to be unpleasant for scammers.

[0474] Record keeping and provision

[0475] 7. Server:

[0476] The data of calls suspected to be fraudulent will be saved and later provided to the police and other relevant authorities. This information will be used to help prevent fraud and to prevent future victims from becoming victims of fraud.

[0477] Specific examples

[0478] Here is a specific scenario:

[0479] 1. User:

[0480] Elderly user receives a call.

[0481] 2. Terminal:

[0482] When a call is initiated, the device automatically starts recording, and the recorded audio data is sent to the server in real time.

[0483] 3. Server:

[0484] The received voice data is converted into text using a generative AI, and keywords related to fraud, such as "depositing money" and "bank account," are analyzed.

[0485] If the server suspects fraud, it will activate an automated response AI that will generate a response such as "Please wait a moment, we're checking."

[0486] 4. Terminal:

[0487] Based on instructions from the server, an automated response is played and the conversation continues to be recorded.

[0488] 5. Server:

[0489] After the service is terminated, conversation data suspected of fraud will be saved and provided to the police on a regular basis.

[0490] This will help prevent victims of special fraud and allow all users, including the elderly, to use the telephone with peace of mind.

[0491] The processing flow will be explained below.

[0492] Step 1:

[0493] User: Start a call.

[0494] Device: Recording will start automatically when a call is initiated.

[0495] Step 2:

[0496] Terminal: Recorded audio data is sent to the server at fixed buffer size intervals.

[0497] Step 3:

[0498] Server: Receives voice data sent from the device and converts the voice into text using generative AI.

[0499] Step 4:

[0500] Server: Analyzes the text of the conversation, retrieves fraud-related keywords and similar words from a database, and determines whether there is any suspicion of fraud.

[0501] Step 5:

[0502] Server: If fraud is suspected, an automated response AI is activated to generate an appropriate response.

[0503] Terminal: Plays back the generated response aloud and acts on behalf of the user.

[0504] Step 6:

[0505] Device: Continue recording and continue sending conversation data to the server.

[0506] Step 7:

[0507] Server: Stores conversation data suspected of fraud in a database. This data will be provided to relevant agencies such as the police at a later date.

[0508] Step 8:

[0509] Server: If necessary, notify the user of suspected fraud.

[0510] This series of steps will protect users from special fraud by monitoring call content in real time and automatically responding if fraud is suspected.

[0511] Example 1

[0512] Next, a description will be given of Example 1. In the following description, the data processing device 12 will be referred to as a "server" and the smart glasses 214 will be referred to as a "terminal."

[0513] Special frauds are becoming more sophisticated every day, and many people, including the elderly, are falling victim to them. Effective measures to prevent these frauds are needed, but existing technology makes it difficult to detect fraud during a call in real time and take appropriate action. Furthermore, there is no system that can instantly recognize and respond to suspected fraudulent calls. Therefore, there is a need to provide a system that prevents fraud during a call and create an environment in which users can use the telephone with peace of mind.

[0514] The specific processing by the specific processing unit 290 of the data processing device 12 in the first embodiment is realized by the following means.

[0515] In this invention, the server includes a means for transmitting the voice data of the call at fixed buffer size intervals, a means for converting the received voice data into text in real time using a generation AI, a means for retrieving and detecting fraud-related keywords from a database, a means for the generation AI to automatically take over the conversation and generate and send an appropriate response if fraud is suspected, and a means for saving records of suspected fraud and providing them to relevant authorities. This makes it possible to detect special frauds in real time during a call and automatically provide an appropriate response, significantly reducing the risk of users becoming victims of fraud.

[0516] The "means for automatically starting recording when a call is started" refers to a device or software that automatically starts a recording function the moment the user starts a call and records the audio data of the call.

[0517] The "means for transmitting voice data of a call at fixed buffer size intervals" refers to a device or software for transmitting collected voice data to a server at fixed volume or time intervals.

[0518] "Means for converting received voice data into text in real time using a generation AI" refers to a device or software in which a server receives voice data sent from a terminal and converts the voice into text in real time using a generation AI.

[0519] "Means for retrieving and detecting fraud-related keywords from a database" refers to a device or software for searching a database for keywords that indicate signs of fraud contained in text-converted voice data and detecting them.

[0520] "Means for a generation AI to automatically take over a conversation when fraud is suspected, and generate and send an appropriate response" refers to a device or software that, when it is determined that fraud is suspected, automatically generates an appropriate countermeasure response from a generation AI and sends it on behalf of the user.

[0521] "Means for storing records of suspected fraud and providing them to relevant authorities" refers to devices or software that record the content of suspected fraudulent calls and the results of their analysis, and provide them to relevant authorities at a later date.

[0522] "Means for generating lines for automated responses using a generation AI and playing them audibly" refers to a device or software that converts the text of an automated response generated by a generation AI into audio and plays it back through a speaker or the like.

[0523] "Means for notifying the user when suspected fraud is detected" refers to a device or software that notifies the user of the information when the system determines that there is a suspicion of fraud.

[0524] This invention relates to a system for preventing special fraud during a call. Each element of this system and its processing will be specifically described below.

[0525] System configuration and processing

[0526] Call recording

[0527] When a user picks up the receiver or presses the call button to start a call, the device's recording function automatically starts and records the audio data of the call. Once the recorded audio data reaches a certain buffer size, it is sent from the device to the server. This is usually done using an HTTP POST request.

[0528] Converting audio data to text

[0529] The server receives the voice data sent from the device and converts it into text in real time using a generation AI (e.g., DeepSpeech, Google Cloud Speech-to-Text, etc.) This generation AI inputs the voice data and outputs the corresponding text data.

[0530] Text data analysis

[0531] The text data is first stored in a database on the server, and then keywords related to fraud are retrieved from the database and detected. This analysis is performed using regular expressions and natural language processing tools (e.g., NLTK). Based on the keyword detection results, a judgment is made as to whether there is suspicion of fraud, and if so, the system proceeds to the next step.

[0532] Generate and send automatic responses

[0533] If fraud is suspected, the server uses a generative AI model (e.g., GPT-3) to generate an appropriate automated response. A specific example of the generative AI's prompt is, "This call is suspected to be fraudulent. Please play the next message." The generated automated response is sent to the device, which plays it aloud. This allows the server to respond appropriately to the fraudster on behalf of the user.

[0534] Record keeping and provision

[0535] Records of calls suspected of fraud are stored on a server. This information will be provided to relevant authorities, such as the police, at a later date. The stored data includes voice data, text data, detected keywords, and responses generated by the AI.

[0536] Specific examples

[0537] The following is a specific example:

[0538] The user is an elderly person and receives a phone call. The device automatically records the call and sends the recorded voice data to the server in real time. The server converts the voice data into text using a generative AI and detects fraud-related keywords such as "depositing money" and "bank account." If the server determines that there is suspicion of fraud, it uses the generative AI to generate an automated response such as "Please wait a moment, we're checking," and sends it to the device. The device then plays this automated response aloud to ensure the user's safety.

[0539] This system will prevent victims of special fraud before they occur, and will enable all users, including the elderly, to use the telephone with peace of mind.

[0540] Prompt Sentence Examples

[0541] 1. "This call is suspected to be fraudulent. Please play the following message."

[0542] 2. "Do not provide bank account information."

[0543] 3. "Always check with family and friends before sending money."

[0544] Using such prompts allows the generative AI to provide appropriate warnings and responses to scammers.

[0545] The flow of the identification process in the first embodiment will be described with reference to FIG.

[0546] Step 1:

[0547] The user answers or makes a call. The input is the action of lifting the telephone receiver or pressing the talk button, and the output is the start of the call signal, which starts recording the call.

[0548] Step 2:

[0549] The device detects that a call has started and automatically starts recording. The input is a call start signal, and the output is recording of audio data. The device's internal recording function is activated, and audio data is collected through the microphone.

[0550] Step 3:

[0551] The device collects recorded audio data in fixed buffer sizes and sends them to the server in bulk. The input is the audio data being recorded, and the output is the audio data buffer. When the buffer reaches a certain size, it is sent to the server via an HTTP POST request.

[0552] Step 4:

[0553] The server receives the audio data sent from the device. The input is the audio data sent via an HTTP POST request, and the output is the audio data stored on the server. The server temporarily stores this audio data and proceeds to the next step.

[0554] Step 5:

[0555] The server converts the received voice data into text in real time using a generation AI (e.g., DeepSpeech, Google Cloud Speech-to-Text). The input is voice data, and the output is text data. The generation AI analyzes the voice data and generates the corresponding text.

[0556] Step 6:

[0557] The server analyzes the data converted to text by the generative AI and detects fraud-related keywords. The input is text data, and the output is the keyword detection results. A natural language processing tool (e.g., NLTK) is used to search for fraud-related keywords.

[0558] Step 7:

[0559] The server determines whether fraud is suspected based on the keyword detection results. The input is the keyword detection results, and the output is a flag indicating whether fraud is suspected or not. A judgment algorithm (e.g., decision tree, SVM) is executed.

[0560] Step 8:

[0561] If the server determines that fraud is suspected, it uses a generative AI model (e.g., GPT-3) to generate an appropriate automated response. The input is a flag indicating suspected fraud, and the output is the automated response. The generative AI inputs a prompt sentence and generates a corresponding response.

[0562] Step 9:

[0563] The server sends the generated automatic response to the terminal. The input is the automatic response text, and the output is the data to be sent to the terminal. An HTTP POST request is used for sending.

[0564] Step 10:

[0565] The terminal plays back the received automated response lines aloud. The input is the data sent from the server, and the output is the audio data. The TTS (Text-to-Speech) function is used to convert the text data into audio, which is then played back through the speaker.

[0566] Step 11:

[0567] The server stores records of suspected fraudulent phone calls and periodically provides them to relevant authorities (e.g., police). The input is the call's audio data, text data, and analysis results, and the output is the stored recorded data and data provided to relevant authorities. The records are stored in a database and provided to relevant authorities via an API or interface.

[0568] (Application example 1)

[0569] Next, a description will be given of Application Example 1. In the following description, the data processing device 12 will be referred to as a "server" and the smart glasses 214 will be referred to as a "terminal."

[0570] In modern society, the number of fraud cases is increasing, and vulnerable people such as the elderly are often the targets. To solve this problem, a system that can identify fraudulent activities in real time and take prompt countermeasures is needed. However, current systems have difficulty monitoring the content of phone calls one by one, making it difficult to prevent fraudulent activities. Therefore, there is a need to provide a system that can detect fraudulent activities early and respond appropriately.

[0571] The specific processing by the specific processing unit 290 of the data processing device 12 in the application example 1 is realized by the following means.

[0572] In this invention, the server includes means for automatically starting recording when a call is initiated, means for transmitting the voice data of the call, means for converting the received voice data into text using a generating AI, means for detecting fraud-related keywords, means for the AI ​​to automatically take over the conversation if fraud is suspected, means for recording data in real time and providing it to relevant authorities, and means for saving and providing records of suspected fraud. This makes it possible to detect special frauds in real time and take appropriate action.

[0573] "Means for automatically starting recording when a call starts" refers to a function that automatically starts recording audio data as soon as a call starts.

[0574] "Means for transmitting voice data of a call" refers to a function for transmitting recorded voice data to a server in real time or at fixed buffer size intervals.

[0575] "Means of converting received voice data into text using generation AI" refers to the function of converting voice data received by the server into text data using generation AI technology.

[0576] "Means for detecting fraud-related keywords" refers to the function of searching for and detecting specific fraud-related keywords from text data.

[0577] "Means for AI to automatically take over the conversation if fraud is suspected" refers to a function in which AI automatically generates and responds to conversations when it is determined that fraud is suspected.

[0578] "Means of recording data in real time and providing it to relevant authorities" refers to the function of recording the contents of phone calls in real time and providing it to relevant authorities such as the police if necessary.

[0579] "Means for storing and providing records of suspected fraud" refers to the function of storing records of phone calls suspected of fraud and providing them to relevant authorities at a later date if necessary.

[0580] "Means for generating lines for automatic responses using generative AI and playing them back as audio" refers to the function of creating the content of automatic responses using generative AI technology and playing them back as audio.

[0581] "Means of notifying users when suspected fraud is detected" refers to a function that warns or notifies users when suspected fraud is detected.

[0582] This invention relates to a system for preventing special fraud during a call. Each element of this system and its processing will be explained below.

[0583] Terminal behavior:

[0584] The device has a means to automatically start recording when a call is initiated, and the recorded audio data is sent to the server in fixed buffer sizes, allowing the content of the call to be analyzed in real time.

[0585] Server behavior:

[0586] The server converts the received voice data into text using a generative AI model, "Google Cloud Speech-to-Text," which allows for accurate and fast conversion of voice to text.

[0587] The text data is then analyzed to detect fraud-related keywords, such as "bank transfer fraud," "depositing money," and "bank account," and an analysis algorithm is used to detect these.

[0588] If fraud is suspected, the server automatically activates AI to generate an appropriate response using a generative AI model such as OpenAI GPT-4. The generated response is sent to the device and played back as audio.

[0589] Data storage and provision:

[0590] The server stores records of calls suspected of fraud and has the means to provide them to relevant authorities if necessary, so that police and other relevant authorities can obtain the necessary information at a later date.

[0591] Examples:

[0592] Here's a specific scenario:

[0593] 1. User:

[0594] An elderly user receives a suspicious phone call.

[0595] 2. Terminal:

[0596] Recording will start automatically as soon as the call is started.

[0597] The recorded audio data is sent to the server in real time.

[0598] 3. Server:

[0599] The received voice data is converted into text using the generative AI model "Google Cloud Speech-to-Text."

[0600] From the text data, keywords related to fraud, such as "depositing money" and "bank account," are detected.

[0601] If it determines that fraud is suspected, an automatic response message is generated using OpenAI GPT-4.

[0602] For example, generate a response like "Please wait, we're checking."

[0603] 4. Terminal:

[0604] The automatic response message sent from the server is played back and the conversation is taken over.

[0605] 5. Server:

[0606] Records of calls suspected of fraud will be saved and provided to the police and other relevant authorities as necessary.

[0607] Example prompt sentence:

[0608] System: "If fraud is detected, generate an appropriate response."

[0609] User: "I'm from the bank. Please give me your account number so I can deposit some money."

[0610] This will enable vulnerable people, such as the elderly, to make calls safely without falling victim to fraud.

[0611] The flow of the specific processing in the application example 1 will be described with reference to FIG.

[0612] Step 1:

[0613] When a user starts a call, the device automatically starts recording. The input is the call audio, and the output is real-time recorded audio data. Specifically, the device's microphone captures the audio and stores it in a buffer.

[0614] Step 2:

[0615] The device sends recorded audio data to the server in fixed buffer size increments. The input is the recorded audio data, and the output is the buffer data sent to the server. Specifically, the device sends the audio data stored in the buffer to the server in real time.

[0616] Step 3:

[0617] The server converts the received voice data into text using a generative AI model. The input is voice data, and the output is text data. Specifically, it uses "Google Cloud Speech-to-Text" to convert the voice data into text data.

[0618] Step 4:

[0619] The server analyzes the text data and detects fraud-related keywords. The input is text data, and the output is the results of fraud keyword detection. Specifically, it uses an algorithm to search for specific keywords (e.g., "transfer fraud" or "bank account") within the text data.

[0620] Step 5:

[0621] If the server determines that fraud is suspected, it activates a generative AI model to generate an appropriate automated response. The input is the fraud keyword detection results and the conversation text, and the output is the generated automated response text. Specific operations include using "OpenAI GPT-4" to generate a response based on the prompt text. Examples of prompt text include "If fraud is detected, please generate an appropriate response" and "This is from a bank. Please tell me your account number so I can deposit the money."

[0622] Step 6:

[0623] The device receives the automated response sent from the server and plays it back as audio. The input is the generated automated response text, and the output is the automated response in audio. Specifically, the AI ​​response is played back as audio using the device's speaker.

[0624] Step 7:

[0625] The server stores call records suspected of fraud and provides them to relevant authorities as necessary. The input is the call records and fraud detection results, and the output is the stored data and submitted data. Specifically, the server stores call records suspected of fraud in a database and provides them to relevant authorities such as the police.

[0626] This will create a system that can detect fraud in real time and respond appropriately.

[0627] Furthermore, an emotion engine that estimates the user's emotion may be further combined. That is, the identification processing unit 290 may estimate the user's emotion using the emotion identification model 59, and perform identification processing using the user's emotion.

[0628] This invention combines an emotion engine with a system for preventing special fraud during phone calls. Below, each element of this system and its processing will be explained in natural language.

[0629] Start recording a call

[0630] 1. User: Start a call.

[0631] 2. Device: Automatically starts recording when a call is initiated. This means that audio data is recorded from the moment the user starts a call.

[0632] Sending audio data

[0633] 3. Terminal: Recorded audio data is sent to the server at fixed buffer size intervals.

[0634] Speech data conversion and sentiment analysis

[0635] 4. Server: Receives the voice data sent from the device, converts the voice to text using generative AI, and analyzes the user's emotions from the voice data using an emotion engine.

[0636] Analysis of text and emotion data

[0637] 5. Server: Analyzes the text of the conversation. Retrieves fraud-related keywords and similar words from the database, and uses an analysis algorithm to determine whether there is suspicion of fraud. Taking into account the emotional data analyzed by the emotion engine, the server strengthens detection of suspected fraud.

[0638] What to do if fraud is suspected

[0639] 6. Server: If fraud is suspected, the automated AI generates an appropriate response. The response is dynamically adjusted based on the user's emotional data obtained from the emotion engine.

[0640] 7. Device: Plays an automated response and responds on behalf of the user. During this time, the call continues to be recorded and the conversation data continues to be sent to the server.

[0641] Record keeping and provision

[0642] 8. Server: Conversation data suspected of fraud is stored in a database. The data will be provided to relevant agencies such as the police at a later date. The analysis results of the emotion engine will also be stored.

[0643] User Notification

[0644] 9. Server: Notify the user of suspected fraud and sentiment analysis results, if necessary.

[0645] Specific examples

[0646] Here is a specific scenario:

[0647] 1. User: An elderly user receives the call.

[0648] 2. Device: When a call is started, the device automatically starts recording. The recorded audio data is sent to the server in real time.

[0649] 3. Server: The received voice data is converted into text using a generative AI, and then the emotion engine analyzes the user's emotions. Fraud-related keywords such as "depositing money" and "bank account" are detected, and a suspicion of fraud is determined based on the analysis results and the user's emotional data.

[0650] 4. Server: If the server detects a suspected fraud, it will activate an automated AI response, generating a response such as "Please wait a moment, we're checking." The response is dynamically adjusted based on the results of sentiment analysis.

[0651] 5. Terminal: Based on instructions from the server, it plays an automated response and continues to record the conversation.

[0652] 6. Server: After the call ends, the server stores the conversation data and sentiment analysis results of any suspected fraud and periodically provides them to the police.

[0653] This series of steps protects users from specialized fraud and enables more advanced fraud detection and response through the emotion engine, providing an environment where users can use their phones with peace of mind.

[0654] The processing flow will be explained below.

[0655] Step 1:

[0656] User: Start a call.

[0657] Device: Automatically starts recording as soon as the call starts.

[0658] Step 2:

[0659] Terminal: Recorded audio data is sent to the server in real time at fixed buffer size intervals.

[0660] Step 3:

[0661] Server: Receives voice data sent from the device and converts it into text using a generation AI. The converted text data is temporarily stored.

[0662] Step 4:

[0663] Server: At the same time, the emotion engine is used to analyze the user's emotions from the voice data, and the analysis results are temporarily saved.

[0664] Step 5:

[0665] Server: The server compares the text of the conversation with a database of fraud-related keywords to determine whether there is any suspicion of fraud. It also considers the analysis results of the emotion engine to enhance detection of suspected fraud.

[0666] Step 6:

[0667] Server: If fraud is suspected, the automated AI responds by generating an appropriate response. The response is dynamically adjusted based on the user's emotional data obtained from the emotion engine.

[0668] Step 7:

[0669] Terminal: Plays back the automated response sent from the server and responds on behalf of the user. Records during the conversation and continues to send voice data to the server.

[0670] Step 8:

[0671] Server: After the call ends, the conversation data suspected of fraud and the results of sentiment analysis are stored in a database.

[0672] Step 9:

[0673] Server: Provides the relevant data to the police and other relevant authorities on a regular basis. If necessary, notifies the user of suspected fraud or the results of sentiment analysis.

[0674] This series of steps will protect users from specialized fraud, enable advanced fraud detection and response using an emotion engine, and provide recorded data to police to help prevent future crimes.

[0675] Example 2

[0676] Next, a description will be given of Example 2. In the following description, the data processing device 12 will be referred to as a "server" and the smart glasses 214 will be referred to as a "terminal."

[0677] Currently, special frauds over the phone are on the rise, and elderly people are particularly vulnerable to these crimes. Conventional countermeasures have limitations, and more effective systems are needed to detect suspected fraud in real time and protect users. It is also important to ensure that records of fraudulent conversations are stored and provided to the necessary authorities.

[0678] The specific processing by the specific processing unit 290 of the data processing device 12 in the second embodiment is realized by the following means.

[0679] In this invention, the server includes a means for converting the voice data of the call into text using a generation AI, a means for analyzing emotions from the converted text, and a means for detecting fraud-related keywords and determining suspicion of fraud based on the emotion analysis results. This makes it possible to analyze the content of the call in real time and strengthen suspicion of fraud based on the emotion data.

[0680] The "means for starting recording" is a device or program that has the function of automatically starting recording of voice data as soon as a call is started.

[0681] The "means for transmitting audio data" is a device or program for transmitting recorded audio data to a server in real time with a certain buffer size.

[0682] "Means for converting audio data into text using generative AI" refers to an artificial intelligence algorithm or software for automatically converting audio data into text data.

[0683] "Means for analyzing emotions" refers to software or algorithms that analyze the user's emotional state from textual data and obtain the results.

[0684] A "means for detecting fraud-related keywords" is a program or algorithm for extracting specific words or phrases related to fraudulent activities from text data.

[0685] "Means for determining suspicion of fraud by taking into account the results of sentiment analysis" refers to software or algorithms that have the function of comprehensively evaluating and determining the possibility of fraud based on the results of sentiment analysis data and keywords related to fraud.

[0686] "Means for automatic AI conversation takeover" refers to a program or device that allows artificial intelligence to automatically continue a conversation on behalf of a user when a suspected fraud is determined.

[0687] "Means for adjusting the content of the automated response based on the results of sentiment analysis" refers to software or algorithms for appropriately changing the content and tone of the automated response message based on the results of sentiment analysis.

[0688] The "means for playing the message aloud" refers to a device or software for converting the generated automated response message into aloud voice and playing it back to the other party on the call.

[0689] The "means of storing and providing records of suspected fraud" refers to a system for storing call records and analysis data of suspected fraud in a database and providing them to relevant authorities as necessary.

[0690] This invention relates to a system for preventing special fraud during phone calls, which operates by combining a generative AI and an emotion engine. This system is realized using the following specific hardware and software.

[0691] Hardware and Software Configuration

[0692] server

[0693] 1. Generative AI model (e.g., Google Cloud's Speech-to-Text API): Converts voice data into text data.

[0694] 2. Sentiment engine (e.g., Microsoft Azure's Text Analytics API): Analyzes emotions from text data.

[0695] 3. Automated response AI (e.g., OpenAI’s GPT-3): Generates appropriate responses in cases of suspected fraud.

[0696] Terminal

[0697] 1. Calling app (e.g. Skype, Google Voice): An application for making calls.

[0698] 2. Recording function (e.g., iOS's Call Recorder app): A function that automatically records phone calls.

[0699] 3. Text-to-speech API (e.g., Amazon Polly): Speech the generated text response.

[0700] Overview of program processing

[0701] Start recording a call

[0702] When a user makes or receives a call using a calling app, the device automatically activates the recording function and records the audio data. This audio data is sent to the server in real time at fixed buffer size intervals.

[0703] Speech data conversion and sentiment analysis

[0704] The server uses a generative AI model to convert the voice data into text, which is then analyzed by an emotion engine to determine the user's emotional state. For example, emotions such as "happy," "sad," and "angry" are detected.

[0705] Determine suspected fraud

[0706] Based on the textual content of the conversation and the emotional data, the server detects keywords related to fraud. For example, it scans for words such as "money," "bank," and "account." At the same time, it also takes into account the results of the emotional analysis and makes a comprehensive judgment on the suspicion of fraud. For example, if the phrase "depositing money" matches the emotional data of "anxiety," it is determined that there is a high suspicion of fraud.

[0707] Auto-response generation and playback

[0708] If suspected fraud is detected, the server activates an automated AI to generate an appropriate response. This response is adjusted based on the results of sentiment analysis. The generated response is sent in text format to the device, which then uses a speech synthesis API to convert it into audio and play it back to the other party. For example, the response could be, "Please wait a moment, we're checking."

[0709] Record keeping and provision

[0710] Conversation data and sentiment analysis results for suspected fraud are stored in a secure database on a server. This data can be provided to relevant authorities as needed, such as the police, to enable further investigation and countermeasures.

[0711] User Notifications

[0712] If a suspected fraud is detected, the server will notify the user via email and / or SMS.

[0713] Examples of concrete examples and prompts

[0714] Specific examples

[0715] When an elderly user answers the phone and the call begins, the device automatically starts recording. This audio data is sent to a server in real time and converted into text by a generative AI model. If the phrase "depositing money" and the emotion "anxiety" are detected, the server activates an automated response AI and generates a response saying, "Please wait a moment, we're checking." This response is then converted into audio and played back to the caller.

[0716] Prompt Sentence Examples

[0717] The voice data is converted into text and analyzed, and if it contains keywords such as "money" or "bank," it is judged to be suspected of fraud. For example, if the phrase "depositing money" is included, it is considered highly suspected of fraud.

[0718] This effectively protects users from specialized fraud while on the phone, while the emotion engine enables more advanced fraud detection and response.

[0719] The flow of the identification process in the second embodiment will be described with reference to FIG.

[0720] Processing Steps

[0721] Step 1: Start recording the call

[0722] User: A user initiates or receives a call using a calling app.

[0723] Terminal: When a call is initiated, the recording function is automatically activated and audio data is recorded.

[0724] Input: Call Start Trigger

[0725] Output: Recorded audio data

[0726] Specific operation: When the device's calling app detects the start of a call, the recording function is activated and the call audio from that point on is saved as digital data.

[0727] Step 2: Sending audio data

[0728] Terminal: Recorded audio data is sent to the server in real time at fixed buffer size intervals.

[0729] Input: Recorded audio data

[0730] Output: Audio data sent to the server

[0731] What it does: The audio data is split into one-minute buffers and sent to the server using an HTTP POST request.

[0732] Step 3: Converting audio data to text

[0733] Server: Uses a generative AI model to convert voice data into text data.

[0734] Input: Audio data received from the device

[0735] Output: Text data

[0736] Specific operation: After receiving the voice data, the server calls Google Cloud's Speech-to-Text API to convert the voice to text.

[0737] Step 4: Sentiment Analysis

[0738] Server: Analyzes emotions from text data using an emotion engine.

[0739] Input: Text data

[0740] Output: Emotion analysis results

[0741] How it works: The text data is sent to Microsoft Azure's Text Analytics API, where it is analyzed to see if it contains emotions such as "happiness," "sadness," or "anger."

[0742] Step 5: Determine if there is any fraud

[0743] Server: Based on the text data and sentiment analysis results, detects fraud-related keywords and determines whether fraud is suspected.

[0744] Input: Text data, sentiment analysis results

[0745] Output: Suspected fraud determination result

[0746] How it works: The server scans the text data and compares it with a list of keywords (e.g., "money," "bank," "account"), while also taking into account the results of sentiment analysis to make a comprehensive judgment on the suspicion of fraud.

[0747] Step 6: Generate an autoresponder

[0748] Server: If fraud is suspected, an automated AI is used to generate an appropriate response.

[0749] Input: Suspected fraud judgment result

[0750] Output: Auto-response text

[0751] Specific operation: The server calls OpenAI's GPT-3 and generates a response message based on the situation, such as "Please wait a moment, checking."

[0752] Step 7: Play Auto Attendant Audio

[0753] Terminal: The text of the automated response is converted into speech using a speech synthesis API and played back to the other party.

[0754] Input: Auto-response text

[0755] Output: Voiced auto attendant

[0756] How it works: The device uses a speech synthesis API such as Amazon Polly to convert the generated text into speech and play it back to the other party in real time.

[0757] Step 8: Record keeping and provision

[0758] Server: Stores suspected fraudulent conversation data and sentiment analysis results in a database and provides them to relevant authorities as needed.

[0759] Input: Suspected fraud voice data, emotion analysis results

[0760] Output: Data stored, data provided

[0761] What it does: The server stores suspected fraudulent conversation data in a secure database and provides it to police and other relevant authorities as needed.

[0762] Step 9: User Notification

[0763] Server: Notifies the user if suspected fraud is detected.

[0764] Input: Suspected fraud judgment result

[0765] Output: User notification

[0766] What it does: The server notifies the user of suspected fraud via email address or mobile phone number. This notification is sent via email using an SMTP server or as a text message using an SMS API.

[0767] Through the above steps, the system can detect suspected special fraud in real time during a call and effectively protect users.

[0768] (Application example 2)

[0769] Next, a description will be given of Application Example 2. In the following description, the data processing device 12 will be referred to as a "server" and the smart glasses 214 will be referred to as a "terminal."

[0770] Conventional call recording systems could record phone conversations and, in some cases, detect keywords that could be suspected of fraud. However, they lacked the ability to analyze user sentiment and use that information to further improve the accuracy of fraud detection. As a result, even if fraud was suspected, it was not possible to evaluate whether the user actually perceived the risk, limiting the accuracy of fraud detection. Furthermore, they lacked the ability to provide appropriate automated responses, making it difficult to respond quickly in situations where there was a high risk of actually being scammed. The challenge is to solve these problems and provide a more advanced fraud prevention system.

[0771] The identification processing by the identification processing unit 290 of the data processing device 12 in Application Example 2 is realized by the following means. In this invention, the server includes means for transmitting voice data of the call, means for converting the received voice data into text using a generation AI, means for detecting fraud-related keywords, means for analyzing user emotions from the voice data using an emotion engine, means for dynamically adjusting automatic responses based on the emotion data, and means for saving and providing records of suspected fraud. This enables advanced fraud detection that takes into account user emotion analysis and the provision of appropriate and dynamically adjusted automatic responses.

[0772] A "call" is an act of exchanging voice information between two people using a telephone line.

[0773] "Recording" is the act of recording sound in digital or analog form.

[0774] "Generative AI" is a system that uses artificial intelligence technology to generate and convert voice and text.

[0775] "Text conversion" is the process of converting audio data into text data.

[0776] "Fraud Keywords" are words or phrases that are likely to be associated with fraudulent activity.

[0777] An "emotion engine" is a system that analyzes user emotions from voice and text data.

[0778] An "auto-response" is a message that the system automatically generates in response to a user.

[0779] "Dynamic adjustment" refers to changing the response in real time based on specific situations and conditions.

[0780] "Record keeping" is the act of storing data in a file or database.

[0781] "Notification" means the act of informing a user of specific information.

[0782] This invention is a system for preventing special fraud during phone calls, and aims to further improve accuracy by analyzing the user's emotions using an emotion engine.

[0783] The system is roughly configured as follows:

[0784] Hardware and Software Configuration

[0785] 1. Device:

[0786] A communication device such as a smartphone that has the ability to record voice data during calls and send it to a server.

[0787] 2. Server:

[0788] The server converts the received voice data into text using a generative AI model, detects fraud-related keywords, and analyzes the user's emotions using an emotion engine.

[0789] Program Overview

[0790] The system operates in the following steps:

[0791] Call Recording:

[0792] When a user starts a call, the device automatically starts recording, and the recorded audio data is sent to the server in real time.

[0793] Speech to text conversion and sentiment analysis:

[0794] The server converts the received voice data into text using a generative AI, and then uses an emotion engine to analyze emotions from the voice data. Specifically, the server uses the speech_recognition library and the transformers library.

[0795] Suspected fraud detection:

[0796] The server analyzes the converted text data to detect fraud-related keywords and, taking into account the results of the emotion engine analysis, determines whether or not there is any suspicion of fraud.

[0797] Generate an auto-response:

[0798] If fraud is suspected, the server will activate an automated AI response that generates a dynamically tailored response based on the results of sentiment analysis. This automated response is played back via voice and acts on behalf of the user.

[0799] Record-Keeping and Notification:

[0800] During and after the call, the server stores the records of suspected fraud and the results of the sentiment analysis. If necessary, the server periodically provides the data to the relevant authorities and notifies the user of the suspected fraud.

[0801] Specific examples

[0802] For example, if an elderly person receives a phone call and is told, "This is confirmation from the bank," the device automatically records the call and sends the audio data to a server. The server converts the audio data into text and analyzes the user's emotions using an emotion engine. If a fraud keyword such as "depositing money" is detected and the analysis finds that the elderly person is anxious, the server generates an automatic response such as "Are you okay? We're checking," and plays it back.

[0803] Prompt Sentence Examples

[0804] Prompt: Convert the speech data into text and detect whether it contains fraud-related keywords. Also, perform sentiment analysis and generate appropriate responses based on that.

[0805] In this way, it provides an advanced prevention system to prevent users from falling victim to fraud during calls.

[0806] The flow of the specific processing in the application example 2 will be described with reference to FIG.

[0807] Step 1:

[0808] Start recording a call

[0809] Input: The user receives a phone call.

[0810] How it works: When you start a call, your device automatically starts recording the call.

[0811] Output: The recorded audio data is generated.

[0812] Step 2:

[0813] Sending audio data

[0814] Input: Recorded audio data.

[0815] How it works: The device sends the recorded audio data to the server in real time.

[0816] Output: Audio data transferred to the server.

[0817] Step 3:

[0818] Converting audio data to text

[0819] Input: The audio data sent to the server.

[0820] How it works: The server uses a generative AI model (e.g., the speech_recognition library) to convert the audio data to text.

[0821] Output: Text data is generated.

[0822] Step 4:

[0823] Emotion analysis

[0824] Input: The text data generated in step 3.

[0825] How it works: The server uses an emotion engine (e.g., the sentiment-analysis model from the transformers library) to analyze user sentiment from text data.

[0826] Output: Emotion data is generated.

[0827] Step 5:

[0828] Fraudulent Keyword Detection

[0829] Input: Text data.

[0830] How it works: The server scans the text data using a pre-defined list of keywords (e.g., "deposit money," "bank account," etc.) to detect fraud-related keywords in the text data.

[0831] Output: The result whether fraud keywords were detected or not.

[0832] Step 6:

[0833] Determining suspected fraud

[0834] Input: Sentiment data and fraud keyword detection results.

[0835] Operation: The server makes a comprehensive judgment on whether there is suspicion of fraud based on the emotional data and the fraud keyword detection results.

[0836] Output: Judgment result on whether fraud is suspected.

[0837] Step 7:

[0838] Generate an auto-response

[0839] Input: Suspected fraud judgment results and sentiment data.

[0840] How it works: If fraud is suspected, the server uses a generative AI model to generate an automated response that dynamically adjusts based on sentiment data.

[0841] Output: A well-tuned auto-response.

[0842] Step 8:

[0843] Playing an Auto Attendant

[0844] Input: The generated auto-response.

[0845] How it works: The generated automated response is output as audio and played back to the user through the device.

[0846] Output: The audio response played to the user.

[0847] Step 9:

[0848] Record keeping and provision

[0849] Input: call content, fraud detection results, emotional data.

[0850] How it works: The server stores suspected fraudulent calls and their analysis results in a database, and provides them to relevant authorities as needed.

[0851] Output: Stored records and provided data.

[0852] Step 10:

[0853] User Notification

[0854] Input: Fraud determination result.

[0855] How it works: The server generates a warning message and sends it to the user via their device to inform them of the suspected fraud.

[0856] Output: The warning message sent to the user.

[0857] These steps will enable an advanced prevention system that analyzes call content in real time and protects users from the risk of special fraud.

[0858] The specific processing unit 290 transmits the result of the specific processing to the smart glasses 214. In the smart glasses 214, the control unit 46A causes the speaker 240 to output the result of the specific processing. The microphone 238 acquires audio indicating a user input regarding the result of the specific processing. The control unit 46A transmits audio data indicating the user input acquired by the microphone 238 to the data processing device 12. In the data processing device 12, the specific processing unit 290 acquires the audio data.

[0859] The data generation model 58 is a so-called generative AI (Artificial Intelligence). An example of the data generation model 58 is ChatGPT (Internet Search<URL: https: / / openai.com / blog / chatgpt> ), Gemini (Internet search <url: https: gemini.google.com ?hl="ja">) and other generation AIs. The data generation model 58 is obtained by performing deep learning on a neural network. A prompt including an instruction is input to the data generation model 58, and inference data such as voice data indicating voice, text data indicating text, and image data indicating an image is also input. The data generation model 58 performs inference on the input inference data in accordance with the instruction indicated by the prompt, and outputs the inference result in a data format such as voice data and text data. Here, inference refers to, for example, analysis, classification, prediction, and / or summarization.

[0860] In the above embodiment, an example in which the specific processing is performed by the data processing device 12 has been given, but the technology of the present disclosure is not limited to this, and the specific processing may be performed by the smart glasses 214.

[0861] [Third embodiment]

[0862] FIG. 5 shows an example of the configuration of a data processing system 310 according to the third embodiment.

[0863] 5, the data processing system 310 includes the data processing device 12 and a headset terminal 314. An example of the data processing device 12 is a server.

[0864] The data processing device 12 includes a computer 22, a database 24, and a communication I / F 26. The computer 22 is an example of a "computer" according to the technology of the present disclosure. The computer 22 includes a processor 28, a RAM 30, and a storage 32. The processor 28, the RAM 30, and the storage 32 are connected to a bus 34. The database 24 and the communication I / F 26 are also connected to the bus 34. The communication I / F 26 is connected to a network 54. Examples of the network 54 include a WAN (Wide Area Network) and / or a LAN (Local Area Network).

[0865] The headset type terminal 314 includes a computer 36, a microphone 238, a speaker 240, a camera 42, a communication I / F 44, and a display 343. The computer 36 includes a processor 46, a RAM 48, and a storage 50. The processor 46, the RAM 48, and the storage 50 are connected to a bus 52. The microphone 238, the speaker 240, the camera 42, and the display 343 are also connected to the bus 52.

[0866] The microphone 238 receives instructions and the like from the user 20 by receiving voice uttered by the user 20. The microphone 238 captures the voice uttered by the user 20, converts the captured voice into audio data, and outputs it to the processor 46. The speaker 240 outputs audio in accordance with instructions from the processor 46.

[0867] Camera 42 is a small digital camera equipped with an optical system including a lens, aperture, and shutter, and an imaging element such as a CMOS (Complementary Metal-Oxide-Semiconductor) image sensor or a CCD (Charge Coupled Device) image sensor, and captures images of the surroundings of user 20 (for example, an imaging range defined by an angle of view equivalent to the field of vision of a typical healthy person).

[0868] The communication I / F 44 is connected to a network 54. The communication I / Fs 44 and 26 are responsible for the exchange of various information between the processor 46 and the processor 28 via the network 54. The exchange of various information between the processor 46 and the processor 28 using the communication I / Fs 44 and 26 is carried out in a secure state.

[0869] Fig. 6 shows an example of the main functions of the data processing device 12 and the headset type terminal 314. As shown in Fig. 6, in the data processing device 12, a specific process is performed by the processor 28. A specific process program 56 is stored in the storage 32.

[0870] The specific processing program 56 is an example of a "program" according to the technology of the present disclosure. The processor 28 reads the specific processing program 56 from the storage 32 and executes the read specific processing program 56 on the RAM 30. The specific processing is realized by the processor 28 operating as a specific processing unit 290 in accordance with the specific processing program 56 executed on the RAM 30.

[0871] The storage 32 stores a data generation model 58 and an emotion identification model 59. The data generation model 58 and the emotion identification model 59 are used by the identification processing unit 290.

[0872] In the headset type terminal 314, a reception output process is performed by the processor 46. A reception output program 60 is stored in the storage 50. The processor 46 reads the reception output program 60 from the storage 50 and executes the read reception output program 60 on the RAM 48. The reception output process is realized by the processor 46 operating as the control unit 46A in accordance with the reception output program 60 executed on the RAM 48.

[0873] Next, a description will be given of the identification process performed by the identification processing unit 290 of the data processing device 12. In the following description, the data processing device 12 will be referred to as the "server" and the headset type terminal 314 will be referred to as the "terminal."

[0874] This invention relates to a system for preventing special fraud during a call. Below, each element of this system and its processing will be explained in natural language.

[0875] Start recording a call

[0876] 1. Device:

[0877] It has a function that automatically starts recording when a call is started, so that audio data is recorded from the moment the user starts a call.

[0878] Sending audio data

[0879] 2. Terminal:

[0880] Recorded audio data is sent to the server in fixed buffer sizes, allowing the content of the call to be analyzed in real time.

[0881] Converting audio data to text

[0882] 3. Server:

[0883] The voice data sent from the device is received and converted into text in real time using a generation AI, which uses voice recognition technology to convert it into text information accurately and quickly.

[0884] Text data analysis

[0885] 4. Server:

[0886] The text of the conversation is analyzed, where fraud-related keywords and similar words are retrieved from a database, and an analysis algorithm is used to determine whether there is any suspicion of fraud.

[0887] What to do if fraud is suspected

[0888] 5. Server:

[0889] If fraud is suspected, the AI ​​automatically takes over the conversation, generating an appropriate response and sending it to the device.

[0890] 6. Terminal:

[0891] It responds on behalf of the user by playing back an automated response received from the server, which is designed to be unpleasant for scammers.

[0892] Record keeping and provision

[0893] 7. Server:

[0894] The data of calls suspected to be fraudulent will be saved and later provided to the police and other relevant authorities. This information will be used to help prevent fraud and to prevent future victims from becoming victims of fraud.

[0895] Specific examples

[0896] Here is a specific scenario:

[0897] 1. User:

[0898] Elderly user receives a call.

[0899] 2. Terminal:

[0900] When a call is initiated, the device automatically starts recording, and the recorded audio data is sent to the server in real time.

[0901] 3. Server:

[0902] The received voice data is converted into text using a generative AI, and keywords related to fraud, such as "depositing money" and "bank account," are analyzed.

[0903] If the server suspects fraud, it will activate an automated response AI that will generate a response such as "Please wait a moment, we're checking."

[0904] 4. Terminal:

[0905] Based on instructions from the server, an automated response is played and the conversation continues to be recorded.

[0906] 5. Server:

[0907] After the service is terminated, conversation data suspected of fraud will be saved and provided to the police on a regular basis.

[0908] This will help prevent victims of special fraud and allow all users, including the elderly, to use the telephone with peace of mind.

[0909] The processing flow will be explained below.

[0910] Step 1:

[0911] User: Start a call.

[0912] Device: Recording will start automatically when a call is initiated.

[0913] Step 2:

[0914] Terminal: Recorded audio data is sent to the server at fixed buffer size intervals.

[0915] Step 3:

[0916] Server: Receives voice data sent from the device and converts the voice into text using generative AI.

[0917] Step 4:

[0918] Server: Analyzes the text of the conversation, retrieves fraud-related keywords and similar words from a database, and determines whether there is any suspicion of fraud.

[0919] Step 5:

[0920] Server: If fraud is suspected, an automated response AI is activated to generate an appropriate response.

[0921] Terminal: Plays back the generated response aloud and acts on behalf of the user.

[0922] Step 6:

[0923] Device: Continue recording and continue sending conversation data to the server.

[0924] Step 7:

[0925] Server: Stores conversation data suspected of fraud in a database. This data will be provided to relevant agencies such as the police at a later date.

[0926] Step 8:

[0927] Server: If necessary, notify the user of suspected fraud.

[0928] This series of steps will protect users from special fraud by monitoring call content in real time and automatically responding if fraud is suspected.

[0929] Example 1

[0930] Next, a description will be given of Example 1. In the following description, the data processing device 12 will be referred to as a "server" and the headset type terminal 314 will be referred to as a "terminal."

[0931] Special frauds are becoming more sophisticated every day, and many people, including the elderly, are falling victim to them. Effective measures to prevent these frauds are needed, but existing technology makes it difficult to detect fraud during a call in real time and take appropriate action. Furthermore, there is no system that can instantly recognize and respond to suspected fraudulent calls. Therefore, there is a need to provide a system that prevents fraud during a call and create an environment in which users can use the telephone with peace of mind.

[0932] The specific processing by the specific processing unit 290 of the data processing device 12 in the first embodiment is realized by the following means.

[0933] In this invention, the server includes a means for transmitting the voice data of the call at fixed buffer size intervals, a means for converting the received voice data into text in real time using a generation AI, a means for retrieving and detecting fraud-related keywords from a database, a means for the generation AI to automatically take over the conversation and generate and send an appropriate response if fraud is suspected, and a means for saving records of suspected fraud and providing them to relevant authorities. This makes it possible to detect special frauds in real time during a call and automatically provide an appropriate response, significantly reducing the risk of users becoming victims of fraud.

[0934] The "means for automatically starting recording when a call is started" refers to a device or software that automatically starts a recording function the moment the user starts a call and records the audio data of the call.

[0935] The "means for transmitting voice data of a call at fixed buffer size intervals" refers to a device or software for transmitting collected voice data to a server at fixed volume or time intervals.

[0936] "Means for converting received voice data into text in real time using a generation AI" refers to a device or software in which a server receives voice data sent from a terminal and converts the voice into text in real time using a generation AI.

[0937] "Means for retrieving and detecting fraud-related keywords from a database" refers to a device or software for searching a database for keywords that indicate signs of fraud contained in text-converted voice data and detecting them.

[0938] "Means for a generation AI to automatically take over a conversation when fraud is suspected, and generate and send an appropriate response" refers to a device or software that, when it is determined that fraud is suspected, automatically generates an appropriate countermeasure response from a generation AI and sends it on behalf of the user.

[0939] "Means for storing records of suspected fraud and providing them to relevant authorities" refers to devices or software that record the content of suspected fraudulent calls and the results of their analysis, and provide them to relevant authorities at a later date.

[0940] "Means for generating lines for automated responses using a generation AI and playing them audibly" refers to a device or software that converts the text of an automated response generated by a generation AI into audio and plays it back through a speaker or the like.

[0941] "Means for notifying the user when suspected fraud is detected" refers to a device or software that notifies the user of the information when the system determines that there is a suspicion of fraud.

[0942] This invention relates to a system for preventing special fraud during a call. Each element of this system and its processing will be specifically described below.

[0943] System configuration and processing

[0944] Call recording

[0945] When a user picks up the receiver or presses the call button to start a call, the device's recording function automatically starts and records the audio data of the call. Once the recorded audio data reaches a certain buffer size, it is sent from the device to the server. This is usually done using an HTTP POST request.

[0946] Converting audio data to text

[0947] The server receives the voice data sent from the device and converts it into text in real time using a generation AI (e.g., DeepSpeech, Google Cloud Speech-to-Text, etc.) This generation AI inputs the voice data and outputs the corresponding text data.

[0948] Text data analysis

[0949] The text data is first stored in a database on the server, and then keywords related to fraud are retrieved from the database and detected. This analysis is performed using regular expressions and natural language processing tools (e.g., NLTK). Based on the keyword detection results, a judgment is made as to whether there is suspicion of fraud, and if so, the system proceeds to the next step.

[0950] Generate and send automatic responses

[0951] If fraud is suspected, the server uses a generative AI model (e.g., GPT-3) to generate an appropriate automated response. A specific example of the generative AI's prompt is, "This call is suspected to be fraudulent. Please play the next message." The generated automated response is sent to the device, which plays it aloud. This allows the server to respond appropriately to the fraudster on behalf of the user.

[0952] Record keeping and provision

[0953] Records of calls suspected of fraud are stored on a server. This information will be provided to relevant authorities, such as the police, at a later date. The stored data includes voice data, text data, detected keywords, and responses generated by the AI.

[0954] Specific examples

[0955] The following is a specific example:

[0956] The user is an elderly person and receives a phone call. The device automatically records the call and sends the recorded voice data to the server in real time. The server converts the voice data into text using a generative AI and detects fraud-related keywords such as "depositing money" and "bank account." If the server determines that there is suspicion of fraud, it uses the generative AI to generate an automated response such as "Please wait a moment, we're checking," and sends it to the device. The device then plays this automated response aloud to ensure the user's safety.

[0957] This system will prevent victims of special fraud before they occur, and will enable all users, including the elderly, to use the telephone with peace of mind.

[0958] Prompt Sentence Examples

[0959] 1. "This call is suspected to be fraudulent. Please play the following message."

[0960] 2. "Do not provide bank account information."

[0961] 3. "Always check with family and friends before sending money."

[0962] Using such prompts allows the generative AI to provide appropriate warnings and responses to scammers.

[0963] The flow of the identification process in the first embodiment will be described with reference to FIG.

[0964] Step 1:

[0965] The user answers or makes a call. The input is the action of lifting the telephone receiver or pressing the talk button, and the output is the start of the call signal, which starts recording the call.

[0966] Step 2:

[0967] The device detects that a call has started and automatically starts recording. The input is a call start signal, and the output is recording of audio data. The device's internal recording function is activated, and audio data is collected through the microphone.

[0968] Step 3:

[0969] The device collects recorded audio data in fixed buffer sizes and sends them to the server in bulk. The input is the audio data being recorded, and the output is the audio data buffer. When the buffer reaches a certain size, it is sent to the server via an HTTP POST request.

[0970] Step 4:

[0971] The server receives the audio data sent from the device. The input is the audio data sent via an HTTP POST request, and the output is the audio data stored on the server. The server temporarily stores this audio data and proceeds to the next step.

[0972] Step 5:

[0973] The server converts the received voice data into text in real time using a generation AI (e.g., DeepSpeech, Google Cloud Speech-to-Text). The input is voice data, and the output is text data. The generation AI analyzes the voice data and generates the corresponding text.

[0974] Step 6:

[0975] The server analyzes the data converted to text by the generative AI and detects fraud-related keywords. The input is text data, and the output is the keyword detection results. A natural language processing tool (e.g., NLTK) is used to search for fraud-related keywords.

[0976] Step 7:

[0977] The server determines whether fraud is suspected based on the keyword detection results. The input is the keyword detection results, and the output is a flag indicating whether fraud is suspected or not. A judgment algorithm (e.g., decision tree, SVM) is executed.

[0978] Step 8:

[0979] If the server determines that fraud is suspected, it uses a generative AI model (e.g., GPT-3) to generate an appropriate automated response. The input is a flag indicating suspected fraud, and the output is the automated response. The generative AI inputs a prompt sentence and generates a corresponding response.

[0980] Step 9:

[0981] The server sends the generated automatic response to the terminal. The input is the automatic response text, and the output is the data to be sent to the terminal. An HTTP POST request is used for sending.

[0982] Step 10:

[0983] The terminal plays back the received automated response lines aloud. The input is the data sent from the server, and the output is the audio data. The TTS (Text-to-Speech) function is used to convert the text data into audio, which is then played back through the speaker.

[0984] Step 11:

[0985] The server stores records of suspected fraudulent phone calls and periodically provides them to relevant authorities (e.g., police). The input is the call's audio data, text data, and analysis results, and the output is the stored recorded data and data provided to relevant authorities. The records are stored in a database and provided to relevant authorities via an API or interface.

[0986] (Application example 1)

[0987] Next, a description will be given of Application Example 1. In the following description, the data processing device 12 will be referred to as a "server" and the headset type terminal 314 will be referred to as a "terminal."

[0988] In modern society, the number of fraud cases is increasing, and vulnerable people such as the elderly are often the targets. To solve this problem, a system that can identify fraudulent activities in real time and take prompt countermeasures is needed. However, current systems have difficulty monitoring the content of phone calls one by one, making it difficult to prevent fraudulent activities. Therefore, there is a need to provide a system that can detect fraudulent activities early and respond appropriately.

[0989] The specific processing by the specific processing unit 290 of the data processing device 12 in the application example 1 is realized by the following means.

[0990] In this invention, the server includes means for automatically starting recording when a call is initiated, means for transmitting the voice data of the call, means for converting the received voice data into text using a generating AI, means for detecting fraud-related keywords, means for the AI ​​to automatically take over the conversation if fraud is suspected, means for recording data in real time and providing it to relevant authorities, and means for saving and providing records of suspected fraud. This makes it possible to detect special frauds in real time and take appropriate action.

[0991] "Means for automatically starting recording when a call starts" refers to a function that automatically starts recording audio data as soon as a call starts.

[0992] "Means for transmitting voice data of a call" refers to a function for transmitting recorded voice data to a server in real time or at fixed buffer size intervals.

[0993] "Means of converting received voice data into text using generation AI" refers to the function of converting voice data received by the server into text data using generation AI technology.

[0994] "Means for detecting fraud-related keywords" refers to the function of searching for and detecting specific fraud-related keywords from text data.

[0995] "Means for AI to automatically take over the conversation if fraud is suspected" refers to a function in which AI automatically generates and responds to conversations when it is determined that fraud is suspected.

[0996] "Means of recording data in real time and providing it to relevant authorities" refers to the function of recording the contents of phone calls in real time and providing it to relevant authorities such as the police if necessary.

[0997] "Means for storing and providing records of suspected fraud" refers to the function of storing records of phone calls suspected of fraud and providing them to relevant authorities at a later date if necessary.

[0998] "Means for generating lines for automatic responses using generative AI and playing them back as audio" refers to the function of creating the content of automatic responses using generative AI technology and playing them back as audio.

[0999] "Means of notifying users when suspected fraud is detected" refers to a function that warns or notifies users when suspected fraud is detected.

[1000] This invention relates to a system for preventing special fraud during a call. Each element of this system and its processing will be explained below.

[1001] Terminal behavior:

[1002] The device has a means to automatically start recording when a call is initiated, and the recorded audio data is sent to the server in fixed buffer sizes, allowing the content of the call to be analyzed in real time.

[1003] Server behavior:

[1004] The server converts the received voice data into text using a generative AI model, "Google Cloud Speech-to-Text," which allows for accurate and fast conversion of voice to text.

[1005] The text data is then analyzed to detect fraud-related keywords, such as "bank transfer fraud," "depositing money," and "bank account," and an analysis algorithm is used to detect these.

[1006] If fraud is suspected, the server automatically activates AI to generate an appropriate response using a generative AI model such as OpenAI GPT-4. The generated response is sent to the device and played back as audio.

[1007] Data storage and provision:

[1008] The server stores records of calls suspected of fraud and has the means to provide them to relevant authorities if necessary, so that police and other relevant authorities can obtain the necessary information at a later date.

[1009] Examples:

[1010] Here's a specific scenario:

[1011] 1. User:

[1012] An elderly user receives a suspicious phone call.

[1013] 2. Terminal:

[1014] Recording will start automatically as soon as the call is started.

[1015] The recorded audio data is sent to the server in real time.

[1016] 3. Server:

[1017] The received voice data is converted into text using the generative AI model "Google Cloud Speech-to-Text."

[1018] From the text data, keywords related to fraud, such as "depositing money" and "bank account," are detected.

[1019] If it determines that fraud is suspected, an automatic response message is generated using OpenAI GPT-4.

[1020] For example, generate a response like "Please wait, we're checking."

[1021] 4. Terminal:

[1022] The automatic response message sent from the server is played back and the conversation is taken over.

[1023] 5. Server:

[1024] Records of calls suspected of fraud will be saved and provided to the police and other relevant authorities as necessary.

[1025] Example prompt sentence:

[1026] System: "If fraud is detected, generate an appropriate response."

[1027] User: "I'm from the bank. Please give me your account number so I can deposit some money."

[1028] This will enable vulnerable people, such as the elderly, to make calls safely without falling victim to fraud.

[1029] The flow of the specific processing in the application example 1 will be described with reference to FIG.

[1030] Step 1:

[1031] When a user starts a call, the device automatically starts recording. The input is the call audio, and the output is real-time recorded audio data. Specifically, the device's microphone captures the audio and stores it in a buffer.

[1032] Step 2:

[1033] The device sends recorded audio data to the server in fixed buffer size increments. The input is the recorded audio data, and the output is the buffer data sent to the server. Specifically, the device sends the audio data stored in the buffer to the server in real time.

[1034] Step 3:

[1035] The server converts the received voice data into text using a generative AI model. The input is voice data, and the output is text data. Specifically, it uses "Google Cloud Speech-to-Text" to convert the voice data into text data.

[1036] Step 4:

[1037] The server analyzes the text data and detects fraud-related keywords. The input is text data, and the output is the results of fraud keyword detection. Specifically, it uses an algorithm to search for specific keywords (e.g., "transfer fraud" or "bank account") within the text data.

[1038] Step 5:

[1039] If the server determines that fraud is suspected, it activates a generative AI model to generate an appropriate automated response. The input is the fraud keyword detection results and the conversation text, and the output is the generated automated response text. Specific operations include using "OpenAI GPT-4" to generate a response based on the prompt text. Examples of prompt text include "If fraud is detected, please generate an appropriate response" and "This is from a bank. Please tell me your account number so I can deposit the money."

[1040] Step 6:

[1041] The device receives the automated response sent from the server and plays it back as audio. The input is the generated automated response text, and the output is the automated response in audio. Specifically, the AI ​​response is played back as audio using the device's speaker.

[1042] Step 7:

[1043] The server stores call records suspected of fraud and provides them to relevant authorities as necessary. The input is the call records and fraud detection results, and the output is the stored data and submitted data. Specifically, the server stores call records suspected of fraud in a database and provides them to relevant authorities such as the police.

[1044] This will create a system that can detect fraud in real time and respond appropriately.

[1045] Furthermore, an emotion engine that estimates the user's emotion may be further combined. That is, the identification processing unit 290 may estimate the user's emotion using the emotion identification model 59, and perform identification processing using the user's emotion.

[1046] This invention combines an emotion engine with a system for preventing special fraud during phone calls. Below, each element of this system and its processing will be explained in natural language.

[1047] Start recording a call

[1048] 1. User: Start a call.

[1049] 2. Device: Automatically starts recording when a call is initiated. This means that audio data is recorded from the moment the user starts a call.

[1050] Sending audio data

[1051] 3. Terminal: Recorded audio data is sent to the server at fixed buffer size intervals.

[1052] Speech data conversion and sentiment analysis

[1053] 4. Server: Receives the voice data sent from the device, converts the voice to text using generative AI, and analyzes the user's emotions from the voice data using an emotion engine.

[1054] Analysis of text and emotion data

[1055] 5. Server: Analyzes the text of the conversation. Retrieves fraud-related keywords and similar words from the database, and uses an analysis algorithm to determine whether there is suspicion of fraud. Taking into account the emotional data analyzed by the emotion engine, the server strengthens detection of suspected fraud.

[1056] What to do if fraud is suspected

[1057] 6. Server: If fraud is suspected, the automated AI generates an appropriate response. The response is dynamically adjusted based on the user's emotional data obtained from the emotion engine.

[1058] 7. Device: Plays an automated response and responds on behalf of the user. During this time, the call continues to be recorded and the conversation data continues to be sent to the server.

[1059] Record keeping and provision

[1060] 8. Server: Conversation data suspected of fraud is stored in a database. The data will be provided to relevant agencies such as the police at a later date. The analysis results of the emotion engine will also be stored.

[1061] User Notification

[1062] 9. Server: Notify the user of suspected fraud and sentiment analysis results, if necessary.

[1063] Specific examples

[1064] Here is a specific scenario:

[1065] 1. User: An elderly user receives the call.

[1066] 2. Device: When a call is started, the device automatically starts recording. The recorded audio data is sent to the server in real time.

[1067] 3. Server: The received voice data is converted into text using a generative AI, and then the emotion engine analyzes the user's emotions. Fraud-related keywords such as "depositing money" and "bank account" are detected, and a suspicion of fraud is determined based on the analysis results and the user's emotional data.

[1068] 4. Server: If the server detects a suspected fraud, it will activate an automated AI response, generating a response such as "Please wait a moment, we're checking." The response is dynamically adjusted based on the results of sentiment analysis.

[1069] 5. Terminal: Based on instructions from the server, it plays an automated response and continues to record the conversation.

[1070] 6. Server: After the call ends, the server stores the conversation data and sentiment analysis results of any suspected fraud and periodically provides them to the police.

[1071] This series of steps protects users from specialized fraud and enables more advanced fraud detection and response through the emotion engine, providing an environment where users can use their phones with peace of mind.

[1072] The processing flow will be explained below.

[1073] Step 1:

[1074] User: Start a call.

[1075] Device: Automatically starts recording as soon as the call starts.

[1076] Step 2:

[1077] Terminal: Recorded audio data is sent to the server in real time at fixed buffer size intervals.

[1078] Step 3:

[1079] Server: Receives voice data sent from the device and converts it into text using a generation AI. The converted text data is temporarily stored.

[1080] Step 4:

[1081] Server: At the same time, the emotion engine is used to analyze the user's emotions from the voice data, and the analysis results are temporarily saved.

[1082] Step 5:

[1083] Server: The server compares the text of the conversation with a database of fraud-related keywords to determine whether there is any suspicion of fraud. It also considers the analysis results of the emotion engine to enhance detection of suspected fraud.

[1084] Step 6:

[1085] Server: If fraud is suspected, the automated AI responds by generating an appropriate response. The response is dynamically adjusted based on the user's emotional data obtained from the emotion engine.

[1086] Step 7:

[1087] Terminal: Plays back the automated response sent from the server and responds on behalf of the user. Records during the conversation and continues to send voice data to the server.

[1088] Step 8:

[1089] Server: After the call ends, the conversation data suspected of fraud and the results of sentiment analysis are stored in a database.

[1090] Step 9:

[1091] Server: Provides the relevant data to the police and other relevant authorities on a regular basis. If necessary, notifies the user of suspected fraud or the results of sentiment analysis.

[1092] This series of steps will protect users from specialized fraud, enable advanced fraud detection and response using an emotion engine, and provide recorded data to police to help prevent future crimes.

[1093] Example 2

[1094] Next, a description will be given of Example 2. In the following description, the data processing device 12 will be referred to as a "server" and the headset type terminal 314 will be referred to as a "terminal."

[1095] Currently, special frauds over the phone are on the rise, and elderly people are particularly vulnerable to these crimes. Conventional countermeasures have limitations, and more effective systems are needed to detect suspected fraud in real time and protect users. It is also important to ensure that records of fraudulent conversations are stored and provided to the necessary authorities.

[1096] The specific processing by the specific processing unit 290 of the data processing device 12 in the second embodiment is realized by the following means.

[1097] In this invention, the server includes a means for converting the voice data of the call into text using a generation AI, a means for analyzing emotions from the converted text, and a means for detecting fraud-related keywords and determining suspicion of fraud based on the emotion analysis results. This makes it possible to analyze the content of the call in real time and strengthen suspicion of fraud based on the emotion data.

[1098] The "means for starting recording" is a device or program that has the function of automatically starting recording of voice data as soon as a call is started.

[1099] The "means for transmitting audio data" is a device or program for transmitting recorded audio data to a server in real time with a certain buffer size.

[1100] "Means for converting audio data into text using generative AI" refers to an artificial intelligence algorithm or software for automatically converting audio data into text data.

[1101] "Means for analyzing emotions" refers to software or algorithms that analyze the user's emotional state from textual data and obtain the results.

[1102] A "means for detecting fraud-related keywords" is a program or algorithm for extracting specific words or phrases related to fraudulent activities from text data.

[1103] "Means for determining suspicion of fraud by taking into account the results of sentiment analysis" refers to software or algorithms that have the function of comprehensively evaluating and determining the possibility of fraud based on the results of sentiment analysis data and keywords related to fraud.

[1104] "Means for automatic AI conversation takeover" refers to a program or device that allows artificial intelligence to automatically continue a conversation on behalf of a user when a suspected fraud is determined.

[1105] "Means for adjusting the content of the automated response based on the results of sentiment analysis" refers to software or algorithms for appropriately changing the content and tone of the automated response message based on the results of sentiment analysis.

[1106] The "means for playing the message aloud" refers to a device or software for converting the generated automated response message into aloud voice and playing it back to the other party on the call.

[1107] The "means of storing and providing records of suspected fraud" refers to a system for storing call records and analysis data of suspected fraud in a database and providing them to relevant authorities as necessary.

[1108] This invention relates to a system for preventing special fraud during phone calls, which operates by combining a generative AI and an emotion engine. This system is realized using the following specific hardware and software.

[1109] Hardware and Software Configuration

[1110] server

[1111] 1. Generative AI model (e.g., Google Cloud's Speech-to-Text API): Converts voice data into text data.

[1112] 2. Sentiment engine (e.g., Microsoft Azure's Text Analytics API): Analyzes emotions from text data.

[1113] 3. Automated response AI (e.g., OpenAI’s GPT-3): Generates appropriate responses in cases of suspected fraud.

[1114] Terminal

[1115] 1. Calling app (e.g. Skype, Google Voice): An application for making calls.

[1116] 2. Recording function (e.g., iOS's Call Recorder app): A function that automatically records phone calls.

[1117] 3. Text-to-speech API (e.g., Amazon Polly): Speech the generated text response.

[1118] Overview of program processing

[1119] Start recording a call

[1120] When a user makes or receives a call using a calling app, the device automatically activates the recording function and records the audio data. This audio data is sent to the server in real time at fixed buffer size intervals.

[1121] Speech data conversion and sentiment analysis

[1122] The server uses a generative AI model to convert the voice data into text, which is then analyzed by an emotion engine to determine the user's emotional state. For example, emotions such as "happy," "sad," and "angry" are detected.

[1123] Determine suspected fraud

[1124] Based on the textual content of the conversation and the emotional data, the server detects keywords related to fraud. For example, it scans for words such as "money," "bank," and "account." At the same time, it also takes into account the results of the emotional analysis and makes a comprehensive judgment on the suspicion of fraud. For example, if the phrase "depositing money" matches the emotional data of "anxiety," it is determined that there is a high suspicion of fraud.

[1125] Auto-response generation and playback

[1126] If suspected fraud is detected, the server activates an automated AI to generate an appropriate response. This response is adjusted based on the results of sentiment analysis. The generated response is sent in text format to the device, which then uses a speech synthesis API to convert it into audio and play it back to the other party. For example, the response could be, "Please wait a moment, we're checking."

[1127] Record keeping and provision

[1128] Conversation data and sentiment analysis results for suspected fraud are stored in a secure database on a server. This data can be provided to relevant authorities as needed, such as the police, to enable further investigation and countermeasures.

[1129] User Notifications

[1130] If a suspected fraud is detected, the server will notify the user via email and / or SMS.

[1131] Examples of concrete examples and prompts

[1132] Specific examples

[1133] When an elderly user answers the phone and the call begins, the device automatically starts recording. This audio data is sent to a server in real time and converted into text by a generative AI model. If the phrase "depositing money" and the emotion "anxiety" are detected, the server activates an automated response AI and generates a response saying, "Please wait a moment, we're checking." This response is then converted into audio and played back to the caller.

[1134] Prompt Sentence Examples

[1135] The voice data is converted into text and analyzed, and if it contains keywords such as "money" or "bank," it is judged to be suspected of fraud. For example, if the phrase "depositing money" is included, it is considered highly suspected of fraud.

[1136] This effectively protects users from specialized fraud while on the phone, while the emotion engine enables more advanced fraud detection and response.

[1137] The flow of the identification process in the second embodiment will be described with reference to FIG.

[1138] Processing Steps

[1139] Step 1: Start recording the call

[1140] User: A user initiates or receives a call using a calling app.

[1141] Terminal: When a call is initiated, the recording function is automatically activated and audio data is recorded.

[1142] Input: Call Start Trigger

[1143] Output: Recorded audio data

[1144] Specific operation: When the device's calling app detects the start of a call, the recording function is activated and the call audio from that point on is saved as digital data.

[1145] Step 2: Sending audio data

[1146] Terminal: Recorded audio data is sent to the server in real time at fixed buffer size intervals.

[1147] Input: Recorded audio data

[1148] Output: Audio data sent to the server

[1149] What it does: The audio data is split into one-minute buffers and sent to the server using an HTTP POST request.

[1150] Step 3: Converting audio data to text

[1151] Server: Uses a generative AI model to convert voice data into text data.

[1152] Input: Audio data received from the device

[1153] Output: Text data

[1154] Specific operation: After receiving the voice data, the server calls Google Cloud's Speech-to-Text API to convert the voice to text.

[1155] Step 4: Sentiment Analysis

[1156] Server: Analyzes emotions from text data using an emotion engine.

[1157] Input: Text data

[1158] Output: Emotion analysis results

[1159] How it works: The text data is sent to Microsoft Azure's Text Analytics API, where it is analyzed to see if it contains emotions such as "happiness," "sadness," or "anger."

[1160] Step 5: Determine if there is any fraud

[1161] Server: Based on the text data and sentiment analysis results, detects fraud-related keywords and determines whether fraud is suspected.

[1162] Input: Text data, sentiment analysis results

[1163] Output: Suspected fraud determination result

[1164] How it works: The server scans the text data and compares it with a list of keywords (e.g., "money," "bank," "account"), while also taking into account the results of sentiment analysis to make a comprehensive judgment on the suspicion of fraud.

[1165] Step 6: Generate an autoresponder

[1166] Server: If fraud is suspected, an automated AI is used to generate an appropriate response.

[1167] Input: Suspected fraud judgment result

[1168] Output: Auto-response text

[1169] Specific operation: The server calls OpenAI's GPT-3 and generates a response message based on the situation, such as "Please wait a moment, checking."

[1170] Step 7: Play Auto Attendant Audio

[1171] Terminal: The text of the automated response is converted into speech using a speech synthesis API and played back to the other party.

[1172] Input: Auto-response text

[1173] Output: Voiced auto attendant

[1174] How it works: The device uses a speech synthesis API such as Amazon Polly to convert the generated text into speech and play it back to the other party in real time.

[1175] Step 8: Record keeping and provision

[1176] Server: Stores suspected fraudulent conversation data and sentiment analysis results in a database and provides them to relevant authorities as needed.

[1177] Input: Suspected fraud voice data, emotion analysis results

[1178] Output: Data stored, data provided

[1179] What it does: The server stores suspected fraudulent conversation data in a secure database and provides it to police and other relevant authorities as needed.

[1180] Step 9: User Notification

[1181] Server: Notifies the user if suspected fraud is detected.

[1182] Input: Suspected fraud judgment result

[1183] Output: User notification

[1184] What it does: The server notifies the user of suspected fraud via email address or mobile phone number. This notification is sent via email using an SMTP server or as a text message using an SMS API.

[1185] Through the above steps, the system can detect suspected special fraud in real time during a call and effectively protect users.

[1186] (Application example 2)

[1187] Next, a description will be given of Application Example 2. In the following description, the data processing device 12 will be referred to as a "server" and the headset type terminal 314 will be referred to as a "terminal."

[1188] Conventional call recording systems could record phone conversations and, in some cases, detect keywords that could be suspected of fraud. However, they lacked the ability to analyze user sentiment and use that information to further improve the accuracy of fraud detection. As a result, even if fraud was suspected, it was not possible to evaluate whether the user actually perceived the risk, limiting the accuracy of fraud detection. Furthermore, they lacked the ability to provide appropriate automated responses, making it difficult to respond quickly in situations where there was a high risk of actually being scammed. The challenge is to solve these problems and provide a more advanced fraud prevention system.

[1189] The identification processing by the identification processing unit 290 of the data processing device 12 in Application Example 2 is realized by the following means. In this invention, the server includes means for transmitting voice data of the call, means for converting the received voice data into text using a generation AI, means for detecting fraud-related keywords, means for analyzing user emotions from the voice data using an emotion engine, means for dynamically adjusting automatic responses based on the emotion data, and means for saving and providing records of suspected fraud. This enables advanced fraud detection that takes into account user emotion analysis and the provision of appropriate and dynamically adjusted automatic responses.

[1190] A "call" is an act of exchanging voice information between two people using a telephone line.

[1191] "Recording" is the act of recording sound in digital or analog form.

[1192] "Generative AI" is a system that uses artificial intelligence technology to generate and convert voice and text.

[1193] "Text conversion" is the process of converting audio data into text data.

[1194] "Fraud Keywords" are words or phrases that are likely to be associated with fraudulent activity.

[1195] An "emotion engine" is a system that analyzes user emotions from voice and text data.

[1196] An "auto-response" is a message that the system automatically generates in response to a user.

[1197] "Dynamic adjustment" refers to changing the response in real time based on specific situations and conditions.

[1198] "Record keeping" is the act of storing data in a file or database.

[1199] "Notification" means the act of informing a user of specific information.

[1200] This invention is a system for preventing special fraud during phone calls, and aims to further improve accuracy by analyzing the user's emotions using an emotion engine.

[1201] The system is roughly configured as follows:

[1202] Hardware and Software Configuration

[1203] 1. Device:

[1204] A communication device such as a smartphone that has the ability to record voice data during calls and send it to a server.

[1205] 2. Server:

[1206] The server converts the received voice data into text using a generative AI model, detects fraud-related keywords, and analyzes the user's emotions using an emotion engine.

[1207] Program Overview

[1208] The system operates in the following steps:

[1209] Call Recording:

[1210] When a user starts a call, the device automatically starts recording, and the recorded audio data is sent to the server in real time.

[1211] Speech to text conversion and sentiment analysis:

[1212] The server converts the received voice data into text using a generative AI, and then uses an emotion engine to analyze emotions from the voice data. Specifically, the server uses the speech_recognition library and the transformers library.

[1213] Suspected fraud detection:

[1214] The server analyzes the converted text data to detect fraud-related keywords and, taking into account the results of the emotion engine analysis, determines whether or not there is any suspicion of fraud.

[1215] Generate an auto-response:

[1216] If fraud is suspected, the server will activate an automated AI response that generates a dynamically tailored response based on the results of sentiment analysis. This automated response is played back via voice and acts on behalf of the user.

[1217] Record-Keeping and Notification:

[1218] During and after the call, the server stores the records of suspected fraud and the results of the sentiment analysis. If necessary, the server periodically provides the data to the relevant authorities and notifies the user of the suspected fraud.

[1219] Specific examples

[1220] For example, if an elderly person receives a phone call and is told, "This is confirmation from the bank," the device automatically records the call and sends the audio data to a server. The server converts the audio data into text and analyzes the user's emotions using an emotion engine. If a fraud keyword such as "depositing money" is detected and the analysis finds that the elderly person is anxious, the server generates an automatic response such as "Are you okay? We're checking," and plays it back.

[1221] Prompt Sentence Examples

[1222] Prompt: Convert the speech data into text and detect whether it contains fraud-related keywords. Also, perform sentiment analysis and generate appropriate responses based on that.

[1223] In this way, it provides an advanced prevention system to prevent users from falling victim to fraud during calls.

[1224] The flow of the specific processing in the application example 2 will be described with reference to FIG.

[1225] Step 1:

[1226] Start recording a call

[1227] Input: The user receives a phone call.

[1228] How it works: When you start a call, your device automatically starts recording the call.

[1229] Output: The recorded audio data is generated.

[1230] Step 2:

[1231] Sending audio data

[1232] Input: Recorded audio data.

[1233] How it works: The device sends the recorded audio data to the server in real time.

[1234] Output: Audio data transferred to the server.

[1235] Step 3:

[1236] Converting audio data to text

[1237] Input: The audio data sent to the server.

[1238] How it works: The server uses a generative AI model (e.g., the speech_recognition library) to convert the audio data to text.

[1239] Output: Text data is generated.

[1240] Step 4:

[1241] Emotion analysis

[1242] Input: The text data generated in step 3.

[1243] How it works: The server uses an emotion engine (e.g., the sentiment-analysis model from the transformers library) to analyze user sentiment from text data.

[1244] Output: Emotion data is generated.

[1245] Step 5:

[1246] Fraudulent Keyword Detection

[1247] Input: Text data.

[1248] How it works: The server scans the text data using a pre-defined list of keywords (e.g., "deposit money," "bank account," etc.) to detect fraud-related keywords in the text data.

[1249] Output: The result whether fraud keywords were detected or not.

[1250] Step 6:

[1251] Determining suspected fraud

[1252] Input: Sentiment data and fraud keyword detection results.

[1253] Operation: The server makes a comprehensive judgment on whether there is suspicion of fraud based on the emotional data and the fraud keyword detection results.

[1254] Output: Judgment result on whether fraud is suspected.

[1255] Step 7:

[1256] Generate an auto-response

[1257] Input: Suspected fraud judgment results and sentiment data.

[1258] How it works: If fraud is suspected, the server uses a generative AI model to generate an automated response that dynamically adjusts based on sentiment data.

[1259] Output: A well-tuned auto-response.

[1260] Step 8:

[1261] Playing an Auto Attendant

[1262] Input: The generated auto-response.

[1263] How it works: The generated automated response is output as audio and played back to the user through the device.

[1264] Output: The audio response played to the user.

[1265] Step 9:

[1266] Record keeping and provision

[1267] Input: call content, fraud detection results, emotional data.

[1268] How it works: The server stores suspected fraudulent calls and their analysis results in a database, and provides them to relevant authorities as needed.

[1269] Output: Stored records and provided data.

[1270] Step 10:

[1271] User Notification

[1272] Input: Fraud determination result.

[1273] How it works: The server generates a warning message and sends it to the user via their device to inform them of the suspected fraud.

[1274] Output: The warning message sent to the user.

[1275] These steps will enable an advanced prevention system that analyzes call content in real time and protects users from the risk of special fraud.

[1276] The specific processing unit 290 transmits the result of the specific processing to the headset type terminal 314. In the headset type terminal 314, the control unit 46A causes the speaker 240 and the display 343 to output the result of the specific processing. The microphone 238 acquires audio indicating a user input regarding the result of the specific processing. The control unit 46A transmits audio data indicating the user input acquired by the microphone 238 to the data processing device 12. In the data processing device 12, the specific processing unit 290 acquires the audio data.

[1277] The data generation model 58 is a so-called generative AI (Artificial Intelligence). An example of the data generation model 58 is ChatGPT (Internet Search<URL: https: / / openai.com / blog / chatgpt> ), Gemini (Internet search <url: https: gemini.google.com ?hl="ja">) and other generation AIs. The data generation model 58 is obtained by performing deep learning on a neural network. A prompt including an instruction is input to the data generation model 58, and inference data such as voice data indicating voice, text data indicating text, and image data indicating an image is also input. The data generation model 58 performs inference on the input inference data in accordance with the instruction indicated by the prompt, and outputs the inference result in a data format such as voice data and text data. Here, inference refers to, for example, analysis, classification, prediction, and / or summarization.

[1278] In the above embodiment, an example was given in which the specific processing is performed by the data processing device 12, but the technology of the present disclosure is not limited to this, and the specific processing may be performed by the headset type terminal 314.

[1279] [Fourth embodiment]

[1280] FIG. 7 shows an example of the configuration of a data processing system 410 according to the fourth embodiment.

[1281] 7, a data processing system 410 includes a data processing device 12 and a robot 414. An example of the data processing device 12 is a server.

[1282] The data processing device 12 includes a computer 22, a database 24, and a communication I / F 26. The computer 22 is an example of a "computer" according to the technology of the present disclosure. The computer 22 includes a processor 28, a RAM 30, and a storage 32. The processor 28, the RAM 30, and the storage 32 are connected to a bus 34. The database 24 and the communication I / F 26 are also connected to the bus 34. The communication I / F 26 is connected to a network 54. Examples of the network 54 include a WAN (Wide Area Network) and / or a LAN (Local Area Network).

[1283] The robot 414 includes a computer 36, a microphone 238, a speaker 240, a camera 42, a communication I / F 44, and a control target 443. The computer 36 includes a processor 46, a RAM 48, and a storage 50. The processor 46, the RAM 48, and the storage 50 are connected to a bus 52. The microphone 238, the speaker 240, the camera 42, and the control target 443 are also connected to the bus 52.

[1284] The microphone 238 receives instructions and the like from the user 20 by receiving voice uttered by the user 20. The microphone 238 captures the voice uttered by the user 20, converts the captured voice into audio data, and outputs it to the processor 46. The speaker 240 outputs audio in accordance with instructions from the processor 46.

[1285] Camera 42 is a small digital camera equipped with an optical system including a lens, aperture, and shutter, and an imaging element such as a CMOS (Complementary Metal-Oxide-Semiconductor) image sensor or a CCD (Charge Coupled Device) image sensor, and captures images of the surroundings of user 20 (for example, an imaging range defined by an angle of view equivalent to the field of vision of a typical healthy person).

[1286] The communication I / F 44 is connected to a network 54. The communication I / Fs 44 and 26 are responsible for the exchange of various information between the processor 46 and the processor 28 via the network 54. The exchange of various information between the processor 46 and the processor 28 using the communication I / Fs 44 and 26 is carried out in a secure state.

[1287] The control object 443 includes a display device, LEDs in the eyes, and motors for driving the arms, hands, and feet. The posture and gestures of the robot 414 are controlled by controlling the motors of the arms, hands, and feet. Some of the emotions of the robot 414 can be expressed by controlling these motors. In addition, the facial expressions of the robot 414 can also be expressed by controlling the light emission state of the LEDs in the eyes of the robot 414.

[1288] Fig. 8 shows an example of the main functions of the data processing device 12 and the robot 414. As shown in Fig. 8, in the data processing device 12, a specific process is performed by the processor 28. A specific process program 56 is stored in the storage 32.

[1289] The specific processing program 56 is an example of a "program" according to the technology of the present disclosure. The processor 28 reads the specific processing program 56 from the storage 32 and executes the read specific processing program 56 on the RAM 30. The specific processing is realized by the processor 28 operating as a specific processing unit 290 in accordance with the specific processing program 56 executed on the RAM 30.

[1290] The storage 32 stores a data generation model 58 and an emotion identification model 59. The data generation model 58 and the emotion identification model 59 are used by the identification processing unit 290.

[1291] In the robot 414, the processor 46 performs the reception output process. A reception output program 60 is stored in the storage 50. The processor 46 reads the reception output program 60 from the storage 50 and executes the read reception output program 60 on the RAM 48. The reception output process is realized by the processor 46 operating as the control unit 46A in accordance with the reception output program 60 executed on the RAM 48.

[1292] Next, a description will be given of the specific processing performed by the specific processing unit 290 of the data processing device 12. In the following description, the data processing device 12 will be referred to as a "server" and the robot 414 will be referred to as a "terminal."

[1293] This invention relates to a system for preventing special fraud during a call. Below, each element of this system and its processing will be explained in natural language.

[1294] Start recording a call

[1295] 1. Device:

[1296] It has a function that automatically starts recording when a call is started, so that audio data is recorded from the moment the user starts a call.

[1297] Sending audio data

[1298] 2. Terminal:

[1299] Recorded audio data is sent to the server in fixed buffer sizes, allowing the content of the call to be analyzed in real time.

[1300] Converting audio data to text

[1301] 3. Server:

[1302] The voice data sent from the device is received and converted into text in real time using a generation AI, which uses voice recognition technology to convert it into text information accurately and quickly.

[1303] Text data analysis

[1304] 4. Server:

[1305] The text of the conversation is analyzed, where fraud-related keywords and similar words are retrieved from a database, and an analysis algorithm is used to determine whether there is any suspicion of fraud.

[1306] What to do if fraud is suspected

[1307] 5. Server:

[1308] If fraud is suspected, the AI ​​automatically takes over the conversation, generating an appropriate response and sending it to the device.

[1309] 6. Terminal:

[1310] It responds on behalf of the user by playing back an automated response received from the server, which is designed to be unpleasant for scammers.

[1311] Record keeping and provision

[1312] 7. Server:

[1313] The data of calls suspected to be fraudulent will be saved and later provided to the police and other relevant authorities. This information will be used to help prevent fraud and to prevent future victims from becoming victims of fraud.

[1314] Specific examples

[1315] Here is a specific scenario:

[1316] 1. User:

[1317] Elderly user receives a call.

[1318] 2. Terminal:

[1319] When a call is initiated, the device automatically starts recording, and the recorded audio data is sent to the server in real time.

[1320] 3. Server:

[1321] The received voice data is converted into text using a generative AI, and keywords related to fraud, such as "depositing money" and "bank account," are analyzed.

[1322] If the server suspects fraud, it will activate an automated response AI that will generate a response such as "Please wait a moment, we're checking."

[1323] 4. Terminal:

[1324] Based on instructions from the server, an automated response is played and the conversation continues to be recorded.

[1325] 5. Server:

[1326] After the service is terminated, conversation data suspected of fraud will be saved and provided to the police on a regular basis.

[1327] This will help prevent victims of special fraud and allow all users, including the elderly, to use the telephone with peace of mind.

[1328] The processing flow will be explained below.

[1329] Step 1:

[1330] User: Start a call.

[1331] Device: Recording will start automatically when a call is initiated.

[1332] Step 2:

[1333] Terminal: Recorded audio data is sent to the server at fixed buffer size intervals.

[1334] Step 3:

[1335] Server: Receives voice data sent from the device and converts the voice into text using generative AI.

[1336] Step 4:

[1337] Server: Analyzes the text of the conversation, retrieves fraud-related keywords and similar words from a database, and determines whether there is any suspicion of fraud.

[1338] Step 5:

[1339] Server: If fraud is suspected, an automated response AI is activated to generate an appropriate response.

[1340] Terminal: Plays back the generated response aloud and acts on behalf of the user.

[1341] Step 6:

[1342] Device: Continue recording and continue sending conversation data to the server.

[1343] Step 7:

[1344] Server: Stores conversation data suspected of fraud in a database. This data will be provided to relevant agencies such as the police at a later date.

[1345] Step 8:

[1346] Server: If necessary, notify the user of suspected fraud.

[1347] This series of steps will protect users from special fraud by monitoring call content in real time and automatically responding if fraud is suspected.

[1348] Example 1

[1349] Next, a description will be given of Example 1. In the following description, the data processing device 12 will be referred to as a "server" and the robot 414 will be referred to as a "terminal."

[1350] Special frauds are becoming more sophisticated every day, and many people, including the elderly, are falling victim to them. Effective measures to prevent these frauds are needed, but existing technology makes it difficult to detect fraud during a call in real time and take appropriate action. Furthermore, there is no system that can instantly recognize and respond to suspected fraudulent calls. Therefore, there is a need to provide a system that prevents fraud during a call and create an environment in which users can use the telephone with peace of mind.

[1351] The specific processing by the specific processing unit 290 of the data processing device 12 in the first embodiment is realized by the following means.

[1352] In this invention, the server includes a means for transmitting the voice data of the call at fixed buffer size intervals, a means for converting the received voice data into text in real time using a generation AI, a means for retrieving and detecting fraud-related keywords from a database, a means for the generation AI to automatically take over the conversation and generate and send an appropriate response if fraud is suspected, and a means for saving records of suspected fraud and providing them to relevant authorities. This makes it possible to detect special frauds in real time during a call and automatically provide an appropriate response, significantly reducing the risk of users becoming victims of fraud.

[1353] The "means for automatically starting recording when a call is started" refers to a device or software that automatically starts a recording function the moment the user starts a call and records the audio data of the call.

[1354] The "means for transmitting voice data of a call at fixed buffer size intervals" refers to a device or software for transmitting collected voice data to a server at fixed volume or time intervals.

[1355] "Means for converting received voice data into text in real time using a generation AI" refers to a device or software in which a server receives voice data sent from a terminal and converts the voice into text in real time using a generation AI.

[1356] "Means for retrieving and detecting fraud-related keywords from a database" refers to a device or software for searching a database for keywords that indicate signs of fraud contained in text-converted voice data and detecting them.

[1357] "Means for a generation AI to automatically take over a conversation when fraud is suspected, and generate and send an appropriate response" refers to a device or software that, when it is determined that fraud is suspected, automatically generates an appropriate countermeasure response from a generation AI and sends it on behalf of the user.

[1358] "Means for storing records of suspected fraud and providing them to relevant authorities" refers to devices or software that record the content of suspected fraudulent calls and the results of their analysis, and provide them to relevant authorities at a later date.

[1359] "Means for generating lines for automated responses using a generation AI and playing them audibly" refers to a device or software that converts the text of an automated response generated by a generation AI into audio and plays it back through a speaker or the like.

[1360] "Means for notifying the user when suspected fraud is detected" refers to a device or software that notifies the user of the information when the system determines that there is a suspicion of fraud.

[1361] This invention relates to a system for preventing special fraud during a call. Each element of this system and its processing will be specifically described below.

[1362] System configuration and processing

[1363] Call recording

[1364] When a user picks up the receiver or presses the call button to start a call, the device's recording function automatically starts and records the audio data of the call. Once the recorded audio data reaches a certain buffer size, it is sent from the device to the server. This is usually done using an HTTP POST request.

[1365] Converting audio data to text

[1366] The server receives the voice data sent from the device and converts it into text in real time using a generation AI (e.g., DeepSpeech, Google Cloud Speech-to-Text, etc.) This generation AI inputs the voice data and outputs the corresponding text data.

[1367] Text data analysis

[1368] The text data is first stored in a database on the server, and then keywords related to fraud are retrieved from the database and detected. This analysis is performed using regular expressions and natural language processing tools (e.g., NLTK). Based on the keyword detection results, a judgment is made as to whether there is suspicion of fraud, and if so, the system proceeds to the next step.

[1369] Generate and send automatic responses

[1370] If fraud is suspected, the server uses a generative AI model (e.g., GPT-3) to generate an appropriate automated response. A specific example of the generative AI's prompt is, "This call is suspected to be fraudulent. Please play the next message." The generated automated response is sent to the device, which plays it aloud. This allows the server to respond appropriately to the fraudster on behalf of the user.

[1371] Record keeping and provision

[1372] Records of calls suspected of fraud are stored on a server. This information will be provided to relevant authorities, such as the police, at a later date. The stored data includes voice data, text data, detected keywords, and responses generated by the AI.

[1373] Specific examples

[1374] The following is a specific example:

[1375] The user is an elderly person and receives a phone call. The device automatically records the call and sends the recorded voice data to the server in real time. The server converts the voice data into text using a generative AI and detects fraud-related keywords such as "depositing money" and "bank account." If the server determines that there is suspicion of fraud, it uses the generative AI to generate an automated response such as "Please wait a moment, we're checking," and sends it to the device. The device then plays this automated response aloud to ensure the user's safety.

[1376] This system will prevent victims of special fraud before they occur, and will enable all users, including the elderly, to use the telephone with peace of mind.

[1377] Prompt Sentence Examples

[1378] 1. "This call is suspected to be fraudulent. Please play the following message."

[1379] 2. "Do not provide bank account information."

[1380] 3. "Always check with family and friends before sending money."

[1381] Using such prompts allows the generative AI to provide appropriate warnings and responses to scammers.

[1382] The flow of the identification process in the first embodiment will be described with reference to FIG.

[1383] Step 1:

[1384] The user answers or makes a call. The input is the action of lifting the telephone receiver or pressing the talk button, and the output is the start of the call signal, which starts recording the call.

[1385] Step 2:

[1386] The device detects that a call has started and automatically starts recording. The input is a call start signal, and the output is recording of audio data. The device's internal recording function is activated, and audio data is collected through the microphone.

[1387] Step 3:

[1388] The device collects recorded audio data in fixed buffer sizes and sends them to the server in bulk. The input is the audio data being recorded, and the output is the audio data buffer. When the buffer reaches a certain size, it is sent to the server via an HTTP POST request.

[1389] Step 4:

[1390] The server receives the audio data sent from the device. The input is the audio data sent via an HTTP POST request, and the output is the audio data stored on the server. The server temporarily stores this audio data and proceeds to the next step.

[1391] Step 5:

[1392] The server converts the received voice data into text in real time using a generation AI (e.g., DeepSpeech, Google Cloud Speech-to-Text). The input is voice data, and the output is text data. The generation AI analyzes the voice data and generates the corresponding text.

[1393] Step 6:

[1394] The server analyzes the data converted to text by the generative AI and detects fraud-related keywords. The input is text data, and the output is the keyword detection results. A natural language processing tool (e.g., NLTK) is used to search for fraud-related keywords.

[1395] Step 7:

[1396] The server determines whether fraud is suspected based on the keyword detection results. The input is the keyword detection results, and the output is a flag indicating whether fraud is suspected or not. A judgment algorithm (e.g., decision tree, SVM) is executed.

[1397] Step 8:

[1398] If the server determines that fraud is suspected, it uses a generative AI model (e.g., GPT-3) to generate an appropriate automated response. The input is a flag indicating suspected fraud, and the output is the automated response. The generative AI inputs a prompt sentence and generates a corresponding response.

[1399] Step 9:

[1400] The server sends the generated automatic response to the terminal. The input is the automatic response text, and the output is the data to be sent to the terminal. An HTTP POST request is used for sending.

[1401] Step 10:

[1402] The terminal plays back the received automated response lines aloud. The input is the data sent from the server, and the output is the audio data. The TTS (Text-to-Speech) function is used to convert the text data into audio, which is then played back through the speaker.

[1403] Step 11:

[1404] The server stores records of suspected fraudulent phone calls and periodically provides them to relevant authorities (e.g., police). The input is the call's audio data, text data, and analysis results, and the output is the stored recorded data and data provided to relevant authorities. The records are stored in a database and provided to relevant authorities via an API or interface.

[1405] (Application example 1)

[1406] Next, a description will be given of Application Example 1. In the following description, the data processing device 12 will be referred to as a "server" and the robot 414 will be referred to as a "terminal."

[1407] In modern society, the number of fraud cases is increasing, and vulnerable people such as the elderly are often the targets. To solve this problem, a system that can identify fraudulent activities in real time and take prompt countermeasures is needed. However, current systems have difficulty monitoring the content of phone calls one by one, making it difficult to prevent fraudulent activities. Therefore, there is a need to provide a system that can detect fraudulent activities early and respond appropriately.

[1408] The specific processing by the specific processing unit 290 of the data processing device 12 in the application example 1 is realized by the following means.

[1409] In this invention, the server includes means for automatically starting recording when a call is initiated, means for transmitting the voice data of the call, means for converting the received voice data into text using a generating AI, means for detecting fraud-related keywords, means for the AI ​​to automatically take over the conversation if fraud is suspected, means for recording data in real time and providing it to relevant authorities, and means for saving and providing records of suspected fraud. This makes it possible to detect special frauds in real time and take appropriate action.

[1410] "Means for automatically starting recording when a call starts" refers to a function that automatically starts recording audio data as soon as a call starts.

[1411] "Means for transmitting voice data of a call" refers to a function for transmitting recorded voice data to a server in real time or at fixed buffer size intervals.

[1412] "Means of converting received voice data into text using generation AI" refers to the function of converting voice data received by the server into text data using generation AI technology.

[1413] "Means for detecting fraud-related keywords" refers to the function of searching for and detecting specific fraud-related keywords from text data.

[1414] "Means for AI to automatically take over the conversation if fraud is suspected" refers to a function in which AI automatically generates and responds to conversations when it is determined that fraud is suspected.

[1415] "Means of recording data in real time and providing it to relevant authorities" refers to the function of recording the contents of phone calls in real time and providing it to relevant authorities such as the police if necessary.

[1416] "Means for storing and providing records of suspected fraud" refers to the function of storing records of phone calls suspected of fraud and providing them to relevant authorities at a later date if necessary.

[1417] "Means for generating lines for automatic responses using generative AI and playing them back as audio" refers to the function of creating the content of automatic responses using generative AI technology and playing them back as audio.

[1418] "Means of notifying users when suspected fraud is detected" refers to a function that warns or notifies users when suspected fraud is detected.

[1419] This invention relates to a system for preventing special fraud during a call. Each element of this system and its processing will be explained below.

[1420] Terminal behavior:

[1421] The device has a means to automatically start recording when a call is initiated, and the recorded audio data is sent to the server in fixed buffer sizes, allowing the content of the call to be analyzed in real time.

[1422] Server behavior:

[1423] The server converts the received voice data into text using a generative AI model, "Google Cloud Speech-to-Text," which allows for accurate and fast conversion of voice to text.

[1424] The text data is then analyzed to detect fraud-related keywords, such as "bank transfer fraud," "depositing money," and "bank account," and an analysis algorithm is used to detect these.

[1425] If fraud is suspected, the server automatically activates AI to generate an appropriate response using a generative AI model such as OpenAI GPT-4. The generated response is sent to the device and played back as audio.

[1426] Data storage and provision:

[1427] The server stores records of calls suspected of fraud and has the means to provide them to relevant authorities if necessary, so that police and other relevant authorities can obtain the necessary information at a later date.

[1428] Examples:

[1429] Here's a specific scenario:

[1430] 1. User:

[1431] An elderly user receives a suspicious phone call.

[1432] 2. Terminal:

[1433] Recording will start automatically as soon as the call is started.

[1434] The recorded audio data is sent to the server in real time.

[1435] 3. Server:

[1436] The received voice data is converted into text using the generative AI model "Google Cloud Speech-to-Text."

[1437] From the text data, keywords related to fraud, such as "depositing money" and "bank account," are detected.

[1438] If it determines that fraud is suspected, an automatic response message is generated using OpenAI GPT-4.

[1439] For example, generate a response like "Please wait, we're checking."

[1440] 4. Terminal:

[1441] The automatic response message sent from the server is played back and the conversation is taken over.

[1442] 5. Server:

[1443] Records of calls suspected of fraud will be saved and provided to the police and other relevant authorities as necessary.

[1444] Example prompt sentence:

[1445] System: "If fraud is detected, generate an appropriate response."

[1446] User: "I'm from the bank. Please give me your account number so I can deposit some money."

[1447] This will enable vulnerable people, such as the elderly, to make calls safely without falling victim to fraud.

[1448] The flow of the specific processing in the application example 1 will be described with reference to FIG.

[1449] Step 1:

[1450] When a user starts a call, the device automatically starts recording. The input is the call audio, and the output is real-time recorded audio data. Specifically, the device's microphone captures the audio and stores it in a buffer.

[1451] Step 2:

[1452] The device sends recorded audio data to the server in fixed buffer size increments. The input is the recorded audio data, and the output is the buffer data sent to the server. Specifically, the device sends the audio data stored in the buffer to the server in real time.

[1453] Step 3:

[1454] The server converts the received voice data into text using a generative AI model. The input is voice data, and the output is text data. Specifically, it uses "Google Cloud Speech-to-Text" to convert the voice data into text data.

[1455] Step 4:

[1456] The server analyzes the text data and detects fraud-related keywords. The input is text data, and the output is the results of fraud keyword detection. Specifically, it uses an algorithm to search for specific keywords (e.g., "transfer fraud" or "bank account") within the text data.

[1457] Step 5:

[1458] If the server determines that fraud is suspected, it activates a generative AI model to generate an appropriate automated response. The input is the fraud keyword detection results and the conversation text, and the output is the generated automated response text. Specific operations include using "OpenAI GPT-4" to generate a response based on the prompt text. Examples of prompt text include "If fraud is detected, please generate an appropriate response" and "This is from a bank. Please tell me your account number so I can deposit the money."

[1459] Step 6:

[1460] The device receives the automated response sent from the server and plays it back as audio. The input is the generated automated response text, and the output is the automated response in audio. Specifically, the AI ​​response is played back as audio using the device's speaker.

[1461] Step 7:

[1462] The server stores call records suspected of fraud and provides them to relevant authorities as necessary. The input is the call records and fraud detection results, and the output is the stored data and submitted data. Specifically, the server stores call records suspected of fraud in a database and provides them to relevant authorities such as the police.

[1463] This will create a system that can detect fraud in real time and respond appropriately.

[1464] Furthermore, an emotion engine that estimates the user's emotion may be further combined. That is, the identification processing unit 290 may estimate the user's emotion using the emotion identification model 59, and perform identification processing using the user's emotion.

[1465] This invention combines an emotion engine with a system for preventing special fraud during phone calls. Below, each element of this system and its processing will be explained in natural language.

[1466] Start recording a call

[1467] 1. User: Start a call.

[1468] 2. Device: Automatically starts recording when a call is initiated. This means that audio data is recorded from the moment the user starts a call.

[1469] Sending audio data

[1470] 3. Terminal: Recorded audio data is sent to the server at fixed buffer size intervals.

[1471] Speech data conversion and sentiment analysis

[1472] 4. Server: Receives the voice data sent from the device, converts the voice to text using generative AI, and analyzes the user's emotions from the voice data using an emotion engine.

[1473] Analysis of text and emotion data

[1474] 5. Server: Analyzes the text of the conversation. Retrieves fraud-related keywords and similar words from the database, and uses an analysis algorithm to determine whether there is suspicion of fraud. Taking into account the emotional data analyzed by the emotion engine, the server strengthens detection of suspected fraud.

[1475] What to do if fraud is suspected

[1476] 6. Server: If fraud is suspected, the automated AI generates an appropriate response. The response is dynamically adjusted based on the user's emotional data obtained from the emotion engine.

[1477] 7. Device: Plays an automated response and responds on behalf of the user. During this time, the call continues to be recorded and the conversation data continues to be sent to the server.

[1478] Record keeping and provision

[1479] 8. Server: Conversation data suspected of fraud is stored in a database. The data will be provided to relevant agencies such as the police at a later date. The analysis results of the emotion engine will also be stored.

[1480] User Notification

[1481] 9. Server: Notify the user of suspected fraud and sentiment analysis results, if necessary.

[1482] Specific examples

[1483] Here is a specific scenario:

[1484] 1. User: An elderly user receives the call.

[1485] 2. Device: When a call is started, the device automatically starts recording. The recorded audio data is sent to the server in real time.

[1486] 3. Server: The received voice data is converted into text using a generative AI, and then the emotion engine analyzes the user's emotions. Fraud-related keywords such as "depositing money" and "bank account" are detected, and a suspicion of fraud is determined based on the analysis results and the user's emotional data.

[1487] 4. Server: If the server detects a suspected fraud, it will activate an automated AI response, generating a response such as "Please wait a moment, we're checking." The response is dynamically adjusted based on the results of sentiment analysis.

[1488] 5. Terminal: Based on instructions from the server, it plays an automated response and continues to record the conversation.

[1489] 6. Server: After the call ends, the server stores the conversation data and sentiment analysis results of any suspected fraud and periodically provides them to the police.

[1490] This series of steps protects users from specialized fraud and enables more advanced fraud detection and response through the emotion engine, providing an environment where users can use their phones with peace of mind.

[1491] The processing flow will be explained below.

[1492] Step 1:

[1493] User: Start a call.

[1494] Device: Automatically starts recording as soon as the call starts.

[1495] Step 2:

[1496] Terminal: Recorded audio data is sent to the server in real time at fixed buffer size intervals.

[1497] Step 3:

[1498] Server: Receives voice data sent from the device and converts it into text using a generation AI. The converted text data is temporarily stored.

[1499] Step 4:

[1500] Server: At the same time, the emotion engine is used to analyze the user's emotions from the voice data, and the analysis results are temporarily saved.

[1501] Step 5:

[1502] Server: The server compares the text of the conversation with a database of fraud-related keywords to determine whether there is any suspicion of fraud. It also considers the analysis results of the emotion engine to enhance detection of suspected fraud.

[1503] Step 6:

[1504] Server: If fraud is suspected, the automated AI responds by generating an appropriate response. The response is dynamically adjusted based on the user's emotional data obtained from the emotion engine.

[1505] Step 7:

[1506] Terminal: Plays back the automated response sent from the server and responds on behalf of the user. Records during the conversation and continues to send voice data to the server.

[1507] Step 8:

[1508] Server: After the call ends, the conversation data suspected of fraud and the results of sentiment analysis are stored in a database.

[1509] Step 9:

[1510] Server: Provides the relevant data to the police and other relevant authorities on a regular basis. If necessary, notifies the user of suspected fraud or the results of sentiment analysis.

[1511] This series of steps will protect users from specialized fraud, enable advanced fraud detection and response using an emotion engine, and provide recorded data to police to help prevent future crimes.

[1512] Example 2

[1513] Next, a description will be given of Example 2. In the following description, the data processing device 12 will be referred to as a "server" and the robot 414 will be referred to as a "terminal."

[1514] Currently, special frauds over the phone are on the rise, and elderly people are particularly vulnerable to these crimes. Conventional countermeasures have limitations, and more effective systems are needed to detect suspected fraud in real time and protect users. It is also important to ensure that records of fraudulent conversations are stored and provided to the necessary authorities.

[1515] The specific processing by the specific processing unit 290 of the data processing device 12 in the second embodiment is realized by the following means.

[1516] In this invention, the server includes a means for converting the voice data of the call into text using a generation AI, a means for analyzing emotions from the converted text, and a means for detecting fraud-related keywords and determining suspicion of fraud based on the emotion analysis results. This makes it possible to analyze the content of the call in real time and strengthen suspicion of fraud based on the emotion data.

[1517] The "means for starting recording" is a device or program that has the function of automatically starting recording of voice data as soon as a call is started.

[1518] The "means for transmitting audio data" is a device or program for transmitting recorded audio data to a server in real time with a certain buffer size.

[1519] "Means for converting audio data into text using generative AI" refers to an artificial intelligence algorithm or software for automatically converting audio data into text data.

[1520] "Means for analyzing emotions" refers to software or algorithms that analyze the user's emotional state from textual data and obtain the results.

[1521] A "means for detecting fraud-related keywords" is a program or algorithm for extracting specific words or phrases related to fraudulent activities from text data.

[1522] "Means for determining suspicion of fraud by taking into account the results of sentiment analysis" refers to software or algorithms that have the function of comprehensively evaluating and determining the possibility of fraud based on the results of sentiment analysis data and keywords related to fraud.

[1523] "Means for automatic AI conversation takeover" refers to a program or device that allows artificial intelligence to automatically continue a conversation on behalf of a user when a suspected fraud is determined.

[1524] "Means for adjusting the content of the automated response based on the results of sentiment analysis" refers to software or algorithms for appropriately changing the content and tone of the automated response message based on the results of sentiment analysis.

[1525] The "means for playing the message aloud" refers to a device or software for converting the generated automated response message into aloud voice and playing it back to the other party on the call.

[1526] The "means of storing and providing records of suspected fraud" refers to a system for storing call records and analysis data of suspected fraud in a database and providing them to relevant authorities as necessary.

[1527] This invention relates to a system for preventing special fraud during phone calls, which operates by combining a generative AI and an emotion engine. This system is realized using the following specific hardware and software.

[1528] Hardware and Software Configuration

[1529] server

[1530] 1. Generative AI model (e.g., Google Cloud's Speech-to-Text API): Converts voice data into text data.

[1531] 2. Sentiment engine (e.g., Microsoft Azure's Text Analytics API): Analyzes emotions from text data.

[1532] 3. Automated response AI (e.g., OpenAI’s GPT-3): Generates appropriate responses in cases of suspected fraud.

[1533] Terminal

[1534] 1. Calling app (e.g. Skype, Google Voice): An application for making calls.

[1535] 2. Recording function (e.g., iOS's Call Recorder app): A function that automatically records phone calls.

[1536] 3. Text-to-speech API (e.g., Amazon Polly): Speech the generated text response.

[1537] Overview of program processing

[1538] Start recording a call

[1539] When a user makes or receives a call using a calling app, the device automatically activates the recording function and records the audio data. This audio data is sent to the server in real time at fixed buffer size intervals.

[1540] Speech data conversion and sentiment analysis

[1541] The server uses a generative AI model to convert the voice data into text, which is then analyzed by an emotion engine to determine the user's emotional state. For example, emotions such as "happy," "sad," and "angry" are detected.

[1542] Determine suspected fraud

[1543] Based on the textual content of the conversation and the emotional data, the server detects keywords related to fraud. For example, it scans for words such as "money," "bank," and "account." At the same time, it also takes into account the results of the emotional analysis and makes a comprehensive judgment on the suspicion of fraud. For example, if the phrase "depositing money" matches the emotional data of "anxiety," it is determined that there is a high suspicion of fraud.

[1544] Auto-response generation and playback

[1545] If suspected fraud is detected, the server activates an automated AI to generate an appropriate response. This response is adjusted based on the results of sentiment analysis. The generated response is sent in text format to the device, which then uses a speech synthesis API to convert it into audio and play it back to the other party. For example, the response could be, "Please wait a moment, we're checking."

[1546] Record keeping and provision

[1547] Conversation data and sentiment analysis results for suspected fraud are stored in a secure database on a server. This data can be provided to relevant authorities as needed, such as the police, to enable further investigation and countermeasures.

[1548] User Notifications

[1549] If a suspected fraud is detected, the server will notify the user via email and / or SMS.

[1550] Examples of concrete examples and prompts

[1551] Specific examples

[1552] When an elderly user answers the phone and the call begins, the device automatically starts recording. This audio data is sent to a server in real time and converted into text by a generative AI model. If the phrase "depositing money" and the emotion "anxiety" are detected, the server activates an automated response AI and generates a response saying, "Please wait a moment, we're checking." This response is then converted into audio and played back to the caller.

[1553] Prompt Sentence Examples

[1554] The voice data is converted into text and analyzed, and if it contains keywords such as "money" or "bank," it is judged to be suspected of fraud. For example, if the phrase "depositing money" is included, it is considered highly suspected of fraud.

[1555] This effectively protects users from specialized fraud while on the phone, while the emotion engine enables more advanced fraud detection and response.

[1556] The flow of the identification process in the second embodiment will be described with reference to FIG.

[1557] Processing Steps

[1558] Step 1: Start recording the call

[1559] User: A user initiates or receives a call using a calling app.

[1560] Terminal: When a call is initiated, the recording function is automatically activated and audio data is recorded.

[1561] Input: Call Start Trigger

[1562] Output: Recorded audio data

[1563] Specific operation: When the device's calling app detects the start of a call, the recording function is activated and the call audio from that point on is saved as digital data.

[1564] Step 2: Sending audio data

[1565] Terminal: Recorded audio data is sent to the server in real time at fixed buffer size intervals.

[1566] Input: Recorded audio data

[1567] Output: Audio data sent to the server

[1568] What it does: The audio data is split into one-minute buffers and sent to the server using an HTTP POST request.

[1569] Step 3: Converting audio data to text

[1570] Server: Uses a generative AI model to convert voice data into text data.

[1571] Input: Audio data received from the device

[1572] Output: Text data

[1573] Specific operation: After receiving the voice data, the server calls Google Cloud's Speech-to-Text API to convert the voice to text.

[1574] Step 4: Sentiment Analysis

[1575] Server: Analyzes emotions from text data using an emotion engine.

[1576] Input: Text data

[1577] Output: Emotion analysis results

[1578] How it works: The text data is sent to Microsoft Azure's Text Analytics API, where it is analyzed to see if it contains emotions such as "happiness," "sadness," or "anger."

[1579] Step 5: Determine if there is any fraud

[1580] Server: Based on the text data and sentiment analysis results, detects fraud-related keywords and determines whether fraud is suspected.

[1581] Input: Text data, sentiment analysis results

[1582] Output: Suspected fraud determination result

[1583] How it works: The server scans the text data and compares it with a list of keywords (e.g., "money," "bank," "account"), while also taking into account the results of sentiment analysis to make a comprehensive judgment on the suspicion of fraud.

[1584] Step 6: Generate an autoresponder

[1585] Server: If fraud is suspected, an automated AI is used to generate an appropriate response.

[1586] Input: Suspected fraud judgment result

[1587] Output: Auto-response text

[1588] Specific operation: The server calls OpenAI's GPT-3 and generates a response message based on the situation, such as "Please wait a moment, checking."

[1589] Step 7: Play Auto Attendant Audio

[1590] Terminal: The text of the automated response is converted into speech using a speech synthesis API and played back to the other party.

[1591] Input: Auto-response text

[1592] Output: Voiced auto attendant

[1593] How it works: The device uses a speech synthesis API such as Amazon Polly to convert the generated text into speech and play it back to the other party in real time.

[1594] Step 8: Record keeping and provision

[1595] Server: Stores suspected fraudulent conversation data and sentiment analysis results in a database and provides them to relevant authorities as needed.

[1596] Input: Suspected fraud voice data, emotion analysis results

[1597] Output: Data stored, data provided

[1598] What it does: The server stores suspected fraudulent conversation data in a secure database and provides it to police and other relevant authorities as needed.

[1599] Step 9: User Notification

[1600] Server: Notifies the user if suspected fraud is detected.

[1601] Input: Suspected fraud judgment result

[1602] Output: User notification

[1603] What it does: The server notifies the user of suspected fraud via email address or mobile phone number. This notification is sent via email using an SMTP server or as a text message using an SMS API.

[1604] Through the above steps, the system can detect suspected special fraud in real time during a call and effectively protect users.

[1605] (Application example 2)

[1606] Next, a description will be given of Application Example 2. In the following description, the data processing device 12 will be referred to as a "server" and the robot 414 will be referred to as a "terminal."

[1607] Conventional call recording systems could record phone conversations and, in some cases, detect keywords that could be suspected of fraud. However, they lacked the ability to analyze user sentiment and use that information to further improve the accuracy of fraud detection. As a result, even if fraud was suspected, it was not possible to evaluate whether the user actually perceived the risk, limiting the accuracy of fraud detection. Furthermore, they lacked the ability to provide appropriate automated responses, making it difficult to respond quickly in situations where there was a high risk of actually being scammed. The challenge is to solve these problems and provide a more advanced fraud prevention system.

[1608] The identification processing by the identification processing unit 290 of the data processing device 12 in Application Example 2 is realized by the following means. In this invention, the server includes means for transmitting voice data of the call, means for converting the received voice data into text using a generation AI, means for detecting fraud-related keywords, means for analyzing user emotions from the voice data using an emotion engine, means for dynamically adjusting automatic responses based on the emotion data, and means for saving and providing records of suspected fraud. This enables advanced fraud detection that takes into account user emotion analysis and the provision of appropriate and dynamically adjusted automatic responses.

[1609] A "call" is an act of exchanging voice information between two people using a telephone line.

[1610] "Recording" is the act of recording sound in digital or analog form.

[1611] "Generative AI" is a system that uses artificial intelligence technology to generate and convert voice and text.

[1612] "Text conversion" is the process of converting audio data into text data.

[1613] "Fraud Keywords" are words or phrases that are likely to be associated with fraudulent activity.

[1614] An "emotion engine" is a system that analyzes user emotions from voice and text data.

[1615] An "auto-response" is a message that the system automatically generates in response to a user.

[1616] "Dynamic adjustment" refers to changing the response in real time based on specific situations and conditions.

[1617] "Record keeping" is the act of storing data in a file or database.

[1618] "Notification" means the act of informing a user of specific information.

[1619] This invention is a system for preventing special fraud during phone calls, and aims to further improve accuracy by analyzing the user's emotions using an emotion engine.

[1620] The system is roughly configured as follows:

[1621] Hardware and Software Configuration

[1622] 1. Device:

[1623] A communication device such as a smartphone that has the ability to record voice data during calls and send it to a server.

[1624] 2. Server:

[1625] The server converts the received voice data into text using a generative AI model, detects fraud-related keywords, and analyzes the user's emotions using an emotion engine.

[1626] Program Overview

[1627] The system operates in the following steps:

[1628] Call Recording:

[1629] When a user starts a call, the device automatically starts recording, and the recorded audio data is sent to the server in real time.

[1630] Speech to text conversion and sentiment analysis:

[1631] The server converts the received voice data into text using a generative AI, and then uses an emotion engine to analyze emotions from the voice data. Specifically, the server uses the speech_recognition library and the transformers library.

[1632] Suspected fraud detection:

[1633] The server analyzes the converted text data to detect fraud-related keywords and, taking into account the results of the emotion engine analysis, determines whether or not there is any suspicion of fraud.

[1634] Generate an auto-response:

[1635] If fraud is suspected, the server will activate an automated AI response that generates a dynamically tailored response based on the results of sentiment analysis. This automated response is played back via voice and acts on behalf of the user.

[1636] Record-Keeping and Notification:

[1637] During and after the call, the server stores the records of suspected fraud and the results of the sentiment analysis. If necessary, the server periodically provides the data to the relevant authorities and notifies the user of the suspected fraud.

[1638] Specific examples

[1639] For example, if an elderly person receives a phone call and is told, "This is confirmation from the bank," the device automatically records the call and sends the audio data to a server. The server converts the audio data into text and analyzes the user's emotions using an emotion engine. If a fraud keyword such as "depositing money" is detected and the analysis finds that the elderly person is anxious, the server generates an automatic response such as "Are you okay? We're checking," and plays it back.

[1640] Prompt Sentence Examples

[1641] Prompt: Convert the speech data into text and detect whether it contains fraud-related keywords. Also, perform sentiment analysis and generate appropriate responses based on that.

[1642] In this way, it provides an advanced prevention system to prevent users from falling victim to fraud during calls.

[1643] The flow of the specific processing in the application example 2 will be described with reference to FIG.

[1644] Step 1:

[1645] Start recording a call

[1646] Input: The user receives a phone call.

[1647] How it works: When you start a call, your device automatically starts recording the call.

[1648] Output: The recorded audio data is generated.

[1649] Step 2:

[1650] Sending audio data

[1651] Input: Recorded audio data.

[1652] How it works: The device sends the recorded audio data to the server in real time.

[1653] Output: Audio data transferred to the server.

[1654] Step 3:

[1655] Converting audio data to text

[1656] Input: The audio data sent to the server.

[1657] How it works: The server uses a generative AI model (e.g., the speech_recognition library) to convert the audio data to text.

[1658] Output: Text data is generated.

[1659] Step 4:

[1660] Emotion analysis

[1661] Input: The text data generated in step 3.

[1662] How it works: The server uses an emotion engine (e.g., the sentiment-analysis model from the transformers library) to analyze user sentiment from text data.

[1663] Output: Emotion data is generated.

[1664] Step 5:

[1665] Fraudulent Keyword Detection

[1666] Input: Text data.

[1667] How it works: The server scans the text data using a pre-defined list of keywords (e.g., "deposit money," "bank account," etc.) to detect fraud-related keywords in the text data.

[1668] Output: The result whether fraud keywords were detected or not.

[1669] Step 6:

[1670] Determining suspected fraud

[1671] Input: Sentiment data and fraud keyword detection results.

[1672] Operation: The server makes a comprehensive judgment on whether there is suspicion of fraud based on the emotional data and the fraud keyword detection results.

[1673] Output: Judgment result on whether fraud is suspected.

[1674] Step 7:

[1675] Generate an auto-response

[1676] Input: Suspected fraud judgment results and sentiment data.

[1677] How it works: If fraud is suspected, the server uses a generative AI model to generate an automated response that dynamically adjusts based on sentiment data.

[1678] Output: A well-tuned auto-response.

[1679] Step 8:

[1680] Playing an Auto Attendant

[1681] Input: The generated auto-response.

[1682] How it works: The generated automated response is output as audio and played back to the user through the device.

[1683] Output: The audio response played to the user.

[1684] Step 9:

[1685] Record keeping and provision

[1686] Input: call content, fraud detection results, emotional data.

[1687] How it works: The server stores suspected fraudulent calls and their analysis results in a database, and provides them to relevant authorities as needed.

[1688] Output: Stored records and provided data.

[1689] Step 10:

[1690] User Notification

[1691] Input: Fraud determination result.

[1692] How it works: The server generates a warning message and sends it to the user via their device to inform them of the suspected fraud.

[1693] Output: The warning message sent to the user.

[1694] These steps will enable an advanced prevention system that analyzes call content in real time and protects users from the risk of special fraud.

[1695] The specific processing unit 290 transmits the result of the specific processing to the robot 414. In the robot 414, the control unit 46A causes the speaker 240 and the control target 443 to output the result of the specific processing. The microphone 238 acquires voice indicating a user input regarding the result of the specific processing. The control unit 46A transmits voice data indicating the user input acquired by the microphone 238 to the data processing device 12. In the data processing device 12, the specific processing unit 290 acquires the voice data.

[1696] The data generation model 58 is a so-called generative AI (Artificial Intelligence). An example of the data generation model 58 is ChatGPT (Internet Search<URL: https: / / openai.com / blog / chatgpt> ), Gemini (Internet search <url: https: gemini.google.com ?hl="ja">) and other generation AIs. The data generation model 58 is obtained by performing deep learning on a neural network. A prompt including an instruction is input to the data generation model 58, and inference data such as voice data indicating voice, text data indicating text, and image data indicating an image is also input. The data generation model 58 performs inference on the input inference data in accordance with the instruction indicated by the prompt, and outputs the inference result in a data format such as voice data and text data. Here, inference refers to, for example, analysis, classification, prediction, and / or summarization.

[1697] In the above embodiment, an example was given in which the specific processing is performed by the data processing device 12, but the technology of the present disclosure is not limited to this, and the specific processing may be performed by the robot 414.

[1698] The emotion identification model 59 as an emotion engine may determine the user's emotion according to a specific mapping. Specifically, the emotion identification model 59 may determine the user's emotion according to an emotion map (see FIG. 9), which is a specific mapping. Similarly, the emotion identification model 59 may determine the robot's emotion, and the identification processing unit 290 may perform identification processing using the robot's emotion.

[1699] FIG. 9 is a diagram illustrating an emotion map 400 on which multiple emotions are mapped. In the emotion map 400, emotions are arranged in concentric circles radiating from the center. Emotions closer to the center of the concentric circles are more primitive. Emotions representing states and actions arising from a state of mind are arranged on the outer edges of the concentric circles. The concept of emotion includes both affect and mental states. Emotions generally generated from reactions occurring in the brain are arranged on the left side of the concentric circles. Emotions generally induced by situational judgment are arranged on the right side of the concentric circles. Emotions generally generated from reactions occurring in the brain and induced by situational judgment are arranged on the upper and lower sides of the concentric circles. Furthermore, the emotion of "pleasure" is arranged on the upper side of the concentric circles, and the emotion of "discomfort" is arranged on the lower side. In this way, in the emotion map 400, multiple emotions are mapped based on the structure by which emotions are generated, and emotions that tend to occur simultaneously are mapped close to each other.

[1700] These emotions are distributed in the 3 o'clock direction on emotion map 400, and typically fluctuate between relief and anxiety. In the right half of emotion map 400, situational awareness dominates over internal sensations, resulting in a sense of calm.

[1701] The inside of emotion map 400 represents what is going on in the mind, and the outside of emotion map 400 represents behavior, so the further you go outside emotion map 400, the more visible the emotions become (the more they are expressed in behavior).

[1702] Human emotions are based on various balances, such as posture and blood sugar levels. When these balances deviate from the ideal, a state of discomfort is indicated, and when they approach the ideal, a state of pleasure is indicated. Emotions can also be created for robots, automobiles, and motorcycles, based on various balances, such as posture and remaining battery life. When these balances deviate from the ideal, a state of discomfort is indicated, and when they approach the ideal, a state of pleasure is indicated. An emotion map can be generated, for example, based on Dr. Mitsuyoshi's emotion map (Research on Voice Emotion Recognition and Emotional Brain Physiological Signal Analysis Systems, Tokushima University, Doctoral Dissertation: https: / / ci.nii.ac.jp / naid / 500000375379). The left half of the emotion map lists emotions belonging to the "reaction" domain, where sensation is dominant. The right half of the emotion map lists emotions belonging to the "situation" domain, where situational awareness is dominant.

[1703] The emotion map defines two emotions that promote learning. One is a negative emotion on the situation side, around the middle of "repentance" or "reflection." In other words, this occurs when the robot experiences negative emotions such as "I never want to feel this way again" or "I don't want to be scolded again." The other is a positive emotion on the response side, around "desire." In other words, this occurs when the robot experiences positive feelings such as "I want more" or "I want to know more."

[1704] The emotion identification model 59 inputs user input into a pre-trained neural network, obtains emotion values ​​indicating each emotion shown in the emotion map 400, and determines the user's emotion. This neural network is pre-trained based on multiple pieces of training data that are combinations of user input and emotion values ​​indicating each emotion shown in the emotion map 400. Furthermore, this neural network is trained so that emotions that are located close to each other have similar values, as in the emotion map 900 shown in FIG. 10. FIG. 10 shows an example in which multiple emotions, "relieved," "calm," and "reassuring," have similar emotion values.

[1705] The system according to the present disclosure has been described above mainly with respect to the functions of the data processing device 12, but the system according to the present disclosure is not necessarily implemented on a server. The system according to the present disclosure may be implemented as a general information processing system. The present disclosure may be implemented, for example, as a software program running on a personal computer or an application running on a smartphone, etc. The method according to the present disclosure may be provided to users in the form of SaaS (Software as a Service).

[1706] In the above embodiment, an example was given in which the specific processing is performed by one computer 22, but the technology of the present disclosure is not limited to this, and the specific processing may be distributed and performed by a plurality of computers including the computer 22. For example, the data generation model 58 may be provided in an external device of the data processing device 12, and data may be generated in the external device in accordance with input data.

[1707] In the above embodiment, an example in which the specific processing program 56 is stored in the storage 32 has been described, but the technology of the present disclosure is not limited to this. For example, the specific processing program 56 may be stored in a portable, computer-readable, non-transitory storage medium such as a USB (Universal Serial Bus) memory. The specific processing program 56 stored in the non-transitory storage medium is installed in the computer 22 of the data processing device 12. The processor 28 executes the specific processing in accordance with the specific processing program 56.

[1708] Alternatively, the specific processing program 56 may be stored in a storage device such as a server connected to the data processing device 12 via the network 54, and the specific processing program 56 may be downloaded and installed on the computer 22 in response to a request from the data processing device 12.

[1709] It is not necessary to store all of the specific processing program 56 in a storage device such as a server connected to the data processing device 12 via the network 54, or to store all of the specific processing program 56 in the storage 32; only a portion of the specific processing program 56 may be stored.

[1710] The hardware resource for executing a specific process can be any of the following processors: An example of a processor is a CPU, which is a general-purpose processor that functions as a hardware resource for executing a specific process by executing software, i.e., a program. Another example of a processor is a dedicated electrical circuit, such as an FPGA (Field-Programmable Gate Array), a PLD (Programmable Logic Device), or an ASIC (Application Specific Integrated Circuit), which is a processor with a circuit configuration designed specifically for executing a specific process. Each processor has built-in or connected memory, and each processor uses the memory to execute the specific process.

[1711] The hardware resource that executes the specific processing may be configured with one of these various processors, or may be configured with a combination of two or more processors of the same or different types (for example, a combination of multiple FPGAs, or a combination of a CPU and an FPGA). Also, the hardware resource that executes the specific processing may be a single processor.

[1712] As an example of a system configured with a single processor, first, one processor is configured by combining one or more CPUs and software, and this processor functions as a hardware resource that executes a specific process. Second, there is a system that uses a processor that realizes the functions of an entire system including multiple hardware resources that execute a specific process on a single IC chip, as typified by SoC (System-on-a-chip). In this way, a specific process is realized using one or more of the above-mentioned various processors as hardware resources.

[1713] Furthermore, the hardware structure of these various processors can be, more specifically, an electric circuit that combines circuit elements such as semiconductor devices. The specific processing described above is merely an example. Therefore, it goes without saying that unnecessary steps may be deleted, new steps may be added, or the processing order may be rearranged, without departing from the spirit of the invention.

[1714] The above-described description and illustrations are a detailed explanation of the parts related to the technology of the present disclosure and are merely an example of the technology of the present disclosure. For example, the above description of the configuration, functions, actions, and effects is an explanation of an example of the configuration, functions, actions, and effects of the parts related to the technology of the present disclosure. Therefore, it goes without saying that unnecessary parts may be deleted, new elements may be added, or replacements may be made to the above-described description and illustrations within the scope of the gist of the technology of the present disclosure. Furthermore, to avoid confusion and facilitate understanding of the parts related to the technology of the present disclosure, the above-described description and illustrations omit explanations of common technical knowledge that do not require particular explanation to enable the implementation of the technology of the present disclosure.

[1715] All publications, patent applications, and technical standards mentioned in this specification are herein incorporated by reference to the same extent as if each individual publication, patent application, or technical standard was specifically and individually indicated to be incorporated by reference.

[1716] The following is further disclosed regarding the above embodiment.

[1717] (Claim 1)

[1718] A way to automatically start recording when a call starts,

[1719] means for transmitting voice data of the call;

[1720] A means for converting received voice data into text using a generation AI;

[1721] a means for detecting keywords related to fraud;

[1722] A way for AI to automatically take over the conversation if fraud is suspected, and

[1723] A system that includes a means to keep and provide records of suspected fraud.

[1724] (Claim 2)

[1725] The system according to claim 1, characterized in that it has a means for generating lines for an automatic response using a generation AI and playing them back aloud.

[1726] (Claim 3)

[1727] 2. The system of claim 1, further comprising means for notifying a user if suspected fraud is detected.

[1728] "Example 1"

[1729] (Claim 1)

[1730] A way to automatically start recording when a call starts,

[1731] means for transmitting audio data of a call in fixed buffer size increments;

[1732] A means to convert received voice data into text in real time using generation AI,

[1733] a means for retrieving and detecting fraud-related keywords from a database;

[1734] If fraud is suspected, an automated AI can take over the conversation and generate and send an appropriate response.

[1735] A system that includes a means to keep records of suspected fraud and provide them to the appropriate authorities.

[1736] (Claim 2)

[1737] The system according to claim 1, characterized in that it has a means for generating lines for an automatic response using a generation AI and playing them back aloud.

[1738] (Claim 3)

[1739] 2. The system of claim 1, further comprising means for notifying a user if suspected fraud is detected.

[1740] "Application Example 1"

[1741] (Claim 1)

[1742] A way to automatically start recording when a call starts,

[1743] means for transmitting voice data of the call;

[1744] A means for converting received voice data into text using a generation AI;

[1745] a means for detecting keywords related to fraud;

[1746] A way for AI to automatically take over the conversation if fraud is suspected, and

[1747] A means of recording data in real time and providing it to relevant authorities;

[1748] A system that includes a means to keep and provide records of suspected fraud.

[1749] (Claim 2)

[1750] The system according to claim 1, characterized in that it has a means for generating lines for an automatic response using a generation AI and playing them back aloud.

[1751] (Claim 3)

[1752] 2. The system of claim 1, further comprising means for notifying a user if suspected fraud is detected.

[1753] "Example 2: Combining Emotion Engines"

[1754] (Claim 1)

[1755] A way to automatically start recording when a call starts,

[1756] means for transmitting voice data of the call;

[1757] A means for converting received voice data into text using a generation AI;

[1758] A means of analyzing emotions from text data,

[1759] a means for detecting fraud-related keywords and determining suspected fraud based on the results of sentiment analysis;

[1760] A way for AI to automatically take over the conversation if fraud is suspected, and

[1761] A means for adjusting the content of the automated response based on the emotion analysis result and playing it back as audio;

[1762] A system that includes a means to keep and provide records of suspected fraud.

[1763] (Claim 2)

[1764] 2. The system according to claim 1, further comprising means for generating and audibly reproducing automatic response lines.

[1765] (Claim 3)

[1766] 10. The system of claim 1, further comprising means for notifying the user if suspected fraud is detected.

[1767] "Application example 2 when combining emotion engines"

[1768] (Claim 1)

[1769] A way to automatically start recording when a call starts,

[1770] means for transmitting voice data of the call;

[1771] A means for converting received voice data into text using a generation AI;

[1772] a means for detecting keywords related to fraud;

[1773] A way for AI to automatically take over the conversation if fraud is suspected, and

[1774] the means to maintain and provide records of suspected fraud;

[1775] a means for analyzing user emotions from the voice data using an emotion engine;

[1776] A system including means for dynamically adjusting automated responses based on emotion data.

[1777] (Claim 2)

[1778] The system according to claim 1, characterized in that it has a means for generating automatic response lines using a generation AI and playing them back in audio, and dynamically adjusts the response based on the results of emotion analysis.

[1779] (Claim 3)

[1780] 2. The system of claim 1, further comprising means for notifying the user if suspected fraud is detected, and for providing a record of suspected fraud generated to relevant authorities on a regular basis. [Explanation of symbols]

[1781] 10, 210, 310, 410 Data Processing Systems 12 Data Processing Device 14 Smart Devices 214 Smart Glasses 314 Headset-type terminal 414 Robot< / url:> < / url:> < / url:> < / url:>

Claims

1. A way to automatically start recording when a call starts, means for transmitting voice data of the call; A means for converting received voice data into text using a generation AI; a means for detecting keywords related to fraud; A way for AI to automatically take over the conversation if fraud is suspected, and A system that includes a means to keep and provide records of suspected fraud.

2. 2. The system according to claim 1, further comprising means for generating lines for an automatic response by a generation AI and reproducing the lines by voice.

3. 10. The system of claim 1, further comprising means for notifying a user if suspected fraud is detected.

Citation Information

Patent Citations

  • Persona chatbot control method and system

    JP2022180282A