system
A system with generative AI analyzes transcribed conversation data to detect fraud patterns and issue alerts, addressing the challenge of recognizing scams in real-time, thereby preventing financial loss.
Patent Information
- Application Number
- JP2024161872
- Authority / Receiving Office
- JP · JP
- Patent Type
- Patents
- Current Assignee / Owner
- Priority Date
- 2023-09-19
- Filing Date
- 2024-09-19
- Publication Date
- 2025-12-01
- Estimated Expiration
- 2044-09-19
AI Technical Summary
Victims of telephone fraud often fail to recognize scams in time, leading to significant losses due to sophisticated fraud methods, especially targeting the elderly.
A system that uses a generative AI to analyze transcribed conversation data from a smartphone or home phone microphone, detecting suspicious fraud patterns and issuing an audio alert when fraud is suspected.
Enables early awareness of potential fraud, allowing users to take preventive measures before significant damage occurs.
Smart Images

Figure 0007778201000001 
Figure 0007778201000002 
Figure 0007778201000003
Abstract
Description
[Technical Field]
[0001] The technology of the present disclosure relates to a system. [Background technology]
[0002] Patent document 1 discloses a persona chatbot control method performed by at least one processor, the method including the steps of receiving a user utterance, adding the user utterance to a prompt including an instruction sentence related to a description of the chatbot character, encoding the prompt, and inputting the encoded prompt into a language model to generate a chatbot utterance in response to the user utterance. [Prior art documents] [Patent documents]
[0003] [Patent Document 1] Japanese Patent Publication No. 2022-180282 Summary of the Invention [Problem to be solved by the invention]
[0004] Special frauds have become a social problem, and it is often difficult for victims to realize that they are being scammed. In particular, in cases of telephone fraud, it can take time for victims to realize that they are being scammed, and in the meantime, they can suffer significant losses. [Means for solving the problem]
[0005] This invention detects incoming fraudulent calls and listens to the conversation through a smartphone application or a microphone installed on a home phone. The conversation data is transcribed and sent to a generative AI. The generative AI determines from the conversation data that fraud is suspected, and if fraud is suspected, issues an audio alert from the microphone. This allows users to become aware of the possibility of fraud early on and prevent damage. [Brief explanation of the drawings]
[0006] [Figure 1] 1 is a conceptual diagram showing an example of the configuration of a data processing system according to a first embodiment. [Figure 2] 1 is a conceptual diagram showing an example of main functions of a data processing device and a smart device according to a first embodiment. [Figure 3] FIG. 10 is a conceptual diagram showing an example of the configuration of a data processing system according to a second embodiment. [Figure 4] FIG. 10 is a conceptual diagram showing an example of main functions of a data processing device and smart glasses according to a second embodiment. [Figure 5] FIG. 10 is a conceptual diagram showing an example of the configuration of a data processing system according to a third embodiment. [Figure 6] FIG. 11 is a conceptual diagram showing an example of main functions of a data processing device and a headset-type terminal according to a third embodiment. [Figure 7] FIG. 10 is a conceptual diagram showing an example of the configuration of a data processing system according to a fourth embodiment. [Figure 8] FIG. 10 is a conceptual diagram showing an example of main functions of a data processing device and a robot according to a fourth embodiment. [Figure 9] 1 shows an emotion map onto which multiple emotions are mapped. [Figure 10] 1 shows an emotion map onto which multiple emotions are mapped. [Figure 11] FIG. 2 is a sequence diagram showing a flow of processing in the data processing system according to the first embodiment of the first form example. [Figure 12] FIG. 10 is a sequence diagram showing the flow of processing in the data processing system in Application Example 1 of Embodiment 1. [Figure 13] FIG. 10 is a sequence diagram showing a processing flow of a data processing system in a second embodiment of the second form example. [Figure 14] FIG. 10 is a sequence diagram showing the flow of processing in the data processing system in Application Example 2 of Embodiment 2. [Figure 15] FIG. 10 is a sequence diagram showing the flow of processing in a data processing system according to a third embodiment of the third embodiment. [Figure 16] FIG. 10 is a sequence diagram showing the flow of processing in the data processing system in Application Example 3 of Embodiment 3. [Figure 17] FIG. 10 is a sequence diagram showing the flow of processing in the data processing system in the first embodiment of the first form example when an emotion engine is combined. [Figure 18] FIG. 10 is a sequence diagram showing the flow of processing in the data processing system in Application Example 1 of Form Example 1 when an emotion engine is combined. [Figure 19] FIG. 10 is a sequence diagram showing the flow of processing in the data processing system in the second embodiment of the second form example when an emotion engine is combined. [Figure 20] FIG. 10 is a sequence diagram showing the flow of processing in the data processing system in Application Example 2 of Form Example 2 when an emotion engine is combined. [Figure 21] FIG. 10 is a sequence diagram showing the flow of processing in the data processing system in the third embodiment of the third form example when an emotion engine is combined. [Figure 22] FIG. 10 is a sequence diagram showing the flow of processing in the data processing system in Application Example 3 of Form Example 3 when an emotion engine is combined. DETAILED DESCRIPTION OF THE INVENTION
[0007] An example of an embodiment of a system according to the technology of the present disclosure will be described below with reference to the accompanying drawings.
[0008] First, the terms used in the following description will be explained.
[0009] In the following embodiments, a coded processor (hereinafter simply referred to as a "processor") may be a single arithmetic device or a combination of multiple arithmetic devices. Furthermore, the processor may be a single type of arithmetic device or a combination of multiple types of arithmetic devices. Examples of arithmetic devices include a CPU (Central Processing Unit), a GPU (Graphics Processing Unit), a GPGPU (General-Purpose computing on Graphics Processing Units), an APU (Accelerated Processing Unit), or a TPU (TENSOR PROCESSING UNIT (registered trademark)).
[0010] In the following embodiments, a coded RAM (Random Access Memory) is a memory in which information is temporarily stored and is used as a working memory by a processor.
[0011] In the following embodiments, the coded storage is one or more non-volatile storage devices that store various programs, various parameters, etc. Examples of non-volatile storage devices include flash memory (SSD (Solid State Drive)), magnetic disks (e.g., hard disks), and magnetic tapes.
[0012] In the following embodiments, a communication I / F (Interface) with a symbol is an interface including a communication processor, an antenna, etc. The communication I / F controls communication between multiple computers. Examples of communication standards applied to the communication I / F include wireless communication standards including 5G (5th Generation Mobile Communication System), Wi-Fi (registered trademark), Bluetooth (registered trademark), etc.
[0013] In the following embodiments, "A and / or B" is synonymous with "at least one of A and B." In other words, "A and / or B" means that it may be only A, only B, or a combination of A and B. Furthermore, in this specification, the same concept as "A and / or B" is also applied when three or more things are expressed connected by "and / or."
[0014] [First embodiment]
[0015] FIG. 1 shows an example of the configuration of a data processing system 10 according to the first embodiment.
[0016] 1, a data processing system 10 includes a data processing device 12 and a smart device 14. An example of the data processing device 12 is a server.
[0017] The data processing device 12 includes a computer 22, a database 24, and a communication I / F 26. The computer 22 is an example of a "computer" according to the technology of the present disclosure. The computer 22 includes a processor 28, a RAM 30, and a storage 32. The processor 28, the RAM 30, and the storage 32 are connected to a bus 34. The database 24 and the communication I / F 26 are also connected to the bus 34. The communication I / F 26 is connected to a network 54. Examples of the network 54 include a WAN (Wide Area Network) and / or a LAN (Local Area Network).
[0018] The smart device 14 includes a computer 36, a reception device 38, an output device 40, a camera 42, and a communication I / F 44. The computer 36 includes a processor 46, a RAM 48, and a storage 50. The processor 46, the RAM 48, and the storage 50 are connected to a bus 52. The reception device 38, the output device 40, and the camera 42 are also connected to the bus 52.
[0019] The reception device 38 includes a touch panel 38A, a microphone 38B, and the like, and receives user input. The touch panel 38A detects contact with an indicator (for example, a pen or a finger) to receive user input by the touch of the indicator. The microphone 38B detects the user's voice to receive user input by voice. The control unit 46A transmits data indicating the user input received by the touch panel 38A and the microphone 38B to the data processing device 12. In the data processing device 12, the specific processing unit 290 acquires the data indicating the user input.
[0020] The output device 40 includes a display 40A and a speaker 40B, and presents data to the user 20 by outputting the data in a form of expression that the user 20 can perceive (for example, audio and / or text). The display 40A displays visible information such as text and images in accordance with instructions from the processor 46. The speaker 40B outputs audio in accordance with instructions from the processor 46. The camera 42 is a compact digital camera equipped with an optical system including a lens, aperture, and shutter, and an imaging element such as a CMOS (Complementary Metal-Oxide-Semiconductor) image sensor or a CCD (Charge Coupled Device) image sensor.
[0021] The communication I / F 44 is connected to a network 54. The communication I / Fs 44 and 26 control the exchange of various information between the processor 46 and the processor 28 via the network 54.
[0022] FIG. 2 shows an example of the main functions of the data processing device 12 and the smart device 14.
[0023] 2, in the data processing device 12, a specific process is performed by the processor 28. A specific processing program 56 is stored in the storage 32. The specific processing program 56 is an example of a "program" according to the technology of the present disclosure. The processor 28 reads the specific processing program 56 from the storage 32 and executes the read specific processing program 56 on the RAM 30. The specific process is realized by the processor 28 operating as a specific processing unit 290 in accordance with the specific processing program 56 executed on the RAM 30.
[0024] The storage 32 stores a data generation model 58 and an emotion identification model 59. The data generation model 58 and the emotion identification model 59 are used by the identification processing unit 290.
[0025] In the smart device 14, the processor 46 performs the reception output process. The storage 50 stores a reception output program 60. The reception output program 60 is used in conjunction with the specific processing program 56 by the data processing system 10. The processor 46 reads the reception output program 60 from the storage 50 and executes the read reception output program 60 on the RAM 48. The reception output process is realized by the processor 46 operating as the control unit 46A in accordance with the reception output program 60 executed on the RAM 48.
[0026] Next, the specific processing by the specific processing unit 290 of the data processing device 12 will be described.
[0027] "Example 1"
[0028] One embodiment of the present invention uses a smartphone application. In this case, the application installed on the smartphone detects an incoming call and listens to the conversation through the microphone. The listened conversation data is transcribed and sent to a generative AI. The generative AI determines from the conversation data that fraud is suspected, and if fraud is suspected, issues a voice alert from the smartphone speaker.
[0029] "Example 2"
[0030] Another embodiment of the present invention involves installing a microphone on a home phone. In this case, the microphone on the home phone listens to the phone conversation, transcribes the conversation data, and sends it to a generative AI. The generative AI determines from the conversation data that fraud is suspected, and if fraud is suspected, issues an audio alert from the home phone's speaker.
[0031] "Example 3"
[0032] In a further embodiment of the present invention, the generative AI detects specific phrases or patterns that indicate the possibility of special fraud. Specifically, the generative AI learns typical phrases for "bank transfer fraud" and typical patterns for "it's me" fraud, and by detecting these, it determines the possibility of fraud.
[0033] The processing flow of each embodiment will be described below.
[0034] "Example 1"
[0035] Step 1: An application installed on the smartphone detects an incoming call.
[0036] Step 2: The application listens to your conversation through your smartphone's microphone.
[0037] Step 3: The conversation data is transcribed and sent to the generative AI.
[0038] Step 4: The generative AI determines that fraud is suspected based on the conversation data.
[0039] Step 5: If fraud is suspected, an audio alert will be issued through the smartphone speaker.
[0040] "Example 2"
[0041] Step 1: A microphone installed in a home phone listens to the phone conversation.
[0042] Step 2: The conversation data is transcribed and sent to the generative AI.
[0043] Step 3: The generative AI determines that fraud is suspected based on the conversation data.
[0044] Step 4: If fraud is suspected, an audio alert will be issued through the home phone speaker.
[0045] "Example 3"
[0046] Step 1: The generative AI learns specific phrases or patterns that indicate potential fraud.
[0047] Step 2: The generative AI detects specific phrases or patterns it has learned to identify suspected fraud in the conversation data.
[0048] Step 3: If fraud is suspected, an audio alert will be issued.
[0049] Example 1
[0050] Next, a description will be given of Example 1 of Form Example 1. In the following description, the data processing device 12 will be referred to as a "server" and the smart device 14 will be referred to as a "terminal."
[0051] In recent years, telephone fraud has been on the rise, and special frauds targeting the elderly in particular have become a social problem. Conventional countermeasures require users to identify fraud themselves when receiving a fraudulent call, but this is becoming more difficult as fraud methods become more sophisticated. Therefore, there is a need for a system that automatically detects the possibility of fraud and issues a warning when a user receives a fraudulent call.
[0052] The identification process by the identification processing unit 290 of the data processing device 12 in Example 1 is realized by the following means. In this invention, the server includes a means for detecting an incoming fraudulent call, a means for listening to the conversation from a microphone installed in the communication terminal, a means for converting the conversation data into text and sending it to the generative AI model, and a means for issuing an audio alert from the speaker of the communication terminal when the generative AI model determines that fraud is suspected based on the conversation data. This makes it possible to automatically detect the possibility of fraud when a user receives a fraudulent call and immediately issue a warning.
[0053] "Means for detecting incoming fraudulent calls" refers to a function or device for detecting incoming calls in a communication terminal.
[0054] "Means for listening to conversations from a microphone installed in a communication terminal" refers to a function or device that uses a microphone built into or connected to a communication terminal to collect voices during a call in real time.
[0055] "Means for converting conversation data into text and sending it to the generative AI model" refers to a function or device for converting collected voice data into text data and sending that text data to the generative AI model.
[0056] A "generative AI model" is an artificial intelligence model that analyzes input data and detects specific patterns or phrases.
[0057] "Means for issuing an audio alert from the speaker of the communication device when it is determined that fraud is suspected" refers to a function or device for playing an alarm sound or message through the speaker of the communication device when the generative AI model detects the possibility of fraud.
[0058] "Specific phrases or patterns that indicate the possibility of fraud" are specific words or sentence structures related to fraudulent acts, and the generative AI model's detection of these serves as the basis for determining the possibility of fraud.
[0059] An "audio alert" is a message or sound that provides an audio warning to the user.
[0060] The present invention is a system that automatically detects fraudulent calls and issues a warning to the user using an application installed on a communication terminal. A specific embodiment of this system will be described below.
[0061] First, the user installs a dedicated application on their communication device. This application has the function of detecting incoming calls. Specifically, it uses the communication device's telephone API to catch incoming call events.
[0062] When a user answers a call, the device listens to the conversation in real time through the microphone. At this time, the application runs in the background and captures the conversation. The listened conversation data is transcribed using the Google® Speech-to-Text API. The voice data is sent to the API and received as text data.
[0063] The transcribed conversation data is sent to a server via the Internet. The HTTPS protocol is used for transmission to ensure data security. The server inputs the received conversation data into a generative AI model (e.g., GPT-4 (registered trademark) from OpenAI (registered trademark)). The generative AI model uses prompt sentences to analyze the possibility of fraud. Specific examples of prompt sentences are as follows:
[0064] Analyze the conversation data below to determine if it is a potential scam.
[0065] Conversation data: "Your bank account has been compromised. Please provide your account number to verify."
[0066] The generative AI model analyzes the conversation data and determines whether fraud is suspected. If so, the server sends that information back to the device, again using the HTTPS protocol.
[0067] The device analyzes the results received from the server, and if fraud is suspected, it issues an audio alert from the device's speaker. Specifically, it plays an audio message such as "There is a possibility of fraud. Please be careful." This alert makes the user aware of the possibility of fraud and allows them to take appropriate action.
[0068] As a concrete example, consider the case where a user receives a phone call saying, "Your bank account has been fraudulently used. Please provide your account number to verify." This conversation is picked up by a microphone and transcribed. The transcribed data becomes, "Your bank account has been fraudulently used. Please provide your account number to verify." This data is sent to a generative AI model, which determines that there is a high possibility of fraud. The server sends the result back to the device, which then issues an audio alert. The user hears this alert and becomes aware of the possibility of fraud.
[0069] In this way, the present invention can automatically detect possible fraud and immediately issue a warning when a user receives a fraudulent call.
[0070] The flow of the identification process in the first embodiment will be described with reference to FIG.
[0071] Step 1: Detecting an incoming call
[0072] The terminal detects incoming calls through an application installed on the communication terminal. Specifically, it uses the communication terminal's telephone API to catch incoming call events. The input is the incoming call signal, and the output is the incoming call detection event.
[0073] Step 2: Listen to the conversation
[0074] When a user answers a call, the device listens to the conversation in real time through the microphone. The application runs in the background and captures the conversation. The input is the voice data during the call, and the output is the collected voice data.
[0075] Step 3: Transcribe the conversation data
[0076] The device uses the Google Speech-to-Text API to transcribe the conversation data it hears. It sends the speech data to the API and receives it as text data. The input is the collected speech data, and the output is the transcribed text data.
[0077] Step 4: Send the transcribed data
[0078] The terminal sends the transcribed conversation data to the server. The HTTPS protocol is used for transmission to ensure data security. The input is the transcribed text data, and the output is the data transmission event to the server.
[0079] Step 5: Fraud detection
[0080] The server inputs the received conversation data into the generative AI model, which analyzes the possibility of fraud using prompt sentences. Specific examples of prompt sentences are as follows:
[0081] Analyze the conversation data below to determine if it is a potential scam.
[0082] Conversation data: "Your bank account has been compromised. Please provide your account number to verify."
[0083] The input is transcribed text data, and the output is a judgment result regarding the possibility of fraud.
[0084] Step 6: Receiving the results
[0085] The server receives the judgment result from the generative AI model and returns the result to the terminal. The return is again via HTTPS. The input is the judgment result from the generative AI model, and the output is a data transmission event to the terminal.
[0086] Step 7: Generate an audio alert
[0087] The terminal analyzes the judgment result received from the server, and if fraud is suspected, it issues an audio alert from the communication terminal's speaker. Specifically, it plays an audio message such as "There is a possibility of fraud. Please be careful." The input is the judgment result from the server, and the output is the generation of an audio alert.
[0088] (Application example 1)
[0089] Next, a description will be given of Application Example 1 of Embodiment Example 1. In the following description, the data processing device 12 will be referred to as a "server" and the smart device 14 will be referred to as a "terminal."
[0090] In recent years, telephone fraud has been on the rise, and special fraud targeting the elderly in particular has become a social problem. Conventional fraud prevention measures have made it difficult for users to recognize the possibility of fraud, making it difficult to prevent damage before it occurs. Therefore, there is a need for a system that analyzes telephone conversations in real time and issues an immediate warning if there is a possibility of fraud.
[0091] The specific processing by the specific processing unit 290 of the data processing device 12 in the application example 1 is realized by the following means.
[0092] In this invention, the server includes means for detecting incoming fraudulent calls, means for listening to conversations from a voice input device installed in the communication terminal, means for converting the conversation data into text and sending it to the generative AI, means for issuing an audio alert from the voice output device if the generative AI determines that fraud is suspected from the conversation data, means for the generative AI to detect specific phrases or patterns that indicate the possibility of fraud, and means for the audio alert to alert the user to the possibility of fraud. This allows the user to immediately become aware of the possibility of fraud and prevent damage before it occurs.
[0093] A "scam call" is a call made with the intent of defrauding someone of money or personal information through fraudulent means.
[0094] An "incoming call" refers to an incoming phone call.
[0095] A "communication terminal" is a device for sending and receiving voice and data, and includes smartphones and home telephones.
[0096] An "audio input device" is a device for converting audio into digital data, and a microphone is an example of this.
[0097] "Conversation data" is data obtained by converting speech acquired by a speech input device into text.
[0098] "Generative AI" is artificial intelligence that analyzes input data and detects specific patterns and phrases.
[0099] An "audio output device" is a device for outputting digital data as audio, and a speaker is an example of this.
[0100] An "audio alert" is a warning sound or message emitted from an audio output device.
[0101] A "particular phrase or pattern" is a characteristic word or sentence structure that indicates possible fraud.
[0102] "User" refers to any individual or organization that uses this system.
[0103] A system for implementing this invention detects incoming fraudulent calls and listens to the conversation through a voice input device installed on a communication terminal. The conversation data is transcribed in real time and sent to a generative AI. The generative AI detects specific phrases or patterns from the conversation data that indicate the possibility of fraud, and if fraud is suspected, issues a voice alert from a voice output device.
[0104] Hardware and software used
[0105] Hardware: Communication devices such as smartphones and home phones, microphones (audio input devices), speakers (audio output devices)
[0106] software:
[0107] SpeechRecognition: A Python library for converting speech to text
[0108] requests: A Python library for sending HTTP requests
[0109] Generative AI: Artificial intelligence that analyzes text data using external APIs
[0110] System Operation
[0111] 1. Incoming call detection: When the communication terminal detects an incoming call, it starts recording the conversation.
[0112] 2. Conversation transcription: The recorded audio data is transcribed in real time using the SpeechRecognition library.
[0113] 3. Generative AI Analysis: The transcribed conversation data is sent via HTTP requests to a generative AI, which detects specific phrases or patterns that indicate potential fraud.
[0114] 4. Alert: If the generative AI detects a potential fraud, an audio alert will be emitted from the device's speaker to alert the user to the potential fraud.
[0115] Specific examples
[0116] For example, if a user receives a phone call and the caller says, "Your bank account has been compromised. Please provide me with your account information immediately," the system will transcribe this conversation in real time and send it to a generative AI that will detect that this phrase matches certain patterns that indicate potential fraud and immediately trigger an audio alert.
[0117] Prompt Sentence Examples
[0118] An example of a prompt for a generative AI model is:
[0119] "Determine whether the following text is a potential scam: 'Your bank account has been compromised. Please provide your account information immediately.'"
[0120] In this way, users can be made aware of potential fraud immediately and prevent damage before it occurs.
[0121] The flow of the specific processing in the application example 1 will be described with reference to FIG.
[0122] Step 1: Detect incoming calls
[0123] The communication terminal detects an incoming call. The input is the incoming call signal, and the output is a trigger to start recording. Specifically, the phone application of the communication terminal detects the incoming call and starts the recording module.
[0124] Step 2: Record the conversation
[0125] The audio input device (microphone) of the communication terminal records the conversation. The input is the audio data during the call, and the output is the recorded audio file. Specifically, the recording module captures the audio during the call and saves it as an audio file.
[0126] Step 3: Transcribe the conversation
[0127] The recorded audio data is transcribed using the SpeechRecognition library. The input is an audio file, and the output is text data. Specifically, the speech recognition engine analyzes the audio file and generates the corresponding text.
[0128] Step 4: Generative AI analysis
[0129] Transcribed text data is sent to the generative AI via an HTTP request. The input is text data, and the output is a flag indicating the possibility of fraud. Specifically, a request containing the text data is sent to the generative AI's API endpoint, and the analysis results are received.
[0130] Step 5: Determine the likelihood of fraud
[0131] Generative AI analyzes text data to detect specific phrases or patterns that indicate potential fraud. The input is text data, and the output is a flag indicating potential fraud. Specifically, generative AI analyzes text data and flags cases where there is a high probability of fraud.
[0132] Step 6: Send an alert
[0133] If a possible fraud is detected, an audio alert is issued from the audio output device (speaker) of the communication terminal. The input is a flag indicating the possible fraud, and the output is an audio alert. Specifically, the audio alert module receives the flag indicating the possible fraud and plays an audio message to warn the user.
[0134] In this way, users can be made aware of potential fraud immediately and prevent damage before it occurs.
[0135] Example 2
[0136] Next, a description will be given of Example 2 of Form Example 2. In the following description, the data processing device 12 will be referred to as a "server" and the smart device 14 will be referred to as a "terminal."
[0137] In recent years, fraudulent phone calls have become more sophisticated, and many people have fallen victim to them. In particular, there has been an increase in special frauds targeting the elderly, and effective countermeasures against this are needed. Conventional countermeasures lack a system that can detect fraudulent phone calls in real time and issue warnings to users, making it difficult to prevent damage before it occurs.
[0138] The specific processing by the specific processing unit 290 of the data processing device 12 in the second embodiment is realized by the following means.
[0139] In this invention, the server includes means for detecting incoming fraudulent calls, means for collecting conversations from a voice input device installed in the communication device, means for converting the conversation data into text data, means for transmitting the text data to the generative AI model, and means for issuing a voice alert from the voice input device when the generative AI model determines that fraud is suspected from the conversation data. This makes it possible to detect fraudulent calls in real time and issue an immediate warning to the user.
[0140] The "means for detecting incoming fraudulent calls" is a function for identifying and detecting incoming calls that may be fraudulent calls to a communication device.
[0141] An "audio input device" is a microphone or other audio capture device installed on a communications device to capture a user's speech.
[0142] "Means for converting speech data into text data" refers to speech recognition software or algorithms used to analyze collected speech data and convert it into text format.
[0143] A "generative AI model" is an artificial intelligence model that analyzes input text data and determines the likelihood of fraud based on specific patterns and phrases.
[0144] "Audio alerting means" refers to a feature that plays an audio message to warn the user if fraud is suspected.
[0145] The present invention is a system for detecting incoming fraudulent calls in real time and issuing a warning to the user. A specific embodiment of this system will be described below.
[0146] 1. System Configuration
[0147] The system consists of the following major components:
[0148] How to detect incoming fraudulent calls
[0149] Voice input device
[0150] A means of converting conversation data into text data
[0151] Generative AI Models
[0152] A means of issuing an audio alert
[0153] 2. Hardware and Software Use
[0154] The voice input device is a microphone installed in a home telephone, which collects the user's speech in real time.
[0155] The means of converting conversation data into text data is through speech recognition software (e.g., Google Speech-to-Text API), which analyzes collected voice data and converts it into text data.
[0156] The generative AI model uses advanced artificial intelligence models such as OpenAI GPT-4, which analyzes input text data and determines the likelihood of fraud.
[0157] The audio alert method uses the speaker on the home phone, which will provide an audio alert if fraud is suspected.
[0158] 3. System Operation
[0159] When a user starts a conversation on their home phone, a voice input device collects the conversation. The collected voice data is converted into text data by speech recognition software. This text data is sent to a server and input into a generative AI model. The generative AI model analyzes the conversation data and, if it determines that fraud is suspected, the server sends an alert signal to the home phone. The home phone's speaker emits an audio alert to warn the user.
[0160] 4. Specific Examples
[0161] For example, imagine a user speaking on their home phone, "Hello, your bank account is at risk. Immediate action is required." A voice input device collects this speech, which speech recognition software converts into text data. This text data is sent to a server and analyzed by a generative AI model. If the generative AI model determines that there is a high possibility of fraud, the server sends an alert signal to the home phone. The speaker on the home phone issues a voice alert saying, "Possible fraud. Be careful."
[0162] 5. Examples of prompts
[0163] "Please analyze the conversation data below and determine whether you suspect fraud. If so, please explain why.
[0164] Conversation data: 'Hello, your bank account has been compromised. We need your immediate attention.'"
[0165] In this way, users can be made aware of their fraud risk in real time.
[0166] The flow of the identification process in the second embodiment will be described with reference to FIG.
[0167] Step 1:
[0168] Audio collection
[0169] Terminal: A voice input device (microphone) installed in a home phone collects the user's conversation in real time.
[0170] Input: User's voice conversation
[0171] Output: Digital audio data
[0172] What it does: When a user starts a phone conversation, the microphone automatically captures the audio and stores it as digital audio data.
[0173] Step 2:
[0174] Transcription of audio data
[0175] Device: Built-in speech recognition software (e.g., Google Speech-to-Text API) in home phones converts collected voice data into text data.
[0176] Input: Digital audio data
[0177] Output: Text data
[0178] What it does: Speech recognition software analyzes the audio and generates text that says something like, "Hello, your bank account has been compromised. Immediate action is required."
[0179] Step 3:
[0180] Sending text data
[0181] Terminal: The home phone sends the transcribed conversation data to the server.
[0182] Input: Text data
[0183] Output: Send data to the server
[0184] How it works: A home phone sends data over the internet to a server, which then passes it on to a generative AI model.
[0185] Step 4:
[0186] Fraud detection
[0187] Server: A generative AI model (e.g., OpenAI GPT-4) analyzes the received conversation data and determines whether fraud is suspected.
[0188] Input: Text data
[0189] Output: Determination of likelihood of fraud
[0190] What it does: A generative AI model analyzes the text "Hello, your bank account has been compromised. Immediate action is required" and determines that it is likely fraudulent.
[0191] Step 5:
[0192] Audio alert occurs
[0193] Server: If fraud is suspected, the server sends an alert signal to the home phone.
[0194] Terminal: The speaker on the home phone receives the alert signal from the server and issues an audio alert.
[0195] Input: Alert signal
[0196] Output: Audio alert
[0197] What it does: The speaker on your home phone plays a voice alert saying, "Possible scam. Be careful."
[0198] (Application example 2)
[0199] Next, a description will be given of Application Example 2 of Form Example 2. In the following description, the data processing device 12 will be referred to as a "server" and the smart device 14 will be referred to as a "terminal."
[0200] In recent years, the number of victims of fraudulent phone calls has been increasing, and special frauds targeting the elderly in particular have become a social problem. Conventional fraud prevention systems have been inadequate in detecting fraudulent phone calls, making it difficult for users to recognize the possibility of fraud. Furthermore, there is a need for a system that can detect fraud in real time and issue a warning to users on home phones and smartphones.
[0201] The specific processing by the specific processing unit 290 of the data processing device 12 in the application example 2 is realized by the following means.
[0202] In this invention, the server includes means for detecting incoming fraudulent calls, means for listening to conversations using a voice recognition device, means for converting the conversation data into text and sending it to a generative AI, means for issuing a voice alert from a voice output device when the generative AI determines that fraud is suspected based on the conversation data, and means including an application to be installed on a smartphone. This enables real-time detection of fraudulent calls and prompt warning to users.
[0203] A "scam call" is a call made with the intent of defrauding someone of money or personal information through fraudulent means.
[0204] An "incoming call" refers to an incoming phone call.
[0205] A "voice recognition device" is a device that converts voice into text data.
[0206] "Conversation data" refers to the content of a conversation that has been converted into text by a voice recognition device.
[0207] "Generative AI" is artificial intelligence that generates new information based on input data.
[0208] An "audio output device" is a device for outputting audio.
[0209] "Voice alert" is a function that issues a warning by voice.
[0210] An "application installed on a smartphone" is a software program that runs on a smartphone.
[0211] To implement this invention, the following hardware and software are required. The hardware requires a smartphone, a voice recognition device, and a voice output device. The software requires a voice recognition library (e.g., speech_recognition), a generative AI library (e.g., openai), and a text-to-speech library (e.g., pyttsx3).
[0212] First, the phone call is recorded in real time using the smartphone's microphone. The recorded voice data is converted into text data using a voice recognition device. The speech_recognition library is used for this voice recognition.
[0213] The converted text data is then sent to the generative AI, which analyzes the input text data and determines whether it is fraudulent. This analysis is performed using the OpenAI library. Specifically, the following prompt is sent to the generative AI:
[0214] Please determine if the following conversation is a scam:
[0215] "Hello, this is bank security. Your account has been compromised. Please provide your account number and PIN."
[0216] If you suspect it is a scam, reply 'scam'.
[0217] If the generative AI determines that there is a possibility of fraud, it will issue a voice alert using the smartphone's audio output device. This voice alert uses the pyttsx3 library. The content of the voice alert is intended to alert the user to the possibility of fraud.
[0218] For example, if a user is on a call on their smartphone and a conversation that could be fraudulent takes place, the application automatically converts the conversation into text and sends it to generative AI to determine whether it is fraudulent. If it is determined that there is a possibility of fraud, an audio alert will sound from the smartphone speaker saying, "Possible fraud!"
[0219] In this way, the present invention provides real-time detection of fraudulent calls and prompt warning to users.
[0220] The flow of the specific processing in the application example 2 will be described with reference to FIG.
[0221] Step 1:
[0222] A user initiates a call on their smartphone, and the smartphone's microphone records the call in real time.
[0223] Input: Call audio
[0224] Output: Recorded audio data
[0225] How it works: The smartphone's microphone captures the call audio and saves it as audio data.
[0226] Step 2:
[0227] The terminal sends the recorded voice data to a voice recognition device, which converts the voice data into text data.
[0228] Input: Recorded audio data
[0229] Output: Text data
[0230] Specific operation: A speech recognition library (e.g., speech_recognition) analyzes the audio data and generates corresponding text data.
[0231] Step 3:
[0232] The device sends the generated text data to the generative AI, which analyzes the text data and determines whether it is fraudulent.
[0233] Input: Text data
[0234] Output: Judgment result on likelihood of fraud
[0235] How it works: A generative AI library (e.g., openai) analyzes the text data based on the prompt and determines whether it is likely to be fraudulent.
[0236] Step 4:
[0237] The server receives the judgment results from the generative AI and generates an audio alert if there is a possibility of fraud.
[0238] Input: Determination result regarding the possibility of fraud
[0239] Output: Audio alert content
[0240] Specific behavior: The generative AI generates a warning message saying, "Possible fraud!"
[0241] Step 5:
[0242] The terminal transmits the contents of the audio alert to the audio output device, and the audio alert is output.
[0243] Input: Audio alert content
[0244] Output: Audio alert
[0245] Specific operation: A text-to-speech library (e.g., pyttsx3) converts the contents of the voice alert into audio, and outputs the warning sound from the smartphone speaker.
[0246] Example 3
[0247] Next, a description will be given of a third embodiment of the third embodiment. In the following description, the data processing device 12 will be referred to as a "server" and the smart device 14 will be referred to as a "terminal."
[0248] Conventional fraud detection systems are limited to detecting incoming fraudulent phone calls, and are inadequate at detecting fraud in text messages and emails received by users. In addition, there are limited means of notifying users of possible fraud, which can delay users' realization of the risk of fraud. This makes it difficult to prevent fraud damage before it happens.
[0249] The specific processing by the specific processing unit 290 of the data processing device 12 in the third embodiment is realized by the following means.
[0250] In this invention, the server includes means for transmitting text data entered by a user to the server, means for the server to analyze the text data using a generative artificial intelligence model and determine the possibility of fraud, and means for the server to transmit the analysis results to the communication terminal, thereby making it possible to quickly detect the possibility of fraud even in text messages or emails received by the user and notify the user.
[0251] "Means for detecting incoming fraudulent calls" means a device or software for identifying and detecting potentially fraudulent calls from among calls coming in through a communication line.
[0252] "A voice input device installed on a communication terminal" refers to a microphone or related hardware for collecting voice that is attached to a communication device such as a smartphone or home phone.
[0253] "Means for converting conversation data into text and transmitting it to the generative AI" refers to software or hardware for converting voice data into text data and transmitting that text data to the generative AI.
[0254] "Generative AI" is an AI model that uses natural language processing technology to analyze text data and detect specific patterns and phrases.
[0255] "Audio alerting means" means a device or software that issues an audio warning when it determines that fraud may be occurring.
[0256] The "means for transmitting text data input by the user to the server" refers to software or hardware for transmitting text data input by the user to the communication terminal to the server via a communication line such as the Internet.
[0257] "Means for the server to analyze text data using a generative artificial intelligence model and determine the possibility of fraud" refers to software or hardware that inputs the text data received by the server into a generative artificial intelligence model and evaluates the possibility of fraud based on the analysis results.
[0258] The "means by which the server transmits the analysis results to the communication terminal" refers to software or hardware for transmitting the analysis results of the generative artificial intelligence model to the communication terminal.
[0259] The "means by which the communication terminal displays the analysis results to the user" refers to software or hardware for visually or audibly notifying the user of the analysis results received from the server.
[0260] The present invention relates to a system for detecting the possibility of special fraud using a generative artificial intelligence model. Specific embodiments of this system will be described below.
[0261] System configuration
[0262] Subject: Server
[0263] The server is responsible for analyzing the text data using a generative artificial intelligence model to determine the likelihood of fraud. Specifically, the server uses the following hardware and software:
[0264] Hardware: High-performance processors, memory, and storage devices
[0265] Software: Generative artificial intelligence models (e.g., AI models using natural language processing technology)
[0266] The server receives the text data sent by the user and inputs it into a generative AI model. The model detects specific phrases and patterns that indicate potential fraud and returns the results to the server. The server then transmits the analysis results to the communication device.
[0267] Subject: Terminal
[0268] The terminal is responsible for sending the text data entered by the user to the server and displaying the analysis results from the server to the user. Specifically, the terminal uses the following hardware and software:
[0269] Hardware: Communication devices such as smartphones, tablets, and PCs
[0270] Software: Dedicated application
[0271] The terminal transmits the text data entered by the user to the server, and displays the analysis results received from the server to the user. If the analysis results indicate a possibility of fraud, the terminal displays a warning message.
[0272] Subject: User
[0273] The user is responsible for inputting text data through the terminal and checking the analysis results. Specifically, the user performs the following operations:
[0274] Enter the contents of a message or email into your device
[0275] Check the analysis results and take appropriate action if necessary
[0276] Specific examples
[0277] Example 1
[0278] If a user receives a message saying "Mom, please transfer the money now," the system will act as follows:
[0279] 1. The user enters a message into a dedicated application on the device.
[0280] 2. The device sends this message to the server.
[0281] 3. The server receives the message and inputs it as a prompt sentence into the generative artificial intelligence model.
[0282] 4. The generative artificial intelligence model analyzes the message and returns the result: "Possible fraud."
[0283] 5. The server sends the analysis results to the device.
[0284] 6. The device displays the analysis results to the user, and the user confirms the warning message.
[0285] Example 2
[0286] If a user receives an email saying "Your bank account has been compromised. Contact us immediately," the system works as follows:
[0287] 1. The user enters the contents of the email into a dedicated application on the device.
[0288] 2. The device sends the contents of this email to the server.
[0289] 3. The server receives the email content and inputs it as a prompt into the generative artificial intelligence model.
[0290] 4. The generative artificial intelligence model analyzes the content of the email and returns the result: "Possible fraud."
[0291] 5. The server sends the analysis results to the device.
[0292] 6. The device displays the analysis results to the user, and the user confirms the warning message.
[0293] Prompt Sentence Examples
[0294] Here are some example prompts to input to a generative AI model:
[0295] "Determine whether the following message is a potential scam: 'Mom, please transfer the money now.'"
[0296] "Determine if the following email is a potential scam: 'Your bank account has been compromised. Contact us now.'"
[0297] In this way, the specific fraud detection system using the generative artificial intelligence model operates specifically. The flow of the identification process in the third embodiment will be described with reference to FIG.
[0298] Step 1:
[0299] The user inputs text data.
[0300] The user enters the contents of the message or email they received into a dedicated application on their device. For example, they copy and paste the message "Mom, please transfer the money right now" into the application. The input data is saved in text format on the device.
[0301] Step 2:
[0302] The terminal transmits the text data to the server.
[0303] The terminal sends text data entered by the user to the server. Specifically, the terminal application generates an HTTP request and sends a payload containing the text data to the server. At this time, the data is encrypted before being sent. The input is the text data entered by the user, and the output is the text data sent to the server.
[0304] Step 3:
[0305] The server receives the text data and inputs it into a generative artificial intelligence model.
[0306] The server receives text data sent from the terminal. The received data is first stored in a database. The server then inputs the text data as a prompt sentence into the generative AI model. An example of a prompt sentence is "Please judge whether the following message is likely to be fraudulent: 'Mom, please transfer the money now.'" The input is the text data received from the terminal, and the output is the prompt sentence input into the generative AI model.
[0307] Step 4:
[0308] A generative artificial intelligence model analyzes the text data and returns the results.
[0309] The generative artificial intelligence model analyzes the input prompt sentence and determines whether it is likely to be fraud. Specifically, the model analyzes it by comparing it with typical phrases and patterns of "bank transfer fraud" and "it's me fraud." If the analysis result indicates a high possibility of fraud, it returns a warning message saying "Possible fraud." The input is the prompt sentence, and the output is the analysis result.
[0310] Step 5:
[0311] The server sends the analysis results to the device.
[0312] The server sends the analysis results obtained from the generative AI model to the terminal. Specifically, the server generates an HTTP response and sends a payload containing the analysis results to the terminal. At this time, the data is encrypted before being sent. The input is the analysis results from the generative AI model, and the output is the analysis results sent to the terminal.
[0313] Step 6:
[0314] The terminal displays the analysis results to the user.
[0315] The terminal displays the analysis results received from the server to the user. For example, if the analysis result is a warning message saying "Possible fraud," the terminal application displays this message to the user as a pop-up window or notification. The input is the analysis result received from the server, and the output is the warning message displayed to the user. The user can check this warning message and take appropriate action.
[0316] (Application example 3)
[0317] Next, a description will be given of Application Example 3 of Form Example 3. In the following description, the data processing device 12 will be referred to as a "server" and the smart device 14 will be referred to as a "terminal."
[0318] In recent years, special fraud methods have become more sophisticated, and many people have fallen victim to them. In particular, "bank transfer fraud" and "it's me" frauds targeting the elderly have become a social problem. In order to prevent these frauds, a system is needed that can detect possible fraud early and issue a warning to users. However, current systems have low fraud detection accuracy, making it difficult for users to recognize the possibility of fraud.
[0319] The specific processing by the specific processing unit 290 of the data processing device 12 in the application example 3 is realized by the following means.
[0320] In this invention, the server includes means for detecting incoming fraudulent calls, means for listening to conversations from a voice input device installed in the communication terminal, means for converting the conversation data into text and sending it to the generative AI, means for issuing a voice alert from the voice input device when the generative AI determines that fraud is suspected from the conversation data, means for the generative AI to detect specific phrases or patterns that indicate the possibility of fraud, and means for issuing a warning to the user when the generative AI detects specific phrases or patterns that indicate the possibility of fraud. This makes it possible to detect the possibility of fraud with high accuracy and issue a warning to the user quickly.
[0321] "Means for detecting incoming fraudulent calls" is a function for identifying and notifying a communication terminal of incoming calls that may be fraudulent.
[0322] A "voice input device installed on a communication terminal" is a device for collecting voice that is attached to communication equipment such as a smartphone or home telephone.
[0323] "Means for converting conversation data into text and sending it to generative AI" refers to a function for converting voice data collected by a voice input device into text data and sending that text data to generative AI.
[0324] "Means for issuing a voice alert from the voice input device when the generative AI determines that fraud is suspected based on the conversation data" is a function for issuing a warning sound through the voice input device when the generative AI determines that there is a possibility of fraud based on the conversation data analyzed.
[0325] "Means for generative AI to detect specific phrases or patterns that indicate the possibility of fraud" refers to a function that allows generative AI to detect typical fraudulent phrases and patterns that it has previously learned from conversation data.
[0326] "Means for issuing a warning to users when generative AI detects specific phrases or patterns that indicate the possibility of fraud" refers to a function that issues a warning to users when generative AI detects phrases or patterns that indicate the possibility of fraud.
[0327] As an embodiment of the present invention, a method for installing a fraud detection assistant system on a smartphone will be described.
[0328] First, the system is equipped with a means for detecting incoming fraudulent calls. This means has the function of identifying and notifying the communication terminal of calls that may be fraudulent. Specifically, the system collects the contents of the call in real time using a voice input device installed on the communication terminal.
[0329] The collected voice data is then converted into text data through a means of transcribing the conversation data and sending it to the generative AI. This conversion is done using speech recognition software (e.g., Google's speech recognition API). The converted text data is then sent to the generative AI model.
[0330] A generative AI model is pre-trained to detect specific phrases and patterns that indicate potential fraud. This model is built using, for example, the Hugging Face transformers library. If the generative AI determines from the conversation data that fraud is suspected, it will warn the user by issuing an audio alert via a voice input device.
[0331] Additionally, if the generative AI detects certain phrases or patterns that indicate potential fraud, it will activate a user alert mechanism that will alert the user to the potential fraud by providing a visual or audio warning.
[0332] As a concrete example, consider the following audio input file:
[0333] Voiceover: "Mom, please transfer the money now. It's urgent."
[0334] An example prompt for parsing this audio file:
[0335] Determine whether the phrase "Mom, please transfer the money now. It's urgent." is a potential scam.
[0336] By feeding this prompt into a generative AI model, it can analyze the possibility of fraud and issue a warning to the user.
[0337] The flow of the specific processing in Application Example 3 will be described with reference to FIG.
[0338] Step 1:
[0339] Detecting incoming fraudulent calls
[0340] The server identifies and notifies the communication terminal of incoming calls that may be fraudulent. Specifically, it monitors the communication terminal's call log and detects possible fraud based on specific numbers or patterns. The input is the call log data, and the output is a notification of a potentially fraudulent call.
[0341] Step 2:
[0342] Listen to conversations from a voice input device
[0343] The terminal collects the contents of the call in real time using an audio input device installed in the communication terminal. Specifically, it acquires audio data through a microphone and saves it as an audio file. The input is the audio during the call, and the output is an audio file.
[0344] Step 3:
[0345] Convert conversation data into text and send it to generative AI
[0346] The device converts the collected voice data into text data and sends it to the generative AI. Specifically, it uses voice recognition software (e.g., Google's voice recognition API) to transcribe the voice data. The input is an audio file, and the output is text data.
[0347] Step 4:
[0348] Generative AI analyzes conversation data
[0349] The server uses a generative AI model to detect specific phrases or patterns in text data that indicate potential fraud. Specifically, it analyzes the text data using a generative AI model (e.g., Hugging Face's transformers library). The input is the text data, and the output is a determination of the likelihood of fraud.
[0350] Step 5:
[0351] Provides audio alerts in case of suspected fraud
[0352] If the generative AI determines that there is a possibility of fraud, the device will issue a voice alert through the voice input device. Specifically, it will play a warning sound or message. The input is the result of the judgment on the possibility of fraud, and the output is a voice alert.
[0353] Step 6:
[0354] Warn the user
[0355] The device will alert the user if the generative AI detects certain phrases or patterns that indicate potential fraud, either by displaying a warning message on the screen or by sending a notification. The input is the determination of potential fraud, and the output is the warning to the user.
[0356] Furthermore, an emotion engine that estimates the user's emotion may be combined. That is, the identification processing unit 290 may estimate the user's emotion using the emotion identification model 59 and perform identification processing using the user's emotion.
[0357] "Example 1"
[0358] One embodiment of the present invention is a system that combines a generative AI with an emotion engine. This system analyzes emotions based on the user's tone of voice, volume, and speaking rate. Specifically, if the user's tone or volume suddenly increases while receiving a scam call, the system senses that the user may be panicking. Based on this information, the generative AI determines that the possibility of a scam has increased and issues an audio alert.
[0359] "Example 2"
[0360] In another embodiment, the emotion engine may generate an audio alert when a certain threshold is exceeded. For example, if the user's tone of voice exceeds a certain threshold or if the user's speaking rate exceeds a certain threshold, the emotion engine may determine that the user is experiencing strong emotion. Based on this determination, the generative AI may generate an audio alert to warn the user of a potential scam.
[0361] "Example 3"
[0362] Another possible embodiment is a system in which the emotion engine and generative AI work together. In this system, the emotion engine analyzes the user's emotions and sends the results to the generative AI. The generative AI then combines the information from the emotion engine with the conversation data it has analyzed to determine the possibility of fraud. This allows for more accurate fraud detection.
[0363] The processing flow of each embodiment will be described below.
[0364] "Example 1"
[0365] Step 1: The user receives a scam call.
[0366] Step 2: A smartphone app or a microphone installed on a home phone listens to the conversation.
[0367] Step 3: Convert the conversation data into text and send it to the generative AI.
[0368] Step 4: Generative AI analyzes the text data to determine the likelihood of fraud.
[0369] Step 5: The emotion engine analyzes the user's emotions based on their tone of voice, volume, speaking rate, etc.
[0370] Step 6: Based on the information from the emotion engine, the generative AI determines that the likelihood of fraud has increased and issues an audio alert.
[0371] "Example 2"
[0372] Step 1: The user receives a scam call.
[0373] Step 2: A smartphone app or a microphone installed on a home phone listens to the conversation.
[0374] Step 3: Convert the conversation data into text and send it to the generative AI.
[0375] Step 4: Generative AI analyzes the text data to determine the likelihood of fraud.
[0376] Step 5: The emotion engine analyzes the user's emotions based on their tone of voice, volume, speaking rate, etc.
[0377] Step 6: If the emotion engine exceeds a certain threshold, notify the generative AI.
[0378] Step 7: Based on the notification from the emotion engine, the generative AI determines that the likelihood of fraud has increased and issues an audio alert.
[0379] "Example 3"
[0380] Step 1: The user receives a scam call.
[0381] Step 2: A smartphone app or a microphone installed on a home phone listens to the conversation.
[0382] Step 3: Convert the conversation data into text and send it to the generative AI.
[0383] Step 4: Generative AI analyzes the text data to determine the likelihood of fraud.
[0384] Step 5: The emotion engine analyzes the user's emotions based on their tone of voice, volume, speaking speed, etc., and sends the results to the generative AI.
[0385] Step 6: The generative AI combines information from the emotion engine with the conversation data it has analyzed to determine the likelihood of fraud.
[0386] Step 7: If fraud is deemed likely, the generative AI will issue an audio alert.
[0387] Example 1
[0388] Next, a description will be given of Example 1 of Form Example 1. In the following description, the data processing device 12 will be referred to as a "server" and the smart device 14 will be referred to as a "terminal."
[0389] Conventional fraud call detection systems analyze only the content of conversations to determine the possibility of fraud, which means they are unable to take into account the emotional state of the user, resulting in low accuracy in detecting fraud. Furthermore, they are unable to respond adequately when the user panics, making it difficult to prevent fraud damage before it happens.
[0390] The specific processing by the specific processing unit 290 of the data processing device 12 in the first embodiment is realized by the following means.
[0391] In this invention, the server includes means for detecting incoming fraudulent calls, means for listening to conversations from a voice input device installed in the communication terminal, means for converting the conversation data into text and sending it to the generative artificial intelligence, means for issuing a voice alert from the voice input device when the generative artificial intelligence determines that fraud is suspected from the conversation data, means for collecting parameters such as the user's voice tone, volume, and speaking rate and analyzing their emotions, and means for reevaluating the possibility of fraud based on the results of the emotion analysis. This makes it possible to detect fraud taking into account the user's emotional state, thereby preventing fraud damage before it occurs.
[0392] "Means for detecting incoming fraudulent calls" refers to a device or software for detecting whether an incoming call to a communication terminal is likely to be fraudulent.
[0393] The "voice input device installed in a communication terminal" refers to a device for inputting voice, such as a microphone attached to a communication device such as a smartphone or home telephone.
[0394] "Means for converting conversation data into text and transmitting it to the generative AI" refers to a device or software for converting voice data into text data and transmitting the text data to the generative AI.
[0395] "Generative AI" is an AI system that analyzes the data it receives and determines the likelihood of fraud by detecting specific patterns and phrases.
[0396] A "means for issuing an audio alert" is a device or software that plays an audio message to warn the user when fraud is deemed likely.
[0397] "Means for collecting parameters such as the tone, volume, and speaking rate of a user's voice and analyzing emotions" refers to a device or software that collects information such as tone, volume, and speaking rate from the user's voice data and analyzes the user's emotional state based on that information.
[0398] "Means for reassessing the possibility of fraud based on the results of sentiment analysis" refers to a device or software that inputs the results of sentiment analysis into generative artificial intelligence and reassess the possibility of fraud.
[0399] MODE FOR CARRYING OUT THE INVENTION
[0400] This invention is a system that detects fraudulent calls and warns users. It is implemented using an application and server installed on a communication terminal. This system detects incoming fraudulent calls, listens to the conversation, transcribes the conversation data, and sends it to a generative artificial intelligence. Furthermore, the system aims to prevent fraud by analyzing the user's emotional state and reassessing the possibility of fraud.
[0401] Hardware and software used
[0402] 1. Communication terminal
[0403] Communication devices such as smartphones and home phones.
[0404] A microphone as an audio input device.
[0405] 2. Server
[0406] Voice recognition software (e.g., Google Cloud Speech-to-Text API) is used to convert voice data into text.
[0407] For sentiment analysis, an emotion engine (e.g., IBM Watson (registered trademark) Tone Analyzer) is used.
[0408] For generative artificial intelligence, we use AI models with natural language processing technology (e.g., OpenAI's GPT-4).
[0409] Specific operation of the system
[0410] 1. Incoming phone call detection
[0411] The server detects incoming calls through an application installed on the communication terminal, which uses the communication terminal's native API to catch incoming call events.
[0412] 2. Listening to conversations
[0413] The device listens to the conversation during the call using a microphone, and the application captures the microphone's audio input in real time and saves it as audio data.
[0414] 3. Transcription of audio data
[0415] The server converts the audio data it hears into text data by calling the Google Cloud Speech-to-Text API.
[0416] 4. Sending to generative AI
[0417] The server sends the transcribed conversation data to the generative AI. The server sends the text data to the generative AI using an HTTP request.
[0418] 5. Fraud Detection
[0419] The generative AI analyzes the received conversation data and determines whether fraud is suspected. The generative AI uses natural language processing technology to analyze the content of the conversation and evaluate whether it matches any fraudulent patterns.
[0420] 6. Sentiment Analysis
[0421] The server collects parameters such as the user's voice tone, volume, and speaking rate and analyzes them using an emotion engine. The server uses voice analysis software to extract emotion parameters from the voice data.
[0422] 7. Panic Detection
[0423] The server determines that a user may be panicking if there is a sudden increase in the tone or volume of their voice. The server analyzes data from the emotion engine and evaluates whether there is a sudden change.
[0424] 8. Reassessing the possibility of fraud
[0425] The generative AI determines that the likelihood of fraud has increased based on information from the emotion engine. The generative AI receives the emotion data as additional input and reassess the likelihood of fraud.
[0426] 9. Sending audio alerts
[0427] If the server determines that there is a high possibility of fraud, it instructs the communication terminal to issue an audio alert from its speaker, and sends a command to the application of the communication terminal to play the audio alert.
[0428] Examples and prompts
[0429] As a concrete example, imagine a user is receiving a phone call on their smartphone. If the caller says, "Your bank account has been fraudulently used. Please tell us your account number and password immediately," the user's voice tone will suddenly rise and the volume will also increase. Based on this information, the generative AI will determine that there is a high possibility of fraud and will issue a voice alert from the communication device's speaker saying, "This may be a fraud. Please be careful."
[0430] Example prompt sentence:
[0431] "Analyze the conversation data below to determine if it is a potential scam. Also consider any changes in the user's tone and volume.
[0432] Conversation data: 'Your bank account has been compromised. Please provide your account number and password immediately.'
[0433] User's tone of voice: Sudden rise
[0434] User volume: louder"
[0435] The flow of the identification process in the first embodiment will be described with reference to FIG.
[0436] Step 1:
[0437] Incoming phone call detection
[0438] The server detects an incoming call through an application installed on the communication terminal.
[0439] Input: Incoming call event of communication terminal.
[0440] Data processing / calculation: Capture incoming call events using the communication device's native API.
[0441] Output: Event information that an incoming call was detected.
[0442] Specific operation: The application uses the native API of ANDROID (registered trademark) or iOS to catch incoming call events and notify the server.
[0443] Step 2:
[0444] Listening to conversations
[0445] The terminal uses a microphone to listen to the conversation during the call.
[0446] Input: Voice data during a call.
[0447] Data processing / calculation: Captures microphone audio input in real time and saves it as audio data.
[0448] Output: Audio data.
[0449] What it does: The application captures audio input from the microphone in real time and saves it as audio data.
[0450] Step 3:
[0451] Transcription of audio data
[0452] The server converts the audio data it hears into text data.
[0453] Input: Audio data.
[0454] Data processing / calculation: Call the Google Cloud Speech-to-Text API to convert the voice data into text.
[0455] Output: Character data.
[0456] What happens: The server converts the audio data into text using the Google Cloud Speech-to-Text API.
[0457] Step 4:
[0458] Sending to generative AI
[0459] The server transmits the transcribed conversation data to the generative artificial intelligence.
[0460] Input: Character data.
[0461] Data processing / calculation: Send text data to the generative AI using an HTTP request.
[0462] Output: The result of the request sent to the generative AI.
[0463] Specific operation: The server sends text data to the generative artificial intelligence using an HTTP request.
[0464] Step 5:
[0465] Fraud detection
[0466] The generative artificial intelligence analyzes the conversation data it receives and determines whether fraud is suspected.
[0467] Input: Character data.
[0468] Data processing / computation: Using natural language processing techniques, the content of conversations is analyzed to assess whether they match fraud patterns.
[0469] Output: Assessment results for likelihood of fraud.
[0470] How it works: Generative AI uses natural language processing technology to analyze the content of a conversation and evaluate whether it matches a fraud pattern.
[0471] Step 6:
[0472] Sentiment analysis
[0473] The server collects parameters such as the user's voice tone, volume, and speaking speed and analyzes them using an emotion engine.
[0474] Input: Audio data.
[0475] Data processing / calculation: Using voice analysis software, emotional parameters are extracted from the voice data.
[0476] Output: Emotion parameters.
[0477] Specific operation: The server uses voice analysis software (e.g., IBM Watson Tone Analyzer) to extract emotion parameters from the voice data.
[0478] Step 7:
[0479] Panic detection
[0480] If the tone or volume of a user's voice suddenly increases, the server determines that the user may be panicking.
[0481] Input: Emotion parameters.
[0482] Data processing / computation: Analyzing data from the emotion engine and assessing whether there are any sudden changes.
[0483] Output: Panic assessment result.
[0484] What it does: The server analyzes the data from the emotion engine and evaluates whether there are any sudden changes.
[0485] Step 8:
[0486] Reassessing the possibility of fraud
[0487] Based on information from the emotion engine, the generative AI determines that the likelihood of fraud has increased.
[0488] Input: Emotion parameters.
[0489] Data processing / computation: Emotional data is taken as additional input to reassess the likelihood of fraud.
[0490] Output: Reassessed fraud probability.
[0491] How it works: Generative AI takes emotional data as additional input and reassess the likelihood of fraud.
[0492] Step 9:
[0493] Sending a voice alert
[0494] If the server determines that there is a high possibility of fraud, it instructs the communication terminal to issue an audio alert through its speaker.
[0495] Enter: Reassessed fraud potential.
[0496] Data processing / calculation: Sends a command to the communication terminal application to play an audio alert.
[0497] Output: Plays an audio alert.
[0498] Specific operation: The server sends a command to the application on the communication terminal to play an audio alert, and the audio alert is played from the speaker of the communication terminal.
[0499] (Application example 1)
[0500] Next, a description will be given of Application Example 1 of Embodiment Example 1. In the following description, the data processing device 12 will be referred to as a "server" and the smart device 14 will be referred to as a "terminal."
[0501] In recent years, fraudulent phone scams have become more sophisticated, and many people have fallen victim to them. Elderly people and those unfamiliar with technology are particularly vulnerable to these scams, as they are less vigilant against them. Furthermore, when users receive a fraudulent call, they often panic, making it difficult for them to make a calm judgment. Effective measures are needed to improve this situation and prevent fraudulent incidents.
[0502] The specific processing by the specific processing unit 290 of the data processing device 12 in the application example 1 is realized by the following means.
[0503] In this invention, the server includes means for detecting incoming fraudulent calls, means for listening to conversations from a voice input device installed in the communication terminal, means for converting the conversation data into text and sending it to the generative AI, means for issuing a voice alert from the voice input device when the generative AI determines that fraud is suspected from the conversation data, means including an emotion engine for analyzing emotions from the user's voice tone, volume, speaking speed, etc., and means for further increasing the possibility of fraud based on the analysis results of the emotion engine. This makes it possible to prevent fraud damage by combining fraud call detection with user emotion analysis.
[0504] A "scam call" is a call made with the intent of defrauding someone of money or personal information through fraudulent means.
[0505] An "incoming call" refers to an incoming phone call.
[0506] A "communication terminal" is an electronic device for transmitting and receiving voice and data.
[0507] An "audio input device" is a device for converting audio into digital data.
[0508] "Conversation data" is data obtained by converting speech acquired by a speech input device into text.
[0509] "Generative AI" is artificial intelligence that makes specific decisions and generates things based on input data.
[0510] A "voice alert" is a means of warning or notifying by voice.
[0511] An "emotion engine" is software or a system for analyzing emotions from voice data.
[0512] "Tone" refers to the pitch and texture of a voice.
[0513] "Volume" refers to the loudness of a sound.
[0514] "Speech rate" refers to the speed at which you speak.
[0515] The following system configuration is used as an embodiment of the present invention.
[0516] System Configuration
[0517] 1. Hardware
[0518] Communication terminal: A device capable of voice communication, such as a smartphone or home telephone.
[0519] Audio input device: A microphone built into the communication terminal.
[0520] Audio output device: A speaker built into a communication terminal.
[0521] 2. Software
[0522] Speech Recognition Library: Use the speech_recognition library to convert speech to text.
[0523] Generative AI model: An AI model that uses the transformers library to determine the likelihood of fraud.
[0524] Sentiment Engine: A model for analyzing user sentiment using the transformers library.
[0525] Audio Alert Library: Uses the pyttsx3 library to emit audio alerts.
[0526] Processing flow
[0527] 1. Incoming call detection
[0528] The server detects that a call has been received by the communication terminal.
[0529] The voice input device of the communication terminal listens to the conversation and acquires voice data.
[0530] 2. Voice Recognition
[0531] The server uses a voice recognition library to convert the acquired voice data into text data.
[0532] 3. Generative AI for Fraud Detection
[0533] The server inputs the text data into a generative AI model to determine the likelihood of fraud.
[0534] If certain phrases or patterns are detected, it is determined to be likely fraud.
[0535] 4. Sentiment analysis
[0536] The server uses an emotion engine to analyze the user's emotions based on the tone of their voice, volume, speaking speed, etc.
[0537] If the user is perceived as panicking, the likelihood of fraud increases even more.
[0538] 5. Audio alerts
[0539] If the server determines that there is a high possibility of fraud, it uses a voice alert library to issue a voice alert from the voice output device of the communication terminal.
[0540] The alerts are intended to raise awareness of potential fraud.
[0541] Specific examples
[0542] For example, if a user receives a phone call and it is determined that the call may be fraudulent, an audio alert will be emitted from the communication device's speaker saying, "Warning! Possible fraud." This alert will make the user aware of the possibility of fraud and allow them to respond calmly.
[0543] Prompt Sentence Examples
[0544] "Could this call be a scam?"
[0545] In this way, this invention can prevent fraud damage by combining fraud call detection with user emotion analysis.
[0546] The flow of the specific processing in the application example 1 will be described with reference to FIG.
[0547] Step 1:
[0548] The server detects that a call has been received by the communication terminal.
[0549] Input: Incoming signal from communication terminal.
[0550] Data processing: Receives an incoming call signal and starts processing when the call starts.
[0551] Output: Flag for starting the call.
[0552] Specific operation: When a communication terminal receives an incoming call, the server detects the signal and recognizes that a call has started.
[0553] Step 2:
[0554] The voice input device of the terminal listens to the conversation and acquires voice data.
[0555] Input: Audio during a call.
[0556] Data processing: The voice input device collects voice data in real time.
[0557] Output: Captured audio data.
[0558] Specific operation: Audio during a call is collected through the microphone and saved as audio data.
[0559] Step 3:
[0560] The server uses a voice recognition library to convert the acquired voice data into text data.
[0561] Input: Captured audio data.
[0562] Data processing: Convert the voice data into text data using the speech recognition library (speech_recognition).
[0563] Output: The converted text data.
[0564] Specific operation: Voice data is input into the voice recognition library and output as text data.
[0565] Step 4:
[0566] The server inputs the text data into a generative AI model to determine the likelihood of fraud.
[0567] Input: The converted text data.
[0568] Data processing: Text data is fed into generative AI models (transformers) to assess the likelihood of fraud.
[0569] Output: Assessment results for likelihood of fraud.
[0570] Specific operation: Text data is input into a generative AI model, which outputs an evaluation result on whether the data is likely to be fraudulent.
[0571] Step 5:
[0572] The server uses an emotion engine to analyze the user's emotions based on the tone of their voice, volume, speaking speed, etc.
[0573] Input: Converted text and audio data.
[0574] Data processing: Using emotion engines (transformers), we analyze user emotions.
[0575] Output: Analysis results on user sentiment.
[0576] Specific operation: Voice and text data are input into the emotion engine, which analyzes the user's emotional state.
[0577] Step 6:
[0578] If the server determines that there is a high possibility of fraud, it uses a voice alert library to issue a voice alert from the voice output device of the communication terminal.
[0579] Input: Fraud likelihood assessment results and user sentiment analysis results.
[0580] Data processing: If a high probability of fraud is determined, a voice alert is generated using the voice alert library (pyttsx3).
[0581] Output: Audio alert.
[0582] Specific operation: If it is determined that there is a high possibility of fraud, an audio alert will be emitted from the communication device's speaker saying, "Warning! Possible fraud."
[0583] Example 2
[0584] Next, a description will be given of Example 2 of Form Example 2. In the following description, the data processing device 12 will be referred to as a "server" and the smart device 14 will be referred to as a "terminal."
[0585] Conventional fraud prevention systems have low accuracy in detecting potential fraud, and are unable to completely eliminate the risk of users falling victim to fraud. In addition, they issue warnings without taking into account the user's emotional state, which can lead to users being unable to respond appropriately. This makes it difficult to prevent fraudulent incidents.
[0586] The specific processing by the specific processing unit 290 of the data processing device 12 in the second embodiment is realized by the following means.
[0587] In this invention, the server includes means for detecting incoming fraudulent calls, means for listening to conversations from a voice input device installed in the communication device, means for converting the conversation data into text and sending it to a generative AI model, means for issuing a voice alert from the voice input device if the generative AI model determines that fraud is suspected from the conversation data, means for an emotion analysis engine to analyze the user's tone of voice and speaking rate, and means for the generative AI model to issue a voice alert if the emotion analysis engine's reading exceeds a specific threshold. This makes it possible to detect the possibility of fraud with high accuracy and issue an appropriate warning according to the user's emotional state.
[0588] The "means for detecting incoming fraudulent calls" is a function for identifying and notifying a communication device of incoming calls that may be fraudulent.
[0589] A "voice input device" is a device that records conversations between a user and another person in real time and stores them in digital format.
[0590] "Means for converting conversation data into text and sending it to the generative AI model" is a function for converting voice data into text data and sending that text data to the generative AI model.
[0591] A "generative AI model" is an artificial intelligence system that analyzes input data, detects specific patterns and phrases, and determines the likelihood of fraud.
[0592] "Means for issuing audio alerts" refers to an audio notification function that warns users when a potential fraud is detected.
[0593] The "emotion analysis engine" is an engine that analyzes the user's tone of voice and speaking speed to determine the user's emotional state.
[0594] "When a specific threshold is exceeded" refers to when the emotion analysis engine analyzes the user's emotional state and the result exceeds a preset reference value.
[0595] A "communication device" is a device for making voice calls, such as a home telephone or smartphone.
[0596] MODE FOR CARRYING OUT THE INVENTION
[0597] The present invention is a system for detecting incoming fraudulent calls and issuing a warning to the user. A specific embodiment of this system will be described below.
[0598] Hardware and software used
[0599] Communication devices: devices used to make voice calls, such as home phones and smartphones
[0600] Audio input device: A microphone installed in the communication device
[0601] Server: A computer system for processing audio data and running generative AI models and sentiment analysis engines.
[0602] Generative AI model: An artificial intelligence system that analyzes conversation data and determines the likelihood of fraud
[0603] Emotion analysis engine: An engine that analyzes the user's tone of voice and speaking speed to determine their emotional state
[0604] Speech recognition software: Software for converting voice data into text data (e.g., Google Speech-to-Text API)
[0605] System Operation
[0606] 1. When a user starts a conversation on a communication device, the device's voice input device records the conversation in real time, and the recorded voice data is stored in digital format.
[0607] 2. The server uses speech recognition software to convert the collected voice data into text data, for example, "Hello, this is a bank representative" is converted into text "Hello, this is a bank representative."
[0608] 3. The server sends the converted text data to the generative AI model, which receives the data and prepares it for analysis.
[0609] 4. A generative AI model analyzes the conversation data to detect specific keywords and phrases (e.g., "bank account" or "password"), which can then be used to determine whether there is potential for fraud.
[0610] 5. If the generative AI model determines that there is a high possibility of fraud, the server will send instructions to the terminal and an audio alert will be issued from the communication device's speaker saying, "There is a possibility of fraud. Please be careful."
[0611] 6. The emotion analysis engine analyzes the user's tone of voice and speaking rate in real time. For example, if the user's voice suddenly gets higher in pitch or speaking rate, the emotion analysis engine will detect this.
[0612] 7. If the sentiment analysis engine determines that the user's emotions exceed a certain threshold, the generative AI model will issue an audio alert saying, "Possible scam. Please stay calm."
[0613] Specific examples
[0614] Example 1: When a user says "Please tell me your bank account information" over the phone, the generative AI model detects the keyword "bank account" and determines that it may be a scam. The speaker on the communication device issues a voice alert saying, "This may be a scam. Please be careful."
[0615] Example 2: If a user says "Hurry, tell me your password" over the phone, and the user's voice tone gets higher and the rate of speech increases, the emotion analysis engine determines that the user is experiencing strong emotions. The generative AI model issues a voice alert, warning, "This may be a scam. Please stay calm."
[0616] Prompt Sentence Examples
[0617] Prompt 1: "If someone calls and asks you to provide your bank account information, determine if this is a potential scam."
[0618] Prompt 2: "Please explain how the emotion engine determines if the user's voice tone gets higher and their speaking rate gets faster."
[0619] The system is designed to reduce the risk of users falling victim to phone scams by combining a generative AI model with a sentiment analysis engine to detect potential scams with greater accuracy and alert users.
[0620] The flow of the identification process in the second embodiment will be described with reference to FIG.
[0621] Step 1:
[0622] A user initiates a conversation on a communication device.
[0623] Specifically, a user initiates a call using a home phone or a smartphone. The input is the user's voice, and the output is voice data collected by a voice input device.
[0624] Step 2:
[0625] The terminal's voice input device listens to the conversation and collects voice data.
[0626] Specifically, the device's microphone records the conversation between the user and the other party in real time. The input is the user's voice and the output is audio data stored in digital format.
[0627] Step 3:
[0628] The server converts the audio data into text data.
[0629] Specifically, the server uses speech recognition software (e.g., Google Speech-to-Text API) to convert the collected voice data into text data. The input is digital voice data, and the output is text conversation data.
[0630] Step 4:
[0631] The server sends the transcribed conversation data to the generative AI model.
[0632] Specifically, the server sends the converted text data to the generative AI model. The input is text-format conversation data, and the output is the data sent to the generative AI model.
[0633] Step 5:
[0634] A generative AI model analyzes conversation data to determine the likelihood of fraud.
[0635] Specifically, the generative AI model analyzes conversation data to detect specific keywords or phrases (e.g., "bank account" or "password"). The input is textual conversation data, and the output is a judgment about the likelihood of fraud.
[0636] Step 6:
[0637] If there is a possibility of fraud, an audio alert will be issued through the device's voice input device.
[0638] Specifically, if the generative AI model determines that there is a high possibility of fraud, the server sends instructions to the terminal, and an audio alert is issued from the communication device's speaker saying, "There is a possibility of fraud. Please be careful." The input is the judgment result of the generative AI model, and the output is an audio alert.
[0639] Step 7:
[0640] The emotion analysis engine analyzes the user's tone of voice and speaking speed.
[0641] Specifically, the emotion analysis engine analyzes the user's tone of voice and speaking rate in real time. The input is the user's voice data, and the output is the analysis result regarding the user's emotional state.
[0642] Step 8:
[0643] If the sentiment analysis engine exceeds a certain threshold, the generative AI model will issue an audio alert.
[0644] Specifically, if the emotion analysis engine determines that the user's emotion exceeds a certain threshold, the generative AI model issues a voice alert saying, "This may be a scam. Please remain calm." The input is the analysis result of the emotion analysis engine, and the output is a voice alert.
[0645] (Application example 2)
[0646] Next, a description will be given of Application Example 2 of Form Example 2. In the following description, the data processing device 12 will be referred to as a "server" and the smart device 14 will be referred to as a "terminal."
[0647] In recent years, fraudulent phone scams have become more sophisticated, and many people have fallen victim to them. In particular, elderly people and people unfamiliar with technology are less vigilant against fraudulent phone calls and are more likely to fall victim to them. Furthermore, although there are systems to detect fraudulent calls, there are few warning systems that take into account the user's emotional state. This increases the risk that users may not be aware of the possibility of fraud and may become victims. Therefore, there is a need to provide a warning system that detects fraudulent calls and takes into account the user's emotional state.
[0648] The specific processing by the specific processing unit 290 of the data processing device 12 in the application example 2 is realized by the following means.
[0649] In this invention, the server includes means for detecting incoming fraudulent calls, means for listening to conversations from a voice input device installed in the smart device, means for converting the conversation data into text and sending it to a generative AI, means for issuing a voice alert from the voice input device when the generative AI determines that fraud is suspected from the conversation data, and means for analyzing the user's emotions using an emotion engine and issuing a warning when a specific threshold is exceeded. This makes it possible to detect fraudulent calls and issue a warning that takes into account the user's emotional state.
[0650] A "scam call" is a call made with the intent of defrauding someone of money or personal information through fraudulent means.
[0651] An "incoming call" refers to an incoming phone call.
[0652] A "smart device" is an electronic device with internet connectivity, such as a smartphone or tablet.
[0653] An "audio input device" is a device such as a microphone for converting audio into digital data.
[0654] "Conversation data" is data obtained by converting voice acquired through a voice input device into text.
[0655] "Generative AI" is a system that uses artificial intelligence to analyze data and make specific decisions or generate results.
[0656] A "voice alert" is a method of warning or notifying a user using voice.
[0657] An "emotion engine" is an algorithm or system for analyzing a user's emotional state from voice and text data.
[0658] A "specific threshold" is a numerical value or condition that serves as a standard when the emotion engine evaluates the user's emotional state.
[0659] A "warning" is a notification or alert that alerts the user.
[0660] The system for carrying out the present invention includes a series of means for detecting incoming fraudulent calls and issuing a warning to the user. A specific embodiment of the system will be described below.
[0661] First, a voice input device (microphone) is installed on a smart device (e.g., a smartphone or tablet). This voice input device captures the user's call content in real time. The captured voice data is converted into text data using voice recognition software (e.g., the speech_recognition library).
[0662] The transcribed conversation data is then sent to a generative AI model (e.g., a model using the Transformers library). This generative AI model analyzes the conversation data and determines whether it is likely to be fraudulent. If it is, audio alert software (e.g., the pyttsx3 library) is triggered to warn the user.
[0663] Furthermore, an emotion engine (e.g., an emotion analysis model using the transformers library) is used to analyze the user's emotional state. The emotion engine evaluates the user's tone of voice, speech rate, etc., and issues a warning if certain thresholds are exceeded. This warning is also notified to the user as an audio alert.
[0664] For example, if a user calls and asks, "Please tell me your bank account information," the generative AI model will detect the possibility of fraud and issue a voice alert. If the user is feeling angry or anxious, the emotion engine will detect this and issue a voice alert.
[0665] An example of a prompt is as follows:
[0666] User says: "What bank account information do I need?"
[0667] Prompt the generative AI model: "Is this call potentially fraudulent?"
[0668] In this way, a system can be realized that can detect fraudulent calls and warn users based on their emotional state.
[0669] The flow of the specific processing in the application example 2 will be described with reference to FIG.
[0670] Step 1:
[0671] A user starts a call on a smart device. The voice input device (microphone) of the smart device captures the call content in real time. The input is the user's voice data, and the output is a stream of voice data.
[0672] Step 2:
[0673] The device uses speech recognition software (speech_recognition library) to convert the captured voice data into text data. The input is a stream of voice data, and the output is transcribed conversation data.
[0674] Step 3:
[0675] The device sends the transcribed conversation data to a generative AI model (a model using the Transformers library). The generative AI model analyzes the conversation data and determines whether there is a possibility of fraud. The input is the transcribed conversation data, and the output is a judgment about the possibility of fraud.
[0676] Step 4:
[0677] The server receives the judgment result of the generative AI model, and if it determines that there is a possibility of fraud, it launches software (pyttsx3 library) to issue an audio alert. The input is the judgment result regarding the possibility of fraud, and the output is an audio alert.
[0678] Step 5:
[0679] The device analyzes the user's emotional state using an emotion engine (an emotion analysis model using the Transformers library). The emotion engine evaluates the user's tone of voice, speech rate, etc., and issues a warning if certain thresholds are exceeded. The input is transcribed conversation data, and the output is an evaluation result of the user's emotional state.
[0680] Step 6:
[0681] The server receives the evaluation results of the emotion engine, and if a certain threshold is exceeded, it launches software (pyttsx3 library) to issue an audio alert. The input is the evaluation result on the emotional state, and the output is an audio alert.
[0682] Step 7:
[0683] The user receives an audio alert and perceives a warning about potential fraud or emotional state. The input is the audio alert and the output is the user's perception.
[0684] Example 3
[0685] Next, a description will be given of a third embodiment of the third embodiment. In the following description, the data processing device 12 will be referred to as a "server" and the smart device 14 will be referred to as a "terminal."
[0686] Conventional fraud detection systems have the problem of low accuracy because they only analyze conversation data when determining the possibility of fraud. Also, because they do not take into account the user's emotional state, there is a high risk of overlooking possible fraud. Furthermore, there are limited means of notifying users of possible fraud, so users may not be aware of the risk of fraud.
[0687] The specific processing by the specific processing unit 290 of the data processing device 12 in the third embodiment is realized by the following means.
[0688] In this invention, the server includes means for detecting incoming fraudulent calls, means for listening to conversations from a voice input device installed in the communication terminal, means for converting the conversation data into text and transmitting it to the generative artificial intelligence, means for issuing a voice alert from the voice input device when the generative artificial intelligence determines that fraud is suspected from the conversation data, means for an emotion analysis engine to analyze the user's emotions, and means for the generative artificial intelligence to integrate information from the emotion analysis engine and the conversation data to determine the possibility of fraud. This makes it possible to detect the possibility of fraud with high accuracy and notify the user promptly.
[0689] A "means for detecting incoming fraudulent calls" is a device or software that monitors a communication line and identifies incoming calls that may be fraudulent.
[0690] The "voice input device installed on a communication terminal" is a microphone for collecting voice that is attached to a communication device such as a smartphone or home telephone.
[0691] "Means for converting conversation data into text and sending it to the generative AI" refers to software or a device for converting voice data into text data and sending that text data to the generative AI.
[0692] "Generative AI" is an AI model that uses natural language processing technology to analyze text data and detect specific patterns and phrases.
[0693] The "means for issuing an audio alert" is an audio output device that issues a warning to the user when the generative artificial intelligence detects the possibility of fraud.
[0694] An "emotion analysis engine" is software or a device that analyzes a user's voice or text data to identify their emotional state.
[0695] The "means for determining the likelihood of fraud" is an algorithm or software that uses generative artificial intelligence to integrate information from the conversational data and sentiment analysis engine to assess the likelihood of fraud.
[0696] The present invention is a system for detecting incoming fraudulent calls and notifying users of the possibility of fraud. A specific embodiment of this system will be described below.
[0697] First, when a user is using a communication terminal (such as a smartphone or home phone), a means for detecting incoming fraudulent calls is activated. This means is a device or software that monitors communication lines and identifies incoming calls that may be fraudulent.
[0698] Next, a voice input device (microphone) installed on the communication terminal listens to the conversation between the user and the fraudulent caller. This voice data is converted into text data by transcribing the conversation data and sending it to generative artificial intelligence. This conversion uses voice recognition technology. Specifically, voice recognition services such as Google Speech-to-Text API and IBM Watson Speech to Text can be used.
[0699] The server analyzes the text data using generative artificial intelligence (e.g., natural language processing models such as BERT or GPT-3 (registered trademark)). The generative artificial intelligence detects specific phrases and patterns that may be fraudulent and evaluates the results. For example, if the phrase "Please transfer money" is included, it will determine that there is a high possibility of fraud.
[0700] Furthermore, the server uses an emotion analysis engine (e.g., IBM Watson or Microsoft® Azure® emotion analysis API) to analyze the user's emotions. It analyzes the user's voice and text data to identify emotional states such as anxiety or tension. The emotion data obtained by the emotion analysis engine is sent to the generative artificial intelligence.
[0701] Generative AI combines information from the sentiment analysis engine with conversational data to more accurately assess the likelihood of fraud. For example, if a message says "Please transfer money" and the user's emotions are anxious, it will conclude that there is a very high likelihood of fraud.
[0702] Finally, if the server determines that the message is likely to be fraudulent, it notifies the user of the result. The notification is made by issuing an audio alert. Specifically, an audio alert is issued from the communication terminal's speaker saying, "This message may be fraudulent. Please be careful."
[0703] As a concrete example, consider the case where a user receives a message saying, "Please transfer money." This message is listened to by the voice input device of the communication terminal and converted into text data. The server analyzes this text data using generative artificial intelligence and determines that there is a high possibility of fraud. At the same time, the emotion analysis engine detects the user's anxiety and sends this information to the generative artificial intelligence. The server combines this information and concludes that there is a very high possibility of fraud. Finally, the server sends a voice alert to the user's communication terminal saying, "This message may be fraudulent. Please be careful."
[0704] Example prompt sentence:
[0705] "Please determine if the following message is a potential scam. Message: 'Please transfer money.'"
[0706] In this way, the server, the terminal, and the user work together to detect the possibility of fraud with high accuracy and notify the user promptly. The flow of the identification process in the third embodiment will be described with reference to FIG.
[0707] Step 1:
[0708] The user enters a message.
[0709] The user inputs the received message using a communication terminal (such as a smartphone or home phone). For example, the user inputs a message such as "Please transfer money." The input message is saved as voice data on the terminal.
[0710] Step 2:
[0711] The device sends a message to the server.
[0712] The terminal transmits the voice data input by the user to the server. At this time, the terminal transmits the voice data to the server via the Internet. The input is the voice data, and the output is the transfer of the voice data to the server.
[0713] Step 3:
[0714] The server converts the voice data into text data.
[0715] The server converts the received voice data into text data using voice recognition technology. Specifically, it uses voice recognition services such as Google Speech-to-Text API and IBM Watson Speech to Text. The input is voice data, and the output is text data.
[0716] Step 4:
[0717] The server analyzes the text data using generative artificial intelligence.
[0718] The server analyzes the text data using generative artificial intelligence (e.g., natural language processing models such as BERT or GPT-3). The generative artificial intelligence detects specific phrases or patterns that may be fraudulent and evaluates the results. The input is the text data, and the output is an evaluation result regarding the likelihood of fraud.
[0719] Step 5:
[0720] The server analyzes the user's emotions using a sentiment analysis engine.
[0721] The server uses an emotion analysis engine (for example, IBM Watson or Microsoft Azure's emotion analysis API) to analyze the user's emotions. It analyzes the user's voice and text data to identify emotional states such as anxiety or tension. The input is voice data or text data, and the output is emotion data.
[0722] Step 6:
[0723] The server combines the results of the generative artificial intelligence and sentiment analysis engine to determine the likelihood of fraud.
[0724] The server combines the analysis results from the generative AI with the emotional data from the emotion analysis engine. This allows for a more accurate assessment of the likelihood of fraud. For example, if the message is "Please transfer money" and the user's emotion is anxiety, it will conclude that there is a very high likelihood of fraud. The inputs are the analysis results and emotional data, and the output is a final assessment of the likelihood of fraud.
[0725] Step 7:
[0726] The server notifies the user of possible fraud.
[0727] If the server determines that the message is likely to be fraudulent, it notifies the user of the result. The notification is performed by issuing an audio alert. Specifically, an audio alert saying "This message may be fraudulent. Please be careful" is issued from the speaker of the communication terminal. The input is the final assessment of the likelihood of fraud, and the output is the audio alert.
[0728] (Application example 3)
[0729] Next, a description will be given of Application Example 3 of Form Example 3. In the following description, the data processing device 12 will be referred to as a "server" and the smart device 14 will be referred to as a "terminal."
[0730] Conventional fraud detection systems only detect specific phrases and patterns and do not take into account the user's emotional state, making it difficult to accurately determine the likelihood of fraud. Furthermore, there is a lack of a way to provide users with appropriate warnings when fraud is likely.
[0731] The identification processing by the identification processing unit 290 of the data processing device 12 in Application Example 3 is realized by the following means. In this invention, the server includes means for detecting incoming fraudulent calls, means for listening to conversations from a smartphone application or a microphone installed on a home phone, means for converting the conversation data into text and sending it to the generative AI, means for issuing an audio alert from the microphone if the generative AI determines that fraud is suspected based on the conversation data, means for analyzing the user's emotions using an emotion engine, means for determining the possibility of fraud by combining information from the emotion engine with the conversation data analyzed by the generative AI, and means for displaying a warning to the user if there is a high possibility of fraud. This makes it possible to more accurately determine the possibility of fraud and provide the user with an appropriate warning.
[0732] "Means for detecting incoming fraudulent calls" refers to devices or software that monitor telephone lines or communications networks to detect incoming calls that may be fraudulent.
[0733] A "microphone installed in a smartphone application or home phone" is a voice input device attached to a smartphone or home phone to collect audio.
[0734] A "means for listening to a conversation" is a device or software that analyzes audio data collected through a microphone and recognizes the content of the conversation.
[0735] "Means for converting conversation data into text and sending it to generative AI" refers to devices or software that convert voice data into text data and send that text data to generative AI.
[0736] "Generative AI" is an artificial intelligence model trained to perform a specific task, in this case to determine the likelihood of fraud.
[0737] "Means for issuing an audio alert from the microphone when it is determined that fraud is suspected" refers to a device or software that issues an alarm or message through the microphone when the generative AI detects the possibility of fraud.
[0738] An "emotion engine" is software or a device for analyzing a user's emotional state from their voice or text data.
[0739] "Means for analyzing user emotions" refers to devices or software that use an emotion engine to analyze the user's emotional state.
[0740] "Means for determining the possibility of fraud by combining information from an emotion engine with conversational data analyzed by a generative AI" refers to devices or software that integrate emotional information obtained from an emotion engine with the results of analysis of conversational data by a generative AI, and make a comprehensive determination of the possibility of fraud.
[0741] "Means for displaying a warning to the user in the event of a high likelihood of fraud" refers to a device or software that provides a visual or audible warning to the user in the event that a high likelihood of fraud is determined.
[0742] The system for carrying out the present invention includes a set of means for detecting an incoming fraudulent call and providing a warning to the user. Specific embodiments will be described below.
[0743] First, the server is equipped with a means for detecting incoming fraudulent calls. This means is a device or software that monitors telephone lines and communication networks to detect incoming calls that may be fraudulent. For example, it is possible to register specific phone numbers or callers on a blacklist and detect fraudulent calls based on that.
[0744] Next, the device (smartphone application or home phone) is equipped with a microphone to listen to the conversation. This microphone collects the voice during the call and sends it to a server as conversation data. The conversation data is converted into text using voice recognition technology and sent to the generative AI.
[0745] Generative AI analyzes conversation data to determine the likelihood of fraud. It uses models trained to detect specific phrases and patterns. For example, it can detect typical phrases for "bank transfer fraud" and "it's me" fraud.
[0746] Furthermore, the emotion engine analyzes the user's emotions. The emotion engine is software that analyzes the user's emotional state from voice and text data. The emotional information obtained from the emotion engine is integrated with the results of the generative AI analysis of the conversation data to comprehensively determine the possibility of fraud.
[0747] If a fraudulent activity is deemed likely, the device will display a warning to the user. This warning can be visual or audible, for example, by displaying a warning message on the smartphone screen or by issuing an audio alert.
[0748] For example, consider the following user input:
[0749] "Your account has been frozen. Please transfer the funds immediately."
[0750] "Hey, I need money right now."
[0751] These prompts can be fed into a generative AI model that analyzes the likelihood of fraud and works in conjunction with an emotion engine to display a warning to the user.
[0752] The flow of the specific processing in Application Example 3 will be described with reference to FIG.
[0753] Step 1:
[0754] The server monitors telephone lines and communication networks to detect incoming calls that may be fraudulent. The input is a signal from the telephone line or communication network, and the output is a flag indicating an incoming fraudulent call. Specifically, it registers specific phone numbers and callers on a blacklist and detects fraudulent calls based on that.
[0755] Step 2:
[0756] The device (smartphone application or home phone) collects the voice during a call through a microphone. The input is the voice signal during the call, and the output is the voice data. Specifically, the microphone converts the voice into a digital signal and sends the data to a server.
[0757] Step 3:
[0758] The server uses voice recognition technology to convert the voice data into text. The input is voice data and the output is text data. Specifically, the voice recognition software analyzes the voice data and converts it into corresponding text.
[0759] Step 4:
[0760] The server sends text data to a generative AI to analyze the likelihood of fraud. The input is text data, and the output is a score or flag indicating the likelihood of fraud. Specifically, the generative AI model detects specific phrases and patterns to determine the likelihood of fraud.
[0761] Step 5:
[0762] The server analyzes the user's emotions using an emotion engine. The input is text data, and the output is data indicating the user's emotional state. Specifically, the emotion engine analyzes the text data and identifies the user's emotional state.
[0763] Step 6:
[0764] The server combines information from the emotion engine with conversation data analyzed by the generative AI to comprehensively determine the likelihood of fraud. The input is emotional state data and a score or flag indicating the likelihood of fraud, and the output is a final flag indicating the likelihood of fraud. Specifically, the server integrates emotional information and the analysis results of the generative AI to determine the likelihood of fraud with high accuracy.
[0765] Step 7:
[0766] If it determines that there is a high possibility of fraud, the device will display a warning to the user. The input is a flag indicating the final possibility of fraud, and the output is a warning message or audio alert. Specifically, a warning message is displayed on the smartphone screen or an audio alert is issued.
[0767] The specific processing unit 290 transmits the result of the specific processing to the smart device 14. In the smart device 14, the control unit 46A causes the output device 40 to output the result of the specific processing. The microphone 38B acquires audio indicating a user input regarding the result of the specific processing. The control unit 46A transmits audio data indicating the user input acquired by the microphone 38B to the data processing device 12. In the data processing device 12, the specific processing unit 290 acquires the audio data.
[0768] The data generation model 58 is a so-called generative AI (Artificial Intelligence). An example of the data generation model 58 is ChatGPT (registered trademark) (Internet search engine).<URL: https: / / openai.com / blog / chatgpt> ) and other generation AIs. The data generation model 58 is obtained by performing deep learning on a neural network. A prompt including an instruction is input to the data generation model 58, and inference data such as voice data indicating voice, text data indicating text, and image data indicating an image is also input. The data generation model 58 performs inference on the input inference data in accordance with the instruction indicated by the prompt, and outputs the inference result in a data format such as voice data and text data. Here, inference refers to, for example, analysis, classification, prediction, and / or summarization.
[0769] Another example of generative AI is Gemini (registered trademark) (Internet search engine). <url: https: gemini.google.com ?hl="ja">) are listed.
[0770] In the above embodiment, an example in which the specific process is performed by the data processing device 12 has been given, but the technology of the present disclosure is not limited to this, and the specific process may be performed by the smart device 14.
[0771] [Second embodiment]
[0772] FIG. 3 shows an example of the configuration of a data processing system 210 according to the second embodiment.
[0773] 3, the data processing system 210 includes the data processing device 12 and smart glasses 214. An example of the data processing device 12 is a server.
[0774] The data processing device 12 includes a computer 22, a database 24, and a communication I / F 26. The computer 22 is an example of a "computer" according to the technology of the present disclosure. The computer 22 includes a processor 28, a RAM 30, and a storage 32. The processor 28, the RAM 30, and the storage 32 are connected to a bus 34. The database 24 and the communication I / F 26 are also connected to the bus 34. The communication I / F 26 is connected to a network 54. Examples of the network 54 include a WAN (Wide Area Network) and / or a LAN (Local Area Network).
[0775] The smart glasses 214 include a computer 36, a microphone 238, a speaker 240, a camera 42, and a communication I / F 44. The computer 36 includes a processor 46, a RAM 48, and a storage 50. The processor 46, the RAM 48, and the storage 50 are connected to a bus 52. The microphone 238, the speaker 240, and the camera 42 are also connected to the bus 52.
[0776] The microphone 238 receives instructions and the like from the user 20 by receiving voice uttered by the user 20. The microphone 238 captures the voice uttered by the user 20, converts the captured voice into audio data, and outputs it to the processor 46. The speaker 240 outputs audio in accordance with instructions from the processor 46.
[0777] Camera 42 is a small digital camera equipped with an optical system including a lens, aperture, and shutter, and an imaging element such as a CMOS (Complementary Metal-Oxide-Semiconductor) image sensor or a CCD (Charge Coupled Device) image sensor, and captures images of the surroundings of user 20 (for example, an imaging range defined by an angle of view equivalent to the field of vision of a typical healthy person).
[0778] The communication I / F 44 is connected to a network 54. The communication I / Fs 44 and 26 are responsible for the exchange of various information between the processor 46 and the processor 28 via the network 54. The exchange of various information between the processor 46 and the processor 28 using the communication I / Fs 44 and 26 is carried out in a secure state.
[0779] Fig. 4 shows an example of the main functions of the data processing device 12 and the smart glasses 214. As shown in Fig. 4, in the data processing device 12, a specific process is performed by the processor 28. A specific process program 56 is stored in the storage 32.
[0780] The specific processing program 56 is an example of a "program" according to the technology of the present disclosure. The processor 28 reads the specific processing program 56 from the storage 32 and executes the read specific processing program 56 on the RAM 30. The specific processing is realized by the processor 28 operating as a specific processing unit 290 in accordance with the specific processing program 56 executed on the RAM 30.
[0781] The storage 32 stores a data generation model 58 and an emotion identification model 59. The data generation model 58 and the emotion identification model 59 are used by the identification processing unit 290.
[0782] In the smart glasses 214, the reception output process is performed by the processor 46. A reception output program 60 is stored in the storage 50. The processor 46 reads the reception output program 60 from the storage 50 and executes the read reception output program 60 on the RAM 48. The reception output process is realized by the processor 46 operating as the control unit 46A in accordance with the reception output program 60 executed on the RAM 48.
[0783] Next, the specific processing by the specific processing unit 290 of the data processing device 12 will be described.
[0784] "Example 1"
[0785] One embodiment of the present invention uses a smartphone application. In this case, the application installed on the smartphone detects an incoming call and listens to the conversation through the microphone. The listened conversation data is transcribed and sent to a generative AI. The generative AI determines from the conversation data that fraud is suspected, and if fraud is suspected, issues a voice alert from the smartphone speaker.
[0786] "Example 2"
[0787] Another embodiment of the present invention involves installing a microphone on a home phone. In this case, the microphone on the home phone listens to the phone conversation, transcribes the conversation data, and sends it to a generative AI. The generative AI determines from the conversation data that fraud is suspected, and if fraud is suspected, issues an audio alert from the home phone's speaker.
[0788] "Example 3"
[0789] In a further embodiment of the present invention, the generative AI detects specific phrases or patterns that indicate the possibility of special fraud. Specifically, the generative AI learns typical phrases for "bank transfer fraud" and typical patterns for "it's me" fraud, and by detecting these, it determines the possibility of fraud.
[0790] The processing flow of each embodiment will be described below.
[0791] "Example 1"
[0792] Step 1: An application installed on the smartphone detects an incoming call.
[0793] Step 2: The application listens to your conversation through your smartphone's microphone.
[0794] Step 3: The conversation data is transcribed and sent to the generative AI.
[0795] Step 4: The generative AI determines that fraud is suspected based on the conversation data.
[0796] Step 5: If fraud is suspected, an audio alert will be issued through the smartphone speaker.
[0797] "Example 2"
[0798] Step 1: A microphone installed in a home phone listens to the phone conversation.
[0799] Step 2: The conversation data is transcribed and sent to the generative AI.
[0800] Step 3: The generative AI determines that fraud is suspected based on the conversation data.
[0801] Step 4: If fraud is suspected, an audio alert will be issued through the home phone speaker.
[0802] "Example 3"
[0803] Step 1: The generative AI learns specific phrases or patterns that indicate potential fraud.
[0804] Step 2: The generative AI detects specific phrases or patterns it has learned to identify suspected fraud in the conversation data.
[0805] Step 3: If fraud is suspected, an audio alert will be issued.
[0806] Example 1
[0807] Next, a description will be given of Example 1 of Form Example 1. In the following description, the data processing device 12 will be referred to as a "server" and the smart glasses 214 will be referred to as a "terminal."
[0808] In recent years, telephone fraud has been on the rise, and special frauds targeting the elderly in particular have become a social problem. Conventional countermeasures require users to identify fraud themselves when receiving a fraudulent call, but this is becoming more difficult as fraud methods become more sophisticated. Therefore, there is a need for a system that automatically detects the possibility of fraud and issues a warning when a user receives a fraudulent call.
[0809] The identification process by the identification processing unit 290 of the data processing device 12 in Example 1 is realized by the following means. In this invention, the server includes a means for detecting an incoming fraudulent call, a means for listening to the conversation from a microphone installed in the communication terminal, a means for converting the conversation data into text and sending it to the generative AI model, and a means for issuing an audio alert from the speaker of the communication terminal when the generative AI model determines that fraud is suspected based on the conversation data. This makes it possible to automatically detect the possibility of fraud when a user receives a fraudulent call and immediately issue a warning.
[0810] "Means for detecting incoming fraudulent calls" refers to a function or device for detecting incoming calls in a communication terminal.
[0811] "Means for listening to conversations from a microphone installed in a communication terminal" refers to a function or device that uses a microphone built into or connected to a communication terminal to collect voices during a call in real time.
[0812] "Means for converting conversation data into text and sending it to the generative AI model" refers to a function or device for converting collected voice data into text data and sending that text data to the generative AI model.
[0813] A "generative AI model" is an artificial intelligence model that analyzes input data and detects specific patterns or phrases.
[0814] "Means for issuing an audio alert from the speaker of the communication device when it is determined that fraud is suspected" refers to a function or device for playing an alarm sound or message through the speaker of the communication device when the generative AI model detects the possibility of fraud.
[0815] "Specific phrases or patterns that indicate the possibility of fraud" are specific words or sentence structures related to fraudulent acts, and the generative AI model's detection of these serves as the basis for determining the possibility of fraud.
[0816] An "audio alert" is a message or sound that provides an audio warning to the user.
[0817] The present invention is a system that automatically detects fraudulent calls and issues a warning to the user using an application installed on a communication terminal. A specific embodiment of this system will be described below.
[0818] First, the user installs a dedicated application on their communication device. This application has the function of detecting incoming calls. Specifically, it uses the communication device's telephone API to catch incoming call events.
[0819] When a user answers a call, the device listens to the conversation in real time through the microphone. The application runs in the background and captures the conversation. The captured conversation data is transcribed using the Google Speech-to-Text API. The voice data is sent to the API and received as text data.
[0820] The transcribed conversation data is sent to a server via the internet. The HTTPS protocol is used for transmission to ensure data security. The server inputs the received conversation data into a generative AI model (e.g., OpenAI's GPT-4). The generative AI model uses prompts to analyze the possibility of fraud. Specific examples of prompts are as follows:
[0821] Analyze the conversation data below to determine if it is a potential scam.
[0822] Conversation data: "Your bank account has been compromised. Please provide your account number to verify."
[0823] The generative AI model analyzes the conversation data and determines whether fraud is suspected. If so, the server sends that information back to the device, again using the HTTPS protocol.
[0824] The device analyzes the results received from the server, and if fraud is suspected, it issues an audio alert from the device's speaker. Specifically, it plays an audio message such as "There is a possibility of fraud. Please be careful." This alert makes the user aware of the possibility of fraud and allows them to take appropriate action.
[0825] As a concrete example, consider the case where a user receives a phone call saying, "Your bank account has been fraudulently used. Please provide your account number to verify." This conversation is picked up by a microphone and transcribed. The transcribed data becomes, "Your bank account has been fraudulently used. Please provide your account number to verify." This data is sent to a generative AI model, which determines that there is a high possibility of fraud. The server sends the result back to the device, which then issues an audio alert. The user hears this alert and becomes aware of the possibility of fraud.
[0826] In this way, the present invention can automatically detect possible fraud and immediately issue a warning when a user receives a fraudulent call.
[0827] The flow of the identification process in the first embodiment will be described with reference to FIG.
[0828] Step 1: Detecting an incoming call
[0829] The terminal detects incoming calls through an application installed on the communication terminal. Specifically, it uses the communication terminal's telephone API to catch incoming call events. The input is the incoming call signal, and the output is the incoming call detection event.
[0830] Step 2: Listen to the conversation
[0831] When a user answers a call, the device listens to the conversation in real time through the microphone. The application runs in the background and captures the conversation. The input is the voice data during the call, and the output is the collected voice data.
[0832] Step 3: Transcribe the conversation data
[0833] The device uses the Google Speech-to-Text API to transcribe the conversation data it hears. It sends the speech data to the API and receives it as text data. The input is the collected speech data, and the output is the transcribed text data.
[0834] Step 4: Send the transcribed data
[0835] The terminal sends the transcribed conversation data to the server. The HTTPS protocol is used for transmission to ensure data security. The input is the transcribed text data, and the output is the data transmission event to the server.
[0836] Step 5: Fraud detection
[0837] The server inputs the received conversation data into the generative AI model, which analyzes the possibility of fraud using prompt sentences. Specific examples of prompt sentences are as follows:
[0838] Analyze the conversation data below to determine if it is a potential scam.
[0839] Conversation data: "Your bank account has been compromised. Please provide your account number to verify."
[0840] The input is transcribed text data, and the output is a judgment result regarding the possibility of fraud.
[0841] Step 6: Receiving the results
[0842] The server receives the judgment result from the generative AI model and returns the result to the terminal. The return is again via HTTPS. The input is the judgment result from the generative AI model, and the output is a data transmission event to the terminal.
[0843] Step 7: Generate an audio alert
[0844] The terminal analyzes the judgment result received from the server, and if fraud is suspected, it issues an audio alert from the communication terminal's speaker. Specifically, it plays an audio message such as "There is a possibility of fraud. Please be careful." The input is the judgment result from the server, and the output is the generation of an audio alert.
[0845] (Application example 1)
[0846] Next, a description will be given of Application Example 1 of Form Example 1. In the following description, the data processing device 12 will be referred to as a "server" and the smart glasses 214 will be referred to as a "terminal."
[0847] In recent years, telephone fraud has been on the rise, and special fraud targeting the elderly in particular has become a social problem. Conventional fraud prevention measures have made it difficult for users to recognize the possibility of fraud, making it difficult to prevent damage before it occurs. Therefore, there is a need for a system that analyzes telephone conversations in real time and issues an immediate warning if there is a possibility of fraud.
[0848] The specific processing by the specific processing unit 290 of the data processing device 12 in the application example 1 is realized by the following means.
[0849] In this invention, the server includes means for detecting incoming fraudulent calls, means for listening to conversations from a voice input device installed in the communication terminal, means for converting the conversation data into text and sending it to the generative AI, means for issuing an audio alert from the voice output device if the generative AI determines that fraud is suspected from the conversation data, means for the generative AI to detect specific phrases or patterns that indicate the possibility of fraud, and means for the audio alert to alert the user to the possibility of fraud. This allows the user to immediately become aware of the possibility of fraud and prevent damage before it occurs.
[0850] A "scam call" is a call made with the intent of defrauding someone of money or personal information through fraudulent means.
[0851] An "incoming call" refers to an incoming phone call.
[0852] A "communication terminal" is a device for sending and receiving voice and data, and includes smartphones and home telephones.
[0853] An "audio input device" is a device for converting audio into digital data, and a microphone is an example of this.
[0854] "Conversation data" is data obtained by converting speech acquired by a speech input device into text.
[0855] "Generative AI" is artificial intelligence that analyzes input data and detects specific patterns and phrases.
[0856] An "audio output device" is a device for outputting digital data as audio, and a speaker is an example of this.
[0857] An "audio alert" is a warning sound or message emitted from an audio output device.
[0858] A "particular phrase or pattern" is a characteristic word or sentence structure that indicates possible fraud.
[0859] "User" refers to any individual or organization that uses this system.
[0860] A system for implementing this invention detects incoming fraudulent calls and listens to the conversation through a voice input device installed on a communication terminal. The conversation data is transcribed in real time and sent to a generative AI. The generative AI detects specific phrases or patterns from the conversation data that indicate the possibility of fraud, and if fraud is suspected, issues a voice alert from a voice output device.
[0861] Hardware and software used
[0862] Hardware: Communication devices such as smartphones and home phones, microphones (audio input devices), speakers (audio output devices)
[0863] software:
[0864] SpeechRecognition: A Python library for converting speech to text
[0865] requests: A Python library for sending HTTP requests
[0866] Generative AI: Artificial intelligence that analyzes text data using external APIs
[0867] System Operation
[0868] 1. Incoming call detection: When the communication terminal detects an incoming call, it starts recording the conversation.
[0869] 2. Conversation transcription: The recorded audio data is transcribed in real time using the SpeechRecognition library.
[0870] 3. Generative AI Analysis: The transcribed conversation data is sent via HTTP requests to a generative AI, which detects specific phrases or patterns that indicate potential fraud.
[0871] 4. Alert: If the generative AI detects a potential fraud, an audio alert will be emitted from the device's speaker to alert the user to the potential fraud.
[0872] Specific examples
[0873] For example, if a user receives a phone call and the caller says, "Your bank account has been compromised. Please provide me with your account information immediately," the system will transcribe this conversation in real time and send it to a generative AI that will detect that this phrase matches certain patterns that indicate potential fraud and immediately trigger an audio alert.
[0874] Prompt Sentence Examples
[0875] An example of a prompt for a generative AI model is:
[0876] "Determine whether the following text is a potential scam: 'Your bank account has been compromised. Please provide your account information immediately.'"
[0877] In this way, users can be made aware of potential fraud immediately and prevent damage before it occurs.
[0878] The flow of the specific processing in the application example 1 will be described with reference to FIG.
[0879] Step 1: Detect incoming calls
[0880] The communication terminal detects an incoming call. The input is the incoming call signal, and the output is a trigger to start recording. Specifically, the phone application of the communication terminal detects the incoming call and starts the recording module.
[0881] Step 2: Record the conversation
[0882] The audio input device (microphone) of the communication terminal records the conversation. The input is the audio data during the call, and the output is the recorded audio file. Specifically, the recording module captures the audio during the call and saves it as an audio file.
[0883] Step 3: Transcribe the conversation
[0884] The recorded audio data is transcribed using the SpeechRecognition library. The input is an audio file, and the output is text data. Specifically, the speech recognition engine analyzes the audio file and generates the corresponding text.
[0885] Step 4: Generative AI analysis
[0886] Transcribed text data is sent to the generative AI via an HTTP request. The input is text data, and the output is a flag indicating the possibility of fraud. Specifically, a request containing the text data is sent to the generative AI's API endpoint, and the analysis results are received.
[0887] Step 5: Determine the likelihood of fraud
[0888] Generative AI analyzes text data to detect specific phrases or patterns that indicate potential fraud. The input is text data, and the output is a flag indicating potential fraud. Specifically, generative AI analyzes text data and flags cases where there is a high probability of fraud.
[0889] Step 6: Send an alert
[0890] If a possible fraud is detected, an audio alert is issued from the audio output device (speaker) of the communication terminal. The input is a flag indicating the possible fraud, and the output is an audio alert. Specifically, the audio alert module receives the flag indicating the possible fraud and plays an audio message to warn the user.
[0891] In this way, users can be made aware of potential fraud immediately and prevent damage before it occurs.
[0892] Example 2
[0893] Next, a description will be given of Example 2 of Form Example 2. In the following description, the data processing device 12 will be referred to as a "server" and the smart glasses 214 will be referred to as a "terminal."
[0894] In recent years, fraudulent phone calls have become more sophisticated, and many people have fallen victim to them. In particular, there has been an increase in special frauds targeting the elderly, and effective countermeasures against this are needed. Conventional countermeasures lack a system that can detect fraudulent phone calls in real time and issue warnings to users, making it difficult to prevent damage before it occurs.
[0895] The specific processing by the specific processing unit 290 of the data processing device 12 in the second embodiment is realized by the following means.
[0896] In this invention, the server includes means for detecting incoming fraudulent calls, means for collecting conversations from a voice input device installed in the communication device, means for converting the conversation data into text data, means for transmitting the text data to the generative AI model, and means for issuing a voice alert from the voice input device when the generative AI model determines that fraud is suspected from the conversation data. This makes it possible to detect fraudulent calls in real time and issue an immediate warning to the user.
[0897] The "means for detecting incoming fraudulent calls" is a function for identifying and detecting incoming calls that may be fraudulent calls to a communication device.
[0898] An "audio input device" is a microphone or other audio capture device installed on a communications device to capture a user's speech.
[0899] "Means for converting speech data into text data" refers to speech recognition software or algorithms used to analyze collected speech data and convert it into text format.
[0900] A "generative AI model" is an artificial intelligence model that analyzes input text data and determines the likelihood of fraud based on specific patterns and phrases.
[0901] "Audio alerting means" refers to a feature that plays an audio message to warn the user if fraud is suspected.
[0902] The present invention is a system for detecting incoming fraudulent calls in real time and issuing a warning to the user. A specific embodiment of this system will be described below.
[0903] 1. System Configuration
[0904] The system consists of the following major components:
[0905] How to detect incoming fraudulent calls
[0906] Voice input device
[0907] A means of converting conversation data into text data
[0908] Generative AI Models
[0909] A means of issuing an audio alert
[0910] 2. Hardware and Software Use
[0911] The voice input device is a microphone installed in a home telephone, which collects the user's speech in real time.
[0912] The means of converting conversation data into text data is through speech recognition software (e.g., Google Speech-to-Text API), which analyzes collected voice data and converts it into text data.
[0913] The generative AI model uses advanced artificial intelligence models such as OpenAI GPT-4, which analyzes input text data and determines the likelihood of fraud.
[0914] The audio alert method uses the speaker on the home phone, which will provide an audio alert if fraud is suspected.
[0915] 3. System Operation
[0916] When a user starts a conversation on their home phone, a voice input device collects the conversation. The collected voice data is converted into text data by speech recognition software. This text data is sent to a server and input into a generative AI model. The generative AI model analyzes the conversation data and, if it determines that fraud is suspected, the server sends an alert signal to the home phone. The home phone's speaker emits an audio alert to warn the user.
[0917] 4. Specific Examples
[0918] For example, imagine a user speaking on their home phone, "Hello, your bank account is at risk. Immediate action is required." A voice input device collects this speech, which speech recognition software converts into text data. This text data is sent to a server and analyzed by a generative AI model. If the generative AI model determines that there is a high possibility of fraud, the server sends an alert signal to the home phone. The speaker on the home phone issues a voice alert saying, "Possible fraud. Be careful."
[0919] 5. Examples of prompts
[0920] "Please analyze the conversation data below and determine whether you suspect fraud. If so, please explain why.
[0921] Conversation data: 'Hello, your bank account has been compromised. We need your immediate attention.'"
[0922] In this way, users can be made aware of their fraud risk in real time.
[0923] The flow of the identification process in the second embodiment will be described with reference to FIG.
[0924] Step 1:
[0925] Audio collection
[0926] Terminal: A voice input device (microphone) installed in a home phone collects the user's conversation in real time.
[0927] Input: User's voice conversation
[0928] Output: Digital audio data
[0929] What it does: When a user starts a phone conversation, the microphone automatically captures the audio and stores it as digital audio data.
[0930] Step 2:
[0931] Transcription of audio data
[0932] Device: Built-in speech recognition software (e.g., Google Speech-to-Text API) in home phones converts collected voice data into text data.
[0933] Input: Digital audio data
[0934] Output: Text data
[0935] What it does: Speech recognition software analyzes the audio and generates text that says something like, "Hello, your bank account has been compromised. Immediate action is required."
[0936] Step 3:
[0937] Sending text data
[0938] Terminal: The home phone sends the transcribed conversation data to the server.
[0939] Input: Text data
[0940] Output: Send data to the server
[0941] How it works: A home phone sends data over the internet to a server, which then passes it on to a generative AI model.
[0942] Step 4:
[0943] Fraud detection
[0944] Server: A generative AI model (e.g., OpenAI GPT-4) analyzes the received conversation data and determines whether fraud is suspected.
[0945] Input: Text data
[0946] Output: Determination of likelihood of fraud
[0947] What it does: A generative AI model analyzes the text "Hello, your bank account has been compromised. Immediate action is required" and determines that it is likely fraudulent.
[0948] Step 5:
[0949] Audio alert occurs
[0950] Server: If fraud is suspected, the server sends an alert signal to the home phone.
[0951] Terminal: The speaker on the home phone receives the alert signal from the server and issues an audio alert.
[0952] Input: Alert signal
[0953] Output: Audio alert
[0954] What it does: The speaker on your home phone plays a voice alert saying, "Possible scam. Be careful."
[0955] (Application example 2)
[0956] Next, a description will be given of Application Example 2 of Form Example 2. In the following description, the data processing device 12 will be referred to as a "server" and the smart glasses 214 will be referred to as a "terminal."
[0957] In recent years, the number of victims of fraudulent phone calls has been increasing, and special frauds targeting the elderly in particular have become a social problem. Conventional fraud prevention systems have been inadequate in detecting fraudulent phone calls, making it difficult for users to recognize the possibility of fraud. Furthermore, there is a need for a system that can detect fraud in real time and issue a warning to users on home phones and smartphones.
[0958] The specific processing by the specific processing unit 290 of the data processing device 12 in the application example 2 is realized by the following means.
[0959] In this invention, the server includes means for detecting incoming fraudulent calls, means for listening to conversations using a voice recognition device, means for converting the conversation data into text and sending it to a generative AI, means for issuing a voice alert from a voice output device when the generative AI determines that fraud is suspected based on the conversation data, and means including an application to be installed on a smartphone. This enables real-time detection of fraudulent calls and prompt warning to users.
[0960] A "scam call" is a call made with the intent of defrauding someone of money or personal information through fraudulent means.
[0961] An "incoming call" refers to an incoming phone call.
[0962] A "voice recognition device" is a device that converts voice into text data.
[0963] "Conversation data" refers to the content of a conversation that has been converted into text by a voice recognition device.
[0964] "Generative AI" is artificial intelligence that generates new information based on input data.
[0965] An "audio output device" is a device for outputting audio.
[0966] "Voice alert" is a function that issues a warning by voice.
[0967] An "application installed on a smartphone" is a software program that runs on a smartphone.
[0968] To implement this invention, the following hardware and software are required. The hardware requires a smartphone, a voice recognition device, and a voice output device. The software requires a voice recognition library (e.g., speech_recognition), a generative AI library (e.g., openai), and a text-to-speech library (e.g., pyttsx3).
[0969] First, the phone call is recorded in real time using the smartphone's microphone. The recorded voice data is converted into text data using a voice recognition device. The speech_recognition library is used for this voice recognition.
[0970] The converted text data is then sent to the generative AI, which analyzes the input text data and determines whether it is fraudulent. This analysis is performed using the OpenAI library. Specifically, the following prompt is sent to the generative AI:
[0971] Please determine if the following conversation is a scam:
[0972] "Hello, this is bank security. Your account has been compromised. Please provide your account number and PIN."
[0973] If you suspect it is a scam, reply 'scam'.
[0974] If the generative AI determines that there is a possibility of fraud, it will issue a voice alert using the smartphone's audio output device. This voice alert uses the pyttsx3 library. The content of the voice alert is intended to alert the user to the possibility of fraud.
[0975] For example, if a user is on a call on their smartphone and a conversation that could be fraudulent takes place, the application automatically converts the conversation into text and sends it to generative AI to determine whether it is fraudulent. If it is determined that there is a possibility of fraud, an audio alert will sound from the smartphone speaker saying, "Possible fraud!"
[0976] In this way, the present invention provides real-time detection of fraudulent calls and prompt warning to users.
[0977] The flow of the specific processing in the application example 2 will be described with reference to FIG.
[0978] Step 1:
[0979] A user initiates a call on their smartphone, and the smartphone's microphone records the call in real time.
[0980] Input: Call audio
[0981] Output: Recorded audio data
[0982] How it works: The smartphone's microphone captures the call audio and saves it as audio data.
[0983] Step 2:
[0984] The terminal sends the recorded voice data to a voice recognition device, which converts the voice data into text data.
[0985] Input: Recorded audio data
[0986] Output: Text data
[0987] Specific operation: A speech recognition library (e.g., speech_recognition) analyzes the audio data and generates corresponding text data.
[0988] Step 3:
[0989] The device sends the generated text data to the generative AI, which analyzes the text data and determines whether it is fraudulent.
[0990] Input: Text data
[0991] Output: Judgment result on likelihood of fraud
[0992] How it works: A generative AI library (e.g., openai) analyzes the text data based on the prompt and determines whether it is likely to be fraudulent.
[0993] Step 4:
[0994] The server receives the judgment results from the generative AI and generates an audio alert if there is a possibility of fraud.
[0995] Input: Determination result regarding the possibility of fraud
[0996] Output: Audio alert content
[0997] Specific behavior: The generative AI generates a warning message saying, "Possible fraud!"
[0998] Step 5:
[0999] The terminal transmits the contents of the audio alert to the audio output device, and the audio alert is output.
[1000] Input: Audio alert content
[1001] Output: Audio alert
[1002] Specific operation: A text-to-speech library (e.g., pyttsx3) converts the contents of the voice alert into audio, and outputs the warning sound from the smartphone speaker.
[1003] Example 3
[1004] Next, a description will be given of Example 3 of Form Example 3. In the following description, the data processing device 12 will be referred to as a "server" and the smart glasses 214 will be referred to as a "terminal."
[1005] Conventional fraud detection systems are limited to detecting incoming fraudulent phone calls, and are inadequate at detecting fraud in text messages and emails received by users. In addition, there are limited means of notifying users of possible fraud, which can delay users' realization of the risk of fraud. This makes it difficult to prevent fraud damage before it happens.
[1006] The specific processing by the specific processing unit 290 of the data processing device 12 in the third embodiment is realized by the following means.
[1007] In this invention, the server includes means for transmitting text data entered by a user to the server, means for the server to analyze the text data using a generative artificial intelligence model and determine the possibility of fraud, and means for the server to transmit the analysis results to the communication terminal, thereby making it possible to quickly detect the possibility of fraud even in text messages or emails received by the user and notify the user.
[1008] "Means for detecting incoming fraudulent calls" means a device or software for identifying and detecting potentially fraudulent calls from among calls coming in through a communication line.
[1009] "A voice input device installed on a communication terminal" refers to a microphone or related hardware for collecting voice that is attached to a communication device such as a smartphone or home phone.
[1010] "Means for converting conversation data into text and transmitting it to the generative AI" refers to software or hardware for converting voice data into text data and transmitting that text data to the generative AI.
[1011] "Generative AI" is an AI model that uses natural language processing technology to analyze text data and detect specific patterns and phrases.
[1012] "Audio alerting means" means a device or software that issues an audio warning when it determines that fraud may be occurring.
[1013] The "means for transmitting text data input by the user to the server" refers to software or hardware for transmitting text data input by the user to the communication terminal to the server via a communication line such as the Internet.
[1014] "Means for the server to analyze text data using a generative artificial intelligence model and determine the possibility of fraud" refers to software or hardware that inputs the text data received by the server into a generative artificial intelligence model and evaluates the possibility of fraud based on the analysis results.
[1015] The "means by which the server transmits the analysis results to the communication terminal" refers to software or hardware for transmitting the analysis results of the generative artificial intelligence model to the communication terminal.
[1016] The "means by which the communication terminal displays the analysis results to the user" refers to software or hardware for visually or audibly notifying the user of the analysis results received from the server.
[1017] The present invention relates to a system for detecting the possibility of special fraud using a generative artificial intelligence model. Specific embodiments of this system will be described below.
[1018] System configuration
[1019] Subject: Server
[1020] The server is responsible for analyzing the text data using a generative artificial intelligence model to determine the likelihood of fraud. Specifically, the server uses the following hardware and software:
[1021] Hardware: High-performance processors, memory, and storage devices
[1022] Software: Generative artificial intelligence models (e.g., AI models using natural language processing technology)
[1023] The server receives the text data sent by the user and inputs it into a generative AI model. The model detects specific phrases and patterns that indicate potential fraud and returns the results to the server. The server then transmits the analysis results to the communication device.
[1024] Subject: Terminal
[1025] The terminal is responsible for sending the text data entered by the user to the server and displaying the analysis results from the server to the user. Specifically, the terminal uses the following hardware and software:
[1026] Hardware: Communication devices such as smartphones, tablets, and PCs
[1027] Software: Dedicated application
[1028] The terminal transmits the text data entered by the user to the server, and displays the analysis results received from the server to the user. If the analysis results indicate a possibility of fraud, the terminal displays a warning message.
[1029] Subject: User
[1030] The user is responsible for inputting text data through the terminal and checking the analysis results. Specifically, the user performs the following operations:
[1031] Enter the contents of a message or email into your device
[1032] Check the analysis results and take appropriate action if necessary
[1033] Specific examples
[1034] Example 1
[1035] If a user receives a message saying "Mom, please transfer the money now," the system will act as follows:
[1036] 1. The user enters a message into a dedicated application on the device.
[1037] 2. The device sends this message to the server.
[1038] 3. The server receives the message and inputs it as a prompt sentence into the generative artificial intelligence model.
[1039] 4. The generative artificial intelligence model analyzes the message and returns the result: "Possible fraud."
[1040] 5. The server sends the analysis results to the device.
[1041] 6. The device displays the analysis results to the user, and the user confirms the warning message.
[1042] Example 2
[1043] If a user receives an email saying "Your bank account has been compromised. Contact us immediately," the system works as follows:
[1044] 1. The user enters the contents of the email into a dedicated application on the device.
[1045] 2. The device sends the contents of this email to the server.
[1046] 3. The server receives the email content and inputs it as a prompt into the generative artificial intelligence model.
[1047] 4. The generative artificial intelligence model analyzes the content of the email and returns the result: "Possible fraud."
[1048] 5. The server sends the analysis results to the device.
[1049] 6. The device displays the analysis results to the user, and the user confirms the warning message.
[1050] Prompt Sentence Examples
[1051] Here are some example prompts to input to a generative AI model:
[1052] "Determine whether the following message is a potential scam: 'Mom, please transfer the money now.'"
[1053] "Determine if the following email is a potential scam: 'Your bank account has been compromised. Contact us now.'"
[1054] In this way, the specific fraud detection system using the generative artificial intelligence model operates specifically. The flow of the identification process in the third embodiment will be described with reference to FIG.
[1055] Step 1:
[1056] The user inputs text data.
[1057] The user enters the contents of the message or email they received into a dedicated application on their device. For example, they copy and paste the message "Mom, please transfer the money right now" into the application. The input data is saved in text format on the device.
[1058] Step 2:
[1059] The terminal transmits the text data to the server.
[1060] The terminal sends text data entered by the user to the server. Specifically, the terminal application generates an HTTP request and sends a payload containing the text data to the server. At this time, the data is encrypted before being sent. The input is the text data entered by the user, and the output is the text data sent to the server.
[1061] Step 3:
[1062] The server receives the text data and inputs it into a generative artificial intelligence model.
[1063] The server receives text data sent from the terminal. The received data is first stored in a database. The server then inputs the text data as a prompt sentence into the generative AI model. An example of a prompt sentence is "Please judge whether the following message is likely to be fraudulent: 'Mom, please transfer the money now.'" The input is the text data received from the terminal, and the output is the prompt sentence input into the generative AI model.
[1064] Step 4:
[1065] A generative artificial intelligence model analyzes the text data and returns the results.
[1066] The generative artificial intelligence model analyzes the input prompt sentence and determines whether it is likely to be fraud. Specifically, the model analyzes it by comparing it with typical phrases and patterns of "bank transfer fraud" and "it's me fraud." If the analysis result indicates a high possibility of fraud, it returns a warning message saying "Possible fraud." The input is the prompt sentence, and the output is the analysis result.
[1067] Step 5:
[1068] The server sends the analysis results to the device.
[1069] The server sends the analysis results obtained from the generative AI model to the terminal. Specifically, the server generates an HTTP response and sends a payload containing the analysis results to the terminal. At this time, the data is encrypted before being sent. The input is the analysis results from the generative AI model, and the output is the analysis results sent to the terminal.
[1070] Step 6:
[1071] The terminal displays the analysis results to the user.
[1072] The terminal displays the analysis results received from the server to the user. For example, if the analysis result is a warning message saying "Possible fraud," the terminal application displays this message to the user as a pop-up window or notification. The input is the analysis result received from the server, and the output is the warning message displayed to the user. The user can check this warning message and take appropriate action.
[1073] (Application example 3)
[1074] Next, a description will be given of Application Example 3 of Form Example 3. In the following description, the data processing device 12 will be referred to as a "server" and the smart glasses 214 will be referred to as a "terminal."
[1075] In recent years, special fraud methods have become more sophisticated, and many people have fallen victim to them. In particular, "bank transfer fraud" and "it's me" frauds targeting the elderly have become a social problem. In order to prevent these frauds, a system is needed that can detect possible fraud early and issue a warning to users. However, current systems have low fraud detection accuracy, making it difficult for users to recognize the possibility of fraud.
[1076] The specific processing by the specific processing unit 290 of the data processing device 12 in the application example 3 is realized by the following means.
[1077] In this invention, the server includes means for detecting incoming fraudulent calls, means for listening to conversations from a voice input device installed in the communication terminal, means for converting the conversation data into text and sending it to the generative AI, means for issuing a voice alert from the voice input device when the generative AI determines that fraud is suspected from the conversation data, means for the generative AI to detect specific phrases or patterns that indicate the possibility of fraud, and means for issuing a warning to the user when the generative AI detects specific phrases or patterns that indicate the possibility of fraud. This makes it possible to detect the possibility of fraud with high accuracy and issue a warning to the user quickly.
[1078] "Means for detecting incoming fraudulent calls" is a function for identifying and notifying a communication terminal of incoming calls that may be fraudulent.
[1079] A "voice input device installed on a communication terminal" is a device for collecting voice that is attached to communication equipment such as a smartphone or home telephone.
[1080] "Means for converting conversation data into text and sending it to generative AI" refers to a function for converting voice data collected by a voice input device into text data and sending that text data to generative AI.
[1081] "Means for issuing a voice alert from the voice input device when the generative AI determines that fraud is suspected based on the conversation data" is a function for issuing a warning sound through the voice input device when the generative AI determines that there is a possibility of fraud based on the conversation data analyzed.
[1082] "Means for generative AI to detect specific phrases or patterns that indicate the possibility of fraud" refers to a function that allows generative AI to detect typical fraudulent phrases and patterns that it has previously learned from conversation data.
[1083] "Means for issuing a warning to users when generative AI detects specific phrases or patterns that indicate the possibility of fraud" refers to a function that issues a warning to users when generative AI detects phrases or patterns that indicate the possibility of fraud.
[1084] As an embodiment of the present invention, a method for installing a fraud detection assistant system on a smartphone will be described.
[1085] First, the system is equipped with a means for detecting incoming fraudulent calls. This means has the function of identifying and notifying the communication terminal of calls that may be fraudulent. Specifically, the system collects the contents of the call in real time using a voice input device installed on the communication terminal.
[1086] The collected voice data is then converted into text data through a means of transcribing the conversation data and sending it to the generative AI. This conversion is done using speech recognition software (e.g., Google's speech recognition API). The converted text data is then sent to the generative AI model.
[1087] A generative AI model is pre-trained to detect specific phrases and patterns that indicate potential fraud. This model is built using, for example, the Hugging Face transformers library. If the generative AI determines from the conversation data that fraud is suspected, it will warn the user by issuing an audio alert via a voice input device.
[1088] Additionally, if the generative AI detects certain phrases or patterns that indicate potential fraud, it will activate a user alert mechanism that will alert the user to the potential fraud by providing a visual or audio warning.
[1089] As a concrete example, consider the following audio input file:
[1090] Voiceover: "Mom, please transfer the money now. It's urgent."
[1091] An example prompt for parsing this audio file:
[1092] Determine whether the phrase "Mom, please transfer the money now. It's urgent." is a potential scam.
[1093] By feeding this prompt into a generative AI model, it can analyze the possibility of fraud and issue a warning to the user.
[1094] The flow of the specific processing in Application Example 3 will be described with reference to FIG.
[1095] Step 1:
[1096] Detecting incoming fraudulent calls
[1097] The server identifies and notifies the communication terminal of incoming calls that may be fraudulent. Specifically, it monitors the communication terminal's call log and detects possible fraud based on specific numbers or patterns. The input is the call log data, and the output is a notification of a potentially fraudulent call.
[1098] Step 2:
[1099] Listen to conversations from a voice input device
[1100] The terminal collects the contents of the call in real time using an audio input device installed in the communication terminal. Specifically, it acquires audio data through a microphone and saves it as an audio file. The input is the audio during the call, and the output is an audio file.
[1101] Step 3:
[1102] Convert conversation data into text and send it to generative AI
[1103] The device converts the collected voice data into text data and sends it to the generative AI. Specifically, it uses voice recognition software (e.g., Google's voice recognition API) to transcribe the voice data. The input is an audio file, and the output is text data.
[1104] Step 4:
[1105] Generative AI analyzes conversation data
[1106] The server uses a generative AI model to detect specific phrases or patterns in text data that indicate potential fraud. Specifically, it analyzes the text data using a generative AI model (e.g., Hugging Face's transformers library). The input is the text data, and the output is a determination of the likelihood of fraud.
[1107] Step 5:
[1108] Provides audio alerts in case of suspected fraud
[1109] If the generative AI determines that there is a possibility of fraud, the device will issue a voice alert through the voice input device. Specifically, it will play a warning sound or message. The input is the result of the judgment on the possibility of fraud, and the output is a voice alert.
[1110] Step 6:
[1111] Warn the user
[1112] The device will alert the user if the generative AI detects certain phrases or patterns that indicate potential fraud, either by displaying a warning message on the screen or by sending a notification. The input is the determination of potential fraud, and the output is the warning to the user.
[1113] Furthermore, an emotion engine that estimates the user's emotion may be further combined. That is, the identification processing unit 290 may estimate the user's emotion using the emotion identification model 59, and perform identification processing using the user's emotion.
[1114] "Example 1"
[1115] One embodiment of the present invention is a system that combines a generative AI with an emotion engine. This system analyzes emotions based on the user's tone of voice, volume, and speaking rate. Specifically, if the user's tone or volume suddenly increases while receiving a scam call, the system senses that the user may be panicking. Based on this information, the generative AI determines that the possibility of a scam has increased and issues an audio alert.
[1116] "Example 2"
[1117] In another embodiment, the emotion engine may generate an audio alert when a certain threshold is exceeded. For example, if the user's tone of voice exceeds a certain threshold or if the user's speaking rate exceeds a certain threshold, the emotion engine may determine that the user is experiencing strong emotion. Based on this determination, the generative AI may generate an audio alert to warn the user of a potential scam.
[1118] "Example 3"
[1119] Another possible embodiment is a system in which the emotion engine and generative AI work together. In this system, the emotion engine analyzes the user's emotions and sends the results to the generative AI. The generative AI then combines the information from the emotion engine with the conversation data it has analyzed to determine the possibility of fraud. This allows for more accurate fraud detection.
[1120] The processing flow of each embodiment will be described below.
[1121] "Example 1"
[1122] Step 1: The user receives a scam call.
[1123] Step 2: A smartphone app or a microphone installed on a home phone listens to the conversation.
[1124] Step 3: Convert the conversation data into text and send it to the generative AI.
[1125] Step 4: Generative AI analyzes the text data to determine the likelihood of fraud.
[1126] Step 5: The emotion engine analyzes the user's emotions based on their tone of voice, volume, speaking rate, etc.
[1127] Step 6: The generative AI uses information from the emotion engine to further increase the likelihood of fraud.
[1128] It will then determine this and issue an audio alert.
[1129] "Example 2"
[1130] Step 1: The user receives a scam call.
[1131] Step 2: A smartphone app or a microphone installed on a home phone listens to the conversation.
[1132] Step 3: Convert the conversation data into text and send it to the generative AI.
[1133] Step 4: Generative AI analyzes the text data to determine the likelihood of fraud.
[1134] Step 5: The emotion engine analyzes the user's emotions based on their tone of voice, volume, speaking rate, etc.
[1135] Step 6: If the emotion engine exceeds a certain threshold, notify the generative AI.
[1136] Step 7: Based on the notification from the emotion engine, the generative AI determines that the likelihood of fraud has increased and issues an audio alert.
[1137] "Example 3"
[1138] Step 1: The user receives a scam call.
[1139] Step 2: A smartphone app or a microphone installed on a home phone listens to the conversation.
[1140] Step 3: Convert the conversation data into text and send it to the generative AI.
[1141] Step 4: Generative AI analyzes the text data to determine the likelihood of fraud.
[1142] Step 5: The emotion engine analyzes the user's emotions based on their tone of voice, volume, speaking speed, etc., and sends the results to the generative AI.
[1143] Step 6: The generative AI combines information from the emotion engine with the conversation data it has analyzed to determine the likelihood of fraud.
[1144] Step 7: If fraud is deemed likely, the generative AI will issue an audio alert.
[1145] Example 1
[1146] Next, a description will be given of Example 1 of Form Example 1. In the following description, the data processing device 12 will be referred to as a "server" and the smart glasses 214 will be referred to as a "terminal."
[1147] Conventional fraud call detection systems analyze only the content of conversations to determine the possibility of fraud, which means they are unable to take into account the emotional state of the user, resulting in low accuracy in detecting fraud. Furthermore, they are unable to respond adequately when the user panics, making it difficult to prevent fraud damage before it happens.
[1148] The specific processing by the specific processing unit 290 of the data processing device 12 in the first embodiment is realized by the following means.
[1149] In this invention, the server includes means for detecting incoming fraudulent calls, means for listening to conversations from a voice input device installed in the communication terminal, means for converting the conversation data into text and sending it to the generative artificial intelligence, means for issuing a voice alert from the voice input device when the generative artificial intelligence determines that fraud is suspected from the conversation data, means for collecting parameters such as the user's voice tone, volume, and speaking rate and analyzing their emotions, and means for reevaluating the possibility of fraud based on the results of the emotion analysis. This makes it possible to detect fraud taking into account the user's emotional state, thereby preventing fraud damage before it occurs.
[1150] "Means for detecting incoming fraudulent calls" refers to a device or software for detecting whether an incoming call to a communication terminal is likely to be fraudulent.
[1151] The "voice input device installed in a communication terminal" refers to a device for inputting voice, such as a microphone attached to a communication device such as a smartphone or home telephone.
[1152] "Means for converting conversation data into text and transmitting it to the generative AI" refers to a device or software for converting voice data into text data and transmitting the text data to the generative AI.
[1153] "Generative AI" is an AI system that analyzes the data it receives and determines the likelihood of fraud by detecting specific patterns and phrases.
[1154] A "means for issuing an audio alert" is a device or software that plays an audio message to warn the user when fraud is deemed likely.
[1155] "Means for collecting parameters such as the tone, volume, and speaking rate of a user's voice and analyzing emotions" refers to a device or software that collects information such as tone, volume, and speaking rate from the user's voice data and analyzes the user's emotional state based on that information.
[1156] "Means for reassessing the possibility of fraud based on the results of sentiment analysis" refers to a device or software that inputs the results of sentiment analysis into generative artificial intelligence and reassess the possibility of fraud.
[1157] MODE FOR CARRYING OUT THE INVENTION
[1158] This invention is a system that detects fraudulent calls and warns users. It is implemented using an application and server installed on a communication terminal. This system detects incoming fraudulent calls, listens to the conversation, transcribes the conversation data, and sends it to a generative artificial intelligence. Furthermore, the system aims to prevent fraud by analyzing the user's emotional state and reassessing the possibility of fraud.
[1159] Hardware and software used
[1160] 1. Communication terminal
[1161] Communication devices such as smartphones and home phones.
[1162] A microphone as an audio input device.
[1163] 2. Server
[1164] Voice recognition software (e.g., Google Cloud Speech-to-Text API) is used to convert voice data into text.
[1165] For sentiment analysis, we use a sentiment engine (e.g., IBM Watson Tone Analyzer).
[1166] For generative artificial intelligence, we use AI models with natural language processing technology (e.g., OpenAI's GPT-4).
[1167] Specific operation of the system
[1168] 1. Incoming phone call detection
[1169] The server detects incoming calls through an application installed on the communication terminal, which uses the communication terminal's native API to catch incoming call events.
[1170] 2. Listening to conversations
[1171] The device listens to the conversation during the call using a microphone, and the application captures the microphone's audio input in real time and saves it as audio data.
[1172] 3. Transcription of audio data
[1173] The server converts the audio data it hears into text data by calling the Google Cloud Speech-to-Text API.
[1174] 4. Sending to generative AI
[1175] The server sends the transcribed conversation data to the generative AI. The server sends the text data to the generative AI using an HTTP request.
[1176] 5. Fraud Detection
[1177] The generative AI analyzes the received conversation data and determines whether fraud is suspected. The generative AI uses natural language processing technology to analyze the content of the conversation and evaluate whether it matches any fraudulent patterns.
[1178] 6. Sentiment Analysis
[1179] The server collects parameters such as the user's voice tone, volume, and speaking rate and analyzes them using an emotion engine. The server uses voice analysis software to extract emotion parameters from the voice data.
[1180] 7. Panic Detection
[1181] The server determines that a user may be panicking if there is a sudden increase in the tone or volume of their voice. The server analyzes data from the emotion engine and evaluates whether there is a sudden change.
[1182] 8. Reassessing the possibility of fraud
[1183] The generative AI determines that the likelihood of fraud has increased based on information from the emotion engine. The generative AI receives the emotion data as additional input and reassess the likelihood of fraud.
[1184] 9. Sending audio alerts
[1185] If the server determines that there is a high possibility of fraud, it instructs the communication terminal to issue an audio alert from its speaker, and sends a command to the application of the communication terminal to play the audio alert.
[1186] Examples and prompts
[1187] As a concrete example, imagine a user is receiving a phone call on their smartphone. If the caller says, "Your bank account has been fraudulently used. Please tell us your account number and password immediately," the user's voice tone will suddenly rise and the volume will also increase. Based on this information, the generative AI will determine that there is a high possibility of fraud and will issue a voice alert from the communication device's speaker saying, "This may be a fraud. Please be careful."
[1188] Example prompt sentence:
[1189] "Analyze the conversation data below to determine if it is a potential scam. Also consider any changes in the user's tone and volume.
[1190] Conversation data: 'Your bank account has been compromised. Please provide your account number and password immediately.'
[1191] User's tone of voice: Sudden rise
[1192] User volume: louder"
[1193] The flow of the identification process in the first embodiment will be described with reference to FIG.
[1194] Step 1:
[1195] Incoming phone call detection
[1196] The server detects an incoming call through an application installed on the communication terminal.
[1197] Input: Incoming call event of communication terminal.
[1198] Data processing / calculation: Capture incoming call events using the communication device's native API.
[1199] Output: Event information that an incoming call was detected.
[1200] Specific behavior: The application uses native APIs for Android and iOS to catch incoming phone calls and notify the server.
[1201] Step 2:
[1202] Listening to conversations
[1203] The terminal uses a microphone to listen to the conversation during the call.
[1204] Input: Voice data during a call.
[1205] Data processing / calculation: Captures microphone audio input in real time and saves it as audio data.
[1206] Output: Audio data.
[1207] What it does: The application captures audio input from the microphone in real time and saves it as audio data.
[1208] Step 3:
[1209] Transcription of audio data
[1210] The server converts the audio data it hears into text data.
[1211] Input: Audio data.
[1212] Data processing / calculation: Call the Google Cloud Speech-to-Text API to convert the voice data into text.
[1213] Output: Character data.
[1214] What happens: The server converts the audio data into text using the Google Cloud Speech-to-Text API.
[1215] Step 4:
[1216] Sending to generative AI
[1217] The server transmits the transcribed conversation data to the generative artificial intelligence.
[1218] Input: Character data.
[1219] Data processing / calculation: Send text data to the generative AI using an HTTP request.
[1220] Output: The result of the request sent to the generative AI.
[1221] Specific operation: The server sends text data to the generative artificial intelligence using an HTTP request.
[1222] Step 5:
[1223] Fraud detection
[1224] The generative artificial intelligence analyzes the conversation data it receives and determines whether fraud is suspected.
[1225] Input: Character data.
[1226] Data processing / computation: Using natural language processing techniques, the content of conversations is analyzed to assess whether they match fraud patterns.
[1227] Output: Assessment results for likelihood of fraud.
[1228] How it works: Generative AI uses natural language processing technology to analyze the content of a conversation and evaluate whether it matches a fraud pattern.
[1229] Step 6:
[1230] Sentiment analysis
[1231] The server collects parameters such as the user's voice tone, volume, and speaking speed and analyzes them using an emotion engine.
[1232] Input: Audio data.
[1233] Data processing / calculation: Using voice analysis software, emotional parameters are extracted from the voice data.
[1234] Output: Emotion parameters.
[1235] Specific operation: The server uses voice analysis software (e.g., IBM Watson Tone Analyzer) to extract emotion parameters from the voice data.
[1236] Step 7:
[1237] Panic detection
[1238] If the tone or volume of a user's voice suddenly increases, the server determines that the user may be panicking.
[1239] Input: Emotion parameters.
[1240] Data processing / computation: Analyzing data from the emotion engine and assessing whether there are any sudden changes.
[1241] Output: Panic assessment result.
[1242] What it does: The server analyzes the data from the emotion engine and evaluates whether there are any sudden changes.
[1243] Step 8:
[1244] Reassessing the possibility of fraud
[1245] Based on information from the emotion engine, the generative AI determines that the likelihood of fraud has increased.
[1246] Input: Emotion parameters.
[1247] Data processing / computation: Emotional data is taken as additional input to reassess the likelihood of fraud.
[1248] Output: Reassessed fraud probability.
[1249] How it works: Generative AI takes emotional data as additional input and reassess the likelihood of fraud.
[1250] Step 9:
[1251] Sending a voice alert
[1252] If the server determines that there is a high possibility of fraud, it instructs the communication terminal to issue an audio alert through its speaker.
[1253] Enter: Reassessed fraud potential.
[1254] Data processing / calculation: Sends a command to the communication terminal application to play an audio alert.
[1255] Output: Plays an audio alert.
[1256] Specific operation: The server sends a command to the application on the communication terminal to play an audio alert, and the audio alert is played from the speaker of the communication terminal.
[1257] (Application example 1)
[1258] Next, a description will be given of Application Example 1 of Form Example 1. In the following description, the data processing device 12 will be referred to as a "server" and the smart glasses 214 will be referred to as a "terminal."
[1259] In recent years, fraudulent phone scams have become more sophisticated, and many people have fallen victim to them. Elderly people and those unfamiliar with technology are particularly vulnerable to these scams, as they are less vigilant against them. Furthermore, when users receive a fraudulent call, they often panic, making it difficult for them to make a calm judgment. Effective measures are needed to improve this situation and prevent fraudulent incidents.
[1260] The specific processing by the specific processing unit 290 of the data processing device 12 in the application example 1 is realized by the following means.
[1261] In this invention, the server includes means for detecting incoming fraudulent calls, means for listening to conversations from a voice input device installed in the communication terminal, means for converting the conversation data into text and sending it to the generative AI, means for issuing a voice alert from the voice input device when the generative AI determines that fraud is suspected from the conversation data, means including an emotion engine for analyzing emotions from the user's voice tone, volume, speaking speed, etc., and means for further increasing the possibility of fraud based on the analysis results of the emotion engine. This makes it possible to prevent fraud damage by combining fraud call detection with user emotion analysis.
[1262] A "scam call" is a call made with the intent of defrauding someone of money or personal information through fraudulent means.
[1263] An "incoming call" refers to an incoming phone call.
[1264] A "communication terminal" is an electronic device for transmitting and receiving voice and data.
[1265] An "audio input device" is a device for converting audio into digital data.
[1266] "Conversation data" is data obtained by converting speech acquired by a speech input device into text.
[1267] "Generative AI" is artificial intelligence that makes specific decisions and generates things based on input data.
[1268] A "voice alert" is a means of warning or notifying by voice.
[1269] An "emotion engine" is software or a system for analyzing emotions from voice data.
[1270] "Tone" refers to the pitch and texture of a voice.
[1271] "Volume" refers to the loudness of a sound.
[1272] "Speech rate" refers to the speed at which you speak.
[1273] The following system configuration is used as an embodiment of the present invention.
[1274] System Configuration
[1275] 1. Hardware
[1276] Communication terminal: A device capable of voice communication, such as a smartphone or home telephone.
[1277] Audio input device: A microphone built into the communication terminal.
[1278] Audio output device: A speaker built into a communication terminal.
[1279] 2. Software
[1280] Speech Recognition Library: Use the speech_recognition library to convert speech to text.
[1281] Generative AI model: An AI model that uses the transformers library to determine the likelihood of fraud.
[1282] Sentiment Engine: A model for analyzing user sentiment using the transformers library.
[1283] Audio Alert Library: Uses the pyttsx3 library to emit audio alerts.
[1284] Processing flow
[1285] 1. Incoming call detection
[1286] The server detects that a call has been received by the communication terminal.
[1287] The voice input device of the communication terminal listens to the conversation and acquires voice data.
[1288] 2. Voice Recognition
[1289] The server uses a voice recognition library to convert the acquired voice data into text data.
[1290] 3. Generative AI for Fraud Detection
[1291] The server inputs the text data into a generative AI model to determine the likelihood of fraud.
[1292] If certain phrases or patterns are detected, it is determined to be likely fraud.
[1293] 4. Sentiment analysis
[1294] The server uses an emotion engine to analyze the user's emotions based on the tone of their voice, volume, speaking speed, etc.
[1295] If the user is perceived as panicking, the likelihood of fraud increases even more.
[1296] 5. Audio alerts
[1297] If the server determines that there is a high possibility of fraud, it uses a voice alert library to issue a voice alert from the voice output device of the communication terminal.
[1298] The alerts are intended to raise awareness of potential fraud.
[1299] Specific examples
[1300] For example, if a user receives a phone call and it is determined that the call may be fraudulent, an audio alert will be emitted from the communication device's speaker saying, "Warning! Possible fraud." This alert will make the user aware of the possibility of fraud and allow them to respond calmly.
[1301] Prompt Sentence Examples
[1302] "Could this call be a scam?"
[1303] In this way, this invention can prevent fraud damage by combining fraud call detection with user emotion analysis.
[1304] The flow of the specific processing in the application example 1 will be described with reference to FIG.
[1305] Step 1:
[1306] The server detects that a call has been received by the communication terminal.
[1307] Input: Incoming signal from communication terminal.
[1308] Data processing: Receives an incoming call signal and starts processing when the call starts.
[1309] Output: Flag for starting the call.
[1310] Specific operation: When a communication terminal receives an incoming call, the server detects the signal and recognizes that a call has started.
[1311] Step 2:
[1312] The voice input device of the terminal listens to the conversation and acquires voice data.
[1313] Input: Audio during a call.
[1314] Data processing: The voice input device collects voice data in real time.
[1315] Output: Captured audio data.
[1316] Specific operation: Audio during a call is collected through the microphone and saved as audio data.
[1317] Step 3:
[1318] The server uses a voice recognition library to convert the acquired voice data into text data.
[1319] Input: Captured audio data.
[1320] Data processing: Convert the voice data into text data using the speech recognition library (speech_recognition).
[1321] Output: The converted text data.
[1322] Specific operation: Voice data is input into the voice recognition library and output as text data.
[1323] Step 4:
[1324] The server inputs the text data into a generative AI model to determine the likelihood of fraud.
[1325] Input: The converted text data.
[1326] Data processing: Text data is fed into generative AI models (transformers) to assess the likelihood of fraud.
[1327] Output: Assessment results for likelihood of fraud.
[1328] Specific operation: Text data is input into a generative AI model, which outputs an evaluation result on whether the data is likely to be fraudulent.
[1329] Step 5:
[1330] The server uses an emotion engine to analyze the user's emotions based on the tone of their voice, volume, speaking speed, etc.
[1331] Input: Converted text and audio data.
[1332] Data processing: Using emotion engines (transformers), we analyze user emotions.
[1333] Output: Analysis results on user sentiment.
[1334] Specific operation: Voice and text data are input into the emotion engine, which analyzes the user's emotional state.
[1335] Step 6:
[1336] If the server determines that there is a high possibility of fraud, it uses a voice alert library to issue a voice alert from the voice output device of the communication terminal.
[1337] Input: Fraud likelihood assessment results and user sentiment analysis results.
[1338] Data processing: If a high probability of fraud is determined, a voice alert is generated using the voice alert library (pyttsx3).
[1339] Output: Audio alert.
[1340] Specific operation: If it is determined that there is a high possibility of fraud, an audio alert will be emitted from the communication device's speaker saying, "Warning! Possible fraud."
[1341] Example 2
[1342] Next, a description will be given of Example 2 of Form Example 2. In the following description, the data processing device 12 will be referred to as a "server" and the smart glasses 214 will be referred to as a "terminal."
[1343] Conventional fraud prevention systems have low accuracy in detecting potential fraud, and are unable to completely eliminate the risk of users falling victim to fraud. In addition, they issue warnings without taking into account the user's emotional state, which can lead to users being unable to respond appropriately. This makes it difficult to prevent fraudulent incidents.
[1344] The specific processing by the specific processing unit 290 of the data processing device 12 in the second embodiment is realized by the following means.
[1345] In this invention, the server includes means for detecting incoming fraudulent calls, means for listening to conversations from a voice input device installed in the communication device, means for converting the conversation data into text and sending it to a generative AI model, means for issuing a voice alert from the voice input device if the generative AI model determines that fraud is suspected from the conversation data, means for an emotion analysis engine to analyze the user's tone of voice and speaking rate, and means for the generative AI model to issue a voice alert if the emotion analysis engine's reading exceeds a specific threshold. This makes it possible to detect the possibility of fraud with high accuracy and issue an appropriate warning according to the user's emotional state.
[1346] The "means for detecting incoming fraudulent calls" is a function for identifying and notifying a communication device of incoming calls that may be fraudulent.
[1347] A "voice input device" is a device that records conversations between a user and another person in real time and stores them in digital format.
[1348] "Means for converting conversation data into text and sending it to the generative AI model" is a function for converting voice data into text data and sending that text data to the generative AI model.
[1349] A "generative AI model" is an artificial intelligence system that analyzes input data, detects specific patterns and phrases, and determines the likelihood of fraud.
[1350] "Means for issuing audio alerts" refers to an audio notification function that warns users when a potential fraud is detected.
[1351] The "emotion analysis engine" is an engine that analyzes the user's tone of voice and speaking speed to determine the user's emotional state.
[1352] "When a specific threshold is exceeded" refers to when the emotion analysis engine analyzes the user's emotional state and the result exceeds a preset reference value.
[1353] A "communication device" is a device for making voice calls, such as a home telephone or smartphone.
[1354] MODE FOR CARRYING OUT THE INVENTION
[1355] The present invention is a system for detecting incoming fraudulent calls and issuing a warning to the user. A specific embodiment of this system will be described below.
[1356] Hardware and software used
[1357] Communication devices: devices used to make voice calls, such as home phones and smartphones
[1358] Audio input device: A microphone installed in the communication device
[1359] Server: A computer system for processing audio data and running generative AI models and sentiment analysis engines.
[1360] Generative AI model: An artificial intelligence system that analyzes conversation data and determines the likelihood of fraud
[1361] Emotion analysis engine: An engine that analyzes the user's tone of voice and speaking speed to determine their emotional state
[1362] Speech recognition software: Software for converting voice data into text data (e.g., Google Speech-to-Text API)
[1363] System Operation
[1364] 1. When a user starts a conversation on a communication device, the device's voice input device records the conversation in real time, and the recorded voice data is stored in digital format.
[1365] 2. The server uses speech recognition software to convert the collected voice data into text data, for example, "Hello, this is a bank representative" is converted into text "Hello, this is a bank representative."
[1366] 3. The server sends the converted text data to the generative AI model, which receives the data and prepares it for analysis.
[1367] 4. A generative AI model analyzes the conversation data to detect specific keywords and phrases (e.g., "bank account" or "password"), which can then be used to determine whether there is potential for fraud.
[1368] 5. If the generative AI model determines that there is a high possibility of fraud, the server will send instructions to the terminal and an audio alert will be issued from the communication device's speaker saying, "There is a possibility of fraud. Please be careful."
[1369] 6. The emotion analysis engine analyzes the user's tone of voice and speaking rate in real time. For example, if the user's voice suddenly gets higher in pitch or speaking rate, the emotion analysis engine will detect this.
[1370] 7. If the sentiment analysis engine determines that the user's emotions exceed a certain threshold, the generative AI model will issue an audio alert saying, "Possible scam. Please stay calm."
[1371] Specific examples
[1372] Example 1: When a user says "Please tell me your bank account information" over the phone, the generative AI model detects the keyword "bank account" and determines that it may be a scam. The speaker on the communication device issues a voice alert saying, "This may be a scam. Please be careful."
[1373] Example 2: If a user says "Hurry, tell me your password" over the phone, and the user's voice tone gets higher and the rate of speech increases, the emotion analysis engine determines that the user is experiencing strong emotions. The generative AI model issues a voice alert, warning, "This may be a scam. Please stay calm."
[1374] Prompt Sentence Examples
[1375] Prompt 1: "If someone calls and asks you to provide your bank account information, determine if this is a potential scam."
[1376] Prompt 2: "Please explain how the emotion engine determines if the user's voice tone gets higher and their speaking rate gets faster."
[1377] The system is designed to reduce the risk of users falling victim to phone scams by combining a generative AI model with a sentiment analysis engine to detect potential scams with greater accuracy and alert users.
[1378] The flow of the identification process in the second embodiment will be described with reference to FIG.
[1379] Step 1:
[1380] A user initiates a conversation on a communication device.
[1381] Specifically, a user initiates a call using a home phone or a smartphone. The input is the user's voice, and the output is voice data collected by a voice input device.
[1382] Step 2:
[1383] The terminal's voice input device listens to the conversation and collects voice data.
[1384] Specifically, the device's microphone records the conversation between the user and the other party in real time. The input is the user's voice and the output is audio data stored in digital format.
[1385] Step 3:
[1386] The server converts the audio data into text data.
[1387] Specifically, the server uses speech recognition software (e.g., Google Speech-to-Text API) to convert the collected voice data into text data. The input is digital voice data, and the output is text conversation data.
[1388] Step 4:
[1389] The server sends the transcribed conversation data to the generative AI model.
[1390] Specifically, the server sends the converted text data to the generative AI model. The input is text-format conversation data, and the output is the data sent to the generative AI model.
[1391] Step 5:
[1392] A generative AI model analyzes conversation data to determine the likelihood of fraud.
[1393] Specifically, the generative AI model analyzes conversation data to detect specific keywords or phrases (e.g., "bank account" or "password"). The input is textual conversation data, and the output is a judgment about the likelihood of fraud.
[1394] Step 6:
[1395] If there is a possibility of fraud, an audio alert will be issued through the device's voice input device.
[1396] Specifically, if the generative AI model determines that there is a high possibility of fraud, the server sends instructions to the terminal, and an audio alert is issued from the communication device's speaker saying, "There is a possibility of fraud. Please be careful." The input is the judgment result of the generative AI model, and the output is an audio alert.
[1397] Step 7:
[1398] The emotion analysis engine analyzes the user's tone of voice and speaking speed.
[1399] Specifically, the emotion analysis engine analyzes the user's tone of voice and speaking rate in real time. The input is the user's voice data, and the output is the analysis result regarding the user's emotional state.
[1400] Step 8:
[1401] If the sentiment analysis engine exceeds a certain threshold, the generative AI model will issue an audio alert.
[1402] Specifically, if the emotion analysis engine determines that the user's emotion exceeds a certain threshold, the generative AI model issues a voice alert saying, "This may be a scam. Please remain calm." The input is the analysis result of the emotion analysis engine, and the output is a voice alert.
[1403] (Application example 2)
[1404] Next, a description will be given of Application Example 2 of Form Example 2. In the following description, the data processing device 12 will be referred to as a "server" and the smart glasses 214 will be referred to as a "terminal."
[1405] In recent years, fraudulent phone scams have become more sophisticated, and many people have fallen victim to them. In particular, elderly people and people unfamiliar with technology are less vigilant against fraudulent phone calls and are more likely to fall victim to them. Furthermore, although there are systems to detect fraudulent calls, there are few warning systems that take into account the user's emotional state. This increases the risk that users may not be aware of the possibility of fraud and may become victims. Therefore, there is a need to provide a warning system that detects fraudulent calls and takes into account the user's emotional state.
[1406] The specific processing by the specific processing unit 290 of the data processing device 12 in the application example 2 is realized by the following means.
[1407] In this invention, the server includes means for detecting incoming fraudulent calls, means for listening to conversations from a voice input device installed in the smart device, means for converting the conversation data into text and sending it to a generative AI, means for issuing a voice alert from the voice input device when the generative AI determines that fraud is suspected from the conversation data, and means for analyzing the user's emotions using an emotion engine and issuing a warning when a specific threshold is exceeded. This makes it possible to detect fraudulent calls and issue a warning that takes into account the user's emotional state.
[1408] A "scam call" is a call made with the intent of defrauding someone of money or personal information through fraudulent means.
[1409] An "incoming call" refers to an incoming phone call.
[1410] A "smart device" is an electronic device with internet connectivity, such as a smartphone or tablet.
[1411] An "audio input device" is a device such as a microphone for converting audio into digital data.
[1412] "Conversation data" is data obtained by converting voice acquired through a voice input device into text.
[1413] "Generative AI" is a system that uses artificial intelligence to analyze data and make specific decisions or generate results.
[1414] A "voice alert" is a method of warning or notifying a user using voice.
[1415] An "emotion engine" is an algorithm or system for analyzing a user's emotional state from voice and text data.
[1416] A "specific threshold" is a numerical value or condition that serves as a standard when the emotion engine evaluates the user's emotional state.
[1417] A "warning" is a notification or alert that alerts the user.
[1418] The system for carrying out the present invention includes a series of means for detecting incoming fraudulent calls and issuing a warning to the user. A specific embodiment of the system will be described below.
[1419] First, a voice input device (microphone) is installed on a smart device (e.g., a smartphone or tablet). This voice input device captures the user's call content in real time. The captured voice data is converted into text data using voice recognition software (e.g., the speech_recognition library).
[1420] The transcribed conversation data is then sent to a generative AI model (e.g., a model using the Transformers library). This generative AI model analyzes the conversation data and determines whether it is likely to be fraudulent. If it is, audio alert software (e.g., the pyttsx3 library) is triggered to warn the user.
[1421] Furthermore, an emotion engine (e.g., an emotion analysis model using the transformers library) is used to analyze the user's emotional state. The emotion engine evaluates the user's tone of voice, speech rate, etc., and issues a warning if certain thresholds are exceeded. This warning is also notified to the user as an audio alert.
[1422] For example, if a user calls and asks, "Please tell me your bank account information," the generative AI model will detect the possibility of fraud and issue a voice alert. If the user is feeling angry or anxious, the emotion engine will detect this and issue a voice alert.
[1423] An example of a prompt is as follows:
[1424] User says: "What bank account information do I need?"
[1425] Prompt the generative AI model: "Is this call potentially fraudulent?"
[1426] In this way, a system can be realized that can detect fraudulent calls and warn users based on their emotional state.
[1427] The flow of the specific processing in the application example 2 will be described with reference to FIG.
[1428] Step 1:
[1429] A user starts a call on a smart device. The voice input device (microphone) of the smart device captures the call content in real time. The input is the user's voice data, and the output is a stream of voice data.
[1430] Step 2:
[1431] The device uses speech recognition software (speech_recognition library) to convert the captured voice data into text data. The input is a stream of voice data, and the output is transcribed conversation data.
[1432] Step 3:
[1433] The device sends the transcribed conversation data to a generative AI model (a model using the Transformers library). The generative AI model analyzes the conversation data and determines whether there is a possibility of fraud. The input is the transcribed conversation data, and the output is a judgment about the possibility of fraud.
[1434] Step 4:
[1435] The server receives the judgment result of the generative AI model, and if it determines that there is a possibility of fraud, it launches software (pyttsx3 library) to issue an audio alert. The input is the judgment result regarding the possibility of fraud, and the output is an audio alert.
[1436] Step 5:
[1437] The device analyzes the user's emotional state using an emotion engine (an emotion analysis model using the Transformers library). The emotion engine evaluates the user's tone of voice, speech rate, etc., and issues a warning if certain thresholds are exceeded. The input is transcribed conversation data, and the output is an evaluation result of the user's emotional state.
[1438] Step 6:
[1439] The server receives the evaluation results of the emotion engine, and if a certain threshold is exceeded, it launches software (pyttsx3 library) to issue an audio alert. The input is the evaluation result on the emotional state, and the output is an audio alert.
[1440] Step 7:
[1441] The user receives an audio alert and perceives a warning about potential fraud or emotional state. The input is the audio alert and the output is the user's perception.
[1442] Example 3
[1443] Next, a description will be given of Example 3 of Form Example 3. In the following description, the data processing device 12 will be referred to as a "server" and the smart glasses 214 will be referred to as a "terminal."
[1444] Conventional fraud detection systems have the problem of low accuracy because they only analyze conversation data when determining the possibility of fraud. Also, because they do not take into account the user's emotional state, there is a high risk of overlooking possible fraud. Furthermore, there are limited means of notifying users of possible fraud, so users may not be aware of the risk of fraud.
[1445] The specific processing by the specific processing unit 290 of the data processing device 12 in the third embodiment is realized by the following means.
[1446] In this invention, the server includes means for detecting incoming fraudulent calls, means for listening to conversations from a voice input device installed in the communication terminal, means for converting the conversation data into text and transmitting it to the generative artificial intelligence, means for issuing a voice alert from the voice input device when the generative artificial intelligence determines that fraud is suspected from the conversation data, means for an emotion analysis engine to analyze the user's emotions, and means for the generative artificial intelligence to integrate information from the emotion analysis engine and the conversation data to determine the possibility of fraud. This makes it possible to detect the possibility of fraud with high accuracy and notify the user promptly.
[1447] A "means for detecting incoming fraudulent calls" is a device or software that monitors a communication line and identifies incoming calls that may be fraudulent.
[1448] The "voice input device installed on a communication terminal" is a microphone for collecting voice that is attached to a communication device such as a smartphone or home telephone.
[1449] "Means for converting conversation data into text and sending it to the generative AI" refers to software or a device for converting voice data into text data and sending that text data to the generative AI.
[1450] "Generative AI" is an AI model that uses natural language processing technology to analyze text data and detect specific patterns and phrases.
[1451] The "means for issuing an audio alert" is an audio output device that issues a warning to the user when the generative artificial intelligence detects the possibility of fraud.
[1452] An "emotion analysis engine" is software or a device that analyzes a user's voice or text data to identify their emotional state.
[1453] The "means for determining the likelihood of fraud" is an algorithm or software that uses generative artificial intelligence to integrate information from the conversational data and sentiment analysis engine to assess the likelihood of fraud.
[1454] The present invention is a system for detecting incoming fraudulent calls and notifying users of the possibility of fraud. A specific embodiment of this system will be described below.
[1455] First, when a user is using a communication terminal (such as a smartphone or home phone), a means for detecting incoming fraudulent calls is activated. This means is a device or software that monitors communication lines and identifies incoming calls that may be fraudulent.
[1456] Next, a voice input device (microphone) installed on the communication terminal listens to the conversation between the user and the fraudulent caller. This voice data is converted into text data by transcribing the conversation data and sending it to generative artificial intelligence. This conversion uses voice recognition technology. Specifically, voice recognition services such as Google Speech-to-Text API and IBM Watson Speech to Text can be used.
[1457] The server analyzes the text data using generative artificial intelligence (e.g., natural language processing models such as BERT and GPT-3). The generative artificial intelligence detects specific phrases and patterns that may be fraudulent and evaluates the results. For example, if the phrase "Please transfer money" is included, it will determine that there is a high possibility of fraud.
[1458] Furthermore, the server uses an emotion analysis engine (for example, IBM Watson or Microsoft Azure's emotion analysis API) to analyze the user's emotions. It analyzes the user's voice and text data to identify emotional states such as anxiety or tension. The emotion data obtained by the emotion analysis engine is sent to the generative artificial intelligence.
[1459] Generative AI combines information from the sentiment analysis engine with conversational data to more accurately assess the likelihood of fraud. For example, if a message says "Please transfer money" and the user's emotions are anxious, it will conclude that there is a very high likelihood of fraud.
[1460] Finally, if the server determines that the message is likely to be fraudulent, it notifies the user of the result. The notification is made by issuing an audio alert. Specifically, an audio alert is issued from the communication terminal's speaker saying, "This message may be fraudulent. Please be careful."
[1461] As a concrete example, consider the case where a user receives a message saying, "Please transfer money." This message is listened to by the voice input device of the communication terminal and converted into text data. The server analyzes this text data using generative artificial intelligence and determines that there is a high possibility of fraud. At the same time, the emotion analysis engine detects the user's anxiety and sends this information to the generative artificial intelligence. The server combines this information and concludes that there is a very high possibility of fraud. Finally, the server sends a voice alert to the user's communication terminal saying, "This message may be fraudulent. Please be careful."
[1462] Example prompt sentence:
[1463] "Please determine if the following message is a potential scam. Message: 'Please transfer money.'"
[1464] In this way, the server, the terminal, and the user work together to detect the possibility of fraud with high accuracy and notify the user promptly. The flow of the identification process in the third embodiment will be described with reference to FIG.
[1465] Step 1:
[1466] The user enters a message.
[1467] The user inputs the received message using a communication terminal (such as a smartphone or home phone). For example, the user inputs a message such as "Please transfer money." The input message is saved as voice data on the terminal.
[1468] Step 2:
[1469] The device sends a message to the server.
[1470] The terminal transmits the voice data input by the user to the server. At this time, the terminal transmits the voice data to the server via the Internet. The input is the voice data, and the output is the transfer of the voice data to the server.
[1471] Step 3:
[1472] The server converts the voice data into text data.
[1473] The server converts the received voice data into text data using voice recognition technology. Specifically, it uses voice recognition services such as Google Speech-to-Text API and IBM Watson Speech to Text. The input is voice data, and the output is text data.
[1474] Step 4:
[1475] The server analyzes the text data using generative artificial intelligence.
[1476] The server analyzes the text data using generative artificial intelligence (e.g., natural language processing models such as BERT or GPT-3). The generative artificial intelligence detects specific phrases or patterns that may be fraudulent and evaluates the results. The input is the text data, and the output is an evaluation result regarding the likelihood of fraud.
[1477] Step 5:
[1478] The server analyzes the user's emotions using a sentiment analysis engine.
[1479] The server uses an emotion analysis engine (for example, IBM Watson or Microsoft Azure's emotion analysis API) to analyze the user's emotions. It analyzes the user's voice and text data to identify emotional states such as anxiety or tension. The input is voice data or text data, and the output is emotion data.
[1480] Step 6:
[1481] The server combines the results of the generative artificial intelligence and sentiment analysis engine to determine the likelihood of fraud.
[1482] The server combines the analysis results from the generative AI with the emotional data from the emotion analysis engine. This allows for a more accurate assessment of the likelihood of fraud. For example, if the message is "Please transfer money" and the user's emotion is anxiety, it will conclude that there is a very high likelihood of fraud. The inputs are the analysis results and emotional data, and the output is a final assessment of the likelihood of fraud.
[1483] Step 7:
[1484] The server notifies the user of possible fraud.
[1485] If the server determines that the message is likely to be fraudulent, it notifies the user of the result. The notification is performed by issuing an audio alert. Specifically, an audio alert saying "This message may be fraudulent. Please be careful" is issued from the speaker of the communication terminal. The input is the final assessment of the likelihood of fraud, and the output is the audio alert.
[1486] (Application example 3)
[1487] Next, a description will be given of Application Example 3 of Form Example 3. In the following description, the data processing device 12 will be referred to as a "server" and the smart glasses 214 will be referred to as a "terminal."
[1488] Conventional fraud detection systems only detect specific phrases and patterns and do not take into account the user's emotional state, making it difficult to accurately determine the likelihood of fraud. Furthermore, there is a lack of a way to provide users with appropriate warnings when fraud is likely.
[1489] The identification processing by the identification processing unit 290 of the data processing device 12 in Application Example 3 is realized by the following means. In this invention, the server includes means for detecting incoming fraudulent calls, means for listening to conversations from a smartphone application or a microphone installed on a home phone, means for converting the conversation data into text and sending it to the generative AI, means for issuing an audio alert from the microphone if the generative AI determines that fraud is suspected based on the conversation data, means for analyzing the user's emotions using an emotion engine, means for determining the possibility of fraud by combining information from the emotion engine with the conversation data analyzed by the generative AI, and means for displaying a warning to the user if there is a high possibility of fraud. This makes it possible to more accurately determine the possibility of fraud and provide the user with an appropriate warning.
[1490] "Means for detecting incoming fraudulent calls" refers to devices or software that monitor telephone lines or communications networks to detect incoming calls that may be fraudulent.
[1491] A "microphone installed in a smartphone application or home phone" is a voice input device attached to a smartphone or home phone to collect audio.
[1492] A "means for listening to a conversation" is a device or software that analyzes audio data collected through a microphone and recognizes the content of the conversation.
[1493] "Means for converting conversation data into text and sending it to generative AI" refers to devices or software that convert voice data into text data and send that text data to generative AI.
[1494] "Generative AI" is an artificial intelligence model trained to perform a specific task, in this case to determine the likelihood of fraud.
[1495] "Means for issuing an audio alert from the microphone when it is determined that fraud is suspected" refers to a device or software that issues an alarm or message through the microphone when the generative AI detects the possibility of fraud.
[1496] An "emotion engine" is software or a device for analyzing a user's emotional state from their voice or text data.
[1497] "Means for analyzing user emotions" refers to devices or software that use an emotion engine to analyze the user's emotional state.
[1498] "Means for determining the possibility of fraud by combining information from an emotion engine with conversational data analyzed by a generative AI" refers to devices or software that integrate emotional information obtained from an emotion engine with the results of analysis of conversational data by a generative AI, and make a comprehensive determination of the possibility of fraud.
[1499] "Means for displaying a warning to the user in the event of a high likelihood of fraud" refers to a device or software that provides a visual or audible warning to the user in the event that a high likelihood of fraud is determined.
[1500] The system for carrying out the present invention includes a set of means for detecting an incoming fraudulent call and providing a warning to the user. Specific embodiments will be described below.
[1501] First, the server is equipped with a means for detecting incoming fraudulent calls. This means is a device or software that monitors telephone lines and communication networks to detect incoming calls that may be fraudulent. For example, it is possible to register specific phone numbers or callers on a blacklist and detect fraudulent calls based on that.
[1502] Next, the device (smartphone application or home phone) is equipped with a microphone to listen to the conversation. This microphone collects the voice during the call and sends it to a server as conversation data. The conversation data is converted into text using voice recognition technology and sent to the generative AI.
[1503] Generative AI analyzes conversation data to determine the likelihood of fraud. It uses models trained to detect specific phrases and patterns. For example, it can detect typical phrases for "bank transfer fraud" and "it's me" fraud.
[1504] Furthermore, the emotion engine analyzes the user's emotions. The emotion engine is software that analyzes the user's emotional state from voice and text data. The emotional information obtained from the emotion engine is integrated with the results of the generative AI analysis of the conversation data to comprehensively determine the possibility of fraud.
[1505] If a fraudulent activity is deemed likely, the device will display a warning to the user. This warning can be visual or audible, for example, by displaying a warning message on the smartphone screen or by issuing an audio alert.
[1506] For example, consider the following user input:
[1507] "Your account has been frozen. Please transfer the funds immediately."
[1508] "Hey, I need money right now."
[1509] These prompts can be fed into a generative AI model that analyzes the likelihood of fraud and works in conjunction with an emotion engine to display a warning to the user.
[1510] The flow of the specific processing in Application Example 3 will be described with reference to FIG.
[1511] Step 1:
[1512] The server monitors telephone lines and communication networks to detect incoming calls that may be fraudulent. The input is a signal from the telephone line or communication network, and the output is a flag indicating an incoming fraudulent call. Specifically, it registers specific phone numbers and callers on a blacklist and detects fraudulent calls based on that.
[1513] Step 2:
[1514] The device (smartphone application or home phone) collects the voice during a call through a microphone. The input is the voice signal during the call, and the output is the voice data. Specifically, the microphone converts the voice into a digital signal and sends the data to a server.
[1515] Step 3:
[1516] The server uses voice recognition technology to convert the voice data into text. The input is voice data and the output is text data. Specifically, the voice recognition software analyzes the voice data and converts it into corresponding text.
[1517] Step 4:
[1518] The server sends text data to a generative AI to analyze the likelihood of fraud. The input is text data, and the output is a score or flag indicating the likelihood of fraud. Specifically, the generative AI model detects specific phrases and patterns to determine the likelihood of fraud.
[1519] Step 5:
[1520] The server analyzes the user's emotions using an emotion engine. The input is text data, and the output is data indicating the user's emotional state. Specifically, the emotion engine analyzes the text data and identifies the user's emotional state.
[1521] Step 6:
[1522] The server combines information from the emotion engine with conversation data analyzed by the generative AI to comprehensively determine the likelihood of fraud. The input is emotional state data and a score or flag indicating the likelihood of fraud, and the output is a final flag indicating the likelihood of fraud. Specifically, the server integrates emotional information and the analysis results of the generative AI to determine the likelihood of fraud with high accuracy.
[1523] Step 7:
[1524] If it determines that there is a high possibility of fraud, the device will display a warning to the user. The input is a flag indicating the final possibility of fraud, and the output is a warning message or audio alert. Specifically, a warning message is displayed on the smartphone screen or an audio alert is issued.
[1525] The specific processing unit 290 transmits the result of the specific processing to the smart glasses 214. In the smart glasses 214, the control unit 46A causes the speaker 240 to output the result of the specific processing. The microphone 238 acquires audio indicating a user input regarding the result of the specific processing. The control unit 46A transmits audio data indicating the user input acquired by the microphone 238 to the data processing device 12. In the data processing device 12, the specific processing unit 290 acquires the audio data.
[1526] The data generation model 58 is a so-called generative AI (Artificial Intelligence). An example of the data generation model 58 is ChatGPT (Internet Search<URL: https: / / openai.com / blog / chatgpt> ) and other generation AIs. The data generation model 58 is obtained by performing deep learning on a neural network. A prompt including an instruction is input to the data generation model 58, and inference data such as voice data indicating voice, text data indicating text, and image data indicating an image is also input. The data generation model 58 performs inference on the input inference data in accordance with the instruction indicated by the prompt, and outputs the inference result in a data format such as voice data and text data. Here, inference refers to, for example, analysis, classification, prediction, and / or summarization.
[1527] Another example of generative AI is Gemini (internet search engine). <url: https: gemini.google.com ?hl="ja">) are listed.
[1528] In the above embodiment, an example in which the specific processing is performed by the data processing device 12 has been given, but the technology of the present disclosure is not limited to this, and the specific processing may be performed by the smart glasses 214.
[1529] [Third embodiment]
[1530] FIG. 5 shows an example of the configuration of a data processing system 310 according to the third embodiment.
[1531] 5, the data processing system 310 includes the data processing device 12 and a headset type terminal 314. An example of the data processing device 12 is a server.
[1532] The data processing device 12 includes a computer 22, a database 24, and a communication I / F 26. The computer 22 is an example of a "computer" according to the technology of the present disclosure. The computer 22 includes a processor 28, a RAM 30, and a storage 32. The processor 28, the RAM 30, and the storage 32 are connected to a bus 34. The database 24 and the communication I / F 26 are also connected to the bus 34. The communication I / F 26 is connected to a network 54. Examples of the network 54 include a WAN (Wide Area Network) and / or a LAN (Local Area Network).
[1533] The headset type terminal 314 includes a computer 36, a microphone 238, a speaker 240, a camera 42, a communication I / F 44, and a display 343. The computer 36 includes a processor 46, a RAM 48, and a storage 50. The processor 46, the RAM 48, and the storage 50 are connected to a bus 52. The microphone 238, the speaker 240, the camera 42, and the display 343 are also connected to the bus 52.
[1534] The microphone 238 receives instructions and the like from the user 20 by receiving voice uttered by the user 20. The microphone 238 captures the voice uttered by the user 20, converts the captured voice into audio data, and outputs it to the processor 46. The speaker 240 outputs audio in accordance with instructions from the processor 46.
[1535] Camera 42 is a small digital camera equipped with an optical system including a lens, aperture, and shutter, and an imaging element such as a CMOS (Complementary Metal-Oxide-Semiconductor) image sensor or a CCD (Charge Coupled Device) image sensor, and captures images of the surroundings of user 20 (for example, an imaging range defined by an angle of view equivalent to the field of vision of a typical healthy person).
[1536] The communication I / F 44 is connected to a network 54. The communication I / Fs 44 and 26 are responsible for the exchange of various information between the processor 46 and the processor 28 via the network 54. The exchange of various information between the processor 46 and the processor 28 using the communication I / Fs 44 and 26 is carried out in a secure state.
[1537] Fig. 6 shows an example of the main functions of the data processing device 12 and the headset type terminal 314. As shown in Fig. 6, in the data processing device 12, a specific process is performed by the processor 28. A specific process program 56 is stored in the storage 32.
[1538] The specific processing program 56 is an example of a "program" according to the technology of the present disclosure. The processor 28 reads the specific processing program 56 from the storage 32 and executes the read specific processing program 56 on the RAM 30. The specific processing is realized by the processor 28 operating as a specific processing unit 290 in accordance with the specific processing program 56 executed on the RAM 30.
[1539] The storage 32 stores a data generation model 58 and an emotion identification model 59. The data generation model 58 and the emotion identification model 59 are used by the identification processing unit 290.
[1540] In the headset type terminal 314, a reception output process is performed by the processor 46. A reception output program 60 is stored in the storage 50. The processor 46 reads the reception output program 60 from the storage 50 and executes the read reception output program 60 on the RAM 48. The reception output process is realized by the processor 46 operating as the control unit 46A in accordance with the reception output program 60 executed on the RAM 48.
[1541] Next, the specific processing by the specific processing unit 290 of the data processing device 12 will be described.
[1542] "Example 1"
[1543] One embodiment of the present invention uses a smartphone application. In this case, the application installed on the smartphone detects an incoming call and listens to the conversation through the microphone. The listened conversation data is transcribed and sent to a generative AI. The generative AI determines from the conversation data that fraud is suspected, and if fraud is suspected, issues a voice alert from the smartphone speaker.
[1544] "Example 2"
[1545] Another embodiment of the present invention involves installing a microphone on a home phone. In this case, the microphone on the home phone listens to the phone conversation, transcribes the conversation data, and sends it to a generative AI. The generative AI determines from the conversation data that fraud is suspected, and if fraud is suspected, issues an audio alert from the home phone's speaker.
[1546] "Example 3"
[1547] In a further embodiment of the present invention, the generative AI detects specific phrases or patterns that indicate the possibility of special fraud. Specifically, the generative AI learns typical phrases for "bank transfer fraud" and typical patterns for "it's me" fraud, and by detecting these, it determines the possibility of fraud.
[1548] The processing flow of each embodiment will be described below.
[1549] "Example 1"
[1550] Step 1: An application installed on the smartphone detects an incoming call.
[1551] Step 2: The application listens to your conversation through your smartphone's microphone.
[1552] Step 3: The conversation data is transcribed and sent to the generative AI.
[1553] Step 4: The generative AI determines that fraud is suspected based on the conversation data.
[1554] Step 5: If fraud is suspected, an audio alert will be issued through the smartphone speaker.
[1555] "Example 2"
[1556] Step 1: A microphone installed in a home phone listens to the phone conversation.
[1557] Step 2: The conversation data is transcribed and sent to the generative AI.
[1558] Step 3: The generative AI determines that fraud is suspected based on the conversation data.
[1559] Step 4: If fraud is suspected, an audio alert will be issued through the home phone speaker.
[1560] "Example 3"
[1561] Step 1: The generative AI learns specific phrases or patterns that indicate potential fraud.
[1562] Step 2: The generative AI detects specific phrases or patterns it has learned to identify suspected fraud in the conversation data.
[1563] Step 3: If fraud is suspected, an audio alert will be issued.
[1564] Example 1
[1565] Next, a description will be given of Example 1 of Form Example 1. In the following description, the data processing device 12 will be referred to as a "server" and the headset type terminal 314 will be referred to as a "terminal."
[1566] In recent years, telephone fraud has been on the rise, and special frauds targeting the elderly in particular have become a social problem. Conventional countermeasures require users to identify fraud themselves when receiving a fraudulent call, but this is becoming more difficult as fraud methods become more sophisticated. Therefore, there is a need for a system that automatically detects the possibility of fraud and issues a warning when a user receives a fraudulent call.
[1567] The identification process by the identification processing unit 290 of the data processing device 12 in Example 1 is realized by the following means. In this invention, the server includes a means for detecting an incoming fraudulent call, a means for listening to the conversation from a microphone installed in the communication terminal, a means for converting the conversation data into text and sending it to the generative AI model, and a means for issuing an audio alert from the speaker of the communication terminal when the generative AI model determines that fraud is suspected based on the conversation data. This makes it possible to automatically detect the possibility of fraud when a user receives a fraudulent call and immediately issue a warning.
[1568] "Means for detecting incoming fraudulent calls" refers to a function or device for detecting incoming calls in a communication terminal.
[1569] "Means for listening to conversations from a microphone installed in a communication terminal" refers to a function or device that uses a microphone built into or connected to a communication terminal to collect voices during a call in real time.
[1570] "Means for converting conversation data into text and sending it to the generative AI model" refers to a function or device for converting collected voice data into text data and sending that text data to the generative AI model.
[1571] A "generative AI model" is an artificial intelligence model that analyzes input data and detects specific patterns or phrases.
[1572] "Means for issuing an audio alert from the speaker of the communication device when it is determined that fraud is suspected" refers to a function or device for playing an alarm sound or message through the speaker of the communication device when the generative AI model detects the possibility of fraud.
[1573] "Specific phrases or patterns that indicate the possibility of fraud" are specific words or sentence structures related to fraudulent acts, and the generative AI model's detection of these serves as the basis for determining the possibility of fraud.
[1574] An "audio alert" is a message or sound that provides an audio warning to the user.
[1575] The present invention is a system that automatically detects fraudulent calls and issues a warning to the user using an application installed on a communication terminal. A specific embodiment of this system will be described below.
[1576] First, the user installs a dedicated application on their communication device. This application has the function of detecting incoming calls. Specifically, it uses the communication device's telephone API to catch incoming call events.
[1577] When a user answers a call, the device listens to the conversation in real time through the microphone. The application runs in the background and captures the conversation. The captured conversation data is transcribed using the Google Speech-to-Text API. The voice data is sent to the API and received as text data.
[1578] The transcribed conversation data is sent to a server via the internet. The HTTPS protocol is used for transmission to ensure data security. The server inputs the received conversation data into a generative AI model (e.g., OpenAI's GPT-4). The generative AI model uses prompts to analyze the possibility of fraud. Specific examples of prompts are as follows:
[1579] Analyze the conversation data below to determine if it is a potential scam.
[1580] Conversation data: "Your bank account has been compromised. Please provide your account number to verify."
[1581] The generative AI model analyzes the conversation data and determines whether fraud is suspected. If so, the server sends that information back to the device, again using the HTTPS protocol.
[1582] The device analyzes the results received from the server, and if fraud is suspected, it issues an audio alert from the device's speaker. Specifically, it plays an audio message such as "There is a possibility of fraud. Please be careful." This alert makes the user aware of the possibility of fraud and allows them to take appropriate action.
[1583] As a concrete example, consider the case where a user receives a phone call saying, "Your bank account has been fraudulently used. Please provide your account number to verify." This conversation is picked up by a microphone and transcribed. The transcribed data becomes, "Your bank account has been fraudulently used. Please provide your account number to verify." This data is sent to a generative AI model, which determines that there is a high possibility of fraud. The server sends the result back to the device, which then issues an audio alert. The user hears this alert and becomes aware of the possibility of fraud.
[1584] In this way, the present invention can automatically detect possible fraud and immediately issue a warning when a user receives a fraudulent call.
[1585] The flow of the identification process in the first embodiment will be described with reference to FIG.
[1586] Step 1: Detecting an incoming call
[1587] The terminal detects incoming calls through an application installed on the communication terminal. Specifically, it uses the communication terminal's telephone API to catch incoming call events. The input is the incoming call signal, and the output is the incoming call detection event.
[1588] Step 2: Listen to the conversation
[1589] When a user answers a call, the device listens to the conversation in real time through the microphone. The application runs in the background and captures the conversation. The input is the voice data during the call, and the output is the collected voice data.
[1590] Step 3: Transcribe the conversation data
[1591] The device uses the Google Speech-to-Text API to transcribe the conversation data it hears. It sends the speech data to the API and receives it as text data. The input is the collected speech data, and the output is the transcribed text data.
[1592] Step 4: Send the transcribed data
[1593] The terminal sends the transcribed conversation data to the server. The HTTPS protocol is used for transmission to ensure data security. The input is the transcribed text data, and the output is the data transmission event to the server.
[1594] Step 5: Fraud detection
[1595] The server inputs the received conversation data into the generative AI model, which analyzes the possibility of fraud using prompt sentences. Specific examples of prompt sentences are as follows:
[1596] Analyze the conversation data below to determine if it is a potential scam.
[1597] Conversation data: "Your bank account has been compromised. Please provide your account number to verify."
[1598] The input is transcribed text data, and the output is a judgment result regarding the possibility of fraud.
[1599] Step 6: Receiving the results
[1600] The server receives the judgment result from the generative AI model and returns the result to the terminal. The return is again via HTTPS. The input is the judgment result from the generative AI model, and the output is a data transmission event to the terminal.
[1601] Step 7: Generate an audio alert
[1602] The terminal analyzes the judgment result received from the server, and if fraud is suspected, it issues an audio alert from the communication terminal's speaker. Specifically, it plays an audio message such as "There is a possibility of fraud. Please be careful." The input is the judgment result from the server, and the output is the generation of an audio alert.
[1603] (Application example 1)
[1604] Next, a description will be given of Application Example 1 of Form Example 1. In the following description, the data processing device 12 will be referred to as a "server" and the headset type terminal 314 will be referred to as a "terminal."
[1605] In recent years, telephone fraud has been on the rise, and special fraud targeting the elderly in particular has become a social problem. Conventional fraud prevention measures have made it difficult for users to recognize the possibility of fraud, making it difficult to prevent damage before it occurs. Therefore, there is a need for a system that analyzes telephone conversations in real time and issues an immediate warning if there is a possibility of fraud.
[1606] The specific processing by the specific processing unit 290 of the data processing device 12 in the application example 1 is realized by the following means.
[1607] In this invention, the server includes means for detecting incoming fraudulent calls, means for listening to conversations from a voice input device installed in the communication terminal, means for converting the conversation data into text and sending it to the generative AI, means for issuing an audio alert from the voice output device if the generative AI determines that fraud is suspected from the conversation data, means for the generative AI to detect specific phrases or patterns that indicate the possibility of fraud, and means for the audio alert to alert the user to the possibility of fraud. This allows the user to immediately become aware of the possibility of fraud and prevent damage before it occurs.
[1608] A "scam call" is a call made with the intent of defrauding someone of money or personal information through fraudulent means.
[1609] An "incoming call" refers to an incoming phone call.
[1610] A "communication terminal" is a device for sending and receiving voice and data, and includes smartphones and home telephones.
[1611] An "audio input device" is a device for converting audio into digital data, and a microphone is an example of this.
[1612] "Conversation data" is data obtained by converting speech acquired by a speech input device into text.
[1613] "Generative AI" is artificial intelligence that analyzes input data and detects specific patterns and phrases.
[1614] An "audio output device" is a device for outputting digital data as audio, and a speaker is an example of this.
[1615] An "audio alert" is a warning sound or message emitted from an audio output device.
[1616] A "particular phrase or pattern" is a characteristic word or sentence structure that indicates possible fraud.
[1617] "User" refers to any individual or organization that uses this system.
[1618] A system for implementing this invention detects incoming fraudulent calls and listens to the conversation through a voice input device installed on a communication terminal. The conversation data is transcribed in real time and sent to a generative AI. The generative AI detects specific phrases or patterns from the conversation data that indicate the possibility of fraud, and if fraud is suspected, issues a voice alert from a voice output device.
[1619] Hardware and software used
[1620] Hardware: Communication devices such as smartphones and home phones, microphones (audio input devices), speakers (audio output devices)
[1621] software:
[1622] SpeechRecognition: A Python library for converting speech to text
[1623] requests: A Python library for sending HTTP requests
[1624] Generative AI: Artificial intelligence that analyzes text data using external APIs
[1625] System Operation
[1626] 1. Incoming call detection: When the communication terminal detects an incoming call, it starts recording the conversation.
[1627] 2. Conversation transcription: The recorded audio data is transcribed in real time using the SpeechRecognition library.
[1628] 3. Generative AI Analysis: The transcribed conversation data is sent via HTTP requests to a generative AI, which detects specific phrases or patterns that indicate potential fraud.
[1629] 4. Alert: If the generative AI detects a potential fraud, an audio alert will be emitted from the device's speaker to alert the user to the potential fraud.
[1630] Specific examples
[1631] For example, if a user receives a phone call and the caller says, "Your bank account has been compromised. Please provide me with your account information immediately," the system will transcribe this conversation in real time and send it to a generative AI that will detect that this phrase matches certain patterns that indicate potential fraud and immediately trigger an audio alert.
[1632] Prompt Sentence Examples
[1633] An example of a prompt for a generative AI model is:
[1634] "Determine whether the following text is a potential scam: 'Your bank account has been compromised. Please provide your account information immediately.'"
[1635] In this way, users can be made aware of potential fraud immediately and prevent damage before it occurs.
[1636] The flow of the specific processing in the application example 1 will be described with reference to FIG.
[1637] Step 1: Detect incoming calls
[1638] The communication terminal detects an incoming call. The input is the incoming call signal, and the output is a trigger to start recording. Specifically, the phone application of the communication terminal detects the incoming call and starts the recording module.
[1639] Step 2: Record the conversation
[1640] The audio input device (microphone) of the communication terminal records the conversation. The input is the audio data during the call, and the output is the recorded audio file. Specifically, the recording module captures the audio during the call and saves it as an audio file.
[1641] Step 3: Transcribe the conversation
[1642] The recorded audio data is transcribed using the SpeechRecognition library. The input is an audio file, and the output is text data. Specifically, the speech recognition engine analyzes the audio file and generates the corresponding text.
[1643] Step 4: Generative AI analysis
[1644] Transcribed text data is sent to the generative AI via an HTTP request. The input is text data, and the output is a flag indicating the possibility of fraud. Specifically, a request containing the text data is sent to the generative AI's API endpoint, and the analysis results are received.
[1645] Step 5: Determine the likelihood of fraud
[1646] Generative AI analyzes text data to detect specific phrases or patterns that indicate potential fraud. The input is text data, and the output is a flag indicating potential fraud. Specifically, generative AI analyzes text data and flags cases where there is a high probability of fraud.
[1647] Step 6: Send an alert
[1648] If a possible fraud is detected, an audio alert is issued from the audio output device (speaker) of the communication terminal. The input is a flag indicating the possible fraud, and the output is an audio alert. Specifically, the audio alert module receives the flag indicating the possible fraud and plays an audio message to warn the user.
[1649] In this way, users can be made aware of potential fraud immediately and prevent damage before it occurs.
[1650] Example 2
[1651] Next, a description will be given of Example 2 of Form Example 2. In the following description, the data processing device 12 will be referred to as a "server" and the headset type terminal 314 will be referred to as a "terminal."
[1652] In recent years, fraudulent phone calls have become more sophisticated, and many people have fallen victim to them. In particular, there has been an increase in special frauds targeting the elderly, and effective countermeasures against this are needed. Conventional countermeasures lack a system that can detect fraudulent phone calls in real time and issue warnings to users, making it difficult to prevent damage before it occurs.
[1653] The specific processing by the specific processing unit 290 of the data processing device 12 in the second embodiment is realized by the following means.
[1654] In this invention, the server includes means for detecting incoming fraudulent calls, means for collecting conversations from a voice input device installed in the communication device, means for converting the conversation data into text data, means for transmitting the text data to the generative AI model, and means for issuing a voice alert from the voice input device when the generative AI model determines that fraud is suspected from the conversation data. This makes it possible to detect fraudulent calls in real time and issue an immediate warning to the user.
[1655] The "means for detecting incoming fraudulent calls" is a function for identifying and detecting incoming calls that may be fraudulent calls to a communication device.
[1656] An "audio input device" is a microphone or other audio capture device installed on a communications device to capture a user's speech.
[1657] "Means for converting speech data into text data" refers to speech recognition software or algorithms used to analyze collected speech data and convert it into text format.
[1658] A "generative AI model" is an artificial intelligence model that analyzes input text data and determines the likelihood of fraud based on specific patterns and phrases.
[1659] "Audio alerting means" refers to a feature that plays an audio message to warn the user if fraud is suspected.
[1660] The present invention is a system for detecting incoming fraudulent calls in real time and issuing a warning to the user. A specific embodiment of this system will be described below.
[1661] 1. System Configuration
[1662] The system consists of the following major components:
[1663] How to detect incoming fraudulent calls
[1664] Voice input device
[1665] A means of converting conversation data into text data
[1666] Generative AI Models
[1667] A means of issuing an audio alert
[1668] 2. Hardware and Software Use
[1669] The voice input device is a microphone installed in a home telephone, which collects the user's speech in real time.
[1670] The means of converting conversation data into text data is through speech recognition software (e.g., Google Speech-to-Text API), which analyzes collected voice data and converts it into text data.
[1671] The generative AI model uses advanced artificial intelligence models such as OpenAI GPT-4, which analyzes input text data and determines the likelihood of fraud.
[1672] The audio alert method uses the speaker on the home phone, which will provide an audio alert if fraud is suspected.
[1673] 3. System Operation
[1674] When a user starts a conversation on their home phone, a voice input device collects the conversation. The collected voice data is converted into text data by speech recognition software. This text data is sent to a server and input into a generative AI model. The generative AI model analyzes the conversation data and, if it determines that fraud is suspected, the server sends an alert signal to the home phone. The home phone's speaker emits an audio alert to warn the user.
[1675] 4. Specific Examples
[1676] For example, imagine a user speaking on their home phone, "Hello, your bank account is at risk. Immediate action is required." A voice input device collects this speech, which speech recognition software converts into text data. This text data is sent to a server and analyzed by a generative AI model. If the generative AI model determines that there is a high possibility of fraud, the server sends an alert signal to the home phone. The speaker on the home phone issues a voice alert saying, "Possible fraud. Be careful."
[1677] 5. Examples of prompts
[1678] "Please analyze the conversation data below and determine whether you suspect fraud. If so, please explain why.
[1679] Conversation data: 'Hello, your bank account has been compromised. We need your immediate attention.'"
[1680] In this way, users can be made aware of their fraud risk in real time.
[1681] The flow of the identification process in the second embodiment will be described with reference to FIG.
[1682] Step 1:
[1683] Audio collection
[1684] Terminal: A voice input device (microphone) installed in a home phone collects the user's conversation in real time.
[1685] Input: User's voice conversation
[1686] Output: Digital audio data
[1687] What it does: When a user starts a phone conversation, the microphone automatically captures the audio and stores it as digital audio data.
[1688] Step 2:
[1689] Transcription of audio data
[1690] Device: Built-in speech recognition software (e.g., Google Speech-to-Text API) in home phones converts collected voice data into text data.
[1691] Input: Digital audio data
[1692] Output: Text data
[1693] What it does: Speech recognition software analyzes the audio and generates text that says something like, "Hello, your bank account has been compromised. Immediate action is required."
[1694] Step 3:
[1695] Sending text data
[1696] Terminal: The home phone sends the transcribed conversation data to the server.
[1697] Input: Text data
[1698] Output: Send data to the server
[1699] How it works: A home phone sends data over the internet to a server, which then passes it on to a generative AI model.
[1700] Step 4:
[1701] Fraud detection
[1702] Server: A generative AI model (e.g., OpenAI GPT-4) analyzes the received conversation data and determines whether fraud is suspected.
[1703] Input: Text data
[1704] Output: Determination of likelihood of fraud
[1705] What it does: A generative AI model analyzes the text "Hello, your bank account has been compromised. Immediate action is required" and determines that it is likely fraudulent.
[1706] Step 5:
[1707] Audio alert occurs
[1708] Server: If fraud is suspected, the server sends an alert signal to the home phone.
[1709] Terminal: The speaker on the home phone receives the alert signal from the server and issues an audio alert.
[1710] Input: Alert signal
[1711] Output: Audio alert
[1712] What it does: The speaker on your home phone plays a voice alert saying, "Possible scam. Be careful."
[1713] (Application example 2)
[1714] Next, a description will be given of Application Example 2 of Form Example 2. In the following description, the data processing device 12 will be referred to as a "server" and the headset type terminal 314 will be referred to as a "terminal."
[1715] In recent years, the number of victims of fraudulent phone calls has been increasing, and special frauds targeting the elderly in particular have become a social problem. Conventional fraud prevention systems have been inadequate in detecting fraudulent phone calls, making it difficult for users to recognize the possibility of fraud. Furthermore, there is a need for a system that can detect fraud in real time and issue a warning to users on home phones and smartphones.
[1716] The specific processing by the specific processing unit 290 of the data processing device 12 in the application example 2 is realized by the following means.
[1717] In this invention, the server includes means for detecting incoming fraudulent calls, means for listening to conversations using a voice recognition device, means for converting the conversation data into text and sending it to a generative AI, means for issuing a voice alert from a voice output device when the generative AI determines that fraud is suspected based on the conversation data, and means including an application to be installed on a smartphone. This enables real-time detection of fraudulent calls and prompt warning to users.
[1718] A "scam call" is a call made with the intent of defrauding someone of money or personal information through fraudulent means.
[1719] An "incoming call" refers to an incoming phone call.
[1720] A "voice recognition device" is a device that converts voice into text data.
[1721] "Conversation data" refers to the content of a conversation that has been converted into text by a voice recognition device.
[1722] "Generative AI" is artificial intelligence that generates new information based on input data.
[1723] An "audio output device" is a device for outputting audio.
[1724] "Voice alert" is a function that issues a warning by voice.
[1725] An "application installed on a smartphone" is a software program that runs on a smartphone.
[1726] To implement this invention, the following hardware and software are required. The hardware requires a smartphone, a voice recognition device, and a voice output device. The software requires a voice recognition library (e.g., speech_recognition), a generative AI library (e.g., openai), and a text-to-speech library (e.g., pyttsx3).
[1727] First, the phone call is recorded in real time using the smartphone's microphone. The recorded voice data is converted into text data using a voice recognition device. The speech_recognition library is used for this voice recognition.
[1728] The converted text data is then sent to the generative AI, which analyzes the input text data and determines whether it is fraudulent. This analysis is performed using the OpenAI library. Specifically, the following prompt is sent to the generative AI:
[1729] Please determine if the following conversation is a scam:
[1730] "Hello, this is bank security. Your account has been compromised. Please provide your account number and PIN."
[1731] If you suspect it is a scam, reply 'scam'.
[1732] If the generative AI determines that there is a possibility of fraud, it will issue a voice alert using the smartphone's audio output device. This voice alert uses the pyttsx3 library. The content of the voice alert is intended to alert the user to the possibility of fraud.
[1733] For example, if a user is on a call on their smartphone and a conversation that could be fraudulent takes place, the application automatically converts the conversation into text and sends it to generative AI to determine whether it is fraudulent. If it is determined that there is a possibility of fraud, an audio alert will sound from the smartphone speaker saying, "Possible fraud!"
[1734] In this way, the present invention provides real-time detection of fraudulent calls and prompt warning to users.
[1735] The flow of the specific processing in the application example 2 will be described with reference to FIG.
[1736] Step 1:
[1737] A user initiates a call on their smartphone, and the smartphone's microphone records the call in real time.
[1738] Input: Call audio
[1739] Output: Recorded audio data
[1740] How it works: The smartphone's microphone captures the call audio and saves it as audio data.
[1741] Step 2:
[1742] The terminal sends the recorded voice data to a voice recognition device, which converts the voice data into text data.
[1743] Input: Recorded audio data
[1744] Output: Text data
[1745] Specific operation: A speech recognition library (e.g., speech_recognition) analyzes the audio data and generates corresponding text data.
[1746] Step 3:
[1747] The device sends the generated text data to the generative AI, which analyzes the text data and determines whether it is fraudulent.
[1748] Input: Text data
[1749] Output: Judgment result on likelihood of fraud
[1750] How it works: A generative AI library (e.g., openai) analyzes the text data based on the prompt and determines whether it is likely to be fraudulent.
[1751] Step 4:
[1752] The server receives the judgment results from the generative AI and generates an audio alert if there is a possibility of fraud.
[1753] Input: Determination result regarding the possibility of fraud
[1754] Output: Audio alert content
[1755] Specific behavior: The generative AI generates a warning message saying, "Possible fraud!"
[1756] Step 5:
[1757] The terminal transmits the contents of the audio alert to the audio output device, and the audio alert is output.
[1758] Input: Audio alert content
[1759] Output: Audio alert
[1760] Specific operation: A text-to-speech library (e.g., pyttsx3) converts the contents of the voice alert into audio, and outputs the warning sound from the smartphone speaker.
[1761] Example 3
[1762] Next, a third embodiment of the third embodiment will be described. In the following description, the data processing device 12 will be referred to as a "server" and the headset type terminal 314 will be referred to as a "terminal."
[1763] Conventional fraud detection systems are limited to detecting incoming fraudulent phone calls, and are inadequate at detecting fraud in text messages and emails received by users. In addition, there are limited means of notifying users of possible fraud, which can delay users' realization of the risk of fraud. This makes it difficult to prevent fraud damage before it happens.
[1764] The specific processing by the specific processing unit 290 of the data processing device 12 in the third embodiment is realized by the following means.
[1765] In this invention, the server includes means for transmitting text data entered by a user to the server, means for the server to analyze the text data using a generative artificial intelligence model and determine the possibility of fraud, and means for the server to transmit the analysis results to the communication terminal, thereby making it possible to quickly detect the possibility of fraud even in text messages or emails received by the user and notify the user.
[1766] "Means for detecting incoming fraudulent calls" means a device or software for identifying and detecting potentially fraudulent calls from among calls coming in through a communication line.
[1767] "A voice input device installed on a communication terminal" refers to a microphone or related hardware for collecting voice that is attached to a communication device such as a smartphone or home phone.
[1768] "Means for converting conversation data into text and transmitting it to the generative AI" refers to software or hardware for converting voice data into text data and transmitting that text data to the generative AI.
[1769] "Generative AI" is an AI model that uses natural language processing technology to analyze text data and detect specific patterns and phrases.
[1770] "Audio alerting means" means a device or software that issues an audio warning when it determines that fraud may be occurring.
[1771] The "means for transmitting text data input by the user to the server" refers to software or hardware for transmitting text data input by the user to the communication terminal to the server via a communication line such as the Internet.
[1772] "Means for the server to analyze text data using a generative artificial intelligence model and determine the possibility of fraud" refers to software or hardware that inputs the text data received by the server into a generative artificial intelligence model and evaluates the possibility of fraud based on the analysis results.
[1773] The "means by which the server transmits the analysis results to the communication terminal" refers to software or hardware for transmitting the analysis results of the generative artificial intelligence model to the communication terminal.
[1774] The "means by which the communication terminal displays the analysis results to the user" refers to software or hardware for visually or audibly notifying the user of the analysis results received from the server.
[1775] The present invention relates to a system for detecting the possibility of special fraud using a generative artificial intelligence model. Specific embodiments of this system will be described below.
[1776] System configuration
[1777] Subject: Server
[1778] The server is responsible for analyzing the text data using a generative artificial intelligence model to determine the likelihood of fraud. Specifically, the server uses the following hardware and software:
[1779] Hardware: High-performance processors, memory, and storage devices
[1780] Software: Generative artificial intelligence models (e.g., AI models using natural language processing technology)
[1781] The server receives the text data sent by the user and inputs it into a generative AI model. The model detects specific phrases and patterns that indicate potential fraud and returns the results to the server. The server then transmits the analysis results to the communication device.
[1782] Subject: Terminal
[1783] The terminal is responsible for sending the text data entered by the user to the server and displaying the analysis results from the server to the user. Specifically, the terminal uses the following hardware and software:
[1784] Hardware: Communication devices such as smartphones, tablets, and PCs
[1785] Software: Dedicated application
[1786] The terminal transmits the text data entered by the user to the server, and displays the analysis results received from the server to the user. If the analysis results indicate a possibility of fraud, the terminal displays a warning message.
[1787] Subject: User
[1788] The user is responsible for inputting text data through the terminal and checking the analysis results. Specifically, the user performs the following operations:
[1789] Enter the contents of a message or email into your device
[1790] Check the analysis results and take appropriate action if necessary
[1791] Specific examples
[1792] Example 1
[1793] If a user receives a message saying "Mom, please transfer the money now," the system will act as follows:
[1794] 1. The user enters a message into a dedicated application on the device.
[1795] 2. The device sends this message to the server.
[1796] 3. The server receives the message and inputs it as a prompt sentence into the generative artificial intelligence model.
[1797] 4. The generative artificial intelligence model analyzes the message and returns the result: "Possible fraud."
[1798] 5. The server sends the analysis results to the device.
[1799] 6. The device displays the analysis results to the user, and the user confirms the warning message.
[1800] Example 2
[1801] If a user receives an email saying "Your bank account has been compromised. Contact us immediately," the system works as follows:
[1802] 1. The user enters the email content into a dedicated application on the device.
[1803] 2. The device sends the contents of this email to the server.
[1804] 3. The server receives the email content and inputs it as a prompt into the generative artificial intelligence model.
[1805] 4. The generative artificial intelligence model analyzes the content of the email and returns the result: "Possible fraud."
[1806] 5. The server sends the analysis results to the device.
[1807] 6. The device displays the analysis results to the user, and the user confirms the warning message.
[1808] Prompt Sentence Examples
[1809] Here are some example prompts to input to a generative AI model:
[1810] "Determine whether the following message is a potential scam: 'Mom, please transfer the money now.'"
[1811] "Determine if the following email is a potential scam: 'Your bank account has been compromised. Contact us now.'"
[1812] In this way, the specific fraud detection system using the generative artificial intelligence model operates specifically. The flow of the identification process in the third embodiment will be described with reference to FIG.
[1813] Step 1:
[1814] The user inputs text data.
[1815] The user enters the contents of the message or email they received into a dedicated application on their device. For example, they copy and paste the message "Mom, please transfer the money right now" into the application. The input data is saved in text format on the device.
[1816] Step 2:
[1817] The terminal transmits the text data to the server.
[1818] The terminal sends text data entered by the user to the server. Specifically, the terminal application generates an HTTP request and sends a payload containing the text data to the server. At this time, the data is encrypted before being sent. The input is the text data entered by the user, and the output is the text data sent to the server.
[1819] Step 3:
[1820] The server receives the text data and inputs it into a generative artificial intelligence model.
[1821] The server receives text data sent from the terminal. The received data is first stored in a database. The server then inputs the text data as a prompt sentence into the generative AI model. An example of a prompt sentence is "Please judge whether the following message is likely to be fraudulent: 'Mom, please transfer the money now.'" The input is the text data received from the terminal, and the output is the prompt sentence input into the generative AI model.
[1822] Step 4:
[1823] A generative artificial intelligence model analyzes the text data and returns the results.
[1824] The generative artificial intelligence model analyzes the input prompt sentence and determines whether it is likely to be fraud. Specifically, the model analyzes it by comparing it with typical phrases and patterns of "bank transfer fraud" and "it's me fraud." If the analysis result indicates a high possibility of fraud, it returns a warning message saying "Possible fraud." The input is the prompt sentence, and the output is the analysis result.
[1825] Step 5:
[1826] The server sends the analysis results to the device.
[1827] The server sends the analysis results obtained from the generative AI model to the terminal. Specifically, the server generates an HTTP response and sends a payload containing the analysis results to the terminal. At this time, the data is encrypted before being sent. The input is the analysis results from the generative AI model, and the output is the analysis results sent to the terminal.
[1828] Step 6:
[1829] The terminal displays the analysis results to the user.
[1830] The terminal displays the analysis results received from the server to the user. For example, if the analysis result is a warning message saying "Possible fraud," the terminal application displays this message to the user as a pop-up window or notification. The input is the analysis result received from the server, and the output is the warning message displayed to the user. The user can check this warning message and take appropriate action.
[1831] (Application example 3)
[1832] Next, a description will be given of Application Example 3 of Form Example 3. In the following description, the data processing device 12 will be referred to as a "server" and the headset type terminal 314 will be referred to as a "terminal."
[1833] In recent years, special fraud methods have become more sophisticated, and many people have fallen victim to them. In particular, "bank transfer fraud" and "it's me" frauds targeting the elderly have become a social problem. In order to prevent these frauds, a system is needed that can detect possible fraud early and issue a warning to users. However, current systems have low fraud detection accuracy, making it difficult for users to recognize the possibility of fraud.
[1834] The specific processing by the specific processing unit 290 of the data processing device 12 in the application example 3 is realized by the following means.
[1835] In this invention, the server includes means for detecting incoming fraudulent calls, means for listening to conversations from a voice input device installed in the communication terminal, means for converting the conversation data into text and sending it to the generative AI, means for issuing a voice alert from the voice input device when the generative AI determines that fraud is suspected from the conversation data, means for the generative AI to detect specific phrases or patterns that indicate the possibility of fraud, and means for issuing a warning to the user when the generative AI detects specific phrases or patterns that indicate the possibility of fraud. This makes it possible to detect the possibility of fraud with high accuracy and issue a warning to the user quickly.
[1836] "Means for detecting incoming fraudulent calls" is a function for identifying and notifying a communication terminal of incoming calls that may be fraudulent.
[1837] A "voice input device installed on a communication terminal" is a device for collecting voice that is attached to communication equipment such as a smartphone or home telephone.
[1838] "Means for converting conversation data into text and sending it to generative AI" refers to a function for converting voice data collected by a voice input device into text data and sending that text data to generative AI.
[1839] "Means for issuing a voice alert from the voice input device when the generative AI determines that fraud is suspected based on the conversation data" is a function for issuing a warning sound through the voice input device when the generative AI determines that there is a possibility of fraud based on the conversation data analyzed.
[1840] "Means for generative AI to detect specific phrases or patterns that indicate the possibility of fraud" refers to a function that allows generative AI to detect typical fraudulent phrases and patterns that it has previously learned from conversation data.
[1841] "Means for issuing a warning to users when generative AI detects specific phrases or patterns that indicate the possibility of fraud" refers to a function that issues a warning to users when generative AI detects phrases or patterns that indicate the possibility of fraud.
[1842] As an embodiment of the present invention, a method for installing a fraud detection assistant system on a smartphone will be described.
[1843] First, the system is equipped with a means for detecting incoming fraudulent calls. This means has the function of identifying and notifying the communication terminal of calls that may be fraudulent. Specifically, the system collects the contents of the call in real time using a voice input device installed on the communication terminal.
[1844] The collected voice data is then converted into text data through a means of transcribing the conversation data and sending it to the generative AI. This conversion is done using speech recognition software (e.g., Google's speech recognition API). The converted text data is then sent to the generative AI model.
[1845] A generative AI model is pre-trained to detect specific phrases and patterns that indicate potential fraud. This model is built using, for example, the Hugging Face transformers library. If the generative AI determines from the conversation data that fraud is suspected, it will warn the user by issuing an audio alert via a voice input device.
[1846] Additionally, if the generative AI detects certain phrases or patterns that indicate potential fraud, it will activate a user alert mechanism that will alert the user to the potential fraud by providing a visual or audio warning.
[1847] As a concrete example, consider the following audio input file:
[1848] Voiceover: "Mom, please transfer the money now. It's urgent."
[1849] An example prompt for parsing this audio file:
[1850] Determine whether the phrase "Mom, please transfer the money now. It's urgent." is a potential scam.
[1851] By feeding this prompt into a generative AI model, it can analyze the possibility of fraud and issue a warning to the user.
[1852] The flow of the specific processing in Application Example 3 will be described with reference to FIG.
[1853] Step 1:
[1854] Detecting incoming fraudulent calls
[1855] The server identifies and notifies the communication terminal of incoming calls that may be fraudulent. Specifically, it monitors the communication terminal's call log and detects possible fraud based on specific numbers or patterns. The input is the call log data, and the output is a notification of a potentially fraudulent call.
[1856] Step 2:
[1857] Listen to conversations from a voice input device
[1858] The terminal collects the contents of the call in real time using an audio input device installed in the communication terminal. Specifically, it acquires audio data through a microphone and saves it as an audio file. The input is the audio during the call, and the output is an audio file.
[1859] Step 3:
[1860] Convert conversation data into text and send it to generative AI
[1861] The device converts the collected voice data into text data and sends it to the generative AI. Specifically, it uses voice recognition software (e.g., Google's voice recognition API) to transcribe the voice data. The input is an audio file, and the output is text data.
[1862] Step 4:
[1863] Generative AI analyzes conversation data
[1864] The server uses a generative AI model to detect specific phrases or patterns in text data that indicate potential fraud. Specifically, it analyzes the text data using a generative AI model (e.g., Hugging Face's transformers library). The input is the text data, and the output is a determination of the likelihood of fraud.
[1865] Step 5:
[1866] Provides audio alerts in case of suspected fraud
[1867] If the generative AI determines that there is a possibility of fraud, the device will issue a voice alert through the voice input device. Specifically, it will play a warning sound or message. The input is the result of the judgment on the possibility of fraud, and the output is a voice alert.
[1868] Step 6:
[1869] Warn the user
[1870] The device will alert the user if the generative AI detects certain phrases or patterns that indicate potential fraud, either by displaying a warning message on the screen or by sending a notification. The input is the determination of potential fraud, and the output is the warning to the user.
[1871] Furthermore, an emotion engine that estimates the user's emotion may be further combined. That is, the identification processing unit 290 may estimate the user's emotion using the emotion identification model 59, and perform identification processing using the user's emotion.
[1872] "Example 1"
[1873] One embodiment of the present invention is a system that combines a generative AI with an emotion engine. This system analyzes emotions based on the user's tone of voice, volume, and speaking rate. Specifically, if the user's tone or volume suddenly increases while receiving a scam call, the system senses that the user may be panicking. Based on this information, the generative AI determines that the possibility of a scam has increased and issues an audio alert.
[1874] "Example 2"
[1875] In another embodiment, the emotion engine may generate an audio alert when a certain threshold is exceeded. For example, if the user's tone of voice exceeds a certain threshold or if the user's speaking rate exceeds a certain threshold, the emotion engine may determine that the user is experiencing strong emotion. Based on this determination, the generative AI may generate an audio alert to warn the user of a potential scam.
[1876] "Example 3"
[1877] Another possible embodiment is a system in which the emotion engine and generative AI work together. In this system, the emotion engine analyzes the user's emotions and sends the results to the generative AI. The generative AI then combines the information from the emotion engine with the conversation data it has analyzed to determine the possibility of fraud. This allows for more accurate fraud detection.
[1878] The processing flow of each embodiment will be described below.
[1879] "Example 1"
[1880] Step 1: The user receives a scam call.
[1881] Step 2: A smartphone app or a microphone installed on a home phone listens to the conversation.
[1882] Step 3: Convert the conversation data into text and send it to the generative AI.
[1883] Step 4: Generative AI analyzes the text data to determine the likelihood of fraud.
[1884] Step 5: The emotion engine analyzes the user's emotions based on their tone of voice, volume, speaking rate, etc.
[1885] Step 6: Based on the information from the emotion engine, the generative AI determines that the likelihood of fraud has increased and issues an audio alert.
[1886] "Example 2"
[1887] Step 1: The user receives a scam call.
[1888] Step 2: A smartphone app or a microphone installed on a home phone listens to the conversation.
[1889] Step 3: Convert the conversation data into text and send it to the generative AI.
[1890] Step 4: Generative AI analyzes the text data to determine the likelihood of fraud.
[1891] Step 5: The emotion engine analyzes the user's emotions based on their tone of voice, volume, speaking rate, etc.
[1892] Step 6: If the emotion engine exceeds a certain threshold, notify the generative AI.
[1893] Step 7: Based on the notification from the emotion engine, the generative AI determines that the likelihood of fraud has increased and issues an audio alert.
[1894] "Example 3"
[1895] Step 1: The user receives a scam call.
[1896] Step 2: A smartphone app or a microphone installed on a home phone listens to the conversation.
[1897] Step 3: Convert the conversation data into text and send it to the generative AI.
[1898] Step 4: Generative AI analyzes the text data to determine the likelihood of fraud.
[1899] Step 5: The emotion engine analyzes the user's emotions based on their tone of voice, volume, speaking speed, etc., and sends the results to the generative AI.
[1900] ...
Claims
1. means for detecting an incoming fraudulent call that may be fraudulent using a predetermined list of incoming calls, and notifying a communication terminal of the incoming fraudulent call; means for collecting and acquiring conversation data of the fraudulent call notified from a voice input device installed in the communication terminal; A means for transcribing the conversation data and inputting instructions for analyzing the transcribed conversation data to a generative AI; means for instructing the communication terminal to issue an audio alert by sending a command to play an audio alert to an audio output device of the communication terminal when the generative AI analyzes the conversation data and determines that fraud is suspected; means for collecting parameters relating to the user's voice in the conversation data and performing sentiment analysis; means for reassessing the likelihood of fraud based on the results of said sentiment analysis; A system including:
2. The system of claim 1 , wherein the generative AI determines that fraud is suspected when it detects a specific phrase or pattern that indicates the possibility of specialized fraud.
3. The system of claim 1 , wherein the audio alert emitted from the audio output device alerts a user to potential fraud.
Citation Information
Patent Citations
Data detection method and device, electronic equipment and storage medium
CN116739769A
Phishing scam prevention system
JP2007323107A
Ill-motivated telephone call prevention device and ill-motivated telephone call prevention system
JP2013005205A
Information processing system, information processing apparatus, information processing method, and program
JP2019153961A
Persona chatbot control method and system
JP2022180282A