System
A speech-to-text and generative AI-based system for real-time fraud detection and alerting family members and authorities addresses the challenge of 'it's me' frauds targeting the elderly, enhancing protection against such scams.
Patent Information
- Application Number
- JP2024133634
- Authority / Receiving Office
- JP · JP
- Patent Type
- Applications
- Current Assignee / Owner
- Filing Date
- 2024-08-08
- Publication Date
- 2026-02-20
AI Technical Summary
The increasing number of elderly victims of special frauds, particularly 'it's me' frauds over the phone, highlights the need for real-time fraud detection and warning systems, as conventional methods are limited to post-fraud warnings and lack effective prevention measures.
A system utilizing speech recognition technology and generative AI models to convert voice into text, analyze the text for fraud patterns, and send alerts to family members and the police if fraud is suspected.
The system effectively prevents elderly individuals from falling victim to fraud by detecting potential scams in real-time and issuing alerts, thereby protecting them from financial and psychological harm.
Smart Images

Figure 2026030650000001_ABST
Abstract
Description
[Technical Field]
[0001] The technology of the present disclosure relates to a system. [Background technology]
[0002] Patent document 1 discloses a persona chatbot control method performed by at least one processor, the method including the steps of receiving a user utterance, adding the user utterance to a prompt including an instruction sentence related to a description of the chatbot character, encoding the prompt, and inputting the encoded prompt into a language model to generate a chatbot utterance in response to the user utterance. [Prior art documents] [Patent documents]
[0003] [Patent Document 1] Japanese Patent Publication No. 2022-180282 Summary of the Invention [Problem to be solved by the invention]
[0004] The number of victims of special frauds targeting the elderly is increasing, especially in cases of "it's my son" frauds over the phone. Conventional countermeasures have been limited to issuing advance warnings, and there have been no effective means of detecting fraud in real time and issuing warnings. For this reason, there is a need for effective ways to prevent elderly people from falling victim to fraud before they become victims. The objective of this invention is to protect elderly people from fraud by utilizing speech recognition technology and generative AI models to detect conversations with a high probability of fraud in real time and send early alerts. [Means for solving the problem]
[0005] The present invention provides a system that includes a speech conversion means that constantly listens to speech and converts it into text data, an analysis means that analyzes the converted text data and determines whether fraud is likely, and a warning means that sends an alert if fraud is determined to be likely. Specifically, the analysis means uses a generative artificial intelligence model to compare the text data with past fraud patterns to evaluate the likelihood of fraud. The warning means is also characterized by having a function that sends an alert message to registered family members and the police if fraud is suspected. This system makes it possible to effectively prevent elderly people from falling victim to fraud before it happens.
[0006] The "voice conversion means" is a device or system that has the function of constantly listening to surrounding voices and converting them into text data.
[0007] An "analysis means" is a device or system that has the function of analyzing text data using a generative artificial intelligence model and determining the possibility of fraud based on the content of that data.
[0008] An "alert means" is a device or system that has the function of sending an alert message to registered family members or the police if it is determined that there is a possibility of fraud.
[0009] A "generative artificial intelligence model" is a type of artificial intelligence that learns from large amounts of data and analyzes and predicts new data, and is used in this invention to assess the possibility of fraud.
[0010] "Text data" is data obtained by converting collected voice information into text information using voice recognition technology, and is data to be analyzed by the analysis means.
[0011] An "alert message" is a warning message sent to family members or the police when a possible fraud situation is detected, and is information that can include specific conversation content.
[0012] "Fraud patterns" are models of typical fraudulent methods and speaking styles collected and analyzed based on past fraud cases, and are data that analysis methods refer to when determining the possibility of fraud. [Brief explanation of the drawings]
[0013] [Figure 1] 1 is a conceptual diagram showing an example of the configuration of a data processing system according to a first embodiment. [Figure 2] 1 is a conceptual diagram showing an example of main functions of a data processing device and a smart device according to a first embodiment. [Figure 3] FIG. 10 is a conceptual diagram showing an example of the configuration of a data processing system according to a second embodiment. [Figure 4] FIG. 10 is a conceptual diagram showing an example of main functions of a data processing device and smart glasses according to a second embodiment. [Figure 5] FIG. 10 is a conceptual diagram showing an example of the configuration of a data processing system according to a third embodiment. [Figure 6] FIG. 11 is a conceptual diagram showing an example of main functions of a data processing device and a headset-type terminal according to a third embodiment. [Figure 7] FIG. 10 is a conceptual diagram showing an example of the configuration of a data processing system according to a fourth embodiment. [Figure 8] FIG. 10 is a conceptual diagram showing an example of main functions of a data processing device and a robot according to a fourth embodiment. [Figure 9] 1 shows an emotion map onto which multiple emotions are mapped. [Figure 10] 1 shows an emotion map onto which multiple emotions are mapped. [Figure 11] FIG. 3 is a sequence diagram showing a processing flow of the data processing system according to the first embodiment. [Figure 12] FIG. 10 is a sequence diagram showing the flow of processing in the data processing system in Application Example 1. [Figure 13] FIG. 10 is a sequence diagram showing the flow of processing in the data processing system according to the second embodiment when an emotion engine is combined. [Figure 14]FIG. 10 is a sequence diagram showing the flow of processing in the data processing system in Application Example 2 when an emotion engine is combined. DETAILED DESCRIPTION OF THE INVENTION
[0014] An example of an embodiment of a system according to the technology of the present disclosure will be described below with reference to the accompanying drawings.
[0015] First, the terms used in the following description will be explained.
[0016] In the following embodiments, a coded processor (hereinafter simply referred to as a "processor") may be a single arithmetic device or a combination of multiple arithmetic devices. Furthermore, a processor may be a single type of arithmetic device or a combination of multiple types of arithmetic devices. Examples of arithmetic devices include a CPU (Central Processing Unit), a GPU (Graphics Processing Unit), a GPGPU (General-Purpose computing on Graphics Processing Units), and an APU (Accelerated Processing Unit).
[0017] In the following embodiments, a coded RAM (Random Access Memory) is a memory in which information is temporarily stored and is used as a working memory by a processor.
[0018] In the following embodiments, the coded storage is one or more non-volatile storage devices that store various programs, various parameters, etc. Examples of non-volatile storage devices include flash memory (SSD (Solid State Drive)), magnetic disks (e.g., hard disks), and magnetic tapes.
[0019] In the following embodiments, a communication I / F (Interface) with a symbol is an interface including a communication processor, an antenna, etc. The communication I / F controls communication between multiple computers. Examples of communication standards applied to the communication I / F include wireless communication standards including 5G (5th Generation Mobile Communication System), Wi-Fi (registered trademark), Bluetooth (registered trademark), etc.
[0020] In the following embodiments, "A and / or B" is synonymous with "at least one of A and B." In other words, "A and / or B" means that it may be only A, only B, or a combination of A and B. Furthermore, in this specification, the same concept as "A and / or B" is also applied when three or more things are expressed connected by "and / or."
[0021] [First embodiment]
[0022] FIG. 1 shows an example of the configuration of a data processing system 10 according to the first embodiment.
[0023] 1, a data processing system 10 includes a data processing device 12 and a smart device 14. An example of the data processing device 12 is a server.
[0024] The data processing device 12 includes a computer 22, a database 24, and a communication I / F 26. The computer 22 is an example of a "computer" according to the technology of the present disclosure. The computer 22 includes a processor 28, a RAM 30, and a storage 32. The processor 28, the RAM 30, and the storage 32 are connected to a bus 34. The database 24 and the communication I / F 26 are also connected to the bus 34. The communication I / F 26 is connected to a network 54. Examples of the network 54 include a WAN (Wide Area Network) and / or a LAN (Local Area Network).
[0025] The smart device 14 includes a computer 36, a reception device 38, an output device 40, a camera 42, and a communication I / F 44. The computer 36 includes a processor 46, a RAM 48, and a storage 50. The processor 46, the RAM 48, and the storage 50 are connected to a bus 52. The reception device 38, the output device 40, and the camera 42 are also connected to the bus 52.
[0026] The reception device 38 includes a touch panel 38A, a microphone 38B, and the like, and receives user input. The touch panel 38A detects contact with an indicator (for example, a pen or a finger) to receive user input by the touch of the indicator. The microphone 38B detects the user's voice to receive user input by voice. The control unit 46A transmits data indicating the user input received by the touch panel 38A and the microphone 38B to the data processing device 12. In the data processing device 12, the specific processing unit 290 acquires the data indicating the user input.
[0027] The output device 40 includes a display 40A and a speaker 40B, and presents data to the user 20 by outputting the data in a form of expression that the user 20 can perceive (for example, audio and / or text). The display 40A displays visible information such as text and images in accordance with instructions from the processor 46. The speaker 40B outputs audio in accordance with instructions from the processor 46. The camera 42 is a compact digital camera equipped with an optical system including a lens, aperture, and shutter, and an imaging element such as a CMOS (Complementary Metal-Oxide-Semiconductor) image sensor or a CCD (Charge Coupled Device) image sensor.
[0028] The communication I / F 44 is connected to a network 54. The communication I / Fs 44 and 26 control the exchange of various information between the processor 46 and the processor 28 via the network 54.
[0029] FIG. 2 shows an example of the main functions of the data processing device 12 and the smart device 14.
[0030] 2, in the data processing device 12, a specific process is performed by the processor 28. A specific processing program 56 is stored in the storage 32. The specific processing program 56 is an example of a "program" according to the technology of the present disclosure. The processor 28 reads the specific processing program 56 from the storage 32 and executes the read specific processing program 56 on the RAM 30. The specific process is realized by the processor 28 operating as a specific processing unit 290 in accordance with the specific processing program 56 executed on the RAM 30.
[0031] The storage 32 stores a data generation model 58 and an emotion identification model 59. The data generation model 58 and the emotion identification model 59 are used by the identification processing unit 290.
[0032] In the smart device 14, the processor 46 performs the reception output process. The storage 50 stores a reception output program 60. The reception output program 60 is used in conjunction with the specific processing program 56 by the data processing system 10. The processor 46 reads the reception output program 60 from the storage 50 and executes the read reception output program 60 on the RAM 48. The reception output process is realized by the processor 46 operating as the control unit 46A in accordance with the reception output program 60 executed on the RAM 48.
[0033] Next, a description will be given of the specific processing performed by the specific processing unit 290 of the data processing device 12. In the following description, the data processing device 12 will be referred to as a "server" and the smart device 14 will be referred to as a "terminal."
[0034] The present invention relates to a system for preventing special frauds, particularly "it's me" frauds, that target elderly people. Hereinafter, an embodiment of the present invention will be described in detail.
[0035] The present invention is a system that includes a voice conversion means that constantly listens to voice and converts it into text data, an analysis means that analyzes the converted text data and determines whether there is a possibility of fraud, and a warning means that sends an alert if it is determined that there is a possibility of fraud.
[0036] Audio conversion means
[0037] Subject: Terminal
[0038] The device constantly collects surrounding sounds through a microphone. This audio data is processed in real time, filtered for noise, and then converted into text data using a speech recognition algorithm. This text data is then sent to the analysis means.
[0039] As a concrete example, if a user is saying on the phone, "My son was in an accident and I need money," the device will collect this speech and convert it into text that says, "My son was in an accident and I need money."
[0040] Analysis means
[0041] Subject: Server
[0042] The server receives the text data sent from the device. The server uses a generative artificial intelligence model to analyze this text data. Specifically, the server compares it with a large amount of past fraud data to determine whether the content of the text data matches a fraud pattern. If it determines that there is a high possibility of fraud, the information is sent to a warning means.
[0043] In the above specific example, if the text data "My son was in an accident and I need money" matches the fraud pattern, the server determines that there is a high possibility of fraud.
[0044] warning means
[0045] Subject: Server
[0046] The server generates an alert message based on information that is judged to be highly likely to be fraudulent. This alert message includes the conversation content that is suspected to be fraudulent. The warning means sends this alert message to registered family members and the police. It also checks whether the alert message was sent successfully and records the sending result.
[0047] For example, family members and police may receive messages like this: "Possible scam. Dialogue: 'My son was in an accident and I need money.'"
[0048] Specific examples
[0049] If a user is having a conversation like this:
[0050] User: "Hello?"
[0051] Scammer: "Mom, help me! I've been in an accident and need money."
[0052] User: "Really? How did that happen?"
[0053] Scammer: "If you don't transfer the money now, I can't have the surgery!"
[0054] Audio conversion means
[0055] The device collects this conversation, filters out noise, and converts it into text data. The converted text reads: "Mom, help me! I've been in an accident and I need money." "Really? How did that happen?" "If you don't transfer the money now, I can't have the surgery!"
[0056] Analysis means
[0057] The server receives this text data and analyzes it using a generative artificial intelligence model, using ChatGPT to detect keywords such as "accident," "money," and "transfer," and determines whether the text matches a fraud pattern.
[0058] warning means
[0059] The server determines that the scam is likely and creates an alert message containing the potentially fraudulent content: "Mom, help me! I've been in an accident and need money." This message is sent to the family and the police.
[0060] This system can quickly alert family members and the police before elderly people become victims of fraud, effectively preventing fraud.
[0061] The processing flow will be explained below.
[0062] Step 1: Collecting audio
[0063] Subject: Terminal
[0064] The device constantly collects surrounding sounds through a microphone.
[0065] The collected audio data is stored in a buffer in real time.
[0066] Step 2: Filtering the noise
[0067] Subject: Terminal
[0068] The terminal performs noise reduction on the audio data in the buffer.
[0069] Filters out background sounds and noise to extract clear audio data.
[0070] Step 3: Speech to text
[0071] Subject: Terminal
[0072] The terminal uses a speech recognition algorithm to convert the filtered voice data into text data.
[0073] The converted text data is sent to the server.
[0074] Step 4: Receiving text data
[0075] Subject: Server
[0076] The server receives the text data sent from the terminal.
[0077] The received text data is stored in memory for analysis.
[0078] Step 5: Analyze potential fraud
[0079] Subject: Server
[0080] The server analyzes the text data using a generative artificial intelligence model (e.g., ChatGPT).
[0081] During the analysis process, the text data is compared with many past fraud patterns.
[0082] Assess whether there is potential for fraud.
[0083] Step 6: Determine the likelihood of fraud
[0084] Subject: Server
[0085] Based on the analysis results, the server determines that there is a possibility of fraud.
[0086] If there is a possibility of fraud, the information is sent to the alert module.
[0087] Step 7: Prepare alert information
[0088] Subject: Server
[0089] The server generates an alert message based on information that is determined to be potentially fraudulent.
[0090] The generated alert message includes the content of the suspected fraudulent conversation.
[0091] Step 8: Sending an alert
[0092] Subject: Server
[0093] The server will then send an alert message to registered family members and police contacts.
[0094] Check the transmission result, and if the transmission is successful, record the result in the log.
[0095] By using the above-mentioned processing steps, the present system can effectively warn elderly people before they fall victim to special fraud, thereby preventing them from becoming victims of fraud.
[0096] Example 1
[0097] Next, a description will be given of Example 1. In the following description, the data processing device 12 will be referred to as a "server" and the smart device 14 will be referred to as a "terminal."
[0098] In the current situation where the number of victims of special frauds targeting the elderly, especially "it's me" frauds, is increasing, there are many cases where victims transfer money without realizing it is a fraud, so a system to prevent this from happening is needed. The purpose of this invention is to protect the elderly from fraud by detecting such frauds in real time and promptly notifying family members and the police of suspected fraud.
[0099] The specific processing by the specific processing unit 290 of the data processing device 12 in the first embodiment is realized by the following means.
[0100] In this invention, the server includes a voice conversion means, an analysis means, and a warning means, which allows the server to convert voice into text data in real time, analyze the content of the text data, determine whether it is a fraud, and promptly notify family members and the police before the elderly person falls victim to fraud.
[0101] The "voice conversion means" is a means for collecting voice in real time, removing noise, and then converting the voice into text data.
[0102] The "analysis means" is a means for analyzing collected text data using a generative artificial intelligence model to determine the possibility of fraud.
[0103] The "warning means" is a means for generating an alert message and sending it to family members or the police if the analysis means determines that there is a possibility of fraud.
[0104] A "generative artificial intelligence model" is an artificial intelligence algorithm that analyzes text data based on a large amount of past fraud data and evaluates whether it matches a fraud pattern.
[0105] "Noise filtering" is a process for removing unwanted background noise from collected audio data.
[0106] A "prompt sentence" is an instruction sentence for a generative artificial intelligence model, and is an input sentence that allows the analysis means to understand the meaning of text data and perform appropriate analysis.
[0107] An "alert message" is a message containing a warning that is generated in the event of a possible fraud and is sent to family members or the police.
[0108] "Real-time" means that the processing occurs immediately, without delay.
[0109] The present invention is a system for preventing special frauds, particularly "it's me" frauds, that target elderly people. Hereinafter, an embodiment of the present invention will be described in detail.
[0110] System Overview
[0111] The system includes a voice conversion means that constantly listens to voice and converts it into text data, an analysis means that analyzes the converted text data and determines whether there is a possibility of fraud, and a warning means that sends an alert if it is determined that there is a possibility of fraud.
[0112] Audio conversion means
[0113] Hardware and Software
[0114] The device collects audio using a microphone. The recommended microphone is the Rode NT-USB, which has high-sensitivity noise-canceling technology. The device uses the open-source software Speex to filter noise from the collected audio data. The data is then converted to text using Google's Speech-to-Text API.
[0115] Example of operation
[0116] If a user receives a call from a scammer,
[0117] User: "Hello?"
[0118] Scammer: "Mom, help me! I've been in an accident and need money."
[0119] User: "Really? How did that happen?"
[0120] Scammer: "If you don't transfer the money now, I can't have the surgery!"
[0121] The device collects this conversation and
[0122] After removing the noise, the text is converted to "Mom, help me! I've been in an accident and need money," "Really? How did that happen?" and "If you don't transfer the money right now, I can't have the surgery!"
[0123] Analysis means
[0124] Hardware and Software
[0125] The server receives the text data sent from the device. The server analyzes it using a generative artificial intelligence model. OpenAI's "ChatGPT" is used as this AI model. The analysis method inputs the text data as a prompt sentence and compares it with a large amount of past fraud data to evaluate the possibility of fraud.
[0126] Prompt Sentence Examples
[0127] Prompt: "Analyze whether this text is a potential scam: 'My son was in an accident and I need money.'"
[0128] Example of operation
[0129] If the server determines that the text matches a fraud pattern based on the prompt above, it will assess the likelihood of fraud.
[0130] warning means
[0131] Hardware and Software
[0132] If the server determines that fraud is likely based on the analysis results, it generates an alert message containing the suspected fraudulent conversation content and sends the alert message to family members or the police using a secure communication protocol (HTTPS).
[0133] Example of operation
[0134] For example, the server could generate a message like this and send it to your family or the police:
[0135] "This could be a scam. The conversation goes like this: 'My son was in an accident and I need money.'"
[0136] Recording transmission results
[0137] The server checks whether the alert message was sent successfully and records the result.
[0138] summary
[0139] This system can collect voice recordings in real time before seniors fall victim to fraud, remove noise, convert the audio into text, and use a generative AI model to identify potential fraud. If a fraud is detected, a warning message can be sent quickly to family members or the police, effectively protecting seniors from fraud.
[0140] The flow of the identification process in the first embodiment will be described with reference to FIG.
[0141] Step 1: Collecting audio
[0142] Subject: Terminal
[0143] The device constantly collects surrounding sounds using a highly sensitive microphone (e.g., a microphone with noise-canceling technology). The microphone is activated and begins collecting audio data immediately after the user answers a call. The input is the audio of the conversation between the user and the fraudster, and the output is raw audio data.
[0144] Specific behavior:
[0145] From the moment the user says "Hello?", the device collects the voice and stores it as audio data.
[0146] Step 2: Noise filtering
[0147] Subject: Terminal
[0148] The device performs noise filtering on the collected audio data. It uses open-source software to remove noise. Raw audio data is used as input, and clear audio data after noise removal is generated as output.
[0149] Specific behavior:
[0150] It filters out background noise and reveals only the conversation between you and the fraudster.
[0151] Step 3: Voice Recognition
[0152] Subject: Terminal
[0153] The device then runs the noise-filtered audio data through a speech recognition algorithm to convert it into text data. This process uses the Speech-to-Text API, which takes the filtered audio data as input and produces the converted text data as output.
[0154] Specific behavior:
[0155] The conversation between the user and the scammer, such as "If you don't transfer the money now, I can't have the surgery!", is recorded as text.
[0156] Step 4: Sending text data
[0157] Subject: Terminal
[0158] The terminal sends the converted text data to the server. HTTPS is used as the communication protocol. The converted text data is sent as input and sent to the server.
[0159] Specific behavior:
[0160] The converted text, such as "Mom, help me! I've been in an accident and need money," is securely sent to a server.
[0161] Step 5: Analyzing the text data
[0162] Subject: Server
[0163] The server receives the text data sent from the device and inputs it into the generative AI model for analysis. The prompt sentence "Analyze whether this text is likely to be fraudulent" is used as input. Based on the input, the generative AI model performs analysis and obtains an output that evaluates the likelihood of fraud.
[0164] Specific behavior:
[0165] Enter the prompt text: "Analyze whether this text is likely to be fraudulent: 'My son was in an accident and I need money'" and match it with fraud patterns.
[0166] Step 6: Fraud detection
[0167] Subject: Server
[0168] The server receives the results of the generative AI model, and if there is a high possibility of fraud, it tags the information and forwards it to the warning means. The analysis results obtained from the generative AI model are used as input, and data with a fraud judgment is generated as output.
[0169] Specific behavior:
[0170] Generate result data such as "This text is likely to be fraudulent" and send it to the warning means.
[0171] Step 7: Generate an alert message
[0172] Subject: Server
[0173] The server generates an alert message based on the fraud detection data. This message contains the conversation content that is suspected to be fraudulent. The fraud detection data is used as input, and the alert message is generated as output.
[0174] Specific behavior:
[0175] Generates messages such as "Possible scam. Conversation: 'My son was in an accident and I need money.'"
[0176] Step 8: Sending an alert message
[0177] Subject: Server
[0178] The server sends the generated alert message to registered family members and the police. The communication method is SMS or email. The generated alert message is used as input and is sent to family members and the police as output.
[0179] Specific behavior:
[0180] The generated alert message is sent to family members or the police by referring to their contact list. For example, a message saying "Your son has been in an accident and you need money" is sent to a mother's smartphone.
[0181] Step 9: Record the transmission results
[0182] Subject: Server
[0183] The server checks whether the alert message was sent successfully and records the result in the database. It uses the sending log as input and generates the sending result record data as output.
[0184] Specific behavior:
[0185] Check whether the transmission was successful or failed and store the details in a recording database. For example, if the transmission was successful, it will be logged as "successful", and if the transmission failed, it will be logged as "failed".
[0186] (Application example 1)
[0187] Next, a description will be given of Application Example 1. In the following description, the data processing device 12 will be referred to as a "server" and the smart device 14 will be referred to as a "terminal."
[0188] In recent years, there has been an increase in special frauds targeting the elderly, particularly "it's me" frauds. This has caused significant economic and psychological damage to many elderly people and their families. Current security measures are often reactive, providing means to prevent damage after the fraud has already progressed. Against this background, there is a need for the development of an effective system that can detect fraudulent activities in real time and respond immediately.
[0189] The specific processing by the specific processing unit 290 of the data processing device 12 in the application example 1 is realized by the following means.
[0190] In this invention, the server includes a conversion means for constantly listening to voice and converting it into text data, an analysis means for analyzing the text data and determining the possibility of fraud, a warning means for sending an alert when it is determined that there is a possibility of fraud, and a communication means for sending the fraud analysis results to others based on the converted text data. This makes it possible to detect signs of fraud in real time and quickly issue a warning before any fraud damage occurs.
[0191] "Speech" refers to acoustic signals such as human speech and words.
[0192] "Conversion means" refers to a device or technology that converts speech into text data.
[0193] "Analysis means" refers to a device or technology that analyzes the converted text data and determines the possibility of fraud.
[0194] An "alert method" is a device or technology that sends an alert when it determines that fraud is likely.
[0195] "Communication means" refers to a device or technology for transmitting analysis results to other devices or users.
[0196] A "generative artificial intelligence model" is an artificial intelligence algorithm that has the ability to generate new data based on past data and recognize specific patterns.
[0197] A "fraud pattern" is a combination of language or phrases characteristic of fraudulent activity.
[0198] An "alert message" is a message containing a warning sent by a warning means.
[0199] "Family" refers to the user's blood relatives or close acquaintances.
[0200] A "police agency" is a public institution with legal authority to investigate crimes and protect the safety of citizens.
[0201] "Text data" is data in which voice is expressed as characters.
[0202] The present invention relates to a system for preventing special frauds targeting elderly people, and aims to prevent fraud in real time by constantly collecting voices, determining the possibility of fraud, and sending warnings to family members or police agencies if necessary.
[0203] Audio conversion means
[0204] The device constantly collects surrounding sounds using a microphone. The collected audio data is processed in real time, noise-filtered, and then converted into text data using a speech recognition algorithm. This text data is then sent to an analysis tool. Specifically, the device uses the Python speech_recognition library and Google's speech recognition service.
[0205] Analysis means
[0206] The server receives the text data sent from the device. The server then uses a generative AI model to analyze this text data. Specifically, it uses OpenAI's API to determine whether the text data matches a fraud pattern. This analysis involves evaluating the text based on past fraud data. The analysis method requests the generative AI model to analyze using a prompt sentence such as the following:
[0207] Prompt Sentence Examples
[0208] "Please determine if the following text is a scam: Mom, help me! I've been in an accident and need money."
[0209] warning means
[0210] If the server determines that there is a high possibility of fraud based on the analysis results, it generates an alert message. This message contains the suspected fraudulent conversation content. The warning mechanism then sends this alert message to registered family members and police agencies. The Python smtplib library is used to send the email. As an example of an alert, the following message is sent to family members:
[0211] "Possible scam. Dialogue: 'Mom, help me! I've been in an accident and need money.'"
[0212] Specific examples
[0213] For example, consider the following situation:
[0214] User: "Hello?"
[0215] Scammer: "Mom, help me! I've been in an accident and need money."
[0216] User: "Really? How did that happen?"
[0217] Scammer: "If you don't transfer the money now, I can't have the surgery!"
[0218] The device collects this conversation, performs speech recognition, and converts it into text data such as, "Mom, help me! I was in an accident and need money," "Really? How did that happen?", and "If you don't transfer the money now, I can't have the surgery!" This text data is then analyzed using a generative artificial intelligence model (such as OpenAI's API), and if it is determined to be a fraud, a warning message is sent to the family and police.
[0219] The system will enable seniors to take prompt action before they become victims of fraud, strengthening security measures.
[0220] The flow of the specific processing in the application example 1 will be described with reference to FIG.
[0221] Step 1:
[0222] The device uses a microphone to collect surrounding audio. The input is surrounding audio data, which the device filters in real time and converts into text data using a speech recognition algorithm. The output is the converted text data.
[0223] Step 2:
[0224] The terminal transmits the text data obtained in step 1 to the server. The input is the converted text data, and the output is the text data transmitted to the server. This data transmission utilizes an internet connection.
[0225] Step 3:
[0226] The server sends a prompt to the generative AI model to analyze the received text data. The input is the received text data, and the output is the analysis result obtained from the AI model. Specifically, the server uses the OpenAI API to generate the prompt text as follows:
[0227] "Please determine if the following text is a scam: [text data]"
[0228] Step 4:
[0229] The generative AI model analyzes the prompts sent from the server and determines whether the text data matches a fraud pattern. The input is a prompt containing the text data to be analyzed, and the output is the analysis result. Specifically, it compares the result with past fraud data.
[0230] Step 5:
[0231] The server receives the analysis results from the generative AI model and evaluates whether there is a high probability of fraud. The input is the analysis results from the AI model, and the output is the evaluation result indicating the possibility of fraud. Specifically, it checks whether the evaluation result meets certain criteria.
[0232] Step 6:
[0233] If it is determined that there is a high possibility of fraud, the server generates an alert message. The input is the evaluation result, and the output is the alert message. The specific operation is to generate a message containing the conversation content that is suspected to be fraudulent.
[0234] Step 7:
[0235] The server sends the generated alert message to the family or police. The input is the alert message, and the output is the message sent to the family or police. Specifically, it uses the Python smtplib library to send the email.
[0236] The above are the specific processing steps of the system embodying the present invention.
[0237] Furthermore, an emotion engine that estimates the user's emotion may be combined. That is, the identification processing unit 290 may estimate the user's emotion using the emotion identification model 59 and perform identification processing using the user's emotion.
[0238] This invention is a system that prevents victims of special frauds targeting the elderly, and in particular combines voice recognition technology, a generative AI model, and an emotion engine that recognizes the user's emotions.
[0239] System configuration
[0240] The system includes the following main means:
[0241] 1. Audio conversion method
[0242] 2. Analysis method
[0243] 3. Warning measures
[0244] 4. Emotion recognition means
[0245] Audio conversion means
[0246] Subject: Terminal
[0247] The device constantly collects surrounding sounds through a microphone. This audio data is processed in real time, noise-filtered, and then converted into text data using a speech recognition algorithm. This text data is then sent to the analysis means.
[0248] As a concrete example, if a user says on the phone, "My son was in an accident and I need money," the device will collect this speech and convert it into text: "My son was in an accident and I need money."
[0249] Analysis means
[0250] Subject: Server
[0251] The server receives text data sent from the device. The server analyzes this text data using a generative artificial intelligence model (e.g., ChatGPT). It compares this text with past fraud patterns and determines whether the text is likely to be fraudulent. The analysis results are sent to the warning mechanism.
[0252] For example, if the text "My son had an accident and I need money" matches a fraud pattern, the server determines that it is likely fraudulent.
[0253] emotion recognition means
[0254] Subject: Server
[0255] The server also has an emotion recognition engine that analyzes the user's emotional state from the voice data. This emotion recognition engine analyzes the user's emotions based on the tone, pitch, speed, etc. of the voice and detects emotional patterns that may indicate fraud. In cooperation with the analysis means, it performs a comprehensive evaluation that includes the user's emotional state.
[0256] For example, if emotions of anxiety or confusion are detected from the user's voice, this can be added to the evaluation to more accurately determine the possibility of fraud.
[0257] warning means
[0258] Subject: Server
[0259] The server generates an alert message based on the results of the analysis and emotion recognition. If it determines that there is a high possibility of fraud, it sends the alert message to family members or the police. This alert message includes the content of the conversation suspected of fraud and the evaluation result of the emotional state. The warning means checks whether the message was sent successfully and records the sending result in a log.
[0260] For example, a message sent to family members or police might read: "Possible scam. Dialogue: 'My son was in an accident and I need money.' User's emotional state: Anxiety."
[0261] Specific examples
[0262] If a user is having a conversation like this:
[0263] User: "Hello?"
[0264] Scammer: "Mom, help me! I've been in an accident and need money."
[0265] User: "Really? How did that happen?"
[0266] Scammer: "If you don't transfer the money now, I can't have the surgery!"
[0267] Audio conversion means
[0268] The device collects this conversation, filters out noise, and then converts it into text data.
[0269] Translated text: "Mom, help me! I was in an accident and need money." "Really? How did that happen?" "If you don't send the money now, I can't have my surgery!"
[0270] Analysis means
[0271] The server receives this text data and analyzes it using a generative artificial intelligence model.
[0272] ChatGPT is used to detect keywords such as "accident," "money," and "transfer," and determine whether this text matches a fraud pattern.
[0273] emotion recognition means
[0274] The server uses the voice data to analyze the user's emotional state.
[0275] The emotion recognition engine detects anxiety in the user's voice.
[0276] warning means
[0277] The server determines that there is a high probability of fraud and creates an alert message.
[0278] The message contains the potentially fraudulent content: "Mom, help me! I've been in an accident and need money" and the user's emotional state: "anxious."
[0279] This message will be sent to the family and police.
[0280] This system can effectively detect fraud before it happens to elderly people and quickly alert family members and the police, thereby preventing fraud from occurring.
[0281] The processing flow will be explained below.
[0282] Step 1: Collecting audio
[0283] Subject: Terminal
[0284] The device constantly collects surrounding sounds through a microphone.
[0285] The collected audio data is stored in a buffer in real time.
[0286] Step 2: Filtering the noise
[0287] Subject: Terminal
[0288] The terminal performs noise reduction on the audio data in the buffer.
[0289] Filters out background sounds and noise to extract clear audio data.
[0290] Step 3: Speech to text
[0291] Subject: Terminal
[0292] The terminal uses a speech recognition algorithm to convert the filtered voice data into text data.
[0293] The converted text data is sent to the server.
[0294] Step 4: Receiving text data
[0295] Subject: Server
[0296] The server receives the text data sent from the terminal.
[0297] The received text data is stored in memory for analysis.
[0298] Step 5: Analyze potential fraud
[0299] Subject: Server
[0300] The server analyzes the text data using a generative artificial intelligence model (e.g., ChatGPT).
[0301] During the analysis process, the text data is compared with many past fraud patterns.
[0302] Assess whether there is potential for fraud.
[0303] Step 6: Determine the likelihood of fraud
[0304] Subject: Server
[0305] Based on the analysis results, the server determines that there is a possibility of fraud.
[0306] If there is a possibility of fraud, the information is sent to the alert module.
[0307] Step 7: Collect emotion data
[0308] Subject: Terminal
[0309] The device uses voice data to collect data to analyze the user's emotional state using an emotion engine.
[0310] The collected emotion data is sent to a server.
[0311] Step 8: Analyze the sentiment data
[0312] Subject: Server
[0313] The server uses an emotion engine to analyze the emotion data sent from the terminal.
[0314] Evaluates the user's emotional state based on the tone, pitch, speed, etc. of the voice.
[0315] Detect emotional patterns that indicate potential fraud.
[0316] Step 9: Integrating and analyzing emotion data
[0317] Subject: Server
[0318] The server integrates the results of the text data analysis and the emotional data analysis, and comprehensively reassess the likelihood of fraud.
[0319] Determine with greater accuracy whether there is a high likelihood of fraud.
[0320] Step 10: Prepare alert information
[0321] Subject: Server
[0322] The server generates an alert message based on information that is determined to be potentially fraudulent.
[0323] The generated alert message includes the suspected fraudulent conversation content and an assessment of the emotional state.
[0324] Step 11: Sending an alert
[0325] Subject: Server
[0326] The server will then send an alert message to registered family members and police contacts.
[0327] Check the transmission result, and if the transmission is successful, record the result in the log.
[0328] Example 2
[0329] Next, a description will be given of Example 2. In the following description, the data processing device 12 will be referred to as a "server" and the smart device 14 will be referred to as a "terminal."
[0330] The number of victims of special frauds targeting the elderly is increasing year by year. Fraudsters use sophisticated methods to deceive the elderly and obtain money illegally. In order to prevent such fraud, a system is needed that can monitor conversations around the elderly in real time and detect suspicious speech. It is also desirable to evaluate the user's emotional state to detect factors that increase the likelihood of fraud.
[0331] The identification process by the identification processing unit 290 of the data processing device 12 in the second embodiment is realized by the following means. In this invention, the server includes a sound collection means that constantly collects environmental sounds and performs noise filtering, a sound recognition means that converts the collected sound data into text data, an analysis means that analyzes the text data and determines the possibility of fraud, an emotion recognition means that analyzes the emotional state of the user, and a warning means that sends an alert when it is determined that there is a possibility of fraud. This makes it possible to detect the risk of elderly people falling victim to fraud early and promptly notify their family or the police.
[0332] "Environmental sound" refers to all sounds occurring around the user, and is a general term for sound data including conversations and background sounds.
[0333] "Noise filtering" is a process that removes unnecessary background sounds and noise from collected audio data and converts it into audio data that is easy to hear.
[0334] "Sound collection means" refers to a device or software that constantly collects environmental sounds using a microphone.
[0335] "Speech recognition means" refers to an algorithm or system for converting collected voice data into text data.
[0336] "Text data" refers to character string data converted from speech by speech recognition means.
[0337] A "generative artificial intelligence model" refers to an algorithm that learns from large amounts of text data and performs natural language processing, and specifically includes models such as ChatGPT.
[0338] A "prompt sentence" refers to input text used to give instructions or ask questions to a generative artificial intelligence model and obtain a specific response.
[0339] "Analysis Method" refers to an algorithm or system that uses a generative artificial intelligence model to analyze text data and determine the likelihood of fraud.
[0340] "Emotion recognition means" refers to an algorithm or system for analyzing a user's emotional state from voice data and detecting emotional patterns.
[0341] "Warning measures" refer to algorithms or systems that send alert messages to family members or police in the event of potential fraud.
[0342] "Alert Message" means a message containing warning content that is generated to notify you of suspected fraud.
[0343] "Transmission result" refers to confirmation information as to whether the alert message was sent successfully.
[0344] "Log" refers to data that records system operations, events, and the results of sending alert messages.
[0345] The present invention is a system for preventing victims of special frauds targeting elderly people. The system includes a voice collection means, a voice recognition means, an analysis means, an emotion recognition means, and a warning means. The details of each means and the specific implementation method are explained below.
[0346] Audio collection method
[0347] The device is equipped with a microphone for constantly collecting ambient sounds. The microphone is highly sensitive and capable of collecting sound over a wide range, such as a condenser microphone or digital microphone. The collected sound data is processed in real time using a noise filtering algorithm, which removes background noise and unnecessary sounds to obtain clear sound data.
[0348] Voice recognition means
[0349] The device inputs the noise-filtered voice data into a voice recognition engine. For example, Google Cloud Speech-to-Text API is used as the voice recognition engine. The voice recognition engine converts the voice data into text data and expresses the content as a string of characters. The converted text data is sent to an analysis means.
[0350] Analysis means
[0351] The server receives the text data sent from the device. The server analyzes the text data using a generative artificial intelligence model (e.g., ChatGPT). It generates a prompt and inputs it into the model to compare it with past fraud patterns and evaluate the likelihood of fraud. For example, the prompt could be set as "Please evaluate the likelihood that this text is a special fraud." When the analysis results are obtained, it determines whether the likelihood of fraud is high and passes the result to the next processing step.
[0352] emotion recognition means
[0353] The server has an emotion recognition engine, such as IBM Watson Tone Analyzer, that analyzes the user's emotional state from the voice data. The emotion recognition engine analyzes the tone, pitch, and speed of the voice to assess whether the user is anxious or confused. These emotional states are also considered as factors that increase the likelihood of fraud.
[0354] warning means
[0355] The server generates an alert message based on the results of the analysis and emotion recognition methods. If it determines that there is a high possibility of fraud, it sends this alert message to the user's family or the police. The alert message includes the content of the conversation that is suspected to be fraudulent and the user's emotional state. For example, a message such as "Possible fraud. Conversation content: 'Mom, help me! I've been in an accident and need money.' User's emotional state: anxious" is generated. The server then checks the transmission result and records in a log that it was sent successfully.
[0356] Specific examples
[0357] Here is a concrete example of the system: If a user is having the following conversation:
[0358] User: "Hello?"
[0359] Scammer: "Mom, help me! I've been in an accident and need money."
[0360] User: "Really? How did that happen?"
[0361] Scammer: "If you don't transfer the money now, I can't have the surgery!"
[0362] 1. Audio collection means: The device collects the conversation and removes noise.
[0363] 2. Voice recognition method: After noise filtering, the conversation is converted into text such as "Mom, help me! I've been in an accident and need money," "Really? How did that happen?" and "If you don't transfer the money now I can't have the surgery!" and sent to the server.
[0364] 3. Analysis method: The server analyzes the text, detects keywords such as "accident," "money," and "transfer," and evaluates the likelihood of fraud. An example of a prompt is, "Please evaluate the likelihood that this text is a special fraud."
[0365] 4. Emotion recognition means: The server analyzes the user's emotion of anxiety from the voice.
[0366] 5. Warning method: The server determines that there is a high possibility of fraud and generates an alert message stating, "Possible fraud. Conversation: 'Mom, help! I've been in an accident and need money.' User's emotional state: Anxiety," and sends it to the family and police.
[0367] This system makes it possible to detect early the risk of elderly people falling victim to fraud and take prompt action.
[0368] The flow of the identification process in the second embodiment will be described with reference to FIG.
[0369] System program processing flow
[0370] Step 1: Audio Collection
[0371] The device collects ambient sounds, acquires audio data in real time through a microphone, and processes the data with a noise filtering algorithm. The input is the audio data of the ambient sounds, and the output is the audio data with the noise removed.
[0372] Specifically, the system captures audio data via a microphone and removes unwanted noise through a noise reduction filter.
[0373] Step 2: Voice Recognition
[0374] The device inputs the noise-filtered voice data into a voice recognition engine and converts it into text data. Here, for example, the Google Cloud Speech-to-Text API is used. The input is the voice data with noise removed, and the output is the converted text data.
[0375] Specifically, the speech data is sent to a speech recognition engine, and text data is obtained as a result.
[0376] Step 3: Text analysis
[0377] The server receives the text data sent from the device. The server analyzes this text data using a generative artificial intelligence model (such as ChatGPT). The input is the text data, and the output is the evaluation result of the likelihood of fraud. The prompt used is "Please evaluate the likelihood that this text is a special fraud."
[0378] Specifically, text data is input into a generative artificial intelligence model and the results of matching with fraud patterns are evaluated.
[0379] Step 4: Emotional state analysis
[0380] The server simultaneously analyzes the user's emotional state from the voice data. Using an emotion recognition engine (such as IBM Watson Tone Analyzer), it analyzes the tone, pitch, and speed of the voice to determine whether the user is anxious or confused. The input is the voice data, and the output is the evaluation result of the user's emotional state.
[0381] Specifically, the voice data is input into an emotion recognition engine to obtain an evaluation result of the emotional state.
[0382] Step 5: Generate and send an alert
[0383] The server generates an alert message based on the results of the analysis and emotion recognition methods. If it determines that there is a high possibility of fraud, it sends the alert message to the user's family or the police. The input is the evaluation result of the possibility of fraud and the evaluation result of the user's emotional state, and the output is the alert message.
[0384] Specifically, the system generates an alert message based on the analysis results and the user's emotional state, and sends it to family members or the police via email or SMS. The system then records the results of the message in a log.
[0385] This series of processes makes it possible to detect early on when an elderly person is at risk of falling victim to fraud and quickly notify their family or the police.
[0386] (Application example 2)
[0387] Next, a description will be given of Application Example 2. In the following description, the data processing device 12 will be referred to as a "server" and the smart device 14 will be referred to as a "terminal."
[0388] Special frauds targeting the elderly remain a social problem, with many elderly people suffering financial losses. The present invention aims to provide an effective system for preventing elderly people from falling victim to fraud. In particular, the objective of the present invention is to protect elderly people from fraud by detecting signs of fraud early and issuing prompt warnings.
[0389] The identification process by the identification processing unit 290 of the data processing device 12 in Application Example 2 is realized by the following means. In this invention, the server includes a voice conversion means that constantly collects voice and converts it into text data, an analysis means that analyzes the text data and evaluates the possibility of fraud, an emotion recognition means that analyzes the emotional state from the voice data, and a warning means that generates and transmits a warning message when it is determined that there is a possibility of fraud. This enables elderly people to quickly recognize the possibility of fraud and issue a warning to their family or public institutions.
[0390] The "voice conversion means" is a unit that constantly collects voice and converts it into text data.
[0391] The "analysis means" is a unit for analyzing the text data and assessing the possibility of fraud.
[0392] The "emotion recognition means" is a unit that analyzes the emotional state from the voice data.
[0393] The "warning means" is a unit for generating and transmitting a warning message when it is determined that there is a possibility of fraud.
[0394] A "generative artificial intelligence model" is an artificial intelligence algorithm generated from past data and learning, and is used as a means of analyzing text data.
[0395] "Text data" refers to data in character format converted from speech by speech conversion means.
[0396] A "fraud pattern" is a collection of characteristic words or phrases that indicate fraud.
[0397] The "emotional state" is the emotional state of the user that is analyzed based on the tone, pitch, speed, etc. of the voice data.
[0398] A "warning message" is a notification that is generated when potential fraud is detected.
[0399] "Family members and public authorities" are those who are notified by the warning measures of possible fraud.
[0400] The present invention is a system for preventing victims of special frauds targeting elderly people, and includes a voice conversion means, an analysis means, an emotion recognition means, and a warning means. Specific embodiments for carrying out the present invention will be described below.
[0401] System configuration
[0402] The system includes a dedicated terminal, a server, and a user terminal (e.g., a smartphone).
[0403] Audio conversion means
[0404] The device constantly collects ambient sound through a microphone. This sound data is processed in real time and noise filtered. Speech recognition software (e.g., the Python speech_recognition library) then converts the sound data into text data, which is then sent to an analysis unit.
[0405] Analysis means
[0406] The server receives text data sent from the device. Using a generative artificial intelligence model (e.g., ChatGPT), it analyzes this text data and matches it with fraud patterns. If the text indicates possible fraud, the results are sent to the emotion recognition means.
[0407] emotion recognition means
[0408] The server then uses an emotion recognition engine (e.g., EmotionRecognizer) to analyze the user's emotional state using the voice data. This engine analyzes the user's emotions based on the tone, pitch, and speed of the voice, and detects emotional patterns that may indicate fraud. It then makes a comprehensive assessment, including the user's emotional state.
[0409] warning means
[0410] If the server determines that a message is likely to be fraudulent, it generates a warning message that is sent to family members or public institutions. The warning message includes the content of the conversation that is suspected to be fraudulent and the result of the evaluation of the emotional state. The warning means checks whether the message was sent successfully and records the sending result in a log.
[0411] Specific examples
[0412] If a user is having a conversation like this:
[0413] User: "Hello?"
[0414] Scammer: "Mom, help me! I've been in an accident and need money."
[0415] User: "Really? How did that happen?"
[0416] Scammer: "If you don't transfer the money now, I can't have the surgery!"
[0417] Audio conversion means
[0418] The device collects this conversation, filters out noise, and converts it into text data, which looks like this:
[0419] "Mom, help me! I was in an accident and I need money." "Really? Why did that happen?" "If you don't transfer the money right now, I can't have the surgery!"
[0420] Analysis means
[0421] The server receives this text data and analyzes it using a generative artificial intelligence model. Because the text contains keywords such as "accident," "money," and "transfer," it is determined to be highly likely to be fraudulent.
[0422] emotion recognition means
[0423] The server uses the voice data to analyze the user's emotional state, and an emotion recognition engine detects anxiety from the user's voice.
[0424] warning means
[0425] The server determines that the fraud is likely and creates an alert message containing the suspected fraudulent content: "Mom, help me! I've been in an accident and need money" and the user's emotional state: "anxious." This message is sent to family members and public authorities.
[0426] Prompt Sentence Examples
[0427] "I received a suspicious call. I would like to provide the following information:
[0428] Conversation: 'My son had an accident and I need money.'
[0429] Emotional state: 'Anxious'
[0430] Please take immediate action as this may be a scam."
[0431] The flow of the specific processing in the application example 2 will be described with reference to FIG.
[0432] Step 1:
[0433] The device constantly collects ambient sound through a microphone. This sound data is processed in real time and noise filtered. The input is raw sound data, and the output is noise-removed sound data. Specifically, the sound signal obtained through the microphone is passed through a digital filter to reduce noise.
[0434] Step 2:
[0435] The device converts the noise-removed audio data into text data using speech recognition software (e.g., Python's speech_recognition library). The input is the noise-removed audio data, and the output is the corresponding text data. Specifically, the speech recognition algorithm analyzes the audio signal and converts it into the corresponding string of characters.
[0436] Step 3:
[0437] The terminal transmits the converted text data to the server. The input is the text data, and the output is data transmission to the server. Specifically, the text data is transferred to the server via a network connection.
[0438] Step 4:
[0439] The server analyzes the text data received from the device using a generative artificial intelligence model (e.g., ChatGPT) to assess the likelihood of fraud. The input is text data, and the output is an assessment of whether the likelihood of fraud is high. Specifically, the chat model compares the text with past fraud patterns to detect signs of fraud.
[0440] Step 5:
[0441] The server analyzes the voice data using an emotion recognition engine (e.g., EmotionRecognizer) to assess the user's emotional state. The input is the voice data, and the output is the assessment of the emotional state. Specifically, the tone, pitch, and speed of the voice are analyzed to identify specific emotions, such as anxiety.
[0442] Step 6:
[0443] The server comprehensively judges the fraud possibility assessment result and emotional state, and generates a warning message if it determines that fraud is highly likely. The input is the assessment result and emotional state, and the output is the warning message. Specifically, the warning message is created using a text template.
[0444] Step 7:
[0445] The server sends the generated warning message to the family or public institution. The input is the warning message, and the output is the message to be sent to the message recipient. Specifically, the server sends the warning message to the specified email address using an email protocol (e.g., SMTP).
[0446] Step 8:
[0447] The server checks whether the alert message was sent successfully and logs the sending result. The input is the sending status and the output is writing to the log file. Specifically, it checks the status of whether the sending was successful and records the result in the log file.
[0448] The specific processing unit 290 transmits the result of the specific processing to the smart device 14. In the smart device 14, the control unit 46A causes the output device 40 to output the result of the specific processing. The microphone 38B acquires audio indicating a user input regarding the result of the specific processing. The control unit 46A transmits audio data indicating the user input acquired by the microphone 38B to the data processing device 12. In the data processing device 12, the specific processing unit 290 acquires the audio data.
[0449] The data generation model 58 is a so-called generative AI (Artificial Intelligence). An example of the data generation model 58 is ChatGPT (Internet Search<URL: https: / / openai.com / blog / chatgpt> ), Gemini (Internet search <url: https: gemini.google.com ?hl="ja">) and other generation AIs. The data generation model 58 is obtained by performing deep learning on a neural network. A prompt including an instruction is input to the data generation model 58, and inference data such as voice data indicating voice, text data indicating text, and image data indicating an image is also input. The data generation model 58 performs inference on the input inference data in accordance with the instruction indicated by the prompt, and outputs the inference result in a data format such as voice data and text data. Here, inference refers to, for example, analysis, classification, prediction, and / or summarization.
[0450] In the above embodiment, an example in which the specific process is performed by the data processing device 12 has been given, but the technology of the present disclosure is not limited to this, and the specific process may be performed by the smart device 14.
[0451] [Second embodiment]
[0452] FIG. 3 shows an example of the configuration of a data processing system 210 according to the second embodiment.
[0453] 3, the data processing system 210 includes the data processing device 12 and smart glasses 214. An example of the data processing device 12 is a server.
[0454] The data processing device 12 includes a computer 22, a database 24, and a communication I / F 26. The computer 22 is an example of a "computer" according to the technology of the present disclosure. The computer 22 includes a processor 28, a RAM 30, and a storage 32. The processor 28, the RAM 30, and the storage 32 are connected to a bus 34. The database 24 and the communication I / F 26 are also connected to the bus 34. The communication I / F 26 is connected to a network 54. Examples of the network 54 include a WAN (Wide Area Network) and / or a LAN (Local Area Network).
[0455] The smart glasses 214 include a computer 36, a microphone 238, a speaker 240, a camera 42, and a communication I / F 44. The computer 36 includes a processor 46, a RAM 48, and a storage 50. The processor 46, the RAM 48, and the storage 50 are connected to a bus 52. The microphone 238, the speaker 240, and the camera 42 are also connected to the bus 52.
[0456] The microphone 238 receives instructions and the like from the user 20 by receiving voice uttered by the user 20. The microphone 238 captures the voice uttered by the user 20, converts the captured voice into audio data, and outputs it to the processor 46. The speaker 240 outputs audio in accordance with instructions from the processor 46.
[0457] Camera 42 is a small digital camera equipped with an optical system including a lens, aperture, and shutter, and an imaging element such as a CMOS (Complementary Metal-Oxide-Semiconductor) image sensor or a CCD (Charge Coupled Device) image sensor, and captures images of the surroundings of user 20 (for example, an imaging range defined by an angle of view equivalent to the field of vision of a typical healthy person).
[0458] The communication I / F 44 is connected to a network 54. The communication I / Fs 44 and 26 are responsible for the exchange of various information between the processor 46 and the processor 28 via the network 54. The exchange of various information between the processor 46 and the processor 28 using the communication I / Fs 44 and 26 is carried out in a secure state.
[0459] Fig. 4 shows an example of the main functions of the data processing device 12 and the smart glasses 214. As shown in Fig. 4, in the data processing device 12, a specific process is performed by the processor 28. A specific process program 56 is stored in the storage 32.
[0460] The specific processing program 56 is an example of a "program" according to the technology of the present disclosure. The processor 28 reads the specific processing program 56 from the storage 32 and executes the read specific processing program 56 on the RAM 30. The specific processing is realized by the processor 28 operating as a specific processing unit 290 in accordance with the specific processing program 56 executed on the RAM 30.
[0461] The storage 32 stores a data generation model 58 and an emotion identification model 59. The data generation model 58 and the emotion identification model 59 are used by the identification processing unit 290.
[0462] In the smart glasses 214, the reception output process is performed by the processor 46. A reception output program 60 is stored in the storage 50. The processor 46 reads the reception output program 60 from the storage 50 and executes the read reception output program 60 on the RAM 48. The reception output process is realized by the processor 46 operating as the control unit 46A in accordance with the reception output program 60 executed on the RAM 48.
[0463] Next, a description will be given of the identification process performed by the identification processing unit 290 of the data processing device 12. In the following description, the data processing device 12 will be referred to as the "server" and the smart glasses 214 will be referred to as the "terminal."
[0464] The present invention relates to a system for preventing special frauds, particularly "it's me" frauds, that target elderly people. Hereinafter, an embodiment of the present invention will be described in detail.
[0465] The present invention is a system that includes a voice conversion means that constantly listens to voice and converts it into text data, an analysis means that analyzes the converted text data and determines whether there is a possibility of fraud, and a warning means that sends an alert if it is determined that there is a possibility of fraud.
[0466] Audio conversion means
[0467] Subject: Terminal
[0468] The device constantly collects surrounding sounds through a microphone. This audio data is processed in real time, filtered for noise, and then converted into text data using a speech recognition algorithm. This text data is then sent to the analysis means.
[0469] As a concrete example, if a user is saying on the phone, "My son was in an accident and I need money," the device will collect this speech and convert it into text that says, "My son was in an accident and I need money."
[0470] Analysis means
[0471] Subject: Server
[0472] The server receives the text data sent from the device. The server uses a generative artificial intelligence model to analyze this text data. Specifically, the server compares it with a large amount of past fraud data to determine whether the content of the text data matches a fraud pattern. If it determines that there is a high possibility of fraud, the information is sent to a warning means.
[0473] In the above specific example, if the text data "My son was in an accident and I need money" matches the fraud pattern, the server determines that there is a high possibility of fraud.
[0474] warning means
[0475] Subject: Server
[0476] The server generates an alert message based on information that is judged to be highly likely to be fraudulent. This alert message includes the conversation content that is suspected to be fraudulent. The warning means sends this alert message to registered family members and the police. It also checks whether the alert message was sent successfully and records the sending result.
[0477] For example, family members and police may receive messages like this: "Possible scam. Dialogue: 'My son was in an accident and I need money.'"
[0478] Specific examples
[0479] If a user is having a conversation like this:
[0480] User: "Hello?"
[0481] Scammer: "Mom, help me! I've been in an accident and need money."
[0482] User: "Really? How did that happen?"
[0483] Scammer: "If you don't transfer the money now, I can't have the surgery!"
[0484] Audio conversion means
[0485] The device collects this conversation, filters out noise, and converts it into text data. The converted text reads: "Mom, help me! I've been in an accident and I need money." "Really? How did that happen?" "If you don't transfer the money now, I can't have the surgery!"
[0486] Analysis means
[0487] The server receives this text data and analyzes it using a generative artificial intelligence model, using ChatGPT to detect keywords such as "accident," "money," and "transfer," and determines whether the text matches a fraud pattern.
[0488] warning means
[0489] The server determines that the scam is likely and creates an alert message containing the potentially fraudulent content: "Mom, help me! I've been in an accident and need money." This message is sent to the family and the police.
[0490] This system can quickly alert family members and the police before elderly people become victims of fraud, effectively preventing fraud.
[0491] The processing flow will be explained below.
[0492] Step 1: Collecting audio
[0493] Subject: Terminal
[0494] The device constantly collects surrounding sounds through a microphone.
[0495] The collected audio data is stored in a buffer in real time.
[0496] Step 2: Filtering the noise
[0497] Subject: Terminal
[0498] The terminal performs noise reduction on the audio data in the buffer.
[0499] Filters out background sounds and noise to extract clear audio data.
[0500] Step 3: Speech to text
[0501] Subject: Terminal
[0502] The terminal uses a speech recognition algorithm to convert the filtered voice data into text data.
[0503] The converted text data is sent to the server.
[0504] Step 4: Receiving text data
[0505] Subject: Server
[0506] The server receives the text data sent from the terminal.
[0507] The received text data is stored in memory for analysis.
[0508] Step 5: Analyze potential fraud
[0509] Subject: Server
[0510] The server analyzes the text data using a generative artificial intelligence model (e.g., ChatGPT).
[0511] During the analysis process, the text data is compared with many past fraud patterns.
[0512] Assess whether there is potential for fraud.
[0513] Step 6: Determine the likelihood of fraud
[0514] Subject: Server
[0515] Based on the analysis results, the server determines that there is a possibility of fraud.
[0516] If there is a possibility of fraud, the information is sent to the alert module.
[0517] Step 7: Prepare alert information
[0518] Subject: Server
[0519] The server generates an alert message based on information that is determined to be potentially fraudulent.
[0520] The generated alert message includes the content of the suspected fraudulent conversation.
[0521] Step 8: Sending an alert
[0522] Subject: Server
[0523] The server will then send an alert message to registered family members and police contacts.
[0524] Check the transmission result, and if the transmission is successful, record the result in the log.
[0525] By using the above-mentioned processing steps, the present system can effectively warn elderly people before they fall victim to special fraud, thereby preventing them from becoming victims of fraud.
[0526] Example 1
[0527] Next, a description will be given of Example 1. In the following description, the data processing device 12 will be referred to as a "server" and the smart glasses 214 will be referred to as a "terminal."
[0528] In the current situation where the number of victims of special frauds targeting the elderly, especially "it's me" frauds, is increasing, there are many cases where victims transfer money without realizing it is a fraud, so a system to prevent this from happening is needed. The purpose of this invention is to protect the elderly from fraud by detecting such frauds in real time and promptly notifying family members and the police of suspected fraud.
[0529] The specific processing by the specific processing unit 290 of the data processing device 12 in the first embodiment is realized by the following means.
[0530] In this invention, the server includes a voice conversion means, an analysis means, and a warning means, which allows the server to convert voice into text data in real time, analyze the content of the text data, determine whether it is a fraud, and promptly notify family members and the police before the elderly person falls victim to fraud.
[0531] The "voice conversion means" is a means for collecting voice in real time, removing noise, and then converting the voice into text data.
[0532] The "analysis means" is a means for analyzing collected text data using a generative artificial intelligence model to determine the possibility of fraud.
[0533] The "warning means" is a means for generating an alert message and sending it to family members or the police if the analysis means determines that there is a possibility of fraud.
[0534] A "generative artificial intelligence model" is an artificial intelligence algorithm that analyzes text data based on a large amount of past fraud data and evaluates whether it matches a fraud pattern.
[0535] "Noise filtering" is a process for removing unwanted background noise from collected audio data.
[0536] A "prompt sentence" is an instruction sentence for a generative artificial intelligence model, and is an input sentence that allows the analysis means to understand the meaning of text data and perform appropriate analysis.
[0537] An "alert message" is a message containing a warning that is generated in the event of a possible fraud and is sent to family members or the police.
[0538] "Real-time" means that the processing occurs immediately, without delay.
[0539] The present invention is a system for preventing special frauds, particularly "it's me" frauds, that target elderly people. Hereinafter, an embodiment of the present invention will be described in detail.
[0540] System Overview
[0541] The system includes a voice conversion means that constantly listens to voice and converts it into text data, an analysis means that analyzes the converted text data and determines whether there is a possibility of fraud, and a warning means that sends an alert if it is determined that there is a possibility of fraud.
[0542] Audio conversion means
[0543] Hardware and Software
[0544] The device collects audio using a microphone. The recommended microphone is the Rode NT-USB, which has high-sensitivity noise-canceling technology. The device uses the open-source software Speex to filter noise from the collected audio data. The data is then converted to text using Google's Speech-to-Text API.
[0545] Example of operation
[0546] If a user receives a call from a scammer,
[0547] User: "Hello?"
[0548] Scammer: "Mom, help me! I've been in an accident and need money."
[0549] User: "Really? How did that happen?"
[0550] Scammer: "If you don't transfer the money now, I can't have the surgery!"
[0551] The device collects this conversation and
[0552] After removing the noise, the text is converted to "Mom, help me! I've been in an accident and need money," "Really? How did that happen?" and "If you don't transfer the money right now, I can't have the surgery!"
[0553] Analysis means
[0554] Hardware and Software
[0555] The server receives the text data sent from the device. The server analyzes it using a generative artificial intelligence model. OpenAI's "ChatGPT" is used as this AI model. The analysis method inputs the text data as a prompt sentence and compares it with a large amount of past fraud data to evaluate the possibility of fraud.
[0556] Prompt Sentence Examples
[0557] Prompt: "Analyze whether this text is a potential scam: 'My son was in an accident and I need money.'"
[0558] Example of operation
[0559] If the server determines that the text matches a fraud pattern based on the prompt above, it will assess the likelihood of fraud.
[0560] warning means
[0561] Hardware and Software
[0562] If the server determines that fraud is likely based on the analysis results, it generates an alert message containing the suspected fraudulent conversation content and sends the alert message to family members or the police using a secure communication protocol (HTTPS).
[0563] Example of operation
[0564] For example, the server could generate a message like this and send it to your family or the police:
[0565] "This could be a scam. The conversation goes like this: 'My son was in an accident and I need money.'"
[0566] Recording transmission results
[0567] The server checks whether the alert message was sent successfully and records the result.
[0568] summary
[0569] This system can collect voice recordings in real time before seniors fall victim to fraud, remove noise, convert the audio into text, and use a generative AI model to identify potential fraud. If a fraud is detected, a warning message can be sent quickly to family members or the police, effectively protecting seniors from fraud.
[0570] The flow of the identification process in the first embodiment will be described with reference to FIG.
[0571] Step 1: Collecting audio
[0572] Subject: Terminal
[0573] The device constantly collects surrounding sounds using a highly sensitive microphone (e.g., a microphone with noise-canceling technology). The microphone is activated and begins collecting audio data immediately after the user answers a call. The input is the audio of the conversation between the user and the fraudster, and the output is raw audio data.
[0574] Specific behavior:
[0575] From the moment the user says "Hello?", the device collects the voice and stores it as audio data.
[0576] Step 2: Noise filtering
[0577] Subject: Terminal
[0578] The device performs noise filtering on the collected audio data. It uses open-source software to remove noise. Raw audio data is used as input, and clear audio data after noise removal is generated as output.
[0579] Specific behavior:
[0580] It filters out background noise and reveals only the conversation between you and the fraudster.
[0581] Step 3: Voice Recognition
[0582] Subject: Terminal
[0583] The device then runs the noise-filtered audio data through a speech recognition algorithm to convert it into text data. This process uses the Speech-to-Text API, which takes the filtered audio data as input and produces the converted text data as output.
[0584] Specific behavior:
[0585] The conversation between the user and the scammer, such as "If you don't transfer the money now, I can't have the surgery!", is recorded as text.
[0586] Step 4: Sending text data
[0587] Subject: Terminal
[0588] The terminal sends the converted text data to the server. HTTPS is used as the communication protocol. The converted text data is sent as input and sent to the server.
[0589] Specific behavior:
[0590] The converted text, such as "Mom, help me! I've been in an accident and need money," is securely sent to a server.
[0591] Step 5: Analyzing the text data
[0592] Subject: Server
[0593] The server receives the text data sent from the device and inputs it into the generative AI model for analysis. The prompt sentence "Analyze whether this text is likely to be fraudulent" is used as input. Based on the input, the generative AI model performs analysis and obtains an output that evaluates the likelihood of fraud.
[0594] Specific behavior:
[0595] Enter the prompt text: "Analyze whether this text is likely to be fraudulent: 'My son was in an accident and I need money'" and match it with fraud patterns.
[0596] Step 6: Fraud detection
[0597] Subject: Server
[0598] The server receives the results of the generative AI model, and if there is a high possibility of fraud, it tags the information and forwards it to the warning means. The analysis results obtained from the generative AI model are used as input, and data with a fraud judgment is generated as output.
[0599] Specific behavior:
[0600] Generate result data such as "This text is likely to be fraudulent" and send it to the warning means.
[0601] Step 7: Generate an alert message
[0602] Subject: Server
[0603] The server generates an alert message based on the fraud detection data. This message contains the conversation content that is suspected to be fraudulent. The fraud detection data is used as input, and the alert message is generated as output.
[0604] Specific behavior:
[0605] Generates messages such as "Possible scam. Conversation: 'My son was in an accident and I need money.'"
[0606] Step 8: Sending an alert message
[0607] Subject: Server
[0608] The server sends the generated alert message to registered family members and the police. The communication method is SMS or email. The generated alert message is used as input and is sent to family members and the police as output.
[0609] Specific behavior:
[0610] The generated alert message is sent to family members or the police by referring to their contact list. For example, a message saying "Your son has been in an accident and you need money" is sent to a mother's smartphone.
[0611] Step 9: Record the transmission results
[0612] Subject: Server
[0613] The server checks whether the alert message was sent successfully and records the result in the database. It uses the sending log as input and generates the sending result record data as output.
[0614] Specific behavior:
[0615] Check whether the transmission was successful or failed and store the details in a recording database. For example, if the transmission was successful, it will be logged as "successful", and if the transmission failed, it will be logged as "failed".
[0616] (Application example 1)
[0617] Next, a description will be given of Application Example 1. In the following description, the data processing device 12 will be referred to as a "server" and the smart glasses 214 will be referred to as a "terminal."
[0618] In recent years, there has been an increase in special frauds targeting the elderly, particularly "it's me" frauds. This has caused significant economic and psychological damage to many elderly people and their families. Current security measures are often reactive, providing means to prevent damage after the fraud has already progressed. Against this background, there is a need for the development of an effective system that can detect fraudulent activities in real time and respond immediately.
[0619] The specific processing by the specific processing unit 290 of the data processing device 12 in the application example 1 is realized by the following means.
[0620] In this invention, the server includes a conversion means for constantly listening to voice and converting it into text data, an analysis means for analyzing the text data and determining the possibility of fraud, a warning means for sending an alert when it is determined that there is a possibility of fraud, and a communication means for sending the fraud analysis results to others based on the converted text data. This makes it possible to detect signs of fraud in real time and quickly issue a warning before any fraud damage occurs.
[0621] "Speech" refers to acoustic signals such as human speech and words.
[0622] "Conversion means" refers to a device or technology that converts speech into text data.
[0623] "Analysis means" refers to a device or technology that analyzes the converted text data and determines the possibility of fraud.
[0624] An "alert method" is a device or technology that sends an alert when it determines that fraud is likely.
[0625] "Communication means" refers to a device or technology for transmitting analysis results to other devices or users.
[0626] A "generative artificial intelligence model" is an artificial intelligence algorithm that has the ability to generate new data based on past data and recognize specific patterns.
[0627] A "fraud pattern" is a combination of language or phrases characteristic of fraudulent activity.
[0628] An "alert message" is a message containing a warning sent by a warning means.
[0629] "Family" refers to the user's blood relatives or close acquaintances.
[0630] A "police agency" is a public institution with legal authority to investigate crimes and protect the safety of citizens.
[0631] "Text data" is data in which voice is expressed as characters.
[0632] The present invention relates to a system for preventing special frauds targeting elderly people, and aims to prevent fraud in real time by constantly collecting voices, determining the possibility of fraud, and sending warnings to family members or police agencies if necessary.
[0633] Audio conversion means
[0634] The device constantly collects surrounding sounds using a microphone. The collected audio data is processed in real time, noise-filtered, and then converted into text data using a speech recognition algorithm. This text data is then sent to an analysis tool. Specifically, the device uses the Python speech_recognition library and Google's speech recognition service.
[0635] Analysis means
[0636] The server receives the text data sent from the device. The server then uses a generative AI model to analyze this text data. Specifically, it uses OpenAI's API to determine whether the text data matches a fraud pattern. This analysis involves evaluating the text based on past fraud data. The analysis method requests the generative AI model to analyze using a prompt sentence such as the following:
[0637] Prompt Sentence Examples
[0638] "Please determine if the following text is a scam: Mom, help me! I've been in an accident and need money."
[0639] warning means
[0640] If the server determines that there is a high possibility of fraud based on the analysis results, it generates an alert message. This message contains the suspected fraudulent conversation content. The warning mechanism then sends this alert message to registered family members and police agencies. The Python smtplib library is used to send the email. As an example of an alert, the following message is sent to family members:
[0641] "Possible scam. Dialogue: 'Mom, help me! I've been in an accident and need money.'"
[0642] Specific examples
[0643] For example, consider the following situation:
[0644] User: "Hello?"
[0645] Scammer: "Mom, help me! I've been in an accident and need money."
[0646] User: "Really? How did that happen?"
[0647] Scammer: "If you don't transfer the money now, I can't have the surgery!"
[0648] The device collects this conversation, performs speech recognition, and converts it into text data such as, "Mom, help me! I was in an accident and need money," "Really? How did that happen?", and "If you don't transfer the money now, I can't have the surgery!" This text data is then analyzed using a generative artificial intelligence model (such as OpenAI's API), and if it is determined to be a fraud, a warning message is sent to the family and police.
[0649] The system will enable seniors to take prompt action before they become victims of fraud, strengthening security measures.
[0650] The flow of the specific processing in the application example 1 will be described with reference to FIG.
[0651] Step 1:
[0652] The device uses a microphone to collect surrounding audio. The input is surrounding audio data, which the device filters in real time and converts into text data using a speech recognition algorithm. The output is the converted text data.
[0653] Step 2:
[0654] The terminal transmits the text data obtained in step 1 to the server. The input is the converted text data, and the output is the text data transmitted to the server. This data transmission utilizes an internet connection.
[0655] Step 3:
[0656] The server sends a prompt to the generative AI model to analyze the received text data. The input is the received text data, and the output is the analysis result obtained from the AI model. Specifically, the server uses the OpenAI API to generate the prompt text as follows:
[0657] "Please determine if the following text is a scam: [text data]"
[0658] Step 4:
[0659] The generative AI model analyzes the prompts sent from the server and determines whether the text data matches a fraud pattern. The input is a prompt containing the text data to be analyzed, and the output is the analysis result. Specifically, it compares the result with past fraud data.
[0660] Step 5:
[0661] The server receives the analysis results from the generative AI model and evaluates whether there is a high probability of fraud. The input is the analysis results from the AI model, and the output is the evaluation result indicating the possibility of fraud. Specifically, it checks whether the evaluation result meets certain criteria.
[0662] Step 6:
[0663] If it is determined that there is a high possibility of fraud, the server generates an alert message. The input is the evaluation result, and the output is the alert message. The specific operation is to generate a message containing the conversation content that is suspected to be fraudulent.
[0664] Step 7:
[0665] The server sends the generated alert message to the family or police. The input is the alert message, and the output is the message sent to the family or police. Specifically, it uses the Python smtplib library to send the email.
[0666] The above are the specific processing steps of the system embodying the present invention.
[0667] Furthermore, an emotion engine that estimates the user's emotion may be further combined. That is, the identification processing unit 290 may estimate the user's emotion using the emotion identification model 59, and perform identification processing using the user's emotion.
[0668] This invention is a system that prevents victims of special frauds targeting the elderly, and in particular combines voice recognition technology, a generative AI model, and an emotion engine that recognizes the user's emotions.
[0669] System configuration
[0670] The system includes the following main means:
[0671] 1. Audio conversion method
[0672] 2. Analysis method
[0673] 3. Warning measures
[0674] 4. Emotion recognition means
[0675] Audio conversion means
[0676] Subject: Terminal
[0677] The device constantly collects surrounding sounds through a microphone. This audio data is processed in real time, noise-filtered, and then converted into text data using a speech recognition algorithm. This text data is then sent to the analysis means.
[0678] As a concrete example, if a user says on the phone, "My son was in an accident and I need money," the device will collect this speech and convert it into text: "My son was in an accident and I need money."
[0679] Analysis means
[0680] Subject: Server
[0681] The server receives text data sent from the device. The server analyzes this text data using a generative artificial intelligence model (e.g., ChatGPT). It compares this text with past fraud patterns and determines whether the text is likely to be fraudulent. The analysis results are sent to the warning mechanism.
[0682] For example, if the text "My son had an accident and I need money" matches a fraud pattern, the server determines that it is likely fraudulent.
[0683] emotion recognition means
[0684] Subject: Server
[0685] The server also has an emotion recognition engine that analyzes the user's emotional state from the voice data. This emotion recognition engine analyzes the user's emotions based on the tone, pitch, speed, etc. of the voice and detects emotional patterns that may indicate fraud. In cooperation with the analysis means, it performs a comprehensive evaluation that includes the user's emotional state.
[0686] For example, if emotions of anxiety or confusion are detected from the user's voice, this can be added to the evaluation to more accurately determine the possibility of fraud.
[0687] warning means
[0688] Subject: Server
[0689] The server generates an alert message based on the results of the analysis and emotion recognition. If it determines that there is a high possibility of fraud, it sends the alert message to family members or the police. This alert message includes the content of the conversation suspected of fraud and the evaluation result of the emotional state. The warning means checks whether the message was sent successfully and records the sending result in a log.
[0690] For example, a message sent to family members or police might read: "Possible scam. Dialogue: 'My son was in an accident and I need money.' User's emotional state: Anxiety."
[0691] Specific examples
[0692] If a user is having a conversation like this:
[0693] User: "Hello?"
[0694] Scammer: "Mom, help me! I've been in an accident and need money."
[0695] User: "Really? How did that happen?"
[0696] Scammer: "If you don't transfer the money now, I can't have the surgery!"
[0697] Audio conversion means
[0698] The device collects this conversation, filters out noise, and then converts it into text data.
[0699] Translated text: "Mom, help me! I was in an accident and need money." "Really? How did that happen?" "If you don't send the money now, I can't have my surgery!"
[0700] Analysis means
[0701] The server receives this text data and analyzes it using a generative artificial intelligence model.
[0702] ChatGPT is used to detect keywords such as "accident," "money," and "transfer," and determine whether this text matches a fraud pattern.
[0703] emotion recognition means
[0704] The server uses the voice data to analyze the user's emotional state.
[0705] The emotion recognition engine detects anxiety in the user's voice.
[0706] warning means
[0707] The server determines that there is a high probability of fraud and creates an alert message.
[0708] The message contains the potentially fraudulent content: "Mom, help me! I've been in an accident and need money" and the user's emotional state: "anxious."
[0709] This message will be sent to the family and police.
[0710] This system can effectively detect fraud before it happens to elderly people and quickly alert family members and the police, thereby preventing fraud from occurring.
[0711] The processing flow will be explained below.
[0712] Step 1: Collecting audio
[0713] Subject: Terminal
[0714] The device constantly collects surrounding sounds through a microphone.
[0715] The collected audio data is stored in a buffer in real time.
[0716] Step 2: Filtering the noise
[0717] Subject: Terminal
[0718] The terminal performs noise reduction on the audio data in the buffer.
[0719] Filters out background sounds and noise to extract clear audio data.
[0720] Step 3: Speech to text
[0721] Subject: Terminal
[0722] The terminal uses a speech recognition algorithm to convert the filtered voice data into text data.
[0723] The converted text data is sent to the server.
[0724] Step 4: Receiving text data
[0725] Subject: Server
[0726] The server receives the text data sent from the terminal.
[0727] The received text data is stored in memory for analysis.
[0728] Step 5: Analyze potential fraud
[0729] Subject: Server
[0730] The server analyzes the text data using a generative artificial intelligence model (e.g., ChatGPT).
[0731] During the analysis process, the text data is compared with many past fraud patterns.
[0732] Assess whether there is potential for fraud.
[0733] Step 6: Determine the likelihood of fraud
[0734] Subject: Server
[0735] Based on the analysis results, the server determines that there is a possibility of fraud.
[0736] If there is a possibility of fraud, the information is sent to the alert module.
[0737] Step 7: Collect emotion data
[0738] Subject: Terminal
[0739] The device uses voice data to collect data to analyze the user's emotional state using an emotion engine.
[0740] The collected emotion data is sent to a server.
[0741] Step 8: Analyze the sentiment data
[0742] Subject: Server
[0743] The server uses an emotion engine to analyze the emotion data sent from the terminal.
[0744] Evaluates the user's emotional state based on the tone, pitch, speed, etc. of the voice.
[0745] Detect emotional patterns that indicate potential fraud.
[0746] Step 9: Integrating and analyzing emotion data
[0747] Subject: Server
[0748] The server integrates the results of the text data analysis and the emotional data analysis, and comprehensively reassess the likelihood of fraud.
[0749] Determine with greater accuracy whether there is a high likelihood of fraud.
[0750] Step 10: Prepare alert information
[0751] Subject: Server
[0752] The server generates an alert message based on information that is determined to be potentially fraudulent.
[0753] The generated alert message includes the suspected fraudulent conversation content and an assessment of the emotional state.
[0754] Step 11: Sending an alert
[0755] Subject: Server
[0756] The server will then send an alert message to registered family members and police contacts.
[0757] Check the transmission result, and if the transmission is successful, record the result in the log.
[0758] Example 2
[0759] Next, a description will be given of Example 2. In the following description, the data processing device 12 will be referred to as a "server" and the smart glasses 214 will be referred to as a "terminal."
[0760] The number of victims of special frauds targeting the elderly is increasing year by year. Fraudsters use sophisticated methods to deceive the elderly and obtain money illegally. In order to prevent such fraud, a system is needed that can monitor conversations around the elderly in real time and detect suspicious speech. It is also desirable to evaluate the user's emotional state to detect factors that increase the likelihood of fraud.
[0761] The identification process by the identification processing unit 290 of the data processing device 12 in the second embodiment is realized by the following means. In this invention, the server includes a sound collection means that constantly collects environmental sounds and performs noise filtering, a sound recognition means that converts the collected sound data into text data, an analysis means that analyzes the text data and determines the possibility of fraud, an emotion recognition means that analyzes the emotional state of the user, and a warning means that sends an alert when it is determined that there is a possibility of fraud. This makes it possible to detect the risk of elderly people falling victim to fraud early and promptly notify their family or the police.
[0762] "Environmental sound" refers to all sounds occurring around the user, and is a general term for sound data including conversations and background sounds.
[0763] "Noise filtering" is a process that removes unnecessary background sounds and noise from collected audio data and converts it into audio data that is easy to hear.
[0764] "Sound collection means" refers to a device or software that constantly collects environmental sounds using a microphone.
[0765] "Speech recognition means" refers to an algorithm or system for converting collected voice data into text data.
[0766] "Text data" refers to character string data converted from speech by speech recognition means.
[0767] A "generative artificial intelligence model" refers to an algorithm that learns from large amounts of text data and performs natural language processing, and specifically includes models such as ChatGPT.
[0768] A "prompt sentence" refers to input text used to give instructions or ask questions to a generative artificial intelligence model and obtain a specific response.
[0769] "Analysis Method" refers to an algorithm or system that uses a generative artificial intelligence model to analyze text data and determine the likelihood of fraud.
[0770] "Emotion recognition means" refers to an algorithm or system for analyzing a user's emotional state from voice data and detecting emotional patterns.
[0771] "Warning measures" refer to algorithms or systems that send alert messages to family members or police in the event of potential fraud.
[0772] "Alert Message" means a message containing warning content that is generated to notify you of suspected fraud.
[0773] "Transmission result" refers to confirmation information as to whether the alert message was sent successfully.
[0774] "Log" refers to data that records system operations, events, and the results of sending alert messages.
[0775] The present invention is a system for preventing victims of special frauds targeting elderly people. The system includes a voice collection means, a voice recognition means, an analysis means, an emotion recognition means, and a warning means. The details of each means and the specific implementation method are explained below.
[0776] Audio collection method
[0777] The device is equipped with a microphone for constantly collecting ambient sounds. The microphone is highly sensitive and capable of collecting sound over a wide range, such as a condenser microphone or digital microphone. The collected sound data is processed in real time using a noise filtering algorithm, which removes background noise and unnecessary sounds to obtain clear sound data.
[0778] Voice recognition means
[0779] The device inputs the noise-filtered voice data into a voice recognition engine. For example, Google Cloud Speech-to-Text API is used as the voice recognition engine. The voice recognition engine converts the voice data into text data and expresses the content as a string of characters. The converted text data is sent to an analysis means.
[0780] Analysis means
[0781] The server receives the text data sent from the device. The server analyzes the text data using a generative artificial intelligence model (e.g., ChatGPT). It generates a prompt and inputs it into the model to compare it with past fraud patterns and evaluate the likelihood of fraud. For example, the prompt could be set as "Please evaluate the likelihood that this text is a special fraud." When the analysis results are obtained, it determines whether the likelihood of fraud is high and passes the result to the next processing step.
[0782] emotion recognition means
[0783] The server has an emotion recognition engine, such as IBM Watson Tone Analyzer, that analyzes the user's emotional state from the voice data. The emotion recognition engine analyzes the tone, pitch, and speed of the voice to assess whether the user is anxious or confused. These emotional states are also considered as factors that increase the likelihood of fraud.
[0784] warning means
[0785] The server generates an alert message based on the results of the analysis and emotion recognition methods. If it determines that there is a high possibility of fraud, it sends this alert message to the user's family or the police. The alert message includes the content of the conversation that is suspected to be fraudulent and the user's emotional state. For example, a message such as "Possible fraud. Conversation content: 'Mom, help me! I've been in an accident and need money.' User's emotional state: anxious" is generated. The server then checks the transmission result and records in a log that it was sent successfully.
[0786] Specific examples
[0787] Here is a concrete example of the system: If a user is having the following conversation:
[0788] User: "Hello?"
[0789] Scammer: "Mom, help me! I've been in an accident and need money."
[0790] User: "Really? How did that happen?"
[0791] Scammer: "If you don't transfer the money now, I can't have the surgery!"
[0792] 1. Audio collection means: The device collects the conversation and removes noise.
[0793] 2. Voice recognition method: After noise filtering, the conversation is converted into text such as "Mom, help me! I've been in an accident and need money," "Really? How did that happen?" and "If you don't transfer the money now I can't have the surgery!" and sent to the server.
[0794] 3. Analysis method: The server analyzes the text, detects keywords such as "accident," "money," and "transfer," and evaluates the likelihood of fraud. An example of a prompt is, "Please evaluate the likelihood that this text is a special fraud."
[0795] 4. Emotion recognition means: The server analyzes the user's emotion of anxiety from the voice.
[0796] 5. Warning method: The server determines that there is a high possibility of fraud and generates an alert message stating, "Possible fraud. Conversation: 'Mom, help! I've been in an accident and need money.' User's emotional state: Anxiety," and sends it to the family and police.
[0797] This system makes it possible to detect early the risk of elderly people falling victim to fraud and take prompt action.
[0798] The flow of the identification process in the second embodiment will be described with reference to FIG.
[0799] System program processing flow
[0800] Step 1: Audio Collection
[0801] The device collects ambient sounds, acquires audio data in real time through a microphone, and processes the data with a noise filtering algorithm. The input is the audio data of the ambient sounds, and the output is the audio data with the noise removed.
[0802] Specifically, the system captures audio data via a microphone and removes unwanted noise through a noise reduction filter.
[0803] Step 2: Voice Recognition
[0804] The device inputs the noise-filtered voice data into a voice recognition engine and converts it into text data. Here, for example, the Google Cloud Speech-to-Text API is used. The input is the voice data with noise removed, and the output is the converted text data.
[0805] Specifically, the speech data is sent to a speech recognition engine, and text data is obtained as a result.
[0806] Step 3: Text analysis
[0807] The server receives the text data sent from the device. The server analyzes this text data using a generative artificial intelligence model (such as ChatGPT). The input is the text data, and the output is the evaluation result of the likelihood of fraud. The prompt used is "Please evaluate the likelihood that this text is a special fraud."
[0808] Specifically, text data is input into a generative artificial intelligence model and the results of matching with fraud patterns are evaluated.
[0809] Step 4: Emotional state analysis
[0810] The server simultaneously analyzes the user's emotional state from the voice data. Using an emotion recognition engine (such as IBM Watson Tone Analyzer), it analyzes the tone, pitch, and speed of the voice to determine whether the user is anxious or confused. The input is the voice data, and the output is the evaluation result of the user's emotional state.
[0811] Specifically, the voice data is input into an emotion recognition engine to obtain an evaluation result of the emotional state.
[0812] Step 5: Generate and send an alert
[0813] The server generates an alert message based on the results of the analysis and emotion recognition methods. If it determines that there is a high possibility of fraud, it sends the alert message to the user's family or the police. The input is the evaluation result of the possibility of fraud and the evaluation result of the user's emotional state, and the output is the alert message.
[0814] Specifically, the system generates an alert message based on the analysis results and the user's emotional state, and sends it to family members or the police via email or SMS. The system then records the results of the message in a log.
[0815] This series of processes makes it possible to detect early on when an elderly person is at risk of falling victim to fraud and quickly notify their family or the police.
[0816] (Application example 2)
[0817] Next, a description will be given of Application Example 2. In the following description, the data processing device 12 will be referred to as a "server" and the smart glasses 214 will be referred to as a "terminal."
[0818] Special frauds targeting the elderly remain a social problem, with many elderly people suffering financial losses. The present invention aims to provide an effective system for preventing elderly people from falling victim to fraud. In particular, the objective of the present invention is to protect elderly people from fraud by detecting signs of fraud early and issuing prompt warnings.
[0819] The identification process by the identification processing unit 290 of the data processing device 12 in Application Example 2 is realized by the following means. In this invention, the server includes a voice conversion means that constantly collects voice and converts it into text data, an analysis means that analyzes the text data and evaluates the possibility of fraud, an emotion recognition means that analyzes the emotional state from the voice data, and a warning means that generates and transmits a warning message when it is determined that there is a possibility of fraud. This enables elderly people to quickly recognize the possibility of fraud and issue a warning to their family or public institutions.
[0820] The "voice conversion means" is a unit that constantly collects voice and converts it into text data.
[0821] The "analysis means" is a unit for analyzing the text data and assessing the possibility of fraud.
[0822] The "emotion recognition means" is a unit that analyzes the emotional state from the voice data.
[0823] The "warning means" is a unit for generating and transmitting a warning message when it is determined that there is a possibility of fraud.
[0824] A "generative artificial intelligence model" is an artificial intelligence algorithm generated from past data and learning, and is used as a means of analyzing text data.
[0825] "Text data" refers to data in character format converted from speech by speech conversion means.
[0826] A "fraud pattern" is a collection of characteristic words or phrases that indicate fraud.
[0827] The "emotional state" is the emotional state of the user that is analyzed based on the tone, pitch, speed, etc. of the voice data.
[0828] A "warning message" is a notification that is generated when potential fraud is detected.
[0829] "Family members and public authorities" are those who are notified by the warning measures of possible fraud.
[0830] The present invention is a system for preventing victims of special frauds targeting elderly people, and includes a voice conversion means, an analysis means, an emotion recognition means, and a warning means. Specific embodiments for carrying out the present invention will be described below.
[0831] System configuration
[0832] The system includes a dedicated terminal, a server, and a user terminal (e.g., a smartphone).
[0833] Audio conversion means
[0834] The device constantly collects ambient sound through a microphone. This sound data is processed in real time and noise filtered. Speech recognition software (e.g., the Python speech_recognition library) then converts the sound data into text data, which is then sent to an analysis unit.
[0835] Analysis means
[0836] The server receives text data sent from the device. Using a generative artificial intelligence model (e.g., ChatGPT), it analyzes this text data and matches it with fraud patterns. If the text indicates possible fraud, the results are sent to the emotion recognition means.
[0837] emotion recognition means
[0838] The server then uses an emotion recognition engine (e.g., EmotionRecognizer) to analyze the user's emotional state using the voice data. This engine analyzes the user's emotions based on the tone, pitch, and speed of the voice, and detects emotional patterns that may indicate fraud. It then makes a comprehensive assessment, including the user's emotional state.
[0839] warning means
[0840] If the server determines that a message is likely to be fraudulent, it generates a warning message that is sent to family members or public institutions. The warning message includes the content of the conversation that is suspected to be fraudulent and the result of the evaluation of the emotional state. The warning means checks whether the message was sent successfully and records the sending result in a log.
[0841] Specific examples
[0842] If a user is having a conversation like this:
[0843] User: "Hello?"
[0844] Scammer: "Mom, help me! I've been in an accident and need money."
[0845] User: "Really? How did that happen?"
[0846] Scammer: "If you don't transfer the money now, I can't have the surgery!"
[0847] Audio conversion means
[0848] The device collects this conversation, filters out noise, and converts it into text data, which looks like this:
[0849] "Mom, help me! I was in an accident and I need money." "Really? Why did that happen?" "If you don't transfer the money right now, I can't have the surgery!"
[0850] Analysis means
[0851] The server receives this text data and analyzes it using a generative artificial intelligence model. Because the text contains keywords such as "accident," "money," and "transfer," it is determined to be highly likely to be fraudulent.
[0852] emotion recognition means
[0853] The server uses the voice data to analyze the user's emotional state, and an emotion recognition engine detects anxiety from the user's voice.
[0854] warning means
[0855] The server determines that the fraud is likely and creates an alert message containing the suspected fraudulent content: "Mom, help me! I've been in an accident and need money" and the user's emotional state: "anxious." This message is sent to family members and public authorities.
[0856] Prompt Sentence Examples
[0857] "I received a suspicious call. I would like to provide the following information:
[0858] Conversation: 'My son had an accident and I need money.'
[0859] Emotional state: 'Anxious'
[0860] Please take immediate action as this may be a scam."
[0861] The flow of the specific processing in the application example 2 will be described with reference to FIG.
[0862] Step 1:
[0863] The device constantly collects ambient sound through a microphone. This sound data is processed in real time and noise filtered. The input is raw sound data, and the output is noise-removed sound data. Specifically, the sound signal obtained through the microphone is passed through a digital filter to reduce noise.
[0864] Step 2:
[0865] The device converts the noise-removed audio data into text data using speech recognition software (e.g., Python's speech_recognition library). The input is the noise-removed audio data, and the output is the corresponding text data. Specifically, the speech recognition algorithm analyzes the audio signal and converts it into the corresponding string of characters.
[0866] Step 3:
[0867] The terminal transmits the converted text data to the server. The input is the text data, and the output is data transmission to the server. Specifically, the text data is transferred to the server via a network connection.
[0868] Step 4:
[0869] The server analyzes the text data received from the device using a generative artificial intelligence model (e.g., ChatGPT) to assess the likelihood of fraud. The input is text data, and the output is an assessment of whether the likelihood of fraud is high. Specifically, the chat model compares the text with past fraud patterns to detect signs of fraud.
[0870] Step 5:
[0871] The server analyzes the voice data using an emotion recognition engine (e.g., EmotionRecognizer) to assess the user's emotional state. The input is the voice data, and the output is the assessment of the emotional state. Specifically, the tone, pitch, and speed of the voice are analyzed to identify specific emotions, such as anxiety.
[0872] Step 6:
[0873] The server comprehensively judges the fraud possibility assessment result and emotional state, and generates a warning message if it determines that fraud is highly likely. The input is the assessment result and emotional state, and the output is the warning message. Specifically, the warning message is created using a text template.
[0874] Step 7:
[0875] The server sends the generated warning message to the family or public institution. The input is the warning message, and the output is the message to be sent to the message recipient. Specifically, the server sends the warning message to the specified email address using an email protocol (e.g., SMTP).
[0876] Step 8:
[0877] The server checks whether the alert message was sent successfully and logs the sending result. The input is the sending status and the output is writing to the log file. Specifically, it checks the status of whether the sending was successful and records the result in the log file.
[0878] The specific processing unit 290 transmits the result of the specific processing to the smart glasses 214. In the smart glasses 214, the control unit 46A causes the speaker 240 to output the result of the specific processing. The microphone 238 acquires audio indicating a user input regarding the result of the specific processing. The control unit 46A transmits audio data indicating the user input acquired by the microphone 238 to the data processing device 12. In the data processing device 12, the specific processing unit 290 acquires the audio data.
[0879] The data generation model 58 is a so-called generative AI (Artificial Intelligence). An example of the data generation model 58 is ChatGPT (Internet Search<URL: https: / / openai.com / blog / chatgpt> ), Gemini (Internet search <url: https: gemini.google.com ?hl="ja">) and other generation AIs. The data generation model 58 is obtained by performing deep learning on a neural network. A prompt including an instruction is input to the data generation model 58, and inference data such as voice data indicating voice, text data indicating text, and image data indicating an image is also input. The data generation model 58 performs inference on the input inference data in accordance with the instruction indicated by the prompt, and outputs the inference result in a data format such as voice data and text data. Here, inference refers to, for example, analysis, classification, prediction, and / or summarization.
[0880] In the above embodiment, an example in which the specific processing is performed by the data processing device 12 has been given, but the technology of the present disclosure is not limited to this, and the specific processing may be performed by the smart glasses 214.
[0881] [Third embodiment]
[0882] FIG. 5 shows an example of the configuration of a data processing system 310 according to the third embodiment.
[0883] 5, the data processing system 310 includes the data processing device 12 and a headset type terminal 314. An example of the data processing device 12 is a server.
[0884] The data processing device 12 includes a computer 22, a database 24, and a communication I / F 26. The computer 22 is an example of a "computer" according to the technology of the present disclosure. The computer 22 includes a processor 28, a RAM 30, and a storage 32. The processor 28, the RAM 30, and the storage 32 are connected to a bus 34. The database 24 and the communication I / F 26 are also connected to the bus 34. The communication I / F 26 is connected to a network 54. Examples of the network 54 include a WAN (Wide Area Network) and / or a LAN (Local Area Network).
[0885] The headset type terminal 314 includes a computer 36, a microphone 238, a speaker 240, a camera 42, a communication I / F 44, and a display 343. The computer 36 includes a processor 46, a RAM 48, and a storage 50. The processor 46, the RAM 48, and the storage 50 are connected to a bus 52. The microphone 238, the speaker 240, the camera 42, and the display 343 are also connected to the bus 52.
[0886] The microphone 238 receives instructions and the like from the user 20 by receiving voice uttered by the user 20. The microphone 238 captures the voice uttered by the user 20, converts the captured voice into audio data, and outputs it to the processor 46. The speaker 240 outputs audio in accordance with instructions from the processor 46.
[0887] Camera 42 is a small digital camera equipped with an optical system including a lens, aperture, and shutter, and an imaging element such as a CMOS (Complementary Metal-Oxide-Semiconductor) image sensor or a CCD (Charge Coupled Device) image sensor, and captures images of the surroundings of user 20 (for example, an imaging range defined by an angle of view equivalent to the field of vision of a typical healthy person).
[0888] The communication I / F 44 is connected to a network 54. The communication I / Fs 44 and 26 are responsible for the exchange of various information between the processor 46 and the processor 28 via the network 54. The exchange of various information between the processor 46 and the processor 28 using the communication I / Fs 44 and 26 is carried out in a secure state.
[0889] Fig. 6 shows an example of the main functions of the data processing device 12 and the headset type terminal 314. As shown in Fig. 6, in the data processing device 12, a specific process is performed by the processor 28. A specific process program 56 is stored in the storage 32.
[0890] The specific processing program 56 is an example of a "program" according to the technology of the present disclosure. The processor 28 reads the specific processing program 56 from the storage 32 and executes the read specific processing program 56 on the RAM 30. The specific processing is realized by the processor 28 operating as a specific processing unit 290 in accordance with the specific processing program 56 executed on the RAM 30.
[0891] The storage 32 stores a data generation model 58 and an emotion identification model 59. The data generation model 58 and the emotion identification model 59 are used by the identification processing unit 290.
[0892] In the headset type terminal 314, a reception output process is performed by the processor 46. A reception output program 60 is stored in the storage 50. The processor 46 reads the reception output program 60 from the storage 50 and executes the read reception output program 60 on the RAM 48. The reception output process is realized by the processor 46 operating as the control unit 46A in accordance with the reception output program 60 executed on the RAM 48.
[0893] Next, a description will be given of the identification process performed by the identification processing unit 290 of the data processing device 12. In the following description, the data processing device 12 will be referred to as the "server" and the headset type terminal 314 will be referred to as the "terminal."
[0894] The present invention relates to a system for preventing special frauds, particularly "it's me" frauds, that target elderly people. Hereinafter, an embodiment of the present invention will be described in detail.
[0895] The present invention is a system that includes a voice conversion means that constantly listens to voice and converts it into text data, an analysis means that analyzes the converted text data and determines whether there is a possibility of fraud, and a warning means that sends an alert if it is determined that there is a possibility of fraud.
[0896] Audio conversion means
[0897] Subject: Terminal
[0898] The device constantly collects surrounding sounds through a microphone. This audio data is processed in real time, filtered for noise, and then converted into text data using a speech recognition algorithm. This text data is then sent to the analysis means.
[0899] As a concrete example, if a user is saying on the phone, "My son was in an accident and I need money," the device will collect this speech and convert it into text that says, "My son was in an accident and I need money."
[0900] Analysis means
[0901] Subject: Server
[0902] The server receives the text data sent from the device. The server uses a generative artificial intelligence model to analyze this text data. Specifically, the server compares it with a large amount of past fraud data to determine whether the content of the text data matches a fraud pattern. If it determines that there is a high possibility of fraud, the information is sent to a warning means.
[0903] In the above specific example, if the text data "My son was in an accident and I need money" matches the fraud pattern, the server determines that there is a high possibility of fraud.
[0904] warning means
[0905] Subject: Server
[0906] The server generates an alert message based on information that is judged to be highly likely to be fraudulent. This alert message includes the conversation content that is suspected to be fraudulent. The warning means sends this alert message to registered family members and the police. It also checks whether the alert message was sent successfully and records the sending result.
[0907] For example, family members and police may receive messages like this: "Possible scam. Dialogue: 'My son was in an accident and I need money.'"
[0908] Specific examples
[0909] If a user is having a conversation like this:
[0910] User: "Hello?"
[0911] Scammer: "Mom, help me! I've been in an accident and need money."
[0912] User: "Really? How did that happen?"
[0913] Scammer: "If you don't transfer the money now, I can't have the surgery!"
[0914] Audio conversion means
[0915] The device collects this conversation, filters out noise, and converts it into text data. The converted text reads: "Mom, help me! I've been in an accident and I need money." "Really? How did that happen?" "If you don't transfer the money now, I can't have the surgery!"
[0916] Analysis means
[0917] The server receives this text data and analyzes it using a generative artificial intelligence model, using ChatGPT to detect keywords such as "accident," "money," and "transfer," and determines whether the text matches a fraud pattern.
[0918] warning means
[0919] The server determines that the scam is likely and creates an alert message containing the potentially fraudulent content: "Mom, help me! I've been in an accident and need money." This message is sent to the family and the police.
[0920] This system can quickly alert family members and the police before elderly people become victims of fraud, effectively preventing fraud.
[0921] The processing flow will be explained below.
[0922] Step 1: Collecting audio
[0923] Subject: Terminal
[0924] The device constantly collects surrounding sounds through a microphone.
[0925] The collected audio data is stored in a buffer in real time.
[0926] Step 2: Filtering the noise
[0927] Subject: Terminal
[0928] The terminal performs noise reduction on the audio data in the buffer.
[0929] Filters out background sounds and noise to extract clear audio data.
[0930] Step 3: Speech to text
[0931] Subject: Terminal
[0932] The terminal uses a speech recognition algorithm to convert the filtered voice data into text data.
[0933] The converted text data is sent to the server.
[0934] Step 4: Receiving text data
[0935] Subject: Server
[0936] The server receives the text data sent from the terminal.
[0937] The received text data is stored in memory for analysis.
[0938] Step 5: Analyze potential fraud
[0939] Subject: Server
[0940] The server analyzes the text data using a generative artificial intelligence model (e.g., ChatGPT).
[0941] During the analysis process, the text data is compared with many past fraud patterns.
[0942] Assess whether there is potential for fraud.
[0943] Step 6: Determine the likelihood of fraud
[0944] Subject: Server
[0945] Based on the analysis results, the server determines that there is a possibility of fraud.
[0946] If there is a possibility of fraud, the information is sent to the alert module.
[0947] Step 7: Prepare alert information
[0948] Subject: Server
[0949] The server generates an alert message based on information that is determined to be potentially fraudulent.
[0950] The generated alert message includes the content of the suspected fraudulent conversation.
[0951] Step 8: Sending an alert
[0952] Subject: Server
[0953] The server will then send an alert message to registered family members and police contacts.
[0954] Check the transmission result, and if the transmission is successful, record the result in the log.
[0955] By using the above-mentioned processing steps, the present system can effectively warn elderly people before they fall victim to special fraud, thereby preventing them from becoming victims of fraud.
[0956] Example 1
[0957] Next, a description will be given of Example 1. In the following description, the data processing device 12 will be referred to as a "server" and the headset type terminal 314 will be referred to as a "terminal."
[0958] In the current situation where the number of victims of special frauds targeting the elderly, especially "it's me" frauds, is increasing, there are many cases where victims transfer money without realizing it is a fraud, so a system to prevent this from happening is needed. The purpose of this invention is to protect the elderly from fraud by detecting such frauds in real time and promptly notifying family members and the police of suspected fraud.
[0959] The specific processing by the specific processing unit 290 of the data processing device 12 in the first embodiment is realized by the following means.
[0960] In this invention, the server includes a voice conversion means, an analysis means, and a warning means, which allows the server to convert voice into text data in real time, analyze the content of the text data, determine whether it is a fraud, and promptly notify family members and the police before the elderly person falls victim to fraud.
[0961] The "voice conversion means" is a means for collecting voice in real time, removing noise, and then converting the voice into text data.
[0962] The "analysis means" is a means for analyzing collected text data using a generative artificial intelligence model to determine the possibility of fraud.
[0963] The "warning means" is a means for generating an alert message and sending it to family members or the police if the analysis means determines that there is a possibility of fraud.
[0964] A "generative artificial intelligence model" is an artificial intelligence algorithm that analyzes text data based on a large amount of past fraud data and evaluates whether it matches a fraud pattern.
[0965] "Noise filtering" is a process for removing unwanted background noise from collected audio data.
[0966] A "prompt sentence" is an instruction sentence for a generative artificial intelligence model, and is an input sentence that allows the analysis means to understand the meaning of text data and perform appropriate analysis.
[0967] An "alert message" is a message containing a warning that is generated in the event of a possible fraud and is sent to family members or the police.
[0968] "Real-time" means that the processing occurs immediately, without delay.
[0969] The present invention is a system for preventing special frauds, particularly "it's me" frauds, that target elderly people. Hereinafter, an embodiment of the present invention will be described in detail.
[0970] System Overview
[0971] The system includes a voice conversion means that constantly listens to voice and converts it into text data, an analysis means that analyzes the converted text data and determines whether there is a possibility of fraud, and a warning means that sends an alert if it is determined that there is a possibility of fraud.
[0972] Audio conversion means
[0973] Hardware and Software
[0974] The device collects audio using a microphone. The recommended microphone is the Rode NT-USB, which has high-sensitivity noise-canceling technology. The device uses the open-source software Speex to filter noise from the collected audio data. The data is then converted to text using Google's Speech-to-Text API.
[0975] Example of operation
[0976] If a user receives a call from a scammer,
[0977] User: "Hello?"
[0978] Scammer: "Mom, help me! I've been in an accident and need money."
[0979] User: "Really? How did that happen?"
[0980] Scammer: "If you don't transfer the money now, I can't have the surgery!"
[0981] The device collects this conversation and
[0982] After removing the noise, the text is converted to "Mom, help me! I've been in an accident and need money," "Really? How did that happen?" and "If you don't transfer the money right now, I can't have the surgery!"
[0983] Analysis means
[0984] Hardware and Software
[0985] The server receives the text data sent from the device. The server analyzes it using a generative artificial intelligence model. OpenAI's "ChatGPT" is used as this AI model. The analysis method inputs the text data as a prompt sentence and compares it with a large amount of past fraud data to evaluate the possibility of fraud.
[0986] Prompt Sentence Examples
[0987] Prompt: "Analyze whether this text is a potential scam: 'My son was in an accident and I need money.'"
[0988] Example of operation
[0989] If the server determines that the text matches a fraud pattern based on the prompt above, it will assess the likelihood of fraud.
[0990] warning means
[0991] Hardware and Software
[0992] If the server determines that fraud is likely based on the analysis results, it generates an alert message containing the suspected fraudulent conversation content and sends the alert message to family members or the police using a secure communication protocol (HTTPS).
[0993] Example of operation
[0994] For example, the server could generate a message like this and send it to your family or the police:
[0995] "This could be a scam. The conversation goes like this: 'My son was in an accident and I need money.'"
[0996] Recording transmission results
[0997] The server checks whether the alert message was sent successfully and records the result.
[0998] summary
[0999] This system can collect voice recordings in real time before seniors fall victim to fraud, remove noise, convert the audio into text, and use a generative AI model to identify potential fraud. If a fraud is detected, a warning message can be sent quickly to family members or the police, effectively protecting seniors from fraud.
[1000] The flow of the identification process in the first embodiment will be described with reference to FIG.
[1001] Step 1: Collecting audio
[1002] Subject: Terminal
[1003] The device constantly collects surrounding sounds using a highly sensitive microphone (e.g., a microphone with noise-canceling technology). The microphone is activated and begins collecting audio data immediately after the user answers a call. The input is the audio of the conversation between the user and the fraudster, and the output is raw audio data.
[1004] Specific behavior:
[1005] From the moment the user says "Hello?", the device collects the voice and stores it as audio data.
[1006] Step 2: Noise filtering
[1007] Subject: Terminal
[1008] The device performs noise filtering on the collected audio data. It uses open-source software to remove noise. Raw audio data is used as input, and clear audio data after noise removal is generated as output.
[1009] Specific behavior:
[1010] It filters out background noise and reveals only the conversation between you and the fraudster.
[1011] Step 3: Voice Recognition
[1012] Subject: Terminal
[1013] The device then runs the noise-filtered audio data through a speech recognition algorithm to convert it into text data. This process uses the Speech-to-Text API, which takes the filtered audio data as input and produces the converted text data as output.
[1014] Specific behavior:
[1015] The conversation between the user and the scammer, such as "If you don't transfer the money now, I can't have the surgery!", is recorded as text.
[1016] Step 4: Sending text data
[1017] Subject: Terminal
[1018] The terminal sends the converted text data to the server. HTTPS is used as the communication protocol. The converted text data is sent as input and sent to the server.
[1019] Specific behavior:
[1020] The converted text, such as "Mom, help me! I've been in an accident and need money," is securely sent to a server.
[1021] Step 5: Analyzing the text data
[1022] Subject: Server
[1023] The server receives the text data sent from the device and inputs it into the generative AI model for analysis. The prompt sentence "Analyze whether this text is likely to be fraudulent" is used as input. Based on the input, the generative AI model performs analysis and obtains an output that evaluates the likelihood of fraud.
[1024] Specific behavior:
[1025] Enter the prompt text: "Analyze whether this text is likely to be fraudulent: 'My son was in an accident and I need money'" and match it with fraud patterns.
[1026] Step 6: Fraud detection
[1027] Subject: Server
[1028] The server receives the results of the generative AI model, and if there is a high possibility of fraud, it tags the information and forwards it to the warning means. The analysis results obtained from the generative AI model are used as input, and data with a fraud judgment is generated as output.
[1029] Specific behavior:
[1030] Generate result data such as "This text is likely to be fraudulent" and send it to the warning means.
[1031] Step 7: Generate an alert message
[1032] Subject: Server
[1033] The server generates an alert message based on the fraud detection data. This message contains the conversation content that is suspected to be fraudulent. The fraud detection data is used as input, and the alert message is generated as output.
[1034] Specific behavior:
[1035] Generates messages such as "Possible scam. Conversation: 'My son was in an accident and I need money.'"
[1036] Step 8: Sending an alert message
[1037] Subject: Server
[1038] The server sends the generated alert message to registered family members and the police. The communication method is SMS or email. The generated alert message is used as input and is sent to family members and the police as output.
[1039] Specific behavior:
[1040] The generated alert message is sent to family members or the police by referring to their contact list. For example, a message saying "Your son has been in an accident and you need money" is sent to a mother's smartphone.
[1041] Step 9: Record the transmission results
[1042] Subject: Server
[1043] The server checks whether the alert message was sent successfully and records the result in the database. It uses the sending log as input and generates the sending result record data as output.
[1044] Specific behavior:
[1045] Check whether the transmission was successful or failed and store the details in a recording database. For example, if the transmission was successful, it will be logged as "successful", and if the transmission failed, it will be logged as "failed".
[1046] (Application example 1)
[1047] Next, a description will be given of Application Example 1. In the following description, the data processing device 12 will be referred to as a "server" and the headset type terminal 314 will be referred to as a "terminal."
[1048] In recent years, there has been an increase in special frauds targeting the elderly, particularly "it's me" frauds. This has caused significant economic and psychological damage to many elderly people and their families. Current security measures are often reactive, providing means to prevent damage after the fraud has already progressed. Against this background, there is a need for the development of an effective system that can detect fraudulent activities in real time and respond immediately.
[1049] The specific processing by the specific processing unit 290 of the data processing device 12 in the application example 1 is realized by the following means.
[1050] In this invention, the server includes a conversion means for constantly listening to voice and converting it into text data, an analysis means for analyzing the text data and determining the possibility of fraud, a warning means for sending an alert when it is determined that there is a possibility of fraud, and a communication means for sending the fraud analysis results to others based on the converted text data. This makes it possible to detect signs of fraud in real time and quickly issue a warning before any fraud damage occurs.
[1051] "Speech" refers to acoustic signals such as human speech and words.
[1052] "Conversion means" refers to a device or technology that converts speech into text data.
[1053] "Analysis means" refers to a device or technology that analyzes the converted text data and determines the possibility of fraud.
[1054] An "alert method" is a device or technology that sends an alert when it determines that fraud is likely.
[1055] "Communication means" refers to a device or technology for transmitting analysis results to other devices or users.
[1056] A "generative artificial intelligence model" is an artificial intelligence algorithm that has the ability to generate new data based on past data and recognize specific patterns.
[1057] A "fraud pattern" is a combination of language or phrases characteristic of fraudulent activity.
[1058] An "alert message" is a message containing a warning sent by a warning means.
[1059] "Family" refers to the user's blood relatives or close acquaintances.
[1060] A "police agency" is a public institution with legal authority to investigate crimes and protect the safety of citizens.
[1061] "Text data" is data in which voice is expressed as characters.
[1062] The present invention relates to a system for preventing special frauds targeting elderly people, and aims to prevent fraud in real time by constantly collecting voices, determining the possibility of fraud, and sending warnings to family members or police agencies if necessary.
[1063] Audio conversion means
[1064] The device constantly collects surrounding sounds using a microphone. The collected audio data is processed in real time, noise-filtered, and then converted into text data using a speech recognition algorithm. This text data is then sent to an analysis tool. Specifically, the device uses the Python speech_recognition library and Google's speech recognition service.
[1065] Analysis means
[1066] The server receives the text data sent from the device. The server then uses a generative AI model to analyze this text data. Specifically, it uses OpenAI's API to determine whether the text data matches a fraud pattern. This analysis involves evaluating the text based on past fraud data. The analysis method requests the generative AI model to analyze using a prompt sentence such as the following:
[1067] Prompt Sentence Examples
[1068] "Please determine if the following text is a scam: Mom, help me! I've been in an accident and need money."
[1069] warning means
[1070] If the server determines that there is a high possibility of fraud based on the analysis results, it generates an alert message. This message contains the suspected fraudulent conversation content. The warning mechanism then sends this alert message to registered family members and police agencies. The Python smtplib library is used to send the email. As an example of an alert, the following message is sent to family members:
[1071] "Possible scam. Dialogue: 'Mom, help me! I've been in an accident and need money.'"
[1072] Specific examples
[1073] For example, consider the following situation:
[1074] User: "Hello?"
[1075] Scammer: "Mom, help me! I've been in an accident and need money."
[1076] User: "Really? How did that happen?"
[1077] Scammer: "If you don't transfer the money now, I can't have the surgery!"
[1078] The device collects this conversation, performs speech recognition, and converts it into text data such as, "Mom, help me! I was in an accident and need money," "Really? How did that happen?", and "If you don't transfer the money now, I can't have the surgery!" This text data is then analyzed using a generative artificial intelligence model (such as OpenAI's API), and if it is determined to be a fraud, a warning message is sent to the family and police.
[1079] The system will enable seniors to take prompt action before they become victims of fraud, strengthening security measures.
[1080] The flow of the specific processing in the application example 1 will be described with reference to FIG.
[1081] Step 1:
[1082] The device uses a microphone to collect surrounding audio. The input is surrounding audio data, which the device filters in real time and converts into text data using a speech recognition algorithm. The output is the converted text data.
[1083] Step 2:
[1084] The terminal transmits the text data obtained in step 1 to the server. The input is the converted text data, and the output is the text data transmitted to the server. This data transmission utilizes an internet connection.
[1085] Step 3:
[1086] The server sends a prompt to the generative AI model to analyze the received text data. The input is the received text data, and the output is the analysis result obtained from the AI model. Specifically, the server uses the OpenAI API to generate the prompt text as follows:
[1087] "Please determine if the following text is a scam: [text data]"
[1088] Step 4:
[1089] The generative AI model analyzes the prompts sent from the server and determines whether the text data matches a fraud pattern. The input is a prompt containing the text data to be analyzed, and the output is the analysis result. Specifically, it compares the result with past fraud data.
[1090] Step 5:
[1091] The server receives the analysis results from the generative AI model and evaluates whether there is a high probability of fraud. The input is the analysis results from the AI model, and the output is the evaluation result indicating the possibility of fraud. Specifically, it checks whether the evaluation result meets certain criteria.
[1092] Step 6:
[1093] If it is determined that there is a high possibility of fraud, the server generates an alert message. The input is the evaluation result, and the output is the alert message. The specific operation is to generate a message containing the conversation content that is suspected to be fraudulent.
[1094] Step 7:
[1095] The server sends the generated alert message to the family or police. The input is the alert message, and the output is the message sent to the family or police. Specifically, it uses the Python smtplib library to send the email.
[1096] The above are the specific processing steps of the system embodying the present invention.
[1097] Furthermore, an emotion engine that estimates the user's emotion may be further combined. That is, the identification processing unit 290 may estimate the user's emotion using the emotion identification model 59, and perform identification processing using the user's emotion.
[1098] This invention is a system that prevents victims of special frauds targeting the elderly, and in particular combines voice recognition technology, a generative AI model, and an emotion engine that recognizes the user's emotions.
[1099] System configuration
[1100] The system includes the following main means:
[1101] 1. Audio conversion method
[1102] 2. Analysis method
[1103] 3. Warning measures
[1104] 4. Emotion recognition means
[1105] Audio conversion means
[1106] Subject: Terminal
[1107] The device constantly collects surrounding sounds through a microphone. This audio data is processed in real time, noise-filtered, and then converted into text data using a speech recognition algorithm. This text data is then sent to the analysis means.
[1108] As a concrete example, if a user says on the phone, "My son was in an accident and I need money," the device will collect this speech and convert it into text: "My son was in an accident and I need money."
[1109] Analysis means
[1110] Subject: Server
[1111] The server receives text data sent from the device. The server analyzes this text data using a generative artificial intelligence model (e.g., ChatGPT). It compares this text with past fraud patterns and determines whether the text is likely to be fraudulent. The analysis results are sent to the warning mechanism.
[1112] For example, if the text "My son had an accident and I need money" matches a fraud pattern, the server determines that it is likely fraudulent.
[1113] emotion recognition means
[1114] Subject: Server
[1115] The server also has an emotion recognition engine that analyzes the user's emotional state from the voice data. This emotion recognition engine analyzes the user's emotions based on the tone, pitch, speed, etc. of the voice and detects emotional patterns that may indicate fraud. In cooperation with the analysis means, it performs a comprehensive evaluation that includes the user's emotional state.
[1116] For example, if emotions of anxiety or confusion are detected from the user's voice, this can be added to the evaluation to more accurately determine the possibility of fraud.
[1117] warning means
[1118] Subject: Server
[1119] The server generates an alert message based on the results of the analysis and emotion recognition. If it determines that there is a high possibility of fraud, it sends the alert message to family members or the police. This alert message includes the content of the conversation suspected of fraud and the evaluation result of the emotional state. The warning means checks whether the message was sent successfully and records the sending result in a log.
[1120] For example, a message sent to family members or police might read: "Possible scam. Dialogue: 'My son was in an accident and I need money.' User's emotional state: Anxiety."
[1121] Specific examples
[1122] If a user is having a conversation like this:
[1123] User: "Hello?"
[1124] Scammer: "Mom, help me! I've been in an accident and need money."
[1125] User: "Really? How did that happen?"
[1126] Scammer: "If you don't transfer the money now, I can't have the surgery!"
[1127] Audio conversion means
[1128] The device collects this conversation, filters out noise, and then converts it into text data.
[1129] Translated text: "Mom, help me! I was in an accident and need money." "Really? How did that happen?" "If you don't send the money now, I can't have my surgery!"
[1130] Analysis means
[1131] The server receives this text data and analyzes it using a generative artificial intelligence model.
[1132] ChatGPT is used to detect keywords such as "accident," "money," and "transfer," and determine whether this text matches a fraud pattern.
[1133] emotion recognition means
[1134] The server uses the voice data to analyze the user's emotional state.
[1135] The emotion recognition engine detects anxiety in the user's voice.
[1136] warning means
[1137] The server determines that there is a high probability of fraud and creates an alert message.
[1138] The message contains the potentially fraudulent content: "Mom, help me! I've been in an accident and need money" and the user's emotional state: "anxious."
[1139] This message will be sent to the family and police.
[1140] This system can effectively detect fraud before it happens to elderly people and quickly alert family members and the police, thereby preventing fraud from occurring.
[1141] The processing flow will be explained below.
[1142] Step 1: Collecting audio
[1143] Subject: Terminal
[1144] The device constantly collects surrounding sounds through a microphone.
[1145] The collected audio data is stored in a buffer in real time.
[1146] Step 2: Filtering the noise
[1147] Subject: Terminal
[1148] The terminal performs noise reduction on the audio data in the buffer.
[1149] Filters out background sounds and noise to extract clear audio data.
[1150] Step 3: Speech to text
[1151] Subject: Terminal
[1152] The terminal uses a speech recognition algorithm to convert the filtered voice data into text data.
[1153] The converted text data is sent to the server.
[1154] Step 4: Receiving text data
[1155] Subject: Server
[1156] The server receives the text data sent from the terminal.
[1157] The received text data is stored in memory for analysis.
[1158] Step 5: Analyze potential fraud
[1159] Subject: Server
[1160] The server analyzes the text data using a generative artificial intelligence model (e.g., ChatGPT).
[1161] During the analysis process, the text data is compared with many past fraud patterns.
[1162] Assess whether there is potential for fraud.
[1163] Step 6: Determine the likelihood of fraud
[1164] Subject: Server
[1165] Based on the analysis results, the server determines that there is a possibility of fraud.
[1166] If there is a possibility of fraud, the information is sent to the alert module.
[1167] Step 7: Collect emotion data
[1168] Subject: Terminal
[1169] The device uses voice data to collect data to analyze the user's emotional state using an emotion engine.
[1170] The collected emotion data is sent to a server.
[1171] Step 8: Analyze the sentiment data
[1172] Subject: Server
[1173] The server uses an emotion engine to analyze the emotion data sent from the terminal.
[1174] Evaluates the user's emotional state based on the tone, pitch, speed, etc. of the voice.
[1175] Detect emotional patterns that indicate potential fraud.
[1176] Step 9: Integrating and analyzing emotion data
[1177] Subject: Server
[1178] The server integrates the results of the text data analysis and the emotional data analysis, and comprehensively reassess the likelihood of fraud.
[1179] Determine with greater accuracy whether there is a high likelihood of fraud.
[1180] Step 10: Prepare alert information
[1181] Subject: Server
[1182] The server generates an alert message based on information that is determined to be potentially fraudulent.
[1183] The generated alert message includes the suspected fraudulent conversation content and an assessment of the emotional state.
[1184] Step 11: Sending an alert
[1185] Subject: Server
[1186] The server will then send an alert message to registered family members and police contacts.
[1187] Check the transmission result, and if the transmission is successful, record the result in the log.
[1188] Example 2
[1189] Next, a description will be given of Example 2. In the following description, the data processing device 12 will be referred to as a "server" and the headset type terminal 314 will be referred to as a "terminal."
[1190] The number of victims of special frauds targeting the elderly is increasing year by year. Fraudsters use sophisticated methods to deceive the elderly and obtain money illegally. In order to prevent such fraud, a system is needed that can monitor conversations around the elderly in real time and detect suspicious speech. It is also desirable to evaluate the user's emotional state to detect factors that increase the likelihood of fraud.
[1191] The identification process by the identification processing unit 290 of the data processing device 12 in the second embodiment is realized by the following means. In this invention, the server includes a sound collection means that constantly collects environmental sounds and performs noise filtering, a sound recognition means that converts the collected sound data into text data, an analysis means that analyzes the text data and determines the possibility of fraud, an emotion recognition means that analyzes the emotional state of the user, and a warning means that sends an alert when it is determined that there is a possibility of fraud. This makes it possible to detect the risk of elderly people falling victim to fraud early and promptly notify their family or the police.
[1192] "Environmental sound" refers to all sounds occurring around the user, and is a general term for sound data including conversations and background sounds.
[1193] "Noise filtering" is a process that removes unnecessary background sounds and noise from collected audio data and converts it into audio data that is easy to hear.
[1194] "Sound collection means" refers to a device or software that constantly collects environmental sounds using a microphone.
[1195] "Speech recognition means" refers to an algorithm or system for converting collected voice data into text data.
[1196] "Text data" refers to character string data converted from speech by speech recognition means.
[1197] A "generative artificial intelligence model" refers to an algorithm that learns from large amounts of text data and performs natural language processing, and specifically includes models such as ChatGPT.
[1198] A "prompt sentence" refers to input text used to give instructions or ask questions to a generative artificial intelligence model and obtain a specific response.
[1199] "Analysis Method" refers to an algorithm or system that uses a generative artificial intelligence model to analyze text data and determine the likelihood of fraud.
[1200] "Emotion recognition means" refers to an algorithm or system for analyzing a user's emotional state from voice data and detecting emotional patterns.
[1201] "Warning measures" refer to algorithms or systems that send alert messages to family members or police in the event of potential fraud.
[1202] "Alert Message" means a message containing warning content that is generated to notify you of suspected fraud.
[1203] "Transmission result" refers to confirmation information as to whether the alert message was sent successfully.
[1204] "Log" refers to data that records system operations, events, and the results of sending alert messages.
[1205] The present invention is a system for preventing victims of special frauds targeting elderly people. The system includes a voice collection means, a voice recognition means, an analysis means, an emotion recognition means, and a warning means. The details of each means and the specific implementation method are explained below.
[1206] Audio collection method
[1207] The device is equipped with a microphone for constantly collecting ambient sounds. The microphone is highly sensitive and capable of collecting sound over a wide range, such as a condenser microphone or digital microphone. The collected sound data is processed in real time using a noise filtering algorithm, which removes background noise and unnecessary sounds to obtain clear sound data.
[1208] Voice recognition means
[1209] The device inputs the noise-filtered voice data into a voice recognition engine. For example, Google Cloud Speech-to-Text API is used as the voice recognition engine. The voice recognition engine converts the voice data into text data and expresses the content as a string of characters. The converted text data is sent to an analysis means.
[1210] Analysis means
[1211] The server receives the text data sent from the device. The server analyzes the text data using a generative artificial intelligence model (e.g., ChatGPT). It generates a prompt and inputs it into the model to compare it with past fraud patterns and evaluate the likelihood of fraud. For example, the prompt could be set as "Please evaluate the likelihood that this text is a special fraud." When the analysis results are obtained, it determines whether the likelihood of fraud is high and passes the result to the next processing step.
[1212] emotion recognition means
[1213] The server has an emotion recognition engine, such as IBM Watson Tone Analyzer, that analyzes the user's emotional state from the voice data. The emotion recognition engine analyzes the tone, pitch, and speed of the voice to assess whether the user is anxious or confused. These emotional states are also considered as factors that increase the likelihood of fraud.
[1214] warning means
[1215] The server generates an alert message based on the results of the analysis and emotion recognition methods. If it determines that there is a high possibility of fraud, it sends this alert message to the user's family or the police. The alert message includes the content of the conversation that is suspected to be fraudulent and the user's emotional state. For example, a message such as "Possible fraud. Conversation content: 'Mom, help me! I've been in an accident and need money.' User's emotional state: anxious" is generated. The server then checks the transmission result and records in a log that it was sent successfully.
[1216] Specific examples
[1217] Here is a concrete example of the system: If a user is having the following conversation:
[1218] User: "Hello?"
[1219] Scammer: "Mom, help me! I've been in an accident and need money."
[1220] User: "Really? How did that happen?"
[1221] Scammer: "If you don't transfer the money now, I can't have the surgery!"
[1222] 1. Audio collection means: The device collects the conversation and removes noise.
[1223] 2. Voice recognition method: After noise filtering, the conversation is converted into text such as "Mom, help me! I've been in an accident and need money," "Really? How did that happen?" and "If you don't transfer the money now I can't have the surgery!" and sent to the server.
[1224] 3. Analysis method: The server analyzes the text, detects keywords such as "accident," "money," and "transfer," and evaluates the likelihood of fraud. An example of a prompt is, "Please evaluate the likelihood that this text is a special fraud."
[1225] 4. Emotion recognition means: The server analyzes the user's emotion of anxiety from the voice.
[1226] 5. Warning method: The server determines that there is a high possibility of fraud and generates an alert message stating, "Possible fraud. Conversation: 'Mom, help! I've been in an accident and need money.' User's emotional state: Anxiety," and sends it to the family and police.
[1227] This system makes it possible to detect early the risk of elderly people falling victim to fraud and take prompt action.
[1228] The flow of the identification process in the second embodiment will be described with reference to FIG.
[1229] System program processing flow
[1230] Step 1: Audio Collection
[1231] The device collects ambient sounds, acquires audio data in real time through a microphone, and processes the data with a noise filtering algorithm. The input is the audio data of the ambient sounds, and the output is the audio data with the noise removed.
[1232] Specifically, the system captures audio data via a microphone and removes unwanted noise through a noise reduction filter.
[1233] Step 2: Voice Recognition
[1234] The device inputs the noise-filtered voice data into a voice recognition engine and converts it into text data. Here, for example, the Google Cloud Speech-to-Text API is used. The input is the voice data with noise removed, and the output is the converted text data.
[1235] Specifically, the speech data is sent to a speech recognition engine, and text data is obtained as a result.
[1236] Step 3: Text analysis
[1237] The server receives the text data sent from the device. The server analyzes this text data using a generative artificial intelligence model (such as ChatGPT). The input is the text data, and the output is the evaluation result of the likelihood of fraud. The prompt used is "Please evaluate the likelihood that this text is a special fraud."
[1238] Specifically, text data is input into a generative artificial intelligence model and the results of matching with fraud patterns are evaluated.
[1239] Step 4: Emotional state analysis
[1240] The server simultaneously analyzes the user's emotional state from the voice data. Using an emotion recognition engine (such as IBM Watson Tone Analyzer), it analyzes the tone, pitch, and speed of the voice to determine whether the user is anxious or confused. The input is the voice data, and the output is the evaluation result of the user's emotional state.
[1241] Specifically, the voice data is input into an emotion recognition engine to obtain an evaluation result of the emotional state.
[1242] Step 5: Generate and send an alert
[1243] The server generates an alert message based on the results of the analysis and emotion recognition methods. If it determines that there is a high possibility of fraud, it sends the alert message to the user's family or the police. The input is the evaluation result of the possibility of fraud and the evaluation result of the user's emotional state, and the output is the alert message.
[1244] Specifically, the system generates an alert message based on the analysis results and the user's emotional state, and sends it to family members or the police via email or SMS. The system then records the results of the message in a log.
[1245] This series of processes makes it possible to detect early on when an elderly person is at risk of falling victim to fraud and quickly notify their family or the police.
[1246] (Application example 2)
[1247] Next, a description will be given of Application Example 2. In the following description, the data processing device 12 will be referred to as a "server" and the headset type terminal 314 will be referred to as a "terminal."
[1248] Special frauds targeting the elderly remain a social problem, with many elderly people suffering financial losses. The present invention aims to provide an effective system for preventing elderly people from falling victim to fraud. In particular, the objective of the present invention is to protect elderly people from fraud by detecting signs of fraud early and issuing prompt warnings.
[1249] The identification process by the identification processing unit 290 of the data processing device 12 in Application Example 2 is realized by the following means. In this invention, the server includes a voice conversion means that constantly collects voice and converts it into text data, an analysis means that analyzes the text data and evaluates the possibility of fraud, an emotion recognition means that analyzes the emotional state from the voice data, and a warning means that generates and transmits a warning message when it is determined that there is a possibility of fraud. This enables elderly people to quickly recognize the possibility of fraud and issue a warning to their family or public institutions.
[1250] The "voice conversion means" is a unit that constantly collects voice and converts it into text data.
[1251] The "analysis means" is a unit for analyzing the text data and assessing the possibility of fraud.
[1252] The "emotion recognition means" is a unit that analyzes the emotional state from the voice data.
[1253] The "warning means" is a unit for generating and transmitting a warning message when it is determined that there is a possibility of fraud.
[1254] A "generative artificial intelligence model" is an artificial intelligence algorithm generated from past data and learning, and is used as a means of analyzing text data.
[1255] "Text data" refers to data in character format converted from speech by speech conversion means.
[1256] A "fraud pattern" is a collection of characteristic words or phrases that indicate fraud.
[1257] The "emotional state" is the emotional state of the user that is analyzed based on the tone, pitch, speed, etc. of the voice data.
[1258] A "warning message" is a notification that is generated when potential fraud is detected.
[1259] "Family members and public authorities" are those who are notified by the warning measures of possible fraud.
[1260] The present invention is a system for preventing victims of special frauds targeting elderly people, and includes a voice conversion means, an analysis means, an emotion recognition means, and a warning means. Specific embodiments for carrying out the present invention will be described below.
[1261] System configuration
[1262] The system includes a dedicated terminal, a server, and a user terminal (e.g., a smartphone).
[1263] Audio conversion means
[1264] The device constantly collects ambient sound through a microphone. This sound data is processed in real time and noise filtered. Speech recognition software (e.g., the Python speech_recognition library) then converts the sound data into text data, which is then sent to an analysis unit.
[1265] Analysis means
[1266] The server receives text data sent from the device. Using a generative artificial intelligence model (e.g., ChatGPT), it analyzes this text data and matches it with fraud patterns. If the text indicates possible fraud, the results are sent to the emotion recognition means.
[1267] emotion recognition means
[1268] The server then uses an emotion recognition engine (e.g., EmotionRecognizer) to analyze the user's emotional state using the voice data. This engine analyzes the user's emotions based on the tone, pitch, and speed of the voice, and detects emotional patterns that may indicate fraud. It then makes a comprehensive assessment, including the user's emotional state.
[1269] warning means
[1270] If the server determines that a message is likely to be fraudulent, it generates a warning message that is sent to family members or public institutions. The warning message includes the content of the conversation that is suspected to be fraudulent and the result of the evaluation of the emotional state. The warning means checks whether the message was sent successfully and records the sending result in a log.
[1271] Specific examples
[1272] If a user is having a conversation like this:
[1273] User: "Hello?"
[1274] Scammer: "Mom, help me! I've been in an accident and need money."
[1275] User: "Really? How did that happen?"
[1276] Scammer: "If you don't transfer the money now, I can't have the surgery!"
[1277] Audio conversion means
[1278] The device collects this conversation, filters out noise, and converts it into text data, which looks like this:
[1279] "Mom, help me! I was in an accident and I need money." "Really? Why did that happen?" "If you don't transfer the money right now, I can't have the surgery!"
[1280] Analysis means
[1281] The server receives this text data and analyzes it using a generative artificial intelligence model. Because the text contains keywords such as "accident," "money," and "transfer," it is determined to be highly likely to be fraudulent.
[1282] emotion recognition means
[1283] The server uses the voice data to analyze the user's emotional state, and an emotion recognition engine detects anxiety from the user's voice.
[1284] warning means
[1285] The server determines that the fraud is likely and creates an alert message containing the suspected fraudulent content: "Mom, help me! I've been in an accident and need money" and the user's emotional state: "anxious." This message is sent to family members and public authorities.
[1286] Prompt Sentence Examples
[1287] "I received a suspicious call. I would like to provide the following information:
[1288] Conversation: 'My son had an accident and I need money.'
[1289] Emotional state: 'Anxious'
[1290] Please take immediate action as this may be a scam."
[1291] The flow of the specific processing in the application example 2 will be described with reference to FIG.
[1292] Step 1:
[1293] The device constantly collects ambient sound through a microphone. This sound data is processed in real time and noise filtered. The input is raw sound data, and the output is noise-removed sound data. Specifically, the sound signal obtained through the microphone is passed through a digital filter to reduce noise.
[1294] Step 2:
[1295] The device converts the noise-removed audio data into text data using speech recognition software (e.g., Python's speech_recognition library). The input is the noise-removed audio data, and the output is the corresponding text data. Specifically, the speech recognition algorithm analyzes the audio signal and converts it into the corresponding string of characters.
[1296] Step 3:
[1297] The terminal transmits the converted text data to the server. The input is the text data, and the output is data transmission to the server. Specifically, the text data is transferred to the server via a network connection.
[1298] Step 4:
[1299] The server analyzes the text data received from the device using a generative artificial intelligence model (e.g., ChatGPT) to assess the likelihood of fraud. The input is text data, and the output is an assessment of whether the likelihood of fraud is high. Specifically, the chat model compares the text with past fraud patterns to detect signs of fraud.
[1300] Step 5:
[1301] The server analyzes the voice data using an emotion recognition engine (e.g., EmotionRecognizer) to assess the user's emotional state. The input is the voice data, and the output is the assessment of the emotional state. Specifically, the tone, pitch, and speed of the voice are analyzed to identify specific emotions, such as anxiety.
[1302] Step 6:
[1303] The server comprehensively judges the fraud possibility assessment result and emotional state, and generates a warning message if it determines that fraud is highly likely. The input is the assessment result and emotional state, and the output is the warning message. Specifically, the warning message is created using a text template.
[1304] Step 7:
[1305] The server sends the generated warning message to the family or public institution. The input is the warning message, and the output is the message to be sent to the message recipient. Specifically, the server sends the warning message to the specified email address using an email protocol (e.g., SMTP).
[1306] Step 8:
[1307] The server checks whether the alert message was sent successfully and logs the sending result. The input is the sending status and the output is writing to the log file. Specifically, it checks the status of whether the sending was successful and records the result in the log file.
[1308] The specific processing unit 290 transmits the result of the specific processing to the headset type terminal 314. In the headset type terminal 314, the control unit 46A causes the speaker 240 and the display 343 to output the result of the specific processing. The microphone 238 acquires audio indicating a user input regarding the result of the specific processing. The control unit 46A transmits audio data indicating the user input acquired by the microphone 238 to the data processing device 12. In the data processing device 12, the specific processing unit 290 acquires the audio data.
[1309] The data generation model 58 is a so-called generative AI (Artificial Intelligence). An example of the data generation model 58 is ChatGPT (Internet Search<URL: https: / / openai.com / blog / chatgpt> ), Gemini (Internet search <url: https: gemini.google.com ?hl="ja">) and other generation AIs. The data generation model 58 is obtained by performing deep learning on a neural network. A prompt including an instruction is input to the data generation model 58, and inference data such as voice data indicating voice, text data indicating text, and image data indicating an image is also input. The data generation model 58 performs inference on the input inference data in accordance with the instruction indicated by the prompt, and outputs the inference result in a data format such as voice data and text data. Here, inference refers to, for example, analysis, classification, prediction, and / or summarization.
[1310] In the above embodiment, an example was given in which the specific processing is performed by the data processing device 12, but the technology of the present disclosure is not limited to this, and the specific processing may be performed by the headset type terminal 314.
[1311] [Fourth embodiment]
[1312] FIG. 7 shows an example of the configuration of a data processing system 410 according to the fourth embodiment.
[1313] 7, a data processing system 410 includes a data processing device 12 and a robot 414. An example of the data processing device 12 is a server.
[1314] The data processing device 12 includes a computer 22, a database 24, and a communication I / F 26. The computer 22 is an example of a "computer" according to the technology of the present disclosure. The computer 22 includes a processor 28, a RAM 30, and a storage 32. The processor 28, the RAM 30, and the storage 32 are connected to a bus 34. The database 24 and the communication I / F 26 are also connected to the bus 34. The communication I / F 26 is connected to a network 54. Examples of the network 54 include a WAN (Wide Area Network) and / or a LAN (Local Area Network).
[1315] The robot 414 includes a computer 36, a microphone 238, a speaker 240, a camera 42, a communication I / F 44, and a control target 443. The computer 36 includes a processor 46, a RAM 48, and a storage 50. The processor 46, the RAM 48, and the storage 50 are connected to a bus 52. The microphone 238, the speaker 240, the camera 42, and the control target 443 are also connected to the bus 52.
[1316] The microphone 238 receives instructions and the like from the user 20 by receiving voice uttered by the user 20. The microphone 238 captures the voice uttered by the user 20, converts the captured voice into audio data, and outputs it to the processor 46. The speaker 240 outputs audio in accordance with instructions from the processor 46.
[1317] Camera 42 is a small digital camera equipped with an optical system including a lens, aperture, and shutter, and an imaging element such as a CMOS (Complementary Metal-Oxide-Semiconductor) image sensor or a CCD (Charge Coupled Device) image sensor, and captures images of the surroundings of user 20 (for example, an imaging range defined by an angle of view equivalent to the field of vision of a typical healthy person).
[1318] The communication I / F 44 is connected to a network 54. The communication I / Fs 44 and 26 are responsible for the exchange of various information between the processor 46 and the processor 28 via the network 54. The exchange of various information between the processor 46 and the processor 28 using the communication I / Fs 44 and 26 is carried out in a secure state.
[1319] The control object 443 includes a display device, LEDs in the eyes, and motors for driving the arms, hands, and feet. The posture and gestures of the robot 414 are controlled by controlling the motors of the arms, hands, and feet. Some of the emotions of the robot 414 can be expressed by controlling these motors. In addition, the facial expressions of the robot 414 can also be expressed by controlling the light emission state of the LEDs in the eyes of the robot 414.
[1320] Fig. 8 shows an example of the main functions of the data processing device 12 and the robot 414. As shown in Fig. 8, in the data processing device 12, a specific process is performed by the processor 28. A specific process program 56 is stored in the storage 32.
[1321] The specific processing program 56 is an example of a "program" according to the technology of the present disclosure. The processor 28 reads the specific processing program 56 from the storage 32 and executes the read specific processing program 56 on the RAM 30. The specific processing is realized by the processor 28 operating as a specific processing unit 290 in accordance with the specific processing program 56 executed on the RAM 30.
[1322] The storage 32 stores a data generation model 58 and an emotion identification model 59. The data generation model 58 and the emotion identification model 59 are used by the identification processing unit 290.
[1323] In the robot 414, the processor 46 performs the reception output process. A reception output program 60 is stored in the storage 50. The processor 46 reads the reception output program 60 from the storage 50 and executes the read reception output program 60 on the RAM 48. The reception output process is realized by the processor 46 operating as the control unit 46A in accordance with the reception output program 60 executed on the RAM 48.
[1324] Next, a description will be given of the specific processing performed by the specific processing unit 290 of the data processing device 12. In the following description, the data processing device 12 will be referred to as a "server" and the robot 414 will be referred to as a "terminal."
[1325] The present invention relates to a system for preventing special frauds, particularly "it's me" frauds, that target elderly people. Hereinafter, an embodiment of the present invention will be described in detail.
[1326] The present invention is a system that includes a voice conversion means that constantly listens to voice and converts it into text data, an analysis means that analyzes the converted text data and determines whether there is a possibility of fraud, and a warning means that sends an alert if it is determined that there is a possibility of fraud.
[1327] Audio conversion means
[1328] Subject: Terminal
[1329] The device constantly collects surrounding sounds through a microphone. This audio data is processed in real time, filtered for noise, and then converted into text data using a speech recognition algorithm. This text data is then sent to the analysis means.
[1330] As a concrete example, if a user is saying on the phone, "My son was in an accident and I need money," the device will collect this speech and convert it into text that says, "My son was in an accident and I need money."
[1331] Analysis means
[1332] Subject: Server
[1333] The server receives the text data sent from the device. The server uses a generative artificial intelligence model to analyze this text data. Specifically, the server compares it with a large amount of past fraud data to determine whether the content of the text data matches a fraud pattern. If it determines that there is a high possibility of fraud, the information is sent to a warning means.
[1334] In the above specific example, if the text data "My son was in an accident and I need money" matches the fraud pattern, the server determines that there is a high possibility of fraud.
[1335] warning means
[1336] Subject: Server
[1337] The server generates an alert message based on information that is judged to be highly likely to be fraudulent. This alert message includes the conversation content that is suspected to be fraudulent. The warning means sends this alert message to registered family members and the police. It also checks whether the alert message was sent successfully and records the sending result.
[1338] For example, family members and police may receive messages like this: "Possible scam. Dialogue: 'My son was in an accident and I need money.'"
[1339] Specific examples
[1340] If a user is having a conversation like this:
[1341] User: "Hello?"
[1342] Scammer: "Mom, help me! I've been in an accident and need money."
[1343] User: "Really? How did that happen?"
[1344] Scammer: "If you don't transfer the money now, I can't have the surgery!"
[1345] Audio conversion means
[1346] The device collects this conversation, filters out noise, and converts it into text data. The converted text reads: "Mom, help me! I've been in an accident and I need money." "Really? How did that happen?" "If you don't transfer the money now, I can't have the surgery!"
[1347] Analysis means
[1348] The server receives this text data and analyzes it using a generative artificial intelligence model, using ChatGPT to detect keywords such as "accident," "money," and "transfer," and determines whether the text matches a fraud pattern.
[1349] warning means
[1350] The server determines that the scam is likely and creates an alert message containing the potentially fraudulent content: "Mom, help me! I've been in an accident and need money." This message is sent to the family and the police.
[1351] This system can quickly alert family members and the police before elderly people become victims of fraud, effectively preventing fraud.
[1352] The processing flow will be explained below.
[1353] Step 1: Collecting audio
[1354] Subject: Terminal
[1355] The device constantly collects surrounding sounds through a microphone.
[1356] The collected audio data is stored in a buffer in real time.
[1357] Step 2: Filtering the noise
[1358] Subject: Terminal
[1359] The terminal performs noise reduction on the audio data in the buffer.
[1360] Filters out background sounds and noise to extract clear audio data.
[1361] Step 3: Speech to text
[1362] Subject: Terminal
[1363] The terminal uses a speech recognition algorithm to convert the filtered voice data into text data.
[1364] The converted text data is sent to the server.
[1365] Step 4: Receiving text data
[1366] Subject: Server
[1367] The server receives the text data sent from the terminal.
[1368] The received text data is stored in memory for analysis.
[1369] Step 5: Analyze potential fraud
[1370] Subject: Server
[1371] The server analyzes the text data using a generative artificial intelligence model (e.g., ChatGPT).
[1372] During the analysis process, the text data is compared with many past fraud patterns.
[1373] Assess whether there is potential for fraud.
[1374] Step 6: Determine the likelihood of fraud
[1375] Subject: Server
[1376] Based on the analysis results, the server determines that there is a possibility of fraud.
[1377] If there is a possibility of fraud, the information is sent to the alert module.
[1378] Step 7: Prepare alert information
[1379] Subject: Server
[1380] The server generates an alert message based on information that is determined to be potentially fraudulent.
[1381] The generated alert message includes the content of the suspected fraudulent conversation.
[1382] Step 8: Sending an alert
[1383] Subject: Server
[1384] The server will then send an alert message to registered family members and police contacts.
[1385] Check the transmission result, and if the transmission is successful, record the result in the log.
[1386] By using the above-mentioned processing steps, the present system can effectively warn elderly people before they fall victim to special fraud, thereby preventing them from becoming victims of fraud.
[1387] Example 1
[1388] Next, a description will be given of Example 1. In the following description, the data processing device 12 will be referred to as a "server" and the robot 414 will be referred to as a "terminal."
[1389] In the current situation where the number of victims of special frauds targeting the elderly, especially "it's me" frauds, is increasing, there are many cases where victims transfer money without realizing it is a fraud, so a system to prevent this from happening is needed. The purpose of this invention is to protect the elderly from fraud by detecting such frauds in real time and promptly notifying family members and the police of suspected fraud.
[1390] The specific processing by the specific processing unit 290 of the data processing device 12 in the first embodiment is realized by the following means.
[1391] In this invention, the server includes a voice conversion means, an analysis means, and a warning means, which allows the server to convert voice into text data in real time, analyze the content of the text data, determine whether it is a fraud, and promptly notify family members and the police before the elderly person falls victim to fraud.
[1392] The "voice conversion means" is a means for collecting voice in real time, removing noise, and then converting the voice into text data.
[1393] The "analysis means" is a means for analyzing collected text data using a generative artificial intelligence model to determine the possibility of fraud.
[1394] The "warning means" is a means for generating an alert message and sending it to family members or the police if the analysis means determines that there is a possibility of fraud.
[1395] A "generative artificial intelligence model" is an artificial intelligence algorithm that analyzes text data based on a large amount of past fraud data and evaluates whether it matches a fraud pattern.
[1396] "Noise filtering" is a process for removing unwanted background noise from collected audio data.
[1397] A "prompt sentence" is an instruction sentence for a generative artificial intelligence model, and is an input sentence that allows the analysis means to understand the meaning of text data and perform appropriate analysis.
[1398] An "alert message" is a message containing a warning that is generated in the event of a possible fraud and is sent to family members or the police.
[1399] "Real-time" means that the processing occurs immediately, without delay.
[1400] The present invention is a system for preventing special frauds, particularly "it's me" frauds, that target elderly people. Hereinafter, an embodiment of the present invention will be described in detail.
[1401] System Overview
[1402] The system includes a voice conversion means that constantly listens to voice and converts it into text data, an analysis means that analyzes the converted text data and determines whether there is a possibility of fraud, and a warning means that sends an alert if it is determined that there is a possibility of fraud.
[1403] Audio conversion means
[1404] Hardware and Software
[1405] The device collects audio using a microphone. The recommended microphone is the Rode NT-USB, which has high-sensitivity noise-canceling technology. The device uses the open-source software Speex to filter noise from the collected audio data. The data is then converted to text using Google's Speech-to-Text API.
[1406] Example of operation
[1407] If a user receives a call from a scammer,
[1408] User: "Hello?"
[1409] Scammer: "Mom, help me! I've been in an accident and need money."
[1410] User: "Really? How did that happen?"
[1411] Scammer: "If you don't transfer the money now, I can't have the surgery!"
[1412] The device collects this conversation and
[1413] After removing the noise, the text is converted to "Mom, help me! I've been in an accident and need money," "Really? How did that happen?" and "If you don't transfer the money right now, I can't have the surgery!"
[1414] Analysis means
[1415] Hardware and Software
[1416] The server receives the text data sent from the device. The server analyzes it using a generative artificial intelligence model. OpenAI's "ChatGPT" is used as this AI model. The analysis method inputs the text data as a prompt sentence and compares it with a large amount of past fraud data to evaluate the possibility of fraud.
[1417] Prompt Sentence Examples
[1418] Prompt: "Analyze whether this text is a potential scam: 'My son was in an accident and I need money.'"
[1419] Example of operation
[1420] If the server determines that the text matches a fraud pattern based on the prompt above, it will assess the likelihood of fraud.
[1421] warning means
[1422] Hardware and Software
[1423] If the server determines that fraud is likely based on the analysis results, it generates an alert message containing the suspected fraudulent conversation content and sends the alert message to family members or the police using a secure communication protocol (HTTPS).
[1424] Example of operation
[1425] For example, the server could generate a message like this and send it to your family or the police:
[1426] "This could be a scam. The conversation goes like this: 'My son was in an accident and I need money.'"
[1427] Recording transmission results
[1428] The server checks whether the alert message was sent successfully and records the result.
[1429] summary
[1430] This system can collect voice recordings in real time before seniors fall victim to fraud, remove noise, convert the audio into text, and use a generative AI model to identify potential fraud. If a fraud is detected, a warning message can be sent quickly to family members or the police, effectively protecting seniors from fraud.
[1431] The flow of the identification process in the first embodiment will be described with reference to FIG.
[1432] Step 1: Collecting audio
[1433] Subject: Terminal
[1434] The device constantly collects surrounding sounds using a highly sensitive microphone (e.g., a microphone with noise-canceling technology). The microphone is activated and begins collecting audio data immediately after the user answers a call. The input is the audio of the conversation between the user and the fraudster, and the output is raw audio data.
[1435] Specific behavior:
[1436] From the moment the user says "Hello?", the device collects the voice and stores it as audio data.
[1437] Step 2: Noise filtering
[1438] Subject: Terminal
[1439] The device performs noise filtering on the collected audio data. It uses open-source software to remove noise. Raw audio data is used as input, and clear audio data after noise removal is generated as output.
[1440] Specific behavior:
[1441] It filters out background noise and reveals only the conversation between you and the fraudster.
[1442] Step 3: Voice Recognition
[1443] Subject: Terminal
[1444] The device then runs the noise-filtered audio data through a speech recognition algorithm to convert it into text data. This process uses the Speech-to-Text API, which takes the filtered audio data as input and produces the converted text data as output.
[1445] Specific behavior:
[1446] The conversation between the user and the scammer, such as "If you don't transfer the money now, I can't have the surgery!", is recorded as text.
[1447] Step 4: Sending text data
[1448] Subject: Terminal
[1449] The terminal sends the converted text data to the server. HTTPS is used as the communication protocol. The converted text data is sent as input and sent to the server.
[1450] Specific behavior:
[1451] The converted text, such as "Mom, help me! I've been in an accident and need money," is securely sent to a server.
[1452] Step 5: Analyzing the text data
[1453] Subject: Server
[1454] The server receives the text data sent from the device and inputs it into the generative AI model for analysis. The prompt sentence "Analyze whether this text is likely to be fraudulent" is used as input. Based on the input, the generative AI model performs analysis and obtains an output that evaluates the likelihood of fraud.
[1455] Specific behavior:
[1456] Enter the prompt text: "Analyze whether this text is likely to be fraudulent: 'My son was in an accident and I need money'" and match it with fraud patterns.
[1457] Step 6: Fraud detection
[1458] Subject: Server
[1459] The server receives the results of the generative AI model, and if there is a high possibility of fraud, it tags the information and forwards it to the warning means. The analysis results obtained from the generative AI model are used as input, and data with a fraud judgment is generated as output.
[1460] Specific behavior:
[1461] Generate result data such as "This text is likely to be fraudulent" and send it to the warning means.
[1462] Step 7: Generate an alert message
[1463] Subject: Server
[1464] The server generates an alert message based on the fraud detection data. This message contains the conversation content that is suspected to be fraudulent. The fraud detection data is used as input, and the alert message is generated as output.
[1465] Specific behavior:
[1466] Generates messages such as "Possible scam. Conversation: 'My son was in an accident and I need money.'"
[1467] Step 8: Sending an alert message
[1468] Subject: Server
[1469] The server sends the generated alert message to registered family members and the police. The communication method is SMS or email. The generated alert message is used as input and is sent to family members and the police as output.
[1470] Specific behavior:
[1471] The generated alert message is sent to family members or the police by referring to their contact list. For example, a message saying "Your son has been in an accident and you need money" is sent to a mother's smartphone.
[1472] Step 9: Record the transmission results
[1473] Subject: Server
[1474] The server checks whether the alert message was sent successfully and records the result in the database. It uses the sending log as input and generates the sending result record data as output.
[1475] Specific behavior:
[1476] Check whether the transmission was successful or failed and store the details in a recording database. For example, if the transmission was successful, it will be logged as "successful", and if the transmission failed, it will be logged as "failed".
[1477] (Application example 1)
[1478] Next, a description will be given of Application Example 1. In the following description, the data processing device 12 will be referred to as a "server" and the robot 414 will be referred to as a "terminal."
[1479] In recent years, there has been an increase in special frauds targeting the elderly, particularly "it's me" frauds. This has caused significant economic and psychological damage to many elderly people and their families. Current security measures are often reactive, providing means to prevent damage after the fraud has already progressed. Against this background, there is a need for the development of an effective system that can detect fraudulent activities in real time and respond immediately.
[1480] The specific processing by the specific processing unit 290 of the data processing device 12 in the application example 1 is realized by the following means.
[1481] In this invention, the server includes a conversion means for constantly listening to voice and converting it into text data, an analysis means for analyzing the text data and determining the possibility of fraud, a warning means for sending an alert when it is determined that there is a possibility of fraud, and a communication means for sending the fraud analysis results to others based on the converted text data. This makes it possible to detect signs of fraud in real time and quickly issue a warning before any fraud damage occurs.
[1482] "Speech" refers to acoustic signals such as human speech and words.
[1483] "Conversion means" refers to a device or technology that converts speech into text data.
[1484] "Analysis means" refers to a device or technology that analyzes the converted text data and determines the possibility of fraud.
[1485] An "alert method" is a device or technology that sends an alert when it determines that fraud is likely.
[1486] "Communication means" refers to a device or technology for transmitting analysis results to other devices or users.
[1487] A "generative artificial intelligence model" is an artificial intelligence algorithm that has the ability to generate new data based on past data and recognize specific patterns.
[1488] A "fraud pattern" is a combination of language or phrases characteristic of fraudulent activity.
[1489] An "alert message" is a message containing a warning sent by a warning means.
[1490] "Family" refers to the user's blood relatives or close acquaintances.
[1491] A "police agency" is a public institution with legal authority to investigate crimes and protect the safety of citizens.
[1492] "Text data" is data in which voice is expressed as characters.
[1493] The present invention relates to a system for preventing special frauds targeting elderly people, and aims to prevent fraud in real time by constantly collecting voices, determining the possibility of fraud, and sending warnings to family members or police agencies if necessary.
[1494] Audio conversion means
[1495] The device constantly collects surrounding sounds using a microphone. The collected audio data is processed in real time, noise-filtered, and then converted into text data using a speech recognition algorithm. This text data is then sent to an analysis tool. Specifically, the device uses the Python speech_recognition library and Google's speech recognition service.
[1496] Analysis means
[1497] The server receives the text data sent from the device. The server then uses a generative AI model to analyze this text data. Specifically, it uses OpenAI's API to determine whether the text data matches a fraud pattern. This analysis involves evaluating the text based on past fraud data. The analysis method requests the generative AI model to analyze using a prompt sentence such as the following:
[1498] Prompt Sentence Examples
[1499] "Please determine if the following text is a scam: Mom, help me! I've been in an accident and need money."
[1500] warning means
[1501] If the server determines that there is a high possibility of fraud based on the analysis results, it generates an alert message. This message contains the suspected fraudulent conversation content. The warning mechanism then sends this alert message to registered family members and police agencies. The Python smtplib library is used to send the email. As an example of an alert, the following message is sent to family members:
[1502] "Possible scam. Dialogue: 'Mom, help me! I've been in an accident and need money.'"
[1503] Specific examples
[1504] For example, consider the following situation:
[1505] User: "Hello?"
[1506] Scammer: "Mom, help me! I've been in an accident and need money."
[1507] User: "Really? How did that happen?"
[1508] Scammer: "If you don't transfer the money now, I can't have the surgery!"
[1509] The device collects this conversation, performs speech recognition, and converts it into text data such as, "Mom, help me! I was in an accident and need money," "Really? How did that happen?", and "If you don't transfer the money now, I can't have the surgery!" This text data is then analyzed using a generative artificial intelligence model (such as OpenAI's API), and if it is determined to be a fraud, a warning message is sent to the family and police.
[1510] The system will enable seniors to take prompt action before they become victims of fraud, strengthening security measures.
[1511] The flow of the specific processing in the application example 1 will be described with reference to FIG.
[1512] Step 1:
[1513] The device uses a microphone to collect surrounding audio. The input is surrounding audio data, which the device filters in real time and converts into text data using a speech recognition algorithm. The output is the converted text data.
[1514] Step 2:
[1515] The terminal transmits the text data obtained in step 1 to the server. The input is the converted text data, and the output is the text data transmitted to the server. This data transmission utilizes an internet connection.
[1516] Step 3:
[1517] The server sends a prompt to the generative AI model to analyze the received text data. The input is the received text data, and the output is the analysis result obtained from the AI model. Specifically, the server uses the OpenAI API to generate the prompt text as follows:
[1518] "Please determine if the following text is a scam: [text data]"
[1519] Step 4:
[1520] The generative AI model analyzes the prompts sent from the server and determines whether the text data matches a fraud pattern. The input is a prompt containing the text data to be analyzed, and the output is the analysis result. Specifically, it compares the result with past fraud data.
[1521] Step 5:
[1522] The server receives the analysis results from the generative AI model and evaluates whether there is a high probability of fraud. The input is the analysis results from the AI model, and the output is the evaluation result indicating the possibility of fraud. Specifically, it checks whether the evaluation result meets certain criteria.
[1523] Step 6:
[1524] If it is determined that there is a high possibility of fraud, the server generates an alert message. The input is the evaluation result, and the output is the alert message. The specific operation is to generate a message containing the conversation content that is suspected to be fraudulent.
[1525] Step 7:
[1526] The server sends the generated alert message to the family or police. The input is the alert message, and the output is the message sent to the family or police. Specifically, it uses the Python smtplib library to send the email.
[1527] The above are the specific processing steps of the system embodying the present invention.
[1528] Furthermore, an emotion engine that estimates the user's emotion may be further combined. That is, the identification processing unit 290 may estimate the user's emotion using the emotion identification model 59, and perform identification processing using the user's emotion.
[1529] This invention is a system that prevents victims of special frauds targeting the elderly, and in particular combines voice recognition technology, a generative AI model, and an emotion engine that recognizes the user's emotions.
[1530] System configuration
[1531] The system includes the following main means:
[1532] 1. Audio conversion method
[1533] 2. Analysis method
[1534] 3. Warning measures
[1535] 4. Emotion recognition means
[1536] Audio conversion means
[1537] Subject: Terminal
[1538] The device constantly collects surrounding sounds through a microphone. This audio data is processed in real time, noise-filtered, and then converted into text data using a speech recognition algorithm. This text data is then sent to the analysis means.
[1539] As a concrete example, if a user says on the phone, "My son was in an accident and I need money," the device will collect this speech and convert it into text: "My son was in an accident and I need money."
[1540] Analysis means
[1541] Subject: Server
[1542] The server receives text data sent from the device. The server analyzes this text data using a generative artificial intelligence model (e.g., ChatGPT). It compares this text with past fraud patterns and determines whether the text is likely to be fraudulent. The analysis results are sent to the warning mechanism.
[1543] For example, if the text "My son had an accident and I need money" matches a fraud pattern, the server determines that it is likely fraudulent.
[1544] emotion recognition means
[1545] Subject: Server
[1546] The server also has an emotion recognition engine that analyzes the user's emotional state from the voice data. This emotion recognition engine analyzes the user's emotions based on the tone, pitch, speed, etc. of the voice and detects emotional patterns that may indicate fraud. In cooperation with the analysis means, it performs a comprehensive evaluation that includes the user's emotional state.
[1547] For example, if emotions of anxiety or confusion are detected from the user's voice, this can be added to the evaluation to more accurately determine the possibility of fraud.
[1548] warning means
[1549] Subject: Server
[1550] The server generates an alert message based on the results of the analysis and emotion recognition. If it determines that there is a high possibility of fraud, it sends the alert message to family members or the police. This alert message includes the content of the conversation suspected of fraud and the evaluation result of the emotional state. The warning means checks whether the message was sent successfully and records the sending result in a log.
[1551] For example, a message sent to family members or police might read: "Possible scam. Dialogue: 'My son was in an accident and I need money.' User's emotional state: Anxiety."
[1552] Specific examples
[1553] If a user is having a conversation like this:
[1554] User: "Hello?"
[1555] Scammer: "Mom, help me! I've been in an accident and need money."
[1556] User: "Really? How did that happen?"
[1557] Scammer: "If you don't transfer the money now, I can't have the surgery!"
[1558] Audio conversion means
[1559] The device collects this conversation, filters out noise, and then converts it into text data.
[1560] Translated text: "Mom, help me! I was in an accident and need money." "Really? How did that happen?" "If you don't send the money now, I can't have my surgery!"
[1561] Analysis means
[1562] The server receives this text data and analyzes it using a generative artificial intelligence model.
[1563] ChatGPT is used to detect keywords such as "accident," "money," and "transfer," and determine whether this text matches a fraud pattern.
[1564] emotion recognition means
[1565] The server uses the voice data to analyze the user's emotional state.
[1566] The emotion recognition engine detects anxiety in the user's voice.
[1567] warning means
[1568] The server determines that there is a high probability of fraud and creates an alert message.
[1569] The message contains the potentially fraudulent content: "Mom, help me! I've been in an accident and need money" and the user's emotional state: "anxious."
[1570] This message will be sent to the family and police.
[1571] This system can effectively detect fraud before it happens to elderly people and quickly alert family members and the police, thereby preventing fraud from occurring.
[1572] The processing flow will be explained below.
[1573] Step 1: Collecting audio
[1574] Subject: Terminal
[1575] The device constantly collects surrounding sounds through a microphone.
[1576] The collected audio data is stored in a buffer in real time.
[1577] Step 2: Filtering the noise
[1578] Subject: Terminal
[1579] The terminal performs noise reduction on the audio data in the buffer.
[1580] Filters out background sounds and noise to extract clear audio data.
[1581] Step 3: Speech to text
[1582] Subject: Terminal
[1583] The terminal uses a speech recognition algorithm to convert the filtered voice data into text data.
[1584] The converted text data is sent to the server.
[1585] Step 4: Receiving text data
[1586] Subject: Server
[1587] The server receives the text data sent from the terminal.
[1588] The received text data is stored in memory for analysis.
[1589] Step 5: Analyze potential fraud
[1590] Subject: Server
[1591] The server analyzes the text data using a generative artificial intelligence model (e.g., ChatGPT).
[1592] During the analysis process, the text data is compared with many past fraud patterns.
[1593] Assess whether there is potential for fraud.
[1594] Step 6: Determine the likelihood of fraud
[1595] Subject: Server
[1596] Based on the analysis results, the server determines that there is a possibility of fraud.
[1597] If there is a possibility of fraud, the information is sent to the alert module.
[1598] Step 7: Collect emotion data
[1599] Subject: Terminal
[1600] The device uses voice data to collect data to analyze the user's emotional state using an emotion engine.
[1601] The collected emotion data is sent to a server.
[1602] Step 8: Analyze the sentiment data
[1603] Subject: Server
[1604] The server uses an emotion engine to analyze the emotion data sent from the terminal.
[1605] Evaluates the user's emotional state based on the tone, pitch, speed, etc. of the voice.
[1606] Detect emotional patterns that indicate potential fraud.
[1607] Step 9: Integrating and analyzing emotion data
[1608] Subject: Server
[1609] The server integrates the results of the text data analysis and the emotional data analysis, and comprehensively reassess the likelihood of fraud.
[1610] Determine with greater accuracy whether there is a high likelihood of fraud.
[1611] Step 10: Prepare alert information
[1612] Subject: Server
[1613] The server generates an alert message based on information that is determined to be potentially fraudulent.
[1614] The generated alert message includes the suspected fraudulent conversation content and an assessment of the emotional state.
[1615] Step 11: Sending an alert
[1616] Subject: Server
[1617] The server will then send an alert message to registered family members and police contacts.
[1618] Check the transmission result, and if the transmission is successful, record the result in the log.
[1619] Example 2
[1620] Next, a description will be given of Example 2. In the following description, the data processing device 12 will be referred to as a "server" and the robot 414 will be referred to as a "terminal."
[1621] The number of victims of special frauds targeting the elderly is increasing year by year. Fraudsters use sophisticated methods to deceive the elderly and obtain money illegally. In order to prevent such fraud, a system is needed that can monitor conversations around the elderly in real time and detect suspicious speech. It is also desirable to evaluate the user's emotional state to detect factors that increase the likelihood of fraud.
[1622] The identification process by the identification processing unit 290 of the data processing device 12 in the second embodiment is realized by the following means. In this invention, the server includes a sound collection means that constantly collects environmental sounds and performs noise filtering, a sound recognition means that converts the collected sound data into text data, an analysis means that analyzes the text data and determines the possibility of fraud, an emotion recognition means that analyzes the emotional state of the user, and a warning means that sends an alert when it is determined that there is a possibility of fraud. This makes it possible to detect the risk of elderly people falling victim to fraud early and promptly notify their family or the police.
[1623] "Environmental sound" refers to all sounds occurring around the user, and is a general term for sound data including conversations and background sounds.
[1624] "Noise filtering" is a process that removes unnecessary background sounds and noise from collected audio data and converts it into audio data that is easy to hear.
[1625] "Sound collection means" refers to a device or software that constantly collects environmental sounds using a microphone.
[1626] "Speech recognition means" refers to an algorithm or system for converting collected voice data into text data.
[1627] "Text data" refers to character string data converted from speech by speech recognition means.
[1628] A "generative artificial intelligence model" refers to an algorithm that learns from large amounts of text data and performs natural language processing, and specifically includes models such as ChatGPT.
[1629] A "prompt sentence" refers to input text used to give instructions or ask questions to a generative artificial intelligence model and obtain a specific response.
[1630] "Analysis Method" refers to an algorithm or system that uses a generative artificial intelligence model to analyze text data and determine the likelihood of fraud.
[1631] "Emotion recognition means" refers to an algorithm or system for analyzing a user's emotional state from voice data and detecting emotional patterns.
[1632] "Warning measures" refer to algorithms or systems that send alert messages to family members or police in the event of potential fraud.
[1633] "Alert Message" means a message containing warning content that is generated to notify you of suspected fraud.
[1634] "Transmission result" refers to confirmation information as to whether the alert message was sent successfully.
[1635] "Log" refers to data that records system operations, events, and the results of sending alert messages.
[1636] The present invention is a system for preventing victims of special frauds targeting elderly people. The system includes a voice collection means, a voice recognition means, an analysis means, an emotion recognition means, and a warning means. The details of each means and the specific implementation method are explained below.
[1637] Audio collection method
[1638] The device is equipped with a microphone for constantly collecting ambient sounds. The microphone is highly sensitive and capable of collecting sound over a wide range, such as a condenser microphone or digital microphone. The collected sound data is processed in real time using a noise filtering algorithm, which removes background noise and unnecessary sounds to obtain clear sound data.
[1639] Voice recognition means
[1640] The device inputs the noise-filtered voice data into a voice recognition engine. For example, Google Cloud Speech-to-Text API is used as the voice recognition engine. The voice recognition engine converts the voice data into text data and expresses the content as a string of characters. The converted text data is sent to an analysis means.
[1641] Analysis means
[1642] The server receives the text data sent from the device. The server analyzes the text data using a generative artificial intelligence model (e.g., ChatGPT). It generates a prompt and inputs it into the model to compare it with past fraud patterns and evaluate the likelihood of fraud. For example, the prompt could be set as "Please evaluate the likelihood that this text is a special fraud." When the analysis results are obtained, it determines whether the likelihood of fraud is high and passes the result to the next processing step.
[1643] emotion recognition means
[1644] The server has an emotion recognition engine, such as IBM Watson Tone Analyzer, that analyzes the user's emotional state from the voice data. The emotion recognition engine analyzes the tone, pitch, and speed of the voice to assess whether the user is anxious or confused. These emotional states are also considered as factors that increase the likelihood of fraud.
[1645] warning means
[1646] The server generates an alert message based on the results of the analysis and emotion recognition methods. If it determines that there is a high possibility of fraud, it sends this alert message to the user's family or the police. The alert message includes the content of the conversation that is suspected to be fraudulent and the user's emotional state. For example, a message such as "Possible fraud. Conversation content: 'Mom, help me! I've been in an accident and need money.' User's emotional state: anxious" is generated. The server then checks the transmission result and records in a log that it was sent successfully.
[1647] Specific examples
[1648] Here is a concrete example of the system: If a user is having the following conversation:
[1649] User: "Hello?"
[1650] Scammer: "Mom, help me! I've been in an accident and need money."
[1651] User: "Really? How did that happen?"
[1652] Scammer: "If you don't transfer the money now, I can't have the surgery!"
[1653] 1. Audio collection means: The device collects the conversation and removes noise.
[1654] 2. Voice recognition method: After noise filtering, the conversation is converted into text such as "Mom, help me! I've been in an accident and need money," "Really? How did that happen?" and "If you don't transfer the money now I can't have the surgery!" and sent to the server.
[1655] 3. Analysis method: The server analyzes the text, detects keywords such as "accident," "money," and "transfer," and evaluates the likelihood of fraud. An example of a prompt is, "Please evaluate the likelihood that this text is a special fraud."
[1656] 4. Emotion recognition means: The server analyzes the user's emotion of anxiety from the voice.
[1657] 5. Warning method: The server determines that there is a high possibility of fraud and generates an alert message stating, "Possible fraud. Conversation: 'Mom, help! I've been in an accident and need money.' User's emotional state: Anxiety," and sends it to the family and police.
[1658] This system makes it possible to detect early the risk of elderly people falling victim to fraud and take prompt action.
[1659] The flow of the identification process in the second embodiment will be described with reference to FIG.
[1660] System program processing flow
[1661] Step 1: Audio Collection
[1662] The device collects ambient sounds, acquires audio data in real time through a microphone, and processes the data with a noise filtering algorithm. The input is the audio data of the ambient sounds, and the output is the audio data with the noise removed.
[1663] Specifically, the system captures audio data via a microphone and removes unwanted noise through a noise reduction filter.
[1664] Step 2: Voice Recognition
[1665] The device inputs the noise-filtered voice data into a voice recognition engine and converts it into text data. Here, for example, the Google Cloud Speech-to-Text API is used. The input is the voice data with noise removed, and the output is the converted text data.
[1666] Specifically, the speech data is sent to a speech recognition engine, and text data is obtained as a result.
[1667] Step 3: Text analysis
[1668] The server receives the text data sent from the device. The server analyzes this text data using a generative artificial intelligence model (such as ChatGPT). The input is the text data, and the output is the evaluation result of the likelihood of fraud. The prompt used is "Please evaluate the likelihood that this text is a special fraud."
[1669] Specifically, text data is input into a generative artificial intelligence model and the results of matching with fraud patterns are evaluated.
[1670] Step 4: Emotional state analysis
[1671] The server simultaneously analyzes the user's emotional state from the voice data. Using an emotion recognition engine (such as IBM Watson Tone Analyzer), it analyzes the tone, pitch, and speed of the voice to determine whether the user is anxious or confused. The input is the voice data, and the output is the evaluation result of the user's emotional state.
[1672] Specifically, the voice data is input into an emotion recognition engine to obtain an evaluation result of the emotional state.
[1673] Step 5: Generate and send an alert
[1674] The server generates an alert message based on the results of the analysis and emotion recognition methods. If it determines that there is a high possibility of fraud, it sends the alert message to the user's family or the police. The input is the evaluation result of the possibility of fraud and the evaluation result of the user's emotional state, and the output is the alert message.
[1675] Specifically, the system generates an alert message based on the analysis results and the user's emotional state, and sends it to family members or the police via email or SMS. The system then records the results of the message in a log.
[1676] This series of processes makes it possible to detect early on when an elderly person is at risk of falling victim to fraud and quickly notify their family or the police.
[1677] (Application example 2)
[1678] Next, a description will be given of Application Example 2. In the following description, the data processing device 12 will be referred to as a "server" and the robot 414 will be referred to as a "terminal."
[1679] Special frauds targeting the elderly remain a social problem, with many elderly people suffering financial losses. The present invention aims to provide an effective system for preventing elderly people from falling victim to fraud. In particular, the objective of the present invention is to protect elderly people from fraud by detecting signs of fraud early and issuing prompt warnings.
[1680] The identification process by the identification processing unit 290 of the data processing device 12 in Application Example 2 is realized by the following means. In this invention, the server includes a voice conversion means that constantly collects voice and converts it into text data, an analysis means that analyzes the text data and evaluates the possibility of fraud, an emotion recognition means that analyzes the emotional state from the voice data, and a warning means that generates and transmits a warning message when it is determined that there is a possibility of fraud. This enables elderly people to quickly recognize the possibility of fraud and issue a warning to their family or public institutions.
[1681] The "voice conversion means" is a unit that constantly collects voice and converts it into text data.
[1682] The "analysis means" is a unit for analyzing the text data and assessing the possibility of fraud.
[1683] The "emotion recognition means" is a unit that analyzes the emotional state from the voice data.
[1684] The "warning means" is a unit for generating and transmitting a warning message when it is determined that there is a possibility of fraud.
[1685] A "generative artificial intelligence model" is an artificial intelligence algorithm generated from past data and learning, and is used as a means of analyzing text data.
[1686] "Text data" refers to data in character format converted from speech by speech conversion means.
[1687] A "fraud pattern" is a collection of characteristic words or phrases that indicate fraud.
[1688] The "emotional state" is the emotional state of the user that is analyzed based on the tone, pitch, speed, etc. of the voice data.
[1689] A "warning message" is a notification that is generated when potential fraud is detected.
[1690] "Family members and public authorities" are those who are notified by the warning measures of possible fraud.
[1691] The present invention is a system for preventing victims of special frauds targeting elderly people, and includes a voice conversion means, an analysis means, an emotion recognition means, and a warning means. Specific embodiments for carrying out the present invention will be described below.
[1692] System configuration
[1693] The system includes a dedicated terminal, a server, and a user terminal (e.g., a smartphone).
[1694] Audio conversion means
[1695] The device constantly collects ambient sound through a microphone. This sound data is processed in real time and noise filtered. Speech recognition software (e.g., the Python speech_recognition library) then converts the sound data into text data, which is then sent to an analysis unit.
[1696] Analysis means
[1697] The server receives text data sent from the device. Using a generative artificial intelligence model (e.g., ChatGPT), it analyzes this text data and matches it with fraud patterns. If the text indicates possible fraud, the results are sent to the emotion recognition means.
[1698] emotion recognition means
[1699] The server then uses an emotion recognition engine (e.g., EmotionRecognizer) to analyze the user's emotional state using the voice data. This engine analyzes the user's emotions based on the tone, pitch, and speed of the voice, and detects emotional patterns that may indicate fraud. It then makes a comprehensive assessment, including the user's emotional state.
[1700] warning means
[1701] If the server determines that a message is likely to be fraudulent, it generates a warning message that is sent to family members or public institutions. The warning message includes the content of the conversation that is suspected to be fraudulent and the result of the evaluation of the emotional state. The warning means checks whether the message was sent successfully and records the sending result in a log.
[1702] Specific examples
[1703] If a user is having a conversation like this:
[1704] User: "Hello?"
[1705] Scammer: "Mom, help me! I've been in an accident and need money."
[1706] User: "Really? How did that happen?"
[1707] Scammer: "If you don't transfer the money now, I can't have the surgery!"
[1708] Audio conversion means
[1709] The device collects this conversation, filters out noise, and converts it into text data, which looks like this:
[1710] "Mom, help me! I was in an accident and I need money." "Really? Why did that happen?" "If you don't transfer the money right now, I can't have the surgery!"
[1711] Analysis means
[1712] The server receives this text data and analyzes it using a generative artificial intelligence model. Because the text contains keywords such as "accident," "money," and "transfer," it is determined to be highly likely to be fraudulent.
[1713] emotion recognition means
[1714] The server uses the voice data to analyze the user's emotional state, and an emotion recognition engine detects anxiety from the user's voice.
[1715] warning means
[1716] The server determines that the fraud is likely and creates an alert message containing the suspected fraudulent content: "Mom, help me! I've been in an accident and need money" and the user's emotional state: "anxious." This message is sent to family members and public authorities.
[1717] Prompt Sentence Examples
[1718] "I received a suspicious call. I would like to provide the following information:
[1719] Conversation: 'My son had an accident and I need money.'
[1720] Emotional state: 'Anxious'
[1721] Please take immediate action as this may be a scam."
[1722] The flow of the specific processing in the application example 2 will be described with reference to FIG.
[1723] Step 1:
[1724] The device constantly collects ambient sound through a microphone. This sound data is processed in real time and noise filtered. The input is raw sound data, and the output is noise-removed sound data. Specifically, the sound signal obtained through the microphone is passed through a digital filter to reduce noise.
[1725] Step 2:
[1726] The device converts the noise-removed audio data into text data using speech recognition software (e.g., Python's speech_recognition library). The input is the noise-removed audio data, and the output is the corresponding text data. Specifically, the speech recognition algorithm analyzes the audio signal and converts it into the corresponding string of characters.
[1727] Step 3:
[1728] The terminal transmits the converted text data to the server. The input is the text data, and the output is data transmission to the server. Specifically, the text data is transferred to the server via a network connection.
[1729] Step 4:
[1730] The server analyzes the text data received from the device using a generative artificial intelligence model (e.g., ChatGPT) to assess the likelihood of fraud. The input is text data, and the output is an assessment of whether the likelihood of fraud is high. Specifically, the chat model compares the text with past fraud patterns to detect signs of fraud.
[1731] Step 5:
[1732] The server analyzes the voice data using an emotion recognition engine (e.g., EmotionRecognizer) to assess the user's emotional state. The input is the voice data, and the output is the assessment of the emotional state. Specifically, the tone, pitch, and speed of the voice are analyzed to identify specific emotions, such as anxiety.
[1733] Step 6:
[1734] The server comprehensively judges the fraud possibility assessment result and emotional state, and generates a warning message if it determines that fraud is highly likely. The input is the assessment result and emotional state, and the output is the warning message. Specifically, the warning message is created using a text template.
[1735] Step 7:
[1736] The server sends the generated warning message to the family or public institution. The input is the warning message, and the output is the message to be sent to the message recipient. Specifically, the server sends the warning message to the specified email address using an email protocol (e.g., SMTP).
[1737] Step 8:
[1738] The server checks whether the alert message was sent successfully and logs the sending result. The input is the sending status and the output is writing to the log file. Specifically, it checks the status of whether the sending was successful and records the result in the log file.
[1739] The specific processing unit 290 transmits the result of the specific processing to the robot 414. In the robot 414, the control unit 46A causes the speaker 240 and the control target 443 to output the result of the specific processing. The microphone 238 acquires voice indicating a user input regarding the result of the specific processing. The control unit 46A transmits voice data indicating the user input acquired by the microphone 238 to the data processing device 12. In the data processing device 12, the specific processing unit 290 acquires the voice data.
[1740] The data generation model 58 is a so-called generative AI (Artificial Intelligence). An example of the data generation model 58 is ChatGPT (Internet Search<URL: https: / / openai.com / blog / chatgpt> ), Gemini (Internet search <url: https: gemini.google.com ?hl="ja">) and other generation AIs. The data generation model 58 is obtained by performing deep learning on a neural network. A prompt including an instruction is input to the data generation model 58, and inference data such as voice data indicating voice, text data indicating text, and image data indicating an image is also input. The data generation model 58 performs inference on the input inference data in accordance with the instruction indicated by the prompt, and outputs the inference result in a data format such as voice data and text data. Here, inference refers to, for example, analysis, classification, prediction, and / or summarization.
[1741] In the above embodiment, an example was given in which the specific processing is performed by the data processing device 12, but the technology of the present disclosure is not limited to this, and the specific processing may be performed by the robot 414.
[1742] The emotion identification model 59 as an emotion engine may determine the user's emotion according to a specific mapping. Specifically, the emotion identification model 59 may determine the user's emotion according to an emotion map (see FIG. 9), which is a specific mapping. Similarly, the emotion identification model 59 may determine the robot's emotion, and the identification processing unit 290 may perform identification processing using the robot's emotion.
[1743] FIG. 9 is a diagram illustrating an emotion map 400 on which multiple emotions are mapped. In the emotion map 400, emotions are arranged in concentric circles radiating from the center. Emotions closer to the center of the concentric circles are more primitive. Emotions representing states and actions arising from a state of mind are arranged on the outer edges of the concentric circles. The concept of emotion includes both affect and mental states. Emotions generally generated from reactions occurring in the brain are arranged on the left side of the concentric circles. Emotions generally induced by situational judgment are arranged on the right side of the concentric circles. Emotions generally generated from reactions occurring in the brain and induced by situational judgment are arranged on the upper and lower sides of the concentric circles. Furthermore, the emotion of "pleasure" is arranged on the upper side of the concentric circles, and the emotion of "discomfort" is arranged on the lower side. In this way, in the emotion map 400, multiple emotions are mapped based on the structure by which emotions are generated, and emotions that tend to occur simultaneously are mapped close to each other.
[1744] These emotions are distributed in the 3 o'clock direction on emotion map 400, and typically fluctuate between relief and anxiety. In the right half of emotion map 400, situational awareness dominates over internal sensations, resulting in a sense of calm.
[1745] The inside of emotion map 400 represents what is going on in the mind, and the outside of emotion map 400 represents behavior, so the further you go outside emotion map 400, the more visible the emotions become (the more they are expressed in behavior).
[1746] Human emotions are based on various balances, such as posture and blood sugar levels. When these balances deviate from the ideal, a state of discomfort is indicated, and when they approach the ideal, a state of pleasure is indicated. Emotions can also be created for robots, automobiles, and motorcycles, based on various balances, such as posture and remaining battery life. When these balances deviate from the ideal, a state of discomfort is indicated, and when they approach the ideal, a state of pleasure is indicated. An emotion map can be generated, for example, based on Dr. Mitsuyoshi's emotion map (Research on Voice Emotion Recognition and Emotional Brain Physiological Signal Analysis Systems, Tokushima University, Doctoral Dissertation: https: / / ci.nii.ac.jp / naid / 500000375379). The left half of the emotion map lists emotions belonging to the "reaction" domain, where sensation is dominant. The right half of the emotion map lists emotions belonging to the "situation" domain, where situational awareness is dominant.
[1747] The emotion map defines two emotions that promote learning. One is a negative emotion on the situation side, around the middle of "repentance" or "reflection." In other words, this occurs when the robot experiences negative emotions such as "I never want to feel this way again" or "I don't want to be scolded again." The other is a positive emotion on the response side, around "desire." In other words, this occurs when the robot experiences positive feelings such as "I want more" or "I want to know more."
[1748] The emotion identification model 59 inputs user input into a pre-trained neural network, obtains emotion values indicating each emotion shown in the emotion map 400, and determines the user's emotion. This neural network is pre-trained based on multiple pieces of training data that are combinations of user input and emotion values indicating each emotion shown in the emotion map 400. Furthermore, this neural network is trained so that emotions that are located close to each other have similar values, as in the emotion map 900 shown in FIG. 10. FIG. 10 shows an example in which multiple emotions, "relieved," "calm," and "reassuring," have similar emotion values.
[1749] The system according to the present disclosure has been described above mainly with respect to the functions of the data processing device 12, but the system according to the present disclosure is not necessarily implemented on a server. The system according to the present disclosure may be implemented as a general information processing system. The present disclosure may be implemented, for example, as a software program running on a personal computer or an application running on a smartphone, etc. The method according to the present disclosure may be provided to users in the form of SaaS (Software as a Service).
[1750] In the above embodiment, an example was given in which the specific processing is performed by one computer 22, but the technology of the present disclosure is not limited to this, and the specific processing may be distributed and performed by a plurality of computers including the computer 22. For example, the data generation model 58 may be provided in an external device of the data processing device 12, and data may be generated in the external device in accordance with input data.
[1751] In the above embodiment, an example in which the specific processing program 56 is stored in the storage 32 has been described, but the technology of the present disclosure is not limited to this. For example, the specific processing program 56 may be stored in a portable, computer-readable, non-transitory storage medium such as a USB (Universal Serial Bus) memory. The specific processing program 56 stored in the non-transitory storage medium is installed in the computer 22 of the data processing device 12. The processor 28 executes the specific processing in accordance with the specific processing program 56.
[1752] Alternatively, the specific processing program 56 may be stored in a storage device such as a server connected to the data processing device 12 via the network 54, and the specific processing program 56 may be downloaded and installed on the computer 22 in response to a request from the data processing device 12.
[1753] It is not necessary to store all of the specific processing program 56 in a storage device such as a server connected to the data processing device 12 via the network 54, or to store all of the specific processing program 56 in the storage 32; only a portion of the specific processing program 56 may be stored.
[1754] The hardware resource for executing a specific process can be any of the following processors: An example of a processor is a CPU, which is a general-purpose processor that functions as a hardware resource for executing a specific process by executing software, i.e., a program. Another example of a processor is a dedicated electrical circuit, such as an FPGA (Field-Programmable Gate Array), a PLD (Programmable Logic Device), or an ASIC (Application Specific Integrated Circuit), which is a processor with a circuit configuration designed specifically for executing a specific process. Each processor has built-in or connected memory, and each processor uses the memory to execute the specific process.
[1755] The hardware resource that executes the specific processing may be configured with one of these various processors, or may be configured with a combination of two or more processors of the same or different types (for example, a combination of multiple FPGAs, or a combination of a CPU and an FPGA). Also, the hardware resource that executes the specific processing may be a single processor.
[1756] As an example of a system configured with a single processor, first, one processor is configured by combining one or more CPUs and software, and this processor functions as a hardware resource that executes a specific process. Second, there is a system that uses a processor that realizes the functions of an entire system including multiple hardware resources that execute a specific process on a single IC chip, as typified by SoC (System-on-a-chip). In this way, a specific process is realized using one or more of the above-mentioned various processors as hardware resources.
[1757] Furthermore, the hardware structure of these various processors can be, more specifically, an electric circuit that combines circuit elements such as semiconductor devices. The specific processing described above is merely an example. Therefore, it goes without saying that unnecessary steps may be deleted, new steps may be added, or the processing order may be rearranged, without departing from the spirit of the invention.
[1758] The above-described description and illustrations are a detailed explanation of the parts related to the technology of the present disclosure and are merely an example of the technology of the present disclosure. For example, the above description of the configuration, functions, actions, and effects is an explanation of an example of the configuration, functions, actions, and effects of the parts related to the technology of the present disclosure. Therefore, it goes without saying that unnecessary parts may be deleted, new elements may be added, or replacements may be made to the above-described description and illustrations within the scope of the gist of the technology of the present disclosure. Furthermore, to avoid confusion and facilitate understanding of the parts related to the technology of the present disclosure, the above-described description and illustrations omit explanations of common technical knowledge that do not require particular explanation to enable the implementation of the technology of the present disclosure.
[1759] All publications, patent applications, and technical standards mentioned in this specification are herein incorporated by reference to the same extent as if each individual publication, patent application, or technical standard was specifically and individually indicated to be incorporated by reference.
[1760] The following is further disclosed regarding the above embodiment.
[1761] (Claim 1)
[1762] a speech conversion means for constantly listening to speech and converting it into text data;
[1763] analysis means for analyzing the text data and determining the possibility of fraud;
[1764] a warning mechanism to send alerts when fraud is suspected;
[1765] A system including:
[1766] (Claim 2)
[1767] The system of claim 1, wherein the analysis means analyzes the text data using a generative artificial intelligence model and compares it with fraud patterns to assess the likelihood of fraud.
[1768] (Claim 3)
[1769] 2. The system according to claim 1, wherein the warning means has a function of sending an alert message to family members or police in the event of a possible fraud.
[1770] "Example 1"
[1771] (Claim 1)
[1772] a speech conversion means for constantly listening to speech and converting it into text data;
[1773] analysis means for analyzing the text data and determining the possibility of fraud;
[1774] a warning mechanism to send alerts when fraud is suspected;
[1775] means for noise filtering the voice in the voice conversion means;
[1776] means for analyzing using a generative artificial intelligence model in the analysis means;
[1777] means for generating and transmitting an alert message when the warning means matches a fraud pattern;
[1778] A system including:
[1779] (Claim 2)
[1780] The system of claim 1, wherein the analysis means utilizes a generative artificial intelligence model to input text data as a prompt sentence, and compares it with fraud patterns to assess the likelihood of fraud.
[1781] (Claim 3)
[1782] 2. The system according to claim 1, wherein the warning means has a function of sending an alert message to family members or the police in the event of a possible fraud, and recording the results of the message transmission.
[1783] "Application Example 1"
[1784] (Claim 1)
[1785] A conversion means for constantly listening to the voice and converting it into text data;
[1786] analysis means for analyzing the text data and determining the possibility of fraud;
[1787] a warning mechanism to send alerts when fraud is suspected;
[1788] A communication means for transmitting the fraud analysis results to another person based on the converted text data;
[1789] A system including:
[1790] (Claim 2)
[1791] The system of claim 1, wherein the analysis means analyzes the text data using a generative artificial intelligence model and compares it with fraud patterns to assess the likelihood of fraud.
[1792] (Claim 3)
[1793] 2. The system of claim 1, wherein the warning means has a function of sending an alert message to family members or police agencies in the event of a possible fraud.
[1794] "Example 2: Combining Emotion Engines"
[1795] (Claim 1)
[1796] an audio collection means for constantly collecting environmental sounds and performing noise filtering;
[1797] A speech recognition means for converting collected speech data into text data;
[1798] analysis means for analyzing the text data and determining the possibility of fraud;
[1799] emotion recognition means for analyzing the emotional state of a user;
[1800] a warning mechanism to send alerts when fraud is suspected;
[1801] A system including:
[1802] (Claim 2)
[1803] The system of claim 1, wherein the analysis means utilizes a generative artificial intelligence model to input text data as a prompt sentence, and compares it with fraud patterns to assess the likelihood of fraud.
[1804] (Claim 3)
[1805] 2. The system according to claim 1, wherein the warning means has a function of sending an alert message to family members or the police in the event of a possible fraud, and records the transmission result in a log.
[1806] "Application example 2 when combining emotion engines"
[1807] (Claim 1)
[1808] A speech conversion means for constantly collecting speech and converting it into text data;
[1809] analysis means for analyzing the text data and assessing the likelihood of fraud;
[1810] emotion recognition means for analyzing an emotional state from voice data;
[1811] a warning means for generating and sending a warning message when a potential fraud is determined;
[1812] A system including:
[1813] (Claim 2)
[1814] 2. The system of claim 1, further comprising an analysis means for analyzing text data using a generative artificial intelligence model and comparing it with fraud patterns to assess the likelihood of fraud.
[1815] (Claim 3)
[1816] 10. The system of claim 1, further comprising a function for sending a warning message to family members or public authorities in the event of a possible fraud. [Explanation of symbols]
[1817] 10, 210, 310, 410 Data Processing Systems 12 Data Processing Device 14 Smart Devices 214 Smart Glasses 314 Headset-type terminal 414 Robot< / url:> < / url:> < / url:> < / url:>
Claims
1. a speech conversion means for constantly listening to speech and converting it into text data; analysis means for analyzing the text data and determining the possibility of fraud; a warning mechanism to send alerts when fraud is suspected; A system including:
2. 2. The system according to claim 1, wherein the analysis means analyzes the text data using a generative artificial intelligence model and compares it with fraud patterns to assess the likelihood of fraud.
3. 2. The system according to claim 1, wherein the warning means has a function of sending an alert message to family members or the police in the event of a possible fraud.
Citation Information
Patent Citations
Persona chatbot control method and system
JP2022180282A