system
A system using location and voice data analysis with a generative AI model generates real-time fraud warnings to prevent remittance fraud, addressing the lack of effective fraud prevention in conventional technologies.
Patent Information
- Authority / Receiving Office
- JP · JP
- Patent Type
- Applications
- Current Assignee / Owner
- Filing Date
- 2024-09-04
- Publication Date
- 2026-03-16
AI Technical Summary
Conventional technologies lack mechanisms to utilize location information and voice data of smartphones effectively for preventing remittance fraud, especially targeting the elderly, leading to high risks of fraud without immediate user intervention.
A system that collects user location information and voice data, encrypts the voice data, and analyzes it using a generative AI model to generate and notify users and pre-registered contacts of potential fraud risks through both voice and text warnings.
Significantly reduces the risk of users falling victim to wire fraud by providing real-time fraud detection and immediate warnings, enabling swift user action to prevent damage.
Smart Images

Figure 2026047904000001_ABST
Abstract
Description
Technical Field
[0001] The technology of the present disclosure relates to a system.
Background Art
[0002] Patent Document 1 discloses a method for controlling a persona chatbot, which is performed by at least one processor, including steps of receiving a user utterance, adding the user utterance to a prompt including an instruction sentence related to an explanation of a chatbot character, encoding the prompt, and inputting the encoded prompt into a language model to generate a chatbot utterance in response to the user utterance.
Prior Art Documents
Patent Documents
[0003]
Patent Document 1
Summary of the Invention
Problems to be Solved by the Invention
[0004] Remittance fraud targeting the elderly has not stopped, especially when using bank ATMs or when money is requested over the phone. However, conventional technologies lack a mechanism to directly utilize location information and voice data of smartphones, etc. for improving user convenience and safety. As a result, the risk for users to suffer fraud is high. Therefore, an object of the present invention is to provide means for users, including the elderly, to confirm and prevent the risk of fraud in advance.
Means for Solving the Problems
[0005] The present invention provides location information collection means for collecting user location information and location information transmission means for transmitting the location information to a server. Furthermore, it provides voice data collection means and voice data encryption means for collecting and encrypting user voice data. The encrypted voice data is transmitted to the server, which includes analysis means for integrating and analyzing the location information and voice data. Based on the analysis results, the server generates a warning message and includes warning notification means for notifying the user terminal of the warning message. It also includes notification means for notifying pre-registered contacts of the generated warning message. In addition, it may include voice notification means and text notification means for notifying the user of the generated warning message in voice and text formats. This provides a system that allows users to recognize the risk of wire fraud in advance and take appropriate action.
[0006] "Location information collection means" refers to devices or functions used to obtain the user's current location.
[0007] "Location information transmission means" refers to a device or function for transmitting acquired user location information to other devices such as servers.
[0008] "Voice data collection means" refers to a device or function for recording the user's voice.
[0009] "Voice data encryption means" refers to devices or functions used to encrypt collected voice data so that it cannot be illegally obtained by third parties.
[0010] "Voice data transmission means" refers to a device or function for transmitting encrypted voice data to other devices such as servers.
[0011] "Analysis means" refers to devices or functions for analyzing collected location information and audio data and extracting necessary information.
[0012] A "warning message generation means" refers to a device or function for creating a warning message for the user based on the analyzed results.
[0013] A "warning notification means" refers to a device or function for notifying a user terminal of a generated warning message.
[0014] "Notification means" refers to a device or function that notifies a warning message to pre-registered contacts (such as family or friends).
[0015] "Voice notification means" refers to a device or function that notifies the user of a generated warning message in voice format.
[0016] A "text notification means" refers to a device or function that notifies the user of a generated warning message in text format. [Brief explanation of the drawing]
[0017] [Figure 1] This is a conceptual diagram showing an example of the configuration of a data processing system according to the first embodiment. [Figure 2] This is a conceptual diagram showing an example of the essential functions of the data processing device and smart device according to the first embodiment. [Figure 3] This is a conceptual diagram showing an example of the configuration of a data processing system according to the second embodiment. [Figure 4] This is a conceptual diagram showing an example of the main functions of a data processing device and smart glasses according to the second embodiment. [Figure 5] This is a conceptual diagram showing an example of the configuration of a data processing system according to the third embodiment. [Figure 6] This is a conceptual diagram showing an example of the main functions of a data processing device and a headset-type terminal according to the third embodiment. [Figure 7] This is a conceptual diagram showing an example of the configuration of a data processing system according to the fourth embodiment. [Figure 8] This is a conceptual diagram showing an example of the main functions of a data processing device and a robot according to the fourth embodiment. [Figure 9] Shows an emotion map to which a plurality of emotions are mapped. [Figure 10] Shows an emotion map to which a plurality of emotions are mapped. [Figure 11] It is a sequence diagram showing the processing flow of the data processing system in Embodiment 1. [Figure 12] It is a sequence diagram showing the processing flow of the data processing system in Application Example 1. [Figure 13] It is a sequence diagram showing the processing flow of the data processing system in Embodiment 2 when the emotion engine is combined. [Figure 14] It is a sequence diagram showing the processing flow of the data processing system in Application Example 2 when the emotion engine is combined.
Mode for Carrying Out the Invention
[0018] Hereinafter, an example of an embodiment of a system according to the technology of the present disclosure will be described with reference to the accompanying drawings.
[0019] First, the language used in the following description will be described.
[0020] In the following embodiments, the numbered processor (hereinafter simply referred to as "processor") may be one arithmetic unit or a combination of a plurality of arithmetic units. Also, the processor may be one type of arithmetic unit or a combination of a plurality of types of arithmetic units. Examples of arithmetic units include a CPU (Central Processing Unit), a GPU (Graphics Processing Unit), a GPGPU (General-Purpose computing on Graphics Processing Units), an APU (Accelerated Processing Unit), and the like.
[0021] In the following embodiments, signed RAM (Random Access Memory) is a memory that temporarily stores information and is used as work memory by the processor.
[0022] In the following embodiments, the signed storage is one or more non-volatile storage devices that store various programs and various parameters. Examples of non-volatile storage devices include flash memory (SSD (Solid State Drive)), magnetic disks (e.g., hard disks), or magnetic tapes.
[0023] In the following embodiments, the signed communication interface (I / F) is an interface that includes a communication processor and an antenna, etc. The communication interface manages communication between multiple computers. Examples of communication standards applicable to the communication interface include wireless communication standards such as 5G (5th Generation Mobile Communication System), Wi-Fi (registered trademark), or Bluetooth (registered trademark).
[0024] In the following embodiments, "A and / or B" is synonymous with "at least one of A and B." That is, "A and / or B" means that it may be A alone, or B alone, or a combination of A and B. Furthermore, in this specification, the same concept as "A and / or B" applies when expressing three or more things linked by "and / or."
[0025] [First Embodiment]
[0026] Figure 1 shows an example of the configuration of the data processing system 10 according to the first embodiment.
[0027] As shown in Figure 1, the data processing system 10 includes a data processing device 12 and a smart device 14. An example of the data processing device 12 is a server.
[0028] The data processing device 12 comprises a computer 22, a database 24, and a communication interface 26. The computer 22 is an example of a "computer" related to the technology of this disclosure. The computer 22 comprises a processor 28, RAM 30, and storage 32. The processor 28, RAM 30, and storage 32 are connected to a bus 34. The database 24 and the communication interface 26 are also connected to the bus 34. The communication interface 26 is connected to a network 54. An example of the network 54 is a WAN (Wide Area Network) and / or a LAN (Local Area Network).
[0029] The smart device 14 comprises a computer 36, a reception device 38, an output device 40, a camera 42, and a communication interface 44. The computer 36 comprises a processor 46, RAM 48, and storage 50. The processor 46, RAM 48, and storage 50 are connected to a bus 52. The reception device 38, output device 40, and camera 42 are also connected to the bus 52.
[0030] The reception device 38 is equipped with a touch panel 38A and a microphone 38B, etc., and receives user input. The touch panel 38A receives user input by detecting contact with an object (e.g., a pen or finger). The microphone 38B receives user input by detecting the user's voice. The control unit 46A transmits data indicating the user input received by the touch panel 38A and microphone 38B to the data processing device 12. In the data processing device 12, the specific processing unit 290 acquires the data indicating the user input.
[0031] The output device 40 includes a display 40A and a speaker 40B, and presents data to the user 20 by outputting the data in a form perceptible to the user 20 (e.g., audio and / or text). The display 40A displays visible information such as text and images according to instructions from the processor 46. The speaker 40B outputs audio according to instructions from the processor 46. The camera 42 is a small digital camera equipped with an optical system such as a lens, aperture, and shutter, and an image sensor such as a CMOS (Complementary Metal-Oxide-Semiconductor) image sensor or a CCD (Charge Coupled Device) image sensor.
[0032] Communication interface 44 is connected to network 54. Communication interfaces 44 and 26 are responsible for the exchange of various types of information between processor 46 and processor 28 via network 54.
[0033] Figure 2 shows an example of the main functions of the data processing device 12 and the smart device 14.
[0034] As shown in Figure 2, in the data processing device 12, a specific processing is performed by the processor 28. A specific processing program 56 is stored in the storage 32. The specific processing program 56 is an example of a "program" related to the technology of this disclosure. The processor 28 reads the specific processing program 56 from the storage 32 and executes the read specific processing program 56 on the RAM 30. The specific processing is realized by the processor 28 operating as a specific processing unit 290 according to the specific processing program 56 executed on the RAM 30.
[0035] The storage 32 stores the data generation model 58 and the emotion identification model 59. The data generation model 58 and the emotion identification model 59 are used by the identification processing unit 290.
[0036] In the smart device 14, the processor 46 performs the reception output processing. The storage 50 stores the reception output program 60. The reception output program 60 is used in conjunction with a specific processing program 56 by the data processing system 10. The processor 46 reads the reception output program 60 from the storage 50 and executes the read reception output program 60 on the RAM 48. The reception output processing is realized by the processor 46 operating as a control unit 46A according to the reception output program 60 executed on the RAM 48.
[0037] Next, the specific processing performed by the specific processing unit 290 of the data processing device 12 will be described. In the following description, the data processing device 12 will be referred to as the "server" and the smart device 14 as the "terminal".
[0038] Modes for carrying out the invention
[0039] This invention is a system for users to recognize and prevent the risk of wire fraud in advance. This system collects the user's location information and voice data, generates a warning message based on the analysis results, and notifies the user's device and pre-registered contacts of the warning message.
[0040] Program Processing Description
[0041] Collection and transmission of location information
[0042] The device constantly monitors its GPS sensor to detect when the user enters a specific area (e.g., near a bank ATM). When the user enters the designated zone, the device collects location information and periodically sends it to the server. This allows the server to accurately determine the user's location.
[0043] Collection and transmission of audio data
[0044] The device collects voice data in real time from the moment the call begins. The voice data generated when the user makes a call is encrypted as needed and sent to the server. This process requires the user's consent, so it is important that consent is obtained in advance.
[0045] Data analysis using analytical methods
[0046] The server integrates and analyzes the received location information and voice data. It analyzes the location information to determine if the user is in a specific location (e.g., a bank ATM) and uses a generative AI model to analyze the voice data. It detects specific keywords and phrases (e.g., "transfer," "urgent," "money," etc.) and assesses the risk of fraud by matching them against a local crime database.
[0047] Generation and notification of warning messages
[0048] The server generates a warning message based on the analysis results. For example, it might create a message such as, "This call may be a scam. Do not transfer any money." The generated warning message is notified to the user in both voice and text formats. In emergencies, the same warning is also sent to pre-registered contacts (family and friends).
[0049] Specific example
[0050] 1. The user approaches the bank ATM.
[0051] The device detects the user's current location and notifies the server that "the user is approaching a bank ATM."
[0052] The specific location information detected is "Latitude: 35.6895, Longitude: 139.6917 (within Tokyo)".
[0053] 2. Data collection during voice calls
[0054] When a user initiates a call, the device starts recording the call, encrypting it in real time, and sending it to the server.
[0055] Example of a phone call: "Grandma, I need money urgently. Could you transfer it to me today?"
[0056] 3. Server-based analysis
[0057] The server checks the user's location and confirms that the user is near a bank ATM.
[0058] The system analyzes the voice data to detect keywords such as "transfer," "urgent," and "money." Furthermore, it cross-references this data with local crime statistics to determine that the risk of fraud is "high."
[0059] 4. Generating and notifying warning messages
[0060] The server generates a warning message saying, "This call may be a scam. Do not transfer any money under any circumstances," and notifies the device via both voice and text.
[0061] Furthermore, a warning will be sent to pre-registered contacts (family and friends).
[0062] The system of this invention operates in this manner, significantly reducing the risk of users falling victim to wire fraud. Furthermore, by immediately notifying users of warning messages, it enables users to take swift action, effectively preventing damage before it occurs.
[0063] The following describes the processing flow.
[0064] Step 1:
[0065] The device uses a GPS sensor to detect the user's current location. When the user enters a specific area (for example, near a bank ATM), the device collects location information. Subsequently, it sends the collected location information to a server, which then accurately determines the user's current location.
[0066] Step 2:
[0067] When a user initiates a call, the device begins collecting the call audio in real time. This audio data is encrypted to protect privacy. The encrypted audio data is then sent to a server.
[0068] Step 3:
[0069] The server integrates and analyzes the received location information and voice data. It analyzes the location information to confirm that the user is in a specific location (e.g., a bank ATM). It also analyzes the voice data using a generation AI model to detect specific keywords or phrases (e.g., "transfer," "hurry," "money," etc.).
[0070] Step 4:
[0071] The server compares the analysis results against a crime database to assess the risk of fraud. A high risk assessment suggests that the user is at risk of wire fraud.
[0072] Step 5:
[0073] The server generates a warning message based on the risk assessment results. The generated warning message will include appropriate content such as, "This call may be a scam. Do not transfer any money."
[0074] Step 6:
[0075] The device notifies the user of the generated warning message. The notification is provided in both audio and text formats so that the user can immediately review it.
[0076] Step 7:
[0077] Furthermore, in the event of an emergency, the server will also send a warning message to pre-registered contacts (family and friends). This will alert those around the user as well.
[0078] (Example 1)
[0079] Next, we will describe Example 1. In the following description, the data processing device 12 will be referred to as the "server," and the smart device 14 will be referred to as the "terminal."
[0080] In modern times, criminal activities such as wire fraud are on the rise, and many people are falling victim to them. Vulnerable individuals, especially the elderly, are particularly vulnerable, making prevention measures an urgent necessity. Conventional wire fraud prevention systems require users to consciously operate them, lacking convenience and immediacy. To address this challenge, a system is needed that minimizes user intervention and detects and warns of fraud risks in real time.
[0081] The identification process performed by the identification processing unit 290 of the data processing device 12 in Example 1 is realized by the following means.
[0082] In this invention, the server includes location information collection means for collecting user location information, location information transmission means for transmitting the location information to a network device, voice data collection means for collecting user voice data, voice data encryption means for encrypting the voice data, voice data transmission means for transmitting the encrypted voice data to a network device, analysis means for analyzing the location information and voice data in the network device, warning message generation means for generating a warning message based on the analysis results of the analysis means, warning notification means for notifying the generated warning message to a user terminal, notification means for notifying the generated warning message to pre-registered contacts, risk assessment means for detecting specific keywords and evaluating risk using a generation artificial intelligence model, range monitoring means for detecting when a user enters a specific range, adjustment means for adjusting the content of the warning message based on the range, and format conversion means for notifying the adjusted message in voice or text format. This makes it possible to detect the risk of fraud in real time and quickly notify warning messages while minimizing user intervention.
[0083] "Location information collection means" refers to sensors or devices used to determine the user's current location, primarily using GPS systems to acquire location information.
[0084] "Location information transmission means" refers to a function or device for transmitting acquired location information to a network device, and transmits data via a communication network.
[0085] "Voice data collection means" refers to devices such as microphones used to collect the user's voice, and to acquire the content of the call as digital voice data.
[0086] "Voice data encryption means" refers to a function or device that performs encryption processing to protect collected voice data, thereby ensuring data security.
[0087] "Voice data transmission means" refers to a function or device for transmitting encrypted voice data to a network device, and which transmits data via a communication network.
[0088] "Analysis means" refers to a function or device for analyzing received location information and audio data, and integrates and analyzes the location data and audio data.
[0089] A "warning message generation means" is a function or device that generates a message to warn the user based on the analysis results, and creates appropriate warning content.
[0090] A "warning notification means" is a function or device for notifying a generated warning message to a user terminal, thereby immediately conveying the warning to the user.
[0091] A "notification method" is a function or device that notifies pre-registered contacts of the generated warning message, thereby informing family and friends of the warning.
[0092] A "generative artificial intelligence model" is a model that analyzes audio data to detect specific keywords and perform risk assessment, using machine learning and natural language processing technologies.
[0093] A "risk assessment tool" is a function or device that analyzes audio data to assess the risk of fraud, and determines the risk level based on specific keywords or phrases.
[0094] "Range monitoring means" refers to a function or device for detecting when a user enters a specific geographical area (e.g., near a bank ATM), and continuously monitors location information.
[0095] "Adjustment means" refers to a function or device for adjusting the content of a warning message based on the user's location and voice data, thereby creating an appropriate message for the situation.
[0096] "Format conversion means" refers to a function or device for notifying warning messages in audio or text format, and for conveying warnings in a format suitable for the user.
[0097] This invention relates to a system for users to recognize and prevent the risk of wire fraud in advance. This system collects and analyzes the user's location information and voice data, and generates and sends a warning message based on the results.
[0098] Collection and transmission of location information
[0099] The terminal constantly monitors its GPS sensor to detect when the user enters a specific area (for example, near a bank ATM). This area monitoring means that when the user enters a designated zone, the terminal collects location information and periodically transmits it to a server via a location information transmission means. The GPS sensor is used in this process.
[0100] Specifically, the device collects location information such as "latitude: 35.6895, longitude: 139.6917 (within Tokyo)" and sends it to the server.
[0101] Collection and transmission of audio data
[0102] The device collects voice data in real time from the moment a call begins using voice data collection means. Voice data generated when a user makes a call is encrypted as needed using voice data encryption means and transmitted to the server via voice data transmission means. This process requires the user's consent.
[0103] Specifically, the device records the conversation as soon as the call starts and sends the audio data to the server while protecting it with AES encryption technology. An example of the conversation content collected might be something like, "Grandma, I need money right away. Can you transfer it to me today?"
[0104] Data Analysis
[0105] The server integrates and analyzes the received location information and encrypted voice data using analysis tools. First, it analyzes the location information to determine if the user is in a specific location (e.g., near a bank ATM). Next, it analyzes the voice data using a generative artificial intelligence model to detect specific keywords and phrases (e.g., "transfer," "urgent," "money," etc.), and assesses the risk of fraud using risk assessment tools. Natural language processing techniques are used in this analysis.
[0106] As an example, the prompt statements for a generative AI model are as follows:
[0107] Analyze the fraud risk from this audio data. Focus on detecting the following keywords and phrases: "transfer," "urgent," and "money." Evaluate the results as a "low," "medium," or "high" risk level to determine the risk level.
[0108] Audio data: ...
[0109] Generation and notification of warning messages
[0110] The server generates a warning message based on the analysis results using a warning message generation system. For example, it might create a message such as, "This call may be a scam. Do not transfer any money." The generated warning message is then notified to the user in both voice and text format using a warning notification system. Depending on the settings, the same warning is also sent to pre-registered contacts (family and friends) via the notification system.
[0111] Specifically, the server generates a warning message stating "This is likely a scam," converts it into voice or text format, and notifies the user and their family.
[0112] This system significantly reduces the risk of users falling victim to wire fraud. Furthermore, by immediately notifying users of warning messages, it enables quick responses, effectively preventing damage before it occurs.
[0113] The flow of the specific processing in Example 1 will be explained using Figure 11.
[0114] Step 1: Collect and transmit location information
[0115] The device constantly monitors the user's location. Specifically, it uses the GPS sensor within the device to obtain the user's current location. When the user enters a specific area (e.g., near a bank ATM), this information is automatically detected. The input is location information from the GPS sensor (e.g., latitude and longitude), and the output is the current user location information. This location information is transmitted to the server using a location information transmission device.
[0116] Step 2: Collect and transmit audio data
[0117] When a user initiates a call, the terminal detects the start of the call. The terminal uses its built-in microphone to collect audio data. The collected audio data is encrypted using an audio data encryption method. The input is the audio data of the call, and the output is encrypted audio data. Subsequently, the encrypted audio data is sent to the server using an audio data transmission method.
[0118] Step 3: Analysis of location and audio data
[0119] The server integrates and analyzes the received location information and voice data. First, it analyzes the location information to determine if the user is in a specific location (e.g., near a bank ATM). Next, it analyzes the voice data using a generative artificial intelligence model. The input consists of decomposed location information and encrypted voice data, and a risk assessment is output as a result of the analysis process. Prompt sentences are input to the generative artificial intelligence model to detect specific keywords (e.g., "transfer", "urgent", "money").
[0120] Step 4: Generating a warning message
[0121] The server generates a warning message based on the analysis results. The AI generation model automatically generates specific warning content tailored to the risk assessment from its output. The input is the result of voice data analysis, and the output is a concrete warning message. For example, a message such as "This call may be a scam. Do not transfer any money under any circumstances" might be generated.
[0122] Step 5: Notification of warning message
[0123] The server notifies the user terminal of the generated warning message. The warning message is sent in both voice and text formats using the warning notification method. The input is the generated warning message, and the output is the notification sent to the user. In addition, the same warning is sent to pre-registered contacts (family and friends). The terminal displays the received warning message to the user and provides voice notifications depending on the content of the warning.
[0124] This entire process makes it possible to detect the risk of wire fraud in real time and to quickly and appropriately notify users of warning messages.
[0125] (Application Example 1)
[0126] Next, we will explain Application Example 1. In the following explanation, the data processing device 12 will be referred to as the "server," and the smart device 14 will be referred to as the "terminal."
[0127] In recent years, there has been a significant increase in wire fraud, particularly targeting the elderly. These scams cleverly exploit the victim's psychology and typically utilize voice communication methods such as telephone calls. Traditional prevention measures, relying solely on vigilance, have been insufficient for effective prevention. Furthermore, there has been a lack of mechanisms to recognize fraud risks in real time and take immediate action. Against this backdrop, there is a need for technology that enables users to detect fraud risks in advance and prevent becoming a victim.
[0128] The specific processing performed by the specific processing unit 290 of the data processing device 12 in Application Example 1 is realized by the following means.
[0129] In this invention, the server includes location information collection means for collecting user location information, voice data collection means for collecting user voice data, and analysis processing means for analyzing using a generative AI model. This makes it possible to analyze both voice data and location information when a user approaches a specific risk area (e.g., near a bank ATM) and immediately recognize and warn of the risk of fraud.
[0130] "Location information collection means" refers to means for obtaining the user's current geographical location.
[0131] "Location information transmission means" refers to a means for transmitting collected location information to a server.
[0132] "Voice data collection means" refers to means for collecting a user's voice communication data.
[0133] "Voice data encryption means" refers to a means for encrypting collected voice data.
[0134] "Voice data transmission means" refers to a means for sending encrypted voice data to a server.
[0135] "Analysis means" refers to methods for evaluating fraud risk by analyzing location information and voice data.
[0136] A "generative AI model" is an artificial intelligence model that learns from large amounts of data and is used to perform specific tasks (e.g., fraud detection).
[0137] The "analysis processing means" is a means for comprehensively analyzing location information and audio data using a generated AI model.
[0138] A "warning message generation means" is a means for creating a warning message for the user based on the analysis results.
[0139] A "warning notification means" is a means for notifying the user terminal of the generated warning message.
[0140] "Notification method" refers to a means of notifying pre-registered contacts of the generated warning message.
[0141] "Voice notification means" refers to a means of notifying the user of the generated warning message in voice format.
[0142] A "text notification method" is a means of notifying the user of the generated warning message in text format.
[0143] Regarding embodiments for carrying out the invention, this system employs a computer-based method that combines multiple means for users to recognize and prevent the risk of wire fraud in advance.
[0144] composition
[0145] This system includes a user terminal, location information collection means, voice data collection means, generation AI model, analysis processing means, warning message generation means, warning notification means, and notification means.
[0146] User terminal: This refers to a device such as a smartphone carried by the user. The terminal is equipped with a GPS sensor and a microphone, and collects location information and voice data.
[0147] Location information collection method: Collects the user's current location. It is possible to monitor location information in real time using GPS.
[0148] Voice data collection method: Voice data during calls is collected in real time. It is collected using a microphone and encrypted for privacy protection.
[0149] Generative AI Models: These models use pre-trained AI models to detect fraudulent activity. They assess the risk of fraud based on a large amount of historical data.
[0150] Analysis processing method: Collected location information and audio data are analyzed using a generating AI model. This allows for the assessment of fraud risk.
[0151] Warning message generation means: Generates a warning message based on the analysis results. Messages are generated in both audio and text formats.
[0152] Warning notification method: The generated warning message is notified to the user's terminal. Specifically, the message is displayed on the smartphone screen and an audible warning is issued.
[0153] Notification method: In an emergency, the same warning message will be sent to pre-registered contacts (family and friends).
[0154] How to use
[0155] This system works as follows:
[0156] 1. When a user approaches a specific area:
[0157] When a user approaches a specific area, such as a bank ATM, the GPS sensor on the user's device collects location information and sends it to a server. The server analyzes this information to detect that the user is near a bank ATM.
[0158] 2. When the user initiates a call:
[0159] When a user initiates a call, the call audio is collected by a voice data collection system, encrypted in real time, and then transmitted to the server. During this process, the voice data is appropriately encrypted to protect the user's privacy.
[0160] 3. Analysis of fraud risk:
[0161] The server integrates collected location and voice data and performs analysis using a generative AI model. When specific keywords or phrases are detected, they are compared against a local crime database to assess the fraud risk.
[0162] Specific example
[0163] Example of a user approaching a bank ATM:
[0164] The terminal notifies the server that "the user is approaching a bank ATM," and specific location information (e.g., latitude: 35.6895, longitude: 139.6917) is detected.
[0165] Examples of data collection during a call:
[0166] When a user calls and says, "Grandma, I need money right away. Can you transfer it to me today?", the device records the audio, encrypts it, and sends it to the server.
[0167] Server-based analysis example:
[0168] The generative AI model detects keywords such as "bank transfer," "urgent," and "money," and compares them with local crime data to determine that the fraud risk is "high."
[0169] Example of a prompt
[0170] Examples of prompt statements are as follows:
[0171] Audio recording: "Grandma, I need money urgently. Can you transfer it to me today?"
[0172] Please evaluate the risk of fraud as part of the analysis. If you find any specific keywords or phrases, return their probability.
[0173] In this way, the system can quickly detect the risk of fraud and prevent damage by warning the user.
[0174] The flow of a specific process in Application Example 1 will be explained using Figure 12.
[0175] Step 1:
[0176] When a user approaches a specific area (e.g., near a bank ATM), the device's GPS sensor collects location information. This information is recorded on the device as numerical data such as "latitude: 35.6895, longitude: 139.6917". The input is the current location information, and the output is the collected location information. The device sends this information to the server.
[0177] Step 2:
[0178] The server analyzes the received location information to determine if the user is in a specific risk area. This analysis confirms that the user is near a bank ATM. The input is location information, and the output is the risk area determination result.
[0179] Step 3:
[0180] When a user initiates a call, the device collects the call audio in real time. The audio data is captured via the microphone and encrypted as needed. The input is the call audio, and the output is encrypted audio data. The device then sends this audio data to the server.
[0181] Step 4:
[0182] The server decrypts encrypted audio data and extracts keywords in order to analyze the received audio data using a generative AI model. The input is encrypted audio data, and the output is string data for analysis. Based on this string data, the generative AI model detects specific keywords (e.g., "transfer", "urgent", "money").
[0183] Step 5:
[0184] Based on the analyzed results, the server cross-references them with a local crime database to assess the fraud risk. The input consists of analyzed audio data, location information, and the crime database; the output is the fraud risk assessment result. If the server determines the risk is high, it generates a warning message.
[0185] Step 6:
[0186] The generated warning message is created by the server in both audio and text formats and immediately notified to the user. The input is the risk assessment results and a warning message template, and the output is the generated warning message. The terminal then informs the user of this message via audio and text.
[0187] Step 7:
[0188] The server also sends warning messages to pre-registered contacts (e.g., family and friends). The input is the generated warning message, and the output is the notification to the registered contacts. This helps users take quick action.
[0189] Through the above processing steps, this system can issue appropriate warnings to users before they are exposed to the risk of fraud, thereby preventing them from becoming victims.
[0190] Furthermore, an emotion engine that estimates the user's emotions may be incorporated. That is, the identification processing unit 290 may use the emotion identification model 59 to estimate the user's emotions and perform identification processing using the user's emotions.
[0191] Modes for carrying out the invention
[0192] This invention is a system for users to recognize and prevent the risk of wire fraud in advance. This system collects the user's location information and voice data, and generates a warning message based on the analysis results. It also utilizes an emotion engine to recognize the user's emotions, thereby improving the accuracy and effectiveness of the warning. The warning message is then sent to the user's device and pre-registered contacts.
[0193] Program Processing Description
[0194] Collection and transmission of location information
[0195] The device uses a GPS sensor to detect the user's current location. When the user enters a specific area (for example, near a bank ATM), the device collects location information and periodically sends it to the server. This allows the server to accurately determine the user's location.
[0196] Collection and transmission of audio data
[0197] When a user initiates a call, the device begins collecting the call audio in real time. This audio data is encrypted to protect privacy. The encrypted audio data is then sent to a server.
[0198] Emotional analysis using an emotion engine
[0199] The terminal or server uses the received audio data to evaluate the user's emotional state. The emotion engine recognizes the user's emotions (e.g., tension, fear, anxiety) from factors such as the tone, speed, and intonation of the voice.
[0200] Data analysis using analytical methods
[0201] The server integrates and analyzes the received location information and voice data. It analyzes the location information to confirm that the user is in a specific location (e.g., a bank ATM). It also uses a generative AI model to analyze the voice data and detect specific keywords or phrases (e.g., "transfer," "hurry," "money," etc.).
[0202] Risk assessment that takes emotional state into account
[0203] The server takes into account the results of the emotion engine analysis and compares them with the crime database to ultimately assess the risk of fraud. If the user's emotions are in a high-risk state such as tension, fear, or anxiety, the server determines that the likelihood of fraud is higher.
[0204] Generation and notification of warning messages
[0205] The server generates a warning message based on the risk assessment results. For example, it might say, "This call may be a scam. Do not transfer any money." The generated warning message is sent to the device in both voice and text formats. In emergencies, the warning message is also sent to pre-registered contacts (family and friends).
[0206] Specific example
[0207] 1. The user approaches the bank ATM.
[0208] The device detects the user's current location and notifies the server that "the user is approaching a bank ATM."
[0209] The specific location information detected is "Latitude: 35.6895, Longitude: 139.6917 (within Tokyo)".
[0210] 2. Data collection during voice calls
[0211] When a user initiates a call, the device starts recording the call, encrypting it in real time, and sending it to the server.
[0212] Example of a phone call: "Grandma, I need money urgently. Could you transfer it to me today?"
[0213] 3. Analysis using an emotion engine
[0214] The server analyzes the user's emotions from the voice data and detects tension and fear.
[0215] 4. Server-based data analysis
[0216] The server checks the user's location and confirms that the user is near a bank ATM.
[0217] The system analyzes voice data using an AI model to detect keywords such as "transfer," "urgent," and "money." Furthermore, it compares this data with local crime data to determine a high risk of fraud.
[0218] 5. Generating and notifying warning messages
[0219] The server generates a warning message stating, "This call may be a scam. Do not transfer any money under any circumstances," and notifies the device via both voice and text. Furthermore, warnings are also sent to pre-registered contacts (family and friends).
[0220] The system of this invention operates in this manner, significantly reducing the risk of users falling victim to wire fraud. Furthermore, by immediately notifying users of warning messages, it enables users to take swift action, effectively preventing damage before it occurs.
[0221] The following describes the processing flow.
[0222] Step 1:
[0223] The device uses a GPS sensor to detect the user's current location. When the user enters a specific area (for example, near a bank ATM), the device collects location information and periodically sends it to the server. This allows the server to accurately determine the user's location.
[0224] Step 2:
[0225] When a user initiates a call, the device begins collecting the call audio in real time. This audio data is encrypted to protect privacy. The encrypted audio data is then sent to a server.
[0226] Step 3:
[0227] The server integrates and analyzes the received location information and voice data. It analyzes the location information to confirm that the user is in a specific location (e.g., a bank ATM). It also uses a generative AI model to analyze the voice data and detect specific keywords or phrases (e.g., "transfer," "hurry," "money," etc.).
[0228] Step 4:
[0229] The terminal or server uses the received audio data to evaluate the user's emotional state. The emotion engine recognizes the user's emotions (e.g., tension, fear, anxiety) from factors such as tone, speed, and intonation of the voice. If the emotional state is determined to be tension, fear, or anxiety, special attention is paid to the analysis results.
[0230] Step 5:
[0231] The server takes into account the results of the emotion engine analysis and compares them with the crime database to ultimately assess the risk of fraud. If the user's emotions are in a high-risk state such as tension, fear, or anxiety, the server determines that the likelihood of fraud is higher. Based on this, the accuracy and reliability of the risk assessment are improved.
[0232] Step 6:
[0233] The server generates a warning message based on the risk assessment results. For example, it might say, "This call may be a scam. Do not transfer any money." The generated warning message is notified to the user in both audio and text format.
[0234] Step 7:
[0235] The device notifies the user of the generated warning message in both voice and text format. It also notifies the user to immediately check the warning message, encouraging a quick response. Furthermore, in emergencies, the warning message is also sent to pre-registered contacts (family and friends). This alerts those around the user as well.
[0236] Specific example:
[0237] 1. The user approaches the bank ATM.
[0238] The device detects the user's current location and notifies the server that "the user is approaching a bank ATM."
[0239] The location information detected is "Latitude: 35.6895, Longitude: 139.6917 (within Tokyo)".
[0240] 2. Data collection during voice calls
[0241] When a user initiates a call, the device begins recording the conversation. The recorded data is encrypted and sent to a server in real time.
[0242] Example of phone call content: "Grandma, I need money urgently. Could you transfer it to me today?"
[0243] 3. Data analysis by the server
[0244] The server analyzes the location information to confirm that the user is near a bank ATM.
[0245] The AI model analyzes the voice data to detect keywords such as "transfer," "urgent," and "money."
[0246] 4. Emotional analysis using an emotion engine
[0247] The server detects the user's emotions from the voice data and recognizes signs of tension or fear.
[0248] 5. Risk Assessment
[0249] The server also considers the results of the emotion engine and assesses the fraud risk as high based on the analysis results.
[0250] 6. Generating and notifying warning messages
[0251] The server generates a warning message stating, "This call may be a scam. Do not transfer any money under any circumstances," and sends it to the device in both voice and text formats.
[0252] 7. Notification to the user
[0253] The device notifies the user of warning messages in both voice and text formats, and in emergencies, it also sends warnings to pre-registered contacts.
[0254] In this way, the system of the present invention significantly reduces the risk of users falling victim to wire fraud and enables rapid countermeasures. Furthermore, by taking into account the user's emotional state, the accuracy and effectiveness of the warning are further improved.
[0255] (Example 2)
[0256] Next, we will describe Example 2. In the following description, the data processing device 12 will be referred to as the "server" and the smart device 14 as the "terminal".
[0257] Currently, the number of victims of wire fraud is increasing, raising the risk of users being deceived and losing their valuable assets. Traditional countermeasures are insufficient because they lack the means for users to recognize the risk of fraud in advance, making it difficult to prevent victims. To solve this problem, there is a need for a system that can accurately detect the risk of wire fraud and provide users with rapid warnings.
[0258] The identification process performed by the identification processing unit 290 of the data processing device 12 in Example 2 is realized by the following means.
[0259] In this invention, the server includes emotion analysis means for analyzing the user's emotional state, analysis means for integrating and analyzing location information and voice data and detecting specific keywords or phrases using a generative AI model, and means for performing risk assessment. This enables users to detect the risk of wire fraud in advance with high accuracy and receive a warning quickly.
[0260] "Location information collection means" refers to means for obtaining the user's current location information.
[0261] "Location information transmission means" refers to a means for transmitting acquired location information to a server.
[0262] "Voice data collection means" refers to a means for collecting user voice data.
[0263] "Voice data encryption means" refers to a method of encrypting collected voice data for privacy protection.
[0264] "Voice data transmission means" refers to a means for sending encrypted voice data to a server.
[0265] "Emotion analysis means" refers to a method for evaluating a user's emotional state from voice data.
[0266] "Analysis means" refers to a means for integrating and analyzing location information and audio data, and for using a generative AI model to detect specific keywords or phrases.
[0267] A "risk assessment tool" is a means for evaluating risk based on the analysis results of an analysis tool.
[0268] A "warning message generation means" is a means for generating a warning message based on the results of a risk assessment.
[0269] A "warning notification means" is a means for notifying the user terminal of the generated warning message.
[0270] "Notification method" refers to a means of notifying pre-registered contacts of the generated warning message.
[0271] "Voice notification means" refers to a means of notifying the user of the generated warning message in voice format.
[0272] A "text notification method" is a means of notifying the user of the generated warning message in text format.
[0273] A "generative AI model" is an artificial intelligence model used to analyze audio data and detect specific keywords or phrases.
[0274] This invention is a system for users to recognize and prevent the risk of wire fraud in advance. The system includes means for collecting the user's location information and voice data and generating a warning message based on the analysis results.
[0275] Hardware Configuration
[0276] The terminal uses a mobile terminal such as a smartphone equipped with a GPS sensor, a microphone, and a data communication function. The GPS sensor collects the user's current location information, and the microphone is used to collect the call voice. Then, the collected data is sent to the server through the Internet.
[0277] The server uses a high-performance computer or cloud service and has the processing ability for data analysis and sentiment analysis. As specific software, the sentiment analysis API of IBM Watson (registered trademark) and the generative AI model of GPT-3 (registered trademark) are used.
[0278] Software Configuration
[0279] The terminal has the following software functions:
[0280] 1. Location information collection means: The built-in GPS sensor is used to obtain the user's current location and generate location data.
[0281] 2. Voice data collection means: The microphone collects voice data in real time during a call.
[0282] 3. Voice data encryption means: The collected voice data is encrypted with AES-256.
[0283] 4. Location information transmission means: The collected location information is sent to the server.
[0284] 5. Voice data transmission means: The encrypted voice data is sent to the server.
[0285] The server has the following software functions:
[0286] 1. Sentiment analysis means: The received voice data is analyzed with the sentiment analysis API of IBM Watson to evaluate the user's sentiment state.
[0287] 2. Analysis means: Integrate the received location information and voice data, and use the GPT-3 generative AI model to detect specific keywords and phrases.
[0288] 3. Risk assessment means: Based on the analysis results and sentiment analysis results, finally evaluate the risk of remittance fraud.
[0289] 4. Warning message generation means: Generate a warning message based on the risk assessment result.
[0290] 5. Warning notification means: Notify the generated warning message to the user terminal and pre-registered contacts.
[0291] Specific example
[0292] 1. The user approaches a bank ATM
[0293] When the terminal uses the GPS sensor to obtain the current location information and detects that "the user has approached the bank ATM", it transmits the location information to the server. The location data example includes "Latitude: 35.6895, Longitude: 139.6917 (in Tokyo)".
[0294] 2. Data collection during a voice call
[0295] When the user starts a call, the terminal collects the call voice, encrypts it with AES-256, and transmits it to the server in real time. Specifically, the call content is "Grandma, I need money right away. Can you transfer it to me today?". The process of voice recording and encryption is executed on the terminal.
[0296] 3. Analysis by the sentiment engine
[0297] The server decrypts the received voice data and evaluates the user's sentiment as "nervous" using the sentiment analysis API. Tone, speed, and intonation of the voice are considered in the analysis.
[0298] 4. Data Analysis by Server
[0299] The server confirms from the location information that the user is near a bank ATM. At the same time, it analyzes the voice data with a generative AI model (e.g., GPT-3) to detect keywords such as "transfer", "hurry", and "money". As a result, the conversion process from voice to text is performed, and specific keywords are emphasized.
[0300] 5. Risk Assessment Considering Emotional State
[0301] The server comprehensively analyzes the voice analysis results and emotional analysis results and determines that the fraud risk is high. This includes comparison with past crime data and calculation of the coincidence rate.
[0302] 6. Generation and Notification of Warning Messages
[0303] The server generates a warning message saying, "This call content may be fraudulent. Do not make any transfers." and also generates a voice message using the text-to-speech API. The generated warning message is sent to the terminal and also notified to the registered contacts via SMS or email as necessary.
[0304] Example of Prompt Sentence
[0305] "Please teach me the processing procedure of the program that notifies that the user has approached the bank ATM. Also, please teach me the procedure for collecting the voice data of the user during the call and evaluating the fraud risk."
[0306] The flow of specific processing in Example 2 will be described using FIG. 13.
[0307] Step 1: Collection and Transmission of Location Information
[0308] The device uses its built-in GPS sensor to acquire the user's current location at regular intervals. Specifically, it acquires location information every minute and generates latitude and longitude data (input: GPS sensor data, output: location information data). When the user enters a specific area (for example, near a bank ATM), it sends that location information to the server (data processing: location information detection, data calculation: range determination, output: location information transmission).
[0309] Step 2: Collect and transmit audio data
[0310] When a user initiates a call, the device uses its microphone to collect the call audio in real time (input: call audio, output: audio data). The collected audio data is immediately encrypted using AES-256 (data processing: audio data encryption), and the encrypted audio data is sent to the server in real time (output: encrypted audio data transmission).
[0311] Step 3: Emotional analysis using the emotion engine
[0312] The server decrypts the received encrypted audio data (input: encrypted audio data, output: decrypted audio data) and sends the audio data to the emotion analysis API. The emotion analysis API analyzes the tone, speed, and intonation of the voice and evaluates the user's emotional state (tension, fear, anxiety, etc.) (data calculation: emotion analysis, output: emotional state data).
[0313] Step 4: Data analysis using analytical tools
[0314] The server integrates and analyzes emotional state data with received location and audio data (input: location data, decoded audio data, emotional state data). It uses a GPT-3 generative AI model to detect specific keywords or phrases (e.g., "transfer", "hurry", "money") from the audio data (data calculation: keyword detection, output: keyword data). It uses location information to determine if the user is in a specific location (e.g., a bank ATM) (data calculation: location confirmation, output: location confirmation data).
[0315] Step 5: Risk assessment considering emotional state
[0316] The server integrates keyword data and location data from the analysis tools, and also considers emotional state data to assess fraud risk (Input: Keyword data, emotional state data, location data; Output: Risk assessment data). For example, if the user's emotions are in a high-risk state such as tension or fear, and specific keywords are included, the risk is judged to be high (Data calculation: Risk assessment).
[0317] Step 6: Generate and send warning messages
[0318] The server generates a warning message based on risk assessment data (input: risk assessment data, output: warning message). For example, it generates a message stating, "This call may be a scam. Do not transfer any money under any circumstances." (data processing: warning message generation). The generated warning message is converted into voice format using a speech synthesis API (output: voice warning message) and also sent to the terminal in text format (output: text warning message). Furthermore, in emergencies, notifications are also sent to pre-registered contacts (family and friends) (output: contact notification data).
[0319] (Application Example 2)
[0320] Next, we will explain application example 2. In the following explanation, the data processing device 12 will be referred to as a "server" and the smart device 14 as a "terminal".
[0321] Fraudulent activities, including wire transfer scams, are increasing, particularly targeting the elderly. Therefore, there is a need to detect the risk of fraud early and prevent it before it occurs. However, current systems lack the means to effectively analyze users' emotional states and real-time call content, making it difficult to accurately determine situations with a high risk of fraud. Therefore, it is necessary to provide a system that can quickly and accurately detect situations with a high risk of fraud and issue warnings to users, their relatives, and friends.
[0322] The specific processing performed by the specific processing unit 290 of the data processing device 12 in Application Example 2 is realized by the following means.
[0323] In this invention, the server includes location information collection means for collecting user location information, voice data collection means for collecting user voice data, emotion analysis means for analyzing the user's emotional state based on the voice data, and keyword detection means for detecting specific keywords. This makes it possible to detect fraud risks such as wire fraud with high accuracy and speed, and to issue warnings to the user and their family and friends.
[0324] "Location information collection means" refers to devices and software used to determine the user's current location.
[0325] "Location information transmission means" refers to devices or software used to transmit collected user location information to a server.
[0326] "Voice data collection means" refers to devices or software used to collect voices emitted by users.
[0327] "Voice data encryption means" refers to devices or software used to encrypt collected voice data so that it cannot be deciphered by third parties.
[0328] "Voice data transmission means" refers to devices or software used to transmit encrypted voice data to a server.
[0329] "Analysis means" refers to devices or software used to analyze location information and voice data collected on a server to assess fraud risk.
[0330] "Warning message generation means" refers to a device or software that automatically generates warning messages based on the analysis results of the analysis means.
[0331] "Warning notification means" refers to devices or software that notify the user's terminal of the generated warning message.
[0332] "Notification means" refers to devices or software used to notify pre-registered contacts of generated warning messages.
[0333] "Emotional analysis means" refers to devices and software that analyze a user's emotional state in real time based on voice data.
[0334] "Keyword detection means" refers to devices or software used to detect specific keywords from audio data.
[0335] This invention is a system for users to recognize and prevent the risk of wire fraud in advance. The system includes means for identifying the user's current location and transmitting that information to a server, and means for collecting, encrypting, and transmitting the user's voice data. Furthermore, the server includes means for analyzing this information and generating a warning message to notify the user's terminal and pre-registered contacts.
[0336] Collection and transmission of location information
[0337] The device is equipped with a GPS sensor that constantly monitors the user's current location and acquires data using location information collection methods. When the user enters a specific area (for example, near a bank ATM), location information is sent to the server using location information transmission methods. This allows the server to determine the user's precise location.
[0338] Collection and transmission of audio data
[0339] When a user starts a phone call, the call content is collected in real time by an audio data collection means and encrypted by an audio data encryption means. The encrypted audio data is then transmitted to the server using an audio data transmission means.
[0340] Sentiment analysis and keyword detection
[0341] The server uses emotion analysis tools to analyze the received audio data. This evaluates the user's emotions (e.g., tension, fear, anxiety) from factors such as tone, speed, and intonation. Furthermore, keyword detection tools detect specific keywords within the audio data (e.g., "transfer," "hurry," "money," etc.).
[0342] Risk assessment and warning message generation
[0343] The server integrates location information, voice data, sentiment analysis results, and keyword detection results to assess the fraud risk. If the risk is determined to be high, a warning message is generated by the warning message generation mechanism. For example, a message such as, "This call may be a scam. Do not transfer any money."
[0344] Warning message notification
[0345] The warning notification system will send the generated warning message to the user's device in both audio and text format. In emergencies, the same warning message will also be sent to pre-registered contacts (family and friends) using the notification system.
[0346] Specific example
[0347] 1. The user approaches the bank ATM.
[0348] The device detects the user's current location and notifies the server that "the user is approaching a bank ATM." The specific location information detected is "Latitude: 35.6895, Longitude: 139.6917 (within Tokyo)."
[0349] 2. Data collection during voice calls
[0350] When a user initiates a call, the device starts recording the conversation, encrypting it in real time, and sending it to the server. Example of a call: "Grandma, I need money urgently. Can you transfer it to me today?"
[0351] 3. Analysis using an emotion engine
[0352] The server analyzes the user's emotions from the voice data and detects tension and fear.
[0353] 4. Server-based data analysis
[0354] The server checks the user's location to confirm they are near a bank ATM. An AI model analyzes the voice data to detect keywords such as "transfer," "urgent," and "money." Furthermore, it cross-references this data with local crime statistics to determine a high risk of fraud.
[0355] 5. Generating and notifying warning messages
[0356] The server generates a warning message stating, "This call may be a scam. Do not transfer any money under any circumstances," and notifies the device via both voice and text. Furthermore, warnings are also sent to pre-registered contacts (family and friends).
[0357] Example of a prompt
[0358] text
[0359] Evaluate the user's emotional state and generate an emergency fraud warning message. For example, use this if you receive audio data such as "Location: 35.6895, 139.6917" and "Grandma, I need money right away. Can you transfer it to me today?". If the emotional score is -0.5 or lower and contains fraud-related keywords such as "transfer" and "money", generate a warning message stating, "This call may be a scam. Do not transfer any money."
[0360] The above describes the embodiments for carrying out this invention.
[0361] The flow of a specific process in Application Example 2 will be explained using Figure 14.
[0362] Step 1:
[0363] Collecting and transmitting user location information
[0364] The device uses a GPS sensor to collect the user's location information. The collected location information (e.g., latitude 35.6895, longitude 139.6917) is transmitted to the server using a location information transmission means.
[0365] Input: Location information from GPS sensor
[0366] Output: User location information sent to the server
[0367] Step 2:
[0368] Collection and transmission of audio data
[0369] When a user initiates a call, the device collects audio in real time using an audio data collection device. The collected audio data is encrypted using an audio data encryption device and transmitted to a server using an audio data transmission device.
[0370] Input: Audio data from a call
[0371] Output: Encrypted audio data sent to the server
[0372] Step 3:
[0373] Location information analysis
[0374] The server analyzes the received location information using an analysis tool to determine whether the user is in a specific risk area (e.g., near a bank ATM).
[0375] Input: Sent location information
[0376] Output: Determination result of whether the user is in a risk area.
[0377] Step 4:
[0378] Sentiment analysis of voice data
[0379] The server analyzes the audio data using emotion analysis tools. Specifically, it evaluates the user's emotions (e.g., tension, fear, anxiety) based on factors such as tone, speed, and intonation of the voice.
[0380] Input: Decrypted audio data
[0381] Output: User sentiment score
[0382] Step 5:
[0383] Detection of specific keywords
[0384] The server uses keyword detection means to detect specific keywords (e.g., "transfer", "hurry", "money") from the audio data.
[0385] Input: Decrypted audio data
[0386] Output: List of detected keywords
[0387] Step 6:
[0388] Risk assessment
[0389] The server integrates location data analysis results, user sentiment scores, and detected keyword lists to assess fraud risk. If a high risk is determined, a warning message is generated using a warning message generation mechanism.
[0390] Input: Location analysis results, sentiment score, keyword list
[0391] Output: Fraud risk assessment results
[0392] Step 7:
[0393] Generation and notification of warning messages
[0394] The server sends the generated warning message to the user's terminal using a warning notification system. Simultaneously, it also sends the same warning message to pre-registered contacts using the notification system.
[0395] Input: Fraud risk assessment results
[0396] Output: Warning message sent to the user's device and registered contacts.
[0397] These steps enable the creation of a system that effectively detects fraud risks and warns users and their associates.
[0398] The specific processing unit 290 transmits the result of the specific processing to the smart device 14. In the smart device 14, the control unit 46A causes the output device 40 to output the result of the specific processing. The microphone 38B acquires audio indicating user input for the result of the specific processing. The control unit 46A transmits the audio data indicating user input acquired by the microphone 38B to the data processing device 12. In the data processing device 12, the specific processing unit 290 acquires the audio data.
[0399] Data generation model 58 is a so-called generative AI (Artificial Intelligence). An example of data generation model 58 is ChatGPT (registered trademark) (Internet search).<URL: https: / / openai.com / blog / chatgpt> ), Gemini (registered trademark) (Internet search) <url: https: gemini.google.com ?hl="ja">Examples of generative AI include the following. The data generation model 58 is obtained by performing deep learning on a neural network. The data generation model 58 is input with prompts containing instructions, and with inference data such as audio data representing speech, text data representing text, and image data representing images. The data generation model 58 infers from the input inference data according to the instructions indicated by the prompts, and outputs the inference results in data formats such as audio data and text data. Here, inference refers to, for example, analysis, classification, prediction, and / or summarization.
[0400] In the above embodiment, an example was given in which specific processing is performed by the data processing device 12, but the technology of this disclosure is not limited thereto, and the specific processing may also be performed by the smart device 14.
[0401] [Second Embodiment]
[0402] Figure 3 shows an example of the configuration of the data processing system 210 according to the second embodiment.
[0403] As shown in Figure 3, the data processing system 210 includes a data processing device 12 and smart glasses 214. An example of the data processing device 12 is a server.
[0404] The data processing device 12 comprises a computer 22, a database 24, and a communication interface 26. The computer 22 is an example of a "computer" related to the technology of this disclosure. The computer 22 comprises a processor 28, RAM 30, and storage 32. The processor 28, RAM 30, and storage 32 are connected to a bus 34. The database 24 and the communication interface 26 are also connected to the bus 34. The communication interface 26 is connected to a network 54. An example of the network 54 is a WAN (Wide Area Network) and / or a LAN (Local Area Network).
[0405] The smart glasses 214 include a computer 36, a microphone 238, a speaker 240, a camera 42, and a communication interface 44. The computer 36 includes a processor 46, RAM 48, and storage 50. The processor 46, RAM 48, and storage 50 are connected to a bus 52. The microphone 238, speaker 240, and camera 42 are also connected to the bus 52.
[0406] The microphone 238 receives voice signals from the user 20 and receives instructions from the user 20. The microphone 238 captures the voice signals from the user 20, converts the captured voice into audio data, and outputs it to the processor 46. The speaker 240 outputs audio according to the instructions from the processor 46.
[0407] Camera 42 is a small digital camera equipped with an optical system including a lens, aperture, and shutter, and an image sensor such as a CMOS (Complementary Metal-Oxide-Semiconductor) image sensor or a CCD (Charge Coupled Device) image sensor, and captures images of the area around the user 20 (for example, an imaging range defined by a field of view equivalent to the width of a typical healthy person's field of vision).
[0408] Communication interface 44 is connected to network 54. Communication interfaces 44 and 26 are responsible for the exchange of various information between processor 46 and processor 28 via network 54. The exchange of various information between processor 46 and processor 28 using communication interfaces 44 and 26 is performed in a secure manner.
[0409] Figure 4 shows an example of the main functions of the data processing device 12 and the smart glasses 214. As shown in Figure 4, the data processing device 12 performs specific processing using the processor 28. The storage 32 stores the specific processing program 56.
[0410] The specific processing program 56 is an example of a "program" relating to the technology of this disclosure. The processor 28 reads the specific processing program 56 from the storage 32 and executes the read specific processing program 56 on the RAM 30. The specific processing is realized by the processor 28 operating as a specific processing unit 290 in accordance with the specific processing program 56 executed on the RAM 30.
[0411] The storage 32 stores the data generation model 58 and the emotion identification model 59. The data generation model 58 and the emotion identification model 59 are used by the identification processing unit 290.
[0412] In the smart glasses 214, the processor 46 performs the reception output processing. The storage 50 stores the reception output program 60. The processor 46 reads the reception output program 60 from the storage 50 and executes the read reception output program 60 on the RAM 48. The reception output processing is realized by the processor 46 operating as a control unit 46A according to the reception output program 60 executed on the RAM 48.
[0413] Next, the identification processing performed by the identification processing unit 290 of the data processing device 12 will be described. In the following description, the data processing device 12 will be referred to as the "server" and the smart glasses 214 will be referred to as the "terminal".
[0414] Modes for carrying out the invention
[0415] This invention is a system for users to recognize and prevent the risk of wire fraud in advance. This system collects the user's location information and voice data, generates a warning message based on the analysis results, and notifies the user's device and pre-registered contacts of the warning message.
[0416] Program Processing Description
[0417] Collection and transmission of location information
[0418] The device constantly monitors its GPS sensor to detect when the user enters a specific area (e.g., near a bank ATM). When the user enters the designated zone, the device collects location information and periodically sends it to the server. This allows the server to accurately determine the user's location.
[0419] Collection and transmission of audio data
[0420] The device collects voice data in real time from the moment the call begins. The voice data generated when the user makes a call is encrypted as needed and sent to the server. This process requires the user's consent, so it is important that consent is obtained in advance.
[0421] Data analysis using analytical methods
[0422] The server integrates and analyzes the received location information and voice data. It analyzes the location information to determine if the user is in a specific location (e.g., a bank ATM) and uses a generative AI model to analyze the voice data. It detects specific keywords and phrases (e.g., "transfer," "urgent," "money," etc.) and assesses the risk of fraud by matching them against a local crime database.
[0423] Generation and notification of warning messages
[0424] The server generates a warning message based on the analysis results. For example, it might create a message such as, "This call may be a scam. Do not transfer any money." The generated warning message is notified to the user in both voice and text formats. In emergencies, the same warning is also sent to pre-registered contacts (family and friends).
[0425] Specific example
[0426] 1. The user approaches the bank ATM.
[0427] The device detects the user's current location and notifies the server that "the user is approaching a bank ATM."
[0428] The specific location information detected is "Latitude: 35.6895, Longitude: 139.6917 (within Tokyo)".
[0429] 2. Data collection during voice calls
[0430] When a user initiates a call, the device starts recording the call, encrypting it in real time, and sending it to the server.
[0431] Example of a phone call: "Grandma, I need money urgently. Could you transfer it to me today?"
[0432] 3. Server-based analysis
[0433] The server checks the user's location and confirms that the user is near a bank ATM.
[0434] The system analyzes the voice data to detect keywords such as "transfer," "urgent," and "money." Furthermore, it cross-references this data with local crime statistics to determine that the risk of fraud is "high."
[0435] 4. Generating and notifying warning messages
[0436] The server generates a warning message saying, "This call may be a scam. Do not transfer any money under any circumstances," and notifies the device via both voice and text.
[0437] Furthermore, a warning will be sent to pre-registered contacts (family and friends).
[0438] The system of this invention operates in this manner, significantly reducing the risk of users falling victim to wire fraud. Furthermore, by immediately notifying users of warning messages, it enables users to take swift action, effectively preventing damage before it occurs.
[0439] The following describes the processing flow.
[0440] Step 1:
[0441] The device uses a GPS sensor to detect the user's current location. When the user enters a specific area (for example, near a bank ATM), the device collects location information. Subsequently, it sends the collected location information to a server, which then accurately determines the user's current location.
[0442] Step 2:
[0443] When a user initiates a call, the device begins collecting the call audio in real time. This audio data is encrypted to protect privacy. The encrypted audio data is then sent to a server.
[0444] Step 3:
[0445] The server integrates and analyzes the received location information and voice data. It analyzes the location information to confirm that the user is in a specific location (e.g., a bank ATM). It also analyzes the voice data using a generation AI model to detect specific keywords or phrases (e.g., "transfer," "hurry," "money," etc.).
[0446] Step 4:
[0447] The server compares the analysis results against a crime database to assess the risk of fraud. A high risk assessment suggests that the user is at risk of wire fraud.
[0448] Step 5:
[0449] The server generates a warning message based on the risk assessment results. The generated warning message will include appropriate content such as, "This call may be a scam. Do not transfer any money."
[0450] Step 6:
[0451] The device notifies the user of the generated warning message. The notification is provided in both audio and text formats so that the user can immediately review it.
[0452] Step 7:
[0453] Furthermore, in the event of an emergency, the server will also send a warning message to pre-registered contacts (family and friends). This will alert those around the user as well.
[0454] (Example 1)
[0455] Next, we will describe Example 1. In the following description, the data processing device 12 will be referred to as the "server" and the smart glasses 214 will be referred to as the "terminal".
[0456] In modern times, criminal activities such as wire fraud are on the rise, and many people are falling victim to them. Vulnerable individuals, especially the elderly, are particularly vulnerable, making prevention measures an urgent necessity. Conventional wire fraud prevention systems require users to consciously operate them, lacking convenience and immediacy. To address this challenge, a system is needed that minimizes user intervention and detects and warns of fraud risks in real time.
[0457] The identification process performed by the identification processing unit 290 of the data processing device 12 in Example 1 is realized by the following means.
[0458] In this invention, the server includes location information collection means for collecting user location information, location information transmission means for transmitting the location information to a network device, voice data collection means for collecting user voice data, voice data encryption means for encrypting the voice data, voice data transmission means for transmitting the encrypted voice data to a network device, analysis means for analyzing the location information and voice data in the network device, warning message generation means for generating a warning message based on the analysis results of the analysis means, warning notification means for notifying the generated warning message to a user terminal, notification means for notifying the generated warning message to pre-registered contacts, risk assessment means for detecting specific keywords and evaluating risk using a generation artificial intelligence model, range monitoring means for detecting when a user enters a specific range, adjustment means for adjusting the content of the warning message based on the range, and format conversion means for notifying the adjusted message in voice or text format. This makes it possible to detect the risk of fraud in real time and quickly notify warning messages while minimizing user intervention.
[0459] "Location information collection means" refers to sensors or devices used to determine the user's current location, primarily using GPS systems to acquire location information.
[0460] "Location information transmission means" refers to a function or device for transmitting acquired location information to a network device, and transmits data via a communication network.
[0461] "Voice data collection means" refers to devices such as microphones used to collect the user's voice, and to acquire the content of the call as digital voice data.
[0462] "Voice data encryption means" refers to a function or device that performs encryption processing to protect collected voice data, thereby ensuring data security.
[0463] "Voice data transmission means" refers to a function or device for transmitting encrypted voice data to a network device, and which transmits data via a communication network.
[0464] "Analysis means" refers to a function or device for analyzing received location information and audio data, and integrates and analyzes the location data and audio data.
[0465] A "warning message generation means" is a function or device that generates a message to warn the user based on the analysis results, and creates appropriate warning content.
[0466] A "warning notification means" is a function or device for notifying a generated warning message to a user terminal, thereby immediately conveying the warning to the user.
[0467] A "notification method" is a function or device that notifies pre-registered contacts of the generated warning message, thereby informing family and friends of the warning.
[0468] A "generative artificial intelligence model" is a model that analyzes audio data to detect specific keywords and perform risk assessment, using machine learning and natural language processing technologies.
[0469] A "risk assessment tool" is a function or device that analyzes audio data to assess the risk of fraud, and determines the risk level based on specific keywords or phrases.
[0470] "Range monitoring means" refers to a function or device for detecting when a user enters a specific geographical area (e.g., near a bank ATM), and continuously monitors location information.
[0471] "Adjustment means" refers to a function or device for adjusting the content of a warning message based on the user's location and voice data, thereby creating an appropriate message for the situation.
[0472] "Format conversion means" refers to a function or device for notifying warning messages in audio or text format, and for conveying warnings in a format suitable for the user.
[0473] This invention relates to a system for users to recognize and prevent the risk of wire fraud in advance. This system collects and analyzes the user's location information and voice data, and generates and sends a warning message based on the results.
[0474] Collection and transmission of location information
[0475] The terminal constantly monitors its GPS sensor to detect when the user enters a specific area (for example, near a bank ATM). This area monitoring means that when the user enters a designated zone, the terminal collects location information and periodically transmits it to a server via a location information transmission means. The GPS sensor is used in this process.
[0476] Specifically, the device collects location information such as "latitude: 35.6895, longitude: 139.6917 (within Tokyo)" and sends it to the server.
[0477] Collection and transmission of audio data
[0478] The device collects voice data in real time from the moment a call begins using voice data collection means. Voice data generated when a user makes a call is encrypted as needed using voice data encryption means and transmitted to the server via voice data transmission means. This process requires the user's consent.
[0479] Specifically, the device records the conversation as soon as the call starts and sends the audio data to the server while protecting it with AES encryption technology. An example of the conversation content collected might be something like, "Grandma, I need money right away. Can you transfer it to me today?"
[0480] Data Analysis
[0481] The server integrates and analyzes the received location information and encrypted voice data using analysis tools. First, it analyzes the location information to determine if the user is in a specific location (e.g., near a bank ATM). Next, it analyzes the voice data using a generative artificial intelligence model to detect specific keywords and phrases (e.g., "transfer," "urgent," "money," etc.), and assesses the risk of fraud using risk assessment tools. Natural language processing techniques are used in this analysis.
[0482] As an example, the prompt statements for a generative AI model are as follows:
[0483] Analyze the fraud risk from this audio data. Focus on detecting the following keywords and phrases: "transfer," "urgent," and "money." Evaluate the results as a "low," "medium," or "high" risk level to determine the risk level.
[0484] Audio data: ...
[0485] Generation and notification of warning messages
[0486] The server generates a warning message based on the analysis results using a warning message generation system. For example, it might create a message such as, "This call may be a scam. Do not transfer any money." The generated warning message is then notified to the user in both voice and text format using a warning notification system. Depending on the settings, the same warning is also sent to pre-registered contacts (family and friends) via the notification system.
[0487] Specifically, the server generates a warning message stating "This is likely a scam," converts it into voice or text format, and notifies the user and their family.
[0488] This system significantly reduces the risk of users falling victim to wire fraud. Furthermore, by immediately notifying users of warning messages, it enables quick responses, effectively preventing damage before it occurs.
[0489] The flow of the specific processing in Example 1 will be explained using Figure 11.
[0490] Step 1: Collect and transmit location information
[0491] The device constantly monitors the user's location. Specifically, it uses the GPS sensor within the device to obtain the user's current location. When the user enters a specific area (e.g., near a bank ATM), this information is automatically detected. The input is location information from the GPS sensor (e.g., latitude and longitude), and the output is the current user location information. This location information is transmitted to the server using a location information transmission device.
[0492] Step 2: Collect and transmit audio data
[0493] When a user initiates a call, the terminal detects the start of the call. The terminal uses its built-in microphone to collect audio data. The collected audio data is encrypted using an audio data encryption method. The input is the audio data of the call, and the output is encrypted audio data. Subsequently, the encrypted audio data is sent to the server using an audio data transmission method.
[0494] Step 3: Analysis of location and audio data
[0495] The server integrates and analyzes the received location information and voice data. First, it analyzes the location information to determine if the user is in a specific location (e.g., near a bank ATM). Next, it analyzes the voice data using a generative artificial intelligence model. The input consists of decomposed location information and encrypted voice data, and a risk assessment is output as a result of the analysis process. Prompt sentences are input to the generative artificial intelligence model to detect specific keywords (e.g., "transfer", "urgent", "money").
[0496] Step 4: Generating a warning message
[0497] The server generates a warning message based on the analysis results. The AI generation model automatically generates specific warning content tailored to the risk assessment from its output. The input is the result of voice data analysis, and the output is a concrete warning message. For example, a message such as "This call may be a scam. Do not transfer any money under any circumstances" might be generated.
[0498] Step 5: Notification of warning message
[0499] The server notifies the user terminal of the generated warning message. The warning message is sent in both voice and text formats using the warning notification method. The input is the generated warning message, and the output is the notification sent to the user. In addition, the same warning is sent to pre-registered contacts (family and friends). The terminal displays the received warning message to the user and provides voice notifications depending on the content of the warning.
[0500] This entire process makes it possible to detect the risk of wire fraud in real time and to quickly and appropriately notify users of warning messages.
[0501] (Application Example 1)
[0502] Next, we will explain Application Example 1. In the following explanation, the data processing device 12 will be referred to as the "server," and the smart glasses 214 will be referred to as the "terminal."
[0503] In recent years, there has been a significant increase in wire fraud, particularly targeting the elderly. These scams cleverly exploit the victim's psychology and typically utilize voice communication methods such as telephone calls. Traditional prevention measures, relying solely on vigilance, have been insufficient for effective prevention. Furthermore, there has been a lack of mechanisms to recognize fraud risks in real time and take immediate action. Against this backdrop, there is a need for technology that enables users to detect fraud risks in advance and prevent becoming a victim.
[0504] The specific processing performed by the specific processing unit 290 of the data processing device 12 in Application Example 1 is realized by the following means.
[0505] In this invention, the server includes location information collection means for collecting user location information, voice data collection means for collecting user voice data, and analysis processing means for analyzing using a generative AI model. This makes it possible to analyze both voice data and location information when a user approaches a specific risk area (e.g., near a bank ATM) and immediately recognize and warn of the risk of fraud.
[0506] "Location information collection means" refers to means for obtaining the user's current geographical location.
[0507] "Location information transmission means" refers to a means for transmitting collected location information to a server.
[0508] "Voice data collection means" refers to means for collecting a user's voice communication data.
[0509] "Voice data encryption means" refers to a means for encrypting collected voice data.
[0510] "Voice data transmission means" refers to a means for sending encrypted voice data to a server.
[0511] "Analysis means" refers to methods for evaluating fraud risk by analyzing location information and voice data.
[0512] A "generative AI model" is an artificial intelligence model that learns from large amounts of data and is used to perform specific tasks (e.g., fraud detection).
[0513] The "analysis processing means" is a means for comprehensively analyzing location information and audio data using a generated AI model.
[0514] A "warning message generation means" is a means for creating a warning message for the user based on the analysis results.
[0515] A "warning notification means" is a means for notifying the user terminal of the generated warning message.
[0516] "Notification method" refers to a means of notifying pre-registered contacts of the generated warning message.
[0517] "Voice notification means" refers to a means of notifying the user of the generated warning message in voice format.
[0518] A "text notification method" is a means of notifying the user of the generated warning message in text format.
[0519] Regarding embodiments for carrying out the invention, this system employs a computer-based method that combines multiple means for users to recognize and prevent the risk of wire fraud in advance.
[0520] composition
[0521] This system includes a user terminal, location information collection means, voice data collection means, generation AI model, analysis processing means, warning message generation means, warning notification means, and notification means.
[0522] User terminal: This refers to a device such as a smartphone carried by the user. The terminal is equipped with a GPS sensor and a microphone, and collects location information and voice data.
[0523] Location information collection method: Collects the user's current location. It is possible to monitor location information in real time using GPS.
[0524] Voice data collection method: Voice data during calls is collected in real time. It is collected using a microphone and encrypted for privacy protection.
[0525] Generative AI Models: These models use pre-trained AI models to detect fraudulent activity. They assess the risk of fraud based on a large amount of historical data.
[0526] Analysis processing method: Collected location information and audio data are analyzed using a generating AI model. This allows for the assessment of fraud risk.
[0527] Warning message generation means: Generates a warning message based on the analysis results. Messages are generated in both audio and text formats.
[0528] Warning notification method: The generated warning message is notified to the user's terminal. Specifically, the message is displayed on the smartphone screen and an audible warning is issued.
[0529] Notification method: In an emergency, the same warning message will be sent to pre-registered contacts (family and friends).
[0530] How to use
[0531] This system works as follows:
[0532] 1. When a user approaches a specific area:
[0533] When a user approaches a specific area, such as a bank ATM, the GPS sensor on the user's device collects location information and sends it to a server. The server analyzes this information to detect that the user is near a bank ATM.
[0534] 2. When the user initiates a call:
[0535] When a user initiates a call, the call audio is collected by a voice data collection system, encrypted in real time, and then transmitted to the server. During this process, the voice data is appropriately encrypted to protect the user's privacy.
[0536] 3. Analysis of fraud risk:
[0537] The server integrates collected location and voice data and performs analysis using a generative AI model. When specific keywords or phrases are detected, they are compared against a local crime database to assess the fraud risk.
[0538] Specific example
[0539] Example of a user approaching a bank ATM:
[0540] The terminal notifies the server that "the user is approaching a bank ATM," and specific location information (e.g., latitude: 35.6895, longitude: 139.6917) is detected.
[0541] Examples of data collection during a call:
[0542] When a user calls and says, "Grandma, I need money right away. Can you transfer it to me today?", the device records the audio, encrypts it, and sends it to the server.
[0543] Server-based analysis example:
[0544] The generative AI model detects keywords such as "bank transfer," "urgent," and "money," and compares them with local crime data to determine that the fraud risk is "high."
[0545] Example of a prompt
[0546] Examples of prompt statements are as follows:
[0547] Audio recording: "Grandma, I need money urgently. Can you transfer it to me today?"
[0548] Please evaluate the risk of fraud as part of the analysis. If you find any specific keywords or phrases, return their probability.
[0549] In this way, the system can quickly detect the risk of fraud and prevent damage by warning the user.
[0550] The flow of a specific process in Application Example 1 will be explained using Figure 12.
[0551] Step 1:
[0552] When a user approaches a specific area (e.g., near a bank ATM), the device's GPS sensor collects location information. This information is recorded on the device as numerical data such as "latitude: 35.6895, longitude: 139.6917". The input is the current location information, and the output is the collected location information. The device sends this information to the server.
[0553] Step 2:
[0554] The server analyzes the received location information to determine if the user is in a specific risk area. This analysis confirms that the user is near a bank ATM. The input is location information, and the output is the risk area determination result.
[0555] Step 3:
[0556] When a user initiates a call, the device collects the call audio in real time. The audio data is captured via the microphone and encrypted as needed. The input is the call audio, and the output is encrypted audio data. The device then sends this audio data to the server.
[0557] Step 4:
[0558] The server decrypts encrypted audio data and extracts keywords in order to analyze the received audio data using a generative AI model. The input is encrypted audio data, and the output is string data for analysis. Based on this string data, the generative AI model detects specific keywords (e.g., "transfer", "urgent", "money").
[0559] Step 5:
[0560] Based on the analyzed results, the server cross-references them with a local crime database to assess the fraud risk. The input consists of analyzed audio data, location information, and the crime database; the output is the fraud risk assessment result. If the server determines the risk is high, it generates a warning message.
[0561] Step 6:
[0562] The generated warning message is created by the server in both audio and text formats and immediately notified to the user. The input is the risk assessment results and a warning message template, and the output is the generated warning message. The terminal then informs the user of this message via audio and text.
[0563] Step 7:
[0564] The server also sends warning messages to pre-registered contacts (e.g., family and friends). The input is the generated warning message, and the output is the notification to the registered contacts. This helps users take quick action.
[0565] Through the above processing steps, this system can issue appropriate warnings to users before they are exposed to the risk of fraud, thereby preventing them from becoming victims.
[0566] Furthermore, an emotion engine that estimates the user's emotions may be incorporated. That is, the identification processing unit 290 may use the emotion identification model 59 to estimate the user's emotions and perform identification processing using the user's emotions.
[0567] Modes for carrying out the invention
[0568] This invention is a system for users to recognize and prevent the risk of wire fraud in advance. This system collects the user's location information and voice data, and generates a warning message based on the analysis results. It also utilizes an emotion engine to recognize the user's emotions, thereby improving the accuracy and effectiveness of the warning. The warning message is then sent to the user's device and pre-registered contacts.
[0569] Program Processing Description
[0570] Collection and transmission of location information
[0571] The device uses a GPS sensor to detect the user's current location. When the user enters a specific area (for example, near a bank ATM), the device collects location information and periodically sends it to the server. This allows the server to accurately determine the user's location.
[0572] Collection and transmission of audio data
[0573] When a user initiates a call, the device begins collecting the call audio in real time. This audio data is encrypted to protect privacy. The encrypted audio data is then sent to a server.
[0574] Emotional analysis using an emotion engine
[0575] The terminal or server uses the received audio data to evaluate the user's emotional state. The emotion engine recognizes the user's emotions (e.g., tension, fear, anxiety) from factors such as the tone, speed, and intonation of the voice.
[0576] Data analysis using analytical methods
[0577] The server integrates and analyzes the received location information and voice data. It analyzes the location information to confirm that the user is in a specific location (e.g., a bank ATM). It also uses a generative AI model to analyze the voice data and detect specific keywords or phrases (e.g., "transfer," "hurry," "money," etc.).
[0578] Risk assessment that takes emotional state into account
[0579] The server takes into account the results of the emotion engine analysis and compares them with the crime database to ultimately assess the risk of fraud. If the user's emotions are in a high-risk state such as tension, fear, or anxiety, the server determines that the likelihood of fraud is higher.
[0580] Generation and notification of warning messages
[0581] The server generates a warning message based on the risk assessment results. For example, it might say, "This call may be a scam. Do not transfer any money." The generated warning message is sent to the device in both voice and text formats. In emergencies, the warning message is also sent to pre-registered contacts (family and friends).
[0582] Specific example
[0583] 1. The user approaches the bank ATM.
[0584] The device detects the user's current location and notifies the server that "the user is approaching a bank ATM."
[0585] The specific location information detected is "Latitude: 35.6895, Longitude: 139.6917 (within Tokyo)".
[0586] 2. Data collection during voice calls
[0587] When a user initiates a call, the device starts recording the call, encrypting it in real time, and sending it to the server.
[0588] Example of a phone call: "Grandma, I need money urgently. Could you transfer it to me today?"
[0589] 3. Analysis using an emotion engine
[0590] The server analyzes the user's emotions from the voice data and detects tension and fear.
[0591] 4. Server-based data analysis
[0592] The server checks the user's location and confirms that the user is near a bank ATM.
[0593] The system analyzes voice data using an AI model to detect keywords such as "transfer," "urgent," and "money." Furthermore, it compares this data with local crime data to determine a high risk of fraud.
[0594] 5. Generating and notifying warning messages
[0595] The server generates a warning message stating, "This call may be a scam. Do not transfer any money under any circumstances," and notifies the device via both voice and text. Furthermore, warnings are also sent to pre-registered contacts (family and friends).
[0596] The system of this invention operates in this manner, significantly reducing the risk of users falling victim to wire fraud. Furthermore, by immediately notifying users of warning messages, it enables users to take swift action, effectively preventing damage before it occurs.
[0597] The following describes the processing flow.
[0598] Step 1:
[0599] The device uses a GPS sensor to detect the user's current location. When the user enters a specific area (for example, near a bank ATM), the device collects location information and periodically sends it to the server. This allows the server to accurately determine the user's location.
[0600] Step 2:
[0601] When a user initiates a call, the device begins collecting the call audio in real time. This audio data is encrypted to protect privacy. The encrypted audio data is then sent to a server.
[0602] Step 3:
[0603] The server integrates and analyzes the received location information and voice data. It analyzes the location information to confirm that the user is in a specific location (e.g., a bank ATM). It also uses a generative AI model to analyze the voice data and detect specific keywords or phrases (e.g., "transfer," "hurry," "money," etc.).
[0604] Step 4:
[0605] The terminal or server uses the received audio data to evaluate the user's emotional state. The emotion engine recognizes the user's emotions (e.g., tension, fear, anxiety) from factors such as tone, speed, and intonation of the voice. If the emotional state is determined to be tension, fear, or anxiety, special attention is paid to the analysis results.
[0606] Step 5:
[0607] The server takes into account the results of the emotion engine analysis and compares them with the crime database to ultimately assess the risk of fraud. If the user's emotions are in a high-risk state such as tension, fear, or anxiety, the server determines that the likelihood of fraud is higher. Based on this, the accuracy and reliability of the risk assessment are improved.
[0608] Step 6:
[0609] The server generates a warning message based on the risk assessment results. For example, it might say, "This call may be a scam. Do not transfer any money." The generated warning message is notified to the user in both audio and text format.
[0610] Step 7:
[0611] The device notifies the user of the generated warning message in both voice and text format. It also notifies the user to immediately check the warning message, encouraging a quick response. Furthermore, in emergencies, the warning message is also sent to pre-registered contacts (family and friends). This alerts those around the user as well.
[0612] Specific example:
[0613] 1. The user approaches the bank ATM.
[0614] The device detects the user's current location and notifies the server that "the user is approaching a bank ATM."
[0615] The location information detected is "Latitude: 35.6895, Longitude: 139.6917 (within Tokyo)".
[0616] 2. Data collection during voice calls
[0617] When a user initiates a call, the device begins recording the conversation. The recorded data is encrypted and sent to a server in real time.
[0618] Example of phone call content: "Grandma, I need money urgently. Could you transfer it to me today?"
[0619] 3. Data analysis by the server
[0620] The server analyzes the location information to confirm that the user is near a bank ATM.
[0621] The AI model analyzes the voice data to detect keywords such as "transfer," "urgent," and "money."
[0622] 4. Emotional analysis using an emotion engine
[0623] The server detects the user's emotions from the voice data and recognizes signs of tension or fear.
[0624] 5. Risk Assessment
[0625] The server also considers the results of the emotion engine and assesses the fraud risk as high based on the analysis results.
[0626] 6. Generating and notifying warning messages
[0627] The server generates a warning message stating, "This call may be a scam. Do not transfer any money under any circumstances," and sends it to the device in both voice and text formats.
[0628] 7. Notification to the user
[0629] The device notifies the user of warning messages in both voice and text formats, and in emergencies, it also sends warnings to pre-registered contacts.
[0630] In this way, the system of the present invention significantly reduces the risk of users falling victim to wire fraud and enables rapid countermeasures. Furthermore, by taking into account the user's emotional state, the accuracy and effectiveness of the warning are further improved.
[0631] (Example 2)
[0632] Next, we will describe Example 2. In the following description, the data processing device 12 will be referred to as the "server" and the smart glasses 214 will be referred to as the "terminal".
[0633] Currently, the number of victims of wire fraud is increasing, raising the risk of users being deceived and losing their valuable assets. Traditional countermeasures are insufficient because they lack the means for users to recognize the risk of fraud in advance, making it difficult to prevent victims. To solve this problem, there is a need for a system that can accurately detect the risk of wire fraud and provide users with rapid warnings.
[0634] The identification process performed by the identification processing unit 290 of the data processing device 12 in Example 2 is realized by the following means.
[0635] In this invention, the server includes emotion analysis means for analyzing the user's emotional state, analysis means for integrating and analyzing location information and voice data and detecting specific keywords or phrases using a generative AI model, and means for performing risk assessment. This enables users to detect the risk of wire fraud in advance with high accuracy and receive a warning quickly.
[0636] "Location information collection means" refers to means for obtaining the user's current location information.
[0637] "Location information transmission means" refers to a means for transmitting acquired location information to a server.
[0638] "Voice data collection means" refers to a means for collecting user voice data.
[0639] "Voice data encryption means" refers to a method of encrypting collected voice data for privacy protection.
[0640] "Voice data transmission means" refers to a means for sending encrypted voice data to a server.
[0641] "Emotion analysis means" refers to a method for evaluating a user's emotional state from voice data.
[0642] "Analysis means" refers to a means for integrating and analyzing location information and audio data, and for using a generative AI model to detect specific keywords or phrases.
[0643] A "risk assessment tool" is a means for evaluating risk based on the analysis results of an analysis tool.
[0644] A "warning message generation means" is a means for generating a warning message based on the results of a risk assessment.
[0645] A "warning notification means" is a means for notifying the user terminal of the generated warning message.
[0646] "Notification method" refers to a means of notifying pre-registered contacts of the generated warning message.
[0647] "Voice notification means" refers to a means of notifying the user of the generated warning message in voice format.
[0648] A "text notification method" is a means of notifying the user of the generated warning message in text format.
[0649] A "generative AI model" is an artificial intelligence model used to analyze audio data and detect specific keywords or phrases.
[0650] This invention is a system for users to recognize and prevent the risk of wire fraud in advance. The system includes means for collecting the user's location information and voice data and generating a warning message based on the analysis results.
[0651] Hardware configuration
[0652] The system utilizes mobile devices such as smartphones equipped with a GPS sensor, microphone, and data communication capabilities. The GPS sensor collects the user's current location, and the microphone collects call audio. The collected data is then transmitted to a server via the internet.
[0653] The server utilizes high-performance computers and cloud services to provide the processing power needed for data analysis and sentiment analysis. Specific software used includes IBM Watson's sentiment analysis API and GPT-3 generative AI models.
[0654] Software Configuration
[0655] The device has the following software features:
[0656] 1. Location information collection method: The user's current location is obtained using the built-in GPS sensor, and location data is generated.
[0657] 2. Voice data collection method: Voice data is collected in real time using a microphone during a call.
[0658] 3. Audio data encryption method: The collected audio data is encrypted using AES-256.
[0659] 4. Location information transmission means: The collected location information is transmitted to the server.
[0660] 5. Means of transmitting audio data: Encrypted audio data is sent to the server.
[0661] The server has the following software features:
[0662] 1. Emotion analysis method: The received audio data is analyzed using IBM Watson's emotion analysis API to evaluate the user's emotional state.
[0663] 2. Analysis method: The received location information and audio data are integrated, and specific keywords and phrases are detected using the GPT-3 generation AI model.
[0664] 3. Risk assessment method: The risk of wire fraud will be finalized based on the analysis results and sentiment analysis results.
[0665] 4. Warning message generation means: Generates a warning message based on the risk assessment results.
[0666] 5. Warning notification means: The generated warning message is sent to the user's terminal and to pre-registered contacts.
[0667] Specific example
[0668] 1. The user approaches the bank ATM.
[0669] The device uses a GPS sensor to obtain its current location and sends the location information to the server when it detects that "the user is approaching a bank ATM." An example of location data is "Latitude: 35.6895, Longitude: 139.6917 (within Tokyo)."
[0670] 2. Data collection during voice calls
[0671] When a user initiates a call, the device collects the call audio, encrypts it with AES-256, and sends it to the server in real time. Specifically, the call content might be something like, "Grandma, I need money urgently. Can you transfer it to me today?" The audio recording and encryption process is performed on the device itself.
[0672] 3. Analysis using an emotion engine
[0673] The server decodes the received audio data and uses an emotion analysis API to assess the user's emotion as "tension." The analysis takes into account the tone, speed, and intonation of the voice.
[0674] 4. Server-based data analysis
[0675] The server uses location data to confirm that the user is near a bank ATM. Simultaneously, it analyzes the voice data using an AI model (e.g., GPT-3) to detect keywords such as "transfer," "urgent," and "money." This then triggers a voice-to-text conversion process, highlighting specific keywords.
[0676] 5. Risk assessment that takes emotional state into consideration
[0677] The server determines a high risk of fraud by combining voice analysis results and sentiment analysis results. This includes matching against past crime data and calculating the degree of match.
[0678] 6. Generating and notifying warning messages
[0679] The server generates a warning message stating, "This call may be a scam. Do not transfer any money under any circumstances," and also generates a voice message using a speech synthesis API. The generated warning message is sent to the device and, if necessary, also notified to registered contacts via SMS or email.
[0680] Example of a prompt
[0681] "Please explain the process for a program that notifies a user when they approach a bank ATM. Also, please explain the procedure for collecting voice data from a user during a call and assessing the risk of fraud."
[0682] The flow of the specific processing in Example 2 will be explained using Figure 13.
[0683] Step 1: Collect and transmit location information
[0684] The device uses its built-in GPS sensor to acquire the user's current location at regular intervals. Specifically, it acquires location information every minute and generates latitude and longitude data (input: GPS sensor data, output: location information data). When the user enters a specific area (for example, near a bank ATM), it sends that location information to the server (data processing: location information detection, data calculation: range determination, output: location information transmission).
[0685] Step 2: Collect and transmit audio data
[0686] When a user initiates a call, the device uses its microphone to collect the call audio in real time (input: call audio, output: audio data). The collected audio data is immediately encrypted using AES-256 (data processing: audio data encryption), and the encrypted audio data is sent to the server in real time (output: encrypted audio data transmission).
[0687] Step 3: Emotional analysis using the emotion engine
[0688] The server decrypts the received encrypted audio data (input: encrypted audio data, output: decrypted audio data) and sends the audio data to the emotion analysis API. The emotion analysis API analyzes the tone, speed, and intonation of the voice and evaluates the user's emotional state (tension, fear, anxiety, etc.) (data calculation: emotion analysis, output: emotional state data).
[0689] Step 4: Data analysis using analytical tools
[0690] The server integrates and analyzes emotional state data with received location and audio data (input: location data, decoded audio data, emotional state data). It uses a GPT-3 generative AI model to detect specific keywords or phrases (e.g., "transfer", "hurry", "money") from the audio data (data calculation: keyword detection, output: keyword data). It uses location information to determine if the user is in a specific location (e.g., a bank ATM) (data calculation: location confirmation, output: location confirmation data).
[0691] Step 5: Risk assessment considering emotional state
[0692] The server integrates keyword data and location data from the analysis tools, and also considers emotional state data to assess fraud risk (Input: Keyword data, emotional state data, location data; Output: Risk assessment data). For example, if the user's emotions are in a high-risk state such as tension or fear, and specific keywords are included, the risk is judged to be high (Data calculation: Risk assessment).
[0693] Step 6: Generate and send warning messages
[0694] The server generates a warning message based on risk assessment data (input: risk assessment data, output: warning message). For example, it generates a message stating, "This call may be a scam. Do not transfer any money under any circumstances." (data processing: warning message generation). The generated warning message is converted into voice format using a speech synthesis API (output: voice warning message) and also sent to the terminal in text format (output: text warning message). Furthermore, in emergencies, notifications are also sent to pre-registered contacts (family and friends) (output: contact notification data).
[0695] (Application Example 2)
[0696] Next, we will explain application example 2. In the following explanation, the data processing device 12 will be referred to as the "server," and the smart glasses 214 will be referred to as the "terminal."
[0697] Fraudulent activities, including wire transfer scams, are increasing, particularly targeting the elderly. Therefore, there is a need to detect the risk of fraud early and prevent it before it occurs. However, current systems lack the means to effectively analyze users' emotional states and real-time call content, making it difficult to accurately determine situations with a high risk of fraud. Therefore, it is necessary to provide a system that can quickly and accurately detect situations with a high risk of fraud and issue warnings to users, their relatives, and friends.
[0698] The specific processing performed by the specific processing unit 290 of the data processing device 12 in Application Example 2 is realized by the following means.
[0699] In this invention, the server includes location information collection means for collecting user location information, voice data collection means for collecting user voice data, emotion analysis means for analyzing the user's emotional state based on the voice data, and keyword detection means for detecting specific keywords. This makes it possible to detect fraud risks such as wire fraud with high accuracy and speed, and to issue warnings to the user and their family and friends.
[0700] "Location information collection means" refers to devices and software used to determine the user's current location.
[0701] "Location information transmission means" refers to devices or software used to transmit collected user location information to a server.
[0702] "Voice data collection means" refers to devices or software used to collect voices emitted by users.
[0703] "Voice data encryption means" refers to devices or software used to encrypt collected voice data so that it cannot be deciphered by third parties.
[0704] "Voice data transmission means" refers to devices or software used to transmit encrypted voice data to a server.
[0705] "Analysis means" refers to devices or software used to analyze location information and voice data collected on a server to assess fraud risk.
[0706] "Warning message generation means" refers to a device or software that automatically generates warning messages based on the analysis results of the analysis means.
[0707] "Warning notification means" refers to devices or software that notify the user's terminal of the generated warning message.
[0708] "Notification means" refers to devices or software used to notify pre-registered contacts of generated warning messages.
[0709] "Emotional analysis means" refers to devices and software that analyze a user's emotional state in real time based on voice data.
[0710] "Keyword detection means" refers to devices or software used to detect specific keywords from audio data.
[0711] This invention is a system for users to recognize and prevent the risk of wire fraud in advance. The system includes means for identifying the user's current location and transmitting that information to a server, and means for collecting, encrypting, and transmitting the user's voice data. Furthermore, the server includes means for analyzing this information and generating a warning message to notify the user's terminal and pre-registered contacts.
[0712] Collection and transmission of location information
[0713] The device is equipped with a GPS sensor that constantly monitors the user's current location and acquires data using location information collection methods. When the user enters a specific area (for example, near a bank ATM), location information is sent to the server using location information transmission methods. This allows the server to determine the user's precise location.
[0714] Collection and transmission of audio data
[0715] When a user starts a phone call, the call content is collected in real time by an audio data collection means and encrypted by an audio data encryption means. The encrypted audio data is then transmitted to the server using an audio data transmission means.
[0716] Sentiment analysis and keyword detection
[0717] The server uses emotion analysis tools to analyze the received audio data. This evaluates the user's emotions (e.g., tension, fear, anxiety) from factors such as tone, speed, and intonation. Furthermore, keyword detection tools detect specific keywords within the audio data (e.g., "transfer," "hurry," "money," etc.).
[0718] Risk assessment and warning message generation
[0719] The server integrates location information, voice data, sentiment analysis results, and keyword detection results to assess the fraud risk. If the risk is determined to be high, a warning message is generated by the warning message generation mechanism. For example, a message such as, "This call may be a scam. Do not transfer any money."
[0720] Warning message notification
[0721] The warning notification system will send the generated warning message to the user's device in both audio and text format. In emergencies, the same warning message will also be sent to pre-registered contacts (family and friends) using the notification system.
[0722] Specific example
[0723] 1. The user approaches the bank ATM.
[0724] The device detects the user's current location and notifies the server that "the user is approaching a bank ATM." The specific location information detected is "Latitude: 35.6895, Longitude: 139.6917 (within Tokyo)."
[0725] 2. Data collection during voice calls
[0726] When a user initiates a call, the device starts recording the conversation, encrypting it in real time, and sending it to the server. Example of a call: "Grandma, I need money urgently. Can you transfer it to me today?"
[0727] 3. Analysis using an emotion engine
[0728] The server analyzes the user's emotions from the voice data and detects tension and fear.
[0729] 4. Server-based data analysis
[0730] The server checks the user's location to confirm they are near a bank ATM. An AI model analyzes the voice data to detect keywords such as "transfer," "urgent," and "money." Furthermore, it cross-references this data with local crime statistics to determine a high risk of fraud.
[0731] 5. Generating and notifying warning messages
[0732] The server generates a warning message stating, "This call may be a scam. Do not transfer any money under any circumstances," and notifies the device via both voice and text. Furthermore, warnings are also sent to pre-registered contacts (family and friends).
[0733] Example of a prompt
[0734] text
[0735] Evaluate the user's emotional state and generate an emergency fraud warning message. For example, use this if you receive audio data such as "Location: 35.6895, 139.6917" and "Grandma, I need money right away. Can you transfer it to me today?". If the emotional score is -0.5 or lower and contains fraud-related keywords such as "transfer" and "money", generate a warning message stating, "This call may be a scam. Do not transfer any money."
[0736] The above describes the embodiments for carrying out this invention.
[0737] The flow of a specific process in Application Example 2 will be explained using Figure 14.
[0738] Step 1:
[0739] Collecting and transmitting user location information
[0740] The device uses a GPS sensor to collect the user's location information. The collected location information (e.g., latitude 35.6895, longitude 139.6917) is transmitted to the server using a location information transmission means.
[0741] Input: Location information from GPS sensor
[0742] Output: User location information sent to the server
[0743] Step 2:
[0744] Collection and transmission of audio data
[0745] When a user initiates a call, the device collects audio in real time using an audio data collection device. The collected audio data is encrypted using an audio data encryption device and transmitted to a server using an audio data transmission device.
[0746] Input: Audio data from a call
[0747] Output: Encrypted audio data sent to the server
[0748] Step 3:
[0749] Location information analysis
[0750] The server analyzes the received location information using an analysis tool to determine whether the user is in a specific risk area (e.g., near a bank ATM).
[0751] Input: Sent location information
[0752] Output: Determination result of whether the user is in a risk area.
[0753] Step 4:
[0754] Sentiment analysis of voice data
[0755] The server analyzes the audio data using emotion analysis tools. Specifically, it evaluates the user's emotions (e.g., tension, fear, anxiety) based on factors such as tone, speed, and intonation of the voice.
[0756] Input: Decrypted audio data
[0757] Output: User sentiment score
[0758] Step 5:
[0759] Detection of specific keywords
[0760] The server uses keyword detection means to detect specific keywords (e.g., "transfer", "hurry", "money") from the audio data.
[0761] Input: Decrypted audio data
[0762] Output: List of detected keywords
[0763] Step 6:
[0764] Risk assessment
[0765] The server integrates location data analysis results, user sentiment scores, and detected keyword lists to assess fraud risk. If a high risk is determined, a warning message is generated using a warning message generation mechanism.
[0766] Input: Location analysis results, sentiment score, keyword list
[0767] Output: Fraud risk assessment results
[0768] Step 7:
[0769] Generation and notification of warning messages
[0770] The server sends the generated warning message to the user's terminal using a warning notification system. Simultaneously, it also sends the same warning message to pre-registered contacts using the notification system.
[0771] Input: Fraud risk assessment results
[0772] Output: Warning message sent to the user's device and registered contacts.
[0773] These steps enable the creation of a system that effectively detects fraud risks and warns users and their associates.
[0774] The specific processing unit 290 transmits the result of the specific processing to the smart glasses 214. In the smart glasses 214, the control unit 46A causes the speaker 240 to output the result of the specific processing. The microphone 238 acquires audio indicating user input for the result of the specific processing. The control unit 46A transmits the audio data indicating user input acquired by the microphone 238 to the data processing unit 12. In the data processing unit 12, the specific processing unit 290 acquires the audio data.
[0775] Data generation model 58 is a type of so-called generative AI (Artificial Intelligence). One example of data generation model 58 is ChatGPT (Internet search<URL: https: / / openai.com / blog / chatgpt> ), Gemini (Internet search) <url: https: gemini.google.com ?hl="ja">Examples of generative AI include the following. The data generation model 58 is obtained by performing deep learning on a neural network. The data generation model 58 is input with prompts containing instructions, and with inference data such as audio data representing speech, text data representing text, and image data representing images. The data generation model 58 infers from the input inference data according to the instructions indicated by the prompts, and outputs the inference results in data formats such as audio data and text data. Here, inference refers to, for example, analysis, classification, prediction, and / or summarization.
[0776] In the above embodiment, an example was given in which specific processing is performed by the data processing device 12, but the technology of this disclosure is not limited thereto, and the specific processing may also be performed by the smart glasses 214.
[0777] [Third Embodiment]
[0778] Figure 5 shows an example of the configuration of the data processing system 310 according to the third embodiment.
[0779] As shown in Figure 5, the data processing system 310 includes a data processing device 12 and a headset terminal 314. An example of the data processing device 12 is a server.
[0780] The data processing device 12 comprises a computer 22, a database 24, and a communication interface 26. The computer 22 is an example of a "computer" related to the technology of this disclosure. The computer 22 comprises a processor 28, RAM 30, and storage 32. The processor 28, RAM 30, and storage 32 are connected to a bus 34. The database 24 and the communication interface 26 are also connected to the bus 34. The communication interface 26 is connected to a network 54. An example of the network 54 is a WAN (Wide Area Network) and / or a LAN (Local Area Network).
[0781] The headset terminal 314 includes a computer 36, a microphone 238, a speaker 240, a camera 42, a communication interface 44, and a display 343. The computer 36 includes a processor 46, RAM 48, and storage 50. The processor 46, RAM 48, and storage 50 are connected to a bus 52. The microphone 238, speaker 240, camera 42, and display 343 are also connected to the bus 52.
[0782] The microphone 238 receives voice signals from the user 20 and receives instructions from the user 20. The microphone 238 captures the voice signals from the user 20, converts the captured voice into audio data, and outputs it to the processor 46. The speaker 240 outputs audio according to the instructions from the processor 46.
[0783] Camera 42 is a small digital camera equipped with an optical system including a lens, aperture, and shutter, and an image sensor such as a CMOS (Complementary Metal-Oxide-Semiconductor) image sensor or a CCD (Charge Coupled Device) image sensor, and captures images of the area around the user 20 (for example, an imaging range defined by a field of view equivalent to the width of a typical healthy person's field of vision).
[0784] Communication interface 44 is connected to network 54. Communication interfaces 44 and 26 are responsible for the exchange of various information between processor 46 and processor 28 via network 54. The exchange of various information between processor 46 and processor 28 using communication interfaces 44 and 26 is performed in a secure manner.
[0785] Figure 6 shows an example of the main functions of the data processing device 12 and the headset terminal 314. As shown in Figure 6, the data processing device 12 performs specific processing using the processor 28. The storage 32 stores the specific processing program 56.
[0786] The specific processing program 56 is an example of a "program" relating to the technology of this disclosure. The processor 28 reads the specific processing program 56 from the storage 32 and executes the read specific processing program 56 on the RAM 30. The specific processing is realized by the processor 28 operating as a specific processing unit 290 in accordance with the specific processing program 56 executed on the RAM 30.
[0787] The storage 32 stores the data generation model 58 and the emotion identification model 59. The data generation model 58 and the emotion identification model 59 are used by the identification processing unit 290.
[0788] In the headset terminal 314, the processor 46 performs the reception output processing. The storage 50 stores the reception output program 60. The processor 46 reads the reception output program 60 from the storage 50 and executes the read reception output program 60 on the RAM 48. The reception output processing is realized by the processor 46 operating as a control unit 46A according to the reception output program 60 executed on the RAM 48.
[0789] Next, the specific processing performed by the specific processing unit 290 of the data processing device 12 will be described. In the following description, the data processing device 12 will be referred to as the "server" and the headset terminal 314 will be referred to as the "terminal".
[0790] Modes for carrying out the invention
[0791] This invention is a system for users to recognize and prevent the risk of wire fraud in advance. This system collects the user's location information and voice data, generates a warning message based on the analysis results, and notifies the user's device and pre-registered contacts of the warning message.
[0792] Program Processing Description
[0793] Collection and transmission of location information
[0794] The device constantly monitors its GPS sensor to detect when the user enters a specific area (e.g., near a bank ATM). When the user enters the designated zone, the device collects location information and periodically sends it to the server. This allows the server to accurately determine the user's location.
[0795] Collection and transmission of audio data
[0796] The device collects voice data in real time from the moment the call begins. The voice data generated when the user makes a call is encrypted as needed and sent to the server. This process requires the user's consent, so it is important that consent is obtained in advance.
[0797] Data analysis using analytical methods
[0798] The server integrates and analyzes the received location information and voice data. It analyzes the location information to determine if the user is in a specific location (e.g., a bank ATM) and uses a generative AI model to analyze the voice data. It detects specific keywords and phrases (e.g., "transfer," "urgent," "money," etc.) and assesses the risk of fraud by matching them against a local crime database.
[0799] Generation and notification of warning messages
[0800] The server generates a warning message based on the analysis results. For example, it might create a message such as, "This call may be a scam. Do not transfer any money." The generated warning message is notified to the user in both voice and text formats. In emergencies, the same warning is also sent to pre-registered contacts (family and friends).
[0801] Specific example
[0802] 1. The user approaches the bank ATM.
[0803] The device detects the user's current location and notifies the server that "the user is approaching a bank ATM."
[0804] The specific location information detected is "Latitude: 35.6895, Longitude: 139.6917 (within Tokyo)".
[0805] 2. Data collection during voice calls
[0806] When a user initiates a call, the device starts recording the call, encrypting it in real time, and sending it to the server.
[0807] Example of a phone call: "Grandma, I need money urgently. Could you transfer it to me today?"
[0808] 3. Server-based analysis
[0809] The server checks the user's location and confirms that the user is near a bank ATM.
[0810] The system analyzes the voice data to detect keywords such as "transfer," "urgent," and "money." Furthermore, it cross-references this data with local crime statistics to determine that the risk of fraud is "high."
[0811] 4. Generating and notifying warning messages
[0812] The server generates a warning message saying, "This call may be a scam. Do not transfer any money under any circumstances," and notifies the device via both voice and text.
[0813] Furthermore, a warning will be sent to pre-registered contacts (family and friends).
[0814] The system of this invention operates in this manner, significantly reducing the risk of users falling victim to wire fraud. Furthermore, by immediately notifying users of warning messages, it enables users to take swift action, effectively preventing damage before it occurs.
[0815] The following describes the processing flow.
[0816] Step 1:
[0817] The device uses a GPS sensor to detect the user's current location. When the user enters a specific area (for example, near a bank ATM), the device collects location information. Subsequently, it sends the collected location information to a server, which then accurately determines the user's current location.
[0818] Step 2:
[0819] When a user initiates a call, the device begins collecting the call audio in real time. This audio data is encrypted to protect privacy. The encrypted audio data is then sent to a server.
[0820] Step 3:
[0821] The server integrates and analyzes the received location information and voice data. It analyzes the location information to confirm that the user is in a specific location (e.g., a bank ATM). It also analyzes the voice data using a generation AI model to detect specific keywords or phrases (e.g., "transfer," "hurry," "money," etc.).
[0822] Step 4:
[0823] The server compares the analysis results against a crime database to assess the risk of fraud. A high risk assessment suggests that the user is at risk of wire fraud.
[0824] Step 5:
[0825] The server generates a warning message based on the risk assessment results. The generated warning message will include appropriate content such as, "This call may be a scam. Do not transfer any money."
[0826] Step 6:
[0827] The device notifies the user of the generated warning message. The notification is provided in both audio and text formats so that the user can immediately review it.
[0828] Step 7:
[0829] Furthermore, in the event of an emergency, the server will also send a warning message to pre-registered contacts (family and friends). This will alert those around the user as well.
[0830] (Example 1)
[0831] Next, we will describe Example 1. In the following description, the data processing device 12 will be referred to as the "server," and the headset-type terminal 314 will be referred to as the "terminal."
[0832] In modern times, criminal activities such as wire fraud are on the rise, and many people are falling victim to them. Vulnerable individuals, especially the elderly, are particularly vulnerable, making prevention measures an urgent necessity. Conventional wire fraud prevention systems require users to consciously operate them, lacking convenience and immediacy. To address this challenge, a system is needed that minimizes user intervention and detects and warns of fraud risks in real time.
[0833] The identification process performed by the identification processing unit 290 of the data processing device 12 in Example 1 is realized by the following means.
[0834] In this invention, the server includes location information collection means for collecting user location information, location information transmission means for transmitting the location information to a network device, voice data collection means for collecting user voice data, voice data encryption means for encrypting the voice data, voice data transmission means for transmitting the encrypted voice data to a network device, analysis means for analyzing the location information and voice data in the network device, warning message generation means for generating a warning message based on the analysis results of the analysis means, warning notification means for notifying the generated warning message to a user terminal, notification means for notifying the generated warning message to pre-registered contacts, risk assessment means for detecting specific keywords and evaluating risk using a generation artificial intelligence model, range monitoring means for detecting when a user enters a specific range, adjustment means for adjusting the content of the warning message based on the range, and format conversion means for notifying the adjusted message in voice or text format. This makes it possible to detect the risk of fraud in real time and quickly notify warning messages while minimizing user intervention.
[0835] "Location information collection means" refers to sensors or devices used to determine the user's current location, primarily using GPS systems to acquire location information.
[0836] "Location information transmission means" refers to a function or device for transmitting acquired location information to a network device, and transmits data via a communication network.
[0837] "Voice data collection means" refers to devices such as microphones used to collect the user's voice, and to acquire the content of the call as digital voice data.
[0838] "Voice data encryption means" refers to a function or device that performs encryption processing to protect collected voice data, thereby ensuring data security.
[0839] "Voice data transmission means" refers to a function or device for transmitting encrypted voice data to a network device, and which transmits data via a communication network.
[0840] "Analysis means" refers to a function or device for analyzing received location information and audio data, and integrates and analyzes the location data and audio data.
[0841] A "warning message generation means" is a function or device that generates a message to warn the user based on the analysis results, and creates appropriate warning content.
[0842] A "warning notification means" is a function or device for notifying a generated warning message to a user terminal, thereby immediately conveying the warning to the user.
[0843] A "notification method" is a function or device that notifies pre-registered contacts of the generated warning message, thereby informing family and friends of the warning.
[0844] A "generative artificial intelligence model" is a model that analyzes audio data to detect specific keywords and perform risk assessment, using machine learning and natural language processing technologies.
[0845] A "risk assessment tool" is a function or device that analyzes audio data to assess the risk of fraud, and determines the risk level based on specific keywords or phrases.
[0846] "Range monitoring means" refers to a function or device for detecting when a user enters a specific geographical area (e.g., near a bank ATM), and continuously monitors location information.
[0847] "Adjustment means" refers to a function or device for adjusting the content of a warning message based on the user's location and voice data, thereby creating an appropriate message for the situation.
[0848] "Format conversion means" refers to a function or device for notifying warning messages in audio or text format, and for conveying warnings in a format suitable for the user.
[0849] This invention relates to a system for users to recognize and prevent the risk of wire fraud in advance. This system collects and analyzes the user's location information and voice data, and generates and sends a warning message based on the results.
[0850] Collection and transmission of location information
[0851] The terminal constantly monitors its GPS sensor to detect when the user enters a specific area (for example, near a bank ATM). This area monitoring means that when the user enters a designated zone, the terminal collects location information and periodically transmits it to a server via a location information transmission means. The GPS sensor is used in this process.
[0852] Specifically, the device collects location information such as "latitude: 35.6895, longitude: 139.6917 (within Tokyo)" and sends it to the server.
[0853] Collection and transmission of audio data
[0854] The device collects voice data in real time from the moment a call begins using voice data collection means. Voice data generated when a user makes a call is encrypted as needed using voice data encryption means and transmitted to the server via voice data transmission means. This process requires the user's consent.
[0855] Specifically, the device records the conversation as soon as the call starts and sends the audio data to the server while protecting it with AES encryption technology. An example of the conversation content collected might be something like, "Grandma, I need money right away. Can you transfer it to me today?"
[0856] Data Analysis
[0857] The server integrates and analyzes the received location information and encrypted voice data using analysis tools. First, it analyzes the location information to determine if the user is in a specific location (e.g., near a bank ATM). Next, it analyzes the voice data using a generative artificial intelligence model to detect specific keywords and phrases (e.g., "transfer," "urgent," "money," etc.), and assesses the risk of fraud using risk assessment tools. Natural language processing techniques are used in this analysis.
[0858] As an example, the prompt statements for a generative AI model are as follows:
[0859] Analyze the fraud risk from this audio data. Focus on detecting the following keywords and phrases: "transfer," "urgent," and "money." Evaluate the results as a "low," "medium," or "high" risk level to determine the risk level.
[0860] Audio data: ...
[0861] Generation and notification of warning messages
[0862] The server generates a warning message based on the analysis results using a warning message generation system. For example, it might create a message such as, "This call may be a scam. Do not transfer any money." The generated warning message is then notified to the user in both voice and text format using a warning notification system. Depending on the settings, the same warning is also sent to pre-registered contacts (family and friends) via the notification system.
[0863] Specifically, the server generates a warning message stating "This is likely a scam," converts it into voice or text format, and notifies the user and their family.
[0864] This system significantly reduces the risk of users falling victim to wire fraud. Furthermore, by immediately notifying users of warning messages, it enables quick responses, effectively preventing damage before it occurs.
[0865] The flow of the specific processing in Example 1 will be explained using Figure 11.
[0866] Step 1: Collect and transmit location information
[0867] The device constantly monitors the user's location. Specifically, it uses the GPS sensor within the device to obtain the user's current location. When the user enters a specific area (e.g., near a bank ATM), this information is automatically detected. The input is location information from the GPS sensor (e.g., latitude and longitude), and the output is the current user location information. This location information is transmitted to the server using a location information transmission device.
[0868] Step 2: Collect and transmit audio data
[0869] When a user initiates a call, the terminal detects the start of the call. The terminal uses its built-in microphone to collect audio data. The collected audio data is encrypted using an audio data encryption method. The input is the audio data of the call, and the output is encrypted audio data. Subsequently, the encrypted audio data is sent to the server using an audio data transmission method.
[0870] Step 3: Analysis of location and audio data
[0871] The server integrates and analyzes the received location information and voice data. First, it analyzes the location information to determine if the user is in a specific location (e.g., near a bank ATM). Next, it analyzes the voice data using a generative artificial intelligence model. The input consists of decomposed location information and encrypted voice data, and a risk assessment is output as a result of the analysis process. Prompt sentences are input to the generative artificial intelligence model to detect specific keywords (e.g., "transfer", "urgent", "money").
[0872] Step 4: Generating a warning message
[0873] The server generates a warning message based on the analysis results. The AI generation model automatically generates specific warning content tailored to the risk assessment from its output. The input is the result of voice data analysis, and the output is a concrete warning message. For example, a message such as "This call may be a scam. Do not transfer any money under any circumstances" might be generated.
[0874] Step 5: Notification of warning message
[0875] The server notifies the user terminal of the generated warning message. The warning message is sent in both voice and text formats using the warning notification method. The input is the generated warning message, and the output is the notification sent to the user. In addition, the same warning is sent to pre-registered contacts (family and friends). The terminal displays the received warning message to the user and provides voice notifications depending on the content of the warning.
[0876] This entire process makes it possible to detect the risk of wire fraud in real time and to quickly and appropriately notify users of warning messages.
[0877] (Application Example 1)
[0878] Next, we will explain Application Example 1. In the following explanation, the data processing device 12 will be referred to as the "server," and the headset-type terminal 314 will be referred to as the "terminal."
[0879] In recent years, there has been a significant increase in wire fraud, particularly targeting the elderly. These scams cleverly exploit the victim's psychology and typically utilize voice communication methods such as telephone calls. Traditional prevention measures, relying solely on vigilance, have been insufficient for effective prevention. Furthermore, there has been a lack of mechanisms to recognize fraud risks in real time and take immediate action. Against this backdrop, there is a need for technology that enables users to detect fraud risks in advance and prevent becoming a victim.
[0880] The specific processing performed by the specific processing unit 290 of the data processing device 12 in Application Example 1 is realized by the following means.
[0881] In this invention, the server includes location information collection means for collecting user location information, voice data collection means for collecting user voice data, and analysis processing means for analyzing using a generative AI model. This makes it possible to analyze both voice data and location information when a user approaches a specific risk area (e.g., near a bank ATM) and immediately recognize and warn of the risk of fraud.
[0882] "Location information collection means" refers to means for obtaining the user's current geographical location.
[0883] "Location information transmission means" refers to a means for transmitting collected location information to a server.
[0884] "Voice data collection means" refers to means for collecting a user's voice communication data.
[0885] "Voice data encryption means" refers to a means for encrypting collected voice data.
[0886] "Voice data transmission means" refers to a means for sending encrypted voice data to a server.
[0887] "Analysis means" refers to methods for evaluating fraud risk by analyzing location information and voice data.
[0888] A "generative AI model" is an artificial intelligence model that learns from large amounts of data and is used to perform specific tasks (e.g., fraud detection).
[0889] The "analysis processing means" is a means for comprehensively analyzing location information and audio data using a generated AI model.
[0890] A "warning message generation means" is a means for creating a warning message for the user based on the analysis results.
[0891] A "warning notification means" is a means for notifying the user terminal of the generated warning message.
[0892] "Notification method" refers to a means of notifying pre-registered contacts of the generated warning message.
[0893] "Voice notification means" refers to a means of notifying the user of the generated warning message in voice format.
[0894] A "text notification method" is a means of notifying the user of the generated warning message in text format.
[0895] Regarding embodiments for carrying out the invention, this system employs a computer-based method that combines multiple means for users to recognize and prevent the risk of wire fraud in advance.
[0896] composition
[0897] This system includes a user terminal, location information collection means, voice data collection means, generation AI model, analysis processing means, warning message generation means, warning notification means, and notification means.
[0898] User terminal: This refers to a device such as a smartphone carried by the user. The terminal is equipped with a GPS sensor and a microphone, and collects location information and voice data.
[0899] Location information collection method: Collects the user's current location. It is possible to monitor location information in real time using GPS.
[0900] Voice data collection method: Voice data during calls is collected in real time. It is collected using a microphone and encrypted for privacy protection.
[0901] Generative AI Models: These models use pre-trained AI models to detect fraudulent activity. They assess the risk of fraud based on a large amount of historical data.
[0902] Analysis processing method: Collected location information and audio data are analyzed using a generating AI model. This allows for the assessment of fraud risk.
[0903] Warning message generation means: Generates a warning message based on the analysis results. Messages are generated in both audio and text formats.
[0904] Warning notification method: The generated warning message is notified to the user's terminal. Specifically, the message is displayed on the smartphone screen and an audible warning is issued.
[0905] Notification method: In an emergency, the same warning message will be sent to pre-registered contacts (family and friends).
[0906] How to use
[0907] This system works as follows:
[0908] 1. When a user approaches a specific area:
[0909] When a user approaches a specific area, such as a bank ATM, the GPS sensor on the user's device collects location information and sends it to a server. The server analyzes this information to detect that the user is near a bank ATM.
[0910] 2. When the user initiates a call:
[0911] When a user initiates a call, the call audio is collected by a voice data collection system, encrypted in real time, and then transmitted to the server. During this process, the voice data is appropriately encrypted to protect the user's privacy.
[0912] 3. Analysis of fraud risk:
[0913] The server integrates collected location and voice data and performs analysis using a generative AI model. When specific keywords or phrases are detected, they are compared against a local crime database to assess the fraud risk.
[0914] Specific example
[0915] Example of a user approaching a bank ATM:
[0916] The terminal notifies the server that "the user is approaching a bank ATM," and specific location information (e.g., latitude: 35.6895, longitude: 139.6917) is detected.
[0917] Examples of data collection during a call:
[0918] When a user calls and says, "Grandma, I need money right away. Can you transfer it to me today?", the device records the audio, encrypts it, and sends it to the server.
[0919] Server-based analysis example:
[0920] The generative AI model detects keywords such as "bank transfer," "urgent," and "money," and compares them with local crime data to determine that the fraud risk is "high."
[0921] Example of a prompt
[0922] Examples of prompt statements are as follows:
[0923] Audio recording: "Grandma, I need money urgently. Can you transfer it to me today?"
[0924] Please evaluate the risk of fraud as part of the analysis. If you find any specific keywords or phrases, return their probability.
[0925] In this way, the system can quickly detect the risk of fraud and prevent damage by warning the user.
[0926] The flow of a specific process in Application Example 1 will be explained using Figure 12.
[0927] Step 1:
[0928] When a user approaches a specific area (e.g., near a bank ATM), the device's GPS sensor collects location information. This information is recorded on the device as numerical data such as "latitude: 35.6895, longitude: 139.6917". The input is the current location information, and the output is the collected location information. The device sends this information to the server.
[0929] Step 2:
[0930] The server analyzes the received location information to determine if the user is in a specific risk area. This analysis confirms that the user is near a bank ATM. The input is location information, and the output is the risk area determination result.
[0931] Step 3:
[0932] When a user initiates a call, the device collects the call audio in real time. The audio data is captured via the microphone and encrypted as needed. The input is the call audio, and the output is encrypted audio data. The device then sends this audio data to the server.
[0933] Step 4:
[0934] The server decrypts encrypted audio data and extracts keywords in order to analyze the received audio data using a generative AI model. The input is encrypted audio data, and the output is string data for analysis. Based on this string data, the generative AI model detects specific keywords (e.g., "transfer", "urgent", "money").
[0935] Step 5:
[0936] Based on the analyzed results, the server cross-references them with a local crime database to assess the fraud risk. The input consists of analyzed audio data, location information, and the crime database; the output is the fraud risk assessment result. If the server determines the risk is high, it generates a warning message.
[0937] Step 6:
[0938] The generated warning message is created by the server in both audio and text formats and immediately notified to the user. The input is the risk assessment results and a warning message template, and the output is the generated warning message. The terminal then informs the user of this message via audio and text.
[0939] Step 7:
[0940] The server also sends warning messages to pre-registered contacts (e.g., family and friends). The input is the generated warning message, and the output is the notification to the registered contacts. This helps users take quick action.
[0941] Through the above processing steps, this system can issue appropriate warnings to users before they are exposed to the risk of fraud, thereby preventing them from becoming victims.
[0942] Furthermore, an emotion engine that estimates the user's emotions may be incorporated. That is, the identification processing unit 290 may use the emotion identification model 59 to estimate the user's emotions and perform identification processing using the user's emotions.
[0943] Modes for carrying out the invention
[0944] This invention is a system for users to recognize and prevent the risk of wire fraud in advance. This system collects the user's location information and voice data, and generates a warning message based on the analysis results. It also utilizes an emotion engine to recognize the user's emotions, thereby improving the accuracy and effectiveness of the warning. The warning message is then sent to the user's device and pre-registered contacts.
[0945] Program Processing Description
[0946] Collection and transmission of location information
[0947] The device uses a GPS sensor to detect the user's current location. When the user enters a specific area (for example, near a bank ATM), the device collects location information and periodically sends it to the server. This allows the server to accurately determine the user's location.
[0948] Collection and transmission of audio data
[0949] When a user initiates a call, the device begins collecting the call audio in real time. This audio data is encrypted to protect privacy. The encrypted audio data is then sent to a server.
[0950] Emotional analysis using an emotion engine
[0951] The terminal or server uses the received audio data to evaluate the user's emotional state. The emotion engine recognizes the user's emotions (e.g., tension, fear, anxiety) from factors such as the tone, speed, and intonation of the voice.
[0952] Data analysis using analytical methods
[0953] The server integrates and analyzes the received location information and voice data. It analyzes the location information to confirm that the user is in a specific location (e.g., a bank ATM). It also uses a generative AI model to analyze the voice data and detect specific keywords or phrases (e.g., "transfer," "hurry," "money," etc.).
[0954] Risk assessment that takes emotional state into account
[0955] The server takes into account the results of the emotion engine analysis and compares them with the crime database to ultimately assess the risk of fraud. If the user's emotions are in a high-risk state such as tension, fear, or anxiety, the server determines that the likelihood of fraud is higher.
[0956] Generation and notification of warning messages
[0957] The server generates a warning message based on the risk assessment results. For example, it might say, "This call may be a scam. Do not transfer any money." The generated warning message is sent to the device in both voice and text formats. In emergencies, the warning message is also sent to pre-registered contacts (family and friends).
[0958] Specific example
[0959] 1. The user approaches the bank ATM.
[0960] The device detects the user's current location and notifies the server that "the user is approaching a bank ATM."
[0961] The specific location information detected is "Latitude: 35.6895, Longitude: 139.6917 (within Tokyo)".
[0962] 2. Data collection during voice calls
[0963] When a user initiates a call, the device starts recording the call, encrypting it in real time, and sending it to the server.
[0964] Example of a phone call: "Grandma, I need money urgently. Could you transfer it to me today?"
[0965] 3. Analysis using an emotion engine
[0966] The server analyzes the user's emotions from the voice data and detects tension and fear.
[0967] 4. Server-based data analysis
[0968] The server checks the user's location and confirms that the user is near a bank ATM.
[0969] The system analyzes voice data using an AI model to detect keywords such as "transfer," "urgent," and "money." Furthermore, it compares this data with local crime data to determine a high risk of fraud.
[0970] 5. Generating and notifying warning messages
[0971] The server generates a warning message stating, "This call may be a scam. Do not transfer any money under any circumstances," and notifies the device via both voice and text. Furthermore, warnings are also sent to pre-registered contacts (family and friends).
[0972] The system of this invention operates in this manner, significantly reducing the risk of users falling victim to wire fraud. Furthermore, by immediately notifying users of warning messages, it enables users to take swift action, effectively preventing damage before it occurs.
[0973] The following describes the processing flow.
[0974] Step 1:
[0975] The device uses a GPS sensor to detect the user's current location. When the user enters a specific area (for example, near a bank ATM), the device collects location information and periodically sends it to the server. This allows the server to accurately determine the user's location.
[0976] Step 2:
[0977] When a user initiates a call, the device begins collecting the call audio in real time. This audio data is encrypted to protect privacy. The encrypted audio data is then sent to a server.
[0978] Step 3:
[0979] The server integrates and analyzes the received location information and voice data. It analyzes the location information to confirm that the user is in a specific location (e.g., a bank ATM). It also uses a generative AI model to analyze the voice data and detect specific keywords or phrases (e.g., "transfer," "hurry," "money," etc.).
[0980] Step 4:
[0981] The terminal or server uses the received audio data to evaluate the user's emotional state. The emotion engine recognizes the user's emotions (e.g., tension, fear, anxiety) from factors such as tone, speed, and intonation of the voice. If the emotional state is determined to be tension, fear, or anxiety, special attention is paid to the analysis results.
[0982] Step 5:
[0983] The server takes into account the results of the emotion engine analysis and compares them with the crime database to ultimately assess the risk of fraud. If the user's emotions are in a high-risk state such as tension, fear, or anxiety, the server determines that the likelihood of fraud is higher. Based on this, the accuracy and reliability of the risk assessment are improved.
[0984] Step 6:
[0985] The server generates a warning message based on the risk assessment results. For example, it might say, "This call may be a scam. Do not transfer any money." The generated warning message is notified to the user in both audio and text format.
[0986] Step 7:
[0987] The device notifies the user of the generated warning message in both voice and text format. It also notifies the user to immediately check the warning message, encouraging a quick response. Furthermore, in emergencies, the warning message is also sent to pre-registered contacts (family and friends). This alerts those around the user as well.
[0988] Specific example:
[0989] 1. The user approaches the bank ATM.
[0990] The device detects the user's current location and notifies the server that "the user is approaching a bank ATM."
[0991] The location information detected is "Latitude: 35.6895, Longitude: 139.6917 (within Tokyo)".
[0992] 2. Data collection during voice calls
[0993] When a user initiates a call, the device begins recording the conversation. The recorded data is encrypted and sent to a server in real time.
[0994] Example of phone call content: "Grandma, I need money urgently. Could you transfer it to me today?"
[0995] 3. Data analysis by the server
[0996] The server analyzes the location information to confirm that the user is near a bank ATM.
[0997] The AI model analyzes the voice data to detect keywords such as "transfer," "urgent," and "money."
[0998] 4. Emotional analysis using an emotion engine
[0999] The server detects the user's emotions from the voice data and recognizes signs of tension or fear.
[1000] 5. Risk Assessment
[1001] The server also considers the results of the emotion engine and assesses the fraud risk as high based on the analysis results.
[1002] 6. Generating and notifying warning messages
[1003] The server generates a warning message stating, "This call may be a scam. Do not transfer any money under any circumstances," and sends it to the device in both voice and text formats.
[1004] 7. Notification to the user
[1005] The device notifies the user of warning messages in both voice and text formats, and in emergencies, it also sends warnings to pre-registered contacts.
[1006] In this way, the system of the present invention significantly reduces the risk of users falling victim to wire fraud and enables rapid countermeasures. Furthermore, by taking into account the user's emotional state, the accuracy and effectiveness of the warning are further improved.
[1007] (Example 2)
[1008] Next, we will describe Example 2. In the following description, the data processing device 12 will be referred to as the "server," and the headset-type terminal 314 will be referred to as the "terminal."
[1009] Currently, the number of victims of wire fraud is increasing, raising the risk of users being deceived and losing their valuable assets. Traditional countermeasures are insufficient because they lack the means for users to recognize the risk of fraud in advance, making it difficult to prevent victims. To solve this problem, there is a need for a system that can accurately detect the risk of wire fraud and provide users with rapid warnings.
[1010] The identification process performed by the identification processing unit 290 of the data processing device 12 in Example 2 is realized by the following means.
[1011] In this invention, the server includes emotion analysis means for analyzing the user's emotional state, analysis means for integrating and analyzing location information and voice data and detecting specific keywords or phrases using a generative AI model, and means for performing risk assessment. This enables users to detect the risk of wire fraud in advance with high accuracy and receive a warning quickly.
[1012] "Location information collection means" refers to means for obtaining the user's current location information.
[1013] "Location information transmission means" refers to a means for transmitting acquired location information to a server.
[1014] "Voice data collection means" refers to a means for collecting user voice data.
[1015] "Voice data encryption means" refers to a method of encrypting collected voice data for privacy protection.
[1016] "Voice data transmission means" refers to a means for sending encrypted voice data to a server.
[1017] "Emotion analysis means" refers to a method for evaluating a user's emotional state from voice data.
[1018] "Analysis means" refers to a means for integrating and analyzing location information and audio data, and for using a generative AI model to detect specific keywords or phrases.
[1019] A "risk assessment tool" is a means for evaluating risk based on the analysis results of an analysis tool.
[1020] A "warning message generation means" is a means for generating a warning message based on the results of a risk assessment.
[1021] A "warning notification means" is a means for notifying the user terminal of the generated warning message.
[1022] "Notification method" refers to a means of notifying pre-registered contacts of the generated warning message.
[1023] "Voice notification means" refers to a means of notifying the user of the generated warning message in voice format.
[1024] A "text notification method" is a means of notifying the user of the generated warning message in text format.
[1025] A "generative AI model" is an artificial intelligence model used to analyze audio data and detect specific keywords or phrases.
[1026] This invention is a system for users to recognize and prevent the risk of wire fraud in advance. The system includes means for collecting the user's location information and voice data and generating a warning message based on the analysis results.
[1027] Hardware configuration
[1028] The system utilizes mobile devices such as smartphones equipped with a GPS sensor, microphone, and data communication capabilities. The GPS sensor collects the user's current location, and the microphone collects call audio. The collected data is then transmitted to a server via the internet.
[1029] The server utilizes high-performance computers and cloud services to provide the processing power needed for data analysis and sentiment analysis. Specific software used includes IBM Watson's sentiment analysis API and GPT-3 generative AI models.
[1030] Software Configuration
[1031] The device has the following software features:
[1032] 1. Location information collection method: The user's current location is obtained using the built-in GPS sensor, and location data is generated.
[1033] 2. Voice data collection method: Voice data is collected in real time using a microphone during a call.
[1034] 3. Audio data encryption method: The collected audio data is encrypted using AES-256.
[1035] 4. Location information transmission means: The collected location information is transmitted to the server.
[1036] 5. Means of transmitting audio data: Encrypted audio data is sent to the server.
[1037] The server has the following software features:
[1038] 1. Emotion analysis method: The received audio data is analyzed using IBM Watson's emotion analysis API to evaluate the user's emotional state.
[1039] 2. Analysis method: The received location information and audio data are integrated, and specific keywords and phrases are detected using the GPT-3 generation AI model.
[1040] 3. Risk assessment method: The risk of wire fraud will be finalized based on the analysis results and sentiment analysis results.
[1041] 4. Warning message generation means: Generates a warning message based on the risk assessment results.
[1042] 5. Warning notification means: The generated warning message is sent to the user's terminal and to pre-registered contacts.
[1043] Specific example
[1044] 1. The user approaches the bank ATM.
[1045] The device uses a GPS sensor to obtain its current location and sends the location information to the server when it detects that "the user is approaching a bank ATM." An example of location data is "Latitude: 35.6895, Longitude: 139.6917 (within Tokyo)."
[1046] 2. Data collection during voice calls
[1047] When a user initiates a call, the device collects the call audio, encrypts it with AES-256, and sends it to the server in real time. Specifically, the call content might be something like, "Grandma, I need money urgently. Can you transfer it to me today?" The audio recording and encryption process is performed on the device itself.
[1048] 3. Analysis using an emotion engine
[1049] The server decodes the received audio data and uses an emotion analysis API to assess the user's emotion as "tension." The analysis takes into account the tone, speed, and intonation of the voice.
[1050] 4. Server-based data analysis
[1051] The server uses location data to confirm that the user is near a bank ATM. Simultaneously, it analyzes the voice data using an AI model (e.g., GPT-3) to detect keywords such as "transfer," "urgent," and "money." This then triggers a voice-to-text conversion process, highlighting specific keywords.
[1052] 5. Risk assessment that takes emotional state into consideration
[1053] The server determines a high risk of fraud by combining voice analysis results and sentiment analysis results. This includes matching against past crime data and calculating the degree of match.
[1054] 6. Generating and notifying warning messages
[1055] The server generates a warning message stating, "This call may be a scam. Do not transfer any money under any circumstances," and also generates a voice message using a speech synthesis API. The generated warning message is sent to the device and, if necessary, also notified to registered contacts via SMS or email.
[1056] Example of a prompt
[1057] "Please explain the process for a program that notifies a user when they approach a bank ATM. Also, please explain the procedure for collecting voice data from a user during a call and assessing the risk of fraud."
[1058] The flow of the specific processing in Example 2 will be explained using Figure 13.
[1059] Step 1: Collect and transmit location information
[1060] The device uses its built-in GPS sensor to acquire the user's current location at regular intervals. Specifically, it acquires location information every minute and generates latitude and longitude data (input: GPS sensor data, output: location information data). When the user enters a specific area (for example, near a bank ATM), it sends that location information to the server (data processing: location information detection, data calculation: range determination, output: location information transmission).
[1061] Step 2: Collect and transmit audio data
[1062] When a user initiates a call, the device uses its microphone to collect the call audio in real time (input: call audio, output: audio data). The collected audio data is immediately encrypted using AES-256 (data processing: audio data encryption), and the encrypted audio data is sent to the server in real time (output: encrypted audio data transmission).
[1063] Step 3: Emotional analysis using the emotion engine
[1064] The server decrypts the received encrypted audio data (input: encrypted audio data, output: decrypted audio data) and sends the audio data to the emotion analysis API. The emotion analysis API analyzes the tone, speed, and intonation of the voice and evaluates the user's emotional state (tension, fear, anxiety, etc.) (data calculation: emotion analysis, output: emotional state data).
[1065] Step 4: Data analysis using analytical tools
[1066] The server integrates and analyzes emotional state data with received location and audio data (input: location data, decoded audio data, emotional state data). It uses a GPT-3 generative AI model to detect specific keywords or phrases (e.g., "transfer", "hurry", "money") from the audio data (data calculation: keyword detection, output: keyword data). It uses location information to determine if the user is in a specific location (e.g., a bank ATM) (data calculation: location confirmation, output: location confirmation data).
[1067] Step 5: Risk assessment considering emotional state
[1068] The server integrates keyword data and location data from the analysis tools, and also considers emotional state data to assess fraud risk (Input: Keyword data, emotional state data, location data; Output: Risk assessment data). For example, if the user's emotions are in a high-risk state such as tension or fear, and specific keywords are included, the risk is judged to be high (Data calculation: Risk assessment).
[1069] Step 6: Generate and send warning messages
[1070] The server generates a warning message based on risk assessment data (input: risk assessment data, output: warning message). For example, it generates a message stating, "This call may be a scam. Do not transfer any money under any circumstances." (data processing: warning message generation). The generated warning message is converted into voice format using a speech synthesis API (output: voice warning message) and also sent to the terminal in text format (output: text warning message). Furthermore, in emergencies, notifications are also sent to pre-registered contacts (family and friends) (output: contact notification data).
[1071] (Application Example 2)
[1072] Next, we will explain application example 2. In the following explanation, the data processing device 12 will be referred to as the "server," and the headset-type terminal 314 will be referred to as the "terminal."
[1073] Fraudulent activities, including wire transfer scams, are increasing, particularly targeting the elderly. Therefore, there is a need to detect the risk of fraud early and prevent it before it occurs. However, current systems lack the means to effectively analyze users' emotional states and real-time call content, making it difficult to accurately determine situations with a high risk of fraud. Therefore, it is necessary to provide a system that can quickly and accurately detect situations with a high risk of fraud and issue warnings to users, their relatives, and friends.
[1074] The specific processing performed by the specific processing unit 290 of the data processing device 12 in Application Example 2 is realized by the following means.
[1075] In this invention, the server includes location information collection means for collecting user location information, voice data collection means for collecting user voice data, emotion analysis means for analyzing the user's emotional state based on the voice data, and keyword detection means for detecting specific keywords. This makes it possible to detect fraud risks such as wire fraud with high accuracy and speed, and to issue warnings to the user and their family and friends.
[1076] "Location information collection means" refers to devices and software used to determine the user's current location.
[1077] "Location information transmission means" refers to devices or software used to transmit collected user location information to a server.
[1078] "Voice data collection means" refers to devices or software used to collect voices emitted by users.
[1079] "Voice data encryption means" refers to devices or software used to encrypt collected voice data so that it cannot be deciphered by third parties.
[1080] "Voice data transmission means" refers to devices or software used to transmit encrypted voice data to a server.
[1081] "Analysis means" refers to devices or software used to analyze location information and voice data collected on a server to assess fraud risk.
[1082] "Warning message generation means" refers to a device or software that automatically generates warning messages based on the analysis results of the analysis means.
[1083] "Warning notification means" refers to devices or software that notify the user's terminal of the generated warning message.
[1084] "Notification means" refers to devices or software used to notify pre-registered contacts of generated warning messages.
[1085] "Emotional analysis means" refers to devices and software that analyze a user's emotional state in real time based on voice data.
[1086] "Keyword detection means" refers to devices or software used to detect specific keywords from audio data.
[1087] This invention is a system for users to recognize and prevent the risk of wire fraud in advance. The system includes means for identifying the user's current location and transmitting that information to a server, and means for collecting, encrypting, and transmitting the user's voice data. Furthermore, the server includes means for analyzing this information and generating a warning message to notify the user's terminal and pre-registered contacts.
[1088] Collection and transmission of location information
[1089] The device is equipped with a GPS sensor that constantly monitors the user's current location and acquires data using location information collection methods. When the user enters a specific area (for example, near a bank ATM), location information is sent to the server using location information transmission methods. This allows the server to determine the user's precise location.
[1090] Collection and transmission of audio data
[1091] When a user starts a phone call, the call content is collected in real time by an audio data collection means and encrypted by an audio data encryption means. The encrypted audio data is then transmitted to the server using an audio data transmission means.
[1092] Sentiment analysis and keyword detection
[1093] The server uses emotion analysis tools to analyze the received audio data. This evaluates the user's emotions (e.g., tension, fear, anxiety) from factors such as tone, speed, and intonation. Furthermore, keyword detection tools detect specific keywords within the audio data (e.g., "transfer," "hurry," "money," etc.).
[1094] Risk assessment and warning message generation
[1095] The server integrates location information, voice data, sentiment analysis results, and keyword detection results to assess the fraud risk. If the risk is determined to be high, a warning message is generated by the warning message generation mechanism. For example, a message such as, "This call may be a scam. Do not transfer any money."
[1096] Warning message notification
[1097] The warning notification system will send the generated warning message to the user's device in both audio and text format. In emergencies, the same warning message will also be sent to pre-registered contacts (family and friends) using the notification system.
[1098] Specific example
[1099] 1. The user approaches the bank ATM.
[1100] The device detects the user's current location and notifies the server that "the user is approaching a bank ATM." The specific location information detected is "Latitude: 35.6895, Longitude: 139.6917 (within Tokyo)."
[1101] 2. Data collection during voice calls
[1102] When a user initiates a call, the device starts recording the conversation, encrypting it in real time, and sending it to the server. Example of a call: "Grandma, I need money urgently. Can you transfer it to me today?"
[1103] 3. Analysis using an emotion engine
[1104] The server analyzes the user's emotions from the voice data and detects tension and fear.
[1105] 4. Server-based data analysis
[1106] The server checks the user's location to confirm they are near a bank ATM. An AI model analyzes the voice data to detect keywords such as "transfer," "urgent," and "money." Furthermore, it cross-references this data with local crime statistics to determine a high risk of fraud.
[1107] 5. Generating and notifying warning messages
[1108] The server generates a warning message stating, "This call may be a scam. Do not transfer any money under any circumstances," and notifies the device via both voice and text. Furthermore, warnings are also sent to pre-registered contacts (family and friends).
[1109] Example of a prompt
[1110] text
[1111] Evaluate the user's emotional state and generate an emergency fraud warning message. For example, use this if you receive audio data such as "Location: 35.6895, 139.6917" and "Grandma, I need money right away. Can you transfer it to me today?". If the emotional score is -0.5 or lower and contains fraud-related keywords such as "transfer" and "money", generate a warning message stating, "This call may be a scam. Do not transfer any money."
[1112] The above describes the embodiments for carrying out this invention.
[1113] The flow of a specific process in Application Example 2 will be explained using Figure 14.
[1114] Step 1:
[1115] Collecting and transmitting user location information
[1116] The device uses a GPS sensor to collect the user's location information. The collected location information (e.g., latitude 35.6895, longitude 139.6917) is transmitted to the server using a location information transmission means.
[1117] Input: Location information from GPS sensor
[1118] Output: User location information sent to the server
[1119] Step 2:
[1120] Collection and transmission of audio data
[1121] When a user initiates a call, the device collects audio in real time using an audio data collection device. The collected audio data is encrypted using an audio data encryption device and transmitted to a server using an audio data transmission device.
[1122] Input: Audio data from a call
[1123] Output: Encrypted audio data sent to the server
[1124] Step 3:
[1125] Location information analysis
[1126] The server analyzes the received location information using an analysis tool to determine whether the user is in a specific risk area (e.g., near a bank ATM).
[1127] Input: Sent location information
[1128] Output: Determination result of whether the user is in a risk area.
[1129] Step 4:
[1130] Sentiment analysis of voice data
[1131] The server analyzes the audio data using emotion analysis tools. Specifically, it evaluates the user's emotions (e.g., tension, fear, anxiety) based on factors such as tone, speed, and intonation of the voice.
[1132] Input: Decrypted audio data
[1133] Output: User sentiment score
[1134] Step 5:
[1135] Detection of specific keywords
[1136] The server uses keyword detection means to detect specific keywords (e.g., "transfer", "hurry", "money") from the audio data.
[1137] Input: Decrypted audio data
[1138] Output: List of detected keywords
[1139] Step 6:
[1140] Risk assessment
[1141] The server integrates location data analysis results, user sentiment scores, and detected keyword lists to assess fraud risk. If a high risk is determined, a warning message is generated using a warning message generation mechanism.
[1142] Input: Location analysis results, sentiment score, keyword list
[1143] Output: Fraud risk assessment results
[1144] Step 7:
[1145] Generation and notification of warning messages
[1146] The server sends the generated warning message to the user's terminal using a warning notification system. Simultaneously, it also sends the same warning message to pre-registered contacts using the notification system.
[1147] Input: Fraud risk assessment results
[1148] Output: Warning message sent to the user's device and registered contacts.
[1149] These steps enable the creation of a system that effectively detects fraud risks and warns users and their associates.
[1150] The specific processing unit 290 transmits the result of the specific processing to the headset terminal 314. In the headset terminal 314, the control unit 46A causes the speaker 240 and display 343 to output the result of the specific processing. The microphone 238 acquires audio indicating user input for the result of the specific processing. The control unit 46A transmits the audio data indicating user input acquired by the microphone 238 to the data processing unit 12. In the data processing unit 12, the specific processing unit 290 acquires the audio data.
[1151] Data generation model 58 is a type of so-called generative AI (Artificial Intelligence). One example of data generation model 58 is ChatGPT (Internet search<URL: https: / / openai.com / blog / chatgpt> ), Gemini (Internet search) <url: https: gemini.google.com ?hl="ja">Examples of generative AI include the following. The data generation model 58 is obtained by performing deep learning on a neural network. The data generation model 58 is input with prompts containing instructions, and with inference data such as audio data representing speech, text data representing text, and image data representing images. The data generation model 58 infers from the input inference data according to the instructions indicated by the prompts, and outputs the inference results in data formats such as audio data and text data. Here, inference refers to, for example, analysis, classification, prediction, and / or summarization.
[1152] In the above embodiment, an example was given in which specific processing is performed by the data processing device 12, but the technology of this disclosure is not limited thereto, and specific processing may also be performed by the headset terminal 314.
[1153] [Fourth Embodiment]
[1154] Figure 7 shows an example of the configuration of the data processing system 410 according to the fourth embodiment.
[1155] As shown in Figure 7, the data processing system 410 includes a data processing device 12 and a robot 414. An example of the data processing device 12 is a server.
[1156] The data processing device 12 comprises a computer 22, a database 24, and a communication interface 26. The computer 22 is an example of a "computer" related to the technology of this disclosure. The computer 22 comprises a processor 28, RAM 30, and storage 32. The processor 28, RAM 30, and storage 32 are connected to a bus 34. The database 24 and the communication interface 26 are also connected to the bus 34. The communication interface 26 is connected to a network 54. An example of the network 54 is a WAN (Wide Area Network) and / or a LAN (Local Area Network).
[1157] The robot 414 includes a computer 36, a microphone 238, a speaker 240, a camera 42, a communication interface 44, and a controlled object 443. The computer 36 includes a processor 46, RAM 48, and storage 50. The processor 46, RAM 48, and storage 50 are connected to a bus 52. The microphone 238, speaker 240, camera 42, and controlled object 443 are also connected to the bus 52.
[1158] The microphone 238 receives voice signals from the user 20 and receives instructions from the user 20. The microphone 238 captures the voice signals from the user 20, converts the captured voice into audio data, and outputs it to the processor 46. The speaker 240 outputs audio according to the instructions from the processor 46.
[1159] Camera 42 is a small digital camera equipped with an optical system including a lens, aperture, and shutter, and an image sensor such as a CMOS (Complementary Metal-Oxide-Semiconductor) image sensor or a CCD (Charge Coupled Device) image sensor, and captures images of the area around the user 20 (for example, an imaging range defined by a field of view equivalent to the width of a typical healthy person's field of vision).
[1160] Communication interface 44 is connected to network 54. Communication interfaces 44 and 26 are responsible for the exchange of various information between processor 46 and processor 28 via network 54. The exchange of various information between processor 46 and processor 28 using communication interfaces 44 and 26 is performed in a secure manner.
[1161] The controlled object 443 includes a display device, LEDs in the eyes, and motors that drive the arms, hands, and feet. The posture and gestures of the robot 414 are controlled by controlling the motors of the arms, hands, and feet. Some of the robot 414's emotions can be expressed by controlling these motors. Furthermore, the robot 414's facial expressions can also be expressed by controlling the illumination state of the LEDs in its eyes.
[1162] Figure 8 shows an example of the main functions of the data processing device 12 and the robot 414. As shown in Figure 8, the data processing device 12 performs specific processing using the processor 28. The storage 32 stores the specific processing program 56.
[1163] The specific processing program 56 is an example of a "program" relating to the technology of this disclosure. The processor 28 reads the specific processing program 56 from the storage 32 and executes the read specific processing program 56 on the RAM 30. The specific processing is realized by the processor 28 operating as a specific processing unit 290 in accordance with the specific processing program 56 executed on the RAM 30.
[1164] The storage 32 stores the data generation model 58 and the emotion identification model 59. The data generation model 58 and the emotion identification model 59 are used by the identification processing unit 290.
[1165] In robot 414, the processor 46 performs the reception output processing. The storage 50 stores the reception output program 60. The processor 46 reads the reception output program 60 from the storage 50 and executes the read reception output program 60 on the RAM 48. The reception output processing is realized by the processor 46 operating as a control unit 46A according to the reception output program 60 executed on the RAM 48.
[1166] Next, the specific processing performed by the specific processing unit 290 of the data processing device 12 will be described. In the following description, the data processing device 12 will be referred to as the "server" and the robot 414 as the "terminal".
[1167] Modes for carrying out the invention
[1168] This invention is a system for users to recognize and prevent the risk of wire fraud in advance. This system collects the user's location information and voice data, generates a warning message based on the analysis results, and notifies the user's device and pre-registered contacts of the warning message.
[1169] Program Processing Description
[1170] Collection and transmission of location information
[1171] The device constantly monitors its GPS sensor to detect when the user enters a specific area (e.g., near a bank ATM). When the user enters the designated zone, the device collects location information and periodically sends it to the server. This allows the server to accurately determine the user's location.
[1172] Collection and transmission of audio data
[1173] The device collects voice data in real time from the moment the call begins. The voice data generated when the user makes a call is encrypted as needed and sent to the server. This process requires the user's consent, so it is important that consent is obtained in advance.
[1174] Data analysis using analytical methods
[1175] The server integrates and analyzes the received location information and voice data. It analyzes the location information to determine if the user is in a specific location (e.g., a bank ATM) and uses a generative AI model to analyze the voice data. It detects specific keywords and phrases (e.g., "transfer," "urgent," "money," etc.) and assesses the risk of fraud by matching them against a local crime database.
[1176] Generation and notification of warning messages
[1177] The server generates a warning message based on the analysis results. For example, it might create a message such as, "This call may be a scam. Do not transfer any money." The generated warning message is notified to the user in both voice and text formats. In emergencies, the same warning is also sent to pre-registered contacts (family and friends).
[1178] Specific example
[1179] 1. The user approaches the bank ATM.
[1180] The device detects the user's current location and notifies the server that "the user is approaching a bank ATM."
[1181] The specific location information detected is "Latitude: 35.6895, Longitude: 139.6917 (within Tokyo)".
[1182] 2. Data collection during voice calls
[1183] When a user initiates a call, the device starts recording the call, encrypting it in real time, and sending it to the server.
[1184] Example of a phone call: "Grandma, I need money urgently. Could you transfer it to me today?"
[1185] 3. Server-based analysis
[1186] The server checks the user's location and confirms that the user is near a bank ATM.
[1187] The system analyzes the voice data to detect keywords such as "transfer," "urgent," and "money." Furthermore, it cross-references this data with local crime statistics to determine that the risk of fraud is "high."
[1188] 4. Generating and notifying warning messages
[1189] The server generates a warning message saying, "This call may be a scam. Do not transfer any money under any circumstances," and notifies the device via both voice and text.
[1190] Furthermore, a warning will be sent to pre-registered contacts (family and friends).
[1191] The system of this invention operates in this manner, significantly reducing the risk of users falling victim to wire fraud. Furthermore, by immediately notifying users of warning messages, it enables users to take swift action, effectively preventing damage before it occurs.
[1192] The following describes the processing flow.
[1193] Step 1:
[1194] The device uses a GPS sensor to detect the user's current location. When the user enters a specific area (for example, near a bank ATM), the device collects location information. Subsequently, it sends the collected location information to a server, which then accurately determines the user's current location.
[1195] Step 2:
[1196] When a user initiates a call, the device begins collecting the call audio in real time. This audio data is encrypted to protect privacy. The encrypted audio data is then sent to a server.
[1197] Step 3:
[1198] The server integrates and analyzes the received location information and voice data. It analyzes the location information to confirm that the user is in a specific location (e.g., a bank ATM). It also analyzes the voice data using a generation AI model to detect specific keywords or phrases (e.g., "transfer," "hurry," "money," etc.).
[1199] Step 4:
[1200] The server compares the analysis results against a crime database to assess the risk of fraud. A high risk assessment suggests that the user is at risk of wire fraud.
[1201] Step 5:
[1202] The server generates a warning message based on the risk assessment results. The generated warning message will include appropriate content such as, "This call may be a scam. Do not transfer any money."
[1203] Step 6:
[1204] The device notifies the user of the generated warning message. The notification is provided in both audio and text formats so that the user can immediately review it.
[1205] Step 7:
[1206] Furthermore, in the event of an emergency, the server will also send a warning message to pre-registered contacts (family and friends). This will alert those around the user as well.
[1207] (Example 1)
[1208] Next, we will describe Example 1. In the following description, the data processing device 12 will be referred to as the "server" and the robot 414 as the "terminal".
[1209] In modern times, criminal activities such as wire fraud are on the rise, and many people are falling victim to them. Vulnerable individuals, especially the elderly, are particularly vulnerable, making prevention measures an urgent necessity. Conventional wire fraud prevention systems require users to consciously operate them, lacking convenience and immediacy. To address this challenge, a system is needed that minimizes user intervention and detects and warns of fraud risks in real time.
[1210] The identification process performed by the identification processing unit 290 of the data processing device 12 in Example 1 is realized by the following means.
[1211] In this invention, the server includes location information collection means for collecting user location information, location information transmission means for transmitting the location information to a network device, voice data collection means for collecting user voice data, voice data encryption means for encrypting the voice data, voice data transmission means for transmitting the encrypted voice data to a network device, analysis means for analyzing the location information and voice data in the network device, warning message generation means for generating a warning message based on the analysis results of the analysis means, warning notification means for notifying the generated warning message to a user terminal, notification means for notifying the generated warning message to pre-registered contacts, risk assessment means for detecting specific keywords and evaluating risk using a generation artificial intelligence model, range monitoring means for detecting when a user enters a specific range, adjustment means for adjusting the content of the warning message based on the range, and format conversion means for notifying the adjusted message in voice or text format. This makes it possible to detect the risk of fraud in real time and quickly notify warning messages while minimizing user intervention.
[1212] "Location information collection means" refers to sensors or devices used to determine the user's current location, primarily using GPS systems to acquire location information.
[1213] "Location information transmission means" refers to a function or device for transmitting acquired location information to a network device, and transmits data via a communication network.
[1214] "Voice data collection means" refers to devices such as microphones used to collect the user's voice, and to acquire the content of the call as digital voice data.
[1215] "Voice data encryption means" refers to a function or device that performs encryption processing to protect collected voice data, thereby ensuring data security.
[1216] "Voice data transmission means" refers to a function or device for transmitting encrypted voice data to a network device, and which transmits data via a communication network.
[1217] "Analysis means" refers to a function or device for analyzing received location information and audio data, and integrates and analyzes the location data and audio data.
[1218] A "warning message generation means" is a function or device that generates a message to warn the user based on the analysis results, and creates appropriate warning content.
[1219] A "warning notification means" is a function or device for notifying a generated warning message to a user terminal, thereby immediately conveying the warning to the user.
[1220] A "notification method" is a function or device that notifies pre-registered contacts of the generated warning message, thereby informing family and friends of the warning.
[1221] A "generative artificial intelligence model" is a model that analyzes audio data to detect specific keywords and perform risk assessment, using machine learning and natural language processing technologies.
[1222] A "risk assessment tool" is a function or device that analyzes audio data to assess the risk of fraud, and determines the risk level based on specific keywords or phrases.
[1223] "Range monitoring means" refers to a function or device for detecting when a user enters a specific geographical area (e.g., near a bank ATM), and continuously monitors location information.
[1224] "Adjustment means" refers to a function or device for adjusting the content of a warning message based on the user's location and voice data, thereby creating an appropriate message for the situation.
[1225] "Format conversion means" refers to a function or device for notifying warning messages in audio or text format, and for conveying warnings in a format suitable for the user.
[1226] This invention relates to a system for users to recognize and prevent the risk of wire fraud in advance. This system collects and analyzes the user's location information and voice data, and generates and sends a warning message based on the results.
[1227] Collection and transmission of location information
[1228] The terminal constantly monitors its GPS sensor to detect when the user enters a specific area (for example, near a bank ATM). This area monitoring means that when the user enters a designated zone, the terminal collects location information and periodically transmits it to a server via a location information transmission means. The GPS sensor is used in this process.
[1229] Specifically, the device collects location information such as "latitude: 35.6895, longitude: 139.6917 (within Tokyo)" and sends it to the server.
[1230] Collection and transmission of audio data
[1231] The device collects voice data in real time from the moment a call begins using voice data collection means. Voice data generated when a user makes a call is encrypted as needed using voice data encryption means and transmitted to the server via voice data transmission means. This process requires the user's consent.
[1232] Specifically, the device records the conversation as soon as the call starts and sends the audio data to the server while protecting it with AES encryption technology. An example of the conversation content collected might be something like, "Grandma, I need money right away. Can you transfer it to me today?"
[1233] Data analysis
[1234] The server integrates and analyzes the received location information and encrypted voice data using analysis tools. First, it analyzes the location information to determine if the user is in a specific location (e.g., near a bank ATM). Next, it analyzes the voice data using a generative artificial intelligence model to detect specific keywords and phrases (e.g., "transfer," "urgent," "money," etc.), and assesses the risk of fraud using risk assessment tools. Natural language processing techniques are used in this analysis.
[1235] As an example, the prompt statements for a generative AI model are as follows:
[1236] Analyze the fraud risk from this audio data. Focus on detecting the following keywords and phrases: "transfer," "urgent," and "money." Evaluate the results as a "low," "medium," or "high" risk level to determine the risk level.
[1237] Audio data: ...
[1238] Generation and notification of warning messages
[1239] The server generates a warning message based on the analysis results using a warning message generation system. For example, it might create a message such as, "This call may be a scam. Do not transfer any money." The generated warning message is then notified to the user in both voice and text format using a warning notification system. Depending on the settings, the same warning is also sent to pre-registered contacts (family and friends) via the notification system.
[1240] Specifically, the server generates a warning message stating "This is likely a scam," converts it into voice or text format, and notifies the user and their family.
[1241] This system significantly reduces the risk of users falling victim to wire fraud. Furthermore, by immediately notifying users of warning messages, it enables quick responses, effectively preventing damage before it occurs.
[1242] The flow of the specific processing in Example 1 will be explained using Figure 11.
[1243] Step 1: Collect and transmit location information
[1244] The device constantly monitors the user's location. Specifically, it uses the GPS sensor within the device to obtain the user's current location. When the user enters a specific area (e.g., near a bank ATM), this information is automatically detected. The input is location information from the GPS sensor (e.g., latitude and longitude), and the output is the current user location information. This location information is transmitted to the server using a location information transmission device.
[1245] Step 2: Collect and transmit audio data
[1246] When a user initiates a call, the terminal detects the start of the call. The terminal uses its built-in microphone to collect audio data. The collected audio data is encrypted using an audio data encryption method. The input is the audio data of the call, and the output is encrypted audio data. Subsequently, the encrypted audio data is sent to the server using an audio data transmission method.
[1247] Step 3: Analysis of location and audio data
[1248] The server integrates and analyzes the received location information and voice data. First, it analyzes the location information to determine if the user is in a specific location (e.g., near a bank ATM). Next, it analyzes the voice data using a generative artificial intelligence model. The input consists of decomposed location information and encrypted voice data, and a risk assessment is output as a result of the analysis process. Prompt sentences are input to the generative artificial intelligence model to detect specific keywords (e.g., "transfer", "urgent", "money").
[1249] Step 4: Generating a warning message
[1250] The server generates a warning message based on the analysis results. The AI generation model automatically generates specific warning content tailored to the risk assessment from its output. The input is the result of voice data analysis, and the output is a concrete warning message. For example, a message such as "This call may be a scam. Do not transfer any money under any circumstances" might be generated.
[1251] Step 5: Notification of warning message
[1252] The server notifies the user terminal of the generated warning message. The warning message is sent in both voice and text formats using the warning notification method. The input is the generated warning message, and the output is the notification sent to the user. In addition, the same warning is sent to pre-registered contacts (family and friends). The terminal displays the received warning message to the user and provides voice notifications depending on the content of the warning.
[1253] This entire process makes it possible to detect the risk of wire fraud in real time and to quickly and appropriately notify users of warning messages.
[1254] (Application Example 1)
[1255] Next, we will explain Application Example 1. In the following explanation, the data processing device 12 will be referred to as the "server" and the robot 414 as the "terminal".
[1256] In recent years, there has been a significant increase in wire fraud, particularly targeting the elderly. These scams cleverly exploit the victim's psychology and typically utilize voice communication methods such as telephone calls. Traditional prevention measures, relying solely on vigilance, have been insufficient for effective prevention. Furthermore, there has been a lack of mechanisms to recognize fraud risks in real time and take immediate action. Against this backdrop, there is a need for technology that enables users to detect fraud risks in advance and prevent becoming a victim.
[1257] The specific processing performed by the specific processing unit 290 of the data processing device 12 in Application Example 1 is realized by the following means.
[1258] In this invention, the server includes location information collection means for collecting user location information, voice data collection means for collecting user voice data, and analysis processing means for analyzing using a generative AI model. This makes it possible to analyze both voice data and location information when a user approaches a specific risk area (e.g., near a bank ATM) and immediately recognize and warn of the risk of fraud.
[1259] "Location information collection means" refers to means for obtaining the user's current geographical location.
[1260] "Location information transmission means" refers to a means for transmitting collected location information to a server.
[1261] "Voice data collection means" refers to means for collecting a user's voice communication data.
[1262] "Voice data encryption means" refers to a means for encrypting collected voice data.
[1263] "Voice data transmission means" refers to a means for sending encrypted voice data to a server.
[1264] "Analysis means" refers to methods for evaluating fraud risk by analyzing location information and voice data.
[1265] A "generative AI model" is an artificial intelligence model that learns from large amounts of data and is used to perform specific tasks (e.g., fraud detection).
[1266] The "analysis processing means" is a means for comprehensively analyzing location information and audio data using a generated AI model.
[1267] A "warning message generation means" is a means for creating a warning message for the user based on the analysis results.
[1268] A "warning notification means" is a means for notifying the user terminal of the generated warning message.
[1269] "Notification method" refers to a means of notifying pre-registered contacts of the generated warning message.
[1270] "Voice notification means" refers to a means of notifying the user of the generated warning message in voice format.
[1271] A "text notification method" is a means of notifying the user of the generated warning message in text format.
[1272] Regarding embodiments for carrying out the invention, this system employs a computer-based method that combines multiple means for users to recognize and prevent the risk of wire fraud in advance.
[1273] composition
[1274] This system includes a user terminal, location information collection means, voice data collection means, generation AI model, analysis processing means, warning message generation means, warning notification means, and notification means.
[1275] User terminal: This refers to a device such as a smartphone carried by the user. The terminal is equipped with a GPS sensor and a microphone, and collects location information and voice data.
[1276] Location information collection method: Collects the user's current location. It is possible to monitor location information in real time using GPS.
[1277] Voice data collection method: Voice data during calls is collected in real time. It is collected using a microphone and encrypted for privacy protection.
[1278] Generative AI Models: These models use pre-trained AI models to detect fraudulent activity. They assess the risk of fraud based on a large amount of historical data.
[1279] Analysis processing method: Collected location information and audio data are analyzed using a generating AI model. This allows for the assessment of fraud risk.
[1280] Warning message generation means: Generates a warning message based on the analysis results. Messages are generated in both audio and text formats.
[1281] Warning notification method: The generated warning message is notified to the user's terminal. Specifically, the message is displayed on the smartphone screen and an audible warning is issued.
[1282] Notification method: In an emergency, the same warning message will be sent to pre-registered contacts (family and friends).
[1283] How to use
[1284] This system works as follows:
[1285] 1. When a user approaches a specific area:
[1286] When a user approaches a specific area, such as a bank ATM, the GPS sensor on the user's device collects location information and sends it to a server. The server analyzes this information to detect that the user is near a bank ATM.
[1287] 2. When the user initiates a call:
[1288] When a user initiates a call, the call audio is collected by a voice data collection system, encrypted in real time, and then transmitted to the server. During this process, the voice data is appropriately encrypted to protect the user's privacy.
[1289] 3. Analysis of fraud risk:
[1290] The server integrates collected location and voice data and performs analysis using a generative AI model. When specific keywords or phrases are detected, they are compared against a local crime database to assess the fraud risk.
[1291] Specific example
[1292] Example of a user approaching a bank ATM:
[1293] The terminal notifies the server that "the user is approaching a bank ATM," and specific location information (e.g., latitude: 35.6895, longitude: 139.6917) is detected.
[1294] Examples of data collection during a call:
[1295] When a user calls and says, "Grandma, I need money right away. Can you transfer it to me today?", the device records the audio, encrypts it, and sends it to the server.
[1296] Server-based analysis example:
[1297] The generative AI model detects keywords such as "bank transfer," "urgent," and "money," and compares them with local crime data to determine that the fraud risk is "high."
[1298] Example of a prompt
[1299] Examples of prompt statements are as follows:
[1300] Audio recording: "Grandma, I need money urgently. Can you transfer it to me today?"
[1301] Please evaluate the risk of fraud as part of the analysis. If you find any specific keywords or phrases, return their probability.
[1302] In this way, the system can quickly detect the risk of fraud and prevent damage by warning the user.
[1303] The flow of a specific process in Application Example 1 will be explained using Figure 12.
[1304] Step 1:
[1305] When a user approaches a specific area (e.g., near a bank ATM), the device's GPS sensor collects location information. This information is recorded on the device as numerical data such as "latitude: 35.6895, longitude: 139.6917". The input is the current location information, and the output is the collected location information. The device sends this information to the server.
[1306] Step 2:
[1307] The server analyzes the received location information to determine if the user is in a specific risk area. This analysis confirms that the user is near a bank ATM. The input is location information, and the output is the risk area determination result.
[1308] Step 3:
[1309] When a user initiates a call, the device collects the call audio in real time. The audio data is captured via the microphone and encrypted as needed. The input is the call audio, and the output is encrypted audio data. The device then sends this audio data to the server.
[1310] Step 4:
[1311] The server decrypts encrypted audio data and extracts keywords in order to analyze the received audio data using a generative AI model. The input is encrypted audio data, and the output is string data for analysis. Based on this string data, the generative AI model detects specific keywords (e.g., "transfer", "urgent", "money").
[1312] Step 5:
[1313] Based on the analyzed results, the server cross-references them with a local crime database to assess the fraud risk. The input consists of analyzed audio data, location information, and the crime database; the output is the fraud risk assessment result. If the server determines the risk is high, it generates a warning message.
[1314] Step 6:
[1315] The generated warning message is created by the server in both audio and text formats and immediately notified to the user. The input is the risk assessment results and a warning message template, and the output is the generated warning message. The terminal then informs the user of this message via audio and text.
[1316] Step 7:
[1317] The server also sends warning messages to pre-registered contacts (e.g., family and friends). The input is the generated warning message, and the output is the notification to the registered contacts. This helps users take quick action.
[1318] Through the above processing steps, this system can issue appropriate warnings to users before they are exposed to the risk of fraud, thereby preventing them from becoming victims.
[1319] Furthermore, an emotion engine that estimates the user's emotions may be incorporated. That is, the identification processing unit 290 may use the emotion identification model 59 to estimate the user's emotions and perform identification processing using the user's emotions.
[1320] Modes for carrying out the invention
[1321] This invention is a system for users to recognize and prevent the risk of wire fraud in advance. This system collects the user's location information and voice data, and generates a warning message based on the analysis results. It also utilizes an emotion engine to recognize the user's emotions, thereby improving the accuracy and effectiveness of the warning. The warning message is then sent to the user's device and pre-registered contacts.
[1322] Program Processing Description
[1323] Collection and transmission of location information
[1324] The device uses a GPS sensor to detect the user's current location. When the user enters a specific area (for example, near a bank ATM), the device collects location information and periodically sends it to the server. This allows the server to accurately determine the user's location.
[1325] Collection and transmission of audio data
[1326] When a user initiates a call, the device begins collecting the call audio in real time. This audio data is encrypted to protect privacy. The encrypted audio data is then sent to a server.
[1327] Emotional analysis using an emotion engine
[1328] The terminal or server uses the received audio data to evaluate the user's emotional state. The emotion engine recognizes the user's emotions (e.g., tension, fear, anxiety) from factors such as the tone, speed, and intonation of the voice.
[1329] Data analysis using analytical methods
[1330] The server integrates and analyzes the received location information and voice data. It analyzes the location information to confirm that the user is in a specific location (e.g., a bank ATM). It also uses a generative AI model to analyze the voice data and detect specific keywords or phrases (e.g., "transfer," "hurry," "money," etc.).
[1331] Risk assessment that takes emotional state into account
[1332] The server takes into account the results of the emotion engine analysis and compares them with the crime database to ultimately assess the risk of fraud. If the user's emotions are in a high-risk state such as tension, fear, or anxiety, the server determines that the likelihood of fraud is higher.
[1333] Generation and notification of warning messages
[1334] The server generates a warning message based on the risk assessment results. For example, it might say, "This call may be a scam. Do not transfer any money." The generated warning message is sent to the device in both voice and text formats. In emergencies, the warning message is also sent to pre-registered contacts (family and friends).
[1335] Specific example
[1336] 1. The user approaches the bank ATM.
[1337] The device detects the user's current location and notifies the server that "the user is approaching a bank ATM."
[1338] The specific location information detected is "Latitude: 35.6895, Longitude: 139.6917 (within Tokyo)".
[1339] 2. Data collection during voice calls
[1340] When a user initiates a call, the device starts recording the call, encrypting it in real time, and sending it to the server.
[1341] Example of a phone call: "Grandma, I need money urgently. Could you transfer it to me today?"
[1342] 3. Analysis using an emotion engine
[1343] The server analyzes the user's emotions from the voice data and detects tension and fear.
[1344] 4. Server-based data analysis
[1345] The server checks the user's location and confirms that the user is near a bank ATM.
[1346] The system analyzes voice data using an AI model to detect keywords such as "transfer," "urgent," and "money." Furthermore, it compares this data with local crime data to determine a high risk of fraud.
[1347] 5. Generating and notifying warning messages
[1348] The server generates a warning message stating, "This call may be a scam. Do not transfer any money under any circumstances," and notifies the device via both voice and text. Furthermore, warnings are also sent to pre-registered contacts (family and friends).
[1349] The system of this invention operates in this manner, significantly reducing the risk of users falling victim to wire fraud. Furthermore, by immediately notifying users of warning messages, it enables users to take swift action, effectively preventing damage before it occurs.
[1350] The following describes the processing flow.
[1351] Step 1:
[1352] The device uses a GPS sensor to detect the user's current location. When the user enters a specific area (for example, near a bank ATM), the device collects location information and periodically sends it to the server. This allows the server to accurately determine the user's location.
[1353] Step 2:
[1354] When a user initiates a call, the device begins collecting the call audio in real time. This audio data is encrypted to protect privacy. The encrypted audio data is then sent to a server.
[1355] Step 3:
[1356] The server integrates and analyzes the received location information and voice data. It analyzes the location information to confirm that the user is in a specific location (e.g., a bank ATM). It also uses a generative AI model to analyze the voice data and detect specific keywords or phrases (e.g., "transfer," "hurry," "money," etc.).
[1357] Step 4:
[1358] The terminal or server uses the received audio data to evaluate the user's emotional state. The emotion engine recognizes the user's emotions (e.g., tension, fear, anxiety) from factors such as tone, speed, and intonation of the voice. If the emotional state is determined to be tension, fear, or anxiety, special attention is paid to the analysis results.
[1359] Step 5:
[1360] The server takes into account the results of the emotion engine analysis and compares them with the crime database to ultimately assess the risk of fraud. If the user's emotions are in a high-risk state such as tension, fear, or anxiety, the server determines that the likelihood of fraud is higher. Based on this, the accuracy and reliability of the risk assessment are improved.
[1361] Step 6:
[1362] The server generates a warning message based on the risk assessment results. For example, it might say, "This call may be a scam. Do not transfer any money." The generated warning message is notified to the user in both audio and text format.
[1363] Step 7:
[1364] The device notifies the user of the generated warning message in both voice and text format. It also notifies the user to immediately check the warning message, encouraging a quick response. Furthermore, in emergencies, the warning message is also sent to pre-registered contacts (family and friends). This alerts those around the user as well.
[1365] Specific example:
[1366] 1. The user approaches the bank ATM.
[1367] The device detects the user's current location and notifies the server that "the user is approaching a bank ATM."
[1368] The location information detected is "Latitude: 35.6895, Longitude: 139.6917 (within Tokyo)".
[1369] 2. Data collection during voice calls
[1370] When a user initiates a call, the device begins recording the conversation. The recorded data is encrypted and sent to a server in real time.
[1371] Example of phone call content: "Grandma, I need money urgently. Could you transfer it to me today?"
[1372] 3. Data analysis by the server
[1373] The server analyzes the location information to confirm that the user is near a bank ATM.
[1374] The AI model analyzes the voice data to detect keywords such as "transfer," "urgent," and "money."
[1375] 4. Emotional analysis using an emotion engine
[1376] The server detects the user's emotions from the voice data and recognizes signs of tension or fear.
[1377] 5. Risk Assessment
[1378] The server also considers the results of the emotion engine and assesses the fraud risk as high based on the analysis results.
[1379] 6. Generating and notifying warning messages
[1380] The server generates a warning message stating, "This call may be a scam. Do not transfer any money under any circumstances," and sends it to the device in both voice and text formats.
[1381] 7. Notification to the user
[1382] The device notifies the user of warning messages in both voice and text formats, and in emergencies, it also sends warnings to pre-registered contacts.
[1383] In this way, the system of the present invention significantly reduces the risk of users falling victim to wire fraud and enables rapid countermeasures. Furthermore, by taking into account the user's emotional state, the accuracy and effectiveness of the warning are further improved.
[1384] (Example 2)
[1385] Next, we will describe Example 2. In the following description, the data processing device 12 will be referred to as the "server" and the robot 414 as the "terminal".
[1386] Currently, the number of victims of wire fraud is increasing, raising the risk of users being deceived and losing their valuable assets. Traditional countermeasures are insufficient because they lack the means for users to recognize the risk of fraud in advance, making it difficult to prevent victims. To solve this problem, there is a need for a system that can accurately detect the risk of wire fraud and provide users with rapid warnings.
[1387] The identification process performed by the identification processing unit 290 of the data processing device 12 in Example 2 is realized by the following means.
[1388] In this invention, the server includes emotion analysis means for analyzing the user's emotional state, analysis means for integrating and analyzing location information and voice data and detecting specific keywords or phrases using a generative AI model, and means for performing risk assessment. This enables users to detect the risk of wire fraud in advance with high accuracy and receive a warning quickly.
[1389] "Location information collection means" refers to means for obtaining the user's current location information.
[1390] "Location information transmission means" refers to a means for transmitting acquired location information to a server.
[1391] "Voice data collection means" refers to a means for collecting user voice data.
[1392] "Voice data encryption means" refers to a method of encrypting collected voice data for privacy protection.
[1393] "Voice data transmission means" refers to a means for sending encrypted voice data to a server.
[1394] "Emotion analysis means" refers to a method for evaluating a user's emotional state from voice data.
[1395] "Analysis means" refers to a means for integrating and analyzing location information and audio data, and for using a generative AI model to detect specific keywords or phrases.
[1396] A "risk assessment tool" is a means for evaluating risk based on the analysis results of an analysis tool.
[1397] A "warning message generation means" is a means for generating a warning message based on the results of a risk assessment.
[1398] A "warning notification means" is a means for notifying the user terminal of the generated warning message.
[1399] "Notification method" refers to a means of notifying pre-registered contacts of the generated warning message.
[1400] "Voice notification means" refers to a means of notifying the user of the generated warning message in voice format.
[1401] A "text notification method" is a means of notifying the user of the generated warning message in text format.
[1402] A "generative AI model" is an artificial intelligence model used to analyze audio data and detect specific keywords or phrases.
[1403] This invention is a system for users to recognize and prevent the risk of wire fraud in advance. The system includes means for collecting the user's location information and voice data and generating a warning message based on the analysis results.
[1404] Hardware configuration
[1405] The system utilizes mobile devices such as smartphones equipped with a GPS sensor, microphone, and data communication capabilities. The GPS sensor collects the user's current location, and the microphone collects call audio. The collected data is then transmitted to a server via the internet.
[1406] The server utilizes high-performance computers and cloud services to provide the processing power needed for data analysis and sentiment analysis. Specific software used includes IBM Watson's sentiment analysis API and GPT-3 generative AI models.
[1407] Software Configuration
[1408] The device has the following software features:
[1409] 1. Location information collection method: The user's current location is obtained using the built-in GPS sensor, and location data is generated.
[1410] 2. Voice data collection method: Voice data is collected in real time using a microphone during a call.
[1411] 3. Audio data encryption method: The collected audio data is encrypted using AES-256.
[1412] 4. Location information transmission means: The collected location information is transmitted to the server.
[1413] 5. Means of transmitting audio data: Encrypted audio data is sent to the server.
[1414] The server has the following software features:
[1415] 1. Emotion analysis method: The received audio data is analyzed using IBM Watson's emotion analysis API to evaluate the user's emotional state.
[1416] 2. Analysis method: The received location information and audio data are integrated, and specific keywords and phrases are detected using the GPT-3 generation AI model.
[1417] 3. Risk assessment method: The risk of wire fraud will be finalized based on the analysis results and sentiment analysis results.
[1418] 4. Warning message generation means: Generates a warning message based on the risk assessment results.
[1419] 5. Warning notification means: The generated warning message is sent to the user's terminal and to pre-registered contacts.
[1420] Specific example
[1421] 1. The user approaches the bank ATM.
[1422] The device uses a GPS sensor to obtain its current location and sends the location information to the server when it detects that "the user is approaching a bank ATM." An example of location data is "Latitude: 35.6895, Longitude: 139.6917 (within Tokyo)."
[1423] 2. Data collection during voice calls
[1424] When a user initiates a call, the device collects the call audio, encrypts it with AES-256, and sends it to the server in real time. Specifically, the call content might be something like, "Grandma, I need money urgently. Can you transfer it to me today?" The audio recording and encryption process is performed on the device itself.
[1425] 3. Analysis using an emotion engine
[1426] The server decodes the received audio data and uses an emotion analysis API to assess the user's emotion as "tension." The analysis takes into account the tone, speed, and intonation of the voice.
[1427] 4. Server-based data analysis
[1428] The server uses location data to confirm that the user is near a bank ATM. Simultaneously, it analyzes the voice data using an AI model (e.g., GPT-3) to detect keywords such as "transfer," "urgent," and "money." This then triggers a voice-to-text conversion process, highlighting specific keywords.
[1429] 5. Risk assessment that takes emotional state into consideration
[1430] The server determines a high risk of fraud by combining voice analysis results and sentiment analysis results. This includes matching against past crime data and calculating the degree of match.
[1431] 6. Generating and notifying warning messages
[1432] The server generates a warning message stating, "This call may be a scam. Do not transfer any money under any circumstances," and also generates a voice message using a speech synthesis API. The generated warning message is sent to the device and, if necessary, also notified to registered contacts via SMS or email.
[1433] Example of a prompt
[1434] "Please explain the process for a program that notifies a user when they approach a bank ATM. Also, please explain the procedure for collecting voice data from a user during a call and assessing the risk of fraud."
[1435] The flow of the specific processing in Example 2 will be explained using Figure 13.
[1436] Step 1: Collect and transmit location information
[1437] The device uses its built-in GPS sensor to acquire the user's current location at regular intervals. Specifically, it acquires location information every minute and generates latitude and longitude data (input: GPS sensor data, output: location information data). When the user enters a specific area (for example, near a bank ATM), it sends that location information to the server (data processing: location information detection, data calculation: range determination, output: location information transmission).
[1438] Step 2: Collect and transmit audio data
[1439] When a user initiates a call, the device uses its microphone to collect the call audio in real time (input: call audio, output: audio data). The collected audio data is immediately encrypted using AES-256 (data processing: audio data encryption), and the encrypted audio data is sent to the server in real time (output: encrypted audio data transmission).
[1440] Step 3: Emotional analysis using the emotion engine
[1441] The server decrypts the received encrypted audio data (input: encrypted audio data, output: decrypted audio data) and sends the audio data to the emotion analysis API. The emotion analysis API analyzes the tone, speed, and intonation of the voice and evaluates the user's emotional state (tension, fear, anxiety, etc.) (data calculation: emotion analysis, output: emotional state data).
[1442] Step 4: Data analysis using analytical tools
[1443] The server integrates and analyzes emotional state data with received location and audio data (input: location data, decoded audio data, emotional state data). It uses a GPT-3 generative AI model to detect specific keywords or phrases (e.g., "transfer", "hurry", "money") from the audio data (data calculation: keyword detection, output: keyword data). It uses location information to determine if the user is in a specific location (e.g., a bank ATM) (data calculation: location confirmation, output: location confirmation data).
[1444] Step 5: Risk assessment considering emotional state
[1445] The server integrates keyword data and location data from the analysis tools, and also considers emotional state data to assess fraud risk (Input: Keyword data, emotional state data, location data; Output: Risk assessment data). For example, if the user's emotions are in a high-risk state such as tension or fear, and specific keywords are included, the risk is judged to be high (Data calculation: Risk assessment).
[1446] Step 6: Generate and send warning messages
[1447] The server generates a warning message based on risk assessment data (input: risk assessment data, output: warning message). For example, it generates a message stating, "This call may be a scam. Do not transfer any money under any circumstances." (data processing: warning message generation). The generated warning message is converted into voice format using a speech synthesis API (output: voice warning message) and also sent to the terminal in text format (output: text warning message). Furthermore, in emergencies, notifications are also sent to pre-registered contacts (family and friends) (output: contact notification data).
[1448] (Application Example 2)
[1449] Next, we will explain application example 2. In the following explanation, the data processing device 12 will be referred to as the "server" and the robot 414 as the "terminal".
[1450] Fraudulent activities, including wire transfer scams, are increasing, particularly targeting the elderly. Therefore, there is a need to detect the risk of fraud early and prevent it before it occurs. However, current systems lack the means to effectively analyze users' emotional states and real-time call content, making it difficult to accurately determine situations with a high risk of fraud. Therefore, it is necessary to provide a system that can quickly and accurately detect situations with a high risk of fraud and issue warnings to users, their relatives, and friends.
[1451] The specific processing performed by the specific processing unit 290 of the data processing device 12 in Application Example 2 is realized by the following means.
[1452] In this invention, the server includes location information collection means for collecting user location information, voice data collection means for collecting user voice data, emotion analysis means for analyzing the user's emotional state based on the voice data, and keyword detection means for detecting specific keywords. This makes it possible to detect fraud risks such as wire fraud with high accuracy and speed, and to issue warnings to the user and their family and friends.
[1453] "Location information collection means" refers to devices and software used to determine the user's current location.
[1454] "Location information transmission means" refers to devices or software used to transmit collected user location information to a server.
[1455] "Voice data collection means" refers to devices or software used to collect voices emitted by users.
[1456] "Voice data encryption means" refers to devices or software used to encrypt collected voice data so that it cannot be deciphered by third parties.
[1457] "Voice data transmission means" refers to devices or software used to transmit encrypted voice data to a server.
[1458] "Analysis means" refers to devices or software used to analyze location information and voice data collected on a server to assess fraud risk.
[1459] "Warning message generation means" refers to a device or software that automatically generates warning messages based on the analysis results of the analysis means.
[1460] "Warning notification means" refers to devices or software that notify the user's terminal of the generated warning message.
[1461] "Notification means" refers to devices or software used to notify pre-registered contacts of generated warning messages.
[1462] "Emotional analysis means" refers to devices and software that analyze a user's emotional state in real time based on voice data.
[1463] "Keyword detection means" refers to devices or software used to detect specific keywords from audio data.
[1464] This invention is a system for users to recognize and prevent the risk of wire fraud in advance. The system includes means for identifying the user's current location and transmitting that information to a server, and means for collecting, encrypting, and transmitting the user's voice data. Furthermore, the server includes means for analyzing this information and generating a warning message to notify the user's terminal and pre-registered contacts.
[1465] Collection and transmission of location information
[1466] The device is equipped with a GPS sensor that constantly monitors the user's current location and acquires data using location information collection methods. When the user enters a specific area (for example, near a bank ATM), location information is sent to the server using location information transmission methods. This allows the server to determine the user's precise location.
[1467] Collection and transmission of audio data
[1468] When a user starts a phone call, the call content is collected in real time by an audio data collection means and encrypted by an audio data encryption means. The encrypted audio data is then transmitted to the server using an audio data transmission means.
[1469] Sentiment analysis and keyword detection
[1470] The server uses emotion analysis tools to analyze the received audio data. This evaluates the user's emotions (e.g., tension, fear, anxiety) from factors such as tone, speed, and intonation. Furthermore, keyword detection tools detect specific keywords within the audio data (e.g., "transfer," "hurry," "money," etc.).
[1471] Risk assessment and warning message generation
[1472] The server integrates location information, voice data, sentiment analysis results, and keyword detection results to assess the fraud risk. If the risk is determined to be high, a warning message is generated by the warning message generation mechanism. For example, a message such as, "This call may be a scam. Do not transfer any money."
[1473] Warning message notification
[1474] The warning notification system will send the generated warning message to the user's device in both audio and text format. In emergencies, the same warning message will also be sent to pre-registered contacts (family and friends) using the notification system.
[1475] Specific example
[1476] 1. The user approaches the bank ATM.
[1477] The device detects the user's current location and notifies the server that "the user is approaching a bank ATM." The specific location information detected is "Latitude: 35.6895, Longitude: 139.6917 (within Tokyo)."
[1478] 2. Data collection during voice calls
[1479] When a user initiates a call, the device starts recording the conversation, encrypting it in real time, and sending it to the server. Example of a call: "Grandma, I need money urgently. Can you transfer it to me today?"
[1480] 3. Analysis using an emotion engine
[1481] The server analyzes the user's emotions from the voice data and detects tension and fear.
[1482] 4. Server-based data analysis
[1483] The server checks the user's location to confirm they are near a bank ATM. An AI model analyzes the voice data to detect keywords such as "transfer," "urgent," and "money." Furthermore, it cross-references this data with local crime statistics to determine a high risk of fraud.
[1484] 5. Generating and notifying warning messages
[1485] The server generates a warning message stating, "This call may be a scam. Do not transfer any money under any circumstances," and notifies the device via both voice and text. Furthermore, warnings are also sent to pre-registered contacts (family and friends).
[1486] Example of a prompt
[1487] text
[1488] Evaluate the user's emotional state and generate an emergency fraud warning message. For example, use this if you receive audio data such as "Location: 35.6895, 139.6917" and "Grandma, I need money right away. Can you transfer it to me today?". If the emotional score is -0.5 or lower and contains fraud-related keywords such as "transfer" and "money", generate a warning message stating, "This call may be a scam. Do not transfer any money."
[1489] The above describes the embodiments for carrying out this invention.
[1490] The flow of a specific process in Application Example 2 will be explained using Figure 14.
[1491] Step 1:
[1492] Collecting and transmitting user location information
[1493] The device uses a GPS sensor to collect the user's location information. The collected location information (e.g., latitude 35.6895, longitude 139.6917) is transmitted to the server using a location information transmission means.
[1494] Input: Location information from GPS sensor
[1495] Output: User location information sent to the server
[1496] Step 2:
[1497] Collection and transmission of audio data
[1498] When a user initiates a call, the device collects audio in real time using an audio data collection device. The collected audio data is encrypted using an audio data encryption device and transmitted to a server using an audio data transmission device.
[1499] Input: Audio data from a call
[1500] Output: Encrypted audio data sent to the server
[1501] Step 3:
[1502] Location information analysis
[1503] The server analyzes the received location information using an analysis tool to determine whether the user is in a specific risk area (e.g., near a bank ATM).
[1504] Input: Sent location information
[1505] Output: Determination result of whether the user is in a risk area.
[1506] Step 4:
[1507] Sentiment analysis of voice data
[1508] The server analyzes the audio data using emotion analysis tools. Specifically, it evaluates the user's emotions (e.g., tension, fear, anxiety) based on factors such as tone, speed, and intonation of the voice.
[1509] Input: Decrypted audio data
[1510] Output: User sentiment score
[1511] Step 5:
[1512] Detection of specific keywords
[1513] The server uses keyword detection means to detect specific keywords (e.g., "transfer", "hurry", "money") from the audio data.
[1514] Input: Decrypted audio data
[1515] Output: List of detected keywords
[1516] Step 6:
[1517] Risk assessment
[1518] The server integrates location data analysis results, user sentiment scores, and detected keyword lists to assess fraud risk. If a high risk is determined, a warning message is generated using a warning message generation mechanism.
[1519] Input: Location analysis results, sentiment score, keyword list
[1520] Output: Fraud risk assessment results
[1521] Step 7:
[1522] Generation and notification of warning messages
[1523] The server sends the generated warning message to the user's terminal using a warning notification system. Simultaneously, it also sends the same warning message to pre-registered contacts using the notification system.
[1524] Input: Fraud risk assessment results
[1525] Output: Warning message sent to the user's device and registered contacts.
[1526] These steps enable the creation of a system that effectively detects fraud risks and warns users and their associates.
[1527] The specific processing unit 290 transmits the result of the specific processing to the robot 414. In the robot 414, the control unit 46A causes the speaker 240 and the controlled object 443 to output the result of the specific processing. The microphone 238 acquires audio indicating user input for the result of the specific processing. The control unit 46A transmits the audio data indicating user input acquired by the microphone 238 to the data processing unit 12. In the data processing unit 12, the specific processing unit 290 acquires the audio data.
[1528] Data generation model 58 is a type of so-called generative AI (Artificial Intelligence). One example of data generation model 58 is ChatGPT (Internet search<URL: https: / / openai.com / blog / chatgpt> ), Gemini (Internet search) <url: https: gemini.google.com ?hl="ja">Examples of generative AI include the following. The data generation model 58 is obtained by performing deep learning on a neural network. The data generation model 58 is input with prompts containing instructions, and with inference data such as audio data representing speech, text data representing text, and image data representing images. The data generation model 58 infers from the input inference data according to the instructions indicated by the prompts, and outputs the inference results in data formats such as audio data and text data. Here, inference refers to, for example, analysis, classification, prediction, and / or summarization.
[1529] In the above embodiment, an example was given in which the specific processing is performed by the data processing device 12, but the technology of this disclosure is not limited thereto, and the specific processing may also be performed by the robot 414.
[1530] Furthermore, the emotion identification model 59, acting as an emotion engine, may determine the user's emotion according to a specific mapping. Specifically, the emotion identification model 59 may determine the user's emotion according to a specific mapping, which is an emotion map (see Figure 9). Similarly, the emotion identification model 59 may also determine the robot's emotion, and the identification processing unit 290 may perform identification processing using the robot's emotion.
[1531] Figure 9 shows an emotion map 400 in which multiple emotions are mapped. In the emotion map 400, emotions are arranged in concentric circles radiating from the center. The closer to the center of the concentric circles, the more primitive the emotions are located. Further out of the concentric circles, emotions representing states and actions arising from mental states are located. Emotion is a concept that includes feelings and mental states. On the left side of the concentric circles, emotions that are generally generated from reactions occurring in the brain are located. On the right side of the concentric circles, emotions that are generally induced by situational judgment are located. Above and below the concentric circles, emotions that are generally generated from reactions occurring in the brain and induced by situational judgment are located. In addition, the emotion of "pleasure" is located on the upper side of the concentric circles, and the emotion of "displeasure" is located on the lower side. Thus, in the emotion map 400, multiple emotions are mapped based on the structure in which emotions arise, and emotions that are likely to occur simultaneously are mapped close together.
[1532] These emotions are distributed at the 3 o'clock position on the Emotion Map 400, and usually fluctuate between feelings of security and anxiety. In the right half of the Emotion Map 400, situational awareness takes precedence over internal feelings, resulting in a calm impression.
[1533] The inside of the Emotion Map 400 represents inner thoughts, while the outside represents actions. Therefore, the further you go from the outside of the Emotion Map 400, the more visible (expressed in actions) your emotions become.
[1534] Here, human emotions are based on various balances, such as posture and blood sugar levels. When these balances deviate from the ideal, it results in discomfort, and when they approach the ideal, it results in pleasure. Similarly, in robots, cars, motorcycles, etc., emotions can be created based on various balances, such as posture and battery level. When these balances deviate from the ideal, it results in discomfort, and when they approach the ideal, it results in pleasure. The emotion map can be generated, for example, based on Dr. Mitsuyoshi's emotion map (Research on a system for analyzing brain physiological signals of speech emotion recognition and emotion, Tokushima University, doctoral dissertation: https: / / ci.nii.ac.jp / naid / 500000375379). The left half of the emotion map contains emotions belonging to a region called "response," where sensation is dominant. The right half of the emotion map contains emotions belonging to a region called "situation," where situational awareness is dominant.
[1535] The emotion map defines two emotions that promote learning. One is the emotion around the middle of the negative "repentance" and "reflection" on the situation side. In other words, it is when the robot experiences negative emotions such as "I never want to feel this way again" or "I don't want to be scolded again." The other is the emotion around the positive "desire" on the reaction side. In other words, it is when the robot has positive feelings such as "I want more" or "I want to know more."
[1536] The emotion identification model 59 inputs user input into a pre-trained neural network, obtains emotion values representing each emotion shown in the emotion map 400, and determines the user's emotion. This neural network is pre-trained based on multiple training data sets, which are combinations of user input and emotion values representing each emotion shown in the emotion map 400. Furthermore, this neural network is trained so that emotions located close together have similar values, as shown in the emotion map 900 in Figure 10. Figure 10 shows an example where multiple emotions such as "reassured," "calm," and "confident" have similar emotion values.
[1537] The above description primarily focuses on the functions of the data processing device 12 in relation to this disclosure. However, the system related to this disclosure is not necessarily implemented on a server. The system related to this disclosure may be implemented as a general information processing system. This disclosure may be implemented, for example, as a software program that runs on a personal computer or as an application that runs on a smartphone. The method related to this disclosure may be provided to users in SaaS (Software as a Service) format.
[1538] In the above embodiment, an example was given in which a specific process is performed by a single computer 22. However, the technology of this disclosure is not limited thereto, and a distributed processing of the specific process may be performed by multiple computers, including computer 22. For example, a data generation model 58 may be provided in an external device of the data processing device 12, and the external device may generate data according to the input data.
[1539] In the above embodiment, an example was given in which the specific processing program 56 is stored in the storage 32, but the technology of this disclosure is not limited thereto. For example, the specific processing program 56 may be stored in a portable, computer-readable, non-temporary storage medium such as a USB (Universal Serial Bus) memory. The specific processing program 56 stored in the non-temporary storage medium is installed in the computer 22 of the data processing device 12. The processor 28 executes specific processing according to the specific processing program 56.
[1540] Alternatively, the specific processing program 56 may be stored in a storage device such as a server connected to the data processing device 12 via the network 54, and the specific processing program 56 may be downloaded and installed on the computer 22 in response to a request from the data processing device 12.
[1541] Furthermore, it is not necessary to store the entirety of the specific processing program 56 in a storage device such as a server connected to the data processing device 12 via the network 54, or to store the entirety of the specific processing program 56 in the storage 32; it is acceptable to store only a portion of the specific processing program 56.
[1542] The following types of processors can be used as hardware resources to perform specific processing. Examples of processors include a CPU, a general-purpose processor that functions as a hardware resource to perform specific processing by executing software, i.e., a program. Other examples of processors include dedicated electrical circuits, such as FPGAs (Field-Programmable Gate Arrays), PLDs (Programmable Logic Devices), or ASICs (Application Specific Integrated Circuits), which have circuit configurations specifically designed to perform specific processing. All of these processors have built-in or connected memory, and all of them perform specific processing by using memory.
[1543] The hardware resource that performs a specific process may consist of one of these various processors, or it may consist of a combination of two or more processors of the same or different types (for example, a combination of multiple FPGAs, or a combination of a CPU and an FPGA). Alternatively, the hardware resource that performs a specific process may consist of a single processor.
[1544] Examples of configurations using a single processor include, firstly, a configuration in which one or more CPUs and software are combined to form a single processor, and this processor functions as a hardware resource that performs a specific process. Secondly, there is a configuration using a processor that realizes the functions of the entire system, including multiple hardware resources that perform a specific process, on a single IC chip, as exemplified by SoCs (System-on-a-chip). In this way, a specific process is realized using one or more of the above types of processors as hardware resources.
[1545] Furthermore, the hardware structure of these various processors can more specifically utilize electrical circuits that combine circuit elements such as semiconductor devices. Also, the specific processing described above is merely an example. Therefore, it goes without saying that unnecessary steps can be deleted, new steps added, or the processing order rearranged, as long as it does not deviate from the main purpose.
[1546] The descriptions and illustrations presented above are detailed explanations of the technical aspects of this disclosure and are merely examples of the technical aspects. For example, the above descriptions of the structure, function, operation, and effect are examples of the structure, function, operation, and effect of the technical aspects of this disclosure. Therefore, it goes without saying that you may delete unnecessary parts, add new elements, or replace elements in the descriptions and illustrations presented above, as long as you do not deviate from the essence of the technical aspects of this disclosure. Furthermore, in order to avoid confusion and facilitate understanding of the technical aspects of this disclosure, explanations of common technical knowledge and the like that do not require special explanation to enable the implementation of the technical aspects of this disclosure have been omitted from the descriptions and illustrations presented above.
[1547] All documents, patent applications, and technical standards described herein are incorporated by reference to the same extent as if each individual document, patent application, and technical standard were specifically and individually noted to be incorporated by reference.
[1548] The following is further disclosed regarding the embodiments described above.
[1549] Claims
[1550] (Claim 1)
[1551] A means for collecting location information to collect user location information,
[1552] A location information transmission means for transmitting the aforementioned location information to a server,
[1553] A means for collecting user voice data,
[1554] A means for encrypting the aforementioned audio data,
[1555] A voice data transmission means for sending the encrypted voice data to a server,
[1556] The server includes an analysis means for analyzing the location information and audio data,
[1557] A warning message generation means for generating a warning message based on the analysis results of the aforementioned analysis means,
[1558] A warning notification means for notifying the user terminal of the generated warning message,
[1559] A notification means for notifying the generated warning message to pre-registered contacts,
[1560] A system that includes this.
[1561] (Claim 2)
[1562] The system according to claim 1, further comprising a voice notification means for notifying the user of the generated warning message in voice format.
[1563] (Claim 3)
[1564] The system according to claim 1, further comprising a text notification means for notifying the user of the generated warning message in text format.
[1565] "Example 1"
[1566] (Claim 1)
[1567] A means for collecting location information to collect user location information,
[1568] A location information transmission means for transmitting the aforementioned location information to a network device,
[1569] A means for collecting user voice data,
[1570] A means for encrypting the aforementioned audio data,
[1571] A voice data transmission means for transmitting the encrypted voice data to a network device,
[1572] The network device includes an analysis means for analyzing the location information and voice data,
[1573] A warning message generation means for generating a warning message based on the analysis results of the aforementioned analysis means,
[1574] A warning notification means for notifying the user terminal of the generated warning message,
[1575] A notification means for notifying the generated warning message to pre-registered contacts,
[1576] A risk assessment method that uses a generative artificial intelligence model to detect specific keywords and assess risk,
[1577] A range monitoring means for detecting when a user enters a specific area,
[1578] Adjustment means for adjusting the content of the warning message based on the aforementioned range,
[1579] A format conversion means for notifying the adjusted message in voice or text format,
[1580] A system that includes this.
[1581] (Claim 2)
[1582] The system according to claim 1, further comprising a voice notification means for notifying the user of the generated warning message in voice format.
[1583] (Claim 3)
[1584] The system according to claim 1, further comprising a text notification means for notifying the user of the generated warning message in text format.
[1585] "Application Example 1"
[1586] (Claim 1)
[1587] A means for collecting location information to collect user location information,
[1588] A location information transmission means for transmitting the aforementioned location information to a server,
[1589] A means for collecting user voice data,
[1590] A means for encrypting the aforementioned audio data,
[1591] A voice data transmission means for sending the encrypted voice data to a server,
[1592] The server includes an analysis means for analyzing the location information and audio data,
[1593] An analytical processing means for performing analysis using a generative AI model,
[1594] A warning message generation means for generating a warning message based on the analysis results of the aforementioned analysis means,
[1595] A warning notification means for notifying the user terminal of the generated warning message,
[1596] A notification means for notifying the generated warning message to pre-registered contacts,
[1597] A system that includes this.
[1598] (Claim 2)
[1599] The system according to claim 1, further comprising a voice notification means for notifying the user of the generated warning message in voice format.
[1600] (Claim 3)
[1601] The system according to claim 1, further comprising a text notification means for notifying the user of the generated warning message in text format.
[1602] "Example 2 of combining an emotion engine"
[1603] (Claim 1)
[1604] A means for collecting location information to collect user location information,
[1605] A location information transmission means for transmitting the aforementioned location information to a server,
[1606] A means for collecting user voice data,
[1607] A means for encrypting the aforementioned audio data,
[1608] A voice data transmission means for sending the encrypted voice data to a server,
[1609] An emotion analysis means for analyzing the user's emotional state from the aforementioned audio data,
[1610] The server includes an analysis means for integrating and analyzing the location information and voice data, and for detecting specific keywords or phrases using a generative AI model.
[1611] A means for performing a risk assessment based on the analysis results of the aforementioned analysis means,
[1612] A warning message generation means for generating a warning message based on the risk assessment results,
[1613] A warning notification means for notifying the user terminal of the generated warning message,
[1614] A notification means for notifying the generated warning message to pre-registered contacts,
[1615] A system that includes this.
[1616] (Claim 2)
[1617] The system according to claim 1, further comprising a voice notification means for notifying the user of the generated warning message in voice format.
[1618] (Claim 3)
[1619] The system according to claim 1, further comprising a text notification means for notifying the user of the generated warning message in text format.
[1620] "Application example 2 when combining with an emotional engine"
[1621] (Claim 1)
[1622] A means for collecting location information to collect user location information,
[1623] A location information transmission means for transmitting the aforementioned location information to a server,
[1624] A means for collecting user voice data,
[1625] A means for encrypting the aforementioned audio data,
[1626] A voice data transmission means for sending the encrypted voice data to a server,
[1627] The server includes an analysis means for analyzing the location information and audio data,
[1628] A warning message generation means for generating a warning message based on the analysis results of the aforementioned analysis means,
[1629] A warning notification means for notifying the user terminal of the generated warning message,
[1630] A notification means for notifying the generated warning message to pre-registered contacts,
[1631] An emotion analysis method for analyzing a user's emotional state based on voice data,
[1632] A keyword detection means for detecting specific keywords,
[1633] A system that includes this.
[1634] (Claim 2)
[1635] The system according to claim 1, further comprising a voice notification means for notifying the user of the generated warning message in voice format.
[1636] (Claim 3)
[1637] The system according to claim 1, further comprising a text notification means for notifying the user of the generated warning message in text format. [Explanation of symbols]
[1638] 10, 210, 310, 410 Data Processing Systems 12 Data Processing Devices 14 Smart Devices 214 Smart Glasses 314 Headset-type terminal 414 Robots< / url:> < / url:> < / url:> < / url:>
Claims
1. A means for collecting location information to collect user location information, A location information transmission means for transmitting the aforementioned location information to a server, A means for collecting user voice data, A means for encrypting the aforementioned audio data, A voice data transmission means for sending the encrypted voice data to a server, The server includes an analysis means for analyzing the location information and audio data, A warning message generation means for generating a warning message based on the analysis results of the aforementioned analysis means, A warning notification means for notifying the user terminal of the generated warning message, A notification means for notifying the generated warning message to pre-registered contacts, A system that includes this.
2. The system according to claim 1, further comprising a voice notification means for notifying the user of the generated warning message in voice format.
3. The system according to claim 1, further comprising a text notification means for notifying the user of the generated warning message in text format.
Citation Information
Patent Citations
Persona chatbot control method and system
JP2022180282A