System

A real-time voice-to-text conversion and analysis system detects fraud in calls, evaluating risk levels and issuing warnings to prevent telephone fraud among the elderly.

JP2026014898APending Publication Date: 2026-01-29SOFTBANK GROUP CORP
View PDF 1 Cites 0 Cited by

Patent Information

Application Number
JP2024116372
Authority / Receiving Office
JP · JP
Patent Type
Applications
Current Assignee / Owner
Filing Date
2024-07-19
Publication Date
2026-01-29

AI Technical Summary

Technical Problem

Conventional systems fail to analyze call content in real time and issue warnings to prevent telephone fraud effectively, particularly targeting the elderly, leading to financial losses and psychological stress.

Method used

A system that converts voice data into text data in real time, analyzes for fraud-related keywords and phrases, evaluates the risk level, and issues warnings when the risk exceeds a threshold, using voice and text notifications.

Benefits of technology

Enables rapid detection and prevention of telephone fraud by issuing intuitive warnings to elderly users, significantly reducing their risk of falling victim to scams.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure 2026014898000001_ABST
    Figure 2026014898000001_ABST
Patent Text Reader

Abstract

A system is provided.SOLUTION: A system comprising: means for converting voice data into text data in real time; means for analyzing the text data and identifying a keyword or phrase related to fraud; means for evaluating a possibility of fraud based on the analysis result and scoring a risk; and means for notifying a user of a warning when the risk exceeds a set threshold.SELECTED DRAWING: Figure 1
Need to check novelty before this filing date? Find Prior Art

Description

[Technical Field]

[0001] The technology of the present disclosure relates to a system. [Background technology]

[0002] Patent document 1 discloses a persona chatbot control method performed by at least one processor, the method including the steps of receiving a user utterance, adding the user utterance to a prompt including an instruction sentence related to a description of the chatbot character, encoding the prompt, and inputting the encoded prompt into a language model to generate a chatbot utterance in response to the user utterance. [Prior art documents] [Patent documents]

[0003] [Patent Document 1] Japanese Patent Publication No. 2022-180282 Summary of the Invention [Problem to be solved by the invention]

[0004] In modern society, the number of victims of telephone fraud, particularly targeting the elderly, is increasing, resulting in serious problems such as financial losses and psychological stress. Conventional technologies and methods are unable to analyze call content in real time and issue warnings to users at appropriate times, making it difficult to prevent fraud damage before it occurs. The present invention aims to solve these problems by providing a system that can detect possible fraud with high accuracy in real time and immediately issue warnings to users. [Means for solving the problem]

[0005] The present invention is a system that includes a means for converting voice data into text data in real time, a means for analyzing the text data and identifying keywords and phrases related to fraud, a means for evaluating the possibility of fraud based on the analysis results and scoring the risk level, and a means for issuing a warning to the user when the risk level exceeds a set threshold. This makes it possible to instantly detect suspicious behavior during a call and issue a fraud warning to the user. Furthermore, because the system includes a means for recording voice data in real time and notifying the user of the warning by voice, it is easy for elderly users to intuitively understand and respond quickly.

[0006] "Voice data" refers to data that records voice information such as telephone calls and conversations in digital format.

[0007] "Text data" refers to data obtained by converting voice data into character information, and is expressed in the form of sentences or words.

[0008] "Real time" means that a process is carried out at the same speed as real time, and refers to a state in which processing is carried out immediately without delay.

[0009] "Analysis" refers to the process of examining data in detail to understand its components and meaning.

[0010] "Keywords" are important words or phrases related to a particular piece of information or topic.

[0011] A "phrase" is a short phrase or phrase that combines multiple words to create a meaning.

[0012] "Scoring" is the act of quantifying an evaluation based on specific criteria.

[0013] "Danger level" is an indicator that indicates the degree of risk in a particular action or situation.

[0014] A "threshold" is a limit value used to determine whether or not a specific condition or standard is exceeded.

[0015] "Warning" refers to a message or notification that alerts the user.

[0016] "User" refers to the target or individual who uses this system, and in this case, we are particularly thinking of the elderly. [Brief explanation of the drawings]

[0017] [Figure 1] 1 is a conceptual diagram showing an example of the configuration of a data processing system according to a first embodiment. [Figure 2] 1 is a conceptual diagram showing an example of main functions of a data processing device and a smart device according to a first embodiment. [Figure 3] FIG. 10 is a conceptual diagram showing an example of the configuration of a data processing system according to a second embodiment. [Figure 4] FIG. 10 is a conceptual diagram showing an example of main functions of a data processing device and smart glasses according to a second embodiment. [Figure 5] FIG. 10 is a conceptual diagram showing an example of the configuration of a data processing system according to a third embodiment. [Figure 6] FIG. 11 is a conceptual diagram showing an example of main functions of a data processing device and a headset-type terminal according to a third embodiment. [Figure 7] FIG. 10 is a conceptual diagram showing an example of the configuration of a data processing system according to a fourth embodiment. [Figure 8] FIG. 10 is a conceptual diagram showing an example of main functions of a data processing device and a robot according to a fourth embodiment. [Figure 9] 1 shows an emotion map onto which multiple emotions are mapped. [Figure 10] 1 shows an emotion map onto which multiple emotions are mapped. [Figure 11] FIG. 3 is a sequence diagram showing a processing flow of the data processing system according to the first embodiment. [Figure 12]FIG. 10 is a sequence diagram showing the flow of processing in the data processing system in Application Example 1. [Figure 13] FIG. 10 is a sequence diagram showing the flow of processing in the data processing system according to the second embodiment when an emotion engine is combined. [Figure 14] FIG. 10 is a sequence diagram showing the flow of processing in the data processing system in Application Example 2 when an emotion engine is combined. DETAILED DESCRIPTION OF THE INVENTION

[0018] An example of an embodiment of a system according to the technology of the present disclosure will be described below with reference to the accompanying drawings.

[0019] First, the terms used in the following description will be explained.

[0020] In the following embodiments, a coded processor (hereinafter simply referred to as a "processor") may be a single arithmetic device or a combination of multiple arithmetic devices. Furthermore, a processor may be a single type of arithmetic device or a combination of multiple types of arithmetic devices. Examples of arithmetic devices include a CPU (Central Processing Unit), a GPU (Graphics Processing Unit), a GPGPU (General-Purpose computing on Graphics Processing Units), and an APU (Accelerated Processing Unit).

[0021] In the following embodiments, a coded RAM (Random Access Memory) is a memory in which information is temporarily stored and is used as a working memory by a processor.

[0022] In the following embodiments, the coded storage is one or more non-volatile storage devices that store various programs, various parameters, etc. Examples of non-volatile storage devices include flash memory (SSD (Solid State Drive)), magnetic disks (e.g., hard disks), and magnetic tapes.

[0023] In the following embodiments, a communication I / F (Interface) with a symbol is an interface including a communication processor, an antenna, etc. The communication I / F controls communication between multiple computers. Examples of communication standards applied to the communication I / F include wireless communication standards including 5G (5th Generation Mobile Communication System), Wi-Fi (registered trademark), Bluetooth (registered trademark), etc.

[0024] In the following embodiments, "A and / or B" is synonymous with "at least one of A and B." In other words, "A and / or B" means that it may be only A, only B, or a combination of A and B. Furthermore, in this specification, the same concept as "A and / or B" is also applied when three or more things are expressed connected by "and / or."

[0025] [First embodiment]

[0026] FIG. 1 shows an example of the configuration of a data processing system 10 according to the first embodiment.

[0027] 1, a data processing system 10 includes a data processing device 12 and a smart device 14. An example of the data processing device 12 is a server.

[0028] The data processing device 12 includes a computer 22, a database 24, and a communication I / F 26. The computer 22 is an example of a "computer" according to the technology of the present disclosure. The computer 22 includes a processor 28, a RAM 30, and a storage 32. The processor 28, the RAM 30, and the storage 32 are connected to a bus 34. The database 24 and the communication I / F 26 are also connected to the bus 34. The communication I / F 26 is connected to a network 54. Examples of the network 54 include a WAN (Wide Area Network) and / or a LAN (Local Area Network).

[0029] The smart device 14 includes a computer 36, a reception device 38, an output device 40, a camera 42, and a communication I / F 44. The computer 36 includes a processor 46, a RAM 48, and a storage 50. The processor 46, the RAM 48, and the storage 50 are connected to a bus 52. The reception device 38, the output device 40, and the camera 42 are also connected to the bus 52.

[0030] The reception device 38 includes a touch panel 38A, a microphone 38B, and the like, and receives user input. The touch panel 38A detects contact with an indicator (for example, a pen or a finger) to receive user input by the touch of the indicator. The microphone 38B detects the user's voice to receive user input by voice. The control unit 46A transmits data indicating the user input received by the touch panel 38A and the microphone 38B to the data processing device 12. In the data processing device 12, the specific processing unit 290 acquires the data indicating the user input.

[0031] The output device 40 includes a display 40A and a speaker 40B, and presents data to the user 20 by outputting the data in a form of expression that the user 20 can perceive (for example, audio and / or text). The display 40A displays visible information such as text and images in accordance with instructions from the processor 46. The speaker 40B outputs audio in accordance with instructions from the processor 46. The camera 42 is a compact digital camera equipped with an optical system including a lens, aperture, and shutter, and an imaging element such as a CMOS (Complementary Metal-Oxide-Semiconductor) image sensor or a CCD (Charge Coupled Device) image sensor.

[0032] The communication I / F 44 is connected to a network 54. The communication I / Fs 44 and 26 control the exchange of various information between the processor 46 and the processor 28 via the network 54.

[0033] FIG. 2 shows an example of the main functions of the data processing device 12 and the smart device 14.

[0034] 2, in the data processing device 12, a specific process is performed by the processor 28. A specific processing program 56 is stored in the storage 32. The specific processing program 56 is an example of a "program" according to the technology of the present disclosure. The processor 28 reads the specific processing program 56 from the storage 32 and executes the read specific processing program 56 on the RAM 30. The specific process is realized by the processor 28 operating as a specific processing unit 290 in accordance with the specific processing program 56 executed on the RAM 30.

[0035] The storage 32 stores a data generation model 58 and an emotion identification model 59. The data generation model 58 and the emotion identification model 59 are used by the identification processing unit 290.

[0036] In the smart device 14, the processor 46 performs the reception output process. The storage 50 stores a reception output program 60. The reception output program 60 is used in conjunction with the specific processing program 56 by the data processing system 10. The processor 46 reads the reception output program 60 from the storage 50 and executes the read reception output program 60 on the RAM 48. The reception output process is realized by the processor 46 operating as the control unit 46A in accordance with the reception output program 60 executed on the RAM 48.

[0037] Next, a description will be given of the specific processing performed by the specific processing unit 290 of the data processing device 12. In the following description, the data processing device 12 will be referred to as a "server" and the smart device 14 will be referred to as a "terminal."

[0038] The present invention is a telephone fraud prevention system for elderly people that converts voice data into text data in real time, evaluates the possibility of fraud, and notifies the user with a warning. The system has the following configuration.

[0039] 1. Voice analysis module (terminal)

[0040] When a call starts, the device activates a voice analysis module and records the call in real time. The recorded voice data is then converted into text data on the spot using voice recognition technology.

[0041] 2. Text data transmission module (terminal)

[0042] The converted text data is sent to the server at regular intervals of less than one second, so that the data is sent in near real time.

[0043] 3. Fraud detection module (server)

[0044] The server analyzes the received text data, specifically matching fraud-related keywords and phrases with a list and scoring them based on their frequency and combinations.

[0045] 4. Fraud Score Evaluation Module (Server)

[0046] The server performs contextual analysis of the text data to detect speech patterns and other characteristics that are particularly characteristic of fraudulent activity. Based on this evaluation, a fraud likelihood score is calculated. If the score exceeds a set threshold, a warning flag is set.

[0047] 5. Alert notification module (server and terminal)

[0048] When the warning flag is set, the server sends a warning message to the device, which receives the message and issues a voice or text warning to the user, saying "Be careful, this may be a scam," or a text message that pops up on the screen.

[0049] Specific examples

[0050] Example 1: Fake billing scam call

[0051] 1. Users

[0052] The elderly user answers the phone and begins the conversation.

[0053] 2. Terminal

[0054] The device's voice analysis module records the call in real time and converts the audio "Payment required now" into text data.

[0055] 3. Terminal

[0056] The text data "Payment is required now" is sent to the server.

[0057] 4. Server

[0058] Analyzes text data to detect and score keywords such as "now" and "payment."

[0059] 5. Server

[0060] The fraud score assessment module gives a high score and sets a warning flag because the threshold has been exceeded.

[0061] 6. Server

[0062] Sends a warning message to the terminal.

[0063] 7. Terminal

[0064] The system will notify the user of the received warning message by voice, warning them with "Be careful, this may be a scam."

[0065] 8. Users

[0066] Users who receive the warning should end the call and consult with their family to prevent further harm.

[0067] In this way, the risk of seniors being scammed can be significantly reduced by this invention. The system operates in real time and issues immediate warnings, allowing for a rapid response.

[0068] The processing flow will be explained below.

[0069] Step 1:

[0070] The user answers the phone and starts the call. The device detects the start of the call and automatically starts the voice analysis module.

[0071] Step 2:

[0072] The device records the contents of the call in real time and instantly converts the acquired voice data into text data using voice recognition technology.

[0073] Step 3:

[0074] The terminal transmits the converted text data to the server at regular intervals (for example, every second).

[0075] Step 4:

[0076] The server receives the text data sent from the terminal and starts analysis using the text data analysis module.

[0077] Step 5:

[0078] The server detects keywords and phrases in the text data, including words related to fraud (e.g., "now," "payment," "legal action").

[0079] Step 6:

[0080] The server then scores the likelihood of fraud based on the frequency and combination of keywords detected, calculated according to pre-defined criteria.

[0081] Step 7:

[0082] The server performs contextual analysis to detect speech patterns that are particularly characteristic of fraudulent activity (forced language and emphasis on urgency).

[0083] Step 8:

[0084] The server evaluates the scoring results and sets a warning flag if the fraud likelihood score exceeds a set threshold.

[0085] Step 9:

[0086] The server notifies the terminal that a warning flag has been set, including a message indicating a high probability of fraud.

[0087] Step 10:

[0088] As soon as the device receives the warning notification from the server, it immediately issues a warning to the user, either by voice or text, saying, "Be careful, this may be a scam."

[0089] Step 11:

[0090] The user receives a warning and becomes suspicious of the call. The user ends the call and consults with family or friends for safety.

[0091] This series of steps enables the system to detect fraudulent activity with high accuracy and reduce the risk of users becoming victims.

[0092] Example 1

[0093] Next, a description will be given of Example 1. In the following description, the data processing device 12 will be referred to as a "server" and the smart device 14 will be referred to as a "terminal."

[0094] Currently, the number of elderly victims of telephone fraud is increasing, and fraudulent acts via telephone in particular has become a social problem. Measures to prevent this are needed. However, current methods are difficult to respond to in real time, and fraud prevention is not sufficient. Therefore, there is a need for a system that can reduce the risk of elderly people becoming victims of telephone fraud and issue warnings quickly and effectively.

[0095] The specific processing by the specific processing unit 290 of the data processing device 12 in the first embodiment is realized by the following means.

[0096] In this invention, the server includes means for converting voice data into text data in real time, means for transmitting the text data to the server at regular intervals, means for analyzing the text data and identifying fraud-related keywords and phrases, means for evaluating the possibility of fraud and scoring a risk level based on the analysis results, means for detecting fraud patterns based on the fraud-related keywords and phrases and calculating a fraud likelihood score, and means for issuing a voice or text warning to the user when the risk level exceeds a set threshold. This significantly reduces the risk of elderly people becoming victims of telephone fraud and enables prompt warnings to be issued in real time.

[0097] "Voice data" refers to data that records or transmits voice information such as telephone calls or conversations in digital format.

[0098] "Text data" is data in a format in which voice data is converted into a string of characters, and is information expressed in a format that can be read by humans.

[0099] "Real-time" means that the system processes voice input immediately and outputs the results without delay.

[0100] "Analysis" is the process of examining the information contained in data in detail and extracting and evaluating that information based on specific criteria and rules.

[0101] "Fraud-related keywords and phrases" refer to expressions and words characteristic of fraudulent activities, and are specific words and phrases that are judged to have a high possibility of being fraudulent.

[0102] "Scoring" refers to the process of quantifying and evaluating the likelihood of fraud based on specific criteria or algorithms based on the analysis results.

[0103] A "threshold" is a numerical value or condition that determines when a system will take a particular action or perform a particular algorithm.

[0104] A "warning" is a message or notification that warns the user and is issued when there is a possibility of fraud.

[0105] "Interval" refers to the timing or cycle at which data is transmitted, and indicates the time difference between successive data transmissions.

[0106] "Patterns" refer to speech patterns and contextual structures that are characteristic of fraudulent behavior, and are standards or models based on which to judge the possibility of fraud.

[0107] "Calculation" is the process by which a system generates a number or evaluation according to a specified algorithm or procedure.

[0108] "Voice or text notification" means that the warning to the user is given in the form of voice output or text display, and the system issues the warning in an appropriate manner.

[0109] This invention is a telephone fraud prevention system aimed at the elderly, which analyzes the voice of a call in real time and issues a warning when there is a high possibility of fraud. The specific configuration is as follows.

[0110] Voice analysis module (terminal)

[0111] User

[0112] When the elderly user receives a call, the call begins. The system starts working when the user presses the call button.

[0113] Terminal

[0114] When a call is received, the device automatically activates a voice analysis module, which records the call in real time. The recorded voice data is then converted into text data on the fly using voice recognition technology. Specifically, real-time voice-to-text conversion is performed using the Google Speech-to-Text API.

[0115] Text data transmission module (terminal)

[0116] Terminal

[0117] The converted text data is sent to the server at regular intervals (less than 1 second). The HTTP protocol is used to send the data, for example, as a POST request. Specifically, the text data is sent to http: / / server_address / text_data.

[0118] Fraud detection module (server)

[0119] server

[0120] The server analyzes the received text data. For fraud detection, it uses a pre-prepared list of fraud-related keywords to scan the text data using natural language processing (NLP) techniques. Specifically, it uses the Python nltk library to tokenize the text and verify whether each token is included in the fraud list.

[0121] For example, if someone sends a text saying "I need payment now," we tokenize it and check if keywords like "now," "payment," and "need" are included in our fraud list.

[0122] Fraud score evaluation module (server)

[0123] server

[0124] The fraud score evaluation module on the server calculates scores based on the frequency of keyword occurrence and their combinations. The scoring is done using machine learning models, utilizing Scikit-learn and TensorFlow.

[0125] "Payment required now" will receive a high score because "now" and "payment" are included in the fraud list. If the score exceeds a set threshold, a warning flag will be set.

[0126] Alert notification module (server and terminal)

[0127] server

[0128] When the warning flag is set, the server sends a warning message to the device, which may include a message like "Possible fraud. Please be careful."

[0129] Terminal

[0130] The device will notify the user of the received warning message. There are two notification methods: voice warning and text message. In the case of a voice warning, the device will use TTS (Text-to-Speech) technology to warn the user by voice, saying "Be careful, this may be a scam." In the case of a text message, it will be displayed as a pop-up on the device screen.

[0131] As a specific example of how this works, elderly users can hear a warning during a call saying, "Be careful, this may be a scam." This allows users to sense the risk of being deceived by a scam in advance and take appropriate action.

[0132] Examples of prompts include:

[0133] Describe a telephone fraud prevention system for seniors. The system converts voice data into text in real time, evaluates the likelihood of fraud, and alerts the user. Explain each step in detail, including specific operations and the names of any hardware or software used.

[0134] The flow of the identification process in the first embodiment will be described with reference to FIG.

[0135] System program processing flow

[0136] Step 1: Start the call

[0137] User

[0138] The elderly user answers the phone and starts the conversation. The user presses the talk button and the system starts working.

[0139] Input: Incoming call / call start signal

[0140] Output: System start signal

[0141] Step 2: Launching Audio Analysis

[0142] Terminal

[0143] Upon receiving the call start signal, the terminal automatically starts the voice analysis module.

[0144] How it works: The audio analysis module records the call.

[0145] Input: Call start signal

[0146] Output: Start recording audio data

[0147] Step 3: Convert the audio to text

[0148] Terminal

[0149] The speech analysis module converts the recorded voice data into text data in real time. The speech-to-text conversion is performed using the Google Speech-to-Text API.

[0150] How it works: Captures audio as digital data and sends it to an API to retrieve text.

[0151] Input: Recorded audio data

[0152] Output: Converted text data

[0153] Step 4: Send text data

[0154] Terminal

[0155] The converted text data is sent to the server at regular intervals (less than 1 second). The data is sent using an HTTP POST request.

[0156] Action: Sends data to http: / / server_address / text_data.

[0157] Input: Text data

[0158] Output: Send data to the server

[0159] Step 5: Analyzing the text data

[0160] server

[0161] The server analyzes the received text data and uses a pre-prepared list of fraud-related keywords to detect fraud. It uses the Python nltk library to tokenize the text and verify whether each token is included in the fraud list.

[0162] How it works: It breaks down text data and matches it with keywords from a scam list.

[0163] Input: Received text data

[0164] Output: Keyword occurrences

[0165] Step 6: Scoring

[0166] server

[0167] Based on the occurrence of fraud keywords, the server scores the likelihood of fraud. It calculates the fraud likelihood score using machine learning models such as Scikit-learn and TensorFlow.

[0168] How it works: Evaluates keyword frequency and patterns to calculate a fraud likelihood score.

[0169] Input: Keyword occurrence information

[0170] Output: Fraud likelihood score

[0171] Step 7: Set warning flags

[0172] server

[0173] If the fraud likelihood score exceeds a set threshold, the server sets a warning flag.

[0174] Behavior: Evaluate the score and set a flag if it exceeds a threshold.

[0175] Input: Fraud likelihood score

[0176] Output: warning flag

[0177] Step 8: Sending a warning message

[0178] server

[0179] When the warning flag is set, the server sends a warning message to the terminal, which may include a message such as "Possible fraud. Please be careful."

[0180] Action: Generates a warning message and sends it to the terminal.

[0181] Input: warning flag

[0182] Output: Warning message

[0183] Step 9: Warning Notification

[0184] Terminal

[0185] The device will notify the user of the received warning message. Notification methods include voice alert and text message. Voice alert uses TTS technology, and text message will be displayed as a pop-up on the screen.

[0186] Action: Outputs a voice or text alert based on the message content.

[0187] Input: warning message

[0188] Output: A warning notice to the user

[0189] (Application example 1)

[0190] Next, a description will be given of Application Example 1. In the following description, the data processing device 12 will be referred to as a "server" and the smart device 14 will be referred to as a "terminal."

[0191] Telephone frauds against the elderly are increasing year by year, and the amount of damage caused is also steadily increasing. Because the elderly are particularly susceptible to fraudulent schemes, there is a need for fast and effective fraud prevention measures. However, many current fraud prevention systems require users to operate and make decisions themselves, which can result in insufficient effectiveness. Furthermore, there are few systems that can prevent telephone fraud in real time, and the risk of being scammed cannot be completely eliminated. Therefore, there is a need for an effective system that can evaluate and notify users of possible fraud in real time while reducing the operational burden on users.

[0192] The specific processing by the specific processing unit 290 of the data processing device 12 in the application example 1 is realized by the following means.

[0193] In this invention, the server includes means for converting voice data into text data in real time, means for analyzing the text data and identifying keywords and phrases related to fraud, means for assessing the likelihood of fraud based on the analysis results and scoring a risk level, means for notifying a user of a warning if the risk level exceeds a set threshold, means for notifying the user of the warning using a visual display means and a voice notification means of the smart glasses, and means for performing contextual analysis of the text data and detecting patterns specific to fraudulent activities using a generative AI model, thereby significantly reducing the risk of users being involved in telephone fraud and enabling fast and effective fraud prevention in real time.

[0194] "Voice data" refers to data obtained by converting voice signals into digital format, and is the content of a conversation saved in digital format as is.

[0195] "Text data" refers to data in the form of a character string converted from audio data.

[0196] "Keywords" are specific words or phrases that are frequently used in contexts or content that are likely to be fraudulent.

[0197] "Scoring" is an evaluation method that numerically represents the likelihood of fraud.

[0198] A "threshold" is a numerical standard used to assess the likelihood of fraud, and if this value is exceeded, a warning will be issued.

[0199] A "warning" is a notification that informs the user of the risk when there is a high possibility of fraud.

[0200] "Smart glasses" are wearable devices that provide information to users through sight and sound.

[0201] "Visual display means" refers to a mechanism for displaying text and images using a device such as smart glasses.

[0202] A "generative AI model" is a natural language processing model that learns from large amounts of language data using machine learning technology.

[0203] "Contextual analysis" is a technology that understands the content of text data and extracts specific meanings and patterns.

[0204] "Pattern detection" is a technique that analyzes text data to identify speech patterns and phrases characteristic of fraudulent activity.

[0205] The present invention is a telephone fraud prevention system for elderly people that converts voice data into text data in real time, evaluates the possibility of fraud, and notifies the user with a warning. The system has the following configuration.

[0206] 1. Voice analysis module (smart glasses)

[0207] The smart glasses use a built-in microphone to capture the voice during a call in real time, and the voice data is converted into text data in real time using a common cloud speech recognition service (e.g., Google Cloud Speech-to-Text API).

[0208] 2. Text data transmission module (smart glasses)

[0209] The converted text data is sent to a cloud server at regular intervals (less than one second). The cloud server is equipped with a high-performance data processing system, enabling rapid data processing.

[0210] 3. Fraud detection module (cloud server)

[0211] The cloud server analyzes the received text data using natural language processing libraries (e.g., spaCy, NLTK) and scores the frequency of occurrence of fraud-related keywords and phrases. This analysis is powered by a powerful text analysis engine.

[0212] 4. Fraud score evaluation module (cloud server)

[0213] The server then performs contextual analysis using generative AI models (e.g., BERT or GPT-4) to detect patterns characteristic of fraudulent activity, which then calculates a fraud likelihood score and sets a warning flag if the score exceeds a threshold.

[0214] 5. Alert notification module (cloud server and smart glasses)

[0215] If the fraud score exceeds a threshold, the cloud server generates a warning message and sends it to the smart glasses, which then alert the user using visual (HUD display) and audio notification means. For example, a text pop-up on the HUD display reads, "Beware, this may be a scam," along with an audio notification stating the same.

[0216] Specific examples

[0217] Example 1: Fake billing scam call

[0218] 1. Users

[0219] The elderly user answers the phone and begins the conversation.

[0220] 2. Smart Glasses

[0221] The voice analysis module records the call in real time and converts the audio "Payment required now" into text data.

[0222] 3. Smart Glasses

[0223] The converted text data "Payment required now" is sent to the cloud server.

[0224] 4. Cloud Server

[0225] The text data is analyzed to detect keywords such as "now" and "payment" and then scored.

[0226] 5. Cloud Server

[0227] The fraud score assessment module gives a high rating and sets a warning flag because the threshold has been exceeded.

[0228] 6. Cloud Server

[0229] Sending a warning message to the smart glasses.

[0230] 7. Smart Glasses

[0231] The user will be notified of the received warning message via voice, warning them "Beware, this may be a scam," and the same text will pop up on the HUD display.

[0232] 8. Users

[0233] Upon receiving the warning, the user ends the call and consults with a family member.

[0234] Prompt Sentence Examples

[0235] "Please transcribe what was said on this call and assess the likelihood of fraud."

[0236] "Please rate the following text for fraud-related keywords: 'Payment required now'"

[0237] Please analyze the following conversation for possible fraud risk.

[0238] With the above configuration, this system significantly reduces the risk of elderly people falling victim to telephone fraud, enabling fast and effective fraud prevention in real time.

[0239] The flow of the specific processing in the application example 1 will be described with reference to FIG.

[0240] Step 1:

[0241] When a user answers a call and starts a conversation, the voice analysis module in the smart glasses is automatically activated. The voice data during the call is taken as input. The microphone in the smart glasses captures the voice and converts the voice data into text data in real time using a cloud speech recognition service (e.g., Google Cloud Speech-to-Text API). Text data is generated as output, which may include a message such as "Payment required now."

[0242] Step 2:

[0243] The text data transmission module of the smart glasses transmits the converted text data to the cloud server at regular intervals (less than 1 second). The text data generated earlier is used as input. The text data is sent to the cloud server as output.

[0244] Step 3:

[0245] The fraud detection module on the cloud server analyzes the received text data. Specifically, it uses natural language processing libraries (e.g., spaCy, NLTK) to identify fraud-related keywords and phrases. The input is the text data sent to the server. The output is the detection results of fraud-related keywords and phrases (e.g., "now," "pay," etc.).

[0246] Step 4:

[0247] The fraud score evaluation module on the cloud server evaluates the likelihood of fraud based on the analysis results and assigns a risk score. A generative AI model (e.g., BERT or GPT-4) is used to perform contextual analysis and pattern detection on the text data. The input is the output of the fraud detection module. The output is a fraud likelihood score, expressed as a number.

[0248] Step 5:

[0249] If the fraud score exceeds a threshold, the cloud server sets a warning flag. The input is the output of the fraud score evaluation module. The output is the setting state of the warning flag.

[0250] Step 6:

[0251] If the warning flag is set, the cloud server generates a warning message and sends it to the smart glasses. The input is the setting state of the warning flag. The output is the warning message (e.g., "Be careful, this may be a scam").

[0252] Step 7:

[0253] The smart glasses notify the user of the received warning message. The HUD display is used as the visual notification means, and the built-in speaker is used as the audio notification means. The input is the warning message sent from the cloud server. The output is to display the warning to the user both audio and visually.

[0254] Step 8:

[0255] When the user receives the warning message, they end the call and consult with their family to avoid the risk of fraud. The input is the warning message from the smart glasses, and the output is the user's action to take to avoid fraud.

[0256] Furthermore, an emotion engine that estimates the user's emotion may be combined. That is, the identification processing unit 290 may estimate the user's emotion using the emotion identification model 59 and perform identification processing using the user's emotion.

[0257] The present invention provides a more accurate warning function by combining a system that converts voice data into text data in real time, evaluates the possibility of fraud, and notifies the user of a warning, with an emotion engine that recognizes the user's emotions. This system has the following configuration.

[0258] 1. Voice analysis module (terminal)

[0259] When a call starts, the device activates a voice analysis module and records the call in real time. The recorded voice data is then converted into text data on the spot using voice recognition technology.

[0260] 2. Text data transmission module (terminal)

[0261] The converted text data is sent to the server at regular intervals of less than one second, so that the data is sent in near real time.

[0262] 3. Fraud detection module (server)

[0263] The server analyzes the received text data, specifically matching fraud-related keywords and phrases with a list and scoring them based on their frequency and combinations.

[0264] 4. Fraud Score Evaluation Module (Server)

[0265] The server performs contextual analysis of the text data to detect speech patterns and other characteristics that are particularly characteristic of fraudulent activity. Based on this evaluation, a fraud likelihood score is calculated. If the score exceeds a set threshold, a warning flag is set.

[0266] 5. Emotion engine (server)

[0267] The server uses the voice data to recognize the user's emotional state: the emotion engine analyzes the tone, rate, and other voice characteristics to determine whether the user is experiencing emotions such as surprise, fear, or anxiety.

[0268] 6. Alert Adjustment Module (Server)

[0269] The server adjusts the content and notification method of the warning based on the analysis results of the emotion engine. For example, if the user is feeling strong anxiety or fear, the warning content can be made more detailed and clearer.

[0270] 7. Alert notification module (server and terminal)

[0271] If the warning flag is set, the server sends a warning message to the device, which receives the message and issues a voice or text warning to the user, either saying "Be careful, this may be a scam" in the case of a voice warning or a text message that pops up on the screen.

[0272] Specific examples

[0273] Example 1: Call fraud using emotion engine

[0274] 1. Users

[0275] The elderly user answers the phone and begins the conversation.

[0276] 2. Terminal

[0277] The device's voice analysis module records the call in real time and converts the audio "Payment required now" into text data.

[0278] 3. Terminal

[0279] The text data "Payment is required now" is sent to the server.

[0280] 4. Server

[0281] Analyzes text data to detect and score keywords such as "now" and "payment."

[0282] 5. Server

[0283] The system analyzes the urgent speech patterns characteristic of fraudulent activity and gives a high rating. If this exceeds a threshold, a warning flag is set.

[0284] 6. Server

[0285] At the same time, the emotion engine analyzes the voice data to recognize the user's emotional state, detecting when the user feels an urgent need and is anxious.

[0286] 7. Server

[0287] The warning content is adjusted based on the results of the emotion engine to generate stronger warning messages.

[0288] 8. Server

[0289] Sends tailored warning messages to the terminal.

[0290] 9. Terminal

[0291] The system will notify the user of the received warning message with a voice message, warning them, "Be careful, this may be a scam. Please tell a family member about this call immediately."

[0292] 10. Users

[0293] Users who receive the warning can end the call and consult with their family to prevent any harm.

[0294] In this way, the present invention can significantly reduce the risk of seniors becoming victims of fraud. The system operates in real time and issues appropriate warnings based on the user's emotional state, allowing for a fast and effective response.

[0295] The processing flow will be explained below.

[0296] Step 1:

[0297] The user answers the phone and starts the call. The device detects the start of the call and automatically starts the voice analysis module.

[0298] Step 2:

[0299] The device records the contents of the call in real time and instantly converts the acquired voice data into text data using voice recognition technology.

[0300] Step 3:

[0301] The terminal transmits the converted text data to the server at regular intervals (for example, every second).

[0302] Step 4:

[0303] The server receives the text data sent from the terminal and starts analysis using the text data analysis module.

[0304] Step 5:

[0305] The server detects keywords and phrases in the text data, including words related to fraud (e.g., "now," "payment," "legal action").

[0306] Step 6:

[0307] The server then scores the likelihood of fraud based on the frequency and combination of keywords detected, calculated according to pre-defined criteria.

[0308] Step 7:

[0309] The server performs contextual analysis to detect speech patterns that are particularly characteristic of fraudulent activity (forced language and emphasis on urgency).

[0310] Step 8:

[0311] The server evaluates the scoring results and sets a warning flag if the fraud likelihood score exceeds a set threshold.

[0312] Step 9:

[0313] The server sends the voice data to the emotion engine, which analyzes the tone, rate, and other voice characteristics of the user's voice.

[0314] Step 10:

[0315] The emotion engine on the server determines the user's emotional state, for example, analyzing whether the user is feeling surprise, fear, anxiety, etc.

[0316] Step 11:

[0317] The server adjusts the content of the warning and notification method based on the analysis results of the emotion engine. If the user is feeling strong anxiety or fear, the warning content will be more detailed and clearer.

[0318] Step 12:

[0319] The server sends a tailored warning message to the device, which includes information about the likely fraud and specific steps to take.

[0320] Step 13:

[0321] As soon as the device receives the warning notification from the server, it immediately issues a warning to the user, either by voice or text, saying, "Caution, this may be a scam. Please inform a family member immediately about this call."

[0322] Step 14:

[0323] The user receives a warning and becomes suspicious of the call. They end the call and consult with family or friends for safety. This process significantly reduces the risk of the user becoming a victim of fraud.

[0324] Example 2

[0325] Next, a description will be given of Example 2. In the following description, the data processing device 12 will be referred to as a "server" and the smart device 14 will be referred to as a "terminal."

[0326] In today's world, fraud is becoming increasingly sophisticated, making the elderly especially vulnerable to fraud. Previous fraud prevention systems primarily relied on text data analysis and did not provide real-time warnings that took the user's emotional state into account. This made it difficult for users to quickly recognize fraud and take appropriate action. Furthermore, because conventional systems issued warnings uniformly, they often failed to respond appropriately to the sense of crisis and anxiety felt by users. This has led to a need for effective methods to prevent fraud before it happens.

[0327] The identification process by the identification processing unit 290 of the data processing device 12 in Example 2 is realized by the following means. In this invention, the server includes means for converting voice data into text data in real time, means for analyzing the text data and identifying keywords and phrases related to fraud, means for evaluating the possibility of fraud based on the analysis results and scoring the risk level, means for recognizing the user's emotional state from the voice analysis and adjusting the warning content based on the emotion, and means for transmitting the warning content to the terminal and notifying the user. This makes it possible to provide an appropriate warning in real time that takes into account the user's emotional state.

[0328] "Voice data" refers to sound waveforms such as those from a call or voice input that have been converted into electrical signals or digital data.

[0329] "Text data" refers to character string data generated as a result of analyzing and converting voice data.

[0330] "Analysis" is the process of dividing data for a specific purpose to reveal meaning and structure.

[0331] "Fraud-related keywords and phrases" refer to specific words and expressions that are likely to be used in fraud, and are used to predict the possibility of fraud by detecting them.

[0332] "Risk level" is a numerical or score that expresses the likelihood of a scam.

[0333] "Scoring" is the process of quantifying risk based on the results of data analysis.

[0334] A "threshold" is a numerical value or condition that triggers a specific action, and when exceeded, a warning is issued.

[0335] A "warning" is information that notifies the user of a potential danger and urges caution.

[0336] "Emotional state" refers to the user's psychological state analyzed from the tone, rate, and other characteristics of the voice.

[0337] "Adjustment" is the process of changing something to an appropriate form depending on the situation or conditions.

[0338] A "notification" is an action or means of informing a user of specific information.

[0339] The present invention provides a more accurate warning function by combining a system that converts voice data into text data in real time, evaluates the possibility of fraud, and notifies the user of a warning with an emotion engine that recognizes the user's emotions.

[0340] System Configuration

[0341] The system includes the following modules:

[0342] Voice analysis module (terminal)

[0343] The device activates a voice analysis module at the start of a call and records the call in real time. The recorded voice data is converted to text data using the Google Cloud Speech-to-Text API.

[0344] Text data transmission module (terminal)

[0345] The converted text data is sent to the server at regular intervals of less than one second using the WebSocket protocol in real time.

[0346] Fraud detection module (server)

[0347] The server analyzes the received text data using a natural language processing library (e.g., NLTK or SpaCy), matching fraud-related keywords and phrases with a list, and scoring them based on their frequency and combinations.

[0348] Fraud score evaluation module (server)

[0349] The server then performs contextual analysis of the text data to detect speech patterns and other characteristics of fraudulent activity. This information is used to calculate a fraud likelihood score. If this score exceeds a set threshold, a warning flag is set.

[0350] Emotion engine (server)

[0351] The server uses the voice data to recognize the user's emotional state, using IBM Watson Tone Analyzer and Microsoft Azure Emotion API to determine whether the user is experiencing emotions such as surprise, fear, or anxiety.

[0352] Alert Adjustment Module (Server)

[0353] The server adjusts the content and notification method of the warning based on the results of the emotion engine. For example, if the user is feeling strong anxiety or fear, the warning content will be more detailed and clearer.

[0354] Alert notification module (server and terminal)

[0355] The server sends the tailored warning message to the device, which receives it and issues a voice or text warning to the user, such as "Be careful, this may be a scam," or a text message that pops up on the screen.

[0356] Specific examples

[0357] Example 1: Call fraud using emotion engine

[0358] 1. Users

[0359] The elderly user answers the phone and begins the conversation.

[0360] 2. Terminal

[0361] The device's voice analysis module records the call in real time and instantly converts the audio "Payment required now" into text data using Google Cloud Speech-to-Text.

[0362] 3. Terminal

[0363] The converted text data "Payment required now" is sent to the server via the WebSocket protocol at intervals of less than one second.

[0364] 4. Server

[0365] The server uses NLTK to analyze the received text data, identifying fraud-related keywords such as "now" and "payment," and analyzing the frequency and context of these keywords.

[0366] 5. Server

[0367] It detects urgent speech patterns characteristic of fraudulent activity and generates a high fraud likelihood score. If this score exceeds a set threshold, it sets a warning flag.

[0368] 6. Server

[0369] The server uses an emotion engine (IBM Watson Tone Analyzer) to analyze the user's emotional state from the voice data. The analysis results indicate that the user feels urgency and is anxious.

[0370] 7. Server

[0371] The content of the warning message is adjusted based on the results of the emotion engine, for example, generating a detailed warning such as "This call may be fraudulent. Please report this call to a family member immediately."

[0372] 8. Server

[0373] Sends tailored warning messages to the terminal.

[0374] 9. Terminal

[0375] The device will then notify the user of the received warning message via voice: "Caution, this may be a scam. Please inform a family member immediately about this call."

[0376] 10. Users

[0377] The user will receive a warning, end the call, and immediately contact their family to prevent any harm.

[0378] In this way, the present invention can significantly reduce the risk of seniors becoming victims of fraud. The system operates in real time and issues appropriate warnings based on the user's emotional state, allowing for a fast and effective response.

[0379] Prompt Sentence Examples

[0380] "Design a system to detect emotions from voice data, assess the likelihood of fraud, and provide a warning. Use natural language processing and emotion recognition technologies for data analysis. Please also specify the specific hardware and software, data flow, and user interface (UI)."

[0381] The flow of the identification process in the second embodiment will be described with reference to FIG.

[0382] Step 1:

[0383] Voice recording and text conversion (device)

[0384] Input: A user receives a call and starts a conversation.

[0385] How it works: The device detects that a call has been initiated and activates the voice analysis module.

[0386] Data processing: Convert the audio data into text data in real time using the Google Cloud Speech-to-Text API.

[0387] Output: The converted text data is generated.

[0388] Step 2:

[0389] Sending text data (terminal)

[0390] Input: The text data generated in step 1.

[0391] Operation: The device sends the converted text data to the server at intervals of less than one second.

[0392] Data calculation: Send data in real time using the WebSocket protocol.

[0393] Output: The server receives text data in real time.

[0394] Step 3:

[0395] Text data analysis (server)

[0396] Input: The text data received in step 2.

[0397] How it works: The server parses the received text data using a natural language processing library (NLTK or SpaCy).

[0398] Data calculations: Fraud-related keywords and phrases are matched against the list and analyzed based on their frequency and combinations.

[0399] Output: A numerical score is generated representing the likelihood of fraud.

[0400] Step 4:

[0401] Fraud score evaluation (server)

[0402] Input: Analysis results (scores) generated in step 3.

[0403] How it works: The server also performs contextual analysis of the text data to detect speech patterns and patterns characteristic of fraudulent activity.

[0404] Data calculation: Calculate a fraud likelihood score and set a warning flag if this score exceeds a set threshold.

[0405] Output: Fraud probability score and warning flag status.

[0406] Step 5:

[0407] User emotion analysis (server)

[0408] Input: The audio data recorded in step 1.

[0409] How it works: The server uses IBM Watson Tone Analyzer and Microsoft Azure Emotion API to analyze the user's emotional state from voice data.

[0410] Data calculations: Analyze the tone, rate, and other voice characteristics of the voice to determine whether the user is surprised, scared, or anxious.

[0411] Output: Generates an analysis of the user's emotional state.

[0412] Step 6:

[0413] Adjustment of warning content (server)

[0414] Input: Warning flags from step 4 and sentiment analysis results from step 5.

[0415] How it works: The server tailors the warning based on the user's emotional state. For example, if the user is feeling strong anxiety or fear, the warning will be more detailed and clearer.

[0416] Data calculation: Optimize the content and format of warning messages based on the results of sentiment analysis.

[0417] Output: A tailored warning message is generated.

[0418] Step 7:

[0419] Notification of warning messages (server and terminal)

[0420] Input: The warning message adjusted in step 6.

[0421] Action: The server sends a tailored alert message to the device.

[0422] Data calculation: The device notifies the user of the received warning message by voice or text.

[0423] Output: A voice or text alert is given to the user.

[0424] These steps enable the system to detect potential fraud in real time and issue warnings that take into account the user's emotional state.

[0425] (Application example 2)

[0426] Next, a description will be given of Application Example 2. In the following description, the data processing device 12 will be referred to as a "server" and the smart device 14 will be referred to as a "terminal."

[0427] In recent years, telephone fraud has been on the rise, with elderly people becoming victims in particular. Conventional fraud prevention measures only issue a warning to users when fraud is suspected. Therefore, even when users receive a warning, they are unable to properly determine how to respond, resulting in a high likelihood of fraud victimization. Furthermore, because warnings are issued uniformly with the same intensity without taking into account the user's emotional state, the effectiveness of the warnings is limited. There is a need to improve this situation and provide more effective fraud prevention measures.

[0428] The specific processing by the specific processing unit 290 of the data processing device 12 in the application example 2 is realized by the following means.

[0429] In this invention, the server includes means for converting voice data into text data in real time, means for analyzing the text data and identifying keywords and phrases related to fraud, means for assessing the likelihood of fraud and scoring the risk based on the analysis results, means for issuing a warning to the user when the risk exceeds a set threshold, means for analyzing the user's emotional state, means for adjusting the content of the warning based on the emotion analysis results, and means for issuing the adjusted warning to the user. This makes it possible to accurately assess the likelihood of fraud in real time based on the voice data and provide an appropriate warning that takes the user's emotional state into consideration.

[0430] "Voice data" means recorded information of a user's voice uttered over the telephone or other voice communication.

[0431] "Real-time" refers to near-instant processing with little actual time or delay.

[0432] "Text data" is voice data converted into character information.

[0433] Analysis is the process of examining data or information in detail to identify important elements or patterns within it.

[0434] "Fraud-related keywords and phrases" are specific words and expressions commonly used in fraudulent activities.

[0435] "Scoring" is the process of assigning points according to specific criteria based on the analysis results.

[0436] A "threshold" is a boundary value for determining a certain condition or standard.

[0437] A "warning" is a notification that notifies the user of an impending danger.

[0438] An "emotional state" is the psychological state that a user is feeling at a particular moment.

[0439] "Notification" is the act of conveying information or a message to a target person.

[0440] "Adjustment" refers to changing something to an optimal state for specific conditions or environments.

[0441] This invention is a system for protecting users from telephone fraud, and provides highly accurate warning functions by combining analysis of voice data and emotion analysis.

[0442] System Overview

[0443] The system mainly consists of three main components.

[0444] 1. Terminal

[0445] 2. Server

[0446] 3. Users

[0447] Hardware and software used

[0448] Hardware:

[0449] Smartphone: Used as a device, requires a microphone and internet connection.

[0450] software:

[0451] Python 3.x: The main programming language.

[0452] speech_recognition library: Uses Google's speech recognition API to convert voice data into text data.

[0453] The requests library: handles HTTP requests to the server.

[0454] Processing Details

[0455] Terminal handling

[0456] When a user initiates a phone call, the device records the voice data in real time and converts it into text data using the speech_recognition library, which is then sent to the server at regular intervals.

[0457] Server Processing

[0458] The server analyzes the audio data and assesses the likelihood of fraud using the following steps:

[0459] 1. Text Analysis: Analyzes incoming text data to identify keywords and phrases related to fraud.

[0460] 2. Scoring: Based on the analysis results, the possibility of fraud is assessed and a risk score is generated. If the score exceeds a set threshold, a warning is generated.

[0461] 3. Emotion analysis: Analyze the user's emotional state from the voice data, determining whether the user is feeling surprised or anxious.

[0462] 4. Adjustment of warning content: Based on the results of sentiment analysis, the warning content is adjusted appropriately.

[0463] Warning Notification

[0464] A tailored warning message is sent to the user's device. For example, if the warning is delivered via voice, the content and intensity of the warning can be adjusted according to the user's emotional state. This is expected to encourage the user to take appropriate action.

[0465] Specific examples

[0466] For example, consider the case of an elderly person receiving a suspected fraudulent phone call. If the user hears phrases such as "Payment is required now," the system immediately transcribes the speech into text and analyzes it. If it detects patterns indicative of fraudulent activity, the system takes into account the user's emotional state and generates a strong warning message. The warning could include specific advice such as "Be careful, this may be a scam. Please report this call to a family member immediately."

[0467] Example prompts for generative AI models

[0468] Your system should take voice data from a user's phone call as input and convert it into text data in real time. It should then assess the likelihood of fraud (by scoring it based on the frequency of fraud-related keywords and phrases) and generate a warning message if necessary. It should also analyze the user's emotional state and adjust the strength and content of the warning. The emotion engine uses the tone, rate, and other voice characteristics to determine whether the user is feeling anxious or scared. It should generate a warning message based on the results of the text and emotion analysis and notify the user.

[0469] The flow of the specific processing in the application example 2 will be described with reference to FIG.

[0470] Step 1:

[0471] A user initiates a phone call.

[0472] Input: User's call audio

[0473] Output: Call audio data

[0474] Specific operation: A user initiates a call and voice communication begins over the phone.

[0475] Step 2:

[0476] The device records the call audio in real time.

[0477] Input: Call audio data

[0478] Output: Recorded audio file

[0479] Specific operation: The device's microphone captures the user's call voice and records the voice data.

[0480] Step 3:

[0481] The device converts recorded voice data into text data in real time.

[0482] Input: Recorded audio file

[0483] Output: Text data

[0484] Specific operation: Using the speech_recognition library installed on the device, the recorded voice data is converted into text data.

[0485] Step 4:

[0486] The terminal transmits the converted text data to the server at regular intervals.

[0487] Input: Text data

[0488] Output: Send text data to the server

[0489] Specific operation: Sends text data to the server using an HTTP request, for example, using the requests library.

[0490] Step 5:

[0491] The server analyzes the received text data to identify keywords and phrases related to fraud.

[0492] Input: Text data sent to the server

[0493] Output: Analysis results (list of fraud-related keywords and phrases)

[0494] How it works: An analytics module on the server scans the text data to identify keywords and phrases related to fraud, matching them against a predefined list of fraud keywords.

[0495] Step 6:

[0496] The server evaluates the likelihood of fraud and assigns a score based on the results of analyzing the text data.

[0497] Input: Analysis results

[0498] Output: Risk score

[0499] Specific operation: The server's scoring module evaluates the likelihood of fraud based on the frequency of fraud keywords and the context of the text, and calculates a risk score.

[0500] Step 7:

[0501] The server generates a warning message if the risk score exceeds a threshold.

[0502] Input: Risk score

[0503] Output: Warning message

[0504] Specific behavior: If the score exceeds a preset threshold, an algorithm is executed that generates a warning message.

[0505] Step 8:

[0506] The server analyzes the user's emotional state from the voice data.

[0507] Input: Audio data

[0508] Output: Emotion analysis results

[0509] Specific behavior: The emotion engine analyzes the tone, rate, and other voice characteristics of the voice to assess the user's emotional state.

[0510] Step 9:

[0511] The server adjusts the warning content based on the results of sentiment analysis.

[0512] Input: Sentiment analysis results, warning message

[0513] Output: Adjusted warning message

[0514] Specific operation: Using the results of the emotion engine, an algorithm is run that adjusts the content and intensity of warning messages.

[0515] Step 10:

[0516] The server sends the adjusted warning message to the terminal, and the terminal notifies the user of the warning.

[0517] Input: Adjusted warning message

[0518] Output: A warning notice to the user

[0519] Specific operation: The server sends a tailored warning message to the terminal, which then notifies the user of the message. The warning is notified as voice or text.

[0520] Each step is performed sequentially, protecting users from the risk of fraud in real time.

[0521] The specific processing unit 290 transmits the result of the specific processing to the smart device 14. In the smart device 14, the control unit 46A causes the output device 40 to output the result of the specific processing. The microphone 38B acquires audio indicating a user input regarding the result of the specific processing. The control unit 46A transmits audio data indicating the user input acquired by the microphone 38B to the data processing device 12. In the data processing device 12, the specific processing unit 290 acquires the audio data.

[0522] The data generation model 58 is a so-called generative AI (Artificial Intelligence). An example of the data generation model 58 is ChatGPT (Internet Search<URL: https: / / openai.com / blog / chatgpt> ), Gemini (Internet search <url: https: gemini.google.com ?hl="ja">) and other generation AIs. The data generation model 58 is obtained by performing deep learning on a neural network. A prompt including an instruction is input to the data generation model 58, and inference data such as voice data indicating voice, text data indicating text, and image data indicating an image is also input. The data generation model 58 performs inference on the input inference data in accordance with the instruction indicated by the prompt, and outputs the inference result in a data format such as voice data and text data. Here, inference refers to, for example, analysis, classification, prediction, and / or summarization.

[0523] In the above embodiment, an example in which the specific process is performed by the data processing device 12 has been given, but the technology of the present disclosure is not limited to this, and the specific process may be performed by the smart device 14.

[0524] [Second embodiment]

[0525] FIG. 3 shows an example of the configuration of a data processing system 210 according to the second embodiment.

[0526] 3, the data processing system 210 includes the data processing device 12 and smart glasses 214. An example of the data processing device 12 is a server.

[0527] The data processing device 12 includes a computer 22, a database 24, and a communication I / F 26. The computer 22 is an example of a "computer" according to the technology of the present disclosure. The computer 22 includes a processor 28, a RAM 30, and a storage 32. The processor 28, the RAM 30, and the storage 32 are connected to a bus 34. The database 24 and the communication I / F 26 are also connected to the bus 34. The communication I / F 26 is connected to a network 54. Examples of the network 54 include a WAN (Wide Area Network) and / or a LAN (Local Area Network).

[0528] The smart glasses 214 include a computer 36, a microphone 238, a speaker 240, a camera 42, and a communication I / F 44. The computer 36 includes a processor 46, a RAM 48, and a storage 50. The processor 46, the RAM 48, and the storage 50 are connected to a bus 52. The microphone 238, the speaker 240, and the camera 42 are also connected to the bus 52.

[0529] The microphone 238 receives instructions and the like from the user 20 by receiving voice uttered by the user 20. The microphone 238 captures the voice uttered by the user 20, converts the captured voice into audio data, and outputs it to the processor 46. The speaker 240 outputs audio in accordance with instructions from the processor 46.

[0530] Camera 42 is a small digital camera equipped with an optical system including a lens, aperture, and shutter, and an imaging element such as a CMOS (Complementary Metal-Oxide-Semiconductor) image sensor or a CCD (Charge Coupled Device) image sensor, and captures images of the surroundings of user 20 (for example, an imaging range defined by an angle of view equivalent to the field of vision of a typical healthy person).

[0531] The communication I / F 44 is connected to a network 54. The communication I / Fs 44 and 26 are responsible for the exchange of various information between the processor 46 and the processor 28 via the network 54. The exchange of various information between the processor 46 and the processor 28 using the communication I / Fs 44 and 26 is carried out in a secure state.

[0532] Fig. 4 shows an example of the main functions of the data processing device 12 and the smart glasses 214. As shown in Fig. 4, in the data processing device 12, a specific process is performed by the processor 28. A specific process program 56 is stored in the storage 32.

[0533] The specific processing program 56 is an example of a "program" according to the technology of the present disclosure. The processor 28 reads the specific processing program 56 from the storage 32 and executes the read specific processing program 56 on the RAM 30. The specific processing is realized by the processor 28 operating as a specific processing unit 290 in accordance with the specific processing program 56 executed on the RAM 30.

[0534] The storage 32 stores a data generation model 58 and an emotion identification model 59. The data generation model 58 and the emotion identification model 59 are used by the identification processing unit 290.

[0535] In the smart glasses 214, the reception output process is performed by the processor 46. A reception output program 60 is stored in the storage 50. The processor 46 reads the reception output program 60 from the storage 50 and executes the read reception output program 60 on the RAM 48. The reception output process is realized by the processor 46 operating as the control unit 46A in accordance with the reception output program 60 executed on the RAM 48.

[0536] Next, a description will be given of the identification process performed by the identification processing unit 290 of the data processing device 12. In the following description, the data processing device 12 will be referred to as the "server" and the smart glasses 214 will be referred to as the "terminal."

[0537] The present invention is a telephone fraud prevention system for elderly people that converts voice data into text data in real time, evaluates the possibility of fraud, and notifies the user with a warning. The system has the following configuration.

[0538] 1. Voice analysis module (terminal)

[0539] When a call starts, the device activates a voice analysis module and records the call in real time. The recorded voice data is then converted into text data on the spot using voice recognition technology.

[0540] 2. Text data transmission module (terminal)

[0541] The converted text data is sent to the server at regular intervals of less than one second, so that the data is sent in near real time.

[0542] 3. Fraud detection module (server)

[0543] The server analyzes the received text data, specifically matching fraud-related keywords and phrases with a list and scoring them based on their frequency and combinations.

[0544] 4. Fraud Score Evaluation Module (Server)

[0545] The server performs contextual analysis of the text data to detect speech patterns and other characteristics that are particularly characteristic of fraudulent activity. Based on this evaluation, a fraud likelihood score is calculated. If the score exceeds a set threshold, a warning flag is set.

[0546] 5. Alert notification module (server and terminal)

[0547] When the warning flag is set, the server sends a warning message to the device, which receives the message and issues a voice or text warning to the user, saying "Be careful, this may be a scam," or a text message that pops up on the screen.

[0548] Specific examples

[0549] Example 1: Fake billing scam call

[0550] 1. Users

[0551] The elderly user answers the phone and begins the conversation.

[0552] 2. Terminal

[0553] The device's voice analysis module records the call in real time and converts the audio "Payment required now" into text data.

[0554] 3. Terminal

[0555] The text data "Payment is required now" is sent to the server.

[0556] 4. Server

[0557] Analyzes text data to detect and score keywords such as "now" and "payment."

[0558] 5. Server

[0559] The fraud score assessment module gives a high score and sets a warning flag because the threshold has been exceeded.

[0560] 6. Server

[0561] Sends a warning message to the terminal.

[0562] 7. Terminal

[0563] The system will notify the user of the received warning message by voice, warning them with "Be careful, this may be a scam."

[0564] 8. Users

[0565] Users who receive the warning should end the call and consult with their family to prevent further harm.

[0566] In this way, the risk of seniors being scammed can be significantly reduced by this invention. The system operates in real time and issues immediate warnings, allowing for a rapid response.

[0567] The processing flow will be explained below.

[0568] Step 1:

[0569] The user answers the phone and starts the call. The device detects the start of the call and automatically starts the voice analysis module.

[0570] Step 2:

[0571] The device records the contents of the call in real time and instantly converts the acquired voice data into text data using voice recognition technology.

[0572] Step 3:

[0573] The terminal transmits the converted text data to the server at regular intervals (for example, every second).

[0574] Step 4:

[0575] The server receives the text data sent from the terminal and starts analysis using the text data analysis module.

[0576] Step 5:

[0577] The server detects keywords and phrases in the text data, including words related to fraud (e.g., "now," "payment," "legal action").

[0578] Step 6:

[0579] The server then scores the likelihood of fraud based on the frequency and combination of keywords detected, calculated according to pre-defined criteria.

[0580] Step 7:

[0581] The server performs contextual analysis to detect speech patterns that are particularly characteristic of fraudulent activity (forced language and emphasis on urgency).

[0582] Step 8:

[0583] The server evaluates the scoring results and sets a warning flag if the fraud likelihood score exceeds a set threshold.

[0584] Step 9:

[0585] The server notifies the terminal that a warning flag has been set, including a message indicating a high probability of fraud.

[0586] Step 10:

[0587] As soon as the device receives the warning notification from the server, it immediately issues a warning to the user, either by voice or text, saying, "Be careful, this may be a scam."

[0588] Step 11:

[0589] The user receives a warning and becomes suspicious of the call. The user ends the call and consults with family or friends for safety.

[0590] This series of steps enables the system to detect fraudulent activity with high accuracy and reduce the risk of users becoming victims.

[0591] Example 1

[0592] Next, a description will be given of Example 1. In the following description, the data processing device 12 will be referred to as a "server" and the smart glasses 214 will be referred to as a "terminal."

[0593] Currently, the number of elderly victims of telephone fraud is increasing, and fraudulent acts via telephone in particular has become a social problem. Measures to prevent this are needed. However, current methods are difficult to respond to in real time, and fraud prevention is not sufficient. Therefore, there is a need for a system that can reduce the risk of elderly people becoming victims of telephone fraud and issue warnings quickly and effectively.

[0594] The specific processing by the specific processing unit 290 of the data processing device 12 in the first embodiment is realized by the following means.

[0595] In this invention, the server includes means for converting voice data into text data in real time, means for transmitting the text data to the server at regular intervals, means for analyzing the text data and identifying fraud-related keywords and phrases, means for evaluating the possibility of fraud and scoring a risk level based on the analysis results, means for detecting fraud patterns based on the fraud-related keywords and phrases and calculating a fraud likelihood score, and means for issuing a voice or text warning to the user when the risk level exceeds a set threshold. This significantly reduces the risk of elderly people becoming victims of telephone fraud and enables prompt warnings to be issued in real time.

[0596] "Voice data" refers to data that records or transmits voice information such as telephone calls or conversations in digital format.

[0597] "Text data" is data in a format in which voice data is converted into a string of characters, and is information expressed in a format that can be read by humans.

[0598] "Real-time" means that the system processes voice input immediately and outputs the results without delay.

[0599] "Analysis" is the process of examining the information contained in data in detail and extracting and evaluating that information based on specific criteria and rules.

[0600] "Fraud-related keywords and phrases" refer to expressions and words characteristic of fraudulent activities, and are specific words and phrases that are judged to have a high possibility of being fraudulent.

[0601] "Scoring" refers to the process of quantifying and evaluating the likelihood of fraud based on specific criteria or algorithms based on the analysis results.

[0602] A "threshold" is a numerical value or condition that determines when a system will take a particular action or perform a particular algorithm.

[0603] A "warning" is a message or notification that warns the user and is issued when there is a possibility of fraud.

[0604] "Interval" refers to the timing or cycle at which data is transmitted, and indicates the time difference between successive data transmissions.

[0605] "Patterns" refer to speech patterns and contextual structures that are characteristic of fraudulent behavior, and are standards or models based on which to judge the possibility of fraud.

[0606] "Calculation" is the process by which a system generates a number or evaluation according to a specified algorithm or procedure.

[0607] "Voice or text notification" means that the warning to the user is given in the form of voice output or text display, and the system issues the warning in an appropriate manner.

[0608] This invention is a telephone fraud prevention system aimed at the elderly, which analyzes the voice of a call in real time and issues a warning when there is a high possibility of fraud. The specific configuration is as follows.

[0609] Voice analysis module (terminal)

[0610] User

[0611] When the elderly user receives a call, the call begins. The system starts working when the user presses the call button.

[0612] Terminal

[0613] When a call is received, the device automatically activates a voice analysis module, which records the call in real time. The recorded voice data is then converted into text data on the fly using voice recognition technology. Specifically, real-time voice-to-text conversion is performed using the Google Speech-to-Text API.

[0614] Text data transmission module (terminal)

[0615] Terminal

[0616] The converted text data is sent to the server at regular intervals (less than 1 second). The HTTP protocol is used to send the data, for example, as a POST request. Specifically, the text data is sent to http: / / server_address / text_data.

[0617] Fraud detection module (server)

[0618] server

[0619] The server analyzes the received text data. For fraud detection, it uses a pre-prepared list of fraud-related keywords to scan the text data using natural language processing (NLP) techniques. Specifically, it uses the Python nltk library to tokenize the text and verify whether each token is included in the fraud list.

[0620] For example, if someone sends a text saying "I need payment now," we tokenize it and check if keywords like "now," "payment," and "need" are included in our fraud list.

[0621] Fraud score evaluation module (server)

[0622] server

[0623] The fraud score evaluation module on the server calculates scores based on the frequency of keyword occurrence and their combinations. The scoring is done using machine learning models, utilizing Scikit-learn and TensorFlow.

[0624] "Payment required now" will receive a high score because "now" and "payment" are included in the fraud list. If the score exceeds a set threshold, a warning flag will be set.

[0625] Alert notification module (server and terminal)

[0626] server

[0627] When the warning flag is set, the server sends a warning message to the device, which may include a message like "Possible fraud. Please be careful."

[0628] Terminal

[0629] The device will notify the user of the received warning message. There are two notification methods: voice warning and text message. In the case of a voice warning, the device will use TTS (Text-to-Speech) technology to warn the user by voice, saying "Be careful, this may be a scam." In the case of a text message, it will be displayed as a pop-up on the device screen.

[0630] As a specific example of how this works, elderly users can hear a warning during a call saying, "Be careful, this may be a scam." This allows users to sense the risk of being deceived by a scam in advance and take appropriate action.

[0631] Examples of prompts include:

[0632] Describe a telephone fraud prevention system for seniors. The system converts voice data into text in real time, evaluates the likelihood of fraud, and alerts the user. Explain each step in detail, including specific operations and the names of any hardware or software used.

[0633] The flow of the identification process in the first embodiment will be described with reference to FIG.

[0634] System program processing flow

[0635] Step 1: Start the call

[0636] User

[0637] The elderly user answers the phone and starts the conversation. The user presses the talk button and the system starts working.

[0638] Input: Incoming call / call start signal

[0639] Output: System start signal

[0640] Step 2: Launching Audio Analysis

[0641] Terminal

[0642] Upon receiving the call start signal, the terminal automatically starts the voice analysis module.

[0643] How it works: The audio analysis module records the call.

[0644] Input: Call start signal

[0645] Output: Start recording audio data

[0646] Step 3: Convert the audio to text

[0647] Terminal

[0648] The speech analysis module converts the recorded voice data into text data in real time. The speech-to-text conversion is performed using the Google Speech-to-Text API.

[0649] How it works: Captures audio as digital data and sends it to an API to retrieve text.

[0650] Input: Recorded audio data

[0651] Output: Converted text data

[0652] Step 4: Send text data

[0653] Terminal

[0654] The converted text data is sent to the server at regular intervals (less than 1 second). The data is sent using an HTTP POST request.

[0655] Action: Sends data to http: / / server_address / text_data.

[0656] Input: Text data

[0657] Output: Send data to the server

[0658] Step 5: Analyzing the text data

[0659] server

[0660] The server analyzes the received text data and uses a pre-prepared list of fraud-related keywords to detect fraud. It uses the Python nltk library to tokenize the text and verify whether each token is included in the fraud list.

[0661] How it works: It breaks down text data and matches it with keywords from a scam list.

[0662] Input: Received text data

[0663] Output: Keyword occurrences

[0664] Step 6: Scoring

[0665] server

[0666] Based on the occurrence of fraud keywords, the server scores the likelihood of fraud. It calculates the fraud likelihood score using machine learning models such as Scikit-learn and TensorFlow.

[0667] How it works: Evaluates keyword frequency and patterns to calculate a fraud likelihood score.

[0668] Input: Keyword occurrence information

[0669] Output: Fraud likelihood score

[0670] Step 7: Set warning flags

[0671] server

[0672] If the fraud likelihood score exceeds a set threshold, the server sets a warning flag.

[0673] Behavior: Evaluate the score and set a flag if it exceeds a threshold.

[0674] Input: Fraud likelihood score

[0675] Output: warning flag

[0676] Step 8: Sending a warning message

[0677] server

[0678] When the warning flag is set, the server sends a warning message to the terminal, which may include a message such as "Possible fraud. Please be careful."

[0679] Action: Generates a warning message and sends it to the terminal.

[0680] Input: warning flag

[0681] Output: Warning message

[0682] Step 9: Warning Notification

[0683] Terminal

[0684] The device will notify the user of the received warning message. Notification methods include voice alert and text message. Voice alert uses TTS technology, and text message will be displayed as a pop-up on the screen.

[0685] Action: Outputs a voice or text alert based on the message content.

[0686] Input: warning message

[0687] Output: A warning notice to the user

[0688] (Application example 1)

[0689] Next, a description will be given of Application Example 1. In the following description, the data processing device 12 will be referred to as a "server" and the smart glasses 214 will be referred to as a "terminal."

[0690] Telephone frauds against the elderly are increasing year by year, and the amount of damage caused is also steadily increasing. Because the elderly are particularly susceptible to fraudulent schemes, there is a need for fast and effective fraud prevention measures. However, many current fraud prevention systems require users to operate and make decisions themselves, which can result in insufficient effectiveness. Furthermore, there are few systems that can prevent telephone fraud in real time, and the risk of being scammed cannot be completely eliminated. Therefore, there is a need for an effective system that can evaluate and notify users of possible fraud in real time while reducing the operational burden on users.

[0691] The specific processing by the specific processing unit 290 of the data processing device 12 in the application example 1 is realized by the following means.

[0692] In this invention, the server includes means for converting voice data into text data in real time, means for analyzing the text data and identifying keywords and phrases related to fraud, means for assessing the likelihood of fraud based on the analysis results and scoring a risk level, means for notifying a user of a warning if the risk level exceeds a set threshold, means for notifying the user of the warning using a visual display means and a voice notification means of the smart glasses, and means for performing contextual analysis of the text data and detecting patterns specific to fraudulent activities using a generative AI model, thereby significantly reducing the risk of users being involved in telephone fraud and enabling fast and effective fraud prevention in real time.

[0693] "Voice data" refers to data obtained by converting voice signals into digital format, and is the content of a conversation saved in digital format as is.

[0694] "Text data" refers to data in the form of a character string converted from audio data.

[0695] "Keywords" are specific words or phrases that are frequently used in contexts or content that are likely to be fraudulent.

[0696] "Scoring" is an evaluation method that numerically represents the likelihood of fraud.

[0697] A "threshold" is a numerical standard used to assess the likelihood of fraud, and if this value is exceeded, a warning will be issued.

[0698] A "warning" is a notification that informs the user of the risk when there is a high possibility of fraud.

[0699] "Smart glasses" are wearable devices that provide information to users through sight and sound.

[0700] "Visual display means" refers to a mechanism for displaying text and images using a device such as smart glasses.

[0701] A "generative AI model" is a natural language processing model that learns from large amounts of language data using machine learning technology.

[0702] "Contextual analysis" is a technology that understands the content of text data and extracts specific meanings and patterns.

[0703] "Pattern detection" is a technique that analyzes text data to identify speech patterns and phrases characteristic of fraudulent activity.

[0704] The present invention is a telephone fraud prevention system for elderly people that converts voice data into text data in real time, evaluates the possibility of fraud, and notifies the user with a warning. The system has the following configuration.

[0705] 1. Voice analysis module (smart glasses)

[0706] The smart glasses use a built-in microphone to capture the voice during a call in real time, and the voice data is converted into text data in real time using a common cloud speech recognition service (e.g., Google Cloud Speech-to-Text API).

[0707] 2. Text data transmission module (smart glasses)

[0708] The converted text data is sent to a cloud server at regular intervals (less than one second). The cloud server is equipped with a high-performance data processing system, enabling rapid data processing.

[0709] 3. Fraud detection module (cloud server)

[0710] The cloud server analyzes the received text data using natural language processing libraries (e.g., spaCy, NLTK) and scores the frequency of occurrence of fraud-related keywords and phrases. This analysis is powered by a powerful text analysis engine.

[0711] 4. Fraud score evaluation module (cloud server)

[0712] The server then performs contextual analysis using generative AI models (e.g., BERT or GPT-4) to detect patterns characteristic of fraudulent activity, which then calculates a fraud likelihood score and sets a warning flag if the score exceeds a threshold.

[0713] 5. Alert notification module (cloud server and smart glasses)

[0714] If the fraud score exceeds a threshold, the cloud server generates a warning message and sends it to the smart glasses, which then alert the user using visual (HUD display) and audio notification means. For example, a text pop-up on the HUD display reads, "Beware, this may be a scam," along with an audio notification stating the same.

[0715] Specific examples

[0716] Example 1: Fake billing scam call

[0717] 1. Users

[0718] The elderly user answers the phone and begins the conversation.

[0719] 2. Smart Glasses

[0720] The voice analysis module records the call in real time and converts the audio "Payment required now" into text data.

[0721] 3. Smart Glasses

[0722] The converted text data "Payment required now" is sent to the cloud server.

[0723] 4. Cloud Server

[0724] The text data is analyzed to detect keywords such as "now" and "payment" and then scored.

[0725] 5. Cloud Server

[0726] The fraud score assessment module gives a high rating and sets a warning flag because the threshold has been exceeded.

[0727] 6. Cloud Server

[0728] Sending a warning message to the smart glasses.

[0729] 7. Smart Glasses

[0730] The user will be notified of the received warning message via voice, warning them "Beware, this may be a scam," and the same text will pop up on the HUD display.

[0731] 8. Users

[0732] Upon receiving the warning, the user ends the call and consults with a family member.

[0733] Prompt Sentence Examples

[0734] "Please transcribe what was said on this call and assess the likelihood of fraud."

[0735] "Please rate the following text for fraud-related keywords: 'Payment required now'"

[0736] Please analyze the following conversation for possible fraud risk.

[0737] With the above configuration, this system significantly reduces the risk of elderly people falling victim to telephone fraud, enabling fast and effective fraud prevention in real time.

[0738] The flow of the specific processing in the application example 1 will be described with reference to FIG.

[0739] Step 1:

[0740] When a user answers a call and starts a conversation, the voice analysis module in the smart glasses is automatically activated. The voice data during the call is taken as input. The microphone in the smart glasses captures the voice and converts the voice data into text data in real time using a cloud speech recognition service (e.g., Google Cloud Speech-to-Text API). Text data is generated as output, which may include a message such as "Payment required now."

[0741] Step 2:

[0742] The text data transmission module of the smart glasses transmits the converted text data to the cloud server at regular intervals (less than 1 second). The text data generated earlier is used as input. The text data is sent to the cloud server as output.

[0743] Step 3:

[0744] The fraud detection module on the cloud server analyzes the received text data. Specifically, it uses natural language processing libraries (e.g., spaCy, NLTK) to identify fraud-related keywords and phrases. The input is the text data sent to the server. The output is the detection results of fraud-related keywords and phrases (e.g., "now," "pay," etc.).

[0745] Step 4:

[0746] The fraud score evaluation module on the cloud server evaluates the likelihood of fraud based on the analysis results and assigns a risk score. A generative AI model (e.g., BERT or GPT-4) is used to perform contextual analysis and pattern detection on the text data. The input is the output of the fraud detection module. The output is a fraud likelihood score, expressed as a number.

[0747] Step 5:

[0748] If the fraud score exceeds a threshold, the cloud server sets a warning flag. The input is the output of the fraud score evaluation module. The output is the setting state of the warning flag.

[0749] Step 6:

[0750] If the warning flag is set, the cloud server generates a warning message and sends it to the smart glasses. The input is the setting state of the warning flag. The output is the warning message (e.g., "Be careful, this may be a scam").

[0751] Step 7:

[0752] The smart glasses notify the user of the received warning message. The HUD display is used as the visual notification means, and the built-in speaker is used as the audio notification means. The input is the warning message sent from the cloud server. The output is to display the warning to the user both audio and visually.

[0753] Step 8:

[0754] When the user receives the warning message, they end the call and consult with their family to avoid the risk of fraud. The input is the warning message from the smart glasses, and the output is the user's action to take to avoid fraud.

[0755] Furthermore, an emotion engine that estimates the user's emotion may be further combined. That is, the identification processing unit 290 may estimate the user's emotion using the emotion identification model 59, and perform identification processing using the user's emotion.

[0756] The present invention provides a more accurate warning function by combining a system that converts voice data into text data in real time, evaluates the possibility of fraud, and notifies the user of a warning, with an emotion engine that recognizes the user's emotions. This system has the following configuration.

[0757] 1. Voice analysis module (terminal)

[0758] When a call starts, the device activates a voice analysis module and records the call in real time. The recorded voice data is then converted into text data on the spot using voice recognition technology.

[0759] 2. Text data transmission module (terminal)

[0760] The converted text data is sent to the server at regular intervals of less than one second, so that the data is sent in near real time.

[0761] 3. Fraud detection module (server)

[0762] The server analyzes the received text data, specifically matching fraud-related keywords and phrases with a list and scoring them based on their frequency and combinations.

[0763] 4. Fraud Score Evaluation Module (Server)

[0764] The server performs contextual analysis of the text data to detect speech patterns and other characteristics that are particularly characteristic of fraudulent activity. Based on this evaluation, a fraud likelihood score is calculated. If the score exceeds a set threshold, a warning flag is set.

[0765] 5. Emotion engine (server)

[0766] The server uses the voice data to recognize the user's emotional state: the emotion engine analyzes the tone, rate, and other voice characteristics to determine whether the user is experiencing emotions such as surprise, fear, or anxiety.

[0767] 6. Alert Adjustment Module (Server)

[0768] The server adjusts the content and notification method of the warning based on the analysis results of the emotion engine. For example, if the user is feeling strong anxiety or fear, the warning content can be made more detailed and clearer.

[0769] 7. Alert notification module (server and terminal)

[0770] If the warning flag is set, the server sends a warning message to the device, which receives the message and issues a voice or text warning to the user, either saying "Be careful, this may be a scam" in the case of a voice warning or a text message that pops up on the screen.

[0771] Specific examples

[0772] Example 1: Call fraud using emotion engine

[0773] 1. Users

[0774] The elderly user answers the phone and begins the conversation.

[0775] 2. Terminal

[0776] The device's voice analysis module records the call in real time and converts the audio "Payment required now" into text data.

[0777] 3. Terminal

[0778] The text data "Payment is required now" is sent to the server.

[0779] 4. Server

[0780] Analyzes text data to detect and score keywords such as "now" and "payment."

[0781] 5. Server

[0782] The system analyzes the urgent speech patterns characteristic of fraudulent activity and gives a high rating. If this exceeds a threshold, a warning flag is set.

[0783] 6. Server

[0784] At the same time, the emotion engine analyzes the voice data to recognize the user's emotional state, detecting when the user feels an urgent need and is anxious.

[0785] 7. Server

[0786] The warning content is adjusted based on the results of the emotion engine to generate stronger warning messages.

[0787] 8. Server

[0788] Sends tailored warning messages to the terminal.

[0789] 9. Terminal

[0790] The system will notify the user of the received warning message with a voice message, warning them, "Be careful, this may be a scam. Please tell a family member about this call immediately."

[0791] 10. Users

[0792] Users who receive the warning can end the call and consult with their family to prevent any harm.

[0793] In this way, the present invention can significantly reduce the risk of seniors becoming victims of fraud. The system operates in real time and issues appropriate warnings based on the user's emotional state, allowing for a fast and effective response.

[0794] The processing flow will be explained below.

[0795] Step 1:

[0796] The user answers the phone and starts the call. The device detects the start of the call and automatically starts the voice analysis module.

[0797] Step 2:

[0798] The device records the contents of the call in real time and instantly converts the acquired voice data into text data using voice recognition technology.

[0799] Step 3:

[0800] The terminal transmits the converted text data to the server at regular intervals (for example, every second).

[0801] Step 4:

[0802] The server receives the text data sent from the terminal and starts analysis using the text data analysis module.

[0803] Step 5:

[0804] The server detects keywords and phrases in the text data, including words related to fraud (e.g., "now," "payment," "legal action").

[0805] Step 6:

[0806] The server then scores the likelihood of fraud based on the frequency and combination of keywords detected, calculated according to pre-defined criteria.

[0807] Step 7:

[0808] The server performs contextual analysis to detect speech patterns that are particularly characteristic of fraudulent activity (forced language and emphasis on urgency).

[0809] Step 8:

[0810] The server evaluates the scoring results and sets a warning flag if the fraud likelihood score exceeds a set threshold.

[0811] Step 9:

[0812] The server sends the voice data to the emotion engine, which analyzes the tone, rate, and other voice characteristics of the user's voice.

[0813] Step 10:

[0814] The emotion engine on the server determines the user's emotional state, for example, analyzing whether the user is feeling surprise, fear, anxiety, etc.

[0815] Step 11:

[0816] The server adjusts the content of the warning and notification method based on the analysis results of the emotion engine. If the user is feeling strong anxiety or fear, the warning content will be more detailed and clearer.

[0817] Step 12:

[0818] The server sends a tailored warning message to the device, which includes information about the likely fraud and specific steps to take.

[0819] Step 13:

[0820] As soon as the device receives the warning notification from the server, it immediately issues a warning to the user, either by voice or text, saying, "Caution, this may be a scam. Please inform a family member immediately about this call."

[0821] Step 14:

[0822] The user receives a warning and becomes suspicious of the call. They end the call and consult with family or friends for safety. This process significantly reduces the risk of the user becoming a victim of fraud.

[0823] Example 2

[0824] Next, a description will be given of Example 2. In the following description, the data processing device 12 will be referred to as a "server" and the smart glasses 214 will be referred to as a "terminal."

[0825] In today's world, fraud is becoming increasingly sophisticated, making the elderly especially vulnerable to fraud. Previous fraud prevention systems primarily relied on text data analysis and did not provide real-time warnings that took the user's emotional state into account. This made it difficult for users to quickly recognize fraud and take appropriate action. Furthermore, because conventional systems issued warnings uniformly, they often failed to respond appropriately to the sense of crisis and anxiety felt by users. This has led to a need for effective methods to prevent fraud before it happens.

[0826] The identification process by the identification processing unit 290 of the data processing device 12 in Example 2 is realized by the following means. In this invention, the server includes means for converting voice data into text data in real time, means for analyzing the text data and identifying keywords and phrases related to fraud, means for evaluating the possibility of fraud based on the analysis results and scoring the risk level, means for recognizing the user's emotional state from the voice analysis and adjusting the warning content based on the emotion, and means for transmitting the warning content to the terminal and notifying the user. This makes it possible to provide an appropriate warning in real time that takes into account the user's emotional state.

[0827] "Voice data" refers to sound waveforms such as those from a call or voice input that have been converted into electrical signals or digital data.

[0828] "Text data" refers to character string data generated as a result of analyzing and converting voice data.

[0829] "Analysis" is the process of dividing data for a specific purpose to reveal meaning and structure.

[0830] "Fraud-related keywords and phrases" refer to specific words and expressions that are likely to be used in fraud, and are used to predict the possibility of fraud by detecting them.

[0831] "Risk level" is a numerical or score that expresses the likelihood of a scam.

[0832] "Scoring" is the process of quantifying risk based on the results of data analysis.

[0833] A "threshold" is a numerical value or condition that triggers a specific action, and when exceeded, a warning is issued.

[0834] A "warning" is information that notifies the user of a potential danger and urges caution.

[0835] "Emotional state" refers to the user's psychological state analyzed from the tone, rate, and other characteristics of the voice.

[0836] "Adjustment" is the process of changing something to an appropriate form depending on the situation or conditions.

[0837] A "notification" is an action or means of informing a user of specific information.

[0838] The present invention provides a more accurate warning function by combining a system that converts voice data into text data in real time, evaluates the possibility of fraud, and notifies the user of a warning with an emotion engine that recognizes the user's emotions.

[0839] System Configuration

[0840] The system includes the following modules:

[0841] Voice analysis module (terminal)

[0842] The device activates a voice analysis module at the start of a call and records the call in real time. The recorded voice data is converted to text data using the Google Cloud Speech-to-Text API.

[0843] Text data transmission module (terminal)

[0844] The converted text data is sent to the server at regular intervals of less than one second using the WebSocket protocol in real time.

[0845] Fraud detection module (server)

[0846] The server analyzes the received text data using a natural language processing library (e.g., NLTK or SpaCy), matching fraud-related keywords and phrases with a list, and scoring them based on their frequency and combinations.

[0847] Fraud score evaluation module (server)

[0848] The server then performs contextual analysis of the text data to detect speech patterns and other characteristics of fraudulent activity. This information is used to calculate a fraud likelihood score. If this score exceeds a set threshold, a warning flag is set.

[0849] Emotion engine (server)

[0850] The server uses the voice data to recognize the user's emotional state, using IBM Watson Tone Analyzer and Microsoft Azure Emotion API to determine whether the user is experiencing emotions such as surprise, fear, or anxiety.

[0851] Alert Adjustment Module (Server)

[0852] The server adjusts the content and notification method of the warning based on the results of the emotion engine. For example, if the user is feeling strong anxiety or fear, the warning content will be more detailed and clearer.

[0853] Alert notification module (server and terminal)

[0854] The server sends the tailored warning message to the device, which receives it and issues a voice or text warning to the user, such as "Be careful, this may be a scam," or a text message that pops up on the screen.

[0855] Specific examples

[0856] Example 1: Call fraud using emotion engine

[0857] 1. Users

[0858] The elderly user answers the phone and begins the conversation.

[0859] 2. Terminal

[0860] The device's voice analysis module records the call in real time and instantly converts the audio "Payment required now" into text data using Google Cloud Speech-to-Text.

[0861] 3. Terminal

[0862] The converted text data "Payment required now" is sent to the server via the WebSocket protocol at intervals of less than one second.

[0863] 4. Server

[0864] The server uses NLTK to analyze the received text data, identifying fraud-related keywords such as "now" and "payment," and analyzing the frequency and context of these keywords.

[0865] 5. Server

[0866] It detects urgent speech patterns characteristic of fraudulent activity and generates a high fraud likelihood score. If this score exceeds a set threshold, it sets a warning flag.

[0867] 6. Server

[0868] The server uses an emotion engine (IBM Watson Tone Analyzer) to analyze the user's emotional state from the voice data. The analysis results indicate that the user feels urgency and is anxious.

[0869] 7. Server

[0870] The content of the warning message is adjusted based on the results of the emotion engine, for example, generating a detailed warning such as "This call may be fraudulent. Please report this call to a family member immediately."

[0871] 8. Server

[0872] Sends tailored warning messages to the terminal.

[0873] 9. Terminal

[0874] The device will then notify the user of the received warning message via voice: "Caution, this may be a scam. Please inform a family member immediately about this call."

[0875] 10. Users

[0876] The user will receive a warning, end the call, and immediately contact their family to prevent any harm.

[0877] In this way, the present invention can significantly reduce the risk of seniors becoming victims of fraud. The system operates in real time and issues appropriate warnings based on the user's emotional state, allowing for a fast and effective response.

[0878] Prompt Sentence Examples

[0879] "Design a system to detect emotions from voice data, assess the likelihood of fraud, and provide a warning. Use natural language processing and emotion recognition technologies for data analysis. Please also specify the specific hardware and software, data flow, and user interface (UI)."

[0880] The flow of the identification process in the second embodiment will be described with reference to FIG.

[0881] Step 1:

[0882] Voice recording and text conversion (device)

[0883] Input: A user receives a call and starts a conversation.

[0884] How it works: The device detects that a call has been initiated and activates the voice analysis module.

[0885] Data processing: Convert the audio data into text data in real time using the Google Cloud Speech-to-Text API.

[0886] Output: The converted text data is generated.

[0887] Step 2:

[0888] Sending text data (terminal)

[0889] Input: The text data generated in step 1.

[0890] Operation: The device sends the converted text data to the server at intervals of less than one second.

[0891] Data calculation: Send data in real time using the WebSocket protocol.

[0892] Output: The server receives text data in real time.

[0893] Step 3:

[0894] Text data analysis (server)

[0895] Input: The text data received in step 2.

[0896] How it works: The server parses the received text data using a natural language processing library (NLTK or SpaCy).

[0897] Data calculations: Fraud-related keywords and phrases are matched against the list and analyzed based on their frequency and combinations.

[0898] Output: A numerical score is generated representing the likelihood of fraud.

[0899] Step 4:

[0900] Fraud score evaluation (server)

[0901] Input: Analysis results (scores) generated in step 3.

[0902] How it works: The server also performs contextual analysis of the text data to detect speech patterns and patterns characteristic of fraudulent activity.

[0903] Data calculation: Calculate a fraud likelihood score and set a warning flag if this score exceeds a set threshold.

[0904] Output: Fraud probability score and warning flag status.

[0905] Step 5:

[0906] User emotion analysis (server)

[0907] Input: The audio data recorded in step 1.

[0908] How it works: The server uses IBM Watson Tone Analyzer and Microsoft Azure Emotion API to analyze the user's emotional state from voice data.

[0909] Data calculations: Analyze the tone, rate, and other voice characteristics of the voice to determine whether the user is surprised, scared, or anxious.

[0910] Output: Generates an analysis of the user's emotional state.

[0911] Step 6:

[0912] Adjustment of warning content (server)

[0913] Input: Warning flags from step 4 and sentiment analysis results from step 5.

[0914] How it works: The server tailors the warning based on the user's emotional state. For example, if the user is feeling strong anxiety or fear, the warning will be more detailed and clearer.

[0915] Data calculation: Optimize the content and format of warning messages based on the results of sentiment analysis.

[0916] Output: A tailored warning message is generated.

[0917] Step 7:

[0918] Notification of warning messages (server and terminal)

[0919] Input: The warning message adjusted in step 6.

[0920] Action: The server sends a tailored alert message to the device.

[0921] Data calculation: The device notifies the user of the received warning message by voice or text.

[0922] Output: A voice or text alert is given to the user.

[0923] These steps enable the system to detect potential fraud in real time and issue warnings that take into account the user's emotional state.

[0924] (Application example 2)

[0925] Next, a description will be given of Application Example 2. In the following description, the data processing device 12 will be referred to as a "server" and the smart glasses 214 will be referred to as a "terminal."

[0926] In recent years, telephone fraud has been on the rise, with elderly people becoming victims in particular. Conventional fraud prevention measures only issue a warning to users when fraud is suspected. Therefore, even when users receive a warning, they are unable to properly determine how to respond, resulting in a high likelihood of fraud victimization. Furthermore, because warnings are issued uniformly with the same intensity without taking into account the user's emotional state, the effectiveness of the warnings is limited. There is a need to improve this situation and provide more effective fraud prevention measures.

[0927] The specific processing by the specific processing unit 290 of the data processing device 12 in the application example 2 is realized by the following means.

[0928] In this invention, the server includes means for converting voice data into text data in real time, means for analyzing the text data and identifying keywords and phrases related to fraud, means for assessing the likelihood of fraud and scoring the risk based on the analysis results, means for issuing a warning to the user when the risk exceeds a set threshold, means for analyzing the user's emotional state, means for adjusting the content of the warning based on the emotion analysis results, and means for issuing the adjusted warning to the user. This makes it possible to accurately assess the likelihood of fraud in real time based on the voice data and provide an appropriate warning that takes the user's emotional state into consideration.

[0929] "Voice data" means recorded information of a user's voice uttered over the telephone or other voice communication.

[0930] "Real-time" refers to near-instant processing with little actual time or delay.

[0931] "Text data" is voice data converted into character information.

[0932] Analysis is the process of examining data or information in detail to identify important elements or patterns within it.

[0933] "Fraud-related keywords and phrases" are specific words and expressions commonly used in fraudulent activities.

[0934] "Scoring" is the process of assigning points according to specific criteria based on the analysis results.

[0935] A "threshold" is a boundary value for determining a certain condition or standard.

[0936] A "warning" is a notification that notifies the user of an impending danger.

[0937] An "emotional state" is the psychological state that a user is feeling at a particular moment.

[0938] "Notification" is the act of conveying information or a message to a target person.

[0939] "Adjustment" refers to changing something to an optimal state for specific conditions or environments.

[0940] This invention is a system for protecting users from telephone fraud, and provides highly accurate warning functions by combining analysis of voice data and emotion analysis.

[0941] System Overview

[0942] The system mainly consists of three main components.

[0943] 1. Terminal

[0944] 2. Server

[0945] 3. Users

[0946] Hardware and software used

[0947] Hardware:

[0948] Smartphone: Used as a device, requires a microphone and internet connection.

[0949] software:

[0950] Python 3.x: The main programming language.

[0951] speech_recognition library: Uses Google's speech recognition API to convert voice data into text data.

[0952] The requests library: handles HTTP requests to the server.

[0953] Processing Details

[0954] Terminal handling

[0955] When a user initiates a phone call, the device records the voice data in real time and converts it into text data using the speech_recognition library, which is then sent to the server at regular intervals.

[0956] Server Processing

[0957] The server analyzes the audio data and assesses the likelihood of fraud using the following steps:

[0958] 1. Text Analysis: Analyzes incoming text data to identify keywords and phrases related to fraud.

[0959] 2. Scoring: Based on the analysis results, the possibility of fraud is assessed and a risk score is generated. If the score exceeds a set threshold, a warning is generated.

[0960] 3. Emotion analysis: Analyze the user's emotional state from the voice data, determining whether the user is feeling surprised or anxious.

[0961] 4. Adjustment of warning content: Based on the results of sentiment analysis, the warning content is adjusted appropriately.

[0962] Warning Notification

[0963] A tailored warning message is sent to the user's device. For example, if the warning is delivered via voice, the content and intensity of the warning can be adjusted according to the user's emotional state. This is expected to encourage the user to take appropriate action.

[0964] Specific examples

[0965] For example, consider the case of an elderly person receiving a suspected fraudulent phone call. If the user hears phrases such as "Payment is required now," the system immediately transcribes the speech into text and analyzes it. If it detects patterns indicative of fraudulent activity, the system takes into account the user's emotional state and generates a strong warning message. The warning could include specific advice such as "Be careful, this may be a scam. Please report this call to a family member immediately."

[0966] Example prompts for generative AI models

[0967] Your system should take voice data from a user's phone call as input and convert it into text data in real time. It should then assess the likelihood of fraud (by scoring it based on the frequency of fraud-related keywords and phrases) and generate a warning message if necessary. It should also analyze the user's emotional state and adjust the strength and content of the warning. The emotion engine uses the tone, rate, and other voice characteristics to determine whether the user is feeling anxious or scared. It should generate a warning message based on the results of the text and emotion analysis and notify the user.

[0968] The flow of the specific processing in the application example 2 will be described with reference to FIG.

[0969] Step 1:

[0970] A user initiates a phone call.

[0971] Input: User's call audio

[0972] Output: Call audio data

[0973] Specific operation: A user initiates a call and voice communication begins over the phone.

[0974] Step 2:

[0975] The device records the call audio in real time.

[0976] Input: Call audio data

[0977] Output: Recorded audio file

[0978] Specific operation: The device's microphone captures the user's call voice and records the voice data.

[0979] Step 3:

[0980] The device converts recorded voice data into text data in real time.

[0981] Input: Recorded audio file

[0982] Output: Text data

[0983] Specific operation: Using the speech_recognition library installed on the device, the recorded voice data is converted into text data.

[0984] Step 4:

[0985] The terminal transmits the converted text data to the server at regular intervals.

[0986] Input: Text data

[0987] Output: Send text data to the server

[0988] Specific operation: Sends text data to the server using an HTTP request, for example, using the requests library.

[0989] Step 5:

[0990] The server analyzes the received text data to identify keywords and phrases related to fraud.

[0991] Input: Text data sent to the server

[0992] Output: Analysis results (list of fraud-related keywords and phrases)

[0993] How it works: An analytics module on the server scans the text data to identify keywords and phrases related to fraud, matching them against a predefined list of fraud keywords.

[0994] Step 6:

[0995] The server evaluates the likelihood of fraud and assigns a score based on the results of analyzing the text data.

[0996] Input: Analysis results

[0997] Output: Risk score

[0998] Specific operation: The server's scoring module evaluates the likelihood of fraud based on the frequency of fraud keywords and the context of the text, and calculates a risk score.

[0999] Step 7:

[1000] The server generates a warning message if the risk score exceeds a threshold.

[1001] Input: Risk score

[1002] Output: Warning message

[1003] Specific behavior: If the score exceeds a preset threshold, an algorithm is executed that generates a warning message.

[1004] Step 8:

[1005] The server analyzes the user's emotional state from the voice data.

[1006] Input: Audio data

[1007] Output: Emotion analysis results

[1008] Specific behavior: The emotion engine analyzes the tone, rate, and other voice characteristics of the voice to assess the user's emotional state.

[1009] Step 9:

[1010] The server adjusts the warning content based on the results of sentiment analysis.

[1011] Input: Sentiment analysis results, warning message

[1012] Output: Adjusted warning message

[1013] Specific operation: Using the results of the emotion engine, an algorithm is run that adjusts the content and intensity of warning messages.

[1014] Step 10:

[1015] The server sends the adjusted warning message to the terminal, and the terminal notifies the user of the warning.

[1016] Input: Adjusted warning message

[1017] Output: A warning notice to the user

[1018] Specific operation: The server sends a tailored warning message to the terminal, which then notifies the user of the message. The warning is notified as voice or text.

[1019] Each step is performed sequentially, protecting users from the risk of fraud in real time.

[1020] The specific processing unit 290 transmits the result of the specific processing to the smart glasses 214. In the smart glasses 214, the control unit 46A causes the speaker 240 to output the result of the specific processing. The microphone 238 acquires audio indicating a user input regarding the result of the specific processing. The control unit 46A transmits audio data indicating the user input acquired by the microphone 238 to the data processing device 12. In the data processing device 12, the specific processing unit 290 acquires the audio data.

[1021] The data generation model 58 is a so-called generative AI (Artificial Intelligence). An example of the data generation model 58 is ChatGPT (Internet Search<URL: https: / / openai.com / blog / chatgpt> ), Gemini (Internet search <url: https: gemini.google.com ?hl="ja">) and other generation AIs. The data generation model 58 is obtained by performing deep learning on a neural network. A prompt including an instruction is input to the data generation model 58, and inference data such as voice data indicating voice, text data indicating text, and image data indicating an image is also input. The data generation model 58 performs inference on the input inference data in accordance with the instruction indicated by the prompt, and outputs the inference result in a data format such as voice data and text data. Here, inference refers to, for example, analysis, classification, prediction, and / or summarization.

[1022] In the above embodiment, an example in which the specific processing is performed by the data processing device 12 has been given, but the technology of the present disclosure is not limited to this, and the specific processing may be performed by the smart glasses 214.

[1023] [Third embodiment]

[1024] FIG. 5 shows an example of the configuration of a data processing system 310 according to the third embodiment.

[1025] 5, the data processing system 310 includes the data processing device 12 and a headset terminal 314. An example of the data processing device 12 is a server.

[1026] The data processing device 12 includes a computer 22, a database 24, and a communication I / F 26. The computer 22 is an example of a "computer" according to the technology of the present disclosure. The computer 22 includes a processor 28, a RAM 30, and a storage 32. The processor 28, the RAM 30, and the storage 32 are connected to a bus 34. The database 24 and the communication I / F 26 are also connected to the bus 34. The communication I / F 26 is connected to a network 54. Examples of the network 54 include a WAN (Wide Area Network) and / or a LAN (Local Area Network).

[1027] The headset type terminal 314 includes a computer 36, a microphone 238, a speaker 240, a camera 42, a communication I / F 44, and a display 343. The computer 36 includes a processor 46, a RAM 48, and a storage 50. The processor 46, the RAM 48, and the storage 50 are connected to a bus 52. The microphone 238, the speaker 240, the camera 42, and the display 343 are also connected to the bus 52.

[1028] The microphone 238 receives instructions and the like from the user 20 by receiving voice uttered by the user 20. The microphone 238 captures the voice uttered by the user 20, converts the captured voice into audio data, and outputs it to the processor 46. The speaker 240 outputs audio in accordance with instructions from the processor 46.

[1029] Camera 42 is a small digital camera equipped with an optical system including a lens, aperture, and shutter, and an imaging element such as a CMOS (Complementary Metal-Oxide-Semiconductor) image sensor or a CCD (Charge Coupled Device) image sensor, and captures images of the surroundings of user 20 (for example, an imaging range defined by an angle of view equivalent to the field of vision of a typical healthy person).

[1030] The communication I / F 44 is connected to a network 54. The communication I / Fs 44 and 26 are responsible for the exchange of various information between the processor 46 and the processor 28 via the network 54. The exchange of various information between the processor 46 and the processor 28 using the communication I / Fs 44 and 26 is carried out in a secure state.

[1031] Fig. 6 shows an example of the main functions of the data processing device 12 and the headset type terminal 314. As shown in Fig. 6, in the data processing device 12, a specific process is performed by the processor 28. A specific process program 56 is stored in the storage 32.

[1032] The specific processing program 56 is an example of a "program" according to the technology of the present disclosure. The processor 28 reads the specific processing program 56 from the storage 32 and executes the read specific processing program 56 on the RAM 30. The specific processing is realized by the processor 28 operating as a specific processing unit 290 in accordance with the specific processing program 56 executed on the RAM 30.

[1033] The storage 32 stores a data generation model 58 and an emotion identification model 59. The data generation model 58 and the emotion identification model 59 are used by the identification processing unit 290.

[1034] In the headset type terminal 314, a reception output process is performed by the processor 46. A reception output program 60 is stored in the storage 50. The processor 46 reads the reception output program 60 from the storage 50 and executes the read reception output program 60 on the RAM 48. The reception output process is realized by the processor 46 operating as the control unit 46A in accordance with the reception output program 60 executed on the RAM 48.

[1035] Next, a description will be given of the identification process performed by the identification processing unit 290 of the data processing device 12. In the following description, the data processing device 12 will be referred to as the "server" and the headset type terminal 314 will be referred to as the "terminal."

[1036] The present invention is a telephone fraud prevention system for elderly people that converts voice data into text data in real time, evaluates the possibility of fraud, and notifies the user with a warning. The system has the following configuration.

[1037] 1. Voice analysis module (terminal)

[1038] When a call starts, the device activates a voice analysis module and records the call in real time. The recorded voice data is then converted into text data on the spot using voice recognition technology.

[1039] 2. Text data transmission module (terminal)

[1040] The converted text data is sent to the server at regular intervals of less than one second, so that the data is sent in near real time.

[1041] 3. Fraud detection module (server)

[1042] The server analyzes the received text data, specifically matching fraud-related keywords and phrases with a list and scoring them based on their frequency and combinations.

[1043] 4. Fraud Score Evaluation Module (Server)

[1044] The server performs contextual analysis of the text data to detect speech patterns and other characteristics that are particularly characteristic of fraudulent activity. Based on this evaluation, a fraud likelihood score is calculated. If the score exceeds a set threshold, a warning flag is set.

[1045] 5. Alert notification module (server and terminal)

[1046] When the warning flag is set, the server sends a warning message to the device, which receives the message and issues a voice or text warning to the user, saying "Be careful, this may be a scam," or a text message that pops up on the screen.

[1047] Specific examples

[1048] Example 1: Fake billing scam call

[1049] 1. Users

[1050] The elderly user answers the phone and begins the conversation.

[1051] 2. Terminal

[1052] The device's voice analysis module records the call in real time and converts the audio "Payment required now" into text data.

[1053] 3. Terminal

[1054] The text data "Payment is required now" is sent to the server.

[1055] 4. Server

[1056] Analyzes text data to detect and score keywords such as "now" and "payment."

[1057] 5. Server

[1058] The fraud score assessment module gives a high score and sets a warning flag because the threshold has been exceeded.

[1059] 6. Server

[1060] Sends a warning message to the terminal.

[1061] 7. Terminal

[1062] The system will notify the user of the received warning message by voice, warning them with "Be careful, this may be a scam."

[1063] 8. Users

[1064] Users who receive the warning should end the call and consult with their family to prevent further harm.

[1065] In this way, the risk of seniors being scammed can be significantly reduced by this invention. The system operates in real time and issues immediate warnings, allowing for a rapid response.

[1066] The processing flow will be explained below.

[1067] Step 1:

[1068] The user answers the phone and starts the call. The device detects the start of the call and automatically starts the voice analysis module.

[1069] Step 2:

[1070] The device records the contents of the call in real time and instantly converts the acquired voice data into text data using voice recognition technology.

[1071] Step 3:

[1072] The terminal transmits the converted text data to the server at regular intervals (for example, every second).

[1073] Step 4:

[1074] The server receives the text data sent from the terminal and starts analysis using the text data analysis module.

[1075] Step 5:

[1076] The server detects keywords and phrases in the text data, including words related to fraud (e.g., "now," "payment," "legal action").

[1077] Step 6:

[1078] The server then scores the likelihood of fraud based on the frequency and combination of keywords detected, calculated according to pre-defined criteria.

[1079] Step 7:

[1080] The server performs contextual analysis to detect speech patterns that are particularly characteristic of fraudulent activity (forced language and emphasis on urgency).

[1081] Step 8:

[1082] The server evaluates the scoring results and sets a warning flag if the fraud likelihood score exceeds a set threshold.

[1083] Step 9:

[1084] The server notifies the terminal that a warning flag has been set, including a message indicating a high probability of fraud.

[1085] Step 10:

[1086] As soon as the device receives the warning notification from the server, it immediately issues a warning to the user, either by voice or text, saying, "Be careful, this may be a scam."

[1087] Step 11:

[1088] The user receives a warning and becomes suspicious of the call. The user ends the call and consults with family or friends for safety.

[1089] This series of steps enables the system to detect fraudulent activity with high accuracy and reduce the risk of users becoming victims.

[1090] Example 1

[1091] Next, a description will be given of Example 1. In the following description, the data processing device 12 will be referred to as a "server" and the headset type terminal 314 will be referred to as a "terminal."

[1092] Currently, the number of elderly victims of telephone fraud is increasing, and fraudulent acts via telephone in particular has become a social problem. Measures to prevent this are needed. However, current methods are difficult to respond to in real time, and fraud prevention is not sufficient. Therefore, there is a need for a system that can reduce the risk of elderly people becoming victims of telephone fraud and issue warnings quickly and effectively.

[1093] The specific processing by the specific processing unit 290 of the data processing device 12 in the first embodiment is realized by the following means.

[1094] In this invention, the server includes means for converting voice data into text data in real time, means for transmitting the text data to the server at regular intervals, means for analyzing the text data and identifying fraud-related keywords and phrases, means for evaluating the possibility of fraud and scoring a risk level based on the analysis results, means for detecting fraud patterns based on the fraud-related keywords and phrases and calculating a fraud likelihood score, and means for issuing a voice or text warning to the user when the risk level exceeds a set threshold. This significantly reduces the risk of elderly people becoming victims of telephone fraud and enables prompt warnings to be issued in real time.

[1095] "Voice data" refers to data that records or transmits voice information such as telephone calls or conversations in digital format.

[1096] "Text data" is data in a format in which voice data is converted into a string of characters, and is information expressed in a format that can be read by humans.

[1097] "Real-time" means that the system processes voice input immediately and outputs the results without delay.

[1098] "Analysis" is the process of examining the information contained in data in detail and extracting and evaluating that information based on specific criteria and rules.

[1099] "Fraud-related keywords and phrases" refer to expressions and words characteristic of fraudulent activities, and are specific words and phrases that are judged to have a high possibility of being fraudulent.

[1100] "Scoring" refers to the process of quantifying and evaluating the likelihood of fraud based on specific criteria or algorithms based on the analysis results.

[1101] A "threshold" is a numerical value or condition that determines when a system will take a particular action or perform a particular algorithm.

[1102] A "warning" is a message or notification that warns the user and is issued when there is a possibility of fraud.

[1103] "Interval" refers to the timing or cycle at which data is transmitted, and indicates the time difference between successive data transmissions.

[1104] "Patterns" refer to speech patterns and contextual structures that are characteristic of fraudulent behavior, and are standards or models based on which to judge the possibility of fraud.

[1105] "Calculation" is the process by which a system generates a number or evaluation according to a specified algorithm or procedure.

[1106] "Voice or text notification" means that the warning to the user is given in the form of voice output or text display, and the system issues the warning in an appropriate manner.

[1107] This invention is a telephone fraud prevention system aimed at the elderly, which analyzes the voice of a call in real time and issues a warning when there is a high possibility of fraud. The specific configuration is as follows.

[1108] Voice analysis module (terminal)

[1109] User

[1110] When the elderly user receives a call, the call begins. The system starts working when the user presses the call button.

[1111] Terminal

[1112] When a call is received, the device automatically activates a voice analysis module, which records the call in real time. The recorded voice data is then converted into text data on the fly using voice recognition technology. Specifically, real-time voice-to-text conversion is performed using the Google Speech-to-Text API.

[1113] Text data transmission module (terminal)

[1114] Terminal

[1115] The converted text data is sent to the server at regular intervals (less than 1 second). The HTTP protocol is used to send the data, for example, as a POST request. Specifically, the text data is sent to http: / / server_address / text_data.

[1116] Fraud detection module (server)

[1117] server

[1118] The server analyzes the received text data. For fraud detection, it uses a pre-prepared list of fraud-related keywords to scan the text data using natural language processing (NLP) techniques. Specifically, it uses the Python nltk library to tokenize the text and verify whether each token is included in the fraud list.

[1119] For example, if someone sends a text saying "I need payment now," we tokenize it and check if keywords like "now," "payment," and "need" are included in our fraud list.

[1120] Fraud score evaluation module (server)

[1121] server

[1122] The fraud score evaluation module on the server calculates scores based on the frequency of keyword occurrence and their combinations. The scoring is done using machine learning models, utilizing Scikit-learn and TensorFlow.

[1123] "Payment required now" will receive a high score because "now" and "payment" are included in the fraud list. If the score exceeds a set threshold, a warning flag will be set.

[1124] Alert notification module (server and terminal)

[1125] server

[1126] When the warning flag is set, the server sends a warning message to the device, which may include a message like "Possible fraud. Please be careful."

[1127] Terminal

[1128] The device will notify the user of the received warning message. There are two notification methods: voice warning and text message. In the case of a voice warning, the device will use TTS (Text-to-Speech) technology to warn the user by voice, saying "Be careful, this may be a scam." In the case of a text message, it will be displayed as a pop-up on the device screen.

[1129] As a specific example of how this works, elderly users can hear a warning during a call saying, "Be careful, this may be a scam." This allows users to sense the risk of being deceived by a scam in advance and take appropriate action.

[1130] Examples of prompts include:

[1131] Describe a telephone fraud prevention system for seniors. The system converts voice data into text in real time, evaluates the likelihood of fraud, and alerts the user. Explain each step in detail, including specific operations and the names of any hardware or software used.

[1132] The flow of the identification process in the first embodiment will be described with reference to FIG.

[1133] System program processing flow

[1134] Step 1: Start the call

[1135] User

[1136] The elderly user answers the phone and starts the conversation. The user presses the talk button and the system starts working.

[1137] Input: Incoming call / call start signal

[1138] Output: System start signal

[1139] Step 2: Launching Audio Analysis

[1140] Terminal

[1141] Upon receiving the call start signal, the terminal automatically starts the voice analysis module.

[1142] How it works: The audio analysis module records the call.

[1143] Input: Call start signal

[1144] Output: Start recording audio data

[1145] Step 3: Convert the audio to text

[1146] Terminal

[1147] The speech analysis module converts the recorded voice data into text data in real time. The speech-to-text conversion is performed using the Google Speech-to-Text API.

[1148] How it works: Captures audio as digital data and sends it to an API to retrieve text.

[1149] Input: Recorded audio data

[1150] Output: Converted text data

[1151] Step 4: Send text data

[1152] Terminal

[1153] The converted text data is sent to the server at regular intervals (less than 1 second). The data is sent using an HTTP POST request.

[1154] Action: Sends data to http: / / server_address / text_data.

[1155] Input: Text data

[1156] Output: Send data to the server

[1157] Step 5: Analyzing the text data

[1158] server

[1159] The server analyzes the received text data and uses a pre-prepared list of fraud-related keywords to detect fraud. It uses the Python nltk library to tokenize the text and verify whether each token is included in the fraud list.

[1160] How it works: It breaks down text data and matches it with keywords from a scam list.

[1161] Input: Received text data

[1162] Output: Keyword occurrences

[1163] Step 6: Scoring

[1164] server

[1165] Based on the occurrence of fraud keywords, the server scores the likelihood of fraud. It calculates the fraud likelihood score using machine learning models such as Scikit-learn and TensorFlow.

[1166] How it works: Evaluates keyword frequency and patterns to calculate a fraud likelihood score.

[1167] Input: Keyword occurrence information

[1168] Output: Fraud likelihood score

[1169] Step 7: Set warning flags

[1170] server

[1171] If the fraud likelihood score exceeds a set threshold, the server sets a warning flag.

[1172] Behavior: Evaluate the score and set a flag if it exceeds a threshold.

[1173] Input: Fraud likelihood score

[1174] Output: warning flag

[1175] Step 8: Sending a warning message

[1176] server

[1177] When the warning flag is set, the server sends a warning message to the terminal, which may include a message such as "Possible fraud. Please be careful."

[1178] Action: Generates a warning message and sends it to the terminal.

[1179] Input: warning flag

[1180] Output: Warning message

[1181] Step 9: Warning Notification

[1182] Terminal

[1183] The device will notify the user of the received warning message. Notification methods include voice alert and text message. Voice alert uses TTS technology, and text message will be displayed as a pop-up on the screen.

[1184] Action: Outputs a voice or text alert based on the message content.

[1185] Input: warning message

[1186] Output: A warning notice to the user

[1187] (Application example 1)

[1188] Next, a description will be given of Application Example 1. In the following description, the data processing device 12 will be referred to as a "server" and the headset type terminal 314 will be referred to as a "terminal."

[1189] Telephone frauds against the elderly are increasing year by year, and the amount of damage caused is also steadily increasing. Because the elderly are particularly susceptible to fraudulent schemes, there is a need for fast and effective fraud prevention measures. However, many current fraud prevention systems require users to operate and make decisions themselves, which can result in insufficient effectiveness. Furthermore, there are few systems that can prevent telephone fraud in real time, and the risk of being scammed cannot be completely eliminated. Therefore, there is a need for an effective system that can evaluate and notify users of possible fraud in real time while reducing the operational burden on users.

[1190] The specific processing by the specific processing unit 290 of the data processing device 12 in the application example 1 is realized by the following means.

[1191] In this invention, the server includes means for converting voice data into text data in real time, means for analyzing the text data and identifying keywords and phrases related to fraud, means for assessing the likelihood of fraud based on the analysis results and scoring a risk level, means for notifying a user of a warning if the risk level exceeds a set threshold, means for notifying the user of the warning using a visual display means and a voice notification means of the smart glasses, and means for performing contextual analysis of the text data and detecting patterns specific to fraudulent activities using a generative AI model, thereby significantly reducing the risk of users being involved in telephone fraud and enabling fast and effective fraud prevention in real time.

[1192] "Voice data" refers to data obtained by converting voice signals into digital format, and is the content of a conversation saved in digital format as is.

[1193] "Text data" refers to data in the form of a character string converted from audio data.

[1194] "Keywords" are specific words or phrases that are frequently used in contexts or content that are likely to be fraudulent.

[1195] "Scoring" is an evaluation method that numerically represents the likelihood of fraud.

[1196] A "threshold" is a numerical standard used to assess the likelihood of fraud, and if this value is exceeded, a warning will be issued.

[1197] A "warning" is a notification that informs the user of the risk when there is a high possibility of fraud.

[1198] "Smart glasses" are wearable devices that provide information to users through sight and sound.

[1199] "Visual display means" refers to a mechanism for displaying text and images using a device such as smart glasses.

[1200] A "generative AI model" is a natural language processing model that learns from large amounts of language data using machine learning technology.

[1201] "Contextual analysis" is a technology that understands the content of text data and extracts specific meanings and patterns.

[1202] "Pattern detection" is a technique that analyzes text data to identify speech patterns and phrases characteristic of fraudulent activity.

[1203] The present invention is a telephone fraud prevention system for elderly people that converts voice data into text data in real time, evaluates the possibility of fraud, and notifies the user with a warning. The system has the following configuration.

[1204] 1. Voice analysis module (smart glasses)

[1205] The smart glasses use a built-in microphone to capture the voice during a call in real time, and the voice data is converted into text data in real time using a common cloud speech recognition service (e.g., Google Cloud Speech-to-Text API).

[1206] 2. Text data transmission module (smart glasses)

[1207] The converted text data is sent to a cloud server at regular intervals (less than one second). The cloud server is equipped with a high-performance data processing system, enabling rapid data processing.

[1208] 3. Fraud detection module (cloud server)

[1209] The cloud server analyzes the received text data using natural language processing libraries (e.g., spaCy, NLTK) and scores the frequency of occurrence of fraud-related keywords and phrases. This analysis is powered by a powerful text analysis engine.

[1210] 4. Fraud score evaluation module (cloud server)

[1211] The server then performs contextual analysis using generative AI models (e.g., BERT or GPT-4) to detect patterns characteristic of fraudulent activity, which then calculates a fraud likelihood score and sets a warning flag if the score exceeds a threshold.

[1212] 5. Alert notification module (cloud server and smart glasses)

[1213] If the fraud score exceeds a threshold, the cloud server generates a warning message and sends it to the smart glasses, which then alert the user using visual (HUD display) and audio notification means. For example, a text pop-up on the HUD display reads, "Beware, this may be a scam," along with an audio notification stating the same.

[1214] Specific examples

[1215] Example 1: Fake billing scam call

[1216] 1. Users

[1217] The elderly user answers the phone and begins the conversation.

[1218] 2. Smart Glasses

[1219] The voice analysis module records the call in real time and converts the audio "Payment required now" into text data.

[1220] 3. Smart Glasses

[1221] The converted text data "Payment required now" is sent to the cloud server.

[1222] 4. Cloud Server

[1223] The text data is analyzed to detect keywords such as "now" and "payment" and then scored.

[1224] 5. Cloud Server

[1225] The fraud score assessment module gives a high rating and sets a warning flag because the threshold has been exceeded.

[1226] 6. Cloud Server

[1227] Sending a warning message to the smart glasses.

[1228] 7. Smart Glasses

[1229] The user will be notified of the received warning message via voice, warning them "Beware, this may be a scam," and the same text will pop up on the HUD display.

[1230] 8. Users

[1231] Upon receiving the warning, the user ends the call and consults with a family member.

[1232] Prompt Sentence Examples

[1233] "Please transcribe what was said on this call and assess the likelihood of fraud."

[1234] "Please rate the following text for fraud-related keywords: 'Payment required now'"

[1235] Please analyze the following conversation for possible fraud risk.

[1236] With the above configuration, this system significantly reduces the risk of elderly people falling victim to telephone fraud, enabling fast and effective fraud prevention in real time.

[1237] The flow of the specific processing in the application example 1 will be described with reference to FIG.

[1238] Step 1:

[1239] When a user answers a call and starts a conversation, the voice analysis module in the smart glasses is automatically activated. The voice data during the call is taken as input. The microphone in the smart glasses captures the voice and converts the voice data into text data in real time using a cloud speech recognition service (e.g., Google Cloud Speech-to-Text API). Text data is generated as output, which may include a message such as "Payment required now."

[1240] Step 2:

[1241] The text data transmission module of the smart glasses transmits the converted text data to the cloud server at regular intervals (less than 1 second). The text data generated earlier is used as input. The text data is sent to the cloud server as output.

[1242] Step 3:

[1243] The fraud detection module on the cloud server analyzes the received text data. Specifically, it uses natural language processing libraries (e.g., spaCy, NLTK) to identify fraud-related keywords and phrases. The input is the text data sent to the server. The output is the detection results of fraud-related keywords and phrases (e.g., "now," "pay," etc.).

[1244] Step 4:

[1245] The fraud score evaluation module on the cloud server evaluates the likelihood of fraud based on the analysis results and assigns a risk score. A generative AI model (e.g., BERT or GPT-4) is used to perform contextual analysis and pattern detection on the text data. The input is the output of the fraud detection module. The output is a fraud likelihood score, expressed as a number.

[1246] Step 5:

[1247] If the fraud score exceeds a threshold, the cloud server sets a warning flag. The input is the output of the fraud score evaluation module. The output is the setting state of the warning flag.

[1248] Step 6:

[1249] If the warning flag is set, the cloud server generates a warning message and sends it to the smart glasses. The input is the setting state of the warning flag. The output is the warning message (e.g., "Be careful, this may be a scam").

[1250] Step 7:

[1251] The smart glasses notify the user of the received warning message. The HUD display is used as the visual notification means, and the built-in speaker is used as the audio notification means. The input is the warning message sent from the cloud server. The output is to display the warning to the user both audio and visually.

[1252] Step 8:

[1253] When the user receives the warning message, they end the call and consult with their family to avoid the risk of fraud. The input is the warning message from the smart glasses, and the output is the user's action to take to avoid fraud.

[1254] Furthermore, an emotion engine that estimates the user's emotion may be further combined. That is, the identification processing unit 290 may estimate the user's emotion using the emotion identification model 59, and perform identification processing using the user's emotion.

[1255] The present invention provides a more accurate warning function by combining a system that converts voice data into text data in real time, evaluates the possibility of fraud, and notifies the user of a warning, with an emotion engine that recognizes the user's emotions. This system has the following configuration.

[1256] 1. Voice analysis module (terminal)

[1257] When a call starts, the device activates a voice analysis module and records the call in real time. The recorded voice data is then converted into text data on the spot using voice recognition technology.

[1258] 2. Text data transmission module (terminal)

[1259] The converted text data is sent to the server at regular intervals of less than one second, so that the data is sent in near real time.

[1260] 3. Fraud detection module (server)

[1261] The server analyzes the received text data, specifically matching fraud-related keywords and phrases with a list and scoring them based on their frequency and combinations.

[1262] 4. Fraud Score Evaluation Module (Server)

[1263] The server performs contextual analysis of the text data to detect speech patterns and other characteristics that are particularly characteristic of fraudulent activity. Based on this evaluation, a fraud likelihood score is calculated. If the score exceeds a set threshold, a warning flag is set.

[1264] 5. Emotion engine (server)

[1265] The server uses the voice data to recognize the user's emotional state: the emotion engine analyzes the tone, rate, and other voice characteristics to determine whether the user is experiencing emotions such as surprise, fear, or anxiety.

[1266] 6. Alert Adjustment Module (Server)

[1267] The server adjusts the content and notification method of the warning based on the analysis results of the emotion engine. For example, if the user is feeling strong anxiety or fear, the warning content can be made more detailed and clearer.

[1268] 7. Alert notification module (server and terminal)

[1269] If the warning flag is set, the server sends a warning message to the device, which receives the message and issues a voice or text warning to the user, either saying "Be careful, this may be a scam" in the case of a voice warning or a text message that pops up on the screen.

[1270] Specific examples

[1271] Example 1: Call fraud using emotion engine

[1272] 1. Users

[1273] The elderly user answers the phone and begins the conversation.

[1274] 2. Terminal

[1275] The device's voice analysis module records the call in real time and converts the audio "Payment required now" into text data.

[1276] 3. Terminal

[1277] The text data "Payment is required now" is sent to the server.

[1278] 4. Server

[1279] Analyzes text data to detect and score keywords such as "now" and "payment."

[1280] 5. Server

[1281] The system analyzes the urgent speech patterns characteristic of fraudulent activity and gives a high rating. If this exceeds a threshold, a warning flag is set.

[1282] 6. Server

[1283] At the same time, the emotion engine analyzes the voice data to recognize the user's emotional state, detecting when the user feels an urgent need and is anxious.

[1284] 7. Server

[1285] The warning content is adjusted based on the results of the emotion engine to generate stronger warning messages.

[1286] 8. Server

[1287] Sends tailored warning messages to the terminal.

[1288] 9. Terminal

[1289] The system will notify the user of the received warning message with a voice message, warning them, "Be careful, this may be a scam. Please tell a family member about this call immediately."

[1290] 10. Users

[1291] Users who receive the warning can end the call and consult with their family to prevent any harm.

[1292] In this way, the present invention can significantly reduce the risk of seniors becoming victims of fraud. The system operates in real time and issues appropriate warnings based on the user's emotional state, allowing for a fast and effective response.

[1293] The processing flow will be explained below.

[1294] Step 1:

[1295] The user answers the phone and starts the call. The device detects the start of the call and automatically starts the voice analysis module.

[1296] Step 2:

[1297] The device records the contents of the call in real time and instantly converts the acquired voice data into text data using voice recognition technology.

[1298] Step 3:

[1299] The terminal transmits the converted text data to the server at regular intervals (for example, every second).

[1300] Step 4:

[1301] The server receives the text data sent from the terminal and starts analysis using the text data analysis module.

[1302] Step 5:

[1303] The server detects keywords and phrases in the text data, including words related to fraud (e.g., "now," "payment," "legal action").

[1304] Step 6:

[1305] The server then scores the likelihood of fraud based on the frequency and combination of keywords detected, calculated according to pre-defined criteria.

[1306] Step 7:

[1307] The server performs contextual analysis to detect speech patterns that are particularly characteristic of fraudulent activity (forced language and emphasis on urgency).

[1308] Step 8:

[1309] The server evaluates the scoring results and sets a warning flag if the fraud likelihood score exceeds a set threshold.

[1310] Step 9:

[1311] The server sends the voice data to the emotion engine, which analyzes the tone, rate, and other voice characteristics of the user's voice.

[1312] Step 10:

[1313] The emotion engine on the server determines the user's emotional state, for example, analyzing whether the user is feeling surprise, fear, anxiety, etc.

[1314] Step 11:

[1315] The server adjusts the content of the warning and notification method based on the analysis results of the emotion engine. If the user is feeling strong anxiety or fear, the warning content will be more detailed and clearer.

[1316] Step 12:

[1317] The server sends a tailored warning message to the device, which includes information about the likely fraud and specific steps to take.

[1318] Step 13:

[1319] As soon as the device receives the warning notification from the server, it immediately issues a warning to the user, either by voice or text, saying, "Caution, this may be a scam. Please inform a family member immediately about this call."

[1320] Step 14:

[1321] The user receives a warning and becomes suspicious of the call. They end the call and consult with family or friends for safety. This process significantly reduces the risk of the user becoming a victim of fraud.

[1322] Example 2

[1323] Next, a description will be given of Example 2. In the following description, the data processing device 12 will be referred to as a "server" and the headset type terminal 314 will be referred to as a "terminal."

[1324] In today's world, fraud is becoming increasingly sophisticated, making the elderly especially vulnerable to fraud. Previous fraud prevention systems primarily relied on text data analysis and did not provide real-time warnings that took the user's emotional state into account. This made it difficult for users to quickly recognize fraud and take appropriate action. Furthermore, because conventional systems issued warnings uniformly, they often failed to respond appropriately to the sense of crisis and anxiety felt by users. This has led to a need for effective methods to prevent fraud before it happens.

[1325] The identification process by the identification processing unit 290 of the data processing device 12 in Example 2 is realized by the following means. In this invention, the server includes means for converting voice data into text data in real time, means for analyzing the text data and identifying keywords and phrases related to fraud, means for evaluating the possibility of fraud based on the analysis results and scoring the risk level, means for recognizing the user's emotional state from the voice analysis and adjusting the warning content based on the emotion, and means for transmitting the warning content to the terminal and notifying the user. This makes it possible to provide an appropriate warning in real time that takes into account the user's emotional state.

[1326] "Voice data" refers to sound waveforms such as those from a call or voice input that have been converted into electrical signals or digital data.

[1327] "Text data" refers to character string data generated as a result of analyzing and converting voice data.

[1328] "Analysis" is the process of dividing data for a specific purpose to reveal meaning and structure.

[1329] "Fraud-related keywords and phrases" refer to specific words and expressions that are likely to be used in fraud, and are used to predict the possibility of fraud by detecting them.

[1330] "Risk level" is a numerical or score that expresses the likelihood of a scam.

[1331] "Scoring" is the process of quantifying risk based on the results of data analysis.

[1332] A "threshold" is a numerical value or condition that triggers a specific action, and when exceeded, a warning is issued.

[1333] A "warning" is information that notifies the user of a potential danger and urges caution.

[1334] "Emotional state" refers to the user's psychological state analyzed from the tone, rate, and other characteristics of the voice.

[1335] "Adjustment" is the process of changing something to an appropriate form depending on the situation or conditions.

[1336] A "notification" is an action or means of informing a user of specific information.

[1337] The present invention provides a more accurate warning function by combining a system that converts voice data into text data in real time, evaluates the possibility of fraud, and notifies the user of a warning with an emotion engine that recognizes the user's emotions.

[1338] System Configuration

[1339] The system includes the following modules:

[1340] Voice analysis module (terminal)

[1341] The device activates a voice analysis module at the start of a call and records the call in real time. The recorded voice data is converted to text data using the Google Cloud Speech-to-Text API.

[1342] Text data transmission module (terminal)

[1343] The converted text data is sent to the server at regular intervals of less than one second using the WebSocket protocol in real time.

[1344] Fraud detection module (server)

[1345] The server analyzes the received text data using a natural language processing library (e.g., NLTK or SpaCy), matching fraud-related keywords and phrases with a list, and scoring them based on their frequency and combinations.

[1346] Fraud score evaluation module (server)

[1347] The server then performs contextual analysis of the text data to detect speech patterns and other characteristics of fraudulent activity. This information is used to calculate a fraud likelihood score. If this score exceeds a set threshold, a warning flag is set.

[1348] Emotion engine (server)

[1349] The server uses the voice data to recognize the user's emotional state, using IBM Watson Tone Analyzer and Microsoft Azure Emotion API to determine whether the user is experiencing emotions such as surprise, fear, or anxiety.

[1350] Alert Adjustment Module (Server)

[1351] The server adjusts the content and notification method of the warning based on the results of the emotion engine. For example, if the user is feeling strong anxiety or fear, the warning content will be more detailed and clearer.

[1352] Alert notification module (server and terminal)

[1353] The server sends the tailored warning message to the device, which receives it and issues a voice or text warning to the user, such as "Be careful, this may be a scam," or a text message that pops up on the screen.

[1354] Specific examples

[1355] Example 1: Call fraud using emotion engine

[1356] 1. Users

[1357] The elderly user answers the phone and begins the conversation.

[1358] 2. Terminal

[1359] The device's voice analysis module records the call in real time and instantly converts the audio "Payment required now" into text data using Google Cloud Speech-to-Text.

[1360] 3. Terminal

[1361] The converted text data "Payment required now" is sent to the server via the WebSocket protocol at intervals of less than one second.

[1362] 4. Server

[1363] The server uses NLTK to analyze the received text data, identifying fraud-related keywords such as "now" and "payment," and analyzing the frequency and context of these keywords.

[1364] 5. Server

[1365] It detects urgent speech patterns characteristic of fraudulent activity and generates a high fraud likelihood score. If this score exceeds a set threshold, it sets a warning flag.

[1366] 6. Server

[1367] The server uses an emotion engine (IBM Watson Tone Analyzer) to analyze the user's emotional state from the voice data. The analysis results indicate that the user feels urgency and is anxious.

[1368] 7. Server

[1369] The content of the warning message is adjusted based on the results of the emotion engine, for example, generating a detailed warning such as "This call may be fraudulent. Please report this call to a family member immediately."

[1370] 8. Server

[1371] Sends tailored warning messages to the terminal.

[1372] 9. Terminal

[1373] The device will then notify the user of the received warning message via voice: "Caution, this may be a scam. Please inform a family member immediately about this call."

[1374] 10. Users

[1375] The user will receive a warning, end the call, and immediately contact their family to prevent any harm.

[1376] In this way, the present invention can significantly reduce the risk of seniors becoming victims of fraud. The system operates in real time and issues appropriate warnings based on the user's emotional state, allowing for a fast and effective response.

[1377] Prompt Sentence Examples

[1378] "Design a system to detect emotions from voice data, assess the likelihood of fraud, and provide a warning. Use natural language processing and emotion recognition technologies for data analysis. Please also specify the specific hardware and software, data flow, and user interface (UI)."

[1379] The flow of the identification process in the second embodiment will be described with reference to FIG.

[1380] Step 1:

[1381] Voice recording and text conversion (device)

[1382] Input: A user receives a call and starts a conversation.

[1383] How it works: The device detects that a call has been initiated and activates the voice analysis module.

[1384] Data processing: Convert the audio data into text data in real time using the Google Cloud Speech-to-Text API.

[1385] Output: The converted text data is generated.

[1386] Step 2:

[1387] Sending text data (terminal)

[1388] Input: The text data generated in step 1.

[1389] Operation: The device sends the converted text data to the server at intervals of less than one second.

[1390] Data calculation: Send data in real time using the WebSocket protocol.

[1391] Output: The server receives text data in real time.

[1392] Step 3:

[1393] Text data analysis (server)

[1394] Input: The text data received in step 2.

[1395] How it works: The server parses the received text data using a natural language processing library (NLTK or SpaCy).

[1396] Data calculations: Fraud-related keywords and phrases are matched against the list and analyzed based on their frequency and combinations.

[1397] Output: A numerical score is generated representing the likelihood of fraud.

[1398] Step 4:

[1399] Fraud score evaluation (server)

[1400] Input: Analysis results (scores) generated in step 3.

[1401] How it works: The server also performs contextual analysis of the text data to detect speech patterns and patterns characteristic of fraudulent activity.

[1402] Data calculation: Calculate a fraud likelihood score and set a warning flag if this score exceeds a set threshold.

[1403] Output: Fraud probability score and warning flag status.

[1404] Step 5:

[1405] User emotion analysis (server)

[1406] Input: The audio data recorded in step 1.

[1407] How it works: The server uses IBM Watson Tone Analyzer and Microsoft Azure Emotion API to analyze the user's emotional state from voice data.

[1408] Data calculations: Analyze the tone, rate, and other voice characteristics of the voice to determine whether the user is surprised, scared, or anxious.

[1409] Output: Generates an analysis of the user's emotional state.

[1410] Step 6:

[1411] Adjustment of warning content (server)

[1412] Input: Warning flags from step 4 and sentiment analysis results from step 5.

[1413] How it works: The server tailors the warning based on the user's emotional state. For example, if the user is feeling strong anxiety or fear, the warning will be more detailed and clearer.

[1414] Data calculation: Optimize the content and format of warning messages based on the results of sentiment analysis.

[1415] Output: A tailored warning message is generated.

[1416] Step 7:

[1417] Notification of warning messages (server and terminal)

[1418] Input: The warning message adjusted in step 6.

[1419] Action: The server sends a tailored alert message to the device.

[1420] Data calculation: The device notifies the user of the received warning message by voice or text.

[1421] Output: A voice or text alert is given to the user.

[1422] These steps enable the system to detect potential fraud in real time and issue warnings that take into account the user's emotional state.

[1423] (Application example 2)

[1424] Next, a description will be given of Application Example 2. In the following description, the data processing device 12 will be referred to as a "server" and the headset type terminal 314 will be referred to as a "terminal."

[1425] In recent years, telephone fraud has been on the rise, with elderly people becoming victims in particular. Conventional fraud prevention measures only issue a warning to users when fraud is suspected. Therefore, even when users receive a warning, they are unable to properly determine how to respond, resulting in a high likelihood of fraud victimization. Furthermore, because warnings are issued uniformly with the same intensity without taking into account the user's emotional state, the effectiveness of the warnings is limited. There is a need to improve this situation and provide more effective fraud prevention measures.

[1426] The specific processing by the specific processing unit 290 of the data processing device 12 in the application example 2 is realized by the following means.

[1427] In this invention, the server includes means for converting voice data into text data in real time, means for analyzing the text data and identifying keywords and phrases related to fraud, means for assessing the likelihood of fraud and scoring the risk based on the analysis results, means for issuing a warning to the user when the risk exceeds a set threshold, means for analyzing the user's emotional state, means for adjusting the content of the warning based on the emotion analysis results, and means for issuing the adjusted warning to the user. This makes it possible to accurately assess the likelihood of fraud in real time based on the voice data and provide an appropriate warning that takes the user's emotional state into consideration.

[1428] "Voice data" means recorded information of a user's voice uttered over the telephone or other voice communication.

[1429] "Real-time" refers to near-instant processing with little actual time or delay.

[1430] "Text data" is voice data converted into character information.

[1431] Analysis is the process of examining data or information in detail to identify important elements or patterns within it.

[1432] "Fraud-related keywords and phrases" are specific words and expressions commonly used in fraudulent activities.

[1433] "Scoring" is the process of assigning points according to specific criteria based on the analysis results.

[1434] A "threshold" is a boundary value for determining a certain condition or standard.

[1435] A "warning" is a notification that notifies the user of an impending danger.

[1436] An "emotional state" is the psychological state that a user is feeling at a particular moment.

[1437] "Notification" is the act of conveying information or a message to a target person.

[1438] "Adjustment" refers to changing something to an optimal state for specific conditions or environments.

[1439] This invention is a system for protecting users from telephone fraud, and provides highly accurate warning functions by combining analysis of voice data and emotion analysis.

[1440] System Overview

[1441] The system mainly consists of three main components.

[1442] 1. Terminal

[1443] 2. Server

[1444] 3. Users

[1445] Hardware and software used

[1446] Hardware:

[1447] Smartphone: Used as a device, requires a microphone and internet connection.

[1448] software:

[1449] Python 3.x: The main programming language.

[1450] speech_recognition library: Uses Google's speech recognition API to convert voice data into text data.

[1451] The requests library: handles HTTP requests to the server.

[1452] Processing Details

[1453] Terminal handling

[1454] When a user initiates a phone call, the device records the voice data in real time and converts it into text data using the speech_recognition library, which is then sent to the server at regular intervals.

[1455] Server Processing

[1456] The server analyzes the audio data and assesses the likelihood of fraud using the following steps:

[1457] 1. Text Analysis: Analyzes incoming text data to identify keywords and phrases related to fraud.

[1458] 2. Scoring: Based on the analysis results, the possibility of fraud is assessed and a risk score is generated. If the score exceeds a set threshold, a warning is generated.

[1459] 3. Emotion analysis: Analyze the user's emotional state from the voice data, determining whether the user is feeling surprised or anxious.

[1460] 4. Adjustment of warning content: Based on the results of sentiment analysis, the warning content is adjusted appropriately.

[1461] Warning Notification

[1462] A tailored warning message is sent to the user's device. For example, if the warning is delivered via voice, the content and intensity of the warning can be adjusted according to the user's emotional state. This is expected to encourage the user to take appropriate action.

[1463] Specific examples

[1464] For example, consider the case of an elderly person receiving a suspected fraudulent phone call. If the user hears phrases such as "Payment is required now," the system immediately transcribes the speech into text and analyzes it. If it detects patterns indicative of fraudulent activity, the system takes into account the user's emotional state and generates a strong warning message. The warning could include specific advice such as "Be careful, this may be a scam. Please report this call to a family member immediately."

[1465] Example prompts for generative AI models

[1466] Your system should take voice data from a user's phone call as input and convert it into text data in real time. It should then assess the likelihood of fraud (by scoring it based on the frequency of fraud-related keywords and phrases) and generate a warning message if necessary. It should also analyze the user's emotional state and adjust the strength and content of the warning. The emotion engine uses the tone, rate, and other voice characteristics to determine whether the user is feeling anxious or scared. It should generate a warning message based on the results of the text and emotion analysis and notify the user.

[1467] The flow of the specific processing in the application example 2 will be described with reference to FIG.

[1468] Step 1:

[1469] A user initiates a phone call.

[1470] Input: User's call audio

[1471] Output: Call audio data

[1472] Specific operation: A user initiates a call and voice communication begins over the phone.

[1473] Step 2:

[1474] The device records the call audio in real time.

[1475] Input: Call audio data

[1476] Output: Recorded audio file

[1477] Specific operation: The device's microphone captures the user's call voice and records the voice data.

[1478] Step 3:

[1479] The device converts recorded voice data into text data in real time.

[1480] Input: Recorded audio file

[1481] Output: Text data

[1482] Specific operation: Using the speech_recognition library installed on the device, the recorded voice data is converted into text data.

[1483] Step 4:

[1484] The terminal transmits the converted text data to the server at regular intervals.

[1485] Input: Text data

[1486] Output: Send text data to the server

[1487] Specific operation: Sends text data to the server using an HTTP request, for example, using the requests library.

[1488] Step 5:

[1489] The server analyzes the received text data to identify keywords and phrases related to fraud.

[1490] Input: Text data sent to the server

[1491] Output: Analysis results (list of fraud-related keywords and phrases)

[1492] How it works: An analytics module on the server scans the text data to identify keywords and phrases related to fraud, matching them against a predefined list of fraud keywords.

[1493] Step 6:

[1494] The server evaluates the likelihood of fraud and assigns a score based on the results of analyzing the text data.

[1495] Input: Analysis results

[1496] Output: Risk score

[1497] Specific operation: The server's scoring module evaluates the likelihood of fraud based on the frequency of fraud keywords and the context of the text, and calculates a risk score.

[1498] Step 7:

[1499] The server generates a warning message if the risk score exceeds a threshold.

[1500] Input: Risk score

[1501] Output: Warning message

[1502] Specific behavior: If the score exceeds a preset threshold, an algorithm is executed that generates a warning message.

[1503] Step 8:

[1504] The server analyzes the user's emotional state from the voice data.

[1505] Input: Audio data

[1506] Output: Emotion analysis results

[1507] Specific behavior: The emotion engine analyzes the tone, rate, and other voice characteristics of the voice to assess the user's emotional state.

[1508] Step 9:

[1509] The server adjusts the warning content based on the results of sentiment analysis.

[1510] Input: Sentiment analysis results, warning message

[1511] Output: Adjusted warning message

[1512] Specific operation: Using the results of the emotion engine, an algorithm is run that adjusts the content and intensity of warning messages.

[1513] Step 10:

[1514] The server sends the adjusted warning message to the terminal, and the terminal notifies the user of the warning.

[1515] Input: Adjusted warning message

[1516] Output: A warning notice to the user

[1517] Specific operation: The server sends a tailored warning message to the terminal, which then notifies the user of the message. The warning is notified as voice or text.

[1518] Each step is performed sequentially, protecting users from the risk of fraud in real time.

[1519] The specific processing unit 290 transmits the result of the specific processing to the headset type terminal 314. In the headset type terminal 314, the control unit 46A causes the speaker 240 and the display 343 to output the result of the specific processing. The microphone 238 acquires audio indicating a user input regarding the result of the specific processing. The control unit 46A transmits audio data indicating the user input acquired by the microphone 238 to the data processing device 12. In the data processing device 12, the specific processing unit 290 acquires the audio data.

[1520] The data generation model 58 is a so-called generative AI (Artificial Intelligence). An example of the data generation model 58 is ChatGPT (Internet Search<URL: https: / / openai.com / blog / chatgpt> ), Gemini (Internet search <url: https: gemini.google.com ?hl="ja">) and other generation AIs. The data generation model 58 is obtained by performing deep learning on a neural network. A prompt including an instruction is input to the data generation model 58, and inference data such as voice data indicating voice, text data indicating text, and image data indicating an image is also input. The data generation model 58 performs inference on the input inference data in accordance with the instruction indicated by the prompt, and outputs the inference result in a data format such as voice data and text data. Here, inference refers to, for example, analysis, classification, prediction, and / or summarization.

[1521] In the above embodiment, an example was given in which the specific processing is performed by the data processing device 12, but the technology of the present disclosure is not limited to this, and the specific processing may be performed by the headset type terminal 314.

[1522] [Fourth embodiment]

[1523] FIG. 7 shows an example of the configuration of a data processing system 410 according to the fourth embodiment.

[1524] 7, a data processing system 410 includes a data processing device 12 and a robot 414. An example of the data processing device 12 is a server.

[1525] The data processing device 12 includes a computer 22, a database 24, and a communication I / F 26. The computer 22 is an example of a "computer" according to the technology of the present disclosure. The computer 22 includes a processor 28, a RAM 30, and a storage 32. The processor 28, the RAM 30, and the storage 32 are connected to a bus 34. The database 24 and the communication I / F 26 are also connected to the bus 34. The communication I / F 26 is connected to a network 54. Examples of the network 54 include a WAN (Wide Area Network) and / or a LAN (Local Area Network).

[1526] The robot 414 includes a computer 36, a microphone 238, a speaker 240, a camera 42, a communication I / F 44, and a control target 443. The computer 36 includes a processor 46, a RAM 48, and a storage 50. The processor 46, the RAM 48, and the storage 50 are connected to a bus 52. The microphone 238, the speaker 240, the camera 42, and the control target 443 are also connected to the bus 52.

[1527] The microphone 238 receives instructions and the like from the user 20 by receiving voice uttered by the user 20. The microphone 238 captures the voice uttered by the user 20, converts the captured voice into audio data, and outputs it to the processor 46. The speaker 240 outputs audio in accordance with instructions from the processor 46.

[1528] Camera 42 is a small digital camera equipped with an optical system including a lens, aperture, and shutter, and an imaging element such as a CMOS (Complementary Metal-Oxide-Semiconductor) image sensor or a CCD (Charge Coupled Device) image sensor, and captures images of the surroundings of user 20 (for example, an imaging range defined by an angle of view equivalent to the field of vision of a typical healthy person).

[1529] The communication I / F 44 is connected to a network 54. The communication I / Fs 44 and 26 are responsible for the exchange of various information between the processor 46 and the processor 28 via the network 54. The exchange of various information between the processor 46 and the processor 28 using the communication I / Fs 44 and 26 is carried out in a secure state.

[1530] The control object 443 includes a display device, LEDs in the eyes, and motors for driving the arms, hands, and feet. The posture and gestures of the robot 414 are controlled by controlling the motors of the arms, hands, and feet. Some of the emotions of the robot 414 can be expressed by controlling these motors. In addition, the facial expressions of the robot 414 can also be expressed by controlling the light emission state of the LEDs in the eyes of the robot 414.

[1531] Fig. 8 shows an example of the main functions of the data processing device 12 and the robot 414. As shown in Fig. 8, in the data processing device 12, a specific process is performed by the processor 28. A specific process program 56 is stored in the storage 32.

[1532] The specific processing program 56 is an example of a "program" according to the technology of the present disclosure. The processor 28 reads the specific processing program 56 from the storage 32 and executes the read specific processing program 56 on the RAM 30. The specific processing is realized by the processor 28 operating as a specific processing unit 290 in accordance with the specific processing program 56 executed on the RAM 30.

[1533] The storage 32 stores a data generation model 58 and an emotion identification model 59. The data generation model 58 and the emotion identification model 59 are used by the identification processing unit 290.

[1534] In the robot 414, the processor 46 performs the reception output process. A reception output program 60 is stored in the storage 50. The processor 46 reads the reception output program 60 from the storage 50 and executes the read reception output program 60 on the RAM 48. The reception output process is realized by the processor 46 operating as the control unit 46A in accordance with the reception output program 60 executed on the RAM 48.

[1535] Next, a description will be given of the specific processing performed by the specific processing unit 290 of the data processing device 12. In the following description, the data processing device 12 will be referred to as a "server" and the robot 414 will be referred to as a "terminal."

[1536] The present invention is a telephone fraud prevention system for elderly people that converts voice data into text data in real time, evaluates the possibility of fraud, and notifies the user with a warning. The system has the following configuration.

[1537] 1. Voice analysis module (terminal)

[1538] When a call starts, the device activates a voice analysis module and records the call in real time. The recorded voice data is then converted into text data on the spot using voice recognition technology.

[1539] 2. Text data transmission module (terminal)

[1540] The converted text data is sent to the server at regular intervals of less than one second, so that the data is sent in near real time.

[1541] 3. Fraud detection module (server)

[1542] The server analyzes the received text data, specifically matching fraud-related keywords and phrases with a list and scoring them based on their frequency and combinations.

[1543] 4. Fraud Score Evaluation Module (Server)

[1544] The server performs contextual analysis of the text data to detect speech patterns and other characteristics that are particularly characteristic of fraudulent activity. Based on this evaluation, a fraud likelihood score is calculated. If the score exceeds a set threshold, a warning flag is set.

[1545] 5. Alert notification module (server and terminal)

[1546] When the warning flag is set, the server sends a warning message to the device, which receives the message and issues a voice or text warning to the user, saying "Be careful, this may be a scam," or a text message that pops up on the screen.

[1547] Specific examples

[1548] Example 1: Fake billing scam call

[1549] 1. Users

[1550] The elderly user answers the phone and begins the conversation.

[1551] 2. Terminal

[1552] The device's voice analysis module records the call in real time and converts the audio "Payment required now" into text data.

[1553] 3. Terminal

[1554] The text data "Payment is required now" is sent to the server.

[1555] 4. Server

[1556] Analyzes text data to detect and score keywords such as "now" and "payment."

[1557] 5. Server

[1558] The fraud score assessment module gives a high score and sets a warning flag because the threshold has been exceeded.

[1559] 6. Server

[1560] Sends a warning message to the terminal.

[1561] 7. Terminal

[1562] The system will notify the user of the received warning message by voice, warning them with "Be careful, this may be a scam."

[1563] 8. Users

[1564] Users who receive the warning should end the call and consult with their family to prevent further harm.

[1565] In this way, the risk of seniors being scammed can be significantly reduced by this invention. The system operates in real time and issues immediate warnings, allowing for a rapid response.

[1566] The processing flow will be explained below.

[1567] Step 1:

[1568] The user answers the phone and starts the call. The device detects the start of the call and automatically starts the voice analysis module.

[1569] Step 2:

[1570] The device records the contents of the call in real time and instantly converts the acquired voice data into text data using voice recognition technology.

[1571] Step 3:

[1572] The terminal transmits the converted text data to the server at regular intervals (for example, every second).

[1573] Step 4:

[1574] The server receives the text data sent from the terminal and starts analysis using the text data analysis module.

[1575] Step 5:

[1576] The server detects keywords and phrases in the text data, including words related to fraud (e.g., "now," "payment," "legal action").

[1577] Step 6:

[1578] The server then scores the likelihood of fraud based on the frequency and combination of keywords detected, calculated according to pre-defined criteria.

[1579] Step 7:

[1580] The server performs contextual analysis to detect speech patterns that are particularly characteristic of fraudulent activity (forced language and emphasis on urgency).

[1581] Step 8:

[1582] The server evaluates the scoring results and sets a warning flag if the fraud likelihood score exceeds a set threshold.

[1583] Step 9:

[1584] The server notifies the terminal that a warning flag has been set, including a message indicating a high probability of fraud.

[1585] Step 10:

[1586] As soon as the device receives the warning notification from the server, it immediately issues a warning to the user, either by voice or text, saying, "Be careful, this may be a scam."

[1587] Step 11:

[1588] The user receives a warning and becomes suspicious of the call. The user ends the call and consults with family or friends for safety.

[1589] This series of steps enables the system to detect fraudulent activity with high accuracy and reduce the risk of users becoming victims.

[1590] Example 1

[1591] Next, a description will be given of Example 1. In the following description, the data processing device 12 will be referred to as a "server" and the robot 414 will be referred to as a "terminal."

[1592] Currently, the number of elderly victims of telephone fraud is increasing, and fraudulent acts via telephone in particular has become a social problem. Measures to prevent this are needed. However, current methods are difficult to respond to in real time, and fraud prevention is not sufficient. Therefore, there is a need for a system that can reduce the risk of elderly people becoming victims of telephone fraud and issue warnings quickly and effectively.

[1593] The specific processing by the specific processing unit 290 of the data processing device 12 in the first embodiment is realized by the following means.

[1594] In this invention, the server includes means for converting voice data into text data in real time, means for transmitting the text data to the server at regular intervals, means for analyzing the text data and identifying fraud-related keywords and phrases, means for evaluating the possibility of fraud and scoring a risk level based on the analysis results, means for detecting fraud patterns based on the fraud-related keywords and phrases and calculating a fraud likelihood score, and means for issuing a voice or text warning to the user when the risk level exceeds a set threshold. This significantly reduces the risk of elderly people becoming victims of telephone fraud and enables prompt warnings to be issued in real time.

[1595] "Voice data" refers to data that records or transmits voice information such as telephone calls or conversations in digital format.

[1596] "Text data" is data in a format in which voice data is converted into a string of characters, and is information expressed in a format that can be read by humans.

[1597] "Real-time" means that the system processes voice input immediately and outputs the results without delay.

[1598] "Analysis" is the process of examining the information contained in data in detail and extracting and evaluating that information based on specific criteria and rules.

[1599] "Fraud-related keywords and phrases" refer to expressions and words characteristic of fraudulent activities, and are specific words and phrases that are judged to have a high possibility of being fraudulent.

[1600] "Scoring" refers to the process of quantifying and evaluating the likelihood of fraud based on specific criteria or algorithms based on the analysis results.

[1601] A "threshold" is a numerical value or condition that determines when a system will take a particular action or perform a particular algorithm.

[1602] A "warning" is a message or notification that warns the user and is issued when there is a possibility of fraud.

[1603] "Interval" refers to the timing or cycle at which data is transmitted, and indicates the time difference between successive data transmissions.

[1604] "Patterns" refer to speech patterns and contextual structures that are characteristic of fraudulent behavior, and are standards or models based on which to judge the possibility of fraud.

[1605] "Calculation" is the process by which a system generates a number or evaluation according to a specified algorithm or procedure.

[1606] "Voice or text notification" means that the warning to the user is given in the form of voice output or text display, and the system issues the warning in an appropriate manner.

[1607] This invention is a telephone fraud prevention system aimed at the elderly, which analyzes the voice of a call in real time and issues a warning when there is a high possibility of fraud. The specific configuration is as follows.

[1608] Voice analysis module (terminal)

[1609] User

[1610] When the elderly user receives a call, the call begins. The system starts working when the user presses the call button.

[1611] Terminal

[1612] When a call is received, the device automatically activates a voice analysis module, which records the call in real time. The recorded voice data is then converted into text data on the fly using voice recognition technology. Specifically, real-time voice-to-text conversion is performed using the Google Speech-to-Text API.

[1613] Text data transmission module (terminal)

[1614] Terminal

[1615] The converted text data is sent to the server at regular intervals (less than 1 second). The HTTP protocol is used to send the data, for example, as a POST request. Specifically, the text data is sent to http: / / server_address / text_data.

[1616] Fraud detection module (server)

[1617] server

[1618] The server analyzes the received text data. For fraud detection, it uses a pre-prepared list of fraud-related keywords to scan the text data using natural language processing (NLP) techniques. Specifically, it uses the Python nltk library to tokenize the text and verify whether each token is included in the fraud list.

[1619] For example, if someone sends a text saying "I need payment now," we tokenize it and check if keywords like "now," "payment," and "need" are included in our fraud list.

[1620] Fraud score evaluation module (server)

[1621] server

[1622] The fraud score evaluation module on the server calculates scores based on the frequency of keyword occurrence and their combinations. The scoring is done using machine learning models, utilizing Scikit-learn and TensorFlow.

[1623] "Payment required now" will receive a high score because "now" and "payment" are included in the fraud list. If the score exceeds a set threshold, a warning flag will be set.

[1624] Alert notification module (server and terminal)

[1625] server

[1626] When the warning flag is set, the server sends a warning message to the device, which may include a message like "Possible fraud. Please be careful."

[1627] Terminal

[1628] The device will notify the user of the received warning message. There are two notification methods: voice warning and text message. In the case of a voice warning, the device will use TTS (Text-to-Speech) technology to warn the user by voice, saying "Be careful, this may be a scam." In the case of a text message, it will be displayed as a pop-up on the device screen.

[1629] As a specific example of how this works, elderly users can hear a warning during a call saying, "Be careful, this may be a scam." This allows users to sense the risk of being deceived by a scam in advance and take appropriate action.

[1630] Examples of prompts include:

[1631] Describe a telephone fraud prevention system for seniors. The system converts voice data into text in real time, evaluates the likelihood of fraud, and alerts the user. Explain each step in detail, including specific operations and the names of any hardware or software used.

[1632] The flow of the identification process in the first embodiment will be described with reference to FIG.

[1633] System program processing flow

[1634] Step 1: Start the call

[1635] User

[1636] The elderly user answers the phone and starts the conversation. The user presses the talk button and the system starts working.

[1637] Input: Incoming call / call start signal

[1638] Output: System start signal

[1639] Step 2: Launching Audio Analysis

[1640] Terminal

[1641] Upon receiving the call start signal, the terminal automatically starts the voice analysis module.

[1642] How it works: The audio analysis module records the call.

[1643] Input: Call start signal

[1644] Output: Start recording audio data

[1645] Step 3: Convert the audio to text

[1646] Terminal

[1647] The speech analysis module converts the recorded voice data into text data in real time. The speech-to-text conversion is performed using the Google Speech-to-Text API.

[1648] How it works: Captures audio as digital data and sends it to an API to retrieve text.

[1649] Input: Recorded audio data

[1650] Output: Converted text data

[1651] Step 4: Send text data

[1652] Terminal

[1653] The converted text data is sent to the server at regular intervals (less than 1 second). The data is sent using an HTTP POST request.

[1654] Action: Sends data to http: / / server_address / text_data.

[1655] Input: Text data

[1656] Output: Send data to the server

[1657] Step 5: Analyzing the text data

[1658] server

[1659] The server analyzes the received text data and uses a pre-prepared list of fraud-related keywords to detect fraud. It uses the Python nltk library to tokenize the text and verify whether each token is included in the fraud list.

[1660] How it works: It breaks down text data and matches it with keywords from a scam list.

[1661] Input: Received text data

[1662] Output: Keyword occurrences

[1663] Step 6: Scoring

[1664] server

[1665] Based on the occurrence of fraud keywords, the server scores the likelihood of fraud. It calculates the fraud likelihood score using machine learning models such as Scikit-learn and TensorFlow.

[1666] How it works: Evaluates keyword frequency and patterns to calculate a fraud likelihood score.

[1667] Input: Keyword occurrence information

[1668] Output: Fraud likelihood score

[1669] Step 7: Set warning flags

[1670] server

[1671] If the fraud likelihood score exceeds a set threshold, the server sets a warning flag.

[1672] Behavior: Evaluate the score and set a flag if it exceeds a threshold.

[1673] Input: Fraud likelihood score

[1674] Output: warning flag

[1675] Step 8: Sending a warning message

[1676] server

[1677] When the warning flag is set, the server sends a warning message to the terminal, which may include a message such as "Possible fraud. Please be careful."

[1678] Action: Generates a warning message and sends it to the terminal.

[1679] Input: warning flag

[1680] Output: Warning message

[1681] Step 9: Warning Notification

[1682] Terminal

[1683] The device will notify the user of the received warning message. Notification methods include voice alert and text message. Voice alert uses TTS technology, and text message will be displayed as a pop-up on the screen.

[1684] Action: Outputs a voice or text alert based on the message content.

[1685] Input: warning message

[1686] Output: A warning notice to the user

[1687] (Application example 1)

[1688] Next, a description will be given of Application Example 1. In the following description, the data processing device 12 will be referred to as a "server" and the robot 414 will be referred to as a "terminal."

[1689] Telephone frauds against the elderly are increasing year by year, and the amount of damage caused is also steadily increasing. Because the elderly are particularly susceptible to fraudulent schemes, there is a need for fast and effective fraud prevention measures. However, many current fraud prevention systems require users to operate and make decisions themselves, which can result in insufficient effectiveness. Furthermore, there are few systems that can prevent telephone fraud in real time, and the risk of being scammed cannot be completely eliminated. Therefore, there is a need for an effective system that can evaluate and notify users of possible fraud in real time while reducing the operational burden on users.

[1690] The specific processing by the specific processing unit 290 of the data processing device 12 in the application example 1 is realized by the following means.

[1691] In this invention, the server includes means for converting voice data into text data in real time, means for analyzing the text data and identifying keywords and phrases related to fraud, means for assessing the likelihood of fraud based on the analysis results and scoring a risk level, means for notifying a user of a warning if the risk level exceeds a set threshold, means for notifying the user of the warning using a visual display means and a voice notification means of the smart glasses, and means for performing contextual analysis of the text data and detecting patterns specific to fraudulent activities using a generative AI model, thereby significantly reducing the risk of users being involved in telephone fraud and enabling fast and effective fraud prevention in real time.

[1692] "Voice data" refers to data obtained by converting voice signals into digital format, and is the content of a conversation saved in digital format as is.

[1693] "Text data" refers to data in the form of a character string converted from audio data.

[1694] "Keywords" are specific words or phrases that are frequently used in contexts or content that are likely to be fraudulent.

[1695] "Scoring" is an evaluation method that numerically represents the likelihood of fraud.

[1696] A "threshold" is a numerical standard used to assess the likelihood of fraud, and if this value is exceeded, a warning will be issued.

[1697] A "warning" is a notification that informs the user of the risk when there is a high possibility of fraud.

[1698] "Smart glasses" are wearable devices that provide information to users through sight and sound.

[1699] "Visual display means" refers to a mechanism for displaying text and images using a device such as smart glasses.

[1700] A "generative AI model" is a natural language processing model that learns from large amounts of language data using machine learning technology.

[1701] "Contextual analysis" is a technology that understands the content of text data and extracts specific meanings and patterns.

[1702] "Pattern detection" is a technique that analyzes text data to identify speech patterns and phrases characteristic of fraudulent activity.

[1703] The present invention is a telephone fraud prevention system for elderly people that converts voice data into text data in real time, evaluates the possibility of fraud, and notifies the user with a warning. The system has the following configuration.

[1704] 1. Voice analysis module (smart glasses)

[1705] The smart glasses use a built-in microphone to capture the voice during a call in real time, and the voice data is converted into text data in real time using a common cloud speech recognition service (e.g., Google Cloud Speech-to-Text API).

[1706] 2. Text data transmission module (smart glasses)

[1707] The converted text data is sent to a cloud server at regular intervals (less than one second). The cloud server is equipped with a high-performance data processing system, enabling rapid data processing.

[1708] 3. Fraud detection module (cloud server)

[1709] The cloud server analyzes the received text data using natural language processing libraries (e.g., spaCy, NLTK) and scores the frequency of occurrence of fraud-related keywords and phrases. This analysis is powered by a powerful text analysis engine.

[1710] 4. Fraud score evaluation module (cloud server)

[1711] The server then performs contextual analysis using generative AI models (e.g., BERT or GPT-4) to detect patterns characteristic of fraudulent activity, which then calculates a fraud likelihood score and sets a warning flag if the score exceeds a threshold.

[1712] 5. Alert notification module (cloud server and smart glasses)

[1713] If the fraud score exceeds a threshold, the cloud server generates a warning message and sends it to the smart glasses, which then alert the user using visual (HUD display) and audio notification means. For example, a text pop-up on the HUD display reads, "Beware, this may be a scam," along with an audio notification stating the same.

[1714] Specific examples

[1715] Example 1: Fake billing scam call

[1716] 1. Users

[1717] The elderly user answers the phone and begins the conversation.

[1718] 2. Smart Glasses

[1719] The voice analysis module records the call in real time and converts the audio "Payment required now" into text data.

[1720] 3. Smart Glasses

[1721] The converted text data "Payment required now" is sent to the cloud server.

[1722] 4. Cloud Server

[1723] The text data is analyzed to detect keywords such as "now" and "payment" and then scored.

[1724] 5. Cloud Server

[1725] The fraud score assessment module gives a high rating and sets a warning flag because the threshold has been exceeded.

[1726] 6. Cloud Server

[1727] Sending a warning message to the smart glasses.

[1728] 7. Smart Glasses

[1729] The user will be notified of the received warning message via voice, warning them "Beware, this may be a scam," and the same text will pop up on the HUD display.

[1730] 8. Users

[1731] Upon receiving the warning, the user ends the call and consults with a family member.

[1732] Prompt Sentence Examples

[1733] "Please transcribe what was said on this call and assess the likelihood of fraud."

[1734] "Please rate the following text for fraud-related keywords: 'Payment required now'"

[1735] Please analyze the following conversation for possible fraud risk.

[1736] With the above configuration, this system significantly reduces the risk of elderly people falling victim to telephone fraud, enabling fast and effective fraud prevention in real time.

[1737] The flow of the specific processing in the application example 1 will be described with reference to FIG.

[1738] Step 1:

[1739] When a user answers a call and starts a conversation, the voice analysis module in the smart glasses is automatically activated. The voice data during the call is taken as input. The microphone in the smart glasses captures the voice and converts the voice data into text data in real time using a cloud speech recognition service (e.g., Google Cloud Speech-to-Text API). Text data is generated as output, which may include a message such as "Payment required now."

[1740] Step 2:

[1741] The text data transmission module of the smart glasses transmits the converted text data to the cloud server at regular intervals (less than 1 second). The text data generated earlier is used as input. The text data is sent to the cloud server as output.

[1742] Step 3:

[1743] The fraud detection module on the cloud server analyzes the received text data. Specifically, it uses natural language processing libraries (e.g., spaCy, NLTK) to identify fraud-related keywords and phrases. The input is the text data sent to the server. The output is the detection results of fraud-related keywords and phrases (e.g., "now," "pay," etc.).

[1744] Step 4:

[1745] The fraud score evaluation module on the cloud server evaluates the likelihood of fraud based on the analysis results and assigns a risk score. A generative AI model (e.g., BERT or GPT-4) is used to perform contextual analysis and pattern detection on the text data. The input is the output of the fraud detection module. The output is a fraud likelihood score, expressed as a number.

[1746] Step 5:

[1747] If the fraud score exceeds a threshold, the cloud server sets a warning flag. The input is the output of the fraud score evaluation module. The output is the setting state of the warning flag.

[1748] Step 6:

[1749] If the warning flag is set, the cloud server generates a warning message and sends it to the smart glasses. The input is the setting state of the warning flag. The output is the warning message (e.g., "Be careful, this may be a scam").

[1750] Step 7:

[1751] The smart glasses notify the user of the received warning message. The HUD display is used as the visual notification means, and the built-in speaker is used as the audio notification means. The input is the warning message sent from the cloud server. The output is to display the warning to the user both audio and visually.

[1752] Step 8:

[1753] When the user receives the warning message, they end the call and consult with their family to avoid the risk of fraud. The input is the warning message from the smart glasses, and the output is the user's action to take to avoid fraud.

[1754] Furthermore, an emotion engine that estimates the user's emotion may be further combined. That is, the identification processing unit 290 may estimate the user's emotion using the emotion identification model 59, and perform identification processing using the user's emotion.

[1755] The present invention provides a more accurate warning function by combining a system that converts voice data into text data in real time, evaluates the possibility of fraud, and notifies the user of a warning, with an emotion engine that recognizes the user's emotions. This system has the following configuration.

[1756] 1. Voice analysis module (terminal)

[1757] When a call starts, the device activates a voice analysis module and records the call in real time. The recorded voice data is then converted into text data on the spot using voice recognition technology.

[1758] 2. Text data transmission module (terminal)

[1759] The converted text data is sent to the server at regular intervals of less than one second, so that the data is sent in near real time.

[1760] 3. Fraud detection module (server)

[1761] The server analyzes the received text data, specifically matching fraud-related keywords and phrases with a list and scoring them based on their frequency and combinations.

[1762] 4. Fraud Score Evaluation Module (Server)

[1763] The server performs contextual analysis of the text data to detect speech patterns and other characteristics that are particularly characteristic of fraudulent activity. Based on this evaluation, a fraud likelihood score is calculated. If the score exceeds a set threshold, a warning flag is set.

[1764] 5. Emotion engine (server)

[1765] The server uses the voice data to recognize the user's emotional state: the emotion engine analyzes the tone, rate, and other voice characteristics to determine whether the user is experiencing emotions such as surprise, fear, or anxiety.

[1766] 6. Alert Adjustment Module (Server)

[1767] The server adjusts the content and notification method of the warning based on the analysis results of the emotion engine. For example, if the user is feeling strong anxiety or fear, the warning content can be made more detailed and clearer.

[1768] 7. Alert notification module (server and terminal)

[1769] If the warning flag is set, the server sends a warning message to the device, which receives the message and issues a voice or text warning to the user, either saying "Be careful, this may be a scam" in the case of a voice warning or a text message that pops up on the screen.

[1770] Specific examples

[1771] Example 1: Call fraud using emotion engine

[1772] 1. Users

[1773] The elderly user answers the phone and begins the conversation.

[1774] 2. Terminal

[1775] The device's voice analysis module records the call in real time and converts the audio "Payment required now" into text data.

[1776] 3. Terminal

[1777] The text data "Payment is required now" is sent to the server.

[1778] 4. Server

[1779] Analyzes text data to detect and score keywords such as "now" and "payment."

[1780] 5. Server

[1781] The system analyzes the urgent speech patterns characteristic of fraudulent activity and gives a high rating. If this exceeds a threshold, a warning flag is set.

[1782] 6. Server

[1783] At the same time, the emotion engine analyzes the voice data to recognize the user's emotional state, detecting when the user feels an urgent need and is anxious.

[1784] 7. Server

[1785] The warning content is adjusted based on the results of the emotion engine to generate stronger warning messages.

[1786] 8. Server

[1787] Sends tailored warning messages to the terminal.

[1788] 9. Terminal

[1789] The system will notify the user of the received warning message with a voice message, warning them, "Be careful, this may be a scam. Please tell a family member about this call immediately."

[1790] 10. Users

[1791] Users who receive the warning can end the call and consult with their family to prevent any harm.

[1792] In this way, the present invention can significantly reduce the risk of seniors becoming victims of fraud. The system operates in real time and issues appropriate warnings based on the user's emotional state, allowing for a fast and effective response.

[1793] The processing flow will be explained below.

[1794] Step 1:

[1795] The user answers the phone and starts the call. The device detects the start of the call and automatically starts the voice analysis module.

[1796] Step 2:

[1797] The device records the contents of the call in real time and instantly converts the acquired voice data into text data using voice recognition technology.

[1798] Step 3:

[1799] The terminal transmits the converted text data to the server at regular intervals (for example, every second).

[1800] Step 4:

[1801] The server receives the text data sent from the terminal and starts analysis using the text data analysis module.

[1802] Step 5:

[1803] The server detects keywords and phrases in the text data, including words related to fraud (e.g., "now," "payment," "legal action").

[1804] Step 6:

[1805] The server then scores the likelihood of fraud based on the frequency and combination of keywords detected, calculated according to pre-defined criteria.

[1806] Step 7:

[1807] The server performs contextual analysis to detect speech patterns that are particularly characteristic of fraudulent activity (forced language and emphasis on urgency).

[1808] Step 8:

[1809] The server evaluates the scoring results and sets a warning flag if the fraud likelihood score exceeds a set threshold.

[1810] Step 9:

[1811] The server sends the voice data to the emotion engine, which analyzes the tone, rate, and other voice characteristics of the user's voice.

[1812] Step 10:

[1813] The emotion engine on the server determines the user's emotional state, for example, analyzing whether the user is feeling surprise, fear, anxiety, etc.

[1814] Step 11:

[1815] The server adjusts the content of the warning and notification method based on the analysis results of the emotion engine. If the user is feeling strong anxiety or fear, the warning content will be more detailed and clearer.

[1816] Step 12:

[1817] The server sends a tailored warning message to the device, which includes information about the likely fraud and specific steps to take.

[1818] Step 13:

[1819] As soon as the device receives the warning notification from the server, it immediately issues a warning to the user, either by voice or text, saying, "Caution, this may be a scam. Please inform a family member immediately about this call."

[1820] Step 14:

[1821] The user receives a warning and becomes suspicious of the call. They end the call and consult with family or friends for safety. This process significantly reduces the risk of the user becoming a victim of fraud.

[1822] Example 2

[1823] Next, a description will be given of Example 2. In the following description, the data processing device 12 will be referred to as a "server" and the robot 414 will be referred to as a "terminal."

[1824] In today's world, fraud is becoming increasingly sophisticated, making the elderly especially vulnerable to fraud. Previous fraud prevention systems primarily relied on text data analysis and did not provide real-time warnings that took the user's emotional state into account. This made it difficult for users to quickly recognize fraud and take appropriate action. Furthermore, because conventional systems issued warnings uniformly, they often failed to respond appropriately to the sense of crisis and anxiety felt by users. This has led to a need for effective methods to prevent fraud before it happens.

[1825] The identification process by the identification processing unit 290 of the data processing device 12 in Example 2 is realized by the following means. In this invention, the server includes means for converting voice data into text data in real time, means for analyzing the text data and identifying keywords and phrases related to fraud, means for evaluating the possibility of fraud based on the analysis results and scoring the risk level, means for recognizing the user's emotional state from the voice analysis and adjusting the warning content based on the emotion, and means for transmitting the warning content to the terminal and notifying the user. This makes it possible to provide an appropriate warning in real time that takes into account the user's emotional state.

[1826] "Voice data" refers to sound waveforms such as those from a call or voice input that have been converted into electrical signals or digital data.

[1827] "Text data" refers to character string data generated as a result of analyzing and converting voice data.

[1828] "Analysis" is the process of dividing data for a specific purpose to reveal meaning and structure.

[1829] "Fraud-related keywords and phrases" refer to specific words and expressions that are likely to be used in fraud, and are used to predict the possibility of fraud by detecting them.

[1830] "Risk level" is a numerical or score that expresses the likelihood of a scam.

[1831] "Scoring" is the process of quantifying risk based on the results of data analysis.

[1832] A "threshold" is a numerical value or condition that triggers a specific action, and when exceeded, a warning is issued.

[1833] A "warning" is information that notifies the user of a potential danger and urges caution.

[1834] "Emotional state" refers to the user's psychological state analyzed from the tone, rate, and other characteristics of the voice.

[1835] "Adjustment" is the process of changing something to an appropriate form depending on the situation or conditions.

[1836] A "notification" is an action or means of informing a user of specific information.

[1837] The present invention provides a more accurate warning function by combining a system that converts voice data into text data in real time, evaluates the possibility of fraud, and notifies the user of a warning with an emotion engine that recognizes the user's emotions.

[1838] System Configuration

[1839] The system includes the following modules:

[1840] Voice analysis module (terminal)

[1841] The device activates a voice analysis module at the start of a call and records the call in real time. The recorded voice data is converted to text data using the Google Cloud Speech-to-Text API.

[1842] Text data transmission module (terminal)

[1843] The converted text data is sent to the server at regular intervals of less than one second using the WebSocket protocol in real time.

[1844] Fraud detection module (server)

[1845] The server analyzes the received text data using a natural language processing library (e.g., NLTK or SpaCy), matching fraud-related keywords and phrases with a list, and scoring them based on their frequency and combinations.

[1846] Fraud score evaluation module (server)

[1847] The server then performs contextual analysis of the text data to detect speech patterns and other characteristics of fraudulent activity. This information is used to calculate a fraud likelihood score. If this score exceeds a set threshold, a warning flag is set.

[1848] Emotion engine (server)

[1849] The server uses the voice data to recognize the user's emotional state, using IBM Watson Tone Analyzer and Microsoft Azure Emotion API to determine whether the user is experiencing emotions such as surprise, fear, or anxiety.

[1850] Alert Adjustment Module (Server)

[1851] The server adjusts the content and notification method of the warning based on the results of the emotion engine. For example, if the user is feeling strong anxiety or fear, the warning content will be more detailed and clearer.

[1852] Alert notification module (server and terminal)

[1853] The server sends the tailored warning message to the device, which receives it and issues a voice or text warning to the user, such as "Be careful, this may be a scam," or a text message that pops up on the screen.

[1854] Specific examples

[1855] Example 1: Call fraud using emotion engine

[1856] 1. Users

[1857] The elderly user answers the phone and begins the conversation.

[1858] 2. Terminal

[1859] The device's voice analysis module records the call in real time and instantly converts the audio "Payment required now" into text data using Google Cloud Speech-to-Text.

[1860] 3. Terminal

[1861] The converted text data "Payment required now" is sent to the server via the WebSocket protocol at intervals of less than one second.

[1862] 4. Server

[1863] The server uses NLTK to analyze the received text data, identifying fraud-related keywords such as "now" and "payment," and analyzing the frequency and context of these keywords.

[1864] 5. Server

[1865] It detects urgent speech patterns characteristic of fraudulent activity and generates a high fraud likelihood score. If this score exceeds a set threshold, it sets a warning flag.

[1866] 6. Server

[1867] The server uses an emotion engine (IBM Watson Tone Analyzer) to analyze the user's emotional state from the voice data. The analysis results indicate that the user feels urgency and is anxious.

[1868] 7. Server

[1869] The content of the warning message is adjusted based on the results of the emotion engine, for example, generating a detailed warning such as "This call may be fraudulent. Please report this call to a family member immediately."

[1870] 8. Server

[1871] Sends tailored warning messages to the terminal.

[1872] 9. Terminal

[1873] The device will then notify the user of the received warning message via voice: "Caution, this may be a scam. Please inform a family member immediately about this call."

[1874] 10. Users

[1875] The user will receive a warning, end the call, and immediately contact their family to prevent any harm.

[1876] In this way, the present invention can significantly reduce the risk of seniors becoming victims of fraud. The system operates in real time and issues appropriate warnings based on the user's emotional state, allowing for a fast and effective response.

[1877] Prompt Sentence Examples

[1878] "Design a system to detect emotions from voice data, assess the likelihood of fraud, and provide a warning. Use natural language processing and emotion recognition technologies for data analysis. Please also specify the specific hardware and software, data flow, and user interface (UI)."

[1879] The flow of the identification process in the second embodiment will be described with reference to FIG.

[1880] Step 1:

[1881] Voice recording and text conversion (device)

[1882] Input: A user receives a call and starts a conversation.

[1883] How it works: The device detects that a call has been initiated and activates the voice analysis module.

[1884] Data processing: Convert the audio data into text data in real time using the Google Cloud Speech-to-Text API.

[1885] Output: The converted text data is generated.

[1886] Step 2:

[1887] Sending text data (terminal)

[1888] Input: The text data generated in step 1.

[1889] Operation: The device sends the converted text data to the server at intervals of less than one second.

[1890] Data calculation: Send data in real time using the WebSocket protocol.

[1891] Output: The server receives text data in real time.

[1892] Step 3:

[1893] Text data analysis (server)

[1894] Input: The text data received in step 2.

[1895] How it works: The server parses the received text data using a natural language processing library (NLTK or SpaCy).

[1896] Data calculations: Fraud-related keywords and phrases are matched against the list and analyzed based on their frequency and combinations.

[1897] Output: A numerical score is generated representing the likelihood of fraud.

[1898] Step 4:

[1899] Fraud score evaluation (server)

[1900] Input: Analysis results (scores) generated in step 3.

[1901] How it works: The server also performs contextual analysis of the text data to detect speech patterns and patterns characteristic of fraudulent activity.

[1902] Data calculation: Calculate a fraud likelihood score and set a warning flag if this score exceeds a set threshold.

[1903] Output: Fraud probability score and warning flag status.

[1904] Step 5:

[1905] User emotion analysis (server)

[1906] Input: The audio data recorded in step 1.

[1907] How it works: The server uses IBM Watson Tone Analyzer and Microsoft Azure Emotion API to analyze the user's emotional state from voice data.

[1908] Data calculations: Analyze the tone, rate, and other voice characteristics of the voice to determine whether the user is surprised, scared, or anxious.

[1909] Output: Generates an analysis of the user's emotional state.

[1910] Step 6:

[1911] Adjustment of warning content (server)

[1912] Input: Warning flags from step 4 and sentiment analysis results from step 5.

[1913] How it works: The server tailors the warning based on the user's emotional state. For example, if the user is feeling strong anxiety or fear, the warning will be more detailed and clearer.

[1914] Data calculation: Optimize the content and format of warning messages based on the results of sentiment analysis.

[1915] Output: A tailored warning message is generated.

[1916] Step 7:

[1917] Notification of warning messages (server and terminal)

[1918] Input: The warning message adjusted in step 6.

[1919] Action: The server sends a tailored alert message to the device.

[1920] Data calculation: The device notifies the user of the received warning message by voice or text.

[1921] Output: A voice or text alert is given to the user.

[1922] These steps enable the system to detect potential fraud in real time and issue warnings that take into account the user's emotional state.

[1923] (Application example 2)

[1924] Next, a description will be given of Application Example 2. In the following description, the data processing device 12 will be referred to as a "server" and the robot 414 will be referred to as a "terminal."

[1925] In recent years, telephone fraud has been on the rise, with elderly people becoming victims in particular. Conventional fraud prevention measures only issue a warning to users when fraud is suspected. Therefore, even when users receive a warning, they are unable to properly determine how to respond, resulting in a high likelihood of fraud victimization. Furthermore, because warnings are issued uniformly with the same intensity without taking into account the user's emotional state, the effectiveness of the warnings is limited. There is a need to improve this situation and provide more effective fraud prevention measures.

[1926] The specific processing by the specific processing unit 290 of the data processing device 12 in the application example 2 is realized by the following means.

[1927] In this invention, the server includes means for converting voice data into text data in real time, means for analyzing the text data and identifying keywords and phrases related to fraud, means for assessing the likelihood of fraud and scoring the risk based on the analysis results, means for issuing a warning to the user when the risk exceeds a set threshold, means for analyzing the user's emotional state, means for adjusting the content of the warning based on the emotion analysis results, and means for issuing the adjusted warning to the user. This makes it possible to accurately assess the likelihood of fraud in real time based on the voice data and provide an appropriate warning that takes the user's emotional state into consideration.

[1928] "Voice data" means recorded information of a user's voice uttered over the telephone or other voice communication.

[1929] "Real-time" refers to near-instant processing with little actual time or delay.

[1930] "Text data" is voice data converted into character information.

[1931] Analysis is the process of examining data or information in detail to identify important elements or patterns within it.

[1932] "Fraud-related keywords and phrases" are specific words and expressions commonly used in fraudulent activities.

[1933] "Scoring" is the process of assigning points according to specific criteria based on the analysis results.

[1934] A "threshold" is a boundary value for determining a certain condition or standard.

[1935] A "warning" is a notification that notifies the user of an impending danger.

[1936] An "emotional state" is the psychological state that a user is feeling at a particular moment.

[1937] "Notification" is the act of conveying information or a message to a target person.

[1938] "Adjustment" refers to changing something to an optimal state for specific conditions or environments.

[1939] This invention is a system for protecting users from telephone fraud, and provides highly accurate warning functions by combining analysis of voice data and emotion analysis.

[1940] System Overview

[1941] The system mainly consists of three main components.

[1942] 1. Terminal

[1943] 2. Server

[1944] 3. Users

[1945] Hardware and software used

[1946] Hardware:

[1947] Smartphone: Used as a device, requires a microphone and internet connection.

[1948] software:

[1949] Python 3.x: The main programming language.

[1950] speech_recognition library: Uses Google's speech recognition API to convert voice data into text data.

[1951] The requests library: handles HTTP requests to the server.

[1952] Processing Details

[1953] Terminal handling

[1954] When a user initiates a phone call, the device records the voice data in real time and converts it into text data using the speech_recognition library, which is then sent to the server at regular intervals.

[1955] Server Processing

[1956] The server analyzes the audio data and assesses the likelihood of fraud using the following steps:

[1957] 1. Text Analysis: Analyzes incoming text data to identify keywords and phrases related to fraud.

[1958] 2. Scoring: Based on the analysis results, the possibility of fraud is assessed and a risk score is generated. If the score exceeds a set threshold, a warning is generated.

[1959] 3. Emotion analysis: Analyze the user's emotional state from the voice data, determining whether the user is feeling surprised or anxious.

[1960] 4. Adjustment of warning content: Based on the results of sentiment analysis, the warning content is adjusted appropriately.

[1961] Warning Notification

[1962] A tailored warning message is sent to the user's device. For example, if the warning is delivered via voice, the content and intensity of the warning can be adjusted according to the user's emotional state. This is expected to encourage the user to take appropriate action.

[1963] Specific examples

[1964] For example, consider the case of an elderly person receiving a suspected fraudulent phone call. If the user hears phrases such as "Payment is required now," the system immediately transcribes the speech into text and analyzes it. If it detects patterns indicative of fraudulent activity, the system takes into account the user's emotional state and generates a strong warning message. The warning could include specific advice such as "Be careful, this may be a scam. Please report this call to a family member immediately."

[1965] Example prompts for generative AI models

[1966] Your system should take voice data from a user's phone call as input and convert it into text data in real time. It should then assess the likelihood of fraud (by scoring it based on the frequency of fraud-related keywords and phrases) and generate a warning message if necessary. It should also analyze the user's emotional state and adjust the strength and content of the warning. The emotion engine uses the tone, rate, and other voice characteristics to determine whether the user is feeling anxious or scared. It should generate a warning message based on the results of the text and emotion analysis and notify the user.

[1967] The flow of the specific processing in the application example 2 will be described with reference to FIG.

[1968] Step 1:

[1969] A user initiates a phone call.

[1970] Input: User's call audio

[1971] Output: Call audio data

[1972] Specific operation: A user initiates a call and voice communication begins over the phone.

[1973] Step 2:

[1974] The device records the call audio in real time.

[1975] Input: Call audio data

[1976] Output: Recorded audio file

[1977] Specific operation: The device's microphone captures the user's call voice and records the voice data.

[1978] Step 3:

[1979] The device converts recorded voice data into text data in real time.

[1980] Input: Recorded audio file

[1981] Output: Text data

[1982] Specific operation: Using the speech_recognition library installed on the device, the recorded voice data is converted into text data.

[1983] Step 4:

[1984] The terminal transmits the converted text data to the server at regular intervals.

[1985] Input: Text data

[1986] Output: Send text data to the server

[1987] Specific operation: Sends text data to the server using an HTTP request, for example, using the requests library.

[1988] Step 5:

[1989] The server analyzes the received text data to identify keywords and phrases related to fraud.

[1990] Input: Text data sent to the server

[1991] Output: Analysis results (list of fraud-related keywords and phrases)

[1992] How it works: An analytics module on the server scans the text data to identify keywords and phrases related to fraud, matching them against a predefined list of fraud keywords.

[1993] Step 6:

[1994] The server evaluates the likelihood of fraud and assigns a score based on the results of analyzing the text data.

[1995] Input: Analysis results

[1996] Output: Risk score

[1997] Specific operation: The server's scoring module evaluates the likelihood of fraud based on the frequency of fraud keywords and the context of the text, and calculates a risk score.

[1998] Step 7:

[1999] The server generates a warning message if the risk score exceeds a threshold.

[2000] Input: Risk score

[2001] Output: Warning message

[2002] Specific behavior: If the score exceeds a preset threshold, an algorithm is executed that generates a warning message.

[2003] Step 8:

[2004] The server analyzes the user's emotional state from the voice data.

[2005] Input: Audio data

[2006] Output: Emotion analysis results

[2007] Specific behavior: The emotion engine analyzes the tone, rate, and other voice characteristics of the voice to assess the user's emotional state.

[2008] Step 9:

[2009] The server adjusts the warning content based on the results of sentiment analysis.

[2010] Input: Sentiment analysis results, warning message

[2011] Output: Adjusted warning message

[2012] Specific operation: Using the results of the emotion engine, an algorithm is run that adjusts the content and intensity of warning messages.

[2013] Step 10:

[2014] The server sends the adjusted warning message to the terminal, and the terminal notifies the user of the warning.

[2015] Input: Adjusted warning message

[2016] Output: A warning notice to the user

[2017] Specific operation: The server sends a tailored warning message to the terminal, which then notifies the user of the message. The warning is notified as voice or text.

[2018] Each step is performed sequentially, protecting users from the risk of fraud in real time.

[2019] The specific processing unit 290 transmits the result of the specific processing to the robot 414. In the robot 414, the control unit 46A causes the speaker 240 and the control target 443 to output the result of the specific processing. The microphone 238 acquires voice indicating a user input regarding the result of the specific processing. The control unit 46A transmits voice data indicating the user input acquired by the microphone 238 to the data processing device 12. In the data processing device 12, the specific processing unit 290 acquires the voice data.

[2020] The data generation model 58 is a so-called generative AI (Artificial Intelligence). An example of the data generation model 58 is ChatGPT (Internet Search<URL: https: / / openai.com / blog / chatgpt> ), Gemini (Internet search <url: https: gemini.google.com ?hl="ja">) and other generation AIs. The data generation model 58 is obtained by performing deep learning on a neural network. A prompt including an instruction is input to the data generation model 58, and inference data such as voice data indicating voice, text data indicating text, and image data indicating an image is also input. The data generation model 58 performs inference on the input inference data in accordance with the instruction indicated by the prompt, and outputs the inference result in a data format such as voice data and text data. Here, inference refers to, for example, analysis, classification, prediction, and / or summarization.

[2021] In the above embodiment, an example was given in which the specific processing is performed by the data processing device 12, but the technology of the present disclosure is not limited to this, and the specific processing may be performed by the robot 414.

[2022] The emotion identification model 59 as an emotion engine may determine the user's emotion according to a specific mapping. Specifically, the emotion identification model 59 may determine the user's emotion according to an emotion map (see FIG. 9), which is a specific mapping. Similarly, the emotion identification model 59 may determine the robot's emotion, and the identification processing unit 290 may perform identification processing using the robot's emotion.

[2023] FIG. 9 is a diagram illustrating an emotion map 400 on which multiple emotions are mapped. In the emotion map 400, emotions are arranged in concentric circles radiating from the center. Emotions closer to the center of the concentric circles are more primitive. Emotions representing states and actions arising from a state of mind are arranged on the outer edges of the concentric circles. The concept of emotion includes both affect and mental states. Emotions generally generated from reactions occurring in the brain are arranged on the left side of the concentric circles. Emotions generally induced by situational judgment are arranged on the right side of the concentric circles. Emotions generally generated from reactions occurring in the brain and induced by situational judgment are arranged on the upper and lower sides of the concentric circles. Furthermore, the emotion of "pleasure" is arranged on the upper side of the concentric circles, and the emotion of "discomfort" is arranged on the lower side. In this way, in the emotion map 400, multiple emotions are mapped based on the structure by which emotions are generated, and emotions that tend to occur simultaneously are mapped close to each other.

[2024] These emotions are distributed in the 3 o'clock direction on emotion map 400, and typically fluctuate between relief and anxiety. In the right half of emotion map 400, situational awareness dominates over internal sensations, resulting in a sense of calm.

[2025] The inside of emotion map 400 represents what is going on in the mind, and the outside of emotion map 400 represents behavior, so the further you go outside emotion map 400, the more visible the emotions become (the more they are expressed in behavior).

[2026] Human emotions are based on various balances, such as posture and blood sugar levels. When these balances deviate from the ideal, a state of discomfort is indicated, and when they approach the ideal, a state of pleasure is indicated. Emotions can also be created for robots, automobiles, and motorcycles, based on various balances, such as posture and remaining battery life. When these balances deviate from the ideal, a state of discomfort is indicated, and when they approach the ideal, a state of pleasure is indicated. An emotion map can be generated, for example, based on Dr. Mitsuyoshi's emotion map (Research on Voice Emotion Recognition and Emotional Brain Physiological Signal Analysis Systems, Tokushima University, Doctoral Dissertation: https: / / ci.nii.ac.jp / naid / 500000375379). The left half of the emotion map lists emotions belonging to the "reaction" domain, where sensation is dominant. The right half of the emotion map lists emotions belonging to the "situation" domain, where situational awareness is dominant.

[2027] The emotion map defines two emotions that promote learning. One is a negative emotion on the situation side, around the middle of "repentance" or "reflection." In other words, this occurs when the robot experiences negative emotions such as "I never want to feel this way again" or "I don't want to be scolded again." The other is a positive emotion on the response side, around "desire." In other words, this occurs when the robot experiences positive feelings such as "I want more" or "I want to know more."

[2028] The emotion identification model 59 inputs user input into a pre-trained neural network, obtains emotion values ​​indicating each emotion shown in the emotion map 400, and determines the user's emotion. This neural network is pre-trained based on multiple pieces of training data that are combinations of user input and emotion values ​​indicating each emotion shown in the emotion map 400. Furthermore, this neural network is trained so that emotions that are located close to each other have similar values, as in the emotion map 900 shown in FIG. 10. FIG. 10 shows an example in which multiple emotions, "relieved," "calm," and "reassuring," have similar emotion values.

[2029] The system according to the present disclosure has been described above mainly with respect to the functions of the data processing device 12, but the system according to the present disclosure is not necessarily implemented on a server. The system according to the present disclosure may be implemented as a general information processing system. The present disclosure may be implemented, for example, as a software program running on a personal computer or an application running on a smartphone, etc. The method according to the present disclosure may be provided to users in the form of SaaS (Software as a Service).

[2030] In the above embodiment, an example was given in which the specific processing is performed by one computer 22, but the technology of the present disclosure is not limited to this, and the specific processing may be distributed and performed by a plurality of computers including the computer 22. For example, the data generation model 58 may be provided in an external device of the data processing device 12, and data may be generated in the external device in accordance with input data.

[2031] In the above embodiment, an example in which the specific processing program 56 is stored in the storage 32 has been described, but the technology of the present disclosure is not limited to this. For example, the specific processing program 56 may be stored in a portable, computer-readable, non-transitory storage medium such as a USB (Universal Serial Bus) memory. The specific processing program 56 stored in the non-transitory storage medium is installed in the computer 22 of the data processing device 12. The processor 28 executes the specific processing in accordance with the specific processing program 56.

[2032] Alternatively, the specific processing program 56 may be stored in a storage device such as a server connected to the data processing device 12 via the network 54, and the specific processing program 56 may be downloaded and installed on the computer 22 in response to a request from the data processing device 12.

[2033] It is not necessary to store all of the specific processing program 56 in a storage device such as a server connected to the data processing device 12 via the network 54, or to store all of the specific processing program 56 in the storage 32; only a portion of the specific processing program 56 may be stored.

[2034] The hardware resource for executing a specific process can be any of the following processors: An example of a processor is a CPU, which is a general-purpose processor that functions as a hardware resource for executing a specific process by executing software, i.e., a program. Another example of a processor is a dedicated electrical circuit, such as an FPGA (Field-Programmable Gate Array), a PLD (Programmable Logic Device), or an ASIC (Application Specific Integrated Circuit), which is a processor with a circuit configuration designed specifically for executing a specific process. Each processor has built-in or connected memory, and each processor uses the memory to execute the specific process.

[2035] The hardware resource that executes the specific processing may be configured with one of these various processors, or may be configured with a combination of two or more processors of the same or different types (for example, a combination of multiple FPGAs, or a combination of a CPU and an FPGA). Also, the hardware resource that executes the specific processing may be a single processor.

[2036] As an example of a system configured with a single processor, first, one processor is configured by combining one or more CPUs and software, and this processor functions as a hardware resource that executes a specific process. Second, there is a system that uses a processor that realizes the functions of an entire system including multiple hardware resources that execute a specific process on a single IC chip, as typified by SoC (System-on-a-chip). In this way, a specific process is realized using one or more of the above-mentioned various processors as hardware resources.

[2037] Furthermore, the hardware structure of these various processors can be, more specifically, an electric circuit that combines circuit elements such as semiconductor devices. The specific processing described above is merely an example. Therefore, it goes without saying that unnecessary steps may be deleted, new steps may be added, or the processing order may be rearranged, without departing from the spirit of the invention.

[2038] The above-described description and illustrations are a detailed explanation of the parts related to the technology of the present disclosure and are merely an example of the technology of the present disclosure. For example, the above description of the configuration, functions, actions, and effects is an explanation of an example of the configuration, functions, actions, and effects of the parts related to the technology of the present disclosure. Therefore, it goes without saying that unnecessary parts may be deleted, new elements may be added, or replacements may be made to the above-described description and illustrations within the scope of the gist of the technology of the present disclosure. Furthermore, to avoid confusion and facilitate understanding of the parts related to the technology of the present disclosure, the above-described description and illustrations omit explanations of common technical knowledge that do not require particular explanation to enable the implementation of the technology of the present disclosure.

[2039] All publications, patent applications, and technical standards mentioned in this specification are herein incorporated by reference to the same extent as if each individual publication, patent application, or technical standard was specifically and individually indicated to be incorporated by reference.

[2040] The following is further disclosed regarding the above embodiment.

[2041] (Claim 1)

[2042] a means for converting voice data into text data in real time;

[2043] means for analyzing the text data to identify keywords and phrases associated with fraud;

[2044] means for assessing the possibility of fraud based on the analysis results and scoring the degree of risk;

[2045] The system further includes a means for notifying a user of a warning when the risk level exceeds a set threshold.

[2046] (Claim 2)

[2047] 10. The system of claim 1, further comprising means for recording said audio data in real time.

[2048] (Claim 3)

[2049] 10. The system of claim 1, further comprising means for notifying the user of the warning by voice.

[2050] "Example 1"

[2051] (Claim 1)

[2052] a means for converting voice data into text data in real time;

[2053] means for analyzing the text data to identify keywords and phrases associated with fraud;

[2054] means for assessing the possibility of fraud based on the analysis results and scoring the degree of risk;

[2055] means for notifying a user of a warning when the risk level exceeds a set threshold;

[2056] means for transmitting the text data to a server at predetermined intervals;

[2057] means for detecting fraud patterns based on said fraud-related keywords and phrases and calculating a fraud likelihood score;

[2058] The system includes means for notifying the user of said warning by voice or text.

[2059] (Claim 2)

[2060] 10. The system of claim 1, further comprising means for recording said audio data in real time.

[2061] (Claim 3)

[2062] 10. The system of claim 1, further comprising means for notifying the user of the warning by voice.

[2063] "Application Example 1"

[2064] (Claim 1)

[2065] a means for converting voice data into text data in real time;

[2066] means for analyzing the text data to identify keywords and phrases associated with fraud;

[2067] means for assessing the possibility of fraud based on the analysis results and scoring the degree of risk;

[2068] means for notifying a user of a warning when the risk level exceeds a set threshold;

[2069] means for notifying the user of the warning using a visual display means and an audio notification means of the smart glasses;

[2070] A system that includes a means for using a generative AI model to perform contextual analysis of text data and detect patterns specific to fraudulent activity.

[2071] (Claim 2)

[2072] 10. The system of claim 1, further comprising means for recording said audio data in real time.

[2073] (Claim 3)

[2074] The system of claim 1, further comprising means for visually notifying the user of the notification using a head-mounted display.

[2075] "Example 2: Combining Emotion Engines"

[2076] (Claim 1)

[2077] a means for converting voice data into text data in real time;

[2078] means for analyzing the text data to identify keywords and phrases associated with fraud;

[2079] means for assessing the possibility of fraud based on the analysis results and scoring the degree of risk;

[2080] means for notifying a user of a warning when the risk level exceeds a set threshold;

[2081] means for recognizing the emotional state of a user from voice analysis and adjusting the content of a warning based on the emotional state;

[2082] and means for transmitting the content of the warning to a terminal and notifying the user.

[2083] (Claim 2)

[2084] 10. The system of claim 1, further comprising means for recording said audio data in real time.

[2085] (Claim 3)

[2086] 10. The system of claim 1, further comprising means for notifying the user of the warning by voice.

[2087] "Application example 2 when combining emotion engines"

[2088] (Claim 1)

[2089] a means for converting voice data into text data in real time;

[2090] means for analyzing the text data to identify keywords and phrases associated with fraud;

[2091] means for assessing the possibility of fraud based on the analysis results and scoring the degree of risk;

[2092] means for notifying a user of a warning when the risk level exceeds a set threshold;

[2093] means for analyzing the emotional state of a user;

[2094] a means for adjusting the warning content based on the sentiment analysis results;

[2095] A system including means for notifying a user of a tailored alert.

[2096] (Claim 2)

[2097] 10. The system of claim 1, further comprising means for recording said audio data in real time.

[2098] (Claim 3)

[2099] 10. The system of claim 1, further comprising means for notifying the user of the warning by voice. [Explanation of symbols]

[2100] 10, 210, 310, 410 Data Processing Systems 12 Data Processing Device 14 Smart Devices 214 Smart Glasses 314 Headset-type terminal 414 Robot< / url:> < / url:> < / url:> < / url:>

Claims

1. a means for converting voice data into text data in real time; means for analyzing the text data to identify keywords and phrases associated with fraud; means for assessing the possibility of fraud based on the analysis results and scoring the degree of risk; The system further includes a means for notifying a user of a warning when the risk level exceeds a set threshold.

2. 2. The system of claim 1, further comprising means for recording said audio data in real time.

3. 2. The system of claim 1, further comprising means for notifying the user of said warning by voice.

Citation Information

Patent Citations

  • Persona chatbot control method and system

    JP2022180282A