System

A real-time fraud detection system monitors user communications, preprocesses data, analyzes it for fraud indicators, and issues warnings, enhancing detection accuracy through user feedback to prevent fraud.

JP2026014290APending Publication Date: 2026-01-29SOFTBANK GROUP CORP
View PDF 1 Cites 0 Cited by

Patent Information

Application Number
JP2024115287
Authority / Receiving Office
JP · JP
Patent Type
Applications
Current Assignee / Owner
Filing Date
2024-07-18
Publication Date
2026-01-29

AI Technical Summary

Technical Problem

Existing systems fail to effectively detect sophisticated fraud schemes, particularly targeting the elderly and those unfamiliar with the internet, leading to financial and psychological damage, as they cannot provide real-time detection and warnings.

Method used

A system that monitors user communications in real-time, preprocesses data, analyzes it using natural language processing, scores fraud likelihood, generates warnings, and improves through user feedback to enhance detection accuracy.

Benefits of technology

The system effectively detects and warns users of potential fraud in real-time, reducing the risk of financial and psychological harm by continuously improving its fraud detection capabilities.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure 2026014290000001_ABST
    Figure 2026014290000001_ABST
Patent Text Reader

Abstract

A system is provided.SOLUTION: A system comprising: means for monitoring a user's communications and capturing data in real-time; means for transmitting the captured data to a server; means for analyzing the data received by the server and detecting possible fraud; means for generating an alert and notifying the user based on the detected possible fraud; and means for collecting feedback information and refining the analyzing means based on the feedback information.SELECTED DRAWING: Figure 1
Need to check novelty before this filing date? Find Prior Art

Description

[Technical Field]

[0001] The technology of the present disclosure relates to a system. [Background technology]

[0002] Patent document 1 discloses a persona chatbot control method performed by at least one processor, the method including the steps of receiving a user utterance, adding the user utterance to a prompt including an instruction sentence related to a description of the chatbot character, encoding the prompt, and inputting the encoded prompt into a language model to generate a chatbot utterance in response to the user utterance. [Prior art documents] [Patent documents]

[0003] [Patent Document 1] Japanese Patent Publication No. 2022-180282 Summary of the Invention [Problem to be solved by the invention]

[0004] In recent years, many people have fallen victim to increasingly sophisticated special fraud schemes. Many users, including the elderly, are unable to detect these scams and suffer financial and psychological damage. New fraud schemes have emerged that cannot be fully addressed by existing prevention methods, and an effective prevention system is needed. The objective of this invention is to provide a system that enables real-time detection of special frauds and promptly warns users, thereby preventing fraud damage before it occurs. [Means for solving the problem]

[0005] The present invention provides a system that includes a means for monitoring user communications and capturing data in real time, a means for transmitting the captured data to a server, a means for analyzing the data received by the server to detect possible fraud, a means for generating a warning based on the detected possible fraud and notifying the user, and a means for collecting feedback information and improving the analysis means based on the feedback information. In the case of voice data, the system also includes a means for converting the voice data into text. The analysis means analyzes the text data using natural language processing technology, and the warning generation means performs fraud scoring based on the analyzed data and generates a warning based on the score. The system also includes a means for continuously updating the machine learning model based on the feedback information. This configuration provides a system that effectively detects special frauds in real time, issues warnings, and protects users.

[0006] A "user" is an individual person or end device owner who uses the system.

[0007] "Communication content" refers to information such as voice data, text messages, and emails sent and received by users.

[0008] "Real-time" means that data is captured and analyzed nearly simultaneously, with very little delay.

[0009] A "data capture device" is a combination of hardware and software for capturing and storing the contents of calls and messages in digital form.

[0010] "Preprocessed data" refers to data that has undergone processing such as noise removal and format standardization.

[0011] A "server" is a remote computer system that performs multiple processes such as analysis and alert generation.

[0012] "Analysis tools" are software used to analyze captured data and detect signs of fraud.

[0013] "Fraud likelihood" is the predicted probability or score of fraud.

[0014] "Means for generating warnings and notifying users" means a system for generating warning messages in cases where there is a high likelihood of fraud, and transmitting and displaying such messages on the user's device.

[0015] "Feedback information" is information about the correctness or incorrectness of fraud detection results provided by a user.

[0016] "Speech recognition technology" is a technology for converting voice data into text data.

[0017] "Natural language processing technology" is a technology for analyzing text data and understanding its context and meaning.

[0018] A "scoring algorithm" is an algorithm that quantifies the likelihood of fraud based on analyzed data.

[0019] A "machine learning model" is a predictive model that learns from large amounts of data and can be applied to new data.

[0020] "Updating" is the process of using feedback information to improve and update the machine learning model. [Brief explanation of the drawings]

[0021] [Figure 1] 1 is a conceptual diagram showing an example of the configuration of a data processing system according to a first embodiment. [Figure 2] 1 is a conceptual diagram showing an example of main functions of a data processing device and a smart device according to a first embodiment. [Figure 3] FIG. 10 is a conceptual diagram showing an example of the configuration of a data processing system according to a second embodiment. [Figure 4] FIG. 10 is a conceptual diagram showing an example of main functions of a data processing device and smart glasses according to a second embodiment. [Figure 5] FIG. 10 is a conceptual diagram showing an example of the configuration of a data processing system according to a third embodiment. [Figure 6] FIG. 11 is a conceptual diagram showing an example of main functions of a data processing device and a headset-type terminal according to a third embodiment. [Figure 7] FIG. 10 is a conceptual diagram showing an example of the configuration of a data processing system according to a fourth embodiment. [Figure 8] FIG. 10 is a conceptual diagram showing an example of main functions of a data processing device and a robot according to a fourth embodiment. [Figure 9] 1 shows an emotion map onto which multiple emotions are mapped. [Figure 10] 1 shows an emotion map onto which multiple emotions are mapped. [Figure 11] FIG. 3 is a sequence diagram showing a processing flow of the data processing system according to the first embodiment. [Figure 12] FIG. 10 is a sequence diagram showing the flow of processing in the data processing system in Application Example 1. [Figure 13] FIG. 10 is a sequence diagram showing the flow of processing in the data processing system according to the second embodiment when an emotion engine is combined. [Figure 14] FIG. 10 is a sequence diagram showing the flow of processing in the data processing system in Application Example 2 when an emotion engine is combined. DETAILED DESCRIPTION OF THE INVENTION

[0022] An example of an embodiment of a system according to the technology of the present disclosure will be described below with reference to the accompanying drawings.

[0023] First, the terms used in the following description will be explained.

[0024] In the following embodiments, a coded processor (hereinafter simply referred to as a "processor") may be a single arithmetic device or a combination of multiple arithmetic devices. Furthermore, a processor may be a single type of arithmetic device or a combination of multiple types of arithmetic devices. Examples of arithmetic devices include a CPU (Central Processing Unit), a GPU (Graphics Processing Unit), a GPGPU (General-Purpose computing on Graphics Processing Units), and an APU (Accelerated Processing Unit).

[0025] In the following embodiments, a coded RAM (Random Access Memory) is a memory in which information is temporarily stored and is used as a working memory by a processor.

[0026] In the following embodiments, the coded storage is one or more non-volatile storage devices that store various programs, various parameters, etc. Examples of non-volatile storage devices include flash memory (SSD (Solid State Drive)), magnetic disks (e.g., hard disks), and magnetic tapes.

[0027] In the following embodiments, a communication I / F (Interface) with a symbol is an interface including a communication processor, an antenna, etc. The communication I / F controls communication between multiple computers. Examples of communication standards applied to the communication I / F include wireless communication standards including 5G (5th Generation Mobile Communication System), Wi-Fi (registered trademark), Bluetooth (registered trademark), etc.

[0028] In the following embodiments, "A and / or B" is synonymous with "at least one of A and B." In other words, "A and / or B" means that it may be only A, only B, or a combination of A and B. Furthermore, in this specification, the same concept as "A and / or B" is also applied when three or more things are expressed connected by "and / or."

[0029] [First embodiment]

[0030] FIG. 1 shows an example of the configuration of a data processing system 10 according to the first embodiment.

[0031] 1, a data processing system 10 includes a data processing device 12 and a smart device 14. An example of the data processing device 12 is a server.

[0032] The data processing device 12 includes a computer 22, a database 24, and a communication I / F 26. The computer 22 is an example of a "computer" according to the technology of the present disclosure. The computer 22 includes a processor 28, a RAM 30, and a storage 32. The processor 28, the RAM 30, and the storage 32 are connected to a bus 34. The database 24 and the communication I / F 26 are also connected to the bus 34. The communication I / F 26 is connected to a network 54. Examples of the network 54 include a WAN (Wide Area Network) and / or a LAN (Local Area Network).

[0033] The smart device 14 includes a computer 36, a reception device 38, an output device 40, a camera 42, and a communication I / F 44. The computer 36 includes a processor 46, a RAM 48, and a storage 50. The processor 46, the RAM 48, and the storage 50 are connected to a bus 52. The reception device 38, the output device 40, and the camera 42 are also connected to the bus 52.

[0034] The reception device 38 includes a touch panel 38A, a microphone 38B, and the like, and receives user input. The touch panel 38A detects contact with an indicator (for example, a pen or a finger) to receive user input by the touch of the indicator. The microphone 38B detects the user's voice to receive user input by voice. The control unit 46A transmits data indicating the user input received by the touch panel 38A and the microphone 38B to the data processing device 12. In the data processing device 12, the specific processing unit 290 acquires the data indicating the user input.

[0035] The output device 40 includes a display 40A and a speaker 40B, and presents data to the user 20 by outputting the data in a form of expression that the user 20 can perceive (for example, audio and / or text). The display 40A displays visible information such as text and images in accordance with instructions from the processor 46. The speaker 40B outputs audio in accordance with instructions from the processor 46. The camera 42 is a compact digital camera equipped with an optical system including a lens, aperture, and shutter, and an imaging element such as a CMOS (Complementary Metal-Oxide-Semiconductor) image sensor or a CCD (Charge Coupled Device) image sensor.

[0036] The communication I / F 44 is connected to a network 54. The communication I / Fs 44 and 26 control the exchange of various information between the processor 46 and the processor 28 via the network 54.

[0037] FIG. 2 shows an example of the main functions of the data processing device 12 and the smart device 14.

[0038] 2, in the data processing device 12, a specific process is performed by the processor 28. A specific processing program 56 is stored in the storage 32. The specific processing program 56 is an example of a "program" according to the technology of the present disclosure. The processor 28 reads the specific processing program 56 from the storage 32 and executes the read specific processing program 56 on the RAM 30. The specific process is realized by the processor 28 operating as a specific processing unit 290 in accordance with the specific processing program 56 executed on the RAM 30.

[0039] The storage 32 stores a data generation model 58 and an emotion identification model 59. The data generation model 58 and the emotion identification model 59 are used by the identification processing unit 290.

[0040] In the smart device 14, the processor 46 performs the reception output process. The storage 50 stores a reception output program 60. The reception output program 60 is used in conjunction with the specific processing program 56 by the data processing system 10. The processor 46 reads the reception output program 60 from the storage 50 and executes the read reception output program 60 on the RAM 48. The reception output process is realized by the processor 46 operating as the control unit 46A in accordance with the reception output program 60 executed on the RAM 48.

[0041] Next, a description will be given of the specific processing performed by the specific processing unit 290 of the data processing device 12. In the following description, the data processing device 12 will be referred to as a "server" and the smart device 14 will be referred to as a "terminal."

[0042] This invention is a system that uses the power of AI to eradicate special frauds. It monitors users' communications in real time, detects signs of fraud, and issues warnings to effectively prevent fraud damage. This system is configured around the user's terminal and server.

[0043] Overall system overview

[0044] 1. Monitoring of outgoing information

[0045] Terminal: Monitors user calls and messages in real time. For calls, the microphone is used to capture audio data, and for messages, the text data is directly acquired.

[0046] 2. Data preprocessing and transmission

[0047] Terminal: For audio data, preprocessing such as noise removal and sampling rate adjustment is performed. The preprocessed data is encrypted and sent to the server.

[0048] 3. Data Analysis

[0049] Server: Analyzes the received data. In the case of voice data, it is first converted into text data using speech recognition technology. Next, it analyzes the text data using natural language processing technology to detect signs of fraud.

[0050] 4. Fraud detection and scoring

[0051] Server: If any fraud indications are detected, a scoring algorithm is used to calculate the likelihood of fraud. Based on this score, the likelihood of fraud is assessed.

[0052] 5. Generating and Sending Alerts

[0053] Server: Generates a warning message when the likelihood of fraud exceeds a certain threshold, and sends the generated warning message to the user's device to notify them in real time.

[0054] 6. Feedback and Learning

[0055] User: Provide feedback on whether it is actually a scam.

[0056] Server: Use this feedback to improve the machine learning model and make the next detection more accurate.

[0057] Specific examples

[0058] Example 1: Detecting fraudulent calls

[0059] User: An elderly person receives a call from someone claiming to be their "son."

[0060] Terminal: Captures the call in real time and sends the audio data to the server.

[0061] Server: Converts voice data into text and uses natural language processing technology to detect signs of fraud, such as "I need money" or "Help me."

[0062] Server: Scores the likelihood of fraud and generates a warning message if it indicates a high probability.

[0063] Devices: Display a warning message on the senior's device to notify them of potential fraud.

[0064] User: Realizes the call is a scam and hangs up.

[0065] User: Provide feedback to the system that it was a scam.

[0066] Server: Improves machine learning models based on feedback.

[0067] Example 2: Message fraud detection

[0068] User: A young person receives a suspicious email.

[0069] Terminal: Captures email text in real time and sends it to the server.

[0070] Server: Analyzes text data using natural language processing technology to detect signs of fraud, such as "urgent" or "click here."

[0071] Server: A scoring algorithm assesses the likelihood of fraud and generates a warning message if it indicates a high probability.

[0072] Devices: Display warning messages on young people's devices to inform them of potential scams.

[0073] User: Ignore the email to prevent unauthorized access.

[0074] User: Provide feedback that the email was fraudulent.

[0075] Server: Improve machine learning models based on feedback.

[0076] In this way, this system can effectively prevent fraud damage through real-time detection and warning of special frauds.

[0077] The processing flow will be explained below.

[0078] Step 1:

[0079] Users: Make and receive calls and messages, some of which may be potentially fraudulent.

[0080] Step 2:

[0081] Terminal: Captures the user's call content as audio data in real time, and also captures the message content directly in the case of text messages.

[0082] Step 3:

[0083] Terminal: For audio data, preprocessing is performed, such as noise removal and unifying the sampling rate. The preprocessed data is prepared for transmission to the server.

[0084] Step 4:

[0085] Terminal: Encrypts the pre-processed data and sends it to the server using a secure communication protocol.

[0086] Step 5:

[0087] Server: Receives data sent from the device, including voice and text data.

[0088] Step 6:

[0089] Server: Converts the received voice data into text data using voice recognition technology. A voice recognition engine is used for this purpose.

[0090] Step 7:

[0091] Server: Analyzes text data using natural language processing (NLP) techniques, such as tokenizing words, tagging parts of speech, and grammar analysis.

[0092] Step 8:

[0093] Server: Detects signs of fraud from the analyzed data, using pre-trained AI models to identify phrases and contexts associated with fraud.

[0094] Step 9:

[0095] Server: Based on the detected fraud indicators, a scoring algorithm is used to assess the likelihood of fraud, using techniques such as logistic regression or support vector machines (SVM).

[0096] Step 10:

[0097] Server: If the score exceeds a set threshold, generate a warning message containing information about the potential for fraud and what to do about it.

[0098] Step 11:

[0099] Server: Sends the generated warning message to the user's terminal.

[0100] Step 12:

[0101] Device: Display an alert message on the user's device, using a notification sound or vibration to get the user's immediate attention.

[0102] Step 13:

[0103] User: Review the warning message and decide how to respond, including whether it is actually a scam.

[0104] Step 14:

[0105] User: Provide feedback to the system as to whether it was a scam or not.

[0106] Step 15:

[0107] Device: Sends user feedback to the server.

[0108] Step 16:

[0109] Server: Receives feedback information and uses it to refine the machine learning model, leveraging a feedback loop to continuously improve the accuracy of the analysis module.

[0110] Through this process, the system can detect signs of special fraud in real time and effectively warn users.

[0111] Example 1

[0112] Next, a description will be given of Example 1. In the following description, the data processing device 12 will be referred to as a "server" and the smart device 14 will be referred to as a "terminal."

[0113] Recently, special frauds have become increasingly sophisticated, increasing the risk of many people becoming victims of fraud. Elderly people and users unfamiliar with the Internet are particularly vulnerable to fraud. Current systems have difficulty detecting fraudulent activities in advance and issuing effective warnings, so effective prevention measures are needed.

[0114] The specific processing by the specific processing unit 290 of the data processing device 12 in the first embodiment is realized by the following means.

[0115] In this invention, the server includes means for monitoring user communications in real time and capturing the data as voice or text data, preprocessing means for performing noise reduction and sampling rate adjustment on the captured data, means for encrypting the preprocessed data and transmitting it to the server, means for analyzing the data received by the server and detecting the possibility of fraud, means for scoring the possibility of fraud based on the analyzed data, means for generating a warning message and notifying the user when the possibility of fraud exceeds a certain threshold, and means for collecting feedback information from users and improving the analysis means based on the feedback information, thereby making it possible to detect fraudulent activity in real time and quickly issue a warning to the user.

[0116] "User" refers to any individual or organization that uses this system.

[0117] "Communication content" refers to data such as voice calls and text messages sent by users.

[0118] "Real-time" refers to processing data almost as soon as the communication content occurs.

[0119] "Means for capturing data" refers to hardware or software for acquiring the contents of user communications.

[0120] "Noise reduction" refers to the process of reducing unnecessary noise from audio data.

[0121] "Sampling rate adjustment" refers to digital processing to convert audio data into a format that is optimal for analysis.

[0122] "Preprocessing means" refers to a system that processes captured data to prepare it in a format suitable for analysis.

[0123] "Means for encrypting data" refers to technology that encrypts transmitted data so that it cannot be deciphered by third parties.

[0124] "Server" refers to a central computing device that receives, analyzes, and processes captured data.

[0125] "Means for analyzing data" refers to technology used to analyze received data and detect signs of fraud.

[0126] "Speech recognition technology" refers to technology for converting voice data into text data.

[0127] "Natural language processing technology" refers to technology for analyzing text data and understanding its meaning and context.

[0128] "Means for scoring the likelihood of fraud" refers to a technology that numerically evaluates the likelihood of fraud from analyzed data.

[0129] "Threshold" refers to a benchmark for assessing the likelihood of fraud.

[0130] "Means for generating a warning message" refers to a technique for creating a message to warn a user when there is a high possibility of fraud.

[0131] "Feedback information" refers to information provided by a user as to whether a communication was fraudulent.

[0132] "Means for improving analytics" refers to techniques for using collected feedback information to improve fraud detection models.

[0133] This invention is a system that uses the power of AI to eradicate special frauds. It monitors users' communications in real time, detects signs of fraud, and issues warnings to effectively prevent fraud damage. This system is configured around the user's terminal and server.

[0134] Specific implementation methods

[0135] Monitoring user communications

[0136] Device: The user's device monitors calls and messages in real time. For calls, the device captures audio data using the smartphone's microphone, and for messages, the device obtains text data from an application on the device. For example, voice calls and text messages are monitored via a communication application on the smartphone.

[0137] Data preprocessing and transmission

[0138] On the device: The audio data is denoised and the sampling rate is adjusted. This process is performed using the Librosa library. The pre-processed data is then encrypted using AES encryption technology. The encrypted data is then securely sent to the server. For example, Librosa functions are used to denoise the audio from a call, and then the audio is encrypted using AES.

[0139] Data analysis

[0140] Server: Receives data sent from the device and begins analysis. In the case of voice data, it is first converted into text data using voice recognition technology. A general-purpose voice recognition API is used here. Next, the text data is analyzed using natural language processing technology to detect signs of fraud. Libraries such as spaCy and NLTK are used for this process. For example, voice data is converted into text using a recognition API, and the text data is then analyzed using spaCy.

[0141] Fraud indicator detection and scoring

[0142] Server: If indicators of fraud are detected from the text data, a scoring algorithm is used to calculate the likelihood of that fraud. Here, the scikit-learn library is used to evaluate the likelihood of fraud using random forests and logistic regression. For example, the likelihood of fraud is scored based on keywords such as "money" and "help" in the text.

[0143] Generate and send alerts

[0144] Server: If the likelihood of fraud exceeds a certain threshold, a warning message is generated for the user. The generated warning message is sent to the user's device and notified in real time. For example, a warning message such as "This call may be fraudulent" is generated and sent to the user's smartphone via push notification.

[0145] Feedback and Learning

[0146] User: Provide feedback on whether it was a scam. For example, after a call ends, the system asks the user, "Was this call a scam?" and the user answers Yes / No.

[0147] Server: Improves the machine learning model based on the feedback information collected from users. This process improves the accuracy of the next detection. This process is performed using TensorFlow and PyTorch. For example, the model is retrained using the collected feedback data.

[0148] Specific examples

[0149] Example 1: Detecting fraudulent calls

[0150] User: An elderly person receives a call from someone claiming to be their "son."

[0151] Device: Captures the call contents in real time through the smartphone's microphone and sends the audio data to the server.

[0152] Server: Converts the received voice data into text and uses natural language processing technology to detect signs of fraud, such as "I need money" or "Help me."

[0153] Server: Scores the likelihood of fraud and generates a warning message if it indicates a high probability.

[0154] Device: Display a warning message on the elderly person's device saying "This call may be fraudulent."

[0155] User: Realizes the call is a scam and hangs up.

[0156] User: Provide feedback to the system that it was a scam.

[0157] Server: Improves machine learning models based on feedback.

[0158] Example 2: Message fraud detection

[0159] Users: Young people receive suspicious emails.

[0160] Terminal: Captures email text in real time and sends it to the server.

[0161] Server: Analyzes text data using natural language processing technology to detect signs of fraud, such as "urgent" or "click here."

[0162] Server: A scoring algorithm assesses the likelihood of fraud and generates a warning message if a high probability is detected.

[0163] Devices: Display a warning message on young people's devices saying, "This email may be fraudulent."

[0164] Users: Ignore the email and avoid becoming a victim of fraud.

[0165] User: Provide feedback that the email was fraudulent.

[0166] Server: Improve machine learning models based on feedback.

[0167] Prompt Sentence Examples

[0168] Prompt Sentence Example 1

[0169] "Is this call potentially fraudulent? Please analyze the following call: 'Mom, I need money. Please send me 100,000 yen right away.'"

[0170] Prompt Sentence Example 2

[0171] "Do you think the following email might be fraudulent? Email: 'URGENT! There has been a suspicious login to your account. Click here to check.'"

[0172] In this way, the system of the present invention can effectively prevent fraud damage through real-time detection and warning of special frauds.

[0173] The flow of the identification process in the first embodiment will be described with reference to FIG.

[0174] Step 1:

[0175] Device: Monitors user communications in real time. For calls, the device captures audio data using the smartphone's microphone, and for messages, it obtains text data from the application. For example, the device can automatically start recording as soon as a call begins, capturing the content of the conversation.

[0176] Input: Call voice or text message

[0177] Output: Captured audio or text data

[0178] Step 2:

[0179] On the device: The captured audio data is denoised and the sampling rate is adjusted. This is done using the Librosa library. Specifically, the audio file is read, and the librosa.effects.reduce_noise function is used to remove noise and adjust the sampling rate.

[0180] Input: Captured audio data

[0181] Output: Denoised and sample-rate adjusted audio data

[0182] Step 3:

[0183] Terminal: Encrypt the preprocessed data using AES encryption technology, for example, using the pycryptodome library to perform AES encryption.

[0184] Input: Denoised and sample-rate adjusted audio data

[0185] Output: Encrypted audio data

[0186] Step 4:

[0187] On the device, the encrypted data is sent to the server, specifically using the HTTP / HTTPS protocol to send the data securely.

[0188] Input: Encrypted audio or text data

[0189] Output: Data sent to the server

[0190] Step 5:

[0191] Server: Decodes the data received from the device and then begins analysis. In the case of voice data, it first converts it into text data using voice recognition technology. This is done using a general-purpose voice recognition API.

[0192] Input: Encrypted data sent to the server

[0193] Output: Decoded audio and text data

[0194] Step 6:

[0195] Server: Analyzes text data using natural language processing techniques to detect signs of fraud. Specifically, it uses spaCy and NLTK to analyze keywords and context within the text.

[0196] Input: Decoded audio and text data

[0197] Output: Analysis results showing signs of fraud

[0198] Step 7:

[0199] Server: If indicators of fraud are detected from the text data, a scoring algorithm is used to calculate the probability of fraud. Random forest and logistic regression models are used using the scikit-learn library.

[0200] Input: Analysis results indicating fraud

[0201] Output: Scoring results

[0202] Step 8:

[0203] Server: If the likelihood of fraud exceeds a certain threshold, a warning message is generated and sent to the user's device. For example, a message saying "This call may be fraudulent" is created and sent to the user's smartphone.

[0204] Input: Scoring results

[0205] Output: The warning message sent to the user.

[0206] Step 9:

[0207] User: Provides feedback on whether the communication was in fact a scam. After the call ends, the system asks the user, "Was this call a scam?" and the user answers Yes / No.

[0208] Input: Communication feedback question

[0209] Output: Feedback information

[0210] Step 10:

[0211] Server: Based on the feedback information collected from users, the machine learning model is retrained and improved, which will improve the detection accuracy next time. This process is performed using TensorFlow and PyTorch.

[0212] Input: Feedback information

[0213] Output: An improved machine learning model

[0214] (Application example 1)

[0215] Next, a description will be given of Application Example 1. In the following description, the data processing device 12 will be referred to as a "server" and the smart device 14 will be referred to as a "terminal."

[0216] Currently, there are many cases of special frauds targeting the elderly and young people, and an effective system to prevent these frauds is needed. However, current fraud detection systems often only notify users on their devices, which makes it easy for users to overlook the warning, and few systems can immediately alert users. Therefore, there is a need for a system that can detect signs of fraud in real time and immediately and effectively warn users.

[0217] The specific processing by the specific processing unit 290 of the data processing device 12 in the application example 1 is realized by the following means.

[0218] In this invention, the server includes means for monitoring the content of user communications and capturing data in real time, means for transmitting the captured data to the server, means for analyzing the data received by the server and detecting the possibility of fraud, means for generating a warning based on the detected possibility of fraud and notifying the user, means for collecting feedback information and improving the analysis means based on the feedback information, and means for displaying a warning message on a head-mounted display in real time. This allows the user to visually confirm fraud warnings in real time, effectively preventing damage from special frauds.

[0219] "User communication content" refers to the content of calls and messages made by the user.

[0220] "Means for capturing data in real time" refers to devices or methods that instantly capture the content of ongoing communications.

[0221] The "means for transmitting to the server" refers to a means for sending the captured data to the central processing unit via a network.

[0222] "Means for analyzing and detecting potential fraud" refers to devices or methods for analyzing acquired data and detecting signs of fraudulent activity therein.

[0223] "Means for generating a warning and notifying a user" refers to means for creating a warning message and notifying a user when a potential fraud is detected.

[0224] "Means for collecting feedback information" refers to means for collecting opinions and evaluations from users.

[0225] "Means for improving analysis means" refers to means for improving data analysis methods based on collected feedback information.

[0226] A "head-mounted display" refers to a device worn by a user that displays information visually.

[0227] "Means for displaying a warning message in real time" refers to means for immediately visually conveying the detected warning content to the user.

[0228] The present invention is a system for monitoring user communications in real time, detecting signs of fraud, and issuing a warning. An embodiment of this system will be described in detail below.

[0229] Overall structure

[0230] This system consists of a user's terminal, a server that analyzes the data, and a head-mounted display (HMD) that notifies the user of warnings.

[0231] The server includes means for monitoring user communications and capturing data in real time, means for transmitting the captured data to the server, means for analyzing the received data to detect possible fraud, means for generating and notifying the user of a warning based on possible fraud, and means for collecting feedback information and improving the analysis means.

[0232] Implementation Details

[0233] 1. A way to monitor user communications and capture data in real time

[0234] The user's device monitors the contents of calls and messages in real time. For calls, audio data is captured using the HMD's built-in microphone. For messages, text data is obtained directly. This data is processed using the AudioCapture API (for audio data) and the TextMessage API (for text data).

[0235] 2. Data preprocessing and transmission

[0236] The captured data is preprocessed on the user's device. Specifically, audio data is subjected to noise removal (NoiseSuppressionLib) and sampling rate adjustment (SoundProcessingLib), and text data is preprocessed using a natural language processing library (e.g., spaCy). The data is then encrypted (SSLTLS) and sent to the server via the HTTPS protocol.

[0237] 3. Data Analysis

[0238] The server runs on a high-performance server (e.g., AWS EC2). The data received on the server side is converted into text using a speech recognition engine (Google Speech-to-Text API), and then analyzed using a natural language processing engine (BERT). This analysis detects signs of fraud.

[0239] 4. Alert Generation and Notification

[0240] The server evaluates the likelihood of fraud using a scoring algorithm (XGBoost). If there is a high likelihood of fraud, a warning message is generated. This warning message is sent in real time to the user's HMD and displayed using the Notification API.

[0241] 5. Feedback and Improvement

[0242] Users provide feedback on the resulting warnings, which is collected on the server and used to improve the analysis method using a machine learning platform (e.g., AWS SageMaker).

[0243] Examples of specific examples and prompts

[0244] As a concrete example, consider a case where an elderly person receives a call from someone claiming to be their "son" saying, "I urgently need money to pay for my hospital bills." In this case, the system will operate according to the following prompt:

[0245] Prompt statement:

[0246] 1. Audio data capture and preprocessing:

[0247] Capture the user's voice in real time through the microphone, remove noise with NoiseSuppressionLib, and adjust the sampling rate with SoundProcessingLib.

[0248] 2. Data transmission:

[0249] The preprocessed data is encrypted using SSLTLS and sent to the server using HTTPS.

[0250] 3. Data Analysis:

[0251] On the server side, the speech is converted to text using the Google Speech-to-Text API and analyzed using BERT.

[0252] 4. Alert Generation and Notification:

[0253] XGBoost scores the likelihood of fraud and generates a warning message, which is displayed on the HMD using the Notification API.

[0254] 5. Feedback Processing:

[0255] Collect user feedback and improve machine learning models with AWS SageMaker.

[0256] This configuration allows users to receive immediate warnings about the risk of special fraud, making it possible to effectively prevent fraud damage.

[0257] The flow of the specific processing in the application example 1 will be described with reference to FIG.

[0258] Step 1:

[0259] The device monitors the user's communication in real time and captures the data. The input is the user's call voice and message content, and the output is the captured raw data. The audio data is obtained through the microphone and processed using the AudioCapture API. The text data is obtained directly using the TextMessage API.

[0260] Step 2:

[0261] Preprocess the data captured by the device. For audio data, use NoiseSuppressionLib to remove noise and SoundProcessingLib to adjust the sampling rate. For text data, use a natural language processing library (e.g. spaCy) for preprocessing. The input is raw data and the output is preprocessed data.

[0262] Step 3:

[0263] The terminal encrypts the preprocessed data and sends it to the server. The input is the preprocessed data and the output is the encrypted data. SSL TLS is used for encryption, and the data is sent via the HTTPS protocol.

[0264] Step 4:

[0265] The server decrypts the received data and performs data analysis. The input is the encrypted data, and the output is the analyzed result. In the case of audio data, the server converts the audio to text using the Google Speech-to-Text API, and then analyzes the text data using BERT.

[0266] Step 5:

[0267] The server scores the likelihood of fraud based on the analysis results and generates a warning message. The input is the analyzed text data, and the output is the warning message. XGBoost is used for scoring.

[0268] Step 6:

[0269] The server generates a warning message and sends it to the user's HMD, where it is displayed using the Notification API. The input is the warning message, and the output is the warning content reflected in the user's visual perception.

[0270] Step 7:

[0271] The user acknowledges the warning message and provides feedback. The input is the user's feedback information, and the output is the collected feedback information.

[0272] Step 8:

[0273] The server improves the analysis method based on the feedback information collected. The input is the feedback information, and the output is the improved analysis method. The feedback information is reflected in the machine learning model using AWS SageMaker.

[0274] Furthermore, an emotion engine that estimates the user's emotion may be combined. That is, the identification processing unit 290 may estimate the user's emotion using the emotion identification model 59 and perform identification processing using the user's emotion.

[0275] This invention realizes more advanced fraud detection by incorporating an emotion engine into a system that uses the power of AI to eradicate special frauds. The system not only monitors users' communications in real time and detects signs of fraud, but also analyzes users' emotions to more accurately assess the possibility of fraud and issue warnings.

[0276] Overall system overview

[0277] 1. Monitoring of outgoing information

[0278] Terminal: Monitors user calls and messages in real time. For calls, the microphone is used to capture audio data, and for messages, the text data is directly acquired.

[0279] 2. Data preprocessing and transmission

[0280] Terminal: For audio data, preprocessing such as noise removal and sampling rate adjustment is performed. The preprocessed data is encrypted and sent to the server.

[0281] 3. Data Analysis

[0282] Server: Analyzes the received data. If it is voice data, it is first converted into text data using speech recognition technology. Next, it analyzes the text data using natural language processing technology to detect signs of fraud.

[0283] 4. Emotional awareness and regulation

[0284] Server: The emotion engine uses the analyzed text and voice data to recognize the user's emotions. The emotion is analyzed based on the tone of voice and the use of words.

[0285] 5. Fraud Indicator Detection and Scoring

[0286] Server: If fraud indicators are detected, adjust the fraud scoring based on the output of the emotion engine, for example increasing the score if the user is feeling anxious or scared.

[0287] 6. Generating and Sending Alerts

[0288] Server: Generates a warning message when the likelihood of fraud exceeds a certain threshold, and sends the generated warning message to the user's device to notify them in real time.

[0289] 7. Feedback and Learning

[0290] User: Provide feedback on whether it is actually a scam.

[0291] Server: Use this feedback to improve the machine learning model and make the next detection more accurate.

[0292] Specific examples

[0293] Example 1: Detecting fraudulent calls using an emotion engine

[0294] User: An elderly person receives a call from someone claiming to be their "son."

[0295] Terminal: Captures the call in real time and sends the audio data to the server.

[0296] Server: Converts voice data into text and uses natural language processing technology to detect signs of fraud, such as "I need money" or "Help me."

[0297] Server: Recognizes when an elderly person is feeling anxious or scared from text and speech analyzed using an emotion engine.

[0298] Server: Score the probability of fraud, increase the score taking into account the results of sentiment analysis, and generate a warning message if the probability is high.

[0299] Devices: Display a warning message on the senior's device to notify them of potential fraud.

[0300] User: Realizes the call is a scam and hangs up.

[0301] User: Provide feedback to the system that it was a scam.

[0302] Server: Improves machine learning models based on feedback.

[0303] Example 2: Detecting message fraud using an emotion engine

[0304] User: A young person receives a suspicious email.

[0305] Terminal: Captures email text in real time and sends it to the server.

[0306] Server: Analyzes text data using natural language processing technology to detect signs of fraud, such as "urgent," "important," and "click here."

[0307] Server: Recognizes from text analyzed using an emotion engine that young people are feeling confused or excited.

[0308] Server: Score the probability of fraud, increase the score taking into account the results of sentiment analysis, and generate a warning message if the probability is high.

[0309] Devices: Display warning messages on young people's devices to inform them of potential scams.

[0310] User: Ignore the email to prevent unauthorized access.

[0311] User: Provide feedback that the email was fraudulent.

[0312] Server: Improve machine learning models based on feedback.

[0313] In this way, this system can effectively prevent fraud damage through real-time detection and warning of special frauds, and even by recognizing user emotions.

[0314] The processing flow will be explained below.

[0315] Step 1:

[0316] User: Makes calls, sends and receives messages, and exchanges necessary information, including potentially fraudulent communications.

[0317] Step 2:

[0318] Terminal: Captures the contents of calls as audio data in real time, and also captures text messages as text data.

[0319] Step 3:

[0320] Terminal: Preprocessing the captured audio data, such as noise removal and unifying the sampling rate, improves the quality of the audio data.

[0321] Step 4:

[0322] Terminal: The pre-processed data is encrypted and sent to the server using a secure communication protocol, along with the communication metadata (source, destination, time, etc.).

[0323] Step 5:

[0324] Server: Receives data sent from the device. In the case of voice data, converts it into text using voice recognition technology. Converts the voice content into text information using a voice recognition engine.

[0325] Step 6:

[0326] Server: Analyzes the received text data using natural language processing (NLP) techniques, such as tokenizing words, tagging parts of speech, and grammar analysis, to detect phrases and structures that may be indicative of fraud.

[0327] Step 7:

[0328] Server: Uses an emotion engine to recognize user emotions from text data. Analyzes user emotions (e.g., anxiety, anger, confusion, etc.) based on voice tone and phrasing.

[0329] Step 8:

[0330] Server: If fraud signs are detected, the server scores the likelihood of fraud by taking into account the user's emotions analyzed by the emotion engine. The likelihood of fraud is adjusted so that the score is higher if the user is feeling anxious or angry.

[0331] Step 9:

[0332] Server: Generates a warning message if the fraud score exceeds a set threshold, including the likelihood of fraud and what to do about it.

[0333] Step 10:

[0334] Server: Sends the generated warning message to the user's device, notifying them in real time and alerting them.

[0335] Step 11:

[0336] Device: Display an alert message on the user's device, using a notification sound or vibration to get the user's attention.

[0337] Step 12:

[0338] Users: Check whether the call or message is fraudulent and take the action outlined in the warning message. If it is fraudulent, immediately stop the transaction or call.

[0339] Step 13:

[0340] Users: Provide feedback to the system on whether the fraud is true or not. User judgment is also important data for improving the system.

[0341] Step 14:

[0342] Device: Sends user feedback information to the server, including feedback content (whether it was a scam, the effectiveness of countermeasures, etc.).

[0343] Step 15:

[0344] Server: Receives feedback information and uses it as data to improve the performance of the machine learning model and emotion engine. Based on the feedback, the detection algorithm is retrained.

[0345] Through this process, the system can detect signs of specialized fraud in real time, and by utilizing an emotion engine, it can accurately assess the possibility of fraud and issue a warning.

[0346] Example 2

[0347] Next, a description will be given of Example 2. In the following description, the data processing device 12 will be referred to as a "server" and the smart device 14 will be referred to as a "terminal."

[0348] In recent years, special frauds have been on the rise, with the damage targeting elderly people and beginners becoming particularly severe. Conventional fraud prevention systems are capable of detecting signs of fraud, but they have limitations in the timing and accuracy of issuing warnings. Furthermore, because they do not take the user's emotional state into account, they have the problem of easily missing important warnings. Therefore, there is a need for a system that can more accurately detect potential fraud and issue warnings promptly by monitoring communication content in real time and analyzing the user's emotions.

[0349] The specific processing by the specific processing unit 290 of the data processing device 12 in the second embodiment is realized by the following means.

[0350] In this invention, the server includes means for monitoring user communications and capturing data in real time, means for preprocessing and encrypting the captured data and transmitting it to the server, means for analyzing the data received by the server and detecting the possibility of fraud, means for recognizing an emotional state from the analyzed data and voice data, means for scoring fraud using the recognized emotional state, means for generating a warning based on the detected possibility of fraud and scoring data and notifying the user, and means for collecting feedback information and improving the analysis means based on the feedback information, thereby enabling more accurate fraud detection and faster warnings that reflect the user's emotional state.

[0351] "User communications" refers to digital communications such as calls, messages, and emails sent or received by a user.

[0352] "Means for capturing data in real time" refers to a piece of equipment or software that acquires the contents of user communications in real time and processes them immediately.

[0353] "Preprocessing" refers to initial processing such as noise removal and sampling rate adjustment that is performed to convert captured data into a form that is easier to analyze.

[0354] "Encryption" is the technology of converting data into cryptographic code to protect the privacy and security of communication data.

[0355] A "server" is a central computer system that receives data from multiple users over a network and processes and analyzes it.

[0356] "Means for detecting potential fraud" are algorithms and technologies that analyze the content of communications to find signs of fraud.

[0357] "Means for recognizing emotional states" refers to technology that analyzes the tone and phrasing of a user's voice or text to identify the emotions the user is feeling.

[0358] A "fraud scoring means" is a technology that numerically calculates the likelihood of fraud based on detected indicators of fraud and the user's emotional state.

[0359] "Means for generating a warning and notifying the user" refers to technology that automatically creates a warning message when a high possibility of fraud is determined and sends it to the user's device.

[0360] "Feedback information" refers to data on reactions and evaluations provided by users in response to warnings issued by the system.

[0361] "Means to improve analysis means" refers to using feedback information to improve fraud detection algorithms and emotion recognition techniques, thereby increasing the accuracy of the overall system.

[0362] The present invention is a system that uses the power of AI to eradicate special frauds, and by incorporating an emotion engine, achieves more advanced fraud detection. The system not only monitors user communications in real time and detects signs of fraud, but also analyzes user emotions to more accurately assess the possibility of fraud and issue warnings. Specific embodiments of the present invention are described below.

[0363] 1. Monitoring of outgoing information

[0364] Device: Monitors user calls and messages in real time. For calls, audio data is captured using the device's built-in microphone, and for messages, text data is directly acquired.

[0365] 2. Data preprocessing and transmission

[0366] Terminal: For audio data, first preprocessing such as noise removal and sampling rate adjustment is performed, and the preprocessed data is encrypted using the AES-256 encryption algorithm and sent to the server via the SSL / TLS protocol.

[0367] 3. Speech data conversion and natural language processing

[0368] Server: Receives the voice data and converts it to text using the Google Cloud Speech-to-Text API. The converted text data is analyzed using a natural language processing library (e.g., spaCy or BERT) to detect signs of fraud, such as "I need money" or "Help me."

[0369] 4. Emotion Recognition and Analysis

[0370] Server: Inputs text and voice data into Affectiva's emotion engine to analyze the user's emotional state, analyzing voice tone and phrasing to identify emotions such as "anxiety," "fear," and "confusion."

[0371] 5. Fraud Indicator Scoring

[0372] Server: Calculates the fraud score based on the detected fraud indicators and the results of sentiment analysis. If the user shows high anxiety or fear, the fraud score is increased.

[0373] 6. Generating and Sending Warning Messages

[0374] Server: When the fraud score exceeds the set threshold, a warning message is automatically generated and sent to the user's device immediately, where the device notifies the user.

[0375] 7. Gather feedback and learn

[0376] User: After receiving the warning message, the user can provide feedback on whether it was indeed a scam, for example, by replying "It was a scam" or "It wasn't a scam."

[0377] Server: Stores the feedback information in a database and updates the machine learning model (e.g., Scikit-learn or TensorFlow) to improve the accuracy of fraud detection next time.

[0378] Specific examples

[0379] Example 1: Detecting fraudulent calls using an emotion engine

[0380] User: An elderly person receives a call from someone claiming to be their "son."

[0381] Terminal: Captures the call in real time and sends the audio data to the server.

[0382] Server: Converts voice data into text and uses natural language processing technology to detect signs of fraud, such as "I need money" or "Help me."

[0383] Server: Recognizes when an elderly person is feeling anxious or scared from text and speech analyzed using an emotion engine.

[0384] Server: Score the likelihood of fraud, increase the score based on the results of sentiment analysis, and generate a warning message if the probability is high.

[0385] Devices: Display a warning message on the senior's device to notify them of potential fraud.

[0386] User: Realizes the call is a scam and hangs up.

[0387] User: Provide feedback to the system that it was a scam.

[0388] Server: Improves machine learning models based on feedback.

[0389] Example prompts to be input to the generative AI model

[0390] "Someone claiming to be my son said he needed money. Please analyze the possibility that this call is fraudulent. I have the audio data and pre-processed text data."

[0391] The flow of the identification process in the second embodiment will be described with reference to FIG.

[0392] Step 1:

[0393] Terminal: Monitors user calls and messages in real time. The input is the user's voice or text message, and the output is captured voice data or text data. For voice, the voice data is captured using the microphone built into the terminal. For messages, the text data is obtained as is.

[0394] Step 2:

[0395] Terminal: Preprocesses the captured audio data. The input is the captured audio data, and the output is the preprocessed audio data. This includes applying noise reduction filters and adjusting the sampling rate. After preprocessing, the data is encrypted using the AES-256 encryption algorithm and sent to the server using the SSL / TLS protocol.

[0396] Step 3:

[0397] Server: Converts audio data to text data. The input is preprocessed audio data and the output is text data. The audio data is converted to text using the Google Cloud Speech-to-Text API. This text data is sent to the next processing step.

[0398] Step 4:

[0399] Server: Analyzes the text data using natural language processing techniques. The input is text data, and the output is the analysis results containing information about signs of fraud. Natural language processing libraries (e.g., spaCy and BERT) are used to detect fraudulent keywords and phrases such as "I need money" and "Help me."

[0400] Step 5:

[0401] Server: Analyzes the user's emotional state using an emotion engine. The input is text data and voice data, and the output is data indicating the user's emotional state. Using Affectiva's emotion engine, it identifies emotions such as "anxiety," "fear," and "confusion" from voice tone and vocabulary.

[0402] Step 6:

[0403] Server: Performs fraud scoring. The input is fraud indicator information and emotional state data, and the output is a fraud score. Based on the detected fraud indicators and the user's emotional state, the possibility of fraud is numerically evaluated. If the emotional state is "anxiety" or "fear," the fraud score increases.

[0404] Step 7:

[0405] Server: Generates a warning message and notifies the user. The input is the fraud score and the output is a warning message. If the fraud score exceeds the set threshold, a warning message is automatically generated. This message is immediately sent to the user's device and the user is notified on the device.

[0406] Step 8:

[0407] User: Provides feedback. The input is the user's reaction to the warning, and the output is feedback information. The user reports whether the warning was correct or not, replying to the system "it was a scam" or "it wasn't a scam."

[0408] Step 9:

[0409] Server: Updates the machine learning model based on the feedback information. The input is the feedback information, and the output is an improved analysis model. The received feedback is stored in a database and used to improve the model using machine learning algorithms (e.g., Scikit-learn or TensorFlow) to improve the fraud detection accuracy next time.

[0410] (Application example 2)

[0411] Next, a description will be given of Application Example 2. In the following description, the data processing device 12 will be referred to as a "server" and the smart device 14 will be referred to as a "terminal."

[0412] Conventional fraud detection systems only detect signs of fraud contained in user communications, and improving detection accuracy by reflecting the user's emotional state is a challenge. Furthermore, there is a lack of real-time warnings and feedback-based model improvements, limiting the effectiveness of fraud prevention.

[0413] The specific processing by the specific processing unit 290 of the data processing device 12 in the application example 2 is realized by the following means.

[0414] In this invention, the server includes means for monitoring user communications and detecting signs of fraud in real time, means for analyzing user emotions using an emotion engine and adjusting the fraud score, means for generating a warning based on the likelihood of fraud and notifying the user, and means for collecting feedback information and improving the analysis means, thereby enabling advanced fraud detection that takes user emotions into account, real-time warnings, and continuous system improvement based on feedback.

[0415] "User communication content" refers to data such as voice calls, messages, and emails exchanged between the user and others.

[0416] "Means of capturing data in real time" refers to technology that instantly collects communication content and stores it as data for later processing.

[0417] A "server" refers to a computer system that receives data from multiple client devices via a network and analyzes and processes the data.

[0418] "Measures to detect potential fraud" refers to algorithms or technologies used to determine whether a user's communications contain fraudulent activity.

[0419] "Means for recognizing emotions and adjusting fraud scores" refers to technology that analyzes a user's emotional state and modifies the indicators used to assess the likelihood of fraud.

[0420] "Means for generating an alert and notifying the user" refers to technology that creates an alert message based on the detected potential fraud and sends it to the user's device.

[0421] "Means for collecting feedback information and improving analytical methods" refers to technologies that collect user feedback and use that information to improve fraud detection algorithms and the overall performance of the system.

[0422] "Means for converting voice data into text" refers to the process of converting voice data into character data using voice recognition technology.

[0423] "Natural language processing technology" refers to technology for analyzing the structure and meaning of language based on text data.

[0424] "Emotion analysis technology using generative models" refers to technology that uses machine learning models to estimate user emotions from text data and voice data.

[0425] This invention is a system that monitors the content of user communications in real time, analyzes signs of fraud and user emotions, and thereby detects the possibility of fraud with high accuracy and issues a warning.

[0426] System configuration

[0427] Hardware

[0428] Devices: Microphones and processing units built into smartphones, smart glasses, and head-mounted displays

[0429] Server: Cloud server capable of high-performance analysis and data processing

[0430] software

[0431] Audio data capture: sounddevice library

[0432] Audio data preprocessing: librosa library

[0433] Speech Recognition: TensorFlow Model

[0434] Sentiment Analysis: Hugging Face transformers library

[0435] Processing Details

[0436] Monitoring user communications and capturing data

[0437] The device monitors the user's calls and messages in real time and captures voice data. In the case of a call, the voice data is collected using a microphone and pre-processed, such as noise reduction and sampling rate adjustment.

[0438] Data transmission and analysis

[0439] The pre-processed data is encrypted and then sent to the server, which then converts the received voice data into text using natural language processing technology.

[0440] Sentiment Analysis and Fraud Scoring

[0441] On the server side, a generative AI model is used to analyze user sentiment from text and voice data, and the fraud score is adjusted based on the analyzed sentiment.

[0442] Generate alerts and notify users

[0443] If the likelihood of fraud exceeds a certain threshold, the server generates a warning message and notifies the user's terminal in real time.

[0444] Feedback and machine learning model improvement

[0445] Users provide feedback to the system, which the server uses to improve the machine learning model and make the next detection more accurate.

[0446] Specific examples

[0447] For example, if an elderly person receives a phone call from someone claiming to be their "son," the following steps may occur:

[0448] 1. The elderly person's device captures the contents of the call in real time and sends it to the server.

[0449] 2. The server converts the voice data into text and uses natural language processing technology to analyze signs of fraud, such as "I need money."

[0450] 3. Furthermore, generative AI models are used to analyze the emotions of older adults (e.g., anxiety and fear) from text and voice data.

[0451] 4. If it is determined that there is a high possibility of fraud, a warning message will be displayed on the elderly person's device.

[0452] Prompt Sentence Examples

[0453] "Please answer the following questions based on your call data:

[0454] 1. Are there any signs of fraud in this conversation?

[0455] 2. What are the user's emotions?

[0456] example:

[0457] Call text: 'Mom, help me! I need money now.'"

[0458] In this way, the invention allows for advanced fraud detection that takes user emotions into account, real-time alerts, and continuous system improvement based on feedback.

[0459] The flow of the specific processing in the application example 2 will be described with reference to FIG.

[0460] Step 1:

[0461] The device monitors the user's communication in real time and captures audio data. Specifically, it collects audio data using a microphone and obtains it through an audio data capture library (e.g., sounddevice). The input is real-time audio, and the output is captured audio data.

[0462] Step 2:

[0463] The device preprocesses the captured audio data. Specifically, it performs operations such as noise reduction and sampling rate adjustment. This process uses an audio data preprocessing library (e.g., librosa). The input is the captured audio data, and the output is the preprocessed audio data.

[0464] Step 3:

[0465] The device encrypts the preprocessed audio data and sends it to the server. Specifically, it uses an encryption library to securely convert the data and send it to the server over the network. The preprocessed audio data is the input, and the encrypted data is sent to the server as the output.

[0466] Step 4:

[0467] The server analyzes the received voice data. Specifically, it uses speech recognition technology (e.g., TensorFlow model) to convert the voice data into text data. The input is encrypted voice data, and the output is text data.

[0468] Step 5:

[0469] The server analyzes the text data using natural language processing technology. Specifically, it uses natural language processing technology (e.g., a generative AI model) to detect signs of fraud from the text data. The input is text data, and the output is a determination of the likelihood of fraud.

[0470] Step 6:

[0471] The server recognizes emotions from the analyzed data and adjusts the fraud score. Specifically, it analyzes the user's emotions using emotion analysis technology (e.g., Hugging Face's transformers library) and updates the fraud score. The input is the analyzed text data and the emotion analysis result, and the output is the adjusted fraud score.

[0472] Step 7:

[0473] The server generates a warning message and notifies the user if the fraud probability exceeds a certain threshold. Specifically, the server automatically generates a warning message and sends it to the user's device. The input is the adjusted fraud score, and the output is the warning message displayed on the user's device.

[0474] Step 8:

[0475] The user provides feedback on whether or not the fraud is actually occurring. Specifically, the user reports whether or not the fraud is occurring within the application, and the feedback is sent to the server. The input is the user's feedback information, and the output is the feedback data provided to the server.

[0476] Step 9:

[0477] The server improves the machine learning model based on the feedback information. Specifically, it analyzes the collected feedback information and runs an algorithm to improve the performance of the fraud detection model. The input is the feedback information and the output is an improved machine learning model.

[0478] The specific processing unit 290 transmits the result of the specific processing to the smart device 14. In the smart device 14, the control unit 46A causes the output device 40 to output the result of the specific processing. The microphone 38B acquires audio indicating a user input regarding the result of the specific processing. The control unit 46A transmits audio data indicating the user input acquired by the microphone 38B to the data processing device 12. In the data processing device 12, the specific processing unit 290 acquires the audio data.

[0479] The data generation model 58 is a so-called generative AI (Artificial Intelligence). An example of the data generation model 58 is ChatGPT (Internet Search<URL: https: / / openai.com / blog / chatgpt> ), Gemini (Internet search <url: https: gemini.google.com ?hl="ja">) and other generation AIs. The data generation model 58 is obtained by performing deep learning on a neural network. A prompt including an instruction is input to the data generation model 58, and inference data such as voice data indicating voice, text data indicating text, and image data indicating an image is also input. The data generation model 58 performs inference on the input inference data in accordance with the instruction indicated by the prompt, and outputs the inference result in a data format such as voice data and text data. Here, inference refers to, for example, analysis, classification, prediction, and / or summarization.

[0480] In the above embodiment, an example in which the specific process is performed by the data processing device 12 has been given, but the technology of the present disclosure is not limited to this, and the specific process may be performed by the smart device 14.

[0481] [Second embodiment]

[0482] FIG. 3 shows an example of the configuration of a data processing system 210 according to the second embodiment.

[0483] 3, the data processing system 210 includes the data processing device 12 and smart glasses 214. An example of the data processing device 12 is a server.

[0484] The data processing device 12 includes a computer 22, a database 24, and a communication I / F 26. The computer 22 is an example of a "computer" according to the technology of the present disclosure. The computer 22 includes a processor 28, a RAM 30, and a storage 32. The processor 28, the RAM 30, and the storage 32 are connected to a bus 34. The database 24 and the communication I / F 26 are also connected to the bus 34. The communication I / F 26 is connected to a network 54. Examples of the network 54 include a WAN (Wide Area Network) and / or a LAN (Local Area Network).

[0485] The smart glasses 214 include a computer 36, a microphone 238, a speaker 240, a camera 42, and a communication I / F 44. The computer 36 includes a processor 46, a RAM 48, and a storage 50. The processor 46, the RAM 48, and the storage 50 are connected to a bus 52. The microphone 238, the speaker 240, and the camera 42 are also connected to the bus 52.

[0486] The microphone 238 receives instructions and the like from the user 20 by receiving voice uttered by the user 20. The microphone 238 captures the voice uttered by the user 20, converts the captured voice into audio data, and outputs it to the processor 46. The speaker 240 outputs audio in accordance with instructions from the processor 46.

[0487] Camera 42 is a small digital camera equipped with an optical system including a lens, aperture, and shutter, and an imaging element such as a CMOS (Complementary Metal-Oxide-Semiconductor) image sensor or a CCD (Charge Coupled Device) image sensor, and captures images of the surroundings of user 20 (for example, an imaging range defined by an angle of view equivalent to the field of vision of a typical healthy person).

[0488] The communication I / F 44 is connected to a network 54. The communication I / Fs 44 and 26 are responsible for the exchange of various information between the processor 46 and the processor 28 via the network 54. The exchange of various information between the processor 46 and the processor 28 using the communication I / Fs 44 and 26 is carried out in a secure state.

[0489] Fig. 4 shows an example of the main functions of the data processing device 12 and the smart glasses 214. As shown in Fig. 4, in the data processing device 12, a specific process is performed by the processor 28. A specific process program 56 is stored in the storage 32.

[0490] The specific processing program 56 is an example of a "program" according to the technology of the present disclosure. The processor 28 reads the specific processing program 56 from the storage 32 and executes the read specific processing program 56 on the RAM 30. The specific processing is realized by the processor 28 operating as a specific processing unit 290 in accordance with the specific processing program 56 executed on the RAM 30.

[0491] The storage 32 stores a data generation model 58 and an emotion identification model 59. The data generation model 58 and the emotion identification model 59 are used by the identification processing unit 290.

[0492] In the smart glasses 214, the reception output process is performed by the processor 46. A reception output program 60 is stored in the storage 50. The processor 46 reads the reception output program 60 from the storage 50 and executes the read reception output program 60 on the RAM 48. The reception output process is realized by the processor 46 operating as the control unit 46A in accordance with the reception output program 60 executed on the RAM 48.

[0493] Next, a description will be given of the identification process performed by the identification processing unit 290 of the data processing device 12. In the following description, the data processing device 12 will be referred to as the "server" and the smart glasses 214 will be referred to as the "terminal."

[0494] This invention is a system that uses the power of AI to eradicate special frauds. It monitors users' communications in real time, detects signs of fraud, and issues warnings to effectively prevent fraud damage. This system is configured around the user's terminal and server.

[0495] Overall system overview

[0496] 1. Monitoring of outgoing information

[0497] Terminal: Monitors user calls and messages in real time. For calls, the microphone is used to capture audio data, and for messages, the text data is directly acquired.

[0498] 2. Data preprocessing and transmission

[0499] Terminal: For audio data, preprocessing such as noise removal and sampling rate adjustment is performed. The preprocessed data is encrypted and sent to the server.

[0500] 3. Data Analysis

[0501] Server: Analyzes the received data. In the case of voice data, it is first converted into text data using speech recognition technology. Next, it analyzes the text data using natural language processing technology to detect signs of fraud.

[0502] 4. Fraud detection and scoring

[0503] Server: If any fraud indications are detected, a scoring algorithm is used to calculate the likelihood of fraud. Based on this score, the likelihood of fraud is assessed.

[0504] 5. Generating and Sending Alerts

[0505] Server: Generates a warning message when the likelihood of fraud exceeds a certain threshold, and sends the generated warning message to the user's device to notify them in real time.

[0506] 6. Feedback and Learning

[0507] User: Provide feedback on whether it is actually a scam.

[0508] Server: Use this feedback to improve the machine learning model and make the next detection more accurate.

[0509] Specific examples

[0510] Example 1: Detecting fraudulent calls

[0511] User: An elderly person receives a call from someone claiming to be their "son."

[0512] Terminal: Captures the call in real time and sends the audio data to the server.

[0513] Server: Converts voice data into text and uses natural language processing technology to detect signs of fraud, such as "I need money" or "Help me."

[0514] Server: Scores the likelihood of fraud and generates a warning message if it indicates a high probability.

[0515] Devices: Display a warning message on the senior's device to notify them of potential fraud.

[0516] User: Realizes the call is a scam and hangs up.

[0517] User: Provide feedback to the system that it was a scam.

[0518] Server: Improves machine learning models based on feedback.

[0519] Example 2: Message fraud detection

[0520] User: A young person receives a suspicious email.

[0521] Terminal: Captures email text in real time and sends it to the server.

[0522] Server: Analyzes text data using natural language processing technology to detect signs of fraud, such as "urgent" or "click here."

[0523] Server: A scoring algorithm assesses the likelihood of fraud and generates a warning message if it indicates a high probability.

[0524] Devices: Display warning messages on young people's devices to inform them of potential scams.

[0525] User: Ignore the email to prevent unauthorized access.

[0526] User: Provide feedback that the email was fraudulent.

[0527] Server: Improve machine learning models based on feedback.

[0528] In this way, this system can effectively prevent fraud damage through real-time detection and warning of special frauds.

[0529] The processing flow will be explained below.

[0530] Step 1:

[0531] Users: Make and receive calls and messages, some of which may be potentially fraudulent.

[0532] Step 2:

[0533] Terminal: Captures the user's call content as audio data in real time, and also captures the message content directly in the case of text messages.

[0534] Step 3:

[0535] Terminal: For audio data, preprocessing is performed, such as noise removal and unifying the sampling rate. The preprocessed data is prepared for transmission to the server.

[0536] Step 4:

[0537] Terminal: Encrypts the pre-processed data and sends it to the server using a secure communication protocol.

[0538] Step 5:

[0539] Server: Receives data sent from the device, including voice and text data.

[0540] Step 6:

[0541] Server: Converts the received voice data into text data using voice recognition technology. A voice recognition engine is used for this purpose.

[0542] Step 7:

[0543] Server: Analyzes text data using natural language processing (NLP) techniques, such as tokenizing words, tagging parts of speech, and grammar analysis.

[0544] Step 8:

[0545] Server: Detects signs of fraud from the analyzed data, using pre-trained AI models to identify phrases and contexts associated with fraud.

[0546] Step 9:

[0547] Server: Based on the detected fraud indicators, a scoring algorithm is used to assess the likelihood of fraud, using techniques such as logistic regression or support vector machines (SVM).

[0548] Step 10:

[0549] Server: If the score exceeds a set threshold, generate a warning message containing information about the potential for fraud and what to do about it.

[0550] Step 11:

[0551] Server: Sends the generated warning message to the user's terminal.

[0552] Step 12:

[0553] Device: Display an alert message on the user's device, using a notification sound or vibration to get the user's immediate attention.

[0554] Step 13:

[0555] User: Review the warning message and decide how to respond, including whether it is actually a scam.

[0556] Step 14:

[0557] User: Provide feedback to the system as to whether it was a scam or not.

[0558] Step 15:

[0559] Device: Sends user feedback to the server.

[0560] Step 16:

[0561] Server: Receives feedback information and uses it to refine the machine learning model, leveraging a feedback loop to continuously improve the accuracy of the analysis module.

[0562] Through this process, the system can detect signs of special fraud in real time and effectively warn users.

[0563] Example 1

[0564] Next, a description will be given of Example 1. In the following description, the data processing device 12 will be referred to as a "server" and the smart glasses 214 will be referred to as a "terminal."

[0565] Recently, special frauds have become increasingly sophisticated, increasing the risk of many people becoming victims of fraud. Elderly people and users unfamiliar with the Internet are particularly vulnerable to fraud. Current systems have difficulty detecting fraudulent activities in advance and issuing effective warnings, so effective prevention measures are needed.

[0566] The specific processing by the specific processing unit 290 of the data processing device 12 in the first embodiment is realized by the following means.

[0567] In this invention, the server includes means for monitoring user communications in real time and capturing the data as voice or text data, preprocessing means for performing noise reduction and sampling rate adjustment on the captured data, means for encrypting the preprocessed data and transmitting it to the server, means for analyzing the data received by the server and detecting the possibility of fraud, means for scoring the possibility of fraud based on the analyzed data, means for generating a warning message and notifying the user when the possibility of fraud exceeds a certain threshold, and means for collecting feedback information from users and improving the analysis means based on the feedback information, thereby making it possible to detect fraudulent activity in real time and quickly issue a warning to the user.

[0568] "User" refers to any individual or organization that uses this system.

[0569] "Communication content" refers to data such as voice calls and text messages sent by users.

[0570] "Real-time" refers to processing data almost as soon as the communication content occurs.

[0571] "Means for capturing data" refers to hardware or software for acquiring the contents of user communications.

[0572] "Noise reduction" refers to the process of reducing unnecessary noise from audio data.

[0573] "Sampling rate adjustment" refers to digital processing to convert audio data into a format that is optimal for analysis.

[0574] "Preprocessing means" refers to a system that processes captured data to prepare it in a format suitable for analysis.

[0575] "Means for encrypting data" refers to technology that encrypts transmitted data so that it cannot be deciphered by third parties.

[0576] "Server" refers to a central computing device that receives, analyzes, and processes captured data.

[0577] "Means for analyzing data" refers to technology used to analyze received data and detect signs of fraud.

[0578] "Speech recognition technology" refers to technology for converting voice data into text data.

[0579] "Natural language processing technology" refers to technology for analyzing text data and understanding its meaning and context.

[0580] "Means for scoring the likelihood of fraud" refers to a technology that numerically evaluates the likelihood of fraud from analyzed data.

[0581] "Threshold" refers to a benchmark for assessing the likelihood of fraud.

[0582] "Means for generating a warning message" refers to a technique for creating a message to warn a user when there is a high possibility of fraud.

[0583] "Feedback information" refers to information provided by a user as to whether a communication was fraudulent.

[0584] "Means for improving analytics" refers to techniques for using collected feedback information to improve fraud detection models.

[0585] This invention is a system that uses the power of AI to eradicate special frauds. It monitors users' communications in real time, detects signs of fraud, and issues warnings to effectively prevent fraud damage. This system is configured around the user's terminal and server.

[0586] Specific implementation methods

[0587] Monitoring user communications

[0588] Device: The user's device monitors calls and messages in real time. For calls, the device captures audio data using the smartphone's microphone, and for messages, the device obtains text data from an application on the device. For example, voice calls and text messages are monitored via a communication application on the smartphone.

[0589] Data preprocessing and transmission

[0590] On the device: The audio data is denoised and the sampling rate is adjusted. This process is performed using the Librosa library. The pre-processed data is then encrypted using AES encryption technology. The encrypted data is then securely sent to the server. For example, Librosa functions are used to denoise the audio from a call, and then the audio is encrypted using AES.

[0591] Data analysis

[0592] Server: Receives data sent from the device and begins analysis. In the case of voice data, it is first converted into text data using voice recognition technology. A general-purpose voice recognition API is used here. Next, the text data is analyzed using natural language processing technology to detect signs of fraud. Libraries such as spaCy and NLTK are used for this process. For example, voice data is converted into text using a recognition API, and the text data is then analyzed using spaCy.

[0593] Fraud indicator detection and scoring

[0594] Server: If indicators of fraud are detected from the text data, a scoring algorithm is used to calculate the likelihood of that fraud. Here, the scikit-learn library is used to evaluate the likelihood of fraud using random forests and logistic regression. For example, the likelihood of fraud is scored based on keywords such as "money" and "help" in the text.

[0595] Generate and send alerts

[0596] Server: If the likelihood of fraud exceeds a certain threshold, a warning message is generated for the user. The generated warning message is sent to the user's device and notified in real time. For example, a warning message such as "This call may be fraudulent" is generated and sent to the user's smartphone via push notification.

[0597] Feedback and Learning

[0598] User: Provide feedback on whether it was a scam. For example, after a call ends, the system asks the user, "Was this call a scam?" and the user answers Yes / No.

[0599] Server: Improves the machine learning model based on the feedback information collected from users. This process improves the accuracy of the next detection. This process is performed using TensorFlow and PyTorch. For example, the model is retrained using the collected feedback data.

[0600] Specific examples

[0601] Example 1: Detecting fraudulent calls

[0602] User: An elderly person receives a call from someone claiming to be their "son."

[0603] Device: Captures the call contents in real time through the smartphone's microphone and sends the audio data to the server.

[0604] Server: Converts the received voice data into text and uses natural language processing technology to detect signs of fraud, such as "I need money" or "Help me."

[0605] Server: Scores the likelihood of fraud and generates a warning message if it indicates a high probability.

[0606] Device: Display a warning message on the elderly person's device saying "This call may be fraudulent."

[0607] User: Realizes the call is a scam and hangs up.

[0608] User: Provide feedback to the system that it was a scam.

[0609] Server: Improves machine learning models based on feedback.

[0610] Example 2: Message fraud detection

[0611] Users: Young people receive suspicious emails.

[0612] Terminal: Captures email text in real time and sends it to the server.

[0613] Server: Analyzes text data using natural language processing technology to detect signs of fraud, such as "urgent" or "click here."

[0614] Server: A scoring algorithm assesses the likelihood of fraud and generates a warning message if a high probability is detected.

[0615] Devices: Display a warning message on young people's devices saying, "This email may be fraudulent."

[0616] Users: Ignore the email and avoid becoming a victim of fraud.

[0617] User: Provide feedback that the email was fraudulent.

[0618] Server: Improve machine learning models based on feedback.

[0619] Prompt Sentence Examples

[0620] Prompt Sentence Example 1

[0621] "Is this call potentially fraudulent? Please analyze the following call: 'Mom, I need money. Please send me 100,000 yen right away.'"

[0622] Prompt Sentence Example 2

[0623] "Do you think the following email might be fraudulent? Email: 'URGENT! There has been a suspicious login to your account. Click here to check.'"

[0624] In this way, the system of the present invention can effectively prevent fraud damage through real-time detection and warning of special frauds.

[0625] The flow of the identification process in the first embodiment will be described with reference to FIG.

[0626] Step 1:

[0627] Device: Monitors user communications in real time. For calls, the device captures audio data using the smartphone's microphone, and for messages, it obtains text data from the application. For example, the device can automatically start recording as soon as a call begins, capturing the content of the conversation.

[0628] Input: Call voice or text message

[0629] Output: Captured audio or text data

[0630] Step 2:

[0631] On the device: The captured audio data is denoised and the sampling rate is adjusted. This is done using the Librosa library. Specifically, the audio file is read, and the librosa.effects.reduce_noise function is used to remove noise and adjust the sampling rate.

[0632] Input: Captured audio data

[0633] Output: Denoised and sample-rate adjusted audio data

[0634] Step 3:

[0635] Terminal: Encrypt the preprocessed data using AES encryption technology, for example, using the pycryptodome library to perform AES encryption.

[0636] Input: Denoised and sample-rate adjusted audio data

[0637] Output: Encrypted audio data

[0638] Step 4:

[0639] On the device, the encrypted data is sent to the server, specifically using the HTTP / HTTPS protocol to send the data securely.

[0640] Input: Encrypted audio or text data

[0641] Output: Data sent to the server

[0642] Step 5:

[0643] Server: Decodes the data received from the device and then begins analysis. In the case of voice data, it first converts it into text data using voice recognition technology. This is done using a general-purpose voice recognition API.

[0644] Input: Encrypted data sent to the server

[0645] Output: Decoded audio and text data

[0646] Step 6:

[0647] Server: Analyzes text data using natural language processing techniques to detect signs of fraud. Specifically, it uses spaCy and NLTK to analyze keywords and context within the text.

[0648] Input: Decoded audio and text data

[0649] Output: Analysis results showing signs of fraud

[0650] Step 7:

[0651] Server: If indicators of fraud are detected from the text data, a scoring algorithm is used to calculate the probability of fraud. Random forest and logistic regression models are used using the scikit-learn library.

[0652] Input: Analysis results indicating fraud

[0653] Output: Scoring results

[0654] Step 8:

[0655] Server: If the likelihood of fraud exceeds a certain threshold, a warning message is generated and sent to the user's device. For example, a message saying "This call may be fraudulent" is created and sent to the user's smartphone.

[0656] Input: Scoring results

[0657] Output: The warning message sent to the user.

[0658] Step 9:

[0659] User: Provides feedback on whether the communication was in fact a scam. After the call ends, the system asks the user, "Was this call a scam?" and the user answers Yes / No.

[0660] Input: Communication feedback question

[0661] Output: Feedback information

[0662] Step 10:

[0663] Server: Based on the feedback information collected from users, the machine learning model is retrained and improved, which will improve the detection accuracy next time. This process is performed using TensorFlow and PyTorch.

[0664] Input: Feedback information

[0665] Output: An improved machine learning model

[0666] (Application example 1)

[0667] Next, a description will be given of Application Example 1. In the following description, the data processing device 12 will be referred to as a "server" and the smart glasses 214 will be referred to as a "terminal."

[0668] Currently, there are many cases of special frauds targeting the elderly and young people, and an effective system to prevent these frauds is needed. However, current fraud detection systems often only notify users on their devices, which makes it easy for users to overlook the warning, and few systems can immediately alert users. Therefore, there is a need for a system that can detect signs of fraud in real time and immediately and effectively warn users.

[0669] The specific processing by the specific processing unit 290 of the data processing device 12 in the application example 1 is realized by the following means.

[0670] In this invention, the server includes means for monitoring the content of user communications and capturing data in real time, means for transmitting the captured data to the server, means for analyzing the data received by the server and detecting the possibility of fraud, means for generating a warning based on the detected possibility of fraud and notifying the user, means for collecting feedback information and improving the analysis means based on the feedback information, and means for displaying a warning message on a head-mounted display in real time. This allows the user to visually confirm fraud warnings in real time, effectively preventing damage from special frauds.

[0671] "User communication content" refers to the content of calls and messages made by the user.

[0672] "Means for capturing data in real time" refers to devices or methods that instantly capture the content of ongoing communications.

[0673] The "means for transmitting to the server" refers to a means for sending the captured data to the central processing unit via a network.

[0674] "Means for analyzing and detecting potential fraud" refers to devices or methods for analyzing acquired data and detecting signs of fraudulent activity therein.

[0675] "Means for generating a warning and notifying a user" refers to means for creating a warning message and notifying a user when a potential fraud is detected.

[0676] "Means for collecting feedback information" refers to means for collecting opinions and evaluations from users.

[0677] "Means for improving analysis means" refers to means for improving data analysis methods based on collected feedback information.

[0678] A "head-mounted display" refers to a device worn by a user that displays information visually.

[0679] "Means for displaying a warning message in real time" refers to means for immediately visually conveying the detected warning content to the user.

[0680] The present invention is a system for monitoring user communications in real time, detecting signs of fraud, and issuing a warning. An embodiment of this system will be described in detail below.

[0681] Overall structure

[0682] This system consists of a user's terminal, a server that analyzes the data, and a head-mounted display (HMD) that notifies the user of warnings.

[0683] The server includes means for monitoring user communications and capturing data in real time, means for transmitting the captured data to the server, means for analyzing the received data to detect possible fraud, means for generating and notifying the user of a warning based on possible fraud, and means for collecting feedback information and improving the analysis means.

[0684] Implementation Details

[0685] 1. A way to monitor user communications and capture data in real time

[0686] The user's device monitors the contents of calls and messages in real time. For calls, audio data is captured using the HMD's built-in microphone. For messages, text data is obtained directly. This data is processed using the AudioCapture API (for audio data) and the TextMessage API (for text data).

[0687] 2. Data preprocessing and transmission

[0688] The captured data is preprocessed on the user's device. Specifically, audio data is subjected to noise removal (NoiseSuppressionLib) and sampling rate adjustment (SoundProcessingLib), and text data is preprocessed using a natural language processing library (e.g., spaCy). The data is then encrypted (SSLTLS) and sent to the server via the HTTPS protocol.

[0689] 3. Data Analysis

[0690] The server runs on a high-performance server (e.g., AWS EC2). The data received on the server side is converted into text using a speech recognition engine (Google Speech-to-Text API), and then analyzed using a natural language processing engine (BERT). This analysis detects signs of fraud.

[0691] 4. Alert Generation and Notification

[0692] The server evaluates the likelihood of fraud using a scoring algorithm (XGBoost). If there is a high likelihood of fraud, a warning message is generated. This warning message is sent in real time to the user's HMD and displayed using the Notification API.

[0693] 5. Feedback and Improvement

[0694] Users provide feedback on the resulting warnings, which is collected on the server and used to improve the analysis method using a machine learning platform (e.g., AWS SageMaker).

[0695] Examples of specific examples and prompts

[0696] As a concrete example, consider a case where an elderly person receives a call from someone claiming to be their "son" saying, "I urgently need money to pay for my hospital bills." In this case, the system will operate according to the following prompt:

[0697] Prompt statement:

[0698] 1. Audio data capture and preprocessing:

[0699] Capture the user's voice in real time through the microphone, remove noise with NoiseSuppressionLib, and adjust the sampling rate with SoundProcessingLib.

[0700] 2. Data transmission:

[0701] The preprocessed data is encrypted using SSLTLS and sent to the server using HTTPS.

[0702] 3. Data Analysis:

[0703] On the server side, the speech is converted to text using the Google Speech-to-Text API and analyzed using BERT.

[0704] 4. Alert Generation and Notification:

[0705] XGBoost scores the likelihood of fraud and generates a warning message, which is displayed on the HMD using the Notification API.

[0706] 5. Feedback Processing:

[0707] Collect user feedback and improve machine learning models with AWS SageMaker.

[0708] This configuration allows users to receive immediate warnings about the risk of special fraud, making it possible to effectively prevent fraud damage.

[0709] The flow of the specific processing in the application example 1 will be described with reference to FIG.

[0710] Step 1:

[0711] The device monitors the user's communication in real time and captures the data. The input is the user's call voice and message content, and the output is the captured raw data. The audio data is obtained through the microphone and processed using the AudioCapture API. The text data is obtained directly using the TextMessage API.

[0712] Step 2:

[0713] Preprocess the data captured by the device. For audio data, use NoiseSuppressionLib to remove noise and SoundProcessingLib to adjust the sampling rate. For text data, use a natural language processing library (e.g. spaCy) for preprocessing. The input is raw data and the output is preprocessed data.

[0714] Step 3:

[0715] The terminal encrypts the preprocessed data and sends it to the server. The input is the preprocessed data and the output is the encrypted data. SSL TLS is used for encryption, and the data is sent via the HTTPS protocol.

[0716] Step 4:

[0717] The server decrypts the received data and performs data analysis. The input is the encrypted data, and the output is the analyzed result. In the case of audio data, the server converts the audio to text using the Google Speech-to-Text API, and then analyzes the text data using BERT.

[0718] Step 5:

[0719] The server scores the likelihood of fraud based on the analysis results and generates a warning message. The input is the analyzed text data, and the output is the warning message. XGBoost is used for scoring.

[0720] Step 6:

[0721] The server generates a warning message and sends it to the user's HMD, where it is displayed using the Notification API. The input is the warning message, and the output is the warning content reflected in the user's visual perception.

[0722] Step 7:

[0723] The user acknowledges the warning message and provides feedback. The input is the user's feedback information, and the output is the collected feedback information.

[0724] Step 8:

[0725] The server improves the analysis method based on the feedback information collected. The input is the feedback information, and the output is the improved analysis method. The feedback information is reflected in the machine learning model using AWS SageMaker.

[0726] Furthermore, an emotion engine that estimates the user's emotion may be further combined. That is, the identification processing unit 290 may estimate the user's emotion using the emotion identification model 59, and perform identification processing using the user's emotion.

[0727] This invention realizes more advanced fraud detection by incorporating an emotion engine into a system that uses the power of AI to eradicate special frauds. The system not only monitors users' communications in real time and detects signs of fraud, but also analyzes users' emotions to more accurately assess the possibility of fraud and issue warnings.

[0728] Overall system overview

[0729] 1. Monitoring of outgoing information

[0730] Terminal: Monitors user calls and messages in real time. For calls, the microphone is used to capture audio data, and for messages, the text data is directly acquired.

[0731] 2. Data preprocessing and transmission

[0732] Terminal: For audio data, preprocessing such as noise removal and sampling rate adjustment is performed. The preprocessed data is encrypted and sent to the server.

[0733] 3. Data Analysis

[0734] Server: Analyzes the received data. If it is voice data, it is first converted into text data using speech recognition technology. Next, it analyzes the text data using natural language processing technology to detect signs of fraud.

[0735] 4. Emotional awareness and regulation

[0736] Server: The emotion engine uses the analyzed text and voice data to recognize the user's emotions. The emotion is analyzed based on the tone of voice and the use of words.

[0737] 5. Fraud Indicator Detection and Scoring

[0738] Server: If fraud indicators are detected, adjust the fraud scoring based on the output of the emotion engine, for example increasing the score if the user is feeling anxious or scared.

[0739] 6. Generating and Sending Alerts

[0740] Server: Generates a warning message when the likelihood of fraud exceeds a certain threshold, and sends the generated warning message to the user's device to notify them in real time.

[0741] 7. Feedback and Learning

[0742] User: Provide feedback on whether it is actually a scam.

[0743] Server: Use this feedback to improve the machine learning model and make the next detection more accurate.

[0744] Specific examples

[0745] Example 1: Detecting fraudulent calls using an emotion engine

[0746] User: An elderly person receives a call from someone claiming to be their "son."

[0747] Terminal: Captures the call in real time and sends the audio data to the server.

[0748] Server: Converts voice data into text and uses natural language processing technology to detect signs of fraud, such as "I need money" or "Help me."

[0749] Server: Recognizes when an elderly person is feeling anxious or scared from text and speech analyzed using an emotion engine.

[0750] Server: Score the probability of fraud, increase the score taking into account the results of sentiment analysis, and generate a warning message if the probability is high.

[0751] Devices: Display a warning message on the senior's device to notify them of potential fraud.

[0752] User: Realizes the call is a scam and hangs up.

[0753] User: Provide feedback to the system that it was a scam.

[0754] Server: Improves machine learning models based on feedback.

[0755] Example 2: Detecting message fraud using an emotion engine

[0756] User: A young person receives a suspicious email.

[0757] Terminal: Captures email text in real time and sends it to the server.

[0758] Server: Analyzes text data using natural language processing technology to detect signs of fraud, such as "urgent," "important," and "click here."

[0759] Server: Recognizes from text analyzed using an emotion engine that young people are feeling confused or excited.

[0760] Server: Score the probability of fraud, increase the score taking into account the results of sentiment analysis, and generate a warning message if the probability is high.

[0761] Devices: Display warning messages on young people's devices to inform them of potential scams.

[0762] User: Ignore the email to prevent unauthorized access.

[0763] User: Provide feedback that the email was fraudulent.

[0764] Server: Improve machine learning models based on feedback.

[0765] In this way, this system can effectively prevent fraud damage through real-time detection and warning of special frauds, and even by recognizing user emotions.

[0766] The processing flow will be explained below.

[0767] Step 1:

[0768] User: Makes calls, sends and receives messages, and exchanges necessary information, including potentially fraudulent communications.

[0769] Step 2:

[0770] Terminal: Captures the contents of calls as audio data in real time, and also captures text messages as text data.

[0771] Step 3:

[0772] Terminal: Preprocessing the captured audio data, such as noise removal and unifying the sampling rate, improves the quality of the audio data.

[0773] Step 4:

[0774] Terminal: The pre-processed data is encrypted and sent to the server using a secure communication protocol, along with the communication metadata (source, destination, time, etc.).

[0775] Step 5:

[0776] Server: Receives data sent from the device. In the case of voice data, converts it into text using voice recognition technology. Converts the voice content into text information using a voice recognition engine.

[0777] Step 6:

[0778] Server: Analyzes the received text data using natural language processing (NLP) techniques, such as tokenizing words, tagging parts of speech, and grammar analysis, to detect phrases and structures that may be indicative of fraud.

[0779] Step 7:

[0780] Server: Uses an emotion engine to recognize user emotions from text data. Analyzes user emotions (e.g., anxiety, anger, confusion, etc.) based on voice tone and phrasing.

[0781] Step 8:

[0782] Server: If fraud signs are detected, the server scores the likelihood of fraud by taking into account the user's emotions analyzed by the emotion engine. The likelihood of fraud is adjusted so that the score is higher if the user is feeling anxious or angry.

[0783] Step 9:

[0784] Server: Generates a warning message if the fraud score exceeds a set threshold, including the likelihood of fraud and what to do about it.

[0785] Step 10:

[0786] Server: Sends the generated warning message to the user's device, notifying them in real time and alerting them.

[0787] Step 11:

[0788] Device: Display an alert message on the user's device, using a notification sound or vibration to get the user's attention.

[0789] Step 12:

[0790] Users: Check whether the call or message is fraudulent and take the action outlined in the warning message. If it is fraudulent, immediately stop the transaction or call.

[0791] Step 13:

[0792] Users: Provide feedback to the system on whether the fraud is true or not. User judgment is also important data for improving the system.

[0793] Step 14:

[0794] Device: Sends user feedback information to the server, including feedback content (whether it was a scam, the effectiveness of countermeasures, etc.).

[0795] Step 15:

[0796] Server: Receives feedback information and uses it as data to improve the performance of the machine learning model and emotion engine. Based on the feedback, the detection algorithm is retrained.

[0797] Through this process, the system can detect signs of specialized fraud in real time, and by utilizing an emotion engine, it can accurately assess the possibility of fraud and issue a warning.

[0798] Example 2

[0799] Next, a description will be given of Example 2. In the following description, the data processing device 12 will be referred to as a "server" and the smart glasses 214 will be referred to as a "terminal."

[0800] In recent years, special frauds have been on the rise, with the damage targeting elderly people and beginners becoming particularly severe. Conventional fraud prevention systems are capable of detecting signs of fraud, but they have limitations in the timing and accuracy of issuing warnings. Furthermore, because they do not take the user's emotional state into account, they have the problem of easily missing important warnings. Therefore, there is a need for a system that can more accurately detect potential fraud and issue warnings promptly by monitoring communication content in real time and analyzing the user's emotions.

[0801] The specific processing by the specific processing unit 290 of the data processing device 12 in the second embodiment is realized by the following means.

[0802] In this invention, the server includes means for monitoring user communications and capturing data in real time, means for preprocessing and encrypting the captured data and transmitting it to the server, means for analyzing the data received by the server and detecting the possibility of fraud, means for recognizing an emotional state from the analyzed data and voice data, means for scoring fraud using the recognized emotional state, means for generating a warning based on the detected possibility of fraud and scoring data and notifying the user, and means for collecting feedback information and improving the analysis means based on the feedback information, thereby enabling more accurate fraud detection and faster warnings that reflect the user's emotional state.

[0803] "User communications" refers to digital communications such as calls, messages, and emails sent or received by a user.

[0804] "Means for capturing data in real time" refers to a piece of equipment or software that acquires the contents of user communications in real time and processes them immediately.

[0805] "Preprocessing" refers to initial processing such as noise removal and sampling rate adjustment that is performed to convert captured data into a form that is easier to analyze.

[0806] "Encryption" is the technology of converting data into cryptographic code to protect the privacy and security of communication data.

[0807] A "server" is a central computer system that receives data from multiple users over a network and processes and analyzes it.

[0808] "Means for detecting potential fraud" are algorithms and technologies that analyze the content of communications to find signs of fraud.

[0809] "Means for recognizing emotional states" refers to technology that analyzes the tone and phrasing of a user's voice or text to identify the emotions the user is feeling.

[0810] A "fraud scoring means" is a technology that numerically calculates the likelihood of fraud based on detected indicators of fraud and the user's emotional state.

[0811] "Means for generating a warning and notifying the user" refers to technology that automatically creates a warning message when a high possibility of fraud is determined and sends it to the user's device.

[0812] "Feedback information" refers to data on reactions and evaluations provided by users in response to warnings issued by the system.

[0813] "Means to improve analysis means" refers to using feedback information to improve fraud detection algorithms and emotion recognition techniques, thereby increasing the accuracy of the overall system.

[0814] The present invention is a system that uses the power of AI to eradicate special frauds, and by incorporating an emotion engine, achieves more advanced fraud detection. The system not only monitors user communications in real time and detects signs of fraud, but also analyzes user emotions to more accurately assess the possibility of fraud and issue warnings. Specific embodiments of the present invention are described below.

[0815] 1. Monitoring of outgoing information

[0816] Device: Monitors user calls and messages in real time. For calls, audio data is captured using the device's built-in microphone, and for messages, text data is directly acquired.

[0817] 2. Data preprocessing and transmission

[0818] Terminal: For audio data, first preprocessing such as noise removal and sampling rate adjustment is performed, and the preprocessed data is encrypted using the AES-256 encryption algorithm and sent to the server via the SSL / TLS protocol.

[0819] 3. Speech data conversion and natural language processing

[0820] Server: Receives the voice data and converts it to text using the Google Cloud Speech-to-Text API. The converted text data is analyzed using a natural language processing library (e.g., spaCy or BERT) to detect signs of fraud, such as "I need money" or "Help me."

[0821] 4. Emotion Recognition and Analysis

[0822] Server: Inputs text and voice data into Affectiva's emotion engine to analyze the user's emotional state, analyzing voice tone and phrasing to identify emotions such as "anxiety," "fear," and "confusion."

[0823] 5. Fraud Indicator Scoring

[0824] Server: Calculates the fraud score based on the detected fraud indicators and the results of sentiment analysis. If the user shows high anxiety or fear, the fraud score is increased.

[0825] 6. Generating and Sending Warning Messages

[0826] Server: When the fraud score exceeds the set threshold, a warning message is automatically generated and sent to the user's device immediately, where the device notifies the user.

[0827] 7. Gather feedback and learn

[0828] User: After receiving the warning message, the user can provide feedback on whether it was indeed a scam, for example, by replying "It was a scam" or "It wasn't a scam."

[0829] Server: Stores the feedback information in a database and updates the machine learning model (e.g., Scikit-learn or TensorFlow) to improve the accuracy of fraud detection next time.

[0830] Specific examples

[0831] Example 1: Detecting fraudulent calls using an emotion engine

[0832] User: An elderly person receives a call from someone claiming to be their "son."

[0833] Terminal: Captures the call in real time and sends the audio data to the server.

[0834] Server: Converts voice data into text and uses natural language processing technology to detect signs of fraud, such as "I need money" or "Help me."

[0835] Server: Recognizes when an elderly person is feeling anxious or scared from text and speech analyzed using an emotion engine.

[0836] Server: Score the likelihood of fraud, increase the score based on the results of sentiment analysis, and generate a warning message if the probability is high.

[0837] Devices: Display a warning message on the senior's device to notify them of potential fraud.

[0838] User: Realizes the call is a scam and hangs up.

[0839] User: Provide feedback to the system that it was a scam.

[0840] Server: Improves machine learning models based on feedback.

[0841] Example prompts to be input to the generative AI model

[0842] "Someone claiming to be my son said he needed money. Please analyze the possibility that this call is fraudulent. I have the audio data and pre-processed text data."

[0843] The flow of the identification process in the second embodiment will be described with reference to FIG.

[0844] Step 1:

[0845] Terminal: Monitors user calls and messages in real time. The input is the user's voice or text message, and the output is captured voice data or text data. For voice, the voice data is captured using the microphone built into the terminal. For messages, the text data is obtained as is.

[0846] Step 2:

[0847] Terminal: Preprocesses the captured audio data. The input is the captured audio data, and the output is the preprocessed audio data. This includes applying noise reduction filters and adjusting the sampling rate. After preprocessing, the data is encrypted using the AES-256 encryption algorithm and sent to the server using the SSL / TLS protocol.

[0848] Step 3:

[0849] Server: Converts audio data to text data. The input is preprocessed audio data and the output is text data. The audio data is converted to text using the Google Cloud Speech-to-Text API. This text data is sent to the next processing step.

[0850] Step 4:

[0851] Server: Analyzes the text data using natural language processing techniques. The input is text data, and the output is the analysis results containing information about signs of fraud. Natural language processing libraries (e.g., spaCy and BERT) are used to detect fraudulent keywords and phrases such as "I need money" and "Help me."

[0852] Step 5:

[0853] Server: Analyzes the user's emotional state using an emotion engine. The input is text data and voice data, and the output is data indicating the user's emotional state. Using Affectiva's emotion engine, it identifies emotions such as "anxiety," "fear," and "confusion" from voice tone and vocabulary.

[0854] Step 6:

[0855] Server: Performs fraud scoring. The input is fraud indicator information and emotional state data, and the output is a fraud score. Based on the detected fraud indicators and the user's emotional state, the possibility of fraud is numerically evaluated. If the emotional state is "anxiety" or "fear," the fraud score increases.

[0856] Step 7:

[0857] Server: Generates a warning message and notifies the user. The input is the fraud score and the output is a warning message. If the fraud score exceeds the set threshold, a warning message is automatically generated. This message is immediately sent to the user's device and the user is notified on the device.

[0858] Step 8:

[0859] User: Provides feedback. The input is the user's reaction to the warning, and the output is feedback information. The user reports whether the warning was correct or not, replying to the system "it was a scam" or "it wasn't a scam."

[0860] Step 9:

[0861] Server: Updates the machine learning model based on the feedback information. The input is the feedback information, and the output is an improved analysis model. The received feedback is stored in a database and used to improve the model using machine learning algorithms (e.g., Scikit-learn or TensorFlow) to improve the fraud detection accuracy next time.

[0862] (Application example 2)

[0863] Next, a description will be given of Application Example 2. In the following description, the data processing device 12 will be referred to as a "server" and the smart glasses 214 will be referred to as a "terminal."

[0864] Conventional fraud detection systems only detect signs of fraud contained in user communications, and improving detection accuracy by reflecting the user's emotional state is a challenge. Furthermore, there is a lack of real-time warnings and feedback-based model improvements, limiting the effectiveness of fraud prevention.

[0865] The specific processing by the specific processing unit 290 of the data processing device 12 in the application example 2 is realized by the following means.

[0866] In this invention, the server includes means for monitoring user communications and detecting signs of fraud in real time, means for analyzing user emotions using an emotion engine and adjusting the fraud score, means for generating a warning based on the likelihood of fraud and notifying the user, and means for collecting feedback information and improving the analysis means, thereby enabling advanced fraud detection that takes user emotions into account, real-time warnings, and continuous system improvement based on feedback.

[0867] "User communication content" refers to data such as voice calls, messages, and emails exchanged between the user and others.

[0868] "Means of capturing data in real time" refers to technology that instantly collects communication content and stores it as data for later processing.

[0869] A "server" refers to a computer system that receives data from multiple client devices via a network and analyzes and processes the data.

[0870] "Measures to detect potential fraud" refers to algorithms or technologies used to determine whether a user's communications contain fraudulent activity.

[0871] "Means for recognizing emotions and adjusting fraud scores" refers to technology that analyzes a user's emotional state and modifies the indicators used to assess the likelihood of fraud.

[0872] "Means for generating an alert and notifying the user" refers to technology that creates an alert message based on the detected potential fraud and sends it to the user's device.

[0873] "Means for collecting feedback information and improving analytical methods" refers to technologies that collect user feedback and use that information to improve fraud detection algorithms and the overall performance of the system.

[0874] "Means for converting voice data into text" refers to the process of converting voice data into character data using voice recognition technology.

[0875] "Natural language processing technology" refers to technology for analyzing the structure and meaning of language based on text data.

[0876] "Emotion analysis technology using generative models" refers to technology that uses machine learning models to estimate user emotions from text data and voice data.

[0877] This invention is a system that monitors the content of user communications in real time, analyzes signs of fraud and user emotions, and thereby detects the possibility of fraud with high accuracy and issues a warning.

[0878] System configuration

[0879] Hardware

[0880] Devices: Microphones and processing units built into smartphones, smart glasses, and head-mounted displays

[0881] Server: Cloud server capable of high-performance analysis and data processing

[0882] software

[0883] Audio data capture: sounddevice library

[0884] Audio data preprocessing: librosa library

[0885] Speech Recognition: TensorFlow Model

[0886] Sentiment Analysis: Hugging Face transformers library

[0887] Processing Details

[0888] Monitoring user communications and capturing data

[0889] The device monitors the user's calls and messages in real time and captures voice data. In the case of a call, the voice data is collected using a microphone and pre-processed, such as noise reduction and sampling rate adjustment.

[0890] Data transmission and analysis

[0891] The pre-processed data is encrypted and then sent to the server, which then converts the received voice data into text using natural language processing technology.

[0892] Sentiment Analysis and Fraud Scoring

[0893] On the server side, a generative AI model is used to analyze user sentiment from text and voice data, and the fraud score is adjusted based on the analyzed sentiment.

[0894] Generate alerts and notify users

[0895] If the likelihood of fraud exceeds a certain threshold, the server generates a warning message and notifies the user's terminal in real time.

[0896] Feedback and machine learning model improvement

[0897] Users provide feedback to the system, which the server uses to improve the machine learning model and make the next detection more accurate.

[0898] Specific examples

[0899] For example, if an elderly person receives a phone call from someone claiming to be their "son," the following steps may occur:

[0900] 1. The elderly person's device captures the contents of the call in real time and sends it to the server.

[0901] 2. The server converts the voice data into text and uses natural language processing technology to analyze signs of fraud, such as "I need money."

[0902] 3. Furthermore, generative AI models are used to analyze the emotions of older adults (e.g., anxiety and fear) from text and voice data.

[0903] 4. If it is determined that there is a high possibility of fraud, a warning message will be displayed on the elderly person's device.

[0904] Prompt Sentence Examples

[0905] "Please answer the following questions based on your call data:

[0906] 1. Are there any signs of fraud in this conversation?

[0907] 2. What are the user's emotions?

[0908] example:

[0909] Call text: 'Mom, help me! I need money now.'"

[0910] In this way, the invention allows for advanced fraud detection that takes user emotions into account, real-time alerts, and continuous system improvement based on feedback.

[0911] The flow of the specific processing in the application example 2 will be described with reference to FIG.

[0912] Step 1:

[0913] The device monitors the user's communication in real time and captures audio data. Specifically, it collects audio data using a microphone and obtains it through an audio data capture library (e.g., sounddevice). The input is real-time audio, and the output is captured audio data.

[0914] Step 2:

[0915] The device preprocesses the captured audio data. Specifically, it performs operations such as noise reduction and sampling rate adjustment. This process uses an audio data preprocessing library (e.g., librosa). The input is the captured audio data, and the output is the preprocessed audio data.

[0916] Step 3:

[0917] The device encrypts the preprocessed audio data and sends it to the server. Specifically, it uses an encryption library to securely convert the data and send it to the server over the network. The preprocessed audio data is the input, and the encrypted data is sent to the server as the output.

[0918] Step 4:

[0919] The server analyzes the received voice data. Specifically, it uses speech recognition technology (e.g., TensorFlow model) to convert the voice data into text data. The input is encrypted voice data, and the output is text data.

[0920] Step 5:

[0921] The server analyzes the text data using natural language processing technology. Specifically, it uses natural language processing technology (e.g., a generative AI model) to detect signs of fraud from the text data. The input is text data, and the output is a determination of the likelihood of fraud.

[0922] Step 6:

[0923] The server recognizes emotions from the analyzed data and adjusts the fraud score. Specifically, it analyzes the user's emotions using emotion analysis technology (e.g., Hugging Face's transformers library) and updates the fraud score. The input is the analyzed text data and the emotion analysis result, and the output is the adjusted fraud score.

[0924] Step 7:

[0925] The server generates a warning message and notifies the user if the fraud probability exceeds a certain threshold. Specifically, the server automatically generates a warning message and sends it to the user's device. The input is the adjusted fraud score, and the output is the warning message displayed on the user's device.

[0926] Step 8:

[0927] The user provides feedback on whether or not the fraud is actually occurring. Specifically, the user reports whether or not the fraud is occurring within the application, and the feedback is sent to the server. The input is the user's feedback information, and the output is the feedback data provided to the server.

[0928] Step 9:

[0929] The server improves the machine learning model based on the feedback information. Specifically, it analyzes the collected feedback information and runs an algorithm to improve the performance of the fraud detection model. The input is the feedback information and the output is an improved machine learning model.

[0930] The specific processing unit 290 transmits the result of the specific processing to the smart glasses 214. In the smart glasses 214, the control unit 46A causes the speaker 240 to output the result of the specific processing. The microphone 238 acquires audio indicating a user input regarding the result of the specific processing. The control unit 46A transmits audio data indicating the user input acquired by the microphone 238 to the data processing device 12. In the data processing device 12, the specific processing unit 290 acquires the audio data.

[0931] The data generation model 58 is a so-called generative AI (Artificial Intelligence). An example of the data generation model 58 is ChatGPT (Internet Search<URL: https: / / openai.com / blog / chatgpt> ), Gemini (Internet search <url: https: gemini.google.com ?hl="ja">) and other generation AIs. The data generation model 58 is obtained by performing deep learning on a neural network. A prompt including an instruction is input to the data generation model 58, and inference data such as voice data indicating voice, text data indicating text, and image data indicating an image is also input. The data generation model 58 performs inference on the input inference data in accordance with the instruction indicated by the prompt, and outputs the inference result in a data format such as voice data and text data. Here, inference refers to, for example, analysis, classification, prediction, and / or summarization.

[0932] In the above embodiment, an example in which the specific processing is performed by the data processing device 12 has been given, but the technology of the present disclosure is not limited to this, and the specific processing may be performed by the smart glasses 214.

[0933] [Third embodiment]

[0934] FIG. 5 shows an example of the configuration of a data processing system 310 according to the third embodiment.

[0935] 5, the data processing system 310 includes the data processing device 12 and a headset type terminal 314. An example of the data processing device 12 is a server.

[0936] The data processing device 12 includes a computer 22, a database 24, and a communication I / F 26. The computer 22 is an example of a "computer" according to the technology of the present disclosure. The computer 22 includes a processor 28, a RAM 30, and a storage 32. The processor 28, the RAM 30, and the storage 32 are connected to a bus 34. The database 24 and the communication I / F 26 are also connected to the bus 34. The communication I / F 26 is connected to a network 54. Examples of the network 54 include a WAN (Wide Area Network) and / or a LAN (Local Area Network).

[0937] The headset type terminal 314 includes a computer 36, a microphone 238, a speaker 240, a camera 42, a communication I / F 44, and a display 343. The computer 36 includes a processor 46, a RAM 48, and a storage 50. The processor 46, the RAM 48, and the storage 50 are connected to a bus 52. The microphone 238, the speaker 240, the camera 42, and the display 343 are also connected to the bus 52.

[0938] The microphone 238 receives instructions and the like from the user 20 by receiving voice uttered by the user 20. The microphone 238 captures the voice uttered by the user 20, converts the captured voice into audio data, and outputs it to the processor 46. The speaker 240 outputs audio in accordance with instructions from the processor 46.

[0939] Camera 42 is a small digital camera equipped with an optical system including a lens, aperture, and shutter, and an imaging element such as a CMOS (Complementary Metal-Oxide-Semiconductor) image sensor or a CCD (Charge Coupled Device) image sensor, and captures images of the surroundings of user 20 (for example, an imaging range defined by an angle of view equivalent to the field of vision of a typical healthy person).

[0940] The communication I / F 44 is connected to a network 54. The communication I / Fs 44 and 26 are responsible for the exchange of various information between the processor 46 and the processor 28 via the network 54. The exchange of various information between the processor 46 and the processor 28 using the communication I / Fs 44 and 26 is carried out in a secure state.

[0941] Fig. 6 shows an example of the main functions of the data processing device 12 and the headset type terminal 314. As shown in Fig. 6, in the data processing device 12, a specific process is performed by the processor 28. A specific process program 56 is stored in the storage 32.

[0942] The specific processing program 56 is an example of a "program" according to the technology of the present disclosure. The processor 28 reads the specific processing program 56 from the storage 32 and executes the read specific processing program 56 on the RAM 30. The specific processing is realized by the processor 28 operating as a specific processing unit 290 in accordance with the specific processing program 56 executed on the RAM 30.

[0943] The storage 32 stores a data generation model 58 and an emotion identification model 59. The data generation model 58 and the emotion identification model 59 are used by the identification processing unit 290.

[0944] In the headset type terminal 314, a reception output process is performed by the processor 46. A reception output program 60 is stored in the storage 50. The processor 46 reads the reception output program 60 from the storage 50 and executes the read reception output program 60 on the RAM 48. The reception output process is realized by the processor 46 operating as the control unit 46A in accordance with the reception output program 60 executed on the RAM 48.

[0945] Next, a description will be given of the identification process performed by the identification processing unit 290 of the data processing device 12. In the following description, the data processing device 12 will be referred to as the "server" and the headset type terminal 314 will be referred to as the "terminal."

[0946] This invention is a system that uses the power of AI to eradicate special frauds. It monitors users' communications in real time, detects signs of fraud, and issues warnings to effectively prevent fraud damage. This system is configured around the user's terminal and server.

[0947] Overall system overview

[0948] 1. Monitoring of outgoing information

[0949] Terminal: Monitors user calls and messages in real time. For calls, the microphone is used to capture audio data, and for messages, the text data is directly acquired.

[0950] 2. Data preprocessing and transmission

[0951] Terminal: For audio data, preprocessing such as noise removal and sampling rate adjustment is performed. The preprocessed data is encrypted and sent to the server.

[0952] 3. Data Analysis

[0953] Server: Analyzes the received data. In the case of voice data, it is first converted into text data using speech recognition technology. Next, it analyzes the text data using natural language processing technology to detect signs of fraud.

[0954] 4. Fraud detection and scoring

[0955] Server: If any fraud indications are detected, a scoring algorithm is used to calculate the likelihood of fraud. Based on this score, the likelihood of fraud is assessed.

[0956] 5. Generating and Sending Alerts

[0957] Server: Generates a warning message when the likelihood of fraud exceeds a certain threshold, and sends the generated warning message to the user's device to notify them in real time.

[0958] 6. Feedback and Learning

[0959] User: Provide feedback on whether it is actually a scam.

[0960] Server: Use this feedback to improve the machine learning model and make the next detection more accurate.

[0961] Specific examples

[0962] Example 1: Detecting fraudulent calls

[0963] User: An elderly person receives a call from someone claiming to be their "son."

[0964] Terminal: Captures the call in real time and sends the audio data to the server.

[0965] Server: Converts voice data into text and uses natural language processing technology to detect signs of fraud, such as "I need money" or "Help me."

[0966] Server: Scores the likelihood of fraud and generates a warning message if it indicates a high probability.

[0967] Devices: Display a warning message on the senior's device to notify them of potential fraud.

[0968] User: Realizes the call is a scam and hangs up.

[0969] User: Provide feedback to the system that it was a scam.

[0970] Server: Improves machine learning models based on feedback.

[0971] Example 2: Message fraud detection

[0972] User: A young person receives a suspicious email.

[0973] Terminal: Captures email text in real time and sends it to the server.

[0974] Server: Analyzes text data using natural language processing technology to detect signs of fraud, such as "urgent" or "click here."

[0975] Server: A scoring algorithm assesses the likelihood of fraud and generates a warning message if it indicates a high probability.

[0976] Devices: Display warning messages on young people's devices to inform them of potential scams.

[0977] User: Ignore the email to prevent unauthorized access.

[0978] User: Provide feedback that the email was fraudulent.

[0979] Server: Improve machine learning models based on feedback.

[0980] In this way, this system can effectively prevent fraud damage through real-time detection and warning of special frauds.

[0981] The processing flow will be explained below.

[0982] Step 1:

[0983] Users: Make and receive calls and messages, some of which may be potentially fraudulent.

[0984] Step 2:

[0985] Terminal: Captures the user's call content as audio data in real time, and also captures the message content directly in the case of text messages.

[0986] Step 3:

[0987] Terminal: For audio data, preprocessing is performed, such as noise removal and unifying the sampling rate. The preprocessed data is prepared for transmission to the server.

[0988] Step 4:

[0989] Terminal: Encrypts the pre-processed data and sends it to the server using a secure communication protocol.

[0990] Step 5:

[0991] Server: Receives data sent from the device, including voice and text data.

[0992] Step 6:

[0993] Server: Converts the received voice data into text data using voice recognition technology. A voice recognition engine is used for this purpose.

[0994] Step 7:

[0995] Server: Analyzes text data using natural language processing (NLP) techniques, such as tokenizing words, tagging parts of speech, and grammar analysis.

[0996] Step 8:

[0997] Server: Detects signs of fraud from the analyzed data, using pre-trained AI models to identify phrases and contexts associated with fraud.

[0998] Step 9:

[0999] Server: Based on the detected fraud indicators, a scoring algorithm is used to assess the likelihood of fraud, using techniques such as logistic regression or support vector machines (SVM).

[1000] Step 10:

[1001] Server: If the score exceeds a set threshold, generate a warning message containing information about the potential for fraud and what to do about it.

[1002] Step 11:

[1003] Server: Sends the generated warning message to the user's terminal.

[1004] Step 12:

[1005] Device: Display an alert message on the user's device, using a notification sound or vibration to get the user's immediate attention.

[1006] Step 13:

[1007] User: Review the warning message and decide how to respond, including whether it is actually a scam.

[1008] Step 14:

[1009] User: Provide feedback to the system as to whether it was a scam or not.

[1010] Step 15:

[1011] Device: Sends user feedback to the server.

[1012] Step 16:

[1013] Server: Receives feedback information and uses it to refine the machine learning model, leveraging a feedback loop to continuously improve the accuracy of the analysis module.

[1014] Through this process, the system can detect signs of special fraud in real time and effectively warn users.

[1015] Example 1

[1016] Next, a description will be given of Example 1. In the following description, the data processing device 12 will be referred to as a "server" and the headset type terminal 314 will be referred to as a "terminal."

[1017] Recently, special frauds have become increasingly sophisticated, increasing the risk of many people becoming victims of fraud. Elderly people and users unfamiliar with the Internet are particularly vulnerable to fraud. Current systems have difficulty detecting fraudulent activities in advance and issuing effective warnings, so effective prevention measures are needed.

[1018] The specific processing by the specific processing unit 290 of the data processing device 12 in the first embodiment is realized by the following means.

[1019] In this invention, the server includes means for monitoring user communications in real time and capturing the data as voice or text data, preprocessing means for performing noise reduction and sampling rate adjustment on the captured data, means for encrypting the preprocessed data and transmitting it to the server, means for analyzing the data received by the server and detecting the possibility of fraud, means for scoring the possibility of fraud based on the analyzed data, means for generating a warning message and notifying the user when the possibility of fraud exceeds a certain threshold, and means for collecting feedback information from users and improving the analysis means based on the feedback information, thereby making it possible to detect fraudulent activity in real time and quickly issue a warning to the user.

[1020] "User" refers to any individual or organization that uses this system.

[1021] "Communication content" refers to data such as voice calls and text messages sent by users.

[1022] "Real-time" refers to processing data almost as soon as the communication content occurs.

[1023] "Means for capturing data" refers to hardware or software for acquiring the contents of user communications.

[1024] "Noise reduction" refers to the process of reducing unnecessary noise from audio data.

[1025] "Sampling rate adjustment" refers to digital processing to convert audio data into a format that is optimal for analysis.

[1026] "Preprocessing means" refers to a system that processes captured data to prepare it in a format suitable for analysis.

[1027] "Means for encrypting data" refers to technology that encrypts transmitted data so that it cannot be deciphered by third parties.

[1028] "Server" refers to a central computing device that receives, analyzes, and processes captured data.

[1029] "Means for analyzing data" refers to technology used to analyze received data and detect signs of fraud.

[1030] "Speech recognition technology" refers to technology for converting voice data into text data.

[1031] "Natural language processing technology" refers to technology for analyzing text data and understanding its meaning and context.

[1032] "Means for scoring the likelihood of fraud" refers to a technology that numerically evaluates the likelihood of fraud from analyzed data.

[1033] "Threshold" refers to a benchmark for assessing the likelihood of fraud.

[1034] "Means for generating a warning message" refers to a technique for creating a message to warn a user when there is a high possibility of fraud.

[1035] "Feedback information" refers to information provided by a user as to whether a communication was fraudulent.

[1036] "Means for improving analytics" refers to techniques for using collected feedback information to improve fraud detection models.

[1037] This invention is a system that uses the power of AI to eradicate special frauds. It monitors users' communications in real time, detects signs of fraud, and issues warnings to effectively prevent fraud damage. This system is configured around the user's terminal and server.

[1038] Specific implementation methods

[1039] Monitoring user communications

[1040] Device: The user's device monitors calls and messages in real time. For calls, the device captures audio data using the smartphone's microphone, and for messages, the device obtains text data from an application on the device. For example, voice calls and text messages are monitored via a communication application on the smartphone.

[1041] Data preprocessing and transmission

[1042] On the device: The audio data is denoised and the sampling rate is adjusted. This process is performed using the Librosa library. The pre-processed data is then encrypted using AES encryption technology. The encrypted data is then securely sent to the server. For example, Librosa functions are used to denoise the audio from a call, and then the audio is encrypted using AES.

[1043] Data analysis

[1044] Server: Receives data sent from the device and begins analysis. In the case of voice data, it is first converted into text data using voice recognition technology. A general-purpose voice recognition API is used here. Next, the text data is analyzed using natural language processing technology to detect signs of fraud. Libraries such as spaCy and NLTK are used for this process. For example, voice data is converted into text using a recognition API, and the text data is then analyzed using spaCy.

[1045] Fraud indicator detection and scoring

[1046] Server: If indicators of fraud are detected from the text data, a scoring algorithm is used to calculate the likelihood of that fraud. Here, the scikit-learn library is used to evaluate the likelihood of fraud using random forests and logistic regression. For example, the likelihood of fraud is scored based on keywords such as "money" and "help" in the text.

[1047] Generate and send alerts

[1048] Server: If the likelihood of fraud exceeds a certain threshold, a warning message is generated for the user. The generated warning message is sent to the user's device and notified in real time. For example, a warning message such as "This call may be fraudulent" is generated and sent to the user's smartphone via push notification.

[1049] Feedback and Learning

[1050] User: Provide feedback on whether it was a scam. For example, after a call ends, the system asks the user, "Was this call a scam?" and the user answers Yes / No.

[1051] Server: Improves the machine learning model based on the feedback information collected from users. This process improves the accuracy of the next detection. This process is performed using TensorFlow and PyTorch. For example, the model is retrained using the collected feedback data.

[1052] Specific examples

[1053] Example 1: Detecting fraudulent calls

[1054] User: An elderly person receives a call from someone claiming to be their "son."

[1055] Device: Captures the call contents in real time through the smartphone's microphone and sends the audio data to the server.

[1056] Server: Converts the received voice data into text and uses natural language processing technology to detect signs of fraud, such as "I need money" or "Help me."

[1057] Server: Scores the likelihood of fraud and generates a warning message if it indicates a high probability.

[1058] Device: Display a warning message on the elderly person's device saying "This call may be fraudulent."

[1059] User: Realizes the call is a scam and hangs up.

[1060] User: Provide feedback to the system that it was a scam.

[1061] Server: Improves machine learning models based on feedback.

[1062] Example 2: Message fraud detection

[1063] Users: Young people receive suspicious emails.

[1064] Terminal: Captures email text in real time and sends it to the server.

[1065] Server: Analyzes text data using natural language processing technology to detect signs of fraud, such as "urgent" or "click here."

[1066] Server: A scoring algorithm assesses the likelihood of fraud and generates a warning message if a high probability is detected.

[1067] Devices: Display a warning message on young people's devices saying, "This email may be fraudulent."

[1068] Users: Ignore the email and avoid becoming a victim of fraud.

[1069] User: Provide feedback that the email was fraudulent.

[1070] Server: Improve machine learning models based on feedback.

[1071] Prompt Sentence Examples

[1072] Prompt Sentence Example 1

[1073] "Is this call potentially fraudulent? Please analyze the following call: 'Mom, I need money. Please send me 100,000 yen right away.'"

[1074] Prompt Sentence Example 2

[1075] "Do you think the following email might be fraudulent? Email: 'URGENT! There has been a suspicious login to your account. Click here to check.'"

[1076] In this way, the system of the present invention can effectively prevent fraud damage through real-time detection and warning of special frauds.

[1077] The flow of the identification process in the first embodiment will be described with reference to FIG.

[1078] Step 1:

[1079] Device: Monitors user communications in real time. For calls, the device captures audio data using the smartphone's microphone, and for messages, it obtains text data from the application. For example, the device can automatically start recording as soon as a call begins, capturing the content of the conversation.

[1080] Input: Call voice or text message

[1081] Output: Captured audio or text data

[1082] Step 2:

[1083] On the device: The captured audio data is denoised and the sampling rate is adjusted. This is done using the Librosa library. Specifically, the audio file is read, and the librosa.effects.reduce_noise function is used to remove noise and adjust the sampling rate.

[1084] Input: Captured audio data

[1085] Output: Denoised and sample-rate adjusted audio data

[1086] Step 3:

[1087] Terminal: Encrypt the preprocessed data using AES encryption technology, for example, using the pycryptodome library to perform AES encryption.

[1088] Input: Denoised and sample-rate adjusted audio data

[1089] Output: Encrypted audio data

[1090] Step 4:

[1091] On the device, the encrypted data is sent to the server, specifically using the HTTP / HTTPS protocol to send the data securely.

[1092] Input: Encrypted audio or text data

[1093] Output: Data sent to the server

[1094] Step 5:

[1095] Server: Decodes the data received from the device and then begins analysis. In the case of voice data, it first converts it into text data using voice recognition technology. This is done using a general-purpose voice recognition API.

[1096] Input: Encrypted data sent to the server

[1097] Output: Decoded audio and text data

[1098] Step 6:

[1099] Server: Analyzes text data using natural language processing techniques to detect signs of fraud. Specifically, it uses spaCy and NLTK to analyze keywords and context within the text.

[1100] Input: Decoded audio and text data

[1101] Output: Analysis results showing signs of fraud

[1102] Step 7:

[1103] Server: If indicators of fraud are detected from the text data, a scoring algorithm is used to calculate the probability of fraud. Random forest and logistic regression models are used using the scikit-learn library.

[1104] Input: Analysis results indicating fraud

[1105] Output: Scoring results

[1106] Step 8:

[1107] Server: If the likelihood of fraud exceeds a certain threshold, a warning message is generated and sent to the user's device. For example, a message saying "This call may be fraudulent" is created and sent to the user's smartphone.

[1108] Input: Scoring results

[1109] Output: The warning message sent to the user.

[1110] Step 9:

[1111] User: Provides feedback on whether the communication was in fact a scam. After the call ends, the system asks the user, "Was this call a scam?" and the user answers Yes / No.

[1112] Input: Communication feedback question

[1113] Output: Feedback information

[1114] Step 10:

[1115] Server: Based on the feedback information collected from users, the machine learning model is retrained and improved, which will improve the detection accuracy next time. This process is performed using TensorFlow and PyTorch.

[1116] Input: Feedback information

[1117] Output: An improved machine learning model

[1118] (Application example 1)

[1119] Next, a description will be given of Application Example 1. In the following description, the data processing device 12 will be referred to as a "server" and the headset type terminal 314 will be referred to as a "terminal."

[1120] Currently, there are many cases of special frauds targeting the elderly and young people, and an effective system to prevent these frauds is needed. However, current fraud detection systems often only notify users on their devices, which makes it easy for users to overlook the warning, and few systems can immediately alert users. Therefore, there is a need for a system that can detect signs of fraud in real time and immediately and effectively warn users.

[1121] The specific processing by the specific processing unit 290 of the data processing device 12 in the application example 1 is realized by the following means.

[1122] In this invention, the server includes means for monitoring the content of user communications and capturing data in real time, means for transmitting the captured data to the server, means for analyzing the data received by the server and detecting the possibility of fraud, means for generating a warning based on the detected possibility of fraud and notifying the user, means for collecting feedback information and improving the analysis means based on the feedback information, and means for displaying a warning message on a head-mounted display in real time. This allows the user to visually confirm fraud warnings in real time, effectively preventing damage from special frauds.

[1123] "User communication content" refers to the content of calls and messages made by the user.

[1124] "Means for capturing data in real time" refers to devices or methods that instantly capture the content of ongoing communications.

[1125] The "means for transmitting to the server" refers to a means for sending the captured data to the central processing unit via a network.

[1126] "Means for analyzing and detecting potential fraud" refers to devices or methods for analyzing acquired data and detecting signs of fraudulent activity therein.

[1127] "Means for generating a warning and notifying a user" refers to means for creating a warning message and notifying a user when a potential fraud is detected.

[1128] "Means for collecting feedback information" refers to means for collecting opinions and evaluations from users.

[1129] "Means for improving analysis means" refers to means for improving data analysis methods based on collected feedback information.

[1130] A "head-mounted display" refers to a device worn by a user that displays information visually.

[1131] "Means for displaying a warning message in real time" refers to means for immediately visually conveying the detected warning content to the user.

[1132] The present invention is a system for monitoring user communications in real time, detecting signs of fraud, and issuing a warning. An embodiment of this system will be described in detail below.

[1133] Overall structure

[1134] This system consists of a user's terminal, a server that analyzes the data, and a head-mounted display (HMD) that notifies the user of warnings.

[1135] The server includes means for monitoring user communications and capturing data in real time, means for transmitting the captured data to the server, means for analyzing the received data to detect possible fraud, means for generating and notifying the user of a warning based on possible fraud, and means for collecting feedback information and improving the analysis means.

[1136] Implementation Details

[1137] 1. A way to monitor user communications and capture data in real time

[1138] The user's device monitors the contents of calls and messages in real time. For calls, audio data is captured using the HMD's built-in microphone. For messages, text data is obtained directly. This data is processed using the AudioCapture API (for audio data) and the TextMessage API (for text data).

[1139] 2. Data preprocessing and transmission

[1140] The captured data is preprocessed on the user's device. Specifically, audio data is subjected to noise removal (NoiseSuppressionLib) and sampling rate adjustment (SoundProcessingLib), and text data is preprocessed using a natural language processing library (e.g., spaCy). The data is then encrypted (SSLTLS) and sent to the server via the HTTPS protocol.

[1141] 3. Data Analysis

[1142] The server runs on a high-performance server (e.g., AWS EC2). The data received on the server side is converted into text using a speech recognition engine (Google Speech-to-Text API), and then analyzed using a natural language processing engine (BERT). This analysis detects signs of fraud.

[1143] 4. Alert Generation and Notification

[1144] The server evaluates the likelihood of fraud using a scoring algorithm (XGBoost). If there is a high likelihood of fraud, a warning message is generated. This warning message is sent in real time to the user's HMD and displayed using the Notification API.

[1145] 5. Feedback and Improvement

[1146] Users provide feedback on the resulting warnings, which is collected on the server and used to improve the analysis method using a machine learning platform (e.g., AWS SageMaker).

[1147] Examples of specific examples and prompts

[1148] As a concrete example, consider a case where an elderly person receives a call from someone claiming to be their "son" saying, "I urgently need money to pay for my hospital bills." In this case, the system will operate according to the following prompt:

[1149] Prompt statement:

[1150] 1. Audio data capture and preprocessing:

[1151] Capture the user's voice in real time through the microphone, remove noise with NoiseSuppressionLib, and adjust the sampling rate with SoundProcessingLib.

[1152] 2. Data transmission:

[1153] The preprocessed data is encrypted using SSLTLS and sent to the server using HTTPS.

[1154] 3. Data Analysis:

[1155] On the server side, the speech is converted to text using the Google Speech-to-Text API and analyzed using BERT.

[1156] 4. Alert Generation and Notification:

[1157] XGBoost scores the likelihood of fraud and generates a warning message, which is displayed on the HMD using the Notification API.

[1158] 5. Feedback Processing:

[1159] Collect user feedback and improve machine learning models with AWS SageMaker.

[1160] This configuration allows users to receive immediate warnings about the risk of special fraud, making it possible to effectively prevent fraud damage.

[1161] The flow of the specific processing in the application example 1 will be described with reference to FIG.

[1162] Step 1:

[1163] The device monitors the user's communication in real time and captures the data. The input is the user's call voice and message content, and the output is the captured raw data. The audio data is obtained through the microphone and processed using the AudioCapture API. The text data is obtained directly using the TextMessage API.

[1164] Step 2:

[1165] Preprocess the data captured by the device. For audio data, use NoiseSuppressionLib to remove noise and SoundProcessingLib to adjust the sampling rate. For text data, use a natural language processing library (e.g. spaCy) for preprocessing. The input is raw data and the output is preprocessed data.

[1166] Step 3:

[1167] The terminal encrypts the preprocessed data and sends it to the server. The input is the preprocessed data and the output is the encrypted data. SSL TLS is used for encryption, and the data is sent via the HTTPS protocol.

[1168] Step 4:

[1169] The server decrypts the received data and performs data analysis. The input is the encrypted data, and the output is the analyzed result. In the case of audio data, the server converts the audio to text using the Google Speech-to-Text API, and then analyzes the text data using BERT.

[1170] Step 5:

[1171] The server scores the likelihood of fraud based on the analysis results and generates a warning message. The input is the analyzed text data, and the output is the warning message. XGBoost is used for scoring.

[1172] Step 6:

[1173] The server generates a warning message and sends it to the user's HMD, where it is displayed using the Notification API. The input is the warning message, and the output is the warning content reflected in the user's visual perception.

[1174] Step 7:

[1175] The user acknowledges the warning message and provides feedback. The input is the user's feedback information, and the output is the collected feedback information.

[1176] Step 8:

[1177] The server improves the analysis method based on the feedback information collected. The input is the feedback information, and the output is the improved analysis method. The feedback information is reflected in the machine learning model using AWS SageMaker.

[1178] Furthermore, an emotion engine that estimates the user's emotion may be further combined. That is, the identification processing unit 290 may estimate the user's emotion using the emotion identification model 59, and perform identification processing using the user's emotion.

[1179] This invention realizes more advanced fraud detection by incorporating an emotion engine into a system that uses the power of AI to eradicate special frauds. The system not only monitors users' communications in real time and detects signs of fraud, but also analyzes users' emotions to more accurately assess the possibility of fraud and issue warnings.

[1180] Overall system overview

[1181] 1. Monitoring of outgoing information

[1182] Terminal: Monitors user calls and messages in real time. For calls, the microphone is used to capture audio data, and for messages, the text data is directly acquired.

[1183] 2. Data preprocessing and transmission

[1184] Terminal: For audio data, preprocessing such as noise removal and sampling rate adjustment is performed. The preprocessed data is encrypted and sent to the server.

[1185] 3. Data Analysis

[1186] Server: Analyzes the received data. If it is voice data, it is first converted into text data using speech recognition technology. Next, it analyzes the text data using natural language processing technology to detect signs of fraud.

[1187] 4. Emotional awareness and regulation

[1188] Server: The emotion engine uses the analyzed text and voice data to recognize the user's emotions. The emotion is analyzed based on the tone of voice and the use of words.

[1189] 5. Fraud Indicator Detection and Scoring

[1190] Server: If fraud indicators are detected, adjust the fraud scoring based on the output of the emotion engine, for example increasing the score if the user is feeling anxious or scared.

[1191] 6. Generating and Sending Alerts

[1192] Server: Generates a warning message when the likelihood of fraud exceeds a certain threshold, and sends the generated warning message to the user's device to notify them in real time.

[1193] 7. Feedback and Learning

[1194] User: Provide feedback on whether it is actually a scam.

[1195] Server: Use this feedback to improve the machine learning model and make the next detection more accurate.

[1196] Specific examples

[1197] Example 1: Detecting fraudulent calls using an emotion engine

[1198] User: An elderly person receives a call from someone claiming to be their "son."

[1199] Terminal: Captures the call in real time and sends the audio data to the server.

[1200] Server: Converts voice data into text and uses natural language processing technology to detect signs of fraud, such as "I need money" or "Help me."

[1201] Server: Recognizes when an elderly person is feeling anxious or scared from text and speech analyzed using an emotion engine.

[1202] Server: Score the probability of fraud, increase the score taking into account the results of sentiment analysis, and generate a warning message if the probability is high.

[1203] Devices: Display a warning message on the senior's device to notify them of potential fraud.

[1204] User: Realizes the call is a scam and hangs up.

[1205] User: Provide feedback to the system that it was a scam.

[1206] Server: Improves machine learning models based on feedback.

[1207] Example 2: Detecting message fraud using an emotion engine

[1208] User: A young person receives a suspicious email.

[1209] Terminal: Captures email text in real time and sends it to the server.

[1210] Server: Analyzes text data using natural language processing technology to detect signs of fraud, such as "urgent," "important," and "click here."

[1211] Server: Recognizes from text analyzed using an emotion engine that young people are feeling confused or excited.

[1212] Server: Score the probability of fraud, increase the score taking into account the results of sentiment analysis, and generate a warning message if the probability is high.

[1213] Devices: Display warning messages on young people's devices to inform them of potential scams.

[1214] User: Ignore the email to prevent unauthorized access.

[1215] User: Provide feedback that the email was fraudulent.

[1216] Server: Improve machine learning models based on feedback.

[1217] In this way, this system can effectively prevent fraud damage through real-time detection and warning of special frauds, and even by recognizing user emotions.

[1218] The processing flow will be explained below.

[1219] Step 1:

[1220] User: Makes calls, sends and receives messages, and exchanges necessary information, including potentially fraudulent communications.

[1221] Step 2:

[1222] Terminal: Captures the contents of calls as audio data in real time, and also captures text messages as text data.

[1223] Step 3:

[1224] Terminal: Preprocessing the captured audio data, such as noise removal and unifying the sampling rate, improves the quality of the audio data.

[1225] Step 4:

[1226] Terminal: The pre-processed data is encrypted and sent to the server using a secure communication protocol, along with the communication metadata (source, destination, time, etc.).

[1227] Step 5:

[1228] Server: Receives data sent from the device. In the case of voice data, converts it into text using voice recognition technology. Converts the voice content into text information using a voice recognition engine.

[1229] Step 6:

[1230] Server: Analyzes the received text data using natural language processing (NLP) techniques, such as tokenizing words, tagging parts of speech, and grammar analysis, to detect phrases and structures that may be indicative of fraud.

[1231] Step 7:

[1232] Server: Uses an emotion engine to recognize user emotions from text data. Analyzes user emotions (e.g., anxiety, anger, confusion, etc.) based on voice tone and phrasing.

[1233] Step 8:

[1234] Server: If fraud signs are detected, the server scores the likelihood of fraud by taking into account the user's emotions analyzed by the emotion engine. The likelihood of fraud is adjusted so that the score is higher if the user is feeling anxious or angry.

[1235] Step 9:

[1236] Server: Generates a warning message if the fraud score exceeds a set threshold, including the likelihood of fraud and what to do about it.

[1237] Step 10:

[1238] Server: Sends the generated warning message to the user's device, notifying them in real time and alerting them.

[1239] Step 11:

[1240] Device: Display an alert message on the user's device, using a notification sound or vibration to get the user's attention.

[1241] Step 12:

[1242] Users: Check whether the call or message is fraudulent and take the action outlined in the warning message. If it is fraudulent, immediately stop the transaction or call.

[1243] Step 13:

[1244] Users: Provide feedback to the system on whether the fraud is true or not. User judgment is also important data for improving the system.

[1245] Step 14:

[1246] Device: Sends user feedback information to the server, including feedback content (whether it was a scam, the effectiveness of countermeasures, etc.).

[1247] Step 15:

[1248] Server: Receives feedback information and uses it as data to improve the performance of the machine learning model and emotion engine. Based on the feedback, the detection algorithm is retrained.

[1249] Through this process, the system can detect signs of specialized fraud in real time, and by utilizing an emotion engine, it can accurately assess the possibility of fraud and issue a warning.

[1250] Example 2

[1251] Next, a description will be given of Example 2. In the following description, the data processing device 12 will be referred to as a "server" and the headset type terminal 314 will be referred to as a "terminal."

[1252] In recent years, special frauds have been on the rise, with the damage targeting elderly people and beginners becoming particularly severe. Conventional fraud prevention systems are capable of detecting signs of fraud, but they have limitations in the timing and accuracy of issuing warnings. Furthermore, because they do not take the user's emotional state into account, they have the problem of easily missing important warnings. Therefore, there is a need for a system that can more accurately detect potential fraud and issue warnings promptly by monitoring communication content in real time and analyzing the user's emotions.

[1253] The specific processing by the specific processing unit 290 of the data processing device 12 in the second embodiment is realized by the following means.

[1254] In this invention, the server includes means for monitoring user communications and capturing data in real time, means for preprocessing and encrypting the captured data and transmitting it to the server, means for analyzing the data received by the server and detecting the possibility of fraud, means for recognizing an emotional state from the analyzed data and voice data, means for scoring fraud using the recognized emotional state, means for generating a warning based on the detected possibility of fraud and scoring data and notifying the user, and means for collecting feedback information and improving the analysis means based on the feedback information, thereby enabling more accurate fraud detection and faster warnings that reflect the user's emotional state.

[1255] "User communications" refers to digital communications such as calls, messages, and emails sent or received by a user.

[1256] "Means for capturing data in real time" refers to a piece of equipment or software that acquires the contents of user communications in real time and processes them immediately.

[1257] "Preprocessing" refers to initial processing such as noise removal and sampling rate adjustment that is performed to convert captured data into a form that is easier to analyze.

[1258] "Encryption" is the technology of converting data into cryptographic code to protect the privacy and security of communication data.

[1259] A "server" is a central computer system that receives data from multiple users over a network and processes and analyzes it.

[1260] "Means for detecting potential fraud" are algorithms and technologies that analyze the content of communications to find signs of fraud.

[1261] "Means for recognizing emotional states" refers to technology that analyzes the tone and phrasing of a user's voice or text to identify the emotions the user is feeling.

[1262] A "fraud scoring means" is a technology that numerically calculates the likelihood of fraud based on detected indicators of fraud and the user's emotional state.

[1263] "Means for generating a warning and notifying the user" refers to technology that automatically creates a warning message when a high possibility of fraud is determined and sends it to the user's device.

[1264] "Feedback information" refers to data on reactions and evaluations provided by users in response to warnings issued by the system.

[1265] "Means to improve analysis means" refers to using feedback information to improve fraud detection algorithms and emotion recognition techniques, thereby increasing the accuracy of the overall system.

[1266] The present invention is a system that uses the power of AI to eradicate special frauds, and by incorporating an emotion engine, achieves more advanced fraud detection. The system not only monitors user communications in real time and detects signs of fraud, but also analyzes user emotions to more accurately assess the possibility of fraud and issue warnings. Specific embodiments of the present invention are described below.

[1267] 1. Monitoring of outgoing information

[1268] Device: Monitors user calls and messages in real time. For calls, audio data is captured using the device's built-in microphone, and for messages, text data is directly acquired.

[1269] 2. Data preprocessing and transmission

[1270] Terminal: For audio data, first preprocessing such as noise removal and sampling rate adjustment is performed, and the preprocessed data is encrypted using the AES-256 encryption algorithm and sent to the server via the SSL / TLS protocol.

[1271] 3. Speech data conversion and natural language processing

[1272] Server: Receives the voice data and converts it to text using the Google Cloud Speech-to-Text API. The converted text data is analyzed using a natural language processing library (e.g., spaCy or BERT) to detect signs of fraud, such as "I need money" or "Help me."

[1273] 4. Emotion Recognition and Analysis

[1274] Server: Inputs text and voice data into Affectiva's emotion engine to analyze the user's emotional state, analyzing voice tone and phrasing to identify emotions such as "anxiety," "fear," and "confusion."

[1275] 5. Fraud Indicator Scoring

[1276] Server: Calculates the fraud score based on the detected fraud indicators and the results of sentiment analysis. If the user shows high anxiety or fear, the fraud score is increased.

[1277] 6. Generating and Sending Warning Messages

[1278] Server: When the fraud score exceeds the set threshold, a warning message is automatically generated and sent to the user's device immediately, where the device notifies the user.

[1279] 7. Gather feedback and learn

[1280] User: After receiving the warning message, the user can provide feedback on whether it was indeed a scam, for example, by replying "It was a scam" or "It wasn't a scam."

[1281] Server: Stores the feedback information in a database and updates the machine learning model (e.g., Scikit-learn or TensorFlow) to improve the accuracy of fraud detection next time.

[1282] Specific examples

[1283] Example 1: Detecting fraudulent calls using an emotion engine

[1284] User: An elderly person receives a call from someone claiming to be their "son."

[1285] Terminal: Captures the call in real time and sends the audio data to the server.

[1286] Server: Converts voice data into text and uses natural language processing technology to detect signs of fraud, such as "I need money" or "Help me."

[1287] Server: Recognizes when an elderly person is feeling anxious or scared from text and speech analyzed using an emotion engine.

[1288] Server: Score the likelihood of fraud, increase the score based on the results of sentiment analysis, and generate a warning message if the probability is high.

[1289] Devices: Display a warning message on the senior's device to notify them of potential fraud.

[1290] User: Realizes the call is a scam and hangs up.

[1291] User: Provide feedback to the system that it was a scam.

[1292] Server: Improves machine learning models based on feedback.

[1293] Example prompts to be input to the generative AI model

[1294] "Someone claiming to be my son said he needed money. Please analyze the possibility that this call is fraudulent. I have the audio data and pre-processed text data."

[1295] The flow of the identification process in the second embodiment will be described with reference to FIG.

[1296] Step 1:

[1297] Terminal: Monitors user calls and messages in real time. The input is the user's voice or text message, and the output is captured voice data or text data. For voice, the voice data is captured using the microphone built into the terminal. For messages, the text data is obtained as is.

[1298] Step 2:

[1299] Terminal: Preprocesses the captured audio data. The input is the captured audio data, and the output is the preprocessed audio data. This includes applying noise reduction filters and adjusting the sampling rate. After preprocessing, the data is encrypted using the AES-256 encryption algorithm and sent to the server using the SSL / TLS protocol.

[1300] Step 3:

[1301] Server: Converts audio data to text data. The input is preprocessed audio data and the output is text data. The audio data is converted to text using the Google Cloud Speech-to-Text API. This text data is sent to the next processing step.

[1302] Step 4:

[1303] Server: Analyzes the text data using natural language processing techniques. The input is text data, and the output is the analysis results containing information about signs of fraud. Natural language processing libraries (e.g., spaCy and BERT) are used to detect fraudulent keywords and phrases such as "I need money" and "Help me."

[1304] Step 5:

[1305] Server: Analyzes the user's emotional state using an emotion engine. The input is text data and voice data, and the output is data indicating the user's emotional state. Using Affectiva's emotion engine, it identifies emotions such as "anxiety," "fear," and "confusion" from voice tone and vocabulary.

[1306] Step 6:

[1307] Server: Performs fraud scoring. The input is fraud indicator information and emotional state data, and the output is a fraud score. Based on the detected fraud indicators and the user's emotional state, the possibility of fraud is numerically evaluated. If the emotional state is "anxiety" or "fear," the fraud score increases.

[1308] Step 7:

[1309] Server: Generates a warning message and notifies the user. The input is the fraud score and the output is a warning message. If the fraud score exceeds the set threshold, a warning message is automatically generated. This message is immediately sent to the user's device and the user is notified on the device.

[1310] Step 8:

[1311] User: Provides feedback. The input is the user's reaction to the warning, and the output is feedback information. The user reports whether the warning was correct or not, replying to the system "it was a scam" or "it wasn't a scam."

[1312] Step 9:

[1313] Server: Updates the machine learning model based on the feedback information. The input is the feedback information, and the output is an improved analysis model. The received feedback is stored in a database and used to improve the model using machine learning algorithms (e.g., Scikit-learn or TensorFlow) to improve the fraud detection accuracy next time.

[1314] (Application example 2)

[1315] Next, a description will be given of Application Example 2. In the following description, the data processing device 12 will be referred to as a "server" and the headset type terminal 314 will be referred to as a "terminal."

[1316] Conventional fraud detection systems only detect signs of fraud contained in user communications, and improving detection accuracy by reflecting the user's emotional state is a challenge. Furthermore, there is a lack of real-time warnings and feedback-based model improvements, limiting the effectiveness of fraud prevention.

[1317] The specific processing by the specific processing unit 290 of the data processing device 12 in the application example 2 is realized by the following means.

[1318] In this invention, the server includes means for monitoring user communications and detecting signs of fraud in real time, means for analyzing user emotions using an emotion engine and adjusting the fraud score, means for generating a warning based on the likelihood of fraud and notifying the user, and means for collecting feedback information and improving the analysis means, thereby enabling advanced fraud detection that takes user emotions into account, real-time warnings, and continuous system improvement based on feedback.

[1319] "User communication content" refers to data such as voice calls, messages, and emails exchanged between the user and others.

[1320] "Means of capturing data in real time" refers to technology that instantly collects communication content and stores it as data for later processing.

[1321] A "server" refers to a computer system that receives data from multiple client devices via a network and analyzes and processes the data.

[1322] "Measures to detect potential fraud" refers to algorithms or technologies used to determine whether a user's communications contain fraudulent activity.

[1323] "Means for recognizing emotions and adjusting fraud scores" refers to technology that analyzes a user's emotional state and modifies the indicators used to assess the likelihood of fraud.

[1324] "Means for generating an alert and notifying the user" refers to technology that creates an alert message based on the detected potential fraud and sends it to the user's device.

[1325] "Means for collecting feedback information and improving analytical methods" refers to technologies that collect user feedback and use that information to improve fraud detection algorithms and the overall performance of the system.

[1326] "Means for converting voice data into text" refers to the process of converting voice data into character data using voice recognition technology.

[1327] "Natural language processing technology" refers to technology for analyzing the structure and meaning of language based on text data.

[1328] "Emotion analysis technology using generative models" refers to technology that uses machine learning models to estimate user emotions from text data and voice data.

[1329] This invention is a system that monitors the content of user communications in real time, analyzes signs of fraud and user emotions, and thereby detects the possibility of fraud with high accuracy and issues a warning.

[1330] System configuration

[1331] Hardware

[1332] Devices: Microphones and processing units built into smartphones, smart glasses, and head-mounted displays

[1333] Server: Cloud server capable of high-performance analysis and data processing

[1334] software

[1335] Audio data capture: sounddevice library

[1336] Audio data preprocessing: librosa library

[1337] Speech Recognition: TensorFlow Model

[1338] Sentiment Analysis: Hugging Face transformers library

[1339] Processing Details

[1340] Monitoring user communications and capturing data

[1341] The device monitors the user's calls and messages in real time and captures voice data. In the case of a call, the voice data is collected using a microphone and pre-processed, such as noise reduction and sampling rate adjustment.

[1342] Data transmission and analysis

[1343] The pre-processed data is encrypted and then sent to the server, which then converts the received voice data into text using natural language processing technology.

[1344] Sentiment Analysis and Fraud Scoring

[1345] On the server side, a generative AI model is used to analyze user sentiment from text and voice data, and the fraud score is adjusted based on the analyzed sentiment.

[1346] Generate alerts and notify users

[1347] If the likelihood of fraud exceeds a certain threshold, the server generates a warning message and notifies the user's terminal in real time.

[1348] Feedback and machine learning model improvement

[1349] Users provide feedback to the system, which the server uses to improve the machine learning model and make the next detection more accurate.

[1350] Specific examples

[1351] For example, if an elderly person receives a phone call from someone claiming to be their "son," the following steps may occur:

[1352] 1. The elderly person's device captures the contents of the call in real time and sends it to the server.

[1353] 2. The server converts the voice data into text and uses natural language processing technology to analyze signs of fraud, such as "I need money."

[1354] 3. Furthermore, generative AI models are used to analyze the emotions of older adults (e.g., anxiety and fear) from text and voice data.

[1355] 4. If it is determined that there is a high possibility of fraud, a warning message will be displayed on the elderly person's device.

[1356] Prompt Sentence Examples

[1357] "Please answer the following questions based on your call data:

[1358] 1. Are there any signs of fraud in this conversation?

[1359] 2. What are the user's emotions?

[1360] example:

[1361] Call text: 'Mom, help me! I need money now.'"

[1362] In this way, the invention allows for advanced fraud detection that takes user emotions into account, real-time alerts, and continuous system improvement based on feedback.

[1363] The flow of the specific processing in the application example 2 will be described with reference to FIG.

[1364] Step 1:

[1365] The device monitors the user's communication in real time and captures audio data. Specifically, it collects audio data using a microphone and obtains it through an audio data capture library (e.g., sounddevice). The input is real-time audio, and the output is captured audio data.

[1366] Step 2:

[1367] The device preprocesses the captured audio data. Specifically, it performs operations such as noise reduction and sampling rate adjustment. This process uses an audio data preprocessing library (e.g., librosa). The input is the captured audio data, and the output is the preprocessed audio data.

[1368] Step 3:

[1369] The device encrypts the preprocessed audio data and sends it to the server. Specifically, it uses an encryption library to securely convert the data and send it to the server over the network. The preprocessed audio data is the input, and the encrypted data is sent to the server as the output.

[1370] Step 4:

[1371] The server analyzes the received voice data. Specifically, it uses speech recognition technology (e.g., TensorFlow model) to convert the voice data into text data. The input is encrypted voice data, and the output is text data.

[1372] Step 5:

[1373] The server analyzes the text data using natural language processing technology. Specifically, it uses natural language processing technology (e.g., a generative AI model) to detect signs of fraud from the text data. The input is text data, and the output is a determination of the likelihood of fraud.

[1374] Step 6:

[1375] The server recognizes emotions from the analyzed data and adjusts the fraud score. Specifically, it analyzes the user's emotions using emotion analysis technology (e.g., Hugging Face's transformers library) and updates the fraud score. The input is the analyzed text data and the emotion analysis result, and the output is the adjusted fraud score.

[1376] Step 7:

[1377] The server generates a warning message and notifies the user if the fraud probability exceeds a certain threshold. Specifically, the server automatically generates a warning message and sends it to the user's device. The input is the adjusted fraud score, and the output is the warning message displayed on the user's device.

[1378] Step 8:

[1379] The user provides feedback on whether or not the fraud is actually occurring. Specifically, the user reports whether or not the fraud is occurring within the application, and the feedback is sent to the server. The input is the user's feedback information, and the output is the feedback data provided to the server.

[1380] Step 9:

[1381] The server improves the machine learning model based on the feedback information. Specifically, it analyzes the collected feedback information and runs an algorithm to improve the performance of the fraud detection model. The input is the feedback information and the output is an improved machine learning model.

[1382] The specific processing unit 290 transmits the result of the specific processing to the headset type terminal 314. In the headset type terminal 314, the control unit 46A causes the speaker 240 and the display 343 to output the result of the specific processing. The microphone 238 acquires audio indicating a user input regarding the result of the specific processing. The control unit 46A transmits audio data indicating the user input acquired by the microphone 238 to the data processing device 12. In the data processing device 12, the specific processing unit 290 acquires the audio data.

[1383] The data generation model 58 is a so-called generative AI (Artificial Intelligence). An example of the data generation model 58 is ChatGPT (Internet Search<URL: https: / / openai.com / blog / chatgpt> ), Gemini (Internet search <url: https: gemini.google.com ?hl="ja">) and other generation AIs. The data generation model 58 is obtained by performing deep learning on a neural network. A prompt including an instruction is input to the data generation model 58, and inference data such as voice data indicating voice, text data indicating text, and image data indicating an image is also input. The data generation model 58 performs inference on the input inference data in accordance with the instruction indicated by the prompt, and outputs the inference result in a data format such as voice data and text data. Here, inference refers to, for example, analysis, classification, prediction, and / or summarization.

[1384] In the above embodiment, an example was given in which the specific processing is performed by the data processing device 12, but the technology of the present disclosure is not limited to this, and the specific processing may be performed by the headset type terminal 314.

[1385] [Fourth embodiment]

[1386] FIG. 7 shows an example of the configuration of a data processing system 410 according to the fourth embodiment.

[1387] 7, a data processing system 410 includes a data processing device 12 and a robot 414. An example of the data processing device 12 is a server.

[1388] The data processing device 12 includes a computer 22, a database 24, and a communication I / F 26. The computer 22 is an example of a "computer" according to the technology of the present disclosure. The computer 22 includes a processor 28, a RAM 30, and a storage 32. The processor 28, the RAM 30, and the storage 32 are connected to a bus 34. The database 24 and the communication I / F 26 are also connected to the bus 34. The communication I / F 26 is connected to a network 54. Examples of the network 54 include a WAN (Wide Area Network) and / or a LAN (Local Area Network).

[1389] The robot 414 includes a computer 36, a microphone 238, a speaker 240, a camera 42, a communication I / F 44, and a control target 443. The computer 36 includes a processor 46, a RAM 48, and a storage 50. The processor 46, the RAM 48, and the storage 50 are connected to a bus 52. The microphone 238, the speaker 240, the camera 42, and the control target 443 are also connected to the bus 52.

[1390] The microphone 238 receives instructions and the like from the user 20 by receiving voice uttered by the user 20. The microphone 238 captures the voice uttered by the user 20, converts the captured voice into audio data, and outputs it to the processor 46. The speaker 240 outputs audio in accordance with instructions from the processor 46.

[1391] Camera 42 is a small digital camera equipped with an optical system including a lens, aperture, and shutter, and an imaging element such as a CMOS (Complementary Metal-Oxide-Semiconductor) image sensor or a CCD (Charge Coupled Device) image sensor, and captures images of the surroundings of user 20 (for example, an imaging range defined by an angle of view equivalent to the field of vision of a typical healthy person).

[1392] The communication I / F 44 is connected to a network 54. The communication I / Fs 44 and 26 are responsible for the exchange of various information between the processor 46 and the processor 28 via the network 54. The exchange of various information between the processor 46 and the processor 28 using the communication I / Fs 44 and 26 is carried out in a secure state.

[1393] The control object 443 includes a display device, LEDs in the eyes, and motors for driving the arms, hands, and feet. The posture and gestures of the robot 414 are controlled by controlling the motors of the arms, hands, and feet. Some of the emotions of the robot 414 can be expressed by controlling these motors. In addition, the facial expressions of the robot 414 can also be expressed by controlling the light emission state of the LEDs in the eyes of the robot 414.

[1394] Fig. 8 shows an example of the main functions of the data processing device 12 and the robot 414. As shown in Fig. 8, in the data processing device 12, a specific process is performed by the processor 28. A specific process program 56 is stored in the storage 32.

[1395] The specific processing program 56 is an example of a "program" according to the technology of the present disclosure. The processor 28 reads the specific processing program 56 from the storage 32 and executes the read specific processing program 56 on the RAM 30. The specific processing is realized by the processor 28 operating as a specific processing unit 290 in accordance with the specific processing program 56 executed on the RAM 30.

[1396] The storage 32 stores a data generation model 58 and an emotion identification model 59. The data generation model 58 and the emotion identification model 59 are used by the identification processing unit 290.

[1397] In the robot 414, the processor 46 performs the reception output process. A reception output program 60 is stored in the storage 50. The processor 46 reads the reception output program 60 from the storage 50 and executes the read reception output program 60 on the RAM 48. The reception output process is realized by the processor 46 operating as the control unit 46A in accordance with the reception output program 60 executed on the RAM 48.

[1398] Next, a description will be given of the specific processing performed by the specific processing unit 290 of the data processing device 12. In the following description, the data processing device 12 will be referred to as a "server" and the robot 414 will be referred to as a "terminal."

[1399] This invention is a system that uses the power of AI to eradicate special frauds. It monitors users' communications in real time, detects signs of fraud, and issues warnings to effectively prevent fraud damage. This system is configured around the user's terminal and server.

[1400] Overall system overview

[1401] 1. Monitoring of outgoing information

[1402] Terminal: Monitors user calls and messages in real time. For calls, the microphone is used to capture audio data, and for messages, the text data is directly acquired.

[1403] 2. Data preprocessing and transmission

[1404] Terminal: For audio data, preprocessing such as noise removal and sampling rate adjustment is performed. The preprocessed data is encrypted and sent to the server.

[1405] 3. Data Analysis

[1406] Server: Analyzes the received data. In the case of voice data, it is first converted into text data using speech recognition technology. Next, it analyzes the text data using natural language processing technology to detect signs of fraud.

[1407] 4. Fraud detection and scoring

[1408] Server: If any fraud indications are detected, a scoring algorithm is used to calculate the likelihood of fraud. Based on this score, the likelihood of fraud is assessed.

[1409] 5. Generating and Sending Alerts

[1410] Server: Generates a warning message when the likelihood of fraud exceeds a certain threshold, and sends the generated warning message to the user's device to notify them in real time.

[1411] 6. Feedback and Learning

[1412] User: Provide feedback on whether it is actually a scam.

[1413] Server: Use this feedback to improve the machine learning model and make the next detection more accurate.

[1414] Specific examples

[1415] Example 1: Detecting fraudulent calls

[1416] User: An elderly person receives a call from someone claiming to be their "son."

[1417] Terminal: Captures the call in real time and sends the audio data to the server.

[1418] Server: Converts voice data into text and uses natural language processing technology to detect signs of fraud, such as "I need money" or "Help me."

[1419] Server: Scores the likelihood of fraud and generates a warning message if it indicates a high probability.

[1420] Devices: Display a warning message on the senior's device to notify them of potential fraud.

[1421] User: Realizes the call is a scam and hangs up.

[1422] User: Provide feedback to the system that it was a scam.

[1423] Server: Improves machine learning models based on feedback.

[1424] Example 2: Message fraud detection

[1425] User: A young person receives a suspicious email.

[1426] Terminal: Captures email text in real time and sends it to the server.

[1427] Server: Analyzes text data using natural language processing technology to detect signs of fraud, such as "urgent" or "click here."

[1428] Server: A scoring algorithm assesses the likelihood of fraud and generates a warning message if it indicates a high probability.

[1429] Devices: Display warning messages on young people's devices to inform them of potential scams.

[1430] User: Ignore the email to prevent unauthorized access.

[1431] User: Provide feedback that the email was fraudulent.

[1432] Server: Improve machine learning models based on feedback.

[1433] In this way, this system can effectively prevent fraud damage through real-time detection and warning of special frauds.

[1434] The processing flow will be explained below.

[1435] Step 1:

[1436] Users: Make and receive calls and messages, some of which may be potentially fraudulent.

[1437] Step 2:

[1438] Terminal: Captures the user's call content as audio data in real time, and also captures the message content directly in the case of text messages.

[1439] Step 3:

[1440] Terminal: For audio data, preprocessing is performed, such as noise removal and unifying the sampling rate. The preprocessed data is prepared for transmission to the server.

[1441] Step 4:

[1442] Terminal: Encrypts the pre-processed data and sends it to the server using a secure communication protocol.

[1443] Step 5:

[1444] Server: Receives data sent from the device, including voice and text data.

[1445] Step 6:

[1446] Server: Converts the received voice data into text data using voice recognition technology. A voice recognition engine is used for this purpose.

[1447] Step 7:

[1448] Server: Analyzes text data using natural language processing (NLP) techniques, such as tokenizing words, tagging parts of speech, and grammar analysis.

[1449] Step 8:

[1450] Server: Detects signs of fraud from the analyzed data, using pre-trained AI models to identify phrases and contexts associated with fraud.

[1451] Step 9:

[1452] Server: Based on the detected fraud indicators, a scoring algorithm is used to assess the likelihood of fraud, using techniques such as logistic regression or support vector machines (SVM).

[1453] Step 10:

[1454] Server: If the score exceeds a set threshold, generate a warning message containing information about the potential for fraud and what to do about it.

[1455] Step 11:

[1456] Server: Sends the generated warning message to the user's terminal.

[1457] Step 12:

[1458] Device: Display an alert message on the user's device, using a notification sound or vibration to get the user's immediate attention.

[1459] Step 13:

[1460] User: Review the warning message and decide how to respond, including whether it is actually a scam.

[1461] Step 14:

[1462] User: Provide feedback to the system as to whether it was a scam or not.

[1463] Step 15:

[1464] Device: Sends user feedback to the server.

[1465] Step 16:

[1466] Server: Receives feedback information and uses it to refine the machine learning model, leveraging a feedback loop to continuously improve the accuracy of the analysis module.

[1467] Through this process, the system can detect signs of special fraud in real time and effectively warn users.

[1468] Example 1

[1469] Next, a description will be given of Example 1. In the following description, the data processing device 12 will be referred to as a "server" and the robot 414 will be referred to as a "terminal."

[1470] Recently, special frauds have become increasingly sophisticated, increasing the risk of many people becoming victims of fraud. Elderly people and users unfamiliar with the Internet are particularly vulnerable to fraud. Current systems have difficulty detecting fraudulent activities in advance and issuing effective warnings, so effective prevention measures are needed.

[1471] The specific processing by the specific processing unit 290 of the data processing device 12 in the first embodiment is realized by the following means.

[1472] In this invention, the server includes means for monitoring user communications in real time and capturing the data as voice or text data, preprocessing means for performing noise reduction and sampling rate adjustment on the captured data, means for encrypting the preprocessed data and transmitting it to the server, means for analyzing the data received by the server and detecting the possibility of fraud, means for scoring the possibility of fraud based on the analyzed data, means for generating a warning message and notifying the user when the possibility of fraud exceeds a certain threshold, and means for collecting feedback information from users and improving the analysis means based on the feedback information, thereby making it possible to detect fraudulent activity in real time and quickly issue a warning to the user.

[1473] "User" refers to any individual or organization that uses this system.

[1474] "Communication content" refers to data such as voice calls and text messages sent by users.

[1475] "Real-time" refers to processing data almost as soon as the communication content occurs.

[1476] "Means for capturing data" refers to hardware or software for acquiring the contents of user communications.

[1477] "Noise reduction" refers to the process of reducing unnecessary noise from audio data.

[1478] "Sampling rate adjustment" refers to digital processing to convert audio data into a format that is optimal for analysis.

[1479] "Preprocessing means" refers to a system that processes captured data to prepare it in a format suitable for analysis.

[1480] "Means for encrypting data" refers to technology that encrypts transmitted data so that it cannot be deciphered by third parties.

[1481] "Server" refers to a central computing device that receives, analyzes, and processes captured data.

[1482] "Means for analyzing data" refers to technology used to analyze received data and detect signs of fraud.

[1483] "Speech recognition technology" refers to technology for converting voice data into text data.

[1484] "Natural language processing technology" refers to technology for analyzing text data and understanding its meaning and context.

[1485] "Means for scoring the likelihood of fraud" refers to a technology that numerically evaluates the likelihood of fraud from analyzed data.

[1486] "Threshold" refers to a benchmark for assessing the likelihood of fraud.

[1487] "Means for generating a warning message" refers to a technique for creating a message to warn a user when there is a high possibility of fraud.

[1488] "Feedback information" refers to information provided by a user as to whether a communication was fraudulent.

[1489] "Means for improving analytics" refers to techniques for using collected feedback information to improve fraud detection models.

[1490] This invention is a system that uses the power of AI to eradicate special frauds. It monitors users' communications in real time, detects signs of fraud, and issues warnings to effectively prevent fraud damage. This system is configured around the user's terminal and server.

[1491] Specific implementation methods

[1492] Monitoring user communications

[1493] Device: The user's device monitors calls and messages in real time. For calls, the device captures audio data using the smartphone's microphone, and for messages, the device obtains text data from an application on the device. For example, voice calls and text messages are monitored via a communication application on the smartphone.

[1494] Data preprocessing and transmission

[1495] On the device: The audio data is denoised and the sampling rate is adjusted. This process is performed using the Librosa library. The pre-processed data is then encrypted using AES encryption technology. The encrypted data is then securely sent to the server. For example, Librosa functions are used to denoise the audio from a call, and then the audio is encrypted using AES.

[1496] Data analysis

[1497] Server: Receives data sent from the device and begins analysis. In the case of voice data, it is first converted into text data using voice recognition technology. A general-purpose voice recognition API is used here. Next, the text data is analyzed using natural language processing technology to detect signs of fraud. Libraries such as spaCy and NLTK are used for this process. For example, voice data is converted into text using a recognition API, and the text data is then analyzed using spaCy.

[1498] Fraud indicator detection and scoring

[1499] Server: If indicators of fraud are detected from the text data, a scoring algorithm is used to calculate the likelihood of that fraud. Here, the scikit-learn library is used to evaluate the likelihood of fraud using random forests and logistic regression. For example, the likelihood of fraud is scored based on keywords such as "money" and "help" in the text.

[1500] Generate and send alerts

[1501] Server: If the likelihood of fraud exceeds a certain threshold, a warning message is generated for the user. The generated warning message is sent to the user's device and notified in real time. For example, a warning message such as "This call may be fraudulent" is generated and sent to the user's smartphone via push notification.

[1502] Feedback and Learning

[1503] User: Provide feedback on whether it was a scam. For example, after a call ends, the system asks the user, "Was this call a scam?" and the user answers Yes / No.

[1504] Server: Improves the machine learning model based on the feedback information collected from users. This process improves the accuracy of the next detection. This process is performed using TensorFlow and PyTorch. For example, the model is retrained using the collected feedback data.

[1505] Specific examples

[1506] Example 1: Detecting fraudulent calls

[1507] User: An elderly person receives a call from someone claiming to be their "son."

[1508] Device: Captures the call contents in real time through the smartphone's microphone and sends the audio data to the server.

[1509] Server: Converts the received voice data into text and uses natural language processing technology to detect signs of fraud, such as "I need money" or "Help me."

[1510] Server: Scores the likelihood of fraud and generates a warning message if it indicates a high probability.

[1511] Device: Display a warning message on the elderly person's device saying "This call may be fraudulent."

[1512] User: Realizes the call is a scam and hangs up.

[1513] User: Provide feedback to the system that it was a scam.

[1514] Server: Improves machine learning models based on feedback.

[1515] Example 2: Message fraud detection

[1516] Users: Young people receive suspicious emails.

[1517] Terminal: Captures email text in real time and sends it to the server.

[1518] Server: Analyzes text data using natural language processing technology to detect signs of fraud, such as "urgent" or "click here."

[1519] Server: A scoring algorithm assesses the likelihood of fraud and generates a warning message if a high probability is detected.

[1520] Devices: Display a warning message on young people's devices saying, "This email may be fraudulent."

[1521] Users: Ignore the email and avoid becoming a victim of fraud.

[1522] User: Provide feedback that the email was fraudulent.

[1523] Server: Improve machine learning models based on feedback.

[1524] Prompt Sentence Examples

[1525] Prompt Sentence Example 1

[1526] "Is this call potentially fraudulent? Please analyze the following call: 'Mom, I need money. Please send me 100,000 yen right away.'"

[1527] Prompt Sentence Example 2

[1528] "Do you think the following email might be fraudulent? Email: 'URGENT! There has been a suspicious login to your account. Click here to check.'"

[1529] In this way, the system of the present invention can effectively prevent fraud damage through real-time detection and warning of special frauds.

[1530] The flow of the identification process in the first embodiment will be described with reference to FIG.

[1531] Step 1:

[1532] Device: Monitors user communications in real time. For calls, the device captures audio data using the smartphone's microphone, and for messages, it obtains text data from the application. For example, the device can automatically start recording as soon as a call begins, capturing the content of the conversation.

[1533] Input: Call voice or text message

[1534] Output: Captured audio or text data

[1535] Step 2:

[1536] On the device: The captured audio data is denoised and the sampling rate is adjusted. This is done using the Librosa library. Specifically, the audio file is read, and the librosa.effects.reduce_noise function is used to remove noise and adjust the sampling rate.

[1537] Input: Captured audio data

[1538] Output: Denoised and sample-rate adjusted audio data

[1539] Step 3:

[1540] Terminal: Encrypt the preprocessed data using AES encryption technology, for example, using the pycryptodome library to perform AES encryption.

[1541] Input: Denoised and sample-rate adjusted audio data

[1542] Output: Encrypted audio data

[1543] Step 4:

[1544] On the device, the encrypted data is sent to the server, specifically using the HTTP / HTTPS protocol to send the data securely.

[1545] Input: Encrypted audio or text data

[1546] Output: Data sent to the server

[1547] Step 5:

[1548] Server: Decodes the data received from the device and then begins analysis. In the case of voice data, it first converts it into text data using voice recognition technology. This is done using a general-purpose voice recognition API.

[1549] Input: Encrypted data sent to the server

[1550] Output: Decoded audio and text data

[1551] Step 6:

[1552] Server: Analyzes text data using natural language processing techniques to detect signs of fraud. Specifically, it uses spaCy and NLTK to analyze keywords and context within the text.

[1553] Input: Decoded audio and text data

[1554] Output: Analysis results showing signs of fraud

[1555] Step 7:

[1556] Server: If indicators of fraud are detected from the text data, a scoring algorithm is used to calculate the probability of fraud. Random forest and logistic regression models are used using the scikit-learn library.

[1557] Input: Analysis results indicating fraud

[1558] Output: Scoring results

[1559] Step 8:

[1560] Server: If the likelihood of fraud exceeds a certain threshold, a warning message is generated and sent to the user's device. For example, a message saying "This call may be fraudulent" is created and sent to the user's smartphone.

[1561] Input: Scoring results

[1562] Output: The warning message sent to the user.

[1563] Step 9:

[1564] User: Provides feedback on whether the communication was in fact a scam. After the call ends, the system asks the user, "Was this call a scam?" and the user answers Yes / No.

[1565] Input: Communication feedback question

[1566] Output: Feedback information

[1567] Step 10:

[1568] Server: Based on the feedback information collected from users, the machine learning model is retrained and improved, which will improve the detection accuracy next time. This process is performed using TensorFlow and PyTorch.

[1569] Input: Feedback information

[1570] Output: An improved machine learning model

[1571] (Application example 1)

[1572] Next, a description will be given of Application Example 1. In the following description, the data processing device 12 will be referred to as a "server" and the robot 414 will be referred to as a "terminal."

[1573] Currently, there are many cases of special frauds targeting the elderly and young people, and an effective system to prevent these frauds is needed. However, current fraud detection systems often only notify users on their devices, which makes it easy for users to overlook the warning, and few systems can immediately alert users. Therefore, there is a need for a system that can detect signs of fraud in real time and immediately and effectively warn users.

[1574] The specific processing by the specific processing unit 290 of the data processing device 12 in the application example 1 is realized by the following means.

[1575] In this invention, the server includes means for monitoring the content of user communications and capturing data in real time, means for transmitting the captured data to the server, means for analyzing the data received by the server and detecting the possibility of fraud, means for generating a warning based on the detected possibility of fraud and notifying the user, means for collecting feedback information and improving the analysis means based on the feedback information, and means for displaying a warning message on a head-mounted display in real time. This allows the user to visually confirm fraud warnings in real time, effectively preventing damage from special frauds.

[1576] "User communication content" refers to the content of calls and messages made by the user.

[1577] "Means for capturing data in real time" refers to devices or methods that instantly capture the content of ongoing communications.

[1578] The "means for transmitting to the server" refers to a means for sending the captured data to the central processing unit via a network.

[1579] "Means for analyzing and detecting potential fraud" refers to devices or methods for analyzing acquired data and detecting signs of fraudulent activity therein.

[1580] "Means for generating a warning and notifying a user" refers to means for creating a warning message and notifying a user when a potential fraud is detected.

[1581] "Means for collecting feedback information" refers to means for collecting opinions and evaluations from users.

[1582] "Means for improving analysis means" refers to means for improving data analysis methods based on collected feedback information.

[1583] A "head-mounted display" refers to a device worn by a user that displays information visually.

[1584] "Means for displaying a warning message in real time" refers to means for immediately visually conveying the detected warning content to the user.

[1585] The present invention is a system for monitoring user communications in real time, detecting signs of fraud, and issuing a warning. An embodiment of this system will be described in detail below.

[1586] Overall structure

[1587] This system consists of a user's terminal, a server that analyzes the data, and a head-mounted display (HMD) that notifies the user of warnings.

[1588] The server includes means for monitoring user communications and capturing data in real time, means for transmitting the captured data to the server, means for analyzing the received data to detect possible fraud, means for generating and notifying the user of a warning based on possible fraud, and means for collecting feedback information and improving the analysis means.

[1589] Implementation Details

[1590] 1. A way to monitor user communications and capture data in real time

[1591] The user's device monitors the contents of calls and messages in real time. For calls, audio data is captured using the HMD's built-in microphone. For messages, text data is obtained directly. This data is processed using the AudioCapture API (for audio data) and the TextMessage API (for text data).

[1592] 2. Data preprocessing and transmission

[1593] The captured data is preprocessed on the user's device. Specifically, audio data is subjected to noise removal (NoiseSuppressionLib) and sampling rate adjustment (SoundProcessingLib), and text data is preprocessed using a natural language processing library (e.g., spaCy). The data is then encrypted (SSLTLS) and sent to the server via the HTTPS protocol.

[1594] 3. Data Analysis

[1595] The server runs on a high-performance server (e.g., AWS EC2). The data received on the server side is converted into text using a speech recognition engine (Google Speech-to-Text API), and then analyzed using a natural language processing engine (BERT). This analysis detects signs of fraud.

[1596] 4. Alert Generation and Notification

[1597] The server evaluates the likelihood of fraud using a scoring algorithm (XGBoost). If there is a high likelihood of fraud, a warning message is generated. This warning message is sent in real time to the user's HMD and displayed using the Notification API.

[1598] 5. Feedback and Improvement

[1599] Users provide feedback on the resulting warnings, which is collected on the server and used to improve the analysis method using a machine learning platform (e.g., AWS SageMaker).

[1600] Examples of specific examples and prompts

[1601] As a concrete example, consider a case where an elderly person receives a call from someone claiming to be their "son" saying, "I urgently need money to pay for my hospital bills." In this case, the system will operate according to the following prompt:

[1602] Prompt statement:

[1603] 1. Audio data capture and preprocessing:

[1604] Capture the user's voice in real time through the microphone, remove noise with NoiseSuppressionLib, and adjust the sampling rate with SoundProcessingLib.

[1605] 2. Data transmission:

[1606] The preprocessed data is encrypted using SSLTLS and sent to the server using HTTPS.

[1607] 3. Data Analysis:

[1608] On the server side, the speech is converted to text using the Google Speech-to-Text API and analyzed using BERT.

[1609] 4. Alert Generation and Notification:

[1610] XGBoost scores the likelihood of fraud and generates a warning message, which is displayed on the HMD using the Notification API.

[1611] 5. Feedback Processing:

[1612] Collect user feedback and improve machine learning models with AWS SageMaker.

[1613] This configuration allows users to receive immediate warnings about the risk of special fraud, making it possible to effectively prevent fraud damage.

[1614] The flow of the specific processing in the application example 1 will be described with reference to FIG.

[1615] Step 1:

[1616] The device monitors the user's communication in real time and captures the data. The input is the user's call voice and message content, and the output is the captured raw data. The audio data is obtained through the microphone and processed using the AudioCapture API. The text data is obtained directly using the TextMessage API.

[1617] Step 2:

[1618] Preprocess the data captured by the device. For audio data, use NoiseSuppressionLib to remove noise and SoundProcessingLib to adjust the sampling rate. For text data, use a natural language processing library (e.g. spaCy) for preprocessing. The input is raw data and the output is preprocessed data.

[1619] Step 3:

[1620] The terminal encrypts the preprocessed data and sends it to the server. The input is the preprocessed data and the output is the encrypted data. SSL TLS is used for encryption, and the data is sent via the HTTPS protocol.

[1621] Step 4:

[1622] The server decrypts the received data and performs data analysis. The input is the encrypted data, and the output is the analyzed result. In the case of audio data, the server converts the audio to text using the Google Speech-to-Text API, and then analyzes the text data using BERT.

[1623] Step 5:

[1624] The server scores the likelihood of fraud based on the analysis results and generates a warning message. The input is the analyzed text data, and the output is the warning message. XGBoost is used for scoring.

[1625] Step 6:

[1626] The server generates a warning message and sends it to the user's HMD, where it is displayed using the Notification API. The input is the warning message, and the output is the warning content reflected in the user's visual perception.

[1627] Step 7:

[1628] The user acknowledges the warning message and provides feedback. The input is the user's feedback information, and the output is the collected feedback information.

[1629] Step 8:

[1630] The server improves the analysis method based on the feedback information collected. The input is the feedback information, and the output is the improved analysis method. The feedback information is reflected in the machine learning model using AWS SageMaker.

[1631] Furthermore, an emotion engine that estimates the user's emotion may be further combined. That is, the identification processing unit 290 may estimate the user's emotion using the emotion identification model 59, and perform identification processing using the user's emotion.

[1632] This invention realizes more advanced fraud detection by incorporating an emotion engine into a system that uses the power of AI to eradicate special frauds. The system not only monitors users' communications in real time and detects signs of fraud, but also analyzes users' emotions to more accurately assess the possibility of fraud and issue warnings.

[1633] Overall system overview

[1634] 1. Monitoring of outgoing information

[1635] Terminal: Monitors user calls and messages in real time. For calls, the microphone is used to capture audio data, and for messages, the text data is directly acquired.

[1636] 2. Data preprocessing and transmission

[1637] Terminal: For audio data, preprocessing such as noise removal and sampling rate adjustment is performed. The preprocessed data is encrypted and sent to the server.

[1638] 3. Data Analysis

[1639] Server: Analyzes the received data. If it is voice data, it is first converted into text data using speech recognition technology. Next, it analyzes the text data using natural language processing technology to detect signs of fraud.

[1640] 4. Emotional awareness and regulation

[1641] Server: The emotion engine uses the analyzed text and voice data to recognize the user's emotions. The emotion is analyzed based on the tone of voice and the use of words.

[1642] 5. Fraud Indicator Detection and Scoring

[1643] Server: If fraud indicators are detected, adjust the fraud scoring based on the output of the emotion engine, for example increasing the score if the user is feeling anxious or scared.

[1644] 6. Generating and Sending Alerts

[1645] Server: Generates a warning message when the likelihood of fraud exceeds a certain threshold, and sends the generated warning message to the user's device to notify them in real time.

[1646] 7. Feedback and Learning

[1647] User: Provide feedback on whether it is actually a scam.

[1648] Server: Use this feedback to improve the machine learning model and make the next detection more accurate.

[1649] Specific examples

[1650] Example 1: Detecting fraudulent calls using an emotion engine

[1651] User: An elderly person receives a call from someone claiming to be their "son."

[1652] Terminal: Captures the call in real time and sends the audio data to the server.

[1653] Server: Converts voice data into text and uses natural language processing technology to detect signs of fraud, such as "I need money" or "Help me."

[1654] Server: Recognizes when an elderly person is feeling anxious or scared from text and speech analyzed using an emotion engine.

[1655] Server: Score the probability of fraud, increase the score taking into account the results of sentiment analysis, and generate a warning message if the probability is high.

[1656] Devices: Display a warning message on the senior's device to notify them of potential fraud.

[1657] User: Realizes the call is a scam and hangs up.

[1658] User: Provide feedback to the system that it was a scam.

[1659] Server: Improves machine learning models based on feedback.

[1660] Example 2: Detecting message fraud using an emotion engine

[1661] User: A young person receives a suspicious email.

[1662] Terminal: Captures email text in real time and sends it to the server.

[1663] Server: Analyzes text data using natural language processing technology to detect signs of fraud, such as "urgent," "important," and "click here."

[1664] Server: Recognizes from text analyzed using an emotion engine that young people are feeling confused or excited.

[1665] Server: Score the probability of fraud, increase the score taking into account the results of sentiment analysis, and generate a warning message if the probability is high.

[1666] Devices: Display warning messages on young people's devices to inform them of potential scams.

[1667] User: Ignore the email to prevent unauthorized access.

[1668] User: Provide feedback that the email was fraudulent.

[1669] Server: Improve machine learning models based on feedback.

[1670] In this way, this system can effectively prevent fraud damage through real-time detection and warning of special frauds, and even by recognizing user emotions.

[1671] The processing flow will be explained below.

[1672] Step 1:

[1673] User: Makes calls, sends and receives messages, and exchanges necessary information, including potentially fraudulent communications.

[1674] Step 2:

[1675] Terminal: Captures the contents of calls as audio data in real time, and also captures text messages as text data.

[1676] Step 3:

[1677] Terminal: Preprocessing the captured audio data, such as noise removal and unifying the sampling rate, improves the quality of the audio data.

[1678] Step 4:

[1679] Terminal: The pre-processed data is encrypted and sent to the server using a secure communication protocol, along with the communication metadata (source, destination, time, etc.).

[1680] Step 5:

[1681] Server: Receives data sent from the device. In the case of voice data, converts it into text using voice recognition technology. Converts the voice content into text information using a voice recognition engine.

[1682] Step 6:

[1683] Server: Analyzes the received text data using natural language processing (NLP) techniques, such as tokenizing words, tagging parts of speech, and grammar analysis, to detect phrases and structures that may be indicative of fraud.

[1684] Step 7:

[1685] Server: Uses an emotion engine to recognize user emotions from text data. Analyzes user emotions (e.g., anxiety, anger, confusion, etc.) based on voice tone and phrasing.

[1686] Step 8:

[1687] Server: If fraud signs are detected, the server scores the likelihood of fraud by taking into account the user's emotions analyzed by the emotion engine. The likelihood of fraud is adjusted so that the score is higher if the user is feeling anxious or angry.

[1688] Step 9:

[1689] Server: Generates a warning message if the fraud score exceeds a set threshold, including the likelihood of fraud and what to do about it.

[1690] Step 10:

[1691] Server: Sends the generated warning message to the user's device, notifying them in real time and alerting them.

[1692] Step 11:

[1693] Device: Display an alert message on the user's device, using a notification sound or vibration to get the user's attention.

[1694] Step 12:

[1695] Users: Check whether the call or message is fraudulent and take the action outlined in the warning message. If it is fraudulent, immediately stop the transaction or call.

[1696] Step 13:

[1697] Users: Provide feedback to the system on whether the fraud is true or not. User judgment is also important data for improving the system.

[1698] Step 14:

[1699] Device: Sends user feedback information to the server, including feedback content (whether it was a scam, the effectiveness of countermeasures, etc.).

[1700] Step 15:

[1701] Server: Receives feedback information and uses it as data to improve the performance of the machine learning model and emotion engine. Based on the feedback, the detection algorithm is retrained.

[1702] Through this process, the system can detect signs of specialized fraud in real time, and by utilizing an emotion engine, it can accurately assess the possibility of fraud and issue a warning.

[1703] Example 2

[1704] Next, a description will be given of Example 2. In the following description, the data processing device 12 will be referred to as a "server" and the robot 414 will be referred to as a "terminal."

[1705] In recent years, special frauds have been on the rise, with the damage targeting elderly people and beginners becoming particularly severe. Conventional fraud prevention systems are capable of detecting signs of fraud, but they have limitations in the timing and accuracy of issuing warnings. Furthermore, because they do not take the user's emotional state into account, they have the problem of easily missing important warnings. Therefore, there is a need for a system that can more accurately detect potential fraud and issue warnings promptly by monitoring communication content in real time and analyzing the user's emotions.

[1706] The specific processing by the specific processing unit 290 of the data processing device 12 in the second embodiment is realized by the following means.

[1707] In this invention, the server includes means for monitoring user communications and capturing data in real time, means for preprocessing and encrypting the captured data and transmitting it to the server, means for analyzing the data received by the server and detecting the possibility of fraud, means for recognizing an emotional state from the analyzed data and voice data, means for scoring fraud using the recognized emotional state, means for generating a warning based on the detected possibility of fraud and scoring data and notifying the user, and means for collecting feedback information and improving the analysis means based on the feedback information, thereby enabling more accurate fraud detection and faster warnings that reflect the user's emotional state.

[1708] "User communications" refers to digital communications such as calls, messages, and emails sent or received by a user.

[1709] "Means for capturing data in real time" refers to a piece of equipment or software that acquires the contents of user communications in real time and processes them immediately.

[1710] "Preprocessing" refers to initial processing such as noise removal and sampling rate adjustment that is performed to convert captured data into a form that is easier to analyze.

[1711] "Encryption" is the technology of converting data into cryptographic code to protect the privacy and security of communication data.

[1712] A "server" is a central computer system that receives data from multiple users over a network and processes and analyzes it.

[1713] "Means for detecting potential fraud" are algorithms and technologies that analyze the content of communications to find signs of fraud.

[1714] "Means for recognizing emotional states" refers to technology that analyzes the tone and phrasing of a user's voice or text to identify the emotions the user is feeling.

[1715] A "fraud scoring means" is a technology that numerically calculates the likelihood of fraud based on detected indicators of fraud and the user's emotional state.

[1716] "Means for generating a warning and notifying the user" refers to technology that automatically creates a warning message when a high possibility of fraud is determined and sends it to the user's device.

[1717] "Feedback information" refers to data on reactions and evaluations provided by users in response to warnings issued by the system.

[1718] "Means to improve analysis means" refers to using feedback information to improve fraud detection algorithms and emotion recognition techniques, thereby increasing the accuracy of the overall system.

[1719] The present invention is a system that uses the power of AI to eradicate special frauds, and by incorporating an emotion engine, achieves more advanced fraud detection. The system not only monitors user communications in real time and detects signs of fraud, but also analyzes user emotions to more accurately assess the possibility of fraud and issue warnings. Specific embodiments of the present invention are described below.

[1720] 1. Monitoring of outgoing information

[1721] Device: Monitors user calls and messages in real time. For calls, audio data is captured using the device's built-in microphone, and for messages, text data is directly acquired.

[1722] 2. Data preprocessing and transmission

[1723] Terminal: For audio data, first preprocessing such as noise removal and sampling rate adjustment is performed, and the preprocessed data is encrypted using the AES-256 encryption algorithm and sent to the server via the SSL / TLS protocol.

[1724] 3. Speech data conversion and natural language processing

[1725] Server: Receives the voice data and converts it to text using the Google Cloud Speech-to-Text API. The converted text data is analyzed using a natural language processing library (e.g., spaCy or BERT) to detect signs of fraud, such as "I need money" or "Help me."

[1726] 4. Emotion Recognition and Analysis

[1727] Server: Inputs text and voice data into Affectiva's emotion engine to analyze the user's emotional state, analyzing voice tone and phrasing to identify emotions such as "anxiety," "fear," and "confusion."

[1728] 5. Fraud Indicator Scoring

[1729] Server: Calculates the fraud score based on the detected fraud indicators and the results of sentiment analysis. If the user shows high anxiety or fear, the fraud score is increased.

[1730] 6. Generating and Sending Warning Messages

[1731] Server: When the fraud score exceeds the set threshold, a warning message is automatically generated and sent to the user's device immediately, where the device notifies the user.

[1732] 7. Gather feedback and learn

[1733] User: After receiving the warning message, the user can provide feedback on whether it was indeed a scam, for example, by replying "It was a scam" or "It wasn't a scam."

[1734] Server: Stores the feedback information in a database and updates the machine learning model (e.g., Scikit-learn or TensorFlow) to improve the accuracy of fraud detection next time.

[1735] Specific examples

[1736] Example 1: Detecting fraudulent calls using an emotion engine

[1737] User: An elderly person receives a call from someone claiming to be their "son."

[1738] Terminal: Captures the call in real time and sends the audio data to the server.

[1739] Server: Converts voice data into text and uses natural language processing technology to detect signs of fraud, such as "I need money" or "Help me."

[1740] Server: Recognizes when an elderly person is feeling anxious or scared from text and speech analyzed using an emotion engine.

[1741] Server: Score the likelihood of fraud, increase the score based on the results of sentiment analysis, and generate a warning message if the probability is high.

[1742] Devices: Display a warning message on the senior's device to notify them of potential fraud.

[1743] User: Realizes the call is a scam and hangs up.

[1744] User: Provide feedback to the system that it was a scam.

[1745] Server: Improves machine learning models based on feedback.

[1746] Example prompts to be input to the generative AI model

[1747] "Someone claiming to be my son said he needed money. Please analyze the possibility that this call is fraudulent. I have the audio data and pre-processed text data."

[1748] The flow of the identification process in the second embodiment will be described with reference to FIG.

[1749] Step 1:

[1750] Terminal: Monitors user calls and messages in real time. The input is the user's voice or text message, and the output is captured voice data or text data. For voice, the voice data is captured using the microphone built into the terminal. For messages, the text data is obtained as is.

[1751] Step 2:

[1752] Terminal: Preprocesses the captured audio data. The input is the captured audio data, and the output is the preprocessed audio data. This includes applying noise reduction filters and adjusting the sampling rate. After preprocessing, the data is encrypted using the AES-256 encryption algorithm and sent to the server using the SSL / TLS protocol.

[1753] Step 3:

[1754] Server: Converts audio data to text data. The input is preprocessed audio data and the output is text data. The audio data is converted to text using the Google Cloud Speech-to-Text API. This text data is sent to the next processing step.

[1755] Step 4:

[1756] Server: Analyzes the text data using natural language processing techniques. The input is text data, and the output is the analysis results containing information about signs of fraud. Natural language processing libraries (e.g., spaCy and BERT) are used to detect fraudulent keywords and phrases such as "I need money" and "Help me."

[1757] Step 5:

[1758] Server: Analyzes the user's emotional state using an emotion engine. The input is text data and voice data, and the output is data indicating the user's emotional state. Using Affectiva's emotion engine, it identifies emotions such as "anxiety," "fear," and "confusion" from voice tone and vocabulary.

[1759] Step 6:

[1760] Server: Performs fraud scoring. The input is fraud indicator information and emotional state data, and the output is a fraud score. Based on the detected fraud indicators and the user's emotional state, the possibility of fraud is numerically evaluated. If the emotional state is "anxiety" or "fear," the fraud score increases.

[1761] Step 7:

[1762] Server: Generates a warning message and notifies the user. The input is the fraud score and the output is a warning message. If the fraud score exceeds the set threshold, a warning message is automatically generated. This message is immediately sent to the user's device and the user is notified on the device.

[1763] Step 8:

[1764] User: Provides feedback. The input is the user's reaction to the warning, and the output is feedback information. The user reports whether the warning was correct or not, replying to the system "it was a scam" or "it wasn't a scam."

[1765] Step 9:

[1766] Server: Updates the machine learning model based on the feedback information. The input is the feedback information, and the output is an improved analysis model. The received feedback is stored in a database and used to improve the model using machine learning algorithms (e.g., Scikit-learn or TensorFlow) to improve the fraud detection accuracy next time.

[1767] (Application example 2)

[1768] Next, a description will be given of Application Example 2. In the following description, the data processing device 12 will be referred to as a "server" and the robot 414 will be referred to as a "terminal."

[1769] Conventional fraud detection systems only detect signs of fraud contained in user communications, and improving detection accuracy by reflecting the user's emotional state is a challenge. Furthermore, there is a lack of real-time warnings and feedback-based model improvements, limiting the effectiveness of fraud prevention.

[1770] The specific processing by the specific processing unit 290 of the data processing device 12 in the application example 2 is realized by the following means.

[1771] In this invention, the server includes means for monitoring user communications and detecting signs of fraud in real time, means for analyzing user emotions using an emotion engine and adjusting the fraud score, means for generating a warning based on the likelihood of fraud and notifying the user, and means for collecting feedback information and improving the analysis means, thereby enabling advanced fraud detection that takes user emotions into account, real-time warnings, and continuous system improvement based on feedback.

[1772] "User communication content" refers to data such as voice calls, messages, and emails exchanged between the user and others.

[1773] "Means of capturing data in real time" refers to technology that instantly collects communication content and stores it as data for later processing.

[1774] A "server" refers to a computer system that receives data from multiple client devices via a network and analyzes and processes the data.

[1775] "Measures to detect potential fraud" refers to algorithms or technologies used to determine whether a user's communications contain fraudulent activity.

[1776] "Means for recognizing emotions and adjusting fraud scores" refers to technology that analyzes a user's emotional state and modifies the indicators used to assess the likelihood of fraud.

[1777] "Means for generating an alert and notifying the user" refers to technology that creates an alert message based on the detected potential fraud and sends it to the user's device.

[1778] "Means for collecting feedback information and improving analytical methods" refers to technologies that collect user feedback and use that information to improve fraud detection algorithms and the overall performance of the system.

[1779] "Means for converting voice data into text" refers to the process of converting voice data into character data using voice recognition technology.

[1780] "Natural language processing technology" refers to technology for analyzing the structure and meaning of language based on text data.

[1781] "Emotion analysis technology using generative models" refers to technology that uses machine learning models to estimate user emotions from text data and voice data.

[1782] This invention is a system that monitors the content of user communications in real time, analyzes signs of fraud and user emotions, and thereby detects the possibility of fraud with high accuracy and issues a warning.

[1783] System configuration

[1784] Hardware

[1785] Devices: Microphones and processing units built into smartphones, smart glasses, and head-mounted displays

[1786] Server: Cloud server capable of high-performance analysis and data processing

[1787] software

[1788] Audio data capture: sounddevice library

[1789] Audio data preprocessing: librosa library

[1790] Speech Recognition: TensorFlow Model

[1791] Sentiment Analysis: Hugging Face transformers library

[1792] Processing Details

[1793] Monitoring user communications and capturing data

[1794] The device monitors the user's calls and messages in real time and captures voice data. In the case of a call, the voice data is collected using a microphone and pre-processed, such as noise reduction and sampling rate adjustment.

[1795] Data transmission and analysis

[1796] The pre-processed data is encrypted and then sent to the server, which then converts the received voice data into text using natural language processing technology.

[1797] Sentiment Analysis and Fraud Scoring

[1798] On the server side, a generative AI model is used to analyze user sentiment from text and voice data, and the fraud score is adjusted based on the analyzed sentiment.

[1799] Generate alerts and notify users

[1800] If the likelihood of fraud exceeds a certain threshold, the server generates a warning message and notifies the user's terminal in real time.

[1801] Feedback and machine learning model improvement

[1802] Users provide feedback to the system, which the server uses to improve the machine learning model and make the next detection more accurate.

[1803] Specific examples

[1804] For example, if an elderly person receives a phone call from someone claiming to be their "son," the following steps may occur:

[1805] 1. The elderly person's device captures the contents of the call in real time and sends it to the server.

[1806] 2. The server converts the voice data into text and uses natural language processing technology to analyze signs of fraud, such as "I need money."

[1807] 3. Furthermore, generative AI models are used to analyze the emotions of older adults (e.g., anxiety and fear) from text and voice data.

[1808] 4. If it is determined that there is a high possibility of fraud, a warning message will be displayed on the elderly person's device.

[1809] Prompt Sentence Examples

[1810] "Please answer the following questions based on your call data:

[1811] 1. Are there any signs of fraud in this conversation?

[1812] 2. What are the user's emotions?

[1813] example:

[1814] Call text: 'Mom, help me! I need money now.'"

[1815] In this way, the invention allows for advanced fraud detection that takes user emotions into account, real-time alerts, and continuous system improvement based on feedback.

[1816] The flow of the specific processing in the application example 2 will be described with reference to FIG.

[1817] Step 1:

[1818] The device monitors the user's communication in real time and captures audio data. Specifically, it collects audio data using a microphone and obtains it through an audio data capture library (e.g., sounddevice). The input is real-time audio, and the output is captured audio data.

[1819] Step 2:

[1820] The device preprocesses the captured audio data. Specifically, it performs operations such as noise reduction and sampling rate adjustment. This process uses an audio data preprocessing library (e.g., librosa). The input is the captured audio data, and the output is the preprocessed audio data.

[1821] Step 3:

[1822] The device encrypts the preprocessed audio data and sends it to the server. Specifically, it uses an encryption library to securely convert the data and send it to the server over the network. The preprocessed audio data is the input, and the encrypted data is sent to the server as the output.

[1823] Step 4:

[1824] The server analyzes the received voice data. Specifically, it uses speech recognition technology (e.g., TensorFlow model) to convert the voice data into text data. The input is encrypted voice data, and the output is text data.

[1825] Step 5:

[1826] The server analyzes the text data using natural language processing technology. Specifically, it uses natural language processing technology (e.g., a generative AI model) to detect signs of fraud from the text data. The input is text data, and the output is a determination of the likelihood of fraud.

[1827] Step 6:

[1828] The server recognizes emotions from the analyzed data and adjusts the fraud score. Specifically, it analyzes the user's emotions using emotion analysis technology (e.g., Hugging Face's transformers library) and updates the fraud score. The input is the analyzed text data and the emotion analysis result, and the output is the adjusted fraud score.

[1829] Step 7:

[1830] The server generates a warning message and notifies the user if the fraud probability exceeds a certain threshold. Specifically, the server automatically generates a warning message and sends it to the user's device. The input is the adjusted fraud score, and the output is the warning message displayed on the user's device.

[1831] Step 8:

[1832] The user provides feedback on whether or not the fraud is actually occurring. Specifically, the user reports whether or not the fraud is occurring within the application, and the feedback is sent to the server. The input is the user's feedback information, and the output is the feedback data provided to the server.

[1833] Step 9:

[1834] The server improves the machine learning model based on the feedback information. Specifically, it analyzes the collected feedback information and runs an algorithm to improve the performance of the fraud detection model. The input is the feedback information and the output is an improved machine learning model.

[1835] The specific processing unit 290 transmits the result of the specific processing to the robot 414. In the robot 414, the control unit 46A causes the speaker 240 and the control target 443 to output the result of the specific processing. The microphone 238 acquires voice indicating a user input regarding the result of the specific processing. The control unit 46A transmits voice data indicating the user input acquired by the microphone 238 to the data processing device 12. In the data processing device 12, the specific processing unit 290 acquires the voice data.

[1836] The data generation model 58 is a so-called generative AI (Artificial Intelligence). An example of the data generation model 58 is ChatGPT (Internet Search<URL: https: / / openai.com / blog / chatgpt> ), Gemini (Internet search <url: https: gemini.google.com ?hl="ja">) and other generation AIs. The data generation model 58 is obtained by performing deep learning on a neural network. A prompt including an instruction is input to the data generation model 58, and inference data such as voice data indicating voice, text data indicating text, and image data indicating an image is also input. The data generation model 58 performs inference on the input inference data in accordance with the instruction indicated by the prompt, and outputs the inference result in a data format such as voice data and text data. Here, inference refers to, for example, analysis, classification, prediction, and / or summarization.

[1837] In the above embodiment, an example was given in which the specific processing is performed by the data processing device 12, but the technology of the present disclosure is not limited to this, and the specific processing may be performed by the robot 414.

[1838] The emotion identification model 59 as an emotion engine may determine the user's emotion according to a specific mapping. Specifically, the emotion identification model 59 may determine the user's emotion according to an emotion map (see FIG. 9), which is a specific mapping. Similarly, the emotion identification model 59 may determine the robot's emotion, and the identification processing unit 290 may perform identification processing using the robot's emotion.

[1839] FIG. 9 is a diagram illustrating an emotion map 400 on which multiple emotions are mapped. In the emotion map 400, emotions are arranged in concentric circles radiating from the center. Emotions closer to the center of the concentric circles are more primitive. Emotions representing states and actions arising from a state of mind are arranged on the outer edges of the concentric circles. The concept of emotion includes both affect and mental states. Emotions generally generated from reactions occurring in the brain are arranged on the left side of the concentric circles. Emotions generally induced by situational judgment are arranged on the right side of the concentric circles. Emotions generally generated from reactions occurring in the brain and induced by situational judgment are arranged on the upper and lower sides of the concentric circles. Furthermore, the emotion of "pleasure" is arranged on the upper side of the concentric circles, and the emotion of "discomfort" is arranged on the lower side. In this way, in the emotion map 400, multiple emotions are mapped based on the structure by which emotions are generated, and emotions that tend to occur simultaneously are mapped close to each other.

[1840] These emotions are distributed in the 3 o'clock direction on emotion map 400, and typically fluctuate between relief and anxiety. In the right half of emotion map 400, situational awareness dominates over internal sensations, resulting in a sense of calm.

[1841] The inside of emotion map 400 represents what is going on in the mind, and the outside of emotion map 400 represents behavior, so the further you go outside emotion map 400, the more visible the emotions become (the more they are expressed in behavior).

[1842] Human emotions are based on various balances, such as posture and blood sugar levels. When these balances deviate from the ideal, a state of discomfort is indicated, and when they approach the ideal, a state of pleasure is indicated. Emotions can also be created for robots, automobiles, and motorcycles, based on various balances, such as posture and remaining battery life. When these balances deviate from the ideal, a state of discomfort is indicated, and when they approach the ideal, a state of pleasure is indicated. An emotion map can be generated, for example, based on Dr. Mitsuyoshi's emotion map (Research on Voice Emotion Recognition and Emotional Brain Physiological Signal Analysis Systems, Tokushima University, Doctoral Dissertation: https: / / ci.nii.ac.jp / naid / 500000375379). The left half of the emotion map lists emotions belonging to the "reaction" domain, where sensation is dominant. The right half of the emotion map lists emotions belonging to the "situation" domain, where situational awareness is dominant.

[1843] The emotion map defines two emotions that promote learning. One is a negative emotion on the situation side, around the middle of "repentance" or "reflection." In other words, this occurs when the robot experiences negative emotions such as "I never want to feel this way again" or "I don't want to be scolded again." The other is a positive emotion on the response side, around "desire." In other words, this occurs when the robot experiences positive feelings such as "I want more" or "I want to know more."

[1844] The emotion identification model 59 inputs user input into a pre-trained neural network, obtains emotion values ​​indicating each emotion shown in the emotion map 400, and determines the user's emotion. This neural network is pre-trained based on multiple pieces of training data that are combinations of user input and emotion values ​​indicating each emotion shown in the emotion map 400. Furthermore, this neural network is trained so that emotions that are located close to each other have similar values, as in the emotion map 900 shown in FIG. 10. FIG. 10 shows an example in which multiple emotions, "relieved," "calm," and "reassuring," have similar emotion values.

[1845] The system according to the present disclosure has been described above mainly with respect to the functions of the data processing device 12, but the system according to the present disclosure is not necessarily implemented on a server. The system according to the present disclosure may be implemented as a general information processing system. The present disclosure may be implemented, for example, as a software program running on a personal computer or an application running on a smartphone, etc. The method according to the present disclosure may be provided to users in the form of SaaS (Software as a Service).

[1846] In the above embodiment, an example was given in which the specific processing is performed by one computer 22, but the technology of the present disclosure is not limited to this, and the specific processing may be distributed and performed by a plurality of computers including the computer 22. For example, the data generation model 58 may be provided in an external device of the data processing device 12, and data may be generated in the external device in accordance with input data.

[1847] In the above embodiment, an example in which the specific processing program 56 is stored in the storage 32 has been described, but the technology of the present disclosure is not limited to this. For example, the specific processing program 56 may be stored in a portable, computer-readable, non-transitory storage medium such as a USB (Universal Serial Bus) memory. The specific processing program 56 stored in the non-transitory storage medium is installed in the computer 22 of the data processing device 12. The processor 28 executes the specific processing in accordance with the specific processing program 56.

[1848] Alternatively, the specific processing program 56 may be stored in a storage device such as a server connected to the data processing device 12 via the network 54, and the specific processing program 56 may be downloaded and installed on the computer 22 in response to a request from the data processing device 12.

[1849] It is not necessary to store all of the specific processing program 56 in a storage device such as a server connected to the data processing device 12 via the network 54, or to store all of the specific processing program 56 in the storage 32; only a portion of the specific processing program 56 may be stored.

[1850] The hardware resource for executing a specific process can be any of the following processors: An example of a processor is a CPU, which is a general-purpose processor that functions as a hardware resource for executing a specific process by executing software, i.e., a program. Another example of a processor is a dedicated electrical circuit, such as an FPGA (Field-Programmable Gate Array), a PLD (Programmable Logic Device), or an ASIC (Application Specific Integrated Circuit), which is a processor with a circuit configuration designed specifically for executing a specific process. Each processor has built-in or connected memory, and each processor uses the memory to execute the specific process.

[1851] The hardware resource that executes the specific processing may be configured with one of these various processors, or may be configured with a combination of two or more processors of the same or different types (for example, a combination of multiple FPGAs, or a combination of a CPU and an FPGA). Also, the hardware resource that executes the specific processing may be a single processor.

[1852] As an example of a system configured with a single processor, first, one processor is configured by combining one or more CPUs and software, and this processor functions as a hardware resource that executes a specific process. Second, there is a system that uses a processor that realizes the functions of an entire system including multiple hardware resources that execute a specific process on a single IC chip, as typified by SoC (System-on-a-chip). In this way, a specific process is realized using one or more of the above-mentioned various processors as hardware resources.

[1853] Furthermore, the hardware structure of these various processors can be, more specifically, an electric circuit that combines circuit elements such as semiconductor devices. The specific processing described above is merely an example. Therefore, it goes without saying that unnecessary steps may be deleted, new steps may be added, or the processing order may be rearranged, without departing from the spirit of the invention.

[1854] The above-described description and illustrations are a detailed explanation of the parts related to the technology of the present disclosure and are merely an example of the technology of the present disclosure. For example, the above description of the configuration, functions, actions, and effects is an explanation of an example of the configuration, functions, actions, and effects of the parts related to the technology of the present disclosure. Therefore, it goes without saying that unnecessary parts may be deleted, new elements may be added, or replacements may be made to the above-described description and illustrations within the scope of the gist of the technology of the present disclosure. Furthermore, to avoid confusion and facilitate understanding of the parts related to the technology of the present disclosure, the above-described description and illustrations omit explanations of common technical knowledge that do not require particular explanation to enable the implementation of the technology of the present disclosure.

[1855] All publications, patent applications, and technical standards mentioned in this specification are herein incorporated by reference to the same extent as if each individual publication, patent application, or technical standard was specifically and individually indicated to be incorporated by reference.

[1856] The following is further disclosed regarding the above embodiment.

[1857] (Claim 1)

[1858] means for monitoring user communications and capturing data in real time;

[1859] means for transmitting the captured data to a server;

[1860] means for analyzing data received by said server to detect possible fraud;

[1861] means for generating and notifying a user of a warning based on the detected likelihood of fraud;

[1862] means for collecting feedback information and improving the analysis means based on the feedback information;

[1863] A system including:

[1864] (Claim 2)

[1865] 10. The system of claim 1, further comprising means for converting the voice data to text if the communication content is voice data.

[1866] (Claim 3)

[1867] 2. The system according to claim 1, wherein the analyzing means analyzes the text data using natural language processing technology.

[1868] (Claim 4)

[1869] 2. The system of claim 1, wherein the means for generating an alert performs fraud scoring based on the analyzed data and generates an alert based on the scoring.

[1870] (Claim 5)

[1871] 10. The system of claim 1, further comprising means for continuously updating the machine learning model based on the feedback information.

[1872] "Example 1"

[1873] (Claim 1)

[1874] means for monitoring user communications in real time and capturing data as voice or text data;

[1875] A pre-processing means for removing noise and adjusting the sampling rate of the captured data;

[1876] means for encrypting the preprocessed data and transmitting it to a server;

[1877] means for analyzing the data received by the server to detect possible fraud;

[1878] A means for scoring the likelihood of fraud based on the analyzed data; and

[1879] means for generating a warning message to notify the user when the likelihood of fraud exceeds a certain threshold;

[1880] means for collecting feedback information from users and improving the analysis means based on the feedback information;

[1881] A system including:

[1882] (Claim 2)

[1883] 10. The system of claim 1, including speech recognition technology for converting the voice data to text.

[1884] (Claim 3)

[1885] 2. The system according to claim 1, wherein the analyzing means analyzes the text data using natural language processing technology.

[1886] "Application Example 1"

[1887] (Claim 1)

[1888] means for monitoring user communications and capturing data in real time;

[1889] means for transmitting the captured data to a server;

[1890] means for analyzing data received by said server to detect possible fraud;

[1891] means for generating and notifying a user of a warning based on the detected likelihood of fraud;

[1892] means for collecting feedback information and improving the analysis means based on the feedback information;

[1893] a means for displaying a warning message on a head mounted display in real time;

[1894] A system including:

[1895] (Claim 2)

[1896] 10. The system of claim 1, further comprising means for converting the voice data to text if the communication content is voice data.

[1897] (Claim 3)

[1898] 2. The system according to claim 1, wherein the analyzing means analyzes the text data using natural language processing technology.

[1899] "Example 2: Combining Emotion Engines"

[1900] (Claim 1)

[1901] means for monitoring user communications and capturing data in real time;

[1902] means for preprocessing the captured data, encrypting it and transmitting it to a server;

[1903] means for analyzing data received by said server to detect possible fraud;

[1904] means for recognizing an emotional state from the analyzed data and the voice data;

[1905] means for scoring fraud using the recognized emotional state;

[1906] means for generating and notifying a user of a warning based on the detected likelihood of fraud and the scoring data;

[1907] means for collecting feedback information and improving the analysis means based on the feedback information;

[1908] A system including:

[1909] (Claim 2)

[1910] 10. The system of claim 1, further comprising means for converting the voice data to text if the communication content is voice data.

[1911] (Claim 3)

[1912] 2. The system according to claim 1, wherein the analyzing means analyzes the text data using natural language processing technology.

[1913] "Application example 2 when combining emotion engines"

[1914] (Claim 1)

[1915] means for monitoring user communications and capturing data in real time;

[1916] means for transmitting the captured data to a server;

[1917] means for analyzing data received by said server to detect possible fraud;

[1918] means for recognizing emotions from the analyzed data and adjusting fraud scores;

[1919] means for generating and notifying a user of a warning based on the detected likelihood of fraud;

[1920] means for collecting feedback information and improving the analysis means based on the feedback information;

[1921] A system including:

[1922] (Claim 2)

[1923] 10. The system of claim 1, further comprising means for converting the voice data to text if the communication content is voice data.

[1924] (Claim 3)

[1925] 2. The system according to claim 1, wherein the analysis means analyzes text data using natural language processing technology and includes sentiment analysis technology using a generative model. [Explanation of symbols]

[1926] 10, 210, 310, 410 Data Processing Systems 12 Data Processing Device 14 Smart Devices 214 Smart Glasses 314 Headset-type terminal 414 Robot< / url:> < / url:> < / url:> < / url:>

Claims

1. means for monitoring user communications and capturing data in real time; means for transmitting the captured data to a server; means for analyzing data received by said server to detect possible fraud; means for generating and notifying a user of a warning based on the detected likelihood of fraud; means for collecting feedback information and improving the analysis means based on the feedback information; A system including:

2. 2. The system of claim 1, further comprising means for converting said voice data to text if said communication is voice data.

3. 2. The system according to claim 1, wherein the analyzing means analyzes the text data using natural language processing techniques.

4. 2. The system of claim 1, wherein the means for generating an alert performs fraud scoring based on the analyzed data and generates an alert based on the scoring.

5. The system of claim 1 , further comprising means for continuously updating the machine learning model based on the feedback information.

Citation Information

Patent Citations

  • Persona chatbot control method and system

    JP2022180282A