system
A system converts user voice to text, uses generative AI for real-time fraud detection, and sends alerts to prevent sophisticated scams, addressing the inadequacies of conventional fraud prevention methods.
Patent Information
- Authority / Receiving Office
- JP · JP
- Patent Type
- Applications
- Current Assignee / Owner
- SOFTBANK GROUP CORP
- Filing Date
- 2024-11-13
- Publication Date
- 2026-05-25
AI Technical Summary
Conventional methods are inadequate in preventing sophisticated frauds such as virtual billing scams and impersonation, particularly affecting the elderly, as they do not provide real-time detection and appropriate warnings.
A system that converts user voice to text, performs real-time analysis using generative AI to detect fraud patterns, and sends alerts to pre-configured contacts.
Enables quick and effective prevention of fraud by warning users and their contacts, enhancing user safety and peace of mind.
Smart Images

Figure 2026085756000001_ABST
Abstract
Description
Technical Field
[0001] The technology of the present disclosure relates to a system.
Background Art
[0002] Patent Document 1 discloses a method for controlling a persona chatbot, which is performed by at least one processor, the method including the steps of receiving a user utterance, adding the user utterance to a prompt including an instruction sentence related to an explanation of a character of the chatbot, encoding the prompt, and inputting the encoded prompt into a language model to generate a chatbot utterance in response to the user utterance.
Prior Art Documents
Patent Documents
[0003]
Patent Document 1
Summary of the Invention
Problems to be Solved by the Invention
[0004] Currently, many people, especially the elderly, are suffering from frauds such as virtual billing fraud and fraud calls pretending to be family members, and the fraud methods are becoming more sophisticated, so the conventional prevention measures are not sufficient to cope with them. Therefore, a new means is required to prevent users from being involved in fraud without noticing. In particular, a system that can detect signs of fraud in real time, warn users, and enable prompt and appropriate countermeasures is needed.
Means for Solving the Problems
[0005] This invention provides a system that acquires a user's voice, converts it into text data, performs real-time analysis using a generation AI based on the text data, and issues a warning if there is a possibility of fraud. This system performs pattern analysis corresponding to fraudulent methods and sends alerts to contacts set in advance by the user, enabling a quick response and preventing fraud by directly warning the user.
[0006] A "user" refers to an individual or their family who uses this system.
[0007] "Voice" refers to sound signals that include conversations made by the user and ambient sounds.
[0008] "Text data" refers to data obtained by converting speech into written information and putting it into a format that can be processed by a computer.
[0009] "Generative AI" refers to artificial intelligence technology used for analyzing voice and text data, and specifically to models used for detecting fraud patterns.
[0010] "Real-time analysis" refers to a process that executes data processing at the moment it is generated or with a very short delay, and outputs immediate results.
[0011] "Potential fraud" indicates that the acquired audio and converted text data may contain patterns or characteristics specific to fraud.
[0012] A "warning" refers to a form of notification or alert intended to draw the user's attention, and is presented through visual or auditory means.
[0013] "Communication methods" refer to the electronic media and protocols used to send alerts.
[0014] "Configured contacts" refers to the contact information of family members and related parties that the user has registered in advance, and it specifies who should receive warning alerts. [Brief explanation of the drawing]
[0015] [Figure 1] This is a conceptual diagram showing an example of the configuration of a data processing system according to the first embodiment. [Figure 2] This is a conceptual diagram showing an example of the essential functions of a data processing device and a smart device according to the first embodiment. [Figure 3] This is a conceptual diagram showing an example of the configuration of a data processing system according to the second embodiment. [Figure 4] This is a conceptual diagram showing an example of the main functions of a data processing device and smart glasses according to the second embodiment. [Figure 5] This is a conceptual diagram showing an example of the configuration of a data processing system according to the third embodiment. [Figure 6] This is a conceptual diagram showing an example of the main functions of a data processing device and a headset-type terminal according to the third embodiment. [Figure 7] This is a conceptual diagram showing an example of the configuration of a data processing system according to the fourth embodiment. [Figure 8] This is a conceptual diagram showing an example of the main functions of a data processing device and a robot according to the fourth embodiment. [Figure 9] This shows an emotion map where multiple emotions are mapped. [Figure 10] This shows an emotion map where multiple emotions are mapped. [Figure 11] This is a sequence diagram showing the processing flow of the data processing system in Example 1. [Figure 12] This is a sequence diagram showing the processing flow of the data processing system in Application Example 1. [Figure 13] This is a sequence diagram showing the processing flow of the data processing system in Example 2, which incorporates an emotion engine. [Figure 14]It is a sequence diagram showing the processing flow of a data processing system in Application Example 2 when a sentiment engine is combined.
Embodiments for Carrying Out the Invention
[0016] Hereinafter, an example of an embodiment of a system according to the technology of the present disclosure will be described with reference to the accompanying drawings.
[0017] First, the terms used in the following description will be explained.
[0018] In the following embodiments, a numbered processor (hereinafter simply referred to as "processor") may be a single arithmetic unit or a combination of multiple arithmetic units. Also, the processor may be a single type of arithmetic unit or a combination of multiple types of arithmetic units. Examples of arithmetic units include a CPU (Central Processing Unit), a GPU (Graphics Processing Unit), a GPGPU (General-Purpose computing on Graphics Processing Units), an APU (Accelerated Processing Unit), and the like.
[0019] In the following embodiments, a numbered RAM (Random Access Memory) is a memory in which information is temporarily stored and is used as a work memory by the processor.
[0020] In the following embodiments, a numbered storage is one or more non-volatile storage devices that store various programs and various parameters, etc. Examples of non-volatile storage devices include flash memory (SSD (Solid State Drive)), magnetic disks (e.g., hard disks), or magnetic tapes, and the like.
[0021] In the following embodiments, the signed communication interface (I / F) is an interface that includes a communication processor and an antenna, etc. The communication interface manages communication between multiple computers. Examples of communication standards applicable to the communication interface include wireless communication standards such as 5G (5th Generation Mobile Communication System), Wi-Fi (registered trademark), or Bluetooth (registered trademark).
[0022] In the following embodiments, "A and / or B" is synonymous with "at least one of A and B." That is, "A and / or B" means that it may be A alone, or B alone, or a combination of A and B. Furthermore, in this specification, the same concept as "A and / or B" applies when expressing three or more things linked by "and / or."
[0023] [First Embodiment]
[0024] Figure 1 shows an example of the configuration of the data processing system 10 according to the first embodiment.
[0025] As shown in Figure 1, the data processing system 10 includes a data processing device 12 and a smart device 14. An example of the data processing device 12 is a server.
[0026] The data processing device 12 comprises a computer 22, a database 24, and a communication interface 26. The computer 22 is an example of a "computer" related to the technology of this disclosure. The computer 22 comprises a processor 28, RAM 30, and storage 32. The processor 28, RAM 30, and storage 32 are connected to a bus 34. The database 24 and the communication interface 26 are also connected to the bus 34. The communication interface 26 is connected to a network 54. An example of the network 54 is a WAN (Wide Area Network) and / or a LAN (Local Area Network).
[0027] The smart device 14 comprises a computer 36, a reception device 38, an output device 40, a camera 42, and a communication interface 44. The computer 36 comprises a processor 46, RAM 48, and storage 50. The processor 46, RAM 48, and storage 50 are connected to a bus 52. The reception device 38, output device 40, and camera 42 are also connected to the bus 52.
[0028] The reception device 38 is equipped with a touch panel 38A and a microphone 38B, etc., and receives user input. The touch panel 38A receives user input by detecting contact with an object (e.g., a pen or finger). The microphone 38B receives user input by detecting the user's voice. The control unit 46A transmits data indicating the user input received by the touch panel 38A and microphone 38B to the data processing device 12. In the data processing device 12, the specific processing unit 290 acquires the data indicating the user input.
[0029] The output device 40 includes a display 40A and a speaker 40B, and presents data to the user 20 by outputting the data in a form perceptible to the user 20 (e.g., audio and / or text). The display 40A displays visible information such as text and images according to instructions from the processor 46. The speaker 40B outputs audio according to instructions from the processor 46. The camera 42 is a small digital camera equipped with an optical system such as a lens, aperture, and shutter, and an image sensor such as a CMOS (Complementary Metal-Oxide-Semiconductor) image sensor or a CCD (Charge Coupled Device) image sensor.
[0030] Communication interface 44 is connected to network 54. Communication interfaces 44 and 26 are responsible for the exchange of various types of information between processor 46 and processor 28 via network 54.
[0031] Figure 2 shows an example of the main functions of the data processing device 12 and the smart device 14.
[0032] As shown in Figure 2, in the data processing device 12, a specific processing is performed by the processor 28. A specific processing program 56 is stored in the storage 32. The specific processing program 56 is an example of a "program" related to the technology of this disclosure. The processor 28 reads the specific processing program 56 from the storage 32 and executes the read specific processing program 56 on the RAM 30. The specific processing is realized by the processor 28 operating as a specific processing unit 290 according to the specific processing program 56 executed on the RAM 30.
[0033] The storage 32 stores the data generation model 58 and the emotion identification model 59. The data generation model 58 and the emotion identification model 59 are used by the identification processing unit 290.
[0034] In the smart device 14, the processor 46 performs the reception output processing. The storage 50 stores the reception output program 60. The reception output program 60 is used in conjunction with a specific processing program 56 by the data processing system 10. The processor 46 reads the reception output program 60 from the storage 50 and executes the read reception output program 60 on the RAM 48. The reception output processing is realized by the processor 46 operating as a control unit 46A according to the reception output program 60 executed on the RAM 48.
[0035] Next, the specific processing performed by the specific processing unit 290 of the data processing device 12 will be described. In the following description, the data processing device 12 will be referred to as the "server" and the smart device 14 as the "terminal".
[0036] This invention is a system designed to reduce the risk of fraudulent billing scams and scams impersonating family members that lurk in users' daily lives. This system primarily acquires the user's voice, converts that voice into text data, and analyzes it to detect potential fraud and issue appropriate warnings.
[0037] First, the device continuously monitors the user's conversation and surrounding sounds in the background to acquire audio data in real time. This audio data is converted into text data using the device's built-in speech recognition function, preparing it for analysis.
[0038] Once the audio is converted into text data, the device uses generative AI to analyze the content of the text data. This generative AI model is trained on a dataset containing characteristic patterns of fraud, and it determines the likelihood of fraudulent activity based on the analysis results.
[0039] If a device is determined to be potentially fraudulent, it will immediately send an alert to the designated contacts. These contacts primarily include family and close friends, but depending on usage, it can also notify law enforcement authorities. Users can choose from a variety of alert delivery methods, including email, SMS, and push notifications.
[0040] For example, if a user receives a phone call stating, "You have outstanding charges," the device immediately analyzes the phrase and detects keywords such as "outstanding" and "charges." The generating AI then determines that there is a high risk of fraud and quickly issues a warning to the user and their registered contacts. This allows the user to take appropriate action without complying with fraudulent demands.
[0041] Furthermore, users can freely customize how the system notifies them of warnings, and if the response is insufficient, they can manually use contact information to provide follow-up support. This flexibility allows the system to be adapted to the user's actual usage and lifestyle.
[0042] This will help prevent fraud and support users in living their daily lives with peace of mind.
[0043] The following describes the processing flow.
[0044] Step 1:
[0045] The device will always have a function enabled to receive the user's voice through the microphone. Voice acquisition will be initiated by a configured trigger event (for example, detection that a conversation has started).
[0046] Step 2:
[0047] The device converts the acquired audio into text data using a speech recognition system. The speech recognition system transcribes the audio in real time using a pre-installed API.
[0048] Step 3:
[0049] The device passes the converted text data to a generating AI model, which analyzes whether the text contains phrases or patterns that indicate fraud. The generating AI is a model that has learned from past fraud cases and has the ability to detect suspicious patterns.
[0050] Step 4:
[0051] If the device receives the results of its analysis of the generated AI model and determines that it may be fraudulent, it prepares to take action to issue a warning according to pre-configured rules.
[0052] Step 5:
[0053] The device will send alerts. These notifications are sent as push notifications to the user's device, as well as via email and SMS to pre-configured contacts (family, police, etc.). The method of sending notifications is pre-configured by the user.
[0054] Step 6:
[0055] Users can receive alerts and follow the instructions displayed on their device to review audio logs and analysis results of suspicious conversations. If necessary, they can also report or request further confirmation from administrators via their device.
[0056] (Example 1)
[0057] Next, we will describe Example 1. In the following description, the data processing device 12 will be referred to as the "server," and the smart device 14 will be referred to as the "terminal."
[0058] The present invention aims to effectively detect the risk of users being subjected to fraudulent activities or unethical solicitations, and to prevent them from becoming victims. Furthermore, it aims to surpass the limitations of conventional technologies by utilizing user feedback to improve detection accuracy.
[0059] The identification process performed by the identification processing unit 290 of the data processing device 12 in Example 1 is realized by the following means.
[0060] In this invention, the server includes means for converting the user's voice into text data, means for using generative artificial intelligence to analyze the text data and detect the possibility of fraudulent activity, and means for notifying the user of a warning based on the detection results. This enables the user to immediately recognize fraudulent activity and take prompt and appropriate action.
[0061] "User" refers to an individual or organization that uses this system and provides voice data for the purpose of detecting and preventing fraudulent activity.
[0062] A "device for acquiring voice" refers to hardware or software that records the user's voice in real time and converts that voice data into a format usable within the system.
[0063] A "device for converting to text data" refers to hardware or software that uses speech recognition technology to convert acquired audio data into analyzable text data.
[0064] "Generative artificial intelligence" refers to artificial intelligence technology equipped with an algorithm that analyzes text data converted from speech and determines the likelihood of fraudulent activity.
[0065] "Warning device" refers to a means or device for notifying a user or relevant contact if potential fraudulent activity is detected.
[0066] A "device for sending notifications" refers to a means of sending alerts to designated contacts using a communication network, and includes a variety of methods such as email, SMS, and push notifications.
[0067] "Devices that improve generative artificial intelligence models by providing feedback" refers to hardware and software that manage the process of improving the system's detection performance through responses and opinions from users.
[0068] This system is designed to help users proactively detect fraudulent activity and avoid its consequences. Specifically, it utilizes speech recognition and generation AI technology to analyze voice data in real time. Details are provided below.
[0069] The device constantly monitors the user's conversations and surrounding sounds in the background. For this purpose, the device uses its built-in microphone and a voice acquisition module (e.g., a common voice recording device or software). The acquired voice data is converted into text data using speech recognition software (e.g., Google® Speech-to-Text or equivalent technology).
[0070] After being converted into text data, the terminal analyzes the content using a generative AI model (e.g., OpenAI's GPT model). During this process, the generative AI model uses prompt messages to assess the likelihood of fraud in real time. Examples of specific prompt messages include "unpaid" and "urgent action required."
[0071] If the generating AI detects a potential scam, the device will immediately issue a warning. The warning will be sent to the designated contacts, but will depend on the user's custom notification method (email, SMS, push notification, etc.). This allows users and stakeholders to immediately understand the situation and take appropriate action.
[0072] Users can freely customize how they receive warning notifications from the system. Furthermore, the device collects user feedback information, which is used to improve the accuracy of the generated AI model. This allows the system to better suit the needs of individual users and enhance the effectiveness of fraud detection. With this flexibility, the system supports users in leading a safer and more secure daily life.
[0073] The flow of the specific processing in Example 1 will be explained using Figure 11.
[0074] Step 1:
[0075] The device acquires the user's voice in real time in the background. This process continuously captures ambient audio signals using the device's built-in microphone. The input is audio signals from the surrounding environment, and the output is digitized audio data. This audio data is divided into short samples for subsequent processing.
[0076] Step 2:
[0077] The device converts the acquired audio data into text data. Here, a speech recognition library (e.g., Google Speech-to-Text) is used to convert the audio signal into a string. The input is digital audio data, and data conversion is performed, including background noise removal and speech recognition. The output is formatted text data, which is used for subsequent analysis.
[0078] Step 3:
[0079] The device analyzes text data using a generating AI model to search for patterns of fraudulent activity. In this step, the AI model (e.g., a GPT model) uses prompt text to assess the likelihood of fraud. The input consists of text data and prompt text, and the AI model performs data exploration and pattern matching. The output is provided as a fraud risk score.
[0080] Step 4:
[0081] If fraud is highly likely, the device generates an alert and sends a notification to the user and their designated contacts. The notification method is pre-specified by the user and can be email, SMS, or push notification. The input is the result of an AI risk assessment, which automatically selects the notification format and issues an alert. The output is the notification message sent to the recipient.
[0082] Step 5:
[0083] The user receives a warning message and takes action. This feedback is collected by the system and used to improve the generated AI model. Depending on the user's actions, notification settings can be adjusted and the AI model can be retrained. The input is the user's response and additional information, and the output is an improvement in the system's analysis accuracy.
[0084] (Application Example 1)
[0085] Next, we will explain Application Example 1. In the following explanation, the data processing device 12 will be referred to as the "server," and the smart device 14 will be referred to as the "terminal."
[0086] In modern society, fraud involving fictitious billing and impersonation of family members is on the rise. In particular, many users, including the elderly, become victims without realizing the risks. In this situation, it is crucial for users to quickly and effectively detect and prevent these fraud risks in their daily lives.
[0087] The specific processing performed by the specific processing unit 290 of the data processing device 12 in Application Example 1 is realized by the following means.
[0088] In this invention, the server includes a device for acquiring voice information, a device for converting it into text data, and a device using generative artificial intelligence that instantly analyzes the risk of fraud. This allows users to detect and be warned in real time about the risk of fraud they may encounter in their daily lives, thereby preventing them from becoming victims of fraud.
[0089] "Audio information" refers to sound data collected from the user and their surroundings, including human conversation and ambient sounds.
[0090] "Text data" refers to data in text format generated based on audio information, and it visually represents the content of speech.
[0091] "Generative artificial intelligence" is an advanced learning model designed to analyze voice information and assess and detect the risk of fraud.
[0092] A "warning device" refers to a device or method used to inform users or designated contacts of the risk of fraud.
[0093] "Communication methods" refer to the infrastructure and technologies used to transmit information, including the internet, telephone lines, and wireless communication.
[0094] "Specified contact information" refers to the contact details of the person or organization that the user has set up in advance to receive warning notifications.
[0095] A "device that continuously monitors ambient sounds" is a device designed to constantly collect sound information and analyze it as needed.
[0096] Embodiments of this invention primarily relate to a system for acquiring audio information and immediately analyzing the risk of fraud. A server or terminal acquires audio information and converts it into text data. Specifically, it monitors surrounding conversations in real time using the microphone of a smartphone or other device and converts the audio data into text using speech recognition technology (e.g., Google Speech-to-Text API).
[0097] The converted text data is analyzed by a generative artificial intelligence model on the server or terminal. This generative AI model (e.g., OpenAI's GPT model) is trained to identify characteristic patterns associated with fraud and assesses the risk in the text data. If the analysis determines that the risk of fraud is high, a warning system is activated, sending a warning to the user or pre-designated contacts. The warning is sent via communication methods such as push notifications, email, or SMS.
[0098] For example, if a user hears a potentially fraudulent phrase such as "You have outstanding charges," the device will detect it and warn them with "Potential scam: Please check how to proceed." This allows users to confidently identify fraudulent requests and prevent themselves from becoming victims of fraud.
[0099] Examples of input prompts for a generative AI model include: "Identify potentially fraudulent phrases in the following conversation: 'You have outstanding charges.'" This allows the system to effectively detect fraud risks and protect users.
[0100] The flow of a specific process in Application Example 1 will be explained using Figure 12.
[0101] Step 1:
[0102] The server or terminal acquires voice information. It collects user and surrounding voices in real time using the device's microphone and inputs them as digital voice data. Voice information acquisition can occur continuously or based on specific triggers.
[0103] Step 2:
[0104] The server or terminal converts the acquired audio information into text data. Using speech recognition technology (e.g., Google Speech-to-Text API), it receives the audio data as input and generates the corresponding text data. This text data is output as the underlying text information for analysis.
[0105] Step 3:
[0106] The server analyzes text data using a generative AI model. The input text data is provided to the AI in the form of a prompt message, "Identify potentially fraudulent phrases in the following conversation," and the risk of fraud is assessed. The generative AI model (e.g., OpenAI's GPT model) identifies characteristic patterns of fraud and outputs a risk score.
[0107] Step 4:
[0108] The server or terminal analyzes the risk score returned by the generated AI model to determine the likelihood of fraud. If the risk score exceeds a certain threshold, it prepares to issue a warning. The output here is the result of the fraud likelihood assessment.
[0109] Step 5:
[0110] The server or terminal sends a warning to the user and their designated contact if there is a possibility of fraud. The warning is sent via a communication method such as push notification, email, or SMS, using the contact information entered as the recipient. The output is a record that the warning was successfully delivered.
[0111] Furthermore, an emotion engine that estimates the user's emotions may be incorporated. That is, the identification processing unit 290 may use the emotion identification model 59 to estimate the user's emotions and perform identification processing using the user's emotions.
[0112] This invention combines a system that analyzes the potential for fraud in real time from a user's voice with an emotion engine that analyzes the user's emotional state, thereby achieving more accurate fraud detection and warnings. Because this invention considers not only the content of the voice but also the user's emotional response, the risk of fraud is assessed more accurately.
[0113] First, the device continuously acquires the user's voice and converts it into text data using speech recognition technology. This text data is then passed to an artificial intelligence model generated for fraud detection.
[0114] Simultaneously, the device is equipped with an emotion engine that analyzes emotional responses from the user's voice tone, speaking patterns, and word choice. The emotion engine can evaluate emotional parameters such as the user's stress level and tension.
[0115] The emotional data obtained by the emotion engine is integrated into the analysis process of the generative AI model. For example, if a user says something like, "Am I really behind on payments?" on the phone, with a nuance of doubt, the device detects this unstable emotional state. This increases the risk score for fraud, and the device determines that it needs to issue an immediate warning.
[0116] Next, if a call is deemed highly likely to be a scam, the device will immediately issue a warning to the user. Upon receiving this warning, the user can quickly take countermeasures such as hanging up the potentially dangerous call or conducting further investigation. The warning includes sending alerts to designated contacts (e.g., family members or the police). The warning method can be pre-configured by the user and can be selected from options such as email, SMS, or push notification.
[0117] In this way, the present invention combines emotion recognition technology with conventional speech recognition-based fraud detection systems, making it possible to identify fraud risks early and with high accuracy, and to provide appropriate warnings. As a result, users can engage in everyday communication with peace of mind, even in situations where emotional responses may indicate the possibility of fraud.
[0118] The following describes the processing flow.
[0119] Step 1:
[0120] The device receives the user's voice through the microphone and begins recording it as audio data. In this initial stage, voice acquisition is triggered when the user starts speaking.
[0121] Step 2:
[0122] The device transmits the acquired voice data to a speech recognition system, which converts it into text data in real time. The converted text data is then prepared for analysis to detect fraud.
[0123] Step 3:
[0124] The device simultaneously uses an emotion engine to perform emotion analysis based on the user's voice. The emotion engine uses factors such as voice tone, speed, and volume to evaluate the degree of stress and tension.
[0125] Step 4:
[0126] The device inputs text data into a generating artificial intelligence model, which analyzes whether it contains patterns that indicate potential fraud. In addition, the generating AI takes into account the results of sentiment analysis to generate an overall fraud risk score.
[0127] Step 5:
[0128] If the fraud risk score exceeds a certain threshold, the device will determine it to be a high-priority situation and immediately prepare a warning.
[0129] Step 6:
[0130] The device issues a warning to the user and sends the alert to the designated contacts using a method selected depending on the situation (push notification, voice alert, email, SMS, etc.). If this information indicates particularly high levels of tension or stress through sentiment analysis, the user will receive a stronger warning.
[0131] Step 7:
[0132] Users can receive warning notifications from their devices and take appropriate measures, such as hanging up the phone early, by following the instructions. Users can also access information through their devices to take further action.
[0133] (Example 2)
[0134] Next, we will describe Example 2. In the following description, the data processing device 12 will be referred to as the "server," and the smart device 14 will be referred to as the "terminal."
[0135] There is a need to provide a system that can detect fraud with high accuracy and issue timely warnings, taking into account not only the user's voice but also their emotional responses, thereby more accurately grasping fraud risks that are often overlooked by existing systems.
[0136] The identification process performed by the identification processing unit 290 of the data processing device 12 in Example 2 is realized by the following means.
[0137] In this invention, the server includes a device for acquiring user voice, a mechanism for converting the voice into data, a configuration using a generative machine learning model for analyzing the likelihood of fraud based on the data, and means for incorporating an emotion analysis engine when evaluating the likelihood of fraud. This makes it possible to comprehensively analyze voice content and emotion data, evaluate fraud risk with high accuracy, and issue timely notifications to the user.
[0138] A "user" is an entity that utilizes the system to provide its own voice data and participate in fraud detection.
[0139] A "device for acquiring voice" refers to hardware or software used to collect a user's voice and convert it into a digital format.
[0140] A "data conversion mechanism" refers to the technology responsible for the process of converting collected audio information into text data or analyzable digital data.
[0141] A "configuration using a generative machine learning model" refers to a system built using an algorithm that analyzes user data and learns specific patterns in order to detect fraudulent activity.
[0142] An "emotion analysis engine" refers to a technology that extracts emotional responses from a user's voice and calculates their parameters.
[0143] The "notification mechanism" refers to a part of the system that sends a warning to the user or designated contacts when potential fraud is detected.
[0144] "Communication channels" refer to electronic means of communication used to transmit warnings and notifications, and include email, SMS, push notifications, etc.
[0145] A "generative AI model" refers to an artificial intelligence model designed to analyze information from data provided as arguments and generate a specific output.
[0146] The present invention is implemented by a system that determines in real time whether or not a voice provided by a user is fraudulent. The terminal includes a microphone or other voice acquisition device for acquiring the user's voice. The acquired voice is converted into text using natural language processing techniques. For example, speech recognition software can be used for this process.
[0147] After converting the audio into data, the device supplies the data to a generative machine learning model. This generative machine learning model is pre-trained to analyze the data and recognize signs of fraud. Specifically, it uses pattern recognition and classification techniques to assess the likelihood of fraud. For example, a machine learning platform can be used for this model.
[0148] Furthermore, the system analyzes the user's emotional responses from voice data through an emotion analysis engine. This engine takes into account factors such as voice tone, pitch, and speed to evaluate emotions like stress and tension. This engine can utilize, for example, emotion analysis software. The emotional data is integrated into a generative machine learning model and used for a comprehensive assessment of fraud risk.
[0149] If a potential scam is detected, the device immediately alerts the user. This alert can be configured to reach the user and their designated contacts via multiple communication channels, including email, SMS, and push notifications. For example, if a user expresses doubt, such as "Am I really behind on payments?", the change in their tone of voice can be detected as high risk by the sentiment analysis engine, and a generative machine learning model will immediately issue a warning.
[0150] The following are specific examples of prompt statements:
[0151] User's voice text: "Am I really behind on payments?" Emotion parameters: Anxiety, tension. Conduct a risk assessment and determine the possibility of fraud.
[0152] With a system configured in this way, users can accurately assess the risk of fraud based on voice content and emotional responses, and take appropriate action.
[0153] The flow of the specific processing in Example 2 will be explained using Figure 13.
[0154] Step 1:
[0155] The device acquires the user's voice through a microphone and stores it as digital audio data. The user's voice is received in real time as input, and a digital audio file is generated as output. This file is then ready for use in subsequent processing.
[0156] Step 2:
[0157] The device acquires digital audio data and converts it into text data using speech recognition software. The digital audio data acquired in step 1 is used as input, and a text-formatted string is generated as output. This conversion includes keyword extraction from the audio and analysis of phonological information.
[0158] Step 3:
[0159] The device inputs the generated text data into a generating AI model to analyze the likelihood of fraud. The text data converted in step 2 is used as input, and a fraud risk score is provided as output. At this stage, natural language processing techniques are used to perform contextual analysis and pattern detection of the text.
[0160] Step 4:
[0161] The device extracts emotional data from the user's voice using an emotion analysis engine. The digital voice data from step 1 is used again as input, and emotional parameters (e.g., stress, tension) are generated as output. This process analyzes the tone, speed, and intonation of the voice.
[0162] Step 5:
[0163] The device integrates emotional data into a generative AI model to perform a comprehensive assessment of fraud risk. The risk score from Step 3 and the emotional parameters from Step 4 are integrated as input, and the final risk assessment result is output. This integration enables a complex analysis that takes into account the impact of emotions on the likelihood of fraud.
[0164] Step 6:
[0165] Based on the final risk assessment results, the device will issue a warning to the user if it determines that there is a high probability of fraud. The risk assessment results generated in step 5 are used as input, and a notification (email, SMS, push notification) is sent to the user as output. Specific actions here include creating and sending the notification message.
[0166] (Application Example 2)
[0167] Next, we will explain application example 2. In the following explanation, the data processing device 12 will be referred to as a "server" and the smart device 14 as a "terminal".
[0168] In today's world, fraud via communications is on the rise, and users are required to protect themselves from fraudulent activities. However, conventional fraud detection systems rely on analyzing voice content and ignore emotional responses, which poses a challenge in terms of accuracy. Furthermore, it is difficult to issue warnings before users recognize actual signs of fraud, making it difficult to quickly reduce fraud risk.
[0169] The specific processing performed by the specific processing unit 290 of the data processing device 12 in Application Example 2 is realized by the following means.
[0170] In this invention, the server includes means for acquiring the user's voice, means for converting the voice into text data, means for using generative artificial intelligence to analyze the likelihood of fraud in real time based on the text data, means for sentiment analysis to evaluate the emotional state from the user's voice, means for integrating the emotional state data into the fraud analysis process, means for issuing a warning when the likelihood of fraud is detected, and means for sending an alert to a set contact via any communication means. This enables highly accurate, real-time fraud detection that takes into account the user's emotional response.
[0171] "Means for acquiring user voice" refers to a function that uses an input device to convert the voice signal emitted by the user into digital data and input it into the system.
[0172] "Means for converting speech to text data" refers to a function that processes speech signals acquired using speech recognition technology into corresponding strings of characters.
[0173] "Methods using generative artificial intelligence" refer to software that uses machine learning algorithms to analyze input data and mimic human judgment to assess the likelihood of fraud.
[0174] "An emotion analysis method for evaluating emotional state from user voice" is an algorithm that extracts the user's emotional characteristics from voice data and evaluates their psychological state.
[0175] "Means of integrating emotional state data into the fraud analysis process" refers to the process of incorporating parameters obtained through emotional analysis into a fraud detection algorithm and using them as part of the analysis.
[0176] "A means of issuing a warning when a potential scam is detected" refers to a function that notifies the user to be cautious when it is determined that there is a high probability of fraud.
[0177] "Means of sending alerts to configured contacts" refers to the process of sending warnings as emails or messages based on pre-registered contact information.
[0178] The system for carrying out the present invention mainly consists of a voice acquisition device, voice recognition software, emotion analysis engine, generative AI model, and warning notification system.
[0179] The server captures the user's voice through a microphone and converts the voice data into text in real time using the Google Speech-to-Text API. This text data is then analyzed using a generative AI model such as the BERT model to assess the likelihood of fraud. During this process, an emotion analysis engine such as openSMILE is used to evaluate the user's emotional state based on factors such as tone, rhythm, and word choice. The resulting emotion parameters are then integrated into the fraud detection algorithm to improve the accuracy of the analysis.
[0180] When a user receives a phone call, for example, if the caller asks for verification of personal information, the user might say, "Is that really true?" At this point, the system detects the user's emotional state, indicating anxiety or suspicion, from their voice and uses that data to assess the risk of fraud.
[0181] Ultimately, if the device is determined to be highly likely to be a scam, it will alert the user via push notification. It can also use communication technologies such as Firebase Cloud Messaging to send alerts via email or message to pre-configured contacts.
[0182] A concrete example of a prompt message from a generated AI model might be text like, "Please evaluate the likelihood that the following message is a scam. User: 'I'm being asked to confirm my personal information, but is this real?'" This allows users to communicate with confidence through everyday communications thanks to highly accurate scam detection.
[0183] The flow of a specific process in Application Example 2 will be explained using Figure 14.
[0184] Step 1:
[0185] The terminal uses a microphone to acquire the user's voice. The input is the user's voice signal, which is captured as digital data. This acquired voice data is then passed on to the next processing step.
[0186] Step 2:
[0187] The server uses the Google Speech-to-Text API to convert the acquired audio data into text data. The input is audio data, and the output is text data generated based on this audio. This conversion process makes the audio information suitable for analysis.
[0188] Step 3:
[0189] The server analyzes text data using a generative AI model. The input is previously converted text data, and the output is a score indicating the likelihood of fraud. This model evaluates the content of the input text based on fraud patterns that it has learned in advance.
[0190] Step 4:
[0191] The device uses an emotion analysis engine such as openSMILE to analyze the user's emotional state from voice data. The input is the initial voice data, and the output is emotion parameters based on voice tone and patterns. This analysis quantifies the user's stress level and anxiety.
[0192] Step 5:
[0193] The server integrates emotional parameters into the fraud analysis process. The emotional parameters are added to the fraud analysis score, and the final fraud risk is recalculated. This results in a more accurate fraud risk assessment that takes the user's psychological state into account.
[0194] Step 6:
[0195] The device issues a warning if the final score it generates exceeds a certain threshold. The input is the final calculated fraud risk score, and the output is a warning message. The device warns the user via push notification and, if necessary, sends an alert to the designated contacts.
[0196] Step 7:
[0197] Upon receiving a warning, the user can confirm whether to end the call. The input is the warning message from the device, and the output is the user's appropriate response. This makes it possible to prevent fraud.
[0198] The specific processing unit 290 transmits the result of the specific processing to the smart device 14. In the smart device 14, the control unit 46A causes the output device 40 to output the result of the specific processing. The microphone 38B acquires audio indicating user input for the result of the specific processing. The control unit 46A transmits the audio data indicating user input acquired by the microphone 38B to the data processing device 12. In the data processing device 12, the specific processing unit 290 acquires the audio data.
[0199] Data generation model 58 is a so-called generative AI (Artificial Intelligence). An example of data generation model 58 is ChatGPT (registered trademark) (Internet search).<URL: https: / / openai.com / blog / chatgpt> ), Gemini (registered trademark) (Internet search) <url: https: gemini.google.com ?hl="ja">Examples of generative AI include the following. The data generation model 58 is obtained by performing deep learning on a neural network. The data generation model 58 is input with prompts containing instructions, and with inference data such as audio data representing speech, text data representing text, and image data representing images. The data generation model 58 infers from the input inference data according to the instructions indicated by the prompts, and outputs the inference results in data formats such as audio data and text data. Here, inference refers to, for example, analysis, classification, prediction, and / or summarization.
[0200] In the above embodiment, an example was given in which specific processing is performed by the data processing device 12, but the technology of this disclosure is not limited thereto, and the specific processing may also be performed by the smart device 14.
[0201] [Second Embodiment]
[0202] Figure 3 shows an example of the configuration of the data processing system 210 according to the second embodiment.
[0203] As shown in Figure 3, the data processing system 210 includes a data processing device 12 and smart glasses 214. An example of the data processing device 12 is a server.
[0204] The data processing device 12 comprises a computer 22, a database 24, and a communication interface 26. The computer 22 is an example of a "computer" related to the technology of this disclosure. The computer 22 comprises a processor 28, RAM 30, and storage 32. The processor 28, RAM 30, and storage 32 are connected to a bus 34. The database 24 and the communication interface 26 are also connected to the bus 34. The communication interface 26 is connected to a network 54. An example of the network 54 is a WAN (Wide Area Network) and / or a LAN (Local Area Network).
[0205] The smart glasses 214 include a computer 36, a microphone 238, a speaker 240, a camera 42, and a communication interface 44. The computer 36 includes a processor 46, RAM 48, and storage 50. The processor 46, RAM 48, and storage 50 are connected to a bus 52. The microphone 238, speaker 240, and camera 42 are also connected to the bus 52.
[0206] The microphone 238 receives voice signals from the user 20 and receives instructions from the user 20. The microphone 238 captures the voice signals from the user 20, converts the captured voice into audio data, and outputs it to the processor 46. The speaker 240 outputs audio according to the instructions from the processor 46.
[0207] Camera 42 is a small digital camera equipped with an optical system including a lens, aperture, and shutter, and an image sensor such as a CMOS (Complementary Metal-Oxide-Semiconductor) image sensor or a CCD (Charge Coupled Device) image sensor, and captures images of the area around the user 20 (for example, an imaging range defined by a field of view equivalent to the width of a typical healthy person's field of vision).
[0208] Communication interface 44 is connected to network 54. Communication interfaces 44 and 26 are responsible for the exchange of various information between processor 46 and processor 28 via network 54. The exchange of various information between processor 46 and processor 28 using communication interfaces 44 and 26 is performed in a secure manner.
[0209] Figure 4 shows an example of the main functions of the data processing device 12 and the smart glasses 214. As shown in Figure 4, the data processing device 12 performs specific processing using the processor 28. The storage 32 stores the specific processing program 56.
[0210] The specific processing program 56 is an example of a "program" relating to the technology of this disclosure. The processor 28 reads the specific processing program 56 from the storage 32 and executes the read specific processing program 56 on the RAM 30. The specific processing is realized by the processor 28 operating as a specific processing unit 290 in accordance with the specific processing program 56 executed on the RAM 30.
[0211] The storage 32 stores the data generation model 58 and the emotion identification model 59. The data generation model 58 and the emotion identification model 59 are used by the identification processing unit 290.
[0212] In the smart glasses 214, the processor 46 performs the reception output processing. The storage 50 stores the reception output program 60. The processor 46 reads the reception output program 60 from the storage 50 and executes the read reception output program 60 on the RAM 48. The reception output processing is realized by the processor 46 operating as a control unit 46A according to the reception output program 60 executed on the RAM 48.
[0213] Next, the identification processing performed by the identification processing unit 290 of the data processing device 12 will be described. In the following description, the data processing device 12 will be referred to as the "server" and the smart glasses 214 will be referred to as the "terminal".
[0214] This invention is a system designed to reduce the risk of fraudulent billing scams and scams impersonating family members that lurk in users' daily lives. This system primarily acquires the user's voice, converts that voice into text data, and analyzes it to detect potential fraud and issue appropriate warnings.
[0215] First, the device continuously monitors the user's conversation and surrounding sounds in the background to acquire audio data in real time. This audio data is converted into text data using the device's built-in speech recognition function, preparing it for analysis.
[0216] Once the audio is converted into text data, the device uses generative AI to analyze the content of the text data. This generative AI model is trained on a dataset containing characteristic patterns of fraud, and it determines the likelihood of fraudulent activity based on the analysis results.
[0217] If a device is determined to be potentially fraudulent, it will immediately send an alert to the designated contacts. These contacts primarily include family and close friends, but depending on usage, it can also notify law enforcement authorities. Users can choose from a variety of alert delivery methods, including email, SMS, and push notifications.
[0218] For example, if a user receives a phone call stating, "You have outstanding charges," the device immediately analyzes the phrase and detects keywords such as "outstanding" and "charges." The generating AI then determines that there is a high risk of fraud and quickly issues a warning to the user and their registered contacts. This allows the user to take appropriate action without complying with fraudulent demands.
[0219] Furthermore, users can freely customize how the system notifies them of warnings, and if the response is insufficient, they can manually use contact information to provide follow-up support. This flexibility allows the system to be adapted to the user's actual usage and lifestyle.
[0220] This will help prevent fraud and support users in living their daily lives with peace of mind.
[0221] The following describes the processing flow.
[0222] Step 1:
[0223] The device will always have a function enabled to receive the user's voice through the microphone. Voice acquisition will be initiated by a configured trigger event (for example, detection that a conversation has started).
[0224] Step 2:
[0225] The device converts the acquired audio into text data using a speech recognition system. The speech recognition system transcribes the audio in real time using a pre-installed API.
[0226] Step 3:
[0227] The device passes the converted text data to a generating AI model, which analyzes whether the text contains phrases or patterns that indicate fraud. The generating AI is a model that has learned from past fraud cases and has the ability to detect suspicious patterns.
[0228] Step 4:
[0229] If the device receives the results of its analysis of the generated AI model and determines that it may be fraudulent, it prepares to take action to issue a warning according to pre-configured rules.
[0230] Step 5:
[0231] The device will send alerts. These notifications are sent as push notifications to the user's device, as well as via email and SMS to pre-configured contacts (family, police, etc.). The method of sending notifications is pre-configured by the user.
[0232] Step 6:
[0233] Users can receive alerts and follow the instructions displayed on their device to review audio logs and analysis results of suspicious conversations. If necessary, they can also report or request further confirmation from administrators via their device.
[0234] (Example 1)
[0235] Next, we will describe Example 1. In the following description, the data processing device 12 will be referred to as the "server," and the smart glasses 214 will be referred to as the "terminal."
[0236] The present invention aims to effectively detect the risk of users being subjected to fraudulent activities or unethical solicitations, and to prevent them from becoming victims. Furthermore, it aims to surpass the limitations of conventional technologies by utilizing user feedback to improve detection accuracy.
[0237] The identification process performed by the identification processing unit 290 of the data processing device 12 in Example 1 is realized by the following means.
[0238] In this invention, the server includes means for converting the user's voice into text data, means for using generative artificial intelligence to analyze the text data and detect the possibility of fraudulent activity, and means for notifying the user of a warning based on the detection results. This enables the user to immediately recognize fraudulent activity and take prompt and appropriate action.
[0239] "User" refers to an individual or organization that uses this system and provides voice data for the purpose of detecting and preventing fraudulent activity.
[0240] A "device for acquiring voice" refers to hardware or software that records the user's voice in real time and converts that voice data into a format usable within the system.
[0241] A "device for converting to text data" refers to hardware or software that uses speech recognition technology to convert acquired audio data into analyzable text data.
[0242] "Generative artificial intelligence" refers to artificial intelligence technology equipped with an algorithm that analyzes text data converted from speech and determines the likelihood of fraudulent activity.
[0243] "Warning device" refers to a means or device for notifying a user or relevant contact if potential fraudulent activity is detected.
[0244] A "device for sending notifications" refers to a means of sending alerts to designated contacts using a communication network, and includes a variety of methods such as email, SMS, and push notifications.
[0245] "Devices that improve generative artificial intelligence models by providing feedback" refers to hardware and software that manage the process of improving the system's detection performance through responses and opinions from users.
[0246] This system is designed to help users proactively detect fraudulent activity and avoid its consequences. Specifically, it utilizes speech recognition and generation AI technology to analyze voice data in real time. Details are provided below.
[0247] The device constantly monitors the user's conversations and surrounding sounds in the background. For this purpose, the device uses its built-in microphone and a voice acquisition module (e.g., a common voice recording device or software). The acquired voice data is converted into text data using speech recognition software (e.g., Google Speech-to-Text or equivalent technology).
[0248] After being converted into text data, the terminal analyzes the content using a generative AI model (e.g., OpenAI's GPT model). In this process, the generative AI model uses prompt messages to assess the likelihood of fraud in real time. Examples of specific prompt messages include "unpaid" and "urgent action required."
[0249] If the generating AI detects a potential scam, the device will immediately issue a warning. The warning will be sent to the designated contacts, but will depend on the user's custom notification method (email, SMS, push notification, etc.). This allows users and stakeholders to immediately understand the situation and take appropriate action.
[0250] Users can freely customize how they receive warning notifications from the system. Furthermore, the device collects user feedback information, which is used to improve the accuracy of the generated AI model. This allows the system to better suit the needs of individual users and enhance the effectiveness of fraud detection. With this flexibility, the system supports users in leading a safer and more secure daily life.
[0251] The flow of the specific processing in Example 1 will be explained using Figure 11.
[0252] Step 1:
[0253] The device acquires the user's voice in real time in the background. This process continuously captures ambient audio signals using the device's built-in microphone. The input is audio signals from the surrounding environment, and the output is digitized audio data. This audio data is divided into short samples for subsequent processing.
[0254] Step 2:
[0255] The device converts the acquired audio data into text data. Here, a speech recognition library (e.g., Google Speech-to-Text) is used to convert the audio signal into a string. The input is digital audio data, and data conversion is performed, including background noise removal and speech recognition. The output is formatted text data, which is used for subsequent analysis.
[0256] Step 3:
[0257] The device analyzes text data using a generating AI model to search for patterns of fraudulent activity. In this step, the AI model (e.g., a GPT model) uses prompt text to assess the likelihood of fraud. The input consists of text data and prompt text, and the AI model performs data exploration and pattern matching. The output is provided as a fraud risk score.
[0258] Step 4:
[0259] If fraud is highly likely, the device generates an alert and sends a notification to the user and their designated contacts. The notification method is pre-specified by the user and can be email, SMS, or push notification. The input is the result of an AI risk assessment, which automatically selects the notification format and issues an alert. The output is the notification message sent to the recipient.
[0260] Step 5:
[0261] The user receives a warning message and takes action. This feedback is collected by the system and used to improve the generated AI model. Depending on the user's actions, notification settings can be adjusted and the AI model can be retrained. The input is the user's response and additional information, and the output is an improvement in the system's analysis accuracy.
[0262] (Application Example 1)
[0263] Next, we will explain Application Example 1. In the following explanation, the data processing device 12 will be referred to as the "server," and the smart glasses 214 will be referred to as the "terminal."
[0264] In modern society, fraud involving fictitious billing and impersonation of family members is on the rise. In particular, many users, including the elderly, become victims without realizing the risks. In this situation, it is crucial for users to quickly and effectively detect and prevent these fraud risks in their daily lives.
[0265] The specific processing performed by the specific processing unit 290 of the data processing device 12 in Application Example 1 is realized by the following means.
[0266] In this invention, the server includes a device for acquiring voice information, a device for converting it into text data, and a device using generative artificial intelligence that instantly analyzes the risk of fraud. This allows users to detect and be warned in real time about the risk of fraud they may encounter in their daily lives, thereby preventing them from becoming victims of fraud.
[0267] "Audio information" refers to sound data collected from the user and their surroundings, including human conversation and ambient sounds.
[0268] "Text data" refers to data in text format generated based on audio information, and it visually represents the content of speech.
[0269] "Generative artificial intelligence" is an advanced learning model designed to analyze voice information and assess and detect the risk of fraud.
[0270] A "warning device" refers to a device or method used to inform users or designated contacts of the risk of fraud.
[0271] "Communication methods" refer to the infrastructure and technologies used to transmit information, including the internet, telephone lines, and wireless communication.
[0272] "Specified contact information" refers to the contact details of the person or organization that the user has set up in advance to receive warning notifications.
[0273] A "device that continuously monitors ambient sounds" is a device designed to constantly collect sound information and analyze it as needed.
[0274] Embodiments of this invention primarily relate to a system for acquiring audio information and immediately analyzing the risk of fraud. A server or terminal acquires audio information and converts it into text data. Specifically, it monitors surrounding conversations in real time using the microphone of a smartphone or other device and converts the audio data into text using speech recognition technology (e.g., Google Speech-to-Text API).
[0275] The converted text data is analyzed by a generative artificial intelligence model on the server or terminal. This generative AI model (e.g., OpenAI's GPT model) is trained to identify characteristic patterns associated with fraud and assesses the risk in the text data. If the analysis determines that the risk of fraud is high, a warning system is activated, sending a warning to the user or pre-designated contacts. The warning is sent via communication methods such as push notifications, email, or SMS.
[0276] For example, if a user hears a potentially fraudulent phrase such as "You have outstanding charges," the device will detect it and warn them with "Potential scam: Please check how to proceed." This allows users to confidently identify fraudulent requests and prevent themselves from becoming victims of fraud.
[0277] Examples of input prompts for a generative AI model include: "Identify potentially fraudulent phrases in the following conversation: 'You have outstanding charges.'" This allows the system to effectively detect fraud risks and protect users.
[0278] The flow of a specific process in Application Example 1 will be explained using Figure 12.
[0279] Step 1:
[0280] The server or terminal acquires voice information. It collects user and surrounding voices in real time using the device's microphone and inputs them as digital voice data. Voice information acquisition can occur continuously or based on specific triggers.
[0281] Step 2:
[0282] The server or terminal converts the acquired voice information into character data. Using voice recognition technology (e.g., Google Speech-to-Text API), it receives voice data as input and generates corresponding character data. This character data is output as text information that serves as the basis for analysis.
[0283] Step 3:
[0284] The server analyzes the character data using a generative AI model. The input character data is provided to the AI in the form of a prompt sentence "Identify phrases that may be fraudulent in the next conversation" to evaluate the risk of fraud. The generative AI model (e.g., OpenAI's GPT model) identifies characteristic patterns of fraud and outputs a risk score.
[0285] Step 4:
[0286] The server or terminal analyzes the risk score returned from the generative AI model to determine the likelihood of fraud. If the risk score exceeds a certain threshold, it prepares to issue a warning. The output here is the result of the determination of the likelihood of fraud.
[0287] Step 5:
[0288] If there is a likelihood of fraud, the server or terminal sends a warning to the user and the designated contact. The warning is sent via communication means such as push notification, email, or SMS, and the contact information entered as the recipient is used. The output is a record that the warning was correctly transmitted.
[0289] Furthermore, an emotion engine for estimating the user's emotion may be combined. That is, the specific processing unit 290 may estimate the user's emotion using the emotion identification model of 59 and perform specific processing using the user's emotion.
[0290] This invention combines a system that analyzes the potential for fraud in real time from a user's voice with an emotion engine that analyzes the user's emotional state, thereby achieving more accurate fraud detection and warnings. Because this invention considers not only the content of the voice but also the user's emotional response, the risk of fraud is assessed more accurately.
[0291] First, the device continuously acquires the user's voice and converts it into text data using speech recognition technology. This text data is then passed to an artificial intelligence model generated for fraud detection.
[0292] Simultaneously, the device is equipped with an emotion engine that analyzes emotional responses from the user's voice tone, speaking patterns, and word choice. The emotion engine can evaluate emotional parameters such as the user's stress level and tension.
[0293] The emotional data obtained by the emotion engine is integrated into the analysis process of the generative AI model. For example, if a user says something like, "Am I really behind on payments?" on the phone, with a nuance of doubt, the device detects this unstable emotional state. This increases the risk score for fraud, and the device determines that it needs to issue an immediate warning.
[0294] Next, if a call is deemed highly likely to be a scam, the device will immediately issue a warning to the user. Upon receiving this warning, the user can quickly take countermeasures such as hanging up the potentially dangerous call or conducting further investigation. The warning includes sending alerts to designated contacts (e.g., family members or the police). The warning method can be pre-configured by the user and can be selected from options such as email, SMS, or push notification.
[0295] In this way, the present invention combines emotion recognition technology with conventional speech recognition-based fraud detection systems, making it possible to identify fraud risks early and with high accuracy, and to provide appropriate warnings. As a result, users can engage in everyday communication with peace of mind, even in situations where emotional responses may indicate the possibility of fraud.
[0296] The following describes the processing flow.
[0297] Step 1:
[0298] The device receives the user's voice through the microphone and begins recording it as audio data. In this initial stage, voice acquisition is triggered when the user starts speaking.
[0299] Step 2:
[0300] The device transmits the acquired voice data to a speech recognition system, which converts it into text data in real time. The converted text data is then prepared for analysis to detect fraud.
[0301] Step 3:
[0302] The device simultaneously uses an emotion engine to perform emotion analysis based on the user's voice. The emotion engine uses factors such as voice tone, speed, and volume to evaluate the degree of stress and tension.
[0303] Step 4:
[0304] The device inputs text data into a generating artificial intelligence model, which analyzes whether it contains patterns that indicate potential fraud. In addition, the generating AI takes into account the results of sentiment analysis to generate an overall fraud risk score.
[0305] Step 5:
[0306] When the fraud risk score of the terminal exceeds a certain threshold, it determines that the situation is urgent and immediately prepares a warning.
[0307] Step 6:
[0308] The terminal issues a warning to the user and sends an alert to the contacts set in the method (push notification, voice alert, email, SMS, etc.) selected according to the situation. This information will result in a stronger warning to the user if sentiment analysis indicates particularly high tension or stress.
[0309] Step 7:
[0310] The user receives the warning notification sent from the terminal and can take appropriate measures such as cutting off the call early according to the instructions. The user can also check through the terminal the information for subsequent necessary responses.
[0311] (Example 2)
[0312] Next, Example 2 will be described. In the following description, the data processing device 12 is referred to as the "server", and the smart glasses 214 are referred to as the "terminal".
[0313] There is a need to more accurately grasp fraud risks that are often overlooked by existing systems by providing a system that can detect fraud with high accuracy and issue warnings in a timely manner, taking into account not only the user's voice but also emotional reactions.
[0314] The specific processing by the specific processing unit 290 of the data processing device 12 in Example 2 is realized by the following means.
[0315] In this invention, the server includes a device for acquiring user voice, a mechanism for converting the voice into data, a configuration using a generative machine learning model for analyzing the likelihood of fraud based on the data, and means for incorporating an emotion analysis engine when evaluating the likelihood of fraud. This makes it possible to comprehensively analyze voice content and emotion data, evaluate fraud risk with high accuracy, and issue timely notifications to the user.
[0316] A "user" is an entity that utilizes the system to provide its own voice data and participate in fraud detection.
[0317] A "device for acquiring voice" refers to hardware or software used to collect a user's voice and convert it into a digital format.
[0318] A "data conversion mechanism" refers to the technology responsible for the process of converting collected audio information into text data or analyzable digital data.
[0319] A "configuration using a generative machine learning model" refers to a system built using an algorithm that analyzes user data and learns specific patterns in order to detect fraudulent activity.
[0320] An "emotion analysis engine" refers to a technology that extracts emotional responses from a user's voice and calculates their parameters.
[0321] The "notification mechanism" refers to a part of the system that sends a warning to the user or designated contacts when potential fraud is detected.
[0322] "Communication channels" refer to electronic means of communication used to transmit warnings and notifications, and include email, SMS, push notifications, etc.
[0323] A "generative AI model" refers to an artificial intelligence model designed to analyze information from data provided as arguments and generate a specific output.
[0324] The present invention is implemented by a system that determines in real time whether or not a voice provided by a user is fraudulent. The terminal includes a microphone or other voice acquisition device for acquiring the user's voice. The acquired voice is converted into text using natural language processing techniques. For example, speech recognition software can be used for this process.
[0325] After converting the audio into data, the device supplies the data to a generative machine learning model. This generative machine learning model is pre-trained to analyze the data and recognize signs of fraud. Specifically, it uses pattern recognition and classification techniques to assess the likelihood of fraud. For example, a machine learning platform can be used for this model.
[0326] Furthermore, the system analyzes the user's emotional responses from voice data through an emotion analysis engine. This engine takes into account factors such as voice tone, pitch, and speed to evaluate emotions like stress and tension. This engine can utilize, for example, emotion analysis software. The emotional data is integrated into a generative machine learning model and used for a comprehensive assessment of fraud risk.
[0327] If a potential scam is detected, the device immediately alerts the user. This alert can be configured to reach the user and their designated contacts via multiple communication channels, including email, SMS, and push notifications. For example, if a user expresses doubt, such as "Am I really behind on payments?", the change in their tone of voice can be detected as high risk by the sentiment analysis engine, and a generative machine learning model will immediately issue a warning.
[0328] The following are specific examples of prompt statements:
[0329] User's voice text: "Am I really behind on payments?" Emotion parameters: Anxiety, tension. Conduct a risk assessment and determine the possibility of fraud.
[0330] With a system configured in this way, users can accurately assess the risk of fraud based on voice content and emotional responses, and take appropriate action.
[0331] The flow of the specific processing in Example 2 will be explained using Figure 13.
[0332] Step 1:
[0333] The device acquires the user's voice through a microphone and stores it as digital audio data. The user's voice is received in real time as input, and a digital audio file is generated as output. This file is then ready for use in subsequent processing.
[0334] Step 2:
[0335] The device acquires digital audio data and converts it into text data using speech recognition software. The digital audio data acquired in step 1 is used as input, and a text-formatted string is generated as output. This conversion includes keyword extraction from the audio and analysis of phonological information.
[0336] Step 3:
[0337] The device inputs the generated text data into a generating AI model to analyze the likelihood of fraud. The text data converted in step 2 is used as input, and a fraud risk score is provided as output. At this stage, natural language processing techniques are used to perform contextual analysis and pattern detection of the text.
[0338] Step 4:
[0339] The device extracts emotional data from the user's voice using an emotion analysis engine. The digital voice data from step 1 is used again as input, and emotional parameters (e.g., stress, tension) are generated as output. This process analyzes the tone, speed, and intonation of the voice.
[0340] Step 5:
[0341] The device integrates emotional data into a generative AI model to perform a comprehensive assessment of fraud risk. The risk score from Step 3 and the emotional parameters from Step 4 are integrated as input, and the final risk assessment result is output. This integration enables a complex analysis that takes into account the impact of emotions on the likelihood of fraud.
[0342] Step 6:
[0343] Based on the final risk assessment results, the device will issue a warning to the user if it determines that there is a high probability of fraud. The risk assessment results generated in step 5 are used as input, and a notification (email, SMS, push notification) is sent to the user as output. Specific actions here include creating and sending the notification message.
[0344] (Application Example 2)
[0345] Next, we will explain application example 2. In the following explanation, the data processing device 12 will be referred to as the "server," and the smart glasses 214 will be referred to as the "terminal."
[0346] In today's world, fraud via communications is on the rise, and users are required to protect themselves from fraudulent activities. However, conventional fraud detection systems rely on analyzing voice content and ignore emotional responses, which poses a challenge in terms of accuracy. Furthermore, it is difficult to issue warnings before users recognize actual signs of fraud, making it difficult to quickly reduce fraud risk.
[0347] The specific processing performed by the specific processing unit 290 of the data processing device 12 in Application Example 2 is realized by the following means.
[0348] In this invention, the server includes means for acquiring the user's voice, means for converting the voice into text data, means for using generative artificial intelligence to analyze the likelihood of fraud in real time based on the text data, means for sentiment analysis to evaluate the emotional state from the user's voice, means for integrating the emotional state data into the fraud analysis process, means for issuing a warning when the likelihood of fraud is detected, and means for sending an alert to a set contact via any communication means. This enables highly accurate, real-time fraud detection that takes into account the user's emotional response.
[0349] "Means for acquiring user voice" refers to a function that uses an input device to convert the voice signal emitted by the user into digital data and input it into the system.
[0350] "Means for converting speech to text data" refers to a function that processes speech signals acquired using speech recognition technology into corresponding strings of characters.
[0351] "Methods using generative artificial intelligence" refer to software that uses machine learning algorithms to analyze input data and mimic human judgment to assess the likelihood of fraud.
[0352] "An emotion analysis method for evaluating emotional state from user voice" is an algorithm that extracts the user's emotional characteristics from voice data and evaluates their psychological state.
[0353] "Means of integrating emotional state data into the fraud analysis process" refers to the process of incorporating parameters obtained through emotional analysis into a fraud detection algorithm and using them as part of the analysis.
[0354] "A means of issuing a warning when a potential scam is detected" refers to a function that notifies the user to be cautious when it is determined that there is a high probability of fraud.
[0355] "Means of sending alerts to configured contacts" refers to the process of sending warnings as emails or messages based on pre-registered contact information.
[0356] The system for carrying out the present invention mainly consists of a voice acquisition device, voice recognition software, emotion analysis engine, generative AI model, and warning notification system.
[0357] The server captures the user's voice through a microphone and converts the voice data into text in real time using the Google Speech-to-Text API. This text data is then analyzed using a generative AI model such as the BERT model to assess the likelihood of fraud. During this process, an emotion analysis engine such as openSMILE is used to evaluate the user's emotional state based on factors such as tone, rhythm, and word choice. The resulting emotion parameters are then integrated into the fraud detection algorithm to improve the accuracy of the analysis.
[0358] When a user receives a phone call, for example, if the caller asks for verification of personal information, the user might say, "Is that really true?" At this point, the system detects the user's emotional state, indicating anxiety or suspicion, from their voice and uses that data to assess the risk of fraud.
[0359] Ultimately, if the device is determined to be highly likely to be a scam, it will alert the user via push notification. It can also use communication technologies such as Firebase Cloud Messaging to send alerts via email or message to pre-configured contacts.
[0360] A concrete example of a prompt message from a generated AI model might be text like, "Please evaluate the likelihood that the following message is a scam. User: 'I'm being asked to confirm my personal information, but is this real?'" This allows users to communicate with confidence through everyday communications thanks to highly accurate scam detection.
[0361] The flow of a specific process in Application Example 2 will be explained using Figure 14.
[0362] Step 1:
[0363] The terminal uses a microphone to acquire the user's voice. The input is the user's voice signal, which is captured as digital data. This acquired voice data is then passed on to the next processing step.
[0364] Step 2:
[0365] The server uses the Google Speech-to-Text API to convert the acquired audio data into text data. The input is audio data, and the output is text data generated based on this audio. This conversion process makes the audio information suitable for analysis.
[0366] Step 3:
[0367] The server analyzes text data using a generative AI model. The input is previously converted text data, and the output is a score indicating the likelihood of fraud. This model evaluates the content of the input text based on fraud patterns that it has learned in advance.
[0368] Step 4:
[0369] The device uses an emotion analysis engine such as openSMILE to analyze the user's emotional state from voice data. The input is the initial voice data, and the output is emotion parameters based on voice tone and patterns. This analysis quantifies the user's stress level and anxiety.
[0370] Step 5:
[0371] The server integrates emotional parameters into the fraud analysis process. The emotional parameters are added to the fraud analysis score, and the final fraud risk is recalculated. This results in a more accurate fraud risk assessment that takes the user's psychological state into account.
[0372] Step 6:
[0373] The device issues a warning if the final score it generates exceeds a certain threshold. The input is the final calculated fraud risk score, and the output is a warning message. The device warns the user via push notification and, if necessary, sends an alert to the designated contacts.
[0374] Step 7:
[0375] Upon receiving a warning, the user can confirm whether to end the call. The input is the warning message from the device, and the output is the user's appropriate response. This makes it possible to prevent fraud.
[0376] The specific processing unit 290 transmits the result of the specific processing to the smart glasses 214. In the smart glasses 214, the control unit 46A causes the speaker 240 to output the result of the specific processing. The microphone 238 acquires audio indicating user input for the result of the specific processing. The control unit 46A transmits the audio data indicating user input acquired by the microphone 238 to the data processing unit 12. In the data processing unit 12, the specific processing unit 290 acquires the audio data.
[0377] Data generation model 58 is a type of so-called generative AI (Artificial Intelligence). One example of data generation model 58 is ChatGPT (Internet search<URL: https: / / openai.com / blog / chatgpt> ), Gemini (Internet search) <url: https: gemini.google.com ?hl="ja">Examples of generative AI include the following. The data generation model 58 is obtained by performing deep learning on a neural network. The data generation model 58 is input with prompts containing instructions, and with inference data such as audio data representing speech, text data representing text, and image data representing images. The data generation model 58 infers from the input inference data according to the instructions indicated by the prompts, and outputs the inference results in data formats such as audio data and text data. Here, inference refers to, for example, analysis, classification, prediction, and / or summarization.
[0378] In the above embodiment, an example was given in which specific processing is performed by the data processing device 12, but the technology of this disclosure is not limited thereto, and the specific processing may also be performed by the smart glasses 214.
[0379] [Third Embodiment]
[0380] Figure 5 shows an example of the configuration of the data processing system 310 according to the third embodiment.
[0381] As shown in Figure 5, the data processing system 310 includes a data processing device 12 and a headset terminal 314. An example of the data processing device 12 is a server.
[0382] The data processing device 12 comprises a computer 22, a database 24, and a communication interface 26. The computer 22 is an example of a "computer" related to the technology of this disclosure. The computer 22 comprises a processor 28, RAM 30, and storage 32. The processor 28, RAM 30, and storage 32 are connected to a bus 34. The database 24 and the communication interface 26 are also connected to the bus 34. The communication interface 26 is connected to a network 54. An example of the network 54 is a WAN (Wide Area Network) and / or a LAN (Local Area Network).
[0383] The headset terminal 314 includes a computer 36, a microphone 238, a speaker 240, a camera 42, a communication interface 44, and a display 343. The computer 36 includes a processor 46, RAM 48, and storage 50. The processor 46, RAM 48, and storage 50 are connected to a bus 52. The microphone 238, speaker 240, camera 42, and display 343 are also connected to the bus 52.
[0384] The microphone 238 receives voice signals from the user 20 and receives instructions from the user 20. The microphone 238 captures the voice signals from the user 20, converts the captured voice into audio data, and outputs it to the processor 46. The speaker 240 outputs audio according to the instructions from the processor 46.
[0385] Camera 42 is a small digital camera equipped with an optical system including a lens, aperture, and shutter, and an image sensor such as a CMOS (Complementary Metal-Oxide-Semiconductor) image sensor or a CCD (Charge Coupled Device) image sensor, and captures images of the area around the user 20 (for example, an imaging range defined by a field of view equivalent to the width of a typical healthy person's field of vision).
[0386] Communication interface 44 is connected to network 54. Communication interfaces 44 and 26 are responsible for the exchange of various information between processor 46 and processor 28 via network 54. The exchange of various information between processor 46 and processor 28 using communication interfaces 44 and 26 is performed in a secure manner.
[0387] Figure 6 shows an example of the main functions of the data processing device 12 and the headset terminal 314. As shown in Figure 6, the data processing device 12 performs specific processing using the processor 28. The storage 32 stores the specific processing program 56.
[0388] The specific processing program 56 is an example of a "program" relating to the technology of this disclosure. The processor 28 reads the specific processing program 56 from the storage 32 and executes the read specific processing program 56 on the RAM 30. The specific processing is realized by the processor 28 operating as a specific processing unit 290 in accordance with the specific processing program 56 executed on the RAM 30.
[0389] The storage 32 stores the data generation model 58 and the emotion identification model 59. The data generation model 58 and the emotion identification model 59 are used by the identification processing unit 290.
[0390] In the headset terminal 314, the processor 46 performs the reception output processing. The storage 50 stores the reception output program 60. The processor 46 reads the reception output program 60 from the storage 50 and executes the read reception output program 60 on the RAM 48. The reception output processing is realized by the processor 46 operating as a control unit 46A according to the reception output program 60 executed on the RAM 48.
[0391] Next, the specific processing performed by the specific processing unit 290 of the data processing device 12 will be described. In the following description, the data processing device 12 will be referred to as the "server" and the headset terminal 314 will be referred to as the "terminal".
[0392] This invention is a system designed to reduce the risk of fraudulent billing scams and scams impersonating family members that lurk in users' daily lives. This system primarily acquires the user's voice, converts that voice into text data, and analyzes it to detect potential fraud and issue appropriate warnings.
[0393] First, the device continuously monitors the user's conversation and surrounding sounds in the background to acquire audio data in real time. This audio data is converted into text data using the device's built-in speech recognition function, preparing it for analysis.
[0394] Once the audio is converted into text data, the device uses generative AI to analyze the content of the text data. This generative AI model is trained on a dataset containing characteristic patterns of fraud, and it determines the likelihood of fraudulent activity based on the analysis results.
[0395] If a device is determined to be potentially fraudulent, it will immediately send an alert to the designated contacts. These contacts primarily include family and close friends, but depending on usage, it can also notify law enforcement authorities. Users can choose from a variety of alert delivery methods, including email, SMS, and push notifications.
[0396] For example, if a user receives a phone call stating, "You have outstanding charges," the device immediately analyzes the phrase and detects keywords such as "outstanding" and "charges." The generating AI then determines that there is a high risk of fraud and quickly issues a warning to the user and their registered contacts. This allows the user to take appropriate action without complying with fraudulent demands.
[0397] Furthermore, users can freely customize how the system notifies them of warnings, and if the response is insufficient, they can manually use contact information to provide follow-up support. This flexibility allows the system to be adapted to the user's actual usage and lifestyle.
[0398] This will help prevent fraud and support users in living their daily lives with peace of mind.
[0399] The following describes the processing flow.
[0400] Step 1:
[0401] The device will always have a function enabled to receive the user's voice through the microphone. Voice acquisition will be initiated by a configured trigger event (for example, detection that a conversation has started).
[0402] Step 2:
[0403] The device converts the acquired audio into text data using a speech recognition system. The speech recognition system transcribes the audio in real time using a pre-installed API.
[0404] Step 3:
[0405] The device passes the converted text data to a generating AI model, which analyzes whether the text contains phrases or patterns that indicate fraud. The generating AI is a model that has learned from past fraud cases and has the ability to detect suspicious patterns.
[0406] Step 4:
[0407] If the device receives the results of its analysis of the generated AI model and determines that it may be fraudulent, it prepares to take action to issue a warning according to pre-configured rules.
[0408] Step 5:
[0409] The device will send alerts. These notifications are sent as push notifications to the user's device, as well as via email and SMS to pre-configured contacts (family, police, etc.). The method of sending notifications is pre-configured by the user.
[0410] Step 6:
[0411] Users can receive alerts and follow the instructions displayed on their device to review audio logs and analysis results of suspicious conversations. If necessary, they can also report or request further confirmation from administrators via their device.
[0412] (Example 1)
[0413] Next, we will describe Example 1. In the following description, the data processing device 12 will be referred to as the "server," and the headset-type terminal 314 will be referred to as the "terminal."
[0414] The present invention aims to effectively detect the risk of users being subjected to fraudulent activities or unethical solicitations, and to prevent them from becoming victims. Furthermore, it aims to surpass the limitations of conventional technologies by utilizing user feedback to improve detection accuracy.
[0415] The identification process performed by the identification processing unit 290 of the data processing device 12 in Example 1 is realized by the following means.
[0416] In this invention, the server includes means for converting the user's voice into text data, means for using generative artificial intelligence to analyze the text data and detect the possibility of fraudulent activity, and means for notifying the user of a warning based on the detection results. This enables the user to immediately recognize fraudulent activity and take prompt and appropriate action.
[0417] "User" refers to an individual or organization that uses this system and provides voice data for the purpose of detecting and preventing fraudulent activity.
[0418] A "device for acquiring voice" refers to hardware or software that records the user's voice in real time and converts that voice data into a format usable within the system.
[0419] A "device for converting to text data" refers to hardware or software that uses speech recognition technology to convert acquired audio data into analyzable text data.
[0420] "Generative artificial intelligence" refers to artificial intelligence technology equipped with an algorithm that analyzes text data converted from speech and determines the likelihood of fraudulent activity.
[0421] "Warning device" refers to a means or device for notifying a user or relevant contact if potential fraudulent activity is detected.
[0422] A "device for sending notifications" refers to a means of sending alerts to designated contacts using a communication network, and includes a variety of methods such as email, SMS, and push notifications.
[0423] "Devices that improve generative artificial intelligence models by providing feedback" refers to hardware and software that manage the process of improving the system's detection performance through responses and opinions from users.
[0424] This system is designed to help users proactively detect fraudulent activity and avoid its consequences. Specifically, it utilizes speech recognition and generation AI technology to analyze voice data in real time. Details are provided below.
[0425] The device constantly monitors the user's conversations and surrounding sounds in the background. For this purpose, the device uses its built-in microphone and a voice acquisition module (e.g., a common voice recording device or software). The acquired voice data is converted into text data using speech recognition software (e.g., Google Speech-to-Text or equivalent technology).
[0426] After being converted into text data, the terminal analyzes the content using a generative AI model (e.g., OpenAI's GPT model). In this process, the generative AI model uses prompt messages to assess the likelihood of fraud in real time. Examples of specific prompt messages include "unpaid" and "urgent action required."
[0427] If the generating AI detects a potential scam, the device will immediately issue a warning. The warning will be sent to the designated contacts, but will depend on the user's custom notification method (email, SMS, push notification, etc.). This allows users and stakeholders to immediately understand the situation and take appropriate action.
[0428] Users can freely customize how they receive warning notifications from the system. Furthermore, the device collects user feedback information, which is used to improve the accuracy of the generated AI model. This allows the system to better suit the needs of individual users and enhance the effectiveness of fraud detection. With this flexibility, the system supports users in leading a safer and more secure daily life.
[0429] The flow of the specific processing in Example 1 will be explained using Figure 11.
[0430] Step 1:
[0431] The device acquires the user's voice in real time in the background. This process continuously captures ambient audio signals using the device's built-in microphone. The input is audio signals from the surrounding environment, and the output is digitized audio data. This audio data is divided into short samples for subsequent processing.
[0432] Step 2:
[0433] The device converts the acquired audio data into text data. Here, a speech recognition library (e.g., Google Speech-to-Text) is used to convert the audio signal into a string. The input is digital audio data, and data conversion is performed, including background noise removal and speech recognition. The output is formatted text data, which is used for subsequent analysis.
[0434] Step 3:
[0435] The device analyzes text data using a generating AI model to search for patterns of fraudulent activity. In this step, the AI model (e.g., a GPT model) uses prompt text to assess the likelihood of fraud. The input consists of text data and prompt text, and the AI model performs data exploration and pattern matching. The output is provided as a fraud risk score.
[0436] Step 4:
[0437] If fraud is highly likely, the device generates an alert and sends a notification to the user and their designated contacts. The notification method is pre-specified by the user and can be email, SMS, or push notification. The input is the result of an AI risk assessment, which automatically selects the notification format and issues an alert. The output is the notification message sent to the recipient.
[0438] Step 5:
[0439] The user receives a warning message and takes action. This feedback is collected by the system and used to improve the generated AI model. Depending on the user's actions, notification settings can be adjusted and the AI model can be retrained. The input is the user's response and additional information, and the output is an improvement in the system's analysis accuracy.
[0440] (Application Example 1)
[0441] Next, we will explain Application Example 1. In the following explanation, the data processing device 12 will be referred to as the "server," and the headset-type terminal 314 will be referred to as the "terminal."
[0442] In modern society, fraud involving fictitious billing and impersonation of family members is on the rise. In particular, many users, including the elderly, become victims without realizing the risks. In this situation, it is crucial for users to quickly and effectively detect and prevent these fraud risks in their daily lives.
[0443] The specific processing performed by the specific processing unit 290 of the data processing device 12 in Application Example 1 is realized by the following means.
[0444] In this invention, the server includes a device for acquiring voice information, a device for converting it into text data, and a device using generative artificial intelligence that instantly analyzes the risk of fraud. This allows users to detect and be warned in real time about the risk of fraud they may encounter in their daily lives, thereby preventing them from becoming victims of fraud.
[0445] "Audio information" refers to sound data collected from the user and their surroundings, including human conversation and ambient sounds.
[0446] "Text data" refers to data in text format generated based on audio information, and it visually represents the content of speech.
[0447] "Generative artificial intelligence" is an advanced learning model designed to analyze voice information and assess and detect the risk of fraud.
[0448] A "warning device" refers to a device or method used to inform users or designated contacts of the risk of fraud.
[0449] "Communication methods" refer to the infrastructure and technologies used to transmit information, including the internet, telephone lines, and wireless communication.
[0450] "Specified contact information" refers to the contact details of the person or organization that the user has set up in advance to receive warning notifications.
[0451] A "device that continuously monitors ambient sounds" is a device designed to constantly collect sound information and analyze it as needed.
[0452] Embodiments of this invention primarily relate to a system for acquiring audio information and immediately analyzing the risk of fraud. A server or terminal acquires audio information and converts it into text data. Specifically, it monitors surrounding conversations in real time using the microphone of a smartphone or other device and converts the audio data into text using speech recognition technology (e.g., Google Speech-to-Text API).
[0453] The converted text data is analyzed by a generative artificial intelligence model on the server or terminal. This generative AI model (e.g., OpenAI's GPT model) is trained to identify characteristic patterns associated with fraud and assesses the risk in the text data. If the analysis determines that the risk of fraud is high, a warning system is activated, sending a warning to the user or pre-designated contacts. The warning is sent via communication methods such as push notifications, email, or SMS.
[0454] For example, if a user hears a potentially fraudulent phrase such as "You have outstanding charges," the device will detect it and warn them with "Potential scam: Please check how to proceed." This allows users to confidently identify fraudulent requests and prevent themselves from becoming victims of fraud.
[0455] Examples of input prompts for a generative AI model include: "Identify potentially fraudulent phrases in the following conversation: 'You have outstanding charges.'" This allows the system to effectively detect fraud risks and protect users.
[0456] The flow of a specific process in Application Example 1 will be explained using Figure 12.
[0457] Step 1:
[0458] The server or terminal acquires voice information. It collects user and surrounding voices in real time using the device's microphone and inputs them as digital voice data. Voice information acquisition can occur continuously or based on specific triggers.
[0459] Step 2:
[0460] The server or terminal converts the acquired audio information into text data. Using speech recognition technology (e.g., Google Speech-to-Text API), it receives the audio data as input and generates the corresponding text data. This text data is output as the underlying text information for analysis.
[0461] Step 3:
[0462] The server analyzes text data using a generative AI model. The input text data is provided to the AI in the form of a prompt message, "Identify potentially fraudulent phrases in the following conversation," and the risk of fraud is assessed. The generative AI model (e.g., OpenAI's GPT model) identifies characteristic patterns of fraud and outputs a risk score.
[0463] Step 4:
[0464] The server or terminal analyzes the risk score returned by the generated AI model to determine the likelihood of fraud. If the risk score exceeds a certain threshold, it prepares to issue a warning. The output here is the result of the fraud likelihood assessment.
[0465] Step 5:
[0466] The server or terminal sends a warning to the user and their designated contact if there is a possibility of fraud. The warning is sent via a communication method such as push notification, email, or SMS, using the contact information entered as the recipient. The output is a record that the warning was successfully delivered.
[0467] Furthermore, an emotion engine that estimates the user's emotions may be incorporated. That is, the identification processing unit 290 may use the emotion identification model 59 to estimate the user's emotions and perform identification processing using the user's emotions.
[0468] This invention combines a system that analyzes the potential for fraud in real time from a user's voice with an emotion engine that analyzes the user's emotional state, thereby achieving more accurate fraud detection and warnings. Because this invention considers not only the content of the voice but also the user's emotional response, the risk of fraud is assessed more accurately.
[0469] First, the device continuously acquires the user's voice and converts it into text data using speech recognition technology. This text data is then passed to an artificial intelligence model generated for fraud detection.
[0470] Simultaneously, the device is equipped with an emotion engine that analyzes emotional responses from the user's voice tone, speaking patterns, and word choice. The emotion engine can evaluate emotional parameters such as the user's stress level and tension.
[0471] The emotional data obtained by the emotion engine is integrated into the analysis process of the generative AI model. For example, if a user says something like, "Am I really behind on payments?" on the phone, with a nuance of doubt, the device detects this unstable emotional state. This increases the risk score for fraud, and the device determines that it needs to issue an immediate warning.
[0472] Next, if a call is deemed highly likely to be a scam, the device will immediately issue a warning to the user. Upon receiving this warning, the user can quickly take countermeasures such as hanging up the potentially dangerous call or conducting further investigation. The warning includes sending alerts to designated contacts (e.g., family members or the police). The warning method can be pre-configured by the user and can be selected from options such as email, SMS, or push notification.
[0473] In this way, the present invention combines emotion recognition technology with conventional speech recognition-based fraud detection systems, making it possible to identify fraud risks early and with high accuracy, and to provide appropriate warnings. As a result, users can engage in everyday communication with peace of mind, even in situations where emotional responses may indicate the possibility of fraud.
[0474] The following describes the processing flow.
[0475] Step 1:
[0476] The device receives the user's voice through the microphone and begins recording it as audio data. In this initial stage, voice acquisition is triggered when the user starts speaking.
[0477] Step 2:
[0478] The device transmits the acquired voice data to a speech recognition system, which converts it into text data in real time. The converted text data is then prepared for analysis to detect fraud.
[0479] Step 3:
[0480] The device simultaneously uses an emotion engine to perform emotion analysis based on the user's voice. The emotion engine uses factors such as voice tone, speed, and volume to evaluate the degree of stress and tension.
[0481] Step 4:
[0482] The device inputs text data into a generating artificial intelligence model, which analyzes whether it contains patterns that indicate potential fraud. In addition, the generating AI takes into account the results of sentiment analysis to generate an overall fraud risk score.
[0483] Step 5:
[0484] If the fraud risk score exceeds a certain threshold, the device will determine it to be a high-priority situation and immediately prepare a warning.
[0485] Step 6:
[0486] The device issues a warning to the user and sends the alert to the designated contacts using a method selected depending on the situation (push notification, voice alert, email, SMS, etc.). If this information indicates particularly high levels of tension or stress through sentiment analysis, the user will receive a stronger warning.
[0487] Step 7:
[0488] Users can receive warning notifications from their devices and take appropriate measures, such as hanging up the phone early, by following the instructions. Users can also access information through their devices to take further action.
[0489] (Example 2)
[0490] Next, we will describe Example 2. In the following description, the data processing device 12 will be referred to as the "server," and the headset-type terminal 314 will be referred to as the "terminal."
[0491] There is a need to provide a system that can detect fraud with high accuracy and issue timely warnings, taking into account not only the user's voice but also their emotional responses, thereby more accurately grasping fraud risks that are often overlooked by existing systems.
[0492] The identification process performed by the identification processing unit 290 of the data processing device 12 in Example 2 is realized by the following means.
[0493] In this invention, the server includes a device for acquiring user voice, a mechanism for converting the voice into data, a configuration using a generative machine learning model for analyzing the likelihood of fraud based on the data, and means for incorporating an emotion analysis engine when evaluating the likelihood of fraud. This makes it possible to comprehensively analyze voice content and emotion data, evaluate fraud risk with high accuracy, and issue timely notifications to the user.
[0494] A "user" is an entity that utilizes the system to provide its own voice data and participate in fraud detection.
[0495] A "device for acquiring voice" refers to hardware or software used to collect a user's voice and convert it into a digital format.
[0496] A "data conversion mechanism" refers to the technology responsible for the process of converting collected audio information into text data or analyzable digital data.
[0497] A "configuration using a generative machine learning model" refers to a system built using an algorithm that analyzes user data and learns specific patterns in order to detect fraudulent activity.
[0498] An "emotion analysis engine" refers to a technology that extracts emotional responses from a user's voice and calculates their parameters.
[0499] The "notification mechanism" refers to a part of the system that sends a warning to the user or designated contacts when potential fraud is detected.
[0500] "Communication channels" refer to electronic means of communication used to transmit warnings and notifications, and include email, SMS, push notifications, etc.
[0501] A "generative AI model" refers to an artificial intelligence model designed to analyze information from data provided as arguments and generate a specific output.
[0502] The present invention is implemented by a system that determines in real time whether or not a voice provided by a user is fraudulent. The terminal includes a microphone or other voice acquisition device for acquiring the user's voice. The acquired voice is converted into text using natural language processing techniques. For example, speech recognition software can be used for this process.
[0503] After converting the audio into data, the device supplies the data to a generative machine learning model. This generative machine learning model is pre-trained to analyze the data and recognize signs of fraud. Specifically, it uses pattern recognition and classification techniques to assess the likelihood of fraud. For example, a machine learning platform can be used for this model.
[0504] Furthermore, the system analyzes the user's emotional responses from voice data through an emotion analysis engine. This engine takes into account factors such as voice tone, pitch, and speed to evaluate emotions like stress and tension. This engine can utilize, for example, emotion analysis software. The emotional data is integrated into a generative machine learning model and used for a comprehensive assessment of fraud risk.
[0505] If a potential scam is detected, the device immediately alerts the user. This alert can be configured to reach the user and their designated contacts via multiple communication channels, including email, SMS, and push notifications. For example, if a user expresses doubt, such as "Am I really behind on payments?", the change in their tone of voice can be detected as high risk by the sentiment analysis engine, and a generative machine learning model will immediately issue a warning.
[0506] The following are specific examples of prompt statements:
[0507] User's voice text: "Am I really behind on payments?" Emotion parameters: Anxiety, tension. Conduct a risk assessment and determine the possibility of fraud.
[0508] With a system configured in this way, users can accurately assess the risk of fraud based on voice content and emotional responses, and take appropriate action.
[0509] The flow of the specific processing in Example 2 will be explained using Figure 13.
[0510] Step 1:
[0511] The device acquires the user's voice through a microphone and stores it as digital audio data. The user's voice is received in real time as input, and a digital audio file is generated as output. This file is then ready for use in subsequent processing.
[0512] Step 2:
[0513] The device acquires digital audio data and converts it into text data using speech recognition software. The digital audio data acquired in step 1 is used as input, and a text-formatted string is generated as output. This conversion includes keyword extraction from the audio and analysis of phonological information.
[0514] Step 3:
[0515] The device inputs the generated text data into a generating AI model to analyze the likelihood of fraud. The text data converted in step 2 is used as input, and a fraud risk score is provided as output. At this stage, natural language processing techniques are used to perform contextual analysis and pattern detection of the text.
[0516] Step 4:
[0517] The device extracts emotional data from the user's voice using an emotion analysis engine. The digital voice data from step 1 is used again as input, and emotional parameters (e.g., stress, tension) are generated as output. This process analyzes the tone, speed, and intonation of the voice.
[0518] Step 5:
[0519] The device integrates emotional data into a generative AI model to perform a comprehensive assessment of fraud risk. The risk score from Step 3 and the emotional parameters from Step 4 are integrated as input, and the final risk assessment result is output. This integration enables a complex analysis that takes into account the impact of emotions on the likelihood of fraud.
[0520] Step 6:
[0521] Based on the final risk assessment results, the device will issue a warning to the user if it determines that there is a high probability of fraud. The risk assessment results generated in step 5 are used as input, and a notification (email, SMS, push notification) is sent to the user as output. Specific actions here include creating and sending the notification message.
[0522] (Application Example 2)
[0523] Next, we will explain Application Example 2. In the following explanation, the data processing device 12 will be referred to as the "server," and the headset-type terminal 314 will be referred to as the "terminal."
[0524] In today's world, fraud via communications is on the rise, and users are required to protect themselves from fraudulent activities. However, conventional fraud detection systems rely on analyzing voice content and ignore emotional responses, which poses a challenge in terms of accuracy. Furthermore, it is difficult to issue warnings before users recognize actual signs of fraud, making it difficult to quickly reduce fraud risk.
[0525] The specific processing performed by the specific processing unit 290 of the data processing device 12 in Application Example 2 is realized by the following means.
[0526] In this invention, the server includes means for acquiring the user's voice, means for converting the voice into text data, means for using generative artificial intelligence to analyze the likelihood of fraud in real time based on the text data, means for sentiment analysis to evaluate the emotional state from the user's voice, means for integrating the emotional state data into the fraud analysis process, means for issuing a warning when the likelihood of fraud is detected, and means for sending an alert to a set contact via any communication means. This enables highly accurate, real-time fraud detection that takes into account the user's emotional response.
[0527] "Means for acquiring user voice" refers to a function that uses an input device to convert the voice signal emitted by the user into digital data and input it into the system.
[0528] "Means for converting speech to text data" refers to a function that processes speech signals acquired using speech recognition technology into corresponding strings of characters.
[0529] "Methods using generative artificial intelligence" refer to software that uses machine learning algorithms to analyze input data and mimic human judgment to assess the likelihood of fraud.
[0530] "An emotion analysis method for evaluating emotional state from user voice" is an algorithm that extracts the user's emotional characteristics from voice data and evaluates their psychological state.
[0531] "Means of integrating emotional state data into the fraud analysis process" refers to the process of incorporating parameters obtained through emotional analysis into a fraud detection algorithm and using them as part of the analysis.
[0532] "A means of issuing a warning when a potential scam is detected" refers to a function that notifies the user to be cautious when it is determined that there is a high probability of fraud.
[0533] "Means of sending alerts to configured contacts" refers to the process of sending warnings as emails or messages based on pre-registered contact information.
[0534] The system for carrying out the present invention mainly consists of a voice acquisition device, voice recognition software, emotion analysis engine, generative AI model, and warning notification system.
[0535] The server captures the user's voice through a microphone and converts the voice data into text in real time using the Google Speech-to-Text API. This text data is then analyzed using a generative AI model such as the BERT model to assess the likelihood of fraud. During this process, an emotion analysis engine such as openSMILE is used to evaluate the user's emotional state based on factors such as tone, rhythm, and word choice. The resulting emotion parameters are then integrated into the fraud detection algorithm to improve the accuracy of the analysis.
[0536] When a user receives a phone call, for example, if the caller asks for verification of personal information, the user might say, "Is that really true?" At this point, the system detects the user's emotional state, indicating anxiety or suspicion, from their voice and uses that data to assess the risk of fraud.
[0537] Ultimately, if the device is determined to be highly likely to be a scam, it will alert the user via push notification. It can also use communication technologies such as Firebase Cloud Messaging to send alerts via email or message to pre-configured contacts.
[0538] A concrete example of a prompt message from a generated AI model might be text like, "Please evaluate the likelihood that the following message is a scam. User: 'I'm being asked to confirm my personal information, but is this real?'" This allows users to communicate with confidence through everyday communications thanks to highly accurate scam detection.
[0539] The flow of a specific process in Application Example 2 will be explained using Figure 14.
[0540] Step 1:
[0541] The terminal uses a microphone to acquire the user's voice. The input is the user's voice signal, which is captured as digital data. This acquired voice data is then passed on to the next processing step.
[0542] Step 2:
[0543] The server uses the Google Speech-to-Text API to convert the acquired audio data into text data. The input is audio data, and the output is text data generated based on this audio. This conversion process makes the audio information suitable for analysis.
[0544] Step 3:
[0545] The server analyzes text data using a generative AI model. The input is previously converted text data, and the output is a score indicating the likelihood of fraud. This model evaluates the content of the input text based on fraud patterns that it has learned in advance.
[0546] Step 4:
[0547] The device uses an emotion analysis engine such as openSMILE to analyze the user's emotional state from voice data. The input is the initial voice data, and the output is emotion parameters based on voice tone and patterns. This analysis quantifies the user's stress level and anxiety.
[0548] Step 5:
[0549] The server integrates emotional parameters into the fraud analysis process. The emotional parameters are added to the fraud analysis score, and the final fraud risk is recalculated. This results in a more accurate fraud risk assessment that takes the user's psychological state into account.
[0550] Step 6:
[0551] The device issues a warning if the final score it generates exceeds a certain threshold. The input is the final calculated fraud risk score, and the output is a warning message. The device warns the user via push notification and, if necessary, sends an alert to the designated contacts.
[0552] Step 7:
[0553] Upon receiving a warning, the user can confirm whether to end the call. The input is the warning message from the device, and the output is the user's appropriate response. This makes it possible to prevent fraud.
[0554] The specific processing unit 290 transmits the result of the specific processing to the headset terminal 314. In the headset terminal 314, the control unit 46A causes the speaker 240 and display 343 to output the result of the specific processing. The microphone 238 acquires audio indicating user input for the result of the specific processing. The control unit 46A transmits the audio data indicating user input acquired by the microphone 238 to the data processing unit 12. In the data processing unit 12, the specific processing unit 290 acquires the audio data.
[0555] Data generation model 58 is a type of so-called generative AI (Artificial Intelligence). One example of data generation model 58 is ChatGPT (Internet search<URL: https: / / openai.com / blog / chatgpt> ), Gemini (Internet search) <url: https: gemini.google.com ?hl="ja">Examples of generative AI include the following. The data generation model 58 is obtained by performing deep learning on a neural network. The data generation model 58 is input with prompts containing instructions, and with inference data such as audio data representing speech, text data representing text, and image data representing images. The data generation model 58 infers from the input inference data according to the instructions indicated by the prompts, and outputs the inference results in data formats such as audio data and text data. Here, inference refers to, for example, analysis, classification, prediction, and / or summarization.
[0556] In the above embodiment, an example was given in which specific processing is performed by the data processing device 12, but the technology of this disclosure is not limited thereto, and specific processing may also be performed by the headset terminal 314.
[0557] [Fourth Embodiment]
[0558] Figure 7 shows an example of the configuration of the data processing system 410 according to the fourth embodiment.
[0559] As shown in Figure 7, the data processing system 410 includes a data processing device 12 and a robot 414. An example of the data processing device 12 is a server.
[0560] The data processing device 12 comprises a computer 22, a database 24, and a communication interface 26. The computer 22 is an example of a "computer" related to the technology of this disclosure. The computer 22 comprises a processor 28, RAM 30, and storage 32. The processor 28, RAM 30, and storage 32 are connected to a bus 34. The database 24 and the communication interface 26 are also connected to the bus 34. The communication interface 26 is connected to a network 54. An example of the network 54 is a WAN (Wide Area Network) and / or a LAN (Local Area Network).
[0561] The robot 414 includes a computer 36, a microphone 238, a speaker 240, a camera 42, a communication interface 44, and a controlled object 443. The computer 36 includes a processor 46, RAM 48, and storage 50. The processor 46, RAM 48, and storage 50 are connected to a bus 52. The microphone 238, speaker 240, camera 42, and controlled object 443 are also connected to the bus 52.
[0562] The microphone 238 receives voice signals from the user 20 and receives instructions from the user 20. The microphone 238 captures the voice signals from the user 20, converts the captured voice into audio data, and outputs it to the processor 46. The speaker 240 outputs audio according to the instructions from the processor 46.
[0563] Camera 42 is a small digital camera equipped with an optical system including a lens, aperture, and shutter, and an image sensor such as a CMOS (Complementary Metal-Oxide-Semiconductor) image sensor or a CCD (Charge Coupled Device) image sensor, and captures images of the area around the user 20 (for example, an imaging range defined by a field of view equivalent to the width of a typical healthy person's field of vision).
[0564] Communication interface 44 is connected to network 54. Communication interfaces 44 and 26 are responsible for the exchange of various information between processor 46 and processor 28 via network 54. The exchange of various information between processor 46 and processor 28 using communication interfaces 44 and 26 is performed in a secure manner.
[0565] The controlled object 443 includes a display device, LEDs in the eyes, and motors that drive the arms, hands, and feet. The posture and gestures of the robot 414 are controlled by controlling the motors of the arms, hands, and feet. Some of the robot 414's emotions can be expressed by controlling these motors. Furthermore, the robot 414's facial expressions can also be expressed by controlling the illumination state of the LEDs in its eyes.
[0566] Figure 8 shows an example of the main functions of the data processing device 12 and the robot 414. As shown in Figure 8, the data processing device 12 performs specific processing using the processor 28. The storage 32 stores the specific processing program 56.
[0567] The specific processing program 56 is an example of a "program" relating to the technology of this disclosure. The processor 28 reads the specific processing program 56 from the storage 32 and executes the read specific processing program 56 on the RAM 30. The specific processing is realized by the processor 28 operating as a specific processing unit 290 in accordance with the specific processing program 56 executed on the RAM 30.
[0568] The storage 32 stores the data generation model 58 and the emotion identification model 59. The data generation model 58 and the emotion identification model 59 are used by the identification processing unit 290.
[0569] In robot 414, the processor 46 performs the reception output processing. The storage 50 stores the reception output program 60. The processor 46 reads the reception output program 60 from the storage 50 and executes the read reception output program 60 on the RAM 48. The reception output processing is realized by the processor 46 operating as a control unit 46A according to the reception output program 60 executed on the RAM 48.
[0570] Next, the specific processing performed by the specific processing unit 290 of the data processing device 12 will be described. In the following description, the data processing device 12 will be referred to as the "server" and the robot 414 as the "terminal".
[0571] This invention is a system designed to reduce the risk of fraudulent billing scams and scams impersonating family members that lurk in users' daily lives. This system primarily acquires the user's voice, converts that voice into text data, and analyzes it to detect potential fraud and issue appropriate warnings.
[0572] First, the device continuously monitors the user's conversation and surrounding sounds in the background to acquire audio data in real time. This audio data is converted into text data using the device's built-in speech recognition function, preparing it for analysis.
[0573] Once the audio is converted into text data, the device uses generative AI to analyze the content of the text data. This generative AI model is trained on a dataset containing characteristic patterns of fraud, and it determines the likelihood of fraudulent activity based on the analysis results.
[0574] If a device is determined to be potentially fraudulent, it will immediately send an alert to the designated contacts. These contacts primarily include family and close friends, but depending on usage, it can also notify law enforcement authorities. Users can choose from a variety of alert delivery methods, including email, SMS, and push notifications.
[0575] For example, if a user receives a phone call stating, "You have outstanding charges," the device immediately analyzes the phrase and detects keywords such as "outstanding" and "charges." The generating AI then determines that there is a high risk of fraud and quickly issues a warning to the user and their registered contacts. This allows the user to take appropriate action without complying with fraudulent demands.
[0576] Furthermore, users can freely customize how the system notifies them of warnings, and if the response is insufficient, they can manually use contact information to provide follow-up support. This flexibility allows the system to be adapted to the user's actual usage and lifestyle.
[0577] This will help prevent fraud and support users in living their daily lives with peace of mind.
[0578] The following describes the processing flow.
[0579] Step 1:
[0580] The device will always have a function enabled to receive the user's voice through the microphone. Voice acquisition will be initiated by a configured trigger event (for example, detection that a conversation has started).
[0581] Step 2:
[0582] The device converts the acquired audio into text data using a speech recognition system. The speech recognition system transcribes the audio in real time using a pre-installed API.
[0583] Step 3:
[0584] The device passes the converted text data to a generating AI model, which analyzes whether the text contains phrases or patterns that indicate fraud. The generating AI is a model that has learned from past fraud cases and has the ability to detect suspicious patterns.
[0585] Step 4:
[0586] If the device receives the results of its analysis of the generated AI model and determines that it may be fraudulent, it prepares to take action to issue a warning according to pre-configured rules.
[0587] Step 5:
[0588] The device will send alerts. These notifications are sent as push notifications to the user's device, as well as via email and SMS to pre-configured contacts (family, police, etc.). The method of sending notifications is pre-configured by the user.
[0589] Step 6:
[0590] Users can receive alerts and follow the instructions displayed on their device to review audio logs and analysis results of suspicious conversations. If necessary, they can also report or request further confirmation from administrators via their device.
[0591] (Example 1)
[0592] Next, we will describe Example 1. In the following description, the data processing device 12 will be referred to as the "server" and the robot 414 as the "terminal".
[0593] The present invention aims to effectively detect the risk of users being subjected to fraudulent activities or unethical solicitations, and to prevent them from becoming victims. Furthermore, it aims to surpass the limitations of conventional technologies by utilizing user feedback to improve detection accuracy.
[0594] The identification process performed by the identification processing unit 290 of the data processing device 12 in Example 1 is realized by the following means.
[0595] In this invention, the server includes means for converting the user's voice into text data, means for using generative artificial intelligence to analyze the text data and detect the possibility of fraudulent activity, and means for notifying the user of a warning based on the detection results. This enables the user to immediately recognize fraudulent activity and take prompt and appropriate action.
[0596] "User" refers to an individual or organization that uses this system and provides voice data for the purpose of detecting and preventing fraudulent activity.
[0597] A "device for acquiring voice" refers to hardware or software that records the user's voice in real time and converts that voice data into a format usable within the system.
[0598] A "device for converting to text data" refers to hardware or software that uses speech recognition technology to convert acquired audio data into analyzable text data.
[0599] "Generative artificial intelligence" refers to artificial intelligence technology equipped with an algorithm that analyzes text data converted from speech and determines the likelihood of fraudulent activity.
[0600] "Warning device" refers to a means or device for notifying a user or relevant contact if potential fraudulent activity is detected.
[0601] A "device for sending notifications" refers to a means of sending alerts to designated contacts using a communication network, and includes a variety of methods such as email, SMS, and push notifications.
[0602] "Devices that improve generative artificial intelligence models by providing feedback" refers to hardware and software that manage the process of improving the system's detection performance through responses and opinions from users.
[0603] This system is designed to help users proactively detect fraudulent activity and avoid its consequences. Specifically, it utilizes speech recognition and generation AI technology to analyze voice data in real time. Details are provided below.
[0604] The device constantly monitors the user's conversations and surrounding sounds in the background. For this purpose, the device uses its built-in microphone and a voice acquisition module (e.g., a common voice recording device or software). The acquired voice data is converted into text data using speech recognition software (e.g., Google Speech-to-Text or equivalent technology).
[0605] After being converted into text data, the terminal analyzes the content using a generative AI model (e.g., OpenAI's GPT model). In this process, the generative AI model uses prompt messages to assess the likelihood of fraud in real time. Examples of specific prompt messages include "unpaid" and "urgent action required."
[0606] If the generating AI detects a potential scam, the device will immediately issue a warning. The warning will be sent to the designated contacts, but will depend on the user's custom notification method (email, SMS, push notification, etc.). This allows users and stakeholders to immediately understand the situation and take appropriate action.
[0607] Users can freely customize how they receive warning notifications from the system. Furthermore, the device collects user feedback information, which is used to improve the accuracy of the generated AI model. This allows the system to better suit the needs of individual users and enhance the effectiveness of fraud detection. With this flexibility, the system supports users in leading a safer and more secure daily life.
[0608] The flow of the specific processing in Example 1 will be explained using Figure 11.
[0609] Step 1:
[0610] The device acquires the user's voice in real time in the background. This process continuously captures ambient audio signals using the device's built-in microphone. The input is audio signals from the surrounding environment, and the output is digitized audio data. This audio data is divided into short samples for subsequent processing.
[0611] Step 2:
[0612] The device converts the acquired audio data into text data. Here, a speech recognition library (e.g., Google Speech-to-Text) is used to convert the audio signal into a string. The input is digital audio data, and data conversion is performed, including background noise removal and speech recognition. The output is formatted text data, which is used for subsequent analysis.
[0613] Step 3:
[0614] The device analyzes text data using a generating AI model to search for patterns of fraudulent activity. In this step, the AI model (e.g., a GPT model) uses prompt text to assess the likelihood of fraud. The input consists of text data and prompt text, and the AI model performs data exploration and pattern matching. The output is provided as a fraud risk score.
[0615] Step 4:
[0616] If fraud is highly likely, the device generates an alert and sends a notification to the user and their designated contacts. The notification method is pre-specified by the user and can be email, SMS, or push notification. The input is the result of an AI risk assessment, which automatically selects the notification format and issues an alert. The output is the notification message sent to the recipient.
[0617] Step 5:
[0618] The user receives a warning message and takes action. This feedback is collected by the system and used to improve the generated AI model. Depending on the user's actions, notification settings can be adjusted and the AI model can be retrained. The input is the user's response and additional information, and the output is an improvement in the system's analysis accuracy.
[0619] (Application Example 1)
[0620] Next, we will explain Application Example 1. In the following explanation, the data processing device 12 will be referred to as the "server" and the robot 414 as the "terminal".
[0621] In modern society, fraud involving fictitious billing and impersonation of family members is on the rise. In particular, many users, including the elderly, become victims without realizing the risks. In this situation, it is crucial for users to quickly and effectively detect and prevent these fraud risks in their daily lives.
[0622] The specific processing performed by the specific processing unit 290 of the data processing device 12 in Application Example 1 is realized by the following means.
[0623] In this invention, the server includes a device for acquiring voice information, a device for converting it into text data, and a device using generative artificial intelligence that instantly analyzes the risk of fraud. This allows users to detect and be warned in real time about the risk of fraud they may encounter in their daily lives, thereby preventing them from becoming victims of fraud.
[0624] "Audio information" refers to sound data collected from the user and their surroundings, including human conversation and ambient sounds.
[0625] "Text data" refers to data in text format generated based on audio information, and it visually represents the content of speech.
[0626] "Generative artificial intelligence" is an advanced learning model designed to analyze voice information and assess and detect the risk of fraud.
[0627] A "warning device" refers to a device or method used to inform users or designated contacts of the risk of fraud.
[0628] "Communication methods" refer to the infrastructure and technologies used to transmit information, including the internet, telephone lines, and wireless communication.
[0629] "Specified contact information" refers to the contact details of the person or organization that the user has set up in advance to receive warning notifications.
[0630] A "device that continuously monitors ambient sounds" is a device designed to constantly collect sound information and analyze it as needed.
[0631] Embodiments of this invention primarily relate to a system for acquiring audio information and immediately analyzing the risk of fraud. A server or terminal acquires audio information and converts it into text data. Specifically, it monitors surrounding conversations in real time using the microphone of a smartphone or other device and converts the audio data into text using speech recognition technology (e.g., Google Speech-to-Text API).
[0632] The converted text data is analyzed by a generative artificial intelligence model on the server or terminal. This generative AI model (e.g., OpenAI's GPT model) is trained to identify characteristic patterns associated with fraud and assesses the risk in the text data. If the analysis determines that the risk of fraud is high, a warning system is activated, sending a warning to the user or pre-designated contacts. The warning is sent via communication methods such as push notifications, email, or SMS.
[0633] For example, if a user hears a potentially fraudulent phrase such as "You have outstanding charges," the device will detect it and warn them with "Potential scam: Please check how to proceed." This allows users to confidently identify fraudulent requests and prevent themselves from becoming victims of fraud.
[0634] Examples of input prompts for a generative AI model include: "Identify potentially fraudulent phrases in the following conversation: 'You have outstanding charges.'" This allows the system to effectively detect fraud risks and protect users.
[0635] The flow of a specific process in Application Example 1 will be explained using Figure 12.
[0636] Step 1:
[0637] The server or terminal acquires voice information. It collects user and surrounding voices in real time using the device's microphone and inputs them as digital voice data. Voice information acquisition can occur continuously or based on specific triggers.
[0638] Step 2:
[0639] The server or terminal converts the acquired audio information into text data. Using speech recognition technology (e.g., Google Speech-to-Text API), it receives the audio data as input and generates the corresponding text data. This text data is output as the underlying text information for analysis.
[0640] Step 3:
[0641] The server analyzes text data using a generative AI model. The input text data is provided to the AI in the form of a prompt message, "Identify potentially fraudulent phrases in the following conversation," and the risk of fraud is assessed. The generative AI model (e.g., OpenAI's GPT model) identifies characteristic patterns of fraud and outputs a risk score.
[0642] Step 4:
[0643] The server or terminal analyzes the risk score returned by the generated AI model to determine the likelihood of fraud. If the risk score exceeds a certain threshold, it prepares to issue a warning. The output here is the result of the fraud likelihood assessment.
[0644] Step 5:
[0645] The server or terminal sends a warning to the user and their designated contact if there is a possibility of fraud. The warning is sent via a communication method such as push notification, email, or SMS, using the contact information entered as the recipient. The output is a record that the warning was successfully delivered.
[0646] Furthermore, an emotion engine that estimates the user's emotions may be incorporated. That is, the identification processing unit 290 may use the emotion identification model 59 to estimate the user's emotions and perform identification processing using the user's emotions.
[0647] This invention combines a system that analyzes the potential for fraud in real time from a user's voice with an emotion engine that analyzes the user's emotional state, thereby achieving more accurate fraud detection and warnings. Because this invention considers not only the content of the voice but also the user's emotional response, the risk of fraud is assessed more accurately.
[0648] First, the device continuously acquires the user's voice and converts it into text data using speech recognition technology. This text data is then passed to an artificial intelligence model generated for fraud detection.
[0649] Simultaneously, the device is equipped with an emotion engine that analyzes emotional responses from the user's voice tone, speaking patterns, and word choice. The emotion engine can evaluate emotional parameters such as the user's stress level and tension.
[0650] The emotional data obtained by the emotion engine is integrated into the analysis process of the generative AI model. For example, if a user says something like, "Am I really behind on payments?" on the phone, with a nuance of doubt, the device detects this unstable emotional state. This increases the risk score for fraud, and the device determines that it needs to issue an immediate warning.
[0651] Next, if a call is deemed highly likely to be a scam, the device will immediately issue a warning to the user. Upon receiving this warning, the user can quickly take countermeasures such as hanging up the potentially dangerous call or conducting further investigation. The warning includes sending alerts to designated contacts (e.g., family members or the police). The warning method can be pre-configured by the user and can be selected from options such as email, SMS, or push notification.
[0652] In this way, the present invention combines emotion recognition technology with conventional speech recognition-based fraud detection systems, making it possible to identify fraud risks early and with high accuracy, and to provide appropriate warnings. As a result, users can engage in everyday communication with peace of mind, even in situations where emotional responses may indicate the possibility of fraud.
[0653] The following describes the processing flow.
[0654] Step 1:
[0655] The device receives the user's voice through the microphone and begins recording it as audio data. In this initial stage, voice acquisition is triggered when the user starts speaking.
[0656] Step 2:
[0657] The device transmits the acquired voice data to a speech recognition system, which converts it into text data in real time. The converted text data is then prepared for analysis to detect fraud.
[0658] Step 3:
[0659] The device simultaneously uses an emotion engine to perform emotion analysis based on the user's voice. The emotion engine uses factors such as voice tone, speed, and volume to evaluate the degree of stress and tension.
[0660] Step 4:
[0661] The device inputs text data into a generating artificial intelligence model, which analyzes whether it contains patterns that indicate potential fraud. In addition, the generating AI takes into account the results of sentiment analysis to generate an overall fraud risk score.
[0662] Step 5:
[0663] If the fraud risk score exceeds a certain threshold, the device will determine it to be a high-priority situation and immediately prepare a warning.
[0664] Step 6:
[0665] The device issues a warning to the user and sends the alert to the designated contacts using a method selected depending on the situation (push notification, voice alert, email, SMS, etc.). If this information indicates particularly high levels of tension or stress through sentiment analysis, the user will receive a stronger warning.
[0666] Step 7:
[0667] Users can receive warning notifications from their devices and take appropriate measures, such as hanging up the phone early, by following the instructions. Users can also access information through their devices to take further action.
[0668] (Example 2)
[0669] Next, we will describe Example 2. In the following description, the data processing device 12 will be referred to as the "server" and the robot 414 as the "terminal".
[0670] There is a need to provide a system that can detect fraud with high accuracy and issue timely warnings, taking into account not only the user's voice but also their emotional responses, thereby more accurately grasping fraud risks that are often overlooked by existing systems.
[0671] The identification process performed by the identification processing unit 290 of the data processing device 12 in Example 2 is realized by the following means.
[0672] In this invention, the server includes a device for acquiring user voice, a mechanism for converting the voice into data, a configuration using a generative machine learning model for analyzing the likelihood of fraud based on the data, and means for incorporating an emotion analysis engine when evaluating the likelihood of fraud. This makes it possible to comprehensively analyze voice content and emotion data, evaluate fraud risk with high accuracy, and issue timely notifications to the user.
[0673] A "user" is an entity that utilizes the system to provide its own voice data and participate in fraud detection.
[0674] A "device for acquiring voice" refers to hardware or software used to collect a user's voice and convert it into a digital format.
[0675] A "data conversion mechanism" refers to the technology responsible for the process of converting collected audio information into text data or analyzable digital data.
[0676] A "configuration using a generative machine learning model" refers to a system built using an algorithm that analyzes user data and learns specific patterns in order to detect fraudulent activity.
[0677] An "emotion analysis engine" refers to a technology that extracts emotional responses from a user's voice and calculates their parameters.
[0678] The "notification mechanism" refers to a part of the system that sends a warning to the user or designated contacts when potential fraud is detected.
[0679] "Communication channels" refer to electronic means of communication used to transmit warnings and notifications, and include email, SMS, push notifications, etc.
[0680] A "generative AI model" refers to an artificial intelligence model designed to analyze information from data provided as arguments and generate a specific output.
[0681] The present invention is implemented by a system that determines in real time whether or not a voice provided by a user is fraudulent. The terminal includes a microphone or other voice acquisition device for acquiring the user's voice. The acquired voice is converted into text using natural language processing techniques. For example, speech recognition software can be used for this process.
[0682] After converting the audio into data, the device supplies the data to a generative machine learning model. This generative machine learning model is pre-trained to analyze the data and recognize signs of fraud. Specifically, it uses pattern recognition and classification techniques to assess the likelihood of fraud. For example, a machine learning platform can be used for this model.
[0683] Furthermore, the system analyzes the user's emotional responses from voice data through an emotion analysis engine. This engine takes into account factors such as voice tone, pitch, and speed to evaluate emotions like stress and tension. This engine can utilize, for example, emotion analysis software. The emotional data is integrated into a generative machine learning model and used for a comprehensive assessment of fraud risk.
[0684] If a potential scam is detected, the device immediately alerts the user. This alert can be configured to reach the user and their designated contacts via multiple communication channels, including email, SMS, and push notifications. For example, if a user expresses doubt, such as "Am I really behind on payments?", the change in their tone of voice can be detected as high risk by the sentiment analysis engine, and a generative machine learning model will immediately issue a warning.
[0685] The following are specific examples of prompt statements:
[0686] User's voice text: "Am I really behind on payments?" Emotion parameters: Anxiety, tension. Conduct a risk assessment and determine the possibility of fraud.
[0687] With a system configured in this way, users can accurately assess the risk of fraud based on voice content and emotional responses, and take appropriate action.
[0688] The flow of the specific processing in Example 2 will be explained using Figure 13.
[0689] Step 1:
[0690] The device acquires the user's voice through a microphone and stores it as digital audio data. The user's voice is received in real time as input, and a digital audio file is generated as output. This file is then ready for use in subsequent processing.
[0691] Step 2:
[0692] The device acquires digital audio data and converts it into text data using speech recognition software. The digital audio data acquired in step 1 is used as input, and a text-formatted string is generated as output. This conversion includes keyword extraction from the audio and analysis of phonological information.
[0693] Step 3:
[0694] The device inputs the generated text data into a generating AI model to analyze the likelihood of fraud. The text data converted in step 2 is used as input, and a fraud risk score is provided as output. At this stage, natural language processing techniques are used to perform contextual analysis and pattern detection of the text.
[0695] Step 4:
[0696] The device extracts emotional data from the user's voice using an emotion analysis engine. The digital voice data from step 1 is used again as input, and emotional parameters (e.g., stress, tension) are generated as output. This process analyzes the tone, speed, and intonation of the voice.
[0697] Step 5:
[0698] The device integrates emotional data into a generative AI model to perform a comprehensive assessment of fraud risk. The risk score from Step 3 and the emotional parameters from Step 4 are integrated as input, and the final risk assessment result is output. This integration enables a complex analysis that takes into account the impact of emotions on the likelihood of fraud.
[0699] Step 6:
[0700] Based on the final risk assessment results, the device will issue a warning to the user if it determines that there is a high probability of fraud. The risk assessment results generated in step 5 are used as input, and a notification (email, SMS, push notification) is sent to the user as output. Specific actions here include creating and sending the notification message.
[0701] (Application Example 2)
[0702] Next, we will explain application example 2. In the following explanation, the data processing device 12 will be referred to as the "server" and the robot 414 as the "terminal".
[0703] In today's world, fraud via communications is on the rise, and users are required to protect themselves from fraudulent activities. However, conventional fraud detection systems rely on analyzing voice content and ignore emotional responses, which poses a challenge in terms of accuracy. Furthermore, it is difficult to issue warnings before users recognize actual signs of fraud, making it difficult to quickly reduce fraud risk.
[0704] The specific processing performed by the specific processing unit 290 of the data processing device 12 in Application Example 2 is realized by the following means.
[0705] In this invention, the server includes means for acquiring the user's voice, means for converting the voice into text data, means for using generative artificial intelligence to analyze the likelihood of fraud in real time based on the text data, means for sentiment analysis to evaluate the emotional state from the user's voice, means for integrating the emotional state data into the fraud analysis process, means for issuing a warning when the likelihood of fraud is detected, and means for sending an alert to a set contact via any communication means. This enables highly accurate, real-time fraud detection that takes into account the user's emotional response.
[0706] "Means for acquiring user voice" refers to a function that uses an input device to convert the voice signal emitted by the user into digital data and input it into the system.
[0707] "Means for converting speech to text data" refers to a function that processes speech signals acquired using speech recognition technology into corresponding strings of characters.
[0708] "Methods using generative artificial intelligence" refer to software that uses machine learning algorithms to analyze input data and mimic human judgment to assess the likelihood of fraud.
[0709] "An emotion analysis method for evaluating emotional state from user voice" is an algorithm that extracts the user's emotional characteristics from voice data and evaluates their psychological state.
[0710] "Means of integrating emotional state data into the fraud analysis process" refers to the process of incorporating parameters obtained through emotional analysis into a fraud detection algorithm and using them as part of the analysis.
[0711] "A means of issuing a warning when a potential scam is detected" refers to a function that notifies the user to be cautious when it is determined that there is a high probability of fraud.
[0712] "Means of sending alerts to configured contacts" refers to the process of sending warnings as emails or messages based on pre-registered contact information.
[0713] The system for carrying out the present invention mainly consists of a voice acquisition device, voice recognition software, emotion analysis engine, generative AI model, and warning notification system.
[0714] The server captures the user's voice through a microphone and converts the voice data into text in real time using the Google Speech-to-Text API. This text data is then analyzed using a generative AI model such as the BERT model to assess the likelihood of fraud. During this process, an emotion analysis engine such as openSMILE is used to evaluate the user's emotional state based on factors such as tone, rhythm, and word choice. The resulting emotion parameters are then integrated into the fraud detection algorithm to improve the accuracy of the analysis.
[0715] When a user receives a phone call, for example, if the caller asks for verification of personal information, the user might say, "Is that really true?" At this point, the system detects the user's emotional state, indicating anxiety or suspicion, from their voice and uses that data to assess the risk of fraud.
[0716] Ultimately, if the device is determined to be highly likely to be a scam, it will alert the user via push notification. It can also use communication technologies such as Firebase Cloud Messaging to send alerts via email or message to pre-configured contacts.
[0717] A concrete example of a prompt message from a generated AI model might be text like, "Please evaluate the likelihood that the following message is a scam. User: 'I'm being asked to confirm my personal information, but is this real?'" This allows users to communicate with confidence through everyday communications thanks to highly accurate scam detection.
[0718] The flow of a specific process in Application Example 2 will be explained using Figure 14.
[0719] Step 1:
[0720] The terminal uses a microphone to acquire the user's voice. The input is the user's voice signal, which is captured as digital data. This acquired voice data is then passed on to the next processing step.
[0721] Step 2:
[0722] The server uses the Google Speech-to-Text API to convert the acquired audio data into text data. The input is audio data, and the output is text data generated based on this audio. This conversion process makes the audio information suitable for analysis.
[0723] Step 3:
[0724] The server analyzes text data using a generative AI model. The input is previously converted text data, and the output is a score indicating the likelihood of fraud. This model evaluates the content of the input text based on fraud patterns that it has learned in advance.
[0725] Step 4:
[0726] The device uses an emotion analysis engine such as openSMILE to analyze the user's emotional state from voice data. The input is the initial voice data, and the output is emotion parameters based on voice tone and patterns. This analysis quantifies the user's stress level and anxiety.
[0727] Step 5:
[0728] The server integrates emotional parameters into the fraud analysis process. The emotional parameters are added to the fraud analysis score, and the final fraud risk is recalculated. This results in a more accurate fraud risk assessment that takes the user's psychological state into account.
[0729] Step 6:
[0730] The device issues a warning if the final score it generates exceeds a certain threshold. The input is the final calculated fraud risk score, and the output is a warning message. The device warns the user via push notification and, if necessary, sends an alert to the designated contacts.
[0731] Step 7:
[0732] Upon receiving a warning, the user can confirm whether to end the call. The input is the warning message from the device, and the output is the user's appropriate response. This makes it possible to prevent fraud.
[0733] The specific processing unit 290 transmits the result of the specific processing to the robot 414. In the robot 414, the control unit 46A causes the speaker 240 and the controlled object 443 to output the result of the specific processing. The microphone 238 acquires audio indicating user input for the result of the specific processing. The control unit 46A transmits the audio data indicating user input acquired by the microphone 238 to the data processing unit 12. In the data processing unit 12, the specific processing unit 290 acquires the audio data.
[0734] Data generation model 58 is a type of so-called generative AI (Artificial Intelligence). One example of data generation model 58 is ChatGPT (Internet search<URL: https: / / openai.com / blog / chatgpt> ), Gemini (Internet search) <url: https: gemini.google.com ?hl="ja">Examples of generative AI include the following. The data generation model 58 is obtained by performing deep learning on a neural network. The data generation model 58 is input with prompts containing instructions, and with inference data such as audio data representing speech, text data representing text, and image data representing images. The data generation model 58 infers from the input inference data according to the instructions indicated by the prompts, and outputs the inference results in data formats such as audio data and text data. Here, inference refers to, for example, analysis, classification, prediction, and / or summarization.
[0735] In the above embodiment, an example was given in which the specific processing is performed by the data processing device 12, but the technology of this disclosure is not limited thereto, and the specific processing may also be performed by the robot 414.
[0736] Furthermore, the emotion identification model 59, acting as an emotion engine, may determine the user's emotion according to a specific mapping. Specifically, the emotion identification model 59 may determine the user's emotion according to a specific mapping, which is an emotion map (see Figure 9). Similarly, the emotion identification model 59 may also determine the robot's emotion, and the identification processing unit 290 may perform identification processing using the robot's emotion.
[0737] Figure 9 shows an emotion map 400 in which multiple emotions are mapped. In the emotion map 400, emotions are arranged in concentric circles radiating from the center. The closer to the center of the concentric circles, the more primitive the emotions are located. Further out of the concentric circles, emotions representing states and actions arising from mental states are located. Emotion is a concept that includes feelings and mental states. On the left side of the concentric circles, emotions that are generally generated from reactions occurring in the brain are located. On the right side of the concentric circles, emotions that are generally induced by situational judgment are located. Above and below the concentric circles, emotions that are generally generated from reactions occurring in the brain and induced by situational judgment are located. In addition, the emotion of "pleasure" is located on the upper side of the concentric circles, and the emotion of "displeasure" is located on the lower side. Thus, in the emotion map 400, multiple emotions are mapped based on the structure in which emotions arise, and emotions that are likely to occur simultaneously are mapped close together.
[0738] These emotions are distributed at the 3 o'clock position on the Emotion Map 400, and usually fluctuate between feelings of security and anxiety. In the right half of the Emotion Map 400, situational awareness takes precedence over internal feelings, resulting in a calm impression.
[0739] The inside of the Emotion Map 400 represents inner thoughts, while the outside represents actions. Therefore, the further you go from the outside of the Emotion Map 400, the more visible (expressed in actions) your emotions become.
[0740] Here, human emotions are based on various balances, such as posture and blood sugar levels. When these balances deviate from the ideal, it results in discomfort, and when they approach the ideal, it results in pleasure. Similarly, in robots, cars, motorcycles, etc., emotions can be created based on various balances, such as posture and battery level. When these balances deviate from the ideal, it results in discomfort, and when they approach the ideal, it results in pleasure. The emotion map can be generated, for example, based on Dr. Mitsuyoshi's emotion map (Research on a system for analyzing brain physiological signals of speech emotion recognition and emotion, Tokushima University, doctoral dissertation: https: / / ci.nii.ac.jp / naid / 500000375379). The left half of the emotion map contains emotions belonging to a region called "response," where sensation is dominant. The right half of the emotion map contains emotions belonging to a region called "situation," where situational awareness is dominant.
[0741] The emotion map defines two emotions that promote learning. One is the emotion around the middle of the negative "repentance" and "reflection" on the situation side. In other words, it is when the robot experiences negative emotions such as "I never want to feel this way again" or "I don't want to be scolded again." The other is the emotion around the positive "desire" on the reaction side. In other words, it is when the robot has positive feelings such as "I want more" or "I want to know more."
[0742] The emotion identification model 59 inputs user input into a pre-trained neural network, obtains emotion values representing each emotion shown in the emotion map 400, and determines the user's emotion. This neural network is pre-trained based on multiple training data sets, which are combinations of user input and emotion values representing each emotion shown in the emotion map 400. Furthermore, this neural network is trained so that emotions located close together have similar values, as shown in the emotion map 900 in Figure 10. Figure 10 shows an example where multiple emotions such as "reassured," "calm," and "confident" have similar emotion values.
[0743] The above description primarily focuses on the functions of the data processing device 12 in relation to this disclosure. However, the system related to this disclosure is not necessarily implemented on a server. The system related to this disclosure may be implemented as a general information processing system. This disclosure may be implemented, for example, as a software program that runs on a personal computer or as an application that runs on a smartphone. The method related to this disclosure may be provided to users in SaaS (Software as a Service) format.
[0744] In the above embodiment, an example was given in which a specific process is performed by a single computer 22. However, the technology of this disclosure is not limited thereto, and a distributed processing of the specific process may be performed by multiple computers, including computer 22. For example, a data generation model 58 may be provided in an external device of the data processing device 12, and the external device may generate data according to the input data.
[0745] In the above embodiment, an example was given in which the specific processing program 56 is stored in the storage 32, but the technology of this disclosure is not limited thereto. For example, the specific processing program 56 may be stored in a portable, computer-readable, non-temporary storage medium such as a USB (Universal Serial Bus) memory. The specific processing program 56 stored in the non-temporary storage medium is installed in the computer 22 of the data processing device 12. The processor 28 executes specific processing according to the specific processing program 56.
[0746] Alternatively, the specific processing program 56 may be stored in a storage device such as a server connected to the data processing device 12 via the network 54, and the specific processing program 56 may be downloaded and installed on the computer 22 in response to a request from the data processing device 12.
[0747] Furthermore, it is not necessary to store the entirety of the specific processing program 56 in a storage device such as a server connected to the data processing device 12 via the network 54, or to store the entirety of the specific processing program 56 in the storage 32; it is acceptable to store only a portion of the specific processing program 56.
[0748] The following types of processors can be used as hardware resources to perform specific processing. Examples of processors include a CPU, a general-purpose processor that functions as a hardware resource to perform specific processing by executing software, i.e., a program. Other examples of processors include dedicated electrical circuits, such as FPGAs (Field-Programmable Gate Arrays), PLDs (Programmable Logic Devices), or ASICs (Application Specific Integrated Circuits), which have circuit configurations specifically designed to perform specific processing. All of these processors have built-in or connected memory, and all of them perform specific processing by using memory.
[0749] The hardware resource that performs a specific process may consist of one of these various processors, or it may consist of a combination of two or more processors of the same or different types (for example, a combination of multiple FPGAs, or a combination of a CPU and an FPGA). Alternatively, the hardware resource that performs a specific process may consist of a single processor.
[0750] Examples of configurations using a single processor include, firstly, a configuration in which one or more CPUs and software are combined to form a single processor, and this processor functions as a hardware resource that performs a specific process. Secondly, there is a configuration using a processor that realizes the functions of the entire system, including multiple hardware resources that perform a specific process, on a single IC chip, as exemplified by SoCs (System-on-a-chip). In this way, a specific process is realized using one or more of the above types of processors as hardware resources.
[0751] Furthermore, the hardware structure of these various processors can more specifically utilize electrical circuits that combine circuit elements such as semiconductor devices. Also, the specific processing described above is merely an example. Therefore, it goes without saying that unnecessary steps can be deleted, new steps added, or the processing order rearranged, as long as it does not deviate from the main purpose.
[0752] The descriptions and illustrations presented above are detailed explanations of the technical aspects of this disclosure and are merely examples of the technical aspects. For example, the above descriptions of the structure, function, operation, and effect are examples of the structure, function, operation, and effect of the technical aspects of this disclosure. Therefore, it goes without saying that you may delete unnecessary parts, add new elements, or replace elements in the descriptions and illustrations presented above, as long as you do not deviate from the essence of the technical aspects of this disclosure. Furthermore, in order to avoid confusion and facilitate understanding of the technical aspects of this disclosure, explanations of common technical knowledge and the like that do not require special explanation to enable the implementation of the technical aspects of this disclosure have been omitted from the descriptions and illustrations presented above.
[0753] All documents, patent applications, and technical standards described herein are incorporated by reference to the same extent as if each individual document, patent application, and technical standard were specifically and individually noted to be incorporated by reference.
[0754] The following is further disclosed regarding the embodiments described above.
[0755] (Claim 1)
[0756] Means for acquiring user voice,
[0757] A means for converting the audio into text data,
[0758] A method using generative artificial intelligence to analyze the likelihood of fraud in real time based on the text data,
[0759] A means of issuing a warning when a potential scam is detected,
[0760] A means of sending an alert to a contact set via any communication method,
[0761] A system that includes this.
[0762] (Claim 2)
[0763] The system according to claim 1, comprising a user-configurable warning notification method.
[0764] (Claim 3)
[0765] The system according to claim 1, comprising means for generating artificial intelligence to score fraud patterns and determine the likelihood of fraud.
[0766] "Example 1"
[0767] (Claim 1)
[0768] A device that acquires the user's voice,
[0769] A device for converting the audio into text data,
[0770] A device using generative artificial intelligence to analyze the possibility of fraudulent activity based on the character data,
[0771] A device that issues a warning when it detects the possibility of fraudulent activity,
[0772] A device that sends notifications to contacts configured via any communication device,
[0773] A device that improves the generated artificial intelligence model by allowing users to provide feedback,
[0774] A system that includes this.
[0775] (Claim 2)
[0776] The system according to claim 1, wherein the user can customize the notification method in advance.
[0777] (Claim 3)
[0778] The system according to claim 1, comprising a device that uses generating artificial intelligence to evaluate fraudulent patterns and determine the likelihood of fraudulent activity.
[0779] "Application Example 1"
[0780] (Claim 1)
[0781] A device for acquiring voice information,
[0782] A device for converting the audio information into text data,
[0783] A device using generative artificial intelligence that instantly analyzes the risk of fraud based on the text data,
[0784] A device that sends out a warning when it detects the risk of fraud,
[0785] A device that transmits a warning to a designated contact via any means of communication,
[0786] A device that continuously monitors ambient sounds,
[0787] A system that includes this.
[0788] (Claim 2)
[0789] The system according to claim 1, including a method for providing warning notifications that can be set in advance by the user.
[0790] (Claim 3)
[0791] The system according to claim 1, comprising a device in which generative artificial intelligence evaluates fraud patterns and determines the risk of fraud.
[0792] "Example 2 of combining an emotion engine"
[0793] (Claim 1)
[0794] A device that acquires the user's voice,
[0795] A mechanism for converting the audio into data,
[0796] A configuration using a generative machine learning model to analyze the likelihood of fraud based on the data,
[0797] A method for incorporating an emotion analysis engine when evaluating the possibility of fraud,
[0798] A system that issues a notification when it detects a potential scam,
[0799] A function to send notifications to configured contacts via any communication channel,
[0800] A system that includes this.
[0801] (Claim 2)
[0802] The system according to claim 1, comprising a notification method that can be set in advance by the user.
[0803] (Claim 3)
[0804] A method in which a generative machine learning model evaluates fraud patterns and determines the likelihood of fraud,
[0805] The system according to claim 1, further comprising a function to integrate emotional data into the generated AI model to perform a comprehensive risk assessment.
[0806] "Application example 2 when combining with an emotional engine"
[0807] (Claim 1)
[0808] Means for acquiring user voice,
[0809] A means for converting the audio into text data,
[0810] A method using generative artificial intelligence to analyze the likelihood of fraud in real time based on the text data,
[0811] A means of sentiment analysis that evaluates the emotional state from the user's voice,
[0812] A means of integrating emotional state data into the fraud analysis process,
[0813] A means of issuing a warning when a potential scam is detected,
[0814] A means of sending an alert to a contact set via any communication method,
[0815] A system that includes this.
[0816] (Claim 2)
[0817] The system according to claim 1, comprising a user-configurable warning notification method.
[0818] (Claim 3)
[0819] The system according to claim 1, comprising means for generating artificial intelligence to score fraud patterns and determine the likelihood of fraud, and means for modifying the risk score using emotional state parameters. [Explanation of symbols]
[0820] 10, 210, 310, 410 Data Processing Systems 12 Data Processing Devices 14 Smart Devices 214 Smart Glasses 314 Headset-type terminal 414 Robots< / url:> < / url:> < / url:> < / url:>
Claims
1. Means for acquiring user voice, A means for converting the audio into text data, A method using generative artificial intelligence to analyze the likelihood of fraud in real time based on the text data, A means of issuing a warning when a potential scam is detected, A means of sending an alert to a contact set via any communication method, A system that includes this.
2. The system according to claim 1, comprising a warning notification method that can be set in advance by the user.
3. The system according to claim 1, comprising means for generating artificial intelligence to score fraud patterns and determine the likelihood of fraud.