system
A security system records and analyzes telephone conversations to detect fraud patterns, converting voice to text and sending immediate warnings, effectively preventing special fraud.
Patent Information
- Application Number
- JP2024141598
- Authority / Receiving Office
- JP · JP
- Patent Type
- Applications
- Current Assignee / Owner
- Filing Date
- 2024-08-22
- Publication Date
- 2026-03-06
AI Technical Summary
Conventional measures against telephone frauds, especially targeting the elderly, are insufficient, and there is a need for a system that can monitor and analyze telephone conversations in real time to detect the possibility of special fraud and issue warnings.
A security system that records telephone voice, converts it into text data, analyzes the text for fraud patterns, and generates and transmits warning messages via push notification, SMS, or email.
The system effectively detects the possibility of special fraud in real time, preventing damage by immediately conveying warnings to users.
Smart Images

Figure 2026038263000001_ABST
Abstract
Description
[Technical Field]
[0001] The technology of the present disclosure relates to a system. [Background technology]
[0002] Patent document 1 discloses a persona chatbot control method performed by at least one processor, the method including the steps of receiving a user utterance, adding the user utterance to a prompt including an instruction sentence related to a description of the chatbot character, encoding the prompt, and inputting the encoded prompt into a language model to generate a chatbot utterance in response to the user utterance. [Prior art documents] [Patent documents]
[0003] [Patent Document 1] Japanese Patent Publication No. 2022-180282 Summary of the Invention [Problem to be solved by the invention]
[0004] In recent years, the number of victims of special frauds has been increasing, and frauds over the phone have become a serious problem, especially among the elderly. Conventional measures against network communications, malware, and viruses are not sufficient to combat such frauds. Therefore, a new type of security service that focuses on the audio of telephone conversations is needed.
[0005] The object of the present invention is to provide a system that monitors and analyzes telephone conversations in real time, detects the possibility of special fraud in advance, and issues a warning. [Means for solving the problem]
[0006] The present invention provides a security system that records telephone voice, converts the recorded voice into text data, analyzes the converted text data to determine the possibility of special fraud, and generates and transmits a warning message based on the determination. In particular, the present invention ensures data security by including a means for encrypting and transmitting the recorded voice. Furthermore, by employing a means for transmitting the warning message via push notification, SMS, email, or operator, it is possible to immediately convey a warning to the user. This makes it possible to prevent damage from special fraud.
[0007] "Telephone voice" refers to a voice signal transmitted over a telephone, including the content of a conversation.
[0008] "Recording" is the process of saving audio or data so that it can be played back or analyzed later.
[0009] "Text data" is audio converted into text information, and is in a format that can be easily analyzed and searched.
[0010] "Analysis" is the process of examining data or information in detail to extract specific patterns or features.
[0011] "Special fraud" refers to fraudulent methods used over the telephone or the Internet to steal personal information or money.
[0012] "Judgment" is the act of drawing a particular conclusion or result based on given information.
[0013] A "warning message" is a message that notifies and calls attention to a particular risk or danger.
[0014] "Transmitting" is the act of transferring data or messages to another terminal or system.
[0015] "Security System" means a system designed to protect information and data and to prevent unauthorized access and fraud.
[0016] "Encryption" is the process of transforming data using specific algorithms to protect it from unauthorized access.
[0017] "Push notification" is a method of sending messages and information to a device in real time, allowing users to be notified immediately.
[0018] "SMS" stands for Short Messaging Service, a method of sending short text messages using a phone number.
[0019] An "operator" is a person or function that manages and maintains the system, and is responsible for user support and emergency response.
[0020] "Transmission means" refers to a method or device for transferring data or messages to other terminals or systems. [Brief explanation of the drawings]
[0021] [Figure 1] 1 is a conceptual diagram showing an example of the configuration of a data processing system according to a first embodiment. [Figure 2] 1 is a conceptual diagram showing an example of main functions of a data processing device and a smart device according to a first embodiment. [Figure 3] FIG. 10 is a conceptual diagram showing an example of the configuration of a data processing system according to a second embodiment. [Figure 4] FIG. 10 is a conceptual diagram showing an example of main functions of a data processing device and smart glasses according to a second embodiment. [Figure 5] FIG. 10 is a conceptual diagram showing an example of the configuration of a data processing system according to a third embodiment. [Figure 6] FIG. 11 is a conceptual diagram showing an example of main functions of a data processing device and a headset-type terminal according to a third embodiment. [Figure 7] FIG. 10 is a conceptual diagram showing an example of the configuration of a data processing system according to a fourth embodiment. [Figure 8] FIG. 10 is a conceptual diagram showing an example of main functions of a data processing device and a robot according to a fourth embodiment. [Figure 9] 1 shows an emotion map onto which multiple emotions are mapped. [Figure 10] 1 shows an emotion map onto which multiple emotions are mapped. [Figure 11] FIG. 3 is a sequence diagram showing a processing flow of the data processing system according to the first embodiment. [Figure 12] FIG. 10 is a sequence diagram showing the flow of processing in the data processing system in Application Example 1. [Figure 13] FIG. 10 is a sequence diagram showing the flow of processing in the data processing system according to the second embodiment when an emotion engine is combined. [Figure 14] FIG. 10 is a sequence diagram showing the flow of processing in the data processing system in Application Example 2 when an emotion engine is combined. DETAILED DESCRIPTION OF THE INVENTION
[0022] An example of an embodiment of a system according to the technology of the present disclosure will be described below with reference to the accompanying drawings.
[0023] First, the terms used in the following description will be explained.
[0024] In the following embodiments, a coded processor (hereinafter simply referred to as a "processor") may be a single arithmetic device or a combination of multiple arithmetic devices. Furthermore, a processor may be a single type of arithmetic device or a combination of multiple types of arithmetic devices. Examples of arithmetic devices include a CPU (Central Processing Unit), a GPU (Graphics Processing Unit), a GPGPU (General-Purpose computing on Graphics Processing Units), and an APU (Accelerated Processing Unit).
[0025] In the following embodiments, a coded RAM (Random Access Memory) is a memory in which information is temporarily stored and is used as a working memory by a processor.
[0026] In the following embodiments, the coded storage is one or more non-volatile storage devices that store various programs, various parameters, etc. Examples of non-volatile storage devices include flash memory (SSD (Solid State Drive)), magnetic disks (e.g., hard disks), and magnetic tapes.
[0027] In the following embodiments, a communication I / F (Interface) with a symbol is an interface including a communication processor, an antenna, etc. The communication I / F controls communication between multiple computers. Examples of communication standards applied to the communication I / F include wireless communication standards including 5G (5th Generation Mobile Communication System), Wi-Fi (registered trademark), Bluetooth (registered trademark), etc.
[0028] In the following embodiments, "A and / or B" is synonymous with "at least one of A and B." In other words, "A and / or B" means that it may be only A, only B, or a combination of A and B. Furthermore, in this specification, the same concept as "A and / or B" is also applied when three or more things are expressed connected by "and / or."
[0029] [First embodiment]
[0030] FIG. 1 shows an example of the configuration of a data processing system 10 according to the first embodiment.
[0031] 1, a data processing system 10 includes a data processing device 12 and a smart device 14. An example of the data processing device 12 is a server.
[0032] The data processing device 12 includes a computer 22, a database 24, and a communication I / F 26. The computer 22 is an example of a "computer" according to the technology of the present disclosure. The computer 22 includes a processor 28, a RAM 30, and a storage 32. The processor 28, the RAM 30, and the storage 32 are connected to a bus 34. The database 24 and the communication I / F 26 are also connected to the bus 34. The communication I / F 26 is connected to a network 54. Examples of the network 54 include a WAN (Wide Area Network) and / or a LAN (Local Area Network).
[0033] The smart device 14 includes a computer 36, a reception device 38, an output device 40, a camera 42, and a communication I / F 44. The computer 36 includes a processor 46, a RAM 48, and a storage 50. The processor 46, the RAM 48, and the storage 50 are connected to a bus 52. The reception device 38, the output device 40, and the camera 42 are also connected to the bus 52.
[0034] The reception device 38 includes a touch panel 38A, a microphone 38B, and the like, and receives user input. The touch panel 38A detects contact with an indicator (for example, a pen or a finger) to receive user input by the touch of the indicator. The microphone 38B detects the user's voice to receive user input by voice. The control unit 46A transmits data indicating the user input received by the touch panel 38A and the microphone 38B to the data processing device 12. In the data processing device 12, the specific processing unit 290 acquires the data indicating the user input.
[0035] The output device 40 includes a display 40A and a speaker 40B, and presents data to the user 20 by outputting the data in a form of expression that the user 20 can perceive (for example, audio and / or text). The display 40A displays visible information such as text and images in accordance with instructions from the processor 46. The speaker 40B outputs audio in accordance with instructions from the processor 46. The camera 42 is a compact digital camera equipped with an optical system including a lens, aperture, and shutter, and an imaging element such as a CMOS (Complementary Metal-Oxide-Semiconductor) image sensor or a CCD (Charge Coupled Device) image sensor.
[0036] The communication I / F 44 is connected to a network 54. The communication I / Fs 44 and 26 control the exchange of various information between the processor 46 and the processor 28 via the network 54.
[0037] FIG. 2 shows an example of the main functions of the data processing device 12 and the smart device 14.
[0038] 2, in the data processing device 12, a specific process is performed by the processor 28. A specific processing program 56 is stored in the storage 32. The specific processing program 56 is an example of a "program" according to the technology of the present disclosure. The processor 28 reads the specific processing program 56 from the storage 32 and executes the read specific processing program 56 on the RAM 30. The specific process is realized by the processor 28 operating as a specific processing unit 290 in accordance with the specific processing program 56 executed on the RAM 30.
[0039] The storage 32 stores a data generation model 58 and an emotion identification model 59. The data generation model 58 and the emotion identification model 59 are used by the identification processing unit 290.
[0040] In the smart device 14, the processor 46 performs the reception output process. The storage 50 stores a reception output program 60. The reception output program 60 is used in conjunction with the specific processing program 56 by the data processing system 10. The processor 46 reads the reception output program 60 from the storage 50 and executes the read reception output program 60 on the RAM 48. The reception output process is realized by the processor 46 operating as the control unit 46A in accordance with the reception output program 60 executed on the RAM 48.
[0041] Next, a description will be given of the specific processing performed by the specific processing unit 290 of the data processing device 12. In the following description, the data processing device 12 will be referred to as a "server" and the smart device 14 will be referred to as a "terminal."
[0042] The system of the present invention is composed of a user terminal, a speech recognition server, a generative AI analysis server, and a notification server. Below, we will explain the program processing and specific examples of this system.
[0043] System program processing
[0044] 1. User Device
[0045] The user terminal has the function to automatically start recording the voice when receiving a call. The voice is recorded in real time and encrypted for security. This encrypted voice data is immediately sent to the voice recognition server.
[0046] 2. Speech Recognition Server
[0047] The speech recognition server decrypts the encrypted speech data received from the user's device. The decrypted speech data is converted into text data using speech recognition technology. The converted text data is then sent to the generative AI analysis server.
[0048] 3. Generative AI Analysis Server
[0049] The generative AI analysis server performs detailed analysis of the text data received from the speech recognition server. Based on the text data, an algorithm is run to determine whether there is a possibility of special fraud. If it determines that there is a high possibility of fraud, the generative AI analysis server generates a warning message and sends it to the notification server.
[0050] 4. Notification Server
[0051] The notification server receives alert messages from the generative AI analysis server and, based on the user's settings, selects push notification, SMS, email, or operator contact method to immediately send the alert message to the user or designated contacts.
[0052] Specific examples
[0053] Example 1: Fraudulent call detection and warning
[0054] When a user answers a call, the user's device automatically records the call audio, encrypts it, and sends it to a speech recognition server. The speech recognition server converts the audio into text and sends it to a generative AI analysis server. The generative AI analysis server analyzes the text, and if it detects a pattern that indicates possible fraud, such as "I was in a traffic accident and need money," it generates a warning message and sends it to a notification server. The notification server then sends a push notification to the user saying, "This may be a scam. Please end the call immediately."
[0055] Example 2: Secure Call
[0056] For other calls, the user's device records the voice in the same way, and the speech recognition server converts the voice into text. The generative AI analysis server determines that the text data does not have the characteristics of a special fraud and does not generate a warning message. In this case, the user continues the call as usual.
[0057] As described above, the present invention is a system that can prevent damage from special frauds by analyzing telephone voices in real time, determining the possibility of special fraud, and immediately sending a warning to the user.
[0058] The processing flow will be explained below.
[0059] Step 1:
[0060] The user terminal detects an incoming or outgoing call, and when the user receives or makes a call, the voice recording function is automatically started.
[0061] Step 2:
[0062] The user device records the call audio in real time, and the recorded audio data is encrypted using an encryption algorithm such as AES-256 to ensure security.
[0063] Step 3:
[0064] The user device sends encrypted voice data to a speech recognition server over an internet connection, using HTTPS to ensure data integrity and security.
[0065] Step 4:
[0066] The voice recognition server decrypts the encrypted voice data received from the user terminal. If the decryption is successful, the original voice data is reproduced.
[0067] Step 5:
[0068] The speech recognition server processes the decoded speech data with a speech recognition engine and converts it into text data, thereby extracting the speech content as text information.
[0069] Step 6:
[0070] The speech recognition server sends the converted text data to the generative AI analysis server, using a secure protocol to ensure data confidentiality.
[0071] Step 7:
[0072] The generative AI analysis server analyzes the text data received from the speech recognition server, applies an algorithm to detect the characteristics of special fraud, and determines the likelihood of fraud based on the text data.
[0073] Step 8:
[0074] If the generative AI analysis server determines that there is a high possibility of fraud, it generates a warning message that indicates the possibility of fraud and provides specific instructions to the user.
[0075] Step 9:
[0076] The generative AI analysis server sends the generated warning message to the notification server, which selects the notification method according to the user's settings.
[0077] Step 10:
[0078] The notification server sends warning messages to users or designated contacts via push notification, SMS, email, or operator, allowing users to receive prompt warnings and take appropriate action.
[0079] Example 1
[0080] Next, a description will be given of Example 1. In the following description, the data processing device 12 will be referred to as a "server" and the smart device 14 will be referred to as a "terminal."
[0081] As the number of victims of special telephone frauds increases, there is a growing need for security systems that use advanced technology to detect fraud in real time and quickly warn users. However, current systems lack the functionality to automatically analyze the content of phone calls, detect possible fraud, and immediately warn users. Therefore, the present invention aims to solve these problems.
[0082] The specific processing by the specific processing unit 290 of the data processing device 12 in the first embodiment is realized by the following means.
[0083] In this invention, the server includes means for detecting an incoming call on the user terminal and automatically recording the call audio, means for encrypting the recorded call audio data, means for transmitting the encrypted audio data to a speech recognition server, means for the speech recognition server to decrypt the encrypted audio data and convert it into text data, means for the generative AI analysis server to analyze the text data and determine the possibility of special fraud, means for generating a warning message based on the determination, and means for transmitting the generated warning message to a notification server and notifying the user. This makes it possible to automatically detect the possibility of fraud during a call and send an appropriate warning to the user in real time.
[0084] A "user terminal" is a device that has the function of detecting an incoming call, automatically recording the call audio, and encrypting and transmitting the recorded audio data.
[0085] "Encryption" is a technology that converts recorded call voice data for security purposes and prevents unwanted access.
[0086] A "voice recognition server" is a server that has the function of receiving encrypted voice data sent from a user terminal, decrypting it, and converting the voice data into text data.
[0087] The "generative AI analysis server" is a server that analyzes text data sent from the voice recognition server and executes an algorithm to determine the possibility of special fraud.
[0088] The "notification server" is a server whose role is to receive warning messages provided by the generative AI analysis server and to send the messages to users using an appropriate notification method.
[0089] A "warning message" is a message containing a warning to the user that is generated when the generative AI analysis server detects the possibility of fraud.
[0090] "Push notifications" are short messages sent directly to a user's device, providing an instant way to communicate important information from applications and services.
[0091] "Short Message Service (SMS)" is a communications protocol for sending and receiving short text messages using mobile devices.
[0092] "Email" is a method of sending and receiving text-based messages over the Internet and is a widely used means of communicating information to users.
[0093] "Communication means" refers to any method for transmitting information from a sender to a receiver, and in this system includes push notifications, short message services, emails, etc.
[0094] The system of the present invention consists of a user terminal, a speech recognition server, a generative AI analysis server, and a notification server. The user terminal automatically records the voice of the call, encrypts the data, and sends it to the speech recognition server. The speech recognition server decrypts the recorded data and converts it into text using speech recognition technology. The converted text data is sent to the generative AI analysis server, where it is analyzed to determine the possibility of fraud. If it is determined that there is a high possibility of fraud, a warning message is sent to the user via the notification server.
[0095] The user device can be a smartphone or tablet. When the user answers a call, the device automatically starts recording the call audio. The recorded data is AES encrypted in real time and sent to the speech recognition server via an HTTP POST request.
[0096] The speech recognition server runs, for example, a Python-based Flask web server. The received encrypted speech data is decrypted using Python's Cryptography library. The speech data is then converted into text data using a service such as the Google® Cloud Speech-to-Text API. This text data is then sent to the generative AI analysis server.
[0097] The generative AI analysis server runs a server built using, for example, Node.js or Python. It analyzes the received text data using a generative AI model, such as OpenAI's GPT-3 model, to detect specific keywords and contexts that indicate the possibility of fraud. For example, it identifies patterns such as "I was in a traffic accident and need money." If it determines that there is a high possibility of fraud, it generates a warning message and sends it to the notification server.
[0098] The notification server runs a server built using, for example, Ruby on Rails or Django. It notifies users of received alert messages using methods such as push notifications, SMS, and email. For example, you can send SMS using the Twilio API or push notifications using Firebase Cloud Messaging.
[0099] Specific examples
[0100] Example 1: Fraudulent call detection and warning
[0101] When a user answers a call, the user's device automatically records the call audio, encrypts it, and sends it to a speech recognition server. The speech recognition server converts the audio into text data and sends the text data to a generative AI analysis server. The generative AI analysis server analyzes the text and detects patterns that indicate possible fraud, such as "I was in a traffic accident and need money." In this case, it generates a fraud warning message and sends it to a notification server. The notification server then sends a push notification to the user saying, "This may be a scam. Please end the call immediately."
[0102] Example 2: Secure Call
[0103] For other calls, the user's device records the voice in the same way, and the speech recognition server converts the voice into text data. The generative AI analysis server analyzes the text data and determines that it does not have the characteristics of a special fraud. In this case, no warning message is generated, and the user can continue the call as usual.
[0104] Example prompt: "Is this a potential scam? 'I've been in a car accident and need money. Please transfer the money.'"
[0105] The flow of the identification process in the first embodiment will be described with reference to FIG.
[0106] Step 1:
[0107] The user device detects an incoming call and automatically starts recording the call audio. The input is the incoming call event. The output is the recorded call audio data. Specifically, the device software catches the incoming call event and starts the recording module. For example, the smartphone's OS notifies the call app of the incoming call, and an application written in Java (registered trademark) starts recording using the ANDROID (registered trademark) MediaRecorder API.
[0108] Step 2:
[0109] The recorded call audio data is encrypted in real time. The input is the audio data obtained in step 1. The output is the encrypted audio data. Specifically, the device's encryption module encrypts the recorded data using an encryption algorithm such as AES. For example, an application written in Java encrypts the data using the AES encryption algorithm.
[0110] Step 3:
[0111] The encrypted voice data is sent to the voice recognition server. The input is the encrypted voice data obtained in step 2. The output is a notification that the data has been sent to the server. Specifically, the device sends the voice data to the server using an HTTP POST request. For example, an application written in Java uses an HTTP POST request.
[0112] Step 4:
[0113] The speech recognition server receives and decrypts the encrypted audio data. The input is the encrypted audio data sent in step 3. The output is the decrypted audio data. Specifically, the server's encryption / decryption module decrypts the data using the AES decryption algorithm. For example, a Flask web server written in Python processes the request and decrypts it using the Cryptography library.
[0114] Step 5:
[0115] The speech recognition server converts the decoded speech data into text data. The input is the decoded speech data obtained in step 4. The output is text data. Specifically, speech-to-text conversion is performed using a speech recognition API. For example, the Google Cloud Speech-to-Text API is used to convert speech data into text.
[0116] Step 6:
[0117] The converted text data is sent to the generative AI analysis server. The input is the text data obtained in step 5. The output is a notification that data transmission to the generative AI analysis server has been completed. Specifically, the speech recognition server sends the data using an HTTP POST request. For example, the Python code sends a POST request to the generative AI analysis server.
[0118] Step 7:
[0119] The generative AI analysis server analyzes the text data and determines the likelihood of fraud. The input is the text data received in step 6. The output is the fraud determination result. Specifically, it uses a generative AI model to analyze the text and detect keywords and contexts that indicate the likelihood of fraud. For example, it uses OpenAI's GPT-3 model to identify phrases such as "I was in a traffic accident and need money."
[0120] Step 8:
[0121] If it is determined that there is a high possibility of fraud, a warning message is generated. The input is the fraud determination result from step 7. The output is a warning message. Specifically, the server runs the message generation module to create a warning message. For example, it creates a message saying, "There is a possibility of fraud. Please end the call immediately."
[0122] Step 9:
[0123] The generated warning message is sent to the notification server. The input is the warning message from step 8. The output is a notification of completion of data transmission to the notification server. Specifically, the generative AI analysis server sends the data using an HTTP POST request. For example, JavaScript (registered trademark) code sends a POST request to the notification server.
[0124] Step 10:
[0125] The notification server sends an alert message to the user. The input is the alert message received in step 9. The output is an alert notification to the user. Specifically, the notification server sends the message to the user using the appropriate notification method. For example, it can send an SMS using the Twilio API or a push notification using Firebase Cloud Messaging.
[0126] (Application example 1)
[0127] Next, a description will be given of Application Example 1. In the following description, the data processing device 12 will be referred to as a "server" and the smart device 14 will be referred to as a "terminal."
[0128] In recent years, special fraud cases using telephones have been increasing, and many people have become victims. For this reason, there is a need for a system that can detect possible fraud in real time when a call is received and issue a warning. However, existing systems have problems such as low accuracy in fraud detection, lack of real-time response, and a lack of diverse means of notification to users. To address this issue, a system is needed that can detect possible fraud more accurately and quickly, and send warnings to users using a variety of means.
[0129] The specific processing by the specific processing unit 290 of the data processing device 12 in the application example 1 is realized by the following means.
[0130] In this invention, the server includes means for recording telephone voice, means for converting the recorded telephone voice into text data, means for analyzing the converted text data and determining the possibility of special fraud, means for generating a warning message based on the determination, means for transmitting the generated warning message, means for analyzing the text data using a generative AI model, means for starting recording in real time on a user terminal, and means for generating and transmitting a warning message by push notification, SMS, or email. This makes it possible to quickly detect the possibility of fraud with high accuracy and to send warnings to users by various means.
[0131] "Telephone voice" refers to voice data transmitted through telephone communication.
[0132] "Recording means" refers to devices or software for saving voice data, and has the function of saving the contents of telephone conversations as files.
[0133] "Means for converting into text data" refers to devices or software that use voice recognition technology to convert voice data into character-based data.
[0134] "Means for analysis" refers to devices or software that have the ability to analyze text data and evaluate its content based on pattern recognition or algorithms.
[0135] "Special fraud" refers to criminal acts that involve deceiving people over the phone and stealing money from them.
[0136] "Means for determining" refers to a device or software that runs an algorithm to determine whether or not a transaction is fraudulent based on the analysis results.
[0137] A "warning message" refers to a message containing information that is notified to the user when a possible fraud is detected.
[0138] "Generating means" refers to a device or software for creating a warning message, and has the function of automatically creating the necessary wording and warning content.
[0139] "Transmitting means" refers to the communication technology and network infrastructure for transferring the generated alert message to a user terminal or other device.
[0140] "Generative AI models" refer to algorithms and tools that use artificial intelligence to generate and analyze data.
[0141] A "prompt" is an instruction or guidance text to be input into a generative AI model, and is text used to obtain a specific analysis result or product.
[0142] "User terminal" refers to a communication device that is directly operated by a user, such as a smartphone, tablet, or computer.
[0143] "Means for starting recording in real time" refers to devices or software that have the function of automatically starting to record audio the moment a call is initiated.
[0144] A "push notification" refers to a notification message sent to a user device in real time from an application or system.
[0145] "SMS" stands for Short Message Service, a service for sending short text messages over mobile phone networks.
[0146] "Mail" is an abbreviation for electronic mail, a means of communication for sending and receiving text and files over the Internet.
[0147] The present invention is a security system that detects potential fraud in real time for telephone calls received by a user and issues a prompt warning. The system includes the following main components:
[0148] User terminal
[0149] The user terminal is equipped with a function for recording telephone calls in real time. The voice is recorded as soon as the call begins, and the voice data is encrypted to ensure security. The encrypted voice data is immediately sent to a voice recognition server. This allows fraud detection to be performed automatically without the user having to perform any special operations during the call.
[0150] Speech Recognition Server
[0151] The speech recognition server decrypts the encrypted voice data received from the user's device. The decrypted voice data is converted into text data using speech recognition technology, such as Google Speech Recognition. The converted text data is then sent to the generative AI analysis server.
[0152] Generative AI analysis server
[0153] The generative AI analysis server analyzes the received text data and runs algorithms to determine the likelihood of fraud. This analysis uses generative AI models such as BERT and GPT. If a high likelihood of fraud is determined, the generative AI analysis server generates a warning message and sends it to the notification server.
[0154] Specific prompt examples:
[0155] Analyze this text and determine if it is potentially fraudulent if it contains phrases like "I was in a car accident" or "I need money."
[0156] Text: "I was in a car accident and need money now. Please transfer it to my bank account."
[0157] Notification Server
[0158] The notification server notifies the user device of the warning message received from the generative AI analysis server via push notification, SMS, email, etc. This allows the user to respond immediately to potentially fraudulent calls.
[0159] Specific examples
[0160] When a user receives a phone call that is suspected to be fraudulent, the user's device automatically records the audio and sends the data to a speech recognition server. The speech recognition server converts the audio data into text data, which is then analyzed by a generative AI analysis server. If the analysis determines that the call is likely to be fraudulent, the notification server sends a warning message to the user, informing them that "This may be a fraudulent call. Please end the call immediately."
[0161] This allows users to prevent themselves from falling victim to fraud. By combining multiple cutting-edge technologies, the system is able to detect potential fraud with high accuracy and speed, and send warnings via a variety of means.
[0162] The flow of the specific processing in the application example 1 will be described with reference to FIG.
[0163] Step 1:
[0164] When the user terminal detects the start of a call, it automatically records the voice in real time. The recorded voice data is encrypted to ensure security. For encryption, for example, the Fernet encryption library is used. The encrypted voice data is then sent to a voice recognition server.
[0165] Input: Phone call audio
[0166] Data processing: Encrypting audio
[0167] Output: Encrypted audio data
[0168] Step 2:
[0169] The voice recognition server decrypts the received encrypted voice data. For decryption, the same encryption key as that used on the user's device is used. The decrypted voice data is converted into text data using voice recognition technology. For example, the Google Speech Recognition API is used. The converted text data is sent to the generative AI analysis server.
[0170] Input: Encrypted audio data
[0171] Data operations: Decryption of encrypted data and conversion of voice data to text
[0172] Output: Text data
[0173] Step 3:
[0174] The generative AI analysis server uses a generative AI model (e.g., BERT or GPT-3) to analyze the received text data. It uses prompt sentences to analyze the text data and determine the likelihood of fraud. Specific examples of prompt sentences are as follows:
[0175] Analyze this text and determine if it is potentially fraudulent if it contains phrases like "I was in a car accident" or "I need money."
[0176] Text: "I was in a car accident and need money now. Please transfer it to my bank account."
[0177] Input: Text data, prompt
[0178] Data Computation: Text Data Analysis with Generative AI Models
[0179] Output: Possibility of fraud determination result
[0180] Step 4:
[0181] If the generative AI analysis server determines that there is a high possibility of fraud, it generates a warning message, such as "This is a possible fraud. Please end the call immediately." This message is sent to the notification server.
[0182] Input: Possibility of fraud determination result
[0183] Data processing: Generate warning messages
[0184] Output: Warning message
[0185] Step 5:
[0186] The notification server notifies the user device of the warning message received from the generative AI analysis server. The warning message is sent to the user immediately using multiple means, such as push notification, SMS, and email, allowing the user to quickly respond to potentially fraudulent calls.
[0187] Input: warning message
[0188] Data processing: Sending notification messages
[0189] Output: Warning message sent to the user
[0190] The above are the specific processing steps of the program for the system that realizes the application example.
[0191] Furthermore, an emotion engine that estimates the user's emotion may be combined. That is, the identification processing unit 290 may estimate the user's emotion using the emotion identification model 59 and perform identification processing using the user's emotion.
[0192] The system of the present invention is composed of a user terminal, a speech recognition server, a generative AI analysis server, a notification server, and an emotion engine. Below, we will explain the program processing and specific examples of this system.
[0193] System program processing
[0194] 1. User Device
[0195] The user terminal has the function to automatically start recording the voice when receiving a call. The voice is recorded in real time and encrypted for security. This encrypted voice data is immediately sent to the voice recognition server.
[0196] 2. Speech Recognition Server
[0197] The speech recognition server decrypts the encrypted speech data received from the user's device. The decrypted speech data is converted into text data using speech recognition technology. The converted text data is then sent to the generative AI analysis server.
[0198] 3. Generative AI Analysis Server
[0199] The generative AI analysis server performs detailed analysis of the text data received from the speech recognition server. Based on the text data, an algorithm is run to determine whether or not there is a possibility of special fraud.
[0200] 4. Emotion Engine
[0201] The emotion engine analyzes the voice data acquired from the user's device to recognize the user's emotional state. The results of the emotion engine are sent to the generative AI analysis server, which influences the determination of the possibility of fraud.
[0202] 5. Correction of analysis results
[0203] The generative AI analysis server receives the analysis results from the emotion engine and amends them based on the user's emotional state. For example, if the user is nervous or confused, the analysis results are strengthened.
[0204] 6. Generating and Sending Warning Messages
[0205] The generative AI analysis server generates warning messages as needed based on the corrected analysis results. The generated warning messages are sent to the notification server, which then sends the warning messages via push notification, SMS, email, or operator notification based on the user's settings.
[0206] Specific examples
[0207] Example 1: Fraudulent call detection and warning
[0208] When a user answers a call, the user's device automatically records the call audio, encrypts it, and sends it to a speech recognition server. The speech recognition server converts the speech into text and sends it to a generative AI analysis server. The generative AI analysis server analyzes the text and detects patterns that indicate the possibility of fraud, such as "I was in a traffic accident and need money." If the emotion engine detects a state of tension in the user's voice, it further determines that the possibility of fraud is high. A warning message is generated, and a push notification is sent to the user via the notification server, stating, "This may be a scam. Please end the call immediately."
[0209] Example 2: Secure Call
[0210] The user receives another call. In this case, the user's device again records the audio, and the speech recognition server converts it into text. The generative AI analysis server analyzes the text data and verifies that it does not have the characteristics of a specialized fraud. The emotion engine also detects that the user is relaxed. In this case, no warning message is generated, and the user continues the call as usual.
[0211] As described above, the present invention is a system that can prevent damage from special frauds by analyzing telephone voices in real time, taking into account the user's emotional state, determining the possibility of special fraud, and sending a warning.
[0212] The processing flow will be explained below.
[0213] Step 1:
[0214] The user receives or makes a call. The user device detects this and automatically starts recording the audio.
[0215] Step 2:
[0216] The user device records the call audio in real time, and the recorded audio data is encrypted using an encryption algorithm such as AES-256.
[0217] Step 3:
[0218] The user terminal sends the encrypted voice data to the voice recognition server using the HTTPS protocol.
[0219] Step 4:
[0220] The voice recognition server decrypts the encrypted voice data received from the user terminal. If the decryption is successful, the voice data returns to its original state.
[0221] Step 5:
[0222] The speech recognition server converts the decoded speech data into text data using speech recognition technology, and the converted text data is sent to the generative AI analysis server.
[0223] Step 6:
[0224] The generative AI analysis server analyzes the text data received from the speech recognition server, and executes a specific algorithm to determine the possibility of special fraud.
[0225] Step 7:
[0226] The user device sends the recorded voice data to the emotion engine, which analyzes the user's emotional state (e.g., tension, confusion, calmness, etc.) and sends the results to the generative AI analysis server.
[0227] Step 8:
[0228] The generative AI analysis server receives the analysis results from the emotion engine and corrects the analysis results for the possibility of special fraud. If the user is nervous or confused, the possibility of fraud is strengthened.
[0229] Step 9:
[0230] Based on the final judgment based on the results of the emotion engine, the generative AI analysis server generates a warning message, which includes information about the suspected fraud and specific instructions for the user.
[0231] Step 10:
[0232] The generative AI analysis server sends the generated warning message to the notification server, which selects the method of notification based on the user's settings: push notification, SMS, email, or operator notification.
[0233] Step 11:
[0234] The notification server sends alert messages to the user or designated contacts via the method of their choice, allowing the user to receive prompt warnings and take appropriate action.
[0235] Example 2
[0236] Next, a description will be given of Example 2. In the following description, the data processing device 12 will be referred to as a "server" and the smart device 14 will be referred to as a "terminal."
[0237] In modern times, special telephone frauds have become a social problem. Fraudulent phone calls targeting the elderly in particular have caused many victims, making countermeasures an urgent necessity. However, conventional security systems lack the functionality to detect potential fraud in real time and issue immediate warnings to users. Furthermore, they lack the functionality to correct analysis results by taking the user's emotional state into account, making them prone to false positives and overreactions. The purpose of this invention is to solve these problems and provide a highly reliable fraud prevention system.
[0238] The identification process by the identification processing unit 290 of the data processing device 12 in the second embodiment is realized by the following means. In this invention, the server includes a means for recording telephone voice, a means for encrypting and transmitting the recorded telephone voice, a means for converting the encrypted telephone voice into text data, a means for analyzing the converted text data and determining the possibility of special fraud, a means for analyzing the user's emotional state and correcting the analysis result, a means for generating a warning message based on the corrected result, and a means for transmitting the generated warning message. This enables real-time detection of the possibility of fraud and correction of the analysis result taking the user's emotional state into consideration.
[0239] "Telephone voice" refers to audio data recorded from a telephone conversation.
[0240] A "recording means" is a device or system that stores the audio during a call as digital data.
[0241] "Encryption and transmission means" refers to the methods and techniques by which stored voice data is encrypted for security purposes and transmitted to the appropriate recipient.
[0242] "Means for converting into text data" refers to technology for converting voice data into character data, such as voice recognition technology.
[0243] "Means for analyzing and determining the possibility of special fraud" refers to algorithms and systems that analyze text data and detect signs of fraud from its content.
[0244] "Means for analyzing emotional state" refers to technology or a system for determining emotions (tension, relaxation, etc.) from the user's voice data.
[0245] "Means for correcting the analysis results" refers to a method for correcting the results as necessary to increase the reliability of the analysis results based on information on emotional state.
[0246] A "means for generating a warning message" is a system that generates a warning message to a user when fraud is likely.
[0247] The "means for sending a warning message" refers to a method for sending the generated warning message to the user using a communication means.
[0248] The system of the present invention is composed of a user terminal, a speech recognition server, a generative AI analysis server, an emotion engine, and a notification server. A specific example of the system will be described below.
[0249] User terminal
[0250] The user device has a function that automatically starts recording voice when a call comes in. The recorded voice data is encrypted in real time using the AES method or similar. This encrypted voice data is immediately sent to the voice recognition server using the HTTPS protocol.
[0251] Speech Recognition Server
[0252] The speech recognition server decrypts the encrypted speech data received from the user's device. This decrypted speech data is converted into text data using speech recognition technology. Specifically, it uses a commonly used speech recognition API, such as the speech recognition function of a cloud service. This converted text data is then sent to the generative AI analysis server.
[0253] Generative AI analysis server
[0254] The generative AI analytics server uses machine learning algorithms to analyze the text data received from the speech recognition server. For example, it uses open-source machine learning libraries or cloud-based AI services to analyze the text data and determine the likelihood of fraud. If patterns indicative of possible fraud are detected, the information is sent to the emotion engine.
[0255] Emotion Engine
[0256] The emotion engine analyzes the user's emotional state based on their voice data. The emotional state is determined using commonly used emotion analysis APIs and software. This information is sent to a generative AI analysis server, which corrects the analysis results. For example, if the user is nervous, it is deemed to be a sign of a high possibility of fraud.
[0257] Correcting analysis results and generating warning messages
[0258] Once the analysis results have been corrected, the generative AI analysis server generates a warning message as needed. This warning message is sent to the user's device in real time. Specifically, the warning message is sent to the notification server and delivered to the user via push notification, SMS, email, or other means. For example, the message might say, "This may be a scam. Please end the call immediately."
[0259] Specific examples
[0260] Example 1: Fraudulent call detection
[0261] When a user answers a call, the user device automatically records the call audio and sends the encrypted audio data to a speech recognition server. The speech recognition server converts the audio into text and sends it to a generative AI analysis server. The generative AI analysis server analyzes the text data and detects signs of fraud, such as "I was in a traffic accident and need money." If the emotion engine detects a state of tension in the user's voice, the possibility of fraud increases. Based on this, a warning message is generated and sent to the user via a notification server.
[0262] Example 2: Secure Call
[0263] The user receives another call and the device begins recording the audio. This audio data is also encrypted and sent to the speech recognition server. The recorded audio data is converted into text and sent to the generative AI analysis server. The generative AI analysis server analyzes this text data and verifies that it does not contain any characteristics of specialized fraud. The emotion engine also detects that the user is relaxed, so no warning message is generated. This allows the user to continue the safe call.
[0264] Prompt Sentence Examples
[0265] "The audio of phone calls received by users is recorded in real time and encrypted for security. This encrypted audio data is converted into text data using speech recognition technology and analyzed by a generative AI analysis server. Based on the analyzed text data, the possibility of fraud is determined and a warning message is generated if necessary. The warning message is sent to the user via a notification server to ensure a safe call. If signs of fraud are detected, how should it be handled?"
[0266] As a result, the present invention is a system that can prevent damage from special fraud by analyzing telephone voices in real time, determining the possibility of special fraud taking into account the user's emotional state, and sending a warning.
[0267] The flow of the identification process in the second embodiment will be described with reference to FIG.
[0268] Processing steps of this system's program
[0269] Step 1
[0270] When a user receives a call, the user terminal automatically records the call audio. This audio recording is initiated using a program that is automatically triggered when the user receives a call. The input is the real-time call audio, and the output is the recorded audio file. Specifically, the user terminal launches a recording application and saves the call audio as a file.
[0271] Step 2
[0272] The user terminal encrypts the recorded voice data using the AES encryption method. At this point, the input is the recorded voice file, and the output is the encrypted voice data. Specifically, the user terminal runs an encryption algorithm to convert the voice data into a secure format.
[0273] Step 3
[0274] The user terminal sends encrypted voice data to the voice recognition server. The HTTPS protocol is used for transmission. The input is encrypted voice data, and the output is the encrypted voice data received by the voice recognition server. Specifically, the user terminal creates an HTTP request and sends the data.
[0275] Step 4
[0276] The speech recognition server decrypts the received encrypted audio data. The input is the encrypted audio data, and the output is the decrypted audio data. Specifically, the server runs the AES decryption algorithm to reconstruct the original audio data.
[0277] Step 5
[0278] The speech recognition server converts the decoded speech data into text data. Specifically, it uses a speech recognition API to convert speech to text. The input is the decoded speech data, and the output is text data. Specifically, the speech recognition server calls the speech recognition API (for example, a cloud-based speech recognition service) and obtains the returned text data.
[0279] Step 6
[0280] The speech recognition server sends the generated text data to the generative AI analysis server. The input is the generated text data, and the output is the text data received by the generative AI analysis server. Specifically, the data is sent via a REST API for server-to-server communication.
[0281] Step 7
[0282] The generative AI analysis server analyzes the received text data and runs machine learning algorithms to determine the likelihood of fraud. The input is text data, and the output is a result indicating the likelihood of fraud. Specifically, the generative AI analysis server uses libraries such as TENSORFLOW (registered trademark) and Hugging Face's Transformers to analyze the text data and detect fraudulent patterns.
[0283] Step 8
[0284] The emotion engine analyzes the user's emotional state based on their voice data. The input is the user's voice data, and the output is data related to the user's emotional state (tension, confusion, etc.). Specifically, the emotion engine calls an emotion analysis API (e.g., IBM Watson (registered trademark) Tone Analyzer) to determine the user's emotional state.
[0285] Step 9
[0286] The generative AI analysis server corrects the analysis results based on the emotional state data received from the emotion engine. The input is the emotional state data and text analysis results, and the output is the corrected analysis results. Specifically, it reevaluates the possibility of fraud taking the emotional state into account and makes a final judgment.
[0287] Step 10
[0288] The generative AI analysis server generates a warning message based on the corrected analysis results. The input is the corrected analysis results, and the output is the generated warning message. Specifically, the server creates a warning message based on the judgment results.
[0289] Step 11
[0290] The notification server sends the generated warning message to the user device. The sending method can be push notification, SMS, email, etc. The input is the generated warning message, and the output is the warning message received by the user device. Specifically, the notification server sends the message to the user using AWS (registered trademark) SNS (Simple Notification Service) or a similar notification service.
[0291] The above are the specific processing steps in this system.
[0292] (Application example 2)
[0293] Next, a description will be given of Application Example 2. In the following description, the data processing device 12 will be referred to as a "server" and the smart device 14 will be referred to as a "terminal."
[0294] In recent years, there has been an increase in special frauds using telephones, and frauds targeting the elderly in particular have become a serious problem. Current security systems lack the means to accurately detect potential fraud in real time and send prompt and appropriate warnings to users. As a result, many users remain at high risk of becoming victims of fraud. Furthermore, few systems take the user's emotional state into account, and the accuracy of fraud detection is low, making it difficult for users to feel safe answering the phone. There is a need for a system that solves this problem.
[0295] The identification process by the identification processing unit 290 of the data processing device 12 in Application Example 2 is realized by the following means. In this invention, the server includes a means for recording telephone voice, a means for encrypting and transmitting the recorded telephone voice, and a means for decrypting the encrypted voice data. This enables secure and highly accurate analysis of telephone voice in real time and fraud detection that takes into account the user's emotional state. In addition, by including a means for quickly sending a warning message via push notification, SMS, email, or operator, it is possible to immediately warn the user of fraud and prevent damage before it occurs.
[0296] "Means for recording telephone voice" refers to technology that automatically records and safely stores the contents of a call when a user initiates a call.
[0297] "Means for encrypting and transmitting recorded telephone voice data" refers to a technology that encrypts recorded telephone voice data to protect it from unauthorized access or eavesdropping by third parties and transmits it to a designated server.
[0298] "Means for decrypting encrypted audio data" refers to a technique for restoring the transmitted encrypted audio data to the original audio data.
[0299] "Means for converting voice data into text data" refers to a technology that analyzes decoded voice data and converts its contents into text information.
[0300] The "means of analyzing the converted text data and determining the possibility of special fraud" is a technology that uses an algorithm to detect signs and patterns of special fraud based on the obtained text data.
[0301] "Means for analyzing the user's emotional state" refers to a technology that analyzes the user's emotions (tension, confusion, fear, etc.) from voice data during a call and determines that state.
[0302] The "means for correcting the determination of the possibility of special fraud based on the analysis results of the emotional state" is a technology for improving the accuracy of the determination of the possibility of special fraud by taking into account the emotional state of the user.
[0303] The "means for generating a warning message" is a technology that automatically generates a notification message to warn the user when it is determined that there is a high possibility of special fraud.
[0304] "Means for sending the generated warning message" refers to a communication technology for immediately delivering the generated warning message to the user, including push notification, SMS, email, or operator notification.
[0305] The purpose of the system of the present invention is to prevent special frauds using voice calls, and it records, encrypts, analyzes telephone voice, and generates and transmits warning messages. The specific system configuration and its program processing are described below.
[0306] System Configuration
[0307] The system of the present invention comprises the following components:
[0308] 1. User device: When a call is initiated, the audio is automatically recorded, and the recorded audio data is encrypted and sent to the server.
[0309] 2. Speech recognition server: Decrypts the encrypted voice data and converts the voice into text data.
[0310] 3. Generative AI analysis server: Analyzes text data to determine the possibility of special fraud. It also analyzes the user's emotional state and adjusts the judgment results based on the results.
[0311] 4. Notification Server: Generates warning messages as needed and sends them to users via push notification, SMS, email or operator.
[0312] Hardware and Software Used
[0313] Smartphone: iOS / Android devices
[0314] Cloud servers: Amazon Web Services (AWS) EC2, Amazon S3
[0315] Speech recognition technology: Google Speech-to-Text API
[0316] Generative AI model: OpenAI GPT-4 (registered trademark) API
[0317] Sentiment analysis engine: IBM Watson Tone Analyzer
[0318] Notification service: Firebase Cloud Messaging (FCM)
[0319] Program Processing Overview
[0320] 1. User Device
[0321] The user device automatically records the voice as soon as the call begins and encrypts the recording using a sophisticated encryption algorithm. The encrypted voice data is then immediately sent to a voice recognition server.
[0322] 2. Speech Recognition Server
[0323] The speech recognition server decrypts the received encrypted voice data and converts it into text using the Google Speech-to-Text API. The converted text data is then sent to the generative AI analysis server.
[0324] 3. Generative AI Analysis Server
[0325] The generative AI analysis server uses the OpenAI GPT-4 API to perform detailed analysis of the text data sent from the speech recognition server. To determine whether there are signs of fraud, it detects specific keywords and phrases and then uses IBM Watson Tone Analyzer to analyze the user's emotional state. Based on the results, it corrects the judgment results as necessary.
[0326] 4. Notification Server
[0327] The notification server receives the results of the generative AI analysis server and generates a warning message if it determines that the call is likely to be fraudulent, and sends it to the user via push notification, SMS, or email via Firebase Cloud Messaging (FCM). This warning message urges the user to end the call.
[0328] Examples and prompts
[0329] Examples:
[0330] For example, if a user receives a call while on a call saying, "My son was in a traffic accident and I need money immediately," the system will analyze the call based on the following prompt:
[0331] Example prompt sentence:
[0332] The user received a call. The text of the call is below:
[0333] "Hello, this is the police. Your son has been in a car accident. We need money."
[0334] The user's emotional state is "tense."
[0335] Determine if this is likely a scam based on the following criteria, and if so, generate a warning message.
[0336] conditions:
[0337] Keywords to detect possible fraud: traffic accident, need money, transfer, etc.
[0338] User's emotional state: nervous, confused, scared
[0339] Output format:
[0340] Possibility of fraud: [Yes / No]
[0341] Warning message: [Possible scam. Please end the call immediately / This call is safe.]
[0342] This system makes it possible to detect the risk of special fraud in real time and quickly protect users.
[0343] The flow of the specific processing in the application example 2 will be described with reference to FIG.
[0344] Step 1:
[0345] When a user receives a call, the user terminal automatically starts recording the call audio. The recorded audio data is encrypted using an encryption algorithm. The input is the call audio data, and the output is the encrypted audio data.
[0346] Step 2:
[0347] The encrypted voice data is sent from the user terminal to the voice recognition server, which then decrypts the received data. The input is the encrypted voice data, and the output is the decrypted voice data.
[0348] Step 3:
[0349] The speech recognition server converts the decoded speech data into text data using the Google Speech-to-Text API. The input is speech data and the output is text data.
[0350] Step 4:
[0351] The generative AI analysis server receives the text data sent from the speech recognition server and uses the OpenAI GPT-4 API to analyze the possibility of fraud by detecting keywords and phrases within the text data. The input is the text data, and the output is the analysis results.
[0352] Step 5:
[0353] The generative AI analysis server analyzes the user's emotional state using the IBM Watson Tone Analyzer, an emotion analysis engine, in addition to analyzing the text data. The input is voice data, and the output is the user's emotional state.
[0354] Step 6:
[0355] The generative AI analysis server corrects the judgment of the possibility of special fraud based on the emotion analysis results. For example, if the user is nervous, it will judge that there is a high possibility of fraud. The input is the analysis result and emotional state, and the output is the corrected judgment result.
[0356] Step 7:
[0357] Based on the corrected judgment results, the generative AI analysis server generates warning messages as necessary. The input is the corrected judgment results, and the output is the warning message.
[0358] Step 8:
[0359] The notification server receives the generated alert messages and uses Firebase Cloud Messaging (FCM) to send them to the user via push notification, SMS, email, or operator. The input is the alert message and the output is the notification to the user.
[0360] The specific processing unit 290 transmits the result of the specific processing to the smart device 14. In the smart device 14, the control unit 46A causes the output device 40 to output the result of the specific processing. The microphone 38B acquires audio indicating a user input regarding the result of the specific processing. The control unit 46A transmits audio data indicating the user input acquired by the microphone 38B to the data processing device 12. In the data processing device 12, the specific processing unit 290 acquires the audio data.
[0361] The data generation model 58 is a so-called generative AI (Artificial Intelligence). An example of the data generation model 58 is ChatGPT (registered trademark) (Internet search engine).<URL: https: / / openai.com / blog / chatgpt> ), Gemini (registered trademark) (Internet search <url: https: gemini.google.com ?hl="ja">) and other generation AIs. The data generation model 58 is obtained by performing deep learning on a neural network. A prompt including an instruction is input to the data generation model 58, and inference data such as voice data indicating voice, text data indicating text, and image data indicating an image is also input. The data generation model 58 performs inference on the input inference data in accordance with the instruction indicated by the prompt, and outputs the inference result in a data format such as voice data and text data. Here, inference refers to, for example, analysis, classification, prediction, and / or summarization.
[0362] In the above embodiment, an example in which the specific process is performed by the data processing device 12 has been given, but the technology of the present disclosure is not limited to this, and the specific process may be performed by the smart device 14.
[0363] [Second embodiment]
[0364] FIG. 3 shows an example of the configuration of a data processing system 210 according to the second embodiment.
[0365] 3, the data processing system 210 includes the data processing device 12 and smart glasses 214. An example of the data processing device 12 is a server.
[0366] The data processing device 12 includes a computer 22, a database 24, and a communication I / F 26. The computer 22 is an example of a "computer" according to the technology of the present disclosure. The computer 22 includes a processor 28, a RAM 30, and a storage 32. The processor 28, the RAM 30, and the storage 32 are connected to a bus 34. The database 24 and the communication I / F 26 are also connected to the bus 34. The communication I / F 26 is connected to a network 54. Examples of the network 54 include a WAN (Wide Area Network) and / or a LAN (Local Area Network).
[0367] The smart glasses 214 include a computer 36, a microphone 238, a speaker 240, a camera 42, and a communication I / F 44. The computer 36 includes a processor 46, a RAM 48, and a storage 50. The processor 46, the RAM 48, and the storage 50 are connected to a bus 52. The microphone 238, the speaker 240, and the camera 42 are also connected to the bus 52.
[0368] The microphone 238 receives instructions and the like from the user 20 by receiving voice uttered by the user 20. The microphone 238 captures the voice uttered by the user 20, converts the captured voice into audio data, and outputs it to the processor 46. The speaker 240 outputs audio in accordance with instructions from the processor 46.
[0369] Camera 42 is a small digital camera equipped with an optical system including a lens, aperture, and shutter, and an imaging element such as a CMOS (Complementary Metal-Oxide-Semiconductor) image sensor or a CCD (Charge Coupled Device) image sensor, and captures images of the surroundings of user 20 (for example, an imaging range defined by an angle of view equivalent to the field of vision of a typical healthy person).
[0370] The communication I / F 44 is connected to a network 54. The communication I / Fs 44 and 26 are responsible for the exchange of various information between the processor 46 and the processor 28 via the network 54. The exchange of various information between the processor 46 and the processor 28 using the communication I / Fs 44 and 26 is carried out in a secure state.
[0371] Fig. 4 shows an example of the main functions of the data processing device 12 and the smart glasses 214. As shown in Fig. 4, in the data processing device 12, a specific process is performed by the processor 28. A specific process program 56 is stored in the storage 32.
[0372] The specific processing program 56 is an example of a "program" according to the technology of the present disclosure. The processor 28 reads the specific processing program 56 from the storage 32 and executes the read specific processing program 56 on the RAM 30. The specific processing is realized by the processor 28 operating as a specific processing unit 290 in accordance with the specific processing program 56 executed on the RAM 30.
[0373] The storage 32 stores a data generation model 58 and an emotion identification model 59. The data generation model 58 and the emotion identification model 59 are used by the identification processing unit 290.
[0374] In the smart glasses 214, the reception output process is performed by the processor 46. A reception output program 60 is stored in the storage 50. The processor 46 reads the reception output program 60 from the storage 50 and executes the read reception output program 60 on the RAM 48. The reception output process is realized by the processor 46 operating as the control unit 46A in accordance with the reception output program 60 executed on the RAM 48.
[0375] Next, a description will be given of the identification process performed by the identification processing unit 290 of the data processing device 12. In the following description, the data processing device 12 will be referred to as the "server" and the smart glasses 214 will be referred to as the "terminal."
[0376] The system of the present invention is composed of a user terminal, a speech recognition server, a generative AI analysis server, and a notification server. Below, we will explain the program processing and specific examples of this system.
[0377] System program processing
[0378] 1. User Device
[0379] The user terminal has the function to automatically start recording the voice when receiving a call. The voice is recorded in real time and encrypted for security. This encrypted voice data is immediately sent to the voice recognition server.
[0380] 2. Speech Recognition Server
[0381] The speech recognition server decrypts the encrypted speech data received from the user's device. The decrypted speech data is converted into text data using speech recognition technology. The converted text data is then sent to the generative AI analysis server.
[0382] 3. Generative AI Analysis Server
[0383] The generative AI analysis server performs detailed analysis of the text data received from the speech recognition server. Based on the text data, an algorithm is run to determine whether there is a possibility of special fraud. If it determines that there is a high possibility of fraud, the generative AI analysis server generates a warning message and sends it to the notification server.
[0384] 4. Notification Server
[0385] The notification server receives alert messages from the generative AI analysis server and, based on the user's settings, selects push notification, SMS, email, or operator contact method to immediately send the alert message to the user or designated contacts.
[0386] Specific examples
[0387] Example 1: Fraudulent call detection and warning
[0388] When a user answers a call, the user's device automatically records the call audio, encrypts it, and sends it to a speech recognition server. The speech recognition server converts the audio into text and sends it to a generative AI analysis server. The generative AI analysis server analyzes the text, and if it detects a pattern that indicates possible fraud, such as "I was in a traffic accident and need money," it generates a warning message and sends it to a notification server. The notification server then sends a push notification to the user saying, "This may be a scam. Please end the call immediately."
[0389] Example 2: Secure Call
[0390] For other calls, the user's device records the voice in the same way, and the speech recognition server converts the voice into text. The generative AI analysis server determines that the text data does not have the characteristics of a special fraud and does not generate a warning message. In this case, the user continues the call as usual.
[0391] As described above, the present invention is a system that can prevent damage from special frauds by analyzing telephone voices in real time, determining the possibility of special fraud, and immediately sending a warning to the user.
[0392] The processing flow will be explained below.
[0393] Step 1:
[0394] The user terminal detects an incoming or outgoing call, and when the user receives or makes a call, the voice recording function is automatically started.
[0395] Step 2:
[0396] The user device records the call audio in real time, and the recorded audio data is encrypted using an encryption algorithm such as AES-256 to ensure security.
[0397] Step 3:
[0398] The user device sends encrypted voice data to a speech recognition server over an internet connection, using HTTPS to ensure data integrity and security.
[0399] Step 4:
[0400] The voice recognition server decrypts the encrypted voice data received from the user terminal. If the decryption is successful, the original voice data is reproduced.
[0401] Step 5:
[0402] The speech recognition server processes the decoded speech data with a speech recognition engine and converts it into text data, thereby extracting the speech content as text information.
[0403] Step 6:
[0404] The speech recognition server sends the converted text data to the generative AI analysis server, using a secure protocol to ensure data confidentiality.
[0405] Step 7:
[0406] The generative AI analysis server analyzes the text data received from the speech recognition server, applies an algorithm to detect the characteristics of special fraud, and determines the likelihood of fraud based on the text data.
[0407] Step 8:
[0408] If the generative AI analysis server determines that there is a high possibility of fraud, it generates a warning message that indicates the possibility of fraud and provides specific instructions to the user.
[0409] Step 9:
[0410] The generative AI analysis server sends the generated warning message to the notification server, which selects the notification method according to the user's settings.
[0411] Step 10:
[0412] The notification server sends warning messages to users or designated contacts via push notification, SMS, email, or operator, allowing users to receive prompt warnings and take appropriate action.
[0413] Example 1
[0414] Next, a description will be given of Example 1. In the following description, the data processing device 12 will be referred to as a "server" and the smart glasses 214 will be referred to as a "terminal."
[0415] As the number of victims of special telephone frauds increases, there is a growing need for security systems that use advanced technology to detect fraud in real time and quickly warn users. However, current systems lack the functionality to automatically analyze the content of phone calls, detect possible fraud, and immediately warn users. Therefore, the present invention aims to solve these problems.
[0416] The specific processing by the specific processing unit 290 of the data processing device 12 in the first embodiment is realized by the following means.
[0417] In this invention, the server includes means for detecting an incoming call on the user terminal and automatically recording the call audio, means for encrypting the recorded call audio data, means for transmitting the encrypted audio data to a speech recognition server, means for the speech recognition server to decrypt the encrypted audio data and convert it into text data, means for the generative AI analysis server to analyze the text data and determine the possibility of special fraud, means for generating a warning message based on the determination, and means for transmitting the generated warning message to a notification server and notifying the user. This makes it possible to automatically detect the possibility of fraud during a call and send an appropriate warning to the user in real time.
[0418] A "user terminal" is a device that has the function of detecting an incoming call, automatically recording the call audio, and encrypting and transmitting the recorded audio data.
[0419] "Encryption" is a technology that converts recorded call voice data for security purposes and prevents unwanted access.
[0420] A "voice recognition server" is a server that has the function of receiving encrypted voice data sent from a user terminal, decrypting it, and converting the voice data into text data.
[0421] The "generative AI analysis server" is a server that analyzes text data sent from the voice recognition server and executes an algorithm to determine the possibility of special fraud.
[0422] The "notification server" is a server whose role is to receive warning messages provided by the generative AI analysis server and to send the messages to users using an appropriate notification method.
[0423] A "warning message" is a message containing a warning to the user that is generated when the generative AI analysis server detects the possibility of fraud.
[0424] "Push notifications" are short messages sent directly to a user's device, providing an instant way to communicate important information from applications and services.
[0425] "Short Message Service (SMS)" is a communications protocol for sending and receiving short text messages using mobile devices.
[0426] "Email" is a method of sending and receiving text-based messages over the Internet and is a widely used means of communicating information to users.
[0427] "Communication means" refers to any method for transmitting information from a sender to a receiver, and in this system includes push notifications, short message services, emails, etc.
[0428] The system of the present invention consists of a user terminal, a speech recognition server, a generative AI analysis server, and a notification server. The user terminal automatically records the voice of the call, encrypts the data, and sends it to the speech recognition server. The speech recognition server decrypts the recorded data and converts it into text using speech recognition technology. The converted text data is sent to the generative AI analysis server, where it is analyzed to determine the possibility of fraud. If it is determined that there is a high possibility of fraud, a warning message is sent to the user via the notification server.
[0429] The user device can be a smartphone or tablet. When the user answers a call, the device automatically starts recording the call audio. The recorded data is AES encrypted in real time and sent to the speech recognition server via an HTTP POST request.
[0430] The speech recognition server runs a Python-based Flask web server. The received encrypted voice data is decrypted using Python's Cryptography library. The voice data is then converted into text data using a service such as the Google Cloud Speech-to-Text API. This text data is then sent to the generative AI analysis server.
[0431] The generative AI analysis server runs a server built using, for example, Node.js or Python. It analyzes the received text data using a generative AI model, such as OpenAI's GPT-3 model, to detect specific keywords and contexts that indicate the possibility of fraud. For example, it identifies patterns such as "I was in a traffic accident and need money." If it determines that there is a high possibility of fraud, it generates a warning message and sends it to the notification server.
[0432] The notification server runs a server built using, for example, Ruby on Rails or Django. It notifies users of received alert messages using methods such as push notifications, SMS, and email. For example, you can send SMS using the Twilio API or push notifications using Firebase Cloud Messaging.
[0433] Specific examples
[0434] Example 1: Fraudulent call detection and warning
[0435] When a user answers a call, the user's device automatically records the call audio, encrypts it, and sends it to a speech recognition server. The speech recognition server converts the audio into text data and sends the text data to a generative AI analysis server. The generative AI analysis server analyzes the text and detects patterns that indicate possible fraud, such as "I was in a traffic accident and need money." In this case, it generates a fraud warning message and sends it to a notification server. The notification server then sends a push notification to the user saying, "This may be a scam. Please end the call immediately."
[0436] Example 2: Secure Call
[0437] For other calls, the user's device records the voice in the same way, and the speech recognition server converts the voice into text data. The generative AI analysis server analyzes the text data and determines that it does not have the characteristics of a special fraud. In this case, no warning message is generated, and the user can continue the call as usual.
[0438] Example prompt: "Is this a potential scam? 'I've been in a car accident and need money. Please transfer the money.'"
[0439] The flow of the identification process in the first embodiment will be described with reference to FIG.
[0440] Step 1:
[0441] The user device detects an incoming call and automatically starts recording the call audio. The input is the incoming call event. The output is the recorded call audio data. Specifically, the device software catches the incoming call event and starts the recording module. For example, the smartphone's OS notifies the calling app of the incoming call, and an application written in Java starts recording using the Android MediaRecorder API.
[0442] Step 2:
[0443] The recorded call audio data is encrypted in real time. The input is the audio data obtained in step 1. The output is the encrypted audio data. Specifically, the device's encryption module encrypts the recorded data using an encryption algorithm such as AES. For example, an application written in Java encrypts the data using the AES encryption algorithm.
[0444] Step 3:
[0445] The encrypted voice data is sent to the voice recognition server. The input is the encrypted voice data obtained in step 2. The output is a notification that the data has been sent to the server. Specifically, the device sends the voice data to the server using an HTTP POST request. For example, an application written in Java uses an HTTP POST request.
[0446] Step 4:
[0447] The speech recognition server receives and decrypts the encrypted audio data. The input is the encrypted audio data sent in step 3. The output is the decrypted audio data. Specifically, the server's encryption / decryption module decrypts the data using the AES decryption algorithm. For example, a Flask web server written in Python processes the request and decrypts it using the Cryptography library.
[0448] Step 5:
[0449] The speech recognition server converts the decoded speech data into text data. The input is the decoded speech data obtained in step 4. The output is text data. Specifically, speech-to-text conversion is performed using a speech recognition API. For example, the Google Cloud Speech-to-Text API is used to convert speech data into text.
[0450] Step 6:
[0451] The converted text data is sent to the generative AI analysis server. The input is the text data obtained in step 5. The output is a notification that data transmission to the generative AI analysis server has been completed. Specifically, the speech recognition server sends the data using an HTTP POST request. For example, the Python code sends a POST request to the generative AI analysis server.
[0452] Step 7:
[0453] The generative AI analysis server analyzes the text data and determines the likelihood of fraud. The input is the text data received in step 6. The output is the fraud determination result. Specifically, it uses a generative AI model to analyze the text and detect keywords and contexts that indicate the likelihood of fraud. For example, it uses OpenAI's GPT-3 model to identify phrases such as "I was in a traffic accident and need money."
[0454] Step 8:
[0455] If it is determined that there is a high possibility of fraud, a warning message is generated. The input is the fraud determination result from step 7. The output is a warning message. Specifically, the server runs the message generation module to create a warning message. For example, it creates a message saying, "There is a possibility of fraud. Please end the call immediately."
[0456] Step 9:
[0457] The generated warning message is sent to the notification server. The input is the warning message from step 8. The output is a notification of completion of data transmission to the notification server. Specifically, the generative AI analysis server sends the data using an HTTP POST request. For example, JavaScript code sends a POST request to the notification server.
[0458] Step 10:
[0459] The notification server sends an alert message to the user. The input is the alert message received in step 9. The output is an alert notification to the user. Specifically, the notification server sends the message to the user using the appropriate notification method. For example, it can send an SMS using the Twilio API or a push notification using Firebase Cloud Messaging.
[0460] (Application example 1)
[0461] Next, a description will be given of Application Example 1. In the following description, the data processing device 12 will be referred to as a "server" and the smart glasses 214 will be referred to as a "terminal."
[0462] In recent years, special fraud cases using telephones have been increasing, and many people have become victims. For this reason, there is a need for a system that can detect possible fraud in real time when a call is received and issue a warning. However, existing systems have problems such as low accuracy in fraud detection, lack of real-time response, and a lack of diverse means of notification to users. To address this issue, a system is needed that can detect possible fraud more accurately and quickly, and send warnings to users using a variety of means.
[0463] The specific processing by the specific processing unit 290 of the data processing device 12 in the application example 1 is realized by the following means.
[0464] In this invention, the server includes means for recording telephone voice, means for converting the recorded telephone voice into text data, means for analyzing the converted text data and determining the possibility of special fraud, means for generating a warning message based on the determination, means for transmitting the generated warning message, means for analyzing the text data using a generative AI model, means for starting recording in real time on a user terminal, and means for generating and transmitting a warning message by push notification, SMS, or email. This makes it possible to quickly detect the possibility of fraud with high accuracy and to send warnings to users by various means.
[0465] "Telephone voice" refers to voice data transmitted through telephone communication.
[0466] "Recording means" refers to devices or software for saving voice data, and has the function of saving the contents of telephone conversations as files.
[0467] "Means for converting into text data" refers to devices or software that use voice recognition technology to convert voice data into character-based data.
[0468] "Means for analysis" refers to devices or software that have the ability to analyze text data and evaluate its content based on pattern recognition or algorithms.
[0469] "Special fraud" refers to criminal acts that involve deceiving people over the phone and stealing money from them.
[0470] "Means for determining" refers to a device or software that runs an algorithm to determine whether or not a transaction is fraudulent based on the analysis results.
[0471] A "warning message" refers to a message containing information that is notified to the user when a possible fraud is detected.
[0472] "Generating means" refers to a device or software for creating a warning message, and has the function of automatically creating the necessary wording and warning content.
[0473] "Transmitting means" refers to the communication technology and network infrastructure for transferring the generated alert message to a user terminal or other device.
[0474] "Generative AI models" refer to algorithms and tools that use artificial intelligence to generate and analyze data.
[0475] A "prompt" is an instruction or guidance text to be input into a generative AI model, and is text used to obtain a specific analysis result or product.
[0476] "User terminal" refers to a communication device that is directly operated by a user, such as a smartphone, tablet, or computer.
[0477] "Means for starting recording in real time" refers to devices or software that have the function of automatically starting to record audio the moment a call is initiated.
[0478] A "push notification" refers to a notification message sent to a user device in real time from an application or system.
[0479] "SMS" stands for Short Message Service, a service for sending short text messages over mobile phone networks.
[0480] "Mail" is an abbreviation for electronic mail, a means of communication for sending and receiving text and files over the Internet.
[0481] The present invention is a security system that detects potential fraud in real time for telephone calls received by a user and issues a prompt warning. The system includes the following main components:
[0482] User terminal
[0483] The user terminal is equipped with a function for recording telephone calls in real time. The voice is recorded as soon as the call begins, and the voice data is encrypted to ensure security. The encrypted voice data is immediately sent to a voice recognition server. This allows fraud detection to be performed automatically without the user having to perform any special operations during the call.
[0484] Speech Recognition Server
[0485] The speech recognition server decrypts the encrypted voice data received from the user's device. The decrypted voice data is converted into text data using speech recognition technology, such as Google Speech Recognition. The converted text data is then sent to the generative AI analysis server.
[0486] Generative AI analysis server
[0487] The generative AI analysis server analyzes the received text data and runs algorithms to determine the likelihood of fraud. This analysis uses generative AI models such as BERT and GPT. If a high likelihood of fraud is determined, the generative AI analysis server generates a warning message and sends it to the notification server.
[0488] Specific prompt examples:
[0489] Analyze this text and determine if it is potentially fraudulent if it contains phrases like "I was in a car accident" or "I need money."
[0490] Text: "I was in a car accident and need money now. Please transfer it to my bank account."
[0491] Notification Server
[0492] The notification server notifies the user device of the warning message received from the generative AI analysis server via push notification, SMS, email, etc. This allows the user to respond immediately to potentially fraudulent calls.
[0493] Specific examples
[0494] When a user receives a phone call that is suspected to be fraudulent, the user's device automatically records the audio and sends the data to a speech recognition server. The speech recognition server converts the audio data into text data, which is then analyzed by a generative AI analysis server. If the analysis determines that the call is likely to be fraudulent, the notification server sends a warning message to the user, informing them that "This may be a fraudulent call. Please end the call immediately."
[0495] This allows users to prevent themselves from falling victim to fraud. By combining multiple cutting-edge technologies, the system is able to detect potential fraud with high accuracy and speed, and send warnings via a variety of means.
[0496] The flow of the specific processing in the application example 1 will be described with reference to FIG.
[0497] Step 1:
[0498] When the user terminal detects the start of a call, it automatically records the voice in real time. The recorded voice data is encrypted to ensure security. For encryption, for example, the Fernet encryption library is used. The encrypted voice data is then sent to a voice recognition server.
[0499] Input: Phone call audio
[0500] Data processing: Encrypting audio
[0501] Output: Encrypted audio data
[0502] Step 2:
[0503] The voice recognition server decrypts the received encrypted voice data. For decryption, the same encryption key as that used on the user's device is used. The decrypted voice data is converted into text data using voice recognition technology. For example, the Google Speech Recognition API is used. The converted text data is sent to the generative AI analysis server.
[0504] Input: Encrypted audio data
[0505] Data operations: Decryption of encrypted data and conversion of voice data to text
[0506] Output: Text data
[0507] Step 3:
[0508] The generative AI analysis server uses a generative AI model (e.g., BERT or GPT-3) to analyze the received text data. It uses prompt sentences to analyze the text data and determine the likelihood of fraud. Specific examples of prompt sentences are as follows:
[0509] Analyze this text and determine if it is potentially fraudulent if it contains phrases like "I was in a car accident" or "I need money."
[0510] Text: "I was in a car accident and need money now. Please transfer it to my bank account."
[0511] Input: Text data, prompt
[0512] Data Computation: Text Data Analysis with Generative AI Models
[0513] Output: Possibility of fraud determination result
[0514] Step 4:
[0515] If the generative AI analysis server determines that there is a high possibility of fraud, it generates a warning message, such as "This is a possible fraud. Please end the call immediately." This message is sent to the notification server.
[0516] Input: Possibility of fraud determination result
[0517] Data processing: Generate warning messages
[0518] Output: Warning message
[0519] Step 5:
[0520] The notification server notifies the user device of the warning message received from the generative AI analysis server. The warning message is sent to the user immediately using multiple means, such as push notification, SMS, and email, allowing the user to quickly respond to potentially fraudulent calls.
[0521] Input: warning message
[0522] Data processing: Sending notification messages
[0523] Output: Warning message sent to the user
[0524] The above are the specific processing steps of the program for the system that realizes the application example.
[0525] Furthermore, an emotion engine that estimates the user's emotion may be further combined. That is, the identification processing unit 290 may estimate the user's emotion using the emotion identification model 59, and perform identification processing using the user's emotion.
[0526] The system of the present invention is composed of a user terminal, a speech recognition server, a generative AI analysis server, a notification server, and an emotion engine. Below, we will explain the program processing and specific examples of this system.
[0527] System program processing
[0528] 1. User Device
[0529] The user terminal has the function to automatically start recording the voice when receiving a call. The voice is recorded in real time and encrypted for security. This encrypted voice data is immediately sent to the voice recognition server.
[0530] 2. Speech Recognition Server
[0531] The speech recognition server decrypts the encrypted speech data received from the user's device. The decrypted speech data is converted into text data using speech recognition technology. The converted text data is then sent to the generative AI analysis server.
[0532] 3. Generative AI Analysis Server
[0533] The generative AI analysis server performs detailed analysis of the text data received from the speech recognition server. Based on the text data, an algorithm is run to determine whether or not there is a possibility of special fraud.
[0534] 4. Emotion Engine
[0535] The emotion engine analyzes the voice data acquired from the user's device to recognize the user's emotional state. The results of the emotion engine are sent to the generative AI analysis server, which influences the determination of the possibility of fraud.
[0536] 5. Correction of analysis results
[0537] The generative AI analysis server receives the analysis results from the emotion engine and amends them based on the user's emotional state. For example, if the user is nervous or confused, the analysis results are strengthened.
[0538] 6. Generating and Sending Warning Messages
[0539] The generative AI analysis server generates warning messages as needed based on the corrected analysis results. The generated warning messages are sent to the notification server, which then sends the warning messages via push notification, SMS, email, or operator notification based on the user's settings.
[0540] Specific examples
[0541] Example 1: Fraudulent call detection and warning
[0542] When a user answers a call, the user's device automatically records the call audio, encrypts it, and sends it to a speech recognition server. The speech recognition server converts the speech into text and sends it to a generative AI analysis server. The generative AI analysis server analyzes the text and detects patterns that indicate the possibility of fraud, such as "I was in a traffic accident and need money." If the emotion engine detects a state of tension in the user's voice, it further determines that the possibility of fraud is high. A warning message is generated, and a push notification is sent to the user via the notification server, stating, "This may be a scam. Please end the call immediately."
[0543] Example 2: Secure Call
[0544] The user receives another call. In this case, the user's device again records the audio, and the speech recognition server converts it into text. The generative AI analysis server analyzes the text data and verifies that it does not have the characteristics of a specialized fraud. The emotion engine also detects that the user is relaxed. In this case, no warning message is generated, and the user continues the call as usual.
[0545] As described above, the present invention is a system that can prevent damage from special frauds by analyzing telephone voices in real time, taking into account the user's emotional state, determining the possibility of special fraud, and sending a warning.
[0546] The processing flow will be explained below.
[0547] Step 1:
[0548] The user receives or makes a call. The user device detects this and automatically starts recording the audio.
[0549] Step 2:
[0550] The user device records the call audio in real time, and the recorded audio data is encrypted using an encryption algorithm such as AES-256.
[0551] Step 3:
[0552] The user terminal sends the encrypted voice data to the voice recognition server using the HTTPS protocol.
[0553] Step 4:
[0554] The voice recognition server decrypts the encrypted voice data received from the user terminal. If the decryption is successful, the voice data returns to its original state.
[0555] Step 5:
[0556] The speech recognition server converts the decoded speech data into text data using speech recognition technology, and the converted text data is sent to the generative AI analysis server.
[0557] Step 6:
[0558] The generative AI analysis server analyzes the text data received from the speech recognition server, and executes a specific algorithm to determine the possibility of special fraud.
[0559] Step 7:
[0560] The user device sends the recorded voice data to the emotion engine, which analyzes the user's emotional state (e.g., tension, confusion, calmness, etc.) and sends the results to the generative AI analysis server.
[0561] Step 8:
[0562] The generative AI analysis server receives the analysis results from the emotion engine and corrects the analysis results for the possibility of special fraud. If the user is nervous or confused, the possibility of fraud is strengthened.
[0563] Step 9:
[0564] Based on the final judgment based on the results of the emotion engine, the generative AI analysis server generates a warning message, which includes information about the suspected fraud and specific instructions for the user.
[0565] Step 10:
[0566] The generative AI analysis server sends the generated warning message to the notification server, which selects the method of notification based on the user's settings: push notification, SMS, email, or operator notification.
[0567] Step 11:
[0568] The notification server sends alert messages to the user or designated contacts via the method of their choice, allowing the user to receive prompt warnings and take appropriate action.
[0569] Example 2
[0570] Next, a description will be given of Example 2. In the following description, the data processing device 12 will be referred to as a "server" and the smart glasses 214 will be referred to as a "terminal."
[0571] In modern times, special telephone frauds have become a social problem. Fraudulent phone calls targeting the elderly in particular have caused many victims, making countermeasures an urgent necessity. However, conventional security systems lack the functionality to detect potential fraud in real time and issue immediate warnings to users. Furthermore, they lack the functionality to correct analysis results by taking the user's emotional state into account, making them prone to false positives and overreactions. The purpose of this invention is to solve these problems and provide a highly reliable fraud prevention system.
[0572] The identification process by the identification processing unit 290 of the data processing device 12 in the second embodiment is realized by the following means. In this invention, the server includes a means for recording telephone voice, a means for encrypting and transmitting the recorded telephone voice, a means for converting the encrypted telephone voice into text data, a means for analyzing the converted text data and determining the possibility of special fraud, a means for analyzing the user's emotional state and correcting the analysis result, a means for generating a warning message based on the corrected result, and a means for transmitting the generated warning message. This enables real-time detection of the possibility of fraud and correction of the analysis result taking the user's emotional state into consideration.
[0573] "Telephone voice" refers to audio data recorded from a telephone conversation.
[0574] A "recording means" is a device or system that stores the audio during a call as digital data.
[0575] "Encryption and transmission means" refers to the methods and techniques by which stored voice data is encrypted for security purposes and transmitted to the appropriate recipient.
[0576] "Means for converting into text data" refers to technology for converting voice data into character data, such as voice recognition technology.
[0577] "Means for analyzing and determining the possibility of special fraud" refers to algorithms and systems that analyze text data and detect signs of fraud from its content.
[0578] "Means for analyzing emotional state" refers to technology or a system for determining emotions (tension, relaxation, etc.) from the user's voice data.
[0579] "Means for correcting the analysis results" refers to a method for correcting the results as necessary to increase the reliability of the analysis results based on information on emotional state.
[0580] A "means for generating a warning message" is a system that generates a warning message to a user when fraud is likely.
[0581] The "means for sending a warning message" refers to a method for sending the generated warning message to the user using a communication means.
[0582] The system of the present invention is composed of a user terminal, a speech recognition server, a generative AI analysis server, an emotion engine, and a notification server. A specific example of the system will be described below.
[0583] User terminal
[0584] The user device has a function that automatically starts recording voice when a call comes in. The recorded voice data is encrypted in real time using the AES method or similar. This encrypted voice data is immediately sent to the voice recognition server using the HTTPS protocol.
[0585] Speech Recognition Server
[0586] The speech recognition server decrypts the encrypted speech data received from the user's device. This decrypted speech data is converted into text data using speech recognition technology. Specifically, it uses a commonly used speech recognition API, such as the speech recognition function of a cloud service. This converted text data is then sent to the generative AI analysis server.
[0587] Generative AI analysis server
[0588] The generative AI analytics server uses machine learning algorithms to analyze the text data received from the speech recognition server. For example, it uses open-source machine learning libraries or cloud-based AI services to analyze the text data and determine the likelihood of fraud. If patterns indicative of possible fraud are detected, the information is sent to the emotion engine.
[0589] Emotion Engine
[0590] The emotion engine analyzes the user's emotional state based on their voice data. The emotional state is determined using commonly used emotion analysis APIs and software. This information is sent to a generative AI analysis server, which corrects the analysis results. For example, if the user is nervous, it is deemed to be a sign of a high possibility of fraud.
[0591] Correcting analysis results and generating warning messages
[0592] Once the analysis results have been corrected, the generative AI analysis server generates a warning message as needed. This warning message is sent to the user's device in real time. Specifically, the warning message is sent to the notification server and delivered to the user via push notification, SMS, email, or other means. For example, the message might say, "This may be a scam. Please end the call immediately."
[0593] Specific examples
[0594] Example 1: Fraudulent call detection
[0595] When a user answers a call, the user device automatically records the call audio and sends the encrypted audio data to a speech recognition server. The speech recognition server converts the audio into text and sends it to a generative AI analysis server. The generative AI analysis server analyzes the text data and detects signs of fraud, such as "I was in a traffic accident and need money." If the emotion engine detects a state of tension in the user's voice, the possibility of fraud increases. Based on this, a warning message is generated and sent to the user via a notification server.
[0596] Example 2: Secure Call
[0597] The user receives another call and the device begins recording the audio. This audio data is also encrypted and sent to the speech recognition server. The recorded audio data is converted into text and sent to the generative AI analysis server. The generative AI analysis server analyzes this text data and verifies that it does not contain any characteristics of specialized fraud. The emotion engine also detects that the user is relaxed, so no warning message is generated. This allows the user to continue the safe call.
[0598] Prompt Sentence Examples
[0599] "The audio of phone calls received by users is recorded in real time and encrypted for security. This encrypted audio data is converted into text data using speech recognition technology and analyzed by a generative AI analysis server. Based on the analyzed text data, the possibility of fraud is determined and a warning message is generated if necessary. The warning message is sent to the user via a notification server to ensure a safe call. If signs of fraud are detected, how should it be handled?"
[0600] As a result, the present invention is a system that can prevent damage from special fraud by analyzing telephone voices in real time, determining the possibility of special fraud taking into account the user's emotional state, and sending a warning.
[0601] The flow of the identification process in the second embodiment will be described with reference to FIG.
[0602] Processing steps of this system's program
[0603] Step 1
[0604] When a user receives a call, the user terminal automatically records the call audio. This audio recording is initiated using a program that is automatically triggered when the user receives a call. The input is the real-time call audio, and the output is the recorded audio file. Specifically, the user terminal launches a recording application and saves the call audio as a file.
[0605] Step 2
[0606] The user terminal encrypts the recorded voice data using the AES encryption method. At this point, the input is the recorded voice file, and the output is the encrypted voice data. Specifically, the user terminal runs an encryption algorithm to convert the voice data into a secure format.
[0607] Step 3
[0608] The user terminal sends encrypted voice data to the voice recognition server. The HTTPS protocol is used for transmission. The input is encrypted voice data, and the output is the encrypted voice data received by the voice recognition server. Specifically, the user terminal creates an HTTP request and sends the data.
[0609] Step 4
[0610] The speech recognition server decrypts the received encrypted audio data. The input is the encrypted audio data, and the output is the decrypted audio data. Specifically, the server runs the AES decryption algorithm to reconstruct the original audio data.
[0611] Step 5
[0612] The speech recognition server converts the decoded speech data into text data. Specifically, it uses a speech recognition API to convert speech to text. The input is the decoded speech data, and the output is text data. Specifically, the speech recognition server calls the speech recognition API (for example, a cloud-based speech recognition service) and obtains the returned text data.
[0613] Step 6
[0614] The speech recognition server sends the generated text data to the generative AI analysis server. The input is the generated text data, and the output is the text data received by the generative AI analysis server. Specifically, the data is sent via a REST API for server-to-server communication.
[0615] Step 7
[0616] The generative AI analysis server analyzes the received text data and runs machine learning algorithms to determine the likelihood of fraud. The input is text data, and the output is a result indicating the likelihood of fraud. Specifically, the generative AI analysis server uses libraries such as TensorFlow and Hugging Face's Transformers to analyze the text data and detect fraudulent patterns.
[0617] Step 8
[0618] The emotion engine analyzes the user's emotional state based on their voice data. The input is the user's voice data, and the output is data about the user's emotional state (tension, confusion, etc.). Specifically, the emotion engine calls an emotion analysis API (e.g., IBM Watson Tone Analyzer) to determine the emotional state.
[0619] Step 9
[0620] The generative AI analysis server corrects the analysis results based on the emotional state data received from the emotion engine. The input is the emotional state data and text analysis results, and the output is the corrected analysis results. Specifically, it reevaluates the possibility of fraud taking the emotional state into account and makes a final judgment.
[0621] Step 10
[0622] The generative AI analysis server generates a warning message based on the corrected analysis results. The input is the corrected analysis results, and the output is the generated warning message. Specifically, the server creates a warning message based on the judgment results.
[0623] Step 11
[0624] The notification server sends the generated warning message to the user device. Methods of sending include push notification, SMS, and email. The input is the generated warning message, and the output is the warning message received by the user device. Specifically, the notification server uses AWS SNS (Simple Notification Service) or a similar notification service to send the message to the user.
[0625] The above are the specific processing steps in this system.
[0626] (Application example 2)
[0627] Next, a description will be given of Application Example 2. In the following description, the data processing device 12 will be referred to as a "server" and the smart glasses 214 will be referred to as a "terminal."
[0628] In recent years, there has been an increase in special frauds using telephones, and frauds targeting the elderly in particular have become a serious problem. Current security systems lack the means to accurately detect potential fraud in real time and send prompt and appropriate warnings to users. As a result, many users remain at high risk of becoming victims of fraud. Furthermore, few systems take the user's emotional state into account, and the accuracy of fraud detection is low, making it difficult for users to feel safe answering the phone. There is a need for a system that solves this problem.
[0629] The identification process by the identification processing unit 290 of the data processing device 12 in Application Example 2 is realized by the following means. In this invention, the server includes a means for recording telephone voice, a means for encrypting and transmitting the recorded telephone voice, and a means for decrypting the encrypted voice data. This enables secure and highly accurate analysis of telephone voice in real time and fraud detection that takes into account the user's emotional state. In addition, by including a means for quickly sending a warning message via push notification, SMS, email, or operator, it is possible to immediately warn the user of fraud and prevent damage before it occurs.
[0630] "Means for recording telephone voice" refers to technology that automatically records and safely stores the contents of a call when a user initiates a call.
[0631] "Means for encrypting and transmitting recorded telephone voice data" refers to a technology that encrypts recorded telephone voice data to protect it from unauthorized access or eavesdropping by third parties and transmits it to a designated server.
[0632] "Means for decrypting encrypted audio data" refers to a technique for restoring the transmitted encrypted audio data to the original audio data.
[0633] "Means for converting voice data into text data" refers to a technology that analyzes decoded voice data and converts its contents into text information.
[0634] The "means of analyzing the converted text data and determining the possibility of special fraud" is a technology that uses an algorithm to detect signs and patterns of special fraud based on the obtained text data.
[0635] "Means for analyzing the user's emotional state" refers to a technology that analyzes the user's emotions (tension, confusion, fear, etc.) from voice data during a call and determines that state.
[0636] The "means for correcting the determination of the possibility of special fraud based on the analysis results of the emotional state" is a technology for improving the accuracy of the determination of the possibility of special fraud by taking into account the emotional state of the user.
[0637] The "means for generating a warning message" is a technology that automatically generates a notification message to warn the user when it is determined that there is a high possibility of special fraud.
[0638] "Means for sending the generated warning message" refers to a communication technology for immediately delivering the generated warning message to the user, including push notification, SMS, email, or operator notification.
[0639] The purpose of the system of the present invention is to prevent special frauds using voice calls, and it records, encrypts, analyzes telephone voice, and generates and transmits warning messages. The specific system configuration and its program processing are described below.
[0640] System Configuration
[0641] The system of the present invention comprises the following components:
[0642] 1. User device: When a call is initiated, the audio is automatically recorded, and the recorded audio data is encrypted and sent to the server.
[0643] 2. Speech recognition server: Decrypts the encrypted voice data and converts the voice into text data.
[0644] 3. Generative AI analysis server: Analyzes text data to determine the possibility of special fraud. It also analyzes the user's emotional state and adjusts the judgment results based on the results.
[0645] 4. Notification Server: Generates warning messages as needed and sends them to users via push notification, SMS, email or operator.
[0646] Hardware and Software Used
[0647] Smartphone: iOS / Android devices
[0648] Cloud servers: Amazon Web Services (AWS) EC2, Amazon S3
[0649] Speech recognition technology: Google Speech-to-Text API
[0650] Generative AI model: OpenAI GPT-4 API
[0651] Sentiment analysis engine: IBM Watson Tone Analyzer
[0652] Notification service: Firebase Cloud Messaging (FCM)
[0653] Program Processing Overview
[0654] 1. User Device
[0655] The user device automatically records the voice as soon as the call begins and encrypts the recording using a sophisticated encryption algorithm. The encrypted voice data is then immediately sent to a voice recognition server.
[0656] 2. Speech Recognition Server
[0657] The speech recognition server decrypts the received encrypted voice data and converts it into text using the Google Speech-to-Text API. The converted text data is then sent to the generative AI analysis server.
[0658] 3. Generative AI Analysis Server
[0659] The generative AI analysis server uses the OpenAI GPT-4 API to perform detailed analysis of the text data sent from the speech recognition server. To determine whether there are signs of fraud, it detects specific keywords and phrases and then uses IBM Watson Tone Analyzer to analyze the user's emotional state. Based on the results, it corrects the judgment results as necessary.
[0660] 4. Notification Server
[0661] The notification server receives the results of the generative AI analysis server and generates a warning message if it determines that the call is likely to be fraudulent, and sends it to the user via push notification, SMS, or email via Firebase Cloud Messaging (FCM). This warning message urges the user to end the call.
[0662] Examples and prompts
[0663] Examples:
[0664] For example, if a user receives a call while on a call saying, "My son was in a traffic accident and I need money immediately," the system will analyze the call based on the following prompt:
[0665] Example prompt sentence:
[0666] The user received a call. The text of the call is below:
[0667] "Hello, this is the police. Your son has been in a car accident. We need money."
[0668] The user's emotional state is "tense."
[0669] Determine if this is likely a scam based on the following criteria, and if so, generate a warning message.
[0670] conditions:
[0671] Keywords to detect possible fraud: traffic accident, need money, transfer, etc.
[0672] User's emotional state: nervous, confused, scared
[0673] Output format:
[0674] Possibility of fraud: [Yes / No]
[0675] Warning message: [Possible scam. Please end the call immediately / This call is safe.]
[0676] This system makes it possible to detect the risk of special fraud in real time and quickly protect users.
[0677] The flow of the specific processing in the application example 2 will be described with reference to FIG.
[0678] Step 1:
[0679] When a user receives a call, the user terminal automatically starts recording the call audio. The recorded audio data is encrypted using an encryption algorithm. The input is the call audio data, and the output is the encrypted audio data.
[0680] Step 2:
[0681] The encrypted voice data is sent from the user terminal to the voice recognition server, which then decrypts the received data. The input is the encrypted voice data, and the output is the decrypted voice data.
[0682] Step 3:
[0683] The speech recognition server converts the decoded speech data into text data using the Google Speech-to-Text API. The input is speech data and the output is text data.
[0684] Step 4:
[0685] The generative AI analysis server receives the text data sent from the speech recognition server and uses the OpenAI GPT-4 API to analyze the possibility of fraud by detecting keywords and phrases within the text data. The input is the text data, and the output is the analysis results.
[0686] Step 5:
[0687] The generative AI analysis server analyzes the user's emotional state using the IBM Watson Tone Analyzer, an emotion analysis engine, in addition to analyzing the text data. The input is voice data, and the output is the user's emotional state.
[0688] Step 6:
[0689] The generative AI analysis server corrects the judgment of the possibility of special fraud based on the emotion analysis results. For example, if the user is nervous, it will judge that there is a high possibility of fraud. The input is the analysis result and emotional state, and the output is the corrected judgment result.
[0690] Step 7:
[0691] Based on the corrected judgment results, the generative AI analysis server generates warning messages as necessary. The input is the corrected judgment results, and the output is the warning message.
[0692] Step 8:
[0693] The notification server receives the generated alert messages and uses Firebase Cloud Messaging (FCM) to send them to the user via push notification, SMS, email, or operator. The input is the alert message and the output is the notification to the user.
[0694] The specific processing unit 290 transmits the result of the specific processing to the smart glasses 214. In the smart glasses 214, the control unit 46A causes the speaker 240 to output the result of the specific processing. The microphone 238 acquires audio indicating a user input regarding the result of the specific processing. The control unit 46A transmits audio data indicating the user input acquired by the microphone 238 to the data processing device 12. In the data processing device 12, the specific processing unit 290 acquires the audio data.
[0695] The data generation model 58 is a so-called generative AI (Artificial Intelligence). An example of the data generation model 58 is ChatGPT (Internet Search<URL: https: / / openai.com / blog / chatgpt> ), Gemini (Internet search <url: https: gemini.google.com ?hl="ja">) and other generation AIs. The data generation model 58 is obtained by performing deep learning on a neural network. A prompt including an instruction is input to the data generation model 58, and inference data such as voice data indicating voice, text data indicating text, and image data indicating an image is also input. The data generation model 58 performs inference on the input inference data in accordance with the instruction indicated by the prompt, and outputs the inference result in a data format such as voice data and text data. Here, inference refers to, for example, analysis, classification, prediction, and / or summarization.
[0696] In the above embodiment, an example in which the specific processing is performed by the data processing device 12 has been given, but the technology of the present disclosure is not limited to this, and the specific processing may be performed by the smart glasses 214.
[0697] [Third embodiment]
[0698] FIG. 5 shows an example of the configuration of a data processing system 310 according to the third embodiment.
[0699] 5, the data processing system 310 includes the data processing device 12 and a headset terminal 314. An example of the data processing device 12 is a server.
[0700] The data processing device 12 includes a computer 22, a database 24, and a communication I / F 26. The computer 22 is an example of a "computer" according to the technology of the present disclosure. The computer 22 includes a processor 28, a RAM 30, and a storage 32. The processor 28, the RAM 30, and the storage 32 are connected to a bus 34. The database 24 and the communication I / F 26 are also connected to the bus 34. The communication I / F 26 is connected to a network 54. Examples of the network 54 include a WAN (Wide Area Network) and / or a LAN (Local Area Network).
[0701] The headset type terminal 314 includes a computer 36, a microphone 238, a speaker 240, a camera 42, a communication I / F 44, and a display 343. The computer 36 includes a processor 46, a RAM 48, and a storage 50. The processor 46, the RAM 48, and the storage 50 are connected to a bus 52. The microphone 238, the speaker 240, the camera 42, and the display 343 are also connected to the bus 52.
[0702] The microphone 238 receives instructions and the like from the user 20 by receiving voice uttered by the user 20. The microphone 238 captures the voice uttered by the user 20, converts the captured voice into audio data, and outputs it to the processor 46. The speaker 240 outputs audio in accordance with instructions from the processor 46.
[0703] Camera 42 is a small digital camera equipped with an optical system including a lens, aperture, and shutter, and an imaging element such as a CMOS (Complementary Metal-Oxide-Semiconductor) image sensor or a CCD (Charge Coupled Device) image sensor, and captures images of the surroundings of user 20 (for example, an imaging range defined by an angle of view equivalent to the field of vision of a typical healthy person).
[0704] The communication I / F 44 is connected to a network 54. The communication I / Fs 44 and 26 are responsible for the exchange of various information between the processor 46 and the processor 28 via the network 54. The exchange of various information between the processor 46 and the processor 28 using the communication I / Fs 44 and 26 is carried out in a secure state.
[0705] Fig. 6 shows an example of the main functions of the data processing device 12 and the headset type terminal 314. As shown in Fig. 6, in the data processing device 12, a specific process is performed by the processor 28. A specific process program 56 is stored in the storage 32.
[0706] The specific processing program 56 is an example of a "program" according to the technology of the present disclosure. The processor 28 reads the specific processing program 56 from the storage 32 and executes the read specific processing program 56 on the RAM 30. The specific processing is realized by the processor 28 operating as a specific processing unit 290 in accordance with the specific processing program 56 executed on the RAM 30.
[0707] The storage 32 stores a data generation model 58 and an emotion identification model 59. The data generation model 58 and the emotion identification model 59 are used by the identification processing unit 290.
[0708] In the headset type terminal 314, a reception output process is performed by the processor 46. A reception output program 60 is stored in the storage 50. The processor 46 reads the reception output program 60 from the storage 50 and executes the read reception output program 60 on the RAM 48. The reception output process is realized by the processor 46 operating as the control unit 46A in accordance with the reception output program 60 executed on the RAM 48.
[0709] Next, a description will be given of the identification process performed by the identification processing unit 290 of the data processing device 12. In the following description, the data processing device 12 will be referred to as the "server" and the headset type terminal 314 will be referred to as the "terminal."
[0710] The system of the present invention is composed of a user terminal, a speech recognition server, a generative AI analysis server, and a notification server. Below, we will explain the program processing and specific examples of this system.
[0711] System program processing
[0712] 1. User Device
[0713] The user terminal has the function to automatically start recording the voice when receiving a call. The voice is recorded in real time and encrypted for security. This encrypted voice data is immediately sent to the voice recognition server.
[0714] 2. Speech Recognition Server
[0715] The speech recognition server decrypts the encrypted speech data received from the user's device. The decrypted speech data is converted into text data using speech recognition technology. The converted text data is then sent to the generative AI analysis server.
[0716] 3. Generative AI Analysis Server
[0717] The generative AI analysis server performs detailed analysis of the text data received from the speech recognition server. Based on the text data, an algorithm is run to determine whether there is a possibility of special fraud. If it determines that there is a high possibility of fraud, the generative AI analysis server generates a warning message and sends it to the notification server.
[0718] 4. Notification Server
[0719] The notification server receives alert messages from the generative AI analysis server and, based on the user's settings, selects push notification, SMS, email, or operator contact method to immediately send the alert message to the user or designated contacts.
[0720] Specific examples
[0721] Example 1: Fraudulent call detection and warning
[0722] When a user answers a call, the user's device automatically records the call audio, encrypts it, and sends it to a speech recognition server. The speech recognition server converts the audio into text and sends it to a generative AI analysis server. The generative AI analysis server analyzes the text, and if it detects a pattern that indicates possible fraud, such as "I was in a traffic accident and need money," it generates a warning message and sends it to a notification server. The notification server then sends a push notification to the user saying, "This may be a scam. Please end the call immediately."
[0723] Example 2: Secure Call
[0724] For other calls, the user's device records the voice in the same way, and the speech recognition server converts the voice into text. The generative AI analysis server determines that the text data does not have the characteristics of a special fraud and does not generate a warning message. In this case, the user continues the call as usual.
[0725] As described above, the present invention is a system that can prevent damage from special frauds by analyzing telephone voices in real time, determining the possibility of special fraud, and immediately sending a warning to the user.
[0726] The processing flow will be explained below.
[0727] Step 1:
[0728] The user terminal detects an incoming or outgoing call, and when the user receives or makes a call, the voice recording function is automatically started.
[0729] Step 2:
[0730] The user device records the call audio in real time, and the recorded audio data is encrypted using an encryption algorithm such as AES-256 to ensure security.
[0731] Step 3:
[0732] The user device sends encrypted voice data to a speech recognition server over an internet connection, using HTTPS to ensure data integrity and security.
[0733] Step 4:
[0734] The voice recognition server decrypts the encrypted voice data received from the user terminal. If the decryption is successful, the original voice data is reproduced.
[0735] Step 5:
[0736] The speech recognition server processes the decoded speech data with a speech recognition engine and converts it into text data, thereby extracting the speech content as text information.
[0737] Step 6:
[0738] The speech recognition server sends the converted text data to the generative AI analysis server, using a secure protocol to ensure data confidentiality.
[0739] Step 7:
[0740] The generative AI analysis server analyzes the text data received from the speech recognition server, applies an algorithm to detect the characteristics of special fraud, and determines the likelihood of fraud based on the text data.
[0741] Step 8:
[0742] If the generative AI analysis server determines that there is a high possibility of fraud, it generates a warning message that indicates the possibility of fraud and provides specific instructions to the user.
[0743] Step 9:
[0744] The generative AI analysis server sends the generated warning message to the notification server, which selects the notification method according to the user's settings.
[0745] Step 10:
[0746] The notification server sends warning messages to users or designated contacts via push notification, SMS, email, or operator, allowing users to receive prompt warnings and take appropriate action.
[0747] Example 1
[0748] Next, a description will be given of Example 1. In the following description, the data processing device 12 will be referred to as a "server" and the headset type terminal 314 will be referred to as a "terminal."
[0749] As the number of victims of special telephone frauds increases, there is a growing need for security systems that use advanced technology to detect fraud in real time and quickly warn users. However, current systems lack the functionality to automatically analyze the content of phone calls, detect possible fraud, and immediately warn users. Therefore, the present invention aims to solve these problems.
[0750] The specific processing by the specific processing unit 290 of the data processing device 12 in the first embodiment is realized by the following means.
[0751] In this invention, the server includes means for detecting an incoming call on the user terminal and automatically recording the call audio, means for encrypting the recorded call audio data, means for transmitting the encrypted audio data to a speech recognition server, means for the speech recognition server to decrypt the encrypted audio data and convert it into text data, means for the generative AI analysis server to analyze the text data and determine the possibility of special fraud, means for generating a warning message based on the determination, and means for transmitting the generated warning message to a notification server and notifying the user. This makes it possible to automatically detect the possibility of fraud during a call and send an appropriate warning to the user in real time.
[0752] A "user terminal" is a device that has the function of detecting an incoming call, automatically recording the call audio, and encrypting and transmitting the recorded audio data.
[0753] "Encryption" is a technology that converts recorded call voice data for security purposes and prevents unwanted access.
[0754] A "voice recognition server" is a server that has the function of receiving encrypted voice data sent from a user terminal, decrypting it, and converting the voice data into text data.
[0755] The "generative AI analysis server" is a server that analyzes text data sent from the voice recognition server and executes an algorithm to determine the possibility of special fraud.
[0756] The "notification server" is a server whose role is to receive warning messages provided by the generative AI analysis server and to send the messages to users using an appropriate notification method.
[0757] A "warning message" is a message containing a warning to the user that is generated when the generative AI analysis server detects the possibility of fraud.
[0758] "Push notifications" are short messages sent directly to a user's device, providing an instant way to communicate important information from applications and services.
[0759] "Short Message Service (SMS)" is a communications protocol for sending and receiving short text messages using mobile devices.
[0760] "Email" is a method of sending and receiving text-based messages over the Internet and is a widely used means of communicating information to users.
[0761] "Communication means" refers to any method for transmitting information from a sender to a receiver, and in this system includes push notifications, short message services, emails, etc.
[0762] The system of the present invention consists of a user terminal, a speech recognition server, a generative AI analysis server, and a notification server. The user terminal automatically records the voice of the call, encrypts the data, and sends it to the speech recognition server. The speech recognition server decrypts the recorded data and converts it into text using speech recognition technology. The converted text data is sent to the generative AI analysis server, where it is analyzed to determine the possibility of fraud. If it is determined that there is a high possibility of fraud, a warning message is sent to the user via the notification server.
[0763] The user device can be a smartphone or tablet. When the user answers a call, the device automatically starts recording the call audio. The recorded data is AES encrypted in real time and sent to the speech recognition server via an HTTP POST request.
[0764] The speech recognition server runs a Python-based Flask web server. The received encrypted voice data is decrypted using Python's Cryptography library. The voice data is then converted into text data using a service such as the Google Cloud Speech-to-Text API. This text data is then sent to the generative AI analysis server.
[0765] The generative AI analysis server runs a server built using, for example, Node.js or Python. It analyzes the received text data using a generative AI model, such as OpenAI's GPT-3 model, to detect specific keywords and contexts that indicate the possibility of fraud. For example, it identifies patterns such as "I was in a traffic accident and need money." If it determines that there is a high possibility of fraud, it generates a warning message and sends it to the notification server.
[0766] The notification server runs a server built using, for example, Ruby on Rails or Django. It notifies users of received alert messages using methods such as push notifications, SMS, and email. For example, you can send SMS using the Twilio API or push notifications using Firebase Cloud Messaging.
[0767] Specific examples
[0768] Example 1: Fraudulent call detection and warning
[0769] When a user answers a call, the user's device automatically records the call audio, encrypts it, and sends it to a speech recognition server. The speech recognition server converts the audio into text data and sends the text data to a generative AI analysis server. The generative AI analysis server analyzes the text and detects patterns that indicate possible fraud, such as "I was in a traffic accident and need money." In this case, it generates a fraud warning message and sends it to a notification server. The notification server then sends a push notification to the user saying, "This may be a scam. Please end the call immediately."
[0770] Example 2: Secure Call
[0771] For other calls, the user's device records the voice in the same way, and the speech recognition server converts the voice into text data. The generative AI analysis server analyzes the text data and determines that it does not have the characteristics of a special fraud. In this case, no warning message is generated, and the user can continue the call as usual.
[0772] Example prompt: "Is this a potential scam? 'I've been in a car accident and need money. Please transfer the money.'"
[0773] The flow of the identification process in the first embodiment will be described with reference to FIG.
[0774] Step 1:
[0775] The user device detects an incoming call and automatically starts recording the call audio. The input is the incoming call event. The output is the recorded call audio data. Specifically, the device software catches the incoming call event and starts the recording module. For example, the smartphone's OS notifies the calling app of the incoming call, and an application written in Java starts recording using the Android MediaRecorder API.
[0776] Step 2:
[0777] The recorded call audio data is encrypted in real time. The input is the audio data obtained in step 1. The output is the encrypted audio data. Specifically, the device's encryption module encrypts the recorded data using an encryption algorithm such as AES. For example, an application written in Java encrypts the data using the AES encryption algorithm.
[0778] Step 3:
[0779] The encrypted voice data is sent to the voice recognition server. The input is the encrypted voice data obtained in step 2. The output is a notification that the data has been sent to the server. Specifically, the device sends the voice data to the server using an HTTP POST request. For example, an application written in Java uses an HTTP POST request.
[0780] Step 4:
[0781] The speech recognition server receives and decrypts the encrypted audio data. The input is the encrypted audio data sent in step 3. The output is the decrypted audio data. Specifically, the server's encryption / decryption module decrypts the data using the AES decryption algorithm. For example, a Flask web server written in Python processes the request and decrypts it using the Cryptography library.
[0782] Step 5:
[0783] The speech recognition server converts the decoded speech data into text data. The input is the decoded speech data obtained in step 4. The output is text data. Specifically, speech-to-text conversion is performed using a speech recognition API. For example, the Google Cloud Speech-to-Text API is used to convert speech data into text.
[0784] Step 6:
[0785] The converted text data is sent to the generative AI analysis server. The input is the text data obtained in step 5. The output is a notification that data transmission to the generative AI analysis server has been completed. Specifically, the speech recognition server sends the data using an HTTP POST request. For example, the Python code sends a POST request to the generative AI analysis server.
[0786] Step 7:
[0787] The generative AI analysis server analyzes the text data and determines the likelihood of fraud. The input is the text data received in step 6. The output is the fraud determination result. Specifically, it uses a generative AI model to analyze the text and detect keywords and contexts that indicate the likelihood of fraud. For example, it uses OpenAI's GPT-3 model to identify phrases such as "I was in a traffic accident and need money."
[0788] Step 8:
[0789] If it is determined that there is a high possibility of fraud, a warning message is generated. The input is the fraud determination result from step 7. The output is a warning message. Specifically, the server runs the message generation module to create a warning message. For example, it creates a message saying, "There is a possibility of fraud. Please end the call immediately."
[0790] Step 9:
[0791] The generated warning message is sent to the notification server. The input is the warning message from step 8. The output is a notification of completion of data transmission to the notification server. Specifically, the generative AI analysis server sends the data using an HTTP POST request. For example, JavaScript code sends a POST request to the notification server.
[0792] Step 10:
[0793] The notification server sends an alert message to the user. The input is the alert message received in step 9. The output is an alert notification to the user. Specifically, the notification server sends the message to the user using the appropriate notification method. For example, it can send an SMS using the Twilio API or a push notification using Firebase Cloud Messaging.
[0794] (Application example 1)
[0795] Next, a description will be given of Application Example 1. In the following description, the data processing device 12 will be referred to as a "server" and the headset type terminal 314 will be referred to as a "terminal."
[0796] In recent years, special fraud cases using telephones have been increasing, and many people have become victims. For this reason, there is a need for a system that can detect possible fraud in real time when a call is received and issue a warning. However, existing systems have problems such as low accuracy in fraud detection, lack of real-time response, and a lack of diverse means of notification to users. To address this issue, a system is needed that can detect possible fraud more accurately and quickly, and send warnings to users using a variety of means.
[0797] The specific processing by the specific processing unit 290 of the data processing device 12 in the application example 1 is realized by the following means.
[0798] In this invention, the server includes means for recording telephone voice, means for converting the recorded telephone voice into text data, means for analyzing the converted text data and determining the possibility of special fraud, means for generating a warning message based on the determination, means for transmitting the generated warning message, means for analyzing the text data using a generative AI model, means for starting recording in real time on a user terminal, and means for generating and transmitting a warning message by push notification, SMS, or email. This makes it possible to quickly detect the possibility of fraud with high accuracy and to send warnings to users by various means.
[0799] "Telephone voice" refers to voice data transmitted through telephone communication.
[0800] "Recording means" refers to devices or software for saving voice data, and has the function of saving the contents of telephone conversations as files.
[0801] "Means for converting into text data" refers to devices or software that use voice recognition technology to convert voice data into character-based data.
[0802] "Means for analysis" refers to devices or software that have the ability to analyze text data and evaluate its content based on pattern recognition or algorithms.
[0803] "Special fraud" refers to criminal acts that involve deceiving people over the phone and stealing money from them.
[0804] "Means for determining" refers to a device or software that runs an algorithm to determine whether or not a transaction is fraudulent based on the analysis results.
[0805] A "warning message" refers to a message containing information that is notified to the user when a possible fraud is detected.
[0806] "Generating means" refers to a device or software for creating a warning message, and has the function of automatically creating the necessary wording and warning content.
[0807] "Transmitting means" refers to the communication technology and network infrastructure for transferring the generated alert message to a user terminal or other device.
[0808] "Generative AI models" refer to algorithms and tools that use artificial intelligence to generate and analyze data.
[0809] A "prompt" is an instruction or guidance text to be input into a generative AI model, and is text used to obtain a specific analysis result or product.
[0810] "User terminal" refers to a communication device that is directly operated by a user, such as a smartphone, tablet, or computer.
[0811] "Means for starting recording in real time" refers to devices or software that have the function of automatically starting to record audio the moment a call is initiated.
[0812] A "push notification" refers to a notification message sent to a user device in real time from an application or system.
[0813] "SMS" stands for Short Message Service, a service for sending short text messages over mobile phone networks.
[0814] "Mail" is an abbreviation for electronic mail, a means of communication for sending and receiving text and files over the Internet.
[0815] The present invention is a security system that detects potential fraud in real time for telephone calls received by a user and issues a prompt warning. The system includes the following main components:
[0816] User terminal
[0817] The user terminal is equipped with a function for recording telephone calls in real time. The voice is recorded as soon as the call begins, and the voice data is encrypted to ensure security. The encrypted voice data is immediately sent to a voice recognition server. This allows fraud detection to be performed automatically without the user having to perform any special operations during the call.
[0818] Speech Recognition Server
[0819] The speech recognition server decrypts the encrypted voice data received from the user's device. The decrypted voice data is converted into text data using speech recognition technology, such as Google Speech Recognition. The converted text data is then sent to the generative AI analysis server.
[0820] Generative AI analysis server
[0821] The generative AI analysis server analyzes the received text data and runs algorithms to determine the likelihood of fraud. This analysis uses generative AI models such as BERT and GPT. If a high likelihood of fraud is determined, the generative AI analysis server generates a warning message and sends it to the notification server.
[0822] Specific prompt examples:
[0823] Analyze this text and determine if it is potentially fraudulent if it contains phrases like "I was in a car accident" or "I need money."
[0824] Text: "I was in a car accident and need money now. Please transfer it to my bank account."
[0825] Notification Server
[0826] The notification server notifies the user device of the warning message received from the generative AI analysis server via push notification, SMS, email, etc. This allows the user to respond immediately to potentially fraudulent calls.
[0827] Specific examples
[0828] When a user receives a phone call that is suspected to be fraudulent, the user's device automatically records the audio and sends the data to a speech recognition server. The speech recognition server converts the audio data into text data, which is then analyzed by a generative AI analysis server. If the analysis determines that the call is likely to be fraudulent, the notification server sends a warning message to the user, informing them that "This may be a fraudulent call. Please end the call immediately."
[0829] This allows users to prevent themselves from falling victim to fraud. By combining multiple cutting-edge technologies, the system is able to detect potential fraud with high accuracy and speed, and send warnings via a variety of means.
[0830] The flow of the specific processing in the application example 1 will be described with reference to FIG.
[0831] Step 1:
[0832] When the user terminal detects the start of a call, it automatically records the voice in real time. The recorded voice data is encrypted to ensure security. For encryption, for example, the Fernet encryption library is used. The encrypted voice data is then sent to a voice recognition server.
[0833] Input: Phone call audio
[0834] Data processing: Encrypting audio
[0835] Output: Encrypted audio data
[0836] Step 2:
[0837] The voice recognition server decrypts the received encrypted voice data. For decryption, the same encryption key as that used on the user's device is used. The decrypted voice data is converted into text data using voice recognition technology. For example, the Google Speech Recognition API is used. The converted text data is sent to the generative AI analysis server.
[0838] Input: Encrypted audio data
[0839] Data operations: Decryption of encrypted data and conversion of voice data to text
[0840] Output: Text data
[0841] Step 3:
[0842] The generative AI analysis server uses a generative AI model (e.g., BERT or GPT-3) to analyze the received text data. It uses prompt sentences to analyze the text data and determine the likelihood of fraud. Specific examples of prompt sentences are as follows:
[0843] Analyze this text and determine if it is potentially fraudulent if it contains phrases like "I was in a car accident" or "I need money."
[0844] Text: "I was in a car accident and need money now. Please transfer it to my bank account."
[0845] Input: Text data, prompt
[0846] Data Computation: Text Data Analysis with Generative AI Models
[0847] Output: Possibility of fraud determination result
[0848] Step 4:
[0849] If the generative AI analysis server determines that there is a high possibility of fraud, it generates a warning message, such as "This is a possible fraud. Please end the call immediately." This message is sent to the notification server.
[0850] Input: Possibility of fraud determination result
[0851] Data processing: Generate warning messages
[0852] Output: Warning message
[0853] Step 5:
[0854] The notification server notifies the user device of the warning message received from the generative AI analysis server. The warning message is sent to the user immediately using multiple means, such as push notification, SMS, and email, allowing the user to quickly respond to potentially fraudulent calls.
[0855] Input: warning message
[0856] Data processing: Sending notification messages
[0857] Output: Warning message sent to the user
[0858] The above are the specific processing steps of the program for the system that realizes the application example.
[0859] Furthermore, an emotion engine that estimates the user's emotion may be further combined. That is, the identification processing unit 290 may estimate the user's emotion using the emotion identification model 59, and perform identification processing using the user's emotion.
[0860] The system of the present invention is composed of a user terminal, a speech recognition server, a generative AI analysis server, a notification server, and an emotion engine. Below, we will explain the program processing and specific examples of this system.
[0861] System program processing
[0862] 1. User Device
[0863] The user terminal has the function to automatically start recording the voice when receiving a call. The voice is recorded in real time and encrypted for security. This encrypted voice data is immediately sent to the voice recognition server.
[0864] 2. Speech Recognition Server
[0865] The speech recognition server decrypts the encrypted speech data received from the user's device. The decrypted speech data is converted into text data using speech recognition technology. The converted text data is then sent to the generative AI analysis server.
[0866] 3. Generative AI Analysis Server
[0867] The generative AI analysis server performs detailed analysis of the text data received from the speech recognition server. Based on the text data, an algorithm is run to determine whether or not there is a possibility of special fraud.
[0868] 4. Emotion Engine
[0869] The emotion engine analyzes the voice data acquired from the user's device to recognize the user's emotional state. The results of the emotion engine are sent to the generative AI analysis server, which influences the determination of the possibility of fraud.
[0870] 5. Correction of analysis results
[0871] The generative AI analysis server receives the analysis results from the emotion engine and amends them based on the user's emotional state. For example, if the user is nervous or confused, the analysis results are strengthened.
[0872] 6. Generating and Sending Warning Messages
[0873] The generative AI analysis server generates warning messages as needed based on the corrected analysis results. The generated warning messages are sent to the notification server, which then sends the warning messages via push notification, SMS, email, or operator notification based on the user's settings.
[0874] Specific examples
[0875] Example 1: Fraudulent call detection and warning
[0876] When a user answers a call, the user's device automatically records the call audio, encrypts it, and sends it to a speech recognition server. The speech recognition server converts the speech into text and sends it to a generative AI analysis server. The generative AI analysis server analyzes the text and detects patterns that indicate the possibility of fraud, such as "I was in a traffic accident and need money." If the emotion engine detects a state of tension in the user's voice, it further determines that the possibility of fraud is high. A warning message is generated, and a push notification is sent to the user via the notification server, stating, "This may be a scam. Please end the call immediately."
[0877] Example 2: Secure Call
[0878] The user receives another call. In this case, the user's device again records the audio, and the speech recognition server converts it into text. The generative AI analysis server analyzes the text data and verifies that it does not have the characteristics of a specialized fraud. The emotion engine also detects that the user is relaxed. In this case, no warning message is generated, and the user continues the call as usual.
[0879] As described above, the present invention is a system that can prevent damage from special frauds by analyzing telephone voices in real time, taking into account the user's emotional state, determining the possibility of special fraud, and sending a warning.
[0880] The processing flow will be explained below.
[0881] Step 1:
[0882] The user receives or makes a call. The user device detects this and automatically starts recording the audio.
[0883] Step 2:
[0884] The user device records the call audio in real time, and the recorded audio data is encrypted using an encryption algorithm such as AES-256.
[0885] Step 3:
[0886] The user terminal sends the encrypted voice data to the voice recognition server using the HTTPS protocol.
[0887] Step 4:
[0888] The voice recognition server decrypts the encrypted voice data received from the user terminal. If the decryption is successful, the voice data returns to its original state.
[0889] Step 5:
[0890] The speech recognition server converts the decoded speech data into text data using speech recognition technology, and the converted text data is sent to the generative AI analysis server.
[0891] Step 6:
[0892] The generative AI analysis server analyzes the text data received from the speech recognition server, and executes a specific algorithm to determine the possibility of special fraud.
[0893] Step 7:
[0894] The user device sends the recorded voice data to the emotion engine, which analyzes the user's emotional state (e.g., tension, confusion, calmness, etc.) and sends the results to the generative AI analysis server.
[0895] Step 8:
[0896] The generative AI analysis server receives the analysis results from the emotion engine and corrects the analysis results for the possibility of special fraud. If the user is nervous or confused, the possibility of fraud is strengthened.
[0897] Step 9:
[0898] Based on the final judgment based on the results of the emotion engine, the generative AI analysis server generates a warning message, which includes information about the suspected fraud and specific instructions for the user.
[0899] Step 10:
[0900] The generative AI analysis server sends the generated warning message to the notification server, which selects the method of notification based on the user's settings: push notification, SMS, email, or operator notification.
[0901] Step 11:
[0902] The notification server sends alert messages to the user or designated contacts via the method of their choice, allowing the user to receive prompt warnings and take appropriate action.
[0903] Example 2
[0904] Next, a description will be given of Example 2. In the following description, the data processing device 12 will be referred to as a "server" and the headset type terminal 314 will be referred to as a "terminal."
[0905] In modern times, special telephone frauds have become a social problem. Fraudulent phone calls targeting the elderly in particular have caused many victims, making countermeasures an urgent necessity. However, conventional security systems lack the functionality to detect potential fraud in real time and issue immediate warnings to users. Furthermore, they lack the functionality to correct analysis results by taking the user's emotional state into account, making them prone to false positives and overreactions. The purpose of this invention is to solve these problems and provide a highly reliable fraud prevention system.
[0906] The identification process by the identification processing unit 290 of the data processing device 12 in the second embodiment is realized by the following means. In this invention, the server includes a means for recording telephone voice, a means for encrypting and transmitting the recorded telephone voice, a means for converting the encrypted telephone voice into text data, a means for analyzing the converted text data and determining the possibility of special fraud, a means for analyzing the user's emotional state and correcting the analysis result, a means for generating a warning message based on the corrected result, and a means for transmitting the generated warning message. This enables real-time detection of the possibility of fraud and correction of the analysis result taking the user's emotional state into consideration.
[0907] "Telephone voice" refers to audio data recorded from a telephone conversation.
[0908] A "recording means" is a device or system that stores the audio during a call as digital data.
[0909] "Encryption and transmission means" refers to the methods and techniques by which stored voice data is encrypted for security purposes and transmitted to the appropriate recipient.
[0910] "Means for converting into text data" refers to technology for converting voice data into character data, such as voice recognition technology.
[0911] "Means for analyzing and determining the possibility of special fraud" refers to algorithms and systems that analyze text data and detect signs of fraud from its content.
[0912] "Means for analyzing emotional state" refers to technology or a system for determining emotions (tension, relaxation, etc.) from the user's voice data.
[0913] "Means for correcting the analysis results" refers to a method for correcting the results as necessary to increase the reliability of the analysis results based on information on emotional state.
[0914] A "means for generating a warning message" is a system that generates a warning message to a user when fraud is likely.
[0915] The "means for sending a warning message" refers to a method for sending the generated warning message to the user using a communication means.
[0916] The system of the present invention is composed of a user terminal, a speech recognition server, a generative AI analysis server, an emotion engine, and a notification server. A specific example of the system will be described below.
[0917] User terminal
[0918] The user device has a function that automatically starts recording voice when a call comes in. The recorded voice data is encrypted in real time using the AES method or similar. This encrypted voice data is immediately sent to the voice recognition server using the HTTPS protocol.
[0919] Speech Recognition Server
[0920] The speech recognition server decrypts the encrypted speech data received from the user's device. This decrypted speech data is converted into text data using speech recognition technology. Specifically, it uses a commonly used speech recognition API, such as the speech recognition function of a cloud service. This converted text data is then sent to the generative AI analysis server.
[0921] Generative AI analysis server
[0922] The generative AI analytics server uses machine learning algorithms to analyze the text data received from the speech recognition server. For example, it uses open-source machine learning libraries or cloud-based AI services to analyze the text data and determine the likelihood of fraud. If patterns indicative of possible fraud are detected, the information is sent to the emotion engine.
[0923] Emotion Engine
[0924] The emotion engine analyzes the user's emotional state based on their voice data. The emotional state is determined using commonly used emotion analysis APIs and software. This information is sent to a generative AI analysis server, which corrects the analysis results. For example, if the user is nervous, it is deemed to be a sign of a high possibility of fraud.
[0925] Correcting analysis results and generating warning messages
[0926] Once the analysis results have been corrected, the generative AI analysis server generates a warning message as needed. This warning message is sent to the user's device in real time. Specifically, the warning message is sent to the notification server and delivered to the user via push notification, SMS, email, or other means. For example, the message might say, "This may be a scam. Please end the call immediately."
[0927] Specific examples
[0928] Example 1: Fraudulent call detection
[0929] When a user answers a call, the user device automatically records the call audio and sends the encrypted audio data to a speech recognition server. The speech recognition server converts the audio into text and sends it to a generative AI analysis server. The generative AI analysis server analyzes the text data and detects signs of fraud, such as "I was in a traffic accident and need money." If the emotion engine detects a state of tension in the user's voice, the possibility of fraud increases. Based on this, a warning message is generated and sent to the user via a notification server.
[0930] Example 2: Secure Call
[0931] The user receives another call and the device begins recording the audio. This audio data is also encrypted and sent to the speech recognition server. The recorded audio data is converted into text and sent to the generative AI analysis server. The generative AI analysis server analyzes this text data and verifies that it does not contain any characteristics of specialized fraud. The emotion engine also detects that the user is relaxed, so no warning message is generated. This allows the user to continue the safe call.
[0932] Prompt Sentence Examples
[0933] "The audio of phone calls received by users is recorded in real time and encrypted for security. This encrypted audio data is converted into text data using speech recognition technology and analyzed by a generative AI analysis server. Based on the analyzed text data, the possibility of fraud is determined and a warning message is generated if necessary. The warning message is sent to the user via a notification server to ensure a safe call. If signs of fraud are detected, how should it be handled?"
[0934] As a result, the present invention is a system that can prevent damage from special fraud by analyzing telephone voices in real time, determining the possibility of special fraud taking into account the user's emotional state, and sending a warning.
[0935] The flow of the identification process in the second embodiment will be described with reference to FIG.
[0936] Processing steps of this system's program
[0937] Step 1
[0938] When a user receives a call, the user terminal automatically records the call audio. This audio recording is initiated using a program that is automatically triggered when the user receives a call. The input is the real-time call audio, and the output is the recorded audio file. Specifically, the user terminal launches a recording application and saves the call audio as a file.
[0939] Step 2
[0940] The user terminal encrypts the recorded voice data using the AES encryption method. At this point, the input is the recorded voice file, and the output is the encrypted voice data. Specifically, the user terminal runs an encryption algorithm to convert the voice data into a secure format.
[0941] Step 3
[0942] The user terminal sends encrypted voice data to the voice recognition server. The HTTPS protocol is used for transmission. The input is encrypted voice data, and the output is the encrypted voice data received by the voice recognition server. Specifically, the user terminal creates an HTTP request and sends the data.
[0943] Step 4
[0944] The speech recognition server decrypts the received encrypted audio data. The input is the encrypted audio data, and the output is the decrypted audio data. Specifically, the server runs the AES decryption algorithm to reconstruct the original audio data.
[0945] Step 5
[0946] The speech recognition server converts the decoded speech data into text data. Specifically, it uses a speech recognition API to convert speech to text. The input is the decoded speech data, and the output is text data. Specifically, the speech recognition server calls the speech recognition API (for example, a cloud-based speech recognition service) and obtains the returned text data.
[0947] Step 6
[0948] The speech recognition server sends the generated text data to the generative AI analysis server. The input is the generated text data, and the output is the text data received by the generative AI analysis server. Specifically, the data is sent via a REST API for server-to-server communication.
[0949] Step 7
[0950] The generative AI analysis server analyzes the received text data and runs machine learning algorithms to determine the likelihood of fraud. The input is text data, and the output is a result indicating the likelihood of fraud. Specifically, the generative AI analysis server uses libraries such as TensorFlow and Hugging Face's Transformers to analyze the text data and detect fraudulent patterns.
[0951] Step 8
[0952] The emotion engine analyzes the user's emotional state based on their voice data. The input is the user's voice data, and the output is data about the user's emotional state (tension, confusion, etc.). Specifically, the emotion engine calls an emotion analysis API (e.g., IBM Watson Tone Analyzer) to determine the emotional state.
[0953] Step 9
[0954] The generative AI analysis server corrects the analysis results based on the emotional state data received from the emotion engine. The input is the emotional state data and text analysis results, and the output is the corrected analysis results. Specifically, it reevaluates the possibility of fraud taking the emotional state into account and makes a final judgment.
[0955] Step 10
[0956] The generative AI analysis server generates a warning message based on the corrected analysis results. The input is the corrected analysis results, and the output is the generated warning message. Specifically, the server creates a warning message based on the judgment results.
[0957] Step 11
[0958] The notification server sends the generated warning message to the user device. Methods of sending include push notification, SMS, and email. The input is the generated warning message, and the output is the warning message received by the user device. Specifically, the notification server uses AWS SNS (Simple Notification Service) or a similar notification service to send the message to the user.
[0959] The above are the specific processing steps in this system.
[0960] (Application example 2)
[0961] Next, a description will be given of Application Example 2. In the following description, the data processing device 12 will be referred to as a "server" and the headset type terminal 314 will be referred to as a "terminal."
[0962] In recent years, there has been an increase in special frauds using telephones, and frauds targeting the elderly in particular have become a serious problem. Current security systems lack the means to accurately detect potential fraud in real time and send prompt and appropriate warnings to users. As a result, many users remain at high risk of becoming victims of fraud. Furthermore, few systems take the user's emotional state into account, and the accuracy of fraud detection is low, making it difficult for users to feel safe answering the phone. There is a need for a system that solves this problem.
[0963] The identification process by the identification processing unit 290 of the data processing device 12 in Application Example 2 is realized by the following means. In this invention, the server includes a means for recording telephone voice, a means for encrypting and transmitting the recorded telephone voice, and a means for decrypting the encrypted voice data. This enables secure and highly accurate analysis of telephone voice in real time and fraud detection that takes into account the user's emotional state. In addition, by including a means for quickly sending a warning message via push notification, SMS, email, or operator, it is possible to immediately warn the user of fraud and prevent damage before it occurs.
[0964] "Means for recording telephone voice" refers to technology that automatically records and safely stores the contents of a call when a user initiates a call.
[0965] "Means for encrypting and transmitting recorded telephone voice data" refers to a technology that encrypts recorded telephone voice data to protect it from unauthorized access or eavesdropping by third parties and transmits it to a designated server.
[0966] "Means for decrypting encrypted audio data" refers to a technique for restoring the transmitted encrypted audio data to the original audio data.
[0967] "Means for converting voice data into text data" refers to a technology that analyzes decoded voice data and converts its contents into text information.
[0968] The "means of analyzing the converted text data and determining the possibility of special fraud" is a technology that uses an algorithm to detect signs and patterns of special fraud based on the obtained text data.
[0969] "Means for analyzing the user's emotional state" refers to a technology that analyzes the user's emotions (tension, confusion, fear, etc.) from voice data during a call and determines that state.
[0970] The "means for correcting the determination of the possibility of special fraud based on the analysis results of the emotional state" is a technology for improving the accuracy of the determination of the possibility of special fraud by taking into account the emotional state of the user.
[0971] The "means for generating a warning message" is a technology that automatically generates a notification message to warn the user when it is determined that there is a high possibility of special fraud.
[0972] "Means for sending the generated warning message" refers to a communication technology for immediately delivering the generated warning message to the user, including push notification, SMS, email, or operator notification.
[0973] The purpose of the system of the present invention is to prevent special frauds using voice calls, and it records, encrypts, analyzes telephone voice, and generates and transmits warning messages. The specific system configuration and its program processing are described below.
[0974] System Configuration
[0975] The system of the present invention comprises the following components:
[0976] 1. User device: When a call is initiated, the audio is automatically recorded, and the recorded audio data is encrypted and sent to the server.
[0977] 2. Speech recognition server: Decrypts the encrypted voice data and converts the voice into text data.
[0978] 3. Generative AI analysis server: Analyzes text data to determine the possibility of special fraud. It also analyzes the user's emotional state and adjusts the judgment results based on the results.
[0979] 4. Notification Server: Generates warning messages as needed and sends them to users via push notification, SMS, email or operator.
[0980] Hardware and Software Used
[0981] Smartphone: iOS / Android devices
[0982] Cloud servers: Amazon Web Services (AWS) EC2, Amazon S3
[0983] Speech recognition technology: Google Speech-to-Text API
[0984] Generative AI model: OpenAI GPT-4 API
[0985] Sentiment analysis engine: IBM Watson Tone Analyzer
[0986] Notification service: Firebase Cloud Messaging (FCM)
[0987] Program Processing Overview
[0988] 1. User Device
[0989] The user device automatically records the voice as soon as the call begins and encrypts the recording using a sophisticated encryption algorithm. The encrypted voice data is then immediately sent to a voice recognition server.
[0990] 2. Speech Recognition Server
[0991] The speech recognition server decrypts the received encrypted voice data and converts it into text using the Google Speech-to-Text API. The converted text data is then sent to the generative AI analysis server.
[0992] 3. Generative AI Analysis Server
[0993] The generative AI analysis server uses the OpenAI GPT-4 API to perform detailed analysis of the text data sent from the speech recognition server. To determine whether there are signs of fraud, it detects specific keywords and phrases and then uses IBM Watson Tone Analyzer to analyze the user's emotional state. Based on the results, it corrects the judgment results as necessary.
[0994] 4. Notification Server
[0995] The notification server receives the results of the generative AI analysis server and generates a warning message if it determines that the call is likely to be fraudulent, and sends it to the user via push notification, SMS, or email via Firebase Cloud Messaging (FCM). This warning message urges the user to end the call.
[0996] Examples and prompts
[0997] Examples:
[0998] For example, if a user receives a call while on a call saying, "My son was in a traffic accident and I need money immediately," the system will analyze the call based on the following prompt:
[0999] Example prompt sentence:
[1000] The user received a call. The text of the call is below:
[1001] "Hello, this is the police. Your son has been in a car accident. We need money."
[1002] The user's emotional state is "tense."
[1003] Determine if this is likely a scam based on the following criteria, and if so, generate a warning message.
[1004] conditions:
[1005] Keywords to detect possible fraud: traffic accident, need money, transfer, etc.
[1006] User's emotional state: nervous, confused, scared
[1007] Output format:
[1008] Possibility of fraud: [Yes / No]
[1009] Warning message: [Possible scam. Please end the call immediately / This call is safe.]
[1010] This system makes it possible to detect the risk of special fraud in real time and quickly protect users.
[1011] The flow of the specific processing in the application example 2 will be described with reference to FIG.
[1012] Step 1:
[1013] When a user receives a call, the user terminal automatically starts recording the call audio. The recorded audio data is encrypted using an encryption algorithm. The input is the call audio data, and the output is the encrypted audio data.
[1014] Step 2:
[1015] The encrypted voice data is sent from the user terminal to the voice recognition server, which then decrypts the received data. The input is the encrypted voice data, and the output is the decrypted voice data.
[1016] Step 3:
[1017] The speech recognition server converts the decoded speech data into text data using the Google Speech-to-Text API. The input is speech data and the output is text data.
[1018] Step 4:
[1019] The generative AI analysis server receives the text data sent from the speech recognition server and uses the OpenAI GPT-4 API to analyze the possibility of fraud by detecting keywords and phrases within the text data. The input is the text data, and the output is the analysis results.
[1020] Step 5:
[1021] The generative AI analysis server analyzes the user's emotional state using the IBM Watson Tone Analyzer, an emotion analysis engine, in addition to analyzing the text data. The input is voice data, and the output is the user's emotional state.
[1022] Step 6:
[1023] The generative AI analysis server corrects the judgment of the possibility of special fraud based on the emotion analysis results. For example, if the user is nervous, it will judge that there is a high possibility of fraud. The input is the analysis result and emotional state, and the output is the corrected judgment result.
[1024] Step 7:
[1025] Based on the corrected judgment results, the generative AI analysis server generates warning messages as necessary. The input is the corrected judgment results, and the output is the warning message.
[1026] Step 8:
[1027] The notification server receives the generated alert messages and uses Firebase Cloud Messaging (FCM) to send them to the user via push notification, SMS, email, or operator. The input is the alert message and the output is the notification to the user.
[1028] The specific processing unit 290 transmits the result of the specific processing to the headset type terminal 314. In the headset type terminal 314, the control unit 46A causes the speaker 240 and the display 343 to output the result of the specific processing. The microphone 238 acquires audio indicating a user input regarding the result of the specific processing. The control unit 46A transmits audio data indicating the user input acquired by the microphone 238 to the data processing device 12. In the data processing device 12, the specific processing unit 290 acquires the audio data.
[1029] The data generation model 58 is a so-called generative AI (Artificial Intelligence). An example of the data generation model 58 is ChatGPT (Internet Search<URL: https: / / openai.com / blog / chatgpt> ), Gemini (Internet search <url: https: gemini.google.com ?hl="ja">) and other generation AIs. The data generation model 58 is obtained by performing deep learning on a neural network. A prompt including an instruction is input to the data generation model 58, and inference data such as voice data indicating voice, text data indicating text, and image data indicating an image is also input. The data generation model 58 performs inference on the input inference data in accordance with the instruction indicated by the prompt, and outputs the inference result in a data format such as voice data and text data. Here, inference refers to, for example, analysis, classification, prediction, and / or summarization.
[1030] In the above embodiment, an example was given in which the specific processing is performed by the data processing device 12, but the technology of the present disclosure is not limited to this, and the specific processing may be performed by the headset type terminal 314.
[1031] [Fourth embodiment]
[1032] FIG. 7 shows an example of the configuration of a data processing system 410 according to the fourth embodiment.
[1033] 7, a data processing system 410 includes a data processing device 12 and a robot 414. An example of the data processing device 12 is a server.
[1034] The data processing device 12 includes a computer 22, a database 24, and a communication I / F 26. The computer 22 is an example of a "computer" according to the technology of the present disclosure. The computer 22 includes a processor 28, a RAM 30, and a storage 32. The processor 28, the RAM 30, and the storage 32 are connected to a bus 34. The database 24 and the communication I / F 26 are also connected to the bus 34. The communication I / F 26 is connected to a network 54. Examples of the network 54 include a WAN (Wide Area Network) and / or a LAN (Local Area Network).
[1035] The robot 414 includes a computer 36, a microphone 238, a speaker 240, a camera 42, a communication I / F 44, and a control target 443. The computer 36 includes a processor 46, a RAM 48, and a storage 50. The processor 46, the RAM 48, and the storage 50 are connected to a bus 52. The microphone 238, the speaker 240, the camera 42, and the control target 443 are also connected to the bus 52.
[1036] The microphone 238 receives instructions and the like from the user 20 by receiving voice uttered by the user 20. The microphone 238 captures the voice uttered by the user 20, converts the captured voice into audio data, and outputs it to the processor 46. The speaker 240 outputs audio in accordance with instructions from the processor 46.
[1037] Camera 42 is a small digital camera equipped with an optical system including a lens, aperture, and shutter, and an imaging element such as a CMOS (Complementary Metal-Oxide-Semiconductor) image sensor or a CCD (Charge Coupled Device) image sensor, and captures images of the surroundings of user 20 (for example, an imaging range defined by an angle of view equivalent to the field of vision of a typical healthy person).
[1038] The communication I / F 44 is connected to a network 54. The communication I / Fs 44 and 26 are responsible for the exchange of various information between the processor 46 and the processor 28 via the network 54. The exchange of various information between the processor 46 and the processor 28 using the communication I / Fs 44 and 26 is carried out in a secure state.
[1039] The control object 443 includes a display device, LEDs in the eyes, and motors for driving the arms, hands, and feet. The posture and gestures of the robot 414 are controlled by controlling the motors of the arms, hands, and feet. Some of the emotions of the robot 414 can be expressed by controlling these motors. In addition, the facial expressions of the robot 414 can also be expressed by controlling the light emission state of the LEDs in the eyes of the robot 414.
[1040] Fig. 8 shows an example of the main functions of the data processing device 12 and the robot 414. As shown in Fig. 8, in the data processing device 12, a specific process is performed by the processor 28. A specific process program 56 is stored in the storage 32.
[1041] The specific processing program 56 is an example of a "program" according to the technology of the present disclosure. The processor 28 reads the specific processing program 56 from the storage 32 and executes the read specific processing program 56 on the RAM 30. The specific processing is realized by the processor 28 operating as a specific processing unit 290 in accordance with the specific processing program 56 executed on the RAM 30.
[1042] The storage 32 stores a data generation model 58 and an emotion identification model 59. The data generation model 58 and the emotion identification model 59 are used by the identification processing unit 290.
[1043] In the robot 414, the processor 46 performs the reception output process. A reception output program 60 is stored in the storage 50. The processor 46 reads the reception output program 60 from the storage 50 and executes the read reception output program 60 on the RAM 48. The reception output process is realized by the processor 46 operating as the control unit 46A in accordance with the reception output program 60 executed on the RAM 48.
[1044] Next, a description will be given of the specific processing performed by the specific processing unit 290 of the data processing device 12. In the following description, the data processing device 12 will be referred to as a "server" and the robot 414 will be referred to as a "terminal."
[1045] The system of the present invention is composed of a user terminal, a speech recognition server, a generative AI analysis server, and a notification server. Below, we will explain the program processing and specific examples of this system.
[1046] System program processing
[1047] 1. User Device
[1048] The user terminal has the function to automatically start recording the voice when receiving a call. The voice is recorded in real time and encrypted for security. This encrypted voice data is immediately sent to the voice recognition server.
[1049] 2. Speech Recognition Server
[1050] The speech recognition server decrypts the encrypted speech data received from the user's device. The decrypted speech data is converted into text data using speech recognition technology. The converted text data is then sent to the generative AI analysis server.
[1051] 3. Generative AI Analysis Server
[1052] The generative AI analysis server performs detailed analysis of the text data received from the speech recognition server. Based on the text data, an algorithm is run to determine whether there is a possibility of special fraud. If it determines that there is a high possibility of fraud, the generative AI analysis server generates a warning message and sends it to the notification server.
[1053] 4. Notification Server
[1054] The notification server receives alert messages from the generative AI analysis server and, based on the user's settings, selects push notification, SMS, email, or operator contact method to immediately send the alert message to the user or designated contacts.
[1055] Specific examples
[1056] Example 1: Fraudulent call detection and warning
[1057] When a user answers a call, the user's device automatically records the call audio, encrypts it, and sends it to a speech recognition server. The speech recognition server converts the audio into text and sends it to a generative AI analysis server. The generative AI analysis server analyzes the text, and if it detects a pattern that indicates possible fraud, such as "I was in a traffic accident and need money," it generates a warning message and sends it to a notification server. The notification server then sends a push notification to the user saying, "This may be a scam. Please end the call immediately."
[1058] Example 2: Secure Call
[1059] For other calls, the user's device records the voice in the same way, and the speech recognition server converts the voice into text. The generative AI analysis server determines that the text data does not have the characteristics of a special fraud and does not generate a warning message. In this case, the user continues the call as usual.
[1060] As described above, the present invention is a system that can prevent damage from special frauds by analyzing telephone voices in real time, determining the possibility of special fraud, and immediately sending a warning to the user.
[1061] The processing flow will be explained below.
[1062] Step 1:
[1063] The user terminal detects an incoming or outgoing call, and when the user receives or makes a call, the voice recording function is automatically started.
[1064] Step 2:
[1065] The user device records the call audio in real time, and the recorded audio data is encrypted using an encryption algorithm such as AES-256 to ensure security.
[1066] Step 3:
[1067] The user device sends encrypted voice data to a speech recognition server over an internet connection, using HTTPS to ensure data integrity and security.
[1068] Step 4:
[1069] The voice recognition server decrypts the encrypted voice data received from the user terminal. If the decryption is successful, the original voice data is reproduced.
[1070] Step 5:
[1071] The speech recognition server processes the decoded speech data with a speech recognition engine and converts it into text data, thereby extracting the speech content as text information.
[1072] Step 6:
[1073] The speech recognition server sends the converted text data to the generative AI analysis server, using a secure protocol to ensure data confidentiality.
[1074] Step 7:
[1075] The generative AI analysis server analyzes the text data received from the speech recognition server, applies an algorithm to detect the characteristics of special fraud, and determines the likelihood of fraud based on the text data.
[1076] Step 8:
[1077] If the generative AI analysis server determines that there is a high possibility of fraud, it generates a warning message that indicates the possibility of fraud and provides specific instructions to the user.
[1078] Step 9:
[1079] The generative AI analysis server sends the generated warning message to the notification server, which selects the notification method according to the user's settings.
[1080] Step 10:
[1081] The notification server sends warning messages to users or designated contacts via push notification, SMS, email, or operator, allowing users to receive prompt warnings and take appropriate action.
[1082] Example 1
[1083] Next, a description will be given of Example 1. In the following description, the data processing device 12 will be referred to as a "server" and the robot 414 will be referred to as a "terminal."
[1084] As the number of victims of special telephone frauds increases, there is a growing need for security systems that use advanced technology to detect fraud in real time and quickly warn users. However, current systems lack the functionality to automatically analyze the content of phone calls, detect possible fraud, and immediately warn users. Therefore, the present invention aims to solve these problems.
[1085] The specific processing by the specific processing unit 290 of the data processing device 12 in the first embodiment is realized by the following means.
[1086] In this invention, the server includes means for detecting an incoming call on the user terminal and automatically recording the call audio, means for encrypting the recorded call audio data, means for transmitting the encrypted audio data to a speech recognition server, means for the speech recognition server to decrypt the encrypted audio data and convert it into text data, means for the generative AI analysis server to analyze the text data and determine the possibility of special fraud, means for generating a warning message based on the determination, and means for transmitting the generated warning message to a notification server and notifying the user. This makes it possible to automatically detect the possibility of fraud during a call and send an appropriate warning to the user in real time.
[1087] A "user terminal" is a device that has the function of detecting an incoming call, automatically recording the call audio, and encrypting and transmitting the recorded audio data.
[1088] "Encryption" is a technology that converts recorded call voice data for security purposes and prevents unwanted access.
[1089] A "voice recognition server" is a server that has the function of receiving encrypted voice data sent from a user terminal, decrypting it, and converting the voice data into text data.
[1090] The "generative AI analysis server" is a server that analyzes text data sent from the voice recognition server and executes an algorithm to determine the possibility of special fraud.
[1091] The "notification server" is a server whose role is to receive warning messages provided by the generative AI analysis server and to send the messages to users using an appropriate notification method.
[1092] A "warning message" is a message containing a warning to the user that is generated when the generative AI analysis server detects the possibility of fraud.
[1093] "Push notifications" are short messages sent directly to a user's device, providing an instant way to communicate important information from applications and services.
[1094] "Short Message Service (SMS)" is a communications protocol for sending and receiving short text messages using mobile devices.
[1095] "Email" is a method of sending and receiving text-based messages over the Internet and is a widely used means of communicating information to users.
[1096] "Communication means" refers to any method for transmitting information from a sender to a receiver, and in this system includes push notifications, short message services, emails, etc.
[1097] The system of the present invention consists of a user terminal, a speech recognition server, a generative AI analysis server, and a notification server. The user terminal automatically records the voice of the call, encrypts the data, and sends it to the speech recognition server. The speech recognition server decrypts the recorded data and converts it into text using speech recognition technology. The converted text data is sent to the generative AI analysis server, where it is analyzed to determine the possibility of fraud. If it is determined that there is a high possibility of fraud, a warning message is sent to the user via the notification server.
[1098] The user device can be a smartphone or tablet. When the user answers a call, the device automatically starts recording the call audio. The recorded data is AES encrypted in real time and sent to the speech recognition server via an HTTP POST request.
[1099] The speech recognition server runs a Python-based Flask web server. The received encrypted voice data is decrypted using Python's Cryptography library. The voice data is then converted into text data using a service such as the Google Cloud Speech-to-Text API. This text data is then sent to the generative AI analysis server.
[1100] The generative AI analysis server runs a server built using, for example, Node.js or Python. It analyzes the received text data using a generative AI model, such as OpenAI's GPT-3 model, to detect specific keywords and contexts that indicate the possibility of fraud. For example, it identifies patterns such as "I was in a traffic accident and need money." If it determines that there is a high possibility of fraud, it generates a warning message and sends it to the notification server.
[1101] The notification server runs a server built using, for example, Ruby on Rails or Django. It notifies users of received alert messages using methods such as push notifications, SMS, and email. For example, you can send SMS using the Twilio API or push notifications using Firebase Cloud Messaging.
[1102] Specific examples
[1103] Example 1: Fraudulent call detection and warning
[1104] When a user answers a call, the user's device automatically records the call audio, encrypts it, and sends it to a speech recognition server. The speech recognition server converts the audio into text data and sends the text data to a generative AI analysis server. The generative AI analysis server analyzes the text and detects patterns that indicate possible fraud, such as "I was in a traffic accident and need money." In this case, it generates a fraud warning message and sends it to a notification server. The notification server then sends a push notification to the user saying, "This may be a scam. Please end the call immediately."
[1105] Example 2: Secure Call
[1106] For other calls, the user's device records the voice in the same way, and the speech recognition server converts the voice into text data. The generative AI analysis server analyzes the text data and determines that it does not have the characteristics of a special fraud. In this case, no warning message is generated, and the user can continue the call as usual.
[1107] Example prompt: "Is this a potential scam? 'I've been in a car accident and need money. Please transfer the money.'"
[1108] The flow of the identification process in the first embodiment will be described with reference to FIG.
[1109] Step 1:
[1110] The user device detects an incoming call and automatically starts recording the call audio. The input is the incoming call event. The output is the recorded call audio data. Specifically, the device software catches the incoming call event and starts the recording module. For example, the smartphone's OS notifies the calling app of the incoming call, and an application written in Java starts recording using the Android MediaRecorder API.
[1111] Step 2:
[1112] The recorded call audio data is encrypted in real time. The input is the audio data obtained in step 1. The output is the encrypted audio data. Specifically, the device's encryption module encrypts the recorded data using an encryption algorithm such as AES. For example, an application written in Java encrypts the data using the AES encryption algorithm.
[1113] Step 3:
[1114] The encrypted voice data is sent to the voice recognition server. The input is the encrypted voice data obtained in step 2. The output is a notification that the data has been sent to the server. Specifically, the device sends the voice data to the server using an HTTP POST request. For example, an application written in Java uses an HTTP POST request.
[1115] Step 4:
[1116] The speech recognition server receives and decrypts the encrypted audio data. The input is the encrypted audio data sent in step 3. The output is the decrypted audio data. Specifically, the server's encryption / decryption module decrypts the data using the AES decryption algorithm. For example, a Flask web server written in Python processes the request and decrypts it using the Cryptography library.
[1117] Step 5:
[1118] The speech recognition server converts the decoded speech data into text data. The input is the decoded speech data obtained in step 4. The output is text data. Specifically, speech-to-text conversion is performed using a speech recognition API. For example, the Google Cloud Speech-to-Text API is used to convert speech data into text.
[1119] Step 6:
[1120] The converted text data is sent to the generative AI analysis server. The input is the text data obtained in step 5. The output is a notification that data transmission to the generative AI analysis server has been completed. Specifically, the speech recognition server sends the data using an HTTP POST request. For example, the Python code sends a POST request to the generative AI analysis server.
[1121] Step 7:
[1122] The generative AI analysis server analyzes the text data and determines the likelihood of fraud. The input is the text data received in step 6. The output is the fraud determination result. Specifically, it uses a generative AI model to analyze the text and detect keywords and contexts that indicate the likelihood of fraud. For example, it uses OpenAI's GPT-3 model to identify phrases such as "I was in a traffic accident and need money."
[1123] Step 8:
[1124] If it is determined that there is a high possibility of fraud, a warning message is generated. The input is the fraud determination result from step 7. The output is a warning message. Specifically, the server runs the message generation module to create a warning message. For example, it creates a message saying, "There is a possibility of fraud. Please end the call immediately."
[1125] Step 9:
[1126] The generated warning message is sent to the notification server. The input is the warning message from step 8. The output is a notification of completion of data transmission to the notification server. Specifically, the generative AI analysis server sends the data using an HTTP POST request. For example, JavaScript code sends a POST request to the notification server.
[1127] Step 10:
[1128] The notification server sends an alert message to the user. The input is the alert message received in step 9. The output is an alert notification to the user. Specifically, the notification server sends the message to the user using the appropriate notification method. For example, it can send an SMS using the Twilio API or a push notification using Firebase Cloud Messaging.
[1129] (Application example 1)
[1130] Next, a description will be given of Application Example 1. In the following description, the data processing device 12 will be referred to as a "server" and the robot 414 will be referred to as a "terminal."
[1131] In recent years, special fraud cases using telephones have been increasing, and many people have become victims. For this reason, there is a need for a system that can detect possible fraud in real time when a call is received and issue a warning. However, existing systems have problems such as low accuracy in fraud detection, lack of real-time response, and a lack of diverse means of notification to users. To address this issue, a system is needed that can detect possible fraud more accurately and quickly, and send warnings to users using a variety of means.
[1132] The specific processing by the specific processing unit 290 of the data processing device 12 in the application example 1 is realized by the following means.
[1133] In this invention, the server includes means for recording telephone voice, means for converting the recorded telephone voice into text data, means for analyzing the converted text data and determining the possibility of special fraud, means for generating a warning message based on the determination, means for transmitting the generated warning message, means for analyzing the text data using a generative AI model, means for starting recording in real time on a user terminal, and means for generating and transmitting a warning message by push notification, SMS, or email. This makes it possible to quickly detect the possibility of fraud with high accuracy and to send warnings to users by various means.
[1134] "Telephone voice" refers to voice data transmitted through telephone communication.
[1135] "Recording means" refers to devices or software for saving voice data, and has the function of saving the contents of telephone conversations as files.
[1136] "Means for converting into text data" refers to devices or software that use voice recognition technology to convert voice data into character-based data.
[1137] "Means for analysis" refers to devices or software that have the ability to analyze text data and evaluate its content based on pattern recognition or algorithms.
[1138] "Special fraud" refers to criminal acts that involve deceiving people over the phone and stealing money from them.
[1139] "Means for determining" refers to a device or software that runs an algorithm to determine whether or not a transaction is fraudulent based on the analysis results.
[1140] A "warning message" refers to a message containing information that is notified to the user when a possible fraud is detected.
[1141] "Generating means" refers to a device or software for creating a warning message, and has the function of automatically creating the necessary wording and warning content.
[1142] "Transmitting means" refers to the communication technology and network infrastructure for transferring the generated alert message to a user terminal or other device.
[1143] "Generative AI models" refer to algorithms and tools that use artificial intelligence to generate and analyze data.
[1144] A "prompt" is an instruction or guidance text to be input into a generative AI model, and is text used to obtain a specific analysis result or product.
[1145] "User terminal" refers to a communication device that is directly operated by a user, such as a smartphone, tablet, or computer.
[1146] "Means for starting recording in real time" refers to devices or software that have the function of automatically starting to record audio the moment a call is initiated.
[1147] A "push notification" refers to a notification message sent to a user device in real time from an application or system.
[1148] "SMS" stands for Short Message Service, a service for sending short text messages over mobile phone networks.
[1149] "Mail" is an abbreviation for electronic mail, a means of communication for sending and receiving text and files over the Internet.
[1150] The present invention is a security system that detects potential fraud in real time for telephone calls received by a user and issues a prompt warning. The system includes the following main components:
[1151] User terminal
[1152] The user terminal is equipped with a function for recording telephone calls in real time. The voice is recorded as soon as the call begins, and the voice data is encrypted to ensure security. The encrypted voice data is immediately sent to a voice recognition server. This allows fraud detection to be performed automatically without the user having to perform any special operations during the call.
[1153] Speech Recognition Server
[1154] The speech recognition server decrypts the encrypted voice data received from the user's device. The decrypted voice data is converted into text data using speech recognition technology, such as Google Speech Recognition. The converted text data is then sent to the generative AI analysis server.
[1155] Generative AI analysis server
[1156] The generative AI analysis server analyzes the received text data and runs algorithms to determine the likelihood of fraud. This analysis uses generative AI models such as BERT and GPT. If a high likelihood of fraud is determined, the generative AI analysis server generates a warning message and sends it to the notification server.
[1157] Specific prompt examples:
[1158] Analyze this text and determine if it is potentially fraudulent if it contains phrases like "I was in a car accident" or "I need money."
[1159] Text: "I was in a car accident and need money now. Please transfer it to my bank account."
[1160] Notification Server
[1161] The notification server notifies the user device of the warning message received from the generative AI analysis server via push notification, SMS, email, etc. This allows the user to respond immediately to potentially fraudulent calls.
[1162] Specific examples
[1163] When a user receives a phone call that is suspected to be fraudulent, the user's device automatically records the audio and sends the data to a speech recognition server. The speech recognition server converts the audio data into text data, which is then analyzed by a generative AI analysis server. If the analysis determines that the call is likely to be fraudulent, the notification server sends a warning message to the user, informing them that "This may be a fraudulent call. Please end the call immediately."
[1164] This allows users to prevent themselves from falling victim to fraud. By combining multiple cutting-edge technologies, the system is able to detect potential fraud with high accuracy and speed, and send warnings via a variety of means.
[1165] The flow of the specific processing in the application example 1 will be described with reference to FIG.
[1166] Step 1:
[1167] When the user terminal detects the start of a call, it automatically records the voice in real time. The recorded voice data is encrypted to ensure security. For encryption, for example, the Fernet encryption library is used. The encrypted voice data is then sent to a voice recognition server.
[1168] Input: Phone call audio
[1169] Data processing: Encrypting audio
[1170] Output: Encrypted audio data
[1171] Step 2:
[1172] The voice recognition server decrypts the received encrypted voice data. For decryption, the same encryption key as that used on the user's device is used. The decrypted voice data is converted into text data using voice recognition technology. For example, the Google Speech Recognition API is used. The converted text data is sent to the generative AI analysis server.
[1173] Input: Encrypted audio data
[1174] Data operations: Decryption of encrypted data and conversion of voice data to text
[1175] Output: Text data
[1176] Step 3:
[1177] The generative AI analysis server uses a generative AI model (e.g., BERT or GPT-3) to analyze the received text data. It uses prompt sentences to analyze the text data and determine the likelihood of fraud. Specific examples of prompt sentences are as follows:
[1178] Analyze this text and determine if it is potentially fraudulent if it contains phrases like "I was in a car accident" or "I need money."
[1179] Text: "I was in a car accident and need money now. Please transfer it to my bank account."
[1180] Input: Text data, prompt
[1181] Data Computation: Text Data Analysis with Generative AI Models
[1182] Output: Possibility of fraud determination result
[1183] Step 4:
[1184] If the generative AI analysis server determines that there is a high possibility of fraud, it generates a warning message, such as "This is a possible fraud. Please end the call immediately." This message is sent to the notification server.
[1185] Input: Possibility of fraud determination result
[1186] Data processing: Generate warning messages
[1187] Output: Warning message
[1188] Step 5:
[1189] The notification server notifies the user device of the warning message received from the generative AI analysis server. The warning message is sent to the user immediately using multiple means, such as push notification, SMS, and email, allowing the user to quickly respond to potentially fraudulent calls.
[1190] Input: warning message
[1191] Data processing: Sending notification messages
[1192] Output: Warning message sent to the user
[1193] The above are the specific processing steps of the program for the system that realizes the application example.
[1194] Furthermore, an emotion engine that estimates the user's emotion may be further combined. That is, the identification processing unit 290 may estimate the user's emotion using the emotion identification model 59, and perform identification processing using the user's emotion.
[1195] The system of the present invention is composed of a user terminal, a speech recognition server, a generative AI analysis server, a notification server, and an emotion engine. Below, we will explain the program processing and specific examples of this system.
[1196] System program processing
[1197] 1. User Device
[1198] The user terminal has the function to automatically start recording the voice when receiving a call. The voice is recorded in real time and encrypted for security. This encrypted voice data is immediately sent to the voice recognition server.
[1199] 2. Speech Recognition Server
[1200] The speech recognition server decrypts the encrypted speech data received from the user's device. The decrypted speech data is converted into text data using speech recognition technology. The converted text data is then sent to the generative AI analysis server.
[1201] 3. Generative AI Analysis Server
[1202] The generative AI analysis server performs detailed analysis of the text data received from the speech recognition server. Based on the text data, an algorithm is run to determine whether or not there is a possibility of special fraud.
[1203] 4. Emotion Engine
[1204] The emotion engine analyzes the voice data acquired from the user's device to recognize the user's emotional state. The results of the emotion engine are sent to the generative AI analysis server, which influences the determination of the possibility of fraud.
[1205] 5. Correction of analysis results
[1206] The generative AI analysis server receives the analysis results from the emotion engine and amends them based on the user's emotional state. For example, if the user is nervous or confused, the analysis results are strengthened.
[1207] 6. Generating and Sending Warning Messages
[1208] The generative AI analysis server generates warning messages as needed based on the corrected analysis results. The generated warning messages are sent to the notification server, which then sends the warning messages via push notification, SMS, email, or operator notification based on the user's settings.
[1209] Specific examples
[1210] Example 1: Fraudulent call detection and warning
[1211] When a user answers a call, the user's device automatically records the call audio, encrypts it, and sends it to a speech recognition server. The speech recognition server converts the speech into text and sends it to a generative AI analysis server. The generative AI analysis server analyzes the text and detects patterns that indicate the possibility of fraud, such as "I was in a traffic accident and need money." If the emotion engine detects a state of tension in the user's voice, it further determines that the possibility of fraud is high. A warning message is generated, and a push notification is sent to the user via the notification server, stating, "This may be a scam. Please end the call immediately."
[1212] Example 2: Secure Call
[1213] The user receives another call. In this case, the user's device again records the audio, and the speech recognition server converts it into text. The generative AI analysis server analyzes the text data and verifies that it does not have the characteristics of a specialized fraud. The emotion engine also detects that the user is relaxed. In this case, no warning message is generated, and the user continues the call as usual.
[1214] As described above, the present invention is a system that can prevent damage from special frauds by analyzing telephone voices in real time, taking into account the user's emotional state, determining the possibility of special fraud, and sending a warning.
[1215] The processing flow will be explained below.
[1216] Step 1:
[1217] The user receives or makes a call. The user device detects this and automatically starts recording the audio.
[1218] Step 2:
[1219] The user device records the call audio in real time, and the recorded audio data is encrypted using an encryption algorithm such as AES-256.
[1220] Step 3:
[1221] The user terminal sends the encrypted voice data to the voice recognition server using the HTTPS protocol.
[1222] Step 4:
[1223] The voice recognition server decrypts the encrypted voice data received from the user terminal. If the decryption is successful, the voice data returns to its original state.
[1224] Step 5:
[1225] The speech recognition server converts the decoded speech data into text data using speech recognition technology, and the converted text data is sent to the generative AI analysis server.
[1226] Step 6:
[1227] The generative AI analysis server analyzes the text data received from the speech recognition server, and executes a specific algorithm to determine the possibility of special fraud.
[1228] Step 7:
[1229] The user device sends the recorded voice data to the emotion engine, which analyzes the user's emotional state (e.g., tension, confusion, calmness, etc.) and sends the results to the generative AI analysis server.
[1230] Step 8:
[1231] The generative AI analysis server receives the analysis results from the emotion engine and corrects the analysis results for the possibility of special fraud. If the user is nervous or confused, the possibility of fraud is strengthened.
[1232] Step 9:
[1233] Based on the final judgment based on the results of the emotion engine, the generative AI analysis server generates a warning message, which includes information about the suspected fraud and specific instructions for the user.
[1234] Step 10:
[1235] The generative AI analysis server sends the generated warning message to the notification server, which selects the method of notification based on the user's settings: push notification, SMS, email, or operator notification.
[1236] Step 11:
[1237] The notification server sends alert messages to the user or designated contacts via the method of their choice, allowing the user to receive prompt warnings and take appropriate action.
[1238] Example 2
[1239] Next, a description will be given of Example 2. In the following description, the data processing device 12 will be referred to as a "server" and the robot 414 will be referred to as a "terminal."
[1240] In modern times, special telephone frauds have become a social problem. Fraudulent phone calls targeting the elderly in particular have caused many victims, making countermeasures an urgent necessity. However, conventional security systems lack the functionality to detect potential fraud in real time and issue immediate warnings to users. Furthermore, they lack the functionality to correct analysis results by taking the user's emotional state into account, making them prone to false positives and overreactions. The purpose of this invention is to solve these problems and provide a highly reliable fraud prevention system.
[1241] The identification process by the identification processing unit 290 of the data processing device 12 in the second embodiment is realized by the following means. In this invention, the server includes a means for recording telephone voice, a means for encrypting and transmitting the recorded telephone voice, a means for converting the encrypted telephone voice into text data, a means for analyzing the converted text data and determining the possibility of special fraud, a means for analyzing the user's emotional state and correcting the analysis result, a means for generating a warning message based on the corrected result, and a means for transmitting the generated warning message. This enables real-time detection of the possibility of fraud and correction of the analysis result taking the user's emotional state into consideration.
[1242] "Telephone voice" refers to audio data recorded from a telephone conversation.
[1243] A "recording means" is a device or system that stores the audio during a call as digital data.
[1244] "Encryption and transmission means" refers to the methods and techniques by which stored voice data is encrypted for security purposes and transmitted to the appropriate recipient.
[1245] "Means for converting into text data" refers to technology for converting voice data into character data, such as voice recognition technology.
[1246] "Means for analyzing and determining the possibility of special fraud" refers to algorithms and systems that analyze text data and detect signs of fraud from its content.
[1247] "Means for analyzing emotional state" refers to technology or a system for determining emotions (tension, relaxation, etc.) from the user's voice data.
[1248] "Means for correcting the analysis results" refers to a method for correcting the results as necessary to increase the reliability of the analysis results based on information on emotional state.
[1249] A "means for generating a warning message" is a system that generates a warning message to a user when fraud is likely.
[1250] The "means for sending a warning message" refers to a method for sending the generated warning message to the user using a communication means.
[1251] The system of the present invention is composed of a user terminal, a speech recognition server, a generative AI analysis server, an emotion engine, and a notification server. A specific example of the system will be described below.
[1252] User terminal
[1253] The user device has a function that automatically starts recording voice when a call comes in. The recorded voice data is encrypted in real time using the AES method or similar. This encrypted voice data is immediately sent to the voice recognition server using the HTTPS protocol.
[1254] Speech Recognition Server
[1255] The speech recognition server decrypts the encrypted speech data received from the user's device. This decrypted speech data is converted into text data using speech recognition technology. Specifically, it uses a commonly used speech recognition API, such as the speech recognition function of a cloud service. This converted text data is then sent to the generative AI analysis server.
[1256] Generative AI analysis server
[1257] The generative AI analytics server uses machine learning algorithms to analyze the text data received from the speech recognition server. For example, it uses open-source machine learning libraries or cloud-based AI services to analyze the text data and determine the likelihood of fraud. If patterns indicative of possible fraud are detected, the information is sent to the emotion engine.
[1258] Emotion Engine
[1259] The emotion engine analyzes the user's emotional state based on their voice data. The emotional state is determined using commonly used emotion analysis APIs and software. This information is sent to a generative AI analysis server, which corrects the analysis results. For example, if the user is nervous, it is deemed to be a sign of a high possibility of fraud.
[1260] Correcting analysis results and generating warning messages
[1261] Once the analysis results have been corrected, the generative AI analysis server generates a warning message as needed. This warning message is sent to the user's device in real time. Specifically, the warning message is sent to the notification server and delivered to the user via push notification, SMS, email, or other means. For example, the message might say, "This may be a scam. Please end the call immediately."
[1262] Specific examples
[1263] Example 1: Fraudulent call detection
[1264] When a user answers a call, the user device automatically records the call audio and sends the encrypted audio data to a speech recognition server. The speech recognition server converts the audio into text and sends it to a generative AI analysis server. The generative AI analysis server analyzes the text data and detects signs of fraud, such as "I was in a traffic accident and need money." If the emotion engine detects a state of tension in the user's voice, the possibility of fraud increases. Based on this, a warning message is generated and sent to the user via a notification server.
[1265] Example 2: Secure Call
[1266] The user receives another call and the device begins recording the audio. This audio data is also encrypted and sent to the speech recognition server. The recorded audio data is converted into text and sent to the generative AI analysis server. The generative AI analysis server analyzes this text data and verifies that it does not contain any characteristics of specialized fraud. The emotion engine also detects that the user is relaxed, so no warning message is generated. This allows the user to continue the safe call.
[1267] Prompt Sentence Examples
[1268] "The audio of phone calls received by users is recorded in real time and encrypted for security. This encrypted audio data is converted into text data using speech recognition technology and analyzed by a generative AI analysis server. Based on the analyzed text data, the possibility of fraud is determined and a warning message is generated if necessary. The warning message is sent to the user via a notification server to ensure a safe call. If signs of fraud are detected, how should it be handled?"
[1269] As a result, the present invention is a system that can prevent damage from special fraud by analyzing telephone voices in real time, determining the possibility of special fraud taking into account the user's emotional state, and sending a warning.
[1270] The flow of the identification process in the second embodiment will be described with reference to FIG.
[1271] Processing steps of this system's program
[1272] Step 1
[1273] When a user receives a call, the user terminal automatically records the call audio. This audio recording is initiated using a program that is automatically triggered when the user receives a call. The input is the real-time call audio, and the output is the recorded audio file. Specifically, the user terminal launches a recording application and saves the call audio as a file.
[1274] Step 2
[1275] The user terminal encrypts the recorded voice data using the AES encryption method. At this point, the input is the recorded voice file, and the output is the encrypted voice data. Specifically, the user terminal runs an encryption algorithm to convert the voice data into a secure format.
[1276] Step 3
[1277] The user terminal sends encrypted voice data to the voice recognition server. The HTTPS protocol is used for transmission. The input is encrypted voice data, and the output is the encrypted voice data received by the voice recognition server. Specifically, the user terminal creates an HTTP request and sends the data.
[1278] Step 4
[1279] The speech recognition server decrypts the received encrypted audio data. The input is the encrypted audio data, and the output is the decrypted audio data. Specifically, the server runs the AES decryption algorithm to reconstruct the original audio data.
[1280] Step 5
[1281] The speech recognition server converts the decoded speech data into text data. Specifically, it uses a speech recognition API to convert speech to text. The input is the decoded speech data, and the output is text data. Specifically, the speech recognition server calls the speech recognition API (for example, a cloud-based speech recognition service) and obtains the returned text data.
[1282] Step 6
[1283] The speech recognition server sends the generated text data to the generative AI analysis server. The input is the generated text data, and the output is the text data received by the generative AI analysis server. Specifically, the data is sent via a REST API for server-to-server communication.
[1284] Step 7
[1285] The generative AI analysis server analyzes the received text data and runs machine learning algorithms to determine the likelihood of fraud. The input is text data, and the output is a result indicating the likelihood of fraud. Specifically, the generative AI analysis server uses libraries such as TensorFlow and Hugging Face's Transformers to analyze the text data and detect fraudulent patterns.
[1286] Step 8
[1287] The emotion engine analyzes the user's emotional state based on their voice data. The input is the user's voice data, and the output is data about the user's emotional state (tension, confusion, etc.). Specifically, the emotion engine calls an emotion analysis API (e.g., IBM Watson Tone Analyzer) to determine the emotional state.
[1288] Step 9
[1289] The generative AI analysis server corrects the analysis results based on the emotional state data received from the emotion engine. The input is the emotional state data and text analysis results, and the output is the corrected analysis results. Specifically, it reevaluates the possibility of fraud taking the emotional state into account and makes a final judgment.
[1290] Step 10
[1291] The generative AI analysis server generates a warning message based on the corrected analysis results. The input is the corrected analysis results, and the output is the generated warning message. Specifically, the server creates a warning message based on the judgment results.
[1292] Step 11
[1293] The notification server sends the generated warning message to the user device. Methods of sending include push notification, SMS, and email. The input is the generated warning message, and the output is the warning message received by the user device. Specifically, the notification server uses AWS SNS (Simple Notification Service) or a similar notification service to send the message to the user.
[1294] The above are the specific processing steps in this system.
[1295] (Application example 2)
[1296] Next, a description will be given of Application Example 2. In the following description, the data processing device 12 will be referred to as a "server" and the robot 414 will be referred to as a "terminal."
[1297] In recent years, there has been an increase in special frauds using telephones, and frauds targeting the elderly in particular have become a serious problem. Current security systems lack the means to accurately detect potential fraud in real time and send prompt and appropriate warnings to users. As a result, many users remain at high risk of becoming victims of fraud. Furthermore, few systems take the user's emotional state into account, and the accuracy of fraud detection is low, making it difficult for users to feel safe answering the phone. There is a need for a system that solves this problem.
[1298] The identification process by the identification processing unit 290 of the data processing device 12 in Application Example 2 is realized by the following means. In this invention, the server includes a means for recording telephone voice, a means for encrypting and transmitting the recorded telephone voice, and a means for decrypting the encrypted voice data. This enables secure and highly accurate analysis of telephone voice in real time and fraud detection that takes into account the user's emotional state. In addition, by including a means for quickly sending a warning message via push notification, SMS, email, or operator, it is possible to immediately warn the user of fraud and prevent damage before it occurs.
[1299] "Means for recording telephone voice" refers to technology that automatically records and safely stores the contents of a call when a user initiates a call.
[1300] "Means for encrypting and transmitting recorded telephone voice data" refers to a technology that encrypts recorded telephone voice data to protect it from unauthorized access or eavesdropping by third parties and transmits it to a designated server.
[1301] "Means for decrypting encrypted audio data" refers to a technique for restoring the transmitted encrypted audio data to the original audio data.
[1302] "Means for converting voice data into text data" refers to a technology that analyzes decoded voice data and converts its contents into text information.
[1303] The "means of analyzing the converted text data and determining the possibility of special fraud" is a technology that uses an algorithm to detect signs and patterns of special fraud based on the obtained text data.
[1304] "Means for analyzing the user's emotional state" refers to a technology that analyzes the user's emotions (tension, confusion, fear, etc.) from voice data during a call and determines that state.
[1305] The "means for correcting the determination of the possibility of special fraud based on the analysis results of the emotional state" is a technology for improving the accuracy of the determination of the possibility of special fraud by taking into account the emotional state of the user.
[1306] The "means for generating a warning message" is a technology that automatically generates a notification message to warn the user when it is determined that there is a high possibility of special fraud.
[1307] "Means for sending the generated warning message" refers to a communication technology for immediately delivering the generated warning message to the user, including push notification, SMS, email, or operator notification.
[1308] The purpose of the system of the present invention is to prevent special frauds using voice calls, and it records, encrypts, analyzes telephone voice, and generates and transmits warning messages. The specific system configuration and its program processing are described below.
[1309] System Configuration
[1310] The system of the present invention comprises the following components:
[1311] 1. User device: When a call is initiated, the audio is automatically recorded, and the recorded audio data is encrypted and sent to the server.
[1312] 2. Speech recognition server: Decrypts the encrypted voice data and converts the voice into text data.
[1313] 3. Generative AI analysis server: Analyzes text data to determine the possibility of special fraud. It also analyzes the user's emotional state and adjusts the judgment results based on the results.
[1314] 4. Notification Server: Generates warning messages as needed and sends them to users via push notification, SMS, email or operator.
[1315] Hardware and Software Used
[1316] Smartphone: iOS / Android devices
[1317] Cloud servers: Amazon Web Services (AWS) EC2, Amazon S3
[1318] Speech recognition technology: Google Speech-to-Text API
[1319] Generative AI model: OpenAI GPT-4 API
[1320] Sentiment analysis engine: IBM Watson Tone Analyzer
[1321] Notification service: Firebase Cloud Messaging (FCM)
[1322] Program Processing Overview
[1323] 1. User Device
[1324] The user device automatically records the voice as soon as the call begins and encrypts the recording using a sophisticated encryption algorithm. The encrypted voice data is then immediately sent to a voice recognition server.
[1325] 2. Speech Recognition Server
[1326] The speech recognition server decrypts the received encrypted voice data and converts it into text using the Google Speech-to-Text API. The converted text data is then sent to the generative AI analysis server.
[1327] 3. Generative AI Analysis Server
[1328] The generative AI analysis server uses the OpenAI GPT-4 API to perform detailed analysis of the text data sent from the speech recognition server. To determine whether there are signs of fraud, it detects specific keywords and phrases and then uses IBM Watson Tone Analyzer to analyze the user's emotional state. Based on the results, it corrects the judgment results as necessary.
[1329] 4. Notification Server
[1330] The notification server receives the results of the generative AI analysis server and generates a warning message if it determines that the call is likely to be fraudulent, and sends it to the user via push notification, SMS, or email via Firebase Cloud Messaging (FCM). This warning message urges the user to end the call.
[1331] Examples and prompts
[1332] Examples:
[1333] For example, if a user receives a call while on a call saying, "My son was in a traffic accident and I need money immediately," the system will analyze the call based on the following prompt:
[1334] Example prompt sentence:
[1335] The user received a call. The text of the call is below:
[1336] "Hello, this is the police. Your son has been in a car accident. We need money."
[1337] The user's emotional state is "tense."
[1338] Determine if this is likely a scam based on the following criteria, and if so, generate a warning message.
[1339] conditions:
[1340] Keywords to detect possible fraud: traffic accident, need money, transfer, etc.
[1341] User's emotional state: nervous, confused, scared
[1342] Output format:
[1343] Possibility of fraud: [Yes / No]
[1344] Warning message: [Possible scam. Please end the call immediately / This call is safe.]
[1345] This system makes it possible to detect the risk of special fraud in real time and quickly protect users.
[1346] The flow of the specific processing in the application example 2 will be described with reference to FIG.
[1347] Step 1:
[1348] When a user receives a call, the user terminal automatically starts recording the call audio. The recorded audio data is encrypted using an encryption algorithm. The input is the call audio data, and the output is the encrypted audio data.
[1349] Step 2:
[1350] The encrypted voice data is sent from the user terminal to the voice recognition server, which then decrypts the received data. The input is the encrypted voice data, and the output is the decrypted voice data.
[1351] Step 3:
[1352] The speech recognition server converts the decoded speech data into text data using the Google Speech-to-Text API. The input is speech data and the output is text data.
[1353] Step 4:
[1354] The generative AI analysis server receives the text data sent from the speech recognition server and uses the OpenAI GPT-4 API to analyze the possibility of fraud by detecting keywords and phrases within the text data. The input is the text data, and the output is the analysis results.
[1355] Step 5:
[1356] The generative AI analysis server analyzes the user's emotional state using the IBM Watson Tone Analyzer, an emotion analysis engine, in addition to analyzing the text data. The input is voice data, and the output is the user's emotional state.
[1357] Step 6:
[1358] The generative AI analysis server corrects the judgment of the possibility of special fraud based on the emotion analysis results. For example, if the user is nervous, it will judge that there is a high possibility of fraud. The input is the analysis result and emotional state, and the output is the corrected judgment result.
[1359] Step 7:
[1360] Based on the corrected judgment results, the generative AI analysis server generates warning messages as necessary. The input is the corrected judgment results, and the output is the warning message.
[1361] Step 8:
[1362] The notification server receives the generated alert messages and uses Firebase Cloud Messaging (FCM) to send them to the user via push notification, SMS, email, or operator. The input is the alert message and the output is the notification to the user.
[1363] The specific processing unit 290 transmits the result of the specific processing to the robot 414. In the robot 414, the control unit 46A causes the speaker 240 and the control target 443 to output the result of the specific processing. The microphone 238 acquires voice indicating a user input regarding the result of the specific processing. The control unit 46A transmits voice data indicating the user input acquired by the microphone 238 to the data processing device 12. In the data processing device 12, the specific processing unit 290 acquires the voice data.
[1364] The data generation model 58 is a so-called generative AI (Artificial Intelligence). An example of the data generation model 58 is ChatGPT (Internet Search<URL: https: / / openai.com / blog / chatgpt> ), Gemini (Internet search <url: https: gemini.google.com ?hl="ja">) and other generation AIs. The data generation model 58 is obtained by performing deep learning on a neural network. A prompt including an instruction is input to the data generation model 58, and inference data such as voice data indicating voice, text data indicating text, and image data indicating an image is also input. The data generation model 58 performs inference on the input inference data in accordance with the instruction indicated by the prompt, and outputs the inference result in a data format such as voice data and text data. Here, inference refers to, for example, analysis, classification, prediction, and / or summarization.
[1365] In the above embodiment, an example was given in which the specific processing is performed by the data processing device 12, but the technology of the present disclosure is not limited to this, and the specific processing may be performed by the robot 414.
[1366] The emotion identification model 59 as an emotion engine may determine the user's emotion according to a specific mapping. Specifically, the emotion identification model 59 may determine the user's emotion according to an emotion map (see FIG. 9), which is a specific mapping. Similarly, the emotion identification model 59 may determine the robot's emotion, and the identification processing unit 290 may perform identification processing using the robot's emotion.
[1367] FIG. 9 is a diagram illustrating an emotion map 400 on which multiple emotions are mapped. In the emotion map 400, emotions are arranged in concentric circles radiating from the center. Emotions closer to the center of the concentric circles are more primitive. Emotions representing states and actions arising from a state of mind are arranged on the outer edges of the concentric circles. The concept of emotion includes both affect and mental states. Emotions generally generated from reactions occurring in the brain are arranged on the left side of the concentric circles. Emotions generally induced by situational judgment are arranged on the right side of the concentric circles. Emotions generally generated from reactions occurring in the brain and induced by situational judgment are arranged on the upper and lower sides of the concentric circles. Furthermore, the emotion of "pleasure" is arranged on the upper side of the concentric circles, and the emotion of "discomfort" is arranged on the lower side. In this way, in the emotion map 400, multiple emotions are mapped based on the structure by which emotions are generated, and emotions that tend to occur simultaneously are mapped close to each other.
[1368] These emotions are distributed in the 3 o'clock direction on emotion map 400, and typically fluctuate between relief and anxiety. In the right half of emotion map 400, situational awareness dominates over internal sensations, resulting in a sense of calm.
[1369] The inside of emotion map 400 represents what is going on in the mind, and the outside of emotion map 400 represents behavior, so the further you go outside emotion map 400, the more visible the emotions become (the more they are expressed in behavior).
[1370] Human emotions are based on various balances, such as posture and blood sugar levels. When these balances deviate from the ideal, a state of discomfort is indicated, and when they approach the ideal, a state of pleasure is indicated. Emotions can also be created for robots, automobiles, and motorcycles, based on various balances, such as posture and remaining battery life. When these balances deviate from the ideal, a state of discomfort is indicated, and when they approach the ideal, a state of pleasure is indicated. An emotion map can be generated, for example, based on Dr. Mitsuyoshi's emotion map (Research on Voice Emotion Recognition and Emotional Brain Physiological Signal Analysis Systems, Tokushima University, Doctoral Dissertation: https: / / ci.nii.ac.jp / naid / 500000375379). The left half of the emotion map lists emotions belonging to the "reaction" domain, where sensation is dominant. The right half of the emotion map lists emotions belonging to the "situation" domain, where situational awareness is dominant.
[1371] The emotion map defines two emotions that promote learning. One is a negative emotion on the situation side, around the middle of "repentance" or "reflection." In other words, this occurs when the robot experiences negative emotions such as "I never want to feel this way again" or "I don't want to be scolded again." The other is a positive emotion on the response side, around "desire." In other words, this occurs when the robot experiences positive feelings such as "I want more" or "I want to know more."
[1372] The emotion identification model 59 inputs user input into a pre-trained neural network, obtains emotion values indicating each emotion shown in the emotion map 400, and determines the user's emotion. This neural network is pre-trained based on multiple pieces of training data that are combinations of user input and emotion values indicating each emotion shown in the emotion map 400. Furthermore, this neural network is trained so that emotions that are located close to each other have similar values, as in the emotion map 900 shown in FIG. 10. FIG. 10 shows an example in which multiple emotions, "relieved," "calm," and "reassuring," have similar emotion values.
[1373] The system according to the present disclosure has been described above mainly with respect to the functions of the data processing device 12, but the system according to the present disclosure is not necessarily implemented on a server. The system according to the present disclosure may be implemented as a general information processing system. The present disclosure may be implemented, for example, as a software program running on a personal computer or an application running on a smartphone, etc. The method according to the present disclosure may be provided to users in the form of SaaS (Software as a Service).
[1374] In the above embodiment, an example was given in which the specific processing is performed by one computer 22, but the technology of the present disclosure is not limited to this, and the specific processing may be distributed and performed by a plurality of computers including the computer 22. For example, the data generation model 58 may be provided in an external device of the data processing device 12, and data may be generated in the external device in accordance with input data.
[1375] In the above embodiment, an example in which the specific processing program 56 is stored in the storage 32 has been described, but the technology of the present disclosure is not limited to this. For example, the specific processing program 56 may be stored in a portable, computer-readable, non-transitory storage medium such as a USB (Universal Serial Bus) memory. The specific processing program 56 stored in the non-transitory storage medium is installed in the computer 22 of the data processing device 12. The processor 28 executes the specific processing in accordance with the specific processing program 56.
[1376] Alternatively, the specific processing program 56 may be stored in a storage device such as a server connected to the data processing device 12 via the network 54, and the specific processing program 56 may be downloaded and installed on the computer 22 in response to a request from the data processing device 12.
[1377] It is not necessary to store all of the specific processing program 56 in a storage device such as a server connected to the data processing device 12 via the network 54, or to store all of the specific processing program 56 in the storage 32; only a portion of the specific processing program 56 may be stored.
[1378] The hardware resource for executing a specific process can be any of the following processors: An example of a processor is a CPU, which is a general-purpose processor that functions as a hardware resource for executing a specific process by executing software, i.e., a program. Another example of a processor is a dedicated electrical circuit, such as an FPGA (Field-Programmable Gate Array), a PLD (Programmable Logic Device), or an ASIC (Application Specific Integrated Circuit), which is a processor with a circuit configuration designed specifically for executing a specific process. Each processor has built-in or connected memory, and each processor uses the memory to execute the specific process.
[1379] The hardware resource that executes the specific processing may be configured with one of these various processors, or may be configured with a combination of two or more processors of the same or different types (for example, a combination of multiple FPGAs, or a combination of a CPU and an FPGA). Also, the hardware resource that executes the specific processing may be a single processor.
[1380] As an example of a system configured with a single processor, first, one processor is configured by combining one or more CPUs and software, and this processor functions as a hardware resource that executes a specific process. Second, there is a system that uses a processor that realizes the functions of an entire system including multiple hardware resources that execute a specific process on a single IC chip, as typified by SoC (System-on-a-chip). In this way, a specific process is realized using one or more of the above-mentioned various processors as hardware resources.
[1381] Furthermore, the hardware structure of these various processors can be, more specifically, an electric circuit that combines circuit elements such as semiconductor devices. The specific processing described above is merely an example. Therefore, it goes without saying that unnecessary steps may be deleted, new steps may be added, or the processing order may be rearranged, without departing from the spirit of the invention.
[1382] The above-described description and illustrations are a detailed explanation of the parts related to the technology of the present disclosure and are merely an example of the technology of the present disclosure. For example, the above description of the configuration, functions, actions, and effects is an explanation of an example of the configuration, functions, actions, and effects of the parts related to the technology of the present disclosure. Therefore, it goes without saying that unnecessary parts may be deleted, new elements may be added, or replacements may be made to the above-described description and illustrations within the scope of the gist of the technology of the present disclosure. Furthermore, to avoid confusion and facilitate understanding of the parts related to the technology of the present disclosure, the above-described description and illustrations omit explanations of common technical knowledge that do not require particular explanation to enable the implementation of the technology of the present disclosure.
[1383] All publications, patent applications, and technical standards mentioned in this specification are herein incorporated by reference to the same extent as if each individual publication, patent application, or technical standard was specifically and individually indicated to be incorporated by reference.
[1384] The following is further disclosed regarding the above embodiment.
[1385] (Claim 1)
[1386] means for recording telephone audio;
[1387] means for converting the recorded telephone voice into text data;
[1388] A means for analyzing the converted text data and determining the possibility of special fraud;
[1389] means for generating a warning message based on the determination;
[1390] means for transmitting the generated warning message;
[1391] Security systems including:
[1392] (Claim 2)
[1393] 10. The security system of claim 1, further comprising means for encrypting and transmitting the recorded telephone voice.
[1394] (Claim 3)
[1395] 10. The security system of claim 1, further comprising means for sending the alert message by push notification, SMS, email, or by an operator.
[1396] "Example 1"
[1397] (Claim 1)
[1398] A means for detecting an incoming call in a user terminal and automatically recording the call audio;
[1399] means for encrypting recorded call audio data;
[1400] means for transmitting the encrypted voice data to a voice recognition server;
[1401] A means for the voice recognition server to decrypt the encrypted voice data and convert it into text data;
[1402] A generative AI analysis server analyzes text data and determines whether it is a special fraud.
[1403] means for generating a warning message based on the determination;
[1404] means for sending the generated warning message to a notification server and notifying a user;
[1405] A system including:
[1406] (Claim 2)
[1407] 10. The system of claim 1, wherein the recorded call audio data is encrypted before being transmitted.
[1408] (Claim 3)
[1409] 10. The system of claim 1, wherein the alert message is sent by push notification, short message service, email, or communication means.
[1410] "Application Example 1"
[1411] (Claim 1)
[1412] means for recording telephone audio;
[1413] means for converting the recorded telephone voice into text data;
[1414] A means for analyzing the converted text data and determining the possibility of special fraud;
[1415] means for generating a warning message based on the determination;
[1416] means for transmitting the generated warning message;
[1417] a means for analyzing text data using a generative AI model;
[1418] a means for initiating recording in real time on a user terminal;
[1419] A means to generate and send alert messages via push notifications, SMS, and email;
[1420] A system including:
[1421] (Claim 2)
[1422] 10. The system of claim 1, further comprising means for encrypting and transmitting the recorded telephone voice.
[1423] (Claim 3)
[1424] 10. The system of claim 1, further comprising means for sending the alert message by push notification, SMS, email, or operator.
[1425] "Example 2: Combining Emotion Engines"
[1426] (Claim 1)
[1427] means for recording telephone audio;
[1428] means for encrypting and transmitting the recorded telephone voice;
[1429] A means for converting the encrypted telephone voice into text data;
[1430] A means for analyzing the converted text data and determining the possibility of special fraud;
[1431] means for analyzing the emotional state of a user and correcting the analysis result;
[1432] means for generating a warning message based on the corrected result;
[1433] means for transmitting the generated warning message;
[1434] A system including:
[1435] (Claim 2)
[1436] 10. The system of claim 1, further comprising means for sending the alert message by push notification, SMS, email, or text message.
[1437] (Claim 3)
[1438] 10. The system of claim 1, further comprising means for analyzing telephone speech using a machine learning algorithm to detect specialized fraud patterns.
[1439] "Application example 2 when combining emotion engines"
[1440] (Claim 1)
[1441] means for recording telephone audio;
[1442] means for encrypting and transmitting the recorded telephone voice;
[1443] means for decrypting the encrypted audio data;
[1444] means for converting the decoded audio data into text data;
[1445] A means for analyzing the converted text data and determining the possibility of special fraud;
[1446] means for analyzing the emotional state of a user;
[1447] A means for correcting a determination of the possibility of special fraud based on the analysis result of the emotional state;
[1448] means for generating a warning message based on the corrected determination;
[1449] means for transmitting the generated warning message;
[1450] A system including:
[1451] (Claim 2)
[1452] 10. The system of claim 1, further comprising means for sending the alert message by push notification, SMS, email, or operator.
[1453] (Claim 3)
[1454] 10. The system of claim 1, further comprising means for detecting tension, confusion, or fear in the user based on the audio data when analyzing the emotional state. [Explanation of symbols]
[1455] 10, 210, 310, 410 Data Processing Systems 12 Data Processing Device 14 Smart Devices 214 Smart Glasses 314 Headset-type terminal 414 Robot< / url:> < / url:> < / url:> < / url:>
Claims
1. means for recording telephone audio; means for converting the recorded telephone voice into text data; A means for analyzing the converted text data and determining the possibility of special fraud; means for generating a warning message based on the determination; means for transmitting the generated warning message; Security systems including:
2. 2. The security system of claim 1, further comprising means for encrypting and transmitting the recorded telephone voice.
3. The security system of claim 1 , further comprising means for sending the warning message by push notification, SMS, email, or operator.
Citation Information
Patent Citations
Persona chatbot control method and system
JP2022180282A