system

A system that converts voice data to text and uses AI to detect fraudulent calls, offering real-time protection and improving its accuracy through user feedback, addresses the evolving challenge of telephone fraud.

JP2026074987APending Publication Date: 2026-05-07SOFTBANK GROUP CORP
View PDF 1 Cites 0 Cited by

Patent Information

Authority / Receiving Office
JP · JP
Patent Type
Applications
Current Assignee / Owner
SOFTBANK GROUP CORP
Filing Date
2024-10-21
Publication Date
2026-05-07

AI Technical Summary

Technical Problem

Conventional methods for preventing fraud through telephone calls are becoming less effective as fraud techniques evolve, necessitating a system that can automatically detect and respond to fraudulent activities in real time.

Method used

A system that acquires voice data in real time, converts it into text, analyzes the text for fraudulent activity, and sends warnings or automatically terminates the call if necessary, utilizing speech recognition and generative AI models.

Benefits of technology

Provides immediate protection against fraud by accurately identifying suspicious calls and allowing users to safely end them, with the system improving its detection capabilities through user feedback.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure 2026074987000001_ABST
    Figure 2026074987000001_ABST
Patent Text Reader

Abstract

We provide the system. [Solution] A means of acquiring audio data in real time, Means for converting the acquired audio data into text data, A means for analyzing the aforementioned text data to evaluate the possibility of fraudulent activity, A means for outputting a warning based on the aforementioned evaluation, A means of terminating a call in response to the aforementioned warning, A system that includes this.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The technology of the present disclosure relates to a system.

Background Art

[0002] Patent Document 1 discloses a persona chatbot control method performed by at least one processor, including steps of receiving a user utterance, adding the user utterance to a prompt including an instruction sentence related to an explanation of a chatbot character, encoding the prompt, and inputting the encoded prompt into a language model to generate a chatbot utterance in response to the user utterance.

Prior Art Documents

Patent Documents

[0003]

Patent Document 1

Summary of the Invention

Problems to be Solved by the Invention

[0004] In order to protect individuals from various illegal acts in society, especially fraud through telephone calls, effective measures to prevent damage are required. However, the methods of fraud have become more sophisticated year by year, and conventional measures for suppression often have limitations. In such a situation, it is necessary to develop a system that can automatically detect illegal acts in real time and respond promptly.

Means for Solving the Problems

[0005] This invention provides a means for acquiring voice data in real time and converting it into text data. Furthermore, it includes means for analyzing the text data to evaluate the possibility of fraudulent activity, and means for sending a warning message via electronic communication means if fraudulent activity is suspected. In addition, the system includes means for automatically terminating the call based on this warning, thereby ensuring the safety of the user.

[0006] "Audio data" refers to the digital representation of audio information exchanged over a communication line.

[0007] "Real-time" refers to a process or response that occurs immediately after an event takes place, without any delay.

[0008] "Text data" refers to digital information converted into character data, in a format that allows for easy analysis and searching.

[0009] "Analysis" refers to the process of breaking down and examining data and information to reveal the underlying patterns and characteristics.

[0010] "Misconduct" is a concept that includes actions that violate laws or regulations, or actions that are ethically unacceptable.

[0011] "Evaluation" refers to judging the value or nature of something based on certain criteria or indicators.

[0012] A "warning" is an act or message that provides advance notice or caution when there is a potential danger or problem.

[0013] "Electronic communication means" refers to means of transmitting information using electronic methods, including, for example, email and messaging services.

[0014] "Ending a call" refers to the action of intentionally stopping an ongoing voice call and disconnecting the connection. [Brief explanation of the drawing]

[0015] [Figure 1] It is a conceptual diagram showing an example of the configuration of a data processing system according to the first embodiment. [Figure 2] It is a conceptual diagram showing an example of the main functions of a data processing device and a smart device according to the first embodiment. [Figure 3] It is a conceptual diagram showing an example of the configuration of a data processing system according to the second embodiment. [Figure 4] It is a conceptual diagram showing an example of the main functions of a data processing device and smart glasses according to the second embodiment. [Figure 5] It is a conceptual diagram showing an example of the configuration of a data processing system according to the third embodiment. [Figure 6] It is a conceptual diagram showing an example of the main functions of a data processing device and a headset-type terminal according to the third embodiment. [Figure 7] It is a conceptual diagram showing an example of the configuration of a data processing system according to the fourth embodiment. [Figure 8] It is a conceptual diagram showing an example of the main functions of a data processing device and a robot according to the fourth embodiment. [Figure 9] Shows an emotion map to which multiple emotions are mapped. [Figure 10] Shows an emotion map to which multiple emotions are mapped. [Figure 11] It is a sequence diagram showing the processing flow of the data processing system in Example 1. [Figure 12] It is a sequence diagram showing the processing flow of the data processing system in Application Example 1. [Figure 13] It is a sequence diagram showing the processing flow of the data processing system in Example 2 when an emotion engine is combined. [Figure 14] It is a sequence diagram showing the processing flow of the data processing system in Application Example 2 when an emotion engine is combined.

Embodiments for Carrying Out the Invention

[0016] An example of an embodiment of the system according to the technology of the present disclosure will be described below with reference to the accompanying drawings.

[0017] First, the terms used in the following description will be explained.

[0018] In the following embodiments, the numbered processor (hereinafter simply referred to as "processor") may be a single arithmetic unit or a combination of multiple arithmetic units. Also, the processor may be a single type of arithmetic unit or a combination of multiple types of arithmetic units. Examples of arithmetic units include a CPU (Central Processing Unit), a GPU (Graphics Processing Unit), a GPGPU (General-Purpose computing on Graphics Processing Units), an APU (Accelerated Processing Unit), and the like.

[0019] In the following embodiments, the numbered RAM (Random Access Memory) is a memory in which information is temporarily stored and is used as a work memory by the processor.

[0020] In the following embodiments, the numbered storage is one or more non-volatile storage devices that store various programs and various parameters, etc. Examples of non-volatile storage devices include flash memory (SSD (Solid State Drive)), magnetic disks (e.g., hard disks), or magnetic tapes, etc.

[0021] In the following embodiments, the signed communication interface (I / F) is an interface that includes a communication processor and an antenna, etc. The communication interface manages communication between multiple computers. Examples of communication standards applicable to the communication interface include wireless communication standards such as 5G (5th Generation Mobile Communication System), Wi-Fi (registered trademark), or Bluetooth (registered trademark).

[0022] In the following embodiments, "A and / or B" is synonymous with "at least one of A and B." That is, "A and / or B" means that it may be A alone, or B alone, or a combination of A and B. Furthermore, in this specification, the same concept as "A and / or B" applies when expressing three or more things linked by "and / or."

[0023] [First Embodiment]

[0024] Figure 1 shows an example of the configuration of the data processing system 10 according to the first embodiment.

[0025] As shown in Figure 1, the data processing system 10 includes a data processing device 12 and a smart device 14. An example of the data processing device 12 is a server.

[0026] The data processing device 12 comprises a computer 22, a database 24, and a communication interface 26. The computer 22 is an example of a "computer" related to the technology of this disclosure. The computer 22 comprises a processor 28, RAM 30, and storage 32. The processor 28, RAM 30, and storage 32 are connected to a bus 34. The database 24 and the communication interface 26 are also connected to the bus 34. The communication interface 26 is connected to a network 54. An example of the network 54 is a WAN (Wide Area Network) and / or a LAN (Local Area Network).

[0027] The smart device 14 comprises a computer 36, a reception device 38, an output device 40, a camera 42, and a communication interface 44. The computer 36 comprises a processor 46, RAM 48, and storage 50. The processor 46, RAM 48, and storage 50 are connected to a bus 52. The reception device 38, output device 40, and camera 42 are also connected to the bus 52.

[0028] The reception device 38 is equipped with a touch panel 38A and a microphone 38B, etc., and receives user input. The touch panel 38A receives user input by detecting contact with an object (e.g., a pen or finger). The microphone 38B receives user input by detecting the user's voice. The control unit 46A transmits data indicating the user input received by the touch panel 38A and microphone 38B to the data processing device 12. In the data processing device 12, the specific processing unit 290 acquires the data indicating the user input.

[0029] The output device 40 includes a display 40A and a speaker 40B, and presents data to the user 20 by outputting the data in a form perceptible to the user 20 (e.g., audio and / or text). The display 40A displays visible information such as text and images according to instructions from the processor 46. The speaker 40B outputs audio according to instructions from the processor 46. The camera 42 is a small digital camera equipped with an optical system such as a lens, aperture, and shutter, and an image sensor such as a CMOS (Complementary Metal-Oxide-Semiconductor) image sensor or a CCD (Charge Coupled Device) image sensor.

[0030] Communication interface 44 is connected to network 54. Communication interfaces 44 and 26 are responsible for the exchange of various types of information between processor 46 and processor 28 via network 54.

[0031] Figure 2 shows an example of the main functions of the data processing device 12 and the smart device 14.

[0032] As shown in Figure 2, in the data processing device 12, a specific processing is performed by the processor 28. A specific processing program 56 is stored in the storage 32. The specific processing program 56 is an example of a "program" related to the technology of this disclosure. The processor 28 reads the specific processing program 56 from the storage 32 and executes the read specific processing program 56 on the RAM 30. The specific processing is realized by the processor 28 operating as a specific processing unit 290 according to the specific processing program 56 executed on the RAM 30.

[0033] The storage 32 stores the data generation model 58 and the emotion identification model 59. The data generation model 58 and the emotion identification model 59 are used by the identification processing unit 290.

[0034] In the smart device 14, the processor 46 performs the reception output processing. The storage 50 stores the reception output program 60. The reception output program 60 is used in conjunction with a specific processing program 56 by the data processing system 10. The processor 46 reads the reception output program 60 from the storage 50 and executes the read reception output program 60 on the RAM 48. The reception output processing is realized by the processor 46 operating as a control unit 46A according to the reception output program 60 executed on the RAM 48.

[0035] Next, the specific processing performed by the specific processing unit 290 of the data processing device 12 will be described. In the following description, the data processing device 12 will be referred to as the "server" and the smart device 14 as the "terminal".

[0036] This invention provides a system for processing voice data in real time during communication in order to protect individuals from fraud. This system consists of a terminal, a server, and a user.

[0037] The device has a function that records audio data in real time when a call is received, and then divides that data and sends it to a server. This provides a foundation for timely analysis of the audio from the phone calls that users receive on a daily basis.

[0038] The server performs a process of converting received audio data into text data using speech recognition technology. Subsequently, it analyzes the text data using a generative AI model and scores the likelihood of fraudulent activity. Based on this score, if the risk of fraudulent activity is high, a warning message is generated and sent to the user or registered third party via electronic communication. For example, if a suspicious request regarding a bank account is detected, a warning such as "This may be a scam, so please be careful" is immediately sent to the user's smartphone.

[0039] If a user receives a warning message, they will be given the option to manually end the call, or the call will be automatically terminated by instructions from the server. In this way, the system prevents users from becoming victims of fraud. In addition, after the call, users can provide feedback on whether or not fraud actually occurred, and this feedback information is used by the server to improve the accuracy of the generated AI model.

[0040] For example, if a user receives a phone call and the caller says, "Please tell me your credit card information immediately," the system transcribes this audio into text, and an AI model determines that it is highly likely to be a scam. This allows the user to receive a real-time warning, preventing them from becoming a victim. This entire process is automated, making the system usable by users without requiring advanced technical knowledge.

[0041] The following describes the processing flow.

[0042] Step 1:

[0043] The device records the audio of calls received by the user in real time. The recording is divided into short time units and immediately prepared as data packets.

[0044] Step 2:

[0045] The device encrypts the recorded audio data packets before sending them to the server. This process is designed to minimize latency.

[0046] Step 3:

[0047] The server converts the received audio data into text data using a speech recognition API. This conversion must be fast and highly accurate.

[0048] Step 4:

[0049] The server analyzes text data using a generative AI model and scores the likelihood of fraudulent activity. This model is designed to identify words and contexts associated with fraud.

[0050] Step 5:

[0051] If the server determines, based on the scoring results, that there is a high probability of fraudulent activity, it will generate a warning message and send it to the user and registered third parties via electronic means. This message will include a warning about fraudulent activity and information on specific countermeasures.

[0052] Step 6:

[0053] If the server determines that fraudulent activity has occurred, it will automatically disconnect the call by sending a call termination command to the terminal.

[0054] Step 7:

[0055] Users send feedback to the server via a simple method, such as the LINE app, to determine whether the call content was truly fraudulent. Based on this feedback, the server improves the detection accuracy of the AI ​​model.

[0056] (Example 1)

[0057] Next, we will describe Example 1. In the following description, the data processing device 12 will be referred to as the "server," and the smart device 14 will be referred to as the "terminal."

[0058] In the communications that individuals receive on a daily basis, there is a need to quickly and efficiently detect the risks of fraudulent activity and prevent its impact. Conventional systems lack sufficient means to identify fraud, and users may easily become involved in fraudulent activities. This invention aims to solve these problems and provide users with a secure communication environment.

[0059] The identification process performed by the identification processing unit 290 of the data processing device 12 in Example 1 is realized by the following means.

[0060] In this invention, the server includes means for acquiring voice information in real time, means for converting it into text information, and means for evaluating the risk of fraudulent activity. This makes it possible to quickly detect the possibility of fraudulent activity during communication and provide appropriate warnings to the user.

[0061] "Voice information" refers to voice data generated by the user during communication, and the purpose is to acquire this data in real time.

[0062] "Real-time" refers to processing information immediately at the moment the communication is taking place, meaning that information is acquired and processed without delay.

[0063] "Textual information" refers to text data generated from audio information through speech recognition technology.

[0064] "Assessing the risk of fraudulent activity" refers to the process of analyzing acquired textual information to identify and quantify the possibility of fraudulent activities such as scams.

[0065] "Generating a warning" means creating a message to alert a user when a high risk of fraudulent activity is detected.

[0066] "Terminating communication" refers to stopping the call or communication session in response to the risk of fraudulent activity detected during the communication.

[0067] "Electronic communication technology" refers to technologies that send and receive information via digital means such as the internet, SMS, and email, and is used for transmitting warning messages in this system.

[0068] This system is built on communication protection technologies that include terminals, servers, and users. It primarily handles the acquisition, conversion, analysis, evaluation, and generation of warnings for voice information.

[0069] When the device receives a phone call or other communication, it uses its built-in voice recording device to capture voice information in real time. This data is then divided into manageable sizes, encrypted, and securely transmitted to the server.

[0070] The server converts received audio information into text using speech recognition technology. This process utilizes speech recognition software such as Google® Speech-to-Text API. The text information is then analyzed by a generative AI model. This model receives the converted text information as prompts and performs a scoring system to assess the risk of fraudulent activity. If the scoring result exceeds a certain threshold, the server automatically generates a warning message and sends it to the user's device. The warning message informs the user of the potential for fraudulent activity and prompts them to take appropriate action.

[0071] The user reviews the received warning message and ends the call if necessary. Furthermore, the system improves the accuracy of the generated AI model through user feedback.

[0072] For example, if a user receives a request such as "Please tell me your credit card information" during an incoming call, the device immediately records the audio and sends it to the server. The server transcribes the audio into text and scores it as highly suspicious. As a result, a warning message such as "This call may be a scam" is immediately sent to the device, allowing the user to safely end the call.

[0073] An example of a prompt for a generative AI model might be: "Analyze the following text data and determine the risk of fraudulent activity. Rate the risk on a scale of 100." This prompt provides the generative AI model with criteria for effectively identifying fraud.

[0074] The flow of the specific processing in Example 1 will be explained using Figure 11.

[0075] Step 1:

[0076] The device automatically acquires audio information when the user receives a call. The call audio is used as input and converted into digital audio data by the device's audio recording device. This audio data is then divided into manageable sizes. The resulting divided audio data is encrypted and ready to be sent to the server. This enables secure, real-time data transfer.

[0077] Step 2:

[0078] The server receives audio data transmitted from the terminal. Encrypted audio data is provided to the server as input and converted into text using speech recognition technology (e.g., Google Speech-to-Text API). This process involves sophisticated calculations to convert speech into text. The resulting text is then sent to the next stage for analysis.

[0079] Step 3:

[0080] The server analyzes the character information converted using a generative AI model. The character information generated in the previous step is provided as input, and a prompt sentence is fed to the AI ​​model. According to this prompt sentence, the AI ​​model scores the likelihood of fraudulent activity. The output is the fraudulent activity risk score corresponding to each text fragment. This score serves as a basis for generating warning messages through several algorithms.

[0081] Step 4:

[0082] The server generates a warning message based on a fraud risk score. If a high risk score is detected, a warning message is generated, and the scoring result is used as input. The output is the generated warning message. This message is sent to the user's terminal or a related third party via electronic means.

[0083] Step 5:

[0084] The user reviews the received warning message. The warning message sent from the server is used as input. The user can manually end the call as needed, or it may be automatically terminated depending on the settings. The output provides a way to securely manage the call status and prevent further harm to the user. Furthermore, the user can contribute to improving the system's accuracy by providing feedback.

[0085] (Application Example 1)

[0086] Next, we will explain Application Example 1. In the following explanation, the data processing device 12 will be referred to as the "server," and the smart device 14 will be referred to as the "terminal."

[0087] It can quickly detect the risk of fraudulent activity during communications and provide a reliable way to protect individuals from fraud. Many people are exposed to the risk of fraud on a daily basis, and there is a need for immediate responses to suspicious requests during calls.

[0088] The specific processing performed by the specific processing unit 290 of the data processing device 12 in Application Example 1 is realized by the following means.

[0089] In this invention, the server includes means for instantly acquiring voice information, means for converting the acquired voice information into text information, and means for analyzing the text information to evaluate the risk of fraudulent activity. This makes it possible to evaluate the risk of fraudulent activity during communication in real time and provide individuals with prompt warnings and protection.

[0090] "Voice information" refers to voice data acquired during a call, which is analyzed in real time.

[0091] "Textual information" refers to data that is generated by analyzing audio information and expressing it as text.

[0092] "Fraudulent activity risk" is an indicator that shows the likelihood of fraudulent activities such as scams and illegal transactions occurring.

[0093] "Evaluation" refers to the act of judging the risk of fraudulent activity based on acquired information, using numerical or other formats.

[0094] A "warning" is information intended to inform users and stakeholders about the risk of fraudulent activity.

[0095] "Feedback" refers to information provided by users after a call, which is used to improve the model.

[0096] A "model" refers to a generative AI model used to assess the risk of fraudulent activity.

[0097] In order to implement this invention, it is necessary to construct a system that analyzes voice information and issues warnings through the cooperation of a terminal, a server, and a user.

[0098] First, when a user initiates a call, the device immediately captures audio information and sends it to the server. The device processes audio data in the background even during a call, so it operates without interfering with the user's actions.

[0099] The server immediately converts the received audio information into text using a speech recognition API (e.g., Google Cloud Speech-to-Text). This text information is then input into a generative AI model (e.g., OpenAI's GPT-4) to assess the risk of fraud. For example, if the message contains phrases like "Please tell me your credit card information quickly," the generative AI model analyzes it and determines that it is highly likely to be fraud.

[0100] If the evaluation determines that there is a high risk of fraudulent activity, the server will generate a warning and send it to the user's terminal via electronic communication. Upon receiving the warning, the user can either manually end the call, or in some cases, the server will automatically terminate the call.

[0101] Users can provide feedback after a call ends, and this feedback is collected on the server to help improve the accuracy of the generated AI model.

[0102] As a concrete example, a prompt such as, "Analyze the text of this conversation and determine the risk of fraud. It contains the statement, 'I urgently need to send money to save my grandmother.'" can be used. This helps improve the accuracy of the model for similar fraud scenarios.

[0103] The flow of a specific process in Application Example 1 will be explained using Figure 12.

[0104] Step 1:

[0105] The terminal captures audio information in real time at the start of a call. It takes audio data from the call as input and prepares it to be sent to the server at regular intervals. The output is a fragment of audio data ready to be sent to the server.

[0106] Step 2:

[0107] The server receives audio data fragments sent from the terminal. It takes these audio data fragments as input and converts them into text information using a speech recognition API. The output is text information derived from the audio data. This process utilizes speech recognition software and leverages natural language processing techniques.

[0108] Step 3:

[0109] The server uses a generative AI model to analyze the converted text information and assess the risk of fraudulent activity. Text information is passed to the generative AI model as input, and analysis is performed based on the prompt text. The output is evaluated in the form of a fraud risk score. This evaluation determines whether the risk exceeds a threshold.

[0110] Step 4:

[0111] The server sends a warning to the user's device if the fraud risk score is high. The server takes a risk score as input and creates a warning message via electronic communication. The output is the warning notification received by the user. The notification contains information indicating a high probability of fraud.

[0112] Step 5:

[0113] The user receives a warning and terminates the call manually or at the server's instruction. The input is a warning notification, and the user decides whether to continue the call. The output is the execution of the call termination. At this point, the call is disconnected either by pressing the call termination button or by the server's automatic operation.

[0114] Step 6:

[0115] After the call ends, the user provides feedback to the server. The server collects user experience information and opinions as input and sends them to the server. The output is feedback data used to improve the accuracy of the generated AI model. The feedback content is used as training data for the model.

[0116] Furthermore, an emotion engine that estimates the user's emotions may be incorporated. That is, the identification processing unit 290 may use the emotion identification model 59 to estimate the user's emotions and perform identification processing using the user's emotions.

[0117] This invention provides a system that prevents fraudulent activity by acquiring voice data during a call in real time and converting it into text data and sentiment data. The system simultaneously monitors the content of the call and the user's emotional state, and takes warnings and preventative measures based on the results.

[0118] The device records voice data in real time during a call and sends the audio to the server. The server converts this voice data into text data using a speech recognition API, and at the same time uses an emotion engine to identify the user's emotional state from the voice. At this time, it analyzes emotions from the tone and tempo of the user's voice, paying particular attention to cases where emotions such as anxiety or doubt are prominent.

[0119] The server's generated AI model analyzes keywords and phrases related to fraud using text data and scores the risk of fraud. Emotional information obtained by the emotion engine is also incorporated into the scoring; if the emotional state is unusual, the likelihood of fraudulent activity is highly rated. This dual analytical approach allows for more accurate risk assessment.

[0120] Based on the evaluation results, the server will send warning messages to the user and registered third parties as needed. If the sentiment engine determines that the situation is particularly dangerous, it can increase the intensity of the message and prompt action. Upon receiving this warning, the user can manually end the call, or the call may be automatically disconnected at the server's direction.

[0121] For example, if a user receives a suspicious phone call and is strongly urged to "check information immediately," the emotion engine will detect that the voice indicates anxiety, and text analysis will determine that it is highly likely to be a scam. As a result, the user will receive an enhanced warning such as "Urgent Warning: This call may be a scam. Take action immediately," prompting the user to act before they become a victim.

[0122] In this way, the present invention provides a system that enables more reliable detection and rapid response to fraudulent activity through a combination of text analysis and sentiment recognition.

[0123] The following describes the processing flow.

[0124] Step 1:

[0125] The device starts recording in real time when the user begins a call. This audio data is divided into small packets and adjusted for processing without delay.

[0126] Step 2:

[0127] The device encrypts the recorded audio data and sends it to the server. This transmission takes place over the network, ensuring security.

[0128] Step 3:

[0129] The server converts the received audio data into text data using a speech recognition API. This process requires that the audio information be converted into text information quickly and accurately.

[0130] Step 4:

[0131] The server simultaneously uses an emotion engine to identify the user's emotional state from the voice data. It analyzes voice tone, tempo, tension, etc., to identify emotions such as joy, anxiety, and anger.

[0132] Step 5:

[0133] The server's generated AI model analyzes text data and scores the likelihood of fraud. It performs risk assessment by detecting specific keywords and contexts.

[0134] Step 6:

[0135] Emotional information from the emotion engine is fed back into the scoring results. If the user's emotional state is different from normal, for example, if they are highly anxious or suspicious, it is determined that there is a high possibility of fraud.

[0136] Step 7:

[0137] Based on the evaluation results, the server generates a warning message and sends it to the user or registered third parties via electronic communication. In some cases, a stronger warning than usual may be issued.

[0138] Step 8:

[0139] The user reviews the received warning message and either manually ends the call or the server automatically disconnects it. This process prompts the user to take swift action.

[0140] Step 9:

[0141] After a call ends, the user sends feedback to the server indicating whether or not the call was a scam. This information is used to further improve the accuracy of the AI ​​model.

[0142] (Example 2)

[0143] Next, we will describe Example 2. In the following description, the data processing device 12 will be referred to as the "server" and the smart device 14 as the "terminal".

[0144] With fraudulent activities using communication methods on the rise, there is a need to monitor conversation content in real time and detect signs of fraud with high accuracy. Furthermore, by considering the emotional state of users, it is necessary to conduct more accurate risk assessments and to have means to quickly protect users from fraudulent activities.

[0145] The identification process performed by the identification processing unit 290 of the data processing device 12 in Example 2 is realized by the following means.

[0146] In this invention, the server includes means for acquiring voice information in real time, means for converting the acquired voice information into text information, means for identifying emotional states from the voice information, and means for analyzing the text information and emotional states to assess the risk of fraudulent activity. This makes it possible to accurately assess the risk of fraudulent activity from a user's conversation and quickly protect the user from potential fraudulent activity.

[0147] "Voice information" refers to digital data related to human voices obtained from phone calls, recordings, etc.

[0148] "Real-time" refers to a state where processing or actions are performed instantly without delay.

[0149] "Textual information" refers to data in text format obtained as a result of converting audio information.

[0150] "Emotional state" refers to the speaker's emotional state and psychological changes, as analyzed from audio information.

[0151] "Fraudulent activity" refers to actions and methods related to fraud or attempted fraud, including acts for fraudulent purposes in communications.

[0152] "Risk assessment" refers to the process of calculating the likelihood of fraudulent activity as a numerical value or evaluation score based on information analysis.

[0153] A "warning" refers to a message or notification intended to inform a user or a third party in advance that a risk exists.

[0154] "Communication" refers to the act or process of sending and receiving voice or digital information.

[0155] This invention is a system that processes voice information during a call in real time, identifies fraudulent activity, and protects the user. This system is realized through cooperation between a terminal and a server.

[0156] The device acquires audio information in real time when a user initiates a call. This function is performed by communication devices such as smartphones and internet phones. The acquired audio information is transmitted to the server via a secure communication protocol.

[0157] The server converts the received audio information into text using a speech recognition API (for example, a common speech recognition technology for converting speech to text). This text conversion process employs an efficient algorithm to minimize latency. Simultaneously, the server analyzes the audio information using an emotion analysis engine to identify the user's emotional state. This emotion analysis incorporates techniques to evaluate the tone, speed, and volume of the voice.

[0158] Next, the server uses a generative AI model to detect terms and context related to fraudulent activity from the textual information. This model is trained to be particularly sensitive to keywords and phrases related to fraud. At the same time, the user's emotional state is also taken into consideration and integrated into the risk assessment.

[0159] Based on the risk assessment performed by the server, a warning will be sent to the user if necessary. The warning will be electronically distributed to the terminal and any relevant third parties to prompt the user to take prompt action. This warning will allow the user to manually end the call, or in some cases, the server may automatically disconnect the call.

[0160] As a concrete example, consider a case where a fictitious scam call is received. If a phrase such as "You need to provide your personal information immediately" is detected, and the user's voice indicates anxiety, the system will send a warning to the user saying, "Caution: This call may be a scam. Please be very careful when answering."

[0161] An example of a prompt to be input to the generating AI model is, "Analyze the content and tone of voice of the call, identify potential fraudulent activity, and assess the risk." In this way, the present invention provides a comprehensive call monitoring function to support user safety.

[0162] The flow of the specific processing in Example 2 will be explained using Figure 13.

[0163] Step 1:

[0164] The device detects the start of a call and acquires audio information in real time. This audio information is collected through the microphone and stored as digital data within the device. The input is the analog audio signal during the call, and the output is the digitized audio data. The device transmits this data to a server via the internet.

[0165] Step 2:

[0166] The server receives audio data from the terminal and inputs it into a speech recognition API, where it converts it into text. The speech recognition software uses an acoustic model and a language model to perform the process of converting the audio data into corresponding text. The input is digital audio data, and the output is text data.

[0167] Step 3:

[0168] The server inputs the received audio data into an emotion analysis engine to identify the user's emotional state. This process analyzes characteristics such as tone, speed, and volume of the voice and classifies them into emotion categories (e.g., anger, anxiety, joy). The input is audio data, and the output is data indicating the user's emotional state.

[0169] Step 4:

[0170] The server uses a generation AI model to analyze the converted text information. This model detects specific keywords and phrases and scores the risk of fraudulent activity based on them. Furthermore, emotional state data is also reflected in this risk assessment, and the score is adjusted if unstable emotions are detected. The input is text data and emotional state data, and the output is a fraudulent activity risk score.

[0171] Step 5:

[0172] The server generates a warning message based on the risk score. If a high risk is detected, the warning is set as urgent and sent to the user's and registered third-party devices. Specifically, a message such as "This call shows signs of fraud. Please be careful" is generated. The input is the risk score, and the output is the warning message.

[0173] Step 6:

[0174] Depending on the warning message received by the user, the call can be terminated either manually or automatically by a disconnection command from the server. This provides swift protection from fraudulent activity. The input is the warning message, and the output is the call termination action.

[0175] (Application Example 2)

[0176] Next, we will explain application example 2. In the following explanation, the data processing device 12 will be referred to as the "server," and the smart device 14 will be referred to as the "terminal."

[0177] In recent years, fraud and illegal activities conducted via voice calls have been increasing, causing harm to many individuals and organizations. There is a need for a system that can detect such fraudulent activities in real time and issue rapid warnings. Furthermore, to prevent harm, highly accurate risk assessments that take into account changes in user emotions are necessary. However, existing systems struggle to meet these requirements with sufficient accuracy.

[0178] The specific processing performed by the specific processing unit 290 of the data processing device 12 in Application Example 2 is realized by the following means.

[0179] In this invention, the server includes means for acquiring voice information in real time, means for converting the acquired voice information into document data, means for analyzing the document data to evaluate the possibility of fraudulent activity, means for performing sentiment analysis and identifying the user's emotional state, and means for scoring the risk of fraudulent activity based on the emotional state. This makes it possible to evaluate the risk of fraud and other fraudulent activities with high accuracy and to quickly warn the user.

[0180] "Voice information" refers to sound signals acquired through phone calls or recordings, and is digital data that includes user speech and background sounds.

[0181] "Real-time" refers to a situation where processing and responses occur instantaneously, with no time lag between voice acquisition and analysis.

[0182] "Document data" refers to information in string format converted from audio information, and is content that has been digitized as text.

[0183] "Analysis" is the process of examining information in detail to understand its structure and meaning, and making decisions according to a specific purpose.

[0184] "Assessing the possibility of fraudulent activity" means using audio and documentary data to determine whether there is a risk of fraud or misconduct.

[0185] "Emotional analysis" is a technique that analyzes characteristics such as tone and rhythm of speech information to infer the speaker's emotional state.

[0186] "Emotional state" refers to information that indicates the psychological emotions expressed during a speaker's utterance, and includes feelings such as joy, anxiety, and anger.

[0187] "Risk scoring" is the process of quantifying the likelihood of fraudulent activity occurring and evaluating the degree of risk.

[0188] A "warning" is a notification or message intended to alert a user when a high risk is identified.

[0189] This section describes the embodiments for carrying out the invention. This invention is a system that evaluates the risk of fraud and illegal activities in real time through voice calls and issues warnings to users. The following shows the specific configuration and operating procedures for realizing this system.

[0190] The device acquires audio information during a voice call and captures it as a digital signal using the microphone. This audio information is immediately sent to a cloud server. The server converts the audio into document data using a speech recognition API (e.g., Google Speech-to-Text API). This process results in the conversation being obtained in text format.

[0191] Next, the server uses an emotion analysis engine (e.g., IBM Watson® Tone Analyzer) to identify the user's emotional state from the acquired audio data. This makes it possible to detect changes in the speaker's emotions in real time.

[0192] The server uses a generative AI model (e.g., OpenAI GPT model) to score the risk of fraudulent activity based on document data and sentiment analysis results. This model analyzes keywords and context within the text and combines them with emotional states to assess the risk.

[0193] As a concrete example, consider a scenario where a user receives a suspicious phone call. The server, through sentiment analysis, recognizes that the user is feeling uneasy and detects that the conversation contains dangerous keywords such as "banking information." As a result, the server generates a high risk score and sends a warning message to the user, such as "Urgent Warning: This call may be a scam. Take immediate action."

[0194] An example of a prompt to input into the generating AI model is: "Calculate a fraud score based on the text extracted from the audio: 'Please verify your bank information now.' Sentiment analysis indicates anxiety."

[0195] This system can detect fraudulent activity during calls with high accuracy and speed, and provide effective warnings to users.

[0196] The flow of a specific process in Application Example 2 will be explained using Figure 14.

[0197] Step 1:

[0198] The device acquires the user's voice information through the microphone during a voice call. This voice information is captured in real time as a digital signal. The input is voice data, and the output is a digitized voice signal. These signals are immediately transmitted to the server via the communication line.

[0199] Step 2:

[0200] The server uses a speech recognition API on the cloud to convert digital audio signals into document data. Specifically, it uses the Google Speech-to-Text API to convert audio into text format. The input is a digital audio signal, and the output is the corresponding document data.

[0201] Step 3:

[0202] The server provides the converted document data to the sentiment analysis engine to identify the user's emotional state. This process uses IBM Watson Tone Analyzer to analyze factors such as speech tone and rhythm. The input is document data, and the output is data indicating the user's emotional state.

[0203] Step 4:

[0204] The server uses a generative AI model to score the risk of fraudulent activity based on document data and sentiment data. Specifically, it uses the OpenAI GPT model to perform a risk assessment by combining the analysis of fraudulent keywords contained in the text with anomalies in sentiment. The input is document data and sentiment data, and the output is a fraud risk score.

[0205] Step 5:

[0206] The server generates a warning message based on the fraud risk score and notifies the user in real time. If the risk score is high, an enhanced warning message is sent to the device as a push notification. The input is the fraud risk score, and the output is the warning message.

[0207] Step 6:

[0208] The user receives a warning message from the server and decides whether to continue or end the call depending on the situation. Depending on the warning, the user is provided with a means to manually end the call. The input is the warning message, and the output is the user's response action.

[0209] The specific processing unit 290 transmits the result of the specific processing to the smart device 14. In the smart device 14, the control unit 46A causes the output device 40 to output the result of the specific processing. The microphone 38B acquires audio indicating user input for the result of the specific processing. The control unit 46A transmits the audio data indicating user input acquired by the microphone 38B to the data processing device 12. In the data processing device 12, the specific processing unit 290 acquires the audio data.

[0210] Data generation model 58 is a so-called generative AI (Artificial Intelligence). An example of data generation model 58 is ChatGPT (registered trademark) (Internet search).<URL: https: / / openai.com / blog / chatgpt> ), Gemini (registered trademark) (Internet search) <url: https: gemini.google.com ?hl="ja">Examples of generative AI include the following. The data generation model 58 is obtained by performing deep learning on a neural network. The data generation model 58 is input with prompts containing instructions, and with inference data such as audio data representing speech, text data representing text, and image data representing images. The data generation model 58 infers from the input inference data according to the instructions indicated by the prompts, and outputs the inference results in data formats such as audio data and text data. Here, inference refers to, for example, analysis, classification, prediction, and / or summarization.

[0211] In the above embodiment, an example was given in which specific processing is performed by the data processing device 12, but the technology of this disclosure is not limited thereto, and the specific processing may also be performed by the smart device 14.

[0212] [Second Embodiment]

[0213] Figure 3 shows an example of the configuration of the data processing system 210 according to the second embodiment.

[0214] As shown in Figure 3, the data processing system 210 includes a data processing device 12 and smart glasses 214. An example of the data processing device 12 is a server.

[0215] The data processing device 12 comprises a computer 22, a database 24, and a communication interface 26. The computer 22 is an example of a "computer" related to the technology of this disclosure. The computer 22 comprises a processor 28, RAM 30, and storage 32. The processor 28, RAM 30, and storage 32 are connected to a bus 34. The database 24 and the communication interface 26 are also connected to the bus 34. The communication interface 26 is connected to a network 54. An example of the network 54 is a WAN (Wide Area Network) and / or a LAN (Local Area Network).

[0216] The smart glasses 214 include a computer 36, a microphone 238, a speaker 240, a camera 42, and a communication interface 44. The computer 36 includes a processor 46, RAM 48, and storage 50. The processor 46, RAM 48, and storage 50 are connected to a bus 52. The microphone 238, speaker 240, and camera 42 are also connected to the bus 52.

[0217] The microphone 238 receives voice signals from the user 20 and receives instructions from the user 20. The microphone 238 captures the voice signals from the user 20, converts the captured voice into audio data, and outputs it to the processor 46. The speaker 240 outputs audio according to the instructions from the processor 46.

[0218] Camera 42 is a small digital camera equipped with an optical system including a lens, aperture, and shutter, and an image sensor such as a CMOS (Complementary Metal-Oxide-Semiconductor) image sensor or a CCD (Charge Coupled Device) image sensor, and captures images of the area around the user 20 (for example, an imaging range defined by a field of view equivalent to the width of a typical healthy person's field of vision).

[0219] Communication interface 44 is connected to network 54. Communication interfaces 44 and 26 are responsible for the exchange of various information between processor 46 and processor 28 via network 54. The exchange of various information between processor 46 and processor 28 using communication interfaces 44 and 26 is performed in a secure manner.

[0220] Figure 4 shows an example of the main functions of the data processing device 12 and the smart glasses 214. As shown in Figure 4, the data processing device 12 performs specific processing using the processor 28. The storage 32 stores the specific processing program 56.

[0221] The specific processing program 56 is an example of a "program" relating to the technology of this disclosure. The processor 28 reads the specific processing program 56 from the storage 32 and executes the read specific processing program 56 on the RAM 30. The specific processing is realized by the processor 28 operating as a specific processing unit 290 in accordance with the specific processing program 56 executed on the RAM 30.

[0222] The storage 32 stores the data generation model 58 and the emotion identification model 59. The data generation model 58 and the emotion identification model 59 are used by the identification processing unit 290.

[0223] In the smart glasses 214, the processor 46 performs the reception output processing. The storage 50 stores the reception output program 60. The processor 46 reads the reception output program 60 from the storage 50 and executes the read reception output program 60 on the RAM 48. The reception output processing is realized by the processor 46 operating as a control unit 46A according to the reception output program 60 executed on the RAM 48.

[0224] Next, the identification processing performed by the identification processing unit 290 of the data processing device 12 will be described. In the following description, the data processing device 12 will be referred to as the "server" and the smart glasses 214 will be referred to as the "terminal".

[0225] This invention provides a system for processing voice data in real time during communication in order to protect individuals from fraud. This system consists of a terminal, a server, and a user.

[0226] The device has a function that records audio data in real time when a call is received, and then divides that data and sends it to a server. This provides a foundation for timely analysis of the audio from the phone calls that users receive on a daily basis.

[0227] The server performs a process of converting received audio data into text data using speech recognition technology. Subsequently, it analyzes the text data using a generative AI model and scores the likelihood of fraudulent activity. Based on this score, if the risk of fraudulent activity is high, a warning message is generated and sent to the user or registered third party via electronic communication. For example, if a suspicious request regarding a bank account is detected, a warning such as "This may be a scam, so please be careful" is immediately sent to the user's smartphone.

[0228] If a user receives a warning message, they will be given the option to manually end the call, or the call will be automatically terminated by instructions from the server. In this way, the system prevents users from becoming victims of fraud. In addition, after the call, users can provide feedback on whether or not fraud actually occurred, and this feedback information is used by the server to improve the accuracy of the generated AI model.

[0229] For example, if a user receives a phone call and the caller says, "Please tell me your credit card information immediately," the system transcribes this audio into text, and an AI model determines that it is highly likely to be a scam. This allows the user to receive a real-time warning, preventing them from becoming a victim. This entire process is automated, making the system usable by users without requiring advanced technical knowledge.

[0230] The following describes the processing flow.

[0231] Step 1:

[0232] The device records the audio of calls received by the user in real time. The recording is divided into short time units and immediately prepared as data packets.

[0233] Step 2:

[0234] The device encrypts the recorded audio data packets before sending them to the server. This process is designed to minimize latency.

[0235] Step 3:

[0236] The server converts the received audio data into text data using a speech recognition API. This conversion must be fast and highly accurate.

[0237] Step 4:

[0238] The server analyzes text data using a generative AI model and scores the likelihood of fraudulent activity. This model is designed to identify words and contexts associated with fraud.

[0239] Step 5:

[0240] If the server determines, based on the scoring results, that there is a high probability of fraudulent activity, it will generate a warning message and send it to the user and registered third parties via electronic means. This message will include a warning about fraudulent activity and information on specific countermeasures.

[0241] Step 6:

[0242] If the server determines that fraudulent activity has occurred, it will automatically disconnect the call by sending a call termination command to the terminal.

[0243] Step 7:

[0244] Users send feedback to the server via a simple method, such as the LINE app, to determine whether the call content was truly fraudulent. Based on this feedback, the server improves the detection accuracy of the AI ​​model.

[0245] (Example 1)

[0246] Next, we will describe Example 1. In the following description, the data processing device 12 will be referred to as the "server," and the smart glasses 214 will be referred to as the "terminal."

[0247] In the communications that individuals receive on a daily basis, there is a need to quickly and efficiently detect the risks of fraudulent activity and prevent its impact. Conventional systems lack sufficient means to identify fraud, and users may easily become involved in fraudulent activities. This invention aims to solve these problems and provide users with a secure communication environment.

[0248] The identification process performed by the identification processing unit 290 of the data processing device 12 in Example 1 is realized by the following means.

[0249] In this invention, the server includes means for acquiring voice information in real time, means for converting it into text information, and means for evaluating the risk of fraudulent activity. This makes it possible to quickly detect the possibility of fraudulent activity during communication and provide appropriate warnings to the user.

[0250] "Voice information" refers to voice data generated by the user during communication, and the purpose is to acquire this data in real time.

[0251] "Real-time" refers to processing information immediately at the moment the communication is taking place, meaning that information is acquired and processed without delay.

[0252] "Textual information" refers to text data generated from audio information through speech recognition technology.

[0253] "Assessing the risk of fraudulent activity" refers to the process of analyzing acquired textual information to identify and quantify the possibility of fraudulent activities such as scams.

[0254] "Generating a warning" means creating a message to alert a user when a high risk of fraudulent activity is detected.

[0255] "Terminating communication" refers to stopping the call or communication session in response to the risk of fraudulent activity detected during the communication.

[0256] "Electronic communication technology" refers to technologies that send and receive information via digital means such as the internet, SMS, and email, and is used for transmitting warning messages in this system.

[0257] This system is built on communication protection technologies that include terminals, servers, and users. It primarily handles the acquisition, conversion, analysis, evaluation, and generation of warnings for voice information.

[0258] When the device receives a phone call or other communication, it uses its built-in voice recording device to capture voice information in real time. This data is then divided into manageable sizes, encrypted, and securely transmitted to the server.

[0259] The server converts the received audio information into text using speech recognition technology. This process utilizes speech recognition software such as the Google Speech-to-Text API. The text information is then analyzed by a generative AI model. This model receives the converted text information as prompts and performs a scoring system to assess the risk of fraudulent activity. If the scoring result exceeds a certain threshold, the server automatically generates a warning message and sends it to the user's device. The warning message informs the user of the potential for fraudulent activity and prompts them to take appropriate action.

[0260] The user reviews the received warning message and ends the call if necessary. Furthermore, the system improves the accuracy of the generated AI model through user feedback.

[0261] For example, if a user receives a request such as "Please tell me your credit card information" during an incoming call, the device immediately records the audio and sends it to the server. The server transcribes the audio into text and scores it as highly suspicious. As a result, a warning message such as "This call may be a scam" is immediately sent to the device, allowing the user to safely end the call.

[0262] An example of a prompt for a generative AI model might be: "Analyze the following text data and determine the risk of fraudulent activity. Rate the risk on a scale of 100." This prompt provides the generative AI model with criteria for effectively identifying fraud.

[0263] The flow of the specific processing in Example 1 will be explained using Figure 11.

[0264] Step 1:

[0265] The device automatically acquires audio information when the user receives a call. The call audio is used as input and converted into digital audio data by the device's audio recording device. This audio data is then divided into manageable sizes. The resulting divided audio data is encrypted and ready to be sent to the server. This enables secure, real-time data transfer.

[0266] Step 2:

[0267] The server receives audio data transmitted from the terminal. Encrypted audio data is provided to the server as input and converted into text using speech recognition technology (e.g., Google Speech-to-Text API). This process involves sophisticated calculations to convert speech into text. The resulting text is then sent to the next stage for analysis.

[0268] Step 3:

[0269] The server analyzes the character information converted using a generative AI model. The character information generated in the previous step is provided as input, and a prompt sentence is fed to the AI ​​model. According to this prompt sentence, the AI ​​model scores the likelihood of fraudulent activity. The output is the fraudulent activity risk score corresponding to each text fragment. This score serves as a basis for generating warning messages through several algorithms.

[0270] Step 4:

[0271] The server generates a warning message based on a fraud risk score. If a high risk score is detected, a warning message is generated, and the scoring result is used as input. The output is the generated warning message. This message is sent to the user's terminal or a related third party via electronic means.

[0272] Step 5:

[0273] The user reviews the received warning message. The warning message sent from the server is used as input. The user can manually end the call as needed, or it may be automatically terminated depending on the settings. The output provides a way to securely manage the call status and prevent further harm to the user. Furthermore, the user can contribute to improving the system's accuracy by providing feedback.

[0274] (Application Example 1)

[0275] Next, we will explain Application Example 1. In the following explanation, the data processing device 12 will be referred to as the "server," and the smart glasses 214 will be referred to as the "terminal."

[0276] It can quickly detect the risk of fraudulent activity during communications and provide a reliable way to protect individuals from fraud. Many people are exposed to the risk of fraud on a daily basis, and there is a need for immediate responses to suspicious requests during calls.

[0277] The specific processing by the specific processing unit 290 of the data processing apparatus 12 in Application Example 1 is realized by the following means.

[0278] In this invention, the server includes means for immediately acquiring voice information, means for converting the acquired voice information into character information, and means for analyzing the character information to evaluate the risk of improper behavior. Thereby, it becomes possible to evaluate the risk of improper behavior during communication in real time and provide prompt warnings and protection to individuals.

[0279] "Voice information" refers to voice data acquired during a call and is the target to be analyzed in real time.

[0280] "Character information" refers to data obtained by analyzing voice information and expressed as characters.

[0281] "Risk of improper behavior" is an indicator indicating the possibility of acts such as fraud and illegal transactions occurring.

[0282] "Evaluation" is an act of judging the risk of improper behavior in numerical or other forms based on the acquired information.

[0283] "Warning" is information for notifying users and related parties that there is a risk of improper behavior.

[0284] "Feedback" is information provided by the user after a call and is used to improve the model.

[0285] "Model" refers to the generative AI model used to evaluate the risk of improper behavior.

[0286] In order to implement this invention, it is necessary to construct a system that analyzes voice information and issues warnings through the cooperation of a terminal, a server, and a user.

[0287] First, when a user initiates a call, the device immediately captures audio information and sends it to the server. The device processes audio data in the background even during a call, so it operates without interfering with the user's actions.

[0288] The server immediately converts the received audio information into text using a speech recognition API (e.g., Google Cloud Speech-to-Text). This text information is then input into a generative AI model (e.g., OpenAI's GPT-4) to assess the risk of fraud. For example, if the message contains phrases like "Please tell me your credit card information quickly," the generative AI model analyzes it and determines that it is highly likely to be fraud.

[0289] If the evaluation determines that there is a high risk of fraudulent activity, the server will generate a warning and send it to the user's terminal via electronic communication. Upon receiving the warning, the user can either manually end the call, or in some cases, the server will automatically terminate the call.

[0290] Users can provide feedback after a call ends, and this feedback is collected on the server to help improve the accuracy of the generated AI model.

[0291] As a concrete example, a prompt such as, "Analyze the text of this conversation and determine the risk of fraud. It contains the statement, 'I urgently need to send money to save my grandmother.'" can be used. This helps improve the accuracy of the model for similar fraud scenarios.

[0292] The flow of a specific process in Application Example 1 will be explained using Figure 12.

[0293] Step 1:

[0294] The terminal captures audio information in real time at the start of a call. It takes audio data from the call as input and prepares it to be sent to the server at regular intervals. The output is a fragment of audio data ready to be sent to the server.

[0295] Step 2:

[0296] The server receives audio data fragments sent from the terminal. It takes these audio data fragments as input and converts them into text information using a speech recognition API. The output is text information derived from the audio data. This process utilizes speech recognition software and leverages natural language processing techniques.

[0297] Step 3:

[0298] The server uses a generative AI model to analyze the converted text information and assess the risk of fraudulent activity. Text information is passed to the generative AI model as input, and analysis is performed based on the prompt text. The output is evaluated in the form of a fraud risk score. This evaluation determines whether the risk exceeds a threshold.

[0299] Step 4:

[0300] The server sends a warning to the user's device if the fraud risk score is high. The server takes a risk score as input and creates a warning message via electronic communication. The output is the warning notification received by the user. The notification contains information indicating a high probability of fraud.

[0301] Step 5:

[0302] The user receives a warning and terminates the call manually or at the server's instruction. The input is a warning notification, and the user decides whether to continue the call. The output is the execution of the call termination. At this point, the call is disconnected either by pressing the call termination button or by the server's automatic operation.

[0303] Step 6:

[0304] After the call ends, the user provides feedback to the server. As input, the user's experience information and opinions are collected and sent to the server. The output is feedback data that is used to improve the accuracy of the generative AI model. The feedback content is used as learning data for the model.

[0305] Furthermore, an emotion engine for estimating the user's emotion may be combined. That is, the specific processing unit 290 may estimate the user's emotion using the emotion specific model 59 and perform specific processing using the user's emotion.

[0306] The present invention provides a system that prevents fraud by acquiring voice data during a call in real time and converting it into text data and emotion data. The system simultaneously monitors the content of the call and the user's emotional state, and takes warnings and avoidance measures based on the results.

[0307] The terminal records voice data in real time during the call and sends the voice to the server. The server converts this voice data into text data using a speech recognition API, and at the same time, uses an emotion engine to identify the user's emotional state from the voice. At this time, emotions are analyzed from the tone and tempo of the user's voice, and special attention is paid when emotions such as uneasiness and suspicion are prominent.

[0308] The generative AI model of the server analyzes keywords and expressions related to fraud using the text data and scores the risk of fraud. The emotion information obtained by the emotion engine is also incorporated into the scoring, and when the emotional state is different from normal, the possibility of fraudulent behavior is highly evaluated. By this dual analysis approach, a more accurate risk assessment can be performed.

[0309] Based on the evaluation results, the server will send warning messages to the user and registered third parties as needed. If the sentiment engine determines that the situation is particularly dangerous, it can increase the intensity of the message and prompt action. Upon receiving this warning, the user can manually end the call, or the call may be automatically disconnected at the server's direction.

[0310] For example, if a user receives a suspicious phone call and is strongly urged to "check information immediately," the emotion engine will detect that the voice indicates anxiety, and text analysis will determine that it is highly likely to be a scam. As a result, the user will receive an enhanced warning such as "Urgent Warning: This call may be a scam. Take action immediately," prompting the user to act before they become a victim.

[0311] In this way, the present invention provides a system that enables more reliable detection and rapid response to fraudulent activity through a combination of text analysis and sentiment recognition.

[0312] The following describes the processing flow.

[0313] Step 1:

[0314] The device starts recording in real time when the user begins a call. This audio data is divided into small packets and adjusted for processing without delay.

[0315] Step 2:

[0316] The device encrypts the recorded audio data and sends it to the server. This transmission takes place over the network, ensuring security.

[0317] Step 3:

[0318] The server converts the received audio data into text data using a speech recognition API. This process requires that the audio information be converted into text information quickly and accurately.

[0319] Step 4:

[0320] The server simultaneously uses an emotion engine to identify the user's emotional state from the voice data. It analyzes voice tone, tempo, tension, etc., to identify emotions such as joy, anxiety, and anger.

[0321] Step 5:

[0322] The server's generated AI model analyzes text data and scores the likelihood of fraud. It performs risk assessment by detecting specific keywords and contexts.

[0323] Step 6:

[0324] Emotional information from the emotion engine is fed back into the scoring results. If the user's emotional state is different from normal, for example, if they are highly anxious or suspicious, it is determined that there is a high possibility of fraud.

[0325] Step 7:

[0326] Based on the evaluation results, the server generates a warning message and sends it to the user or registered third parties via electronic communication. In some cases, a stronger warning than usual may be issued.

[0327] Step 8:

[0328] The user reviews the received warning message and either manually ends the call or the server automatically disconnects it. This process prompts the user to take swift action.

[0329] Step 9:

[0330] After a call ends, the user sends feedback to the server indicating whether or not the call was a scam. This information is used to further improve the accuracy of the AI ​​model.

[0331] (Example 2)

[0332] Next, we will describe Example 2. In the following description, the data processing device 12 will be referred to as the "server" and the smart glasses 214 will be referred to as the "terminal".

[0333] With fraudulent activities using communication methods on the rise, there is a need to monitor conversation content in real time and detect signs of fraud with high accuracy. Furthermore, by considering the emotional state of users, it is necessary to conduct more accurate risk assessments and to have means to quickly protect users from fraudulent activities.

[0334] The identification process performed by the identification processing unit 290 of the data processing device 12 in Example 2 is realized by the following means.

[0335] In this invention, the server includes means for acquiring voice information in real time, means for converting the acquired voice information into text information, means for identifying emotional states from the voice information, and means for analyzing the text information and emotional states to assess the risk of fraudulent activity. This makes it possible to accurately assess the risk of fraudulent activity from a user's conversation and quickly protect the user from potential fraudulent activity.

[0336] "Voice information" refers to digital data related to human voices obtained from phone calls, recordings, etc.

[0337] "Real-time" refers to a state where processing or actions are performed instantly without delay.

[0338] "Textual information" refers to data in text format obtained as a result of converting audio information.

[0339] "Emotional state" refers to the speaker's emotional state and psychological changes, as analyzed from audio information.

[0340] "Fraudulent activity" refers to actions and methods related to fraud or attempted fraud, including acts for fraudulent purposes in communications.

[0341] "Risk assessment" refers to the process of calculating the likelihood of fraudulent activity as a numerical value or evaluation score based on information analysis.

[0342] A "warning" refers to a message or notification intended to inform a user or a third party in advance that a risk exists.

[0343] "Communication" refers to the act or process of sending and receiving voice or digital information.

[0344] This invention is a system that processes voice information during a call in real time, identifies fraudulent activity, and protects the user. This system is realized through cooperation between a terminal and a server.

[0345] The device acquires audio information in real time when a user initiates a call. This function is performed by communication devices such as smartphones and internet phones. The acquired audio information is transmitted to the server via a secure communication protocol.

[0346] The server converts the received audio information into text using a speech recognition API (for example, a common speech recognition technology for converting speech to text). This text conversion process employs an efficient algorithm to minimize latency. Simultaneously, the server analyzes the audio information using an emotion analysis engine to identify the user's emotional state. This emotion analysis incorporates techniques to evaluate the tone, speed, and volume of the voice.

[0347] Next, the server uses a generative AI model to detect terms and context related to fraudulent activity from the textual information. This model is trained to be particularly sensitive to keywords and phrases related to fraud. At the same time, the user's emotional state is also taken into consideration and integrated into the risk assessment.

[0348] Based on the risk assessment performed by the server, a warning will be sent to the user if necessary. The warning will be electronically distributed to the terminal and any relevant third parties to prompt the user to take prompt action. This warning will allow the user to manually end the call, or in some cases, the server may automatically disconnect the call.

[0349] As a concrete example, consider a case where a fictitious scam call is received. If a phrase such as "You need to provide your personal information immediately" is detected, and the user's voice indicates anxiety, the system will send a warning to the user saying, "Caution: This call may be a scam. Please be very careful when answering."

[0350] An example of a prompt to be input to the generating AI model is, "Analyze the content and tone of voice of the call, identify potential fraudulent activity, and assess the risk." In this way, the present invention provides a comprehensive call monitoring function to support user safety.

[0351] The flow of the specific processing in Example 2 will be explained using Figure 13.

[0352] Step 1:

[0353] The device detects the start of a call and acquires audio information in real time. This audio information is collected through the microphone and stored as digital data within the device. The input is the analog audio signal during the call, and the output is the digitized audio data. The device transmits this data to a server via the internet.

[0354] Step 2:

[0355] The server receives audio data from the terminal and inputs it into a speech recognition API, where it converts it into text. The speech recognition software uses an acoustic model and a language model to perform the process of converting the audio data into corresponding text. The input is digital audio data, and the output is text data.

[0356] Step 3:

[0357] The server inputs the received audio data into an emotion analysis engine to identify the user's emotional state. This process analyzes characteristics such as tone, speed, and volume of the voice and classifies them into emotion categories (e.g., anger, anxiety, joy). The input is audio data, and the output is data indicating the user's emotional state.

[0358] Step 4:

[0359] The server uses a generation AI model to analyze the converted text information. This model detects specific keywords and phrases and scores the risk of fraudulent activity based on them. Furthermore, emotional state data is also reflected in this risk assessment, and the score is adjusted if unstable emotions are detected. The input is text data and emotional state data, and the output is a fraudulent activity risk score.

[0360] Step 5:

[0361] The server generates a warning message based on the risk score. If a high risk is detected, the warning is set as urgent and sent to the user's and registered third-party devices. Specifically, a message such as "This call shows signs of fraud. Please be careful" is generated. The input is the risk score, and the output is the warning message.

[0362] Step 6:

[0363] Depending on the warning message received by the user, the call can be terminated either manually or automatically by a disconnection command from the server. This provides swift protection from fraudulent activity. The input is the warning message, and the output is the call termination action.

[0364] (Application Example 2)

[0365] Next, we will explain application example 2. In the following explanation, the data processing device 12 will be referred to as the "server," and the smart glasses 214 will be referred to as the "terminal."

[0366] In recent years, fraud and illegal activities conducted via voice calls have been increasing, causing harm to many individuals and organizations. There is a need for a system that can detect such fraudulent activities in real time and issue rapid warnings. Furthermore, to prevent harm, highly accurate risk assessments that take into account changes in user emotions are necessary. However, existing systems struggle to meet these requirements with sufficient accuracy.

[0367] The specific processing performed by the specific processing unit 290 of the data processing device 12 in Application Example 2 is realized by the following means.

[0368] In this invention, the server includes means for acquiring voice information in real time, means for converting the acquired voice information into document data, means for analyzing the document data to evaluate the possibility of fraudulent activity, means for performing sentiment analysis and identifying the user's emotional state, and means for scoring the risk of fraudulent activity based on the emotional state. This makes it possible to evaluate the risk of fraud and other fraudulent activities with high accuracy and to quickly warn the user.

[0369] "Voice information" refers to sound signals acquired through phone calls or recordings, and is digital data that includes user speech and background sounds.

[0370] "Real-time" refers to a situation where processing and responses occur instantaneously, with no time lag between voice acquisition and analysis.

[0371] "Document data" refers to information in string format converted from audio information, and is content that has been digitized as text.

[0372] "Analysis" is the process of examining information in detail to understand its structure and meaning, and making decisions according to a specific purpose.

[0373] "Assessing the possibility of fraudulent activity" means using audio and documentary data to determine whether there is a risk of fraud or misconduct.

[0374] "Emotional analysis" is a technique that analyzes characteristics such as tone and rhythm of speech information to infer the speaker's emotional state.

[0375] "Emotional state" refers to information that indicates the psychological emotions expressed during a speaker's utterance, and includes feelings such as joy, anxiety, and anger.

[0376] "Risk scoring" is the process of quantifying the likelihood of fraudulent activity occurring and evaluating the degree of risk.

[0377] A "warning" is a notification or message intended to alert a user when a high risk is identified.

[0378] This section describes the embodiments for carrying out the invention. This invention is a system that evaluates the risk of fraud and illegal activities in real time through voice calls and issues warnings to users. The following shows the specific configuration and operating procedures for realizing this system.

[0379] The device acquires audio information during a voice call and captures it as a digital signal using the microphone. This audio information is immediately sent to a cloud server. The server converts the audio into document data using a speech recognition API (e.g., Google Speech-to-Text API). This process results in the conversation being obtained in text format.

[0380] Next, the server uses an emotion analysis engine (e.g., IBM Watson Tone Analyzer) to identify the user's emotional state from the acquired audio data. This makes it possible to detect changes in the speaker's emotions in real time.

[0381] The server uses a generative AI model (e.g., OpenAI GPT model) to score the risk of fraudulent activity based on document data and sentiment analysis results. This model analyzes keywords and context within the text and combines them with emotional states to assess the risk.

[0382] As a concrete example, consider a scenario where a user receives a suspicious phone call. The server, through sentiment analysis, recognizes that the user is feeling uneasy and detects that the conversation contains dangerous keywords such as "banking information." As a result, the server generates a high risk score and sends a warning message to the user, such as "Urgent Warning: This call may be a scam. Take immediate action."

[0383] An example of a prompt to input into the generating AI model is: "Calculate a fraud score based on the text extracted from the audio: 'Please verify your bank information now.' Sentiment analysis indicates anxiety."

[0384] This system can detect fraudulent activity during calls with high accuracy and speed, and provide effective warnings to users.

[0385] The flow of a specific process in Application Example 2 will be explained using Figure 14.

[0386] Step 1:

[0387] The device acquires the user's voice information through the microphone during a voice call. This voice information is captured in real time as a digital signal. The input is voice data, and the output is a digitized voice signal. These signals are immediately transmitted to the server via the communication line.

[0388] Step 2:

[0389] The server uses a speech recognition API on the cloud to convert digital audio signals into document data. Specifically, it uses the Google Speech-to-Text API to convert audio into text format. The input is a digital audio signal, and the output is the corresponding document data.

[0390] Step 3:

[0391] The server provides the converted document data to the sentiment analysis engine to identify the user's emotional state. This process uses IBM Watson Tone Analyzer to analyze factors such as speech tone and rhythm. The input is document data, and the output is data indicating the user's emotional state.

[0392] Step 4:

[0393] The server uses a generative AI model to score the risk of fraudulent activity based on document data and sentiment data. Specifically, it uses the OpenAI GPT model to perform a risk assessment by combining the analysis of fraudulent keywords contained in the text with anomalies in sentiment. The input is document data and sentiment data, and the output is a fraud risk score.

[0394] Step 5:

[0395] The server generates a warning message based on the fraud risk score and notifies the user in real time. If the risk score is high, an enhanced warning message is sent to the device as a push notification. The input is the fraud risk score, and the output is the warning message.

[0396] Step 6:

[0397] The user receives a warning message from the server and decides whether to continue or end the call depending on the situation. Depending on the warning, the user is provided with a means to manually end the call. The input is the warning message, and the output is the user's response action.

[0398] The specific processing unit 290 transmits the result of the specific processing to the smart glasses 214. In the smart glasses 214, the control unit 46A causes the speaker 240 to output the result of the specific processing. The microphone 238 acquires audio indicating user input for the result of the specific processing. The control unit 46A transmits the audio data indicating user input acquired by the microphone 238 to the data processing unit 12. In the data processing unit 12, the specific processing unit 290 acquires the audio data.

[0399] Data generation model 58 is a type of so-called generative AI (Artificial Intelligence). One example of data generation model 58 is ChatGPT (Internet search<URL: https: / / openai.com / blog / chatgpt> ), Gemini (Internet search) <url: https: gemini.google.com ?hl="ja">Examples of generative AI include the following. The data generation model 58 is obtained by performing deep learning on a neural network. The data generation model 58 is input with prompts containing instructions, and with inference data such as audio data representing speech, text data representing text, and image data representing images. The data generation model 58 infers from the input inference data according to the instructions indicated by the prompts, and outputs the inference results in data formats such as audio data and text data. Here, inference refers to, for example, analysis, classification, prediction, and / or summarization.

[0400] In the above embodiment, an example was given in which specific processing is performed by the data processing device 12, but the technology of this disclosure is not limited thereto, and the specific processing may also be performed by the smart glasses 214.

[0401] [Third Embodiment]

[0402] Figure 5 shows an example of the configuration of the data processing system 310 according to the third embodiment.

[0403] As shown in Figure 5, the data processing system 310 includes a data processing device 12 and a headset terminal 314. An example of the data processing device 12 is a server.

[0404] The data processing device 12 comprises a computer 22, a database 24, and a communication interface 26. The computer 22 is an example of a "computer" related to the technology of this disclosure. The computer 22 comprises a processor 28, RAM 30, and storage 32. The processor 28, RAM 30, and storage 32 are connected to a bus 34. The database 24 and the communication interface 26 are also connected to the bus 34. The communication interface 26 is connected to a network 54. An example of the network 54 is a WAN (Wide Area Network) and / or a LAN (Local Area Network).

[0405] The headset terminal 314 includes a computer 36, a microphone 238, a speaker 240, a camera 42, a communication interface 44, and a display 343. The computer 36 includes a processor 46, RAM 48, and storage 50. The processor 46, RAM 48, and storage 50 are connected to a bus 52. The microphone 238, speaker 240, camera 42, and display 343 are also connected to the bus 52.

[0406] The microphone 238 receives voice signals from the user 20 and receives instructions from the user 20. The microphone 238 captures the voice signals from the user 20, converts the captured voice into audio data, and outputs it to the processor 46. The speaker 240 outputs audio according to the instructions from the processor 46.

[0407] Camera 42 is a small digital camera equipped with an optical system including a lens, aperture, and shutter, and an image sensor such as a CMOS (Complementary Metal-Oxide-Semiconductor) image sensor or a CCD (Charge Coupled Device) image sensor, and captures images of the area around the user 20 (for example, an imaging range defined by a field of view equivalent to the width of a typical healthy person's field of vision).

[0408] Communication interface 44 is connected to network 54. Communication interfaces 44 and 26 are responsible for the exchange of various information between processor 46 and processor 28 via network 54. The exchange of various information between processor 46 and processor 28 using communication interfaces 44 and 26 is performed in a secure manner.

[0409] Figure 6 shows an example of the main functions of the data processing device 12 and the headset terminal 314. As shown in Figure 6, the data processing device 12 performs specific processing using the processor 28. The storage 32 stores the specific processing program 56.

[0410] The specific processing program 56 is an example of a "program" relating to the technology of this disclosure. The processor 28 reads the specific processing program 56 from the storage 32 and executes the read specific processing program 56 on the RAM 30. The specific processing is realized by the processor 28 operating as a specific processing unit 290 in accordance with the specific processing program 56 executed on the RAM 30.

[0411] The storage 32 stores the data generation model 58 and the emotion identification model 59. The data generation model 58 and the emotion identification model 59 are used by the identification processing unit 290.

[0412] In the headset terminal 314, the processor 46 performs the reception output processing. The storage 50 stores the reception output program 60. The processor 46 reads the reception output program 60 from the storage 50 and executes the read reception output program 60 on the RAM 48. The reception output processing is realized by the processor 46 operating as a control unit 46A according to the reception output program 60 executed on the RAM 48.

[0413] Next, the specific processing performed by the specific processing unit 290 of the data processing device 12 will be described. In the following description, the data processing device 12 will be referred to as the "server" and the headset terminal 314 will be referred to as the "terminal".

[0414] This invention provides a system for processing voice data in real time during communication in order to protect individuals from fraud. This system consists of a terminal, a server, and a user.

[0415] The device has a function that records audio data in real time when a call is received, and then divides that data and sends it to a server. This provides a foundation for timely analysis of the audio from the phone calls that users receive on a daily basis.

[0416] The server performs a process of converting received audio data into text data using speech recognition technology. Subsequently, it analyzes the text data using a generative AI model and scores the likelihood of fraudulent activity. Based on this score, if the risk of fraudulent activity is high, a warning message is generated and sent to the user or registered third party via electronic communication. For example, if a suspicious request regarding a bank account is detected, a warning such as "This may be a scam, so please be careful" is immediately sent to the user's smartphone.

[0417] If a user receives a warning message, they will be given the option to manually end the call, or the call will be automatically terminated by instructions from the server. In this way, the system prevents users from becoming victims of fraud. In addition, after the call, users can provide feedback on whether or not fraud actually occurred, and this feedback information is used by the server to improve the accuracy of the generated AI model.

[0418] For example, if a user receives a phone call and the caller says, "Please tell me your credit card information immediately," the system transcribes this audio into text, and an AI model determines that it is highly likely to be a scam. This allows the user to receive a real-time warning, preventing them from becoming a victim. This entire process is automated, making the system usable by users without requiring advanced technical knowledge.

[0419] The following describes the processing flow.

[0420] Step 1:

[0421] The device records the audio of calls received by the user in real time. The recording is divided into short time units and immediately prepared as data packets.

[0422] Step 2:

[0423] The device encrypts the recorded audio data packets before sending them to the server. This process is designed to minimize latency.

[0424] Step 3:

[0425] The server converts the received audio data into text data using a speech recognition API. This conversion must be fast and highly accurate.

[0426] Step 4:

[0427] The server analyzes text data using a generative AI model and scores the likelihood of fraudulent activity. This model is designed to identify words and contexts associated with fraud.

[0428] Step 5:

[0429] If the server determines, based on the scoring results, that there is a high probability of fraudulent activity, it will generate a warning message and send it to the user and registered third parties via electronic means. This message will include a warning about fraudulent activity and information on specific countermeasures.

[0430] Step 6:

[0431] If the server determines that fraudulent activity has occurred, it will automatically disconnect the call by sending a call termination command to the terminal.

[0432] Step 7:

[0433] Users send feedback to the server via a simple method, such as the LINE app, to determine whether the call content was truly fraudulent. Based on this feedback, the server improves the detection accuracy of the AI ​​model.

[0434] (Example 1)

[0435] Next, we will describe Example 1. In the following description, the data processing device 12 will be referred to as the "server," and the headset-type terminal 314 will be referred to as the "terminal."

[0436] In the communications that individuals receive on a daily basis, there is a need to quickly and efficiently detect the risks of fraudulent activity and prevent its impact. Conventional systems lack sufficient means to identify fraud, and users may easily become involved in fraudulent activities. This invention aims to solve these problems and provide users with a secure communication environment.

[0437] The identification process performed by the identification processing unit 290 of the data processing device 12 in Example 1 is realized by the following means.

[0438] In this invention, the server includes means for acquiring voice information in real time, means for converting it into text information, and means for evaluating the risk of fraudulent activity. This makes it possible to quickly detect the possibility of fraudulent activity during communication and provide appropriate warnings to the user.

[0439] "Voice information" refers to voice data generated by the user during communication, and the purpose is to acquire this data in real time.

[0440] "Real-time" refers to processing information immediately at the moment the communication is taking place, meaning that information is acquired and processed without delay.

[0441] "Textual information" refers to text data generated from audio information through speech recognition technology.

[0442] "Assessing the risk of fraudulent activity" refers to the process of analyzing acquired textual information to identify and quantify the possibility of fraudulent activities such as scams.

[0443] "Generating a warning" means creating a message to alert a user when a high risk of fraudulent activity is detected.

[0444] "Terminating communication" refers to stopping the call or communication session in response to the risk of fraudulent activity detected during the communication.

[0445] "Electronic communication technology" refers to technologies that send and receive information via digital means such as the internet, SMS, and email, and is used for transmitting warning messages in this system.

[0446] This system is built on communication protection technologies that include terminals, servers, and users. It primarily handles the acquisition, conversion, analysis, evaluation, and generation of warnings for voice information.

[0447] When the device receives a phone call or other communication, it uses its built-in voice recording device to capture voice information in real time. This data is then divided into manageable sizes, encrypted, and securely transmitted to the server.

[0448] The server converts the received audio information into text using speech recognition technology. This process utilizes speech recognition software such as the Google Speech-to-Text API. The text information is then analyzed by a generative AI model. This model receives the converted text information as prompts and performs a scoring system to assess the risk of fraudulent activity. If the scoring result exceeds a certain threshold, the server automatically generates a warning message and sends it to the user's device. The warning message informs the user of the potential for fraudulent activity and prompts them to take appropriate action.

[0449] The user reviews the received warning message and ends the call if necessary. Furthermore, the system improves the accuracy of the generated AI model through user feedback.

[0450] For example, if a user receives a request such as "Please tell me your credit card information" during an incoming call, the device immediately records the audio and sends it to the server. The server transcribes the audio into text and scores it as highly suspicious. As a result, a warning message such as "This call may be a scam" is immediately sent to the device, allowing the user to safely end the call.

[0451] An example of a prompt for a generative AI model might be: "Analyze the following text data and determine the risk of fraudulent activity. Rate the risk on a scale of 100." This prompt provides the generative AI model with criteria for effectively identifying fraud.

[0452] The flow of the specific processing in Example 1 will be explained using Figure 11.

[0453] Step 1:

[0454] The device automatically acquires audio information when the user receives a call. The call audio is used as input and converted into digital audio data by the device's audio recording device. This audio data is then divided into manageable sizes. The resulting divided audio data is encrypted and ready to be sent to the server. This enables secure, real-time data transfer.

[0455] Step 2:

[0456] The server receives audio data transmitted from the terminal. Encrypted audio data is provided to the server as input and converted into text using speech recognition technology (e.g., Google Speech-to-Text API). This process involves sophisticated calculations to convert speech into text. The resulting text is then sent to the next stage for analysis.

[0457] Step 3:

[0458] The server analyzes the character information converted using a generative AI model. The character information generated in the previous step is provided as input, and a prompt sentence is fed to the AI ​​model. According to this prompt sentence, the AI ​​model scores the likelihood of fraudulent activity. The output is the fraudulent activity risk score corresponding to each text fragment. This score serves as a basis for generating warning messages through several algorithms.

[0459] Step 4:

[0460] The server generates a warning message based on a fraud risk score. If a high risk score is detected, a warning message is generated, and the scoring result is used as input. The output is the generated warning message. This message is sent to the user's terminal or a related third party via electronic means.

[0461] Step 5:

[0462] The user reviews the received warning message. The warning message sent from the server is used as input. The user can manually end the call as needed, or it may be automatically terminated depending on the settings. The output provides a way to securely manage the call status and prevent further harm to the user. Furthermore, the user can contribute to improving the system's accuracy by providing feedback.

[0463] (Application Example 1)

[0464] Next, we will explain Application Example 1. In the following explanation, the data processing device 12 will be referred to as the "server," and the headset-type terminal 314 will be referred to as the "terminal."

[0465] It can quickly detect the risk of fraudulent activity during communications and provide a reliable way to protect individuals from fraud. Many people are exposed to the risk of fraud on a daily basis, and there is a need for immediate responses to suspicious requests during calls.

[0466] The specific processing performed by the specific processing unit 290 of the data processing device 12 in Application Example 1 is realized by the following means.

[0467] In this invention, the server includes means for instantly acquiring voice information, means for converting the acquired voice information into text information, and means for analyzing the text information to evaluate the risk of fraudulent activity. This makes it possible to evaluate the risk of fraudulent activity during communication in real time and provide individuals with prompt warnings and protection.

[0468] "Voice information" refers to voice data acquired during a call, which is analyzed in real time.

[0469] "Textual information" refers to data that is generated by analyzing audio information and expressing it as text.

[0470] "Fraudulent activity risk" is an indicator that shows the likelihood of fraudulent activities such as scams and illegal transactions occurring.

[0471] "Evaluation" refers to the act of judging the risk of fraudulent activity based on acquired information, using numerical or other formats.

[0472] A "warning" is information intended to inform users and stakeholders about the risk of fraudulent activity.

[0473] "Feedback" refers to information provided by users after a call, which is used to improve the model.

[0474] A "model" refers to a generative AI model used to assess the risk of fraudulent activity.

[0475] In order to implement this invention, it is necessary to construct a system that analyzes voice information and issues warnings through the cooperation of a terminal, a server, and a user.

[0476] First, when a user initiates a call, the device immediately captures audio information and sends it to the server. The device processes audio data in the background even during a call, so it operates without interfering with the user's actions.

[0477] The server immediately converts the received audio information into text using a speech recognition API (e.g., Google Cloud Speech-to-Text). This text information is then input into a generative AI model (e.g., OpenAI's GPT-4) to assess the risk of fraud. For example, if the message contains phrases like "Please tell me your credit card information quickly," the generative AI model analyzes it and determines that it is highly likely to be fraud.

[0478] If the evaluation determines that there is a high risk of fraudulent activity, the server will generate a warning and send it to the user's terminal via electronic communication. Upon receiving the warning, the user can either manually end the call, or in some cases, the server will automatically terminate the call.

[0479] Users can provide feedback after a call ends, and this feedback is collected on the server to help improve the accuracy of the generated AI model.

[0480] As a concrete example, a prompt such as, "Analyze the text of this conversation and determine the risk of fraud. It contains the statement, 'I urgently need to send money to save my grandmother.'" can be used. This helps improve the accuracy of the model for similar fraud scenarios.

[0481] The flow of a specific process in Application Example 1 will be explained using Figure 12.

[0482] Step 1:

[0483] The terminal captures audio information in real time at the start of a call. It takes audio data from the call as input and prepares it to be sent to the server at regular intervals. The output is a fragment of audio data ready to be sent to the server.

[0484] Step 2:

[0485] The server receives audio data fragments sent from the terminal. It takes these audio data fragments as input and converts them into text information using a speech recognition API. The output is text information derived from the audio data. This process utilizes speech recognition software and leverages natural language processing techniques.

[0486] Step 3:

[0487] The server uses a generative AI model to analyze the converted text information and assess the risk of fraudulent activity. Text information is passed to the generative AI model as input, and analysis is performed based on the prompt text. The output is evaluated in the form of a fraud risk score. This evaluation determines whether the risk exceeds a threshold.

[0488] Step 4:

[0489] The server sends a warning to the user's device if the fraud risk score is high. The server takes a risk score as input and creates a warning message via electronic communication. The output is the warning notification received by the user. The notification contains information indicating a high probability of fraud.

[0490] Step 5:

[0491] The user receives a warning and terminates the call manually or at the server's instruction. The input is a warning notification, and the user decides whether to continue the call. The output is the execution of the call termination. At this point, the call is disconnected either by pressing the call termination button or by the server's automatic operation.

[0492] Step 6:

[0493] After the call ends, the user provides feedback to the server. The server collects user experience information and opinions as input and sends them to the server. The output is feedback data used to improve the accuracy of the generated AI model. The feedback content is used as training data for the model.

[0494] Furthermore, an emotion engine that estimates the user's emotions may be incorporated. That is, the identification processing unit 290 may use the emotion identification model 59 to estimate the user's emotions and perform identification processing using the user's emotions.

[0495] This invention provides a system that prevents fraudulent activity by acquiring voice data during a call in real time and converting it into text data and sentiment data. The system simultaneously monitors the content of the call and the user's emotional state, and takes warnings and preventative measures based on the results.

[0496] The device records voice data in real time during a call and sends the audio to the server. The server converts this voice data into text data using a speech recognition API, and at the same time uses an emotion engine to identify the user's emotional state from the voice. At this time, it analyzes emotions from the tone and tempo of the user's voice, paying particular attention to cases where emotions such as anxiety or doubt are prominent.

[0497] The server's generated AI model analyzes keywords and phrases related to fraud using text data and scores the risk of fraud. Emotional information obtained by the emotion engine is also incorporated into the scoring; if the emotional state is unusual, the likelihood of fraudulent activity is highly rated. This dual analytical approach allows for more accurate risk assessment.

[0498] Based on the evaluation results, the server will send warning messages to the user and registered third parties as needed. If the sentiment engine determines that the situation is particularly dangerous, it can increase the intensity of the message and prompt action. Upon receiving this warning, the user can manually end the call, or the call may be automatically disconnected at the server's direction.

[0499] For example, if a user receives a suspicious phone call and is strongly urged to "check information immediately," the emotion engine will detect that the voice indicates anxiety, and text analysis will determine that it is highly likely to be a scam. As a result, the user will receive an enhanced warning such as "Urgent Warning: This call may be a scam. Take action immediately," prompting the user to act before they become a victim.

[0500] In this way, the present invention provides a system that enables more reliable detection and rapid response to fraudulent activity through a combination of text analysis and sentiment recognition.

[0501] The following describes the processing flow.

[0502] Step 1:

[0503] The device starts recording in real time when the user begins a call. This audio data is divided into small packets and adjusted for processing without delay.

[0504] Step 2:

[0505] The device encrypts the recorded audio data and sends it to the server. This transmission takes place over the network, ensuring security.

[0506] Step 3:

[0507] The server converts the received audio data into text data using a speech recognition API. This process requires that the audio information be converted into text information quickly and accurately.

[0508] Step 4:

[0509] The server simultaneously uses an emotion engine to identify the user's emotional state from the voice data. It analyzes voice tone, tempo, tension, etc., to identify emotions such as joy, anxiety, and anger.

[0510] Step 5:

[0511] The server's generated AI model analyzes text data and scores the likelihood of fraud. It performs risk assessment by detecting specific keywords and contexts.

[0512] Step 6:

[0513] Emotional information from the emotion engine is fed back into the scoring results. If the user's emotional state is different from normal, for example, if they are highly anxious or suspicious, it is determined that there is a high possibility of fraud.

[0514] Step 7:

[0515] Based on the evaluation results, the server generates a warning message and sends it to the user or registered third parties via electronic communication. In some cases, a stronger warning than usual may be issued.

[0516] Step 8:

[0517] The user reviews the received warning message and either manually ends the call or the server automatically disconnects it. This process prompts the user to take swift action.

[0518] Step 9:

[0519] After a call ends, the user sends feedback to the server indicating whether or not the call was a scam. This information is used to further improve the accuracy of the AI ​​model.

[0520] (Example 2)

[0521] Next, we will describe Example 2. In the following description, the data processing device 12 will be referred to as the "server," and the headset-type terminal 314 will be referred to as the "terminal."

[0522] With fraudulent activities using communication methods on the rise, there is a need to monitor conversation content in real time and detect signs of fraud with high accuracy. Furthermore, by considering the emotional state of users, it is necessary to conduct more accurate risk assessments and to have means to quickly protect users from fraudulent activities.

[0523] The identification process performed by the identification processing unit 290 of the data processing device 12 in Example 2 is realized by the following means.

[0524] In this invention, the server includes means for acquiring voice information in real time, means for converting the acquired voice information into text information, means for identifying emotional states from the voice information, and means for analyzing the text information and emotional states to assess the risk of fraudulent activity. This makes it possible to accurately assess the risk of fraudulent activity from a user's conversation and quickly protect the user from potential fraudulent activity.

[0525] "Voice information" refers to digital data related to human voices obtained from phone calls, recordings, etc.

[0526] "Real-time" refers to a state where processing or actions are performed instantly without delay.

[0527] "Textual information" refers to data in text format obtained as a result of converting audio information.

[0528] "Emotional state" refers to the speaker's emotional state and psychological changes, as analyzed from audio information.

[0529] "Fraudulent activity" refers to actions and methods related to fraud or attempted fraud, including acts for fraudulent purposes in communications.

[0530] "Risk assessment" refers to the process of calculating the likelihood of fraudulent activity as a numerical value or evaluation score based on information analysis.

[0531] A "warning" refers to a message or notification intended to inform a user or a third party in advance that a risk exists.

[0532] "Communication" refers to the act or process of sending and receiving voice or digital information.

[0533] This invention is a system that processes voice information during a call in real time, identifies fraudulent activity, and protects the user. This system is realized through cooperation between a terminal and a server.

[0534] The device acquires audio information in real time when a user initiates a call. This function is performed by communication devices such as smartphones and internet phones. The acquired audio information is transmitted to the server via a secure communication protocol.

[0535] The server converts the received audio information into text using a speech recognition API (for example, a common speech recognition technology for converting speech to text). This text conversion process employs an efficient algorithm to minimize latency. Simultaneously, the server analyzes the audio information using an emotion analysis engine to identify the user's emotional state. This emotion analysis incorporates techniques to evaluate the tone, speed, and volume of the voice.

[0536] Next, the server uses a generative AI model to detect terms and context related to fraudulent activity from the textual information. This model is trained to be particularly sensitive to keywords and phrases related to fraud. At the same time, the user's emotional state is also taken into consideration and integrated into the risk assessment.

[0537] Based on the risk assessment performed by the server, a warning will be sent to the user if necessary. The warning will be electronically distributed to the terminal and any relevant third parties to prompt the user to take prompt action. This warning will allow the user to manually end the call, or in some cases, the server may automatically disconnect the call.

[0538] As a concrete example, consider a case where a fictitious scam call is received. If a phrase such as "You need to provide your personal information immediately" is detected, and the user's voice indicates anxiety, the system will send a warning to the user saying, "Caution: This call may be a scam. Please be very careful when answering."

[0539] An example of a prompt to be input to the generating AI model is, "Analyze the content and tone of voice of the call, identify potential fraudulent activity, and assess the risk." In this way, the present invention provides a comprehensive call monitoring function to support user safety.

[0540] The flow of the specific processing in Example 2 will be explained using Figure 13.

[0541] Step 1:

[0542] The device detects the start of a call and acquires audio information in real time. This audio information is collected through the microphone and stored as digital data within the device. The input is the analog audio signal during the call, and the output is the digitized audio data. The device transmits this data to a server via the internet.

[0543] Step 2:

[0544] The server receives audio data from the terminal and inputs it into a speech recognition API, where it converts it into text. The speech recognition software uses an acoustic model and a language model to perform the process of converting the audio data into corresponding text. The input is digital audio data, and the output is text data.

[0545] Step 3:

[0546] The server inputs the received audio data into an emotion analysis engine to identify the user's emotional state. This process analyzes characteristics such as tone, speed, and volume of the voice and classifies them into emotion categories (e.g., anger, anxiety, joy). The input is audio data, and the output is data indicating the user's emotional state.

[0547] Step 4:

[0548] The server uses a generation AI model to analyze the converted text information. This model detects specific keywords and phrases and scores the risk of fraudulent activity based on them. Furthermore, emotional state data is also reflected in this risk assessment, and the score is adjusted if unstable emotions are detected. The input is text data and emotional state data, and the output is a fraudulent activity risk score.

[0549] Step 5:

[0550] The server generates a warning message based on the risk score. If a high risk is detected, the warning is set as urgent and sent to the user's and registered third-party devices. Specifically, a message such as "This call shows signs of fraud. Please be careful" is generated. The input is the risk score, and the output is the warning message.

[0551] Step 6:

[0552] Depending on the warning message received by the user, the call can be terminated either manually or automatically by a disconnection command from the server. This provides swift protection from fraudulent activity. The input is the warning message, and the output is the call termination action.

[0553] (Application Example 2)

[0554] Next, we will explain application example 2. In the following explanation, the data processing device 12 will be referred to as the "server," and the headset-type terminal 314 will be referred to as the "terminal."

[0555] In recent years, fraud and illegal activities conducted via voice calls have been increasing, causing harm to many individuals and organizations. There is a need for a system that can detect such fraudulent activities in real time and issue rapid warnings. Furthermore, to prevent harm, highly accurate risk assessments that take into account changes in user emotions are necessary. However, existing systems struggle to meet these requirements with sufficient accuracy.

[0556] The specific processing performed by the specific processing unit 290 of the data processing device 12 in Application Example 2 is realized by the following means.

[0557] In this invention, the server includes means for acquiring voice information in real time, means for converting the acquired voice information into document data, means for analyzing the document data to evaluate the possibility of fraudulent activity, means for performing sentiment analysis and identifying the user's emotional state, and means for scoring the risk of fraudulent activity based on the emotional state. This makes it possible to evaluate the risk of fraud and other fraudulent activities with high accuracy and to quickly warn the user.

[0558] "Voice information" refers to sound signals acquired through phone calls or recordings, and is digital data that includes user speech and background sounds.

[0559] "Real-time" refers to a situation where processing and responses occur instantaneously, with no time lag between voice acquisition and analysis.

[0560] "Document data" refers to information in string format converted from audio information, and is content that has been digitized as text.

[0561] "Analysis" is the process of examining information in detail to understand its structure and meaning, and making decisions according to a specific purpose.

[0562] "Assessing the possibility of fraudulent activity" means using audio and documentary data to determine whether there is a risk of fraud or misconduct.

[0563] "Emotional analysis" is a technique that analyzes characteristics such as tone and rhythm of speech information to infer the speaker's emotional state.

[0564] "Emotional state" refers to information that indicates the psychological emotions expressed during a speaker's utterance, and includes feelings such as joy, anxiety, and anger.

[0565] "Risk scoring" is the process of quantifying the likelihood of fraudulent activity occurring and evaluating the degree of risk.

[0566] A "warning" is a notification or message intended to alert a user when a high risk is identified.

[0567] This section describes the embodiments for carrying out the invention. This invention is a system that evaluates the risk of fraud and illegal activities in real time through voice calls and issues warnings to users. The following shows the specific configuration and operating procedures for realizing this system.

[0568] The device acquires audio information during a voice call and captures it as a digital signal using the microphone. This audio information is immediately sent to a cloud server. The server converts the audio into document data using a speech recognition API (e.g., Google Speech-to-Text API). This process results in the conversation being obtained in text format.

[0569] Next, the server uses an emotion analysis engine (e.g., IBM Watson Tone Analyzer) to identify the user's emotional state from the acquired audio data. This makes it possible to detect changes in the speaker's emotions in real time.

[0570] The server uses a generative AI model (e.g., OpenAI GPT model) to score the risk of fraudulent activity based on document data and sentiment analysis results. This model analyzes keywords and context within the text and combines them with emotional states to assess the risk.

[0571] As a concrete example, consider a scenario where a user receives a suspicious phone call. The server, through sentiment analysis, recognizes that the user is feeling uneasy and detects that the conversation contains dangerous keywords such as "banking information." As a result, the server generates a high risk score and sends a warning message to the user, such as "Urgent Warning: This call may be a scam. Take immediate action."

[0572] An example of a prompt to input into the generating AI model is: "Calculate a fraud score based on the text extracted from the audio: 'Please verify your bank information now.' Sentiment analysis indicates anxiety."

[0573] This system can detect fraudulent activity during calls with high accuracy and speed, and provide effective warnings to users.

[0574] The flow of a specific process in Application Example 2 will be explained using Figure 14.

[0575] Step 1:

[0576] The device acquires the user's voice information through the microphone during a voice call. This voice information is captured in real time as a digital signal. The input is voice data, and the output is a digitized voice signal. These signals are immediately transmitted to the server via the communication line.

[0577] Step 2:

[0578] The server uses a speech recognition API on the cloud to convert digital audio signals into document data. Specifically, it uses the Google Speech-to-Text API to convert audio into text format. The input is a digital audio signal, and the output is the corresponding document data.

[0579] Step 3:

[0580] The server provides the converted document data to the sentiment analysis engine to identify the user's emotional state. This process uses IBM Watson Tone Analyzer to analyze factors such as speech tone and rhythm. The input is document data, and the output is data indicating the user's emotional state.

[0581] Step 4:

[0582] The server uses a generative AI model to score the risk of fraudulent activity based on document data and sentiment data. Specifically, it uses the OpenAI GPT model to perform a risk assessment by combining the analysis of fraudulent keywords contained in the text with anomalies in sentiment. The input is document data and sentiment data, and the output is a fraud risk score.

[0583] Step 5:

[0584] The server generates a warning message based on the fraud risk score and notifies the user in real time. If the risk score is high, an enhanced warning message is sent to the device as a push notification. The input is the fraud risk score, and the output is the warning message.

[0585] Step 6:

[0586] The user receives a warning message from the server and decides whether to continue or end the call depending on the situation. Depending on the warning, the user is provided with a means to manually end the call. The input is the warning message, and the output is the user's response action.

[0587] The specific processing unit 290 transmits the result of the specific processing to the headset terminal 314. In the headset terminal 314, the control unit 46A causes the speaker 240 and display 343 to output the result of the specific processing. The microphone 238 acquires audio indicating user input for the result of the specific processing. The control unit 46A transmits the audio data indicating user input acquired by the microphone 238 to the data processing unit 12. In the data processing unit 12, the specific processing unit 290 acquires the audio data.

[0588] Data generation model 58 is a type of so-called generative AI (Artificial Intelligence). One example of data generation model 58 is ChatGPT (Internet search<URL: https: / / openai.com / blog / chatgpt> ), Gemini (Internet search) <url: https: gemini.google.com ?hl="ja">Examples of generative AI include the following. The data generation model 58 is obtained by performing deep learning on a neural network. The data generation model 58 is input with prompts containing instructions, and with inference data such as audio data representing speech, text data representing text, and image data representing images. The data generation model 58 infers from the input inference data according to the instructions indicated by the prompts, and outputs the inference results in data formats such as audio data and text data. Here, inference refers to, for example, analysis, classification, prediction, and / or summarization.

[0589] In the above embodiment, an example was given in which specific processing is performed by the data processing device 12, but the technology of this disclosure is not limited thereto, and specific processing may also be performed by the headset terminal 314.

[0590] [Fourth Embodiment]

[0591] Figure 7 shows an example of the configuration of the data processing system 410 according to the fourth embodiment.

[0592] As shown in Figure 7, the data processing system 410 includes a data processing device 12 and a robot 414. An example of the data processing device 12 is a server.

[0593] The data processing device 12 comprises a computer 22, a database 24, and a communication interface 26. The computer 22 is an example of a "computer" related to the technology of this disclosure. The computer 22 comprises a processor 28, RAM 30, and storage 32. The processor 28, RAM 30, and storage 32 are connected to a bus 34. The database 24 and the communication interface 26 are also connected to the bus 34. The communication interface 26 is connected to a network 54. An example of the network 54 is a WAN (Wide Area Network) and / or a LAN (Local Area Network).

[0594] The robot 414 includes a computer 36, a microphone 238, a speaker 240, a camera 42, a communication interface 44, and a controlled object 443. The computer 36 includes a processor 46, RAM 48, and storage 50. The processor 46, RAM 48, and storage 50 are connected to a bus 52. The microphone 238, speaker 240, camera 42, and controlled object 443 are also connected to the bus 52.

[0595] The microphone 238 receives voice signals from the user 20 and receives instructions from the user 20. The microphone 238 captures the voice signals from the user 20, converts the captured voice into audio data, and outputs it to the processor 46. The speaker 240 outputs audio according to the instructions from the processor 46.

[0596] Camera 42 is a small digital camera equipped with an optical system including a lens, aperture, and shutter, and an image sensor such as a CMOS (Complementary Metal-Oxide-Semiconductor) image sensor or a CCD (Charge Coupled Device) image sensor, and captures images of the area around the user 20 (for example, an imaging range defined by a field of view equivalent to the width of a typical healthy person's field of vision).

[0597] Communication interface 44 is connected to network 54. Communication interfaces 44 and 26 are responsible for the exchange of various information between processor 46 and processor 28 via network 54. The exchange of various information between processor 46 and processor 28 using communication interfaces 44 and 26 is performed in a secure manner.

[0598] The controlled object 443 includes a display device, LEDs in the eyes, and motors that drive the arms, hands, and feet. The posture and gestures of the robot 414 are controlled by controlling the motors of the arms, hands, and feet. Some of the robot 414's emotions can be expressed by controlling these motors. Furthermore, the robot 414's facial expressions can also be expressed by controlling the illumination state of the LEDs in its eyes.

[0599] Figure 8 shows an example of the main functions of the data processing device 12 and the robot 414. As shown in Figure 8, the data processing device 12 performs specific processing using the processor 28. The storage 32 stores the specific processing program 56.

[0600] The specific processing program 56 is an example of a "program" relating to the technology of this disclosure. The processor 28 reads the specific processing program 56 from the storage 32 and executes the read specific processing program 56 on the RAM 30. The specific processing is realized by the processor 28 operating as a specific processing unit 290 in accordance with the specific processing program 56 executed on the RAM 30.

[0601] The storage 32 stores the data generation model 58 and the emotion identification model 59. The data generation model 58 and the emotion identification model 59 are used by the identification processing unit 290.

[0602] In robot 414, the processor 46 performs the reception output processing. The storage 50 stores the reception output program 60. The processor 46 reads the reception output program 60 from the storage 50 and executes the read reception output program 60 on the RAM 48. The reception output processing is realized by the processor 46 operating as a control unit 46A according to the reception output program 60 executed on the RAM 48.

[0603] Next, the specific processing performed by the specific processing unit 290 of the data processing device 12 will be described. In the following description, the data processing device 12 will be referred to as the "server" and the robot 414 as the "terminal".

[0604] This invention provides a system for processing voice data in real time during communication in order to protect individuals from fraud. This system consists of a terminal, a server, and a user.

[0605] The device has a function that records audio data in real time when a call is received, and then divides that data and sends it to a server. This provides a foundation for timely analysis of the audio from the phone calls that users receive on a daily basis.

[0606] The server performs a process of converting received audio data into text data using speech recognition technology. Subsequently, it analyzes the text data using a generative AI model and scores the likelihood of fraudulent activity. Based on this score, if the risk of fraudulent activity is high, a warning message is generated and sent to the user or registered third party via electronic communication. For example, if a suspicious request regarding a bank account is detected, a warning such as "This may be a scam, so please be careful" is immediately sent to the user's smartphone.

[0607] If a user receives a warning message, they will be given the option to manually end the call, or the call will be automatically terminated by instructions from the server. In this way, the system prevents users from becoming victims of fraud. In addition, after the call, users can provide feedback on whether or not fraud actually occurred, and this feedback information is used by the server to improve the accuracy of the generated AI model.

[0608] For example, if a user receives a phone call and the caller says, "Please tell me your credit card information immediately," the system transcribes this audio into text, and an AI model determines that it is highly likely to be a scam. This allows the user to receive a real-time warning, preventing them from becoming a victim. This entire process is automated, making the system usable by users without requiring advanced technical knowledge.

[0609] The following describes the processing flow.

[0610] Step 1:

[0611] The device records the audio of calls received by the user in real time. The recording is divided into short time units and immediately prepared as data packets.

[0612] Step 2:

[0613] The device encrypts the recorded audio data packets before sending them to the server. This process is designed to minimize latency.

[0614] Step 3:

[0615] The server converts the received audio data into text data using a speech recognition API. This conversion must be fast and highly accurate.

[0616] Step 4:

[0617] The server analyzes text data using a generative AI model and scores the likelihood of fraudulent activity. This model is designed to identify words and contexts associated with fraud.

[0618] Step 5:

[0619] If the server determines, based on the scoring results, that there is a high probability of fraudulent activity, it will generate a warning message and send it to the user and registered third parties via electronic means. This message will include a warning about fraudulent activity and information on specific countermeasures.

[0620] Step 6:

[0621] If the server determines that fraudulent activity has occurred, it will automatically disconnect the call by sending a call termination command to the terminal.

[0622] Step 7:

[0623] Users send feedback to the server via a simple method, such as the LINE app, to determine whether the call content was truly fraudulent. Based on this feedback, the server improves the detection accuracy of the AI ​​model.

[0624] (Example 1)

[0625] Next, we will describe Example 1. In the following description, the data processing device 12 will be referred to as the "server" and the robot 414 as the "terminal".

[0626] In the communications that individuals receive on a daily basis, there is a need to quickly and efficiently detect the risks of fraudulent activity and prevent its impact. Conventional systems lack sufficient means to identify fraud, and users may easily become involved in fraudulent activities. This invention aims to solve these problems and provide users with a secure communication environment.

[0627] The identification process performed by the identification processing unit 290 of the data processing device 12 in Example 1 is realized by the following means.

[0628] In this invention, the server includes means for acquiring voice information in real time, means for converting it into text information, and means for evaluating the risk of fraudulent activity. This makes it possible to quickly detect the possibility of fraudulent activity during communication and provide appropriate warnings to the user.

[0629] "Voice information" refers to voice data generated by the user during communication, and the purpose is to acquire this data in real time.

[0630] "Real-time" refers to processing information immediately at the moment the communication is taking place, meaning that information is acquired and processed without delay.

[0631] "Textual information" refers to text data generated from audio information through speech recognition technology.

[0632] "Assessing the risk of fraudulent activity" refers to the process of analyzing acquired textual information to identify and quantify the possibility of fraudulent activities such as scams.

[0633] "Generating a warning" means creating a message to alert a user when a high risk of fraudulent activity is detected.

[0634] "Terminating communication" refers to stopping the call or communication session in response to the risk of fraudulent activity detected during the communication.

[0635] "Electronic communication technology" refers to technologies that send and receive information via digital means such as the internet, SMS, and email, and is used for transmitting warning messages in this system.

[0636] This system is built on communication protection technologies that include terminals, servers, and users. It primarily handles the acquisition, conversion, analysis, evaluation, and generation of warnings for voice information.

[0637] When the device receives a phone call or other communication, it uses its built-in voice recording device to capture voice information in real time. This data is then divided into manageable sizes, encrypted, and securely transmitted to the server.

[0638] The server converts the received audio information into text using speech recognition technology. This process utilizes speech recognition software such as the Google Speech-to-Text API. The text information is then analyzed by a generative AI model. This model receives the converted text information as prompts and performs a scoring system to assess the risk of fraudulent activity. If the scoring result exceeds a certain threshold, the server automatically generates a warning message and sends it to the user's device. The warning message informs the user of the potential for fraudulent activity and prompts them to take appropriate action.

[0639] The user reviews the received warning message and ends the call if necessary. Furthermore, the system improves the accuracy of the generated AI model through user feedback.

[0640] For example, if a user receives a request such as "Please tell me your credit card information" during an incoming call, the device immediately records the audio and sends it to the server. The server transcribes the audio into text and scores it as highly suspicious. As a result, a warning message such as "This call may be a scam" is immediately sent to the device, allowing the user to safely end the call.

[0641] An example of a prompt for a generative AI model might be: "Analyze the following text data and determine the risk of fraudulent activity. Rate the risk on a scale of 100." This prompt provides the generative AI model with criteria for effectively identifying fraud.

[0642] The flow of the specific processing in Example 1 will be explained using Figure 11.

[0643] Step 1:

[0644] The device automatically acquires audio information when the user receives a call. The call audio is used as input and converted into digital audio data by the device's audio recording device. This audio data is then divided into manageable sizes. The resulting divided audio data is encrypted and ready to be sent to the server. This enables secure, real-time data transfer.

[0645] Step 2:

[0646] The server receives audio data transmitted from the terminal. Encrypted audio data is provided to the server as input and converted into text using speech recognition technology (e.g., Google Speech-to-Text API). This process involves sophisticated calculations to convert speech into text. The resulting text is then sent to the next stage for analysis.

[0647] Step 3:

[0648] The server analyzes the character information converted using a generative AI model. The character information generated in the previous step is provided as input, and a prompt sentence is fed to the AI ​​model. According to this prompt sentence, the AI ​​model scores the likelihood of fraudulent activity. The output is the fraudulent activity risk score corresponding to each text fragment. This score serves as a basis for generating warning messages through several algorithms.

[0649] Step 4:

[0650] The server generates a warning message based on a fraud risk score. If a high risk score is detected, a warning message is generated, and the scoring result is used as input. The output is the generated warning message. This message is sent to the user's terminal or a related third party via electronic means.

[0651] Step 5:

[0652] The user reviews the received warning message. The warning message sent from the server is used as input. The user can manually end the call as needed, or it may be automatically terminated depending on the settings. The output provides a way to securely manage the call status and prevent further harm to the user. Furthermore, the user can contribute to improving the system's accuracy by providing feedback.

[0653] (Application Example 1)

[0654] Next, we will explain Application Example 1. In the following explanation, the data processing device 12 will be referred to as the "server" and the robot 414 as the "terminal".

[0655] It can quickly detect the risk of fraudulent activity during communications and provide a reliable way to protect individuals from fraud. Many people are exposed to the risk of fraud on a daily basis, and there is a need for immediate responses to suspicious requests during calls.

[0656] The specific processing performed by the specific processing unit 290 of the data processing device 12 in Application Example 1 is realized by the following means.

[0657] In this invention, the server includes means for instantly acquiring voice information, means for converting the acquired voice information into text information, and means for analyzing the text information to evaluate the risk of fraudulent activity. This makes it possible to evaluate the risk of fraudulent activity during communication in real time and provide individuals with prompt warnings and protection.

[0658] "Voice information" refers to voice data acquired during a call, which is analyzed in real time.

[0659] "Textual information" refers to data that is generated by analyzing audio information and expressing it as text.

[0660] "Fraudulent activity risk" is an indicator that shows the likelihood of fraudulent activities such as scams and illegal transactions occurring.

[0661] "Evaluation" refers to the act of judging the risk of fraudulent activity based on acquired information, using numerical or other formats.

[0662] A "warning" is information intended to inform users and stakeholders about the risk of fraudulent activity.

[0663] "Feedback" refers to information provided by users after a call, which is used to improve the model.

[0664] A "model" refers to a generative AI model used to assess the risk of fraudulent activity.

[0665] In order to implement this invention, it is necessary to construct a system that analyzes voice information and issues warnings through the cooperation of a terminal, a server, and a user.

[0666] First, when a user initiates a call, the device immediately captures audio information and sends it to the server. The device processes audio data in the background even during a call, so it operates without interfering with the user's actions.

[0667] The server immediately converts the received audio information into text using a speech recognition API (e.g., Google Cloud Speech-to-Text). This text information is then input into a generative AI model (e.g., OpenAI's GPT-4) to assess the risk of fraud. For example, if the message contains phrases like "Please tell me your credit card information quickly," the generative AI model analyzes it and determines that it is highly likely to be fraud.

[0668] If the evaluation determines that there is a high risk of fraudulent activity, the server will generate a warning and send it to the user's terminal via electronic communication. Upon receiving the warning, the user can either manually end the call, or in some cases, the server will automatically terminate the call.

[0669] Users can provide feedback after a call ends, and this feedback is collected on the server to help improve the accuracy of the generated AI model.

[0670] As a concrete example, a prompt such as, "Analyze the text of this conversation and determine the risk of fraud. It contains the statement, 'I urgently need to send money to save my grandmother.'" can be used. This helps improve the accuracy of the model for similar fraud scenarios.

[0671] The flow of a specific process in Application Example 1 will be explained using Figure 12.

[0672] Step 1:

[0673] The terminal captures audio information in real time at the start of a call. It takes audio data from the call as input and prepares it to be sent to the server at regular intervals. The output is a fragment of audio data ready to be sent to the server.

[0674] Step 2:

[0675] The server receives audio data fragments sent from the terminal. It takes these audio data fragments as input and converts them into text information using a speech recognition API. The output is text information derived from the audio data. This process utilizes speech recognition software and leverages natural language processing techniques.

[0676] Step 3:

[0677] The server uses a generative AI model to analyze the converted text information and assess the risk of fraudulent activity. Text information is passed to the generative AI model as input, and analysis is performed based on the prompt text. The output is evaluated in the form of a fraud risk score. This evaluation determines whether the risk exceeds a threshold.

[0678] Step 4:

[0679] The server sends a warning to the user's device if the fraud risk score is high. The server takes a risk score as input and creates a warning message via electronic communication. The output is the warning notification received by the user. The notification contains information indicating a high probability of fraud.

[0680] Step 5:

[0681] The user receives a warning and terminates the call manually or at the server's instruction. The input is a warning notification, and the user decides whether to continue the call. The output is the execution of the call termination. At this point, the call is disconnected either by pressing the call termination button or by the server's automatic operation.

[0682] Step 6:

[0683] After the call ends, the user provides feedback to the server. The server collects user experience information and opinions as input and sends them to the server. The output is feedback data used to improve the accuracy of the generated AI model. The feedback content is used as training data for the model.

[0684] Furthermore, an emotion engine that estimates the user's emotions may be incorporated. That is, the identification processing unit 290 may use the emotion identification model 59 to estimate the user's emotions and perform identification processing using the user's emotions.

[0685] This invention provides a system that prevents fraudulent activity by acquiring voice data during a call in real time and converting it into text data and sentiment data. The system simultaneously monitors the content of the call and the user's emotional state, and takes warnings and preventative measures based on the results.

[0686] The device records voice data in real time during a call and sends the audio to the server. The server converts this voice data into text data using a speech recognition API, and at the same time uses an emotion engine to identify the user's emotional state from the voice. At this time, it analyzes emotions from the tone and tempo of the user's voice, paying particular attention to cases where emotions such as anxiety or doubt are prominent.

[0687] The server's generated AI model analyzes keywords and phrases related to fraud using text data and scores the risk of fraud. Emotional information obtained by the emotion engine is also incorporated into the scoring; if the emotional state is unusual, the likelihood of fraudulent activity is highly rated. This dual analytical approach allows for more accurate risk assessment.

[0688] Based on the evaluation results, the server will send warning messages to the user and registered third parties as needed. If the sentiment engine determines that the situation is particularly dangerous, it can increase the intensity of the message and prompt action. Upon receiving this warning, the user can manually end the call, or the call may be automatically disconnected at the server's direction.

[0689] For example, if a user receives a suspicious phone call and is strongly urged to "check information immediately," the emotion engine will detect that the voice indicates anxiety, and text analysis will determine that it is highly likely to be a scam. As a result, the user will receive an enhanced warning such as "Urgent Warning: This call may be a scam. Take action immediately," prompting the user to act before they become a victim.

[0690] In this way, the present invention provides a system that enables more reliable detection and rapid response to fraudulent activity through a combination of text analysis and sentiment recognition.

[0691] The following describes the processing flow.

[0692] Step 1:

[0693] The device starts recording in real time when the user begins a call. This audio data is divided into small packets and adjusted for processing without delay.

[0694] Step 2:

[0695] The device encrypts the recorded audio data and sends it to the server. This transmission takes place over the network, ensuring security.

[0696] Step 3:

[0697] The server converts the received audio data into text data using a speech recognition API. This process requires that the audio information be converted into text information quickly and accurately.

[0698] Step 4:

[0699] The server simultaneously uses an emotion engine to identify the user's emotional state from the voice data. It analyzes voice tone, tempo, tension, etc., to identify emotions such as joy, anxiety, and anger.

[0700] Step 5:

[0701] The server's generated AI model analyzes text data and scores the likelihood of fraud. It performs risk assessment by detecting specific keywords and contexts.

[0702] Step 6:

[0703] Emotional information from the emotion engine is fed back into the scoring results. If the user's emotional state is different from normal, for example, if they are highly anxious or suspicious, it is determined that there is a high possibility of fraud.

[0704] Step 7:

[0705] Based on the evaluation results, the server generates a warning message and sends it to the user or registered third parties via electronic communication. In some cases, a stronger warning than usual may be issued.

[0706] Step 8:

[0707] The user reviews the received warning message and either manually ends the call or the server automatically disconnects it. This process prompts the user to take swift action.

[0708] Step 9:

[0709] After a call ends, the user sends feedback to the server indicating whether or not the call was a scam. This information is used to further improve the accuracy of the AI ​​model.

[0710] (Example 2)

[0711] Next, we will describe Example 2. In the following description, the data processing device 12 will be referred to as the "server" and the robot 414 as the "terminal".

[0712] With fraudulent activities using communication methods on the rise, there is a need to monitor conversation content in real time and detect signs of fraud with high accuracy. Furthermore, by considering the emotional state of users, it is necessary to conduct more accurate risk assessments and to have means to quickly protect users from fraudulent activities.

[0713] The identification process performed by the identification processing unit 290 of the data processing device 12 in Example 2 is realized by the following means.

[0714] In this invention, the server includes means for acquiring voice information in real time, means for converting the acquired voice information into text information, means for identifying emotional states from the voice information, and means for analyzing the text information and emotional states to assess the risk of fraudulent activity. This makes it possible to accurately assess the risk of fraudulent activity from a user's conversation and quickly protect the user from potential fraudulent activity.

[0715] "Voice information" refers to digital data related to human voices obtained from phone calls, recordings, etc.

[0716] "Real-time" refers to a state where processing or actions are performed instantly without delay.

[0717] "Textual information" refers to data in text format obtained as a result of converting audio information.

[0718] "Emotional state" refers to the speaker's emotional state and psychological changes, as analyzed from audio information.

[0719] "Fraudulent activity" refers to actions and methods related to fraud or attempted fraud, including acts for fraudulent purposes in communications.

[0720] "Risk assessment" refers to the process of calculating the likelihood of fraudulent activity as a numerical value or evaluation score based on information analysis.

[0721] A "warning" refers to a message or notification intended to inform a user or a third party in advance that a risk exists.

[0722] "Communication" refers to the act or process of sending and receiving voice or digital information.

[0723] This invention is a system that processes voice information during a call in real time, identifies fraudulent activity, and protects the user. This system is realized through cooperation between a terminal and a server.

[0724] The device acquires audio information in real time when a user initiates a call. This function is performed by communication devices such as smartphones and internet phones. The acquired audio information is transmitted to the server via a secure communication protocol.

[0725] The server converts the received audio information into text using a speech recognition API (for example, a common speech recognition technology for converting speech to text). This text conversion process employs an efficient algorithm to minimize latency. Simultaneously, the server analyzes the audio information using an emotion analysis engine to identify the user's emotional state. This emotion analysis incorporates techniques to evaluate the tone, speed, and volume of the voice.

[0726] Next, the server uses a generative AI model to detect terms and context related to fraudulent activity from the textual information. This model is trained to be particularly sensitive to keywords and phrases related to fraud. At the same time, the user's emotional state is also taken into consideration and integrated into the risk assessment.

[0727] Based on the risk assessment performed by the server, a warning will be sent to the user if necessary. The warning will be electronically distributed to the terminal and any relevant third parties to prompt the user to take prompt action. This warning will allow the user to manually end the call, or in some cases, the server may automatically disconnect the call.

[0728] As a concrete example, consider a case where a fictitious scam call is received. If a phrase such as "You need to provide your personal information immediately" is detected, and the user's voice indicates anxiety, the system will send a warning to the user saying, "Caution: This call may be a scam. Please be very careful when answering."

[0729] An example of a prompt to be input to the generating AI model is, "Analyze the content and tone of voice of the call, identify potential fraudulent activity, and assess the risk." In this way, the present invention provides a comprehensive call monitoring function to support user safety.

[0730] The flow of the specific processing in Example 2 will be explained using Figure 13.

[0731] Step 1:

[0732] The device detects the start of a call and acquires audio information in real time. This audio information is collected through the microphone and stored as digital data within the device. The input is the analog audio signal during the call, and the output is the digitized audio data. The device transmits this data to a server via the internet.

[0733] Step 2:

[0734] The server receives audio data from the terminal and inputs it into a speech recognition API, where it converts it into text. The speech recognition software uses an acoustic model and a language model to perform the process of converting the audio data into corresponding text. The input is digital audio data, and the output is text data.

[0735] Step 3:

[0736] The server inputs the received audio data into an emotion analysis engine to identify the user's emotional state. This process analyzes characteristics such as tone, speed, and volume of the voice and classifies them into emotion categories (e.g., anger, anxiety, joy). The input is audio data, and the output is data indicating the user's emotional state.

[0737] Step 4:

[0738] The server uses a generation AI model to analyze the converted text information. This model detects specific keywords and phrases and scores the risk of fraudulent activity based on them. Furthermore, emotional state data is also reflected in this risk assessment, and the score is adjusted if unstable emotions are detected. The input is text data and emotional state data, and the output is a fraudulent activity risk score.

[0739] Step 5:

[0740] The server generates a warning message based on the risk score. If a high risk is detected, the warning is set as urgent and sent to the user's and registered third-party devices. Specifically, a message such as "This call shows signs of fraud. Please be careful" is generated. The input is the risk score, and the output is the warning message.

[0741] Step 6:

[0742] Depending on the warning message received by the user, the call can be terminated either manually or automatically by a disconnection command from the server. This provides swift protection from fraudulent activity. The input is the warning message, and the output is the call termination action.

[0743] (Application Example 2)

[0744] Next, we will explain application example 2. In the following explanation, the data processing device 12 will be referred to as the "server" and the robot 414 as the "terminal".

[0745] In recent years, fraud and illegal activities conducted via voice calls have been increasing, causing harm to many individuals and organizations. There is a need for a system that can detect such fraudulent activities in real time and issue rapid warnings. Furthermore, to prevent harm, highly accurate risk assessments that take into account changes in user emotions are necessary. However, existing systems struggle to meet these requirements with sufficient accuracy.

[0746] The specific processing performed by the specific processing unit 290 of the data processing device 12 in Application Example 2 is realized by the following means.

[0747] In this invention, the server includes means for acquiring voice information in real time, means for converting the acquired voice information into document data, means for analyzing the document data to evaluate the possibility of fraudulent activity, means for performing sentiment analysis and identifying the user's emotional state, and means for scoring the risk of fraudulent activity based on the emotional state. This makes it possible to evaluate the risk of fraud and other fraudulent activities with high accuracy and to quickly warn the user.

[0748] "Voice information" refers to sound signals acquired through phone calls or recordings, and is digital data that includes user speech and background sounds.

[0749] "Real-time" refers to a situation where processing and responses occur instantaneously, with no time lag between voice acquisition and analysis.

[0750] "Document data" refers to information in string format converted from audio information, and is content that has been digitized as text.

[0751] "Analysis" is the process of examining information in detail to understand its structure and meaning, and making decisions according to a specific purpose.

[0752] "Assessing the possibility of fraudulent activity" means using audio and documentary data to determine whether there is a risk of fraud or misconduct.

[0753] "Emotional analysis" is a technique that analyzes characteristics such as tone and rhythm of speech information to infer the speaker's emotional state.

[0754] "Emotional state" refers to information that indicates the psychological emotions expressed during a speaker's utterance, and includes feelings such as joy, anxiety, and anger.

[0755] "Risk scoring" is the process of quantifying the likelihood of fraudulent activity occurring and evaluating the degree of risk.

[0756] A "warning" is a notification or message intended to alert a user when a high risk is identified.

[0757] This section describes the embodiments for carrying out the invention. This invention is a system that evaluates the risk of fraud and illegal activities in real time through voice calls and issues warnings to users. The following shows the specific configuration and operating procedures for realizing this system.

[0758] The device acquires audio information during a voice call and captures it as a digital signal using the microphone. This audio information is immediately sent to a cloud server. The server converts the audio into document data using a speech recognition API (e.g., Google Speech-to-Text API). This process results in the conversation being obtained in text format.

[0759] Next, the server uses an emotion analysis engine (e.g., IBM Watson Tone Analyzer) to identify the user's emotional state from the acquired audio data. This makes it possible to detect changes in the speaker's emotions in real time.

[0760] The server uses a generative AI model (e.g., OpenAI GPT model) to score the risk of fraudulent activity based on document data and sentiment analysis results. This model analyzes keywords and context within the text and combines them with emotional states to assess the risk.

[0761] As a concrete example, consider a scenario where a user receives a suspicious phone call. The server, through sentiment analysis, recognizes that the user is feeling uneasy and detects that the conversation contains dangerous keywords such as "banking information." As a result, the server generates a high risk score and sends a warning message to the user, such as "Urgent Warning: This call may be a scam. Take immediate action."

[0762] An example of a prompt to input into the generating AI model is: "Calculate a fraud score based on the text extracted from the audio: 'Please verify your bank information now.' Sentiment analysis indicates anxiety."

[0763] This system can detect fraudulent activity during calls with high accuracy and speed, and provide effective warnings to users.

[0764] The flow of a specific process in Application Example 2 will be explained using Figure 14.

[0765] Step 1:

[0766] The device acquires the user's voice information through the microphone during a voice call. This voice information is captured in real time as a digital signal. The input is voice data, and the output is a digitized voice signal. These signals are immediately transmitted to the server via the communication line.

[0767] Step 2:

[0768] The server uses a speech recognition API on the cloud to convert digital audio signals into document data. Specifically, it uses the Google Speech-to-Text API to convert audio into text format. The input is a digital audio signal, and the output is the corresponding document data.

[0769] Step 3:

[0770] The server provides the converted document data to the sentiment analysis engine to identify the user's emotional state. This process uses IBM Watson Tone Analyzer to analyze factors such as speech tone and rhythm. The input is document data, and the output is data indicating the user's emotional state.

[0771] Step 4:

[0772] The server uses a generative AI model to score the risk of fraudulent activity based on document data and sentiment data. Specifically, it uses the OpenAI GPT model to perform a risk assessment by combining the analysis of fraudulent keywords contained in the text with anomalies in sentiment. The input is document data and sentiment data, and the output is a fraud risk score.

[0773] Step 5:

[0774] The server generates a warning message based on the fraud risk score and notifies the user in real time. If the risk score is high, an enhanced warning message is sent to the device as a push notification. The input is the fraud risk score, and the output is the warning message.

[0775] Step 6:

[0776] The user receives a warning message from the server and decides whether to continue or end the call depending on the situation. Depending on the warning, the user is provided with a means to manually end the call. The input is the warning message, and the output is the user's response action.

[0777] The specific processing unit 290 transmits the result of the specific processing to the robot 414. In the robot 414, the control unit 46A causes the speaker 240 and the controlled object 443 to output the result of the specific processing. The microphone 238 acquires audio indicating user input for the result of the specific processing. The control unit 46A transmits the audio data indicating user input acquired by the microphone 238 to the data processing unit 12. In the data processing unit 12, the specific processing unit 290 acquires the audio data.

[0778] Data generation model 58 is a type of so-called generative AI (Artificial Intelligence). One example of data generation model 58 is ChatGPT (Internet search<URL: https: / / openai.com / blog / chatgpt> ), Gemini (Internet search) <url: https: gemini.google.com ?hl="ja">Examples of generative AI include the following. The data generation model 58 is obtained by performing deep learning on a neural network. The data generation model 58 is input with prompts containing instructions, and with inference data such as audio data representing speech, text data representing text, and image data representing images. The data generation model 58 infers from the input inference data according to the instructions indicated by the prompts, and outputs the inference results in data formats such as audio data and text data. Here, inference refers to, for example, analysis, classification, prediction, and / or summarization.

[0779] In the above embodiment, an example was given in which specific processing is performed by the data processing device 12, but the technology of this disclosure is not limited thereto, and the specific processing may also be performed by the robot 414.

[0780] Furthermore, the emotion identification model 59, acting as an emotion engine, may determine the user's emotion according to a specific mapping. Specifically, the emotion identification model 59 may determine the user's emotion according to a specific mapping, which is an emotion map (see Figure 9). Similarly, the emotion identification model 59 may also determine the robot's emotion, and the identification processing unit 290 may perform identification processing using the robot's emotion.

[0781] Figure 9 shows an emotion map 400 in which multiple emotions are mapped. In the emotion map 400, emotions are arranged in concentric circles radiating from the center. The closer to the center of the concentric circles, the more primitive the emotions are located. Further out of the concentric circles, emotions representing states and actions arising from mental states are located. Emotion is a concept that includes feelings and mental states. On the left side of the concentric circles, emotions that are generally generated from reactions occurring in the brain are located. On the right side of the concentric circles, emotions that are generally induced by situational judgment are located. Above and below the concentric circles, emotions that are generally generated from reactions occurring in the brain and induced by situational judgment are located. In addition, the emotion of "pleasure" is located on the upper side of the concentric circles, and the emotion of "displeasure" is located on the lower side. Thus, in the emotion map 400, multiple emotions are mapped based on the structure in which emotions arise, and emotions that are likely to occur simultaneously are mapped close together.

[0782] These emotions are distributed at the 3 o'clock position on the Emotion Map 400, and usually fluctuate between feelings of security and anxiety. In the right half of the Emotion Map 400, situational awareness takes precedence over internal feelings, resulting in a calm impression.

[0783] The inside of the Emotion Map 400 represents inner thoughts, while the outside represents actions. Therefore, the further you go from the outside of the Emotion Map 400, the more visible (expressed in actions) your emotions become.

[0784] Here, human emotions are based on various balances, such as posture and blood sugar levels. When these balances deviate from the ideal, it results in discomfort, and when they approach the ideal, it results in pleasure. Similarly, in robots, cars, motorcycles, etc., emotions can be created based on various balances, such as posture and battery level. When these balances deviate from the ideal, it results in discomfort, and when they approach the ideal, it results in pleasure. The emotion map can be generated, for example, based on Dr. Mitsuyoshi's emotion map (Research on a system for analyzing brain physiological signals of speech emotion recognition and emotion, Tokushima University, doctoral dissertation: https: / / ci.nii.ac.jp / naid / 500000375379). The left half of the emotion map contains emotions belonging to a region called "response," where sensation is dominant. The right half of the emotion map contains emotions belonging to a region called "situation," where situational awareness is dominant.

[0785] The emotion map defines two emotions that promote learning. One is the emotion around the middle of the negative "repentance" and "reflection" on the situation side. In other words, it is when the robot experiences negative emotions such as "I never want to feel this way again" or "I don't want to be scolded again." The other is the emotion around the positive "desire" on the reaction side. In other words, it is when the robot has positive feelings such as "I want more" or "I want to know more."

[0786] The emotion identification model 59 inputs user input into a pre-trained neural network, obtains emotion values ​​representing each emotion shown in the emotion map 400, and determines the user's emotion. This neural network is pre-trained based on multiple training data sets, which are combinations of user input and emotion values ​​representing each emotion shown in the emotion map 400. Furthermore, this neural network is trained so that emotions located close together have similar values, as shown in the emotion map 900 in Figure 10. Figure 10 shows an example where multiple emotions such as "reassured," "calm," and "confident" have similar emotion values.

[0787] The above description primarily focuses on the functions of the data processing device 12 in relation to this disclosure. However, the system related to this disclosure is not necessarily implemented on a server. The system related to this disclosure may be implemented as a general information processing system. This disclosure may be implemented, for example, as a software program that runs on a personal computer or as an application that runs on a smartphone. The method related to this disclosure may be provided to users in SaaS (Software as a Service) format.

[0788] In the above embodiment, an example was given in which a specific process is performed by a single computer 22. However, the technology of this disclosure is not limited thereto, and a distributed processing of the specific process may be performed by multiple computers, including computer 22. For example, a data generation model 58 may be provided in an external device of the data processing device 12, and the external device may generate data according to the input data.

[0789] In the above embodiment, an example was given in which the specific processing program 56 is stored in the storage 32, but the technology of this disclosure is not limited thereto. For example, the specific processing program 56 may be stored in a portable, computer-readable, non-temporary storage medium such as a USB (Universal Serial Bus) memory. The specific processing program 56 stored in the non-temporary storage medium is installed in the computer 22 of the data processing device 12. The processor 28 executes specific processing according to the specific processing program 56.

[0790] Alternatively, the specific processing program 56 may be stored in a storage device such as a server connected to the data processing device 12 via the network 54, and the specific processing program 56 may be downloaded and installed on the computer 22 in response to a request from the data processing device 12.

[0791] Furthermore, it is not necessary to store the entirety of the specific processing program 56 in a storage device such as a server connected to the data processing device 12 via the network 54, or to store the entirety of the specific processing program 56 in the storage 32; it is acceptable to store only a portion of the specific processing program 56.

[0792] The following types of processors can be used as hardware resources to perform specific processing. Examples of processors include a CPU, a general-purpose processor that functions as a hardware resource to perform specific processing by executing software, i.e., a program. Other examples of processors include dedicated electrical circuits, such as FPGAs (Field-Programmable Gate Arrays), PLDs (Programmable Logic Devices), or ASICs (Application Specific Integrated Circuits), which have circuit configurations specifically designed to perform specific processing. All of these processors have built-in or connected memory, and all of them perform specific processing by using memory.

[0793] The hardware resource that performs a specific process may consist of one of these various processors, or it may consist of a combination of two or more processors of the same or different types (for example, a combination of multiple FPGAs, or a combination of a CPU and an FPGA). Alternatively, the hardware resource that performs a specific process may consist of a single processor.

[0794] Examples of configurations using a single processor include, firstly, a configuration in which one or more CPUs and software are combined to form a single processor, and this processor functions as a hardware resource that performs a specific process. Secondly, there is a configuration using a processor that realizes the functions of the entire system, including multiple hardware resources that perform a specific process, on a single IC chip, as exemplified by SoCs (System-on-a-chip). In this way, a specific process is realized using one or more of the above types of processors as hardware resources.

[0795] Furthermore, the hardware structure of these various processors can more specifically utilize electrical circuits that combine circuit elements such as semiconductor devices. Also, the specific processing described above is merely an example. Therefore, it goes without saying that unnecessary steps can be deleted, new steps added, or the processing order rearranged, as long as it does not deviate from the main purpose.

[0796] The descriptions and illustrations presented above are detailed explanations of the technical aspects of this disclosure and are merely examples of the technical aspects. For example, the above descriptions of the structure, function, operation, and effect are examples of the structure, function, operation, and effect of the technical aspects of this disclosure. Therefore, it goes without saying that you may delete unnecessary parts, add new elements, or replace elements in the descriptions and illustrations presented above, as long as you do not deviate from the essence of the technical aspects of this disclosure. Furthermore, in order to avoid confusion and facilitate understanding of the technical aspects of this disclosure, explanations of common technical knowledge and the like that do not require special explanation to enable the implementation of the technical aspects of this disclosure have been omitted from the descriptions and illustrations presented above.

[0797] All documents, patent applications, and technical standards described herein are incorporated by reference to the same extent as if each individual document, patent application, and technical standard were specifically and individually noted to be incorporated by reference.

[0798] The following is further disclosed regarding the embodiments described above.

[0799] (Claim 1)

[0800] A means of acquiring audio data in real time,

[0801] Means for converting the acquired audio data into text data,

[0802] A means for analyzing the aforementioned text data to evaluate the possibility of fraudulent activity,

[0803] A means for outputting a warning based on the aforementioned evaluation,

[0804] A means of terminating a call in response to the aforementioned warning,

[0805] A system that includes this.

[0806] (Claim 2)

[0807] The system according to claim 1, wherein the evaluation means includes means for scoring the likelihood of fraudulent activity by detecting specific keywords or contexts related to fraudulent activity.

[0808] (Claim 3)

[0809] The system according to claim 1, wherein the warning means includes means for transmitting a warning message to a user or an associated third party via electronic communication means.

[0810] "Example 1"

[0811] (Claim 1)

[0812] A means of acquiring audio information in real time,

[0813] Means for converting the acquired audio information into text information,

[0814] A means for analyzing the aforementioned textual information to evaluate the risk of fraudulent activity,

[0815] Means for generating a warning based on the aforementioned evaluation,

[0816] Means for terminating communication in response to the aforementioned warning,

[0817] A system that includes this.

[0818] (Claim 2)

[0819] The system according to claim 1, wherein the evaluation means includes means for scoring the risk of fraudulent activity by identifying specific keywords or contexts related to fraudulent activity.

[0820] (Claim 3)

[0821] The system according to claim 1, wherein the warning means includes means for transmitting a warning message to a user or a related third party via electronic communication technology.

[0822] "Application Example 1"

[0823] (Claim 1)

[0824] A means of instantly acquiring audio information,

[0825] Means for converting the acquired audio information into text information,

[0826] A means for analyzing the aforementioned textual information to evaluate the risk of fraudulent activity,

[0827] Means for sending a warning based on the aforementioned evaluation,

[0828] A means for automatically terminating a call based on the aforementioned warning,

[0829] A means for users to provide feedback after a call,

[0830] This feedback provides a means to improve the accuracy of the model,

[0831] Information processing device including

[0832] (Claim 2)

[0833] The information processing apparatus according to claim 1, wherein the evaluation means has means for scoring the likelihood of fraudulent activity by detecting words and contexts related to fraudulent activity using a generative AI model.

[0834] (Claim 3)

[0835] The information processing apparatus according to claim 1, wherein the warning means includes means for transmitting warning information to a user or a related third party via electronic communication means.

[0836] "Example 2 of combining an emotion engine"

[0837] (Claim 1)

[0838] A means of acquiring audio information in real time,

[0839] Means for converting the acquired audio information into text information,

[0840] A means for identifying an emotional state from the aforementioned audio information,

[0841] A means for analyzing the aforementioned textual information and emotional state to assess the risk of fraudulent activity,

[0842] A means for outputting a warning based on the aforementioned evaluation,

[0843] Means for terminating communication in response to the aforementioned warning,

[0844] A system that includes this.

[0845] (Claim 2)

[0846] The system according to claim 1, wherein the evaluation means includes means for detecting specific terms and contexts related to fraudulent activity and scoring risk based thereon, and further includes means for incorporating the results of an analysis of emotional state into the risk assessment.

[0847] (Claim 3)

[0848] The system according to claim 1, wherein the warning means includes means for transmitting a warning message to a user or a related third party via digital communication means.

[0849] "Application example 2 when combining with an emotional engine"

[0850] (Claim 1)

[0851] A means of acquiring audio information in real time,

[0852] Means for converting the acquired audio information into document data,

[0853] A means for analyzing the aforementioned document data to evaluate the possibility of fraudulent activity,

[0854] A means for outputting a warning based on the aforementioned evaluation,

[0855] Means for terminating communication in response to the aforementioned warning,

[0856] A means of performing emotion analysis to identify the user's emotional state,

[0857] A means for scoring the risk of fraudulent activity based on the aforementioned emotional state,

[0858] A system that includes this.

[0859] (Claim 2)

[0860] The system according to claim 1, comprising means for scoring the likelihood of fraudulent activity by detecting specific indicators or contexts related to fraudulent activity.

[0861] (Claim 3)

[0862] The system according to claim 1, comprising means for transmitting a warning notice to a user or related stakeholders via electronic communication means. [Explanation of symbols]

[0863] 10, 210, 310, 410 Data Processing Systems 12 Data Processing Devices 14 Smart Devices 214 Smart Glasses 314 Headset-type terminal 414 Robots< / url:> < / url:> < / url:> < / url:>

Claims

1. A means of acquiring audio data in real time, Means for converting the acquired audio data into text data, A means for analyzing the aforementioned text data to evaluate the possibility of fraudulent activity, A means for outputting a warning based on the aforementioned evaluation, A means of terminating a call in response to the aforementioned warning, A system that includes this.

2. The system according to claim 1, wherein the evaluation means has means for scoring the likelihood of fraudulent activity by detecting specific keywords or contexts related to fraudulent activity.

3. The system according to claim 1, wherein the warning means includes means for transmitting a warning message to a user or a related third party via electronic communication means.

Citation Information

Patent Citations

  • Persona chatbot control method and system

    JP2022180282A