system

A system converts voice data to digital format, uses a generative model to detect fraud, sends immediate warnings, and records calls, effectively preventing telephone scams with enhanced emotion recognition.

JP2026068391APending Publication Date: 2026-04-22SOFTBANK GROUP CORP
View PDF 1 Cites 0 Cited by

Patent Information

Authority / Receiving Office
JP · JP
Patent Type
Applications
Current Assignee / Owner
SOFTBANK GROUP CORP
Filing Date
2024-10-10
Publication Date
2026-04-22

Smart Images

  • Figure 2026068391000001_ABST
    Figure 2026068391000001_ABST
Patent Text Reader

Abstract

We provide the system. [Solution] A terminal means for receiving user communications as voice signals, Means for converting the aforementioned audio signal into digital data and transmitting it over a network, An analytical means for analyzing transmitted digital data and identifying signs of fraudulent activity, A warning system that issues a warning when signs of such fraudulent activity are identified, A recording means for recording an audio signal and retaining it for a specified period of time, A system that includes this.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The technology of the present disclosure relates to a system.

Background Art

[0002] Patent Document 1 discloses a persona chatbot control method performed by at least one processor, including steps of receiving a user utterance, adding the user utterance to a prompt including an instruction sentence related to an explanation of a chatbot character, encoding the prompt, and inputting the encoded prompt into a language model to generate a chatbot utterance in response to the user utterance.

Prior Art Documents

Patent Documents

[0003]

Patent Document 1

Summary of the Invention

Problems to be Solved by the Invention

[0004] Illegal acts via telephone, especially fraud tactics, are widely recognized as a social problem, and while the sophistication of such tactics continues, new victims keep emerging. In order to prevent damage, it is necessary to build a real-time detection and warning system in addition to existing monitoring methods. However, with current technology, it is difficult to effectively meet these requirements, and there is a demand for a system that enables practical and rapid intervention.

Means for Solving the Problems

[0005] This invention provides a system that converts voice data into digital data when communication is initiated by a user's telephone terminal and transmits it to a central server via a network. This system utilizes a generative model to generate text data from voice signals and detects keywords and phrases related to fraudulent activities in real time. Based on these detection results, if it is determined that fraud is possible, a warning notification is sent to the user and registered recipients such as family members. In addition, all calls are recorded and stored for a certain period of time so that the content can be reviewed later and used as evidence if necessary. In this way, an advanced monitoring and warning system is realized that enables the prevention of fraudulent activity.

[0006] A "terminal device" is a device that receives audio signals when a user communicates and converts them into digital data.

[0007] "Means of converting to digital data" refers to a device or method that converts an analog audio signal into a digital format, making it usable for transmission over a communication network.

[0008] "Analysis means" refers to a device or program that has the function of analyzing received digital data and identifying signs of fraudulent activity from it.

[0009] A "warning device" is a device or system that sends a warning notification to the user and other recipients when fraudulent activity is detected based on the results of the analysis.

[0010] "Recording means" refers to a device or method that has the function of recording communication audio signals as digital data and storing them for a certain period of time.

[0011] A "generative model" is an algorithm or statistical model used to convert audio signals into text data and identify specific keywords or phrases. [Brief explanation of the drawing]

[0012] [Figure 1] This is a conceptual diagram showing an example of the configuration of a data processing system according to the first embodiment. [Figure 2] This is a conceptual diagram showing an example of the essential functions of a data processing device and a smart device according to the first embodiment. [Figure 3] This is a conceptual diagram showing an example of the configuration of a data processing system according to the second embodiment. [Figure 4] This is a conceptual diagram showing an example of the main functions of a data processing device and smart glasses according to the second embodiment. [Figure 5] This is a conceptual diagram showing an example of the configuration of a data processing system according to the third embodiment. [Figure 6] This is a conceptual diagram showing an example of the main functions of a data processing device and a headset-type terminal according to the third embodiment. [Figure 7] This is a conceptual diagram showing an example of the configuration of a data processing system according to the fourth embodiment. [Figure 8] This is a conceptual diagram showing an example of the main functions of a data processing device and a robot according to the fourth embodiment. [Figure 9] This shows an emotion map where multiple emotions are mapped. [Figure 10] This shows an emotion map where multiple emotions are mapped. [Figure 11] This is a sequence diagram showing the processing flow of the data processing system in Example 1. [Figure 12] This is a sequence diagram showing the processing flow of the data processing system in Application Example 1. [Figure 13] This is a sequence diagram showing the processing flow of the data processing system in Example 2, which incorporates an emotion engine. [Figure 14] This is a sequence diagram showing the processing flow of the data processing system in Application Example 2, which combines an emotion engine. [Modes for carrying out the invention]

[0013] An example of an embodiment of the system according to the technology of the present disclosure will be described below with reference to the accompanying drawings.

[0014] First, the terms used in the following description will be explained.

[0015] In the following embodiments, the numbered processor (hereinafter simply referred to as "processor") may be a single arithmetic unit or a combination of multiple arithmetic units. Also, the processor may be a single type of arithmetic unit or a combination of multiple types of arithmetic units. Examples of arithmetic units include a CPU (Central Processing Unit), a GPU (Graphics Processing Unit), a GPGPU (General-Purpose computing on Graphics Processing Units), an APU (Accelerated Processing Unit), and the like.

[0016] [[ID=?]]

[0017] In the following embodiments, the numbered RAM (Random Access Memory) is a memory in which information is temporarily stored and is used as a work memory by the processor.

[0018] It seems there is a formatting error in the original text where the tag is missing its corresponding text. I've translated the text as accurately as possible based on the provided content.In the following embodiments, the signed communication interface (I / F) is an interface that includes a communication processor and an antenna, etc. The communication interface manages communication between multiple computers. Examples of communication standards applicable to the communication interface include wireless communication standards such as 5G (5th Generation Mobile Communication System), Wi-Fi (registered trademark), or Bluetooth (registered trademark).

[0019] In the following embodiments, "A and / or B" is synonymous with "at least one of A and B." That is, "A and / or B" means that it may be A alone, or B alone, or a combination of A and B. Furthermore, in this specification, the same concept as "A and / or B" applies when expressing three or more things linked by "and / or."

[0020] [First Embodiment]

[0021] Figure 1 shows an example of the configuration of the data processing system 10 according to the first embodiment.

[0022] As shown in Figure 1, the data processing system 10 includes a data processing device 12 and a smart device 14. An example of the data processing device 12 is a server.

[0023] The data processing device 12 comprises a computer 22, a database 24, and a communication interface 26. The computer 22 is an example of a "computer" related to the technology of this disclosure. The computer 22 comprises a processor 28, RAM 30, and storage 32. The processor 28, RAM 30, and storage 32 are connected to a bus 34. The database 24 and the communication interface 26 are also connected to the bus 34. The communication interface 26 is connected to a network 54. An example of the network 54 is a WAN (Wide Area Network) and / or a LAN (Local Area Network).

[0024] The smart device 14 comprises a computer 36, a reception device 38, an output device 40, a camera 42, and a communication interface 44. The computer 36 comprises a processor 46, RAM 48, and storage 50. The processor 46, RAM 48, and storage 50 are connected to a bus 52. The reception device 38, output device 40, and camera 42 are also connected to the bus 52.

[0025] The reception device 38 is equipped with a touch panel 38A and a microphone 38B, etc., and receives user input. The touch panel 38A receives user input by detecting contact with an object (e.g., a pen or finger). The microphone 38B receives user input by detecting the user's voice. The control unit 46A transmits data indicating the user input received by the touch panel 38A and microphone 38B to the data processing device 12. In the data processing device 12, the specific processing unit 290 acquires the data indicating the user input.

[0026] The output device 40 includes a display 40A and a speaker 40B, and presents data to the user 20 by outputting the data in a form perceptible to the user 20 (e.g., audio and / or text). The display 40A displays visible information such as text and images according to instructions from the processor 46. The speaker 40B outputs audio according to instructions from the processor 46. The camera 42 is a small digital camera equipped with an optical system such as a lens, aperture, and shutter, and an image sensor such as a CMOS (Complementary Metal-Oxide-Semiconductor) image sensor or a CCD (Charge Coupled Device) image sensor.

[0027] Communication interface 44 is connected to network 54. Communication interfaces 44 and 26 are responsible for the exchange of various types of information between processor 46 and processor 28 via network 54.

[0028] Figure 2 shows an example of the main functions of the data processing device 12 and the smart device 14.

[0029] As shown in Figure 2, in the data processing device 12, specific processing is performed by the processor 28. A specific processing program 56 is stored in the storage 32. The specific processing program 56 is an example of a "program" related to the technology of this disclosure. The processor 28 reads the specific processing program 56 from the storage 32 and executes the read specific processing program 56 on the RAM 30. The specific processing is realized by the processor 28 operating as a specific processing unit 290 according to the specific processing program 56 executed on the RAM 30.

[0030] The storage 32 stores the data generation model 58 and the emotion identification model 59. The data generation model 58 and the emotion identification model 59 are used by the identification processing unit 290.

[0031] In the smart device 14, the processor 46 performs the reception output processing. The storage 50 stores the reception output program 60. The reception output program 60 is used in conjunction with a specific processing program 56 by the data processing system 10. The processor 46 reads the reception output program 60 from the storage 50 and executes the read reception output program 60 on the RAM 48. The reception output processing is realized by the processor 46 operating as a control unit 46A according to the reception output program 60 executed on the RAM 48.

[0032] Next, the specific processing performed by the specific processing unit 290 of the data processing device 12 will be described. In the following description, the data processing device 12 will be referred to as the "server" and the smart device 14 as the "terminal".

[0033] This invention provides a system for preventing fraudulent activities via voice communication from a terminal used by a user. The main components of this system are based on three entities: a terminal, a server, and a user.

[0034] When a user initiates a phone call, the terminal converts the voice signal into digital data in real time. The converted digital data is then sent to the server using a secure protocol.

[0035] On the server, a generative model runs to analyze the received digital data. This generative model has the ability to convert voice signals into text and identify keywords and phrases that may be related to fraud. For example, if phrases such as "ATM," "bank transfer," and "urgent" are detected, the likelihood that the call is associated with fraudulent activity is assessed.

[0036] If potential fraudulent activity is detected, the server will generate an alert through its warning system and send a warning notification to the user's device and the devices of their pre-registered family members. This notification will display a message such as, "This call may contain fraudulent activity. Please review the content."

[0037] The user or their family can immediately review the content of a call based on the received warning notification. Calls are always recorded, and the recorded audio files are stored on the server for a certain period, allowing them to be played back and reviewed as needed.

[0038] Through the process described above, this system aims to prevent telephone-based fraud by detecting fraudulent activity in real time and enabling a rapid response.

[0039] The following describes the processing flow.

[0040] Step 1:

[0041] The terminal receives the user's call and converts the audio signal into digital data. The converted digital data is then sent to the server via a secure protocol.

[0042] Step 2:

[0043] The server analyzes the received digital data and uses a generative model to convert the audio signal into text data. From this text data, keywords and phrases related to fraudulent activities are detected.

[0044] Step 3:

[0045] The server scores the likelihood of fraud and prepares to generate a warning alert if the likelihood of fraud exceeds a threshold.

[0046] Step 4:

[0047] The server uses a warning mechanism to send a warning notification to the user's device and the devices of registered family members, alerting the user.

[0048] Step 5:

[0049] The system utilizes a function that allows users to review warning notifications they receive and then play back recorded call data stored on the server to examine its contents.

[0050] Step 6:

[0051] If a user or their family member suspects fraud, they should take action, such as reporting it to the police, to resolve the issue.

[0052] (Example 1)

[0053] Next, we will describe Example 1. In the following description, the data processing device 12 will be referred to as the "server," and the smart device 14 will be referred to as the "terminal."

[0054] Addressing fraud and malicious communications via telephone is crucial, but currently, there are limited means to detect these activities in real time and issue rapid warnings. This means many users could become victims of fraud and suffer serious damage. To solve this problem, an effective system is needed that can quickly and accurately detect fraudulent activity through voice communications and immediately notify users.

[0055] The identification process performed by the identification processing unit 290 of the data processing device 12 in Example 1 is realized by the following means.

[0056] In this invention, the server includes means for receiving voice signals and converting them into digital information, processing means for analyzing the transmitted digital information to identify signs of fraudulent activity, and notification means for issuing warnings to appropriate recipients. This makes it possible to detect fraudulent activity via voice communication in real time and immediately issue warnings to users and related parties.

[0057] "Electronic device" refers to a device that has the function of receiving communications from a user as audio signals and converting them into digital information.

[0058] A "data network" refers to a communication network used to securely transmit digital information from electronic devices to servers.

[0059] "Processing means" refers to equipment equipped with the function of analyzing transmitted digital information and performing a process to identify signs of specific fraudulent activity.

[0060] "Notification means" refers to a system that has the function of issuing warnings to users and related parties based on identified signs of fraudulent activity.

[0061] "Storage means" refers to a device or system that records audio signals and retains that recorded data for a predetermined period of time.

[0062] This invention is a system that analyzes the content of a voice communication in real time when the user initiates a voice communication, thereby preventing fraudulent activity. The system mainly consists of three elements: a terminal, a server, and a user.

[0063] When a user initiates voice communication, the terminal captures the audio signal using its built-in microphone and converts the analog signal into digital data using audio processing software such as "FFmpeg". This digital data is then transmitted to the server via the data network using the SSL / TLS protocol.

[0064] The server utilizes a natural language processing API to process the received digital information. This generative AI model converts the audio signal into text data. Subsequently, a script developed in Python is used to identify specific keywords and phrases (e.g., "ATM," "bank transfer," "urgent") from the generated text data. If potential fraud is detected, the server uses services such as Twilio to send warning notifications to the user and pre-registered recipients.

[0065] As an example, suppose a user receives a suspicious telemarketing call instructing them to "transfer the money immediately." In this case, the system detects the relevant phrase and promptly sends a warning notification to the user. This allows the user to recognize the danger of the call and take swift action.

[0066] Examples of prompt statements for a generative AI model are as follows:

[0067] "Is this call potentially a scam? Detected keywords: [ATM, bank transfer, urgent]"

[0068] This system allows for real-time monitoring of fraudulent activity via voice communication, enabling users to take prompt and appropriate action.

[0069] The flow of the specific processing in Example 1 will be explained using Figure 11.

[0070] Step 1:

[0071] When a user initiates voice communication, the device captures the audio signal using its built-in microphone. The input is an analog audio signal, and the device prepares to convert the acquired audio into digital data. Specifically, it uses a process such as FFmpeg to convert the analog signal into digital information, and the output is digital data in PCM format.

[0072] Step 2:

[0073] The terminal transmits the converted digital data to the server using a security protocol (SSL / TLS). The input is digital data in PCM format; this data is encrypted, packetized, and securely transferred to the server via the data network. The output is encrypted data packets.

[0074] Step 3:

[0075] The server receives digital data transmitted from the terminal and converts it from speech to text. The input is encrypted data packets, which are decrypted and then converted into text data using a natural language processing API. As a result of the data analysis, the output is text data corresponding to the speech.

[0076] Step 4:

[0077] The server uses a generative AI model to detect specific keywords and phrases from the obtained text data. The input is text data, and a script written in Python searches for pre-configured fraud-related keywords (e.g., "ATM," "bank transfer," "urgent"). The output is a list of the keywords found.

[0078] Step 5:

[0079] The server evaluates the potential for malicious activity and generates a warning notification when specific keywords are detected. The input is a list of detected keywords, and warning messages are sent to the user and registered recipients using services such as Twilio. The output is the sending of warning messages and a record of them.

[0080] Step 6:

[0081] The user or recipient receives a warning and reviews the content of the call. The input is the received warning message, and the server plays back the recorded data to examine its contents. The server stores the call records, which can be played back and reviewed as needed. The output is the reviewed call content and its evaluation result.

[0082] (Application Example 1)

[0083] Next, we will explain Application Example 1. In the following explanation, the data processing device 12 will be referred to as the "server," and the smart device 14 will be referred to as the "terminal."

[0084] In modern society, fraudulent activities conducted via acoustic communication are increasingly causing serious harm to users. Existing technologies have made it difficult to detect and warn of fraudulent activities in real time, therefore, there is a need to provide effective means to prevent fraud before it occurs.

[0085] The specific processing performed by the specific processing unit 290 of the data processing device 12 in Application Example 1 is realized by the following means.

[0086] In this invention, the server includes means for converting acoustic signals into digital data, means for analyzing the transmitted digital data and identifying signs of fraudulent activity using a generative model, and means for displaying a warning notification when fraudulent activity is identified and sending the warning to the recipient. This makes it possible to detect fraudulent activity during acoustic communication in real time and immediately warn the user.

[0087] A "user" is an individual or group that uses the system to conduct acoustic communication.

[0088] "Communication" is the process of sending and receiving acoustic signals between users.

[0089] An "acoustic signal" refers to a signal that represents the user's speech or voice.

[0090] "Device means" refers to hardware or software components used to receive the user's acoustic signals.

[0091] "Digitized data" refers to data obtained after converting an acoustic signal into a digital format.

[0092] A "wide-area communication network" is a network infrastructure used to transmit digitized data.

[0093] "Analysis means" refers to the processes and tools used to analyze the transmitted digitized data.

[0094] A "generative model" is a machine learning algorithm used to generate character data from received acoustic signals.

[0095] "Signs of fraudulent activity" refers to the identification of keywords or patterns that may be associated with fraud or crime.

[0096] A "warning mechanism" refers to a structure or method for issuing a warning when signs of fraudulent activity are detected.

[0097] "Visualization means" refers to devices or software that visually display warning notifications to the user.

[0098] "Recording means" refers to equipment or technology for recording acoustic signals and retaining them for a specified period of time.

[0099] A "recipient" is a user or their associate who has been pre-registered to receive warning notifications.

[0100] The system for realizing this invention provides a function to detect fraudulent activity during acoustic communication in real time.

[0101] First, when a user performs acoustic communication, the terminal receives an acoustic signal. The received acoustic signal is converted into digital data using the terminal's speech recognition engine. It is expected that existing speech recognition services such as Google® Speech-to-Text API will be used for this conversion. The converted digital data is then transmitted to a server via a wide-area communication network (e.g., the internet). In this process, the secure protocol HTTPS is used to ensure that the data is transferred safely.

[0102] On the server, a generative model operates to analyze the received digitized data. The generative model uses machine learning algorithms such as OpenAI® GPT-3® to generate text data from acoustic signals. Next, this text data is analyzed to identify signs of fraudulent activity. The analysis process monitors for predefined fraud-related keywords, and if these keywords are present, they are identified as signs of fraudulent activity.

[0103] When identification occurs, the server sends a warning notification to the user and pre-registered recipients via a warning mechanism. The warning notification is immediately displayed on the terminal's visualization mechanism, allowing the user to check the communication content and take necessary measures. In addition, all acoustic signals are recorded and stored by a recording mechanism, and can be played back and reviewed later as needed.

[0104] For example, if a user receives suspicious instructions during an audio communication, such as "Please make a transfer immediately," the system will display a warning that it "may be fraudulent." This allows the user to immediately realize it is a scam and prevent becoming a victim.

[0105] An example of a prompt to a generative AI model might be something like, "Analyze the phrases used in this call and assess the likelihood of fraud."

[0106] Therefore, this system provides an effective means of reducing the risk of fraudulent activity via acoustic communication and ensuring user safety.

[0107] The flow of a specific process in Application Example 1 will be explained using Figure 12.

[0108] Step 1:

[0109] The terminal receives an acoustic signal from the user. The input is the user's voice, and the output is a raw acoustic signal. The terminal captures the acoustic signal using a microphone.

[0110] Step 2:

[0111] The device converts the received acoustic signal into digital data. The input is the acoustic signal acquired in step 1, and the output is the digital data. This conversion is performed using a speech recognition engine (e.g., Google Speech-to-Text API). The device encodes the audio features into a digital format.

[0112] Step 3:

[0113] The terminal transmits the converted digital data to the server via a wide-area communication network. The input is digital data, and the output is data transfer via a secure protocol (e.g., HTTPS). The terminal transmits the data using a network interface.

[0114] Step 4:

[0115] The server launches a generative model to analyze the received digitized data. The input is digitized data, and the output is the analysis result. The server uses a generative AI model (e.g., OpenAI GPT-3) to generate character data from acoustic signals.

[0116] Step 5:

[0117] The server identifies signs of fraudulent activity based on the generated text data. The input is text data, and the output is the identification result of the potential for fraudulent activity. The server uses an algorithm that checks for fraud-related keywords to identify specific phrases and patterns.

[0118] Step 6:

[0119] The server generates a warning notification and sends it to the user and registered recipients if it detects signs of fraudulent activity. The input is the identification result, and the output is the warning notification. The server generates the notification message and delivers it to the specified recipients.

[0120] Step 7:

[0121] The terminal displays received warning notifications. The input is the warning notification, and the output is the display of the warning message on the user interface. The terminal communicates the warning to the user using either a display or audio output.

[0122] Step 8:

[0123] The server records all audio signals and retains them for a specified period. The input is the audio signal, and the output is the stored audio recording. The server uses a database or storage system to save the data in a format that can be accessed later.

[0124] Furthermore, an emotion engine that estimates the user's emotions may be incorporated. That is, the identification processing unit 290 may use the emotion identification model 59 to estimate the user's emotions and perform identification processing using the user's emotions.

[0125] This invention provides a voice communication fraud prevention system that combines an emotion engine. This system focuses on the terminal, server, and user, and achieves more precise fraud detection and response by comprehensively analyzing the voice communications made by the user.

[0126] When a user initiates communication, the terminal converts the voice signal into digital data and sends it to the server in real time. The server receives this digital data and performs analysis. For the analysis, a generative model is used to convert the voice signal into text data and detect keywords and phrases that may be related to fraud.

[0127] Furthermore, this system incorporates an emotion engine that estimates emotions from the user's voice. The emotion engine particularly identifies emotions such as anxiety, tension, and agitation, and uses this to improve the accuracy of the analysis. For example, when a user utters words expressing anxiety such as "I'm in trouble" or "What should I do?", the system analyzes the tone, speed, and volume of their voice to understand their emotional state.

[0128] The server generates alerts based on signs of fraudulent activity and emotional shifts. If the emotion engine detects strong signs of anxiety or tension, it adjusts the level and urgency of the warning notification and promptly alerts the user and registered recipients. This warning may include specific details such as, "This call may be fraudulent and anxiety has been detected. Please investigate immediately."

[0129] The user or their family can review the content of the call and play back the recording for further analysis based on the warning notification sent. They can also take appropriate measures, such as reporting to the police, if necessary.

[0130] The system of this invention makes it possible to prevent damage from telephone-based fraud by providing dual protection through real-time analysis and emotion recognition. For example, if a user suddenly becomes anxious and starts worrying about their money, the system will highly assess the possibility of fraud, immediately generate a warning, and notify family members to encourage appropriate action.

[0131] The following describes the processing flow.

[0132] Step 1:

[0133] When the terminal initiates a call with a user, it converts the audio signal into a digital format and transmits it to the server in real time using a secure protocol.

[0134] Step 2:

[0135] The server analyzes the received digital data using a generative model and converts the audio signal into text data. It then processes the text data to detect fraud-related keywords and phrases.

[0136] Step 3:

[0137] An emotion engine installed on the server estimates the user's emotions from their voice. This engine analyzes the tone, speed, and volume of the voice to identify the emotions the user is expressing (such as anxiety or tension).

[0138] Step 4:

[0139] The server integrates the analysis results and sentiment estimation results to score the likelihood of fraud. If signs of fraudulent activity or emotional instability are detected, a warning alert is prepared.

[0140] Step 5:

[0141] The server sends a warning notification to the user's device and the devices of registered recipients. The notification includes warnings about potential fraud and emotional instability.

[0142] Step 6:

[0143] The system will review the warning notification received by the user or a family member, and play back the recorded data stored on the server to examine the call content. If fraud is suspected, the system will promptly request appropriate action.

[0144] (Example 2)

[0145] Next, we will describe Example 2. In the following description, the data processing device 12 will be referred to as the "server" and the smart device 14 as the "terminal".

[0146] Fraudulent activity in voice communications presents a challenge because the content of these communications is diverse, making it difficult to prevent simply by detecting keywords. Furthermore, since a user's psychological state and emotions may be related to fraudulent activity, there is a need for highly accurate fraud detection that also takes emotional changes into account.

[0147] The identification process performed by the identification processing unit 290 of the data processing device 12 in Example 2 is realized by the following means.

[0148] In this invention, the server includes means for generating electrical data, means for estimating emotional states using emotion analysis means, and means for issuing warnings based on signs of misconduct and changes in emotion. This makes it possible to detect misconduct in voice communications with high accuracy while taking emotional states into consideration and to respond quickly.

[0149] "Acoustic signal" refers to the physical sound vibrations generated when the user's voice is captured.

[0150] "Device means" refers to a device for receiving acoustic signals and converting them into digital data.

[0151] "Electrical data" refers to audio information in digital format that has been converted by a device or means.

[0152] A "data path" refers to a means of communication for transmitting electrical data.

[0153] "Analysis" refers to the process of processing received electrical data and evaluating signs of fraudulent activity.

[0154] "Analysis means" refers to a method or device for analyzing electrical data and identifying fraudulent activity.

[0155] "Recording means" refers to a function that saves audio signals in digital format and allows them to be played back later.

[0156] "Emotion analysis means" refers to a method or device for estimating emotions from a user's voice and recognizing changes in those emotions.

[0157] A "warning" refers to a message that alerts users or other relevant parties to suspected fraudulent activity.

[0158] This invention provides a system for preventing fraudulent activities in voice communications, and implements this using a user, a terminal, and a server.

[0159] The user initiates voice communication and inputs their voice through the terminal. The terminal converts the input acoustic signal into electrical data. This process is achieved using the audio processor built into the terminal.

[0160] The terminal transmits the converted electrical data to the server. A secure and efficient data path is ensured for communication, and generally, communication methods using encryption technology are employed.

[0161] The server analyzes the received electrical data. A generative AI model is used for the analysis, generating text data from the audio. Based on this text data, natural language processing (NLP) techniques are used to detect keywords related to fraud. The server also utilizes sentiment analysis to estimate the user's emotional state from the audio data. Sentiment analysis determines the psychological state from the tone, speed, and volume of the voice, identifying emotions such as anxiety and tension.

[0162] Based on the analysis above, the server generates a warning when it detects signs of fraudulent activity and changes in mood. The warning is sent to the user and registered recipients, prompting them to take necessary action. For example, the warning may include a specific message such as, "This call may be fraudulent and anxiety has been detected. Please check immediately."

[0163] This system's real-time fraud detection and rapid response capabilities allow users and their families to prevent fraud and improve the security of voice communications. A specific use case would be if a user, speaking in a tense tone during a call, says, "I'm worried about my finances." The system would immediately detect the risk and send a warning notification to the family.

[0164] An example of a prompt message would be, "Analyze the voice that expresses the user's anxiety and assess the likelihood of fraud." This prompt allows the generative AI model to appropriately evaluate the user's psychological state and the content of their statements, enabling sophisticated fraud detection.

[0165] The flow of the specific processing in Example 2 will be explained using Figure 13.

[0166] Step 1:

[0167] The user inputs an acoustic signal to the terminal to initiate voice communication. This acoustic signal is then generated.

[0168] Step 2:

[0169] The terminal converts the received acoustic signal into electrical data using a digital audio processor. This process transforms the signal into digital data. The output is the converted electrical data.

[0170] Step 3:

[0171] The terminal transmits the converted electrical data to the server via the data path. This allows the server to obtain the data necessary for analysis.

[0172] Step 4:

[0173] The server converts received electrical data into text data using a generation AI model. The input is electrical data, and the output is the generated text data. This conversion process is performed via a speech recognition engine.

[0174] Step 5:

[0175] The server uses natural language processing techniques to detect fraud-related keywords from text data. The input is text data, and the output is the identified keywords or their absence.

[0176] Step 6:

[0177] The server estimates emotions using emotion analysis tools. The input is electrical data, and the output is the estimated emotional state. Specifically, it analyzes the tone and speed of speech to identify feelings of anxiety and tension.

[0178] Step 7:

[0179] The server generates an alert based on an assessment of signs of misconduct and changes in sentiment. The alert is customized using prompts as needed. The content of the alert message is generated at this stage.

[0180] Step 8:

[0181] The server sends the generated warnings to the user and registered recipients. The output is the specific warning message that the user and recipients receive. This warning gives the user an opportunity to take corrective action.

[0182] (Application Example 2)

[0183] Next, we will explain application example 2. In the following explanation, the data processing device 12 will be referred to as a "server" and the smart device 14 as a "terminal".

[0184] In recent years, with the advancement of communication technology, voice-based fraud has been increasing, posing a serious problem, especially for vulnerable individuals such as the elderly. To effectively protect users from such fraud, there is a need for a system that can quickly and accurately detect signs of fraudulent activity and also provide warnings in response to changes in the user's emotional state.

[0185] The specific processing performed by the specific processing unit 290 of the data processing device 12 in Application Example 2 is realized by the following means.

[0186] In this invention, the server includes a device means for converting voice signals into digital data and transmitting them over a network, an analysis means for analyzing the transmitted digital data to identify signs of fraudulent activity, and a device means for estimating the user's emotional state using an emotion analysis engine and detecting anxiety and tension. This makes it possible to detect signs that should raise suspicion of fraudulent activity during the user's voice communication, understand the user's emotional changes at that time, and quickly display and transmit appropriate warnings to the user's visual device and pre-registered receiving equipment.

[0187] "Device means" refers to equipment that has the function of receiving communications as audio signals, converting them into digital data, and transmitting them over a network.

[0188] An "analysis device" is a device that has the function of identifying signs of fraudulent activity by analyzing transmitted digital data.

[0189] A "notification system" is a system that sends a warning to users or registered receiving equipment when signs of fraudulent activity are identified.

[0190] A "memory device" is a device that has the function of recording audio signals and retaining them for a specified period of time.

[0191] An "emotion analysis engine" is a system that estimates a user's emotional state and detects anxiety or tension by analyzing the tone, speed, and volume of their voice.

[0192] A "visual device" is a device used to display visual information to a user.

[0193] This invention provides a system for preventing fraudulent activities in voice communications, which has a configuration that comprehensively analyzes the content of user communications. The main components of the system are digitization of voice signals, data analysis, fraud detection, emotion recognition, and warning issuance.

[0194] The server receives the audio signal sent from the terminal when the user initiates communication. The received audio signal is first converted into digital data. This process uses a real-time audio analysis engine (e.g., Nuance Dragon) to handle the audio data.

[0195] The server then uses a generative AI model to generate text data from the audio data. This generative AI model (e.g., OpenAI GPT) has the ability to analyze the text data and identify keywords and phrases related to fraudulent activity. Simultaneously, the server uses an emotion analysis engine (e.g., Affectiva) to estimate the user's emotional state from their voice and detect anxiety and tension.

[0196] The user's device is equipped with a visual device (e.g., smart glasses) to display warnings visually. When the system detects signs of misconduct and changes in the user's emotional state, it displays a warning on the user's visual device and immediately sends a notification to registered receiving equipment.

[0197] As a concrete example, consider a scenario where a user is on a regular phone call and the other party approaches them with a "good offer about investing in a new financial product." In this case, the server transcribes the other party's words in real time, detects the possibility of fraud, and also detects feelings of tension or anxiety from the user's voice. Based on this, the system displays a warning such as "This may be a scam, please be careful" on the user's visual device and sends a notification to the devices of registered recipients.

[0198] Examples of prompts for the generative AI model include, "Identify potentially fraudulent phrases that can be detected from this audio data," and "Analyze the emotional shifts following the user's statements and report any signs of anxiety or tension."

[0199] In this way, the server can ensure user safety and prevent fraud.

[0200] The flow of a specific process in Application Example 2 will be explained using Figure 14.

[0201] Step 1:

[0202] When a user initiates communication, the terminal captures the audio signal using its built-in microphone. This signal is then converted into digital data in real time. This digital data is then input to the server and sent to the audio analysis engine.

[0203] Step 2:

[0204] The server analyzes the received digital data and uses a generative AI model to convert the audio data into text data. The generative AI model extracts features from the audio signal and uses prompt sentences to generate text. The output of this step is the transcribed audio data.

[0205] Step 3:

[0206] The server analyzes the generated text data to identify keywords and phrases related to fraud. This process compares specific words indicating fraudulent activity against a database to generate identification results. The output is a flag indicating the presence or absence of fraudulent activity.

[0207] Step 4:

[0208] Simultaneously, the server processes the user's voice signal using an emotion analysis engine to estimate their emotional state. Specifically, it analyzes the tone, speed, and volume of the voice to estimate the level of anxiety and tension. The output is an indicator of the user's emotional state.

[0209] Step 5:

[0210] The server issues warnings based on flags indicating fraudulent activity and emotional state indicators. It instructs visual devices to display warning messages notifying them of potential fraud and emotional changes. It also sends similar warning notifications to registered receiving devices.

[0211] Step 6:

[0212] Users receive a warning and review feedback regarding changes in call content and emotional state. This allows them to determine whether actual misconduct occurred and take appropriate action.

[0213] This allows the entire system to work together to protect users from fraudulent voice communication.

[0214] The specific processing unit 290 transmits the result of the specific processing to the smart device 14. In the smart device 14, the control unit 46A causes the output device 40 to output the result of the specific processing. The microphone 38B acquires audio indicating user input for the result of the specific processing. The control unit 46A transmits the audio data indicating user input acquired by the microphone 38B to the data processing device 12. In the data processing device 12, the specific processing unit 290 acquires the audio data.

[0215] Data generation model 58 is a so-called generative AI (Artificial Intelligence). An example of data generation model 58 is ChatGPT (registered trademark) (Internet search).<URL: https: / / openai.com / blog / chatgpt> ), Gemini (registered trademark) (Internet search) <url: https: gemini.google.com ?hl="ja">Examples of generative AI include the following. The data generation model 58 is obtained by performing deep learning on a neural network. The data generation model 58 is input with prompts containing instructions, and with inference data such as audio data representing speech, text data representing text, and image data representing images. The data generation model 58 infers from the input inference data according to the instructions indicated by the prompts, and outputs the inference results in data formats such as audio data and text data. Here, inference refers to, for example, analysis, classification, prediction, and / or summarization.

[0216] In the above embodiment, an example was given in which specific processing is performed by the data processing device 12, but the technology of this disclosure is not limited thereto, and the specific processing may also be performed by the smart device 14.

[0217] [Second Embodiment]

[0218] Figure 3 shows an example of the configuration of the data processing system 210 according to the second embodiment.

[0219] As shown in Figure 3, the data processing system 210 includes a data processing device 12 and smart glasses 214. An example of the data processing device 12 is a server.

[0220] The data processing device 12 comprises a computer 22, a database 24, and a communication interface 26. The computer 22 is an example of a "computer" related to the technology of this disclosure. The computer 22 comprises a processor 28, RAM 30, and storage 32. The processor 28, RAM 30, and storage 32 are connected to a bus 34. The database 24 and the communication interface 26 are also connected to the bus 34. The communication interface 26 is connected to a network 54. An example of the network 54 is a WAN (Wide Area Network) and / or a LAN (Local Area Network).

[0221] The smart glasses 214 include a computer 36, a microphone 238, a speaker 240, a camera 42, and a communication interface 44. The computer 36 includes a processor 46, RAM 48, and storage 50. The processor 46, RAM 48, and storage 50 are connected to a bus 52. The microphone 238, speaker 240, and camera 42 are also connected to the bus 52.

[0222] The microphone 238 receives voice signals from the user 20 and receives instructions from the user 20. The microphone 238 captures the voice signals from the user 20, converts the captured voice into audio data, and outputs it to the processor 46. The speaker 240 outputs audio according to the instructions from the processor 46.

[0223] Camera 42 is a small digital camera equipped with an optical system including a lens, aperture, and shutter, and an image sensor such as a CMOS (Complementary Metal-Oxide-Semiconductor) image sensor or a CCD (Charge Coupled Device) image sensor, and captures images of the area around the user 20 (for example, an imaging range defined by a field of view equivalent to the width of a typical healthy person's field of vision).

[0224] Communication interface 44 is connected to network 54. Communication interfaces 44 and 26 are responsible for the exchange of various information between processor 46 and processor 28 via network 54. The exchange of various information between processor 46 and processor 28 using communication interfaces 44 and 26 is performed in a secure manner.

[0225] Figure 4 shows an example of the main functions of the data processing device 12 and the smart glasses 214. As shown in Figure 4, the data processing device 12 performs specific processing using the processor 28. The storage 32 stores the specific processing program 56.

[0226] The specific processing program 56 is an example of a "program" relating to the technology of this disclosure. The processor 28 reads the specific processing program 56 from the storage 32 and executes the read specific processing program 56 on the RAM 30. The specific processing is realized by the processor 28 operating as a specific processing unit 290 in accordance with the specific processing program 56 executed on the RAM 30.

[0227] The storage 32 stores the data generation model 58 and the emotion identification model 59. The data generation model 58 and the emotion identification model 59 are used by the identification processing unit 290.

[0228] In the smart glasses 214, the processor 46 performs the reception output processing. The storage 50 stores the reception output program 60. The processor 46 reads the reception output program 60 from the storage 50 and executes the read reception output program 60 on the RAM 48. The reception output processing is realized by the processor 46 operating as a control unit 46A according to the reception output program 60 executed on the RAM 48.

[0229] Next, the identification processing performed by the identification processing unit 290 of the data processing device 12 will be described. In the following description, the data processing device 12 will be referred to as the "server" and the smart glasses 214 will be referred to as the "terminal".

[0230] This invention provides a system for preventing fraudulent activities via voice communication from a terminal used by a user. The main components of this system are based on three entities: a terminal, a server, and a user.

[0231] When a user initiates a phone call, the terminal converts the voice signal into digital data in real time. The converted digital data is then sent to the server using a secure protocol.

[0232] On the server, a generative model runs to analyze the received digital data. This generative model has the ability to convert voice signals into text and identify keywords and phrases that may be related to fraud. For example, if phrases such as "ATM," "bank transfer," and "urgent" are detected, the likelihood that the call is associated with fraudulent activity is assessed.

[0233] If potential fraudulent activity is detected, the server will generate an alert through its warning system and send a warning notification to the user's device and the devices of their pre-registered family members. This notification will display a message such as, "This call may contain fraudulent activity. Please review the content."

[0234] The user or their family can immediately review the content of a call based on the received warning notification. Calls are always recorded, and the recorded audio files are stored on the server for a certain period, allowing them to be played back and reviewed as needed.

[0235] Through the process described above, this system aims to prevent telephone-based fraud by detecting fraudulent activity in real time and enabling a rapid response.

[0236] The following describes the processing flow.

[0237] Step 1:

[0238] The terminal receives the user's call and converts the audio signal into digital data. The converted digital data is then sent to the server via a secure protocol.

[0239] Step 2:

[0240] The server analyzes the received digital data and uses a generative model to convert the audio signal into text data. From this text data, keywords and phrases related to fraudulent activities are detected.

[0241] Step 3:

[0242] The server scores the likelihood of fraud and prepares to generate a warning alert if the likelihood of fraud exceeds a threshold.

[0243] Step 4:

[0244] The server uses a warning mechanism to send a warning notification to the user's device and the devices of registered family members, alerting the user.

[0245] Step 5:

[0246] The system utilizes a function that allows users to review warning notifications they receive and then play back recorded call data stored on the server to examine its contents.

[0247] Step 6:

[0248] If a user or their family member suspects fraud, they should take action, such as reporting it to the police, to resolve the issue.

[0249] (Example 1)

[0250] Next, we will describe Example 1. In the following description, the data processing device 12 will be referred to as the "server," and the smart glasses 214 will be referred to as the "terminal."

[0251] Addressing fraud and malicious communications via telephone is crucial, but currently, there are limited means to detect these activities in real time and issue rapid warnings. This means many users could become victims of fraud and suffer serious damage. To solve this problem, an effective system is needed that can quickly and accurately detect fraudulent activity through voice communications and immediately notify users.

[0252] The identification process performed by the identification processing unit 290 of the data processing device 12 in Example 1 is realized by the following means.

[0253] In this invention, the server includes means for receiving voice signals and converting them into digital information, processing means for analyzing the transmitted digital information to identify signs of fraudulent activity, and notification means for issuing warnings to appropriate recipients. This makes it possible to detect fraudulent activity via voice communication in real time and immediately issue warnings to users and related parties.

[0254] "Electronic device" refers to a device that has the function of receiving communications from a user as audio signals and converting them into digital information.

[0255] A "data network" refers to a communication network used to securely transmit digital information from electronic devices to servers.

[0256] "Processing means" refers to equipment equipped with the function of analyzing transmitted digital information and performing a process to identify signs of specific fraudulent activity.

[0257] "Notification means" refers to a system that has the function of issuing warnings to users and related parties based on identified signs of fraudulent activity.

[0258] "Storage means" refers to a device or system that records audio signals and retains that recorded data for a predetermined period of time.

[0259] This invention is a system that analyzes the content of a voice communication in real time when the user initiates a voice communication, thereby preventing fraudulent activity. The system mainly consists of three elements: a terminal, a server, and a user.

[0260] When a user initiates voice communication, the terminal captures the audio signal using its built-in microphone and converts the analog signal into digital data using audio processing software such as "FFmpeg". This digital data is then transmitted to the server via the data network using the SSL / TLS protocol.

[0261] The server utilizes a natural language processing API to process the received digital information. This generative AI model converts the audio signal into text data. Subsequently, a script developed in Python is used to identify specific keywords and phrases (e.g., "ATM," "bank transfer," "urgent") from the generated text data. If potential fraud is detected, the server uses services such as Twilio to send warning notifications to the user and pre-registered recipients.

[0262] As an example, suppose a user receives a suspicious telemarketing call instructing them to "transfer the money immediately." In this case, the system detects the relevant phrase and promptly sends a warning notification to the user. This allows the user to recognize the danger of the call and take swift action.

[0263] Examples of prompt statements for a generative AI model are as follows:

[0264] "Is this call potentially a scam? Detected keywords: [ATM, bank transfer, urgent]"

[0265] This system allows for real-time monitoring of fraudulent activity via voice communication, enabling users to take prompt and appropriate action.

[0266] The flow of the specific processing in Example 1 will be explained using Figure 11.

[0267] Step 1:

[0268] When a user initiates voice communication, the device captures the audio signal using its built-in microphone. The input is an analog audio signal, and the device prepares to convert the acquired audio into digital data. Specifically, it uses a process such as FFmpeg to convert the analog signal into digital information, and the output is digital data in PCM format.

[0269] Step 2:

[0270] The terminal transmits the converted digital data to the server using a security protocol (SSL / TLS). The input is digital data in PCM format; this data is encrypted, packetized, and securely transferred to the server via the data network. The output is encrypted data packets.

[0271] Step 3:

[0272] The server receives digital data transmitted from the terminal and converts it from speech to text. The input is encrypted data packets, which are decrypted and then converted into text data using a natural language processing API. As a result of the data analysis, the output is text data corresponding to the speech.

[0273] Step 4:

[0274] The server uses a generative AI model to detect specific keywords and phrases from the obtained text data. The input is text data, and a script written in Python searches for pre-configured fraud-related keywords (e.g., "ATM," "bank transfer," "urgent"). The output is a list of the keywords found.

[0275] Step 5:

[0276] The server evaluates the potential for malicious activity and generates a warning notification when specific keywords are detected. The input is a list of detected keywords, and warning messages are sent to the user and registered recipients using services such as Twilio. The output is the sending of warning messages and a record of them.

[0277] Step 6:

[0278] The user or recipient receives a warning and checks the content of the call. The input is the received warning message, and the server plays the recorded data it holds to scrutinize the content. Call records are stored in the server and can be played and checked as needed. The output is the confirmed call content and its evaluation result.

[0279] (Application Example 1)

[0280] Next, Application Example 1 will be described. In the following description, the data processing device 12 is referred to as the "server", and the smart glasses 214 are referred to as the "terminal".

[0281] In modern society, illegal acts through acoustic communication are increasingly causing serious harm to users. With previous technologies, it has been difficult to detect illegal acts in real time and issue warnings, so there is a need to provide an effective means to prevent fraud.

[0282] The specific processing by the specific processing unit 290 of the data processing device 12 in Application Example 1 is realized by the following means.

[0283] In this invention, the server includes means for converting an acoustic signal into digitized data, means for analyzing the transmitted digitized data and identifying signs of illegal acts using a generation model, and means for displaying a warning notification and transmitting a warning to the recipient when identified. This makes it possible to detect illegal acts during acoustic communication in real time and immediately issue a warning to the user.

[0284] The "user" is an individual or group that uses the system and conducts acoustic communication.

[0285] "Communication" is a process of transmitting and receiving acoustic signals between users.

[0286] The "acoustic signal" is a signal that refers to the speech or voice of the user.

[0287] "Device means" refers to hardware or software components used to receive the user's acoustic signals.

[0288] "Digitized data" refers to data obtained after converting an acoustic signal into a digital format.

[0289] A "wide-area communication network" is a network infrastructure used to transmit digitized data.

[0290] "Analysis means" refers to the processes and tools used to analyze the transmitted digitized data.

[0291] A "generative model" is a machine learning algorithm used to generate character data from received acoustic signals.

[0292] "Signs of fraudulent activity" refers to the identification of keywords or patterns that may be associated with fraud or crime.

[0293] A "warning mechanism" refers to a structure or method for issuing a warning when signs of fraudulent activity are detected.

[0294] "Visualization means" refers to devices or software that visually display warning notifications to the user.

[0295] "Recording means" refers to equipment or technology for recording acoustic signals and retaining them for a specified period of time.

[0296] A "recipient" is a user or their associate who has been pre-registered to receive warning notifications.

[0297] The system for realizing this invention provides a function to detect fraudulent activity during acoustic communication in real time.

[0298] First, when a user performs acoustic communication, the terminal receives an acoustic signal. The received acoustic signal is converted into digital data using the terminal's speech recognition engine. It is expected that existing speech recognition services such as the Google Speech-to-Text API will be used for this conversion. The converted digital data is then sent to a server via a wide-area communication network (e.g., the internet). In this process, the secure protocol HTTPS is used to ensure that the data is transferred safely.

[0299] On the server, a generative model operates to analyze the received digitized data. The generative model uses machine learning algorithms such as OpenAI GPT-3 to generate character data from acoustic signals. Next, this character data is analyzed to identify signs of fraudulent activity. The analysis process monitors for predefined fraud-related keywords, and if these keywords are present, they are identified as signs of fraudulent activity.

[0300] When identification occurs, the server sends a warning notification to the user and pre-registered recipients via a warning mechanism. The warning notification is immediately displayed on the terminal's visualization mechanism, allowing the user to check the communication content and take necessary measures. In addition, all acoustic signals are recorded and stored by a recording mechanism, and can be played back and reviewed later as needed.

[0301] For example, if a user receives suspicious instructions during an audio communication, such as "Please make a transfer immediately," the system will display a warning that it "may be fraudulent." This allows the user to immediately realize it is a scam and prevent becoming a victim.

[0302] An example of a prompt to a generative AI model might be something like, "Analyze the phrases used in this call and assess the likelihood of fraud."

[0303] As described above, this system provides an effective means to reduce the risk of illegal activities through acoustic communication and ensure the safety of users.

[0304] The flow of the specific process in Application Example 1 will be described using FIG. 12.

[0305] Step 1:

[0306] The terminal receives an acoustic signal from the user. The input at this time is the user's voice, and the output is the raw acoustic signal. The terminal captures the acoustic signal using a microphone.

[0307] Step 2:

[0308] The terminal converts the received acoustic signal into digitized data. The input is the acoustic signal obtained in Step 1, and the output is digitized data. This conversion is performed using a speech recognition engine (e.g., Google Speech-to-Text API). The terminal encodes the features of the voice into a digital format.

[0309] Step 3:

[0310] The terminal transmits the converted digitized data to the server via a wide-area communication network. The input is digitized data, and the output is data transfer via a secure protocol (e.g., HTTPS). The terminal uses a network interface to transmit the data.

[0311] Step 4:

[0312] The server activates a generation model for analyzing the received digitized data. The input is digitized data, and the output is the analysis result. The server uses a generative AI model (e.g., OpenAI GPT-3) to generate character data from the acoustic signal.

[0313] Step 5:

[0314] The server identifies signs of fraudulent activity based on the generated text data. The input is text data, and the output is the identification result of the potential for fraudulent activity. The server uses an algorithm that checks for fraud-related keywords to identify specific phrases and patterns.

[0315] Step 6:

[0316] The server generates a warning notification and sends it to the user and registered recipients if it detects signs of fraudulent activity. The input is the identification result, and the output is the warning notification. The server generates the notification message and delivers it to the specified recipients.

[0317] Step 7:

[0318] The terminal displays received warning notifications. The input is the warning notification, and the output is the display of the warning message on the user interface. The terminal communicates the warning to the user using either a display or audio output.

[0319] Step 8:

[0320] The server records all audio signals and retains them for a specified period. The input is the audio signal, and the output is the stored audio recording. The server uses a database or storage system to save the data in a format that can be accessed later.

[0321] Furthermore, an emotion engine that estimates the user's emotions may be incorporated. That is, the identification processing unit 290 may use the emotion identification model 59 to estimate the user's emotions and perform identification processing using the user's emotions.

[0322] This invention provides a voice communication fraud prevention system that combines an emotion engine. This system focuses on the terminal, server, and user, and achieves more precise fraud detection and response by comprehensively analyzing the voice communications made by the user.

[0323] When a user initiates communication, the terminal converts the voice signal into digital data and sends it to the server in real time. The server receives this digital data and performs analysis. For the analysis, a generative model is used to convert the voice signal into text data and detect keywords and phrases that may be related to fraud.

[0324] Furthermore, this system incorporates an emotion engine that estimates emotions from the user's voice. The emotion engine particularly identifies emotions such as anxiety, tension, and agitation, and uses this to improve the accuracy of the analysis. For example, when a user utters words expressing anxiety such as "I'm in trouble" or "What should I do?", the system analyzes the tone, speed, and volume of their voice to understand their emotional state.

[0325] The server generates alerts based on signs of fraudulent activity and emotional shifts. If the emotion engine detects strong signs of anxiety or tension, it adjusts the level and urgency of the warning notification and promptly alerts the user and registered recipients. This warning may include specific details such as, "This call may be fraudulent and anxiety has been detected. Please investigate immediately."

[0326] The user or their family can review the content of the call and play back the recording for further analysis based on the warning notification sent. They can also take appropriate measures, such as reporting to the police, if necessary.

[0327] The system of this invention makes it possible to prevent damage from telephone-based fraud by providing dual protection through real-time analysis and emotion recognition. For example, if a user suddenly becomes anxious and starts worrying about their money, the system will highly assess the possibility of fraud, immediately generate a warning, and notify family members to encourage appropriate action.

[0328] The following describes the processing flow.

[0329] Step 1:

[0330] When the terminal initiates a call with a user, it converts the audio signal into a digital format and transmits it to the server in real time using a secure protocol.

[0331] Step 2:

[0332] The server analyzes the received digital data using a generative model and converts the audio signal into text data. It then processes the text data to detect fraud-related keywords and phrases.

[0333] Step 3:

[0334] An emotion engine installed on the server estimates the user's emotions from their voice. This engine analyzes the tone, speed, and volume of the voice to identify the emotions the user is expressing (such as anxiety or tension).

[0335] Step 4:

[0336] The server integrates the analysis results and sentiment estimation results to score the likelihood of fraud. If signs of fraudulent activity or emotional instability are detected, a warning alert is prepared.

[0337] Step 5:

[0338] The server sends a warning notification to the user's device and the devices of registered recipients. The notification includes warnings about potential fraud and emotional instability.

[0339] Step 6:

[0340] The system will review the warning notification received by the user or a family member, and play back the recorded data stored on the server to examine the call content. If fraud is suspected, the system will promptly request appropriate action.

[0341] (Example 2)

[0342] Next, we will describe Example 2. In the following description, the data processing device 12 will be referred to as the "server" and the smart glasses 214 will be referred to as the "terminal".

[0343] Fraudulent activity in voice communications presents a challenge because the content of these communications is diverse, making it difficult to prevent simply by detecting keywords. Furthermore, since a user's psychological state and emotions may be related to fraudulent activity, there is a need for highly accurate fraud detection that also takes emotional changes into account.

[0344] The identification process performed by the identification processing unit 290 of the data processing device 12 in Example 2 is realized by the following means.

[0345] In this invention, the server includes means for generating electrical data, means for estimating emotional states using emotion analysis means, and means for issuing warnings based on signs of misconduct and changes in emotion. This makes it possible to detect misconduct in voice communications with high accuracy while taking emotional states into consideration and to respond quickly.

[0346] "Acoustic signal" refers to the physical sound vibrations generated when the user's voice is captured.

[0347] "Device means" refers to a device for receiving acoustic signals and converting them into digital data.

[0348] "Electrical data" refers to audio information in digital format that has been converted by a device or means.

[0349] A "data path" refers to a means of communication for transmitting electrical data.

[0350] "Analysis" refers to the process of processing received electrical data and evaluating signs of fraudulent activity.

[0351] "Analysis means" refers to a method or device for analyzing electrical data and identifying fraudulent activity.

[0352] "Recording means" refers to a function that saves audio signals in digital format and allows them to be played back later.

[0353] "Emotion analysis means" refers to a method or device for estimating emotions from a user's voice and recognizing changes in those emotions.

[0354] A "warning" refers to a message that alerts users or other relevant parties to suspected fraudulent activity.

[0355] This invention provides a system for preventing fraudulent activities in voice communications, and implements this using a user, a terminal, and a server.

[0356] The user initiates voice communication and inputs their voice through the terminal. The terminal converts the input acoustic signal into electrical data. This process is achieved using the audio processor built into the terminal.

[0357] The terminal transmits the converted electrical data to the server. A secure and efficient data path is ensured for communication, and generally, communication methods using encryption technology are employed.

[0358] The server analyzes the received electrical data. A generative AI model is used for the analysis, generating text data from the audio. Based on this text data, natural language processing (NLP) techniques are used to detect keywords related to fraud. The server also utilizes sentiment analysis to estimate the user's emotional state from the audio data. Sentiment analysis determines the psychological state from the tone, speed, and volume of the voice, identifying emotions such as anxiety and tension.

[0359] Based on the analysis above, the server generates a warning when it detects signs of fraudulent activity and changes in mood. The warning is sent to the user and registered recipients, prompting them to take necessary action. For example, the warning may include a specific message such as, "This call may be fraudulent and anxiety has been detected. Please check immediately."

[0360] This system's real-time fraud detection and rapid response capabilities allow users and their families to prevent fraud and improve the security of voice communications. A specific use case would be if a user, speaking in a tense tone during a call, says, "I'm worried about my finances." The system would immediately detect the risk and send a warning notification to the family.

[0361] An example of a prompt message would be, "Analyze the voice that expresses the user's anxiety and assess the likelihood of fraud." This prompt allows the generative AI model to appropriately evaluate the user's psychological state and the content of their statements, enabling sophisticated fraud detection.

[0362] The flow of the specific processing in Example 2 will be explained using Figure 13.

[0363] Step 1:

[0364] The user inputs an acoustic signal to the terminal to initiate voice communication. This acoustic signal is then generated.

[0365] Step 2:

[0366] The terminal converts the received acoustic signal into electrical data using a digital audio processor. This process transforms the signal into digital data. The output is the converted electrical data.

[0367] Step 3:

[0368] The terminal transmits the converted electrical data to the server via the data path. This allows the server to obtain the data necessary for analysis.

[0369] Step 4:

[0370] The server converts received electrical data into text data using a generation AI model. The input is electrical data, and the output is the generated text data. This conversion process is performed via a speech recognition engine.

[0371] Step 5:

[0372] The server uses natural language processing techniques to detect fraud-related keywords from text data. The input is text data, and the output is the identified keywords or their absence.

[0373] Step 6:

[0374] The server estimates emotions using emotion analysis tools. The input is electrical data, and the output is the estimated emotional state. Specifically, it analyzes the tone and speed of speech to identify feelings of anxiety and tension.

[0375] Step 7:

[0376] The server generates an alert based on an assessment of signs of misconduct and changes in sentiment. The alert is customized using prompts as needed. The content of the alert message is generated at this stage.

[0377] Step 8:

[0378] The server sends the generated warnings to the user and registered recipients. The output is the specific warning message that the user and recipients receive. This warning gives the user an opportunity to take corrective action.

[0379] (Application Example 2)

[0380] Next, we will explain application example 2. In the following explanation, the data processing device 12 will be referred to as the "server," and the smart glasses 214 will be referred to as the "terminal."

[0381] In recent years, with the advancement of communication technology, voice-based fraud has been increasing, posing a serious problem, especially for vulnerable individuals such as the elderly. To effectively protect users from such fraud, there is a need for a system that can quickly and accurately detect signs of fraudulent activity and also provide warnings in response to changes in the user's emotional state.

[0382] The specific processing performed by the specific processing unit 290 of the data processing device 12 in Application Example 2 is realized by the following means.

[0383] In this invention, the server includes a device means for converting voice signals into digital data and transmitting them over a network, an analysis means for analyzing the transmitted digital data to identify signs of fraudulent activity, and a device means for estimating the user's emotional state using an emotion analysis engine and detecting anxiety and tension. This makes it possible to detect signs that should raise suspicion of fraudulent activity during the user's voice communication, understand the user's emotional changes at that time, and quickly display and transmit appropriate warnings to the user's visual device and pre-registered receiving equipment.

[0384] "Device means" refers to equipment that has the function of receiving communications as audio signals, converting them into digital data, and transmitting them over a network.

[0385] An "analysis device" is a device that has the function of identifying signs of fraudulent activity by analyzing transmitted digital data.

[0386] A "notification system" is a system that sends a warning to users or registered receiving equipment when signs of fraudulent activity are identified.

[0387] A "memory device" is a device that has the function of recording audio signals and retaining them for a specified period of time.

[0388] An "emotion analysis engine" is a system that estimates a user's emotional state and detects anxiety or tension by analyzing the tone, speed, and volume of their voice.

[0389] A "visual device" is a device used to display visual information to a user.

[0390] This invention provides a system for preventing fraudulent activities in voice communications, which has a configuration that comprehensively analyzes the content of user communications. The main components of the system are digitization of voice signals, data analysis, fraud detection, emotion recognition, and warning issuance.

[0391] The server receives the audio signal sent from the terminal when the user initiates communication. The received audio signal is first converted into digital data. This process uses a real-time audio analysis engine (e.g., Nuance Dragon) to handle the audio data.

[0392] The server then uses a generative AI model to generate text data from the audio data. This generative AI model (e.g., OpenAI GPT) has the ability to analyze the text data and identify keywords and phrases related to fraudulent activity. Simultaneously, the server uses an emotion analysis engine (e.g., Affectiva) to estimate the user's emotional state from their voice and detect anxiety and tension.

[0393] The user's device is equipped with a visual device (e.g., smart glasses) to display warnings visually. When the system detects signs of misconduct and changes in the user's emotional state, it displays a warning on the user's visual device and immediately sends a notification to registered receiving equipment.

[0394] As a concrete example, consider a scenario where a user is on a regular phone call and the other party approaches them with a "good offer about investing in a new financial product." In this case, the server transcribes the other party's words in real time, detects the possibility of fraud, and also detects feelings of tension or anxiety from the user's voice. Based on this, the system displays a warning such as "This may be a scam, please be careful" on the user's visual device and sends a notification to the devices of registered recipients.

[0395] Examples of prompts for the generative AI model include, "Identify potentially fraudulent phrases that can be detected from this audio data," and "Analyze the emotional shifts following the user's statements and report any signs of anxiety or tension."

[0396] In this way, the server can ensure user safety and prevent fraud.

[0397] The flow of a specific process in Application Example 2 will be explained using Figure 14.

[0398] Step 1:

[0399] When a user initiates communication, the terminal captures the audio signal using its built-in microphone. This signal is then converted into digital data in real time. This digital data is then input to the server and sent to the audio analysis engine.

[0400] Step 2:

[0401] The server analyzes the received digital data and uses a generative AI model to convert the audio data into text data. The generative AI model extracts features from the audio signal and uses prompt sentences to generate text. The output of this step is the transcribed audio data.

[0402] Step 3:

[0403] The server analyzes the generated text data to identify keywords and phrases related to fraud. This process compares specific words indicating fraudulent activity against a database to generate identification results. The output is a flag indicating the presence or absence of fraudulent activity.

[0404] Step 4:

[0405] Simultaneously, the server processes the user's voice signal using an emotion analysis engine to estimate their emotional state. Specifically, it analyzes the tone, speed, and volume of the voice to estimate the level of anxiety and tension. The output is an indicator of the user's emotional state.

[0406] Step 5:

[0407] The server issues warnings based on flags indicating fraudulent activity and emotional state indicators. It instructs visual devices to display warning messages notifying them of potential fraud and emotional changes. It also sends similar warning notifications to registered receiving devices.

[0408] Step 6:

[0409] Users receive a warning and review feedback regarding changes in call content and emotional state. This allows them to determine whether actual misconduct occurred and take appropriate action.

[0410] This allows the entire system to work together to protect users from fraudulent voice communication.

[0411] The specific processing unit 290 transmits the result of the specific processing to the smart glasses 214. In the smart glasses 214, the control unit 46A causes the speaker 240 to output the result of the specific processing. The microphone 238 acquires audio indicating user input for the result of the specific processing. The control unit 46A transmits the audio data indicating user input acquired by the microphone 238 to the data processing unit 12. In the data processing unit 12, the specific processing unit 290 acquires the audio data.

[0412] Data generation model 58 is a type of so-called generative AI (Artificial Intelligence). One example of data generation model 58 is ChatGPT (Internet search<URL: https: / / openai.com / blog / chatgpt> ), Gemini (Internet search) <url: https: gemini.google.com ?hl="ja">Examples of generative AI include the following. The data generation model 58 is obtained by performing deep learning on a neural network. The data generation model 58 is input with prompts containing instructions, and with inference data such as audio data representing speech, text data representing text, and image data representing images. The data generation model 58 infers from the input inference data according to the instructions indicated by the prompts, and outputs the inference results in data formats such as audio data and text data. Here, inference refers to, for example, analysis, classification, prediction, and / or summarization.

[0413] In the above embodiment, an example was given in which specific processing is performed by the data processing device 12, but the technology of this disclosure is not limited thereto, and the specific processing may also be performed by the smart glasses 214.

[0414] [Third Embodiment]

[0415] Figure 5 shows an example of the configuration of the data processing system 310 according to the third embodiment.

[0416] As shown in Figure 5, the data processing system 310 includes a data processing device 12 and a headset terminal 314. An example of the data processing device 12 is a server.

[0417] The data processing device 12 comprises a computer 22, a database 24, and a communication interface 26. The computer 22 is an example of a "computer" related to the technology of this disclosure. The computer 22 comprises a processor 28, RAM 30, and storage 32. The processor 28, RAM 30, and storage 32 are connected to a bus 34. The database 24 and the communication interface 26 are also connected to the bus 34. The communication interface 26 is connected to a network 54. An example of the network 54 is a WAN (Wide Area Network) and / or a LAN (Local Area Network).

[0418] The headset terminal 314 includes a computer 36, a microphone 238, a speaker 240, a camera 42, a communication interface 44, and a display 343. The computer 36 includes a processor 46, RAM 48, and storage 50. The processor 46, RAM 48, and storage 50 are connected to a bus 52. The microphone 238, speaker 240, camera 42, and display 343 are also connected to the bus 52.

[0419] The microphone 238 receives voice signals from the user 20 and receives instructions from the user 20. The microphone 238 captures the voice signals from the user 20, converts the captured voice into audio data, and outputs it to the processor 46. The speaker 240 outputs audio according to the instructions from the processor 46.

[0420] Camera 42 is a small digital camera equipped with an optical system including a lens, aperture, and shutter, and an image sensor such as a CMOS (Complementary Metal-Oxide-Semiconductor) image sensor or a CCD (Charge Coupled Device) image sensor, and captures images of the area around the user 20 (for example, an imaging range defined by a field of view equivalent to the width of a typical healthy person's field of vision).

[0421] Communication interface 44 is connected to network 54. Communication interfaces 44 and 26 are responsible for the exchange of various information between processor 46 and processor 28 via network 54. The exchange of various information between processor 46 and processor 28 using communication interfaces 44 and 26 is performed in a secure manner.

[0422] Figure 6 shows an example of the main functions of the data processing device 12 and the headset terminal 314. As shown in Figure 6, the data processing device 12 performs specific processing using the processor 28. The storage 32 stores the specific processing program 56.

[0423] The specific processing program 56 is an example of a "program" relating to the technology of this disclosure. The processor 28 reads the specific processing program 56 from the storage 32 and executes the read specific processing program 56 on the RAM 30. The specific processing is realized by the processor 28 operating as a specific processing unit 290 in accordance with the specific processing program 56 executed on the RAM 30.

[0424] The storage 32 stores the data generation model 58 and the emotion identification model 59. The data generation model 58 and the emotion identification model 59 are used by the identification processing unit 290.

[0425] In the headset terminal 314, the processor 46 performs the reception output processing. The storage 50 stores the reception output program 60. The processor 46 reads the reception output program 60 from the storage 50 and executes the read reception output program 60 on the RAM 48. The reception output processing is realized by the processor 46 operating as a control unit 46A according to the reception output program 60 executed on the RAM 48.

[0426] Next, the specific processing performed by the specific processing unit 290 of the data processing device 12 will be described. In the following description, the data processing device 12 will be referred to as the "server" and the headset terminal 314 will be referred to as the "terminal".

[0427] This invention provides a system for preventing fraudulent activities via voice communication from a terminal used by a user. The main components of this system are based on three entities: a terminal, a server, and a user.

[0428] When a user initiates a phone call, the terminal converts the voice signal into digital data in real time. The converted digital data is then sent to the server using a secure protocol.

[0429] On the server, a generative model runs to analyze the received digital data. This generative model has the ability to convert voice signals into text and identify keywords and phrases that may be related to fraud. For example, if phrases such as "ATM," "bank transfer," and "urgent" are detected, the likelihood that the call is associated with fraudulent activity is assessed.

[0430] If potential fraudulent activity is detected, the server will generate an alert through its warning system and send a warning notification to the user's device and the devices of their pre-registered family members. This notification will display a message such as, "This call may contain fraudulent activity. Please review the content."

[0431] The user or their family can immediately review the content of a call based on the received warning notification. Calls are always recorded, and the recorded audio files are stored on the server for a certain period, allowing them to be played back and reviewed as needed.

[0432] Through the process described above, this system aims to prevent telephone-based fraud by detecting fraudulent activity in real time and enabling a rapid response.

[0433] The following describes the processing flow.

[0434] Step 1:

[0435] The terminal receives the user's call and converts the audio signal into digital data. The converted digital data is then sent to the server via a secure protocol.

[0436] Step 2:

[0437] The server analyzes the received digital data and uses a generative model to convert the audio signal into text data. From this text data, keywords and phrases related to fraudulent activities are detected.

[0438] Step 3:

[0439] The server scores the likelihood of fraud and prepares to generate a warning alert if the likelihood of fraud exceeds a threshold.

[0440] Step 4:

[0441] The server uses a warning mechanism to send a warning notification to the user's device and the devices of registered family members, alerting the user.

[0442] Step 5:

[0443] The system utilizes a function that allows users to review warning notifications they receive and then play back recorded call data stored on the server to examine its contents.

[0444] Step 6:

[0445] If a user or their family member suspects fraud, they should take action, such as reporting it to the police, to resolve the issue.

[0446] (Example 1)

[0447] Next, we will describe Example 1. In the following description, the data processing device 12 will be referred to as the "server," and the headset-type terminal 314 will be referred to as the "terminal."

[0448] Addressing fraud and malicious communications via telephone is crucial, but currently, there are limited means to detect these activities in real time and issue rapid warnings. This means many users could become victims of fraud and suffer serious damage. To solve this problem, an effective system is needed that can quickly and accurately detect fraudulent activity through voice communications and immediately notify users.

[0449] The identification process performed by the identification processing unit 290 of the data processing device 12 in Example 1 is realized by the following means.

[0450] In this invention, the server includes means for receiving voice signals and converting them into digital information, processing means for analyzing the transmitted digital information to identify signs of fraudulent activity, and notification means for issuing warnings to appropriate recipients. This makes it possible to detect fraudulent activity via voice communication in real time and immediately issue warnings to users and related parties.

[0451] "Electronic device" refers to a device that has the function of receiving communications from a user as audio signals and converting them into digital information.

[0452] A "data network" refers to a communication network used to securely transmit digital information from electronic devices to servers.

[0453] "Processing means" refers to equipment equipped with the function of analyzing transmitted digital information and performing a process to identify signs of specific fraudulent activity.

[0454] "Notification means" refers to a system that has the function of issuing warnings to users and related parties based on identified signs of fraudulent activity.

[0455] "Storage means" refers to a device or system that records audio signals and retains that recorded data for a predetermined period of time.

[0456] This invention is a system that analyzes the content of a voice communication in real time when the user initiates a voice communication, thereby preventing fraudulent activity. The system mainly consists of three elements: a terminal, a server, and a user.

[0457] When a user initiates voice communication, the terminal captures the audio signal using its built-in microphone and converts the analog signal into digital data using audio processing software such as "FFmpeg". This digital data is then transmitted to the server via the data network using the SSL / TLS protocol.

[0458] The server utilizes a natural language processing API to process the received digital information. This generative AI model converts the audio signal into text data. Subsequently, a script developed in Python is used to identify specific keywords and phrases (e.g., "ATM," "bank transfer," "urgent") from the generated text data. If potential fraud is detected, the server uses services such as Twilio to send warning notifications to the user and pre-registered recipients.

[0459] As an example, suppose a user receives a suspicious telemarketing call instructing them to "transfer the money immediately." In this case, the system detects the relevant phrase and promptly sends a warning notification to the user. This allows the user to recognize the danger of the call and take swift action.

[0460] Examples of prompt statements for a generative AI model are as follows:

[0461] "Is this call potentially a scam? Detected keywords: [ATM, bank transfer, urgent]"

[0462] This system allows for real-time monitoring of fraudulent activity via voice communication, enabling users to take prompt and appropriate action.

[0463] The flow of the specific processing in Example 1 will be explained using Figure 11.

[0464] Step 1:

[0465] When a user initiates voice communication, the device captures the audio signal using its built-in microphone. The input is an analog audio signal, and the device prepares to convert the acquired audio into digital data. Specifically, it uses a process such as FFmpeg to convert the analog signal into digital information, and the output is digital data in PCM format.

[0466] Step 2:

[0467] The terminal transmits the converted digital data to the server using a security protocol (SSL / TLS). The input is digital data in PCM format; this data is encrypted, packetized, and securely transferred to the server via the data network. The output is encrypted data packets.

[0468] Step 3:

[0469] The server receives digital data transmitted from the terminal and converts it from speech to text. The input is encrypted data packets, which are decrypted and then converted into text data using a natural language processing API. As a result of the data analysis, the output is text data corresponding to the speech.

[0470] Step 4:

[0471] The server uses a generative AI model to detect specific keywords and phrases from the obtained text data. The input is text data, and a script written in Python searches for pre-configured fraud-related keywords (e.g., "ATM," "bank transfer," "urgent"). The output is a list of the keywords found.

[0472] Step 5:

[0473] The server evaluates the potential for malicious activity and generates a warning notification when specific keywords are detected. The input is a list of detected keywords, and warning messages are sent to the user and registered recipients using services such as Twilio. The output is the sending of warning messages and a record of them.

[0474] Step 6:

[0475] The user or recipient receives a warning and reviews the content of the call. The input is the received warning message, and the server plays back the recorded data to examine its contents. The server stores the call records, which can be played back and reviewed as needed. The output is the reviewed call content and its evaluation result.

[0476] (Application Example 1)

[0477] Next, we will explain Application Example 1. In the following explanation, the data processing device 12 will be referred to as the "server," and the headset-type terminal 314 will be referred to as the "terminal."

[0478] In modern society, fraudulent activities conducted via acoustic communication are increasingly causing serious harm to users. Existing technologies have made it difficult to detect and warn of fraudulent activities in real time, therefore, there is a need to provide effective means to prevent fraud before it occurs.

[0479] The specific processing performed by the specific processing unit 290 of the data processing device 12 in Application Example 1 is realized by the following means.

[0480] In this invention, the server includes means for converting acoustic signals into digital data, means for analyzing the transmitted digital data and identifying signs of fraudulent activity using a generative model, and means for displaying a warning notification when fraudulent activity is identified and sending the warning to the recipient. This makes it possible to detect fraudulent activity during acoustic communication in real time and immediately warn the user.

[0481] A "user" is an individual or group that uses the system to conduct acoustic communication.

[0482] "Communication" is the process of sending and receiving acoustic signals between users.

[0483] An "acoustic signal" refers to a signal that represents the user's speech or voice.

[0484] "Device means" refers to hardware or software components used to receive the user's acoustic signals.

[0485] "Digitized data" refers to data obtained after converting an acoustic signal into a digital format.

[0486] A "wide-area communication network" is a network infrastructure used to transmit digitized data.

[0487] "Analysis means" refers to the processes and tools used to analyze the transmitted digitized data.

[0488] A "generative model" is a machine learning algorithm used to generate character data from received acoustic signals.

[0489] "Signs of fraudulent activity" refers to the identification of keywords or patterns that may be associated with fraud or crime.

[0490] A "warning mechanism" refers to a structure or method for issuing a warning when signs of fraudulent activity are detected.

[0491] "Visualization means" refers to devices or software that visually display warning notifications to the user.

[0492] "Recording means" refers to equipment or technology for recording acoustic signals and retaining them for a specified period of time.

[0493] A "recipient" is a user or their associate who has been pre-registered to receive warning notifications.

[0494] The system for realizing this invention provides a function to detect fraudulent activity during acoustic communication in real time.

[0495] First, when a user performs acoustic communication, the terminal receives an acoustic signal. The received acoustic signal is converted into digital data using the terminal's speech recognition engine. It is expected that existing speech recognition services such as the Google Speech-to-Text API will be used for this conversion. The converted digital data is then sent to a server via a wide-area communication network (e.g., the internet). In this process, the secure protocol HTTPS is used to ensure that the data is transferred safely.

[0496] On the server, a generative model operates to analyze the received digitized data. The generative model uses machine learning algorithms such as OpenAI GPT-3 to generate character data from acoustic signals. Next, this character data is analyzed to identify signs of fraudulent activity. The analysis process monitors for predefined fraud-related keywords, and if these keywords are present, they are identified as signs of fraudulent activity.

[0497] When identification occurs, the server sends a warning notification to the user and pre-registered recipients via a warning mechanism. The warning notification is immediately displayed on the terminal's visualization mechanism, allowing the user to check the communication content and take necessary measures. In addition, all acoustic signals are recorded and stored by a recording mechanism, and can be played back and reviewed later as needed.

[0498] For example, if a user receives suspicious instructions during an audio communication, such as "Please make a transfer immediately," the system will display a warning that it "may be fraudulent." This allows the user to immediately realize it is a scam and prevent becoming a victim.

[0499] An example of a prompt to a generative AI model might be something like, "Analyze the phrases used in this call and assess the likelihood of fraud."

[0500] Therefore, this system provides an effective means of reducing the risk of fraudulent activity via acoustic communication and ensuring user safety.

[0501] The flow of a specific process in Application Example 1 will be explained using Figure 12.

[0502] Step 1:

[0503] The terminal receives an acoustic signal from the user. The input is the user's voice, and the output is a raw acoustic signal. The terminal captures the acoustic signal using a microphone.

[0504] Step 2:

[0505] The device converts the received acoustic signal into digital data. The input is the acoustic signal acquired in step 1, and the output is the digital data. This conversion is performed using a speech recognition engine (e.g., Google Speech-to-Text API). The device encodes the audio features into a digital format.

[0506] Step 3:

[0507] The terminal transmits the converted digital data to the server via a wide-area communication network. The input is digital data, and the output is data transfer via a secure protocol (e.g., HTTPS). The terminal transmits the data using a network interface.

[0508] Step 4:

[0509] The server launches a generative model to analyze the received digitized data. The input is digitized data, and the output is the analysis result. The server uses a generative AI model (e.g., OpenAI GPT-3) to generate character data from acoustic signals.

[0510] Step 5:

[0511] The server identifies signs of fraudulent activity based on the generated text data. The input is text data, and the output is the identification result of the potential for fraudulent activity. The server uses an algorithm that checks for fraud-related keywords to identify specific phrases and patterns.

[0512] Step 6:

[0513] The server generates a warning notification and sends it to the user and registered recipients if it detects signs of fraudulent activity. The input is the identification result, and the output is the warning notification. The server generates the notification message and delivers it to the specified recipients.

[0514] Step 7:

[0515] The terminal displays received warning notifications. The input is the warning notification, and the output is the display of the warning message on the user interface. The terminal communicates the warning to the user using either a display or audio output.

[0516] Step 8:

[0517] The server records all audio signals and retains them for a specified period. The input is the audio signal, and the output is the stored audio recording. The server uses a database or storage system to save the data in a format that can be accessed later.

[0518] Furthermore, an emotion engine that estimates the user's emotions may be incorporated. That is, the identification processing unit 290 may use the emotion identification model 59 to estimate the user's emotions and perform identification processing using the user's emotions.

[0519] This invention provides a voice communication fraud prevention system that combines an emotion engine. This system focuses on the terminal, server, and user, and achieves more precise fraud detection and response by comprehensively analyzing the voice communications made by the user.

[0520] When a user initiates communication, the terminal converts the voice signal into digital data and sends it to the server in real time. The server receives this digital data and performs analysis. For the analysis, a generative model is used to convert the voice signal into text data and detect keywords and phrases that may be related to fraud.

[0521] Furthermore, this system incorporates an emotion engine that estimates emotions from the user's voice. The emotion engine particularly identifies emotions such as anxiety, tension, and agitation, and uses this to improve the accuracy of the analysis. For example, when a user utters words expressing anxiety such as "I'm in trouble" or "What should I do?", the system analyzes the tone, speed, and volume of their voice to understand their emotional state.

[0522] The server generates alerts based on signs of fraudulent activity and emotional shifts. If the emotion engine detects strong signs of anxiety or tension, it adjusts the level and urgency of the warning notification and promptly alerts the user and registered recipients. This warning may include specific details such as, "This call may be fraudulent and anxiety has been detected. Please investigate immediately."

[0523] The user or their family can review the content of the call and play back the recording for further analysis based on the warning notification sent. They can also take appropriate measures, such as reporting to the police, if necessary.

[0524] The system of this invention makes it possible to prevent damage from telephone-based fraud by providing dual protection through real-time analysis and emotion recognition. For example, if a user suddenly becomes anxious and starts worrying about their money, the system will highly assess the possibility of fraud, immediately generate a warning, and notify family members to encourage appropriate action.

[0525] The following describes the processing flow.

[0526] Step 1:

[0527] When the terminal initiates a call with a user, it converts the audio signal into a digital format and transmits it to the server in real time using a secure protocol.

[0528] Step 2:

[0529] The server analyzes the received digital data using a generative model and converts the audio signal into text data. It then processes the text data to detect fraud-related keywords and phrases.

[0530] Step 3:

[0531] An emotion engine installed on the server estimates the user's emotions from their voice. This engine analyzes the tone, speed, and volume of the voice to identify the emotions the user is expressing (such as anxiety or tension).

[0532] Step 4:

[0533] The server integrates the analysis results and sentiment estimation results to score the likelihood of fraud. If signs of fraudulent activity or emotional instability are detected, a warning alert is prepared.

[0534] Step 5:

[0535] The server sends a warning notification to the user's device and the devices of registered recipients. The notification includes warnings about potential fraud and emotional instability.

[0536] Step 6:

[0537] The system will review the warning notification received by the user or a family member, and play back the recorded data stored on the server to examine the call content. If fraud is suspected, the system will promptly request appropriate action.

[0538] (Example 2)

[0539] Next, we will describe Example 2. In the following description, the data processing device 12 will be referred to as the "server," and the headset-type terminal 314 will be referred to as the "terminal."

[0540] Fraudulent activity in voice communications presents a challenge because the content of these communications is diverse, making it difficult to prevent simply by detecting keywords. Furthermore, since a user's psychological state and emotions may be related to fraudulent activity, there is a need for highly accurate fraud detection that also takes emotional changes into account.

[0541] The identification process performed by the identification processing unit 290 of the data processing device 12 in Example 2 is realized by the following means.

[0542] In this invention, the server includes means for generating electrical data, means for estimating emotional states using emotion analysis means, and means for issuing warnings based on signs of misconduct and changes in emotion. This makes it possible to detect misconduct in voice communications with high accuracy while taking emotional states into consideration and to respond quickly.

[0543] "Acoustic signal" refers to the physical sound vibrations generated when the user's voice is captured.

[0544] "Device means" refers to a device for receiving acoustic signals and converting them into digital data.

[0545] "Electrical data" refers to audio information in digital format that has been converted by a device or means.

[0546] A "data path" refers to a means of communication for transmitting electrical data.

[0547] "Analysis" refers to the process of processing received electrical data and evaluating signs of fraudulent activity.

[0548] "Analysis means" refers to a method or device for analyzing electrical data and identifying fraudulent activity.

[0549] "Recording means" refers to a function that saves audio signals in digital format and allows them to be played back later.

[0550] "Emotion analysis means" refers to a method or device for estimating emotions from a user's voice and recognizing changes in those emotions.

[0551] A "warning" refers to a message that alerts users or other relevant parties to suspected fraudulent activity.

[0552] This invention provides a system for preventing fraudulent activities in voice communications, and implements this using a user, a terminal, and a server.

[0553] The user initiates voice communication and inputs their voice through the terminal. The terminal converts the input acoustic signal into electrical data. This process is achieved using the audio processor built into the terminal.

[0554] The terminal transmits the converted electrical data to the server. A secure and efficient data path is ensured for communication, and generally, communication methods using encryption technology are employed.

[0555] The server analyzes the received electrical data. A generative AI model is used for the analysis, generating text data from the audio. Based on this text data, natural language processing (NLP) techniques are used to detect keywords related to fraud. The server also utilizes sentiment analysis to estimate the user's emotional state from the audio data. Sentiment analysis determines the psychological state from the tone, speed, and volume of the voice, identifying emotions such as anxiety and tension.

[0556] Based on the analysis above, the server generates a warning when it detects signs of fraudulent activity and changes in mood. The warning is sent to the user and registered recipients, prompting them to take necessary action. For example, the warning may include a specific message such as, "This call may be fraudulent and anxiety has been detected. Please check immediately."

[0557] This system's real-time fraud detection and rapid response capabilities allow users and their families to prevent fraud and improve the security of voice communications. A specific use case would be if a user, speaking in a tense tone during a call, says, "I'm worried about my finances." The system would immediately detect the risk and send a warning notification to the family.

[0558] An example of a prompt message would be, "Analyze the voice that expresses the user's anxiety and assess the likelihood of fraud." This prompt allows the generative AI model to appropriately evaluate the user's psychological state and the content of their statements, enabling sophisticated fraud detection.

[0559] The flow of the specific processing in Example 2 will be explained using Figure 13.

[0560] Step 1:

[0561] The user inputs an acoustic signal to the terminal to initiate voice communication. This acoustic signal is then generated.

[0562] Step 2:

[0563] The terminal converts the received acoustic signal into electrical data using a digital audio processor. This process transforms the signal into digital data. The output is the converted electrical data.

[0564] Step 3:

[0565] The terminal transmits the converted electrical data to the server via the data path. This allows the server to obtain the data necessary for analysis.

[0566] Step 4:

[0567] The server converts received electrical data into text data using a generation AI model. The input is electrical data, and the output is the generated text data. This conversion process is performed via a speech recognition engine.

[0568] Step 5:

[0569] The server uses natural language processing techniques to detect fraud-related keywords from text data. The input is text data, and the output is the identified keywords or their absence.

[0570] Step 6:

[0571] The server estimates emotions using emotion analysis tools. The input is electrical data, and the output is the estimated emotional state. Specifically, it analyzes the tone and speed of speech to identify feelings of anxiety and tension.

[0572] Step 7:

[0573] The server generates an alert based on an assessment of signs of misconduct and changes in sentiment. The alert is customized using prompts as needed. The content of the alert message is generated at this stage.

[0574] Step 8:

[0575] The server sends the generated warnings to the user and registered recipients. The output is the specific warning message that the user and recipients receive. This warning gives the user an opportunity to take corrective action.

[0576] (Application Example 2)

[0577] Next, we will explain application example 2. In the following explanation, the data processing device 12 will be referred to as the "server," and the headset-type terminal 314 will be referred to as the "terminal."

[0578] In recent years, with the advancement of communication technology, voice-based fraud has been increasing, posing a serious problem, especially for vulnerable individuals such as the elderly. To effectively protect users from such fraud, there is a need for a system that can quickly and accurately detect signs of fraudulent activity and also provide warnings in response to changes in the user's emotional state.

[0579] The specific processing performed by the specific processing unit 290 of the data processing device 12 in Application Example 2 is realized by the following means.

[0580] In this invention, the server includes a device means for converting voice signals into digital data and transmitting them over a network, an analysis means for analyzing the transmitted digital data to identify signs of fraudulent activity, and a device means for estimating the user's emotional state using an emotion analysis engine and detecting anxiety and tension. This makes it possible to detect signs that should raise suspicion of fraudulent activity during the user's voice communication, understand the user's emotional changes at that time, and quickly display and transmit appropriate warnings to the user's visual device and pre-registered receiving equipment.

[0581] "Device means" refers to equipment that has the function of receiving communications as audio signals, converting them into digital data, and transmitting them over a network.

[0582] An "analysis device" is a device that has the function of identifying signs of fraudulent activity by analyzing transmitted digital data.

[0583] A "notification system" is a system that sends a warning to users or registered receiving equipment when signs of fraudulent activity are identified.

[0584] A "memory device" is a device that has the function of recording audio signals and retaining them for a specified period of time.

[0585] An "emotion analysis engine" is a system that estimates a user's emotional state and detects anxiety or tension by analyzing the tone, speed, and volume of their voice.

[0586] A "visual device" is a device used to display visual information to a user.

[0587] This invention provides a system for preventing fraudulent activities in voice communications, which has a configuration that comprehensively analyzes the content of user communications. The main components of the system are digitization of voice signals, data analysis, fraud detection, emotion recognition, and warning issuance.

[0588] The server receives the audio signal sent from the terminal when the user initiates communication. The received audio signal is first converted into digital data. This process uses a real-time audio analysis engine (e.g., Nuance Dragon) to handle the audio data.

[0589] The server then uses a generative AI model to generate text data from the audio data. This generative AI model (e.g., OpenAI GPT) has the ability to analyze the text data and identify keywords and phrases related to fraudulent activity. Simultaneously, the server uses an emotion analysis engine (e.g., Affectiva) to estimate the user's emotional state from their voice and detect anxiety and tension.

[0590] The user's device is equipped with a visual device (e.g., smart glasses) to display warnings visually. When the system detects signs of misconduct and changes in the user's emotional state, it displays a warning on the user's visual device and immediately sends a notification to registered receiving equipment.

[0591] As a concrete example, consider a scenario where a user is on a regular phone call and the other party approaches them with a "good offer about investing in a new financial product." In this case, the server transcribes the other party's words in real time, detects the possibility of fraud, and also detects feelings of tension or anxiety from the user's voice. Based on this, the system displays a warning such as "This may be a scam, please be careful" on the user's visual device and sends a notification to the devices of registered recipients.

[0592] Examples of prompts for the generative AI model include, "Identify potentially fraudulent phrases that can be detected from this audio data," and "Analyze the emotional shifts following the user's statements and report any signs of anxiety or tension."

[0593] In this way, the server can ensure user safety and prevent fraud.

[0594] The flow of a specific process in Application Example 2 will be explained using Figure 14.

[0595] Step 1:

[0596] When a user initiates communication, the terminal captures the audio signal using its built-in microphone. This signal is then converted into digital data in real time. This digital data is then input to the server and sent to the audio analysis engine.

[0597] Step 2:

[0598] The server analyzes the received digital data and uses a generative AI model to convert the audio data into text data. The generative AI model extracts features from the audio signal and uses prompt sentences to generate text. The output of this step is the transcribed audio data.

[0599] Step 3:

[0600] The server analyzes the generated text data to identify keywords and phrases related to fraud. This process compares specific words indicating fraudulent activity against a database to generate identification results. The output is a flag indicating the presence or absence of fraudulent activity.

[0601] Step 4:

[0602] Simultaneously, the server processes the user's voice signal using an emotion analysis engine to estimate their emotional state. Specifically, it analyzes the tone, speed, and volume of the voice to estimate the level of anxiety and tension. The output is an indicator of the user's emotional state.

[0603] Step 5:

[0604] The server issues warnings based on flags indicating fraudulent activity and emotional state indicators. It instructs visual devices to display warning messages notifying them of potential fraud and emotional changes. It also sends similar warning notifications to registered receiving devices.

[0605] Step 6:

[0606] Users receive a warning and review feedback regarding changes in call content and emotional state. This allows them to determine whether actual misconduct occurred and take appropriate action.

[0607] This allows the entire system to work together to protect users from fraudulent voice communication.

[0608] The specific processing unit 290 transmits the result of the specific processing to the headset terminal 314. In the headset terminal 314, the control unit 46A causes the speaker 240 and display 343 to output the result of the specific processing. The microphone 238 acquires audio indicating user input for the result of the specific processing. The control unit 46A transmits the audio data indicating user input acquired by the microphone 238 to the data processing unit 12. In the data processing unit 12, the specific processing unit 290 acquires the audio data.

[0609] Data generation model 58 is a type of so-called generative AI (Artificial Intelligence). One example of data generation model 58 is ChatGPT (Internet search<URL: https: / / openai.com / blog / chatgpt> ), Gemini (Internet search) <url: https: gemini.google.com ?hl="ja">Examples of generative AI include the following. The data generation model 58 is obtained by performing deep learning on a neural network. The data generation model 58 is input with prompts containing instructions, and with inference data such as audio data representing speech, text data representing text, and image data representing images. The data generation model 58 infers from the input inference data according to the instructions indicated by the prompts, and outputs the inference results in data formats such as audio data and text data. Here, inference refers to, for example, analysis, classification, prediction, and / or summarization.

[0610] In the above embodiment, an example was given in which specific processing is performed by the data processing device 12, but the technology of this disclosure is not limited thereto, and specific processing may also be performed by the headset terminal 314.

[0611] [Fourth Embodiment]

[0612] Figure 7 shows an example of the configuration of the data processing system 410 according to the fourth embodiment.

[0613] As shown in Figure 7, the data processing system 410 includes a data processing device 12 and a robot 414. An example of the data processing device 12 is a server.

[0614] The data processing device 12 comprises a computer 22, a database 24, and a communication interface 26. The computer 22 is an example of a "computer" related to the technology of this disclosure. The computer 22 comprises a processor 28, RAM 30, and storage 32. The processor 28, RAM 30, and storage 32 are connected to a bus 34. The database 24 and the communication interface 26 are also connected to the bus 34. The communication interface 26 is connected to a network 54. An example of the network 54 is a WAN (Wide Area Network) and / or a LAN (Local Area Network).

[0615] The robot 414 includes a computer 36, a microphone 238, a speaker 240, a camera 42, a communication interface 44, and a controlled object 443. The computer 36 includes a processor 46, RAM 48, and storage 50. The processor 46, RAM 48, and storage 50 are connected to a bus 52. The microphone 238, speaker 240, camera 42, and controlled object 443 are also connected to the bus 52.

[0616] The microphone 238 receives voice signals from the user 20 and receives instructions from the user 20. The microphone 238 captures the voice signals from the user 20, converts the captured voice into audio data, and outputs it to the processor 46. The speaker 240 outputs audio according to the instructions from the processor 46.

[0617] Camera 42 is a small digital camera equipped with an optical system including a lens, aperture, and shutter, and an image sensor such as a CMOS (Complementary Metal-Oxide-Semiconductor) image sensor or a CCD (Charge Coupled Device) image sensor, and captures images of the area around the user 20 (for example, an imaging range defined by a field of view equivalent to the width of a typical healthy person's field of vision).

[0618] Communication interface 44 is connected to network 54. Communication interfaces 44 and 26 are responsible for the exchange of various information between processor 46 and processor 28 via network 54. The exchange of various information between processor 46 and processor 28 using communication interfaces 44 and 26 is performed in a secure manner.

[0619] The controlled object 443 includes a display device, LEDs in the eyes, and motors that drive the arms, hands, and feet. The posture and gestures of the robot 414 are controlled by controlling the motors of the arms, hands, and feet. Some of the robot 414's emotions can be expressed by controlling these motors. Furthermore, the robot 414's facial expressions can also be expressed by controlling the illumination state of the LEDs in its eyes.

[0620] Figure 8 shows an example of the main functions of the data processing device 12 and the robot 414. As shown in Figure 8, the data processing device 12 performs specific processing using the processor 28. The storage 32 stores the specific processing program 56.

[0621] The specific processing program 56 is an example of a "program" relating to the technology of this disclosure. The processor 28 reads the specific processing program 56 from the storage 32 and executes the read specific processing program 56 on the RAM 30. The specific processing is realized by the processor 28 operating as a specific processing unit 290 in accordance with the specific processing program 56 executed on the RAM 30.

[0622] The storage 32 stores the data generation model 58 and the emotion identification model 59. The data generation model 58 and the emotion identification model 59 are used by the identification processing unit 290.

[0623] In robot 414, the processor 46 performs the reception output processing. The storage 50 stores the reception output program 60. The processor 46 reads the reception output program 60 from the storage 50 and executes the read reception output program 60 on the RAM 48. The reception output processing is realized by the processor 46 operating as a control unit 46A according to the reception output program 60 executed on the RAM 48.

[0624] Next, the specific processing performed by the specific processing unit 290 of the data processing device 12 will be described. In the following description, the data processing device 12 will be referred to as the "server" and the robot 414 as the "terminal".

[0625] This invention provides a system for preventing fraudulent activities via voice communication from a terminal used by a user. The main components of this system are based on three entities: a terminal, a server, and a user.

[0626] When a user initiates a phone call, the terminal converts the voice signal into digital data in real time. The converted digital data is then sent to the server using a secure protocol.

[0627] On the server, a generative model runs to analyze the received digital data. This generative model has the ability to convert voice signals into text and identify keywords and phrases that may be related to fraud. For example, if phrases such as "ATM," "bank transfer," and "urgent" are detected, the likelihood that the call is associated with fraudulent activity is assessed.

[0628] If potential fraudulent activity is detected, the server will generate an alert through its warning system and send a warning notification to the user's device and the devices of their pre-registered family members. This notification will display a message such as, "This call may contain fraudulent activity. Please review the content."

[0629] The user or their family can immediately review the content of a call based on the received warning notification. Calls are always recorded, and the recorded audio files are stored on the server for a certain period, allowing them to be played back and reviewed as needed.

[0630] Through the process described above, this system aims to prevent telephone-based fraud by detecting fraudulent activity in real time and enabling a rapid response.

[0631] The following describes the processing flow.

[0632] Step 1:

[0633] The terminal receives the user's call and converts the audio signal into digital data. The converted digital data is then sent to the server via a secure protocol.

[0634] Step 2:

[0635] The server analyzes the received digital data and uses a generative model to convert the audio signal into text data. From this text data, keywords and phrases related to fraudulent activities are detected.

[0636] Step 3:

[0637] The server scores the likelihood of fraud and prepares to generate a warning alert if the likelihood of fraud exceeds a threshold.

[0638] Step 4:

[0639] The server uses a warning mechanism to send a warning notification to the user's device and the devices of registered family members, alerting the user.

[0640] Step 5:

[0641] The system utilizes a function that allows users to review warning notifications they receive and then play back recorded call data stored on the server to examine its contents.

[0642] Step 6:

[0643] If a user or their family member suspects fraud, they should take action, such as reporting it to the police, to resolve the issue.

[0644] (Example 1)

[0645] Next, we will describe Example 1. In the following description, the data processing device 12 will be referred to as the "server" and the robot 414 as the "terminal".

[0646] Addressing fraud and malicious communications via telephone is crucial, but currently, there are limited means to detect these activities in real time and issue rapid warnings. This means many users could become victims of fraud and suffer serious damage. To solve this problem, an effective system is needed that can quickly and accurately detect fraudulent activity through voice communications and immediately notify users.

[0647] The identification process performed by the identification processing unit 290 of the data processing device 12 in Example 1 is realized by the following means.

[0648] In this invention, the server includes means for receiving voice signals and converting them into digital information, processing means for analyzing the transmitted digital information to identify signs of fraudulent activity, and notification means for issuing warnings to appropriate recipients. This makes it possible to detect fraudulent activity via voice communication in real time and immediately issue warnings to users and related parties.

[0649] "Electronic device" refers to a device that has the function of receiving communications from a user as audio signals and converting them into digital information.

[0650] A "data network" refers to a communication network used to securely transmit digital information from electronic devices to servers.

[0651] "Processing means" refers to equipment equipped with the function of analyzing transmitted digital information and performing a process to identify signs of specific fraudulent activity.

[0652] "Notification means" refers to a system that has the function of issuing warnings to users and related parties based on identified signs of fraudulent activity.

[0653] "Storage means" refers to a device or system that records audio signals and retains that recorded data for a predetermined period of time.

[0654] This invention is a system that analyzes the content of a voice communication in real time when the user initiates a voice communication, thereby preventing fraudulent activity. The system mainly consists of three elements: a terminal, a server, and a user.

[0655] When a user initiates voice communication, the terminal captures the audio signal using its built-in microphone and converts the analog signal into digital data using audio processing software such as "FFmpeg". This digital data is then transmitted to the server via the data network using the SSL / TLS protocol.

[0656] The server utilizes a natural language processing API to process the received digital information. This generative AI model converts the audio signal into text data. Subsequently, a script developed in Python is used to identify specific keywords and phrases (e.g., "ATM," "bank transfer," "urgent") from the generated text data. If potential fraud is detected, the server uses services such as Twilio to send warning notifications to the user and pre-registered recipients.

[0657] As an example, suppose a user receives a suspicious telemarketing call instructing them to "transfer the money immediately." In this case, the system detects the relevant phrase and promptly sends a warning notification to the user. This allows the user to recognize the danger of the call and take swift action.

[0658] Examples of prompt statements for a generative AI model are as follows:

[0659] "Is this call potentially a scam? Detected keywords: [ATM, bank transfer, urgent]"

[0660] This system allows for real-time monitoring of fraudulent activity via voice communication, enabling users to take prompt and appropriate action.

[0661] The flow of the specific processing in Example 1 will be explained using Figure 11.

[0662] Step 1:

[0663] When a user initiates voice communication, the device captures the audio signal using its built-in microphone. The input is an analog audio signal, and the device prepares to convert the acquired audio into digital data. Specifically, it uses a process such as FFmpeg to convert the analog signal into digital information, and the output is digital data in PCM format.

[0664] Step 2:

[0665] The terminal transmits the converted digital data to the server using a security protocol (SSL / TLS). The input is digital data in PCM format; this data is encrypted, packetized, and securely transferred to the server via the data network. The output is encrypted data packets.

[0666] Step 3:

[0667] The server receives digital data transmitted from the terminal and converts it from speech to text. The input is encrypted data packets, which are decrypted and then converted into text data using a natural language processing API. As a result of the data analysis, the output is text data corresponding to the speech.

[0668] Step 4:

[0669] The server uses a generative AI model to detect specific keywords and phrases from the obtained text data. The input is text data, and a script written in Python searches for pre-configured fraud-related keywords (e.g., "ATM," "bank transfer," "urgent"). The output is a list of the keywords found.

[0670] Step 5:

[0671] The server evaluates the potential for malicious activity and generates a warning notification when specific keywords are detected. The input is a list of detected keywords, and warning messages are sent to the user and registered recipients using services such as Twilio. The output is the sending of warning messages and a record of them.

[0672] Step 6:

[0673] The user or recipient receives a warning and reviews the content of the call. The input is the received warning message, and the server plays back the recorded data to examine its contents. The server stores the call records, which can be played back and reviewed as needed. The output is the reviewed call content and its evaluation result.

[0674] (Application Example 1)

[0675] Next, we will explain Application Example 1. In the following explanation, the data processing device 12 will be referred to as the "server" and the robot 414 as the "terminal".

[0676] In modern society, fraudulent activities conducted via acoustic communication are increasingly causing serious harm to users. Existing technologies have made it difficult to detect and warn of fraudulent activities in real time, therefore, there is a need to provide effective means to prevent fraud before it occurs.

[0677] The specific processing performed by the specific processing unit 290 of the data processing device 12 in Application Example 1 is realized by the following means.

[0678] In this invention, the server includes means for converting acoustic signals into digital data, means for analyzing the transmitted digital data and identifying signs of fraudulent activity using a generative model, and means for displaying a warning notification when fraudulent activity is identified and sending the warning to the recipient. This makes it possible to detect fraudulent activity during acoustic communication in real time and immediately warn the user.

[0679] A "user" is an individual or group that uses the system to conduct acoustic communication.

[0680] "Communication" is the process of sending and receiving acoustic signals between users.

[0681] An "acoustic signal" refers to a signal that represents the user's speech or voice.

[0682] "Device means" refers to hardware or software components used to receive the user's acoustic signals.

[0683] "Digitized data" refers to data obtained after converting an acoustic signal into a digital format.

[0684] A "wide-area communication network" is a network infrastructure used to transmit digitized data.

[0685] "Analysis means" refers to the processes and tools used to analyze the transmitted digitized data.

[0686] A "generative model" is a machine learning algorithm used to generate character data from received acoustic signals.

[0687] "Signs of fraudulent activity" refers to the identification of keywords or patterns that may be associated with fraud or crime.

[0688] A "warning mechanism" refers to a structure or method for issuing a warning when signs of fraudulent activity are detected.

[0689] "Visualization means" refers to devices or software that visually display warning notifications to the user.

[0690] "Recording means" refers to equipment or technology for recording acoustic signals and retaining them for a specified period of time.

[0691] A "recipient" is a user or their associate who has been pre-registered to receive warning notifications.

[0692] The system for realizing this invention provides a function to detect fraudulent activity during acoustic communication in real time.

[0693] First, when a user performs acoustic communication, the terminal receives an acoustic signal. The received acoustic signal is converted into digital data using the terminal's speech recognition engine. It is expected that existing speech recognition services such as the Google Speech-to-Text API will be used for this conversion. The converted digital data is then sent to a server via a wide-area communication network (e.g., the internet). In this process, the secure protocol HTTPS is used to ensure that the data is transferred safely.

[0694] On the server, a generative model operates to analyze the received digitized data. The generative model uses machine learning algorithms such as OpenAI GPT-3 to generate character data from acoustic signals. Next, this character data is analyzed to identify signs of fraudulent activity. The analysis process monitors for predefined fraud-related keywords, and if these keywords are present, they are identified as signs of fraudulent activity.

[0695] When identification occurs, the server sends a warning notification to the user and pre-registered recipients via a warning mechanism. The warning notification is immediately displayed on the terminal's visualization mechanism, allowing the user to check the communication content and take necessary measures. In addition, all acoustic signals are recorded and stored by a recording mechanism, and can be played back and reviewed later as needed.

[0696] For example, if a user receives suspicious instructions during an audio communication, such as "Please make a transfer immediately," the system will display a warning that it "may be fraudulent." This allows the user to immediately realize it is a scam and prevent becoming a victim.

[0697] An example of a prompt to a generative AI model might be something like, "Analyze the phrases used in this call and assess the likelihood of fraud."

[0698] Therefore, this system provides an effective means of reducing the risk of fraudulent activity via acoustic communication and ensuring user safety.

[0699] The flow of a specific process in Application Example 1 will be explained using Figure 12.

[0700] Step 1:

[0701] The terminal receives an acoustic signal from the user. The input is the user's voice, and the output is a raw acoustic signal. The terminal captures the acoustic signal using a microphone.

[0702] Step 2:

[0703] The device converts the received acoustic signal into digital data. The input is the acoustic signal acquired in step 1, and the output is the digital data. This conversion is performed using a speech recognition engine (e.g., Google Speech-to-Text API). The device encodes the audio features into a digital format.

[0704] Step 3:

[0705] The terminal transmits the converted digital data to the server via a wide-area communication network. The input is digital data, and the output is data transfer via a secure protocol (e.g., HTTPS). The terminal transmits the data using a network interface.

[0706] Step 4:

[0707] The server launches a generative model to analyze the received digitized data. The input is digitized data, and the output is the analysis result. The server uses a generative AI model (e.g., OpenAI GPT-3) to generate character data from acoustic signals.

[0708] Step 5:

[0709] The server identifies signs of fraudulent activity based on the generated text data. The input is text data, and the output is the identification result of the potential for fraudulent activity. The server uses an algorithm that checks for fraud-related keywords to identify specific phrases and patterns.

[0710] Step 6:

[0711] The server generates a warning notification and sends it to the user and registered recipients if it detects signs of fraudulent activity. The input is the identification result, and the output is the warning notification. The server generates the notification message and delivers it to the specified recipients.

[0712] Step 7:

[0713] The terminal displays received warning notifications. The input is the warning notification, and the output is the display of the warning message on the user interface. The terminal communicates the warning to the user using either a display or audio output.

[0714] Step 8:

[0715] The server records all audio signals and retains them for a specified period. The input is the audio signal, and the output is the stored audio recording. The server uses a database or storage system to save the data in a format that can be accessed later.

[0716] Furthermore, an emotion engine that estimates the user's emotions may be incorporated. That is, the identification processing unit 290 may use the emotion identification model 59 to estimate the user's emotions and perform identification processing using the user's emotions.

[0717] This invention provides a voice communication fraud prevention system that combines an emotion engine. This system focuses on the terminal, server, and user, and achieves more precise fraud detection and response by comprehensively analyzing the voice communications made by the user.

[0718] When a user initiates communication, the terminal converts the voice signal into digital data and sends it to the server in real time. The server receives this digital data and performs analysis. For the analysis, a generative model is used to convert the voice signal into text data and detect keywords and phrases that may be related to fraud.

[0719] Furthermore, this system incorporates an emotion engine that estimates emotions from the user's voice. The emotion engine particularly identifies emotions such as anxiety, tension, and agitation, and uses this to improve the accuracy of the analysis. For example, when a user utters words expressing anxiety such as "I'm in trouble" or "What should I do?", the system analyzes the tone, speed, and volume of their voice to understand their emotional state.

[0720] The server generates alerts based on signs of fraudulent activity and emotional shifts. If the emotion engine detects strong signs of anxiety or tension, it adjusts the level and urgency of the warning notification and promptly alerts the user and registered recipients. This warning may include specific details such as, "This call may be fraudulent and anxiety has been detected. Please investigate immediately."

[0721] The user or their family can review the content of the call and play back the recording for further analysis based on the warning notification sent. They can also take appropriate measures, such as reporting to the police, if necessary.

[0722] The system of this invention makes it possible to prevent damage from telephone-based fraud by providing dual protection through real-time analysis and emotion recognition. For example, if a user suddenly becomes anxious and starts worrying about their money, the system will highly assess the possibility of fraud, immediately generate a warning, and notify family members to encourage appropriate action.

[0723] The following describes the processing flow.

[0724] Step 1:

[0725] When the terminal initiates a call with a user, it converts the audio signal into a digital format and transmits it to the server in real time using a secure protocol.

[0726] Step 2:

[0727] The server analyzes the received digital data using a generative model and converts the audio signal into text data. It then processes the text data to detect fraud-related keywords and phrases.

[0728] Step 3:

[0729] An emotion engine installed on the server estimates the user's emotions from their voice. This engine analyzes the tone, speed, and volume of the voice to identify the emotions the user is expressing (such as anxiety or tension).

[0730] Step 4:

[0731] The server integrates the analysis results and sentiment estimation results to score the likelihood of fraud. If signs of fraudulent activity or emotional instability are detected, a warning alert is prepared.

[0732] Step 5:

[0733] The server sends a warning notification to the user's device and the devices of registered recipients. The notification includes warnings about potential fraud and emotional instability.

[0734] Step 6:

[0735] The system will review the warning notification received by the user or a family member, and play back the recorded data stored on the server to examine the call content. If fraud is suspected, the system will promptly request appropriate action.

[0736] (Example 2)

[0737] Next, we will describe Example 2. In the following description, the data processing device 12 will be referred to as the "server" and the robot 414 as the "terminal".

[0738] Fraudulent activity in voice communications presents a challenge because the content of these communications is diverse, making it difficult to prevent simply by detecting keywords. Furthermore, since a user's psychological state and emotions may be related to fraudulent activity, there is a need for highly accurate fraud detection that also takes emotional changes into account.

[0739] The identification process performed by the identification processing unit 290 of the data processing device 12 in Example 2 is realized by the following means.

[0740] In this invention, the server includes means for generating electrical data, means for estimating emotional states using emotion analysis means, and means for issuing warnings based on signs of misconduct and changes in emotion. This makes it possible to detect misconduct in voice communications with high accuracy while taking emotional states into consideration and to respond quickly.

[0741] "Acoustic signal" refers to the physical sound vibrations generated when the user's voice is captured.

[0742] "Device means" refers to a device for receiving acoustic signals and converting them into digital data.

[0743] "Electrical data" refers to audio information in digital format that has been converted by a device or means.

[0744] A "data path" refers to a means of communication for transmitting electrical data.

[0745] "Analysis" refers to the process of processing received electrical data and evaluating signs of fraudulent activity.

[0746] "Analysis means" refers to a method or device for analyzing electrical data and identifying fraudulent activity.

[0747] "Recording means" refers to a function that saves audio signals in digital format and allows them to be played back later.

[0748] "Emotion analysis means" refers to a method or device for estimating emotions from a user's voice and recognizing changes in those emotions.

[0749] A "warning" refers to a message that alerts users or other relevant parties to suspected fraudulent activity.

[0750] This invention provides a system for preventing fraudulent activities in voice communications, and implements this using a user, a terminal, and a server.

[0751] The user initiates voice communication and inputs their voice through the terminal. The terminal converts the input acoustic signal into electrical data. This process is achieved using the audio processor built into the terminal.

[0752] The terminal transmits the converted electrical data to the server. A secure and efficient data path is ensured for communication, and generally, communication methods using encryption technology are employed.

[0753] The server analyzes the received electrical data. A generative AI model is used for the analysis, generating text data from the audio. Based on this text data, natural language processing (NLP) techniques are used to detect keywords related to fraud. The server also utilizes sentiment analysis to estimate the user's emotional state from the audio data. Sentiment analysis determines the psychological state from the tone, speed, and volume of the voice, identifying emotions such as anxiety and tension.

[0754] Based on the analysis above, the server generates a warning when it detects signs of fraudulent activity and changes in mood. The warning is sent to the user and registered recipients, prompting them to take necessary action. For example, the warning may include a specific message such as, "This call may be fraudulent and anxiety has been detected. Please check immediately."

[0755] This system's real-time fraud detection and rapid response capabilities allow users and their families to prevent fraud and improve the security of voice communications. A specific use case would be if a user, speaking in a tense tone during a call, says, "I'm worried about my finances." The system would immediately detect the risk and send a warning notification to the family.

[0756] An example of a prompt message would be, "Analyze the voice that expresses the user's anxiety and assess the likelihood of fraud." This prompt allows the generative AI model to appropriately evaluate the user's psychological state and the content of their statements, enabling sophisticated fraud detection.

[0757] The flow of the specific processing in Example 2 will be explained using Figure 13.

[0758] Step 1:

[0759] The user inputs an acoustic signal to the terminal to initiate voice communication. This acoustic signal is then generated.

[0760] Step 2:

[0761] The terminal converts the received acoustic signal into electrical data using a digital audio processor. This process transforms the signal into digital data. The output is the converted electrical data.

[0762] Step 3:

[0763] The terminal transmits the converted electrical data to the server via the data path. This allows the server to obtain the data necessary for analysis.

[0764] Step 4:

[0765] The server converts received electrical data into text data using a generation AI model. The input is electrical data, and the output is the generated text data. This conversion process is performed via a speech recognition engine.

[0766] Step 5:

[0767] The server uses natural language processing techniques to detect fraud-related keywords from text data. The input is text data, and the output is the identified keywords or their absence.

[0768] Step 6:

[0769] The server estimates emotions using emotion analysis tools. The input is electrical data, and the output is the estimated emotional state. Specifically, it analyzes the tone and speed of speech to identify feelings of anxiety and tension.

[0770] Step 7:

[0771] The server generates an alert based on an assessment of signs of misconduct and changes in sentiment. The alert is customized using prompts as needed. The content of the alert message is generated at this stage.

[0772] Step 8:

[0773] The server sends the generated warnings to the user and registered recipients. The output is the specific warning message that the user and recipients receive. This warning gives the user an opportunity to take corrective action.

[0774] (Application Example 2)

[0775] Next, we will explain application example 2. In the following explanation, the data processing device 12 will be referred to as the "server" and the robot 414 as the "terminal".

[0776] In recent years, with the advancement of communication technology, voice-based fraud has been increasing, posing a serious problem, especially for vulnerable individuals such as the elderly. To effectively protect users from such fraud, there is a need for a system that can quickly and accurately detect signs of fraudulent activity and also provide warnings in response to changes in the user's emotional state.

[0777] The specific processing performed by the specific processing unit 290 of the data processing device 12 in Application Example 2 is realized by the following means.

[0778] In this invention, the server includes a device means for converting voice signals into digital data and transmitting them over a network, an analysis means for analyzing the transmitted digital data to identify signs of fraudulent activity, and a device means for estimating the user's emotional state using an emotion analysis engine and detecting anxiety and tension. This makes it possible to detect signs that should raise suspicion of fraudulent activity during the user's voice communication, understand the user's emotional changes at that time, and quickly display and transmit appropriate warnings to the user's visual device and pre-registered receiving equipment.

[0779] "Device means" refers to equipment that has the function of receiving communications as audio signals, converting them into digital data, and transmitting them over a network.

[0780] An "analysis device" is a device that has the function of identifying signs of fraudulent activity by analyzing transmitted digital data.

[0781] A "notification system" is a system that sends a warning to users or registered receiving equipment when signs of fraudulent activity are identified.

[0782] A "memory device" is a device that has the function of recording audio signals and retaining them for a specified period of time.

[0783] An "emotion analysis engine" is a system that estimates a user's emotional state and detects anxiety or tension by analyzing the tone, speed, and volume of their voice.

[0784] A "visual device" is a device used to display visual information to a user.

[0785] This invention provides a system for preventing fraudulent activities in voice communications, which has a configuration that comprehensively analyzes the content of user communications. The main components of the system are digitization of voice signals, data analysis, fraud detection, emotion recognition, and warning issuance.

[0786] The server receives the audio signal sent from the terminal when the user initiates communication. The received audio signal is first converted into digital data. This process uses a real-time audio analysis engine (e.g., Nuance Dragon) to handle the audio data.

[0787] The server then uses a generative AI model to generate text data from the audio data. This generative AI model (e.g., OpenAI GPT) has the ability to analyze the text data and identify keywords and phrases related to fraudulent activity. Simultaneously, the server uses an emotion analysis engine (e.g., Affectiva) to estimate the user's emotional state from their voice and detect anxiety and tension.

[0788] The user's device is equipped with a visual device (e.g., smart glasses) to display warnings visually. When the system detects signs of misconduct and changes in the user's emotional state, it displays a warning on the user's visual device and immediately sends a notification to registered receiving equipment.

[0789] As a concrete example, consider a scenario where a user is on a regular phone call and the other party approaches them with a "good offer about investing in a new financial product." In this case, the server transcribes the other party's words in real time, detects the possibility of fraud, and also detects feelings of tension or anxiety from the user's voice. Based on this, the system displays a warning such as "This may be a scam, please be careful" on the user's visual device and sends a notification to the devices of registered recipients.

[0790] Examples of prompts for the generative AI model include, "Identify potentially fraudulent phrases that can be detected from this audio data," and "Analyze the emotional shifts following the user's statements and report any signs of anxiety or tension."

[0791] In this way, the server can ensure user safety and prevent fraud.

[0792] The flow of a specific process in Application Example 2 will be explained using Figure 14.

[0793] Step 1:

[0794] When a user initiates communication, the terminal captures the audio signal using its built-in microphone. This signal is then converted into digital data in real time. This digital data is then input to the server and sent to the audio analysis engine.

[0795] Step 2:

[0796] The server analyzes the received digital data and uses a generative AI model to convert the audio data into text data. The generative AI model extracts features from the audio signal and uses prompt sentences to generate text. The output of this step is the transcribed audio data.

[0797] Step 3:

[0798] The server analyzes the generated text data to identify keywords and phrases related to fraud. This process compares specific words indicating fraudulent activity against a database to generate identification results. The output is a flag indicating the presence or absence of fraudulent activity.

[0799] Step 4:

[0800] Simultaneously, the server processes the user's voice signal using an emotion analysis engine to estimate their emotional state. Specifically, it analyzes the tone, speed, and volume of the voice to estimate the level of anxiety and tension. The output is an indicator of the user's emotional state.

[0801] Step 5:

[0802] The server issues warnings based on flags indicating fraudulent activity and emotional state indicators. It instructs visual devices to display warning messages notifying them of potential fraud and emotional changes. It also sends similar warning notifications to registered receiving devices.

[0803] Step 6:

[0804] Users receive a warning and review feedback regarding changes in call content and emotional state. This allows them to determine whether actual misconduct occurred and take appropriate action.

[0805] This allows the entire system to work together to protect users from fraudulent voice communication.

[0806] The specific processing unit 290 transmits the result of the specific processing to the robot 414. In the robot 414, the control unit 46A causes the speaker 240 and the controlled object 443 to output the result of the specific processing. The microphone 238 acquires audio indicating user input for the result of the specific processing. The control unit 46A transmits the audio data indicating user input acquired by the microphone 238 to the data processing unit 12. In the data processing unit 12, the specific processing unit 290 acquires the audio data.

[0807] Data generation model 58 is a type of so-called generative AI (Artificial Intelligence). One example of data generation model 58 is ChatGPT (Internet search<URL: https: / / openai.com / blog / chatgpt> ), Gemini (Internet search) <url: https: gemini.google.com ?hl="ja">Examples of generative AI include the following. The data generation model 58 is obtained by performing deep learning on a neural network. The data generation model 58 is input with prompts containing instructions, and with inference data such as audio data representing speech, text data representing text, and image data representing images. The data generation model 58 infers from the input inference data according to the instructions indicated by the prompts, and outputs the inference results in data formats such as audio data and text data. Here, inference refers to, for example, analysis, classification, prediction, and / or summarization.

[0808] In the above embodiment, an example was given in which the specific processing is performed by the data processing device 12, but the technology of this disclosure is not limited thereto, and the specific processing may also be performed by the robot 414.

[0809] Furthermore, the emotion identification model 59, acting as an emotion engine, may determine the user's emotion according to a specific mapping. Specifically, the emotion identification model 59 may determine the user's emotion according to a specific mapping, which is an emotion map (see Figure 9). Similarly, the emotion identification model 59 may also determine the robot's emotion, and the identification processing unit 290 may perform identification processing using the robot's emotion.

[0810] Figure 9 shows an emotion map 400 in which multiple emotions are mapped. In the emotion map 400, emotions are arranged in concentric circles radiating from the center. The closer to the center of the concentric circles, the more primitive the emotions are located. Further out of the concentric circles, emotions representing states and actions arising from mental states are located. Emotion is a concept that includes feelings and mental states. On the left side of the concentric circles, emotions that are generally generated from reactions occurring in the brain are located. On the right side of the concentric circles, emotions that are generally induced by situational judgment are located. Above and below the concentric circles, emotions that are generally generated from reactions occurring in the brain and induced by situational judgment are located. In addition, the emotion of "pleasure" is located on the upper side of the concentric circles, and the emotion of "displeasure" is located on the lower side. Thus, in the emotion map 400, multiple emotions are mapped based on the structure in which emotions arise, and emotions that are likely to occur simultaneously are mapped close together.

[0811] These emotions are distributed at the 3 o'clock position on the Emotion Map 400, and usually fluctuate between feelings of security and anxiety. In the right half of the Emotion Map 400, situational awareness takes precedence over internal feelings, resulting in a calm impression.

[0812] The inside of the Emotion Map 400 represents inner thoughts, while the outside represents actions. Therefore, the further you go from the outside of the Emotion Map 400, the more visible (expressed in actions) your emotions become.

[0813] Here, human emotions are based on various balances, such as posture and blood sugar levels. When these balances deviate from the ideal, it results in discomfort, and when they approach the ideal, it results in pleasure. Similarly, in robots, cars, motorcycles, etc., emotions can be created based on various balances, such as posture and battery level. When these balances deviate from the ideal, it results in discomfort, and when they approach the ideal, it results in pleasure. The emotion map can be generated, for example, based on Dr. Mitsuyoshi's emotion map (Research on a system for analyzing brain physiological signals of speech emotion recognition and emotion, Tokushima University, doctoral dissertation: https: / / ci.nii.ac.jp / naid / 500000375379). The left half of the emotion map contains emotions belonging to a region called "response," where sensation is dominant. The right half of the emotion map contains emotions belonging to a region called "situation," where situational awareness is dominant.

[0814] The emotion map defines two emotions that promote learning. One is the emotion around the middle of the negative "repentance" and "reflection" on the situation side. In other words, it is when the robot experiences negative emotions such as "I never want to feel this way again" or "I don't want to be scolded again." The other is the emotion around the positive "desire" on the reaction side. In other words, it is when the robot has positive feelings such as "I want more" or "I want to know more."

[0815] The emotion identification model 59 inputs user input into a pre-trained neural network, obtains emotion values ​​representing each emotion shown in the emotion map 400, and determines the user's emotion. This neural network is pre-trained based on multiple training data sets, which are combinations of user input and emotion values ​​representing each emotion shown in the emotion map 400. Furthermore, this neural network is trained so that emotions located close together have similar values, as shown in the emotion map 900 in Figure 10. Figure 10 shows an example where multiple emotions such as "reassured," "calm," and "confident" have similar emotion values.

[0816] The above description primarily focuses on the functions of the data processing device 12 in relation to this disclosure. However, the system related to this disclosure is not necessarily implemented on a server. The system related to this disclosure may be implemented as a general information processing system. This disclosure may be implemented, for example, as a software program that runs on a personal computer or as an application that runs on a smartphone. The method related to this disclosure may be provided to users in SaaS (Software as a Service) format.

[0817] In the above embodiment, an example was given in which a specific process is performed by a single computer 22. However, the technology of this disclosure is not limited thereto, and a distributed processing of the specific process may be performed by multiple computers, including computer 22. For example, a data generation model 58 may be provided in an external device of the data processing device 12, and the external device may generate data according to the input data.

[0818] In the above embodiment, an example was given in which the specific processing program 56 is stored in the storage 32, but the technology of this disclosure is not limited thereto. For example, the specific processing program 56 may be stored in a portable, computer-readable, non-temporary storage medium such as a USB (Universal Serial Bus) memory. The specific processing program 56 stored in the non-temporary storage medium is installed in the computer 22 of the data processing device 12. The processor 28 executes specific processing according to the specific processing program 56.

[0819] Alternatively, the specific processing program 56 may be stored in a storage device such as a server connected to the data processing device 12 via the network 54, and the specific processing program 56 may be downloaded and installed on the computer 22 in response to a request from the data processing device 12.

[0820] Furthermore, it is not necessary to store the entirety of the specific processing program 56 in a storage device such as a server connected to the data processing device 12 via the network 54, or to store the entirety of the specific processing program 56 in the storage 32; it is acceptable to store only a portion of the specific processing program 56.

[0821] The following types of processors can be used as hardware resources to perform specific processing. Examples of processors include a CPU, a general-purpose processor that functions as a hardware resource to perform specific processing by executing software, i.e., a program. Other examples of processors include dedicated electrical circuits, such as FPGAs (Field-Programmable Gate Arrays), PLDs (Programmable Logic Devices), or ASICs (Application Specific Integrated Circuits), which have circuit configurations specifically designed to perform specific processing. All of these processors have built-in or connected memory, and all of them perform specific processing by using memory.

[0822] The hardware resource that performs a specific process may consist of one of these various processors, or it may consist of a combination of two or more processors of the same or different types (for example, a combination of multiple FPGAs, or a combination of a CPU and an FPGA). Alternatively, the hardware resource that performs a specific process may consist of a single processor.

[0823] Examples of configurations using a single processor include, firstly, a configuration in which one or more CPUs and software are combined to form a single processor, and this processor functions as a hardware resource that performs a specific process. Secondly, there is a configuration using a processor that realizes the functions of the entire system, including multiple hardware resources that perform a specific process, on a single IC chip, as exemplified by SoCs (System-on-a-chip). In this way, a specific process is realized using one or more of the above types of processors as hardware resources.

[0824] Furthermore, the hardware structure of these various processors can more specifically utilize electrical circuits that combine circuit elements such as semiconductor devices. Also, the specific processing described above is merely an example. Therefore, it goes without saying that unnecessary steps can be deleted, new steps added, or the processing order rearranged, as long as it does not deviate from the main purpose.

[0825] The descriptions and illustrations presented above are detailed explanations of the technical aspects of this disclosure and are merely examples of the technical aspects. For example, the above descriptions of the structure, function, operation, and effect are examples of the structure, function, operation, and effect of the technical aspects of this disclosure. Therefore, it goes without saying that you may delete unnecessary parts, add new elements, or replace elements in the descriptions and illustrations presented above, as long as you do not deviate from the essence of the technical aspects of this disclosure. Furthermore, in order to avoid confusion and facilitate understanding of the technical aspects of this disclosure, explanations of common technical knowledge and the like that do not require special explanation to enable the implementation of the technical aspects of this disclosure have been omitted from the descriptions and illustrations presented above.

[0826] All documents, patent applications, and technical standards described herein are incorporated by reference to the same extent as if each individual document, patent application, and technical standard were specifically and individually noted as being incorporated by reference.

[0827] The following is further disclosed regarding the embodiments described above.

[0828] (Claim 1)

[0829] A terminal means for receiving user communications as voice signals,

[0830] Means for converting the aforementioned audio signal into digital data and transmitting it over a network,

[0831] An analytical means for analyzing transmitted digital data and identifying signs of fraudulent activity,

[0832] A warning system that issues a warning when signs of such fraudulent activity are identified,

[0833] A recording means for recording an audio signal and retaining it for a specified period of time,

[0834] A system that includes this.

[0835] (Claim 2)

[0836] The analysis means generates text data from the audio signal using a generative model,

[0837] A system according to claim 1 for identifying fraud-related keywords.

[0838] (Claim 3)

[0839] The warning system sends warning notifications to users and pre-registered recipients, prompting them to take appropriate action.

[0840] The system according to claim 1.

[0841] "Example 1"

[0842] (Claim 1)

[0843] Electronic means for receiving communications as voice signals,

[0844] Means for converting the aforementioned audio signal into digital information and transmitting it via a data network,

[0845] A processing means for analyzing transmitted digital information and identifying signs of fraudulent activity,

[0846] A notification mechanism for issuing a warning when signs of such fraudulent activity are identified,

[0847] A storage means for recording audio signals and retaining them for a specified period,

[0848] A system that includes this.

[0849] (Claim 2)

[0850] The processing means generates character data from the audio signal using a generation model,

[0851] A system according to claim 1 for identifying fraud-related keywords.

[0852] (Claim 3)

[0853] The notification system sends warning notifications to users and pre-registered recipients, prompting them to take appropriate action.

[0854] The system according to claim 1.

[0855] "Application Example 1"

[0856] (Claim 1)

[0857] A device and means for receiving user communications as acoustic signals,

[0858] Means for converting the aforementioned acoustic signal into digital data and transmitting it via a wide-area communication network,

[0859] An analytical means for analyzing transmitted digitized data and identifying signs of fraudulent activity,

[0860] A warning system that issues a warning when signs of the fraudulent activity are identified,

[0861] A recording means for recording an acoustic signal and retaining it for a specified period of time,

[0862] A means of displaying warning notifications,

[0863] A system that includes this.

[0864] (Claim 2)

[0865] The system according to claim 1, wherein the analysis means generates character data from the acoustic signal using a generative model and identifies fraud-related keywords.

[0866] (Claim 3)

[0867] The system according to claim 1, wherein the warning means sends a warning notification to the user and pre-registered recipients, and prompts them to take action based on this notification.

[0868] "Example 2 of combining an emotion engine"

[0869] (Claim 1)

[0870] A device and means for receiving user communications as acoustic signals,

[0871] Means for converting the aforementioned acoustic signal into electrical data and transmitting it via a data path,

[0872] An analytical means for analyzing transmitted electrical data and identifying signs of fraudulent activity,

[0873] A recording means for recording an acoustic signal and retaining it for a specified period of time,

[0874] The aforementioned analysis means uses emotion analysis means for estimating emotional states from electrical data, and means for identifying fraudulent behavior based on changes in emotion.

[0875] Means of issuing warnings based on signs of misconduct and changes in mood,

[0876] A system that includes this.

[0877] (Claim 2)

[0878] The system according to claim 1, wherein the analysis means generates textual data from the electrical data using a generative model and identifies fraud-related terms.

[0879] (Claim 3)

[0880] The system according to claim 1, wherein the warning means sends a warning notification to the user and pre-registered recipients, prompting them to take action based on this notification.

[0881] "Application example 2 when combining with an emotional engine"

[0882] (Claim 1)

[0883] A device and means for receiving user communications as audio signals,

[0884] A device means for converting the aforementioned audio signal into digital data and transmitting it via a network,

[0885] An analytical means for analyzing transmitted digital data to identify signs of fraudulent activity,

[0886] A notification mechanism for issuing a warning when signs of such fraudulent activity are identified,

[0887] A storage means for recording audio signals and retaining them for a specified period,

[0888] A device means that estimates the user's emotional state using an emotion analysis engine and detects anxiety and tension,

[0889] A functional means that displays a warning message on the user's visual device and sends a notification to a pre-registered receiving device,

[0890] A system that includes this.

[0891] (Claim 2)

[0892] The system according to claim 1, wherein the analysis means generates text data from the audio signal using a generating AI and identifies fraud-related keywords.

[0893] (Claim 3)

[0894] The system according to claim 1, wherein the notification means transmits a warning notification to the user's visual device and pre-registered receiving equipment, and recommends a response based on this notification. [Explanation of Symbols]

[0895] 10, 210, 310, 410 Data Processing Systems 12 Data Processing Devices 14 Smart Devices 214 Smart Glasses 314 Headset-type terminal 414 Robots< / url:> < / url:> < / url:> < / url:>

Claims

1. A terminal means for receiving user communications as voice signals, Means for converting the aforementioned audio signal into digital data and transmitting it over a network, An analytical means for analyzing transmitted digital data and identifying signs of fraudulent activity, A warning system that issues a warning when signs of such fraudulent activity are identified, A recording means for recording an audio signal and retaining it for a specified period of time, A system that includes this.

2. The analysis means generates text data from the audio signal using a generative model, A system according to claim 1 for identifying keywords related to fraud.

3. The warning system sends warning notifications to users and pre-registered recipients, prompting them to take appropriate action. The system according to claim 1.

Citation Information

Patent Citations

  • Persona chatbot control method and system

    JP2022180282A