system

The system addresses the challenge of unauthorized voice generation by capturing, preprocessing, and analyzing voice data to detect anomalies, ensuring real-time and accurate detection of fraudulent speech, thereby enhancing user safety.

JP2026038010APending Publication Date: 2026-03-06SOFTBANK GROUP CORP
View PDF 1 Cites 0 Cited by

Patent Information

Authority / Receiving Office
JP · JP
Patent Type
Applications
Current Assignee / Owner
Filing Date
2024-08-22
Publication Date
2026-03-06

AI Technical Summary

Technical Problem

Current voice generation AI technologies lack the ability to detect unauthorized voice generation in real time and respond quickly, posing a risk to user safety due to the inability to accurately identify unnatural changes or traces of synthesis in voice data.

Method used

A system that captures voice data, performs preprocessing, extracts acoustic features, compares them with existing voice profiles, and generates alerts for anomalies, ensuring real-time detection and notification of fraudulent speech generation.

Benefits of technology

The system provides highly accurate and timely detection of fraudulent voice generation, protecting user safety by generating alerts via SMS, email, or in-app notifications.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure 2026038010000001
    Figure 2026038010000001
  • Figure 2026038010000002
    Figure 2026038010000002
  • Figure 2026038010000003
    Figure 2026038010000003
Patent Text Reader

Abstract

To provide a system that protects users from fraudulent voice generation. [Solution] The system includes a means for capturing voice data, a means for preprocessing the captured voice data, a means for extracting acoustic features from the preprocessed voice data, a means for comparing the extracted acoustic features with an existing voice profile, a means for analyzing the comparison results to detect anomalies, and a means for generating and notifying an alert if an anomaly is detected.
Need to check novelty before this filing date? Find Prior Art

Description

[Technical Field]

[0001] The technology of the present disclosure relates to a system. [Background technology]

[0002] Patent document 1 discloses a persona chatbot control method performed by at least one processor, the method including the steps of receiving a user utterance, adding the user utterance to a prompt including an instruction sentence related to a description of the chatbot character, encoding the prompt, and inputting the encoded prompt into a language model to generate a chatbot utterance in response to the user utterance. [Prior art documents] [Patent documents]

[0003] [Patent Document 1] Japanese Patent Publication No. 2022-180282 Summary of the Invention [Problem to be solved by the invention]

[0004] In recent years, advances in voice generation AI technology have led to an increase in crimes involving the unauthorized generation and misuse of voice. In this context, it is necessary to ensure the reliability of voice data and provide an environment in which users can safely use voice communications. However, current technology lacks the means to detect unauthorized voice generation in real time and respond quickly. In particular, there is a need for highly accurate detection of unnatural changes in voice data and traces of synthesis. The present invention aims to solve these issues and provide a system that protects users from unauthorized voice generation. [Means for solving the problem]

[0005] This invention provides a system that performs an integrated process from capturing voice data to detecting anomalies. Specifically, it includes a means for capturing voice data and performing preprocessing (noise removal and normalization). Next, it includes a means for extracting acoustic features from the preprocessed voice data and comparing those features with an existing voice profile. It also includes a means for analyzing the comparison results to detect unnatural changes or traces of synthesis in the voice data. If an anomaly is detected, an alert is generated to notify the user or administrator, enabling a prompt response. In this way, a system is provided that detects fraudulent speech generation AI with high accuracy and in real time, protecting user safety.

[0006] "Audio data" refers to data in which an audio signal is recorded in digital format.

[0007] "Capture" refers to the act of acquiring and saving audio data in real time.

[0008] "Preprocessing" refers to performing necessary processing such as noise removal and normalization on the captured audio data.

[0009] "Acoustic features" refer to characteristic values ​​such as Mel-Frequency Cepstrum Coefficients (MFCCs) and spectral features extracted from speech data.

[0010] A "voice profile" is a collection of voice samples that a user pre-registers.

[0011] "Comparison" refers to the act of evaluating the similarity between the extracted acoustic features and the voice profile using cosine similarity, Euclidean distance, or the like.

[0012] "Anomaly detection" involves analyzing the comparison results and identifying unnatural voice changes or signs of synthesis.

[0013] An "alert" is a warning or notification information that is generated when fraud is confirmed as a result of an abnormality detection.

[0014] "Notification" means the act of communicating a generated alert to a user or administrator via SMS, email, or in-app notification. [Brief explanation of the drawings]

[0015] [Figure 1] 1 is a conceptual diagram showing an example of the configuration of a data processing system according to a first embodiment. [Figure 2] 1 is a conceptual diagram showing an example of main functions of a data processing device and a smart device according to a first embodiment. [Figure 3] FIG. 10 is a conceptual diagram showing an example of the configuration of a data processing system according to a second embodiment. [Figure 4] FIG. 10 is a conceptual diagram showing an example of main functions of a data processing device and smart glasses according to a second embodiment. [Figure 5] FIG. 10 is a conceptual diagram showing an example of the configuration of a data processing system according to a third embodiment. [Figure 6] FIG. 11 is a conceptual diagram showing an example of main functions of a data processing device and a headset-type terminal according to a third embodiment. [Figure 7] FIG. 10 is a conceptual diagram showing an example of the configuration of a data processing system according to a fourth embodiment. [Figure 8] FIG. 10 is a conceptual diagram showing an example of main functions of a data processing device and a robot according to a fourth embodiment. [Figure 9] 1 shows an emotion map onto which multiple emotions are mapped. [Figure 10] 1 shows an emotion map onto which multiple emotions are mapped. [Figure 11] FIG. 3 is a sequence diagram showing a processing flow of the data processing system according to the first embodiment. [Figure 12] FIG. 10 is a sequence diagram showing the flow of processing in the data processing system in Application Example 1. [Figure 13] FIG. 10 is a sequence diagram showing the flow of processing in the data processing system according to the second embodiment when an emotion engine is combined. [Figure 14] FIG. 10 is a sequence diagram showing the flow of processing in the data processing system in Application Example 2 when an emotion engine is combined. DETAILED DESCRIPTION OF THE INVENTION

[0016] An example of an embodiment of a system according to the technology of the present disclosure will be described below with reference to the accompanying drawings.

[0017] First, the terms used in the following description will be explained.

[0018] In the following embodiments, a coded processor (hereinafter simply referred to as a "processor") may be a single arithmetic device or a combination of multiple arithmetic devices. Furthermore, a processor may be a single type of arithmetic device or a combination of multiple types of arithmetic devices. Examples of arithmetic devices include a CPU (Central Processing Unit), a GPU (Graphics Processing Unit), a GPGPU (General-Purpose computing on Graphics Processing Units), and an APU (Accelerated Processing Unit).

[0019] In the following embodiments, a coded RAM (Random Access Memory) is a memory in which information is temporarily stored and is used as a working memory by a processor.

[0020] In the following embodiments, the coded storage is one or more non-volatile storage devices that store various programs, various parameters, etc. Examples of non-volatile storage devices include flash memory (SSD (Solid State Drive)), magnetic disks (e.g., hard disks), and magnetic tapes.

[0021] In the following embodiments, a communication I / F (Interface) with a symbol is an interface including a communication processor, an antenna, etc. The communication I / F controls communication between multiple computers. Examples of communication standards applied to the communication I / F include wireless communication standards including 5G (5th Generation Mobile Communication System), Wi-Fi (registered trademark), Bluetooth (registered trademark), etc.

[0022] In the following embodiments, "A and / or B" is synonymous with "at least one of A and B." In other words, "A and / or B" means that it may be only A, only B, or a combination of A and B. Furthermore, in this specification, the same concept as "A and / or B" is also applied when three or more things are expressed connected by "and / or."

[0023] [First embodiment]

[0024] FIG. 1 shows an example of the configuration of a data processing system 10 according to the first embodiment.

[0025] 1, a data processing system 10 includes a data processing device 12 and a smart device 14. An example of the data processing device 12 is a server.

[0026] The data processing device 12 includes a computer 22, a database 24, and a communication I / F 26. The computer 22 is an example of a "computer" according to the technology of the present disclosure. The computer 22 includes a processor 28, a RAM 30, and a storage 32. The processor 28, the RAM 30, and the storage 32 are connected to a bus 34. The database 24 and the communication I / F 26 are also connected to the bus 34. The communication I / F 26 is connected to a network 54. Examples of the network 54 include a WAN (Wide Area Network) and / or a LAN (Local Area Network).

[0027] The smart device 14 includes a computer 36, a reception device 38, an output device 40, a camera 42, and a communication I / F 44. The computer 36 includes a processor 46, a RAM 48, and a storage 50. The processor 46, the RAM 48, and the storage 50 are connected to a bus 52. The reception device 38, the output device 40, and the camera 42 are also connected to the bus 52.

[0028] The reception device 38 includes a touch panel 38A, a microphone 38B, and the like, and receives user input. The touch panel 38A detects contact with an indicator (for example, a pen or a finger) to receive user input by the touch of the indicator. The microphone 38B detects the user's voice to receive user input by voice. The control unit 46A transmits data indicating the user input received by the touch panel 38A and the microphone 38B to the data processing device 12. In the data processing device 12, the specific processing unit 290 acquires the data indicating the user input.

[0029] The output device 40 includes a display 40A and a speaker 40B, and presents data to the user 20 by outputting the data in a form of expression that the user 20 can perceive (for example, audio and / or text). The display 40A displays visible information such as text and images in accordance with instructions from the processor 46. The speaker 40B outputs audio in accordance with instructions from the processor 46. The camera 42 is a compact digital camera equipped with an optical system including a lens, aperture, and shutter, and an imaging element such as a CMOS (Complementary Metal-Oxide-Semiconductor) image sensor or a CCD (Charge Coupled Device) image sensor.

[0030] The communication I / F 44 is connected to a network 54. The communication I / Fs 44 and 26 control the exchange of various information between the processor 46 and the processor 28 via the network 54.

[0031] FIG. 2 shows an example of the main functions of the data processing device 12 and the smart device 14.

[0032] 2, in the data processing device 12, a specific process is performed by the processor 28. A specific processing program 56 is stored in the storage 32. The specific processing program 56 is an example of a "program" according to the technology of the present disclosure. The processor 28 reads the specific processing program 56 from the storage 32 and executes the read specific processing program 56 on the RAM 30. The specific process is realized by the processor 28 operating as a specific processing unit 290 in accordance with the specific processing program 56 executed on the RAM 30.

[0033] The storage 32 stores a data generation model 58 and an emotion identification model 59. The data generation model 58 and the emotion identification model 59 are used by the identification processing unit 290.

[0034] In the smart device 14, the processor 46 performs the reception output process. The storage 50 stores a reception output program 60. The reception output program 60 is used in conjunction with the specific processing program 56 by the data processing system 10. The processor 46 reads the reception output program 60 from the storage 50 and executes the read reception output program 60 on the RAM 48. The reception output process is realized by the processor 46 operating as the control unit 46A in accordance with the reception output program 60 executed on the RAM 48.

[0035] Next, a description will be given of the specific processing performed by the specific processing unit 290 of the data processing device 12. In the following description, the data processing device 12 will be referred to as a "server" and the smart device 14 will be referred to as a "terminal."

[0036] This invention relates to a system for detecting fraudulent voice data generated by a voice generation AI in real time and ensuring the safety of users. Specific embodiments of the system are described below.

[0037] System Configuration

[0038] The system consists of the following major components:

[0039] 1. Device: The device that captures the audio data (e.g., a smartphone or PC)

[0040] 2. Server: The unit that processes and analyzes the captured audio data.

[0041] 3. User: A person who uses the system and checks alerts

[0042] Specific processing of the program

[0043] Capture and transmit audio data

[0044] When a device starts a voice call or message transmission, it immediately captures the voice data, which is then sent to the server in real time.

[0045] Examples:

[0046] When the user starts a call, the terminal automatically starts monitoring the voice data and sends that data to the server.

[0047] Audio data preprocessing

[0048] The server performs the necessary pre-processing on the received audio data, which includes the following steps:

[0049] Noise Reduction: Reduce background noise from audio data.

[0050] Normalization: Adjusting audio levels to a certain standard.

[0051] Acoustic feature extraction

[0052] The server extracts acoustic features from the preprocessed speech data, including Mel-Frequency Cepstral Coefficients (MFCCs) and spectral features.

[0053] Examples:

[0054] The server calculates MFCCs from the noise-removed and normalized audio data and extracts them as features.

[0055] Comparison with audio profiles

[0056] The server compares the extracted acoustic features with existing voice profiles using statistical methods such as cosine similarity and Euclidean distance.

[0057] Examples:

[0058] The server compares the voice samples previously registered by the user and calculates how well the current voice data matches.

[0059] Anomaly detection and alerting

[0060] If the comparison detects any abnormal changes or signs of unnatural synthesis, the server generates an alert, which is immediately sent to the user via SMS, email, or in-app notification.

[0061] Examples:

[0062] The server detects unnatural voice changes and sends a warning message to the user's smartphone.

[0063] Saving logs

[0064] All processing results and alert information are stored as logs by the server and used for later analysis and auditing, which increases the transparency and reliability of the system.

[0065] In this way, the system analyzes the user's voice data in real time, detects fraudulent voice production with high accuracy, and protects the user's safety.

[0066] The processing flow will be explained below.

[0067] Step 1:

[0068] The device captures voice data in real time when the user initiates a call or voice message.

[0069] Save the audio data in a specified format (e.g. WAV, MP3).

[0070] Step 2:

[0071] The device transmits the captured audio data to the server.

[0072] Encrypted communications are used for data transmission to ensure data security.

[0073] Step 3:

[0074] The server stores the received audio data and starts pre-processing.

[0075] Noise Reduction: Performs a filtering process to reduce background noise.

[0076] Normalize: Adjust the audio level to a certain standard.

[0077] Step 4:

[0078] The server extracts acoustic features from the preprocessed speech data.

[0079] Calculates Mel-Frequency Cepstral Coefficients (MFCC).

[0080] Analyze the spectral features.

[0081] Step 5:

[0082] The server compares the extracted acoustic features with existing voice profiles.

[0083] Statistical methods such as cosine similarity or Euclidean distance are used.

[0084] Step 6:

[0085] The server analyzes the comparison results to detect any unnatural voice changes or signs of synthesis.

[0086] Apply anomaly detection algorithms and flag any anomalies.

[0087] Step 7:

[0088] The server generates an alert if an anomaly is detected.

[0089] The alert content includes a warning message and details of the abnormality detection.

[0090] Step 8:

[0091] The server notifies the user of the generated alert.

[0092] Notification methods include SMS, email, or in-app notifications.

[0093] Step 9:

[0094] The user receives and confirms the notification.

[0095] If necessary, take appropriate action (e.g., report, suspend use of services).

[0096] Step 10:

[0097] The server stores all processing results and alert information as logs.

[0098] The saved logs can be used for later analysis and auditing.

[0099] The above is the specific processing flow of this system.

[0100] Example 1

[0101] Next, a description will be given of Example 1. In the following description, the data processing device 12 will be referred to as a "server" and the smart device 14 will be referred to as a "terminal."

[0102] In current voice communication systems, the generation of fraudulent voice data using generative AI models is a problem, threatening user safety. Effective methods for detecting such fraudulent voice generation in real time and immediately notifying users are needed. However, existing systems lack sufficient reliability and speed because they are unable to integrate multiple processes, such as voice data capture, preprocessing, anomaly detection, alert generation, and log storage.

[0103] The specific processing by the specific processing unit 290 of the data processing device 12 in the first embodiment is realized by the following means.

[0104] In this invention, the server includes means for capturing voice data, means for preprocessing the captured voice data, means for extracting acoustic features from the preprocessed voice data, means for comparing the extracted acoustic features with an existing voice profile, means for analyzing the comparison result to detect anomalies, means for generating and notifying an alert when an anomaly is detected, and means for storing the processing results and alert information in a database. This enables highly accurate real-time detection and immediate notification of improper voice generation, and recording of all processing results.

[0105] "Audio data" refers to digital or analog data that records sound.

[0106] "Capture" is the process of acquiring audio data.

[0107] "Preprocessing" refers to processing performed to make captured audio data easier to analyze, and includes operations such as noise removal and normalization.

[0108] "Acoustic features" are numerical representations of the characteristics of speech data, and include Mel-Frequency Cepstrum Coefficients (MFCCs) and spectral features.

[0109] An "audio profile" is reference data constructed using features of pre-registered audio data.

[0110] "Comparison" is a process of evaluating the extracted acoustic features and the features of the voice profile using statistical methods.

[0111] "Anomaly detection" is the process of analyzing the comparison results to find incorrect speech production or unusual changes.

[0112] An "alert" is a warning or notification that is generated when an abnormality is detected.

[0113] "Notification" is the process of communicating generated alerts to users, including by means of SMS, email, in-app notifications, etc.

[0114] A "database" is a system for storing processing results and alert information.

[0115] "Storage" is the process of recording processing results and alert information in a database.

[0116] This invention relates to a system that detects fraudulent voice data generation by voice generation AI in real time and ensures user safety.

[0117] System Configuration

[0118] The system consists of the following major components:

[0119] 1. Device: The device that captures the audio data (e.g., a smartphone or PC)

[0120] 2. Server: The unit that processes and analyzes the captured audio data.

[0121] 3. User: A person who uses the system and checks alerts

[0122] Specific processing of the program

[0123] Capture and transmit audio data

[0124] When a user initiates a voice call or sends a message, the device captures the voice data using the built-in microphone and transmits it to the server in real time. The hardware used is the microphone of a smartphone or PC, and the software is a dedicated voice capture application (e.g., a homemade VoiceCaptureApp).

[0125] Specific behavior:

[0126] The device automatically starts capturing audio when the user starts a call and sends the data to a remote server using SSL / TLS encryption, ensuring data security.

[0127] Audio data preprocessing

[0128] The server performs preprocessing on the received audio data, which mainly includes noise removal and voice level normalization. Specifically, it applies a filter to the audio data to remove background noise, and then normalizes the voice level.

[0129] Specific behavior:

[0130] The server uses Python's SciPy library to apply a bandpass filter to remove background noise.

[0131] After noise removal, the audio levels are normalized to the range 0 to 1 using the normalize function from the Librosa library.

[0132] Acoustic feature extraction

[0133] The server extracts acoustic features from the preprocessed speech data, including Mel-Frequency Cepstral Coefficients (MFCCs) and spectral features.

[0134] Specific behavior:

[0135] The server extracts MFCCs from the audio data using the mfcc function in the Librosa library and stores them in a database as acoustic features.

[0136] Comparison with audio profiles

[0137] The server compares the extracted acoustic features with pre-trained voice profiles using cosine similarity or Euclidean distance.

[0138] Specific behavior:

[0139] The server uses the cosine function in the SciPy library to calculate the similarity between the features of the current voice data and the features of the pre-registered voice profile.

[0140] If the calculation results below a certain threshold, it is flagged as a likely incorrect speech production.

[0141] Anomaly detection and alerting

[0142] If the server detects an anomaly in the audio data, it generates an alert and notifies the user immediately via SMS, email, or in-app notification.

[0143] Specific behavior:

[0144] The server uses an anomaly detection algorithm (e.g., Scikit-learn's Isolation Forest) to detect unnatural changes in the audio data.

[0145] If an anomaly is detected, a warning message is sent to the user's smartphone using Twilio's SMS API, and email notifications are sent using the SendGrid API.

[0146] Saving logs

[0147] The server stores all processing results and alert information in a database, allowing for later analysis and auditing, improving system reliability.

[0148] Specific behavior:

[0149] The server stores the processing results and alert information in a MySQL® database using SQLAlchemy to create database entries.

[0150] Example prompts to input to the generative AI model

[0151] "Please explain how speech generation AI detects fraudulent audio data. Please include specific steps for capturing audio data, preprocessing, extracting acoustic features, comparing with audio profiles, detecting anomalies, generating alerts, and saving logs."

[0152] In this way, the system can analyze the user's voice data in real time, detect fraudulent voice production with high accuracy, and ensure the user's safety.

[0153] The flow of the identification process in the first embodiment will be described with reference to FIG.

[0154] Processing Steps

[0155] Step 1: Capture and send audio data

[0156] The device uses a built-in microphone to capture voice data when a user initiates a voice call or sends a message.

[0157] input:

[0158] Initiate voice calls and send messages

[0159] Specific behavior:

[0160] The device starts a dedicated voice capture application (e.g., a custom VoiceCaptureApp) and starts microphone input. The captured voice data is sent to the server using encryption (SSL / TLS).

[0161] output:

[0162] Encrypted and transmitted voice data

[0163] Step 2: Preprocessing the audio data

[0164] The server performs preprocessing on the received audio data.

[0165] input:

[0166] Received audio data

[0167] Specific behavior:

[0168] The server uses Python's SciPy library to apply a bandpass filter to remove background noise, then normalizes the audio level to the range 0 to 1 using the normalize function from the Librosa library.

[0169] output:

[0170] Denoised and normalized audio data

[0171] Step 3: Extraction of acoustic features

[0172] The server extracts acoustic features from the preprocessed speech data.

[0173] input:

[0174] Denoised and normalized audio data

[0175] Specific behavior:

[0176] The server extracts MFCCs using the mfcc function in the Librosa library and stores them in a database as acoustic features.

[0177] output:

[0178] Extracted acoustic features

[0179] Step 4: Compare with the audio profile

[0180] The server compares the extracted acoustic features with pre-registered voice profiles.

[0181] input:

[0182] Extracted acoustic features

[0183] Pre-registered audio profiles

[0184] Specific behavior:

[0185] The server uses the cosine function in the SciPy library to calculate the similarity between the features of the current voice data and the features of the existing voice profile.

[0186] output:

[0187] Similarity calculation results

[0188] Step 5: Detect anomalies and generate alerts

[0189] If the server detects an abnormality in the voice data, it generates an alert and immediately notifies the user.

[0190] input:

[0191] Similarity calculation results

[0192] Specific behavior:

[0193] The server uses an anomaly detection algorithm (e.g., Scikit-learn's Isolation Forest) to detect anomalies. If an anomaly is detected, it uses Twilio's SMS API to send an alert message to the user's smartphone. For email notifications, it uses the SendGrid API.

[0194] output:

[0195] Warning message

[0196] Step 6: Save the logs

[0197] The server stores all processing results and alert information in a database.

[0198] input:

[0199] Similarity calculation results

[0200] Warning message

[0201] Specific behavior:

[0202] The server uses SQLAlchemy to create and save database entries to store the processing results and alert information in a MySQL database.

[0203] output:

[0204] Saved processing results and alert information

[0205] (Application example 1)

[0206] Next, a description will be given of Application Example 1. In the following description, the data processing device 12 will be referred to as a "server" and the smart device 14 will be referred to as a "terminal."

[0207] With the advancement of AI voice generation technology, it has become easier to generate voice data fraudulently. As a result, the risk of personal and corporate voice communications being used fraudulently and important call content being tampered with is increasing. This has a particularly severe impact on business communications that contain confidential information. Current systems face the challenge of being unable to detect such fraudulent activity in real time and respond immediately.

[0208] The specific processing by the specific processing unit 290 of the data processing device 12 in the application example 1 is realized by the following means.

[0209] In this invention, the server includes means for capturing voice data, means for preprocessing the captured voice data, means for extracting acoustic features from the preprocessed voice data, means for comparing the extracted acoustic features with a voice profile, means for detecting anomalies based on the analysis results, means for generating and notifying an alarm when an anomaly is detected, means for recording the alarm information, means for detecting fraudulent generation of voice data in real time, an application for monitoring call data, and a cloud-based processing server. This makes it possible to detect fraudulent generation of voice data in real time and respond quickly.

[0210] "Audio Data" means data for recording, storing, and transmitting audio in digital format.

[0211] A "capturing means" is a device or software for capturing audio data in real time.

[0212] The "pre-processing means" refers to a method or device for removing noise from the captured audio data and normalizing it.

[0213] "Means for extracting acoustic features" refers to a method or device that calculates and extracts characteristic information (e.g., Mel-Frequency Cepstral Coefficients (MFCCs)) from preprocessed speech data.

[0214] An "audio profile" is a collection of feature data generated based on specific audio data.

[0215] The "comparison means" refers to a method or device for comparing the extracted acoustic features with a voice profile and evaluating the degree of similarity.

[0216] The "analysis result" is information indicating the degree of agreement of the data obtained from the comparison of the acoustic features and the presence or absence of abnormalities.

[0217] "Means for detecting anomalies" refers to a method or device that determines whether there are unnatural changes or signs of fraudulent generation in the audio data based on the analysis results.

[0218] The "means for generating and notifying a warning" refers to a method or device for generating and notifying a warning message to a user when an abnormality is detected.

[0219] The "means for recording warning information" refers to a method or device for saving generated warning messages and analysis results so that they can be checked at a later date.

[0220] "Real-time detection means" refers to a method or device that can detect fraud immediately without interrupting the time when voice data is being generated.

[0221] An "application that monitors call data" is software that runs on a device such as a smartphone and monitors voice data during a call.

[0222] A "cloud-based processing server" is a remote server accessible via the Internet that pre-processes, analyzes, and compares audio data.

[0223] This invention relates to a system that detects fraudulent voice data generated by voice generation AI in real time during phone calls to ensure user security. This system automatically captures voice data, preprocesses it, extracts acoustic features, compares it with existing voice profiles, detects anomalies, generates warnings, and records it.

[0224] System Configuration

[0225] The system consists of the following major components:

[0226] 1. Terminal: A device for capturing audio data (e.g., a smartphone)

[0227] 2. Server: A unit for processing and analyzing captured audio data. A cloud-based processing server is used (e.g., AWS (registered trademark), Google (registered trademark) Cloud Platform).

[0228] 3. Users: People who use the system and receive alerts

[0229] Specific processing of the program

[0230] Capture audio data

[0231] The device uses the Twilio Voice SDK to capture voice data during a call in real time and immediately transmits that data to the server.

[0232] Audio data preprocessing

[0233] The server uses libraries (e.g., PyDub, Librosa) to denoise and normalize the audio data. This preprocessing produces clean audio data suitable for analysis.

[0234] Acoustic feature extraction

[0235] From the preprocessed audio data, the server uses Librosa to extract acoustic features such as Mel-Frequency Cepstral Coefficients (MFCCs).

[0236] Comparison with audio profiles

[0237] The extracted acoustic features are compared to existing voice profiles using machine learning frameworks such as TENSORFLOW®, using statistical methods such as cosine similarity.

[0238] Anomaly detection and alert generation

[0239] Based on the analysis results, the server determines whether the voice data has any unnatural variations or signs of fraudulent generation, and if an anomaly is detected, the server sends a warning to the user via SMS or in-app notification.

[0240] Recording warning information

[0241] All processing results and warning information are logged by the server for later analysis and auditing.

[0242] Specific examples

[0243] For example, suppose a user is using the Secure Call Detector app to make a call for a business meeting and the call content is tampered with. The system detects this tampering in real time and immediately generates an alert. The user receives the following message on their smartphone:

[0244] "Warning: Indications of unauthorized audio have been detected. The following message will appear on your screen: 'Warning: Unauthorized audio has been detected. Please ensure your call is secure.' Call recording will automatically stop and your system administrator will be notified."

[0245] This allows the user to respond quickly and ensures the safety of the call.

[0246] The system of the present invention can detect fraudulent generation of voice data in real time and respond quickly, thereby enhancing the security of individuals and businesses.

[0247] The flow of the specific processing in the application example 1 will be described with reference to FIG.

[0248] Step 1:

[0249] The terminal uses the Twilio Voice SDK to detect when a user starts a call. When a call starts, it immediately starts capturing voice data and sends that data to the server in real time. The input is the user's call voice data, and the output is the raw voice data sent to the server.

[0250] Step 2:

[0251] The server preprocesses the audio data received from the device using a library (e.g., PyDub, Librosa). This preprocessing removes noise from the audio data and normalizes the audio level. The input is raw audio data, and the output is noise-removed and normalized audio data.

[0252] Step 3:

[0253] The server extracts acoustic features from the preprocessed speech data. It uses Librosa to extract features such as Mel-Frequency Cepstral Coefficients (MFCCs). The input is the preprocessed speech data, and the output is the extracted acoustic features.

[0254] Step 4:

[0255] The server compares the extracted acoustic features with existing voice profiles. This comparison uses TensorFlow's machine learning model to evaluate similarity using statistical methods such as cosine similarity and Euclidean distance. The input is the acoustic features and the voice profile, and the output is a similarity score.

[0256] Step 5:

[0257] The server analyzes the similarity score and determines whether or not unauthorized voice data has been generated. If an abnormality is detected based on the analysis results, a warning message is generated. The input is the similarity score, and the output is the warning message.

[0258] Step 6:

[0259] The server notifies the user via SMS or in-app notification with the generated alert message. This notification is sent using Twilio's messaging API. The input is the alert message, and the output is the alert message displayed on the user's device.

[0260] Step 7:

[0261] The server stores all processing results and warning information in a database. This information is recorded using a database such as MySQL for later analysis and auditing. The input is the processing results and warning information, and the output is the log information stored in the database.

[0262] Furthermore, an emotion engine that estimates the user's emotion may be combined. That is, the identification processing unit 290 may estimate the user's emotion using the emotion identification model 59 and perform identification processing using the user's emotion.

[0263] This invention relates to a system that ensures the safety and comfort of users by detecting fraudulent voice data generated by voice generation AI in real time and recognizing the emotional state of the user from the voice data. Specific embodiments of the system are described below.

[0264] System Configuration

[0265] The system consists of the following major components:

[0266] 1. Device: The device that captures the audio data (e.g., a smartphone or PC)

[0267] 2. Server: The unit that processes, analyzes, and recognizes emotions from captured audio data.

[0268] 3. User: A person who uses the system and checks alerts

[0269] Specific processing of the program

[0270] Capture and transmit audio data

[0271] The device captures voice data in real time when the user initiates a call or voice message, and the captured voice data is sent to the server via encrypted communication means.

[0272] Examples:

[0273] The terminal starts monitoring the voice data at the same time as the user starts a call, and transmits the voice data to the server via secure communication.

[0274] Audio data preprocessing

[0275] The server performs pre-processing on the received audio data, including noise reduction to reduce background noise and normalization to adjust the audio level to a certain standard.

[0276] Acoustic feature extraction

[0277] The server extracts acoustic features from the preprocessed speech data, including Mel-Frequency Cepstral Coefficients (MFCCs) and spectral features.

[0278] Examples:

[0279] The server calculates MFCCs from the noise-removed and normalized audio data and obtains them as feature data.

[0280] Comparison with audio profiles

[0281] The server compares the extracted acoustic features with existing voice profiles using statistical methods such as cosine similarity and Euclidean distance.

[0282] Examples:

[0283] The server compares the voice samples previously registered by the user and calculates how well the current voice data matches.

[0284] Anomaly detection and alerting

[0285] If the comparison detects unnatural voice changes or signs of synthesis, the server generates an alert, which is promptly sent to the user via SMS, email, or in-app notification.

[0286] Examples:

[0287] The server detects abnormal audio changes and sends a warning message to the user's smartphone.

[0288] Additional analysis with emotion recognition

[0289] The server uses an emotion engine to analyze the user's emotional state from the voice data, and determines whether the voice data matches the user's typical emotional state.

[0290] Examples:

[0291] The server analyzes the voice data to determine whether the user is expressing emotions such as anxiety, anger, or joy.

[0292] Selecting notification methods based on emotional state

[0293] The server selects the optimal notification method based on the detected emotional state. For example, if the user is determined to be in an anxious state, an immediate notification via SMS can be sent to enhance safety.

[0294] Examples:

[0295] Based on the results of the emotion engine, if the user is feeling stressed, the server determines that a prompt response is required and selects an SMS notification.

[0296] Saving logs

[0297] All processing results and alert information are stored as logs by the server, which can be used for later analysis and auditing, ensuring the transparency and reliability of the system.

[0298] In this way, the system analyzes the user's voice data in real time, detects fraudulent voice production with high accuracy, and by recognizing the user's emotional state, selects an appropriate notification method to maintain the user's safety and comfort.

[0299] The processing flow will be explained below.

[0300] Step 1:

[0301] The device automatically captures voice data in real time when the user initiates a call or voice message.

[0302] Save the audio data in a specified format (e.g. WAV, MP3).

[0303] Step 2:

[0304] The device transmits the captured audio data to the server.

[0305] Encrypted communications are used for data transmission to ensure data security.

[0306] Step 3:

[0307] The server stores the received audio data and starts pre-processing.

[0308] The server first applies a noise reduction filter to reduce background noise.

[0309] The server then normalizes the audio levels to ensure consistency.

[0310] Step 4:

[0311] The server extracts acoustic features from the preprocessed speech data.

[0312] The server calculates Mel-Frequency Cepstral Coefficients (MFCCs) and extracts them as speech features.

[0313] The server also analyzes the spectral features and stores them as supplementary features.

[0314] Step 5:

[0315] The server compares the extracted acoustic features with existing voice profiles.

[0316] The server calculates cosine similarity and Euclidean distance to evaluate the degree of matching of the voices.

[0317] Step 6:

[0318] The server analyzes the comparison results to detect any unnatural voice changes or signs of synthesis.

[0319] The server applies an anomaly detection algorithm and flags any anomalies detected.

[0320] Step 7:

[0321] The server generates an alert if an anomaly is detected.

[0322] The server generates alert information including warning messages and details of anomaly detections.

[0323] Step 8:

[0324] The server notifies the user of the generated alert.

[0325] Notification methods include SMS, email, or in-app notifications.

[0326] Step 9:

[0327] The user receives and confirms the notification.

[0328] Based on the content of the notification, the user will take appropriate action as necessary (e.g., report the issue, stop using the service).

[0329] Step 10:

[0330] The server uses an emotion engine to analyze the user's emotional state from the voice data.

[0331] The server identifies the emotional state (e.g., anxiety, anger, joy, etc.) from the voice data.

[0332] Step 11:

[0333] The server selects the most appropriate notification method based on the detected emotional state.

[0334] For example, if a user is in a state of anxiety, they will be notified immediately via SMS to enhance safety.

[0335] Step 12:

[0336] The server stores all processing results and alert information as logs.

[0337] The saved logs can be used for later analysis and auditing.

[0338] This is the specific processing flow of this system. This system allows users to detect fraudulent generation of voice data with high accuracy, and by recognizing the user's emotional state, it enables more appropriate responses.

[0339] Example 2

[0340] Next, a description will be given of Example 2. In the following description, the data processing device 12 will be referred to as a "server" and the smart device 14 will be referred to as a "terminal."

[0341] Conventional voice data analysis systems have difficulty detecting fraudulent voice generation in real time, and there are no systems that can recognize the user's emotional state and take appropriate action. This increases the risk of voice data being manipulated or used fraudulently, and poses a challenge to ensuring user safety and comfort.

[0342] The identification process by the identification processing unit 290 of the data processing device 12 in the second embodiment is realized by the following means. In this invention, the server includes means for preprocessing voice data, means for extracting acoustic features from the preprocessed voice data, means for comparing the extracted acoustic features with an existing voice profile, and means for selecting a notification method based on the recognized emotional state. This enables highly accurate detection of fraudulent voice generation and appropriate response based on the user's emotional state.

[0343] "Audio data" is a digital representation of an acoustic signal captured using a microphone or other audio input device.

[0344] "Capture" is the process of capturing audio data in real time via a microphone or other input device.

[0345] "Preprocessing" refers to the process of applying early-stage processing such as noise removal and normalization to speech data to facilitate analysis and feature extraction.

[0346] "Encryption" is a method of converting audio data using a specific algorithm to protect it from third parties.

[0347] "Transmission" refers to the act of transferring captured audio data to another system or server via a network.

[0348] "Acoustic features" are computed physical or statistical properties, such as Mel-Frequency Cepstral Coefficients (MFCCs) or spectral features, extracted from speech data.

[0349] "Comparison" is a process in which the extracted acoustic features are matched with an existing data set called a voice profile, and the degree of match is evaluated.

[0350] "Anomaly detection" is a function that detects and reports unnatural voice changes and traces of synthesized voice as a result of voice data analysis.

[0351] An "alert" is a message or signal that notifies a user or system administrator when an abnormality is detected.

[0352] "Emotion recognition" is a technology that analyzes and determines a user's emotional state (e.g., joy, anxiety, anger, etc.) from voice data.

[0353] "Notification method selection" is the process of selecting the optimal method (e.g., SMS, email, in-app notification, etc.) to notify the user based on the recognized emotional state.

[0354] This invention relates to a system that ensures the safety and comfort of users by detecting fraudulent voice data generated by voice generation AI in real time and recognizing the emotional state of the user from the voice data. Specific embodiments are described below.

[0355] System Configuration

[0356] The system consists of the following major components:

[0357] 1. Device: The device that captures the audio data (e.g., a smartphone or computer)

[0358] 2. Server: The unit that processes, analyzes, and recognizes emotions from captured audio data.

[0359] 3. User: A person who uses the system and checks alerts

[0360] Capture and transmit audio data

[0361] When a user initiates a call or voice message, the device captures audio data in real time through the microphone and transmits the captured audio data to the server using SSL / TLS encryption.

[0362] Examples:

[0363] The device collects the audio as soon as the user starts a call on their smartphone and sends it to the server as encrypted data.

[0364] Audio data preprocessing

[0365] The server performs pre-processing on the received audio data, including removing background noise using a noise attenuation algorithm (e.g., Wiener filter) and normalizing the volume of the audio data.

[0366] Examples:

[0367] The server applies a Wiener filter to the received audio data to reduce background noise and normalize the audio level to 0.5.

[0368] Acoustic feature extraction

[0369] From the preprocessed audio data, the server extracts Mel-Frequency Cepstral Coefficients (MFCCs) and spectral features using the Python library Librosa.

[0370] Examples:

[0371] The server extracts 13-dimensional MFCC features from the preprocessed audio data using Librosa and stores them in memory as an array.

[0372] Comparison with audio profiles

[0373] The server compares the extracted acoustic features with the voice profile registered by the user in advance, and detects fraudulent voices by calculating cosine similarity and Euclidean distance using SciPy to evaluate the degree of match.

[0374] Examples:

[0375] The server passes the generated MFCC features to the SciPy cosine similarity function to compare the similarity with pre-registered voice profiles. It also calculates the Euclidean distance to evaluate the match from different angles.

[0376] Anomaly detection and alerting

[0377] If the comparison detects unnatural voice changes or signs of synthetic speech, the server detects the anomaly and generates an alert. The alert is sent via email using the SMTP protocol and also via push notifications within the app.

[0378] Examples:

[0379] If the cosine similarity is below a certain level, the server determines that an abnormality has occurred and sends a warning message to the user's email address via SMTP. At the same time, a warning message is also sent to the user's smartphone via push notification.

[0380] Additional analysis with emotion recognition

[0381] The server uses a voice emotion recognition engine (e.g., Google Cloud Speech-to-Text API) to analyze the user's emotional state from the voice data and determine whether it matches their normal emotional state.

[0382] Examples:

[0383] The server uses the Google Cloud Speech-to-Text API to perform emotion recognition and detect emotions such as "anger," "sadness," and "joy." The detection results are saved in tensor format.

[0384] Selecting notification methods based on emotional state

[0385] The server selects the optimal notification method based on the detected emotional state. For example, if the user is determined to be in an "anxious" state, it will send an SMS notification immediately using the Twilio API.

[0386] Examples:

[0387] If the server detects an "uneasy" state, it will use the Twilio API to send an SMS notification with the message "urgent action required."

[0388] Saving logs

[0389] All processing results and generated alert information are stored as logs in a PostgreSQL database by the server, which can be used for later analysis and auditing, increasing the transparency and security of the system.

[0390] Examples:

[0391] The server inserts all data, including the results of speech preprocessing, acoustic features, emotion recognition results, and alert information, into a PostgreSQL database and stores it along with a timestamp.

[0392] Prompt Sentence Examples

[0393] As an example of a prompt for this system, the following query is input to the generative AI model:

[0394] "When a user sends a voice message, we want to identify the emotion from that data and activate fraudulent voice detection. Please suggest an algorithm."

[0395] The flow of the identification process in the second embodiment will be described with reference to FIG.

[0396] Step 1: Capture audio data

[0397] When a user initiates a call or voice message, the device uses the built-in microphone to capture voice data in real time, which is then recorded in WAV format and temporarily stored in the device's memory.

[0398] Specific behavior:

[0399] The moment a user starts a call on their smartphone, the device captures the audio data at a sampling rate of 44.1 kHz and saves it as a temporary file.

[0400] input:

[0401] Call start triggers

[0402] output:

[0403] Audio data captured by the built-in microphone (WAV format)

[0404] Step 2: Sending audio data

[0405] The device protects the captured audio data with AES-256 encryption and sends it to the server using SSL / TLS, preventing data leakage during transmission.

[0406] Specific behavior:

[0407] The device encrypts the voice data with AES-256 and transmits it in real time to the server via the HTTPS protocol.

[0408] input:

[0409] Captured audio data (WAV format)

[0410] output:

[0411] The encrypted audio data is sent to the server

[0412] Step 3: Preprocessing the audio data

[0413] The server applies a noise attenuation algorithm (e.g., a Wiener filter) to the received audio data to normalize the audio level.

[0414] Specific behavior:

[0415] The server applies a Wiener filter to the received audio data to reduce noise and normalize the volume level of the audio data to 0.5.

[0416] input:

[0417] Encrypted and decrypted audio data

[0418] output:

[0419] Noise-reduced and normalized audio data

[0420] Step 4: Extraction of acoustic features

[0421] The server extracts Mel-Frequency Cepstral Coefficients (MFCCs) and spectral features from the preprocessed audio data using the Librosa library.

[0422] Specific behavior:

[0423] The server uses Librosa to calculate 13-dimensional MFCC features from the preprocessed audio data and saves them as an array.

[0424] input:

[0425] Normalized audio data

[0426] output:

[0427] Extracted acoustic features (MFCC)

[0428] Step 5: Compare with the audio profile

[0429] The server uses SciPy to calculate the cosine similarity and Euclidean distance between the extracted acoustic features and the voice profile previously registered by the user, and evaluates the degree of match.

[0430] Specific behavior:

[0431] The server passes the generated MFCC features to SciPy's cosine similarity function to compare them with existing audio profiles and calculates the degree of match. It also calculates the Euclidean distance.

[0432] input:

[0433] Extracted acoustic features

[0434] Existing Audio Profiles

[0435] output:

[0436] Evaluation results of match (cosine similarity and Euclidean distance)

[0437] Step 6: Detect anomalies and generate alerts

[0438] The server determines whether there are any invalid voices or abnormalities based on the results of the evaluation of the degree of matching of the voice features, and generates an alert if necessary. The generated alert is sent via email using the SMTP protocol, and also sent in-app via push notification.

[0439] Specific behavior:

[0440] The server detects anomalies based on criteria such as a cosine similarity of 0.7 or less, and sends a warning message to the user's email address via SMTP. At the same time, a warning message is also sent to the user's smartphone via push notification.

[0441] input:

[0442] Matching evaluation results

[0443] output:

[0444] An alert is generated and an email and push notification is sent

[0445] Step 7: Further analysis with emotion recognition

[0446] The server analyzes the user's emotional state from the voice data using an emotion recognition engine (e.g., Google Cloud Speech-to-Text API), and the analyzed emotional state is stored in a database.

[0447] Specific behavior:

[0448] The server sends the voice data to the Google Cloud Speech-to-Text API, which performs emotion recognition and detects emotions such as "anger," "sadness," and "joy."

[0449] input:

[0450] Normalized audio data

[0451] output:

[0452] Detected emotional state

[0453] Step 8: Select notification method based on emotional state

[0454] The server selects the optimal notification method based on the detection result. For example, if the user is determined to be in an "anxious state," it will immediately send an SMS notification using the Twilio API.

[0455] Specific behavior:

[0456] If the emotion recognition result is "anxiety," the server uses the Twilio API to quickly send an SMS containing the message "urgent action required."

[0457] input:

[0458] Detected emotional state

[0459] output:

[0460] Messages sent via the most appropriate notification method

[0461] Step 9: Save the logs

[0462] The server logs all processing results and generated alert information in a PostgreSQL database for later analysis and auditing.

[0463] Specific behavior:

[0464] The server stores all data, including audio preprocessing results, acoustic features, emotion recognition results, and alert information, in a PostgreSQL database and records them with timestamps.

[0465] input:

[0466] Various processing results and alert information

[0467] output:

[0468] Logs stored in the database

[0469] (Application example 2)

[0470] Next, a description will be given of Application Example 2. In the following description, the data processing device 12 will be referred to as a "server" and the smart device 14 will be referred to as a "terminal."

[0471] Conventional voice recognition systems have difficulty detecting fraudulent voice data generated by voice generation AI in real time, making it impossible to ensure user safety. Furthermore, they lack the ability to recognize the user's emotional state, making it difficult to respond appropriately based on the user's psychological state. Therefore, there is a demand for a system that combines voice data anomaly detection and user emotion recognition.

[0472] The identification process by the identification processing unit 290 of the data processing device 12 in Application Example 2 is realized by the following means. In this invention, the server includes means for capturing voice data, means for preprocessing the captured voice data, means for comparing extracted acoustic features with existing voice profiles, means for generating and notifying an alert when an abnormality is detected, means for analyzing the user's emotional state from the voice data, and means for selecting a notification method for the alert based on the analysis results. This makes it possible not only to detect abnormalities in voice data in real time, but also to provide an appropriate notification according to the user's emotional state.

[0473] "Voice data" refers to data that records a user's voice information in digital format.

[0474] "Capturing means" refers to equipment or technology used to record or capture audio data.

[0475] "Preprocessing means" refers to the techniques and processes that perform preprocessing such as noise removal and normalization on the captured audio data.

[0476] "Acoustic features" are acoustic feature information extracted from speech data, and include Mel-Frequency Cepstrum Coefficients (MFCCs).

[0477] A "profile" is a database that records information about specific users and their voice characteristics.

[0478] "Means of comparison" refers to the technology or algorithm that matches the extracted acoustic features with existing voice profiles.

[0479] "Means for detecting anomalies" refers to technologies and algorithms that analyze the results of comparing acoustic features to identify fraudulent audio data.

[0480] "Means for generating and notifying alerts" refers to the technology and methods for generating a warning message and notifying the user when an abnormality is detected.

[0481] "Means for analyzing emotional state" refers to technologies and algorithms for recognizing a user's emotions from voice data and assessing their state.

[0482] "Means for selecting a notification method" refers to a technology or method for selecting the most appropriate alert notification method (e.g., SMS, email, etc.) based on the analyzed emotional state.

[0483] MODE FOR CARRYING OUT THE INVENTION

[0484] This invention relates to a system that detects fraudulent voice data generated by a voice generation AI in real time and analyzes the user's emotional state from the voice data to ensure the safety and comfort of the user. Specific embodiments of the system are described below.

[0485] System Configuration

[0486] The system consists of the following major components:

[0487] 1. Terminal

[0488] The terminal is a device that captures the user's voice data. This includes smartphones, tablets, PCs, etc. When the user makes a call or sends a voice message, the voice data is captured in real time and sent to the server via a secure communication method.

[0489] 2. Server

[0490] The server is a unit that processes, analyzes, and recognizes emotions from received voice data. It is mainly responsible for the following tasks:

[0491] Audio data preprocessing

[0492] Acoustic feature extraction

[0493] Comparison with audio profiles

[0494] Anomaly detection and alerting

[0495] emotion recognition

[0496] Selecting notification methods based on emotional state

[0497] Saving logs

[0498] 3. Users

[0499] The user is the person who interacts with the system and checks the alerts: they receive alerts when an anomaly is detected or when a particular emotional state is recognized.

[0500] Specific processing of the program

[0501] This system performs a series of processes in real time, from capturing voice data to analyzing it and generating alerts. The specific process is explained below.

[0502] 1. Capture and transmit audio data

[0503] The device captures voice data in real time when the user initiates a call or voice message, and the captured voice data is sent to the server via encrypted communication means.

[0504] 2. Preprocessing of audio data

[0505] The server performs preprocessing on the received audio data, including noise reduction to reduce background noise and normalization to adjust the audio level to a certain standard. This processing is performed using the Librosa library.

[0506] 3. Acoustic feature extraction

[0507] From the preprocessed audio data, the server extracts acoustic features, including Mel-Frequency Cepstral Coefficients (MFCCs) and spectral features, also using the Librosa library.

[0508] 4. Comparison with audio profiles

[0509] The server compares the extracted acoustic features with existing voice profiles using statistical methods such as cosine similarity and Euclidean distance, and generates an alert if an anomaly is detected.

[0510] 5. Anomaly detection and alert generation

[0511] If the comparison detects unnatural voice changes or signs of synthesis, the server generates an alert, which is promptly sent to the user via SMS, email, or in-app notification.

[0512] 6. Additional analysis using emotion recognition

[0513] The server uses an emotion recognition engine to analyze the user's emotional state from the voice data. Based on the analysis results, it determines whether the voice data matches the user's normal emotional state. DeepMoji and other emotion recognition APIs are used for emotion recognition.

[0514] 7. Notification method selection based on emotional state

[0515] The server selects the optimal notification method based on the detected emotional state. For example, if the user is determined to be in an anxious state, an immediate notification via SMS can be sent to enhance safety.

[0516] Specific examples

[0517] Specific application examples of this system are shown below.

[0518] Example 1:

[0519] If fraudulent voice generation occurs while a user is making an important call at work, the server will detect the anomaly and send a warning notification to the user's smartphone.

[0520] Example 2:

[0521] If a user feels stressed while talking to a family member, the server will recognize the emotion and send a notification encouraging the user to "relax."

[0522] Example prompts for generative AI models

[0523] "What is the execution flow of the application notification when fraudulent voice is detected while a user is on a call with a bank's call center?"

[0524] The flow of the specific processing in the application example 2 will be described with reference to FIG.

[0525] Program processing steps

[0526] Step 1: Capture and send audio data

[0527] Input: A user initiates a call or voice message.

[0528] Specific operation: The device uses the built-in microphone to capture audio data in real time.

[0529] Data processing: Captured audio data is encrypted immediately.

[0530] Output: Encrypted audio data is generated.

[0531] Processing flow: The device sends encrypted audio data to the server via a secure communication method.

[0532] Step 2: Preprocessing the audio data

[0533] Input: Encrypted audio data is sent to the server.

[0534] Specific operation: The server decodes the received audio data.

[0535] Data processing: Denoising and normalization are performed on the decoded audio data using the Librosa library.

[0536] Output: Preprocessed audio data is generated.

[0537] Processing flow: The server sends the noise-removed and normalized audio data to the next step.

[0538] Step 3: Extraction of acoustic features

[0539] Input: Preprocessed audio data is sent to the server.

[0540] Specific operation: The server uses the Librosa library to calculate Mel-Frequency Cepstral Coefficients (MFCCs).

[0541] Data processing: MFCC and spectral features are extracted from the audio data.

[0542] Output: The extracted acoustic features are generated.

[0543] Processing flow: The server sends the extracted features to the next step.

[0544] Step 4: Compare with the audio profile

[0545] Input: The extracted acoustic features are sent to the server.

[0546] Specific behavior: The server loads the existing audio profile.

[0547] Data analysis: Compare acoustic features and voice profiles using cosine similarity and Euclidean distance.

[0548] Output: The comparison results are generated.

[0549] Process flow: The server sends the comparison result to the next step.

[0550] Step 5: Detect anomalies and generate alerts

[0551] Input: The comparison result is sent to the server.

[0552] Specific operation: The server analyzes the comparison results and determines whether an anomaly is detected.

[0553] Data judgment: If an abnormality is detected, an alert message is generated.

[0554] Output: An alert message is generated.

[0555] Process flow: The server sends an alert message to the user's terminal.

[0556] Step 6: Further analysis with emotion recognition

[0557] Input: Preprocessed audio data is sent to the server.

[0558] Specific operation: The server analyzes the voice data using an emotion recognition engine.

[0559] Data analysis: Implementing specific algorithms to identify the user's emotional state from the audio data.

[0560] Output: The emotion recognition results are generated.

[0561] Processing flow: The server refers to the emotion recognition results and sends them to the next step.

[0562] Step 7: Select notification method based on emotional state

[0563] Input: Emotion recognition results are sent to the server.

[0564] Specific operation: The server selects the optimal notification method based on the emotional state.

[0565] Data judgment: Apply the conditions to select the notification method (e.g., SMS notification if the emotional state is stressed).

[0566] Output: The best notification method is selected.

[0567] Process flow: The server notifies the user of the alert according to the selected notification method.

[0568] The specific processing unit 290 transmits the result of the specific processing to the smart device 14. In the smart device 14, the control unit 46A causes the output device 40 to output the result of the specific processing. The microphone 38B acquires audio indicating a user input regarding the result of the specific processing. The control unit 46A transmits audio data indicating the user input acquired by the microphone 38B to the data processing device 12. In the data processing device 12, the specific processing unit 290 acquires the audio data.

[0569] The data generation model 58 is a so-called generative AI (Artificial Intelligence). An example of the data generation model 58 is ChatGPT (registered trademark) (Internet search engine).<URL: https: / / openai.com / blog / chatgpt> ), Gemini (registered trademark) (Internet search <url: https: gemini.google.com ?hl="ja">) and other generation AIs. The data generation model 58 is obtained by performing deep learning on a neural network. A prompt including an instruction is input to the data generation model 58, and inference data such as voice data indicating voice, text data indicating text, and image data indicating an image is also input. The data generation model 58 performs inference on the input inference data in accordance with the instruction indicated by the prompt, and outputs the inference result in a data format such as voice data and text data. Here, inference refers to, for example, analysis, classification, prediction, and / or summarization.

[0570] In the above embodiment, an example in which the specific process is performed by the data processing device 12 has been given, but the technology of the present disclosure is not limited to this, and the specific process may be performed by the smart device 14.

[0571] [Second embodiment]

[0572] FIG. 3 shows an example of the configuration of a data processing system 210 according to the second embodiment.

[0573] 3, the data processing system 210 includes the data processing device 12 and smart glasses 214. An example of the data processing device 12 is a server.

[0574] The data processing device 12 includes a computer 22, a database 24, and a communication I / F 26. The computer 22 is an example of a "computer" according to the technology of the present disclosure. The computer 22 includes a processor 28, a RAM 30, and a storage 32. The processor 28, the RAM 30, and the storage 32 are connected to a bus 34. The database 24 and the communication I / F 26 are also connected to the bus 34. The communication I / F 26 is connected to a network 54. Examples of the network 54 include a WAN (Wide Area Network) and / or a LAN (Local Area Network).

[0575] The smart glasses 214 include a computer 36, a microphone 238, a speaker 240, a camera 42, and a communication I / F 44. The computer 36 includes a processor 46, a RAM 48, and a storage 50. The processor 46, the RAM 48, and the storage 50 are connected to a bus 52. The microphone 238, the speaker 240, and the camera 42 are also connected to the bus 52.

[0576] The microphone 238 receives instructions and the like from the user 20 by receiving voice uttered by the user 20. The microphone 238 captures the voice uttered by the user 20, converts the captured voice into audio data, and outputs it to the processor 46. The speaker 240 outputs audio in accordance with instructions from the processor 46.

[0577] Camera 42 is a small digital camera equipped with an optical system including a lens, aperture, and shutter, and an imaging element such as a CMOS (Complementary Metal-Oxide-Semiconductor) image sensor or a CCD (Charge Coupled Device) image sensor, and captures images of the surroundings of user 20 (for example, an imaging range defined by an angle of view equivalent to the field of vision of a typical healthy person).

[0578] The communication I / F 44 is connected to a network 54. The communication I / Fs 44 and 26 control the exchange of various information between the processor 46 and the processor 28 via the network 54. The exchange of various information between the processor 46 and the processor 28 using the communication I / Fs 44 and 26 is carried out in a secure state.

[0579] Fig. 4 shows an example of the main functions of the data processing device 12 and the smart glasses 214. As shown in Fig. 4, in the data processing device 12, a specific process is performed by the processor 28. A specific process program 56 is stored in the storage 32.

[0580] The specific processing program 56 is an example of a "program" according to the technology of the present disclosure. The processor 28 reads the specific processing program 56 from the storage 32 and executes the read specific processing program 56 on the RAM 30. The specific processing is realized by the processor 28 operating as a specific processing unit 290 in accordance with the specific processing program 56 executed on the RAM 30.

[0581] The storage 32 stores a data generation model 58 and an emotion identification model 59. The data generation model 58 and the emotion identification model 59 are used by the identification processing unit 290.

[0582] In the smart glasses 214, the reception output process is performed by the processor 46. A reception output program 60 is stored in the storage 50. The processor 46 reads the reception output program 60 from the storage 50 and executes the read reception output program 60 on the RAM 48. The reception output process is realized by the processor 46 operating as the control unit 46A in accordance with the reception output program 60 executed on the RAM 48.

[0583] Next, a description will be given of the identification process performed by the identification processing unit 290 of the data processing device 12. In the following description, the data processing device 12 will be referred to as the "server" and the smart glasses 214 will be referred to as the "terminal."

[0584] This invention relates to a system for detecting fraudulent voice data generated by a voice generation AI in real time and ensuring the safety of users. Specific embodiments of the system are described below.

[0585] System Configuration

[0586] The system consists of the following major components:

[0587] 1. Device: The device that captures the audio data (e.g., a smartphone or PC)

[0588] 2. Server: The unit that processes and analyzes the captured audio data.

[0589] 3. User: A person who uses the system and checks alerts

[0590] Specific processing of the program

[0591] Capture and transmit audio data

[0592] When a device starts a voice call or message transmission, it immediately captures the voice data, which is then sent to the server in real time.

[0593] Examples:

[0594] When the user starts a call, the terminal automatically starts monitoring the voice data and sends that data to the server.

[0595] Audio data preprocessing

[0596] The server performs the necessary pre-processing on the received audio data, which includes the following steps:

[0597] Noise Reduction: Reduce background noise from audio data.

[0598] Normalization: Adjusting audio levels to a certain standard.

[0599] Acoustic feature extraction

[0600] The server extracts acoustic features from the preprocessed speech data, including Mel-Frequency Cepstral Coefficients (MFCCs) and spectral features.

[0601] Examples:

[0602] The server calculates MFCCs from the noise-removed and normalized audio data and extracts them as features.

[0603] Comparison with audio profiles

[0604] The server compares the extracted acoustic features with existing voice profiles using statistical methods such as cosine similarity and Euclidean distance.

[0605] Examples:

[0606] The server compares the voice samples previously registered by the user and calculates how well the current voice data matches.

[0607] Anomaly detection and alerting

[0608] If the comparison detects any abnormal changes or signs of unnatural synthesis, the server generates an alert, which is immediately sent to the user via SMS, email, or in-app notification.

[0609] Examples:

[0610] The server detects unnatural voice changes and sends a warning message to the user's smartphone.

[0611] Saving logs

[0612] All processing results and alert information are stored as logs by the server and used for later analysis and auditing, which increases the transparency and reliability of the system.

[0613] In this way, the system analyzes the user's voice data in real time, detects fraudulent voice production with high accuracy, and protects the user's safety.

[0614] The processing flow will be explained below.

[0615] Step 1:

[0616] The device captures voice data in real time when the user initiates a call or voice message.

[0617] Save the audio data in a specified format (e.g. WAV, MP3).

[0618] Step 2:

[0619] The device transmits the captured audio data to the server.

[0620] Encrypted communications are used for data transmission to ensure data security.

[0621] Step 3:

[0622] The server stores the received audio data and starts pre-processing.

[0623] Noise Reduction: Performs a filtering process to reduce background noise.

[0624] Normalize: Adjust the audio level to a certain standard.

[0625] Step 4:

[0626] The server extracts acoustic features from the preprocessed speech data.

[0627] Calculates Mel-Frequency Cepstral Coefficients (MFCC).

[0628] Analyze the spectral features.

[0629] Step 5:

[0630] The server compares the extracted acoustic features with existing voice profiles.

[0631] Statistical methods such as cosine similarity or Euclidean distance are used.

[0632] Step 6:

[0633] The server analyzes the comparison results to detect any unnatural voice changes or signs of synthesis.

[0634] Apply anomaly detection algorithms and flag any anomalies.

[0635] Step 7:

[0636] The server generates an alert if an anomaly is detected.

[0637] The alert content includes a warning message and details of the abnormality detection.

[0638] Step 8:

[0639] The server notifies the user of the generated alert.

[0640] Notification methods include SMS, email, or in-app notifications.

[0641] Step 9:

[0642] The user receives and confirms the notification.

[0643] If necessary, take appropriate action (e.g., report, suspend use of services).

[0644] Step 10:

[0645] The server stores all processing results and alert information as logs.

[0646] The saved logs can be used for later analysis and auditing.

[0647] The above is the specific processing flow of this system.

[0648] Example 1

[0649] Next, a description will be given of Example 1. In the following description, the data processing device 12 will be referred to as a "server" and the smart glasses 214 will be referred to as a "terminal."

[0650] In current voice communication systems, the generation of fraudulent voice data using generative AI models is a problem, threatening user safety. Effective methods for detecting such fraudulent voice generation in real time and immediately notifying users are needed. However, existing systems lack sufficient reliability and speed because they are unable to integrate multiple processes, such as voice data capture, preprocessing, anomaly detection, alert generation, and log storage.

[0651] The specific processing by the specific processing unit 290 of the data processing device 12 in the first embodiment is realized by the following means.

[0652] In this invention, the server includes means for capturing voice data, means for preprocessing the captured voice data, means for extracting acoustic features from the preprocessed voice data, means for comparing the extracted acoustic features with an existing voice profile, means for analyzing the comparison result to detect anomalies, means for generating and notifying an alert when an anomaly is detected, and means for storing the processing results and alert information in a database. This enables highly accurate real-time detection and immediate notification of improper voice generation, and recording of all processing results.

[0653] "Audio data" refers to digital or analog data that records sound.

[0654] "Capture" is the process of acquiring audio data.

[0655] "Preprocessing" refers to processing performed to make captured audio data easier to analyze, and includes operations such as noise removal and normalization.

[0656] "Acoustic features" are numerical representations of the characteristics of speech data, and include Mel-Frequency Cepstrum Coefficients (MFCCs) and spectral features.

[0657] An "audio profile" is reference data constructed using features of pre-registered audio data.

[0658] "Comparison" is a process of evaluating the extracted acoustic features and the features of the voice profile using statistical methods.

[0659] "Anomaly detection" is the process of analyzing the comparison results to find incorrect speech production or unusual changes.

[0660] An "alert" is a warning or notification that is generated when an abnormality is detected.

[0661] "Notification" is the process of communicating generated alerts to users, including by means of SMS, email, in-app notifications, etc.

[0662] A "database" is a system for storing processing results and alert information.

[0663] "Storage" is the process of recording processing results and alert information in a database.

[0664] This invention relates to a system that detects fraudulent voice data generation by voice generation AI in real time and ensures user safety.

[0665] System Configuration

[0666] The system consists of the following major components:

[0667] 1. Device: The device that captures the audio data (e.g., a smartphone or PC)

[0668] 2. Server: The unit that processes and analyzes the captured audio data.

[0669] 3. User: A person who uses the system and checks alerts

[0670] Specific processing of the program

[0671] Capture and transmit audio data

[0672] When a user initiates a voice call or sends a message, the device captures the voice data using the built-in microphone and transmits it to the server in real time. The hardware used is the microphone of a smartphone or PC, and the software is a dedicated voice capture application (e.g., a homemade VoiceCaptureApp).

[0673] Specific behavior:

[0674] The device automatically starts capturing audio when the user starts a call and sends the data to a remote server using SSL / TLS encryption, ensuring data security.

[0675] Audio data preprocessing

[0676] The server performs preprocessing on the received audio data, which mainly includes noise removal and voice level normalization. Specifically, it applies a filter to the audio data to remove background noise, and then normalizes the voice level.

[0677] Specific behavior:

[0678] The server uses Python's SciPy library to apply a bandpass filter to remove background noise.

[0679] After noise removal, the audio levels are normalized to the range 0 to 1 using the normalize function from the Librosa library.

[0680] Acoustic feature extraction

[0681] The server extracts acoustic features from the preprocessed speech data, including Mel-Frequency Cepstral Coefficients (MFCCs) and spectral features.

[0682] Specific behavior:

[0683] The server extracts MFCCs from the audio data using the mfcc function in the Librosa library and stores them in a database as acoustic features.

[0684] Comparison with audio profiles

[0685] The server compares the extracted acoustic features with pre-trained voice profiles using cosine similarity or Euclidean distance.

[0686] Specific behavior:

[0687] The server uses the cosine function in the SciPy library to calculate the similarity between the features of the current voice data and the features of the pre-registered voice profile.

[0688] If the calculation results below a certain threshold, it is flagged as a likely incorrect speech production.

[0689] Anomaly detection and alerting

[0690] If the server detects an anomaly in the audio data, it generates an alert and notifies the user immediately via SMS, email, or in-app notification.

[0691] Specific behavior:

[0692] The server uses an anomaly detection algorithm (e.g., Scikit-learn's Isolation Forest) to detect unnatural changes in the audio data.

[0693] If an anomaly is detected, a warning message is sent to the user's smartphone using Twilio's SMS API, and email notifications are sent using the SendGrid API.

[0694] Saving logs

[0695] The server stores all processing results and alert information in a database, allowing for later analysis and auditing, improving system reliability.

[0696] Specific behavior:

[0697] The server stores the processing results and alert information in a MySQL database using SQLAlchemy to create database entries.

[0698] Example prompts to input to the generative AI model

[0699] "Please explain how speech generation AI detects fraudulent audio data. Please include specific steps for capturing audio data, preprocessing, extracting acoustic features, comparing with audio profiles, detecting anomalies, generating alerts, and saving logs."

[0700] In this way, the system can analyze the user's voice data in real time, detect fraudulent voice production with high accuracy, and ensure the user's safety.

[0701] The flow of the identification process in the first embodiment will be described with reference to FIG.

[0702] Processing Steps

[0703] Step 1: Capture and send audio data

[0704] The device uses a built-in microphone to capture voice data when a user initiates a voice call or sends a message.

[0705] input:

[0706] Initiate voice calls and send messages

[0707] Specific behavior:

[0708] The device starts a dedicated voice capture application (e.g., a custom VoiceCaptureApp) and starts microphone input. The captured voice data is sent to the server using encryption (SSL / TLS).

[0709] output:

[0710] Encrypted and transmitted voice data

[0711] Step 2: Preprocessing the audio data

[0712] The server performs preprocessing on the received audio data.

[0713] input:

[0714] Received audio data

[0715] Specific behavior:

[0716] The server uses Python's SciPy library to apply a bandpass filter to remove background noise, then normalizes the audio level to the range 0 to 1 using the normalize function from the Librosa library.

[0717] output:

[0718] Denoised and normalized audio data

[0719] Step 3: Extraction of acoustic features

[0720] The server extracts acoustic features from the preprocessed speech data.

[0721] input:

[0722] Denoised and normalized audio data

[0723] Specific behavior:

[0724] The server extracts MFCCs using the mfcc function in the Librosa library and stores them in a database as acoustic features.

[0725] output:

[0726] Extracted acoustic features

[0727] Step 4: Compare with the audio profile

[0728] The server compares the extracted acoustic features with pre-registered voice profiles.

[0729] input:

[0730] Extracted acoustic features

[0731] Pre-registered audio profiles

[0732] Specific behavior:

[0733] The server uses the cosine function in the SciPy library to calculate the similarity between the features of the current voice data and the features of the existing voice profile.

[0734] output:

[0735] Similarity calculation results

[0736] Step 5: Detect anomalies and generate alerts

[0737] If the server detects an abnormality in the voice data, it generates an alert and immediately notifies the user.

[0738] input:

[0739] Similarity calculation results

[0740] Specific behavior:

[0741] The server uses an anomaly detection algorithm (e.g., Scikit-learn's Isolation Forest) to detect anomalies. If an anomaly is detected, it uses Twilio's SMS API to send an alert message to the user's smartphone. For email notifications, it uses the SendGrid API.

[0742] output:

[0743] Warning message

[0744] Step 6: Save the logs

[0745] The server stores all processing results and alert information in a database.

[0746] input:

[0747] Similarity calculation results

[0748] Warning message

[0749] Specific behavior:

[0750] The server uses SQLAlchemy to create and save database entries to store the processing results and alert information in a MySQL database.

[0751] output:

[0752] Saved processing results and alert information

[0753] (Application example 1)

[0754] Next, a description will be given of Application Example 1. In the following description, the data processing device 12 will be referred to as a "server" and the smart glasses 214 will be referred to as a "terminal."

[0755] With the advancement of AI voice generation technology, it has become easier to generate voice data fraudulently. As a result, the risk of personal and corporate voice communications being used fraudulently and important call content being tampered with is increasing. This has a particularly severe impact on business communications that contain confidential information. Current systems face the challenge of being unable to detect such fraudulent activity in real time and respond immediately.

[0756] The specific processing by the specific processing unit 290 of the data processing device 12 in the application example 1 is realized by the following means.

[0757] In this invention, the server includes means for capturing voice data, means for preprocessing the captured voice data, means for extracting acoustic features from the preprocessed voice data, means for comparing the extracted acoustic features with a voice profile, means for detecting anomalies based on the analysis results, means for generating and notifying an alarm when an anomaly is detected, means for recording the alarm information, means for detecting fraudulent generation of voice data in real time, an application for monitoring call data, and a cloud-based processing server. This makes it possible to detect fraudulent generation of voice data in real time and respond quickly.

[0758] "Audio Data" means data for recording, storing, and transmitting audio in digital format.

[0759] A "capturing means" is a device or software for capturing audio data in real time.

[0760] The "pre-processing means" refers to a method or device for removing noise from the captured audio data and normalizing it.

[0761] "Means for extracting acoustic features" refers to a method or device that calculates and extracts characteristic information (e.g., Mel-Frequency Cepstral Coefficients (MFCCs)) from preprocessed speech data.

[0762] An "audio profile" is a collection of feature data generated based on specific audio data.

[0763] The "comparison means" refers to a method or device for comparing the extracted acoustic features with a voice profile and evaluating the degree of similarity.

[0764] The "analysis result" is information indicating the degree of agreement of the data obtained from the comparison of the acoustic features and the presence or absence of abnormalities.

[0765] "Means for detecting anomalies" refers to a method or device that determines whether there are unnatural changes or signs of fraudulent generation in the audio data based on the analysis results.

[0766] The "means for generating and notifying a warning" refers to a method or device for generating and notifying a warning message to a user when an abnormality is detected.

[0767] The "means for recording warning information" refers to a method or device for saving generated warning messages and analysis results so that they can be checked at a later date.

[0768] "Real-time detection means" refers to a method or device that can detect fraud immediately without interrupting the time when voice data is being generated.

[0769] An "application that monitors call data" is software that runs on a device such as a smartphone and monitors voice data during a call.

[0770] A "cloud-based processing server" is a remote server accessible via the Internet that pre-processes, analyzes, and compares audio data.

[0771] This invention relates to a system that detects fraudulent voice data generated by voice generation AI in real time during phone calls to ensure user security. This system automatically captures voice data, preprocesses it, extracts acoustic features, compares it with existing voice profiles, detects anomalies, generates warnings, and records it.

[0772] System Configuration

[0773] The system consists of the following major components:

[0774] 1. Terminal: A device for capturing audio data (e.g., a smartphone)

[0775] 2. Server: A unit for processing and analyzing the captured audio data. Uses a cloud-based processing server (e.g., AWS, Google Cloud Platform).

[0776] 3. Users: People who use the system and receive alerts

[0777] Specific processing of the program

[0778] Capture audio data

[0779] The device uses the Twilio Voice SDK to capture voice data during a call in real time and immediately transmits that data to the server.

[0780] Audio data preprocessing

[0781] The server uses libraries (e.g., PyDub, Librosa) to denoise and normalize the audio data. This preprocessing produces clean audio data suitable for analysis.

[0782] Acoustic feature extraction

[0783] From the preprocessed audio data, the server uses Librosa to extract acoustic features such as Mel-Frequency Cepstral Coefficients (MFCCs).

[0784] Comparison with audio profiles

[0785] The extracted acoustic features are compared to existing voice profiles using machine learning frameworks such as TensorFlow, using statistical methods such as cosine similarity.

[0786] Anomaly detection and alert generation

[0787] Based on the analysis results, the server determines whether the voice data has any unnatural variations or signs of fraudulent generation, and if an anomaly is detected, the server sends a warning to the user via SMS or in-app notification.

[0788] Recording warning information

[0789] All processing results and warning information are logged by the server for later analysis and auditing.

[0790] Specific examples

[0791] For example, suppose a user is using the Secure Call Detector app to make a call for a business meeting and the call content is tampered with. The system detects this tampering in real time and immediately generates an alert. The user receives the following message on their smartphone:

[0792] "Warning: Indications of unauthorized audio have been detected. The following message will appear on your screen: 'Warning: Unauthorized audio has been detected. Please ensure your call is secure.' Call recording will automatically stop and your system administrator will be notified."

[0793] This allows the user to respond quickly and ensures the safety of the call.

[0794] The system of the present invention can detect fraudulent generation of voice data in real time and respond quickly, thereby enhancing the security of individuals and businesses.

[0795] The flow of the specific processing in the application example 1 will be described with reference to FIG.

[0796] Step 1:

[0797] The terminal uses the Twilio Voice SDK to detect when a user starts a call. When a call starts, it immediately starts capturing voice data and sends that data to the server in real time. The input is the user's call voice data, and the output is the raw voice data sent to the server.

[0798] Step 2:

[0799] The server preprocesses the audio data received from the device using a library (e.g., PyDub, Librosa). This preprocessing removes noise from the audio data and normalizes the audio level. The input is raw audio data, and the output is noise-removed and normalized audio data.

[0800] Step 3:

[0801] The server extracts acoustic features from the preprocessed speech data. It uses Librosa to extract features such as Mel-Frequency Cepstral Coefficients (MFCCs). The input is the preprocessed speech data, and the output is the extracted acoustic features.

[0802] Step 4:

[0803] The server compares the extracted acoustic features with existing voice profiles. This comparison uses TensorFlow's machine learning model to evaluate similarity using statistical methods such as cosine similarity and Euclidean distance. The input is the acoustic features and the voice profile, and the output is a similarity score.

[0804] Step 5:

[0805] The server analyzes the similarity score and determines whether or not unauthorized voice data has been generated. If an abnormality is detected based on the analysis results, a warning message is generated. The input is the similarity score, and the output is the warning message.

[0806] Step 6:

[0807] The server notifies the user via SMS or in-app notification with the generated alert message. This notification is sent using Twilio's messaging API. The input is the alert message, and the output is the alert message displayed on the user's device.

[0808] Step 7:

[0809] The server stores all processing results and warning information in a database. This information is recorded using a database such as MySQL for later analysis and auditing. The input is the processing results and warning information, and the output is the log information stored in the database.

[0810] Furthermore, an emotion engine that estimates the user's emotion may be further combined. That is, the identification processing unit 290 may estimate the user's emotion using the emotion identification model 59, and perform identification processing using the user's emotion.

[0811] This invention relates to a system that ensures the safety and comfort of users by detecting fraudulent voice data generated by voice generation AI in real time and recognizing the emotional state of the user from the voice data. Specific embodiments of the system are described below.

[0812] System Configuration

[0813] The system consists of the following major components:

[0814] 1. Device: The device that captures the audio data (e.g., a smartphone or PC)

[0815] 2. Server: The unit that processes, analyzes, and recognizes emotions from captured audio data.

[0816] 3. User: A person who uses the system and checks alerts

[0817] Specific processing of the program

[0818] Capture and transmit audio data

[0819] The device captures voice data in real time when the user initiates a call or voice message, and the captured voice data is sent to the server via encrypted communication means.

[0820] Examples:

[0821] The terminal starts monitoring the voice data at the same time as the user starts a call, and transmits the voice data to the server via secure communication.

[0822] Audio data preprocessing

[0823] The server performs pre-processing on the received audio data, including noise reduction to reduce background noise and normalization to adjust the audio level to a certain standard.

[0824] Acoustic feature extraction

[0825] The server extracts acoustic features from the preprocessed speech data, including Mel-Frequency Cepstral Coefficients (MFCCs) and spectral features.

[0826] Examples:

[0827] The server calculates MFCCs from the noise-removed and normalized audio data and obtains them as feature data.

[0828] Comparison with audio profiles

[0829] The server compares the extracted acoustic features with existing voice profiles using statistical methods such as cosine similarity and Euclidean distance.

[0830] Examples:

[0831] The server compares the voice samples previously registered by the user and calculates how well the current voice data matches.

[0832] Anomaly detection and alerting

[0833] If the comparison detects unnatural voice changes or signs of synthesis, the server generates an alert, which is promptly sent to the user via SMS, email, or in-app notification.

[0834] Examples:

[0835] The server detects abnormal audio changes and sends a warning message to the user's smartphone.

[0836] Additional analysis with emotion recognition

[0837] The server uses an emotion engine to analyze the user's emotional state from the voice data, and determines whether the voice data matches the user's typical emotional state.

[0838] Examples:

[0839] The server analyzes the voice data to determine whether the user is expressing emotions such as anxiety, anger, or joy.

[0840] Selecting notification methods based on emotional state

[0841] The server selects the optimal notification method based on the detected emotional state. For example, if the user is determined to be in an anxious state, an immediate notification via SMS can be sent to enhance safety.

[0842] Examples:

[0843] Based on the results of the emotion engine, if the user is feeling stressed, the server determines that a prompt response is required and selects an SMS notification.

[0844] Saving logs

[0845] All processing results and alert information are stored as logs by the server, which can be used for later analysis and auditing, ensuring the transparency and reliability of the system.

[0846] In this way, the system analyzes the user's voice data in real time, detects fraudulent voice production with high accuracy, and by recognizing the user's emotional state, selects an appropriate notification method to maintain the user's safety and comfort.

[0847] The processing flow will be explained below.

[0848] Step 1:

[0849] The device automatically captures voice data in real time when the user initiates a call or voice message.

[0850] Save the audio data in a specified format (e.g. WAV, MP3).

[0851] Step 2:

[0852] The device transmits the captured audio data to the server.

[0853] Encrypted communications are used for data transmission to ensure data security.

[0854] Step 3:

[0855] The server stores the received audio data and starts pre-processing.

[0856] The server first applies a noise reduction filter to reduce background noise.

[0857] The server then normalizes the audio levels to ensure consistency.

[0858] Step 4:

[0859] The server extracts acoustic features from the preprocessed speech data.

[0860] The server calculates Mel-Frequency Cepstral Coefficients (MFCCs) and extracts them as speech features.

[0861] The server also analyzes the spectral features and stores them as supplementary features.

[0862] Step 5:

[0863] The server compares the extracted acoustic features with existing voice profiles.

[0864] The server calculates cosine similarity and Euclidean distance to evaluate the degree of matching of the voices.

[0865] Step 6:

[0866] The server analyzes the comparison results to detect any unnatural voice changes or signs of synthesis.

[0867] The server applies an anomaly detection algorithm and flags any anomalies detected.

[0868] Step 7:

[0869] The server generates an alert if an anomaly is detected.

[0870] The server generates alert information including warning messages and details of anomaly detections.

[0871] Step 8:

[0872] The server notifies the user of the generated alert.

[0873] Notification methods include SMS, email, or in-app notifications.

[0874] Step 9:

[0875] The user receives and confirms the notification.

[0876] Based on the content of the notification, the user will take appropriate action as necessary (e.g., report the issue, stop using the service).

[0877] Step 10:

[0878] The server uses an emotion engine to analyze the user's emotional state from the voice data.

[0879] The server identifies the emotional state (e.g., anxiety, anger, joy, etc.) from the voice data.

[0880] Step 11:

[0881] The server selects the most appropriate notification method based on the detected emotional state.

[0882] For example, if a user is in a state of anxiety, they will be notified immediately via SMS to enhance safety.

[0883] Step 12:

[0884] The server stores all processing results and alert information as logs.

[0885] The saved logs can be used for later analysis and auditing.

[0886] This is the specific processing flow of this system. This system allows users to detect fraudulent generation of voice data with high accuracy, and by recognizing the user's emotional state, it enables more appropriate responses.

[0887] Example 2

[0888] Next, a description will be given of Example 2. In the following description, the data processing device 12 will be referred to as a "server" and the smart glasses 214 will be referred to as a "terminal."

[0889] Conventional voice data analysis systems have difficulty detecting fraudulent voice generation in real time, and there are no systems that can recognize the user's emotional state and take appropriate action. This increases the risk of voice data being manipulated or used fraudulently, and poses a challenge to ensuring user safety and comfort.

[0890] The identification process by the identification processing unit 290 of the data processing device 12 in the second embodiment is realized by the following means. In this invention, the server includes means for preprocessing voice data, means for extracting acoustic features from the preprocessed voice data, means for comparing the extracted acoustic features with an existing voice profile, and means for selecting a notification method based on the recognized emotional state. This enables highly accurate detection of fraudulent voice generation and appropriate response based on the user's emotional state.

[0891] "Audio data" is a digital representation of an acoustic signal captured using a microphone or other audio input device.

[0892] "Capture" is the process of capturing audio data in real time via a microphone or other input device.

[0893] "Preprocessing" refers to the process of applying early-stage processing such as noise removal and normalization to speech data to facilitate analysis and feature extraction.

[0894] "Encryption" is a method of converting audio data using a specific algorithm to protect it from third parties.

[0895] "Transmission" refers to the act of transferring captured audio data to another system or server via a network.

[0896] "Acoustic features" are computed physical or statistical properties, such as Mel-Frequency Cepstral Coefficients (MFCCs) or spectral features, extracted from speech data.

[0897] "Comparison" is a process in which the extracted acoustic features are matched with an existing data set called a voice profile, and the degree of match is evaluated.

[0898] "Anomaly detection" is a function that detects and reports unnatural voice changes and traces of synthesized voice as a result of voice data analysis.

[0899] An "alert" is a message or signal that notifies a user or system administrator when an abnormality is detected.

[0900] "Emotion recognition" is a technology that analyzes and determines a user's emotional state (e.g., joy, anxiety, anger, etc.) from voice data.

[0901] "Notification method selection" is the process of selecting the optimal method (e.g., SMS, email, in-app notification, etc.) to notify the user based on the recognized emotional state.

[0902] This invention relates to a system that ensures the safety and comfort of users by detecting fraudulent voice data generated by voice generation AI in real time and recognizing the emotional state of the user from the voice data. Specific embodiments are described below.

[0903] System Configuration

[0904] The system consists of the following major components:

[0905] 1. Device: The device that captures the audio data (e.g., a smartphone or computer)

[0906] 2. Server: The unit that processes, analyzes, and recognizes emotions from captured audio data.

[0907] 3. User: A person who uses the system and checks alerts

[0908] Capture and transmit audio data

[0909] When a user initiates a call or voice message, the device captures audio data in real time through the microphone and transmits the captured audio data to the server using SSL / TLS encryption.

[0910] Examples:

[0911] The device collects the audio as soon as the user starts a call on their smartphone and sends it to the server as encrypted data.

[0912] Audio data preprocessing

[0913] The server performs pre-processing on the received audio data, including removing background noise using a noise attenuation algorithm (e.g., Wiener filter) and normalizing the volume of the audio data.

[0914] Examples:

[0915] The server applies a Wiener filter to the received audio data to reduce background noise and normalize the audio level to 0.5.

[0916] Acoustic feature extraction

[0917] From the preprocessed audio data, the server extracts Mel-Frequency Cepstral Coefficients (MFCCs) and spectral features using the Python library Librosa.

[0918] Examples:

[0919] The server extracts 13-dimensional MFCC features from the preprocessed audio data using Librosa and stores them in memory as an array.

[0920] Comparison with audio profiles

[0921] The server compares the extracted acoustic features with the voice profile registered by the user in advance, and detects fraudulent voices by calculating cosine similarity and Euclidean distance using SciPy to evaluate the degree of match.

[0922] Examples:

[0923] The server passes the generated MFCC features to the SciPy cosine similarity function to compare the similarity with pre-registered voice profiles. It also calculates the Euclidean distance to evaluate the match from different angles.

[0924] Anomaly detection and alerting

[0925] If the comparison detects unnatural voice changes or signs of synthetic speech, the server detects the anomaly and generates an alert. The alert is sent via email using the SMTP protocol and also via push notifications within the app.

[0926] Examples:

[0927] If the cosine similarity is below a certain level, the server determines that an abnormality has occurred and sends a warning message to the user's email address via SMTP. At the same time, a warning message is also sent to the user's smartphone via push notification.

[0928] Additional analysis with emotion recognition

[0929] The server uses a voice emotion recognition engine (e.g., Google Cloud Speech-to-Text API) to analyze the user's emotional state from the voice data and determine whether it matches their normal emotional state.

[0930] Examples:

[0931] The server uses the Google Cloud Speech-to-Text API to perform emotion recognition and detect emotions such as "anger," "sadness," and "joy." The detection results are saved in tensor format.

[0932] Selecting notification methods based on emotional state

[0933] The server selects the optimal notification method based on the detected emotional state. For example, if the user is determined to be in an "anxious" state, it will send an SMS notification immediately using the Twilio API.

[0934] Examples:

[0935] If the server detects an "uneasy" state, it will use the Twilio API to send an SMS notification with the message "urgent action required."

[0936] Saving logs

[0937] All processing results and generated alert information are stored as logs in a PostgreSQL database by the server, which can be used for later analysis and auditing, increasing the transparency and security of the system.

[0938] Examples:

[0939] The server inserts all data, including the results of speech preprocessing, acoustic features, emotion recognition results, and alert information, into a PostgreSQL database and stores it along with a timestamp.

[0940] Prompt Sentence Examples

[0941] As an example of a prompt for this system, the following query is input to the generative AI model:

[0942] "When a user sends a voice message, we want to identify the emotion from that data and activate fraudulent voice detection. Please suggest an algorithm."

[0943] The flow of the identification process in the second embodiment will be described with reference to FIG.

[0944] Step 1: Capture audio data

[0945] When a user initiates a call or voice message, the device uses the built-in microphone to capture voice data in real time, which is then recorded in WAV format and temporarily stored in the device's memory.

[0946] Specific behavior:

[0947] The moment a user starts a call on their smartphone, the device captures the audio data at a sampling rate of 44.1 kHz and saves it as a temporary file.

[0948] input:

[0949] Call start triggers

[0950] output:

[0951] Audio data captured by the built-in microphone (WAV format)

[0952] Step 2: Sending audio data

[0953] The device protects the captured audio data with AES-256 encryption and sends it to the server using SSL / TLS, preventing data leakage during transmission.

[0954] Specific behavior:

[0955] The device encrypts the voice data with AES-256 and transmits it in real time to the server via the HTTPS protocol.

[0956] input:

[0957] Captured audio data (WAV format)

[0958] output:

[0959] The encrypted audio data is sent to the server

[0960] Step 3: Preprocessing the audio data

[0961] The server applies a noise attenuation algorithm (e.g., a Wiener filter) to the received audio data to normalize the audio level.

[0962] Specific behavior:

[0963] The server applies a Wiener filter to the received audio data to reduce noise and normalize the volume level of the audio data to 0.5.

[0964] input:

[0965] Encrypted and decrypted audio data

[0966] output:

[0967] Noise-reduced and normalized audio data

[0968] Step 4: Extraction of acoustic features

[0969] The server extracts Mel-Frequency Cepstral Coefficients (MFCCs) and spectral features from the preprocessed audio data using the Librosa library.

[0970] Specific behavior:

[0971] The server uses Librosa to calculate 13-dimensional MFCC features from the preprocessed audio data and saves them as an array.

[0972] input:

[0973] Normalized audio data

[0974] output:

[0975] Extracted acoustic features (MFCC)

[0976] Step 5: Compare with the audio profile

[0977] The server uses SciPy to calculate the cosine similarity and Euclidean distance between the extracted acoustic features and the voice profile previously registered by the user, and evaluates the degree of match.

[0978] Specific behavior:

[0979] The server passes the generated MFCC features to SciPy's cosine similarity function to compare them with existing audio profiles and calculates the degree of match. It also calculates the Euclidean distance.

[0980] input:

[0981] Extracted acoustic features

[0982] Existing Audio Profiles

[0983] output:

[0984] Evaluation results of match (cosine similarity and Euclidean distance)

[0985] Step 6: Detect anomalies and generate alerts

[0986] The server determines whether there are any invalid voices or abnormalities based on the results of the evaluation of the degree of matching of the voice features, and generates an alert if necessary. The generated alert is sent via email using the SMTP protocol, and also sent in-app via push notification.

[0987] Specific behavior:

[0988] The server detects anomalies based on criteria such as a cosine similarity of 0.7 or less, and sends a warning message to the user's email address via SMTP. At the same time, a warning message is also sent to the user's smartphone via push notification.

[0989] input:

[0990] Matching evaluation results

[0991] output:

[0992] An alert is generated and an email and push notification is sent

[0993] Step 7: Further analysis with emotion recognition

[0994] The server analyzes the user's emotional state from the voice data using an emotion recognition engine (e.g., Google Cloud Speech-to-Text API), and the analyzed emotional state is stored in a database.

[0995] Specific behavior:

[0996] The server sends the voice data to the Google Cloud Speech-to-Text API, which performs emotion recognition and detects emotions such as "anger," "sadness," and "joy."

[0997] input:

[0998] Normalized audio data

[0999] output:

[1000] Detected emotional state

[1001] Step 8: Select notification method based on emotional state

[1002] The server selects the optimal notification method based on the detection result. For example, if the user is determined to be in an "anxious state," it will immediately send an SMS notification using the Twilio API.

[1003] Specific behavior:

[1004] If the emotion recognition result is "anxiety," the server uses the Twilio API to quickly send an SMS containing the message "urgent action required."

[1005] input:

[1006] Detected emotional state

[1007] output:

[1008] Messages sent via the most appropriate notification method

[1009] Step 9: Save the logs

[1010] The server logs all processing results and generated alert information in a PostgreSQL database for later analysis and auditing.

[1011] Specific behavior:

[1012] The server stores all data, including audio preprocessing results, acoustic features, emotion recognition results, and alert information, in a PostgreSQL database and records them with timestamps.

[1013] input:

[1014] Various processing results and alert information

[1015] output:

[1016] Logs stored in the database

[1017] (Application example 2)

[1018] Next, a description will be given of Application Example 2. In the following description, the data processing device 12 will be referred to as a "server" and the smart glasses 214 will be referred to as a "terminal."

[1019] Conventional voice recognition systems have difficulty detecting fraudulent voice data generated by voice generation AI in real time, making it impossible to ensure user safety. Furthermore, they lack the ability to recognize the user's emotional state, making it difficult to respond appropriately based on the user's psychological state. Therefore, there is a demand for a system that combines voice data anomaly detection and user emotion recognition.

[1020] The identification process by the identification processing unit 290 of the data processing device 12 in Application Example 2 is realized by the following means. In this invention, the server includes means for capturing voice data, means for preprocessing the captured voice data, means for comparing extracted acoustic features with existing voice profiles, means for generating and notifying an alert when an abnormality is detected, means for analyzing the user's emotional state from the voice data, and means for selecting a notification method for the alert based on the analysis results. This makes it possible not only to detect abnormalities in voice data in real time, but also to provide an appropriate notification according to the user's emotional state.

[1021] "Voice data" refers to data that records a user's voice information in digital format.

[1022] "Capturing means" refers to equipment or technology used to record or capture audio data.

[1023] "Preprocessing means" refers to the techniques and processes that perform preprocessing such as noise removal and normalization on the captured audio data.

[1024] "Acoustic features" are acoustic feature information extracted from speech data, and include Mel-Frequency Cepstrum Coefficients (MFCCs).

[1025] A "profile" is a database that records information about specific users and their voice characteristics.

[1026] "Means of comparison" refers to the technology or algorithm that matches the extracted acoustic features with existing voice profiles.

[1027] "Means for detecting anomalies" refers to technologies and algorithms that analyze the results of comparing acoustic features to identify fraudulent audio data.

[1028] "Means for generating and notifying alerts" refers to the technology and methods for generating a warning message and notifying the user when an abnormality is detected.

[1029] "Means for analyzing emotional state" refers to technologies and algorithms for recognizing a user's emotions from voice data and assessing their state.

[1030] "Means for selecting a notification method" refers to a technology or method for selecting the most appropriate alert notification method (e.g., SMS, email, etc.) based on the analyzed emotional state.

[1031] MODE FOR CARRYING OUT THE INVENTION

[1032] This invention relates to a system that detects fraudulent voice data generated by a voice generation AI in real time and analyzes the user's emotional state from the voice data to ensure the safety and comfort of the user. Specific embodiments of the system are described below.

[1033] System Configuration

[1034] The system consists of the following major components:

[1035] 1. Terminal

[1036] The terminal is a device that captures the user's voice data. This includes smartphones, tablets, PCs, etc. When the user makes a call or sends a voice message, the voice data is captured in real time and sent to the server via a secure communication method.

[1037] 2. Server

[1038] The server is a unit that processes, analyzes, and recognizes emotions from received voice data. It is mainly responsible for the following tasks:

[1039] Audio data preprocessing

[1040] Acoustic feature extraction

[1041] Comparison with audio profiles

[1042] Anomaly detection and alerting

[1043] emotion recognition

[1044] Selecting notification methods based on emotional state

[1045] Saving logs

[1046] 3. Users

[1047] The user is the person who interacts with the system and checks the alerts: they receive alerts when an anomaly is detected or when a particular emotional state is recognized.

[1048] Specific processing of the program

[1049] This system performs a series of processes in real time, from capturing voice data to analyzing it and generating alerts. The specific process is explained below.

[1050] 1. Capture and transmit audio data

[1051] The device captures voice data in real time when the user initiates a call or voice message, and the captured voice data is sent to the server via encrypted communication means.

[1052] 2. Preprocessing of audio data

[1053] The server performs preprocessing on the received audio data, including noise reduction to reduce background noise and normalization to adjust the audio level to a certain standard. This processing is performed using the Librosa library.

[1054] 3. Acoustic feature extraction

[1055] From the preprocessed audio data, the server extracts acoustic features, including Mel-Frequency Cepstral Coefficients (MFCCs) and spectral features, also using the Librosa library.

[1056] 4. Comparison with audio profiles

[1057] The server compares the extracted acoustic features with existing voice profiles using statistical methods such as cosine similarity and Euclidean distance, and generates an alert if an anomaly is detected.

[1058] 5. Anomaly detection and alert generation

[1059] If the comparison detects unnatural voice changes or signs of synthesis, the server generates an alert, which is promptly sent to the user via SMS, email, or in-app notification.

[1060] 6. Additional analysis using emotion recognition

[1061] The server uses an emotion recognition engine to analyze the user's emotional state from the voice data. Based on the analysis results, it determines whether the voice data matches the user's normal emotional state. DeepMoji and other emotion recognition APIs are used for emotion recognition.

[1062] 7. Notification method selection based on emotional state

[1063] The server selects the optimal notification method based on the detected emotional state. For example, if the user is determined to be in an anxious state, an immediate notification via SMS can be sent to enhance safety.

[1064] Specific examples

[1065] Specific application examples of this system are shown below.

[1066] Example 1:

[1067] If fraudulent voice generation occurs while a user is making an important call at work, the server will detect the anomaly and send a warning notification to the user's smartphone.

[1068] Example 2:

[1069] If a user feels stressed while talking to a family member, the server will recognize the emotion and send a notification encouraging the user to "relax."

[1070] Example prompts for generative AI models

[1071] "What is the execution flow of the application notification when fraudulent voice is detected while a user is on a call with a bank's call center?"

[1072] The flow of the specific processing in the application example 2 will be described with reference to FIG.

[1073] Program processing steps

[1074] Step 1: Capture and send audio data

[1075] Input: A user initiates a call or voice message.

[1076] Specific operation: The device uses the built-in microphone to capture audio data in real time.

[1077] Data processing: Captured audio data is encrypted immediately.

[1078] Output: Encrypted audio data is generated.

[1079] Processing flow: The device sends encrypted audio data to the server via a secure communication method.

[1080] Step 2: Preprocessing the audio data

[1081] Input: Encrypted audio data is sent to the server.

[1082] Specific operation: The server decodes the received audio data.

[1083] Data processing: Denoising and normalization are performed on the decoded audio data using the Librosa library.

[1084] Output: Preprocessed audio data is generated.

[1085] Processing flow: The server sends the noise-removed and normalized audio data to the next step.

[1086] Step 3: Extraction of acoustic features

[1087] Input: Preprocessed audio data is sent to the server.

[1088] Specific operation: The server uses the Librosa library to calculate Mel-Frequency Cepstral Coefficients (MFCCs).

[1089] Data processing: MFCC and spectral features are extracted from the audio data.

[1090] Output: The extracted acoustic features are generated.

[1091] Processing flow: The server sends the extracted features to the next step.

[1092] Step 4: Compare with the audio profile

[1093] Input: The extracted acoustic features are sent to the server.

[1094] Specific behavior: The server loads the existing audio profile.

[1095] Data analysis: Compare acoustic features and voice profiles using cosine similarity and Euclidean distance.

[1096] Output: The comparison results are generated.

[1097] Process flow: The server sends the comparison result to the next step.

[1098] Step 5: Detect anomalies and generate alerts

[1099] Input: The comparison result is sent to the server.

[1100] Specific operation: The server analyzes the comparison results and determines whether an anomaly is detected.

[1101] Data judgment: If an abnormality is detected, an alert message is generated.

[1102] Output: An alert message is generated.

[1103] Process flow: The server sends an alert message to the user's terminal.

[1104] Step 6: Further analysis with emotion recognition

[1105] Input: Preprocessed audio data is sent to the server.

[1106] Specific operation: The server analyzes the voice data using an emotion recognition engine.

[1107] Data analysis: Implementing specific algorithms to identify the user's emotional state from the audio data.

[1108] Output: The emotion recognition results are generated.

[1109] Processing flow: The server refers to the emotion recognition results and sends them to the next step.

[1110] Step 7: Select notification method based on emotional state

[1111] Input: Emotion recognition results are sent to the server.

[1112] Specific operation: The server selects the optimal notification method based on the emotional state.

[1113] Data judgment: Apply the conditions to select the notification method (e.g., SMS notification if the emotional state is stressed).

[1114] Output: The best notification method is selected.

[1115] Process flow: The server notifies the user of the alert according to the selected notification method.

[1116] The specific processing unit 290 transmits the result of the specific processing to the smart glasses 214. In the smart glasses 214, the control unit 46A causes the speaker 240 to output the result of the specific processing. The microphone 238 acquires audio indicating a user input regarding the result of the specific processing. The control unit 46A transmits audio data indicating the user input acquired by the microphone 238 to the data processing device 12. In the data processing device 12, the specific processing unit 290 acquires the audio data.

[1117] The data generation model 58 is a so-called generative AI (Artificial Intelligence). An example of the data generation model 58 is ChatGPT (Internet Search<URL: https: / / openai.com / blog / chatgpt> ), Gemini (Internet search <url: https: gemini.google.com ?hl="ja">) and other generation AIs. The data generation model 58 is obtained by performing deep learning on a neural network. A prompt including an instruction is input to the data generation model 58, and inference data such as voice data indicating voice, text data indicating text, and image data indicating an image is also input. The data generation model 58 performs inference on the input inference data in accordance with the instruction indicated by the prompt, and outputs the inference result in a data format such as voice data and text data. Here, inference refers to, for example, analysis, classification, prediction, and / or summarization.

[1118] In the above embodiment, an example in which the specific processing is performed by the data processing device 12 has been given, but the technology of the present disclosure is not limited to this, and the specific processing may be performed by the smart glasses 214.

[1119] [Third embodiment]

[1120] FIG. 5 shows an example of the configuration of a data processing system 310 according to the third embodiment.

[1121] 5, the data processing system 310 includes the data processing device 12 and a headset type terminal 314. An example of the data processing device 12 is a server.

[1122] The data processing device 12 includes a computer 22, a database 24, and a communication I / F 26. The computer 22 is an example of a "computer" according to the technology of the present disclosure. The computer 22 includes a processor 28, a RAM 30, and a storage 32. The processor 28, the RAM 30, and the storage 32 are connected to a bus 34. The database 24 and the communication I / F 26 are also connected to the bus 34. The communication I / F 26 is connected to a network 54. Examples of the network 54 include a WAN (Wide Area Network) and / or a LAN (Local Area Network).

[1123] The headset type terminal 314 includes a computer 36, a microphone 238, a speaker 240, a camera 42, a communication I / F 44, and a display 343. The computer 36 includes a processor 46, a RAM 48, and a storage 50. The processor 46, the RAM 48, and the storage 50 are connected to a bus 52. The microphone 238, the speaker 240, the camera 42, and the display 343 are also connected to the bus 52.

[1124] The microphone 238 receives instructions and the like from the user 20 by receiving voice uttered by the user 20. The microphone 238 captures the voice uttered by the user 20, converts the captured voice into audio data, and outputs it to the processor 46. The speaker 240 outputs audio in accordance with instructions from the processor 46.

[1125] Camera 42 is a small digital camera equipped with an optical system including a lens, aperture, and shutter, and an imaging element such as a CMOS (Complementary Metal-Oxide-Semiconductor) image sensor or a CCD (Charge Coupled Device) image sensor, and captures images of the surroundings of user 20 (for example, an imaging range defined by an angle of view equivalent to the field of vision of a typical healthy person).

[1126] The communication I / F 44 is connected to a network 54. The communication I / Fs 44 and 26 control the exchange of various information between the processor 46 and the processor 28 via the network 54. The exchange of various information between the processor 46 and the processor 28 using the communication I / Fs 44 and 26 is carried out in a secure state.

[1127] Fig. 6 shows an example of the main functions of the data processing device 12 and the headset type terminal 314. As shown in Fig. 6, in the data processing device 12, a specific process is performed by the processor 28. A specific process program 56 is stored in the storage 32.

[1128] The specific processing program 56 is an example of a "program" according to the technology of the present disclosure. The processor 28 reads the specific processing program 56 from the storage 32 and executes the read specific processing program 56 on the RAM 30. The specific processing is realized by the processor 28 operating as a specific processing unit 290 in accordance with the specific processing program 56 executed on the RAM 30.

[1129] The storage 32 stores a data generation model 58 and an emotion identification model 59. The data generation model 58 and the emotion identification model 59 are used by the identification processing unit 290.

[1130] In the headset type terminal 314, a reception output process is performed by the processor 46. A reception output program 60 is stored in the storage 50. The processor 46 reads the reception output program 60 from the storage 50 and executes the read reception output program 60 on the RAM 48. The reception output process is realized by the processor 46 operating as the control unit 46A in accordance with the reception output program 60 executed on the RAM 48.

[1131] Next, a description will be given of the identification process performed by the identification processing unit 290 of the data processing device 12. In the following description, the data processing device 12 will be referred to as the "server" and the headset type terminal 314 will be referred to as the "terminal."

[1132] This invention relates to a system for detecting fraudulent voice data generated by a voice generation AI in real time and ensuring the safety of users. Specific embodiments of the system are described below.

[1133] System Configuration

[1134] The system consists of the following major components:

[1135] 1. Device: The device that captures the audio data (e.g., a smartphone or PC)

[1136] 2. Server: The unit that processes and analyzes the captured audio data.

[1137] 3. User: A person who uses the system and checks alerts

[1138] Specific processing of the program

[1139] Capture and transmit audio data

[1140] When a device starts a voice call or message transmission, it immediately captures the voice data, which is then sent to the server in real time.

[1141] Examples:

[1142] When the user starts a call, the terminal automatically starts monitoring the voice data and sends that data to the server.

[1143] Audio data preprocessing

[1144] The server performs the necessary pre-processing on the received audio data, which includes the following steps:

[1145] Noise Reduction: Reduce background noise from audio data.

[1146] Normalization: Adjusting audio levels to a certain standard.

[1147] Acoustic feature extraction

[1148] The server extracts acoustic features from the preprocessed speech data, including Mel-Frequency Cepstral Coefficients (MFCCs) and spectral features.

[1149] Examples:

[1150] The server calculates MFCCs from the noise-removed and normalized audio data and extracts them as features.

[1151] Comparison with audio profiles

[1152] The server compares the extracted acoustic features with existing voice profiles using statistical methods such as cosine similarity and Euclidean distance.

[1153] Examples:

[1154] The server compares the voice samples previously registered by the user and calculates how well the current voice data matches.

[1155] Anomaly detection and alerting

[1156] If the comparison detects any abnormal changes or signs of unnatural synthesis, the server generates an alert, which is immediately sent to the user via SMS, email, or in-app notification.

[1157] Examples:

[1158] The server detects unnatural voice changes and sends a warning message to the user's smartphone.

[1159] Saving logs

[1160] All processing results and alert information are stored as logs by the server and used for later analysis and auditing, which increases the transparency and reliability of the system.

[1161] In this way, the system analyzes the user's voice data in real time, detects fraudulent voice production with high accuracy, and protects the user's safety.

[1162] The processing flow will be explained below.

[1163] Step 1:

[1164] The device captures voice data in real time when the user initiates a call or voice message.

[1165] Save the audio data in a specified format (e.g. WAV, MP3).

[1166] Step 2:

[1167] The device transmits the captured audio data to the server.

[1168] Encrypted communications are used for data transmission to ensure data security.

[1169] Step 3:

[1170] The server stores the received audio data and starts pre-processing.

[1171] Noise Reduction: Performs a filtering process to reduce background noise.

[1172] Normalize: Adjust the audio level to a certain standard.

[1173] Step 4:

[1174] The server extracts acoustic features from the preprocessed speech data.

[1175] Calculates Mel-Frequency Cepstral Coefficients (MFCC).

[1176] Analyze the spectral features.

[1177] Step 5:

[1178] The server compares the extracted acoustic features with existing voice profiles.

[1179] Statistical methods such as cosine similarity or Euclidean distance are used.

[1180] Step 6:

[1181] The server analyzes the comparison results to detect any unnatural voice changes or signs of synthesis.

[1182] Apply anomaly detection algorithms and flag any anomalies.

[1183] Step 7:

[1184] The server generates an alert if an anomaly is detected.

[1185] The alert content includes a warning message and details of the abnormality detection.

[1186] Step 8:

[1187] The server notifies the user of the generated alert.

[1188] Notification methods include SMS, email, or in-app notifications.

[1189] Step 9:

[1190] The user receives and confirms the notification.

[1191] If necessary, take appropriate action (e.g., report, suspend use of services).

[1192] Step 10:

[1193] The server stores all processing results and alert information as logs.

[1194] The saved logs can be used for later analysis and auditing.

[1195] The above is the specific processing flow of this system.

[1196] Example 1

[1197] Next, a description will be given of Example 1. In the following description, the data processing device 12 will be referred to as a "server" and the headset type terminal 314 will be referred to as a "terminal."

[1198] In current voice communication systems, the generation of fraudulent voice data using generative AI models is a problem, threatening user safety. Effective methods for detecting such fraudulent voice generation in real time and immediately notifying users are needed. However, existing systems lack sufficient reliability and speed because they are unable to integrate multiple processes, such as voice data capture, preprocessing, anomaly detection, alert generation, and log storage.

[1199] The specific processing by the specific processing unit 290 of the data processing device 12 in the first embodiment is realized by the following means.

[1200] In this invention, the server includes means for capturing voice data, means for preprocessing the captured voice data, means for extracting acoustic features from the preprocessed voice data, means for comparing the extracted acoustic features with an existing voice profile, means for analyzing the comparison result to detect anomalies, means for generating and notifying an alert when an anomaly is detected, and means for storing the processing results and alert information in a database. This enables highly accurate real-time detection and immediate notification of improper voice generation, and recording of all processing results.

[1201] "Audio data" refers to digital or analog data that records sound.

[1202] "Capture" is the process of acquiring audio data.

[1203] "Preprocessing" refers to processing performed to make captured audio data easier to analyze, and includes operations such as noise removal and normalization.

[1204] "Acoustic features" are numerical representations of the characteristics of speech data, and include Mel-Frequency Cepstrum Coefficients (MFCCs) and spectral features.

[1205] An "audio profile" is reference data constructed using features of pre-registered audio data.

[1206] "Comparison" is a process of evaluating the extracted acoustic features and the features of the voice profile using statistical methods.

[1207] "Anomaly detection" is the process of analyzing the comparison results to find incorrect speech production or unusual changes.

[1208] An "alert" is a warning or notification that is generated when an abnormality is detected.

[1209] "Notification" is the process of communicating generated alerts to users, including by means of SMS, email, in-app notifications, etc.

[1210] A "database" is a system for storing processing results and alert information.

[1211] "Storage" is the process of recording processing results and alert information in a database.

[1212] This invention relates to a system that detects fraudulent voice data generation by voice generation AI in real time and ensures user safety.

[1213] System Configuration

[1214] The system consists of the following major components:

[1215] 1. Device: The device that captures the audio data (e.g., a smartphone or PC)

[1216] 2. Server: The unit that processes and analyzes the captured audio data.

[1217] 3. User: A person who uses the system and checks alerts

[1218] Specific processing of the program

[1219] Capture and transmit audio data

[1220] When a user initiates a voice call or sends a message, the device captures the voice data using the built-in microphone and transmits it to the server in real time. The hardware used is the microphone of a smartphone or PC, and the software is a dedicated voice capture application (e.g., a homemade VoiceCaptureApp).

[1221] Specific behavior:

[1222] The device automatically starts capturing audio when the user starts a call and sends the data to a remote server using SSL / TLS encryption, ensuring data security.

[1223] Audio data preprocessing

[1224] The server performs preprocessing on the received audio data, which mainly includes noise removal and voice level normalization. Specifically, it applies a filter to the audio data to remove background noise, and then normalizes the voice level.

[1225] Specific behavior:

[1226] The server uses Python's SciPy library to apply a bandpass filter to remove background noise.

[1227] After noise removal, the audio levels are normalized to the range 0 to 1 using the normalize function from the Librosa library.

[1228] Acoustic feature extraction

[1229] The server extracts acoustic features from the preprocessed speech data, including Mel-Frequency Cepstral Coefficients (MFCCs) and spectral features.

[1230] Specific behavior:

[1231] The server extracts MFCCs from the audio data using the mfcc function in the Librosa library and stores them in a database as acoustic features.

[1232] Comparison with audio profiles

[1233] The server compares the extracted acoustic features with pre-trained voice profiles using cosine similarity or Euclidean distance.

[1234] Specific behavior:

[1235] The server uses the cosine function in the SciPy library to calculate the similarity between the features of the current voice data and the features of the pre-registered voice profile.

[1236] If the calculation results below a certain threshold, it is flagged as a likely incorrect speech production.

[1237] Anomaly detection and alerting

[1238] If the server detects an anomaly in the audio data, it generates an alert and notifies the user immediately via SMS, email, or in-app notification.

[1239] Specific behavior:

[1240] The server uses an anomaly detection algorithm (e.g., Scikit-learn's Isolation Forest) to detect unnatural changes in the audio data.

[1241] If an anomaly is detected, a warning message is sent to the user's smartphone using Twilio's SMS API, and email notifications are sent using the SendGrid API.

[1242] Saving logs

[1243] The server stores all processing results and alert information in a database, allowing for later analysis and auditing, improving system reliability.

[1244] Specific behavior:

[1245] The server stores the processing results and alert information in a MySQL database using SQLAlchemy to create database entries.

[1246] Example prompts to input to the generative AI model

[1247] "Please explain how speech generation AI detects fraudulent audio data. Please include specific steps for capturing audio data, preprocessing, extracting acoustic features, comparing with audio profiles, detecting anomalies, generating alerts, and saving logs."

[1248] In this way, the system can analyze the user's voice data in real time, detect fraudulent voice production with high accuracy, and ensure the user's safety.

[1249] The flow of the identification process in the first embodiment will be described with reference to FIG.

[1250] Processing Steps

[1251] Step 1: Capture and send audio data

[1252] The device uses a built-in microphone to capture voice data when a user initiates a voice call or sends a message.

[1253] input:

[1254] Initiate voice calls and send messages

[1255] Specific behavior:

[1256] The device starts a dedicated voice capture application (e.g., a custom VoiceCaptureApp) and starts microphone input. The captured voice data is sent to the server using encryption (SSL / TLS).

[1257] output:

[1258] Encrypted and transmitted voice data

[1259] Step 2: Preprocessing the audio data

[1260] The server performs preprocessing on the received audio data.

[1261] input:

[1262] Received audio data

[1263] Specific behavior:

[1264] The server uses Python's SciPy library to apply a bandpass filter to remove background noise, then normalizes the audio level to the range 0 to 1 using the normalize function from the Librosa library.

[1265] output:

[1266] Denoised and normalized audio data

[1267] Step 3: Extraction of acoustic features

[1268] The server extracts acoustic features from the preprocessed speech data.

[1269] input:

[1270] Denoised and normalized audio data

[1271] Specific behavior:

[1272] The server extracts MFCCs using the mfcc function in the Librosa library and stores them in a database as acoustic features.

[1273] output:

[1274] Extracted acoustic features

[1275] Step 4: Compare with the audio profile

[1276] The server compares the extracted acoustic features with pre-registered voice profiles.

[1277] input:

[1278] Extracted acoustic features

[1279] Pre-registered audio profiles

[1280] Specific behavior:

[1281] The server uses the cosine function in the SciPy library to calculate the similarity between the features of the current voice data and the features of the existing voice profile.

[1282] output:

[1283] Similarity calculation results

[1284] Step 5: Detect anomalies and generate alerts

[1285] If the server detects an abnormality in the voice data, it generates an alert and immediately notifies the user.

[1286] input:

[1287] Similarity calculation results

[1288] Specific behavior:

[1289] The server uses an anomaly detection algorithm (e.g., Scikit-learn's Isolation Forest) to detect anomalies. If an anomaly is detected, it uses Twilio's SMS API to send an alert message to the user's smartphone. For email notifications, it uses the SendGrid API.

[1290] output:

[1291] Warning message

[1292] Step 6: Save the logs

[1293] The server stores all processing results and alert information in a database.

[1294] input:

[1295] Similarity calculation results

[1296] Warning message

[1297] Specific behavior:

[1298] The server uses SQLAlchemy to create and save database entries to store the processing results and alert information in a MySQL database.

[1299] output:

[1300] Saved processing results and alert information

[1301] (Application example 1)

[1302] Next, a description will be given of Application Example 1. In the following description, the data processing device 12 will be referred to as a "server" and the headset type terminal 314 will be referred to as a "terminal."

[1303] With the advancement of AI voice generation technology, it has become easier to generate voice data fraudulently. As a result, the risk of personal and corporate voice communications being used fraudulently and important call content being tampered with is increasing. This has a particularly severe impact on business communications that contain confidential information. Current systems face the challenge of being unable to detect such fraudulent activity in real time and respond immediately.

[1304] The specific processing by the specific processing unit 290 of the data processing device 12 in the application example 1 is realized by the following means.

[1305] In this invention, the server includes means for capturing voice data, means for preprocessing the captured voice data, means for extracting acoustic features from the preprocessed voice data, means for comparing the extracted acoustic features with a voice profile, means for detecting anomalies based on the analysis results, means for generating and notifying an alarm when an anomaly is detected, means for recording the alarm information, means for detecting fraudulent generation of voice data in real time, an application for monitoring call data, and a cloud-based processing server. This makes it possible to detect fraudulent generation of voice data in real time and respond quickly.

[1306] "Audio Data" means data for recording, storing, and transmitting audio in digital format.

[1307] A "capturing means" is a device or software for capturing audio data in real time.

[1308] The "pre-processing means" refers to a method or device for removing noise from the captured audio data and normalizing it.

[1309] "Means for extracting acoustic features" refers to a method or device that calculates and extracts characteristic information (e.g., Mel-Frequency Cepstral Coefficients (MFCCs)) from preprocessed speech data.

[1310] An "audio profile" is a collection of feature data generated based on specific audio data.

[1311] The "comparison means" refers to a method or device for comparing the extracted acoustic features with a voice profile and evaluating the degree of similarity.

[1312] The "analysis result" is information indicating the degree of agreement of the data obtained from the comparison of the acoustic features and the presence or absence of abnormalities.

[1313] "Means for detecting anomalies" refers to a method or device that determines whether there are unnatural changes or signs of fraudulent generation in the audio data based on the analysis results.

[1314] The "means for generating and notifying a warning" refers to a method or device for generating and notifying a warning message to a user when an abnormality is detected.

[1315] The "means for recording warning information" refers to a method or device for saving generated warning messages and analysis results so that they can be checked at a later date.

[1316] "Real-time detection means" refers to a method or device that can detect fraud immediately without interrupting the time when voice data is being generated.

[1317] An "application that monitors call data" is software that runs on a device such as a smartphone and monitors voice data during a call.

[1318] A "cloud-based processing server" is a remote server accessible via the Internet that pre-processes, analyzes, and compares audio data.

[1319] This invention relates to a system that detects fraudulent voice data generated by voice generation AI in real time during phone calls to ensure user security. This system automatically captures voice data, preprocesses it, extracts acoustic features, compares it with existing voice profiles, detects anomalies, generates warnings, and records it.

[1320] System Configuration

[1321] The system consists of the following major components:

[1322] 1. Terminal: A device for capturing audio data (e.g., a smartphone)

[1323] 2. Server: A unit for processing and analyzing the captured audio data. Uses a cloud-based processing server (e.g., AWS, Google Cloud Platform).

[1324] 3. Users: People who use the system and receive alerts

[1325] Specific processing of the program

[1326] Capture audio data

[1327] The device uses the Twilio Voice SDK to capture voice data during a call in real time and immediately transmits that data to the server.

[1328] Audio data preprocessing

[1329] The server uses libraries (e.g., PyDub, Librosa) to denoise and normalize the audio data. This preprocessing produces clean audio data suitable for analysis.

[1330] Acoustic feature extraction

[1331] From the preprocessed audio data, the server uses Librosa to extract acoustic features such as Mel-Frequency Cepstral Coefficients (MFCCs).

[1332] Comparison with audio profiles

[1333] The extracted acoustic features are compared to existing voice profiles using machine learning frameworks such as TensorFlow, using statistical methods such as cosine similarity.

[1334] Anomaly detection and alert generation

[1335] Based on the analysis results, the server determines whether the voice data has any unnatural variations or signs of fraudulent generation, and if an anomaly is detected, the server sends a warning to the user via SMS or in-app notification.

[1336] Recording warning information

[1337] All processing results and warning information are logged by the server for later analysis and auditing.

[1338] Specific examples

[1339] For example, suppose a user is using the Secure Call Detector app to make a call for a business meeting and the call content is tampered with. The system detects this tampering in real time and immediately generates an alert. The user receives the following message on their smartphone:

[1340] "Warning: Indications of unauthorized audio have been detected. The following message will appear on your screen: 'Warning: Unauthorized audio has been detected. Please ensure your call is secure.' Call recording will automatically stop and your system administrator will be notified."

[1341] This allows the user to respond quickly and ensures the safety of the call.

[1342] The system of the present invention can detect fraudulent generation of voice data in real time and respond quickly, thereby enhancing the security of individuals and businesses.

[1343] The flow of the specific processing in the application example 1 will be described with reference to FIG.

[1344] Step 1:

[1345] The terminal uses the Twilio Voice SDK to detect when a user starts a call. When a call starts, it immediately starts capturing voice data and sends that data to the server in real time. The input is the user's call voice data, and the output is the raw voice data sent to the server.

[1346] Step 2:

[1347] The server preprocesses the audio data received from the device using a library (e.g., PyDub, Librosa). This preprocessing removes noise from the audio data and normalizes the audio level. The input is raw audio data, and the output is noise-removed and normalized audio data.

[1348] Step 3:

[1349] The server extracts acoustic features from the preprocessed speech data. It uses Librosa to extract features such as Mel-Frequency Cepstral Coefficients (MFCCs). The input is the preprocessed speech data, and the output is the extracted acoustic features.

[1350] Step 4:

[1351] The server compares the extracted acoustic features with existing voice profiles. This comparison uses TensorFlow's machine learning model to evaluate similarity using statistical methods such as cosine similarity and Euclidean distance. The input is the acoustic features and the voice profile, and the output is a similarity score.

[1352] Step 5:

[1353] The server analyzes the similarity score and determines whether or not unauthorized voice data has been generated. If an abnormality is detected based on the analysis results, a warning message is generated. The input is the similarity score, and the output is the warning message.

[1354] Step 6:

[1355] The server notifies the user via SMS or in-app notification with the generated alert message. This notification is sent using Twilio's messaging API. The input is the alert message, and the output is the alert message displayed on the user's device.

[1356] Step 7:

[1357] The server stores all processing results and warning information in a database. This information is recorded using a database such as MySQL for later analysis and auditing. The input is the processing results and warning information, and the output is the log information stored in the database.

[1358] Furthermore, an emotion engine that estimates the user's emotion may be further combined. That is, the identification processing unit 290 may estimate the user's emotion using the emotion identification model 59, and perform identification processing using the user's emotion.

[1359] This invention relates to a system that ensures the safety and comfort of users by detecting fraudulent voice data generated by voice generation AI in real time and recognizing the emotional state of the user from the voice data. Specific embodiments of the system are described below.

[1360] System Configuration

[1361] The system consists of the following major components:

[1362] 1. Device: The device that captures the audio data (e.g., a smartphone or PC)

[1363] 2. Server: The unit that processes, analyzes, and recognizes emotions from captured audio data.

[1364] 3. User: A person who uses the system and checks alerts

[1365] Specific processing of the program

[1366] Capture and transmit audio data

[1367] The device captures voice data in real time when the user initiates a call or voice message, and the captured voice data is sent to the server via encrypted communication means.

[1368] Examples:

[1369] The terminal starts monitoring the voice data at the same time as the user starts a call, and transmits the voice data to the server via secure communication.

[1370] Audio data preprocessing

[1371] The server performs pre-processing on the received audio data, including noise reduction to reduce background noise and normalization to adjust the audio level to a certain standard.

[1372] Acoustic feature extraction

[1373] The server extracts acoustic features from the preprocessed speech data, including Mel-Frequency Cepstral Coefficients (MFCCs) and spectral features.

[1374] Examples:

[1375] The server calculates MFCCs from the noise-removed and normalized audio data and obtains them as feature data.

[1376] Comparison with audio profiles

[1377] The server compares the extracted acoustic features with existing voice profiles using statistical methods such as cosine similarity and Euclidean distance.

[1378] Examples:

[1379] The server compares the voice samples previously registered by the user and calculates how well the current voice data matches.

[1380] Anomaly detection and alerting

[1381] If the comparison detects unnatural voice changes or signs of synthesis, the server generates an alert, which is promptly sent to the user via SMS, email, or in-app notification.

[1382] Examples:

[1383] The server detects abnormal audio changes and sends a warning message to the user's smartphone.

[1384] Additional analysis with emotion recognition

[1385] The server uses an emotion engine to analyze the user's emotional state from the voice data, and determines whether the voice data matches the user's typical emotional state.

[1386] Examples:

[1387] The server analyzes the voice data to determine whether the user is expressing emotions such as anxiety, anger, or joy.

[1388] Selecting notification methods based on emotional state

[1389] The server selects the optimal notification method based on the detected emotional state. For example, if the user is determined to be in an anxious state, an immediate notification via SMS can be sent to enhance safety.

[1390] Examples:

[1391] Based on the results of the emotion engine, if the user is feeling stressed, the server determines that a prompt response is required and selects an SMS notification.

[1392] Saving logs

[1393] All processing results and alert information are stored as logs by the server, which can be used for later analysis and auditing, ensuring the transparency and reliability of the system.

[1394] In this way, the system analyzes the user's voice data in real time, detects fraudulent voice production with high accuracy, and by recognizing the user's emotional state, selects an appropriate notification method to maintain the user's safety and comfort.

[1395] The processing flow will be explained below.

[1396] Step 1:

[1397] The device automatically captures voice data in real time when the user initiates a call or voice message.

[1398] Save the audio data in a specified format (e.g. WAV, MP3).

[1399] Step 2:

[1400] The device transmits the captured audio data to the server.

[1401] Encrypted communications are used for data transmission to ensure data security.

[1402] Step 3:

[1403] The server stores the received audio data and starts pre-processing.

[1404] The server first applies a noise reduction filter to reduce background noise.

[1405] The server then normalizes the audio levels to ensure consistency.

[1406] Step 4:

[1407] The server extracts acoustic features from the preprocessed speech data.

[1408] The server calculates Mel-Frequency Cepstral Coefficients (MFCCs) and extracts them as speech features.

[1409] The server also analyzes the spectral features and stores them as supplementary features.

[1410] Step 5:

[1411] The server compares the extracted acoustic features with existing voice profiles.

[1412] The server calculates cosine similarity and Euclidean distance to evaluate the degree of matching of the voices.

[1413] Step 6:

[1414] The server analyzes the comparison results to detect any unnatural voice changes or signs of synthesis.

[1415] The server applies an anomaly detection algorithm and flags any anomalies detected.

[1416] Step 7:

[1417] The server generates an alert if an anomaly is detected.

[1418] The server generates alert information including warning messages and details of anomaly detections.

[1419] Step 8:

[1420] The server notifies the user of the generated alert.

[1421] Notification methods include SMS, email, or in-app notifications.

[1422] Step 9:

[1423] The user receives and confirms the notification.

[1424] Based on the content of the notification, the user will take appropriate action as necessary (e.g., report the issue, stop using the service).

[1425] Step 10:

[1426] The server uses an emotion engine to analyze the user's emotional state from the voice data.

[1427] The server identifies the emotional state (e.g., anxiety, anger, joy, etc.) from the voice data.

[1428] Step 11:

[1429] The server selects the most appropriate notification method based on the detected emotional state.

[1430] For example, if a user is in a state of anxiety, they will be notified immediately via SMS to enhance safety.

[1431] Step 12:

[1432] The server stores all processing results and alert information as logs.

[1433] The saved logs can be used for later analysis and auditing.

[1434] This is the specific processing flow of this system. This system allows users to detect fraudulent generation of voice data with high accuracy, and by recognizing the user's emotional state, it enables more appropriate responses.

[1435] Example 2

[1436] Next, a description will be given of Example 2. In the following description, the data processing device 12 will be referred to as a "server" and the headset type terminal 314 will be referred to as a "terminal."

[1437] Conventional voice data analysis systems have difficulty detecting fraudulent voice generation in real time, and there are no systems that can recognize the user's emotional state and take appropriate action. This increases the risk of voice data being manipulated or used fraudulently, and poses a challenge to ensuring user safety and comfort.

[1438] The identification process by the identification processing unit 290 of the data processing device 12 in the second embodiment is realized by the following means. In this invention, the server includes means for preprocessing voice data, means for extracting acoustic features from the preprocessed voice data, means for comparing the extracted acoustic features with an existing voice profile, and means for selecting a notification method based on the recognized emotional state. This enables highly accurate detection of fraudulent voice generation and appropriate response based on the user's emotional state.

[1439] "Audio data" is a digital representation of an acoustic signal captured using a microphone or other audio input device.

[1440] "Capture" is the process of capturing audio data in real time via a microphone or other input device.

[1441] "Preprocessing" refers to the process of applying early-stage processing such as noise removal and normalization to speech data to facilitate analysis and feature extraction.

[1442] "Encryption" is a method of converting audio data using a specific algorithm to protect it from third parties.

[1443] "Transmission" refers to the act of transferring captured audio data to another system or server via a network.

[1444] "Acoustic features" are computed physical or statistical properties, such as Mel-Frequency Cepstral Coefficients (MFCCs) or spectral features, extracted from speech data.

[1445] "Comparison" is a process in which the extracted acoustic features are matched with an existing data set called a voice profile, and the degree of match is evaluated.

[1446] "Anomaly detection" is a function that detects and reports unnatural voice changes and traces of synthesized voice as a result of voice data analysis.

[1447] An "alert" is a message or signal that notifies a user or system administrator when an abnormality is detected.

[1448] "Emotion recognition" is a technology that analyzes and determines a user's emotional state (e.g., joy, anxiety, anger, etc.) from voice data.

[1449] "Notification method selection" is the process of selecting the optimal method (e.g., SMS, email, in-app notification, etc.) to notify the user based on the recognized emotional state.

[1450] This invention relates to a system that ensures the safety and comfort of users by detecting fraudulent voice data generated by voice generation AI in real time and recognizing the emotional state of the user from the voice data. Specific embodiments are described below.

[1451] System Configuration

[1452] The system consists of the following major components:

[1453] 1. Device: The device that captures the audio data (e.g., a smartphone or computer)

[1454] 2. Server: The unit that processes, analyzes, and recognizes emotions from captured audio data.

[1455] 3. User: A person who uses the system and checks alerts

[1456] Capture and transmit audio data

[1457] When a user initiates a call or voice message, the device captures audio data in real time through the microphone and transmits the captured audio data to the server using SSL / TLS encryption.

[1458] Examples:

[1459] The device collects the audio as soon as the user starts a call on their smartphone and sends it to the server as encrypted data.

[1460] Audio data preprocessing

[1461] The server performs pre-processing on the received audio data, including removing background noise using a noise attenuation algorithm (e.g., Wiener filter) and normalizing the volume of the audio data.

[1462] Examples:

[1463] The server applies a Wiener filter to the received audio data to reduce background noise and normalize the audio level to 0.5.

[1464] Acoustic feature extraction

[1465] From the preprocessed audio data, the server extracts Mel-Frequency Cepstral Coefficients (MFCCs) and spectral features using the Python library Librosa.

[1466] Examples:

[1467] The server extracts 13-dimensional MFCC features from the preprocessed audio data using Librosa and stores them in memory as an array.

[1468] Comparison with audio profiles

[1469] The server compares the extracted acoustic features with the voice profile registered by the user in advance, and detects fraudulent voices by calculating cosine similarity and Euclidean distance using SciPy to evaluate the degree of match.

[1470] Examples:

[1471] The server passes the generated MFCC features to the SciPy cosine similarity function to compare the similarity with pre-registered voice profiles. It also calculates the Euclidean distance to evaluate the match from different angles.

[1472] Anomaly detection and alerting

[1473] If the comparison detects unnatural voice changes or signs of synthetic speech, the server detects the anomaly and generates an alert. The alert is sent via email using the SMTP protocol and also via push notifications within the app.

[1474] Examples:

[1475] If the cosine similarity is below a certain level, the server determines that an abnormality has occurred and sends a warning message to the user's email address via SMTP. At the same time, a warning message is also sent to the user's smartphone via push notification.

[1476] Additional analysis with emotion recognition

[1477] The server uses a voice emotion recognition engine (e.g., Google Cloud Speech-to-Text API) to analyze the user's emotional state from the voice data and determine whether it matches their normal emotional state.

[1478] Examples:

[1479] The server uses the Google Cloud Speech-to-Text API to perform emotion recognition and detect emotions such as "anger," "sadness," and "joy." The detection results are saved in tensor format.

[1480] Selecting notification methods based on emotional state

[1481] The server selects the optimal notification method based on the detected emotional state. For example, if the user is determined to be in an "anxious" state, it will send an SMS notification immediately using the Twilio API.

[1482] Examples:

[1483] If the server detects an "uneasy" state, it will use the Twilio API to send an SMS notification with the message "urgent action required."

[1484] Saving logs

[1485] All processing results and generated alert information are stored as logs in a PostgreSQL database by the server, which can be used for later analysis and auditing, increasing the transparency and security of the system.

[1486] Examples:

[1487] The server inserts all data, including the results of speech preprocessing, acoustic features, emotion recognition results, and alert information, into a PostgreSQL database and stores it along with a timestamp.

[1488] Prompt Sentence Examples

[1489] As an example of a prompt for this system, the following query is input to the generative AI model:

[1490] "When a user sends a voice message, we want to identify the emotion from that data and activate fraudulent voice detection. Please suggest an algorithm."

[1491] The flow of the identification process in the second embodiment will be described with reference to FIG.

[1492] Step 1: Capture audio data

[1493] When a user initiates a call or voice message, the device uses the built-in microphone to capture voice data in real time, which is then recorded in WAV format and temporarily stored in the device's memory.

[1494] Specific behavior:

[1495] The moment a user starts a call on their smartphone, the device captures the audio data at a sampling rate of 44.1 kHz and saves it as a temporary file.

[1496] input:

[1497] Call start triggers

[1498] output:

[1499] Audio data captured by the built-in microphone (WAV format)

[1500] Step 2: Sending audio data

[1501] The device protects the captured audio data with AES-256 encryption and sends it to the server using SSL / TLS, preventing data leakage during transmission.

[1502] Specific behavior:

[1503] The device encrypts the voice data with AES-256 and transmits it in real time to the server via the HTTPS protocol.

[1504] input:

[1505] Captured audio data (WAV format)

[1506] output:

[1507] The encrypted audio data is sent to the server

[1508] Step 3: Preprocessing the audio data

[1509] The server applies a noise attenuation algorithm (e.g., a Wiener filter) to the received audio data to normalize the audio level.

[1510] Specific behavior:

[1511] The server applies a Wiener filter to the received audio data to reduce noise and normalize the volume level of the audio data to 0.5.

[1512] input:

[1513] Encrypted and decrypted audio data

[1514] output:

[1515] Noise-reduced and normalized audio data

[1516] Step 4: Extraction of acoustic features

[1517] The server extracts Mel-Frequency Cepstral Coefficients (MFCCs) and spectral features from the preprocessed audio data using the Librosa library.

[1518] Specific behavior:

[1519] The server uses Librosa to calculate 13-dimensional MFCC features from the preprocessed audio data and saves them as an array.

[1520] input:

[1521] Normalized audio data

[1522] output:

[1523] Extracted acoustic features (MFCC)

[1524] Step 5: Compare with the audio profile

[1525] The server uses SciPy to calculate the cosine similarity and Euclidean distance between the extracted acoustic features and the voice profile previously registered by the user, and evaluates the degree of match.

[1526] Specific behavior:

[1527] The server passes the generated MFCC features to SciPy's cosine similarity function to compare them with existing audio profiles and calculates the degree of match. It also calculates the Euclidean distance.

[1528] input:

[1529] Extracted acoustic features

[1530] Existing Audio Profiles

[1531] output:

[1532] Evaluation results of match (cosine similarity and Euclidean distance)

[1533] Step 6: Detect anomalies and generate alerts

[1534] The server determines whether there are any invalid voices or abnormalities based on the results of the evaluation of the degree of matching of the voice features, and generates an alert if necessary. The generated alert is sent via email using the SMTP protocol, and also sent in-app via push notification.

[1535] Specific behavior:

[1536] The server detects anomalies based on criteria such as a cosine similarity of 0.7 or less, and sends a warning message to the user's email address via SMTP. At the same time, a warning message is also sent to the user's smartphone via push notification.

[1537] input:

[1538] Matching evaluation results

[1539] output:

[1540] An alert is generated and an email and push notification is sent

[1541] Step 7: Further analysis with emotion recognition

[1542] The server analyzes the user's emotional state from the voice data using an emotion recognition engine (e.g., Google Cloud Speech-to-Text API), and the analyzed emotional state is stored in a database.

[1543] Specific behavior:

[1544] The server sends the voice data to the Google Cloud Speech-to-Text API, which performs emotion recognition and detects emotions such as "anger," "sadness," and "joy."

[1545] input:

[1546] Normalized audio data

[1547] output:

[1548] Detected emotional state

[1549] Step 8: Select notification method based on emotional state

[1550] The server selects the optimal notification method based on the detection result. For example, if the user is determined to be in an "anxious state," it will immediately send an SMS notification using the Twilio API.

[1551] Specific behavior:

[1552] If the emotion recognition result is "anxiety," the server uses the Twilio API to quickly send an SMS containing the message "urgent action required."

[1553] input:

[1554] Detected emotional state

[1555] output:

[1556] Messages sent via the most appropriate notification method

[1557] Step 9: Save the logs

[1558] The server logs all processing results and generated alert information in a PostgreSQL database for later analysis and auditing.

[1559] Specific behavior:

[1560] The server stores all data, including audio preprocessing results, acoustic features, emotion recognition results, and alert information, in a PostgreSQL database and records them with timestamps.

[1561] input:

[1562] Various processing results and alert information

[1563] output:

[1564] Logs stored in the database

[1565] (Application example 2)

[1566] Next, a description will be given of Application Example 2. In the following description, the data processing device 12 will be referred to as a "server" and the headset type terminal 314 will be referred to as a "terminal."

[1567] Conventional voice recognition systems have difficulty detecting fraudulent voice data generated by voice generation AI in real time, making it impossible to ensure user safety. Furthermore, they lack the ability to recognize the user's emotional state, making it difficult to respond appropriately based on the user's psychological state. Therefore, there is a demand for a system that combines voice data anomaly detection and user emotion recognition.

[1568] The identification process by the identification processing unit 290 of the data processing device 12 in Application Example 2 is realized by the following means. In this invention, the server includes means for capturing voice data, means for preprocessing the captured voice data, means for comparing extracted acoustic features with existing voice profiles, means for generating and notifying an alert when an abnormality is detected, means for analyzing the user's emotional state from the voice data, and means for selecting a notification method for the alert based on the analysis results. This makes it possible not only to detect abnormalities in voice data in real time, but also to provide an appropriate notification according to the user's emotional state.

[1569] "Voice data" refers to data that records a user's voice information in digital format.

[1570] "Capturing means" refers to equipment or technology used to record or capture audio data.

[1571] "Preprocessing means" refers to the techniques and processes that perform preprocessing such as noise removal and normalization on the captured audio data.

[1572] "Acoustic features" are acoustic feature information extracted from speech data, and include Mel-Frequency Cepstrum Coefficients (MFCCs).

[1573] A "profile" is a database that records information about specific users and their voice characteristics.

[1574] "Means of comparison" refers to the technology or algorithm that matches the extracted acoustic features with existing voice profiles.

[1575] "Means for detecting anomalies" refers to technologies and algorithms that analyze the results of comparing acoustic features to identify fraudulent audio data.

[1576] "Means for generating and notifying alerts" refers to the technology and methods for generating a warning message and notifying the user when an abnormality is detected.

[1577] "Means for analyzing emotional state" refers to technologies and algorithms for recognizing a user's emotions from voice data and assessing their state.

[1578] "Means for selecting a notification method" refers to a technology or method for selecting the most appropriate alert notification method (e.g., SMS, email, etc.) based on the analyzed emotional state.

[1579] MODE FOR CARRYING OUT THE INVENTION

[1580] This invention relates to a system that detects fraudulent voice data generated by a voice generation AI in real time and analyzes the user's emotional state from the voice data to ensure the safety and comfort of the user. Specific embodiments of the system are described below.

[1581] System Configuration

[1582] The system consists of the following major components:

[1583] 1. Terminal

[1584] The terminal is a device that captures the user's voice data. This includes smartphones, tablets, PCs, etc. When the user makes a call or sends a voice message, the voice data is captured in real time and sent to the server via a secure communication method.

[1585] 2. Server

[1586] The server is a unit that processes, analyzes, and recognizes emotions from received voice data. It is mainly responsible for the following tasks:

[1587] Audio data preprocessing

[1588] Acoustic feature extraction

[1589] Comparison with audio profiles

[1590] Anomaly detection and alerting

[1591] emotion recognition

[1592] Selecting notification methods based on emotional state

[1593] Saving logs

[1594] 3. Users

[1595] The user is the person who interacts with the system and checks the alerts: they receive alerts when an anomaly is detected or when a particular emotional state is recognized.

[1596] Specific processing of the program

[1597] This system performs a series of processes in real time, from capturing voice data to analyzing it and generating alerts. The specific process is explained below.

[1598] 1. Capture and transmit audio data

[1599] The device captures voice data in real time when the user initiates a call or voice message, and the captured voice data is sent to the server via encrypted communication means.

[1600] 2. Preprocessing of audio data

[1601] The server performs preprocessing on the received audio data, including noise reduction to reduce background noise and normalization to adjust the audio level to a certain standard. This processing is performed using the Librosa library.

[1602] 3. Acoustic feature extraction

[1603] From the preprocessed audio data, the server extracts acoustic features, including Mel-Frequency Cepstral Coefficients (MFCCs) and spectral features, also using the Librosa library.

[1604] 4. Comparison with audio profiles

[1605] The server compares the extracted acoustic features with existing voice profiles using statistical methods such as cosine similarity and Euclidean distance, and generates an alert if an anomaly is detected.

[1606] 5. Anomaly detection and alert generation

[1607] If the comparison detects unnatural voice changes or signs of synthesis, the server generates an alert, which is promptly sent to the user via SMS, email, or in-app notification.

[1608] 6. Additional analysis using emotion recognition

[1609] The server uses an emotion recognition engine to analyze the user's emotional state from the voice data. Based on the analysis results, it determines whether the voice data matches the user's normal emotional state. DeepMoji and other emotion recognition APIs are used for emotion recognition.

[1610] 7. Notification method selection based on emotional state

[1611] The server selects the optimal notification method based on the detected emotional state. For example, if the user is determined to be in an anxious state, an immediate notification via SMS can be sent to enhance safety.

[1612] Specific examples

[1613] Specific application examples of this system are shown below.

[1614] Example 1:

[1615] If fraudulent voice generation occurs while a user is making an important call at work, the server will detect the anomaly and send a warning notification to the user's smartphone.

[1616] Example 2:

[1617] If a user feels stressed while talking to a family member, the server will recognize the emotion and send a notification encouraging the user to "relax."

[1618] Example prompts for generative AI models

[1619] "What is the execution flow of the application notification when fraudulent voice is detected while a user is on a call with a bank's call center?"

[1620] The flow of the specific processing in the application example 2 will be described with reference to FIG.

[1621] Program processing steps

[1622] Step 1: Capture and send audio data

[1623] Input: A user initiates a call or voice message.

[1624] Specific operation: The device uses the built-in microphone to capture audio data in real time.

[1625] Data processing: Captured audio data is encrypted immediately.

[1626] Output: Encrypted audio data is generated.

[1627] Processing flow: The device sends encrypted audio data to the server via a secure communication method.

[1628] Step 2: Preprocessing the audio data

[1629] Input: Encrypted audio data is sent to the server.

[1630] Specific operation: The server decodes the received audio data.

[1631] Data processing: Denoising and normalization are performed on the decoded audio data using the Librosa library.

[1632] Output: Preprocessed audio data is generated.

[1633] Processing flow: The server sends the noise-removed and normalized audio data to the next step.

[1634] Step 3: Extraction of acoustic features

[1635] Input: Preprocessed audio data is sent to the server.

[1636] Specific operation: The server uses the Librosa library to calculate Mel-Frequency Cepstral Coefficients (MFCCs).

[1637] Data processing: MFCC and spectral features are extracted from the audio data.

[1638] Output: The extracted acoustic features are generated.

[1639] Processing flow: The server sends the extracted features to the next step.

[1640] Step 4: Compare with the audio profile

[1641] Input: The extracted acoustic features are sent to the server.

[1642] Specific behavior: The server loads the existing audio profile.

[1643] Data analysis: Compare acoustic features and voice profiles using cosine similarity and Euclidean distance.

[1644] Output: The comparison results are generated.

[1645] Process flow: The server sends the comparison result to the next step.

[1646] Step 5: Detect anomalies and generate alerts

[1647] Input: The comparison result is sent to the server.

[1648] Specific operation: The server analyzes the comparison results and determines whether an anomaly is detected.

[1649] Data judgment: If an abnormality is detected, an alert message is generated.

[1650] Output: An alert message is generated.

[1651] Process flow: The server sends an alert message to the user's terminal.

[1652] Step 6: Further analysis with emotion recognition

[1653] Input: Preprocessed audio data is sent to the server.

[1654] Specific operation: The server analyzes the voice data using an emotion recognition engine.

[1655] Data analysis: Implementing specific algorithms to identify the user's emotional state from the audio data.

[1656] Output: The emotion recognition results are generated.

[1657] Processing flow: The server refers to the emotion recognition results and sends them to the next step.

[1658] Step 7: Select notification method based on emotional state

[1659] Input: Emotion recognition results are sent to the server.

[1660] Specific operation: The server selects the optimal notification method based on the emotional state.

[1661] Data judgment: Apply the conditions to select the notification method (e.g., SMS notification if the emotional state is stressed).

[1662] Output: The best notification method is selected.

[1663] Process flow: The server notifies the user of the alert according to the selected notification method.

[1664] The specific processing unit 290 transmits the result of the specific processing to the headset type terminal 314. In the headset type terminal 314, the control unit 46A causes the speaker 240 and the display 343 to output the result of the specific processing. The microphone 238 acquires audio indicating a user input regarding the result of the specific processing. The control unit 46A transmits audio data indicating the user input acquired by the microphone 238 to the data processing device 12. In the data processing device 12, the specific processing unit 290 acquires the audio data.

[1665] The data generation model 58 is a so-called generative AI (Artificial Intelligence). An example of the data generation model 58 is ChatGPT (Internet Search<URL: https: / / openai.com / blog / chatgpt> ), Gemini (Internet search <url: https: gemini.google.com ?hl="ja">) and other generation AIs. The data generation model 58 is obtained by performing deep learning on a neural network. A prompt including an instruction is input to the data generation model 58, and inference data such as voice data indicating voice, text data indicating text, and image data indicating an image is also input. The data generation model 58 performs inference on the input inference data in accordance with the instruction indicated by the prompt, and outputs the inference result in a data format such as voice data and text data. Here, inference refers to, for example, analysis, classification, prediction, and / or summarization.

[1666] In the above embodiment, an example was given in which the specific processing is performed by the data processing device 12, but the technology of the present disclosure is not limited to this, and the specific processing may be performed by the headset type terminal 314.

[1667] [Fourth embodiment]

[1668] FIG. 7 shows an example of the configuration of a data processing system 410 according to the fourth embodiment.

[1669] 7, a data processing system 410 includes a data processing device 12 and a robot 414. An example of the data processing device 12 is a server.

[1670] The data processing device 12 includes a computer 22, a database 24, and a communication I / F 26. The computer 22 is an example of a "computer" according to the technology of the present disclosure. The computer 22 includes a processor 28, a RAM 30, and a storage 32. The processor 28, the RAM 30, and the storage 32 are connected to a bus 34. The database 24 and the communication I / F 26 are also connected to the bus 34. The communication I / F 26 is connected to a network 54. Examples of the network 54 include a WAN (Wide Area Network) and / or a LAN (Local Area Network).

[1671] The robot 414 includes a computer 36, a microphone 238, a speaker 240, a camera 42, a communication I / F 44, and a control target 443. The computer 36 includes a processor 46, a RAM 48, and a storage 50. The processor 46, the RAM 48, and the storage 50 are connected to a bus 52. The microphone 238, the speaker 240, the camera 42, and the control target 443 are also connected to the bus 52.

[1672] The microphone 238 receives instructions and the like from the user 20 by receiving voice uttered by the user 20. The microphone 238 captures the voice uttered by the user 20, converts the captured voice into audio data, and outputs it to the processor 46. The speaker 240 outputs audio in accordance with instructions from the processor 46.

[1673] Camera 42 is a small digital camera equipped with an optical system including a lens, aperture, and shutter, and an imaging element such as a CMOS (Complementary Metal-Oxide-Semiconductor) image sensor or a CCD (Charge Coupled Device) image sensor, and captures images of the surroundings of user 20 (for example, an imaging range defined by an angle of view equivalent to the field of vision of a typical healthy person).

[1674] The communication I / F 44 is connected to a network 54. The communication I / Fs 44 and 26 control the exchange of various information between the processor 46 and the processor 28 via the network 54. The exchange of various information between the processor 46 and the processor 28 using the communication I / Fs 44 and 26 is carried out in a secure state.

[1675] The control object 443 includes a display device, LEDs in the eyes, and motors for driving the arms, hands, and feet. The posture and gestures of the robot 414 are controlled by controlling the motors of the arms, hands, and feet. Some of the emotions of the robot 414 can be expressed by controlling these motors. In addition, the facial expressions of the robot 414 can also be expressed by controlling the light emission state of the LEDs in the eyes of the robot 414.

[1676] Fig. 8 shows an example of the main functions of the data processing device 12 and the robot 414. As shown in Fig. 8, in the data processing device 12, a specific process is performed by the processor 28. A specific process program 56 is stored in the storage 32.

[1677] The specific processing program 56 is an example of a "program" according to the technology of the present disclosure. The processor 28 reads the specific processing program 56 from the storage 32 and executes the read specific processing program 56 on the RAM 30. The specific processing is realized by the processor 28 operating as a specific processing unit 290 in accordance with the specific processing program 56 executed on the RAM 30.

[1678] The storage 32 stores a data generation model 58 and an emotion identification model 59. The data generation model 58 and the emotion identification model 59 are used by the identification processing unit 290.

[1679] In the robot 414, the processor 46 performs the reception output process. A reception output program 60 is stored in the storage 50. The processor 46 reads the reception output program 60 from the storage 50 and executes the read reception output program 60 on the RAM 48. The reception output process is realized by the processor 46 operating as the control unit 46A in accordance with the reception output program 60 executed on the RAM 48.

[1680] Next, a description will be given of the specific processing performed by the specific processing unit 290 of the data processing device 12. In the following description, the data processing device 12 will be referred to as a "server" and the robot 414 will be referred to as a "terminal."

[1681] This invention relates to a system for detecting fraudulent voice data generated by a voice generation AI in real time and ensuring the safety of users. Specific embodiments of the system are described below.

[1682] System Configuration

[1683] The system consists of the following major components:

[1684] 1. Device: The device that captures the audio data (e.g., a smartphone or PC)

[1685] 2. Server: The unit that processes and analyzes the captured audio data.

[1686] 3. User: A person who uses the system and checks alerts

[1687] Specific processing of the program

[1688] Capture and transmit audio data

[1689] When a device starts a voice call or message transmission, it immediately captures the voice data, which is then sent to the server in real time.

[1690] Examples:

[1691] When the user starts a call, the terminal automatically starts monitoring the voice data and sends that data to the server.

[1692] Audio data preprocessing

[1693] The server performs the necessary pre-processing on the received audio data, which includes the following steps:

[1694] Noise Reduction: Reduce background noise from audio data.

[1695] Normalization: Adjusting audio levels to a certain standard.

[1696] Acoustic feature extraction

[1697] The server extracts acoustic features from the preprocessed speech data, including Mel-Frequency Cepstral Coefficients (MFCCs) and spectral features.

[1698] Examples:

[1699] The server calculates MFCCs from the noise-removed and normalized audio data and extracts them as features.

[1700] Comparison with audio profiles

[1701] The server compares the extracted acoustic features with existing voice profiles using statistical methods such as cosine similarity and Euclidean distance.

[1702] Examples:

[1703] The server compares the voice samples previously registered by the user and calculates how well the current voice data matches.

[1704] Anomaly detection and alerting

[1705] If the comparison detects any abnormal changes or signs of unnatural synthesis, the server generates an alert, which is immediately sent to the user via SMS, email, or in-app notification.

[1706] Examples:

[1707] The server detects unnatural voice changes and sends a warning message to the user's smartphone.

[1708] Saving logs

[1709] All processing results and alert information are stored as logs by the server and used for later analysis and auditing, which increases the transparency and reliability of the system.

[1710] In this way, the system analyzes the user's voice data in real time, detects fraudulent voice production with high accuracy, and protects the user's safety.

[1711] The processing flow will be explained below.

[1712] Step 1:

[1713] The device captures voice data in real time when the user initiates a call or voice message.

[1714] Save the audio data in a specified format (e.g. WAV, MP3).

[1715] Step 2:

[1716] The device transmits the captured audio data to the server.

[1717] Encrypted communications are used for data transmission to ensure data security.

[1718] Step 3:

[1719] The server stores the received audio data and starts pre-processing.

[1720] Noise Reduction: Performs a filtering process to reduce background noise.

[1721] Normalize: Adjust the audio level to a certain standard.

[1722] Step 4:

[1723] The server extracts acoustic features from the preprocessed speech data.

[1724] Calculates Mel-Frequency Cepstral Coefficients (MFCC).

[1725] Analyze the spectral features.

[1726] Step 5:

[1727] The server compares the extracted acoustic features with existing voice profiles.

[1728] Statistical methods such as cosine similarity or Euclidean distance are used.

[1729] Step 6:

[1730] The server analyzes the comparison results to detect any unnatural voice changes or signs of synthesis.

[1731] Apply anomaly detection algorithms and flag any anomalies.

[1732] Step 7:

[1733] The server generates an alert if an anomaly is detected.

[1734] The alert content includes a warning message and details of the abnormality detection.

[1735] Step 8:

[1736] The server notifies the user of the generated alert.

[1737] Notification methods include SMS, email, or in-app notifications.

[1738] Step 9:

[1739] The user receives and confirms the notification.

[1740] If necessary, take appropriate action (e.g., report, suspend use of services).

[1741] Step 10:

[1742] The server stores all processing results and alert information as logs.

[1743] The saved logs can be used for later analysis and auditing.

[1744] The above is the specific processing flow of this system.

[1745] Example 1

[1746] Next, a description will be given of Example 1. In the following description, the data processing device 12 will be referred to as a "server" and the robot 414 will be referred to as a "terminal."

[1747] In current voice communication systems, the generation of fraudulent voice data using generative AI models is a problem, threatening user safety. Effective methods for detecting such fraudulent voice generation in real time and immediately notifying users are needed. However, existing systems lack sufficient reliability and speed because they are unable to integrate multiple processes, such as voice data capture, preprocessing, anomaly detection, alert generation, and log storage.

[1748] The specific processing by the specific processing unit 290 of the data processing device 12 in the first embodiment is realized by the following means.

[1749] In this invention, the server includes means for capturing voice data, means for preprocessing the captured voice data, means for extracting acoustic features from the preprocessed voice data, means for comparing the extracted acoustic features with an existing voice profile, means for analyzing the comparison result to detect anomalies, means for generating and notifying an alert when an anomaly is detected, and means for storing the processing results and alert information in a database. This enables highly accurate real-time detection and immediate notification of improper voice generation, and recording of all processing results.

[1750] "Audio data" refers to digital or analog data that records sound.

[1751] "Capture" is the process of acquiring audio data.

[1752] "Preprocessing" refers to processing performed to make captured audio data easier to analyze, and includes operations such as noise removal and normalization.

[1753] "Acoustic features" are numerical representations of the characteristics of speech data, and include Mel-Frequency Cepstrum Coefficients (MFCCs) and spectral features.

[1754] An "audio profile" is reference data constructed using features of pre-registered audio data.

[1755] "Comparison" is a process of evaluating the extracted acoustic features and the features of the voice profile using statistical methods.

[1756] "Anomaly detection" is the process of analyzing the comparison results to find incorrect speech production or unusual changes.

[1757] An "alert" is a warning or notification that is generated when an abnormality is detected.

[1758] "Notification" is the process of communicating generated alerts to users, including by means of SMS, email, in-app notifications, etc.

[1759] A "database" is a system for storing processing results and alert information.

[1760] "Storage" is the process of recording processing results and alert information in a database.

[1761] This invention relates to a system that detects fraudulent voice data generation by voice generation AI in real time and ensures user safety.

[1762] System Configuration

[1763] The system consists of the following major components:

[1764] 1. Device: The device that captures the audio data (e.g., a smartphone or PC)

[1765] 2. Server: The unit that processes and analyzes the captured audio data.

[1766] 3. User: A person who uses the system and checks alerts

[1767] Specific processing of the program

[1768] Capture and transmit audio data

[1769] When a user initiates a voice call or sends a message, the device captures the voice data using the built-in microphone and transmits it to the server in real time. The hardware used is the microphone of a smartphone or PC, and the software is a dedicated voice capture application (e.g., a homemade VoiceCaptureApp).

[1770] Specific behavior:

[1771] The device automatically starts capturing audio when the user starts a call and sends the data to a remote server using SSL / TLS encryption, ensuring data security.

[1772] Audio data preprocessing

[1773] The server performs preprocessing on the received audio data, which mainly includes noise removal and voice level normalization. Specifically, it applies a filter to the audio data to remove background noise, and then normalizes the voice level.

[1774] Specific behavior:

[1775] The server uses Python's SciPy library to apply a bandpass filter to remove background noise.

[1776] After noise removal, the audio levels are normalized to the range 0 to 1 using the normalize function from the Librosa library.

[1777] Acoustic feature extraction

[1778] The server extracts acoustic features from the preprocessed speech data, including Mel-Frequency Cepstral Coefficients (MFCCs) and spectral features.

[1779] Specific behavior:

[1780] The server extracts MFCCs from the audio data using the mfcc function in the Librosa library and stores them in a database as acoustic features.

[1781] Comparison with audio profiles

[1782] The server compares the extracted acoustic features with pre-trained voice profiles using cosine similarity or Euclidean distance.

[1783] Specific behavior:

[1784] The server uses the cosine function in the SciPy library to calculate the similarity between the features of the current voice data and the features of the pre-registered voice profile.

[1785] If the calculation results below a certain threshold, it is flagged as a likely incorrect speech production.

[1786] Anomaly detection and alerting

[1787] If the server detects an anomaly in the audio data, it generates an alert and notifies the user immediately via SMS, email, or in-app notification.

[1788] Specific behavior:

[1789] The server uses an anomaly detection algorithm (e.g., Scikit-learn's Isolation Forest) to detect unnatural changes in the audio data.

[1790] If an anomaly is detected, a warning message is sent to the user's smartphone using Twilio's SMS API, and email notifications are sent using the SendGrid API.

[1791] Saving logs

[1792] The server stores all processing results and alert information in a database, allowing for later analysis and auditing, improving system reliability.

[1793] Specific behavior:

[1794] The server stores the processing results and alert information in a MySQL database using SQLAlchemy to create database entries.

[1795] Example prompts to input to the generative AI model

[1796] "Please explain how speech generation AI detects fraudulent audio data. Please include specific steps for capturing audio data, preprocessing, extracting acoustic features, comparing with audio profiles, detecting anomalies, generating alerts, and saving logs."

[1797] In this way, the system can analyze the user's voice data in real time, detect fraudulent voice production with high accuracy, and ensure the user's safety.

[1798] The flow of the identification process in the first embodiment will be described with reference to FIG.

[1799] Processing Steps

[1800] Step 1: Capture and send audio data

[1801] The device uses a built-in microphone to capture voice data when a user initiates a voice call or sends a message.

[1802] input:

[1803] Initiate voice calls and send messages

[1804] Specific behavior:

[1805] The device starts a dedicated voice capture application (e.g., a custom VoiceCaptureApp) and starts microphone input. The captured voice data is sent to the server using encryption (SSL / TLS).

[1806] output:

[1807] Encrypted and transmitted voice data

[1808] Step 2: Preprocessing the audio data

[1809] The server performs preprocessing on the received audio data.

[1810] input:

[1811] Received audio data

[1812] Specific behavior:

[1813] The server uses Python's SciPy library to apply a bandpass filter to remove background noise, then normalizes the audio level to the range 0 to 1 using the normalize function from the Librosa library.

[1814] output:

[1815] Denoised and normalized audio data

[1816] Step 3: Extraction of acoustic features

[1817] The server extracts acoustic features from the preprocessed speech data.

[1818] input:

[1819] Denoised and normalized audio data

[1820] Specific behavior:

[1821] The server extracts MFCCs using the mfcc function in the Librosa library and stores them in a database as acoustic features.

[1822] output:

[1823] Extracted acoustic features

[1824] Step 4: Compare with the audio profile

[1825] The server compares the extracted acoustic features with pre-registered voice profiles.

[1826] input:

[1827] Extracted acoustic features

[1828] Pre-registered audio profiles

[1829] Specific behavior:

[1830] The server uses the cosine function in the SciPy library to calculate the similarity between the features of the current voice data and the features of the existing voice profile.

[1831] output:

[1832] Similarity calculation results

[1833] Step 5: Detect anomalies and generate alerts

[1834] If the server detects an abnormality in the voice data, it generates an alert and immediately notifies the user.

[1835] input:

[1836] Similarity calculation results

[1837] Specific behavior:

[1838] The server uses an anomaly detection algorithm (e.g., Scikit-learn's Isolation Forest) to detect anomalies. If an anomaly is detected, it uses Twilio's SMS API to send an alert message to the user's smartphone. For email notifications, it uses the SendGrid API.

[1839] output:

[1840] Warning message

[1841] Step 6: Save the logs

[1842] The server stores all processing results and alert information in a database.

[1843] input:

[1844] Similarity calculation results

[1845] Warning message

[1846] Specific behavior:

[1847] The server uses SQLAlchemy to create and save database entries to store the processing results and alert information in a MySQL database.

[1848] output:

[1849] Saved processing results and alert information

[1850] (Application example 1)

[1851] Next, a description will be given of Application Example 1. In the following description, the data processing device 12 will be referred to as a "server" and the robot 414 will be referred to as a "terminal."

[1852] With the advancement of AI voice generation technology, it has become easier to generate voice data fraudulently. As a result, the risk of personal and corporate voice communications being used fraudulently and important call content being tampered with is increasing. This has a particularly severe impact on business communications that contain confidential information. Current systems face the challenge of being unable to detect such fraudulent activity in real time and respond immediately.

[1853] The specific processing by the specific processing unit 290 of the data processing device 12 in the application example 1 is realized by the following means.

[1854] In this invention, the server includes means for capturing voice data, means for preprocessing the captured voice data, means for extracting acoustic features from the preprocessed voice data, means for comparing the extracted acoustic features with a voice profile, means for detecting anomalies based on the analysis results, means for generating and notifying an alarm when an anomaly is detected, means for recording the alarm information, means for detecting fraudulent generation of voice data in real time, an application for monitoring call data, and a cloud-based processing server. This makes it possible to detect fraudulent generation of voice data in real time and respond quickly.

[1855] "Audio Data" means data for recording, storing, and transmitting audio in digital format.

[1856] A "capturing means" is a device or software for capturing audio data in real time.

[1857] The "pre-processing means" refers to a method or device for removing noise from the captured audio data and normalizing it.

[1858] "Means for extracting acoustic features" refers to a method or device that calculates and extracts characteristic information (e.g., Mel-Frequency Cepstral Coefficients (MFCCs)) from preprocessed speech data.

[1859] An "audio profile" is a collection of feature data generated based on specific audio data.

[1860] The "comparison means" refers to a method or device for comparing the extracted acoustic features with a voice profile and evaluating the degree of similarity.

[1861] The "analysis result" is information indicating the degree of agreement of the data obtained from the comparison of the acoustic features and the presence or absence of abnormalities.

[1862] "Means for detecting anomalies" refers to a method or device that determines whether there are unnatural changes or signs of fraudulent generation in the audio data based on the analysis results.

[1863] The "means for generating and notifying a warning" refers to a method or device for generating and notifying a warning message to a user when an abnormality is detected.

[1864] The "means for recording warning information" refers to a method or device for saving generated warning messages and analysis results so that they can be checked at a later date.

[1865] "Real-time detection means" refers to a method or device that can detect fraud immediately without interrupting the time when voice data is being generated.

[1866] An "application that monitors call data" is software that runs on a device such as a smartphone and monitors voice data during a call.

[1867] A "cloud-based processing server" is a remote server accessible via the Internet that pre-processes, analyzes, and compares audio data.

[1868] This invention relates to a system that detects fraudulent voice data generated by voice generation AI in real time during phone calls to ensure user security. This system automatically captures voice data, preprocesses it, extracts acoustic features, compares it with existing voice profiles, detects anomalies, generates warnings, and records it.

[1869] System Configuration

[1870] The system consists of the following major components:

[1871] 1. Terminal: A device for capturing audio data (e.g., a smartphone)

[1872] 2. Server: A unit for processing and analyzing the captured audio data. Uses a cloud-based processing server (e.g., AWS, Google Cloud Platform).

[1873] 3. Users: People who use the system and receive alerts

[1874] Specific processing of the program

[1875] Capture audio data

[1876] The device uses the Twilio Voice SDK to capture voice data during a call in real time and immediately transmits that data to the server.

[1877] Audio data preprocessing

[1878] The server uses libraries (e.g., PyDub, Librosa) to denoise and normalize the audio data. This preprocessing produces clean audio data suitable for analysis.

[1879] Acoustic feature extraction

[1880] From the preprocessed audio data, the server uses Librosa to extract acoustic features such as Mel-Frequency Cepstral Coefficients (MFCCs).

[1881] Comparison with audio profiles

[1882] The extracted acoustic features are compared to existing voice profiles using machine learning frameworks such as TensorFlow, using statistical methods such as cosine similarity.

[1883] Anomaly detection and alert generation

[1884] Based on the analysis results, the server determines whether the voice data has any unnatural variations or signs of fraudulent generation, and if an anomaly is detected, the server sends a warning to the user via SMS or in-app notification.

[1885] Recording warning information

[1886] All processing results and warning information are logged by the server for later analysis and auditing.

[1887] Specific examples

[1888] For example, suppose a user is using the Secure Call Detector app to make a call for a business meeting and the call content is tampered with. The system detects this tampering in real time and immediately generates an alert. The user receives the following message on their smartphone:

[1889] "Warning: Indications of unauthorized audio have been detected. The following message will appear on your screen: 'Warning: Unauthorized audio has been detected. Please ensure your call is secure.' Call recording will automatically stop and your system administrator will be notified."

[1890] This allows the user to respond quickly and ensures the safety of the call.

[1891] The system of the present invention can detect fraudulent generation of voice data in real time and respond quickly, thereby enhancing the security of individuals and businesses.

[1892] The flow of the specific processing in the application example 1 will be described with reference to FIG.

[1893] Step 1:

[1894] The terminal uses the Twilio Voice SDK to detect when a user starts a call. When a call starts, it immediately starts capturing voice data and sends that data to the server in real time. The input is the user's call voice data, and the output is the raw voice data sent to the server.

[1895] Step 2:

[1896] The server preprocesses the audio data received from the device using a library (e.g., PyDub, Librosa). This preprocessing removes noise from the audio data and normalizes the audio level. The input is raw audio data, and the output is noise-removed and normalized audio data.

[1897] Step 3:

[1898] The server extracts acoustic features from the preprocessed speech data. It uses Librosa to extract features such as Mel-Frequency Cepstral Coefficients (MFCCs). The input is the preprocessed speech data, and the output is the extracted acoustic features.

[1899] Step 4:

[1900] The server compares the extracted acoustic features with existing voice profiles. This comparison uses TensorFlow's machine learning model to evaluate similarity using statistical methods such as cosine similarity and Euclidean distance. The input is the acoustic features and the voice profile, and the output is a similarity score.

[1901] Step 5:

[1902] The server analyzes the similarity score and determines whether or not unauthorized voice data has been generated. If an abnormality is detected based on the analysis results, a warning message is generated. The input is the similarity score, and the output is the warning message.

[1903] Step 6:

[1904] The server notifies the user via SMS or in-app notification with the generated alert message. This notification is sent using Twilio's messaging API. The input is the alert message, and the output is the alert message displayed on the user's device.

[1905] Step 7:

[1906] The server stores all processing results and warning information in a database. This information is recorded using a database such as MySQL for later analysis and auditing. The input is the processing results and warning information, and the output is the log information stored in the database.

[1907] Furthermore, an emotion engine that estimates the user's emotion may be further combined. That is, the identification processing unit 290 may estimate the user's emotion using the emotion identification model 59, and perform identification processing using the user's emotion.

[1908] This invention relates to a system that ensures the safety and comfort of users by detecting fraudulent voice data generated by voice generation AI in real time and recognizing the emotional state of the user from the voice data. Specific embodiments of the system are described below.

[1909] System Configuration

[1910] The system consists of the following major components:

[1911] 1. Device: The device that captures the audio data (e.g., a smartphone or PC)

[1912] 2. Server: The unit that processes, analyzes, and recognizes emotions from captured audio data.

[1913] 3. User: A person who uses the system and checks alerts

[1914] Specific processing of the program

[1915] Capture and transmit audio data

[1916] The device captures voice data in real time when the user initiates a call or voice message, and the captured voice data is sent to the server via encrypted communication means.

[1917] Examples:

[1918] The terminal starts monitoring the voice data at the same time as the user starts a call, and transmits the voice data to the server via secure communication.

[1919] Audio data preprocessing

[1920] The server performs pre-processing on the received audio data, including noise reduction to reduce background noise and normalization to adjust the audio level to a certain standard.

[1921] Acoustic feature extraction

[1922] The server extracts acoustic features from the preprocessed speech data, including Mel-Frequency Cepstral Coefficients (MFCCs) and spectral features.

[1923] Examples:

[1924] The server calculates MFCCs from the noise-removed and normalized audio data and obtains them as feature data.

[1925] Comparison with audio profiles

[1926] The server compares the extracted acoustic features with existing voice profiles using statistical methods such as cosine similarity and Euclidean distance.

[1927] Examples:

[1928] The server compares the voice samples previously registered by the user and calculates how well the current voice data matches.

[1929] Anomaly detection and alerting

[1930] If the comparison detects unnatural voice changes or signs of synthesis, the server generates an alert, which is promptly sent to the user via SMS, email, or in-app notification.

[1931] Examples:

[1932] The server detects abnormal audio changes and sends a warning message to the user's smartphone.

[1933] Additional analysis with emotion recognition

[1934] The server uses an emotion engine to analyze the user's emotional state from the voice data, and determines whether the voice data matches the user's typical emotional state.

[1935] Examples:

[1936] The server analyzes the voice data to determine whether the user is expressing emotions such as anxiety, anger, or joy.

[1937] Selecting notification methods based on emotional state

[1938] The server selects the optimal notification method based on the detected emotional state. For example, if the user is determined to be in an anxious state, an immediate notification via SMS can be sent to enhance safety.

[1939] Examples:

[1940] Based on the results of the emotion engine, if the user is feeling stressed, the server determines that a prompt response is required and selects an SMS notification.

[1941] Saving logs

[1942] All processing results and alert information are stored as logs by the server, which can be used for later analysis and auditing, ensuring the transparency and reliability of the system.

[1943] In this way, the system analyzes the user's voice data in real time, detects fraudulent voice production with high accuracy, and by recognizing the user's emotional state, selects an appropriate notification method to maintain the user's safety and comfort.

[1944] The processing flow will be explained below.

[1945] Step 1:

[1946] The device automatically captures voice data in real time when the user initiates a call or voice message.

[1947] Save the audio data in a specified format (e.g. WAV, MP3).

[1948] Step 2:

[1949] The device transmits the captured audio data to the server.

[1950] Encrypted communications are used for data transmission to ensure data security.

[1951] Step 3:

[1952] The server stores the received audio data and starts pre-processing.

[1953] The server first applies a noise reduction filter to reduce background noise.

[1954] The server then normalizes the audio levels to ensure consistency.

[1955] Step 4:

[1956] The server extracts acoustic features from the preprocessed speech data.

[1957] The server calculates Mel-Frequency Cepstral Coefficients (MFCCs) and extracts them as speech features.

[1958] The server also analyzes the spectral features and stores them as supplementary features.

[1959] Step 5:

[1960] The server compares the extracted acoustic features with existing voice profiles.

[1961] The server calculates cosine similarity and Euclidean distance to evaluate the degree of matching of the voices.

[1962] Step 6:

[1963] The server analyzes the comparison results to detect any unnatural voice changes or signs of synthesis.

[1964] The server applies an anomaly detection algorithm and flags any anomalies detected.

[1965] Step 7:

[1966] The server generates an alert if an anomaly is detected.

[1967] The server generates alert information including warning messages and details of anomaly detections.

[1968] Step 8:

[1969] The server notifies the user of the generated alert.

[1970] Notification methods include SMS, email, or in-app notifications.

[1971] Step 9:

[1972] The user receives and confirms the notification.

[1973] Based on the content of the notification, the user will take appropriate action as necessary (e.g., report the issue, stop using the service).

[1974] Step 10:

[1975] The server uses an emotion engine to analyze the user's emotional state from the voice data.

[1976] The server identifies the emotional state (e.g., anxiety, anger, joy, etc.) from the voice data.

[1977] Step 11:

[1978] The server selects the most appropriate notification method based on the detected emotional state.

[1979] For example, if a user is in a state of anxiety, they will be notified immediately via SMS to enhance safety.

[1980] Step 12:

[1981] The server stores all processing results and alert information as logs.

[1982] The saved logs can be used for later analysis and auditing.

[1983] This is the specific processing flow of this system. This system allows users to detect fraudulent generation of voice data with high accuracy, and by recognizing the user's emotional state, it enables more appropriate responses.

[1984] Example 2

[1985] Next, a description will be given of Example 2. In the following description, the data processing device 12 will be referred to as a "server" and the robot 414 will be referred to as a "terminal."

[1986] Conventional voice data analysis systems have difficulty detecting fraudulent voice generation in real time, and there are no systems that can recognize the user's emotional state and take appropriate action. This increases the risk of voice data being manipulated or used fraudulently, and poses a challenge to ensuring user safety and comfort.

[1987] The identification process by the identification processing unit 290 of the data processing device 12 in the second embodiment is realized by the following means. In this invention, the server includes means for preprocessing voice data, means for extracting acoustic features from the preprocessed voice data, means for comparing the extracted acoustic features with an existing voice profile, and means for selecting a notification method based on the recognized emotional state. This enables highly accurate detection of fraudulent voice generation and appropriate response based on the user's emotional state.

[1988] "Audio data" is a digital representation of an acoustic signal captured using a microphone or other audio input device.

[1989] "Capture" is the process of capturing audio data in real time via a microphone or other input device.

[1990] "Preprocessing" refers to the process of applying early-stage processing such as noise removal and normalization to speech data to facilitate analysis and feature extraction.

[1991] "Encryption" is a method of converting audio data using a specific algorithm to protect it from third parties.

[1992] "Transmission" refers to the act of transferring captured audio data to another system or server via a network.

[1993] "Acoustic features" are computed physical or statistical properties, such as Mel-Frequency Cepstral Coefficients (MFCCs) or spectral features, extracted from speech data.

[1994] "Comparison" is a process in which the extracted acoustic features are matched with an existing data set called a voice profile, and the degree of match is evaluated.

[1995] "Anomaly detection" is a function that detects and reports unnatural voice changes and traces of synthesized voice as a result of voice data analysis.

[1996] An "alert" is a message or signal that notifies a user or system administrator when an abnormality is detected.

[1997] "Emotion recognition" is a technology that analyzes and determines a user's emotional state (e.g., joy, anxiety, anger, etc.) from voice data.

[1998] "Notification method selection" is the process of selecting the optimal method (e.g., SMS, email, in-app notification, etc.) to notify the user based on the recognized emotional state.

[1999] This invention relates to a system that ensures the safety and comfort of users by detecting fraudulent voice data generated by voice generation AI in real time and recognizing the emotional state of the user from the voice data. Specific embodiments are described below.

[2000] System Configuration

[2001] The system consists of the following major components:

[2002] 1. Device: The device that captures the audio data (e.g., a smartphone or computer)

[2003] 2. Server: The unit that processes, analyzes, and recognizes emotions from captured audio data.

[2004] 3. User: A person who uses the system and checks alerts

[2005] Capture and transmit audio data

[2006] When a user initiates a call or voice message, the device captures audio data in real time through the microphone and transmits the captured audio data to the server using SSL / TLS encryption.

[2007] Examples:

[2008] The device collects the audio as soon as the user starts a call on their smartphone and sends it to the server as encrypted data.

[2009] Audio data preprocessing

[2010] The server performs pre-processing on the received audio data, including removing background noise using a noise attenuation algorithm (e.g., Wiener filter) and normalizing the volume of the audio data.

[2011] Examples:

[2012] The server applies a Wiener filter to the received audio data to reduce background noise and normalize the audio level to 0.5.

[2013] Acoustic feature extraction

[2014] From the preprocessed audio data, the server extracts Mel-Frequency Cepstral Coefficients (MFCCs) and spectral features using the Python library Librosa.

[2015] Examples:

[2016] The server extracts 13-dimensional MFCC features from the preprocessed audio data using Librosa and stores them in memory as an array.

[2017] Comparison with audio profiles

[2018] The server compares the extracted acoustic features with the voice profile registered by the user in advance, and detects fraudulent voices by calculating cosine similarity and Euclidean distance using SciPy to evaluate the degree of match.

[2019] Examples:

[2020] The server passes the generated MFCC features to the SciPy cosine similarity function to compare the similarity with pre-registered voice profiles. It also calculates the Euclidean distance to evaluate the match from different angles.

[2021] Anomaly detection and alerting

[2022] If the comparison detects unnatural voice changes or signs of synthetic speech, the server detects the anomaly and generates an alert. The alert is sent via email using the SMTP protocol and also via push notifications within the app.

[2023] Examples:

[2024] If the cosine similarity is below a certain level, the server determines that an abnormality has occurred and sends a warning message to the user's email address via SMTP. At the same time, a warning message is also sent to the user's smartphone via push notification.

[2025] Additional analysis with emotion recognition

[2026] The server uses a voice emotion recognition engine (e.g., Google Cloud Speech-to-Text API) to analyze the user's emotional state from the voice data and determine whether it matches their normal emotional state.

[2027] Examples:

[2028] The server uses the Google Cloud Speech-to-Text API to perform emotion recognition and detect emotions such as "anger," "sadness," and "joy." The detection results are saved in tensor format.

[2029] Selecting notification methods based on emotional state

[2030] The server selects the optimal notification method based on the detected emotional state. For example, if the user is determined to be in an "anxious" state, it will send an SMS notification immediately using the Twilio API.

[2031] Examples:

[2032] If the server detects an "uneasy" state, it will use the Twilio API to send an SMS notification with the message "urgent action required."

[2033] Saving logs

[2034] All processing results and generated alert information are stored as logs in a PostgreSQL database by the server, which can be used for later analysis and auditing, increasing the transparency and security of the system.

[2035] Examples:

[2036] The server inserts all data, including the results of speech preprocessing, acoustic features, emotion recognition results, and alert information, into a PostgreSQL database and stores it along with a timestamp.

[2037] Prompt Sentence Examples

[2038] As an example of a prompt for this system, the following query is input to the generative AI model:

[2039] "When a user sends a voice message, we want to identify the emotion from that data and activate fraudulent voice detection. Please suggest an algorithm."

[2040] The flow of the identification process in the second embodiment will be described with reference to FIG.

[2041] Step 1: Capture audio data

[2042] When a user initiates a call or voice message, the device uses the built-in microphone to capture voice data in real time, which is then recorded in WAV format and temporarily stored in the device's memory.

[2043] Specific behavior:

[2044] The moment a user starts a call on their smartphone, the device captures the audio data at a sampling rate of 44.1 kHz and saves it as a temporary file.

[2045] input:

[2046] Call start triggers

[2047] output:

[2048] Audio data captured by the built-in microphone (WAV format)

[2049] Step 2: Sending audio data

[2050] The device protects the captured audio data with AES-256 encryption and sends it to the server using SSL / TLS, preventing data leakage during transmission.

[2051] Specific behavior:

[2052] The device encrypts the voice data with AES-256 and transmits it in real time to the server via the HTTPS protocol.

[2053] input:

[2054] Captured audio data (WAV format)

[2055] output:

[2056] The encrypted audio data is sent to the server

[2057] Step 3: Preprocessing the audio data

[2058] The server applies a noise attenuation algorithm (e.g., a Wiener filter) to the received audio data to normalize the audio level.

[2059] Specific behavior:

[2060] The server applies a Wiener filter to the received audio data to reduce noise and normalize the volume level of the audio data to 0.5.

[2061] input:

[2062] Encrypted and decrypted audio data

[2063] output:

[2064] Noise-reduced and normalized audio data

[2065] Step 4: Extraction of acoustic features

[2066] The server extracts Mel-Frequency Cepstral Coefficients (MFCCs) and spectral features from the preprocessed audio data using the Librosa library.

[2067] Specific behavior:

[2068] The server uses Librosa to calculate 13-dimensional MFCC features from the preprocessed audio data and saves them as an array.

[2069] input:

[2070] Normalized audio data

[2071] output:

[2072] Extracted acoustic features (MFCC)

[2073] Step 5: Compare with the audio profile

[2074] The server uses SciPy to calculate the cosine similarity and Euclidean distance between the extracted acoustic features and the voice profile previously registered by the user, and evaluates the degree of match.

[2075] Specific behavior:

[2076] The server passes the generated MFCC features to SciPy's cosine similarity function to compare them with existing audio profiles and calculates the degree of match. It also calculates the Euclidean distance.

[2077] input:

[2078] Extracted acoustic features

[2079] Existing Audio Profiles

[2080] output:

[2081] Evaluation results of match (cosine similarity and Euclidean distance)

[2082] Step 6: Detect anomalies and generate alerts

[2083] The server determines whether there are any invalid voices or abnormalities based on the results of the evaluation of the degree of matching of the voice features, and generates an alert if necessary. The generated alert is sent via email using the SMTP protocol, and also sent in-app via push notification.

[2084] Specific behavior:

[2085] The server detects anomalies based on criteria such as a cosine similarity of 0.7 or less, and sends a warning message to the user's email address via SMTP. At the same time, a warning message is also sent to the user's smartphone via push notification.

[2086] input:

[2087] Matching evaluation results

[2088] output:

[2089] An alert is generated and an email and push notification is sent

[2090] Step 7: Further analysis with emotion recognition

[2091] The server analyzes the user's emotional state from the voice data using an emotion recognition engine (e.g., Google Cloud Speech-to-Text API), and the analyzed emotional state is stored in a database.

[2092] Specific behavior:

[2093] The server sends the voice data to the Google Cloud Speech-to-Text API, which performs emotion recognition and detects emotions such as "anger," "sadness," and "joy."

[2094] input:

[2095] Normalized audio data

[2096] output:

[2097] Detected emotional state

[2098] Step 8: Select notification method based on emotional state

[2099] The server selects the optimal notification method based on the detection result. For example, if the user is determined to be in an "anxious state," it will immediately send an SMS notification using the Twilio API.

[2100] Specific behavior:

[2101] If the emotion recognition result is "anxiety," the server uses the Twilio API to quickly send an SMS containing the message "urgent action required."

[2102] input:

[2103] Detected emotional state

[2104] output:

[2105] Messages sent via the most appropriate notification method

[2106] Step 9: Save the logs

[2107] The server logs all processing results and generated alert information in a PostgreSQL database for later analysis and auditing.

[2108] Specific behavior:

[2109] The server stores all data, including audio preprocessing results, acoustic features, emotion recognition results, and alert information, in a PostgreSQL database and records them with timestamps.

[2110] input:

[2111] Various processing results and alert information

[2112] output:

[2113] Logs stored in the database

[2114] (Application example 2)

[2115] Next, a description will be given of Application Example 2. In the following description, the data processing device 12 will be referred to as a "server" and the robot 414 will be referred to as a "terminal."

[2116] Conventional voice recognition systems have difficulty detecting fraudulent voice data generated by voice generation AI in real time, making it impossible to ensure user safety. Furthermore, they lack the ability to recognize the user's emotional state, making it difficult to respond appropriately based on the user's psychological state. Therefore, there is a demand for a system that combines voice data anomaly detection and user emotion recognition.

[2117] The identification process by the identification processing unit 290 of the data processing device 12 in Application Example 2 is realized by the following means. In this invention, the server includes means for capturing voice data, means for preprocessing the captured voice data, means for comparing extracted acoustic features with existing voice profiles, means for generating and notifying an alert when an abnormality is detected, means for analyzing the user's emotional state from the voice data, and means for selecting a notification method for the alert based on the analysis results. This makes it possible not only to detect abnormalities in voice data in real time, but also to provide an appropriate notification according to the user's emotional state.

[2118] "Voice data" refers to data that records a user's voice information in digital format.

[2119] "Capturing means" refers to equipment or technology used to record or capture audio data.

[2120] "Preprocessing means" refers to the techniques and processes that perform preprocessing such as noise removal and normalization on the captured audio data.

[2121] "Acoustic features" are acoustic feature information extracted from speech data, and include Mel-Frequency Cepstrum Coefficients (MFCCs).

[2122] A "profile" is a database that records information about specific users and their voice characteristics.

[2123] "Means of comparison" refers to the technology or algorithm that matches the extracted acoustic features with existing voice profiles.

[2124] "Means for detecting anomalies" refers to technologies and algorithms that analyze the results of comparing acoustic features to identify fraudulent audio data.

[2125] "Means for generating and notifying alerts" refers to the technology and methods for generating a warning message and notifying the user when an abnormality is detected.

[2126] "Means for analyzing emotional state" refers to technologies and algorithms for recognizing a user's emotions from voice data and assessing their state.

[2127] "Means for selecting a notification method" refers to a technology or method for selecting the most appropriate alert notification method (e.g., SMS, email, etc.) based on the analyzed emotional state.

[2128] MODE FOR CARRYING OUT THE INVENTION

[2129] This invention relates to a system that detects fraudulent voice data generated by a voice generation AI in real time and analyzes the user's emotional state from the voice data to ensure the safety and comfort of the user. Specific embodiments of the system are described below.

[2130] System Configuration

[2131] The system consists of the following major components:

[2132] 1. Terminal

[2133] The terminal is a device that captures the user's voice data. This includes smartphones, tablets, PCs, etc. When the user makes a call or sends a voice message, the voice data is captured in real time and sent to the server via a secure communication method.

[2134] 2. Server

[2135] The server is a unit that processes, analyzes, and recognizes emotions from received voice data. It is mainly responsible for the following tasks:

[2136] Audio data preprocessing

[2137] Acoustic feature extraction

[2138] Comparison with audio profiles

[2139] Anomaly detection and alerting

[2140] emotion recognition

[2141] Selecting notification methods based on emotional state

[2142] Saving logs

[2143] 3. Users

[2144] The user is the person who interacts with the system and checks the alerts: they receive alerts when an anomaly is detected or when a particular emotional state is recognized.

[2145] Specific processing of the program

[2146] This system performs a series of processes in real time, from capturing voice data to analyzing it and generating alerts. The specific process is explained below.

[2147] 1. Capture and transmit audio data

[2148] The device captures voice data in real time when the user initiates a call or voice message, and the captured voice data is sent to the server via encrypted communication means.

[2149] 2. Preprocessing of audio data

[2150] The server performs preprocessing on the received audio data, including noise reduction to reduce background noise and normalization to adjust the audio level to a certain standard. This processing is performed using the Librosa library.

[2151] 3. Acoustic feature extraction

[2152] From the preprocessed audio data, the server extracts acoustic features, including Mel-Frequency Cepstral Coefficients (MFCCs) and spectral features, also using the Librosa library.

[2153] 4. Comparison with audio profiles

[2154] The server compares the extracted acoustic features with existing voice profiles using statistical methods such as cosine similarity and Euclidean distance, and generates an alert if an anomaly is detected.

[2155] 5. Anomaly detection and alert generation

[2156] If the comparison detects unnatural voice changes or signs of synthesis, the server generates an alert, which is promptly sent to the user via SMS, email, or in-app notification.

[2157] 6. Additional analysis using emotion recognition

[2158] The server uses an emotion recognition engine to analyze the user's emotional state from the voice data. Based on the analysis results, it determines whether the voice data matches the user's normal emotional state. DeepMoji and other emotion recognition APIs are used for emotion recognition.

[2159] 7. Notification method selection based on emotional state

[2160] The server selects the optimal notification method based on the detected emotional state. For example, if the user is determined to be in an anxious state, an immediate notification via SMS can be sent to enhance safety.

[2161] Specific examples

[2162] Specific application examples of this system are shown below.

[2163] Example 1:

[2164] If fraudulent voice generation occurs while a user is making an important call at work, the server will detect the anomaly and send a warning notification to the user's smartphone.

[2165] Example 2:

[2166] If a user feels stressed while talking to a family member, the server will recognize the emotion and send a notification encouraging the user to "relax."

[2167] Example prompts for generative AI models

[2168] "What is the execution flow of the application notification when fraudulent voice is detected while a user is on a call with a bank's call center?"

[2169] The flow of the specific processing in the application example 2 will be described with reference to FIG.

[2170] Program processing steps

[2171] Step 1: Capture and send audio data

[2172] Input: A user initiates a call or voice message.

[2173] Specific operation: The device uses the built-in microphone to capture audio data in real time.

[2174] Data processing: Captured audio data is encrypted immediately.

[2175] Output: Encrypted audio data is generated.

[2176] Processing flow: The device sends encrypted audio data to the server via a secure communication method.

[2177] Step 2: Preprocessing the audio data

[2178] Input: Encrypted audio data is sent to the server.

[2179] Specific operation: The server decodes the received audio data.

[2180] Data processing: Denoising and normalization are performed on the decoded audio data using the Librosa library.

[2181] Output: Preprocessed audio data is generated.

[2182] Processing flow: The server sends the noise-removed and normalized audio data to the next step.

[2183] Step 3: Extraction of acoustic features

[2184] Input: Preprocessed audio data is sent to the server.

[2185] Specific operation: The server uses the Librosa library to calculate Mel-Frequency Cepstral Coefficients (MFCCs).

[2186] Data processing: MFCC and spectral features are extracted from the audio data.

[2187] Output: The extracted acoustic features are generated.

[2188] Processing flow: The server sends the extracted features to the next step.

[2189] Step 4: Compare with the audio profile

[2190] Input: The extracted acoustic features are sent to the server.

[2191] Specific behavior: The server loads the existing audio profile.

[2192] Data analysis: Compare acoustic features and voice profiles using cosine similarity and Euclidean distance.

[2193] Output: The comparison results are generated.

[2194] Process flow: The server sends the comparison result to the next step.

[2195] Step 5: Detect anomalies and generate alerts

[2196] Input: The comparison result is sent to the server.

[2197] Specific operation: The server analyzes the comparison results and determines whether an anomaly is detected.

[2198] Data judgment: If an abnormality is detected, an alert message is generated.

[2199] Output: An alert message is generated.

[2200] Process flow: The server sends an alert message to the user's terminal.

[2201] Step 6: Further analysis with emotion recognition

[2202] Input: Preprocessed audio data is sent to the server.

[2203] Specific operation: The server analyzes the voice data using an emotion recognition engine.

[2204] Data analysis: Implementing specific algorithms to identify the user's emotional state from the audio data.

[2205] Output: The emotion recognition results are generated.

[2206] Processing flow: The server refers to the emotion recognition results and sends them to the next step.

[2207] Step 7: Select notification method based on emotional state

[2208] Input: Emotion recognition results are sent to the server.

[2209] Specific operation: The server selects the optimal notification method based on the emotional state.

[2210] Data judgment: Apply the conditions to select the notification method (e.g., SMS notification if the emotional state is stressed).

[2211] Output: The best notification method is selected.

[2212] Process flow: The server notifies the user of the alert according to the selected notification method.

[2213] The specific processing unit 290 transmits the result of the specific processing to the robot 414. In the robot 414, the control unit 46A causes the speaker 240 and the control target 443 to output the result of the specific processing. The microphone 238 acquires voice indicating a user input regarding the result of the specific processing. The control unit 46A transmits voice data indicating the user input acquired by the microphone 238 to the data processing device 12. In the data processing device 12, the specific processing unit 290 acquires the voice data.

[2214] The data generation model 58 is a so-called generative AI (Artificial Intelligence). An example of the data generation model 58 is ChatGPT (Internet Search<URL: https: / / openai.com / blog / chatgpt> ), Gemini (Internet search <url: https: gemini.google.com ?hl="ja">) and other generation AIs. The data generation model 58 is obtained by performing deep learning on a neural network. A prompt including an instruction is input to the data generation model 58, and inference data such as voice data indicating voice, text data indicating text, and image data indicating an image is also input. The data generation model 58 performs inference on the input inference data in accordance with the instruction indicated by the prompt, and outputs the inference result in a data format such as voice data and text data. Here, inference refers to, for example, analysis, classification, prediction, and / or summarization.

[2215] In the above embodiment, an example was given in which the specific processing is performed by the data processing device 12, but the technology of the present disclosure is not limited to this, and the specific processing may be performed by the robot 414.

[2216] The emotion identification model 59 as an emotion engine may determine the user's emotion according to a specific mapping. Specifically, the emotion identification model 59 may determine the user's emotion according to an emotion map (see FIG. 9), which is a specific mapping. Similarly, the emotion identification model 59 may determine the robot's emotion, and the identification processing unit 290 may perform identification processing using the robot's emotion.

[2217] FIG. 9 is a diagram illustrating an emotion map 400 on which multiple emotions are mapped. In the emotion map 400, emotions are arranged in concentric circles radiating from the center. Emotions closer to the center of the concentric circles are more primitive. Emotions representing states and actions arising from a state of mind are arranged on the outer edges of the concentric circles. The concept of emotion includes both affect and mental states. Emotions generally generated from reactions occurring in the brain are arranged on the left side of the concentric circles. Emotions generally induced by situational judgment are arranged on the right side of the concentric circles. Emotions generally generated from reactions occurring in the brain and induced by situational judgment are arranged on the upper and lower sides of the concentric circles. Furthermore, the emotion of "pleasure" is arranged on the upper side of the concentric circles, and the emotion of "discomfort" is arranged on the lower side. In this way, in the emotion map 400, multiple emotions are mapped based on the structure by which emotions are generated, and emotions that tend to occur simultaneously are mapped close to each other.

[2218] These emotions are distributed in the 3 o'clock direction on emotion map 400, and typically fluctuate between relief and anxiety. In the right half of emotion map 400, situational awareness dominates over internal sensations, resulting in a sense of calm.

[2219] The inside of emotion map 400 represents what is going on in the mind, and the outside of emotion map 400 represents behavior, so the further you go outside emotion map 400, the more visible the emotions become (the more they are expressed in behavior).

[2220] Human emotions are based on various balances, such as posture and blood sugar levels. When these balances deviate from the ideal, a state of discomfort is indicated, and when they approach the ideal, a state of pleasure is indicated. Emotions can also be created for robots, automobiles, and motorcycles, based on various balances, such as posture and remaining battery life. When these balances deviate from the ideal, a state of discomfort is indicated, and when they approach the ideal, a state of pleasure is indicated. An emotion map can be generated, for example, based on Dr. Mitsuyoshi's emotion map (Research on Voice Emotion Recognition and Emotional Brain Physiological Signal Analysis Systems, Tokushima University, Doctoral Dissertation: https: / / ci.nii.ac.jp / naid / 500000375379). The left half of the emotion map lists emotions belonging to the "reaction" domain, where sensation is dominant. The right half of the emotion map lists emotions belonging to the "situation" domain, where situational awareness is dominant.

[2221] The emotion map defines two emotions that promote learning. One is a negative emotion on the situation side, around the middle of "repentance" or "reflection." In other words, this occurs when the robot experiences negative emotions such as "I never want to feel this way again" or "I don't want to be scolded again." The other is a positive emotion on the response side, around "desire." In other words, this occurs when the robot experiences positive feelings such as "I want more" or "I want to know more."

[2222] The emotion identification model 59 inputs user input into a pre-trained neural network, obtains emotion values ​​indicating each emotion shown in the emotion map 400, and determines the user's emotion. This neural network is pre-trained based on multiple pieces of training data that are combinations of user input and emotion values ​​indicating each emotion shown in the emotion map 400. Furthermore, this neural network is trained so that emotions that are located close to each other have similar values, as in the emotion map 900 shown in FIG. 10. FIG. 10 shows an example in which multiple emotions, "relieved," "calm," and "reassuring," have similar emotion values.

[2223] The system according to the present disclosure has been described above mainly with respect to the functions of the data processing device 12, but the system according to the present disclosure is not necessarily implemented on a server. The system according to the present disclosure may be implemented as a general information processing system. The present disclosure may be implemented, for example, as a software program running on a personal computer or an application running on a smartphone, etc. The method according to the present disclosure may be provided to users in the form of SaaS (Software as a Service).

[2224] In the above embodiment, an example was given in which the specific processing is performed by one computer 22, but the technology of the present disclosure is not limited to this, and the specific processing may be distributed and performed by a plurality of computers including the computer 22. For example, the data generation model 58 may be provided in an external device of the data processing device 12, and data may be generated in the external device in accordance with input data.

[2225] In the above embodiment, an example in which the specific processing program 56 is stored in the storage 32 has been described, but the technology of the present disclosure is not limited to this. For example, the specific processing program 56 may be stored in a portable, computer-readable, non-transitory storage medium such as a USB (Universal Serial Bus) memory. The specific processing program 56 stored in the non-transitory storage medium is installed in the computer 22 of the data processing device 12. The processor 28 executes the specific processing in accordance with the specific processing program 56.

[2226] Alternatively, the specific processing program 56 may be stored in a storage device such as a server connected to the data processing device 12 via the network 54, and the specific processing program 56 may be downloaded and installed on the computer 22 in response to a request from the data processing device 12.

[2227] It is not necessary to store all of the specific processing program 56 in a storage device such as a server connected to the data processing device 12 via the network 54, or to store all of the specific processing program 56 in the storage 32; only a portion of the specific processing program 56 may be stored.

[2228] The hardware resource for executing a specific process can be any of the following processors: An example of a processor is a CPU, which is a general-purpose processor that functions as a hardware resource for executing a specific process by executing software, i.e., a program. Another example of a processor is a dedicated electrical circuit, such as an FPGA (Field-Programmable Gate Array), a PLD (Programmable Logic Device), or an ASIC (Application Specific Integrated Circuit), which is a processor with a circuit configuration designed specifically for executing a specific process. Each processor has built-in or connected memory, and each processor uses the memory to execute the specific process.

[2229] The hardware resource that executes the specific processing may be configured with one of these various processors, or may be configured with a combination of two or more processors of the same or different types (for example, a combination of multiple FPGAs, or a combination of a CPU and an FPGA). Also, the hardware resource that executes the specific processing may be a single processor.

[2230] As an example of a system configured with a single processor, first, one processor is configured by combining one or more CPUs and software, and this processor functions as a hardware resource that executes a specific process. Second, there is a system that uses a processor that realizes the functions of an entire system including multiple hardware resources that execute a specific process on a single IC chip, as typified by SoC (System-on-a-chip). In this way, a specific process is realized using one or more of the above-mentioned various processors as hardware resources.

[2231] Furthermore, the hardware structure of these various processors can be, more specifically, an electric circuit that combines circuit elements such as semiconductor devices. The specific processing described above is merely an example. Therefore, it goes without saying that unnecessary steps may be deleted, new steps may be added, or the processing order may be rearranged, without departing from the spirit of the invention.

[2232] The above-described description and illustrations are a detailed explanation of the parts related to the technology of the present disclosure and are merely an example of the technology of the present disclosure. For example, the above description of the configuration, functions, actions, and effects is an explanation of an example of the configuration, functions, actions, and effects of the parts related to the technology of the present disclosure. Therefore, it goes without saying that unnecessary parts may be deleted, new elements may be added, or replacements may be made to the above-described description and illustrations within the scope of the gist of the technology of the present disclosure. Furthermore, to avoid confusion and facilitate understanding of the parts related to the technology of the present disclosure, the above-described description and illustrations omit explanations of common technical knowledge that do not require particular explanation to enable the implementation of the technology of the present disclosure.

[2233] All publications, patent applications, and technical standards mentioned in this specification are herein incorporated by reference to the same extent as if each individual publication, patent application, or technical standard was specifically and individually indicated to be incorporated by reference.

[2234] The following is further disclosed regarding the above embodiment.

[2235] (Claim 1)

[2236] means for capturing audio data;

[2237] means for pre-processing the captured audio data;

[2238] means for extracting acoustic features from the preprocessed speech data;

[2239] means for comparing the extracted acoustic features with an existing speech profile;

[2240] a means for analyzing the comparison results to detect anomalies;

[2241] A system that includes a means for generating and notifying alerts when an anomaly is detected.

[2242] (Claim 2)

[2243] 10. The system of claim 1,

[2244] The system includes pre-processing means for denoising and normalising the captured audio data.

[2245] (Claim 3)

[2246] 10. The system of claim 1,

[2247] The system includes a means for using cosine similarity or Euclidean distance to compare extracted acoustic features.

[2248] "Example 1"

[2249] (Claim 1)

[2250] means for capturing audio data;

[2251] means for pre-processing the captured audio data;

[2252] means for extracting acoustic features from the preprocessed speech data;

[2253] means for comparing the extracted acoustic features with an existing speech profile;

[2254] a means for analyzing the comparison results to detect anomalies;

[2255] A means for generating and notifying alerts when an anomaly is detected;

[2256] The system includes a means for storing processing results and alert information in a database.

[2257] (Claim 2)

[2258] 10. The system of claim 1, further comprising pre-processing means for performing noise reduction and normalization on the captured audio data.

[2259] (Claim 3)

[2260] 2. The system according to claim 1, further comprising means for using cosine similarity or Euclidean distance to compare the extracted acoustic features.

[2261] "Application Example 1"

[2262] (Claim 1)

[2263] means for capturing audio data;

[2264] means for pre-processing the captured audio data;

[2265] means for extracting acoustic features from the preprocessed speech data;

[2266] means for comparing the extracted acoustic features with a speech profile;

[2267] A means of detecting anomalies based on the analysis results;

[2268] means for generating and notifying an alert when an anomaly is detected;

[2269] means for recording warning information;

[2270] a means for detecting in real time when the voice data is fraudulently generated;

[2271] an application for monitoring call data;

[2272] A system including a cloud-based processing server.

[2273] (Claim 2)

[2274] 10. The system of claim 1, wherein the system automatically monitors and transmits voice data upon initiation of a call.

[2275] (Claim 3)

[2276] 2. The system of claim 1, further comprising a pre-processing means, wherein the pre-processing includes noise removal and normalization.

[2277] "Example 2: Combining Emotion Engines"

[2278] (Claim 1)

[2279] means for capturing audio data;

[2280] means for encrypting and transmitting the captured audio data;

[2281] means for preprocessing received audio data;

[2282] means for extracting acoustic features from the preprocessed speech data;

[2283] means for comparing the extracted acoustic features with an existing speech profile;

[2284] a means for analyzing the comparison results to detect anomalies;

[2285] A means for generating and notifying alerts when an anomaly is detected;

[2286] means for recognizing an emotional state from speech data;

[2287] The system includes a means for selecting a notification method based on a recognized emotional state.

[2288] (Claim 2)

[2289] 10. The system of claim 1, further comprising pre-processing means for performing noise reduction and normalization on the captured audio data.

[2290] (Claim 3)

[2291] 2. The system according to claim 1, further comprising means for using cosine similarity or Euclidean distance to compare the extracted acoustic features.

[2292] "Application example 2 when combining emotion engines"

[2293] Claims (modified)

[2294] (Claim 1)

[2295] means for capturing audio data;

[2296] means for pre-processing the captured audio data;

[2297] means for extracting acoustic features from the preprocessed speech data;

[2298] means for comparing the extracted acoustic features with an existing speech profile;

[2299] a means for analyzing the comparison results to detect anomalies;

[2300] A means for generating and notifying alerts when an anomaly is detected;

[2301] means for analyzing a user's emotional state from the voice data;

[2302] A system including a means for selecting a notification method for an alert based on the analysis results.

[2303] (Claim 2)

[2304] 10. The system of claim 1, further comprising pre-processing means for performing noise reduction and normalization on the captured audio data.

[2305] (Claim 3)

[2306] 2. The system according to claim 1, further comprising means for using cosine similarity or Euclidean distance to compare the extracted acoustic features. [Explanation of symbols]

[2307] 10, 210, 310, 410 Data Processing Systems 12 Data Processing Device 14 Smart Devices 214 Smart Glasses 314 Headset-type terminal 414 Robot< / url:> < / url:> < / url:> < / url:>

Claims

1. means for capturing audio data; means for pre-processing the captured audio data; means for extracting acoustic features from the preprocessed speech data; means for comparing the extracted acoustic features with an existing speech profile; a means for analyzing the comparison results to detect anomalies; A system that includes a means for generating and notifying alerts when an anomaly is detected.

2. 10. The system of claim 1, The system includes pre-processing means for denoising and normalising the captured audio data.

3. 10. The system of claim 1, The system includes a means for using cosine similarity or Euclidean distance to compare extracted acoustic features.

Citation Information

Patent Citations

  • Persona chatbot control method and system

    JP2022180282A