system
A system using a terminal and machine learning model to analyze voice data in real time detects synthesized speech and provides warnings, addressing the challenge of voice fraud by enhancing user safety and accuracy through feedback mechanisms.
Patent Information
- Authority / Receiving Office
- JP · JP
- Patent Type
- Applications
- Current Assignee / Owner
- SOFTBANK GROUP CORP
- Filing Date
- 2024-12-12
- Publication Date
- 2026-06-24
AI Technical Summary
The increasing difficulty in distinguishing between synthesized and natural voice due to advancements in voice synthesis technology has led to a rise in fraud and malicious use, particularly affecting users unfamiliar with technology and the elderly, necessitating a system to analyze voice in real time and provide warnings.
A terminal collects audio data, transmits it to a processing unit for analysis using a machine learning model to determine synthesized speech, generates a warning signal if processed, and allows users to provide feedback for model improvement.
Enables users to communicate with confidence by detecting fraudulent audio and improving the system's accuracy through user feedback, thereby preventing voice scams and ensuring safe voice interactions.
Smart Images

Figure 2026103495000001_ABST
Abstract
Description
Technical Field
[0001] The technology of the present disclosure relates to a system.
Background Art
[0002] Patent Document 1 discloses a method for controlling a persona chatbot, which is performed by at least one processor, including the steps of receiving a user utterance, adding the user utterance to a prompt including an instruction sentence related to an explanation of a chatbot character, encoding the prompt, and inputting the encoded prompt into a language model to generate a chatbot utterance that responds to the user utterance.
Prior Art Documents
Patent Documents
[0003]
Patent Document 1
Summary of the Invention
Problems to be Solved by the Invention
[0004] With the evolution of voice synthesis technology, it has become difficult to distinguish between synthesized voice and natural voice. For this reason, the risk of fraud and malicious use has increased, and especially users who are not familiar with technology and the elderly may be victimized. To address this issue, it is necessary to provide an environment where users can communicate by voice with confidence by analyzing voice in real time, detecting the possibility of synthesized voice, and issuing a warning to the user.
Means for Solving the Problems
[0005] This invention provides a terminal that collects audio data and transmits it to a processing unit for analysis. The processing unit analyzes the received audio data using a machine learning model to determine whether the audio has been processed. If the audio is recognized as processed based on this determination, it generates a warning signal and transmits the warning signal to the terminal. The terminal displays a warning to the user in response to the warning signal and identifies the fraudulent audio. Furthermore, the invention provides a means to continuously improve the machine learning model by receiving feedback from the user to report misjudgments in response to the warning message and using this feedback.
[0006] "Audio data" refers to data that represents audio information in a digital format.
[0007] A "terminal" is a device used by a user to communicate through direct operation.
[0008] A "processing device" is a computer that receives, analyzes, and transmits data.
[0009] A "machine learning model" is an algorithm that learns from data and performs analysis and judgment on new data.
[0010] "Synthesized speech" refers to artificially generated voices that imitate natural human speech.
[0011] A "warning signal" is a notification method generated to alert a user when certain conditions are met.
[0012] "Feedback" refers to operational results and suggestions for improvement obtained from users, and contributes to improving the accuracy of the system. [Brief explanation of the drawing]
[0013] [Figure 1] This is a conceptual diagram showing an example of the configuration of a data processing system according to the first embodiment. [Figure 2]It is a conceptual diagram showing an example of the main functions of a data processing device and a smart device according to the first embodiment. [Figure 3] It is a conceptual diagram showing an example of the configuration of a data processing system according to the second embodiment. [Figure 4] It is a conceptual diagram showing an example of the main functions of a data processing device and smart glasses according to the second embodiment. [Figure 5] It is a conceptual diagram showing an example of the configuration of a data processing system according to the third embodiment. [Figure 6] It is a conceptual diagram showing an example of the main functions of a data processing device and a headset-type terminal according to the third embodiment. [Figure 7] It is a conceptual diagram showing an example of the configuration of a data processing system according to the fourth embodiment. [Figure 8] It is a conceptual diagram showing an example of the main functions of a data processing device and a robot according to the fourth embodiment. [Figure 9] It shows an emotion map to which multiple emotions are mapped. [Figure 10] It shows an emotion map to which multiple emotions are mapped. [Figure 11] It is a sequence diagram showing the processing flow of the data processing system in Example 1. [Figure 12] It is a sequence diagram showing the processing flow of the data processing system in Application Example 1. [Figure 13] It is a sequence diagram showing the processing flow of the data processing system in Example 2 when an emotion engine is combined. [Figure 14] It is a sequence diagram showing the processing flow of the data processing system in Application Example 2 when an emotion engine is combined.
MODE FOR CARRYING OUT THE INVENTION
[0014] Hereinafter, an example of an embodiment of a system according to the technology of the present disclosure will be described according to the accompanying drawings.
[0015] First, the terms used in the following description will be explained.
[0016] In the following embodiments, the labeled processor (hereinafter simply referred to as "processor") may be a single arithmetic unit or a combination of multiple arithmetic units. Also, the processor may be a single type of arithmetic unit or a combination of multiple types of arithmetic units. Examples of arithmetic units include a CPU (Central Processing Unit), a GPU (Graphics Processing Unit), a GPGPU (General-Purpose computing on Graphics Processing Units), an APU (Accelerated Processing Unit), and the like.
[0017] In the following embodiments, the labeled RAM (Random Access Memory) is a memory in which information is temporarily stored and is used as a work memory by the processor.
[0018] In the following embodiments, the labeled storage is one or more non-volatile storage devices that store various programs and various parameters, etc. Examples of non-volatile storage devices include flash memory (SSD (Solid State Drive)), magnetic disks (e.g., hard disks), or magnetic tapes, and the like.
[0019] In the following embodiments, the labeled communication I / F (Interface) is an interface including a communication processor and an antenna, etc. The communication I / F controls communication between multiple computers. Examples of communication standards applied to the communication I / F include wireless communication standards including 5G (5th Generation Mobile Communication System), Wi-Fi (registered trademark), or Bluetooth (registered trademark), and the like.
[0020] In the following embodiments, "A and / or B" is synonymous with "at least one of A and B." That is, "A and / or B" means that it may be A alone, or B alone, or a combination of A and B. Furthermore, in this specification, the same concept as "A and / or B" applies when expressing three or more things linked by "and / or."
[0021] [First Embodiment]
[0022] Figure 1 shows an example of the configuration of the data processing system 10 according to the first embodiment.
[0023] As shown in Figure 1, the data processing system 10 includes a data processing device 12 and a smart device 14. An example of the data processing device 12 is a server.
[0024] The data processing device 12 comprises a computer 22, a database 24, and a communication interface 26. The computer 22 is an example of a "computer" related to the technology of this disclosure. The computer 22 comprises a processor 28, RAM 30, and storage 32. The processor 28, RAM 30, and storage 32 are connected to a bus 34. The database 24 and the communication interface 26 are also connected to the bus 34. The communication interface 26 is connected to a network 54. An example of the network 54 is a WAN (Wide Area Network) and / or a LAN (Local Area Network).
[0025] The smart device 14 comprises a computer 36, a reception device 38, an output device 40, a camera 42, and a communication interface 44. The computer 36 comprises a processor 46, RAM 48, and storage 50. The processor 46, RAM 48, and storage 50 are connected to a bus 52. The reception device 38, output device 40, and camera 42 are also connected to the bus 52.
[0026] The reception device 38 is equipped with a touch panel 38A and a microphone 38B, etc., and receives user input. The touch panel 38A receives user input by detecting contact with an object (e.g., a pen or finger). The microphone 38B receives user input by detecting the user's voice. The control unit 46A transmits data indicating the user input received by the touch panel 38A and microphone 38B to the data processing device 12. In the data processing device 12, the specific processing unit 290 acquires the data indicating the user input.
[0027] The output device 40 includes a display 40A and a speaker 40B, and presents data to the user 20 by outputting the data in a form perceptible to the user 20 (e.g., audio and / or text). The display 40A displays visible information such as text and images according to instructions from the processor 46. The speaker 40B outputs audio according to instructions from the processor 46. The camera 42 is a small digital camera equipped with an optical system such as a lens, aperture, and shutter, and an image sensor such as a CMOS (Complementary Metal-Oxide-Semiconductor) image sensor or a CCD (Charge Coupled Device) image sensor.
[0028] Communication interface 44 is connected to network 54. Communication interfaces 44 and 26 are responsible for the exchange of various types of information between processor 46 and processor 28 via network 54.
[0029] Figure 2 shows an example of the main functions of the data processing device 12 and the smart device 14.
[0030] As shown in Figure 2, in the data processing device 12, a specific processing is performed by the processor 28. A specific processing program 56 is stored in the storage 32. The specific processing program 56 is an example of a "program" related to the technology of this disclosure. The processor 28 reads the specific processing program 56 from the storage 32 and executes the read specific processing program 56 on the RAM 30. The specific processing is realized by the processor 28 operating as a specific processing unit 290 according to the specific processing program 56 executed on the RAM 30.
[0031] The storage 32 stores the data generation model 58 and the emotion identification model 59. The data generation model 58 and the emotion identification model 59 are used by the identification processing unit 290.
[0032] In the smart device 14, the processor 46 performs the reception output processing. The storage 50 stores the reception output program 60. The reception output program 60 is used in conjunction with a specific processing program 56 by the data processing system 10. The processor 46 reads the reception output program 60 from the storage 50 and executes the read reception output program 60 on the RAM 48. The reception output processing is realized by the processor 46 operating as a control unit 46A according to the reception output program 60 executed on the RAM 48.
[0033] Next, the specific processing performed by the specific processing unit 290 of the data processing device 12 will be described. In the following description, the data processing device 12 will be referred to as the "server" and the smart device 14 as the "terminal".
[0034] This invention is a system in which a terminal, a server, and a user work together. The terminal, with the user's consent, collects voice data from voice calls and voice messages in real time. The collected voice data is temporarily stored in a buffer, and the packetized voice data is sent to the server at regular intervals.
[0035] The server processes the received audio data in real time and inputs it into a machine learning model. This model analyzes the characteristics of the audio and scores the likelihood that it is synthesized speech. If the score exceeds a set threshold, the server determines that the audio is processed speech.
[0036] Once a determination is made, the server generates a warning signal and sends it to the terminal. The terminal receives this signal and displays a warning message to the user visually or audibly. The user reviews the warning and decides whether to continue or end the call based on that information.
[0037] For example, if a user receives a phone call that they suspect is a scam, their device sends the call to a server. If the server determines that it is a synthesized voice, a warning is displayed on the device, and the user reviews the warning and makes an appropriate decision. If there is a misidentification, the user can send feedback to the server through their device, and the server can use this feedback to improve its machine learning model and perform more accurate analysis.
[0038] This system allows users to protect themselves from voice scams and use voice services safely.
[0039] The following describes the processing flow.
[0040] Step 1:
[0041] The device monitors audio data in real time while the user is receiving phone calls or voice messages. This data is temporarily stored in a buffer and prepared for processing.
[0042] Step 2:
[0043] The terminal divides the voice data stored in the buffer into packets at regular time intervals and sends them to the server. At this time, voice metadata (such as call ID and timestamp) is also sent.
[0044] Step 3:
[0045] The server sequentially inputs the received audio packets into a machine learning model. The model then begins the process of analyzing the features of the audio data and scoring the likelihood that it is synthesized speech.
[0046] Step 4:
[0047] The server evaluates the score obtained through analysis and determines that the audio is processed if it exceeds a set threshold. The determination result is recorded as a log.
[0048] Step 5:
[0049] If the server detects that the audio has been altered, it generates a warning signal and sends this signal to the terminal. This warning signal includes a message to inform the user.
[0050] Step 6:
[0051] The device displays a visual or auditory warning to the user based on the received warning signal. The user reviews the warning and decides whether to end the call.
[0052] Step 7:
[0053] After making a decision based on a warning, users can send feedback from their device to the server to report any misjudgments, if necessary. This feedback is used to improve the machine learning model.
[0054] (Example 1)
[0055] Next, we will describe Example 1. In the following description, the data processing device 12 will be referred to as the "server," and the smart device 14 will be referred to as the "terminal."
[0056] In today's world, fraud and information manipulation using synthesized voices are on the rise, making it crucial to determine the authenticity of voice data in real time. However, current technology makes it difficult to distinguish synthesized voices simply by listening, potentially causing harm to users. Solving this problem and enabling users to communicate with voices with confidence is essential.
[0057] The identification process performed by the identification processing unit 290 of the data processing device 12 in Example 1 is realized by the following means.
[0058] In this invention, the server includes means for transmitting voice information to an analysis device, means for analyzing the received voice information using a machine learning algorithm to determine whether it is processed voice, and means for generating a warning signal based on the determination result and transmitting the warning signal to a device. This prevents fraud using synthesized voices and enables users to use voice information safely in real time.
[0059] "Audio information" refers to data that represents audio in digital format, and this generally includes the content of voice calls and voice messages.
[0060] "Device" refers to an electronic device that collects voice information and transmits it to a server, and includes mobile devices such as smartphones and tablets.
[0061] An "analysis device" refers to a computer system that receives audio information and analyzes its content.
[0062] A "machine learning algorithm" refers to a model used to analyze audio data, understand its characteristics, and determine whether it is synthesized speech.
[0063] A "warning signal" refers to a signal generated to alert the user when there is a possibility that the audio information is synthesized.
[0064] "Response" refers to information that users can use to provide feedback on the synthesized speech's evaluation results and the system's operation.
[0065] This invention relates to a system for determining the authenticity of audio information in real time. The system operates in cooperation with three parties: a terminal used by the user, a server that analyzes the audio data, and the user who receives the warning.
[0066] The device collects voice information from voice calls and voice messages in real time with the user's consent. The device used is a portable electronic device such as a smartphone or tablet. The collected voice information is packetized at regular intervals and sent to a server via a communication protocol. To ensure smooth data transmission, the TCP / IP protocol is commonly used.
[0067] The server applies machine learning algorithms to analyze the received audio information. These algorithms are developed using frameworks such as TENSORFLOW® or PyTorch. The server analyzes the characteristics of the audio information and quantifies the likelihood of it being synthesized speech through scoring. If the score exceeds a certain threshold, it is determined to be synthesized speech and a warning signal is generated. This warning signal is sent to the terminal via the HTTP protocol or similar.
[0068] Based on the received warning signal, the device notifies the user of a warning message visually or audibly. For example, it can display "Warning: This may be a synthesized voice" on the device's screen, or emit a warning sound. This allows the user to make informed decisions to prevent fraud.
[0069] For example, if a user receives a suspicious call, the system analyzes the audio information of the call and makes a judgment. If the server determines that the voice is synthesized speech, a warning is displayed on the terminal, allowing the user to take appropriate action. The user can also provide feedback on the system's judgment, which is used to improve the machine learning algorithm.
[0070] Examples of prompt messages could include, "How should we warn the user if the call is potentially fraudulent?" or "Please tell us how to improve the voice analysis system using machine learning models." By using this system, users can communicate via voice with peace of mind.
[0071] The flow of the specific processing in Example 1 will be explained using Figure 11.
[0072] Step 1:
[0073] With the user's consent, the device collects voice information from voice calls and voice messages in real time. Specifically, the device uses the smartphone's built-in microphone to convert the voice information into a digital format. As a result, the input acquired by the device is analog voice data, and the output is digitized voice information.
[0074] Step 2:
[0075] The device temporarily stores the collected audio information in memory. Specifically, it packets the real-time audio data at regular intervals (e.g., every second). The input is the real-time collected audio information, and the output is the data of that audio information divided into smaller packets.
[0076] Step 3:
[0077] The terminal sends packetized voice information to the server. Specifically, this involves sending data to the server's receiving port using the TCP / IP protocol. The input is packetized voice information, and the output is the voice data sent to the server via the network.
[0078] Step 4:
[0079] The server inputs the received audio data into a machine learning algorithm. The server provides the audio data to a model developed using, for example, TensorFlow, and extracts audio features. The input is the audio data received by the server, and the output is the audio data features and analysis results.
[0080] Step 5:
[0081] The server uses the output of a machine learning algorithm to score whether the speech is likely to be synthesized. Here, if the score exceeds a certain threshold, it is classified as "synthesized speech." The input is the analyzed features, and the output is the scoring result and the classification result.
[0082] Step 6:
[0083] The server generates a warning signal if it determines that the voice is synthesized. This signal contains a message such as "Caution: This voice may have been synthesized." In its specific operation, it generates a warning signal and creates a data structure to send it to the terminal. The input is the detection result, and the output is the warning signal.
[0084] Step 7:
[0085] The terminal notifies the user based on the received warning signal. Specifically, the terminal attracts the user's attention by displaying text on the screen or emitting a warning sound. The input is a warning signal sent from the server, and the output is a warning message that the user sees or hears.
[0086] (Application Example 1)
[0087] Next, we will explain Application Example 1. In the following explanation, the data processing device 12 will be referred to as the "server," and the smart device 14 will be referred to as the "terminal."
[0088] In recent years, advancements in voice technology have led to an increase in fraud and illegal activities using synthesized voices. This increases the likelihood of users finding themselves in dangerous situations. However, technology to accurately determine whether voice information is synthesized is still insufficient, and users lack effective countermeasures against fraudulent voices. Against this backdrop, the present invention aims to provide a system that can quickly and accurately determine whether voice information is synthesized and warn users.
[0089] The specific processing performed by the specific processing unit 290 of the data processing device 12 in Application Example 1 is realized by the following means.
[0090] In this invention, the server includes information processing means for analyzing voice information using a learning algorithm and determining whether it is modified voice; signal generation means for generating a warning signal based on the determination result and transmitting it to an information terminal; and means for receiving improvement reports from users who have received warnings and applying them to the learning algorithm. This enables users to quickly detect fraudulent synthesized voices and take appropriate action.
[0091] "Voice information" refers to information data recorded as sound waves, including voice calls and voice messages.
[0092] An "information terminal" is a device that has the function of collecting voice information and transmitting it to a data processing device.
[0093] A "data processing device" is an electronic device with computing power for analyzing received audio information.
[0094] A "learning algorithm" is an algorithm that uses machine learning to analyze the characteristics of audio information and determine whether it is likely to be synthesized speech.
[0095] "Information processing means" refers to a function that performs a process of analyzing received audio information and determining its validity.
[0096] A "signal generation means" is a means that has the function of generating a warning signal when it is determined that audio information has been altered.
[0097] A "notification means" is a function that transmits a warning to the user visually or audibly via an information terminal that receives a warning signal.
[0098] A "user" is a person who receives and manipulates voice information, or an entity intended to use such information.
[0099] An "improvement report" is user feedback based on warning messages, and it provides information to improve the accuracy of the learning algorithm.
[0100] This invention relates to a system for detecting fraudulent voice information, and consists of three components: an information terminal, a data processing device, and a user. The information terminal is comprised of a device such as a smartphone or wearable device, through which voice information is collected. The voice information is temporarily stored in a buffer in real time and periodically transmitted to the data processing device as data packets.
[0101] The data processing unit functions as a server and, upon receiving audio information, uses a pre-trained learning algorithm to analyze the characteristics of that audio information in detail. The analysis quantifies the likelihood that the audio is synthesized speech, and if it exceeds a certain threshold, it is determined to be modified speech. At this point, the server generates a warning signal, which is transmitted to the information terminal. The information terminal receives this warning signal and provides the user with a visual or auditory warning.
[0102] Users are alerted to warnings and encouraged to be vigilant against fraud and misconduct. If a misidentification occurs during actual use, users can send feedback from their information terminal to the server. This feedback is received by the server and used to improve the learning algorithm, enabling more accurate analysis.
[0103] As a concrete example, when a user receives a phone call, the system analyzes the call and, if it detects a synthesized voice suspected of being fraudulent, the user receives an instant warning, thus preventing fraud. An example of a prompt message for this voice analysis would be, "Please determine whether this voice is genuine or synthesized. If the score exceeds the threshold, generate a warning message."
[0104] The flow of a specific process in Application Example 1 will be explained using Figure 12.
[0105] Step 1:
[0106] The device collects voice information in real time. When a voice call or voice message begins, the voice signal is stored in a buffer as digital data. At this stage, the input is voice information, and the output is digital voice data. The collected voice data is packetized at regular intervals.
[0107] Step 2:
[0108] The terminal sends packetized voice data to the server. The data is securely transferred over the network and received by the server. The input is packetized voice data, and the output is the data received on the server side. The transfer is secured using encrypted communication.
[0109] Step 3:
[0110] The server inputs the received audio data into a machine learning model for analysis. It extracts audio characteristics as features and scores the likelihood that the audio is synthesized speech. The input is the received audio data, and the output is the synthesized speech score. This analysis is performed using a sophisticated algorithm based on a generative AI model.
[0111] Step 4:
[0112] The server generates a warning signal if the score exceeds a set threshold. This signal indicates a high probability of synthesized speech. The input is the analysis score, and the output is the warning signal. Signal generation is performed in real time, requiring immediate response.
[0113] Step 5:
[0114] The terminal receives a warning signal sent from the server and notifies the user. The warning is presented visually as an on-screen alert and audibly as an audio notification. The input is the warning signal, and the output is a warning display to the user. The most suitable notification method is selected depending on the situation.
[0115] Step 6:
[0116] The user receives a warning and determines the veracity of the audio information. They review the warning and, if necessary, interrupt the call or send feedback. The input is the warning to the user, and the output is the user's response. Feedback helps improve the system in case of misjudgments.
[0117] Step 7:
[0118] The server receives feedback from users and uses it to improve the machine learning model. The model is retrained with new data to achieve more accurate analysis. The input is user feedback, and the output is the improved learning algorithm. This process is continuous, improving the model's performance.
[0119] Furthermore, an emotion engine that estimates the user's emotions may be incorporated. That is, the identification processing unit 290 may use the emotion identification model 59 to estimate the user's emotions and perform identification processing using the user's emotions.
[0120] The present invention is a system that combines a terminal for collecting and analyzing voice data, a server equipped with means for determining whether the voice data is processed voice, and an emotion engine that recognizes the user's emotions and appropriately adjusts warning messages.
[0121] The terminal monitors voice calls and messages received by the user in real time and temporarily stores the voice data in a buffer. This data is packetized at regular intervals and sent to a server for analysis.
[0122] The server feeds the received audio data into a machine learning model. This model analyzes the characteristics of the audio and scores the likelihood that it is synthesized speech. If the score exceeds a threshold, the server determines that the audio is processed speech and sends a warning signal to the terminal.
[0123] The emotion engine analyzes the user's voice tone, speaking speed, facial expressions, etc., to recognize the user's emotional state in real time. Based on the recognized emotional information, it is possible to adjust how warning messages are presented (for example, the emphasis and type of the message).
[0124] For example, during a normal phone call, the server analyzes the received audio and determines that it may be synthesized speech. If the emotion engine then recognizes that the user's voice sounds tense, it will alert the user by displaying a more emphasized warning message on the device.
[0125] This system not only combats the risk of voice fraud but also enables flexible responses that respond to the user's emotions, providing more personalized protection.
[0126] The following describes the processing flow.
[0127] Step 1:
[0128] The device monitors the audio data in real time when the user is receiving a phone call or voice message, and temporarily stores it in a buffer. This audio data is then packetized in preparation for subsequent processing.
[0129] Step 2:
[0130] The terminal sends the voice data stored in the buffer to the server at regular time intervals. At this time, the call ID and timestamp are also sent as metadata.
[0131] Step 3:
[0132] The server inputs the received audio data into a machine learning model to analyze the audio's characteristics. The model scores the likelihood that the audio is synthesized speech, and if it exceeds a threshold, it is classified as processed speech.
[0133] Step 4:
[0134] If the server determines that the audio is processed, it generates a warning signal and sends it to the terminal. The warning signal includes instructions to alert the user.
[0135] Step 5:
[0136] The device activates the emotion engine based on the received warning signal. The emotion engine analyzes the user's voice tone, speaking speed, and facial expressions to recognize the user's emotional state.
[0137] Step 6:
[0138] Based on the emotional information recognized by the emotion engine, the device adjusts how warning messages are presented. For example, if the user is showing signs of anxiety, the warning message is visually highlighted.
[0139] Step 7:
[0140] The user reviews the warning message from their device and decides whether to continue or end the call. If necessary, they can send feedback from their device to the server if there is a misjudgment. This feedback is used by the server to improve the machine learning model.
[0141] (Example 2)
[0142] Next, we will describe Example 2. In the following description, the data processing device 12 will be referred to as the "server" and the smart device 14 as the "terminal".
[0143] In modern society, voice-based fraud is on the rise, but effective detection systems to combat it are still not well-established. Furthermore, existing systems do not take into account the user's emotional state, and inappropriate warnings can lead to unnecessary stress and confusion. Therefore, there is a need for technology that can quickly and accurately identify manipulated voices and provide warnings tailored to the user's emotional state.
[0144] The identification process performed by the identification processing unit 290 of the data processing device 12 in Example 2 is realized by the following means.
[0145] In this invention, the server includes means for analyzing voice information using a computational model and determining whether the voice has been processed; means for generating and transmitting a warning signal based on the determination result; and means for evaluating the user's emotional state using an emotion analysis device and adjusting the method of presenting the warning based on the emotional state. This enables rapid and accurate detection of processed voices and personalized warning presentation that takes the user's emotions into consideration.
[0146] "Audio information" refers to the digital or analog representation of sound waves that have been transmitted or stored.
[0147] An "information processing device" is a combination of hardware and software used to compute, analyze, or transform data.
[0148] A "computational model" is an algorithm or mathematical formula used to analyze the characteristics of data and make specific decisions based on the results.
[0149] A "warning signal" is a signal generated to inform the user or system of a specific condition or abnormality.
[0150] An "emotion analysis device" is a device or system that evaluates or estimates a user's emotional state based on information such as their voice and facial expressions.
[0151] "User" refers to the individual who operates or uses the system or product.
[0152] "Numerical evaluation" refers to numerical results calculated to quantify specific characteristics or states.
[0153] "Opinions" refers to feedback and reactions provided by users.
[0154] This invention is a system for monitoring voice calls and messages received by a user and analyzing the voice information. Its primary purpose is to collect and analyze voice information and provide warnings that take into account the user's emotional state. The following describes a specific implementation of this system.
[0155] The terminal acquires the user's voice calls and messages in real time. To do this, it captures voice information using a voice input device and digital signal processing technology. The collected voice information is temporarily stored in the terminal's memory and then transmitted to a server using network communication technology.
[0156] The server is equipped with a computing device that analyzes transmitted audio information using a computational model. This analysis utilizes machine learning libraries (e.g., TensorFlow and PyTorch) to analyze the characteristics of the audio and determine whether it is synthesized speech. If it is determined to be synthesized speech, a warning signal is generated and sent to the terminal.
[0157] The emotion analysis device analyzes emotional data from voice information and the user's facial expressions to understand the user's emotional state. This information is transmitted to a server and used to display warning signals. For example, if the device detects a state of tension, it will display a more emphasized warning message to prompt appropriate action.
[0158] As a concrete example, suppose a user makes a voice call, and the server analyzes the voice information and determines that it is synthesized speech. In this situation, the emotion analysis device can sense the tension in the user's voice and issue a command to display an emphasized warning message on the device. This allows the user to immediately respond to the risk of voice fraud.
[0159] An example of a prompt to input into a generative AI model is, "Analyze the user's emotion from their voice tone and facial expressions, and create a message appropriate to the situation." This prompt helps improve the accuracy of emotion analysis and make warning messages more appropriate.
[0160] The flow of the specific processing in Example 2 will be explained using Figure 13.
[0161] Step 1:
[0162] The terminal monitors incoming voice calls and messages in real time. When voice input occurs, the terminal uses the microphone to capture sound wave data and converts it into a digital format. This data is temporarily stored in memory and packetized. The input is voice data, and the output is packetized digital data.
[0163] Step 2:
[0164] The terminal sends packetized voice data to the server at regular intervals. Efficient data transfer using network protocols is required to minimize latency. The input is packetized digital data, and the output is voice data received by the server.
[0165] Step 3:
[0166] The server acquires the received audio data and feeds it into a machine learning model. The computational model used here is, for example, trained using TensorFlow or PyTorch. It analyzes the characteristics of the audio, such as frequency components and waveform features, and scores the likelihood of it being synthesized speech. The input is the audio data sent to the server, and the output is the scored likelihood of it being synthesized speech.
[0167] Step 4:
[0168] The server compares the scoring result to a threshold. If the threshold is exceeded, the server determines that the voice is synthesized speech and generates a warning signal. This warning signal is then prepared to be sent to the terminal. The input is the scoring result, and the output is the warning signal.
[0169] Step 5:
[0170] The emotion analysis device acquires the user's voice tone and facial expression data to evaluate the user's emotional state. Data is acquired from audio or video cameras and analyzed in real time. Emotional states are classified, for example, into categories such as tension or calmness. The input is the user's voice and facial expression data, and the output is the evaluated emotional state.
[0171] Step 6:
[0172] The terminal references the warning signal received from the server and the emotional state from the emotion analyzer, and adjusts how the warning message is displayed. For example, if it determines that the user is in a state of tension, the terminal will highlight the warning message, providing an appropriate warning to the user. The input is the warning signal and emotional state, and the output is the adjusted warning message.
[0173] (Application Example 2)
[0174] Next, we will explain application example 2. In the following explanation, the data processing device 12 will be referred to as the "server," and the smart device 14 will be referred to as the "terminal."
[0175] In recent years, advancements in speech synthesis technology have increased the risk of general users being affected by voice fraud and disguised voices. However, currently, there are insufficient means to determine in real time whether a received voice is synthesized or not. Furthermore, there is a need for a system that can appropriately adjust warning messages by taking into account the user's emotional state, thereby preventing excessive warnings and providing effective alerts.
[0176] The specific processing performed by the specific processing unit 290 of the data processing device 12 in Application Example 2 is realized by the following means.
[0177] In this invention, the server includes means for a terminal that collects audio data and transmits it to an information processing device for analysis; means for analyzing the received audio data using a machine learning model and determining whether it is synthesized audio; and emotion analysis means for recognizing the user's emotional state and adjusting the warning message based on that emotional state. This makes it possible to determine whether the audio data is synthesized or natural and to provide an optimal warning based on the user's emotional state.
[0178] "Audio data" refers to information recorded in digital format, which is used for analysis and transmission.
[0179] An "information processing device" is an electronic device used to perform data analysis and calculations, and includes servers and computers.
[0180] A "machine learning model" is a collection of algorithms and computational methods used for analysis and prediction, which learn features from large amounts of data.
[0181] "Emotional state" refers to the user's emotional response and is determined by indicators such as tone of voice, speaking speed, and facial expressions.
[0182] A "warning signal" is a signal used to alert the user when certain conditions are met, and is communicated through audio, on-screen displays, or other means.
[0183] A "terminal" is an input / output device used by a user, and includes smartphones and tablets.
[0184] The system that implements this application involves real-time analysis of audio data and adjustment of warning messages based on the user's emotional state. The server receives audio data and performs analysis using a machine learning model. Specifically, it extracts features from the audio data using digital signal processing technology, applies these features to a machine learning algorithm, and scores the likelihood that the audio is synthesized speech.
[0185] The terminal receives a warning signal sent from the server and uses an emotion analysis engine to analyze the user's voice tone and speaking speed. This allows the terminal to recognize the user's emotional state and adjust the warning message accordingly. For example, if the user is feeling anxious, the warning message can be displayed with greater emphasis.
[0186] The program is implemented using programming languages such as Python, and utilizes the speech_recognition library for speech analysis. Data communication is performed via the HTTP protocol. The user interface is designed to display both audio and visual warnings.
[0187] For example, if a phone call received by a user is suspected to be using synthesized speech, the terminal receives a corresponding warning signal. At this time, the emotion engine analyzes the user's emotional state, and only if it recognizes anxiety or tension, it highlights the warning message to prompt the user to take immediate action. The prompt text to be input to the generating AI model is as follows:
[0188] "Please consider specific implementation methods for a function that acquires audio data in real time and analyzes the potential of synthesized speech."
[0189] The flow of a specific process in Application Example 2 will be explained using Figure 14.
[0190] Step 1:
[0191] The terminal temporarily stores the voice call data received from the user in a buffer. It receives real-time captured voice data as input, converts it to a digital data format, and saves it in preparation for subsequent processing.
[0192] Step 2:
[0193] The terminal packets the voice data stored in a buffer at regular intervals and sends it to the server. The input is voice call data, which is then divided into small data packets and sent to the server via the network—a data processing step.
[0194] Step 3:
[0195] The server interprets the received audio data and feeds it into a machine learning model. The input here is data packets sent from the terminal, and the server performs data calculations by extracting this audio data and analyzing it using a machine learning model.
[0196] Step 4:
[0197] The server uses a machine learning model to analyze the audio data and score its likelihood of being synthesized speech. Based on the analysis, it determines whether the audio is synthesized speech, and if the score exceeds a threshold, it generates a warning signal. The output of this step is the scoring based on the analyzed data.
[0198] Step 5:
[0199] The server sends the generated warning signal to the terminal. It takes a scoring-based judgment result as input and performs data communication to prompt appropriate action by sending that result to the user's terminal.
[0200] Step 6:
[0201] The terminal analyzes the user's emotional state using an emotion analysis engine based on the received warning signal. The input is the warning signal received from the server, and in addition, the user's voice tone and speaking speed are analyzed to process the data in order to identify the user's emotions.
[0202] Step 7:
[0203] The device adjusts and displays warning messages on the screen based on the user's emotional state. The input is the result of the analysis of the user's emotional state, and the output is a warning message that matches the result. Specifically, if the user is feeling anxious, the message will be highlighted to draw their attention.
[0204] The specific processing unit 290 transmits the result of the specific processing to the smart device 14. In the smart device 14, the control unit 46A causes the output device 40 to output the result of the specific processing. The microphone 38B acquires audio indicating user input for the result of the specific processing. The control unit 46A transmits the audio data indicating user input acquired by the microphone 38B to the data processing device 12. In the data processing device 12, the specific processing unit 290 acquires the audio data.
[0205] Data generation model 58 is a so-called generative AI (Artificial Intelligence). An example of data generation model 58 is ChatGPT (registered trademark) (Internet search).<URL: https: / / openai.com / blog / chatgpt> ), Gemini (registered trademark) (Internet search) <url: https: gemini.google.com ?hl="ja">Examples of generative AI include the following. The data generation model 58 is obtained by performing deep learning on a neural network. The data generation model 58 is input with prompts containing instructions, and with inference data such as audio data representing speech, text data representing text, and image data representing images. The data generation model 58 infers from the input inference data according to the instructions indicated by the prompts, and outputs the inference results in data formats such as audio data and text data. Here, inference refers to, for example, analysis, classification, prediction, and / or summarization.
[0206] In the above embodiment, an example was given in which specific processing is performed by the data processing device 12, but the technology of this disclosure is not limited thereto, and the specific processing may also be performed by the smart device 14.
[0207] [Second Embodiment]
[0208] Figure 3 shows an example of the configuration of the data processing system 210 according to the second embodiment.
[0209] As shown in Figure 3, the data processing system 210 includes a data processing device 12 and smart glasses 214. An example of the data processing device 12 is a server.
[0210] The data processing device 12 comprises a computer 22, a database 24, and a communication interface 26. The computer 22 is an example of a "computer" related to the technology of this disclosure. The computer 22 comprises a processor 28, RAM 30, and storage 32. The processor 28, RAM 30, and storage 32 are connected to a bus 34. The database 24 and the communication interface 26 are also connected to the bus 34. The communication interface 26 is connected to a network 54. An example of the network 54 is a WAN (Wide Area Network) and / or a LAN (Local Area Network).
[0211] The smart glasses 214 include a computer 36, a microphone 238, a speaker 240, a camera 42, and a communication interface 44. The computer 36 includes a processor 46, RAM 48, and storage 50. The processor 46, RAM 48, and storage 50 are connected to a bus 52. The microphone 238, speaker 240, and camera 42 are also connected to the bus 52.
[0212] The microphone 238 receives voice signals from the user 20 and receives instructions from the user 20. The microphone 238 captures the voice signals from the user 20, converts the captured voice into audio data, and outputs it to the processor 46. The speaker 240 outputs audio according to the instructions from the processor 46.
[0213] Camera 42 is a small digital camera equipped with an optical system including a lens, aperture, and shutter, and an image sensor such as a CMOS (Complementary Metal-Oxide-Semiconductor) image sensor or a CCD (Charge Coupled Device) image sensor, and captures images of the area around the user 20 (for example, an imaging range defined by a field of view equivalent to the width of a typical healthy person's field of vision).
[0214] Communication interface 44 is connected to network 54. Communication interfaces 44 and 26 are responsible for the exchange of various information between processor 46 and processor 28 via network 54. The exchange of various information between processor 46 and processor 28 using communication interfaces 44 and 26 is performed in a secure manner.
[0215] Figure 4 shows an example of the main functions of the data processing device 12 and the smart glasses 214. As shown in Figure 4, the data processing device 12 performs specific processing using the processor 28. The storage 32 stores the specific processing program 56.
[0216] The specific processing program 56 is an example of a "program" relating to the technology of this disclosure. The processor 28 reads the specific processing program 56 from the storage 32 and executes the read specific processing program 56 on the RAM 30. The specific processing is realized by the processor 28 operating as a specific processing unit 290 in accordance with the specific processing program 56 executed on the RAM 30.
[0217] The storage 32 stores the data generation model 58 and the emotion identification model 59. The data generation model 58 and the emotion identification model 59 are used by the identification processing unit 290.
[0218] In the smart glasses 214, the processor 46 performs the reception output processing. The storage 50 stores the reception output program 60. The processor 46 reads the reception output program 60 from the storage 50 and executes the read reception output program 60 on the RAM 48. The reception output processing is realized by the processor 46 operating as a control unit 46A according to the reception output program 60 executed on the RAM 48.
[0219] Next, the identification processing performed by the identification processing unit 290 of the data processing device 12 will be described. In the following description, the data processing device 12 will be referred to as the "server" and the smart glasses 214 will be referred to as the "terminal".
[0220] This invention is a system in which a terminal, a server, and a user work together. The terminal, with the user's consent, collects voice data from voice calls and voice messages in real time. The collected voice data is temporarily stored in a buffer, and the packetized voice data is sent to the server at regular intervals.
[0221] The server processes the received audio data in real time and inputs it into a machine learning model. This model analyzes the characteristics of the audio and scores the likelihood that it is synthesized speech. If the score exceeds a set threshold, the server determines that the audio is processed speech.
[0222] Once a determination is made, the server generates a warning signal and sends it to the terminal. The terminal receives this signal and displays a warning message to the user visually or audibly. The user reviews the warning and decides whether to continue or end the call based on that information.
[0223] For example, if a user receives a phone call that they suspect is a scam, their device sends the call to a server. If the server determines that it is a synthesized voice, a warning is displayed on the device, and the user reviews the warning and makes an appropriate decision. If there is a misidentification, the user can send feedback to the server through their device, and the server can use this feedback to improve its machine learning model and perform more accurate analysis.
[0224] This system allows users to protect themselves from voice scams and use voice services safely.
[0225] The following describes the processing flow.
[0226] Step 1:
[0227] The device monitors audio data in real time while the user is receiving phone calls or voice messages. This data is temporarily stored in a buffer and prepared for processing.
[0228] Step 2:
[0229] The terminal divides the voice data stored in the buffer into packets at regular time intervals and sends them to the server. At this time, voice metadata (such as call ID and timestamp) is also sent.
[0230] Step 3:
[0231] The server sequentially inputs the received audio packets into a machine learning model. The model then begins the process of analyzing the features of the audio data and scoring the likelihood that it is synthesized speech.
[0232] Step 4:
[0233] The server evaluates the score obtained through analysis and determines that the audio is processed if it exceeds a set threshold. The determination result is recorded as a log.
[0234] Step 5:
[0235] If the server detects that the audio has been altered, it generates a warning signal and sends this signal to the terminal. This warning signal includes a message to inform the user.
[0236] Step 6:
[0237] The device displays a visual or audible warning to the user based on the received warning signal. The user reviews the warning and decides whether to end the call.
[0238] Step 7:
[0239] After making a decision based on a warning, users can send feedback from their device to the server to report any misjudgments, if necessary. This feedback is used to improve the machine learning model.
[0240] (Example 1)
[0241] Next, we will describe Example 1. In the following description, the data processing device 12 will be referred to as the "server" and the smart glasses 214 will be referred to as the "terminal".
[0242] In today's world, fraud and information manipulation using synthesized voices are on the rise, making it crucial to determine the authenticity of voice data in real time. However, current technology makes it difficult to distinguish synthesized voices simply by listening, potentially causing harm to users. Solving this problem and enabling users to communicate with voices with confidence is essential.
[0243] The identification process performed by the identification processing unit 290 of the data processing device 12 in Example 1 is realized by the following means.
[0244] In this invention, the server includes means for transmitting voice information to an analysis device, means for analyzing the received voice information using a machine learning algorithm to determine whether it is processed voice, and means for generating a warning signal based on the determination result and transmitting the warning signal to a device. This prevents fraud using synthesized voices and enables users to use voice information safely in real time.
[0245] "Audio information" refers to data that represents audio in digital format, and this generally includes the content of voice calls and voice messages.
[0246] "Device" refers to an electronic device that collects voice information and transmits it to a server, and includes mobile devices such as smartphones and tablets.
[0247] An "analysis device" refers to a computer system that receives audio information and analyzes its content.
[0248] A "machine learning algorithm" refers to a model used to analyze audio data, understand its characteristics, and determine whether it is synthesized speech.
[0249] A "warning signal" refers to a signal generated to alert the user when there is a possibility that the audio information is synthesized.
[0250] "Response" refers to information that users can use to provide feedback on the synthesized speech's evaluation results and the system's operation.
[0251] This invention relates to a system for determining the authenticity of audio information in real time. The system operates in cooperation with three parties: a terminal used by the user, a server that analyzes the audio data, and the user who receives the warning.
[0252] The device collects voice information from voice calls and voice messages in real time with the user's consent. The device used is a portable electronic device such as a smartphone or tablet. The collected voice information is packetized at regular intervals and sent to a server via a communication protocol. To ensure smooth data transmission, the TCP / IP protocol is commonly used.
[0253] The server applies machine learning algorithms to analyze the received audio information. These algorithms are developed using frameworks such as TensorFlow or PyTorch. The server analyzes the characteristics of the audio information and quantifies the likelihood that it is synthesized speech through scoring. If the score exceeds a certain threshold, it is determined to be synthesized speech and a warning signal is generated. This warning signal is sent to the terminal via the HTTP protocol or similar.
[0254] Based on the received warning signal, the device notifies the user of a warning message visually or audibly. For example, it can display "Warning: This may be a synthesized voice" on the device's screen, or emit a warning sound. This allows the user to make informed decisions to prevent fraud.
[0255] For example, if a user receives a suspicious call, the system analyzes the audio information of the call and makes a judgment. If the server determines that the voice is synthesized speech, a warning is displayed on the terminal, allowing the user to take appropriate action. The user can also provide feedback on the system's judgment, which is used to improve the machine learning algorithm.
[0256] Examples of prompt messages could include, "How should we warn the user if the call is potentially fraudulent?" or "Please tell us how to improve the voice analysis system using machine learning models." By using this system, users can communicate via voice with peace of mind.
[0257] The flow of the specific processing in Example 1 will be explained using Figure 11.
[0258] Step 1:
[0259] With the user's consent, the device collects voice information from voice calls and voice messages in real time. Specifically, the device uses the smartphone's built-in microphone to convert the voice information into a digital format. As a result, the input acquired by the device is analog voice data, and the output is digitized voice information.
[0260] Step 2:
[0261] The terminal temporarily stores the collected audio information in memory. Specifically, it packets the real-time acquired audio data at regular intervals (e.g., every second). The input is the real-time collected audio information, and the output is the data of that audio information divided into smaller packets.
[0262] Step 3:
[0263] The terminal sends packetized voice information to the server. Specifically, this involves sending data to the server's receiving port using the TCP / IP protocol. The input is packetized voice information, and the output is the voice data sent to the server via the network.
[0264] Step 4:
[0265] The server inputs the received audio data into a machine learning algorithm. The server provides the audio data to a model developed using, for example, TensorFlow, and extracts audio features. The input is the audio data received by the server, and the output is the audio data features and analysis results.
[0266] Step 5:
[0267] The server uses the output of a machine learning algorithm to score whether the speech is likely to be synthesized. Here, if the score exceeds a certain threshold, it is classified as "synthesized speech." The input is the analyzed features, and the output is the scoring result and the classification result.
[0268] Step 6:
[0269] The server generates a warning signal if it determines that the voice is synthesized. This signal contains a message such as "Caution: This voice may have been synthesized." In its specific operation, it generates a warning signal and creates a data structure to send it to the terminal. The input is the detection result, and the output is the warning signal.
[0270] Step 7:
[0271] The terminal notifies the user based on the received warning signal. Specifically, the terminal attracts the user's attention by displaying text on the screen or emitting a warning sound. The input is a warning signal sent from the server, and the output is a warning message that the user sees or hears.
[0272] (Application Example 1)
[0273] Next, we will explain Application Example 1. In the following explanation, the data processing device 12 will be referred to as the "server," and the smart glasses 214 will be referred to as the "terminal."
[0274] In recent years, advancements in voice technology have led to an increase in fraud and illegal activities using synthesized voices. This increases the likelihood of users finding themselves in dangerous situations. However, technology to accurately determine whether voice information is synthesized is still insufficient, and users lack effective countermeasures against fraudulent voices. Against this backdrop, the present invention aims to provide a system that can quickly and accurately determine whether voice information is synthesized and warn users.
[0275] The specific processing by the specific processing unit 290 of the data processing device 12 in Application Example 1 is realized by the following means.
[0276] In this invention, the server includes information processing means for analyzing voice information by a learning algorithm to determine whether it is modified voice, signal generation means for generating a warning signal based on the determination result and transmitting it to an information terminal, and means for receiving an improvement report from a user who has received a warning and applying it to the learning algorithm. Thereby, the user can quickly detect an unauthorized synthesized voice and take appropriate measures.
[0277] "Voice information" is information data recorded as sound waves, including voice calls and voice messages.
[0278] "Information terminal" is a device having a function of collecting voice information and transmitting it to a data processing device.
[0279] "Data processing device" is an electronic device having computing power for analyzing received voice information.
[0280] "Learning algorithm" is an algorithm for analyzing the characteristics of voice information using machine learning to determine the possibility of it being synthesized voice.
[0281] "Information processing means" is a function for analyzing received voice information and executing a process for determining its validity
[0282] "Signal generation means" is means having a function of generating a warning signal when it is determined that voice information has been modified.
[0283] "Notification means" is a function for visually or auditorily transmitting a warning to a user on an information terminal that has received a warning signal.
[0284] "User" is a human who receives and operates voice information or an entity intended for its use.
[0285] The "Improvement Report" is the user feedback based on the warning message and is information for improving the accuracy of the learning algorithm.
[0286] The present invention is a system for detecting fraud in voice information, which consists of an information terminal, a data processing device, and a user. The information terminal is composed of devices such as smartphones and wearable devices, and voice information is collected thereby. The voice information is temporarily stored in a buffer in real time and is periodically transmitted to the data processing device as data packets.
[0287] The data processing device functions as a server. When it receives voice information, it uses a pre-trained learning algorithm to analyze the characteristics of the voice information in detail. The possibility that it is synthetic voice is quantified by the analysis, and when it exceeds a specific threshold, it is determined as modified voice. At this time, the server generates a warning signal, and the warning signal is transmitted to the information terminal. The information terminal receives this warning signal and gives a visual or auditory warning to the user.
[0288] The user confirms the warning and is prompted to be cautious about fraud and irregularities. If a misjudgment occurs in actual use, the user can send feedback from the information terminal to the server. This feedback is received by the server and utilized for improving the learning algorithm, enabling more accurate analysis.
[0289] As a specific example, when the user answers a call and the system analyzes the call and detects synthetic voice suspected of fraud, the user can prevent fraud by receiving a warning immediately. An example of the prompt sentence for voice analysis at this time is "Please determine whether this voice is real or synthetic. If the score exceeds the threshold, please generate a warning message."
[0290] The flow of the specific process in Application Example 1 will be described with reference to FIG. 12.
[0291] Step 1:
[0292] The device collects voice information in real time. When a voice call or voice message begins, the voice signal is stored in a buffer as digital data. At this stage, the input is voice information, and the output is digital voice data. The collected voice data is packetized at regular intervals.
[0293] Step 2:
[0294] The terminal sends packetized voice data to the server. The data is securely transferred over the network and received by the server. The input is packetized voice data, and the output is the data received on the server side. The transfer is secured using encrypted communication.
[0295] Step 3:
[0296] The server inputs the received audio data into a machine learning model for analysis. It extracts audio characteristics as features and scores the likelihood that the audio is synthesized speech. The input is the received audio data, and the output is the synthesized speech score. This analysis is performed using a sophisticated algorithm based on a generative AI model.
[0297] Step 4:
[0298] The server generates a warning signal if the score exceeds a set threshold. This signal indicates a high probability of synthesized speech. The input is the analysis score, and the output is the warning signal. Signal generation is performed in real time, requiring immediate response.
[0299] Step 5:
[0300] The terminal receives a warning signal sent from the server and notifies the user. The warning is presented visually as an on-screen alert and audibly as an audio notification. The input is the warning signal, and the output is a warning display to the user. The most suitable notification method is selected depending on the situation.
[0301] Step 6:
[0302] The user receives a warning and determines the authenticity of the voice information. The user checks the warning and interrupts the call or sends feedback as necessary. The input is the warning to the user, and the output is the user's corresponding actions. In case of misjudgment due to feedback, it helps improve the system.
[0303] Step 7:
[0304] The server receives feedback from the user and uses it to improve the machine learning model. The model is retrained using new data to achieve more accurate analysis. The input is the feedback from the user, and the output is the improved learning algorithm. This process is continuously carried out to improve the performance of the model.
[0305] Furthermore, an emotion engine for estimating the user's emotion may be combined. That is, the specific processing unit 290 may estimate the user's emotion using the emotion recognition model 59 and perform specific processing using the user's emotion.
[0306] The present invention is a system that combines a terminal for collecting and analyzing voice data, a server having means for determining whether the voice data is processed voice, and an emotion engine for recognizing the user's emotion and appropriately adjusting a warning message.
[0307] The terminal monitors in real time voice calls and messages received by the user and temporarily stores the voice data in a buffer. These data are packetized at regular intervals and transmitted to the server for analysis.
[0308] The server inputs the received voice data into a machine learning model. This model analyzes the characteristics of the voice and scores the possibility of it being synthesized voice. If the score exceeds the threshold, the voice is determined to be processed voice, and a warning signal is transmitted to the terminal.
[0309] The emotion engine analyzes the user's voice tone, speaking speed, facial expressions, etc., to recognize the user's emotional state in real time. Based on the recognized emotional information, it is possible to adjust how warning messages are presented (for example, the emphasis and type of the message).
[0310] For example, during a normal phone call, the server analyzes the received audio and determines that it may be synthesized speech. If the emotion engine then recognizes that the user's voice sounds tense, it will alert the user by displaying a more emphasized warning message on the device.
[0311] This system not only combats the risk of voice fraud but also enables flexible responses that respond to the user's emotions, providing more personalized protection.
[0312] The following describes the processing flow.
[0313] Step 1:
[0314] The device monitors the audio data in real time when the user is receiving a phone call or voice message, and temporarily stores it in a buffer. This audio data is then packetized in preparation for subsequent processing.
[0315] Step 2:
[0316] The terminal sends the voice data stored in the buffer to the server at regular time intervals. At this time, the call ID and timestamp are also sent as metadata.
[0317] Step 3:
[0318] The server inputs the received audio data into a machine learning model to analyze the audio's characteristics. The model scores the likelihood that the audio is synthesized speech, and if it exceeds a threshold, it is classified as processed speech.
[0319] Step 4:
[0320] If the server determines that the audio is processed, it generates a warning signal and sends it to the terminal. The warning signal includes instructions to alert the user.
[0321] Step 5:
[0322] The device activates the emotion engine based on the received warning signal. The emotion engine analyzes the user's voice tone, speaking speed, and facial expressions to recognize the user's emotional state.
[0323] Step 6:
[0324] Based on the emotional information recognized by the emotion engine, the device adjusts how warning messages are presented. For example, if the user is showing signs of anxiety, the warning message is visually highlighted.
[0325] Step 7:
[0326] The user reviews the warning message from their device and decides whether to continue or end the call. If necessary, they can send feedback from their device to the server if there is a misjudgment. This feedback is used by the server to improve the machine learning model.
[0327] (Example 2)
[0328] Next, we will describe Example 2. In the following description, the data processing device 12 will be referred to as the "server" and the smart glasses 214 will be referred to as the "terminal".
[0329] In modern society, voice-based fraud is on the rise, but effective detection systems to combat it are still not well-established. Furthermore, existing systems do not take into account the user's emotional state, and inappropriate warnings can lead to unnecessary stress and confusion. Therefore, there is a need for technology that can quickly and accurately identify manipulated voices and provide warnings tailored to the user's emotional state.
[0330] The identification process performed by the identification processing unit 290 of the data processing device 12 in Example 2 is realized by the following means.
[0331] In this invention, the server includes means for analyzing voice information using a computational model and determining whether the voice has been processed; means for generating and transmitting a warning signal based on the determination result; and means for evaluating the user's emotional state using an emotion analysis device and adjusting the method of presenting the warning based on the emotional state. This enables rapid and accurate detection of processed voices and personalized warning presentation that takes the user's emotions into consideration.
[0332] "Audio information" refers to the digital or analog representation of sound waves that have been transmitted or stored.
[0333] An "information processing device" is a combination of hardware and software used to compute, analyze, or transform data.
[0334] A "computational model" is an algorithm or mathematical formula used to analyze the characteristics of data and make specific decisions based on the results.
[0335] A "warning signal" is a signal generated to inform the user or system of a specific condition or abnormality.
[0336] An "emotion analysis device" is a device or system that evaluates or estimates a user's emotional state based on information such as their voice and facial expressions.
[0337] "User" refers to the individual who operates or uses the system or product.
[0338] "Numerical evaluation" refers to numerical results calculated to quantify specific characteristics or states.
[0339] "Opinions" refers to feedback and reactions provided by users.
[0340] This invention is a system for monitoring voice calls and messages received by a user and analyzing the voice information. Its primary purpose is to collect and analyze voice information and provide warnings that take into account the user's emotional state. The following describes a specific implementation of this system.
[0341] The terminal acquires the user's voice calls and messages in real time. To do this, it captures voice information using a voice input device and digital signal processing technology. The collected voice information is temporarily stored in the terminal's memory and then transmitted to a server using network communication technology.
[0342] The server is equipped with a computing device that analyzes transmitted audio information using a computational model. This analysis utilizes machine learning libraries (e.g., TensorFlow and PyTorch) to analyze the characteristics of the audio and determine whether it is synthesized speech. If it is determined to be synthesized speech, a warning signal is generated and sent to the terminal.
[0343] The emotion analysis device analyzes emotional data from voice information and the user's facial expressions to understand the user's emotional state. This information is transmitted to a server and used to display warning signals. For example, if the device detects a state of tension, it will display a more emphasized warning message to prompt appropriate action.
[0344] As a concrete example, suppose a user makes a voice call, and the server analyzes the voice information and determines that it is synthesized speech. In this situation, the emotion analysis device can sense the tension in the user's voice and issue a command to display an emphasized warning message on the device. This allows the user to immediately respond to the risk of voice fraud.
[0345] An example of a prompt to input into a generative AI model is, "Analyze the user's emotion from their voice tone and facial expressions, and create a message appropriate to the situation." This prompt helps improve the accuracy of emotion analysis and make warning messages more appropriate.
[0346] The flow of the specific processing in Example 2 will be explained using Figure 13.
[0347] Step 1:
[0348] The terminal monitors incoming voice calls and messages in real time. When voice input occurs, the terminal uses the microphone to capture sound wave data and converts it into a digital format. This data is temporarily stored in memory and packetized. The input is voice data, and the output is packetized digital data.
[0349] Step 2:
[0350] The terminal sends packetized voice data to the server at regular intervals. Efficient data transfer using network protocols is required to minimize latency. The input is packetized digital data, and the output is voice data received by the server.
[0351] Step 3:
[0352] The server acquires the received audio data and feeds it into a machine learning model. The computational model used here is, for example, trained using TensorFlow or PyTorch. It analyzes the characteristics of the audio, such as frequency components and waveform features, and scores the likelihood of it being synthesized speech. The input is the audio data sent to the server, and the output is the scored likelihood of it being synthesized speech.
[0353] Step 4:
[0354] The server compares the scoring result to a threshold. If the threshold is exceeded, the server determines that the voice is synthesized speech and generates a warning signal. This warning signal is then prepared to be sent to the terminal. The input is the scoring result, and the output is the warning signal.
[0355] Step 5:
[0356] The emotion analysis device acquires the user's voice tone and facial expression data to evaluate the user's emotional state. Data is acquired from audio or video cameras and analyzed in real time. Emotional states are classified, for example, into categories such as tension or calmness. The input is the user's voice and facial expression data, and the output is the evaluated emotional state.
[0357] Step 6:
[0358] The terminal references the warning signal received from the server and the emotional state from the emotion analyzer, and adjusts how the warning message is displayed. For example, if it determines that the user is in a state of tension, the terminal will highlight the warning message, providing an appropriate warning to the user. The input is the warning signal and emotional state, and the output is the adjusted warning message.
[0359] (Application Example 2)
[0360] Next, we will explain application example 2. In the following explanation, the data processing device 12 will be referred to as the "server," and the smart glasses 214 will be referred to as the "terminal."
[0361] In recent years, advancements in speech synthesis technology have increased the risk of general users being affected by voice fraud and disguised voices. However, currently, there are insufficient means to determine in real time whether a received voice is synthesized or not. Furthermore, there is a need for a system that can appropriately adjust warning messages by taking into account the user's emotional state, thereby preventing excessive warnings and providing effective alerts.
[0362] The specific processing performed by the specific processing unit 290 of the data processing device 12 in Application Example 2 is realized by the following means.
[0363] In this invention, the server includes means for a terminal that collects audio data and transmits it to an information processing device for analysis; means for analyzing the received audio data using a machine learning model and determining whether it is synthesized audio; and emotion analysis means for recognizing the user's emotional state and adjusting the warning message based on that emotional state. This makes it possible to determine whether the audio data is synthesized or natural and to provide an optimal warning based on the user's emotional state.
[0364] "Audio data" refers to information recorded in digital format, which is used for analysis and transmission.
[0365] An "information processing device" is an electronic device used to perform data analysis and calculations, and includes servers and computers.
[0366] A "machine learning model" is a collection of algorithms and computational methods used for analysis and prediction, which learn features from large amounts of data.
[0367] "Emotional state" refers to the user's emotional response and is determined by indicators such as tone of voice, speaking speed, and facial expressions.
[0368] A "warning signal" is a signal used to alert the user when certain conditions are met, and is communicated through audio, on-screen displays, or other means.
[0369] A "terminal" is an input / output device used by a user, and includes smartphones and tablets.
[0370] The system that implements this application involves real-time analysis of audio data and adjustment of warning messages based on the user's emotional state. The server receives audio data and performs analysis using a machine learning model. Specifically, it extracts features from the audio data using digital signal processing technology, applies these features to a machine learning algorithm, and scores the likelihood that the audio is synthesized speech.
[0371] The terminal receives a warning signal sent from the server and uses an emotion analysis engine to analyze the user's voice tone and speaking speed. This allows the terminal to recognize the user's emotional state and adjust the warning message accordingly. For example, if the user is feeling anxious, the warning message can be displayed with greater emphasis.
[0372] The program is implemented using programming languages such as Python, and utilizes the speech_recognition library for speech analysis. Data communication is performed via the HTTP protocol. The user interface is designed to display both audio and visual warnings.
[0373] For example, if a phone call received by a user is suspected to be using synthesized speech, the terminal receives a corresponding warning signal. At this time, the emotion engine analyzes the user's emotional state, and only if it recognizes anxiety or tension, it highlights the warning message to prompt the user to take immediate action. The prompt text to be input to the generating AI model is as follows:
[0374] "Please consider specific implementation methods for a function that acquires audio data in real time and analyzes the potential of synthesized speech."
[0375] The flow of a specific process in Application Example 2 will be explained using Figure 14.
[0376] Step 1:
[0377] The terminal temporarily stores the voice call data received from the user in a buffer. It receives real-time captured voice data as input, converts it to a digital data format, and saves it in preparation for subsequent processing.
[0378] Step 2:
[0379] The terminal packets the voice data stored in a buffer at regular intervals and sends it to the server. The input is voice call data, which is then divided into small data packets and sent to the server via the network—a data processing step.
[0380] Step 3:
[0381] The server interprets the received audio data and feeds it into a machine learning model. The input here is data packets sent from the terminal, and the server performs data calculations by extracting this audio data and analyzing it using a machine learning model.
[0382] Step 4:
[0383] The server uses a machine learning model to analyze the audio data and score its likelihood of being synthesized speech. Based on the analysis, it determines whether the audio is synthesized speech, and if the score exceeds a threshold, it generates a warning signal. The output of this step is the scoring based on the analyzed data.
[0384] Step 5:
[0385] The server sends the generated warning signal to the terminal. It takes a scoring-based judgment result as input and performs data communication to prompt appropriate action by sending that result to the user's terminal.
[0386] Step 6:
[0387] The terminal analyzes the user's emotional state using an emotion analysis engine based on the received warning signal. The input is the warning signal received from the server, and in addition, the user's voice tone and speaking speed are analyzed to process the data in order to identify the user's emotions.
[0388] Step 7:
[0389] The device adjusts and displays warning messages on the screen based on the user's emotional state. The input is the result of the analysis of the user's emotional state, and the output is a warning message that matches the result. Specifically, if the user is feeling anxious, the message will be highlighted to draw their attention.
[0390] The specific processing unit 290 transmits the result of the specific processing to the smart glasses 214. In the smart glasses 214, the control unit 46A causes the speaker 240 to output the result of the specific processing. The microphone 238 acquires audio indicating user input for the result of the specific processing. The control unit 46A transmits the audio data indicating user input acquired by the microphone 238 to the data processing unit 12. In the data processing unit 12, the specific processing unit 290 acquires the audio data.
[0391] Data generation model 58 is a type of so-called generative AI (Artificial Intelligence). One example of data generation model 58 is ChatGPT (Internet search<URL: https: / / openai.com / blog / chatgpt> ), Gemini (Internet search) <url: https: gemini.google.com ?hl="ja">Examples of generative AI include the following. The data generation model 58 is obtained by performing deep learning on a neural network. The data generation model 58 is input with prompts containing instructions, and with inference data such as audio data representing speech, text data representing text, and image data representing images. The data generation model 58 infers from the input inference data according to the instructions indicated by the prompts, and outputs the inference results in data formats such as audio data and text data. Here, inference refers to, for example, analysis, classification, prediction, and / or summarization.
[0392] In the above embodiment, an example was given in which specific processing is performed by the data processing device 12, but the technology of this disclosure is not limited thereto, and the specific processing may also be performed by the smart glasses 214.
[0393] [Third Embodiment]
[0394] Figure 5 shows an example of the configuration of the data processing system 310 according to the third embodiment.
[0395] As shown in Figure 5, the data processing system 310 includes a data processing device 12 and a headset terminal 314. An example of the data processing device 12 is a server.
[0396] The data processing device 12 comprises a computer 22, a database 24, and a communication interface 26. The computer 22 is an example of a "computer" related to the technology of this disclosure. The computer 22 comprises a processor 28, RAM 30, and storage 32. The processor 28, RAM 30, and storage 32 are connected to a bus 34. The database 24 and the communication interface 26 are also connected to the bus 34. The communication interface 26 is connected to a network 54. An example of the network 54 is a WAN (Wide Area Network) and / or a LAN (Local Area Network).
[0397] The headset terminal 314 includes a computer 36, a microphone 238, a speaker 240, a camera 42, a communication interface 44, and a display 343. The computer 36 includes a processor 46, RAM 48, and storage 50. The processor 46, RAM 48, and storage 50 are connected to a bus 52. The microphone 238, speaker 240, camera 42, and display 343 are also connected to the bus 52.
[0398] The microphone 238 receives voice signals from the user 20 and receives instructions from the user 20. The microphone 238 captures the voice signals from the user 20, converts the captured voice into audio data, and outputs it to the processor 46. The speaker 240 outputs audio according to the instructions from the processor 46.
[0399] Camera 42 is a small digital camera equipped with an optical system including a lens, aperture, and shutter, and an image sensor such as a CMOS (Complementary Metal-Oxide-Semiconductor) image sensor or a CCD (Charge Coupled Device) image sensor, and captures images of the area around the user 20 (for example, an imaging range defined by a field of view equivalent to the width of a typical healthy person's field of vision).
[0400] Communication interface 44 is connected to network 54. Communication interfaces 44 and 26 are responsible for the exchange of various information between processor 46 and processor 28 via network 54. The exchange of various information between processor 46 and processor 28 using communication interfaces 44 and 26 is performed in a secure manner.
[0401] Figure 6 shows an example of the main functions of the data processing device 12 and the headset terminal 314. As shown in Figure 6, the data processing device 12 performs specific processing using the processor 28. The storage 32 stores the specific processing program 56.
[0402] The specific processing program 56 is an example of a "program" relating to the technology of this disclosure. The processor 28 reads the specific processing program 56 from the storage 32 and executes the read specific processing program 56 on the RAM 30. The specific processing is realized by the processor 28 operating as a specific processing unit 290 in accordance with the specific processing program 56 executed on the RAM 30.
[0403] The storage 32 stores the data generation model 58 and the emotion identification model 59. The data generation model 58 and the emotion identification model 59 are used by the identification processing unit 290.
[0404] In the headset terminal 314, the processor 46 performs the reception output processing. The storage 50 stores the reception output program 60. The processor 46 reads the reception output program 60 from the storage 50 and executes the read reception output program 60 on the RAM 48. The reception output processing is realized by the processor 46 operating as a control unit 46A according to the reception output program 60 executed on the RAM 48.
[0405] Next, the specific processing performed by the specific processing unit 290 of the data processing device 12 will be described. In the following description, the data processing device 12 will be referred to as the "server" and the headset terminal 314 will be referred to as the "terminal".
[0406] This invention is a system in which a terminal, a server, and a user work together. The terminal, with the user's consent, collects voice data from voice calls and voice messages in real time. The collected voice data is temporarily stored in a buffer, and the packetized voice data is sent to the server at regular intervals.
[0407] The server processes the received audio data in real time and inputs it into a machine learning model. This model analyzes the characteristics of the audio and scores the likelihood that it is synthesized speech. If the score exceeds a set threshold, the server determines that the audio is processed speech.
[0408] Once a determination is made, the server generates a warning signal and sends it to the terminal. The terminal receives this signal and displays a warning message to the user visually or audibly. The user reviews the warning and decides whether to continue or end the call based on that information.
[0409] For example, if a user receives a phone call that they suspect is a scam, their device sends the call to a server. If the server determines that it is a synthesized voice, a warning is displayed on the device, and the user reviews the warning and makes an appropriate decision. If there is a misidentification, the user can send feedback to the server through their device, and the server can use this feedback to improve its machine learning model and perform more accurate analysis.
[0410] This system allows users to protect themselves from voice scams and use voice services safely.
[0411] The following describes the processing flow.
[0412] Step 1:
[0413] The device monitors audio data in real time while the user is receiving phone calls or voice messages. This data is temporarily stored in a buffer and prepared for processing.
[0414] Step 2:
[0415] The terminal divides the voice data stored in the buffer into packets at regular time intervals and sends them to the server. At this time, voice metadata (such as call ID and timestamp) is also sent.
[0416] Step 3:
[0417] The server sequentially inputs the received audio packets into a machine learning model. The model then begins the process of analyzing the features of the audio data and scoring the likelihood that it is synthesized speech.
[0418] Step 4:
[0419] The server evaluates the score obtained through analysis and determines that the audio is processed if it exceeds a set threshold. The determination result is recorded as a log.
[0420] Step 5:
[0421] If the server detects that the audio has been altered, it generates a warning signal and sends this signal to the terminal. This warning signal includes a message to inform the user.
[0422] Step 6:
[0423] The device displays a visual or audible warning to the user based on the received warning signal. The user reviews the warning and decides whether to end the call.
[0424] Step 7:
[0425] After making a decision based on a warning, users can send feedback from their device to the server to report any misjudgments, if necessary. This feedback is used to improve the machine learning model.
[0426] (Example 1)
[0427] Next, we will describe Example 1. In the following description, the data processing device 12 will be referred to as the "server," and the headset-type terminal 314 will be referred to as the "terminal."
[0428] In today's world, fraud and information manipulation using synthesized voices are on the rise, making it crucial to determine the authenticity of voice data in real time. However, current technology makes it difficult to distinguish synthesized voices simply by listening, potentially causing harm to users. Solving this problem and enabling users to communicate with voices with confidence is essential.
[0429] The identification process performed by the identification processing unit 290 of the data processing device 12 in Example 1 is realized by the following means.
[0430] In this invention, the server includes means for transmitting voice information to an analysis device, means for analyzing the received voice information using a machine learning algorithm to determine whether it is processed voice, and means for generating a warning signal based on the determination result and transmitting the warning signal to a device. This prevents fraud using synthesized voices and enables users to use voice information safely in real time.
[0431] "Audio information" refers to data that represents audio in digital format, and this generally includes the content of voice calls and voice messages.
[0432] "Device" refers to an electronic device that collects voice information and transmits it to a server, and includes mobile devices such as smartphones and tablets.
[0433] An "analysis device" refers to a computer system that receives audio information and analyzes its content.
[0434] A "machine learning algorithm" refers to a model used to analyze audio data, understand its characteristics, and determine whether it is synthesized speech.
[0435] A "warning signal" refers to a signal generated to alert the user when there is a possibility that the audio information is synthesized.
[0436] "Response" refers to information that users can use to provide feedback on the synthesized speech's evaluation results and the system's operation.
[0437] This invention relates to a system for determining the authenticity of audio information in real time. The system operates in cooperation with three parties: a terminal used by the user, a server that analyzes the audio data, and the user who receives the warning.
[0438] The device collects voice information from voice calls and voice messages in real time with the user's consent. The device used is a portable electronic device such as a smartphone or tablet. The collected voice information is packetized at regular intervals and sent to a server via a communication protocol. To ensure smooth data transmission, the TCP / IP protocol is commonly used.
[0439] The server applies machine learning algorithms to analyze the received audio information. These algorithms are developed using frameworks such as TensorFlow or PyTorch. The server analyzes the characteristics of the audio information and quantifies the likelihood that it is synthesized speech through scoring. If the score exceeds a certain threshold, it is determined to be synthesized speech and a warning signal is generated. This warning signal is sent to the terminal via the HTTP protocol or similar.
[0440] Based on the received warning signal, the device notifies the user of a warning message visually or audibly. For example, it can display "Warning: This may be a synthesized voice" on the device's screen, or emit a warning sound. This allows the user to make informed decisions to prevent fraud.
[0441] For example, if a user receives a suspicious call, the system analyzes the audio information of the call and makes a judgment. If the server determines that the voice is synthesized speech, a warning is displayed on the terminal, allowing the user to take appropriate action. The user can also provide feedback on the system's judgment, which is used to improve the machine learning algorithm.
[0442] Examples of prompt messages could include, "How should we warn the user if the call is potentially fraudulent?" or "Please tell us how to improve the voice analysis system using machine learning models." By using this system, users can communicate via voice with peace of mind.
[0443] The flow of the specific processing in Example 1 will be explained using Figure 11.
[0444] Step 1:
[0445] With the user's consent, the device collects voice information from voice calls and voice messages in real time. Specifically, the device uses the smartphone's built-in microphone to convert the voice information into a digital format. As a result, the input acquired by the device is analog voice data, and the output is digitized voice information.
[0446] Step 2:
[0447] The terminal temporarily stores the collected audio information in memory. Specifically, it packets the real-time acquired audio data at regular intervals (e.g., every second). The input is the real-time collected audio information, and the output is the data of that audio information divided into smaller packets.
[0448] Step 3:
[0449] The terminal sends packetized voice information to the server. Specifically, this involves sending data to the server's receiving port using the TCP / IP protocol. The input is packetized voice information, and the output is the voice data sent to the server via the network.
[0450] Step 4:
[0451] The server inputs the received audio data into a machine learning algorithm. The server provides the audio data to a model developed using, for example, TensorFlow, and extracts audio features. The input is the audio data received by the server, and the output is the audio data features and analysis results.
[0452] Step 5:
[0453] The server uses the output of a machine learning algorithm to score whether the speech is likely to be synthesized. Here, if the score exceeds a certain threshold, it is classified as "synthesized speech." The input is the analyzed features, and the output is the scoring result and the classification result.
[0454] Step 6:
[0455] The server generates a warning signal if it determines that the voice is synthesized. This signal contains a message such as "Caution: This voice may have been synthesized." In its specific operation, it generates a warning signal and creates a data structure to send it to the terminal. The input is the detection result, and the output is the warning signal.
[0456] Step 7:
[0457] The terminal notifies the user based on the received warning signal. Specifically, the terminal attracts the user's attention by displaying text on the screen or emitting a warning sound. The input is a warning signal sent from the server, and the output is a warning message that the user sees or hears.
[0458] (Application Example 1)
[0459] Next, we will explain Application Example 1. In the following explanation, the data processing device 12 will be referred to as the "server," and the headset-type terminal 314 will be referred to as the "terminal."
[0460] In recent years, advancements in voice technology have led to an increase in fraud and illegal activities using synthesized voices. This increases the likelihood of users finding themselves in dangerous situations. However, technology to accurately determine whether voice information is synthesized is still insufficient, and users lack effective countermeasures against fraudulent voices. Against this backdrop, the present invention aims to provide a system that can quickly and accurately determine whether voice information is synthesized and warn users.
[0461] The specific processing performed by the specific processing unit 290 of the data processing device 12 in Application Example 1 is realized by the following means.
[0462] In this invention, the server includes information processing means for analyzing voice information using a learning algorithm and determining whether it is modified voice; signal generation means for generating a warning signal based on the determination result and transmitting it to an information terminal; and means for receiving improvement reports from users who have received warnings and applying them to the learning algorithm. This enables users to quickly detect fraudulent synthesized voices and take appropriate action.
[0463] "Voice information" refers to information data recorded as sound waves, including voice calls and voice messages.
[0464] An "information terminal" is a device that has the function of collecting voice information and transmitting it to a data processing device.
[0465] A "data processing device" is an electronic device with computing power for analyzing received audio information.
[0466] A "learning algorithm" is an algorithm that uses machine learning to analyze the characteristics of audio information and determine whether it is likely to be synthesized speech.
[0467] "Information processing means" refers to a function that performs a process of analyzing received audio information and determining its validity.
[0468] A "signal generation means" is a means that has the function of generating a warning signal when it is determined that audio information has been altered.
[0469] A "notification means" is a function that transmits a warning to the user visually or audibly via an information terminal that receives a warning signal.
[0470] A "user" is a person who receives and manipulates voice information, or an entity intended to use such information.
[0471] An "improvement report" is user feedback based on warning messages, and it provides information to improve the accuracy of the learning algorithm.
[0472] This invention relates to a system for detecting fraudulent voice information, and consists of three components: an information terminal, a data processing device, and a user. The information terminal is comprised of a device such as a smartphone or wearable device, through which voice information is collected. The voice information is temporarily stored in a buffer in real time and periodically transmitted to the data processing device as data packets.
[0473] The data processing unit functions as a server and, upon receiving audio information, uses a pre-trained learning algorithm to analyze the characteristics of that audio information in detail. The analysis quantifies the likelihood that the audio is synthesized speech, and if it exceeds a certain threshold, it is determined to be modified speech. At this point, the server generates a warning signal, which is transmitted to the information terminal. The information terminal receives this warning signal and provides the user with a visual or auditory warning.
[0474] Users are alerted to warnings and encouraged to be vigilant against fraud and misconduct. If a misidentification occurs during actual use, users can send feedback from their information terminal to the server. This feedback is received by the server and used to improve the learning algorithm, enabling more accurate analysis.
[0475] As a concrete example, when a user receives a phone call, the system analyzes the call and, if it detects a synthesized voice suspected of being fraudulent, the user receives an instant warning, thus preventing fraud. An example of a prompt message for this voice analysis would be, "Please determine whether this voice is genuine or synthesized. If the score exceeds the threshold, generate a warning message."
[0476] The flow of a specific process in Application Example 1 will be explained using Figure 12.
[0477] Step 1:
[0478] The device collects voice information in real time. When a voice call or voice message begins, the voice signal is stored in a buffer as digital data. At this stage, the input is voice information, and the output is digital voice data. The collected voice data is packetized at regular intervals.
[0479] Step 2:
[0480] The terminal sends packetized voice data to the server. The data is securely transferred over the network and received by the server. The input is packetized voice data, and the output is the data received on the server side. The transfer is secured using encrypted communication.
[0481] Step 3:
[0482] The server inputs the received audio data into a machine learning model for analysis. It extracts audio characteristics as features and scores the likelihood that the audio is synthesized speech. The input is the received audio data, and the output is the synthesized speech score. This analysis is performed using a sophisticated algorithm based on a generative AI model.
[0483] Step 4:
[0484] The server generates a warning signal if the score exceeds a set threshold. This signal indicates a high probability of synthesized speech. The input is the analysis score, and the output is the warning signal. Signal generation is performed in real time, requiring immediate response.
[0485] Step 5:
[0486] The terminal receives a warning signal sent from the server and notifies the user. The warning is presented visually as an on-screen alert and audibly as an audio notification. The input is the warning signal, and the output is a warning display to the user. The most suitable notification method is selected depending on the situation.
[0487] Step 6:
[0488] The user receives a warning and determines the veracity of the audio information. They review the warning and, if necessary, interrupt the call or send feedback. The input is the warning to the user, and the output is the user's response. Feedback helps improve the system in case of misjudgments.
[0489] Step 7:
[0490] The server receives feedback from users and uses it to improve the machine learning model. The model is retrained with new data to achieve more accurate analysis. The input is user feedback, and the output is the improved learning algorithm. This process is continuous, improving the model's performance.
[0491] Furthermore, an emotion engine that estimates the user's emotions may be incorporated. That is, the identification processing unit 290 may use the emotion identification model 59 to estimate the user's emotions and perform identification processing using the user's emotions.
[0492] The present invention is a system that combines a terminal for collecting and analyzing voice data, a server equipped with means for determining whether the voice data is processed voice, and an emotion engine that recognizes the user's emotions and appropriately adjusts warning messages.
[0493] The terminal monitors voice calls and messages received by the user in real time and temporarily stores the voice data in a buffer. This data is packetized at regular intervals and sent to a server for analysis.
[0494] The server feeds the received audio data into a machine learning model. This model analyzes the characteristics of the audio and scores the likelihood that it is synthesized speech. If the score exceeds a threshold, the server determines that the audio is processed speech and sends a warning signal to the terminal.
[0495] The emotion engine analyzes the user's voice tone, speaking speed, facial expressions, etc., to recognize the user's emotional state in real time. Based on the recognized emotional information, it is possible to adjust how warning messages are presented (for example, the emphasis and type of the message).
[0496] For example, during a normal phone call, the server analyzes the received audio and determines that it may be synthesized speech. If the emotion engine then recognizes that the user's voice sounds tense, it will alert the user by displaying a more emphasized warning message on the device.
[0497] This system not only combats the risk of voice fraud but also enables flexible responses that respond to the user's emotions, providing more personalized protection.
[0498] The following describes the processing flow.
[0499] Step 1:
[0500] The device monitors the audio data in real time when the user is receiving a phone call or voice message, and temporarily stores it in a buffer. This audio data is then packetized in preparation for subsequent processing.
[0501] Step 2:
[0502] The terminal sends the voice data stored in the buffer to the server at regular time intervals. At this time, the call ID and timestamp are also sent as metadata.
[0503] Step 3:
[0504] The server inputs the received audio data into a machine learning model to analyze the audio's characteristics. The model scores the likelihood that the audio is synthesized speech, and if it exceeds a threshold, it is classified as processed speech.
[0505] Step 4:
[0506] If the server determines that the audio is processed, it generates a warning signal and sends it to the terminal. The warning signal includes instructions to alert the user.
[0507] Step 5:
[0508] The device activates the emotion engine based on the received warning signal. The emotion engine analyzes the user's voice tone, speaking speed, and facial expressions to recognize the user's emotional state.
[0509] Step 6:
[0510] Based on the emotional information recognized by the emotion engine, the device adjusts how warning messages are presented. For example, if the user is showing signs of anxiety, the warning message is visually highlighted.
[0511] Step 7:
[0512] The user reviews the warning message from their device and decides whether to continue or end the call. If necessary, they can send feedback from their device to the server if there is a misjudgment. This feedback is used by the server to improve the machine learning model.
[0513] (Example 2)
[0514] Next, we will describe Example 2. In the following description, the data processing device 12 will be referred to as the "server," and the headset-type terminal 314 will be referred to as the "terminal."
[0515] In modern society, voice-based fraud is on the rise, but effective detection systems to combat it are still not well-established. Furthermore, existing systems do not take into account the user's emotional state, and inappropriate warnings can lead to unnecessary stress and confusion. Therefore, there is a need for technology that can quickly and accurately identify manipulated voices and provide warnings tailored to the user's emotional state.
[0516] The identification process performed by the identification processing unit 290 of the data processing device 12 in Example 2 is realized by the following means.
[0517] In this invention, the server includes means for analyzing voice information using a computational model and determining whether the voice has been processed; means for generating and transmitting a warning signal based on the determination result; and means for evaluating the user's emotional state using an emotion analysis device and adjusting the method of presenting the warning based on the emotional state. This enables rapid and accurate detection of processed voices and personalized warning presentation that takes the user's emotions into consideration.
[0518] "Audio information" refers to the digital or analog representation of sound waves that have been transmitted or stored.
[0519] An "information processing device" is a combination of hardware and software used to compute, analyze, or transform data.
[0520] A "computational model" is an algorithm or mathematical formula used to analyze the characteristics of data and make specific decisions based on the results.
[0521] A "warning signal" is a signal generated to inform the user or system of a specific condition or abnormality.
[0522] An "emotion analysis device" is a device or system that evaluates or estimates a user's emotional state based on information such as their voice and facial expressions.
[0523] "User" refers to the individual who operates or uses the system or product.
[0524] "Numerical evaluation" refers to numerical results calculated to quantify specific characteristics or states.
[0525] "Opinions" refers to feedback and reactions provided by users.
[0526] This invention is a system for monitoring voice calls and messages received by a user and analyzing the voice information. Its primary purpose is to collect and analyze voice information and provide warnings that take into account the user's emotional state. The following describes a specific implementation of this system.
[0527] The terminal acquires the user's voice calls and messages in real time. To do this, it captures voice information using a voice input device and digital signal processing technology. The collected voice information is temporarily stored in the terminal's memory and then transmitted to a server using network communication technology.
[0528] The server is equipped with a computing device that analyzes transmitted audio information using a computational model. This analysis utilizes machine learning libraries (e.g., TensorFlow and PyTorch) to analyze the characteristics of the audio and determine whether it is synthesized speech. If it is determined to be synthesized speech, a warning signal is generated and sent to the terminal.
[0529] The emotion analysis device analyzes emotional data from voice information and the user's facial expressions to understand the user's emotional state. This information is transmitted to a server and used to display warning signals. For example, if the device detects a state of tension, it will display a more emphasized warning message to prompt appropriate action.
[0530] As a concrete example, suppose a user makes a voice call, and the server analyzes the voice information and determines that it is synthesized speech. In this situation, the emotion analysis device can sense the tension in the user's voice and issue a command to display an emphasized warning message on the device. This allows the user to immediately respond to the risk of voice fraud.
[0531] An example of a prompt to input into a generative AI model is, "Analyze the user's emotion from their voice tone and facial expressions, and create a message appropriate to the situation." This prompt helps improve the accuracy of emotion analysis and make warning messages more appropriate.
[0532] The flow of the specific processing in Example 2 will be explained using Figure 13.
[0533] Step 1:
[0534] The terminal monitors incoming voice calls and messages in real time. When voice input occurs, the terminal uses the microphone to capture sound wave data and converts it into a digital format. This data is temporarily stored in memory and packetized. The input is voice data, and the output is packetized digital data.
[0535] Step 2:
[0536] The terminal sends packetized voice data to the server at regular intervals. Efficient data transfer using network protocols is required to minimize latency. The input is packetized digital data, and the output is voice data received by the server.
[0537] Step 3:
[0538] The server acquires the received audio data and feeds it into a machine learning model. The computational model used here is, for example, trained using TensorFlow or PyTorch. It analyzes the characteristics of the audio, such as frequency components and waveform features, and scores the likelihood of it being synthesized speech. The input is the audio data sent to the server, and the output is the scored likelihood of it being synthesized speech.
[0539] Step 4:
[0540] The server compares the scoring result to a threshold. If the threshold is exceeded, the server determines that the voice is synthesized speech and generates a warning signal. This warning signal is then prepared to be sent to the terminal. The input is the scoring result, and the output is the warning signal.
[0541] Step 5:
[0542] The emotion analysis device acquires the user's voice tone and facial expression data to evaluate the user's emotional state. Data is acquired from audio or video cameras and analyzed in real time. Emotional states are classified, for example, into categories such as tension or calmness. The input is the user's voice and facial expression data, and the output is the evaluated emotional state.
[0543] Step 6:
[0544] The terminal references the warning signal received from the server and the emotional state from the emotion analyzer, and adjusts how the warning message is displayed. For example, if it determines that the user is in a state of tension, the terminal will highlight the warning message, providing an appropriate warning to the user. The input is the warning signal and emotional state, and the output is the adjusted warning message.
[0545] (Application Example 2)
[0546] Next, we will explain application example 2. In the following explanation, the data processing device 12 will be referred to as the "server," and the headset-type terminal 314 will be referred to as the "terminal."
[0547] In recent years, advancements in speech synthesis technology have increased the risk of general users being affected by voice fraud and disguised voices. However, currently, there are insufficient means to determine in real time whether a received voice is synthesized or not. Furthermore, there is a need for a system that can appropriately adjust warning messages by taking into account the user's emotional state, thereby preventing excessive warnings and providing effective alerts.
[0548] The specific processing performed by the specific processing unit 290 of the data processing device 12 in Application Example 2 is realized by the following means.
[0549] In this invention, the server includes means for a terminal that collects audio data and transmits it to an information processing device for analysis; means for analyzing the received audio data using a machine learning model and determining whether it is synthesized audio; and emotion analysis means for recognizing the user's emotional state and adjusting the warning message based on that emotional state. This makes it possible to determine whether the audio data is synthesized or natural and to provide an optimal warning based on the user's emotional state.
[0550] "Audio data" refers to information recorded in digital format, which is used for analysis and transmission.
[0551] An "information processing device" is an electronic device used to perform data analysis and calculations, and includes servers and computers.
[0552] A "machine learning model" is a collection of algorithms and computational methods used for analysis and prediction, which learn features from large amounts of data.
[0553] "Emotional state" refers to the user's emotional response and is determined by indicators such as tone of voice, speaking speed, and facial expressions.
[0554] A "warning signal" is a signal used to alert the user when certain conditions are met, and is communicated through audio, on-screen displays, or other means.
[0555] A "terminal" is an input / output device used by a user, and includes smartphones and tablets.
[0556] The system that implements this application involves real-time analysis of audio data and adjustment of warning messages based on the user's emotional state. The server receives audio data and performs analysis using a machine learning model. Specifically, it extracts features from the audio data using digital signal processing technology, applies these features to a machine learning algorithm, and scores the likelihood that the audio is synthesized speech.
[0557] The terminal receives a warning signal sent from the server and uses an emotion analysis engine to analyze the user's voice tone and speaking speed. This allows the terminal to recognize the user's emotional state and adjust the warning message accordingly. For example, if the user is feeling anxious, the warning message can be displayed with greater emphasis.
[0558] The program is implemented using programming languages such as Python, and utilizes the speech_recognition library for speech analysis. Data communication is performed via the HTTP protocol. The user interface is designed to display both audio and visual warnings.
[0559] For example, if a phone call received by a user is suspected to be using synthesized speech, the terminal receives a corresponding warning signal. At this time, the emotion engine analyzes the user's emotional state, and only if it recognizes anxiety or tension, it highlights the warning message to prompt the user to take immediate action. The prompt text to be input to the generating AI model is as follows:
[0560] "Please consider specific implementation methods for a function that acquires audio data in real time and analyzes the potential of synthesized speech."
[0561] The flow of a specific process in Application Example 2 will be explained using Figure 14.
[0562] Step 1:
[0563] The terminal temporarily stores the voice call data received from the user in a buffer. It receives real-time captured voice data as input, converts it to a digital data format, and saves it in preparation for subsequent processing.
[0564] Step 2:
[0565] The terminal packets the voice data stored in a buffer at regular intervals and sends it to the server. The input is voice call data, which is then divided into small data packets and sent to the server via the network—a data processing step.
[0566] Step 3:
[0567] The server interprets the received audio data and feeds it into a machine learning model. The input here is data packets sent from the terminal, and the server performs data calculations by extracting this audio data and analyzing it using a machine learning model.
[0568] Step 4:
[0569] The server uses a machine learning model to analyze the audio data and score its likelihood of being synthesized speech. Based on the analysis, it determines whether the audio is synthesized speech, and if the score exceeds a threshold, it generates a warning signal. The output of this step is the scoring based on the analyzed data.
[0570] Step 5:
[0571] The server sends the generated warning signal to the terminal. It takes a scoring-based judgment result as input and performs data communication to prompt appropriate action by sending that result to the user's terminal.
[0572] Step 6:
[0573] The terminal analyzes the user's emotional state using an emotion analysis engine based on the received warning signal. The input is the warning signal received from the server, and in addition, the user's voice tone and speaking speed are analyzed to process the data in order to identify the user's emotions.
[0574] Step 7:
[0575] The device adjusts and displays warning messages on the screen based on the user's emotional state. The input is the result of the analysis of the user's emotional state, and the output is a warning message that matches the result. Specifically, if the user is feeling anxious, the message will be highlighted to draw their attention.
[0576] The specific processing unit 290 transmits the result of the specific processing to the headset terminal 314. In the headset terminal 314, the control unit 46A causes the speaker 240 and display 343 to output the result of the specific processing. The microphone 238 acquires audio indicating user input for the result of the specific processing. The control unit 46A transmits the audio data indicating user input acquired by the microphone 238 to the data processing unit 12. In the data processing unit 12, the specific processing unit 290 acquires the audio data.
[0577] Data generation model 58 is a type of so-called generative AI (Artificial Intelligence). One example of data generation model 58 is ChatGPT (Internet search<URL: https: / / openai.com / blog / chatgpt> ), Gemini (Internet search) <url: https: gemini.google.com ?hl="ja">Examples of generative AI include the following. The data generation model 58 is obtained by performing deep learning on a neural network. The data generation model 58 is input with prompts containing instructions, and with inference data such as audio data representing speech, text data representing text, and image data representing images. The data generation model 58 infers from the input inference data according to the instructions indicated by the prompts, and outputs the inference results in data formats such as audio data and text data. Here, inference refers to, for example, analysis, classification, prediction, and / or summarization.
[0578] In the above embodiment, an example was given in which specific processing is performed by the data processing device 12, but the technology of this disclosure is not limited thereto, and specific processing may also be performed by the headset terminal 314.
[0579] [Fourth Embodiment]
[0580] Figure 7 shows an example of the configuration of the data processing system 410 according to the fourth embodiment.
[0581] As shown in Figure 7, the data processing system 410 includes a data processing device 12 and a robot 414. An example of the data processing device 12 is a server.
[0582] The data processing device 12 comprises a computer 22, a database 24, and a communication interface 26. The computer 22 is an example of a "computer" related to the technology of this disclosure. The computer 22 comprises a processor 28, RAM 30, and storage 32. The processor 28, RAM 30, and storage 32 are connected to a bus 34. The database 24 and the communication interface 26 are also connected to the bus 34. The communication interface 26 is connected to a network 54. An example of the network 54 is a WAN (Wide Area Network) and / or a LAN (Local Area Network).
[0583] The robot 414 includes a computer 36, a microphone 238, a speaker 240, a camera 42, a communication interface 44, and a controlled object 443. The computer 36 includes a processor 46, RAM 48, and storage 50. The processor 46, RAM 48, and storage 50 are connected to a bus 52. The microphone 238, speaker 240, camera 42, and controlled object 443 are also connected to the bus 52.
[0584] The microphone 238 receives voice signals from the user 20 and receives instructions from the user 20. The microphone 238 captures the voice signals from the user 20, converts the captured voice into audio data, and outputs it to the processor 46. The speaker 240 outputs audio according to the instructions from the processor 46.
[0585] Camera 42 is a small digital camera equipped with an optical system including a lens, aperture, and shutter, and an image sensor such as a CMOS (Complementary Metal-Oxide-Semiconductor) image sensor or a CCD (Charge Coupled Device) image sensor, and captures images of the area around the user 20 (for example, an imaging range defined by a field of view equivalent to the width of a typical healthy person's field of vision).
[0586] Communication interface 44 is connected to network 54. Communication interfaces 44 and 26 are responsible for the exchange of various information between processor 46 and processor 28 via network 54. The exchange of various information between processor 46 and processor 28 using communication interfaces 44 and 26 is performed in a secure manner.
[0587] The controlled object 443 includes a display device, LEDs in the eyes, and motors that drive the arms, hands, and feet. The posture and gestures of the robot 414 are controlled by controlling the motors of the arms, hands, and feet. Some of the robot 414's emotions can be expressed by controlling these motors. Furthermore, the robot 414's facial expressions can also be expressed by controlling the illumination state of the LEDs in its eyes.
[0588] Figure 8 shows an example of the main functions of the data processing device 12 and the robot 414. As shown in Figure 8, the data processing device 12 performs specific processing using the processor 28. The storage 32 stores the specific processing program 56.
[0589] The specific processing program 56 is an example of a "program" relating to the technology of this disclosure. The processor 28 reads the specific processing program 56 from the storage 32 and executes the read specific processing program 56 on the RAM 30. The specific processing is realized by the processor 28 operating as a specific processing unit 290 in accordance with the specific processing program 56 executed on the RAM 30.
[0590] The storage 32 stores the data generation model 58 and the emotion identification model 59. The data generation model 58 and the emotion identification model 59 are used by the identification processing unit 290.
[0591] In robot 414, the processor 46 performs the reception output processing. The storage 50 stores the reception output program 60. The processor 46 reads the reception output program 60 from the storage 50 and executes the read reception output program 60 on the RAM 48. The reception output processing is realized by the processor 46 operating as a control unit 46A according to the reception output program 60 executed on the RAM 48.
[0592] Next, the specific processing performed by the specific processing unit 290 of the data processing device 12 will be described. In the following description, the data processing device 12 will be referred to as the "server" and the robot 414 as the "terminal".
[0593] This invention is a system in which a terminal, a server, and a user work together. The terminal, with the user's consent, collects voice data from voice calls and voice messages in real time. The collected voice data is temporarily stored in a buffer, and the packetized voice data is sent to the server at regular intervals.
[0594] The server processes the received audio data in real time and inputs it into a machine learning model. This model analyzes the characteristics of the audio and scores the likelihood that it is synthesized speech. If the score exceeds a set threshold, the server determines that the audio is processed speech.
[0595] Once a determination is made, the server generates a warning signal and sends it to the terminal. The terminal receives this signal and displays a warning message to the user visually or audibly. The user reviews the warning and decides whether to continue or end the call based on that information.
[0596] For example, if a user receives a phone call that they suspect is a scam, their device sends the call to a server. If the server determines that it is a synthesized voice, a warning is displayed on the device, and the user reviews the warning and makes an appropriate decision. If there is a misidentification, the user can send feedback to the server through their device, and the server can use this feedback to improve its machine learning model and perform more accurate analysis.
[0597] This system allows users to protect themselves from voice scams and use voice services safely.
[0598] The following describes the processing flow.
[0599] Step 1:
[0600] The device monitors audio data in real time while the user is receiving phone calls or voice messages. This data is temporarily stored in a buffer and prepared for processing.
[0601] Step 2:
[0602] The terminal divides the voice data stored in the buffer into packets at regular time intervals and sends them to the server. At this time, voice metadata (such as call ID and timestamp) is also sent.
[0603] Step 3:
[0604] The server sequentially inputs the received audio packets into a machine learning model. The model then begins the process of analyzing the features of the audio data and scoring the likelihood that it is synthesized speech.
[0605] Step 4:
[0606] The server evaluates the score obtained through analysis and determines that the audio is processed if it exceeds a set threshold. The determination result is recorded as a log.
[0607] Step 5:
[0608] If the server detects that the audio has been altered, it generates a warning signal and sends this signal to the terminal. This warning signal includes a message to inform the user.
[0609] Step 6:
[0610] The device displays a visual or audible warning to the user based on the received warning signal. The user reviews the warning and decides whether to end the call.
[0611] Step 7:
[0612] After making a decision based on a warning, users can send feedback from their device to the server to report any misjudgments, if necessary. This feedback is used to improve the machine learning model.
[0613] (Example 1)
[0614] Next, we will describe Example 1. In the following description, the data processing device 12 will be referred to as the "server" and the robot 414 as the "terminal".
[0615] In today's world, fraud and information manipulation using synthesized voices are on the rise, making it crucial to determine the authenticity of voice data in real time. However, current technology makes it difficult to distinguish synthesized voices simply by listening, potentially causing harm to users. Solving this problem and enabling users to communicate with voices with confidence is essential.
[0616] The identification process performed by the identification processing unit 290 of the data processing device 12 in Example 1 is realized by the following means.
[0617] In this invention, the server includes means for transmitting voice information to an analysis device, means for analyzing the received voice information using a machine learning algorithm to determine whether it is processed voice, and means for generating a warning signal based on the determination result and transmitting the warning signal to a device. This prevents fraud using synthesized voices and enables users to use voice information safely in real time.
[0618] "Audio information" refers to data that represents audio in digital format, and this generally includes the content of voice calls and voice messages.
[0619] "Device" refers to an electronic device that collects voice information and transmits it to a server, and includes mobile devices such as smartphones and tablets.
[0620] An "analysis device" refers to a computer system that receives audio information and analyzes its content.
[0621] A "machine learning algorithm" refers to a model used to analyze audio data, understand its characteristics, and determine whether it is synthesized speech.
[0622] A "warning signal" refers to a signal generated to alert the user when there is a possibility that the audio information is synthesized.
[0623] "Response" refers to information that users can use to provide feedback on the synthesized speech's evaluation results and the system's operation.
[0624] This invention relates to a system for determining the authenticity of audio information in real time. The system operates in cooperation with three parties: a terminal used by the user, a server that analyzes the audio data, and the user who receives the warning.
[0625] The device collects voice information from voice calls and voice messages in real time with the user's consent. The device used is a portable electronic device such as a smartphone or tablet. The collected voice information is packetized at regular intervals and sent to a server via a communication protocol. To ensure smooth data transmission, the TCP / IP protocol is commonly used.
[0626] The server applies machine learning algorithms to analyze the received audio information. These algorithms are developed using frameworks such as TensorFlow or PyTorch. The server analyzes the characteristics of the audio information and quantifies the likelihood that it is synthesized speech through scoring. If the score exceeds a certain threshold, it is determined to be synthesized speech and a warning signal is generated. This warning signal is sent to the terminal via the HTTP protocol or similar.
[0627] Based on the received warning signal, the device notifies the user of a warning message visually or audibly. For example, it can display "Warning: This may be a synthesized voice" on the device's screen, or emit a warning sound. This allows the user to make informed decisions to prevent fraud.
[0628] For example, if a user receives a suspicious call, the system analyzes the audio information of the call and makes a judgment. If the server determines that the voice is synthesized speech, a warning is displayed on the terminal, allowing the user to take appropriate action. The user can also provide feedback on the system's judgment, which is used to improve the machine learning algorithm.
[0629] Examples of prompt messages could include, "How should we warn the user if the call is potentially fraudulent?" or "Please tell us how to improve the voice analysis system using machine learning models." By using this system, users can communicate via voice with peace of mind.
[0630] The flow of the specific processing in Example 1 will be explained using Figure 11.
[0631] Step 1:
[0632] With the user's consent, the device collects voice information from voice calls and voice messages in real time. Specifically, the device uses the smartphone's built-in microphone to convert the voice information into a digital format. As a result, the input acquired by the device is analog voice data, and the output is digitized voice information.
[0633] Step 2:
[0634] The terminal temporarily stores the collected audio information in memory. Specifically, it packets the real-time acquired audio data at regular intervals (e.g., every second). The input is the real-time collected audio information, and the output is the data of that audio information divided into smaller packets.
[0635] Step 3:
[0636] The terminal sends packetized voice information to the server. Specifically, this involves sending data to the server's receiving port using the TCP / IP protocol. The input is packetized voice information, and the output is the voice data sent to the server via the network.
[0637] Step 4:
[0638] The server inputs the received audio data into a machine learning algorithm. The server provides the audio data to a model developed using, for example, TensorFlow, and extracts audio features. The input is the audio data received by the server, and the output is the audio data features and analysis results.
[0639] Step 5:
[0640] The server uses the output of a machine learning algorithm to score whether the speech is likely to be synthesized. Here, if the score exceeds a certain threshold, it is classified as "synthesized speech." The input is the analyzed features, and the output is the scoring result and the classification result.
[0641] Step 6:
[0642] The server generates a warning signal if it determines that the voice is synthesized. This signal contains a message such as "Caution: This voice may have been synthesized." In its specific operation, it generates a warning signal and creates a data structure to send it to the terminal. The input is the detection result, and the output is the warning signal.
[0643] Step 7:
[0644] The terminal notifies the user based on the received warning signal. Specifically, the terminal attracts the user's attention by displaying text on the screen or emitting a warning sound. The input is a warning signal sent from the server, and the output is a warning message that the user sees or hears.
[0645] (Application Example 1)
[0646] Next, we will explain Application Example 1. In the following explanation, the data processing device 12 will be referred to as the "server" and the robot 414 as the "terminal".
[0647] In recent years, advancements in voice technology have led to an increase in fraud and illegal activities using synthesized voices. This increases the likelihood of users finding themselves in dangerous situations. However, technology to accurately determine whether voice information is synthesized is still insufficient, and users lack effective countermeasures against fraudulent voices. Against this backdrop, the present invention aims to provide a system that can quickly and accurately determine whether voice information is synthesized and warn users.
[0648] The specific processing performed by the specific processing unit 290 of the data processing device 12 in Application Example 1 is realized by the following means.
[0649] In this invention, the server includes information processing means for analyzing voice information using a learning algorithm and determining whether it is modified voice; signal generation means for generating a warning signal based on the determination result and transmitting it to an information terminal; and means for receiving improvement reports from users who have received warnings and applying them to the learning algorithm. This enables users to quickly detect fraudulent synthesized voices and take appropriate action.
[0650] "Voice information" refers to information data recorded as sound waves, including voice calls and voice messages.
[0651] An "information terminal" is a device that has the function of collecting voice information and transmitting it to a data processing device.
[0652] A "data processing device" is an electronic device with computing power for analyzing received audio information.
[0653] A "learning algorithm" is an algorithm that uses machine learning to analyze the characteristics of audio information and determine whether it is likely to be synthesized speech.
[0654] "Information processing means" refers to a function that performs a process of analyzing received audio information and determining its validity.
[0655] A "signal generation means" is a means that has the function of generating a warning signal when it is determined that audio information has been altered.
[0656] A "notification means" is a function that transmits a warning to the user visually or audibly via an information terminal that receives a warning signal.
[0657] A "user" is a person who receives and manipulates voice information, or an entity intended to use such information.
[0658] An "improvement report" is user feedback based on warning messages, and it provides information to improve the accuracy of the learning algorithm.
[0659] This invention relates to a system for detecting fraudulent voice information, and consists of three components: an information terminal, a data processing device, and a user. The information terminal is comprised of a device such as a smartphone or wearable device, through which voice information is collected. The voice information is temporarily stored in a buffer in real time and periodically transmitted to the data processing device as data packets.
[0660] The data processing unit functions as a server and, upon receiving audio information, uses a pre-trained learning algorithm to analyze the characteristics of that audio information in detail. The analysis quantifies the likelihood that the audio is synthesized speech, and if it exceeds a certain threshold, it is determined to be modified speech. At this point, the server generates a warning signal, which is transmitted to the information terminal. The information terminal receives this warning signal and provides the user with a visual or auditory warning.
[0661] Users are alerted to warnings and encouraged to be vigilant against fraud and misconduct. If a misidentification occurs during actual use, users can send feedback from their information terminal to the server. This feedback is received by the server and used to improve the learning algorithm, enabling more accurate analysis.
[0662] As a concrete example, when a user receives a phone call, the system analyzes the call and, if it detects a synthesized voice suspected of being fraudulent, the user receives an instant warning, thus preventing fraud. An example of a prompt message for this voice analysis would be, "Please determine whether this voice is genuine or synthesized. If the score exceeds the threshold, generate a warning message."
[0663] The flow of a specific process in Application Example 1 will be explained using Figure 12.
[0664] Step 1:
[0665] The device collects voice information in real time. When a voice call or voice message begins, the voice signal is stored in a buffer as digital data. At this stage, the input is voice information, and the output is digital voice data. The collected voice data is packetized at regular intervals.
[0666] Step 2:
[0667] The terminal sends packetized voice data to the server. The data is securely transferred over the network and received by the server. The input is packetized voice data, and the output is the data received on the server side. The transfer is secured using encrypted communication.
[0668] Step 3:
[0669] The server inputs the received audio data into a machine learning model for analysis. It extracts audio characteristics as features and scores the likelihood that the audio is synthesized speech. The input is the received audio data, and the output is the synthesized speech score. This analysis is performed using a sophisticated algorithm based on a generative AI model.
[0670] Step 4:
[0671] The server generates a warning signal if the score exceeds a set threshold. This signal indicates a high probability of synthesized speech. The input is the analysis score, and the output is the warning signal. Signal generation is performed in real time, requiring immediate response.
[0672] Step 5:
[0673] The terminal receives a warning signal sent from the server and notifies the user. The warning is presented visually as an on-screen alert and audibly as an audio notification. The input is the warning signal, and the output is a warning display to the user. The most suitable notification method is selected depending on the situation.
[0674] Step 6:
[0675] The user receives a warning and determines the veracity of the audio information. They review the warning and, if necessary, interrupt the call or send feedback. The input is the warning to the user, and the output is the user's response. Feedback helps improve the system in case of misjudgments.
[0676] Step 7:
[0677] The server receives feedback from users and uses it to improve the machine learning model. The model is retrained with new data to achieve more accurate analysis. The input is user feedback, and the output is the improved learning algorithm. This process is continuous, improving the model's performance.
[0678] Furthermore, an emotion engine that estimates the user's emotions may be incorporated. That is, the identification processing unit 290 may use the emotion identification model 59 to estimate the user's emotions and perform identification processing using the user's emotions.
[0679] The present invention is a system that combines a terminal for collecting and analyzing voice data, a server equipped with means for determining whether the voice data is processed voice, and an emotion engine that recognizes the user's emotions and appropriately adjusts warning messages.
[0680] The terminal monitors voice calls and messages received by the user in real time and temporarily stores the voice data in a buffer. This data is packetized at regular intervals and sent to a server for analysis.
[0681] The server feeds the received audio data into a machine learning model. This model analyzes the characteristics of the audio and scores the likelihood that it is synthesized speech. If the score exceeds a threshold, the server determines that the audio is processed speech and sends a warning signal to the terminal.
[0682] The emotion engine analyzes the user's voice tone, speaking speed, facial expressions, etc., to recognize the user's emotional state in real time. Based on the recognized emotional information, it is possible to adjust how warning messages are presented (for example, the emphasis and type of the message).
[0683] For example, during a normal phone call, the server analyzes the received audio and determines that it may be synthesized speech. If the emotion engine then recognizes that the user's voice sounds tense, it will alert the user by displaying a more emphasized warning message on the device.
[0684] This system not only combats the risk of voice fraud but also enables flexible responses that respond to the user's emotions, providing more personalized protection.
[0685] The following describes the processing flow.
[0686] Step 1:
[0687] The device monitors the audio data in real time when the user is receiving a phone call or voice message, and temporarily stores it in a buffer. This audio data is then packetized in preparation for subsequent processing.
[0688] Step 2:
[0689] The terminal sends the voice data stored in the buffer to the server at regular time intervals. At this time, the call ID and timestamp are also sent as metadata.
[0690] Step 3:
[0691] The server inputs the received audio data into a machine learning model to analyze the audio's characteristics. The model scores the likelihood that the audio is synthesized speech, and if it exceeds a threshold, it is classified as processed speech.
[0692] Step 4:
[0693] If the server determines that the audio is processed, it generates a warning signal and sends it to the terminal. The warning signal includes instructions to alert the user.
[0694] Step 5:
[0695] The device activates the emotion engine based on the received warning signal. The emotion engine analyzes the user's voice tone, speaking speed, and facial expressions to recognize the user's emotional state.
[0696] Step 6:
[0697] Based on the emotional information recognized by the emotion engine, the device adjusts how warning messages are presented. For example, if the user is showing signs of anxiety, the warning message is visually highlighted.
[0698] Step 7:
[0699] The user reviews the warning message from their device and decides whether to continue or end the call. If necessary, they can send feedback from their device to the server if there is a misjudgment. This feedback is used by the server to improve the machine learning model.
[0700] (Example 2)
[0701] Next, we will describe Example 2. In the following description, the data processing device 12 will be referred to as the "server" and the robot 414 as the "terminal".
[0702] In modern society, voice-based fraud is on the rise, but effective detection systems to combat it are still not well-established. Furthermore, existing systems do not take into account the user's emotional state, and inappropriate warnings can lead to unnecessary stress and confusion. Therefore, there is a need for technology that can quickly and accurately identify manipulated voices and provide warnings tailored to the user's emotional state.
[0703] The identification process performed by the identification processing unit 290 of the data processing device 12 in Example 2 is realized by the following means.
[0704] In this invention, the server includes means for analyzing voice information using a computational model and determining whether the voice has been processed; means for generating and transmitting a warning signal based on the determination result; and means for evaluating the user's emotional state using an emotion analysis device and adjusting the method of presenting the warning based on the emotional state. This enables rapid and accurate detection of processed voices and personalized warning presentation that takes the user's emotions into consideration.
[0705] "Audio information" refers to the digital or analog representation of sound waves that have been transmitted or stored.
[0706] An "information processing device" is a combination of hardware and software used to compute, analyze, or transform data.
[0707] A "computational model" is an algorithm or mathematical formula used to analyze the characteristics of data and make specific decisions based on the results.
[0708] A "warning signal" is a signal generated to inform the user or system of a specific condition or abnormality.
[0709] An "emotion analysis device" is a device or system that evaluates or estimates a user's emotional state based on information such as their voice and facial expressions.
[0710] "User" refers to the individual who operates or uses the system or product.
[0711] "Numerical evaluation" refers to numerical results calculated to quantify specific characteristics or states.
[0712] "Opinions" refers to feedback and reactions provided by users.
[0713] This invention is a system for monitoring voice calls and messages received by a user and analyzing the voice information. Its primary purpose is to collect and analyze voice information and provide warnings that take into account the user's emotional state. The following describes a specific implementation of this system.
[0714] The terminal acquires the user's voice calls and messages in real time. To do this, it captures voice information using a voice input device and digital signal processing technology. The collected voice information is temporarily stored in the terminal's memory and then transmitted to a server using network communication technology.
[0715] The server is equipped with a computing device that analyzes transmitted audio information using a computational model. This analysis utilizes machine learning libraries (e.g., TensorFlow and PyTorch) to analyze the characteristics of the audio and determine whether it is synthesized speech. If it is determined to be synthesized speech, a warning signal is generated and sent to the terminal.
[0716] The emotion analysis device analyzes emotional data from voice information and the user's facial expressions to understand the user's emotional state. This information is transmitted to a server and used to display warning signals. For example, if the device detects a state of tension, it will display a more emphasized warning message to prompt appropriate action.
[0717] As a concrete example, suppose a user makes a voice call, and the server analyzes the voice information and determines that it is synthesized speech. In this situation, the emotion analysis device can sense the tension in the user's voice and issue a command to display an emphasized warning message on the device. This allows the user to immediately respond to the risk of voice fraud.
[0718] An example of a prompt to input into a generative AI model is, "Analyze the user's emotion from their voice tone and facial expressions, and create a message appropriate to the situation." This prompt helps improve the accuracy of emotion analysis and make warning messages more appropriate.
[0719] The flow of the specific processing in Example 2 will be explained using Figure 13.
[0720] Step 1:
[0721] The terminal monitors incoming voice calls and messages in real time. When voice input occurs, the terminal uses the microphone to capture sound wave data and converts it into a digital format. This data is temporarily stored in memory and packetized. The input is voice data, and the output is packetized digital data.
[0722] Step 2:
[0723] The terminal sends packetized voice data to the server at regular intervals. Efficient data transfer using network protocols is required to minimize latency. The input is packetized digital data, and the output is voice data received by the server.
[0724] Step 3:
[0725] The server acquires the received audio data and feeds it into a machine learning model. The computational model used here is, for example, trained using TensorFlow or PyTorch. It analyzes the characteristics of the audio, such as frequency components and waveform features, and scores the likelihood of it being synthesized speech. The input is the audio data sent to the server, and the output is the scored likelihood of it being synthesized speech.
[0726] Step 4:
[0727] The server compares the scoring result to a threshold. If the threshold is exceeded, the server determines that the voice is synthesized speech and generates a warning signal. This warning signal is then prepared to be sent to the terminal. The input is the scoring result, and the output is the warning signal.
[0728] Step 5:
[0729] The emotion analysis device acquires the user's voice tone and facial expression data to evaluate the user's emotional state. Data is acquired from audio or video cameras and analyzed in real time. Emotional states are classified, for example, into categories such as tension or calmness. The input is the user's voice and facial expression data, and the output is the evaluated emotional state.
[0730] Step 6:
[0731] The terminal references the warning signal received from the server and the emotional state from the emotion analyzer, and adjusts how the warning message is displayed. For example, if it determines that the user is in a state of tension, the terminal will highlight the warning message, providing an appropriate warning to the user. The input is the warning signal and emotional state, and the output is the adjusted warning message.
[0732] (Application Example 2)
[0733] Next, we will explain application example 2. In the following explanation, the data processing device 12 will be referred to as the "server" and the robot 414 as the "terminal".
[0734] In recent years, advancements in speech synthesis technology have increased the risk of general users being affected by voice fraud and disguised voices. However, currently, there are insufficient means to determine in real time whether a received voice is synthesized or not. Furthermore, there is a need for a system that can appropriately adjust warning messages by taking into account the user's emotional state, thereby preventing excessive warnings and providing effective alerts.
[0735] The specific processing performed by the specific processing unit 290 of the data processing device 12 in Application Example 2 is realized by the following means.
[0736] In this invention, the server includes means for a terminal that collects audio data and transmits it to an information processing device for analysis; means for analyzing the received audio data using a machine learning model and determining whether it is synthesized audio; and emotion analysis means for recognizing the user's emotional state and adjusting the warning message based on that emotional state. This makes it possible to determine whether the audio data is synthesized or natural and to provide an optimal warning based on the user's emotional state.
[0737] "Audio data" refers to information recorded in digital format, which is used for analysis and transmission.
[0738] An "information processing device" is an electronic device used to perform data analysis and calculations, and includes servers and computers.
[0739] A "machine learning model" is a collection of algorithms and computational methods used for analysis and prediction, which learn features from large amounts of data.
[0740] "Emotional state" refers to the user's emotional response and is determined by indicators such as tone of voice, speaking speed, and facial expressions.
[0741] A "warning signal" is a signal used to alert the user when certain conditions are met, and is communicated through audio, on-screen displays, or other means.
[0742] A "terminal" is an input / output device used by a user, and includes smartphones and tablets.
[0743] The system that implements this application involves real-time analysis of audio data and adjustment of warning messages based on the user's emotional state. The server receives audio data and performs analysis using a machine learning model. Specifically, it extracts features from the audio data using digital signal processing technology, applies these features to a machine learning algorithm, and scores the likelihood that the audio is synthesized speech.
[0744] The terminal receives a warning signal sent from the server and uses an emotion analysis engine to analyze the user's voice tone and speaking speed. This allows the terminal to recognize the user's emotional state and adjust the warning message accordingly. For example, if the user is feeling anxious, the warning message can be displayed with greater emphasis.
[0745] The program is implemented using programming languages such as Python, and utilizes the speech_recognition library for speech analysis. Data communication is performed via the HTTP protocol. The user interface is designed to display both audio and visual warnings.
[0746] For example, if a phone call received by a user is suspected to be using synthesized speech, the terminal receives a corresponding warning signal. At this time, the emotion engine analyzes the user's emotional state, and only if it recognizes anxiety or tension, it highlights the warning message to prompt the user to take immediate action. The prompt text to be input to the generating AI model is as follows:
[0747] "Please consider specific implementation methods for a function that acquires audio data in real time and analyzes the potential of synthesized speech."
[0748] The flow of a specific process in Application Example 2 will be explained using Figure 14.
[0749] Step 1:
[0750] The terminal temporarily stores the voice call data received from the user in a buffer. It receives real-time captured voice data as input, converts it to a digital data format, and saves it in preparation for subsequent processing.
[0751] Step 2:
[0752] The terminal packets the voice data stored in a buffer at regular intervals and sends it to the server. The input is voice call data, which is then divided into small data packets and sent to the server via the network—a data processing step.
[0753] Step 3:
[0754] The server interprets the received audio data and feeds it into a machine learning model. The input here is data packets sent from the terminal, and the server performs data calculations by extracting this audio data and analyzing it using a machine learning model.
[0755] Step 4:
[0756] The server uses a machine learning model to analyze the audio data and score its likelihood of being synthesized speech. Based on the analysis, it determines whether the audio is synthesized speech, and if the score exceeds a threshold, it generates a warning signal. The output of this step is the scoring based on the analyzed data.
[0757] Step 5:
[0758] The server sends the generated warning signal to the terminal. It takes a scoring-based judgment result as input and performs data communication to prompt appropriate action by sending that result to the user's terminal.
[0759] Step 6:
[0760] The terminal analyzes the user's emotional state using an emotion analysis engine based on the received warning signal. The input is the warning signal received from the server, and in addition, the user's voice tone and speaking speed are analyzed to process the data in order to identify the user's emotions.
[0761] Step 7:
[0762] The device adjusts and displays warning messages on the screen based on the user's emotional state. The input is the result of the analysis of the user's emotional state, and the output is a warning message that matches the result. Specifically, if the user is feeling anxious, the message will be highlighted to draw their attention.
[0763] The specific processing unit 290 transmits the result of the specific processing to the robot 414. In the robot 414, the control unit 46A causes the speaker 240 and the controlled object 443 to output the result of the specific processing. The microphone 238 acquires audio indicating user input for the result of the specific processing. The control unit 46A transmits the audio data indicating user input acquired by the microphone 238 to the data processing unit 12. In the data processing unit 12, the specific processing unit 290 acquires the audio data.
[0764] Data generation model 58 is a type of so-called generative AI (Artificial Intelligence). One example of data generation model 58 is ChatGPT (Internet search<URL: https: / / openai.com / blog / chatgpt> ), Gemini (Internet search) <url: https: gemini.google.com ?hl="ja">Examples of generative AI include the following. The data generation model 58 is obtained by performing deep learning on a neural network. The data generation model 58 is input with prompts containing instructions, and with inference data such as audio data representing speech, text data representing text, and image data representing images. The data generation model 58 infers from the input inference data according to the instructions indicated by the prompts, and outputs the inference results in data formats such as audio data and text data. Here, inference refers to, for example, analysis, classification, prediction, and / or summarization.
[0765] In the above embodiment, an example was given in which the specific processing is performed by the data processing device 12, but the technology of this disclosure is not limited thereto, and the specific processing may also be performed by the robot 414.
[0766] Furthermore, the emotion identification model 59, acting as an emotion engine, may determine the user's emotion according to a specific mapping. Specifically, the emotion identification model 59 may determine the user's emotion according to a specific mapping, which is an emotion map (see Figure 9). Similarly, the emotion identification model 59 may also determine the robot's emotion, and the identification processing unit 290 may perform identification processing using the robot's emotion.
[0767] Figure 9 shows an emotion map 400 in which multiple emotions are mapped. In the emotion map 400, emotions are arranged in concentric circles radiating from the center. The closer to the center of the concentric circles, the more primitive the emotions are located. Further out of the concentric circles, emotions representing states and actions arising from mental states are located. Emotion is a concept that includes feelings and mental states. On the left side of the concentric circles, emotions that are generally generated from reactions occurring in the brain are located. On the right side of the concentric circles, emotions that are generally induced by situational judgment are located. Above and below the concentric circles, emotions that are generally generated from reactions occurring in the brain and induced by situational judgment are located. In addition, the emotion of "pleasure" is located on the upper side of the concentric circles, and the emotion of "displeasure" is located on the lower side. Thus, in the emotion map 400, multiple emotions are mapped based on the structure in which emotions arise, and emotions that are likely to occur simultaneously are mapped close together.
[0768] These emotions are distributed at the 3 o'clock position on the Emotion Map 400, and usually fluctuate between feelings of security and anxiety. In the right half of the Emotion Map 400, situational awareness takes precedence over internal feelings, resulting in a calm impression.
[0769] The inside of the Emotion Map 400 represents inner thoughts, while the outside represents actions. Therefore, the further you go from the outside of the Emotion Map 400, the more visible (expressed in actions) your emotions become.
[0770] Here, human emotions are based on various balances, such as posture and blood sugar levels. When these balances deviate from the ideal, it results in discomfort, and when they approach the ideal, it results in pleasure. Similarly, in robots, cars, motorcycles, etc., emotions can be created based on various balances, such as posture and battery level. When these balances deviate from the ideal, it results in discomfort, and when they approach the ideal, it results in pleasure. The emotion map can be generated, for example, based on Dr. Mitsuyoshi's emotion map (Research on a system for analyzing brain physiological signals of speech emotion recognition and emotion, Tokushima University, doctoral dissertation: https: / / ci.nii.ac.jp / naid / 500000375379). The left half of the emotion map contains emotions belonging to a region called "response," where sensation is dominant. The right half of the emotion map contains emotions belonging to a region called "situation," where situational awareness is dominant.
[0771] The emotion map defines two emotions that promote learning. One is the emotion around the middle of the negative "repentance" and "reflection" on the situation side. In other words, it is when the robot experiences negative emotions such as "I never want to feel this way again" or "I don't want to be scolded again." The other is the emotion around the positive "desire" on the reaction side. In other words, it is when the robot has positive feelings such as "I want more" or "I want to know more."
[0772] The emotion identification model 59 inputs user input into a pre-trained neural network, obtains emotion values representing each emotion shown in the emotion map 400, and determines the user's emotion. This neural network is pre-trained based on multiple training data sets, which are combinations of user input and emotion values representing each emotion shown in the emotion map 400. Furthermore, this neural network is trained so that emotions located close together have similar values, as shown in the emotion map 900 in Figure 10. Figure 10 shows an example where multiple emotions such as "reassured," "calm," and "confident" have similar emotion values.
[0773] The above description primarily focuses on the functions of the data processing device 12 in relation to this disclosure. However, the system related to this disclosure is not necessarily implemented on a server. The system related to this disclosure may be implemented as a general information processing system. This disclosure may be implemented, for example, as a software program that runs on a personal computer or as an application that runs on a smartphone. The method related to this disclosure may be provided to users in SaaS (Software as a Service) format.
[0774] In the above embodiment, an example was given in which a specific process is performed by a single computer 22. However, the technology of this disclosure is not limited thereto, and a distributed processing of the specific process may be performed by multiple computers, including computer 22. For example, a data generation model 58 may be provided in an external device of the data processing device 12, and the external device may generate data according to the input data.
[0775] In the above embodiment, an example was given in which the specific processing program 56 is stored in the storage 32, but the technology of this disclosure is not limited thereto. For example, the specific processing program 56 may be stored in a portable, computer-readable, non-temporary storage medium such as a USB (Universal Serial Bus) memory. The specific processing program 56 stored in the non-temporary storage medium is installed in the computer 22 of the data processing device 12. The processor 28 executes specific processing according to the specific processing program 56.
[0776] Alternatively, the specific processing program 56 may be stored in a storage device such as a server connected to the data processing device 12 via the network 54, and the specific processing program 56 may be downloaded and installed on the computer 22 in response to a request from the data processing device 12.
[0777] Furthermore, it is not necessary to store the entirety of the specific processing program 56 in a storage device such as a server connected to the data processing device 12 via the network 54, or to store the entirety of the specific processing program 56 in the storage 32; it is acceptable to store only a portion of the specific processing program 56.
[0778] The following types of processors can be used as hardware resources to perform specific processing. Examples of processors include a CPU, a general-purpose processor that functions as a hardware resource to perform specific processing by executing software, i.e., a program. Other examples of processors include dedicated electrical circuits, such as FPGAs (Field-Programmable Gate Arrays), PLDs (Programmable Logic Devices), or ASICs (Application Specific Integrated Circuits), which have circuit configurations specifically designed to perform specific processing. All of these processors have built-in or connected memory, and all of them perform specific processing by using memory.
[0779] The hardware resource that performs a specific process may consist of one of these various processors, or it may consist of a combination of two or more processors of the same or different types (for example, a combination of multiple FPGAs, or a combination of a CPU and an FPGA). Alternatively, the hardware resource that performs a specific process may consist of a single processor.
[0780] Examples of configurations using a single processor include, firstly, a configuration in which one or more CPUs and software are combined to form a single processor, and this processor functions as a hardware resource that performs a specific process. Secondly, there is a configuration using a processor that realizes the functions of the entire system, including multiple hardware resources that perform a specific process, on a single IC chip, as exemplified by SoCs (System-on-a-chip). In this way, a specific process is realized using one or more of the above types of processors as hardware resources.
[0781] Furthermore, the hardware structure of these various processors can more specifically utilize electrical circuits that combine circuit elements such as semiconductor devices. Also, the specific processing described above is merely an example. Therefore, it goes without saying that unnecessary steps can be deleted, new steps added, or the processing order rearranged, as long as it does not deviate from the main purpose.
[0782] The descriptions and illustrations presented above are detailed explanations of the technical aspects of this disclosure and are merely examples of the technical aspects. For example, the above descriptions of the structure, function, operation, and effect are examples of the structure, function, operation, and effect of the technical aspects of this disclosure. Therefore, it goes without saying that you may delete unnecessary parts, add new elements, or replace elements in the descriptions and illustrations presented above, as long as you do not deviate from the essence of the technical aspects of this disclosure. Furthermore, in order to avoid confusion and facilitate understanding of the technical aspects of this disclosure, explanations of common technical knowledge and the like that do not require special explanation to enable the implementation of the technical aspects of this disclosure have been omitted from the descriptions and illustrations presented above.
[0783] All documents, patent applications, and technical standards described herein are incorporated by reference to the same extent as if each individual document, patent application, and technical standard were specifically and individually noted to be incorporated by reference.
[0784] The following is further disclosed regarding the embodiments described above.
[0785] (Claim 1)
[0786] A terminal that collects audio data and transmits the audio data to a processing device for analysis,
[0787] The processing apparatus includes means for analyzing received audio data using a machine learning model and determining whether the audio has been processed,
[0788] A means for generating a warning signal if the audio is processed based on the judgment result, and transmitting the warning signal to the terminal,
[0789] Means for displaying a warning on the terminal in response to the warning signal,
[0790] A system that includes this.
[0791] (Claim 2)
[0792] The system according to claim 1, characterized in that, in the analysis of speech data, the machine learning model analyzes the characteristics of the speech and uses means to score the likelihood that it is synthesized speech.
[0793] (Claim 3)
[0794] The system according to claim 1, characterized in that it includes means for receiving feedback from a user to report a misjudgment after they have taken action in response to a warning message, and for using said feedback to improve the machine learning model.
[0795] "Example 1"
[0796] (Claim 1)
[0797] A device that collects audio information and transmits the audio information to an analysis device,
[0798] The analysis device includes means for analyzing received audio information using a machine learning algorithm and determining whether the audio has been processed,
[0799] A means for generating a warning signal if the audio is processed based on the judgment result, and transmitting the warning signal to the device,
[0800] Means for displaying a warning on the device in response to the warning signal,
[0801] A system that includes this.
[0802] (Claim 2)
[0803] The system according to claim 1, characterized in that, in the analysis of speech information, the machine learning algorithm analyzes the characteristics of the speech and uses means to quantify the possibility that it is synthesized speech.
[0804] (Claim 3)
[0805] The system according to claim 1, characterized in that it includes means for receiving a response from a user to report a misjudgment after the user has taken action in response to a warning notification, and for improving the machine learning algorithm using said response.
[0806] "Application Example 1"
[0807] (Claim 1)
[0808] An information terminal that collects voice information and transmits said voice information to a data processing device for analysis,
[0809] The data processing device includes an information processing means that analyzes the received audio information using a learning algorithm and determines whether the audio has been altered,
[0810] A signal generation means that generates a warning signal if the audio is modified based on the judgment result, and transmits the warning signal to the information terminal,
[0811] A notification means that displays a warning on the information terminal in response to the warning signal,
[0812] A means for receiving a report when a user who has received a warning through the notification means makes a false report, and for using the report to improve the learning algorithm,
[0813] An information processing system that includes this.
[0814] (Claim 2)
[0815] The information processing system according to claim 1, characterized in that, in the analysis of speech information, the learning algorithm analyzes the characteristics of the speech and uses means to quantify the possibility that it is generative speech.
[0816] (Claim 3)
[0817] The information processing system according to claim 1, characterized in that it includes means for receiving improvement reports after a user takes action in response to a warning message, and for using said improvement reports to improve the accuracy of the learning algorithm.
[0818] "Example 2 of combining an emotion engine"
[0819] (Claim 1)
[0820] An information terminal that acquires audio information and transmits said audio information to an information processing device for analysis,
[0821] The information processing device includes means for analyzing received audio information using a computational model and determining whether it is processed audio,
[0822] A means for generating a warning signal if the audio is processed based on the judgment result, and transmitting the warning signal to the information terminal,
[0823] A means for evaluating the user's emotional state using an emotion analysis device and adjusting the method of presenting warnings based on that emotional state,
[0824] Means for displaying a warning on the information terminal in response to the warning signal,
[0825] A system that includes this.
[0826] (Claim 2)
[0827] The system according to claim 1, characterized in that, in the analysis of speech information, the computational model analyzes the characteristics of the speech and uses means to numerically evaluate the possibility that it is synthesized speech.
[0828] (Claim 3)
[0829] The system according to claim 1, characterized in that it includes means for receiving feedback from users to report misjudgments after they have taken action in response to a warning, and for improving the calculation model using said feedback.
[0830] "Application example 2 when combining with an emotional engine"
[0831] (Claim 1)
[0832] A terminal that collects audio data and transmits the audio data to an information processing device for analysis,
[0833] The information processing device includes means for analyzing received audio data using a machine learning model and determining whether the audio has been processed,
[0834] A means for generating a warning signal if the audio is processed based on the judgment result, and for transmitting the warning signal to the terminal,
[0835] A sentiment analysis means that recognizes the user's emotional state and adjusts the warning message based on that emotional state,
[0836] Means for displaying a warning on the terminal in response to the warning signal,
[0837] ...
[0838] A system that includes this.
[0839] (Claim 2)
[0840] The system according to claim 1, characterized in that, in the analysis of speech data, the machine learning model analyzes the characteristics of the speech and uses means to score the likelihood that it is synthesized speech.
[0841] (Claim 3)
[0842] The system according to claim 1, characterized in that it includes means for receiving feedback from a user to report a misjudgment after they have taken action in response to a warning message, and for using said feedback to improve the machine learning model. [Explanation of symbols]
[0843] 10, 210, 310, 410 Data Processing Systems 12 Data Processing Devices 14 Smart Devices 214 Smart Glasses 314 Headset-type terminal 414 Robots< / url:> < / url:> < / url:> < / url:>
Claims
1. An information terminal that collects voice information and transmits said voice information to a data processing device for analysis, The data processing device includes an information processing means that analyzes the received audio information using a learning algorithm and determines whether the audio has been altered, A signal generation means that generates a warning signal if the audio is modified based on the judgment result, and transmits the warning signal to the information terminal, A notification means that displays a warning on the information terminal in response to the warning signal, A means for receiving a report when a user who has received a warning through the notification means makes a false report, and for using the report to improve the learning algorithm, An information processing system that includes this.
2. The information processing system according to claim 1, characterized in that, in the analysis of speech information, the learning algorithm analyzes the characteristics of the speech and uses means to quantify the possibility that it is generative speech.
3. The information processing system according to claim 1, characterized in that it includes means for receiving improvement reports after a user takes action in response to a warning message, and for using said improvement reports to improve the accuracy of the learning algorithm.
Citation Information
Patent Citations
Persona chatbot control method and system
JP2022180282A