System
A voice-based system quickly and reliably verifies user identity to release credit card security locks, addressing the inefficiencies of manual verification methods by using voice authentication and interface controls for seamless operation.
Patent Information
- Application Number
- JP2024116320
- Authority / Receiving Office
- JP · JP
- Patent Type
- Applications
- Current Assignee / Owner
- Filing Date
- 2024-07-19
- Publication Date
- 2026-01-29
AI Technical Summary
Conventional methods for releasing a credit card security lock due to fraudulent activity detection require time-consuming and laborious manual identity verification processes.
A system that records a user's voice, transmits the data to a server for analysis using a voice authentication model, evaluates the authentication score, and releases the security lock based on the score, with user interface controls for recording and notification of results.
Enables quick and reliable verification of user identity to release the security lock, improving user convenience and efficiency in unlocking credit card usage.
Smart Images

Figure 2026014846000001_ABST
Abstract
Description
[Technical Field]
[0001] The technology of the present disclosure relates to a system. [Background technology]
[0002] Patent document 1 discloses a persona chatbot control method performed by at least one processor, the method including the steps of receiving a user utterance, adding the user utterance to a prompt including an instruction sentence related to a description of the chatbot character, encoding the prompt, and inputting the encoded prompt into a language model to generate a chatbot utterance in response to the user utterance. [Prior art documents] [Patent documents]
[0003] [Patent Document 1] Japanese Patent Publication No. 2022-180282 Summary of the Invention [Problem to be solved by the invention]
[0004] Even when a credit card is used by the cardholder when making a payment, the card may be suspended due to the card company's logic to prevent fraudulent activity. In such a situation, conventional methods require the user to call the card company and go through a complicated identity verification process, which is time-consuming and laborious. The present invention aims to solve these problems by providing a system that quickly and reliably verifies the user's identity and releases the security lock. [Means for solving the problem]
[0005] The present invention solves the above problems by providing a system including the following means: means for recording a user's voice, means for transmitting the recorded voice data to a server, means for analyzing the received voice data with a voice authentication model, means for evaluating the authentication score of the analyzed voice data, and means for releasing a security lock based on the evaluated authentication score. Furthermore, by including means for providing an interface for the user to start and stop recording, and means for notifying the user of the results of voice authentication, these problems can be effectively solved.
[0006] An "audio recording device" is any device or software that collects and stores a user's voice in digital form.
[0007] The "means for transmitting recorded voice data to the server" refers to a communication protocol or software for transferring voice data from the terminal to the server.
[0008] The "analysis means" is hardware or software that includes a voice authentication model for analyzing received voice data and generating an authentication score.
[0009] The "means for evaluating the authentication score" is software that includes algorithms and logic for authenticating a user based on the authentication score generated by the analysis means.
[0010] "Means to release security lock" refers to the process or system for releasing credit card usage restrictions based on the evaluation results of the authentication score.
[0011] The "interface for the user to operate to start and stop recording" refers to a UI (user interface) or button that allows the user to start and stop audio recording.
[0012] The "notification means" is a system including a message transmission function and a display device for notifying the user of the results of voice authentication. [Brief explanation of the drawings]
[0013] [Figure 1] 1 is a conceptual diagram showing an example of the configuration of a data processing system according to a first embodiment. [Figure 2] 1 is a conceptual diagram showing an example of main functions of a data processing device and a smart device according to a first embodiment. [Figure 3] FIG. 10 is a conceptual diagram showing an example of the configuration of a data processing system according to a second embodiment. [Figure 4] FIG. 10 is a conceptual diagram showing an example of main functions of a data processing device and smart glasses according to a second embodiment. [Figure 5] FIG. 10 is a conceptual diagram showing an example of the configuration of a data processing system according to a third embodiment. [Figure 6] FIG. 11 is a conceptual diagram showing an example of main functions of a data processing device and a headset-type terminal according to a third embodiment. [Figure 7] FIG. 10 is a conceptual diagram showing an example of the configuration of a data processing system according to a fourth embodiment. [Figure 8] FIG. 10 is a conceptual diagram showing an example of main functions of a data processing device and a robot according to a fourth embodiment. [Figure 9] 1 shows an emotion map onto which multiple emotions are mapped. [Figure 10] 1 shows an emotion map onto which multiple emotions are mapped. [Figure 11] FIG. 3 is a sequence diagram showing a processing flow of the data processing system according to the first embodiment. [Figure 12] FIG. 10 is a sequence diagram showing the flow of processing in the data processing system in Application Example 1. [Figure 13] FIG. 10 is a sequence diagram showing the flow of processing in the data processing system according to the second embodiment when an emotion engine is combined. [Figure 14] FIG. 10 is a sequence diagram showing the flow of processing in the data processing system in Application Example 2 when an emotion engine is combined. DETAILED DESCRIPTION OF THE INVENTION
[0014] An example of an embodiment of a system according to the technology of the present disclosure will be described below with reference to the accompanying drawings.
[0015] First, the terms used in the following description will be explained.
[0016] In the following embodiments, a coded processor (hereinafter simply referred to as a "processor") may be a single arithmetic device or a combination of multiple arithmetic devices. Furthermore, a processor may be a single type of arithmetic device or a combination of multiple types of arithmetic devices. Examples of arithmetic devices include a CPU (Central Processing Unit), a GPU (Graphics Processing Unit), a GPGPU (General-Purpose computing on Graphics Processing Units), and an APU (Accelerated Processing Unit).
[0017] In the following embodiments, a coded RAM (Random Access Memory) is a memory in which information is temporarily stored and is used as a working memory by a processor.
[0018] In the following embodiments, the coded storage is one or more non-volatile storage devices that store various programs, various parameters, etc. Examples of non-volatile storage devices include flash memory (SSD (Solid State Drive)), magnetic disks (e.g., hard disks), and magnetic tapes.
[0019] In the following embodiments, a communication I / F (Interface) with a symbol is an interface including a communication processor, an antenna, etc. The communication I / F controls communication between multiple computers. Examples of communication standards applied to the communication I / F include wireless communication standards including 5G (5th Generation Mobile Communication System), Wi-Fi (registered trademark), Bluetooth (registered trademark), etc.
[0020] In the following embodiments, "A and / or B" is synonymous with "at least one of A and B." In other words, "A and / or B" means that it may be only A, only B, or a combination of A and B. Furthermore, in this specification, the same concept as "A and / or B" is also applied when three or more things are expressed connected by "and / or."
[0021] [First embodiment]
[0022] FIG. 1 shows an example of the configuration of a data processing system 10 according to the first embodiment.
[0023] 1, a data processing system 10 includes a data processing device 12 and a smart device 14. An example of the data processing device 12 is a server.
[0024] The data processing device 12 includes a computer 22, a database 24, and a communication I / F 26. The computer 22 is an example of a "computer" according to the technology of the present disclosure. The computer 22 includes a processor 28, a RAM 30, and a storage 32. The processor 28, the RAM 30, and the storage 32 are connected to a bus 34. The database 24 and the communication I / F 26 are also connected to the bus 34. The communication I / F 26 is connected to a network 54. Examples of the network 54 include a WAN (Wide Area Network) and / or a LAN (Local Area Network).
[0025] The smart device 14 includes a computer 36, a reception device 38, an output device 40, a camera 42, and a communication I / F 44. The computer 36 includes a processor 46, a RAM 48, and a storage 50. The processor 46, the RAM 48, and the storage 50 are connected to a bus 52. The reception device 38, the output device 40, and the camera 42 are also connected to the bus 52.
[0026] The reception device 38 includes a touch panel 38A, a microphone 38B, and the like, and receives user input. The touch panel 38A detects contact with an indicator (for example, a pen or a finger) to receive user input by the touch of the indicator. The microphone 38B detects the user's voice to receive user input by voice. The control unit 46A transmits data indicating the user input received by the touch panel 38A and the microphone 38B to the data processing device 12. In the data processing device 12, the specific processing unit 290 acquires the data indicating the user input.
[0027] The output device 40 includes a display 40A and a speaker 40B, and presents data to the user 20 by outputting the data in a form of expression that the user 20 can perceive (for example, audio and / or text). The display 40A displays visible information such as text and images in accordance with instructions from the processor 46. The speaker 40B outputs audio in accordance with instructions from the processor 46. The camera 42 is a compact digital camera equipped with an optical system including a lens, aperture, and shutter, and an imaging element such as a CMOS (Complementary Metal-Oxide-Semiconductor) image sensor or a CCD (Charge Coupled Device) image sensor.
[0028] The communication I / F 44 is connected to a network 54. The communication I / Fs 44 and 26 control the exchange of various information between the processor 46 and the processor 28 via the network 54.
[0029] FIG. 2 shows an example of the main functions of the data processing device 12 and the smart device 14.
[0030] 2, in the data processing device 12, a specific process is performed by the processor 28. A specific processing program 56 is stored in the storage 32. The specific processing program 56 is an example of a "program" according to the technology of the present disclosure. The processor 28 reads the specific processing program 56 from the storage 32 and executes the read specific processing program 56 on the RAM 30. The specific process is realized by the processor 28 operating as a specific processing unit 290 in accordance with the specific processing program 56 executed on the RAM 30.
[0031] The storage 32 stores a data generation model 58 and an emotion identification model 59. The data generation model 58 and the emotion identification model 59 are used by the identification processing unit 290.
[0032] In the smart device 14, the processor 46 performs the reception output process. The storage 50 stores a reception output program 60. The reception output program 60 is used in conjunction with the specific processing program 56 by the data processing system 10. The processor 46 reads the reception output program 60 from the storage 50 and executes the read reception output program 60 on the RAM 48. The reception output process is realized by the processor 46 operating as the control unit 46A in accordance with the reception output program 60 executed on the RAM 48.
[0033] Next, a description will be given of the specific processing performed by the specific processing unit 290 of the data processing device 12. In the following description, the data processing device 12 will be referred to as a "server" and the smart device 14 will be referred to as a "terminal."
[0034] The system according to the present invention provides a means for quickly and reliably releasing a security lock on a credit card to prevent fraudulent use through voice authentication. This system is implemented in the following procedure.
[0035] User voice recordings
[0036] The user accesses a dedicated web page and clicks the "Start Recording" button. The device obtains the user's microphone input using navigator.mediaDevices.getUserMedia, creates a MediaRecorder object, and starts recording audio. When the user clicks the "Stop Recording" button, the device stops recording and obtains the recorded audio data in Blob format.
[0037] Sending voice data to the server
[0038] The recorded audio data is sent from the device to the server by sending the audio file to the server via an HTTP POST request via a FormData object.
[0039] Receiving and storing audio data on the server
[0040] The server receives the HTTP POST request for the audio data and stores it in a temporary directory, preparing it for the subsequent audio authentication process.
[0041] Voice Authentication
[0042] The server uses an AI voice authentication model to obtain an authentication score based on the saved voice file. This AI model compares the voice data of pre-registered users to calculate an authentication score. The server evaluates this authentication score, and if the score exceeds a certain threshold, the user's security lock is released.
[0043] Unlocking the security lock
[0044] If the score exceeds the threshold, the server will unlock the credit card security lock based on its internal processing logic, and will also remove any restrictions on card usage after going through the necessary procedures.
[0045] Notification of results
[0046] The server notifies the user of the authentication result. The notification is sent back to the terminal in JSON format, and the terminal parses this message and displays the result to the user. The user can check the voice authentication result and find out whether they can resume using their card.
[0047] Explanation with concrete examples
[0048] Scenario: When a user is using their credit card and their card is suspended due to the card company's logic for detecting fraudulent activity, the user can use this system to safely and quickly release the security lock.
[0049] 1. User opens a web page:
[0050] Users can record their own voice by operating the recording start / stop button.
[0051] 2. Sending audio data to the server:
[0052] The device converts the recorded data into Blob format and sends it to the server.
[0053] 3. Server performs voice authentication:
[0054] The server receives the voice data and uses an AI model to calculate an authentication score.
[0055] 4. Security unlock and notifications:
[0056] - The server determines that the authentication score criteria have been met and releases the security lock. The user is then shown a message indicating successful authentication, allowing them to resume using their credit card.
[0057] In this way, the system is designed to enable users to smoothly release the security lock and quickly resume using their credit card.
[0058] The processing flow will be explained below.
[0059] Step 1:
[0060] A user accesses a web page and clicks the "Start Recording" button. The device uses navigator.mediaDevices.getUserMedia to obtain the user's microphone input. At this time, the user is prompted for microphone access permission in the browser and must grant it.
[0061] Step 2:
[0062] The device starts recording by creating a MediaRecorder object, which records the audio stream from the microphone in real time and collects the audio data.
[0063] Step 3:
[0064] The user clicks the "Stop Recording" button. The device stops recording and the audio data captured by MediaRecorder is retrieved in Blob format. This aggregates the audio data into a single object.
[0065] Step 4:
[0066] The device adds the acquired blob-formatted audio data to a FormData object and sends it to the server using an HTTP POST request. A FormData object is a data structure for sending files or data to the server in multiple key-value formats.
[0067] Step 5:
[0068] The server receives the HTTP POST request, extracts the audio file from the request, stores it in a temporary directory, and later uses it in the authentication process.
[0069] Step 6:
[0070] The server inputs the saved audio file into the authenticate method of the AI voice authentication model, which uses an algorithm to calculate a similarity score with the voice of a pre-registered user.
[0071] Step 7:
[0072] The server evaluates the authentication score returned by the AI model, which includes determining whether the score exceeds a certain threshold, which is set according to the security level.
[0073] Step 8:
[0074] If the score exceeds a threshold, the server will release the security lock based on internal logic, a process that includes updating security settings and removing card usage restrictions.
[0075] Step 9:
[0076] The server returns the authentication result to the user as a JSON message, which includes whether the authentication was successful or not, as well as any additional information required.
[0077] Step 10:
[0078] The user's device receives and analyzes the JSON message returned from the server. Based on the analysis results, the device notifies the user of the authentication result. For example, it displays a message in the browser saying, "Authentication successful. Security lock has been released."
[0079] In this way, a voice authentication system is realized that allows the user to smoothly release the security lock and resume use of the credit card.
[0080] Example 1
[0081] Next, a description will be given of Example 1. In the following description, the data processing device 12 will be referred to as a "server" and the smart device 14 will be referred to as a "terminal."
[0082] Conventional credit card security lock release systems require users to manually perform authentication methods and procedures, which are time-consuming and laborious. Therefore, there was a need for a system that would allow users to quickly and reliably release security locks.
[0083] The specific processing by the specific processing unit 290 of the data processing device 12 in the first embodiment is realized by the following means.
[0084] In this invention, the server includes means for recording the user's voice, means for transmitting the recorded voice data to the server, analysis means having an AI voice authentication model for analyzing the received voice data, means for evaluating the authentication score of the analyzed voice data, means for releasing the security lock based on the evaluated authentication score, means for notifying the user of the result of the voice authentication, and means for transmitting the voice data via an HTTP POST request. This enables the user to easily, quickly, and reliably release the security lock using their voice.
[0085] "User" refers to an individual who uses the System to unlock a credit card security lock.
[0086] "Audio recording means" refers to a device or function for obtaining a user's microphone input and recording that audio data.
[0087] "Means for transmitting audio data to a server" refers to a device or function for transferring recorded audio data to a server using an appropriate protocol.
[0088] "AI voice authentication model" refers to a machine learning algorithm or program that analyzes voice data and performs specific identity verification.
[0089] "Analysis means" refers to a device or function that processes received voice data using an AI voice authentication model and calculates an authentication score.
[0090] "Means for assessing the authentication score" refers to a device or function for determining the calculated authentication score based on the analyzed voice data and determining the next processing step accordingly.
[0091] "Means for unlocking security lock" refers to a device or function for unlocking credit card usage restrictions when the authentication score exceeds a threshold value.
[0092] "Means for notifying the result of voice authentication" refers to a device or function for notifying the user of the result of voice authentication.
[0093] "Means for sending via HTTP POST request" refers to a device or function for sending recorded audio data to a server using the POST method of the HTTP protocol.
[0094] "Interface" refers to a graphical user interface that provides a means for a user to start and stop recording.
[0095] The present invention relates to a system for quickly and reliably releasing a security lock to prevent fraudulent use of a credit card through voice authentication. The system is specifically implemented using the following hardware and software.
[0096] Hardware and Software Used
[0097] 1. User voice recordings
[0098] The user accesses a dedicated web page and operates the "Start Recording" and "Stop Recording" buttons in the browser. The browser uses JavaScript's navigator.mediaDevices.getUserMedia API to obtain the user's microphone input. The audio is recorded using the MediaRecorder object.
[0099] 2. Sending audio data to the server
[0100] After finishing recording, the device converts the recorded audio data into Blob format and sends it to the server via an HTTP POST request using the FormData object. This communication uses the JavaScript fetch API.
[0101] 3. Receiving and storing audio data by the server
[0102] The server receives the audio data via an HTTP POST request and temporarily stores it in a file system managed, for example, using the Node.js fs module.
[0103] 4. Voice Authentication Process
[0104] The server inputs the saved audio file into an AI voice recognition model. This model is pre-trained and compares the user's pre-registered voice data with the target voice data to calculate a recognition score. The AI model is run using libraries such as TensorFlow and PyTorch.
[0105] 5. Unlocking the security lock
[0106] If the authentication score calculated based on the voice data exceeds the threshold, the server executes the procedure to release the credit card security lock. In this step, the API of the card company is called and the necessary procedures are followed to remove the usage restrictions.
[0107] 6. Notification of authentication results
[0108] If authentication is successful, the server notifies the result in JSON format, and the browser receives the response and displays the result to the user.
[0109] Specific examples
[0110] For example, suppose a user is using a credit card and the card is suspended due to the card company's logic for detecting fraudulent activity. In this case, the security lock can be released by following the procedure below.
[0111] 1. Users access a dedicated web page and click the "Start Recording" button to record their own voice.
[0112] 2. When you finish recording, click the "Stop Recording" button and the audio data will be converted to Blob format and sent to the server.
[0113] 3. The server receives the audio data and temporarily stores it.
[0114] 4. The saved voice data is analyzed using an AI voice authentication model to calculate an authentication score.
[0115] 5. If the authentication score exceeds the threshold, the server executes the procedure to unlock the credit card security lock.
[0116] 6. Finally, the server notifies the user of the result, and the user confirms that the credit card can be used again.
[0117] This allows users to quickly and reliably release security locks using voice commands.
[0118] The flow of the identification process in the first embodiment will be described with reference to FIG.
[0119] Step 1:
[0120] The user accesses a dedicated web page and clicks the "Start Recording" button. The device uses the navigator.mediaDevices.getUserMedia API to obtain the user's microphone input, creates a MediaRecorder object, and begins recording audio. Specifically, the browser displays a dialog requesting permission to use the microphone. Once audio input is obtained, MediaRecorder starts and recording begins. The input at this time is the user's voice, and the output is the audio data being recorded.
[0121] Step 2:
[0122] When the user clicks the "Stop Recording" button, the device stops recording and obtains the recorded audio data in Blob format. Specifically, the MediaRecorder object stops recording data, and the resulting audio data is saved in Blob format on the device. The input at this time is the audio data being recorded, and the output is audio data in Blob format.
[0123] Step 3:
[0124] The device adds the recorded audio data in Blob format to a FormData object and sends it to the server via an HTTP POST request. This communication is performed using JavaScript's fetch API. Specifically, the audio data is added to the FormData object, and the fetch API sends an HTTP POST request to the server. The input at this time is audio data in Blob format, and the output is an HTTP POST request to the server.
[0125] Step 4:
[0126] The server receives an HTTP POST request containing audio data and temporarily stores the data in a directory. The data received on the server side is saved in a local temporary directory using, for example, the Node.js fs module. The input is the HTTP POST request sent to the server, and the output is the audio data saved in the temporary directory.
[0127] Step 5:
[0128] The server inputs the saved voice data into an AI voice recognition model and calculates a recognition score. The AI voice recognition model is a pre-trained model that compares the voice data to calculate a recognition score. Specifically, the AI model is run using libraries such as TensorFlow and PyTorch. The input is the voice data saved in a temporary directory, and the output is a recognition score.
[0129] Step 6:
[0130] The server evaluates the authentication score and, if the score exceeds the threshold, releases the credit card's security lock. The server then calls the card company's API and executes the procedure to release the usage restriction. Specifically, it sends an HTTP request to the card company's API endpoint and obtains confirmation of the release. The input is the authentication score, and the output is confirmation of the release from the card company.
[0131] Step 7:
[0132] The server notifies the user of the authentication result in a JSON format message. The device parses this message and displays the authentication result to the user. Specifically, the server generates a JSON format response and sends it to the device as an HTTP response. The device receives this response and displays the authentication result on the screen. The input at this time is the JSON message of the authentication result generated by the server, and the output is the authentication result displayed on the device.
[0133] (Application example 1)
[0134] Next, a description will be given of Application Example 1. In the following description, the data processing device 12 will be referred to as a "server" and the smart device 14 will be referred to as a "terminal."
[0135] Conventional credit card fraud prevention systems require complicated procedures to activate security locks, making it difficult to quickly unlock them. Furthermore, the interface and notification methods for users to authenticate using their own voice are inadequate, reducing user convenience. Given these circumstances, there is a demand for a system that can quickly and reliably unlock credit card security locks and instantly notify users of the results.
[0136] The specific processing by the specific processing unit 290 of the data processing device 12 in the application example 1 is realized by the following means.
[0137] In this invention, the server includes means for recording the user's voice, means for transmitting the recorded voice data to the server, analysis means having a voice authentication model for analyzing the received voice data, means for evaluating the authentication score of the analyzed voice data, means for releasing the security lock based on the evaluated authentication score, display means including operation buttons for recording and transmitting the voice, and means for notifying the user of the result. This allows the user to quickly release the security lock on the credit card using their own voice and immediately confirm the result.
[0138] "User" means an individual who uses the system to record voice and attempt to unlock a credit card security lock.
[0139] "Audio recording means" refers to a device or application for recording a user's voice.
[0140] The "means for transmitting recorded voice data to a server" refers to a communication means for uploading recorded voice data to a server via a network.
[0141] A "voice authentication model" is an algorithm or AI model that analyzes recorded voice data and authenticates whether it is the user's voice.
[0142] The "analysis means" is a device that includes hardware and software for analyzing the voice data sent to the server.
[0143] The "means for evaluating the authentication score" refers to a device or program for determining whether authentication is successful or unsuccessful based on the score calculated by the voice authentication model.
[0144] A "means for unlocking security lock" is a device or program that executes a procedure to unlock a credit card when the authentication score exceeds a reference value.
[0145] The "display means including operation buttons" refers to a screen on which buttons for the user to operate to start and stop voice recording are displayed.
[0146] The "notification means" is a device or program for notifying the user of the result of voice authentication.
[0147] The term "computer system" refers to a collection of hardware and software that includes the above-mentioned means and operates as a whole.
[0148] This invention is a security system that allows users to quickly and reliably release the security lock on their credit cards using their own voice. This system is realized by combining the procedures of voice recording, data transmission, voice authentication, and result notification. The details are described below.
[0149] server
[0150] The server first receives the voice data sent by the user. After receiving the voice data, the AI voice authentication model within the server analyzes the voice data and calculates an authentication score. In this process, a machine learning model (e.g., TensorFlow, PyTorch) is used to determine whether the specific voice data matches the voice data of a pre-registered user. If the authentication score exceeds a threshold, the server releases the credit card security lock based on its internal logic. The authentication result is then returned to the user in JSON format, allowing interest confirmation.
[0151] Terminal
[0152] The device (such as the user's smartphone or computer) provides a user interface and displays operation buttons (record start button and record stop button) that the user can use to record audio. It records audio in response to specific operations, converts the audio data into Blob format, and sends it to the server. It also displays the authentication result to the user based on the authentication result returned from the server.
[0153] User
[0154] When a security lock is activated, the user opens the provided browser-based application and records their voice. After recording, the voice data is sent to the server, after which the authentication result is notified. This process allows the user to smoothly unlock the security lock on their credit card and resume use.
[0155] Specific examples
[0156] A user on a business trip attempts to use their credit card in a new city, but the card is locked due to fraud prevention logic. In this case, the user launches the smartphone app, records their voice, and sends it to the server. The server receives the voice data and calculates an authentication score using an AI voice authentication model. If authentication is successful, the server releases the card's security lock and notifies the user of the result. The user receives the notification and can resume using their credit card while on a business trip.
[0157] Prompt Sentence Examples
[0158] User: "Please press the start recording button and say the following phrase: 'This audio will verify my identity.'"
[0159] Input prompt to AI model: "Please authenticate the voice data below to confirm whether this voice belongs to a registered user."
[0160] The flow of the specific processing in the application example 1 will be described with reference to FIG.
[0161] Step 1:
[0162] The user opens a security application or web page, which displays a "Start Recording" button to start recording and a "Stop Recording" button to stop recording. The input of this step is the user's operation, and the output is the display of the user interface.
[0163] Step 2:
[0164] When the user presses the "Start Recording" button, the device accesses the microphone using navigator.mediaDevices.getUserMedia and starts recording. The recorded audio is temporarily saved on the device. The inputs to this step are the user's operation and audio data, and the outputs are the recording status and the temporarily saved audio data.
[0165] Step 3:
[0166] When the user presses the "Stop Recording" button, the device stops recording using the MediaRecorder object and converts the audio data to Blob format, which is then prepared for sending to the server. The input to this step is the audio data being recorded, and the output is Blob format audio data.
[0167] Step 4:
[0168] The device adds the audio data in Blob format to a FormData object and sends it to the server using an HTTP POST request using the fetch API. The input of this step is the audio data in Blob format, and the output is the data sent to the server.
[0169] Step 5:
[0170] The server saves the audio data received via the HTTP POST request in a temporary directory. Here, it checks whether the received audio data has been received and saved correctly. The input of this step is the HTTP request, and the output is the audio file on the server.
[0171] Step 6:
[0172] The server analyzes the received voice data using an AI voice authentication model (e.g., TensorFlow, PyTorch) and calculates an authentication score. The voice authentication model compares the voice data with that of registered users to measure the degree of match. The input for this step is the received voice data, and the output is an authentication score.
[0173] Step 7:
[0174] The server evaluates the calculated authentication score to determine if it exceeds a threshold. If so, it executes a procedure to unlock the credit card security lock. The input to this step is the authentication score, and the output is the execution of the unlock.
[0175] Step 8:
[0176] The server returns the authentication result (success or failure) to the device as a JSON-formatted message. The input to this step is the result of the authentication process, and the output is a JSON-formatted message.
[0177] Step 9:
[0178] The terminal analyzes the authentication result received from the server and notifies the user through the user interface. The user can check on the screen whether the credit card can be used or not. The input of this step is the authentication result in JSON format, and the output is the message displayed to the user.
[0179] Furthermore, an emotion engine that estimates the user's emotion may be combined. That is, the identification processing unit 290 may estimate the user's emotion using the emotion identification model 59 and perform identification processing using the user's emotion.
[0180] The system according to the present invention provides a means for quickly and reliably unlocking a credit card's security lock to prevent fraudulent use through voice authentication and emotion recognition. This system is implemented in the following steps.
[0181] User voice recordings
[0182] The user accesses a dedicated web page and clicks the "Start Recording" button. The device uses navigator.mediaDevices.getUserMedia to obtain the user's microphone input, creates a MediaRecorder object, and starts recording audio. When the user clicks the "Stop Recording" button, the device stops recording and obtains the recorded audio data in Blob format.
[0183] Sending voice data to the server
[0184] The recorded audio data is sent from the device to the server by sending the audio file to the server via an HTTP POST request via a FormData object.
[0185] Receiving and storing audio data on the server
[0186] The server receives an HTTP POST request for audio data, extracts the audio file from the request, and stores it in a temporary directory for later use in the authentication process.
[0187] Voice Authentication
[0188] The server inputs the saved audio file into the authenticate method of the AI voice authentication model. This model compares it with the voice data of pre-registered users and calculates an authentication score. In parallel, the emotion engine analyzes the user's emotions from the audio data and obtains the results.
[0189] Authentication score and sentiment rating
[0190] The server evaluates the authentication score returned by the voice authentication model and the emotion data obtained from the emotion engine. This evaluation includes calculating an overall score that combines the authentication score and the emotion data. For example, if the user expresses anger or tension, this will be taken into account as a factor that affects the overall score.
[0191] Unlocking the security lock
[0192] If the total score exceeds a certain threshold, the server will release the security lock based on internal logic, a process that includes updating security settings and removing card usage restrictions.
[0193] Notification of results
[0194] The server returns the authentication result to the user as a JSON message. The message contains information about whether the authentication was successful or not, as well as any additional information required. The user's device receives this message, parses it, and notifies the user of the result.
[0195] Explanation with concrete examples
[0196] Scenario: When a user is using their credit card and their card is suspended due to the card company's logic for detecting fraudulent activity, the user can use this system to safely and quickly release the security lock.
[0197] 1. User opens a web page:
[0198] The user can record their own voice by operating the recording start / stop button. After recording is complete, the device sends the voice data to the server.
[0199] 2. Receiving and storing audio data by the server:
[0200] The server receives the audio data and stores it in a temporary directory.
[0201] 3. Server performs voice authentication and emotion recognition:
[0202] The server analyzes the voice data using an AI voice recognition model to calculate a recognition score. At the same time, the emotion engine analyzes emotions from the voice data and obtains emotion data.
[0203] 4. Overall score evaluation and security unlock:
[0204] The server evaluates the authentication score and emotional data together, and unlocks the credit card security lock if the criteria are met.
[0205] 5. Notification of Results:
[0206] The server notifies the user of the authentication result, and the user can confirm the result and resume using the card.
[0207] In this way, the system is designed to combine voice authentication and emotion recognition to enable users to smoothly release security locks and quickly resume using their credit cards.
[0208] The processing flow will be explained below.
[0209] Step 1:
[0210] A user accesses a web page and clicks the "Start Recording" button. The device uses navigator.mediaDevices.getUserMedia to obtain the user's microphone input. At this time, the user is prompted for microphone access permission in the browser and must grant it.
[0211] Step 2:
[0212] The device starts recording by creating a MediaRecorder object, which records the audio stream from the microphone in real time and collects the audio data.
[0213] Step 3:
[0214] The user clicks the "Stop Recording" button. The device stops recording and the audio data captured by MediaRecorder is retrieved in Blob format. This aggregates the audio data into a single object.
[0215] Step 4:
[0216] The device adds the acquired blob-formatted audio data to a FormData object and sends it to the server using an HTTP POST request. A FormData object is a data structure for sending files or data to the server in multiple key-value formats.
[0217] Step 5:
[0218] The server receives the HTTP POST request, extracts the audio file from the request, stores it in a temporary directory, and later uses it in the authentication process.
[0219] Step 6:
[0220] The server inputs the saved audio file into the authenticate method of the AI voice authentication model, which uses an algorithm to calculate a similarity score with the voice of a pre-registered user.
[0221] Step 7:
[0222] The server also sends the voice data to the emotion engine, which analyzes the user's emotions. The emotion engine analyzes the voice characteristics and determines whether the user is expressing anger, joy, sadness, or other emotions.
[0223] Step 8:
[0224] The server combines the authentication score returned by the AI model with the emotion data obtained from the emotion engine to calculate an overall score, which is generated using a calculation algorithm based on the voice authentication score and emotion data.
[0225] Step 9:
[0226] If the total score exceeds a certain threshold, the server will release the security lock based on internal logic, a process that includes updating security settings and removing card usage restrictions.
[0227] Step 10:
[0228] The server returns the authentication result to the user as a JSON message, which includes whether the authentication was successful or not, as well as any additional information required.
[0229] Step 11:
[0230] The user's device receives and analyzes the JSON message returned from the server. Based on the analysis results, the device notifies the user of the authentication result. For example, it displays a message in the browser saying, "Authentication successful. Security lock has been released."
[0231] Explanation with concrete examples
[0232] If a user is using a credit card and their card is suspended due to the card company's logic for detecting fraudulent activity, the user can use this system to safely and quickly release the security lock.
[0233] example:
[0234] The user opens a web page and records their own voice by operating the recording start / stop button. After recording is complete, the device sends the voice data to the server. The server receives the voice data and saves it in a temporary directory. The server analyzes the voice data using an AI voice authentication model and calculates an authentication score. At the same time, the emotion engine analyzes emotions from the voice data and obtains emotional data. The server evaluates the authentication score and emotional data together, and if the criteria are met, the credit card security lock is released. Finally, the server notifies the user of the authentication result, who can confirm the result and resume using their card.
[0235] In this way, the system is designed to combine voice authentication and emotion recognition to enable users to smoothly release security locks and quickly resume using their credit cards.
[0236] Example 2
[0237] Next, a description will be given of Example 2. In the following description, the data processing device 12 will be referred to as a "server" and the smart device 14 will be referred to as a "terminal."
[0238] In conventional credit card fraud prevention systems, voice authentication alone does not provide sufficient security, making it difficult to quickly and safely release the security lock. In particular, the user's emotional state, such as tension or anger, can affect the authentication results, raising the risk of incorrect judgment.
[0239] The specific processing by the specific processing unit 290 of the data processing device 12 in the second embodiment is realized by the following means.
[0240] In this invention, the server includes means for recording the user's voice, means for transmitting the recorded voice data to the server, analysis means having a voice authentication model for analyzing the received voice data, analysis means having an emotion recognition engine for analyzing emotions from the voice data, means for evaluating the authentication score and emotion data of the analyzed voice data, means for releasing the security lock based on a total score of the evaluated authentication score and emotion data, and means for notifying the user of the authentication result. This enables a comprehensive evaluation based on a combination of voice authentication and emotion recognition, thereby enabling the security lock to be released more quickly and safely.
[0241] "User" refers to a person who attempts to use this system to unlock the security lock on a credit card.
[0242] "Means for recording audio" refers to the function for capturing the user's audio and saving it as digital data.
[0243] "Server" refers to a computer system that includes a central processing unit that receives and analyzes audio data.
[0244] "Audio data" refers to data that represents audio information recorded by a user in digital form.
[0245] "Voice authentication model" refers to an algorithm or program that analyzes voice data and calculates an authentication score based on the user's voice characteristics.
[0246] "Analysis means" refers to a program or process for evaluating audio data and other input data.
[0247] "Emotion recognition engine" refers to a program or algorithm that analyzes a user's emotional state from voice data and generates emotional data.
[0248] "Authentication score" refers to a numerical value calculated by a voice authentication model that indicates the degree to which a user's voice matches a pre-enrolled voice.
[0249] "Emotion data" refers to information about emotions determined from the user's voice as analyzed by an emotion recognition engine.
[0250] "Total score" refers to an evaluation value calculated by integrating the authentication score and emotional data.
[0251] "Means to unlock security lock" refers to the function for lifting restrictions on a user's credit card usage based on the overall score.
[0252] "Notification means" refers to the interface or program used to communicate authentication results and system status to users.
[0253] The system based on this invention combines voice authentication and emotion recognition to provide a means for quickly and reliably unlocking security locks to prevent fraudulent use of credit cards. This system uses the following hardware and software to process data and perform calculations.
[0254] Hardware and software used
[0255] Hardware:
[0256] Microphone: A device that inputs the user's voice.
[0257] Server: A central processing unit that receives, stores, and analyzes voice data.
[0258] software:
[0259] Web browsers: Browsers that use navigator.mediaDevices.getUserMedia to get microphone input.
[0260] Voice authentication model: An algorithm that analyzes voice data and calculates an authentication score.
[0261] Emotion recognition engine: A program that analyzes emotions from voice data.
[0262] Example of system operation
[0263] 1. User visits a web page:
[0264] The user opens a browser and accesses the specified web page.
[0265] For example, a user visits the URL https: / / securitylock.example.com.
[0266] 2. Audio Recording:
[0267] The user clicks the "Start Recording" button and follows the instructions provided to record their voice.
[0268] The device uses navigator.mediaDevices.getUserMedia to obtain the user's microphone input, creates a MediaRecorder object, and starts recording audio.
[0269] When the user clicks the "Stop Recording" button, the device stops recording and obtains the recorded audio data in Blob format.
[0270] 3. Sending audio data to the server:
[0271] The terminal sends the recorded audio data to the server via a FormData object.
[0272] Use the fetch method to send it to the server as an HTTP POST request.
[0273] 4. Receiving and storing audio data by the server:
[0274] The server receives the HTTP POST request and retrieves the audio data.
[0275] The acquired audio data is saved in a temporary directory.
[0276] 5. Voice Verification and Emotion Recognition:
[0277] The server inputs the saved audio file into an AI voice recognition model and calculates a recognition score.
[0278] At the same time, an emotion recognition engine is used to analyze the user's emotional state from the voice data and obtain emotion data.
[0279] 6. Overall score evaluation and security unlock:
[0280] The server calculates a total score based on the authentication score from the voice authentication model and the emotion data from the emotion recognition engine.
[0281] If the total score meets a certain standard, the server will unlock the credit card's security lock.
[0282] 7. Notification of Results:
[0283] The server returns the authentication result to the user as a JSON message.
[0284] The terminal analyzes the result message received from the server and notifies the user of the result.
[0285] Prompt Sentence Examples
[0286] Please send the audio data recorded by the user in the following format.
[0287] Format: Blob format
[0288] Sending method: HTTP POST
[0289] Object used: FormData
[0290] The system is designed to combine voice authentication and emotion recognition to seamlessly unlock credit card security locks, allowing users to quickly resume using their cards.
[0291] The flow of the identification process in the second embodiment will be described with reference to FIG.
[0292] Step 1:
[0293] A user visits a web page:
[0294] A user opens a browser and accesses a specified web page. For example, a user accesses the URL https: / / securitylock.example.com. The input is the user's action (accessing the URL) and the output is the displayed web page.
[0295] Step 2:
[0296] User starts audio recording:
[0297] The user clicks the "Start Recording" button on a web page. The device uses navigator.mediaDevices.getUserMedia to obtain the user's microphone input. Specifically, the browser asks the user for permission to use the microphone, and if permission is granted, the microphone stream is obtained. The input is the user's click operation and microphone input, and the output is the audio stream.
[0298] Step 3:
[0299] The device records audio:
[0300] The device creates a MediaRecorder object and starts recording audio. The data being recorded is temporarily stored in memory. The input is the microphone stream, and the output is a buffer of the data being recorded.
[0301] Step 4:
[0302] User stops recording:
[0303] The user clicks the "Stop Recording" button. This causes the device to stop recording and obtain the recorded audio data in Blob format. Specifically, the recorder.stop() method is called and the audio data is obtained in the ondataavailable event. The input is the user's click operation, and the output is audio data in Blob format.
[0304] Step 5:
[0305] The device sends the audio data to the server:
[0306] The device adds the acquired audio data to a FormData object and uses the fetch method to send it to the server as an HTTP POST request. Specifically, the following operations are performed: let formData = new FormData(); formData.append('audio', audioBlob, 'audio.wav'); fetch(' / upload', { method: 'POST', body: formData}); The input is audio data in Blob format, and the output is an HTTP request to the server.
[0307] Step 6:
[0308] The server receives the audio data:
[0309] The server receives the HTTP POST request and extracts the audio data. The extracted audio data is saved in a temporary directory. Specifically, the server-side code retrieves the audio file from req.file or similar and saves it using the fs.writeFile method. The input is the HTTP request, and the output is the saved audio file.
[0310] Step 7:
[0311] The server performs voice authentication and emotion recognition:
[0312] The server inputs the saved audio file into the AI voice authentication model and calculates an authentication score. For example, the code "let score = voiceAuthModel.authenticate(' / temp / audio.wav');" is executed. At the same time, the emotion engine analyzes the user's emotional state from the audio data and executes the operation "let emotionData = emotionEngine.analyze(' / temp / audio.wav');" to obtain emotion data. The input is the audio file, and the output is the authentication score and emotion data.
[0313] Step 8:
[0314] The server evaluates the overall score:
[0315] The server calculates the total score based on the authentication score from the voice authentication model and the emotion data. For example, it executes the operation let totalScore = calculateTotalScore(score, emotionData);. The input is the authentication score and emotion data, and the output is the total score.
[0316] Step 9:
[0317] Server removes security lock:
[0318] If the total score meets a certain criterion, the server unlocks the credit card's security lock. For example, the following operation is performed: if (totalScore > threshold) { unlockSecurityLock();}. The input is the total score, and the output is the unlock status of the credit card.
[0319] Step 10:
[0320] The server notifies the user of the authentication result:
[0321] The server returns the authentication result to the user as a JSON message. Specifically, the following operation is performed: res.json({ success: true, message: 'Security lock has been released'});. The input is the authentication result, and the output is the message sent to the user.
[0322] Step 11:
[0323] User sees the results:
[0324] The terminal analyzes the result message received from the server and notifies the user. For example, it performs the following operation: const result = await response.json(); alert(result.message);. The input is the message from the server, and the output is the notification to the user.
[0325] (Application example 2)
[0326] Next, a description will be given of Application Example 2. In the following description, the data processing device 12 will be referred to as a "server" and the smart device 14 will be referred to as a "terminal."
[0327] Conventional security systems can cause inconvenience to users due to delays in detecting unauthorized use and unlocking the device. Furthermore, simple voice authentication alone cannot take into account user conditions such as emotional changes and stress. As a result, legitimate users are unable to unlock the security lock, resulting in a poor user experience.
[0328] The identification process by the identification processing unit 290 of the data processing device 12 in Application Example 2 is realized by the following means. In this invention, the server includes means for recording the user's voice, means for transmitting the recorded voice data to the server, analysis means having a voice authentication model and an emotion recognition engine for analyzing the received voice data, means for evaluating the authentication score and emotion data of the analyzed voice data, and means for calculating a total score based on the evaluated authentication score and emotion data and for releasing the security lock using the total score. This allows the legitimate user to release the security lock quickly and safely.
[0329] A "user" is an individual who records voice and uses the recorded voice data.
[0330] "Audio data" means sound in the form of data recorded by a user and transmitted to a server.
[0331] A "server" is a computer system for analyzing received voice data and performing voice authentication and emotion recognition.
[0332] "Analysis means" refers to a set of functions for analyzing audio data using a voice authentication model and an emotion recognition engine.
[0333] The "voice authentication model" is a machine learning model that analyzes received voice data and calculates an authentication score.
[0334] An "emotion recognition engine" is software that analyzes and obtains a user's emotional data from voice data.
[0335] The "authentication score" is a score calculated by the voice authentication model that indicates the degree to which the user's voice matches pre-registered data.
[0336] "Emotion data" is data that indicates the emotional state of the user, analyzed from the voice data.
[0337] The "total score" is an evaluation score calculated by combining the authentication score of the voice authentication model and emotion data.
[0338] A "security lock" is a security measure that limits the use of credit cards.
[0339] "Unlocking means" refers to a set of functions for unlocking a security lock based on the overall score.
[0340] "Notification means" refers to a method or technical means for notifying the user of the results of voice authentication and emotion recognition.
[0341] This invention relates to a system that records a user's voice, transmits the voice data to a server for analysis, and quickly and safely releases a security lock through voice authentication and emotion recognition. This system is implemented with the following configuration and procedure.
[0342] Voice recording and data transmission
[0343] The device is equipped with a microphone for recording the user's voice, a function for acquiring voice input using navigator.mediaDevices.getUserMedia, and a function for recording and saving voice data using the MediaRecorder object. The user records voice through an interface for starting and stopping recording, and sends the voice data to the server after recording is complete. To send data to the server, the voice data is added to a FormData object and sent to the server using an HTTP POST request.
[0344] Receiving and analyzing audio data
[0345] The server has the ability to store the received voice data in a temporary directory and analyze it. A voice authentication model and an emotion recognition engine are used for the analysis. The voice authentication model analyzes the user's voice data and calculates an authentication score that indicates the degree of match with pre-registered data. Meanwhile, the emotion recognition engine extracts the user's emotion data from the voice data.
[0346] Authentication score and emotion data evaluation
[0347] The server evaluates the authentication score of the voice authentication model and the emotion data obtained from the emotion recognition engine. If the overall score exceeds a certain threshold, the security lock is released.
[0348] Notification of results
[0349] The server returns the results of voice authentication and emotion recognition to the user as a JSON-formatted message. The message includes the result of authentication (success or failure) and any additional information required. The device receives this message, analyzes it, and notifies the user of the result.
[0350] Specific examples
[0351] For example, if a user tries to use their credit card but finds it is locked, they can launch the "Safe Payment Release App" on their smartphone and press the "Start Recording" button to record their voice. The recorded data is automatically sent to a server, where voice authentication and emotion recognition are performed. If the security lock can be released, the user is notified of the result and can use their credit card again.
[0352] Prompt Sentence Examples
[0353] "Please use this user's voice recording data to perform voice authentication and emotion recognition to determine whether the security lock can be released. Please use the following audio file for voice authentication and emotion recognition."
[0354] The system allows the server and terminal to work together to provide users with a quick and reliable security unlocking experience.
[0355] The flow of the specific processing in the application example 2 will be described with reference to FIG.
[0356] Step 1:
[0357] The user launches the "Safe Payment Cancellation App" on their smartphone and presses the "Start Recording" button. The device obtains microphone input using navigator.mediaDevices.getUserMedia. The device then creates a MediaRecorder object and begins recording audio. The input here is the user's voice, and the output is audio data obtained from the microphone. Specifically, when the user presses the button, the device begins microphone input and collects audio data in real time.
[0358] Step 2:
[0359] When the user presses the "Stop Recording" button, the device stops recording and obtains the recorded audio data in Blob format. The input here is the user operation that triggers the end of recording, and the output is audio data in Blob format. Specifically, when the user presses the button, the device stops MediaRecorder and converts the audio data stored in the internal buffer into Blob format.
[0360] Step 3:
[0361] The device adds this audio data in Blob format to a FormData object and sends it to the server using an HTTP POST request. The input is the audio data in Blob format, and the output is a FormData object sent to the server. Specifically, the device adds the audio data to FormData and issues an HTTP POST request to the endpoint to the server.
[0362] Step 4:
[0363] The server saves the received audio data in a temporary directory. The input is the audio data sent in the HTTP POST request, and the output is an audio file saved in the temporary directory. Specifically, the server receives the request, extracts the audio data, and saves it in a temporary directory.
[0364] Step 5:
[0365] The server inputs the saved audio file into the authenticate method of the voice authentication model and calculates an authentication score. The emotion recognition engine also analyzes and obtains the user's emotional data from the audio data. The input is the audio file, and the output is an authentication score and emotional data. Specifically, the server calls the voice authentication model and emotion recognition engine to perform the analysis.
[0366] Step 6:
[0367] The server calculates an overall score by combining the authentication score from the voice authentication model and the emotion data from the emotion recognition engine, and releases the security lock if the overall score exceeds a threshold. The input is the authentication score and emotion data, and the output is an instruction to release the security lock. Specifically, the server evaluates the score and updates the credit card security settings as necessary.
[0368] Step 7:
[0369] The server returns the results of voice authentication and emotion recognition to the user as a JSON-formatted message. The input is the overall score and evaluation result, and the output is a JSON-formatted result message sent to the user. Specifically, the server formats the results and sends an HTTP response to the user's device.
[0370] Step 8:
[0371] The user's device parses the JSON-formatted message received from the server and notifies the user of the results. The input is a JSON-formatted message, and the output is a notification to the user. Specifically, the device parses the message and displays the results to the user through an alert or notification.
[0372] The above are the specific processing steps of this system.
[0373] The specific processing unit 290 transmits the result of the specific processing to the smart device 14. In the smart device 14, the control unit 46A causes the output device 40 to output the result of the specific processing. The microphone 38B acquires audio indicating a user input regarding the result of the specific processing. The control unit 46A transmits audio data indicating the user input acquired by the microphone 38B to the data processing device 12. In the data processing device 12, the specific processing unit 290 acquires the audio data.
[0374] The data generation model 58 is a so-called generative AI (Artificial Intelligence). An example of the data generation model 58 is ChatGPT (Internet Search<URL: https: / / openai.com / blog / chatgpt> ), Gemini (Internet search <url: https: gemini.google.com ?hl="ja">) and other generation AIs. The data generation model 58 is obtained by performing deep learning on a neural network. A prompt including an instruction is input to the data generation model 58, and inference data such as voice data indicating voice, text data indicating text, and image data indicating an image is also input. The data generation model 58 performs inference on the input inference data in accordance with the instruction indicated by the prompt, and outputs the inference result in a data format such as voice data and text data. Here, inference refers to, for example, analysis, classification, prediction, and / or summarization.
[0375] In the above embodiment, an example in which the specific process is performed by the data processing device 12 has been given, but the technology of the present disclosure is not limited to this, and the specific process may be performed by the smart device 14.
[0376] [Second embodiment]
[0377] FIG. 3 shows an example of the configuration of a data processing system 210 according to the second embodiment.
[0378] 3, the data processing system 210 includes the data processing device 12 and smart glasses 214. An example of the data processing device 12 is a server.
[0379] The data processing device 12 includes a computer 22, a database 24, and a communication I / F 26. The computer 22 is an example of a "computer" according to the technology of the present disclosure. The computer 22 includes a processor 28, a RAM 30, and a storage 32. The processor 28, the RAM 30, and the storage 32 are connected to a bus 34. The database 24 and the communication I / F 26 are also connected to the bus 34. The communication I / F 26 is connected to a network 54. Examples of the network 54 include a WAN (Wide Area Network) and / or a LAN (Local Area Network).
[0380] The smart glasses 214 include a computer 36, a microphone 238, a speaker 240, a camera 42, and a communication I / F 44. The computer 36 includes a processor 46, a RAM 48, and a storage 50. The processor 46, the RAM 48, and the storage 50 are connected to a bus 52. The microphone 238, the speaker 240, and the camera 42 are also connected to the bus 52.
[0381] The microphone 238 receives instructions and the like from the user 20 by receiving voice uttered by the user 20. The microphone 238 captures the voice uttered by the user 20, converts the captured voice into audio data, and outputs it to the processor 46. The speaker 240 outputs audio in accordance with instructions from the processor 46.
[0382] Camera 42 is a small digital camera equipped with an optical system including a lens, aperture, and shutter, and an imaging element such as a CMOS (Complementary Metal-Oxide-Semiconductor) image sensor or a CCD (Charge Coupled Device) image sensor, and captures images of the surroundings of user 20 (for example, an imaging range defined by an angle of view equivalent to the field of vision of a typical healthy person).
[0383] The communication I / F 44 is connected to a network 54. The communication I / Fs 44 and 26 are responsible for the exchange of various information between the processor 46 and the processor 28 via the network 54. The exchange of various information between the processor 46 and the processor 28 using the communication I / Fs 44 and 26 is carried out in a secure state.
[0384] Fig. 4 shows an example of the main functions of the data processing device 12 and the smart glasses 214. As shown in Fig. 4, in the data processing device 12, a specific process is performed by the processor 28. A specific process program 56 is stored in the storage 32.
[0385] The specific processing program 56 is an example of a "program" according to the technology of the present disclosure. The processor 28 reads the specific processing program 56 from the storage 32 and executes the read specific processing program 56 on the RAM 30. The specific processing is realized by the processor 28 operating as a specific processing unit 290 in accordance with the specific processing program 56 executed on the RAM 30.
[0386] The storage 32 stores a data generation model 58 and an emotion identification model 59. The data generation model 58 and the emotion identification model 59 are used by the identification processing unit 290.
[0387] In the smart glasses 214, the reception output process is performed by the processor 46. A reception output program 60 is stored in the storage 50. The processor 46 reads the reception output program 60 from the storage 50 and executes the read reception output program 60 on the RAM 48. The reception output process is realized by the processor 46 operating as the control unit 46A in accordance with the reception output program 60 executed on the RAM 48.
[0388] Next, a description will be given of the identification process performed by the identification processing unit 290 of the data processing device 12. In the following description, the data processing device 12 will be referred to as the "server" and the smart glasses 214 will be referred to as the "terminal."
[0389] The system according to the present invention provides a means for quickly and reliably releasing a security lock on a credit card to prevent fraudulent use through voice authentication. This system is implemented in the following procedure.
[0390] User voice recordings
[0391] The user accesses a dedicated web page and clicks the "Start Recording" button. The device obtains the user's microphone input using navigator.mediaDevices.getUserMedia, creates a MediaRecorder object, and starts recording audio. When the user clicks the "Stop Recording" button, the device stops recording and obtains the recorded audio data in Blob format.
[0392] Sending voice data to the server
[0393] The recorded audio data is sent from the device to the server by sending the audio file to the server via an HTTP POST request via a FormData object.
[0394] Receiving and storing audio data on the server
[0395] The server receives the HTTP POST request for the audio data and stores it in a temporary directory, preparing it for the subsequent audio authentication process.
[0396] Voice Authentication
[0397] The server uses an AI voice authentication model to obtain an authentication score based on the saved voice file. This AI model compares the voice data of pre-registered users to calculate an authentication score. The server evaluates this authentication score, and if the score exceeds a certain threshold, the user's security lock is released.
[0398] Unlocking the security lock
[0399] If the score exceeds the threshold, the server will unlock the credit card security lock based on its internal processing logic, and will also remove any restrictions on card usage after going through the necessary procedures.
[0400] Notification of results
[0401] The server notifies the user of the authentication result. The notification is sent back to the terminal in JSON format, and the terminal parses this message and displays the result to the user. The user can check the voice authentication result and find out whether they can resume using their card.
[0402] Explanation with concrete examples
[0403] Scenario: When a user is using their credit card and their card is suspended due to the card company's logic for detecting fraudulent activity, the user can use this system to safely and quickly release the security lock.
[0404] 1. User opens a web page:
[0405] Users can record their own voice by operating the recording start / stop button.
[0406] 2. Sending audio data to the server:
[0407] The device converts the recorded data into Blob format and sends it to the server.
[0408] 3. Server performs voice authentication:
[0409] The server receives the voice data and uses an AI model to calculate an authentication score.
[0410] 4. Security unlock and notifications:
[0411] - The server determines that the authentication score criteria have been met and releases the security lock. The user is then shown a message indicating successful authentication, allowing them to resume using their credit card.
[0412] In this way, the system is designed to enable users to smoothly release the security lock and quickly resume using their credit card.
[0413] The processing flow will be explained below.
[0414] Step 1:
[0415] A user accesses a web page and clicks the "Start Recording" button. The device uses navigator.mediaDevices.getUserMedia to obtain the user's microphone input. At this time, the user is prompted for microphone access permission in the browser and must grant it.
[0416] Step 2:
[0417] The device starts recording by creating a MediaRecorder object, which records the audio stream from the microphone in real time and collects the audio data.
[0418] Step 3:
[0419] The user clicks the "Stop Recording" button. The device stops recording and the audio data captured by MediaRecorder is retrieved in Blob format. This aggregates the audio data into a single object.
[0420] Step 4:
[0421] The device adds the acquired blob-formatted audio data to a FormData object and sends it to the server using an HTTP POST request. A FormData object is a data structure for sending files or data to the server in multiple key-value formats.
[0422] Step 5:
[0423] The server receives the HTTP POST request, extracts the audio file from the request, stores it in a temporary directory, and later uses it in the authentication process.
[0424] Step 6:
[0425] The server inputs the saved audio file into the authenticate method of the AI voice authentication model, which uses an algorithm to calculate a similarity score with the voice of a pre-registered user.
[0426] Step 7:
[0427] The server evaluates the authentication score returned by the AI model, which includes determining whether the score exceeds a certain threshold, which is set according to the security level.
[0428] Step 8:
[0429] If the score exceeds a threshold, the server will release the security lock based on internal logic, a process that includes updating security settings and removing card usage restrictions.
[0430] Step 9:
[0431] The server returns the authentication result to the user as a JSON message, which includes whether the authentication was successful or not, as well as any additional information required.
[0432] Step 10:
[0433] The user's device receives and analyzes the JSON message returned from the server. Based on the analysis results, the device notifies the user of the authentication result. For example, it displays a message in the browser saying, "Authentication successful. Security lock has been released."
[0434] In this way, a voice authentication system is realized that allows the user to smoothly release the security lock and resume use of the credit card.
[0435] Example 1
[0436] Next, a description will be given of Example 1. In the following description, the data processing device 12 will be referred to as a "server" and the smart glasses 214 will be referred to as a "terminal."
[0437] Conventional credit card security lock release systems require users to manually perform authentication methods and procedures, which are time-consuming and laborious. Therefore, there was a need for a system that would allow users to quickly and reliably release security locks.
[0438] The specific processing by the specific processing unit 290 of the data processing device 12 in the first embodiment is realized by the following means.
[0439] In this invention, the server includes means for recording the user's voice, means for transmitting the recorded voice data to the server, analysis means having an AI voice authentication model for analyzing the received voice data, means for evaluating the authentication score of the analyzed voice data, means for releasing the security lock based on the evaluated authentication score, means for notifying the user of the result of the voice authentication, and means for transmitting the voice data via an HTTP POST request. This enables the user to easily, quickly, and reliably release the security lock using their voice.
[0440] "User" refers to an individual who uses the System to unlock a credit card security lock.
[0441] "Audio recording means" refers to a device or function for obtaining a user's microphone input and recording that audio data.
[0442] "Means for transmitting audio data to a server" refers to a device or function for transferring recorded audio data to a server using an appropriate protocol.
[0443] "AI voice authentication model" refers to a machine learning algorithm or program that analyzes voice data and performs specific identity verification.
[0444] "Analysis means" refers to a device or function that processes received voice data using an AI voice authentication model and calculates an authentication score.
[0445] "Means for assessing the authentication score" refers to a device or function for determining the calculated authentication score based on the analyzed voice data and determining the next processing step accordingly.
[0446] "Means for unlocking security lock" refers to a device or function for unlocking credit card usage restrictions when the authentication score exceeds a threshold value.
[0447] "Means for notifying the result of voice authentication" refers to a device or function for notifying the user of the result of voice authentication.
[0448] "Means for sending via HTTP POST request" refers to a device or function for sending recorded audio data to a server using the POST method of the HTTP protocol.
[0449] "Interface" refers to a graphical user interface that provides a means for a user to start and stop recording.
[0450] The present invention relates to a system for quickly and reliably releasing a security lock to prevent fraudulent use of a credit card through voice authentication. The system is specifically implemented using the following hardware and software.
[0451] Hardware and Software Used
[0452] 1. User voice recordings
[0453] The user accesses a dedicated web page and operates the "Start Recording" and "Stop Recording" buttons in the browser. The browser uses JavaScript's navigator.mediaDevices.getUserMedia API to obtain the user's microphone input. The audio is recorded using the MediaRecorder object.
[0454] 2. Sending audio data to the server
[0455] After finishing recording, the device converts the recorded audio data into Blob format and sends it to the server via an HTTP POST request using the FormData object. This communication uses the JavaScript fetch API.
[0456] 3. Receiving and storing audio data by the server
[0457] The server receives the audio data via an HTTP POST request and temporarily stores it in a file system managed, for example, using the Node.js fs module.
[0458] 4. Voice Authentication Process
[0459] The server inputs the saved audio file into an AI voice recognition model. This model is pre-trained and compares the user's pre-registered voice data with the target voice data to calculate a recognition score. The AI model is run using libraries such as TensorFlow and PyTorch.
[0460] 5. Unlocking the security lock
[0461] If the authentication score calculated based on the voice data exceeds the threshold, the server executes the procedure to release the credit card security lock. In this step, the API of the card company is called and the necessary procedures are followed to remove the usage restrictions.
[0462] 6. Notification of authentication results
[0463] If authentication is successful, the server notifies the result in JSON format, and the browser receives the response and displays the result to the user.
[0464] Specific examples
[0465] For example, suppose a user is using a credit card and the card is suspended due to the card company's logic for detecting fraudulent activity. In this case, the security lock can be released by following the procedure below.
[0466] 1. Users access a dedicated web page and click the "Start Recording" button to record their own voice.
[0467] 2. When you finish recording, click the "Stop Recording" button and the audio data will be converted to Blob format and sent to the server.
[0468] 3. The server receives the audio data and temporarily stores it.
[0469] 4. The saved voice data is analyzed using an AI voice authentication model to calculate an authentication score.
[0470] 5. If the authentication score exceeds the threshold, the server executes the procedure to unlock the credit card security lock.
[0471] 6. Finally, the server notifies the user of the result, and the user confirms that the credit card can be used again.
[0472] This allows users to quickly and reliably release security locks using voice commands.
[0473] The flow of the identification process in the first embodiment will be described with reference to FIG.
[0474] Step 1:
[0475] The user accesses a dedicated web page and clicks the "Start Recording" button. The device uses the navigator.mediaDevices.getUserMedia API to obtain the user's microphone input, creates a MediaRecorder object, and begins recording audio. Specifically, the browser displays a dialog requesting permission to use the microphone. Once audio input is obtained, MediaRecorder starts and recording begins. The input at this time is the user's voice, and the output is the audio data being recorded.
[0476] Step 2:
[0477] When the user clicks the "Stop Recording" button, the device stops recording and obtains the recorded audio data in Blob format. Specifically, the MediaRecorder object stops recording data, and the resulting audio data is saved in Blob format on the device. The input at this time is the audio data being recorded, and the output is audio data in Blob format.
[0478] Step 3:
[0479] The device adds the recorded audio data in Blob format to a FormData object and sends it to the server via an HTTP POST request. This communication is performed using JavaScript's fetch API. Specifically, the audio data is added to the FormData object, and the fetch API sends an HTTP POST request to the server. The input at this time is audio data in Blob format, and the output is an HTTP POST request to the server.
[0480] Step 4:
[0481] The server receives an HTTP POST request containing audio data and temporarily stores the data in a directory. The data received on the server side is saved in a local temporary directory using, for example, the Node.js fs module. The input is the HTTP POST request sent to the server, and the output is the audio data saved in the temporary directory.
[0482] Step 5:
[0483] The server inputs the saved voice data into an AI voice recognition model and calculates a recognition score. The AI voice recognition model is a pre-trained model that compares the voice data to calculate a recognition score. Specifically, the AI model is run using libraries such as TensorFlow and PyTorch. The input is the voice data saved in a temporary directory, and the output is a recognition score.
[0484] Step 6:
[0485] The server evaluates the authentication score and, if the score exceeds the threshold, releases the credit card's security lock. The server then calls the card company's API and executes the procedure to release the usage restriction. Specifically, it sends an HTTP request to the card company's API endpoint and obtains confirmation of the release. The input is the authentication score, and the output is confirmation of the release from the card company.
[0486] Step 7:
[0487] The server notifies the user of the authentication result in a JSON format message. The device parses this message and displays the authentication result to the user. Specifically, the server generates a JSON format response and sends it to the device as an HTTP response. The device receives this response and displays the authentication result on the screen. The input at this time is the JSON message of the authentication result generated by the server, and the output is the authentication result displayed on the device.
[0488] (Application example 1)
[0489] Next, a description will be given of Application Example 1. In the following description, the data processing device 12 will be referred to as a "server" and the smart glasses 214 will be referred to as a "terminal."
[0490] Conventional credit card fraud prevention systems require complicated procedures to activate security locks, making it difficult to quickly unlock them. Furthermore, the interface and notification methods for users to authenticate using their own voice are inadequate, reducing user convenience. Given these circumstances, there is a demand for a system that can quickly and reliably unlock credit card security locks and instantly notify users of the results.
[0491] The specific processing by the specific processing unit 290 of the data processing device 12 in the application example 1 is realized by the following means.
[0492] In this invention, the server includes means for recording the user's voice, means for transmitting the recorded voice data to the server, analysis means having a voice authentication model for analyzing the received voice data, means for evaluating the authentication score of the analyzed voice data, means for releasing the security lock based on the evaluated authentication score, display means including operation buttons for recording and transmitting the voice, and means for notifying the user of the result. This allows the user to quickly release the security lock on the credit card using their own voice and immediately confirm the result.
[0493] "User" means an individual who uses the system to record voice and attempt to unlock a credit card security lock.
[0494] "Audio recording means" refers to a device or application for recording a user's voice.
[0495] The "means for transmitting recorded voice data to a server" refers to a communication means for uploading recorded voice data to a server via a network.
[0496] A "voice authentication model" is an algorithm or AI model that analyzes recorded voice data and authenticates whether it is the user's voice.
[0497] The "analysis means" is a device that includes hardware and software for analyzing the voice data sent to the server.
[0498] The "means for evaluating the authentication score" refers to a device or program for determining whether authentication is successful or unsuccessful based on the score calculated by the voice authentication model.
[0499] A "means for unlocking security lock" is a device or program that executes a procedure to unlock a credit card when the authentication score exceeds a reference value.
[0500] The "display means including operation buttons" refers to a screen on which buttons for the user to operate to start and stop voice recording are displayed.
[0501] The "notification means" is a device or program for notifying the user of the result of voice authentication.
[0502] The term "computer system" refers to a collection of hardware and software that includes the above-mentioned means and operates as a whole.
[0503] This invention is a security system that allows users to quickly and reliably release the security lock on their credit cards using their own voice. This system is realized by combining the procedures of voice recording, data transmission, voice authentication, and result notification. The details are described below.
[0504] server
[0505] The server first receives the voice data sent by the user. After receiving the voice data, the AI voice authentication model within the server analyzes the voice data and calculates an authentication score. In this process, a machine learning model (e.g., TensorFlow, PyTorch) is used to determine whether the specific voice data matches the voice data of a pre-registered user. If the authentication score exceeds a threshold, the server releases the credit card security lock based on its internal logic. The authentication result is then returned to the user in JSON format, allowing interest confirmation.
[0506] Terminal
[0507] The device (such as the user's smartphone or computer) provides a user interface and displays operation buttons (record start button and record stop button) that the user can use to record audio. It records audio in response to specific operations, converts the audio data into Blob format, and sends it to the server. It also displays the authentication result to the user based on the authentication result returned from the server.
[0508] User
[0509] When a security lock is activated, the user opens the provided browser-based application and records their voice. After recording, the voice data is sent to the server, after which the authentication result is notified. This process allows the user to smoothly unlock the security lock on their credit card and resume use.
[0510] Specific examples
[0511] A user on a business trip attempts to use their credit card in a new city, but the card is locked due to fraud prevention logic. In this case, the user launches the smartphone app, records their voice, and sends it to the server. The server receives the voice data and calculates an authentication score using an AI voice authentication model. If authentication is successful, the server releases the card's security lock and notifies the user of the result. The user receives the notification and can resume using their credit card while on a business trip.
[0512] Prompt Sentence Examples
[0513] User: "Please press the start recording button and say the following phrase: 'This audio will verify my identity.'"
[0514] Input prompt to AI model: "Please authenticate the voice data below to confirm whether this voice belongs to a registered user."
[0515] The flow of the specific processing in the application example 1 will be described with reference to FIG.
[0516] Step 1:
[0517] The user opens a security application or web page, which displays a "Start Recording" button to start recording and a "Stop Recording" button to stop recording. The input of this step is the user's operation, and the output is the display of the user interface.
[0518] Step 2:
[0519] When the user presses the "Start Recording" button, the device accesses the microphone using navigator.mediaDevices.getUserMedia and starts recording. The recorded audio is temporarily saved on the device. The inputs to this step are the user's operation and audio data, and the outputs are the recording status and the temporarily saved audio data.
[0520] Step 3:
[0521] When the user presses the "Stop Recording" button, the device stops recording using the MediaRecorder object and converts the audio data to Blob format, which is then prepared for sending to the server. The input to this step is the audio data being recorded, and the output is Blob format audio data.
[0522] Step 4:
[0523] The device adds the audio data in Blob format to a FormData object and sends it to the server using an HTTP POST request using the fetch API. The input of this step is the audio data in Blob format, and the output is the data sent to the server.
[0524] Step 5:
[0525] The server saves the audio data received via the HTTP POST request in a temporary directory. Here, it checks whether the received audio data has been received and saved correctly. The input of this step is the HTTP request, and the output is the audio file on the server.
[0526] Step 6:
[0527] The server analyzes the received voice data using an AI voice authentication model (e.g., TensorFlow, PyTorch) and calculates an authentication score. The voice authentication model compares the voice data with that of registered users to measure the degree of match. The input for this step is the received voice data, and the output is an authentication score.
[0528] Step 7:
[0529] The server evaluates the calculated authentication score to determine if it exceeds a threshold. If so, it executes a procedure to unlock the credit card security lock. The input to this step is the authentication score, and the output is the execution of the unlock.
[0530] Step 8:
[0531] The server returns the authentication result (success or failure) to the device as a JSON-formatted message. The input to this step is the result of the authentication process, and the output is a JSON-formatted message.
[0532] Step 9:
[0533] The terminal analyzes the authentication result received from the server and notifies the user through the user interface. The user can check on the screen whether the credit card can be used or not. The input of this step is the authentication result in JSON format, and the output is the message displayed to the user.
[0534] Furthermore, an emotion engine that estimates the user's emotion may be further combined. That is, the identification processing unit 290 may estimate the user's emotion using the emotion identification model 59, and perform identification processing using the user's emotion.
[0535] The system according to the present invention provides a means for quickly and reliably unlocking a credit card's security lock to prevent fraudulent use through voice authentication and emotion recognition. This system is implemented in the following steps.
[0536] User voice recordings
[0537] The user accesses a dedicated web page and clicks the "Start Recording" button. The device uses navigator.mediaDevices.getUserMedia to obtain the user's microphone input, creates a MediaRecorder object, and starts recording audio. When the user clicks the "Stop Recording" button, the device stops recording and obtains the recorded audio data in Blob format.
[0538] Sending voice data to the server
[0539] The recorded audio data is sent from the device to the server by sending the audio file to the server via an HTTP POST request via a FormData object.
[0540] Receiving and storing audio data on the server
[0541] The server receives an HTTP POST request for audio data, extracts the audio file from the request, and stores it in a temporary directory for later use in the authentication process.
[0542] Voice Authentication
[0543] The server inputs the saved audio file into the authenticate method of the AI voice authentication model. This model compares it with the voice data of pre-registered users and calculates an authentication score. In parallel, the emotion engine analyzes the user's emotions from the audio data and obtains the results.
[0544] Authentication score and sentiment rating
[0545] The server evaluates the authentication score returned by the voice authentication model and the emotion data obtained from the emotion engine. This evaluation includes calculating an overall score that combines the authentication score and the emotion data. For example, if the user expresses anger or tension, this will be taken into account as a factor that affects the overall score.
[0546] Unlocking the security lock
[0547] If the total score exceeds a certain threshold, the server will release the security lock based on internal logic, a process that includes updating security settings and removing card usage restrictions.
[0548] Notification of results
[0549] The server returns the authentication result to the user as a JSON message. The message contains information about whether the authentication was successful or not, as well as any additional information required. The user's device receives this message, parses it, and notifies the user of the result.
[0550] Explanation with concrete examples
[0551] Scenario: When a user is using their credit card and their card is suspended due to the card company's logic for detecting fraudulent activity, the user can use this system to safely and quickly release the security lock.
[0552] 1. User opens a web page:
[0553] The user can record their own voice by operating the recording start / stop button. After recording is complete, the device sends the voice data to the server.
[0554] 2. Receiving and storing audio data by the server:
[0555] The server receives the audio data and stores it in a temporary directory.
[0556] 3. Server performs voice authentication and emotion recognition:
[0557] The server analyzes the voice data using an AI voice recognition model to calculate a recognition score. At the same time, the emotion engine analyzes emotions from the voice data and obtains emotion data.
[0558] 4. Overall score evaluation and security unlock:
[0559] The server evaluates the authentication score and emotional data together, and unlocks the credit card security lock if the criteria are met.
[0560] 5. Notification of Results:
[0561] The server notifies the user of the authentication result, and the user can confirm the result and resume using the card.
[0562] In this way, the system is designed to combine voice authentication and emotion recognition to enable users to smoothly release security locks and quickly resume using their credit cards.
[0563] The processing flow will be explained below.
[0564] Step 1:
[0565] A user accesses a web page and clicks the "Start Recording" button. The device uses navigator.mediaDevices.getUserMedia to obtain the user's microphone input. At this time, the user is prompted for microphone access permission in the browser and must grant it.
[0566] Step 2:
[0567] The device starts recording by creating a MediaRecorder object, which records the audio stream from the microphone in real time and collects the audio data.
[0568] Step 3:
[0569] The user clicks the "Stop Recording" button. The device stops recording and the audio data captured by MediaRecorder is retrieved in Blob format. This aggregates the audio data into a single object.
[0570] Step 4:
[0571] The device adds the acquired blob-formatted audio data to a FormData object and sends it to the server using an HTTP POST request. A FormData object is a data structure for sending files or data to the server in multiple key-value formats.
[0572] Step 5:
[0573] The server receives the HTTP POST request, extracts the audio file from the request, stores it in a temporary directory, and later uses it in the authentication process.
[0574] Step 6:
[0575] The server inputs the saved audio file into the authenticate method of the AI voice authentication model, which uses an algorithm to calculate a similarity score with the voice of a pre-registered user.
[0576] Step 7:
[0577] The server also sends the voice data to the emotion engine, which analyzes the user's emotions. The emotion engine analyzes the voice characteristics and determines whether the user is expressing anger, joy, sadness, or other emotions.
[0578] Step 8:
[0579] The server combines the authentication score returned by the AI model with the emotion data obtained from the emotion engine to calculate an overall score, which is generated using a calculation algorithm based on the voice authentication score and emotion data.
[0580] Step 9:
[0581] If the total score exceeds a certain threshold, the server will release the security lock based on internal logic, a process that includes updating security settings and removing card usage restrictions.
[0582] Step 10:
[0583] The server returns the authentication result to the user as a JSON message, which includes whether the authentication was successful or not, as well as any additional information required.
[0584] Step 11:
[0585] The user's device receives and analyzes the JSON message returned from the server. Based on the analysis results, the device notifies the user of the authentication result. For example, it displays a message in the browser saying, "Authentication successful. Security lock has been released."
[0586] Explanation with concrete examples
[0587] If a user is using a credit card and their card is suspended due to the card company's logic for detecting fraudulent activity, the user can use this system to safely and quickly release the security lock.
[0588] example:
[0589] The user opens a web page and records their own voice by operating the recording start / stop button. After recording is complete, the device sends the voice data to the server. The server receives the voice data and saves it in a temporary directory. The server analyzes the voice data using an AI voice authentication model and calculates an authentication score. At the same time, the emotion engine analyzes emotions from the voice data and obtains emotional data. The server evaluates the authentication score and emotional data together, and if the criteria are met, the credit card security lock is released. Finally, the server notifies the user of the authentication result, who can confirm the result and resume using their card.
[0590] In this way, the system is designed to combine voice authentication and emotion recognition to enable users to smoothly release security locks and quickly resume using their credit cards.
[0591] Example 2
[0592] Next, a description will be given of Example 2. In the following description, the data processing device 12 will be referred to as a "server" and the smart glasses 214 will be referred to as a "terminal."
[0593] In conventional credit card fraud prevention systems, voice authentication alone does not provide sufficient security, making it difficult to quickly and safely release the security lock. In particular, the user's emotional state, such as tension or anger, can affect the authentication results, raising the risk of incorrect judgment.
[0594] The specific processing by the specific processing unit 290 of the data processing device 12 in the second embodiment is realized by the following means.
[0595] In this invention, the server includes means for recording the user's voice, means for transmitting the recorded voice data to the server, analysis means having a voice authentication model for analyzing the received voice data, analysis means having an emotion recognition engine for analyzing emotions from the voice data, means for evaluating the authentication score and emotion data of the analyzed voice data, means for releasing the security lock based on a total score of the evaluated authentication score and emotion data, and means for notifying the user of the authentication result. This enables a comprehensive evaluation based on a combination of voice authentication and emotion recognition, thereby enabling the security lock to be released more quickly and safely.
[0596] "User" refers to a person who attempts to use this system to unlock the security lock on a credit card.
[0597] "Means for recording audio" refers to the function for capturing the user's audio and saving it as digital data.
[0598] "Server" refers to a computer system that includes a central processing unit that receives and analyzes audio data.
[0599] "Audio data" refers to data that represents audio information recorded by a user in digital form.
[0600] "Voice authentication model" refers to an algorithm or program that analyzes voice data and calculates an authentication score based on the user's voice characteristics.
[0601] "Analysis means" refers to a program or process for evaluating audio data and other input data.
[0602] "Emotion recognition engine" refers to a program or algorithm that analyzes a user's emotional state from voice data and generates emotional data.
[0603] "Authentication score" refers to a numerical value calculated by a voice authentication model that indicates the degree to which a user's voice matches a pre-enrolled voice.
[0604] "Emotion data" refers to information about emotions determined from the user's voice as analyzed by an emotion recognition engine.
[0605] "Total score" refers to an evaluation value calculated by integrating the authentication score and emotional data.
[0606] "Means to unlock security lock" refers to the function for lifting restrictions on a user's credit card usage based on the overall score.
[0607] "Notification means" refers to the interface or program used to communicate authentication results and system status to users.
[0608] The system based on this invention combines voice authentication and emotion recognition to provide a means for quickly and reliably unlocking security locks to prevent fraudulent use of credit cards. This system uses the following hardware and software to process data and perform calculations.
[0609] Hardware and software used
[0610] Hardware:
[0611] Microphone: A device that inputs the user's voice.
[0612] Server: A central processing unit that receives, stores, and analyzes voice data.
[0613] software:
[0614] Web browsers: Browsers that use navigator.mediaDevices.getUserMedia to get microphone input.
[0615] Voice authentication model: An algorithm that analyzes voice data and calculates an authentication score.
[0616] Emotion recognition engine: A program that analyzes emotions from voice data.
[0617] Example of system operation
[0618] 1. User visits a web page:
[0619] The user opens a browser and accesses the specified web page.
[0620] For example, a user visits the URL https: / / securitylock.example.com.
[0621] 2. Audio Recording:
[0622] The user clicks the "Start Recording" button and follows the instructions provided to record their voice.
[0623] The device uses navigator.mediaDevices.getUserMedia to obtain the user's microphone input, creates a MediaRecorder object, and starts recording audio.
[0624] When the user clicks the "Stop Recording" button, the device stops recording and obtains the recorded audio data in Blob format.
[0625] 3. Sending audio data to the server:
[0626] The terminal sends the recorded audio data to the server via a FormData object.
[0627] Use the fetch method to send it to the server as an HTTP POST request.
[0628] 4. Receiving and storing audio data by the server:
[0629] The server receives the HTTP POST request and retrieves the audio data.
[0630] The acquired audio data is saved in a temporary directory.
[0631] 5. Voice Verification and Emotion Recognition:
[0632] The server inputs the saved audio file into an AI voice recognition model and calculates a recognition score.
[0633] At the same time, an emotion recognition engine is used to analyze the user's emotional state from the voice data and obtain emotion data.
[0634] 6. Overall score evaluation and security unlock:
[0635] The server calculates a total score based on the authentication score from the voice authentication model and the emotion data from the emotion recognition engine.
[0636] If the total score meets a certain standard, the server will unlock the credit card's security lock.
[0637] 7. Notification of Results:
[0638] The server returns the authentication result to the user as a JSON message.
[0639] The terminal analyzes the result message received from the server and notifies the user of the result.
[0640] Prompt Sentence Examples
[0641] Please send the audio data recorded by the user in the following format.
[0642] Format: Blob format
[0643] Sending method: HTTP POST
[0644] Object used: FormData
[0645] The system is designed to combine voice authentication and emotion recognition to seamlessly unlock credit card security locks, allowing users to quickly resume using their cards.
[0646] The flow of the identification process in the second embodiment will be described with reference to FIG.
[0647] Step 1:
[0648] A user visits a web page:
[0649] A user opens a browser and accesses a specified web page. For example, a user accesses the URL https: / / securitylock.example.com. The input is the user's action (accessing the URL) and the output is the displayed web page.
[0650] Step 2:
[0651] User starts audio recording:
[0652] The user clicks the "Start Recording" button on a web page. The device uses navigator.mediaDevices.getUserMedia to obtain the user's microphone input. Specifically, the browser asks the user for permission to use the microphone, and if permission is granted, the microphone stream is obtained. The input is the user's click operation and microphone input, and the output is the audio stream.
[0653] Step 3:
[0654] The device records audio:
[0655] The device creates a MediaRecorder object and starts recording audio. The data being recorded is temporarily stored in memory. The input is the microphone stream, and the output is a buffer of the data being recorded.
[0656] Step 4:
[0657] User stops recording:
[0658] The user clicks the "Stop Recording" button. This causes the device to stop recording and obtain the recorded audio data in Blob format. Specifically, the recorder.stop() method is called and the audio data is obtained in the ondataavailable event. The input is the user's click operation, and the output is audio data in Blob format.
[0659] Step 5:
[0660] The device sends the audio data to the server:
[0661] The device adds the acquired audio data to a FormData object and uses the fetch method to send it to the server as an HTTP POST request. Specifically, the following operations are performed: let formData = new FormData(); formData.append('audio', audioBlob, 'audio.wav'); fetch(' / upload', { method: 'POST', body: formData}); The input is audio data in Blob format, and the output is an HTTP request to the server.
[0662] Step 6:
[0663] The server receives the audio data:
[0664] The server receives the HTTP POST request and extracts the audio data. The extracted audio data is saved in a temporary directory. Specifically, the server-side code retrieves the audio file from req.file or similar and saves it using the fs.writeFile method. The input is the HTTP request, and the output is the saved audio file.
[0665] Step 7:
[0666] The server performs voice authentication and emotion recognition:
[0667] The server inputs the saved audio file into the AI voice authentication model and calculates an authentication score. For example, the code "let score = voiceAuthModel.authenticate(' / temp / audio.wav');" is executed. At the same time, the emotion engine analyzes the user's emotional state from the audio data and executes the operation "let emotionData = emotionEngine.analyze(' / temp / audio.wav');" to obtain emotion data. The input is the audio file, and the output is the authentication score and emotion data.
[0668] Step 8:
[0669] The server evaluates the overall score:
[0670] The server calculates the total score based on the authentication score from the voice authentication model and the emotion data. For example, it executes the operation let totalScore = calculateTotalScore(score, emotionData);. The input is the authentication score and emotion data, and the output is the total score.
[0671] Step 9:
[0672] Server removes security lock:
[0673] If the total score meets a certain criterion, the server unlocks the credit card's security lock. For example, the following operation is performed: if (totalScore > threshold) { unlockSecurityLock();}. The input is the total score, and the output is the unlock status of the credit card.
[0674] Step 10:
[0675] The server notifies the user of the authentication result:
[0676] The server returns the authentication result to the user as a JSON message. Specifically, the following operation is performed: res.json({ success: true, message: 'Security lock has been released'});. The input is the authentication result, and the output is the message sent to the user.
[0677] Step 11:
[0678] User sees the results:
[0679] The terminal analyzes the result message received from the server and notifies the user. For example, it performs the following operation: const result = await response.json(); alert(result.message);. The input is the message from the server, and the output is the notification to the user.
[0680] (Application example 2)
[0681] Next, a description will be given of Application Example 2. In the following description, the data processing device 12 will be referred to as a "server" and the smart glasses 214 will be referred to as a "terminal."
[0682] Conventional security systems can cause inconvenience to users due to delays in detecting unauthorized use and unlocking the device. Furthermore, simple voice authentication alone cannot take into account user conditions such as emotional changes and stress. As a result, legitimate users are unable to unlock the security lock, resulting in a poor user experience.
[0683] The identification process by the identification processing unit 290 of the data processing device 12 in Application Example 2 is realized by the following means. In this invention, the server includes means for recording the user's voice, means for transmitting the recorded voice data to the server, analysis means having a voice authentication model and an emotion recognition engine for analyzing the received voice data, means for evaluating the authentication score and emotion data of the analyzed voice data, and means for calculating a total score based on the evaluated authentication score and emotion data and for releasing the security lock using the total score. This allows the legitimate user to release the security lock quickly and safely.
[0684] A "user" is an individual who records voice and uses the recorded voice data.
[0685] "Audio data" means sound in the form of data recorded by a user and transmitted to a server.
[0686] A "server" is a computer system for analyzing received voice data and performing voice authentication and emotion recognition.
[0687] "Analysis means" refers to a set of functions for analyzing audio data using a voice authentication model and an emotion recognition engine.
[0688] The "voice authentication model" is a machine learning model that analyzes received voice data and calculates an authentication score.
[0689] An "emotion recognition engine" is software that analyzes and obtains a user's emotional data from voice data.
[0690] The "authentication score" is a score calculated by the voice authentication model that indicates the degree to which the user's voice matches pre-registered data.
[0691] "Emotion data" is data that indicates the emotional state of the user, analyzed from the voice data.
[0692] The "total score" is an evaluation score calculated by combining the authentication score of the voice authentication model and emotion data.
[0693] A "security lock" is a security measure that limits the use of credit cards.
[0694] "Unlocking means" refers to a set of functions for unlocking a security lock based on the overall score.
[0695] "Notification means" refers to a method or technical means for notifying the user of the results of voice authentication and emotion recognition.
[0696] This invention relates to a system that records a user's voice, transmits the voice data to a server for analysis, and quickly and safely releases a security lock through voice authentication and emotion recognition. This system is implemented with the following configuration and procedure.
[0697] Voice recording and data transmission
[0698] The device is equipped with a microphone for recording the user's voice, a function for acquiring voice input using navigator.mediaDevices.getUserMedia, and a function for recording and saving voice data using the MediaRecorder object. The user records voice through an interface for starting and stopping recording, and sends the voice data to the server after recording is complete. To send data to the server, the voice data is added to a FormData object and sent to the server using an HTTP POST request.
[0699] Receiving and analyzing audio data
[0700] The server has the ability to store the received voice data in a temporary directory and analyze it. A voice authentication model and an emotion recognition engine are used for the analysis. The voice authentication model analyzes the user's voice data and calculates an authentication score that indicates the degree of match with pre-registered data. Meanwhile, the emotion recognition engine extracts the user's emotion data from the voice data.
[0701] Authentication score and emotion data evaluation
[0702] The server evaluates the authentication score of the voice authentication model and the emotion data obtained from the emotion recognition engine. If the overall score exceeds a certain threshold, the security lock is released.
[0703] Notification of results
[0704] The server returns the results of voice authentication and emotion recognition to the user as a JSON-formatted message. The message includes the result of authentication (success or failure) and any additional information required. The device receives this message, analyzes it, and notifies the user of the result.
[0705] Specific examples
[0706] For example, if a user tries to use their credit card but finds it is locked, they can launch the "Safe Payment Release App" on their smartphone and press the "Start Recording" button to record their voice. The recorded data is automatically sent to a server, where voice authentication and emotion recognition are performed. If the security lock can be released, the user is notified of the result and can use their credit card again.
[0707] Prompt Sentence Examples
[0708] "Please use this user's voice recording data to perform voice authentication and emotion recognition to determine whether the security lock can be released. Please use the following audio file for voice authentication and emotion recognition."
[0709] The system allows the server and terminal to work together to provide users with a quick and reliable security unlocking experience.
[0710] The flow of the specific processing in the application example 2 will be described with reference to FIG.
[0711] Step 1:
[0712] The user launches the "Safe Payment Cancellation App" on their smartphone and presses the "Start Recording" button. The device obtains microphone input using navigator.mediaDevices.getUserMedia. The device then creates a MediaRecorder object and begins recording audio. The input here is the user's voice, and the output is audio data obtained from the microphone. Specifically, when the user presses the button, the device begins microphone input and collects audio data in real time.
[0713] Step 2:
[0714] When the user presses the "Stop Recording" button, the device stops recording and obtains the recorded audio data in Blob format. The input here is the user operation that triggers the end of recording, and the output is audio data in Blob format. Specifically, when the user presses the button, the device stops MediaRecorder and converts the audio data stored in the internal buffer into Blob format.
[0715] Step 3:
[0716] The device adds this audio data in Blob format to a FormData object and sends it to the server using an HTTP POST request. The input is the audio data in Blob format, and the output is a FormData object sent to the server. Specifically, the device adds the audio data to FormData and issues an HTTP POST request to the endpoint to the server.
[0717] Step 4:
[0718] The server saves the received audio data in a temporary directory. The input is the audio data sent in the HTTP POST request, and the output is an audio file saved in the temporary directory. Specifically, the server receives the request, extracts the audio data, and saves it in a temporary directory.
[0719] Step 5:
[0720] The server inputs the saved audio file into the authenticate method of the voice authentication model and calculates an authentication score. The emotion recognition engine also analyzes and obtains the user's emotional data from the audio data. The input is the audio file, and the output is an authentication score and emotional data. Specifically, the server calls the voice authentication model and emotion recognition engine to perform the analysis.
[0721] Step 6:
[0722] The server calculates an overall score by combining the authentication score from the voice authentication model and the emotion data from the emotion recognition engine, and releases the security lock if the overall score exceeds a threshold. The input is the authentication score and emotion data, and the output is an instruction to release the security lock. Specifically, the server evaluates the score and updates the credit card security settings as necessary.
[0723] Step 7:
[0724] The server returns the results of voice authentication and emotion recognition to the user as a JSON-formatted message. The input is the overall score and evaluation result, and the output is a JSON-formatted result message sent to the user. Specifically, the server formats the results and sends an HTTP response to the user's device.
[0725] Step 8:
[0726] The user's device parses the JSON-formatted message received from the server and notifies the user of the results. The input is a JSON-formatted message, and the output is a notification to the user. Specifically, the device parses the message and displays the results to the user through an alert or notification.
[0727] The above are the specific processing steps of this system.
[0728] The specific processing unit 290 transmits the result of the specific processing to the smart glasses 214. In the smart glasses 214, the control unit 46A causes the speaker 240 to output the result of the specific processing. The microphone 238 acquires audio indicating a user input regarding the result of the specific processing. The control unit 46A transmits audio data indicating the user input acquired by the microphone 238 to the data processing device 12. In the data processing device 12, the specific processing unit 290 acquires the audio data.
[0729] The data generation model 58 is a so-called generative AI (Artificial Intelligence). An example of the data generation model 58 is ChatGPT (Internet Search<URL: https: / / openai.com / blog / chatgpt> ), Gemini (Internet search <url: https: gemini.google.com ?hl="ja">) and other generation AIs. The data generation model 58 is obtained by performing deep learning on a neural network. A prompt including an instruction is input to the data generation model 58, and inference data such as voice data indicating voice, text data indicating text, and image data indicating an image is also input. The data generation model 58 performs inference on the input inference data in accordance with the instruction indicated by the prompt, and outputs the inference result in a data format such as voice data and text data. Here, inference refers to, for example, analysis, classification, prediction, and / or summarization.
[0730] In the above embodiment, an example in which the specific processing is performed by the data processing device 12 has been given, but the technology of the present disclosure is not limited to this, and the specific processing may be performed by the smart glasses 214.
[0731] [Third embodiment]
[0732] FIG. 5 shows an example of the configuration of a data processing system 310 according to the third embodiment.
[0733] 5, the data processing system 310 includes the data processing device 12 and a headset type terminal 314. An example of the data processing device 12 is a server.
[0734] The data processing device 12 includes a computer 22, a database 24, and a communication I / F 26. The computer 22 is an example of a "computer" according to the technology of the present disclosure. The computer 22 includes a processor 28, a RAM 30, and a storage 32. The processor 28, the RAM 30, and the storage 32 are connected to a bus 34. The database 24 and the communication I / F 26 are also connected to the bus 34. The communication I / F 26 is connected to a network 54. Examples of the network 54 include a WAN (Wide Area Network) and / or a LAN (Local Area Network).
[0735] The headset type terminal 314 includes a computer 36, a microphone 238, a speaker 240, a camera 42, a communication I / F 44, and a display 343. The computer 36 includes a processor 46, a RAM 48, and a storage 50. The processor 46, the RAM 48, and the storage 50 are connected to a bus 52. The microphone 238, the speaker 240, the camera 42, and the display 343 are also connected to the bus 52.
[0736] The microphone 238 receives instructions and the like from the user 20 by receiving voice uttered by the user 20. The microphone 238 captures the voice uttered by the user 20, converts the captured voice into audio data, and outputs it to the processor 46. The speaker 240 outputs audio in accordance with instructions from the processor 46.
[0737] Camera 42 is a small digital camera equipped with an optical system including a lens, aperture, and shutter, and an imaging element such as a CMOS (Complementary Metal-Oxide-Semiconductor) image sensor or a CCD (Charge Coupled Device) image sensor, and captures images of the surroundings of user 20 (for example, an imaging range defined by an angle of view equivalent to the field of vision of a typical healthy person).
[0738] The communication I / F 44 is connected to a network 54. The communication I / Fs 44 and 26 are responsible for the exchange of various information between the processor 46 and the processor 28 via the network 54. The exchange of various information between the processor 46 and the processor 28 using the communication I / Fs 44 and 26 is carried out in a secure state.
[0739] Fig. 6 shows an example of the main functions of the data processing device 12 and the headset type terminal 314. As shown in Fig. 6, in the data processing device 12, a specific process is performed by the processor 28. A specific process program 56 is stored in the storage 32.
[0740] The specific processing program 56 is an example of a "program" according to the technology of the present disclosure. The processor 28 reads the specific processing program 56 from the storage 32 and executes the read specific processing program 56 on the RAM 30. The specific processing is realized by the processor 28 operating as a specific processing unit 290 in accordance with the specific processing program 56 executed on the RAM 30.
[0741] The storage 32 stores a data generation model 58 and an emotion identification model 59. The data generation model 58 and the emotion identification model 59 are used by the identification processing unit 290.
[0742] In the headset type terminal 314, a reception output process is performed by the processor 46. A reception output program 60 is stored in the storage 50. The processor 46 reads the reception output program 60 from the storage 50 and executes the read reception output program 60 on the RAM 48. The reception output process is realized by the processor 46 operating as the control unit 46A in accordance with the reception output program 60 executed on the RAM 48.
[0743] Next, a description will be given of the identification process performed by the identification processing unit 290 of the data processing device 12. In the following description, the data processing device 12 will be referred to as the "server" and the headset type terminal 314 will be referred to as the "terminal."
[0744] The system according to the present invention provides a means for quickly and reliably releasing a security lock on a credit card to prevent fraudulent use through voice authentication. This system is implemented in the following procedure.
[0745] User voice recordings
[0746] The user accesses a dedicated web page and clicks the "Start Recording" button. The device obtains the user's microphone input using navigator.mediaDevices.getUserMedia, creates a MediaRecorder object, and starts recording audio. When the user clicks the "Stop Recording" button, the device stops recording and obtains the recorded audio data in Blob format.
[0747] Sending voice data to the server
[0748] The recorded audio data is sent from the device to the server by sending the audio file to the server via an HTTP POST request via a FormData object.
[0749] Receiving and storing audio data on the server
[0750] The server receives the HTTP POST request for the audio data and stores it in a temporary directory, preparing it for the subsequent audio authentication process.
[0751] Voice Authentication
[0752] The server uses an AI voice authentication model to obtain an authentication score based on the saved voice file. This AI model compares the voice data of pre-registered users to calculate an authentication score. The server evaluates this authentication score, and if the score exceeds a certain threshold, the user's security lock is released.
[0753] Unlocking the security lock
[0754] If the score exceeds the threshold, the server will unlock the credit card security lock based on its internal processing logic, and will also remove any restrictions on card usage after going through the necessary procedures.
[0755] Notification of results
[0756] The server notifies the user of the authentication result. The notification is sent back to the terminal in JSON format, and the terminal parses this message and displays the result to the user. The user can check the voice authentication result and find out whether they can resume using their card.
[0757] Explanation with concrete examples
[0758] Scenario: When a user is using their credit card and their card is suspended due to the card company's logic for detecting fraudulent activity, the user can use this system to safely and quickly release the security lock.
[0759] 1. User opens a web page:
[0760] Users can record their own voice by operating the recording start / stop button.
[0761] 2. Sending audio data to the server:
[0762] The device converts the recorded data into Blob format and sends it to the server.
[0763] 3. Server performs voice authentication:
[0764] The server receives the voice data and uses an AI model to calculate an authentication score.
[0765] 4. Security unlock and notifications:
[0766] - The server determines that the authentication score criteria have been met and releases the security lock. The user is then shown a message indicating successful authentication, allowing them to resume using their credit card.
[0767] In this way, the system is designed to enable users to smoothly release the security lock and quickly resume using their credit card.
[0768] The processing flow will be explained below.
[0769] Step 1:
[0770] A user accesses a web page and clicks the "Start Recording" button. The device uses navigator.mediaDevices.getUserMedia to obtain the user's microphone input. At this time, the user is prompted for microphone access permission in the browser and must grant it.
[0771] Step 2:
[0772] The device starts recording by creating a MediaRecorder object, which records the audio stream from the microphone in real time and collects the audio data.
[0773] Step 3:
[0774] The user clicks the "Stop Recording" button. The device stops recording and the audio data captured by MediaRecorder is retrieved in Blob format. This aggregates the audio data into a single object.
[0775] Step 4:
[0776] The device adds the acquired blob-formatted audio data to a FormData object and sends it to the server using an HTTP POST request. A FormData object is a data structure for sending files or data to the server in multiple key-value formats.
[0777] Step 5:
[0778] The server receives the HTTP POST request, extracts the audio file from the request, stores it in a temporary directory, and later uses it in the authentication process.
[0779] Step 6:
[0780] The server inputs the saved audio file into the authenticate method of the AI voice authentication model, which uses an algorithm to calculate a similarity score with the voice of a pre-registered user.
[0781] Step 7:
[0782] The server evaluates the authentication score returned by the AI model, which includes determining whether the score exceeds a certain threshold, which is set according to the security level.
[0783] Step 8:
[0784] If the score exceeds a threshold, the server will release the security lock based on internal logic, a process that includes updating security settings and removing card usage restrictions.
[0785] Step 9:
[0786] The server returns the authentication result to the user as a JSON message, which includes whether the authentication was successful or not, as well as any additional information required.
[0787] Step 10:
[0788] The user's device receives and analyzes the JSON message returned from the server. Based on the analysis results, the device notifies the user of the authentication result. For example, it displays a message in the browser saying, "Authentication successful. Security lock has been released."
[0789] In this way, a voice authentication system is realized that allows the user to smoothly release the security lock and resume use of the credit card.
[0790] Example 1
[0791] Next, a description will be given of Example 1. In the following description, the data processing device 12 will be referred to as a "server" and the headset type terminal 314 will be referred to as a "terminal."
[0792] Conventional credit card security lock release systems require users to manually perform authentication methods and procedures, which are time-consuming and laborious. Therefore, there was a need for a system that would allow users to quickly and reliably release security locks.
[0793] The specific processing by the specific processing unit 290 of the data processing device 12 in the first embodiment is realized by the following means.
[0794] In this invention, the server includes means for recording the user's voice, means for transmitting the recorded voice data to the server, analysis means having an AI voice authentication model for analyzing the received voice data, means for evaluating the authentication score of the analyzed voice data, means for releasing the security lock based on the evaluated authentication score, means for notifying the user of the result of the voice authentication, and means for transmitting the voice data via an HTTP POST request. This enables the user to easily, quickly, and reliably release the security lock using their voice.
[0795] "User" refers to an individual who uses the System to unlock a credit card security lock.
[0796] "Audio recording means" refers to a device or function for obtaining a user's microphone input and recording that audio data.
[0797] "Means for transmitting audio data to a server" refers to a device or function for transferring recorded audio data to a server using an appropriate protocol.
[0798] "AI voice authentication model" refers to a machine learning algorithm or program that analyzes voice data and performs specific identity verification.
[0799] "Analysis means" refers to a device or function that processes received voice data using an AI voice authentication model and calculates an authentication score.
[0800] "Means for assessing the authentication score" refers to a device or function for determining the calculated authentication score based on the analyzed voice data and determining the next processing step accordingly.
[0801] "Means for unlocking security lock" refers to a device or function for unlocking credit card usage restrictions when the authentication score exceeds a threshold value.
[0802] "Means for notifying the result of voice authentication" refers to a device or function for notifying the user of the result of voice authentication.
[0803] "Means for sending via HTTP POST request" refers to a device or function for sending recorded audio data to a server using the POST method of the HTTP protocol.
[0804] "Interface" refers to a graphical user interface that provides a means for a user to start and stop recording.
[0805] The present invention relates to a system for quickly and reliably releasing a security lock to prevent fraudulent use of a credit card through voice authentication. The system is specifically implemented using the following hardware and software.
[0806] Hardware and Software Used
[0807] 1. User voice recordings
[0808] The user accesses a dedicated web page and operates the "Start Recording" and "Stop Recording" buttons in the browser. The browser uses JavaScript's navigator.mediaDevices.getUserMedia API to obtain the user's microphone input. The audio is recorded using the MediaRecorder object.
[0809] 2. Sending audio data to the server
[0810] After finishing recording, the device converts the recorded audio data into Blob format and sends it to the server via an HTTP POST request using the FormData object. This communication uses the JavaScript fetch API.
[0811] 3. Receiving and storing audio data by the server
[0812] The server receives the audio data via an HTTP POST request and temporarily stores it in a file system managed, for example, using the Node.js fs module.
[0813] 4. Voice Authentication Process
[0814] The server inputs the saved audio file into an AI voice recognition model. This model is pre-trained and compares the user's pre-registered voice data with the target voice data to calculate a recognition score. The AI model is run using libraries such as TensorFlow and PyTorch.
[0815] 5. Unlocking the security lock
[0816] If the authentication score calculated based on the voice data exceeds the threshold, the server executes the procedure to release the credit card security lock. In this step, the API of the card company is called and the necessary procedures are followed to remove the usage restrictions.
[0817] 6. Notification of authentication results
[0818] If authentication is successful, the server notifies the result in JSON format, and the browser receives the response and displays the result to the user.
[0819] Specific examples
[0820] For example, suppose a user is using a credit card and the card is suspended due to the card company's logic for detecting fraudulent activity. In this case, the security lock can be released by following the procedure below.
[0821] 1. Users access a dedicated web page and click the "Start Recording" button to record their own voice.
[0822] 2. When you finish recording, click the "Stop Recording" button and the audio data will be converted to Blob format and sent to the server.
[0823] 3. The server receives the audio data and temporarily stores it.
[0824] 4. The saved voice data is analyzed using an AI voice authentication model to calculate an authentication score.
[0825] 5. If the authentication score exceeds the threshold, the server executes the procedure to unlock the credit card security lock.
[0826] 6. Finally, the server notifies the user of the result, and the user confirms that the credit card can be used again.
[0827] This allows users to quickly and reliably release security locks using voice commands.
[0828] The flow of the identification process in the first embodiment will be described with reference to FIG.
[0829] Step 1:
[0830] The user accesses a dedicated web page and clicks the "Start Recording" button. The device uses the navigator.mediaDevices.getUserMedia API to obtain the user's microphone input, creates a MediaRecorder object, and begins recording audio. Specifically, the browser displays a dialog requesting permission to use the microphone. Once audio input is obtained, MediaRecorder starts and recording begins. The input at this time is the user's voice, and the output is the audio data being recorded.
[0831] Step 2:
[0832] When the user clicks the "Stop Recording" button, the device stops recording and obtains the recorded audio data in Blob format. Specifically, the MediaRecorder object stops recording data, and the resulting audio data is saved in Blob format on the device. The input at this time is the audio data being recorded, and the output is audio data in Blob format.
[0833] Step 3:
[0834] The device adds the recorded audio data in Blob format to a FormData object and sends it to the server via an HTTP POST request. This communication is performed using JavaScript's fetch API. Specifically, the audio data is added to the FormData object, and the fetch API sends an HTTP POST request to the server. The input at this time is audio data in Blob format, and the output is an HTTP POST request to the server.
[0835] Step 4:
[0836] The server receives an HTTP POST request containing audio data and temporarily stores the data in a directory. The data received on the server side is saved in a local temporary directory using, for example, the Node.js fs module. The input is the HTTP POST request sent to the server, and the output is the audio data saved in the temporary directory.
[0837] Step 5:
[0838] The server inputs the saved voice data into an AI voice recognition model and calculates a recognition score. The AI voice recognition model is a pre-trained model that compares the voice data to calculate a recognition score. Specifically, the AI model is run using libraries such as TensorFlow and PyTorch. The input is the voice data saved in a temporary directory, and the output is a recognition score.
[0839] Step 6:
[0840] The server evaluates the authentication score and, if the score exceeds the threshold, releases the credit card's security lock. The server then calls the card company's API and executes the procedure to release the usage restriction. Specifically, it sends an HTTP request to the card company's API endpoint and obtains confirmation of the release. The input is the authentication score, and the output is confirmation of the release from the card company.
[0841] Step 7:
[0842] The server notifies the user of the authentication result in a JSON format message. The device parses this message and displays the authentication result to the user. Specifically, the server generates a JSON format response and sends it to the device as an HTTP response. The device receives this response and displays the authentication result on the screen. The input at this time is the JSON message of the authentication result generated by the server, and the output is the authentication result displayed on the device.
[0843] (Application example 1)
[0844] Next, a description will be given of Application Example 1. In the following description, the data processing device 12 will be referred to as a "server" and the headset type terminal 314 will be referred to as a "terminal."
[0845] Conventional credit card fraud prevention systems require complicated procedures to activate security locks, making it difficult to quickly unlock them. Furthermore, the interface and notification methods for users to authenticate using their own voice are inadequate, reducing user convenience. Given these circumstances, there is a demand for a system that can quickly and reliably unlock credit card security locks and instantly notify users of the results.
[0846] The specific processing by the specific processing unit 290 of the data processing device 12 in the application example 1 is realized by the following means.
[0847] In this invention, the server includes means for recording the user's voice, means for transmitting the recorded voice data to the server, analysis means having a voice authentication model for analyzing the received voice data, means for evaluating the authentication score of the analyzed voice data, means for releasing the security lock based on the evaluated authentication score, display means including operation buttons for recording and transmitting the voice, and means for notifying the user of the result. This allows the user to quickly release the security lock on the credit card using their own voice and immediately confirm the result.
[0848] "User" means an individual who uses the system to record voice and attempt to unlock a credit card security lock.
[0849] "Audio recording means" refers to a device or application for recording a user's voice.
[0850] The "means for transmitting recorded voice data to a server" refers to a communication means for uploading recorded voice data to a server via a network.
[0851] A "voice authentication model" is an algorithm or AI model that analyzes recorded voice data and authenticates whether it is the user's voice.
[0852] The "analysis means" is a device that includes hardware and software for analyzing the voice data sent to the server.
[0853] The "means for evaluating the authentication score" refers to a device or program for determining whether authentication is successful or unsuccessful based on the score calculated by the voice authentication model.
[0854] A "means for unlocking security lock" is a device or program that executes a procedure to unlock a credit card when the authentication score exceeds a reference value.
[0855] The "display means including operation buttons" refers to a screen on which buttons for the user to operate to start and stop voice recording are displayed.
[0856] The "notification means" is a device or program for notifying the user of the result of voice authentication.
[0857] The term "computer system" refers to a collection of hardware and software that includes the above-mentioned means and operates as a whole.
[0858] This invention is a security system that allows users to quickly and reliably release the security lock on their credit cards using their own voice. This system is realized by combining the procedures of voice recording, data transmission, voice authentication, and result notification. The details are described below.
[0859] server
[0860] The server first receives the voice data sent by the user. After receiving the voice data, the AI voice authentication model within the server analyzes the voice data and calculates an authentication score. In this process, a machine learning model (e.g., TensorFlow, PyTorch) is used to determine whether the specific voice data matches the voice data of a pre-registered user. If the authentication score exceeds a threshold, the server releases the credit card security lock based on its internal logic. The authentication result is then returned to the user in JSON format, allowing interest confirmation.
[0861] Terminal
[0862] The device (such as the user's smartphone or computer) provides a user interface and displays operation buttons (record start button and record stop button) that the user can use to record audio. It records audio in response to specific operations, converts the audio data into Blob format, and sends it to the server. It also displays the authentication result to the user based on the authentication result returned from the server.
[0863] User
[0864] When a security lock is activated, the user opens the provided browser-based application and records their voice. After recording, the voice data is sent to the server, after which the authentication result is notified. This process allows the user to smoothly unlock the security lock on their credit card and resume use.
[0865] Specific examples
[0866] A user on a business trip attempts to use their credit card in a new city, but the card is locked due to fraud prevention logic. In this case, the user launches the smartphone app, records their voice, and sends it to the server. The server receives the voice data and calculates an authentication score using an AI voice authentication model. If authentication is successful, the server releases the card's security lock and notifies the user of the result. The user receives the notification and can resume using their credit card while on a business trip.
[0867] Prompt Sentence Examples
[0868] User: "Please press the start recording button and say the following phrase: 'This audio will verify my identity.'"
[0869] Input prompt to AI model: "Please authenticate the voice data below to confirm whether this voice belongs to a registered user."
[0870] The flow of the specific processing in the application example 1 will be described with reference to FIG.
[0871] Step 1:
[0872] The user opens a security application or web page, which displays a "Start Recording" button to start recording and a "Stop Recording" button to stop recording. The input of this step is the user's operation, and the output is the display of the user interface.
[0873] Step 2:
[0874] When the user presses the "Start Recording" button, the device accesses the microphone using navigator.mediaDevices.getUserMedia and starts recording. The recorded audio is temporarily saved on the device. The inputs to this step are the user's operation and audio data, and the outputs are the recording status and the temporarily saved audio data.
[0875] Step 3:
[0876] When the user presses the "Stop Recording" button, the device stops recording using the MediaRecorder object and converts the audio data to Blob format, which is then prepared for sending to the server. The input to this step is the audio data being recorded, and the output is Blob format audio data.
[0877] Step 4:
[0878] The device adds the audio data in Blob format to a FormData object and sends it to the server using an HTTP POST request using the fetch API. The input of this step is the audio data in Blob format, and the output is the data sent to the server.
[0879] Step 5:
[0880] The server saves the audio data received via the HTTP POST request in a temporary directory. Here, it checks whether the received audio data has been received and saved correctly. The input of this step is the HTTP request, and the output is the audio file on the server.
[0881] Step 6:
[0882] The server analyzes the received voice data using an AI voice authentication model (e.g., TensorFlow, PyTorch) and calculates an authentication score. The voice authentication model compares the voice data with that of registered users to measure the degree of match. The input for this step is the received voice data, and the output is an authentication score.
[0883] Step 7:
[0884] The server evaluates the calculated authentication score to determine if it exceeds a threshold. If so, it executes a procedure to unlock the credit card security lock. The input to this step is the authentication score, and the output is the execution of the unlock.
[0885] Step 8:
[0886] The server returns the authentication result (success or failure) to the device as a JSON-formatted message. The input to this step is the result of the authentication process, and the output is a JSON-formatted message.
[0887] Step 9:
[0888] The terminal analyzes the authentication result received from the server and notifies the user through the user interface. The user can check on the screen whether the credit card can be used or not. The input of this step is the authentication result in JSON format, and the output is the message displayed to the user.
[0889] Furthermore, an emotion engine that estimates the user's emotion may be further combined. That is, the identification processing unit 290 may estimate the user's emotion using the emotion identification model 59, and perform identification processing using the user's emotion.
[0890] The system according to the present invention provides a means for quickly and reliably unlocking a credit card's security lock to prevent fraudulent use through voice authentication and emotion recognition. This system is implemented in the following steps.
[0891] User voice recordings
[0892] The user accesses a dedicated web page and clicks the "Start Recording" button. The device uses navigator.mediaDevices.getUserMedia to obtain the user's microphone input, creates a MediaRecorder object, and starts recording audio. When the user clicks the "Stop Recording" button, the device stops recording and obtains the recorded audio data in Blob format.
[0893] Sending voice data to the server
[0894] The recorded audio data is sent from the device to the server by sending the audio file to the server via an HTTP POST request via a FormData object.
[0895] Receiving and storing audio data on the server
[0896] The server receives an HTTP POST request for audio data, extracts the audio file from the request, and stores it in a temporary directory for later use in the authentication process.
[0897] Voice Authentication
[0898] The server inputs the saved audio file into the authenticate method of the AI voice authentication model. This model compares it with the voice data of pre-registered users and calculates an authentication score. In parallel, the emotion engine analyzes the user's emotions from the audio data and obtains the results.
[0899] Authentication score and sentiment rating
[0900] The server evaluates the authentication score returned by the voice authentication model and the emotion data obtained from the emotion engine. This evaluation includes calculating an overall score that combines the authentication score and the emotion data. For example, if the user expresses anger or tension, this will be taken into account as a factor that affects the overall score.
[0901] Unlocking the security lock
[0902] If the total score exceeds a certain threshold, the server will release the security lock based on internal logic, a process that includes updating security settings and removing card usage restrictions.
[0903] Notification of results
[0904] The server returns the authentication result to the user as a JSON message. The message contains information about whether the authentication was successful or not, as well as any additional information required. The user's device receives this message, parses it, and notifies the user of the result.
[0905] Explanation with concrete examples
[0906] Scenario: When a user is using their credit card and their card is suspended due to the card company's logic for detecting fraudulent activity, the user can use this system to safely and quickly release the security lock.
[0907] 1. User opens a web page:
[0908] The user can record their own voice by operating the recording start / stop button. After recording is complete, the device sends the voice data to the server.
[0909] 2. Receiving and storing audio data by the server:
[0910] The server receives the audio data and stores it in a temporary directory.
[0911] 3. Server performs voice authentication and emotion recognition:
[0912] The server analyzes the voice data using an AI voice recognition model to calculate a recognition score. At the same time, the emotion engine analyzes emotions from the voice data and obtains emotion data.
[0913] 4. Overall score evaluation and security unlock:
[0914] The server evaluates the authentication score and emotional data together, and unlocks the credit card security lock if the criteria are met.
[0915] 5. Notification of Results:
[0916] The server notifies the user of the authentication result, and the user can confirm the result and resume using the card.
[0917] In this way, the system is designed to combine voice authentication and emotion recognition to enable users to smoothly release security locks and quickly resume using their credit cards.
[0918] The processing flow will be explained below.
[0919] Step 1:
[0920] A user accesses a web page and clicks the "Start Recording" button. The device uses navigator.mediaDevices.getUserMedia to obtain the user's microphone input. At this time, the user is prompted for microphone access permission in the browser and must grant it.
[0921] Step 2:
[0922] The device starts recording by creating a MediaRecorder object, which records the audio stream from the microphone in real time and collects the audio data.
[0923] Step 3:
[0924] The user clicks the "Stop Recording" button. The device stops recording and the audio data captured by MediaRecorder is retrieved in Blob format. This aggregates the audio data into a single object.
[0925] Step 4:
[0926] The device adds the acquired blob-formatted audio data to a FormData object and sends it to the server using an HTTP POST request. A FormData object is a data structure for sending files or data to the server in multiple key-value formats.
[0927] Step 5:
[0928] The server receives the HTTP POST request, extracts the audio file from the request, stores it in a temporary directory, and later uses it in the authentication process.
[0929] Step 6:
[0930] The server inputs the saved audio file into the authenticate method of the AI voice authentication model, which uses an algorithm to calculate a similarity score with the voice of a pre-registered user.
[0931] Step 7:
[0932] The server also sends the voice data to the emotion engine, which analyzes the user's emotions. The emotion engine analyzes the voice characteristics and determines whether the user is expressing anger, joy, sadness, or other emotions.
[0933] Step 8:
[0934] The server combines the authentication score returned by the AI model with the emotion data obtained from the emotion engine to calculate an overall score, which is generated using a calculation algorithm based on the voice authentication score and emotion data.
[0935] Step 9:
[0936] If the total score exceeds a certain threshold, the server will release the security lock based on internal logic, a process that includes updating security settings and removing card usage restrictions.
[0937] Step 10:
[0938] The server returns the authentication result to the user as a JSON message, which includes whether the authentication was successful or not, as well as any additional information required.
[0939] Step 11:
[0940] The user's device receives and analyzes the JSON message returned from the server. Based on the analysis results, the device notifies the user of the authentication result. For example, it displays a message in the browser saying, "Authentication successful. Security lock has been released."
[0941] Explanation with concrete examples
[0942] If a user is using a credit card and their card is suspended due to the card company's logic for detecting fraudulent activity, the user can use this system to safely and quickly release the security lock.
[0943] example:
[0944] The user opens a web page and records their own voice by operating the recording start / stop button. After recording is complete, the device sends the voice data to the server. The server receives the voice data and saves it in a temporary directory. The server analyzes the voice data using an AI voice authentication model and calculates an authentication score. At the same time, the emotion engine analyzes emotions from the voice data and obtains emotional data. The server evaluates the authentication score and emotional data together, and if the criteria are met, the credit card security lock is released. Finally, the server notifies the user of the authentication result, who can confirm the result and resume using their card.
[0945] In this way, the system is designed to combine voice authentication and emotion recognition to enable users to smoothly release security locks and quickly resume using their credit cards.
[0946] Example 2
[0947] Next, a description will be given of Example 2. In the following description, the data processing device 12 will be referred to as a "server" and the headset type terminal 314 will be referred to as a "terminal."
[0948] In conventional credit card fraud prevention systems, voice authentication alone does not provide sufficient security, making it difficult to quickly and safely release the security lock. In particular, the user's emotional state, such as tension or anger, can affect the authentication results, raising the risk of incorrect judgment.
[0949] The specific processing by the specific processing unit 290 of the data processing device 12 in the second embodiment is realized by the following means.
[0950] In this invention, the server includes means for recording the user's voice, means for transmitting the recorded voice data to the server, analysis means having a voice authentication model for analyzing the received voice data, analysis means having an emotion recognition engine for analyzing emotions from the voice data, means for evaluating the authentication score and emotion data of the analyzed voice data, means for releasing the security lock based on a total score of the evaluated authentication score and emotion data, and means for notifying the user of the authentication result. This enables a comprehensive evaluation based on a combination of voice authentication and emotion recognition, thereby enabling the security lock to be released more quickly and safely.
[0951] "User" refers to a person who attempts to use this system to unlock the security lock on a credit card.
[0952] "Means for recording audio" refers to the function for capturing the user's audio and saving it as digital data.
[0953] "Server" refers to a computer system that includes a central processing unit that receives and analyzes audio data.
[0954] "Audio data" refers to data that represents audio information recorded by a user in digital form.
[0955] "Voice authentication model" refers to an algorithm or program that analyzes voice data and calculates an authentication score based on the user's voice characteristics.
[0956] "Analysis means" refers to a program or process for evaluating audio data and other input data.
[0957] "Emotion recognition engine" refers to a program or algorithm that analyzes a user's emotional state from voice data and generates emotional data.
[0958] "Authentication score" refers to a numerical value calculated by a voice authentication model that indicates the degree to which a user's voice matches a pre-enrolled voice.
[0959] "Emotion data" refers to information about emotions determined from the user's voice as analyzed by an emotion recognition engine.
[0960] "Total score" refers to an evaluation value calculated by integrating the authentication score and emotional data.
[0961] "Means to unlock security lock" refers to the function for lifting restrictions on a user's credit card usage based on the overall score.
[0962] "Notification means" refers to the interface or program used to communicate authentication results and system status to users.
[0963] The system based on this invention combines voice authentication and emotion recognition to provide a means for quickly and reliably unlocking security locks to prevent fraudulent use of credit cards. This system uses the following hardware and software to process data and perform calculations.
[0964] Hardware and software used
[0965] Hardware:
[0966] Microphone: A device that inputs the user's voice.
[0967] Server: A central processing unit that receives, stores, and analyzes voice data.
[0968] software:
[0969] Web browsers: Browsers that use navigator.mediaDevices.getUserMedia to get microphone input.
[0970] Voice authentication model: An algorithm that analyzes voice data and calculates an authentication score.
[0971] Emotion recognition engine: A program that analyzes emotions from voice data.
[0972] Example of system operation
[0973] 1. User visits a web page:
[0974] The user opens a browser and accesses the specified web page.
[0975] For example, a user visits the URL https: / / securitylock.example.com.
[0976] 2. Audio Recording:
[0977] The user clicks the "Start Recording" button and follows the instructions provided to record their voice.
[0978] The device uses navigator.mediaDevices.getUserMedia to obtain the user's microphone input, creates a MediaRecorder object, and starts recording audio.
[0979] When the user clicks the "Stop Recording" button, the device stops recording and obtains the recorded audio data in Blob format.
[0980] 3. Sending audio data to the server:
[0981] The terminal sends the recorded audio data to the server via a FormData object.
[0982] Use the fetch method to send it to the server as an HTTP POST request.
[0983] 4. Receiving and storing audio data by the server:
[0984] The server receives the HTTP POST request and retrieves the audio data.
[0985] The acquired audio data is saved in a temporary directory.
[0986] 5. Voice Verification and Emotion Recognition:
[0987] The server inputs the saved audio file into an AI voice recognition model and calculates a recognition score.
[0988] At the same time, an emotion recognition engine is used to analyze the user's emotional state from the voice data and obtain emotion data.
[0989] 6. Overall score evaluation and security unlock:
[0990] The server calculates a total score based on the authentication score from the voice authentication model and the emotion data from the emotion recognition engine.
[0991] If the total score meets a certain standard, the server will unlock the credit card's security lock.
[0992] 7. Notification of Results:
[0993] The server returns the authentication result to the user as a JSON message.
[0994] The terminal analyzes the result message received from the server and notifies the user of the result.
[0995] Prompt Sentence Examples
[0996] Please send the audio data recorded by the user in the following format.
[0997] Format: Blob format
[0998] Sending method: HTTP POST
[0999] Object used: FormData
[1000] The system is designed to combine voice authentication and emotion recognition to seamlessly unlock credit card security locks, allowing users to quickly resume using their cards.
[1001] The flow of the identification process in the second embodiment will be described with reference to FIG.
[1002] Step 1:
[1003] A user visits a web page:
[1004] A user opens a browser and accesses a specified web page. For example, a user accesses the URL https: / / securitylock.example.com. The input is the user's action (accessing the URL) and the output is the displayed web page.
[1005] Step 2:
[1006] User starts audio recording:
[1007] The user clicks the "Start Recording" button on a web page. The device uses navigator.mediaDevices.getUserMedia to obtain the user's microphone input. Specifically, the browser asks the user for permission to use the microphone, and if permission is granted, the microphone stream is obtained. The input is the user's click operation and microphone input, and the output is the audio stream.
[1008] Step 3:
[1009] The device records audio:
[1010] The device creates a MediaRecorder object and starts recording audio. The data being recorded is temporarily stored in memory. The input is the microphone stream, and the output is a buffer of the data being recorded.
[1011] Step 4:
[1012] User stops recording:
[1013] The user clicks the "Stop Recording" button. This causes the device to stop recording and obtain the recorded audio data in Blob format. Specifically, the recorder.stop() method is called and the audio data is obtained in the ondataavailable event. The input is the user's click operation, and the output is audio data in Blob format.
[1014] Step 5:
[1015] The device sends the audio data to the server:
[1016] The device adds the acquired audio data to a FormData object and uses the fetch method to send it to the server as an HTTP POST request. Specifically, the following operations are performed: let formData = new FormData(); formData.append('audio', audioBlob, 'audio.wav'); fetch(' / upload', { method: 'POST', body: formData}); The input is audio data in Blob format, and the output is an HTTP request to the server.
[1017] Step 6:
[1018] The server receives the audio data:
[1019] The server receives the HTTP POST request and extracts the audio data. The extracted audio data is saved in a temporary directory. Specifically, the server-side code retrieves the audio file from req.file or similar and saves it using the fs.writeFile method. The input is the HTTP request, and the output is the saved audio file.
[1020] Step 7:
[1021] The server performs voice authentication and emotion recognition:
[1022] The server inputs the saved audio file into the AI voice authentication model and calculates an authentication score. For example, the code "let score = voiceAuthModel.authenticate(' / temp / audio.wav');" is executed. At the same time, the emotion engine analyzes the user's emotional state from the audio data and executes the operation "let emotionData = emotionEngine.analyze(' / temp / audio.wav');" to obtain emotion data. The input is the audio file, and the output is the authentication score and emotion data.
[1023] Step 8:
[1024] The server evaluates the overall score:
[1025] The server calculates the total score based on the authentication score from the voice authentication model and the emotion data. For example, it executes the operation let totalScore = calculateTotalScore(score, emotionData);. The input is the authentication score and emotion data, and the output is the total score.
[1026] Step 9:
[1027] Server removes security lock:
[1028] If the total score meets a certain criterion, the server unlocks the credit card's security lock. For example, the following operation is performed: if (totalScore > threshold) { unlockSecurityLock();}. The input is the total score, and the output is the unlock status of the credit card.
[1029] Step 10:
[1030] The server notifies the user of the authentication result:
[1031] The server returns the authentication result to the user as a JSON message. Specifically, the following operation is performed: res.json({ success: true, message: 'Security lock has been released'});. The input is the authentication result, and the output is the message sent to the user.
[1032] Step 11:
[1033] User sees the results:
[1034] The terminal analyzes the result message received from the server and notifies the user. For example, it performs the following operation: const result = await response.json(); alert(result.message);. The input is the message from the server, and the output is the notification to the user.
[1035] (Application example 2)
[1036] Next, a description will be given of Application Example 2. In the following description, the data processing device 12 will be referred to as a "server" and the headset type terminal 314 will be referred to as a "terminal."
[1037] Conventional security systems can cause inconvenience to users due to delays in detecting unauthorized use and unlocking the device. Furthermore, simple voice authentication alone cannot take into account user conditions such as emotional changes and stress. As a result, legitimate users are unable to unlock the security lock, resulting in a poor user experience.
[1038] The identification process by the identification processing unit 290 of the data processing device 12 in Application Example 2 is realized by the following means. In this invention, the server includes means for recording the user's voice, means for transmitting the recorded voice data to the server, analysis means having a voice authentication model and an emotion recognition engine for analyzing the received voice data, means for evaluating the authentication score and emotion data of the analyzed voice data, and means for calculating a total score based on the evaluated authentication score and emotion data and for releasing the security lock using the total score. This allows the legitimate user to release the security lock quickly and safely.
[1039] A "user" is an individual who records voice and uses the recorded voice data.
[1040] "Audio data" means sound in the form of data recorded by a user and transmitted to a server.
[1041] A "server" is a computer system for analyzing received voice data and performing voice authentication and emotion recognition.
[1042] "Analysis means" refers to a set of functions for analyzing audio data using a voice authentication model and an emotion recognition engine.
[1043] The "voice authentication model" is a machine learning model that analyzes received voice data and calculates an authentication score.
[1044] An "emotion recognition engine" is software that analyzes and obtains a user's emotional data from voice data.
[1045] The "authentication score" is a score calculated by the voice authentication model that indicates the degree to which the user's voice matches pre-registered data.
[1046] "Emotion data" is data that indicates the emotional state of the user, analyzed from the voice data.
[1047] The "total score" is an evaluation score calculated by combining the authentication score of the voice authentication model and emotion data.
[1048] A "security lock" is a security measure that limits the use of credit cards.
[1049] "Unlocking means" refers to a set of functions for unlocking a security lock based on the overall score.
[1050] "Notification means" refers to a method or technical means for notifying the user of the results of voice authentication and emotion recognition.
[1051] This invention relates to a system that records a user's voice, transmits the voice data to a server for analysis, and quickly and safely releases a security lock through voice authentication and emotion recognition. This system is implemented with the following configuration and procedure.
[1052] Voice recording and data transmission
[1053] The device is equipped with a microphone for recording the user's voice, a function for acquiring voice input using navigator.mediaDevices.getUserMedia, and a function for recording and saving voice data using the MediaRecorder object. The user records voice through an interface for starting and stopping recording, and sends the voice data to the server after recording is complete. To send data to the server, the voice data is added to a FormData object and sent to the server using an HTTP POST request.
[1054] Receiving and analyzing audio data
[1055] The server has the ability to store the received voice data in a temporary directory and analyze it. A voice authentication model and an emotion recognition engine are used for the analysis. The voice authentication model analyzes the user's voice data and calculates an authentication score that indicates the degree of match with pre-registered data. Meanwhile, the emotion recognition engine extracts the user's emotion data from the voice data.
[1056] Authentication score and emotion data evaluation
[1057] The server evaluates the authentication score of the voice authentication model and the emotion data obtained from the emotion recognition engine. If the overall score exceeds a certain threshold, the security lock is released.
[1058] Notification of results
[1059] The server returns the results of voice authentication and emotion recognition to the user as a JSON-formatted message. The message includes the result of authentication (success or failure) and any additional information required. The device receives this message, analyzes it, and notifies the user of the result.
[1060] Specific examples
[1061] For example, if a user tries to use their credit card but finds it is locked, they can launch the "Safe Payment Release App" on their smartphone and press the "Start Recording" button to record their voice. The recorded data is automatically sent to a server, where voice authentication and emotion recognition are performed. If the security lock can be released, the user is notified of the result and can use their credit card again.
[1062] Prompt Sentence Examples
[1063] "Please use this user's voice recording data to perform voice authentication and emotion recognition to determine whether the security lock can be released. Please use the following audio file for voice authentication and emotion recognition."
[1064] The system allows the server and terminal to work together to provide users with a quick and reliable security unlocking experience.
[1065] The flow of the specific processing in the application example 2 will be described with reference to FIG.
[1066] Step 1:
[1067] The user launches the "Safe Payment Cancellation App" on their smartphone and presses the "Start Recording" button. The device obtains microphone input using navigator.mediaDevices.getUserMedia. The device then creates a MediaRecorder object and begins recording audio. The input here is the user's voice, and the output is audio data obtained from the microphone. Specifically, when the user presses the button, the device begins microphone input and collects audio data in real time.
[1068] Step 2:
[1069] When the user presses the "Stop Recording" button, the device stops recording and obtains the recorded audio data in Blob format. The input here is the user operation that triggers the end of recording, and the output is audio data in Blob format. Specifically, when the user presses the button, the device stops MediaRecorder and converts the audio data stored in the internal buffer into Blob format.
[1070] Step 3:
[1071] The device adds this audio data in Blob format to a FormData object and sends it to the server using an HTTP POST request. The input is the audio data in Blob format, and the output is a FormData object sent to the server. Specifically, the device adds the audio data to FormData and issues an HTTP POST request to the endpoint to the server.
[1072] Step 4:
[1073] The server saves the received audio data in a temporary directory. The input is the audio data sent in the HTTP POST request, and the output is an audio file saved in the temporary directory. Specifically, the server receives the request, extracts the audio data, and saves it in a temporary directory.
[1074] Step 5:
[1075] The server inputs the saved audio file into the authenticate method of the voice authentication model and calculates an authentication score. The emotion recognition engine also analyzes and obtains the user's emotional data from the audio data. The input is the audio file, and the output is an authentication score and emotional data. Specifically, the server calls the voice authentication model and emotion recognition engine to perform the analysis.
[1076] Step 6:
[1077] The server calculates an overall score by combining the authentication score from the voice authentication model and the emotion data from the emotion recognition engine, and releases the security lock if the overall score exceeds a threshold. The input is the authentication score and emotion data, and the output is an instruction to release the security lock. Specifically, the server evaluates the score and updates the credit card security settings as necessary.
[1078] Step 7:
[1079] The server returns the results of voice authentication and emotion recognition to the user as a JSON-formatted message. The input is the overall score and evaluation result, and the output is a JSON-formatted result message sent to the user. Specifically, the server formats the results and sends an HTTP response to the user's device.
[1080] Step 8:
[1081] The user's device parses the JSON-formatted message received from the server and notifies the user of the results. The input is a JSON-formatted message, and the output is a notification to the user. Specifically, the device parses the message and displays the results to the user through an alert or notification.
[1082] The above are the specific processing steps of this system.
[1083] The specific processing unit 290 transmits the result of the specific processing to the headset type terminal 314. In the headset type terminal 314, the control unit 46A causes the speaker 240 and the display 343 to output the result of the specific processing. The microphone 238 acquires audio indicating a user input regarding the result of the specific processing. The control unit 46A transmits audio data indicating the user input acquired by the microphone 238 to the data processing device 12. In the data processing device 12, the specific processing unit 290 acquires the audio data.
[1084] The data generation model 58 is a so-called generative AI (Artificial Intelligence). An example of the data generation model 58 is ChatGPT (Internet Search<URL: https: / / openai.com / blog / chatgpt> ), Gemini (Internet search <url: https: gemini.google.com ?hl="ja">) and other generation AIs. The data generation model 58 is obtained by performing deep learning on a neural network. A prompt including an instruction is input to the data generation model 58, and inference data such as voice data indicating voice, text data indicating text, and image data indicating an image is also input. The data generation model 58 performs inference on the input inference data in accordance with the instruction indicated by the prompt, and outputs the inference result in a data format such as voice data and text data. Here, inference refers to, for example, analysis, classification, prediction, and / or summarization.
[1085] In the above embodiment, an example was given in which the specific processing is performed by the data processing device 12, but the technology of the present disclosure is not limited to this, and the specific processing may be performed by the headset type terminal 314.
[1086] [Fourth embodiment]
[1087] FIG. 7 shows an example of the configuration of a data processing system 410 according to the fourth embodiment.
[1088] 7, a data processing system 410 includes a data processing device 12 and a robot 414. An example of the data processing device 12 is a server.
[1089] The data processing device 12 includes a computer 22, a database 24, and a communication I / F 26. The computer 22 is an example of a "computer" according to the technology of the present disclosure. The computer 22 includes a processor 28, a RAM 30, and a storage 32. The processor 28, the RAM 30, and the storage 32 are connected to a bus 34. The database 24 and the communication I / F 26 are also connected to the bus 34. The communication I / F 26 is connected to a network 54. Examples of the network 54 include a WAN (Wide Area Network) and / or a LAN (Local Area Network).
[1090] The robot 414 includes a computer 36, a microphone 238, a speaker 240, a camera 42, a communication I / F 44, and a control target 443. The computer 36 includes a processor 46, a RAM 48, and a storage 50. The processor 46, the RAM 48, and the storage 50 are connected to a bus 52. The microphone 238, the speaker 240, the camera 42, and the control target 443 are also connected to the bus 52.
[1091] The microphone 238 receives instructions and the like from the user 20 by receiving voice uttered by the user 20. The microphone 238 captures the voice uttered by the user 20, converts the captured voice into audio data, and outputs it to the processor 46. The speaker 240 outputs audio in accordance with instructions from the processor 46.
[1092] Camera 42 is a small digital camera equipped with an optical system including a lens, aperture, and shutter, and an imaging element such as a CMOS (Complementary Metal-Oxide-Semiconductor) image sensor or a CCD (Charge Coupled Device) image sensor, and captures images of the surroundings of user 20 (for example, an imaging range defined by an angle of view equivalent to the field of vision of a typical healthy person).
[1093] The communication I / F 44 is connected to a network 54. The communication I / Fs 44 and 26 are responsible for the exchange of various information between the processor 46 and the processor 28 via the network 54. The exchange of various information between the processor 46 and the processor 28 using the communication I / Fs 44 and 26 is carried out in a secure state.
[1094] The control object 443 includes a display device, LEDs in the eyes, and motors for driving the arms, hands, and feet. The posture and gestures of the robot 414 are controlled by controlling the motors of the arms, hands, and feet. Some of the emotions of the robot 414 can be expressed by controlling these motors. In addition, the facial expressions of the robot 414 can also be expressed by controlling the light emission state of the LEDs in the eyes of the robot 414.
[1095] Fig. 8 shows an example of the main functions of the data processing device 12 and the robot 414. As shown in Fig. 8, in the data processing device 12, a specific process is performed by the processor 28. A specific process program 56 is stored in the storage 32.
[1096] The specific processing program 56 is an example of a "program" according to the technology of the present disclosure. The processor 28 reads the specific processing program 56 from the storage 32 and executes the read specific processing program 56 on the RAM 30. The specific processing is realized by the processor 28 operating as a specific processing unit 290 in accordance with the specific processing program 56 executed on the RAM 30.
[1097] The storage 32 stores a data generation model 58 and an emotion identification model 59. The data generation model 58 and the emotion identification model 59 are used by the identification processing unit 290.
[1098] In the robot 414, the processor 46 performs the reception output process. A reception output program 60 is stored in the storage 50. The processor 46 reads the reception output program 60 from the storage 50 and executes the read reception output program 60 on the RAM 48. The reception output process is realized by the processor 46 operating as the control unit 46A in accordance with the reception output program 60 executed on the RAM 48.
[1099] Next, a description will be given of the specific processing performed by the specific processing unit 290 of the data processing device 12. In the following description, the data processing device 12 will be referred to as a "server" and the robot 414 will be referred to as a "terminal."
[1100] The system according to the present invention provides a means for quickly and reliably releasing a security lock on a credit card to prevent fraudulent use through voice authentication. This system is implemented in the following procedure.
[1101] User voice recordings
[1102] The user accesses a dedicated web page and clicks the "Start Recording" button. The device obtains the user's microphone input using navigator.mediaDevices.getUserMedia, creates a MediaRecorder object, and starts recording audio. When the user clicks the "Stop Recording" button, the device stops recording and obtains the recorded audio data in Blob format.
[1103] Sending voice data to the server
[1104] The recorded audio data is sent from the device to the server by sending the audio file to the server via an HTTP POST request via a FormData object.
[1105] Receiving and storing audio data on the server
[1106] The server receives the HTTP POST request for the audio data and stores it in a temporary directory, preparing it for the subsequent audio authentication process.
[1107] Voice Authentication
[1108] The server uses an AI voice authentication model to obtain an authentication score based on the saved voice file. This AI model compares the voice data of pre-registered users to calculate an authentication score. The server evaluates this authentication score, and if the score exceeds a certain threshold, the user's security lock is released.
[1109] Unlocking the security lock
[1110] If the score exceeds the threshold, the server will unlock the credit card security lock based on its internal processing logic, and will also remove any restrictions on card usage after going through the necessary procedures.
[1111] Notification of results
[1112] The server notifies the user of the authentication result. The notification is sent back to the terminal in JSON format, and the terminal parses this message and displays the result to the user. The user can check the voice authentication result and find out whether they can resume using their card.
[1113] Explanation with concrete examples
[1114] Scenario: When a user is using their credit card and their card is suspended due to the card company's logic for detecting fraudulent activity, the user can use this system to safely and quickly release the security lock.
[1115] 1. User opens a web page:
[1116] Users can record their own voice by operating the recording start / stop button.
[1117] 2. Sending audio data to the server:
[1118] The device converts the recorded data into Blob format and sends it to the server.
[1119] 3. Server performs voice authentication:
[1120] The server receives the voice data and uses an AI model to calculate an authentication score.
[1121] 4. Security unlock and notifications:
[1122] - The server determines that the authentication score criteria have been met and releases the security lock. The user is then shown a message indicating successful authentication, allowing them to resume using their credit card.
[1123] In this way, the system is designed to enable users to smoothly release the security lock and quickly resume using their credit card.
[1124] The processing flow will be explained below.
[1125] Step 1:
[1126] A user accesses a web page and clicks the "Start Recording" button. The device uses navigator.mediaDevices.getUserMedia to obtain the user's microphone input. At this time, the user is prompted for microphone access permission in the browser and must grant it.
[1127] Step 2:
[1128] The device starts recording by creating a MediaRecorder object, which records the audio stream from the microphone in real time and collects the audio data.
[1129] Step 3:
[1130] The user clicks the "Stop Recording" button. The device stops recording and the audio data captured by MediaRecorder is retrieved in Blob format. This aggregates the audio data into a single object.
[1131] Step 4:
[1132] The device adds the acquired blob-formatted audio data to a FormData object and sends it to the server using an HTTP POST request. A FormData object is a data structure for sending files or data to the server in multiple key-value formats.
[1133] Step 5:
[1134] The server receives the HTTP POST request, extracts the audio file from the request, stores it in a temporary directory, and later uses it in the authentication process.
[1135] Step 6:
[1136] The server inputs the saved audio file into the authenticate method of the AI voice authentication model, which uses an algorithm to calculate a similarity score with the voice of a pre-registered user.
[1137] Step 7:
[1138] The server evaluates the authentication score returned by the AI model, which includes determining whether the score exceeds a certain threshold, which is set according to the security level.
[1139] Step 8:
[1140] If the score exceeds a threshold, the server will release the security lock based on internal logic, a process that includes updating security settings and removing card usage restrictions.
[1141] Step 9:
[1142] The server returns the authentication result to the user as a JSON message, which includes whether the authentication was successful or not, as well as any additional information required.
[1143] Step 10:
[1144] The user's device receives and analyzes the JSON message returned from the server. Based on the analysis results, the device notifies the user of the authentication result. For example, it displays a message in the browser saying, "Authentication successful. Security lock has been released."
[1145] In this way, a voice authentication system is realized that allows the user to smoothly release the security lock and resume use of the credit card.
[1146] Example 1
[1147] Next, a description will be given of Example 1. In the following description, the data processing device 12 will be referred to as a "server" and the robot 414 will be referred to as a "terminal."
[1148] Conventional credit card security lock release systems require users to manually perform authentication methods and procedures, which are time-consuming and laborious. Therefore, there was a need for a system that would allow users to quickly and reliably release security locks.
[1149] The specific processing by the specific processing unit 290 of the data processing device 12 in the first embodiment is realized by the following means.
[1150] In this invention, the server includes means for recording the user's voice, means for transmitting the recorded voice data to the server, analysis means having an AI voice authentication model for analyzing the received voice data, means for evaluating the authentication score of the analyzed voice data, means for releasing the security lock based on the evaluated authentication score, means for notifying the user of the result of the voice authentication, and means for transmitting the voice data via an HTTP POST request. This enables the user to easily, quickly, and reliably release the security lock using their voice.
[1151] "User" refers to an individual who uses the System to unlock a credit card security lock.
[1152] "Audio recording means" refers to a device or function for obtaining a user's microphone input and recording that audio data.
[1153] "Means for transmitting audio data to a server" refers to a device or function for transferring recorded audio data to a server using an appropriate protocol.
[1154] "AI voice authentication model" refers to a machine learning algorithm or program that analyzes voice data and performs specific identity verification.
[1155] "Analysis means" refers to a device or function that processes received voice data using an AI voice authentication model and calculates an authentication score.
[1156] "Means for assessing the authentication score" refers to a device or function for determining the calculated authentication score based on the analyzed voice data and determining the next processing step accordingly.
[1157] "Means for unlocking security lock" refers to a device or function for unlocking credit card usage restrictions when the authentication score exceeds a threshold value.
[1158] "Means for notifying the result of voice authentication" refers to a device or function for notifying the user of the result of voice authentication.
[1159] "Means for sending via HTTP POST request" refers to a device or function for sending recorded audio data to a server using the POST method of the HTTP protocol.
[1160] "Interface" refers to a graphical user interface that provides a means for a user to start and stop recording.
[1161] The present invention relates to a system for quickly and reliably releasing a security lock to prevent fraudulent use of a credit card through voice authentication. The system is specifically implemented using the following hardware and software.
[1162] Hardware and Software Used
[1163] 1. User voice recordings
[1164] The user accesses a dedicated web page and operates the "Start Recording" and "Stop Recording" buttons in the browser. The browser uses JavaScript's navigator.mediaDevices.getUserMedia API to obtain the user's microphone input. The audio is recorded using the MediaRecorder object.
[1165] 2. Sending audio data to the server
[1166] After finishing recording, the device converts the recorded audio data into Blob format and sends it to the server via an HTTP POST request using the FormData object. This communication uses the JavaScript fetch API.
[1167] 3. Receiving and storing audio data by the server
[1168] The server receives the audio data via an HTTP POST request and temporarily stores it in a file system managed, for example, using the Node.js fs module.
[1169] 4. Voice Authentication Process
[1170] The server inputs the saved audio file into an AI voice recognition model. This model is pre-trained and compares the user's pre-registered voice data with the target voice data to calculate a recognition score. The AI model is run using libraries such as TensorFlow and PyTorch.
[1171] 5. Unlocking the security lock
[1172] If the authentication score calculated based on the voice data exceeds the threshold, the server executes the procedure to release the credit card security lock. In this step, the API of the card company is called and the necessary procedures are followed to remove the usage restrictions.
[1173] 6. Notification of authentication results
[1174] If authentication is successful, the server notifies the result in JSON format, and the browser receives the response and displays the result to the user.
[1175] Specific examples
[1176] For example, suppose a user is using a credit card and the card is suspended due to the card company's logic for detecting fraudulent activity. In this case, the security lock can be released by following the procedure below.
[1177] 1. Users access a dedicated web page and click the "Start Recording" button to record their own voice.
[1178] 2. When you finish recording, click the "Stop Recording" button and the audio data will be converted to Blob format and sent to the server.
[1179] 3. The server receives the audio data and temporarily stores it.
[1180] 4. The saved voice data is analyzed using an AI voice authentication model to calculate an authentication score.
[1181] 5. If the authentication score exceeds the threshold, the server executes the procedure to unlock the credit card security lock.
[1182] 6. Finally, the server notifies the user of the result, and the user confirms that the credit card can be used again.
[1183] This allows users to quickly and reliably release security locks using voice commands.
[1184] The flow of the identification process in the first embodiment will be described with reference to FIG.
[1185] Step 1:
[1186] The user accesses a dedicated web page and clicks the "Start Recording" button. The device uses the navigator.mediaDevices.getUserMedia API to obtain the user's microphone input, creates a MediaRecorder object, and begins recording audio. Specifically, the browser displays a dialog requesting permission to use the microphone. Once audio input is obtained, MediaRecorder starts and recording begins. The input at this time is the user's voice, and the output is the audio data being recorded.
[1187] Step 2:
[1188] When the user clicks the "Stop Recording" button, the device stops recording and obtains the recorded audio data in Blob format. Specifically, the MediaRecorder object stops recording data, and the resulting audio data is saved in Blob format on the device. The input at this time is the audio data being recorded, and the output is audio data in Blob format.
[1189] Step 3:
[1190] The device adds the recorded audio data in Blob format to a FormData object and sends it to the server via an HTTP POST request. This communication is performed using JavaScript's fetch API. Specifically, the audio data is added to the FormData object, and the fetch API sends an HTTP POST request to the server. The input at this time is audio data in Blob format, and the output is an HTTP POST request to the server.
[1191] Step 4:
[1192] The server receives an HTTP POST request containing audio data and temporarily stores the data in a directory. The data received on the server side is saved in a local temporary directory using, for example, the Node.js fs module. The input is the HTTP POST request sent to the server, and the output is the audio data saved in the temporary directory.
[1193] Step 5:
[1194] The server inputs the saved voice data into an AI voice recognition model and calculates a recognition score. The AI voice recognition model is a pre-trained model that compares the voice data to calculate a recognition score. Specifically, the AI model is run using libraries such as TensorFlow and PyTorch. The input is the voice data saved in a temporary directory, and the output is a recognition score.
[1195] Step 6:
[1196] The server evaluates the authentication score and, if the score exceeds the threshold, releases the credit card's security lock. The server then calls the card company's API and executes the procedure to release the usage restriction. Specifically, it sends an HTTP request to the card company's API endpoint and obtains confirmation of the release. The input is the authentication score, and the output is confirmation of the release from the card company.
[1197] Step 7:
[1198] The server notifies the user of the authentication result in a JSON format message. The device parses this message and displays the authentication result to the user. Specifically, the server generates a JSON format response and sends it to the device as an HTTP response. The device receives this response and displays the authentication result on the screen. The input at this time is the JSON message of the authentication result generated by the server, and the output is the authentication result displayed on the device.
[1199] (Application example 1)
[1200] Next, a description will be given of Application Example 1. In the following description, the data processing device 12 will be referred to as a "server" and the robot 414 will be referred to as a "terminal."
[1201] Conventional credit card fraud prevention systems require complicated procedures to activate security locks, making it difficult to quickly unlock them. Furthermore, the interface and notification methods for users to authenticate using their own voice are inadequate, reducing user convenience. Given these circumstances, there is a demand for a system that can quickly and reliably unlock credit card security locks and instantly notify users of the results.
[1202] The specific processing by the specific processing unit 290 of the data processing device 12 in the application example 1 is realized by the following means.
[1203] In this invention, the server includes means for recording the user's voice, means for transmitting the recorded voice data to the server, analysis means having a voice authentication model for analyzing the received voice data, means for evaluating the authentication score of the analyzed voice data, means for releasing the security lock based on the evaluated authentication score, display means including operation buttons for recording and transmitting the voice, and means for notifying the user of the result. This allows the user to quickly release the security lock on the credit card using their own voice and immediately confirm the result.
[1204] "User" means an individual who uses the system to record voice and attempt to unlock a credit card security lock.
[1205] "Audio recording means" refers to a device or application for recording a user's voice.
[1206] The "means for transmitting recorded voice data to a server" refers to a communication means for uploading recorded voice data to a server via a network.
[1207] A "voice authentication model" is an algorithm or AI model that analyzes recorded voice data and authenticates whether it is the user's voice.
[1208] The "analysis means" is a device that includes hardware and software for analyzing the voice data sent to the server.
[1209] The "means for evaluating the authentication score" refers to a device or program for determining whether authentication is successful or unsuccessful based on the score calculated by the voice authentication model.
[1210] A "means for unlocking security lock" is a device or program that executes a procedure to unlock a credit card when the authentication score exceeds a reference value.
[1211] The "display means including operation buttons" refers to a screen on which buttons for the user to operate to start and stop voice recording are displayed.
[1212] The "notification means" is a device or program for notifying the user of the result of voice authentication.
[1213] The term "computer system" refers to a collection of hardware and software that includes the above-mentioned means and operates as a whole.
[1214] This invention is a security system that allows users to quickly and reliably release the security lock on their credit cards using their own voice. This system is realized by combining the procedures of voice recording, data transmission, voice authentication, and result notification. The details are described below.
[1215] server
[1216] The server first receives the voice data sent by the user. After receiving the voice data, the AI voice authentication model within the server analyzes the voice data and calculates an authentication score. In this process, a machine learning model (e.g., TensorFlow, PyTorch) is used to determine whether the specific voice data matches the voice data of a pre-registered user. If the authentication score exceeds a threshold, the server releases the credit card security lock based on its internal logic. The authentication result is then returned to the user in JSON format, allowing interest confirmation.
[1217] Terminal
[1218] The device (such as the user's smartphone or computer) provides a user interface and displays operation buttons (record start button and record stop button) that the user can use to record audio. It records audio in response to specific operations, converts the audio data into Blob format, and sends it to the server. It also displays the authentication result to the user based on the authentication result returned from the server.
[1219] User
[1220] When a security lock is activated, the user opens the provided browser-based application and records their voice. After recording, the voice data is sent to the server, after which the authentication result is notified. This process allows the user to smoothly unlock the security lock on their credit card and resume use.
[1221] Specific examples
[1222] A user on a business trip attempts to use their credit card in a new city, but the card is locked due to fraud prevention logic. In this case, the user launches the smartphone app, records their voice, and sends it to the server. The server receives the voice data and calculates an authentication score using an AI voice authentication model. If authentication is successful, the server releases the card's security lock and notifies the user of the result. The user receives the notification and can resume using their credit card while on a business trip.
[1223] Prompt Sentence Examples
[1224] User: "Please press the start recording button and say the following phrase: 'This audio will verify my identity.'"
[1225] Input prompt to AI model: "Please authenticate the voice data below to confirm whether this voice belongs to a registered user."
[1226] The flow of the specific processing in the application example 1 will be described with reference to FIG.
[1227] Step 1:
[1228] The user opens a security application or web page, which displays a "Start Recording" button to start recording and a "Stop Recording" button to stop recording. The input of this step is the user's operation, and the output is the display of the user interface.
[1229] Step 2:
[1230] When the user presses the "Start Recording" button, the device accesses the microphone using navigator.mediaDevices.getUserMedia and starts recording. The recorded audio is temporarily saved on the device. The inputs to this step are the user's operation and audio data, and the outputs are the recording status and the temporarily saved audio data.
[1231] Step 3:
[1232] When the user presses the "Stop Recording" button, the device stops recording using the MediaRecorder object and converts the audio data to Blob format, which is then prepared for sending to the server. The input to this step is the audio data being recorded, and the output is Blob format audio data.
[1233] Step 4:
[1234] The device adds the audio data in Blob format to a FormData object and sends it to the server using an HTTP POST request using the fetch API. The input of this step is the audio data in Blob format, and the output is the data sent to the server.
[1235] Step 5:
[1236] The server saves the audio data received via the HTTP POST request in a temporary directory. Here, it checks whether the received audio data has been received and saved correctly. The input of this step is the HTTP request, and the output is the audio file on the server.
[1237] Step 6:
[1238] The server analyzes the received voice data using an AI voice authentication model (e.g., TensorFlow, PyTorch) and calculates an authentication score. The voice authentication model compares the voice data with that of registered users to measure the degree of match. The input for this step is the received voice data, and the output is an authentication score.
[1239] Step 7:
[1240] The server evaluates the calculated authentication score to determine if it exceeds a threshold. If so, it executes a procedure to unlock the credit card security lock. The input to this step is the authentication score, and the output is the execution of the unlock.
[1241] Step 8:
[1242] The server returns the authentication result (success or failure) to the device as a JSON-formatted message. The input to this step is the result of the authentication process, and the output is a JSON-formatted message.
[1243] Step 9:
[1244] The terminal analyzes the authentication result received from the server and notifies the user through the user interface. The user can check on the screen whether the credit card can be used or not. The input of this step is the authentication result in JSON format, and the output is the message displayed to the user.
[1245] Furthermore, an emotion engine that estimates the user's emotion may be further combined. That is, the identification processing unit 290 may estimate the user's emotion using the emotion identification model 59, and perform identification processing using the user's emotion.
[1246] The system according to the present invention provides a means for quickly and reliably unlocking a credit card's security lock to prevent fraudulent use through voice authentication and emotion recognition. This system is implemented in the following steps.
[1247] User voice recordings
[1248] The user accesses a dedicated web page and clicks the "Start Recording" button. The device uses navigator.mediaDevices.getUserMedia to obtain the user's microphone input, creates a MediaRecorder object, and starts recording audio. When the user clicks the "Stop Recording" button, the device stops recording and obtains the recorded audio data in Blob format.
[1249] Sending voice data to the server
[1250] The recorded audio data is sent from the device to the server by sending the audio file to the server via an HTTP POST request via a FormData object.
[1251] Receiving and storing audio data on the server
[1252] The server receives an HTTP POST request for audio data, extracts the audio file from the request, and stores it in a temporary directory for later use in the authentication process.
[1253] Voice Authentication
[1254] The server inputs the saved audio file into the authenticate method of the AI voice authentication model. This model compares it with the voice data of pre-registered users and calculates an authentication score. In parallel, the emotion engine analyzes the user's emotions from the audio data and obtains the results.
[1255] Authentication score and sentiment rating
[1256] The server evaluates the authentication score returned by the voice authentication model and the emotion data obtained from the emotion engine. This evaluation includes calculating an overall score that combines the authentication score and the emotion data. For example, if the user expresses anger or tension, this will be taken into account as a factor that affects the overall score.
[1257] Unlocking the security lock
[1258] If the total score exceeds a certain threshold, the server will release the security lock based on internal logic, a process that includes updating security settings and removing card usage restrictions.
[1259] Notification of results
[1260] The server returns the authentication result to the user as a JSON message. The message contains information about whether the authentication was successful or not, as well as any additional information required. The user's device receives this message, parses it, and notifies the user of the result.
[1261] Explanation with concrete examples
[1262] Scenario: When a user is using their credit card and their card is suspended due to the card company's logic for detecting fraudulent activity, the user can use this system to safely and quickly release the security lock.
[1263] 1. User opens a web page:
[1264] The user can record their own voice by operating the recording start / stop button. After recording is complete, the device sends the voice data to the server.
[1265] 2. Receiving and storing audio data by the server:
[1266] The server receives the audio data and stores it in a temporary directory.
[1267] 3. Server performs voice authentication and emotion recognition:
[1268] The server analyzes the voice data using an AI voice recognition model to calculate a recognition score. At the same time, the emotion engine analyzes emotions from the voice data and obtains emotion data.
[1269] 4. Overall score evaluation and security unlock:
[1270] The server evaluates the authentication score and emotional data together, and unlocks the credit card security lock if the criteria are met.
[1271] 5. Notification of Results:
[1272] The server notifies the user of the authentication result, and the user can confirm the result and resume using the card.
[1273] In this way, the system is designed to combine voice authentication and emotion recognition to enable users to smoothly release security locks and quickly resume using their credit cards.
[1274] The processing flow will be explained below.
[1275] Step 1:
[1276] A user accesses a web page and clicks the "Start Recording" button. The device uses navigator.mediaDevices.getUserMedia to obtain the user's microphone input. At this time, the user is prompted for microphone access permission in the browser and must grant it.
[1277] Step 2:
[1278] The device starts recording by creating a MediaRecorder object, which records the audio stream from the microphone in real time and collects the audio data.
[1279] Step 3:
[1280] The user clicks the "Stop Recording" button. The device stops recording and the audio data captured by MediaRecorder is retrieved in Blob format. This aggregates the audio data into a single object.
[1281] Step 4:
[1282] The device adds the acquired blob-formatted audio data to a FormData object and sends it to the server using an HTTP POST request. A FormData object is a data structure for sending files or data to the server in multiple key-value formats.
[1283] Step 5:
[1284] The server receives the HTTP POST request, extracts the audio file from the request, stores it in a temporary directory, and later uses it in the authentication process.
[1285] Step 6:
[1286] The server inputs the saved audio file into the authenticate method of the AI voice authentication model, which uses an algorithm to calculate a similarity score with the voice of a pre-registered user.
[1287] Step 7:
[1288] The server also sends the voice data to the emotion engine, which analyzes the user's emotions. The emotion engine analyzes the voice characteristics and determines whether the user is expressing anger, joy, sadness, or other emotions.
[1289] Step 8:
[1290] The server combines the authentication score returned by the AI model with the emotion data obtained from the emotion engine to calculate an overall score, which is generated using a calculation algorithm based on the voice authentication score and emotion data.
[1291] Step 9:
[1292] If the total score exceeds a certain threshold, the server will release the security lock based on internal logic, a process that includes updating security settings and removing card usage restrictions.
[1293] Step 10:
[1294] The server returns the authentication result to the user as a JSON message, which includes whether the authentication was successful or not, as well as any additional information required.
[1295] Step 11:
[1296] The user's device receives and analyzes the JSON message returned from the server. Based on the analysis results, the device notifies the user of the authentication result. For example, it displays a message in the browser saying, "Authentication successful. Security lock has been released."
[1297] Explanation with concrete examples
[1298] If a user is using a credit card and their card is suspended due to the card company's logic for detecting fraudulent activity, the user can use this system to safely and quickly release the security lock.
[1299] example:
[1300] The user opens a web page and records their own voice by operating the recording start / stop button. After recording is complete, the device sends the voice data to the server. The server receives the voice data and saves it in a temporary directory. The server analyzes the voice data using an AI voice authentication model and calculates an authentication score. At the same time, the emotion engine analyzes emotions from the voice data and obtains emotional data. The server evaluates the authentication score and emotional data together, and if the criteria are met, the credit card security lock is released. Finally, the server notifies the user of the authentication result, who can confirm the result and resume using their card.
[1301] In this way, the system is designed to combine voice authentication and emotion recognition to enable users to smoothly release security locks and quickly resume using their credit cards.
[1302] Example 2
[1303] Next, a description will be given of Example 2. In the following description, the data processing device 12 will be referred to as a "server" and the robot 414 will be referred to as a "terminal."
[1304] In conventional credit card fraud prevention systems, voice authentication alone does not provide sufficient security, making it difficult to quickly and safely release the security lock. In particular, the user's emotional state, such as tension or anger, can affect the authentication results, raising the risk of incorrect judgment.
[1305] The specific processing by the specific processing unit 290 of the data processing device 12 in the second embodiment is realized by the following means.
[1306] In this invention, the server includes means for recording the user's voice, means for transmitting the recorded voice data to the server, analysis means having a voice authentication model for analyzing the received voice data, analysis means having an emotion recognition engine for analyzing emotions from the voice data, means for evaluating the authentication score and emotion data of the analyzed voice data, means for releasing the security lock based on a total score of the evaluated authentication score and emotion data, and means for notifying the user of the authentication result. This enables a comprehensive evaluation based on a combination of voice authentication and emotion recognition, thereby enabling the security lock to be released more quickly and safely.
[1307] "User" refers to a person who attempts to use this system to unlock the security lock on a credit card.
[1308] "Means for recording audio" refers to the function for capturing the user's audio and saving it as digital data.
[1309] "Server" refers to a computer system that includes a central processing unit that receives and analyzes audio data.
[1310] "Audio data" refers to data that represents audio information recorded by a user in digital form.
[1311] "Voice authentication model" refers to an algorithm or program that analyzes voice data and calculates an authentication score based on the user's voice characteristics.
[1312] "Analysis means" refers to a program or process for evaluating audio data and other input data.
[1313] "Emotion recognition engine" refers to a program or algorithm that analyzes a user's emotional state from voice data and generates emotional data.
[1314] "Authentication score" refers to a numerical value calculated by a voice authentication model that indicates the degree to which a user's voice matches a pre-enrolled voice.
[1315] "Emotion data" refers to information about emotions determined from the user's voice as analyzed by an emotion recognition engine.
[1316] "Total score" refers to an evaluation value calculated by integrating the authentication score and emotional data.
[1317] "Means to unlock security lock" refers to the function for lifting restrictions on a user's credit card usage based on the overall score.
[1318] "Notification means" refers to the interface or program used to communicate authentication results and system status to users.
[1319] The system based on this invention combines voice authentication and emotion recognition to provide a means for quickly and reliably unlocking security locks to prevent fraudulent use of credit cards. This system uses the following hardware and software to process data and perform calculations.
[1320] Hardware and software used
[1321] Hardware:
[1322] Microphone: A device that inputs the user's voice.
[1323] Server: A central processing unit that receives, stores, and analyzes voice data.
[1324] software:
[1325] Web browsers: Browsers that use navigator.mediaDevices.getUserMedia to get microphone input.
[1326] Voice authentication model: An algorithm that analyzes voice data and calculates an authentication score.
[1327] Emotion recognition engine: A program that analyzes emotions from voice data.
[1328] Example of system operation
[1329] 1. User visits a web page:
[1330] The user opens a browser and accesses the specified web page.
[1331] For example, a user visits the URL https: / / securitylock.example.com.
[1332] 2. Audio Recording:
[1333] The user clicks the "Start Recording" button and follows the instructions provided to record their voice.
[1334] The device uses navigator.mediaDevices.getUserMedia to obtain the user's microphone input, creates a MediaRecorder object, and starts recording audio.
[1335] When the user clicks the "Stop Recording" button, the device stops recording and obtains the recorded audio data in Blob format.
[1336] 3. Sending audio data to the server:
[1337] The terminal sends the recorded audio data to the server via a FormData object.
[1338] Use the fetch method to send it to the server as an HTTP POST request.
[1339] 4. Receiving and storing audio data by the server:
[1340] The server receives the HTTP POST request and retrieves the audio data.
[1341] The acquired audio data is saved in a temporary directory.
[1342] 5. Voice Verification and Emotion Recognition:
[1343] The server inputs the saved audio file into an AI voice recognition model and calculates a recognition score.
[1344] At the same time, an emotion recognition engine is used to analyze the user's emotional state from the voice data and obtain emotion data.
[1345] 6. Overall score evaluation and security unlock:
[1346] The server calculates a total score based on the authentication score from the voice authentication model and the emotion data from the emotion recognition engine.
[1347] If the total score meets a certain standard, the server will unlock the credit card's security lock.
[1348] 7. Notification of Results:
[1349] The server returns the authentication result to the user as a JSON message.
[1350] The terminal analyzes the result message received from the server and notifies the user of the result.
[1351] Prompt Sentence Examples
[1352] Please send the audio data recorded by the user in the following format.
[1353] Format: Blob format
[1354] Sending method: HTTP POST
[1355] Object used: FormData
[1356] The system is designed to combine voice authentication and emotion recognition to seamlessly unlock credit card security locks, allowing users to quickly resume using their cards.
[1357] The flow of the identification process in the second embodiment will be described with reference to FIG.
[1358] Step 1:
[1359] A user visits a web page:
[1360] A user opens a browser and accesses a specified web page. For example, a user accesses the URL https: / / securitylock.example.com. The input is the user's action (accessing the URL) and the output is the displayed web page.
[1361] Step 2:
[1362] User starts audio recording:
[1363] The user clicks the "Start Recording" button on a web page. The device uses navigator.mediaDevices.getUserMedia to obtain the user's microphone input. Specifically, the browser asks the user for permission to use the microphone, and if permission is granted, the microphone stream is obtained. The input is the user's click operation and microphone input, and the output is the audio stream.
[1364] Step 3:
[1365] The device records audio:
[1366] The device creates a MediaRecorder object and starts recording audio. The data being recorded is temporarily stored in memory. The input is the microphone stream, and the output is a buffer of the data being recorded.
[1367] Step 4:
[1368] User stops recording:
[1369] The user clicks the "Stop Recording" button. This causes the device to stop recording and obtain the recorded audio data in Blob format. Specifically, the recorder.stop() method is called and the audio data is obtained in the ondataavailable event. The input is the user's click operation, and the output is audio data in Blob format.
[1370] Step 5:
[1371] The device sends the audio data to the server:
[1372] The device adds the acquired audio data to a FormData object and uses the fetch method to send it to the server as an HTTP POST request. Specifically, the following operations are performed: let formData = new FormData(); formData.append('audio', audioBlob, 'audio.wav'); fetch(' / upload', { method: 'POST', body: formData}); The input is audio data in Blob format, and the output is an HTTP request to the server.
[1373] Step 6:
[1374] The server receives the audio data:
[1375] The server receives the HTTP POST request and extracts the audio data. The extracted audio data is saved in a temporary directory. Specifically, the server-side code retrieves the audio file from req.file or similar and saves it using the fs.writeFile method. The input is the HTTP request, and the output is the saved audio file.
[1376] Step 7:
[1377] The server performs voice authentication and emotion recognition:
[1378] The server inputs the saved audio file into the AI voice authentication model and calculates an authentication score. For example, the code "let score = voiceAuthModel.authenticate(' / temp / audio.wav');" is executed. At the same time, the emotion engine analyzes the user's emotional state from the audio data and executes the operation "let emotionData = emotionEngine.analyze(' / temp / audio.wav');" to obtain emotion data. The input is the audio file, and the output is the authentication score and emotion data.
[1379] Step 8:
[1380] The server evaluates the overall score:
[1381] The server calculates the total score based on the authentication score from the voice authentication model and the emotion data. For example, it executes the operation let totalScore = calculateTotalScore(score, emotionData);. The input is the authentication score and emotion data, and the output is the total score.
[1382] Step 9:
[1383] Server removes security lock:
[1384] If the total score meets a certain criterion, the server unlocks the credit card's security lock. For example, the following operation is performed: if (totalScore > threshold) { unlockSecurityLock();}. The input is the total score, and the output is the unlock status of the credit card.
[1385] Step 10:
[1386] The server notifies the user of the authentication result:
[1387] The server returns the authentication result to the user as a JSON message. Specifically, the following operation is performed: res.json({ success: true, message: 'Security lock has been released'});. The input is the authentication result, and the output is the message sent to the user.
[1388] Step 11:
[1389] User sees the results:
[1390] The terminal analyzes the result message received from the server and notifies the user. For example, it performs the following operation: const result = await response.json(); alert(result.message);. The input is the message from the server, and the output is the notification to the user.
[1391] (Application example 2)
[1392] Next, a description will be given of Application Example 2. In the following description, the data processing device 12 will be referred to as a "server" and the robot 414 will be referred to as a "terminal."
[1393] Conventional security systems can cause inconvenience to users due to delays in detecting unauthorized use and unlocking the device. Furthermore, simple voice authentication alone cannot take into account user conditions such as emotional changes and stress. As a result, legitimate users are unable to unlock the security lock, resulting in a poor user experience.
[1394] The identification process by the identification processing unit 290 of the data processing device 12 in Application Example 2 is realized by the following means. In this invention, the server includes means for recording the user's voice, means for transmitting the recorded voice data to the server, analysis means having a voice authentication model and an emotion recognition engine for analyzing the received voice data, means for evaluating the authentication score and emotion data of the analyzed voice data, and means for calculating a total score based on the evaluated authentication score and emotion data and for releasing the security lock using the total score. This allows the legitimate user to release the security lock quickly and safely.
[1395] A "user" is an individual who records voice and uses the recorded voice data.
[1396] "Audio data" means sound in the form of data recorded by a user and transmitted to a server.
[1397] A "server" is a computer system for analyzing received voice data and performing voice authentication and emotion recognition.
[1398] "Analysis means" refers to a set of functions for analyzing audio data using a voice authentication model and an emotion recognition engine.
[1399] The "voice authentication model" is a machine learning model that analyzes received voice data and calculates an authentication score.
[1400] An "emotion recognition engine" is software that analyzes and obtains a user's emotional data from voice data.
[1401] The "authentication score" is a score calculated by the voice authentication model that indicates the degree to which the user's voice matches pre-registered data.
[1402] "Emotion data" is data that indicates the emotional state of the user, analyzed from the voice data.
[1403] The "total score" is an evaluation score calculated by combining the authentication score of the voice authentication model and emotion data.
[1404] A "security lock" is a security measure that limits the use of credit cards.
[1405] "Unlocking means" refers to a set of functions for unlocking a security lock based on the overall score.
[1406] "Notification means" refers to a method or technical means for notifying the user of the results of voice authentication and emotion recognition.
[1407] This invention relates to a system that records a user's voice, transmits the voice data to a server for analysis, and quickly and safely releases a security lock through voice authentication and emotion recognition. This system is implemented with the following configuration and procedure.
[1408] Voice recording and data transmission
[1409] The device is equipped with a microphone for recording the user's voice, a function for acquiring voice input using navigator.mediaDevices.getUserMedia, and a function for recording and saving voice data using the MediaRecorder object. The user records voice through an interface for starting and stopping recording, and sends the voice data to the server after recording is complete. To send data to the server, the voice data is added to a FormData object and sent to the server using an HTTP POST request.
[1410] Receiving and analyzing audio data
[1411] The server has the ability to store the received voice data in a temporary directory and analyze it. A voice authentication model and an emotion recognition engine are used for the analysis. The voice authentication model analyzes the user's voice data and calculates an authentication score that indicates the degree of match with pre-registered data. Meanwhile, the emotion recognition engine extracts the user's emotion data from the voice data.
[1412] Authentication score and emotion data evaluation
[1413] The server evaluates the authentication score of the voice authentication model and the emotion data obtained from the emotion recognition engine. If the overall score exceeds a certain threshold, the security lock is released.
[1414] Notification of results
[1415] The server returns the results of voice authentication and emotion recognition to the user as a JSON-formatted message. The message includes the result of authentication (success or failure) and any additional information required. The device receives this message, analyzes it, and notifies the user of the result.
[1416] Specific examples
[1417] For example, if a user tries to use their credit card but finds it is locked, they can launch the "Safe Payment Release App" on their smartphone and press the "Start Recording" button to record their voice. The recorded data is automatically sent to a server, where voice authentication and emotion recognition are performed. If the security lock can be released, the user is notified of the result and can use their credit card again.
[1418] Prompt Sentence Examples
[1419] "Please use this user's voice recording data to perform voice authentication and emotion recognition to determine whether the security lock can be released. Please use the following audio file for voice authentication and emotion recognition."
[1420] The system allows the server and terminal to work together to provide users with a quick and reliable security unlocking experience.
[1421] The flow of the specific processing in the application example 2 will be described with reference to FIG.
[1422] Step 1:
[1423] The user launches the "Safe Payment Cancellation App" on their smartphone and presses the "Start Recording" button. The device obtains microphone input using navigator.mediaDevices.getUserMedia. The device then creates a MediaRecorder object and begins recording audio. The input here is the user's voice, and the output is audio data obtained from the microphone. Specifically, when the user presses the button, the device begins microphone input and collects audio data in real time.
[1424] Step 2:
[1425] When the user presses the "Stop Recording" button, the device stops recording and obtains the recorded audio data in Blob format. The input here is the user operation that triggers the end of recording, and the output is audio data in Blob format. Specifically, when the user presses the button, the device stops MediaRecorder and converts the audio data stored in the internal buffer into Blob format.
[1426] Step 3:
[1427] The device adds this audio data in Blob format to a FormData object and sends it to the server using an HTTP POST request. The input is the audio data in Blob format, and the output is a FormData object sent to the server. Specifically, the device adds the audio data to FormData and issues an HTTP POST request to the endpoint to the server.
[1428] Step 4:
[1429] The server saves the received audio data in a temporary directory. The input is the audio data sent in the HTTP POST request, and the output is an audio file saved in the temporary directory. Specifically, the server receives the request, extracts the audio data, and saves it in a temporary directory.
[1430] Step 5:
[1431] The server inputs the saved audio file into the authenticate method of the voice authentication model and calculates an authentication score. The emotion recognition engine also analyzes and obtains the user's emotional data from the audio data. The input is the audio file, and the output is an authentication score and emotional data. Specifically, the server calls the voice authentication model and emotion recognition engine to perform the analysis.
[1432] Step 6:
[1433] The server calculates an overall score by combining the authentication score from the voice authentication model and the emotion data from the emotion recognition engine, and releases the security lock if the overall score exceeds a threshold. The input is the authentication score and emotion data, and the output is an instruction to release the security lock. Specifically, the server evaluates the score and updates the credit card security settings as necessary.
[1434] Step 7:
[1435] The server returns the results of voice authentication and emotion recognition to the user as a JSON-formatted message. The input is the overall score and evaluation result, and the output is a JSON-formatted result message sent to the user. Specifically, the server formats the results and sends an HTTP response to the user's device.
[1436] Step 8:
[1437] The user's device parses the JSON-formatted message received from the server and notifies the user of the results. The input is a JSON-formatted message, and the output is a notification to the user. Specifically, the device parses the message and displays the results to the user through an alert or notification.
[1438] The above are the specific processing steps of this system.
[1439] The specific processing unit 290 transmits the result of the specific processing to the robot 414. In the robot 414, the control unit 46A causes the speaker 240 and the control target 443 to output the result of the specific processing. The microphone 238 acquires voice indicating a user input regarding the result of the specific processing. The control unit 46A transmits voice data indicating the user input acquired by the microphone 238 to the data processing device 12. In the data processing device 12, the specific processing unit 290 acquires the voice data.
[1440] The data generation model 58 is a so-called generative AI (Artificial Intelligence). An example of the data generation model 58 is ChatGPT (Internet Search<URL: https: / / openai.com / blog / chatgpt> ), Gemini (Internet search <url: https: gemini.google.com ?hl="ja">) and other generation AIs. The data generation model 58 is obtained by performing deep learning on a neural network. A prompt including an instruction is input to the data generation model 58, and inference data such as voice data indicating voice, text data indicating text, and image data indicating an image is also input. The data generation model 58 performs inference on the input inference data in accordance with the instruction indicated by the prompt, and outputs the inference result in a data format such as voice data and text data. Here, inference refers to, for example, analysis, classification, prediction, and / or summarization.
[1441] In the above embodiment, an example was given in which the specific processing is performed by the data processing device 12, but the technology of the present disclosure is not limited to this, and the specific processing may be performed by the robot 414.
[1442] The emotion identification model 59 as an emotion engine may determine the user's emotion according to a specific mapping. Specifically, the emotion identification model 59 may determine the user's emotion according to an emotion map (see FIG. 9), which is a specific mapping. Similarly, the emotion identification model 59 may determine the robot's emotion, and the identification processing unit 290 may perform identification processing using the robot's emotion.
[1443] FIG. 9 is a diagram illustrating an emotion map 400 on which multiple emotions are mapped. In the emotion map 400, emotions are arranged in concentric circles radiating from the center. Emotions closer to the center of the concentric circles are more primitive. Emotions representing states and actions arising from a state of mind are arranged on the outer edges of the concentric circles. The concept of emotion includes both affect and mental states. Emotions generally generated from reactions occurring in the brain are arranged on the left side of the concentric circles. Emotions generally induced by situational judgment are arranged on the right side of the concentric circles. Emotions generally generated from reactions occurring in the brain and induced by situational judgment are arranged on the upper and lower sides of the concentric circles. Furthermore, the emotion of "pleasure" is arranged on the upper side of the concentric circles, and the emotion of "discomfort" is arranged on the lower side. In this way, in the emotion map 400, multiple emotions are mapped based on the structure by which emotions are generated, and emotions that tend to occur simultaneously are mapped close to each other.
[1444] These emotions are distributed in the 3 o'clock direction on emotion map 400, and typically fluctuate between relief and anxiety. In the right half of emotion map 400, situational awareness dominates over internal sensations, resulting in a sense of calm.
[1445] The inside of emotion map 400 represents what is going on in the mind, and the outside of emotion map 400 represents behavior, so the further you go outside emotion map 400, the more visible the emotions become (the more they are expressed in behavior).
[1446] Human emotions are based on various balances, such as posture and blood sugar levels. When these balances deviate from the ideal, a state of discomfort is indicated, and when they approach the ideal, a state of pleasure is indicated. Emotions can also be created for robots, automobiles, and motorcycles, based on various balances, such as posture and remaining battery life. When these balances deviate from the ideal, a state of discomfort is indicated, and when they approach the ideal, a state of pleasure is indicated. An emotion map can be generated, for example, based on Dr. Mitsuyoshi's emotion map (Research on Voice Emotion Recognition and Emotional Brain Physiological Signal Analysis Systems, Tokushima University, Doctoral Dissertation: https: / / ci.nii.ac.jp / naid / 500000375379). The left half of the emotion map lists emotions belonging to the "reaction" domain, where sensation is dominant. The right half of the emotion map lists emotions belonging to the "situation" domain, where situational awareness is dominant.
[1447] The emotion map defines two emotions that promote learning. One is a negative emotion on the situation side, around the middle of "repentance" or "reflection." In other words, this occurs when the robot experiences negative emotions such as "I never want to feel this way again" or "I don't want to be scolded again." The other is a positive emotion on the response side, around "desire." In other words, this occurs when the robot experiences positive feelings such as "I want more" or "I want to know more."
[1448] The emotion identification model 59 inputs user input into a pre-trained neural network, obtains emotion values indicating each emotion shown in the emotion map 400, and determines the user's emotion. This neural network is pre-trained based on multiple pieces of training data that are combinations of user input and emotion values indicating each emotion shown in the emotion map 400. Furthermore, this neural network is trained so that emotions that are located close to each other have similar values, as in the emotion map 900 shown in FIG. 10. FIG. 10 shows an example in which multiple emotions, "relieved," "calm," and "reassuring," have similar emotion values.
[1449] The system according to the present disclosure has been described above mainly with respect to the functions of the data processing device 12, but the system according to the present disclosure is not necessarily implemented on a server. The system according to the present disclosure may be implemented as a general information processing system. The present disclosure may be implemented, for example, as a software program running on a personal computer or an application running on a smartphone, etc. The method according to the present disclosure may be provided to users in the form of SaaS (Software as a Service).
[1450] In the above embodiment, an example was given in which the specific processing is performed by one computer 22, but the technology of the present disclosure is not limited to this, and the specific processing may be distributed and performed by a plurality of computers including the computer 22. For example, the data generation model 58 may be provided in an external device of the data processing device 12, and data may be generated in the external device in accordance with input data.
[1451] In the above embodiment, an example in which the specific processing program 56 is stored in the storage 32 has been described, but the technology of the present disclosure is not limited to this. For example, the specific processing program 56 may be stored in a portable, computer-readable, non-transitory storage medium such as a USB (Universal Serial Bus) memory. The specific processing program 56 stored in the non-transitory storage medium is installed in the computer 22 of the data processing device 12. The processor 28 executes the specific processing in accordance with the specific processing program 56.
[1452] Alternatively, the specific processing program 56 may be stored in a storage device such as a server connected to the data processing device 12 via the network 54, and the specific processing program 56 may be downloaded and installed on the computer 22 in response to a request from the data processing device 12.
[1453] It is not necessary to store all of the specific processing program 56 in a storage device such as a server connected to the data processing device 12 via the network 54, or to store all of the specific processing program 56 in the storage 32; only a portion of the specific processing program 56 may be stored.
[1454] The hardware resource for executing a specific process can be any of the following processors: An example of a processor is a CPU, which is a general-purpose processor that functions as a hardware resource for executing a specific process by executing software, i.e., a program. Another example of a processor is a dedicated electrical circuit, such as an FPGA (Field-Programmable Gate Array), a PLD (Programmable Logic Device), or an ASIC (Application Specific Integrated Circuit), which is a processor with a circuit configuration designed specifically for executing a specific process. Each processor has built-in or connected memory, and each processor uses the memory to execute the specific process.
[1455] The hardware resource that executes the specific processing may be configured with one of these various processors, or may be configured with a combination of two or more processors of the same or different types (for example, a combination of multiple FPGAs, or a combination of a CPU and an FPGA). Also, the hardware resource that executes the specific processing may be a single processor.
[1456] As an example of a system configured with a single processor, first, one processor is configured by combining one or more CPUs and software, and this processor functions as a hardware resource that executes a specific process. Second, there is a system that uses a processor that realizes the functions of an entire system including multiple hardware resources that execute a specific process on a single IC chip, as typified by SoC (System-on-a-chip). In this way, a specific process is realized using one or more of the above-mentioned various processors as hardware resources.
[1457] Furthermore, the hardware structure of these various processors can be, more specifically, an electric circuit that combines circuit elements such as semiconductor devices. The specific processing described above is merely an example. Therefore, it goes without saying that unnecessary steps may be deleted, new steps may be added, or the processing order may be rearranged, without departing from the spirit of the invention.
[1458] The above-described description and illustrations are a detailed explanation of the parts related to the technology of the present disclosure and are merely an example of the technology of the present disclosure. For example, the above description of the configuration, functions, actions, and effects is an explanation of an example of the configuration, functions, actions, and effects of the parts related to the technology of the present disclosure. Therefore, it goes without saying that unnecessary parts may be deleted, new elements may be added, or replacements may be made to the above-described description and illustrations within the scope of the gist of the technology of the present disclosure. Furthermore, to avoid confusion and facilitate understanding of the parts related to the technology of the present disclosure, the above-described description and illustrations omit explanations of common technical knowledge that do not require particular explanation to enable the implementation of the technology of the present disclosure.
[1459] All publications, patent applications, and technical standards mentioned in this specification are herein incorporated by reference to the same extent as if each individual publication, patent application, or technical standard was specifically and individually indicated to be incorporated by reference.
[1460] The following is further disclosed regarding the above embodiment.
[1461] (Claim 1)
[1462] a means for recording the user's voice;
[1463] means for transmitting the recorded voice data to a server;
[1464] analysis means for analyzing the received voice data, the analysis means including a voice authentication model;
[1465] means for assessing an authentication score for the analyzed voice data;
[1466] A system including a means for unlocking a security lock based on an assessed authentication score.
[1467] (Claim 2)
[1468] 10. The system of claim 1, further comprising means for providing an interface for a user to start and stop recording.
[1469] (Claim 3)
[1470] 2. The system according to claim 1, further comprising a notification means for notifying the user of the result of the voice authentication.
[1471] "Example 1"
[1472] (Claim 1)
[1473] a means for recording the user's voice;
[1474] means for transmitting the recorded voice data to a server;
[1475] An analysis means having an AI voice authentication model that analyzes the received voice data;
[1476] means for assessing an authentication score for the analyzed voice data;
[1477] A means for unlocking security based on the assessed authentication score;
[1478] A system including means for notifying a user of the results of voice authentication.
[1479] (Claim 2)
[1480] 10. The system of claim 1, further comprising means for providing an interface for a user to start and stop recording.
[1481] (Claim 3)
[1482] 10. The system of claim 1, further comprising means for transmitting the audio data in an HTTP POST request.
[1483] "Application Example 1"
[1484] (Claim 1)
[1485] a means for recording the user's voice;
[1486] means for transmitting the recorded voice data to a server;
[1487] analysis means for analyzing the received voice data, the analysis means including a voice authentication model;
[1488] means for assessing an authentication score for the analyzed voice data;
[1489] A means for unlocking security based on the assessed authentication score;
[1490] a display means including operation buttons for recording and transmitting voice;
[1491] A computer system including a means for notifying the user of the results.
[1492] (Claim 2)
[1493] 2. The computer system according to claim 1, further comprising means for providing a UI for a user to operate start and stop of recording.
[1494] (Claim 3)
[1495] 10. The computer system of claim 1, further comprising a notification interface for displaying the results of the voice authentication to a user.
[1496] "Example 2: Combining Emotion Engines"
[1497] (Claim 1)
[1498] a means for recording the user's voice;
[1499] means for transmitting the recorded voice data to a server;
[1500] analysis means for analyzing the received voice data, the analysis means including a voice authentication model;
[1501] an analysis means having an emotion recognition engine that analyzes emotions from voice data;
[1502] means for evaluating the authentication score and emotion data of the analyzed voice data;
[1503] A means for unlocking a security lock based on the evaluated authentication score and a comprehensive score of the emotion data;
[1504] a means for notifying the user of the authentication result;
[1505] A system including:
[1506] (Claim 2)
[1507] 2. The system according to claim 1, further comprising means for providing an interface for a user to operate start and stop of recording.
[1508] (Claim 3)
[1509] 10. The system of claim 1, further comprising a notification means for notifying a user of the results of the voice authentication and emotion recognition.
[1510] "Application example 2 when combining emotion engines"
[1511] (Claim 1)
[1512] a means for recording the user's voice;
[1513] means for transmitting the recorded voice data to a server;
[1514] analysis means for analyzing the received voice data, the analysis means including a voice authentication model;
[1515] means for evaluating the authentication score and emotion data of the analyzed voice data;
[1516] The system includes a means for calculating a total score based on the evaluated authentication score and emotion data, and for releasing the security lock based on the total score.
[1517] (Claim 2)
[1518] 2. The system according to claim 1, further comprising means for providing an interface for a user to operate start and stop of recording, and for inputting recorded voice data into an emotion recognition engine to obtain emotion data of the user.
[1519] (Claim 3)
[1520] 10. The system according to claim 1, further comprising a notification means for notifying a user of the results of the voice authentication and the emotion recognition. [Explanation of symbols]
[1521] 10, 210, 310, 410 Data Processing Systems 12 Data Processing Device 14 Smart Devices 214 Smart Glasses 314 Headset-type terminal 414 Robot< / url:> < / url:> < / url:> < / url:>
Claims
1. a means for recording the user's voice; means for transmitting the recorded voice data to a server; analysis means for analyzing the received voice data, the analysis means including a voice authentication model; means for assessing an authentication score for the analyzed voice data; A system including a means for unlocking a security lock based on an assessed authentication score.
2. 2. The system according to claim 1, further comprising means for providing an interface for a user to operate start and stop recording.
3. 2. The system according to claim 1, further comprising a notification means for notifying the user of the result of the voice authentication.
Citation Information
Patent Citations
Persona chatbot control method and system
JP2022180282A