Method and system for providing high-dimensional payment solution using generative artificial intelligence-based voiceprint authentication technology

The system uses generative AI-based voice fingerprint authentication with voice pattern and keyword verification to address security vulnerabilities in contactless payments, ensuring secure and convenient transactions for all users.

WO2026095640A1PCT designated stage Publication Date: 2026-05-07NEXTPAYMENTS INC
View PDF 6 Cites 0 Cited by

Patent Information

Authority / Receiving Office
WO · WO
Patent Type
Applications
Current Assignee / Owner
NEXTPAYMENTS INC
Filing Date
2025-10-29
Publication Date
2026-05-07

AI Technical Summary

Technical Problem

Existing contactless payment methods, such as facial or voice recognition, pose security vulnerabilities and are difficult for visually impaired individuals to use, and general users face security threats when using touchscreens.

Method used

A system utilizing generative AI-based voice fingerprint authentication that requires users to pronounce a pre-set password, combining voice pattern matching and keyword verification to ensure secure and touch-free payments.

Benefits of technology

Provides high security and convenience by accurately verifying user identity through voice patterns, preventing attacks from recorded or forged voices, and enabling safe payments for both visually impaired and general users.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure KR2025017504_07052026_PF_FP_ABST
    Figure KR2025017504_07052026_PF_FP_ABST
Patent Text Reader

Abstract

This payment solution provision system using generative artificial intelligence (AI)-based voiceprint authentication may comprise a memory and a processor. The processor receives and preprocesses a voice input of a user, and then extracts a voice pattern through generative AI, and compares same with a pre-registered pattern so as to determine matching thereof. If there is matching, a set password pronunciation is requested from the user, and the pitch, speed and tone of the corresponding pronunciation are analyzed to check whether the user is a registered user. If it is determined that both the voice pattern and the password pronunciation match, a payment request is processed, and a payment completion message can be controlled to be transmitted to a user terminal and a payment platform.
Need to check novelty before this filing date? Find Prior Art

Description

Method and system for providing a high-level payment solution utilizing generative AI-based voice fingerprint authentication technology

[0001] The present invention relates to biometric authentication technology, and more specifically, to a high-level payment solution that combines Generative AI-based voice fingerprints and voice keywords to enable users to authenticate their identity and perform payments without touch. This technology targets both visually impaired and sighted individuals and provides security and convenience that go beyond existing simple facial recognition technologies.

[0002] Existing contactless payment methods are based on facial or voice recognition, but these technologies have security vulnerabilities or limitations in user experience. In particular, visually impaired individuals may find it difficult to perform simple facial recognition or enter passwords, and even general users may be exposed to security threats while using touchscreens.

[0003] Voiceprint is a technology that authenticates identity by analyzing an individual's unique voice patterns, and it can provide higher accuracy and security through generative AI. Additionally, performing secondary authentication by entering keywords via voice can enhance security while improving the user experience.

[0004] The objective of the present invention is to provide a system that enables users to perform safe and rapid payments without touch by utilizing generative AI-based voice fingerprints and voice keywords. In particular, it is designed to be easily used by both the visually impaired and the general public, with a focus on maximizing user convenience while maintaining high security.

[0005] A system providing a high-level payment solution utilizing generative artificial intelligence (AI)-based voice fingerprint authentication technology may include memory and a processor for storing instructions. When the instructions are executed by the processor, the system receives a user's voice input and performs preprocessing; provides the preprocessed voice data to the generative AI to extract a voice pattern; compares the extracted voice pattern with a voice pattern of a user registered in advance to verify a match; displays a notification requesting the user to pronounce a pre-set password based on the fact that the extracted voice pattern matches the voice pattern of the registered user; receives the user's pronunciation and analyzes the pitch, speed, and timbre of the voice; determines whether the voice belongs to the registered user based on the analysis results; processes a payment request based on the fact that the extracted voice pattern matches the voice pattern of the registered user and the pronunciation of the pre-set password is determined to belong to the registered user; and controls the system to transmit a message indicating that the payment has been completed to the user terminal and the payment platform in real time.

[0006] When a user pronounces a command such as "Please authenticate by speech for payment verification," the system collects the corresponding voice in real time and performs a preprocessing step. For example, it removes background noise and refines the voice signal.

[0007] Generative AI analyzes preprocessed voice data to extract the user's unique voice pattern. This pattern is compared with a user voice fingerprint registered in advance, and if they match, the first authentication is successfully completed.

[0008] The result of the first verification is immediately provided to the user, allowing them to proceed to the next step. For example, a voice message saying, "First verification is complete. Please continue," is provided.

[0009] Once the first authentication is complete, the system requests that the user pronounce a pre-set voice password. This password is not a simple word, but consists of a unique voice pattern that includes the user's voice fingerprint.

[0010] Generative AI analyzes the input voice password to verify whether the patterns of the registered voice fingerprint and keywords match. In this process, the unique pitch, pronunciation speed, and timbre of the voice are taken into account.

[0011] The double-protected voice authentication process determines not only whether the keywords pronounced by the user are correct but also whether the voice belongs to the actual user, thereby effectively blocking attacks using recorded or forged voices.

[0012] Once both primary and secondary authentication are complete, the payment authorization module operates to process the payment request. During this process, the user's payment information is protected through advanced encryption technology, and the transaction is completed securely.

[0013] The present invention provides a system that enables both visually impaired and general users to safely perform identity authentication and payment using only voice without touch.

[0014] Advanced voice fingerprint and keyword authentication utilizing generative AI provides high security, enabling the implementation of payment solutions that enhance security while improving the user experience through a dual authentication process.

[0015] FIG. 1 is a diagram illustrating a system that provides a high-dimensional payment solution utilizing generative artificial intelligence-based voice fingerprint authentication technology according to one embodiment.

[0016] FIG. 2 is a block diagram showing the configuration of a system providing a high-dimensional payment solution utilizing generative artificial intelligence (AI)-based voice fingerprint authentication technology according to one embodiment.

[0017] FIG. 3 is a flowchart illustrating a method for providing a high-dimensional payment solution utilizing generative artificial intelligence (AI)-based voice fingerprint authentication technology according to one embodiment.

[0018] FIG. 4 is a flowchart illustrating a method for providing a high-dimensional payment solution utilizing generative artificial intelligence (AI)-based voice fingerprint authentication technology according to one embodiment.

[0019] FIG. 5 is a flowchart illustrating a method for providing a high-dimensional payment solution utilizing generative artificial intelligence (AI)-based voice fingerprint authentication technology according to one embodiment.

[0020] Hereinafter, embodiments are described in detail with reference to the attached drawings. However, various modifications may be made to the embodiments, and thus the scope of the patent application is not limited or restricted by these embodiments. It should be understood that all modifications, equivalents, and substitutions to the embodiments are included within the scope of the rights.

[0021] Specific structural or functional descriptions of the embodiments are disclosed for illustrative purposes only and may be modified and implemented in various forms. Accordingly, the embodiments are not limited to the specific disclosed forms, and the scope of this specification includes modifications, equivalents, or substitutions that fall within the technical concept.

[0022] Terms such as "first" or "second" may be used to describe various components, but these terms should be interpreted solely for the purpose of distinguishing one component from another. For example, the first component may be named the second component, and similarly, the second component may be named the first component.

[0023] When it is stated that a component is "connected" to another component, it should be understood that it may be directly connected to or joined to that other component, or that there may be other components in between.

[0024] The terms used in the embodiments are for illustrative purposes only and should not be interpreted as intended to be limiting. Singular expressions include plural expressions unless the context clearly indicates otherwise. In this specification, terms such as "comprising" or "having" are intended to indicate the existence of the features, numbers, steps, actions, components, parts, or combinations thereof described in the specification, and should be understood as not precluding the existence or addition of one or more other features, numbers, steps, actions, components, parts, or combinations thereof.

[0025] Unless otherwise defined, all terms used herein, including technical or scientific terms, have the same meaning as generally understood by those skilled in the art to which the embodiments pertain. Terms such as those defined in commonly used dictionaries should be interpreted as having a meaning consistent with their meaning in the context of the relevant technology, and should not be interpreted in an ideal or overly formal sense unless explicitly defined in this application.

[0026] In addition, when describing with reference to the attached drawings, identical components are assigned the same reference numeral regardless of drawing symbols, and redundant descriptions thereof are omitted. In describing the embodiments, if it is determined that a detailed description of related prior art could unnecessarily obscure the essence of the embodiments, such detailed description is omitted.

[0027] The embodiments can be implemented in various forms of products such as personal computers, laptop computers, tablet computers, smartphones, televisions, smart home appliances, intelligent automobiles, kiosks, and wearable devices.

[0028]

[0029] FIG. 1 is a diagram illustrating a system for providing a high-dimensional payment solution utilizing generative artificial intelligence-based voice fingerprint authentication technology according to one embodiment.

[0030] As illustrated in FIG. 1, a system (100) providing a high-level payment solution utilizing generative AI-based voice fingerprint authentication technology may include a plurality of user terminals (110-1,…), a server (120), and a database (130). According to one embodiment, the database (130) is depicted as being configured separately from the server (120), but is not limited thereto, and the database (130) may be provided within the server (120). For example, the server (120) may include a plurality of AIs for performing machine learning algorithms. According to one embodiment, the plurality of user terminals (110-1,…), the server (120), and the database (130) may be connected to communicate with each other through a network (N).

[0031] A network (N) can perform wireless or wired communication between multiple user terminals (110-1,…), a server (120), a database (130), etc. For example, the network can perform wireless communication according to methods such as LTE (long-term evolution), LTE-A (LTE Advanced), CDMA (code division multiple access), WCDMA (wideband CDMA), WiBro (Wireless BroadBand), WiFi (wireless fidelity), Bluetooth, NFC (near field communication), GPS (Global Positioning System), or GNSS (global navigation satellite system). For example, the network (N) can perform wired communication according to methods such as USB (universal serial bus), HDMI (high definition multimedia interface), RS-232 (recommended standard 232), or POTS (plain old telephone service).

[0032] The database (130) can store various data. Data stored in the database (130) may include software (e.g., programs) as data acquired, processed, or used by a plurality of user terminals (110-1,…) and at least one component of the server (120). The database (130) may include volatile and / or non-volatile memory.

[0033] In the present invention, Artificial Intelligence (AI) refers to a technology that imitates human learning ability, reasoning ability, and perceptual ability, and implements them on a computer, and may include concepts such as machine learning and symbolic logic. Machine Learning (ML) is an algorithmic technology that classifies or learns the characteristics of input data on its own. AI technology can analyze input data as a machine learning algorithm, learn from the results of the analysis, and make judgments or predictions based on the results of the learning. Furthermore, technologies that mimic the functions of the human brain, such as cognition and judgment, by utilizing machine learning algorithms can also be understood as falling within the category of AI. For example, technological fields such as linguistic understanding, visual understanding, reasoning / prediction, knowledge representation, and motion control may be included.

[0034] Machine learning can refer to the process of training neural network models using experience in processing data. It implies that through machine learning, computer software improves its own data processing capabilities. A neural network model is constructed by modeling the correlations between data, and these correlations can be expressed by multiple parameters. A neural network model extracts and analyzes features from given data to derive correlations between them; machine learning can be defined as the process of optimizing the model's parameters by repeating this process. For example, a neural network model can learn the mapping (correlation) between inputs and outputs for data given as input-output pairs. Alternatively, even when only input data is provided, a neural network model can derive regularities between the given data and learn those relationships.

[0035] An artificial intelligence learning model or neural network model can be designed to implement the structure of the human brain on a computer and may include multiple network nodes that have weights and simulate neurons of a human neural network. The multiple network nodes may have interconnected relationships by simulating the synaptic activity of neurons, where neurons exchange signals through synapses. In an artificial intelligence learning model, multiple network nodes may be located in layers of different depths and exchange data according to convolutional connections. The artificial intelligence learning model may be, for example, an Artificial Neural Network (ANN) or a Convolutional Neural Network (CNN). As an embodiment, the artificial intelligence learning model may be machine learned according to methods such as supervised learning, unsupervised learning, and reinforcement learning. Machine learning algorithms for performing machine learning may include Decision Tree, Bayesian Network, Support Vector Machine, Artificial Neural Network, Ada-boost, Perceptron, Genetic Programming, and Clustering.

[0036] Among these, CNNs are a type of multilayer perceptron designed to use minimal preprocessing. CNNs consist of one or more convolutional layers and standard artificial neural network layers stacked on top, additionally utilizing weights and pooling layers. Thanks to this structure, CNNs can fully utilize two-dimensional input data. Compared to other deep learning architectures, CNNs demonstrate good performance in both image and audio fields. CNNs can also be trained using standard backpropagation. CNNs have the advantage of being easier to train than other feedforward artificial neural network techniques and using a small number of parameters.

[0037] Convolutional networks are neural networks comprising sets of nodes with bounded parameters. Many computer vision tasks have been significantly improved, driven by the increased size of available training data and the availability of computational power, combined with algorithmic advancements such as discriminative linear units and dropout training. In the case of massive datasets, such as those available for many tasks today, outfitting is not critical, and increasing the network size improves test accuracy. Optimal utilization of computing resources becomes a limiting factor. To address this, distributed, scalable implementations of deep neural networks can be employed.

[0038] The learning device can train a neural network to process review responses received from multiple user terminals (110-1,…) by item. Additionally, the learning device can train a neural network to extract user dwell history from user movement path information. According to one embodiment, the learning device may be a separate entity from the server (120), but is not limited thereto.

[0039] A neural network includes an input layer into which training samples are input and an output layer that outputs training outputs, and can be trained based on the difference between the training outputs and the labels. Here, the labels are defined based on items corresponding to review responses and can be defined based on user dwell history corresponding to movement path information. The neural network is connected as a group of multiple nodes and is defined by weights between the connected nodes and activation functions that activate the nodes.

[0040] The learning device can train a neural network using the Gradient Descent (GD) technique or the Stochastic Gradient Descent (SGD) technique. The learning device can use a loss function designed based on the outputs and labels of the neural network.

[0041] The learning device can calculate the training error using a predefined loss function. The loss function can be predefined with labels, outputs, and parameters as input variables, where the parameters can be set by weights within the neural network. For example, the loss function can be designed in the form of Mean Square Error (MSE), entropy, etc., and various techniques or methods may be employed in the embodiments in which the loss function is designed.

[0042] The learning device can identify weights that influence the training error using the backpropagation technique. Here, weights are relationships between nodes within a neural network. The learning device can use the SGD technique with labels and outputs to optimize the weights found through the backpropagation technique. For example, the learning device can update the weights of a loss function defined based on labels, outputs, and weights using the SGD technique.

[0043] According to one embodiment, the learning device extracts first objects from review responses, obtains first labels which are items corresponding to the first objects, applies the first objects to a first neural network to generate first training outputs corresponding to the first objects, and can train the first neural network based on the first training outputs and the first labels.

[0044] The learning device extracts second objects from movement path information, obtains second labels which are user dwell history corresponding to the second objects, applies the second objects to a second neural network to generate second training outputs corresponding to the second objects, and can train the second neural network based on the second training outputs and second labels.

[0045] According to one embodiment, the learning device can generate first training feature vectors based on the constituent features, location features, and pattern features of the review response. Various methods may be employed for extracting features.

[0046] According to one embodiment, the learning device can generate second training feature vectors based on the constituent features, length features, and pattern features of the movement path information. Various methods may be employed for extracting features.

[0047] According to one embodiment, the learning device can obtain training outputs by applying first training feature vectors to a neural network. The learning device can train a review item extraction algorithm of the neural network based on the training outputs and first labels. The learning device can train a review item extraction algorithm of the neural network by calculating training errors corresponding to the training outputs and optimizing the connection relationships of nodes within the neural network to minimize the training errors. The server (120) can extract items from review responses using the first neural network that has completed training. For example, the extracted items may include, but are not limited to, friendliness, store cleanliness, service satisfaction, etc.

[0048] According to one embodiment, the learning device can obtain training outputs by applying second training feature vectors to a neural network. The learning device can train a user stay history acquisition algorithm of the neural network based on the training outputs and second labels. The learning device can train a user stay history acquisition algorithm of the neural network by calculating training errors corresponding to the training outputs and optimizing the connection relationships of nodes within the neural network to minimize the training errors. The server (120) can obtain user stay history from movement path information using the second neural network that has been trained.

[0049] An artificial intelligence model according to one embodiment may include an input layer, a hidden layer, and an output layer.

[0050] The input layer is the layer associated with the input values ​​fed into the artificial intelligence model.

[0051] In the hidden layer, a feature map can be output by performing MAC (multiply-accumulate) and activation operations on the input values.

[0052] A MAC operation can be an operation that multiplies each input value by its corresponding weight and sums the multiplied values.

[0053] The activation operation may be an operation that inputs the result of the MAC operation into an activation function and outputs a result value. The activation function may be of various types. For example, the activation function may include a sigmoid function, a tangent function, a ReLU function, a Leaky ReLU function, a Max Out function, and / or an ELU function, but there are no restrictions on the types.

[0054] A hidden layer may consist of at least one layer. For example, if the hidden layer consists of a first hidden layer and a second hidden layer, the first hidden layer performs MAC operations and activation operations based on input values ​​of an input system to output a feature map, and the feature map, which is the result value of the first hidden layer, may become the input value of the second hidden layer. The second hidden layer may perform MAC operations and activation operations based on the feature map, which is the result value of the first hidden layer.

[0055] The output layer may be a layer associated with the result of an operation performed in the hidden layer.

[0056] In one embodiment, the learning model learns syllable (character) patterns that are frequently combined and used in a given corpus to automatically learn the boundaries of compound words and named entities, integrates object information from a first UI source with object information rendered in a browser to create a learning object information file, uses the learning object information file to create learning data for the learning of a deep learning network, receives data from various domains of a support system, standardizes the data from the various domains into an integrated format based on at least one standardization method corresponding to each of the various domains, learns and infers data from a specific domain, determines information to be transmitted for standardization from the data of the specific domain, and can perform post-processing on the data from the various domains. The first UI source includes an XML file, and the learning object information file includes an input JSON file for learning features and an output JSON file that is the correct answer (label) data during learning, and the output JSON file includes a file containing HTML DOM Tree information implemented in compliance with web standards, and the various domains include at least one of a RAN (radio access network), a transport, or a core, and the post-processing may include a correlation function.

[0057]

[0058] FIG. 2 is a block diagram showing the configuration of a system providing a high-dimensional payment solution utilizing generative artificial intelligence (AI)-based voice fingerprint authentication technology according to one embodiment.

[0059] A system (200) according to one embodiment may include a processor (220) and a memory (230), and some of the illustrated components may be omitted or substituted. A system (200) according to one embodiment may be a server or a terminal. According to one embodiment, the processor (220) is a component capable of performing operations or data processing regarding the control and / or communication of each component of the system (200), and may be composed of one or more processors. The memory (230) may store information related to the method described above or store a program in which the method described above is implemented. The memory (230) may be volatile memory or non-volatile memory. The memory (230) may store various file data, and the stored file data may be updated according to the operation of the processor (220).

[0060] According to one embodiment, the processor (220) can execute a program and control the device (400). The code of the program executed by the processor (220) can be stored in memory (230). Operations of the processor (220) can be performed by loading instructions stored in memory (230). The system (200) can be connected to an external device (e.g., a personal computer or a network) through an input / output device (not shown in the drawing) and exchange data.

[0061] According to one embodiment, there are no limitations on the computation and data processing functions that the processor (220) can implement on the system (200), but below, a function providing a high-level payment solution utilizing generative artificial intelligence (AI)-based voice fingerprint authentication technology will be described.

[0062]

[0063] FIG. 3 is a flowchart illustrating a method for providing a high-dimensional payment solution utilizing generative artificial intelligence (AI)-based voice fingerprint authentication technology according to one embodiment.

[0064] Although process steps, method steps, algorithms, etc. are described in a sequential order in the flowchart of FIG. 3, such processes, methods, and algorithms may be configured to operate in any suitable order. In other words, the steps of the processes, methods, and algorithms described in various embodiments of the present invention do not need to be performed in the order described in the present invention. Furthermore, even if some steps are described as being performed asynchronously, in other embodiments, such steps may be performed simultaneously. Also, the example of a process by the depiction in the drawings does not mean that the exampled process excludes other variations and modifications therefrom, does not mean that any of the exampled process or any of its steps are essential to one or more of the various embodiments of the present invention, and does not mean that the exampled process is desirable.

[0065] In operation 310, the system (e.g., the system (200) of FIG. 2) can receive voice input from a user and perform preprocessing under the control of a processor (e.g., the processor (220) of FIG. 2).

[0066] For example, when a user inputs "Please make a payment" in voice, the system undergoes preprocessing steps such as removing background noise from the voice data and amplifying the voice signal. For example, when a user says "Please speak for payment authentication," the system (200) receives this voice input. Subsequently, a preprocessing step is performed to ensure that the voice is not affected by noise, and in this process, background noise is removed and the voice is converted into a clean form. In this way, through preprocessing, the system clearly prepares the voice data.

[0067] In operation 320, the system (200) can provide preprocessed voice data to a generative artificial intelligence to extract and compare voice patterns. The system (200) provides preprocessed voice data to a generative artificial intelligence to extract voice patterns, and can compare the extracted voice patterns with the voice patterns of a user registered in advance to check for a match.

[0068] Generative AI analyzes features such as the user's voice tone, pronunciation habits, and voice frequency from preprocessed voice data to extract voice patterns. It then compares these extracted patterns with pre-registered voice patterns of the user to determine if they match. For example, if patterns such as "high tone, fast pronunciation speed, and specific frequency stress" are stored in User A's voice pattern database, it determines whether they match by comparing them with the patterns extracted from the input voice data.

[0069] Preprocessed voice data is provided to generative AI, which is used to extract unique voice patterns. For example, if a user's voice contains the password "1234," the generative AI analyzes the voice to generate a unique voice pattern based on the frequency, pitch, and pronunciation speed. This pattern is then compared with the user's previously registered voice patterns to check for a match.

[0070] In operation 330, the system (200) may display a notification requesting the pronunciation of a preset password. The system (200) may display a notification requesting the pronunciation of a preset password based on the fact that the extracted voice pattern matches the voice pattern of the registered user.

[0071] If the system determines, based on the voice pattern analysis results, that the input voice matches the voice pattern of User A, it displays a notification on the user's terminal saying, "Please pronounce the pre-set password for user authentication." For example, the system (200) delivers a voice notification to the user saying, "Please pronounce the password." The user understands the system's request and pronounces, "The password is 1234." This notification guides the user to the next steps required to proceed with the authentication process.

[0072] In operation 340, the system (200) can receive the user's pronunciation and analyze the pitch, speed, and timbre of the speech. The system (200) can receive the user's pronunciation and analyze the pitch, speed, and timbre of the speech and determine, based on the analysis results, that the speech belongs to a registered user.

[0073] When a user pronounces the password "1234," the system analyzes the pitch, pronunciation speed, and timbre (the degree of lightness or darkness of the voice) from this voice data. Based on the analysis results, it checks for a match by comparing them with the password pronunciation data from when User A registered. For example, if User A has a pattern of pronouncing "1234" with "low pitch, slow pronunciation speed, and dark timbre," it determines whether it is indeed User A by comparing it with the input pronunciation data.

[0074] The system (200) receives the content pronounced by the user and analyzes the pitch (highness), speed (how fast it was spoken), and timbre (characteristics of the voice) of the pronunciation. For example, when the user pronounces "1234," the system analyzes the characteristics of the pronunciation and compares it with the voice of a user registered in the past. If the analyzed voice is similar to the characteristics of a registered user, the system determines that the voice belongs to that user.

[0075] In operation 350, the system (200) can process a payment request. The system (200) processes the payment request based on the fact that the extracted voice pattern matches the voice pattern of the registered user and that the pronunciation of the pre-set password is determined to be that of the registered user, and can send a message indicating that the payment has been completed to the user terminal and the payment platform in real time.

[0076] The system processes a payment request if user authentication is successful by combining the voice pattern comparison results and the password pronunciation analysis results. When the payment is completed, a message stating "Payment has been completed" is sent in real time to the user's smartphone and the payment platform (e.g., an online shopping mall). To make the payment, the system (200) proceeds to the step of processing the payment request when the user's voice is confirmed. For example, if the user pronounces "The password is 1234" and this pronunciation matches the registered voice, the system processes the payment based on the user's payment information (e.g., card information). Then, a message stating that the payment has been successfully completed is sent in real time to the user's terminal and the payment platform, so that the user immediately receives a payment completion notification.

[0077] In this manner, each operation proceeds sequentially, and the system provides safe and efficient payment authentication based on the user's voice.

[0078]

[0079] According to one embodiment, the system (200) can be controlled to perform voice pattern comparison using a Dynamic Time Warping (DTW) algorithm, extract voice patterns based on time-frequency analysis of a voice signal using generative artificial intelligence, and perform liveness verification to prevent fake attacks. The preprocessing process may be characterized by including steps of noise removal, normalization, and feature extraction of the voice signal, and a step of extracting Mel-Frequency Cepstral Coefficients (MFCC).

[0080] For example, when a user pronounces "Hello," the system (200) uses a DTW algorithm to compare the registered voice with the real-time voice. The DTW corrects temporal distortion of the voice pattern to measure the similarity between the two signals. Generative artificial intelligence analyzes this voice to extract features in the time-frequency domain and detects changes in specific frequencies. Additionally, to verify whether the user's voice actually occurred through liveness verification, the system checks if the voice is real by requiring the user to repeat a specific sound during pronunciation or by requesting various pronunciations.

[0081] When user A pronounces "payment," even if the pronunciation time becomes slightly longer or shorter, the DTW algorithm is used to accurately compare the similarity with A's previously registered "payment" pronunciation pattern to determine whether they match.

[0082] Generative AI analyzes speech signals by dividing them into the time domain and the frequency domain. For example, it analyzes which frequency components appear strongly over time in the pronunciation of "payment" to extract a unique pronunciation pattern specific to User A. This enables more accurate voice authentication than simply comparing voice waveforms.

[0083] It verifies that the voice is that of an actual user, not a recorded or synthesized voice, by having the user repeat specific words randomly or read specific sentences. For example, if the system presents the sentence "The weather is really nice today," the user must repeat this sentence in real time.

[0084] MFCC features that reflect human auditory characteristics are extracted from the voice signal and used for voice pattern analysis. Since MFCCs represent voice data in a manner similar to how humans perceive sound, they reduce the influence of background noise or emotional changes, thereby enabling more accurate voice authentication. For example, when a user says "Please make a payment," the system (200) first performs noise removal to filter out background noise (e.g., the sound of people talking). Then, the voice signal is normalized to adjust the volume to a constant level, and MFCCs are extracted to convert the frequency components of the voice into numerical values. MFCCs are widely used features in voice recognition and play an important role in uniquely representing the user's voice.

[0085]

[0086] According to one embodiment, the system (200) controls the output of a message saying "Please authenticate by speaking for payment authentication" before receiving the user's voice input, checks whether the extracted voice pattern matches the voice pattern of the registered user, and then feeds the result to the user terminal in real time. When receiving the user's pronunciation and analyzing the pitch, pronunciation speed, and timbre of the voice, the system performs the analysis using generative artificial intelligence, and when processing a payment request, it can control the system to support one or more payment methods among credit cards, electronic money, and mobile payments.

[0087] For example, when a user attempts to proceed with payment, the system (200) outputs a voice message saying, "Please authenticate by speaking for payment authentication" to guide the user through the voice recognition procedure. This message encourages the user to understand the system's request and proceed to the next step. The user hears this message and prepares to input voice. The system (200) analyzes what the user says, compares it with a registered pattern, and immediately delivers a message such as "Authentication successful" or "Authentication failed" to the user terminal. For example, when the user says, "My password is 1234," the system analyzes this and, if it matches, outputs a message saying "Authentication successful" so that the user can proceed to the next step.

[0088] Generative AI analyzes pitch, speed, and timbre from a user's password pronunciation and compares them to the pronunciation patterns of previously registered users. For example, even if User A has a cold and their voice sounds different than usual, the AI ​​can detect this change and authenticate that it is User A.

[0089] The system (200) receives the content pronounced by the user and analyzes how high the user's voice is by measuring the pitch of the voice through generative artificial intelligence. In addition, it analyzes the pronunciation speed (e.g., how many words were pronounced per minute) and the timbre (e.g., whether the voice is soft or rough) and comprehensively evaluates all this information. For example, if the user pronounces "1234" quickly, the system verifies this and determines that the user is an actual registered user.

[0090] When payment is approved, the system (200) processes the payment according to the payment method selected by the user. For example, if the user selects a credit card, the system verifies the card information and proceeds with the payment. If the user selects electronic money, the system deducts the amount from the electronic money wallet. In addition, if mobile payment is selected, the system processes the payment by linking with a mobile payment platform. In this process, the user can choose a convenient method among various payment methods.

[0091] Users can make payments using various payment methods, such as credit cards, electronic money (e.g., T-money), and mobile payments (e.g., Kakao Pay), through voice authentication.

[0092]

[0093] According to one embodiment, the system (200) controls all voice fingerprint data and payment information to be encrypted using the AES-256 encryption method, monitors all authentication and payment activities in real time and stores logs to detect abnormal activities, and uses generative artificial intelligence to analyze fine features of the voice pronounced by the user to block attacks using recorded or synthesized voice.

[0094] Important information, such as users' voice pattern data, credit card details, and payment history, is securely stored using AES-256 encryption. This prevents hackers from stealing or exploiting information even if they infiltrate the database.

[0095] For example, voice fingerprint data and payment information entered by the user are encrypted using the AES-256 encryption algorithm and stored securely. In this process, the system (200) encrypts all data to protect it from external attacks. For example, even if a hacker attempts to access the database, the encrypted data cannot be decrypted, thus keeping user information safe.

[0096] The system (200) observes all authentication and payment activities of the user in real time and, for example, immediately issues a warning if an abnormal pattern is detected. If the user attempts to make a payment at an unusual time or if multiple authentication failures occur, the system records this and sends a warning message to the administrator so that additional measures can be taken.

[0097] The system (200) analyzes the user's voice and detects subtle differences in the voice (e.g., subtle changes in intonation of pronunciation). For example, if someone attempts authentication by playing a recorded voice, the system detects that the voice is different from the actual user and rejects the authentication. In this process, the generative artificial intelligence analyzes the user's unique pronunciation pattern to effectively block such attacks.

[0098] The system monitors all user authentication attempts and payment activities in real time and stores log data. For example, if multiple authentication failures occur within a short period, the system detects this as abnormal activity and temporarily locks the corresponding account to enhance security. Generative AI analyzes even the most minute features of a user's voice to distinguish between recorded and synthesized speech. For instance, it determines whether the voice is that of the actual user by analyzing subtle noises, breathing sounds, and pronunciation tremors that are difficult for humans to detect.

[0099]

[0100] According to one embodiment, the system (200) detects ambient noise when receiving voice input from a user, requests re-input from the user if the noise level exceeds a preset threshold, analyzes the emotional state when extracting the user's voice pattern, performs an additional verification procedure if stress or coercive situation is detected, receives input through multiple microphones when receiving voice input from a user, verifies the actual location of the user by analyzing the direction and distance of the voice, detects abnormal transactions by comparing and analyzing the user's usual payment pattern and current payment details when processing a payment request, and controls to require additional authentication if necessary.

[0101] When a user attempts to make a payment by voice on a noisy roadside, the system detects ambient noise, and if it determines that the noise level is too high, it displays a message saying, "The surrounding noise is too loud. Please try again in a quiet place," and requests voice re-entry. If emotions such as anxiety, tension, or fear are detected in the user's voice, the system blocks the attempt at payment by coercion by presenting additional questions (e.g., "What is your mother's name?") along with a message saying, "Please answer additional questions for identity verification."

[0102] It receives voice input through multiple microphones built into the smartphone and identifies the actual user's location by analyzing the direction and distance of the voice. For example, it prevents fraudulent payments using a stolen smartphone by allowing authentication only when the user's voice is heard within 1 meter of the smartphone and input from the front. If a user who usually makes purchases of 100,000 won or less at an online shopping mall suddenly attempts to purchase a high-priced item exceeding 1 million won, the system detects this as an abnormal transaction and requests additional authentication. For instance, it strengthens identity verification by displaying a message such as, "Please complete fingerprint authentication."

[0103] For example, when the user pronounces "Please say for payment authentication," the system (200) detects ambient noise in real time, and if the noise level exceeds 60 dB, outputs a message saying "The noise is too loud. Please say it again" to request the user to re-enter it in a quiet environment. When the user pronounces a password, the system analyzes the pitch and speed of the voice to determine the user's emotional state. For example, if the user pronounces "1234" in a nervous voice, the system detects this as a signal indicating stress and requests an additional confirmation procedure saying "Please pronounce it again in a more comfortable environment."

[0104] The system (200) uses multiple microphones to receive the user's voice and analyzes the direction from which the voice is coming. For example, if the user speaks from a location close to the system, the system processes this voice first and verifies the user's identity based on the information that "the user is speaking from a close location." The system (200) analyzes the user's past payment history and, for example, if a user who typically makes payments of 50,000 won or less attempts to make a payment of 500,000 won, the system outputs a message saying "An abnormal transaction has been detected. Please proceed with additional authentication" to request additional authentication.

[0105]

[0106] According to one embodiment, the system (200) improves authentication accuracy by considering age-specific characteristics when analyzing the user's voice pattern, periodically updates the user's voice pattern data, and controls the system to maintain authentication accuracy by learning voice changes over time, prevents the use of pre-recorded voice by randomly requesting specific words or phrases when the user inputs voice and checking real-time responses, and provides a backup authentication method in the event of a user's voice authentication failure, while controlling the system so that the method is automatically selected according to the user's physical characteristics or situation.

[0107] Since the elderly often have inaccurate pronunciation or trembling voices, the system adjusts the voice pattern analysis algorithm by taking into account these age-specific characteristics to improve authentication accuracy. For example, when the system (200) analyzes the voice patterns of people in their 20s and 60s, it learns the voice characteristics of each age group and adjusts to recognize the lower voice of users in their 60s more accurately. The system (200) analyzes the user's voice annually to reflect changing voice characteristics. For example, if the tone of the user's voice lowers as they age, the system detects this and registers a new voice pattern to improve authentication accuracy. When the user says, "Please enter your password," the system (200) randomly requests, "Please say a specific phrase, 'Hello.'" It verifies whether the user can respond immediately to this phrase to check whether a pre-recorded voice is being used.

[0108] For example, if a user fails voice authentication, the system displays a message asking, "Would you like to authenticate via fingerprint recognition?" and suggests an additional authentication method based on the user's physical characteristics (such as whether fingerprint recognition is possible). If the user is unable to perform fingerprint recognition, the system automatically suggests an alternative method, such as "Please enter your password."

[0109] Since a user's voice can change over time, the system periodically requests voice data updates (e.g., every three months) and learns from the new data to maintain authentication accuracy. During user authentication, the system displays random words or phrases that cannot be pre-recorded, such as "Please read the numbers you see now," on the screen and verifies the user's real-time response to prevent recorded voice playback attacks. If a user loses their voice and fails voice authentication, the system automatically suggests fingerprint authentication if the user's smartphone has a fingerprint sensor, or facial authentication if it has a facial recognition camera, providing a backup authentication method appropriate to the situation.

[0110]

[0111] FIG. 4 is a flowchart illustrating a method for providing a high-dimensional payment solution utilizing generative artificial intelligence (AI)-based voice fingerprint authentication technology according to one embodiment.

[0112] Although process steps, method steps, algorithms, etc. are described in a sequential order in the flowchart of FIG. 4, such processes, methods, and algorithms may be configured to operate in any suitable order. In other words, the steps of the processes, methods, and algorithms described in various embodiments of the present invention do not need to be performed in the order described in the present invention. Furthermore, even if some steps are described as being performed asynchronously, in other embodiments, such steps may be performed simultaneously. Also, the example of a process by the depiction in the drawings does not mean that the exampled process excludes other variations and modifications therefrom, does not mean that any of the exampled process or its steps is essential to one or more of the various embodiments of the present invention, and does not mean that the exampled process is desirable.

[0113] In operation 410, the system (e.g., the system (200) of FIG. 2) can assign weights based on the number of recognized characters for each of the first utterance period and the second utterance period, under the control of the processor (e.g., the processor (220) of FIG. 2). The system (200) assigns weights based on the number of recognized characters for each of the first utterance period and the second utterance period of the voice signal, calculates the total number of characters for the first utterance period and the second utterance period, and can adjust the weight of the first utterance period according to the ratio of the first number of characters to the total number of characters.

[0114] For example, if the user pronounces "Hello" in the first utterance, the system counts the characters recognized in this utterance and records that 5 characters were recognized. If the user says "Nice to meet you" in the second utterance, the number of characters recognized in this utterance becomes 6. In this case, the system calculates the total number of characters as 11 and can adjust the weight of the first utterance to 5 / 11.

[0115] In operation 420, the system (200) can authenticate a user by comparing voice signals of the first speech period and the second speech period. The system (200) can adjust the weight of the second speech period according to the ratio of the number of characters in the second speech period to the total number of characters, and the authentication device authenticates a user by comparing the voice signals of the first speech period and the second speech period with the weight applied, detects noise cycles in the voice signals of the first speech period and the second speech period, performs authentication by comparing after removing the detected noise cycles, and performs a first authentication that identifies the first user by a voice authentication and an authentication method different from the authentication method.

[0116] For example, if the voice recognized in the first utterance is "Hello" and the voice in the second utterance is "Nice to meet you," the system evaluates similarity by comparing the two voices. At this time, the weight of the second utterance is adjusted to 6 / 11, and the authentication device performs authentication by comparing the two voices with the applied weights. In addition, the system detects noise cycles; for example, if a user speaks in an environment with ambient noise, the system analyzes and removes this noise and performs authentication by comparing only the remaining voice signal.

[0117] For example, if a user first pronounces "100,000 won from my account" (first utterance period) and then, shortly thereafter, pronounces "transfer it to Hong Gil-dong" (second utterance period), the system calculates the number of characters recognized in each utterance period. In the first utterance period, 10 characters ("100,000 won from my account") and in the second utterance period, 7 characters ("transfer it to Hong Gil-dong") were recognized. The total number of characters is 17, and the weight for the first utterance period is calculated as 10 / 17, while the weight for the second utterance period is calculated as 7 / 17. In other words, a higher weight is assigned to the first utterance period, which has more characters.

[0118] The weights calculated in the above example (first: 10 / 17, second: 7 / 17) are applied to the voice signals of each utterance period for comparison. In other words, the voice signal of the first utterance period is reflected in the comparison with greater weight. If a sudden loud noise occurs during the second utterance period, the system detects this as noise and compares it with the voice signal of the first utterance period using only the remaining voice signal, excluding that part. When the user selects "Log in with voice" through the smartphone app, the first authentication is performed by entering a password. After successful password authentication, the process proceeds to the voice authentication stage.

[0119] In operation 430, the system (200) may perform a second authentication to confirm that the second user is the first user identified in the first authentication. The system (200) obtains the first registered voice of the first user identified in the first authentication from a server separated from the system, but does not obtain the registered voices of other users, receives the second voice spoken by the second user during the operation and performs preprocessing, provides the preprocessed second voice data to a generative artificial intelligence to extract a voice pattern, and performs a second authentication to confirm that the second user is the first user identified in the first authentication by comparing the features of the extracted second voice with the features of the first voice, but does not compare with other registered voices, and when the second authentication is successfully completed, it may perform a process according to the operation associated with the second voice.

[0120] For example, the system obtains the first user's registered voice from a separate server and ignores the voices of other users. Subsequently, when the second user pronounces "Hello," the system preprocesses the voice to, for instance, remove noise and refine the voice data. The refined voice data is then provided to a generative AI to extract voice patterns, which are then compared with the first user's voice. If, during this process, the second user's voice is determined to be similar to the first user's voice, the second authentication is successfully completed, and a specific action associated with the second voice is subsequently performed.

[0121] The system (200) uses the ID information of a user who has passed the first authentication (password authentication) to retrieve only the voice data of that user stored on the server. It does not access the voice data of other users. The system requests voice input from the user with the guidance, "Please say 'Have a good day today' to verify your identity." It receives the user's voice input and checks for a match by comparing it with the voice data of that user retrieved from the server. At this time, it does not compare it with the voice data of other users, but only with the voice data of that user to enhance personal information protection. Once the second authentication (voice authentication) is successfully completed, it performs the "voice login" action requested by the user.

[0122]

[0123] In operation 440, the system (200) may display a notification requesting the pronunciation of a preset password. When the system (200) performs a process according to an operation associated with the second voice, if the second voice is associated with an operation that sets some of a plurality of setting values, it performs the process using the said some setting values ​​and other preset setting values, and transmits the second voice to a server that performs a voice recognition process only when the second authentication is successfully completed, and if the features of the second voice include the features of the first voice, it considers the second authentication to be successfully completed, and may display a notification requesting the pronunciation of a preset password based on the fact that the extracted second voice pattern matches the voice pattern of the registered first user.

[0124] When the user's voice command "Turn on the air conditioner" is executed, the system first checks the settings related to the action of turning on the air conditioner (e.g., temperature, airflow). If the user only says "Turn on the air conditioner" and does not specify the temperature or airflow, the system turns on the air conditioner using predetermined default settings (e.g., temperature 24 degrees, airflow medium).

[0125] If the user's voice command involves an action related to personal information (e.g., account transfer, personal information lookup) rather than simply turning the device on or off, the system transmits the user's voice data to the server to perform the voice recognition process only after a second authentication (voice authentication) has been successfully completed.

[0126] The system considers the second authentication to be successfully completed if the voice pattern extracted from the user's voice data is similar to the voice pattern of an existing registered user. For example, authentication is successful if the user's pronunciation, intonation, speaking speed, etc., match the existing data.

[0127] After voice authentication, the system may additionally request the pronunciation of the password to enhance security. For example, it displays a message saying "Please say your password for payment" and receives the user's pronunciation of the password.

[0128] The system displays a notification requesting the user to pronounce a pre-set password. For example, after the user is successfully authenticated, the system displays a message on the screen saying, "Please pronounce the password." In this case, if the second voice is related to an action of setting some of multiple configuration values, the system performs the process by combining those values ​​with other pre-defined configuration values. If the user pronounces "2023 password is secret," the system analyzes this voice and transmits it to a server performing a voice recognition process only if the second authentication is successfully completed. The system evaluates whether the features of the second voice include the features of the first voice, and if it determines that the two voices are similar, it displays a notification to pronounce the password.

[0129] In operation 450, the system (200) can process a payment request. The system (200) receives the user's pronunciation, analyzes the pitch, pronunciation speed, and timbre of the voice, and determines based on the analysis results that the voice belongs to a registered user. It processes the payment request based on the fact that the extracted voice pattern matches the voice pattern of the registered user and that the pronunciation for a pre-set password is determined to belong to the registered user, and can send a message indicating that the payment has been completed to the user terminal and the payment platform in real time.

[0130] When a user pronounces the password "7890", the system (200) analyzes the pitch, pronunciation speed, and tone of the voice and compares it with the password pronunciation data of a previously registered user. The system (200) combines the voice pattern comparison and password pronunciation analysis results, and if user authentication is successful, processes the payment request. For example, if the user gives a voice command saying "Pay 50,000 won via Samsung Pay" and authentication is successful, the system (200) proceeds with the payment of 50,000 won via Samsung Pay. When the payment is completed, a message saying "Payment completed" is displayed on the user's smartphone screen, and at the same time, payment completion information is transmitted in real time to a payment platform (e.g., an online shopping mall).

[0131] When the user pronounces "Please make a payment," the system receives the voice and analyzes its pitch, pronunciation speed, and timbre. For example, if the user pronounces it in a tone similar to their usual voice, the system verifies that the voice belongs to a registered user. If the extracted voice pattern matches that of a registered user and it is determined that the user has correctly pronounced a pre-set password, the system processes the payment request. Once the payment is complete, a message stating "Payment completed" is sent in real-time to the user's terminal and the payment platform to notify the user of the result.

[0132]

[0133] FIG. 5 is a flowchart illustrating a method for providing a high-dimensional payment solution utilizing generative artificial intelligence (AI)-based voice fingerprint authentication technology according to one embodiment.

[0134] Although process steps, method steps, algorithms, etc. are described in a sequential order in the flowchart of FIG. 5, such processes, methods, and algorithms may be configured to operate in any suitable order. In other words, the steps of the processes, methods, and algorithms described in various embodiments of the present invention do not need to be performed in the order described in the present invention. Furthermore, even if some steps are described as being performed asynchronously, in other embodiments, such steps may be performed simultaneously. Also, the example of a process by the depiction in the drawings does not mean that the exampled process excludes other variations and modifications therefrom, does not mean that any of the exampled process or any of its steps is essential to one or more of the various embodiments of the present invention, and does not mean that the exampled process is desirable.

[0135] In operation 510, the system (e.g., the system (200) of FIG. 2) can sample the user's voice under the control of a processor (e.g., the processor (220) of FIG. 2) and build a voice fingerprint characteristic model therefrom. The system (200) can identify the source of the voice based on voice data collected from the user.

[0136] For example, when a user pronounces "Hello, I am the user," the system samples the voice and analyzes its frequency, pitch, timbre, and other characteristics. Based on this collected voice data, the system models the user's voice characteristics, thereby identifying the specific features of that user's voice. This voice fingerprint characteristic model is utilized to recognize the user's unique voice.

[0137] When the user presses the "Voice Registration" button and says "Hello. I am Hong Gil-dong," the system samples the voice input through the microphone at regular intervals and converts it into digital data. The system extracts unique features from the sampled voice data, such as the user's timbre, intonation, and pronunciation habits, to create a voice fingerprint characteristic model. For example, it analyzes specific frequency patterns that appear when pronouncing "Ah" and changes in intonation that appear in the sentence "Hello" and stores them in the model.

[0138] The system analyzes the collected voice data to identify the source of the voice, such as whether it was input from the smartphone's built-in microphone, via an external Bluetooth device, or played back from a recorded file.

[0139] In operation 520, the system (200) can verify identity by comparing the user's voice with a registered voice. For example, if the user pronounces "I will make a payment," the system evaluates similarity by comparing this voice with a registered voice. If the two voices are similar, the system confirms that the user is a registered user and proceeds with authentication. In this process, the pitch, speed, and nuances of the pronunciation of the voice are analyzed to determine identity.

[0140] When a user attempts voice authentication, the system compares the currently input voice with a previously registered voice fingerprint characteristic model to verify if it is the same user. For example, when a user says "Please make the payment," the system analyzes the pronunciation, intonation, tone, etc., calculates the similarity with the registered model, and determines that it is the same user if it matches above a certain threshold.

[0141] In operation 530, the system (200) can improve the accuracy of speech recognition by applying an acoustic model. The system (200) can improve the accuracy of speech recognition by applying an acoustic model composed of a sampling convolutional layer, an intermediate layer, an attention pooling layer, a segment layer, and a softmax classification layer. The sampling convolutional layer includes a plurality of sampling bandpass filters, and each speech frame is converted into a bandpass signal of multiple channels, and the bandpass signal is output as a T-frame level vector through the intermediate layer, and the attention pooling layer can be configured to convert the T-frame level vector into a 1-frame vector through an attention introduction mechanism, and the 1-frame vector is transmitted to the softmax classification layer through the segment layer.

[0142] A sampling bandpass filter may include two learnable parameters having a high cutoff frequency and a low cutoff frequency. The form of the filter is characterized by being defined by the following: g[n, f1, f2]=2f2sinc(2πf2n)-2f1sinc(2πf1n)

[0143] Here, n represents the sampling point, f1 represents the low-frequency cutoff frequency of the bandpass filter, and f2 represents the high-frequency cutoff frequency of the bandpass filter.

[0144] For example, the sampling convolutional layer includes multiple sampling bandpass filters to convert the user's voice into a multi-channel bandpass signal. At this time, each voice frame is processed through the filter and output as a T-frame level vector at the intermediate layer. Subsequently, the attention pooling layer converts this vector into a single-frame vector through an attention mechanism, and the converted vector passes through the segment layer and is transmitted to the softmax classification layer.

[0145] A sampling bandpass filter includes two learnable parameters with a high cutoff frequency and a low cutoff frequency. For example, assuming f1 is 300 Hz and f2 is 3400 Hz, the filter improves the quality of the speech signal by passing only signals within a specific frequency range and blocking the rest. The form of the filter is defined as g[n, f1, f2] = 2f2sinc(2πf2n) - 2f1sinc(2πf1n), where n represents the sampling point. This formula allows for the calculation of the filter's characteristics to enable more sophisticated speech recognition.

[0146] Voice data is input, and multiple sampling bandpass filters are applied. Each filter is designed to allow only signals of a specific frequency band to pass through. For example, a filter that passes the 100Hz to 500Hz frequency band, a filter that passes the 500Hz to 1000Hz frequency band, etc., are used to separate the voice signal into multiple frequency bands.

[0147] The signals passing through each filter are analyzed and converted into a vector consisting of T frames. Each frame represents speech information at a specific time interval. Only important information is extracted from the T-frame vectors and compressed into a single frame vector. For example, if the "je" part of the word "payment" is more important for user identification than the "gyeol" part in pronunciation, a higher weight is assigned to the frame corresponding to the "je" part. The single frame vector is divided into phonemes. For example, the word "payment" is divided into phonemes such as "ㄱ," "ㅕ," "ㄹ," "ㅈ," and "ㅔ." The softmax classification layer classifies each phoneme and ultimately recognizes which word was pronounced.

[0148] In a sampling bandpass filter, each filter serves to pass a specific frequency band. For example, the g[n, 100, 500] filter passes only signals in the frequency band of 100Hz to 500Hz. The form of the filter is defined as g[n, f1, f2] = 2f2sinc(2πf2n) - 2f1sinc(2πf1n), and it uses the sinc function to produce the effect of smoothly cutting out a specific frequency band.

[0149]

[0150] According to one embodiment, the system (200) can remove silent segments and preserve valid voice segments from voice fingerprint signals through voice activity detection, divide the valid voice segments into voice segments of equal length, perform a short-time Fourier transform of each voice signal to generate a spectrogram matrix S_i, and perform a difference according to the time order of the spectrogram matrix to generate a difference matrix D_i.

[0151] For example, while a user pronounces "Hello," the system detects the start and end points of the speech and removes the silent "..." segment. Then, the valid speech segment is divided into speech segments of equal length. If the entire utterance is 3 seconds long, it is divided into 3 speech segments of 1 second each. A short-time Fourier transform is performed on each speech signal to generate a spectrogram matrix S_i. In this process, for example, the speech of the first 1-second segment is transformed to generate data visualized in the form of a spectrogram. Subsequently, a difference matrix D_i is generated by performing a difference operation on the generated spectrogram matrix according to its chronological order. This difference matrix visually represents changes in speech, contributing to the improvement of speech recognition accuracy.

[0152] For example, while the user initiates "voice authentication" and says "password 1234," the system removes the silent intervals before and after through voice activity detection and leaves only the "password 1234" pronunciation interval. The system (200) divides the "password 1234" pronunciation interval into voice intervals of 0.1 seconds in length.

[0153] The system (200) performs a Short Time Fourier Transform (STFT) on each 0.1-second speech segment to analyze frequency components that change over time and generates a spectrogram matrix S_i that visually represents this. Each row of the S_i matrix represents a time frame, and each column represents a frequency component. The values ​​of the matrix represent the intensity of the corresponding frequency component in the corresponding time frame. The system (200) generates a difference matrix D_i by calculating the difference between adjacent frames according to the time order of the spectrogram matrix S_i. This is to analyze the changes in the speech signal over time more clearly. For example, the spectrogram change between the "1" pronunciation and the "2" pronunciation is analyzed and stored in the difference matrix.

[0154] According to one embodiment, the system (200) is configured to set a threshold value to compare the values ​​of each coordinate and then convert them into a pulse matrix to be used as input to a pulse neural network, and when a user inputs a voice signal, it authenticates the identity using the voice fingerprint signal, and if authentication is passed, it grants the authority to execute the corresponding command, and if authentication fails, it provides an interface to select whether the user is a new user at the user terminal, and if no user inputs, it controls the system to deny access.

[0155] For example, when the system analyzes spectrogram data of a voice signal, it determines whether values ​​within a specific frequency range exceed a set threshold. When a user inputs a voice signal, the system authenticates their identity based on this voice fingerprint signal. If authentication is successful, the user is granted permission to execute specific commands. For instance, if the user pronounces "payment," the system recognizes this and allows the payment process to proceed. Conversely, if authentication fails, the system provides an interface on the user's terminal to select whether they are a new user. At this point, if the user wishes to register as a new user, they can click the "Register New User" button. If the user provides no input, the system controls access to be denied, thereby enhancing security.

[0156] The system (200) generates a pulse matrix by comparing each coordinate value of the difference matrix D_i with a preset threshold value, converting it to 1 if it is greater than the threshold value and 0 if it is smaller. This pulse matrix is ​​used as an input to a pulse neural network (SNN).

[0157] When a user says the voice command "Make a call," the system (200) inputs the pulse matrix generated through the above process into a pulse neural network to authenticate the user. A pulse neural network is a neural network that processes information in the form of pulses, is energy-efficient, and is suitable for real-time processing. If voice authentication is successful, the system grants the user permission to execute the "Make a call" command. If voice authentication fails, the system displays a message on the smartphone screen saying "No registered user found. Would you like to register as a new user?" to allow the user to choose whether to register as a new user. If the user does not choose to register as a new user, the system (200) displays a message saying "Voice authentication failed. Access denied." and blocks access to the corresponding function.

[0158]

[0159] According to one embodiment, the system (200) receives a dynamic authentication character generated in response to a user's voice payment request, the dynamic authentication character is randomly selected from a fixed-length character data template and transmitted to the user, provides guidance information to the user, collects voice input including voice information corresponding to user information and the dynamic authentication character, displays a message to the user to perform voice input before collecting voice input, and displays a message to the user to select and confirm multiple pronunciations when collecting voice input.

[0160] For example, when a user says "payment request," the system detects this request and generates a dynamic authentication code. This dynamic authentication code is randomly selected from a fixed-length character data template and sent to the user. For instance, the system selects the authentication code "ABC123" and sends it to the user. Subsequently, the system provides the user with instructions such as, "Please perform voice input now." The user prepares to collect voice input containing their voice information. Before collecting voice input, the system displays a message stating, "Please speak when you are ready to input voice." When collecting voice input, the system displays an additional message to the user stating, "Please select multiple pronunciations and confirm," prompting the user to choose from various pronunciations, such as "payment" or "cancel."

[0161] For example, when a user requests "Pay 50,000 won via OO Pay" by voice, the system receives the payment request and generates a dynamic authentication character such as "7835". At this time, the dynamic authentication character is generated by randomly selecting 4 characters from a pre-stored fixed-length character data template "0123456789".

[0162] The system displays the generated dynamic authentication characters "Chilpalsam-o" on the user's smartphone screen and provides a guidance message such as "Please read the numbers displayed on the screen." When the user pronounces "Chilpalsam-o" according to the guidance message, the system collects voice input through the microphone. Before collecting voice input, the system guides the user on how to input voice by displaying a message such as "Touch the screen to start voice input." If there is a character that can be pronounced in multiple ways, such as "Chil," the system displays a message such as "Which pronunciation would you like to select? 1) Chil 2) Chil" to allow the user to select the desired pronunciation.

[0163] According to one embodiment, the system (200) displays a message to the user to check the input result after collecting voice input, performs identity authentication and voice fingerprint authentication based on voice input including dynamic authentication characters and voice information corresponding to user information, verifies the user's bank card information, obtains seller information and transaction amount if identity authentication and voice fingerprint authentication are passed, performs payment verification by entering a password, and controls the payment of the transaction amount to the seller corresponding to the seller information if payment verification is passed.

[0164] For example, when the user pronounces "payment request," the system displays a message asking, "Is the entered content correct? Would you like to confirm as 'payment request'?" Subsequently, the system performs identity authentication and voice fingerprint authentication based on voice input containing dynamic authentication characters and voice information corresponding to user information. During this process, the system verifies the user's identity by comparing the user's voice data with registered voice data and checks the user's bank card information. If identity authentication and voice fingerprint authentication are passed, the system obtains seller information and the transaction amount. For example, if the transaction amount is 50,000 won, the system provides an announcement stating, "The transaction amount is 50,000 won," based on this information. The user then enters a password to verify the payment, and the system performs payment verification after confirming that the password is correct. If payment verification is passed, the system controls the payment of the transaction amount to the seller corresponding to the aforementioned seller information. For example, it outputs a message stating, "Paying 50,000 won to Seller A," and completes the transaction.

[0165] When the user finishes voice input, the system displays a message such as, "The entered number is seven-eight-three-five. Is that correct?" to allow the user to confirm the input result. The system verifies the user's identity using voice fingerprint information extracted from the user's voice. For example, it analyzes the tone, pronunciation habits, and frequency characteristics of the voice and compares them with previously registered user voice data. The system checks the user's bank card information to verify whether the card is valid and whether the balance is sufficient. If identity authentication and voice fingerprint authentication are passed, the system obtains seller information and the transaction amount from the user's payment request information. For example, from a request such as "Pay 50,000 won via OO Pay," it extracts the seller information "OO Pay" and the transaction amount "50,000 won."

[0166] The system requests the user to enter a payment password and performs a payment verification process to confirm whether the entered password matches. Once payment verification is complete, the system pays the transaction amount to the relevant merchant. For example, payment information for 50,000 won is transmitted to the OO Pay system.

Claims

1. In a system providing a high-level payment solution utilizing generative artificial intelligence (AI)-based voice fingerprint authentication technology, Memory for storing instructions; and Includes a processor, When the above instructions are executed by the processor, the system Receive user voice input and perform preprocessing, Preprocessed voice data is provided to generative artificial intelligence to extract voice patterns, The extracted voice pattern is compared with the voice pattern of a user registered in advance to check for a match, Displays a notification requesting the pronunciation of a preset password based on the match between the extracted voice pattern and the registered user's voice pattern, and It receives the user's pronunciation, analyzes the pitch, speed, and timbre of the voice, and determines based on the analysis results whether the voice belongs to a registered user, The extracted voice pattern matches the voice pattern of the registered user, and Processing a payment request based on the fact that the sequence of numbers of a pre-set password matches and the voice pattern of the password pronunciation is confirmed to belong to a registered user, and A system that controls the real-time transmission of a message indicating that payment has been completed to a user terminal and a payment platform.

2. In Paragraph 1, When the above instructions are executed by the processor, the system Voice pattern comparison is performed using the Dynamic Time Warping (DTW) algorithm, and Using the above generative artificial intelligence, voice patterns are extracted based on time-frequency analysis of voice signals, and To prevent fake attacks, control liveness verification to be performed, and A system characterized by a preprocessing process that includes steps for noise removal, normalization, and feature extraction of a speech signal, and a step for extracting Mel-Frequency Cepstral Coefficients (MFCC).

3. In Paragraph 1, When the above instructions are executed by the processor, the system Control to output the message "Please authenticate by speaking for payment authentication" before receiving the user's voice input, and After verifying whether the extracted voice pattern matches the voice pattern of a registered user, the result is fed back to the user terminal in real time, and When receiving a user's pronunciation and analyzing the pitch, speed, and timbre of the voice, the analysis is performed using generative artificial intelligence, and A system that controls support for one or more payment methods among credit cards, electronic money, and mobile payments when processing a payment request.

4. In Paragraph 1, When the above instructions are executed by the processor, the system Controls all voice fingerprint data and payment information to be encrypted using the AES-256 encryption method, and It monitors all authentication and payment activities in real time and stores logs to detect abnormal activity, and A system that uses generative artificial intelligence to analyze the minute features of a voice pronounced by a user and controls it to block attacks using recorded or synthesized voice.

5. In Paragraph 1, When the above instructions are executed by the processor, the system When receiving user voice input, it detects ambient noise, and if the noise level exceeds a preset threshold, it requests re-input from the user, and Analyzes the emotional state when extracting the user's voice patterns, and performs additional verification procedures if stress or coercive situations are detected, When receiving user voice input, it receives input through multiple microphones, analyzes the direction and distance of the voice to determine the actual user's location, and A system that detects abnormal transactions by comparing and analyzing the user's usual payment patterns with current payment details when processing payment requests, and controls the system to require additional authentication if necessary.

6. In Paragraph 1, When the above instructions are executed by the processor, the system When analyzing users' voice patterns, consider age-specific characteristics to improve authentication accuracy, and It periodically updates the user's voice pattern data and learns voice changes over time to control and maintain authentication accuracy, By randomly requesting specific words or phrases during user voice input and verifying real-time responses, the use of pre-recorded voice is prevented, and A system that provides a backup authentication method in the event of a user's voice authentication failure, and controls the method to be automatically selected according to the user's physical characteristics or situation.

7. In Paragraph 1, When the above instructions are executed by the processor, the system Weights are assigned to the first and second utterance periods of the speech signal, respectively, based on the number of recognized characters, and Calculate the total number of characters in the first utterance period and the second utterance period, adjust the weight of the first utterance period according to the ratio of the first number of characters to the total number of characters, and By adjusting the weight of the second utterance period according to the ratio of the second character count to the total character count, the authentication device authenticates the user by comparing the voice signals of the first utterance period and the second utterance period with the weights applied. Detect noise cycles in the voice signals of the first and second utterance periods, perform authentication by removing the detected noise cycles and comparing, and Performing a first authentication that identifies a first user using voice authentication and other authentication methods, and Obtain the first registered voice of the first user identified in the first authentication from a server separated from the above system, but not the registered voices of other users, and receive the second voice spoken by the second user during operation and perform preprocessing, Preprocessed second voice data is provided to a generative artificial intelligence to extract voice patterns, and a second authentication is performed by comparing the features of the extracted second voice with the features of the first voice to confirm that the second user is the first user identified in the first authentication, without comparing with other registered voices, and if the second authentication is successfully completed, a process according to an action associated with the second voice is performed. When performing a process according to an action associated with the second voice, if the second voice is associated with an action of setting some of a plurality of setting values, the process is performed using the said some setting values ​​and other predetermined setting values. The second voice is transmitted to a server performing a voice recognition process only when the second authentication is successfully completed, and If the features of the second voice include the features of the first voice, the second authentication is considered to have been successfully completed, and a notification is displayed requesting the pronunciation of a preset password based on the fact that the extracted second voice pattern matches the voice pattern of the registered first user, and It receives the user's pronunciation, analyzes the pitch, speed, and timbre of the voice, and determines based on the analysis results whether the voice belongs to a registered user, Processing a payment request based on the fact that the extracted voice pattern matches the voice pattern of the registered user and the pronunciation of the preset password is determined to be that of the registered user, and A system that controls the real-time transmission of a message indicating that payment has been completed to the user terminal and the payment platform.

8. In Paragraph 1, When the above instructions are executed by the processor, the system It samples the user's voice and builds a voice fingerprint feature model based on it, identifies the source of the voice based on voice data collected from the current user, and It verifies identity by comparing the user's voice with the registered voice through a voice recognition function, and Improve the accuracy of speech recognition by applying an acoustic model composed of a sampling convolutional layer, an intermediate layer, an attention pooling layer, a segmentation layer, and a softmax classification layer, and The sampling convolutional layer includes a plurality of sampling bandpass filters, and Each voice frame is converted into a multi-channel bandpass signal, and The bandpass signal is output as a T-frame level vector through the intermediate layer, and The above attention pooling layer is configured to convert a T-frame level vector into a 1-frame vector through an attention introduction mechanism, and the 1-frame vector is transmitted to a softmax classification layer via a segment layer, and The above sampling bandpass filter includes two learnable parameters having a high cutoff frequency and a low cutoff frequency, and The shape of the filter is characterized by being defined by the following, and g[n, f1, f2]=2f2sinc(2πf2n)-2f1sinc(2πf1n) A system characterized in that n is a sampling point, f1 is a low-frequency cutoff frequency of a bandpass filter, and f2 is a high-frequency cutoff frequency of a bandpass filter.

9. In Paragraph 1, When the above instructions are executed by the processor, the system The voice fingerprint signal removes silent intervals and preserves valid voice intervals through voice activity detection, Divide the above valid voice segments into voice segments of equal length, and A spectrogram matrix S_i is generated by performing a short-time Fourier transform on each speech signal, and Difference matrix D_i is generated by performing difference according to the time order of the above spectrogram matrix, and It is configured to compare the values ​​of each coordinate by setting a threshold, and then convert them into a pulse matrix to be used as input to a pulse neural network, When a user inputs a voice signal, the identity is authenticated using the voice fingerprint signal, and if authentication is successful, permission to execute the corresponding command is granted. If authentication fails, an interface is provided on the user terminal to select whether it is a new user, and A system that controls access so that it is denied if no user inputs.

10. In Paragraph 1, When the above instructions are executed by the processor, the system A dynamic authentication character generated in response to a user's voice payment request is received, said dynamic authentication character is randomly selected from a fixed-length character data template and sent to the user, and provides guidance information to the user. Collect voice input including voice information corresponding to user information and the dynamic authentication character, and Before collecting voice input, display a message to the user to perform voice input, and When collecting voice input, a message is displayed to the user prompting them to select multiple pronunciations and confirm, After collecting voice input, it displays a message to the user asking them to check the input results, Based on voice input including voice information corresponding to the dynamic authentication characters and user information, identity authentication and voice fingerprint authentication are performed, and the user's bank card information is verified. Upon passing the above identity verification and voice fingerprint verification, seller information and transaction amount are obtained, and A system that performs payment verification by entering a password, and controls the payment of a transaction amount to the seller corresponding to the above seller information if the payment verification is passed.

Citation Information

Patent Citations

  • Method for processing voice payment

    KR1020090039693A

  • Payment method using voice information and payment relay server thereof

    KR1020140143047A

  • Voice elctronic settlement service

    KR1020180058150A

  • Plate heat exchanger

    KR1020240175690A

  • Aluminum-based amorphous alloy and electronic devices using the same

    KR102614141B1