Internet-based AI voice payment system and use method thereof

Through an Internet-based AI voice payment system, deep learning models are used for speech recognition and natural language understanding, combined with security verification and abnormal detection mechanisms, the recognition problem of existing voice payment systems in complex sentences and long voice input is solved, and an efficient and secure payment experience is achieved.

CN120013542AInactive Publication Date: 2025-05-16深圳毕加索电子有限公司

Patent Information

Application Number
CN202510347413.5
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-03-24
Publication Date
2025-05-16
Estimated Expiration
Not applicable · inactive patent

AI Technical Summary

Technical Problem

The existing voice payment system has problems such as unclear understanding of intentions and inaccurate slot extraction when recognizing complex sentences and long voice input, resulting in the inability to proceed smoothly.

Method used

An Internet-based AI voice payment system is adopted, including a voice input module, a voice recognition module, a natural language understanding module, a security verification module, a payment execution module and a feedback module. The system uses deep learning models to perform speech recognition, combines natural language understanding technology to analyze the intentions in user speech, uses voiceprint recognition and face recognition for security verification, and detects abnormal speech through machine learning algorithms to block high-risk transactions in real time.

Benefits of technology

It improves the accuracy and robustness of voice recognition, ensures the legality of payment requests, prevents malicious users from impersonating voice for payment operations, realizes efficient payment execution and real-time feedback, and improves the security and convenience of the payment experience.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120013542A_ABST
    Figure CN120013542A_ABST
Patent Text Reader

Abstract

The invention relates to the technical field, in particular to an Internet-based AI voice payment system and a use method thereof. According to the technical scheme, a voice input module is responsible for collecting voice input signals of a user, a voice recognition module converts the voice signals into text information, a natural language understanding module conducts semantic analysis on the text information and understands the payment intention of the user, and a safety verification module ensures the safety of the identity of the user. Comprising verification modes of face recognition, voiceprint recognition and the like, the payment execution module is in butt joint with a bank and a payment platform through a payment interface to perform actual payment operation, and the feedback module returns a payment result to inform a user of whether payment succeeds or not and information such as account balance. According to the invention, a multi-modal speech recognition and security verification mechanism is integrated, the accuracy and robustness of speech recognition are improved, and payment instructions and key information related to payment are accurately recognized in the aspects of natural language understanding and intention recognition.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the technical field, and in particular to an Internet-based AI voice payment system and a method for using the same. Background Art

[0002] With the rapid development of Internet technology and artificial intelligence (AI) technology, smart payment has gradually become an indispensable part of daily life. Although traditional payment methods, such as cash, bank card payment, QR code scanning payment, etc., are convenient and fast, they still have some shortcomings. For example, cash payment is easily restricted by region and convenience; bank card payment is easy to be stolen or lost; QR code payment requires the use of mobile terminals, and its payment process may have certain security risks and inconveniences.

[0003] To solve these problems, voice payment has gradually emerged as a new payment method. Voice payment uses AI voice recognition technology to enable users to complete payments with voice commands without entering passwords or using hardware devices such as mobile phones, which is extremely convenient and secure.

[0004] In this context, an Internet-based AI voice payment system is proposed. Through the collaborative work of AI voice recognition, natural language understanding, payment execution and other technical modules, users can realize remote payment through simple voice commands, avoiding the complexity of manual input and device operation in traditional payment methods, thereby improving the efficiency and convenience of payment.

[0005] Although some voice payment systems have attempted to enter the market, current technology still faces the following problems:

[0006] Existing speech recognition technology may be affected by factors such as noise, dialects, and speaking speed, resulting in the inability to guarantee the accuracy of speech recognition results, which in turn affects the payment experience. Existing voice payment systems have problems with unclear intention understanding and inaccurate slot extraction when recognizing complex sentences and long voice inputs, resulting in payment operations being unable to proceed smoothly.

[0007] To this end, we propose an Internet-based AI voice payment system and its usage method to solve existing problems. Summary of the invention

[0008] The purpose of the present invention is to address the problems existing in the background technology and to propose an Internet-based AI voice payment system and a method for using the same.

[0009] To achieve the above-mentioned purpose, the present invention provides the following technical solutions: an AI voice payment system based on the Internet, comprising a voice input module, a voice recognition module, a natural language understanding module, a security verification module, a payment execution module and a feedback module, wherein the voice input module is responsible for collecting the user's voice input signal, the voice recognition module converts the voice signal into text information, the natural language understanding module performs semantic analysis on the text information and understands the user's payment intention, the security verification module ensures the security of the user's identity, including verification methods such as face recognition and voiceprint recognition, the payment execution module connects with the bank and the payment platform through the payment interface to perform actual payment operations, the feedback module returns the payment result to inform the user whether the payment is successful and the account balance and other information, the security verification module also includes an abnormal voice detection module, a risk scoring module and a real-time monitoring module, the abnormal voice detection module detects abnormal voice through a machine learning algorithm, the risk scoring module performs risk scoring according to user behavior and voice characteristics, and blocks high-risk transactions in real time, and the real-time monitoring module is used to monitor abnormal behavior in the payment process in real time, and to promptly warn and handle it.

[0010] Preferably, the voice input module uses a microphone with high sensitivity and dynamic range to capture the user's voice, converts the analog signal into a digital signal, ensures sufficient sampling rate and quantization accuracy, identifies and removes the silent part in the voice signal, reduces the burden of subsequent processing, enhances the high-frequency part of the voice to compensate for the attenuation of the high-frequency components during the propagation process, analyzes the characteristics of the ambient noise, estimates the power spectrum of the noise, subtracts the noise spectrum from the spectrum of the voice signal to reduce the impact of the noise, and uses a nonlinear processor filter to predict and eliminate echoes.

[0011] Preferably, the speech recognition module includes converting the speech signal processed by the speech input module into a Mel-scale spectrum through Mel-frequency cepstral coefficients and extracting the cepstral coefficients, thereby retaining the main information of the speech, and then using a long short-term memory network acoustic model to model the relationship between the input audio signal and words or phonemes, outputting the score or state sequence of the acoustic model, and then using an N-gram language model to combine the output of the acoustic model and the language model to find the most likely word sequence through a decoding algorithm.

[0012] Preferably, the natural language understanding module, i.e., the NLU module, after receiving the user's voice signal and converting it into text through the speech recognition module, first needs to perform a series of preprocessing operations on the original text to ensure text quality, remove noise, and normalize it into a standard format. This includes removing irrelevant characters: removing possible noise characters or meaningless symbols, such as filler words like "um", "ah", etc.; case conversion: unifying the case in the text to ensure consistency; word segmentation and part-of-speech tagging: splitting the text into words and tagging the part of speech of each word to support subsequent semantic analysis and slot extraction; spelling correction: automatically correcting possible spelling mistakes according to the context, especially in cases of accents or rapid voice input. Then, intent recognition is performed to identify the specific operations the user wants to execute from the user's text, such as payment, balance query, etc. Intent recognition uses a neural network model learning algorithm. Subsequently, a conditional random field (CRF) pre-trained language model is used for slot filling. The task of slot filling is to extract specific entity information from the user's text, such as amount, payee, payment method, etc. Finally, context understanding is carried out to track the user's historical intentions and adjust the results of intent recognition according to changes in the context.

[0013] Preferably, the security verification module includes voiceprint recognition: verifying identity through the user's voice characteristics (such as pitch, speech rate, pronunciation method, etc.) and face recognition: capturing the user's facial image through a camera and comparing it with the face features in the database.

[0014] Preferably, the payment execution module calls the payment interface of the payment platform through an API. The system transmits the payment request and related information (such as amount, payee) to the payment platform. After the payment platform verifies the payment request, it returns the payment result. First, the system checks the user's balance to ensure that the account balance is sufficient for payment. If the balance is sufficient, the payment platform performs a deduction operation to complete the payment. If the balance is insufficient, the system returns a prompt message indicating insufficient balance.

[0015] Preferably, the feedback module generates feedback information based on the payment execution result and presents it to the user in the form of voice or text.

[0016] A method for using an Internet-based AI voice payment system includes the following steps:

[0017] Step 1: User voice input

[0018] The user inputs a payment instruction through a voice assistant such as a smartphone, smart speaker, etc. The user says, "Pay 100 yuan to Zhang San."

[0019] Step 2: Speech recognition

[0020] The system converts the user's voice signal into text through the voice recognition module. The voice command "Pay 100 yuan to Zhang San" will be converted into text: "Pay 100 yuan to Zhang San".

[0021] Step 3: Natural Language Understanding

[0022] The system inputs the text into the NLU module for analysis, extracting the user's payment intention, payment amount, and payee information:

[0023] Payment Intent: Pay

[0024] Payment amount: 100 yuan

[0025] Payee: Zhang San.

[0026] Step 4: Security Verification

[0027] According to the security policy, the system requires users to authenticate their identities, which may be voiceprint recognition or face recognition. If the verification is successful, the system continues with the payment operation; if the verification fails, the payment is rejected.

[0028] Step 5: Payment Execution

[0029] The system checks the user's account balance to ensure that the balance is sufficient for payment. If the balance is sufficient, the system completes the payment operation through the payment platform API. After the payment is completed, the system updates the user's account balance and returns the payment status.

[0030] Step 6: Feedback

[0031] The system will provide feedback to the user on the payment result, including whether the payment is successful, the payment amount, the account balance, etc. If the payment is successful, the user will receive feedback that "payment is successful, the account balance is 400 yuan"; if the payment fails, the system will provide the reason for the failure, such as insufficient balance, etc.

[0032] Compared with the prior art, the present invention has the following beneficial effects:

[0033] Multimodal speech recognition and security verification mechanism: The system uses the advanced deep learning model LSTM for speech recognition, which improves recognition accuracy and robustness. Especially in complex noise environments, the system can effectively recognize the user's voice commands, introduce voiceprint recognition technology to verify the user's identity, ensure the legitimacy of payment requests, and prevent malicious users from using voice to make payment operations. It uses machine learning algorithms to detect abnormal voices (such as recordings and synthesized voices), and can perform risk scoring based on user behavior and voice characteristics, blocking high-risk transactions in real time, and can monitor abnormal behaviors in the payment process in real time, and promptly warn and handle them;

[0034] Natural language understanding and intent recognition: The system uses natural language understanding technology to analyze the intent in the user's voice, accurately identify payment instructions and key payment-related information (such as amount, payee, etc.), support multiple payment scenarios, and can handle cross-platform payment requests and multiple payment methods (such as bank card payment, third-party payment platform, etc.);

[0035] Efficient payment execution and real-time feedback mechanism: The system can complete the processing of payment requests in a very short time and provide real-time feedback of payment results to users after the payment is completed, ensuring the efficiency and transparency of the payment process. BRIEF DESCRIPTION OF THE DRAWINGS

[0036] Figure 1 It is a schematic diagram of the system flow of the present invention; DETAILED DESCRIPTION

[0037] The following will be combined with the drawings in the embodiments of the present invention to clearly and completely describe the technical solutions in the embodiments of the present invention. Obviously, the described embodiments are only part of the embodiments of the present invention, not all of the embodiments. Based on the embodiments of the present invention, all other embodiments obtained by ordinary technicians in this field without creative work are within the scope of protection of the present invention.

[0038] Embodiment 1

[0039] like Figure 1 As shown, the present invention proposes an Internet-based AI voice payment system and a method of using the same, including a voice input module, a voice recognition module, a natural language understanding module, a security verification module, a payment execution module and a feedback module.

[0040] The voice input module is responsible for collecting the user's voice input signal, using a microphone with high sensitivity and dynamic range to capture the user's voice, converting the analog signal into a digital signal, ensuring sufficient sampling rate and quantization accuracy, identifying and removing the silent part in the voice signal, reducing the burden of subsequent processing, enhancing the high-frequency part of the voice to compensate for the attenuation of the high-frequency components during the propagation process, analyzing the characteristics of the ambient noise, estimating the power spectrum of the noise, subtracting the noise spectrum from the spectrum of the voice signal to reduce the impact of noise, and using a nonlinear processor filter to predict and eliminate echoes. Through the above steps, the user's voice commands can be effectively extracted and analyzed even in a noisy environment.

[0041] The voice recognition module is a crucial part of the entire Internet-based AI voice payment system. It is responsible for converting the user's voice input signal into text information that the machine can understand. This is the first step in the entire payment process.

[0042] The speech recognition module includes converting the speech signal processed by the speech input module into a Mel-scale spectrum through the Mel-frequency cepstral coefficient and extracting the cepstral coefficient, which retains the main information of the speech. Its mathematical formula is:

[0043] MFCC=DCT(log(|S(f)| 2 ))

[0044] Among them, S(f) is the short-time Fourier transform of the signal, and DCT is the discrete cosine transform. Through MFCC, the spectral features of the speech signal can be effectively extracted, making the subsequent recognition process more efficient.

[0045] Then, the long short-term memory network acoustic model LSTM is used to model the relationship between the input audio signal and words or phonemes, and the score or state sequence of the acoustic model is output. Then, the N-gram language model is used to combine the output of the acoustic model and the language model, and the most likely word sequence is found through a decoding algorithm.

[0046] The speech recognition process can be expressed as:

[0047] T=f sr (V,θ)

[0048] in:

[0049] T is the transformed text,

[0050] V is the input speech signal,

[0051] f sr is the speech recognition function,

[0052] θ is the parameter of the AI ​​model.

[0053] The acoustic model output can be expressed as:

[0054] The goal of the acoustic model is to learn the mapping relationship between audio signals and text. Assume that the input is a speech signal V = {v1, v2, ..., v T}, and its corresponding text sequence is T = {t1, t2, …, t N}, where t i is the recognition result at each moment. The output of the acoustic model is i The probability distribution of .

[0055] P(t i |v1,v2,…,v T )=P(t i |V);

[0056] Among them, P(t i|V) is a given input audio signal V, generating a specific character t i probability.

[0057] The audio signal V is converted into text T through an end-to-end neural network architecture, where the end-to-end neural network architecture is the CTC loss function. In speech recognition, the CTC (connected temporal classification) loss function is used to train the model so that it can handle problems of different lengths between audio signals and text. The core idea of ​​CTC loss is to calculate the probability of the entire sequence, and the formula is as follows:

[0058]

[0059] in:

[0060] Z is the set of all possible character label sequences (including blank labels).

[0061] P(z|V) represents the probability of generating the sequence z (including blank characters) given the audio signal V.

[0062] Minimize the CTC loss L by maximizing P(z|V) CTC , thereby training the model for accurate speech recognition.

[0063] Mathematical model T = f sr (V,θ) describes the entire speech recognition process, where V is the input speech signal, T is the output text sequence, and f sr is the function that realizes this transformation, which depends on the model parameters θ. The acoustic model is f sr is responsible for generating the probability distribution of T.

[0064] During training, we need to optimize the model parameters θ to minimize the CTC loss function L CTC This process is achieved through the gradient descent optimization algorithm, and the output probability distribution of the acoustic model directly affects the value of the CTC loss function.

[0065] The CTC loss function allows alignment between input and output sequences of different lengths by introducing blank labels. The output of the acoustic model is the character probability for each time step, and the CTC loss function accumulates these probabilities, considering all possible alignments, to obtain the probability of the entire output sequence.

[0066] The natural language understanding module, or NLU module, receives the user's voice signal and converts it into text through the speech recognition module. It first needs to perform a series of preprocessing operations on the original text to ensure the text quality, remove noise and normalize it to a standard format, including:

[0067] Remove irrelevant characters: Remove possible noise characters or meaningless symbols, such as filler words like "um", "ah", etc.;

[0068] Case conversion: Unify the case in the text to ensure consistency;

[0069] Word segmentation and part-of-speech tagging: Segment the text into words and tag the part of speech of each word to support subsequent semantic analysis and slot extraction;

[0070] Spelling correction: Automatically correct possible spelling mistakes according to the context, especially in the case of accents or rapid speech input.

[0071] Then perform intent recognition to identify the specific operations the user wants to perform from the text, such as payment, balance inquiry, etc. Intent recognition uses a neural network model for learning.

[0072] In deep learning, intent recognition can be completed through the neural network model f intent Given the cleaned text T clean , the model outputs the corresponding intent label I intent :

[0073] I intent = f intent (T clean )

[0074] Where:

[0075] I intent represents the intent label recognized from the text, such as "payment", "balance inquiry", etc.

[0076] f intent is a classification model, usually a deep neural network, which maps the text to predefined intent categories.

[0077] For example, the user's speech "Pay 200 yuan to Zhang San" may be recognized as the "payment" intent, and information such as the amount and payee is extracted.

[0078] Subsequently, a conditional random field (CRF) pre-trained language model is used for slot filling. The task of slot filling is to extract specific entity information from the user's text, such as amount, payee, payment method, etc. Common slots include:

[0079] Amount: Identify the payment amount in the text, such as "100 yuan".

[0080] Payee: Identify the name of the payment recipient, such as "Zhang San".

[0081] Payment method: Identify the payment method, such as "bank card", "Alipay", etc.

[0082] Mathematically, slot filling can be expressed as a function f slot , which is based on the cleaned text T clean To extract information about each slot:

[0083] Slots = f slot (T clean )

[0084] in:

[0085] Slots is a collection containing all extraction slot information, for example:

[0086] {Amount:100,Recipient:Zhang San,PaymentMethod:Bank Card}.

[0087] Finally, context understanding is performed to help the system make more accurate judgments in subsequent conversations by recording the user's historical conversation status (such as payment amount, payee, etc.). In multiple rounds of conversations, the system can understand changes in user needs and dynamically adjust its intent recognition and slot filling strategies.

[0088] Mathematically, this can be achieved by introducing the dialogue state S context To describe the system's understanding of the current conversation:

[0089] S context =f context (T clean ,S previous )

[0090] in:

[0091] S context Indicates the current dialog state of the system, including recognized intents and slots.

[0092] S previous Represents the previous state of the conversation, helping the system understand changes in user intent.

[0093] The security verification module includes voiceprint recognition: verifying identity through the user's voice characteristics (such as tone, speaking speed, pronunciation, etc.) and face recognition: capturing the user's facial image through the camera and comparing it with the facial features in the database.

[0094] Voiceprint recognition function

[0095] Assume X u represents the user's identity information (such as voiceprint template), and Y v represents the verification signal collected in real time (such as the user's current voiceprint). We use the voiceprint recognition function f id (X u ,Yv ) to determine the user's identity.

[0096]

[0097] in:

[0098] S=1 means verification passed, S=0 means verification failed;

[0099] f id (X u ,Y v ) is the similarity score of voiceprint recognition;

[0100] λ is the threshold for passing the verification.

[0101] If the similarity score of voiceprint recognition exceeds the threshold λ, the verification is considered to be passed and the system continues to perform the payment operation; if it fails, the payment is rejected.

[0102] Face recognition function

[0103] Face recognition is similar to voiceprint recognition and also involves identity verification. u Facial features stored for the user, Y v For the currently captured facial features, use the face recognition function for identity authentication:

[0104]

[0105] in:

[0106] f face (X u ,Y v ) represents the similarity measure of facial features (such as cosine similarity, Euclidean distance, etc.).

[0107] The payment execution module is responsible for connecting with the payment platform (such as Alipay, WeChat Pay, banks, etc.) to complete the actual payment operation. It calls the payment interface of the payment platform through the API. The system passes the payment request and related information (such as amount, payee) to the payment platform. After the payment platform verifies the payment request, it returns the payment result.

[0108] Payment process:

[0109] First, the system checks the balance of user B to ensure that the account balance is sufficient for payment:

[0110] B′=BP (if B≥P)

[0111] If the balance is sufficient, the payment platform will deduct the money and complete the payment.

[0112] If the balance is insufficient, the system returns a prompt message indicating insufficient balance.

[0113] After the payment is completed, the system returns the payment result to the user through the feedback module. The feedback content includes whether the payment is successful, the payment amount, the current account balance, etc.

[0114] The feedback module generates feedback information based on the payment execution results and displays it to the user in the form of voice or text.

[0115] Feedback function:

[0116] R = f feedback (B′, state

[0117] Among them, R is the feedback content, B′ is the account balance after payment, and the status is whether the payment is successful.

[0118] The method of using this system includes the following steps:

[0119] Step 1: User voice input

[0120] The user enters payment instructions through voice assistants such as smartphones, smart speakers and other devices, and the user says: "Pay 100 yuan to Zhang San."

[0121] Step 2: Voice Recognition

[0122] The system converts the user's voice signal into text through the voice recognition module. The voice command "Pay 100 yuan to Zhang San" will be converted into text: "Pay 100 yuan to Zhang San".

[0123] Step 3: Natural Language Understanding

[0124] The system inputs the text into the NLU module for analysis, extracting the user's payment intention, payment amount, and payee information:

[0125] Payment Intent: Pay

[0126] Payment amount: 100 yuan

[0127] Payee: Zhang San.

[0128] Step 4: Security Verification

[0129] According to the security policy, the system requires users to authenticate their identities, which may be voiceprint recognition or face recognition. If the verification is successful, the system continues with the payment operation; if the verification fails, the payment is rejected.

[0130] Step 5: Payment Execution

[0131] The system checks the user's account balance to ensure that the balance is sufficient for payment. If the balance is sufficient, the system completes the payment operation through the payment platform API. After the payment is completed, the system updates the user's account balance and returns the payment status.

[0132] Step 6: Feedback

[0133] The system will provide feedback to the user on the payment result, including whether the payment is successful, the payment amount, the account balance, etc. If the payment is successful, the user will receive feedback that "payment is successful, the account balance is 400 yuan"; if the payment fails, the system will provide the reason for the failure, such as insufficient balance, etc.

[0134] The above-mentioned specific embodiments are only several preferred embodiments of the present invention. Based on the technical solutions of the present invention and the relevant inspirations of the above-mentioned embodiments, those skilled in the art can make various alternative improvements and combinations to the above-mentioned specific embodiments.

[0135] It will be apparent to those skilled in the art that the present invention is not limited to the details of the exemplary embodiments described above and that the present invention can be implemented in other specific forms without departing from the spirit or essential features of the present invention. Therefore, the embodiments should be considered exemplary and non-restrictive in all respects, and the scope of the present invention is defined by the appended claims rather than the above description, and it is intended that all changes falling within the meaning and scope of the equivalent elements of the claims be included in the present invention.

Claims

1. An AI voice payment system based on the Internet, comprising a voice input module, a voice recognition module, a natural language understanding module, a security verification module, a payment execution module and a feedback module, characterized in that: The voice input module is responsible for collecting the user's voice input signal, the voice recognition module converts the voice signal into text information, the natural language understanding module performs semantic analysis on the text information and understands the user's payment intention, the security verification module ensures the security of the user's identity, including verification methods such as face recognition and voiceprint recognition, the payment execution module connects with the bank and payment platform through the payment interface to perform actual payment operations, the feedback module returns the payment result to inform the user whether the payment is successful and information such as the account balance, the security verification module also includes an abnormal voice detection module, a risk scoring module and a real-time monitoring module, the abnormal voice detection module detects abnormal voice through a machine learning algorithm, the risk scoring module performs risk scoring based on user behavior and voice characteristics, and blocks high-risk transactions in real time, and the real-time monitoring module is used to monitor abnormal behavior during the payment process in real time, and to provide timely warnings and processing.

2. The Internet-based AI voice payment system according to claim 1, characterized in that: The voice input module uses a microphone with high sensitivity and dynamic range to capture the user's voice, converts analog signals into digital signals, identifies and removes the silent part of the voice signal, enhances the high-frequency part of the voice, analyzes the characteristics of ambient noise, estimates the power spectrum of the noise, subtracts the noise spectrum from the spectrum of the voice signal, and uses a nonlinear processor filter to predict and eliminate echoes.

3. The Internet-based AI voice payment system according to claim 1, characterized in that: The speech recognition module includes converting the speech signal processed by the speech input module into a Mel-scale spectrum through Mel-frequency cepstral coefficients and extracting the cepstral coefficients, thereby retaining the main information of the speech, and then using a long short-term memory network acoustic model to model the relationship between the input audio signal and words or phonemes, outputting the score or state sequence of the acoustic model, and then using an N-gram language model to combine the output of the acoustic model and the language model to find the most likely word sequence through a decoding algorithm.

4. The Internet-based AI voice payment system according to claim 1, characterized in that: The natural language understanding module, namely the NLU module, after receiving the user's voice signal and converting it into text through the voice recognition module, first needs to perform a series of preprocessing operations on the original text, and then perform intent recognition to identify the specific operations the user wants to perform from the text. Intent recognition uses a neural network model learning algorithm, and then uses a conditional random field CRF pre-trained language model to fill in the slots. Finally, context understanding is performed to track the user's historical intentions and adjust the intent recognition results according to changes in the context.

5. The Internet-based AI voice payment system according to claim 1, characterized in that: The security verification module includes voiceprint recognition: verifying identity through the user's voice characteristics (such as tone, speaking speed, pronunciation method, etc.) and face recognition: capturing the user's facial image through a camera and comparing it with the facial features in the database.

6. The Internet-based AI voice payment system according to claim 1, characterized in that: The payment execution module calls the payment interface of the payment platform through the API, and the system passes the payment request and related information (such as amount, payee) to the payment platform. After the payment platform verifies the payment request, it returns the payment result. First, the system checks the user balance to ensure that the account balance is sufficient for payment. If the balance is sufficient, the payment platform deducts the money and completes the payment. If the balance is insufficient, the system returns a prompt message indicating insufficient balance.

7. The Internet-based AI voice payment system according to claim 1, characterized in that: The feedback module generates feedback information according to the payment execution result and displays it to the user in the form of voice or text.

8. A method for using an Internet-based AI voice payment system, characterized in that: The following steps are involved: Step 1: User voice input The user enters payment instructions through voice assistants such as smartphones, smart speakers and other devices, and the user says: "Pay 100 yuan to Zhang San." Step 2: Voice Recognition The system converts the user's voice signal into text through the voice recognition module. The voice command "Pay 100 yuan to Zhang San" will be converted into text: "Pay 100 yuan to Zhang San". Step 3: Natural Language Understanding The system inputs the text into the NLU module for analysis, extracting the user's payment intention, payment amount, and payee information: Payment Intent: Pay Payment amount: 100 yuan Payee: Zhang San. Step 4: Security Verification According to the security policy, the system requires users to authenticate their identities, which may be voiceprint recognition or face recognition. If the verification is successful, the system continues with the payment operation; if the verification fails, the payment is rejected. Step 5: Payment Execution The system checks the user's account balance to ensure that the balance is sufficient for payment. If the balance is sufficient, the system completes the payment operation through the payment platform API. After the payment is completed, the system updates the user's account balance and returns the payment status. Step 6: Feedback The system will provide feedback to the user on the payment result, including whether the payment is successful, the payment amount, the account balance, etc. If the payment is successful, the user will receive feedback that "payment is successful, the account balance is 400 yuan"; if the payment fails, the system will provide the reason for the failure, such as insufficient balance, etc.

Citation Information

Patent Citations

  • On-site calling management voice assistant system, management method, medium and terminal

    CN114596852A

  • Event triggering processing method based on voice recognition

    CN118486305A

  • Payment method and device

    CN118608141A

  • Payment security system based on AI artificial intelligence and use method

    CN118967134A

  • Method and device applying artificial intelligence to send money by using voice input

    US20180144346A1

Cited By

  • AI mobile banking construction method and system based on artificial intelligence algorithm

    CN122222714A