system

The voice-based payment system addresses the inconvenience of QR codes by using speech recognition and natural language processing for secure and emotional intelligence, ensuring efficient and secure transactions.

JP2026068313APending Publication Date: 2026-04-22SOFTBANK GROUP CORP
View PDF 1 Cites 0 Cited by

Patent Information

Authority / Receiving Office
JP · JP
Patent Type
Applications
Current Assignee / Owner
SOFTBANK GROUP CORP
Filing Date
2024-10-10
Publication Date
2026-04-22

Smart Images

  • Figure 2026068313000001_ABST
    Figure 2026068313000001_ABST
Patent Text Reader

Abstract

We provide the system. [Solution] A voice acquisition means for receiving voice input, A speech recognition means that converts speech data acquired by the speech acquisition means into text data, A natural language processing means that analyzes the text data generated by the speech recognition means to extract transaction details, An authentication method that authenticates transactions based on extracted transaction details, A payment method that executes settlement processing when a transaction is authenticated, A notification method for notifying the transaction results, A system that includes this.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The technology of the present disclosure relates to a system.

Background Art

[0002] Patent Document 1 discloses a method for controlling a persona chatbot, which is performed by at least one processor, the method including steps of receiving a user utterance, adding the user utterance to a prompt including an instruction sentence related to an explanation of a character of the chatbot, encoding the prompt, and inputting the encoded prompt into a language model to generate a chatbot utterance in response to the user utterance.

Prior Art Documents

Patent Documents

[0003]

Patent Document 1

Summary of the Invention

Problems to be Solved by the Invention

[0004] Conventional payment methods often use QR codes (registered trademarks), which may be time-consuming for reading and displaying QR codes. This time-consuming process is particularly inconvenient for visually impaired or elderly people and hinders a smooth payment experience. Also, QR codes may sometimes not be read properly or scanning may become difficult due to environmental factors. Due to these problems, the speed and accuracy of information transmission are impaired, so a more convenient and universal payment method using only voice is required.

Means for Solving the Problems

[0005] This invention is characterized by the use of speech recognition technology that acquires speech and converts the speech data into text. Furthermore, it utilizes natural language processing technology to analyze the converted text data and understand the transaction details. This accurately grasps the user's payment intention and verifies the user's identity through an authentication method. After authentication, such as voice authentication, is completed, the transaction is confirmed using the payment method, and the user is notified of the completion of the payment using a notification method. Through this entire process, it is possible to provide an intuitive and smooth payment experience using only voice.

[0006] "Voice acquisition means" refers to a device or function for receiving voices emitted by a user and incorporating that voice data into the system.

[0007] "Speech recognition means" refers to a device or function that analyzes acquired speech data and converts it into text data.

[0008] "Natural language processing means" refers to a device or function that analyzes text data generated by speech recognition and extracts transaction details.

[0009] "Authentication means" refers to a device or function used to verify the legitimacy of a transaction and to verify the identity of the user.

[0010] A "payment method" is a device or function for executing payment processing based on authenticated transaction details.

[0011] "Notification means" refers to a device or function for communicating the results of a transaction to the user. [Brief explanation of the drawing]

[0012] [Figure 1] This is a conceptual diagram showing an example of the configuration of a data processing system according to the first embodiment. [Figure 2] This is a conceptual diagram showing an example of the essential functions of a data processing device and a smart device according to the first embodiment. [Figure 3]It is a conceptual diagram showing an example of the configuration of a data processing system according to the second embodiment. [Figure 4] It is a conceptual diagram showing an example of the main functions of a data processing device and smart glasses according to the second embodiment. [Figure 5] It is a conceptual diagram showing an example of the configuration of a data processing system according to the third embodiment. [Figure 6] It is a conceptual diagram showing an example of the main functions of a data processing device and a headset-type terminal according to the third embodiment. [Figure 7] It is a conceptual diagram showing an example of the configuration of a data processing system according to the fourth embodiment. [Figure 8] It is a conceptual diagram showing an example of the main functions of a data processing device and a robot according to the fourth embodiment. [Figure 9] It shows an emotion map to which a plurality of emotions are mapped. [Figure 10] It shows an emotion map to which a plurality of emotions are mapped. [Figure 11] It is a sequence diagram showing the processing flow of the data processing system in Example 1. [Figure 12] It is a sequence diagram showing the processing flow of the data processing system in Application Example 1. [Figure 13] It is a sequence diagram showing the processing flow of the data processing system in Example 2 when an emotion engine is combined. [Figure 14] It is a sequence diagram showing the processing flow of the data processing system in Application Example 2 when an emotion engine is combined.

Embodiments for Carrying Out the Invention

[0013] Hereinafter, an example of an embodiment of a system according to the technology of the present disclosure will be described with reference to the accompanying drawings.

[0014] First, the language used in the following description will be explained.

[0015] In the following embodiments, the numbered processor (hereinafter simply referred to as "processor") may be a single arithmetic unit or a combination of multiple arithmetic units. Also, the processor may be a single type of arithmetic unit or a combination of multiple types of arithmetic units. Examples of arithmetic units include a CPU (Central Processing Unit), a GPU (Graphics Processing Unit), a GPGPU (General-Purpose computing on Graphics Processing Units), an APU (Accelerated Processing Unit), and the like.

[0016] In the following embodiments, the numbered RAM (Random Access Memory) is a memory in which information is temporarily stored and is used as a work memory by the processor.

[0017] In the following embodiments, the numbered storage is one or more non-volatile storage devices that store various programs, various parameters, and the like. Examples of non-volatile storage devices include flash memory (SSD (Solid State Drive)), magnetic disks (e.g., hard disks), or magnetic tapes, and the like.

[0018] In the following embodiments, the numbered communication I / F (Interface) is an interface including a communication processor, an antenna, and the like. The communication I / F controls communication between multiple computers. Examples of communication standards applied to the communication I / F include wireless communication standards including 5G (5th Generation Mobile Communication System), Wi-Fi (registered trademark), or Bluetooth (registered trademark).

[0019] In the following embodiments, "A and / or B" is synonymous with "at least one of A and B." That is, "A and / or B" means that it may be A alone, or B alone, or a combination of A and B. Furthermore, in this specification, the same concept as "A and / or B" applies when expressing three or more things linked by "and / or."

[0020] [First Embodiment]

[0021] Figure 1 shows an example of the configuration of the data processing system 10 according to the first embodiment.

[0022] As shown in Figure 1, the data processing system 10 includes a data processing device 12 and a smart device 14. An example of the data processing device 12 is a server.

[0023] The data processing device 12 comprises a computer 22, a database 24, and a communication interface 26. The computer 22 is an example of a "computer" related to the technology of this disclosure. The computer 22 comprises a processor 28, RAM 30, and storage 32. The processor 28, RAM 30, and storage 32 are connected to a bus 34. The database 24 and the communication interface 26 are also connected to the bus 34. The communication interface 26 is connected to a network 54. An example of the network 54 is a WAN (Wide Area Network) and / or a LAN (Local Area Network).

[0024] The smart device 14 comprises a computer 36, a reception device 38, an output device 40, a camera 42, and a communication interface 44. The computer 36 comprises a processor 46, RAM 48, and storage 50. The processor 46, RAM 48, and storage 50 are connected to a bus 52. The reception device 38, output device 40, and camera 42 are also connected to the bus 52.

[0025] The reception device 38 is equipped with a touch panel 38A and a microphone 38B, etc., and receives user input. The touch panel 38A receives user input by detecting contact with an object (e.g., a pen or finger). The microphone 38B receives user input by detecting the user's voice. The control unit 46A transmits data indicating the user input received by the touch panel 38A and microphone 38B to the data processing device 12. In the data processing device 12, the specific processing unit 290 acquires the data indicating the user input.

[0026] The output device 40 includes a display 40A and a speaker 40B, and presents data to the user 20 by outputting the data in a form perceptible to the user 20 (e.g., audio and / or text). The display 40A displays visible information such as text and images according to instructions from the processor 46. The speaker 40B outputs audio according to instructions from the processor 46. The camera 42 is a small digital camera equipped with an optical system such as a lens, aperture, and shutter, and an image sensor such as a CMOS (Complementary Metal-Oxide-Semiconductor) image sensor or a CCD (Charge Coupled Device) image sensor.

[0027] Communication interface 44 is connected to network 54. Communication interfaces 44 and 26 are responsible for the exchange of various types of information between processor 46 and processor 28 via network 54.

[0028] Figure 2 shows an example of the main functions of the data processing device 12 and the smart device 14.

[0029] As shown in Figure 2, in the data processing device 12, a specific processing is performed by the processor 28. A specific processing program 56 is stored in the storage 32. The specific processing program 56 is an example of a "program" related to the technology of this disclosure. The processor 28 reads the specific processing program 56 from the storage 32 and executes the read specific processing program 56 on the RAM 30. The specific processing is realized by the processor 28 operating as a specific processing unit 290 according to the specific processing program 56 executed on the RAM 30.

[0030] The storage 32 stores the data generation model 58 and the emotion identification model 59. The data generation model 58 and the emotion identification model 59 are used by the identification processing unit 290.

[0031] In the smart device 14, the processor 46 performs the reception output processing. The storage 50 stores the reception output program 60. The reception output program 60 is used in conjunction with a specific processing program 56 by the data processing system 10. The processor 46 reads the reception output program 60 from the storage 50 and executes the read reception output program 60 on the RAM 48. The reception output processing is realized by the processor 46 operating as a control unit 46A according to the reception output program 60 executed on the RAM 48.

[0032] Next, the specific processing performed by the specific processing unit 290 of the data processing device 12 will be described. In the following description, the data processing device 12 will be referred to as the "server" and the smart device 14 as the "terminal".

[0033] This invention is a system for making payments by voice, and its configuration includes a user, a terminal, and a server.

[0034] User-side operations

[0035] After finding the product they want to purchase, the user communicates their intention to the device by voice. For example, they can say a voice command such as, "I want to buy this product."

[0036] Terminal operation

[0037] The terminal receives voice from the user using a highly sensitive microphone and immediately sends the acquired voice data to the server. The terminal compresses the voice into an appropriate format and transfers the data using secure communication. It then receives a reply from the server and provides feedback to the user.

[0038] Server Processing

[0039] The server first uses a speech recognition engine to convert the received audio data into text. This engine, based on a machine learning model, accurately captures what the user is saying as text. Next, the converted text data is analyzed using natural language processing (NLP) technology. NLP accurately identifies what the user is looking for (product name, purchase intention, etc.).

[0040] Subsequently, the server proceeds with the authentication process, verifying the user's identity based on the acquired information. This authentication can utilize voice authentication or other authentication methods. If authentication is successful, the server immediately executes the payment process. The payment process uses the user's registered payment information to ensure a secure and fast transaction.

[0041] Finally, the server sends a notification to the device to inform the user that the payment has been completed. The device then communicates this notification to the user via voice or on-screen display. A message such as "Purchase complete" is sent to the user.

[0042] Thus, this technology provides high convenience by quickly understanding the user's intent through voice recognition and completing transactions safely and efficiently. For example, when purchasing a product at a coffee shop, the user can say "I want to buy a coffee" by voice, and the system can automatically carry out the coffee purchase procedure according to this flow.

[0043] The following describes the processing flow.

[0044] Step 1:

[0045] The user communicates their intention to purchase the item by voice to the device. For example, they might say, "I will purchase this item."

[0046] Step 2:

[0047] The device acquires the user's voice in real time through the microphone and captures it as audio data.

[0048] Step 3:

[0049] The device transmits the acquired audio data to the server using a secure protocol. During this process, the data is encrypted and compressed.

[0050] Step 4:

[0051] The server passes the received audio data to the speech recognition engine, which converts it into high-precision text data.

[0052] Step 5:

[0053] The server analyzes the text data using natural language processing technology to extract the user's intent (product name, transaction details, etc.).

[0054] Step 6:

[0055] The server authenticates the user based on the transaction details. It uses voice authentication or other authentication methods to verify the user's identity.

[0056] Step 7:

[0057] Upon successful authentication, the server processes the payment. It uses the registered payment information to execute a secure transaction.

[0058] Step 8:

[0059] The server sends the payment completion status to the terminal, including information on whether the payment was successful.

[0060] Step 9:

[0061] The device receives a notification from the server and informs the user of the completion of the payment via voice or screen display. For example, it might notify the user that "Your purchase is complete."

[0062] (Example 1)

[0063] Next, we will describe Example 1. In the following description, the data processing device 12 will be referred to as the "server," and the smart device 14 will be referred to as the "terminal."

[0064] Traditional payment systems have limitations in terms of user input methods, resulting in low convenience and difficulty in processing transactions quickly. Furthermore, they suffer from insufficient security and user authentication accuracy, leading to a lack of transaction reliability.

[0065] The identification process performed by the identification processing unit 290 of the data processing device 12 in Example 1 is realized by the following means.

[0066] In this invention, the server includes an information acquisition means for receiving voice input, a voice recognition means for converting the voice information acquired by the information acquisition means into text information, and a language processing means for analyzing the text information generated by the voice recognition means and extracting its content. This enables the rapid and highly accurate processing of the user's voice, making secure and reliable payment possible.

[0067] "Information acquisition means" refers to a device or method for receiving user input, and in this case, it refers to technology for receiving voice input.

[0068] "Speech recognition means" refers to technology that analyzes speech information and converts it into text information, and has the function of accurately converting speech to text using machine learning models, etc.

[0069] "Language processing means" refers to techniques for further analyzing generated textual information and extracting necessary transaction details from it, and this involves using natural language processing.

[0070] "Authentication methods" refer to technologies used to verify the identity of users in order to ensure the security of transactions, and include voice authentication and biometric authentication.

[0071] A "payment method" is a system that, after a transaction has been authenticated, executes the payment process using the user's registered payment information.

[0072] "Notification means" refers to technologies for communicating transaction results to users, and includes methods for providing information through voice or screen display.

[0073] "Communication methods" refer to technologies for securely transmitting data by compressing it into an appropriate format, and utilize secure communication protocols.

[0074] This invention is a voice-based payment system in which a user, a terminal, and a server work together. The user gives voice instructions for purchasing goods, and the terminal is equipped with information acquisition means to receive this voice. The terminal collects voice using a high-sensitivity microphone and has communication means to transmit the data to the server, using a secure communication protocol.

[0075] The server converts speech information into text information as a speech recognition tool, for example, by using a machine learning-based speech recognition engine. Furthermore, the server uses language processing tools to analyze this text information using natural language processing technology (e.g., spaCy) and extract transaction details. Subsequently, identity verification is performed using authentication tools. User authentication can be performed quickly and securely using voice authentication or biometric authentication technology. After successful authentication, the server executes a secure transaction using the user's registered payment information with payment tools. It is desirable to use a payment service (e.g., Stripe) at this stage.

[0076] Once a transaction is complete, the server notifies the terminal of the result using a notification system, and the terminal communicates this information to the user via voice or screen display. This notification allows the user to confirm that the transaction was completed successfully.

[0077] As a concrete example, if a user says "I want to buy a coffee" by voice at a coffee shop, this system can automatically complete the coffee purchase process. The system analyzes and authenticates the voice, completes the payment, and notifies the user of the result by voice or display.

[0078] An example of a prompt in a generative AI model is a request for a detailed explanation of the system, such as, "Please explain in detail the processing flow of the system for purchasing products by voice."

[0079] The flow of the specific processing in Example 1 will be explained using Figure 11.

[0080] Step 1:

[0081] The user communicates their intention to purchase a product via voice. Specifically, the user issues a voice command such as "I want to buy this product" into the device. The input is the user's voice, which the device receives. The output is voice data.

[0082] Step 2:

[0083] The terminal collects the received audio using an information acquisition device and employs a high-sensitivity microphone. Subsequently, this audio data is converted into a compressed format and transmitted to the server using a secure communication method. The input is audio data, and the output is compressed audio data.

[0084] Step 3:

[0085] The server converts the received audio data into text information using a machine learning-based speech recognition engine as a speech recognition method. In this process, the speech recognition model analyzes the audio and obtains output as text data. The input is compressed audio data, and the output is text data.

[0086] Step 4:

[0087] The server uses language processing tools to analyze the generated text data through natural language processing techniques to identify transaction details. Specifically, it performs data calculations to identify product names and transaction intent. The input is text data, and the output is the analyzed transaction information.

[0088] Step 5:

[0089] The server uses authentication methods to verify the user's identity. This process utilizes voice authentication and biometric authentication technologies to enhance the accuracy of user identification and authentication. The input is the analyzed transaction information, and the output is the authentication result.

[0090] Step 6:

[0091] If the transaction is authenticated, the server executes the purchase process using the payment method. Specifically, it completes the economic transaction through a secure payment service. The input is the authentication result, and the output is information about the completion of the payment.

[0092] Step 7:

[0093] The server notifies the terminal that the payment has been completed via a notification system. The terminal then informs the user of this information via voice or display. The input is the payment completion information, and the output is the transaction completion notification to the user.

[0094] (Application Example 1)

[0095] Next, we will explain Application Example 1. In the following explanation, the data processing device 12 will be referred to as the "server," and the smart device 14 will be referred to as the "terminal."

[0096] In modern brick-and-mortar stores, there is a demand for customers to complete purchases quickly and efficiently. However, current methods result in waiting times at the checkout, reducing the efficiency of the shopping experience. In particular, there is a need for a new interface that allows payment to be completed by voice without going through a physical checkout system.

[0097] The specific processing performed by the specific processing unit 290 of the data processing device 12 in Application Example 1 is realized by the following means.

[0098] In this invention, the server includes sound acquisition means, sound recognition means, natural language processing means, settlement means, and notification means. This allows users to quickly complete purchase procedures using voice input in physical stores, and after confirming the notified transaction result, they can leave the store with their purchased items without going through a cash register.

[0099] "Sound acquisition means" refers to a device or system that receives audio emitted by a user and converts it into data for subsequent processing.

[0100] "Acoustic recognition means" refers to a processing device that has the function of analyzing acquired audio data and converting it into a corresponding text format.

[0101] "Natural language processing means" refers to a technology that analyzes text data acquired by acoustic recognition means to extract user intent and transaction details.

[0102] "Verification methods" refer to the process of verifying the legitimacy of a transaction and confirming the user's identity based on the extracted transaction details.

[0103] A "settlement instrument" is a device or system that, based on the results of a verification instrument, executes settlement processing to complete a transaction.

[0104] A "notification system" is a system that has audio and visual notification functions to communicate the results of a transaction to the user.

[0105] This invention is a system that enables instant voice payment in physical stores. Users use a terminal such as a smartphone. The terminal uses a high-sensitivity microphone to acquire voice data. This voice data is converted into text data by an acoustic recognition means, and the transaction details are further analyzed by a natural language processing means.

[0106] To facilitate these processes, the server utilizes a speech recognition engine (e.g., Google® Speech-to-Text API) to generate text data. The analyzed content is then used to verify the user's identity and determine the validity of the transaction using voice authentication or other methods. Once authentication is complete, a secure payment process is executed using a settlement method. This uses a payment gateway API (e.g., Stripe).

[0107] Once payment is complete, a notification system will alert the user to the transaction result. This notification will be made via audio or on-screen display. It will inform the user that the purchase of the product has been completed in the physical store and confirm that the transaction has been successfully completed.

[0108] As a concrete example, when a user at a cafe says "I'll buy this coffee" using voice input on their smartphone, the device connects with a server and completes the payment process. The server authenticates the transaction using a verification mechanism and notifies the user via a notification mechanism that "the purchase is complete."

[0109] An example of a prompt message would be: "Design a program that uses a speech recognition system to convert what the user says into text in real time and then processes the payment based on that information. Specifically, please show the flow when the user says, 'I will buy this product.'"

[0110] The flow of a specific process in Application Example 1 will be explained using Figure 12.

[0111] Step 1:

[0112] The user makes a voice input near the smartphone. The device receives this voice using a high-sensitivity microphone and acquires it as digital audio data. The input is the user's voice, and the output is digital audio data.

[0113] Step 2:

[0114] The terminal immediately transmits the acquired digital audio data to the server. During this process, the terminal compresses the audio data and transfers it using a secure communication method. The input is digital audio data, and the output is the transmitted audio data.

[0115] Step 3:

[0116] The server uses a speech recognition engine to convert received audio data into text data. This process utilizes a generative AI model to convert speech to text. The input is audio data, and the output is text data.

[0117] Step 4:

[0118] The server performs natural language processing on the text data to extract the transaction details intended by the user. Through this analysis, product names and purchase intentions are identified. The input is text data, and the output is the extracted transaction details.

[0119] Step 5:

[0120] The server uses verification methods to authenticate the user based on the extracted transaction details. This process utilizes voice authentication and other methods to determine the user's legitimacy. The input is the transaction details, and the output is the authentication result.

[0121] Step 6:

[0122] If the verification process is successful, the server will use the settlement method to process the payment. In this process, the user's payment information is used to execute a secure transaction via the payment gateway API. The input is the authenticated transaction details, and the output is the settlement result.

[0123] Step 7:

[0124] The server transmits the payment result to the terminal via a notification system. The terminal notifies the user of a message such as "Purchase complete" via voice or screen display. The input is the payment result, and the output is the notification to the user.

[0125] Furthermore, an emotion engine that estimates the user's emotions may be incorporated. That is, the identification processing unit 290 may use the emotion identification model 59 to estimate the user's emotions and perform identification processing using the user's emotions.

[0126] This invention incorporates an emotion engine that recognizes user emotions into a voice-based payment system. Its aim is to enhance the conventional process of understanding transaction intent from voice input and to provide a more human-like interaction and experience by taking the user's emotional state into consideration during the payment process.

[0127] User operation

[0128] The user expresses their intention to purchase the product by voice, as is typical. For example, they might say, "I want to buy this product." The user's emotions, such as joy or anxiety, may also be included in their voice.

[0129] Terminal operation

[0130] The terminal captures the user's voice directly and transfers it to the server. The terminal's processing involves the clear and accurate collection and transmission of the voice data.

[0131] Processing on the server

[0132] The server uses speech recognition to convert the audio data received from the terminal into text data. This text data is then analyzed using natural language processing (NLP) techniques to confirm the user's purchase intent. Next, the emotion engine identifies emotions from the user's audio data. For example, it infers the user's emotions from changes in tone, volume, and tempo of their voice.

[0133] Once emotions are identified, that data is associated with the transaction intent and passed to the authentication mechanism. The authentication process, in addition to standard identity verification, considers the user's emotional state to perform a more precise intent check. If authentication is successful, the server proceeds with the payment process and securely completes the transaction.

[0134] Notification process

[0135] Once the payment is complete, the server uses the emotion engine's analysis results to notify the device of the outcome using an appropriate tone. The device then verbally communicates this information to the user. For example, if the user is nervous, the server can deliver a message in a calm tone such as, "Your purchase has been successfully completed."

[0136] This invention allows users to enjoy more personalized interactions that take their emotions into account, making the service experience more comfortable and reassuring. A concrete example is a scenario where the system provides reassuring voice support to a user who is feeling anxious while shopping.

[0137] The following describes the processing flow.

[0138] Step 1:

[0139] The user expresses the product they want to purchase by voice, using phrases such as "I want to buy this product."

[0140] Step 2:

[0141] The device acquires the user's voice through the microphone and captures it as audio data.

[0142] Step 3:

[0143] The device sends the acquired audio data to the server. The data is encrypted during this process.

[0144] Step 4:

[0145] The server analyzes the received audio data using a speech recognition engine and converts the audio into text.

[0146] Step 5:

[0147] The server analyzes the text data using natural language processing technology to identify the transaction details.

[0148] Step 6:

[0149] The server further analyzes the voice data using an emotion engine to recognize the user's emotions. This process includes analyzing the tone and volume of the voice.

[0150] Step 7:

[0151] The server performs an authentication process and approves the transaction, taking into account the user's identity and perceived emotions.

[0152] Step 8:

[0153] Upon successful authentication, the server processes the payment and securely completes the transaction.

[0154] Step 9:

[0155] The server generates an appropriate notification tone based on the payment result and the user's sentiment, and sends it to the device.

[0156] Step 10:

[0157] The terminal notifies the user of the payment result via voice. For example, it might say in a calm voice, "Your purchase is complete."

[0158] (Example 2)

[0159] Next, we will describe Example 2. In the following description, the data processing device 12 will be referred to as the "server" and the smart device 14 as the "terminal".

[0160] Conventional voice recognition systems can understand a user's transaction intent, but they do not consider the user's emotional state, which can result in a lack of human-like interaction and experience. Furthermore, in authentication processes, ignoring the user's emotional state can lead to insufficient identity verification. Therefore, there is a need for a voice payment system that takes user emotions into account and provides a more natural and secure transaction experience.

[0161] The identification process performed by the identification processing unit 290 of the data processing device 12 in Example 2 is realized by the following means.

[0162] In this invention, the server includes an input means for receiving voice, a conversion means for converting the voice acquired by the input means into text, a processing means for analyzing the text data generated by the conversion means to extract transaction details, an emotion analysis means for identifying emotions from the voice data acquired by the processing means, a verification means for performing authentication based on the extracted transaction details and identified emotions, an operation means for executing processing when the transaction is authenticated, and a communication means for notifying the processing results. This enables a more natural and secure voice-based transaction experience while taking into account the user's emotional state.

[0163] "Input means" refers to a function that acquires audio and provides it in a format that the system can process.

[0164] "Conversion means" refers to the process of converting acquired audio into text format and generating text data.

[0165] The "processing means" refers to a function that analyzes the generated text data to extract the user's transaction intent.

[0166] "Emotion analysis means" refers to a process for identifying a user's emotions from voice data and inferring their emotional state.

[0167] "Verification means" refers to the process of authentication based on extracted transaction details and identified sentiments.

[0168] "Operational means" refers to functions that execute settlement and other related processes once a transaction is authenticated.

[0169] "Means of communication" refers to the process of notifying the user of the processing results, and the results can be conveyed through means such as voice.

[0170] This invention incorporates an emotion analysis function that recognizes user emotions into a voice-based payment system. The aim is to understand the user's transaction intent and provide a natural dialogue and experience that takes their emotional state into account. Specific embodiments are described below.

[0171] User operation

[0172] When purchasing a product, users express their intentions verbally. For example, they might say, "I want to buy this product." The tone and tempo of their voice may reflect their emotions, such as joy or anxiety.

[0173] Terminal operation

[0174] The terminal uses a microphone or similar device to acquire the user's voice information and transmits it to the server as digital audio data. Clear and accurate audio capture is crucial on the terminal side.

[0175] Processing on the server

[0176] The server converts the received audio data into text data using a speech recognition engine. General-purpose speech recognition software, such as a "speech recognition API," is used for this process. The text data is then analyzed using natural language processing techniques to confirm the user's transaction intent. Natural language processing libraries are used for this analysis.

[0177] Next, the emotion analysis engine analyzes changes in tone, intensity, and tempo of the voice to identify the user's emotions. For example, it uses an "emotion analysis tool" to infer emotions such as joy or anxiety. The authentication process provides highly accurate authentication using voice and emotion information.

[0178] Notification process

[0179] Based on the sentiment analysis results, the server creates a message in an appropriate tone and sends it back to the device. The device then informs the user of the message aloud. For example, if the user is nervous, it will play a calm message such as, "Your purchase has been successfully completed."

[0180] This system allows users to enjoy a consistent and personalized experience that takes their emotions into account. As a concrete example, consider the case of online shopping. An example of a prompt message might be: "Analyze the emotions expressed by the user during product purchase and advise on appropriate responses. Specifically, identify kindness, anxiety, and joy, and suggest corresponding voice messages."

[0181] In this way, the present invention makes it possible to provide a more enriching user experience by considering and analyzing emotions during the process of transactions conducted via voice.

[0182] The flow of the specific processing in Example 2 will be explained using Figure 13.

[0183] Step 1:

[0184] The user expresses their intention to purchase a product by voice. For example, they might say, "I want to buy this product." This voice information is acquired by the terminal and sent to the server as digital voice data. The terminal captures the voice input using a microphone and sends it to the server in the form of an audio file. During this process, efforts are made to minimize noise and maintain the clarity and accuracy of the voice.

[0185] Step 2:

[0186] The server converts the audio data received from the terminal into text data using a speech recognition engine. Specifically, it analyzes the audio waveform using speech recognition software and converts it into a corresponding string of characters. Through this process, the text "I want to buy this product" is obtained from the audio.

[0187] Step 3:

[0188] The server analyzes the converted text data using natural language processing techniques to extract transaction intent. For example, it uses a natural language processing library to perform grammatical analysis and keyword extraction to identify the intent to "purchase a product." Text data is used as input, and intent and related information are included in the output.

[0189] Step 4:

[0190] The server uses an emotion analysis engine to identify emotions from audio data. Using an audio analysis tool, it evaluates the tone, intensity, and tempo of the voice to infer the emotional state (e.g., joy, anxiety). The emotion analysis engine takes audio data as input and outputs emotion tags such as "joy."

[0191] Step 5:

[0192] The server performs authentication based on identified transaction details and emotions. The verification process uses voice authentication technology to verify identity and enhance reliability. For example, it compares voiceprints and outputs a result indicating whether voice authentication was successful or not.

[0193] Step 6:

[0194] If the transaction is authenticated, the server will execute the payment through the payment processing service. Payment information is entered, and the transaction ID and payment status are received as output. Encryption technology may be used to ensure the security of the payment process.

[0195] Step 7:

[0196] The server creates a message based on the sentiment analysis results and sends it to the device. The tone of the message is adjusted according to the user's emotions, and it is presented in a way that is easy for the user to understand. The device notifies the user of the received message via voice, confirming the completion of the purchase. For example, if the user is nervous, a message such as, "Your purchase has been successfully completed. Please rest assured," is played in a gentle tone.

[0197] (Application Example 2)

[0198] Next, we will explain application example 2. In the following explanation, the data processing device 12 will be referred to as the "server," and the smart device 14 will be referred to as the "terminal."

[0199] In voice-based payment systems, there is a need to provide sophisticated communication and reassuring interactions that take into account not only the transaction details intended by the user but also their emotional state at the time. However, conventional voice payment systems have the challenge of not being able to adequately reflect the user's emotional state and not being able to provide sufficient reassurance to users who feel anxious.

[0200] The specific processing performed by the specific processing unit 290 of the data processing device 12 in Application Example 2 is realized by the following means.

[0201] In this invention, the server includes data acquisition means for receiving voice input, voice recognition means for converting voice data into text data, and emotion analysis means for analyzing the user's emotional state. This makes it possible to analyze the transaction content and emotions from the user's voice and provide voice feedback that takes into account a sense of security.

[0202] "Data acquisition means" refers to a device or function that receives voice input and collects it as voice data.

[0203] "Speech recognition means" refers to a technology or system that analyzes acquired speech data and converts it into corresponding text data.

[0204] "Natural language processing means" refers to software or algorithms that analyze text data generated by speech recognition means to accurately extract transaction details.

[0205] "Emotion analysis tools" are technologies that analyze a user's emotional state from their speech and identify emotions based on changes in voice tone and tempo.

[0206] "Authentication means" refers to a process or function that verifies the validity of a transaction based on the transaction details and emotional state, and verifies the user's identity as necessary.

[0207] A "data processing system" is a system for executing authenticated transactions and ensuring settlement.

[0208] An "information transmission means" is a mechanism that transmits the results of a transaction to the user and notifies them through voice or other interfaces.

[0209] This invention is a system that, when conducting transactions based on user voice input, analyzes the user's emotions using voice data and provides personalized feedback based on that analysis. The server is responsible for acquiring, processing, and analyzing the voice data, confirming the transaction details, and then authenticating and executing the payment.

[0210] The terminal captures high-quality voice input from the user and sends it to the server via a data acquisition method. The server uses the Google Speech-to-Text API to convert the speech to text and extracts transaction details through natural language processing. Furthermore, it uses the Google Cloud Natural Language API to perform sentiment analysis based on factors such as tone and intensity of the speech.

[0211] The analysis results are used for identity verification through authentication methods, and if the transaction is successfully authenticated, the payment process is executed via the Stripe API. Based on the sentiment results obtained during this process, feedback is provided to the user in an appropriate tone as a means of communication.

[0212] For example, if a user says in voice, "I want to buy this product, but I'm a little worried," the emotion analysis system can recognize the user's anxiety and respond with a calm voice message such as, "Please relax, the purchase process is going smoothly."

[0213] Example prompt: "The user is attempting to purchase a specific product with an anxious voice. The sentiment analysis engine should recognize this anxiety and generate reassuring voice feedback."

[0214] The flow of a specific process in Application Example 2 will be explained using Figure 14.

[0215] Step 1:

[0216] The user performs voice input. The user uses their smartphone to express their intention to purchase a product by voice. The input is the user's speech and is recorded directly on the device.

[0217] Step 2:

[0218] The terminal acquires audio data and sends it to the server. The terminal uses its built-in microphone to save the audio input as a digital audio file and sends it to the server via a data acquisition device. The input is audio data, and the output is a digital audio file sent to the server.

[0219] Step 3:

[0220] The server converts audio data into text using speech recognition technology. Using the Google Speech-to-Text API, the server generates corresponding text data from the audio data. The input is audio data, and the output is text data.

[0221] Step 4:

[0222] The server analyzes text data using natural language processing techniques and extracts transaction details. The server applies natural language processing algorithms to extract transaction intent and product characteristics from the text. The input is text data, and the output is structured data containing transaction details.

[0223] Step 5:

[0224] The server identifies the user's emotional state using sentiment analysis tools. Leveraging the Google Cloud Natural Language API, the server analyzes changes in tone and tempo of speech to determine the emotional state. The input is audio data, and the output is information about the user's emotional state.

[0225] Step 6:

[0226] The server uses authentication methods to authenticate the transaction. It takes into account the transaction details and sentiment status, and verifies the user's identity through the authentication process. The input is the transaction details and sentiment information, and the output is the authentication result.

[0227] Step 7:

[0228] If the server authenticates the transaction, it initiates the settlement process using data processing tools. The settlement is securely executed using the Stripe API. The input is the authentication result, and the output is a settlement completion notification.

[0229] Step 8:

[0230] The server notifies the user of the payment result via a communication method. The server generates voice feedback reflecting the sentiment analysis results and delivers it to the user through the terminal. The input is the payment completion notification and sentiment information, and the output is a personalized voice message.

[0231] The specific processing unit 290 transmits the result of the specific processing to the smart device 14. In the smart device 14, the control unit 46A causes the output device 40 to output the result of the specific processing. The microphone 38B acquires audio indicating user input for the result of the specific processing. The control unit 46A transmits the audio data indicating user input acquired by the microphone 38B to the data processing device 12. In the data processing device 12, the specific processing unit 290 acquires the audio data.

[0232] Data generation model 58 is a so-called generative AI (Artificial Intelligence). An example of data generation model 58 is ChatGPT (registered trademark) (Internet search).<URL: https: / / openai.com / blog / chatgpt> ), Gemini (registered trademark) (Internet search) <url: https: gemini.google.com ?hl="ja">Examples of generative AI include the following. The data generation model 58 is obtained by performing deep learning on a neural network. The data generation model 58 is input with prompts containing instructions, and with inference data such as audio data representing speech, text data representing text, and image data representing images. The data generation model 58 infers from the input inference data according to the instructions indicated by the prompts, and outputs the inference results in data formats such as audio data and text data. Here, inference refers to, for example, analysis, classification, prediction, and / or summarization.

[0233] In the above embodiment, an example was given in which specific processing is performed by the data processing device 12, but the technology of this disclosure is not limited thereto, and the specific processing may also be performed by the smart device 14.

[0234] [Second Embodiment]

[0235] Figure 3 shows an example of the configuration of the data processing system 210 according to the second embodiment.

[0236] As shown in Figure 3, the data processing system 210 includes a data processing device 12 and smart glasses 214. An example of the data processing device 12 is a server.

[0237] The data processing device 12 comprises a computer 22, a database 24, and a communication interface 26. The computer 22 is an example of a "computer" related to the technology of this disclosure. The computer 22 comprises a processor 28, RAM 30, and storage 32. The processor 28, RAM 30, and storage 32 are connected to a bus 34. The database 24 and the communication interface 26 are also connected to the bus 34. The communication interface 26 is connected to a network 54. An example of the network 54 is a WAN (Wide Area Network) and / or a LAN (Local Area Network).

[0238] The smart glasses 214 include a computer 36, a microphone 238, a speaker 240, a camera 42, and a communication interface 44. The computer 36 includes a processor 46, RAM 48, and storage 50. The processor 46, RAM 48, and storage 50 are connected to a bus 52. The microphone 238, speaker 240, and camera 42 are also connected to the bus 52.

[0239] The microphone 238 receives voice signals from the user 20 and receives instructions from the user 20. The microphone 238 captures the voice signals from the user 20, converts the captured voice into audio data, and outputs it to the processor 46. The speaker 240 outputs audio according to the instructions from the processor 46.

[0240] Camera 42 is a small digital camera equipped with an optical system including a lens, aperture, and shutter, and an image sensor such as a CMOS (Complementary Metal-Oxide-Semiconductor) image sensor or a CCD (Charge Coupled Device) image sensor, and captures images of the area around the user 20 (for example, an imaging range defined by a field of view equivalent to the width of a typical healthy person's field of vision).

[0241] Communication interface 44 is connected to network 54. Communication interfaces 44 and 26 are responsible for the exchange of various information between processor 46 and processor 28 via network 54. The exchange of various information between processor 46 and processor 28 using communication interfaces 44 and 26 is performed in a secure manner.

[0242] Figure 4 shows an example of the main functions of the data processing device 12 and the smart glasses 214. As shown in Figure 4, the data processing device 12 performs specific processing using the processor 28. The storage 32 stores the specific processing program 56.

[0243] The specific processing program 56 is an example of a "program" relating to the technology of this disclosure. The processor 28 reads the specific processing program 56 from the storage 32 and executes the read specific processing program 56 on the RAM 30. The specific processing is realized by the processor 28 operating as a specific processing unit 290 in accordance with the specific processing program 56 executed on the RAM 30.

[0244] The storage 32 stores the data generation model 58 and the emotion identification model 59. The data generation model 58 and the emotion identification model 59 are used by the identification processing unit 290.

[0245] In the smart glasses 214, the processor 46 performs the reception output processing. The storage 50 stores the reception output program 60. The processor 46 reads the reception output program 60 from the storage 50 and executes the read reception output program 60 on the RAM 48. The reception output processing is realized by the processor 46 operating as a control unit 46A according to the reception output program 60 executed on the RAM 48.

[0246] Next, the identification processing performed by the identification processing unit 290 of the data processing device 12 will be described. In the following description, the data processing device 12 will be referred to as the "server" and the smart glasses 214 will be referred to as the "terminal".

[0247] This invention is a system for making payments by voice, and its configuration includes a user, a terminal, and a server.

[0248] User-side operations

[0249] After finding the product they want to purchase, the user communicates their intention to the device by voice. For example, they can say a voice command such as, "I want to buy this product."

[0250] Terminal operation

[0251] The terminal receives voice from the user using a highly sensitive microphone and immediately sends the acquired voice data to the server. The terminal compresses the voice into an appropriate format and transfers the data using secure communication. It then receives a reply from the server and provides feedback to the user.

[0252] Server Processing

[0253] The server first uses a speech recognition engine to convert the received audio data into text. This engine, based on a machine learning model, accurately captures what the user is saying as text. Next, the converted text data is analyzed using natural language processing (NLP) technology. NLP accurately identifies what the user is looking for (product name, purchase intention, etc.).

[0254] Subsequently, the server proceeds with the authentication process, verifying the user's identity based on the acquired information. This authentication can utilize voice authentication or other authentication methods. If authentication is successful, the server immediately executes the payment process. The payment process uses the user's registered payment information to ensure a secure and fast transaction.

[0255] Finally, the server sends a notification to the device to inform the user that the payment has been completed. The device then communicates this notification to the user via voice or on-screen display. A message such as "Purchase complete" is sent to the user.

[0256] Thus, this technology provides high convenience by quickly understanding the user's intent through voice recognition and completing transactions safely and efficiently. For example, when purchasing a product at a coffee shop, the user can say "I want to buy a coffee" by voice, and the system can automatically carry out the coffee purchase procedure according to this flow.

[0257] The following describes the processing flow.

[0258] Step 1:

[0259] The user communicates their intention to purchase the item by voice to the device. For example, they might say, "I will purchase this item."

[0260] Step 2:

[0261] The device acquires the user's voice in real time through the microphone and captures it as audio data.

[0262] Step 3:

[0263] The device transmits the acquired audio data to the server using a secure protocol. During this process, the data is encrypted and compressed.

[0264] Step 4:

[0265] The server passes the received audio data to the speech recognition engine, which converts it into high-precision text data.

[0266] Step 5:

[0267] The server analyzes the text data using natural language processing technology to extract the user's intent (product name, transaction details, etc.).

[0268] Step 6:

[0269] The server authenticates the user based on the transaction details. It uses voice authentication or other authentication methods to verify the user's identity.

[0270] Step 7:

[0271] Upon successful authentication, the server processes the payment. It uses the registered payment information to execute a secure transaction.

[0272] Step 8:

[0273] The server sends the payment completion status to the terminal, including information on whether the payment was successful.

[0274] Step 9:

[0275] The device receives a notification from the server and informs the user of the completion of the payment via voice or screen display. For example, it might notify the user that "Your purchase is complete."

[0276] (Example 1)

[0277] Next, we will describe Example 1. In the following description, the data processing device 12 will be referred to as the "server," and the smart glasses 214 will be referred to as the "terminal."

[0278] Traditional payment systems have limitations in terms of user input methods, resulting in low convenience and difficulty in processing transactions quickly. Furthermore, they suffer from insufficient security and user authentication accuracy, leading to a lack of transaction reliability.

[0279] The identification process performed by the identification processing unit 290 of the data processing device 12 in Example 1 is realized by the following means.

[0280] In this invention, the server includes an information acquisition means for receiving voice input, a voice recognition means for converting the voice information acquired by the information acquisition means into text information, and a language processing means for analyzing the text information generated by the voice recognition means and extracting its content. This enables the rapid and highly accurate processing of the user's voice, making secure and reliable payment possible.

[0281] "Information acquisition means" refers to a device or method for receiving user input, and in this case, it refers to technology for receiving voice input.

[0282] "Speech recognition means" refers to technology that analyzes speech information and converts it into text information, and has the function of accurately converting speech to text using machine learning models, etc.

[0283] The "language processing means" refers to a technology for further analyzing the generated character information and extracting necessary transaction contents therefrom, using natural language processing.

[0284] The "authentication means" is a technology for verifying the identity of the user to ensure the security of transactions, including voice authentication and biometric authentication.

[0285] The "settlement means" is a system that executes settlement processing using the user's registered payment information after the transaction is authenticated.

[0286] The "notification means" is a technology for communicating the transaction result to the user, including methods of providing information through voice or screen display.

[0287] The "communication means" refers to a technology for compressing data into an appropriate format and transmitting it securely, using a secure communication protocol.

[0288] This invention is a system for performing voice-based settlement, in which the user, terminal, and server cooperate to operate. The user gives an instruction to purchase a product by voice, and the terminal has information acquisition means for receiving this voice. The terminal collects voice using a high-sensitivity microphone, has communication means for transmitting data to the server, and uses a secure communication protocol.

[0289] The server converts voice information into character information as voice recognition means, for example, using a voice recognition engine based on machine learning. Furthermore, the server uses language processing means for this character information, analyzes it with natural language processing technology (e.g., spaCy), and extracts transaction contents. Subsequently, identity verification is performed by the authentication means. The user can be authenticated quickly and securely by voice authentication or biometric authentication technology. After the authentication is successful, the server executes a secure transaction using the user's registered payment information with the settlement means. At this time, it is desirable to use a settlement service (e.g., Stripe).

[0290] Once a transaction is complete, the server notifies the terminal of the result using a notification system, and the terminal communicates this information to the user via voice or screen display. This notification allows the user to confirm that the transaction was completed successfully.

[0291] As a concrete example, if a user says "I want to buy a coffee" by voice at a coffee shop, this system can automatically complete the coffee purchase process. The system analyzes and authenticates the voice, completes the payment, and notifies the user of the result by voice or display.

[0292] An example of a prompt in a generative AI model is a request for a detailed explanation of the system, such as, "Please explain in detail the processing flow of the system for purchasing products by voice."

[0293] The flow of the specific processing in Example 1 will be explained using Figure 11.

[0294] Step 1:

[0295] The user communicates their intention to purchase a product via voice. Specifically, the user issues a voice command such as "I want to buy this product" into the device. The input is the user's voice, which the device receives. The output is voice data.

[0296] Step 2:

[0297] The terminal collects the received audio using an information acquisition device and employs a high-sensitivity microphone. Subsequently, this audio data is converted into a compressed format and transmitted to the server using a secure communication method. The input is audio data, and the output is compressed audio data.

[0298] Step 3:

[0299] The server uses a speech recognition engine based on machine learning as the speech recognition means to convert the received speech data into character information. In this process, the speech recognition model analyzes the speech and obtains an output as text data. The input is compressed speech data, and the output is text data.

[0300] Step 4:

[0301] The server uses language processing means to analyze the generated text data through natural language processing technology to identify the transaction content. Specifically, data operations are performed to identify the product name and the intention of the transaction. The input is text data, and the output is the analyzed transaction information.

[0302] Step 5:

[0303] The server uses authentication means to verify the identity of the user. In this process, speech authentication and biometric authentication technologies are used to improve the accuracy of user identification and authentication. The input is the analyzed transaction information, and the output is the authentication result.

[0304] Step 6:

[0305] If the transaction is authenticated, the server uses settlement means to execute the purchase process. Specifically, economic transactions are completed through a secure settlement service. The input is the authentication result, and the output is the information indicating the completion of the settlement.

[0306] Step 7:

[0307] The server transmits to the terminal through notification means that the settlement has been completed. The terminal notifies the user of this information through voice or display. The input is the information indicating the completion of the settlement, and the output is a notification of the completion of the transaction to the user.

[0308] (Application Example 1)

[0309] Next, Application Example 1 will be described. In the following description, the data processing device 12 is referred to as the "server", and the smart glasses 214 are referred to as the "terminal".

[0310] In modern brick-and-mortar stores, there is a demand for customers to complete purchases quickly and efficiently. However, current methods result in waiting times at the checkout, reducing the efficiency of the shopping experience. In particular, there is a need for a new interface that allows payment to be completed by voice without going through a physical checkout system.

[0311] The specific processing performed by the specific processing unit 290 of the data processing device 12 in Application Example 1 is realized by the following means.

[0312] In this invention, the server includes sound acquisition means, sound recognition means, natural language processing means, settlement means, and notification means. This allows users to quickly complete purchase procedures using voice input in physical stores, and after confirming the notified transaction result, they can leave the store with their purchased items without going through a cash register.

[0313] "Sound acquisition means" refers to a device or system that receives audio emitted by a user and converts it into data for subsequent processing.

[0314] "Acoustic recognition means" refers to a processing device that has the function of analyzing acquired audio data and converting it into a corresponding text format.

[0315] "Natural language processing means" refers to a technology that analyzes text data acquired by acoustic recognition means to extract user intent and transaction details.

[0316] "Verification methods" refer to the process of verifying the legitimacy of a transaction and confirming the user's identity based on the extracted transaction details.

[0317] A "settlement instrument" is a device or system that, based on the results of a verification instrument, executes settlement processing to complete a transaction.

[0318] A "notification system" is a system that has audio and visual notification functions to communicate the results of a transaction to the user.

[0319] This invention is a system that enables instant voice payment in physical stores. Users use a terminal such as a smartphone. The terminal uses a high-sensitivity microphone to acquire voice data. This voice data is converted into text data by an acoustic recognition means, and the transaction details are further analyzed by a natural language processing means.

[0320] To facilitate these processes, the server utilizes a speech recognition engine (e.g., Google Speech-to-Text API) to generate text data. The analyzed content is then used to verify the user's identity and determine the validity of the transaction using voice authentication or other methods. Once authentication is complete, a secure payment process is executed using a settlement method. This uses a payment gateway API (e.g., Stripe).

[0321] Once payment is complete, a notification system will alert the user to the transaction result. This notification will be made via audio or on-screen display. It will inform the user that the purchase of the product has been completed in the physical store and confirm that the transaction has been successfully completed.

[0322] As a concrete example, when a user at a cafe says "I'll buy this coffee" using voice input on their smartphone, the device connects with a server and completes the payment process. The server authenticates the transaction using a verification mechanism and notifies the user via a notification mechanism that "the purchase is complete."

[0323] An example of a prompt message would be: "Design a program that uses a speech recognition system to convert what the user says into text in real time and then processes the payment based on that information. Specifically, please show the flow when the user says, 'I will buy this product.'"

[0324] The flow of a specific process in Application Example 1 will be explained using Figure 12.

[0325] Step 1:

[0326] The user makes a voice input near the smartphone. The device receives this voice using a high-sensitivity microphone and acquires it as digital audio data. The input is the user's voice, and the output is digital audio data.

[0327] Step 2:

[0328] The terminal immediately transmits the acquired digital audio data to the server. During this process, the terminal compresses the audio data and transfers it using a secure communication method. The input is digital audio data, and the output is the transmitted audio data.

[0329] Step 3:

[0330] The server uses a speech recognition engine to convert received audio data into text data. This process utilizes a generative AI model to convert speech to text. The input is audio data, and the output is text data.

[0331] Step 4:

[0332] The server performs natural language processing on the text data to extract the transaction details intended by the user. Through this analysis, product names and purchase intentions are identified. The input is text data, and the output is the extracted transaction details.

[0333] Step 5:

[0334] The server uses verification methods to authenticate the user based on the extracted transaction details. This process utilizes voice authentication and other methods to determine the user's legitimacy. The input is the transaction details, and the output is the authentication result.

[0335] Step 6:

[0336] If the verification process is successful, the server will use the settlement method to process the payment. In this process, the user's payment information is used to execute a secure transaction via the payment gateway API. The input is the authenticated transaction details, and the output is the settlement result.

[0337] Step 7:

[0338] The server transmits the payment result to the terminal via a notification system. The terminal notifies the user of a message such as "Purchase complete" via voice or screen display. The input is the payment result, and the output is the notification to the user.

[0339] Furthermore, an emotion engine that estimates the user's emotions may be incorporated. That is, the identification processing unit 290 may use the emotion identification model 59 to estimate the user's emotions and perform identification processing using the user's emotions.

[0340] This invention incorporates an emotion engine that recognizes user emotions into a voice-based payment system. Its aim is to enhance the conventional process of understanding transaction intent from voice input and to provide a more human-like interaction and experience by taking the user's emotional state into consideration during the payment process.

[0341] User operation

[0342] The user expresses their intention to purchase the product by voice, as is typical. For example, they might say, "I want to buy this product." The user's emotions, such as joy or anxiety, may also be included in their voice.

[0343] Terminal operation

[0344] The terminal captures the user's voice directly and transfers it to the server. The terminal's processing involves the clear and accurate collection and transmission of the voice data.

[0345] Processing on the server

[0346] The server uses speech recognition to convert the audio data received from the terminal into text data. This text data is then analyzed using natural language processing (NLP) techniques to confirm the user's purchase intent. Next, the emotion engine identifies emotions from the user's audio data. For example, it infers the user's emotions from changes in tone, volume, and tempo of their voice.

[0347] Once emotions are identified, that data is associated with the transaction intent and passed to the authentication mechanism. The authentication process, in addition to standard identity verification, considers the user's emotional state to perform a more precise intent check. If authentication is successful, the server proceeds with the payment process and securely completes the transaction.

[0348] Notification process

[0349] Once the payment is complete, the server uses the emotion engine's analysis results to notify the device of the outcome using an appropriate tone. The device then verbally communicates this information to the user. For example, if the user is nervous, the server can deliver a message in a calm tone such as, "Your purchase has been successfully completed."

[0350] This invention allows users to enjoy more personalized interactions that take their emotions into account, making the service experience more comfortable and reassuring. A concrete example is a scenario where the system provides reassuring voice support to a user who is feeling anxious while shopping.

[0351] The following describes the processing flow.

[0352] Step 1:

[0353] The user expresses the product they want to purchase by voice, using phrases such as "I want to buy this product."

[0354] Step 2:

[0355] The device acquires the user's voice through the microphone and captures it as audio data.

[0356] Step 3:

[0357] The device sends the acquired audio data to the server. The data is encrypted during this process.

[0358] Step 4:

[0359] The server analyzes the received audio data using a speech recognition engine and converts the audio into text.

[0360] Step 5:

[0361] The server analyzes the text data using natural language processing technology to identify the transaction details.

[0362] Step 6:

[0363] The server further analyzes the voice data using an emotion engine to recognize the user's emotions. This process includes analyzing the tone and volume of the voice.

[0364] Step 7:

[0365] The server performs an authentication process and approves the transaction, taking into account the user's identity and perceived emotions.

[0366] Step 8:

[0367] Upon successful authentication, the server processes the payment and securely completes the transaction.

[0368] Step 9:

[0369] The server generates an appropriate notification tone based on the payment result and the user's sentiment, and sends it to the device.

[0370] Step 10:

[0371] The terminal notifies the user of the payment result via voice. For example, it might say in a calm voice, "Your purchase is complete."

[0372] (Example 2)

[0373] Next, we will describe Example 2. In the following description, the data processing device 12 will be referred to as the "server" and the smart glasses 214 will be referred to as the "terminal".

[0374] Conventional voice recognition systems can understand a user's transaction intent, but they do not consider the user's emotional state, which can result in a lack of human-like interaction and experience. Furthermore, in authentication processes, ignoring the user's emotional state can lead to insufficient identity verification. Therefore, there is a need for a voice payment system that takes user emotions into account and provides a more natural and secure transaction experience.

[0375] The identification process performed by the identification processing unit 290 of the data processing device 12 in Example 2 is realized by the following means.

[0376] In this invention, the server includes an input means for receiving voice, a conversion means for converting the voice acquired by the input means into text, a processing means for analyzing the text data generated by the conversion means to extract transaction details, an emotion analysis means for identifying emotions from the voice data acquired by the processing means, a verification means for performing authentication based on the extracted transaction details and identified emotions, an operation means for executing processing when the transaction is authenticated, and a communication means for notifying the processing results. This enables a more natural and secure voice-based transaction experience while taking into account the user's emotional state.

[0377] "Input means" refers to a function that acquires audio and provides it in a format that the system can process.

[0378] "Conversion means" refers to the process of converting acquired audio into text format and generating text data.

[0379] The "processing means" refers to a function that analyzes the generated text data to extract the user's transaction intent.

[0380] "Emotion analysis means" refers to a process for identifying a user's emotions from voice data and inferring their emotional state.

[0381] "Verification means" refers to the process of authentication based on extracted transaction details and identified sentiments.

[0382] "Operational means" refers to functions that execute settlement and other related processes once a transaction is authenticated.

[0383] "Means of communication" refers to the process of notifying the user of the processing results, and the results can be conveyed through means such as voice.

[0384] This invention incorporates an emotion analysis function that recognizes user emotions into a voice-based payment system. The aim is to understand the user's transaction intent and provide a natural dialogue and experience that takes their emotional state into account. Specific embodiments are described below.

[0385] User operation

[0386] When purchasing a product, users express their intentions verbally. For example, they might say, "I want to buy this product." The tone and tempo of their voice may reflect their emotions, such as joy or anxiety.

[0387] Terminal operation

[0388] The terminal uses a microphone or similar device to acquire the user's voice information and transmits it to the server as digital audio data. Clear and accurate audio capture is crucial on the terminal side.

[0389] Processing on the server

[0390] The server converts the received audio data into text data using a speech recognition engine. General-purpose speech recognition software, such as a "speech recognition API," is used for this process. The text data is then analyzed using natural language processing techniques to confirm the user's transaction intent. Natural language processing libraries are used for this analysis.

[0391] Next, the emotion analysis engine analyzes changes in tone, intensity, and tempo of the voice to identify the user's emotions. For example, it uses an "emotion analysis tool" to infer emotions such as joy or anxiety. The authentication process provides highly accurate authentication using voice and emotion information.

[0392] Notification process

[0393] Based on the sentiment analysis results, the server creates a message in an appropriate tone and sends it back to the device. The device then informs the user of the message aloud. For example, if the user is nervous, it will play a calm message such as, "Your purchase has been successfully completed."

[0394] This system allows users to enjoy a consistent and personalized experience that takes their emotions into account. As a concrete example, consider the case of online shopping. An example of a prompt message might be: "Analyze the emotions expressed by the user during product purchase and advise on appropriate responses. Specifically, identify kindness, anxiety, and joy, and suggest corresponding voice messages."

[0395] In this way, the present invention makes it possible to provide a more enriching user experience by considering and analyzing emotions during the process of transactions conducted via voice.

[0396] The flow of the specific processing in Example 2 will be explained using Figure 13.

[0397] Step 1:

[0398] The user expresses their intention to purchase a product by voice. For example, they might say, "I want to buy this product." This voice information is acquired by the terminal and sent to the server as digital voice data. The terminal captures the voice input using a microphone and sends it to the server in the form of an audio file. During this process, efforts are made to minimize noise and maintain the clarity and accuracy of the voice.

[0399] Step 2:

[0400] The server converts the audio data received from the terminal into text data using a speech recognition engine. Specifically, it analyzes the audio waveform using speech recognition software and converts it into a corresponding string of characters. Through this process, the text "I want to buy this product" is obtained from the audio.

[0401] Step 3:

[0402] The server analyzes the converted text data using natural language processing techniques to extract transaction intent. For example, it uses a natural language processing library to perform grammatical analysis and keyword extraction to identify the intent to "purchase a product." Text data is used as input, and intent and related information are included in the output.

[0403] Step 4:

[0404] The server uses an emotion analysis engine to identify emotions from audio data. Using an audio analysis tool, it evaluates the tone, intensity, and tempo of the voice to infer the emotional state (e.g., joy, anxiety). The emotion analysis engine takes audio data as input and outputs emotion tags such as "joy."

[0405] Step 5:

[0406] The server performs authentication based on identified transaction details and emotions. The verification process uses voice authentication technology to verify identity and enhance reliability. For example, it compares voiceprints and outputs a result indicating whether voice authentication was successful or not.

[0407] Step 6:

[0408] If the transaction is authenticated, the server will execute the payment through the payment processing service. Payment information is entered, and the transaction ID and payment status are received as output. Encryption technology may be used to ensure the security of the payment process.

[0409] Step 7:

[0410] The server creates a message based on the sentiment analysis results and sends it to the device. The tone of the message is adjusted according to the user's emotions, and it is presented in a way that is easy for the user to understand. The device notifies the user of the received message via voice, confirming the completion of the purchase. For example, if the user is nervous, a message such as, "Your purchase has been successfully completed. Please rest assured," is played in a gentle tone.

[0411] (Application Example 2)

[0412] Next, we will explain application example 2. In the following explanation, the data processing device 12 will be referred to as the "server" and the smart glasses 214 as the "terminal".

[0413] In voice-based payment systems, there is a need to provide sophisticated communication and reassuring interactions that take into account not only the transaction details intended by the user but also their emotional state at the time. However, conventional voice payment systems have the challenge of not being able to adequately reflect the user's emotional state and not being able to provide sufficient reassurance to users who feel anxious.

[0414] The specific processing performed by the specific processing unit 290 of the data processing device 12 in Application Example 2 is realized by the following means.

[0415] In this invention, the server includes data acquisition means for receiving voice input, voice recognition means for converting voice data into text data, and emotion analysis means for analyzing the user's emotional state. This makes it possible to analyze the transaction content and emotions from the user's voice and provide voice feedback that takes into account a sense of security.

[0416] "Data acquisition means" refers to a device or function that receives voice input and collects it as voice data.

[0417] "Speech recognition means" refers to a technology or system that analyzes acquired speech data and converts it into corresponding text data.

[0418] "Natural language processing means" refers to software or algorithms that analyze text data generated by speech recognition means to accurately extract transaction details.

[0419] "Emotion analysis tools" are technologies that analyze a user's emotional state from their speech and identify emotions based on changes in voice tone and tempo.

[0420] "Authentication means" refers to a process or function that verifies the validity of a transaction based on the transaction details and emotional state, and verifies the user's identity as necessary.

[0421] A "data processing system" is a system for executing authenticated transactions and ensuring settlement.

[0422] An "information transmission means" is a mechanism that transmits the results of a transaction to the user and notifies them through voice or other interfaces.

[0423] This invention is a system that, when conducting transactions based on user voice input, analyzes the user's emotions using voice data and provides personalized feedback based on that analysis. The server is responsible for acquiring, processing, and analyzing the voice data, confirming the transaction details, and then authenticating and executing the payment.

[0424] The terminal captures high-quality voice input from the user and sends it to the server via a data acquisition method. The server uses the Google Speech-to-Text API to convert the speech to text and extracts transaction details through natural language processing. Furthermore, it uses the Google Cloud Natural Language API to perform sentiment analysis based on factors such as tone and intensity of the speech.

[0425] The analysis results are used for identity verification through authentication methods, and if the transaction is successfully authenticated, the payment process is executed via the Stripe API. Based on the sentiment results obtained during this process, feedback is provided to the user in an appropriate tone as a means of communication.

[0426] For example, if a user says in voice, "I want to buy this product, but I'm a little worried," the emotion analysis system can recognize the user's anxiety and respond with a calm voice message such as, "Please relax, the purchase process is going smoothly."

[0427] Example prompt: "The user is attempting to purchase a specific product with an anxious voice. The sentiment analysis engine should recognize this anxiety and generate reassuring voice feedback."

[0428] The flow of a specific process in Application Example 2 will be explained using Figure 14.

[0429] Step 1:

[0430] The user performs voice input. The user uses their smartphone to express their intention to purchase a product by voice. The input is the user's speech and is recorded directly on the device.

[0431] Step 2:

[0432] The terminal acquires audio data and sends it to the server. The terminal uses its built-in microphone to save the audio input as a digital audio file and sends it to the server via a data acquisition device. The input is audio data, and the output is a digital audio file sent to the server.

[0433] Step 3:

[0434] The server converts audio data into text using speech recognition technology. Using the Google Speech-to-Text API, the server generates corresponding text data from the audio data. The input is audio data, and the output is text data.

[0435] Step 4:

[0436] The server analyzes text data using natural language processing techniques and extracts transaction details. The server applies natural language processing algorithms to extract transaction intent and product characteristics from the text. The input is text data, and the output is structured data containing transaction details.

[0437] Step 5:

[0438] The server identifies the user's emotional state using sentiment analysis tools. Leveraging the Google Cloud Natural Language API, the server analyzes changes in tone and tempo of speech to determine the emotional state. The input is audio data, and the output is information about the user's emotional state.

[0439] Step 6:

[0440] The server uses authentication methods to authenticate the transaction. It takes into account the transaction details and sentiment status, and verifies the user's identity through the authentication process. The input is the transaction details and sentiment information, and the output is the authentication result.

[0441] Step 7:

[0442] If the server authenticates the transaction, it initiates the settlement process using data processing tools. The settlement is securely executed using the Stripe API. The input is the authentication result, and the output is a settlement completion notification.

[0443] Step 8:

[0444] The server notifies the user of the payment result via a communication method. The server generates voice feedback reflecting the sentiment analysis results and delivers it to the user through the terminal. The input is the payment completion notification and sentiment information, and the output is a personalized voice message.

[0445] The specific processing unit 290 transmits the result of the specific processing to the smart glasses 214. In the smart glasses 214, the control unit 46A causes the speaker 240 to output the result of the specific processing. The microphone 238 acquires audio indicating user input for the result of the specific processing. The control unit 46A transmits the audio data indicating user input acquired by the microphone 238 to the data processing unit 12. In the data processing unit 12, the specific processing unit 290 acquires the audio data.

[0446] Data generation model 58 is a type of so-called generative AI (Artificial Intelligence). One example of data generation model 58 is ChatGPT (Internet search<URL: https: / / openai.com / blog / chatgpt> ), Gemini (Internet search) <url: https: gemini.google.com ?hl="ja">Examples of generative AI include the following. The data generation model 58 is obtained by performing deep learning on a neural network. The data generation model 58 is input with prompts containing instructions, and with inference data such as audio data representing speech, text data representing text, and image data representing images. The data generation model 58 infers from the input inference data according to the instructions indicated by the prompts, and outputs the inference results in data formats such as audio data and text data. Here, inference refers to, for example, analysis, classification, prediction, and / or summarization.

[0447] In the above embodiment, an example was given in which specific processing is performed by the data processing device 12, but the technology of this disclosure is not limited thereto, and the specific processing may also be performed by the smart glasses 214.

[0448] [Third Embodiment]

[0449] Figure 5 shows an example of the configuration of the data processing system 310 according to the third embodiment.

[0450] As shown in Figure 5, the data processing system 310 includes a data processing device 12 and a headset terminal 314. An example of the data processing device 12 is a server.

[0451] The data processing device 12 comprises a computer 22, a database 24, and a communication interface 26. The computer 22 is an example of a "computer" related to the technology of this disclosure. The computer 22 comprises a processor 28, RAM 30, and storage 32. The processor 28, RAM 30, and storage 32 are connected to a bus 34. The database 24 and the communication interface 26 are also connected to the bus 34. The communication interface 26 is connected to a network 54. An example of the network 54 is a WAN (Wide Area Network) and / or a LAN (Local Area Network).

[0452] The headset terminal 314 includes a computer 36, a microphone 238, a speaker 240, a camera 42, a communication interface 44, and a display 343. The computer 36 includes a processor 46, RAM 48, and storage 50. The processor 46, RAM 48, and storage 50 are connected to a bus 52. The microphone 238, speaker 240, camera 42, and display 343 are also connected to the bus 52.

[0453] The microphone 238 receives voice signals from the user 20 and receives instructions from the user 20. The microphone 238 captures the voice signals from the user 20, converts the captured voice into audio data, and outputs it to the processor 46. The speaker 240 outputs audio according to the instructions from the processor 46.

[0454] Camera 42 is a small digital camera equipped with an optical system including a lens, aperture, and shutter, and an image sensor such as a CMOS (Complementary Metal-Oxide-Semiconductor) image sensor or a CCD (Charge Coupled Device) image sensor, and captures images of the area around the user 20 (for example, an imaging range defined by a field of view equivalent to the width of a typical healthy person's field of vision).

[0455] Communication interface 44 is connected to network 54. Communication interfaces 44 and 26 are responsible for the exchange of various information between processor 46 and processor 28 via network 54. The exchange of various information between processor 46 and processor 28 using communication interfaces 44 and 26 is performed in a secure manner.

[0456] Figure 6 shows an example of the main functions of the data processing device 12 and the headset terminal 314. As shown in Figure 6, the data processing device 12 performs specific processing using the processor 28. The storage 32 stores the specific processing program 56.

[0457] The specific processing program 56 is an example of a "program" relating to the technology of this disclosure. The processor 28 reads the specific processing program 56 from the storage 32 and executes the read specific processing program 56 on the RAM 30. The specific processing is realized by the processor 28 operating as a specific processing unit 290 in accordance with the specific processing program 56 executed on the RAM 30.

[0458] The storage 32 stores the data generation model 58 and the emotion identification model 59. The data generation model 58 and the emotion identification model 59 are used by the identification processing unit 290.

[0459] In the headset terminal 314, the processor 46 performs the reception output processing. The storage 50 stores the reception output program 60. The processor 46 reads the reception output program 60 from the storage 50 and executes the read reception output program 60 on the RAM 48. The reception output processing is realized by the processor 46 operating as a control unit 46A according to the reception output program 60 executed on the RAM 48.

[0460] Next, the specific processing performed by the specific processing unit 290 of the data processing device 12 will be described. In the following description, the data processing device 12 will be referred to as the "server" and the headset terminal 314 will be referred to as the "terminal".

[0461] This invention is a system for making payments by voice, and its configuration includes a user, a terminal, and a server.

[0462] User-side operations

[0463] After finding the product they want to purchase, the user communicates their intention to the device by voice. For example, they can say a voice command such as, "I want to buy this product."

[0464] Terminal operation

[0465] The terminal receives voice from the user using a highly sensitive microphone and immediately sends the acquired voice data to the server. The terminal compresses the voice into an appropriate format and transfers the data using secure communication. It then receives a reply from the server and provides feedback to the user.

[0466] Server Processing

[0467] The server first uses a speech recognition engine to convert the received audio data into text. This engine, based on a machine learning model, accurately captures what the user is saying as text. Next, the converted text data is analyzed using natural language processing (NLP) technology. NLP accurately identifies what the user is looking for (product name, purchase intention, etc.).

[0468] Subsequently, the server proceeds with the authentication process, verifying the user's identity based on the acquired information. This authentication can utilize voice authentication or other authentication methods. If authentication is successful, the server immediately executes the payment process. The payment process uses the user's registered payment information to ensure a secure and fast transaction.

[0469] Finally, the server sends a notification to the device to inform the user that the payment has been completed. The device then communicates this notification to the user via voice or on-screen display. A message such as "Purchase complete" is sent to the user.

[0470] Thus, this technology provides high convenience by quickly understanding the user's intent through voice recognition and completing transactions safely and efficiently. For example, when purchasing a product at a coffee shop, the user can say "I want to buy a coffee" by voice, and the system can automatically carry out the coffee purchase procedure according to this flow.

[0471] The following describes the processing flow.

[0472] Step 1:

[0473] The user communicates their intention to purchase the item by voice to the device. For example, they might say, "I will purchase this item."

[0474] Step 2:

[0475] The device acquires the user's voice in real time through the microphone and captures it as audio data.

[0476] Step 3:

[0477] The device transmits the acquired audio data to the server using a secure protocol. During this process, the data is encrypted and compressed.

[0478] Step 4:

[0479] The server passes the received audio data to the speech recognition engine, which converts it into high-precision text data.

[0480] Step 5:

[0481] The server analyzes the text data using natural language processing technology to extract the user's intent (product name, transaction details, etc.).

[0482] Step 6:

[0483] The server authenticates the user based on the transaction details. It uses voice authentication or other authentication methods to verify the user's identity.

[0484] Step 7:

[0485] Upon successful authentication, the server processes the payment. It uses the registered payment information to execute a secure transaction.

[0486] Step 8:

[0487] The server sends the payment completion status to the terminal, including information on whether the payment was successful.

[0488] Step 9:

[0489] The device receives a notification from the server and informs the user of the completion of the payment via voice or screen display. For example, it might notify the user that "Your purchase is complete."

[0490] (Example 1)

[0491] Next, we will describe Example 1. In the following description, the data processing device 12 will be referred to as the "server," and the headset-type terminal 314 will be referred to as the "terminal."

[0492] Traditional payment systems have limitations in terms of user input methods, resulting in low convenience and difficulty in processing transactions quickly. Furthermore, they suffer from insufficient security and user authentication accuracy, leading to a lack of transaction reliability.

[0493] The identification process performed by the identification processing unit 290 of the data processing device 12 in Example 1 is realized by the following means.

[0494] In this invention, the server includes an information acquisition means for receiving voice input, a voice recognition means for converting the voice information acquired by the information acquisition means into text information, and a language processing means for analyzing the text information generated by the voice recognition means and extracting its content. This enables the rapid and highly accurate processing of the user's voice, making secure and reliable payment possible.

[0495] "Information acquisition means" refers to a device or method for receiving user input, and in this case, it refers to technology for receiving voice input.

[0496] "Speech recognition means" refers to technology that analyzes speech information and converts it into text information, and has the function of accurately converting speech to text using machine learning models, etc.

[0497] "Language processing means" refers to techniques for further analyzing generated textual information and extracting necessary transaction details from it, and this involves using natural language processing.

[0498] "Authentication methods" refer to technologies used to verify the identity of users in order to ensure the security of transactions, and include voice authentication and biometric authentication.

[0499] A "payment method" is a system that, after a transaction has been authenticated, executes the payment process using the user's registered payment information.

[0500] "Notification means" refers to technologies for communicating transaction results to users, and includes methods for providing information through voice or screen display.

[0501] "Communication methods" refer to technologies for securely transmitting data by compressing it into an appropriate format, and utilize secure communication protocols.

[0502] This invention is a voice-based payment system in which a user, a terminal, and a server work together. The user gives voice instructions for purchasing goods, and the terminal is equipped with information acquisition means to receive this voice. The terminal collects voice using a high-sensitivity microphone and has communication means to transmit the data to the server, using a secure communication protocol.

[0503] The server converts speech information into text information as a speech recognition tool, for example, by using a machine learning-based speech recognition engine. Furthermore, the server uses language processing tools to analyze this text information using natural language processing technology (e.g., spaCy) and extract transaction details. Subsequently, identity verification is performed using authentication tools. User authentication can be performed quickly and securely using voice authentication or biometric authentication technology. After successful authentication, the server executes a secure transaction using the user's registered payment information with payment tools. It is desirable to use a payment service (e.g., Stripe) at this stage.

[0504] Once a transaction is complete, the server notifies the terminal of the result using a notification system, and the terminal communicates this information to the user via voice or screen display. This notification allows the user to confirm that the transaction was completed successfully.

[0505] As a concrete example, if a user says "I want to buy a coffee" by voice at a coffee shop, this system can automatically complete the coffee purchase process. The system analyzes and authenticates the voice, completes the payment, and notifies the user of the result by voice or display.

[0506] An example of a prompt in a generative AI model is a request for a detailed explanation of the system, such as, "Please explain in detail the processing flow of the system for purchasing products by voice."

[0507] The flow of the specific processing in Example 1 will be explained using Figure 11.

[0508] Step 1:

[0509] The user communicates their intention to purchase a product via voice. Specifically, the user issues a voice command such as "I want to buy this product" into the device. The input is the user's voice, which the device receives. The output is voice data.

[0510] Step 2:

[0511] The terminal collects the received audio using an information acquisition device and employs a high-sensitivity microphone. Subsequently, this audio data is converted into a compressed format and transmitted to the server using a secure communication method. The input is audio data, and the output is compressed audio data.

[0512] Step 3:

[0513] The server converts the received audio data into text information using a machine learning-based speech recognition engine as a speech recognition method. In this process, the speech recognition model analyzes the audio and obtains output as text data. The input is compressed audio data, and the output is text data.

[0514] Step 4:

[0515] The server uses language processing tools to analyze the generated text data through natural language processing techniques to identify transaction details. Specifically, it performs data calculations to identify product names and transaction intent. The input is text data, and the output is the analyzed transaction information.

[0516] Step 5:

[0517] The server uses authentication methods to verify the user's identity. This process utilizes voice authentication and biometric authentication technologies to enhance the accuracy of user identification and authentication. The input is the analyzed transaction information, and the output is the authentication result.

[0518] Step 6:

[0519] If the transaction is authenticated, the server executes the purchase process using the payment method. Specifically, it completes the economic transaction through a secure payment service. The input is the authentication result, and the output is information about the completion of the payment.

[0520] Step 7:

[0521] The server notifies the terminal that the payment has been completed via a notification system. The terminal then informs the user of this information via voice or display. The input is the payment completion information, and the output is the transaction completion notification to the user.

[0522] (Application Example 1)

[0523] Next, we will explain Application Example 1. In the following explanation, the data processing device 12 will be referred to as the "server," and the headset-type terminal 314 will be referred to as the "terminal."

[0524] In modern brick-and-mortar stores, there is a demand for customers to complete purchases quickly and efficiently. However, current methods result in waiting times at the checkout, reducing the efficiency of the shopping experience. In particular, there is a need for a new interface that allows payment to be completed by voice without going through a physical checkout system.

[0525] The specific processing performed by the specific processing unit 290 of the data processing device 12 in Application Example 1 is realized by the following means.

[0526] In this invention, the server includes sound acquisition means, sound recognition means, natural language processing means, settlement means, and notification means. This allows users to quickly complete purchase procedures using voice input in physical stores, and after confirming the notified transaction result, they can leave the store with their purchased items without going through a cash register.

[0527] "Sound acquisition means" refers to a device or system that receives audio emitted by a user and converts it into data for subsequent processing.

[0528] "Acoustic recognition means" refers to a processing device that has the function of analyzing acquired audio data and converting it into a corresponding text format.

[0529] "Natural language processing means" refers to a technology that analyzes text data acquired by acoustic recognition means to extract user intent and transaction details.

[0530] "Verification methods" refer to the process of verifying the legitimacy of a transaction and confirming the user's identity based on the extracted transaction details.

[0531] A "settlement instrument" is a device or system that, based on the results of a verification instrument, executes settlement processing to complete a transaction.

[0532] A "notification system" is a system that has audio and visual notification functions to communicate the results of a transaction to the user.

[0533] This invention is a system that enables instant voice payment in physical stores. Users use a terminal such as a smartphone. The terminal uses a high-sensitivity microphone to acquire voice data. This voice data is converted into text data by an acoustic recognition means, and the transaction details are further analyzed by a natural language processing means.

[0534] To facilitate these processes, the server utilizes a speech recognition engine (e.g., Google Speech-to-Text API) to generate text data. The analyzed content is then used to verify the user's identity and determine the validity of the transaction using voice authentication or other methods. Once authentication is complete, a secure payment process is executed using a settlement method. This uses a payment gateway API (e.g., Stripe).

[0535] Once payment is complete, a notification system will alert the user to the transaction result. This notification will be made via audio or on-screen display. It will inform the user that the purchase of the product has been completed in the physical store and confirm that the transaction has been successfully completed.

[0536] As a concrete example, when a user at a cafe says "I'll buy this coffee" using voice input on their smartphone, the device connects with a server and completes the payment process. The server authenticates the transaction using a verification mechanism and notifies the user via a notification mechanism that "the purchase is complete."

[0537] An example of a prompt message would be: "Design a program that uses a speech recognition system to convert what the user says into text in real time and then processes the payment based on that information. Specifically, please show the flow when the user says, 'I will buy this product.'"

[0538] The flow of a specific process in Application Example 1 will be explained using Figure 12.

[0539] Step 1:

[0540] The user makes a voice input near the smartphone. The device receives this voice using a high-sensitivity microphone and acquires it as digital audio data. The input is the user's voice, and the output is digital audio data.

[0541] Step 2:

[0542] The terminal immediately transmits the acquired digital audio data to the server. During this process, the terminal compresses the audio data and transfers it using a secure communication method. The input is digital audio data, and the output is the transmitted audio data.

[0543] Step 3:

[0544] The server uses a speech recognition engine to convert received audio data into text data. This process utilizes a generative AI model to convert speech to text. The input is audio data, and the output is text data.

[0545] Step 4:

[0546] The server performs natural language processing on the text data to extract the transaction details intended by the user. Through this analysis, product names and purchase intentions are identified. The input is text data, and the output is the extracted transaction details.

[0547] Step 5:

[0548] The server uses verification methods to authenticate the user based on the extracted transaction details. This process utilizes voice authentication and other methods to determine the user's legitimacy. The input is the transaction details, and the output is the authentication result.

[0549] Step 6:

[0550] If the verification process is successful, the server will use the settlement method to process the payment. In this process, the user's payment information is used to execute a secure transaction via the payment gateway API. The input is the authenticated transaction details, and the output is the settlement result.

[0551] Step 7:

[0552] The server transmits the payment result to the terminal via a notification system. The terminal notifies the user of a message such as "Purchase complete" via voice or screen display. The input is the payment result, and the output is the notification to the user.

[0553] Furthermore, an emotion engine that estimates the user's emotions may be incorporated. That is, the identification processing unit 290 may use the emotion identification model 59 to estimate the user's emotions and perform identification processing using the user's emotions.

[0554] This invention incorporates an emotion engine that recognizes user emotions into a voice-based payment system. Its aim is to enhance the conventional process of understanding transaction intent from voice input and to provide a more human-like interaction and experience by taking the user's emotional state into consideration during the payment process.

[0555] User operation

[0556] The user expresses their intention to purchase the product by voice, as is typical. For example, they might say, "I want to buy this product." The user's emotions, such as joy or anxiety, may also be included in their voice.

[0557] Terminal operation

[0558] The terminal captures the user's voice directly and transfers it to the server. The terminal's processing involves the clear and accurate collection and transmission of the voice data.

[0559] Processing on the server

[0560] The server uses speech recognition to convert the audio data received from the terminal into text data. This text data is then analyzed using natural language processing (NLP) techniques to confirm the user's purchase intent. Next, the emotion engine identifies emotions from the user's audio data. For example, it infers the user's emotions from changes in tone, volume, and tempo of their voice.

[0561] Once emotions are identified, that data is associated with the transaction intent and passed to the authentication mechanism. The authentication process, in addition to standard identity verification, considers the user's emotional state to perform a more precise intent check. If authentication is successful, the server proceeds with the payment process and securely completes the transaction.

[0562] Notification process

[0563] Once the payment is complete, the server uses the emotion engine's analysis results to notify the device of the outcome using an appropriate tone. The device then verbally communicates this information to the user. For example, if the user is nervous, the server can deliver a message in a calm tone such as, "Your purchase has been successfully completed."

[0564] This invention allows users to enjoy more personalized interactions that take their emotions into account, making the service experience more comfortable and reassuring. A concrete example is a scenario where the system provides reassuring voice support to a user who is feeling anxious while shopping.

[0565] The following describes the processing flow.

[0566] Step 1:

[0567] The user expresses the product they want to purchase by voice, using phrases such as "I want to buy this product."

[0568] Step 2:

[0569] The device acquires the user's voice through the microphone and captures it as audio data.

[0570] Step 3:

[0571] The device sends the acquired audio data to the server. The data is encrypted during this process.

[0572] Step 4:

[0573] The server analyzes the received audio data using a speech recognition engine and converts the audio into text.

[0574] Step 5:

[0575] The server analyzes the text data using natural language processing technology to identify the transaction details.

[0576] Step 6:

[0577] The server further analyzes the voice data using an emotion engine to recognize the user's emotions. This process includes analyzing the tone and volume of the voice.

[0578] Step 7:

[0579] The server performs an authentication process and approves the transaction, taking into account the user's identity and perceived emotions.

[0580] Step 8:

[0581] Upon successful authentication, the server processes the payment and securely completes the transaction.

[0582] Step 9:

[0583] The server generates an appropriate notification tone based on the payment result and the user's sentiment, and sends it to the device.

[0584] Step 10:

[0585] The terminal notifies the user of the payment result via voice. For example, it might say in a calm voice, "Your purchase is complete."

[0586] (Example 2)

[0587] Next, we will describe Example 2. In the following description, the data processing device 12 will be referred to as the "server," and the headset-type terminal 314 will be referred to as the "terminal."

[0588] Conventional voice recognition systems can understand a user's transaction intent, but they do not consider the user's emotional state, which can result in a lack of human-like interaction and experience. Furthermore, in authentication processes, ignoring the user's emotional state can lead to insufficient identity verification. Therefore, there is a need for a voice payment system that takes user emotions into account and provides a more natural and secure transaction experience.

[0589] The identification process performed by the identification processing unit 290 of the data processing device 12 in Example 2 is realized by the following means.

[0590] In this invention, the server includes an input means for receiving voice, a conversion means for converting the voice acquired by the input means into text, a processing means for analyzing the text data generated by the conversion means to extract transaction details, an emotion analysis means for identifying emotions from the voice data acquired by the processing means, a verification means for performing authentication based on the extracted transaction details and identified emotions, an operation means for executing processing when the transaction is authenticated, and a communication means for notifying the processing results. This enables a more natural and secure voice-based transaction experience while taking into account the user's emotional state.

[0591] "Input means" refers to a function that acquires audio and provides it in a format that the system can process.

[0592] "Conversion means" refers to the process of converting acquired audio into text format and generating text data.

[0593] The "processing means" refers to a function that analyzes the generated text data to extract the user's transaction intent.

[0594] "Emotion analysis means" refers to a process for identifying a user's emotions from voice data and inferring their emotional state.

[0595] "Verification means" refers to the process of authentication based on extracted transaction details and identified sentiments.

[0596] "Operational means" refers to functions that execute settlement and other related processes once a transaction is authenticated.

[0597] "Means of communication" refers to the process of notifying the user of the processing results, and the results can be conveyed through means such as voice.

[0598] This invention incorporates an emotion analysis function that recognizes user emotions into a voice-based payment system. The aim is to understand the user's transaction intent and provide a natural dialogue and experience that takes their emotional state into account. Specific embodiments are described below.

[0599] User operation

[0600] When purchasing a product, users express their intentions verbally. For example, they might say, "I want to buy this product." The tone and tempo of their voice may reflect their emotions, such as joy or anxiety.

[0601] Terminal operation

[0602] The terminal uses a microphone or similar device to acquire the user's voice information and transmits it to the server as digital audio data. Clear and accurate audio capture is crucial on the terminal side.

[0603] Processing on the server

[0604] The server converts the received audio data into text data using a speech recognition engine. General-purpose speech recognition software, such as a "speech recognition API," is used for this process. The text data is then analyzed using natural language processing techniques to confirm the user's transaction intent. Natural language processing libraries are used for this analysis.

[0605] Next, the emotion analysis engine analyzes changes in tone, intensity, and tempo of the voice to identify the user's emotions. For example, it uses an "emotion analysis tool" to infer emotions such as joy or anxiety. The authentication process provides highly accurate authentication using voice and emotion information.

[0606] Notification process

[0607] Based on the sentiment analysis results, the server creates a message in an appropriate tone and sends it back to the device. The device then informs the user of the message aloud. For example, if the user is nervous, it will play a calm message such as, "Your purchase has been successfully completed."

[0608] This system allows users to enjoy a consistent and personalized experience that takes their emotions into account. As a concrete example, consider the case of online shopping. An example of a prompt message might be: "Analyze the emotions expressed by the user during product purchase and advise on appropriate responses. Specifically, identify kindness, anxiety, and joy, and suggest corresponding voice messages."

[0609] In this way, the present invention makes it possible to provide a more enriching user experience by considering and analyzing emotions during the process of transactions conducted via voice.

[0610] The flow of the specific processing in Example 2 will be explained using Figure 13.

[0611] Step 1:

[0612] The user expresses their intention to purchase a product by voice. For example, they might say, "I want to buy this product." This voice information is acquired by the terminal and sent to the server as digital voice data. The terminal captures the voice input using a microphone and sends it to the server in the form of an audio file. During this process, efforts are made to minimize noise and maintain the clarity and accuracy of the voice.

[0613] Step 2:

[0614] The server converts the audio data received from the terminal into text data using a speech recognition engine. Specifically, it analyzes the audio waveform using speech recognition software and converts it into a corresponding string of characters. Through this process, the text "I want to buy this product" is obtained from the audio.

[0615] Step 3:

[0616] The server analyzes the converted text data using natural language processing techniques to extract transaction intent. For example, it uses a natural language processing library to perform grammatical analysis and keyword extraction to identify the intent to "purchase a product." Text data is used as input, and intent and related information are included in the output.

[0617] Step 4:

[0618] The server uses an emotion analysis engine to identify emotions from audio data. Using an audio analysis tool, it evaluates the tone, intensity, and tempo of the voice to infer the emotional state (e.g., joy, anxiety). The emotion analysis engine takes audio data as input and outputs emotion tags such as "joy."

[0619] Step 5:

[0620] The server performs authentication based on identified transaction details and emotions. The verification process uses voice authentication technology to verify identity and enhance reliability. For example, it compares voiceprints and outputs a result indicating whether voice authentication was successful or not.

[0621] Step 6:

[0622] If the transaction is authenticated, the server will execute the payment through the payment processing service. Payment information is entered, and the transaction ID and payment status are received as output. Encryption technology may be used to ensure the security of the payment process.

[0623] Step 7:

[0624] The server creates a message based on the sentiment analysis results and sends it to the device. The tone of the message is adjusted according to the user's emotions, and it is presented in a way that is easy for the user to understand. The device notifies the user of the received message via voice, confirming the completion of the purchase. For example, if the user is nervous, a message such as, "Your purchase has been successfully completed. Please rest assured," is played in a gentle tone.

[0625] (Application Example 2)

[0626] Next, we will explain application example 2. In the following explanation, the data processing device 12 will be referred to as the "server," and the headset-type terminal 314 will be referred to as the "terminal."

[0627] In voice-based payment systems, there is a need to provide sophisticated communication and reassuring interactions that take into account not only the transaction details intended by the user but also their emotional state at the time. However, conventional voice payment systems have the challenge of not being able to adequately reflect the user's emotional state and not being able to provide sufficient reassurance to users who feel anxious.

[0628] The specific processing performed by the specific processing unit 290 of the data processing device 12 in Application Example 2 is realized by the following means.

[0629] In this invention, the server includes data acquisition means for receiving voice input, voice recognition means for converting voice data into text data, and emotion analysis means for analyzing the user's emotional state. This makes it possible to analyze the transaction content and emotions from the user's voice and provide voice feedback that takes into account a sense of security.

[0630] "Data acquisition means" refers to a device or function that receives voice input and collects it as voice data.

[0631] "Speech recognition means" refers to a technology or system that analyzes acquired speech data and converts it into corresponding text data.

[0632] "Natural language processing means" refers to software or algorithms that analyze text data generated by speech recognition means to accurately extract transaction details.

[0633] "Emotion analysis tools" are technologies that analyze a user's emotional state from their speech and identify emotions based on changes in voice tone and tempo.

[0634] "Authentication means" refers to a process or function that verifies the validity of a transaction based on the transaction details and emotional state, and verifies the user's identity as necessary.

[0635] A "data processing system" is a system for executing authenticated transactions and ensuring settlement.

[0636] An "information transmission means" is a mechanism that transmits the results of a transaction to the user and notifies them through voice or other interfaces.

[0637] This invention is a system that, when conducting transactions based on user voice input, analyzes the user's emotions using voice data and provides personalized feedback based on that analysis. The server is responsible for acquiring, processing, and analyzing the voice data, confirming the transaction details, and then authenticating and executing the payment.

[0638] The terminal captures high-quality voice input from the user and sends it to the server via a data acquisition method. The server uses the Google Speech-to-Text API to convert the speech to text and extracts transaction details through natural language processing. Furthermore, it uses the Google Cloud Natural Language API to perform sentiment analysis based on factors such as tone and intensity of the speech.

[0639] The analysis results are used for identity verification through authentication methods, and if the transaction is successfully authenticated, the payment process is executed via the Stripe API. Based on the sentiment results obtained during this process, feedback is provided to the user in an appropriate tone as a means of communication.

[0640] For example, if a user says in voice, "I want to buy this product, but I'm a little worried," the emotion analysis system can recognize the user's anxiety and respond with a calm voice message such as, "Please relax, the purchase process is going smoothly."

[0641] Example prompt: "The user is attempting to purchase a specific product with an anxious voice. The sentiment analysis engine should recognize this anxiety and generate reassuring voice feedback."

[0642] The flow of a specific process in Application Example 2 will be explained using Figure 14.

[0643] Step 1:

[0644] The user performs voice input. The user uses their smartphone to express their intention to purchase a product by voice. The input is the user's speech and is recorded directly on the device.

[0645] Step 2:

[0646] The terminal acquires audio data and sends it to the server. The terminal uses its built-in microphone to save the audio input as a digital audio file and sends it to the server via a data acquisition device. The input is audio data, and the output is a digital audio file sent to the server.

[0647] Step 3:

[0648] The server converts audio data into text using speech recognition technology. Using the Google Speech-to-Text API, the server generates corresponding text data from the audio data. The input is audio data, and the output is text data.

[0649] Step 4:

[0650] The server analyzes text data using natural language processing techniques and extracts transaction details. The server applies natural language processing algorithms to extract transaction intent and product characteristics from the text. The input is text data, and the output is structured data containing transaction details.

[0651] Step 5:

[0652] The server identifies the user's emotional state using sentiment analysis tools. Leveraging the Google Cloud Natural Language API, the server analyzes changes in tone and tempo of speech to determine the emotional state. The input is audio data, and the output is information about the user's emotional state.

[0653] Step 6:

[0654] The server uses authentication methods to authenticate the transaction. It takes into account the transaction details and sentiment status, and verifies the user's identity through the authentication process. The input is the transaction details and sentiment information, and the output is the authentication result.

[0655] Step 7:

[0656] If the server authenticates the transaction, it initiates the settlement process using data processing tools. The settlement is securely executed using the Stripe API. The input is the authentication result, and the output is a settlement completion notification.

[0657] Step 8:

[0658] The server notifies the user of the payment result via a communication method. The server generates voice feedback reflecting the sentiment analysis results and delivers it to the user through the terminal. The input is the payment completion notification and sentiment information, and the output is a personalized voice message.

[0659] The specific processing unit 290 transmits the result of the specific processing to the headset terminal 314. In the headset terminal 314, the control unit 46A causes the speaker 240 and display 343 to output the result of the specific processing. The microphone 238 acquires audio indicating user input for the result of the specific processing. The control unit 46A transmits the audio data indicating user input acquired by the microphone 238 to the data processing unit 12. In the data processing unit 12, the specific processing unit 290 acquires the audio data.

[0660] Data generation model 58 is a type of so-called generative AI (Artificial Intelligence). One example of data generation model 58 is ChatGPT (Internet search<URL: https: / / openai.com / blog / chatgpt> ), Gemini (Internet search) <url: https: gemini.google.com ?hl="ja">Examples of generative AI include the following. The data generation model 58 is obtained by performing deep learning on a neural network. The data generation model 58 is input with prompts containing instructions, and with inference data such as audio data representing speech, text data representing text, and image data representing images. The data generation model 58 infers from the input inference data according to the instructions indicated by the prompts, and outputs the inference results in data formats such as audio data and text data. Here, inference refers to, for example, analysis, classification, prediction, and / or summarization.

[0661] In the above embodiment, an example was given in which specific processing is performed by the data processing device 12, but the technology of this disclosure is not limited thereto, and specific processing may also be performed by the headset terminal 314.

[0662] [Fourth Embodiment]

[0663] Figure 7 shows an example of the configuration of the data processing system 410 according to the fourth embodiment.

[0664] As shown in Figure 7, the data processing system 410 includes a data processing device 12 and a robot 414. An example of the data processing device 12 is a server.

[0665] The data processing device 12 comprises a computer 22, a database 24, and a communication interface 26. The computer 22 is an example of a "computer" related to the technology of this disclosure. The computer 22 comprises a processor 28, RAM 30, and storage 32. The processor 28, RAM 30, and storage 32 are connected to a bus 34. The database 24 and the communication interface 26 are also connected to the bus 34. The communication interface 26 is connected to a network 54. An example of the network 54 is a WAN (Wide Area Network) and / or a LAN (Local Area Network).

[0666] The robot 414 includes a computer 36, a microphone 238, a speaker 240, a camera 42, a communication interface 44, and a controlled object 443. The computer 36 includes a processor 46, RAM 48, and storage 50. The processor 46, RAM 48, and storage 50 are connected to a bus 52. The microphone 238, speaker 240, camera 42, and controlled object 443 are also connected to the bus 52.

[0667] The microphone 238 receives voice signals from the user 20 and receives instructions from the user 20. The microphone 238 captures the voice signals from the user 20, converts the captured voice into audio data, and outputs it to the processor 46. The speaker 240 outputs audio according to the instructions from the processor 46.

[0668] Camera 42 is a small digital camera equipped with an optical system including a lens, aperture, and shutter, and an image sensor such as a CMOS (Complementary Metal-Oxide-Semiconductor) image sensor or a CCD (Charge Coupled Device) image sensor, and captures images of the area around the user 20 (for example, an imaging range defined by a field of view equivalent to the width of a typical healthy person's field of vision).

[0669] Communication interface 44 is connected to network 54. Communication interfaces 44 and 26 are responsible for the exchange of various information between processor 46 and processor 28 via network 54. The exchange of various information between processor 46 and processor 28 using communication interfaces 44 and 26 is performed in a secure manner.

[0670] The controlled object 443 includes a display device, LEDs in the eyes, and motors that drive the arms, hands, and feet. The posture and gestures of the robot 414 are controlled by controlling the motors of the arms, hands, and feet. Some of the robot 414's emotions can be expressed by controlling these motors. Furthermore, the robot 414's facial expressions can also be expressed by controlling the illumination state of the LEDs in its eyes.

[0671] Figure 8 shows an example of the main functions of the data processing device 12 and the robot 414. As shown in Figure 8, the data processing device 12 performs specific processing using the processor 28. The storage 32 stores the specific processing program 56.

[0672] The specific processing program 56 is an example of a "program" relating to the technology of this disclosure. The processor 28 reads the specific processing program 56 from the storage 32 and executes the read specific processing program 56 on the RAM 30. The specific processing is realized by the processor 28 operating as a specific processing unit 290 in accordance with the specific processing program 56 executed on the RAM 30.

[0673] The storage 32 stores the data generation model 58 and the emotion identification model 59. The data generation model 58 and the emotion identification model 59 are used by the identification processing unit 290.

[0674] In robot 414, the processor 46 performs the reception output processing. The storage 50 stores the reception output program 60. The processor 46 reads the reception output program 60 from the storage 50 and executes the read reception output program 60 on the RAM 48. The reception output processing is realized by the processor 46 operating as a control unit 46A according to the reception output program 60 executed on the RAM 48.

[0675] Next, the specific processing performed by the specific processing unit 290 of the data processing device 12 will be described. In the following description, the data processing device 12 will be referred to as the "server" and the robot 414 as the "terminal".

[0676] This invention is a system for making payments by voice, and its configuration includes a user, a terminal, and a server.

[0677] User-side operations

[0678] After finding the product they want to purchase, the user communicates their intention to the device by voice. For example, they can say a voice command such as, "I want to buy this product."

[0679] Terminal operation

[0680] The terminal receives voice from the user using a highly sensitive microphone and immediately sends the acquired voice data to the server. The terminal compresses the voice into an appropriate format and transfers the data using secure communication. It then receives a reply from the server and provides feedback to the user.

[0681] Server Processing

[0682] The server first uses a speech recognition engine to convert the received audio data into text. This engine, based on a machine learning model, accurately captures what the user is saying as text. Next, the converted text data is analyzed using natural language processing (NLP) technology. NLP accurately identifies what the user is looking for (product name, purchase intention, etc.).

[0683] Subsequently, the server proceeds with the authentication process, verifying the user's identity based on the acquired information. This authentication can utilize voice authentication or other authentication methods. If authentication is successful, the server immediately executes the payment process. The payment process uses the user's registered payment information to ensure a secure and fast transaction.

[0684] Finally, the server sends a notification to the device to inform the user that the payment has been completed. The device then communicates this notification to the user via voice or on-screen display. A message such as "Purchase complete" is sent to the user.

[0685] Thus, this technology provides high convenience by quickly understanding the user's intent through voice recognition and completing transactions safely and efficiently. For example, when purchasing a product at a coffee shop, the user can say "I want to buy a coffee" by voice, and the system can automatically carry out the coffee purchase procedure according to this flow.

[0686] The following describes the processing flow.

[0687] Step 1:

[0688] The user communicates their intention to purchase the item by voice to the device. For example, they might say, "I will purchase this item."

[0689] Step 2:

[0690] The device acquires the user's voice in real time through the microphone and captures it as audio data.

[0691] Step 3:

[0692] The device transmits the acquired audio data to the server using a secure protocol. During this process, the data is encrypted and compressed.

[0693] Step 4:

[0694] The server passes the received audio data to the speech recognition engine, which converts it into high-precision text data.

[0695] Step 5:

[0696] The server analyzes the text data using natural language processing technology to extract the user's intent (product name, transaction details, etc.).

[0697] Step 6:

[0698] The server authenticates the user based on the transaction details. It uses voice authentication or other authentication methods to verify the user's identity.

[0699] Step 7:

[0700] Upon successful authentication, the server processes the payment. It uses the registered payment information to execute a secure transaction.

[0701] Step 8:

[0702] The server sends the payment completion status to the terminal, including information on whether the payment was successful.

[0703] Step 9:

[0704] The device receives a notification from the server and informs the user of the completion of the payment via voice or screen display. For example, it might notify the user that "Your purchase is complete."

[0705] (Example 1)

[0706] Next, we will describe Example 1. In the following description, the data processing device 12 will be referred to as the "server" and the robot 414 as the "terminal".

[0707] Traditional payment systems have limitations in terms of user input methods, resulting in low convenience and difficulty in processing transactions quickly. Furthermore, they suffer from insufficient security and user authentication accuracy, leading to a lack of transaction reliability.

[0708] The identification process performed by the identification processing unit 290 of the data processing device 12 in Example 1 is realized by the following means.

[0709] In this invention, the server includes an information acquisition means for receiving voice input, a voice recognition means for converting the voice information acquired by the information acquisition means into text information, and a language processing means for analyzing the text information generated by the voice recognition means and extracting its content. This enables the rapid and highly accurate processing of the user's voice, making secure and reliable payment possible.

[0710] "Information acquisition means" refers to a device or method for receiving user input, and in this case, it refers to technology for receiving voice input.

[0711] "Speech recognition means" refers to technology that analyzes speech information and converts it into text information, and has the function of accurately converting speech to text using machine learning models, etc.

[0712] "Language processing means" refers to techniques for further analyzing generated textual information and extracting necessary transaction details from it, and this involves using natural language processing.

[0713] "Authentication methods" refer to technologies used to verify the identity of users in order to ensure the security of transactions, and include voice authentication and biometric authentication.

[0714] A "payment method" is a system that, after a transaction has been authenticated, executes the payment process using the user's registered payment information.

[0715] "Notification means" refers to technologies for communicating transaction results to users, and includes methods for providing information through voice or screen display.

[0716] "Communication methods" refer to technologies for securely transmitting data by compressing it into an appropriate format, and utilize secure communication protocols.

[0717] This invention is a voice-based payment system in which a user, a terminal, and a server work together. The user gives voice instructions for purchasing goods, and the terminal is equipped with information acquisition means to receive this voice. The terminal collects voice using a high-sensitivity microphone and has communication means to transmit the data to the server, using a secure communication protocol.

[0718] The server converts speech information into text information as a speech recognition tool, for example, by using a machine learning-based speech recognition engine. Furthermore, the server uses language processing tools to analyze this text information using natural language processing technology (e.g., spaCy) and extract transaction details. Subsequently, identity verification is performed using authentication tools. User authentication can be performed quickly and securely using voice authentication or biometric authentication technology. After successful authentication, the server executes a secure transaction using the user's registered payment information with payment tools. It is desirable to use a payment service (e.g., Stripe) at this stage.

[0719] Once a transaction is complete, the server notifies the terminal of the result using a notification system, and the terminal communicates this information to the user via voice or screen display. This notification allows the user to confirm that the transaction was completed successfully.

[0720] As a concrete example, if a user says "I want to buy a coffee" by voice at a coffee shop, this system can automatically complete the coffee purchase process. The system analyzes and authenticates the voice, completes the payment, and notifies the user of the result by voice or display.

[0721] An example of a prompt in a generative AI model is a request for a detailed explanation of the system, such as, "Please explain in detail the processing flow of the system for purchasing products by voice."

[0722] The flow of the specific processing in Example 1 will be explained using Figure 11.

[0723] Step 1:

[0724] The user communicates their intention to purchase a product via voice. Specifically, the user issues a voice command such as "I want to buy this product" into the device. The input is the user's voice, which the device receives. The output is voice data.

[0725] Step 2:

[0726] The terminal collects the received audio using an information acquisition device and employs a high-sensitivity microphone. Subsequently, this audio data is converted into a compressed format and transmitted to the server using a secure communication method. The input is audio data, and the output is compressed audio data.

[0727] Step 3:

[0728] The server converts the received audio data into text information using a machine learning-based speech recognition engine as a speech recognition method. In this process, the speech recognition model analyzes the audio and obtains output as text data. The input is compressed audio data, and the output is text data.

[0729] Step 4:

[0730] The server uses language processing tools to analyze the generated text data through natural language processing techniques to identify transaction details. Specifically, it performs data calculations to identify product names and transaction intent. The input is text data, and the output is the analyzed transaction information.

[0731] Step 5:

[0732] The server uses authentication methods to verify the user's identity. This process utilizes voice authentication and biometric authentication technologies to enhance the accuracy of user identification and authentication. The input is the analyzed transaction information, and the output is the authentication result.

[0733] Step 6:

[0734] If the transaction is authenticated, the server executes the purchase process using the payment method. Specifically, it completes the economic transaction through a secure payment service. The input is the authentication result, and the output is information about the completion of the payment.

[0735] Step 7:

[0736] The server notifies the terminal that the payment has been completed via a notification system. The terminal then informs the user of this information via voice or display. The input is the payment completion information, and the output is the transaction completion notification to the user.

[0737] (Application Example 1)

[0738] Next, we will explain Application Example 1. In the following explanation, the data processing device 12 will be referred to as the "server" and the robot 414 as the "terminal".

[0739] In modern brick-and-mortar stores, there is a demand for customers to complete purchases quickly and efficiently. However, current methods result in waiting times at the checkout, reducing the efficiency of the shopping experience. In particular, there is a need for a new interface that allows payment to be completed by voice without going through a physical checkout system.

[0740] The specific processing performed by the specific processing unit 290 of the data processing device 12 in Application Example 1 is realized by the following means.

[0741] In this invention, the server includes sound acquisition means, sound recognition means, natural language processing means, settlement means, and notification means. This allows users to quickly complete purchase procedures using voice input in physical stores, and after confirming the notified transaction result, they can leave the store with their purchased items without going through a cash register.

[0742] "Sound acquisition means" refers to a device or system that receives audio emitted by a user and converts it into data for subsequent processing.

[0743] "Acoustic recognition means" refers to a processing device that has the function of analyzing acquired audio data and converting it into a corresponding text format.

[0744] "Natural language processing means" refers to a technology that analyzes text data acquired by acoustic recognition means to extract user intent and transaction details.

[0745] "Verification methods" refer to the process of verifying the legitimacy of a transaction and confirming the user's identity based on the extracted transaction details.

[0746] A "settlement instrument" is a device or system that, based on the results of a verification instrument, executes settlement processing to complete a transaction.

[0747] A "notification system" is a system that has audio and visual notification functions to communicate the results of a transaction to the user.

[0748] This invention is a system that enables instant voice payment in physical stores. Users use a terminal such as a smartphone. The terminal uses a high-sensitivity microphone to acquire voice data. This voice data is converted into text data by an acoustic recognition means, and the transaction details are further analyzed by a natural language processing means.

[0749] To facilitate these processes, the server utilizes a speech recognition engine (e.g., Google Speech-to-Text API) to generate text data. The analyzed content is then used to verify the user's identity and determine the validity of the transaction using voice authentication or other methods. Once authentication is complete, a secure payment process is executed using a settlement method. This uses a payment gateway API (e.g., Stripe).

[0750] Once payment is complete, a notification system will alert the user to the transaction result. This notification will be made via audio or on-screen display. It will inform the user that the purchase of the product has been completed in the physical store and confirm that the transaction has been successfully completed.

[0751] As a concrete example, when a user at a cafe says "I'll buy this coffee" using voice input on their smartphone, the device connects with a server and completes the payment process. The server authenticates the transaction using a verification mechanism and notifies the user via a notification mechanism that "the purchase is complete."

[0752] An example of a prompt message would be: "Design a program that uses a speech recognition system to convert what the user says into text in real time and then processes the payment based on that information. Specifically, please show the flow when the user says, 'I will buy this product.'"

[0753] The flow of a specific process in Application Example 1 will be explained using Figure 12.

[0754] Step 1:

[0755] The user makes a voice input near the smartphone. The device receives this voice using a high-sensitivity microphone and acquires it as digital audio data. The input is the user's voice, and the output is digital audio data.

[0756] Step 2:

[0757] The terminal immediately transmits the acquired digital audio data to the server. During this process, the terminal compresses the audio data and transfers it using a secure communication method. The input is digital audio data, and the output is the transmitted audio data.

[0758] Step 3:

[0759] The server uses a speech recognition engine to convert received audio data into text data. This process utilizes a generative AI model to convert speech to text. The input is audio data, and the output is text data.

[0760] Step 4:

[0761] The server performs natural language processing on the text data to extract the transaction details intended by the user. Through this analysis, product names and purchase intentions are identified. The input is text data, and the output is the extracted transaction details.

[0762] Step 5:

[0763] The server uses verification methods to authenticate the user based on the extracted transaction details. This process utilizes voice authentication and other methods to determine the user's legitimacy. The input is the transaction details, and the output is the authentication result.

[0764] Step 6:

[0765] If the verification process is successful, the server will use the settlement method to process the payment. In this process, the user's payment information is used to execute a secure transaction via the payment gateway API. The input is the authenticated transaction details, and the output is the settlement result.

[0766] Step 7:

[0767] The server transmits the payment result to the terminal via a notification system. The terminal notifies the user of a message such as "Purchase complete" via voice or screen display. The input is the payment result, and the output is the notification to the user.

[0768] Furthermore, an emotion engine that estimates the user's emotions may be incorporated. That is, the identification processing unit 290 may use the emotion identification model 59 to estimate the user's emotions and perform identification processing using the user's emotions.

[0769] This invention incorporates an emotion engine that recognizes user emotions into a voice-based payment system. Its aim is to enhance the conventional process of understanding transaction intent from voice input and to provide a more human-like interaction and experience by taking the user's emotional state into consideration during the payment process.

[0770] User operation

[0771] The user expresses their intention to purchase the product by voice, as is typical. For example, they might say, "I want to buy this product." The user's emotions, such as joy or anxiety, may also be included in their voice.

[0772] Terminal operation

[0773] The terminal captures the user's voice directly and transfers it to the server. The terminal's processing involves the clear and accurate collection and transmission of the voice data.

[0774] Processing on the server

[0775] The server uses speech recognition to convert the audio data received from the terminal into text data. This text data is then analyzed using natural language processing (NLP) techniques to confirm the user's purchase intent. Next, the emotion engine identifies emotions from the user's audio data. For example, it infers the user's emotions from changes in tone, volume, and tempo of their voice.

[0776] Once emotions are identified, that data is associated with the transaction intent and passed to the authentication mechanism. The authentication process, in addition to standard identity verification, considers the user's emotional state to perform a more precise intent check. If authentication is successful, the server proceeds with the payment process and securely completes the transaction.

[0777] Notification process

[0778] Once the payment is complete, the server uses the emotion engine's analysis results to notify the device of the outcome using an appropriate tone. The device then verbally communicates this information to the user. For example, if the user is nervous, the server can deliver a message in a calm tone such as, "Your purchase has been successfully completed."

[0779] This invention allows users to enjoy more personalized interactions that take their emotions into account, making the service experience more comfortable and reassuring. A concrete example is a scenario where the system provides reassuring voice support to a user who is feeling anxious while shopping.

[0780] The following describes the processing flow.

[0781] Step 1:

[0782] The user expresses the product they want to purchase by voice, using phrases such as "I want to buy this product."

[0783] Step 2:

[0784] The device acquires the user's voice through the microphone and captures it as audio data.

[0785] Step 3:

[0786] The device sends the acquired audio data to the server. The data is encrypted during this process.

[0787] Step 4:

[0788] The server analyzes the received audio data using a speech recognition engine and converts the audio into text.

[0789] Step 5:

[0790] The server analyzes the text data using natural language processing technology to identify the transaction details.

[0791] Step 6:

[0792] The server further analyzes the voice data using an emotion engine to recognize the user's emotions. This process includes analyzing the tone and volume of the voice.

[0793] Step 7:

[0794] The server performs an authentication process and approves the transaction, taking into account the user's identity and perceived emotions.

[0795] Step 8:

[0796] Upon successful authentication, the server processes the payment and securely completes the transaction.

[0797] Step 9:

[0798] The server generates an appropriate notification tone based on the payment result and the user's sentiment, and sends it to the device.

[0799] Step 10:

[0800] The terminal notifies the user of the payment result via voice. For example, it might say in a calm voice, "Your purchase is complete."

[0801] (Example 2)

[0802] Next, we will describe Example 2. In the following description, the data processing device 12 will be referred to as the "server" and the robot 414 as the "terminal".

[0803] Conventional voice recognition systems can understand a user's transaction intent, but they do not consider the user's emotional state, which can result in a lack of human-like interaction and experience. Furthermore, in authentication processes, ignoring the user's emotional state can lead to insufficient identity verification. Therefore, there is a need for a voice payment system that takes user emotions into account and provides a more natural and secure transaction experience.

[0804] The identification process performed by the identification processing unit 290 of the data processing device 12 in Example 2 is realized by the following means.

[0805] In this invention, the server includes an input means for receiving voice, a conversion means for converting the voice acquired by the input means into text, a processing means for analyzing the text data generated by the conversion means to extract transaction details, an emotion analysis means for identifying emotions from the voice data acquired by the processing means, a verification means for performing authentication based on the extracted transaction details and identified emotions, an operation means for executing processing when the transaction is authenticated, and a communication means for notifying the processing results. This enables a more natural and secure voice-based transaction experience while taking into account the user's emotional state.

[0806] "Input means" refers to a function that acquires audio and provides it in a format that the system can process.

[0807] "Conversion means" refers to the process of converting acquired audio into text format and generating text data.

[0808] The "processing means" refers to a function that analyzes the generated text data to extract the user's transaction intent.

[0809] "Emotion analysis means" refers to a process for identifying a user's emotions from voice data and inferring their emotional state.

[0810] "Verification means" refers to the process of authentication based on extracted transaction details and identified sentiments.

[0811] "Operational means" refers to functions that execute settlement and other related processes once a transaction is authenticated.

[0812] "Means of communication" refers to the process of notifying the user of the processing results, and the results can be conveyed through means such as voice.

[0813] This invention incorporates an emotion analysis function that recognizes user emotions into a voice-based payment system. The aim is to understand the user's transaction intent and provide a natural dialogue and experience that takes their emotional state into account. Specific embodiments are described below.

[0814] User operation

[0815] When purchasing a product, users express their intentions verbally. For example, they might say, "I want to buy this product." The tone and tempo of their voice may reflect their emotions, such as joy or anxiety.

[0816] Terminal operation

[0817] The terminal uses a microphone or similar device to acquire the user's voice information and transmits it to the server as digital audio data. Clear and accurate audio capture is crucial on the terminal side.

[0818] Processing on the server

[0819] The server converts the received audio data into text data using a speech recognition engine. General-purpose speech recognition software, such as a "speech recognition API," is used for this process. The text data is then analyzed using natural language processing techniques to confirm the user's transaction intent. Natural language processing libraries are used for this analysis.

[0820] Next, the emotion analysis engine analyzes changes in tone, intensity, and tempo of the voice to identify the user's emotions. For example, it uses an "emotion analysis tool" to infer emotions such as joy or anxiety. The authentication process provides highly accurate authentication using voice and emotion information.

[0821] Notification process

[0822] Based on the sentiment analysis results, the server creates a message in an appropriate tone and sends it back to the device. The device then informs the user of the message aloud. For example, if the user is nervous, it will play a calm message such as, "Your purchase has been successfully completed."

[0823] This system allows users to enjoy a consistent and personalized experience that takes their emotions into account. As a concrete example, consider the case of online shopping. An example of a prompt message might be: "Analyze the emotions expressed by the user during product purchase and advise on appropriate responses. Specifically, identify kindness, anxiety, and joy, and suggest corresponding voice messages."

[0824] In this way, the present invention makes it possible to provide a more enriching user experience by considering and analyzing emotions during the process of transactions conducted via voice.

[0825] The flow of the specific processing in Example 2 will be explained using Figure 13.

[0826] Step 1:

[0827] The user expresses their intention to purchase a product by voice. For example, they might say, "I want to buy this product." This voice information is acquired by the terminal and sent to the server as digital voice data. The terminal captures the voice input using a microphone and sends it to the server in the form of an audio file. During this process, efforts are made to minimize noise and maintain the clarity and accuracy of the voice.

[0828] Step 2:

[0829] The server converts the audio data received from the terminal into text data using a speech recognition engine. Specifically, it analyzes the audio waveform using speech recognition software and converts it into a corresponding string of characters. Through this process, the text "I want to buy this product" is obtained from the audio.

[0830] Step 3:

[0831] The server analyzes the converted text data using natural language processing techniques to extract transaction intent. For example, it uses a natural language processing library to perform grammatical analysis and keyword extraction to identify the intent to "purchase a product." Text data is used as input, and intent and related information are included in the output.

[0832] Step 4:

[0833] The server uses an emotion analysis engine to identify emotions from audio data. Using an audio analysis tool, it evaluates the tone, intensity, and tempo of the voice to infer the emotional state (e.g., joy, anxiety). The emotion analysis engine takes audio data as input and outputs emotion tags such as "joy."

[0834] Step 5:

[0835] The server performs authentication based on identified transaction details and emotions. The verification process uses voice authentication technology to verify identity and enhance reliability. For example, it compares voiceprints and outputs a result indicating whether voice authentication was successful or not.

[0836] Step 6:

[0837] If the transaction is authenticated, the server will execute the payment through the payment processing service. Payment information is entered, and the transaction ID and payment status are received as output. Encryption technology may be used to ensure the security of the payment process.

[0838] Step 7:

[0839] The server creates a message based on the sentiment analysis results and sends it to the device. The tone of the message is adjusted according to the user's emotions, and it is presented in a way that is easy for the user to understand. The device notifies the user of the received message via voice, confirming the completion of the purchase. For example, if the user is nervous, a message such as, "Your purchase has been successfully completed. Please rest assured," is played in a gentle tone.

[0840] (Application Example 2)

[0841] Next, we will explain application example 2. In the following explanation, the data processing device 12 will be referred to as the "server" and the robot 414 as the "terminal".

[0842] In voice-based payment systems, there is a need to provide sophisticated communication and reassuring interactions that take into account not only the transaction details intended by the user but also their emotional state at the time. However, conventional voice payment systems have the challenge of not being able to adequately reflect the user's emotional state and not being able to provide sufficient reassurance to users who feel anxious.

[0843] The specific processing performed by the specific processing unit 290 of the data processing device 12 in Application Example 2 is realized by the following means.

[0844] In this invention, the server includes data acquisition means for receiving voice input, voice recognition means for converting voice data into text data, and emotion analysis means for analyzing the user's emotional state. This makes it possible to analyze the transaction content and emotions from the user's voice and provide voice feedback that takes into account a sense of security.

[0845] "Data acquisition means" refers to a device or function that receives voice input and collects it as voice data.

[0846] "Speech recognition means" refers to a technology or system that analyzes acquired speech data and converts it into corresponding text data.

[0847] "Natural language processing means" refers to software or algorithms that analyze text data generated by speech recognition means to accurately extract transaction details.

[0848] "Emotion analysis tools" are technologies that analyze a user's emotional state from their speech and identify emotions based on changes in voice tone and tempo.

[0849] "Authentication means" refers to a process or function that verifies the validity of a transaction based on the transaction details and emotional state, and verifies the user's identity as necessary.

[0850] A "data processing system" is a system for executing authenticated transactions and ensuring settlement.

[0851] An "information transmission means" is a mechanism that transmits the results of a transaction to the user and notifies them through voice or other interfaces.

[0852] This invention is a system that, when conducting transactions based on user voice input, analyzes the user's emotions using voice data and provides personalized feedback based on that analysis. The server is responsible for acquiring, processing, and analyzing the voice data, confirming the transaction details, and then authenticating and executing the payment.

[0853] The terminal captures high-quality voice input from the user and sends it to the server via a data acquisition method. The server uses the Google Speech-to-Text API to convert the speech to text and extracts transaction details through natural language processing. Furthermore, it uses the Google Cloud Natural Language API to perform sentiment analysis based on factors such as tone and intensity of the speech.

[0854] The analysis results are used for identity verification through authentication methods, and if the transaction is successfully authenticated, the payment process is executed via the Stripe API. Based on the sentiment results obtained during this process, feedback is provided to the user in an appropriate tone as a means of communication.

[0855] For example, if a user says in voice, "I want to buy this product, but I'm a little worried," the emotion analysis system can recognize the user's anxiety and respond with a calm voice message such as, "Please relax, the purchase process is going smoothly."

[0856] Example prompt: "The user is attempting to purchase a specific product with an anxious voice. The sentiment analysis engine should recognize this anxiety and generate reassuring voice feedback."

[0857] The flow of a specific process in Application Example 2 will be explained using Figure 14.

[0858] Step 1:

[0859] The user performs voice input. The user uses their smartphone to express their intention to purchase a product by voice. The input is the user's speech and is recorded directly on the device.

[0860] Step 2:

[0861] The terminal acquires audio data and sends it to the server. The terminal uses its built-in microphone to save the audio input as a digital audio file and sends it to the server via a data acquisition device. The input is audio data, and the output is a digital audio file sent to the server.

[0862] Step 3:

[0863] The server converts audio data into text using speech recognition technology. Using the Google Speech-to-Text API, the server generates corresponding text data from the audio data. The input is audio data, and the output is text data.

[0864] Step 4:

[0865] The server analyzes text data using natural language processing techniques and extracts transaction details. The server applies natural language processing algorithms to extract transaction intent and product characteristics from the text. The input is text data, and the output is structured data containing transaction details.

[0866] Step 5:

[0867] The server identifies the user's emotional state using sentiment analysis tools. Leveraging the Google Cloud Natural Language API, the server analyzes changes in tone and tempo of speech to determine the emotional state. The input is audio data, and the output is information about the user's emotional state.

[0868] Step 6:

[0869] The server uses authentication methods to authenticate the transaction. It takes into account the transaction details and sentiment status, and verifies the user's identity through the authentication process. The input is the transaction details and sentiment information, and the output is the authentication result.

[0870] Step 7:

[0871] If the server authenticates the transaction, it initiates the settlement process using data processing tools. The settlement is securely executed using the Stripe API. The input is the authentication result, and the output is a settlement completion notification.

[0872] Step 8:

[0873] The server notifies the user of the payment result via a communication method. The server generates voice feedback reflecting the sentiment analysis results and delivers it to the user through the terminal. The input is the payment completion notification and sentiment information, and the output is a personalized voice message.

[0874] The specific processing unit 290 transmits the result of the specific processing to the robot 414. In the robot 414, the control unit 46A causes the speaker 240 and the controlled object 443 to output the result of the specific processing. The microphone 238 acquires audio indicating user input for the result of the specific processing. The control unit 46A transmits the audio data indicating user input acquired by the microphone 238 to the data processing unit 12. In the data processing unit 12, the specific processing unit 290 acquires the audio data.

[0875] Data generation model 58 is a type of so-called generative AI (Artificial Intelligence). One example of data generation model 58 is ChatGPT (Internet search<URL: https: / / openai.com / blog / chatgpt> ), Gemini (Internet search) <url: https: gemini.google.com ?hl="ja">Examples of generative AI include the following. The data generation model 58 is obtained by performing deep learning on a neural network. The data generation model 58 is input with prompts containing instructions, and with inference data such as audio data representing speech, text data representing text, and image data representing images. The data generation model 58 infers from the input inference data according to the instructions indicated by the prompts, and outputs the inference results in data formats such as audio data and text data. Here, inference refers to, for example, analysis, classification, prediction, and / or summarization.

[0876] In the above embodiment, an example was given in which specific processing is performed by the data processing device 12, but the technology of this disclosure is not limited thereto, and the specific processing may also be performed by the robot 414.

[0877] Furthermore, the emotion identification model 59, acting as an emotion engine, may determine the user's emotion according to a specific mapping. Specifically, the emotion identification model 59 may determine the user's emotion according to a specific mapping, which is an emotion map (see Figure 9). Similarly, the emotion identification model 59 may also determine the robot's emotion, and the identification processing unit 290 may perform identification processing using the robot's emotion.

[0878] Figure 9 shows an emotion map 400 in which multiple emotions are mapped. In the emotion map 400, emotions are arranged in concentric circles radiating from the center. The closer to the center of the concentric circles, the more primitive the emotions are located. Further out of the concentric circles, emotions representing states and actions arising from mental states are located. Emotion is a concept that includes feelings and mental states. On the left side of the concentric circles, emotions that are generally generated from reactions occurring in the brain are located. On the right side of the concentric circles, emotions that are generally induced by situational judgment are located. Above and below the concentric circles, emotions that are generally generated from reactions occurring in the brain and induced by situational judgment are located. In addition, the emotion of "pleasure" is located on the upper side of the concentric circles, and the emotion of "displeasure" is located on the lower side. Thus, in the emotion map 400, multiple emotions are mapped based on the structure in which emotions arise, and emotions that are likely to occur simultaneously are mapped close together.

[0879] These emotions are distributed at the 3 o'clock position on the Emotion Map 400, and usually fluctuate between feelings of security and anxiety. In the right half of the Emotion Map 400, situational awareness takes precedence over internal feelings, resulting in a calm impression.

[0880] The inside of the Emotion Map 400 represents inner thoughts, while the outside represents actions. Therefore, the further you go from the outside of the Emotion Map 400, the more visible (expressed in actions) your emotions become.

[0881] Here, human emotions are based on various balances, such as posture and blood sugar levels. When these balances deviate from the ideal, it results in discomfort, and when they approach the ideal, it results in pleasure. Similarly, in robots, cars, motorcycles, etc., emotions can be created based on various balances, such as posture and battery level. When these balances deviate from the ideal, it results in discomfort, and when they approach the ideal, it results in pleasure. The emotion map can be generated, for example, based on Dr. Mitsuyoshi's emotion map (Research on a system for analyzing brain physiological signals of speech emotion recognition and emotion, Tokushima University, doctoral dissertation: https: / / ci.nii.ac.jp / naid / 500000375379). The left half of the emotion map contains emotions belonging to a region called "response," where sensation is dominant. The right half of the emotion map contains emotions belonging to a region called "situation," where situational awareness is dominant.

[0882] The emotion map defines two emotions that promote learning. One is the emotion around the middle of the negative "repentance" and "reflection" on the situation side. In other words, it is when the robot experiences negative emotions such as "I never want to feel this way again" or "I don't want to be scolded again." The other is the emotion around the positive "desire" on the reaction side. In other words, it is when the robot has positive feelings such as "I want more" or "I want to know more."

[0883] The emotion identification model 59 inputs user input into a pre-trained neural network, obtains emotion values ​​representing each emotion shown in the emotion map 400, and determines the user's emotion. This neural network is pre-trained based on multiple training data sets, which are combinations of user input and emotion values ​​representing each emotion shown in the emotion map 400. Furthermore, this neural network is trained so that emotions located close together have similar values, as shown in the emotion map 900 in Figure 10. Figure 10 shows an example where multiple emotions such as "reassured," "calm," and "confident" have similar emotion values.

[0884] The above description primarily focuses on the functions of the data processing device 12 in relation to this disclosure. However, the system related to this disclosure is not necessarily implemented on a server. The system related to this disclosure may be implemented as a general information processing system. This disclosure may be implemented, for example, as a software program that runs on a personal computer or as an application that runs on a smartphone. The method related to this disclosure may be provided to users in SaaS (Software as a Service) format.

[0885] In the above embodiment, an example was given in which a specific process is performed by a single computer 22. However, the technology of this disclosure is not limited thereto, and a distributed processing of the specific process may be performed by multiple computers, including computer 22. For example, a data generation model 58 may be provided in an external device of the data processing device 12, and the external device may generate data according to the input data.

[0886] In the above embodiment, an example was given in which the specific processing program 56 is stored in the storage 32, but the technology of this disclosure is not limited thereto. For example, the specific processing program 56 may be stored in a portable, computer-readable, non-temporary storage medium such as a USB (Universal Serial Bus) memory. The specific processing program 56 stored in the non-temporary storage medium is installed in the computer 22 of the data processing device 12. The processor 28 executes specific processing according to the specific processing program 56.

[0887] Alternatively, the specific processing program 56 may be stored in a storage device such as a server connected to the data processing device 12 via the network 54, and the specific processing program 56 may be downloaded and installed on the computer 22 in response to a request from the data processing device 12.

[0888] Furthermore, it is not necessary to store the entirety of the specific processing program 56 in a storage device such as a server connected to the data processing device 12 via the network 54, or to store the entirety of the specific processing program 56 in the storage 32; it is acceptable to store only a portion of the specific processing program 56.

[0889] The following types of processors can be used as hardware resources to perform specific processing. Examples of processors include a CPU, a general-purpose processor that functions as a hardware resource to perform specific processing by executing software, i.e., a program. Other examples of processors include dedicated electrical circuits, such as FPGAs (Field-Programmable Gate Arrays), PLDs (Programmable Logic Devices), or ASICs (Application Specific Integrated Circuits), which have circuit configurations specifically designed to perform specific processing. All of these processors have built-in or connected memory, and all of them perform specific processing by using memory.

[0890] The hardware resource that performs a specific process may consist of one of these various processors, or it may consist of a combination of two or more processors of the same or different types (for example, a combination of multiple FPGAs, or a combination of a CPU and an FPGA). Alternatively, the hardware resource that performs a specific process may consist of a single processor.

[0891] Examples of configurations using a single processor include, firstly, a configuration in which one or more CPUs and software are combined to form a single processor, and this processor functions as a hardware resource that performs a specific process. Secondly, there is a configuration using a processor that realizes the functions of the entire system, including multiple hardware resources that perform a specific process, on a single IC chip, as exemplified by SoCs (System-on-a-chip). In this way, a specific process is realized using one or more of the above types of processors as hardware resources.

[0892] Furthermore, the hardware structure of these various processors can more specifically utilize electrical circuits that combine circuit elements such as semiconductor devices. Also, the specific processing described above is merely an example. Therefore, it goes without saying that unnecessary steps can be deleted, new steps added, or the processing order rearranged, as long as it does not deviate from the main purpose.

[0893] The descriptions and illustrations presented above are detailed explanations of the technical aspects of this disclosure and are merely examples of the technical aspects. For example, the above descriptions of the structure, function, operation, and effect are examples of the structure, function, operation, and effect of the technical aspects of this disclosure. Therefore, it goes without saying that you may delete unnecessary parts, add new elements, or replace elements in the descriptions and illustrations presented above, as long as you do not deviate from the essence of the technical aspects of this disclosure. Furthermore, in order to avoid confusion and facilitate understanding of the technical aspects of this disclosure, explanations of common technical knowledge and the like that do not require special explanation to enable the implementation of the technical aspects of this disclosure have been omitted from the descriptions and illustrations presented above.

[0894] All documents, patent applications, and technical standards described herein are incorporated by reference to the same extent as if each individual document, patent application, and technical standard were specifically and individually noted as being incorporated by reference.

[0895] The following is further disclosed regarding the embodiments described above.

[0896] (Claim 1)

[0897] A voice acquisition means for receiving voice input,

[0898] A speech recognition means that converts speech data acquired by the speech acquisition means into text data,

[0899] A natural language processing means that analyzes the text data generated by the speech recognition means to extract transaction details,

[0900] An authentication method that authenticates transactions based on extracted transaction details,

[0901] A payment method that executes settlement processing when a transaction is authenticated,

[0902] A notification method for notifying the transaction results,

[0903] A system that includes this.

[0904] (Claim 2)

[0905] The system according to claim 1, wherein the authentication means uses voice authentication to verify the identity of the user.

[0906] (Claim 3)

[0907] The system according to claim 1, wherein the notification means provides voice notification of the settlement result.

[0908] "Example 1"

[0909] (Claim 1)

[0910] Information acquisition means for receiving voice input,

[0911] A speech recognition means that converts speech information acquired by the aforementioned information acquisition means into text information,

[0912] A language processing means that analyzes the character information generated by the speech recognition means and extracts its content,

[0913] An authentication method that verifies identity based on the extracted information,

[0914] A payment method that executes processing once the transaction is authenticated,

[0915] A notification method for providing transaction results,

[0916] A communication method that compresses audio into an appropriate format and transfers information using secure communication,

[0917] A system that includes this.

[0918] (Claim 2)

[0919] The system according to claim 1, wherein the authentication means uses voice authentication to verify the identity of the individual.

[0920] (Claim 3)

[0921] The system according to claim 1, wherein the notification means provides voice notification of the transaction result.

[0922] "Application Example 1"

[0923] (Claim 1)

[0924] A means for receiving audio input and an acoustic acquisition means,

[0925] An acoustic recognition means that converts acoustic data acquired by the acoustic acquisition means into character data,

[0926] A natural language processing means that analyzes the character data generated by the aforementioned acoustic recognition means to extract transaction details,

[0927] A verification method for confirming transactions based on extracted transaction details,

[0928] A settlement method that executes settlement processing when a transaction is confirmed,

[0929] A notification method for informing about the transaction results,

[0930] The notification means includes means for notifying the user of the completion of a transaction in a physical store, both audibly and visually.

[0931] A system that includes this.

[0932] (Claim 2)

[0933] The system according to claim 1, wherein the verification means uses acoustic verification to verify the identity of the user.

[0934] (Claim 3)

[0935] The system according to claim 1, wherein voice input includes instructions for transactions at a physical store.

[0936] "Example 2 of combining an emotion engine"

[0937] (Claim 1)

[0938] An input means for receiving sound,

[0939] A conversion means that converts the audio acquired by the input means into text,

[0940] A processing means for analyzing the text data generated by the conversion means and extracting transaction details,

[0941] An emotion analysis means for identifying emotions from audio data acquired by the processing means,

[0942] A verification method that performs authentication based on extracted transaction details and identified sentiment,

[0943] An operational means to execute a process when a transaction is authenticated,

[0944] A means of communication for notifying the processing results,

[0945] A system that includes this.

[0946] (Claim 2)

[0947] The system according to claim 1, wherein the verification means uses voice to verify the user.

[0948] (Claim 3)

[0949] The system according to claim 1, wherein the communication means notifies the result by voice.

[0950] "Application example 2 of combining emotional engines"

[0951] (Claim 1)

[0952] A data acquisition means for receiving voice input,

[0953] A speech recognition means that converts the speech data acquired by the data acquisition means into text data,

[0954] A natural language processing means that analyzes the text data generated by the speech recognition means to extract transaction details,

[0955] A means of analyzing the emotional state of a user,

[0956] An authentication method that authenticates transactions based on extracted transaction details and the user's emotional state,

[0957] A data processing means that executes settlement processing when a transaction is authenticated,

[0958] A means of transmitting information to notify the results of a transaction,

[0959] A system that includes this.

[0960] (Claim 2)

[0961] The system according to claim 1, wherein the authentication means uses voice authentication to verify the identity of the user.

[0962] (Claim 3)

[0963] The system according to claim 1, wherein the information transmission means provides voice notification of the settlement result based on the analysis result of the emotion analysis means. [Explanation of Symbols]

[0964] 10, 210, 310, 410 Data Processing Systems 12 Data Processing Devices 14 Smart Devices 214 Smart Glasses 314 Headset-type terminal 414 Robots< / url:> < / url:> < / url:> < / url:>

Claims

1. A voice acquisition means for receiving voice input, A speech recognition means that converts speech data acquired by the speech acquisition means into text data, A natural language processing means that analyzes the text data generated by the speech recognition means to extract transaction details, An authentication method that authenticates transactions based on extracted transaction details, A payment method that executes settlement processing when a transaction is authenticated, A notification method for notifying the transaction results, A system that includes this.

2. The system according to claim 1, wherein the authentication means uses voice authentication to verify the identity of the user.

3. The system according to claim 1, wherein the notification means provides voice notification of the settlement result.

Citation Information

Patent Citations

  • Persona chatbot control method and system

    JP2022180282A