System
A voice-based payment system using AI-generated authentication words and voiceprint authentication addresses the challenge of QR payment accessibility for visually impaired and elderly users, providing secure and convenient transactions.
Patent Information
- Application Number
- JP2024136424
- Authority / Receiving Office
- JP · JP
- Patent Type
- Applications
- Current Assignee / Owner
- Filing Date
- 2024-08-16
- Publication Date
- 2026-02-27
AI Technical Summary
Conventional QR payments are difficult for individuals with visual impairments and the elderly, especially in situations where smartphones cannot be used.
A voice-based payment system utilizing a generation AI to generate authentication words, which are communicated to the user, and voiceprint authentication is performed to confirm the user's identity for secure payment transactions.
Enables secure and convenient voice-based payments for individuals who cannot use smartphones or have difficulty with QR code payments, including the visually impaired and elderly, in various situations.
Smart Images

Figure 2026033382000001_ABST
Abstract
Description
[Technical Field]
[0001] The technology of the present disclosure relates to a system. [Background technology]
[0002] Patent document 1 discloses a persona chatbot control method performed by at least one processor, the method including the steps of receiving a user utterance, adding the user utterance to a prompt including an instruction sentence related to a description of the chatbot character, encoding the prompt, and inputting the encoded prompt into a language model to generate a chatbot utterance in response to the user utterance. [Prior art documents] [Patent documents]
[0003] [Patent Document 1] Japanese Patent Publication No. 2022-180282 Summary of the Invention [Problem to be solved by the invention]
[0004] With conventional technology, QR payments were difficult to use in situations where smartphones could not be used, for people with visual impairments, and for the elderly.
[0005] The system according to the embodiment aims to realize voice-based payment that can be used by anyone in any situation. [Means for solving the problem]
[0006] The system according to the embodiment includes a generation unit, a transmission unit, a voiceprint authentication unit, and a payment unit. The generation unit generates an authentication word. The transmission unit communicates the authentication word generated by the generation unit to the purchaser. The voiceprint authentication unit performs voiceprint authentication on the purchaser's voice when the purchaser speaks the authentication word communicated by the transmission unit. The payment unit performs payment based on the result of authentication by the voiceprint authentication unit. [Effects of the Invention]
[0007] The system according to the embodiment can realize voice-based payment that can be used by anyone in any situation. [Brief explanation of the drawings]
[0008] [Figure 1] 1 is a conceptual diagram showing an example of the configuration of a data processing system according to a first embodiment. [Figure 2] 1 is a conceptual diagram showing an example of main functions of a data processing device and a smart device according to a first embodiment. [Figure 3] FIG. 10 is a conceptual diagram showing an example of the configuration of a data processing system according to a second embodiment. [Figure 4] FIG. 10 is a conceptual diagram showing an example of main functions of a data processing device and smart glasses according to a second embodiment. [Figure 5] FIG. 10 is a conceptual diagram showing an example of the configuration of a data processing system according to a third embodiment. [Figure 6] FIG. 11 is a conceptual diagram showing an example of main functions of a data processing device and a headset-type terminal according to a third embodiment. [Figure 7] FIG. 10 is a conceptual diagram showing an example of the configuration of a data processing system according to a fourth embodiment. [Figure 8] FIG. 10 is a conceptual diagram showing an example of main functions of a data processing device and a robot according to a fourth embodiment. [Figure 9] 1 shows an emotion map onto which multiple emotions are mapped. [Figure 10] 1 shows an emotion map onto which multiple emotions are mapped. DETAILED DESCRIPTION OF THE INVENTION
[0009] An example of an embodiment of a system according to the technology of the present disclosure will be described below with reference to the accompanying drawings.
[0010] First, the terms used in the following description will be explained.
[0011] In the following embodiments, a coded processor (hereinafter simply referred to as a "processor") may be a single arithmetic device or a combination of multiple arithmetic devices. Furthermore, the processor may be a single type of arithmetic device or a combination of multiple types of arithmetic devices. Examples of arithmetic devices include a CPU (Central Processing Unit), a GPU (Graphics Processing Unit), a GPGPU (General-Purpose computing on Graphics Processing Units), an APU (Accelerated Processing Unit), or a TPU (Tensor Processing Unit).
[0012] In the following embodiments, a coded RAM (Random Access Memory) is a memory in which information is temporarily stored and is used as a working memory by a processor.
[0013] In the following embodiments, the coded storage is one or more non-volatile storage devices that store various programs, various parameters, etc. Examples of non-volatile storage devices include flash memory (SSD (Solid State Drive)), magnetic disks (e.g., hard disks), and magnetic tapes.
[0014] In the following embodiments, a communication I / F (Interface) with a symbol is an interface including a communication processor, an antenna, etc. The communication I / F controls communication between multiple computers. Examples of communication standards applied to the communication I / F include wireless communication standards including 5G (5th Generation Mobile Communication System), Wi-Fi (registered trademark), and Bluetooth (registered trademark).
[0015] In the following embodiments, "A and / or B" is synonymous with "at least one of A and B." In other words, "A and / or B" means that it may be only A, only B, or a combination of A and B. Furthermore, in this specification, the same concept as "A and / or B" is also applied when three or more things are expressed connected by "and / or."
[0016] [First embodiment] FIG. 1 shows an example of the configuration of a data processing system 10 according to the first embodiment.
[0017] 1, a data processing system 10 includes a data processing device 12 and a smart device 14. An example of the data processing device 12 is a server.
[0018] The data processing device 12 includes a computer 22, a database 24, and a communication I / F 26. The computer 22 includes a processor 28, a RAM 30, and a storage 32. The processor 28, RAM 30, and storage 32 are connected to a bus 34. The database 24 and the communication I / F 26 are also connected to the bus 34. The communication I / F 26 is connected to a network 54. Examples of the network 54 include a WAN (Wide Area Network) and / or a LAN (Local Area Network).
[0019] The smart device 14 includes a computer 36, a reception device 38, an output device 40, a camera 42, and a communication I / F 44. The computer 36 includes a processor 46, a RAM 48, and a storage 50. The processor 46, the RAM 48, and the storage 50 are connected to a bus 52. The reception device 38, the output device 40, and the camera 42 are also connected to the bus 52.
[0020] The reception device 38 includes a touch panel 38A and a microphone 38B, and receives user input. The touch panel 38A detects contact with a pointer (for example, a pen or a finger) to receive user input by the touch of the pointer. The microphone 38B detects the user's voice to receive user input by voice. The control unit 46A transmits data indicating the user input received by the touch panel 38A and the microphone 38B to the data processing device 12. In the data processing device 12, the specific processing unit 290 (see FIG. 2) acquires the data indicating the user input.
[0021] Output device 40 includes a display 40A and a speaker 40B, and presents data to a user by outputting the data in a form of expression that the user can perceive (e.g., audio and / or text). Display 40A displays visible information such as text and images in accordance with instructions from processor 46. Speaker 40B outputs audio in accordance with instructions from processor 46. Camera 42 is a compact digital camera equipped with an optical system including a lens, aperture, and shutter, and an imaging element such as a CMOS (Complementary Metal-Oxide-Semiconductor) image sensor or a CCD (Charge Coupled Device) image sensor.
[0022] The communication I / F 44 is connected to a network 54. The communication I / Fs 44 and 26 control the exchange of various information between the processor 46 and the processor 28 via the network 54.
[0023] FIG. 2 shows an example of the main functions of the data processing device 12 and the smart device 14.
[0024] 2, in the data processing device 12, a specific process is performed by the processor 28. A specific processing program 56 is stored in the storage 32. The specific processing program 56 is an example of a "program" according to the technology of the present disclosure. The processor 28 reads the specific processing program 56 from the storage 32 and executes the read specific processing program 56 on the RAM 30. The specific process is realized by the processor 28 operating as a specific processing unit 290 in accordance with the specific processing program 56 executed on the RAM 30.
[0025] The storage 32 stores a data generation model 58 and an emotion identification model 59. The data generation model 58 and the emotion identification model 59 are used by the identification processing unit 290. The identification processing unit 290 can estimate a user's emotion using the emotion identification model 59 and perform identification processing using the user's emotion. The emotion estimation function (emotion identification function) using the emotion identification model 59 performs various estimations and predictions regarding the user's emotion, including estimation and prediction of the user's emotion, but is not limited to these examples. Furthermore, the estimation and prediction of emotion also includes, for example, emotion analysis.
[0026] In the smart device 14, the specific processing is performed by the processor 46. The storage 50 stores a specific processing program 60. The specific processing program 60 is used together with the specific processing program 56 by the data processing system 10. The processor 46 reads the specific processing program 60 from the storage 50 and executes the read specific processing program 60 on the RAM 48. The specific processing is realized by the processor 46 operating as the control unit 46A in accordance with the specific processing program 60 executed on the RAM 48. Note that the smart device 14 has a data generation model and an emotion identification model similar to the data generation model 58 and the emotion identification model 59, and can also perform processing similar to that of the specific processing unit 290 using these models.
[0027] Note that a device other than the data processing device 12 may have the data generation model 58. For example, a server device (e.g., a generation server) may have the data generation model 58. In this case, the data processing device 12 obtains a processing result (prediction result, etc.) using the data generation model 58 by communicating with the server device having the data generation model 58. Furthermore, the data processing device 12 may be a server device, or may be a terminal device owned by a user (e.g., a mobile phone, a robot, a home appliance, etc.). Next, an example of processing by the data processing system 10 according to the first embodiment will be described.
[0028] (Example 1) A payment system according to an embodiment of the present invention enables voice-based payments for those who cannot use smartphones, such as those with visual impairments or the elderly, or those in situations where QR code payments are difficult to use. In this payment system, a generation AI generates an authentication word at checkout, and when the customer speaks the authentication word, voiceprint authentication is performed on the customer's voice to complete the payment. For example, in this payment system, a generation AI randomly generates an authentication word at checkout and communicates it to the customer. When the customer speaks the generated authentication word, the customer's voice is recorded and voiceprint authentication is performed. Voiceprint authentication analyzes the customer's voice and compares it with pre-registered voiceprint data. If the comparison is successful, the payment is completed. This allows for easy payments even for those who cannot use smartphones, such as those without smartphones, visual impairments, or the elderly. For example, visually impaired customers can complete payments using only their voice when purchasing items in a store, without using a smartphone. Furthermore, elderly customers who are unfamiliar with smartphones can easily make payments using only their voice. Furthermore, payments can be made using only their voice, even when their hands are full or they do not have a smartphone. This allows for convenient payment access for a wider range of people. This makes the payment system easy to use, even for those who have difficulty using QR payments, or those who are unable to use smartphones, such as the visually impaired or elderly. For example, when purchasing goods in a store, a visually impaired person can complete the payment using only their voice, without using a smartphone. Also, even elderly people who are unfamiliar with using smartphones can easily make payments using only their voice. Furthermore, even if their hands are full or they do not have a smartphone, they can make payments using only their voice. This makes payment convenient for more people.
[0029] The payment system according to the embodiment includes a generation unit, a transmission unit, a voiceprint authentication unit, and a payment unit. The generation unit generates an authentication word. The generation unit randomly generates the authentication word using, for example, a generation AI. The generation AI may be a text generation AI (e.g., LLM) or a multimodal generation AI. The generation unit may generate words such as "apple pie" or "sunflower." The generation unit may also generate the authentication word based on the customer's emotions and past authentication history. For example, if the customer is nervous, it may generate a simple word. The transmission unit communicates the generated authentication word to the customer. The transmission unit may communicate the authentication word aloud using, for example, speech synthesis technology. The transmission unit may also adjust the tone and speed of the voice based on the customer's hearing characteristics and emotions. For example, the authentication word may be communicated at a slower speed to an elderly customer. The voiceprint authentication unit analyzes the customer's voice and compares it with pre-registered voiceprint data. The voiceprint authentication unit may analyze the customer's voice using, for example, speech signal processing technology, and extract features. The voiceprint authentication unit can also optimize the authentication algorithm by taking into account the buyer's emotions and voice fluctuations. For example, if the buyer has a cold, authentication is performed taking these fluctuations into account. The payment unit performs payment if voiceprint authentication is successful. The payment unit performs payment using means such as credit card payment or electronic money payment. The payment unit can also select the optimal payment method based on the buyer's emotions and past payment history. For example, if the buyer is in a hurry, a quick payment method can be provided. As a result, the payment system according to the embodiment allows anyone to easily make payments in any situation.
[0030] The generation unit randomly generates an authentication word. The generation unit randomly generates an authentication word using, for example, a random number generation algorithm. The generation unit can also randomly generate an authentication word using a generation AI. The generation AI can be a text generation AI (e.g., LLM) or a multimodal generation AI. For example, the generation unit randomly generates words such as "apple pie" or "sunflower." The generation unit can also generate an authentication word based on the buyer's emotions or past authentication history. For example, if the buyer is nervous, a simple word can be generated. This improves security by randomly generating authentication words.
[0031] The communication unit communicates the generated authentication word to the purchaser by voice. The communication unit communicates the authentication word by voice, for example, using speech synthesis technology. The communication unit can also communicate the authentication word by voice using generation AI. The generation AI can be a text generation AI (e.g., LLM) or a multimodal generation AI. For example, the communication unit communicates words such as "apple pie" or "sunflower" by voice. The communication unit can also adjust the tone and speed of the voice based on the purchaser's hearing characteristics and emotions. For example, for elderly people, the authentication word can be communicated at a slower speed. As a result, communicating the authentication word by voice makes it easier for visually impaired people and elderly people to use.
[0032] The voiceprint authentication unit analyzes the purchaser's voice and compares it with pre-registered voiceprint data. The voiceprint authentication unit may, for example, use voice signal processing technology to analyze the purchaser's voice and extract features. The voiceprint authentication unit may also analyze the purchaser's voice using generation AI. The generation AI may be a text generation AI (e.g., LLM) or a multimodal generation AI. For example, the voiceprint authentication unit may record the purchaser's voice and compare it with pre-registered voiceprint data. The voiceprint authentication unit may also optimize the authentication algorithm by taking into account the purchaser's emotions and voice fluctuations. For example, if the purchaser has a cold, authentication may take these fluctuations into account. This ensures that the purchaser's identity is confirmed through voiceprint authentication.
[0033] The payment unit executes payment if the voiceprint authentication is successful. The payment unit executes payment using means such as credit card payment or electronic money payment. The payment unit can also execute payment using generation AI. The generation AI can be a text generation AI (e.g., LLM) or a multimodal generation AI. For example, the payment unit executes payment if the voiceprint authentication is successful. The payment unit can also select the optimal payment method based on the buyer's emotions and past payment history. For example, if the buyer is in a hurry, a quick payment method can be provided. This improves security because payment is executed only if the voiceprint authentication is successful.
[0034] The payment unit can provide a payment method that can be used by people who cannot use smartphones, such as those who are unable to own a smartphone, those who are visually impaired, or the elderly. The payment unit can also provide a payment method using generation AI. Generation AI includes text generation AI (e.g., LLM) and multimodal generation AI. For example, the payment unit provides a means for visually impaired people to complete payment using only their voice when purchasing items in a store, without using a smartphone. The payment unit can provide a means for elderly people who are unfamiliar with using smartphones to easily make payments using only their voice. This allows people who cannot use smartphones to make payments easily.
[0035] The generation unit can analyze the purchaser's past authentication history and generate optimal authentication words. For example, the generation unit analyzes the purchaser's past authentication history and generates optimal authentication words. The generation unit can also analyze past authentication history using generation AI. The generation AI is a text generation AI (e.g., LLM) or multimodal generation AI. For example, the generation unit analyzes patterns of authentication words that the purchaser has used successfully in the past and generates authentication words with similar patterns. The generation unit can also avoid authentication words that the purchaser has used unsuccessfully in the past and generate authentication words with a high success rate. The generation unit can also generate optimal authentication words for specific time periods or situations from the purchaser's past authentication history. This generates optimal authentication words based on the past authentication history, improving the authentication success rate.
[0036] When generating authentication words, the generation unit can generate authentication words in different languages based on the language settings of the purchaser. For example, the generation unit generates authentication words in different languages based on the language settings of the purchaser's device. The generation unit can also generate authentication words based on the language settings using generation AI. The generation AI can be a text generation AI (e.g., LLM) or a multimodal generation AI. For example, if the purchaser speaks English, the generation unit generates authentication words in English. For example, it selects words such as "apple" or "banana." If the purchaser speaks Japanese, the generation unit can also generate authentication words in Japanese. For example, it selects words such as "ringo" or "banana." If the purchaser speaks multiple languages, the generation unit can also randomly generate authentication words from multiple languages. For example, it selects words in English and Japanese alternately. This facilitates authentication because authentication words are generated according to the purchaser's language settings.
[0037] When generating authentication words, the generation unit can generate the authentication words based on the purchaser's current environmental sounds. For example, the generation unit generates authentication words taking into account the purchaser's current environmental sounds. The generation unit can also generate authentication words taking into account environmental sounds using generation AI. The generation AI can be a text generation AI (e.g., LLM) or a multimodal generation AI. For example, if the purchaser is in a quiet environment, the generation unit generates complex authentication words. For example, it selects long words such as "elephant" or "sunflower." If the purchaser is in a noisy environment, the generation unit can also generate simple authentication words. For example, it selects short words such as "cat" or "dog." If the purchaser is listening to music, the generation unit can also generate authentication words related to the music genre. For example, if the purchaser is listening to rock music, it selects words such as "guitar" or "drums." This makes authentication easier because authentication words are generated according to the environmental sounds.
[0038] When generating authentication words, the generation unit can generate region-specific authentication words taking into account the geographic location information of the purchaser. For example, the generation unit generates region-specific authentication words taking into account the geographic location information of the purchaser. The generation unit can also generate authentication words taking into account the geographic location information using generation AI. The generation AI can be a text generation AI (e.g., LLM) or a multimodal generation AI. For example, if the purchaser is in Japan, the generation unit generates authentication words related to Japanese place names and culture. For example, it selects words such as "Mount Fuji" and "sushi." If the purchaser is in the United States, the generation unit can also generate authentication words related to American place names and culture. For example, it selects words such as "New York" and "hamburger." If the purchaser is in Europe, the generation unit can also generate authentication words related to European place names and culture. For example, it selects words such as "Eiffel Tower" and "pizza." This generates region-specific authentication words, making authentication easier.
[0039] When generating authentication words, the generation unit can generate related authentication words by referencing the purchaser's past purchase history. For example, the generation unit generates related authentication words by referencing the purchaser's past purchase history. The generation unit can also generate authentication words by referencing the purchaser's past purchase history using generation AI. The generation AI can be a text generation AI (e.g., LLM) or a multimodal generation AI. For example, the generation unit generates authentication words related to products the purchaser has previously purchased. For example, if the purchaser has previously purchased "apple pie," the generation unit selects words such as "apple" or "pie." The generation unit can also generate authentication words related to stores the purchaser has previously visited. For example, if the purchaser has previously visited a "cafe," the generation unit selects words such as "coffee" or "latte." The generation unit can also generate authentication words related to specific brands or products based on the purchaser's past purchase history. For example, if the purchaser has previously purchased "Nike" products, the generation unit selects words such as "shoes" or "running." This facilitates authentication because authentication words are generated based on the purchase history.
[0040] When generating authentication words, the generator can customize the authentication words based on the purchaser's age and gender. For example, the generator customizes the authentication words based on the purchaser's age and gender. The generator can also customize the authentication words based on age and gender using generation AI. The generation AI can be a text generation AI (e.g., LLM) or a multimodal generation AI. For example, if the purchaser is young, the generator generates words popular among young people as authentication words. For example, it selects words such as "Instagram" or "TikTok." If the purchaser is elderly, the generator can also generate simple, easy-to-remember words as authentication words. For example, it selects words such as "flower" or "bird." The generator can also generate related words as authentication words based on the purchaser's gender. For example, it selects words such as "flower" or "dress" for women and "sports" or "car" for men. This facilitates authentication by generating authentication words according to age and gender.
[0041] When transmitting the authentication word, the transmission unit can adjust the tone and speed of the voice based on the hearing characteristics of the purchaser. For example, the transmission unit adjusts the tone and speed of the voice based on the hearing characteristics of the purchaser. The transmission unit can also adjust the tone and speed of the voice based on the hearing characteristics using generation AI. The generation AI can be a text generation AI (e.g., LLM) or a multimodal generation AI. For example, if the purchaser is elderly and has poor hearing, the transmission unit can transmit the authentication word at a slower speed. For example, it can pronounce "apple pie" slowly. If the purchaser is young and has good hearing, the transmission unit can also transmit the authentication word at a normal speed. For example, it can pronounce "apple pie" at a normal speed. If the purchaser has hearing impairment, the transmission unit can also adjust the tone of the voice to transmit the authentication word. For example, if the purchaser can easily hear low-pitched sounds, it can pronounce "apple pie" at a lower pitch. This makes authentication easier because the voice is adjusted according to the hearing characteristics.
[0042] When transmitting the authentication word, the transmission unit can select the optimal transmission method by referring to the purchaser's past responses. For example, the transmission unit selects the optimal transmission method by referring to the purchaser's past responses. The transmission unit can also select the transmission method by referring to past responses using a generation AI. The generation AI can be a text generation AI (e.g., LLM) or a multimodal generation AI. For example, if the purchaser has successfully transmitted the authentication word by voice in the past, the transmission unit transmits the authentication word by the same method. For example, if the purchaser previously transmitted the authentication word by voice, the transmission unit will transmit it by voice this time as well. If the purchaser previously transmitted the authentication word visually, the transmission unit can also transmit the authentication word visually. For example, if the purchaser previously transmitted the authentication word by text, the transmission unit will transmit it by text this time as well. The transmission unit can also select the most effective transmission method based on the purchaser's past responses. For example, if the purchaser previously transmitted the authentication word by both voice and text, the transmission unit will transmit it by both methods this time as well. This makes authentication easier because the transmission method is selected based on past responses.
[0043] When transmitting the authentication word, the transmission unit can adjust the clarity of the voice taking into account the sound in the purchaser's environment. For example, the transmission unit adjusts the clarity of the voice taking into account the sound in the purchaser's environment. The transmission unit can also adjust the clarity of the voice taking into account the sound in the purchaser's environment using a generation AI. The generation AI can be a text generation AI (e.g., LLM) or a multimodal generation AI. For example, if the purchaser is in a quiet environment, the transmission unit transmits the authentication word in a normal voice. For example, it pronounces "apple pie" at a normal volume. If the purchaser is in a noisy environment, the transmission unit can also increase the clarity of the voice to transmit the authentication word. For example, it pronounces "apple pie" in a clear voice. If the purchaser is listening to music, the transmission unit can also adjust the clarity of the voice to match the volume of the music. For example, if the music volume is loud, it pronounces "apple pie" at a loud volume. This makes authentication easier because the voice is adjusted according to the sound in the environment.
[0044] When transmitting the authentication word, the transmission unit can select the optimal transmission means by taking into consideration the device information of the purchaser. For example, the transmission unit selects the optimal transmission means by taking into consideration the device information of the purchaser. The transmission unit can also select the transmission means by taking into consideration the device information using a generation AI. The generation AI can be a text generation AI (e.g., LLM) or a multimodal generation AI. For example, if the purchaser is using a smartphone, the transmission unit can transmit the authentication word by voice. For example, it can transmit "apple pie" by voice. If the purchaser is using a tablet, the transmission unit can also transmit the authentication word by text. For example, it can display "apple pie" in text. If the purchaser is using a smartwatch, the transmission unit can also transmit the authentication word by vibration or voice. For example, it can transmit "apple pie" by vibration. This makes authentication easier because the transmission means is selected according to the device information.
[0045] When transmitting the authentication word, the transmission unit can provide multilingual audio according to the language setting of the purchaser. For example, the transmission unit provides multilingual audio based on the language setting of the purchaser's device. The transmission unit can also provide multilingual audio based on the language setting using a generation AI. The generation AI can be a text generation AI (e.g., LLM) or a multimodal generation AI. For example, if the purchaser speaks English, the transmission unit transmits the authentication word in English. For example, it pronounces "apple pie" in English. If the purchaser speaks Japanese, the transmission unit can also transmit the authentication word in Japanese. For example, it pronounces "apple pie" in Japanese. If the purchaser speaks multiple languages, the transmission unit can provide a language switching function and transmit the authentication word in the selected language. For example, it can provide a function that allows switching between English and Japanese. This facilitates authentication by providing multilingual audio according to the language setting.
[0046] When transmitting the authentication word, the transmission unit can adjust the volume of the voice according to the degree of the purchaser's visual impairment. For example, the transmission unit adjusts the volume of the voice according to the degree of the purchaser's visual impairment. The transmission unit can also adjust the volume of the voice according to the degree of the visual impairment using a generation AI. The generation AI can be a text generation AI (e.g., LLM) or a multimodal generation AI. For example, if the purchaser has a mild visual impairment, the transmission unit can transmit the authentication word at a normal volume. For example, it can pronounce "apple pie" at a normal volume. If the purchaser has a moderate visual impairment, the transmission unit can also transmit the authentication word at a slightly higher volume. For example, it can pronounce "apple pie" at a slightly louder volume. If the purchaser has a severe visual impairment, the transmission unit can also transmit the authentication word at a higher volume. For example, it can pronounce "apple pie" at a louder volume. This makes authentication easier because the voice volume is adjusted according to the degree of visual impairment.
[0047] The voiceprint authentication unit can improve the accuracy of voiceprint authentication by referring to the purchaser's past voiceprint data during voiceprint authentication. For example, the voiceprint authentication unit improves authentication accuracy by referring to the purchaser's past voiceprint data. The voiceprint authentication unit can also improve authentication accuracy by using generation AI to refer to past voiceprint data. The generation AI can be a text generation AI (e.g., LLM) or a multimodal generation AI. For example, the voiceprint authentication unit refers to the purchaser's past voiceprint data and compares it with current voiceprint data. The voiceprint authentication unit can also extract specific patterns from the purchaser's past voiceprint data to improve authentication accuracy. For example, it can analyze the pronunciation patterns of specific sounds. The voiceprint authentication unit can also optimize the authentication algorithm based on the purchaser's past voiceprint data. For example, it can adjust the algorithm by learning from past data. This improves authentication accuracy based on past voiceprint data, thereby improving the success rate of authentication.
[0048] The voiceprint authentication unit can optimize the authentication algorithm during voiceprint authentication by taking into account fluctuations in the purchaser's voice. For example, the voiceprint authentication unit optimizes the authentication algorithm by taking into account fluctuations in the purchaser's voice. The voiceprint authentication unit can also optimize the authentication algorithm by using generation AI to take into account voice fluctuations. The generation AI can be a text generation AI (e.g., LLM) or a multimodal generation AI. For example, if the purchaser's voice fluctuates due to a cold or other condition, the voiceprint authentication unit adjusts the authentication algorithm by taking into account such fluctuations. For example, fluctuations in voice pitch and tone can be tolerated. The voiceprint authentication unit can also adjust the authentication algorithm by taking into account such fluctuations if the purchaser's voice fluctuates due to emotional state. For example, voice fluctuations due to nervousness or excitement can be tolerated. The voiceprint authentication unit can also adjust the authentication algorithm by taking into account such fluctuations if the purchaser's voice fluctuates due to age or health condition. For example, voice changes due to age can be tolerated. This improves the success rate of authentication by using an authentication algorithm that takes voice fluctuations into account.
[0049] The voiceprint authentication unit can improve the accuracy of voiceprint authentication by filtering the purchaser's environmental sounds during voiceprint authentication. For example, the voiceprint authentication unit improves the accuracy of authentication by filtering the purchaser's environmental sounds. The voiceprint authentication unit can also improve the accuracy of authentication by filtering the environmental sounds using generation AI. The generation AI can be a text generation AI (e.g., LLM) or a multimodal generation AI. For example, if the purchaser is in a noisy environment, the voiceprint authentication unit filters the environmental sounds to perform voiceprint authentication. For example, background noise is removed. If the purchaser is in a quiet environment, the voiceprint authentication unit can also perform voiceprint authentication using a normal method. For example, the voiceprint authentication unit can prevent the purchaser from being affected by environmental sounds. If the purchaser is listening to music, the voiceprint authentication unit can also filter the music to perform voiceprint authentication. For example, the frequency band of the music is removed. This filters the environmental sounds and improves the accuracy of authentication.
[0050] The voiceprint authentication unit can improve the accuracy of voiceprint authentication by taking into account the geographic location information of the purchaser. For example, the voiceprint authentication unit improves the accuracy of authentication by taking into account the geographic location information of the purchaser. The voiceprint authentication unit can also improve the accuracy of authentication by taking into account the geographic location information using generation AI. The generation AI can be a text generation AI (e.g., LLM) or a multimodal generation AI. For example, if the purchaser is in a specific region, the voiceprint authentication unit performs voiceprint authentication taking into account the characteristics of that region. For example, it can take into account accents unique to that region. If the purchaser is traveling, the voiceprint authentication unit can also adjust the accuracy of authentication based on the geographic location information. For example, it can filter out environmental sounds at the travel destination. If the purchaser is at home, the voiceprint authentication unit can also perform voiceprint authentication in the usual way. For example, it can perform authentication assuming a quiet environment at home. This improves the accuracy of authentication by taking into account the geographic location information.
[0051] The voiceprint authentication unit can improve the reliability of authentication by referring to the purchaser's past purchase history during voiceprint authentication. For example, the voiceprint authentication unit improves the reliability of authentication by referring to the purchaser's past purchase history. The voiceprint authentication unit can also improve the reliability of authentication by using generation AI to refer to the past purchase history. The generation AI can be a text generation AI (e.g., LLM) or a multimodal generation AI. For example, the voiceprint authentication unit improves the reliability of authentication by referring to the purchaser's past purchase history. For example, it can check the history of past purchases at the same store. The voiceprint authentication unit can also extract specific patterns from the purchaser's past purchase history to improve the reliability of authentication. For example, it can analyze patterns of frequent purchases of specific products. The voiceprint authentication unit can also optimize the authentication algorithm based on the purchaser's past purchase history. For example, it can adjust the algorithm by learning from past data. This improves the reliability of authentication based on the purchase history, thereby improving the success rate of authentication.
[0052] The voiceprint authentication unit can customize the authentication algorithm based on the age and gender of the purchaser during voiceprint authentication. The voiceprint authentication unit customizes the authentication algorithm based on, for example, the age and gender of the purchaser. The voiceprint authentication unit can also customize the authentication algorithm based on age and gender using generation AI. The generation AI can be a text generation AI (e.g., LLM) or a multimodal generation AI. For example, if the purchaser is young, the voiceprint authentication unit customizes the authentication algorithm by taking into account voice characteristics unique to young people. For example, it can emphasize high-pitched voices. If the purchaser is elderly, the voiceprint authentication unit can also customize the authentication algorithm by taking into account voice characteristics unique to elderly people. For example, it can emphasize low-pitched voices. The voiceprint authentication unit can also customize the authentication algorithm by taking into account related voice characteristics based on the gender of the purchaser. For example, it can emphasize high-pitched voices for women and low-pitched voices for men. This improves the success rate of authentication by using an authentication algorithm that is tailored to age and gender.
[0053] At the time of payment, the payment department can select the optimal payment method by referring to the purchaser's past payment history. For example, the payment department selects the optimal payment method by referring to the purchaser's past payment history. The payment department can also select a payment method by referring to the past payment history using generation AI. The generation AI can be a text generation AI (e.g., LLM) or a multimodal generation AI. For example, the payment department can select the most frequently used payment method by referring to the purchaser's past payment history. For example, if a credit card has been used in the past, a credit card will be suggested. The payment department can also select the optimal payment method for a specific time period or situation based on the purchaser's past payment history. For example, it can suggest a payment method that was previously used overnight. The payment department can also evaluate the reliability of payment methods based on the purchaser's past payment history and select the optimal method. For example, it can suggest a payment method with a high success rate in the past. This makes payment easier because the optimal payment method is selected based on the purchaser's past payment history.
[0054] The payment unit can customize the payment method based on the buyer's current purchasing situation at the time of payment. For example, the payment unit customizes the payment method based on the buyer's current purchasing situation. The payment unit can also customize the payment method based on the current purchasing situation using generation AI. The generation AI can be a text generation AI (e.g., LLM) or a multimodal generation AI. For example, if the buyer is purchasing a large number of items, the payment unit can provide a quick payment method. For example, the payment can be completed with voice authentication alone. If the buyer is purchasing a small number of items, the payment unit can also provide a payment method that includes detailed confirmation procedures. For example, the payment can be made after displaying a confirmation screen. If the buyer is purchasing from a specific store, the payment unit can also provide a payment method specialized for that store. For example, the payment unit can suggest payment using the store's point card. This makes payment easier because a payment method is selected according to the current purchasing situation.
[0055] The payment department can improve the payment method by reflecting buyer feedback at the time of payment. For example, the payment department improves the payment method by reflecting buyer feedback. The payment department can also improve the payment method by using generation AI to reflect feedback. The generation AI can be a text generation AI (e.g., LLM) or a multimodal generation AI. For example, the payment department improves the payment method based on feedback provided by the buyer in the past. For example, it resolves issues pointed out in the feedback. The payment department can also reflect feedback provided by the buyer after payment in real time to improve the next payment method. For example, it adds features suggested in the feedback. The payment department can also analyze buyer feedback and implement improvements to resolve common issues. For example, it resolves issues pointed out by multiple buyers. This improves the payment method based on feedback, making payment easier.
[0056] The payment unit can select the optimal payment method by taking into account the geographic location information of the purchaser at the time of payment. For example, the payment unit selects the optimal payment method by taking into account the geographic location information of the purchaser. The payment unit can also select the payment method by taking into account the geographic location information using generation AI. The generation AI can be a text generation AI (e.g., LLM) or a multimodal generation AI. For example, if the purchaser is in a specific region, the payment unit can suggest payment methods commonly used in that region. For example, if the purchaser is in Japan, the payment unit can suggest "Suica (registered trademark)" or "PayPay (registered trademark)." If the purchaser is traveling, the payment unit can also suggest payment methods available at the purchaser's destination. For example, if the purchaser is in the United States, the payment unit can suggest "Apple Pay" or "Google Pay (registered trademark)." If the purchaser is at home, the payment unit can also suggest payment methods that the purchaser normally uses. For example, the payment unit can suggest credit cards that the purchaser has previously used at home. In this way, the optimal payment method is selected by taking into account the geographic location information.
[0057] At the time of payment, the payment unit can analyze the buyer's social media activity to suggest payment methods. For example, the payment unit analyzes the buyer's social media activity to suggest payment methods. The payment unit can also use generation AI to analyze social media activity to suggest payment methods. The generation AI can be a text generation AI (e.g., LLM) or a multimodal generation AI. For example, the payment unit can suggest relevant payment methods based on the location where the buyer checked in on social media. For example, if the buyer checked in at a specific cafe, the payment unit can suggest payment methods available at that cafe. The payment unit can also analyze the content of the buyer's social media posts to suggest relevant payment methods. For example, if the buyer posted about a specific product, the payment unit can suggest payment methods for purchasing that product. The payment unit can also suggest relevant payment methods based on the activity of the buyer's friends on social media. For example, the payment unit can suggest payment methods used by the friends. This makes payment easier by suggesting the optimal payment method based on social media activity.
[0058] The payment department can customize the payment method by reflecting the buyer's past feedback at the time of payment. For example, the payment department customizes the payment method by reflecting the buyer's past feedback. The payment department can also customize the payment method by using generation AI to reflect past feedback. The generation AI can be a text generation AI (e.g., LLM) or a multimodal generation AI. For example, the payment department customizes the payment method based on feedback provided by the buyer in the past. For example, it resolves issues pointed out in the feedback. The payment department can also reflect feedback provided by the buyer after payment in real time to customize the next payment method. For example, it adds features suggested in the feedback. The payment department can also analyze buyer feedback and perform customization to resolve common issues. For example, it resolves issues pointed out by multiple buyers. This makes payment easier because the payment method is customized based on past feedback.
[0059] The system according to the embodiment is not limited to the above-described example, and various modifications are possible, for example, as follows.
[0060] The payment system may further include a purchasing pattern analysis unit that analyzes a user's purchasing patterns. The purchasing pattern analysis unit analyzes the user's past purchasing history and understands the user's purchasing trends. For example, if the user tends to purchase specific products on specific days of the week or at specific times of the day, the purchasing pattern analysis unit provides that information to the payment unit, which can then suggest the optimal payment method based on that trend. If the user frequently purchases products from a specific brand or category, the purchasing pattern analysis unit provides that information to the generation unit, which can then generate authentication keywords related to that brand or category. Furthermore, the purchasing pattern analysis unit may also provide benefits or discounts based on the user's purchasing patterns. This allows for personalized services to be provided in accordance with the user's purchasing behavior.
[0061] The payment system may further include a location information utilization unit that utilizes the user's location information. The location information utilization unit obtains the user's current location information and provides that information to other elements. For example, if the user is in a specific store, the location information utilization unit provides that information to a payment unit, which can suggest payment methods available at that store. If the user is traveling, the location information utilization unit provides that information to a generation unit, which can generate authentication words related to the language and culture of the destination. Furthermore, if the user is in a specific region, the location information utilization unit can provide information about special offers and discounts for that region. This enables flexible responses based on the user's location information and provides a more convenient payment experience.
[0062] The payment system may further include a voice command analysis unit that analyzes a user's voice command. The voice command analysis unit analyzes the voice command issued by the user and performs appropriate processing based on the content of the command. For example, if the user utters "check balance," the voice command analysis unit analyzes the command and provides balance information to the providing unit. The providing unit can convey the balance information to the user by voice. If the user utters "cancel," the voice command analysis unit can analyze the command and send a cancellation instruction to the payment unit. Furthermore, the voice command analysis unit can automate certain operations based on the user's voice command. This enables flexible responses based on the user's voice command, providing a more intuitive and convenient payment experience.
[0063] The payment system may further include a behavior prediction unit that predicts user behavior. The behavior prediction unit analyzes the user's past behavior data and predicts future behavior. For example, if the user tends to visit a particular store on a particular day or time, the behavior prediction unit provides that information to the payment unit, which can then prepare payment methods available at that store in advance. If the user plans to attend a particular event, the behavior prediction unit provides that information to the generation unit, which can then generate authentication words related to the event. Furthermore, the behavior prediction unit may provide benefits or discounts based on the user's behavioral patterns. This allows for personalized services to be provided according to the user's behavior.
[0064] The payment system can further include an incentive providing unit to increase the user's purchasing motivation. The incentive providing unit analyzes the user's purchasing history and behavioral patterns and provides optimal incentives. For example, if the user frequently purchases a particular product, the incentive providing unit can provide a discount coupon for that product. If the user frequently visits a particular store, the incentive providing unit can provide points that can be used at that store. Furthermore, the incentive providing unit can also provide limited-time benefits and campaigns to increase the user's purchasing motivation. This can increase the user's purchasing motivation and encourage more purchasing behavior.
[0065] The processing flow of the first embodiment will be briefly explained below.
[0066] Step 1: The generator generates authentication words. The generator randomly generates authentication words using, for example, a generation AI. The generation AI may be a text generation AI (e.g., LLM) or a multimodal generation AI. The generator can generate words such as "apple pie" or "sunflower." The generator can also generate authentication words based on the buyer's emotions and past authentication history. For example, if the buyer is nervous, it can generate a simple word. Step 2: The communication unit communicates the generated authentication word to the purchaser. The communication unit communicates the authentication word aloud, for example, using voice synthesis technology. The communication unit can also adjust the tone and speed of the voice based on the hearing characteristics and emotions of the purchaser. For example, the authentication word can be communicated at a slower speed to an elderly person. Step 3: The voiceprint authentication unit analyzes the customer's voice and compares it with pre-registered voiceprint data. For example, the voiceprint authentication unit uses voice signal processing technology to analyze the customer's voice and extract features. The voiceprint authentication unit can also optimize the authentication algorithm by taking into account the customer's emotions and voice fluctuations. For example, if the customer has a cold, authentication is performed taking these fluctuations into account. Step 4: The payment unit executes payment if the voiceprint authentication is successful. The payment unit performs payment using means such as credit card payment or electronic money payment. The payment unit can also select the optimal payment method based on the buyer's emotions and past payment history. For example, if the buyer is in a hurry, it can provide a quick payment method.
[0067] (Example 2) A payment system according to an embodiment of the present invention enables voice-based payments for those who cannot use smartphones, such as those with visual impairments or the elderly, or those in situations where QR code payments are difficult to use. In this payment system, a generation AI generates an authentication word at checkout, and when the customer speaks the authentication word, voiceprint authentication is performed on the customer's voice to complete the payment. For example, in this payment system, a generation AI randomly generates an authentication word at checkout and communicates it to the customer. When the customer speaks the generated authentication word, the customer's voice is recorded and voiceprint authentication is performed. Voiceprint authentication analyzes the customer's voice and compares it with pre-registered voiceprint data. If the comparison is successful, the payment is completed. This allows for easy payments even for those who cannot use smartphones, such as those without smartphones, visual impairments, or the elderly. For example, visually impaired customers can complete payments using only their voice when purchasing items in a store, without using a smartphone. Furthermore, elderly customers who are unfamiliar with smartphones can easily make payments using only their voice. Furthermore, payments can be made using only their voice, even when their hands are full or they do not have a smartphone. This allows for convenient payment access for a wider range of people. This makes the payment system easy to use, even for those who have difficulty using QR payments, or those who are unable to use smartphones, such as the visually impaired or elderly. For example, when purchasing goods in a store, a visually impaired person can complete the payment using only their voice, without using a smartphone. Also, even elderly people who are unfamiliar with using smartphones can easily make payments using only their voice. Furthermore, even if their hands are full or they do not have a smartphone, they can make payments using only their voice. This makes payment convenient for more people.
[0068] The payment system according to the embodiment includes a generation unit, a transmission unit, a voiceprint authentication unit, and a payment unit. The generation unit generates an authentication word. The generation unit randomly generates the authentication word using, for example, a generation AI. The generation AI may be a text generation AI (e.g., LLM) or a multimodal generation AI. The generation unit may generate words such as "apple pie" or "sunflower." The generation unit may also generate the authentication word based on the customer's emotions and past authentication history. For example, if the customer is nervous, it may generate a simple word. The transmission unit communicates the generated authentication word to the customer. The transmission unit may communicate the authentication word aloud using, for example, speech synthesis technology. The transmission unit may also adjust the tone and speed of the voice based on the customer's hearing characteristics and emotions. For example, the authentication word may be communicated at a slower speed to an elderly customer. The voiceprint authentication unit analyzes the customer's voice and compares it with pre-registered voiceprint data. The voiceprint authentication unit may analyze the customer's voice using, for example, speech signal processing technology, and extract features. The voiceprint authentication unit can also optimize the authentication algorithm by taking into account the buyer's emotions and voice fluctuations. For example, if the buyer has a cold, authentication is performed taking these fluctuations into account. The payment unit performs payment if voiceprint authentication is successful. The payment unit performs payment using means such as credit card payment or electronic money payment. The payment unit can also select the optimal payment method based on the buyer's emotions and past payment history. For example, if the buyer is in a hurry, a quick payment method can be provided. As a result, the payment system according to the embodiment allows anyone to easily make payments in any situation.
[0069] The generation unit randomly generates an authentication word. The generation unit randomly generates an authentication word using, for example, a random number generation algorithm. The generation unit can also randomly generate an authentication word using a generation AI. The generation AI can be a text generation AI (e.g., LLM) or a multimodal generation AI. For example, the generation unit randomly generates words such as "apple pie" or "sunflower." The generation unit can also generate an authentication word based on the buyer's emotions or past authentication history. For example, if the buyer is nervous, a simple word can be generated. This improves security by randomly generating authentication words.
[0070] The communication unit communicates the generated authentication word to the purchaser by voice. The communication unit communicates the authentication word by voice, for example, using speech synthesis technology. The communication unit can also communicate the authentication word by voice using generation AI. The generation AI can be a text generation AI (e.g., LLM) or a multimodal generation AI. For example, the communication unit communicates words such as "apple pie" or "sunflower" by voice. The communication unit can also adjust the tone and speed of the voice based on the purchaser's hearing characteristics and emotions. For example, for elderly people, the authentication word can be communicated at a slower speed. As a result, communicating the authentication word by voice makes it easier for visually impaired people and elderly people to use.
[0071] The voiceprint authentication unit analyzes the purchaser's voice and compares it with pre-registered voiceprint data. The voiceprint authentication unit may, for example, use voice signal processing technology to analyze the purchaser's voice and extract features. The voiceprint authentication unit may also analyze the purchaser's voice using generation AI. The generation AI may be a text generation AI (e.g., LLM) or a multimodal generation AI. For example, the voiceprint authentication unit may record the purchaser's voice and compare it with pre-registered voiceprint data. The voiceprint authentication unit may also optimize the authentication algorithm by taking into account the purchaser's emotions and voice fluctuations. For example, if the purchaser has a cold, authentication may take these fluctuations into account. This ensures that the purchaser's identity is confirmed through voiceprint authentication.
[0072] The payment unit executes payment if the voiceprint authentication is successful. The payment unit executes payment using means such as credit card payment or electronic money payment. The payment unit can also execute payment using generation AI. The generation AI can be a text generation AI (e.g., LLM) or a multimodal generation AI. For example, the payment unit executes payment if the voiceprint authentication is successful. The payment unit can also select the optimal payment method based on the buyer's emotions and past payment history. For example, if the buyer is in a hurry, a quick payment method can be provided. This improves security because payment is executed only if the voiceprint authentication is successful.
[0073] The payment unit can provide a payment method that can be used by people who cannot use smartphones, such as those who are unable to own a smartphone, those who are visually impaired, or the elderly. The payment unit can also provide a payment method using generation AI. Generation AI includes text generation AI (e.g., LLM) and multimodal generation AI. For example, the payment unit provides a means for visually impaired people to complete payment using only their voice when purchasing items in a store, without using a smartphone. The payment unit can provide a means for elderly people who are unfamiliar with using smartphones to easily make payments using only their voice. This allows people who cannot use smartphones to make payments easily.
[0074] The generator can estimate the buyer's emotions and adjust the difficulty of the authentication words based on the estimated emotions. The generator can estimate the buyer's emotions using, for example, voice analysis technology. The generator can also estimate the buyer's emotions using generation AI. The generation AI can be a text generation AI (e.g., LLM) or a multimodal generation AI. For example, if the buyer is nervous, the generator can generate simple words as authentication words. For example, it can select short words such as "cat" or "dog." If the buyer is relaxed, the generator can generate slightly more difficult words as authentication words. For example, it can select longer words such as "elephant" or "sunflower." If the buyer is in a hurry, the generator can generate short, easy-to-remember words as authentication words. For example, it can select simple words such as "apple" or "banana." This adjusts the difficulty of the authentication words according to the buyer's emotions, thereby improving the success rate of authentication.
[0075] The generation unit can analyze the purchaser's past authentication history and generate optimal authentication words. For example, the generation unit analyzes the purchaser's past authentication history and generates optimal authentication words. The generation unit can also analyze past authentication history using generation AI. The generation AI is a text generation AI (e.g., LLM) or multimodal generation AI. For example, the generation unit analyzes patterns of authentication words that the purchaser has used successfully in the past and generates authentication words with similar patterns. The generation unit can also avoid authentication words that the purchaser has used unsuccessfully in the past and generate authentication words with a high success rate. The generation unit can also generate optimal authentication words for specific time periods or situations from the purchaser's past authentication history. This generates optimal authentication words based on the past authentication history, improving the authentication success rate.
[0076] When generating authentication words, the generation unit can generate authentication words in different languages based on the language settings of the purchaser. For example, the generation unit generates authentication words in different languages based on the language settings of the purchaser's device. The generation unit can also generate authentication words based on the language settings using generation AI. The generation AI can be a text generation AI (e.g., LLM) or a multimodal generation AI. For example, if the purchaser speaks English, the generation unit generates authentication words in English. For example, it selects words such as "apple" or "banana." If the purchaser speaks Japanese, the generation unit can also generate authentication words in Japanese. For example, it selects words such as "ringo" or "banana." If the purchaser speaks multiple languages, the generation unit can also randomly generate authentication words from multiple languages. For example, it selects words in English and Japanese alternately. This facilitates authentication because authentication words are generated according to the purchaser's language settings.
[0077] When generating authentication words, the generation unit can generate the authentication words based on the purchaser's current environmental sounds. For example, the generation unit generates authentication words taking into account the purchaser's current environmental sounds. The generation unit can also generate authentication words taking into account environmental sounds using generation AI. The generation AI can be a text generation AI (e.g., LLM) or a multimodal generation AI. For example, if the purchaser is in a quiet environment, the generation unit generates complex authentication words. For example, it selects long words such as "elephant" or "sunflower." If the purchaser is in a noisy environment, the generation unit can also generate simple authentication words. For example, it selects short words such as "cat" or "dog." If the purchaser is listening to music, the generation unit can also generate authentication words related to the music genre. For example, if the purchaser is listening to rock music, it selects words such as "guitar" or "drums." This makes authentication easier because authentication words are generated according to the environmental sounds.
[0078] The generator can estimate the buyer's emotions and adjust the length of the authentication words based on the estimated emotions. The generator can estimate the buyer's emotions using, for example, voice analysis technology. The generator can also estimate the buyer's emotions using generation AI. The generation AI can be a text generation AI (e.g., LLM) or a multimodal generation AI. For example, if the buyer is nervous, the generator can generate short authentication words. For example, it can select short words such as "cat" or "dog." If the buyer is relaxed, the generator can also generate long authentication words. For example, it can select long words such as "elephant" or "sunflower." If the buyer is in a hurry, the generator can also generate short, easy-to-remember authentication words. For example, it can select simple words such as "apple" or "banana." This adjusts the length of the authentication words according to the buyer's emotions, thereby improving the success rate of authentication.
[0079] When generating authentication words, the generation unit can generate region-specific authentication words taking into account the geographic location information of the purchaser. For example, the generation unit generates region-specific authentication words taking into account the geographic location information of the purchaser. The generation unit can also generate authentication words taking into account the geographic location information using generation AI. The generation AI can be a text generation AI (e.g., LLM) or a multimodal generation AI. For example, if the purchaser is in Japan, the generation unit generates authentication words related to Japanese place names and culture. For example, it selects words such as "Mount Fuji" and "sushi." If the purchaser is in the United States, the generation unit can also generate authentication words related to American place names and culture. For example, it selects words such as "New York" and "hamburger." If the purchaser is in Europe, the generation unit can also generate authentication words related to European place names and culture. For example, it selects words such as "Eiffel Tower" and "pizza." This generates region-specific authentication words, making authentication easier.
[0080] When generating authentication words, the generation unit can generate related authentication words by referencing the purchaser's past purchase history. For example, the generation unit generates related authentication words by referencing the purchaser's past purchase history. The generation unit can also generate authentication words by referencing the purchaser's past purchase history using generation AI. The generation AI can be a text generation AI (e.g., LLM) or a multimodal generation AI. For example, the generation unit generates authentication words related to products the purchaser has previously purchased. For example, if the purchaser has previously purchased "apple pie," the generation unit selects words such as "apple" or "pie." The generation unit can also generate authentication words related to stores the purchaser has previously visited. For example, if the purchaser has previously visited a "cafe," the generation unit selects words such as "coffee" or "latte." The generation unit can also generate authentication words related to specific brands or products based on the purchaser's past purchase history. For example, if the purchaser has previously purchased "Nike" products, the generation unit selects words such as "shoes" or "running." This facilitates authentication because authentication words are generated based on the purchase history.
[0081] When generating authentication words, the generator can customize the authentication words based on the purchaser's age and gender. For example, the generator customizes the authentication words based on the purchaser's age and gender. The generator can also customize the authentication words based on age and gender using generation AI. The generation AI can be a text generation AI (e.g., LLM) or a multimodal generation AI. For example, if the purchaser is young, the generator generates words popular among young people as authentication words. For example, it selects words such as "Instagram" or "TikTok." If the purchaser is elderly, the generator can also generate simple, easy-to-remember words as authentication words. For example, it selects words such as "flower" or "bird." The generator can also generate related words as authentication words based on the purchaser's gender. For example, it selects words such as "flower" or "dress" for women and "sports" or "car" for men. This facilitates authentication by generating authentication words according to age and gender.
[0082] The communication unit can estimate the buyer's emotions and adjust the method of communicating the authentication word based on the estimated buyer's emotions. The communication unit can estimate the buyer's emotions using, for example, voice analysis technology. The communication unit can also estimate the buyer's emotions using generation AI. The generation AI can be a text generation AI (e.g., LLM) or a multimodal generation AI. For example, if the buyer is nervous, the communication unit can communicate the authentication word in a calm voice. For example, it can communicate "apple pie" in a gentle tone. If the buyer is relaxed, the communication unit can also communicate the authentication word in a cheerful voice. For example, it can communicate "sunflower" in a cheerful tone. If the buyer is in a hurry, the communication unit can also communicate the authentication word in a quick and concise manner. For example, it can select a short, easy-to-remember word and communicate it quickly. This allows the communication method to be selected according to the buyer's emotions, making authentication easier.
[0083] When transmitting the authentication word, the transmission unit can adjust the tone and speed of the voice based on the hearing characteristics of the purchaser. For example, the transmission unit adjusts the tone and speed of the voice based on the hearing characteristics of the purchaser. The transmission unit can also adjust the tone and speed of the voice based on the hearing characteristics using generation AI. The generation AI can be a text generation AI (e.g., LLM) or a multimodal generation AI. For example, if the purchaser is elderly and has poor hearing, the transmission unit can transmit the authentication word at a slower speed. For example, it can pronounce "apple pie" slowly. If the purchaser is young and has good hearing, the transmission unit can also transmit the authentication word at a normal speed. For example, it can pronounce "apple pie" at a normal speed. If the purchaser has hearing impairment, the transmission unit can also adjust the tone of the voice to transmit the authentication word. For example, if the purchaser can easily hear low-pitched sounds, it can pronounce "apple pie" at a lower pitch. This makes authentication easier because the voice is adjusted according to the hearing characteristics.
[0084] When transmitting the authentication word, the transmission unit can select the optimal transmission method by referring to the purchaser's past responses. For example, the transmission unit selects the optimal transmission method by referring to the purchaser's past responses. The transmission unit can also select the transmission method by referring to past responses using a generation AI. The generation AI can be a text generation AI (e.g., LLM) or a multimodal generation AI. For example, if the purchaser has successfully transmitted the authentication word by voice in the past, the transmission unit transmits the authentication word by the same method. For example, if the purchaser previously transmitted the authentication word by voice, the transmission unit will transmit it by voice this time as well. If the purchaser previously transmitted the authentication word visually, the transmission unit can also transmit the authentication word visually. For example, if the purchaser previously transmitted the authentication word by text, the transmission unit will transmit it by text this time as well. The transmission unit can also select the most effective transmission method based on the purchaser's past responses. For example, if the purchaser previously transmitted the authentication word by both voice and text, the transmission unit will transmit it by both methods this time as well. This makes authentication easier because the transmission method is selected based on past responses.
[0085] When transmitting the authentication word, the transmission unit can adjust the clarity of the voice taking into account the sound in the purchaser's environment. For example, the transmission unit adjusts the clarity of the voice taking into account the sound in the purchaser's environment. The transmission unit can also adjust the clarity of the voice taking into account the sound in the purchaser's environment using a generation AI. The generation AI can be a text generation AI (e.g., LLM) or a multimodal generation AI. For example, if the purchaser is in a quiet environment, the transmission unit transmits the authentication word in a normal voice. For example, it pronounces "apple pie" at a normal volume. If the purchaser is in a noisy environment, the transmission unit can also increase the clarity of the voice to transmit the authentication word. For example, it pronounces "apple pie" in a clear voice. If the purchaser is listening to music, the transmission unit can also adjust the clarity of the voice to match the volume of the music. For example, if the music volume is loud, it pronounces "apple pie" at a loud volume. This makes authentication easier because the voice is adjusted according to the sound in the environment.
[0086] The transmission unit can estimate the buyer's emotions and adjust the timing of transmitting the authentication word based on the estimated buyer's emotions. The transmission unit estimates the buyer's emotions using, for example, voice analysis technology. The transmission unit can also estimate the buyer's emotions using generation AI. The generation AI can be a text generation AI (for example, LLM) or a multimodal generation AI. For example, if the buyer is nervous, the transmission unit can transmit the authentication word at a slower timing. For example, it can pronounce "apple pie" slowly. If the buyer is relaxed, the transmission unit can also transmit the authentication word at a normal timing. For example, it can pronounce "apple pie" at a normal speed. If the buyer is in a hurry, the transmission unit can also transmit the authentication word at a quicker timing. For example, it can pronounce "apple pie" quickly. This allows the transmission timing to be selected according to the buyer's emotions, making authentication easier.
[0087] When transmitting the authentication word, the transmission unit can select the optimal transmission means by taking into consideration the device information of the purchaser. For example, the transmission unit selects the optimal transmission means by taking into consideration the device information of the purchaser. The transmission unit can also select the transmission means by taking into consideration the device information using a generation AI. The generation AI can be a text generation AI (e.g., LLM) or a multimodal generation AI. For example, if the purchaser is using a smartphone, the transmission unit can transmit the authentication word by voice. For example, it can transmit "apple pie" by voice. If the purchaser is using a tablet, the transmission unit can also transmit the authentication word by text. For example, it can display "apple pie" in text. If the purchaser is using a smartwatch, the transmission unit can also transmit the authentication word by vibration or voice. For example, it can transmit "apple pie" by vibration. This makes authentication easier because the transmission means is selected according to the device information.
[0088] When transmitting the authentication word, the transmission unit can provide multilingual audio according to the language setting of the purchaser. For example, the transmission unit provides multilingual audio based on the language setting of the purchaser's device. The transmission unit can also provide multilingual audio based on the language setting using a generation AI. The generation AI can be a text generation AI (e.g., LLM) or a multimodal generation AI. For example, if the purchaser speaks English, the transmission unit transmits the authentication word in English. For example, it pronounces "apple pie" in English. If the purchaser speaks Japanese, the transmission unit can also transmit the authentication word in Japanese. For example, it pronounces "apple pie" in Japanese. If the purchaser speaks multiple languages, the transmission unit can provide a language switching function and transmit the authentication word in the selected language. For example, it can provide a function that allows switching between English and Japanese. This facilitates authentication by providing multilingual audio according to the language setting.
[0089] When transmitting the authentication word, the transmission unit can adjust the volume of the voice according to the degree of the purchaser's visual impairment. For example, the transmission unit adjusts the volume of the voice according to the degree of the purchaser's visual impairment. The transmission unit can also adjust the volume of the voice according to the degree of the visual impairment using a generation AI. The generation AI can be a text generation AI (e.g., LLM) or a multimodal generation AI. For example, if the purchaser has a mild visual impairment, the transmission unit can transmit the authentication word at a normal volume. For example, it can pronounce "apple pie" at a normal volume. If the purchaser has a moderate visual impairment, the transmission unit can also transmit the authentication word at a slightly higher volume. For example, it can pronounce "apple pie" at a slightly louder volume. If the purchaser has a severe visual impairment, the transmission unit can also transmit the authentication word at a higher volume. For example, it can pronounce "apple pie" at a louder volume. This makes authentication easier because the voice volume is adjusted according to the degree of visual impairment.
[0090] The voiceprint authentication unit can estimate the buyer's emotions and adjust the accuracy of the voiceprint authentication based on the estimated emotions. The voiceprint authentication unit, for example, uses voice analysis technology to estimate the buyer's emotions. The voiceprint authentication unit can also estimate the buyer's emotions using generation AI. The generation AI can be a text generation AI (e.g., LLM) or a multimodal generation AI. For example, if the buyer is nervous, the voiceprint authentication unit references multiple voiceprint data to improve the accuracy of voiceprint authentication. For example, it compares past voiceprint data with current voiceprint data. If the buyer is relaxed, the voiceprint authentication unit can also perform voiceprint authentication with normal accuracy. For example, it references only current voiceprint data. If the buyer is in a hurry, the voiceprint authentication unit can also use a simplified algorithm to perform voiceprint authentication quickly. For example, it compares only key features. This adjusts the accuracy of voiceprint authentication according to the buyer's emotions, thereby improving the success rate of authentication.
[0091] The voiceprint authentication unit can improve the accuracy of voiceprint authentication by referring to the purchaser's past voiceprint data during voiceprint authentication. For example, the voiceprint authentication unit improves authentication accuracy by referring to the purchaser's past voiceprint data. The voiceprint authentication unit can also improve authentication accuracy by using generation AI to refer to past voiceprint data. The generation AI can be a text generation AI (e.g., LLM) or a multimodal generation AI. For example, the voiceprint authentication unit refers to the purchaser's past voiceprint data and compares it with current voiceprint data. The voiceprint authentication unit can also extract specific patterns from the purchaser's past voiceprint data to improve authentication accuracy. For example, it can analyze the pronunciation patterns of specific sounds. The voiceprint authentication unit can also optimize the authentication algorithm based on the purchaser's past voiceprint data. For example, it can adjust the algorithm by learning from past data. This improves authentication accuracy based on past voiceprint data, thereby improving the success rate of authentication.
[0092] The voiceprint authentication unit can optimize the authentication algorithm during voiceprint authentication by taking into account fluctuations in the purchaser's voice. For example, the voiceprint authentication unit optimizes the authentication algorithm by taking into account fluctuations in the purchaser's voice. The voiceprint authentication unit can also optimize the authentication algorithm by using generation AI to take into account voice fluctuations. The generation AI can be a text generation AI (e.g., LLM) or a multimodal generation AI. For example, if the purchaser's voice fluctuates due to a cold or other condition, the voiceprint authentication unit adjusts the authentication algorithm by taking into account such fluctuations. For example, fluctuations in voice pitch and tone can be tolerated. The voiceprint authentication unit can also adjust the authentication algorithm by taking into account such fluctuations if the purchaser's voice fluctuates due to emotional state. For example, voice fluctuations due to nervousness or excitement can be tolerated. The voiceprint authentication unit can also adjust the authentication algorithm by taking into account such fluctuations if the purchaser's voice fluctuates due to age or health condition. For example, voice changes due to age can be tolerated. This improves the success rate of authentication by using an authentication algorithm that takes voice fluctuations into account.
[0093] The voiceprint authentication unit can improve the accuracy of voiceprint authentication by filtering the purchaser's environmental sounds during voiceprint authentication. For example, the voiceprint authentication unit improves the accuracy of authentication by filtering the purchaser's environmental sounds. The voiceprint authentication unit can also improve the accuracy of authentication by filtering the environmental sounds using generation AI. The generation AI can be a text generation AI (e.g., LLM) or a multimodal generation AI. For example, if the purchaser is in a noisy environment, the voiceprint authentication unit filters the environmental sounds to perform voiceprint authentication. For example, background noise is removed. If the purchaser is in a quiet environment, the voiceprint authentication unit can also perform voiceprint authentication using a normal method. For example, the voiceprint authentication unit can prevent the purchaser from being affected by environmental sounds. If the purchaser is listening to music, the voiceprint authentication unit can also filter the music to perform voiceprint authentication. For example, the frequency band of the music is removed. This filters the environmental sounds and improves the accuracy of authentication.
[0094] The voiceprint authentication unit can estimate the buyer's emotions and adjust the speed of voiceprint authentication based on the estimated buyer's emotions. The voiceprint authentication unit estimates the buyer's emotions using, for example, voice analysis technology. The voiceprint authentication unit can also estimate the buyer's emotions using generation AI. The generation AI can be a text generation AI (e.g., LLM) or a multimodal generation AI. For example, if the buyer is nervous, the voiceprint authentication unit can slow down the speed of voiceprint authentication to improve accuracy. For example, detailed analysis can be performed. If the buyer is relaxed, the voiceprint authentication unit can also perform voiceprint authentication at a normal speed. For example, standard analysis can be performed. If the buyer is in a hurry, the voiceprint authentication unit can also speed up the speed of voiceprint authentication to quickly authenticate. For example, simplified analysis can be performed. This adjusts the speed of voiceprint authentication according to the buyer's emotions, thereby improving the success rate of authentication.
[0095] The voiceprint authentication unit can improve the accuracy of voiceprint authentication by taking into account the geographic location information of the purchaser. For example, the voiceprint authentication unit improves the accuracy of authentication by taking into account the geographic location information of the purchaser. The voiceprint authentication unit can also improve the accuracy of authentication by taking into account the geographic location information using generation AI. The generation AI can be a text generation AI (e.g., LLM) or a multimodal generation AI. For example, if the purchaser is in a specific region, the voiceprint authentication unit performs voiceprint authentication taking into account the characteristics of that region. For example, it can take into account accents unique to that region. If the purchaser is traveling, the voiceprint authentication unit can also adjust the accuracy of authentication based on the geographic location information. For example, it can filter out environmental sounds at the travel destination. If the purchaser is at home, the voiceprint authentication unit can also perform voiceprint authentication in the usual way. For example, it can perform authentication assuming a quiet environment at home. This improves the accuracy of authentication by taking into account the geographic location information.
[0096] The voiceprint authentication unit can improve the reliability of authentication by referring to the purchaser's past purchase history during voiceprint authentication. For example, the voiceprint authentication unit improves the reliability of authentication by referring to the purchaser's past purchase history. The voiceprint authentication unit can also improve the reliability of authentication by using generation AI to refer to the past purchase history. The generation AI can be a text generation AI (e.g., LLM) or a multimodal generation AI. For example, the voiceprint authentication unit improves the reliability of authentication by referring to the purchaser's past purchase history. For example, it can check the history of past purchases at the same store. The voiceprint authentication unit can also extract specific patterns from the purchaser's past purchase history to improve the reliability of authentication. For example, it can analyze patterns of frequent purchases of specific products. The voiceprint authentication unit can also optimize the authentication algorithm based on the purchaser's past purchase history. For example, it can adjust the algorithm by learning from past data. This improves the reliability of authentication based on the purchase history, thereby improving the success rate of authentication.
[0097] The voiceprint authentication unit can customize the authentication algorithm based on the age and gender of the purchaser during voiceprint authentication. The voiceprint authentication unit customizes the authentication algorithm based on, for example, the age and gender of the purchaser. The voiceprint authentication unit can also customize the authentication algorithm based on age and gender using generation AI. The generation AI can be a text generation AI (e.g., LLM) or a multimodal generation AI. For example, if the purchaser is young, the voiceprint authentication unit customizes the authentication algorithm by taking into account voice characteristics unique to young people. For example, it can emphasize high-pitched voices. If the purchaser is elderly, the voiceprint authentication unit can also customize the authentication algorithm by taking into account voice characteristics unique to elderly people. For example, it can emphasize low-pitched voices. The voiceprint authentication unit can also customize the authentication algorithm by taking into account related voice characteristics based on the gender of the purchaser. For example, it can emphasize high-pitched voices for women and low-pitched voices for men. This improves the success rate of authentication by using an authentication algorithm that is tailored to age and gender.
[0098] The payment unit can estimate the buyer's emotions and adjust the payment method based on the estimated buyer's emotions. The payment unit can estimate the buyer's emotions using, for example, voice analysis technology. The payment unit can also estimate the buyer's emotions using generation AI. The generation AI can be a text generation AI (e.g., LLM) or a multimodal generation AI. For example, if the buyer is nervous, the payment unit can provide a simple and quick payment method, such as completing the payment with one tap. If the buyer is relaxed, the payment unit can provide a payment method that includes detailed confirmation steps, such as displaying a confirmation screen before making the payment. If the buyer is in a hurry, the payment unit can also provide a quick payment method, such as completing the payment with voice authentication only. This allows the payment method to be selected according to the buyer's emotions, making payment easier.
[0099] At the time of payment, the payment department can select the optimal payment method by referring to the purchaser's past payment history. For example, the payment department selects the optimal payment method by referring to the purchaser's past payment history. The payment department can also select a payment method by referring to the past payment history using generation AI. The generation AI can be a text generation AI (e.g., LLM) or a multimodal generation AI. For example, the payment department can select the most frequently used payment method by referring to the purchaser's past payment history. For example, if a credit card has been used in the past, a credit card will be suggested. The payment department can also select the optimal payment method for a specific time period or situation based on the purchaser's past payment history. For example, it can suggest a payment method that was previously used overnight. The payment department can also evaluate the reliability of payment methods based on the purchaser's past payment history and select the optimal method. For example, it can suggest a payment method with a high success rate in the past. This makes payment easier because the optimal payment method is selected based on the purchaser's past payment history.
[0100] The payment unit can customize the payment method based on the buyer's current purchasing situation at the time of payment. For example, the payment unit customizes the payment method based on the buyer's current purchasing situation. The payment unit can also customize the payment method based on the current purchasing situation using generation AI. The generation AI can be a text generation AI (e.g., LLM) or a multimodal generation AI. For example, if the buyer is purchasing a large number of items, the payment unit can provide a quick payment method. For example, the payment can be completed with voice authentication alone. If the buyer is purchasing a small number of items, the payment unit can also provide a payment method that includes detailed confirmation procedures. For example, the payment can be made after displaying a confirmation screen. If the buyer is purchasing from a specific store, the payment unit can also provide a payment method specialized for that store. For example, the payment unit can suggest payment using the store's point card. This makes payment easier because a payment method is selected according to the current purchasing situation.
[0101] The payment department can improve the payment method by reflecting buyer feedback at the time of payment. For example, the payment department improves the payment method by reflecting buyer feedback. The payment department can also improve the payment method by using generation AI to reflect feedback. The generation AI can be a text generation AI (e.g., LLM) or a multimodal generation AI. For example, the payment department improves the payment method based on feedback provided by the buyer in the past. For example, it resolves issues pointed out in the feedback. The payment department can also reflect feedback provided by the buyer after payment in real time to improve the next payment method. For example, it adds features suggested in the feedback. The payment department can also analyze buyer feedback and implement improvements to resolve common issues. For example, it resolves issues pointed out by multiple buyers. This improves the payment method based on feedback, making payment easier.
[0102] The payment unit can estimate the buyer's emotions and determine the priority of the payment based on the estimated buyer's emotions. The payment unit, for example, estimates the buyer's emotions using voice analysis technology. The payment unit can also estimate the buyer's emotions using generation AI. The generation AI can be a text generation AI (e.g., LLM) or a multimodal generation AI. For example, if the buyer is nervous, the payment unit can set a high priority to complete the payment quickly. For example, it can prioritize the payment by postponing other processes. If the buyer is relaxed, the payment unit can also perform the payment with normal priority. For example, it can perform the payment with the same priority as other processes. If the buyer is in a hurry, the payment unit can also set the priority to the highest to complete the payment as the top priority. For example, it can pause all processes and perform the payment as the top priority. This allows the payment priority to be determined according to the buyer's emotions, resulting in a quick payment.
[0103] The payment unit can select the optimal payment method at the time of payment, taking into account the geographic location information of the purchaser. For example, the payment unit selects the optimal payment method by taking into account the geographic location information of the purchaser. The payment unit can also select the payment method by taking into account the geographic location information using generation AI. The generation AI can be a text generation AI (e.g., LLM) or a multimodal generation AI. For example, if the purchaser is in a specific region, the payment unit can suggest payment methods commonly used in that region. For example, if the purchaser is in Japan, the payment unit can suggest "Suica" or "PayPay." If the purchaser is traveling, the payment unit can also suggest payment methods available at the travel destination. For example, if the purchaser is in the United States, the payment unit can suggest "Apple Pay" or "Google Pay." If the purchaser is at home, the payment unit can also suggest payment methods that the purchaser normally uses. For example, the payment unit can suggest credit cards that the purchaser has previously used at home. In this way, the optimal payment method is selected by taking into account the geographic location information.
[0104] At the time of payment, the payment unit can analyze the buyer's social media activity to suggest payment methods. For example, the payment unit analyzes the buyer's social media activity to suggest payment methods. The payment unit can also use generation AI to analyze social media activity to suggest payment methods. The generation AI can be a text generation AI (e.g., LLM) or a multimodal generation AI. For example, the payment unit can suggest relevant payment methods based on the location where the buyer checked in on social media. For example, if the buyer checked in at a specific cafe, the payment unit can suggest payment methods available at that cafe. The payment unit can also analyze the content of the buyer's social media posts to suggest relevant payment methods. For example, if the buyer posted about a specific product, the payment unit can suggest payment methods for purchasing that product. The payment unit can also suggest relevant payment methods based on the activity of the buyer's friends on social media. For example, the payment unit can suggest payment methods used by the friends. This makes payment easier by suggesting the optimal payment method based on social media activity.
[0105] The payment department can customize the payment method by reflecting the buyer's past feedback at the time of payment. For example, the payment department customizes the payment method by reflecting the buyer's past feedback. The payment department can also customize the payment method by using generation AI to reflect past feedback. The generation AI can be a text generation AI (e.g., LLM) or a multimodal generation AI. For example, the payment department customizes the payment method based on feedback provided by the buyer in the past. For example, it resolves issues pointed out in the feedback. The payment department can also reflect feedback provided by the buyer after payment in real time to customize the next payment method. For example, it adds features suggested in the feedback. The payment department can also analyze buyer feedback and perform customization to resolve common issues. For example, it resolves issues pointed out by multiple buyers. This makes payment easier because the payment method is customized based on past feedback. === Hard Collateral 1-1 === For example, the generation unit is realized by the processor 46 of the smart device 14 or the specific processing unit 290 of the data processing device 12. For example, the transmission unit is realized by the output device 40 of the smart device 14 or the specific processing unit 290 of the data processing device 12. For example, the voiceprint authentication unit is realized by the processor 46 of the smart device 14 or the specific processing unit 290 of the data processing device 12. For example, the payment unit is realized by the processor 46 of the smart device 14 or the specific processing unit 290 of the data processing device 12. === Hard Collateral 1-2 === For example, the generation unit is realized by the processor 46 of the smart glasses 214 or the specific processing unit 290 of the data processing device 12. For example, the transmission unit is realized by the speaker 240 of the smart glasses 214 or the specific processing unit 290 of the data processing device 12. For example, the voiceprint authentication unit is realized by the processor 46 of the smart glasses 214 or the specific processing unit 290 of the data processing device 12. For example, the payment unit is realized by the processor 46 of the smart glasses 214 or the specific processing unit 290 of the data processing device 12. === Hard Collateral 1-3 === For example, the generation unit is realized by the processor 46 of the headset type terminal 314 or the specific processing unit 290 of the data processing device 12. For example, the transmission unit is realized by the speaker 240 of the headset type terminal 314 or the specific processing unit 290 of the data processing device 12. For example, the voiceprint authentication unit is realized by the processor 46 of the headset type terminal 314 or the specific processing unit 290 of the data processing device 12. For example, the payment unit is realized by the processor 46 of the headset type terminal 314 or the specific processing unit 290 of the data processing device 12. === Hard Collateral 1-4 === For example, the generation unit is realized by the processor 46 of the robot 414 or the specific processing unit 290 of the data processing device 12. For example, the transmission unit is realized by the speaker 240 of the robot 414 or the specific processing unit 290 of the data processing device 12. For example, the voiceprint authentication unit is realized by the processor 46 of the robot 414 or the specific processing unit 290 of the data processing device 12. For example, the payment unit is realized by the processor 46 of the robot 414 or the specific processing unit 290 of the data processing device 12.
[0106] The system according to the embodiment is not limited to the above-described example, and various modifications are possible, for example, as follows.
[0107] The payment system may further include a health monitoring unit that monitors the user's health condition. The health monitoring unit acquires the user's vital signs, such as heart rate and blood pressure, in real time and evaluates the user's health condition. For example, if the user has high blood pressure, the health monitoring unit provides that information to the payment unit, which can select a payment method appropriate to the user's health condition. If the user is feeling stressed, the health monitoring unit provides that information to the generation unit, which can then generate a simple authentication phrase that reduces stress. Furthermore, the health monitoring unit can issue a warning if the user has a health problem. This allows for flexible responses based on the user's health condition, providing a safer and more comfortable payment experience.
[0108] The payment system may further include a purchasing pattern analysis unit that analyzes a user's purchasing patterns. The purchasing pattern analysis unit analyzes the user's past purchasing history and understands the user's purchasing trends. For example, if the user tends to purchase specific products on specific days of the week or at specific times of the day, the purchasing pattern analysis unit provides that information to the payment unit, which can then suggest the optimal payment method based on that trend. If the user frequently purchases products from a specific brand or category, the purchasing pattern analysis unit provides that information to the generation unit, which can then generate authentication keywords related to that brand or category. Furthermore, the purchasing pattern analysis unit may also provide benefits or discounts based on the user's purchasing patterns. This allows for personalized services to be provided in accordance with the user's purchasing behavior.
[0109] The payment system may further include a location information utilization unit that utilizes the user's location information. The location information utilization unit obtains the user's current location information and provides that information to other elements. For example, if the user is in a specific store, the location information utilization unit provides that information to a payment unit, which can suggest payment methods available at that store. If the user is traveling, the location information utilization unit provides that information to a generation unit, which can generate authentication words related to the language and culture of the destination. Furthermore, if the user is in a specific region, the location information utilization unit can provide information about special offers and discounts for that region. This enables flexible responses based on the user's location information and provides a more convenient payment experience.
[0110] The payment system may further include a voice command analysis unit that analyzes a user's voice command. The voice command analysis unit analyzes the voice command issued by the user and performs appropriate processing based on the content of the command. For example, if the user utters "check balance," the voice command analysis unit analyzes the command and provides balance information to the providing unit. The providing unit can convey the balance information to the user by voice. If the user utters "cancel," the voice command analysis unit can analyze the command and send a cancellation instruction to the payment unit. Furthermore, the voice command analysis unit can automate certain operations based on the user's voice command. This enables flexible responses based on the user's voice command, providing a more intuitive and convenient payment experience.
[0111] The payment system may further include an advertisement providing unit that estimates the user's emotions and provides advertisements based on the emotions. The advertisement providing unit estimates the user's emotions and selects the most appropriate advertisement based on the emotions. For example, if the user is relaxed, the advertisement providing unit may provide advertisements for products related to relaxation. If the user is excited, the advertisement providing unit may provide advertisements for products related to entertainment. Furthermore, the advertisement providing unit may provide personalized advertisements based on the user's past purchase history and interests. This provides advertisements that correspond to the user's emotions, improving the effectiveness of the advertisements and providing useful information to the user.
[0112] The payment system may further include a behavior prediction unit that predicts user behavior. The behavior prediction unit analyzes the user's past behavior data and predicts future behavior. For example, if the user tends to visit a particular store on a particular day or time, the behavior prediction unit provides that information to the payment unit, which can then prepare payment methods available at that store in advance. If the user plans to attend a particular event, the behavior prediction unit provides that information to the generation unit, which can then generate authentication words related to the event. Furthermore, the behavior prediction unit may provide benefits or discounts based on the user's behavioral patterns. This allows for personalized services to be provided according to the user's behavior.
[0113] The payment system may further include a music providing unit that estimates the user's emotions and provides music based on the emotions. The music providing unit estimates the user's emotions and selects optimal music based on the emotions. For example, if the user is relaxed, the music providing unit may provide music suitable for relaxation. If the user is excited, the music providing unit may provide energetic music. Furthermore, the music providing unit may provide personalized music based on the user's past music history and preferences. This provides music that matches the user's emotions, improving the user's experience and providing a more comfortable payment environment.
[0114] The payment system can further include an incentive providing unit to increase the user's purchasing motivation. The incentive providing unit analyzes the user's purchasing history and behavioral patterns and provides optimal incentives. For example, if the user frequently purchases a particular product, the incentive providing unit can provide a discount coupon for that product. If the user frequently visits a particular store, the incentive providing unit can provide points that can be used at that store. Furthermore, the incentive providing unit can also provide limited-time benefits and campaigns to increase the user's purchasing motivation. This can increase the user's purchasing motivation and encourage more purchasing behavior.
[0115] The payment system may further include a feedback providing unit that estimates the user's emotions and provides customized feedback based on the emotions. The feedback providing unit estimates the user's emotions and provides optimal feedback based on the emotions. For example, if the user is nervous, the feedback providing unit may provide feedback that encourages the user to relax. If the user is relaxed, the feedback providing unit may provide feedback that helps the user maintain that state. Furthermore, the feedback providing unit may provide personalized feedback based on the user's past feedback history and behavior. This provides feedback that corresponds to the user's emotions, improving the user's experience and providing a more comfortable payment environment.
[0116] The payment system may further include a notification providing unit that estimates a user's emotion and provides a customized notification based on the emotion. The notification providing unit estimates a user's emotion and provides an optimal notification based on the emotion. For example, if the user is nervous, the notification providing unit may provide a notification encouraging the user to relax. If the user is relaxed, the notification providing unit may provide a notification to maintain that state. Furthermore, the notification providing unit may provide personalized notifications based on the user's past notification history and behavior. This provides notifications according to the user's emotion, improving the user's experience and providing a more comfortable payment environment.
[0117] The processing flow of the second embodiment will be briefly explained below.
[0118] Step 1: The generator generates authentication words. The generator randomly generates authentication words using, for example, a generation AI. The generation AI may be a text generation AI (e.g., LLM) or a multimodal generation AI. The generator can generate words such as "apple pie" or "sunflower." The generator can also generate authentication words based on the buyer's emotions and past authentication history. For example, if the buyer is nervous, it can generate a simple word. Step 2: The communication unit communicates the generated authentication word to the purchaser. The communication unit communicates the authentication word aloud, for example, using voice synthesis technology. The communication unit can also adjust the tone and speed of the voice based on the hearing characteristics and emotions of the purchaser. For example, the authentication word can be communicated at a slower speed to an elderly person. Step 3: The voiceprint authentication unit analyzes the customer's voice and compares it with pre-registered voiceprint data. For example, the voiceprint authentication unit uses voice signal processing technology to analyze the customer's voice and extract features. The voiceprint authentication unit can also optimize the authentication algorithm by taking into account the customer's emotions and voice fluctuations. For example, if the customer has a cold, authentication is performed taking these fluctuations into account. Step 4: The payment unit executes payment if the voiceprint authentication is successful. The payment unit performs payment using means such as credit card payment or electronic money payment. The payment unit can also select the optimal payment method based on the buyer's emotions and past payment history. For example, if the buyer is in a hurry, it can provide a quick payment method.
[0119] The specific processing unit 290 transmits the result of the specific processing to the smart device 14. In the smart device 14, the control unit 46A causes the output device 40 to output the result of the specific processing. The microphone 38B acquires audio indicating a user input regarding the result of the specific processing. The control unit 46A transmits audio data indicating the user input acquired by the microphone 38B to the data processing device 12. In the data processing device 12, the specific processing unit 290 acquires the audio data.
[0120] The data generation model 58 is a so-called generative AI (Artificial Intelligence). An example of the data generation model 58 is ChatGPT (registered trademark) (Internet search engine).<URL: https: / / openai.com / blog / chatgpt> Examples of generative AIs include the data generation model 58, such as a neural network model (e.g., a neural network model), and a neural network model (e.g., a neural network model). The data generation model 58 is obtained by performing deep learning on a neural network. A prompt including an instruction is input to the data generation model 58, and inference data such as voice data indicating speech, text data indicating text, and image data indicating an image is also input to the data generation model 58. The data generation model 58 performs inference on the input inference data in accordance with the instruction indicated by the prompt and outputs the inference result in a data format such as voice data and text data. Here, inference refers to, for example, analysis, classification, prediction, and / or summarization. The specification processing unit 290 performs the above-mentioned specification processing using the data generation model 58. The data generation model 58 may be a fine-tuned model so as to output an inference result from a prompt that does not include an instruction. In this case, the data generation model 58 can output an inference result from a prompt that does not include an instruction. The data processing device 12 and the like include multiple types of data generation models 58, and the data generation model 58 includes AIs other than the generative AI. The AI other than the generative AI may be, for example, linear regression, logistic regression, decision tree, random forest, support vector machine (SVM), k-means clustering, convolutional neural network (CNN), recurrent neural network (RNN), generative adversarial network (GAN), or naive Bayes, and can perform various processes, but is not limited to these examples. The AI may also be an AI agent. When the processes of each of the above-mentioned parts are performed by AI, the processes may be performed in part or entirely by AI, but are not limited to these examples. The processes performed by AI, including the generative AI, may be replaced with rule-based processes.
[0121] Furthermore, the processing by the data processing system 10 described above is executed by the specific processing unit 290 of the data processing device 12 or the control unit 46A of the smart device 14, but may also be executed by the specific processing unit 290 of the data processing device 12 and the control unit 46A of the smart device 14. Furthermore, the specific processing unit 290 of the data processing device 12 acquires or collects information necessary for processing from the smart device 14 or an external device, and the smart device 14 acquires or collects information necessary for processing from the data processing device 12 or an external device.
[0122] The correspondence between each part and the device or control part is not limited to the example described above, and various modifications are possible.
[0123] [Second embodiment] FIG. 3 shows an example of the configuration of a data processing system 210 according to the second embodiment.
[0124] 3, the data processing system 210 includes the data processing device 12 and smart glasses 214. An example of the data processing device 12 is a server.
[0125] The data processing device 12 includes a computer 22, a database 24, and a communication I / F 26. The computer 22 includes a processor 28, a RAM 30, and a storage 32. The processor 28, RAM 30, and storage 32 are connected to a bus 34. The database 24 and the communication I / F 26 are also connected to the bus 34. The communication I / F 26 is connected to a network 54. Examples of the network 54 include a WAN and / or a LAN.
[0126] The smart glasses 214 include a computer 36, a microphone 238, a speaker 240, a camera 42, and a communication I / F 44. The computer 36 includes a processor 46, a RAM 48, and a storage 50. The processor 46, the RAM 48, and the storage 50 are connected to a bus 52. The microphone 238, the speaker 240, and the camera 42 are also connected to the bus 52.
[0127] The microphone 238 receives instructions and the like from the user by receiving voice uttered by the user. The microphone 238 captures the voice uttered by the user, converts the captured voice into audio data, and outputs it to the processor 46. The speaker 240 outputs audio according to instructions from the processor 46.
[0128] Camera 42 is a small digital camera equipped with an optical system including a lens, aperture, and shutter, and an imaging element such as a CMOS (Complementary Metal-Oxide-Semiconductor) image sensor or a CCD (Charge Coupled Device) image sensor, and captures images of the user's surroundings (for example, an imaging range defined by an angle of view equivalent to the field of vision of a typical healthy person).
[0129] The communication I / F 44 is connected to a network 54. The communication I / Fs 44 and 26 are responsible for the exchange of various information between the processor 46 and the processor 28 via the network 54. The exchange of various information between the processor 46 and the processor 28 using the communication I / Fs 44 and 26 is carried out in a secure state.
[0130] Fig. 4 shows an example of the main functions of the data processing device 12 and the smart glasses 214. As shown in Fig. 4, in the data processing device 12, a specific process is performed by the processor 28. A specific process program 56 is stored in the storage 32.
[0131] The processor 28 reads the specific processing program 56 from the storage 32 and executes the read specific processing program 56 on the RAM 30. The specific processing is realized by the processor 28 operating as a specific processing unit 290 in accordance with the specific processing program 56 executed on the RAM 30.
[0132] The storage 32 stores a data generation model 58 and an emotion identification model 59. The data generation model 58 and the emotion identification model 59 are used by the identification processing unit 290. The identification processing unit 290 can estimate a user's emotion using the emotion identification model 59 and perform identification processing using the user's emotion. The emotion estimation function (emotion identification function) using the emotion identification model 59 performs various estimations and predictions regarding the user's emotion, including estimation and prediction of the user's emotion, but is not limited to these examples. Furthermore, the estimation and prediction of emotion also includes, for example, emotion analysis.
[0133] In the smart glasses 214, the specific processing is performed by the processor 46. A specific processing program 60 is stored in the storage 50. The processor 46 reads the specific processing program 60 from the storage 50 and executes the read specific processing program 60 on the RAM 48. The specific processing is realized by the processor 46 operating as the control unit 46A in accordance with the specific processing program 60 executed on the RAM 48. The smart glasses 214 also have a data generation model and an emotion identification model similar to the data generation model 58 and the emotion identification model 59, and can perform processing similar to that of the specific processing unit 290 using these models.
[0134] Note that a device other than the data processing device 12 may have the data generation model 58. For example, a server device may have the data generation model 58. In this case, the data processing device 12 communicates with the server device having the data generation model 58 to obtain a processing result (such as a prediction result) using the data generation model 58. Furthermore, the data processing device 12 may be a server device, or may be a terminal device (for example, a mobile phone, a robot, a home appliance, etc.) owned by a user.
[0135] The specific processing unit 290 transmits the result of the specific processing to the smart glasses 214. In the smart glasses 214, the control unit 46A causes the speaker 240 to output the result of the specific processing. The microphone 238 acquires audio indicating a user input regarding the result of the specific processing. The control unit 46A transmits audio data indicating the user input acquired by the microphone 238 to the data processing device 12. In the data processing device 12, the specific processing unit 290 acquires the audio data.
[0136] The data generation model 58 is a so-called generative AI. An example of the data generation model 58 is a generative AI such as ChatGPT. The data generation model 58 is obtained by performing deep learning on a neural network. The data generation model 58 receives a prompt containing an instruction, as well as inference data such as voice data representing speech, text data representing text, and image data representing an image. The data generation model 58 performs inference on the input inference data in accordance with the instruction indicated by the prompt and outputs the inference result in a data format such as voice data and text data. Here, inference refers to, for example, analysis, classification, prediction, and / or summarization. The identification processing unit 290 performs the above-mentioned identification processing using the data generation model 58. The data generation model 58 may be a fine-tuned model so as to output an inference result from a prompt that does not include an instruction. In this case, the data generation model 58 can output an inference result from a prompt that does not include an instruction. The data processing device 12 and the like include multiple types of data generation models 58, and the data generation model 58 includes AI other than the generative AI. The AI other than the generative AI may be, for example, linear regression, logistic regression, decision tree, random forest, support vector machine (SVM), k-means clustering, convolutional neural network (CNN), recurrent neural network (RNN), generative adversarial network (GAN), or naive Bayes, and can perform various processes, but is not limited to these examples. The AI may also be an AI agent. When the processes of each of the above-mentioned parts are performed by AI, the processes may be performed in part or entirely by AI, but are not limited to these examples. The processes performed by AI, including the generative AI, may be replaced with rule-based processes.
[0137] The data processing system 210 according to the second embodiment performs the same processing as the data processing system 10 according to the first embodiment. The processing by the data processing system 210 is executed by the specific processing unit 290 of the data processing device 12 or the control unit 46A of the smart glasses 214, but may also be executed by the specific processing unit 290 of the data processing device 12 and the control unit 46A of the smart glasses 214. Furthermore, the specific processing unit 290 of the data processing device 12 acquires or collects information required for processing from the smart glasses 214 or an external device, etc., and the smart glasses 214 acquires or collects information required for processing from the data processing device 12 or an external device, etc.
[0138] The correspondence between each part and the device or control part is not limited to the example described above, and various modifications are possible.
[0139] [Third embodiment] FIG. 5 shows an example of the configuration of a data processing system 310 according to the third embodiment.
[0140] 5, the data processing system 310 includes the data processing device 12 and a headset terminal 314. An example of the data processing device 12 is a server.
[0141] The data processing device 12 includes a computer 22, a database 24, and a communication I / F 26. The computer 22 includes a processor 28, a RAM 30, and a storage 32. The processor 28, RAM 30, and storage 32 are connected to a bus 34. The database 24 and the communication I / F 26 are also connected to the bus 34. The communication I / F 26 is connected to a network 54. Examples of the network 54 include a WAN and / or a LAN.
[0142] The headset type terminal 314 includes a computer 36, a microphone 238, a speaker 240, a camera 42, a communication I / F 44, and a display 343. The computer 36 includes a processor 46, a RAM 48, and a storage 50. The processor 46, the RAM 48, and the storage 50 are connected to a bus 52. The microphone 238, the speaker 240, the camera 42, and the display 343 are also connected to the bus 52.
[0143] The microphone 238 receives instructions and the like from the user by receiving voice uttered by the user. The microphone 238 captures the voice uttered by the user, converts the captured voice into audio data, and outputs it to the processor 46. The speaker 240 outputs audio according to instructions from the processor 46.
[0144] Camera 42 is a small digital camera equipped with an optical system including a lens, aperture, and shutter, and an imaging element such as a CMOS (Complementary Metal-Oxide-Semiconductor) image sensor or a CCD (Charge Coupled Device) image sensor, and captures images of the user's surroundings (for example, an imaging range defined by an angle of view equivalent to the field of vision of a typical healthy person).
[0145] The communication I / F 44 is connected to a network 54. The communication I / Fs 44 and 26 are responsible for the exchange of various information between the processor 46 and the processor 28 via the network 54. The exchange of various information between the processor 46 and the processor 28 using the communication I / Fs 44 and 26 is carried out in a secure state.
[0146] Fig. 6 shows an example of the main functions of the data processing device 12 and the headset type terminal 314. As shown in Fig. 6, in the data processing device 12, a specific process is performed by the processor 28. A specific process program 56 is stored in the storage 32.
[0147] The processor 28 reads the specific processing program 56 from the storage 32 and executes the read specific processing program 56 on the RAM 30. The specific processing is realized by the processor 28 operating as a specific processing unit 290 in accordance with the specific processing program 56 executed on the RAM 30.
[0148] The storage 32 stores a data generation model 58 and an emotion identification model 59. The data generation model 58 and the emotion identification model 59 are used by the identification processing unit 290. The identification processing unit 290 can estimate a user's emotion using the emotion identification model 59 and perform identification processing using the user's emotion. The emotion estimation function (emotion identification function) using the emotion identification model 59 performs various estimations and predictions regarding the user's emotion, including estimation and prediction of the user's emotion, but is not limited to these examples. Furthermore, the estimation and prediction of emotion also includes, for example, emotion analysis.
[0149] In the headset type terminal 314, the identification process is performed by the processor 46. A identification program 60 is stored in the storage 50. The processor 46 reads the identification program 60 from the storage 50 and executes the read identification program 60 on the RAM 48. The identification process is realized by the processor 46 operating as a control unit 46A in accordance with the identification program 60 executed on the RAM 48. Note that the headset type terminal 314 has a data generation model and an emotion identification model similar to the data generation model 58 and the emotion identification model 59, and can also perform processing similar to that of the identification processing unit 290 using these models.
[0150] Note that a device other than the data processing device 12 may have the data generation model 58. For example, a server device may have the data generation model 58. In this case, the data processing device 12 communicates with the server device having the data generation model 58 to obtain a processing result (such as a prediction result) using the data generation model 58. Furthermore, the data processing device 12 may be a server device, or may be a terminal device (for example, a mobile phone, a robot, a home appliance, etc.) owned by a user.
[0151] The specific processing unit 290 transmits the result of the specific processing to the headset type terminal 314. In the headset type terminal 314, the control unit 46A causes the speaker 240 and the display 343 to output the result of the specific processing. The microphone 238 acquires audio indicating a user input regarding the result of the specific processing. The control unit 46A transmits audio data indicating the user input acquired by the microphone 238 to the data processing device 12. In the data processing device 12, the specific processing unit 290 acquires the audio data.
[0152] The data generation model 58 is a so-called generative AI. An example of the data generation model 58 is a generative AI such as ChatGPT. The data generation model 58 is obtained by performing deep learning on a neural network. The data generation model 58 receives a prompt containing an instruction, as well as inference data such as voice data representing speech, text data representing text, and image data representing an image. The data generation model 58 performs inference on the input inference data in accordance with the instruction indicated by the prompt and outputs the inference result in a data format such as voice data and text data. Here, inference refers to, for example, analysis, classification, prediction, and / or summarization. The identification processing unit 290 performs the above-mentioned identification processing using the data generation model 58. The data generation model 58 may be a fine-tuned model so as to output an inference result from a prompt that does not include an instruction. In this case, the data generation model 58 can output an inference result from a prompt that does not include an instruction. The data processing device 12 and the like include multiple types of data generation models 58, and the data generation model 58 includes AI other than the generative AI. The AI other than the generative AI may be, for example, linear regression, logistic regression, decision tree, random forest, support vector machine (SVM), k-means clustering, convolutional neural network (CNN), recurrent neural network (RNN), generative adversarial network (GAN), or naive Bayes, and can perform various processes, but is not limited to these examples. The AI may also be an AI agent. When the processes of each of the above-mentioned parts are performed by AI, the processes may be performed in part or entirely by AI, but are not limited to these examples. The processes performed by AI, including the generative AI, may be replaced with rule-based processes.
[0153] The data processing system 310 according to the third embodiment performs the same processing as the data processing system 10 according to the first embodiment. The processing by the data processing system 310 is executed by the specific processing unit 290 of the data processing device 12 or the control unit 46A of the headset type terminal 314, but may also be executed by the specific processing unit 290 of the data processing device 12 and the control unit 46A of the headset type terminal 314. Furthermore, the specific processing unit 290 of the data processing device 12 acquires or collects information required for processing from the headset type terminal 314 or an external device, etc., and the headset type terminal 314 acquires or collects information required for processing from the data processing device 12 or an external device, etc.
[0154] The correspondence between each part and the device or control part is not limited to the example described above, and various modifications are possible.
[0155] [Fourth embodiment] FIG. 7 shows an example of the configuration of a data processing system 410 according to the fourth embodiment.
[0156] 7, a data processing system 410 includes a data processing device 12 and a robot 414. An example of the data processing device 12 is a server.
[0157] The data processing device 12 includes a computer 22, a database 24, and a communication I / F 26. The computer 22 includes a processor 28, a RAM 30, and a storage 32. The processor 28, RAM 30, and storage 32 are connected to a bus 34. The database 24 and the communication I / F 26 are also connected to the bus 34. The communication I / F 26 is connected to a network 54. Examples of the network 54 include a WAN and / or a LAN.
[0158] The robot 414 includes a computer 36, a microphone 238, a speaker 240, a camera 42, a communication I / F 44, and a control target 443. The computer 36 includes a processor 46, a RAM 48, and a storage 50. The processor 46, the RAM 48, and the storage 50 are connected to a bus 52. The microphone 238, the speaker 240, the camera 42, and the control target 443 are also connected to the bus 52.
[0159] The microphone 238 receives instructions and the like from the user by receiving voice uttered by the user. The microphone 238 captures the voice uttered by the user, converts the captured voice into audio data, and outputs it to the processor 46. The speaker 240 outputs audio according to instructions from the processor 46.
[0160] Camera 42 is a small digital camera equipped with an optical system including a lens, aperture, and shutter, and an imaging element such as a CMOS image sensor or a CCD image sensor, and captures images of the user's surroundings (for example, an imaging range defined by an angle of view equivalent to the field of vision of a typical healthy person).
[0161] The communication I / F 44 is connected to a network 54. The communication I / Fs 44 and 26 are responsible for the exchange of various information between the processor 46 and the processor 28 via the network 54. The exchange of various information between the processor 46 and the processor 28 using the communication I / Fs 44 and 26 is carried out in a secure state.
[0162] The control object 443 includes a display device, LEDs in the eyes, and motors that drive the arms, hands, and feet. The posture and gestures of the robot 414 are controlled by controlling the motors of the arms, hands, and feet. Some of the emotions of the robot 414 can be expressed by controlling these motors. In addition, the facial expressions of the robot 414 can also be expressed by controlling the light emission state of the LEDs in the eyes of the robot 414.
[0163] Fig. 8 shows an example of the main functions of the data processing device 12 and the robot 414. As shown in Fig. 8, in the data processing device 12, a specific process is performed by the processor 28. A specific process program 56 is stored in the storage 32.
[0164] The processor 28 reads the specific processing program 56 from the storage 32 and executes the read specific processing program 56 on the RAM 30. The specific processing is realized by the processor 28 operating as a specific processing unit 290 in accordance with the specific processing program 56 executed on the RAM 30.
[0165] The storage 32 stores a data generation model 58 and an emotion identification model 59. The data generation model 58 and the emotion identification model 59 are used by the identification processing unit 290. The identification processing unit 290 can estimate a user's emotion using the emotion identification model 59 and perform identification processing using the user's emotion. The emotion estimation function (emotion identification function) using the emotion identification model 59 performs various estimations and predictions regarding the user's emotion, including estimation and prediction of the user's emotion, but is not limited to these examples. Furthermore, the estimation and prediction of emotion also includes, for example, emotion analysis.
[0166] In the robot 414, the processor 46 performs the identification process. The storage 50 stores the identification program 60. The processor 46 reads the identification program 60 from the storage 50 and executes the read identification program 60 on the RAM 48. The identification process is realized by the processor 46 operating as the control unit 46A in accordance with the identification program 60 executed on the RAM 48. The robot 414 also has a data generation model and an emotion identification model similar to the data generation model 58 and the emotion identification model 59, and can perform the same process as the identification processing unit 290 using these models.
[0167] Note that a device other than the data processing device 12 may have the data generation model 58. For example, a server device may have the data generation model 58. In this case, the data processing device 12 communicates with the server device having the data generation model 58 to obtain a processing result (such as a prediction result) using the data generation model 58. Furthermore, the data processing device 12 may be a server device, or may be a terminal device (for example, a mobile phone, a robot, a home appliance, etc.) owned by a user.
[0168] The specific processing unit 290 transmits the result of the specific processing to the robot 414. In the robot 414, the control unit 46A causes the speaker 240 and the control target 443 to output the result of the specific processing. The microphone 238 acquires voice indicating a user input regarding the result of the specific processing. The control unit 46A transmits voice data indicating the user input acquired by the microphone 238 to the data processing device 12. In the data processing device 12, the specific processing unit 290 acquires the voice data.
[0169] The data generation model 58 is a so-called generative AI. An example of the data generation model 58 is a generative AI such as ChatGPT. The data generation model 58 is obtained by performing deep learning on a neural network. The data generation model 58 receives a prompt containing an instruction, as well as inference data such as voice data representing speech, text data representing text, and image data representing an image. The data generation model 58 performs inference on the input inference data in accordance with the instruction indicated by the prompt and outputs the inference result in a data format such as voice data and text data. Here, inference refers to, for example, analysis, classification, prediction, and / or summarization. The identification processing unit 290 performs the above-mentioned identification processing using the data generation model 58. The data generation model 58 may be a fine-tuned model so as to output an inference result from a prompt that does not include an instruction. In this case, the data generation model 58 can output an inference result from a prompt that does not include an instruction. The data processing device 12 and the like include multiple types of data generation models 58, and the data generation model 58 includes AI other than the generative AI. The AI other than the generative AI may be, for example, linear regression, logistic regression, decision tree, random forest, support vector machine (SVM), k-means clustering, convolutional neural network (CNN), recurrent neural network (RNN), generative adversarial network (GAN), or naive Bayes, and can perform various processes, but is not limited to these examples. The AI may also be an AI agent. When the processes of each of the above-mentioned parts are performed by AI, the processes may be performed in part or entirely by AI, but are not limited to these examples. The processes performed by AI, including the generative AI, may be replaced with rule-based processes.
[0170] The data processing system 410 according to the fourth embodiment performs the same processing as the data processing system 10 according to the first embodiment. The processing by the data processing system 410 is executed by the specific processing unit 290 of the data processing device 12 or the control unit 46A of the robot 414, but may also be executed by the specific processing unit 290 of the data processing device 12 and the control unit 46A of the robot 414. Furthermore, the specific processing unit 290 of the data processing device 12 acquires or collects information required for processing from the robot 414 or an external device, etc., and the robot 414 acquires or collects information required for processing from the data processing device 12 or an external device, etc.
[0171] The correspondence between each part and the device or control part is not limited to the example described above, and various modifications are possible.
[0172] The emotion identification model 59 as an emotion engine may determine the user's emotion according to a specific mapping. Specifically, the emotion identification model 59 may determine the user's emotion according to an emotion map (see FIG. 9), which is a specific mapping. Similarly, the emotion identification model 59 may determine the robot's emotion, and the identification processing unit 290 may perform identification processing using the robot's emotion.
[0173] FIG. 9 illustrates an emotion map 400 on which multiple emotions are mapped. In the emotion map 400, emotions are arranged in concentric circles radiating from the center. Emotions closer to the center of the concentric circles are more primitive. Emotions representing states and behaviors arising from a state of mind are arranged on the outer edges of the concentric circles. The concept of emotion encompasses both emotions and mental states. Emotions generally generated from reactions occurring in the brain are arranged on the left side of the concentric circles. Emotions generally induced by situational judgment are arranged on the right side of the concentric circles. Emotions generally generated from reactions occurring in the brain and induced by situational judgment are arranged on the upper and lower sides of the concentric circles. Furthermore, the emotion of "pleasure" is arranged on the upper side of the concentric circles, and the emotion of "discomfort" is arranged on the lower side. In this way, in the emotion map 400, multiple emotions are mapped based on the structure by which emotions are generated, and emotions that tend to occur simultaneously are mapped close to each other.
[0174] These emotions are distributed in the 3 o'clock direction on emotion map 400, and typically fluctuate between relief and anxiety. In the right half of emotion map 400, situational awareness dominates over internal sensations, resulting in a sense of calm.
[0175] The inside of emotion map 400 represents what is going on in the mind, and the outside of emotion map 400 represents behavior, so the further you go outside emotion map 400, the more visible the emotions become (the more they are expressed in behavior).
[0176] Human emotions are based on various balances, such as posture and blood sugar levels. When these balances deviate from the ideal, a state of discomfort is expressed, and when they approach the ideal, a state of pleasure is expressed. Emotions can also be created for robots, cars, and motorcycles, based on various balances, such as posture and remaining battery life. When these balances deviate from the ideal, a state of discomfort is expressed, and when they approach the ideal, a state of pleasure is expressed. An emotion map can be generated, for example, based on Dr. Mitsuyoshi's emotion map (Research on speech emotion recognition and brain physiological signal analysis systems for emotions, Tokushima University, doctoral dissertation: https: / / ci.nii.ac.jp / naid / 500000375379). The left half of the emotion map lists emotions belonging to the area called "reaction," where sensation is dominant. The right half of the emotion map lists emotions belonging to the area called "situation," where situational awareness is dominant.
[0177] The emotion map defines two emotions that promote learning. One is a negative emotion on the situation side, around the middle of "repentance" or "reflection." In other words, this occurs when the robot experiences negative emotions such as "I never want to feel this way again" or "I don't want to be scolded again." The other is a positive emotion on the response side, around "desire." In other words, this occurs when the robot experiences positive feelings such as "I want more" or "I want to know more."
[0178] The emotion identification model 59 inputs user input into a pre-trained neural network, obtains emotion values indicating each emotion shown in the emotion map 400, and determines the user's emotion. This neural network is pre-trained based on multiple pieces of training data that are combinations of user input and emotion values indicating each emotion shown in the emotion map 400. Furthermore, this neural network is trained so that emotions that are located close to each other have similar values, as in the emotion map 900 shown in FIG. 10. FIG. 10 shows an example in which multiple emotions, "relieved," "calm," and "reassuring," have similar emotion values.
[0179] In the above embodiment, an example was given in which a specific process is performed by one computer 22, but the technology disclosed herein is not limited to this, and distributed processing of the specific process may be performed by multiple computers including computer 22.
[0180] In the above embodiment, an example in which the specific processing program 56 is stored in the storage 32 has been described, but the technology of the present disclosure is not limited to this. For example, the specific processing program 56 may be stored in a portable, computer-readable, non-transitory storage medium such as a USB (Universal Serial Bus) memory. The specific processing program 56 stored in the non-transitory storage medium is installed in the computer 22 of the data processing device 12. The processor 28 executes the specific processing in accordance with the specific processing program 56.
[0181] Alternatively, the specific processing program 56 may be stored in a storage device such as a server connected to the data processing device 12 via the network 54, and the specific processing program 56 may be downloaded and installed on the computer 22 in response to a request from the data processing device 12.
[0182] It is not necessary to store all of the specific processing program 56 in a storage device such as a server connected to the data processing device 12 via the network 54, or to store all of the specific processing program 56 in the storage 32; only a portion of the specific processing program 56 may be stored.
[0183] The hardware resource for executing a specific process can be any of the following types of processors: A processor, for example, is a CPU, which is a general-purpose processor that functions as a hardware resource for executing a specific process by executing software, i.e., a program. A processor also includes a dedicated electrical circuit, such as an FPGA (Field-Programmable Gate Array), a PLD (Programmable Logic Device), or an ASIC (Application Specific Integrated Circuit), which is a processor with a circuit configuration designed specifically for executing a specific process. Each processor has built-in or connected memory, and each processor uses the memory to execute the specific process.
[0184] The hardware resource that executes the specific process may be configured with one of these various processors, or may be configured with a combination of two or more processors of the same or different types (for example, a combination of multiple FPGAs, or a combination of a CPU and an FPGA). Also, the hardware resource that executes the specific process may be a single processor.
[0185] As an example of a system configured with a single processor, first, one processor is configured by combining one or more CPUs and software, and this processor functions as a hardware resource that executes a specific process. Second, there is a system that uses a processor that realizes the functions of an entire system including multiple hardware resources that execute a specific process on a single IC chip, as typified by SoC (System-on-a-chip). In this way, a specific process is realized using one or more of the above-mentioned various processors as hardware resources.
[0186] Furthermore, the hardware structure of these various processors can be, more specifically, an electric circuit that combines circuit elements such as semiconductor devices. The specific processing described above is merely an example. Therefore, it goes without saying that unnecessary steps may be deleted, new steps may be added, or the processing order may be rearranged, without departing from the spirit of the invention.
[0187] In the above example, the first to fourth embodiments have been described separately, but some or all of these embodiments may be combined. The smart device 14, smart glasses 214, headset terminal 314, and robot 414 are merely examples, and they may be combined, or other devices may be used. In the above example, the first and second embodiments have been described separately, but they may be combined.
[0188] The above-described description and illustrations are a detailed explanation of the parts related to the technology of the present disclosure and are merely an example of the technology of the present disclosure. For example, the above description of the configuration, functions, actions, and effects is an explanation of an example of the configuration, functions, actions, and effects of the parts related to the technology of the present disclosure. Therefore, it goes without saying that unnecessary parts may be deleted, new elements may be added, or replacements may be made to the above-described description and illustrations within the scope of the gist of the technology of the present disclosure. Furthermore, to avoid confusion and facilitate understanding of the parts related to the technology of the present disclosure, the above-described description and illustrations omit explanations of common technical knowledge that do not require particular explanation to enable the implementation of the technology of the present disclosure.
[0189] All publications, patent applications, and technical standards mentioned in this specification are herein incorporated by reference to the same extent as if each individual publication, patent application, or technical standard was specifically and individually indicated to be incorporated by reference.
[0190] [Explanation of symbols]
[0191] 10, 210, 310, 410 Data Processing Systems 12 Data Processing Device 14 Smart Devices 214 Smart Glasses 314 Headset-type terminal 414 Robot
Claims
1. a generator for generating an authentication word; a communication unit that communicates the authentication word generated by the generation unit to a purchaser; a voiceprint authentication unit that performs voiceprint authentication of the purchaser's voice when the purchaser speaks the authentication word transmitted by the transmission unit; a payment unit that performs payment based on the result of authentication by the voiceprint authentication unit. A system characterized by:
2. The generation unit Randomly generate authentication words The system of claim 1 .
3. The transmission unit is Speak the generated authentication word to the buyer The system of claim 1 .
4. The voiceprint authentication unit Analyze the buyer's voice and compare it with pre-registered voiceprint data The system of claim 1 .
5. The settlement unit If voice authentication is successful, payment is made. The system of claim 1 .
6. The settlement unit Providing payment methods that can be used by people who cannot own a smartphone, or who are visually impaired or elderly, etc. The system of claim 1 .
7. The generation unit Estimate buyer sentiment and adjust the difficulty of authentication words based on the estimated buyer sentiment. The system of claim 1 .
8. The generation unit Analyze the purchaser's past authentication history and generate optimal authentication words The system of claim 1 .
Citation Information
Patent Citations
Persona chatbot control method and system
JP2022180282A