System
The system uses smart glasses to address language barriers and facilitate product location, currency conversion, and price comparison, enhancing the shopping experience abroad by providing real-time voice guidance and efficient navigation.
Patent Information
- Application Number
- JP2024118182
- Authority / Receiving Office
- JP · JP
- Patent Type
- Applications
- Current Assignee / Owner
- Filing Date
- 2024-07-23
- Publication Date
- 2026-02-04
AI Technical Summary
Shopping abroad poses challenges such as language barriers, locating products, currency conversion, and price comparisons, which are particularly stressful for non-multilingual consumers, making the shopping experience less satisfying.
A system that includes voice input acquisition, voice recognition, real-time location identification, product database access, route generation, currency conversion using external APIs, and price comparison across multiple online stores, all provided as voice guidance through smart glasses.
The system effectively addresses language barriers and facilitates efficient product location, currency conversion, and price comparison, enhancing the shopping experience abroad by providing intuitive and convenient guidance.
Smart Images

Figure 2026017400000001_ABST
Abstract
Description
[Technical Field]
[0001] The technology of the present disclosure relates to a system. [Background technology]
[0002] Patent document 1 discloses a persona chatbot control method performed by at least one processor, the method including the steps of receiving a user utterance, adding the user utterance to a prompt including an instruction sentence related to a description of the chatbot character, encoding the prompt, and inputting the encoded prompt into a language model to generate a chatbot utterance in response to the user utterance. [Prior art documents] [Patent documents]
[0003] [Patent Document 1] Japanese Patent Publication No. 2022-180282 Summary of the Invention [Problem to be solved by the invention]
[0004] Shopping abroad poses many challenges, including language barriers, locating products, and currency conversion. This can be particularly stressful for non-multilingual consumers and shoppers who prioritize price comparisons. Additionally, finding the desired product section in a large store can be difficult, reducing satisfaction with the shopping experience. The present invention aims to resolve these challenges and improve the shopping experience abroad. [Means for solving the problem]
[0005] The present invention provides a system including a means for acquiring a user's voice input, a means for converting the acquired voice input into text using voice recognition technology, a means for analyzing the user's request from the converted text, a means for acquiring product location information from a database based on the analysis results, and a means for providing the acquired location information as voice guidance.
[0006] Furthermore, a system is provided that includes a means for identifying a user's current location in real time, a means for generating route information to a product section based on the identified current location, and a means for providing the generated route information as voice guidance.A system is also provided that includes a means for obtaining the latest exchange rate from an external API, converting product prices into the user's home currency using the obtained exchange rate, and a means for providing the converted price information by voice.
[0007] Additionally, the system includes a means for querying price databases of multiple online stores, acquiring and comparing prices from each online store, and providing the lowest price and that information as voice guidance. The system also includes a means for analyzing a user's movement data to identify their current location, and a means for providing real-time guidance to the product section based on the identified current location. This effectively solves many of the challenges associated with shopping abroad.
[0008] "User" refers to an individual who wears smart glasses and uses voice input.
[0009] "Voice input" refers to the method by which users input information by speaking into smart glasses.
[0010] "Voice recognition technology" refers to the technology that converts acquired voice data into text.
[0011] "Database" refers to an information aggregation system that stores product location and price information and can be accessed as needed.
[0012] "Voice guidance" refers to the technology or results of conveying textual information to users as voice.
[0013] "Current location" refers to information indicating where the customer is within the store.
[0014] "Real-time" refers to the near-simultaneous acquisition, processing, and provision of data.
[0015] "Route information" refers to information that indicates the route from the user's current location to the desired product section.
[0016] "Exchange rate" refers to a number that indicates the exchange rate between different currencies.
[0017] An "external API" refers to an interface for accessing external services and data.
[0018] "Price Database" means a system that maintains and makes accessible pricing information for products.
[0019] "Online store" refers to a website or platform that sells products over the Internet.
[0020] "Product Identifier" means a unique code or identification that distinguishes a particular product from other products.
[0021] "Lowest price information" refers to information about the lowest price among multiple prices. [Brief explanation of the drawings]
[0022] [Figure 1] 1 is a conceptual diagram showing an example of the configuration of a data processing system according to a first embodiment. [Figure 2] 1 is a conceptual diagram showing an example of main functions of a data processing device and a smart device according to a first embodiment. [Figure 3]FIG. 10 is a conceptual diagram showing an example of the configuration of a data processing system according to a second embodiment. [Figure 4] FIG. 10 is a conceptual diagram showing an example of main functions of a data processing device and smart glasses according to a second embodiment. [Figure 5] FIG. 10 is a conceptual diagram showing an example of the configuration of a data processing system according to a third embodiment. [Figure 6] FIG. 11 is a conceptual diagram showing an example of main functions of a data processing device and a headset-type terminal according to a third embodiment. [Figure 7] FIG. 10 is a conceptual diagram showing an example of the configuration of a data processing system according to a fourth embodiment. [Figure 8] FIG. 10 is a conceptual diagram showing an example of main functions of a data processing device and a robot according to a fourth embodiment. [Figure 9] 1 shows an emotion map onto which multiple emotions are mapped. [Figure 10] 1 shows an emotion map onto which multiple emotions are mapped. [Figure 11] FIG. 3 is a sequence diagram showing a processing flow of the data processing system according to the first embodiment. [Figure 12] FIG. 10 is a sequence diagram showing the flow of processing in the data processing system in Application Example 1. [Figure 13] FIG. 10 is a sequence diagram showing the flow of processing in the data processing system according to the second embodiment when an emotion engine is combined. [Figure 14] FIG. 10 is a sequence diagram showing the flow of processing in the data processing system in Application Example 2 when an emotion engine is combined. DETAILED DESCRIPTION OF THE INVENTION
[0023] An example of an embodiment of a system according to the technology of the present disclosure will be described below with reference to the accompanying drawings.
[0024] First, the terms used in the following description will be explained.
[0025] In the following embodiments, a coded processor (hereinafter simply referred to as a "processor") may be a single arithmetic device or a combination of multiple arithmetic devices. Furthermore, a processor may be a single type of arithmetic device or a combination of multiple types of arithmetic devices. Examples of arithmetic devices include a CPU (Central Processing Unit), a GPU (Graphics Processing Unit), a GPGPU (General-Purpose computing on Graphics Processing Units), and an APU (Accelerated Processing Unit).
[0026] In the following embodiments, a coded RAM (Random Access Memory) is a memory in which information is temporarily stored and is used as a working memory by a processor.
[0027] In the following embodiments, the coded storage is one or more non-volatile storage devices that store various programs, various parameters, etc. Examples of non-volatile storage devices include flash memory (SSD (Solid State Drive)), magnetic disks (e.g., hard disks), and magnetic tapes.
[0028] In the following embodiments, a communication I / F (Interface) with a symbol is an interface including a communication processor, an antenna, etc. The communication I / F controls communication between multiple computers. Examples of communication standards applied to the communication I / F include wireless communication standards including 5G (5th Generation Mobile Communication System), Wi-Fi (registered trademark), Bluetooth (registered trademark), etc.
[0029] In the following embodiments, "A and / or B" is synonymous with "at least one of A and B." In other words, "A and / or B" means that it may be only A, only B, or a combination of A and B. Furthermore, in this specification, the same concept as "A and / or B" is also applied when three or more things are expressed connected by "and / or."
[0030] [First embodiment]
[0031] FIG. 1 shows an example of the configuration of a data processing system 10 according to the first embodiment.
[0032] 1, a data processing system 10 includes a data processing device 12 and a smart device 14. An example of the data processing device 12 is a server.
[0033] The data processing device 12 includes a computer 22, a database 24, and a communication I / F 26. The computer 22 is an example of a "computer" according to the technology of the present disclosure. The computer 22 includes a processor 28, a RAM 30, and a storage 32. The processor 28, the RAM 30, and the storage 32 are connected to a bus 34. The database 24 and the communication I / F 26 are also connected to the bus 34. The communication I / F 26 is connected to a network 54. Examples of the network 54 include a WAN (Wide Area Network) and / or a LAN (Local Area Network).
[0034] The smart device 14 includes a computer 36, a reception device 38, an output device 40, a camera 42, and a communication I / F 44. The computer 36 includes a processor 46, a RAM 48, and a storage 50. The processor 46, the RAM 48, and the storage 50 are connected to a bus 52. The reception device 38, the output device 40, and the camera 42 are also connected to the bus 52.
[0035] The reception device 38 includes a touch panel 38A, a microphone 38B, and the like, and receives user input. The touch panel 38A detects contact with an indicator (for example, a pen or a finger) to receive user input by the touch of the indicator. The microphone 38B detects the user's voice to receive user input by voice. The control unit 46A transmits data indicating the user input received by the touch panel 38A and the microphone 38B to the data processing device 12. In the data processing device 12, the specific processing unit 290 acquires the data indicating the user input.
[0036] The output device 40 includes a display 40A and a speaker 40B, and presents data to the user 20 by outputting the data in a form of expression that the user 20 can perceive (for example, audio and / or text). The display 40A displays visible information such as text and images in accordance with instructions from the processor 46. The speaker 40B outputs audio in accordance with instructions from the processor 46. The camera 42 is a compact digital camera equipped with an optical system including a lens, aperture, and shutter, and an imaging element such as a CMOS (Complementary Metal-Oxide-Semiconductor) image sensor or a CCD (Charge Coupled Device) image sensor.
[0037] The communication I / F 44 is connected to a network 54. The communication I / Fs 44 and 26 control the exchange of various information between the processor 46 and the processor 28 via the network 54.
[0038] FIG. 2 shows an example of the main functions of the data processing device 12 and the smart device 14.
[0039] 2, in the data processing device 12, a specific process is performed by the processor 28. A specific processing program 56 is stored in the storage 32. The specific processing program 56 is an example of a "program" according to the technology of the present disclosure. The processor 28 reads the specific processing program 56 from the storage 32 and executes the read specific processing program 56 on the RAM 30. The specific process is realized by the processor 28 operating as a specific processing unit 290 in accordance with the specific processing program 56 executed on the RAM 30.
[0040] The storage 32 stores a data generation model 58 and an emotion identification model 59. The data generation model 58 and the emotion identification model 59 are used by the identification processing unit 290.
[0041] In the smart device 14, the processor 46 performs the reception output process. The storage 50 stores a reception output program 60. The reception output program 60 is used in conjunction with the specific processing program 56 by the data processing system 10. The processor 46 reads the reception output program 60 from the storage 50 and executes the read reception output program 60 on the RAM 48. The reception output process is realized by the processor 46 operating as the control unit 46A in accordance with the reception output program 60 executed on the RAM 48.
[0042] Next, a description will be given of the specific processing performed by the specific processing unit 290 of the data processing device 12. In the following description, the data processing device 12 will be referred to as a "server" and the smart device 14 will be referred to as a "terminal."
[0043] The present invention is a voice guidance system using smart glasses that solves problems such as language barriers when shopping abroad, product location, currency conversion, price comparison, etc. Hereinafter, an embodiment of the present invention will be described in detail.
[0044] Voice guidance function
[0045] Device (smart glasses)
[0046] 1. The user asks the smart glasses, "Where is the shampoo shelf?"
[0047] 2. The smart glasses capture voice input via a built-in microphone and convert the voice data into a digital format.
[0048] 3. Send the digital audio data to the server.
[0049] server
[0050] 1. Receives digital voice data and converts it into text using speech recognition technology.
[0051] 2. Parse the user's request (in this case, the shelf location of shampoo) from the converted text.
[0052] 3. Obtain the location information of the shampoo shelves from the database and generate text for voice guidance.
[0053] 4. The generated voice guidance text is sent to the terminal.
[0054] Terminal
[0055] 1. Receive the voice guidance text sent from the server and use the voice output device to instruct the user, "Go 10 meters to the right and turn left at the next shelf."
[0056] Real-time location information function
[0057] Terminal
[0058] 1. Determine the user's current location using built-in sensors (GPS, WiFi triangulation, etc.).
[0059] 2. The determined current location data is sent to the server.
[0060] server
[0061] 1. Receive current location data and determine the user's location based on that data.
[0062] 2. Obtain the location information of the desired product from the database and generate route information from the current location to the destination.
[0063] 3. The generated route information is sent to the terminal.
[0064] Terminal
[0065] 1. Route information sent from the server is provided to the user via voice.
[0066] Currency conversion function
[0067] Terminal
[0068] 1. Scan the product barcode to get price information.
[0069] 2. Send price information and user's home currency information to the server.
[0070] server
[0071] 1. Receive price information and get the latest exchange rates using an external API.
[0072] 2. Convert the product price into the user's home currency using the obtained exchange rate.
[0073] 3. The converted price information is sent to the terminal.
[0074] Terminal
[0075] 1. The converted price information is notified to the user via voice, such as "The price of this shampoo is 500 yen."
[0076] Price comparison feature
[0077] Terminal
[0078] 1. Scan the product barcode to get price information.
[0079] 2. Send the price information and product identifier to the server.
[0080] server
[0081] 1. Receives price information and product identifiers and queries price databases from multiple online stores.
[0082] 2. Get prices from each online store and identify the lowest price.
[0083] 3. Send the cheapest price information to the terminal.
[0084] Terminal
[0085] 1. The lowest price information is announced to the user via voice, saying, "It's on sale online for 450 yen."
[0086] Specific examples
[0087] If a user is in a supermarket in France and asks the smart glasses, "Where is the shampoo shelf?", the smart glasses will respond with, "Go 10 meters to the right and turn left at the next shelf." After reaching the shampoo section, the user picks up an item and wants to check the price. They can ask, "How much is this shampoo?" and the smart glasses will respond, "This shampoo costs 5 euros (approximately 620 yen)." Furthermore, to compare the price with online prices, they can ask, "How much does this shampoo cost online?" and the smart glasses will respond, "It's selling for 4.5 euros online." Through this process, users can enjoy an efficient and comfortable shopping experience.
[0088] The processing flow will be explained below.
[0089] Voice guidance function
[0090] Specific processing steps from voice input to guidance
[0091] Step 1:
[0092] The user asks the smart glasses, "Where is the shampoo shelf?"
[0093] Step 2:
[0094] The terminal uses a built-in microphone to capture the user's voice and converts the voice into digital data.
[0095] Step 3:
[0096] The terminal transmits the converted digital audio data to the server.
[0097] Step 4:
[0098] The server receives the digital voice data and converts the voice data into text using voice recognition technology.
[0099] Step 5:
[0100] The server parses the text and understands the user's request (in this case, a request to know the shelf location of shampoo).
[0101] Step 6:
[0102] The server retrieves the location information of the shampoo shelves from the database.
[0103] Step 7:
[0104] Based on the acquired location information, the server generates text for voice guidance such as, "Go 10 meters to the right and turn left at the next shelf."
[0105] Step 8:
[0106] The server transmits the generated voice guidance text to the terminal.
[0107] Step 9:
[0108] The device converts the received voice guidance text into speech and begins providing guidance through the speaker.
[0109] Step 10:
[0110] The user moves as instructed.
[0111] Real-time location information function
[0112] A processing step to guide you from your current location to the product location
[0113] Step 1:
[0114] The user speaks to the smart glasses, saying, "Tell me where I am."
[0115] Step 2:
[0116] The device determines the user's current location using built-in sensors (e.g., GPS and WiFi triangulation).
[0117] Step 3:
[0118] The terminal transmits the identified current location data to the server.
[0119] Step 4:
[0120] The server receives the location data and confirms the user's current location.
[0121] Step 5:
[0122] The server retrieves route information to the product section from the database.
[0123] Step 6:
[0124] The server generates route information from the current location to the destination product section.
[0125] Step 7:
[0126] The server transmits the generated route information to the terminal.
[0127] Step 8:
[0128] The terminal converts the received route information into text for voice guidance.
[0129] Step 9:
[0130] The device converts the voice guidance text into speech and begins providing guidance through the speaker.
[0131] Step 10:
[0132] The user moves along the guided route.
[0133] Currency conversion function
[0134] Processing steps to convert product prices into your home currency
[0135] Step 1:
[0136] Users scan the barcode of the product they pick up with the smart glasses.
[0137] Step 2:
[0138] The terminal scans the barcode and obtains the product's price information.
[0139] Step 3:
[0140] The terminal transmits product price information and the user's home currency information to the server.
[0141] Step 4:
[0142] The server receives the price information and uses an external API to get the latest exchange rates.
[0143] Step 5:
[0144] The server uses the obtained exchange rate to convert the product price into the user's home currency.
[0145] Step 6:
[0146] The server transmits the converted price information to the terminal.
[0147] Step 7:
[0148] The terminal converts the received converted price information into text for voice output.
[0149] Step 8:
[0150] The device converts the text to speech and announces the price through the speaker.
[0151] Step 9:
[0152] The user receives the price information sent via the smart glasses.
[0153] Price comparison feature
[0154] Processing steps to compare in-store and online prices
[0155] Step 1:
[0156] Users scan the barcode of the product they pick up with the smart glasses.
[0157] Step 2:
[0158] The terminal scans the barcode and obtains the product's price information.
[0159] Step 3:
[0160] The terminal transmits the price information and the product identifier to the server.
[0161] Step 4:
[0162] The server receives the price information and the product identifier and queries the online store's price database.
[0163] Step 5:
[0164] The server retrieves the prices from each online store and identifies the lowest price.
[0165] Step 6:
[0166] The server transmits the cheapest price information to the terminal.
[0167] Step 7:
[0168] The terminal converts the received information about the cheapest price into text for voice output.
[0169] Step 8:
[0170] The device converts the text to speech and announces price information through the speaker.
[0171] Step 9:
[0172] The user receives price information from the smart glasses and decides to purchase.
[0173] Through the above processing steps, this system supports users in shopping abroad efficiently and comfortably.
[0174] Example 1
[0175] Next, a description will be given of Example 1. In the following description, the data processing device 12 will be referred to as a "server" and the smart device 14 will be referred to as a "terminal."
[0176] In today's global society, shopping in a multilingual environment can be extremely difficult for users. When shopping abroad, language barriers make it particularly difficult to accurately determine product locations. It's also difficult to convert prices displayed in local currencies into one's own currency, or to quickly compare prices across different stores. A system that solves these problems is needed.
[0177] The specific processing by the specific processing unit 290 of the data processing device 12 in the first embodiment is realized by the following means.
[0178] In this invention, the server includes: means for acquiring a user's voice input; means for converting the acquired voice input into digital format; means for transmitting the converted digital voice input to the server; means for converting the voice input into text on the server using voice recognition technology; means for analyzing the user's request from the converted text; means for acquiring product location information from a database based on the analysis results; means for converting the acquired location information into text for voice guidance and transmitting it to the terminal; and means for converting the transmitted text for voice guidance into speech and providing guidance to the user. This allows users to easily find products even in different language environments. Furthermore, by providing functions such as real-time location information identification and route guidance, currency conversion using the latest exchange rates, and price comparison across multiple stores, the server achieves an efficient and convenient shopping experience.
[0179] The "means for acquiring user's voice input" is a combination of hardware and software that recognizes the voice uttered by the user and that the terminal acquires as digital data.
[0180] The "means for converting to digital form" refers to an audio signal processing device and conversion algorithm for converting analog audio signals into digital data.
[0181] The "means for transmitting to a server" refers to a communication module and protocol that allows the terminal to transfer digital data over the Internet to a remote server.
[0182] "Means for converting speech to text using speech recognition technology" refers to software and algorithms that analyze speech data on a server and convert it into corresponding text.
[0183] The "means for analyzing user requests" is a natural language processing system that extracts and analyzes the user's intentions and requests from the text obtained by speech recognition.
[0184] The "means for acquiring product location information from a database" is a query processing system for searching and acquiring product placement information from a database based on the analyzed user request.
[0185] The "means for converting the information into text for voice guidance and sending it to the terminal" is a server-side function for converting the acquired information into text in a format that is easy for the user to understand and sending this text to the terminal.
[0186] The "means for converting the text information into audio and providing the user with guidance" refers to the software and hardware within the terminal for playing back the received text information as audio.
[0187] "Means of real-time identification" refers to a system that instantly identifies a user's current location using sensor technology such as GPS and Wi-Fi triangulation.
[0188] The "means for generating route information" refers to algorithms and software for calculating a route from the specified current position to the destination and generating that information.
[0189] The "means for obtaining the latest exchange rates from an external data source" is an interface for obtaining the latest currency exchange information from an external financial data provider service.
[0190] The "means for converting the product price using the exchange rate" is a calculation algorithm for converting the product price into the user's home currency using the acquired exchange rate information.
[0191] MODE FOR CARRYING OUT THE INVENTION
[0192] The present invention is a voice guidance system using smart glasses that solves problems such as language barriers when shopping abroad, product location, currency conversion, price comparison, etc. Hereinafter, an embodiment of the present invention will be described in detail.
[0193] Overall system overview
[0194] This system is composed of a combination of smart glasses (terminals), a cloud server, a database, and an external API. The smart glasses receive voice input from the user and send it to the cloud server. The cloud server uses voice recognition technology to convert the voice data into text, and then analyzes the user's request through natural language processing. Based on the analysis results, it provides information such as product location, currency conversion, and price comparison.
[0195] Hardware and software used
[0196] Device (smart glasses)
[0197] Built-in microphone: Captures the user's voice.
[0198] Communication module: Sends digital data to a cloud server (e.g., Wi-Fi, mobile data).
[0199] Voice output device: Converts guided text into voice.
[0200] Sensors: Determine current location (e.g. GPS, Wi-Fi triangulation).
[0201] server
[0202] Speech recognition software: Google Cloud Speech-to-Text API.
[0203] Natural language processing engine: Uses SpaCy or similar.
[0204] Database: Uses MySQL to manage product location and price information.
[0205] External API: Uses Open Exchange Rates API for currency conversion.
[0206] Program processing
[0207] 1. Voice guidance function
[0208] The user asks the smart glasses, "Where is the shampoo shelf?"
[0209] The smart glasses pick up sound through a built-in microphone and convert it into a digital format.
[0210] The digital audio data is transmitted to a cloud server.
[0211] The server uses speech recognition software to convert the voice data into text.
[0212] The text is analyzed using a natural language processing engine to understand the request.
[0213] Retrieve product location information from the database and generate guide text.
[0214] The guidance text is sent to the smart glasses and the user is guided by voice.
[0215] 2. Real-time location information function
[0216] The smart glasses use built-in sensors to determine the user's current location.
[0217] The current location data is sent to a cloud server.
[0218] The server generates route information to the product section based on the current location data.
[0219] The generated route information is sent to smart glasses and guidance is provided via voice.
[0220] 3. Currency conversion function
[0221] Smart glasses scan product barcodes to obtain price information.
[0222] Price information and the user's home currency information are sent to the server.
[0223] The server uses an external API to get the latest exchange rates.
[0224] Convert product prices into the user's home currency using exchange rates.
[0225] The converted price information is sent to the smart glasses and notified via voice.
[0226] 4. Price comparison feature
[0227] The smart glasses scan the product's barcode and retrieve price information.
[0228] Send the price information and product identifier to the server.
[0229] The server queries a price database of multiple online stores.
[0230] The lowest price information is identified and sent to the smart glasses.
[0231] The cheapest price information is notified to the user by voice.
[0232] Specific examples
[0233] If a user is in a supermarket in France and asks the smart glasses, "Where is the shampoo shelf?", the smart glasses will respond with, "Go 10 meters to the right and turn left at the next shelf." After reaching the shampoo section, the user picks up an item and wants to check the price. They can ask, "How much is this shampoo?" and the smart glasses will respond, "This shampoo costs 5 euros (approximately 620 yen)." Furthermore, to compare the price with online prices, they can ask, "How much does this shampoo cost online?" and the smart glasses will respond, "It's selling for 4.5 euros online." Through this process, users can enjoy an efficient and comfortable shopping experience.
[0234] Example prompts to input to the generative AI model
[0235] Please explain in detail the process by which a user queries the smart glasses for product location, product price, currency conversion, and price comparison. Please also include specific hardware, software, data processing, and calculation methods.
[0236] The flow of the identification process in the first embodiment will be described with reference to FIG.
[0237] Step 1:
[0238] A user speaks to the smart glasses, saying, "Where is the shampoo shelf?" The user's voice is picked up by the smart glasses' built-in microphone. This voice input is an analog signal, so it is converted into digital format. Specifically, A / D conversion is performed to convert the voice waveform into a digital signal. The input is the user's voice signal, and the output is digital voice data.
[0239] Step 2:
[0240] The device transmits digital audio data to a cloud server via Wi-Fi or a mobile data network using a communication module. The input is digital audio data, and the output is confirmation of data transmission to the server.
[0241] Step 3:
[0242] The server receives the digital voice data and converts it into text using speech recognition technology. The Google Cloud Speech-to-Text API is used. The server analyzes the voice data and generates corresponding text data. The input is digital voice data, and the output is the converted text data.
[0243] Step 4:
[0244] The server sends the converted text data to a natural language processing engine to analyze the user's request. For example, SpaCy can be used to extract the meaning of the text. The input is text data, and the output is the analysis result of the user's request ("where shampoo is on the shelf"). Specifically, the keyword "shampoo" is extracted from the text, and the request content is determined based on that.
[0245] Step 5:
[0246] The server retrieves product location information from the database based on the analysis results. In this example, it queries a MySQL database to retrieve the location information of the shampoo shelf. The input is the analysis result of the user request, and the output is the product location information. The database query is executed and the location information is returned.
[0247] Step 6:
[0248] The server converts the acquired location information into text for voice guidance and sends it to the terminal. An algorithm for generating voice guidance text is applied to generate text such as "Go 10 meters to the right and turn left at the next shelf." The input is the product location information, and the output is the text for voice guidance. The generated text is sent to the terminal as an HTTP response.
[0249] Step 7:
[0250] The device receives the text for voice guidance sent from the server. The received text is then used to provide guidance to the user using a voice output device. Specifically, the text is converted to speech using the Google Text-to-Speech API. The input is the text for voice guidance, and the output is the voice guidance. "Go 10 meters to the right and turn left at the next shelf," is the voice guidance that comes out of the smart glasses.
[0251] (Application example 1)
[0252] Next, a description will be given of Application Example 1. In the following description, the data processing device 12 will be referred to as a "server" and the smart device 14 will be referred to as a "terminal."
[0253] The present invention aims to solve problems associated with shopping abroad, such as language barriers, product location, currency conversion, and price comparisons. Specifically, the problem is to provide a support system that enables users to intuitively and efficiently find products in physical stores, understand prices, and purchase them at the most advantageous prices.
[0254] The specific processing by the specific processing unit 290 of the data processing device 12 in the application example 1 is realized by the following means.
[0255] In this invention, the server includes means for acquiring a user's voice input, means for converting the acquired voice input into text using voice recognition technology, means for analyzing the user's request from the converted text, means for acquiring product location information from a database based on the analysis result, means for providing the acquired location information as voice guidance, means for converting product prices into the user's home currency using a currency conversion function, and means for acquiring price information from multiple online stores using a price comparison function to identify the lowest price. This allows users to easily find products even in stores in a foreign country, and also makes it easy to convert currencies and compare prices.
[0256] "User voice input" refers to voice commands issued by a user to a device.
[0257] "Speech recognition technology" refers to machine learning and natural language processing technology for converting speech into text data.
[0258] "Means for converting to text" refers to a device or software that performs a process of converting captured voice input into digital data and then parsing that digital data into text.
[0259] "Means for analyzing user requests" refers to a device or software that processes voice commands converted into text to understand the information or action the user is seeking.
[0260] "Means for obtaining product location information from a database" refers to a device or software for searching and obtaining information about the product location stored in a database.
[0261] "Means for providing audio guidance" refers to a device or software that provides audio guidance to the user based on analyzed information and acquired data.
[0262] "Currency Conversion Facility" means a device or software that performs a process to convert product prices displayed in one currency into another currency selected by the user.
[0263] "Price comparison tool" means a device or software that performs a process to obtain and compare price information from multiple sources or online stores for a particular product to identify the most favorable price.
[0264] "Means for determining current location in real time" refers to devices or software that obtain a user's current location in real time using technologies such as GPS or WiFi triangulation.
[0265] The "means for generating route information" refers to a device or software for calculating the optimal route from the user's current location to the location of the desired product and generating the route information.
[0266] "Means for obtaining the latest exchange rates from an external API" refers to a device or software for obtaining exchange rate information from an external server via the Internet.
[0267] The present invention relates to a system that uses a voice guidance system in a specified target environment to provide users with various information such as product location information, currency conversion, price comparison, etc. Specific embodiments of the system are described below.
[0268] Generating a Program
[0269] The system consists of the following main functions:
[0270] 1. Voice Input Capture and Recognition
[0271] 2. Location information acquisition and guidance
[0272] 3. Currency conversion function
[0273] 4. Price comparison feature
[0274] Program processing explanation
[0275] Voice Input Capture and Recognition
[0276] The device (e.g., smart glasses) captures the user's voice input with a microphone and converts the voice data into a digital format. The converted digital voice data is sent to the server using the os and wave libraries. The server receives the voice data via a flask-based endpoint and converts it into text using the pydub and speech_recognition libraries.
[0277] Next, a natural language processing library such as nltk is used to analyze the user's request, and the next steps are taken based on the analysis results.
[0278] Location information acquisition and guidance
[0279] The device acquires its current location in real time using its built-in GPS and WiFi triangulation, and sends it to the server. The server processes the location data using the geopy library and generates the optimal route to the product section based on the user's current location. This route information is then provided as audio guidance.
[0280] Currency conversion function
[0281] When a user scans a product's barcode, the device retrieves the price information and sends it to the server. The server then uses an external API to retrieve the latest exchange rate and converts the product price based on the obtained rate. The converted price information is retrieved using the requests library and notified to the user via voice.
[0282] Price comparison feature
[0283] When a user scans a product's barcode, the device sends the price information to a server, which queries a price database of multiple online stores to identify the lowest price, which is then provided to the user via voice.
[0284] Specific examples
[0285] If a user is in a supermarket in France and asks the smart glasses, "Where is the shampoo shelf?", the smart glasses will respond with, "Go 10 meters to the right and turn left at the next shelf." After reaching the shampoo section, the user picks up an item and wants to check the price. They can ask, "How much is this shampoo?" and the smart glasses will respond, "This shampoo costs 5 euros (approximately 620 yen)." Furthermore, to compare the price with online prices, they can ask, "How much does this shampoo cost online?" and the smart glasses will respond, "It's selling for 4.5 euros online." Through this process, users can enjoy an efficient and comfortable shopping experience.
[0286] Prompt Sentence Examples
[0287] "This smartglasses-based shopping app allows users to easily navigate in-store, convert currencies, compare prices, and more, all by voice. For example, if a user is in a supermarket in France and asks the smartglasses, 'Where is the shampoo shelf?', the smartglasses will respond with, 'Go 10 meters to the right and then turn left at the next shelf.' Please generate a program that achieves this function."
[0288] The flow of the specific processing in the application example 1 will be described with reference to FIG.
[0289] Step 1:
[0290] The device (smart glasses) captures the user's voice input with a microphone. The user speaks, "Where is the shampoo shelf?", and this voice signal is acquired. The voice input is converted into digital voice data using the OS and Wave libraries. The output is digital voice data.
[0291] Step 2:
[0292] The device sends digital audio data to the server. This is done using the requests library. The input is the digital audio data, and the output is a request to the server to send the audio data.
[0293] Step 3:
[0294] The server receives the received digital audio data at the flask endpoint and starts analyzing the data. The input is the digital audio data. The server converts it into text using the pydub and speech_recognition libraries. The output is the converted text data.
[0295] Step 4:
[0296] The server analyzes the converted text data using a natural language processing library such as nltk to identify the user's request. The converted text data is input, and the user's request extracted through analysis is output. Specifically, "the location of the shampoo shelf" is identified as the user's request.
[0297] Step 5:
[0298] The server retrieves the shampoo shelf location information from the database. The input is the user's request (the shampoo shelf location), and the output is the product location information. This location information indicates where the product is located in the store.
[0299] Step 6:
[0300] The server calculates the location using the geopy library based on the product location information and generates route information. The input is the product location information and the user's current location, and the output is route information. Specifically, the guidance will be something like "Go 10 meters to the right and turn left at the next shelf."
[0301] Step 7:
[0302] The server converts the generated route information into text for voice guidance and sends it to the terminal. The input is route information, and a process is performed to convert this information into text for voice guidance. The output is text for voice guidance.
[0303] Step 8:
[0304] The terminal notifies the user of the voice guidance text received from the server using a voice output device. The input is the voice guidance text, and the output is the voice guidance ("Go 10 meters to the right and turn left at the next shelf").
[0305] Step 9:
[0306] When a user scans a product barcode, the terminal obtains the price information and sends it to the server. The input is the price information output by scanning the product barcode.
[0307] Step 10:
[0308] The server uses an external API to obtain the latest exchange rate based on the acquired price information. The input is the price information, and the output is the latest exchange rate.
[0309] Step 11:
[0310] The server converts the price into the user's home currency based on the exchange rate. The input is the acquired exchange rate and price information, and the output is the converted price information.
[0311] Step 12:
[0312] The server converts the converted price information into text for voice guidance and sends it to the terminal. The input is the converted price information, and the output is the text for voice guidance.
[0313] Step 13:
[0314] The terminal notifies the user of the voice guidance text received from the server using a voice output device. The input is the voice guidance text, and the output is the voice guidance ("The price of this shampoo is XX yen").
[0315] Step 14:
[0316] When a user speaks "What is the price of this shampoo online?", the device sends this speech input to the server. The input is the user's speech input, and the output is digital voice data.
[0317] Step 15:
[0318] The server analyzes the received voice data to identify it as a price comparison request and queries a price database of multiple online stores. The input is the analyzed request, and the output is the price information of each online store.
[0319] Step 16:
[0320] The server identifies the cheapest price information, converts it into a voice guidance text, and sends it to the terminal. The input is the price information of the online store, and the output is a voice guidance text containing the cheapest price information.
[0321] Step 17:
[0322] The terminal notifies the user of the voice guidance text received from the server using a voice output device. The input is the voice guidance text, and the output is the voice guidance ("This item is sold online for XX yen").
[0323] Furthermore, an emotion engine that estimates the user's emotion may be combined. That is, the identification processing unit 290 may estimate the user's emotion using the emotion identification model 59 and perform identification processing using the user's emotion.
[0324] The present invention provides a system that combines a voice guidance system using smart glasses with an emotion engine that recognizes the user's emotions, providing a more personalized shopping experience. The following describes in detail an embodiment of the present invention.
[0325] Voice guidance function
[0326] Device (smart glasses)
[0327] 1. The user asks the smart glasses, "Where is the shampoo shelf?"
[0328] 2. The smart glasses capture voice input via a built-in microphone and convert the voice data into a digital format.
[0329] 3. Send the digital audio data to the server.
[0330] server
[0331] 1. Receive digital voice data and convert the voice data into text using voice recognition technology.
[0332] 2. Understand the user's request from the converted text (in this case, the request to know the shelf location of shampoo).
[0333] 3. Obtain the location information of the shampoo shelves from the database and generate text for voice guidance.
[0334] 4. Use the emotion engine to recognize emotions from the user's voice and adjust the guidance content.
[0335] 5. The adjusted voice guidance text is sent to the device.
[0336] Terminal
[0337] 1. Receive the voice guidance text sent from the server and use the voice output device to instruct the user, "Go 10 meters to the right and turn left at the next shelf."
[0338] Real-time location information function
[0339] Terminal
[0340] 1. Determine the user's current location using built-in sensors (GPS, WiFi triangulation, etc.).
[0341] 2. The determined current location data is sent to the server.
[0342] server
[0343] 1. Receive current location data and determine the user's location based on that data.
[0344] 2. Obtain the location information of the desired product from the database and generate route information from the current location to the destination.
[0345] 3. Use an emotion engine to recognize emotions from the user's voice and facial expressions and adjust the guidance content accordingly.
[0346] 4. The adjusted route information is sent to the terminal.
[0347] Terminal
[0348] 1. The route information sent from the server is converted into text for voice guidance, and guidance to the user begins.
[0349] Currency conversion function
[0350] Terminal
[0351] 1. Scan the product barcode to get price information.
[0352] 2. Send price information and user's home currency information to the server.
[0353] server
[0354] 1. Receive price information and get the latest exchange rates using an external API.
[0355] 2. Convert the product price into the user's home currency using the obtained exchange rate.
[0356] 3. Use an emotion engine to recognize emotions from the user's voice and facial expressions and adjust the content of price information notifications.
[0357] 4. The adjusted price information is sent to the terminal.
[0358] Terminal
[0359] 1. The converted price information is converted into text for voice output and announced through the speaker: "The price of this shampoo is 500 yen."
[0360] Price comparison feature
[0361] Terminal
[0362] 1. Scan the product barcode to get price information.
[0363] 2. Send the price information and product identifier to the server.
[0364] server
[0365] 1. Receives price information and product identifiers and queries price databases from multiple online stores.
[0366] 2. Get the prices from each online store and identify the cheapest price.
[0367] 3. Use an emotion engine to recognize emotions from the user's voice and facial expressions and adjust the content of the notification about the cheapest price.
[0368] 4. The adjusted lowest price information is sent to the terminal.
[0369] Terminal
[0370] 1. Convert the cheapest price information into text for voice output and announce through the speaker, "It's on sale online for 450 yen."
[0371] Use of emotion engine
[0372] server
[0373] 1. Receives voice data and facial expression data and recognizes the user's emotions using an emotion engine.
[0374] 2. Tailor voice prompts and pricing notifications based on the recognized emotion.
[0375] 3. The adjusted information is sent to the device.
[0376] Specific examples
[0377] If a user is in a supermarket in France and asks the smart glasses, "Where is the shampoo shelf?", the smart glasses will respond with, "Go 10 meters to the right and turn left at the next shelf." After reaching the shampoo section, the user picks up an item and asks, "How much is this shampoo?" to check the price. The smart glasses will respond, "This shampoo costs 5 euros (approximately 620 yen)." Furthermore, if the user asks, "How much does this shampoo cost online?" to compare it with online prices, the smart glasses will respond, "It's selling for 4.5 euros online." This invention utilizes an emotion engine to optimize the user experience, such as providing more attentive guidance if the user shows a confused expression.
[0378] The processing flow will be explained below.
[0379] Voice guidance function
[0380] Specific processing steps from voice input to guidance
[0381] Step 1:
[0382] The user asks the smart glasses, "Where is the shampoo shelf?"
[0383] Step 2:
[0384] The terminal uses a built-in microphone to capture the user's voice and converts the voice into digital data.
[0385] Step 3:
[0386] The terminal transmits the converted digital audio data to the server.
[0387] Step 4:
[0388] The server receives the digital voice data and converts the voice data into text using voice recognition technology.
[0389] Step 5:
[0390] The server parses the text and understands the user's request (in this case, a request to know the shelf location of shampoo).
[0391] Step 6:
[0392] The server retrieves the location information of the shampoo shelves from the database.
[0393] Step 7:
[0394] Based on the acquired location information, the server generates text for voice guidance such as, "Go 10 meters to the right and turn left at the next shelf."
[0395] Step 8:
[0396] The server sends the voice data to an emotion engine to analyze the user's emotion from the voice data.
[0397] Step 9:
[0398] The server receives the emotion recognition results from the emotion engine and adjusts the content and tone of the announcement based on the user's emotion.
[0399] Step 10:
[0400] The server transmits the adjusted voice guidance text to the terminal.
[0401] Step 11:
[0402] The device converts the received voice guidance text into speech and begins providing guidance through the speaker.
[0403] Step 12:
[0404] The user moves as instructed.
[0405] Real-time location information function
[0406] Specific processing steps to guide you from your current location to the product location
[0407] Step 1:
[0408] The user speaks to the smart glasses, saying, "Tell me where I am."
[0409] Step 2:
[0410] The device determines the user's current location using built-in sensors (e.g., GPS and WiFi triangulation).
[0411] Step 3:
[0412] The terminal transmits the identified current location data to the server.
[0413] Step 4:
[0414] The server receives the location data and determines the user's location based on it.
[0415] Step 5:
[0416] The server retrieves route information to the product section from the database.
[0417] Step 6:
[0418] The server generates route information from the current location to the destination product section.
[0419] Step 7:
[0420] The server sends the user's voice data and facial expression data to the emotion engine to analyze the user's emotions.
[0421] Step 8:
[0422] The server receives the emotion recognition results from the emotion engine and adjusts the content and tone of the route guidance based on the emotion.
[0423] Step 9:
[0424] The server transmits the adjusted route information to the terminal.
[0425] Step 10:
[0426] The terminal converts the received route information into text for voice guidance.
[0427] Step 11:
[0428] The terminal converts the voice guidance text into voice and starts providing guidance to the user through the speaker.
[0429] Step 12:
[0430] The user moves along the guided route.
[0431] Currency conversion function
[0432] Specific steps to convert product prices into your home currency
[0433] Step 1:
[0434] Users scan the barcode of the product they pick up with the smart glasses.
[0435] Step 2:
[0436] The terminal scans the barcode and obtains the product's price information.
[0437] Step 3:
[0438] The terminal transmits product price information and the user's home currency information to the server.
[0439] Step 4:
[0440] The server receives the price information and uses an external API to get the latest exchange rates.
[0441] Step 5:
[0442] The server uses the obtained exchange rate to convert the product price into the user's home currency.
[0443] Step 6:
[0444] The server sends the user's voice data and facial expression data to the emotion engine to analyze the user's emotions.
[0445] Step 7:
[0446] The server receives emotion recognition results from the emotion engine and adjusts the content and tone of price information notifications based on the emotion.
[0447] Step 8:
[0448] The server transmits the adjusted price information to the terminal.
[0449] Step 9:
[0450] The terminal converts the received price information into text for voice output.
[0451] Step 10:
[0452] The device converts the text for voice output into speech and announces through the speaker, "The price of this shampoo is 500 yen."
[0453] Step 11:
[0454] The user receives the price information sent via the smart glasses.
[0455] Price comparison feature
[0456] Specific steps to compare in-store and online prices
[0457] Step 1:
[0458] Users scan the barcode of the product they pick up with the smart glasses.
[0459] Step 2:
[0460] The terminal scans the barcode and obtains the product's price information.
[0461] Step 3:
[0462] The terminal transmits the price information and the product identifier to the server.
[0463] Step 4:
[0464] The server receives the price information and the product identifier and queries a price database of multiple online stores.
[0465] Step 5:
[0466] The server retrieves the prices from each online store and identifies the lowest price.
[0467] Step 6:
[0468] The server sends the user's voice data and facial expression data to the emotion engine for emotion analysis.
[0469] Step 7:
[0470] The server receives the emotion recognition result from the emotion engine and adjusts the content and tone of the notification of the cheapest price information based on the emotion recognition result.
[0471] Step 8:
[0472] The server transmits the adjusted lowest price information to the terminal.
[0473] Step 9:
[0474] The terminal converts the received information about the cheapest price into text for voice output.
[0475] Step 10:
[0476] The device converts the text to speech and announces through the speaker, "It's on sale online for 450 yen."
[0477] Step 11:
[0478] The user receives price information from the smart glasses and decides to purchase.
[0479] Through the above processing steps, this system supports users in shopping abroad efficiently and comfortably. By utilizing the emotion engine, guidance and information are provided according to the user's emotional state, providing a more personalized experience.
[0480] Example 2
[0481] Next, a description will be given of Example 2. In the following description, the data processing device 12 will be referred to as a "server" and the smart device 14 will be referred to as a "terminal."
[0482] The current shopping experience is burdensome for users, as it is difficult to locate products in physical stores. Furthermore, price conversion and comparison are time-consuming, making shopping especially complicated in foreign countries. Furthermore, conventional voice guidance systems do not take user emotions into account, so guidance information is not optimized according to the user's state. To solve these problems, a system that can recognize user emotions and provide more personalized guidance is needed.
[0483] The identification process by the identification processing unit 290 of the data processing device 12 in the second embodiment is realized by the following means. In this invention, the server includes a means for converting acquired voice input into text using voice recognition technology, a means for analyzing the user's request from the converted text, and a means for acquiring product location information from a database based on the analysis result, adjusting the information based on an emotion recognition engine, and providing voice guidance. This makes it possible to provide optimal guidance according to the user's emotions.
[0484] "User" refers to the person who uses this system and is the subject who performs voice input.
[0485] "Voice input" refers to voice data provided by a user through a voice input device such as a microphone.
[0486] "Speech recognition technology" refers to technology for converting voice data into text data, and refers to technology that uses, for example, voice signal processing algorithms or machine learning models.
[0487] "Text" is character string data converted using speech recognition technology.
[0488] "Request" refers to various questions and instructions that a user gives to the system.
[0489] "Analysis" refers to the process of understanding user requests and responding appropriately to them.
[0490] "Location information" is data that indicates the physical location of a particular product or location.
[0491] A "database" is an information management system for storing and managing location information and related data.
[0492] An "emotion recognition engine" is software or hardware that recognizes a user's emotions from their voice or facial expressions.
[0493] "Voice guidance" refers to guidance information provided to users via voice.
[0494] "Real-time" refers to near-instant processing or response.
[0495] "Current location" is data indicating the physical location where the user is currently located.
[0496] "Route information" is data that indicates the route and instructions from the current position to the destination.
[0497] An "external API" is an application program interface provided by a third party that is a means of connecting with other systems.
[0498] "Exchange rate" is data that indicates the exchange rate between one country's currency and another country's currency.
[0499] "Home currency" refers to the currency of the country or region to which the user belongs.
[0500] This invention is a voice guidance system using smart glasses that combines an emotion engine that recognizes the user's emotions to provide a more personalized shopping experience.
[0501] Voice guidance function
[0502] Device (smart glasses)
[0503] The user speaks to the smart glasses, saying, "Where is the shampoo shelf?" The smart glasses receive voice input via a built-in microphone and convert the voice data into a digital format. The converted digital voice data is then transmitted to the server via wireless communication (WiFi or Bluetooth).
[0504] server
[0505] The server converts the received digital voice data into text using voice recognition technology (e.g., technology using a voice signal processing algorithm or a machine learning model), then analyzes the user's request from the text (e.g., using natural language processing technology) and retrieves the location information of the corresponding product from a database (e.g., an information management system).
[0506] The acquired location information is adjusted by an emotion recognition engine (e.g., software or hardware for recognizing emotions from voice and facial expressions) to take into account the user's emotions. Finally, an adjusted voice guidance text is generated and sent to the device.
[0507] Device (smart glasses)
[0508] The voice guidance text sent from the server is received and the voice output device (speaker or earphone) is used to guide the user, for example, "Go 10 meters to the right and turn left at the next shelf."
[0509] Real-time location information function
[0510] Device (smart glasses)
[0511] The user's current location is determined using built-in sensors (e.g., GPS or WiFi triangulation), and the determined current location data is transmitted to a server.
[0512] server
[0513] The server identifies the user's location based on the received current location data. Next, it retrieves the desired product's location information from the database and generates route information from the user's current location to the destination. This route information is also adjusted by an emotion recognition engine, taking into account the user's emotions. Finally, the adjusted route information is sent to the device.
[0514] Device (smart glasses)
[0515] The route information sent from the server is converted into text for voice guidance, and guidance to the user is started.
[0516] Currency conversion function
[0517] Device (smart glasses)
[0518] When a user picks up a product and wants to check the price, the smart glasses scan the product barcode to obtain price information, which is then sent to the server along with the user's home currency.
[0519] server
[0520] The server obtains the latest exchange rate from an external API (e.g., an application program interface for obtaining exchange rates) and uses it to convert the product price into the user's home currency. The converted price information is adjusted by an emotion recognition engine and sent to the device.
[0521] Device (smart glasses)
[0522] The converted price information is converted into text for voice output to the user, and the user is notified through the speaker that "The price of this shampoo is 500 yen."
[0523] Price comparison feature
[0524] Device (smart glasses)
[0525] When a user wants to compare prices at different stores, they scan the product barcode to get the price information, which is then sent to the server along with the product identifier.
[0526] server
[0527] The server queries price databases of multiple online stores, retrieves prices from each online store, and identifies the cheapest price. This cheapest price information is also adjusted by the emotion recognition engine. Finally, the adjusted cheapest price information is sent to the device.
[0528] Device (smart glasses)
[0529] The lowest price information is converted into text for voice output and announced through the speaker: "It's on sale online for 450 yen."
[0530] Using the Emotion Recognition Engine
[0531] server
[0532] The server receives the voice data and facial expression data and uses an emotion recognition engine to recognize the user's emotions. Based on the recognized emotions, the voice guidance and price information notifications are adjusted. Finally, the adjusted information is sent to the device.
[0533] Prompt Sentence Examples
[0534] If a user is in a supermarket in France, they can ask their smart glasses to:
[0535] "Where's the shampoo shelf?"
[0536] "How much is this shampoo?"
[0537] "How much does this shampoo cost online?"
[0538] Entering these prompts into the system will provide corresponding guidance and pricing information.
[0539] As described above, the system combines multiple functions such as voice input, location information, currency conversion, price comparison, and emotion recognition to provide people with an advanced and personalized shopping experience.
[0540] The flow of the identification process in the second embodiment will be described with reference to FIG.
[0541] Voice guidance function
[0542] Step 1:
[0543] The device (smart glasses) receives the user's voice input. When a user speaks to the smart glasses, saying, "Where is the shampoo shelf?", this voice data is captured by the device's built-in microphone.
[0544] Input: User's voice data.
[0545] Output: Audio data.
[0546] Step 2:
[0547] The audio data acquired by the device (smart glasses) is converted into a digital format. Specifically, AD conversion is performed to convert the audio data into a digital signal.
[0548] Input: Analog audio data.
[0549] Output: Digital audio data.
[0550] Step 3:
[0551] The device (smart glasses) sends digital audio data to the server. The audio data is transferred to the server via wireless communication (WiFi or Bluetooth).
[0552] Input: Digital audio data.
[0553] Output: Audio data transfer to the server.
[0554] Step 4:
[0555] The server receives the digital audio data. The audio data is received at the server.
[0556] Input: Digital audio data.
[0557] Output: Audio data in the server.
[0558] Step 5:
[0559] The server converts the voice data into text using speech recognition technology. During this process, a speech signal processing algorithm extracts the characteristics of the voice waveform and converts it into text.
[0560] Input: Audio data.
[0561] Output: Text data (e.g., "Where is the shampoo shelf?").
[0562] Step 6:
[0563] The server analyzes the user's request from the converted text and uses natural language processing techniques to extract the intent (e.g., where to find shampoo on the shelf) from the text.
[0564] Input: Text data.
[0565] Output: The intent of the text (e.g., I want to know where shampoo is on the shelf).
[0566] Step 7:
[0567] The server retrieves product location information from the database. Using an SQL query, the server retrieves shampoo shelf location information from the database.
[0568] Input: User request information.
[0569] Output: Product location.
[0570] Step 8:
[0571] The server uses an emotion recognition engine to recognize emotions from the user's voice. In this process, the server analyzes the voice features and estimates the user's emotions.
[0572] Input: Audio data.
[0573] Output: Perceived emotion (e.g., confused).
[0574] Step 9:
[0575] The server adjusts the guidance content based on the recognized emotion. For example, if the user is confused, the guidance content will be revised to be more polite.
[0576] Input: Product location and perceived sentiment.
[0577] Output: Adjusted voice guidance text (e.g. "Slowly walk 10 meters to the right, then turn left at the next ledge.").
[0578] Step 10:
[0579] The server transmits the adjusted voice guidance text to the terminal.
[0580] Input: Adjusted voice prompt text.
[0581] Output: Transfer of guidance text to the terminal.
[0582] Step 11:
[0583] The device (smart glasses) receives the voice guidance text sent from the server.
[0584] Input: Adjusted voice prompt text.
[0585] Output: Voice guidance text in the device.
[0586] Step 12:
[0587] The device (smart glasses) uses a voice output device to provide guidance to the user. For example, it may use a speaker to say, "Go 10 meters to the right and then turn left at the next shelf."
[0588] Input: Voice prompt text.
[0589] Output: Audio instructions to the user.
[0590] Real-time location information function
[0591] Step 1:
[0592] The device (smart glasses) determines the user's current location using built-in sensors (e.g., GPS or WiFi triangulation).
[0593] Input: Position sensor data.
[0594] Output: Current location data.
[0595] Step 2:
[0596] The device (smart glasses) sends the determined current location data to the server via Wi-Fi or Bluetooth.
[0597] Input: Current location data.
[0598] Output: Send current location data to the server.
[0599] Step 3:
[0600] The server receives the current location data and determines the user's current location.
[0601] Input: Current location data.
[0602] Output: The current location determined.
[0603] Step 4:
[0604] The server retrieves the location information of the target product from the database and generates route information from the current location to the destination using a routing algorithm such as Dijkstra's algorithm.
[0605] Input: Current location and product location information.
[0606] Output: Route information.
[0607] Step 5:
[0608] The server recognizes the user's emotions using an emotion recognition engine and adjusts the route information.
[0609] Input: Voice data and route information.
[0610] Output: Adjusted route information.
[0611] Step 6:
[0612] The server sends the adjusted route information to the terminal.
[0613] Input: Adjusted route information.
[0614] Output: Sending route information to the terminal.
[0615] Step 7:
[0616] The terminal (smart glasses) converts the route information sent from the server into text for voice guidance and begins providing guidance to the user.
[0617] Input: Adjusted route information.
[0618] Output: Audio instructions to the user.
[0619] Currency conversion function
[0620] Step 1:
[0621] The device (smart glasses) scans the product barcode and obtains price information. A barcode reader is used to read the barcode and obtain price information.
[0622] Input: Barcode data.
[0623] Output: Price information.
[0624] Step 2:
[0625] Price information and the user's home currency information are sent to the server.
[0626] Input: Price information and home currency information.
[0627] Output: Sending data to the server.
[0628] Step 3:
[0629] The server receives the price information and retrieves the latest exchange rates from an external API, for example, using the API of a currency exchange rate provider.
[0630] Input: Price information.
[0631] Output: Exchange rate from external API.
[0632] Step 4:
[0633] The obtained exchange rate is used to convert the product price into the user's home currency.
[0634] Input: Exchange rate and price information.
[0635] Output: The converted price information.
[0636] Step 5:
[0637] The server recognizes the user's emotions using an emotion recognition engine and adjusts the price information.
[0638] Input: Audio data and converted price information.
[0639] Output: Adjusted price information.
[0640] Step 6:
[0641] The server sends the adjusted price information to the terminal.
[0642] Input: Adjusted price information.
[0643] Output: Sending price information to the terminal.
[0644] Step 7:
[0645] The device (smart glasses) converts the price information into text for voice output and notifies the user of the price.
[0646] Input: Adjusted price information.
[0647] Output: Audio notification to the user.
[0648] Price comparison feature
[0649] Step 1:
[0650] The terminal (smart glasses) scans the product barcode and obtains price information.
[0651] Input: Barcode data.
[0652] Output: Price information.
[0653] Step 2:
[0654] Send the price information and product identifier to the server.
[0655] Inputs: Price information and product identifiers.
[0656] Output: Sending data to the server.
[0657] Step 3:
[0658] A server receives the price information and the product identifier and queries a price database of multiple online stores.
[0659] Inputs: Price information and product identifiers.
[0660] Output: Price data for each online store.
[0661] Step 4:
[0662] Get prices from each online store and find the cheapest price.
[0663] Input: Online store price data.
[0664] Output: Cheapest price information.
[0665] Step 5:
[0666] The server recognizes the user's emotions using an emotion recognition engine and adjusts the cheapest information.
[0667] Input: Voice data and lowest price information.
[0668] Output: Adjusted lowest price information.
[0669] Step 6:
[0670] The server transmits the adjusted lowest price information to the terminal.
[0671] Input: Adjusted lowest price information.
[0672] Output: Sending price information to the terminal.
[0673] Step 7:
[0674] The device (smart glasses) converts the cheapest price information into text for voice output and notifies the user.
[0675] Input: Adjusted lowest price information.
[0676] Output: Audio notification to the user.
[0677] By following these steps, the system performs a series of processes, from voice input to data acquisition, location identification, currency conversion, price comparison, and emotion recognition, to provide users with an advanced shopping experience.
[0678] (Application example 2)
[0679] Next, a description will be given of Application Example 2. In the following description, the data processing device 12 will be referred to as a "server" and the smart device 14 will be referred to as a "terminal."
[0680] Conventional voice guidance systems using smart glasses could provide guidance based on user requests, but they could not adjust the guidance content according to the user's emotions. As a result, uniform guidance and information was provided without taking the user's emotional state into consideration, resulting in a suboptimal user experience. Particularly when shopping in a store, there was an issue of not being able to provide appropriate guidance when the user was confused or in a hurry.
[0681] The specific processing by the specific processing unit 290 of the data processing device 12 in the application example 2 is realized by the following means.
[0682] In this invention, the server includes emotion recognition means for recognizing emotions from the user's voice and facial expressions, means for converting voice input into text using voice recognition technology, and means for retrieving product location information from a database based on the analysis results, thereby enabling personalized voice guidance and information provision that reflects the user's emotional state.
[0683] "User" means an individual who operates a system or device.
[0684] "Voice input" refers to commands, questions, etc. that a user communicates to a system by speaking.
[0685] "Voice recognition technology" is a technology that analyzes voice signals and converts them into text data.
[0686] "Text conversion" refers to the process of converting captured voice input into text data.
[0687] "Request analysis" is the process of understanding the intent and content of a user's request from the converted text of their voice.
[0688] "Location information" refers to information that indicates where a particular product or section is located within a store.
[0689] A "database" is an information system that structures and stores a collection of related data so that it can be easily searched and retrieved.
[0690] "Voice guidance" refers to a method of providing information to users using a voice source.
[0691] "Emotion recognition means" refers to technology or equipment for detecting and classifying a user's emotional state from data such as the user's voice and facial expressions.
[0692] "Route information" refers to information relating to directions and directions from a specific starting point to a destination.
[0693] An "external API" is an interface for connecting with external software applications and services.
[0694] An "exchange rate" is a ratio that represents the exchange value between different currencies.
[0695] "Price conversion" refers to the process of converting product prices into other currency units.
[0696] "Facial expressions" refer to the movements and appearances that appear on the surface of the face and indicate emotions and intentions.
[0697] "Adjusting guidance content" means appropriately changing the information and direction guidance provided depending on the user's situation and emotions.
[0698] This invention relates to a system that allows users to use smart glasses in a store to receive voice guidance and obtain product information. Specifically, this system provides a more personalized shopping experience by combining voice input, voice recognition, emotion recognition, and location-based services.
[0699] Voice guidance function
[0700] The server uses the microphone built into the smart glasses to capture the user's voice input. The voice input is converted into text using voice recognition technology, and the user's request is analyzed from that text. Based on the results of this analysis, the server retrieves the product's location information from a database and provides the retrieved location information as voice guidance. An emotion recognition means is used to recognize the user's emotions from their voice and facial expressions, and the guidance content is adjusted based on the recognition results. This makes it possible to provide more detailed guidance if the user is confused, and faster guidance if the user is in a hurry.
[0701] Real-time location information function
[0702] The server uses sensors (such as GPS and WiFi triangulation) built into the smart glasses to identify the user's current location in real time. Based on the identified current location, the server generates route information to the product section and provides that route information as voice guidance. By adjusting the route guidance based on the user's location information, the user can move smoothly through the store. During route guidance, emotion recognition means is used to analyze the user's emotions and adjust the guidance content appropriately, providing a more personalized service.
[0703] Currency conversion function
[0704] The server retrieves the latest exchange rates from an external API, converts the product price into the user's home currency using the obtained exchange rate, and provides the converted price information via voice. During the price notification, the system recognizes the user's emotions and adjusts the guidance, such as providing a detailed explanation, if it detects doubt or confusion.
[0705] Specific examples
[0706] For example, if a user speaks to a smart glass in a store and asks, "Where is the shampoo shelf?", the smart glass will recognize the voice and send it to the server. The server will convert the voice into text and identify the location of the shampoo shelf. The emotion recognition means will analyze the user's emotions, and if the user appears confused, the smart glass will provide polite guidance such as, "Go 10 meters to the right and then turn left at the next shelf." On the other hand, if the system detects that the user is in a hurry, it will provide simple guidance such as, "It's 10 meters to the right."
[0707] If you pick up a product and ask, "How much is this product?", the device will retrieve price information and announce in voice, "The price of this shampoo is 5 euros (approximately 620 yen)." If you ask, "How much does this shampoo cost online?", the device will also provide the latest information on online prices, telling you, "It's being sold online for 4.5 euros." At the same time, it can analyze the user's emotions and adjust the guidance accordingly.
[0708] Prompt Sentence Examples
[0709] "Where's the shampoo shelf?"
[0710] "How much is this item?"
[0711] "How much does this shampoo cost online?"
[0712] The flow of the specific processing in the application example 2 will be described with reference to FIG.
[0713] Step 1:
[0714] The user provides voice input to the smart glasses.
[0715] Input: User speech (e.g., "Where is the shampoo shelf?")
[0716] How it works: The user's voice is captured using the microphone built into the smart glasses.
[0717] Output: Analog audio data
[0718] Step 2:
[0719] The device converts the voice data into a digital format.
[0720] Input: Analog audio data
[0721] How it works: The audio processing module inside the smart glasses converts analog audio data into digital audio data.
[0722] Output: Digital audio data
[0723] Step 3:
[0724] The terminal transmits digital audio data to the server.
[0725] Input: Digital audio data
[0726] How it works: Smart glasses send digital audio data to a server.
[0727] Output: Digital audio data sent to the server
[0728] Step 4:
[0729] The server converts the digital voice data into text.
[0730] Input: Digital audio data
[0731] How it works: The server's speech recognition technology analyzes the digital voice data and converts it into corresponding text.
[0732] Output: The transformed text (e.g., "Where is the shampoo shelf?")
[0733] Step 5:
[0734] The server parses the user's request from the text.
[0735] Input: Translated text
[0736] How it works: The server parses the text and determines what the user is requesting (e.g., I want to know where the shampoo shelf is).
[0737] Output: Analysis results (e.g., obtain product location information)
[0738] Step 6:
[0739] The server retrieves the product location information from the database.
[0740] Input: Analysis results (request for product location information)
[0741] Operation: The server queries the database and obtains the location information of the relevant product.
[0742] Output: Product location information (e.g., go 10 meters to the right and turn left at the next shelf)
[0743] Step 7:
[0744] The server recognizes the user's emotions using an emotion recognition means.
[0745] Input: User's voice and facial expression data
[0746] How it works: The server uses emotion recognition technology to analyze the user's emotions from their voice and facial expressions.
[0747] Output: User's emotional state (e.g. confused, in a hurry)
[0748] Step 8:
[0749] The server adjusts the guidance content based on the emotion.
[0750] Input: Product location, user emotional state
[0751] How it works: The server takes into account the user's emotional state and adjusts the guidance content (e.g., polite guidance if the user is confused, brief guidance if the user is in a hurry).
[0752] Output: Adjusted guidance text (e.g. "Go right 10 meters and turn left at the next ledge")
[0753] Step 9:
[0754] The server sends the adjusted guidance text to the smart glasses.
[0755] Input: Adjusted guidance text
[0756] Operation: The server sends the adjusted guidance text to the smart glasses.
[0757] Output: Guidance text sent to smart glasses
[0758] Step 10:
[0759] The terminal provides the user with the guidance text by voice.
[0760] Input: Guidance text
[0761] How it works: The audio output device of the smart glasses plays the guidance text as audio.
[0762] Output: Guidance speech (e.g. "Go 10 meters to the right and turn left at the next shelf")
[0763] The specific processing unit 290 transmits the result of the specific processing to the smart device 14. In the smart device 14, the control unit 46A causes the output device 40 to output the result of the specific processing. The microphone 38B acquires audio indicating a user input regarding the result of the specific processing. The control unit 46A transmits audio data indicating the user input acquired by the microphone 38B to the data processing device 12. In the data processing device 12, the specific processing unit 290 acquires the audio data.
[0764] The data generation model 58 is a so-called generative AI (Artificial Intelligence). An example of the data generation model 58 is ChatGPT (Internet Search<URL: https: / / openai.com / blog / chatgpt> ), Gemini (Internet search <url: https: gemini.google.com ?hl="ja">) and other generation AIs. The data generation model 58 is obtained by performing deep learning on a neural network. A prompt including an instruction is input to the data generation model 58, and inference data such as voice data indicating voice, text data indicating text, and image data indicating an image is also input. The data generation model 58 performs inference on the input inference data in accordance with the instruction indicated by the prompt, and outputs the inference result in a data format such as voice data and text data. Here, inference refers to, for example, analysis, classification, prediction, and / or summarization.
[0765] In the above embodiment, an example in which the specific process is performed by the data processing device 12 has been given, but the technology of the present disclosure is not limited to this, and the specific process may be performed by the smart device 14.
[0766] [Second embodiment]
[0767] FIG. 3 shows an example of the configuration of a data processing system 210 according to the second embodiment.
[0768] 3, the data processing system 210 includes the data processing device 12 and smart glasses 214. An example of the data processing device 12 is a server.
[0769] The data processing device 12 includes a computer 22, a database 24, and a communication I / F 26. The computer 22 is an example of a "computer" according to the technology of the present disclosure. The computer 22 includes a processor 28, a RAM 30, and a storage 32. The processor 28, the RAM 30, and the storage 32 are connected to a bus 34. The database 24 and the communication I / F 26 are also connected to the bus 34. The communication I / F 26 is connected to a network 54. Examples of the network 54 include a WAN (Wide Area Network) and / or a LAN (Local Area Network).
[0770] The smart glasses 214 include a computer 36, a microphone 238, a speaker 240, a camera 42, and a communication I / F 44. The computer 36 includes a processor 46, a RAM 48, and a storage 50. The processor 46, the RAM 48, and the storage 50 are connected to a bus 52. The microphone 238, the speaker 240, and the camera 42 are also connected to the bus 52.
[0771] The microphone 238 receives instructions and the like from the user 20 by receiving voice uttered by the user 20. The microphone 238 captures the voice uttered by the user 20, converts the captured voice into audio data, and outputs it to the processor 46. The speaker 240 outputs audio in accordance with instructions from the processor 46.
[0772] Camera 42 is a small digital camera equipped with an optical system including a lens, aperture, and shutter, and an imaging element such as a CMOS (Complementary Metal-Oxide-Semiconductor) image sensor or a CCD (Charge Coupled Device) image sensor, and captures images of the surroundings of user 20 (for example, an imaging range defined by an angle of view equivalent to the field of vision of a typical healthy person).
[0773] The communication I / F 44 is connected to a network 54. The communication I / Fs 44 and 26 are responsible for the exchange of various information between the processor 46 and the processor 28 via the network 54. The exchange of various information between the processor 46 and the processor 28 using the communication I / Fs 44 and 26 is carried out in a secure state.
[0774] Fig. 4 shows an example of the main functions of the data processing device 12 and the smart glasses 214. As shown in Fig. 4, in the data processing device 12, a specific process is performed by the processor 28. A specific process program 56 is stored in the storage 32.
[0775] The specific processing program 56 is an example of a "program" according to the technology of the present disclosure. The processor 28 reads the specific processing program 56 from the storage 32 and executes the read specific processing program 56 on the RAM 30. The specific processing is realized by the processor 28 operating as a specific processing unit 290 in accordance with the specific processing program 56 executed on the RAM 30.
[0776] The storage 32 stores a data generation model 58 and an emotion identification model 59. The data generation model 58 and the emotion identification model 59 are used by the identification processing unit 290.
[0777] In the smart glasses 214, the reception output process is performed by the processor 46. A reception output program 60 is stored in the storage 50. The processor 46 reads the reception output program 60 from the storage 50 and executes the read reception output program 60 on the RAM 48. The reception output process is realized by the processor 46 operating as the control unit 46A in accordance with the reception output program 60 executed on the RAM 48.
[0778] Next, a description will be given of the identification process performed by the identification processing unit 290 of the data processing device 12. In the following description, the data processing device 12 will be referred to as the "server" and the smart glasses 214 will be referred to as the "terminal."
[0779] The present invention is a voice guidance system using smart glasses that solves problems such as language barriers when shopping abroad, product location, currency conversion, price comparison, etc. Hereinafter, an embodiment of the present invention will be described in detail.
[0780] Voice guidance function
[0781] Device (smart glasses)
[0782] 1. The user asks the smart glasses, "Where is the shampoo shelf?"
[0783] 2. The smart glasses capture voice input via a built-in microphone and convert the voice data into a digital format.
[0784] 3. Send the digital audio data to the server.
[0785] server
[0786] 1. Receives digital voice data and converts it into text using speech recognition technology.
[0787] 2. Parse the user's request (in this case, the shelf location of shampoo) from the converted text.
[0788] 3. Obtain the location information of the shampoo shelves from the database and generate text for voice guidance.
[0789] 4. The generated voice guidance text is sent to the terminal.
[0790] Terminal
[0791] 1. Receive the voice guidance text sent from the server and use the voice output device to instruct the user, "Go 10 meters to the right and turn left at the next shelf."
[0792] Real-time location information function
[0793] Terminal
[0794] 1. Determine the user's current location using built-in sensors (GPS, WiFi triangulation, etc.).
[0795] 2. The determined current location data is sent to the server.
[0796] server
[0797] 1. Receive current location data and determine the user's location based on that data.
[0798] 2. Obtain the location information of the desired product from the database and generate route information from the current location to the destination.
[0799] 3. The generated route information is sent to the terminal.
[0800] Terminal
[0801] 1. Route information sent from the server is provided to the user via voice.
[0802] Currency conversion function
[0803] Terminal
[0804] 1. Scan the product barcode to get price information.
[0805] 2. Send price information and user's home currency information to the server.
[0806] server
[0807] 1. Receive price information and get the latest exchange rates using an external API.
[0808] 2. Convert the product price into the user's home currency using the obtained exchange rate.
[0809] 3. The converted price information is sent to the terminal.
[0810] Terminal
[0811] 1. The converted price information is notified to the user via voice, such as "The price of this shampoo is 500 yen."
[0812] Price comparison feature
[0813] Terminal
[0814] 1. Scan the product barcode to get price information.
[0815] 2. Send the price information and product identifier to the server.
[0816] server
[0817] 1. Receives price information and product identifiers and queries price databases from multiple online stores.
[0818] 2. Get prices from each online store and identify the lowest price.
[0819] 3. Send the cheapest price information to the terminal.
[0820] Terminal
[0821] 1. The lowest price information is announced to the user via voice, saying, "It's on sale online for 450 yen."
[0822] Specific examples
[0823] If a user is in a supermarket in France and asks the smart glasses, "Where is the shampoo shelf?", the smart glasses will respond with, "Go 10 meters to the right and turn left at the next shelf." After reaching the shampoo section, the user picks up an item and wants to check the price. They can ask, "How much is this shampoo?" and the smart glasses will respond, "This shampoo costs 5 euros (approximately 620 yen)." Furthermore, to compare the price with online prices, they can ask, "How much does this shampoo cost online?" and the smart glasses will respond, "It's selling for 4.5 euros online." Through this process, users can enjoy an efficient and comfortable shopping experience.
[0824] The processing flow will be explained below.
[0825] Voice guidance function
[0826] Specific processing steps from voice input to guidance
[0827] Step 1:
[0828] The user asks the smart glasses, "Where is the shampoo shelf?"
[0829] Step 2:
[0830] The terminal uses a built-in microphone to capture the user's voice and converts the voice into digital data.
[0831] Step 3:
[0832] The terminal transmits the converted digital audio data to the server.
[0833] Step 4:
[0834] The server receives the digital voice data and converts the voice data into text using voice recognition technology.
[0835] Step 5:
[0836] The server parses the text and understands the user's request (in this case, a request to know the shelf location of shampoo).
[0837] Step 6:
[0838] The server retrieves the location information of the shampoo shelves from the database.
[0839] Step 7:
[0840] Based on the acquired location information, the server generates text for voice guidance such as, "Go 10 meters to the right and turn left at the next shelf."
[0841] Step 8:
[0842] The server transmits the generated voice guidance text to the terminal.
[0843] Step 9:
[0844] The device converts the received voice guidance text into speech and begins providing guidance through the speaker.
[0845] Step 10:
[0846] The user moves as instructed.
[0847] Real-time location information function
[0848] A processing step to guide you from your current location to the product location
[0849] Step 1:
[0850] The user speaks to the smart glasses, saying, "Tell me where I am."
[0851] Step 2:
[0852] The device determines the user's current location using built-in sensors (e.g., GPS and WiFi triangulation).
[0853] Step 3:
[0854] The terminal transmits the identified current location data to the server.
[0855] Step 4:
[0856] The server receives the location data and confirms the user's current location.
[0857] Step 5:
[0858] The server retrieves route information to the product section from the database.
[0859] Step 6:
[0860] The server generates route information from the current location to the destination product section.
[0861] Step 7:
[0862] The server transmits the generated route information to the terminal.
[0863] Step 8:
[0864] The terminal converts the received route information into text for voice guidance.
[0865] Step 9:
[0866] The device converts the voice guidance text into speech and begins providing guidance through the speaker.
[0867] Step 10:
[0868] The user moves along the guided route.
[0869] Currency conversion function
[0870] Processing steps to convert product prices into your home currency
[0871] Step 1:
[0872] Users scan the barcode of the product they pick up with the smart glasses.
[0873] Step 2:
[0874] The terminal scans the barcode and obtains the product's price information.
[0875] Step 3:
[0876] The terminal transmits product price information and the user's home currency information to the server.
[0877] Step 4:
[0878] The server receives the price information and uses an external API to get the latest exchange rates.
[0879] Step 5:
[0880] The server uses the obtained exchange rate to convert the product price into the user's home currency.
[0881] Step 6:
[0882] The server transmits the converted price information to the terminal.
[0883] Step 7:
[0884] The terminal converts the received converted price information into text for voice output.
[0885] Step 8:
[0886] The device converts the text to speech and announces the price through the speaker.
[0887] Step 9:
[0888] The user receives the price information sent via the smart glasses.
[0889] Price comparison feature
[0890] Processing steps to compare in-store and online prices
[0891] Step 1:
[0892] Users scan the barcode of the product they pick up with the smart glasses.
[0893] Step 2:
[0894] The terminal scans the barcode and obtains the product's price information.
[0895] Step 3:
[0896] The terminal transmits the price information and the product identifier to the server.
[0897] Step 4:
[0898] The server receives the price information and the product identifier and queries the online store's price database.
[0899] Step 5:
[0900] The server retrieves the prices from each online store and identifies the lowest price.
[0901] Step 6:
[0902] The server transmits the cheapest price information to the terminal.
[0903] Step 7:
[0904] The terminal converts the received information about the cheapest price into text for voice output.
[0905] Step 8:
[0906] The device converts the text to speech and announces price information through the speaker.
[0907] Step 9:
[0908] The user receives price information from the smart glasses and decides to purchase.
[0909] Through the above processing steps, this system supports users in shopping abroad efficiently and comfortably.
[0910] Example 1
[0911] Next, a description will be given of Example 1. In the following description, the data processing device 12 will be referred to as a "server" and the smart glasses 214 will be referred to as a "terminal."
[0912] In today's global society, shopping in a multilingual environment can be extremely difficult for users. When shopping abroad, language barriers make it particularly difficult to accurately determine product locations. It's also difficult to convert prices displayed in local currencies into one's own currency, or to quickly compare prices across different stores. A system that solves these problems is needed.
[0913] The specific processing by the specific processing unit 290 of the data processing device 12 in the first embodiment is realized by the following means.
[0914] In this invention, the server includes: means for acquiring a user's voice input; means for converting the acquired voice input into digital format; means for transmitting the converted digital voice input to the server; means for converting the voice input into text on the server using voice recognition technology; means for analyzing the user's request from the converted text; means for acquiring product location information from a database based on the analysis results; means for converting the acquired location information into text for voice guidance and transmitting it to the terminal; and means for converting the transmitted text for voice guidance into speech and providing guidance to the user. This allows users to easily find products even in different language environments. Furthermore, by providing functions such as real-time location information identification and route guidance, currency conversion using the latest exchange rates, and price comparison across multiple stores, the server achieves an efficient and convenient shopping experience.
[0915] The "means for acquiring user's voice input" is a combination of hardware and software that recognizes the voice uttered by the user and that the terminal acquires as digital data.
[0916] The "means for converting to digital form" refers to an audio signal processing device and conversion algorithm for converting analog audio signals into digital data.
[0917] The "means for transmitting to a server" refers to a communication module and protocol that allows the terminal to transfer digital data over the Internet to a remote server.
[0918] "Means for converting speech to text using speech recognition technology" refers to software and algorithms that analyze speech data on a server and convert it into corresponding text.
[0919] The "means for analyzing user requests" is a natural language processing system that extracts and analyzes the user's intentions and requests from the text obtained by speech recognition.
[0920] The "means for acquiring product location information from a database" is a query processing system for searching and acquiring product placement information from a database based on the analyzed user request.
[0921] The "means for converting the information into text for voice guidance and sending it to the terminal" is a server-side function for converting the acquired information into text in a format that is easy for the user to understand and sending this text to the terminal.
[0922] The "means for converting the text information into audio and providing the user with guidance" refers to the software and hardware within the terminal for playing back the received text information as audio.
[0923] "Means of real-time identification" refers to a system that instantly identifies a user's current location using sensor technology such as GPS and Wi-Fi triangulation.
[0924] The "means for generating route information" refers to algorithms and software for calculating a route from the specified current position to the destination and generating that information.
[0925] The "means for obtaining the latest exchange rates from an external data source" is an interface for obtaining the latest currency exchange information from an external financial data provider service.
[0926] The "means for converting the product price using the exchange rate" is a calculation algorithm for converting the product price into the user's home currency using the acquired exchange rate information.
[0927] MODE FOR CARRYING OUT THE INVENTION
[0928] The present invention is a voice guidance system using smart glasses that solves problems such as language barriers when shopping abroad, product location, currency conversion, price comparison, etc. Hereinafter, an embodiment of the present invention will be described in detail.
[0929] Overall system overview
[0930] This system is composed of a combination of smart glasses (terminals), a cloud server, a database, and an external API. The smart glasses receive voice input from the user and send it to the cloud server. The cloud server uses voice recognition technology to convert the voice data into text, and then analyzes the user's request through natural language processing. Based on the analysis results, it provides information such as product location, currency conversion, and price comparison.
[0931] Hardware and software used
[0932] Device (smart glasses)
[0933] Built-in microphone: Captures the user's voice.
[0934] Communication module: Sends digital data to a cloud server (e.g., Wi-Fi, mobile data).
[0935] Voice output device: Converts guided text into voice.
[0936] Sensors: Determine current location (e.g. GPS, Wi-Fi triangulation).
[0937] server
[0938] Speech recognition software: Google Cloud Speech-to-Text API.
[0939] Natural language processing engine: Uses SpaCy or similar.
[0940] Database: Uses MySQL to manage product location and price information.
[0941] External API: Uses Open Exchange Rates API for currency conversion.
[0942] Program processing
[0943] 1. Voice guidance function
[0944] The user asks the smart glasses, "Where is the shampoo shelf?"
[0945] The smart glasses pick up sound through a built-in microphone and convert it into a digital format.
[0946] The digital audio data is transmitted to a cloud server.
[0947] The server uses speech recognition software to convert the voice data into text.
[0948] The text is analyzed using a natural language processing engine to understand the request.
[0949] Retrieve product location information from the database and generate guide text.
[0950] The guidance text is sent to the smart glasses and the user is guided by voice.
[0951] 2. Real-time location information function
[0952] The smart glasses use built-in sensors to determine the user's current location.
[0953] The current location data is sent to a cloud server.
[0954] The server generates route information to the product section based on the current location data.
[0955] The generated route information is sent to smart glasses and guidance is provided via voice.
[0956] 3. Currency conversion function
[0957] Smart glasses scan product barcodes to obtain price information.
[0958] Price information and the user's home currency information are sent to the server.
[0959] The server uses an external API to get the latest exchange rates.
[0960] Convert product prices into the user's home currency using exchange rates.
[0961] The converted price information is sent to the smart glasses and notified via voice.
[0962] 4. Price comparison feature
[0963] The smart glasses scan the product's barcode and retrieve price information.
[0964] Send the price information and product identifier to the server.
[0965] The server queries a price database of multiple online stores.
[0966] The lowest price information is identified and sent to the smart glasses.
[0967] The cheapest price information is notified to the user by voice.
[0968] Specific examples
[0969] If a user is in a supermarket in France and asks the smart glasses, "Where is the shampoo shelf?", the smart glasses will respond with, "Go 10 meters to the right and turn left at the next shelf." After reaching the shampoo section, the user picks up an item and wants to check the price. They can ask, "How much is this shampoo?" and the smart glasses will respond, "This shampoo costs 5 euros (approximately 620 yen)." Furthermore, to compare the price with online prices, they can ask, "How much does this shampoo cost online?" and the smart glasses will respond, "It's selling for 4.5 euros online." Through this process, users can enjoy an efficient and comfortable shopping experience.
[0970] Example prompts to input to the generative AI model
[0971] Please explain in detail the process by which a user queries the smart glasses for product location, product price, currency conversion, and price comparison. Please also include specific hardware, software, data processing, and calculation methods.
[0972] The flow of the identification process in the first embodiment will be described with reference to FIG.
[0973] Step 1:
[0974] A user speaks to the smart glasses, saying, "Where is the shampoo shelf?" The user's voice is picked up by the smart glasses' built-in microphone. This voice input is an analog signal, so it is converted into digital format. Specifically, A / D conversion is performed to convert the voice waveform into a digital signal. The input is the user's voice signal, and the output is digital voice data.
[0975] Step 2:
[0976] The device transmits digital audio data to a cloud server via Wi-Fi or a mobile data network using a communication module. The input is digital audio data, and the output is confirmation of data transmission to the server.
[0977] Step 3:
[0978] The server receives the digital voice data and converts it into text using speech recognition technology. The Google Cloud Speech-to-Text API is used. The server analyzes the voice data and generates corresponding text data. The input is digital voice data, and the output is the converted text data.
[0979] Step 4:
[0980] The server sends the converted text data to a natural language processing engine to analyze the user's request. For example, SpaCy can be used to extract the meaning of the text. The input is text data, and the output is the analysis result of the user's request ("where shampoo is on the shelf"). Specifically, the keyword "shampoo" is extracted from the text, and the request content is determined based on that.
[0981] Step 5:
[0982] The server retrieves product location information from the database based on the analysis results. In this example, it queries a MySQL database to retrieve the location information of the shampoo shelf. The input is the analysis result of the user request, and the output is the product location information. The database query is executed and the location information is returned.
[0983] Step 6:
[0984] The server converts the acquired location information into text for voice guidance and sends it to the terminal. An algorithm for generating voice guidance text is applied to generate text such as "Go 10 meters to the right and turn left at the next shelf." The input is the product location information, and the output is the text for voice guidance. The generated text is sent to the terminal as an HTTP response.
[0985] Step 7:
[0986] The device receives the text for voice guidance sent from the server. The received text is then used to provide guidance to the user using a voice output device. Specifically, the text is converted to speech using the Google Text-to-Speech API. The input is the text for voice guidance, and the output is the voice guidance. "Go 10 meters to the right and turn left at the next shelf," is the voice guidance that comes out of the smart glasses.
[0987] (Application example 1)
[0988] Next, a description will be given of Application Example 1. In the following description, the data processing device 12 will be referred to as a "server" and the smart glasses 214 will be referred to as a "terminal."
[0989] The present invention aims to solve problems associated with shopping abroad, such as language barriers, product location, currency conversion, and price comparisons. Specifically, the problem is to provide a support system that enables users to intuitively and efficiently find products in physical stores, understand prices, and purchase them at the most advantageous prices.
[0990] The specific processing by the specific processing unit 290 of the data processing device 12 in the application example 1 is realized by the following means.
[0991] In this invention, the server includes means for acquiring a user's voice input, means for converting the acquired voice input into text using voice recognition technology, means for analyzing the user's request from the converted text, means for acquiring product location information from a database based on the analysis result, means for providing the acquired location information as voice guidance, means for converting product prices into the user's home currency using a currency conversion function, and means for acquiring price information from multiple online stores using a price comparison function to identify the lowest price. This allows users to easily find products even in stores in a foreign country, and also makes it easy to convert currencies and compare prices.
[0992] "User voice input" refers to voice commands issued by a user to a device.
[0993] "Speech recognition technology" refers to machine learning and natural language processing technology for converting speech into text data.
[0994] "Means for converting to text" refers to a device or software that performs a process of converting captured voice input into digital data and then parsing that digital data into text.
[0995] "Means for analyzing user requests" refers to a device or software that processes voice commands converted into text to understand the information or action the user is seeking.
[0996] "Means for obtaining product location information from a database" refers to a device or software for searching and obtaining information about the product location stored in a database.
[0997] "Means for providing audio guidance" refers to a device or software that provides audio guidance to the user based on analyzed information and acquired data.
[0998] "Currency Conversion Facility" means a device or software that performs a process to convert product prices displayed in one currency into another currency selected by the user.
[0999] "Price comparison tool" means a device or software that performs a process to obtain and compare price information from multiple sources or online stores for a particular product to identify the most favorable price.
[1000] "Means for determining current location in real time" refers to devices or software that obtain a user's current location in real time using technologies such as GPS or WiFi triangulation.
[1001] The "means for generating route information" refers to a device or software for calculating the optimal route from the user's current location to the location of the desired product and generating the route information.
[1002] "Means for obtaining the latest exchange rates from an external API" refers to a device or software for obtaining exchange rate information from an external server via the Internet.
[1003] The present invention relates to a system that uses a voice guidance system in a specified target environment to provide users with various information such as product location information, currency conversion, price comparison, etc. Specific embodiments of the system are described below.
[1004] Generating a Program
[1005] The system consists of the following main functions:
[1006] 1. Voice Input Capture and Recognition
[1007] 2. Location information acquisition and guidance
[1008] 3. Currency conversion function
[1009] 4. Price comparison feature
[1010] Program processing explanation
[1011] Voice Input Capture and Recognition
[1012] The device (e.g., smart glasses) captures the user's voice input with a microphone and converts the voice data into a digital format. The converted digital voice data is sent to the server using the os and wave libraries. The server receives the voice data via a flask-based endpoint and converts it into text using the pydub and speech_recognition libraries.
[1013] Next, a natural language processing library such as nltk is used to analyze the user's request, and the next steps are taken based on the analysis results.
[1014] Location information acquisition and guidance
[1015] The device acquires its current location in real time using its built-in GPS and WiFi triangulation, and sends it to the server. The server processes the location data using the geopy library and generates the optimal route to the product section based on the user's current location. This route information is then provided as audio guidance.
[1016] Currency conversion function
[1017] When a user scans a product's barcode, the device retrieves the price information and sends it to the server. The server then uses an external API to retrieve the latest exchange rate and converts the product price based on the obtained rate. The converted price information is retrieved using the requests library and notified to the user via voice.
[1018] Price comparison feature
[1019] When a user scans a product's barcode, the device sends the price information to a server, which queries a price database of multiple online stores to identify the lowest price, which is then provided to the user via voice.
[1020] Specific examples
[1021] If a user is in a supermarket in France and asks the smart glasses, "Where is the shampoo shelf?", the smart glasses will respond with, "Go 10 meters to the right and turn left at the next shelf." After reaching the shampoo section, the user picks up an item and wants to check the price. They can ask, "How much is this shampoo?" and the smart glasses will respond, "This shampoo costs 5 euros (approximately 620 yen)." Furthermore, to compare the price with online prices, they can ask, "How much does this shampoo cost online?" and the smart glasses will respond, "It's selling for 4.5 euros online." Through this process, users can enjoy an efficient and comfortable shopping experience.
[1022] Prompt Sentence Examples
[1023] "This smartglasses-based shopping app allows users to easily navigate in-store, convert currencies, compare prices, and more, all by voice. For example, if a user is in a supermarket in France and asks the smartglasses, 'Where is the shampoo shelf?', the smartglasses will respond with, 'Go 10 meters to the right and then turn left at the next shelf.' Please generate a program that achieves this function."
[1024] The flow of the specific processing in the application example 1 will be described with reference to FIG.
[1025] Step 1:
[1026] The device (smart glasses) captures the user's voice input with a microphone. The user speaks, "Where is the shampoo shelf?", and this voice signal is acquired. The voice input is converted into digital voice data using the OS and Wave libraries. The output is digital voice data.
[1027] Step 2:
[1028] The device sends digital audio data to the server. This is done using the requests library. The input is the digital audio data, and the output is a request to the server to send the audio data.
[1029] Step 3:
[1030] The server receives the received digital audio data at the flask endpoint and starts analyzing the data. The input is the digital audio data. The server converts it into text using the pydub and speech_recognition libraries. The output is the converted text data.
[1031] Step 4:
[1032] The server analyzes the converted text data using a natural language processing library such as nltk to identify the user's request. The converted text data is input, and the user's request extracted through analysis is output. Specifically, "the location of the shampoo shelf" is identified as the user's request.
[1033] Step 5:
[1034] The server retrieves the shampoo shelf location information from the database. The input is the user's request (the shampoo shelf location), and the output is the product location information. This location information indicates where the product is located in the store.
[1035] Step 6:
[1036] The server calculates the location using the geopy library based on the product location information and generates route information. The input is the product location information and the user's current location, and the output is route information. Specifically, the guidance will be something like "Go 10 meters to the right and turn left at the next shelf."
[1037] Step 7:
[1038] The server converts the generated route information into text for voice guidance and sends it to the terminal. The input is route information, and a process is performed to convert this information into text for voice guidance. The output is text for voice guidance.
[1039] Step 8:
[1040] The terminal notifies the user of the voice guidance text received from the server using a voice output device. The input is the voice guidance text, and the output is the voice guidance ("Go 10 meters to the right and turn left at the next shelf").
[1041] Step 9:
[1042] When a user scans a product barcode, the terminal obtains the price information and sends it to the server. The input is the price information output by scanning the product barcode.
[1043] Step 10:
[1044] The server uses an external API to obtain the latest exchange rate based on the acquired price information. The input is the price information, and the output is the latest exchange rate.
[1045] Step 11:
[1046] The server converts the price into the user's home currency based on the exchange rate. The input is the acquired exchange rate and price information, and the output is the converted price information.
[1047] Step 12:
[1048] The server converts the converted price information into text for voice guidance and sends it to the terminal. The input is the converted price information, and the output is the text for voice guidance.
[1049] Step 13:
[1050] The terminal notifies the user of the voice guidance text received from the server using a voice output device. The input is the voice guidance text, and the output is the voice guidance ("The price of this shampoo is XX yen").
[1051] Step 14:
[1052] When a user speaks "What is the price of this shampoo online?", the device sends this speech input to the server. The input is the user's speech input, and the output is digital voice data.
[1053] Step 15:
[1054] The server analyzes the received voice data to identify it as a price comparison request and queries a price database of multiple online stores. The input is the analyzed request, and the output is the price information of each online store.
[1055] Step 16:
[1056] The server identifies the cheapest price information, converts it into a voice guidance text, and sends it to the terminal. The input is the price information of the online store, and the output is a voice guidance text containing the cheapest price information.
[1057] Step 17:
[1058] The terminal notifies the user of the voice guidance text received from the server using a voice output device. The input is the voice guidance text, and the output is the voice guidance ("This item is sold online for XX yen").
[1059] Furthermore, an emotion engine that estimates the user's emotion may be further combined. That is, the identification processing unit 290 may estimate the user's emotion using the emotion identification model 59, and perform identification processing using the user's emotion.
[1060] The present invention provides a system that combines a voice guidance system using smart glasses with an emotion engine that recognizes the user's emotions, providing a more personalized shopping experience. The following describes in detail an embodiment of the present invention.
[1061] Voice guidance function
[1062] Device (smart glasses)
[1063] 1. The user asks the smart glasses, "Where is the shampoo shelf?"
[1064] 2. The smart glasses capture voice input via a built-in microphone and convert the voice data into a digital format.
[1065] 3. Send the digital audio data to the server.
[1066] server
[1067] 1. Receive digital voice data and convert the voice data into text using voice recognition technology.
[1068] 2. Understand the user's request from the converted text (in this case, the request to know the shelf location of shampoo).
[1069] 3. Obtain the location information of the shampoo shelves from the database and generate text for voice guidance.
[1070] 4. Use the emotion engine to recognize emotions from the user's voice and adjust the guidance content.
[1071] 5. The adjusted voice guidance text is sent to the device.
[1072] Terminal
[1073] 1. Receive the voice guidance text sent from the server and use the voice output device to instruct the user, "Go 10 meters to the right and turn left at the next shelf."
[1074] Real-time location information function
[1075] Terminal
[1076] 1. Determine the user's current location using built-in sensors (GPS, WiFi triangulation, etc.).
[1077] 2. The determined current location data is sent to the server.
[1078] server
[1079] 1. Receive current location data and determine the user's location based on that data.
[1080] 2. Obtain the location information of the desired product from the database and generate route information from the current location to the destination.
[1081] 3. Use an emotion engine to recognize emotions from the user's voice and facial expressions and adjust the guidance content accordingly.
[1082] 4. The adjusted route information is sent to the terminal.
[1083] Terminal
[1084] 1. The route information sent from the server is converted into text for voice guidance, and guidance to the user begins.
[1085] Currency conversion function
[1086] Terminal
[1087] 1. Scan the product barcode to get price information.
[1088] 2. Send price information and user's home currency information to the server.
[1089] server
[1090] 1. Receive price information and get the latest exchange rates using an external API.
[1091] 2. Convert the product price into the user's home currency using the obtained exchange rate.
[1092] 3. Use an emotion engine to recognize emotions from the user's voice and facial expressions and adjust the content of price information notifications.
[1093] 4. The adjusted price information is sent to the terminal.
[1094] Terminal
[1095] 1. The converted price information is converted into text for voice output and announced through the speaker: "The price of this shampoo is 500 yen."
[1096] Price comparison feature
[1097] Terminal
[1098] 1. Scan the product barcode to get price information.
[1099] 2. Send the price information and product identifier to the server.
[1100] server
[1101] 1. Receives price information and product identifiers and queries price databases from multiple online stores.
[1102] 2. Get the prices from each online store and identify the cheapest price.
[1103] 3. Use an emotion engine to recognize emotions from the user's voice and facial expressions and adjust the content of the notification about the cheapest price.
[1104] 4. The adjusted lowest price information is sent to the terminal.
[1105] Terminal
[1106] 1. Convert the cheapest price information into text for voice output and announce through the speaker, "It's on sale online for 450 yen."
[1107] Use of emotion engine
[1108] server
[1109] 1. Receives voice data and facial expression data and recognizes the user's emotions using an emotion engine.
[1110] 2. Tailor voice prompts and pricing notifications based on the recognized emotion.
[1111] 3. The adjusted information is sent to the device.
[1112] Specific examples
[1113] If a user is in a supermarket in France and asks the smart glasses, "Where is the shampoo shelf?", the smart glasses will respond with, "Go 10 meters to the right and turn left at the next shelf." After reaching the shampoo section, the user picks up an item and asks, "How much is this shampoo?" to check the price. The smart glasses will respond, "This shampoo costs 5 euros (approximately 620 yen)." Furthermore, if the user asks, "How much does this shampoo cost online?" to compare it with online prices, the smart glasses will respond, "It's selling for 4.5 euros online." This invention utilizes an emotion engine to optimize the user experience, such as providing more attentive guidance if the user shows a confused expression.
[1114] The processing flow will be explained below.
[1115] Voice guidance function
[1116] Specific processing steps from voice input to guidance
[1117] Step 1:
[1118] The user asks the smart glasses, "Where is the shampoo shelf?"
[1119] Step 2:
[1120] The terminal uses a built-in microphone to capture the user's voice and converts the voice into digital data.
[1121] Step 3:
[1122] The terminal transmits the converted digital audio data to the server.
[1123] Step 4:
[1124] The server receives the digital voice data and converts the voice data into text using voice recognition technology.
[1125] Step 5:
[1126] The server parses the text and understands the user's request (in this case, a request to know the shelf location of shampoo).
[1127] Step 6:
[1128] The server retrieves the location information of the shampoo shelves from the database.
[1129] Step 7:
[1130] Based on the acquired location information, the server generates text for voice guidance such as, "Go 10 meters to the right and turn left at the next shelf."
[1131] Step 8:
[1132] The server sends the voice data to an emotion engine to analyze the user's emotion from the voice data.
[1133] Step 9:
[1134] The server receives the emotion recognition results from the emotion engine and adjusts the content and tone of the announcement based on the user's emotion.
[1135] Step 10:
[1136] The server transmits the adjusted voice guidance text to the terminal.
[1137] Step 11:
[1138] The device converts the received voice guidance text into speech and begins providing guidance through the speaker.
[1139] Step 12:
[1140] The user moves as instructed.
[1141] Real-time location information function
[1142] Specific processing steps to guide you from your current location to the product location
[1143] Step 1:
[1144] The user speaks to the smart glasses, saying, "Tell me where I am."
[1145] Step 2:
[1146] The device determines the user's current location using built-in sensors (e.g., GPS and WiFi triangulation).
[1147] Step 3:
[1148] The terminal transmits the identified current location data to the server.
[1149] Step 4:
[1150] The server receives the location data and determines the user's location based on it.
[1151] Step 5:
[1152] The server retrieves route information to the product section from the database.
[1153] Step 6:
[1154] The server generates route information from the current location to the destination product section.
[1155] Step 7:
[1156] The server sends the user's voice data and facial expression data to the emotion engine to analyze the user's emotions.
[1157] Step 8:
[1158] The server receives the emotion recognition results from the emotion engine and adjusts the content and tone of the route guidance based on the emotion.
[1159] Step 9:
[1160] The server transmits the adjusted route information to the terminal.
[1161] Step 10:
[1162] The terminal converts the received route information into text for voice guidance.
[1163] Step 11:
[1164] The terminal converts the voice guidance text into voice and starts providing guidance to the user through the speaker.
[1165] Step 12:
[1166] The user moves along the guided route.
[1167] Currency conversion function
[1168] Specific steps to convert product prices into your home currency
[1169] Step 1:
[1170] Users scan the barcode of the product they pick up with the smart glasses.
[1171] Step 2:
[1172] The terminal scans the barcode and obtains the product's price information.
[1173] Step 3:
[1174] The terminal transmits product price information and the user's home currency information to the server.
[1175] Step 4:
[1176] The server receives the price information and uses an external API to get the latest exchange rates.
[1177] Step 5:
[1178] The server uses the obtained exchange rate to convert the product price into the user's home currency.
[1179] Step 6:
[1180] The server sends the user's voice data and facial expression data to the emotion engine to analyze the user's emotions.
[1181] Step 7:
[1182] The server receives emotion recognition results from the emotion engine and adjusts the content and tone of price information notifications based on the emotion.
[1183] Step 8:
[1184] The server transmits the adjusted price information to the terminal.
[1185] Step 9:
[1186] The terminal converts the received price information into text for voice output.
[1187] Step 10:
[1188] The device converts the text for voice output into speech and announces through the speaker, "The price of this shampoo is 500 yen."
[1189] Step 11:
[1190] The user receives the price information sent via the smart glasses.
[1191] Price comparison feature
[1192] Specific steps to compare in-store and online prices
[1193] Step 1:
[1194] Users scan the barcode of the product they pick up with the smart glasses.
[1195] Step 2:
[1196] The terminal scans the barcode and obtains the product's price information.
[1197] Step 3:
[1198] The terminal transmits the price information and the product identifier to the server.
[1199] Step 4:
[1200] The server receives the price information and the product identifier and queries a price database of multiple online stores.
[1201] Step 5:
[1202] The server retrieves the prices from each online store and identifies the lowest price.
[1203] Step 6:
[1204] The server sends the user's voice data and facial expression data to the emotion engine for emotion analysis.
[1205] Step 7:
[1206] The server receives the emotion recognition result from the emotion engine and adjusts the content and tone of the notification of the cheapest price information based on the emotion recognition result.
[1207] Step 8:
[1208] The server transmits the adjusted lowest price information to the terminal.
[1209] Step 9:
[1210] The terminal converts the received information about the cheapest price into text for voice output.
[1211] Step 10:
[1212] The device converts the text to speech and announces through the speaker, "It's on sale online for 450 yen."
[1213] Step 11:
[1214] The user receives price information from the smart glasses and decides to purchase.
[1215] Through the above processing steps, this system supports users in shopping abroad efficiently and comfortably. By utilizing the emotion engine, guidance and information are provided according to the user's emotional state, providing a more personalized experience.
[1216] Example 2
[1217] Next, a description will be given of Example 2. In the following description, the data processing device 12 will be referred to as a "server" and the smart glasses 214 will be referred to as a "terminal."
[1218] The current shopping experience is burdensome for users, as it is difficult to locate products in physical stores. Furthermore, price conversion and comparison are time-consuming, making shopping especially complicated in foreign countries. Furthermore, conventional voice guidance systems do not take user emotions into account, so guidance information is not optimized according to the user's state. To solve these problems, a system that can recognize user emotions and provide more personalized guidance is needed.
[1219] The identification process by the identification processing unit 290 of the data processing device 12 in the second embodiment is realized by the following means. In this invention, the server includes a means for converting acquired voice input into text using voice recognition technology, a means for analyzing the user's request from the converted text, and a means for acquiring product location information from a database based on the analysis result, adjusting the information based on an emotion recognition engine, and providing voice guidance. This makes it possible to provide optimal guidance according to the user's emotions.
[1220] "User" refers to the person who uses this system and is the subject who performs voice input.
[1221] "Voice input" refers to voice data provided by a user through a voice input device such as a microphone.
[1222] "Speech recognition technology" refers to technology for converting voice data into text data, and refers to technology that uses, for example, voice signal processing algorithms or machine learning models.
[1223] "Text" is character string data converted using speech recognition technology.
[1224] "Request" refers to various questions and instructions that a user gives to the system.
[1225] "Analysis" refers to the process of understanding user requests and responding appropriately to them.
[1226] "Location information" is data that indicates the physical location of a particular product or location.
[1227] A "database" is an information management system for storing and managing location information and related data.
[1228] An "emotion recognition engine" is software or hardware that recognizes a user's emotions from their voice or facial expressions.
[1229] "Voice guidance" refers to guidance information provided to users via voice.
[1230] "Real-time" refers to near-instant processing or response.
[1231] "Current location" is data indicating the physical location where the user is currently located.
[1232] "Route information" is data that indicates the route and instructions from the current position to the destination.
[1233] An "external API" is an application program interface provided by a third party that is a means of connecting with other systems.
[1234] "Exchange rate" is data that indicates the exchange rate between one country's currency and another country's currency.
[1235] "Home currency" refers to the currency of the country or region to which the user belongs.
[1236] This invention is a voice guidance system using smart glasses that combines an emotion engine that recognizes the user's emotions to provide a more personalized shopping experience.
[1237] Voice guidance function
[1238] Device (smart glasses)
[1239] The user speaks to the smart glasses, saying, "Where is the shampoo shelf?" The smart glasses receive voice input via a built-in microphone and convert the voice data into a digital format. The converted digital voice data is then transmitted to the server via wireless communication (WiFi or Bluetooth).
[1240] server
[1241] The server converts the received digital voice data into text using voice recognition technology (e.g., technology using a voice signal processing algorithm or a machine learning model), then analyzes the user's request from the text (e.g., using natural language processing technology) and retrieves the location information of the corresponding product from a database (e.g., an information management system).
[1242] The acquired location information is adjusted by an emotion recognition engine (e.g., software or hardware for recognizing emotions from voice and facial expressions) to take into account the user's emotions. Finally, an adjusted voice guidance text is generated and sent to the device.
[1243] Device (smart glasses)
[1244] The voice guidance text sent from the server is received and the voice output device (speaker or earphone) is used to guide the user, for example, "Go 10 meters to the right and turn left at the next shelf."
[1245] Real-time location information function
[1246] Device (smart glasses)
[1247] The user's current location is determined using built-in sensors (e.g., GPS or WiFi triangulation), and the determined current location data is transmitted to a server.
[1248] server
[1249] The server identifies the user's location based on the received current location data. Next, it retrieves the desired product's location information from the database and generates route information from the user's current location to the destination. This route information is also adjusted by an emotion recognition engine, taking into account the user's emotions. Finally, the adjusted route information is sent to the device.
[1250] Device (smart glasses)
[1251] The route information sent from the server is converted into text for voice guidance, and guidance to the user is started.
[1252] Currency conversion function
[1253] Device (smart glasses)
[1254] When a user picks up a product and wants to check the price, the smart glasses scan the product barcode to obtain price information, which is then sent to the server along with the user's home currency.
[1255] server
[1256] The server obtains the latest exchange rate from an external API (e.g., an application program interface for obtaining exchange rates) and uses it to convert the product price into the user's home currency. The converted price information is adjusted by an emotion recognition engine and sent to the device.
[1257] Device (smart glasses)
[1258] The converted price information is converted into text for voice output to the user, and the user is notified through the speaker that "The price of this shampoo is 500 yen."
[1259] Price comparison feature
[1260] Device (smart glasses)
[1261] When a user wants to compare prices at different stores, they scan the product barcode to get the price information, which is then sent to the server along with the product identifier.
[1262] server
[1263] The server queries price databases of multiple online stores, retrieves prices from each online store, and identifies the cheapest price. This cheapest price information is also adjusted by the emotion recognition engine. Finally, the adjusted cheapest price information is sent to the device.
[1264] Device (smart glasses)
[1265] The lowest price information is converted into text for voice output and announced through the speaker: "It's on sale online for 450 yen."
[1266] Using the Emotion Recognition Engine
[1267] server
[1268] The server receives the voice data and facial expression data and uses an emotion recognition engine to recognize the user's emotions. Based on the recognized emotions, the voice guidance and price information notifications are adjusted. Finally, the adjusted information is sent to the device.
[1269] Prompt Sentence Examples
[1270] If a user is in a supermarket in France, they can ask their smart glasses to:
[1271] "Where's the shampoo shelf?"
[1272] "How much is this shampoo?"
[1273] "How much does this shampoo cost online?"
[1274] Entering these prompts into the system will provide corresponding guidance and pricing information.
[1275] As described above, the system combines multiple functions such as voice input, location information, currency conversion, price comparison, and emotion recognition to provide people with an advanced and personalized shopping experience.
[1276] The flow of the identification process in the second embodiment will be described with reference to FIG.
[1277] Voice guidance function
[1278] Step 1:
[1279] The device (smart glasses) receives the user's voice input. When a user speaks to the smart glasses, saying, "Where is the shampoo shelf?", this voice data is captured by the device's built-in microphone.
[1280] Input: User's voice data.
[1281] Output: Audio data.
[1282] Step 2:
[1283] The audio data acquired by the device (smart glasses) is converted into a digital format. Specifically, AD conversion is performed to convert the audio data into a digital signal.
[1284] Input: Analog audio data.
[1285] Output: Digital audio data.
[1286] Step 3:
[1287] The device (smart glasses) sends digital audio data to the server. The audio data is transferred to the server via wireless communication (WiFi or Bluetooth).
[1288] Input: Digital audio data.
[1289] Output: Audio data transfer to the server.
[1290] Step 4:
[1291] The server receives the digital audio data. The audio data is received at the server.
[1292] Input: Digital audio data.
[1293] Output: Audio data in the server.
[1294] Step 5:
[1295] The server converts the voice data into text using speech recognition technology. During this process, a speech signal processing algorithm extracts the characteristics of the voice waveform and converts it into text.
[1296] Input: Audio data.
[1297] Output: Text data (e.g., "Where is the shampoo shelf?").
[1298] Step 6:
[1299] The server analyzes the user's request from the converted text and uses natural language processing techniques to extract the intent (e.g., where to find shampoo on the shelf) from the text.
[1300] Input: Text data.
[1301] Output: The intent of the text (e.g., I want to know where shampoo is on the shelf).
[1302] Step 7:
[1303] The server retrieves product location information from the database. Using an SQL query, the server retrieves shampoo shelf location information from the database.
[1304] Input: User request information.
[1305] Output: Product location.
[1306] Step 8:
[1307] The server uses an emotion recognition engine to recognize emotions from the user's voice. In this process, the server analyzes the voice features and estimates the user's emotions.
[1308] Input: Audio data.
[1309] Output: Perceived emotion (e.g., confused).
[1310] Step 9:
[1311] The server adjusts the guidance content based on the recognized emotion. For example, if the user is confused, the guidance content will be revised to be more polite.
[1312] Input: Product location and perceived sentiment.
[1313] Output: Adjusted voice guidance text (e.g. "Slowly walk 10 meters to the right, then turn left at the next ledge.").
[1314] Step 10:
[1315] The server transmits the adjusted voice guidance text to the terminal.
[1316] Input: Adjusted voice prompt text.
[1317] Output: Transfer of guidance text to the terminal.
[1318] Step 11:
[1319] The device (smart glasses) receives the voice guidance text sent from the server.
[1320] Input: Adjusted voice prompt text.
[1321] Output: Voice guidance text in the device.
[1322] Step 12:
[1323] The device (smart glasses) uses a voice output device to provide guidance to the user. For example, it may use a speaker to say, "Go 10 meters to the right and then turn left at the next shelf."
[1324] Input: Voice prompt text.
[1325] Output: Audio instructions to the user.
[1326] Real-time location information function
[1327] Step 1:
[1328] The device (smart glasses) determines the user's current location using built-in sensors (e.g., GPS or WiFi triangulation).
[1329] Input: Position sensor data.
[1330] Output: Current location data.
[1331] Step 2:
[1332] The device (smart glasses) sends the determined current location data to the server via Wi-Fi or Bluetooth.
[1333] Input: Current location data.
[1334] Output: Send current location data to the server.
[1335] Step 3:
[1336] The server receives the current location data and determines the user's current location.
[1337] Input: Current location data.
[1338] Output: The current location determined.
[1339] Step 4:
[1340] The server retrieves the location information of the target product from the database and generates route information from the current location to the destination using a routing algorithm such as Dijkstra's algorithm.
[1341] Input: Current location and product location information.
[1342] Output: Route information.
[1343] Step 5:
[1344] The server recognizes the user's emotions using an emotion recognition engine and adjusts the route information.
[1345] Input: Voice data and route information.
[1346] Output: Adjusted route information.
[1347] Step 6:
[1348] The server sends the adjusted route information to the terminal.
[1349] Input: Adjusted route information.
[1350] Output: Sending route information to the terminal.
[1351] Step 7:
[1352] The terminal (smart glasses) converts the route information sent from the server into text for voice guidance and begins providing guidance to the user.
[1353] Input: Adjusted route information.
[1354] Output: Audio instructions to the user.
[1355] Currency conversion function
[1356] Step 1:
[1357] The device (smart glasses) scans the product barcode and obtains price information. A barcode reader is used to read the barcode and obtain price information.
[1358] Input: Barcode data.
[1359] Output: Price information.
[1360] Step 2:
[1361] Price information and the user's home currency information are sent to the server.
[1362] Input: Price information and home currency information.
[1363] Output: Sending data to the server.
[1364] Step 3:
[1365] The server receives the price information and retrieves the latest exchange rates from an external API, for example, using the API of a currency exchange rate provider.
[1366] Input: Price information.
[1367] Output: Exchange rate from external API.
[1368] Step 4:
[1369] The obtained exchange rate is used to convert the product price into the user's home currency.
[1370] Input: Exchange rate and price information.
[1371] Output: The converted price information.
[1372] Step 5:
[1373] The server recognizes the user's emotions using an emotion recognition engine and adjusts the price information.
[1374] Input: Audio data and converted price information.
[1375] Output: Adjusted price information.
[1376] Step 6:
[1377] The server sends the adjusted price information to the terminal.
[1378] Input: Adjusted price information.
[1379] Output: Sending price information to the terminal.
[1380] Step 7:
[1381] The device (smart glasses) converts the price information into text for voice output and notifies the user of the price.
[1382] Input: Adjusted price information.
[1383] Output: Audio notification to the user.
[1384] Price comparison feature
[1385] Step 1:
[1386] The terminal (smart glasses) scans the product barcode and obtains price information.
[1387] Input: Barcode data.
[1388] Output: Price information.
[1389] Step 2:
[1390] Send the price information and product identifier to the server.
[1391] Inputs: Price information and product identifiers.
[1392] Output: Sending data to the server.
[1393] Step 3:
[1394] A server receives the price information and the product identifier and queries a price database of multiple online stores.
[1395] Inputs: Price information and product identifiers.
[1396] Output: Price data for each online store.
[1397] Step 4:
[1398] Get prices from each online store and find the cheapest price.
[1399] Input: Online store price data.
[1400] Output: Cheapest price information.
[1401] Step 5:
[1402] The server recognizes the user's emotions using an emotion recognition engine and adjusts the cheapest information.
[1403] Input: Voice data and lowest price information.
[1404] Output: Adjusted lowest price information.
[1405] Step 6:
[1406] The server transmits the adjusted lowest price information to the terminal.
[1407] Input: Adjusted lowest price information.
[1408] Output: Sending price information to the terminal.
[1409] Step 7:
[1410] The device (smart glasses) converts the cheapest price information into text for voice output and notifies the user.
[1411] Input: Adjusted lowest price information.
[1412] Output: Audio notification to the user.
[1413] By following these steps, the system performs a series of processes, from voice input to data acquisition, location identification, currency conversion, price comparison, and emotion recognition, to provide users with an advanced shopping experience.
[1414] (Application example 2)
[1415] Next, a description will be given of Application Example 2. In the following description, the data processing device 12 will be referred to as a "server" and the smart glasses 214 will be referred to as a "terminal."
[1416] Conventional voice guidance systems using smart glasses could provide guidance based on user requests, but they could not adjust the guidance content according to the user's emotions. As a result, uniform guidance and information was provided without taking the user's emotional state into consideration, resulting in a suboptimal user experience. Particularly when shopping in a store, there was an issue of not being able to provide appropriate guidance when the user was confused or in a hurry.
[1417] The specific processing by the specific processing unit 290 of the data processing device 12 in the application example 2 is realized by the following means.
[1418] In this invention, the server includes emotion recognition means for recognizing emotions from the user's voice and facial expressions, means for converting voice input into text using voice recognition technology, and means for retrieving product location information from a database based on the analysis results, thereby enabling personalized voice guidance and information provision that reflects the user's emotional state.
[1419] "User" means an individual who operates a system or device.
[1420] "Voice input" refers to commands, questions, etc. that a user communicates to a system by speaking.
[1421] "Voice recognition technology" is a technology that analyzes voice signals and converts them into text data.
[1422] "Text conversion" refers to the process of converting captured voice input into text data.
[1423] "Request analysis" is the process of understanding the intent and content of a user's request from the converted text of their voice.
[1424] "Location information" refers to information that indicates where a particular product or section is located within a store.
[1425] A "database" is an information system that structures and stores a collection of related data so that it can be easily searched and retrieved.
[1426] "Voice guidance" refers to a method of providing information to users using a voice source.
[1427] "Emotion recognition means" refers to technology or equipment for detecting and classifying a user's emotional state from data such as the user's voice and facial expressions.
[1428] "Route information" refers to information relating to directions and directions from a specific starting point to a destination.
[1429] An "external API" is an interface for connecting with external software applications and services.
[1430] An "exchange rate" is a ratio that represents the exchange value between different currencies.
[1431] "Price conversion" refers to the process of converting product prices into other currency units.
[1432] "Facial expressions" refer to the movements and appearances that appear on the surface of the face and indicate emotions and intentions.
[1433] "Adjusting guidance content" means appropriately changing the information and direction guidance provided depending on the user's situation and emotions.
[1434] This invention relates to a system that allows users to use smart glasses in a store to receive voice guidance and obtain product information. Specifically, this system provides a more personalized shopping experience by combining voice input, voice recognition, emotion recognition, and location-based services.
[1435] Voice guidance function
[1436] The server uses the microphone built into the smart glasses to capture the user's voice input. The voice input is converted into text using voice recognition technology, and the user's request is analyzed from that text. Based on the results of this analysis, the server retrieves the product's location information from a database and provides the retrieved location information as voice guidance. An emotion recognition means is used to recognize the user's emotions from their voice and facial expressions, and the guidance content is adjusted based on the recognition results. This makes it possible to provide more detailed guidance if the user is confused, and faster guidance if the user is in a hurry.
[1437] Real-time location information function
[1438] The server uses sensors (such as GPS and WiFi triangulation) built into the smart glasses to identify the user's current location in real time. Based on the identified current location, the server generates route information to the product section and provides that route information as voice guidance. By adjusting the route guidance based on the user's location information, the user can move smoothly through the store. During route guidance, emotion recognition means is used to analyze the user's emotions and adjust the guidance content appropriately, providing a more personalized service.
[1439] Currency conversion function
[1440] The server retrieves the latest exchange rates from an external API, converts the product price into the user's home currency using the obtained exchange rate, and provides the converted price information via voice. During the price notification, the system recognizes the user's emotions and adjusts the guidance, such as providing a detailed explanation, if it detects doubt or confusion.
[1441] Specific examples
[1442] For example, if a user speaks to a smart glass in a store and asks, "Where is the shampoo shelf?", the smart glass will recognize the voice and send it to the server. The server will convert the voice into text and identify the location of the shampoo shelf. The emotion recognition means will analyze the user's emotions, and if the user appears confused, the smart glass will provide polite guidance such as, "Go 10 meters to the right and then turn left at the next shelf." On the other hand, if the system detects that the user is in a hurry, it will provide simple guidance such as, "It's 10 meters to the right."
[1443] If you pick up a product and ask, "How much is this product?", the device will retrieve price information and announce in voice, "The price of this shampoo is 5 euros (approximately 620 yen)." If you ask, "How much does this shampoo cost online?", the device will also provide the latest information on online prices, telling you, "It's being sold online for 4.5 euros." At the same time, it can analyze the user's emotions and adjust the guidance accordingly.
[1444] Prompt Sentence Examples
[1445] "Where's the shampoo shelf?"
[1446] "How much is this item?"
[1447] "How much does this shampoo cost online?"
[1448] The flow of the specific processing in the application example 2 will be described with reference to FIG.
[1449] Step 1:
[1450] The user provides voice input to the smart glasses.
[1451] Input: User speech (e.g., "Where is the shampoo shelf?")
[1452] How it works: The user's voice is captured using the microphone built into the smart glasses.
[1453] Output: Analog audio data
[1454] Step 2:
[1455] The device converts the voice data into a digital format.
[1456] Input: Analog audio data
[1457] How it works: The audio processing module inside the smart glasses converts analog audio data into digital audio data.
[1458] Output: Digital audio data
[1459] Step 3:
[1460] The terminal transmits digital audio data to the server.
[1461] Input: Digital audio data
[1462] How it works: Smart glasses send digital audio data to a server.
[1463] Output: Digital audio data sent to the server
[1464] Step 4:
[1465] The server converts the digital voice data into text.
[1466] Input: Digital audio data
[1467] How it works: The server's speech recognition technology analyzes the digital voice data and converts it into corresponding text.
[1468] Output: The transformed text (e.g., "Where is the shampoo shelf?")
[1469] Step 5:
[1470] The server parses the user's request from the text.
[1471] Input: Translated text
[1472] How it works: The server parses the text and determines what the user is requesting (e.g., I want to know where the shampoo shelf is).
[1473] Output: Analysis results (e.g., obtain product location information)
[1474] Step 6:
[1475] The server retrieves the product location information from the database.
[1476] Input: Analysis results (request for product location information)
[1477] Operation: The server queries the database and obtains the location information of the relevant product.
[1478] Output: Product location information (e.g., go 10 meters to the right and turn left at the next shelf)
[1479] Step 7:
[1480] The server recognizes the user's emotions using an emotion recognition means.
[1481] Input: User's voice and facial expression data
[1482] How it works: The server uses emotion recognition technology to analyze the user's emotions from their voice and facial expressions.
[1483] Output: User's emotional state (e.g. confused, in a hurry)
[1484] Step 8:
[1485] The server adjusts the guidance content based on the emotion.
[1486] Input: Product location, user emotional state
[1487] How it works: The server takes into account the user's emotional state and adjusts the guidance content (e.g., polite guidance if the user is confused, brief guidance if the user is in a hurry).
[1488] Output: Adjusted guidance text (e.g. "Go right 10 meters and turn left at the next ledge")
[1489] Step 9:
[1490] The server sends the adjusted guidance text to the smart glasses.
[1491] Input: Adjusted guidance text
[1492] Operation: The server sends the adjusted guidance text to the smart glasses.
[1493] Output: Guidance text sent to smart glasses
[1494] Step 10:
[1495] The terminal provides the user with the guidance text by voice.
[1496] Input: Guidance text
[1497] How it works: The audio output device of the smart glasses plays the guidance text as audio.
[1498] Output: Guidance speech (e.g. "Go 10 meters to the right and turn left at the next shelf")
[1499] The specific processing unit 290 transmits the result of the specific processing to the smart glasses 214. In the smart glasses 214, the control unit 46A causes the speaker 240 to output the result of the specific processing. The microphone 238 acquires audio indicating a user input regarding the result of the specific processing. The control unit 46A transmits audio data indicating the user input acquired by the microphone 238 to the data processing device 12. In the data processing device 12, the specific processing unit 290 acquires the audio data.
[1500] The data generation model 58 is a so-called generative AI (Artificial Intelligence). An example of the data generation model 58 is ChatGPT (Internet Search<URL: https: / / openai.com / blog / chatgpt> ), Gemini (Internet search <url: https: gemini.google.com ?hl="ja">) and other generation AIs. The data generation model 58 is obtained by performing deep learning on a neural network. A prompt including an instruction is input to the data generation model 58, and inference data such as voice data indicating voice, text data indicating text, and image data indicating an image is also input. The data generation model 58 performs inference on the input inference data in accordance with the instruction indicated by the prompt, and outputs the inference result in a data format such as voice data and text data. Here, inference refers to, for example, analysis, classification, prediction, and / or summarization.
[1501] In the above embodiment, an example in which the specific processing is performed by the data processing device 12 has been given, but the technology of the present disclosure is not limited to this, and the specific processing may be performed by the smart glasses 214.
[1502] [Third embodiment]
[1503] FIG. 5 shows an example of the configuration of a data processing system 310 according to the third embodiment.
[1504] 5, the data processing system 310 includes the data processing device 12 and a headset terminal 314. An example of the data processing device 12 is a server.
[1505] The data processing device 12 includes a computer 22, a database 24, and a communication I / F 26. The computer 22 is an example of a "computer" according to the technology of the present disclosure. The computer 22 includes a processor 28, a RAM 30, and a storage 32. The processor 28, the RAM 30, and the storage 32 are connected to a bus 34. The database 24 and the communication I / F 26 are also connected to the bus 34. The communication I / F 26 is connected to a network 54. Examples of the network 54 include a WAN (Wide Area Network) and / or a LAN (Local Area Network).
[1506] The headset type terminal 314 includes a computer 36, a microphone 238, a speaker 240, a camera 42, a communication I / F 44, and a display 343. The computer 36 includes a processor 46, a RAM 48, and a storage 50. The processor 46, the RAM 48, and the storage 50 are connected to a bus 52. The microphone 238, the speaker 240, the camera 42, and the display 343 are also connected to the bus 52.
[1507] The microphone 238 receives instructions and the like from the user 20 by receiving voice uttered by the user 20. The microphone 238 captures the voice uttered by the user 20, converts the captured voice into audio data, and outputs it to the processor 46. The speaker 240 outputs audio in accordance with instructions from the processor 46.
[1508] Camera 42 is a small digital camera equipped with an optical system including a lens, aperture, and shutter, and an imaging element such as a CMOS (Complementary Metal-Oxide-Semiconductor) image sensor or a CCD (Charge Coupled Device) image sensor, and captures images of the surroundings of user 20 (for example, an imaging range defined by an angle of view equivalent to the field of vision of a typical healthy person).
[1509] The communication I / F 44 is connected to a network 54. The communication I / Fs 44 and 26 are responsible for the exchange of various information between the processor 46 and the processor 28 via the network 54. The exchange of various information between the processor 46 and the processor 28 using the communication I / Fs 44 and 26 is carried out in a secure state.
[1510] Fig. 6 shows an example of the main functions of the data processing device 12 and the headset type terminal 314. As shown in Fig. 6, in the data processing device 12, a specific process is performed by the processor 28. A specific process program 56 is stored in the storage 32.
[1511] The specific processing program 56 is an example of a "program" according to the technology of the present disclosure. The processor 28 reads the specific processing program 56 from the storage 32 and executes the read specific processing program 56 on the RAM 30. The specific processing is realized by the processor 28 operating as a specific processing unit 290 in accordance with the specific processing program 56 executed on the RAM 30.
[1512] The storage 32 stores a data generation model 58 and an emotion identification model 59. The data generation model 58 and the emotion identification model 59 are used by the identification processing unit 290.
[1513] In the headset type terminal 314, a reception output process is performed by the processor 46. A reception output program 60 is stored in the storage 50. The processor 46 reads the reception output program 60 from the storage 50 and executes the read reception output program 60 on the RAM 48. The reception output process is realized by the processor 46 operating as the control unit 46A in accordance with the reception output program 60 executed on the RAM 48.
[1514] Next, a description will be given of the identification process performed by the identification processing unit 290 of the data processing device 12. In the following description, the data processing device 12 will be referred to as the "server" and the headset type terminal 314 will be referred to as the "terminal."
[1515] The present invention is a voice guidance system using smart glasses that solves problems such as language barriers when shopping abroad, product location, currency conversion, price comparison, etc. Hereinafter, an embodiment of the present invention will be described in detail.
[1516] Voice guidance function
[1517] Device (smart glasses)
[1518] 1. The user asks the smart glasses, "Where is the shampoo shelf?"
[1519] 2. The smart glasses capture voice input via a built-in microphone and convert the voice data into a digital format.
[1520] 3. Send the digital audio data to the server.
[1521] server
[1522] 1. Receives digital voice data and converts it into text using speech recognition technology.
[1523] 2. Parse the user's request (in this case, the shelf location of shampoo) from the converted text.
[1524] 3. Obtain the location information of the shampoo shelves from the database and generate text for voice guidance.
[1525] 4. The generated voice guidance text is sent to the terminal.
[1526] Terminal
[1527] 1. Receive the voice guidance text sent from the server and use the voice output device to instruct the user, "Go 10 meters to the right and turn left at the next shelf."
[1528] Real-time location information function
[1529] Terminal
[1530] 1. Determine the user's current location using built-in sensors (GPS, WiFi triangulation, etc.).
[1531] 2. The determined current location data is sent to the server.
[1532] server
[1533] 1. Receive current location data and determine the user's location based on that data.
[1534] 2. Obtain the location information of the desired product from the database and generate route information from the current location to the destination.
[1535] 3. The generated route information is sent to the terminal.
[1536] Terminal
[1537] 1. Route information sent from the server is provided to the user via voice.
[1538] Currency conversion function
[1539] Terminal
[1540] 1. Scan the product barcode to get price information.
[1541] 2. Send price information and user's home currency information to the server.
[1542] server
[1543] 1. Receive price information and get the latest exchange rates using an external API.
[1544] 2. Convert the product price into the user's home currency using the obtained exchange rate.
[1545] 3. The converted price information is sent to the terminal.
[1546] Terminal
[1547] 1. The converted price information is notified to the user via voice, such as "The price of this shampoo is 500 yen."
[1548] Price comparison feature
[1549] Terminal
[1550] 1. Scan the product barcode to get price information.
[1551] 2. Send the price information and product identifier to the server.
[1552] server
[1553] 1. Receives price information and product identifiers and queries price databases from multiple online stores.
[1554] 2. Get prices from each online store and identify the lowest price.
[1555] 3. Send the cheapest price information to the terminal.
[1556] Terminal
[1557] 1. The lowest price information is announced to the user via voice, saying, "It's on sale online for 450 yen."
[1558] Specific examples
[1559] If a user is in a supermarket in France and asks the smart glasses, "Where is the shampoo shelf?", the smart glasses will respond with, "Go 10 meters to the right and turn left at the next shelf." After reaching the shampoo section, the user picks up an item and wants to check the price. They can ask, "How much is this shampoo?" and the smart glasses will respond, "This shampoo costs 5 euros (approximately 620 yen)." Furthermore, to compare the price with online prices, they can ask, "How much does this shampoo cost online?" and the smart glasses will respond, "It's selling for 4.5 euros online." Through this process, users can enjoy an efficient and comfortable shopping experience.
[1560] The processing flow will be explained below.
[1561] Voice guidance function
[1562] Specific processing steps from voice input to guidance
[1563] Step 1:
[1564] The user asks the smart glasses, "Where is the shampoo shelf?"
[1565] Step 2:
[1566] The terminal uses a built-in microphone to capture the user's voice and converts the voice into digital data.
[1567] Step 3:
[1568] The terminal transmits the converted digital audio data to the server.
[1569] Step 4:
[1570] The server receives the digital voice data and converts the voice data into text using voice recognition technology.
[1571] Step 5:
[1572] The server parses the text and understands the user's request (in this case, a request to know the shelf location of shampoo).
[1573] Step 6:
[1574] The server retrieves the location information of the shampoo shelves from the database.
[1575] Step 7:
[1576] Based on the acquired location information, the server generates text for voice guidance such as, "Go 10 meters to the right and turn left at the next shelf."
[1577] Step 8:
[1578] The server transmits the generated voice guidance text to the terminal.
[1579] Step 9:
[1580] The device converts the received voice guidance text into speech and begins providing guidance through the speaker.
[1581] Step 10:
[1582] The user moves as instructed.
[1583] Real-time location information function
[1584] A processing step to guide you from your current location to the product location
[1585] Step 1:
[1586] The user speaks to the smart glasses, saying, "Tell me where I am."
[1587] Step 2:
[1588] The device determines the user's current location using built-in sensors (e.g., GPS and WiFi triangulation).
[1589] Step 3:
[1590] The terminal transmits the identified current location data to the server.
[1591] Step 4:
[1592] The server receives the location data and confirms the user's current location.
[1593] Step 5:
[1594] The server retrieves route information to the product section from the database.
[1595] Step 6:
[1596] The server generates route information from the current location to the destination product section.
[1597] Step 7:
[1598] The server transmits the generated route information to the terminal.
[1599] Step 8:
[1600] The terminal converts the received route information into text for voice guidance.
[1601] Step 9:
[1602] The device converts the voice guidance text into speech and begins providing guidance through the speaker.
[1603] Step 10:
[1604] The user moves along the guided route.
[1605] Currency conversion function
[1606] Processing steps to convert product prices into your home currency
[1607] Step 1:
[1608] Users scan the barcode of the product they pick up with the smart glasses.
[1609] Step 2:
[1610] The terminal scans the barcode and obtains the product's price information.
[1611] Step 3:
[1612] The terminal transmits product price information and the user's home currency information to the server.
[1613] Step 4:
[1614] The server receives the price information and uses an external API to get the latest exchange rates.
[1615] Step 5:
[1616] The server uses the obtained exchange rate to convert the product price into the user's home currency.
[1617] Step 6:
[1618] The server transmits the converted price information to the terminal.
[1619] Step 7:
[1620] The terminal converts the received converted price information into text for voice output.
[1621] Step 8:
[1622] The device converts the text to speech and announces the price through the speaker.
[1623] Step 9:
[1624] The user receives the price information sent via the smart glasses.
[1625] Price comparison feature
[1626] Processing steps to compare in-store and online prices
[1627] Step 1:
[1628] Users scan the barcode of the product they pick up with the smart glasses.
[1629] Step 2:
[1630] The terminal scans the barcode and obtains the product's price information.
[1631] Step 3:
[1632] The terminal transmits the price information and the product identifier to the server.
[1633] Step 4:
[1634] The server receives the price information and the product identifier and queries the online store's price database.
[1635] Step 5:
[1636] The server retrieves the prices from each online store and identifies the lowest price.
[1637] Step 6:
[1638] The server transmits the cheapest price information to the terminal.
[1639] Step 7:
[1640] The terminal converts the received information about the cheapest price into text for voice output.
[1641] Step 8:
[1642] The device converts the text to speech and announces price information through the speaker.
[1643] Step 9:
[1644] The user receives price information from the smart glasses and decides to purchase.
[1645] Through the above processing steps, this system supports users in shopping abroad efficiently and comfortably.
[1646] Example 1
[1647] Next, a description will be given of Example 1. In the following description, the data processing device 12 will be referred to as a "server" and the headset type terminal 314 will be referred to as a "terminal."
[1648] In today's global society, shopping in a multilingual environment can be extremely difficult for users. When shopping abroad, language barriers make it particularly difficult to accurately determine product locations. It's also difficult to convert prices displayed in local currencies into one's own currency, or to quickly compare prices across different stores. A system that solves these problems is needed.
[1649] The specific processing by the specific processing unit 290 of the data processing device 12 in the first embodiment is realized by the following means.
[1650] In this invention, the server includes: means for acquiring a user's voice input; means for converting the acquired voice input into digital format; means for transmitting the converted digital voice input to the server; means for converting the voice input into text on the server using voice recognition technology; means for analyzing the user's request from the converted text; means for acquiring product location information from a database based on the analysis results; means for converting the acquired location information into text for voice guidance and transmitting it to the terminal; and means for converting the transmitted text for voice guidance into speech and providing guidance to the user. This allows users to easily find products even in different language environments. Furthermore, by providing functions such as real-time location information identification and route guidance, currency conversion using the latest exchange rates, and price comparison across multiple stores, the server achieves an efficient and convenient shopping experience.
[1651] The "means for acquiring user's voice input" is a combination of hardware and software that recognizes the voice uttered by the user and that the terminal acquires as digital data.
[1652] The "means for converting to digital form" refers to an audio signal processing device and conversion algorithm for converting analog audio signals into digital data.
[1653] The "means for transmitting to a server" refers to a communication module and protocol that allows the terminal to transfer digital data over the Internet to a remote server.
[1654] "Means for converting speech to text using speech recognition technology" refers to software and algorithms that analyze speech data on a server and convert it into corresponding text.
[1655] The "means for analyzing user requests" is a natural language processing system that extracts and analyzes the user's intentions and requests from the text obtained by speech recognition.
[1656] The "means for acquiring product location information from a database" is a query processing system for searching and acquiring product placement information from a database based on the analyzed user request.
[1657] The "means for converting the information into text for voice guidance and sending it to the terminal" is a server-side function for converting the acquired information into text in a format that is easy for the user to understand and sending this text to the terminal.
[1658] The "means for converting the text information into audio and providing the user with guidance" refers to the software and hardware within the terminal for playing back the received text information as audio.
[1659] "Means of real-time identification" refers to a system that instantly identifies a user's current location using sensor technology such as GPS and Wi-Fi triangulation.
[1660] The "means for generating route information" refers to algorithms and software for calculating a route from the specified current position to the destination and generating that information.
[1661] The "means for obtaining the latest exchange rates from an external data source" is an interface for obtaining the latest currency exchange information from an external financial data provider service.
[1662] The "means for converting the product price using the exchange rate" is a calculation algorithm for converting the product price into the user's home currency using the acquired exchange rate information.
[1663] MODE FOR CARRYING OUT THE INVENTION
[1664] The present invention is a voice guidance system using smart glasses that solves problems such as language barriers when shopping abroad, product location, currency conversion, price comparison, etc. Hereinafter, an embodiment of the present invention will be described in detail.
[1665] Overall system overview
[1666] This system is composed of a combination of smart glasses (terminals), a cloud server, a database, and an external API. The smart glasses receive voice input from the user and send it to the cloud server. The cloud server uses voice recognition technology to convert the voice data into text, and then analyzes the user's request through natural language processing. Based on the analysis results, it provides information such as product location, currency conversion, and price comparison.
[1667] Hardware and software used
[1668] Device (smart glasses)
[1669] Built-in microphone: Captures the user's voice.
[1670] Communication module: Sends digital data to a cloud server (e.g., Wi-Fi, mobile data).
[1671] Voice output device: Converts guided text into voice.
[1672] Sensors: Determine current location (e.g. GPS, Wi-Fi triangulation).
[1673] server
[1674] Speech recognition software: Google Cloud Speech-to-Text API.
[1675] Natural language processing engine: Uses SpaCy or similar.
[1676] Database: Uses MySQL to manage product location and price information.
[1677] External API: Uses Open Exchange Rates API for currency conversion.
[1678] Program processing
[1679] 1. Voice guidance function
[1680] The user asks the smart glasses, "Where is the shampoo shelf?"
[1681] The smart glasses pick up sound through a built-in microphone and convert it into a digital format.
[1682] The digital audio data is transmitted to a cloud server.
[1683] The server uses speech recognition software to convert the voice data into text.
[1684] The text is analyzed using a natural language processing engine to understand the request.
[1685] Retrieve product location information from the database and generate guide text.
[1686] The guidance text is sent to the smart glasses and the user is guided by voice.
[1687] 2. Real-time location information function
[1688] The smart glasses use built-in sensors to determine the user's current location.
[1689] The current location data is sent to a cloud server.
[1690] The server generates route information to the product section based on the current location data.
[1691] The generated route information is sent to smart glasses and guidance is provided via voice.
[1692] 3. Currency conversion function
[1693] Smart glasses scan product barcodes to obtain price information.
[1694] Price information and the user's home currency information are sent to the server.
[1695] The server uses an external API to get the latest exchange rates.
[1696] Convert product prices into the user's home currency using exchange rates.
[1697] The converted price information is sent to the smart glasses and notified via voice.
[1698] 4. Price comparison feature
[1699] The smart glasses scan the product's barcode and retrieve price information.
[1700] Send the price information and product identifier to the server.
[1701] The server queries a price database of multiple online stores.
[1702] The lowest price information is identified and sent to the smart glasses.
[1703] The cheapest price information is notified to the user by voice.
[1704] Specific examples
[1705] If a user is in a supermarket in France and asks the smart glasses, "Where is the shampoo shelf?", the smart glasses will respond with, "Go 10 meters to the right and turn left at the next shelf." After reaching the shampoo section, the user picks up an item and wants to check the price. They can ask, "How much is this shampoo?" and the smart glasses will respond, "This shampoo costs 5 euros (approximately 620 yen)." Furthermore, to compare the price with online prices, they can ask, "How much does this shampoo cost online?" and the smart glasses will respond, "It's selling for 4.5 euros online." Through this process, users can enjoy an efficient and comfortable shopping experience.
[1706] Example prompts to input to the generative AI model
[1707] Please explain in detail the process by which a user queries the smart glasses for product location, product price, currency conversion, and price comparison. Please also include specific hardware, software, data processing, and calculation methods.
[1708] The flow of the identification process in the first embodiment will be described with reference to FIG.
[1709] Step 1:
[1710] A user speaks to the smart glasses, saying, "Where is the shampoo shelf?" The user's voice is picked up by the smart glasses' built-in microphone. This voice input is an analog signal, so it is converted into digital format. Specifically, A / D conversion is performed to convert the voice waveform into a digital signal. The input is the user's voice signal, and the output is digital voice data.
[1711] Step 2:
[1712] The device transmits digital audio data to a cloud server via Wi-Fi or a mobile data network using a communication module. The input is digital audio data, and the output is confirmation of data transmission to the server.
[1713] Step 3:
[1714] The server receives the digital voice data and converts it into text using speech recognition technology. The Google Cloud Speech-to-Text API is used. The server analyzes the voice data and generates corresponding text data. The input is digital voice data, and the output is the converted text data.
[1715] Step 4:
[1716] The server sends the converted text data to a natural language processing engine to analyze the user's request. For example, SpaCy can be used to extract the meaning of the text. The input is text data, and the output is the analysis result of the user's request ("where shampoo is on the shelf"). Specifically, the keyword "shampoo" is extracted from the text, and the request content is determined based on that.
[1717] Step 5:
[1718] The server retrieves product location information from the database based on the analysis results. In this example, it queries a MySQL database to retrieve the location information of the shampoo shelf. The input is the analysis result of the user request, and the output is the product location information. The database query is executed and the location information is returned.
[1719] Step 6:
[1720] The server converts the acquired location information into text for voice guidance and sends it to the terminal. An algorithm for generating voice guidance text is applied to generate text such as "Go 10 meters to the right and turn left at the next shelf." The input is the product location information, and the output is the text for voice guidance. The generated text is sent to the terminal as an HTTP response.
[1721] Step 7:
[1722] The device receives the text for voice guidance sent from the server. The received text is then used to provide guidance to the user using a voice output device. Specifically, the text is converted to speech using the Google Text-to-Speech API. The input is the text for voice guidance, and the output is the voice guidance. "Go 10 meters to the right and turn left at the next shelf," is the voice guidance that comes out of the smart glasses.
[1723] (Application example 1)
[1724] Next, a description will be given of Application Example 1. In the following description, the data processing device 12 will be referred to as a "server" and the headset type terminal 314 will be referred to as a "terminal."
[1725] The present invention aims to solve problems associated with shopping abroad, such as language barriers, product location, currency conversion, and price comparisons. Specifically, the problem is to provide a support system that enables users to intuitively and efficiently find products in physical stores, understand prices, and purchase them at the most advantageous prices.
[1726] The specific processing by the specific processing unit 290 of the data processing device 12 in the application example 1 is realized by the following means.
[1727] In this invention, the server includes means for acquiring a user's voice input, means for converting the acquired voice input into text using voice recognition technology, means for analyzing the user's request from the converted text, means for acquiring product location information from a database based on the analysis result, means for providing the acquired location information as voice guidance, means for converting product prices into the user's home currency using a currency conversion function, and means for acquiring price information from multiple online stores using a price comparison function to identify the lowest price. This allows users to easily find products even in stores in a foreign country, and also makes it easy to convert currencies and compare prices.
[1728] "User voice input" refers to voice commands issued by a user to a device.
[1729] "Speech recognition technology" refers to machine learning and natural language processing technology for converting speech into text data.
[1730] "Means for converting to text" refers to a device or software that performs a process of converting captured voice input into digital data and then parsing that digital data into text.
[1731] "Means for analyzing user requests" refers to a device or software that processes voice commands converted into text to understand the information or action the user is seeking.
[1732] "Means for obtaining product location information from a database" refers to a device or software for searching and obtaining information about the product location stored in a database.
[1733] "Means for providing audio guidance" refers to a device or software that provides audio guidance to the user based on analyzed information and acquired data.
[1734] "Currency Conversion Facility" means a device or software that performs a process to convert product prices displayed in one currency into another currency selected by the user.
[1735] "Price comparison tool" means a device or software that performs a process to obtain and compare price information from multiple sources or online stores for a particular product to identify the most favorable price.
[1736] "Means for determining current location in real time" refers to devices or software that obtain a user's current location in real time using technologies such as GPS or WiFi triangulation.
[1737] The "means for generating route information" refers to a device or software for calculating the optimal route from the user's current location to the location of the desired product and generating the route information.
[1738] "Means for obtaining the latest exchange rates from an external API" refers to a device or software for obtaining exchange rate information from an external server via the Internet.
[1739] The present invention relates to a system that uses a voice guidance system in a specified target environment to provide users with various information such as product location information, currency conversion, price comparison, etc. Specific embodiments of the system are described below.
[1740] Generating a Program
[1741] The system consists of the following main functions:
[1742] 1. Voice Input Capture and Recognition
[1743] 2. Location information acquisition and guidance
[1744] 3. Currency conversion function
[1745] 4. Price comparison feature
[1746] Program processing explanation
[1747] Voice Input Capture and Recognition
[1748] The device (e.g., smart glasses) captures the user's voice input with a microphone and converts the voice data into a digital format. The converted digital voice data is sent to the server using the os and wave libraries. The server receives the voice data via a flask-based endpoint and converts it into text using the pydub and speech_recognition libraries.
[1749] Next, a natural language processing library such as nltk is used to analyze the user's request, and the next steps are taken based on the analysis results.
[1750] Location information acquisition and guidance
[1751] The device acquires its current location in real time using its built-in GPS and WiFi triangulation, and sends it to the server. The server processes the location data using the geopy library and generates the optimal route to the product section based on the user's current location. This route information is then provided as audio guidance.
[1752] Currency conversion function
[1753] When a user scans a product's barcode, the device retrieves the price information and sends it to the server. The server then uses an external API to retrieve the latest exchange rate and converts the product price based on the obtained rate. The converted price information is retrieved using the requests library and notified to the user via voice.
[1754] Price comparison feature
[1755] When a user scans a product's barcode, the device sends the price information to a server, which queries a price database of multiple online stores to identify the lowest price, which is then provided to the user via voice.
[1756] Specific examples
[1757] If a user is in a supermarket in France and asks the smart glasses, "Where is the shampoo shelf?", the smart glasses will respond with, "Go 10 meters to the right and turn left at the next shelf." After reaching the shampoo section, the user picks up an item and wants to check the price. They can ask, "How much is this shampoo?" and the smart glasses will respond, "This shampoo costs 5 euros (approximately 620 yen)." Furthermore, to compare the price with online prices, they can ask, "How much does this shampoo cost online?" and the smart glasses will respond, "It's selling for 4.5 euros online." Through this process, users can enjoy an efficient and comfortable shopping experience.
[1758] Prompt Sentence Examples
[1759] "This smartglasses-based shopping app allows users to easily navigate in-store, convert currencies, compare prices, and more, all by voice. For example, if a user is in a supermarket in France and asks the smartglasses, 'Where is the shampoo shelf?', the smartglasses will respond with, 'Go 10 meters to the right and then turn left at the next shelf.' Please generate a program that achieves this function."
[1760] The flow of the specific processing in the application example 1 will be described with reference to FIG.
[1761] Step 1:
[1762] The device (smart glasses) captures the user's voice input with a microphone. The user speaks, "Where is the shampoo shelf?", and this voice signal is acquired. The voice input is converted into digital voice data using the OS and Wave libraries. The output is digital voice data.
[1763] Step 2:
[1764] The device sends digital audio data to the server. This is done using the requests library. The input is the digital audio data, and the output is a request to the server to send the audio data.
[1765] Step 3:
[1766] The server receives the received digital audio data at the flask endpoint and starts analyzing the data. The input is the digital audio data. The server converts it into text using the pydub and speech_recognition libraries. The output is the converted text data.
[1767] Step 4:
[1768] The server analyzes the converted text data using a natural language processing library such as nltk to identify the user's request. The converted text data is input, and the user's request extracted through analysis is output. Specifically, "the location of the shampoo shelf" is identified as the user's request.
[1769] Step 5:
[1770] The server retrieves the shampoo shelf location information from the database. The input is the user's request (the shampoo shelf location), and the output is the product location information. This location information indicates where the product is located in the store.
[1771] Step 6:
[1772] The server calculates the location using the geopy library based on the product location information and generates route information. The input is the product location information and the user's current location, and the output is route information. Specifically, the guidance will be something like "Go 10 meters to the right and turn left at the next shelf."
[1773] Step 7:
[1774] The server converts the generated route information into text for voice guidance and sends it to the terminal. The input is route information, and a process is performed to convert this information into text for voice guidance. The output is text for voice guidance.
[1775] Step 8:
[1776] The terminal notifies the user of the voice guidance text received from the server using a voice output device. The input is the voice guidance text, and the output is the voice guidance ("Go 10 meters to the right and turn left at the next shelf").
[1777] Step 9:
[1778] When a user scans a product barcode, the terminal obtains the price information and sends it to the server. The input is the price information output by scanning the product barcode.
[1779] Step 10:
[1780] The server uses an external API to obtain the latest exchange rate based on the acquired price information. The input is the price information, and the output is the latest exchange rate.
[1781] Step 11:
[1782] The server converts the price into the user's home currency based on the exchange rate. The input is the acquired exchange rate and price information, and the output is the converted price information.
[1783] Step 12:
[1784] The server converts the converted price information into text for voice guidance and sends it to the terminal. The input is the converted price information, and the output is the text for voice guidance.
[1785] Step 13:
[1786] The terminal notifies the user of the voice guidance text received from the server using a voice output device. The input is the voice guidance text, and the output is the voice guidance ("The price of this shampoo is XX yen").
[1787] Step 14:
[1788] When a user speaks "What is the price of this shampoo online?", the device sends this speech input to the server. The input is the user's speech input, and the output is digital voice data.
[1789] Step 15:
[1790] The server analyzes the received voice data to identify it as a price comparison request and queries a price database of multiple online stores. The input is the analyzed request, and the output is the price information of each online store.
[1791] Step 16:
[1792] The server identifies the cheapest price information, converts it into a voice guidance text, and sends it to the terminal. The input is the price information of the online store, and the output is a voice guidance text containing the cheapest price information.
[1793] Step 17:
[1794] The terminal notifies the user of the voice guidance text received from the server using a voice output device. The input is the voice guidance text, and the output is the voice guidance ("This item is sold online for XX yen").
[1795] Furthermore, an emotion engine that estimates the user's emotion may be further combined. That is, the identification processing unit 290 may estimate the user's emotion using the emotion identification model 59, and perform identification processing using the user's emotion.
[1796] The present invention provides a system that combines a voice guidance system using smart glasses with an emotion engine that recognizes the user's emotions, providing a more personalized shopping experience. The following describes in detail an embodiment of the present invention.
[1797] Voice guidance function
[1798] Device (smart glasses)
[1799] 1. The user asks the smart glasses, "Where is the shampoo shelf?"
[1800] 2. The smart glasses capture voice input via a built-in microphone and convert the voice data into a digital format.
[1801] 3. Send the digital audio data to the server.
[1802] server
[1803] 1. Receive digital voice data and convert the voice data into text using voice recognition technology.
[1804] 2. Understand the user's request from the converted text (in this case, the request to know the shelf location of shampoo).
[1805] 3. Obtain the location information of the shampoo shelves from the database and generate text for voice guidance.
[1806] 4. Use the emotion engine to recognize emotions from the user's voice and adjust the guidance content.
[1807] 5. The adjusted voice guidance text is sent to the device.
[1808] Terminal
[1809] 1. Receive the voice guidance text sent from the server and use the voice output device to instruct the user, "Go 10 meters to the right and turn left at the next shelf."
[1810] Real-time location information function
[1811] Terminal
[1812] 1. Determine the user's current location using built-in sensors (GPS, WiFi triangulation, etc.).
[1813] 2. The determined current location data is sent to the server.
[1814] server
[1815] 1. Receive current location data and determine the user's location based on that data.
[1816] 2. Obtain the location information of the desired product from the database and generate route information from the current location to the destination.
[1817] 3. Use an emotion engine to recognize emotions from the user's voice and facial expressions and adjust the guidance content accordingly.
[1818] 4. The adjusted route information is sent to the terminal.
[1819] Terminal
[1820] 1. The route information sent from the server is converted into text for voice guidance, and guidance to the user begins.
[1821] Currency conversion function
[1822] Terminal
[1823] 1. Scan the product barcode to get price information.
[1824] 2. Send price information and user's home currency information to the server.
[1825] server
[1826] 1. Receive price information and get the latest exchange rates using an external API.
[1827] 2. Convert the product price into the user's home currency using the obtained exchange rate.
[1828] 3. Use an emotion engine to recognize emotions from the user's voice and facial expressions and adjust the content of price information notifications.
[1829] 4. The adjusted price information is sent to the terminal.
[1830] Terminal
[1831] 1. The converted price information is converted into text for voice output and announced through the speaker: "The price of this shampoo is 500 yen."
[1832] Price comparison feature
[1833] Terminal
[1834] 1. Scan the product barcode to get price information.
[1835] 2. Send the price information and product identifier to the server.
[1836] server
[1837] 1. Receives price information and product identifiers and queries price databases from multiple online stores.
[1838] 2. Get the prices from each online store and identify the cheapest price.
[1839] 3. Use an emotion engine to recognize emotions from the user's voice and facial expressions and adjust the content of the notification about the cheapest price.
[1840] 4. The adjusted lowest price information is sent to the terminal.
[1841] Terminal
[1842] 1. Convert the cheapest price information into text for voice output and announce through the speaker, "It's on sale online for 450 yen."
[1843] Use of emotion engine
[1844] server
[1845] 1. Receives voice data and facial expression data and recognizes the user's emotions using an emotion engine.
[1846] 2. Tailor voice prompts and pricing notifications based on the recognized emotion.
[1847] 3. The adjusted information is sent to the device.
[1848] Specific examples
[1849] If a user is in a supermarket in France and asks the smart glasses, "Where is the shampoo shelf?", the smart glasses will respond with, "Go 10 meters to the right and turn left at the next shelf." After reaching the shampoo section, the user picks up an item and asks, "How much is this shampoo?" to check the price. The smart glasses will respond, "This shampoo costs 5 euros (approximately 620 yen)." Furthermore, if the user asks, "How much does this shampoo cost online?" to compare it with online prices, the smart glasses will respond, "It's selling for 4.5 euros online." This invention utilizes an emotion engine to optimize the user experience, such as providing more attentive guidance if the user shows a confused expression.
[1850] The processing flow will be explained below.
[1851] Voice guidance function
[1852] Specific processing steps from voice input to guidance
[1853] Step 1:
[1854] The user asks the smart glasses, "Where is the shampoo shelf?"
[1855] Step 2:
[1856] The terminal uses a built-in microphone to capture the user's voice and converts the voice into digital data.
[1857] Step 3:
[1858] The terminal transmits the converted digital audio data to the server.
[1859] Step 4:
[1860] The server receives the digital voice data and converts the voice data into text using voice recognition technology.
[1861] Step 5:
[1862] The server parses the text and understands the user's request (in this case, a request to know the shelf location of shampoo).
[1863] Step 6:
[1864] The server retrieves the location information of the shampoo shelves from the database.
[1865] Step 7:
[1866] Based on the acquired location information, the server generates text for voice guidance such as, "Go 10 meters to the right and turn left at the next shelf."
[1867] Step 8:
[1868] The server sends the voice data to an emotion engine to analyze the user's emotion from the voice data.
[1869] Step 9:
[1870] The server receives the emotion recognition results from the emotion engine and adjusts the content and tone of the announcement based on the user's emotion.
[1871] Step 10:
[1872] The server transmits the adjusted voice guidance text to the terminal.
[1873] Step 11:
[1874] The device converts the received voice guidance text into speech and begins providing guidance through the speaker.
[1875] Step 12:
[1876] The user moves as instructed.
[1877] Real-time location information function
[1878] Specific processing steps to guide you from your current location to the product location
[1879] Step 1:
[1880] The user speaks to the smart glasses, saying, "Tell me where I am."
[1881] Step 2:
[1882] The device determines the user's current location using built-in sensors (e.g., GPS and WiFi triangulation).
[1883] Step 3:
[1884] The terminal transmits the identified current location data to the server.
[1885] Step 4:
[1886] The server receives the location data and determines the user's location based on it.
[1887] Step 5:
[1888] The server retrieves route information to the product section from the database.
[1889] Step 6:
[1890] The server generates route information from the current location to the destination product section.
[1891] Step 7:
[1892] The server sends the user's voice data and facial expression data to the emotion engine to analyze the user's emotions.
[1893] Step 8:
[1894] The server receives the emotion recognition results from the emotion engine and adjusts the content and tone of the route guidance based on the emotion.
[1895] Step 9:
[1896] The server transmits the adjusted route information to the terminal.
[1897] Step 10:
[1898] The terminal converts the received route information into text for voice guidance.
[1899] Step 11:
[1900] The terminal converts the voice guidance text into voice and starts providing guidance to the user through the speaker.
[1901] Step 12:
[1902] The user moves along the guided route.
[1903] Currency conversion function
[1904] Specific steps to convert product prices into your home currency
[1905] Step 1:
[1906] Users scan the barcode of the product they pick up with the smart glasses.
[1907] Step 2:
[1908] The terminal scans the barcode and obtains the product's price information.
[1909] Step 3:
[1910] The terminal transmits product price information and the user's home currency information to the server.
[1911] Step 4:
[1912] The server receives the price information and uses an external API to get the latest exchange rates.
[1913] Step 5:
[1914] The server uses the obtained exchange rate to convert the product price into the user's home currency.
[1915] Step 6:
[1916] The server sends the user's voice data and facial expression data to the emotion engine to analyze the user's emotions.
[1917] Step 7:
[1918] The server receives emotion recognition results from the emotion engine and adjusts the content and tone of price information notifications based on the emotion.
[1919] Step 8:
[1920] The server transmits the adjusted price information to the terminal.
[1921] Step 9:
[1922] The terminal converts the received price information into text for voice output.
[1923] Step 10:
[1924] The device converts the text for voice output into speech and announces through the speaker, "The price of this shampoo is 500 yen."
[1925] Step 11:
[1926] The user receives the price information sent via the smart glasses.
[1927] Price comparison feature
[1928] Specific steps to compare in-store and online prices
[1929] Step 1:
[1930] Users scan the barcode of the product they pick up with the smart glasses.
[1931] Step 2:
[1932] The terminal scans the barcode and obtains the product's price information.
[1933] Step 3:
[1934] The terminal transmits the price information and the product identifier to the server.
[1935] Step 4:
[1936] The server receives the price information and the product identifier and queries a price database of multiple online stores.
[1937] Step 5:
[1938] The server retrieves the prices from each online store and identifies the lowest price.
[1939] Step 6:
[1940] The server sends the user's voice data and facial expression data to the emotion engine for emotion analysis.
[1941] Step 7:
[1942] The server receives the emotion recognition result from the emotion engine and adjusts the content and tone of the notification of the cheapest price information based on the emotion recognition result.
[1943] Step 8:
[1944] The server transmits the adjusted lowest price information to the terminal.
[1945] Step 9:
[1946] The terminal converts the received information about the cheapest price into text for voice output.
[1947] Step 10:
[1948] The device converts the text to speech and announces through the speaker, "It's on sale online for 450 yen."
[1949] Step 11:
[1950] The user receives price information from the smart glasses and decides to purchase.
[1951] Through the above processing steps, this system supports users in shopping abroad efficiently and comfortably. By utilizing the emotion engine, guidance and information are provided according to the user's emotional state, providing a more personalized experience.
[1952] Example 2
[1953] Next, a description will be given of Example 2. In the following description, the data processing device 12 will be referred to as a "server" and the headset type terminal 314 will be referred to as a "terminal."
[1954] The current shopping experience is burdensome for users, as it is difficult to locate products in physical stores. Furthermore, price conversion and comparison are time-consuming, making shopping especially complicated in foreign countries. Furthermore, conventional voice guidance systems do not take user emotions into account, so guidance information is not optimized according to the user's state. To solve these problems, a system that can recognize user emotions and provide more personalized guidance is needed.
[1955] The identification process by the identification processing unit 290 of the data processing device 12 in the second embodiment is realized by the following means. In this invention, the server includes a means for converting acquired voice input into text using voice recognition technology, a means for analyzing the user's request from the converted text, and a means for acquiring product location information from a database based on the analysis result, adjusting the information based on an emotion recognition engine, and providing voice guidance. This makes it possible to provide optimal guidance according to the user's emotions.
[1956] "User" refers to the person who uses this system and is the subject who performs voice input.
[1957] "Voice input" refers to voice data provided by a user through a voice input device such as a microphone.
[1958] "Speech recognition technology" refers to technology for converting voice data into text data, and refers to technology that uses, for example, voice signal processing algorithms or machine learning models.
[1959] "Text" is character string data converted using speech recognition technology.
[1960] "Request" refers to various questions and instructions that a user gives to the system.
[1961] "Analysis" refers to the process of understanding user requests and responding appropriately to them.
[1962] "Location information" is data that indicates the physical location of a particular product or location.
[1963] A "database" is an information management system for storing and managing location information and related data.
[1964] An "emotion recognition engine" is software or hardware that recognizes a user's emotions from their voice or facial expressions.
[1965] "Voice guidance" refers to guidance information provided to users via voice.
[1966] "Real-time" refers to near-instant processing or response.
[1967] "Current location" is data indicating the physical location where the user is currently located.
[1968] "Route information" is data that indicates the route and instructions from the current position to the destination.
[1969] An "external API" is an application program interface provided by a third party that is a means of connecting with other systems.
[1970] "Exchange rate" is data that indicates the exchange rate between one country's currency and another country's currency.
[1971] "Home currency" refers to the currency of the country or region to which the user belongs.
[1972] This invention is a voice guidance system using smart glasses that combines an emotion engine that recognizes the user's emotions to provide a more personalized shopping experience.
[1973] Voice guidance function
[1974] Device (smart glasses)
[1975] The user speaks to the smart glasses, saying, "Where is the shampoo shelf?" The smart glasses receive voice input via a built-in microphone and convert the voice data into a digital format. The converted digital voice data is then transmitted to the server via wireless communication (WiFi or Bluetooth).
[1976] server
[1977] The server converts the received digital voice data into text using voice recognition technology (e.g., technology using a voice signal processing algorithm or a machine learning model), then analyzes the user's request from the text (e.g., using natural language processing technology) and retrieves the location information of the corresponding product from a database (e.g., an information management system).
[1978] The acquired location information is adjusted by an emotion recognition engine (e.g., software or hardware for recognizing emotions from voice and facial expressions) to take into account the user's emotions. Finally, an adjusted voice guidance text is generated and sent to the device.
[1979] Device (smart glasses)
[1980] The voice guidance text sent from the server is received and the voice output device (speaker or earphone) is used to guide the user, for example, "Go 10 meters to the right and turn left at the next shelf."
[1981] Real-time location information function
[1982] Device (smart glasses)
[1983] The user's current location is determined using built-in sensors (e.g., GPS or WiFi triangulation), and the determined current location data is transmitted to a server.
[1984] server
[1985] The server identifies the user's location based on the received current location data. Next, it retrieves the desired product's location information from the database and generates route information from the user's current location to the destination. This route information is also adjusted by an emotion recognition engine, taking into account the user's emotions. Finally, the adjusted route information is sent to the device.
[1986] Device (smart glasses)
[1987] The route information sent from the server is converted into text for voice guidance, and guidance to the user is started.
[1988] Currency conversion function
[1989] Device (smart glasses)
[1990] When a user picks up a product and wants to check the price, the smart glasses scan the product barcode to obtain price information, which is then sent to the server along with the user's home currency.
[1991] server
[1992] The server obtains the latest exchange rate from an external API (e.g., an application program interface for obtaining exchange rates) and uses it to convert the product price into the user's home currency. The converted price information is adjusted by an emotion recognition engine and sent to the device.
[1993] Device (smart glasses)
[1994] The converted price information is converted into text for voice output to the user, and the user is notified through the speaker that "The price of this shampoo is 500 yen."
[1995] Price comparison feature
[1996] Device (smart glasses)
[1997] When a user wants to compare prices at different stores, they scan the product barcode to get the price information, which is then sent to the server along with the product identifier.
[1998] server
[1999] The server queries price databases of multiple online stores, retrieves prices from each online store, and identifies the cheapest price. This cheapest price information is also adjusted by the emotion recognition engine. Finally, the adjusted cheapest price information is sent to the device.
[2000] Device (smart glasses)
[2001] The lowest price information is converted into text for voice output and announced through the speaker: "It's on sale online for 450 yen."
[2002] Using the Emotion Recognition Engine
[2003] server
[2004] The server receives the voice data and facial expression data and uses an emotion recognition engine to recognize the user's emotions. Based on the recognized emotions, the voice guidance and price information notifications are adjusted. Finally, the adjusted information is sent to the device.
[2005] Prompt Sentence Examples
[2006] If a user is in a supermarket in France, they can ask their smart glasses to:
[2007] "Where's the shampoo shelf?"
[2008] "How much is this shampoo?"
[2009] "How much does this shampoo cost online?"
[2010] Entering these prompts into the system will provide corresponding guidance and pricing information.
[2011] As described above, the system combines multiple functions such as voice input, location information, currency conversion, price comparison, and emotion recognition to provide people with an advanced and personalized shopping experience.
[2012] The flow of the identification process in the second embodiment will be described with reference to FIG.
[2013] Voice guidance function
[2014] Step 1:
[2015] The device (smart glasses) receives the user's voice input. When a user speaks to the smart glasses, saying, "Where is the shampoo shelf?", this voice data is captured by the device's built-in microphone.
[2016] Input: User's voice data.
[2017] Output: Audio data.
[2018] Step 2:
[2019] The audio data acquired by the device (smart glasses) is converted into a digital format. Specifically, AD conversion is performed to convert the audio data into a digital signal.
[2020] Input: Analog audio data.
[2021] Output: Digital audio data.
[2022] Step 3:
[2023] The device (smart glasses) sends digital audio data to the server. The audio data is transferred to the server via wireless communication (WiFi or Bluetooth).
[2024] Input: Digital audio data.
[2025] Output: Audio data transfer to the server.
[2026] Step 4:
[2027] The server receives the digital audio data. The audio data is received at the server.
[2028] Input: Digital audio data.
[2029] Output: Audio data in the server.
[2030] Step 5:
[2031] The server converts the voice data into text using speech recognition technology. During this process, a speech signal processing algorithm extracts the characteristics of the voice waveform and converts it into text.
[2032] Input: Audio data.
[2033] Output: Text data (e.g., "Where is the shampoo shelf?").
[2034] Step 6:
[2035] The server analyzes the user's request from the converted text and uses natural language processing techniques to extract the intent (e.g., where to find shampoo on the shelf) from the text.
[2036] Input: Text data.
[2037] Output: The intent of the text (e.g., I want to know where shampoo is on the shelf).
[2038] Step 7:
[2039] The server retrieves product location information from the database. Using an SQL query, the server retrieves shampoo shelf location information from the database.
[2040] Input: User request information.
[2041] Output: Product location.
[2042] Step 8:
[2043] The server uses an emotion recognition engine to recognize emotions from the user's voice. In this process, the server analyzes the voice features and estimates the user's emotions.
[2044] Input: Audio data.
[2045] Output: Perceived emotion (e.g., confused).
[2046] Step 9:
[2047] The server adjusts the guidance content based on the recognized emotion. For example, if the user is confused, the guidance content will be revised to be more polite.
[2048] Input: Product location and perceived sentiment.
[2049] Output: Adjusted voice guidance text (e.g. "Slowly walk 10 meters to the right, then turn left at the next ledge.").
[2050] Step 10:
[2051] The server transmits the adjusted voice guidance text to the terminal.
[2052] Input: Adjusted voice prompt text.
[2053] Output: Transfer of guidance text to the terminal.
[2054] Step 11:
[2055] The device (smart glasses) receives the voice guidance text sent from the server.
[2056] Input: Adjusted voice prompt text.
[2057] Output: Voice guidance text in the device.
[2058] Step 12:
[2059] The device (smart glasses) uses a voice output device to provide guidance to the user. For example, it may use a speaker to say, "Go 10 meters to the right and then turn left at the next shelf."
[2060] Input: Voice prompt text.
[2061] Output: Audio instructions to the user.
[2062] Real-time location information function
[2063] Step 1:
[2064] The device (smart glasses) determines the user's current location using built-in sensors (e.g., GPS or WiFi triangulation).
[2065] Input: Position sensor data.
[2066] Output: Current location data.
[2067] Step 2:
[2068] The device (smart glasses) sends the determined current location data to the server via Wi-Fi or Bluetooth.
[2069] Input: Current location data.
[2070] Output: Send current location data to the server.
[2071] Step 3:
[2072] The server receives the current location data and determines the user's current location.
[2073] Input: Current location data.
[2074] Output: The current location determined.
[2075] Step 4:
[2076] The server retrieves the location information of the target product from the database and generates route information from the current location to the destination using a routing algorithm such as Dijkstra's algorithm.
[2077] Input: Current location and product location information.
[2078] Output: Route information.
[2079] Step 5:
[2080] The server recognizes the user's emotions using an emotion recognition engine and adjusts the route information.
[2081] Input: Voice data and route information.
[2082] Output: Adjusted route information.
[2083] Step 6:
[2084] The server sends the adjusted route information to the terminal.
[2085] Input: Adjusted route information.
[2086] Output: Sending route information to the terminal.
[2087] Step 7:
[2088] The terminal (smart glasses) converts the route information sent from the server into text for voice guidance and begins providing guidance to the user.
[2089] Input: Adjusted route information.
[2090] Output: Audio instructions to the user.
[2091] Currency conversion function
[2092] Step 1:
[2093] The device (smart glasses) scans the product barcode and obtains price information. A barcode reader is used to read the barcode and obtain price information.
[2094] Input: Barcode data.
[2095] Output: Price information.
[2096] Step 2:
[2097] Price information and the user's home currency information are sent to the server.
[2098] Input: Price information and home currency information.
[2099] Output: Sending data to the server.
[2100] Step 3:
[2101] The server receives the price information and retrieves the latest exchange rates from an external API, for example, using the API of a currency exchange rate provider.
[2102] Input: Price information.
[2103] Output: Exchange rate from external API.
[2104] Step 4:
[2105] The obtained exchange rate is used to convert the product price into the user's home currency.
[2106] Input: Exchange rate and price information.
[2107] Output: The converted price information.
[2108] Step 5:
[2109] The server recognizes the user's emotions using an emotion recognition engine and adjusts the price information.
[2110] Input: Audio data and converted price information.
[2111] Output: Adjusted price information.
[2112] Step 6:
[2113] The server sends the adjusted price information to the terminal.
[2114] Input: Adjusted price information.
[2115] Output: Sending price information to the terminal.
[2116] Step 7:
[2117] The device (smart glasses) converts the price information into text for voice output and notifies the user of the price.
[2118] Input: Adjusted price information.
[2119] Output: Audio notification to the user.
[2120] Price comparison feature
[2121] Step 1:
[2122] The terminal (smart glasses) scans the product barcode and obtains price information.
[2123] Input: Barcode data.
[2124] Output: Price information.
[2125] Step 2:
[2126] Send the price information and product identifier to the server.
[2127] Inputs: Price information and product identifiers.
[2128] Output: Sending data to the server.
[2129] Step 3:
[2130] A server receives the price information and the product identifier and queries a price database of multiple online stores.
[2131] Inputs: Price information and product identifiers.
[2132] Output: Price data for each online store.
[2133] Step 4:
[2134] Get prices from each online store and find the cheapest price.
[2135] Input: Online store price data.
[2136] Output: Cheapest price information.
[2137] Step 5:
[2138] The server recognizes the user's emotions using an emotion recognition engine and adjusts the cheapest information.
[2139] Input: Voice data and lowest price information.
[2140] Output: Adjusted lowest price information.
[2141] Step 6:
[2142] The server transmits the adjusted lowest price information to the terminal.
[2143] Input: Adjusted lowest price information.
[2144] Output: Sending price information to the terminal.
[2145] Step 7:
[2146] The device (smart glasses) converts the cheapest price information into text for voice output and notifies the user.
[2147] Input: Adjusted lowest price information.
[2148] Output: Audio notification to the user.
[2149] By following these steps, the system performs a series of processes, from voice input to data acquisition, location identification, currency conversion, price comparison, and emotion recognition, to provide users with an advanced shopping experience.
[2150] (Application example 2)
[2151] Next, a description will be given of Application Example 2. In the following description, the data processing device 12 will be referred to as a "server" and the headset type terminal 314 will be referred to as a "terminal."
[2152] Conventional voice guidance systems using smart glasses could provide guidance based on user requests, but they could not adjust the guidance content according to the user's emotions. As a result, uniform guidance and information was provided without taking the user's emotional state into consideration, resulting in a suboptimal user experience. Particularly when shopping in a store, there was an issue of not being able to provide appropriate guidance when the user was confused or in a hurry.
[2153] The specific processing by the specific processing unit 290 of the data processing device 12 in the application example 2 is realized by the following means.
[2154] In this invention, the server includes emotion recognition means for recognizing emotions from the user's voice and facial expressions, means for converting voice input into text using voice recognition technology, and means for retrieving product location information from a database based on the analysis results, thereby enabling personalized voice guidance and information provision that reflects the user's emotional state.
[2155] "User" means an individual who operates a system or device.
[2156] "Voice input" refers to commands, questions, etc. that a user communicates to a system by speaking.
[2157] "Voice recognition technology" is a technology that analyzes voice signals and converts them into text data.
[2158] "Text conversion" refers to the process of converting captured voice input into text data.
[2159] "Request analysis" is the process of understanding the intent and content of a user's request from the converted text of their voice.
[2160] "Location information" refers to information that indicates where a particular product or section is located within a store.
[2161] A "database" is an information system that structures and stores a collection of related data so that it can be easily searched and retrieved.
[2162] "Voice guidance" refers to a method of providing information to users using a voice source.
[2163] "Emotion recognition means" refers to technology or equipment for detecting and classifying a user's emotional state from data such as the user's voice and facial expressions.
[2164] "Route information" refers to information relating to directions and directions from a specific starting point to a destination.
[2165] An "external API" is an interface for connecting with external software applications and services.
[2166] An "exchange rate" is a ratio that represents the exchange value between different currencies.
[2167] "Price conversion" refers to the process of converting product prices into other currency units.
[2168] "Facial expressions" refer to the movements and appearances that appear on the surface of the face and indicate emotions and intentions.
[2169] "Adjusting guidance content" means appropriately changing the information and direction guidance provided depending on the user's situation and emotions.
[2170] This invention relates to a system that allows users to use smart glasses in a store to receive voice guidance and obtain product information. Specifically, this system provides a more personalized shopping experience by combining voice input, voice recognition, emotion recognition, and location-based services.
[2171] Voice guidance function
[2172] The server uses the microphone built into the smart glasses to capture the user's voice input. The voice input is converted into text using voice recognition technology, and the user's request is analyzed from that text. Based on the results of this analysis, the server retrieves the product's location information from a database and provides the retrieved location information as voice guidance. An emotion recognition means is used to recognize the user's emotions from their voice and facial expressions, and the guidance content is adjusted based on the recognition results. This makes it possible to provide more detailed guidance if the user is confused, and faster guidance if the user is in a hurry.
[2173] Real-time location information function
[2174] The server uses sensors (such as GPS and WiFi triangulation) built into the smart glasses to identify the user's current location in real time. Based on the identified current location, the server generates route information to the product section and provides that route information as voice guidance. By adjusting the route guidance based on the user's location information, the user can move smoothly through the store. During route guidance, emotion recognition means is used to analyze the user's emotions and adjust the guidance content appropriately, providing a more personalized service.
[2175] Currency conversion function
[2176] The server retrieves the latest exchange rates from an external API, converts the product price into the user's home currency using the obtained exchange rate, and provides the converted price information via voice. During the price notification, the system recognizes the user's emotions and adjusts the guidance, such as providing a detailed explanation, if it detects doubt or confusion.
[2177] Specific examples
[2178] For example, if a user speaks to a smart glass in a store and asks, "Where is the shampoo shelf?", the smart glass will recognize the voice and send it to the server. The server will convert the voice into text and identify the location of the shampoo shelf. The emotion recognition means will analyze the user's emotions, and if the user appears confused, the smart glass will provide polite guidance such as, "Go 10 meters to the right and then turn left at the next shelf." On the other hand, if the system detects that the user is in a hurry, it will provide simple guidance such as, "It's 10 meters to the right."
[2179] If you pick up a product and ask, "How much is this product?", the device will retrieve price information and announce in voice, "The price of this shampoo is 5 euros (approximately 620 yen)." If you ask, "How much does this shampoo cost online?", the device will also provide the latest information on online prices, telling you, "It's being sold online for 4.5 euros." At the same time, it can analyze the user's emotions and adjust the guidance accordingly.
[2180] Prompt Sentence Examples
[2181] "Where's the shampoo shelf?"
[2182] "How much is this item?"
[2183] "How much does this shampoo cost online?"
[2184] The flow of the specific processing in the application example 2 will be described with reference to FIG.
[2185] Step 1:
[2186] The user provides voice input to the smart glasses.
[2187] Input: User speech (e.g., "Where is the shampoo shelf?")
[2188] How it works: The user's voice is captured using the microphone built into the smart glasses.
[2189] Output: Analog audio data
[2190] Step 2:
[2191] The device converts the voice data into a digital format.
[2192] Input: Analog audio data
[2193] How it works: The audio processing module inside the smart glasses converts analog audio data into digital audio data.
[2194] Output: Digital audio data
[2195] Step 3:
[2196] The terminal transmits digital audio data to the server.
[2197] Input: Digital audio data
[2198] How it works: Smart glasses send digital audio data to a server.
[2199] Output: Digital audio data sent to the server
[2200] Step 4:
[2201] The server converts the digital voice data into text.
[2202] Input: Digital audio data
[2203] How it works: The server's speech recognition technology analyzes the digital voice data and converts it into corresponding text.
[2204] Output: The transformed text (e.g., "Where is the shampoo shelf?")
[2205] Step 5:
[2206] The server parses the user's request from the text.
[2207] Input: Translated text
[2208] How it works: The server parses the text and determines what the user is requesting (e.g., I want to know where the shampoo shelf is).
[2209] Output: Analysis results (e.g., obtain product location information)
[2210] Step 6:
[2211] The server retrieves the product location information from the database.
[2212] Input: Analysis results (request for product location information)
[2213] Operation: The server queries the database and obtains the location information of the relevant product.
[2214] Output: Product location information (e.g., go 10 meters to the right and turn left at the next shelf)
[2215] Step 7:
[2216] The server recognizes the user's emotions using an emotion recognition means.
[2217] Input: User's voice and facial expression data
[2218] How it works: The server uses emotion recognition technology to analyze the user's emotions from their voice and facial expressions.
[2219] Output: User's emotional state (e.g. confused, in a hurry)
[2220] Step 8:
[2221] The server adjusts the guidance content based on the emotion.
[2222] Input: Product location, user emotional state
[2223] How it works: The server takes into account the user's emotional state and adjusts the guidance content (e.g., polite guidance if the user is confused, brief guidance if the user is in a hurry).
[2224] Output: Adjusted guidance text (e.g. "Go right 10 meters and turn left at the next ledge")
[2225] Step 9:
[2226] The server sends the adjusted guidance text to the smart glasses.
[2227] Input: Adjusted guidance text
[2228] Operation: The server sends the adjusted guidance text to the smart glasses.
[2229] Output: Guidance text sent to smart glasses
[2230] Step 10:
[2231] The terminal provides the user with the guidance text by voice.
[2232] Input: Guidance text
[2233] How it works: The audio output device of the smart glasses plays the guidance text as audio.
[2234] Output: Guidance speech (e.g. "Go 10 meters to the right and turn left at the next shelf")
[2235] The specific processing unit 290 transmits the result of the specific processing to the headset type terminal 314. In the headset type terminal 314, the control unit 46A causes the speaker 240 and the display 343 to output the result of the specific processing. The microphone 238 acquires audio indicating a user input regarding the result of the specific processing. The control unit 46A transmits audio data indicating the user input acquired by the microphone 238 to the data processing device 12. In the data processing device 12, the specific processing unit 290 acquires the audio data.
[2236] The data generation model 58 is a so-called generative AI (Artificial Intelligence). An example of the data generation model 58 is ChatGPT (Internet Search<URL: https: / / openai.com / blog / chatgpt> ), Gemini (Internet search <url: https: gemini.google.com ?hl="ja">) and other generation AIs. The data generation model 58 is obtained by performing deep learning on a neural network. A prompt including an instruction is input to the data generation model 58, and inference data such as voice data indicating voice, text data indicating text, and image data indicating an image is also input. The data generation model 58 performs inference on the input inference data in accordance with the instruction indicated by the prompt, and outputs the inference result in a data format such as voice data and text data. Here, inference refers to, for example, analysis, classification, prediction, and / or summarization.
[2237] In the above embodiment, an example was given in which the specific processing is performed by the data processing device 12, but the technology of the present disclosure is not limited to this, and the specific processing may be performed by the headset type terminal 314.
[2238] [Fourth embodiment]
[2239] FIG. 7 shows an example of the configuration of a data processing system 410 according to the fourth embodiment.
[2240] 7, a data processing system 410 includes a data processing device 12 and a robot 414. An example of the data processing device 12 is a server.
[2241] The data processing device 12 includes a computer 22, a database 24, and a communication I / F 26. The computer 22 is an example of a "computer" according to the technology of the present disclosure. The computer 22 includes a processor 28, a RAM 30, and a storage 32. The processor 28, the RAM 30, and the storage 32 are connected to a bus 34. The database 24 and the communication I / F 26 are also connected to the bus 34. The communication I / F 26 is connected to a network 54. Examples of the network 54 include a WAN (Wide Area Network) and / or a LAN (Local Area Network).
[2242] The robot 414 includes a computer 36, a microphone 238, a speaker 240, a camera 42, a communication I / F 44, and a control target 443. The computer 36 includes a processor 46, a RAM 48, and a storage 50. The processor 46, the RAM 48, and the storage 50 are connected to a bus 52. The microphone 238, the speaker 240, the camera 42, and the control target 443 are also connected to the bus 52.
[2243] The microphone 238 receives instructions and the like from the user 20 by receiving voice uttered by the user 20. The microphone 238 captures the voice uttered by the user 20, converts the captured voice into audio data, and outputs it to the processor 46. The speaker 240 outputs audio in accordance with instructions from the processor 46.
[2244] Camera 42 is a small digital camera equipped with an optical system including a lens, aperture, and shutter, and an imaging element such as a CMOS (Complementary Metal-Oxide-Semiconductor) image sensor or a CCD (Charge Coupled Device) image sensor, and captures images of the surroundings of user 20 (for example, an imaging range defined by an angle of view equivalent to the field of vision of a typical healthy person).
[2245] The communication I / F 44 is connected to a network 54. The communication I / Fs 44 and 26 are responsible for the exchange of various information between the processor 46 and the processor 28 via the network 54. The exchange of various information between the processor 46 and the processor 28 using the communication I / Fs 44 and 26 is carried out in a secure state.
[2246] The control object 443 includes a display device, LEDs in the eyes, and motors for driving the arms, hands, and feet. The posture and gestures of the robot 414 are controlled by controlling the motors of the arms, hands, and feet. Some of the emotions of the robot 414 can be expressed by controlling these motors. In addition, the facial expressions of the robot 414 can also be expressed by controlling the light emission state of the LEDs in the eyes of the robot 414.
[2247] Fig. 8 shows an example of the main functions of the data processing device 12 and the robot 414. As shown in Fig. 8, in the data processing device 12, a specific process is performed by the processor 28. A specific process program 56 is stored in the storage 32.
[2248] The specific processing program 56 is an example of a "program" according to the technology of the present disclosure. The processor 28 reads the specific processing program 56 from the storage 32 and executes the read specific processing program 56 on the RAM 30. The specific processing is realized by the processor 28 operating as a specific processing unit 290 in accordance with the specific processing program 56 executed on the RAM 30.
[2249] The storage 32 stores a data generation model 58 and an emotion identification model 59. The data generation model 58 and the emotion identification model 59 are used by the identification processing unit 290.
[2250] In the robot 414, the processor 46 performs the reception output process. A reception output program 60 is stored in the storage 50. The processor 46 reads the reception output program 60 from the storage 50 and executes the read reception output program 60 on the RAM 48. The reception output process is realized by the processor 46 operating as the control unit 46A in accordance with the reception output program 60 executed on the RAM 48.
[2251] Next, a description will be given of the specific processing performed by the specific processing unit 290 of the data processing device 12. In the following description, the data processing device 12 will be referred to as a "server" and the robot 414 will be referred to as a "terminal."
[2252] The present invention is a voice guidance system using smart glasses that solves problems such as language barriers when shopping abroad, product location, currency conversion, price comparison, etc. Hereinafter, an embodiment of the present invention will be described in detail.
[2253] Voice guidance function
[2254] Device (smart glasses)
[2255] 1. The user asks the smart glasses, "Where is the shampoo shelf?"
[2256] 2. The smart glasses capture voice input via a built-in microphone and convert the voice data into a digital format.
[2257] 3. Send the digital audio data to the server.
[2258] server
[2259] 1. Receives digital voice data and converts it into text using speech recognition technology.
[2260] 2. Parse the user's request (in this case, the shelf location of shampoo) from the converted text.
[2261] 3. Obtain the location information of the shampoo shelves from the database and generate text for voice guidance.
[2262] 4. The generated voice guidance text is sent to the terminal.
[2263] Terminal
[2264] 1. Receive the voice guidance text sent from the server and use the voice output device to instruct the user, "Go 10 meters to the right and turn left at the next shelf."
[2265] Real-time location information function
[2266] Terminal
[2267] 1. Determine the user's current location using built-in sensors (GPS, WiFi triangulation, etc.).
[2268] 2. The determined current location data is sent to the server.
[2269] server
[2270] 1. Receive current location data and determine the user's location based on that data.
[2271] 2. Obtain the location information of the desired product from the database and generate route information from the current location to the destination.
[2272] 3. The generated route information is sent to the terminal.
[2273] Terminal
[2274] 1. Route information sent from the server is provided to the user via voice.
[2275] Currency conversion function
[2276] Terminal
[2277] 1. Scan the product barcode to get price information.
[2278] 2. Send price information and user's home currency information to the server.
[2279] server
[2280] 1. Receive price information and get the latest exchange rates using an external API.
[2281] 2. Convert the product price into the user's home currency using the obtained exchange rate.
[2282] 3. The converted price information is sent to the terminal.
[2283] Terminal
[2284] 1. The converted price information is notified to the user via voice, such as "The price of this shampoo is 500 yen."
[2285] Price comparison feature
[2286] Terminal
[2287] 1. Scan the product barcode to get price information.
[2288] 2. Send the price information and product identifier to the server.
[2289] server
[2290] 1. Receives price information and product identifiers and queries price databases from multiple online stores.
[2291] 2. Get prices from each online store and identify the lowest price.
[2292] 3. Send the cheapest price information to the terminal.
[2293] Terminal
[2294] 1. The lowest price information is announced to the user via voice, saying, "It's on sale online for 450 yen."
[2295] Specific examples
[2296] If a user is in a supermarket in France and asks the smart glasses, "Where is the shampoo shelf?", the smart glasses will respond with, "Go 10 meters to the right and turn left at the next shelf." After reaching the shampoo section, the user picks up an item and wants to check the price. They can ask, "How much is this shampoo?" and the smart glasses will respond, "This shampoo costs 5 euros (approximately 620 yen)." Furthermore, to compare the price with online prices, they can ask, "How much does this shampoo cost online?" and the smart glasses will respond, "It's selling for 4.5 euros online." Through this process, users can enjoy an efficient and comfortable shopping experience.
[2297] The processing flow will be explained below.
[2298] Voice guidance function
[2299] Specific processing steps from voice input to guidance
[2300] Step 1:
[2301] The user asks the smart glasses, "Where is the shampoo shelf?"
[2302] Step 2:
[2303] The terminal uses a built-in microphone to capture the user's voice and converts the voice into digital data.
[2304] Step 3:
[2305] The terminal transmits the converted digital audio data to the server.
[2306] Step 4:
[2307] The server receives the digital voice data and converts the voice data into text using voice recognition technology.
[2308] Step 5:
[2309] The server parses the text and understands the user's request (in this case, a request to know the shelf location of shampoo).
[2310] Step 6:
[2311] The server retrieves the location information of the shampoo shelves from the database.
[2312] Step 7:
[2313] Based on the acquired location information, the server generates text for voice guidance such as, "Go 10 meters to the right and turn left at the next shelf."
[2314] Step 8:
[2315] The server transmits the generated voice guidance text to the terminal.
[2316] Step 9:
[2317] The device converts the received voice guidance text into speech and begins providing guidance through the speaker.
[2318] Step 10:
[2319] The user moves as instructed.
[2320] Real-time location information function
[2321] A processing step to guide you from your current location to the product location
[2322] Step 1:
[2323] The user speaks to the smart glasses, saying, "Tell me where I am."
[2324] Step 2:
[2325] The device determines the user's current location using built-in sensors (e.g., GPS and WiFi triangulation).
[2326] Step 3:
[2327] The terminal transmits the identified current location data to the server.
[2328] Step 4:
[2329] The server receives the location data and confirms the user's current location.
[2330] Step 5:
[2331] The server retrieves route information to the product section from the database.
[2332] Step 6:
[2333] The server generates route information from the current location to the destination product section.
[2334] Step 7:
[2335] The server transmits the generated route information to the terminal.
[2336] Step 8:
[2337] The terminal converts the received route information into text for voice guidance.
[2338] Step 9:
[2339] The device converts the voice guidance text into speech and begins providing guidance through the speaker.
[2340] Step 10:
[2341] The user moves along the guided route.
[2342] Currency conversion function
[2343] Processing steps to convert product prices into your home currency
[2344] Step 1:
[2345] Users scan the barcode of the product they pick up with the smart glasses.
[2346] Step 2:
[2347] The terminal scans the barcode and obtains the product's price information.
[2348] Step 3:
[2349] The terminal transmits product price information and the user's home currency information to the server.
[2350] Step 4:
[2351] The server receives the price information and uses an external API to get the latest exchange rates.
[2352] Step 5:
[2353] The server uses the obtained exchange rate to convert the product price into the user's home currency.
[2354] Step 6:
[2355] The server transmits the converted price information to the terminal.
[2356] Step 7:
[2357] The terminal converts the received converted price information into text for voice output.
[2358] Step 8:
[2359] The device converts the text to speech and announces the price through the speaker.
[2360] Step 9:
[2361] The user receives the price information sent via the smart glasses.
[2362] Price comparison feature
[2363] Processing steps to compare in-store and online prices
[2364] Step 1:
[2365] Users scan the barcode of the product they pick up with the smart glasses.
[2366] Step 2:
[2367] The terminal scans the barcode and obtains the product's price information.
[2368] Step 3:
[2369] The terminal transmits the price information and the product identifier to the server.
[2370] Step 4:
[2371] The server receives the price information and the product identifier and queries the online store's price database.
[2372] Step 5:
[2373] The server retrieves the prices from each online store and identifies the lowest price.
[2374] Step 6:
[2375] The server transmits the cheapest price information to the terminal.
[2376] Step 7:
[2377] The terminal converts the received information about the cheapest price into text for voice output.
[2378] Step 8:
[2379] The device converts the text to speech and announces price information through the speaker.
[2380] Step 9:
[2381] The user receives price information from the smart glasses and decides to purchase.
[2382] Through the above processing steps, this system supports users in shopping abroad efficiently and comfortably.
[2383] Example 1
[2384] Next, a description will be given of Example 1. In the following description, the data processing device 12 will be referred to as a "server" and the robot 414 will be referred to as a "terminal."
[2385] In today's global society, shopping in a multilingual environment can be extremely difficult for users. When shopping abroad, language barriers make it particularly difficult to accurately determine product locations. It's also difficult to convert prices displayed in local currencies into one's own currency, or to quickly compare prices across different stores. A system that solves these problems is needed.
[2386] The specific processing by the specific processing unit 290 of the data processing device 12 in the first embodiment is realized by the following means.
[2387] In this invention, the server includes: means for acquiring a user's voice input; means for converting the acquired voice input into digital format; means for transmitting the converted digital voice input to the server; means for converting the voice input into text on the server using voice recognition technology; means for analyzing the user's request from the converted text; means for acquiring product location information from a database based on the analysis results; means for converting the acquired location information into text for voice guidance and transmitting it to the terminal; and means for converting the transmitted text for voice guidance into speech and providing guidance to the user. This allows users to easily find products even in different language environments. Furthermore, by providing functions such as real-time location information identification and route guidance, currency conversion using the latest exchange rates, and price comparison across multiple stores, the server achieves an efficient and convenient shopping experience.
[2388] The "means for acquiring user's voice input" is a combination of hardware and software that recognizes the voice uttered by the user and that the terminal acquires as digital data.
[2389] The "means for converting to digital form" refers to an audio signal processing device and conversion algorithm for converting analog audio signals into digital data.
[2390] The "means for transmitting to a server" refers to a communication module and protocol that allows the terminal to transfer digital data over the Internet to a remote server.
[2391] "Means for converting speech to text using speech recognition technology" refers to software and algorithms that analyze speech data on a server and convert it into corresponding text.
[2392] The "means for analyzing user requests" is a natural language processing system that extracts and analyzes the user's intentions and requests from the text obtained by speech recognition.
[2393] The "means for acquiring product location information from a database" is a query processing system for searching and acquiring product placement information from a database based on the analyzed user request.
[2394] The "means for converting the information into text for voice guidance and sending it to the terminal" is a server-side function for converting the acquired information into text in a format that is easy for the user to understand and sending this text to the terminal.
[2395] The "means for converting the text information into audio and providing the user with guidance" refers to the software and hardware within the terminal for playing back the received text information as audio.
[2396] "Means of real-time identification" refers to a system that instantly identifies a user's current location using sensor technology such as GPS and Wi-Fi triangulation.
[2397] The "means for generating route information" refers to algorithms and software for calculating a route from the specified current position to the destination and generating that information.
[2398] The "means for obtaining the latest exchange rates from an external data source" is an interface for obtaining the latest currency exchange information from an external financial data provider service.
[2399] The "means for converting the product price using the exchange rate" is a calculation algorithm for converting the product price into the user's home currency using the acquired exchange rate information.
[2400] MODE FOR CARRYING OUT THE INVENTION
[2401] The present invention is a voice guidance system using smart glasses that solves problems such as language barriers when shopping abroad, product location, currency conversion, price comparison, etc. Hereinafter, an embodiment of the present invention will be described in detail.
[2402] Overall system overview
[2403] This system is composed of a combination of smart glasses (terminals), a cloud server, a database, and an external API. The smart glasses receive voice input from the user and send it to the cloud server. The cloud server uses voice recognition technology to convert the voice data into text, and then analyzes the user's request through natural language processing. Based on the analysis results, it provides information such as product location, currency conversion, and price comparison.
[2404] Hardware and software used
[2405] Device (smart glasses)
[2406] Built-in microphone: Captures the user's voice.
[2407] Communication module: Sends digital data to a cloud server (e.g., Wi-Fi, mobile data).
[2408] Voice output device: Converts guided text into voice.
[2409] Sensors: Determine current location (e.g. GPS, Wi-Fi triangulation).
[2410] server
[2411] Speech recognition software: Google Cloud Speech-to-Text API.
[2412] Natural language processing engine: Uses SpaCy or similar.
[2413] Database: Uses MySQL to manage product location and price information.
[2414] External API: Uses Open Exchange Rates API for currency conversion.
[2415] Program processing
[2416] 1. Voice guidance function
[2417] The user asks the smart glasses, "Where is the shampoo shelf?"
[2418] The smart glasses pick up sound through a built-in microphone and convert it into a digital format.
[2419] The digital audio data is transmitted to a cloud server.
[2420] The server uses speech recognition software to convert the voice data into text.
[2421] The text is analyzed using a natural language processing engine to understand the request.
[2422] Retrieve product location information from the database and generate guide text.
[2423] The guidance text is sent to the smart glasses and the user is guided by voice.
[2424] 2. Real-time location information function
[2425] The smart glasses use built-in sensors to determine the user's current location.
[2426] The current location data is sent to a cloud server.
[2427] The server generates route information to the product section based on the current location data.
[2428] The generated route information is sent to smart glasses and guidance is provided via voice.
[2429] 3. Currency conversion function
[2430] Smart glasses scan product barcodes to obtain price information.
[2431] Price information and the user's home currency information are sent to the server.
[2432] The server uses an external API to get the latest exchange rates.
[2433] Convert product prices into the user's home currency using exchange rates.
[2434] The converted price information is sent to the smart glasses and notified via voice.
[2435] 4. Price comparison feature
[2436] The smart glasses scan the product's barcode and retrieve price information.
[2437] Send the price information and product identifier to the server.
[2438] The server queries a price database of multiple online stores.
[2439] The lowest price information is identified and sent to the smart glasses.
[2440] The cheapest price information is notified to the user by voice.
[2441] Specific examples
[2442] If a user is in a supermarket in France and asks the smart glasses, "Where is the shampoo shelf?", the smart glasses will respond with, "Go 10 meters to the right and turn left at the next shelf." After reaching the shampoo section, the user picks up an item and wants to check the price. They can ask, "How much is this shampoo?" and the smart glasses will respond, "This shampoo costs 5 euros (approximately 620 yen)." Furthermore, to compare the price with online prices, they can ask, "How much does this shampoo cost online?" and the smart glasses will respond, "It's selling for 4.5 euros online." Through this process, users can enjoy an efficient and comfortable shopping experience.
[2443] Example prompts to input to the generative AI model
[2444] Please explain in detail the process by which a user queries the smart glasses for product location, product price, currency conversion, and price comparison. Please also include specific hardware, software, data processing, and calculation methods.
[2445] The flow of the identification process in the first embodiment will be described with reference to FIG.
[2446] Step 1:
[2447] A user speaks to the smart glasses, saying, "Where is the shampoo shelf?" The user's voice is picked up by the smart glasses' built-in microphone. This voice input is an analog signal, so it is converted into digital format. Specifically, A / D conversion is performed to convert the voice waveform into a digital signal. The input is the user's voice signal, and the output is digital voice data.
[2448] Step 2:
[2449] The device transmits digital audio data to a cloud server via Wi-Fi or a mobile data network using a communication module. The input is digital audio data, and the output is confirmation of data transmission to the server.
[2450] Step 3:
[2451] The server receives the digital voice data and converts it into text using speech recognition technology. The Google Cloud Speech-to-Text API is used. The server analyzes the voice data and generates corresponding text data. The input is digital voice data, and the output is the converted text data.
[2452] Step 4:
[2453] The server sends the converted text data to a natural language processing engine to analyze the user's request. For example, SpaCy can be used to extract the meaning of the text. The input is text data, and the output is the analysis result of the user's request ("where shampoo is on the shelf"). Specifically, the keyword "shampoo" is extracted from the text, and the request content is determined based on that.
[2454] Step 5:
[2455] The server retrieves product location information from the database based on the analysis results. In this example, it queries a MySQL database to retrieve the location information of the shampoo shelf. The input is the analysis result of the user request, and the output is the product location information. The database query is executed and the location information is returned.
[2456] Step 6:
[2457] The server converts the acquired location information into text for voice guidance and sends it to the terminal. An algorithm for generating voice guidance text is applied to generate text such as "Go 10 meters to the right and turn left at the next shelf." The input is the product location information, and the output is the text for voice guidance. The generated text is sent to the terminal as an HTTP response.
[2458] Step 7:
[2459] The device receives the text for voice guidance sent from the server. The received text is then used to provide guidance to the user using a voice output device. Specifically, the text is converted to speech using the Google Text-to-Speech API. The input is the text for voice guidance, and the output is the voice guidance. "Go 10 meters to the right and turn left at the next shelf," is the voice guidance that comes out of the smart glasses.
[2460] (Application example 1)
[2461] Next, a description will be given of Application Example 1. In the following description, the data processing device 12 will be referred to as a "server" and the robot 414 will be referred to as a "terminal."
[2462] The present invention aims to solve problems associated with shopping abroad, such as language barriers, product location, currency conversion, and price comparisons. Specifically, the problem is to provide a support system that enables users to intuitively and efficiently find products in physical stores, understand prices, and purchase them at the most advantageous prices.
[2463] The specific processing by the specific processing unit 290 of the data processing device 12 in the application example 1 is realized by the following means.
[2464] In this invention, the server includes means for acquiring a user's voice input, means for converting the acquired voice input into text using voice recognition technology, means for analyzing the user's request from the converted text, means for acquiring product location information from a database based on the analysis result, means for providing the acquired location information as voice guidance, means for converting product prices into the user's home currency using a currency conversion function, and means for acquiring price information from multiple online stores using a price comparison function to identify the lowest price. This allows users to easily find products even in stores in a foreign country, and also makes it easy to convert currencies and compare prices.
[2465] "User voice input" refers to voice commands issued by a user to a device.
[2466] "Speech recognition technology" refers to machine learning and natural language processing technology for converting speech into text data.
[2467] "Means for converting to text" refers to a device or software that performs a process of converting captured voice input into digital data and then parsing that digital data into text.
[2468] "Means for analyzing user requests" refers to a device or software that processes voice commands converted into text to understand the information or action the user is seeking.
[2469] "Means for obtaining product location information from a database" refers to a device or software for searching and obtaining information about the product location stored in a database.
[2470] "Means for providing audio guidance" refers to a device or software that provides audio guidance to the user based on analyzed information and acquired data.
[2471] "Currency Conversion Facility" means a device or software that performs a process to convert product prices displayed in one currency into another currency selected by the user.
[2472] "Price comparison tool" means a device or software that performs a process to obtain and compare price information from multiple sources or online stores for a particular product to identify the most favorable price.
[2473] "Means for determining current location in real time" refers to devices or software that obtain a user's current location in real time using technologies such as GPS or WiFi triangulation.
[2474] The "means for generating route information" refers to a device or software for calculating the optimal route from the user's current location to the location of the desired product and generating the route information.
[2475] "Means for obtaining the latest exchange rates from an external API" refers to a device or software for obtaining exchange rate information from an external server via the Internet.
[2476] The present invention relates to a system that uses a voice guidance system in a specified target environment to provide users with various information such as product location information, currency conversion, price comparison, etc. Specific embodiments of the system are described below.
[2477] Generating a Program
[2478] The system consists of the following main functions:
[2479] 1. Voice Input Capture and Recognition
[2480] 2. Location information acquisition and guidance
[2481] 3. Currency conversion function
[2482] 4. Price comparison feature
[2483] Program processing explanation
[2484] Voice Input Capture and Recognition
[2485] The device (e.g., smart glasses) captures the user's voice input with a microphone and converts the voice data into a digital format. The converted digital voice data is sent to the server using the os and wave libraries. The server receives the voice data via a flask-based endpoint and converts it into text using the pydub and speech_recognition libraries.
[2486] Next, a natural language processing library such as nltk is used to analyze the user's request, and the next steps are taken based on the analysis results.
[2487] Location information acquisition and guidance
[2488] The device acquires its current location in real time using its built-in GPS and WiFi triangulation, and sends it to the server. The server processes the location data using the geopy library and generates the optimal route to the product section based on the user's current location. This route information is then provided as audio guidance.
[2489] Currency conversion function
[2490] When a user scans a product's barcode, the device retrieves the price information and sends it to the server. The server then uses an external API to retrieve the latest exchange rate and converts the product price based on the obtained rate. The converted price information is retrieved using the requests library and notified to the user via voice.
[2491] Price comparison feature
[2492] When a user scans a product's barcode, the device sends the price information to a server, which queries a price database of multiple online stores to identify the lowest price, which is then provided to the user via voice.
[2493] Specific examples
[2494] If a user is in a supermarket in France and asks the smart glasses, "Where is the shampoo shelf?", the smart glasses will respond with, "Go 10 meters to the right and turn left at the next shelf." After reaching the shampoo section, the user picks up an item and wants to check the price. They can ask, "How much is this shampoo?" and the smart glasses will respond, "This shampoo costs 5 euros (approximately 620 yen)." Furthermore, to compare the price with online prices, they can ask, "How much does this shampoo cost online?" and the smart glasses will respond, "It's selling for 4.5 euros online." Through this process, users can enjoy an efficient and comfortable shopping experience.
[2495] Prompt Sentence Examples
[2496] "This smartglasses-based shopping app allows users to easily navigate in-store, convert currencies, compare prices, and more, all by voice. For example, if a user is in a supermarket in France and asks the smartglasses, 'Where is the shampoo shelf?', the smartglasses will respond with, 'Go 10 meters to the right and then turn left at the next shelf.' Please generate a program that achieves this function."
[2497] The flow of the specific processing in the application example 1 will be described with reference to FIG.
[2498] Step 1:
[2499] The device (smart glasses) captures the user's voice input with a microphone. The user speaks, "Where is the shampoo shelf?", and this voice signal is acquired. The voice input is converted into digital voice data using the OS and Wave libraries. The output is digital voice data.
[2500] Step 2:
[2501] The device sends digital audio data to the server. This is done using the requests library. The input is the digital audio data, and the output is a request to the server to send the audio data.
[2502] Step 3:
[2503] The server receives the received digital audio data at the flask endpoint and starts analyzing the data. The input is the digital audio data. The server converts it into text using the pydub and speech_recognition libraries. The output is the converted text data.
[2504] Step 4:
[2505] The server analyzes the converted text data using a natural language processing library such as nltk to identify the user's request. The converted text data is input, and the user's request extracted through analysis is output. Specifically, "the location of the shampoo shelf" is identified as the user's request.
[2506] Step 5:
[2507] The server retrieves the shampoo shelf location information from the database. The input is the user's request (the shampoo shelf location), and the output is the product location information. This location information indicates where the product is located in the store.
[2508] Step 6:
[2509] The server calculates the location using the geopy library based on the product location information and generates route information. The input is the product location information and the user's current location, and the output is route information. Specifically, the guidance will be something like "Go 10 meters to the right and turn left at the next shelf."
[2510] Step 7:
[2511] The server converts the generated route information into text for voice guidance and sends it to the terminal. The input is route information, and a process is performed to convert this information into text for voice guidance. The output is text for voice guidance.
[2512] Step 8:
[2513] The terminal notifies the user of the voice guidance text received from the server using a voice output device. The input is the voice guidance text, and the output is the voice guidance ("Go 10 meters to the right and turn left at the next shelf").
[2514] Step 9:
[2515] When a user scans a product barcode, the terminal obtains the price information and sends it to the server. The input is the price information output by scanning the product barcode.
[2516] Step 10:
[2517] The server uses an external API to obtain the latest exchange rate based on the acquired price information. The input is the price information, and the output is the latest exchange rate.
[2518] Step 11:
[2519] The server converts the price into the user's home currency based on the exchange rate. The input is the acquired exchange rate and price information, and the output is the converted price information.
[2520] Step 12:
[2521] The server converts the converted price information into text for voice guidance and sends it to the terminal. The input is the converted price information, and the output is the text for voice guidance.
[2522] Step 13:
[2523] The terminal notifies the user of the voice guidance text received from the server using a voice output device. The input is the voice guidance text, and the output is the voice guidance ("The price of this shampoo is XX yen").
[2524] Step 14:
[2525] When a user speaks "What is the price of this shampoo online?", the device sends this speech input to the server. The input is the user's speech input, and the output is digital voice data.
[2526] Step 15:
[2527] The server analyzes the received voice data to identify it as a price comparison request and queries a price database of multiple online stores. The input is the analyzed request, and the output is the price information of each online store.
[2528] Step 16:
[2529] The server identifies the cheapest price information, converts it into a voice guidance text, and sends it to the terminal. The input is the price information of the online store, and the output is a voice guidance text containing the cheapest price information.
[2530] Step 17:
[2531] The terminal notifies the user of the voice guidance text received from the server using a voice output device. The input is the voice guidance text, and the output is the voice guidance ("This item is sold online for XX yen").
[2532] Furthermore, an emotion engine that estimates the user's emotion may be further combined. That is, the identification processing unit 290 may estimate the user's emotion using the emotion identification model 59, and perform identification processing using the user's emotion.
[2533] The present invention provides a system that combines a voice guidance system using smart glasses with an emotion engine that recognizes the user's emotions, providing a more personalized shopping experience. The following describes in detail an embodiment of the present invention.
[2534] Voice guidance function
[2535] Device (smart glasses)
[2536] 1. The user asks the smart glasses, "Where is the shampoo shelf?"
[2537] 2. The smart glasses capture voice input via a built-in microphone and convert the voice data into a digital format.
[2538] 3. Send the digital audio data to the server.
[2539] server
[2540] 1. Receive digital voice data and convert the voice data into text using voice recognition technology.
[2541] 2. Understand the user's request from the converted text (in this case, the request to know the shelf location of shampoo).
[2542] 3. Obtain the location information of the shampoo shelves from the database and generate text for voice guidance.
[2543] 4. Use the emotion engine to recognize emotions from the user's voice and adjust the guidance content.
[2544] 5. The adjusted voice guidance text is sent to the device.
[2545] Terminal
[2546] 1. Receive the voice guidance text sent from the server and use the voice output device to instruct the user, "Go 10 meters to the right and turn left at the next shelf."
[2547] Real-time location information function
[2548] Terminal
[2549] 1. Determine the user's current location using built-in sensors (GPS, WiFi triangulation, etc.).
[2550] 2. The determined current location data is sent to the server.
[2551] server
[2552] 1. Receive current location data and determine the user's location based on that data.
[2553] 2. Obtain the location information of the desired product from the database and generate route information from the current location to the destination.
[2554] 3. Use an emotion engine to recognize emotions from the user's voice and facial expressions and adjust the guidance content accordingly.
[2555] 4. The adjusted route information is sent to the terminal.
[2556] Terminal
[2557] 1. The route information sent from the server is converted into text for voice guidance, and guidance to the user begins.
[2558] Currency conversion function
[2559] Terminal
[2560] 1. Scan the product barcode to get price information.
[2561] 2. Send price information and user's home currency information to the server.
[2562] server
[2563] 1. Receive price information and get the latest exchange rates using an external API.
[2564] 2. Convert the product price into the user's home currency using the obtained exchange rate.
[2565] 3. Use an emotion engine to recognize emotions from the user's voice and facial expressions and adjust the content of price information notifications.
[2566] 4. The adjusted price information is sent to the terminal.
[2567] Terminal
[2568] 1. The converted price information is converted into text for voice output and announced through the speaker: "The price of this shampoo is 500 yen."
[2569] Price comparison feature
[2570] Terminal
[2571] 1. Scan the product barcode to get price information.
[2572] 2. Send the price information and product identifier to the server.
[2573] server
[2574] 1. Receives price information and product identifiers and queries price databases from multiple online stores.
[2575] 2. Get the prices from each online store and identify the cheapest price.
[2576] 3. Use an emotion engine to recognize emotions from the user's voice and facial expressions and adjust the content of the notification about the cheapest price.
[2577] 4. The adjusted lowest price information is sent to the terminal.
[2578] Terminal
[2579] 1. Convert the cheapest price information into text for voice output and announce through the speaker, "It's on sale online for 450 yen."
[2580] Use of emotion engine
[2581] server
[2582] 1. Receives voice data and facial expression data and recognizes the user's emotions using an emotion engine.
[2583] 2. Tailor voice prompts and pricing notifications based on the recognized emotion.
[2584] 3. The adjusted information is sent to the device.
[2585] Specific examples
[2586] If a user is in a supermarket in France and asks the smart glasses, "Where is the shampoo shelf?", the smart glasses will respond with, "Go 10 meters to the right and turn left at the next shelf." After reaching the shampoo section, the user picks up an item and asks, "How much is this shampoo?" to check the price. The smart glasses will respond, "This shampoo costs 5 euros (approximately 620 yen)." Furthermore, if the user asks, "How much does this shampoo cost online?" to compare it with online prices, the smart glasses will respond, "It's selling for 4.5 euros online." This invention utilizes an emotion engine to optimize the user experience, such as providing more attentive guidance if the user shows a confused expression.
[2587] The processing flow will be explained below.
[2588] Voice guidance function
[2589] Specific processing steps from voice input to guidance
[2590] Step 1:
[2591] The user asks the smart glasses, "Where is the shampoo shelf?"
[2592] Step 2:
[2593] The terminal uses a built-in microphone to capture the user's voice and converts the voice into digital data.
[2594] Step 3:
[2595] The terminal transmits the converted digital audio data to the server.
[2596] Step 4:
[2597] The server receives the digital voice data and converts the voice data into text using voice recognition technology.
[2598] Step 5:
[2599] The server parses the text and understands the user's request (in this case, a request to know the shelf location of shampoo).
[2600] Step 6:
[2601] The server retrieves the location information of the shampoo shelves from the database.
[2602] Step 7:
[2603] Based on the acquired location information, the server generates text for voice guidance such as, "Go 10 meters to the right and turn left at the next shelf."
[2604] Step 8:
[2605] The server sends the voice data to an emotion engine to analyze the user's emotion from the voice data.
[2606] Step 9:
[2607] The server receives the emotion recognition results from the emotion engine and adjusts the content and tone of the announcement based on the user's emotion.
[2608] Step 10:
[2609] The server transmits the adjusted voice guidance text to the terminal.
[2610] Step 11:
[2611] The device converts the received voice guidance text into speech and begins providing guidance through the speaker.
[2612] Step 12:
[2613] The user moves as instructed.
[2614] Real-time location information function
[2615] Specific processing steps to guide you from your current location to the product location
[2616] Step 1:
[2617] The user speaks to the smart glasses, saying, "Tell me where I am."
[2618] Step 2:
[2619] The device determines the user's current location using built-in sensors (e.g., GPS and WiFi triangulation).
[2620] Step 3:
[2621] The terminal transmits the identified current location data to the server.
[2622] Step 4:
[2623] The server receives the location data and determines the user's location based on it.
[2624] Step 5:
[2625] The server retrieves route information to the product section from the database.
[2626] Step 6:
[2627] The server generates route information from the current location to the destination product section.
[2628] Step 7:
[2629] The server sends the user's voice data and facial expression data to the emotion engine to analyze the user's emotions.
[2630] Step 8:
[2631] The server receives the emotion recognition results from the emotion engine and adjusts the content and tone of the route guidance based on the emotion.
[2632] Step 9:
[2633] The server transmits the adjusted route information to the terminal.
[2634] Step 10:
[2635] The terminal converts the received route information into text for voice guidance.
[2636] Step 11:
[2637] The terminal converts the voice guidance text into voice and starts providing guidance to the user through the speaker.
[2638] Step 12:
[2639] The user moves along the guided route.
[2640] Currency conversion function
[2641] Specific steps to convert product prices into your home currency
[2642] Step 1:
[2643] Users scan the barcode of the product they pick up with the smart glasses.
[2644] Step 2:
[2645] The terminal scans the barcode and obtains the product's price information.
[2646] Step 3: 【264...
Claims
1. means for obtaining a user's voice input; means for converting the acquired voice input into text using voice recognition technology; means for parsing a user request from the converted text; A means for acquiring product location information from a database based on the analysis results; means for providing the acquired location information as voice guidance; A system including:
2. a means for determining the user's current location in real time; a means for generating route information to a product section based on the identified current location; means for providing the generated route information as voice guidance; The system of claim 1 , comprising:
3. A means to retrieve the latest exchange rates from an external API, A means for converting the price of the product into the user's home currency using the obtained exchange rate; a means for providing the converted price information by voice; The system of claim 1 , comprising:
4. means for querying price databases of multiple online stores; A way to get and compare prices from different online stores, A means of providing the lowest price and its information as a voice guide; The system of claim 1 , comprising:
5. A means for analyzing user movement data and identifying a current location; means for providing real-time directions to a merchandise section based on the identified current location; The system of claim 2 , comprising:
Citation Information
Patent Citations
Persona chatbot control method and system
JP2022180282A