system
The system addresses inefficiencies in store operations by converting voice and image data to automate customer service, sales, and inventory management, improving service quality and operational efficiency.
Patent Information
- Authority / Receiving Office
- JP · JP
- Patent Type
- Applications
- Current Assignee / Owner
- SOFTBANK GROUP CORP
- Filing Date
- 2024-10-02
- Publication Date
- 2026-04-14
AI Technical Summary
Conventional store operations rely heavily on manual labor for customer service, sales, operation guidance, cash register management, and inventory management, leading to inefficiencies, labor shortages, and varying service quality, which can result in decreased customer satisfaction and operational inefficiencies.
A system that converts user voice data into text, analyzes it to determine purchase intent and operation requests, generates optimal suggestions or procedures, and provides them via speech synthesis, while also using image recognition to enhance guidance and automates inventory and cash management by analyzing sales data and generating alerts and order plans.
The system streamlines store operations, improves service quality, ensures quick and accurate information provision, and addresses labor shortages by automating tasks, thereby enhancing customer satisfaction and operational efficiency.
Smart Images

Figure 2026064590000001_ABST
Abstract
Description
Technical Field
[0001] The technology of the present disclosure relates to a system.
Background Art
[0002] Patent Document 1 discloses a method for controlling a persona chatbot, which is performed by at least one processor, and includes steps of receiving a user utterance, adding the user utterance to a prompt including an instruction sentence related to an explanation of a character of the chatbot, encoding the prompt, and inputting the encoded prompt into a language model to generate a chatbot utterance as a response to the user utterance.
Prior Art Documents
Patent Documents
[0003]
Patent Document 1
Summary of the Invention
Problems to be Solved by the Invention
[0004] In conventional store operations, various operations such as customer service and sales, operation guidance, cash register money management, and inventory management rely on manual labor, so there is a problem that it is difficult to secure labor and reduce costs. Also, during a specific time period, it may not be possible to quickly respond to customers, which may lead to a decrease in customer satisfaction. Furthermore, the quality of services provided varies depending on the skill levels of the staff, which is also an issue. It is an object of the present invention to solve the above problems and achieve efficient and high-quality store operations.
Means for Solving the Problems
[0005] The present invention solves the above problems by providing a system having the following configuration: a system comprising means for converting user voice data into text data, means for analyzing the text data to determine purchase intent and operation requests, means for generating optimal suggestions or operation procedures based on the determination results, and means for providing the generated suggestions or operation procedures to the user by speech synthesis. Furthermore, the accuracy of operation guidance is improved by further including means for identifying the user's mobile terminal by image recognition, obtaining an operation manual corresponding to the model of the mobile terminal from a database, and generating operation procedures according to the user's requests from the obtained operation manual.
[0006] Furthermore, the system includes means for collecting sales data from POS terminals in real time, analyzing the sales data to calculate the amount of cash in stock, and generating an alert and notification when the calculated amount of cash in stock falls below an appropriate level, thereby improving the efficiency of cash in-store management. In addition, it includes means for registering product arrival and sales information in real time, calculating the current amount of stock, and generating an order plan when the calculated amount of stock falls below a set threshold, thereby optimizing inventory management.
[0007] "User" refers to a general customer who uses the system.
[0008] "Audio data" refers to data that records what a user says in its original audio form.
[0009] "Text data" refers to data obtained by converting audio data into a string of characters.
[0010] "Analysis" refers to the process performed to understand the user's intent and the content of their questions from text data.
[0011] "Purchase intent" refers to information that indicates what kind of product a user is planning to buy.
[0012] An "operation request" refers to a situation where a user is inquiring about a specific operation method or procedure.
[0013] "Recommendation" refers to the act of recommending the most suitable product or plan based on the user's purchase intent.
[0014] "Operating procedures" refer to the specific steps a user takes to use a particular function or service.
[0015] "Speech synthesis" refers to the technology that converts text data into speech and plays it back.
[0016] "Mobile devices" refer to portable devices such as smartphones and tablets.
[0017] "Image recognition" refers to the technology that identifies objects and characters from images captured by a camera.
[0018] A "database" refers to a system for efficiently managing and retrieving large amounts of information.
[0019] "Sales data" refers to data that records information about products and services that have been sold.
[0020] "Cash inventory" refers to the total amount of cash that is tracked in real time.
[0021] An "alert" refers to a warning that the system uses to notify the user of an abnormal or attention-grabbing situation.
[0022] "Product data" refers to data that records information about a product, such as its type, quantity, and price.
[0023] An "ordering plan" refers to a plan for ordering goods necessary to maintain an appropriate inventory level.
[0024] "Real-time" refers to the process or display of data at the very moment it is collected.
[0025] The "threshold value" refers to a value set to meet specific conditions or criteria.
Brief Description of Drawings
[0026] [Figure 1] It is a conceptual diagram showing an example of the configuration of a data processing system according to the first embodiment. [Figure 2] It is a conceptual diagram showing an example of the main functions of a data processing device and a smart device according to the first embodiment. [Figure 3] It is a conceptual diagram showing an example of the configuration of a data processing system according to the second embodiment. [Figure 4] It is a conceptual diagram showing an example of the main functions of a data processing device and smart glasses according to the second embodiment. [Figure 5] It is a conceptual diagram showing an example of the configuration of a data processing system according to the third embodiment. [Figure 6] It is a conceptual diagram showing an example of the main functions of a data processing device and a headset-type terminal according to the third embodiment. [Figure 7] It is a conceptual diagram showing an example of the configuration of a data processing system according to the fourth embodiment. [Figure 8] It is a conceptual diagram showing an example of the main functions of a data processing device and a robot according to the fourth embodiment. [Figure 9] It shows an emotion map to which multiple emotions are mapped. [Figure 10] It shows an emotion map to which multiple emotions are mapped. [Figure 11] It is a sequence diagram showing the processing flow of the data processing system in Example 1. [Figure 12] It is a sequence diagram showing the processing flow of the data processing system in Application Example 1. [Figure 13] It is a sequence diagram showing the processing flow of the data processing system in Example 2 when the emotion engine is combined. [Figure 14] It is a sequence diagram showing the processing flow of the data processing system in Application Example 2 when the emotion engine is combined. [Modes for carrying out the invention]
[0027] Hereinafter, an example of an embodiment of the system relating to the technology of this disclosure will be described with reference to the attached drawings.
[0028] First, let's explain the terminology used in the following explanation.
[0029] In the following embodiments, the signed processor (hereinafter simply referred to as "processor") may be a single arithmetic unit or a combination of multiple arithmetic units. Furthermore, the processor may be a single type of arithmetic unit or a combination of multiple types of arithmetic units. Examples of arithmetic units include CPU (Central Processing Unit), GPU (Graphics Processing Unit), GPGPU (General-Purpose computing on Graphics Processing Units), and APU (Accelerated Processing Unit).
[0030] In the following embodiments, signed RAM (Random Access Memory) is a memory that temporarily stores information and is used as work memory by the processor.
[0031] In the following embodiments, the signed storage is one or more non-volatile storage devices that store various programs and various parameters. Examples of non-volatile storage devices include flash memory (SSD (Solid State Drive)), magnetic disks (e.g., hard disks), or magnetic tapes.
[0032] In the following embodiments, the signed communication interface (I / F) is an interface that includes a communication processor and an antenna, etc. The communication interface manages communication between multiple computers. Examples of communication standards applicable to the communication interface include wireless communication standards such as 5G (5th Generation Mobile Communication System), Wi-Fi (registered trademark), or Bluetooth (registered trademark).
[0033] In the following embodiments, "A and / or B" is synonymous with "at least one of A and B." That is, "A and / or B" means that it may be A alone, or B alone, or a combination of A and B. Furthermore, in this specification, the same concept as "A and / or B" applies when expressing three or more things linked by "and / or."
[0034] [First Embodiment]
[0035] Figure 1 shows an example of the configuration of the data processing system 10 according to the first embodiment.
[0036] As shown in Figure 1, the data processing system 10 includes a data processing device 12 and a smart device 14. An example of the data processing device 12 is a server.
[0037] The data processing device 12 comprises a computer 22, a database 24, and a communication interface 26. The computer 22 is an example of a "computer" related to the technology of this disclosure. The computer 22 comprises a processor 28, RAM 30, and storage 32. The processor 28, RAM 30, and storage 32 are connected to a bus 34. The database 24 and the communication interface 26 are also connected to the bus 34. The communication interface 26 is connected to a network 54. An example of the network 54 is a WAN (Wide Area Network) and / or a LAN (Local Area Network).
[0038] The smart device 14 comprises a computer 36, a reception device 38, an output device 40, a camera 42, and a communication interface 44. The computer 36 comprises a processor 46, RAM 48, and storage 50. The processor 46, RAM 48, and storage 50 are connected to a bus 52. The reception device 38, output device 40, and camera 42 are also connected to the bus 52.
[0039] The reception device 38 is equipped with a touch panel 38A and a microphone 38B, etc., and receives user input. The touch panel 38A receives user input by detecting contact with an object (e.g., a pen or finger). The microphone 38B receives user input by detecting the user's voice. The control unit 46A transmits data indicating the user input received by the touch panel 38A and microphone 38B to the data processing device 12. In the data processing device 12, the specific processing unit 290 acquires the data indicating the user input.
[0040] The output device 40 includes a display 40A and a speaker 40B, and presents data to the user 20 by outputting the data in a form perceptible to the user 20 (e.g., audio and / or text). The display 40A displays visible information such as text and images according to instructions from the processor 46. The speaker 40B outputs audio according to instructions from the processor 46. The camera 42 is a small digital camera equipped with an optical system such as a lens, aperture, and shutter, and an image sensor such as a CMOS (Complementary Metal-Oxide-Semiconductor) image sensor or a CCD (Charge Coupled Device) image sensor.
[0041] Communication interface 44 is connected to network 54. Communication interfaces 44 and 26 are responsible for the exchange of various types of information between processor 46 and processor 28 via network 54.
[0042] Figure 2 shows an example of the main functions of the data processing device 12 and the smart device 14.
[0043] As shown in Figure 2, in the data processing device 12, a specific processing is performed by the processor 28. A specific processing program 56 is stored in the storage 32. The specific processing program 56 is an example of a "program" related to the technology of this disclosure. The processor 28 reads the specific processing program 56 from the storage 32 and executes the read specific processing program 56 on the RAM 30. The specific processing is realized by the processor 28 operating as a specific processing unit 290 according to the specific processing program 56 executed on the RAM 30.
[0044] The storage 32 stores the data generation model 58 and the emotion identification model 59. The data generation model 58 and the emotion identification model 59 are used by the identification processing unit 290.
[0045] In the smart device 14, the processor 46 performs the reception output processing. The storage 50 stores the reception output program 60. The reception output program 60 is used in conjunction with a specific processing program 56 by the data processing system 10. The processor 46 reads the reception output program 60 from the storage 50 and executes the read reception output program 60 on the RAM 48. The reception output processing is realized by the processor 46 operating as a control unit 46A according to the reception output program 60 executed on the RAM 48.
[0046] Next, the specific processing performed by the specific processing unit 290 of the data processing device 12 will be described. In the following description, the data processing device 12 will be referred to as the "server" and the smart device 14 as the "terminal".
[0047] The system of the present invention is configured to streamline customer service, sales, operation guidance, cash register management, and inventory management tasks that users experience in stores. The following describes in detail the embodiments for implementing the system of the present invention.
[0048] Customer service and sales
[0049] User: The user speaks to the robot to ask questions or provide information about products they want to purchase. For example, "I want to buy this smartphone, which plan do you recommend?"
[0050] Terminal: The speech recognition system converts the user's voice data into text data. The text data is then sent to the server.
[0051] Server: Analyzes text data to determine the user's purchase intent. For example, it recognizes needs regarding smartphone models and plans.
[0052] Server: Based on purchase intent, the system retrieves information from the database to suggest the optimal plan and generates a recommended plan. The recommended plan includes details such as price, data capacity, and benefits.
[0053] Terminal: The recommended plan information is converted into speech using a speech synthesis system, and the robot communicates the proposal to the user.
[0054] User: Review the proposal, ask further questions, or decide to purchase.
[0055] Instructions
[0056] User: The user asks the robot a question about a specific smartphone operation. For example, they might say, "I don't know how to send a photo, could you please tell me how?"
[0057] Terminal: The voice recognition system converts the user's voice data into text data. It also uses the camera function to recognize the model of the user's smartphone.
[0058] Server: Based on the image recognition results and text data, retrieves the operation manual for the corresponding model from the database.
[0059] Server: Generates specific operating procedures based on user requests from the operation manual.
[0060] Terminal: The generated operating procedure is converted into speech using a speech synthesis system, and the robot explains it to the user.
[0061] User: Operate the smartphone following the instructions provided.
[0062] Cash register management
[0063] Terminal: The POS terminal collects daily sales data and sends it to the server.
[0064] Server: Analyzes sales data and calculates the current cash inventory level.
[0065] Server: Generates an alert and sends a notification if the cash inventory falls below the appropriate level.
[0066] Terminal: Sends alert notifications to store administrators in real time.
[0067] User: The store manager who receives the notification will take action to address any cash discrepancies.
[0068] Inventory Management
[0069] Terminal: Registers product arrival and sales data in real time and sends it to the server.
[0070] Server: Calculates current inventory levels based on incoming and outgoing data.
[0071] Server: When inventory levels fall below a set threshold, it generates feedback and creates an ordering plan.
[0072] Terminal: Sends notifications to staff via robot to confirm the generated order plan.
[0073] User: Staff members who receive the notification will proceed with placing the order.
[0074] The system of this invention streamlines store operations and improves the quality of customer service. In particular, it enables users to receive quick and accurate information at stores, leading to high customer satisfaction. Furthermore, the automation of cash register management and inventory management solves the problem of labor shortages.
[0075] The following describes the processing flow.
[0076] Customer service and sales
[0077] Step 1:
[0078] Users can ask the robot questions or request information about products they want to buy. For example, they might ask, "I want to buy this smartphone, which plan do you recommend?"
[0079] Step 2:
[0080] The device uses a speech recognition system to convert the user's voice data into text data.
[0081] Step 3:
[0082] The terminal sends the converted text data to the server.
[0083] Step 4:
[0084] The server analyzes text data to determine the user's purchase intent. For example, it recognizes their needs regarding smartphone models and plans.
[0085] Step 5:
[0086] The server retrieves relevant information (product data and plan information) from the database.
[0087] Step 6:
[0088] The server generates the optimal plan based on the user's needs and creates text data containing detailed information about the recommended plan.
[0089] Step 7:
[0090] The text data generated by the server is returned to the terminal and converted into speech using a speech synthesis system.
[0091] Step 8:
[0092] The terminal, via a robot, provides the user with generated suggestions via voice.
[0093] Step 9:
[0094] The user reviews the proposal and decides on their next action (e.g., continue asking questions, make a purchase).
[0095] Instructions
[0096] Step 1:
[0097] The user asks the robot questions about specific smartphone operations. For example, they might say, "I don't know how to send a photo, could you please tell me how?"
[0098] Step 2:
[0099] The device uses a speech recognition system to convert the user's voice data into text data.
[0100] Step 3:
[0101] The device uses the smartphone's camera to perform image recognition to determine the model of the user's mobile device.
[0102] Step 4:
[0103] The terminal sends the converted text data and image recognition results to the server.
[0104] Step 5:
[0105] Based on the image recognition results, the server retrieves the corresponding model's operation manual from the database.
[0106] Step 6:
[0107] The server generates specific operating procedures based on user requests from the acquired operation manual.
[0108] Step 7:
[0109] The server returns the generated operating procedure as text data to the terminal, and the speech synthesis system converts it into speech.
[0110] Step 8:
[0111] The terminal provides the user with voice instructions on how to operate it via a robot.
[0112] Step 9:
[0113] The user operates the smartphone according to the instructions provided.
[0114] Cash register management
[0115] Step 1:
[0116] The terminal collects daily sales data from the POS terminal.
[0117] Step 2:
[0118] The terminal sends the collected sales data to the server.
[0119] Step 3:
[0120] The server analyzes sales data and calculates the current cash inventory level.
[0121] Step 4:
[0122] Based on the cash inventory amount calculated by the server, the appropriate inventory level is calculated.
[0123] Step 5:
[0124] The server compares the appropriate inventory level with the actual inventory level and generates an alert if there is a surplus or shortage.
[0125] Step 6:
[0126] The server sends the generated alert to the terminal.
[0127] Step 7:
[0128] The device sends alert notifications to the store administrator in real time.
[0129] Step 8:
[0130] The store manager who receives the user notification will then take action to address any cash discrepancies.
[0131] Inventory Management
[0132] Step 1:
[0133] The terminal registers product arrival and sales data in real time.
[0134] Step 2:
[0135] The device sends the registered data to the server.
[0136] Step 3:
[0137] The server calculates the current inventory level based on incoming and outgoing data.
[0138] Step 4:
[0139] The server generates feedback when the inventory level falls below a set threshold.
[0140] Step 5:
[0141] The server generates an order plan based on the feedback.
[0142] Step 6:
[0143] The server sends the generated order plan to the terminal.
[0144] Step 7:
[0145] The device sends notifications to staff via a robot.
[0146] Step 8:
[0147] The staff member who receives the user notification will then proceed with the ordering process.
[0148] Thus, the system of the present invention achieves efficient and high-quality store operations through each step.
[0149] (Example 1)
[0150] Next, we will describe Example 1. In the following description, the data processing device 12 will be referred to as the "server," and the smart device 14 will be referred to as the "terminal."
[0151] In traditional retail operations, a wide range of tasks, including customer service, inventory management, and cash register management, are performed manually. This is not only inefficient but also prone to human error and wasted time. As a result, problems such as decreased customer satisfaction, reduced operational efficiency, and increased burden due to labor shortages arise. A system is needed to solve these problems and provide timely and accurate information while improving operational efficiency.
[0152] The identification process performed by the identification processing unit 290 of the data processing device 12 in Example 1 is realized by the following means.
[0153] In this invention, the server includes means for converting user voice data into text data, means for analyzing the text data to determine purchase intent and operation requests, means for generating optimal suggestions or operation procedures based on the determination results, means for providing the generated suggestions or operation procedures to the user by speech synthesis, means for registering product arrival and sales data in real time and analyzing the data to calculate the current inventory level, means for formulating an order plan and notifying the user when the calculated inventory level falls below a set threshold, means for collecting sales data from a cash register terminal and analyzing the sales data to calculate the cash inventory level, means for generating an alert and notifying the user when the calculated cash inventory level falls below an appropriate level, and means for streamlining customer service, sales, operation guidance, cash register management, and inventory management that the user receives in the store. This enables automation and efficiency of store operations, improving the quality of user service and operational efficiency.
[0154] "Audio data" refers to sound information acquired digitally via a microphone or other audio input device, such as user speech or instructions.
[0155] "Text data" is a data format that converts audio data into written text.
[0156] "Analysis" is the process of understanding the meaning and intent of input data, and often involves using technologies such as natural language processing.
[0157] "Purchase intent" refers to the user's wishes and desires regarding the type and conditions of the product they are trying to acquire.
[0158] An "operation request" refers to a request from a user for assistance in performing a specific operation.
[0159] "Speech synthesis" is a technology that converts text data into speech that sounds like human speech.
[0160] "Product arrival data" refers to data that records information about when new products arrive at a store.
[0161] "Sales data" refers to data that records information about when a product was sold in a store.
[0162] "Inventory level" refers to the quantity of goods currently stored in the store or warehouse.
[0163] An "ordering plan" refers to a plan for purchasing new goods when inventory falls below a certain level.
[0164] A "cash register terminal" is an electronic device used for selling goods and processing payments.
[0165] "Sales data" refers to data that records information such as the quantity, price, and date and time of the transaction of the goods sold.
[0166] "Cash inventory" refers to the total amount of cash stored in cash registers and within the store.
[0167] An "alert" refers to a warning or cautionary message that is sent when certain conditions are met.
[0168] "Customer service" refers to the work of providing product descriptions, guidance, and answering questions from customers.
[0169] "Operation guidance" refers to the service of providing customers with instructions on how to use the equipment and services they use.
[0170] "Cash register management" refers to the task of managing daily sales and cash inflows and outflows.
[0171] "Inventory management" refers to the overall management of receiving, storing, and shipping goods in order to maintain an appropriate level of inventory.
[0172] The system of the present invention aims to streamline store operations and improve customer satisfaction. This system is configured to support the various tasks that users perform in stores, including customer service and sales, operation guidance, cash register management, and inventory management. The following describes specific embodiments for carrying out the present invention.
[0173] Customer service and sales
[0174] User: The user speaks to the robot to ask questions or provide information about products they want to purchase. For example, they might say, "I want to buy this smartphone, which plan do you recommend?"
[0175] Terminal: Uses a speech recognition system (e.g., a cloud-based speech recognition service) to convert the user's voice data into text data. The converted text data is sent to a server via the internet.
[0176] Server: Uses a Natural Language Processing (NLP) engine (e.g., a cloud-based natural language processing service) to analyze text data and determine the user's purchase intent. This purchase intent includes needs regarding smartphone models and plans.
[0177] Server: Based on the determination result, the server retrieves information on the corresponding plan from the database (e.g., relational database management system) and recommends the optimal plan. The recommended plan includes pricing, data capacity, and benefits.
[0178] Terminal: A speech synthesis system (e.g., speech synthesis API) converts the recommended plan information into speech, and a robot communicates the proposal to the user.
[0179] User: Review the proposal, ask further questions, or decide to purchase.
[0180] Specific example:
[0181] User: "I want to buy this smartphone, which plan would you recommend?"
[0182] Robot: "The best plan for you is the one with X GB of data for Y yen per month. This plan comes with the following benefits."
[0183] Examples of prompts to input into a generative AI model:
[0184] "Please generate a database query to analyze the user's purchase intent for the specified product and propose the optimal plan."
[0185] Instructions
[0186] User: The user asks the robot a question about a specific smartphone operation. For example, they might say, "I don't know how to send a photo, could you please tell me how?"
[0187] Terminal: A voice recognition system (e.g., a cloud-based voice recognition service) converts the user's voice data into text data. It also uses the camera function to recognize the model of the user's smartphone.
[0188] Server: Based on image recognition results (e.g., an image recognition model using deep learning) and text data, retrieves the operation manual for the corresponding model from the database.
[0189] Server: Generates specific operating procedures based on user requests from the operation manual.
[0190] Terminal: The generated operating procedure is converted into speech using a speech synthesis system (e.g., speech synthesis API), and the robot explains it to the user.
[0191] User: Operate the smartphone following the instructions provided.
[0192] Specific example:
[0193] User: "I don't know how to send photos, could you please tell me how?"
[0194] Robot: "Your device is a [model name]. To send a photo, first open your gallery, then press the share button and select the recipient."
[0195] Examples of prompts to input into a generative AI model:
[0196] "Generate specific steps to explain how to operate the smartphone as instructed by the user."
[0197] Cash register management
[0198] Terminal: A point-of-sale terminal (e.g., a point-of-sale management system) collects daily sales data and transmits it to a server via the internet.
[0199] Server: Aggregates and analyzes sales data to calculate the current cash inventory level.
[0200] Server: Generates an alert if the cash inventory falls below the appropriate level.
[0201] Terminal: Sends alert notifications to store administrators in real time.
[0202] User: The store manager who receives the notification will take action to address any cash discrepancies.
[0203] Specific example:
[0204] Server: "Cash inventory is insufficient. Please take immediate action."
[0205] Manager: "We'll replenish the cash immediately."
[0206] Examples of prompts to input into a generative AI model:
[0207] "Calculate cash inventory levels based on daily sales data and generate alerts when inventory levels are low."
[0208] Inventory Management
[0209] Terminal: Product arrival and sales data (e.g., sales management system) are registered in real time and transmitted to the server via the internet.
[0210] Server: Calculates current inventory levels based on incoming and outgoing data.
[0211] Server: When inventory levels fall below a set threshold, it generates feedback and develops an ordering plan.
[0212] Terminal: Sends notifications to staff via robot to confirm the generated order plan.
[0213] User: Staff members who receive the notification will proceed with placing the order.
[0214] Specific example:
[0215] Server: "Inventory levels have fallen below the threshold. An order is required."
[0216] Staff: "We have confirmed your order. We have placed the necessary items."
[0217] Examples of prompts to input into a generative AI model:
[0218] "Monitor current inventory levels and generate reordering plans when they fall below the set threshold."
[0219] The system of this invention is expected to automate many store operations and improve the quality of customer service. In particular, it will enable users to receive quick and accurate information at stores, resulting in high customer satisfaction. Furthermore, the automation of cash register management and inventory management will solve the problem of labor shortages.
[0220] The flow of the specific processing in Example 1 will be explained using Figure 11.
[0221] Customer service and sales
[0222] Step 1:
[0223] Users can ask the robot questions or request information about products they want to buy. For example, they might say, "I want to buy this smartphone, which plan do you recommend?"
[0224] Input: User's voice
[0225] Output: Audio data
[0226] Step 2:
[0227] The device uses a speech recognition system (e.g., a cloud-based speech recognition service) to convert speech data into text data. The converted text data is then sent to the server.
[0228] Input: Audio data
[0229] Output: Text data
[0230] Step 3:
[0231] The server uses a Natural Language Processing (NLP) engine (e.g., a cloud-based natural language processing service) to analyze text data and determine the user's purchase intent.
[0232] Input: Text data
[0233] Output: Classification result (user's purchase intent)
[0234] Step 4:
[0235] Based on the determination result, the server retrieves information on the corresponding plan from the database (e.g., a relational database management system) and recommends the optimal plan.
[0236] Input: Discrimination result
[0237] Output: Recommended plan (price, data capacity, benefits, etc.)
[0238] Step 5:
[0239] The terminal uses a speech synthesis system (e.g., a speech synthesis API) to convert the recommended plan information into speech, and a robot then communicates the proposal to the user.
[0240] Input: Recommended plan
[0241] Output: Audio data (proposed content)
[0242] Step 6:
[0243] The user reviews the proposal, asks further questions, or decides to make a purchase.
[0244] Input: Audio data (proposal)
[0245] Output: User decisions and questions
[0246] Adding specific actions
[0247] When a user asks, "I want to buy this smartphone, which plan do you recommend?", the robot suggests, "The best plan for you is the one with X GB of data for Y yen per month. This plan comes with the following benefits."
[0248] Instructions
[0249] Step 1:
[0250] The user asks the robot questions about specific smartphone operations. For example, they might say, "I don't know how to send a photo, could you please tell me how?"
[0251] Input: User's voice
[0252] Output: Audio data
[0253] Step 2:
[0254] The device converts the user's voice data into text data using a voice recognition system (e.g., a cloud-based voice recognition service). It also utilizes its camera function to recognize the model of the user's smartphone.
[0255] Input: Audio data and image data
[0256] Output: Text data and image recognition results
[0257] Step 3:
[0258] The server retrieves the operation manual for the corresponding model from the database based on the image recognition results (e.g., an image recognition model using deep learning) and text data.
[0259] Input: Text data and image recognition results
[0260] Output: Operation Manual
[0261] Step 4:
[0262] The server generates specific operating procedures based on the user's request, using the operation manual.
[0263] Input: Operation Manual
[0264] Output: Operating Procedure
[0265] Step 5:
[0266] The terminal converts the generated operating procedures into speech using a speech synthesis system (e.g., a speech synthesis API), and the robot explains them to the user.
[0267] Input: Operating Procedure
[0268] Output: Audio data (operating instructions)
[0269] Step 6:
[0270] The user operates the smartphone according to the instructions provided.
[0271] Input: Audio data (operating instructions)
[0272] Output: Smartphone operation results
[0273] Adding specific actions
[0274] When a user asks, "I don't know how to send photos, could you please tell me how?", the robot responds, "Your device is a XX. To send photos, first open your gallery, then press the share button and select the recipient."
[0275] Cash register management
[0276] Step 1:
[0277] A point-of-sale terminal (e.g., a point-of-sale management system) collects daily sales data and transmits it to a server via the internet.
[0278] Input: Sales data
[0279] Output: Transmission to server
[0280] Step 2:
[0281] The server aggregates and analyzes the sales data and calculates the current cash inventory level.
[0282] Input: Sales data
[0283] Output: Cash inventory level
[0284] Step 3:
[0285] If the cash inventory level is below the appropriate level, the server generates an alert.
[0286] Input: Cash inventory level
[0287] Output: Alert
[0288] Step 4:
[0289] The terminal sends the alert notification to the store manager in real time.
[0290] Input: Alert
[0291] Output: Alert notification
[0292] Step 5:
[0293] The store manager responds to the cash shortage or surplus upon receiving the notification.
[0294] Input: Alert notification
[0295] Output: Cash replenishment or reduction response
[0296] Addition of specific operations
[0297] The server sends an alert saying "The cash inventory is insufficient. Please respond immediately", and the store manager checks it and responds with "I will replenish the cash immediately".
[0298] Inventory Management
[0299] Step 1:
[0300] The terminal registers the incoming and sales data of products (e.g., sales management system) in real time and sends it to the server via the Internet.
[0301] Input: Incoming data and sales data
[0302] Output: Transmission to the server
[0303] Step 2:
[0304] The server calculates the current inventory based on the incoming and sales data.
[0305] Input: Incoming data and sales data
[0306] Output: Inventory quantity
[0307] Step 3:
[0308] When the inventory quantity falls below the set threshold, the server generates feedback and formulates an ordering plan.
[0309] Input: Inventory quantity
[0310] Output: Feedback and ordering plan
[0311] Step 4:
[0312] The terminal sends a notification to the staff via the robot to confirm the generated ordering plan.
[0313] Input: Ordering plan
[0314] Output: Notification
[0315] Step 5:
[0316] After receiving notification, the staff will proceed with placing the order.
[0317] Input: Notification
[0318] Output: Order result
[0319] Adding specific actions
[0320] The server notifies, "Inventory levels have fallen below the threshold. An order is required," and the staff responds, "We have confirmed the order. We have placed an order for the necessary items."
[0321] This invention is expected to automate many store operations and improve the quality of customer service. In particular, it will enable users to receive quick and accurate information at stores, resulting in high customer satisfaction. Furthermore, the automation of cash register management and inventory management will help solve the problem of labor shortages.
[0322] (Application Example 1)
[0323] Next, we will explain Application Example 1. In the following explanation, the data processing device 12 will be referred to as the "server," and the smart device 14 will be referred to as the "terminal."
[0324] In recent years, brick-and-mortar stores have been required to provide customers with the information they need quickly and accurately. However, variations in the number and knowledge levels of store staff can lead to delays in customer service and inability to provide appropriate information. Furthermore, internal operations such as inventory management and cash register management are also reliant on human labor, which can lead to errors and decreased efficiency. It is necessary to solve these problems and improve customer satisfaction while also increasing the efficiency of store operations.
[0325] The specific processing performed by the specific processing unit 290 of the data processing device 12 in Application Example 1 is realized by the following means.
[0326] In this invention, the server includes means for converting user voice data into text data, means for analyzing the text data to determine purchase intent and operation requests, means for generating optimal suggestions or operation procedures based on the determination results, means for providing the generated suggestions or operation procedures to the user by speech synthesis, means for recognizing questions about the user's smart device using voice and camera and providing information based on those questions, and means for acquiring product information from a database in real time and providing that information to the user by a speech synthesis system. This enables the rapid and accurate provision of information that customers seek, improving the efficiency of store operations and enhancing customer satisfaction.
[0327] "User voice data" refers to the voice information that the user speaks.
[0328] "Text data" refers to audio data converted into written information.
[0329] "Purchase intent" refers to a user's desire to buy a particular product or service.
[0330] An "operation request" is a user's desire to know a specific operation method or procedure.
[0331] A "proposal" refers to the optimal plan or options provided based on the user's needs.
[0332] "Operating procedures" refer to the specific steps or steps required for a user to perform a particular operation.
[0333] A "speech synthesis system" is a system that converts text data into speech and conveys information to the user audibly.
[0334] A "smart device" refers to a mobile device with advanced functions, such as a smartphone or tablet.
[0335] A "database" is an information management system that stores and allows searching for product information, operation manuals, and other data.
[0336] "Real-time" refers to a state where the current situation and data can be reflected immediately.
[0337] The system of this invention aims to improve the efficiency of customer service and internal operations in physical stores. This system utilizes speech recognition, text analysis, image recognition, and speech synthesis to provide users with optimal suggestions and operating procedures. The specific configuration and operation of the system are described below.
[0338] First, the user asks a question to a smart robot in the store. For example, they might say, "I want to buy this smartphone, which plan do you recommend?" The smart robot collects the voice data using its microphone. The speech recognition system installed in the device (for example, Google® Speech Recognition API) converts the voice data into text data. This text data is then sent to a server.
[0339] The server analyzes text data to determine the user's purchase intent and operational requests. For example, if a user asks about a smartphone plan, the server recognizes the user's needs from the content of the question. Next, the server retrieves relevant information from the database to generate the best possible suggestions based on this recognition. For example, it generates recommended smartphone plans and attaches detailed information such as price, data capacity, and benefits.
[0340] The generated suggestions are converted into speech using a speech synthesis system (e.g., Pyttsx3). The smart robot communicates the suggestions to the user through the converted speech. The user can then review them, ask further questions, or make a purchase decision.
[0341] The same applies when a user asks a question about operating a specific smartphone. For example, it can respond to prompts such as, "I don't know how to send photos, please tell me how." In addition to voice, the terminal uses its camera function to recognize the model of the user's smartphone. Based on the image recognition results and text data, the server retrieves the operation manual for that model from its database. Based on this operation manual, it generates specific operating procedures that meet the user's request and explains them to the user using speech synthesis.
[0342] Furthermore, this system contributes to the efficiency of internal operations. For example, it collects daily sales data from POS terminals and sends it to a server to calculate the cash inventory level. If the cash inventory level falls below the appropriate level, an alert is generated and a notification is sent. This notification reaches the store manager in real time, allowing for quick action to address any cash shortages or surpluses. In addition, for inventory management, it registers product arrival and sales data in real time and calculates the current inventory level. If the inventory level falls below a set threshold, it automatically generates feedback and creates an ordering plan.
[0343] This system enables the rapid and accurate provision of information that customers need, leading to increased efficiency in store operations and improved customer satisfaction.
[0344] For example:
[0345] "I'd like to buy this smartphone. Which plan would you recommend?"
[0346] "Do you have this item in stock?"
[0347] "How do I send a photo?"
[0348] These prompt messages will be recognized by the smart robot, enabling it to provide appropriate information and operating guidance.
[0349] The flow of a specific process in Application Example 1 will be explained using Figure 12.
[0350] Step 1:
[0351] The user asks a question to the smart robot. For example, they might say a prompt like, "I want to buy this smartphone, which plan do you recommend?" This voice data is then input into the robot.
[0352] Step 2:
[0353] The device uses a speech recognition system (Google Speech Recognition API) to convert the user's voice data into text data. This process converts the voice data into text data. This text data is output and sent to the server.
[0354] Step 3:
[0355] The server analyzes the input text data. Specifically, it identifies purchase intent and operational requests from the text data and recognizes the corresponding questions and intentions. This discrimination process yields the discrimination result.
[0356] Step 4:
[0357] The server generates optimal suggestions or operating procedures based on the determination results. For example, if there is a question about a smartphone plan, the server retrieves the plan information from the database and generates a recommended plan. This data processing generates suggestions and operating procedures. These generated results are output.
[0358] Step 5:
[0359] The server converts the generated suggestions or operating procedures into speech using a speech synthesis system (Pyttsx3). This process converts text data into speech data.
[0360] Step 6:
[0361] The terminal provides the user with synthesized speech suggestions. These suggestions are transmitted to the user as voice from the smart robot. Through this process, the user confirms the information provided.
[0362] Step 7:
[0363] The user asks additional questions or makes a purchase or action decision based on the information provided. Depending on the user's actions, the process either returns to step 1 or ends.
[0364] Step 8:
[0365] When a user asks a question about a specific smartphone operation (for example, "I don't know how to send a photo, can you tell me how?"), the device uses its camera function to perform image recognition on the user's smart device. This input allows the device's model information to be obtained.
[0366] Step 9:
[0367] The server retrieves the corresponding operation manual from the database based on the image recognition results and text data. This data processing identifies the appropriate operation manual. This information is then output.
[0368] Step 10:
[0369] The server generates specific operating procedures based on the user's request, using the operation manual. For example, it generates procedures for sending a photograph. This processing yields specific operating procedures. These generated results are then output.
[0370] Step 11:
[0371] The server converts the generated operating instructions into speech using a speech synthesis system and provides them to the user. This allows the user to operate their smart device by following the voice guidance.
[0372] Furthermore, an emotion engine that estimates the user's emotions may be incorporated. That is, the identification processing unit 290 may use the emotion identification model 59 to estimate the user's emotions and perform identification processing using the user's emotions.
[0373] The system of the present invention aims to streamline customer service, sales, operation guidance, cash register management, and inventory management tasks that users experience in stores, and further to recognize user emotions and provide optimal responses. The following describes in detail the embodiments for implementing the system of the present invention.
[0374] Customer service and sales
[0375] User: The user speaks to the robot to ask questions or provide information about products they want to purchase. For example, "I want to buy this smartphone, which plan do you recommend?"
[0376] Terminal: The speech recognition system converts the user's voice data into text data. Furthermore, an emotion engine is used to recognize emotions from the user's voice.
[0377] Terminal: Sends converted text data and sentiment data to the server.
[0378] Server: Analyzes text and sentiment data to determine the user's purchase intent and emotional state. For example, it recognizes the user's needs regarding smartphone models and plans, and determines whether the user is excited or calm.
[0379] Server: Retrieves relevant information (product data and plan information) from the database. Adjusts suggestions based on sentiment data.
[0380] Server: Sends the adjusted suggestions to the terminal. The speech synthesis system converts them into speech using appropriate tone and expression.
[0381] Terminal: The robot provides the user with generated suggestions via voice.
[0382] User: Review the proposal, ask further questions, or decide to purchase.
[0383] Instructions
[0384] User: The user asks the robot a question about a specific smartphone operation. For example, they might say, "I don't know how to send a photo, could you please tell me how?"
[0385] Terminal: The speech recognition system converts the user's voice data into text data. Simultaneously, the emotion engine recognizes the user's emotions.
[0386] Terminal: Sends converted text data and sentiment data to the server. It also uses the camera function to recognize the model of the user's smartphone.
[0387] Server: Based on image recognition results, text data, and emotion data, retrieves the corresponding model's operation manual from the database.
[0388] Server: Generates specific operating procedures based on user requests from the operation manual, and further adjusts the difficulty and tone of the explanations by considering sentiment data.
[0389] Server: Returns the adjusted operating procedure to the terminal and converts it into speech using the speech synthesis system.
[0390] Terminal: A robot provides voice instructions to the user. Tones and expressions are based on emotional data.
[0391] User: Operate the smartphone following the instructions provided.
[0392] Cash register management
[0393] Terminal: Collects daily sales data from the POS terminal.
[0394] Terminal: Sends collected sales data to the server.
[0395] Server: Analyzes sales data and calculates the current cash inventory level.
[0396] Server: Generates an alert and sends a notification if the cash inventory falls below the appropriate level.
[0397] Terminal: Sends alert notifications to store administrators in real time.
[0398] User: The store manager who receives the notification will take action to address any cash discrepancies.
[0399] Inventory Management
[0400] Terminal: Registers product arrival and sales data in real time.
[0401] Terminal: Sends registered data to the server.
[0402] Server: Calculates current inventory levels based on incoming and outgoing data.
[0403] Server: Generates feedback when inventory levels fall below a set threshold. Sentiment data is also used as needed.
[0404] Server: Generates order plans based on feedback.
[0405] Server: Sends the generated order plan to the terminal.
[0406] Terminal: Sends notifications to staff via a robot. It can also adjust tone and expression based on emotional data.
[0407] User: Staff members who receive the notification will proceed with placing the order.
[0408] This system will streamline store operations and provide high-quality service to both employees and customers. Furthermore, because it can recognize user emotions and adjust responses in real time, it is expected to further improve customer satisfaction.
[0409] The following describes the processing flow.
[0410] Customer service and sales
[0411] Step 1:
[0412] Users can ask the robot questions or request information about products they want to buy. For example, they might ask, "I want to buy this smartphone, which plan do you recommend?"
[0413] Step 2:
[0414] The device uses a speech recognition system to convert the user's voice data into text data. Simultaneously, an emotion engine recognizes the user's emotions from their voice.
[0415] Step 3:
[0416] The device sends the converted text data and sentiment data to the server.
[0417] Step 4:
[0418] The server analyzes text data to determine the user's purchase intent. For example, it recognizes their needs regarding smartphone models and plans.
[0419] Step 5:
[0420] The server retrieves relevant information (product data and plan information) from the database.
[0421] Step 6:
[0422] The server generates optimal suggestions while considering the user's emotions. For example, if the user is undecided, it adjusts the suggestions, such as emphasizing special offers.
[0423] Step 7:
[0424] The server generates suggestions and sends them to the terminal, where they are converted into speech using a speech synthesis system. The tone and expression of the converted speech are adjusted according to the user's emotions.
[0425] Step 8:
[0426] The terminal provides the user with generated suggestions via voice through a robot.
[0427] Step 9:
[0428] The user reviews the proposal and decides whether to continue asking questions or to make a purchase.
[0429] Instructions
[0430] Step 1:
[0431] The user asks the robot questions about specific smartphone operations. For example, they might say, "I don't know how to send a photo, could you please tell me how?"
[0432] Step 2:
[0433] The device uses a speech recognition system to convert the user's voice data into text data. Simultaneously, an emotion engine recognizes the user's emotions from their voice.
[0434] Step 3:
[0435] The device sends text and sentiment data to the server, and uses its camera function to perform image recognition to determine the model of the user's smartphone.
[0436] Step 4:
[0437] The server analyzes the image recognition results and text data to determine the user's operation request.
[0438] Step 5:
[0439] The server retrieves the operation manual for the corresponding model from the database.
[0440] Step 6:
[0441] The server generates specific operating procedures based on user requests from the acquired operation manual, and further adjusts the difficulty level and tone of the explanations by taking sentiment data into consideration.
[0442] Step 7:
[0443] The server generates operating instructions and sends them to the terminal as text data, which are then converted into speech using a speech synthesis system. The tone and expression of the speech are adjusted according to the user's emotions.
[0444] Step 8:
[0445] The terminal provides the user with voice instructions on how to operate it via a robot.
[0446] Step 9:
[0447] The user operates the smartphone according to the instructions provided.
[0448] Cash register management
[0449] Step 1:
[0450] The terminal collects daily sales data from the POS terminal.
[0451] Step 2:
[0452] The terminal sends the collected sales data to the server.
[0453] Step 3:
[0454] The server analyzes sales data and calculates the current cash inventory level.
[0455] Step 4:
[0456] Based on the cash inventory amount calculated by the server, the appropriate inventory level is calculated.
[0457] Step 5:
[0458] The server compares the appropriate inventory level with the actual inventory level and generates an alert if there is a surplus or shortage.
[0459] Step 6:
[0460] The server sends the generated alert to the terminal.
[0461] Step 7:
[0462] The device sends alert notifications to the store administrator in real time.
[0463] Step 8:
[0464] The store manager who receives the user notification will then take action to address any cash discrepancies.
[0465] Inventory Management
[0466] Step 1:
[0467] The terminal registers product arrival and sales data in real time.
[0468] Step 2:
[0469] The device sends the registered data to the server.
[0470] Step 3:
[0471] The server calculates the current inventory level based on incoming and outgoing data.
[0472] Step 4:
[0473] The server generates feedback when the inventory level falls below a set threshold.
[0474] Step 5:
[0475] The server generates an order plan based on the feedback.
[0476] Step 6:
[0477] The server sends the generated order plan to the terminal.
[0478] Step 7:
[0479] The device sends notifications to staff via a robot. It can also adjust the tone and expression based on emotional data.
[0480] Step 8:
[0481] The staff member who receives the user notification will then proceed with the ordering process.
[0482] Thus, the system of the present invention achieves efficient and high-quality store operations through each step and has the ability to recognize user emotions and adjust responses in real time. As a result, an improvement in customer satisfaction is expected.
[0483] (Example 2)
[0484] Next, we will describe Example 2. In the following description, the data processing device 12 will be referred to as the "server" and the smart device 14 as the "terminal".
[0485] Existing store management systems have limitations in improving customer satisfaction because they respond to user purchase intentions and operational requests in a uniform and mechanical manner. Furthermore, real-time responses to cash and inventory management are difficult, which can lead to decreased operational efficiency. Additionally, the lack of consideration for user emotions makes it difficult to maximize individual customer satisfaction.
[0486] The identification process performed by the identification processing unit 290 of the data processing device 12 in Example 2 is realized by the following means.
[0487] In this invention, the server includes means for converting user voice data into text data, means for analyzing the text data to determine purchase intent and operation requests, means for generating optimal suggestions or operation procedures based on the determination results, means for providing the generated suggestions or operation procedures to the user by speech synthesis, means for recognizing emotions from the user's voice, means for adjusting the content of suggestions or operation procedures based on the emotion recognition results, and means for acquiring relevant information based on the content of the user's purchase or operation request and emotional state. This enables flexible responses that take into account not only the user's purchase intent but also their emotional state, and is expected to significantly improve customer satisfaction. In addition, real-time cash inventory management and inventory management can be performed efficiently, improving the efficiency of store operations.
[0488] "User voice data" refers to the audio data emitted when a user speaks to the system.
[0489] "Text data" refers to string data obtained by converting audio data.
[0490] "Emotion recognition" refers to the process of analyzing and identifying a user's emotional state from their voice data.
[0491] "Suggested content" refers to appropriate answers and recommendation information that the system generates in response to user questions and requests.
[0492] "Operating procedures" refer to the specific methods of operation that the system provides in response to user questions or requests.
[0493] "Speech synthesis" refers to the technology that converts text data into speech and conveys it to the user in a natural-sounding voice.
[0494] "Purchase intent" refers to a user's intention or purpose when purchasing a particular product or service.
[0495] An "operation request" refers to a request from a user who wants to know a specific operation method or information.
[0496] "Related information" refers to product data and service information that is applicable to the user's questions or requests.
[0497] "Image recognition" refers to the technology that identifies specific objects or models from image data captured by a camera.
[0498] "Sales data" refers to data that records sales information from the store's cash register.
[0499] "Cash inventory" refers to numerical data that indicates the current amount of cash on hand.
[0500] An "alert" refers to a warning system that notifies you when certain conditions occur.
[0501] "Inventory level" refers to data that shows the current number of items in stock at a store.
[0502] "Feedback" refers to re-evaluation and instructional information generated by the system based on the analysis results.
[0503] An "ordering plan" refers to a plan for ordering new goods based on the results of inventory and cash management.
[0504] The system of this invention aims to provide users with information to guide them through purchasing and operating products, and to efficiently manage cash inventory and stock. The system's implementation primarily involves servers, terminals, and users. The specific configuration of the system is as follows:
[0505] System Configuration
[0506] hardware
[0507] Device: A device equipped with voice recognition, emotion recognition, and camera capabilities. Examples include smartphones and robotic devices for specific purposes.
[0508] Server: Consists of a database server and an analysis server. MySQL (registered trademark) is used as the database.
[0509] Database: Storage for product data, operation manuals, sales data, and inventory data.
[0510] software
[0511] Speech recognition system: Uses Google Speech-to-Text, etc. This converts the user's speech into text data.
[0512] Emotion Engine: Uses Emotion AI, which recognizes the user's emotions from their voice.
[0513] Speech synthesis system: Uses a text-to-speech engine such as Amazon Polly. This converts text data back into speech.
[0514] Detailed description of the system
[0515] Customer service and sales
[0516] 1. User: The user speaks into the device to ask questions or provide information about the product they want to purchase. For example, "I want to buy this smartphone, which plan do you recommend?"
[0517] 2. Terminal: The speech recognition system converts the user's voice data into text data. Furthermore, an emotion engine is used to recognize emotions from the user's voice.
[0518] 3. Terminal: Sends the converted text data and sentiment data to the server.
[0519] 4. Server: Analyzes text and sentiment data to determine the user's purchase intent and emotional state. For example, it recognizes the user's needs regarding smartphone models and plans, and determines whether the user is excited or calm.
[0520] 5. Server: Retrieves relevant information (product data and plan information) from the database. Adjusts suggestions based on sentiment data.
[0521] 6. Server: Sends the adjusted suggestions to the terminal. The speech synthesis system converts them into speech using appropriate tone and expression.
[0522] 7. Terminal: The robot provides the user with generated suggestions via voice.
[0523] 8. User: Review the proposal, ask further questions, or decide to purchase.
[0524] Specific example:
[0525] The user asks the robot, "I want to buy this smartphone, which plan do you recommend?"
[0526] The robot responds, "For this smartphone, we recommend a plan with a large data allowance."
[0527] Examples of prompts to input into a generative AI model:
[0528] A user asked, "I want to buy this smartphone, which plan do you recommend?" Please generate the best answer suggesting smartphone plans. The user seems a little excited.
[0529] Instructions
[0530] 1. User: The user asks the robot a question about a specific smartphone operation. For example, they might say, "I don't know how to send a photo, could you please tell me how?"
[0531] 2. Terminal: The speech recognition system converts the user's voice data into text data. Simultaneously, the emotion engine recognizes the user's emotions.
[0532] 3. Terminal: Sends the converted text data and sentiment data to the server. It also uses the camera function to recognize the model of the user's smartphone.
[0533] 4. Server: Based on image recognition results, text data, and emotion data, the server retrieves the corresponding model's operation manual from the database.
[0534] 5. Server: Generates specific operating procedures based on user requests from the operation manual, and further adjusts the difficulty and tone of the explanations by taking sentiment data into consideration.
[0535] 6. Server: Returns the adjusted operating procedure to the terminal and converts it into speech using the speech synthesis system.
[0536] 7. Terminal: The robot provides voice instructions to the user. Tones and expressions are based on emotional data.
[0537] 8. User: Operate the smartphone following the provided instructions.
[0538] Specific example:
[0539] The user asks the robot, "I don't know how to send a photo, could you please tell me how?"
[0540] The robot explains, "First, open the camera app, then tap the send button."
[0541] Examples of prompts to input into a generative AI model:
[0542] A user asked, "I don't know how to send photos, could you please explain?" Please provide clear instructions on how to do this on a smartphone. The user seems a little confused.
[0543] Cash register management
[0544] 1. Terminal: Collect daily sales data from the POS terminal.
[0545] 2. Terminal: Sends the collected sales data to the server.
[0546] 3. Server: Analyzes sales data and calculates the current cash inventory level.
[0547] 4. Server: Generates an alert and sends a notification if the cash inventory falls below the appropriate level.
[0548] 5. Terminal: Sends alert notifications to store administrators in real time.
[0549] 6. User: The store manager who receives the notification will take action to address any cash discrepancies.
[0550] Specific example:
[0551] If cash inventory falls below the appropriate level, the system will notify the store manager with the message, "Cash inventory has fallen below the appropriate level. Please check."
[0552] Examples of prompts to input into a generative AI model:
[0553] Analyze cash inventory levels based on sales data and generate an alert if they fall below the appropriate level.
[0554] Inventory Management
[0555] 1. Terminal: Registers product arrival and sales data in real time.
[0556] 2. Terminal: Sends registered data to the server.
[0557] 3. Server: Calculates the current inventory level based on incoming and outgoing data.
[0558] 4. Server: Generates feedback when inventory levels fall below a set threshold. Sentiment data is also used as needed.
[0559] 5. Server: Generates order plans based on feedback.
[0560] 6. Server: Sends the generated order plan to the terminal.
[0561] 7. Terminal: Sends notifications to staff via a robot. It can also adjust tone and expression based on emotional data.
[0562] 8. User: The staff member who receives the notification will proceed with the ordering process.
[0563] Specific example:
[0564] If inventory falls below a set threshold, the system will notify staff with a message saying, "Inventory is low. Please place an order."
[0565] Examples of prompts to input into a generative AI model:
[0566] Based on inventory data, analyze inventory levels and generate and notify us of an order plan if the levels fall below a set threshold.
[0567] Implementing this system requires a combination of appropriate speech recognition, emotion recognition, and speech synthesis systems. Furthermore, infrastructure development is necessary to ensure smooth data communication between the server and terminals. This configuration and operation will enable the provision of high-quality services that take user emotions into consideration, leading to improved efficiency in store operations and increased customer satisfaction.
[0568] The flow of the specific processing in Example 2 will be explained using Figure 13.
[0569] Customer service and sales
[0570] Step 1:
[0571] Users speak into the device to provide information about the products they want to purchase or to ask questions.
[0572] Input: User's voice data.
[0573] Output: Audio data is input to the terminal.
[0574] Specific action: The user says, "I want to buy this smartphone, which plan do you recommend?"
[0575] Step 2:
[0576] The device converts voice data into text data using a speech recognition system, and simultaneously recognizes emotions using an emotion engine.
[0577] Input: User's voice data.
[0578] Output: Converted text data and sentiment data.
[0579] Data processing / data calculation: Convert speech to text using Google Speech-to-Text and analyze emotions using Emotion AI.
[0580] Specific operation: From the voice data, text data such as "I want to buy this smartphone, which plan do you recommend?" and emotion data such as "excitement" are obtained.
[0581] Step 3:
[0582] The device sends text data and sentiment data to the server.
[0583] Input: Text data and sentiment data.
[0584] Output: Data sent to the server.
[0585] Specific operation: The device sends the converted text data and sentiment data to the server.
[0586] Step 4:
[0587] The server analyzes text data and sentiment data to determine the user's purchase intent and emotional state.
[0588] Input: Text data and sentiment data.
[0589] Output: User purchase intent and emotional state information.
[0590] Data processing / data calculation: Perform text data analysis and sentiment data analysis to determine user needs.
[0591] Specific actions: The system distinguishes between a purchase intention ("looking for a smartphone plan") and an emotional state ("excited").
[0592] Step 5:
[0593] The server retrieves relevant information (product data and plan information) from the database.
[0594] Input: Purchase intent and emotional state information.
[0595] Output: Related product information and plan information.
[0596] Data processing / data calculation: Search and retrieve relevant information from a MySQL database.
[0597] Specific action: Retrieve related information such as "Smartphone Plan A, B, C".
[0598] Step 6:
[0599] The server adjusts the suggestions based on sentiment data.
[0600] Input: Related product information, plan information, and sentiment data.
[0601] Output: Revised proposal.
[0602] Data processing / data calculation: Generate optimal suggestions for users based on emotional data.
[0603] Specific action: For excited users, propose a "simple and intuitive plan."
[0604] Step 7:
[0605] The server sends the adjusted proposal to the terminal, which then converts it into speech using a speech synthesis system.
[0606] Input: Adjusted proposal.
[0607] Output: Audio data.
[0608] Data processing / data calculation: Convert the proposed content into audio data using Amazon Polly.
[0609] Specific operation: The server sends a suggestion message saying "Simple Plan A is recommended," and the terminal converts it into speech.
[0610] Step 8:
[0611] The device provides the user with generated suggestions via voice.
[0612] Input: Audio data.
[0613] Output: Audio output to the user.
[0614] Specific action: The robot says, "I recommend the simple Plan A."
[0615] Step 9:
[0616] The user reviews the proposal, asks further questions, or decides to make a purchase.
[0617] Input: Audio of the proposed content.
[0618] Output: User's next action (continue questioning or make a purchase decision).
[0619] Specific action: The user accepts the offer and responds, "I will purchase it."
[0620] Instructions
[0621] Step 1:
[0622] A user asks the robot questions about specific smartphone operations.
[0623] Input: User's voice data.
[0624] Output: Audio data is input to the terminal.
[0625] Specific action: The user says, "I don't know how to send photos, could you please tell me how?"
[0626] Step 2:
[0627] The device converts voice data into text data using a speech recognition system, and simultaneously recognizes emotions using an emotion engine.
[0628] Input: User's voice data.
[0629] Output: Converted text data and sentiment data.
[0630] Data processing / data calculation: Convert speech to text using Google Speech-to-Text and analyze emotions using Emotion AI.
[0631] Specific operation: From the audio data, text data saying "I don't know how to send a photo, please tell me" and emotion data indicating "confusion" are obtained.
[0632] Step 3:
[0633] The device sends text data and sentiment data to the server, and the camera function is used to recognize the smartphone model.
[0634] Input: Text data, sentiment data, and smartphone image data.
[0635] Output: Data sent to the server and recognized model information.
[0636] Specific operation: The device sends the user's question to the server, and the smartphone image captured by the camera is analyzed to identify the model.
[0637] Step 4:
[0638] The server retrieves the corresponding model's operation manual from the database based on the image recognition results, text data, and emotion data.
[0639] Input: Image recognition results, text data, sentiment data.
[0640] Output: Operation manual for the applicable model.
[0641] Data processing / data calculation: Obtain the operation manual corresponding to the identified model from the database.
[0642] Specific operation: The server searches for and retrieves the user manual for the corresponding smartphone.
[0643] Step 5:
[0644] The server generates specific operating procedures from the user manual based on the user's request, and adjusts the difficulty and tone of the explanation based on sentiment data.
[0645] Input: Operation manual, emotion data.
[0646] Output: Adjusted operating instructions.
[0647] Data processing / data calculation: Generate operating procedures while considering emotional data, and explain them to the user in the most appropriate tone.
[0648] Specific actions: Explain "simple and intuitive steps" to confused users.
[0649] Step 6:
[0650] The server returns the adjusted operating procedure to the terminal, which then converts it into speech using a speech synthesis system.
[0651] Input: Adjusted operating procedure.
[0652] Output: Audio data.
[0653] Data processing / data calculation: Convert the operation procedure into audio data using Amazon Polly.
[0654] Specific operation: The server sends instructions such as "First open the camera app, then tap the send button," which the device then converts to audio.
[0655] Step 7:
[0656] The device provides the user with voice instructions on how to operate it. Tones and expressions based on emotional data are used.
[0657] Input: Audio data.
[0658] Output: Audio output to the user.
[0659] Specific instructions: The robot explains, "First, open the camera app, then tap the send button."
[0660] Step 8:
[0661] The user operates the smartphone according to the instructions provided.
[0662] Input: Audio instructions for the operating procedure.
[0663] Output: Smartphone operation complete.
[0664] Specific operation: The user follows the robot's instructions and operates their smartphone to send a photo.
[0665] Cash register management
[0666] Step 1:
[0667] The terminal collects daily sales data from the POS terminal.
[0668] Input: Sales data.
[0669] Output: Collected sales data.
[0670] Specific operation: The POS terminal saves daily sales data to the terminal.
[0671] Step 2:
[0672] The terminal sends the collected sales data to the server.
[0673] Input: Collected sales data.
[0674] Output: Sales data sent to the server.
[0675] Specific action: The terminal transfers sales data to the server.
[0676] Step 3:
[0677] The server analyzes sales data and calculates the current cash inventory level.
[0678] Input: Sales data.
[0679] Output: Cash inventory.
[0680] Data processing / calculation: Calculate cash inventory based on sales data.
[0681] Specific operation: The server analyzes sales data and calculates the amount of cash inventory.
[0682] Step 4:
[0683] The server generates an alert and sends a notification if the cash inventory falls below the appropriate level.
[0684] Input: Cash inventory quantity.
[0685] Output: Alert notification.
[0686] Data processing / calculation: Generate an alert when it is confirmed that the cash inventory level falls below the appropriate level.
[0687] Specific action: The server prepares the alert notification.
[0688] Step 5:
[0689] The device sends an alert notification to the store administrator.
[0690] Input: Alert notification data.
[0691] Output: Notification to the store manager.
[0692] Specific operation: The device notifies the store manager of alerts in real time.
[0693] Step 6:
[0694] The store manager who receives the notification from the user will then take action to address any cash discrepancies.
[0695] Input: Alert notification.
[0696] Output: Issue resolved.
[0697] Specific action: The store manager checks the notification and takes action to address any cash discrepancies.
[0698] Inventory Management
[0699] Step 1:
[0700] The terminal registers product data in real time.
[0701] Input: Incoming and sales data.
[0702] Output: Registered data.
[0703] Specific operation: The terminal updates the database with arrival and sales information in real time.
[0704] Step 2:
[0705] The device sends the registration data to the server.
[0706] Input: Registration data.
[0707] Output: Data sent to the server.
[0708] Specific action: The device transfers the registration data to the server.
[0709] Step 3:
[0710] The server calculates the current inventory level based on incoming and sales data.
[0711] Input: Registration data.
[0712] Output: Current inventory level.
[0713] Data processing / data calculation: Calculate inventory levels based on incoming and outgoing data.
[0714] Specific operation: The server analyzes incoming and outgoing data and calculates the inventory level.
[0715] Step 4:
[0716] When the server's inventory level falls below a set threshold, it generates feedback. Sentiment data is also used as needed.
[0717] Input: Inventory quantity.
[0718] Output: Feedback.
[0719] Data processing / data calculation: Generate feedback when inventory levels fall below a threshold.
[0720] Specific actions: Generate feedback and determine whether an order is necessary.
[0721] Step 5:
[0722] The server generates an order plan based on the feedback.
[0723] Input: Feedback.
[0724] Output: Ordering plan.
[0725] Data processing / data calculation: Create order plans and make necessary adjustments.
[0726] Specific operation: The server generates the order plan.
[0727] Step 6:
[0728] The server sends the generated order plan to the terminal.
[0729] Input: Order plan.
[0730] Output: Data sent to the terminal.
[0731] Specific action: The server sends the order plan to the terminal.
[0732] Step 7:
[0733] The device sends notifications to staff via a robot. It can also adjust the tone and expression based on emotional data.
[0734] Input: Order planning data.
[0735] Output: Notification to staff.
[0736] Specific operation: The terminal passes the order plan to the robot, and the robot notifies the staff.
[0737] Step 8:
[0738] The staff member who receives the user notification will then proceed with the ordering process.
[0739] Input: Notification data.
[0740] Output: Order completed.
[0741] Specific action: The staff member who receives the notification will carry out the necessary ordering tasks.
[0742] (Application Example 2)
[0743] Next, we will explain application example 2. In the following explanation, the data processing device 12 will be referred to as a "server" and the smart device 14 as a "terminal".
[0744] Traditional store operations have been cumbersome, involving tasks such as customer service, sales, operation guidance, cash register management, and inventory management, making efficient operation difficult. Furthermore, it has been challenging to respond flexibly to user emotions and needs, leading to a demand for improved customer satisfaction. This invention aims to solve these problems by providing a system that recognizes user emotions and provides optimal responses, thereby improving the efficiency of store operations and enhancing customer satisfaction.
[0745] In Application Example 2, the identification processing by the identification processing unit 290 of the data processing device 12 is realized by the following means. In this invention, the server includes means for converting the user's voice data into text data, means for analyzing the text data to determine purchase intent and operation requests, means for generating optimal suggestions or operation procedures based on the determination results and the user's emotional data, means for providing the generated suggestions or operation procedures to the user by speech synthesis, means for identifying the user's mobile terminal by image recognition and obtaining an operation manual corresponding to the model of the mobile terminal from a database, means for generating operation procedures according to the user's requests from the obtained operation manual and adjusting the generated operation procedures based on the user's emotional data, means for collecting sales data from a cash register terminal in real time and analyzing the sales data to calculate the amount of cash inventory, means for generating and notifying an alert when the calculated amount of cash inventory falls below an appropriate amount, and means for adjusting the generated alert notification based on the user's emotional data. This makes it possible to provide more personalized responses by reflecting the user's emotional data in the overall system adjustments, suggestions, and notifications.
[0746] "Audio data" refers to the audio signals spoken by the user.
[0747] "Text data" refers to digital information obtained by converting audio data into a string of characters.
[0748] "Emotional data" refers to data that indicates the emotional state of a user, as recognized from their voice and facial expressions.
[0749] "Purchase intent" refers to a user's intention to purchase a particular product or service.
[0750] An "operation request" refers to a user's desire to learn about a specific operation or usage method.
[0751] A "proposal" refers to an introduction or advice regarding specific products or services provided in response to user requests.
[0752] "Operating procedures" refer to specific methods of operation provided in response to user requests.
[0753] "Speech synthesis" is a technology that converts text data into audio signals for playback.
[0754] "Image recognition" is a technology that recognizes specific objects from image data acquired using input devices such as cameras.
[0755] A "database" is a collection of information stored in digital format, making it easy to search and retrieve.
[0756] "Sales data" refers to digital information about sales collected through point-of-sale terminals.
[0757] "Cash inventory" refers to the total amount of cash present in a store.
[0758] An "alert" is a warning signal used to notify users of important changes or anomalies in the situation.
[0759] This invention is a system that streamlines various tasks that users experience in stores, such as customer service, sales, operation guidance, cash register management, and inventory management, and further recognizes user emotions to provide optimal responses. The following describes the specific forms for implementing the system.
[0760] 1. Customer service and sales
[0761] User: The user speaks to the robot and asks questions or requests information about products they want to purchase. For example, "I want to buy this smartphone, which plan do you recommend?"
[0762] Terminal: The robot's microphone acquires voice data, and a speech recognition system (e.g., Python's speech_recognition library) converts the voice data into text data. It also uses an emotion engine (e.g., EmotionEngine) to recognize emotions from the user's voice.
[0763] Server: Analyzes text and sentiment data to determine the user's purchase intent and emotional state. For example, it recognizes the user's needs regarding smartphone models and plans, and determines whether the user is excited or calm.
[0764] Server: Retrieves relevant information (product data and plan information) from a database (e.g., a relational database management system) and adjusts the suggested content based on sentiment data.
[0765] Terminal: The adjusted suggestions are converted into speech by a speech synthesis system (e.g., TextToSpeech module) and provided to the user by the robot.
[0766] For example, if a user asks, "I want to buy this smartphone, which plan do you recommend?", the robot will respond to the excited user, "The unlimited data plan is the best!"
[0767] 2. Instructions for Use
[0768] User: The user asks the robot questions about how to use their smartphone. For example, they might say, "I don't know how to send a photo, could you please show me how?"
[0769] Terminal: The voice recognition system converts voice data into text data, and the emotion engine recognizes the user's emotions. It also uses the camera function to recognize the model of the user's smartphone.
[0770] Server: Based on image recognition results, text data, and emotion data, retrieves the corresponding model's operation manual from the database.
[0771] Server: Generates specific operating procedures based on user requests from the operation manual, and adjusts the difficulty and tone of the explanations considering sentiment data.
[0772] Terminal: The robot converts the adjusted operating procedures into speech using speech synthesis and explains them to the user.
[0773] For example, if a user asks, "I don't know how to send a photo, could you tell me how?", the robot calmly guides them by saying, "Select the photo you took, press the share button, and then choose the recipient."
[0774] 3. Cash register management
[0775] Terminal: The POS terminal collects daily sales data.
[0776] Server: Analyzes collected sales data and calculates cash inventory levels. A relational database management system is used.
[0777] Server: Generates an alert when cash inventory falls below the appropriate level and makes adjustments based on user sentiment data.
[0778] Terminal: Sends tailored alert notifications to store managers in real time. Notifications use a tone based on sentiment data.
[0779] 4. Inventory Management
[0780] Terminal: Registers product arrival and sales data in real time.
[0781] Server: Calculates the current inventory level based on registered data and generates feedback if the inventory level falls below a threshold. A relational database management system is used.
[0782] Server: Generates order plans based on feedback and adjusts them considering sentiment data.
[0783] Terminal: Notifies staff of the adjusted order plan via robot.
[0784] Examples of prompts for generative AI models
[0785] "It converts user questions into text and recognizes their emotions. Example: 'I want to buy this smartphone, which plan do you recommend?' Emotion: Excitement. Then, it retrieves relevant plan information from the product information database and generates the best recommendation based on the emotion."
[0786] This allows for more personalized responses by incorporating user sentiment data into overall system adjustments, suggestions, and notifications.
[0787] The flow of a specific process in Application Example 2 will be explained using Figure 14.
[0788] Step 1:
[0789] User: Speaks to the robot about the products they want to buy or any questions they have. The voice data from this interaction becomes the input.
[0790] Step 2:
[0791] Terminal: Audio data acquired via the microphone is converted into text data using a speech recognition system (such as Python's speech_recognition library). The converted text data is obtained as output.
[0792] Step 3:
[0793] Terminal: Uses an emotion engine (such as EmotionEngine) to recognize emotion data from the user's voice data. The recognized emotion data is obtained as output.
[0794] Step 4:
[0795] Terminal: Sends text data and sentiment data obtained through speech recognition to the server. Input data consists of text data and sentiment data, which are sent to the server.
[0796] Step 5:
[0797] Server: Analyzes received text and sentiment data to determine the user's purchase intent and action requests. The results obtained (e.g., purchase intent and action requests) become the output of the analysis.
[0798] Step 6:
[0799] Server: Based on relevant information (such as product data and plan information) obtained from the database (relational database management system), the server generates optimal suggestions based on the user's sentiment data. The generated suggestions are then output.
[0800] Step 7:
[0801] Server: Sends the generated proposal content to the terminal. The input data is the proposal content, which is then sent to the terminal.
[0802] Step 8:
[0803] Terminal: The received suggestion content is converted into speech using a speech synthesis system (such as the TextToSpeech module). Audio data is obtained as output.
[0804] Step 9:
[0805] Terminal: The robot plays voice data and provides suggestions to the user. The user reviews the suggestions and decides on their next action.
[0806] The specific processing unit 290 transmits the result of the specific processing to the smart device 14. In the smart device 14, the control unit 46A causes the output device 40 to output the result of the specific processing. The microphone 38B acquires audio indicating user input for the result of the specific processing. The control unit 46A transmits the audio data indicating user input acquired by the microphone 38B to the data processing device 12. In the data processing device 12, the specific processing unit 290 acquires the audio data.
[0807] Data generation model 58 is a so-called generative AI (Artificial Intelligence). An example of data generation model 58 is ChatGPT (registered trademark) (Internet search).<URL: https: / / openai.com / blog / chatgpt> ), Gemini (registered trademark) (Internet search) <url: https: gemini.google.com ?hl="ja">Examples of generative AI include the following. The data generation model 58 is obtained by performing deep learning on a neural network. The data generation model 58 is input with prompts containing instructions, and with inference data such as audio data representing speech, text data representing text, and image data representing images. The data generation model 58 infers from the input inference data according to the instructions indicated by the prompts, and outputs the inference results in data formats such as audio data and text data. Here, inference refers to, for example, analysis, classification, prediction, and / or summarization.
[0808] In the above embodiment, an example was given in which specific processing is performed by the data processing device 12, but the technology of this disclosure is not limited thereto, and the specific processing may also be performed by the smart device 14.
[0809] [Second Embodiment]
[0810] Figure 3 shows an example of the configuration of the data processing system 210 according to the second embodiment.
[0811] As shown in Figure 3, the data processing system 210 includes a data processing device 12 and smart glasses 214. An example of the data processing device 12 is a server.
[0812] The data processing device 12 comprises a computer 22, a database 24, and a communication interface 26. The computer 22 is an example of a "computer" related to the technology of this disclosure. The computer 22 comprises a processor 28, RAM 30, and storage 32. The processor 28, RAM 30, and storage 32 are connected to a bus 34. The database 24 and the communication interface 26 are also connected to the bus 34. The communication interface 26 is connected to a network 54. An example of the network 54 is a WAN (Wide Area Network) and / or a LAN (Local Area Network).
[0813] The smart glasses 214 include a computer 36, a microphone 238, a speaker 240, a camera 42, and a communication interface 44. The computer 36 includes a processor 46, RAM 48, and storage 50. The processor 46, RAM 48, and storage 50 are connected to a bus 52. The microphone 238, speaker 240, and camera 42 are also connected to the bus 52.
[0814] The microphone 238 receives voice signals from the user 20 and receives instructions from the user 20. The microphone 238 captures the voice signals from the user 20, converts the captured voice into audio data, and outputs it to the processor 46. The speaker 240 outputs audio according to the instructions from the processor 46.
[0815] Camera 42 is a small digital camera equipped with an optical system including a lens, aperture, and shutter, and an image sensor such as a CMOS (Complementary Metal-Oxide-Semiconductor) image sensor or a CCD (Charge Coupled Device) image sensor, and captures images of the area around the user 20 (for example, an imaging range defined by a field of view equivalent to the width of a typical healthy person's field of vision).
[0816] Communication interface 44 is connected to network 54. Communication interfaces 44 and 26 are responsible for the exchange of various information between processor 46 and processor 28 via network 54. The exchange of various information between processor 46 and processor 28 using communication interfaces 44 and 26 is performed in a secure manner.
[0817] Figure 4 shows an example of the main functions of the data processing device 12 and the smart glasses 214. As shown in Figure 4, the data processing device 12 performs specific processing using the processor 28. The storage 32 stores the specific processing program 56.
[0818] The specific processing program 56 is an example of a "program" relating to the technology of this disclosure. The processor 28 reads the specific processing program 56 from the storage 32 and executes the read specific processing program 56 on the RAM 30. The specific processing is realized by the processor 28 operating as a specific processing unit 290 in accordance with the specific processing program 56 executed on the RAM 30.
[0819] The storage 32 stores the data generation model 58 and the emotion identification model 59. The data generation model 58 and the emotion identification model 59 are used by the identification processing unit 290.
[0820] In the smart glasses 214, the processor 46 performs the reception output processing. The storage 50 stores the reception output program 60. The processor 46 reads the reception output program 60 from the storage 50 and executes the read reception output program 60 on the RAM 48. The reception output processing is realized by the processor 46 operating as a control unit 46A according to the reception output program 60 executed on the RAM 48.
[0821] Next, the identification processing performed by the identification processing unit 290 of the data processing device 12 will be described. In the following description, the data processing device 12 will be referred to as the "server" and the smart glasses 214 will be referred to as the "terminal".
[0822] The system of the present invention is configured to streamline customer service, sales, operation guidance, cash register management, and inventory management tasks that users experience in stores. The following describes in detail the embodiments for implementing the system of the present invention.
[0823] Customer service and sales
[0824] User: The user speaks to the robot to ask questions or provide information about products they want to purchase. For example, "I want to buy this smartphone, which plan do you recommend?"
[0825] Terminal: The speech recognition system converts the user's voice data into text data. The text data is then sent to the server.
[0826] Server: Analyzes text data to determine the user's purchase intent. For example, it recognizes needs regarding smartphone models and plans.
[0827] Server: Based on purchase intent, the system retrieves information from the database to suggest the optimal plan and generates a recommended plan. The recommended plan includes details such as price, data capacity, and benefits.
[0828] Terminal: The recommended plan information is converted into speech using a speech synthesis system, and the robot communicates the proposal to the user.
[0829] User: Review the proposal, ask further questions, or decide to purchase.
[0830] Instructions
[0831] User: The user asks the robot a question about a specific smartphone operation. For example, they might say, "I don't know how to send a photo, could you please tell me how?"
[0832] Terminal: The voice recognition system converts the user's voice data into text data. It also uses the camera function to recognize the model of the user's smartphone.
[0833] Server: Based on the image recognition results and text data, retrieves the operation manual for the corresponding model from the database.
[0834] Server: Generates specific operating procedures based on user requests from the operation manual.
[0835] Terminal: The generated operating procedure is converted into speech using a speech synthesis system, and the robot explains it to the user.
[0836] User: Operate the smartphone following the instructions provided.
[0837] Cash register management
[0838] Terminal: The POS terminal collects daily sales data and sends it to the server.
[0839] Server: Analyzes sales data and calculates the current cash inventory level.
[0840] Server: Generates an alert and sends a notification if the cash inventory falls below the appropriate level.
[0841] Terminal: Sends alert notifications to store administrators in real time.
[0842] User: The store manager who receives the notification will take action to address any cash discrepancies.
[0843] Inventory Management
[0844] Terminal: Registers product arrival and sales data in real time and sends it to the server.
[0845] Server: Calculates current inventory levels based on incoming and outgoing data.
[0846] Server: When inventory levels fall below a set threshold, it generates feedback and creates an ordering plan.
[0847] Terminal: Sends notifications to staff via robot to confirm the generated order plan.
[0848] User: Staff members who receive the notification will proceed with placing the order.
[0849] The system of this invention streamlines store operations and improves the quality of customer service. In particular, it enables users to receive quick and accurate information at stores, leading to high customer satisfaction. Furthermore, the automation of cash register management and inventory management solves the problem of labor shortages.
[0850] The following describes the processing flow.
[0851] Customer service and sales
[0852] Step 1:
[0853] Users can ask the robot questions or request information about products they want to buy. For example, they might ask, "I want to buy this smartphone, which plan do you recommend?"
[0854] Step 2:
[0855] The device uses a speech recognition system to convert the user's voice data into text data.
[0856] Step 3:
[0857] The terminal sends the converted text data to the server.
[0858] Step 4:
[0859] The server analyzes text data to determine the user's purchase intent. For example, it recognizes their needs regarding smartphone models and plans.
[0860] Step 5:
[0861] The server retrieves relevant information (product data and plan information) from the database.
[0862] Step 6:
[0863] The server generates the optimal plan based on the user's needs and creates text data containing detailed information about the recommended plan.
[0864] Step 7:
[0865] The text data generated by the server is returned to the terminal and converted into speech using a speech synthesis system.
[0866] Step 8:
[0867] The terminal, via a robot, provides the user with generated suggestions via voice.
[0868] Step 9:
[0869] The user reviews the proposal and decides on their next action (e.g., continue asking questions, make a purchase).
[0870] Instructions
[0871] Step 1:
[0872] The user asks the robot questions about specific smartphone operations. For example, they might say, "I don't know how to send a photo, could you please tell me how?"
[0873] Step 2:
[0874] The device uses a speech recognition system to convert the user's voice data into text data.
[0875] Step 3:
[0876] The device uses the smartphone's camera to perform image recognition to determine the model of the user's mobile device.
[0877] Step 4:
[0878] The terminal sends the converted text data and image recognition results to the server.
[0879] Step 5:
[0880] Based on the image recognition results, the server retrieves the corresponding model's operation manual from the database.
[0881] Step 6:
[0882] The server generates specific operating procedures based on user requests from the acquired operation manual.
[0883] Step 7:
[0884] The server returns the generated operating procedure as text data to the terminal, and the speech synthesis system converts it into speech.
[0885] Step 8:
[0886] The terminal provides the user with voice instructions on how to operate it via a robot.
[0887] Step 9:
[0888] The user operates the smartphone according to the instructions provided.
[0889] Cash register management
[0890] Step 1:
[0891] The terminal collects daily sales data from the POS terminal.
[0892] Step 2:
[0893] The terminal sends the collected sales data to the server.
[0894] Step 3:
[0895] The server analyzes sales data and calculates the current cash inventory level.
[0896] Step 4:
[0897] Based on the cash inventory amount calculated by the server, the appropriate inventory level is calculated.
[0898] Step 5:
[0899] The server compares the appropriate inventory level with the actual inventory level and generates an alert if there is a surplus or shortage.
[0900] Step 6:
[0901] The server sends the generated alert to the terminal.
[0902] Step 7:
[0903] The device sends alert notifications to the store administrator in real time.
[0904] Step 8:
[0905] The store manager who receives the user notification will then take action to address any cash discrepancies.
[0906] Inventory Management
[0907] Step 1:
[0908] The terminal registers product arrival and sales data in real time.
[0909] Step 2:
[0910] The device sends the registered data to the server.
[0911] Step 3:
[0912] The server calculates the current inventory level based on incoming and outgoing data.
[0913] Step 4:
[0914] The server generates feedback when the inventory level falls below a set threshold.
[0915] Step 5:
[0916] The server generates an order plan based on the feedback.
[0917] Step 6:
[0918] The server sends the generated order plan to the terminal.
[0919] Step 7:
[0920] The device sends notifications to staff via a robot.
[0921] Step 8:
[0922] The staff member who receives the user notification will then proceed with the ordering process.
[0923] Thus, the system of the present invention achieves efficient and high-quality store operations through each step.
[0924] (Example 1)
[0925] Next, we will describe Example 1. In the following description, the data processing device 12 will be referred to as the "server" and the smart glasses 214 will be referred to as the "terminal".
[0926] In traditional retail operations, a wide range of tasks, including customer service, inventory management, and cash register management, are performed manually. This is not only inefficient but also prone to human error and wasted time. As a result, problems such as decreased customer satisfaction, reduced operational efficiency, and increased burden due to labor shortages arise. A system is needed to solve these problems and provide timely and accurate information while improving operational efficiency.
[0927] The identification process performed by the identification processing unit 290 of the data processing device 12 in Example 1 is realized by the following means.
[0928] In this invention, the server includes means for converting user voice data into text data, means for analyzing the text data to determine purchase intent and operation requests, means for generating optimal suggestions or operation procedures based on the determination results, means for providing the generated suggestions or operation procedures to the user by speech synthesis, means for registering product arrival and sales data in real time and analyzing the data to calculate the current inventory level, means for formulating an order plan and notifying the user when the calculated inventory level falls below a set threshold, means for collecting sales data from a cash register terminal and analyzing the sales data to calculate the cash inventory level, means for generating an alert and notifying the user when the calculated cash inventory level falls below an appropriate level, and means for streamlining customer service, sales, operation guidance, cash register management, and inventory management that the user receives in the store. This enables automation and efficiency of store operations, improving the quality of user service and operational efficiency.
[0929] "Audio data" refers to sound information acquired digitally via a microphone or other audio input device, such as user speech or instructions.
[0930] "Text data" is a data format that converts audio data into written text.
[0931] "Analysis" is the process of understanding the meaning and intent of input data, and often involves using technologies such as natural language processing.
[0932] "Purchase intent" refers to the user's wishes and desires regarding the type and conditions of the product they are trying to acquire.
[0933] An "operation request" refers to a request from a user for assistance in performing a specific operation.
[0934] "Speech synthesis" is a technology that converts text data into speech that sounds like human speech.
[0935] "Product arrival data" refers to data that records information about when new products arrive at a store.
[0936] "Sales data" refers to data that records information about when a product was sold in a store.
[0937] "Inventory level" refers to the quantity of goods currently stored in the store or warehouse.
[0938] An "ordering plan" refers to a plan for purchasing new goods when inventory falls below a certain level.
[0939] A "cash register terminal" is an electronic device used for selling goods and processing payments.
[0940] "Sales data" refers to data that records information such as the quantity, price, and date and time of the transaction of the goods sold.
[0941] "Cash inventory" refers to the total amount of cash stored in cash registers and within the store.
[0942] An "alert" refers to a warning or cautionary message that is sent when certain conditions are met.
[0943] "Customer service" refers to the work of providing product descriptions, guidance, and answering questions from customers.
[0944] "Operation guidance" refers to the service of providing customers with instructions on how to use the equipment and services they use.
[0945] "Cash register management" refers to the task of managing daily sales and cash inflows and outflows.
[0946] "Inventory management" refers to the overall management of receiving, storing, and shipping goods in order to maintain an appropriate level of inventory.
[0947] The system of the present invention aims to streamline store operations and improve customer satisfaction. This system is configured to support the various tasks that users perform in stores, including customer service and sales, operation guidance, cash register management, and inventory management. The following describes specific embodiments for carrying out the present invention.
[0948] Customer service and sales
[0949] User: The user speaks to the robot to ask questions or provide information about products they want to purchase. For example, they might say, "I want to buy this smartphone, which plan do you recommend?"
[0950] Terminal: Uses a speech recognition system (e.g., a cloud-based speech recognition service) to convert the user's voice data into text data. The converted text data is sent to a server via the internet.
[0951] Server: Uses a Natural Language Processing (NLP) engine (e.g., a cloud-based natural language processing service) to analyze text data and determine the user's purchase intent. This purchase intent includes needs regarding smartphone models and plans.
[0952] Server: Based on the determination result, the server retrieves information on the corresponding plan from the database (e.g., relational database management system) and recommends the optimal plan. The recommended plan includes pricing, data capacity, and benefits.
[0953] Terminal: A speech synthesis system (e.g., speech synthesis API) converts the recommended plan information into speech, and a robot communicates the proposal to the user.
[0954] User: Review the proposal, ask further questions, or decide to purchase.
[0955] Specific example:
[0956] User: "I want to buy this smartphone, which plan would you recommend?"
[0957] Robot: "The best plan for you is the one with X GB of data for Y yen per month. This plan comes with the following benefits."
[0958] Examples of prompts to input into a generative AI model:
[0959] "Please generate a database query to analyze the user's purchase intent for the specified product and propose the optimal plan."
[0960] Instructions
[0961] User: The user asks the robot a question about a specific smartphone operation. For example, they might say, "I don't know how to send a photo, could you please tell me how?"
[0962] Terminal: A voice recognition system (e.g., a cloud-based voice recognition service) converts the user's voice data into text data. It also uses the camera function to recognize the model of the user's smartphone.
[0963] Server: Based on image recognition results (e.g., an image recognition model using deep learning) and text data, retrieves the operation manual for the corresponding model from the database.
[0964] Server: Generates specific operating procedures based on user requests from the operation manual.
[0965] Terminal: The generated operating procedure is converted into speech using a speech synthesis system (e.g., speech synthesis API), and the robot explains it to the user.
[0966] User: Operate the smartphone following the instructions provided.
[0967] Specific example:
[0968] User: "I don't know how to send photos, could you please tell me how?"
[0969] Robot: "Your device is a [model name]. To send a photo, first open your gallery, then press the share button and select the recipient."
[0970] Examples of prompts to input into a generative AI model:
[0971] "Generate specific steps to explain how to operate the smartphone as instructed by the user."
[0972] Cash register management
[0973] Terminal: A point-of-sale terminal (e.g., a point-of-sale management system) collects daily sales data and transmits it to a server via the internet.
[0974] Server: Aggregates and analyzes sales data to calculate the current cash inventory level.
[0975] Server: Generates an alert if the cash inventory falls below the appropriate level.
[0976] Terminal: Sends alert notifications to store administrators in real time.
[0977] User: The store manager who receives the notification will take action to address any cash discrepancies.
[0978] Specific example:
[0979] Server: "Cash inventory is insufficient. Please take immediate action."
[0980] Manager: "We'll replenish the cash immediately."
[0981] Examples of prompts to input into a generative AI model:
[0982] "Calculate cash inventory levels based on daily sales data and generate alerts when inventory levels are low."
[0983] Inventory Management
[0984] Terminal: Product arrival and sales data (e.g., sales management system) are registered in real time and transmitted to the server via the internet.
[0985] Server: Calculates current inventory levels based on incoming and outgoing data.
[0986] Server: When inventory levels fall below a set threshold, it generates feedback and develops an ordering plan.
[0987] Terminal: Sends notifications to staff via robot to confirm the generated order plan.
[0988] User: Staff members who receive the notification will proceed with placing the order.
[0989] Specific example:
[0990] Server: "Inventory levels have fallen below the threshold. An order is required."
[0991] Staff: "We have confirmed your order. We have placed the necessary items."
[0992] Examples of prompts to input into a generative AI model:
[0993] "Monitor current inventory levels and generate reordering plans when they fall below the set threshold."
[0994] The system of this invention is expected to automate many store operations and improve the quality of customer service. In particular, it will enable users to receive quick and accurate information at stores, resulting in high customer satisfaction. Furthermore, the automation of cash register management and inventory management will solve the problem of labor shortages.
[0995] The flow of the specific processing in Example 1 will be explained using Figure 11.
[0996] Customer service and sales
[0997] Step 1:
[0998] Users can ask the robot questions or request information about products they want to buy. For example, they might say, "I want to buy this smartphone, which plan do you recommend?"
[0999] Input: User's voice
[1000] Output: Audio data
[1001] Step 2:
[1002] The device uses a speech recognition system (e.g., a cloud-based speech recognition service) to convert speech data into text data. The converted text data is then sent to the server.
[1003] Input: Audio data
[1004] Output: Text data
[1005] Step 3:
[1006] The server uses a Natural Language Processing (NLP) engine (e.g., a cloud-based natural language processing service) to analyze text data and determine the user's purchase intent.
[1007] Input: Text data
[1008] Output: Classification result (user's purchase intent)
[1009] Step 4:
[1010] Based on the determination result, the server retrieves information on the corresponding plan from the database (e.g., a relational database management system) and recommends the optimal plan.
[1011] Input: Discrimination result
[1012] Output: Recommended plan (price, data capacity, benefits, etc.)
[1013] Step 5:
[1014] The terminal uses a speech synthesis system (e.g., a speech synthesis API) to convert the recommended plan information into speech, and a robot then communicates the proposal to the user.
[1015] Input: Recommended plan
[1016] Output: Audio data (proposed content)
[1017] Step 6:
[1018] The user reviews the proposal, asks further questions, or decides to make a purchase.
[1019] Input: Audio data (proposal)
[1020] Output: User decisions and questions
[1021] Adding specific actions
[1022] When a user asks, "I want to buy this smartphone, which plan do you recommend?", the robot suggests, "The best plan for you is the one with X GB of data for Y yen per month. This plan comes with the following benefits."
[1023] Instructions
[1024] Step 1:
[1025] The user asks the robot questions about specific smartphone operations. For example, they might say, "I don't know how to send a photo, could you please tell me how?"
[1026] Input: User's voice
[1027] Output: Audio data
[1028] Step 2:
[1029] The device converts the user's voice data into text data using a voice recognition system (e.g., a cloud-based voice recognition service). It also utilizes its camera function to recognize the model of the user's smartphone.
[1030] Input: Audio data and image data
[1031] Output: Text data and image recognition results
[1032] Step 3:
[1033] The server retrieves the operation manual for the corresponding model from the database based on the image recognition results (e.g., an image recognition model using deep learning) and text data.
[1034] Input: Text data and image recognition results
[1035] Output: Operation Manual
[1036] Step 4:
[1037] The server generates specific operating procedures based on the user's request, using the operation manual.
[1038] Input: Operation Manual
[1039] Output: Operating Procedure
[1040] Step 5:
[1041] The terminal converts the generated operating procedures into speech using a speech synthesis system (e.g., a speech synthesis API), and the robot explains them to the user.
[1042] Input: Operating Procedure
[1043] Output: Audio data (operating instructions)
[1044] Step 6:
[1045] The user operates the smartphone according to the instructions provided.
[1046] Input: Audio data (operating instructions)
[1047] Output: Smartphone operation results
[1048] Adding specific actions
[1049] When a user asks, "I don't know how to send photos, could you please tell me how?", the robot responds, "Your device is a XX. To send photos, first open your gallery, then press the share button and select the recipient."
[1050] Cash register management
[1051] Step 1:
[1052] A point-of-sale terminal (e.g., a point-of-sale management system) collects daily sales data and transmits it to a server via the internet.
[1053] Input: Sales data
[1054] Output: Send to server
[1055] Step 2:
[1056] The server aggregates and analyzes sales data to calculate the current cash inventory level.
[1057] Input: Sales data
[1058] Output: Cash Inventory
[1059] Step 3:
[1060] The server generates an alert if the cash inventory falls below the appropriate level.
[1061] Input: Cash inventory
[1062] Output: Alert
[1063] Step 4:
[1064] The device sends alert notifications to the store administrator in real time.
[1065] Input: Alert
[1066] Output: Alert notification
[1067] Step 5:
[1068] The store manager will take action to address any cash discrepancies upon receiving notification.
[1069] Input: Alert notification
[1070] Output: Cash replenishment and cash depletion response
[1071] Adding specific actions
[1072] The server sends an alert saying, "Cash inventory is low. Please take immediate action," and the store manager checks it and responds, "We will replenish the cash immediately."
[1073] Inventory Management
[1074] Step 1:
[1075] The terminal registers product arrival and sales data (e.g., sales management system) in real time and transmits it to the server via the internet.
[1076] Input: Incoming data and sales data
[1077] Output: Send to server
[1078] Step 2:
[1079] The server calculates the current inventory level based on incoming and outgoing data.
[1080] Input: Incoming data and sales data
[1081] Output: Inventory Quantity
[1082] Step 3:
[1083] The server generates feedback and develops an ordering plan when inventory levels fall below a set threshold.
[1084] Input: Inventory Quantity
[1085] Output: Feedback and ordering plan
[1086] Step 4:
[1087] The terminal sends a notification to staff via a robot for them to review the generated order plan.
[1088] Input: Ordering plan
[1089] Output: Notification
[1090] Step 5:
[1091] After receiving notification, the staff will proceed with placing the order.
[1092] Input: Notification
[1093] Output: Order result
[1094] Adding specific actions
[1095] The server notifies, "Inventory levels have fallen below the threshold. An order is required," and the staff responds, "We have confirmed the order. We have placed an order for the necessary items."
[1096] This invention is expected to automate many store operations and improve the quality of customer service. In particular, it will enable users to receive quick and accurate information at stores, resulting in high customer satisfaction. Furthermore, the automation of cash register management and inventory management will help solve the problem of labor shortages.
[1097] (Application Example 1)
[1098] Next, we will explain Application Example 1. In the following explanation, the data processing device 12 will be referred to as the "server," and the smart glasses 214 will be referred to as the "terminal."
[1099] In recent years, brick-and-mortar stores have been required to provide customers with the information they need quickly and accurately. However, variations in the number and knowledge levels of store staff can lead to delays in customer service and inability to provide appropriate information. Furthermore, internal operations such as inventory management and cash register management are also reliant on human labor, which can lead to errors and decreased efficiency. It is necessary to solve these problems and improve customer satisfaction while also increasing the efficiency of store operations.
[1100] The specific processing performed by the specific processing unit 290 of the data processing device 12 in Application Example 1 is realized by the following means.
[1101] In this invention, the server includes means for converting user voice data into text data, means for analyzing the text data to determine purchase intent and operation requests, means for generating optimal suggestions or operation procedures based on the determination results, means for providing the generated suggestions or operation procedures to the user by speech synthesis, means for recognizing questions about the user's smart device using voice and camera and providing information based on those questions, and means for acquiring product information from a database in real time and providing that information to the user by a speech synthesis system. This enables the rapid and accurate provision of information that customers seek, improving the efficiency of store operations and enhancing customer satisfaction.
[1102] "User voice data" refers to the voice information that the user speaks.
[1103] "Text data" refers to audio data converted into written information.
[1104] "Purchase intent" refers to a user's desire to buy a particular product or service.
[1105] An "operation request" is a user's desire to know a specific operation method or procedure.
[1106] A "proposal" refers to the optimal plan or options provided based on the user's needs.
[1107] "Operating procedures" refer to the specific steps or steps required for a user to perform a particular operation.
[1108] A "speech synthesis system" is a system that converts text data into speech and conveys information to the user audibly.
[1109] A "smart device" refers to a mobile device with advanced functions, such as a smartphone or tablet.
[1110] A "database" is an information management system that stores and allows searching for product information, operation manuals, and other data.
[1111] "Real-time" refers to a state where the current situation and data can be reflected immediately.
[1112] The system of this invention aims to improve the efficiency of customer service and internal operations in physical stores. This system utilizes speech recognition, text analysis, image recognition, and speech synthesis to provide users with optimal suggestions and operating procedures. The specific configuration and operation of the system are described below.
[1113] First, the user asks a question to a smart robot in the store. For example, they might say, "I want to buy this smartphone, which plan do you recommend?" The smart robot collects the voice data using its microphone. The speech recognition system installed in the device (for example, Google Speech Recognition API) converts the voice data into text data. This text data is then sent to a server.
[1114] The server analyzes text data to determine the user's purchase intent and operational requests. For example, if a user asks about a smartphone plan, the server recognizes the user's needs from the content of the question. Next, the server retrieves relevant information from the database to generate the best possible suggestions based on this recognition. For example, it generates recommended smartphone plans and attaches detailed information such as price, data capacity, and benefits.
[1115] The generated suggestions are converted into speech using a speech synthesis system (e.g., Pyttsx3). The smart robot communicates the suggestions to the user through the converted speech. The user can then review them, ask further questions, or make a purchase decision.
[1116] The same applies when a user asks a question about operating a specific smartphone. For example, it can respond to prompts such as, "I don't know how to send photos, please tell me how." In addition to voice, the terminal uses its camera function to recognize the model of the user's smartphone. Based on the image recognition results and text data, the server retrieves the operation manual for that model from its database. Based on this operation manual, it generates specific operating procedures that meet the user's request and explains them to the user using speech synthesis.
[1117] Furthermore, this system contributes to the efficiency of internal operations. For example, it collects daily sales data from POS terminals and sends it to a server to calculate the cash inventory level. If the cash inventory level falls below the appropriate level, an alert is generated and a notification is sent. This notification reaches the store manager in real time, allowing for quick action to address any cash shortages or surpluses. In addition, for inventory management, it registers product arrival and sales data in real time and calculates the current inventory level. If the inventory level falls below a set threshold, it automatically generates feedback and creates an ordering plan.
[1118] This system enables the rapid and accurate provision of information that customers need, leading to increased efficiency in store operations and improved customer satisfaction.
[1119] For example:
[1120] "I'd like to buy this smartphone. Which plan would you recommend?"
[1121] "Do you have this item in stock?"
[1122] "How do I send a photo?"
[1123] These prompt messages will be recognized by the smart robot, enabling it to provide appropriate information and operating guidance.
[1124] The flow of a specific process in Application Example 1 will be explained using Figure 12.
[1125] Step 1:
[1126] The user asks a question to the smart robot. For example, they might say a prompt like, "I want to buy this smartphone, which plan do you recommend?" This voice data is then input into the robot.
[1127] Step 2:
[1128] The device uses a speech recognition system (Google Speech Recognition API) to convert the user's voice data into text data. This process converts the voice data into text data. This text data is output and sent to the server.
[1129] Step 3:
[1130] The server analyzes the input text data. Specifically, it identifies purchase intent and operational requests from the text data and recognizes the corresponding questions and intentions. This discrimination process yields the discrimination result.
[1131] Step 4:
[1132] The server generates optimal suggestions or operating procedures based on the determination results. For example, if there is a question about a smartphone plan, the server retrieves the plan information from the database and generates a recommended plan. This data processing generates suggestions and operating procedures. These generated results are output.
[1133] Step 5:
[1134] The server converts the generated suggestions or operating procedures into speech using a speech synthesis system (Pyttsx3). This process converts text data into speech data.
[1135] Step 6:
[1136] The terminal provides the user with synthesized speech suggestions. These suggestions are transmitted to the user as voice from the smart robot. Through this process, the user confirms the information provided.
[1137] Step 7:
[1138] The user asks additional questions or makes a purchase or action decision based on the information provided. Depending on the user's actions, the process either returns to step 1 or ends.
[1139] Step 8:
[1140] When a user asks a question about a specific smartphone operation (for example, "I don't know how to send a photo, can you tell me how?"), the device uses its camera function to perform image recognition on the user's smart device. This input allows the device's model information to be obtained.
[1141] Step 9:
[1142] The server retrieves the corresponding operation manual from the database based on the image recognition results and text data. This data processing identifies the appropriate operation manual. This information is then output.
[1143] Step 10:
[1144] The server generates specific operating procedures based on the user's request, using the operation manual. For example, it generates procedures for sending a photograph. This processing yields specific operating procedures. These generated results are then output.
[1145] Step 11:
[1146] The server converts the generated operating instructions into speech using a speech synthesis system and provides them to the user. This allows the user to operate their smart device by following the voice guidance.
[1147] Furthermore, an emotion engine that estimates the user's emotions may be incorporated. That is, the identification processing unit 290 may use the emotion identification model 59 to estimate the user's emotions and perform identification processing using the user's emotions.
[1148] The system of the present invention aims to streamline customer service, sales, operation guidance, cash register management, and inventory management tasks that users experience in stores, and further to recognize user emotions and provide optimal responses. The following describes in detail the embodiments for implementing the system of the present invention.
[1149] Customer service and sales
[1150] User: The user speaks to the robot to ask questions or provide information about products they want to purchase. For example, "I want to buy this smartphone, which plan do you recommend?"
[1151] Terminal: The speech recognition system converts the user's voice data into text data. Furthermore, an emotion engine is used to recognize emotions from the user's voice.
[1152] Terminal: Sends converted text data and sentiment data to the server.
[1153] Server: Analyzes text and sentiment data to determine the user's purchase intent and emotional state. For example, it recognizes the user's needs regarding smartphone models and plans, and determines whether the user is excited or calm.
[1154] Server: Retrieves relevant information (product data and plan information) from the database. Adjusts suggestions based on sentiment data.
[1155] Server: Sends the adjusted suggestions to the terminal. The speech synthesis system converts them into speech using appropriate tone and expression.
[1156] Terminal: The robot provides the user with generated suggestions via voice.
[1157] User: Review the proposal, ask further questions, or decide to purchase.
[1158] Instructions
[1159] User: The user asks the robot a question about a specific smartphone operation. For example, they might say, "I don't know how to send a photo, could you please tell me how?"
[1160] Terminal: The speech recognition system converts the user's voice data into text data. Simultaneously, the emotion engine recognizes the user's emotions.
[1161] Terminal: Sends converted text data and sentiment data to the server. It also uses the camera function to recognize the model of the user's smartphone.
[1162] Server: Based on image recognition results, text data, and emotion data, retrieves the corresponding model's operation manual from the database.
[1163] Server: Generates specific operating procedures based on user requests from the operation manual, and further adjusts the difficulty and tone of the explanations by considering sentiment data.
[1164] Server: Returns the adjusted operating procedure to the terminal and converts it into speech using the speech synthesis system.
[1165] Terminal: A robot provides voice instructions to the user. Tones and expressions are based on emotional data.
[1166] User: Operate the smartphone following the instructions provided.
[1167] Cash register management
[1168] Terminal: Collects daily sales data from the POS terminal.
[1169] Terminal: Sends collected sales data to the server.
[1170] Server: Analyzes sales data and calculates the current cash inventory level.
[1171] Server: Generates an alert and sends a notification if the cash inventory falls below the appropriate level.
[1172] Terminal: Sends alert notifications to store administrators in real time.
[1173] User: The store manager who receives the notification will take action to address any cash discrepancies.
[1174] Inventory Management
[1175] Terminal: Registers product arrival and sales data in real time.
[1176] Terminal: Sends registered data to the server.
[1177] Server: Calculates current inventory levels based on incoming and outgoing data.
[1178] Server: Generates feedback when inventory levels fall below a set threshold. Sentiment data is also used as needed.
[1179] Server: Generates order plans based on feedback.
[1180] Server: Sends the generated order plan to the terminal.
[1181] Terminal: Sends notifications to staff via a robot. It can also adjust tone and expression based on emotional data.
[1182] User: Staff members who receive the notification will proceed with placing the order.
[1183] This system will streamline store operations and provide high-quality service to both employees and customers. Furthermore, because it can recognize user emotions and adjust responses in real time, it is expected to further improve customer satisfaction.
[1184] The following describes the processing flow.
[1185] Customer service and sales
[1186] Step 1:
[1187] Users can ask the robot questions or request information about products they want to buy. For example, they might ask, "I want to buy this smartphone, which plan do you recommend?"
[1188] Step 2:
[1189] The device uses a speech recognition system to convert the user's voice data into text data. Simultaneously, an emotion engine recognizes the user's emotions from their voice.
[1190] Step 3:
[1191] The device sends the converted text data and sentiment data to the server.
[1192] Step 4:
[1193] The server analyzes text data to determine the user's purchase intent. For example, it recognizes their needs regarding smartphone models and plans.
[1194] Step 5:
[1195] The server retrieves relevant information (product data and plan information) from the database.
[1196] Step 6:
[1197] The server generates optimal suggestions while considering the user's emotions. For example, if the user is undecided, it adjusts the suggestions, such as emphasizing special offers.
[1198] Step 7:
[1199] The server generates suggestions and sends them to the terminal, where they are converted into speech using a speech synthesis system. The tone and expression of the converted speech are adjusted according to the user's emotions.
[1200] Step 8:
[1201] The terminal provides the user with generated suggestions via voice through a robot.
[1202] Step 9:
[1203] The user reviews the proposal and decides whether to continue asking questions or to make a purchase.
[1204] Instructions
[1205] Step 1:
[1206] The user asks the robot questions about specific smartphone operations. For example, they might say, "I don't know how to send a photo, could you please tell me how?"
[1207] Step 2:
[1208] The device uses a speech recognition system to convert the user's voice data into text data. Simultaneously, an emotion engine recognizes the user's emotions from their voice.
[1209] Step 3:
[1210] The device sends text and sentiment data to the server, and uses its camera function to perform image recognition to determine the model of the user's smartphone.
[1211] Step 4:
[1212] The server analyzes the image recognition results and text data to determine the user's operation request.
[1213] Step 5:
[1214] The server retrieves the operation manual for the corresponding model from the database.
[1215] Step 6:
[1216] The server generates specific operating procedures based on user requests from the acquired operation manual, and further adjusts the difficulty level and tone of the explanations by taking sentiment data into consideration.
[1217] Step 7:
[1218] The server generates operating instructions and sends them to the terminal as text data, which are then converted into speech using a speech synthesis system. The tone and expression of the speech are adjusted according to the user's emotions.
[1219] Step 8:
[1220] The terminal provides the user with voice instructions on how to operate it via a robot.
[1221] Step 9:
[1222] The user operates the smartphone according to the instructions provided.
[1223] Cash register management
[1224] Step 1:
[1225] The terminal collects daily sales data from the POS terminal.
[1226] Step 2:
[1227] The terminal sends the collected sales data to the server.
[1228] Step 3:
[1229] The server analyzes sales data and calculates the current cash inventory level.
[1230] Step 4:
[1231] Based on the cash inventory amount calculated by the server, the appropriate inventory level is calculated.
[1232] Step 5:
[1233] The server compares the appropriate inventory level with the actual inventory level and generates an alert if there is a surplus or shortage.
[1234] Step 6:
[1235] The server sends the generated alert to the terminal.
[1236] Step 7:
[1237] The device sends alert notifications to the store administrator in real time.
[1238] Step 8:
[1239] The store manager who receives the user notification will then take action to address any cash discrepancies.
[1240] Inventory Management
[1241] Step 1:
[1242] The terminal registers product arrival and sales data in real time.
[1243] Step 2:
[1244] The device sends the registered data to the server.
[1245] Step 3:
[1246] The server calculates the current inventory level based on incoming and outgoing data.
[1247] Step 4:
[1248] The server generates feedback when the inventory level falls below a set threshold.
[1249] Step 5:
[1250] The server generates an order plan based on the feedback.
[1251] Step 6:
[1252] The server sends the generated order plan to the terminal.
[1253] Step 7:
[1254] The device sends notifications to staff via a robot. It can also adjust the tone and expression based on emotional data.
[1255] Step 8:
[1256] The staff member who receives the user notification will then proceed with the ordering process.
[1257] Thus, the system of the present invention achieves efficient and high-quality store operations through each step and has the ability to recognize user emotions and adjust responses in real time. As a result, an improvement in customer satisfaction is expected.
[1258] (Example 2)
[1259] Next, we will describe Example 2. In the following description, the data processing device 12 will be referred to as the "server" and the smart glasses 214 will be referred to as the "terminal".
[1260] Existing store management systems have limitations in improving customer satisfaction because they respond to user purchase intentions and operational requests in a uniform and mechanical manner. Furthermore, real-time responses to cash and inventory management are difficult, which can lead to decreased operational efficiency. Additionally, the lack of consideration for user emotions makes it difficult to maximize individual customer satisfaction.
[1261] The identification process performed by the identification processing unit 290 of the data processing device 12 in Example 2 is realized by the following means.
[1262] In this invention, the server includes means for converting user voice data into text data, means for analyzing the text data to determine purchase intent and operation requests, means for generating optimal suggestions or operation procedures based on the determination results, means for providing the generated suggestions or operation procedures to the user by speech synthesis, means for recognizing emotions from the user's voice, means for adjusting the content of suggestions or operation procedures based on the emotion recognition results, and means for acquiring relevant information based on the content of the user's purchase or operation request and emotional state. This enables flexible responses that take into account not only the user's purchase intent but also their emotional state, and is expected to significantly improve customer satisfaction. In addition, real-time cash inventory management and inventory management can be performed efficiently, improving the efficiency of store operations.
[1263] "User voice data" refers to the audio data emitted when a user speaks to the system.
[1264] "Text data" refers to string data obtained by converting audio data.
[1265] "Emotion recognition" refers to the process of analyzing and identifying a user's emotional state from their voice data.
[1266] "Suggested content" refers to appropriate answers and recommendation information that the system generates in response to user questions and requests.
[1267] "Operating procedures" refer to the specific methods of operation that the system provides in response to user questions or requests.
[1268] "Speech synthesis" refers to the technology that converts text data into speech and conveys it to the user in a natural-sounding voice.
[1269] "Purchase intent" refers to a user's intention or purpose when purchasing a particular product or service.
[1270] An "operation request" refers to a request from a user who wants to know a specific operation method or information.
[1271] "Related information" refers to product data and service information that is applicable to the user's questions or requests.
[1272] "Image recognition" refers to the technology that identifies specific objects or models from image data captured by a camera.
[1273] "Sales data" refers to data that records sales information from the store's cash register.
[1274] "Cash inventory" refers to numerical data that indicates the current amount of cash on hand.
[1275] An "alert" refers to a warning system that notifies you when certain conditions occur.
[1276] "Inventory level" refers to data that shows the current number of items in stock at a store.
[1277] "Feedback" refers to re-evaluation and instructional information generated by the system based on the analysis results.
[1278] An "ordering plan" refers to a plan for ordering new goods based on the results of inventory and cash management.
[1279] The system of this invention aims to provide users with information to guide them through purchasing and operating products, and to efficiently manage cash inventory and stock. The system's implementation primarily involves servers, terminals, and users. The specific configuration of the system is as follows:
[1280] System Configuration
[1281] hardware
[1282] Device: A device equipped with voice recognition, emotion recognition, and camera capabilities. Examples include smartphones and robotic devices for specific purposes.
[1283] The server consists of a database server and an analysis server. MySQL is used as the database.
[1284] Database: Storage for product data, operation manuals, sales data, and inventory data.
[1285] software
[1286] Speech recognition system: Uses Google Speech-to-Text, etc. This converts the user's speech into text data.
[1287] Emotion Engine: Uses Emotion AI, which recognizes the user's emotions from their voice.
[1288] Speech synthesis system: Uses a text-to-speech engine such as Amazon Polly. This converts text data back into speech.
[1289] Detailed description of the system
[1290] Customer service and sales
[1291] 1. User: The user speaks into the device to ask questions or provide information about the product they want to purchase. For example, "I want to buy this smartphone, which plan do you recommend?"
[1292] 2. Terminal: The speech recognition system converts the user's voice data into text data. Furthermore, an emotion engine is used to recognize emotions from the user's voice.
[1293] 3. Terminal: Sends the converted text data and sentiment data to the server.
[1294] 4. Server: Analyzes text and sentiment data to determine the user's purchase intent and emotional state. For example, it recognizes the user's needs regarding smartphone models and plans, and determines whether the user is excited or calm.
[1295] 5. Server: Retrieves relevant information (product data and plan information) from the database. Adjusts suggestions based on sentiment data.
[1296] 6. Server: Sends the adjusted suggestions to the terminal. The speech synthesis system converts them into speech using appropriate tone and expression.
[1297] 7. Terminal: The robot provides the user with generated suggestions via voice.
[1298] 8. User: Review the proposal, ask further questions, or decide to purchase.
[1299] Specific example:
[1300] The user asks the robot, "I want to buy this smartphone, which plan do you recommend?"
[1301] The robot responds, "For this smartphone, we recommend a plan with a large data allowance."
[1302] Examples of prompts to input into a generative AI model:
[1303] A user asked, "I want to buy this smartphone, which plan do you recommend?" Please generate the best answer suggesting smartphone plans. The user seems a little excited.
[1304] Instructions
[1305] 1. User: The user asks the robot a question about a specific smartphone operation. For example, they might say, "I don't know how to send a photo, could you please tell me how?"
[1306] 2. Terminal: The speech recognition system converts the user's voice data into text data. Simultaneously, the emotion engine recognizes the user's emotions.
[1307] 3. Terminal: Sends the converted text data and sentiment data to the server. It also uses the camera function to recognize the model of the user's smartphone.
[1308] 4. Server: Based on image recognition results, text data, and emotion data, the server retrieves the corresponding model's operation manual from the database.
[1309] 5. Server: Generates specific operating procedures based on user requests from the operation manual, and further adjusts the difficulty and tone of the explanations by taking sentiment data into consideration.
[1310] 6. Server: Returns the adjusted operating procedure to the terminal and converts it into speech using the speech synthesis system.
[1311] 7. Terminal: The robot provides voice instructions to the user. Tones and expressions are based on emotional data.
[1312] 8. User: Operate the smartphone following the provided instructions.
[1313] Specific example:
[1314] The user asks the robot, "I don't know how to send a photo, could you please tell me how?"
[1315] The robot explains, "First, open the camera app, then tap the send button."
[1316] Examples of prompts to input into a generative AI model:
[1317] A user asked, "I don't know how to send photos, could you please explain?" Please provide clear instructions on how to do this on a smartphone. The user seems a little confused.
[1318] Cash register management
[1319] 1. Terminal: Collect daily sales data from the POS terminal.
[1320] 2. Terminal: Sends the collected sales data to the server.
[1321] 3. Server: Analyzes sales data and calculates the current cash inventory level.
[1322] 4. Server: Generates an alert and sends a notification if the cash inventory falls below the appropriate level.
[1323] 5. Terminal: Sends alert notifications to store administrators in real time.
[1324] 6. User: The store manager who receives the notification will take action to address any cash discrepancies.
[1325] Specific example:
[1326] If cash inventory falls below the appropriate level, the system will notify the store manager with the message, "Cash inventory has fallen below the appropriate level. Please check."
[1327] Examples of prompts to input into a generative AI model:
[1328] Analyze cash inventory levels based on sales data and generate an alert if they fall below the appropriate level.
[1329] Inventory Management
[1330] 1. Terminal: Registers product arrival and sales data in real time.
[1331] 2. Terminal: Sends registered data to the server.
[1332] 3. Server: Calculates the current inventory level based on incoming and outgoing data.
[1333] 4. Server: Generates feedback when inventory levels fall below a set threshold. Sentiment data is also used as needed.
[1334] 5. Server: Generates order plans based on feedback.
[1335] 6. Server: Sends the generated order plan to the terminal.
[1336] 7. Terminal: Sends notifications to staff via a robot. It can also adjust tone and expression based on emotional data.
[1337] 8. User: The staff member who receives the notification will proceed with the ordering process.
[1338] Specific example:
[1339] If inventory falls below a set threshold, the system will notify staff with a message saying, "Inventory is low. Please place an order."
[1340] Examples of prompts to input into a generative AI model:
[1341] Based on inventory data, analyze inventory levels and generate and notify us of an order plan if the levels fall below a set threshold.
[1342] Implementing this system requires a combination of appropriate speech recognition, emotion recognition, and speech synthesis systems. Furthermore, infrastructure development is necessary to ensure smooth data communication between the server and terminals. This configuration and operation will enable the provision of high-quality services that take user emotions into consideration, leading to improved efficiency in store operations and increased customer satisfaction.
[1343] The flow of the specific processing in Example 2 will be explained using Figure 13.
[1344] Customer service and sales
[1345] Step 1:
[1346] Users speak into the device to provide information about the products they want to purchase or to ask questions.
[1347] Input: User's voice data.
[1348] Output: Audio data is input to the terminal.
[1349] Specific action: The user says, "I want to buy this smartphone, which plan do you recommend?"
[1350] Step 2:
[1351] The device converts voice data into text data using a speech recognition system, and simultaneously recognizes emotions using an emotion engine.
[1352] Input: User's voice data.
[1353] Output: Converted text data and sentiment data.
[1354] Data processing / data calculation: Convert speech to text using Google Speech-to-Text and analyze emotions using Emotion AI.
[1355] Specific operation: From the voice data, text data such as "I want to buy this smartphone, which plan do you recommend?" and emotion data such as "excitement" are obtained.
[1356] Step 3:
[1357] The device sends text data and sentiment data to the server.
[1358] Input: Text data and sentiment data.
[1359] Output: Data sent to the server.
[1360] Specific operation: The device sends the converted text data and sentiment data to the server.
[1361] Step 4:
[1362] The server analyzes text data and sentiment data to determine the user's purchase intent and emotional state.
[1363] Input: Text data and sentiment data.
[1364] Output: User purchase intent and emotional state information.
[1365] Data processing / data calculation: Perform text data analysis and sentiment data analysis to determine user needs.
[1366] Specific actions: The system distinguishes between a purchase intention ("looking for a smartphone plan") and an emotional state ("excited").
[1367] Step 5:
[1368] The server retrieves relevant information (product data and plan information) from the database.
[1369] Input: Purchase intent and emotional state information.
[1370] Output: Related product information and plan information.
[1371] Data processing / data calculation: Search and retrieve relevant information from a MySQL database.
[1372] Specific action: Retrieve related information such as "Smartphone Plan A, B, C".
[1373] Step 6:
[1374] The server adjusts the suggestions based on sentiment data.
[1375] Input: Related product information, plan information, and sentiment data.
[1376] Output: Revised proposal.
[1377] Data processing / data calculation: Generate optimal suggestions for users based on emotional data.
[1378] Specific action: For excited users, propose a "simple and intuitive plan."
[1379] Step 7:
[1380] The server sends the adjusted proposal to the terminal, which then converts it into speech using a speech synthesis system.
[1381] Input: Adjusted proposal.
[1382] Output: Audio data.
[1383] Data processing / data calculation: Convert the proposed content into audio data using Amazon Polly.
[1384] Specific operation: The server sends a suggestion message saying "Simple Plan A is recommended," and the terminal converts it into speech.
[1385] Step 8:
[1386] The device provides the user with generated suggestions via voice.
[1387] Input: Audio data.
[1388] Output: Audio output to the user.
[1389] Specific action: The robot says, "I recommend the simple Plan A."
[1390] Step 9:
[1391] The user reviews the proposal, asks further questions, or decides to make a purchase.
[1392] Input: Audio of the proposed content.
[1393] Output: User's next action (continue questioning or make a purchase decision).
[1394] Specific action: The user accepts the offer and responds, "I will purchase it."
[1395] Instructions
[1396] Step 1:
[1397] A user asks the robot questions about specific smartphone operations.
[1398] Input: User's voice data.
[1399] Output: Audio data is input to the terminal.
[1400] Specific action: The user says, "I don't know how to send photos, could you please tell me how?"
[1401] Step 2:
[1402] The device converts voice data into text data using a speech recognition system, and simultaneously recognizes emotions using an emotion engine.
[1403] Input: User's voice data.
[1404] Output: Converted text data and sentiment data.
[1405] Data processing / data calculation: Convert speech to text using Google Speech-to-Text and analyze emotions using Emotion AI.
[1406] Specific operation: From the audio data, text data saying "I don't know how to send a photo, please tell me" and emotion data indicating "confusion" are obtained.
[1407] Step 3:
[1408] The device sends text data and sentiment data to the server, and the camera function is used to recognize the smartphone model.
[1409] Input: Text data, sentiment data, and smartphone image data.
[1410] Output: Data sent to the server and recognized model information.
[1411] Specific operation: The device sends the user's question to the server, and the smartphone image captured by the camera is analyzed to identify the model.
[1412] Step 4:
[1413] The server retrieves the corresponding model's operation manual from the database based on the image recognition results, text data, and emotion data.
[1414] Input: Image recognition results, text data, sentiment data.
[1415] Output: Operation manual for the applicable model.
[1416] Data processing / data calculation: Obtain the operation manual corresponding to the identified model from the database.
[1417] Specific operation: The server searches for and retrieves the user manual for the corresponding smartphone.
[1418] Step 5:
[1419] The server generates specific operating procedures from the user manual based on the user's request, and adjusts the difficulty and tone of the explanation based on sentiment data.
[1420] Input: Operation manual, emotion data.
[1421] Output: Adjusted operating instructions.
[1422] Data processing / data calculation: Generate operating procedures while considering emotional data, and explain them to the user in the most appropriate tone.
[1423] Specific actions: Explain "simple and intuitive steps" to confused users.
[1424] Step 6:
[1425] The server returns the adjusted operating procedure to the terminal, which then converts it into speech using a speech synthesis system.
[1426] Input: Adjusted operating procedure.
[1427] Output: Audio data.
[1428] Data processing / data calculation: Convert the operation procedure into audio data using Amazon Polly.
[1429] Specific operation: The server sends instructions such as "First open the camera app, then tap the send button," which the device then converts to audio.
[1430] Step 7:
[1431] The device provides the user with voice instructions on how to operate it. Tones and expressions based on emotional data are used.
[1432] Input: Audio data.
[1433] Output: Audio output to the user.
[1434] Specific instructions: The robot explains, "First, open the camera app, then tap the send button."
[1435] Step 8:
[1436] The user operates the smartphone according to the instructions provided.
[1437] Input: Audio instructions for the operating procedure.
[1438] Output: Smartphone operation complete.
[1439] Specific operation: The user follows the robot's instructions and operates their smartphone to send a photo.
[1440] Cash register management
[1441] Step 1:
[1442] The terminal collects daily sales data from the POS terminal.
[1443] Input: Sales data.
[1444] Output: Collected sales data.
[1445] Specific operation: The POS terminal saves daily sales data to the terminal.
[1446] Step 2:
[1447] The terminal sends the collected sales data to the server.
[1448] Input: Collected sales data.
[1449] Output: Sales data sent to the server.
[1450] Specific action: The terminal transfers sales data to the server.
[1451] Step 3:
[1452] The server analyzes sales data and calculates the current cash inventory level.
[1453] Input: Sales data.
[1454] Output: Cash inventory.
[1455] Data processing / calculation: Calculate cash inventory based on sales data.
[1456] Specific operation: The server analyzes sales data and calculates the amount of cash inventory.
[1457] Step 4:
[1458] The server generates an alert and sends a notification if the cash inventory falls below the appropriate level.
[1459] Input: Cash inventory quantity.
[1460] Output: Alert notification.
[1461] Data processing / calculation: Generate an alert when it is confirmed that the cash inventory level falls below the appropriate level.
[1462] Specific action: The server prepares the alert notification.
[1463] Step 5:
[1464] The device sends an alert notification to the store administrator.
[1465] Input: Alert notification data.
[1466] Output: Notification to the store manager.
[1467] Specific operation: The device notifies the store manager of alerts in real time.
[1468] Step 6:
[1469] The store manager who receives the notification from the user will then take action to address any cash discrepancies.
[1470] Input: Alert notification.
[1471] Output: Issue resolved.
[1472] Specific action: The store manager checks the notification and takes action to address any cash discrepancies.
[1473] Inventory Management
[1474] Step 1:
[1475] The terminal registers product data in real time.
[1476] Input: Incoming and sales data.
[1477] Output: Registered data.
[1478] Specific operation: The terminal updates the database with arrival and sales information in real time.
[1479] Step 2:
[1480] The device sends the registration data to the server.
[1481] Input: Registration data.
[1482] Output: Data sent to the server.
[1483] Specific action: The device transfers the registration data to the server.
[1484] Step 3:
[1485] The server calculates the current inventory level based on incoming and sales data.
[1486] Input: Registration data.
[1487] Output: Current inventory level.
[1488] Data processing / data calculation: Calculate inventory levels based on incoming and outgoing data.
[1489] Specific operation: The server analyzes incoming and outgoing data and calculates the inventory level.
[1490] Step 4:
[1491] When the server's inventory level falls below a set threshold, it generates feedback. Sentiment data is also used as needed.
[1492] Input: Inventory quantity.
[1493] Output: Feedback.
[1494] Data processing / data calculation: Generate feedback when inventory levels fall below a threshold.
[1495] Specific actions: Generate feedback and determine whether an order is necessary.
[1496] Step 5:
[1497] The server generates an order plan based on the feedback.
[1498] Input: Feedback.
[1499] Output: Ordering plan.
[1500] Data processing / data calculation: Create order plans and make necessary adjustments.
[1501] Specific operation: The server generates the order plan.
[1502] Step 6:
[1503] The server sends the generated order plan to the terminal.
[1504] Input: Order plan.
[1505] Output: Data sent to the terminal.
[1506] Specific action: The server sends the order plan to the terminal.
[1507] Step 7:
[1508] The device sends notifications to staff via a robot. It can also adjust the tone and expression based on emotional data.
[1509] Input: Order planning data.
[1510] Output: Notification to staff.
[1511] Specific operation: The terminal passes the order plan to the robot, and the robot notifies the staff.
[1512] Step 8:
[1513] The staff member who receives the user notification will then proceed with the ordering process.
[1514] Input: Notification data.
[1515] Output: Order completed.
[1516] Specific action: The staff member who receives the notification will carry out the necessary ordering tasks.
[1517] (Application Example 2)
[1518] Next, we will explain application example 2. In the following explanation, the data processing device 12 will be referred to as the "server," and the smart glasses 214 will be referred to as the "terminal."
[1519] Traditional store operations have been cumbersome, involving tasks such as customer service, sales, operation guidance, cash register management, and inventory management, making efficient operation difficult. Furthermore, it has been challenging to respond flexibly to user emotions and needs, leading to a demand for improved customer satisfaction. This invention aims to solve these problems by providing a system that recognizes user emotions and provides optimal responses, thereby improving the efficiency of store operations and enhancing customer satisfaction.
[1520] In Application Example 2, the identification processing by the identification processing unit 290 of the data processing device 12 is realized by the following means. In this invention, the server includes means for converting the user's voice data into text data, means for analyzing the text data to determine purchase intent and operation requests, means for generating optimal suggestions or operation procedures based on the determination results and the user's emotional data, means for providing the generated suggestions or operation procedures to the user by speech synthesis, means for identifying the user's mobile terminal by image recognition and obtaining an operation manual corresponding to the model of the mobile terminal from a database, means for generating operation procedures according to the user's requests from the obtained operation manual and adjusting the generated operation procedures based on the user's emotional data, means for collecting sales data from a cash register terminal in real time and analyzing the sales data to calculate the amount of cash inventory, means for generating and notifying an alert when the calculated amount of cash inventory falls below an appropriate amount, and means for adjusting the generated alert notification based on the user's emotional data. This makes it possible to provide more personalized responses by reflecting the user's emotional data in the overall system adjustments, suggestions, and notifications.
[1521] "Audio data" refers to the audio signals spoken by the user.
[1522] "Text data" refers to digital information obtained by converting audio data into a string of characters.
[1523] "Emotional data" refers to data that indicates the emotional state of a user, as recognized from their voice and facial expressions.
[1524] "Purchase intent" refers to a user's intention to purchase a particular product or service.
[1525] An "operation request" refers to a user's desire to learn about a specific operation or usage method.
[1526] A "proposal" refers to an introduction or advice regarding specific products or services provided in response to user requests.
[1527] "Operating procedures" refer to specific methods of operation provided in response to user requests.
[1528] "Speech synthesis" is a technology that converts text data into audio signals for playback.
[1529] "Image recognition" is a technology that recognizes specific objects from image data acquired using input devices such as cameras.
[1530] A "database" is a collection of information stored in digital format, making it easy to search and retrieve.
[1531] "Sales data" refers to digital information about sales collected through point-of-sale terminals.
[1532] "Cash inventory" refers to the total amount of cash present in a store.
[1533] An "alert" is a warning signal used to notify users of important changes or anomalies in the situation.
[1534] This invention is a system that streamlines various tasks that users experience in stores, such as customer service, sales, operation guidance, cash register management, and inventory management, and further recognizes user emotions to provide optimal responses. The following describes the specific forms for implementing the system.
[1535] 1. Customer service and sales
[1536] User: The user speaks to the robot and asks questions or requests information about products they want to purchase. For example, "I want to buy this smartphone, which plan do you recommend?"
[1537] Terminal: The robot's microphone acquires voice data, and a speech recognition system (e.g., Python's speech_recognition library) converts the voice data into text data. It also uses an emotion engine (e.g., EmotionEngine) to recognize emotions from the user's voice.
[1538] Server: Analyzes text and sentiment data to determine the user's purchase intent and emotional state. For example, it recognizes the user's needs regarding smartphone models and plans, and determines whether the user is excited or calm.
[1539] Server: Retrieves relevant information (product data and plan information) from a database (e.g., a relational database management system) and adjusts the suggested content based on sentiment data.
[1540] Terminal: The adjusted suggestions are converted into speech by a speech synthesis system (e.g., TextToSpeech module) and provided to the user by the robot.
[1541] For example, if a user asks, "I want to buy this smartphone, which plan do you recommend?", the robot will respond to the excited user, "The unlimited data plan is the best!"
[1542] 2. Instructions for Use
[1543] User: The user asks the robot questions about how to use their smartphone. For example, they might say, "I don't know how to send a photo, could you please show me how?"
[1544] Terminal: The voice recognition system converts voice data into text data, and the emotion engine recognizes the user's emotions. It also uses the camera function to recognize the model of the user's smartphone.
[1545] Server: Based on image recognition results, text data, and emotion data, retrieves the corresponding model's operation manual from the database.
[1546] Server: Generates specific operating procedures based on user requests from the operation manual, and adjusts the difficulty and tone of the explanations considering sentiment data.
[1547] Terminal: The robot converts the adjusted operating procedures into speech using speech synthesis and explains them to the user.
[1548] For example, if a user asks, "I don't know how to send a photo, could you tell me how?", the robot calmly guides them by saying, "Select the photo you took, press the share button, and then choose the recipient."
[1549] 3. Cash register management
[1550] Terminal: The POS terminal collects daily sales data.
[1551] Server: Analyzes collected sales data and calculates cash inventory levels. A relational database management system is used.
[1552] Server: Generates an alert when cash inventory falls below the appropriate level and makes adjustments based on user sentiment data.
[1553] Terminal: Sends tailored alert notifications to store managers in real time. Notifications use a tone based on sentiment data.
[1554] 4. Inventory Management
[1555] Terminal: Registers product arrival and sales data in real time.
[1556] Server: Calculates the current inventory level based on registered data and generates feedback if the inventory level falls below a threshold. A relational database management system is used.
[1557] Server: Generates order plans based on feedback and adjusts them considering sentiment data.
[1558] Terminal: Notifies staff of the adjusted order plan via robot.
[1559] Examples of prompts for generative AI models
[1560] "It converts user questions into text and recognizes their emotions. Example: 'I want to buy this smartphone, which plan do you recommend?' Emotion: Excitement. Then, it retrieves relevant plan information from the product information database and generates the best recommendation based on the emotion."
[1561] This allows for more personalized responses by incorporating user sentiment data into overall system adjustments, suggestions, and notifications.
[1562] The flow of a specific process in Application Example 2 will be explained using Figure 14.
[1563] Step 1:
[1564] User: Speaks to the robot about the products they want to buy or any questions they have. The voice data from this interaction becomes the input.
[1565] Step 2:
[1566] Terminal: Audio data acquired via the microphone is converted into text data using a speech recognition system (such as Python's speech_recognition library). The converted text data is obtained as output.
[1567] Step 3:
[1568] Terminal: Uses an emotion engine (such as EmotionEngine) to recognize emotion data from the user's voice data. The recognized emotion data is obtained as output.
[1569] Step 4:
[1570] Terminal: Sends text data and sentiment data obtained through speech recognition to the server. Input data consists of text data and sentiment data, which are sent to the server.
[1571] Step 5:
[1572] Server: Analyzes received text and sentiment data to determine the user's purchase intent and action requests. The results obtained (e.g., purchase intent and action requests) become the output of the analysis.
[1573] Step 6:
[1574] Server: Based on relevant information (such as product data and plan information) obtained from the database (relational database management system), the server generates optimal suggestions based on the user's sentiment data. The generated suggestions are then output.
[1575] Step 7:
[1576] Server: Sends the generated proposal content to the terminal. The input data is the proposal content, which is then sent to the terminal.
[1577] Step 8:
[1578] Terminal: The received suggestion content is converted into speech using a speech synthesis system (such as the TextToSpeech module). Audio data is obtained as output.
[1579] Step 9:
[1580] Terminal: The robot plays voice data and provides suggestions to the user. The user reviews the suggestions and decides on their next action.
[1581] The specific processing unit 290 transmits the result of the specific processing to the smart glasses 214. In the smart glasses 214, the control unit 46A causes the speaker 240 to output the result of the specific processing. The microphone 238 acquires audio indicating user input for the result of the specific processing. The control unit 46A transmits the audio data indicating user input acquired by the microphone 238 to the data processing unit 12. In the data processing unit 12, the specific processing unit 290 acquires the audio data.
[1582] Data generation model 58 is a type of so-called generative AI (Artificial Intelligence). One example of data generation model 58 is ChatGPT (Internet search<URL: https: / / openai.com / blog / chatgpt> ), Gemini (Internet search) <url: https: gemini.google.com ?hl="ja">Examples of generative AI include the following. The data generation model 58 is obtained by performing deep learning on a neural network. The data generation model 58 is input with prompts containing instructions, and with inference data such as audio data representing speech, text data representing text, and image data representing images. The data generation model 58 infers from the input inference data according to the instructions indicated by the prompts, and outputs the inference results in data formats such as audio data and text data. Here, inference refers to, for example, analysis, classification, prediction, and / or summarization.
[1583] In the above embodiment, an example was given in which specific processing is performed by the data processing device 12, but the technology of this disclosure is not limited thereto, and the specific processing may also be performed by the smart glasses 214.
[1584] [Third Embodiment]
[1585] Figure 5 shows an example of the configuration of the data processing system 310 according to the third embodiment.
[1586] As shown in Figure 5, the data processing system 310 includes a data processing device 12 and a headset terminal 314. An example of the data processing device 12 is a server.
[1587] The data processing device 12 comprises a computer 22, a database 24, and a communication interface 26. The computer 22 is an example of a "computer" related to the technology of this disclosure. The computer 22 comprises a processor 28, RAM 30, and storage 32. The processor 28, RAM 30, and storage 32 are connected to a bus 34. The database 24 and the communication interface 26 are also connected to the bus 34. The communication interface 26 is connected to a network 54. An example of the network 54 is a WAN (Wide Area Network) and / or a LAN (Local Area Network).
[1588] The headset terminal 314 includes a computer 36, a microphone 238, a speaker 240, a camera 42, a communication interface 44, and a display 343. The computer 36 includes a processor 46, RAM 48, and storage 50. The processor 46, RAM 48, and storage 50 are connected to a bus 52. The microphone 238, speaker 240, camera 42, and display 343 are also connected to the bus 52.
[1589] The microphone 238 receives voice signals from the user 20 and receives instructions from the user 20. The microphone 238 captures the voice signals from the user 20, converts the captured voice into audio data, and outputs it to the processor 46. The speaker 240 outputs audio according to the instructions from the processor 46.
[1590] Camera 42 is a small digital camera equipped with an optical system including a lens, aperture, and shutter, and an image sensor such as a CMOS (Complementary Metal-Oxide-Semiconductor) image sensor or a CCD (Charge Coupled Device) image sensor, and captures images of the area around the user 20 (for example, an imaging range defined by a field of view equivalent to the width of a typical healthy person's field of vision).
[1591] Communication interface 44 is connected to network 54. Communication interfaces 44 and 26 are responsible for the exchange of various information between processor 46 and processor 28 via network 54. The exchange of various information between processor 46 and processor 28 using communication interfaces 44 and 26 is performed in a secure manner.
[1592] Figure 6 shows an example of the main functions of the data processing device 12 and the headset terminal 314. As shown in Figure 6, the data processing device 12 performs specific processing using the processor 28. The storage 32 stores the specific processing program 56.
[1593] The specific processing program 56 is an example of a "program" relating to the technology of this disclosure. The processor 28 reads the specific processing program 56 from the storage 32 and executes the read specific processing program 56 on the RAM 30. The specific processing is realized by the processor 28 operating as a specific processing unit 290 in accordance with the specific processing program 56 executed on the RAM 30.
[1594] The storage 32 stores the data generation model 58 and the emotion identification model 59. The data generation model 58 and the emotion identification model 59 are used by the identification processing unit 290.
[1595] In the headset terminal 314, the processor 46 performs the reception output processing. The storage 50 stores the reception output program 60. The processor 46 reads the reception output program 60 from the storage 50 and executes the read reception output program 60 on the RAM 48. The reception output processing is realized by the processor 46 operating as a control unit 46A according to the reception output program 60 executed on the RAM 48.
[1596] Next, the specific processing performed by the specific processing unit 290 of the data processing device 12 will be described. In the following description, the data processing device 12 will be referred to as the "server" and the headset terminal 314 will be referred to as the "terminal".
[1597] The system of the present invention is configured to streamline customer service, sales, operation guidance, cash register management, and inventory management tasks that users experience in stores. The following describes in detail the embodiments for implementing the system of the present invention.
[1598] Customer service and sales
[1599] User: The user speaks to the robot to ask questions or provide information about products they want to purchase. For example, "I want to buy this smartphone, which plan do you recommend?"
[1600] Terminal: The speech recognition system converts the user's voice data into text data. The text data is then sent to the server.
[1601] Server: Analyzes text data to determine the user's purchase intent. For example, it recognizes needs regarding smartphone models and plans.
[1602] Server: Based on purchase intent, the system retrieves information from the database to suggest the optimal plan and generates a recommended plan. The recommended plan includes details such as price, data capacity, and benefits.
[1603] Terminal: The recommended plan information is converted into speech using a speech synthesis system, and the robot communicates the proposal to the user.
[1604] User: Review the proposal, ask further questions, or decide to purchase.
[1605] Instructions
[1606] User: The user asks the robot a question about a specific smartphone operation. For example, they might say, "I don't know how to send a photo, could you please tell me how?"
[1607] Terminal: The voice recognition system converts the user's voice data into text data. It also uses the camera function to recognize the model of the user's smartphone.
[1608] Server: Based on the image recognition results and text data, retrieves the operation manual for the corresponding model from the database.
[1609] Server: Generates specific operating procedures based on user requests from the operation manual.
[1610] Terminal: The generated operating procedure is converted into speech using a speech synthesis system, and the robot explains it to the user.
[1611] User: Operate the smartphone following the instructions provided.
[1612] Cash register management
[1613] Terminal: The POS terminal collects daily sales data and sends it to the server.
[1614] Server: Analyzes sales data and calculates the current cash inventory level.
[1615] Server: Generates an alert and sends a notification if the cash inventory falls below the appropriate level.
[1616] Terminal: Sends alert notifications to store administrators in real time.
[1617] User: The store manager who receives the notification will take action to address any cash discrepancies.
[1618] Inventory Management
[1619] Terminal: Registers product arrival and sales data in real time and sends it to the server.
[1620] Server: Calculates current inventory levels based on incoming and outgoing data.
[1621] Server: When inventory levels fall below a set threshold, it generates feedback and creates an ordering plan.
[1622] Terminal: Sends notifications to staff via robot to confirm the generated order plan.
[1623] User: Staff members who receive the notification will proceed with placing the order.
[1624] The system of this invention streamlines store operations and improves the quality of customer service. In particular, it enables users to receive quick and accurate information at stores, leading to high customer satisfaction. Furthermore, the automation of cash register management and inventory management solves the problem of labor shortages.
[1625] The following describes the processing flow.
[1626] Customer service and sales
[1627] Step 1:
[1628] Users can ask the robot questions or request information about products they want to buy. For example, they might ask, "I want to buy this smartphone, which plan do you recommend?"
[1629] Step 2:
[1630] The device uses a speech recognition system to convert the user's voice data into text data.
[1631] Step 3:
[1632] The terminal sends the converted text data to the server.
[1633] Step 4:
[1634] The server analyzes text data to determine the user's purchase intent. For example, it recognizes their needs regarding smartphone models and plans.
[1635] Step 5:
[1636] The server retrieves relevant information (product data and plan information) from the database.
[1637] Step 6:
[1638] The server generates the optimal plan based on the user's needs and creates text data containing detailed information about the recommended plan.
[1639] Step 7:
[1640] The text data generated by the server is returned to the terminal and converted into speech using a speech synthesis system.
[1641] Step 8:
[1642] The terminal, via a robot, provides the user with generated suggestions via voice.
[1643] Step 9:
[1644] The user reviews the proposal and decides on their next action (e.g., continue asking questions, make a purchase).
[1645] Instructions
[1646] Step 1:
[1647] The user asks the robot questions about specific smartphone operations. For example, they might say, "I don't know how to send a photo, could you please tell me how?"
[1648] Step 2:
[1649] The device uses a speech recognition system to convert the user's voice data into text data.
[1650] Step 3:
[1651] The device uses the smartphone's camera to perform image recognition to determine the model of the user's mobile device.
[1652] Step 4:
[1653] The terminal sends the converted text data and image recognition results to the server.
[1654] Step 5:
[1655] Based on the image recognition results, the server retrieves the corresponding model's operation manual from the database.
[1656] Step 6:
[1657] The server generates specific operating procedures based on user requests from the acquired operation manual.
[1658] Step 7:
[1659] The server returns the generated operating procedure as text data to the terminal, and the speech synthesis system converts it into speech.
[1660] Step 8:
[1661] The terminal provides the user with voice instructions on how to operate it via a robot.
[1662] Step 9:
[1663] The user operates the smartphone according to the instructions provided.
[1664] Cash register management
[1665] Step 1:
[1666] The terminal collects daily sales data from the POS terminal.
[1667] Step 2:
[1668] The terminal sends the collected sales data to the server.
[1669] Step 3:
[1670] The server analyzes sales data and calculates the current cash inventory level.
[1671] Step 4:
[1672] Based on the cash inventory amount calculated by the server, the appropriate inventory level is calculated.
[1673] Step 5:
[1674] The server compares the appropriate inventory level with the actual inventory level and generates an alert if there is a surplus or shortage.
[1675] Step 6:
[1676] The server sends the generated alert to the terminal.
[1677] Step 7:
[1678] The device sends alert notifications to the store administrator in real time.
[1679] Step 8:
[1680] The store manager who receives the user notification will then take action to address any cash discrepancies.
[1681] Inventory Management
[1682] Step 1:
[1683] The terminal registers product arrival and sales data in real time.
[1684] Step 2:
[1685] The device sends the registered data to the server.
[1686] Step 3:
[1687] The server calculates the current inventory level based on incoming and outgoing data.
[1688] Step 4:
[1689] The server generates feedback when the inventory level falls below a set threshold.
[1690] Step 5:
[1691] The server generates an order plan based on the feedback.
[1692] Step 6:
[1693] The server sends the generated order plan to the terminal.
[1694] Step 7:
[1695] The device sends notifications to staff via a robot.
[1696] Step 8:
[1697] The staff member who receives the user notification will then proceed with the ordering process.
[1698] Thus, the system of the present invention achieves efficient and high-quality store operations through each step.
[1699] (Example 1)
[1700] Next, we will describe Example 1. In the following description, the data processing device 12 will be referred to as the "server," and the headset-type terminal 314 will be referred to as the "terminal."
[1701] In traditional retail operations, a wide range of tasks, including customer service, inventory management, and cash register management, are performed manually. This is not only inefficient but also prone to human error and wasted time. As a result, problems such as decreased customer satisfaction, reduced operational efficiency, and increased burden due to labor shortages arise. A system is needed to solve these problems and provide timely and accurate information while improving operational efficiency.
[1702] The identification process performed by the identification processing unit 290 of the data processing device 12 in Example 1 is realized by the following means.
[1703] In this invention, the server includes means for converting user voice data into text data, means for analyzing the text data to determine purchase intent and operation requests, means for generating optimal suggestions or operation procedures based on the determination results, means for providing the generated suggestions or operation procedures to the user by speech synthesis, means for registering product arrival and sales data in real time and analyzing the data to calculate the current inventory level, means for formulating an order plan and notifying the user when the calculated inventory level falls below a set threshold, means for collecting sales data from a cash register terminal and analyzing the sales data to calculate the cash inventory level, means for generating an alert and notifying the user when the calculated cash inventory level falls below an appropriate level, and means for streamlining customer service, sales, operation guidance, cash register management, and inventory management that the user receives in the store. This enables automation and efficiency of store operations, improving the quality of user service and operational efficiency.
[1704] "Audio data" refers to sound information acquired digitally via a microphone or other audio input device, such as user speech or instructions.
[1705] "Text data" is a data format that converts audio data into written text.
[1706] "Analysis" is the process of understanding the meaning and intent of input data, and often involves using technologies such as natural language processing.
[1707] "Purchase intent" refers to the user's wishes and desires regarding the type and conditions of the product they are trying to acquire.
[1708] An "operation request" refers to a request from a user for assistance in performing a specific operation.
[1709] "Speech synthesis" is a technology that converts text data into speech that sounds like human speech.
[1710] "Product arrival data" refers to data that records information about when new products arrive at a store.
[1711] "Sales data" refers to data that records information about when a product was sold in a store.
[1712] "Inventory level" refers to the quantity of goods currently stored in the store or warehouse.
[1713] An "ordering plan" refers to a plan for purchasing new goods when inventory falls below a certain level.
[1714] A "cash register terminal" is an electronic device used for selling goods and processing payments.
[1715] "Sales data" refers to data that records information such as the quantity, price, and date and time of the transaction of the goods sold.
[1716] "Cash inventory" refers to the total amount of cash stored in cash registers and within the store.
[1717] An "alert" refers to a warning or cautionary message that is sent when certain conditions are met.
[1718] "Customer service" refers to the work of providing product descriptions, guidance, and answering questions from customers.
[1719] "Operation guidance" refers to the service of providing customers with instructions on how to use the equipment and services they use.
[1720] "Cash register management" refers to the task of managing daily sales and cash inflows and outflows.
[1721] "Inventory management" refers to the overall management of receiving, storing, and shipping goods in order to maintain an appropriate level of inventory.
[1722] The system of the present invention aims to streamline store operations and improve customer satisfaction. This system is configured to support the various tasks that users perform in stores, including customer service and sales, operation guidance, cash register management, and inventory management. The following describes specific embodiments for carrying out the present invention.
[1723] Customer service and sales
[1724] User: The user speaks to the robot to ask questions or provide information about products they want to purchase. For example, they might say, "I want to buy this smartphone, which plan do you recommend?"
[1725] Terminal: Uses a speech recognition system (e.g., a cloud-based speech recognition service) to convert the user's voice data into text data. The converted text data is sent to a server via the internet.
[1726] Server: Uses a Natural Language Processing (NLP) engine (e.g., a cloud-based natural language processing service) to analyze text data and determine the user's purchase intent. This purchase intent includes needs regarding smartphone models and plans.
[1727] Server: Based on the determination result, the server retrieves information on the corresponding plan from the database (e.g., relational database management system) and recommends the optimal plan. The recommended plan includes pricing, data capacity, and benefits.
[1728] Terminal: A speech synthesis system (e.g., speech synthesis API) converts the recommended plan information into speech, and a robot communicates the proposal to the user.
[1729] User: Review the proposal, ask further questions, or decide to purchase.
[1730] Specific example:
[1731] User: "I want to buy this smartphone, which plan would you recommend?"
[1732] Robot: "The best plan for you is the one with X GB of data for Y yen per month. This plan comes with the following benefits."
[1733] Examples of prompts to input into a generative AI model:
[1734] "Please generate a database query to analyze the user's purchase intent for the specified product and propose the optimal plan."
[1735] Instructions
[1736] User: The user asks the robot a question about a specific smartphone operation. For example, they might say, "I don't know how to send a photo, could you please tell me how?"
[1737] Terminal: A voice recognition system (e.g., a cloud-based voice recognition service) converts the user's voice data into text data. It also uses the camera function to recognize the model of the user's smartphone.
[1738] Server: Based on image recognition results (e.g., an image recognition model using deep learning) and text data, retrieves the operation manual for the corresponding model from the database.
[1739] Server: Generates specific operating procedures based on user requests from the operation manual.
[1740] Terminal: The generated operating procedure is converted into speech using a speech synthesis system (e.g., speech synthesis API), and the robot explains it to the user.
[1741] User: Operate the smartphone following the instructions provided.
[1742] Specific example:
[1743] User: "I don't know how to send photos, could you please tell me how?"
[1744] Robot: "Your device is a [model name]. To send a photo, first open your gallery, then press the share button and select the recipient."
[1745] Examples of prompts to input into a generative AI model:
[1746] "Generate specific steps to explain how to operate the smartphone as instructed by the user."
[1747] Cash register management
[1748] Terminal: A point-of-sale terminal (e.g., a point-of-sale management system) collects daily sales data and transmits it to a server via the internet.
[1749] Server: Aggregates and analyzes sales data to calculate the current cash inventory level.
[1750] Server: Generates an alert if the cash inventory falls below the appropriate level.
[1751] Terminal: Sends alert notifications to store administrators in real time.
[1752] User: The store manager who receives the notification will take action to address any cash discrepancies.
[1753] Specific example:
[1754] Server: "Cash inventory is insufficient. Please take immediate action."
[1755] Manager: "We'll replenish the cash immediately."
[1756] Examples of prompts to input into a generative AI model:
[1757] "Calculate cash inventory levels based on daily sales data and generate alerts when inventory levels are low."
[1758] Inventory Management
[1759] Terminal: Product arrival and sales data (e.g., sales management system) are registered in real time and transmitted to the server via the internet.
[1760] Server: Calculates current inventory levels based on incoming and outgoing data.
[1761] Server: When inventory levels fall below a set threshold, it generates feedback and develops an ordering plan.
[1762] Terminal: Sends notifications to staff via robot to confirm the generated order plan.
[1763] User: Staff members who receive the notification will proceed with placing the order.
[1764] Specific example:
[1765] Server: "Inventory levels have fallen below the threshold. An order is required."
[1766] Staff: "We have confirmed your order. We have placed the necessary items."
[1767] Examples of prompts to input into a generative AI model:
[1768] "Monitor current inventory levels and generate reordering plans when they fall below the set threshold."
[1769] The system of this invention is expected to automate many store operations and improve the quality of customer service. In particular, it will enable users to receive quick and accurate information at stores, resulting in high customer satisfaction. Furthermore, the automation of cash register management and inventory management will solve the problem of labor shortages.
[1770] The flow of the specific processing in Example 1 will be explained using Figure 11.
[1771] Customer service and sales
[1772] Step 1:
[1773] Users can ask the robot questions or request information about products they want to buy. For example, they might say, "I want to buy this smartphone, which plan do you recommend?"
[1774] Input: User's voice
[1775] Output: Audio data
[1776] Step 2:
[1777] The device uses a speech recognition system (e.g., a cloud-based speech recognition service) to convert speech data into text data. The converted text data is then sent to the server.
[1778] Input: Audio data
[1779] Output: Text data
[1780] Step 3:
[1781] The server uses a Natural Language Processing (NLP) engine (e.g., a cloud-based natural language processing service) to analyze text data and determine the user's purchase intent.
[1782] Input: Text data
[1783] Output: Classification result (user's purchase intent)
[1784] Step 4:
[1785] Based on the determination result, the server retrieves information on the corresponding plan from the database (e.g., a relational database management system) and recommends the optimal plan.
[1786] Input: Discrimination result
[1787] Output: Recommended plan (price, data capacity, benefits, etc.)
[1788] Step 5:
[1789] The terminal uses a speech synthesis system (e.g., a speech synthesis API) to convert the recommended plan information into speech, and a robot then communicates the proposal to the user.
[1790] Input: Recommended plan
[1791] Output: Audio data (proposed content)
[1792] Step 6:
[1793] The user reviews the proposal, asks further questions, or decides to make a purchase.
[1794] Input: Audio data (proposal)
[1795] Output: User decisions and questions
[1796] Adding specific actions
[1797] When a user asks, "I want to buy this smartphone, which plan do you recommend?", the robot suggests, "The best plan for you is the one with X GB of data for Y yen per month. This plan comes with the following benefits."
[1798] Instructions
[1799] Step 1:
[1800] The user asks the robot questions about specific smartphone operations. For example, they might say, "I don't know how to send a photo, could you please tell me how?"
[1801] Input: User's voice
[1802] Output: Audio data
[1803] Step 2:
[1804] The device converts the user's voice data into text data using a voice recognition system (e.g., a cloud-based voice recognition service). It also utilizes its camera function to recognize the model of the user's smartphone.
[1805] Input: Audio data and image data
[1806] Output: Text data and image recognition results
[1807] Step 3:
[1808] The server retrieves the operation manual for the corresponding model from the database based on the image recognition results (e.g., an image recognition model using deep learning) and text data.
[1809] Input: Text data and image recognition results
[1810] Output: Operation Manual
[1811] Step 4:
[1812] The server generates specific operating procedures based on the user's request, using the operation manual.
[1813] Input: Operation Manual
[1814] Output: Operating Procedure
[1815] Step 5:
[1816] The terminal converts the generated operating procedures into speech using a speech synthesis system (e.g., a speech synthesis API), and the robot explains them to the user.
[1817] Input: Operating Procedure
[1818] Output: Audio data (operating instructions)
[1819] Step 6:
[1820] The user operates the smartphone according to the instructions provided.
[1821] Input: Audio data (operating instructions)
[1822] Output: Smartphone operation results
[1823] Adding specific actions
[1824] When a user asks, "I don't know how to send photos, could you please tell me how?", the robot responds, "Your device is a XX. To send photos, first open your gallery, then press the share button and select the recipient."
[1825] Cash register management
[1826] Step 1:
[1827] A point-of-sale terminal (e.g., a point-of-sale management system) collects daily sales data and transmits it to a server via the internet.
[1828] Input: Sales data
[1829] Output: Send to server
[1830] Step 2:
[1831] The server aggregates and analyzes sales data to calculate the current cash inventory level.
[1832] Input: Sales data
[1833] Output: Cash Inventory
[1834] Step 3:
[1835] The server generates an alert if the cash inventory falls below the appropriate level.
[1836] Input: Cash inventory
[1837] Output: Alert
[1838] Step 4:
[1839] The device sends alert notifications to the store administrator in real time.
[1840] Input: Alert
[1841] Output: Alert notification
[1842] Step 5:
[1843] The store manager will take action to address any cash discrepancies upon receiving notification.
[1844] Input: Alert notification
[1845] Output: Cash replenishment and cash depletion response
[1846] Adding specific actions
[1847] The server sends an alert saying, "Cash inventory is low. Please take immediate action," and the store manager checks it and responds, "We will replenish the cash immediately."
[1848] Inventory Management
[1849] Step 1:
[1850] The terminal registers product arrival and sales data (e.g., sales management system) in real time and transmits it to the server via the internet.
[1851] Input: Incoming data and sales data
[1852] Output: Send to server
[1853] Step 2:
[1854] The server calculates the current inventory level based on incoming and outgoing data.
[1855] Input: Incoming data and sales data
[1856] Output: Inventory Quantity
[1857] Step 3:
[1858] The server generates feedback and develops an ordering plan when inventory levels fall below a set threshold.
[1859] Input: Inventory Quantity
[1860] Output: Feedback and ordering plan
[1861] Step 4:
[1862] The terminal sends a notification to staff via a robot for them to review the generated order plan.
[1863] Input: Ordering plan
[1864] Output: Notification
[1865] Step 5:
[1866] After receiving notification, the staff will proceed with placing the order.
[1867] Input: Notification
[1868] Output: Order result
[1869] Adding specific actions
[1870] The server notifies, "Inventory levels have fallen below the threshold. An order is required," and the staff responds, "We have confirmed the order. We have placed an order for the necessary items."
[1871] This invention is expected to automate many store operations and improve the quality of customer service. In particular, it will enable users to receive quick and accurate information at stores, resulting in high customer satisfaction. Furthermore, the automation of cash register management and inventory management will help solve the problem of labor shortages.
[1872] (Application Example 1)
[1873] Next, we will explain Application Example 1. In the following explanation, the data processing device 12 will be referred to as the "server," and the headset-type terminal 314 will be referred to as the "terminal."
[1874] In recent years, brick-and-mortar stores have been required to provide customers with the information they need quickly and accurately. However, variations in the number and knowledge levels of store staff can lead to delays in customer service and inability to provide appropriate information. Furthermore, internal operations such as inventory management and cash register management are also reliant on human labor, which can lead to errors and decreased efficiency. It is necessary to solve these problems and improve customer satisfaction while also increasing the efficiency of store operations.
[1875] The specific processing performed by the specific processing unit 290 of the data processing device 12 in Application Example 1 is realized by the following means.
[1876] In this invention, the server includes means for converting user voice data into text data, means for analyzing the text data to determine purchase intent and operation requests, means for generating optimal suggestions or operation procedures based on the determination results, means for providing the generated suggestions or operation procedures to the user by speech synthesis, means for recognizing questions about the user's smart device using voice and camera and providing information based on those questions, and means for acquiring product information from a database in real time and providing that information to the user by a speech synthesis system. This enables the rapid and accurate provision of information that customers seek, improving the efficiency of store operations and enhancing customer satisfaction.
[1877] "User voice data" refers to the voice information that the user speaks.
[1878] "Text data" refers to audio data converted into written information.
[1879] "Purchase intent" refers to a user's desire to buy a particular product or service.
[1880] An "operation request" is a user's desire to know a specific operation method or procedure.
[1881] A "proposal" refers to the optimal plan or options provided based on the user's needs.
[1882] "Operating procedures" refer to the specific steps or steps required for a user to perform a particular operation.
[1883] A "speech synthesis system" is a system that converts text data into speech and conveys information to the user audibly.
[1884] A "smart device" refers to a mobile device with advanced functions, such as a smartphone or tablet.
[1885] A "database" is an information management system that stores and allows searching for product information, operation manuals, and other data.
[1886] "Real-time" refers to a state where the current situation and data can be reflected immediately.
[1887] The system of this invention aims to improve the efficiency of customer service and internal operations in physical stores. This system utilizes speech recognition, text analysis, image recognition, and speech synthesis to provide users with optimal suggestions and operating procedures. The specific configuration and operation of the system are described below.
[1888] First, the user asks a question to a smart robot in the store. For example, they might say, "I want to buy this smartphone, which plan do you recommend?" The smart robot collects the voice data using its microphone. The speech recognition system installed in the device (for example, Google Speech Recognition API) converts the voice data into text data. This text data is then sent to a server.
[1889] The server analyzes text data to determine the user's purchase intent and operational requests. For example, if a user asks about a smartphone plan, the server recognizes the user's needs from the content of the question. Next, the server retrieves relevant information from the database to generate the best possible suggestions based on this recognition. For example, it generates recommended smartphone plans and attaches detailed information such as price, data capacity, and benefits.
[1890] The generated suggestions are converted into speech using a speech synthesis system (e.g., Pyttsx3). The smart robot communicates the suggestions to the user through the converted speech. The user can then review them, ask further questions, or make a purchase decision.
[1891] The same applies when a user asks a question about operating a specific smartphone. For example, it can respond to prompts such as, "I don't know how to send photos, please tell me how." In addition to voice, the terminal uses its camera function to recognize the model of the user's smartphone. Based on the image recognition results and text data, the server retrieves the operation manual for that model from its database. Based on this operation manual, it generates specific operating procedures that meet the user's request and explains them to the user using speech synthesis.
[1892] Furthermore, this system contributes to the efficiency of internal operations. For example, it collects daily sales data from POS terminals and sends it to a server to calculate the cash inventory level. If the cash inventory level falls below the appropriate level, an alert is generated and a notification is sent. This notification reaches the store manager in real time, allowing for quick action to address any cash shortages or surpluses. In addition, for inventory management, it registers product arrival and sales data in real time and calculates the current inventory level. If the inventory level falls below a set threshold, it automatically generates feedback and creates an ordering plan.
[1893] This system enables the rapid and accurate provision of information that customers need, leading to increased efficiency in store operations and improved customer satisfaction.
[1894] For example:
[1895] "I'd like to buy this smartphone. Which plan would you recommend?"
[1896] "Do you have this item in stock?"
[1897] "How do I send a photo?"
[1898] These prompt messages will be recognized by the smart robot, enabling it to provide appropriate information and operating guidance.
[1899] The flow of a specific process in Application Example 1 will be explained using Figure 12.
[1900] Step 1:
[1901] The user asks a question to the smart robot. For example, they might say a prompt like, "I want to buy this smartphone, which plan do you recommend?" This voice data is then input into the robot.
[1902] Step 2:
[1903] The device uses a speech recognition system (Google Speech Recognition API) to convert the user's voice data into text data. This process converts the voice data into text data. This text data is output and sent to the server.
[1904] Step 3:
[1905] The server analyzes the input text data. Specifically, it identifies purchase intent and operational requests from the text data and recognizes the corresponding questions and intentions. This discrimination process yields the discrimination result.
[1906] Step 4:
[1907] The server generates optimal suggestions or operating procedures based on the determination results. For example, if there is a question about a smartphone plan, the server retrieves the plan information from the database and generates a recommended plan. This data processing generates suggestions and operating procedures. These generated results are output.
[1908] Step 5:
[1909] The server converts the generated suggestions or operating procedures into speech using a speech synthesis system (Pyttsx3). This process converts text data into speech data.
[1910] Step 6:
[1911] The terminal provides the user with synthesized speech suggestions. These suggestions are transmitted to the user as voice from the smart robot. Through this process, the user confirms the information provided.
[1912] Step 7:
[1913] The user asks additional questions or makes a purchase or action decision based on the information provided. Depending on the user's actions, the process either returns to step 1 or ends.
[1914] Step 8:
[1915] When a user asks a question about a specific smartphone operation (for example, "I don't know how to send a photo, can you tell me how?"), the device uses its camera function to perform image recognition on the user's smart device. This input allows the device's model information to be obtained.
[1916] Step 9:
[1917] The server retrieves the corresponding operation manual from the database based on the image recognition results and text data. This data processing identifies the appropriate operation manual. This information is then output.
[1918] Step 10:
[1919] The server generates specific operating procedures based on the user's request, using the operation manual. For example, it generates procedures for sending a photograph. This processing yields specific operating procedures. These generated results are then output.
[1920] Step 11:
[1921] The server converts the generated operating instructions into speech using a speech synthesis system and provides them to the user. This allows the user to operate their smart device by following the voice guidance.
[1922] Furthermore, an emotion engine that estimates the user's emotions may be incorporated. That is, the identification processing unit 290 may use the emotion identification model 59 to estimate the user's emotions and perform identification processing using the user's emotions.
[1923] The system of the present invention aims to streamline customer service, sales, operation guidance, cash register management, and inventory management tasks that users experience in stores, and further to recognize user emotions and provide optimal responses. The following describes in detail the embodiments for implementing the system of the present invention.
[1924] Customer service and sales
[1925] User: The user speaks to the robot to ask questions or provide information about products they want to purchase. For example, "I want to buy this smartphone, which plan do you recommend?"
[1926] Terminal: The speech recognition system converts the user's voice data into text data. Furthermore, an emotion engine is used to recognize emotions from the user's voice.
[1927] Terminal: Sends converted text data and sentiment data to the server.
[1928] Server: Analyzes text and sentiment data to determine the user's purchase intent and emotional state. For example, it recognizes the user's needs regarding smartphone models and plans, and determines whether the user is excited or calm.
[1929] Server: Retrieves relevant information (product data and plan information) from the database. Adjusts suggestions based on sentiment data.
[1930] Server: Sends the adjusted suggestions to the terminal. The speech synthesis system converts them into speech using appropriate tone and expression.
[1931] Terminal: The robot provides the user with generated suggestions via voice.
[1932] User: Review the proposal, ask further questions, or decide to purchase.
[1933] Instructions
[1934] User: The user asks the robot a question about a specific smartphone operation. For example, they might say, "I don't know how to send a photo, could you please tell me how?"
[1935] Terminal: The speech recognition system converts the user's voice data into text data. Simultaneously, the emotion engine recognizes the user's emotions.
[1936] Terminal: Sends converted text data and sentiment data to the server. It also uses the camera function to recognize the model of the user's smartphone.
[1937] Server: Based on image recognition results, text data, and emotion data, retrieves the corresponding model's operation manual from the database.
[1938] Server: Generates specific operating procedures based on user requests from the operation manual, and further adjusts the difficulty and tone of the explanations by considering sentiment data.
[1939] Server: Returns the adjusted operating procedure to the terminal and converts it into speech using the speech synthesis system.
[1940] Terminal: A robot provides voice instructions to the user. Tones and expressions are based on emotional data.
[1941] User: Operate the smartphone following the instructions provided.
[1942] Cash register management
[1943] Terminal: Collects daily sales data from the POS terminal.
[1944] Terminal: Sends collected sales data to the server.
[1945] Server: Analyzes sales data and calculates the current cash inventory level.
[1946] Server: Generates an alert and sends a notification if the cash inventory falls below the appropriate level.
[1947] Terminal: Sends alert notifications to store administrators in real time.
[1948] User: The store manager who receives the notification will take action to address any cash discrepancies.
[1949] Inventory Management
[1950] Terminal: Registers product arrival and sales data in real time.
[1951] Terminal: Sends registered data to the server.
[1952] Server: Calculates current inventory levels based on incoming and outgoing data.
[1953] Server: Generates feedback when inventory levels fall below a set threshold. Sentiment data is also used as needed.
[1954] Server: Generates order plans based on feedback.
[1955] Server: Sends the generated order plan to the terminal.
[1956] Terminal: Sends notifications to staff via a robot. It can also adjust tone and expression based on emotional data.
[1957] User: Staff members who receive the notification will proceed with placing the order.
[1958] This system will streamline store operations and provide high-quality service to both employees and customers. Furthermore, because it can recognize user emotions and adjust responses in real time, it is expected to further improve customer satisfaction.
[1959] The following describes the processing flow.
[1960] Customer service and sales
[1961] Step 1:
[1962] Users can ask the robot questions or request information about products they want to buy. For example, they might ask, "I want to buy this smartphone, which plan do you recommend?"
[1963] Step 2:
[1964] The device uses a speech recognition system to convert the user's voice data into text data. Simultaneously, an emotion engine recognizes the user's emotions from their voice.
[1965] Step 3:
[1966] The device sends the converted text data and sentiment data to the server.
[1967] Step 4:
[1968] The server analyzes text data to determine the user's purchase intent. For example, it recognizes their needs regarding smartphone models and plans.
[1969] Step 5:
[1970] The server retrieves relevant information (product data and plan information) from the database.
[1971] Step 6:
[1972] The server generates optimal suggestions while considering the user's emotions. For example, if the user is undecided, it adjusts the suggestions, such as emphasizing special offers.
[1973] Step 7:
[1974] The server generates suggestions and sends them to the terminal, where they are converted into speech using a speech synthesis system. The tone and expression of the converted speech are adjusted according to the user's emotions.
[1975] Step 8:
[1976] The terminal provides the user with generated suggestions via voice through a robot.
[1977] Step 9:
[1978] The user reviews the proposal and decides whether to continue asking questions or to make a purchase.
[1979] Instructions
[1980] Step 1:
[1981] The user asks the robot questions about specific smartphone operations. For example, they might say, "I don't know how to send a photo, could you please tell me how?"
[1982] Step 2:
[1983] The device uses a speech recognition system to convert the user's voice data into text data. Simultaneously, an emotion engine recognizes the user's emotions from their voice.
[1984] Step 3:
[1985] The device sends text and sentiment data to the server, and uses its camera function to perform image recognition to determine the model of the user's smartphone.
[1986] Step 4:
[1987] The server analyzes the image recognition results and text data to determine the user's operation request.
[1988] Step 5:
[1989] The server retrieves the operation manual for the corresponding model from the database.
[1990] Step 6:
[1991] The server generates specific operating procedures based on user requests from the acquired operation manual, and further adjusts the difficulty level and tone of the explanations by taking sentiment data into consideration.
[1992] Step 7:
[1993] The server generates operating instructions and sends them to the terminal as text data, which are then converted into speech using a speech synthesis system. The tone and expression of the speech are adjusted according to the user's emotions.
[1994] Step 8:
[1995] The terminal provides the user with voice instructions on how to operate it via a robot.
[1996] Step 9:
[1997] The user operates the smartphone according to the instructions provided.
[1998] Cash register management
[1999] Step 1:
[2000] The terminal collects daily sales data from the POS terminal.
[2001] Step 2:
[2002] The terminal sends the collected sales data to the server.
[2003] Step 3:
[2004] The server analyzes sales data and calculates the current cash inventory level.
[2005] Step 4:
[2006] Based on the cash inventory amount calculated by the server, the appropriate inventory level is calculated.
[2007] Step 5:
[2008] The server compares the appropriate inventory level with the actual inventory level and generates an alert if there is a surplus or shortage.
[2009] Step 6:
[2010] The server sends the generated alert to the terminal.
[2011] Step 7:
[2012] The device sends alert notifications to the store administrator in real time.
[2013] Step 8:
[2014] The store manager who receives the user notification will then take action to address any cash discrepancies.
[2015] Inventory Management
[2016] Step 1:
[2017] The terminal registers product arrival and sales data in real time.
[2018] Step 2:
[2019] The device sends the registered data to the server.
[2020] Step 3:
[2021] The server calculates the current inventory level based on incoming and outgoing data.
[2022] Step 4:
[2023] The server generates feedback when the inventory level falls below a set threshold.
[2024] Step 5:
[2025] The server generates an order plan based on the feedback.
[2026] Step 6:
[2027] The server sends the generated order plan to the terminal.
[2028] Step 7:
[2029] The device sends notifications to staff via a robot. It can also adjust the tone and expression based on emotional data.
[2030] Step 8:
[2031] The staff member who receives the user notification will then proceed with the ordering process.
[2032] Thus, the system of the present invention achieves efficient and high-quality store operations through each step and has the ability to recognize user emotions and adjust responses in real time. As a result, an improvement in customer satisfaction is expected.
[2033] (Example 2)
[2034] Next, we will describe Example 2. In the following description, the data processing device 12 will be referred to as the "server," and the headset-type terminal 314 will be referred to as the "terminal."
[2035] Existing store management systems have limitations in improving customer satisfaction because they respond to user purchase intentions and operational requests in a uniform and mechanical manner. Furthermore, real-time responses to cash and inventory management are difficult, which can lead to decreased operational efficiency. Additionally, the lack of consideration for user emotions makes it difficult to maximize individual customer satisfaction.
[2036] The identification process performed by the identification processing unit 290 of the data processing device 12 in Example 2 is realized by the following means.
[2037] In this invention, the server includes means for converting user voice data into text data, means for analyzing the text data to determine purchase intent and operation requests, means for generating optimal suggestions or operation procedures based on the determination results, means for providing the generated suggestions or operation procedures to the user by speech synthesis, means for recognizing emotions from the user's voice, means for adjusting the content of suggestions or operation procedures based on the emotion recognition results, and means for acquiring relevant information based on the content of the user's purchase or operation request and emotional state. This enables flexible responses that take into account not only the user's purchase intent but also their emotional state, and is expected to significantly improve customer satisfaction. In addition, real-time cash inventory management and inventory management can be performed efficiently, improving the efficiency of store operations.
[2038] "User voice data" refers to the audio data emitted when a user speaks to the system.
[2039] "Text data" refers to string data obtained by converting audio data.
[2040] "Emotion recognition" refers to the process of analyzing and identifying a user's emotional state from their voice data.
[2041] "Suggested content" refers to appropriate answers and recommendation information that the system generates in response to user questions and requests.
[2042] "Operating procedures" refer to the specific methods of operation that the system provides in response to user questions or requests.
[2043] "Speech synthesis" refers to the technology that converts text data into speech and conveys it to the user in a natural-sounding voice.
[2044] "Purchase intent" refers to a user's intention or purpose when purchasing a particular product or service.
[2045] An "operation request" refers to a request from a user who wants to know a specific operation method or information.
[2046] "Related information" refers to product data and service information that is applicable to the user's questions or requests.
[2047] "Image recognition" refers to the technology that identifies specific objects or models from image data captured by a camera.
[2048] "Sales data" refers to data that records sales information from the store's cash register.
[2049] "Cash inventory" refers to numerical data that indicates the current amount of cash on hand.
[2050] An "alert" refers to a warning system that notifies you when certain conditions occur.
[2051] "Inventory level" refers to data that shows the current number of items in stock at a store.
[2052] "Feedback" refers to re-evaluation and instructional information generated by the system based on the analysis results.
[2053] An "ordering plan" refers to a plan for ordering new goods based on the results of inventory and cash management.
[2054] The system of this invention aims to provide users with information to guide them through purchasing and operating products, and to efficiently manage cash inventory and stock. The system's implementation primarily involves servers, terminals, and users. The specific configuration of the system is as follows:
[2055] System Configuration
[2056] hardware
[2057] Device: A device equipped with voice recognition, emotion recognition, and camera capabilities. Examples include smartphones and robotic devices for specific purposes.
[2058] The server consists of a database server and an analysis server. MySQL is used as the database.
[2059] Database: Storage for product data, operation manuals, sales data, and inventory data.
[2060] software
[2061] Speech recognition system: Uses Google Speech-to-Text, etc. This converts the user's speech into text data.
[2062] Emotion Engine: Uses Emotion AI, which recognizes the user's emotions from their voice.
[2063] Speech synthesis system: Uses a text-to-speech engine such as Amazon Polly. This converts text data back into speech.
[2064] Detailed description of the system
[2065] Customer service and sales
[2066] 1. User: The user speaks into the device to ask questions or provide information about the product they want to purchase. For example, "I want to buy this smartphone, which plan do you recommend?"
[2067] 2. Terminal: The speech recognition system converts the user's voice data into text data. Furthermore, an emotion engine is used to recognize emotions from the user's voice.
[2068] 3. Terminal: Sends the converted text data and sentiment data to the server.
[2069] 4. Server: Analyzes text and sentiment data to determine the user's purchase intent and emotional state. For example, it recognizes the user's needs regarding smartphone models and plans, and determines whether the user is excited or calm.
[2070] 5. Server: Retrieves relevant information (product data and plan information) from the database. Adjusts suggestions based on sentiment data.
[2071] 6. Server: Sends the adjusted suggestions to the terminal. The speech synthesis system converts them into speech using appropriate tone and expression.
[2072] 7. Terminal: The robot provides the user with generated suggestions via voice.
[2073] 8. User: Review the proposal, ask further questions, or decide to purchase.
[2074] Specific example:
[2075] The user asks the robot, "I want to buy this smartphone, which plan do you recommend?"
[2076] The robot responds, "For this smartphone, we recommend a plan with a large data allowance."
[2077] Examples of prompts to input into a generative AI model:
[2078] A user asked, "I want to buy this smartphone, which plan do you recommend?" Please generate the best answer suggesting smartphone plans. The user seems a little excited.
[2079] Instructions
[2080] 1. User: The user asks the robot a question about a specific smartphone operation. For example, they might say, "I don't know how to send a photo, could you please tell me how?"
[2081] 2. Terminal: The speech recognition system converts the user's voice data into text data. Simultaneously, the emotion engine recognizes the user's emotions.
[2082] 3. Terminal: Sends the converted text data and sentiment data to the server. It also uses the camera function to recognize the model of the user's smartphone.
[2083] 4. Server: Based on image recognition results, text data, and emotion data, the server retrieves the corresponding model's operation manual from the database.
[2084] 5. Server: Generates specific operating procedures based on user requests from the operation manual, and further adjusts the difficulty and tone of the explanations by taking sentiment data into consideration.
[2085] 6. Server: Returns the adjusted operating procedure to the terminal and converts it into speech using the speech synthesis system.
[2086] 7. Terminal: The robot provides voice instructions to the user. Tones and expressions are based on emotional data.
[2087] 8. User: Operate the smartphone following the provided instructions.
[2088] Specific example:
[2089] The user asks the robot, "I don't know how to send a photo, could you please tell me how?"
[2090] The robot explains, "First, open the camera app, then tap the send button."
[2091] Examples of prompts to input into a generative AI model:
[2092] A user asked, "I don't know how to send photos, could you please explain?" Please provide clear instructions on how to do this on a smartphone. The user seems a little confused.
[2093] Cash register management
[2094] 1. Terminal: Collect daily sales data from the POS terminal.
[2095] 2. Terminal: Sends the collected sales data to the server.
[2096] 3. Server: Analyzes sales data and calculates the current cash inventory level.
[2097] 4. Server: Generates an alert and sends a notification if the cash inventory falls below the appropriate level.
[2098] 5. Terminal: Sends alert notifications to store administrators in real time.
[2099] 6. User: The store manager who receives the notification will take action to address any cash discrepancies.
[2100] Specific example:
[2101] If cash inventory falls below the appropriate level, the system will notify the store manager with the message, "Cash inventory has fallen below the appropriate level. Please check."
[2102] Examples of prompts to input into a generative AI model:
[2103] Analyze cash inventory levels based on sales data and generate an alert if they fall below the appropriate level.
[2104] Inventory Management
[2105] 1. Terminal: Registers product arrival and sales data in real time.
[2106] 2. Terminal: Sends registered data to the server.
[2107] 3. Server: Calculates the current inventory level based on incoming and outgoing data.
[2108] 4. Server: Generates feedback when inventory levels fall below a set threshold. Sentiment data is also used as needed.
[2109] 5. Server: Generates order plans based on feedback.
[2110] 6. Server: Sends the generated order plan to the terminal.
[2111] 7. Terminal: Sends notifications to staff via a robot. It can also adjust tone and expression based on emotional data.
[2112] 8. User: The staff member who receives the notification will proceed with the ordering process.
[2113] Specific example:
[2114] If inventory falls below a set threshold, the system will notify staff with a message saying, "Inventory is low. Please place an order."
[2115] Examples of prompts to input into a generative AI model:
[2116] Based on inventory data, analyze inventory levels and generate and notify us of an order plan if the levels fall below a set threshold.
[2117] Implementing this system requires a combination of appropriate speech recognition, emotion recognition, and speech synthesis systems. Furthermore, infrastructure development is necessary to ensure smooth data communication between the server and terminals. This configuration and operation will enable the provision of high-quality services that take user emotions into consideration, leading to improved efficiency in store operations and increased customer satisfaction.
[2118] The flow of the specific processing in Example 2 will be explained using Figure 13.
[2119] Customer service and sales
[2120] Step 1:
[2121] Users speak into the device to provide information about the products they want to purchase or to ask questions.
[2122] Input: User's voice data.
[2123] Output: Audio data is input to the terminal.
[2124] Specific action: The user says, "I want to buy this smartphone, which plan do you recommend?"
[2125] Step 2:
[2126] The device converts voice data into text data using a speech recognition system, and simultaneously recognizes emotions using an emotion engine.
[2127] Input: User's voice data.
[2128] Output: Converted text data and sentiment data.
[2129] Data processing / data calculation: Convert speech to text using Google Speech-to-Text and analyze emotions using Emotion AI.
[2130] Specific operation: From the voice data, text data such as "I want to buy this smartphone, which plan do you recommend?" and emotion data such as "excitement" are obtained.
[2131] Step 3:
[2132] The device sends text data and sentiment data to the server.
[2133] Input: Text data and sentiment data.
[2134] Output: Data sent to the server.
[2135] Specific operation: The device sends the converted text data and sentiment data to the server.
[2136] Step 4:
[2137] The server analyzes text data and sentiment data to determine the user's purchase intent and emotional state.
[2138] Input: Text data and sentiment data.
[2139] Output: User purchase intent and emotional state information.
[2140] Data processing / data calculation: Perform text data analysis and sentiment data analysis to determine user needs.
[2141] Specific actions: The system distinguishes between a purchase intention ("looking for a smartphone plan") and an emotional state ("excited").
[2142] Step 5:
[2143] The server retrieves relevant information (product data and plan information) from the database.
[2144] Input: Purchase intent and emotional state information.
[2145] Output: Related product information and plan information.
[2146] Data processing / data calculation: Search and retrieve relevant information from a MySQL database.
[2147] Specific action: Retrieve related information such as "Smartphone Plan A, B, C".
[2148] Step 6:
[2149] The server adjusts the suggestions based on sentiment data.
[2150] Input: Related product information, plan information, and sentiment data.
[2151] Output: Revised proposal.
[2152] Data processing / data calculation: Generate optimal suggestions for users based on emotional data.
[2153] Specific action: For excited users, propose a "simple and intuitive plan."
[2154] Step 7:
[2155] The server sends the adjusted proposal to the terminal, which then converts it into speech using a speech synthesis system.
[2156] Input: Adjusted proposal.
[2157] Output: Audio data.
[2158] Data processing / data calculation: Convert the proposed content into audio data using Amazon Polly.
[2159] Specific operation: The server sends a suggestion message saying "Simple Plan A is recommended," and the terminal converts it into speech.
[2160] Step 8:
[2161] The device provides the user with generated suggestions via voice.
[2162] Input: Audio data.
[2163] Output: Audio output to the user.
[2164] Specific action: The robot says, "I recommend the simple Plan A."
[2165] Step 9:
[2166] The user reviews the proposal, asks further questions, or decides to make a purchase.
[2167] Input: Audio of the proposed content.
[2168] Output: User's next action (continue questioning or make a purchase decision).
[2169] Specific action: The user accepts the offer and responds, "I will purchase it."
[2170] Instructions
[2171] Step 1:
[2172] A user asks the robot questions about specific smartphone operations.
[2173] Input: User's voice data.
[2174] Output: Audio data is input to the terminal.
[2175] Specific action: The user says, "I don't know how to send photos, could you please tell me how?"
[2176] Step 2:
[2177] The device converts voice data into text data using a speech recognition system, and simultaneously recognizes emotions using an emotion engine.
[2178] Input: User's voice data.
[2179] Output: Converted text data and sentiment data.
[2180] Data processing / data calculation: Convert speech to text using Google Speech-to-Text and analyze emotions using Emotion AI.
[2181] Specific operation: From the audio data, text data saying "I don't know how to send a photo, please tell me" and emotion data indicating "confusion" are obtained.
[2182] Step 3:
[2183] The device sends text data and sentiment data to the server, and the camera function is used to recognize the smartphone model.
[2184] Input: Text data, sentiment data, and smartphone image data.
[2185] Output: Data sent to the server and recognized model information.
[2186] Specific operation: The device sends the user's question to the server, and the smartphone image captured by the camera is analyzed to identify the model.
[2187] Step 4:
[2188] The server retrieves the corresponding model's operation manual from the database based on the image recognition results, text data, and emotion data.
[2189] Input: Image recognition results, text data, sentiment data.
[2190] Output: Operation manual for the applicable model.
[2191] Data processing / data calculation: Obtain the operation manual corresponding to the identified model from the database.
[2192] Specific operation: The server searches for and retrieves the user manual for the corresponding smartphone.
[2193] Step 5:
[2194] The server generates specific operating procedures from the user manual based on the user's request, and adjusts the difficulty and tone of the explanation based on sentiment data.
[2195] Input: Operation manual, emotion data.
[2196] Output: Adjusted operating instructions.
[2197] Data processing / data calculation: Generate operating procedures while considering emotional data, and explain them to the user in the most appropriate tone.
[2198] Specific actions: Explain "simple and intuitive steps" to confused users.
[2199] Step 6:
[2200] The server returns the adjusted operating procedure to the terminal, which then converts it into speech using a speech synthesis system.
[2201] Input: Adjusted operating procedure.
[2202] Output: Audio data.
[2203] Data processing / data calculation: Convert the operation procedure into audio data using Amazon Polly.
[2204] Specific operation: The server sends instructions such as "First open the camera app, then tap the send button," which the device then converts to audio.
[2205] Step 7:
[2206] The device provides the user with voice instructions on how to operate it. Tones and expressions based on emotional data are used.
[2207] Input: Audio data.
[2208] Output: Audio output to the user.
[2209] Specific instructions: The robot explains, "First, open the camera app, then tap the send button."
[2210] Step 8:
[2211] The user operates the smartphone according to the instructions provided.
[2212] Input: Audio instructions for the operating procedure.
[2213] Output: Smartphone operation complete.
[2214] Specific operation: The user follows the robot's instructions and operates their smartphone to send a photo.
[2215] Cash register management
[2216] Step 1:
[2217] The terminal collects daily sales data from the POS terminal.
[2218] Input: Sales data.
[2219] Output: Collected sales data.
[2220] Specific operation: The POS terminal saves daily sales data to the terminal.
[2221] Step 2:
[2222] The terminal sends the collected sales data to the server.
[2223] Input: Collected sales data.
[2224] Output: Sales data sent to the server.
[2225] Specific action: The terminal transfers sales data to the server.
[2226] Step 3:
[2227] The server analyzes sales data and calculates the current cash inventory level.
[2228] Input: Sales data.
[2229] Output: Cash inventory.
[2230] Data processing / calculation: Calculate cash inventory based on sales data.
[2231] Specific operation: The server analyzes sales data and calculates the amount of cash inventory.
[2232] Step 4:
[2233] The server generates an alert and sends a notification if the cash inventory falls below the appropriate level.
[2234] Input: Cash inventory quantity.
[2235] Output: Alert notification.
[2236] Data processing / calculation: Generate an alert when it is confirmed that the cash inventory level falls below the appropriate level.
[2237] Specific action: The server prepares the alert notification.
[2238] Step 5:
[2239] The device sends an alert notification to the store administrator.
[2240] Input: Alert notification data.
[2241] Output: Notification to the store manager.
[2242] Specific operation: The device notifies the store manager of alerts in real time.
[2243] Step 6:
[2244] The store manager who receives the notification from the user will then take action to address any cash discrepancies.
[2245] Input: Alert notification.
[2246] Output: Issue resolved.
[2247] Specific action: The store manager checks the notification and takes action to address any cash discrepancies.
[2248] Inventory Management
[2249] Step 1:
[2250] The terminal registers product data in real time.
[2251] Input: Incoming and sales data.
[2252] Output: Registered data.
[2253] Specific operation: The terminal updates the database with arrival and sales information in real time.
[2254] Step 2:
[2255] The device sends the registration data to the server.
[2256] Input: Registration data.
[2257] Output: Data sent to the server.
[2258] Specific action: The device transfers the registration data to the server.
[2259] Step 3:
[2260] The server calculates the current inventory level based on incoming and sales data.
[2261] Input: Registration data.
[2262] Output: Current inventory level.
[2263] Data processing / data calculation: Calculate inventory levels based on incoming and outgoing data.
[2264] Specific operation: The server analyzes incoming and outgoing data and calculates the inventory level.
[2265] Step 4:
[2266] When the server's inventory level falls below a set threshold, it generates feedback. Sentiment data is also used as needed.
[2267] Input: Inventory quantity.
[2268] Output: Feedback.
[2269] Data processing / data calculation: Generate feedback when inventory levels fall below a threshold.
[2270] Specific actions: Generate feedback and determine whether an order is necessary.
[2271] Step 5:
[2272] The server generates an order plan based on the feedback.
[2273] Input: Feedback.
[2274] Output: Ordering plan.
[2275] Data processing / data calculation: Create order plans and make necessary adjustments.
[2276] Specific operation: The server generates the order plan.
[2277] Step 6:
[2278] The server sends the generated order plan to the terminal.
[2279] Input: Order plan.
[2280] Output: Data sent to the terminal.
[2281] Specific action: The server sends the order plan to the terminal.
[2282] Step 7:
[2283] The device sends notifications to staff via a robot. It can also adjust the tone and expression based on emotional data.
[2284] Input: Order planning data.
[2285] Output: Notification to staff.
[2286] Specific operation: The terminal passes the order plan to the robot, and the robot notifies the staff.
[2287] Step 8:
[2288] The staff member who receives the user notification will then proceed with the ordering process.
[2289] Input: Notification data.
[2290] Output: Order completed.
[2291] Specific action: The staff member who receives the notification will carry out the necessary ordering tasks.
[2292] (Application Example 2)
[2293] Next, we will explain application example 2. In the following explanation, the data processing device 12 will be referred to as the "server," and the headset-type terminal 314 will be referred to as the "terminal."
[2294] Traditional store operations have been cumbersome, involving tasks such as customer service, sales, operation guidance, cash register management, and inventory management, making efficient operation difficult. Furthermore, it has been challenging to respond flexibly to user emotions and needs, leading to a demand for improved customer satisfaction. This invention aims to solve these problems by providing a system that recognizes user emotions and provides optimal responses, thereby improving the efficiency of store operations and enhancing customer satisfaction.
[2295] In Application Example 2, the identification processing by the identification processing unit 290 of the data processing device 12 is realized by the following means. In this invention, the server includes means for converting the user's voice data into text data, means for analyzing the text data to determine purchase intent and operation requests, means for generating optimal suggestions or operation procedures based on the determination results and the user's emotional data, means for providing the generated suggestions or operation procedures to the user by speech synthesis, means for identifying the user's mobile terminal by image recognition and obtaining an operation manual corresponding to the model of the mobile terminal from a database, means for generating operation procedures according to the user's requests from the obtained operation manual and adjusting the generated operation procedures based on the user's emotional data, means for collecting sales data from a cash register terminal in real time and analyzing the sales data to calculate the amount of cash inventory, means for generating and notifying an alert when the calculated amount of cash inventory falls below an appropriate amount, and means for adjusting the generated alert notification based on the user's emotional data. This makes it possible to provide more personalized responses by reflecting the user's emotional data in the overall system adjustments, suggestions, and notifications.
[2296] "Audio data" refers to the audio signals spoken by the user.
[2297] "Text data" refers to digital information obtained by converting audio data into a string of characters.
[2298] "Emotional data" refers to data that indicates the emotional state of a user, as recognized from their voice and facial expressions.
[2299] "Purchase intent" refers to a user's intention to purchase a particular product or service.
[2300] An "operation request" refers to a user's desire to learn about a specific operation or usage method.
[2301] A "proposal" refers to an introduction or advice regarding specific products or services provided in response to user requests.
[2302] "Operating procedures" refer to specific methods of operation provided in response to user requests.
[2303] "Speech synthesis" is a technology that converts text data into audio signals for playback.
[2304] "Image recognition" is a technology that recognizes specific objects from image data acquired using input devices such as cameras.
[2305] A "database" is a collection of information stored in digital format, making it easy to search and retrieve.
[2306] "Sales data" refers to digital information about sales collected through point-of-sale terminals.
[2307] "Cash inventory" refers to the total amount of cash present in a store.
[2308] An "alert" is a warning signal used to notify users of important changes or anomalies in the situation.
[2309] This invention is a system that streamlines various tasks that users experience in stores, such as customer service, sales, operation guidance, cash register management, and inventory management, and further recognizes user emotions to provide optimal responses. The following describes the specific forms for implementing the system.
[2310] 1. Customer service and sales
[2311] User: The user speaks to the robot and asks questions or requests information about products they want to purchase. For example, "I want to buy this smartphone, which plan do you recommend?"
[2312] Terminal: The robot's microphone acquires voice data, and a speech recognition system (e.g., Python's speech_recognition library) converts the voice data into text data. It also uses an emotion engine (e.g., EmotionEngine) to recognize emotions from the user's voice.
[2313] Server: Analyzes text and sentiment data to determine the user's purchase intent and emotional state. For example, it recognizes the user's needs regarding smartphone models and plans, and determines whether the user is excited or calm.
[2314] Server: Retrieves relevant information (product data and plan information) from a database (e.g., a relational database management system) and adjusts the suggested content based on sentiment data.
[2315] Terminal: The adjusted suggestions are converted into speech by a speech synthesis system (e.g., TextToSpeech module) and provided to the user by the robot.
[2316] For example, if a user asks, "I want to buy this smartphone, which plan do you recommend?", the robot will respond to the excited user, "The unlimited data plan is the best!"
[2317] 2. Instructions for Use
[2318] User: The user asks the robot questions about how to use their smartphone. For example, they might say, "I don't know how to send a photo, could you please show me how?"
[2319] Terminal: The voice recognition system converts voice data into text data, and the emotion engine recognizes the user's emotions. It also uses the camera function to recognize the model of the user's smartphone.
[2320] Server: Based on image recognition results, text data, and emotion data, retrieves the corresponding model's operation manual from the database.
[2321] Server: Generates specific operating procedures based on user requests from the operation manual, and adjusts the difficulty and tone of the explanations considering sentiment data.
[2322] Terminal: The robot converts the adjusted operating procedures into speech using speech synthesis and explains them to the user.
[2323] For example, if a user asks, "I don't know how to send a photo, could you tell me how?", the robot calmly guides them by saying, "Select the photo you took, press the share button, and then choose the recipient."
[2324] 3. Cash register management
[2325] Terminal: The POS terminal collects daily sales data.
[2326] Server: Analyzes collected sales data and calculates cash inventory levels. A relational database management system is used.
[2327] Server: Generates an alert when cash inventory falls below the appropriate level and makes adjustments based on user sentiment data.
[2328] Terminal: Sends tailored alert notifications to store managers in real time. Notifications use a tone based on sentiment data.
[2329] 4. Inventory Management
[2330] Terminal: Registers product arrival and sales data in real time.
[2331] Server: Calculates the current inventory level based on registered data and generates feedback if the inventory level falls below a threshold. A relational database management system is used.
[2332] Server: Generates order plans based on feedback and adjusts them considering sentiment data.
[2333] Terminal: Notifies staff of the adjusted order plan via robot.
[2334] Examples of prompts for generative AI models
[2335] "It converts user questions into text and recognizes their emotions. Example: 'I want to buy this smartphone, which plan do you recommend?' Emotion: Excitement. Then, it retrieves relevant plan information from the product information database and generates the best recommendation based on the emotion."
[2336] This allows for more personalized responses by incorporating user sentiment data into overall system adjustments, suggestions, and notifications.
[2337] The flow of a specific process in Application Example 2 will be explained using Figure 14.
[2338] Step 1:
[2339] User: Speaks to the robot about the products they want to buy or any questions they have. The voice data from this interaction becomes the input.
[2340] Step 2:
[2341] Terminal: Audio data acquired via the microphone is converted into text data using a speech recognition system (such as Python's speech_recognition library). The converted text data is obtained as output.
[2342] Step 3:
[2343] Terminal: Uses an emotion engine (such as EmotionEngine) to recognize emotion data from the user's voice data. The recognized emotion data is obtained as output.
[2344] Step 4:
[2345] Terminal: Sends text data and sentiment data obtained through speech recognition to the server. Input data consists of text data and sentiment data, which are sent to the server.
[2346] Step 5:
[2347] Server: Analyzes received text and sentiment data to determine the user's purchase intent and action requests. The results obtained (e.g., purchase intent and action requests) become the output of the analysis.
[2348] Step 6:
[2349] Server: Based on relevant information (such as product data and plan information) obtained from the database (relational database management system), the server generates optimal suggestions based on the user's sentiment data. The generated suggestions are then output.
[2350] Step 7:
[2351] Server: Sends the generated proposal content to the terminal. The input data is the proposal content, which is then sent to the terminal.
[2352] Step 8:
[2353] Terminal: The received suggestion content is converted into speech using a speech synthesis system (such as the TextToSpeech module). Audio data is obtained as output.
[2354] Step 9:
[2355] Terminal: The robot plays voice data and provides suggestions to the user. The user reviews the suggestions and decides on their next action.
[2356] The specific processing unit 290 transmits the result of the specific processing to the headset terminal 314. In the headset terminal 314, the control unit 46A causes the speaker 240 and display 343 to output the result of the specific processing. The microphone 238 acquires audio indicating user input for the result of the specific processing. The control unit 46A transmits the audio data indicating user input acquired by the microphone 238 to the data processing unit 12. In the data processing unit 12, the specific processing unit 290 acquires the audio data.
[2357] Data generation model 58 is a type of so-called generative AI (Artificial Intelligence). One example of data generation model 58 is ChatGPT (Internet search<URL: https: / / openai.com / blog / chatgpt> ), Gemini (Internet search) <url: https: gemini.google.com ?hl="ja">Examples of generative AI include the following. The data generation model 58 is obtained by performing deep learning on a neural network. The data generation model 58 is input with prompts containing instructions, and with inference data such as audio data representing speech, text data representing text, and image data representing images. The data generation model 58 infers from the input inference data according to the instructions indicated by the prompts, and outputs the inference results in data formats such as audio data and text data. Here, inference refers to, for example, analysis, classification, prediction, and / or summarization.
[2358] In the above embodiment, an example was given in which specific processing is performed by the data processing device 12, but the technology of this disclosure is not limited thereto, and specific processing may also be performed by the headset terminal 314.
[2359] [Fourth Embodiment]
[2360] Figure 7 shows an example of the configuration of the data processing system 410 according to the fourth embodiment.
[2361] As shown in Figure 7, the data processing system 410 includes a data processing device 12 and a robot 414. An example of the data processing device 12 is a server.
[2362] The data processing device 12 comprises a computer 22, a database 24, and a communication interface 26. The computer 22 is an example of a "computer" related to the technology of this disclosure. The computer 22 comprises a processor 28, RAM 30, and storage 32. The processor 28, RAM 30, and storage 32 are connected to a bus 34. The database 24 and the communication interface 26 are also connected to the bus 34. The communication interface 26 is connected to a network 54. An example of the network 54 is a WAN (Wide Area Network) and / or a LAN (Local Area Network).
[2363] The robot 414 includes a computer 36, a microphone 238, a speaker 240, a camera 42, a communication interface 44, and a controlled object 443. The computer 36 includes a processor 46, RAM 48, and storage 50. The processor 46, RAM 48, and storage 50 are connected to a bus 52. The microphone 238, speaker 240, camera 42, and controlled object 443 are also connected to the bus 52.
[2364] The microphone 238 receives voice signals from the user 20 and receives instructions from the user 20. The microphone 238 captures the voice signals from the user 20, converts the captured voice into audio data, and outputs it to the processor 46. The speaker 240 outputs audio according to the instructions from the processor 46.
[2365] Camera 42 is a small digital camera equipped with an optical system including a lens, aperture, and shutter, and an image sensor such as a CMOS (Complementary Metal-Oxide-Semiconductor) image sensor or a CCD (Charge Coupled Device) image sensor, and captures images of the area around the user 20 (for example, an imaging range defined by a field of view equivalent to the width of a typical healthy person's field of vision).
[2366] Communication interface 44 is connected to network 54. Communication interfaces 44 and 26 are responsible for the exchange of various information between processor 46 and processor 28 via network 54. The exchange of various information between processor 46 and processor 28 using communication interfaces 44 and 26 is performed in a secure manner.
[2367] The controlled object 443 includes a display device, LEDs in the eyes, and motors that drive the arms, hands, and feet. The posture and gestures of the robot 414 are controlled by controlling the motors of the arms, hands, and feet. Some of the robot 414's emotions can be expressed by controlling these motors. Furthermore, the robot 414's facial expressions can also be expressed by controlling the illumination state of the LEDs in its eyes.
[2368] Figure 8 shows an example of the main functions of the data processing device 12 and the robot 414. As shown in Figure 8, the data processing device 12 performs specific processing using the processor 28. The storage 32 stores the specific processing program 56.
[2369] The specific processing program 56 is an example of a "program" relating to the technology of this disclosure. The processor 28 reads the specific processing program 56 from the storage 32 and executes the read specific processing program 56 on the RAM 30. The specific processing is realized by the processor 28 operating as a specific processing unit 290 in accordance with the specific processing program 56 executed on the RAM 30.
[2370] The storage 32 stores the data generation model 58 and the emotion identification model 59. The data generation model 58 and the emotion identification model 59 are used by the identification processing unit 290.
[2371] In robot 414, the processor 46 performs the reception output processing. The storage 50 stores the reception output program 60. The processor 46 reads the reception output program 60 from the storage 50 and executes the read reception output program 60 on the RAM 48. The reception output processing is realized by the processor 46 operating as a control unit 46A according to the reception output program 60 executed on the RAM 48.
[2372] Next, the specific processing performed by the specific processing unit 290 of the data processing device 12 will be described. In the following description, the data processing device 12 will be referred to as the "server" and the robot 414 as the "terminal".
[2373] The system of the present invention is configured to streamline customer service, sales, operation guidance, cash register management, and inventory management tasks that users experience in stores. The following describes in detail the embodiments for implementing the system of the present invention.
[2374] Customer service and sales
[2375] User: The user speaks to the robot to ask questions or provide information about products they want to purchase. For example, "I want to buy this smartphone, which plan do you recommend?"
[2376] Terminal: The speech recognition system converts the user's voice data into text data. The text data is then sent to the server.
[2377] Server: Analyzes text data to determine the user's purchase intent. For example, it recognizes needs regarding smartphone models and plans.
[2378] Server: Based on purchase intent, the system retrieves information from the database to suggest the optimal plan and generates a recommended plan. The recommended plan includes details such as price, data capacity, and benefits.
[2379] Terminal: The recommended plan information is converted into speech using a speech synthesis system, and the robot communicates the proposal to the user.
[2380] User: Review the proposal, ask further questions, or decide to purchase.
[2381] Instructions
[2382] User: The user asks the robot a question about a specific smartphone operation. For example, they might say, "I don't know how to send a photo, could you please tell me how?"
[2383] Terminal: The voice recognition system converts the user's voice data into text data. It also uses the camera function to recognize the model of the user's smartphone.
[2384] Server: Based on the image recognition results and text data, retrieves the operation manual for the corresponding model from the database.
[2385] Server: Generates specific operating procedures based on user requests from the operation manual.
[2386] Terminal: The generated operating procedure is converted into speech using a speech synthesis system, and the robot explains it to the user.
[2387] User: Operate the smartphone following the instructions provided.
[2388] Cash register management
[2389] Terminal: The POS terminal collects daily sales data and sends it to the server.
[2390] Server: Analyzes sales data and calculates the current cash inventory level.
[2391] Server: Generates an alert and sends a notification if the cash inventory falls below the appropriate level.
[2392] Terminal: Sends alert notifications to store administrators in real time.
[2393] User: The store manager who receives the notification will take action to address any cash discrepancies.
[2394] Inventory Management
[2395] Terminal: Registers product arrival and sales data in real time and sends it to the server.
[2396] Server: Calculates current inventory levels based on incoming and outgoing data.
[2397] Server: When inventory levels fall below a set threshold, it generates feedback and creates an ordering plan.
[2398] Terminal: Sends notifications to staff via robot to confirm the generated order plan.
[2399] User: Staff members who receive the notification will proceed with placing the order.
[2400] The system of this invention streamlines store operations and improves the quality of customer service. In particular, it enables users to receive quick and accurate information at stores, leading to high customer satisfaction. Furthermore, the automation of cash register management and inventory management solves the problem of labor shortages.
[2401] The following describes the processing flow.
[2402] Customer service and sales
[2403] Step 1:
[2404] Users can ask the robot questions or request information about products they want to buy. For example, they might ask, "I want to buy this smartphone, which plan do you recommend?"
[2405] Step 2:
[2406] The device uses a speech recognition system to convert the user's voice data into text data.
[2407] Step 3:
[2408] The terminal sends the converted text data to the server.
[2409] Step 4:
[2410] The server analyzes text data to determine the user's purchase intent. For example, it recognizes their needs regarding smartphone models and plans.
[2411] Step 5:
[2412] The server retrieves relevant information (product data and plan information) from the database.
[2413] Step 6:
[2414] The server generates the optimal plan based on the user's needs and creates text data containing detailed information about the recommended plan.
[2415] Step 7:
[2416] The text data generated by the server is returned to the terminal and converted into speech using a speech synthesis system.
[2417] Step 8:
[2418] The terminal, via a robot, provides the user with generated suggestions via voice.
[2419] Step 9:
[2420] The user reviews the proposal and decides on their next action (e.g., continue asking questions, make a purchase).
[2421] Instructions
[2422] Step 1:
[2423] The user asks the robot questions about specific smartphone operations. For example, they might say, "I don't know how to send a photo, could you please tell me how?"
[2424] Step 2:
[2425] The device uses a speech recognition system to convert the user's voice data into text data.
[2426] Step 3:
[2427] The device uses the smartphone's camera to perform image recognition to determine the model of the user's mobile device.
[2428] Step 4:
[2429] The terminal sends the converted text data and image recognition results to the server.
[2430] Step 5:
[2431] Based on the image recognition results, the server retrieves the corresponding model's operation manual from the database.
[2432] Step 6:
[2433] The server generates specific operating procedures based on user requests from the acquired operation manual.
[2434] Step 7:
[2435] The server returns the generated operating procedure as text data to the terminal, and the speech synthesis system converts it into speech.
[2436] Step 8:
[2437] The terminal provides the user with voice instructions on how to operate it via a robot.
[2438] Step 9:
[2439] The user operates the smartphone according to the instructions provided.
[2440] Cash register management
[2441] Step 1:
[2442] The terminal collects daily sales data from the POS terminal.
[2443] Step 2:
[2444] The terminal sends the collected sales data to the server.
[2445] Step 3:
[2446] The server analyzes sales data and calculates the current cash inventory level.
[2447] Step 4:
[2448] Based on the cash inventory amount calculated by the server, the appropriate inventory level is calculated.
[2449] Step 5:
[2450] The server compares the appropriate inventory level with the actual inventory level and generates an alert if there is a surplus or shortage.
[2451] Step 6:
[2452] The server sends the generated alert to the terminal.
[2453] Step 7:
[2454] The device sends alert notifications to the store administrator in real time.
[2455] Step 8:
[2456] The store manager who receives the user notification will then take action to address any cash discrepancies.
[2457] Inventory Management
[2458] Step 1:
[2459] The terminal registers product arrival and sales data in real time.
[2460] Step 2:
[2461] The device sends the registered data to the server.
[2462] Step 3:
[2463] The server calculates the current inventory level based on incoming and outgoing data.
[2464] Step 4:
[2465] The server generates feedback when the inventory level falls below a set threshold.
[2466] Step 5:
[2467] The server generates an order plan based on the feedback.
[2468] Step 6:
[2469] The server sends the generated order plan to the terminal.
[2470] Step 7:
[2471] The device sends notifications to staff via a robot.
[2472] Step 8:
[2473] The staff member who receives the user notification will then proceed with the ordering process.
[2474] Thus, the system of the present invention achieves efficient and high-quality store operations through each step.
[2475] (Example 1)
[2476] Next, we will describe Example 1. In the following description, the data processing device 12 will be referred to as the "server" and the robot 414 as the "terminal".
[2477] In traditional retail operations, a wide range of tasks, including customer service, inventory management, and cash register management, are performed manually. This is not only inefficient but also prone to human error and wasted time. As a result, problems such as decreased customer satisfaction, reduced operational efficiency, and increased burden due to labor shortages arise. A system is needed to solve these problems and provide timely and accurate information while improving operational efficiency.
[2478] The identification process performed by the identification processing unit 290 of the data processing device 12 in Example 1 is realized by the following means.
[2479] In this invention, the server includes means for converting user voice data into text data, means for analyzing the text data to determine purchase intent and operation requests, means for generating optimal suggestions or operation procedures based on the determination results, means for providing the generated suggestions or operation procedures to the user by speech synthesis, means for registering product arrival and sales data in real time and analyzing the data to calculate the current inventory level, means for formulating an order plan and notifying the user when the calculated inventory level falls below a set threshold, means for collecting sales data from a cash register terminal and analyzing the sales data to calculate the cash inventory level, means for generating an alert and notifying the user when the calculated cash inventory level falls below an appropriate level, and means for streamlining customer service, sales, operation guidance, cash register management, and inventory management that the user receives in the store. This enables automation and efficiency of store operations, improving the quality of user service and operational efficiency.
[2480] "Audio data" refers to sound information acquired digitally via a microphone or other audio input device, such as user speech or instructions.
[2481] "Text data" is a data format that converts audio data into written text.
[2482] "Analysis" is the process of understanding the meaning and intent of input data, and often involves using technologies such as natural language processing.
[2483] "Purchase intent" refers to the user's wishes and desires regarding the type and conditions of the product they are trying to acquire.
[2484] An "operation request" refers to a request from a user for assistance in performing a specific operation.
[2485] "Speech synthesis" is a technology that converts text data into speech that sounds like human speech.
[2486] "Product arrival data" refers to data that records information about when new products arrive at a store.
[2487] "Sales data" refers to data that records information about when a product was sold in a store.
[2488] "Inventory level" refers to the quantity of goods currently stored in the store or warehouse.
[2489] An "ordering plan" refers to a plan for purchasing new goods when inventory falls below a certain level.
[2490] A "cash register terminal" is an electronic device used for selling goods and processing payments.
[2491] "Sales data" refers to data that records information such as the quantity, price, and date and time of the transaction of the goods sold.
[2492] "Cash inventory" refers to the total amount of cash stored in cash registers and within the store.
[2493] An "alert" refers to a warning or cautionary message that is sent when certain conditions are met.
[2494] "Customer service" refers to the work of providing product descriptions, guidance, and answering questions from customers.
[2495] "Operation guidance" refers to the service of providing customers with instructions on how to use the equipment and services they use.
[2496] "Cash register management" refers to the task of managing daily sales and cash inflows and outflows.
[2497] "Inventory management" refers to the overall management of receiving, storing, and shipping goods in order to maintain an appropriate level of inventory.
[2498] The system of the present invention aims to streamline store operations and improve customer satisfaction. This system is configured to support the various tasks that users perform in stores, including customer service and sales, operation guidance, cash register management, and inventory management. The following describes specific embodiments for carrying out the present invention.
[2499] Customer service and sales
[2500] User: The user speaks to the robot to ask questions or provide information about products they want to purchase. For example, they might say, "I want to buy this smartphone, which plan do you recommend?"
[2501] Terminal: Uses a speech recognition system (e.g., a cloud-based speech recognition service) to convert the user's voice data into text data. The converted text data is sent to a server via the internet.
[2502] Server: Uses a Natural Language Processing (NLP) engine (e.g., a cloud-based natural language processing service) to analyze text data and determine the user's purchase intent. This purchase intent includes needs regarding smartphone models and plans.
[2503] Server: Based on the determination result, the server retrieves information on the corresponding plan from the database (e.g., relational database management system) and recommends the optimal plan. The recommended plan includes pricing, data capacity, and benefits.
[2504] Terminal: A speech synthesis system (e.g., speech synthesis API) converts the recommended plan information into speech, and a robot communicates the proposal to the user.
[2505] User: Review the proposal, ask further questions, or decide to purchase.
[2506] Specific example:
[2507] User: "I want to buy this smartphone, which plan would you recommend?"
[2508] Robot: "The best plan for you is the one with X GB of data for Y yen per month. This plan comes with the following benefits."
[2509] Examples of prompts to input into a generative AI model:
[2510] "Please generate a database query to analyze the user's purchase intent for the specified product and propose the optimal plan."
[2511] Instructions
[2512] User: The user asks the robot a question about a specific smartphone operation. For example, they might say, "I don't know how to send a photo, could you please tell me how?"
[2513] Terminal: A voice recognition system (e.g., a cloud-based voice recognition service) converts the user's voice data into text data. It also uses the camera function to recognize the model of the user's smartphone.
[2514] Server: Based on image recognition results (e.g., an image recognition model using deep learning) and text data, retrieves the operation manual for the corresponding model from the database.
[2515] Server: Generates specific operating procedures based on user requests from the operation manual.
[2516] Terminal: The generated operating procedure is converted into speech using a speech synthesis system (e.g., speech synthesis API), and the robot explains it to the user.
[2517] User: Operate the smartphone following the instructions provided.
[2518] Specific example:
[2519] User: "I don't know how to send photos, could you please tell me how?"
[2520] Robot: "Your device is a [model name]. To send a photo, first open your gallery, then press the share button and select the recipient."
[2521] Examples of prompts to input into a generative AI model:
[2522] "Generate specific steps to explain how to operate the smartphone as instructed by the user."
[2523] Cash register management
[2524] Terminal: A point-of-sale terminal (e.g., a point-of-sale management system) collects daily sales data and transmits it to a server via the internet.
[2525] Server: Aggregates and analyzes sales data to calculate the current cash inventory level.
[2526] Server: Generates an alert if the cash inventory falls below the appropriate level.
[2527] Terminal: Sends alert notifications to store administrators in real time.
[2528] User: The store manager who receives the notification will take action to address any cash discrepancies.
[2529] Specific example:
[2530] Server: "Cash inventory is insufficient. Please take immediate action."
[2531] Manager: "We'll replenish the cash immediately."
[2532] Examples of prompts to input into a generative AI model:
[2533] "Calculate cash inventory levels based on daily sales data and generate alerts when inventory levels are low."
[2534] Inventory Management
[2535] Terminal: Product arrival and sales data (e.g., sales management system) are registered in real time and transmitted to the server via the internet.
[2536] Server: Calculates current inventory levels based on incoming and outgoing data.
[2537] Server: When inventory levels fall below a set threshold, it generates feedback and develops an ordering plan.
[2538] Terminal: Sends notifications to staff via robot to confirm the generated order plan.
[2539] User: Staff members who receive the notification will proceed with placing the order.
[2540] Specific example:
[2541] Server: "Inventory levels have fallen below the threshold. An order is required."
[2542] Staff: "We have confirmed your order. We have placed the necessary items."
[2543] Examples of prompts to input into a generative AI model:
[2544] "Monitor current inventory levels and generate reordering plans when they fall below the set threshold."
[2545] The system of this invention is expected to automate many store operations and improve the quality of customer service. In particular, it will enable users to receive quick and accurate information at stores, resulting in high customer satisfaction. Furthermore, the automation of cash register management and inventory management will solve the problem of labor shortages.
[2546] The flow of the specific processing in Example 1 will be explained using Figure 11.
[2547] Customer service and sales
[2548] Step 1:
[2549] Users can ask the robot questions or request information about products they want to buy. For example, they might say, "I want to buy this smartphone, which plan do you recommend?"
[2550] Input: User's voice
[2551] Output: Audio data
[2552] Step 2:
[2553] The device uses a speech recognition system (e.g., a cloud-based speech recognition service) to convert speech data into text data. The converted text data is then sent to the server.
[2554] Input: Audio data
[2555] Output: Text data
[2556] Step 3:
[2557] The server uses a Natural Language Processing (NLP) engine (e.g., a cloud-based natural language processing service) to analyze text data and determine the user's purchase intent.
[2558] Input: Text data
[2559] Output: Classification result (user's purchase intent)
[2560] Step 4:
[2561] Based on the determination result, the server retrieves information on the corresponding plan from the database (e.g., a relational database management system) and recommends the optimal plan.
[2562] Input: Discrimination result
[2563] Output: Recommended plan (price, data capacity, benefits, etc.)
[2564] Step 5:
[2565] The terminal uses a speech synthesis system (e.g., a speech synthesis API) to convert the recommended plan information into speech, and a robot then communicates the proposal to the user.
[2566] Input: Recommended plan
[2567] Output: Audio data (proposed content)
[2568] Step 6:
[2569] The user reviews the proposal, asks further questions, or decides to make a purchase.
[2570] Input: Audio data (proposal)
[2571] Output: User decisions and questions
[2572] Adding specific actions
[2573] When a user asks, "I want to buy this smartphone, which plan do you recommend?", the robot suggests, "The best plan for you is the one with X GB of data for Y yen per month. This plan comes with the following benefits."
[2574] Instructions
[2575] Step 1:
[2576] The user asks the robot questions about specific smartphone operations. For example, they might say, "I don't know how to send a photo, could you please tell me how?"
[2577] Input: User's voice
[2578] Output: Audio data
[2579] Step 2:
[2580] The device converts the user's voice data into text data using a voice recognition system (e.g., a cloud-based voice recognition service). It also utilizes its camera function to recognize the model of the user's smartphone.
[2581] Input: Audio data and image data
[2582] Output: Text data and image recognition results
[2583] Step 3:
[2584] The server retrieves the operation manual for the corresponding model from the database based on the image recognition results (e.g., an image recognition model using deep learning) and text data.
[2585] Input: Text data and image recognition results
[2586] Output: Operation Manual
[2587] Step 4:
[2588] The server generates specific operating procedures based on the user's request, using the operation manual.
[2589] Input: Operation Manual
[2590] Output: Operating Procedure
[2591] Step 5:
[2592] The terminal converts the generated operating procedures into speech using a speech synthesis system (e.g., a speech synthesis API), and the robot explains them to the user.
[2593] Input: Operating Procedure
[2594] Output: Audio data (operating instructions)
[2595] Step 6:
[2596] The user operates the smartphone according to the instructions provided.
[2597] Input: Audio data (operating instructions)
[2598] Output: Smartphone operation results
[2599] Adding specific actions
[2600] When a user asks, "I don't know how to send photos, could you please tell me how?", the robot responds, "Your device is a XX. To send photos, first open your gallery, then press the share button and select the recipient."
[2601] Cash register management
[2602] Step 1:
[2603] A point-of-sale terminal (e.g., a point-of-sale management system) collects daily sales data and transmits it to a server via the internet.
[2604] Input: Sales data
[2605] Output: Send to server
[2606] Step 2:
[2607] The server aggregates and analyzes sales data to calculate the current cash inventory level.
[2608] Input: Sales data
[2609] Output: Cash Inventory
[2610] Step 3:
[2611] The server generates an alert if the cash inventory falls below the appropriate level.
[2612] Input: Cash inventory
[2613] Output: Alert
[2614] Step 4:
[2615] The device sends alert notifications to the store administrator in real time.
[2616] Input: Alert
[2617] Output: Alert notification
[2618] Step 5:
[2619] The store manager will take action to address any cash discrepancies upon receiving notification.
[2620] Input: Alert notification
[2621] Output: Cash replenishment and cash depletion response
[2622] Adding specific actions
[2623] The server sends an alert saying, "Cash inventory is low. Please take immediate action," and the store manager checks it and responds, "We will replenish the cash immediately."
[2624] Inventory Management
[2625] Step 1:
[2626] The terminal registers product arrival and sales data (e.g., sales management system) in real time and transmits it to the server via the internet.
[2627] Input: Incoming data and sales data
[2628] Output: Send to server
[2629] Step 2:
[2630] The server calculates the current inventory level based on incoming and outgoing data.
[2631] Input: Incoming data and sales data
[2632] Output: Inventory Quantity
[2633] Step 3:
[2634] The server generates feedback and develops an ordering plan when inventory levels fall below a set threshold.
[2635] Input: Inventory Quantity
[2636] Output: Feedback and...
Claims
1. A means of converting user voice data into text data, A means of analyzing text data to determine purchase intent and operational requests, A means for generating an optimal suggestion or operating procedure based on the determination result, A system including means for providing the generated suggestions or operating procedures to the user by speech synthesis.
2. A means for identifying the user's mobile device using image recognition and obtaining an operation manual corresponding to the model of that mobile device from a database, The system according to claim 1, further comprising means for generating operating procedures in response to user requests from the acquired operating manual.
3. A means for collecting sales data from a cash register terminal in real time, analyzing the sales data, and calculating the amount of cash inventory, The system according to claim 1, further comprising means for generating and notifying an alert when the calculated cash inventory falls below an appropriate amount.
4. A means of registering product data arrival and sales information in real time and calculating the current inventory level, The system according to claim 1, further comprising means for generating an order plan when the calculated inventory quantity falls below a set threshold.
Citation Information
Patent Citations
Persona chatbot control method and system
JP2022180282A