system

The system automates concierge services by converting voice or text inputs to text, analyzing with natural language processing, and generating responses using generative AI, addressing the inefficiencies and quality issues in conventional concierge services.

JP2026062165APending Publication Date: 2026-04-09SOFTBANK GROUP CORP
View PDF 1 Cites 0 Cited by

Patent Information

Authority / Receiving Office
JP · JP
Patent Type
Applications
Current Assignee / Owner
Filing Date
2024-09-30
Publication Date
2026-04-09

AI Technical Summary

Technical Problem

Conventional concierge services require significant time and effort from human concierges, leading to inconsistent quality and delayed responses, especially for complex requests.

Method used

A system that converts user voice or text inputs into text data using speech recognition, analyzes the data with natural language processing, generates responses using generative AI, and allows concierges to confirm and correct these responses before sending them back to the user.

Benefits of technology

This system enables quick and accurate responses to user requests, reducing the burden on concierges and improving service quality by automating the processing of requests.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure 2026062165000001_ABST
    Figure 2026062165000001_ABST
Patent Text Reader

Abstract

We provide the system. [Solution] A user terminal for the user to input requests in voice or text format, A conversion means for converting audio data received from a user into text data, An analysis means for analyzing text data generated by a conversion means as a request using a natural language processing algorithm, A generation means for generating example answers using a generative AI based on the requirements obtained by the analysis means, A transmission means for sending the example answer generated by the generation means to a concierge terminal for confirmation and correction by the concierge, A means for sending corrected responses returned from the concierge terminal to the user terminal, A system that includes this.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The technology of this disclosure relates to a system.

Background Art

[0002] Patent Document 1 discloses a method for controlling a persona chatbot, which is performed by at least one processor, including steps of receiving a user utterance, adding the user utterance to a prompt including an instruction sentence related to an explanation of a chatbot character, encoding the prompt, and inputting the encoded prompt into a language model to generate a chatbot utterance in response to the user utterance.

Prior Art Documents

Patent Documents

[0003]

Patent Document 1

Summary of the Invention

Problems to be Solved by the Invention

[0004] In the conventional concierge service, since the concierge manually responds to the user's requests, a great deal of time and effort are required, and it may be difficult to keep the quality of the service constant. Also, when the user's requests are complicated or a prompt response is required, it is inevitable that the response will be delayed. There is a need to solve such problems and reduce the load on the concierge and improve the quality of the service.

Means for Solving the Problems

[0005] The present invention provides a system for efficiently processing requests entered by users in voice or text format. Specifically, the system includes a conversion means for receiving voice data from a user terminal and converting it into text data using speech recognition technology; an analysis means for analyzing the text data with a natural language processing algorithm and extracting requests; a generation means for generating example answers to requests using a generative AI; a transmission means for sending the example answers to a concierge terminal for confirmation and correction; and a transmission means for sending the corrected answers to the user terminal. As a result, it is possible to respond to user requests quickly and accurately, reduce the burden on concierges, and improve the quality of service.

[0006] A "user" is someone who uses a system to input requests and receive services.

[0007] A "user terminal" is a communication device used by a user to input requests in voice or text format and communicate with the server.

[0008] "Audio data" refers to information that digitally represents the voice spoken by a user.

[0009] "Text data" refers to digital information expressed as a string of characters, which can be converted from audio data or directly entered by the user.

[0010] "Conversion means" refers to a device or software that converts speech data into text data using speech recognition technology.

[0011] "Analysis means" refers to a device or software that uses natural language processing algorithms to analyze text data and extract user requests.

[0012] "Generation means" refers to a device or software that utilizes a generative AI to generate example answers based on requests.

[0013] "Generative AI" refers to algorithms and models that use artificial intelligence technology to generate appropriate responses based on input data.

[0014] A "concierge terminal" is a communication terminal used by concierges to review, correct, and resend sample answers sent from the server.

[0015] "Transmission means" refers to a device or software for sending and receiving text data and sample answers between related terminals.

[0016] A "natural language processing algorithm" is an algorithm or technology used to analyze human language and understand its meaning. [Brief explanation of the drawing]

[0017] [Figure 1] This is a conceptual diagram showing an example of the configuration of a data processing system according to the first embodiment. [Figure 2] This is a conceptual diagram showing an example of the essential functions of a data processing device and a smart device according to the first embodiment. [Figure 3] This is a conceptual diagram showing an example of the configuration of a data processing system according to the second embodiment. [Figure 4] This is a conceptual diagram showing an example of the main functions of a data processing device and smart glasses according to the second embodiment. [Figure 5] This is a conceptual diagram showing an example of the configuration of a data processing system according to the third embodiment. [Figure 6] This is a conceptual diagram showing an example of the main functions of a data processing device and a headset-type terminal according to the third embodiment. [Figure 7] This is a conceptual diagram showing an example of the configuration of a data processing system according to the fourth embodiment. [Figure 8] This is a conceptual diagram showing an example of the main functions of a data processing device and a robot according to the fourth embodiment. [Figure 9] This shows an emotion map where multiple emotions are mapped. [Figure 10] It shows an emotion map to which a plurality of emotions are mapped. [Figure 11] It is a sequence diagram showing the processing flow of the data processing system in Embodiment 1. [Figure 12] It is a sequence diagram showing the processing flow of the data processing system in Application Example 1. [Figure 13] It is a sequence diagram showing the processing flow of the data processing system in Embodiment 2 when an emotion engine is combined. [Figure 14] It is a sequence diagram showing the processing flow of the data processing system in Application Example 2 when an emotion engine is combined.

Mode for Carrying Out the Invention

[0018] Hereinafter, an example of an embodiment of a system according to the technology of the present disclosure will be described with reference to the accompanying drawings.

[0019] First, the terms used in the following description will be described.

[0020] In the following embodiments, the numbered processor (hereinafter simply referred to as "processor") may be one arithmetic unit or a combination of a plurality of arithmetic units. Also, the processor may be one type of arithmetic unit or a combination of a plurality of types of arithmetic units. Examples of arithmetic units include a CPU (Central Processing Unit), a GPU (Graphics Processing Unit), a GPGPU (General-Purpose computing on Graphics Processing Units), an APU (Accelerated Processing Unit), and the like.

[0021] In the following embodiments, the numbered RAM (Random Access Memory) is a memory in which information is temporarily stored and is used as a work memory by the processor.

[0022] In the following embodiments, the signed storage is one or more non-volatile storage devices that store various programs and various parameters. Examples of non-volatile storage devices include flash memory (SSD (Solid State Drive)), magnetic disks (e.g., hard disks), or magnetic tapes.

[0023] In the following embodiments, the signed communication interface (I / F) is an interface that includes a communication processor and an antenna, etc. The communication interface manages communication between multiple computers. Examples of communication standards applicable to the communication interface include wireless communication standards such as 5G (5th Generation Mobile Communication System), Wi-Fi (registered trademark), or Bluetooth (registered trademark).

[0024] In the following embodiments, "A and / or B" is synonymous with "at least one of A and B." That is, "A and / or B" means that it may be A alone, or B alone, or a combination of A and B. Furthermore, in this specification, the same concept as "A and / or B" applies when expressing three or more things linked by "and / or."

[0025] [First Embodiment]

[0026] Figure 1 shows an example of the configuration of the data processing system 10 according to the first embodiment.

[0027] As shown in Figure 1, the data processing system 10 includes a data processing device 12 and a smart device 14. An example of the data processing device 12 is a server.

[0028] The data processing device 12 comprises a computer 22, a database 24, and a communication interface 26. The computer 22 is an example of a "computer" related to the technology of this disclosure. The computer 22 comprises a processor 28, RAM 30, and storage 32. The processor 28, RAM 30, and storage 32 are connected to a bus 34. The database 24 and the communication interface 26 are also connected to the bus 34. The communication interface 26 is connected to a network 54. An example of the network 54 is a WAN (Wide Area Network) and / or a LAN (Local Area Network).

[0029] The smart device 14 comprises a computer 36, a reception device 38, an output device 40, a camera 42, and a communication interface 44. The computer 36 comprises a processor 46, RAM 48, and storage 50. The processor 46, RAM 48, and storage 50 are connected to a bus 52. The reception device 38, output device 40, and camera 42 are also connected to the bus 52.

[0030] The reception device 38 is equipped with a touch panel 38A and a microphone 38B, etc., and receives user input. The touch panel 38A receives user input by detecting contact with an object (e.g., a pen or finger). The microphone 38B receives user input by detecting the user's voice. The control unit 46A transmits data indicating the user input received by the touch panel 38A and microphone 38B to the data processing device 12. In the data processing device 12, the specific processing unit 290 acquires the data indicating the user input.

[0031] The output device 40 includes a display 40A and a speaker 40B, and presents data to the user 20 by outputting the data in a form perceptible to the user 20 (e.g., audio and / or text). The display 40A displays visible information such as text and images according to instructions from the processor 46. The speaker 40B outputs audio according to instructions from the processor 46. The camera 42 is a small digital camera equipped with an optical system such as a lens, aperture, and shutter, and an image sensor such as a CMOS (Complementary Metal-Oxide-Semiconductor) image sensor or a CCD (Charge Coupled Device) image sensor.

[0032] Communication interface 44 is connected to network 54. Communication interfaces 44 and 26 are responsible for the exchange of various types of information between processor 46 and processor 28 via network 54.

[0033] Figure 2 shows an example of the main functions of the data processing device 12 and the smart device 14.

[0034] As shown in Figure 2, in the data processing device 12, a specific processing is performed by the processor 28. A specific processing program 56 is stored in the storage 32. The specific processing program 56 is an example of a "program" related to the technology of this disclosure. The processor 28 reads the specific processing program 56 from the storage 32 and executes the read specific processing program 56 on the RAM 30. The specific processing is realized by the processor 28 operating as a specific processing unit 290 according to the specific processing program 56 executed on the RAM 30.

[0035] The storage 32 stores the data generation model 58 and the emotion identification model 59. The data generation model 58 and the emotion identification model 59 are used by the identification processing unit 290.

[0036] In the smart device 14, the processor 46 performs the reception output processing. The storage 50 stores the reception output program 60. The reception output program 60 is used in conjunction with a specific processing program 56 by the data processing system 10. The processor 46 reads the reception output program 60 from the storage 50 and executes the read reception output program 60 on the RAM 48. The reception output processing is realized by the processor 46 operating as a control unit 46A according to the reception output program 60 executed on the RAM 48.

[0037] Next, the specific processing performed by the specific processing unit 290 of the data processing device 12 will be described. In the following description, the data processing device 12 will be referred to as the "server" and the smart device 14 as the "terminal".

[0038] This invention is a system for efficiently processing user requests entered in voice or text format and improving the quality of concierge services. Specific embodiments of the system are described below.

[0039] How users can enter their requests

[0040] The process begins with the user entering their request using a device (e.g., a smartphone or PC). Users can enter their request via voice or text. For example, a user might voice-input, "I would like to make a dinner reservation at a high-end restaurant in Tokyo next Saturday night."

[0041] Converting audio data to text

[0042] The device sends the user's voice data to the server. The server uses speech recognition technology to convert the voice data into text data. For example, it might generate the text "I would like to make a dinner reservation at a high-end restaurant in Tokyo next Saturday night" from the voice data.

[0043] Request analysis

[0044] The server analyzes the generated text data using a natural language processing algorithm and extracts specific elements representing the user's request. This yields information such as "Date: Next Saturday," "Time: Evening," "Location: Tokyo," "Genre: High-end restaurant," and "Content: Dinner reservation."

[0045] Generating example answers

[0046] The server uses generative AI to generate example answers based on the extracted elements. For example, the AI ​​might generate an example answer such as, "We suggest the following restaurants: 1. Restaurant A (reservations available) 2. Restaurant B (waitlist)."

[0047] Concierge confirmation

[0048] The server sends the generated sample response to the concierge's terminal, where the concierge reviews and corrects the content. The concierge then actually contacts, for example, "Restaurant A" and "Restaurant B" to check the latest reservation status and corrects the sample response as needed.

[0049] Submit your final response

[0050] The revised response is sent back to the server, which then sends the final response to the user's device. The user can receive the final response via email or app notification. For example, the notification might say, "Here are some high-end restaurants available next Saturday evening. Restaurant A is fully booked, while Restaurant B is on the waiting list."

[0051] Specific example

[0052] The user speaks into their smartphone's microphone and says, "I'd like to make a dinner reservation at a high-end restaurant in Tokyo next Saturday night."

[0053] The device sends the voice data to the server. The server performs speech recognition and obtains text data that reads, "I would like to make a dinner reservation at a high-end restaurant in Tokyo next Saturday night."

[0054] The server analyzes the text data using a natural language processing algorithm and extracts requests such as "next Saturday," "evening," "in Tokyo," "high-end restaurant," and "dinner reservation."

[0055] The server uses generative AI to generate a list of restaurants that match the user's request. For example, "Restaurant A (Reservations Available)" and "Restaurant B (Waiting List)".

[0056] This list is sent to the concierge terminal, where the concierge reviews it and makes corrections as needed.

[0057] The revised response is sent back to the server, which then sends the final response to the user's device. The user is notified in a format such as, "The only upscale restaurants available next Saturday evening are Restaurant C (reservations available) and Restaurant B (waitlist)."

[0058] In this way, the system can respond quickly and accurately to user requests while reducing the burden on concierges.

[0059] The following describes the processing flow.

[0060] Step 1:

[0061] The user uses their device to input their request in voice or text format. For example, the user might say into their smartphone's microphone, "I'd like to make a dinner reservation at a high-end restaurant in Tokyo next Saturday night."

[0062] Step 2:

[0063] The terminal sends the input voice data to the server. The voice data is transmitted in digital format.

[0064] Step 3:

[0065] The server uses speech recognition technology to convert the received audio data into text data. For example, it might produce text data such as, "I would like to make a dinner reservation at a high-end restaurant in Tokyo next Saturday night."

[0066] Step 4:

[0067] The server uses a natural language processing algorithm to analyze the generated text data. This analysis extracts elements such as "next Saturday," "evening," "in Tokyo," "high-end restaurant," and "dinner reservation."

[0068] Step 5:

[0069] The server uses generative AI to generate example answers based on the extracted elements. For example, the AI ​​might generate an example answer such as, "We suggest the following restaurants: 1. Restaurant A (reservations available) 2. Restaurant B (waitlist)."

[0070] Step 6:

[0071] The server sends the generated example answer to the concierge terminal. The concierge terminal displays the received content.

[0072] Step 7:

[0073] The concierge reviews the sample responses received and makes corrections as needed. For example, the concierge will actually contact "Restaurant A" and "Restaurant B" to check the latest reservation status and update the list if necessary.

[0074] Step 8:

[0075] The concierge sends the revised answer back to the server. The server verifies the revised content.

[0076] Step 9:

[0077] The server sends the final response to the user's terminal. For example, the user might be notified that "The only upscale restaurants available next Saturday evening are Restaurant C (reservations available) and Restaurant B (waitlist)."

[0078] Through the steps outlined above, this system efficiently responds to user requests, reduces the burden on concierges, and improves service quality.

[0079] (Example 1)

[0080] Next, we will describe Example 1. In the following description, the data processing device 12 will be referred to as the "server," and the smart device 14 will be referred to as the "terminal."

[0081] Traditional concierge services have faced challenges in quickly and accurately understanding user requests and generating appropriate responses. In particular, voice input requires significant time and manpower for the process of converting speech to text and then understanding that text. Furthermore, maintaining the quality of the generated responses necessitates manual review and correction by concierges, which further reduces efficiency.

[0082] The identification process performed by the identification processing unit 290 of the data processing device 12 in Example 1 is realized by the following means.

[0083] In this invention, the server includes a user terminal for the user to input requests in voice or text format, a conversion means for converting voice data received from the user into text data, an analysis means for analyzing the text data generated by the conversion means as a request using a natural language processing algorithm, a generation means for generating example answers using generative artificial intelligence based on the requests obtained by the analysis means, a transmission means for sending the example answers generated by the generation means to an administrator terminal for confirmation and correction by the administrator, a transmission means for sending the corrected answers returned from the administrator terminal to the user terminal, and a computer for automatically generating information that matches the user's requests. This makes it possible to quickly and accurately understand the user's requests and efficiently provide appropriate answers.

[0084] A "user terminal" is a device used by users to input requests in voice or text format, and specifically includes smartphones and personal computers.

[0085] "Conversion means" refers to a function or device for converting audio data received from a user into text data, and includes those that utilize speech recognition technology.

[0086] "Analysis means" refers to a function or device for analyzing text data generated by the conversion means as a request using a natural language processing algorithm.

[0087] "Generation means" refers to a function or device for generating example answers using generative artificial intelligence based on requests obtained by analysis means.

[0088] "Transmission means" refers to a function or device for transmitting example answers and modified answers generated by the generation means to administrator terminals and user terminals.

[0089] An "administrator terminal" is a device used by administrators to review and modify generated sample answers, and specifically includes personal computers and tablets.

[0090] An "electronic computing device" is a device that performs functions and processes to automatically generate information that meets the user's requirements.

[0091] System Configuration

[0092] This invention is a system for efficiently processing user requests entered in voice or text format and improving the quality of concierge services. This system includes the following elements:

[0093] 1. User terminal:

[0094] These are devices that allow users to input requests in voice or text format, such as smartphones and personal computers.

[0095] 2. Server:

[0096] It has a conversion means for converting audio data received from a user into text data.

[0097] We will use Google® Cloud Speech-to-Text API as our speech recognition technology.

[0098] The system includes an analysis means for analyzing text data generated by a conversion means as a request using a natural language processing algorithm.

[0099] We will use spaCy or the NLTK library as natural language processing algorithms.

[0100] It has a generation means for generating example answers using generative artificial intelligence based on the request.

[0101] As a generative artificial intelligence, we will use the GPT-3® model from OpenAI®.

[0102] The system has a means for sending the generated sample answers to the administrator's terminal for review and correction by the administrator.

[0103] It has a means for sending the corrected answer to the user's terminal.

[0104] 3. Administrator terminal:

[0105] This device allows administrators to review and correct generated sample answers, and examples include personal computers and tablets.

[0106] Operation details

[0107] First, the user uses a device (smartphone or personal computer) to input their request in voice or text format. For example, the user might say into their smartphone's microphone, "I would like to make a dinner reservation at a high-end restaurant in Tokyo next Saturday night."

[0108] Next, the device sends the audio data to the server. The server uses the Google Cloud Speech-to-Text API to convert the audio data into text data, obtaining the text "I would like to make a dinner reservation at a high-end restaurant in Tokyo next Saturday night."

[0109] The server then analyzes the text data using a natural language processing algorithm (for example, the spaCy library) to extract requests such as "next Saturday," "evening," "in Tokyo," "high-end restaurant," and "dinner reservation."

[0110] The server then generates example answers based on the elements extracted using OpenAI's GPT-3 model. An example of the prompt text used here is as follows:

[0111] "The user said, 'I want to make a dinner reservation at a high-end restaurant in Tokyo next Saturday night.' Please create a list of restaurant suggestions based on this."

[0112] The generated example responses will be in the format of "Restaurant A (Reservations Available)" or "Restaurant B (Waiting List)". The server sends these example responses to the administrator's terminal, where the administrator actually contacts the restaurants to check the latest reservation status and corrects the example responses as needed.

[0113] The revised response is sent back to the server, and the final response is sent to the user's device (smartphone or personal computer). The user can receive the final response via email or app notification, such as, "The only upscale restaurants available next Saturday evening are Restaurant C (reservations available) and Restaurant B (waitlist)."

[0114] In this way, the system can respond quickly and accurately to user requests while reducing the burden on concierges.

[0115] The flow of the specific processing in Example 1 will be explained using Figure 11.

[0116] Step 1:

[0117] Users use their devices (smartphones or personal computers) to input their requests in voice or text format.

[0118] Specific action: The user speaks into their smartphone's microphone, saying, "I would like to make a dinner reservation at a high-end restaurant in Tokyo next Saturday night." Alternatively, they type a similar request as text.

[0119] Input: User's request in voice or text format.

[0120] Output: Audio or text data generated within the device.

[0121] Step 2:

[0122] The device sends the user's voice data to the server.

[0123] Specific operation: The smartphone app uploads the recorded audio file to the server. Text files are sent as is.

[0124] Input: User's voice data or text data.

[0125] Output: Audio or text data sent to the server.

[0126] Step 3:

[0127] The server converts the received audio data into text data using speech recognition technology.

[0128] Specific operation: The server calls the Google Cloud Speech-to-Text API and converts the audio data into text, "I would like to make a dinner reservation at a high-end restaurant in Tokyo next Saturday night."

[0129] Input: Audio data.

[0130] Output: Text data.

[0131] Step 4:

[0132] The server analyzes the text data using a natural language processing algorithm to extract the specific elements of the request.

[0133] Specific operation: The server uses the spaCy library to parse the text and extract elements such as "next Saturday," "evening," "in Tokyo," "high-end restaurant," and "dinner reservation."

[0134] Input: Text data.

[0135] Output: Extracted elements ("Date: Next Saturday", "Time: Evening", "Location: Tokyo", "Genre: High-end restaurant", "Content: Dinner reservation").

[0136] Step 5:

[0137] The server uses generative artificial intelligence to generate example answers based on the extracted elements.

[0138] Specific operation: The server inputs the following prompt to the OpenAI GPT-3 model: "The user said, 'I would like to make a dinner reservation at a high-end restaurant in Tokyo next Saturday night.' Please create a list of restaurant suggestions based on this."

[0139] Input: Extracted elements.

[0140] Output: Generated example answers (e.g., "Restaurant A (Reservations Available)", "Restaurant B (Waiting List)").

[0141] Step 6:

[0142] The server sends the generated sample answers to the administrator's terminal, where the administrator reviews and corrects the content.

[0143] Specific operation: The server sends a sample response to the administrator's terminal, and the administrator checks the latest reservation status with the pre-designated restaurant by phone or web and modifies the sample response as needed.

[0144] Input: Generated example answer.

[0145] Output: Example answer corrected by the administrator.

[0146] Step 7:

[0147] The server sends the corrected answer example to the user's terminal.

[0148] Specific operation: The server sends the corrected answer example to the user's smartphone or personal computer via email or app notification.

[0149] Input: Example answer corrected by the administrator.

[0150] Output: The final response sent to the user (e.g., "The only upscale restaurants available next Saturday evening are Restaurant C (reservations available) and Restaurant B (waitlist)").

[0151] (Application Example 1)

[0152] Next, we will explain Application Example 1. In the following explanation, the data processing device 12 will be referred to as the "server," and the smart device 14 will be referred to as the "terminal."

[0153] In modern brick-and-mortar stores, customers often demand prompt and accurate product guidance and recommendations, but there is a lack of high-quality and efficient service to meet this demand. Furthermore, manual service is costly and places a heavy burden on employees, leading to decreased customer satisfaction. In particular, while there is a growing demand for interactive service delivery using smart devices, appropriate systems are still lacking.

[0154] The specific processing performed by the specific processing unit 290 of the data processing device 12 in Application Example 1 is realized by the following means.

[0155] In this invention, the server includes a user terminal for the user to input requests in voice or text format; a conversion means for converting voice data received from the user into text data; an analysis means for analyzing the text data generated by the conversion means as a request using a natural language processing algorithm; a generation means for generating example answers using a generative AI based on the requests obtained by the analysis means; a transmission means for sending the example answers generated by the generation means to a concierge terminal for confirmation and correction by the concierge; a transmission means for sending the corrected answers returned from the concierge terminal to the user terminal; and a notification means that allows the user terminal to input requests via smart glasses and notifies the user of the final answer via the smart glasses. This makes it possible for customers in physical stores to receive quick and appropriate product information through voice input.

[0156] A "user terminal" is a device used by users to input requests in voice or text format.

[0157] "Conversion means" refers to a device or method for converting audio data received from a user into text data.

[0158] "Analysis means" refers to a device or method for analyzing text data generated by a conversion means as a request using a natural language processing algorithm.

[0159] "Generation means" refers to an apparatus or method for generating example answers using a generative AI based on requests obtained by an analysis means.

[0160] "Transmission means" refers to a device or method for transmitting example answers generated by the generation means to a concierge terminal, and for transmitting corrected answers returned from the concierge terminal to a user terminal.

[0161] "Smart glasses" are devices that users wear to input requests and receive responses visually.

[0162] "Notification means" refers to a device or method for notifying a user of the final response to a request through a user terminal (including smart glasses).

[0163] Modes for carrying out the invention

[0164] This invention is a system for providing customers with quick and appropriate product information in physical stores. Specific embodiments are described below.

[0165] System Program Overview

[0166] 1. Voice input and data conversion:

[0167] The user performs voice input while wearing smart glasses. For example, they might voice a request such as, "I want to find a shirt that matches these shoes."

[0168] The smart glasses incorporate voice recognition technology, converting received voice data into text data. This process utilizes Google's voice recognition API.

[0169] 2. Analysis of requests:

[0170] The converted text data is sent to the server. On the server side, a natural language processing algorithm (BERT, a transformer model from Hugging Face) is used to analyze the text data and extract the customer's request as specific elements. For example, information such as "Item: Shoes," "Request: Shirt," and "Purpose: Coordination support" is obtained.

[0171] 3. Generating example answers:

[0172] Based on the extracted elements, a generative AI (e.g., GPT-3) is used to generate example answers. The generated example answers include specific product information, such as "the blue shirt on shelf A" or "the white shirt on shelf B." An example of this prompt is: "User request: I want to find a shirt that matches these shoes. Please suggest specific products that match this request."

[0173] 4. Concierge review and correction:

[0174] The generated sample responses are sent to the concierge terminal. The concierge reviews the list of suggestions and makes revisions as needed. For example, they may update the list to reflect inventory status or additional product information, and then send the revised responses back to the server.

[0175] 5. Notification of final response:

[0176] The server notifies the user of the corrected final answer via the smart glasses they are wearing. This allows the user to receive the notification visually. For example, the guidance might say, "The shirts that match these shoes are the blue shirt on shelf A and the white shirt on shelf B."

[0177] Hardware and software to be used

[0178] Smart glasses: A device for voice input and notification display.

[0179] Server: The central hub responsible for converting audio data to text, analyzing text data, and managing generated response examples. It runs Google's speech recognition API, Hugging Face's BERT, and generative AI (GPT-3).

[0180] Natural Language Processing (NLP): These algorithms run on a server and are used to analyze and extract user requests.

[0181] Adding specific examples

[0182] For example, if a user enters a store and uses smart glasses to voice-input "I want to find a shirt that goes with these shoes," they will receive a notification a few seconds later saying, "We recommend the blue shirt on shelf A and the white shirt on shelf B." This experience allows users to quickly find products that meet their needs.

[0183] In this way, this system can improve the customer experience in physical stores and provide high-quality service while reducing the burden on employees.

[0184] The flow of a specific process in Application Example 1 will be explained using Figure 12.

[0185] Step 1:

[0186] The user wears smart glasses and inputs their request by voice. For example, they might say, "I want to find a shirt that matches these shoes." This input is collected via the smart glasses' microphone. The voice signal is then used as input for speech recognition processing to convert it into text.

[0187] Step 2:

[0188] Smart glasses convert audio data into text data. Specifically, they use Google's speech recognition API to convert audio data into text data. The input to this process is audio data, and the output is text data.

[0189] Step 3:

[0190] The converted text data is sent to the server. The server analyzes the received text data using a natural language processing algorithm (Hugging Face's BERT model). In this step, the input is text data, and the output is the specific elements of the analyzed request (e.g., "Item: Shoes", "Request: Shirt", "Purpose: Coordination Support").

[0191] Step 4:

[0192] The server uses a generative AI (GPT-3) to generate example responses based on the elements of the analyzed request. The input used in this process is the specific elements of the analyzed request, and the output is a specific product suggestion. For example, example responses might be "a blue shirt on shelf A" or "a white shirt on shelf B."

[0193] Step 5:

[0194] The generated sample responses are sent to the concierge terminal. The concierge reviews the list of suggestions and makes corrections as needed. Based on inventory information and new product information, they create more accurate suggestions. The input in this step is the generated sample response, and the output is the corrected sample response.

[0195] Step 6:

[0196] The revised answer examples are sent back to the server and compiled into the final answer. This final answer is then sent to the smart glasses worn by the user. For example, the notification might say, "The shirts that go with these shoes are the blue shirt on shelf A and the white shirt on shelf B." In this step, the input is the revised answer examples, and the output is the notification message displayed on the smart glasses.

[0197] In this way, the input data is sequentially processed and analyzed at each processing step, and finally, appropriate product information is provided to the user.

[0198] Furthermore, an emotion engine that estimates the user's emotions may be incorporated. That is, the identification processing unit 290 may use the emotion identification model 59 to estimate the user's emotions and perform identification processing using the user's emotions.

[0199] This invention is a system for efficiently processing user requests in voice or text format and further improving the quality of concierge services by combining it with an emotion engine that recognizes the user's emotions. Specific embodiments of the system are shown below.

[0200] How users can enter their requests

[0201] The process begins with the user entering their request using a device (e.g., a smartphone or PC). Users can enter their request via voice or text. For example, a user might voice-input, "I would like to make a dinner reservation at a high-end restaurant in Tokyo next Saturday night."

[0202] Converting audio data to text

[0203] The device sends the user's voice data to the server. The server uses speech recognition technology to convert the voice data into text data. For example, it might generate the text "I would like to make a dinner reservation at a high-end restaurant in Tokyo next Saturday night" from the voice data.

[0204] Request analysis

[0205] The server analyzes the generated text data using a natural language processing algorithm and extracts specific elements representing the user's request. This yields information such as "Date: Next Saturday," "Time: Evening," "Location: Tokyo," "Genre: High-end restaurant," and "Content: Dinner reservation."

[0206] Using an Emotion Engine

[0207] The server then uses an emotion engine to analyze the user's emotions from voice or text data. For example, it extracts emotional information such as "excited," "anxious," or "hurried." This emotional information influences the generative AI in subsequent processes, which is used to generate more appropriate responses that align with the user's feelings.

[0208] Generating example answers

[0209] The server uses generative AI to generate example responses based on extracted elements and emotional information. For example, the AI ​​might generate a response like, "We suggest the following restaurants: 1. Restaurant A (reservations available) 2. Restaurant B (waitlist)." By considering emotional information, it's possible to include additional perks for users who are excited, for example.

[0210] Concierge confirmation

[0211] The server sends the generated sample response to the concierge's terminal, where the concierge reviews and corrects the content. The concierge then contacts, for example, "Restaurant A" and "Restaurant B" to check the latest reservation status and updates the list if necessary.

[0212] Submit your final response

[0213] The revised response is sent back to the server, which then sends the final response to the user's device. The user can receive the final response via email or app notification. For example, the notification might say, "Here are some high-end restaurants available next Saturday evening. Restaurant A is fully booked, while Restaurant B is on the waiting list."

[0214] Specific example

[0215] The user speaks into their smartphone's microphone and says, "I'd like to make a dinner reservation at a high-end restaurant in Tokyo next Saturday night."

[0216] The device sends the voice data to the server. The server performs speech recognition and obtains text data that reads, "I would like to make a dinner reservation at a high-end restaurant in Tokyo next Saturday night."

[0217] The server analyzes the text data using a natural language processing algorithm and extracts requests such as "next Saturday," "evening," "in Tokyo," "high-end restaurant," and "dinner reservation."

[0218] The server uses an emotion engine to extract emotional information from the user's voice and text data, recognizing emotions such as "excited" or "hurried."

[0219] The server uses generative AI to generate a list based on the user's requests and recognized emotions. For example, "Restaurant A (Reservations Available)" and "Restaurant B (Waiting List)." Depending on the emotion, for instance, if the emotion is excitement, the system will generate a response that includes additional information about special offers.

[0220] This list is sent to the concierge terminal, where the concierge reviews it and makes corrections as needed.

[0221] The revised response is sent back to the server, which then sends the final response to the user's device. The user is notified that "The only upscale restaurants available next Saturday evening are Restaurant A (reservations available) and Restaurant B (waitlist)."

[0222] This system responds quickly and accurately to user requests and provides more personalized service by taking emotions into consideration. This reduces the burden on concierges and improves the quality of service.

[0223] The following describes the processing flow.

[0224] Step 1:

[0225] The user uses their device to input their request in voice or text format. For example, the user might say into their smartphone's microphone, "I'd like to make a dinner reservation at a high-end restaurant in Tokyo next Saturday night."

[0226] Step 2:

[0227] The terminal sends the input voice data to the server. The voice data is transmitted in digital format.

[0228] Step 3:

[0229] The server uses speech recognition technology to convert the received audio data into text data. For example, it might produce text data such as, "I would like to make a dinner reservation at a high-end restaurant in Tokyo next Saturday night."

[0230] Step 4:

[0231] The server uses a natural language processing algorithm to analyze the generated text data. This analysis extracts elements such as "next Saturday," "evening," "in Tokyo," "high-end restaurant," and "dinner reservation."

[0232] Step 5:

[0233] The server uses an emotion engine to analyze the user's emotions from voice or text data. For example, the emotion engine extracts emotional information such as "excited," "anxious," or "hurried."

[0234] Step 6:

[0235] The server uses generative AI to generate example responses based on emotional information and desired elements. By considering emotional information, more personalized responses are generated. For example, the AI ​​might generate a response like, "We suggest the following restaurants: 1. Restaurant A (reservations available) 2. Restaurant B (waitlist)." If the user is looking forward to something, special offers may also be added.

[0236] Step 7:

[0237] The server sends the generated example answer to the concierge terminal. The concierge terminal displays the received content.

[0238] Step 8:

[0239] The concierge reviews the sample responses received and makes corrections as needed. For example, the concierge will actually contact "Restaurant A" and "Restaurant B" to check the latest reservation status and update the list if necessary.

[0240] Step 9:

[0241] The concierge sends the revised answer back to the server. The server verifies the revised content.

[0242] Step 10:

[0243] The server sends the final response to the user's terminal. For example, the user might be notified that "The only upscale restaurants available next Saturday evening are Restaurant A (reservations available) and Restaurant B (waitlist)."

[0244] Through the steps outlined above, this system can efficiently respond to user requests and, by taking user emotions into consideration, improve the quality of concierge services.

[0245] (Example 2)

[0246] Next, we will describe Example 2. In the following description, the data processing device 12 will be referred to as the "server" and the smart device 14 as the "terminal".

[0247] Traditional concierge services lacked the technology to appropriately and quickly handle specific user requests, and particularly struggled to respond while considering emotions. Therefore, improving service quality requires analyzing user requests, including their emotions, and generating appropriate responses based on that analysis.

[0248] The specific processing performed by the specific processing unit 290 of the data processing device 12 in Example 2 is realized by the following means. In this invention, the server includes an information terminal for the user to input requests in voice or text format, a conversion means for converting voice data received from the user into text data, an analysis means for analyzing the text data generated by the conversion means using a natural language processing algorithm, a generation means for generating example answers using generative artificial intelligence based on the requests and sentiment information obtained by the analysis means, a transmission means for sending the example answers generated by the generation means to a service provider terminal for confirmation and correction by the service provider, and a transmission means for sending the corrected answers returned from the service provider terminal to the user terminal. This enables the provision of quick and appropriate services that take into account the user's requests and sentiments.

[0249] An "information terminal" is a device used by users to input requests in voice or text format, and includes smartphones and personal computers.

[0250] "Conversion means" refers to a technology or method for converting audio data received from a user into text data, and which uses speech recognition technology.

[0251] "Analysis means" refers to a technology or method that analyzes text data generated by the conversion means using a natural language processing algorithm to extract user requests and emotional information.

[0252] "Generation means" refers to a technology or method that generates example responses using generative artificial intelligence based on requests and emotional information obtained by analysis means.

[0253] "Transmission means" refers to a technology or method for transmitting example answers generated by the generation means to a service provider terminal, and for transmitting the corrected answers to a user terminal.

[0254] A "service provider terminal" refers to a device used to review and correct generated sample answers, specifically the device used by the concierge.

[0255] A "natural language processing algorithm" is an algorithm used to analyze text data and extract information about desires and emotions, and it includes machine learning and data analysis techniques.

[0256] "Generative artificial intelligence" refers to artificial intelligence technology that generates example responses based on analyzed requests and sentiment information, such as using a generative language model.

[0257] "Emotional information" refers to emotional information analyzed from the user's input voice or text data, and includes emotions such as excitement, anxiety, and urgency.

[0258] Modes for carrying out the invention

[0259] This invention is a system for efficiently processing user requests in voice or text format and further improving the quality of concierge services by combining it with an emotion engine that recognizes the user's emotions. Specific embodiments of the system are shown below.

[0260] How users can enter their requests

[0261] The process begins with the user entering their request in voice or text format using an information terminal (such as a smartphone or personal computer). For example, a user could voice-input a request like, "I would like to make a dinner reservation at a high-end restaurant in Tokyo next Saturday night." Users can input voice data using the terminal's microphone or input their request as text data using a keyboard.

[0262] Converting audio data to text

[0263] The device sends the user's voice data to the server. The server uses speech recognition technology to convert the voice data into text data. Specifically, it uses a speech recognition service such as the Google Cloud Speech-to-Text API to convert the voice data into text data. For example, it generates text data such as "I would like to make a dinner reservation at a high-end restaurant in Tokyo next Saturday night" from the voice data.

[0264] Request analysis

[0265] The server analyzes the generated text data using natural language processing algorithms to extract specific elements representing the user's request. Algorithms used include SpaCy and the Google Cloud Natural Language API. This analysis yields information such as "Date: Next Saturday," "Time: Evening," "Location: Tokyo," "Genre: High-end restaurant," and "Content: Dinner reservation."

[0266] Using an Emotion Engine

[0267] The server further utilizes an emotion engine to analyze the user's emotions from voice or text data. For example, it uses IBM Watson® Tone Analyzer to extract emotional information such as "excited," "anxious," and "hurried." This emotional information influences the generative artificial intelligence in subsequent processes, which is used to generate more appropriate responses that align with the user's feelings.

[0268] Generating example answers

[0269] The server generates example responses using generative artificial intelligence (e.g., OpenAI GPT-3) based on extracted elements and emotional information. For example, the generative AI might generate a response such as, "We suggest the following restaurants: 1. Restaurant A (reservations available) 2. Restaurant B (waitlist)." By considering emotional information, it's possible to include additional perks for users who are excited, for example.

[0270] Concierge confirmation

[0271] The server sends the generated sample response to the service provider's terminal, where the service provider reviews and corrects the content. The service provider (e.g., a concierge) then actually contacts restaurants such as "Restaurant A" and "Restaurant B" to check their latest reservation status and updates the list if necessary.

[0272] Submit your final response

[0273] The revised response from the service provider's terminal is sent back to the server, which then sends the final response to the user's information terminal. The user can receive the final response via email or app notification. For example, the notification might say, "We have a list of upscale restaurants available next Saturday evening. Restaurant A is fully booked, while Restaurant B is on the waiting list."

[0274] Examples of prompt statements

[0275] If a user speaks into their smartphone's microphone and says, "I'd like to make a dinner reservation at a high-end restaurant in Tokyo next Saturday night," the following prompt will be used for the generative artificial intelligence.

[0276] "User request: "I want to make a dinner reservation at a high-end restaurant in Tokyo next Saturday night." Sentiment information: "Excited, in a hurry." Please generate a list of responses: Example) - Restaurant A (Reservations available) - Restaurant B (Waiting list) Perks information: Add perks information based on the excited sentiment."

[0277] This allows the system to respond quickly and accurately to user requests and provide more personalized service by taking emotions into consideration. In this way, it is possible to reduce the burden on concierges and improve the quality of service.

[0278] The flow of the specific process in Example 2 will be described with reference to FIG. 13.

[0279] Step 1: Input of requests

[0280] User

[0281] The user uses an information terminal (e.g., a smartphone, a personal computer) to input requests in voice or text form.

[0282] Input: Voice data or text data

[0283] Specific operation: For example, the user says to the smartphone, "I want to make a dinner reservation at a high-class restaurant in the city on Saturday night next week."

[0284] Step 2: Conversion of voice data to text

[0285] Terminal

[0286] The terminal sends the voice data to the server.

[0287] Input: Voice data

[0288] Output: Sending of voice data to the server

[0289] Specific operation: Send the voice data to the server in digital form.

[0290] Server

[0291] The server uses speech recognition technology such as Google Cloud Speech-to-Text API to convert the voice data into text data.

[0292] Input: Voice data

[0293] Output: Text data

[0294] Specific operation: Generate the text "I want to make a dinner reservation at a high-class restaurant in Tokyo next Saturday night" from voice data.

[0295] Step 3: Analysis of requests

[0296] Server

[0297] The server analyzes the generated text data using natural language processing algorithms (e.g., SpaCy, Google Cloud Natural Language API) and extracts specific elements of the request.

[0298] Input: Text data

[0299] Output: Elements of the request (e.g., "Date: Next Saturday", "Time: Night", "Location: Tokyo", "Genre: High-class restaurant", "Content: Dinner reservation")

[0300] Specific operation: Analyze the specific elements of the request and store them in the database.

[0301] Step 4: Use of the emotion engine

[0302] Server

[0303] The server uses an emotion engine (e.g., IBM Watson Tone Analyzer) to analyze emotion information from voice data or text data.

[0304] Input: Voice data or text data

[0305] Output: Emotion information (e.g., "Happy", "Anxious", "In a hurry")

[0306] Specific operation: Use the emotion engine to analyze the emotion from voice or text and store it in the database.

[0307] Step 5: Generation of example answers

[0308] server

[0309] The server generates example responses using a generative AI (e.g., OpenAI GPT-3) based on the elements of the request and emotional information.

[0310] Input: Elements of requests, emotional information

[0311] Output: Example answer (Example: "Restaurant A (Reservations accepted)", "Restaurant B (Waiting list)")

[0312] Specific operation: Use generative AI to generate responses that respond to requests and emotions.

[0313] Step 6: Concierge Confirmation

[0314] server

[0315] The server sends the generated example answers to the service provider's terminal.

[0316] Input: Example Answer

[0317] Output: Sent to the service provider's terminal

[0318] Specific action: Send example answers to the service provider's terminal.

[0319] Service provider

[0320] The service provider will review the responses and make corrections as necessary.

[0321] Input: Example Answer

[0322] Output: Revised answer example

[0323] Specific action: The service provider inquires about the latest booking status and corrects the response.

[0324] Step 7: Submit your final response

[0325] Service provider

[0326] The revised answer example is sent back to the server.

[0327] Input: Revised answer example

[0328] Output: Send to server

[0329] Specific action: Send the corrected response to the server.

[0330] server

[0331] The server sends the final response to the user's information terminal.

[0332] Input: Revised answer example

[0333] Output: Sending the final response to the user's information terminal.

[0334] Specific action: Send the final response to the user via email or app notification.

[0335] In this way, by responding quickly and accurately to user requests and taking their emotions into consideration, it becomes possible to provide more personalized services.

[0336] (Application Example 2)

[0337] Next, we will explain application example 2. In the following explanation, the data processing device 12 will be referred to as the "server," and the smart device 14 will be referred to as the "terminal."

[0338] Traditional concierge services faced challenges in providing prompt service due to the extensive work required to accurately understand and incorporate user requests. Furthermore, they lacked consideration for user emotions, resulting in a lack of personalization and low satisfaction. In short, the uniform approach to users and the inability to provide personalized suggestions based on emotions led to a poor user experience.

[0339] In Application Example 2, the specific processing performed by the specific processing unit 290 of the data processing device 12 is realized by the following means. In this invention, the server includes a user terminal for the user to input requests in voice or text format, a conversion means for converting voice data received from the user into text data, an analysis means for analyzing the text data generated by the conversion means as a request using a natural language processing algorithm, a generation means for generating example answers using a generative AI based on the requests and emotional information from the analysis means, a transmission means for sending the example answers generated by the generation means to a concierge terminal for confirmation and correction by the concierge, a transmission means for sending the corrected answers returned from the concierge terminal to the user terminal, and an emotional analysis means for extracting emotional information from the user's voice data and text data. This makes it possible to quickly and accurately analyze user requests and provide personalized services based on emotions.

[0340] A "user terminal" is a device used by users to input requests in voice or text format.

[0341] "Conversion means" refers to technical means for converting audio data received from a user into text data.

[0342] "Analysis means" refers to technical means for analyzing text data generated by the conversion means as a request using a natural language processing algorithm.

[0343] "Generation means" refers to technical means for generating example responses using a generative AI based on requests and emotional information obtained through analysis means.

[0344] "Transmission means" refers to a technical means for sending example answers generated by the generation means to a concierge terminal, for confirmation and correction by the concierge, and for sending the corrected answers returned from the concierge terminal to the user terminal.

[0345] "Emotion analysis methods" refer to technical means for extracting emotional information from user voice data and text data.

[0346] A "natural language processing algorithm" is an algorithm that analyzes text data to understand and extract user requests.

[0347] "Generative AI" is an artificial intelligence technology that generates appropriate response examples based on user requests and emotional information.

[0348] A "concierge terminal" is a device used to review and correct generated sample answers.

[0349] "Emotional information" refers to information about a user's emotions extracted from their voice data and text data.

[0350] This invention is a system for further improving the quality of concierge services by efficiently processing requests entered by users in voice or text format, and by combining this with an emotion engine that recognizes the user's emotions.

[0351] The system includes the following components:

[0352] 1. User terminal: A device used by users to input requests in voice or text format. This includes smartphones, PCs, and smart glasses.

[0353] 2. Conversion method: Speech recognition technology is used to convert the audio data received from the user into text data. Specifically, the SpeechRecognition library is used to convert the audio data into text data.

[0354] 3. Analysis Method: The text data generated by the transformation method is analyzed as a request using a natural language processing algorithm. This uses the natural language processing model from the Transformers library.

[0355] 4. Emotion Analysis Method: An emotion analysis model is used to extract emotional information from the user's voice data and text data. Specifically, the Transformers emotion analysis model is used to obtain emotional information.

[0356] 5. Generation Method: Based on the requests and sentiment information obtained through the analysis method, a generative AI model (such as OpenAI GPT-3) is used to generate example responses. This process uses OpenAI's GPT-3 as an API and takes into account the relationship between the prompt text and the generated text.

[0357] 6. Transmission method: The example answer generated by the generation method is sent to the concierge terminal for confirmation and correction by the concierge. Furthermore, the corrected answer returned from the concierge terminal is sent to the user terminal.

[0358] Explanation of specific processing examples:

[0359] The user inputs their request by voice into their smartphone's microphone, saying, "I would like to make a dinner reservation at a high-end restaurant in Tokyo next Saturday night."

[0360] The user's terminal sends the voice data to the server, which uses the SpeechRecognition library to convert the voice data into text data that reads, "I would like to make a dinner reservation at a high-end restaurant in Tokyo next Saturday night."

[0361] The converted text data is analyzed by Transformers' natural language processing model, and requests such as "Date: Next Saturday", "Time: Evening", "Location: Tokyo", "Genre: High-end restaurant", and "Content: Dinner reservation" are extracted.

[0362] An emotion analysis model analyzes the same text data and obtains emotional information such as "excited" or "hurried."

[0363] Based on the analyzed requests and sentiment information, the server uses GPT-3 to generate example responses with the following prompts:

[0364] User Input: I'm looking for recommended baby products in the store.

[0365] Emotion: Enjoyment

[0366] Response:

[0367] For example, the AI ​​can generate sample responses such as, "In this baby products section, we recommend our new skin-friendly diapers and soft swaddles."

[0368] The generated sample answers are sent to the concierge terminal, where the concierge reviews the content and makes corrections as needed.

[0369] The corrected response is sent back to the server and finally delivered to the user's terminal.

[0370] This system enables the rapid and accurate analysis of user requests and the provision of personalized, emotion-based services. This improves the user experience and reduces the workload on concierges.

[0371] The flow of a specific process in Application Example 2 will be explained using Figure 14.

[0372] Step 1:

[0373] The user performs voice input. The user uses a smartphone or smart glasses to input their request by voice. For example, they might say, "I would like to make a dinner reservation at a high-end restaurant in Tokyo next Saturday night." This voice data is recorded on the device. The input is voice data, and the output is the voice file recorded on the device.

[0374] Step 2:

[0375] The device sends audio data to the server. The device sends the recorded audio data to the server. The server uses speech recognition technology to convert this audio data into text data. The SpeechRecognition library is used for this conversion. The input is audio data, and the output is text data that says, "I would like to make a dinner reservation at a high-end restaurant in Tokyo next Saturday night."

[0376] Step 3:

[0377] The server analyzes the text data. The server then analyzes the transformed text data using a natural language processing algorithm. This process uses the Transformers library to extract user requests (e.g., "Date: Next Saturday", "Time: Evening", "Location: Tokyo", "Genre: High-end restaurant", "Content: Dinner reservation"). The input is text data, and the output is the analyzed elemental information.

[0378] Step 4:

[0379] The server extracts emotional information. The server then uses emotion analysis tools to analyze the user's emotional information from the text data. This process utilizes the Transformers emotion analysis model. For example, emotions such as "excited" and "hurried" are extracted. The input is text data, and the output is emotional information.

[0380] Step 5:

[0381] The server generates example answers using generative AI. Based on the analyzed requests and sentiment information, the server generates example answers using a generative AI model such as GPT-3. This process uses OpenAI's GPT-3 as an API, sending the generated text as a prompt. The input is requests and sentiment information, and the output is the suggested answer. For example, the prompt might look like this:

[0382] User Input: I'm looking for recommended baby products in the store.

[0383] Emotion: Enjoyment

[0384] Response:

[0385] Step 6:

[0386] The server sends example answers to the concierge terminal. The generated example answers are sent to the concierge terminal, where the concierge reviews the content and makes corrections as needed. The input is the suggested example answer, and the output is the corrected example answer.

[0387] Step 7:

[0388] The concierge terminal sends the corrected answer back to the server. After the concierge reviews and makes corrections, the corrected answer is sent back to the server. The input is an example of the corrected answer, and the output is the final example answer.

[0389] Step 8:

[0390] The server sends the final answer to the user's device. The user can receive the final answer via email or app notification. The input is an example of the final answer, and the output is the final answer notified to the user. For example, the user might be notified of the final answer, "The upscale restaurants available next Saturday evening are Restaurant A (reservations available) and Restaurant B (waitlist)."

[0391] The specific processing unit 290 transmits the result of the specific processing to the smart device 14. In the smart device 14, the control unit 46A causes the output device 40 to output the result of the specific processing. The microphone 38B acquires audio indicating user input for the result of the specific processing. The control unit 46A transmits the audio data indicating user input acquired by the microphone 38B to the data processing device 12. In the data processing device 12, the specific processing unit 290 acquires the audio data.

[0392] Data generation model 58 is a so-called generative AI (Artificial Intelligence). An example of data generation model 58 is ChatGPT (registered trademark) (Internet search).<URL: https: / / openai.com / blog / chatgpt> ), Gemini (registered trademark) (Internet search) <url: https: gemini.google.com ?hl="ja">Examples of generative AI include the following. The data generation model 58 is obtained by performing deep learning on a neural network. The data generation model 58 is input with prompts containing instructions, and with inference data such as audio data representing speech, text data representing text, and image data representing images. The data generation model 58 infers from the input inference data according to the instructions indicated by the prompts, and outputs the inference results in data formats such as audio data and text data. Here, inference refers to, for example, analysis, classification, prediction, and / or summarization.

[0393] In the above embodiment, an example was given in which specific processing is performed by the data processing device 12, but the technology of this disclosure is not limited thereto, and the specific processing may also be performed by the smart device 14.

[0394] [Second Embodiment]

[0395] Figure 3 shows an example of the configuration of the data processing system 210 according to the second embodiment.

[0396] As shown in Figure 3, the data processing system 210 includes a data processing device 12 and smart glasses 214. An example of the data processing device 12 is a server.

[0397] The data processing device 12 comprises a computer 22, a database 24, and a communication interface 26. The computer 22 is an example of a "computer" related to the technology of this disclosure. The computer 22 comprises a processor 28, RAM 30, and storage 32. The processor 28, RAM 30, and storage 32 are connected to a bus 34. The database 24 and the communication interface 26 are also connected to the bus 34. The communication interface 26 is connected to a network 54. An example of the network 54 is a WAN (Wide Area Network) and / or a LAN (Local Area Network).

[0398] The smart glasses 214 include a computer 36, a microphone 238, a speaker 240, a camera 42, and a communication interface 44. The computer 36 includes a processor 46, RAM 48, and storage 50. The processor 46, RAM 48, and storage 50 are connected to a bus 52. The microphone 238, speaker 240, and camera 42 are also connected to the bus 52.

[0399] The microphone 238 receives voice signals from the user 20 and receives instructions from the user 20. The microphone 238 captures the voice signals from the user 20, converts the captured voice into audio data, and outputs it to the processor 46. The speaker 240 outputs audio according to the instructions from the processor 46.

[0400] Camera 42 is a small digital camera equipped with an optical system including a lens, aperture, and shutter, and an image sensor such as a CMOS (Complementary Metal-Oxide-Semiconductor) image sensor or a CCD (Charge Coupled Device) image sensor, and captures images of the area around the user 20 (for example, an imaging range defined by a field of view equivalent to the width of a typical healthy person's field of vision).

[0401] Communication interface 44 is connected to network 54. Communication interfaces 44 and 26 are responsible for the exchange of various information between processor 46 and processor 28 via network 54. The exchange of various information between processor 46 and processor 28 using communication interfaces 44 and 26 is performed in a secure manner.

[0402] Figure 4 shows an example of the main functions of the data processing device 12 and the smart glasses 214. As shown in Figure 4, the data processing device 12 performs specific processing using the processor 28. The storage 32 stores the specific processing program 56.

[0403] The specific processing program 56 is an example of a "program" relating to the technology of this disclosure. The processor 28 reads the specific processing program 56 from the storage 32 and executes the read specific processing program 56 on the RAM 30. The specific processing is realized by the processor 28 operating as a specific processing unit 290 in accordance with the specific processing program 56 executed on the RAM 30.

[0404] The storage 32 stores the data generation model 58 and the emotion identification model 59. The data generation model 58 and the emotion identification model 59 are used by the identification processing unit 290.

[0405] In the smart glasses 214, the processor 46 performs the reception output processing. The storage 50 stores the reception output program 60. The processor 46 reads the reception output program 60 from the storage 50 and executes the read reception output program 60 on the RAM 48. The reception output processing is realized by the processor 46 operating as a control unit 46A according to the reception output program 60 executed on the RAM 48.

[0406] Next, the identification processing performed by the identification processing unit 290 of the data processing device 12 will be described. In the following description, the data processing device 12 will be referred to as the "server" and the smart glasses 214 will be referred to as the "terminal".

[0407] This invention is a system for efficiently processing user requests entered in voice or text format and improving the quality of concierge services. Specific embodiments of the system are described below.

[0408] How users can enter their requests

[0409] The process begins with the user entering their request using a device (e.g., a smartphone or PC). Users can enter their request via voice or text. For example, a user might voice-input, "I would like to make a dinner reservation at a high-end restaurant in Tokyo next Saturday night."

[0410] Converting audio data to text

[0411] The device sends the user's voice data to the server. The server uses speech recognition technology to convert the voice data into text data. For example, it might generate the text "I would like to make a dinner reservation at a high-end restaurant in Tokyo next Saturday night" from the voice data.

[0412] Request analysis

[0413] The server analyzes the generated text data using a natural language processing algorithm and extracts specific elements representing the user's request. This yields information such as "Date: Next Saturday," "Time: Evening," "Location: Tokyo," "Genre: High-end restaurant," and "Content: Dinner reservation."

[0414] Generating example answers

[0415] The server uses generative AI to generate example answers based on the extracted elements. For example, the AI ​​might generate an example answer such as, "We suggest the following restaurants: 1. Restaurant A (reservations available) 2. Restaurant B (waitlist)."

[0416] Concierge confirmation

[0417] The server sends the generated sample response to the concierge's terminal, where the concierge reviews and corrects the content. The concierge then actually contacts, for example, "Restaurant A" and "Restaurant B" to check the latest reservation status and corrects the sample response as needed.

[0418] Submit your final response

[0419] The revised response is sent back to the server, which then sends the final response to the user's device. The user can receive the final response via email or app notification. For example, the notification might say, "Here are some high-end restaurants available next Saturday evening. Restaurant A is fully booked, while Restaurant B is on the waiting list."

[0420] Specific example

[0421] The user speaks into their smartphone's microphone and says, "I'd like to make a dinner reservation at a high-end restaurant in Tokyo next Saturday night."

[0422] The device sends the voice data to the server. The server performs speech recognition and obtains text data that reads, "I would like to make a dinner reservation at a high-end restaurant in Tokyo next Saturday night."

[0423] The server analyzes the text data using a natural language processing algorithm and extracts requests such as "next Saturday," "evening," "in Tokyo," "high-end restaurant," and "dinner reservation."

[0424] The server uses generative AI to generate a list of restaurants that match the user's request. For example, "Restaurant A (Reservations Available)" and "Restaurant B (Waiting List)".

[0425] This list is sent to the concierge terminal, where the concierge reviews it and makes corrections as needed.

[0426] The revised response is sent back to the server, which then sends the final response to the user's device. The user is notified in a format such as, "The only upscale restaurants available next Saturday evening are Restaurant C (reservations available) and Restaurant B (waitlist)."

[0427] In this way, the system can respond quickly and accurately to user requests while reducing the burden on concierges.

[0428] The following describes the processing flow.

[0429] Step 1:

[0430] The user uses their device to input their request in voice or text format. For example, the user might say into their smartphone's microphone, "I'd like to make a dinner reservation at a high-end restaurant in Tokyo next Saturday night."

[0431] Step 2:

[0432] The terminal sends the input voice data to the server. The voice data is transmitted in digital format.

[0433] Step 3:

[0434] The server uses speech recognition technology to convert the received audio data into text data. For example, it might produce text data such as, "I would like to make a dinner reservation at a high-end restaurant in Tokyo next Saturday night."

[0435] Step 4:

[0436] The server uses a natural language processing algorithm to analyze the generated text data. This analysis extracts elements such as "next Saturday," "evening," "in Tokyo," "high-end restaurant," and "dinner reservation."

[0437] Step 5:

[0438] The server uses generative AI to generate example answers based on the extracted elements. For example, the AI ​​might generate an example answer such as, "We suggest the following restaurants: 1. Restaurant A (reservations available) 2. Restaurant B (waitlist)."

[0439] Step 6:

[0440] The server sends the generated example answer to the concierge terminal. The concierge terminal displays the received content.

[0441] Step 7:

[0442] The concierge reviews the sample responses received and makes corrections as needed. For example, the concierge will actually contact "Restaurant A" and "Restaurant B" to check the latest reservation status and update the list if necessary.

[0443] Step 8:

[0444] The concierge sends the revised answer back to the server. The server verifies the revised content.

[0445] Step 9:

[0446] The server sends the final response to the user's terminal. For example, the user might be notified that "The only upscale restaurants available next Saturday evening are Restaurant C (reservations available) and Restaurant B (waitlist)."

[0447] Through the steps outlined above, this system efficiently responds to user requests, reduces the burden on concierges, and improves service quality.

[0448] (Example 1)

[0449] Next, we will describe Example 1. In the following description, the data processing device 12 will be referred to as the "server" and the smart glasses 214 will be referred to as the "terminal".

[0450] Traditional concierge services have faced challenges in quickly and accurately understanding user requests and generating appropriate responses. In particular, voice input requires significant time and manpower for the process of converting speech to text and then understanding that text. Furthermore, maintaining the quality of the generated responses necessitates manual review and correction by concierges, which further reduces efficiency.

[0451] The identification process performed by the identification processing unit 290 of the data processing device 12 in Example 1 is realized by the following means.

[0452] In this invention, the server includes a user terminal for the user to input requests in voice or text format, a conversion means for converting voice data received from the user into text data, an analysis means for analyzing the text data generated by the conversion means as a request using a natural language processing algorithm, a generation means for generating example answers using generative artificial intelligence based on the requests obtained by the analysis means, a transmission means for sending the example answers generated by the generation means to an administrator terminal for confirmation and correction by the administrator, a transmission means for sending the corrected answers returned from the administrator terminal to the user terminal, and a computer for automatically generating information that matches the user's requests. This makes it possible to quickly and accurately understand the user's requests and efficiently provide appropriate answers.

[0453] A "user terminal" is a device used by users to input requests in voice or text format, and specifically includes smartphones and personal computers.

[0454] "Conversion means" refers to a function or device for converting audio data received from a user into text data, and includes those that utilize speech recognition technology.

[0455] "Analysis means" refers to a function or device for analyzing text data generated by the conversion means as a request using a natural language processing algorithm.

[0456] "Generation means" refers to a function or device for generating example answers using generative artificial intelligence based on requests obtained by analysis means.

[0457] "Transmission means" refers to a function or device for transmitting example answers and modified answers generated by the generation means to administrator terminals and user terminals.

[0458] An "administrator terminal" is a device used by administrators to review and modify generated sample answers, and specifically includes personal computers and tablets.

[0459] An "electronic computing device" is a device that performs functions and processes to automatically generate information that meets the user's requirements.

[0460] System Configuration

[0461] This invention is a system for efficiently processing user requests entered in voice or text format and improving the quality of concierge services. This system includes the following elements:

[0462] 1. User terminal:

[0463] These are devices that allow users to input requests in voice or text format, such as smartphones and personal computers.

[0464] 2. Server:

[0465] It has a conversion means for converting audio data received from a user into text data.

[0466] We will use the Google Cloud Speech-to-Text API as our speech recognition technology.

[0467] The system includes an analysis means for analyzing text data generated by a conversion means as a request using a natural language processing algorithm.

[0468] We will use spaCy or the NLTK library as natural language processing algorithms.

[0469] It has a generation means for generating example answers using generative artificial intelligence based on the request.

[0470] As a generative artificial intelligence, we will use the OpenAI GPT-3 model.

[0471] The system has a means for sending the generated sample answers to the administrator's terminal for review and correction by the administrator.

[0472] It has a means for sending the corrected answer to the user's terminal.

[0473] 3. Administrator terminal:

[0474] This device allows administrators to review and correct generated sample answers, and examples include personal computers and tablets.

[0475] Operation details

[0476] First, the user uses a device (smartphone or personal computer) to input their request in voice or text format. For example, the user might say into their smartphone's microphone, "I would like to make a dinner reservation at a high-end restaurant in Tokyo next Saturday night."

[0477] Next, the device sends the audio data to the server. The server uses the Google Cloud Speech-to-Text API to convert the audio data into text data, obtaining the text "I would like to make a dinner reservation at a high-end restaurant in Tokyo next Saturday night."

[0478] The server then analyzes the text data using a natural language processing algorithm (for example, the spaCy library) to extract requests such as "next Saturday," "evening," "in Tokyo," "high-end restaurant," and "dinner reservation."

[0479] The server then generates example answers based on the elements extracted using OpenAI's GPT-3 model. An example of the prompt text used here is as follows:

[0480] "The user said, 'I want to make a dinner reservation at a high-end restaurant in Tokyo next Saturday night.' Please create a list of restaurant suggestions based on this."

[0481] The generated example responses will be in the format of "Restaurant A (Reservations Available)" or "Restaurant B (Waiting List)". The server sends these example responses to the administrator's terminal, where the administrator actually contacts the restaurants to check the latest reservation status and corrects the example responses as needed.

[0482] The revised response is sent back to the server, and the final response is sent to the user's device (smartphone or personal computer). The user can receive the final response via email or app notification, such as, "The only upscale restaurants available next Saturday evening are Restaurant C (reservations available) and Restaurant B (waitlist)."

[0483] In this way, the system can respond quickly and accurately to user requests while reducing the burden on concierges.

[0484] The flow of the specific processing in Example 1 will be explained using Figure 11.

[0485] Step 1:

[0486] Users use their devices (smartphones or personal computers) to input their requests in voice or text format.

[0487] Specific action: The user speaks into their smartphone's microphone, saying, "I would like to make a dinner reservation at a high-end restaurant in Tokyo next Saturday night." Alternatively, they type a similar request as text.

[0488] Input: User's request in voice or text format.

[0489] Output: Audio or text data generated within the device.

[0490] Step 2:

[0491] The device sends the user's voice data to the server.

[0492] Specific operation: The smartphone app uploads the recorded audio file to the server. Text files are sent as is.

[0493] Input: User's voice data or text data.

[0494] Output: Audio or text data sent to the server.

[0495] Step 3:

[0496] The server converts the received audio data into text data using speech recognition technology.

[0497] Specific operation: The server calls the Google Cloud Speech-to-Text API and converts the audio data into text, "I would like to make a dinner reservation at a high-end restaurant in Tokyo next Saturday night."

[0498] Input: Audio data.

[0499] Output: Text data.

[0500] Step 4:

[0501] The server analyzes the text data using a natural language processing algorithm to extract the specific elements of the request.

[0502] Specific operation: The server uses the spaCy library to parse the text and extract elements such as "next Saturday," "evening," "in Tokyo," "high-end restaurant," and "dinner reservation."

[0503] Input: Text data.

[0504] Output: Extracted elements ("Date: Next Saturday", "Time: Evening", "Location: Tokyo", "Genre: High-end restaurant", "Content: Dinner reservation").

[0505] Step 5:

[0506] The server uses generative artificial intelligence to generate example answers based on the extracted elements.

[0507] Specific operation: The server inputs the following prompt to the OpenAI GPT-3 model: "The user said, 'I would like to make a dinner reservation at a high-end restaurant in Tokyo next Saturday night.' Please create a list of restaurant suggestions based on this."

[0508] Input: Extracted elements.

[0509] Output: Generated example answers (e.g., "Restaurant A (Reservations Available)", "Restaurant B (Waiting List)").

[0510] Step 6:

[0511] The server sends the generated sample answers to the administrator's terminal, where the administrator reviews and corrects the content.

[0512] Specific operation: The server sends a sample response to the administrator's terminal, and the administrator checks the latest reservation status with the pre-designated restaurant by phone or web and modifies the sample response as needed.

[0513] Input: Generated example answer.

[0514] Output: Example answer corrected by the administrator.

[0515] Step 7:

[0516] The server sends the corrected answer example to the user's terminal.

[0517] Specific operation: The server sends the corrected answer example to the user's smartphone or personal computer via email or app notification.

[0518] Input: Example answer corrected by the administrator.

[0519] Output: The final response sent to the user (e.g., "The only upscale restaurants available next Saturday evening are Restaurant C (reservations available) and Restaurant B (waitlist)").

[0520] (Application Example 1)

[0521] Next, we will explain Application Example 1. In the following explanation, the data processing device 12 will be referred to as the "server," and the smart glasses 214 will be referred to as the "terminal."

[0522] In modern brick-and-mortar stores, customers often demand prompt and accurate product guidance and recommendations, but there is a lack of high-quality and efficient service to meet this demand. Furthermore, manual service is costly and places a heavy burden on employees, leading to decreased customer satisfaction. In particular, while there is a growing demand for interactive service delivery using smart devices, appropriate systems are still lacking.

[0523] The specific processing performed by the specific processing unit 290 of the data processing device 12 in Application Example 1 is realized by the following means.

[0524] In this invention, the server includes a user terminal for the user to input requests in voice or text format; a conversion means for converting voice data received from the user into text data; an analysis means for analyzing the text data generated by the conversion means as a request using a natural language processing algorithm; a generation means for generating example answers using a generative AI based on the requests obtained by the analysis means; a transmission means for sending the example answers generated by the generation means to a concierge terminal for confirmation and correction by the concierge; a transmission means for sending the corrected answers returned from the concierge terminal to the user terminal; and a notification means that allows the user terminal to input requests via smart glasses and notifies the user of the final answer via the smart glasses. This makes it possible for customers in physical stores to receive quick and appropriate product information through voice input.

[0525] A "user terminal" is a device used by users to input requests in voice or text format.

[0526] "Conversion means" refers to a device or method for converting audio data received from a user into text data.

[0527] "Analysis means" refers to a device or method for analyzing text data generated by a conversion means as a request using a natural language processing algorithm.

[0528] "Generation means" refers to an apparatus or method for generating example answers using a generative AI based on requests obtained by an analysis means.

[0529] "Transmission means" refers to a device or method for transmitting example answers generated by the generation means to a concierge terminal, and for transmitting corrected answers returned from the concierge terminal to a user terminal.

[0530] "Smart glasses" are devices that users wear to input requests and receive responses visually.

[0531] "Notification means" refers to a device or method for notifying a user of the final response to a request through a user terminal (including smart glasses).

[0532] Modes for carrying out the invention

[0533] This invention is a system for providing customers with quick and appropriate product information in physical stores. Specific embodiments are described below.

[0534] System Program Overview

[0535] 1. Voice input and data conversion:

[0536] The user performs voice input while wearing smart glasses. For example, they might voice a request such as, "I want to find a shirt that matches these shoes."

[0537] The smart glasses incorporate voice recognition technology, converting received voice data into text data. This process utilizes Google's voice recognition API.

[0538] 2. Analysis of requests:

[0539] The converted text data is sent to the server. On the server side, a natural language processing algorithm (BERT, a transformer model from Hugging Face) is used to analyze the text data and extract the customer's request as specific elements. For example, information such as "Item: Shoes," "Request: Shirt," and "Purpose: Coordination support" is obtained.

[0540] 3. Generating example answers:

[0541] Based on the extracted elements, a generative AI (e.g., GPT-3) is used to generate example answers. The generated example answers include specific product information, such as "the blue shirt on shelf A" or "the white shirt on shelf B." An example of this prompt is: "User request: I want to find a shirt that matches these shoes. Please suggest specific products that match this request."

[0542] 4. Concierge review and correction:

[0543] The generated sample responses are sent to the concierge terminal. The concierge reviews the list of suggestions and makes revisions as needed. For example, they may update the list to reflect inventory status or additional product information, and then send the revised responses back to the server.

[0544] 5. Notification of final response:

[0545] The server notifies the user of the corrected final answer via the smart glasses they are wearing. This allows the user to receive the notification visually. For example, the guidance might say, "The shirts that match these shoes are the blue shirt on shelf A and the white shirt on shelf B."

[0546] Hardware and software to be used

[0547] Smart glasses: A device for voice input and notification display.

[0548] Server: The central hub responsible for converting audio data to text, analyzing text data, and managing generated response examples. It runs Google's speech recognition API, Hugging Face's BERT, and generative AI (GPT-3).

[0549] Natural Language Processing (NLP): These algorithms run on a server and are used to analyze and extract user requests.

[0550] Adding specific examples

[0551] For example, if a user enters a store and uses smart glasses to voice-input "I want to find a shirt that goes with these shoes," they will receive a notification a few seconds later saying, "We recommend the blue shirt on shelf A and the white shirt on shelf B." This experience allows users to quickly find products that meet their needs.

[0552] In this way, this system can improve the customer experience in physical stores and provide high-quality service while reducing the burden on employees.

[0553] The flow of a specific process in Application Example 1 will be explained using Figure 12.

[0554] Step 1:

[0555] The user wears smart glasses and inputs their request by voice. For example, they might say, "I want to find a shirt that matches these shoes." This input is collected via the smart glasses' microphone. The voice signal is then used as input for speech recognition processing to convert it into text.

[0556] Step 2:

[0557] Smart glasses convert audio data into text data. Specifically, they use Google's speech recognition API to convert audio data into text data. The input to this process is audio data, and the output is text data.

[0558] Step 3:

[0559] The converted text data is sent to the server. The server analyzes the received text data using a natural language processing algorithm (Hugging Face's BERT model). In this step, the input is text data, and the output is the specific elements of the analyzed request (e.g., "Item: Shoes", "Request: Shirt", "Purpose: Coordination Support").

[0560] Step 4:

[0561] The server uses a generative AI (GPT-3) to generate example responses based on the elements of the analyzed request. The input used in this process is the specific elements of the analyzed request, and the output is a specific product suggestion. For example, example responses might be "a blue shirt on shelf A" or "a white shirt on shelf B."

[0562] Step 5:

[0563] The generated sample responses are sent to the concierge terminal. The concierge reviews the list of suggestions and makes corrections as needed. Based on inventory information and new product information, they create more accurate suggestions. The input in this step is the generated sample response, and the output is the corrected sample response.

[0564] Step 6:

[0565] The revised answer examples are sent back to the server and compiled into the final answer. This final answer is then sent to the smart glasses worn by the user. For example, the notification might say, "The shirts that go with these shoes are the blue shirt on shelf A and the white shirt on shelf B." In this step, the input is the revised answer examples, and the output is the notification message displayed on the smart glasses.

[0566] In this way, the input data is sequentially processed and analyzed at each processing step, and finally, appropriate product information is provided to the user.

[0567] Furthermore, an emotion engine that estimates the user's emotions may be incorporated. That is, the identification processing unit 290 may use the emotion identification model 59 to estimate the user's emotions and perform identification processing using the user's emotions.

[0568] This invention is a system for efficiently processing user requests in voice or text format and further improving the quality of concierge services by combining it with an emotion engine that recognizes the user's emotions. Specific embodiments of the system are shown below.

[0569] How users can enter their requests

[0570] The process begins with the user entering their request using a device (e.g., a smartphone or PC). Users can enter their request via voice or text. For example, a user might voice-input, "I would like to make a dinner reservation at a high-end restaurant in Tokyo next Saturday night."

[0571] Converting audio data to text

[0572] The device sends the user's voice data to the server. The server uses speech recognition technology to convert the voice data into text data. For example, it might generate the text "I would like to make a dinner reservation at a high-end restaurant in Tokyo next Saturday night" from the voice data.

[0573] Request analysis

[0574] The server analyzes the generated text data using a natural language processing algorithm and extracts specific elements representing the user's request. This yields information such as "Date: Next Saturday," "Time: Evening," "Location: Tokyo," "Genre: High-end restaurant," and "Content: Dinner reservation."

[0575] Using an Emotion Engine

[0576] The server then uses an emotion engine to analyze the user's emotions from voice or text data. For example, it extracts emotional information such as "excited," "anxious," or "hurried." This emotional information influences the generative AI in subsequent processes, which is used to generate more appropriate responses that align with the user's feelings.

[0577] Generating example answers

[0578] The server uses generative AI to generate example responses based on extracted elements and emotional information. For example, the AI ​​might generate a response like, "We suggest the following restaurants: 1. Restaurant A (reservations available) 2. Restaurant B (waitlist)." By considering emotional information, it's possible to include additional perks for users who are excited, for example.

[0579] Concierge confirmation

[0580] The server sends the generated sample response to the concierge's terminal, where the concierge reviews and corrects the content. The concierge then contacts, for example, "Restaurant A" and "Restaurant B" to check the latest reservation status and updates the list if necessary.

[0581] Submit your final response

[0582] The revised response is sent back to the server, which then sends the final response to the user's device. The user can receive the final response via email or app notification. For example, the notification might say, "Here are some high-end restaurants available next Saturday evening. Restaurant A is fully booked, while Restaurant B is on the waiting list."

[0583] Specific example

[0584] The user speaks into their smartphone's microphone and says, "I'd like to make a dinner reservation at a high-end restaurant in Tokyo next Saturday night."

[0585] The device sends the voice data to the server. The server performs speech recognition and obtains text data that reads, "I would like to make a dinner reservation at a high-end restaurant in Tokyo next Saturday night."

[0586] The server analyzes the text data using a natural language processing algorithm and extracts requests such as "next Saturday," "evening," "in Tokyo," "high-end restaurant," and "dinner reservation."

[0587] The server uses an emotion engine to extract emotional information from the user's voice and text data, recognizing emotions such as "excited" or "hurried."

[0588] The server uses generative AI to generate a list based on the user's requests and recognized emotions. For example, "Restaurant A (Reservations Available)" and "Restaurant B (Waiting List)." Depending on the emotion, for instance, if the emotion is excitement, the system will generate a response that includes additional information about special offers.

[0589] This list is sent to the concierge terminal, where the concierge reviews it and makes corrections as needed.

[0590] The revised response is sent back to the server, which then sends the final response to the user's device. The user is notified that "The only upscale restaurants available next Saturday evening are Restaurant A (reservations available) and Restaurant B (waitlist)."

[0591] This system responds quickly and accurately to user requests and provides more personalized service by taking emotions into consideration. This reduces the burden on concierges and improves the quality of service.

[0592] The following describes the processing flow.

[0593] Step 1:

[0594] The user uses their device to input their request in voice or text format. For example, the user might say into their smartphone's microphone, "I'd like to make a dinner reservation at a high-end restaurant in Tokyo next Saturday night."

[0595] Step 2:

[0596] The terminal sends the input voice data to the server. The voice data is transmitted in digital format.

[0597] Step 3:

[0598] The server uses speech recognition technology to convert the received audio data into text data. For example, it might produce text data such as, "I would like to make a dinner reservation at a high-end restaurant in Tokyo next Saturday night."

[0599] Step 4:

[0600] The server uses a natural language processing algorithm to analyze the generated text data. This analysis extracts elements such as "next Saturday," "evening," "in Tokyo," "high-end restaurant," and "dinner reservation."

[0601] Step 5:

[0602] The server uses an emotion engine to analyze the user's emotions from voice or text data. For example, the emotion engine extracts emotional information such as "excited," "anxious," or "hurried."

[0603] Step 6:

[0604] The server uses generative AI to generate example responses based on emotional information and desired elements. By considering emotional information, more personalized responses are generated. For example, the AI ​​might generate a response like, "We suggest the following restaurants: 1. Restaurant A (reservations available) 2. Restaurant B (waitlist)." If the user is looking forward to something, special offers may also be added.

[0605] Step 7:

[0606] The server sends the generated example answer to the concierge terminal. The concierge terminal displays the received content.

[0607] Step 8:

[0608] The concierge reviews the sample responses received and makes corrections as needed. For example, the concierge will actually contact "Restaurant A" and "Restaurant B" to check the latest reservation status and update the list if necessary.

[0609] Step 9:

[0610] The concierge sends the revised answer back to the server. The server verifies the revised content.

[0611] Step 10:

[0612] The server sends the final response to the user's terminal. For example, the user might be notified that "The only upscale restaurants available next Saturday evening are Restaurant A (reservations available) and Restaurant B (waitlist)."

[0613] Through the steps outlined above, this system can efficiently respond to user requests and, by taking user emotions into consideration, improve the quality of concierge services.

[0614] (Example 2)

[0615] Next, we will describe Example 2. In the following description, the data processing device 12 will be referred to as the "server" and the smart glasses 214 will be referred to as the "terminal".

[0616] Traditional concierge services lacked the technology to appropriately and quickly handle specific user requests, and particularly struggled to respond while considering emotions. Therefore, improving service quality requires analyzing user requests, including their emotions, and generating appropriate responses based on that analysis.

[0617] The specific processing performed by the specific processing unit 290 of the data processing device 12 in Example 2 is realized by the following means. In this invention, the server includes an information terminal for the user to input requests in voice or text format, a conversion means for converting voice data received from the user into text data, an analysis means for analyzing the text data generated by the conversion means using a natural language processing algorithm, a generation means for generating example answers using generative artificial intelligence based on the requests and sentiment information obtained by the analysis means, a transmission means for sending the example answers generated by the generation means to a service provider terminal for confirmation and correction by the service provider, and a transmission means for sending the corrected answers returned from the service provider terminal to the user terminal. This enables the provision of quick and appropriate services that take into account the user's requests and sentiments.

[0618] An "information terminal" is a device used by users to input requests in voice or text format, and includes smartphones and personal computers.

[0619] "Conversion means" refers to a technology or method for converting audio data received from a user into text data, and which uses speech recognition technology.

[0620] "Analysis means" refers to a technology or method that analyzes text data generated by the conversion means using a natural language processing algorithm to extract user requests and emotional information.

[0621] "Generation means" refers to a technology or method that generates example responses using generative artificial intelligence based on requests and emotional information obtained by analysis means.

[0622] "Transmission means" refers to a technology or method for transmitting example answers generated by the generation means to a service provider terminal, and for transmitting the corrected answers to a user terminal.

[0623] A "service provider terminal" refers to a device used to review and correct generated sample answers, specifically the device used by the concierge.

[0624] A "natural language processing algorithm" is an algorithm used to analyze text data and extract information about desires and emotions, and it includes machine learning and data analysis techniques.

[0625] "Generative artificial intelligence" refers to artificial intelligence technology that generates example responses based on analyzed requests and sentiment information, such as using a generative language model.

[0626] "Emotional information" refers to emotional information analyzed from the user's input voice or text data, and includes emotions such as excitement, anxiety, and urgency.

[0627] Modes for carrying out the invention

[0628] This invention is a system for efficiently processing user requests in voice or text format and further improving the quality of concierge services by combining it with an emotion engine that recognizes the user's emotions. Specific embodiments of the system are shown below.

[0629] How users can enter their requests

[0630] The process begins with the user entering their request in voice or text format using an information terminal (such as a smartphone or personal computer). For example, a user could voice-input a request like, "I would like to make a dinner reservation at a high-end restaurant in Tokyo next Saturday night." Users can input voice data using the terminal's microphone or input their request as text data using a keyboard.

[0631] Converting audio data to text

[0632] The device sends the user's voice data to the server. The server uses speech recognition technology to convert the voice data into text data. Specifically, it uses a speech recognition service such as the Google Cloud Speech-to-Text API to convert the voice data into text data. For example, it generates text data such as "I would like to make a dinner reservation at a high-end restaurant in Tokyo next Saturday night" from the voice data.

[0633] Request analysis

[0634] The server analyzes the generated text data using natural language processing algorithms to extract specific elements representing the user's request. Algorithms used include SpaCy and the Google Cloud Natural Language API. This analysis yields information such as "Date: Next Saturday," "Time: Evening," "Location: Tokyo," "Genre: High-end restaurant," and "Content: Dinner reservation."

[0635] Using an Emotion Engine

[0636] The server then uses an emotion engine to analyze the user's emotions from voice or text data. For example, it uses IBM Watson Tone Analyzer to extract emotional information such as "excited," "anxious," and "hurried." This emotional information influences the generative artificial intelligence in subsequent processes, which is used to generate more appropriate responses that align with the user's feelings.

[0637] Generating example answers

[0638] The server generates example responses using generative artificial intelligence (e.g., OpenAI GPT-3) based on extracted elements and emotional information. For example, the generative AI might generate a response such as, "We suggest the following restaurants: 1. Restaurant A (reservations available) 2. Restaurant B (waitlist)." By considering emotional information, it's possible to include additional perks for users who are excited, for example.

[0639] Concierge confirmation

[0640] The server sends the generated sample response to the service provider's terminal, where the service provider reviews and corrects the content. The service provider (e.g., a concierge) then actually contacts restaurants such as "Restaurant A" and "Restaurant B" to check their latest reservation status and updates the list if necessary.

[0641] Submit your final response

[0642] The revised response from the service provider's terminal is sent back to the server, which then sends the final response to the user's information terminal. The user can receive the final response via email or app notification. For example, the notification might say, "We have a list of upscale restaurants available next Saturday evening. Restaurant A is fully booked, while Restaurant B is on the waiting list."

[0643] Examples of prompt statements

[0644] If a user speaks into their smartphone's microphone and says, "I'd like to make a dinner reservation at a high-end restaurant in Tokyo next Saturday night," the following prompt will be used for the generative artificial intelligence.

[0645] "User request: "I want to make a dinner reservation at a high-end restaurant in Tokyo next Saturday night." Sentiment information: "Excited, in a hurry." Please generate a list of responses: Example) - Restaurant A (Reservations available) - Restaurant B (Waiting list) Perks information: Add perks information based on the excited sentiment."

[0646] This allows the system to respond quickly and accurately to user requests and provide more personalized service by taking emotions into consideration. In this way, it is possible to reduce the burden on concierges and improve the quality of service.

[0647] The flow of the specific processing in Example 2 will be explained using Figure 13.

[0648] Step 1: Enter your request

[0649] User

[0650] Users input their requests in voice or text format using an information terminal (e.g., smartphone, personal computer).

[0651] Input: Audio data or text data

[0652] Specific action: For example, the user speaks into their smartphone and says, "I want to make a dinner reservation at a high-end restaurant in Tokyo next Saturday night."

[0653] Step 2: Converting audio data to text

[0654] terminal

[0655] The device sends the audio data to the server.

[0656] Input: Audio data

[0657] Output: Sending audio data to the server

[0658] Specific action: Send audio data to the server in digital format.

[0659] server

[0660] The server uses speech recognition technologies such as the Google Cloud Speech-to-Text API to convert audio data into text data.

[0661] Input: Audio data

[0662] Output: Text data

[0663] Specific operation: Generate the text "I would like to make a dinner reservation at a high-end restaurant in Tokyo next Saturday night" from the audio data.

[0664] Step 3: Analyzing the requirements

[0665] server

[0666] The server analyzes the generated text data using natural language processing algorithms (e.g., SpaCy, Google Cloud Natural Language API) to extract the specific elements of the request.

[0667] Input: Text data

[0668] Output: Elements of the request (e.g., "Date: Next Saturday", "Time: Evening", "Location: Tokyo", "Genre: High-end restaurant", "Content: Dinner reservation")

[0669] Specific operation: Analyze the specific elements of the request and store them in the database.

[0670] Step 4: Utilizing the Emotional Engine

[0671] server

[0672] The server uses an emotion engine (e.g., IBM Watson Tone Analyzer) to analyze emotional information from audio or text data.

[0673] Input: Audio data or text data

[0674] Output: Emotional information (e.g., "excited," "anxious," "hurried")

[0675] Specific operation: Use an emotion engine to analyze emotions from speech or text and store them in a database.

[0676] Step 5: Generating example answers

[0677] server

[0678] The server generates example responses using a generative AI (e.g., OpenAI GPT-3) based on the elements of the request and emotional information.

[0679] Input: Elements of requests, emotional information

[0680] Output: Example answer (Example: "Restaurant A (Reservations accepted)", "Restaurant B (Waiting list)")

[0681] Specific operation: Use generative AI to generate responses that respond to requests and emotions.

[0682] Step 6: Concierge Confirmation

[0683] server

[0684] The server sends the generated example answers to the service provider's terminal.

[0685] Input: Example Answer

[0686] Output: Sent to the service provider's terminal

[0687] Specific action: Send example answers to the service provider's terminal.

[0688] Service provider

[0689] The service provider will review the responses and make corrections as necessary.

[0690] Input: Example Answer

[0691] Output: Revised answer example

[0692] Specific action: The service provider inquires about the latest booking status and corrects the response.

[0693] Step 7: Submit your final response

[0694] Service provider

[0695] The revised answer example is sent back to the server.

[0696] Input: Revised answer example

[0697] Output: Send to server

[0698] Specific action: Send the corrected response to the server.

[0699] server

[0700] The server sends the final response to the user's information terminal.

[0701] Input: Revised answer example

[0702] Output: Sending the final response to the user's information terminal.

[0703] Specific action: Send the final response to the user via email or app notification.

[0704] In this way, by responding quickly and accurately to user requests and taking their emotions into consideration, it becomes possible to provide more personalized services.

[0705] (Application Example 2)

[0706] Next, we will explain application example 2. In the following explanation, the data processing device 12 will be referred to as the "server," and the smart glasses 214 will be referred to as the "terminal."

[0707] Traditional concierge services faced challenges in providing prompt service due to the extensive work required to accurately understand and incorporate user requests. Furthermore, they lacked consideration for user emotions, resulting in a lack of personalization and low satisfaction. In short, the uniform approach to users and the inability to provide personalized suggestions based on emotions led to a poor user experience.

[0708] In Application Example 2, the specific processing performed by the specific processing unit 290 of the data processing device 12 is realized by the following means. In this invention, the server includes a user terminal for the user to input requests in voice or text format, a conversion means for converting voice data received from the user into text data, an analysis means for analyzing the text data generated by the conversion means as a request using a natural language processing algorithm, a generation means for generating example answers using a generative AI based on the requests and emotional information from the analysis means, a transmission means for sending the example answers generated by the generation means to a concierge terminal for confirmation and correction by the concierge, a transmission means for sending the corrected answers returned from the concierge terminal to the user terminal, and an emotional analysis means for extracting emotional information from the user's voice data and text data. This makes it possible to quickly and accurately analyze user requests and provide personalized services based on emotions.

[0709] A "user terminal" is a device used by users to input requests in voice or text format.

[0710] "Conversion means" refers to technical means for converting audio data received from a user into text data.

[0711] "Analysis means" refers to technical means for analyzing text data generated by the conversion means as a request using a natural language processing algorithm.

[0712] "Generation means" refers to technical means for generating example responses using a generative AI based on requests and emotional information obtained through analysis means.

[0713] "Transmission means" refers to a technical means for sending example answers generated by the generation means to a concierge terminal, for confirmation and correction by the concierge, and for sending the corrected answers returned from the concierge terminal to the user terminal.

[0714] "Emotion analysis methods" refer to technical means for extracting emotional information from user voice data and text data.

[0715] A "natural language processing algorithm" is an algorithm that analyzes text data to understand and extract user requests.

[0716] "Generative AI" is an artificial intelligence technology that generates appropriate response examples based on user requests and emotional information.

[0717] A "concierge terminal" is a device used to review and correct generated sample answers.

[0718] "Emotional information" refers to information about a user's emotions extracted from their voice data and text data.

[0719] This invention is a system for further improving the quality of concierge services by efficiently processing requests entered by users in voice or text format, and by combining this with an emotion engine that recognizes the user's emotions.

[0720] The system includes the following components:

[0721] 1. User terminal: A device used by users to input requests in voice or text format. This includes smartphones, PCs, and smart glasses.

[0722] 2. Conversion method: Speech recognition technology is used to convert the audio data received from the user into text data. Specifically, the SpeechRecognition library is used to convert the audio data into text data.

[0723] 3. Analysis Method: The text data generated by the transformation method is analyzed as a request using a natural language processing algorithm. This uses the natural language processing model from the Transformers library.

[0724] 4. Emotion Analysis Method: An emotion analysis model is used to extract emotional information from the user's voice data and text data. Specifically, the Transformers emotion analysis model is used to obtain emotional information.

[0725] 5. Generation Method: Based on the requests and sentiment information obtained through the analysis method, a generative AI model (such as OpenAI GPT-3) is used to generate example responses. This process uses OpenAI's GPT-3 as an API and takes into account the relationship between the prompt text and the generated text.

[0726] 6. Transmission method: The example answer generated by the generation method is sent to the concierge terminal for confirmation and correction by the concierge. Furthermore, the corrected answer returned from the concierge terminal is sent to the user terminal.

[0727] Explanation of specific processing examples:

[0728] The user inputs their request by voice into their smartphone's microphone, saying, "I would like to make a dinner reservation at a high-end restaurant in Tokyo next Saturday night."

[0729] The user's terminal sends the voice data to the server, which uses the SpeechRecognition library to convert the voice data into text data that reads, "I would like to make a dinner reservation at a high-end restaurant in Tokyo next Saturday night."

[0730] The converted text data is analyzed by Transformers' natural language processing model, and requests such as "Date: Next Saturday", "Time: Evening", "Location: Tokyo", "Genre: High-end restaurant", and "Content: Dinner reservation" are extracted.

[0731] An emotion analysis model analyzes the same text data and obtains emotional information such as "excited" or "hurried."

[0732] Based on the analyzed requests and sentiment information, the server uses GPT-3 to generate example responses with the following prompts:

[0733] User Input: I'm looking for recommended baby products in the store.

[0734] Emotion: Enjoyment

[0735] Response:

[0736] For example, the AI ​​can generate sample responses such as, "In this baby products section, we recommend our new skin-friendly diapers and soft swaddles."

[0737] The generated sample answers are sent to the concierge terminal, where the concierge reviews the content and makes corrections as needed.

[0738] The corrected response is sent back to the server and finally delivered to the user's terminal.

[0739] This system enables the rapid and accurate analysis of user requests and the provision of personalized, emotion-based services. This improves the user experience and reduces the workload on concierges.

[0740] The flow of a specific process in Application Example 2 will be explained using Figure 14.

[0741] Step 1:

[0742] The user performs voice input. The user uses a smartphone or smart glasses to input their request by voice. For example, they might say, "I would like to make a dinner reservation at a high-end restaurant in Tokyo next Saturday night." This voice data is recorded on the device. The input is voice data, and the output is the voice file recorded on the device.

[0743] Step 2:

[0744] The device sends audio data to the server. The device sends the recorded audio data to the server. The server uses speech recognition technology to convert this audio data into text data. The SpeechRecognition library is used for this conversion. The input is audio data, and the output is text data that says, "I would like to make a dinner reservation at a high-end restaurant in Tokyo next Saturday night."

[0745] Step 3:

[0746] The server analyzes the text data. The server then analyzes the transformed text data using a natural language processing algorithm. This process uses the Transformers library to extract user requests (e.g., "Date: Next Saturday", "Time: Evening", "Location: Tokyo", "Genre: High-end restaurant", "Content: Dinner reservation"). The input is text data, and the output is the analyzed elemental information.

[0747] Step 4:

[0748] The server extracts emotional information. The server then uses emotion analysis tools to analyze the user's emotional information from the text data. This process utilizes the Transformers emotion analysis model. For example, emotions such as "excited" and "hurried" are extracted. The input is text data, and the output is emotional information.

[0749] Step 5:

[0750] The server generates example answers using generative AI. Based on the analyzed requests and sentiment information, the server generates example answers using a generative AI model such as GPT-3. This process uses OpenAI's GPT-3 as an API, sending the generated text as a prompt. The input is requests and sentiment information, and the output is the suggested answer. For example, the prompt might look like this:

[0751] User Input: I'm looking for recommended baby products in the store.

[0752] Emotion: Enjoyment

[0753] Response:

[0754] Step 6:

[0755] The server sends example answers to the concierge terminal. The generated example answers are sent to the concierge terminal, where the concierge reviews the content and makes corrections as needed. The input is the suggested example answer, and the output is the corrected example answer.

[0756] Step 7:

[0757] The concierge terminal sends the corrected answer back to the server. After the concierge reviews and makes corrections, the corrected answer is sent back to the server. The input is an example of the corrected answer, and the output is the final example answer.

[0758] Step 8:

[0759] The server sends the final answer to the user's device. The user can receive the final answer via email or app notification. The input is an example of the final answer, and the output is the final answer notified to the user. For example, the user might be notified of the final answer, "The upscale restaurants available next Saturday evening are Restaurant A (reservations available) and Restaurant B (waitlist)."

[0760] The specific processing unit 290 transmits the result of the specific processing to the smart glasses 214. In the smart glasses 214, the control unit 46A causes the speaker 240 to output the result of the specific processing. The microphone 238 acquires audio indicating user input for the result of the specific processing. The control unit 46A transmits the audio data indicating user input acquired by the microphone 238 to the data processing unit 12. In the data processing unit 12, the specific processing unit 290 acquires the audio data.

[0761] Data generation model 58 is a type of so-called generative AI (Artificial Intelligence). One example of data generation model 58 is ChatGPT (Internet search<URL: https: / / openai.com / blog / chatgpt> ), Gemini (Internet search) <url: https: gemini.google.com ?hl="ja">Examples of generative AI include the following. The data generation model 58 is obtained by performing deep learning on a neural network. The data generation model 58 is input with prompts containing instructions, and with inference data such as audio data representing speech, text data representing text, and image data representing images. The data generation model 58 infers from the input inference data according to the instructions indicated by the prompts, and outputs the inference results in data formats such as audio data and text data. Here, inference refers to, for example, analysis, classification, prediction, and / or summarization.

[0762] In the above embodiment, an example was given in which specific processing is performed by the data processing device 12, but the technology of this disclosure is not limited thereto, and the specific processing may also be performed by the smart glasses 214.

[0763] [Third Embodiment]

[0764] Figure 5 shows an example of the configuration of the data processing system 310 according to the third embodiment.

[0765] As shown in Figure 5, the data processing system 310 includes a data processing device 12 and a headset terminal 314. An example of the data processing device 12 is a server.

[0766] The data processing device 12 comprises a computer 22, a database 24, and a communication interface 26. The computer 22 is an example of a "computer" related to the technology of this disclosure. The computer 22 comprises a processor 28, RAM 30, and storage 32. The processor 28, RAM 30, and storage 32 are connected to a bus 34. The database 24 and the communication interface 26 are also connected to the bus 34. The communication interface 26 is connected to a network 54. An example of the network 54 is a WAN (Wide Area Network) and / or a LAN (Local Area Network).

[0767] The headset terminal 314 includes a computer 36, a microphone 238, a speaker 240, a camera 42, a communication interface 44, and a display 343. The computer 36 includes a processor 46, RAM 48, and storage 50. The processor 46, RAM 48, and storage 50 are connected to a bus 52. The microphone 238, speaker 240, camera 42, and display 343 are also connected to the bus 52.

[0768] The microphone 238 receives voice signals from the user 20 and receives instructions from the user 20. The microphone 238 captures the voice signals from the user 20, converts the captured voice into audio data, and outputs it to the processor 46. The speaker 240 outputs audio according to the instructions from the processor 46.

[0769] Camera 42 is a small digital camera equipped with an optical system including a lens, aperture, and shutter, and an image sensor such as a CMOS (Complementary Metal-Oxide-Semiconductor) image sensor or a CCD (Charge Coupled Device) image sensor, and captures images of the area around the user 20 (for example, an imaging range defined by a field of view equivalent to the width of a typical healthy person's field of vision).

[0770] Communication interface 44 is connected to network 54. Communication interfaces 44 and 26 are responsible for the exchange of various information between processor 46 and processor 28 via network 54. The exchange of various information between processor 46 and processor 28 using communication interfaces 44 and 26 is performed in a secure manner.

[0771] Figure 6 shows an example of the main functions of the data processing device 12 and the headset terminal 314. As shown in Figure 6, the data processing device 12 performs specific processing using the processor 28. The storage 32 stores the specific processing program 56.

[0772] The specific processing program 56 is an example of a "program" relating to the technology of this disclosure. The processor 28 reads the specific processing program 56 from the storage 32 and executes the read specific processing program 56 on the RAM 30. The specific processing is realized by the processor 28 operating as a specific processing unit 290 in accordance with the specific processing program 56 executed on the RAM 30.

[0773] The storage 32 stores the data generation model 58 and the emotion identification model 59. The data generation model 58 and the emotion identification model 59 are used by the identification processing unit 290.

[0774] In the headset terminal 314, the processor 46 performs the reception output processing. The storage 50 stores the reception output program 60. The processor 46 reads the reception output program 60 from the storage 50 and executes the read reception output program 60 on the RAM 48. The reception output processing is realized by the processor 46 operating as a control unit 46A according to the reception output program 60 executed on the RAM 48.

[0775] Next, the specific processing performed by the specific processing unit 290 of the data processing device 12 will be described. In the following description, the data processing device 12 will be referred to as the "server" and the headset terminal 314 will be referred to as the "terminal".

[0776] This invention is a system for efficiently processing user requests entered in voice or text format and improving the quality of concierge services. Specific embodiments of the system are described below.

[0777] How users can enter their requests

[0778] The process begins with the user entering their request using a device (e.g., a smartphone or PC). Users can enter their request via voice or text. For example, a user might voice-input, "I would like to make a dinner reservation at a high-end restaurant in Tokyo next Saturday night."

[0779] Converting audio data to text

[0780] The device sends the user's voice data to the server. The server uses speech recognition technology to convert the voice data into text data. For example, it might generate the text "I would like to make a dinner reservation at a high-end restaurant in Tokyo next Saturday night" from the voice data.

[0781] Request analysis

[0782] The server analyzes the generated text data using a natural language processing algorithm and extracts specific elements representing the user's request. This yields information such as "Date: Next Saturday," "Time: Evening," "Location: Tokyo," "Genre: High-end restaurant," and "Content: Dinner reservation."

[0783] Generating example answers

[0784] The server uses generative AI to generate example answers based on the extracted elements. For example, the AI ​​might generate an example answer such as, "We suggest the following restaurants: 1. Restaurant A (reservations available) 2. Restaurant B (waitlist)."

[0785] Concierge confirmation

[0786] The server sends the generated sample response to the concierge's terminal, where the concierge reviews and corrects the content. The concierge then actually contacts, for example, "Restaurant A" and "Restaurant B" to check the latest reservation status and corrects the sample response as needed.

[0787] Submit your final response

[0788] The revised response is sent back to the server, which then sends the final response to the user's device. The user can receive the final response via email or app notification. For example, the notification might say, "Here are some high-end restaurants available next Saturday evening. Restaurant A is fully booked, while Restaurant B is on the waiting list."

[0789] Specific example

[0790] The user speaks into their smartphone's microphone and says, "I'd like to make a dinner reservation at a high-end restaurant in Tokyo next Saturday night."

[0791] The device sends the voice data to the server. The server performs speech recognition and obtains text data that reads, "I would like to make a dinner reservation at a high-end restaurant in Tokyo next Saturday night."

[0792] The server analyzes the text data using a natural language processing algorithm and extracts requests such as "next Saturday," "evening," "in Tokyo," "high-end restaurant," and "dinner reservation."

[0793] The server uses generative AI to generate a list of restaurants that match the user's request. For example, "Restaurant A (Reservations Available)" and "Restaurant B (Waiting List)".

[0794] This list is sent to the concierge terminal, where the concierge reviews it and makes corrections as needed.

[0795] The revised response is sent back to the server, which then sends the final response to the user's device. The user is notified in a format such as, "The only upscale restaurants available next Saturday evening are Restaurant C (reservations available) and Restaurant B (waitlist)."

[0796] In this way, the system can respond quickly and accurately to user requests while reducing the burden on concierges.

[0797] The following describes the processing flow.

[0798] Step 1:

[0799] The user uses their device to input their request in voice or text format. For example, the user might say into their smartphone's microphone, "I'd like to make a dinner reservation at a high-end restaurant in Tokyo next Saturday night."

[0800] Step 2:

[0801] The terminal sends the input voice data to the server. The voice data is transmitted in digital format.

[0802] Step 3:

[0803] The server uses speech recognition technology to convert the received audio data into text data. For example, it might produce text data such as, "I would like to make a dinner reservation at a high-end restaurant in Tokyo next Saturday night."

[0804] Step 4:

[0805] The server uses a natural language processing algorithm to analyze the generated text data. This analysis extracts elements such as "next Saturday," "evening," "in Tokyo," "high-end restaurant," and "dinner reservation."

[0806] Step 5:

[0807] The server uses generative AI to generate example answers based on the extracted elements. For example, the AI ​​might generate an example answer such as, "We suggest the following restaurants: 1. Restaurant A (reservations available) 2. Restaurant B (waitlist)."

[0808] Step 6:

[0809] The server sends the generated example answer to the concierge terminal. The concierge terminal displays the received content.

[0810] Step 7:

[0811] The concierge reviews the sample responses received and makes corrections as needed. For example, the concierge will actually contact "Restaurant A" and "Restaurant B" to check the latest reservation status and update the list if necessary.

[0812] Step 8:

[0813] The concierge sends the revised answer back to the server. The server verifies the revised content.

[0814] Step 9:

[0815] The server sends the final response to the user's terminal. For example, the user might be notified that "The only upscale restaurants available next Saturday evening are Restaurant C (reservations available) and Restaurant B (waitlist)."

[0816] Through the steps outlined above, this system efficiently responds to user requests, reduces the burden on concierges, and improves service quality.

[0817] (Example 1)

[0818] Next, we will describe Example 1. In the following description, the data processing device 12 will be referred to as the "server," and the headset-type terminal 314 will be referred to as the "terminal."

[0819] Traditional concierge services have faced challenges in quickly and accurately understanding user requests and generating appropriate responses. In particular, voice input requires significant time and manpower for the process of converting speech to text and then understanding that text. Furthermore, maintaining the quality of the generated responses necessitates manual review and correction by concierges, which further reduces efficiency.

[0820] The identification process performed by the identification processing unit 290 of the data processing device 12 in Example 1 is realized by the following means.

[0821] In this invention, the server includes a user terminal for the user to input requests in voice or text format, a conversion means for converting voice data received from the user into text data, an analysis means for analyzing the text data generated by the conversion means as a request using a natural language processing algorithm, a generation means for generating example answers using generative artificial intelligence based on the requests obtained by the analysis means, a transmission means for sending the example answers generated by the generation means to an administrator terminal for confirmation and correction by the administrator, a transmission means for sending the corrected answers returned from the administrator terminal to the user terminal, and a computer for automatically generating information that matches the user's requests. This makes it possible to quickly and accurately understand the user's requests and efficiently provide appropriate answers.

[0822] A "user terminal" is a device used by users to input requests in voice or text format, and specifically includes smartphones and personal computers.

[0823] "Conversion means" refers to a function or device for converting audio data received from a user into text data, and includes those that utilize speech recognition technology.

[0824] "Analysis means" refers to a function or device for analyzing text data generated by the conversion means as a request using a natural language processing algorithm.

[0825] "Generation means" refers to a function or device for generating example answers using generative artificial intelligence based on requests obtained by analysis means.

[0826] "Transmission means" refers to a function or device for transmitting example answers and modified answers generated by the generation means to administrator terminals and user terminals.

[0827] An "administrator terminal" is a device used by administrators to review and modify generated sample answers, and specifically includes personal computers and tablets.

[0828] An "electronic computing device" is a device that performs functions and processes to automatically generate information that meets the user's requirements.

[0829] System Configuration

[0830] This invention is a system for efficiently processing user requests entered in voice or text format and improving the quality of concierge services. This system includes the following elements:

[0831] 1. User terminal:

[0832] These are devices that allow users to input requests in voice or text format, such as smartphones and personal computers.

[0833] 2. Server:

[0834] It has a conversion means for converting audio data received from a user into text data.

[0835] We will use the Google Cloud Speech-to-Text API as our speech recognition technology.

[0836] The system includes an analysis means for analyzing text data generated by a conversion means as a request using a natural language processing algorithm.

[0837] We will use spaCy or the NLTK library as natural language processing algorithms.

[0838] It has a generation means for generating example answers using generative artificial intelligence based on the request.

[0839] As a generative artificial intelligence, we will use the OpenAI GPT-3 model.

[0840] The system has a means for sending the generated sample answers to the administrator's terminal for review and correction by the administrator.

[0841] It has a means for sending the corrected answer to the user's terminal.

[0842] 3. Administrator terminal:

[0843] This device allows administrators to review and correct generated sample answers, and examples include personal computers and tablets.

[0844] Operation details

[0845] First, the user uses a device (smartphone or personal computer) to input their request in voice or text format. For example, the user might say into their smartphone's microphone, "I would like to make a dinner reservation at a high-end restaurant in Tokyo next Saturday night."

[0846] Next, the device sends the audio data to the server. The server uses the Google Cloud Speech-to-Text API to convert the audio data into text data, obtaining the text "I would like to make a dinner reservation at a high-end restaurant in Tokyo next Saturday night."

[0847] The server then analyzes the text data using a natural language processing algorithm (for example, the spaCy library) to extract requests such as "next Saturday," "evening," "in Tokyo," "high-end restaurant," and "dinner reservation."

[0848] The server then generates example answers based on the elements extracted using OpenAI's GPT-3 model. An example of the prompt text used here is as follows:

[0849] "The user said, 'I want to make a dinner reservation at a high-end restaurant in Tokyo next Saturday night.' Please create a list of restaurant suggestions based on this."

[0850] The generated example responses will be in the format of "Restaurant A (Reservations Available)" or "Restaurant B (Waiting List)". The server sends these example responses to the administrator's terminal, where the administrator actually contacts the restaurants to check the latest reservation status and corrects the example responses as needed.

[0851] The revised response is sent back to the server, and the final response is sent to the user's device (smartphone or personal computer). The user can receive the final response via email or app notification, such as, "The only upscale restaurants available next Saturday evening are Restaurant C (reservations available) and Restaurant B (waitlist)."

[0852] In this way, the system can respond quickly and accurately to user requests while reducing the burden on concierges.

[0853] The flow of the specific processing in Example 1 will be explained using Figure 11.

[0854] Step 1:

[0855] Users use their devices (smartphones or personal computers) to input their requests in voice or text format.

[0856] Specific action: The user speaks into their smartphone's microphone, saying, "I would like to make a dinner reservation at a high-end restaurant in Tokyo next Saturday night." Alternatively, they type a similar request as text.

[0857] Input: User's request in voice or text format.

[0858] Output: Audio or text data generated within the device.

[0859] Step 2:

[0860] The device sends the user's voice data to the server.

[0861] Specific operation: The smartphone app uploads the recorded audio file to the server. Text files are sent as is.

[0862] Input: User's voice data or text data.

[0863] Output: Audio or text data sent to the server.

[0864] Step 3:

[0865] The server converts the received audio data into text data using speech recognition technology.

[0866] Specific operation: The server calls the Google Cloud Speech-to-Text API and converts the audio data into text, "I would like to make a dinner reservation at a high-end restaurant in Tokyo next Saturday night."

[0867] Input: Audio data.

[0868] Output: Text data.

[0869] Step 4:

[0870] The server analyzes the text data using a natural language processing algorithm to extract the specific elements of the request.

[0871] Specific operation: The server uses the spaCy library to parse the text and extract elements such as "next Saturday," "evening," "in Tokyo," "high-end restaurant," and "dinner reservation."

[0872] Input: Text data.

[0873] Output: Extracted elements ("Date: Next Saturday", "Time: Evening", "Location: Tokyo", "Genre: High-end restaurant", "Content: Dinner reservation").

[0874] Step 5:

[0875] The server uses generative artificial intelligence to generate example answers based on the extracted elements.

[0876] Specific operation: The server inputs the following prompt to the OpenAI GPT-3 model: "The user said, 'I would like to make a dinner reservation at a high-end restaurant in Tokyo next Saturday night.' Please create a list of restaurant suggestions based on this."

[0877] Input: Extracted elements.

[0878] Output: Generated example answers (e.g., "Restaurant A (Reservations Available)", "Restaurant B (Waiting List)").

[0879] Step 6:

[0880] The server sends the generated sample answers to the administrator's terminal, where the administrator reviews and corrects the content.

[0881] Specific operation: The server sends a sample response to the administrator's terminal, and the administrator checks the latest reservation status with the pre-designated restaurant by phone or web and modifies the sample response as needed.

[0882] Input: Generated example answer.

[0883] Output: Example answer corrected by the administrator.

[0884] Step 7:

[0885] The server sends the corrected answer example to the user's terminal.

[0886] Specific operation: The server sends the corrected answer example to the user's smartphone or personal computer via email or app notification.

[0887] Input: Example answer corrected by the administrator.

[0888] Output: The final response sent to the user (e.g., "The only upscale restaurants available next Saturday evening are Restaurant C (reservations available) and Restaurant B (waitlist)").

[0889] (Application Example 1)

[0890] Next, we will explain Application Example 1. In the following explanation, the data processing device 12 will be referred to as the "server," and the headset-type terminal 314 will be referred to as the "terminal."

[0891] In modern brick-and-mortar stores, customers often demand prompt and accurate product guidance and recommendations, but there is a lack of high-quality and efficient service to meet this demand. Furthermore, manual service is costly and places a heavy burden on employees, leading to decreased customer satisfaction. In particular, while there is a growing demand for interactive service delivery using smart devices, appropriate systems are still lacking.

[0892] The specific processing performed by the specific processing unit 290 of the data processing device 12 in Application Example 1 is realized by the following means.

[0893] In this invention, the server includes a user terminal for the user to input requests in voice or text format; a conversion means for converting voice data received from the user into text data; an analysis means for analyzing the text data generated by the conversion means as a request using a natural language processing algorithm; a generation means for generating example answers using a generative AI based on the requests obtained by the analysis means; a transmission means for sending the example answers generated by the generation means to a concierge terminal for confirmation and correction by the concierge; a transmission means for sending the corrected answers returned from the concierge terminal to the user terminal; and a notification means that allows the user terminal to input requests via smart glasses and notifies the user of the final answer via the smart glasses. This makes it possible for customers in physical stores to receive quick and appropriate product information through voice input.

[0894] A "user terminal" is a device used by users to input requests in voice or text format.

[0895] "Conversion means" refers to a device or method for converting audio data received from a user into text data.

[0896] "Analysis means" refers to a device or method for analyzing text data generated by a conversion means as a request using a natural language processing algorithm.

[0897] "Generation means" refers to an apparatus or method for generating example answers using a generative AI based on requests obtained by an analysis means.

[0898] "Transmission means" refers to a device or method for transmitting example answers generated by the generation means to a concierge terminal, and for transmitting corrected answers returned from the concierge terminal to a user terminal.

[0899] "Smart glasses" are devices that users wear to input requests and receive responses visually.

[0900] "Notification means" refers to a device or method for notifying a user of the final response to a request through a user terminal (including smart glasses).

[0901] Modes for carrying out the invention

[0902] This invention is a system for providing customers with quick and appropriate product information in physical stores. Specific embodiments are described below.

[0903] System Program Overview

[0904] 1. Voice input and data conversion:

[0905] The user performs voice input while wearing smart glasses. For example, they might voice a request such as, "I want to find a shirt that matches these shoes."

[0906] The smart glasses incorporate voice recognition technology, converting received voice data into text data. This process utilizes Google's voice recognition API.

[0907] 2. Analysis of requests:

[0908] The converted text data is sent to the server. On the server side, a natural language processing algorithm (BERT, a transformer model from Hugging Face) is used to analyze the text data and extract the customer's request as specific elements. For example, information such as "Item: Shoes," "Request: Shirt," and "Purpose: Coordination support" is obtained.

[0909] 3. Generating example answers:

[0910] Based on the extracted elements, a generative AI (e.g., GPT-3) is used to generate example answers. The generated example answers include specific product information, such as "the blue shirt on shelf A" or "the white shirt on shelf B." An example of this prompt is: "User request: I want to find a shirt that matches these shoes. Please suggest specific products that match this request."

[0911] 4. Concierge review and correction:

[0912] The generated sample responses are sent to the concierge terminal. The concierge reviews the list of suggestions and makes revisions as needed. For example, they may update the list to reflect inventory status or additional product information, and then send the revised responses back to the server.

[0913] 5. Notification of final response:

[0914] The server notifies the user of the corrected final answer via the smart glasses they are wearing. This allows the user to receive the notification visually. For example, the guidance might say, "The shirts that match these shoes are the blue shirt on shelf A and the white shirt on shelf B."

[0915] Hardware and software to be used

[0916] Smart glasses: A device for voice input and notification display.

[0917] Server: The central hub responsible for converting audio data to text, analyzing text data, and managing generated response examples. It runs Google's speech recognition API, Hugging Face's BERT, and generative AI (GPT-3).

[0918] Natural Language Processing (NLP): These algorithms run on a server and are used to analyze and extract user requests.

[0919] Adding specific examples

[0920] For example, if a user enters a store and uses smart glasses to voice-input "I want to find a shirt that goes with these shoes," they will receive a notification a few seconds later saying, "We recommend the blue shirt on shelf A and the white shirt on shelf B." This experience allows users to quickly find products that meet their needs.

[0921] In this way, this system can improve the customer experience in physical stores and provide high-quality service while reducing the burden on employees.

[0922] The flow of a specific process in Application Example 1 will be explained using Figure 12.

[0923] Step 1:

[0924] The user wears smart glasses and inputs their request by voice. For example, they might say, "I want to find a shirt that matches these shoes." This input is collected via the smart glasses' microphone. The voice signal is then used as input for speech recognition processing to convert it into text.

[0925] Step 2:

[0926] Smart glasses convert audio data into text data. Specifically, they use Google's speech recognition API to convert audio data into text data. The input to this process is audio data, and the output is text data.

[0927] Step 3:

[0928] The converted text data is sent to the server. The server analyzes the received text data using a natural language processing algorithm (Hugging Face's BERT model). In this step, the input is text data, and the output is the specific elements of the analyzed request (e.g., "Item: Shoes", "Request: Shirt", "Purpose: Coordination Support").

[0929] Step 4:

[0930] The server uses a generative AI (GPT-3) to generate example responses based on the elements of the analyzed request. The input used in this process is the specific elements of the analyzed request, and the output is a specific product suggestion. For example, example responses might be "a blue shirt on shelf A" or "a white shirt on shelf B."

[0931] Step 5:

[0932] The generated sample responses are sent to the concierge terminal. The concierge reviews the list of suggestions and makes corrections as needed. Based on inventory information and new product information, they create more accurate suggestions. The input in this step is the generated sample response, and the output is the corrected sample response.

[0933] Step 6:

[0934] The revised answer examples are sent back to the server and compiled into the final answer. This final answer is then sent to the smart glasses worn by the user. For example, the notification might say, "The shirts that go with these shoes are the blue shirt on shelf A and the white shirt on shelf B." In this step, the input is the revised answer examples, and the output is the notification message displayed on the smart glasses.

[0935] In this way, the input data is sequentially processed and analyzed at each processing step, and finally, appropriate product information is provided to the user.

[0936] Furthermore, an emotion engine that estimates the user's emotions may be incorporated. That is, the identification processing unit 290 may use the emotion identification model 59 to estimate the user's emotions and perform identification processing using the user's emotions.

[0937] This invention is a system for efficiently processing user requests in voice or text format and further improving the quality of concierge services by combining it with an emotion engine that recognizes the user's emotions. Specific embodiments of the system are shown below.

[0938] How users can enter their requests

[0939] The process begins with the user entering their request using a device (e.g., a smartphone or PC). Users can enter their request via voice or text. For example, a user might voice-input, "I would like to make a dinner reservation at a high-end restaurant in Tokyo next Saturday night."

[0940] Converting audio data to text

[0941] The device sends the user's voice data to the server. The server uses speech recognition technology to convert the voice data into text data. For example, it might generate the text "I would like to make a dinner reservation at a high-end restaurant in Tokyo next Saturday night" from the voice data.

[0942] Request analysis

[0943] The server analyzes the generated text data using a natural language processing algorithm and extracts specific elements representing the user's request. This yields information such as "Date: Next Saturday," "Time: Evening," "Location: Tokyo," "Genre: High-end restaurant," and "Content: Dinner reservation."

[0944] Using an Emotion Engine

[0945] The server then uses an emotion engine to analyze the user's emotions from voice or text data. For example, it extracts emotional information such as "excited," "anxious," or "hurried." This emotional information influences the generative AI in subsequent processes, which is used to generate more appropriate responses that align with the user's feelings.

[0946] Generating example answers

[0947] The server uses generative AI to generate example responses based on extracted elements and emotional information. For example, the AI ​​might generate a response like, "We suggest the following restaurants: 1. Restaurant A (reservations available) 2. Restaurant B (waitlist)." By considering emotional information, it's possible to include additional perks for users who are excited, for example.

[0948] Concierge confirmation

[0949] The server sends the generated sample response to the concierge's terminal, where the concierge reviews and corrects the content. The concierge then contacts, for example, "Restaurant A" and "Restaurant B" to check the latest reservation status and updates the list if necessary.

[0950] Submit your final response

[0951] The revised response is sent back to the server, which then sends the final response to the user's device. The user can receive the final response via email or app notification. For example, the notification might say, "Here are some high-end restaurants available next Saturday evening. Restaurant A is fully booked, while Restaurant B is on the waiting list."

[0952] Specific example

[0953] The user speaks into their smartphone's microphone and says, "I'd like to make a dinner reservation at a high-end restaurant in Tokyo next Saturday night."

[0954] The device sends the voice data to the server. The server performs speech recognition and obtains text data that reads, "I would like to make a dinner reservation at a high-end restaurant in Tokyo next Saturday night."

[0955] The server analyzes the text data using a natural language processing algorithm and extracts requests such as "next Saturday," "evening," "in Tokyo," "high-end restaurant," and "dinner reservation."

[0956] The server uses an emotion engine to extract emotional information from the user's voice and text data, recognizing emotions such as "excited" or "hurried."

[0957] The server uses generative AI to generate a list based on the user's requests and recognized emotions. For example, "Restaurant A (Reservations Available)" and "Restaurant B (Waiting List)." Depending on the emotion, for instance, if the emotion is excitement, the system will generate a response that includes additional information about special offers.

[0958] This list is sent to the concierge terminal, where the concierge reviews it and makes corrections as needed.

[0959] The revised response is sent back to the server, which then sends the final response to the user's device. The user is notified that "The only upscale restaurants available next Saturday evening are Restaurant A (reservations available) and Restaurant B (waitlist)."

[0960] This system responds quickly and accurately to user requests and provides more personalized service by taking emotions into consideration. This reduces the burden on concierges and improves the quality of service.

[0961] The following describes the processing flow.

[0962] Step 1:

[0963] The user uses their device to input their request in voice or text format. For example, the user might say into their smartphone's microphone, "I'd like to make a dinner reservation at a high-end restaurant in Tokyo next Saturday night."

[0964] Step 2:

[0965] The terminal sends the input voice data to the server. The voice data is transmitted in digital format.

[0966] Step 3:

[0967] The server uses speech recognition technology to convert the received audio data into text data. For example, it might produce text data such as, "I would like to make a dinner reservation at a high-end restaurant in Tokyo next Saturday night."

[0968] Step 4:

[0969] The server uses a natural language processing algorithm to analyze the generated text data. This analysis extracts elements such as "next Saturday," "evening," "in Tokyo," "high-end restaurant," and "dinner reservation."

[0970] Step 5:

[0971] The server uses an emotion engine to analyze the user's emotions from voice or text data. For example, the emotion engine extracts emotional information such as "excited," "anxious," or "hurried."

[0972] Step 6:

[0973] The server uses generative AI to generate example responses based on emotional information and desired elements. By considering emotional information, more personalized responses are generated. For example, the AI ​​might generate a response like, "We suggest the following restaurants: 1. Restaurant A (reservations available) 2. Restaurant B (waitlist)." If the user is looking forward to something, special offers may also be added.

[0974] Step 7:

[0975] The server sends the generated example answer to the concierge terminal. The concierge terminal displays the received content.

[0976] Step 8:

[0977] The concierge reviews the sample responses received and makes corrections as needed. For example, the concierge will actually contact "Restaurant A" and "Restaurant B" to check the latest reservation status and update the list if necessary.

[0978] Step 9:

[0979] The concierge sends the revised answer back to the server. The server verifies the revised content.

[0980] Step 10:

[0981] The server sends the final response to the user's terminal. For example, the user might be notified that "The only upscale restaurants available next Saturday evening are Restaurant A (reservations available) and Restaurant B (waitlist)."

[0982] Through the steps outlined above, this system can efficiently respond to user requests and, by taking user emotions into consideration, improve the quality of concierge services.

[0983] (Example 2)

[0984] Next, we will describe Example 2. In the following description, the data processing device 12 will be referred to as the "server," and the headset-type terminal 314 will be referred to as the "terminal."

[0985] Traditional concierge services lacked the technology to appropriately and quickly handle specific user requests, and particularly struggled to respond while considering emotions. Therefore, improving service quality requires analyzing user requests, including their emotions, and generating appropriate responses based on that analysis.

[0986] The specific processing performed by the specific processing unit 290 of the data processing device 12 in Example 2 is realized by the following means. In this invention, the server includes an information terminal for the user to input requests in voice or text format, a conversion means for converting voice data received from the user into text data, an analysis means for analyzing the text data generated by the conversion means using a natural language processing algorithm, a generation means for generating example answers using generative artificial intelligence based on the requests and sentiment information obtained by the analysis means, a transmission means for sending the example answers generated by the generation means to a service provider terminal for confirmation and correction by the service provider, and a transmission means for sending the corrected answers returned from the service provider terminal to the user terminal. This enables the provision of quick and appropriate services that take into account the user's requests and sentiments.

[0987] An "information terminal" is a device used by users to input requests in voice or text format, and includes smartphones and personal computers.

[0988] "Conversion means" refers to a technology or method for converting audio data received from a user into text data, and which uses speech recognition technology.

[0989] "Analysis means" refers to a technology or method that analyzes text data generated by the conversion means using a natural language processing algorithm to extract user requests and emotional information.

[0990] "Generation means" refers to a technology or method that generates example responses using generative artificial intelligence based on requests and emotional information obtained by analysis means.

[0991] "Transmission means" refers to a technology or method for transmitting example answers generated by the generation means to a service provider terminal, and for transmitting the corrected answers to a user terminal.

[0992] A "service provider terminal" refers to a device used to review and correct generated sample answers, specifically the device used by the concierge.

[0993] A "natural language processing algorithm" is an algorithm used to analyze text data and extract information about desires and emotions, and it includes machine learning and data analysis techniques.

[0994] "Generative artificial intelligence" refers to artificial intelligence technology that generates example responses based on analyzed requests and sentiment information, such as using a generative language model.

[0995] "Emotional information" refers to emotional information analyzed from the user's input voice or text data, and includes emotions such as excitement, anxiety, and urgency.

[0996] Modes for carrying out the invention

[0997] This invention is a system for efficiently processing user requests in voice or text format and further improving the quality of concierge services by combining it with an emotion engine that recognizes the user's emotions. Specific embodiments of the system are shown below.

[0998] How users can enter their requests

[0999] The process begins with the user entering their request in voice or text format using an information terminal (such as a smartphone or personal computer). For example, a user could voice-input a request like, "I would like to make a dinner reservation at a high-end restaurant in Tokyo next Saturday night." Users can input voice data using the terminal's microphone or input their request as text data using a keyboard.

[1000] Converting audio data to text

[1001] The device sends the user's voice data to the server. The server uses speech recognition technology to convert the voice data into text data. Specifically, it uses a speech recognition service such as the Google Cloud Speech-to-Text API to convert the voice data into text data. For example, it generates text data such as "I would like to make a dinner reservation at a high-end restaurant in Tokyo next Saturday night" from the voice data.

[1002] Request analysis

[1003] The server analyzes the generated text data using natural language processing algorithms to extract specific elements representing the user's request. Algorithms used include SpaCy and the Google Cloud Natural Language API. This analysis yields information such as "Date: Next Saturday," "Time: Evening," "Location: Tokyo," "Genre: High-end restaurant," and "Content: Dinner reservation."

[1004] Using an Emotion Engine

[1005] The server then uses an emotion engine to analyze the user's emotions from voice or text data. For example, it uses IBM Watson Tone Analyzer to extract emotional information such as "excited," "anxious," and "hurried." This emotional information influences the generative artificial intelligence in subsequent processes, which is used to generate more appropriate responses that align with the user's feelings.

[1006] Generating example answers

[1007] The server generates example responses using generative artificial intelligence (e.g., OpenAI GPT-3) based on extracted elements and emotional information. For example, the generative AI might generate a response such as, "We suggest the following restaurants: 1. Restaurant A (reservations available) 2. Restaurant B (waitlist)." By considering emotional information, it's possible to include additional perks for users who are excited, for example.

[1008] Concierge confirmation

[1009] The server sends the generated sample response to the service provider's terminal, where the service provider reviews and corrects the content. The service provider (e.g., a concierge) then actually contacts restaurants such as "Restaurant A" and "Restaurant B" to check their latest reservation status and updates the list if necessary.

[1010] Submit your final response

[1011] The revised response from the service provider's terminal is sent back to the server, which then sends the final response to the user's information terminal. The user can receive the final response via email or app notification. For example, the notification might say, "We have a list of upscale restaurants available next Saturday evening. Restaurant A is fully booked, while Restaurant B is on the waiting list."

[1012] Examples of prompt statements

[1013] If a user speaks into their smartphone's microphone and says, "I'd like to make a dinner reservation at a high-end restaurant in Tokyo next Saturday night," the following prompt will be used for the generative artificial intelligence.

[1014] "User request: "I want to make a dinner reservation at a high-end restaurant in Tokyo next Saturday night." Sentiment information: "Excited, in a hurry." Please generate a list of responses: Example) - Restaurant A (Reservations available) - Restaurant B (Waiting list) Perks information: Add perks information based on the excited sentiment."

[1015] This allows the system to respond quickly and accurately to user requests and provide more personalized service by taking emotions into consideration. In this way, it is possible to reduce the burden on concierges and improve the quality of service.

[1016] The flow of the specific processing in Example 2 will be explained using Figure 13.

[1017] Step 1: Enter your request

[1018] User

[1019] Users input their requests in voice or text format using an information terminal (e.g., smartphone, personal computer).

[1020] Input: Audio data or text data

[1021] Specific action: For example, the user speaks into their smartphone and says, "I want to make a dinner reservation at a high-end restaurant in Tokyo next Saturday night."

[1022] Step 2: Converting audio data to text

[1023] terminal

[1024] The device sends the audio data to the server.

[1025] Input: Audio data

[1026] Output: Sending audio data to the server

[1027] Specific action: Send audio data to the server in digital format.

[1028] server

[1029] The server uses speech recognition technologies such as the Google Cloud Speech-to-Text API to convert audio data into text data.

[1030] Input: Audio data

[1031] Output: Text data

[1032] Specific operation: Generate the text "I would like to make a dinner reservation at a high-end restaurant in Tokyo next Saturday night" from the audio data.

[1033] Step 3: Analyzing the requirements

[1034] server

[1035] The server analyzes the generated text data using natural language processing algorithms (e.g., SpaCy, Google Cloud Natural Language API) to extract the specific elements of the request.

[1036] Input: Text data

[1037] Output: Elements of the request (e.g., "Date: Next Saturday", "Time: Evening", "Location: Tokyo", "Genre: High-end restaurant", "Content: Dinner reservation")

[1038] Specific operation: Analyze the specific elements of the request and store them in the database.

[1039] Step 4: Utilizing the Emotional Engine

[1040] server

[1041] The server uses an emotion engine (e.g., IBM Watson Tone Analyzer) to analyze emotional information from audio or text data.

[1042] Input: Audio data or text data

[1043] Output: Emotional information (e.g., "excited," "anxious," "hurried")

[1044] Specific operation: Use an emotion engine to analyze emotions from speech or text and store them in a database.

[1045] Step 5: Generating example answers

[1046] server

[1047] The server generates example responses using a generative AI (e.g., OpenAI GPT-3) based on the elements of the request and emotional information.

[1048] Input: Elements of requests, emotional information

[1049] Output: Example answer (Example: "Restaurant A (Reservations accepted)", "Restaurant B (Waiting list)")

[1050] Specific operation: Use generative AI to generate responses that respond to requests and emotions.

[1051] Step 6: Concierge Confirmation

[1052] server

[1053] The server sends the generated example answers to the service provider's terminal.

[1054] Input: Example Answer

[1055] Output: Sent to the service provider's terminal

[1056] Specific action: Send example answers to the service provider's terminal.

[1057] Service provider

[1058] The service provider will review the responses and make corrections as necessary.

[1059] Input: Example Answer

[1060] Output: Revised answer example

[1061] Specific action: The service provider inquires about the latest booking status and corrects the response.

[1062] Step 7: Submit your final response

[1063] Service provider

[1064] The revised answer example is sent back to the server.

[1065] Input: Revised answer example

[1066] Output: Send to server

[1067] Specific action: Send the corrected response to the server.

[1068] server

[1069] The server sends the final response to the user's information terminal.

[1070] Input: Revised answer example

[1071] Output: Sending the final response to the user's information terminal.

[1072] Specific action: Send the final response to the user via email or app notification.

[1073] In this way, by responding quickly and accurately to user requests and taking their emotions into consideration, it becomes possible to provide more personalized services.

[1074] (Application Example 2)

[1075] Next, we will explain application example 2. In the following explanation, the data processing device 12 will be referred to as the "server," and the headset-type terminal 314 will be referred to as the "terminal."

[1076] Traditional concierge services faced challenges in providing prompt service due to the extensive work required to accurately understand and incorporate user requests. Furthermore, they lacked consideration for user emotions, resulting in a lack of personalization and low satisfaction. In short, the uniform approach to users and the inability to provide personalized suggestions based on emotions led to a poor user experience.

[1077] In Application Example 2, the specific processing performed by the specific processing unit 290 of the data processing device 12 is realized by the following means. In this invention, the server includes a user terminal for the user to input requests in voice or text format, a conversion means for converting voice data received from the user into text data, an analysis means for analyzing the text data generated by the conversion means as a request using a natural language processing algorithm, a generation means for generating example answers using a generative AI based on the requests and emotional information from the analysis means, a transmission means for sending the example answers generated by the generation means to a concierge terminal for confirmation and correction by the concierge, a transmission means for sending the corrected answers returned from the concierge terminal to the user terminal, and an emotional analysis means for extracting emotional information from the user's voice data and text data. This makes it possible to quickly and accurately analyze user requests and provide personalized services based on emotions.

[1078] A "user terminal" is a device used by users to input requests in voice or text format.

[1079] "Conversion means" refers to technical means for converting audio data received from a user into text data.

[1080] "Analysis means" refers to technical means for analyzing text data generated by the conversion means as a request using a natural language processing algorithm.

[1081] "Generation means" refers to technical means for generating example responses using a generative AI based on requests and emotional information obtained through analysis means.

[1082] "Transmission means" refers to a technical means for sending example answers generated by the generation means to a concierge terminal, for confirmation and correction by the concierge, and for sending the corrected answers returned from the concierge terminal to the user terminal.

[1083] "Emotion analysis methods" refer to technical means for extracting emotional information from user voice data and text data.

[1084] A "natural language processing algorithm" is an algorithm that analyzes text data to understand and extract user requests.

[1085] "Generative AI" is an artificial intelligence technology that generates appropriate response examples based on user requests and emotional information.

[1086] A "concierge terminal" is a device used to review and correct generated sample answers.

[1087] "Emotional information" refers to information about a user's emotions extracted from their voice data and text data.

[1088] This invention is a system for further improving the quality of concierge services by efficiently processing requests entered by users in voice or text format, and by combining this with an emotion engine that recognizes the user's emotions.

[1089] The system includes the following components:

[1090] 1. User terminal: A device used by users to input requests in voice or text format. This includes smartphones, PCs, and smart glasses.

[1091] 2. Conversion method: Speech recognition technology is used to convert the audio data received from the user into text data. Specifically, the SpeechRecognition library is used to convert the audio data into text data.

[1092] 3. Analysis Method: The text data generated by the transformation method is analyzed as a request using a natural language processing algorithm. This uses the natural language processing model from the Transformers library.

[1093] 4. Emotion Analysis Method: An emotion analysis model is used to extract emotional information from the user's voice data and text data. Specifically, the Transformers emotion analysis model is used to obtain emotional information.

[1094] 5. Generation Method: Based on the requests and sentiment information obtained through the analysis method, a generative AI model (such as OpenAI GPT-3) is used to generate example responses. This process uses OpenAI's GPT-3 as an API and takes into account the relationship between the prompt text and the generated text.

[1095] 6. Transmission method: The example answer generated by the generation method is sent to the concierge terminal for confirmation and correction by the concierge. Furthermore, the corrected answer returned from the concierge terminal is sent to the user terminal.

[1096] Explanation of specific processing examples:

[1097] The user inputs their request by voice into their smartphone's microphone, saying, "I would like to make a dinner reservation at a high-end restaurant in Tokyo next Saturday night."

[1098] The user's terminal sends the voice data to the server, which uses the SpeechRecognition library to convert the voice data into text data that reads, "I would like to make a dinner reservation at a high-end restaurant in Tokyo next Saturday night."

[1099] The converted text data is analyzed by Transformers' natural language processing model, and requests such as "Date: Next Saturday", "Time: Evening", "Location: Tokyo", "Genre: High-end restaurant", and "Content: Dinner reservation" are extracted.

[1100] An emotion analysis model analyzes the same text data and obtains emotional information such as "excited" or "hurried."

[1101] Based on the analyzed requests and sentiment information, the server uses GPT-3 to generate example responses with the following prompts:

[1102] User Input: I'm looking for recommended baby products in the store.

[1103] Emotion: Enjoyment

[1104] Response:

[1105] For example, the AI ​​can generate sample responses such as, "In this baby products section, we recommend our new skin-friendly diapers and soft swaddles."

[1106] The generated sample answers are sent to the concierge terminal, where the concierge reviews the content and makes corrections as needed.

[1107] The corrected response is sent back to the server and finally delivered to the user's terminal.

[1108] This system enables the rapid and accurate analysis of user requests and the provision of personalized, emotion-based services. This improves the user experience and reduces the workload on concierges.

[1109] The flow of a specific process in Application Example 2 will be explained using Figure 14.

[1110] Step 1:

[1111] The user performs voice input. The user uses a smartphone or smart glasses to input their request by voice. For example, they might say, "I would like to make a dinner reservation at a high-end restaurant in Tokyo next Saturday night." This voice data is recorded on the device. The input is voice data, and the output is the voice file recorded on the device.

[1112] Step 2:

[1113] The device sends audio data to the server. The device sends the recorded audio data to the server. The server uses speech recognition technology to convert this audio data into text data. The SpeechRecognition library is used for this conversion. The input is audio data, and the output is text data that says, "I would like to make a dinner reservation at a high-end restaurant in Tokyo next Saturday night."

[1114] Step 3:

[1115] The server analyzes the text data. The server then analyzes the transformed text data using a natural language processing algorithm. This process uses the Transformers library to extract user requests (e.g., "Date: Next Saturday", "Time: Evening", "Location: Tokyo", "Genre: High-end restaurant", "Content: Dinner reservation"). The input is text data, and the output is the analyzed elemental information.

[1116] Step 4:

[1117] The server extracts emotional information. The server then uses emotion analysis tools to analyze the user's emotional information from the text data. This process utilizes the Transformers emotion analysis model. For example, emotions such as "excited" and "hurried" are extracted. The input is text data, and the output is emotional information.

[1118] Step 5:

[1119] The server generates example answers using generative AI. Based on the analyzed requests and sentiment information, the server generates example answers using a generative AI model such as GPT-3. This process uses OpenAI's GPT-3 as an API, sending the generated text as a prompt. The input is requests and sentiment information, and the output is the suggested answer. For example, the prompt might look like this:

[1120] User Input: I'm looking for recommended baby products in the store.

[1121] Emotion: Enjoyment

[1122] Response:

[1123] Step 6:

[1124] The server sends example answers to the concierge terminal. The generated example answers are sent to the concierge terminal, where the concierge reviews the content and makes corrections as needed. The input is the suggested example answer, and the output is the corrected example answer.

[1125] Step 7:

[1126] The concierge terminal sends the corrected answer back to the server. After the concierge reviews and makes corrections, the corrected answer is sent back to the server. The input is an example of the corrected answer, and the output is the final example answer.

[1127] Step 8:

[1128] The server sends the final answer to the user's device. The user can receive the final answer via email or app notification. The input is an example of the final answer, and the output is the final answer notified to the user. For example, the user might be notified of the final answer, "The upscale restaurants available next Saturday evening are Restaurant A (reservations available) and Restaurant B (waitlist)."

[1129] The specific processing unit 290 transmits the result of the specific processing to the headset terminal 314. In the headset terminal 314, the control unit 46A causes the speaker 240 and display 343 to output the result of the specific processing. The microphone 238 acquires audio indicating user input for the result of the specific processing. The control unit 46A transmits the audio data indicating user input acquired by the microphone 238 to the data processing unit 12. In the data processing unit 12, the specific processing unit 290 acquires the audio data.

[1130] Data generation model 58 is a type of so-called generative AI (Artificial Intelligence). One example of data generation model 58 is ChatGPT (Internet search<URL: https: / / openai.com / blog / chatgpt> ), Gemini (Internet search) <url: https: gemini.google.com ?hl="ja">Examples of generative AI include the following. The data generation model 58 is obtained by performing deep learning on a neural network. The data generation model 58 is input with prompts containing instructions, and with inference data such as audio data representing speech, text data representing text, and image data representing images. The data generation model 58 infers from the input inference data according to the instructions indicated by the prompts, and outputs the inference results in data formats such as audio data and text data. Here, inference refers to, for example, analysis, classification, prediction, and / or summarization.

[1131] In the above embodiment, an example was given in which specific processing is performed by the data processing device 12, but the technology of this disclosure is not limited thereto, and specific processing may also be performed by the headset terminal 314.

[1132] [Fourth Embodiment]

[1133] Figure 7 shows an example of the configuration of the data processing system 410 according to the fourth embodiment.

[1134] As shown in Figure 7, the data processing system 410 includes a data processing device 12 and a robot 414. An example of the data processing device 12 is a server.

[1135] The data processing device 12 comprises a computer 22, a database 24, and a communication interface 26. The computer 22 is an example of a "computer" related to the technology of this disclosure. The computer 22 comprises a processor 28, RAM 30, and storage 32. The processor 28, RAM 30, and storage 32 are connected to a bus 34. The database 24 and the communication interface 26 are also connected to the bus 34. The communication interface 26 is connected to a network 54. An example of the network 54 is a WAN (Wide Area Network) and / or a LAN (Local Area Network).

[1136] The robot 414 includes a computer 36, a microphone 238, a speaker 240, a camera 42, a communication interface 44, and a controlled object 443. The computer 36 includes a processor 46, RAM 48, and storage 50. The processor 46, RAM 48, and storage 50 are connected to a bus 52. The microphone 238, speaker 240, camera 42, and controlled object 443 are also connected to the bus 52.

[1137] The microphone 238 receives voice signals from the user 20 and receives instructions from the user 20. The microphone 238 captures the voice signals from the user 20, converts the captured voice into audio data, and outputs it to the processor 46. The speaker 240 outputs audio according to the instructions from the processor 46.

[1138] Camera 42 is a small digital camera equipped with an optical system including a lens, aperture, and shutter, and an image sensor such as a CMOS (Complementary Metal-Oxide-Semiconductor) image sensor or a CCD (Charge Coupled Device) image sensor, and captures images of the area around the user 20 (for example, an imaging range defined by a field of view equivalent to the width of a typical healthy person's field of vision).

[1139] Communication interface 44 is connected to network 54. Communication interfaces 44 and 26 are responsible for the exchange of various information between processor 46 and processor 28 via network 54. The exchange of various information between processor 46 and processor 28 using communication interfaces 44 and 26 is performed in a secure manner.

[1140] The controlled object 443 includes a display device, LEDs in the eyes, and motors that drive the arms, hands, and feet. The posture and gestures of the robot 414 are controlled by controlling the motors of the arms, hands, and feet. Some of the robot 414's emotions can be expressed by controlling these motors. Furthermore, the robot 414's facial expressions can also be expressed by controlling the illumination state of the LEDs in its eyes.

[1141] Figure 8 shows an example of the main functions of the data processing device 12 and the robot 414. As shown in Figure 8, the data processing device 12 performs specific processing using the processor 28. The storage 32 stores the specific processing program 56.

[1142] The specific processing program 56 is an example of a "program" relating to the technology of this disclosure. The processor 28 reads the specific processing program 56 from the storage 32 and executes the read specific processing program 56 on the RAM 30. The specific processing is realized by the processor 28 operating as a specific processing unit 290 in accordance with the specific processing program 56 executed on the RAM 30.

[1143] The storage 32 stores the data generation model 58 and the emotion identification model 59. The data generation model 58 and the emotion identification model 59 are used by the identification processing unit 290.

[1144] In robot 414, the processor 46 performs the reception output processing. The storage 50 stores the reception output program 60. The processor 46 reads the reception output program 60 from the storage 50 and executes the read reception output program 60 on the RAM 48. The reception output processing is realized by the processor 46 operating as a control unit 46A according to the reception output program 60 executed on the RAM 48.

[1145] Next, the specific processing performed by the specific processing unit 290 of the data processing device 12 will be described. In the following description, the data processing device 12 will be referred to as the "server" and the robot 414 as the "terminal".

[1146] This invention is a system for efficiently processing user requests entered in voice or text format and improving the quality of concierge services. Specific embodiments of the system are described below.

[1147] How users can enter their requests

[1148] The process begins with the user entering their request using a device (e.g., a smartphone or PC). Users can enter their request via voice or text. For example, a user might voice-input, "I would like to make a dinner reservation at a high-end restaurant in Tokyo next Saturday night."

[1149] Converting audio data to text

[1150] The device sends the user's voice data to the server. The server uses speech recognition technology to convert the voice data into text data. For example, it might generate the text "I would like to make a dinner reservation at a high-end restaurant in Tokyo next Saturday night" from the voice data.

[1151] Request analysis

[1152] The server analyzes the generated text data using a natural language processing algorithm and extracts specific elements representing the user's request. This yields information such as "Date: Next Saturday," "Time: Evening," "Location: Tokyo," "Genre: High-end restaurant," and "Content: Dinner reservation."

[1153] Generating example answers

[1154] The server uses generative AI to generate example answers based on the extracted elements. For example, the AI ​​might generate an example answer such as, "We suggest the following restaurants: 1. Restaurant A (reservations available) 2. Restaurant B (waitlist)."

[1155] Concierge confirmation

[1156] The server sends the generated sample response to the concierge's terminal, where the concierge reviews and corrects the content. The concierge then actually contacts, for example, "Restaurant A" and "Restaurant B" to check the latest reservation status and corrects the sample response as needed.

[1157] Submit your final response

[1158] The revised response is sent back to the server, which then sends the final response to the user's device. The user can receive the final response via email or app notification. For example, the notification might say, "Here are some high-end restaurants available next Saturday evening. Restaurant A is fully booked, while Restaurant B is on the waiting list."

[1159] Specific example

[1160] The user speaks into their smartphone's microphone and says, "I'd like to make a dinner reservation at a high-end restaurant in Tokyo next Saturday night."

[1161] The device sends the voice data to the server. The server performs speech recognition and obtains text data that reads, "I would like to make a dinner reservation at a high-end restaurant in Tokyo next Saturday night."

[1162] The server analyzes the text data using a natural language processing algorithm and extracts requests such as "next Saturday," "evening," "in Tokyo," "high-end restaurant," and "dinner reservation."

[1163] The server uses generative AI to generate a list of restaurants that match the user's request. For example, "Restaurant A (Reservations Available)" and "Restaurant B (Waiting List)".

[1164] This list is sent to the concierge terminal, where the concierge reviews it and makes corrections as needed.

[1165] The revised response is sent back to the server, which then sends the final response to the user's device. The user is notified in a format such as, "The only upscale restaurants available next Saturday evening are Restaurant C (reservations available) and Restaurant B (waitlist)."

[1166] In this way, the system can respond quickly and accurately to user requests while reducing the burden on concierges.

[1167] The following describes the processing flow.

[1168] Step 1:

[1169] The user uses their device to input their request in voice or text format. For example, the user might say into their smartphone's microphone, "I'd like to make a dinner reservation at a high-end restaurant in Tokyo next Saturday night."

[1170] Step 2:

[1171] The terminal sends the input voice data to the server. The voice data is transmitted in digital format.

[1172] Step 3:

[1173] The server uses speech recognition technology to convert the received audio data into text data. For example, it might produce text data such as, "I would like to make a dinner reservation at a high-end restaurant in Tokyo next Saturday night."

[1174] Step 4:

[1175] The server uses a natural language processing algorithm to analyze the generated text data. This analysis extracts elements such as "next Saturday," "evening," "in Tokyo," "high-end restaurant," and "dinner reservation."

[1176] Step 5:

[1177] The server uses generative AI to generate example answers based on the extracted elements. For example, the AI ​​might generate an example answer such as, "We suggest the following restaurants: 1. Restaurant A (reservations available) 2. Restaurant B (waitlist)."

[1178] Step 6:

[1179] The server sends the generated example answer to the concierge terminal. The concierge terminal displays the received content.

[1180] Step 7:

[1181] The concierge reviews the sample responses received and makes corrections as needed. For example, the concierge will actually contact "Restaurant A" and "Restaurant B" to check the latest reservation status and update the list if necessary.

[1182] Step 8:

[1183] The concierge sends the revised answer back to the server. The server verifies the revised content.

[1184] Step 9:

[1185] The server sends the final response to the user's terminal. For example, the user might be notified that "The only upscale restaurants available next Saturday evening are Restaurant C (reservations available) and Restaurant B (waitlist)."

[1186] Through the steps outlined above, this system efficiently responds to user requests, reduces the burden on concierges, and improves service quality.

[1187] (Example 1)

[1188] Next, we will describe Example 1. In the following description, the data processing device 12 will be referred to as the "server" and the robot 414 as the "terminal".

[1189] Traditional concierge services have faced challenges in quickly and accurately understanding user requests and generating appropriate responses. In particular, voice input requires significant time and manpower for the process of converting speech to text and then understanding that text. Furthermore, maintaining the quality of the generated responses necessitates manual review and correction by concierges, which further reduces efficiency.

[1190] The identification process performed by the identification processing unit 290 of the data processing device 12 in Example 1 is realized by the following means.

[1191] In this invention, the server includes a user terminal for the user to input requests in voice or text format, a conversion means for converting voice data received from the user into text data, an analysis means for analyzing the text data generated by the conversion means as a request using a natural language processing algorithm, a generation means for generating example answers using generative artificial intelligence based on the requests obtained by the analysis means, a transmission means for sending the example answers generated by the generation means to an administrator terminal for confirmation and correction by the administrator, a transmission means for sending the corrected answers returned from the administrator terminal to the user terminal, and a computer for automatically generating information that matches the user's requests. This makes it possible to quickly and accurately understand the user's requests and efficiently provide appropriate answers.

[1192] A "user terminal" is a device used by users to input requests in voice or text format, and specifically includes smartphones and personal computers.

[1193] "Conversion means" refers to a function or device for converting audio data received from a user into text data, and includes those that utilize speech recognition technology.

[1194] "Analysis means" refers to a function or device for analyzing text data generated by the conversion means as a request using a natural language processing algorithm.

[1195] "Generation means" refers to a function or device for generating example answers using generative artificial intelligence based on requests obtained by analysis means.

[1196] "Transmission means" refers to a function or device for transmitting example answers and modified answers generated by the generation means to administrator terminals and user terminals.

[1197] An "administrator terminal" is a device used by administrators to review and modify generated sample answers, and specifically includes personal computers and tablets.

[1198] An "electronic computing device" is a device that performs functions and processes to automatically generate information that meets the user's requirements.

[1199] System Configuration

[1200] This invention is a system for efficiently processing user requests entered in voice or text format and improving the quality of concierge services. This system includes the following elements:

[1201] 1. User terminal:

[1202] These are devices that allow users to input requests in voice or text format, such as smartphones and personal computers.

[1203] 2. Server:

[1204] It has a conversion means for converting audio data received from a user into text data.

[1205] We will use the Google Cloud Speech-to-Text API as our speech recognition technology.

[1206] The system includes an analysis means for analyzing text data generated by a conversion means as a request using a natural language processing algorithm.

[1207] We will use spaCy or the NLTK library as natural language processing algorithms.

[1208] It has a generation means for generating example answers using generative artificial intelligence based on the request.

[1209] As a generative artificial intelligence, we will use the OpenAI GPT-3 model.

[1210] The system has a means for sending the generated sample answers to the administrator's terminal for review and correction by the administrator.

[1211] It has a means for sending the corrected answer to the user's terminal.

[1212] 3. Administrator terminal:

[1213] This device allows administrators to review and correct generated sample answers, and examples include personal computers and tablets.

[1214] Operation details

[1215] First, the user uses a device (smartphone or personal computer) to input their request in voice or text format. For example, the user might say into their smartphone's microphone, "I would like to make a dinner reservation at a high-end restaurant in Tokyo next Saturday night."

[1216] Next, the device sends the audio data to the server. The server uses the Google Cloud Speech-to-Text API to convert the audio data into text data, obtaining the text "I would like to make a dinner reservation at a high-end restaurant in Tokyo next Saturday night."

[1217] The server then analyzes the text data using a natural language processing algorithm (for example, the spaCy library) to extract requests such as "next Saturday," "evening," "in Tokyo," "high-end restaurant," and "dinner reservation."

[1218] The server then generates example answers based on the elements extracted using OpenAI's GPT-3 model. An example of the prompt text used here is as follows:

[1219] "The user said, 'I want to make a dinner reservation at a high-end restaurant in Tokyo next Saturday night.' Please create a list of restaurant suggestions based on this."

[1220] The generated example responses will be in the format of "Restaurant A (Reservations Available)" or "Restaurant B (Waiting List)". The server sends these example responses to the administrator's terminal, where the administrator actually contacts the restaurants to check the latest reservation status and corrects the example responses as needed.

[1221] The revised response is sent back to the server, and the final response is sent to the user's device (smartphone or personal computer). The user can receive the final response via email or app notification, such as, "The only upscale restaurants available next Saturday evening are Restaurant C (reservations available) and Restaurant B (waitlist)."

[1222] In this way, the system can respond quickly and accurately to user requests while reducing the burden on concierges.

[1223] The flow of the specific processing in Example 1 will be explained using Figure 11.

[1224] Step 1:

[1225] Users use their devices (smartphones or personal computers) to input their requests in voice or text format.

[1226] Specific action: The user speaks into their smartphone's microphone, saying, "I would like to make a dinner reservation at a high-end restaurant in Tokyo next Saturday night." Alternatively, they type a similar request as text.

[1227] Input: User's request in voice or text format.

[1228] Output: Audio or text data generated within the device.

[1229] Step 2:

[1230] The device sends the user's voice data to the server.

[1231] Specific operation: The smartphone app uploads the recorded audio file to the server. Text files are sent as is.

[1232] Input: User's voice data or text data.

[1233] Output: Audio or text data sent to the server.

[1234] Step 3:

[1235] The server converts the received audio data into text data using speech recognition technology.

[1236] Specific operation: The server calls the Google Cloud Speech-to-Text API and converts the audio data into text, "I would like to make a dinner reservation at a high-end restaurant in Tokyo next Saturday night."

[1237] Input: Audio data.

[1238] Output: Text data.

[1239] Step 4:

[1240] The server analyzes the text data using a natural language processing algorithm to extract the specific elements of the request.

[1241] Specific operation: The server uses the spaCy library to parse the text and extract elements such as "next Saturday," "evening," "in Tokyo," "high-end restaurant," and "dinner reservation."

[1242] Input: Text data.

[1243] Output: Extracted elements ("Date: Next Saturday", "Time: Evening", "Location: Tokyo", "Genre: High-end restaurant", "Content: Dinner reservation").

[1244] Step 5:

[1245] The server uses generative artificial intelligence to generate example answers based on the extracted elements.

[1246] Specific operation: The server inputs the following prompt to the OpenAI GPT-3 model: "The user said, 'I would like to make a dinner reservation at a high-end restaurant in Tokyo next Saturday night.' Please create a list of restaurant suggestions based on this."

[1247] Input: Extracted elements.

[1248] Output: Generated example answers (e.g., "Restaurant A (Reservations Available)", "Restaurant B (Waiting List)").

[1249] Step 6:

[1250] The server sends the generated sample answers to the administrator's terminal, where the administrator reviews and corrects the content.

[1251] Specific operation: The server sends a sample response to the administrator's terminal, and the administrator checks the latest reservation status with the pre-designated restaurant by phone or web and modifies the sample response as needed.

[1252] Input: Generated example answer.

[1253] Output: Example answer corrected by the administrator.

[1254] Step 7:

[1255] The server sends the corrected answer example to the user's terminal.

[1256] Specific operation: The server sends the corrected answer example to the user's smartphone or personal computer via email or app notification.

[1257] Input: Example answer corrected by the administrator.

[1258] Output: The final response sent to the user (e.g., "The only upscale restaurants available next Saturday evening are Restaurant C (reservations available) and Restaurant B (waitlist)").

[1259] (Application Example 1)

[1260] Next, we will explain Application Example 1. In the following explanation, the data processing device 12 will be referred to as the "server" and the robot 414 as the "terminal".

[1261] In modern brick-and-mortar stores, customers often demand prompt and accurate product guidance and recommendations, but there is a lack of high-quality and efficient service to meet this demand. Furthermore, manual service is costly and places a heavy burden on employees, leading to decreased customer satisfaction. In particular, while there is a growing demand for interactive service delivery using smart devices, appropriate systems are still lacking.

[1262] The specific processing performed by the specific processing unit 290 of the data processing device 12 in Application Example 1 is realized by the following means.

[1263] In this invention, the server includes a user terminal for the user to input requests in voice or text format; a conversion means for converting voice data received from the user into text data; an analysis means for analyzing the text data generated by the conversion means as a request using a natural language processing algorithm; a generation means for generating example answers using a generative AI based on the requests obtained by the analysis means; a transmission means for sending the example answers generated by the generation means to a concierge terminal for confirmation and correction by the concierge; a transmission means for sending the corrected answers returned from the concierge terminal to the user terminal; and a notification means that allows the user terminal to input requests via smart glasses and notifies the user of the final answer via the smart glasses. This makes it possible for customers in physical stores to receive quick and appropriate product information through voice input.

[1264] A "user terminal" is a device used by users to input requests in voice or text format.

[1265] "Conversion means" refers to a device or method for converting audio data received from a user into text data.

[1266] "Analysis means" refers to a device or method for analyzing text data generated by a conversion means as a request using a natural language processing algorithm.

[1267] "Generation means" refers to an apparatus or method for generating example answers using a generative AI based on requests obtained by an analysis means.

[1268] "Transmission means" refers to a device or method for transmitting example answers generated by the generation means to a concierge terminal, and for transmitting corrected answers returned from the concierge terminal to a user terminal.

[1269] "Smart glasses" are devices that users wear to input requests and receive responses visually.

[1270] "Notification means" refers to a device or method for notifying a user of the final response to a request through a user terminal (including smart glasses).

[1271] Modes for carrying out the invention

[1272] This invention is a system for providing customers with quick and appropriate product information in physical stores. Specific embodiments are described below.

[1273] System Program Overview

[1274] 1. Voice input and data conversion:

[1275] The user performs voice input while wearing smart glasses. For example, they might voice a request such as, "I want to find a shirt that matches these shoes."

[1276] The smart glasses incorporate voice recognition technology, converting received voice data into text data. This process utilizes Google's voice recognition API.

[1277] 2. Analysis of requests:

[1278] The converted text data is sent to the server. On the server side, a natural language processing algorithm (BERT, a transformer model from Hugging Face) is used to analyze the text data and extract the customer's request as specific elements. For example, information such as "Item: Shoes," "Request: Shirt," and "Purpose: Coordination support" is obtained.

[1279] 3. Generating example answers:

[1280] Based on the extracted elements, a generative AI (e.g., GPT-3) is used to generate example answers. The generated example answers include specific product information, such as "the blue shirt on shelf A" or "the white shirt on shelf B." An example of this prompt is: "User request: I want to find a shirt that matches these shoes. Please suggest specific products that match this request."

[1281] 4. Concierge review and correction:

[1282] The generated sample responses are sent to the concierge terminal. The concierge reviews the list of suggestions and makes revisions as needed. For example, they may update the list to reflect inventory status or additional product information, and then send the revised responses back to the server.

[1283] 5. Notification of final response:

[1284] The server notifies the user of the corrected final answer via the smart glasses they are wearing. This allows the user to receive the notification visually. For example, the guidance might say, "The shirts that match these shoes are the blue shirt on shelf A and the white shirt on shelf B."

[1285] Hardware and software to be used

[1286] Smart glasses: A device for voice input and notification display.

[1287] Server: The central hub responsible for converting audio data to text, analyzing text data, and managing generated response examples. It runs Google's speech recognition API, Hugging Face's BERT, and generative AI (GPT-3).

[1288] Natural Language Processing (NLP): These algorithms run on a server and are used to analyze and extract user requests.

[1289] Adding specific examples

[1290] For example, if a user enters a store and uses smart glasses to voice-input "I want to find a shirt that goes with these shoes," they will receive a notification a few seconds later saying, "We recommend the blue shirt on shelf A and the white shirt on shelf B." This experience allows users to quickly find products that meet their needs.

[1291] In this way, this system can improve the customer experience in physical stores and provide high-quality service while reducing the burden on employees.

[1292] The flow of a specific process in Application Example 1 will be explained using Figure 12.

[1293] Step 1:

[1294] The user wears smart glasses and inputs their request by voice. For example, they might say, "I want to find a shirt that matches these shoes." This input is collected via the smart glasses' microphone. The voice signal is then used as input for speech recognition processing to convert it into text.

[1295] Step 2:

[1296] Smart glasses convert audio data into text data. Specifically, they use Google's speech recognition API to convert audio data into text data. The input to this process is audio data, and the output is text data.

[1297] Step 3:

[1298] The converted text data is sent to the server. The server analyzes the received text data using a natural language processing algorithm (Hugging Face's BERT model). In this step, the input is text data, and the output is the specific elements of the analyzed request (e.g., "Item: Shoes", "Request: Shirt", "Purpose: Coordination Support").

[1299] Step 4:

[1300] The server uses a generative AI (GPT-3) to generate example responses based on the elements of the analyzed request. The input used in this process is the specific elements of the analyzed request, and the output is a specific product suggestion. For example, example responses might be "a blue shirt on shelf A" or "a white shirt on shelf B."

[1301] Step 5:

[1302] The generated sample responses are sent to the concierge terminal. The concierge reviews the list of suggestions and makes corrections as needed. Based on inventory information and new product information, they create more accurate suggestions. The input in this step is the generated sample response, and the output is the corrected sample response.

[1303] Step 6:

[1304] The revised answer examples are sent back to the server and compiled into the final answer. This final answer is then sent to the smart glasses worn by the user. For example, the notification might say, "The shirts that go with these shoes are the blue shirt on shelf A and the white shirt on shelf B." In this step, the input is the revised answer examples, and the output is the notification message displayed on the smart glasses.

[1305] In this way, the input data is sequentially processed and analyzed at each processing step, and finally, appropriate product information is provided to the user.

[1306] Furthermore, an emotion engine that estimates the user's emotions may be incorporated. That is, the identification processing unit 290 may use the emotion identification model 59 to estimate the user's emotions and perform identification processing using the user's emotions.

[1307] This invention is a system for efficiently processing user requests in voice or text format and further improving the quality of concierge services by combining it with an emotion engine that recognizes the user's emotions. Specific embodiments of the system are shown below.

[1308] How users can enter their requests

[1309] The process begins with the user entering their request using a device (e.g., a smartphone or PC). Users can enter their request via voice or text. For example, a user might voice-input, "I would like to make a dinner reservation at a high-end restaurant in Tokyo next Saturday night."

[1310] Converting audio data to text

[1311] The device sends the user's voice data to the server. The server uses speech recognition technology to convert the voice data into text data. For example, it might generate the text "I would like to make a dinner reservation at a high-end restaurant in Tokyo next Saturday night" from the voice data.

[1312] Request analysis

[1313] The server analyzes the generated text data using a natural language processing algorithm and extracts specific elements representing the user's request. This yields information such as "Date: Next Saturday," "Time: Evening," "Location: Tokyo," "Genre: High-end restaurant," and "Content: Dinner reservation."

[1314] Using an Emotion Engine

[1315] The server then uses an emotion engine to analyze the user's emotions from voice or text data. For example, it extracts emotional information such as "excited," "anxious," or "hurried." This emotional information influences the generative AI in subsequent processes, which is used to generate more appropriate responses that align with the user's feelings.

[1316] Generating example answers

[1317] The server uses generative AI to generate example responses based on extracted elements and emotional information. For example, the AI ​​might generate a response like, "We suggest the following restaurants: 1. Restaurant A (reservations available) 2. Restaurant B (waitlist)." By considering emotional information, it's possible to include additional perks for users who are excited, for example.

[1318] Concierge confirmation

[1319] The server sends the generated sample response to the concierge's terminal, where the concierge reviews and corrects the content. The concierge then contacts, for example, "Restaurant A" and "Restaurant B" to check the latest reservation status and updates the list if necessary.

[1320] Submit your final response

[1321] The revised response is sent back to the server, which then sends the final response to the user's device. The user can receive the final response via email or app notification. For example, the notification might say, "Here are some high-end restaurants available next Saturday evening. Restaurant A is fully booked, while Restaurant B is on the waiting list."

[1322] Specific example

[1323] The user speaks into their smartphone's microphone and says, "I'd like to make a dinner reservation at a high-end restaurant in Tokyo next Saturday night."

[1324] The device sends the voice data to the server. The server performs speech recognition and obtains text data that reads, "I would like to make a dinner reservation at a high-end restaurant in Tokyo next Saturday night."

[1325] The server analyzes the text data using a natural language processing algorithm and extracts requests such as "next Saturday," "evening," "in Tokyo," "high-end restaurant," and "dinner reservation."

[1326] The server uses an emotion engine to extract emotional information from the user's voice and text data, recognizing emotions such as "excited" or "hurried."

[1327] The server uses generative AI to generate a list based on the user's requests and recognized emotions. For example, "Restaurant A (Reservations Available)" and "Restaurant B (Waiting List)." Depending on the emotion, for instance, if the emotion is excitement, the system will generate a response that includes additional information about special offers.

[1328] This list is sent to the concierge terminal, where the concierge reviews it and makes corrections as needed.

[1329] The revised response is sent back to the server, which then sends the final response to the user's device. The user is notified that "The only upscale restaurants available next Saturday evening are Restaurant A (reservations available) and Restaurant B (waitlist)."

[1330] This system responds quickly and accurately to user requests and provides more personalized service by taking emotions into consideration. This reduces the burden on concierges and improves the quality of service.

[1331] The following describes the processing flow.

[1332] Step 1:

[1333] The user uses their device to input their request in voice or text format. For example, the user might say into their smartphone's microphone, "I'd like to make a dinner reservation at a high-end restaurant in Tokyo next Saturday night."

[1334] Step 2:

[1335] The terminal sends the input voice data to the server. The voice data is transmitted in digital format.

[1336] Step 3:

[1337] The server uses speech recognition technology to convert the received audio data into text data. For example, it might produce text data such as, "I would like to make a dinner reservation at a high-end restaurant in Tokyo next Saturday night."

[1338] Step 4:

[1339] The server uses a natural language processing algorithm to analyze the generated text data. This analysis extracts elements such as "next Saturday," "evening," "in Tokyo," "high-end restaurant," and "dinner reservation."

[1340] Step 5:

[1341] The server uses an emotion engine to analyze the user's emotions from voice or text data. For example, the emotion engine extracts emotional information such as "excited," "anxious," or "hurried."

[1342] Step 6:

[1343] The server uses generative AI to generate example responses based on emotional information and desired elements. By considering emotional information, more personalized responses are generated. For example, the AI ​​might generate a response like, "We suggest the following restaurants: 1. Restaurant A (reservations available) 2. Restaurant B (waitlist)." If the user is looking forward to something, special offers may also be added.

[1344] Step 7:

[1345] The server sends the generated example answer to the concierge terminal. The concierge terminal displays the received content.

[1346] Step 8:

[1347] The concierge reviews the sample responses received and makes corrections as needed. For example, the concierge will actually contact "Restaurant A" and "Restaurant B" to check the latest reservation status and update the list if necessary.

[1348] Step 9:

[1349] The concierge sends the revised answer back to the server. The server verifies the revised content.

[1350] Step 10:

[1351] The server sends the final response to the user's terminal. For example, the user might be notified that "The only upscale restaurants available next Saturday evening are Restaurant A (reservations available) and Restaurant B (waitlist)."

[1352] Through the steps outlined above, this system can efficiently respond to user requests and, by taking user emotions into consideration, improve the quality of concierge services.

[1353] (Example 2)

[1354] Next, we will describe Example 2. In the following description, the data processing device 12 will be referred to as the "server" and the robot 414 as the "terminal".

[1355] Traditional concierge services lacked the technology to appropriately and quickly handle specific user requests, and particularly struggled to respond while considering emotions. Therefore, improving service quality requires analyzing user requests, including their emotions, and generating appropriate responses based on that analysis.

[1356] The specific processing performed by the specific processing unit 290 of the data processing device 12 in Example 2 is realized by the following means. In this invention, the server includes an information terminal for the user to input requests in voice or text format, a conversion means for converting voice data received from the user into text data, an analysis means for analyzing the text data generated by the conversion means using a natural language processing algorithm, a generation means for generating example answers using generative artificial intelligence based on the requests and sentiment information obtained by the analysis means, a transmission means for sending the example answers generated by the generation means to a service provider terminal for confirmation and correction by the service provider, and a transmission means for sending the corrected answers returned from the service provider terminal to the user terminal. This enables the provision of quick and appropriate services that take into account the user's requests and sentiments.

[1357] An "information terminal" is a device used by users to input requests in voice or text format, and includes smartphones and personal computers.

[1358] "Conversion means" refers to a technology or method for converting audio data received from a user into text data, and which uses speech recognition technology.

[1359] "Analysis means" refers to a technology or method that analyzes text data generated by the conversion means using a natural language processing algorithm to extract user requests and emotional information.

[1360] "Generation means" refers to a technology or method that generates example responses using generative artificial intelligence based on requests and emotional information obtained by analysis means.

[1361] "Transmission means" refers to a technology or method for transmitting example answers generated by the generation means to a service provider terminal, and for transmitting the corrected answers to a user terminal.

[1362] A "service provider terminal" refers to a device used to review and correct generated sample answers, specifically the device used by the concierge.

[1363] A "natural language processing algorithm" is an algorithm used to analyze text data and extract information about desires and emotions, and it includes machine learning and data analysis techniques.

[1364] "Generative artificial intelligence" refers to artificial intelligence technology that generates example responses based on analyzed requests and sentiment information, such as using a generative language model.

[1365] "Emotional information" refers to emotional information analyzed from the user's input voice or text data, and includes emotions such as excitement, anxiety, and urgency.

[1366] Modes for carrying out the invention

[1367] This invention is a system for efficiently processing user requests in voice or text format and further improving the quality of concierge services by combining it with an emotion engine that recognizes the user's emotions. Specific embodiments of the system are shown below.

[1368] How users can enter their requests

[1369] The process begins with the user entering their request in voice or text format using an information terminal (such as a smartphone or personal computer). For example, a user could voice-input a request like, "I would like to make a dinner reservation at a high-end restaurant in Tokyo next Saturday night." Users can input voice data using the terminal's microphone or input their request as text data using a keyboard.

[1370] Converting audio data to text

[1371] The device sends the user's voice data to the server. The server uses speech recognition technology to convert the voice data into text data. Specifically, it uses a speech recognition service such as the Google Cloud Speech-to-Text API to convert the voice data into text data. For example, it generates text data such as "I would like to make a dinner reservation at a high-end restaurant in Tokyo next Saturday night" from the voice data.

[1372] Request analysis

[1373] The server analyzes the generated text data using natural language processing algorithms to extract specific elements representing the user's request. Algorithms used include SpaCy and the Google Cloud Natural Language API. This analysis yields information such as "Date: Next Saturday," "Time: Evening," "Location: Tokyo," "Genre: High-end restaurant," and "Content: Dinner reservation."

[1374] Using an Emotion Engine

[1375] The server then uses an emotion engine to analyze the user's emotions from voice or text data. For example, it uses IBM Watson Tone Analyzer to extract emotional information such as "excited," "anxious," and "hurried." This emotional information influences the generative artificial intelligence in subsequent processes, which is used to generate more appropriate responses that align with the user's feelings.

[1376] Generating example answers

[1377] The server generates example responses using generative artificial intelligence (e.g., OpenAI GPT-3) based on extracted elements and emotional information. For example, the generative AI might generate a response such as, "We suggest the following restaurants: 1. Restaurant A (reservations available) 2. Restaurant B (waitlist)." By considering emotional information, it's possible to include additional perks for users who are excited, for example.

[1378] Concierge confirmation

[1379] The server sends the generated sample response to the service provider's terminal, where the service provider reviews and corrects the content. The service provider (e.g., a concierge) then actually contacts restaurants such as "Restaurant A" and "Restaurant B" to check their latest reservation status and updates the list if necessary.

[1380] Submit your final response

[1381] The revised response from the service provider's terminal is sent back to the server, which then sends the final response to the user's information terminal. The user can receive the final response via email or app notification. For example, the notification might say, "We have a list of upscale restaurants available next Saturday evening. Restaurant A is fully booked, while Restaurant B is on the waiting list."

[1382] Examples of prompt statements

[1383] If a user speaks into their smartphone's microphone and says, "I'd like to make a dinner reservation at a high-end restaurant in Tokyo next Saturday night," the following prompt will be used for the generative artificial intelligence.

[1384] "User request: "I want to make a dinner reservation at a high-end restaurant in Tokyo next Saturday night." Sentiment information: "Excited, in a hurry." Please generate a list of responses: Example) - Restaurant A (Reservations available) - Restaurant B (Waiting list) Perks information: Add perks information based on the excited sentiment."

[1385] This allows the system to respond quickly and accurately to user requests and provide more personalized service by taking emotions into consideration. In this way, it is possible to reduce the burden on concierges and improve the quality of service.

[1386] The flow of the specific processing in Example 2 will be explained using Figure 13.

[1387] Step 1: Enter your request

[1388] User

[1389] Users input their requests in voice or text format using an information terminal (e.g., smartphone, personal computer).

[1390] Input: Audio data or text data

[1391] Specific action: For example, the user speaks into their smartphone and says, "I want to make a dinner reservation at a high-end restaurant in Tokyo next Saturday night."

[1392] Step 2: Converting audio data to text

[1393] terminal

[1394] The device sends the audio data to the server.

[1395] Input: Audio data

[1396] Output: Sending audio data to the server

[1397] Specific action: Send audio data to the server in digital format.

[1398] server

[1399] The server uses speech recognition technologies such as the Google Cloud Speech-to-Text API to convert audio data into text data.

[1400] Input: Audio data

[1401] Output: Text data

[1402] Specific operation: Generate the text "I would like to make a dinner reservation at a high-end restaurant in Tokyo next Saturday night" from the audio data.

[1403] Step 3: Analyzing the requirements

[1404] server

[1405] The server analyzes the generated text data using natural language processing algorithms (e.g., SpaCy, Google Cloud Natural Language API) to extract the specific elements of the request.

[1406] Input: Text data

[1407] Output: Elements of the request (e.g., "Date: Next Saturday", "Time: Evening", "Location: Tokyo", "Genre: High-end restaurant", "Content: Dinner reservation")

[1408] Specific operation: Analyze the specific elements of the request and store them in the database.

[1409] Step 4: Utilizing the Emotional Engine

[1410] server

[1411] The server uses an emotion engine (e.g., IBM Watson Tone Analyzer) to analyze emotional information from audio or text data.

[1412] Input: Audio data or text data

[1413] Output: Emotional information (e.g., "excited," "anxious," "hurried")

[1414] Specific operation: Use an emotion engine to analyze emotions from speech or text and store them in a database.

[1415] Step 5: Generating example answers

[1416] server

[1417] The server generates example responses using a generative AI (e.g., OpenAI GPT-3) based on the elements of the request and emotional information.

[1418] Input: Elements of requests, emotional information

[1419] Output: Example answer (Example: "Restaurant A (Reservations accepted)", "Restaurant B (Waiting list)")

[1420] Specific operation: Use generative AI to generate responses that respond to requests and emotions.

[1421] Step 6: Concierge Confirmation

[1422] server

[1423] The server sends the generated example answers to the service provider's terminal.

[1424] Input: Example Answer

[1425] Output: Sent to the service provider's terminal

[1426] Specific action: Send example answers to the service provider's terminal.

[1427] Service provider

[1428] The service provider will review the responses and make corrections as necessary.

[1429] Input: Example Answer

[1430] Output: Revised answer example

[1431] Specific action: The service provider inquires about the latest booking status and corrects the response.

[1432] Step 7: Submit your final response

[1433] Service provider

[1434] The revised answer example is sent back to the server.

[1435] Input: Revised answer example

[1436] Output: Send to server

[1437] Specific action: Send the corrected response to the server.

[1438] server

[1439] The server sends the final response to the user's information terminal.

[1440] Input: Revised answer example

[1441] Output: Sending the final response to the user's information terminal.

[1442] Specific action: Send the final response to the user via email or app notification.

[1443] In this way, by responding quickly and accurately to user requests and taking their emotions into consideration, it becomes possible to provide more personalized services.

[1444] (Application Example 2)

[1445] Next, we will explain application example 2. In the following explanation, the data processing device 12 will be referred to as the "server" and the robot 414 as the "terminal".

[1446] Traditional concierge services faced challenges in providing prompt service due to the extensive work required to accurately understand and incorporate user requests. Furthermore, they lacked consideration for user emotions, resulting in a lack of personalization and low satisfaction. In short, the uniform approach to users and the inability to provide personalized suggestions based on emotions led to a poor user experience.

[1447] In Application Example 2, the specific processing performed by the specific processing unit 290 of the data processing device 12 is realized by the following means. In this invention, the server includes a user terminal for the user to input requests in voice or text format, a conversion means for converting voice data received from the user into text data, an analysis means for analyzing the text data generated by the conversion means as a request using a natural language processing algorithm, a generation means for generating example answers using a generative AI based on the requests and emotional information from the analysis means, a transmission means for sending the example answers generated by the generation means to a concierge terminal for confirmation and correction by the concierge, a transmission means for sending the corrected answers returned from the concierge terminal to the user terminal, and an emotional analysis means for extracting emotional information from the user's voice data and text data. This makes it possible to quickly and accurately analyze user requests and provide personalized services based on emotions.

[1448] A "user terminal" is a device used by users to input requests in voice or text format.

[1449] "Conversion means" refers to technical means for converting audio data received from a user into text data.

[1450] "Analysis means" refers to technical means for analyzing text data generated by the conversion means as a request using a natural language processing algorithm.

[1451] "Generation means" refers to technical means for generating example responses using a generative AI based on requests and emotional information obtained through analysis means.

[1452] "Transmission means" refers to a technical means for sending example answers generated by the generation means to a concierge terminal, for confirmation and correction by the concierge, and for sending the corrected answers returned from the concierge terminal to the user terminal.

[1453] "Emotion analysis methods" refer to technical means for extracting emotional information from user voice data and text data.

[1454] A "natural language processing algorithm" is an algorithm that analyzes text data to understand and extract user requests.

[1455] "Generative AI" is an artificial intelligence technology that generates appropriate response examples based on user requests and emotional information.

[1456] A "concierge terminal" is a device used to review and correct generated sample answers.

[1457] "Emotional information" refers to information about a user's emotions extracted from their voice data and text data.

[1458] This invention is a system for further improving the quality of concierge services by efficiently processing requests entered by users in voice or text format, and by combining this with an emotion engine that recognizes the user's emotions.

[1459] The system includes the following components:

[1460] 1. User terminal: A device used by users to input requests in voice or text format. This includes smartphones, PCs, and smart glasses.

[1461] 2. Conversion method: Speech recognition technology is used to convert the audio data received from the user into text data. Specifically, the SpeechRecognition library is used to convert the audio data into text data.

[1462] 3. Analysis Method: The text data generated by the transformation method is analyzed as a request using a natural language processing algorithm. This uses the natural language processing model from the Transformers library.

[1463] 4. Emotion Analysis Method: An emotion analysis model is used to extract emotional information from the user's voice data and text data. Specifically, the Transformers emotion analysis model is used to obtain emotional information.

[1464] 5. Generation Method: Based on the requests and sentiment information obtained through the analysis method, a generative AI model (such as OpenAI GPT-3) is used to generate example responses. This process uses OpenAI's GPT-3 as an API and takes into account the relationship between the prompt text and the generated text.

[1465] 6. Transmission method: The example answer generated by the generation method is sent to the concierge terminal for confirmation and correction by the concierge. Furthermore, the corrected answer returned from the concierge terminal is sent to the user terminal.

[1466] Explanation of specific processing examples:

[1467] The user inputs their request by voice into their smartphone's microphone, saying, "I would like to make a dinner reservation at a high-end restaurant in Tokyo next Saturday night."

[1468] The user's terminal sends the voice data to the server, which uses the SpeechRecognition library to convert the voice data into text data that reads, "I would like to make a dinner reservation at a high-end restaurant in Tokyo next Saturday night."

[1469] The converted text data is analyzed by Transformers' natural language processing model, and requests such as "Date: Next Saturday", "Time: Evening", "Location: Tokyo", "Genre: High-end restaurant", and "Content: Dinner reservation" are extracted.

[1470] An emotion analysis model analyzes the same text data and obtains emotional information such as "excited" or "hurried."

[1471] Based on the analyzed requests and sentiment information, the server uses GPT-3 to generate example responses with the following prompts:

[1472] User Input: I'm looking for recommended baby products in the store.

[1473] Emotion: Enjoyment

[1474] Response:

[1475] For example, the AI ​​can generate sample responses such as, "In this baby products section, we recommend our new skin-friendly diapers and soft swaddles."

[1476] The generated sample answers are sent to the concierge terminal, where the concierge reviews the content and makes corrections as needed.

[1477] The corrected response is sent back to the server and finally delivered to the user's terminal.

[1478] This system enables the rapid and accurate analysis of user requests and the provision of personalized, emotion-based services. This improves the user experience and reduces the workload on concierges.

[1479] The flow of a specific process in Application Example 2 will be explained using Figure 14.

[1480] Step 1:

[1481] The user performs voice input. The user uses a smartphone or smart glasses to input their request by voice. For example, they might say, "I would like to make a dinner reservation at a high-end restaurant in Tokyo next Saturday night." This voice data is recorded on the device. The input is voice data, and the output is the voice file recorded on the device.

[1482] Step 2:

[1483] The device sends audio data to the server. The device sends the recorded audio data to the server. The server uses speech recognition technology to convert this audio data into text data. The SpeechRecognition library is used for this conversion. The input is audio data, and the output is text data that says, "I would like to make a dinner reservation at a high-end restaurant in Tokyo next Saturday night."

[1484] Step 3:

[1485] The server analyzes the text data. The server then analyzes the transformed text data using a natural language processing algorithm. This process uses the Transformers library to extract user requests (e.g., "Date: Next Saturday", "Time: Evening", "Location: Tokyo", "Genre: High-end restaurant", "Content: Dinner reservation"). The input is text data, and the output is the analyzed elemental information.

[1486] Step 4:

[1487] The server extracts emotional information. The server then uses emotion analysis tools to analyze the user's emotional information from the text data. This process utilizes the Transformers emotion analysis model. For example, emotions such as "excited" and "hurried" are extracted. The input is text data, and the output is emotional information.

[1488] Step 5:

[1489] The server generates example answers using generative AI. Based on the analyzed requests and sentiment information, the server generates example answers using a generative AI model such as GPT-3. This process uses OpenAI's GPT-3 as an API, sending the generated text as a prompt. The input is requests and sentiment information, and the output is the suggested answer. For example, the prompt might look like this:

[1490] User Input: I'm looking for recommended baby products in the store.

[1491] Emotion: Enjoyment

[1492] Response:

[1493] Step 6:

[1494] The server sends example answers to the concierge terminal. The generated example answers are sent to the concierge terminal, where the concierge reviews the content and makes corrections as needed. The input is the suggested example answer, and the output is the corrected example answer.

[1495] Step 7:

[1496] The concierge terminal sends the corrected answer back to the server. After the concierge reviews and makes corrections, the corrected answer is sent back to the server. The input is an example of the corrected answer, and the output is the final example answer.

[1497] Step 8:

[1498] The server sends the final answer to the user's device. The user can receive the final answer via email or app notification. The input is an example of the final answer, and the output is the final answer notified to the user. For example, the user might be notified of the final answer, "The upscale restaurants available next Saturday evening are Restaurant A (reservations available) and Restaurant B (waitlist)."

[1499] The specific processing unit 290 transmits the result of the specific processing to the robot 414. In the robot 414, the control unit 46A causes the speaker 240 and the controlled object 443 to output the result of the specific processing. The microphone 238 acquires audio indicating user input for the result of the specific processing. The control unit 46A transmits the audio data indicating user input acquired by the microphone 238 to the data processing unit 12. In the data processing unit 12, the specific processing unit 290 acquires the audio data.

[1500] Data generation model 58 is a type of so-called generative AI (Artificial Intelligence). One example of data generation model 58 is ChatGPT (Internet search<URL: https: / / openai.com / blog / chatgpt> ), Gemini (Internet search) <url: https: gemini.google.com ?hl="ja">Examples of generative AI include the following. The data generation model 58 is obtained by performing deep learning on a neural network. The data generation model 58 is input with prompts containing instructions, and with inference data such as audio data representing speech, text data representing text, and image data representing images. The data generation model 58 infers from the input inference data according to the instructions indicated by the prompts, and outputs the inference results in data formats such as audio data and text data. Here, inference refers to, for example, analysis, classification, prediction, and / or summarization.

[1501] In the above embodiment, an example was given in which specific processing is performed by the data processing device 12, but the technology of this disclosure is not limited thereto, and the specific processing may also be performed by the robot 414.

[1502] Furthermore, the emotion identification model 59, acting as an emotion engine, may determine the user's emotion according to a specific mapping. Specifically, the emotion identification model 59 may determine the user's emotion according to a specific mapping, which is an emotion map (see Figure 9). Similarly, the emotion identification model 59 may also determine the robot's emotion, and the identification processing unit 290 may perform identification processing using the robot's emotion.

[1503] Figure 9 shows an emotion map 400 in which multiple emotions are mapped. In the emotion map 400, emotions are arranged in concentric circles radiating from the center. The closer to the center of the concentric circles, the more primitive the emotions are located. Further out of the concentric circles, emotions representing states and actions arising from mental states are located. Emotion is a concept that includes feelings and mental states. On the left side of the concentric circles, emotions that are generally generated from reactions occurring in the brain are located. On the right side of the concentric circles, emotions that are generally induced by situational judgment are located. Above and below the concentric circles, emotions that are generally generated from reactions occurring in the brain and induced by situational judgment are located. In addition, the emotion of "pleasure" is located on the upper side of the concentric circles, and the emotion of "displeasure" is located on the lower side. Thus, in the emotion map 400, multiple emotions are mapped based on the structure in which emotions arise, and emotions that are likely to occur simultaneously are mapped close together.

[1504] These emotions are distributed at the 3 o'clock position on the Emotion Map 400, and usually fluctuate between feelings of security and anxiety. In the right half of the Emotion Map 400, situational awareness takes precedence over internal feelings, resulting in a calm impression.

[1505] The inside of the Emotion Map 400 represents inner thoughts, while the outside represents actions. Therefore, the further you go from the outside of the Emotion Map 400, the more visible (expressed in actions) your emotions become.

[1506] Here, human emotions are based on various balances, such as posture and blood sugar levels. When these balances deviate from the ideal, it results in discomfort, and when they approach the ideal, it results in pleasure. Similarly, in robots, cars, motorcycles, etc., emotions can be created based on various balances, such as posture and battery level. When these balances deviate from the ideal, it results in discomfort, and when they approach the ideal, it results in pleasure. The emotion map can be generated, for example, based on Dr. Mitsuyoshi's emotion map (Research on a system for analyzing brain physiological signals of speech emotion recognition and emotion, Tokushima University, doctoral dissertation: https: / / ci.nii.ac.jp / naid / 500000375379). The left half of the emotion map contains emotions belonging to a region called "response," where sensation is dominant. The right half of the emotion map contains emotions belonging to a region called "situation," where situational awareness is dominant.

[1507] The emotion map defines two emotions that promote learning. One is the emotion around the middle of the negative "repentance" and "reflection" on the situation side. In other words, it is when the robot experiences negative emotions such as "I never want to feel this way again" or "I don't want to be scolded again." The other is the emotion around the positive "desire" on the reaction side. In other words, it is when the robot has positive feelings such as "I want more" or "I want to know more."

[1508] The emotion identification model 59 inputs user input into a pre-trained neural network, obtains emotion values ​​representing each emotion shown in the emotion map 400, and determines the user's emotion. This neural network is pre-trained based on multiple training data sets, which are combinations of user input and emotion values ​​representing each emotion shown in the emotion map 400. Furthermore, this neural network is trained so that emotions located close together have similar values, as shown in the emotion map 900 in Figure 10. Figure 10 shows an example where multiple emotions such as "reassured," "calm," and "confident" have similar emotion values.

[1509] The above description primarily focuses on the functions of the data processing device 12 in relation to this disclosure. However, the system related to this disclosure is not necessarily implemented on a server. The system related to this disclosure may be implemented as a general information processing system. This disclosure may be implemented, for example, as a software program that runs on a personal computer or as an application that runs on a smartphone. The method related to this disclosure may be provided to users in SaaS (Software as a Service) format.

[1510] In the above embodiment, an example was given in which a specific process is performed by a single computer 22. However, the technology of this disclosure is not limited thereto, and a distributed processing of the specific process may be performed by multiple computers, including computer 22. For example, a data generation model 58 may be provided in an external device of the data processing device 12, and the external device may generate data according to the input data.

[1511] In the above embodiment, an example was given in which the specific processing program 56 is stored in the storage 32, but the technology of this disclosure is not limited thereto. For example, the specific processing program 56 may be stored in a portable, computer-readable, non-temporary storage medium such as a USB (Universal Serial Bus) memory. The specific processing program 56 stored in the non-temporary storage medium is installed in the computer 22 of the data processing device 12. The processor 28 executes specific processing according to the specific processing program 56.

[1512] Alternatively, the specific processing program 56 may be stored in a storage device such as a server connected to the data processing device 12 via the network 54, and the specific processing program 56 may be downloaded and installed on the computer 22 in response to a request from the data processing device 12.

[1513] Furthermore, it is not necessary to store the entirety of the specific processing program 56 in a storage device such as a server connected to the data processing device 12 via the network 54, or to store the entirety of the specific processing program 56 in the storage 32; it is acceptable to store only a portion of the specific processing program 56.

[1514] The following types of processors can be used as hardware resources to perform specific processing. Examples of processors include a CPU, a general-purpose processor that functions as a hardware resource to perform specific processing by executing software, i.e., a program. Other examples of processors include dedicated electrical circuits, such as FPGAs (Field-Programmable Gate Arrays), PLDs (Programmable Logic Devices), or ASICs (Application Specific Integrated Circuits), which have circuit configurations specifically designed to perform specific processing. All of these processors have built-in or connected memory, and all of them perform specific processing by using memory.

[1515] The hardware resource that performs a specific process may consist of one of these various processors, or it may consist of a combination of two or more processors of the same or different types (for example, a combination of multiple FPGAs, or a combination of a CPU and an FPGA). Alternatively, the hardware resource that performs a specific process may consist of a single processor.

[1516] Examples of configurations using a single processor include, firstly, a configuration in which one or more CPUs and software are combined to form a single processor, and this processor functions as a hardware resource that performs a specific process. Secondly, there is a configuration using a processor that realizes the functions of the entire system, including multiple hardware resources that perform a specific process, on a single IC chip, as exemplified by SoCs (System-on-a-chip). In this way, a specific process is realized using one or more of the above types of processors as hardware resources.

[1517] Furthermore, the hardware structure of these various processors can more specifically utilize electrical circuits that combine circuit elements such as semiconductor devices. Also, the specific processing described above is merely an example. Therefore, it goes without saying that unnecessary steps can be deleted, new steps added, or the processing order rearranged, as long as it does not deviate from the main purpose.

[1518] The descriptions and illustrations presented above are detailed explanations of the technical aspects of this disclosure and are merely examples of the technical aspects. For example, the above descriptions of the structure, function, operation, and effect are examples of the structure, function, operation, and effect of the technical aspects of this disclosure. Therefore, it goes without saying that you may delete unnecessary parts, add new elements, or replace elements in the descriptions and illustrations presented above, as long as you do not deviate from the essence of the technical aspects of this disclosure. Furthermore, in order to avoid confusion and facilitate understanding of the technical aspects of this disclosure, explanations of common technical knowledge and the like that do not require special explanation to enable the implementation of the technical aspects of this disclosure have been omitted from the descriptions and illustrations presented above.

[1519] All documents, patent applications, and technical standards described herein are incorporated by reference to the same extent as if each individual document, patent application, and technical standard were specifically and individually noted as being incorporated by reference.

[1520] The following is further disclosed regarding the embodiments described above.

[1521] (Claim 1)

[1522] A user terminal for users to input requests in voice or text format,

[1523] A conversion means for converting audio data received from a user into text data,

[1524] An analysis means for analyzing text data generated by a conversion means as a request using a natural language processing algorithm,

[1525] A generation means for generating example answers using a generative AI based on the requirements obtained by the analysis means,

[1526] A transmission means for sending the example answer generated by the generation means to a concierge terminal for confirmation and correction by the concierge,

[1527] A means for sending corrected responses returned from the concierge terminal to the user terminal,

[1528] A system that includes this.

[1529] (Claim 2)

[1530] The system according to claim 1, comprising speech recognition technology for recognizing speech data.

[1531] (Claim 3)

[1532] The system according to claim 1, comprising a natural language processing algorithm for extracting user requests.

[1533] "Example 1"

[1534] (Claim 1)

[1535] A user terminal for users to input requests in voice or text format,

[1536] A conversion means for converting audio data received from a user into text data,

[1537] An analysis means for analyzing text data generated by a conversion means as a request using a natural language processing algorithm,

[1538] A generation means for generating example answers using generative artificial intelligence based on the requirements obtained by the analysis means,

[1539] A transmission means for sending the example answer generated by the generation means to the administrator terminal for confirmation and correction by the administrator,

[1540] A means for sending corrected responses returned from the administrator terminal to the user terminal,

[1541] An electronic computing device for automatically generating information that matches the user's requirements,

[1542] A system that includes this.

[1543] (Claim 2)

[1544] The system according to claim 1, comprising speech recognition technology for recognizing speech data.

[1545] (Claim 3)

[1546] The system according to claim 1, comprising a natural language processing algorithm for extracting user requests.

[1547] "Application Example 1"

[1548] (Claim 1)

[1549] A user terminal for users to input requests in voice or text format,

[1550] A conversion means for converting audio data received from a user into text data,

[1551] An analysis means for analyzing text data generated by a conversion means as a request using a natural language processing algorithm,

[1552] A generation means for generating example answers using a generative AI based on the requirements obtained by the analysis means,

[1553] A transmission means for sending the example answer generated by the generation means to a concierge terminal for confirmation and correction by the concierge,

[1554] A means for sending corrected responses returned from the concierge terminal to the user terminal,

[1555] A notification method that allows a user terminal to input a request via smart glasses and to notify the user of the final answer via smart glasses,

[1556] A system that includes this.

[1557] (Claim 2)

[1558] The system according to claim 1, comprising speech recognition technology for recognizing speech data.

[1559] (Claim 3)

[1560] The system according to claim 1, comprising a natural language processing algorithm for extracting user requests.

[1561] "Example 2 of combining an emotion engine"

[1562] (Claim 1)

[1563] An information terminal for users to input requests in voice or text format,

[1564] A conversion means for converting audio data received from a user into text data,

[1565] An analysis means for analyzing text data generated by a conversion means using a natural language processing algorithm,

[1566] A generation means for generating example answers using generative artificial intelligence based on requests and emotional information obtained by an analysis means,

[1567] A transmission means for sending the example answer generated by the generation means to the service provider's terminal and for receiving confirmation and correction from the service provider,

[1568] A means for sending the corrected response returned from the service provider's terminal to the user's terminal,

[1569] A system that includes this.

[1570] (Claim 2)

[1571] The system according to claim 1, comprising speech recognition technology for recognizing speech data.

[1572] (Claim 3)

[1573] The system according to claim 1, comprising a natural language processing algorithm and sentiment analysis technique for extracting user requests and emotional information.

[1574] "Application example 2 when combining with an emotional engine"

[1575] (Claim 1)

[1576] A user terminal for users to input requests in voice or text format,

[1577] A conversion means for converting audio data received from a user into text data,

[1578] An analysis means for analyzing text data generated by a conversion means as a request using a natural language processing algorithm,

[1579] A generation method for generating example answers using a generative AI based on requests and emotional information obtained through analysis,

[1580] A transmission means for sending the example answer generated by the generation means to a concierge terminal for confirmation and correction by the concierge,

[1581] A means for sending corrected responses returned from the concierge terminal to the user terminal,

[1582] A means for analyzing emotions to extract emotional information from user voice data and text data,

[1583] A system that includes this.

[1584] (Claim 2)

[1585] The system according to claim 1, comprising speech recognition technology for recognizing speech data.

[1586] (Claim 3)

[1587] The system according to claim 1, comprising a natural language processing algorithm for extracting user requests. [Explanation of Symbols]

[1588] 10, 210, 310, 410 Data Processing Systems 12 Data Processing Devices 14 Smart Devices 214 Smart Glasses 314 Headset-type terminal 414 Robots< / url:> < / url:> < / url:> < / url:>

Claims

1. A user terminal for users to input requests in voice or text format, A conversion means for converting audio data received from a user into text data, An analysis means for analyzing text data generated by a conversion means as a request using a natural language processing algorithm, A generation means for generating example answers using a generative AI based on the requirements obtained by the analysis means, A transmission means for sending the example answer generated by the generation means to a concierge terminal for confirmation and correction by the concierge, A means for sending corrected responses returned from the concierge terminal to the user terminal, A system that includes this.

2. The system according to claim 1, comprising speech recognition technology for recognizing speech data.

3. The system according to claim 1, comprising a natural language processing algorithm for extracting user requests.

Citation Information

Patent Citations

  • Persona chatbot control method and system

    JP2022180282A