system
A system with a terminal interface and generative machine learning model facilitates detailed product explanations, improving customer satisfaction and operational efficiency by enabling new employees to provide accurate descriptions and customers to access product details easily.
Patent Information
- Authority / Receiving Office
- JP · JP
- Patent Type
- Applications
- Current Assignee / Owner
- Filing Date
- 2024-10-01
- Publication Date
- 2026-04-13
AI Technical Summary
The challenge of providing detailed product explanations to customers and new employees in retail settings due to insufficient product knowledge and lack of efficient information systems.
A system equipped with a terminal interface for product identification, a generative machine learning model to generate detailed product information, and communication means to transmit this information to the terminal, allowing new employees to provide accurate descriptions and customers to access product details easily.
Enhances customer satisfaction and operational efficiency by enabling new employees to provide detailed product information without specialized knowledge and allowing customers to obtain information quickly and accurately.
Smart Images

Figure 2026063793000001_ABST
Abstract
Description
Technical Field
[0001] The technology of the present disclosure relates to a system.
Background Art
[0002] Patent Document 1 discloses a method for controlling a persona chatbot, which is performed by at least one processor and includes steps of receiving a user utterance, adding the user utterance to a prompt including an instruction sentence related to an explanation of a chatbot character, encoding the prompt, and inputting the encoded prompt into a language model to generate a chatbot utterance in response to the user utterance.
Prior Art Documents
Patent Documents
[0003]
Patent Document 1
Summary of the Invention
Problems to be Solved by the Invention
[0004] In recent years, the types and functions of household appliances have diversified, and the descriptions of products in stores have become increasingly complex. Therefore, especially for new employees, there is a problem that it is difficult to give appropriate product explanations to customers due to lack of detailed knowledge about products. In addition, customers themselves may not be able to obtain sufficient information about products, which may make it difficult to make purchase decisions. To solve these problems, a system that can easily explain product details even for new employees is required.
Means for Solving the Problems
[0005] The present invention solves the above problem with a system that includes a terminal means equipped with a user interface for a user to input product identification information, a processing means including a generative machine learning model that receives product identification information from the terminal means and generates detailed product information based on the product identification information, and a communication means that transmits the detailed product information generated by the generative machine learning model to the terminal means. Specifically, when a user inputs product identification information on the terminal, the processing means automatically generates a product description using the generative machine learning model based on that information and displays it on the terminal. As a result, even new crew members can provide product descriptions without needing specialized knowledge, and customers can easily obtain detailed product information themselves.
[0006] "User interface" refers to all screens and input methods that users use to interact with a system.
[0007] "Terminal means" refers to an electronic device equipped with a user interface that has the function of inputting and receiving information.
[0008] "Product identification information" refers to unique information used to identify a specific product, and includes product IDs, barcodes, and other similar information.
[0009] "Processing means" refers to the entire system that performs calculations and data processing based on information received from terminal means.
[0010] A "generative machine learning model" refers to a model that uses machine learning algorithms to generate new information from data.
[0011] "Communication means" refers to all devices and methods that have the function of transmitting data to other systems or devices.
[0012] "Product details" refers to all information regarding the product's characteristics, functions, specifications, etc. [Brief explanation of the drawing]
[0013] [Figure 1] This is a conceptual diagram showing an example of the configuration of a data processing system according to the first embodiment. [Figure 2] This is a conceptual diagram showing an example of the essential functions of a data processing device and a smart device according to the first embodiment. [Figure 3] This is a conceptual diagram showing an example of the configuration of a data processing system according to the second embodiment. [Figure 4] This is a conceptual diagram showing an example of the main functions of a data processing device and smart glasses according to the second embodiment. [Figure 5] This is a conceptual diagram showing an example of the configuration of a data processing system according to the third embodiment. [Figure 6] This is a conceptual diagram showing an example of the main functions of a data processing device and a headset-type terminal according to the third embodiment. [Figure 7] This is a conceptual diagram showing an example of the configuration of a data processing system according to the fourth embodiment. [Figure 8] This is a conceptual diagram showing an example of the main functions of a data processing device and a robot according to the fourth embodiment. [Figure 9] This shows an emotion map where multiple emotions are mapped. [Figure 10] This shows an emotion map where multiple emotions are mapped. [Figure 11] This is a sequence diagram showing the processing flow of the data processing system in Example 1. [Figure 12] This is a sequence diagram showing the processing flow of the data processing system in Application Example 1. [Figure 13] This is a sequence diagram showing the processing flow of the data processing system in Example 2, when an emotion engine is combined. [Figure 14] This is a sequence diagram showing the processing flow of the data processing system in Application Example 2, which combines an emotion engine. [Modes for carrying out the invention]
[0014] An example of an embodiment of the system according to the technology of the present disclosure will be described below with reference to the accompanying drawings.
[0015] First, the terms used in the following description will be explained.
[0016] In the following embodiments, a tagged processor (hereinafter simply referred to as "processor") may be a single arithmetic unit or a combination of multiple arithmetic units. Also, the processor may be a single type of arithmetic unit or a combination of multiple types of arithmetic units. Examples of arithmetic units include a CPU (Central Processing Unit), a GPU (Graphics Processing Unit), a GPGPU (General-Purpose computing on Graphics Processing Units), an APU (Accelerated Processing Unit), and the like.
[0017] In the following embodiments, a tagged RAM (Random Access Memory) is a memory in which information is temporarily stored and is used as a work memory by the processor.
[0018] In the following embodiments, a tagged storage is one or more non-volatile storage devices that store various programs and various parameters, etc. Examples of non-volatile storage devices include flash memory (SSD (Solid State Drive)), magnetic disks (e.g., hard disks), or magnetic tapes, etc.
[0019] In the following embodiments, the signed communication interface (I / F) is an interface that includes a communication processor and an antenna, etc. The communication interface manages communication between multiple computers. Examples of communication standards applicable to the communication interface include wireless communication standards such as 5G (5th Generation Mobile Communication System), Wi-Fi (registered trademark), or Bluetooth (registered trademark).
[0020] In the following embodiments, "A and / or B" is synonymous with "at least one of A and B." That is, "A and / or B" means that it may be A alone, or B alone, or a combination of A and B. Furthermore, in this specification, the same concept as "A and / or B" applies when expressing three or more things linked by "and / or."
[0021] [First Embodiment]
[0022] Figure 1 shows an example of the configuration of the data processing system 10 according to the first embodiment.
[0023] As shown in Figure 1, the data processing system 10 includes a data processing device 12 and a smart device 14. An example of the data processing device 12 is a server.
[0024] The data processing device 12 comprises a computer 22, a database 24, and a communication interface 26. The computer 22 is an example of a "computer" related to the technology of this disclosure. The computer 22 comprises a processor 28, RAM 30, and storage 32. The processor 28, RAM 30, and storage 32 are connected to a bus 34. The database 24 and the communication interface 26 are also connected to the bus 34. The communication interface 26 is connected to a network 54. An example of the network 54 is a WAN (Wide Area Network) and / or a LAN (Local Area Network).
[0025] The smart device 14 comprises a computer 36, a reception device 38, an output device 40, a camera 42, and a communication interface 44. The computer 36 comprises a processor 46, RAM 48, and storage 50. The processor 46, RAM 48, and storage 50 are connected to a bus 52. The reception device 38, output device 40, and camera 42 are also connected to the bus 52.
[0026] The reception device 38 is equipped with a touch panel 38A and a microphone 38B, etc., and receives user input. The touch panel 38A receives user input by detecting contact with an object (e.g., a pen or finger). The microphone 38B receives user input by detecting the user's voice. The control unit 46A transmits data indicating the user input received by the touch panel 38A and microphone 38B to the data processing device 12. In the data processing device 12, the specific processing unit 290 acquires the data indicating the user input.
[0027] The output device 40 includes a display 40A and a speaker 40B, and presents data to the user 20 by outputting the data in a form perceptible to the user 20 (e.g., audio and / or text). The display 40A displays visible information such as text and images according to instructions from the processor 46. The speaker 40B outputs audio according to instructions from the processor 46. The camera 42 is a small digital camera equipped with an optical system such as a lens, aperture, and shutter, and an image sensor such as a CMOS (Complementary Metal-Oxide-Semiconductor) image sensor or a CCD (Charge Coupled Device) image sensor.
[0028] Communication interface 44 is connected to network 54. Communication interfaces 44 and 26 are responsible for the exchange of various types of information between processor 46 and processor 28 via network 54.
[0029] Figure 2 shows an example of the main functions of the data processing device 12 and the smart device 14.
[0030] As shown in Figure 2, in the data processing device 12, a specific processing is performed by the processor 28. A specific processing program 56 is stored in the storage 32. The specific processing program 56 is an example of a "program" related to the technology of this disclosure. The processor 28 reads the specific processing program 56 from the storage 32 and executes the read specific processing program 56 on the RAM 30. The specific processing is realized by the processor 28 operating as a specific processing unit 290 according to the specific processing program 56 executed on the RAM 30.
[0031] The storage 32 stores the data generation model 58 and the emotion identification model 59. The data generation model 58 and the emotion identification model 59 are used by the identification processing unit 290.
[0032] In the smart device 14, the processor 46 performs the reception output processing. The storage 50 stores the reception output program 60. The reception output program 60 is used in conjunction with a specific processing program 56 by the data processing system 10. The processor 46 reads the reception output program 60 from the storage 50 and executes the read reception output program 60 on the RAM 48. The reception output processing is realized by the processor 46 operating as a control unit 46A according to the reception output program 60 executed on the RAM 48.
[0033] Next, the specific processing performed by the specific processing unit 290 of the data processing device 12 will be described. In the following description, the data processing device 12 will be referred to as the "server" and the smart device 14 as the "terminal".
[0034] The embodiments for carrying out the present invention are described in detail below. This system consists of a terminal, a server, a generative machine learning model, and a communication means. The terminal, equipped with a user interface, receives input from the user, and the server is responsible for generating product descriptions.
[0035] First, the user enters product identification information (e.g., product ID) using a device (e.g., tablet, smartphone, digital signage). After entering this information, the user requests product information from the server by pressing the search button.
[0036] The terminal sends an HTTP request to the server based on the product identification information received from the user. The server parses the received request and begins the process of retrieving the corresponding product information from the database. At this stage, the server searches the database for product data that matches the product identification information and passes the retrieved data to a generative machine learning model.
[0037] The generative machine learning model installed on the server generates detailed product descriptions based on the provided product data. This model has the ability to automatically generate user-friendly and detailed product descriptions using natural language processing techniques.
[0038] The generated product description is transmitted from the server to the terminal using a communication method. The terminal displays the received product description on its user interface. This allows the user to view a detailed product description in real time based on the product identification information they entered.
[0039] As a concrete example, consider the case where a user enters the product ID "CAM123". When the user enters the product ID and presses the search button, the terminal sends the product ID to the server. The server retrieves information about the camera corresponding to "CAM123" from the database (for example, "high-resolution 4K camera, 20 megapixels, with zoom function") and passes this information to a generative machine learning model. The model generates a product description such as, "This camera can shoot high-resolution 4K video and take 20-megapixel photos. It also has a zoom function, so you can take clear pictures of distant subjects." This product description is sent to the terminal and displayed on the user's screen.
[0040] This system allows even new crew members to provide detailed and accurate product descriptions without requiring extensive product knowledge. Furthermore, users can easily access product details by entering their own information. Therefore, this invention has the effect of improving customer satisfaction in the sale of home appliances, as well as enhancing the efficiency of crew members' work.
[0041] The following describes the processing flow.
[0042] Step 1:
[0043] The user enters product identification information (e.g., product ID "CAM123") using a device (e.g., tablet or smartphone). After entering this information, the user taps the "Search" button on the screen.
[0044] Step 2:
[0045] The terminal retrieves the product identification information entered by the user. This information is stored in a variable and used for the next process.
[0046] Step 3:
[0047] The device sends an HTTP request to the server based on the acquired product identification information. The request is sent in the format of a GET request, for example, with the product identification information included in the URL.
[0048] Step 4:
[0049] The server parses the request received from the terminal and extracts product identification information from the URL parameters. The extracted product identification information is stored in a variable.
[0050] Step 5:
[0051] The server executes queries against the product database based on the extracted product identification information. The queries are configured to retrieve information about products that match the product identification information.
[0052] Step 6:
[0053] The server retrieves relevant product information from the database (e.g., "High-resolution 4K camera, 20 megapixels, with zoom function"). The retrieved product information is then passed to a generative machine learning model.
[0054] Step 7:
[0055] The generative machine learning model installed on the server generates a detailed product description based on the product information provided. For example, the generated description might be: "This camera can shoot high-resolution 4K video and take 20-megapixel photos. It also has a zoom function, allowing you to take clear photos of distant subjects."
[0056] Step 8:
[0057] The server creates an HTTP response containing the generated product description and sends it to the terminal. The response format is, for example, JSON.
[0058] Step 9:
[0059] The terminal analyzes the response received from the server and extracts the product description text. The extracted text is then placed on a specific element of the user interface (for example, HTML).Display in the tags.
[0060] Step 10:
[0061] Users can view detailed product descriptions displayed on their device screens in real time. This allows users to obtain sufficient information about the product, making it easier for them to make a purchase decision.
[0062] Through the above processing steps, users can easily receive detailed product explanations, and a system is created that allows even new crew members to confidently serve customers.
[0063] (Example 1)
[0064] Next, we will describe Example 1. In the following description, the data processing device 12 will be referred to as the "server," and the smart device 14 will be referred to as the "terminal."
[0065] In the conventional system, obtaining detailed product information required staff with specialized knowledge, making it difficult for new or less knowledgeable crew members to handle such requests. Furthermore, it was difficult for customers to input product information themselves and receive detailed explanations in real time. This led to a decline in the quality of customer service and a risk of decreased customer satisfaction. Additionally, the inability to efficiently provide product information resulted in reduced operational efficiency.
[0066] The identification process performed by the identification processing unit 290 of the data processing device 12 in Example 1 is realized by the following means.
[0067] In this invention, the server includes terminal means equipped with a user interface for users to input product identification information, server means that receives product identification information from the terminal means, analyzes the product identification information, and obtains product information from a database, processing means including a generative machine learning model that generates a detailed product description using the product information obtained by the server means, and communication means that transmits the detailed product information generated by the generative machine learning model to the terminal means. As a result, even new crew members can provide detailed product descriptions without specialized knowledge, and users can input product information themselves and immediately obtain detailed explanations, thereby improving customer satisfaction and operational efficiency.
[0068] A "terminal device" is an electronic device equipped with a user interface for users to input product identification information.
[0069] A "server means" is a device or system that analyzes product identification information received from a terminal means and retrieves the corresponding product information from a database.
[0070] A "generative machine learning model" is a processing method for generating detailed product descriptions based on acquired product information.
[0071] "Communication means" refers to means for transmitting detailed information about the generated product to a terminal device.
[0072] "Product identification information" refers to information used to individually identify a product, such as a product ID or barcode.
[0073] A "database" is a system or location for storing product information.
[0074] A "user interface" refers to an interface that allows users to input information and view output results.
[0075] "Product information" refers to detailed information about a product, including its features and specifications.
[0076] The embodiments for carrying out the present invention are described in detail below. This system consists of terminal means, server means, a generative machine learning model, and communication means.
[0077] The terminal device is equipped with a user interface for the user to input product identification information. Specifically, electronic display devices (signage), personal digital assistants (tablets), and smart devices (smartphones) are used. The user uses these devices to input product identification information (e.g., product ID).
[0078] Once the user completes the input, the terminal sends the product identification information to the server as an HTTP request. The server receives this HTTP request and begins parsing. Specifically, it parses the product identification information and retrieves the corresponding product information from the database.
[0079] The database stores detailed product information, and the server retrieves the necessary product information using SQL queries and other methods. The retrieved product information is then passed to a generative machine learning model. This model uses natural language processing techniques to generate detailed product descriptions. Specifically, it has the ability to automatically generate user-friendly and detailed product descriptions based on the product information.
[0080] The generated product description is transmitted from the server to the terminal using a communication method. HTTP communication is used as the communication method. Finally, the terminal displays the received product details on the user interface. This allows the user to view detailed product descriptions in real time based on the product identification information they entered.
[0081] As a concrete example, consider the case where a user enters the product ID "CAM123". When the user enters the product ID "CAM123" and presses the search button, the terminal sends this information to the server. The server retrieves information about the camera corresponding to "CAM123" from its database (for example, "high-resolution 4K camera, 20 megapixels, with zoom function") and passes this information to a generative machine learning model. Based on this data, the generative machine learning model automatically generates a product description such as, "This camera can shoot high-resolution 4K video and take 20-megapixel photos. It also has a zoom function, so you can take clear pictures of distant subjects." This product description is finally sent to the terminal and displayed on the user's screen.
[0082] As an example of a prompt, the text "Generate a detailed product description for product ID 'CAM123'" is input to the generative machine learning model. A specific example of input to the model would be "High-resolution 4K camera, 20 megapixels, with zoom function."
[0083] This system allows users to check detailed product information in real time by entering their own product identification information. Furthermore, even new crew members can provide quick and accurate product descriptions without needing specialized knowledge. This is expected to improve customer satisfaction and operational efficiency in the sale of home appliances.
[0084] The flow of the specific processing in Example 1 will be explained using Figure 11.
[0085] Step 1:
[0086] The user enters product identification information.
[0087] Specific explanation: The user enters product identification information (e.g., product ID) into an input field on the terminal's user interface and presses the search button. The entered product identification information is, for example, "CAM123".
[0088] Input: Product identification information (product ID)
[0089] Output: Input product identification information
[0090] Step 2:
[0091] The device sends product identification information to the server.
[0092] Detailed explanation: The terminal sends an HTTP request to the server containing the entered product identification information. Internally, the terminal generates a request in the format "GET / search?product_id=CAM123".
[0093] Input: Entered product identification information (product ID)
[0094] Output: HTTP request to the server
[0095] Step 3:
[0096] The server parses the request.
[0097] Specific explanation: The server parses the received HTTP request and extracts product identification information from the query parameters. For example, it extracts the product ID "CAM123" from the request "GET / search?product_id=CAM123".
[0098] Input: HTTP Request
[0099] Output: Analyzed product identification information (product ID)
[0100] Step 4:
[0101] The server searches the database.
[0102] Specific explanation: The server executes an SQL query against the database based on the analyzed product identification information. It retrieves the relevant product information using the query "SELECT FROM products WHERE product_id = 'CAM123'".
[0103] Input: Analyzed product identification information (product ID)
[0104] Output: Acquired product information (e.g., "High-resolution 4K camera, 20 megapixels, with zoom function")
[0105] Step 5:
[0106] The server passes data to the generative machine learning model.
[0107] Detailed explanation: The server passes the acquired product information to a generative machine learning model. The generative machine learning model generates a detailed product description based on the given product information. The model incorporates natural language processing techniques.
[0108] Input: Acquired product information
[0109] Output: Generated product description (Example: "This camera can shoot high-resolution 4K video and take 20-megapixel photos. It also has a zoom function, allowing you to capture distant subjects clearly.")
[0110] Step 6:
[0111] The server sends the generated product description to the terminal.
[0112] Specific explanation: The server generates an HTTP response containing the product description generated by the generative machine learning model and sends it to the terminal. The response includes the description along with "200 OK".
[0113] Input: Generated product description
[0114] Output: HTTP response
[0115] Step 7:
[0116] The device displays the product description.
[0117] Specific explanation: The device parses the received HTTP response and displays the product description on the user interface. The user can view the detailed product description displayed on the device screen in real time.
[0118] Input: HTTP response
[0119] Output: Product description displayed on the user interface
[0120] (Application Example 1)
[0121] Next, we will explain Application Example 1. In the following explanation, the data processing device 12 will be referred to as the "server," and the smart device 14 will be referred to as the "terminal."
[0122] Traditional methods of product description in physical stores rely on the skill level and consistency of staff explanations, making it difficult to provide customers with appropriate and detailed information quickly. Furthermore, even when systems existed that automatically generated product descriptions, they lacked features such as voice assistants or scanning capabilities, failing to adequately consider user convenience. As a result, challenges remain in improving customer satisfaction and reducing staff workload.
[0123] The specific processing performed by the specific processing unit 290 of the data processing device 12 in Application Example 1 is realized by the following means.
[0124] In this invention, the server includes terminal means equipped with a user interface for users to input product identification information; processing means including a generative machine learning model that receives product identification information from the terminal means and generates detailed product information based on the product identification information; communication means for transmitting the detailed product information generated by the generative machine learning model to the terminal means; voice output means equipped with a voice assistant function for providing product descriptions by voice; and scanning means for customers to obtain product identification information by scanning a QR code (registered trademark) or barcode and display it on the terminal. This makes it possible to provide customers with real-time, consistent, and detailed product descriptions, regardless of the skill level of the staff. In addition, the voice assistant function and scanning function allow customers to easily obtain product information themselves, simultaneously improving customer satisfaction and reducing the burden on staff.
[0125] A "terminal device" refers to a device equipped with an interface for users to input product identification information. Examples include digital signage, tablets, smartphones, and smart glasses.
[0126] A "generative machine learning model" refers to artificial intelligence technology that automatically generates detailed product descriptions based on input product identification information.
[0127] "Communication means" refers to the technology and methods for transmitting detailed product information generated by a generative machine learning model to a terminal device.
[0128] "Voice output means" refers to a function that provides generated product descriptions in audio format. This allows users to obtain product information through audio.
[0129] "Scanning method" refers to technology that allows customers to obtain product identification information by scanning a QR code or barcode and display it on their device.
[0130] "User interface" is a general term for the display and input methods used by users to input product identification information and to review the generated product description.
[0131] The system of this invention consists of terminal means, a server, a generative machine learning model, communication means, voice output means, and scanning means. Specific embodiments of the system using these components will be described in detail below.
[0132] Hardware and software
[0133] Terminal devices: Digital signage, tablets, smartphones, smart glasses, and other devices.
[0134] Server: A cloud server using Amazon Web Services (AWS®) or Google® Cloud Platform (GCP).
[0135] Database: Use MySQL (registered trademark) or PostgreSQL.
[0136] Generative machine learning models: Generative machine learning models including OpenAI® GPT-4® or Google BERT.
[0137] Audio output method: The speaker or earphones built into the device.
[0138] Scanning method: Camera on the device or a dedicated barcode scanner.
[0139] Program Overview
[0140] The server receives product identification information from a terminal device equipped with a user interface for users to input product identification information. For example, the user can obtain product information by entering a product ID on a tablet's input screen or by scanning a QR code or barcode with a smartphone's camera. This allows the product identification information to be acquired through the scanning device.
[0141] The received product identification information is transmitted to the server via a communication method. The server searches the database based on this product identification information and retrieves the corresponding product data. The retrieved product data is passed to a generative machine learning model, which uses natural language processing techniques to generate a detailed product description.
[0142] The generated product description is then transmitted back to the terminal via a communication device. The user can view the detailed product description in real time on the terminal screen. Furthermore, it is possible to listen to the generated product description aloud using an audio output device. This allows users to obtain product information both visually and aurally, improving convenience.
[0143] Specific example
[0144] For example, consider a case where a user enters the product ID "CAM123". When the user enters "CAM123" into a tablet and presses the search button, the device sends the product ID to the server. The server retrieves information about the camera corresponding to "CAM123" from its database (for example, "High-resolution 4K camera, 20 megapixels, with zoom function") and passes this information to a generative machine learning model. The model generates a product description like the following:
[0145] Product ID: CAM123
[0146] This camera can shoot high-resolution 4K video and take 20-megapixel photos. It also features a zoom function, allowing you to capture distant subjects clearly.
[0147] The generated product description is displayed on the device screen and also provided to the user via audio output. In this way, the user can obtain product information through both sight and sound.
[0148] The flow of a specific process in Application Example 1 will be explained using Figure 12.
[0149] Step 1: Enter product identification information
[0150] The user inputs product identification information (e.g., product ID, QR code, barcode) using a terminal device (e.g., smartphone or tablet). At this time, the user interface saves the entered product identification information to the terminal device's memory.
[0151] Input: Product ID, QR code, or barcode
[0152] Output: Product identification information
[0153] Step 2: Submit product identification information
[0154] The terminal device transmits the product identification information obtained in step 1 to the server using the communication device. Specifically, an HTTP request is generated and sent to the server's endpoint.
[0155] Input: Product identification information
[0156] Output: HTTP request sent to the server
[0157] Step 3: Search the database
[0158] The server queries the database based on the received product identification information to retrieve the corresponding product data. It executes an SQL query to fetch the relevant product data.
[0159] Input: Product identification information
[0160] Output: Product data
[0161] Step 4: Explanation generation using generative machine learning models
[0162] The server inputs the product data obtained in step 3 into a generative machine learning model to generate detailed product descriptions. The model uses natural language processing techniques to convert the product data into easy-to-understand text.
[0163] Input: Product data
[0164] Output: Product Description
[0165] Step 5: Send the generated product description.
[0166] The server sends the generated product description to the terminal device via a communication means. An HTTP response is generated and sent to the endpoint of the terminal device.
[0167] Input: Product Description
[0168] Output: HTTP response sent to the terminal device
[0169] Step 6: Display and audio output of product description
[0170] The terminal device displays the received product description on the user interface and also provides the product description in audio using the audio output device.
[0171] Input: Product Description
[0172] Output: Screen display and audio output
[0173] As a concrete example, consider the case where a user enters the product ID "CAM123". When the user enters the product ID and presses the search button, the terminal sends the product ID to the server. The server retrieves information about the camera corresponding to "CAM123" from its database (for example, "High-resolution 4K camera, 20 megapixels, with zoom function") and passes this information to a generative machine learning model. The model generates a product description like the following:
[0174] Product ID: CAM123
[0175] This camera can shoot high-resolution 4K video and take 20-megapixel photos. It also features a zoom function, allowing you to capture distant subjects clearly.
[0176] The generated product description is displayed on the device screen and also provided to the user via audio output. In this way, the user can obtain product information through both sight and sound.
[0177] Furthermore, an emotion engine that estimates the user's emotions may be incorporated. That is, the identification processing unit 290 may use the emotion identification model 59 to estimate the user's emotions and perform identification processing using the user's emotions.
[0178] The embodiments for carrying out the present invention are described in detail below. This system consists of a terminal, a server, a generative machine learning model, an emotion engine, and a communication means. The terminal, equipped with a user interface, receives input from the user, and the server is responsible for generating product descriptions.
[0179] First, the user enters product identification information (e.g., product ID) using a device (e.g., tablet, smartphone, digital signage). After entering this information, the user requests product information from the server by pressing the search button.
[0180] The device acquires emotional data from product identification information received from the user and from an emotion engine that recognizes emotions from the user's facial expressions and voice. The emotion engine uses facial recognition technology and voice recognition technology to evaluate the user's emotions in real time.
[0181] The device sends product identification information and sentiment data to the server. The server analyzes the received request, accesses the database using the product identification information, and retrieves the corresponding product information. This information, along with the sentiment data, is then passed to a generative machine learning model.
[0182] Generative machine learning models can generate product descriptions adapted to the user's emotional state based on the provided product information and sentiment data. For example, if the user shows an interested expression, it will provide a more detailed and technical explanation. On the other hand, if the user shows an anxious expression, it will generate a simple and reassuring explanation.
[0183] The generated product description is transmitted from the server to the terminal using a communication method. The terminal displays the received product description on its user interface. This allows the user to view a detailed product description in real time based on the product identification information they entered.
[0184] As a concrete example, consider the case where a user enters the product ID "CAM123". When the user enters the product ID and presses the search button, the terminal sends the product ID along with emotional data acquired from the user's facial expressions and voice to the server. The server retrieves information about the camera corresponding to "CAM123" from the database (for example, "high-resolution 4K camera, 20 megapixels, with zoom function") and passes this information and emotional data to a generative machine learning model. The model generates an appropriate product description according to the user's emotional state, adjusting content such as, "This camera can shoot high-resolution 4K video and take 20-megapixel photos. It also has a zoom function, so you can take clear pictures of distant subjects." For example, if the user seems interested, detailed technical explanations will be added, and if they are feeling anxious, more approachable terminology will be used.
[0185] This system allows even new crew members, without requiring extensive product knowledge, to provide detailed and accurate product descriptions tailored to the user's emotional state. Furthermore, users can easily learn about product details in an emotionally sensitive way by inputting product information themselves. Therefore, this invention has the effect of improving customer satisfaction in the sale of home appliances and enhancing the efficiency of crew members' work.
[0186] The following describes the processing flow.
[0187] Step 1:
[0188] The user enters product identification information (e.g., product ID "CAM123") using a device (e.g., tablet or smartphone). After entering this information, the user taps the "Search" button on the screen.
[0189] Step 2:
[0190] The terminal retrieves the product identification information entered by the user. This information is stored in a variable and used for the next process.
[0191] Step 3:
[0192] The device analyzes the user's facial expressions and voice using an emotion engine to obtain the user's emotional data (e.g., "interest," "anxiety," "confusion," etc.). This data is also stored in variables.
[0193] Step 4:
[0194] The device sends an HTTP request to the server based on the acquired product identification information and sentiment data. The request is sent in the format of a POST request, for example, containing the product identification information and sentiment data.
[0195] Step 5:
[0196] The server analyzes the request received from the terminal and extracts product identification information and sentiment data. The extracted data is then stored in variables.
[0197] Step 6:
[0198] The server uses product identification information to query the product database. The query is configured to retrieve information about products that match the product identification information.
[0199] Step 7:
[0200] The server retrieves relevant product information from the database (for example, "high-resolution 4K camera, 20 megapixels, with zoom function"). The retrieved product information and sentiment data are then passed to a generative machine learning model.
[0201] Step 8:
[0202] Generative machine learning models generate detailed product descriptions based on the product information and sentiment data they receive. For example, the model adjusts a description such as "This camera can shoot high-resolution 4K video and take 20-megapixel photos. It also has a zoom function, allowing you to capture distant subjects clearly" according to the user's emotional state.
[0203] Step 9:
[0204] The server creates an HTTP response containing the generated product description and sends it to the terminal. The response format is, for example, JSON.
[0205] Step 10:
[0206] The terminal analyzes the response received from the server and extracts the product description text. The extracted text is then placed on a specific element of the user interface (for example, HTML). Display in the tags.
[0207] Step 11:
[0208] Users can view detailed product descriptions displayed on their device screens in real time. This allows users to obtain sufficient information about the product, making it easier for them to make a purchase decision.
[0209] These processing steps enable users to easily receive detailed product descriptions, creating a system that allows even new crew members to confidently interact with customers. Furthermore, by utilizing emotional data, product descriptions optimized for the user's emotional state are provided.
[0210] (Example 2)
[0211] Next, we will describe Example 2. In the following description, the data processing device 12 will be referred to as the "server" and the smart device 14 as the "terminal".
[0212] In recent years, providing users with appropriate product information has become increasingly important in e-commerce and brick-and-mortar retail. However, many systems provide uniform information without considering user emotions, resulting in insufficient user experience and customer satisfaction. Furthermore, a shortage of staff and sales personnel who understand and can provide detailed technical information about products is also a problem. Therefore, there is a need for a system that provides effective product descriptions in real time, tailored to the user's emotional state.
[0213] The identification processing performed by the identification processing unit 290 of the data processing device 12 in Example 2 is realized by the following means. In this invention, the server includes terminal means equipped with a user interface for the user to input product identification information and emotion data; processing means including a generative machine learning model that receives product identification information and emotion data from the terminal means and generates product information based on the product identification information and detailed product information adapted to the user's emotional state; and communication means that transmits the detailed product information generated by the generative machine learning model to the terminal means. This makes it possible to provide detailed and appropriate product descriptions in real time that correspond to the user's emotions.
[0214] "Terminal means" refers to devices used by users to input product identification information and emotional data, and includes tablets, smartphones, and digital signage.
[0215] "Product identification information" refers to information used to identify a product, and includes, for example, product IDs and barcodes.
[0216] "Emotional data" refers to data that indicates the emotional state of a user, analyzed from their facial expressions and voice.
[0217] "User interface" is a general term for the screens and input devices that users use to operate a device, and includes touch panels and button interfaces.
[0218] A "generative machine learning model" is a model that uses machine learning techniques to generate appropriate product descriptions based on product information and sentiment data.
[0219] "Processing means" refers to a device or system that processes information received from a terminal means, and includes a generative machine learning model.
[0220] "Communication means" refers to technologies and devices for sending and receiving data between a server and a terminal, and includes the internet and local networks.
[0221] A "database" is a system or device that manages and stores data such as product information.
[0222] An "emotion engine" is a system or device that analyzes a user's facial expressions and voice to generate emotional data.
[0223] This invention is a system for improving the user experience when acquiring product information. This system consists of a terminal, an emotion engine, a server, a generative machine learning model, and communication means. The following describes specific implementations of this system.
[0224] System Configuration
[0225] The terminal device is a device equipped with a user interface, used by users to input product identification information and emotional data. Specifically, this includes tablets, smartphones, and digital signage. The terminal device uses a built-in camera and microphone to analyze the user's facial expressions and voice, and generates emotional data using an emotion engine.
[0226] The emotion engine uses facial recognition and speech recognition technologies to analyze emotional data from the user's facial expressions and voice in real time. This allows for an accurate understanding of the user's current emotional state.
[0227] The server receives product identification information and sentiment data transmitted from the terminal device. Furthermore, the server accesses a predefined database to retrieve the relevant product information. The retrieved product information and sentiment data are then passed to a generative machine learning model.
[0228] Generative machine learning models generate detailed product information tailored to the user's emotional state, based on product information and sentiment data. For example, if the user shows an interested expression, it provides a more detailed and technical explanation. On the other hand, if the user shows an anxious expression, it generates a simple and reassuring explanation.
[0229] The communication means plays the role of transmitting detailed information about the generated product from the server to the terminal means.
[0230] Specific example
[0231] For example, consider a case where a user enters the product ID "CAM123". When the user enters the product ID and presses the search button, the terminal sends the product ID along with emotional data acquired from the user's facial expressions and voice to the server. The server retrieves product information for the camera corresponding to "CAM123" from the database (for example, "High-resolution 4K camera, 20 megapixels, with zoom function") and passes this information and emotional data to a generative machine learning model. The model generates an appropriate product description according to the user's emotional state, adjusting content such as, "This camera can shoot high-resolution 4K video and take 20-megapixel photos. It also has a zoom function, so you can take clear pictures of distant subjects." For example, if the user seems interested, detailed technical explanations will be added, and if they are anxious, more approachable terminology will be used.
[0232] Example of a prompt
[0233] Create a description of the high-resolution 4K camera that matches the user's expression of interest.
[0234] Please create a description of the 20-megapixel camera that is tailored to the user's anxious expression.
[0235] This system allows even new crew members to provide detailed and accurate product descriptions tailored to the user's emotional state, even without extensive product knowledge. Furthermore, users can easily learn about product details in an emotionally sensitive way by inputting product information themselves. Therefore, it improves customer satisfaction in home appliance sales and enhances the efficiency of crew members' work.
[0236] The flow of the specific processing in Example 2 will be explained using Figure 13.
[0237] Step 1:
[0238] The user enters product identification information (e.g., product ID "CAM123") using the terminal. When the user presses the search button, this product ID is saved as input data on the terminal. The terminal's user interface receives this product identification information, preparing the data to be used in the next step.
[0239] Step 2:
[0240] The device uses its built-in camera and microphone to capture the user's facial expressions and voice. The device sends this data to an emotion engine, which analyzes the user's emotional data in real time. Specifically, the device captures the user's face with its camera and acquires emotional data such as "interested facial expressions" and "cheerful voice tones."
[0241] Step 3:
[0242] The terminal sends product identification information (e.g., "CAM123") and acquired emotion data to the server. By receiving product identification information and emotion data as input data, the server analyzes this data and prepares it for use in the next step.
[0243] Step 4:
[0244] The server accesses the database based on the product identification information and retrieves the corresponding product information. Specifically, the server executes a database query to retrieve product information for cameras corresponding to "CAM123" (e.g., "High-resolution 4K camera, 20 megapixels, with zoom function"). The input data is the product identification information, and the output data is the product information.
[0245] Step 5:
[0246] The server passes the acquired product information and sentiment data to a generative machine learning model. The generative machine learning model receives the product information and sentiment data as input and generates detailed product information adapted to the user's emotional state. Specifically, the generative machine learning model considers that "the user appears interested" and generates a product description that includes detailed technical explanations. The output data is a sentiment-based, adjusted product description.
[0247] Step 6:
[0248] The generated product description is transmitted from the server to the terminal via a communication method. Specifically, the server generates the product description and sends it to the terminal as a data packet over the internet. The input data is the generated product description, and the output data is the result of the transmission to the terminal.
[0249] Step 7:
[0250] The device displays the received product description on its user interface. Users can view detailed product descriptions in real time on the device's screen. Specifically, the device's touchscreen displays a text message such as, "This camera can shoot high-resolution 4K video and 20-megapixel photos. It also has a zoom function, allowing you to take clear photos of distant subjects." The output data is the displayed product description.
[0251] (Application Example 2)
[0252] Next, we will explain application example 2. In the following explanation, the data processing device 12 will be referred to as a "server" and the smart device 14 as a "terminal".
[0253] When customers choose products in physical stores, they may feel anxious because they lack detailed information about the products. Furthermore, the uniform nature of product descriptions makes it difficult to provide personalized information tailored to individual customers' emotions and needs. This situation can lead to decreased customer satisfaction and reduced purchasing intent. To address this, there is a need for a system that analyzes customer emotions from their facial expressions and voice, and generates and displays product descriptions adapted to those emotions in real time.
[0254] The specific processing performed by the specific processing unit 290 of the data processing device 12 in Application Example 2 is realized by the following means.
[0255] In this invention, the server includes emotion recognition means for analyzing the user's emotions from their facial expressions and voice, processing means including a generative machine learning model that receives product identification information and emotion data and generates a detailed product description, and communication means for transmitting the generated product details to a terminal means. This makes it possible to provide personalized product descriptions that are adapted to the customer's emotions.
[0256] A "terminal device" is a device equipped with an interface for the user to input product identification information.
[0257] "Processing means" refers to a device that includes a generative machine learning model that generates detailed product information based on product identification information and sentiment data.
[0258] "Communication means" refers to a device that transmits detailed product information generated by a generative machine learning model to a terminal means.
[0259] "Emotion recognition means" refers to technologies and devices used to analyze emotions from a customer's facial expressions and voice.
[0260] "Product identification information" refers to information that a user enters in order to identify a specific product.
[0261] A "generative machine learning model" is a machine learning algorithm that generates detailed product information based on product identification information and sentiment data.
[0262] A "database" is a collection of information where product information is stored in a predefined format.
[0263] A "user interface" is a means of interaction that allows a user to input information into a system.
[0264] "Emotional data" refers to numerical and categorical data of emotional states analyzed from a user's facial expressions and voice.
[0265] "Product description" refers to a detailed description of a product generated by a generative machine learning model.
[0266] The embodiments for carrying out the present invention are described in detail below. This system consists of a terminal, a server, a generative machine learning model, an emotion recognition means, and a communication means. The terminal, equipped with a user interface, receives input from the user, and the server is responsible for generating product descriptions.
[0267] The process begins when the user enters product identification information using a terminal device (e.g., smartphone, tablet, or digital signage). Once the user enters the product identification information and presses the search button, the terminal requests this information from the server. At the same time, the user's facial expressions and voice are also captured, and this data is analyzed in real time by emotion recognition technology.
[0268] The emotion recognition means acquires user emotion data using facial recognition technology and speech recognition technology. This emotion data is numerical data representing the user's emotional state, such as interest, doubt, and anxiety. The terminal means transmits this emotion data to the server along with product identification information.
[0269] The server parses the received request and retrieves the corresponding product information from the database based on the product identification information. It then passes the product information, along with sentiment data, to a generative machine learning model. The generative machine learning model uses the product information and sentiment data to generate a product description adapted to the user's emotional state. For example, if the user shows an interested expression, it provides a more detailed and technical explanation; if they show an anxious expression, it generates a simple and friendly explanation.
[0270] The generated product description is transmitted from the server to the terminal using a communication method and displayed on the terminal's user interface. This allows the user to view a detailed product description based on the entered product identification information in real time.
[0271] The hardware required to implement this system includes terminal devices such as smartphones and tablets, as well as servers equipped with GPUs. The software used includes OpenCV (for image processing), TENSORFLOW (registered trademark) (for loading and using emotion recognition models), Transformers (for loading and using GPT-3 (registered trademark) models), and requests (for API communication).
[0272] Specific example:
[0273] For example, if a user enters the product ID "CAM123" and looks interested, the generated description will be: "This camera can shoot high-resolution 4K video and take 20-megapixel photos. It also has a zoom function, allowing you to take clear photos of distant subjects." In this case, an example of the prompt message would be as follows:
[0274] Example of a prompt sentence:
[0275] Product: 4K Video Camera
[0276] Description: A high-resolution 4K camera that can capture 20-megapixel photos and features zoom functionality.
[0277] Emotion: Curious
[0278] With this format, it is possible for a generative machine learning model to provide a detailed product description adapted to the user's emotional state.
[0279] The flow of the specific process in Application Example 2 will be described using FIG. 14.
[0280] Step 1:
[0281] The terminal acquires product identification information and emotion data
[0282] The user inputs product identification information using terminal means (such as a smartphone, tablet, signage, etc.). Also, the user's expression and voice are acquired through the camera and microphone of the terminal. Thereby, product identification information and emotion data are obtained as inputs. This emotion data is analyzed in real time using image processing tools (e.g., OpenCV) and voice analysis tools.
[0283] Step 2:
[0284] Analysis and integration of emotion data
[0285] The terminal passes the acquired facial expression data and voice data to the emotion recognition means (emotion recognition model using TensorFlow) to convert the user's emotional state into numerical data. For example, when the user has an interested expression, a category value of "interesting" is obtained as the emotion data. Thereby, emotion data analyzed from the input data (facial expression data, voice data) is generated.
[0286] Step 3:
[0287] Transmission of product identification information and emotion data to the server
[0288] The terminal transmits the product identification information and the analyzed emotion data to the server. Thereby, the product identification information and the emotion data are passed to the server as inputs transmitted from the terminal.
[0289] Step 4:
[0290] Acquisition of product information
[0291] The server obtains the corresponding product information from a predefined database based on the received product identification information. For example, product information (high-resolution 4K camera, 20 megapixels, with zoom function, etc.) for product ID "CAM123" is obtained. Thereby, product information is obtained from the input data (product identification information).
[0292] Step 5:
[0293] Generation of product description by generative machine learning model
[0294] The server passes the acquired product information and emotion data to a generative machine learning model (for example, a generative AI model using TensorFlow) to create a prompt sentence. Based on this prompt sentence, the generative machine learning model generates a product description adapted to the user's emotional state. For example, it generates a description such as "This camera can shoot high-resolution 4K videos and take 20-megapixel photos. It also has a zoom function and can clearly shoot distant subjects."
[0295] Step 6:
[0296] Sending and displaying product descriptions to the device
[0297] Product descriptions generated by generative machine learning models are transmitted from the server to the terminal via communication. The terminal displays the received product descriptions on its user interface. This ensures that the output is presented in a format that the user can verify, based on the input (generated product descriptions).
[0298] Through these steps, users can receive detailed product descriptions tailored to their emotional state in real time.
[0299] The specific processing unit 290 transmits the result of the specific processing to the smart device 14. In the smart device 14, the control unit 46A causes the output device 40 to output the result of the specific processing. The microphone 38B acquires audio indicating user input for the result of the specific processing. The control unit 46A transmits the audio data indicating user input acquired by the microphone 38B to the data processing device 12. In the data processing device 12, the specific processing unit 290 acquires the audio data.
[0300] Data generation model 58 is a so-called generative AI (Artificial Intelligence). An example of data generation model 58 is ChatGPT (registered trademark) (Internet search).<URL: https: / / openai.com / blog / chatgpt> ), Gemini (registered trademark) (Internet search) <url: https: gemini.google.com ?hl="ja">Examples of generative AI include the following. The data generation model 58 is obtained by performing deep learning on a neural network. The data generation model 58 is input with prompts containing instructions, and with inference data such as audio data representing speech, text data representing text, and image data representing images. The data generation model 58 infers from the input inference data according to the instructions indicated by the prompts, and outputs the inference results in data formats such as audio data and text data. Here, inference refers to, for example, analysis, classification, prediction, and / or summarization.
[0301] In the above embodiment, an example was given in which specific processing is performed by the data processing device 12, but the technology of this disclosure is not limited thereto, and the specific processing may also be performed by the smart device 14.
[0302] [Second Embodiment]
[0303] Figure 3 shows an example of the configuration of the data processing system 210 according to the second embodiment.
[0304] As shown in Figure 3, the data processing system 210 includes a data processing device 12 and smart glasses 214. An example of the data processing device 12 is a server.
[0305] The data processing device 12 comprises a computer 22, a database 24, and a communication interface 26. The computer 22 is an example of a "computer" related to the technology of this disclosure. The computer 22 comprises a processor 28, RAM 30, and storage 32. The processor 28, RAM 30, and storage 32 are connected to a bus 34. The database 24 and the communication interface 26 are also connected to the bus 34. The communication interface 26 is connected to a network 54. An example of the network 54 is a WAN (Wide Area Network) and / or a LAN (Local Area Network).
[0306] The smart glasses 214 include a computer 36, a microphone 238, a speaker 240, a camera 42, and a communication I / F 44. The computer 36 includes a processor 46, a RAM 48, and a storage 50. The processor 46, the RAM 48, and the storage 50 are connected to a bus 52. Also, the microphone 238, the speaker 240, and the camera 42 are connected to the bus 52.
[0307] The microphone 238 receives instructions and the like from the user 20 by receiving the voice uttered by the user 20. The microphone 238 captures the voice uttered by the user 20, converts the captured voice into voice data, and outputs the voice data to the processor 46. The speaker 240 outputs voice according to instructions from the processor 46.
[0308] The camera 42 is a small digital camera equipped with an optical system such as a lens, an aperture, and a shutter, and an imaging device such as a CMOS (Complementary Metal-Oxide-Semiconductor) image sensor or a CCD (Charge Coupled Device) image sensor, and images the surroundings of the user 20 (for example, an imaging range defined by an angle of view corresponding to the field of view of a generally healthy person).
[0309] The communication I / F 44 is connected to a network 54. The communication I / Fs 44 and 26 control the exchange of various information between the processor 46 and the processor 28 via the network 54. The exchange of various information between the processor 46 and the processor 28 using the communication I / Fs 44 and 26 is performed in a secure state.
[0310] FIG. 4 shows an example of the main functions of the data processing device 12 and the smart glasses 214. As shown in FIG. 4, in the data processing device 12, specific processing is performed by the processor 28. A specific processing program 56 is stored in the storage 32.
[0311] The specific processing program 56 is an example of a "program" relating to the technology of this disclosure. The processor 28 reads the specific processing program 56 from the storage 32 and executes the read specific processing program 56 on the RAM 30. The specific processing is realized by the processor 28 operating as a specific processing unit 290 in accordance with the specific processing program 56 executed on the RAM 30.
[0312] The storage 32 stores the data generation model 58 and the emotion identification model 59. The data generation model 58 and the emotion identification model 59 are used by the identification processing unit 290.
[0313] In the smart glasses 214, the processor 46 performs the reception output processing. The storage 50 stores the reception output program 60. The processor 46 reads the reception output program 60 from the storage 50 and executes the read reception output program 60 on the RAM 48. The reception output processing is realized by the processor 46 operating as a control unit 46A according to the reception output program 60 executed on the RAM 48.
[0314] Next, the identification processing performed by the identification processing unit 290 of the data processing device 12 will be described. In the following description, the data processing device 12 will be referred to as the "server" and the smart glasses 214 will be referred to as the "terminal".
[0315] The embodiments for carrying out the present invention are described in detail below. This system consists of a terminal, a server, a generative machine learning model, and a communication means. The terminal, equipped with a user interface, receives input from the user, and the server is responsible for generating product descriptions.
[0316] First, the user enters product identification information (e.g., product ID) using a device (e.g., tablet, smartphone, digital signage). After entering this information, the user requests product information from the server by pressing the search button.
[0317] The terminal sends an HTTP request to the server based on the product identification information received from the user. The server parses the received request and begins the process of retrieving the corresponding product information from the database. At this stage, the server searches the database for product data that matches the product identification information and passes the retrieved data to a generative machine learning model.
[0318] The generative machine learning model installed on the server generates detailed product descriptions based on the provided product data. This model has the ability to automatically generate user-friendly and detailed product descriptions using natural language processing techniques.
[0319] The generated product description is transmitted from the server to the terminal using a communication method. The terminal displays the received product description on its user interface. This allows the user to view a detailed product description in real time based on the product identification information they entered.
[0320] As a concrete example, consider the case where a user enters the product ID "CAM123". When the user enters the product ID and presses the search button, the terminal sends the product ID to the server. The server retrieves information about the camera corresponding to "CAM123" from the database (for example, "high-resolution 4K camera, 20 megapixels, with zoom function") and passes this information to a generative machine learning model. The model generates a product description such as, "This camera can shoot high-resolution 4K video and take 20-megapixel photos. It also has a zoom function, so you can take clear pictures of distant subjects." This product description is sent to the terminal and displayed on the user's screen.
[0321] This system allows even new crew members to provide detailed and accurate product descriptions without requiring extensive product knowledge. Furthermore, users can easily access product details by entering their own information. Therefore, this invention has the effect of improving customer satisfaction in the sale of home appliances, as well as enhancing the efficiency of crew members' work.
[0322] The following describes the processing flow.
[0323] Step 1:
[0324] The user enters product identification information (e.g., product ID "CAM123") using a device (e.g., tablet or smartphone). After entering this information, the user taps the "Search" button on the screen.
[0325] Step 2:
[0326] The terminal retrieves the product identification information entered by the user. This information is stored in a variable and used for the next process.
[0327] Step 3:
[0328] The device sends an HTTP request to the server based on the acquired product identification information. The request is sent in the format of a GET request, for example, with the product identification information included in the URL.
[0329] Step 4:
[0330] The server parses the request received from the terminal and extracts product identification information from the URL parameters. The extracted product identification information is stored in a variable.
[0331] Step 5:
[0332] The server executes queries against the product database based on the extracted product identification information. The queries are configured to retrieve information about products that match the product identification information.
[0333] Step 6:
[0334] The server retrieves relevant product information from the database (e.g., "High-resolution 4K camera, 20 megapixels, with zoom function"). The retrieved product information is then passed to a generative machine learning model.
[0335] Step 7:
[0336] The generative machine learning model installed on the server generates a detailed product description based on the product information provided. For example, the generated description might be: "This camera can shoot high-resolution 4K video and take 20-megapixel photos. It also has a zoom function, allowing you to take clear photos of distant subjects."
[0337] Step 8:
[0338] The server creates an HTTP response containing the generated product description and sends it to the terminal. The response format is, for example, JSON.
[0339] Step 9:
[0340] The terminal analyzes the response received from the server and extracts the product description text. The extracted text is then placed on a specific element of the user interface (for example, HTML). Display in the tags.
[0341] Step 10:
[0342] Users can view detailed product descriptions displayed on their device screens in real time. This allows users to obtain sufficient information about the product, making it easier for them to make a purchase decision.
[0343] Through the above processing steps, users can easily receive detailed product explanations, and a system is created that allows even new crew members to confidently serve customers.
[0344] (Example 1)
[0345] Next, we will describe Example 1. In the following description, the data processing device 12 will be referred to as the "server," and the smart glasses 214 will be referred to as the "terminal."
[0346] In the conventional system, obtaining detailed product information required staff with specialized knowledge, making it difficult for new or less knowledgeable crew members to handle such requests. Furthermore, it was difficult for customers to input product information themselves and receive detailed explanations in real time. This led to a decline in the quality of customer service and a risk of decreased customer satisfaction. Additionally, the inability to efficiently provide product information resulted in reduced operational efficiency.
[0347] The identification process performed by the identification processing unit 290 of the data processing device 12 in Example 1 is realized by the following means.
[0348] In this invention, the server includes terminal means equipped with a user interface for users to input product identification information, server means that receives product identification information from the terminal means, analyzes the product identification information, and obtains product information from a database, processing means including a generative machine learning model that generates a detailed product description using the product information obtained by the server means, and communication means that transmits the detailed product information generated by the generative machine learning model to the terminal means. As a result, even new crew members can provide detailed product descriptions without specialized knowledge, and users can input product information themselves and immediately obtain detailed explanations, thereby improving customer satisfaction and operational efficiency.
[0349] A "terminal device" is an electronic device equipped with a user interface for users to input product identification information.
[0350] A "server means" is a device or system that analyzes product identification information received from a terminal means and retrieves the corresponding product information from a database.
[0351] A "generative machine learning model" is a processing method for generating detailed product descriptions based on acquired product information.
[0352] "Communication means" refers to means for transmitting detailed information about the generated product to a terminal device.
[0353] "Product identification information" refers to information used to individually identify a product, such as a product ID or barcode.
[0354] A "database" is a system or location for storing product information.
[0355] A "user interface" refers to an interface that allows users to input information and view output results.
[0356] "Product information" refers to detailed information about a product, including its features and specifications.
[0357] The embodiments for carrying out the present invention are described in detail below. This system consists of terminal means, server means, a generative machine learning model, and communication means.
[0358] The terminal device is equipped with a user interface for the user to input product identification information. Specifically, electronic display devices (signage), personal digital assistants (tablets), and smart devices (smartphones) are used. The user uses these devices to input product identification information (e.g., product ID).
[0359] Once the user completes the input, the terminal sends the product identification information to the server as an HTTP request. The server receives this HTTP request and begins parsing. Specifically, it parses the product identification information and retrieves the corresponding product information from the database.
[0360] The database stores detailed product information, and the server retrieves the necessary product information using SQL queries and other methods. The retrieved product information is then passed to a generative machine learning model. This model uses natural language processing techniques to generate detailed product descriptions. Specifically, it has the ability to automatically generate user-friendly and detailed product descriptions based on the product information.
[0361] The generated product description is transmitted from the server to the terminal using a communication method. HTTP communication is used as the communication method. Finally, the terminal displays the received product details on the user interface. This allows the user to view detailed product descriptions in real time based on the product identification information they entered.
[0362] As a concrete example, consider the case where a user enters the product ID "CAM123". When the user enters the product ID "CAM123" and presses the search button, the terminal sends this information to the server. The server retrieves information about the camera corresponding to "CAM123" from its database (for example, "high-resolution 4K camera, 20 megapixels, with zoom function") and passes this information to a generative machine learning model. Based on this data, the generative machine learning model automatically generates a product description such as, "This camera can shoot high-resolution 4K video and take 20-megapixel photos. It also has a zoom function, so you can take clear pictures of distant subjects." This product description is finally sent to the terminal and displayed on the user's screen.
[0363] As an example of a prompt, the text "Generate a detailed product description for product ID 'CAM123'" is input to the generative machine learning model. A specific example of input to the model would be "High-resolution 4K camera, 20 megapixels, with zoom function."
[0364] This system allows users to check detailed product information in real time by entering their own product identification information. Furthermore, even new crew members can provide quick and accurate product descriptions without needing specialized knowledge. This is expected to improve customer satisfaction and operational efficiency in the sale of home appliances.
[0365] The flow of the specific processing in Example 1 will be explained using Figure 11.
[0366] Step 1:
[0367] The user enters product identification information.
[0368] Specific explanation: The user enters product identification information (e.g., product ID) into an input field on the terminal's user interface and presses the search button. The entered product identification information is, for example, "CAM123".
[0369] Input: Product identification information (product ID)
[0370] Output: Input product identification information
[0371] Step 2:
[0372] The device sends product identification information to the server.
[0373] Detailed explanation: The terminal sends an HTTP request to the server containing the entered product identification information. Internally, the terminal generates a request in the format "GET / search?product_id=CAM123".
[0374] Input: Entered product identification information (product ID)
[0375] Output: HTTP request to the server
[0376] Step 3:
[0377] The server parses the request.
[0378] Specific explanation: The server parses the received HTTP request and extracts product identification information from the query parameters. For example, it extracts the product ID "CAM123" from the request "GET / search?product_id=CAM123".
[0379] Input: HTTP Request
[0380] Output: Analyzed product identification information (product ID)
[0381] Step 4:
[0382] The server searches the database.
[0383] Specific explanation: The server executes an SQL query against the database based on the analyzed product identification information. It retrieves the relevant product information using the query "SELECT FROM products WHERE product_id = 'CAM123'".
[0384] Input: Analyzed product identification information (product ID)
[0385] Output: Acquired product information (e.g., "High-resolution 4K camera, 20 megapixels, with zoom function")
[0386] Step 5:
[0387] The server passes data to the generative machine learning model.
[0388] Detailed explanation: The server passes the acquired product information to a generative machine learning model. The generative machine learning model generates a detailed product description based on the given product information. The model incorporates natural language processing techniques.
[0389] Input: Acquired product information
[0390] Output: Generated product description (Example: "This camera can shoot high-resolution 4K video and take 20-megapixel photos. It also has a zoom function, allowing you to capture distant subjects clearly.")
[0391] Step 6:
[0392] The server sends the generated product description to the terminal.
[0393] Specific explanation: The server generates an HTTP response containing the product description generated by the generative machine learning model and sends it to the terminal. The response includes the description along with "200 OK".
[0394] Input: Generated product description
[0395] Output: HTTP response
[0396] Step 7:
[0397] The device displays the product description.
[0398] Specific explanation: The device parses the received HTTP response and displays the product description on the user interface. The user can view the detailed product description displayed on the device screen in real time.
[0399] Input: HTTP response
[0400] Output: Product description displayed on the user interface
[0401] (Application Example 1)
[0402] Next, we will explain Application Example 1. In the following explanation, the data processing device 12 will be referred to as the "server," and the smart glasses 214 will be referred to as the "terminal."
[0403] Traditional methods of product description in physical stores rely on the skill level and consistency of staff explanations, making it difficult to provide customers with appropriate and detailed information quickly. Furthermore, even when systems existed that automatically generated product descriptions, they lacked features such as voice assistants or scanning capabilities, failing to adequately consider user convenience. As a result, challenges remain in improving customer satisfaction and reducing staff workload.
[0404] The specific processing performed by the specific processing unit 290 of the data processing device 12 in Application Example 1 is realized by the following means.
[0405] In this invention, the server includes terminal means equipped with a user interface for users to input product identification information; processing means including a generative machine learning model that receives product identification information from the terminal means and generates detailed product information based on the product identification information; communication means for transmitting the detailed product information generated by the generative machine learning model to the terminal means; voice output means equipped with a voice assistant function for providing product descriptions by voice; and scanning means for obtaining product identification information by having customers scan a QR code or barcode and displaying it on the terminal. This makes it possible to provide customers with real-time, consistent, and detailed product descriptions, regardless of the skill level of the staff. In addition, the voice assistant function and scanning function allow customers to easily obtain product information themselves, simultaneously improving customer satisfaction and reducing the burden on staff.
[0406] A "terminal device" refers to a device equipped with an interface for users to input product identification information. Examples include digital signage, tablets, smartphones, and smart glasses.
[0407] A "generative machine learning model" refers to artificial intelligence technology that automatically generates detailed product descriptions based on input product identification information.
[0408] "Communication means" refers to the technology and methods for transmitting detailed product information generated by a generative machine learning model to a terminal device.
[0409] "Voice output means" refers to a function that provides generated product descriptions in audio format. This allows users to obtain product information through audio.
[0410] "Scanning method" refers to technology that allows customers to obtain product identification information by scanning a QR code or barcode and display it on their device.
[0411] "User interface" is a general term for the display and input methods used by users to input product identification information and to review the generated product description.
[0412] The system of this invention consists of terminal means, a server, a generative machine learning model, communication means, voice output means, and scanning means. Specific embodiments of the system using these components will be described in detail below.
[0413] Hardware and software
[0414] Terminal devices: Digital signage, tablets, smartphones, smart glasses, and other devices.
[0415] Server: A cloud server using Amazon Web Services (AWS) or Google Cloud Platform (GCP).
[0416] Use either MySQL or PostgreSQL as the database.
[0417] Generative machine learning models: Generative machine learning models including OpenAI GPT-4 or Google BERT.
[0418] Audio output method: The speaker or earphones built into the device.
[0419] Scanning method: Camera on the device or a dedicated barcode scanner.
[0420] Program Overview
[0421] The server receives product identification information from a terminal device equipped with a user interface for users to input product identification information. For example, the user can obtain product information by entering a product ID on a tablet's input screen or by scanning a QR code or barcode with a smartphone's camera. This allows the product identification information to be acquired through the scanning device.
[0422] The received product identification information is transmitted to the server via a communication method. The server searches the database based on this product identification information and retrieves the corresponding product data. The retrieved product data is passed to a generative machine learning model, which uses natural language processing techniques to generate a detailed product description.
[0423] The generated product description is then transmitted back to the terminal via a communication device. The user can view the detailed product description in real time on the terminal screen. Furthermore, it is possible to listen to the generated product description aloud using an audio output device. This allows users to obtain product information both visually and aurally, improving convenience.
[0424] Specific example
[0425] For example, consider a case where a user enters the product ID "CAM123". When the user enters "CAM123" into a tablet and presses the search button, the device sends the product ID to the server. The server retrieves information about the camera corresponding to "CAM123" from its database (for example, "High-resolution 4K camera, 20 megapixels, with zoom function") and passes this information to a generative machine learning model. The model generates a product description like the following:
[0426] Product ID: CAM123
[0427] This camera can shoot high-resolution 4K video and take 20-megapixel photos. It also features a zoom function, allowing you to capture distant subjects clearly.
[0428] The generated product description is displayed on the device screen and also provided to the user via audio output. In this way, the user can obtain product information through both sight and sound.
[0429] The flow of a specific process in Application Example 1 will be explained using Figure 12.
[0430] Step 1: Enter product identification information
[0431] The user inputs product identification information (e.g., product ID, QR code, barcode) using a terminal device (e.g., smartphone or tablet). At this time, the user interface saves the entered product identification information to the terminal device's memory.
[0432] Input: Product ID, QR code, or barcode
[0433] Output: Product identification information
[0434] Step 2: Submit product identification information
[0435] The terminal device transmits the product identification information obtained in step 1 to the server using the communication device. Specifically, an HTTP request is generated and sent to the server's endpoint.
[0436] Input: Product identification information
[0437] Output: HTTP request sent to the server
[0438] Step 3: Search the database
[0439] The server queries the database based on the received product identification information to retrieve the corresponding product data. It executes an SQL query to fetch the relevant product data.
[0440] Input: Product identification information
[0441] Output: Product data
[0442] Step 4: Explanation generation using generative machine learning models
[0443] The server inputs the product data obtained in step 3 into a generative machine learning model to generate detailed product descriptions. The model uses natural language processing techniques to convert the product data into easy-to-understand text.
[0444] Input: Product data
[0445] Output: Product Description
[0446] Step 5: Send the generated product description.
[0447] The server sends the generated product description to the terminal device via a communication means. An HTTP response is generated and sent to the endpoint of the terminal device.
[0448] Input: Product Description
[0449] Output: HTTP response sent to the terminal device
[0450] Step 6: Display and audio output of product description
[0451] The terminal device displays the received product description on the user interface and also provides the product description in audio using the audio output device.
[0452] Input: Product Description
[0453] Output: Screen display and audio output
[0454] As a concrete example, consider the case where a user enters the product ID "CAM123". When the user enters the product ID and presses the search button, the terminal sends the product ID to the server. The server retrieves information about the camera corresponding to "CAM123" from its database (for example, "High-resolution 4K camera, 20 megapixels, with zoom function") and passes this information to a generative machine learning model. The model generates a product description like the following:
[0455] Product ID: CAM123
[0456] This camera can shoot high-resolution 4K video and take 20-megapixel photos. It also features a zoom function, allowing you to capture distant subjects clearly.
[0457] The generated product description is displayed on the device screen and also provided to the user via audio output. In this way, the user can obtain product information through both sight and sound.
[0458] Furthermore, an emotion engine that estimates the user's emotions may be incorporated. That is, the identification processing unit 290 may use the emotion identification model 59 to estimate the user's emotions and perform identification processing using the user's emotions.
[0459] The embodiments for carrying out the present invention are described in detail below. This system consists of a terminal, a server, a generative machine learning model, an emotion engine, and a communication means. The terminal, equipped with a user interface, receives input from the user, and the server is responsible for generating product descriptions.
[0460] First, the user enters product identification information (e.g., product ID) using a device (e.g., tablet, smartphone, digital signage). After entering this information, the user requests product information from the server by pressing the search button.
[0461] The device acquires emotional data from product identification information received from the user and from an emotion engine that recognizes emotions from the user's facial expressions and voice. The emotion engine uses facial recognition technology and voice recognition technology to evaluate the user's emotions in real time.
[0462] The device sends product identification information and sentiment data to the server. The server analyzes the received request, accesses the database using the product identification information, and retrieves the corresponding product information. This information, along with the sentiment data, is then passed to a generative machine learning model.
[0463] Generative machine learning models can generate product descriptions adapted to the user's emotional state based on the provided product information and sentiment data. For example, if the user shows an interested expression, it will provide a more detailed and technical explanation. On the other hand, if the user shows an anxious expression, it will generate a simple and reassuring explanation.
[0464] The generated product description is transmitted from the server to the terminal using a communication method. The terminal displays the received product description on its user interface. This allows the user to view a detailed product description in real time based on the product identification information they entered.
[0465] As a concrete example, consider the case where a user enters the product ID "CAM123". When the user enters the product ID and presses the search button, the terminal sends the product ID along with emotional data acquired from the user's facial expressions and voice to the server. The server retrieves information about the camera corresponding to "CAM123" from the database (for example, "high-resolution 4K camera, 20 megapixels, with zoom function") and passes this information and emotional data to a generative machine learning model. The model generates an appropriate product description according to the user's emotional state, adjusting content such as, "This camera can shoot high-resolution 4K video and take 20-megapixel photos. It also has a zoom function, so you can take clear pictures of distant subjects." For example, if the user seems interested, detailed technical explanations will be added, and if they are feeling anxious, more approachable terminology will be used.
[0466] This system allows even new crew members, without requiring extensive product knowledge, to provide detailed and accurate product descriptions tailored to the user's emotional state. Furthermore, users can easily learn about product details in an emotionally sensitive way by inputting product information themselves. Therefore, this invention has the effect of improving customer satisfaction in the sale of home appliances and enhancing the efficiency of crew members' work.
[0467] The following describes the processing flow.
[0468] Step 1:
[0469] The user enters product identification information (e.g., product ID "CAM123") using a device (e.g., tablet or smartphone). After entering this information, the user taps the "Search" button on the screen.
[0470] Step 2:
[0471] The terminal retrieves the product identification information entered by the user. This information is stored in a variable and used for the next process.
[0472] Step 3:
[0473] The device analyzes the user's facial expressions and voice using an emotion engine to obtain the user's emotional data (e.g., "interest," "anxiety," "confusion," etc.). This data is also stored in variables.
[0474] Step 4:
[0475] The device sends an HTTP request to the server based on the acquired product identification information and sentiment data. The request is sent in the format of a POST request, for example, containing the product identification information and sentiment data.
[0476] Step 5:
[0477] The server analyzes the request received from the terminal and extracts product identification information and sentiment data. The extracted data is then stored in variables.
[0478] Step 6:
[0479] The server uses product identification information to query the product database. The query is configured to retrieve information about products that match the product identification information.
[0480] Step 7:
[0481] The server retrieves relevant product information from the database (for example, "high-resolution 4K camera, 20 megapixels, with zoom function"). The retrieved product information and sentiment data are then passed to a generative machine learning model.
[0482] Step 8:
[0483] Generative machine learning models generate detailed product descriptions based on the product information and sentiment data they receive. For example, the model adjusts a description such as "This camera can shoot high-resolution 4K video and take 20-megapixel photos. It also has a zoom function, allowing you to capture distant subjects clearly" according to the user's emotional state.
[0484] Step 9:
[0485] The server creates an HTTP response containing the generated product description and sends it to the terminal. The response format is, for example, JSON.
[0486] Step 10:
[0487] The terminal analyzes the response received from the server and extracts the product description text. The extracted text is then placed on a specific element of the user interface (for example, HTML). Display in the tags.
[0488] Step 11:
[0489] Users can view detailed product descriptions displayed on their device screens in real time. This allows users to obtain sufficient information about the product, making it easier for them to make a purchase decision.
[0490] These processing steps enable users to easily receive detailed product descriptions, creating a system that allows even new crew members to confidently interact with customers. Furthermore, by utilizing emotional data, product descriptions optimized for the user's emotional state are provided.
[0491] (Example 2)
[0492] Next, we will describe Example 2. In the following description, the data processing device 12 will be referred to as the "server" and the smart glasses 214 will be referred to as the "terminal".
[0493] In recent years, providing users with appropriate product information has become increasingly important in e-commerce and brick-and-mortar retail. However, many systems provide uniform information without considering user emotions, resulting in insufficient user experience and customer satisfaction. Furthermore, a shortage of staff and sales personnel who understand and can provide detailed technical information about products is also a problem. Therefore, there is a need for a system that provides effective product descriptions in real time, tailored to the user's emotional state.
[0494] The identification processing performed by the identification processing unit 290 of the data processing device 12 in Example 2 is realized by the following means. In this invention, the server includes terminal means equipped with a user interface for the user to input product identification information and emotion data; processing means including a generative machine learning model that receives product identification information and emotion data from the terminal means and generates product information based on the product identification information and detailed product information adapted to the user's emotional state; and communication means that transmits the detailed product information generated by the generative machine learning model to the terminal means. This makes it possible to provide detailed and appropriate product descriptions in real time that correspond to the user's emotions.
[0495] "Terminal means" refers to devices used by users to input product identification information and emotional data, and includes tablets, smartphones, and digital signage.
[0496] "Product identification information" refers to information used to identify a product, and includes, for example, product IDs and barcodes.
[0497] "Emotional data" refers to data that indicates the emotional state of a user, analyzed from their facial expressions and voice.
[0498] "User interface" is a general term for the screens and input devices that users use to operate a device, and includes touch panels and button interfaces.
[0499] A "generative machine learning model" is a model that uses machine learning techniques to generate appropriate product descriptions based on product information and sentiment data.
[0500] "Processing means" refers to a device or system that processes information received from a terminal means, and includes a generative machine learning model.
[0501] "Communication means" refers to technologies and devices for sending and receiving data between a server and a terminal, and includes the internet and local networks.
[0502] A "database" is a system or device that manages and stores data such as product information.
[0503] An "emotion engine" is a system or device that analyzes a user's facial expressions and voice to generate emotional data.
[0504] This invention is a system for improving the user experience when acquiring product information. This system consists of a terminal, an emotion engine, a server, a generative machine learning model, and communication means. The following describes specific implementations of this system.
[0505] System Configuration
[0506] The terminal device is a device equipped with a user interface, used by users to input product identification information and emotional data. Specifically, this includes tablets, smartphones, and digital signage. The terminal device uses a built-in camera and microphone to analyze the user's facial expressions and voice, and generates emotional data using an emotion engine.
[0507] The emotion engine uses facial recognition and speech recognition technologies to analyze emotional data from the user's facial expressions and voice in real time. This allows for an accurate understanding of the user's current emotional state.
[0508] The server receives product identification information and sentiment data transmitted from the terminal device. Furthermore, the server accesses a predefined database to retrieve the relevant product information. The retrieved product information and sentiment data are then passed to a generative machine learning model.
[0509] Generative machine learning models generate detailed product information tailored to the user's emotional state, based on product information and sentiment data. For example, if the user shows an interested expression, it provides a more detailed and technical explanation. On the other hand, if the user shows an anxious expression, it generates a simple and reassuring explanation.
[0510] The communication means plays the role of transmitting detailed information about the generated product from the server to the terminal means.
[0511] Specific example
[0512] For example, consider a case where a user enters the product ID "CAM123". When the user enters the product ID and presses the search button, the terminal sends the product ID along with emotional data acquired from the user's facial expressions and voice to the server. The server retrieves product information for the camera corresponding to "CAM123" from the database (for example, "High-resolution 4K camera, 20 megapixels, with zoom function") and passes this information and emotional data to a generative machine learning model. The model generates an appropriate product description according to the user's emotional state, adjusting content such as, "This camera can shoot high-resolution 4K video and take 20-megapixel photos. It also has a zoom function, so you can take clear pictures of distant subjects." For example, if the user seems interested, detailed technical explanations will be added, and if they are anxious, more approachable terminology will be used.
[0513] Example of a prompt
[0514] Create a description of the high-resolution 4K camera that matches the user's expression of interest.
[0515] Please create a description of the 20-megapixel camera that is tailored to the user's anxious expression.
[0516] This system allows even new crew members to provide detailed and accurate product descriptions tailored to the user's emotional state, even without extensive product knowledge. Furthermore, users can easily learn about product details in an emotionally sensitive way by inputting product information themselves. Therefore, it improves customer satisfaction in home appliance sales and enhances the efficiency of crew members' work.
[0517] The flow of the specific processing in Example 2 will be explained using Figure 13.
[0518] Step 1:
[0519] The user enters product identification information (e.g., product ID "CAM123") using the terminal. When the user presses the search button, this product ID is saved as input data on the terminal. The terminal's user interface receives this product identification information, preparing the data to be used in the next step.
[0520] Step 2:
[0521] The device uses its built-in camera and microphone to capture the user's facial expressions and voice. The device sends this data to an emotion engine, which analyzes the user's emotional data in real time. Specifically, the device captures the user's face with its camera and acquires emotional data such as "interested facial expressions" and "cheerful voice tones."
[0522] Step 3:
[0523] The terminal sends product identification information (e.g., "CAM123") and acquired emotion data to the server. By receiving product identification information and emotion data as input data, the server analyzes this data and prepares it for use in the next step.
[0524] Step 4:
[0525] The server accesses the database based on the product identification information and retrieves the corresponding product information. Specifically, the server executes a database query to retrieve product information for cameras corresponding to "CAM123" (e.g., "High-resolution 4K camera, 20 megapixels, with zoom function"). The input data is the product identification information, and the output data is the product information.
[0526] Step 5:
[0527] The server passes the acquired product information and sentiment data to a generative machine learning model. The generative machine learning model receives the product information and sentiment data as input and generates detailed product information adapted to the user's emotional state. Specifically, the generative machine learning model considers that "the user appears interested" and generates a product description that includes detailed technical explanations. The output data is a sentiment-based, adjusted product description.
[0528] Step 6:
[0529] The generated product description is transmitted from the server to the terminal via a communication method. Specifically, the server generates the product description and sends it to the terminal as a data packet over the internet. The input data is the generated product description, and the output data is the result of the transmission to the terminal.
[0530] Step 7:
[0531] The device displays the received product description on its user interface. Users can view detailed product descriptions in real time on the device's screen. Specifically, the device's touchscreen displays a text message such as, "This camera can shoot high-resolution 4K video and 20-megapixel photos. It also has a zoom function, allowing you to take clear photos of distant subjects." The output data is the displayed product description.
[0532] (Application Example 2)
[0533] Next, we will explain application example 2. In the following explanation, the data processing device 12 will be referred to as the "server," and the smart glasses 214 will be referred to as the "terminal."
[0534] When customers choose products in physical stores, they may feel anxious because they lack detailed information about the products. Furthermore, the uniform nature of product descriptions makes it difficult to provide personalized information tailored to individual customers' emotions and needs. This situation can lead to decreased customer satisfaction and reduced purchasing intent. To address this, there is a need for a system that analyzes customer emotions from their facial expressions and voice, and generates and displays product descriptions adapted to those emotions in real time.
[0535] The specific processing performed by the specific processing unit 290 of the data processing device 12 in Application Example 2 is realized by the following means.
[0536] In this invention, the server includes emotion recognition means for analyzing the user's emotions from their facial expressions and voice, processing means including a generative machine learning model that receives product identification information and emotion data and generates a detailed product description, and communication means for transmitting the generated product details to a terminal means. This makes it possible to provide personalized product descriptions that are adapted to the customer's emotions.
[0537] A "terminal device" is a device equipped with an interface for the user to input product identification information.
[0538] "Processing means" refers to a device that includes a generative machine learning model that generates detailed product information based on product identification information and sentiment data.
[0539] "Communication means" refers to a device that transmits detailed product information generated by a generative machine learning model to a terminal means.
[0540] "Emotion recognition means" refers to technologies and devices used to analyze emotions from a customer's facial expressions and voice.
[0541] "Product identification information" refers to information that a user enters in order to identify a specific product.
[0542] A "generative machine learning model" is a machine learning algorithm that generates detailed product information based on product identification information and sentiment data.
[0543] A "database" is a collection of information where product information is stored in a predefined format.
[0544] A "user interface" is a means of interaction that allows a user to input information into a system.
[0545] "Emotional data" refers to numerical and categorical data of emotional states analyzed from a user's facial expressions and voice.
[0546] "Product description" refers to a detailed description of a product generated by a generative machine learning model.
[0547] The embodiments for carrying out the present invention are described in detail below. This system consists of a terminal, a server, a generative machine learning model, an emotion recognition means, and a communication means. The terminal, equipped with a user interface, receives input from the user, and the server is responsible for generating product descriptions.
[0548] The process begins when the user enters product identification information using a terminal device (e.g., smartphone, tablet, or digital signage). Once the user enters the product identification information and presses the search button, the terminal requests this information from the server. At the same time, the user's facial expressions and voice are also captured, and this data is analyzed in real time by emotion recognition technology.
[0549] The emotion recognition means acquires user emotion data using facial recognition technology and speech recognition technology. This emotion data is numerical data representing the user's emotional state, such as interest, doubt, and anxiety. The terminal means transmits this emotion data to the server along with product identification information.
[0550] The server parses the received request and retrieves the corresponding product information from the database based on the product identification information. It then passes the product information, along with sentiment data, to a generative machine learning model. The generative machine learning model uses the product information and sentiment data to generate a product description adapted to the user's emotional state. For example, if the user shows an interested expression, it provides a more detailed and technical explanation; if they show an anxious expression, it generates a simple and friendly explanation.
[0551] The generated product description is transmitted from the server to the terminal using a communication method and displayed on the terminal's user interface. This allows the user to view a detailed product description based on the entered product identification information in real time.
[0552] The hardware required to implement this system includes terminal devices such as smartphones and tablets, as well as servers equipped with GPUs. The software used includes OpenCV (for image processing), TensorFlow (for loading and using emotion recognition models), Transformers (for loading and using GPT-3 models), and requests (for API communication).
[0553] Specific example:
[0554] For example, if a user enters the product ID "CAM123" and looks interested, the generated description will be: "This camera can shoot high-resolution 4K video and take 20-megapixel photos. It also has a zoom function, allowing you to take clear photos of distant subjects." In this case, an example of the prompt message would be as follows:
[0555] Example of a prompt:
[0556] Product: 4K Video Camera
[0557] Description: A high-resolution 4K camera that can capture 20-megapixel photos and features zoom functionality.
[0558] Emotion: Curious
[0559] This format allows generative machine learning models to provide detailed product descriptions that are adapted to the user's emotional state.
[0560] The flow of a specific process in Application Example 2 will be explained using Figure 14.
[0561] Step 1:
[0562] The device acquires product identification information and sentiment data.
[0563] Users input product identification information using a terminal device (smartphone, tablet, digital signage, etc.). The system also captures the user's facial expressions and voice through the terminal's camera and microphone. This results in the acquisition of both product identification information and emotion data. This emotion data is analyzed in real time using image processing tools (e.g., OpenCV) and voice analysis tools.
[0564] Step 2:
[0565] Analysis and integration of emotional data
[0566] The device passes the acquired facial expression data and audio data to an emotion recognition system (an emotion recognition model using TensorFlow), which converts the user's emotional state into numerical data. For example, if the user has an expression that suggests interest, the category value "interesting" is obtained as emotion data. In this way, emotion data is generated by analyzing the input data (facial expression data and audio data).
[0567] Step 3:
[0568] Sending product identification information and emotional data to the server
[0569] The terminal sends product identification information and analyzed sentiment data to the server. This means that the product identification information and sentiment data are passed to the server as input from the terminal.
[0570] Step 4:
[0571] Product information acquisition
[0572] The server retrieves the corresponding product information from a predefined database based on the received product identification information. For example, it retrieves product information for product ID "CAM123" (e.g., high-resolution 4K camera, 20 megapixels, with zoom function). This allows the server to obtain product information from the input data (product identification information).
[0573] Step 5:
[0574] Product description generation using generative machine learning models
[0575] The server passes the acquired product information and sentiment data to a generative machine learning model (for example, a generative AI model using TensorFlow) to create a prompt. Based on this prompt, the generative machine learning model generates a product description adapted to the user's emotional state. For example, it might generate a description such as, "This camera can shoot high-resolution 4K video and take 20-megapixel photos. It also has a zoom function, allowing you to take clear photos of distant subjects."
[0576] Step 6:
[0577] Sending and displaying product descriptions to the device
[0578] Product descriptions generated by generative machine learning models are transmitted from the server to the terminal via communication. The terminal displays the received product descriptions on its user interface. This ensures that the output is presented in a format that the user can verify, based on the input (generated product descriptions).
[0579] Through these steps, users can receive detailed product descriptions tailored to their emotional state in real time.
[0580] The specific processing unit 290 transmits the result of the specific processing to the smart glasses 214. In the smart glasses 214, the control unit 46A causes the speaker 240 to output the result of the specific processing. The microphone 238 acquires audio indicating user input for the result of the specific processing. The control unit 46A transmits the audio data indicating user input acquired by the microphone 238 to the data processing unit 12. In the data processing unit 12, the specific processing unit 290 acquires the audio data.
[0581] Data generation model 58 is a type of so-called generative AI (Artificial Intelligence). One example of data generation model 58 is ChatGPT (Internet search<URL: https: / / openai.com / blog / chatgpt> ), Gemini (Internet search) <url: https: gemini.google.com ?hl="ja">Examples of generative AI include the following. The data generation model 58 is obtained by performing deep learning on a neural network. The data generation model 58 is input with prompts containing instructions, and with inference data such as audio data representing speech, text data representing text, and image data representing images. The data generation model 58 infers from the input inference data according to the instructions indicated by the prompts, and outputs the inference results in data formats such as audio data and text data. Here, inference refers to, for example, analysis, classification, prediction, and / or summarization.
[0582] In the above embodiment, an example was given in which specific processing is performed by the data processing device 12, but the technology of this disclosure is not limited thereto, and the specific processing may also be performed by the smart glasses 214.
[0583] [Third Embodiment]
[0584] Figure 5 shows an example of the configuration of the data processing system 310 according to the third embodiment.
[0585] As shown in Figure 5, the data processing system 310 includes a data processing device 12 and a headset terminal 314. An example of the data processing device 12 is a server.
[0586] The data processing device 12 comprises a computer 22, a database 24, and a communication interface 26. The computer 22 is an example of a "computer" related to the technology of this disclosure. The computer 22 comprises a processor 28, RAM 30, and storage 32. The processor 28, RAM 30, and storage 32 are connected to a bus 34. The database 24 and the communication interface 26 are also connected to the bus 34. The communication interface 26 is connected to a network 54. An example of the network 54 is a WAN (Wide Area Network) and / or a LAN (Local Area Network).
[0587] The headset terminal 314 includes a computer 36, a microphone 238, a speaker 240, a camera 42, a communication interface 44, and a display 343. The computer 36 includes a processor 46, RAM 48, and storage 50. The processor 46, RAM 48, and storage 50 are connected to a bus 52. The microphone 238, speaker 240, camera 42, and display 343 are also connected to the bus 52.
[0588] The microphone 238 receives voice signals from the user 20 and receives instructions from the user 20. The microphone 238 captures the voice signals from the user 20, converts the captured voice into audio data, and outputs it to the processor 46. The speaker 240 outputs audio according to the instructions from the processor 46.
[0589] Camera 42 is a small digital camera equipped with an optical system including a lens, aperture, and shutter, and an image sensor such as a CMOS (Complementary Metal-Oxide-Semiconductor) image sensor or a CCD (Charge Coupled Device) image sensor, and captures images of the area around the user 20 (for example, an imaging range defined by a field of view equivalent to the width of a typical healthy person's field of vision).
[0590] Communication interface 44 is connected to network 54. Communication interfaces 44 and 26 are responsible for the exchange of various information between processor 46 and processor 28 via network 54. The exchange of various information between processor 46 and processor 28 using communication interfaces 44 and 26 is performed in a secure manner.
[0591] Figure 6 shows an example of the main functions of the data processing device 12 and the headset terminal 314. As shown in Figure 6, the data processing device 12 performs specific processing using the processor 28. The storage 32 stores the specific processing program 56.
[0592] The specific processing program 56 is an example of a "program" relating to the technology of this disclosure. The processor 28 reads the specific processing program 56 from the storage 32 and executes the read specific processing program 56 on the RAM 30. The specific processing is realized by the processor 28 operating as a specific processing unit 290 in accordance with the specific processing program 56 executed on the RAM 30.
[0593] The storage 32 stores the data generation model 58 and the emotion identification model 59. The data generation model 58 and the emotion identification model 59 are used by the identification processing unit 290.
[0594] In the headset terminal 314, the processor 46 performs the reception output processing. The storage 50 stores the reception output program 60. The processor 46 reads the reception output program 60 from the storage 50 and executes the read reception output program 60 on the RAM 48. The reception output processing is realized by the processor 46 operating as a control unit 46A according to the reception output program 60 executed on the RAM 48.
[0595] Next, the specific processing performed by the specific processing unit 290 of the data processing device 12 will be described. In the following description, the data processing device 12 will be referred to as the "server" and the headset terminal 314 will be referred to as the "terminal".
[0596] The embodiments for carrying out the present invention are described in detail below. This system consists of a terminal, a server, a generative machine learning model, and a communication means. The terminal, equipped with a user interface, receives input from the user, and the server is responsible for generating product descriptions.
[0597] First, the user enters product identification information (e.g., product ID) using a device (e.g., tablet, smartphone, digital signage). After entering this information, the user requests product information from the server by pressing the search button.
[0598] The terminal sends an HTTP request to the server based on the product identification information received from the user. The server parses the received request and begins the process of retrieving the corresponding product information from the database. At this stage, the server searches the database for product data that matches the product identification information and passes the retrieved data to a generative machine learning model.
[0599] The generative machine learning model installed on the server generates detailed product descriptions based on the provided product data. This model has the ability to automatically generate user-friendly and detailed product descriptions using natural language processing techniques.
[0600] The generated product description is transmitted from the server to the terminal using a communication method. The terminal displays the received product description on its user interface. This allows the user to view a detailed product description in real time based on the product identification information they entered.
[0601] As a concrete example, consider the case where a user enters the product ID "CAM123". When the user enters the product ID and presses the search button, the terminal sends the product ID to the server. The server retrieves information about the camera corresponding to "CAM123" from the database (for example, "high-resolution 4K camera, 20 megapixels, with zoom function") and passes this information to a generative machine learning model. The model generates a product description such as, "This camera can shoot high-resolution 4K video and take 20-megapixel photos. It also has a zoom function, so you can take clear pictures of distant subjects." This product description is sent to the terminal and displayed on the user's screen.
[0602] This system allows even new crew members to provide detailed and accurate product descriptions without requiring extensive product knowledge. Furthermore, users can easily access product details by entering their own information. Therefore, this invention has the effect of improving customer satisfaction in the sale of home appliances, as well as enhancing the efficiency of crew members' work.
[0603] The following describes the processing flow.
[0604] Step 1:
[0605] The user enters product identification information (e.g., product ID "CAM123") using a device (e.g., tablet or smartphone). After entering this information, the user taps the "Search" button on the screen.
[0606] Step 2:
[0607] The terminal retrieves the product identification information entered by the user. This information is stored in a variable and used for the next process.
[0608] Step 3:
[0609] The device sends an HTTP request to the server based on the acquired product identification information. The request is sent in the format of a GET request, for example, with the product identification information included in the URL.
[0610] Step 4:
[0611] The server parses the request received from the terminal and extracts product identification information from the URL parameters. The extracted product identification information is stored in a variable.
[0612] Step 5:
[0613] The server executes queries against the product database based on the extracted product identification information. The queries are configured to retrieve information about products that match the product identification information.
[0614] Step 6:
[0615] The server retrieves relevant product information from the database (e.g., "High-resolution 4K camera, 20 megapixels, with zoom function"). The retrieved product information is then passed to a generative machine learning model.
[0616] Step 7:
[0617] The generative machine learning model installed on the server generates a detailed product description based on the product information provided. For example, the generated description might be: "This camera can shoot high-resolution 4K video and take 20-megapixel photos. It also has a zoom function, allowing you to take clear photos of distant subjects."
[0618] Step 8:
[0619] The server creates an HTTP response containing the generated product description and sends it to the terminal. The response format is, for example, JSON.
[0620] Step 9:
[0621] The terminal analyzes the response received from the server and extracts the product description text. The extracted text is then placed on a specific element of the user interface (for example, HTML). Display in the tags.
[0622] Step 10:
[0623] Users can view detailed product descriptions displayed on their device screens in real time. This allows users to obtain sufficient information about the product, making it easier for them to make a purchase decision.
[0624] Through the above processing steps, users can easily receive detailed product explanations, and a system is created that allows even new crew members to confidently serve customers.
[0625] (Example 1)
[0626] Next, we will describe Example 1. In the following description, the data processing device 12 will be referred to as the "server," and the headset-type terminal 314 will be referred to as the "terminal."
[0627] In the conventional system, obtaining detailed product information required staff with specialized knowledge, making it difficult for new or less knowledgeable crew members to handle such requests. Furthermore, it was difficult for customers to input product information themselves and receive detailed explanations in real time. This led to a decline in the quality of customer service and a risk of decreased customer satisfaction. Additionally, the inability to efficiently provide product information resulted in reduced operational efficiency.
[0628] The identification process performed by the identification processing unit 290 of the data processing device 12 in Example 1 is realized by the following means.
[0629] In this invention, the server includes terminal means equipped with a user interface for users to input product identification information, server means that receives product identification information from the terminal means, analyzes the product identification information, and obtains product information from a database, processing means including a generative machine learning model that generates a detailed product description using the product information obtained by the server means, and communication means that transmits the detailed product information generated by the generative machine learning model to the terminal means. As a result, even new crew members can provide detailed product descriptions without specialized knowledge, and users can input product information themselves and immediately obtain detailed explanations, thereby improving customer satisfaction and operational efficiency.
[0630] A "terminal device" is an electronic device equipped with a user interface for users to input product identification information.
[0631] A "server means" is a device or system that analyzes product identification information received from a terminal means and retrieves the corresponding product information from a database.
[0632] A "generative machine learning model" is a processing method for generating detailed product descriptions based on acquired product information.
[0633] "Communication means" refers to means for transmitting detailed information about the generated product to a terminal device.
[0634] "Product identification information" refers to information used to individually identify a product, such as a product ID or barcode.
[0635] A "database" is a system or location for storing product information.
[0636] A "user interface" refers to an interface that allows users to input information and view output results.
[0637] "Product information" refers to detailed information about a product, including its features and specifications.
[0638] The embodiments for carrying out the present invention are described in detail below. This system consists of terminal means, server means, a generative machine learning model, and communication means.
[0639] The terminal device is equipped with a user interface for the user to input product identification information. Specifically, electronic display devices (signage), personal digital assistants (tablets), and smart devices (smartphones) are used. The user uses these devices to input product identification information (e.g., product ID).
[0640] Once the user completes the input, the terminal sends the product identification information to the server as an HTTP request. The server receives this HTTP request and begins parsing. Specifically, it parses the product identification information and retrieves the corresponding product information from the database.
[0641] The database stores detailed product information, and the server retrieves the necessary product information using SQL queries and other methods. The retrieved product information is then passed to a generative machine learning model. This model uses natural language processing techniques to generate detailed product descriptions. Specifically, it has the ability to automatically generate user-friendly and detailed product descriptions based on the product information.
[0642] The generated product description is transmitted from the server to the terminal using a communication method. HTTP communication is used as the communication method. Finally, the terminal displays the received product details on the user interface. This allows the user to view detailed product descriptions in real time based on the product identification information they entered.
[0643] As a concrete example, consider the case where a user enters the product ID "CAM123". When the user enters the product ID "CAM123" and presses the search button, the terminal sends this information to the server. The server retrieves information about the camera corresponding to "CAM123" from its database (for example, "high-resolution 4K camera, 20 megapixels, with zoom function") and passes this information to a generative machine learning model. Based on this data, the generative machine learning model automatically generates a product description such as, "This camera can shoot high-resolution 4K video and take 20-megapixel photos. It also has a zoom function, so you can take clear pictures of distant subjects." This product description is finally sent to the terminal and displayed on the user's screen.
[0644] As an example of a prompt, the text "Generate a detailed product description for product ID 'CAM123'" is input to the generative machine learning model. A specific example of input to the model would be "High-resolution 4K camera, 20 megapixels, with zoom function."
[0645] This system allows users to check detailed product information in real time by entering their own product identification information. Furthermore, even new crew members can provide quick and accurate product descriptions without needing specialized knowledge. This is expected to improve customer satisfaction and operational efficiency in the sale of home appliances.
[0646] The flow of the specific processing in Example 1 will be explained using Figure 11.
[0647] Step 1:
[0648] The user enters product identification information.
[0649] Specific explanation: The user enters product identification information (e.g., product ID) into an input field on the terminal's user interface and presses the search button. The entered product identification information is, for example, "CAM123".
[0650] Input: Product identification information (product ID)
[0651] Output: Input product identification information
[0652] Step 2:
[0653] The device sends product identification information to the server.
[0654] Detailed explanation: The terminal sends an HTTP request to the server containing the entered product identification information. Internally, the terminal generates a request in the format "GET / search?product_id=CAM123".
[0655] Input: Entered product identification information (product ID)
[0656] Output: HTTP request to the server
[0657] Step 3:
[0658] The server parses the request.
[0659] Specific explanation: The server parses the received HTTP request and extracts product identification information from the query parameters. For example, it extracts the product ID "CAM123" from the request "GET / search?product_id=CAM123".
[0660] Input: HTTP Request
[0661] Output: Analyzed product identification information (product ID)
[0662] Step 4:
[0663] The server searches the database.
[0664] Specific explanation: The server executes an SQL query against the database based on the analyzed product identification information. It retrieves the relevant product information using the query "SELECT FROM products WHERE product_id = 'CAM123'".
[0665] Input: Analyzed product identification information (product ID)
[0666] Output: Acquired product information (e.g., "High-resolution 4K camera, 20 megapixels, with zoom function")
[0667] Step 5:
[0668] The server passes data to the generative machine learning model.
[0669] Detailed explanation: The server passes the acquired product information to a generative machine learning model. The generative machine learning model generates a detailed product description based on the given product information. The model incorporates natural language processing techniques.
[0670] Input: Acquired product information
[0671] Output: Generated product description (Example: "This camera can shoot high-resolution 4K video and take 20-megapixel photos. It also has a zoom function, allowing you to capture distant subjects clearly.")
[0672] Step 6:
[0673] The server sends the generated product description to the terminal.
[0674] Specific explanation: The server generates an HTTP response containing the product description generated by the generative machine learning model and sends it to the terminal. The response includes the description along with "200 OK".
[0675] Input: Generated product description
[0676] Output: HTTP response
[0677] Step 7:
[0678] The device displays the product description.
[0679] Specific explanation: The device parses the received HTTP response and displays the product description on the user interface. The user can view the detailed product description displayed on the device screen in real time.
[0680] Input: HTTP response
[0681] Output: Product description displayed on the user interface
[0682] (Application Example 1)
[0683] Next, we will explain Application Example 1. In the following explanation, the data processing device 12 will be referred to as the "server," and the headset-type terminal 314 will be referred to as the "terminal."
[0684] Traditional methods of product description in physical stores rely on the skill level and consistency of staff explanations, making it difficult to provide customers with appropriate and detailed information quickly. Furthermore, even when systems existed that automatically generated product descriptions, they lacked features such as voice assistants or scanning capabilities, failing to adequately consider user convenience. As a result, challenges remain in improving customer satisfaction and reducing staff workload.
[0685] The specific processing performed by the specific processing unit 290 of the data processing device 12 in Application Example 1 is realized by the following means.
[0686] In this invention, the server includes terminal means equipped with a user interface for users to input product identification information; processing means including a generative machine learning model that receives product identification information from the terminal means and generates detailed product information based on the product identification information; communication means for transmitting the detailed product information generated by the generative machine learning model to the terminal means; voice output means equipped with a voice assistant function for providing product descriptions by voice; and scanning means for obtaining product identification information by having customers scan a QR code or barcode and displaying it on the terminal. This makes it possible to provide customers with real-time, consistent, and detailed product descriptions, regardless of the skill level of the staff. In addition, the voice assistant function and scanning function allow customers to easily obtain product information themselves, simultaneously improving customer satisfaction and reducing the burden on staff.
[0687] A "terminal device" refers to a device equipped with an interface for users to input product identification information. Examples include digital signage, tablets, smartphones, and smart glasses.
[0688] A "generative machine learning model" refers to artificial intelligence technology that automatically generates detailed product descriptions based on input product identification information.
[0689] "Communication means" refers to the technology and methods for transmitting detailed product information generated by a generative machine learning model to a terminal device.
[0690] "Voice output means" refers to a function that provides generated product descriptions in audio format. This allows users to obtain product information through audio.
[0691] "Scanning method" refers to technology that allows customers to obtain product identification information by scanning a QR code or barcode and display it on their device.
[0692] "User interface" is a general term for the display and input methods used by users to input product identification information and to review the generated product description.
[0693] The system of this invention consists of terminal means, a server, a generative machine learning model, communication means, voice output means, and scanning means. Specific embodiments of the system using these components will be described in detail below.
[0694] Hardware and software
[0695] Terminal devices: Digital signage, tablets, smartphones, smart glasses, and other devices.
[0696] Server: A cloud server using Amazon Web Services (AWS) or Google Cloud Platform (GCP).
[0697] Use either MySQL or PostgreSQL as the database.
[0698] Generative machine learning models: Generative machine learning models including OpenAI GPT-4 or Google BERT.
[0699] Audio output method: The speaker or earphones built into the device.
[0700] Scanning method: Camera on the device or a dedicated barcode scanner.
[0701] Program Overview
[0702] The server receives product identification information from a terminal device equipped with a user interface for users to input product identification information. For example, the user can obtain product information by entering a product ID on a tablet's input screen or by scanning a QR code or barcode with a smartphone's camera. This allows the product identification information to be acquired through the scanning device.
[0703] The received product identification information is transmitted to the server via a communication method. The server searches the database based on this product identification information and retrieves the corresponding product data. The retrieved product data is passed to a generative machine learning model, which uses natural language processing techniques to generate a detailed product description.
[0704] The generated product description is then transmitted back to the terminal via a communication device. The user can view the detailed product description in real time on the terminal screen. Furthermore, it is possible to listen to the generated product description aloud using an audio output device. This allows users to obtain product information both visually and aurally, improving convenience.
[0705] Specific example
[0706] For example, consider a case where a user enters the product ID "CAM123". When the user enters "CAM123" into a tablet and presses the search button, the device sends the product ID to the server. The server retrieves information about the camera corresponding to "CAM123" from its database (for example, "High-resolution 4K camera, 20 megapixels, with zoom function") and passes this information to a generative machine learning model. The model generates a product description like the following:
[0707] Product ID: CAM123
[0708] This camera can shoot high-resolution 4K video and take 20-megapixel photos. It also features a zoom function, allowing you to capture distant subjects clearly.
[0709] The generated product description is displayed on the device screen and also provided to the user via audio output. In this way, the user can obtain product information through both sight and sound.
[0710] The flow of a specific process in Application Example 1 will be explained using Figure 12.
[0711] Step 1: Enter product identification information
[0712] The user inputs product identification information (e.g., product ID, QR code, barcode) using a terminal device (e.g., smartphone or tablet). At this time, the user interface saves the entered product identification information to the terminal device's memory.
[0713] Input: Product ID, QR code, or barcode
[0714] Output: Product identification information
[0715] Step 2: Submit product identification information
[0716] The terminal device transmits the product identification information obtained in step 1 to the server using the communication device. Specifically, an HTTP request is generated and sent to the server's endpoint.
[0717] Input: Product identification information
[0718] Output: HTTP request sent to the server
[0719] Step 3: Search the database
[0720] The server queries the database based on the received product identification information to retrieve the corresponding product data. It executes an SQL query to fetch the relevant product data.
[0721] Input: Product identification information
[0722] Output: Product data
[0723] Step 4: Explanation generation using generative machine learning models
[0724] The server inputs the product data obtained in step 3 into a generative machine learning model to generate detailed product descriptions. The model uses natural language processing techniques to convert the product data into easy-to-understand text.
[0725] Input: Product data
[0726] Output: Product Description
[0727] Step 5: Send the generated product description.
[0728] The server sends the generated product description to the terminal device via a communication means. An HTTP response is generated and sent to the endpoint of the terminal device.
[0729] Input: Product Description
[0730] Output: HTTP response sent to the terminal device
[0731] Step 6: Display and audio output of product description
[0732] The terminal device displays the received product description on the user interface and also provides the product description in audio using the audio output device.
[0733] Input: Product Description
[0734] Output: Screen display and audio output
[0735] As a concrete example, consider the case where a user enters the product ID "CAM123". When the user enters the product ID and presses the search button, the terminal sends the product ID to the server. The server retrieves information about the camera corresponding to "CAM123" from its database (for example, "High-resolution 4K camera, 20 megapixels, with zoom function") and passes this information to a generative machine learning model. The model generates a product description like the following:
[0736] Product ID: CAM123
[0737] This camera can shoot high-resolution 4K video and take 20-megapixel photos. It also features a zoom function, allowing you to capture distant subjects clearly.
[0738] The generated product description is displayed on the device screen and also provided to the user via audio output. In this way, the user can obtain product information through both sight and sound.
[0739] Furthermore, an emotion engine that estimates the user's emotions may be incorporated. That is, the identification processing unit 290 may use the emotion identification model 59 to estimate the user's emotions and perform identification processing using the user's emotions.
[0740] The embodiments for carrying out the present invention are described in detail below. This system consists of a terminal, a server, a generative machine learning model, an emotion engine, and a communication means. The terminal, equipped with a user interface, receives input from the user, and the server is responsible for generating product descriptions.
[0741] First, the user enters product identification information (e.g., product ID) using a device (e.g., tablet, smartphone, digital signage). After entering this information, the user requests product information from the server by pressing the search button.
[0742] The device acquires emotional data from product identification information received from the user and from an emotion engine that recognizes emotions from the user's facial expressions and voice. The emotion engine uses facial recognition technology and voice recognition technology to evaluate the user's emotions in real time.
[0743] The device sends product identification information and sentiment data to the server. The server analyzes the received request, accesses the database using the product identification information, and retrieves the corresponding product information. This information, along with the sentiment data, is then passed to a generative machine learning model.
[0744] Generative machine learning models can generate product descriptions adapted to the user's emotional state based on the provided product information and sentiment data. For example, if the user shows an interested expression, it will provide a more detailed and technical explanation. On the other hand, if the user shows an anxious expression, it will generate a simple and reassuring explanation.
[0745] The generated product description is transmitted from the server to the terminal using a communication method. The terminal displays the received product description on its user interface. This allows the user to view a detailed product description in real time based on the product identification information they entered.
[0746] As a concrete example, consider the case where a user enters the product ID "CAM123". When the user enters the product ID and presses the search button, the terminal sends the product ID along with emotional data acquired from the user's facial expressions and voice to the server. The server retrieves information about the camera corresponding to "CAM123" from the database (for example, "high-resolution 4K camera, 20 megapixels, with zoom function") and passes this information and emotional data to a generative machine learning model. The model generates an appropriate product description according to the user's emotional state, adjusting content such as, "This camera can shoot high-resolution 4K video and take 20-megapixel photos. It also has a zoom function, so you can take clear pictures of distant subjects." For example, if the user seems interested, detailed technical explanations will be added, and if they are feeling anxious, more approachable terminology will be used.
[0747] This system allows even new crew members, without requiring extensive product knowledge, to provide detailed and accurate product descriptions tailored to the user's emotional state. Furthermore, users can easily learn about product details in an emotionally sensitive way by inputting product information themselves. Therefore, this invention has the effect of improving customer satisfaction in the sale of home appliances and enhancing the efficiency of crew members' work.
[0748] The following describes the processing flow.
[0749] Step 1:
[0750] The user enters product identification information (e.g., product ID "CAM123") using a device (e.g., tablet or smartphone). After entering this information, the user taps the "Search" button on the screen.
[0751] Step 2:
[0752] The terminal retrieves the product identification information entered by the user. This information is stored in a variable and used for the next process.
[0753] Step 3:
[0754] The device analyzes the user's facial expressions and voice using an emotion engine to obtain the user's emotional data (e.g., "interest," "anxiety," "confusion," etc.). This data is also stored in variables.
[0755] Step 4:
[0756] The device sends an HTTP request to the server based on the acquired product identification information and sentiment data. The request is sent in the format of a POST request, for example, containing the product identification information and sentiment data.
[0757] Step 5:
[0758] The server analyzes the request received from the terminal and extracts product identification information and sentiment data. The extracted data is then stored in variables.
[0759] Step 6:
[0760] The server uses product identification information to query the product database. The query is configured to retrieve information about products that match the product identification information.
[0761] Step 7:
[0762] The server retrieves relevant product information from the database (for example, "high-resolution 4K camera, 20 megapixels, with zoom function"). The retrieved product information and sentiment data are then passed to a generative machine learning model.
[0763] Step 8:
[0764] Generative machine learning models generate detailed product descriptions based on the product information and sentiment data they receive. For example, the model adjusts a description such as "This camera can shoot high-resolution 4K video and take 20-megapixel photos. It also has a zoom function, allowing you to capture distant subjects clearly" according to the user's emotional state.
[0765] Step 9:
[0766] The server creates an HTTP response containing the generated product description and sends it to the terminal. The response format is, for example, JSON.
[0767] Step 10:
[0768] The terminal analyzes the response received from the server and extracts the product description text. The extracted text is then placed on a specific element of the user interface (for example, HTML). Display in the tags.
[0769] Step 11:
[0770] Users can view detailed product descriptions displayed on their device screens in real time. This allows users to obtain sufficient information about the product, making it easier for them to make a purchase decision.
[0771] These processing steps enable users to easily receive detailed product descriptions, creating a system that allows even new crew members to confidently interact with customers. Furthermore, by utilizing emotional data, product descriptions optimized for the user's emotional state are provided.
[0772] (Example 2)
[0773] Next, we will describe Example 2. In the following description, the data processing device 12 will be referred to as the "server," and the headset-type terminal 314 will be referred to as the "terminal."
[0774] In recent years, providing users with appropriate product information has become increasingly important in e-commerce and brick-and-mortar retail. However, many systems provide uniform information without considering user emotions, resulting in insufficient user experience and customer satisfaction. Furthermore, a shortage of staff and sales personnel who understand and can provide detailed technical information about products is also a problem. Therefore, there is a need for a system that provides effective product descriptions in real time, tailored to the user's emotional state.
[0775] The identification processing performed by the identification processing unit 290 of the data processing device 12 in Example 2 is realized by the following means. In this invention, the server includes terminal means equipped with a user interface for the user to input product identification information and emotion data; processing means including a generative machine learning model that receives product identification information and emotion data from the terminal means and generates product information based on the product identification information and detailed product information adapted to the user's emotional state; and communication means that transmits the detailed product information generated by the generative machine learning model to the terminal means. This makes it possible to provide detailed and appropriate product descriptions in real time that correspond to the user's emotions.
[0776] "Terminal means" refers to devices used by users to input product identification information and emotional data, and includes tablets, smartphones, and digital signage.
[0777] "Product identification information" refers to information used to identify a product, and includes, for example, product IDs and barcodes.
[0778] "Emotional data" refers to data that indicates the emotional state of a user, analyzed from their facial expressions and voice.
[0779] "User interface" is a general term for the screens and input devices that users use to operate a device, and includes touch panels and button interfaces.
[0780] A "generative machine learning model" is a model that uses machine learning techniques to generate appropriate product descriptions based on product information and sentiment data.
[0781] "Processing means" refers to a device or system that processes information received from a terminal means, and includes a generative machine learning model.
[0782] "Communication means" refers to technologies and devices for sending and receiving data between a server and a terminal, and includes the internet and local networks.
[0783] A "database" is a system or device that manages and stores data such as product information.
[0784] An "emotion engine" is a system or device that analyzes a user's facial expressions and voice to generate emotional data.
[0785] This invention is a system for improving the user experience when acquiring product information. This system consists of a terminal, an emotion engine, a server, a generative machine learning model, and communication means. The following describes specific implementations of this system.
[0786] System Configuration
[0787] The terminal device is a device equipped with a user interface, used by users to input product identification information and emotional data. Specifically, this includes tablets, smartphones, and digital signage. The terminal device uses a built-in camera and microphone to analyze the user's facial expressions and voice, and generates emotional data using an emotion engine.
[0788] The emotion engine uses facial recognition and speech recognition technologies to analyze emotional data from the user's facial expressions and voice in real time. This allows for an accurate understanding of the user's current emotional state.
[0789] The server receives product identification information and sentiment data transmitted from the terminal device. Furthermore, the server accesses a predefined database to retrieve the relevant product information. The retrieved product information and sentiment data are then passed to a generative machine learning model.
[0790] Generative machine learning models generate detailed product information tailored to the user's emotional state, based on product information and sentiment data. For example, if the user shows an interested expression, it provides a more detailed and technical explanation. On the other hand, if the user shows an anxious expression, it generates a simple and reassuring explanation.
[0791] The communication means plays the role of transmitting detailed information about the generated product from the server to the terminal means.
[0792] Specific example
[0793] For example, consider a case where a user enters the product ID "CAM123". When the user enters the product ID and presses the search button, the terminal sends the product ID along with emotional data acquired from the user's facial expressions and voice to the server. The server retrieves product information for the camera corresponding to "CAM123" from the database (for example, "High-resolution 4K camera, 20 megapixels, with zoom function") and passes this information and emotional data to a generative machine learning model. The model generates an appropriate product description according to the user's emotional state, adjusting content such as, "This camera can shoot high-resolution 4K video and take 20-megapixel photos. It also has a zoom function, so you can take clear pictures of distant subjects." For example, if the user seems interested, detailed technical explanations will be added, and if they are anxious, more approachable terminology will be used.
[0794] Example of a prompt
[0795] Create a description of the high-resolution 4K camera that matches the user's expression of interest.
[0796] Please create a description of the 20-megapixel camera that is tailored to the user's anxious expression.
[0797] This system allows even new crew members to provide detailed and accurate product descriptions tailored to the user's emotional state, even without extensive product knowledge. Furthermore, users can easily learn about product details in an emotionally sensitive way by inputting product information themselves. Therefore, it improves customer satisfaction in home appliance sales and enhances the efficiency of crew members' work.
[0798] The flow of the specific processing in Example 2 will be explained using Figure 13.
[0799] Step 1:
[0800] The user enters product identification information (e.g., product ID "CAM123") using the terminal. When the user presses the search button, this product ID is saved as input data on the terminal. The terminal's user interface receives this product identification information, preparing the data to be used in the next step.
[0801] Step 2:
[0802] The device uses its built-in camera and microphone to capture the user's facial expressions and voice. The device sends this data to an emotion engine, which analyzes the user's emotional data in real time. Specifically, the device captures the user's face with its camera and acquires emotional data such as "interested facial expressions" and "cheerful voice tones."
[0803] Step 3:
[0804] The terminal sends product identification information (e.g., "CAM123") and acquired emotion data to the server. By receiving product identification information and emotion data as input data, the server analyzes this data and prepares it for use in the next step.
[0805] Step 4:
[0806] The server accesses the database based on the product identification information and retrieves the corresponding product information. Specifically, the server executes a database query to retrieve product information for cameras corresponding to "CAM123" (e.g., "High-resolution 4K camera, 20 megapixels, with zoom function"). The input data is the product identification information, and the output data is the product information.
[0807] Step 5:
[0808] The server passes the acquired product information and sentiment data to a generative machine learning model. The generative machine learning model receives the product information and sentiment data as input and generates detailed product information adapted to the user's emotional state. Specifically, the generative machine learning model considers that "the user appears interested" and generates a product description that includes detailed technical explanations. The output data is a sentiment-based, adjusted product description.
[0809] Step 6:
[0810] The generated product description is transmitted from the server to the terminal via a communication method. Specifically, the server generates the product description and sends it to the terminal as a data packet over the internet. The input data is the generated product description, and the output data is the result of the transmission to the terminal.
[0811] Step 7:
[0812] The device displays the received product description on its user interface. Users can view detailed product descriptions in real time on the device's screen. Specifically, the device's touchscreen displays a text message such as, "This camera can shoot high-resolution 4K video and 20-megapixel photos. It also has a zoom function, allowing you to take clear photos of distant subjects." The output data is the displayed product description.
[0813] (Application Example 2)
[0814] Next, we will explain application example 2. In the following explanation, the data processing device 12 will be referred to as the "server," and the headset-type terminal 314 will be referred to as the "terminal."
[0815] When customers choose products in physical stores, they may feel anxious because they lack detailed information about the products. Furthermore, the uniform nature of product descriptions makes it difficult to provide personalized information tailored to individual customers' emotions and needs. This situation can lead to decreased customer satisfaction and reduced purchasing intent. To address this, there is a need for a system that analyzes customer emotions from their facial expressions and voice, and generates and displays product descriptions adapted to those emotions in real time.
[0816] The specific processing performed by the specific processing unit 290 of the data processing device 12 in Application Example 2 is realized by the following means.
[0817] In this invention, the server includes emotion recognition means for analyzing the user's emotions from their facial expressions and voice, processing means including a generative machine learning model that receives product identification information and emotion data and generates a detailed product description, and communication means for transmitting the generated product details to a terminal means. This makes it possible to provide personalized product descriptions that are adapted to the customer's emotions.
[0818] A "terminal device" is a device equipped with an interface for the user to input product identification information.
[0819] "Processing means" refers to a device that includes a generative machine learning model that generates detailed product information based on product identification information and sentiment data.
[0820] "Communication means" refers to a device that transmits detailed product information generated by a generative machine learning model to a terminal means.
[0821] "Emotion recognition means" refers to technologies and devices used to analyze emotions from a customer's facial expressions and voice.
[0822] "Product identification information" refers to information that a user enters in order to identify a specific product.
[0823] A "generative machine learning model" is a machine learning algorithm that generates detailed product information based on product identification information and sentiment data.
[0824] A "database" is a collection of information where product information is stored in a predefined format.
[0825] A "user interface" is a means of interaction that allows a user to input information into a system.
[0826] "Emotional data" refers to numerical and categorical data of emotional states analyzed from a user's facial expressions and voice.
[0827] "Product description" refers to a detailed description of a product generated by a generative machine learning model.
[0828] The embodiments for carrying out the present invention are described in detail below. This system consists of a terminal, a server, a generative machine learning model, an emotion recognition means, and a communication means. The terminal, equipped with a user interface, receives input from the user, and the server is responsible for generating product descriptions.
[0829] The process begins when the user enters product identification information using a terminal device (e.g., smartphone, tablet, or digital signage). Once the user enters the product identification information and presses the search button, the terminal requests this information from the server. At the same time, the user's facial expressions and voice are also captured, and this data is analyzed in real time by emotion recognition technology.
[0830] The emotion recognition means acquires user emotion data using facial recognition technology and speech recognition technology. This emotion data is numerical data representing the user's emotional state, such as interest, doubt, and anxiety. The terminal means transmits this emotion data to the server along with product identification information.
[0831] The server parses the received request and retrieves the corresponding product information from the database based on the product identification information. It then passes the product information, along with sentiment data, to a generative machine learning model. The generative machine learning model uses the product information and sentiment data to generate a product description adapted to the user's emotional state. For example, if the user shows an interested expression, it provides a more detailed and technical explanation; if they show an anxious expression, it generates a simple and friendly explanation.
[0832] The generated product description is transmitted from the server to the terminal using a communication method and displayed on the terminal's user interface. This allows the user to view a detailed product description based on the entered product identification information in real time.
[0833] The hardware required to implement this system includes terminal devices such as smartphones and tablets, as well as servers equipped with GPUs. The software used includes OpenCV (for image processing), TensorFlow (for loading and using emotion recognition models), Transformers (for loading and using GPT-3 models), and requests (for API communication).
[0834] Specific example:
[0835] For example, if a user enters the product ID "CAM123" and looks interested, the generated description will be: "This camera can shoot high-resolution 4K video and take 20-megapixel photos. It also has a zoom function, allowing you to take clear photos of distant subjects." In this case, an example of the prompt message would be as follows:
[0836] Example of a prompt:
[0837] Product: 4K Video Camera
[0838] Description: A high-resolution 4K camera that can capture 20-megapixel photos and features zoom functionality.
[0839] Emotion: Curious
[0840] This format allows generative machine learning models to provide detailed product descriptions that are adapted to the user's emotional state.
[0841] The flow of a specific process in Application Example 2 will be explained using Figure 14.
[0842] Step 1:
[0843] The device acquires product identification information and sentiment data.
[0844] Users input product identification information using a terminal device (smartphone, tablet, digital signage, etc.). The system also captures the user's facial expressions and voice through the terminal's camera and microphone. This results in the acquisition of both product identification information and emotion data. This emotion data is analyzed in real time using image processing tools (e.g., OpenCV) and voice analysis tools.
[0845] Step 2:
[0846] Analysis and integration of emotional data
[0847] The device passes the acquired facial expression data and audio data to an emotion recognition system (an emotion recognition model using TensorFlow), which converts the user's emotional state into numerical data. For example, if the user has an expression that suggests interest, the category value "interesting" is obtained as emotion data. In this way, emotion data is generated by analyzing the input data (facial expression data and audio data).
[0848] Step 3:
[0849] Sending product identification information and emotional data to the server
[0850] The terminal sends product identification information and analyzed sentiment data to the server. This means that the product identification information and sentiment data are passed to the server as input from the terminal.
[0851] Step 4:
[0852] Product information acquisition
[0853] The server retrieves the corresponding product information from a predefined database based on the received product identification information. For example, it retrieves product information for product ID "CAM123" (e.g., high-resolution 4K camera, 20 megapixels, with zoom function). This allows the server to obtain product information from the input data (product identification information).
[0854] Step 5:
[0855] Product description generation using generative machine learning models
[0856] The server passes the acquired product information and sentiment data to a generative machine learning model (for example, a generative AI model using TensorFlow) to create a prompt. Based on this prompt, the generative machine learning model generates a product description adapted to the user's emotional state. For example, it might generate a description such as, "This camera can shoot high-resolution 4K video and take 20-megapixel photos. It also has a zoom function, allowing you to take clear photos of distant subjects."
[0857] Step 6:
[0858] Sending and displaying product descriptions to the device
[0859] Product descriptions generated by generative machine learning models are transmitted from the server to the terminal via communication. The terminal displays the received product descriptions on its user interface. This ensures that the output is presented in a format that the user can verify, based on the input (generated product descriptions).
[0860] Through these steps, users can receive detailed product descriptions tailored to their emotional state in real time.
[0861] The specific processing unit 290 transmits the result of the specific processing to the headset terminal 314. In the headset terminal 314, the control unit 46A causes the speaker 240 and display 343 to output the result of the specific processing. The microphone 238 acquires audio indicating user input for the result of the specific processing. The control unit 46A transmits the audio data indicating user input acquired by the microphone 238 to the data processing unit 12. In the data processing unit 12, the specific processing unit 290 acquires the audio data.
[0862] Data generation model 58 is a type of so-called generative AI (Artificial Intelligence). One example of data generation model 58 is ChatGPT (Internet search<URL: https: / / openai.com / blog / chatgpt> ), Gemini (Internet search) <url: https: gemini.google.com ?hl="ja">Examples of generative AI include the following. The data generation model 58 is obtained by performing deep learning on a neural network. The data generation model 58 is input with prompts containing instructions, and with inference data such as audio data representing speech, text data representing text, and image data representing images. The data generation model 58 infers from the input inference data according to the instructions indicated by the prompts, and outputs the inference results in data formats such as audio data and text data. Here, inference refers to, for example, analysis, classification, prediction, and / or summarization.
[0863] In the above embodiment, an example was given in which specific processing is performed by the data processing device 12, but the technology of this disclosure is not limited thereto, and specific processing may also be performed by the headset terminal 314.
[0864] [Fourth Embodiment]
[0865] Figure 7 shows an example of the configuration of the data processing system 410 according to the fourth embodiment.
[0866] As shown in Figure 7, the data processing system 410 includes a data processing device 12 and a robot 414. An example of the data processing device 12 is a server.
[0867] The data processing device 12 comprises a computer 22, a database 24, and a communication interface 26. The computer 22 is an example of a "computer" related to the technology of this disclosure. The computer 22 comprises a processor 28, RAM 30, and storage 32. The processor 28, RAM 30, and storage 32 are connected to a bus 34. The database 24 and the communication interface 26 are also connected to the bus 34. The communication interface 26 is connected to a network 54. An example of the network 54 is a WAN (Wide Area Network) and / or a LAN (Local Area Network).
[0868] The robot 414 includes a computer 36, a microphone 238, a speaker 240, a camera 42, a communication interface 44, and a controlled object 443. The computer 36 includes a processor 46, RAM 48, and storage 50. The processor 46, RAM 48, and storage 50 are connected to a bus 52. The microphone 238, speaker 240, camera 42, and controlled object 443 are also connected to the bus 52.
[0869] The microphone 238 receives voice signals from the user 20 and receives instructions from the user 20. The microphone 238 captures the voice signals from the user 20, converts the captured voice into audio data, and outputs it to the processor 46. The speaker 240 outputs audio according to the instructions from the processor 46.
[0870] Camera 42 is a small digital camera equipped with an optical system including a lens, aperture, and shutter, and an image sensor such as a CMOS (Complementary Metal-Oxide-Semiconductor) image sensor or a CCD (Charge Coupled Device) image sensor, and captures images of the area around the user 20 (for example, an imaging range defined by a field of view equivalent to the width of a typical healthy person's field of vision).
[0871] Communication interface 44 is connected to network 54. Communication interfaces 44 and 26 are responsible for the exchange of various information between processor 46 and processor 28 via network 54. The exchange of various information between processor 46 and processor 28 using communication interfaces 44 and 26 is performed in a secure manner.
[0872] The controlled object 443 includes a display device, LEDs in the eyes, and motors that drive the arms, hands, and feet. The posture and gestures of the robot 414 are controlled by controlling the motors of the arms, hands, and feet. Some of the robot 414's emotions can be expressed by controlling these motors. Furthermore, the robot 414's facial expressions can also be expressed by controlling the illumination state of the LEDs in its eyes.
[0873] Figure 8 shows an example of the main functions of the data processing device 12 and the robot 414. As shown in Figure 8, the data processing device 12 performs specific processing using the processor 28. The storage 32 stores the specific processing program 56.
[0874] The specific processing program 56 is an example of a "program" relating to the technology of this disclosure. The processor 28 reads the specific processing program 56 from the storage 32 and executes the read specific processing program 56 on the RAM 30. The specific processing is realized by the processor 28 operating as a specific processing unit 290 in accordance with the specific processing program 56 executed on the RAM 30.
[0875] The storage 32 stores the data generation model 58 and the emotion identification model 59. The data generation model 58 and the emotion identification model 59 are used by the identification processing unit 290.
[0876] In robot 414, the processor 46 performs the reception output processing. The storage 50 stores the reception output program 60. The processor 46 reads the reception output program 60 from the storage 50 and executes the read reception output program 60 on the RAM 48. The reception output processing is realized by the processor 46 operating as a control unit 46A according to the reception output program 60 executed on the RAM 48.
[0877] Next, the specific processing performed by the specific processing unit 290 of the data processing device 12 will be described. In the following description, the data processing device 12 will be referred to as the "server" and the robot 414 as the "terminal".
[0878] The embodiments for carrying out the present invention are described in detail below. This system consists of a terminal, a server, a generative machine learning model, and a communication means. The terminal, equipped with a user interface, receives input from the user, and the server is responsible for generating product descriptions.
[0879] First, the user enters product identification information (e.g., product ID) using a device (e.g., tablet, smartphone, digital signage). After entering this information, the user requests product information from the server by pressing the search button.
[0880] The terminal sends an HTTP request to the server based on the product identification information received from the user. The server parses the received request and begins the process of retrieving the corresponding product information from the database. At this stage, the server searches the database for product data that matches the product identification information and passes the retrieved data to a generative machine learning model.
[0881] The generative machine learning model installed on the server generates detailed product descriptions based on the provided product data. This model has the ability to automatically generate user-friendly and detailed product descriptions using natural language processing techniques.
[0882] The generated product description is transmitted from the server to the terminal using a communication method. The terminal displays the received product description on its user interface. This allows the user to view a detailed product description in real time based on the product identification information they entered.
[0883] As a concrete example, consider the case where a user enters the product ID "CAM123". When the user enters the product ID and presses the search button, the terminal sends the product ID to the server. The server retrieves information about the camera corresponding to "CAM123" from the database (for example, "high-resolution 4K camera, 20 megapixels, with zoom function") and passes this information to a generative machine learning model. The model generates a product description such as, "This camera can shoot high-resolution 4K video and take 20-megapixel photos. It also has a zoom function, so you can take clear pictures of distant subjects." This product description is sent to the terminal and displayed on the user's screen.
[0884] This system allows even new crew members to provide detailed and accurate product descriptions without requiring extensive product knowledge. Furthermore, users can easily access product details by entering their own information. Therefore, this invention has the effect of improving customer satisfaction in the sale of home appliances, as well as enhancing the efficiency of crew members' work.
[0885] The following describes the processing flow.
[0886] Step 1:
[0887] The user enters product identification information (e.g., product ID "CAM123") using a device (e.g., tablet or smartphone). After entering this information, the user taps the "Search" button on the screen.
[0888] Step 2:
[0889] The terminal retrieves the product identification information entered by the user. This information is stored in a variable and used for the next process.
[0890] Step 3:
[0891] The device sends an HTTP request to the server based on the acquired product identification information. The request is sent in the format of a GET request, for example, with the product identification information included in the URL.
[0892] Step 4:
[0893] The server parses the request received from the terminal and extracts product identification information from the URL parameters. The extracted product identification information is stored in a variable.
[0894] Step 5:
[0895] The server executes queries against the product database based on the extracted product identification information. The queries are configured to retrieve information about products that match the product identification information.
[0896] Step 6:
[0897] The server retrieves relevant product information from the database (e.g., "High-resolution 4K camera, 20 megapixels, with zoom function"). The retrieved product information is then passed to a generative machine learning model.
[0898] Step 7:
[0899] The generative machine learning model installed on the server generates a detailed product description based on the product information provided. For example, the generated description might be: "This camera can shoot high-resolution 4K video and take 20-megapixel photos. It also has a zoom function, allowing you to take clear photos of distant subjects."
[0900] Step 8:
[0901] The server creates an HTTP response containing the generated product description and sends it to the terminal. The response format is, for example, JSON.
[0902] Step 9:
[0903] The terminal analyzes the response received from the server and extracts the product description text. The extracted text is then placed on a specific element of the user interface (for example, HTML). Display in the tags.
[0904] Step 10:
[0905] Users can view detailed product descriptions displayed on their device screens in real time. This allows users to obtain sufficient information about the product, making it easier for them to make a purchase decision.
[0906] Through the above processing steps, users can easily receive detailed product explanations, and a system is created that allows even new crew members to confidently serve customers.
[0907] (Example 1)
[0908] Next, we will describe Example 1. In the following description, the data processing device 12 will be referred to as the "server" and the robot 414 as the "terminal".
[0909] In the conventional system, obtaining detailed product information required staff with specialized knowledge, making it difficult for new or less knowledgeable crew members to handle such requests. Furthermore, it was difficult for customers to input product information themselves and receive detailed explanations in real time. This led to a decline in the quality of customer service and a risk of decreased customer satisfaction. Additionally, the inability to efficiently provide product information resulted in reduced operational efficiency.
[0910] The identification process performed by the identification processing unit 290 of the data processing device 12 in Example 1 is realized by the following means.
[0911] In this invention, the server includes terminal means equipped with a user interface for users to input product identification information, server means that receives product identification information from the terminal means, analyzes the product identification information, and obtains product information from a database, processing means including a generative machine learning model that generates a detailed product description using the product information obtained by the server means, and communication means that transmits the detailed product information generated by the generative machine learning model to the terminal means. As a result, even new crew members can provide detailed product descriptions without specialized knowledge, and users can input product information themselves and immediately obtain detailed explanations, thereby improving customer satisfaction and operational efficiency.
[0912] A "terminal device" is an electronic device equipped with a user interface for users to input product identification information.
[0913] A "server means" is a device or system that analyzes product identification information received from a terminal means and retrieves the corresponding product information from a database.
[0914] A "generative machine learning model" is a processing method for generating detailed product descriptions based on acquired product information.
[0915] "Communication means" refers to means for transmitting detailed information about the generated product to a terminal device.
[0916] "Product identification information" refers to information used to individually identify a product, such as a product ID or barcode.
[0917] A "database" is a system or location for storing product information.
[0918] A "user interface" refers to an interface that allows users to input information and view output results.
[0919] "Product information" refers to detailed information about a product, including its features and specifications.
[0920] The embodiments for carrying out the present invention are described in detail below. This system consists of terminal means, server means, a generative machine learning model, and communication means.
[0921] The terminal device is equipped with a user interface for the user to input product identification information. Specifically, electronic display devices (signage), personal digital assistants (tablets), and smart devices (smartphones) are used. The user uses these devices to input product identification information (e.g., product ID).
[0922] Once the user completes the input, the terminal sends the product identification information to the server as an HTTP request. The server receives this HTTP request and begins parsing. Specifically, it parses the product identification information and retrieves the corresponding product information from the database.
[0923] The database stores detailed product information, and the server retrieves the necessary product information using SQL queries and other methods. The retrieved product information is then passed to a generative machine learning model. This model uses natural language processing techniques to generate detailed product descriptions. Specifically, it has the ability to automatically generate user-friendly and detailed product descriptions based on the product information.
[0924] The generated product description is transmitted from the server to the terminal using a communication method. HTTP communication is used as the communication method. Finally, the terminal displays the received product details on the user interface. This allows the user to view detailed product descriptions in real time based on the product identification information they entered.
[0925] As a concrete example, consider the case where a user enters the product ID "CAM123". When the user enters the product ID "CAM123" and presses the search button, the terminal sends this information to the server. The server retrieves information about the camera corresponding to "CAM123" from its database (for example, "high-resolution 4K camera, 20 megapixels, with zoom function") and passes this information to a generative machine learning model. Based on this data, the generative machine learning model automatically generates a product description such as, "This camera can shoot high-resolution 4K video and take 20-megapixel photos. It also has a zoom function, so you can take clear pictures of distant subjects." This product description is finally sent to the terminal and displayed on the user's screen.
[0926] As an example of a prompt, the text "Generate a detailed product description for product ID 'CAM123'" is input to the generative machine learning model. A specific example of input to the model would be "High-resolution 4K camera, 20 megapixels, with zoom function."
[0927] This system allows users to check detailed product information in real time by entering their own product identification information. Furthermore, even new crew members can provide quick and accurate product descriptions without needing specialized knowledge. This is expected to improve customer satisfaction and operational efficiency in the sale of home appliances.
[0928] The flow of the specific processing in Example 1 will be explained using Figure 11.
[0929] Step 1:
[0930] The user enters product identification information.
[0931] Specific explanation: The user enters product identification information (e.g., product ID) into an input field on the terminal's user interface and presses the search button. The entered product identification information is, for example, "CAM123".
[0932] Input: Product identification information (product ID)
[0933] Output: Input product identification information
[0934] Step 2:
[0935] The device sends product identification information to the server.
[0936] Detailed explanation: The terminal sends an HTTP request to the server containing the entered product identification information. Internally, the terminal generates a request in the format "GET / search?product_id=CAM123".
[0937] Input: Entered product identification information (product ID)
[0938] Output: HTTP request to the server
[0939] Step 3:
[0940] The server parses the request.
[0941] Specific explanation: The server parses the received HTTP request and extracts product identification information from the query parameters. For example, it extracts the product ID "CAM123" from the request "GET / search?product_id=CAM123".
[0942] Input: HTTP Request
[0943] Output: Analyzed product identification information (product ID)
[0944] Step 4:
[0945] The server searches the database.
[0946] Specific explanation: The server executes an SQL query against the database based on the analyzed product identification information. It retrieves the relevant product information using the query "SELECT FROM products WHERE product_id = 'CAM123'".
[0947] Input: Analyzed product identification information (product ID)
[0948] Output: Acquired product information (e.g., "High-resolution 4K camera, 20 megapixels, with zoom function")
[0949] Step 5:
[0950] The server passes data to the generative machine learning model.
[0951] Detailed explanation: The server passes the acquired product information to a generative machine learning model. The generative machine learning model generates a detailed product description based on the given product information. The model incorporates natural language processing techniques.
[0952] Input: Acquired product information
[0953] Output: Generated product description (Example: "This camera can shoot high-resolution 4K video and take 20-megapixel photos. It also has a zoom function, allowing you to capture distant subjects clearly.")
[0954] Step 6:
[0955] The server sends the generated product description to the terminal.
[0956] Specific explanation: The server generates an HTTP response containing the product description generated by the generative machine learning model and sends it to the terminal. The response includes the description along with "200 OK".
[0957] Input: Generated product description
[0958] Output: HTTP response
[0959] Step 7:
[0960] The device displays the product description.
[0961] Specific explanation: The device parses the received HTTP response and displays the product description on the user interface. The user can view the detailed product description displayed on the device screen in real time.
[0962] Input: HTTP response
[0963] Output: Product description displayed on the user interface
[0964] (Application Example 1)
[0965] Next, we will explain Application Example 1. In the following explanation, the data processing device 12 will be referred to as the "server" and the robot 414 as the "terminal".
[0966] Traditional methods of product description in physical stores rely on the skill level and consistency of staff explanations, making it difficult to provide customers with appropriate and detailed information quickly. Furthermore, even when systems existed that automatically generated product descriptions, they lacked features such as voice assistants or scanning capabilities, failing to adequately consider user convenience. As a result, challenges remain in improving customer satisfaction and reducing staff workload.
[0967] The specific processing performed by the specific processing unit 290 of the data processing device 12 in Application Example 1 is realized by the following means.
[0968] In this invention, the server includes terminal means equipped with a user interface for users to input product identification information; processing means including a generative machine learning model that receives product identification information from the terminal means and generates detailed product information based on the product identification information; communication means for transmitting the detailed product information generated by the generative machine learning model to the terminal means; voice output means equipped with a voice assistant function for providing product descriptions by voice; and scanning means for obtaining product identification information by having customers scan a QR code or barcode and displaying it on the terminal. This makes it possible to provide customers with real-time, consistent, and detailed product descriptions, regardless of the skill level of the staff. In addition, the voice assistant function and scanning function allow customers to easily obtain product information themselves, simultaneously improving customer satisfaction and reducing the burden on staff.
[0969] A "terminal device" refers to a device equipped with an interface for users to input product identification information. Examples include digital signage, tablets, smartphones, and smart glasses.
[0970] A "generative machine learning model" refers to artificial intelligence technology that automatically generates detailed product descriptions based on input product identification information.
[0971] "Communication means" refers to the technology and methods for transmitting detailed product information generated by a generative machine learning model to a terminal device.
[0972] "Voice output means" refers to a function that provides generated product descriptions in audio format. This allows users to obtain product information through audio.
[0973] "Scanning method" refers to technology that allows customers to obtain product identification information by scanning a QR code or barcode and display it on their device.
[0974] "User interface" is a general term for the display and input methods used by users to input product identification information and to review the generated product description.
[0975] The system of this invention consists of terminal means, a server, a generative machine learning model, communication means, voice output means, and scanning means. Specific embodiments of the system using these components will be described in detail below.
[0976] Hardware and software
[0977] Terminal devices: Digital signage, tablets, smartphones, smart glasses, and other devices.
[0978] Server: A cloud server using Amazon Web Services (AWS) or Google Cloud Platform (GCP).
[0979] Use either MySQL or PostgreSQL as the database.
[0980] Generative machine learning models: Generative machine learning models including OpenAI GPT-4 or Google BERT.
[0981] Audio output method: The speaker or earphones built into the device.
[0982] Scanning method: Camera on the device or a dedicated barcode scanner.
[0983] Program Overview
[0984] The server receives product identification information from a terminal device equipped with a user interface for users to input product identification information. For example, the user can obtain product information by entering a product ID on a tablet's input screen or by scanning a QR code or barcode with a smartphone's camera. This allows the product identification information to be acquired through the scanning device.
[0985] The received product identification information is transmitted to the server via a communication method. The server searches the database based on this product identification information and retrieves the corresponding product data. The retrieved product data is passed to a generative machine learning model, which uses natural language processing techniques to generate a detailed product description.
[0986] The generated product description is then transmitted back to the terminal via a communication device. The user can view the detailed product description in real time on the terminal screen. Furthermore, it is possible to listen to the generated product description aloud using an audio output device. This allows users to obtain product information both visually and aurally, improving convenience.
[0987] Specific example
[0988] For example, consider a case where a user enters the product ID "CAM123". When the user enters "CAM123" into a tablet and presses the search button, the device sends the product ID to the server. The server retrieves information about the camera corresponding to "CAM123" from its database (for example, "High-resolution 4K camera, 20 megapixels, with zoom function") and passes this information to a generative machine learning model. The model generates a product description like the following:
[0989] Product ID: CAM123
[0990] This camera can shoot high-resolution 4K video and take 20-megapixel photos. It also features a zoom function, allowing you to capture distant subjects clearly.
[0991] The generated product description is displayed on the device screen and also provided to the user via audio output. In this way, the user can obtain product information through both sight and sound.
[0992] The flow of a specific process in Application Example 1 will be explained using Figure 12.
[0993] Step 1: Enter product identification information
[0994] The user inputs product identification information (e.g., product ID, QR code, barcode) using a terminal device (e.g., smartphone or tablet). At this time, the user interface saves the entered product identification information to the terminal device's memory.
[0995] Input: Product ID, QR code, or barcode
[0996] Output: Product identification information
[0997] Step 2: Submit product identification information
[0998] The terminal device transmits the product identification information obtained in step 1 to the server using the communication device. Specifically, an HTTP request is generated and sent to the server's endpoint.
[0999] Input: Product identification information
[1000] Output: HTTP request sent to the server
[1001] Step 3: Search the database
[1002] The server queries the database based on the received product identification information to retrieve the corresponding product data. It executes an SQL query to fetch the relevant product data.
[1003] Input: Product identification information
[1004] Output: Product data
[1005] Step 4: Explanation generation using generative machine learning models
[1006] The server inputs the product data obtained in step 3 into a generative machine learning model to generate detailed product descriptions. The model uses natural language processing techniques to convert the product data into easy-to-understand text.
[1007] Input: Product data
[1008] Output: Product Description
[1009] Step 5: Send the generated product description.
[1010] The server sends the generated product description to the terminal device via a communication means. An HTTP response is generated and sent to the endpoint of the terminal device.
[1011] Input: Product Description
[1012] Output: HTTP response sent to the terminal device
[1013] Step 6: Display and audio output of product description
[1014] The terminal device displays the received product description on the user interface and also provides the product description in audio using the audio output device.
[1015] Input: Product Description
[1016] Output: Screen display and audio output
[1017] As a concrete example, consider the case where a user enters the product ID "CAM123". When the user enters the product ID and presses the search button, the terminal sends the product ID to the server. The server retrieves information about the camera corresponding to "CAM123" from its database (for example, "High-resolution 4K camera, 20 megapixels, with zoom function") and passes this information to a generative machine learning model. The model generates a product description like the following:
[1018] Product ID: CAM123
[1019] This camera can shoot high-resolution 4K video and take 20-megapixel photos. It also features a zoom function, allowing you to capture distant subjects clearly.
[1020] The generated product description is displayed on the device screen and also provided to the user via audio output. In this way, the user can obtain product information through both sight and sound.
[1021] Furthermore, an emotion engine that estimates the user's emotions may be incorporated. That is, the identification processing unit 290 may use the emotion identification model 59 to estimate the user's emotions and perform identification processing using the user's emotions.
[1022] The embodiments for carrying out the present invention are described in detail below. This system consists of a terminal, a server, a generative machine learning model, an emotion engine, and a communication means. The terminal, equipped with a user interface, receives input from the user, and the server is responsible for generating product descriptions.
[1023] First, the user enters product identification information (e.g., product ID) using a device (e.g., tablet, smartphone, digital signage). After entering this information, the user requests product information from the server by pressing the search button.
[1024] The device acquires emotional data from product identification information received from the user and from an emotion engine that recognizes emotions from the user's facial expressions and voice. The emotion engine uses facial recognition technology and voice recognition technology to evaluate the user's emotions in real time.
[1025] The device sends product identification information and sentiment data to the server. The server analyzes the received request, accesses the database using the product identification information, and retrieves the corresponding product information. This information, along with the sentiment data, is then passed to a generative machine learning model.
[1026] Generative machine learning models can generate product descriptions adapted to the user's emotional state based on the provided product information and sentiment data. For example, if the user shows an interested expression, it will provide a more detailed and technical explanation. On the other hand, if the user shows an anxious expression, it will generate a simple and reassuring explanation.
[1027] The generated product description is transmitted from the server to the terminal using a communication method. The terminal displays the received product description on its user interface. This allows the user to view a detailed product description in real time based on the product identification information they entered.
[1028] As a concrete example, consider the case where a user enters the product ID "CAM123". When the user enters the product ID and presses the search button, the terminal sends the product ID along with emotional data acquired from the user's facial expressions and voice to the server. The server retrieves information about the camera corresponding to "CAM123" from the database (for example, "high-resolution 4K camera, 20 megapixels, with zoom function") and passes this information and emotional data to a generative machine learning model. The model generates an appropriate product description according to the user's emotional state, adjusting content such as, "This camera can shoot high-resolution 4K video and take 20-megapixel photos. It also has a zoom function, so you can take clear pictures of distant subjects." For example, if the user seems interested, detailed technical explanations will be added, and if they are feeling anxious, more approachable terminology will be used.
[1029] This system allows even new crew members, without requiring extensive product knowledge, to provide detailed and accurate product descriptions tailored to the user's emotional state. Furthermore, users can easily learn about product details in an emotionally sensitive way by inputting product information themselves. Therefore, this invention has the effect of improving customer satisfaction in the sale of home appliances and enhancing the efficiency of crew members' work.
[1030] The following describes the processing flow.
[1031] Step 1:
[1032] The user enters product identification information (e.g., product ID "CAM123") using a device (e.g., tablet or smartphone). After entering this information, the user taps the "Search" button on the screen.
[1033] Step 2:
[1034] The terminal retrieves the product identification information entered by the user. This information is stored in a variable and used for the next process.
[1035] Step 3:
[1036] The device analyzes the user's facial expressions and voice using an emotion engine to obtain the user's emotional data (e.g., "interest," "anxiety," "confusion," etc.). This data is also stored in variables.
[1037] Step 4:
[1038] The device sends an HTTP request to the server based on the acquired product identification information and sentiment data. The request is sent in the format of a POST request, for example, containing the product identification information and sentiment data.
[1039] Step 5:
[1040] The server analyzes the request received from the terminal and extracts product identification information and sentiment data. The extracted data is then stored in variables.
[1041] Step 6:
[1042] The server uses product identification information to query the product database. The query is configured to retrieve information about products that match the product identification information.
[1043] Step 7:
[1044] The server retrieves relevant product information from the database (for example, "high-resolution 4K camera, 20 megapixels, with zoom function"). The retrieved product information and sentiment data are then passed to a generative machine learning model.
[1045] Step 8:
[1046] Generative machine learning models generate detailed product descriptions based on the product information and sentiment data they receive. For example, the model adjusts a description such as "This camera can shoot high-resolution 4K video and take 20-megapixel photos. It also has a zoom function, allowing you to capture distant subjects clearly" according to the user's emotional state.
[1047] Step 9:
[1048] The server creates an HTTP response containing the generated product description and sends it to the terminal. The response format is, for example, JSON.
[1049] Step 10:
[1050] The terminal analyzes the response received from the server and extracts the product description text. The extracted text is then placed on a specific element of the user interface (for example, HTML). Display in the tags.
[1051] Step 11:
[1052] Users can view detailed product descriptions displayed on their device screens in real time. This allows users to obtain sufficient information about the product, making it easier for them to make a purchase decision.
[1053] These processing steps enable users to easily receive detailed product descriptions, creating a system that allows even new crew members to confidently interact with customers. Furthermore, by utilizing emotional data, product descriptions optimized for the user's emotional state are provided.
[1054] (Example 2)
[1055] Next, we will describe Example 2. In the following description, the data processing device 12 will be referred to as the "server" and the robot 414 as the "terminal".
[1056] In recent years, providing users with appropriate product information has become increasingly important in e-commerce and brick-and-mortar retail. However, many systems provide uniform information without considering user emotions, resulting in insufficient user experience and customer satisfaction. Furthermore, a shortage of staff and sales personnel who understand and can provide detailed technical information about products is also a problem. Therefore, there is a need for a system that provides effective product descriptions in real time, tailored to the user's emotional state.
[1057] The identification processing performed by the identification processing unit 290 of the data processing device 12 in Example 2 is realized by the following means. In this invention, the server includes terminal means equipped with a user interface for the user to input product identification information and emotion data; processing means including a generative machine learning model that receives product identification information and emotion data from the terminal means and generates product information based on the product identification information and detailed product information adapted to the user's emotional state; and communication means that transmits the detailed product information generated by the generative machine learning model to the terminal means. This makes it possible to provide detailed and appropriate product descriptions in real time that correspond to the user's emotions.
[1058] "Terminal means" refers to devices used by users to input product identification information and emotional data, and includes tablets, smartphones, and digital signage.
[1059] "Product identification information" refers to information used to identify a product, and includes, for example, product IDs and barcodes.
[1060] "Emotional data" refers to data that indicates the emotional state of a user, analyzed from their facial expressions and voice.
[1061] "User interface" is a general term for the screens and input devices that users use to operate a device, and includes touch panels and button interfaces.
[1062] A "generative machine learning model" is a model that uses machine learning techniques to generate appropriate product descriptions based on product information and sentiment data.
[1063] "Processing means" refers to a device or system that processes information received from a terminal means, and includes a generative machine learning model.
[1064] "Communication means" refers to technologies and devices for sending and receiving data between a server and a terminal, and includes the internet and local networks.
[1065] A "database" is a system or device that manages and stores data such as product information.
[1066] An "emotion engine" is a system or device that analyzes a user's facial expressions and voice to generate emotional data.
[1067] This invention is a system for improving the user experience when acquiring product information. This system consists of a terminal, an emotion engine, a server, a generative machine learning model, and communication means. The following describes specific implementations of this system.
[1068] System Configuration
[1069] The terminal device is a device equipped with a user interface, used by users to input product identification information and emotional data. Specifically, this includes tablets, smartphones, and digital signage. The terminal device uses a built-in camera and microphone to analyze the user's facial expressions and voice, and generates emotional data using an emotion engine.
[1070] The emotion engine uses facial recognition and speech recognition technologies to analyze emotional data from the user's facial expressions and voice in real time. This allows for an accurate understanding of the user's current emotional state.
[1071] The server receives product identification information and sentiment data transmitted from the terminal device. Furthermore, the server accesses a predefined database to retrieve the relevant product information. The retrieved product information and sentiment data are then passed to a generative machine learning model.
[1072] Generative machine learning models generate detailed product information tailored to the user's emotional state, based on product information and sentiment data. For example, if the user shows an interested expression, it provides a more detailed and technical explanation. On the other hand, if the user shows an anxious expression, it generates a simple and reassuring explanation.
[1073] The communication means plays the role of transmitting detailed information about the generated product from the server to the terminal means.
[1074] Specific example
[1075] For example, consider a case where a user enters the product ID "CAM123". When the user enters the product ID and presses the search button, the terminal sends the product ID along with emotional data acquired from the user's facial expressions and voice to the server. The server retrieves product information for the camera corresponding to "CAM123" from the database (for example, "High-resolution 4K camera, 20 megapixels, with zoom function") and passes this information and emotional data to a generative machine learning model. The model generates an appropriate product description according to the user's emotional state, adjusting content such as, "This camera can shoot high-resolution 4K video and take 20-megapixel photos. It also has a zoom function, so you can take clear pictures of distant subjects." For example, if the user seems interested, detailed technical explanations will be added, and if they are anxious, more approachable terminology will be used.
[1076] Example of a prompt
[1077] Create a description of the high-resolution 4K camera that matches the user's expression of interest.
[1078] Please create a description of the 20-megapixel camera that is tailored to the user's anxious expression.
[1079] This system allows even new crew members to provide detailed and accurate product descriptions tailored to the user's emotional state, even without extensive product knowledge. Furthermore, users can easily learn about product details in an emotionally sensitive way by inputting product information themselves. Therefore, it improves customer satisfaction in home appliance sales and enhances the efficiency of crew members' work.
[1080] The flow of the specific processing in Example 2 will be explained using Figure 13.
[1081] Step 1:
[1082] The user enters product identification information (e.g., product ID "CAM123") using the terminal. When the user presses the search button, this product ID is saved as input data on the terminal. The terminal's user interface receives this product identification information, preparing the data to be used in the next step.
[1083] Step 2:
[1084] The device uses its built-in camera and microphone to capture the user's facial expressions and voice. The device sends this data to an emotion engine, which analyzes the user's emotional data in real time. Specifically, the device captures the user's face with its camera and acquires emotional data such as "interested facial expressions" and "cheerful voice tones."
[1085] Step 3:
[1086] The terminal sends product identification information (e.g., "CAM123") and acquired emotion data to the server. By receiving product identification information and emotion data as input data, the server analyzes this data and prepares it for use in the next step.
[1087] Step 4:
[1088] The server accesses the database based on the product identification information and retrieves the corresponding product information. Specifically, the server executes a database query to retrieve product information for cameras corresponding to "CAM123" (e.g., "High-resolution 4K camera, 20 megapixels, with zoom function"). The input data is the product identification information, and the output data is the product information.
[1089] Step 5:
[1090] The server passes the acquired product information and sentiment data to a generative machine learning model. The generative machine learning model receives the product information and sentiment data as input and generates detailed product information adapted to the user's emotional state. Specifically, the generative machine learning model considers that "the user appears interested" and generates a product description that includes detailed technical explanations. The output data is a sentiment-based, adjusted product description.
[1091] Step 6:
[1092] The generated product description is transmitted from the server to the terminal via a communication method. Specifically, the server generates the product description and sends it to the terminal as a data packet over the internet. The input data is the generated product description, and the output data is the result of the transmission to the terminal.
[1093] Step 7:
[1094] The device displays the received product description on its user interface. Users can view detailed product descriptions in real time on the device's screen. Specifically, the device's touchscreen displays a text message such as, "This camera can shoot high-resolution 4K video and 20-megapixel photos. It also has a zoom function, allowing you to take clear photos of distant subjects." The output data is the displayed product description.
[1095] (Application Example 2)
[1096] Next, we will explain application example 2. In the following explanation, the data processing device 12 will be referred to as the "server" and the robot 414 as the "terminal".
[1097] When customers choose products in physical stores, they may feel anxious because they lack detailed information about the products. Furthermore, the uniform nature of product descriptions makes it difficult to provide personalized information tailored to individual customers' emotions and needs. This situation can lead to decreased customer satisfaction and reduced purchasing intent. To address this, there is a need for a system that analyzes customer emotions from their facial expressions and voice, and generates and displays product descriptions adapted to those emotions in real time.
[1098] The specific processing performed by the specific processing unit 290 of the data processing device 12 in Application Example 2 is realized by the following means.
[1099] In this invention, the server includes emotion recognition means for analyzing the user's emotions from their facial expressions and voice, processing means including a generative machine learning model that receives product identification information and emotion data and generates a detailed product description, and communication means for transmitting the generated product details to a terminal means. This makes it possible to provide personalized product descriptions that are adapted to the customer's emotions.
[1100] A "terminal device" is a device equipped with an interface for the user to input product identification information.
[1101] "Processing means" refers to a device that includes a generative machine learning model that generates detailed product information based on product identification information and sentiment data.
[1102] "Communication means" refers to a device that transmits detailed product information generated by a generative machine learning model to a terminal means.
[1103] "Emotion recognition means" refers to technologies and devices used to analyze emotions from a customer's facial expressions and voice.
[1104] "Product identification information" refers to information that a user enters in order to identify a specific product.
[1105] A "generative machine learning model" is a machine learning algorithm that generates detailed product information based on product identification information and sentiment data.
[1106] A "database" is a collection of information where product information is stored in a predefined format.
[1107] A "user interface" is a means of interaction that allows a user to input information into a system.
[1108] "Emotional data" refers to numerical and categorical data of emotional states analyzed from a user's facial expressions and voice.
[1109] "Product description" refers to a detailed description of a product generated by a generative machine learning model.
[1110] The embodiments for carrying out the present invention are described in detail below. This system consists of a terminal, a server, a generative machine learning model, an emotion recognition means, and a communication means. The terminal, equipped with a user interface, receives input from the user, and the server is responsible for generating product descriptions.
[1111] The process begins when the user enters product identification information using a terminal device (e.g., smartphone, tablet, or digital signage). Once the user enters the product identification information and presses the search button, the terminal requests this information from the server. At the same time, the user's facial expressions and voice are also captured, and this data is analyzed in real time by emotion recognition technology.
[1112] The emotion recognition means acquires user emotion data using facial recognition technology and speech recognition technology. This emotion data is numerical data representing the user's emotional state, such as interest, doubt, and anxiety. The terminal means transmits this emotion data to the server along with product identification information.
[1113] The server parses the received request and retrieves the corresponding product information from the database based on the product identification information. It then passes the product information, along with sentiment data, to a generative machine learning model. The generative machine learning model uses the product information and sentiment data to generate a product description adapted to the user's emotional state. For example, if the user shows an interested expression, it provides a more detailed and technical explanation; if they show an anxious expression, it generates a simple and friendly explanation.
[1114] The generated product description is transmitted from the server to the terminal using a communication method and displayed on the terminal's user interface. This allows the user to view a detailed product description based on the entered product identification information in real time.
[1115] The hardware required to implement this system includes terminal devices such as smartphones and tablets, as well as servers equipped with GPUs. The software used includes OpenCV (for image processing), TensorFlow (for loading and using emotion recognition models), Transformers (for loading and using GPT-3 models), and requests (for API communication).
[1116] Specific example:
[1117] For example, if a user enters the product ID "CAM123" and looks interested, the generated description will be: "This camera can shoot high-resolution 4K video and take 20-megapixel photos. It also has a zoom function, allowing you to take clear photos of distant subjects." In this case, an example of the prompt message would be as follows:
[1118] Example of a prompt:
[1119] Product: 4K Video Camera
[1120] Description: A high-resolution 4K camera that can capture 20-megapixel photos and features zoom functionality.
[1121] Emotion: Curious
[1122] This format allows generative machine learning models to provide detailed product descriptions that are adapted to the user's emotional state.
[1123] The flow of a specific process in Application Example 2 will be explained using Figure 14.
[1124] Step 1:
[1125] The device acquires product identification information and sentiment data.
[1126] Users input product identification information using a terminal device (smartphone, tablet, digital signage, etc.). The system also captures the user's facial expressions and voice through the terminal's camera and microphone. This results in the acquisition of both product identification information and emotion data. This emotion data is analyzed in real time using image processing tools (e.g., OpenCV) and voice analysis tools.
[1127] Step 2:
[1128] Analysis and integration of emotional data
[1129] The device passes the acquired facial expression data and audio data to an emotion recognition system (an emotion recognition model using TensorFlow), which converts the user's emotional state into numerical data. For example, if the user has an expression that suggests interest, the category value "interesting" is obtained as emotion data. In this way, emotion data is generated by analyzing the input data (facial expression data and audio data).
[1130] Step 3:
[1131] Sending product identification information and emotional data to the server
[1132] The terminal sends product identification information and analyzed sentiment data to the server. This means that the product identification information and sentiment data are passed to the server as input from the terminal.
[1133] Step 4:
[1134] Product information acquisition
[1135] The server retrieves the corresponding product information from a predefined database based on the received product identification information. For example, it retrieves product information for product ID "CAM123" (e.g., high-resolution 4K camera, 20 megapixels, with zoom function). This allows the server to obtain product information from the input data (product identification information).
[1136] Step 5:
[1137] Product description generation using generative machine learning models
[1138] The server passes the acquired product information and sentiment data to a generative machine learning model (for example, a generative AI model using TensorFlow) to create a prompt. Based on this prompt, the generative machine learning model generates a product description adapted to the user's emotional state. For example, it might generate a description such as, "This camera can shoot high-resolution 4K video and take 20-megapixel photos. It also has a zoom function, allowing you to take clear photos of distant subjects."
[1139] Step 6:
[1140] Sending and displaying product descriptions to the device
[1141] Product descriptions generated by generative machine learning models are transmitted from the server to the terminal via communication. The terminal displays the received product descriptions on its user interface. This ensures that the output is presented in a format that the user can verify, based on the input (generated product descriptions).
[1142] Through these steps, users can receive detailed product descriptions tailored to their emotional state in real time.
[1143] The specific processing unit 290 transmits the result of the specific processing to the robot 414. In the robot 414, the control unit 46A causes the speaker 240 and the controlled object 443 to output the result of the specific processing. The microphone 238 acquires audio indicating user input for the result of the specific processing. The control unit 46A transmits the audio data indicating user input acquired by the microphone 238 to the data processing unit 12. In the data processing unit 12, the specific processing unit 290 acquires the audio data.
[1144] Data generation model 58 is a type of so-called generative AI (Artificial Intelligence). One example of data generation model 58 is ChatGPT (Internet search<URL: https: / / openai.com / blog / chatgpt> ), Gemini (Internet search) <url: https: gemini.google.com ?hl="ja">Examples of generative AI include the following. The data generation model 58 is obtained by performing deep learning on a neural network. The data generation model 58 is input with prompts containing instructions, and with inference data such as audio data representing speech, text data representing text, and image data representing images. The data generation model 58 infers from the input inference data according to the instructions indicated by the prompts, and outputs the inference results in data formats such as audio data and text data. Here, inference refers to, for example, analysis, classification, prediction, and / or summarization.
[1145] In the above embodiment, an example was given in which specific processing is performed by the data processing device 12, but the technology of this disclosure is not limited thereto, and the specific processing may also be performed by the robot 414.
[1146] Furthermore, the emotion identification model 59, acting as an emotion engine, may determine the user's emotion according to a specific mapping. Specifically, the emotion identification model 59 may determine the user's emotion according to a specific mapping, which is an emotion map (see Figure 9). Similarly, the emotion identification model 59 may also determine the robot's emotion, and the identification processing unit 290 may perform identification processing using the robot's emotion.
[1147] Figure 9 shows an emotion map 400 in which multiple emotions are mapped. In the emotion map 400, emotions are arranged in concentric circles radiating from the center. The closer to the center of the concentric circles, the more primitive the emotions are located. Further out of the concentric circles, emotions representing states and actions arising from mental states are located. Emotion is a concept that includes feelings and mental states. On the left side of the concentric circles, emotions that are generally generated from reactions occurring in the brain are located. On the right side of the concentric circles, emotions that are generally induced by situational judgment are located. Above and below the concentric circles, emotions that are generally generated from reactions occurring in the brain and induced by situational judgment are located. In addition, the emotion of "pleasure" is located on the upper side of the concentric circles, and the emotion of "displeasure" is located on the lower side. Thus, in the emotion map 400, multiple emotions are mapped based on the structure in which emotions arise, and emotions that are likely to occur simultaneously are mapped close together.
[1148] These emotions are distributed at the 3 o'clock position on the Emotion Map 400, and usually fluctuate between feelings of security and anxiety. In the right half of the Emotion Map 400, situational awareness takes precedence over internal feelings, resulting in a calm impression.
[1149] The inside of the Emotion Map 400 represents inner thoughts, while the outside represents actions. Therefore, the further you go from the outside of the Emotion Map 400, the more visible (expressed in actions) your emotions become.
[1150] Here, human emotions are based on various balances, such as posture and blood sugar levels. When these balances deviate from the ideal, it results in discomfort, and when they approach the ideal, it results in pleasure. Similarly, in robots, cars, motorcycles, etc., emotions can be created based on various balances, such as posture and battery level. When these balances deviate from the ideal, it results in discomfort, and when they approach the ideal, it results in pleasure. The emotion map can be generated, for example, based on Dr. Mitsuyoshi's emotion map (Research on a system for analyzing brain physiological signals of speech emotion recognition and emotion, Tokushima University, doctoral dissertation: https: / / ci.nii.ac.jp / naid / 500000375379). The left half of the emotion map contains emotions belonging to a region called "response," where sensation is dominant. The right half of the emotion map contains emotions belonging to a region called "situation," where situational awareness is dominant.
[1151] The emotion map defines two emotions that promote learning. One is the emotion around the middle of the negative "repentance" and "reflection" on the situation side. In other words, it is when the robot experiences negative emotions such as "I never want to feel this way again" or "I don't want to be scolded again." The other is the emotion around the positive "desire" on the reaction side. In other words, it is when the robot has positive feelings such as "I want more" or "I want to know more."
[1152] The emotion identification model 59 inputs user input into a pre-trained neural network, obtains emotion values representing each emotion shown in the emotion map 400, and determines the user's emotion. This neural network is pre-trained based on multiple training data sets, which are combinations of user input and emotion values representing each emotion shown in the emotion map 400. Furthermore, this neural network is trained so that emotions located close together have similar values, as shown in the emotion map 900 in Figure 10. Figure 10 shows an example where multiple emotions such as "reassured," "calm," and "confident" have similar emotion values.
[1153] The above description primarily focuses on the functions of the data processing device 12 in relation to this disclosure. However, the system related to this disclosure is not necessarily implemented on a server. The system related to this disclosure may be implemented as a general information processing system. This disclosure may be implemented, for example, as a software program that runs on a personal computer or as an application that runs on a smartphone. The method related to this disclosure may be provided to users in SaaS (Software as a Service) format.
[1154] In the above embodiment, an example was given in which a specific process is performed by a single computer 22. However, the technology of this disclosure is not limited thereto, and a distributed processing of the specific process may be performed by multiple computers, including computer 22. For example, a data generation model 58 may be provided in an external device of the data processing device 12, and the external device may generate data according to the input data.
[1155] In the above embodiment, an example was given in which the specific processing program 56 is stored in the storage 32, but the technology of this disclosure is not limited thereto. For example, the specific processing program 56 may be stored in a portable, computer-readable, non-temporary storage medium such as a USB (Universal Serial Bus) memory. The specific processing program 56 stored in the non-temporary storage medium is installed in the computer 22 of the data processing device 12. The processor 28 executes specific processing according to the specific processing program 56.
[1156] Alternatively, the specific processing program 56 may be stored in a storage device such as a server connected to the data processing device 12 via the network 54, and the specific processing program 56 may be downloaded and installed on the computer 22 in response to a request from the data processing device 12.
[1157] Furthermore, it is not necessary to store the entirety of the specific processing program 56 in a storage device such as a server connected to the data processing device 12 via the network 54, or to store the entirety of the specific processing program 56 in the storage 32; it is acceptable to store only a portion of the specific processing program 56.
[1158] The following types of processors can be used as hardware resources to perform specific processing. Examples of processors include a CPU, a general-purpose processor that functions as a hardware resource to perform specific processing by executing software, i.e., a program. Other examples of processors include dedicated electrical circuits, such as FPGAs (Field-Programmable Gate Arrays), PLDs (Programmable Logic Devices), or ASICs (Application Specific Integrated Circuits), which have circuit configurations specifically designed to perform specific processing. All of these processors have built-in or connected memory, and all of them perform specific processing by using memory.
[1159] The hardware resource that performs a specific process may consist of one of these various processors, or it may consist of a combination of two or more processors of the same or different types (for example, a combination of multiple FPGAs, or a combination of a CPU and an FPGA). Alternatively, the hardware resource that performs a specific process may consist of a single processor.
[1160] Examples of configurations using a single processor include, firstly, a configuration in which one or more CPUs and software are combined to form a single processor, and this processor functions as a hardware resource that performs a specific process. Secondly, there is a configuration using a processor that realizes the functions of the entire system, including multiple hardware resources that perform a specific process, on a single IC chip, as exemplified by SoCs (System-on-a-chip). In this way, a specific process is realized using one or more of the above types of processors as hardware resources.
[1161] Furthermore, the hardware structure of these various processors can more specifically utilize electrical circuits that combine circuit elements such as semiconductor devices. Also, the specific processing described above is merely an example. Therefore, it goes without saying that unnecessary steps can be deleted, new steps added, or the processing order rearranged, as long as it does not deviate from the main purpose.
[1162] The descriptions and illustrations presented above are detailed explanations of the technical aspects of this disclosure and are merely examples of the technical aspects. For example, the above descriptions of the structure, function, operation, and effect are examples of the structure, function, operation, and effect of the technical aspects of this disclosure. Therefore, it goes without saying that you may delete unnecessary parts, add new elements, or replace elements in the descriptions and illustrations presented above, as long as you do not deviate from the essence of the technical aspects of this disclosure. Furthermore, in order to avoid confusion and facilitate understanding of the technical aspects of this disclosure, explanations of common technical knowledge and the like that do not require special explanation to enable the implementation of the technical aspects of this disclosure have been omitted from the descriptions and illustrations presented above.
[1163] All documents, patent applications, and technical standards described herein are incorporated by reference to the same extent as if each individual document, patent application, and technical standard were specifically and individually noted to be incorporated by reference.
[1164] The following is further disclosed regarding the embodiments described above.
[1165] (Claim 1)
[1166] A terminal device equipped with a user interface for the user to input product identification information,
[1167] Processing means including a generative machine learning model that receives product identification information from the terminal means and generates detailed product information based on said product identification information,
[1168] A communication means for transmitting detailed information of products generated by the generative machine learning model to the terminal means,
[1169] A system that includes this.
[1170] (Claim 2)
[1171] The system according to claim 1, wherein the processing means further includes means for obtaining detailed information of a product from a predefined database based on the product identification information.
[1172] (Claim 3)
[1173] The system according to claim 1, wherein the user interface is one of a digital signage, a tablet, or a smartphone.
[1174] "Example 1"
[1175] (Claim 1)
[1176] A terminal device equipped with a user interface for the user to input product identification information,
[1177] A server means that receives product identification information from the terminal means, analyzes the product identification information, and obtains product information from a database,
[1178] Processing means including a generative machine learning model that generates a detailed product description using product information acquired by the server means,
[1179] A communication means for transmitting detailed information of products generated by the generative machine learning model to the terminal means,
[1180] A system that includes this.
[1181] (Claim 2)
[1182] The system according to claim 1, wherein the processing means further includes means for obtaining detailed information of a product from a predefined database based on the product identification information.
[1183] (Claim 3)
[1184] The system according to claim 1, wherein the user interface is one of an electronic display device, a portable information terminal, and a smart device.
[1185] "Application Example 1"
[1186] (Claim 1)
[1187] A terminal device equipped with a user interface for the user to input product identification information,
[1188] Processing means including a generative machine learning model that receives product identification information from the terminal means and generates detailed product information based on said product identification information,
[1189] A communication means for transmitting detailed information of products generated by the generative machine learning model to the terminal means,
[1190] Equipped with a voice assistant function, it provides a voice output means for delivering product descriptions by voice,
[1191] A scanning means for obtaining product identification information by having a customer scan a QR code or barcode and displaying it on a terminal,
[1192] A system that includes this.
[1193] (Claim 2)
[1194] The system according to claim 1, wherein the processing means further includes means for obtaining detailed information of a product from a predefined database based on the product identification information.
[1195] (Claim 3)
[1196] The system according to claim 1, wherein the user interface is one of a digital signage, a tablet, a smartphone, and smart glasses.
[1197] "Example 2 of combining an emotion engine"
[1198] (Claim 1)
[1199] A terminal device equipped with a user interface for the user to input product identification information and sentiment data,
[1200] Processing means including a generative machine learning model that receives product identification information and emotion data from the terminal means and generates product information based on the product identification information and detailed product information that is adapted to the user's emotional state,
[1201] A communication means for transmitting detailed information of products generated by the generative machine learning model to the terminal means,
[1202] A system that includes this.
[1203] (Claim 2)
[1204] The system according to claim 1, wherein the processing means further includes means for obtaining product information from a predefined database based on the product identification information.
[1205] (Claim 3)
[1206] The system according to claim 1, wherein the processing means includes an emotion engine that analyzes the user's facial expressions and voice to acquire emotion data.
[1207] (Claim 4)
[1208] The system according to claim 1, wherein the user interface is one of a digital signage, a tablet, or a smartphone.
[1209] "Application example 2 when combining with an emotional engine"
[1210] (Claim 1)
[1211] A terminal device equipped with a user interface for the user to input product identification information,
[1212] Processing means including a generative machine learning model that receives product identification information and sentiment data from the terminal means and generates detailed product information based on the product identification information and sentiment data,
[1213] A communication means for transmitting detailed information of products generated by the generative machine learning model to the terminal means,
[1214] An emotion recognition method that analyzes emotions from the customer's facial expressions and voice,
[1215] The means for generating a product description adapted to the emotion based on the emotion recognition means and product identification information,
[1216] A system that includes this.
[1217] (Claim 2)
[1218] The system according to claim 1, wherein the processing means further includes means for obtaining detailed information of a product from a predefined database based on the product identification information.
[1219] (Claim 3)
[1220] The system according to claim 1, wherein the user interface is one of the user interfaces of an electronic device. [Explanation of symbols]
[1221] 10, 210, 310, 410 Data Processing Systems 12 Data Processing Devices 14 Smart Devices 214 Smart Glasses 314 Headset-type terminal 414 Robots< / url:> < / url:> < / url:> < / url:>
Claims
1. A terminal device equipped with a user interface for the user to input product identification information, Processing means including a generative machine learning model that receives product identification information from the terminal means and generates detailed product information based on said product identification information, A communication means for transmitting detailed information of products generated by the generative machine learning model to the terminal means, A system that includes this.
2. The system according to claim 1, wherein the processing means further includes means for obtaining detailed information of a product from a predefined database based on the product identification information.
3. The system according to claim 1, wherein the user interface is one of a digital signage, a tablet, or a smartphone.
Citation Information
Patent Citations
Persona chatbot control method and system
JP2022180282A