System
A system that analyzes product information and sales interactions helps elderly and children make informed shopping decisions, reducing waste and expenses by providing real-time notifications on product quality and sales pressure.
Patent Information
- Application Number
- JP2024125377
- Authority / Receiving Office
- JP · JP
- Patent Type
- Applications
- Current Assignee / Owner
- Filing Date
- 2024-07-31
- Publication Date
- 2026-02-13
AI Technical Summary
Elderly people and children face challenges in shopping appropriately, often purchasing products close to their expiration date or being pressured by salespeople, leading to food waste and unnecessary expenses, particularly for those with mobility issues.
A system that captures product information using a user terminal, analyzes image and audio data to evaluate freshness, price, and expiration date, and detects inappropriate sales tactics, sending real-time notifications to users.
Enables elderly and children to shop with peace of mind by accurately assessing product quality and avoiding unnecessary purchases and sales pressure.
Smart Images

Figure 2026023442000001_ABST
Abstract
Description
[Technical Field]
[0001] The technology of the present disclosure relates to a system. [Background technology]
[0002] Patent document 1 discloses a persona chatbot control method performed by at least one processor, the method including the steps of receiving a user utterance, adding the user utterance to a prompt including an instruction sentence related to a description of the chatbot character, encoding the prompt, and inputting the encoded prompt into a language model to generate a chatbot utterance in response to the user utterance. [Prior art documents] [Patent documents]
[0003] [Patent Document 1] Japanese Patent Publication No. 2022-180282 Summary of the Invention [Problem to be solved by the invention]
[0004] In modern society, there is a lack of support for elderly people and children to shop appropriately. Furthermore, these users are likely to mistakenly purchase products that are close to their expiration date or be pressured by salespeople. This leads to food waste and unnecessary expenses. This problem is particularly prevalent and serious for elderly people with mobility issues and children with limited mobility. Therefore, there is a need for a system that allows elderly people and children to shop with peace of mind by selecting appropriate products. [Means for solving the problem]
[0005] The present invention is a system that includes a means for capturing product information using a user terminal, a means for transmitting image data and audio data of the captured product information to a server, a means for analyzing the received image data in the server to evaluate the product's quality, price, and expiration date, a means for analyzing the received audio data in the server to detect whether a salesperson is trying to pressure the user, and a means for sending a notification to the user based on the analysis results. This allows the user to check in real time the freshness, reasonable price, expiration date, etc. of the product they are about to purchase. It can also detect inappropriate solicitations from salespersons and prompt appropriate action. This makes it possible to provide an environment where the elderly and children can shop more safely.
[0006] A "user terminal" is a communication device that is equipped with a camera and a microphone and that can be carried by a user.
[0007] "Product information" refers to information about the label, price, expiration date, freshness, etc. of the product to be purchased.
[0008] "Image data" is digital data containing visual information of a product captured by a camera.
[0009] "Voice data" is digital data that includes the speech of a store clerk or user captured by a microphone.
[0010] "Server" means a computer system that receives and analyzes data sent from a user terminal and sends the analysis results to the user terminal.
[0011] "Analysis" refers to data processing to determine the quality, price, expiration date, and whether or not there was any hard sell of the product from the received image data and audio data.
[0012] "Notification" is a message that notifies the user of the analysis results. [Brief explanation of the drawings]
[0013] [Figure 1] 1 is a conceptual diagram showing an example of the configuration of a data processing system according to a first embodiment. [Figure 2] 1 is a conceptual diagram showing an example of main functions of a data processing device and a smart device according to a first embodiment. [Figure 3] FIG. 10 is a conceptual diagram showing an example of the configuration of a data processing system according to a second embodiment. [Figure 4] FIG. 10 is a conceptual diagram showing an example of main functions of a data processing device and smart glasses according to a second embodiment. [Figure 5] FIG. 10 is a conceptual diagram showing an example of the configuration of a data processing system according to a third embodiment. [Figure 6] FIG. 11 is a conceptual diagram showing an example of main functions of a data processing device and a headset-type terminal according to a third embodiment. [Figure 7] FIG. 10 is a conceptual diagram showing an example of the configuration of a data processing system according to a fourth embodiment. [Figure 8] FIG. 10 is a conceptual diagram showing an example of main functions of a data processing device and a robot according to a fourth embodiment. [Figure 9] 1 shows an emotion map onto which multiple emotions are mapped. [Figure 10] 1 shows an emotion map onto which multiple emotions are mapped. [Figure 11] FIG. 3 is a sequence diagram showing a processing flow of the data processing system according to the first embodiment. [Figure 12] FIG. 10 is a sequence diagram showing the flow of processing in the data processing system in Application Example 1. [Figure 13] FIG. 10 is a sequence diagram showing the flow of processing in the data processing system according to the second embodiment when an emotion engine is combined. [Figure 14] FIG. 10 is a sequence diagram showing the flow of processing in the data processing system in Application Example 2 when an emotion engine is combined. DETAILED DESCRIPTION OF THE INVENTION
[0014] An example of an embodiment of a system according to the technology of the present disclosure will be described below with reference to the accompanying drawings.
[0015] First, the terms used in the following description will be explained.
[0016] In the following embodiments, a coded processor (hereinafter simply referred to as a "processor") may be a single arithmetic device or a combination of multiple arithmetic devices. Furthermore, a processor may be a single type of arithmetic device or a combination of multiple types of arithmetic devices. Examples of arithmetic devices include a CPU (Central Processing Unit), a GPU (Graphics Processing Unit), a GPGPU (General-Purpose computing on Graphics Processing Units), and an APU (Accelerated Processing Unit).
[0017] In the following embodiments, a coded RAM (Random Access Memory) is a memory in which information is temporarily stored and is used as a working memory by a processor.
[0018] In the following embodiments, the coded storage is one or more non-volatile storage devices that store various programs, various parameters, etc. Examples of non-volatile storage devices include flash memory (SSD (Solid State Drive)), magnetic disks (e.g., hard disks), and magnetic tapes.
[0019] In the following embodiments, a communication I / F (Interface) with a symbol is an interface including a communication processor, an antenna, etc. The communication I / F controls communication between multiple computers. Examples of communication standards applied to the communication I / F include wireless communication standards including 5G (5th Generation Mobile Communication System), Wi-Fi (registered trademark), Bluetooth (registered trademark), etc.
[0020] In the following embodiments, "A and / or B" is synonymous with "at least one of A and B." In other words, "A and / or B" means that it may be only A, only B, or a combination of A and B. Furthermore, in this specification, the same concept as "A and / or B" is also applied when three or more things are expressed connected by "and / or."
[0021] [First embodiment]
[0022] FIG. 1 shows an example of the configuration of a data processing system 10 according to the first embodiment.
[0023] 1, a data processing system 10 includes a data processing device 12 and a smart device 14. An example of the data processing device 12 is a server.
[0024] The data processing device 12 includes a computer 22, a database 24, and a communication I / F 26. The computer 22 is an example of a "computer" according to the technology of the present disclosure. The computer 22 includes a processor 28, a RAM 30, and a storage 32. The processor 28, the RAM 30, and the storage 32 are connected to a bus 34. The database 24 and the communication I / F 26 are also connected to the bus 34. The communication I / F 26 is connected to a network 54. Examples of the network 54 include a WAN (Wide Area Network) and / or a LAN (Local Area Network).
[0025] The smart device 14 includes a computer 36, a reception device 38, an output device 40, a camera 42, and a communication I / F 44. The computer 36 includes a processor 46, a RAM 48, and a storage 50. The processor 46, the RAM 48, and the storage 50 are connected to a bus 52. The reception device 38, the output device 40, and the camera 42 are also connected to the bus 52.
[0026] The reception device 38 includes a touch panel 38A, a microphone 38B, and the like, and receives user input. The touch panel 38A detects contact with an indicator (for example, a pen or a finger) to receive user input by the touch of the indicator. The microphone 38B detects the user's voice to receive user input by voice. The control unit 46A transmits data indicating the user input received by the touch panel 38A and the microphone 38B to the data processing device 12. In the data processing device 12, the specific processing unit 290 acquires the data indicating the user input.
[0027] The output device 40 includes a display 40A and a speaker 40B, and presents data to the user 20 by outputting the data in a form of expression that the user 20 can perceive (for example, audio and / or text). The display 40A displays visible information such as text and images in accordance with instructions from the processor 46. The speaker 40B outputs audio in accordance with instructions from the processor 46. The camera 42 is a compact digital camera equipped with an optical system including a lens, aperture, and shutter, and an imaging element such as a CMOS (Complementary Metal-Oxide-Semiconductor) image sensor or a CCD (Charge Coupled Device) image sensor.
[0028] The communication I / F 44 is connected to a network 54. The communication I / Fs 44 and 26 control the exchange of various information between the processor 46 and the processor 28 via the network 54.
[0029] FIG. 2 shows an example of the main functions of the data processing device 12 and the smart device 14.
[0030] 2, in the data processing device 12, a specific process is performed by the processor 28. A specific processing program 56 is stored in the storage 32. The specific processing program 56 is an example of a "program" according to the technology of the present disclosure. The processor 28 reads the specific processing program 56 from the storage 32 and executes the read specific processing program 56 on the RAM 30. The specific process is realized by the processor 28 operating as a specific processing unit 290 in accordance with the specific processing program 56 executed on the RAM 30.
[0031] The storage 32 stores a data generation model 58 and an emotion identification model 59. The data generation model 58 and the emotion identification model 59 are used by the identification processing unit 290.
[0032] In the smart device 14, the processor 46 performs the reception output process. The storage 50 stores a reception output program 60. The reception output program 60 is used in conjunction with the specific processing program 56 by the data processing system 10. The processor 46 reads the reception output program 60 from the storage 50 and executes the read reception output program 60 on the RAM 48. The reception output process is realized by the processor 46 operating as the control unit 46A in accordance with the reception output program 60 executed on the RAM 48.
[0033] Next, a description will be given of the specific processing performed by the specific processing unit 290 of the data processing device 12. In the following description, the data processing device 12 will be referred to as a "server" and the smart device 14 will be referred to as a "terminal."
[0034] System Overview
[0035] This system supports elderly people and children in choosing appropriate products and shopping with peace of mind. Product information captured from a user's device is sent to a server as image data and audio data, which is then analyzed by the server. Based on the analysis results, the system sends appropriate notifications to the user to support their shopping.
[0036] What the program does
[0037] 1. Data capture
[0038] The device uses a camera and microphone to capture images of the products the user picks up and audio of the conversation with the store clerk in real time.
[0039] The images include product labels, expiration dates, prices, etc.
[0040] The device converts images to the appropriate format (e.g., JPEG, PNG) and saves audio in the appropriate format (e.g., WAV, MP3).
[0041] 2. Data transmission
[0042] The device sends the captured image and audio data to the server using a secure communication protocol (e.g., HTTPS), where the data is encrypted to prevent tampering during transmission.
[0043] 3. Receiving data and preparing for analysis
[0044] The server receives the image data and audio data sent from the device, and stores the received data in temporary storage.
[0045] The server checks the integrity of the received data and checks whether it can be parsed.
[0046] 4. Analysis of image data
[0047] The server passes the image data to the Generative AI module, which extracts the following information from the image:
[0048] Product name
[0049] price
[0050] expiration date
[0051] Appearance of the product (whether damaged or not)
[0052] The generative AI will use the extracted data to make the following decisions:
[0053] Is the expiration date approaching?
[0054] Is there any damage to the product's appearance?
[0055] Is the price justified?
[0056] 5. Analysis of audio data
[0057] The server passes the voice data to the generative AI module, which extracts the following information from the voice and performs natural language processing:
[0058] What the store clerk said
[0059] Tone and degree of emphasis
[0060] The generative AI will use the extracted data to make the following decisions:
[0061] Whether the salesperson is pushing unnecessary sales
[0062] 6. Synthesis and evaluation of results
[0063] The server combines the results of image analysis and audio analysis to evaluate whether the product the user is about to purchase is appropriate.
[0064] The evaluation produces the following information:
[0065] Decision to recommend or not recommend purchase
[0066] Things to note when purchasing
[0067] 7. Notification of Results
[0068] The server formats the evaluation results and sends them to the terminal.
[0069] The device will notify the user of the received evaluation results. Notification will be done as follows:
[0070] Displayed as a pop-up message on the screen
[0071] Communicate content through audio
[0072] Specific examples
[0073] Scenario 1: Purchasing an item that is close to its expiration date
[0074] 1. The device captures a milk label that says "Best before: October 10, 2023."
[0075] 2. The device receives the image data and sends it to the server.
[0076] 3. The server receives the image and passes it to the generation AI.
[0077] 4. The generative AI detects the expiration date and recognizes that it is approaching.
[0078] 5. The server formats the results and sends them to the device.
[0079] 6. The device will notify the user that "This milk is nearing its expiration date. It is not recommended that you purchase it."
[0080] 7. The user receives a notification and returns the milk to its original position.
[0081] Scenario 2: The salesperson is trying to push you too hard
[0082] 1. The device captures the store clerk's statement, "Would you like some of this sweets as well?"
[0083] 2. The device receives the voice data and sends it to the server.
[0084] 3. The server receives the audio and passes it to the generation AI.
[0085] 4. Generative AI detects potential hard sales.
[0086] 5. The server formats the results and sends them to the device.
[0087] 6. The device will notify the user, "Please be careful as store clerks may be trying to force you to buy more than you need."
[0088] 7. The user receives a notification and responds to the store clerk, "I don't need it right now."
[0089] Through the above process, the system of the present invention supports elderly people and children in selecting appropriate products and shopping with peace of mind.
[0090] The processing flow will be explained below.
[0091] Step 1:
[0092] The device uses a camera and microphone to capture images of the products the user picks up and audio of the conversation with the store clerk in real time. The images include product labels, expiration dates, prices, etc. The images are converted to JPEG format, and the audio is converted to WAV format.
[0093] Step 2:
[0094] The device sends the captured image and audio data to the server using a secure communication protocol (HTTPS). The data is encrypted during transmission to ensure the safety of the communication.
[0095] Step 3:
[0096] The server receives the image and audio data sent from the device and stores them in temporary storage. It checks the integrity of the received data and prepares it for analysis.
[0097] Step 4:
[0098] The server passes the image data to the generation AI module, which analyzes the image data to extract information such as the product name, price, expiration date, and appearance (damage).
[0099] Step 5:
[0100] Based on the extracted data, the generative AI evaluates whether the product's expiration date is approaching, whether the product's appearance is damaged, and whether the price is appropriate.
[0101] Step 6:
[0102] The server passes the voice data to a generation AI module, which analyzes the data, extracts the content and tone of the salesperson's speech, and uses natural language processing to assess the likelihood of a hard sell.
[0103] Step 7:
[0104] Based on the results of voice data analysis, the generative AI determines whether the store clerk is making unnecessary sales pitches.
[0105] Step 8:
[0106] The server combines the results of image and audio analysis to evaluate whether the product the user is about to purchase is appropriate, and generates information such as whether to recommend or not to purchase it, as well as points to be aware of.
[0107] Step 9:
[0108] The server then formats the evaluation results and sends them to the device, using a secure communication protocol to ensure data safety.
[0109] Step 10:
[0110] The device will then notify the user of the evaluation results, which will be displayed as a pop-up message on the screen and explained to them via audio.
[0111] Step 11:
[0112] The user receives a notification from the device and can decide whether to decline the purchase or respond appropriately to the store clerk. For example, if the product is close to its expiration date, the user can decline the purchase and return it to the store clerk.
[0113] Example 1
[0114] Next, a description will be given of Example 1. In the following description, the data processing device 12 will be referred to as a "server" and the smart device 14 will be referred to as a "terminal."
[0115] Elderly people and children often have difficulty choosing the right products when shopping. Specifically, it can be difficult to check the quality, expiration date, and price of a product, and it can be difficult to avoid sales tactics from salespeople. In these situations, a method is needed to help them choose products safely and with peace of mind.
[0116] The specific processing by the specific processing unit 290 of the data processing device 12 in the first embodiment is realized by the following means.
[0117] In this invention, the server includes means for analyzing image data received by the server using a generative AI model to evaluate the product name, price, expiration date, and appearance of the product, means for analyzing voice data received by the server using a generative AI model to detect whether or not a salesperson is trying to pressure a user, and means for sending a notification to the user based on the analysis results, which makes it easier for the user to judge the quality, expiration date, and price of a product and to avoid unnecessary pressure from a salesperson.
[0118] A "user terminal" is a device equipped with a camera and a microphone, and is capable of capturing product information as image data and audio data and transmitting the data to a server.
[0119] "Capture" refers to the process of obtaining information about the product a user picks up as image data or audio data using a camera or microphone.
[0120] "Image data" is information that is digitally recorded using a camera to capture product labels and product appearances.
[0121] "Voice data" refers to information recorded in digital format using a microphone to record what a user or store clerk says.
[0122] A "server" is a device or system that receives image data and audio data sent from a user terminal and performs analysis processing.
[0123] A "generative AI model" is a model that uses machine learning algorithms to analyze image data and audio data and make decisions based on product information and user requests.
[0124] "Analysis" is the act of processing received data and extracting specific information.
[0125] "Product quality" refers to the criteria for evaluating a product's appearance, expiration date, price, etc.
[0126] "Hard selling" is when a salesperson pushes more product on a user than they need.
[0127] "Notifications" are messages or alerts that convey analysis results to users.
[0128] System Overview
[0129] This system supports elderly people and children in choosing appropriate products and shopping with peace of mind. Product information captured from a user's device is sent to a server as image data and audio data, which is then analyzed by the server. Based on the analysis results, the system sends appropriate notifications to the user to support their shopping.
[0130] Hardware and Software Configuration
[0131] This system consists of a user terminal, a server, and a generative AI model.
[0132] User Device
[0133] It is equipped with a camera to capture the product's label and appearance.
[0134] It is equipped with a microphone that records conversations with store staff and the user's voice commands.
[0135] Converting the captured data into an appropriate format (e.g. JPEG, PNG, WAV, MP3).
[0136] server
[0137] Receive data sent from the device via a secure communication protocol (e.g. HTTPS).
[0138] The integrity of the received data is checked and analysis processing is performed.
[0139] Data analysis is performed using generative AI models.
[0140] Data analysis details
[0141] Image data analysis
[0142] The server passes the image data to the generative AI model and extracts the product name, price, expiration date, and product appearance (whether damaged or not).
[0143] Based on the extracted data, the generative AI model determines whether the expiration date is approaching, whether the product is damaged, and whether the price is appropriate.
[0144] Analysis of audio data
[0145] The server passes the audio data to a generative AI model, which analyzes the content and tone of the clerk's speech.
[0146] Generative AI models detect potential hard sales.
[0147] Notification details
[0148] The server formalizes the evaluation results based on the analysis results and notifies the user.
[0149] The device will notify the user of the received evaluation results by displaying a pop-up on the screen or by voice.
[0150] Specific examples
[0151] Scenario 1: Purchasing an item that is close to its expiration date
[0152] 1. A user picks up a bottle of milk that says "Best before: October 10, 2023."
[0153] 2. The device captures the image of the milk label and sends the data to the server.
[0154] 3. The server analyzes the image data, and the generative AI model recognizes that the expiration date is approaching.
[0155] 4. The server formats the results and sends a notification to the device saying, "This milk is nearing its expiration date. It is not recommended for purchase."
[0156] 5. The user receives a notification and returns the milk to its original position.
[0157] Scenario 2: The salesperson is trying to push you too hard
[0158] 1. The user hears the store clerk say, "Would you like some of this candy as well?"
[0159] 2. The device captures this speech and sends the audio data to the server.
[0160] 3. The server analyzes the voice data and a generative AI model detects potential hard sales.
[0161] 4. The server formats the results and sends a notification to the device saying, "The store clerk may be trying to force you to buy something. Please be careful."
[0162] 5. The user receives a notification and replies to the store clerk, "I don't need it right now."
[0163] By using generative AI models for analysis, users can accurately judge the quality of products and whether or not store clerks are trying to push products. This system allows elderly people and children to shop with peace of mind by selecting appropriate products.
[0164] The flow of the identification process in the first embodiment will be described with reference to FIG.
[0165] Step 1: Capture the data
[0166] How it works: Uses the device's camera and microphone to capture information about the product the user picks up.
[0167] Input: Images and audio of the product picked up by the user (e.g., product label, conversation with the store clerk).
[0168] Processing: The device uses its camera to capture the product label (product name, price, expiration date, etc.) and uses a microphone to record the conversation with the store clerk and the user's voice commands. The image data is converted to JPEG or PNG format, and the voice data is converted to WAV or MP3 format.
[0169] Output: Image data and audio data are generated.
[0170] Step 2: Sending data
[0171] Operation: The device sends the captured image and audio data to the server.
[0172] Input: Image and audio data generated in step 1.
[0173] Processing: The device sends encrypted image and audio data to the server using a secure communication protocol such as HTTPS.
[0174] Output: Securely transmitted image and audio data.
[0175] Step 3: Receive data and prepare for analysis
[0176] Operation: The server receives image data and audio data sent from the device and temporarily stores them in storage.
[0177] Input: Image and audio data sent from the device.
[0178] Processing: When the server receives the data, it checks its integrity and whether it can be analyzed. If there are no problems, it temporarily stores it in storage.
[0179] Output: Data ready for analysis.
[0180] Step 4: Analyzing the image data
[0181] How it works: The server passes image data to the generative AI model and extracts product information.
[0182] Input: Image data stored in storage.
[0183] Processing: The generative AI model analyzes the image data and extracts the following information: product name, price, expiration date, and product appearance (damaged or not).
[0184] Output: Extracted product information (e.g. product name, price, expiry date, product appearance).
[0185] Step 5: Analyze the audio data
[0186] How it works: The server passes the audio data to a generative AI model, which analyzes the content and tone of what the clerk is saying.
[0187] Input: Audio data stored in storage.
[0188] Processing: A generative AI model analyzes the audio data and extracts information about the salesperson's speech, including their content, tone, and emphasis. Based on this, it can detect potential pressure.
[0189] Output: Extracted audio information and a judgment on whether there was a hard sell.
[0190] Step 6: Synthesis and evaluation of results
[0191] Operation: The server integrates the image analysis results and the audio analysis results to evaluate whether or not to purchase the item.
[0192] Input: Extracted product information and audio information.
[0193] Processing: The server combines the results of image analysis (e.g., expiration date, product appearance) and voice analysis (e.g., whether there is any hard sell) to evaluate whether the product the user is trying to purchase is appropriate.
[0194] Output: Evaluation result (e.g., recommended / not recommended for purchase, points to note).
[0195] Step 7: Notification of results
[0196] Operation: The server formats the evaluation results and sends them to the terminal.
[0197] Input: Consolidated evaluation results.
[0198] Processing: The server converts the evaluation results into a format that is easy for the user to understand and sends them to the terminal. The terminal receives the evaluation results and notifies the user.
[0199] Output: A notification message that is displayed to the user and / or an audio notification (e.g., "This milk is nearing its expiration date. It is not recommended for purchase.").
[0200] The above is the specific processing flow of this system. By performing appropriate data processing and calculations at each step, users can select products and shop with confidence.
[0201] (Application example 1)
[0202] Next, a description will be given of Application Example 1. In the following description, the data processing device 12 will be referred to as a "server" and the smart device 14 will be referred to as a "terminal."
[0203] There is a challenge for certain user groups, such as the elderly and children, to choose appropriate products and shop without anxiety. In particular, there is a risk that they may have difficulty determining the quality, expiration date, price, etc. of a product, or that they may be pressured by store clerks into buying unnecessary products. This creates an environment in which it is difficult to enjoy shopping with peace of mind.
[0204] The specific processing by the specific processing unit 290 of the data processing device 12 in the application example 1 is realized by the following means.
[0205] In this invention, the server includes means for capturing product information using a user terminal, means for transmitting image data and audio data of the captured product information to the server, means for analyzing the received image data in the server to evaluate the quality, price, and expiration date of the product, means for analyzing the received audio data in the server to detect whether the salesperson is trying to pressure the user, means for sending a notification to the user based on the analysis results, and means for capturing and notifying the user of the product information using an application installed on a smartphone. This allows users such as the elderly and children to select appropriate products and shop with peace of mind.
[0206] A "user terminal" is a portable computing device such as a smartphone or tablet.
[0207] "Capture" is the action of acquiring image data or audio data using a camera or microphone.
[0208] "Image data" is visual information captured by a camera or other device expressed in data format.
[0209] "Audio data" is sound information recorded by a microphone or the like expressed in data format.
[0210] A "server" is a computing device that receives data from user terminals over a network and analyzes and processes the data.
[0211] "Analysis" is the process of examining acquired data in detail for evaluation or judgment.
[0212] "Quality" is a property that describes the condition and performance of a product.
[0213] "Price" is the amount paid for a product.
[0214] The "best before" date indicates the period during which a product is safe and delicious to eat.
[0215] "Hard selling" is when a salesperson tries to force you into buying a product.
[0216] "Notifications" are messages or alerts that communicate analysis results to users.
[0217] An "application" is a software program that is installed on a user device, such as a smartphone or tablet, and performs a specific function.
[0218] In order to support specific user groups such as the elderly and children in selecting appropriate products and shopping with peace of mind, the system of the present invention is configured by the following means.
[0219] The system includes a user terminal, a server, and a program for linking them. The user terminal can be a smartphone or tablet. The user terminal is equipped with a camera and microphone to capture product information as image and audio data. The captured data is then sent to the server using a secure protocol.
[0220] The server analyzes the received image and audio data to evaluate the product's quality, price, and expiration date. This evaluation uses image and audio analysis technologies. Image analysis evaluates the product's label, expiration date, price, and appearance (whether damaged or not). Deep learning models and generative AI models are used for this. For example, machine learning frameworks such as TensorFlow and PyTorch can be used.
[0221] Voice data analysis analyzes conversations with store clerks to detect whether or not there is any hard sell. Natural language processing (NLP) technology, such as Google's Dialogflow or OpenAI's GPT model, is used to convert the voice file into text, and the likelihood of hard sell is assessed based on that text.
[0222] The analysis results are integrated and notifications are sent to users based on the evaluation. Notifications are sent through an application installed on the smartphone, and users can check the results via pop-up messages or voice notifications. Based on the analysis results, purchase recommendations or non-recommendations are also notified.
[0223] Specific use cases include the following scenarios:
[0224] Scenario 1: Purchasing an item that is close to its expiration date
[0225] 1. The user device captures a product label that displays "Best before: October 10, 2023."
[0226] 2. The device sends the captured image data to the server.
[0227] 3. The server analyzes the image, and the generative AI model detects the expiration date and recognizes that it is approaching.
[0228] 4. The server sends the results to the device.
[0229] 5. The device will notify the user that "This product's expiration date is approaching. It is not recommended that you purchase it."
[0230] Scenario 2: The salesperson is trying to push you too hard
[0231] 1. The user device captures the store clerk's statement, "Would you like some of this sweets as well?"
[0232] 2. The device sends the captured audio data to the server.
[0233] 3. The server analyzes the audio and a generative AI model detects potential hard sales attempts.
[0234] 4. The server sends the results to the device.
[0235] 5. The device will notify the user, "Please be careful as store clerks may be trying to force you to buy more than you need."
[0236] Prompt Sentence Examples
[0237] Image data: Product label image showing the expiration date "October 10, 2023"
[0238] Audio data: A recording of the store clerk saying, "Would you like some of these sweets as well?"
[0239] Analysis results:
[0240] 1. This product is nearing its expiration date and is not recommended for purchase.
[0241] 2. Be careful, as store clerks may try to pressure you into buying something you don't need.
[0242] In this way, the system of the present invention supports elderly people and children in choosing appropriate products and shopping with peace of mind.
[0243] The flow of the specific processing in the application example 1 will be described with reference to FIG.
[0244] Step 1:
[0245] The user picks up a product, captures the product label with the smartphone camera, and records the conversation with the store clerk using the microphone. The input is the image data captured by the camera and the audio data recorded by the microphone. The output is an image file (e.g., JPEG) and an audio file (e.g., WAV).
[0246] Step 2:
[0247] The device encrypts the captured image and audio data and sends them to the server using a secure communication protocol (e.g., HTTPS) to maintain security. The input is an image file and an audio file, and the output is an encrypted data packet.
[0248] Step 3:
[0249] The server receives the encrypted data sent from the terminal, decrypts it, and temporarily stores it in storage. The input is the encrypted data packet, and the output is the decrypted image data and audio data.
[0250] Step 4:
[0251] The server passes the image data to a generative AI model, which analyzes the product's quality, price, and expiration date. The input is the decoded image data, and the output is the analysis results for the product's quality, price, and expiration date. For example, the generative AI model reads the expiration date label from the image and determines whether the expiration date is approaching.
[0252] Step 5:
[0253] The server passes the audio data to a generative AI model, which analyzes the content and tone of the salesperson's speech. The input is the decoded audio data, and the output is an analysis of the salesperson's likelihood of aggressive sales. For example, the generative AI model converts the audio file into text and detects aggressive sales cues from the text.
[0254] Step 6:
[0255] The server integrates the results of image and audio analysis and evaluates the appropriate notification content for the user. The input is the analysis results of the image and audio data, and the output is the notification content. For example, notification content such as "This product's expiration date is approaching. We do not recommend purchasing it" or "The store clerk may be trying to pressure you into buying something unnecessarily" may be generated.
[0256] Step 7:
[0257] The server then formattes the evaluation results and sends them to the user's device. The input is the consolidated analysis result, and the output is a formatted notification message. For example, the server converts the results into JSON format and sends them to the device using a secure protocol.
[0258] Step 8:
[0259] The device receives the notification from the server and displays it to the user as a pop-up message or a sound notification. The input is the notification message sent from the server, and the output is the notification displayed to the user. For example, "This product is nearing its expiration date. It is not recommended to purchase it."
[0260] By checking this notification, users can choose the right product and shop with peace of mind.
[0261] Furthermore, an emotion engine that estimates the user's emotion may be combined. That is, the identification processing unit 290 may estimate the user's emotion using the emotion identification model 59 and perform identification processing using the user's emotion.
[0262] System Overview
[0263] This system captures product information from a user's device and sends it to a server as image data or voice data, where the server analyzes the data. It also incorporates an emotion engine that recognizes the user's emotions, and sends appropriate notifications to the user in real time based on the analysis results. This allows elderly people and children to choose appropriate products and shop with peace of mind.
[0264] What the program does
[0265] 1. Data capture
[0266] The device uses a camera and microphone to capture images of the products the user picks up and audio of the conversation with the store associate in real time, including product labels, expiration dates, and prices.
[0267] The device converts and saves images in JPEG format and audio in WAV format.
[0268] 2. Data transmission
[0269] The device sends the captured image and audio data to the server using a secure communication protocol (HTTPS). The data is encrypted during transmission to ensure the safety of the communication.
[0270] 3. Receiving data and preparing for analysis
[0271] The server receives the image and audio data sent from the device and stores them in temporary storage. It checks the integrity of the received data and prepares it for analysis.
[0272] 4. Analysis of image data
[0273] The server passes the image data to the generation AI module, which analyzes the image data to extract information such as the product name, price, expiration date, and appearance (damage).
[0274] The generative AI evaluates whether the product is nearing its expiration date, whether the product's appearance is damaged, and whether the price is appropriate.
[0275] 5. Analysis of audio data
[0276] The server passes the voice data to a generation AI module, which extracts the content and tone of the salesperson's speech from the voice data and uses natural language processing to evaluate the likelihood of hard sales.
[0277] Based on the results of voice data analysis, the generative AI determines whether the store clerk is making unnecessary sales pitches.
[0278] 6. Emotion Data Analysis
[0279] The server uses an emotion engine to analyze the user's facial expressions and tone of voice based on data acquired from the camera and microphone, which allows the emotion engine to recognize the user's emotions and determine whether the user is expressing discomfort.
[0280] 7. Synthesis and evaluation of results
[0281] The server combines the results of image analysis, audio analysis, and emotion analysis to evaluate the suitability of the product the user is about to purchase, and generates information such as whether to recommend or not to purchase it, as well as points to be aware of.
[0282] 8. Notification of Results
[0283] The server then formats the evaluation results and sends them to the device, using a secure communication protocol to ensure data safety.
[0284] The device will then notify the user of the evaluation results, which will be displayed as a pop-up message on the screen and explained to them via audio.
[0285] Specific examples
[0286] Scenario 1: Purchasing an item that is close to its expiration date
[0287] 1. The device captures a milk label that says "Best before: October 10, 2023."
[0288] 2. The device receives the image data and sends it to the server.
[0289] 3. The server receives the image and passes it to the generation AI.
[0290] 4. The generative AI detects the expiration date and recognizes that it is approaching.
[0291] 5. The server formats the results and sends them to the device.
[0292] 6. The device will notify the user that "This milk is nearing its expiration date. It is not recommended that you purchase it."
[0293] 7. The user receives a notification and returns the milk to its original position.
[0294] Scenario 2: The salesperson is trying to push you too hard
[0295] 1. The device captures the store clerk's statement, "Would you like some of this sweets as well?"
[0296] 2. The device receives the voice data and sends it to the server.
[0297] 3. The server receives the audio and passes it to the generation AI.
[0298] 4. Generative AI detects potential hard sales.
[0299] 5. The server formats the results and sends them to the device.
[0300] 6. The device will notify the user, "Please be careful as store clerks may be trying to force you to buy more than you need."
[0301] 7. The user receives a notification and responds to the store clerk, "I don't need it right now."
[0302] Scenario 3: Adjusting notifications based on the user's emotional state
[0303] 1. Your device uses its camera and microphone to capture your facial expressions and tone of voice.
[0304] 2. The device receives the emotion data and sends it to the server.
[0305] 3. The server uses the emotion engine to analyze the user's emotions.
[0306] 4. The emotion engine detects the user's discomfort.
[0307] 5. The server reflects the analysis results in the evaluation and adjusts the notification content.
[0308] 6. The device will then display a tailored notification to the user, saying, "We understand you're feeling annoyed. Choose only what you need."
[0309] Through the above process, the system of the present invention supports elderly people and children in selecting appropriate products and shopping with peace of mind. The introduction of an emotion engine enables more precise support based on the user's emotional state.
[0310] The processing flow will be explained below.
[0311] Step 1:
[0312] The device uses a camera and microphone to capture images of the products the user picks up and audio of the conversation with the store clerk in real time. The images include the product label, price, expiration date, etc. The image data is converted to JPEG format, and the audio data is converted to WAV format.
[0313] Step 2:
[0314] The device sends the captured image and audio data to the server using a secure communication protocol (HTTPS). The data is encrypted during transmission, protecting it from unauthorized access.
[0315] Step 3:
[0316] The server receives the image and audio data sent from the device and stores them in temporary storage. It checks the integrity and completeness of the received data and prepares it for analysis.
[0317] Step 4:
[0318] The server passes the image data to the generation AI module, which analyzes the image and extracts information such as the product name, price, expiration date, and appearance (whether damaged or not).
[0319] Step 5:
[0320] Based on the extracted data, the generative AI evaluates whether the product's expiration date is approaching, whether the product's appearance is damaged, and whether the price is appropriate.
[0321] Step 6:
[0322] The server passes the voice data to the generative AI module, which analyzes the voice and extracts the content and tone of the clerk's speech. Natural language processing (NLP) is also performed on the speech.
[0323] Step 7:
[0324] The AI generator uses voice analysis to determine whether a salesperson is trying to push a customer too hard, and tone analysis is also taken into account.
[0325] Step 8:
[0326] The server passes data captured by the camera and microphone to the emotion engine, which analyzes the user's facial expressions and tone of voice to detect whether the user is feeling uncomfortable or stressed.
[0327] Step 9:
[0328] The server combines the results of image analysis, audio analysis, and emotion analysis to evaluate whether the product the user is about to purchase is appropriate. As a result of the evaluation, it generates information such as whether to recommend or not to purchase the product and points to be careful about. Emotional data is also taken into consideration, and the content of notifications is adjusted as necessary.
[0329] Step 10:
[0330] The server then formats the evaluation results and sends them to the device, using a secure communication protocol to ensure data safety.
[0331] Step 11:
[0332] The device will then notify the user of the evaluation results. The notification will appear as a pop-up message on the screen and will also be announced via audio. Based on the emotional data, a gentle, encouraging message may also be displayed.
[0333] Specific examples
[0334] Scenario 1: Purchasing an item that is close to its expiration date
[0335] 1. The device captures a milk label that says "Best before: October 10, 2023."
[0336] 2. The device receives the image data and sends it to the server.
[0337] 3. The server receives the image and passes it to the generation AI.
[0338] 4. The generative AI detects the expiration date and recognizes that it is approaching.
[0339] 5. The server formats the results and sends them to the device.
[0340] 6. The device will notify the user that "This milk is nearing its expiration date. It is not recommended that you purchase it."
[0341] 7. The user receives a notification and returns the milk to its original position.
[0342] Scenario 2: The salesperson is trying to push you too hard
[0343] 1. The device captures the store clerk's statement, "Would you like some of this sweets as well?"
[0344] 2. The device receives the voice data and sends it to the server.
[0345] 3. The server receives the audio and passes it to the generation AI.
[0346] 4. Generative AI detects potential hard sales.
[0347] 5. The server formats the results and sends them to the device.
[0348] 6. The device will notify the user, "Please be careful as store clerks may be trying to force you to buy more than you need."
[0349] 7. The user receives a notification and responds to the store clerk, "I don't need it right now."
[0350] Scenario 3: Adjusting notifications based on the user's emotional state
[0351] 1. Your device uses its camera and microphone to capture your facial expressions and tone of voice.
[0352] 2. The device receives the emotion data and sends it to the server.
[0353] 3. The server uses the emotion engine to analyze the user's emotions.
[0354] 4. The emotion engine detects the user's discomfort.
[0355] 5. The server reflects the analysis results in the evaluation and adjusts the notification content.
[0356] 6. The device will then display a tailored notification to the user, saying, "We understand you're feeling annoyed. Choose only what you need."
[0357] Example 2
[0358] Next, a description will be given of Example 2. In the following description, the data processing device 12 will be referred to as a "server" and the smart device 14 will be referred to as a "terminal."
[0359] When users, such as the elderly and children, purchase products, they face challenges such as difficulty in accurately assessing product information, salesperson pressure, and even their own emotional state. In particular, there is a need for systems that can quickly and accurately evaluate product quality, price, and expiration dates, detect salesperson pressure, and provide real-time notifications that take into account the user's emotional state. Furthermore, there is a lack of a means to comprehensively analyze and evaluate this information in a single system and provide appropriate advice to users.
[0360] The specific processing by the specific processing unit 290 of the data processing device 12 in the second embodiment is realized by the following means.
[0361] In this invention, the server includes means for analyzing received image data and evaluating the quality, price, and expiration date of the product, means for analyzing received voice data and detecting whether the salesperson is trying to pressure you, and means for analyzing the user's facial expression and voice tone using an emotion engine to recognize the user's emotions. This makes it possible to support elderly people and children in choosing appropriate products and shopping with peace of mind.
[0362] "User terminal" means an electronic device used by a user to capture product information and process and transmit that information to a server.
[0363] "Product information" is image data and audio data that includes information about the product's quality, price, expiration date, and so on.
[0364] A "server" is a computer system for receiving, storing, and analyzing data sent from a user terminal.
[0365] "Image data" is electronic data that contains visual information about a product, such as the product label, price, and expiration date.
[0366] "Voice data" refers to electronic data containing the content of a conversation between a salesperson and a user.
[0367] "Analysis" is the process of processing data and extracting specific information or patterns.
[0368] "Quality" is a standard for evaluating a product's condition, performance, reliability, etc.
[0369] "Price" means the amount you are willing to pay for the Goods.
[0370] "Best before date" is information indicating the expiration date of food products and the like.
[0371] "Presence or absence of hard selling" is a state that indicates whether the salesperson is trying to forcefully sell the product.
[0372] An "emotion engine" is software or hardware that analyzes a user's facial expressions and voice to recognize their emotional state.
[0373] "Notification" refers to an informational message sent to the user based on the analysis results.
[0374] This invention is a system that captures product information from a user's device and sends it to a server as image and voice data, where the server analyzes the data. It also incorporates an emotion engine that recognizes the user's emotions, and sends appropriate notifications to the user in real time based on the analysis results. This allows elderly people and children to choose appropriate products and shop with peace of mind.
[0375] System Configuration
[0376] User Device
[0377] A user device is an electronic device equipped with a camera and microphone. This includes smartphones, tablets, electronic devices, etc. A user device:
[0378] Capture of product information (image data and audio data).
[0379] Convert captured data to JPEG and WAV formats.
[0380] Sending data to the server (using the HTTPS protocol).
[0381] server
[0382] The server is a computer system that receives, stores, and analyzes data sent from user terminals. The server performs the following functions:
[0383] Check the integrity of received data and temporarily store it.
[0384] Analysis of image data (using generative AI modules).
[0385] Analysis of voice data (using generative AI modules and natural language processing techniques).
[0386] Parsing sentiment data (using the sentiment engine).
[0387] Integrating analytical results and generating evaluation results.
[0388] Sending the results to the user's device.
[0389] Specific examples of data analysis
[0390] Image data analysis
[0391] The server analyzes the image data using a generative AI module (e.g., TensorFlow or PyTorch). The generative AI extracts the product name, price, expiration date, and appearance (whether damaged or not) from the image data. For example, it analyzes a milk label that reads "Best before: October 10, 2023" and recognizes that the expiration date is approaching. Based on this, it notifies the user that "This milk's expiration date is approaching. It is not recommended that you purchase it."
[0392] Analysis of audio data
[0393] The server analyzes the voice data using automatic speech recognition (ASR) technology (e.g., Google Speech-to-Text API). The generative AI converts the voice data into text and uses natural language processing (NLP) technology (e.g., the BERT model) to analyze the content and tone of the salesperson's speech. For example, it analyzes a salesperson's statement, "Would you like to buy this candy as well?" and evaluates the possibility of a hard sell. Based on this, it notifies the customer, "The salesperson may be trying to force an unnecessary sale. Please be careful."
[0394] Emotional Data Analysis
[0395] The server uses an emotion engine (e.g., OpenCV or DeepFace) to analyze the user's facial expressions and voice tone. This allows the generative AI to recognize the user's emotional state (e.g., displeasure, joy). For example, if the user expresses displeasure, the server notifies them by saying, "It seems you are displeased. Please choose only what you need."
[0396] Prompt Sentence Examples
[0397] Below are some examples of prompt sentences:
[0398] Example prompt to detect products nearing their expiration date:
[0399] "This milk's expiration date is October 10, 2023. It is not recommended for purchase."
[0400] Example prompt to detect salespeople making hard sales pitches:
[0401] "The salesperson asked me if I wanted to buy some sweets while I was there. It might be a case of pushy sales."
[0402] In this way, the system of the present invention combines hardware such as cameras and microphones with software such as generative AI modules and emotion engines to help elderly people and children shop with peace of mind.
[0403] The flow of the identification process in the second embodiment will be described with reference to FIG.
[0404] Program processing flow
[0405] Step 1: Capture the data
[0406] The device uses a camera and microphone to capture images of the products the user picks up and audio of conversations with store staff in real time.
[0407] Input: Product images from the camera, audio data from the microphone
[0408] Output: JPEG image data, WAV audio data
[0409] Specific operation: The camera takes pictures of product labels and prices, and the microphone records conversations. The recorded data is automatically converted to JPEG and WAV format, respectively, and temporarily saved to the device's storage device.
[0410] Step 2: Sending data
[0411] The device sends the captured image and audio data to the server using the HTTPS protocol.
[0412] Input: JPEG image data, WAV audio data
[0413] Output: Encrypted data packet
[0414] Specific operation: Data is encrypted and sent to the server using a secure communication protocol (HTTPS). A hash value is generated during transmission to ensure data integrity.
[0415] Step 3: Receive data and prepare for analysis
[0416] The server receives the image data and audio data sent from the terminal and stores them in temporary storage.
[0417] Input: Encrypted data packet
[0418] Output: Image and audio data with integrity confirmed
[0419] Specific operation: Decrypts encrypted data, compares hash values to verify data integrity, and then stores the data in temporary storage (e.g., a database).
[0420] Step 4: Analyzing the image data
[0421] The server passes the image data to a generative AI module, which evaluates the product's quality, price, and expiration date.
[0422] Input: JPEG format image data
[0423] Output: Product name, price, expiration date, appearance condition evaluation result
[0424] How it works: The generative AI uses image recognition algorithms to extract information from product labels, converts it into text using OCR technology, and then uses quality assessment algorithms to evaluate the product's expiration date and appearance damage.
[0425] Step 5: Analyze the audio data
[0426] The server passes the voice data to a generation AI module to detect whether the salesperson is trying to force a sale.
[0427] Input: WAV format audio data
[0428] Output: Evaluation results of hard selling using natural language processing
[0429] Specific operations: The system converts voice data into text using automatic speech recognition (ASR) technology, and then analyzes the content and tone of the salesperson's speech using natural language processing (NLP) technology. It evaluates the likelihood of a hard sell and stores the results in a database.
[0430] Step 6: Analyze the sentiment data
[0431] The server uses an emotion engine to analyze the user's facial expressions and tone of voice to recognize the user's emotions.
[0432] Input: JPEG image data, WAV audio data
[0433] Output: Emotional state evaluation result
[0434] Specific operation: Analyzes image data from the camera and audio data from the microphone, evaluates the user's facial expressions and tone of voice using an emotion engine, analyzes whether the user is showing signs of discomfort, and records the results.
[0435] Step 7: Synthesis and evaluation of results
[0436] The server combines the results of image analysis, audio analysis, and emotion analysis to evaluate whether the product the user is about to purchase is appropriate.
[0437] Input: Image analysis results, audio analysis results, emotion analysis results
[0438] Output: Evaluation results such as purchase recommendation / non-recommendation and points to note
[0439] Specific operation: The results of each analysis are integrated and a comprehensive evaluation is performed based on a rule-based evaluation model. Information such as purchase recommendations, non-recommendations, and points to be aware of is generated and formalized as evaluation results.
[0440] Step 8: Notification of results
[0441] The server formalizes the evaluation results and sends them to the user terminal.
[0442] Input: Evaluation result
[0443] Output: Notification to user device
[0444] Specific operation: The evaluation results are re-encrypted and sent to the user's device using a secure communication protocol. The device then displays the received evaluation results as a pop-up message and announces the contents via audio.
[0445] In this way, a system can be constructed that processes data and performs data calculations at each step while providing appropriate notifications to the user in real time.
[0446] (Application example 2)
[0447] Next, a description will be given of Application Example 2. In the following description, the data processing device 12 will be referred to as a "server" and the smart device 14 will be referred to as a "terminal."
[0448] When purchasing products in physical stores, certain user groups, such as the elderly and children, often find it difficult to evaluate product quality, expiration dates, and prices, and are often annoyed by salespeople's pushy sales tactics. In such situations, it is difficult for them to select the right product and they are unable to shop with peace of mind. Therefore, there is a need for a system that can comprehensively analyze the user's emotional state and detailed product information and provide appropriate notifications.
[0449] The specification process by the specification processing unit 290 of the data processing device 12 in Application Example 2 is realized by the following means. In this invention, the server includes means for capturing product information using a user terminal, means for transmitting image data and audio data of the captured product information to the server, means for analyzing the received image data in the server and evaluating the product quality, price, and expiration date, means for analyzing the received audio data in the server and detecting whether or not the salesperson is trying to pressure the user, means for sending a notification to the user based on the analysis results, and means for analyzing the user's emotions using an emotion engine in the server and adjusting the content of the notification based on the user's emotional state. This allows the user to select appropriate products with confidence and enjoy comfortable shopping.
[0450] "User terminal" refers to an electronic device that has the function of capturing product information and transmitting it to a server.
[0451] "Capture" refers to the act of acquiring image data or audio data using a camera or microphone.
[0452] "Image data" refers to data that represents product photos and labels taken with a camera in digital format.
[0453] "Audio data" refers to data that represents conversations and environmental sounds recorded by a microphone in digital form.
[0454] A "server" refers to a computer system that receives data sent from a user terminal, analyzes it, and returns the results.
[0455] "Analysis" refers to the process of evaluating and judging received image data and audio data using a program.
[0456] "Product quality" refers to the standard by which a product is evaluated based on its condition and appearance.
[0457] "Price" refers to the price of the product.
[0458] "Best before date" refers to the period during which a product such as food will retain its quality.
[0459] "Hard selling" refers to the act of a salesperson forcibly pushing a product on a customer.
[0460] An "emotion engine" is a mechanism that analyzes a user's facial expressions and tone of voice to recognize their emotions.
[0461] "Notification" refers to messages or alerts sent to users based on analysis results.
[0462] "Adjustment" refers to the process of appropriately changing the content of notifications based on analysis results and the user's emotional state.
[0463] This invention is a system that, when a user purchases a product in a physical store, uses a smart device (e.g., a smartphone or smart glasses) to capture product information, sends it to a server, which analyzes it and sends appropriate notifications.
[0464] The main flow of the system is as follows:
[0465] Data capture
[0466] The device uses a camera and microphone to capture images of the products the user picks up and audio of the conversation with the store clerk in real time. The images include product labels, expiration dates, prices, etc. The device converts the images to JPEG format and the audio to WAV format and saves them.
[0467] Sending data
[0468] The device sends the captured image and audio data to the server using a secure communication protocol (HTTPS). The data is encrypted during transmission to ensure the safety of the communication.
[0469] Receiving data and preparing for analysis
[0470] The server receives the image and audio data sent from the device and stores them in temporary storage. It checks the integrity of the received data and prepares it for analysis.
[0471] Image data analysis
[0472] The server passes the image data to the generation AI module. The generation AI performs image analysis to extract information such as the product name, price, expiration date, and appearance (damage or not) from the image data. The generation AI evaluates whether the product's expiration date is approaching, whether the product's appearance is intact, and whether the price is appropriate.
[0473] Analysis of audio data
[0474] The server passes the voice data to the generation AI module. The generation AI extracts the clerk's speech and its tone from the voice data and evaluates the possibility of hard selling using natural language processing. Based on the results of the voice data analysis, the generation AI determines whether the clerk is trying to force a sale.
[0475] Emotional Data Analysis
[0476] The server uses an emotion engine to analyze the user's facial expressions and tone of voice based on data acquired from the camera and microphone, which allows the emotion engine to recognize the user's emotions and determine whether the user is expressing discomfort.
[0477] Consolidating and communicating results
[0478] The server combines the results of image analysis, audio analysis, and emotion analysis to evaluate whether the product the user is about to purchase is appropriate. The evaluation results include information such as whether to recommend or not to purchase, and points to note. The server then formats the evaluation results and sends them to the device. A secure communication protocol is used during transmission to ensure data safety. The device then notifies the user of the received evaluation results. The notification is displayed as a pop-up message on the screen and the content is announced via audio.
[0479] Specific examples
[0480] For example, if a user picks up milk that is close to its expiration date, the device captures the label and sends it to the server, which analyzes the expiration date and notifies the user, "This milk is close to its expiration date. We do not recommend purchasing it." If a store clerk tries to pressure the user to buy more than they need, the device captures the conversation and sends it to the server, which analyzes the content and notifies the user, "The store clerk may be trying to pressure the user to buy more than they need. Please be careful."
[0481] Prompt Sentence Examples
[0482] An example of a prompt to be input to the generative AI model is as follows:
[0483] "Binary data of product images," "Binary data of conversational audio"
[0484] This allows users to choose the right product with confidence and enjoy a comfortable shopping experience.The main hardware and software used include a smart device (with camera and microphone), a secure communication protocol (HTTPS), an image analysis module, a voice analysis module, and an emotion engine.
[0485] The flow of the specific processing in the application example 2 will be described with reference to FIG.
[0486] Step 1:
[0487] When a user picks up a product, the device uses its camera to capture an image of the product. The input is the product image, and the output is a JPEG image file. Specifically, the device's camera operates to take a photo of the product label or packaging.
[0488] Step 2:
[0489] The device uses a microphone to capture the audio of the conversation with the store clerk. The input is the audio of the conversation with the store clerk, and the output is a WAV format audio file. Specifically, the device's microphone works and records the content and tone of the clerk's speech.
[0490] Step 3:
[0491] The captured image and audio data are sent to the server via a secure communication protocol (HTTPS). The input is JPEG image data and WAV audio data, and the output is a confirmation of receipt of the sent data. Specifically, the device encrypts the data and sends it to the server.
[0492] Step 4:
[0493] The server temporarily stores the received image data and audio data in storage. The input is the transmitted JPEG image data and WAV audio data, and the output is the data stored in the server's memory area. Specifically, the server checks the integrity of the received data before storing it.
[0494] Step 5:
[0495] The server passes the image data to the generation AI module, which evaluates the product's quality, price, and expiration date. The input is JPEG image data, and the output is the product name, price, expiration date, and quality evaluation results. Specifically, the generation AI analyzes the image and extracts product information.
[0496] Step 6:
[0497] The server passes the voice data to the generation AI module, which analyzes the clerk's remarks to evaluate whether or not there was a hard sell. The input is WAV-format voice data, and the output is the evaluation result of whether or not there was a hard sell. Specifically, the generation AI analyzes the voice data and detects the possibility of a hard sell from the tone and content.
[0498] Step 7:
[0499] The server uses an emotion engine to analyze the user's facial expressions and tone of voice. The input is the user's facial expression data and tone of voice, and the output is the user's emotional assessment result. Specifically, the emotion engine recognizes the user's emotions and determines the degree of discomfort or stress.
[0500] Step 8:
[0501] The server integrates the image analysis results, audio analysis results, and emotion analysis results to generate an integrated evaluation result. The input is the various analysis results, and the output is the integrated evaluation result. Specifically, the server comprehensively evaluates each analysis result and determines the appropriate notification content.
[0502] Step 9:
[0503] The server sends the integrated evaluation results to the user terminal. The input is the integrated evaluation results, and the output is notification data that arrives at the user terminal. Specifically, the server formalizes the evaluation results and sends them to the terminal using a secure communication protocol.
[0504] Step 10:
[0505] The user device then notifies the user of the received evaluation results. The input is notification data, and the output is a pop-up message and a voice notification to the user. Specifically, the device displays the notification content on the screen and communicates it to the user by voice.
[0506] Through the above steps, the present invention allows the user to select appropriate products with confidence and enjoy comfortable shopping.
[0507] The specific processing unit 290 transmits the result of the specific processing to the smart device 14. In the smart device 14, the control unit 46A causes the output device 40 to output the result of the specific processing. The microphone 38B acquires audio indicating a user input regarding the result of the specific processing. The control unit 46A transmits audio data indicating the user input acquired by the microphone 38B to the data processing device 12. In the data processing device 12, the specific processing unit 290 acquires the audio data.
[0508] The data generation model 58 is a so-called generative AI (Artificial Intelligence). An example of the data generation model 58 is ChatGPT (Internet Search<URL: https: / / openai.com / blog / chatgpt> ), Gemini (Internet search <url: https: gemini.google.com ?hl="ja">) and other generation AIs. The data generation model 58 is obtained by performing deep learning on a neural network. A prompt including an instruction is input to the data generation model 58, and inference data such as voice data indicating voice, text data indicating text, and image data indicating an image is also input. The data generation model 58 performs inference on the input inference data in accordance with the instruction indicated by the prompt, and outputs the inference result in a data format such as voice data and text data. Here, inference refers to, for example, analysis, classification, prediction, and / or summarization.
[0509] In the above embodiment, an example in which the specific process is performed by the data processing device 12 has been given, but the technology of the present disclosure is not limited to this, and the specific process may be performed by the smart device 14.
[0510] [Second embodiment]
[0511] FIG. 3 shows an example of the configuration of a data processing system 210 according to the second embodiment.
[0512] 3, the data processing system 210 includes the data processing device 12 and smart glasses 214. An example of the data processing device 12 is a server.
[0513] The data processing device 12 includes a computer 22, a database 24, and a communication I / F 26. The computer 22 is an example of a "computer" according to the technology of the present disclosure. The computer 22 includes a processor 28, a RAM 30, and a storage 32. The processor 28, the RAM 30, and the storage 32 are connected to a bus 34. The database 24 and the communication I / F 26 are also connected to the bus 34. The communication I / F 26 is connected to a network 54. Examples of the network 54 include a WAN (Wide Area Network) and / or a LAN (Local Area Network).
[0514] The smart glasses 214 include a computer 36, a microphone 238, a speaker 240, a camera 42, and a communication I / F 44. The computer 36 includes a processor 46, a RAM 48, and a storage 50. The processor 46, the RAM 48, and the storage 50 are connected to a bus 52. The microphone 238, the speaker 240, and the camera 42 are also connected to the bus 52.
[0515] The microphone 238 receives instructions and the like from the user 20 by receiving voice uttered by the user 20. The microphone 238 captures the voice uttered by the user 20, converts the captured voice into audio data, and outputs it to the processor 46. The speaker 240 outputs audio in accordance with instructions from the processor 46.
[0516] Camera 42 is a small digital camera equipped with an optical system including a lens, aperture, and shutter, and an imaging element such as a CMOS (Complementary Metal-Oxide-Semiconductor) image sensor or a CCD (Charge Coupled Device) image sensor, and captures images of the surroundings of user 20 (for example, an imaging range defined by an angle of view equivalent to the field of vision of a typical healthy person).
[0517] The communication I / F 44 is connected to a network 54. The communication I / Fs 44 and 26 are responsible for the exchange of various information between the processor 46 and the processor 28 via the network 54. The exchange of various information between the processor 46 and the processor 28 using the communication I / Fs 44 and 26 is carried out in a secure state.
[0518] Fig. 4 shows an example of the main functions of the data processing device 12 and the smart glasses 214. As shown in Fig. 4, in the data processing device 12, a specific process is performed by the processor 28. A specific process program 56 is stored in the storage 32.
[0519] The specific processing program 56 is an example of a "program" according to the technology of the present disclosure. The processor 28 reads the specific processing program 56 from the storage 32 and executes the read specific processing program 56 on the RAM 30. The specific processing is realized by the processor 28 operating as a specific processing unit 290 in accordance with the specific processing program 56 executed on the RAM 30.
[0520] The storage 32 stores a data generation model 58 and an emotion identification model 59. The data generation model 58 and the emotion identification model 59 are used by the identification processing unit 290.
[0521] In the smart glasses 214, the reception output process is performed by the processor 46. A reception output program 60 is stored in the storage 50. The processor 46 reads the reception output program 60 from the storage 50 and executes the read reception output program 60 on the RAM 48. The reception output process is realized by the processor 46 operating as the control unit 46A in accordance with the reception output program 60 executed on the RAM 48.
[0522] Next, a description will be given of the identification process performed by the identification processing unit 290 of the data processing device 12. In the following description, the data processing device 12 will be referred to as the "server" and the smart glasses 214 will be referred to as the "terminal."
[0523] System Overview
[0524] This system supports elderly people and children in choosing appropriate products and shopping with peace of mind. Product information captured from a user's device is sent to a server as image data and audio data, which is then analyzed by the server. Based on the analysis results, the system sends appropriate notifications to the user to support their shopping.
[0525] What the program does
[0526] 1. Data capture
[0527] The device uses a camera and microphone to capture images of the products the user picks up and audio of the conversation with the store clerk in real time.
[0528] The images include product labels, expiration dates, prices, etc.
[0529] The device converts images to the appropriate format (e.g., JPEG, PNG) and saves audio in the appropriate format (e.g., WAV, MP3).
[0530] 2. Data transmission
[0531] The device sends the captured image and audio data to the server using a secure communication protocol (e.g., HTTPS), where the data is encrypted to prevent tampering during transmission.
[0532] 3. Receiving data and preparing for analysis
[0533] The server receives the image data and audio data sent from the device, and stores the received data in temporary storage.
[0534] The server checks the integrity of the received data and checks whether it can be parsed.
[0535] 4. Analysis of image data
[0536] The server passes the image data to the Generative AI module, which extracts the following information from the image:
[0537] Product name
[0538] price
[0539] expiration date
[0540] Appearance of the product (whether damaged or not)
[0541] The generative AI will use the extracted data to make the following decisions:
[0542] Is the expiration date approaching?
[0543] Is there any damage to the product's appearance?
[0544] Is the price justified?
[0545] 5. Analysis of audio data
[0546] The server passes the voice data to the generative AI module, which extracts the following information from the voice and performs natural language processing:
[0547] What the store clerk said
[0548] Tone and degree of emphasis
[0549] The generative AI will use the extracted data to make the following decisions:
[0550] Whether the salesperson is pushing unnecessary sales
[0551] 6. Synthesis and evaluation of results
[0552] The server combines the results of image analysis and audio analysis to evaluate whether the product the user is about to purchase is appropriate.
[0553] The evaluation produces the following information:
[0554] Decision to recommend or not recommend purchase
[0555] Things to note when purchasing
[0556] 7. Notification of Results
[0557] The server formats the evaluation results and sends them to the terminal.
[0558] The device will notify the user of the received evaluation results. Notification will be done as follows:
[0559] Displayed as a pop-up message on the screen
[0560] Communicate content through audio
[0561] Specific examples
[0562] Scenario 1: Purchasing an item that is close to its expiration date
[0563] 1. The device captures a milk label that says "Best before: October 10, 2023."
[0564] 2. The device receives the image data and sends it to the server.
[0565] 3. The server receives the image and passes it to the generation AI.
[0566] 4. The generative AI detects the expiration date and recognizes that it is approaching.
[0567] 5. The server formats the results and sends them to the device.
[0568] 6. The device will notify the user that "This milk is nearing its expiration date. It is not recommended that you purchase it."
[0569] 7. The user receives a notification and returns the milk to its original position.
[0570] Scenario 2: The salesperson is trying to push you too hard
[0571] 1. The device captures the store clerk's statement, "Would you like some of this sweets as well?"
[0572] 2. The device receives the voice data and sends it to the server.
[0573] 3. The server receives the audio and passes it to the generation AI.
[0574] 4. Generative AI detects potential hard sales.
[0575] 5. The server formats the results and sends them to the device.
[0576] 6. The device will notify the user, "Please be careful as store clerks may be trying to force you to buy more than you need."
[0577] 7. The user receives a notification and responds to the store clerk, "I don't need it right now."
[0578] Through the above process, the system of the present invention supports elderly people and children in selecting appropriate products and shopping with peace of mind.
[0579] The processing flow will be explained below.
[0580] Step 1:
[0581] The device uses a camera and microphone to capture images of the products the user picks up and audio of the conversation with the store clerk in real time. The images include product labels, expiration dates, prices, etc. The images are converted to JPEG format, and the audio is converted to WAV format.
[0582] Step 2:
[0583] The device sends the captured image and audio data to the server using a secure communication protocol (HTTPS). The data is encrypted during transmission to ensure the safety of the communication.
[0584] Step 3:
[0585] The server receives the image and audio data sent from the device and stores them in temporary storage. It checks the integrity of the received data and prepares it for analysis.
[0586] Step 4:
[0587] The server passes the image data to the generation AI module, which analyzes the image data to extract information such as the product name, price, expiration date, and appearance (damage).
[0588] Step 5:
[0589] Based on the extracted data, the generative AI evaluates whether the product's expiration date is approaching, whether the product's appearance is damaged, and whether the price is appropriate.
[0590] Step 6:
[0591] The server passes the voice data to a generation AI module, which analyzes the data, extracts the content and tone of the salesperson's speech, and uses natural language processing to assess the likelihood of a hard sell.
[0592] Step 7:
[0593] Based on the results of voice data analysis, the generative AI determines whether the store clerk is making unnecessary sales pitches.
[0594] Step 8:
[0595] The server combines the results of image and audio analysis to evaluate whether the product the user is about to purchase is appropriate, and generates information such as whether to recommend or not to purchase it, as well as points to be aware of.
[0596] Step 9:
[0597] The server then formats the evaluation results and sends them to the device, using a secure communication protocol to ensure data safety.
[0598] Step 10:
[0599] The device will then notify the user of the evaluation results, which will be displayed as a pop-up message on the screen and explained to them via audio.
[0600] Step 11:
[0601] The user receives a notification from the device and can decide whether to decline the purchase or respond appropriately to the store clerk. For example, if the product is close to its expiration date, the user can decline the purchase and return it to the store clerk.
[0602] Example 1
[0603] Next, a description will be given of Example 1. In the following description, the data processing device 12 will be referred to as a "server" and the smart glasses 214 will be referred to as a "terminal."
[0604] Elderly people and children often have difficulty choosing the right products when shopping. Specifically, it can be difficult to check the quality, expiration date, and price of a product, and it can be difficult to avoid sales tactics from salespeople. In these situations, a method is needed to help them choose products safely and with peace of mind.
[0605] The specific processing by the specific processing unit 290 of the data processing device 12 in the first embodiment is realized by the following means.
[0606] In this invention, the server includes means for analyzing image data received by the server using a generative AI model to evaluate the product name, price, expiration date, and appearance of the product, means for analyzing voice data received by the server using a generative AI model to detect whether or not a salesperson is trying to pressure a user, and means for sending a notification to the user based on the analysis results, which makes it easier for the user to judge the quality, expiration date, and price of a product and to avoid unnecessary pressure from a salesperson.
[0607] A "user terminal" is a device equipped with a camera and a microphone, and is capable of capturing product information as image data and audio data and transmitting the data to a server.
[0608] "Capture" refers to the process of obtaining information about the product a user picks up as image data or audio data using a camera or microphone.
[0609] "Image data" is information that is digitally recorded using a camera to capture product labels and product appearances.
[0610] "Voice data" refers to information recorded in digital format using a microphone to record what a user or store clerk says.
[0611] A "server" is a device or system that receives image data and audio data sent from a user terminal and performs analysis processing.
[0612] A "generative AI model" is a model that uses machine learning algorithms to analyze image data and audio data and make decisions based on product information and user requests.
[0613] "Analysis" is the act of processing received data and extracting specific information.
[0614] "Product quality" refers to the criteria for evaluating a product's appearance, expiration date, price, etc.
[0615] "Hard selling" is when a salesperson pushes more product on a user than they need.
[0616] "Notifications" are messages or alerts that convey analysis results to users.
[0617] System Overview
[0618] This system supports elderly people and children in choosing appropriate products and shopping with peace of mind. Product information captured from a user's device is sent to a server as image data and audio data, which is then analyzed by the server. Based on the analysis results, the system sends appropriate notifications to the user to support their shopping.
[0619] Hardware and Software Configuration
[0620] This system consists of a user terminal, a server, and a generative AI model.
[0621] User Device
[0622] It is equipped with a camera to capture the product's label and appearance.
[0623] It is equipped with a microphone that records conversations with store staff and the user's voice commands.
[0624] Converting the captured data into an appropriate format (e.g. JPEG, PNG, WAV, MP3).
[0625] server
[0626] Receive data sent from the device via a secure communication protocol (e.g. HTTPS).
[0627] The integrity of the received data is checked and analysis processing is performed.
[0628] Data analysis is performed using generative AI models.
[0629] Data analysis details
[0630] Image data analysis
[0631] The server passes the image data to the generative AI model and extracts the product name, price, expiration date, and product appearance (whether damaged or not).
[0632] Based on the extracted data, the generative AI model determines whether the expiration date is approaching, whether the product is damaged, and whether the price is appropriate.
[0633] Analysis of audio data
[0634] The server passes the audio data to a generative AI model, which analyzes the content and tone of the clerk's speech.
[0635] Generative AI models detect potential hard sales.
[0636] Notification details
[0637] The server formalizes the evaluation results based on the analysis results and notifies the user.
[0638] The device will notify the user of the received evaluation results by displaying a pop-up on the screen or by voice.
[0639] Specific examples
[0640] Scenario 1: Purchasing an item that is close to its expiration date
[0641] 1. A user picks up a bottle of milk that says "Best before: October 10, 2023."
[0642] 2. The device captures the image of the milk label and sends the data to the server.
[0643] 3. The server analyzes the image data, and the generative AI model recognizes that the expiration date is approaching.
[0644] 4. The server formats the results and sends a notification to the device saying, "This milk is nearing its expiration date. It is not recommended for purchase."
[0645] 5. The user receives a notification and returns the milk to its original position.
[0646] Scenario 2: The salesperson is trying to push you too hard
[0647] 1. The user hears the store clerk say, "Would you like some of this candy as well?"
[0648] 2. The device captures this speech and sends the audio data to the server.
[0649] 3. The server analyzes the voice data and a generative AI model detects potential hard sales.
[0650] 4. The server formats the results and sends a notification to the device saying, "The store clerk may be trying to force you to buy something. Please be careful."
[0651] 5. The user receives a notification and replies to the store clerk, "I don't need it right now."
[0652] By using generative AI models for analysis, users can accurately judge the quality of products and whether or not store clerks are trying to push products. This system allows elderly people and children to shop with peace of mind by selecting appropriate products.
[0653] The flow of the identification process in the first embodiment will be described with reference to FIG.
[0654] Step 1: Capture the data
[0655] How it works: Uses the device's camera and microphone to capture information about the product the user picks up.
[0656] Input: Images and audio of the product picked up by the user (e.g., product label, conversation with the store clerk).
[0657] Processing: The device uses its camera to capture the product label (product name, price, expiration date, etc.) and uses a microphone to record the conversation with the store clerk and the user's voice commands. The image data is converted to JPEG or PNG format, and the voice data is converted to WAV or MP3 format.
[0658] Output: Image data and audio data are generated.
[0659] Step 2: Sending data
[0660] Operation: The device sends the captured image and audio data to the server.
[0661] Input: Image and audio data generated in step 1.
[0662] Processing: The device sends encrypted image and audio data to the server using a secure communication protocol such as HTTPS.
[0663] Output: Securely transmitted image and audio data.
[0664] Step 3: Receive data and prepare for analysis
[0665] Operation: The server receives image data and audio data sent from the device and temporarily stores them in storage.
[0666] Input: Image and audio data sent from the device.
[0667] Processing: When the server receives the data, it checks its integrity and whether it can be analyzed. If there are no problems, it temporarily stores it in storage.
[0668] Output: Data ready for analysis.
[0669] Step 4: Analyzing the image data
[0670] How it works: The server passes image data to the generative AI model and extracts product information.
[0671] Input: Image data stored in storage.
[0672] Processing: The generative AI model analyzes the image data and extracts the following information: product name, price, expiration date, and product appearance (damaged or not).
[0673] Output: Extracted product information (e.g. product name, price, expiry date, product appearance).
[0674] Step 5: Analyze the audio data
[0675] How it works: The server passes the audio data to a generative AI model, which analyzes the content and tone of what the clerk is saying.
[0676] Input: Audio data stored in storage.
[0677] Processing: A generative AI model analyzes the audio data and extracts information about the salesperson's speech, including their content, tone, and emphasis. Based on this, it can detect potential pressure.
[0678] Output: Extracted audio information and a judgment on whether there was a hard sell.
[0679] Step 6: Synthesis and evaluation of results
[0680] Operation: The server integrates the image analysis results and the audio analysis results to evaluate whether or not to purchase the item.
[0681] Input: Extracted product information and audio information.
[0682] Processing: The server combines the results of image analysis (e.g., expiration date, product appearance) and voice analysis (e.g., whether there is any hard sell) to evaluate whether the product the user is trying to purchase is appropriate.
[0683] Output: Evaluation result (e.g., recommended / not recommended for purchase, points to note).
[0684] Step 7: Notification of results
[0685] Operation: The server formats the evaluation results and sends them to the terminal.
[0686] Input: Consolidated evaluation results.
[0687] Processing: The server converts the evaluation results into a format that is easy for the user to understand and sends them to the terminal. The terminal receives the evaluation results and notifies the user.
[0688] Output: A notification message that is displayed to the user and / or an audio notification (e.g., "This milk is nearing its expiration date. It is not recommended for purchase.").
[0689] The above is the specific processing flow of this system. By performing appropriate data processing and calculations at each step, users can select products and shop with confidence.
[0690] (Application example 1)
[0691] Next, a description will be given of Application Example 1. In the following description, the data processing device 12 will be referred to as a "server" and the smart glasses 214 will be referred to as a "terminal."
[0692] There is a challenge for certain user groups, such as the elderly and children, to choose appropriate products and shop without anxiety. In particular, there is a risk that they may have difficulty determining the quality, expiration date, price, etc. of a product, or that they may be pressured by store clerks into buying unnecessary products. This creates an environment in which it is difficult to enjoy shopping with peace of mind.
[0693] The specific processing by the specific processing unit 290 of the data processing device 12 in the application example 1 is realized by the following means.
[0694] In this invention, the server includes means for capturing product information using a user terminal, means for transmitting image data and audio data of the captured product information to the server, means for analyzing the received image data in the server to evaluate the quality, price, and expiration date of the product, means for analyzing the received audio data in the server to detect whether the salesperson is trying to pressure the user, means for sending a notification to the user based on the analysis results, and means for capturing and notifying the user of the product information using an application installed on a smartphone. This allows users such as the elderly and children to select appropriate products and shop with peace of mind.
[0695] A "user terminal" is a portable computing device such as a smartphone or tablet.
[0696] "Capture" is the action of acquiring image data or audio data using a camera or microphone.
[0697] "Image data" is visual information captured by a camera or other device expressed in data format.
[0698] "Audio data" is sound information recorded by a microphone or the like expressed in data format.
[0699] A "server" is a computing device that receives data from user terminals over a network and analyzes and processes the data.
[0700] "Analysis" is the process of examining acquired data in detail for evaluation or judgment.
[0701] "Quality" is a property that describes the condition and performance of a product.
[0702] "Price" is the amount paid for a product.
[0703] The "best before" date indicates the period during which a product is safe and delicious to eat.
[0704] "Hard selling" is when a salesperson tries to force you into buying a product.
[0705] "Notifications" are messages or alerts that communicate analysis results to users.
[0706] An "application" is a software program that is installed on a user device, such as a smartphone or tablet, and performs a specific function.
[0707] In order to support specific user groups such as the elderly and children in selecting appropriate products and shopping with peace of mind, the system of the present invention is configured by the following means.
[0708] The system includes a user terminal, a server, and a program for linking them. The user terminal can be a smartphone or tablet. The user terminal is equipped with a camera and microphone to capture product information as image and audio data. The captured data is then sent to the server using a secure protocol.
[0709] The server analyzes the received image and audio data to evaluate the product's quality, price, and expiration date. This evaluation uses image and audio analysis technologies. Image analysis evaluates the product's label, expiration date, price, and appearance (whether damaged or not). Deep learning models and generative AI models are used for this. For example, machine learning frameworks such as TensorFlow and PyTorch can be used.
[0710] Voice data analysis analyzes conversations with store clerks to detect whether or not there is any hard sell. Natural language processing (NLP) technology, such as Google's Dialogflow or OpenAI's GPT model, is used to convert the voice file into text, and the likelihood of hard sell is assessed based on that text.
[0711] The analysis results are integrated and notifications are sent to users based on the evaluation. Notifications are sent through an application installed on the smartphone, and users can check the results via pop-up messages or voice notifications. Based on the analysis results, purchase recommendations or non-recommendations are also notified.
[0712] Specific use cases include the following scenarios:
[0713] Scenario 1: Purchasing an item that is close to its expiration date
[0714] 1. The user device captures a product label that displays "Best before: October 10, 2023."
[0715] 2. The device sends the captured image data to the server.
[0716] 3. The server analyzes the image, and the generative AI model detects the expiration date and recognizes that it is approaching.
[0717] 4. The server sends the results to the device.
[0718] 5. The device will notify the user that "This product's expiration date is approaching. It is not recommended that you purchase it."
[0719] Scenario 2: The salesperson is trying to push you too hard
[0720] 1. The user device captures the store clerk's statement, "Would you like some of this sweets as well?"
[0721] 2. The device sends the captured audio data to the server.
[0722] 3. The server analyzes the audio and a generative AI model detects potential hard sales attempts.
[0723] 4. The server sends the results to the device.
[0724] 5. The device will notify the user, "Please be careful as store clerks may be trying to force you to buy more than you need."
[0725] Prompt Sentence Examples
[0726] Image data: Product label image showing the expiration date "October 10, 2023"
[0727] Audio data: A recording of the store clerk saying, "Would you like some of these sweets as well?"
[0728] Analysis results:
[0729] 1. This product is nearing its expiration date and is not recommended for purchase.
[0730] 2. Be careful, as store clerks may try to pressure you into buying something you don't need.
[0731] In this way, the system of the present invention supports elderly people and children in choosing appropriate products and shopping with peace of mind.
[0732] The flow of the specific processing in the application example 1 will be described with reference to FIG.
[0733] Step 1:
[0734] The user picks up a product, captures the product label with the smartphone camera, and records the conversation with the store clerk using the microphone. The input is the image data captured by the camera and the audio data recorded by the microphone. The output is an image file (e.g., JPEG) and an audio file (e.g., WAV).
[0735] Step 2:
[0736] The device encrypts the captured image and audio data and sends them to the server using a secure communication protocol (e.g., HTTPS) to maintain security. The input is an image file and an audio file, and the output is an encrypted data packet.
[0737] Step 3:
[0738] The server receives the encrypted data sent from the terminal, decrypts it, and temporarily stores it in storage. The input is the encrypted data packet, and the output is the decrypted image data and audio data.
[0739] Step 4:
[0740] The server passes the image data to a generative AI model, which analyzes the product's quality, price, and expiration date. The input is the decoded image data, and the output is the analysis results for the product's quality, price, and expiration date. For example, the generative AI model reads the expiration date label from the image and determines whether the expiration date is approaching.
[0741] Step 5:
[0742] The server passes the audio data to a generative AI model, which analyzes the content and tone of the salesperson's speech. The input is the decoded audio data, and the output is an analysis of the salesperson's likelihood of aggressive sales. For example, the generative AI model converts the audio file into text and detects aggressive sales cues from the text.
[0743] Step 6:
[0744] The server integrates the results of image and audio analysis and evaluates the appropriate notification content for the user. The input is the analysis results of the image and audio data, and the output is the notification content. For example, notification content such as "This product's expiration date is approaching. We do not recommend purchasing it" or "The store clerk may be trying to pressure you into buying something unnecessarily" may be generated.
[0745] Step 7:
[0746] The server then formattes the evaluation results and sends them to the user's device. The input is the consolidated analysis result, and the output is a formatted notification message. For example, the server converts the results into JSON format and sends them to the device using a secure protocol.
[0747] Step 8:
[0748] The device receives the notification from the server and displays it to the user as a pop-up message or a sound notification. The input is the notification message sent from the server, and the output is the notification displayed to the user. For example, "This product is nearing its expiration date. It is not recommended to purchase it."
[0749] By checking this notification, users can choose the right product and shop with peace of mind.
[0750] Furthermore, an emotion engine that estimates the user's emotion may be further combined. That is, the identification processing unit 290 may estimate the user's emotion using the emotion identification model 59, and perform identification processing using the user's emotion.
[0751] System Overview
[0752] This system captures product information from a user's device and sends it to a server as image data or voice data, where the server analyzes the data. It also incorporates an emotion engine that recognizes the user's emotions, and sends appropriate notifications to the user in real time based on the analysis results. This allows elderly people and children to choose appropriate products and shop with peace of mind.
[0753] What the program does
[0754] 1. Data capture
[0755] The device uses a camera and microphone to capture images of the products the user picks up and audio of the conversation with the store associate in real time, including product labels, expiration dates, and prices.
[0756] The device converts and saves images in JPEG format and audio in WAV format.
[0757] 2. Data transmission
[0758] The device sends the captured image and audio data to the server using a secure communication protocol (HTTPS). The data is encrypted during transmission to ensure the safety of the communication.
[0759] 3. Receiving data and preparing for analysis
[0760] The server receives the image and audio data sent from the device and stores them in temporary storage. It checks the integrity of the received data and prepares it for analysis.
[0761] 4. Analysis of image data
[0762] The server passes the image data to the generation AI module, which analyzes the image data to extract information such as the product name, price, expiration date, and appearance (damage).
[0763] The generative AI evaluates whether the product is nearing its expiration date, whether the product's appearance is damaged, and whether the price is appropriate.
[0764] 5. Analysis of audio data
[0765] The server passes the voice data to a generation AI module, which extracts the content and tone of the salesperson's speech from the voice data and uses natural language processing to evaluate the likelihood of hard sales.
[0766] Based on the results of voice data analysis, the generative AI determines whether the store clerk is making unnecessary sales pitches.
[0767] 6. Emotion Data Analysis
[0768] The server uses an emotion engine to analyze the user's facial expressions and tone of voice based on data acquired from the camera and microphone, which allows the emotion engine to recognize the user's emotions and determine whether the user is expressing discomfort.
[0769] 7. Synthesis and evaluation of results
[0770] The server combines the results of image analysis, audio analysis, and emotion analysis to evaluate the suitability of the product the user is about to purchase, and generates information such as whether to recommend or not to purchase it, as well as points to be aware of.
[0771] 8. Notification of Results
[0772] The server then formats the evaluation results and sends them to the device, using a secure communication protocol to ensure data safety.
[0773] The device will then notify the user of the evaluation results, which will be displayed as a pop-up message on the screen and explained to them via audio.
[0774] Specific examples
[0775] Scenario 1: Purchasing an item that is close to its expiration date
[0776] 1. The device captures a milk label that says "Best before: October 10, 2023."
[0777] 2. The device receives the image data and sends it to the server.
[0778] 3. The server receives the image and passes it to the generation AI.
[0779] 4. The generative AI detects the expiration date and recognizes that it is approaching.
[0780] 5. The server formats the results and sends them to the device.
[0781] 6. The device will notify the user that "This milk is nearing its expiration date. It is not recommended that you purchase it."
[0782] 7. The user receives a notification and returns the milk to its original position.
[0783] Scenario 2: The salesperson is trying to push you too hard
[0784] 1. The device captures the store clerk's statement, "Would you like some of this sweets as well?"
[0785] 2. The device receives the voice data and sends it to the server.
[0786] 3. The server receives the audio and passes it to the generation AI.
[0787] 4. Generative AI detects potential hard sales.
[0788] 5. The server formats the results and sends them to the device.
[0789] 6. The device will notify the user, "Please be careful as store clerks may be trying to force you to buy more than you need."
[0790] 7. The user receives a notification and responds to the store clerk, "I don't need it right now."
[0791] Scenario 3: Adjusting notifications based on the user's emotional state
[0792] 1. Your device uses its camera and microphone to capture your facial expressions and tone of voice.
[0793] 2. The device receives the emotion data and sends it to the server.
[0794] 3. The server uses the emotion engine to analyze the user's emotions.
[0795] 4. The emotion engine detects the user's discomfort.
[0796] 5. The server reflects the analysis results in the evaluation and adjusts the notification content.
[0797] 6. The device will then display a tailored notification to the user, saying, "We understand you're feeling annoyed. Choose only what you need."
[0798] Through the above process, the system of the present invention supports elderly people and children in selecting appropriate products and shopping with peace of mind. The introduction of an emotion engine enables more precise support based on the user's emotional state.
[0799] The processing flow will be explained below.
[0800] Step 1:
[0801] The device uses a camera and microphone to capture images of the products the user picks up and audio of the conversation with the store clerk in real time. The images include the product label, price, expiration date, etc. The image data is converted to JPEG format, and the audio data is converted to WAV format.
[0802] Step 2:
[0803] The device sends the captured image and audio data to the server using a secure communication protocol (HTTPS). The data is encrypted during transmission, protecting it from unauthorized access.
[0804] Step 3:
[0805] The server receives the image and audio data sent from the device and stores them in temporary storage. It checks the integrity and completeness of the received data and prepares it for analysis.
[0806] Step 4:
[0807] The server passes the image data to the generation AI module, which analyzes the image and extracts information such as the product name, price, expiration date, and appearance (whether damaged or not).
[0808] Step 5:
[0809] Based on the extracted data, the generative AI evaluates whether the product's expiration date is approaching, whether the product's appearance is damaged, and whether the price is appropriate.
[0810] Step 6:
[0811] The server passes the voice data to the generative AI module, which analyzes the voice and extracts the content and tone of the clerk's speech. Natural language processing (NLP) is also performed on the speech.
[0812] Step 7:
[0813] The AI generator uses voice analysis to determine whether a salesperson is trying to push a customer too hard, and tone analysis is also taken into account.
[0814] Step 8:
[0815] The server passes data captured by the camera and microphone to the emotion engine, which analyzes the user's facial expressions and tone of voice to detect whether the user is feeling uncomfortable or stressed.
[0816] Step 9:
[0817] The server combines the results of image analysis, audio analysis, and emotion analysis to evaluate whether the product the user is about to purchase is appropriate. As a result of the evaluation, it generates information such as whether to recommend or not to purchase the product and points to be careful about. Emotional data is also taken into consideration, and the content of notifications is adjusted as necessary.
[0818] Step 10:
[0819] The server then formats the evaluation results and sends them to the device, using a secure communication protocol to ensure data safety.
[0820] Step 11:
[0821] The device will then notify the user of the evaluation results. The notification will appear as a pop-up message on the screen and will also be announced via audio. Based on the emotional data, a gentle, encouraging message may also be displayed.
[0822] Specific examples
[0823] Scenario 1: Purchasing an item that is close to its expiration date
[0824] 1. The device captures a milk label that says "Best before: October 10, 2023."
[0825] 2. The device receives the image data and sends it to the server.
[0826] 3. The server receives the image and passes it to the generation AI.
[0827] 4. The generative AI detects the expiration date and recognizes that it is approaching.
[0828] 5. The server formats the results and sends them to the device.
[0829] 6. The device will notify the user that "This milk is nearing its expiration date. It is not recommended that you purchase it."
[0830] 7. The user receives a notification and returns the milk to its original position.
[0831] Scenario 2: The salesperson is trying to push you too hard
[0832] 1. The device captures the store clerk's statement, "Would you like some of this sweets as well?"
[0833] 2. The device receives the voice data and sends it to the server.
[0834] 3. The server receives the audio and passes it to the generation AI.
[0835] 4. Generative AI detects potential hard sales.
[0836] 5. The server formats the results and sends them to the device.
[0837] 6. The device will notify the user, "Please be careful as store clerks may be trying to force you to buy more than you need."
[0838] 7. The user receives a notification and responds to the store clerk, "I don't need it right now."
[0839] Scenario 3: Adjusting notifications based on the user's emotional state
[0840] 1. Your device uses its camera and microphone to capture your facial expressions and tone of voice.
[0841] 2. The device receives the emotion data and sends it to the server.
[0842] 3. The server uses the emotion engine to analyze the user's emotions.
[0843] 4. The emotion engine detects the user's discomfort.
[0844] 5. The server reflects the analysis results in the evaluation and adjusts the notification content.
[0845] 6. The device will then display a tailored notification to the user, saying, "We understand you're feeling annoyed. Choose only what you need."
[0846] Example 2
[0847] Next, a description will be given of Example 2. In the following description, the data processing device 12 will be referred to as a "server" and the smart glasses 214 will be referred to as a "terminal."
[0848] When users, such as the elderly and children, purchase products, they face challenges such as difficulty in accurately assessing product information, salesperson pressure, and even their own emotional state. In particular, there is a need for systems that can quickly and accurately evaluate product quality, price, and expiration dates, detect salesperson pressure, and provide real-time notifications that take into account the user's emotional state. Furthermore, there is a lack of a means to comprehensively analyze and evaluate this information in a single system and provide appropriate advice to users.
[0849] The specific processing by the specific processing unit 290 of the data processing device 12 in the second embodiment is realized by the following means.
[0850] In this invention, the server includes means for analyzing received image data and evaluating the quality, price, and expiration date of the product, means for analyzing received voice data and detecting whether the salesperson is trying to pressure you, and means for analyzing the user's facial expression and voice tone using an emotion engine to recognize the user's emotions. This makes it possible to support elderly people and children in choosing appropriate products and shopping with peace of mind.
[0851] "User terminal" means an electronic device used by a user to capture product information and process and transmit that information to a server.
[0852] "Product information" is image data and audio data that includes information about the product's quality, price, expiration date, and so on.
[0853] A "server" is a computer system for receiving, storing, and analyzing data sent from a user terminal.
[0854] "Image data" is electronic data that contains visual information about a product, such as the product label, price, and expiration date.
[0855] "Voice data" refers to electronic data containing the content of a conversation between a salesperson and a user.
[0856] "Analysis" is the process of processing data and extracting specific information or patterns.
[0857] "Quality" is a standard for evaluating a product's condition, performance, reliability, etc.
[0858] "Price" means the amount you are willing to pay for the Goods.
[0859] "Best before date" is information indicating the expiration date of food products and the like.
[0860] "Presence or absence of hard selling" is a state that indicates whether the salesperson is trying to forcefully sell the product.
[0861] An "emotion engine" is software or hardware that analyzes a user's facial expressions and voice to recognize their emotional state.
[0862] "Notification" refers to an informational message sent to the user based on the analysis results.
[0863] This invention is a system that captures product information from a user's device and sends it to a server as image and voice data, where the server analyzes the data. It also incorporates an emotion engine that recognizes the user's emotions, and sends appropriate notifications to the user in real time based on the analysis results. This allows elderly people and children to choose appropriate products and shop with peace of mind.
[0864] System Configuration
[0865] User Device
[0866] A user device is an electronic device equipped with a camera and microphone. This includes smartphones, tablets, electronic devices, etc. A user device:
[0867] Capture of product information (image data and audio data).
[0868] Convert captured data to JPEG and WAV formats.
[0869] Sending data to the server (using the HTTPS protocol).
[0870] server
[0871] The server is a computer system that receives, stores, and analyzes data sent from user terminals. The server performs the following functions:
[0872] Check the integrity of received data and temporarily store it.
[0873] Analysis of image data (using generative AI modules).
[0874] Analysis of voice data (using generative AI modules and natural language processing techniques).
[0875] Parsing sentiment data (using the sentiment engine).
[0876] Integrating analytical results and generating evaluation results.
[0877] Sending the results to the user's device.
[0878] Specific examples of data analysis
[0879] Image data analysis
[0880] The server analyzes the image data using a generative AI module (e.g., TensorFlow or PyTorch). The generative AI extracts the product name, price, expiration date, and appearance (whether damaged or not) from the image data. For example, it analyzes a milk label that reads "Best before: October 10, 2023" and recognizes that the expiration date is approaching. Based on this, it notifies the user that "This milk's expiration date is approaching. It is not recommended that you purchase it."
[0881] Analysis of audio data
[0882] The server analyzes the voice data using automatic speech recognition (ASR) technology (e.g., Google Speech-to-Text API). The generative AI converts the voice data into text and uses natural language processing (NLP) technology (e.g., the BERT model) to analyze the content and tone of the salesperson's speech. For example, it analyzes a salesperson's statement, "Would you like to buy this candy as well?" and evaluates the possibility of a hard sell. Based on this, it notifies the customer, "The salesperson may be trying to force an unnecessary sale. Please be careful."
[0883] Emotional Data Analysis
[0884] The server uses an emotion engine (e.g., OpenCV or DeepFace) to analyze the user's facial expressions and voice tone. This allows the generative AI to recognize the user's emotional state (e.g., displeasure, joy). For example, if the user expresses displeasure, the server notifies them by saying, "It seems you are displeased. Please choose only what you need."
[0885] Prompt Sentence Examples
[0886] Below are some examples of prompt sentences:
[0887] Example prompt to detect products nearing their expiration date:
[0888] "This milk's expiration date is October 10, 2023. It is not recommended for purchase."
[0889] Example prompt to detect salespeople making hard sales pitches:
[0890] "The salesperson asked me if I wanted to buy some sweets while I was there. It might be a case of pushy sales."
[0891] In this way, the system of the present invention combines hardware such as cameras and microphones with software such as generative AI modules and emotion engines to help elderly people and children shop with peace of mind.
[0892] The flow of the identification process in the second embodiment will be described with reference to FIG.
[0893] Program processing flow
[0894] Step 1: Capture the data
[0895] The device uses a camera and microphone to capture images of the products the user picks up and audio of conversations with store staff in real time.
[0896] Input: Product images from the camera, audio data from the microphone
[0897] Output: JPEG image data, WAV audio data
[0898] Specific operation: The camera takes pictures of product labels and prices, and the microphone records conversations. The recorded data is automatically converted to JPEG and WAV format, respectively, and temporarily saved to the device's storage device.
[0899] Step 2: Sending data
[0900] The device sends the captured image and audio data to the server using the HTTPS protocol.
[0901] Input: JPEG image data, WAV audio data
[0902] Output: Encrypted data packet
[0903] Specific operation: Data is encrypted and sent to the server using a secure communication protocol (HTTPS). A hash value is generated during transmission to ensure data integrity.
[0904] Step 3: Receive data and prepare for analysis
[0905] The server receives the image data and audio data sent from the terminal and stores them in temporary storage.
[0906] Input: Encrypted data packet
[0907] Output: Image and audio data with integrity confirmed
[0908] Specific operation: Decrypts encrypted data, compares hash values to verify data integrity, and then stores the data in temporary storage (e.g., a database).
[0909] Step 4: Analyzing the image data
[0910] The server passes the image data to a generative AI module, which evaluates the product's quality, price, and expiration date.
[0911] Input: JPEG format image data
[0912] Output: Product name, price, expiration date, appearance condition evaluation result
[0913] How it works: The generative AI uses image recognition algorithms to extract information from product labels, converts it into text using OCR technology, and then uses quality assessment algorithms to evaluate the product's expiration date and appearance damage.
[0914] Step 5: Analyze the audio data
[0915] The server passes the voice data to a generation AI module to detect whether the salesperson is trying to force a sale.
[0916] Input: WAV format audio data
[0917] Output: Evaluation results of hard selling using natural language processing
[0918] Specific operations: The system converts voice data into text using automatic speech recognition (ASR) technology, and then analyzes the content and tone of the salesperson's speech using natural language processing (NLP) technology. It evaluates the likelihood of a hard sell and stores the results in a database.
[0919] Step 6: Analyze the sentiment data
[0920] The server uses an emotion engine to analyze the user's facial expressions and tone of voice to recognize the user's emotions.
[0921] Input: JPEG image data, WAV audio data
[0922] Output: Emotional state evaluation result
[0923] Specific operation: Analyzes image data from the camera and audio data from the microphone, evaluates the user's facial expressions and tone of voice using an emotion engine, analyzes whether the user is showing signs of discomfort, and records the results.
[0924] Step 7: Synthesis and evaluation of results
[0925] The server combines the results of image analysis, audio analysis, and emotion analysis to evaluate whether the product the user is about to purchase is appropriate.
[0926] Input: Image analysis results, audio analysis results, emotion analysis results
[0927] Output: Evaluation results such as purchase recommendation / non-recommendation and points to note
[0928] Specific operation: The results of each analysis are integrated and a comprehensive evaluation is performed based on a rule-based evaluation model. Information such as purchase recommendations, non-recommendations, and points to be aware of is generated and formalized as evaluation results.
[0929] Step 8: Notification of results
[0930] The server formalizes the evaluation results and sends them to the user terminal.
[0931] Input: Evaluation result
[0932] Output: Notification to user device
[0933] Specific operation: The evaluation results are re-encrypted and sent to the user's device using a secure communication protocol. The device then displays the received evaluation results as a pop-up message and announces the contents via audio.
[0934] In this way, a system can be constructed that processes data and performs data calculations at each step while providing appropriate notifications to the user in real time.
[0935] (Application example 2)
[0936] Next, a description will be given of Application Example 2. In the following description, the data processing device 12 will be referred to as a "server" and the smart glasses 214 will be referred to as a "terminal."
[0937] When purchasing products in physical stores, certain user groups, such as the elderly and children, often find it difficult to evaluate product quality, expiration dates, and prices, and are often annoyed by salespeople's pushy sales tactics. In such situations, it is difficult for them to select the right product and they are unable to shop with peace of mind. Therefore, there is a need for a system that can comprehensively analyze the user's emotional state and detailed product information and provide appropriate notifications.
[0938] The specification process by the specification processing unit 290 of the data processing device 12 in Application Example 2 is realized by the following means. In this invention, the server includes means for capturing product information using a user terminal, means for transmitting image data and audio data of the captured product information to the server, means for analyzing the received image data in the server and evaluating the product quality, price, and expiration date, means for analyzing the received audio data in the server and detecting whether or not the salesperson is trying to pressure the user, means for sending a notification to the user based on the analysis results, and means for analyzing the user's emotions using an emotion engine in the server and adjusting the content of the notification based on the user's emotional state. This allows the user to select appropriate products with confidence and enjoy comfortable shopping.
[0939] "User terminal" refers to an electronic device that has the function of capturing product information and transmitting it to a server.
[0940] "Capture" refers to the act of acquiring image data or audio data using a camera or microphone.
[0941] "Image data" refers to data that represents product photos and labels taken with a camera in digital format.
[0942] "Audio data" refers to data that represents conversations and environmental sounds recorded by a microphone in digital form.
[0943] A "server" refers to a computer system that receives data sent from a user terminal, analyzes it, and returns the results.
[0944] "Analysis" refers to the process of evaluating and judging received image data and audio data using a program.
[0945] "Product quality" refers to the standard by which a product is evaluated based on its condition and appearance.
[0946] "Price" refers to the price of the product.
[0947] "Best before date" refers to the period during which a product such as food will retain its quality.
[0948] "Hard selling" refers to the act of a salesperson forcibly pushing a product on a customer.
[0949] An "emotion engine" is a mechanism that analyzes a user's facial expressions and tone of voice to recognize their emotions.
[0950] "Notification" refers to messages or alerts sent to users based on analysis results.
[0951] "Adjustment" refers to the process of appropriately changing the content of notifications based on analysis results and the user's emotional state.
[0952] This invention is a system that, when a user purchases a product in a physical store, uses a smart device (e.g., a smartphone or smart glasses) to capture product information, sends it to a server, which analyzes it and sends appropriate notifications.
[0953] The main flow of the system is as follows:
[0954] Data capture
[0955] The device uses a camera and microphone to capture images of the products the user picks up and audio of the conversation with the store clerk in real time. The images include product labels, expiration dates, prices, etc. The device converts the images to JPEG format and the audio to WAV format and saves them.
[0956] Sending data
[0957] The device sends the captured image and audio data to the server using a secure communication protocol (HTTPS). The data is encrypted during transmission to ensure the safety of the communication.
[0958] Receiving data and preparing for analysis
[0959] The server receives the image and audio data sent from the device and stores them in temporary storage. It checks the integrity of the received data and prepares it for analysis.
[0960] Image data analysis
[0961] The server passes the image data to the generation AI module. The generation AI performs image analysis to extract information such as the product name, price, expiration date, and appearance (damage or not) from the image data. The generation AI evaluates whether the product's expiration date is approaching, whether the product's appearance is intact, and whether the price is appropriate.
[0962] Analysis of audio data
[0963] The server passes the voice data to the generation AI module. The generation AI extracts the clerk's speech and its tone from the voice data and evaluates the possibility of hard selling using natural language processing. Based on the results of the voice data analysis, the generation AI determines whether the clerk is trying to force a sale.
[0964] Emotional Data Analysis
[0965] The server uses an emotion engine to analyze the user's facial expressions and tone of voice based on data acquired from the camera and microphone, which allows the emotion engine to recognize the user's emotions and determine whether the user is expressing discomfort.
[0966] Consolidating and communicating results
[0967] The server combines the results of image analysis, audio analysis, and emotion analysis to evaluate whether the product the user is about to purchase is appropriate. The evaluation results include information such as whether to recommend or not to purchase, and points to note. The server then formats the evaluation results and sends them to the device. A secure communication protocol is used during transmission to ensure data safety. The device then notifies the user of the received evaluation results. The notification is displayed as a pop-up message on the screen and the content is announced via audio.
[0968] Specific examples
[0969] For example, if a user picks up milk that is close to its expiration date, the device captures the label and sends it to the server, which analyzes the expiration date and notifies the user, "This milk is close to its expiration date. We do not recommend purchasing it." If a store clerk tries to pressure the user to buy more than they need, the device captures the conversation and sends it to the server, which analyzes the content and notifies the user, "The store clerk may be trying to pressure the user to buy more than they need. Please be careful."
[0970] Prompt Sentence Examples
[0971] An example of a prompt to be input to the generative AI model is as follows:
[0972] "Binary data of product images," "Binary data of conversational audio"
[0973] This allows users to choose the right product with confidence and enjoy a comfortable shopping experience.The main hardware and software used include a smart device (with camera and microphone), a secure communication protocol (HTTPS), an image analysis module, a voice analysis module, and an emotion engine.
[0974] The flow of the specific processing in the application example 2 will be described with reference to FIG.
[0975] Step 1:
[0976] When a user picks up a product, the device uses its camera to capture an image of the product. The input is the product image, and the output is a JPEG image file. Specifically, the device's camera operates to take a photo of the product label or packaging.
[0977] Step 2:
[0978] The device uses a microphone to capture the audio of the conversation with the store clerk. The input is the audio of the conversation with the store clerk, and the output is a WAV format audio file. Specifically, the device's microphone works and records the content and tone of the clerk's speech.
[0979] Step 3:
[0980] The captured image and audio data are sent to the server via a secure communication protocol (HTTPS). The input is JPEG image data and WAV audio data, and the output is a confirmation of receipt of the sent data. Specifically, the device encrypts the data and sends it to the server.
[0981] Step 4:
[0982] The server temporarily stores the received image data and audio data in storage. The input is the transmitted JPEG image data and WAV audio data, and the output is the data stored in the server's memory area. Specifically, the server checks the integrity of the received data before storing it.
[0983] Step 5:
[0984] The server passes the image data to the generation AI module, which evaluates the product's quality, price, and expiration date. The input is JPEG image data, and the output is the product name, price, expiration date, and quality evaluation results. Specifically, the generation AI analyzes the image and extracts product information.
[0985] Step 6:
[0986] The server passes the voice data to the generation AI module, which analyzes the clerk's remarks to evaluate whether or not there was a hard sell. The input is WAV-format voice data, and the output is the evaluation result of whether or not there was a hard sell. Specifically, the generation AI analyzes the voice data and detects the possibility of a hard sell from the tone and content.
[0987] Step 7:
[0988] The server uses an emotion engine to analyze the user's facial expressions and tone of voice. The input is the user's facial expression data and tone of voice, and the output is the user's emotional assessment result. Specifically, the emotion engine recognizes the user's emotions and determines the degree of discomfort or stress.
[0989] Step 8:
[0990] The server integrates the image analysis results, audio analysis results, and emotion analysis results to generate an integrated evaluation result. The input is the various analysis results, and the output is the integrated evaluation result. Specifically, the server comprehensively evaluates each analysis result and determines the appropriate notification content.
[0991] Step 9:
[0992] The server sends the integrated evaluation results to the user terminal. The input is the integrated evaluation results, and the output is notification data that arrives at the user terminal. Specifically, the server formalizes the evaluation results and sends them to the terminal using a secure communication protocol.
[0993] Step 10:
[0994] The user device then notifies the user of the received evaluation results. The input is notification data, and the output is a pop-up message and a voice notification to the user. Specifically, the device displays the notification content on the screen and communicates it to the user by voice.
[0995] Through the above steps, the present invention allows the user to select appropriate products with confidence and enjoy comfortable shopping.
[0996] The specific processing unit 290 transmits the result of the specific processing to the smart glasses 214. In the smart glasses 214, the control unit 46A causes the speaker 240 to output the result of the specific processing. The microphone 238 acquires audio indicating a user input regarding the result of the specific processing. The control unit 46A transmits audio data indicating the user input acquired by the microphone 238 to the data processing device 12. In the data processing device 12, the specific processing unit 290 acquires the audio data.
[0997] The data generation model 58 is a so-called generative AI (Artificial Intelligence). An example of the data generation model 58 is ChatGPT (Internet Search<URL: https: / / openai.com / blog / chatgpt> ), Gemini (Internet search <url: https: gemini.google.com ?hl="ja">) and other generation AIs. The data generation model 58 is obtained by performing deep learning on a neural network. A prompt including an instruction is input to the data generation model 58, and inference data such as voice data indicating voice, text data indicating text, and image data indicating an image is also input. The data generation model 58 performs inference on the input inference data in accordance with the instruction indicated by the prompt, and outputs the inference result in a data format such as voice data and text data. Here, inference refers to, for example, analysis, classification, prediction, and / or summarization.
[0998] In the above embodiment, an example in which the specific processing is performed by the data processing device 12 has been given, but the technology of the present disclosure is not limited to this, and the specific processing may be performed by the smart glasses 214.
[0999] [Third embodiment]
[1000] FIG. 5 shows an example of the configuration of a data processing system 310 according to the third embodiment.
[1001] 5, the data processing system 310 includes the data processing device 12 and a headset type terminal 314. An example of the data processing device 12 is a server.
[1002] The data processing device 12 includes a computer 22, a database 24, and a communication I / F 26. The computer 22 is an example of a "computer" according to the technology of the present disclosure. The computer 22 includes a processor 28, a RAM 30, and a storage 32. The processor 28, the RAM 30, and the storage 32 are connected to a bus 34. The database 24 and the communication I / F 26 are also connected to the bus 34. The communication I / F 26 is connected to a network 54. Examples of the network 54 include a WAN (Wide Area Network) and / or a LAN (Local Area Network).
[1003] The headset type terminal 314 includes a computer 36, a microphone 238, a speaker 240, a camera 42, a communication I / F 44, and a display 343. The computer 36 includes a processor 46, a RAM 48, and a storage 50. The processor 46, the RAM 48, and the storage 50 are connected to a bus 52. The microphone 238, the speaker 240, the camera 42, and the display 343 are also connected to the bus 52.
[1004] The microphone 238 receives instructions and the like from the user 20 by receiving voice uttered by the user 20. The microphone 238 captures the voice uttered by the user 20, converts the captured voice into audio data, and outputs it to the processor 46. The speaker 240 outputs audio in accordance with instructions from the processor 46.
[1005] Camera 42 is a small digital camera equipped with an optical system including a lens, aperture, and shutter, and an imaging element such as a CMOS (Complementary Metal-Oxide-Semiconductor) image sensor or a CCD (Charge Coupled Device) image sensor, and captures images of the surroundings of user 20 (for example, an imaging range defined by an angle of view equivalent to the field of vision of a typical healthy person).
[1006] The communication I / F 44 is connected to a network 54. The communication I / Fs 44 and 26 are responsible for the exchange of various information between the processor 46 and the processor 28 via the network 54. The exchange of various information between the processor 46 and the processor 28 using the communication I / Fs 44 and 26 is carried out in a secure state.
[1007] Fig. 6 shows an example of the main functions of the data processing device 12 and the headset type terminal 314. As shown in Fig. 6, in the data processing device 12, a specific process is performed by the processor 28. A specific process program 56 is stored in the storage 32.
[1008] The specific processing program 56 is an example of a "program" according to the technology of the present disclosure. The processor 28 reads the specific processing program 56 from the storage 32 and executes the read specific processing program 56 on the RAM 30. The specific processing is realized by the processor 28 operating as a specific processing unit 290 in accordance with the specific processing program 56 executed on the RAM 30.
[1009] The storage 32 stores a data generation model 58 and an emotion identification model 59. The data generation model 58 and the emotion identification model 59 are used by the identification processing unit 290.
[1010] In the headset type terminal 314, a reception output process is performed by the processor 46. A reception output program 60 is stored in the storage 50. The processor 46 reads the reception output program 60 from the storage 50 and executes the read reception output program 60 on the RAM 48. The reception output process is realized by the processor 46 operating as the control unit 46A in accordance with the reception output program 60 executed on the RAM 48.
[1011] Next, a description will be given of the identification process performed by the identification processing unit 290 of the data processing device 12. In the following description, the data processing device 12 will be referred to as the "server" and the headset type terminal 314 will be referred to as the "terminal."
[1012] System Overview
[1013] This system supports elderly people and children in choosing appropriate products and shopping with peace of mind. Product information captured from a user's device is sent to a server as image data and audio data, which is then analyzed by the server. Based on the analysis results, the system sends appropriate notifications to the user to support their shopping.
[1014] What the program does
[1015] 1. Data capture
[1016] The device uses a camera and microphone to capture images of the products the user picks up and audio of the conversation with the store clerk in real time.
[1017] The images include product labels, expiration dates, prices, etc.
[1018] The device converts images to the appropriate format (e.g., JPEG, PNG) and saves audio in the appropriate format (e.g., WAV, MP3).
[1019] 2. Data transmission
[1020] The device sends the captured image and audio data to the server using a secure communication protocol (e.g., HTTPS), where the data is encrypted to prevent tampering during transmission.
[1021] 3. Receiving data and preparing for analysis
[1022] The server receives the image data and audio data sent from the device, and stores the received data in temporary storage.
[1023] The server checks the integrity of the received data and checks whether it can be parsed.
[1024] 4. Analysis of image data
[1025] The server passes the image data to the Generative AI module, which extracts the following information from the image:
[1026] Product name
[1027] price
[1028] expiration date
[1029] Appearance of the product (whether damaged or not)
[1030] The generative AI will use the extracted data to make the following decisions:
[1031] Is the expiration date approaching?
[1032] Is there any damage to the product's appearance?
[1033] Is the price justified?
[1034] 5. Analysis of audio data
[1035] The server passes the voice data to the generative AI module, which extracts the following information from the voice and performs natural language processing:
[1036] What the store clerk said
[1037] Tone and degree of emphasis
[1038] The generative AI will use the extracted data to make the following decisions:
[1039] Whether the salesperson is pushing unnecessary sales
[1040] 6. Synthesis and evaluation of results
[1041] The server combines the results of image analysis and audio analysis to evaluate whether the product the user is about to purchase is appropriate.
[1042] The evaluation produces the following information:
[1043] Decision to recommend or not recommend purchase
[1044] Things to note when purchasing
[1045] 7. Notification of Results
[1046] The server formats the evaluation results and sends them to the terminal.
[1047] The device will notify the user of the received evaluation results. Notification will be done as follows:
[1048] Displayed as a pop-up message on the screen
[1049] Communicate content through audio
[1050] Specific examples
[1051] Scenario 1: Purchasing an item that is close to its expiration date
[1052] 1. The device captures a milk label that says "Best before: October 10, 2023."
[1053] 2. The device receives the image data and sends it to the server.
[1054] 3. The server receives the image and passes it to the generation AI.
[1055] 4. The generative AI detects the expiration date and recognizes that it is approaching.
[1056] 5. The server formats the results and sends them to the device.
[1057] 6. The device will notify the user that "This milk is nearing its expiration date. It is not recommended that you purchase it."
[1058] 7. The user receives a notification and returns the milk to its original position.
[1059] Scenario 2: The salesperson is trying to push you too hard
[1060] 1. The device captures the store clerk's statement, "Would you like some of this sweets as well?"
[1061] 2. The device receives the voice data and sends it to the server.
[1062] 3. The server receives the audio and passes it to the generation AI.
[1063] 4. Generative AI detects potential hard sales.
[1064] 5. The server formats the results and sends them to the device.
[1065] 6. The device will notify the user, "Please be careful as store clerks may be trying to force you to buy more than you need."
[1066] 7. The user receives a notification and responds to the store clerk, "I don't need it right now."
[1067] Through the above process, the system of the present invention supports elderly people and children in selecting appropriate products and shopping with peace of mind.
[1068] The processing flow will be explained below.
[1069] Step 1:
[1070] The device uses a camera and microphone to capture images of the products the user picks up and audio of the conversation with the store clerk in real time. The images include product labels, expiration dates, prices, etc. The images are converted to JPEG format, and the audio is converted to WAV format.
[1071] Step 2:
[1072] The device sends the captured image and audio data to the server using a secure communication protocol (HTTPS). The data is encrypted during transmission to ensure the safety of the communication.
[1073] Step 3:
[1074] The server receives the image and audio data sent from the device and stores them in temporary storage. It checks the integrity of the received data and prepares it for analysis.
[1075] Step 4:
[1076] The server passes the image data to the generation AI module, which analyzes the image data to extract information such as the product name, price, expiration date, and appearance (damage).
[1077] Step 5:
[1078] Based on the extracted data, the generative AI evaluates whether the product's expiration date is approaching, whether the product's appearance is damaged, and whether the price is appropriate.
[1079] Step 6:
[1080] The server passes the voice data to a generation AI module, which analyzes the data, extracts the content and tone of the salesperson's speech, and uses natural language processing to assess the likelihood of a hard sell.
[1081] Step 7:
[1082] Based on the results of voice data analysis, the generative AI determines whether the store clerk is making unnecessary sales pitches.
[1083] Step 8:
[1084] The server combines the results of image and audio analysis to evaluate whether the product the user is about to purchase is appropriate, and generates information such as whether to recommend or not to purchase it, as well as points to be aware of.
[1085] Step 9:
[1086] The server then formats the evaluation results and sends them to the device, using a secure communication protocol to ensure data safety.
[1087] Step 10:
[1088] The device will then notify the user of the evaluation results, which will be displayed as a pop-up message on the screen and explained to them via audio.
[1089] Step 11:
[1090] The user receives a notification from the device and can decide whether to decline the purchase or respond appropriately to the store clerk. For example, if the product is close to its expiration date, the user can decline the purchase and return it to the store clerk.
[1091] Example 1
[1092] Next, a description will be given of Example 1. In the following description, the data processing device 12 will be referred to as a "server" and the headset type terminal 314 will be referred to as a "terminal."
[1093] Elderly people and children often have difficulty choosing the right products when shopping. Specifically, it can be difficult to check the quality, expiration date, and price of a product, and it can be difficult to avoid sales tactics from salespeople. In these situations, a method is needed to help them choose products safely and with peace of mind.
[1094] The specific processing by the specific processing unit 290 of the data processing device 12 in the first embodiment is realized by the following means.
[1095] In this invention, the server includes means for analyzing image data received by the server using a generative AI model to evaluate the product name, price, expiration date, and appearance of the product, means for analyzing voice data received by the server using a generative AI model to detect whether or not a salesperson is trying to pressure a user, and means for sending a notification to the user based on the analysis results, which makes it easier for the user to judge the quality, expiration date, and price of a product and to avoid unnecessary pressure from a salesperson.
[1096] A "user terminal" is a device equipped with a camera and a microphone, and is capable of capturing product information as image data and audio data and transmitting the data to a server.
[1097] "Capture" refers to the process of obtaining information about the product a user picks up as image data or audio data using a camera or microphone.
[1098] "Image data" is information that is digitally recorded using a camera to capture product labels and product appearances.
[1099] "Voice data" refers to information recorded in digital format using a microphone to record what a user or store clerk says.
[1100] A "server" is a device or system that receives image data and audio data sent from a user terminal and performs analysis processing.
[1101] A "generative AI model" is a model that uses machine learning algorithms to analyze image data and audio data and make decisions based on product information and user requests.
[1102] "Analysis" is the act of processing received data and extracting specific information.
[1103] "Product quality" refers to the criteria for evaluating a product's appearance, expiration date, price, etc.
[1104] "Hard selling" is when a salesperson pushes more product on a user than they need.
[1105] "Notifications" are messages or alerts that convey analysis results to users.
[1106] System Overview
[1107] This system supports elderly people and children in choosing appropriate products and shopping with peace of mind. Product information captured from a user's device is sent to a server as image data and audio data, which is then analyzed by the server. Based on the analysis results, the system sends appropriate notifications to the user to support their shopping.
[1108] Hardware and Software Configuration
[1109] This system consists of a user terminal, a server, and a generative AI model.
[1110] User Device
[1111] It is equipped with a camera to capture the product's label and appearance.
[1112] It is equipped with a microphone that records conversations with store staff and the user's voice commands.
[1113] Converting the captured data into an appropriate format (e.g. JPEG, PNG, WAV, MP3).
[1114] server
[1115] Receive data sent from the device via a secure communication protocol (e.g. HTTPS).
[1116] The integrity of the received data is checked and analysis processing is performed.
[1117] Data analysis is performed using generative AI models.
[1118] Data analysis details
[1119] Image data analysis
[1120] The server passes the image data to the generative AI model and extracts the product name, price, expiration date, and product appearance (whether damaged or not).
[1121] Based on the extracted data, the generative AI model determines whether the expiration date is approaching, whether the product is damaged, and whether the price is appropriate.
[1122] Analysis of audio data
[1123] The server passes the audio data to a generative AI model, which analyzes the content and tone of the clerk's speech.
[1124] Generative AI models detect potential hard sales.
[1125] Notification details
[1126] The server formalizes the evaluation results based on the analysis results and notifies the user.
[1127] The device will notify the user of the received evaluation results by displaying a pop-up on the screen or by voice.
[1128] Specific examples
[1129] Scenario 1: Purchasing an item that is close to its expiration date
[1130] 1. A user picks up a bottle of milk that says "Best before: October 10, 2023."
[1131] 2. The device captures the image of the milk label and sends the data to the server.
[1132] 3. The server analyzes the image data, and the generative AI model recognizes that the expiration date is approaching.
[1133] 4. The server formats the results and sends a notification to the device saying, "This milk is nearing its expiration date. It is not recommended for purchase."
[1134] 5. The user receives a notification and returns the milk to its original position.
[1135] Scenario 2: The salesperson is trying to push you too hard
[1136] 1. The user hears the store clerk say, "Would you like some of this candy as well?"
[1137] 2. The device captures this speech and sends the audio data to the server.
[1138] 3. The server analyzes the voice data and a generative AI model detects potential hard sales.
[1139] 4. The server formats the results and sends a notification to the device saying, "The store clerk may be trying to force you to buy something. Please be careful."
[1140] 5. The user receives a notification and replies to the store clerk, "I don't need it right now."
[1141] By using generative AI models for analysis, users can accurately judge the quality of products and whether or not store clerks are trying to push products. This system allows elderly people and children to shop with peace of mind by selecting appropriate products.
[1142] The flow of the identification process in the first embodiment will be described with reference to FIG.
[1143] Step 1: Capture the data
[1144] How it works: Uses the device's camera and microphone to capture information about the product the user picks up.
[1145] Input: Images and audio of the product picked up by the user (e.g., product label, conversation with the store clerk).
[1146] Processing: The device uses its camera to capture the product label (product name, price, expiration date, etc.) and uses a microphone to record the conversation with the store clerk and the user's voice commands. The image data is converted to JPEG or PNG format, and the voice data is converted to WAV or MP3 format.
[1147] Output: Image data and audio data are generated.
[1148] Step 2: Sending data
[1149] Operation: The device sends the captured image and audio data to the server.
[1150] Input: Image and audio data generated in step 1.
[1151] Processing: The device sends encrypted image and audio data to the server using a secure communication protocol such as HTTPS.
[1152] Output: Securely transmitted image and audio data.
[1153] Step 3: Receive data and prepare for analysis
[1154] Operation: The server receives image data and audio data sent from the device and temporarily stores them in storage.
[1155] Input: Image and audio data sent from the device.
[1156] Processing: When the server receives the data, it checks its integrity and whether it can be analyzed. If there are no problems, it temporarily stores it in storage.
[1157] Output: Data ready for analysis.
[1158] Step 4: Analyzing the image data
[1159] How it works: The server passes image data to the generative AI model and extracts product information.
[1160] Input: Image data stored in storage.
[1161] Processing: The generative AI model analyzes the image data and extracts the following information: product name, price, expiration date, and product appearance (damaged or not).
[1162] Output: Extracted product information (e.g. product name, price, expiry date, product appearance).
[1163] Step 5: Analyze the audio data
[1164] How it works: The server passes the audio data to a generative AI model, which analyzes the content and tone of what the clerk is saying.
[1165] Input: Audio data stored in storage.
[1166] Processing: A generative AI model analyzes the audio data and extracts information about the salesperson's speech, including their content, tone, and emphasis. Based on this, it can detect potential pressure.
[1167] Output: Extracted audio information and a judgment on whether there was a hard sell.
[1168] Step 6: Synthesis and evaluation of results
[1169] Operation: The server integrates the image analysis results and the audio analysis results to evaluate whether or not to purchase the item.
[1170] Input: Extracted product information and audio information.
[1171] Processing: The server combines the results of image analysis (e.g., expiration date, product appearance) and voice analysis (e.g., whether there is any hard sell) to evaluate whether the product the user is trying to purchase is appropriate.
[1172] Output: Evaluation result (e.g., recommended / not recommended for purchase, points to note).
[1173] Step 7: Notification of results
[1174] Operation: The server formats the evaluation results and sends them to the terminal.
[1175] Input: Consolidated evaluation results.
[1176] Processing: The server converts the evaluation results into a format that is easy for the user to understand and sends them to the terminal. The terminal receives the evaluation results and notifies the user.
[1177] Output: A notification message that is displayed to the user and / or an audio notification (e.g., "This milk is nearing its expiration date. It is not recommended for purchase.").
[1178] The above is the specific processing flow of this system. By performing appropriate data processing and calculations at each step, users can select products and shop with confidence.
[1179] (Application example 1)
[1180] Next, a description will be given of Application Example 1. In the following description, the data processing device 12 will be referred to as a "server" and the headset type terminal 314 will be referred to as a "terminal."
[1181] There is a challenge for certain user groups, such as the elderly and children, to choose appropriate products and shop without anxiety. In particular, there is a risk that they may have difficulty determining the quality, expiration date, price, etc. of a product, or that they may be pressured by store clerks into buying unnecessary products. This creates an environment in which it is difficult to enjoy shopping with peace of mind.
[1182] The specific processing by the specific processing unit 290 of the data processing device 12 in the application example 1 is realized by the following means.
[1183] In this invention, the server includes means for capturing product information using a user terminal, means for transmitting image data and audio data of the captured product information to the server, means for analyzing the received image data in the server to evaluate the quality, price, and expiration date of the product, means for analyzing the received audio data in the server to detect whether the salesperson is trying to pressure the user, means for sending a notification to the user based on the analysis results, and means for capturing and notifying the user of the product information using an application installed on a smartphone. This allows users such as the elderly and children to select appropriate products and shop with peace of mind.
[1184] A "user terminal" is a portable computing device such as a smartphone or tablet.
[1185] "Capture" is the action of acquiring image data or audio data using a camera or microphone.
[1186] "Image data" is visual information captured by a camera or other device expressed in data format.
[1187] "Audio data" is sound information recorded by a microphone or the like expressed in data format.
[1188] A "server" is a computing device that receives data from user terminals over a network and analyzes and processes the data.
[1189] "Analysis" is the process of examining acquired data in detail for evaluation or judgment.
[1190] "Quality" is a property that describes the condition and performance of a product.
[1191] "Price" is the amount paid for a product.
[1192] The "best before" date indicates the period during which a product is safe and delicious to eat.
[1193] "Hard selling" is when a salesperson tries to force you into buying a product.
[1194] "Notifications" are messages or alerts that communicate analysis results to users.
[1195] An "application" is a software program that is installed on a user device, such as a smartphone or tablet, and performs a specific function.
[1196] In order to support specific user groups such as the elderly and children in selecting appropriate products and shopping with peace of mind, the system of the present invention is configured by the following means.
[1197] The system includes a user terminal, a server, and a program for linking them. The user terminal can be a smartphone or tablet. The user terminal is equipped with a camera and microphone to capture product information as image and audio data. The captured data is then sent to the server using a secure protocol.
[1198] The server analyzes the received image and audio data to evaluate the product's quality, price, and expiration date. This evaluation uses image and audio analysis technologies. Image analysis evaluates the product's label, expiration date, price, and appearance (whether damaged or not). Deep learning models and generative AI models are used for this. For example, machine learning frameworks such as TensorFlow and PyTorch can be used.
[1199] Voice data analysis analyzes conversations with store clerks to detect whether or not there is any hard sell. Natural language processing (NLP) technology, such as Google's Dialogflow or OpenAI's GPT model, is used to convert the voice file into text, and the likelihood of hard sell is assessed based on that text.
[1200] The analysis results are integrated and notifications are sent to users based on the evaluation. Notifications are sent through an application installed on the smartphone, and users can check the results via pop-up messages or voice notifications. Based on the analysis results, purchase recommendations or non-recommendations are also notified.
[1201] Specific use cases include the following scenarios:
[1202] Scenario 1: Purchasing an item that is close to its expiration date
[1203] 1. The user device captures a product label that displays "Best before: October 10, 2023."
[1204] 2. The device sends the captured image data to the server.
[1205] 3. The server analyzes the image, and the generative AI model detects the expiration date and recognizes that it is approaching.
[1206] 4. The server sends the results to the device.
[1207] 5. The device will notify the user that "This product's expiration date is approaching. It is not recommended that you purchase it."
[1208] Scenario 2: The salesperson is trying to push you too hard
[1209] 1. The user device captures the store clerk's statement, "Would you like some of this sweets as well?"
[1210] 2. The device sends the captured audio data to the server.
[1211] 3. The server analyzes the audio and a generative AI model detects potential hard sales attempts.
[1212] 4. The server sends the results to the device.
[1213] 5. The device will notify the user, "Please be careful as store clerks may be trying to force you to buy more than you need."
[1214] Prompt Sentence Examples
[1215] Image data: Product label image showing the expiration date "October 10, 2023"
[1216] Audio data: A recording of the store clerk saying, "Would you like some of these sweets as well?"
[1217] Analysis results:
[1218] 1. This product is nearing its expiration date and is not recommended for purchase.
[1219] 2. Be careful, as store clerks may try to pressure you into buying something you don't need.
[1220] In this way, the system of the present invention supports elderly people and children in choosing appropriate products and shopping with peace of mind.
[1221] The flow of the specific processing in the application example 1 will be described with reference to FIG.
[1222] Step 1:
[1223] The user picks up a product, captures the product label with the smartphone camera, and records the conversation with the store clerk using the microphone. The input is the image data captured by the camera and the audio data recorded by the microphone. The output is an image file (e.g., JPEG) and an audio file (e.g., WAV).
[1224] Step 2:
[1225] The device encrypts the captured image and audio data and sends them to the server using a secure communication protocol (e.g., HTTPS) to maintain security. The input is an image file and an audio file, and the output is an encrypted data packet.
[1226] Step 3:
[1227] The server receives the encrypted data sent from the terminal, decrypts it, and temporarily stores it in storage. The input is the encrypted data packet, and the output is the decrypted image data and audio data.
[1228] Step 4:
[1229] The server passes the image data to a generative AI model, which analyzes the product's quality, price, and expiration date. The input is the decoded image data, and the output is the analysis results for the product's quality, price, and expiration date. For example, the generative AI model reads the expiration date label from the image and determines whether the expiration date is approaching.
[1230] Step 5:
[1231] The server passes the audio data to a generative AI model, which analyzes the content and tone of the salesperson's speech. The input is the decoded audio data, and the output is an analysis of the salesperson's likelihood of aggressive sales. For example, the generative AI model converts the audio file into text and detects aggressive sales cues from the text.
[1232] Step 6:
[1233] The server integrates the results of image and audio analysis and evaluates the appropriate notification content for the user. The input is the analysis results of the image and audio data, and the output is the notification content. For example, notification content such as "This product's expiration date is approaching. We do not recommend purchasing it" or "The store clerk may be trying to pressure you into buying something unnecessarily" may be generated.
[1234] Step 7:
[1235] The server then formattes the evaluation results and sends them to the user's device. The input is the consolidated analysis result, and the output is a formatted notification message. For example, the server converts the results into JSON format and sends them to the device using a secure protocol.
[1236] Step 8:
[1237] The device receives the notification from the server and displays it to the user as a pop-up message or a sound notification. The input is the notification message sent from the server, and the output is the notification displayed to the user. For example, "This product is nearing its expiration date. It is not recommended to purchase it."
[1238] By checking this notification, users can choose the right product and shop with peace of mind.
[1239] Furthermore, an emotion engine that estimates the user's emotion may be further combined. That is, the identification processing unit 290 may estimate the user's emotion using the emotion identification model 59, and perform identification processing using the user's emotion.
[1240] System Overview
[1241] This system captures product information from a user's device and sends it to a server as image data or voice data, where the server analyzes the data. It also incorporates an emotion engine that recognizes the user's emotions, and sends appropriate notifications to the user in real time based on the analysis results. This allows elderly people and children to choose appropriate products and shop with peace of mind.
[1242] What the program does
[1243] 1. Data capture
[1244] The device uses a camera and microphone to capture images of the products the user picks up and audio of the conversation with the store associate in real time, including product labels, expiration dates, and prices.
[1245] The device converts and saves images in JPEG format and audio in WAV format.
[1246] 2. Data transmission
[1247] The device sends the captured image and audio data to the server using a secure communication protocol (HTTPS). The data is encrypted during transmission to ensure the safety of the communication.
[1248] 3. Receiving data and preparing for analysis
[1249] The server receives the image and audio data sent from the device and stores them in temporary storage. It checks the integrity of the received data and prepares it for analysis.
[1250] 4. Analysis of image data
[1251] The server passes the image data to the generation AI module, which analyzes the image data to extract information such as the product name, price, expiration date, and appearance (damage).
[1252] The generative AI evaluates whether the product is nearing its expiration date, whether the product's appearance is damaged, and whether the price is appropriate.
[1253] 5. Analysis of audio data
[1254] The server passes the voice data to a generation AI module, which extracts the content and tone of the salesperson's speech from the voice data and uses natural language processing to evaluate the likelihood of hard sales.
[1255] Based on the results of voice data analysis, the generative AI determines whether the store clerk is making unnecessary sales pitches.
[1256] 6. Emotion Data Analysis
[1257] The server uses an emotion engine to analyze the user's facial expressions and tone of voice based on data acquired from the camera and microphone, which allows the emotion engine to recognize the user's emotions and determine whether the user is expressing discomfort.
[1258] 7. Synthesis and evaluation of results
[1259] The server combines the results of image analysis, audio analysis, and emotion analysis to evaluate the suitability of the product the user is about to purchase, and generates information such as whether to recommend or not to purchase it, as well as points to be aware of.
[1260] 8. Notification of Results
[1261] The server then formats the evaluation results and sends them to the device, using a secure communication protocol to ensure data safety.
[1262] The device will then notify the user of the evaluation results, which will be displayed as a pop-up message on the screen and explained to them via audio.
[1263] Specific examples
[1264] Scenario 1: Purchasing an item that is close to its expiration date
[1265] 1. The device captures a milk label that says "Best before: October 10, 2023."
[1266] 2. The device receives the image data and sends it to the server.
[1267] 3. The server receives the image and passes it to the generation AI.
[1268] 4. The generative AI detects the expiration date and recognizes that it is approaching.
[1269] 5. The server formats the results and sends them to the device.
[1270] 6. The device will notify the user that "This milk is nearing its expiration date. It is not recommended that you purchase it."
[1271] 7. The user receives a notification and returns the milk to its original position.
[1272] Scenario 2: The salesperson is trying to push you too hard
[1273] 1. The device captures the store clerk's statement, "Would you like some of this sweets as well?"
[1274] 2. The device receives the voice data and sends it to the server.
[1275] 3. The server receives the audio and passes it to the generation AI.
[1276] 4. Generative AI detects potential hard sales.
[1277] 5. The server formats the results and sends them to the device.
[1278] 6. The device will notify the user, "Please be careful as store clerks may be trying to force you to buy more than you need."
[1279] 7. The user receives a notification and responds to the store clerk, "I don't need it right now."
[1280] Scenario 3: Adjusting notifications based on the user's emotional state
[1281] 1. Your device uses its camera and microphone to capture your facial expressions and tone of voice.
[1282] 2. The device receives the emotion data and sends it to the server.
[1283] 3. The server uses the emotion engine to analyze the user's emotions.
[1284] 4. The emotion engine detects the user's discomfort.
[1285] 5. The server reflects the analysis results in the evaluation and adjusts the notification content.
[1286] 6. The device will then display a tailored notification to the user, saying, "We understand you're feeling annoyed. Choose only what you need."
[1287] Through the above process, the system of the present invention supports elderly people and children in selecting appropriate products and shopping with peace of mind. The introduction of an emotion engine enables more precise support based on the user's emotional state.
[1288] The processing flow will be explained below.
[1289] Step 1:
[1290] The device uses a camera and microphone to capture images of the products the user picks up and audio of the conversation with the store clerk in real time. The images include the product label, price, expiration date, etc. The image data is converted to JPEG format, and the audio data is converted to WAV format.
[1291] Step 2:
[1292] The device sends the captured image and audio data to the server using a secure communication protocol (HTTPS). The data is encrypted during transmission, protecting it from unauthorized access.
[1293] Step 3:
[1294] The server receives the image and audio data sent from the device and stores them in temporary storage. It checks the integrity and completeness of the received data and prepares it for analysis.
[1295] Step 4:
[1296] The server passes the image data to the generation AI module, which analyzes the image and extracts information such as the product name, price, expiration date, and appearance (whether damaged or not).
[1297] Step 5:
[1298] Based on the extracted data, the generative AI evaluates whether the product's expiration date is approaching, whether the product's appearance is damaged, and whether the price is appropriate.
[1299] Step 6:
[1300] The server passes the voice data to the generative AI module, which analyzes the voice and extracts the content and tone of the clerk's speech. Natural language processing (NLP) is also performed on the speech.
[1301] Step 7:
[1302] The AI generator uses voice analysis to determine whether a salesperson is trying to push a customer too hard, and tone analysis is also taken into account.
[1303] Step 8:
[1304] The server passes data captured by the camera and microphone to the emotion engine, which analyzes the user's facial expressions and tone of voice to detect whether the user is feeling uncomfortable or stressed.
[1305] Step 9:
[1306] The server combines the results of image analysis, audio analysis, and emotion analysis to evaluate whether the product the user is about to purchase is appropriate. As a result of the evaluation, it generates information such as whether to recommend or not to purchase the product and points to be careful about. Emotional data is also taken into consideration, and the content of notifications is adjusted as necessary.
[1307] Step 10:
[1308] The server then formats the evaluation results and sends them to the device, using a secure communication protocol to ensure data safety.
[1309] Step 11:
[1310] The device will then notify the user of the evaluation results. The notification will appear as a pop-up message on the screen and will also be announced via audio. Based on the emotional data, a gentle, encouraging message may also be displayed.
[1311] Specific examples
[1312] Scenario 1: Purchasing an item that is close to its expiration date
[1313] 1. The device captures a milk label that says "Best before: October 10, 2023."
[1314] 2. The device receives the image data and sends it to the server.
[1315] 3. The server receives the image and passes it to the generation AI.
[1316] 4. The generative AI detects the expiration date and recognizes that it is approaching.
[1317] 5. The server formats the results and sends them to the device.
[1318] 6. The device will notify the user that "This milk is nearing its expiration date. It is not recommended that you purchase it."
[1319] 7. The user receives a notification and returns the milk to its original position.
[1320] Scenario 2: The salesperson is trying to push you too hard
[1321] 1. The device captures the store clerk's statement, "Would you like some of this sweets as well?"
[1322] 2. The device receives the voice data and sends it to the server.
[1323] 3. The server receives the audio and passes it to the generation AI.
[1324] 4. Generative AI detects potential hard sales.
[1325] 5. The server formats the results and sends them to the device.
[1326] 6. The device will notify the user, "Please be careful as store clerks may be trying to force you to buy more than you need."
[1327] 7. The user receives a notification and responds to the store clerk, "I don't need it right now."
[1328] Scenario 3: Adjusting notifications based on the user's emotional state
[1329] 1. Your device uses its camera and microphone to capture your facial expressions and tone of voice.
[1330] 2. The device receives the emotion data and sends it to the server.
[1331] 3. The server uses the emotion engine to analyze the user's emotions.
[1332] 4. The emotion engine detects the user's discomfort.
[1333] 5. The server reflects the analysis results in the evaluation and adjusts the notification content.
[1334] 6. The device will then display a tailored notification to the user, saying, "We understand you're feeling annoyed. Choose only what you need."
[1335] Example 2
[1336] Next, a description will be given of Example 2. In the following description, the data processing device 12 will be referred to as a "server" and the headset type terminal 314 will be referred to as a "terminal."
[1337] When users, such as the elderly and children, purchase products, they face challenges such as difficulty in accurately assessing product information, salesperson pressure, and even their own emotional state. In particular, there is a need for systems that can quickly and accurately evaluate product quality, price, and expiration dates, detect salesperson pressure, and provide real-time notifications that take into account the user's emotional state. Furthermore, there is a lack of a means to comprehensively analyze and evaluate this information in a single system and provide appropriate advice to users.
[1338] The specific processing by the specific processing unit 290 of the data processing device 12 in the second embodiment is realized by the following means.
[1339] In this invention, the server includes means for analyzing received image data and evaluating the quality, price, and expiration date of the product, means for analyzing received voice data and detecting whether the salesperson is trying to pressure you, and means for analyzing the user's facial expression and voice tone using an emotion engine to recognize the user's emotions. This makes it possible to support elderly people and children in choosing appropriate products and shopping with peace of mind.
[1340] "User terminal" means an electronic device used by a user to capture product information and process and transmit that information to a server.
[1341] "Product information" is image data and audio data that includes information about the product's quality, price, expiration date, and so on.
[1342] A "server" is a computer system for receiving, storing, and analyzing data sent from a user terminal.
[1343] "Image data" is electronic data that contains visual information about a product, such as the product label, price, and expiration date.
[1344] "Voice data" refers to electronic data containing the content of a conversation between a salesperson and a user.
[1345] "Analysis" is the process of processing data and extracting specific information or patterns.
[1346] "Quality" is a standard for evaluating a product's condition, performance, reliability, etc.
[1347] "Price" means the amount you are willing to pay for the Goods.
[1348] "Best before date" is information indicating the expiration date of food products and the like.
[1349] "Presence or absence of hard selling" is a state that indicates whether the salesperson is trying to forcefully sell the product.
[1350] An "emotion engine" is software or hardware that analyzes a user's facial expressions and voice to recognize their emotional state.
[1351] "Notification" refers to an informational message sent to the user based on the analysis results.
[1352] This invention is a system that captures product information from a user's device and sends it to a server as image and voice data, where the server analyzes the data. It also incorporates an emotion engine that recognizes the user's emotions, and sends appropriate notifications to the user in real time based on the analysis results. This allows elderly people and children to choose appropriate products and shop with peace of mind.
[1353] System Configuration
[1354] User Device
[1355] A user device is an electronic device equipped with a camera and microphone. This includes smartphones, tablets, electronic devices, etc. A user device:
[1356] Capture of product information (image data and audio data).
[1357] Convert captured data to JPEG and WAV formats.
[1358] Sending data to the server (using the HTTPS protocol).
[1359] server
[1360] The server is a computer system that receives, stores, and analyzes data sent from user terminals. The server performs the following functions:
[1361] Check the integrity of received data and temporarily store it.
[1362] Analysis of image data (using generative AI modules).
[1363] Analysis of voice data (using generative AI modules and natural language processing techniques).
[1364] Parsing sentiment data (using the sentiment engine).
[1365] Integrating analytical results and generating evaluation results.
[1366] Sending the results to the user's device.
[1367] Specific examples of data analysis
[1368] Image data analysis
[1369] The server analyzes the image data using a generative AI module (e.g., TensorFlow or PyTorch). The generative AI extracts the product name, price, expiration date, and appearance (whether damaged or not) from the image data. For example, it analyzes a milk label that reads "Best before: October 10, 2023" and recognizes that the expiration date is approaching. Based on this, it notifies the user that "This milk's expiration date is approaching. It is not recommended that you purchase it."
[1370] Analysis of audio data
[1371] The server analyzes the voice data using automatic speech recognition (ASR) technology (e.g., Google Speech-to-Text API). The generative AI converts the voice data into text and uses natural language processing (NLP) technology (e.g., the BERT model) to analyze the content and tone of the salesperson's speech. For example, it analyzes a salesperson's statement, "Would you like to buy this candy as well?" and evaluates the possibility of a hard sell. Based on this, it notifies the customer, "The salesperson may be trying to force an unnecessary sale. Please be careful."
[1372] Emotional Data Analysis
[1373] The server uses an emotion engine (e.g., OpenCV or DeepFace) to analyze the user's facial expressions and voice tone. This allows the generative AI to recognize the user's emotional state (e.g., displeasure, joy). For example, if the user expresses displeasure, the server notifies them by saying, "It seems you are displeased. Please choose only what you need."
[1374] Prompt Sentence Examples
[1375] Below are some examples of prompt sentences:
[1376] Example prompt to detect products nearing their expiration date:
[1377] "This milk's expiration date is October 10, 2023. It is not recommended for purchase."
[1378] Example prompt to detect salespeople making hard sales pitches:
[1379] "The salesperson asked me if I wanted to buy some sweets while I was there. It might be a case of pushy sales."
[1380] In this way, the system of the present invention combines hardware such as cameras and microphones with software such as generative AI modules and emotion engines to help elderly people and children shop with peace of mind.
[1381] The flow of the identification process in the second embodiment will be described with reference to FIG.
[1382] Program processing flow
[1383] Step 1: Capture the data
[1384] The device uses a camera and microphone to capture images of the products the user picks up and audio of conversations with store staff in real time.
[1385] Input: Product images from the camera, audio data from the microphone
[1386] Output: JPEG image data, WAV audio data
[1387] Specific operation: The camera takes pictures of product labels and prices, and the microphone records conversations. The recorded data is automatically converted to JPEG and WAV format, respectively, and temporarily saved to the device's storage device.
[1388] Step 2: Sending data
[1389] The device sends the captured image and audio data to the server using the HTTPS protocol.
[1390] Input: JPEG image data, WAV audio data
[1391] Output: Encrypted data packet
[1392] Specific operation: Data is encrypted and sent to the server using a secure communication protocol (HTTPS). A hash value is generated during transmission to ensure data integrity.
[1393] Step 3: Receive data and prepare for analysis
[1394] The server receives the image data and audio data sent from the terminal and stores them in temporary storage.
[1395] Input: Encrypted data packet
[1396] Output: Image and audio data with integrity confirmed
[1397] Specific operation: Decrypts encrypted data, compares hash values to verify data integrity, and then stores the data in temporary storage (e.g., a database).
[1398] Step 4: Analyzing the image data
[1399] The server passes the image data to a generative AI module, which evaluates the product's quality, price, and expiration date.
[1400] Input: JPEG format image data
[1401] Output: Product name, price, expiration date, appearance condition evaluation result
[1402] How it works: The generative AI uses image recognition algorithms to extract information from product labels, converts it into text using OCR technology, and then uses quality assessment algorithms to evaluate the product's expiration date and appearance damage.
[1403] Step 5: Analyze the audio data
[1404] The server passes the voice data to a generation AI module to detect whether the salesperson is trying to force a sale.
[1405] Input: WAV format audio data
[1406] Output: Evaluation results of hard selling using natural language processing
[1407] Specific operations: The system converts voice data into text using automatic speech recognition (ASR) technology, and then analyzes the content and tone of the salesperson's speech using natural language processing (NLP) technology. It evaluates the likelihood of a hard sell and stores the results in a database.
[1408] Step 6: Analyze the sentiment data
[1409] The server uses an emotion engine to analyze the user's facial expressions and tone of voice to recognize the user's emotions.
[1410] Input: JPEG image data, WAV audio data
[1411] Output: Emotional state evaluation result
[1412] Specific operation: Analyzes image data from the camera and audio data from the microphone, evaluates the user's facial expressions and tone of voice using an emotion engine, analyzes whether the user is showing signs of discomfort, and records the results.
[1413] Step 7: Synthesis and evaluation of results
[1414] The server combines the results of image analysis, audio analysis, and emotion analysis to evaluate whether the product the user is about to purchase is appropriate.
[1415] Input: Image analysis results, audio analysis results, emotion analysis results
[1416] Output: Evaluation results such as purchase recommendation / non-recommendation and points to note
[1417] Specific operation: The results of each analysis are integrated and a comprehensive evaluation is performed based on a rule-based evaluation model. Information such as purchase recommendations, non-recommendations, and points to be aware of is generated and formalized as evaluation results.
[1418] Step 8: Notification of results
[1419] The server formalizes the evaluation results and sends them to the user terminal.
[1420] Input: Evaluation result
[1421] Output: Notification to user device
[1422] Specific operation: The evaluation results are re-encrypted and sent to the user's device using a secure communication protocol. The device then displays the received evaluation results as a pop-up message and announces the contents via audio.
[1423] In this way, a system can be constructed that processes data and performs data calculations at each step while providing appropriate notifications to the user in real time.
[1424] (Application example 2)
[1425] Next, a description will be given of Application Example 2. In the following description, the data processing device 12 will be referred to as a "server" and the headset type terminal 314 will be referred to as a "terminal."
[1426] When purchasing products in physical stores, certain user groups, such as the elderly and children, often find it difficult to evaluate product quality, expiration dates, and prices, and are often annoyed by salespeople's pushy sales tactics. In such situations, it is difficult for them to select the right product and they are unable to shop with peace of mind. Therefore, there is a need for a system that can comprehensively analyze the user's emotional state and detailed product information and provide appropriate notifications.
[1427] The specification process by the specification processing unit 290 of the data processing device 12 in Application Example 2 is realized by the following means. In this invention, the server includes means for capturing product information using a user terminal, means for transmitting image data and audio data of the captured product information to the server, means for analyzing the received image data in the server and evaluating the product quality, price, and expiration date, means for analyzing the received audio data in the server and detecting whether or not the salesperson is trying to pressure the user, means for sending a notification to the user based on the analysis results, and means for analyzing the user's emotions using an emotion engine in the server and adjusting the content of the notification based on the user's emotional state. This allows the user to select appropriate products with confidence and enjoy comfortable shopping.
[1428] "User terminal" refers to an electronic device that has the function of capturing product information and transmitting it to a server.
[1429] "Capture" refers to the act of acquiring image data or audio data using a camera or microphone.
[1430] "Image data" refers to data that represents product photos and labels taken with a camera in digital format.
[1431] "Audio data" refers to data that represents conversations and environmental sounds recorded by a microphone in digital form.
[1432] A "server" refers to a computer system that receives data sent from a user terminal, analyzes it, and returns the results.
[1433] "Analysis" refers to the process of evaluating and judging received image data and audio data using a program.
[1434] "Product quality" refers to the standard by which a product is evaluated based on its condition and appearance.
[1435] "Price" refers to the price of the product.
[1436] "Best before date" refers to the period during which a product such as food will retain its quality.
[1437] "Hard selling" refers to the act of a salesperson forcibly pushing a product on a customer.
[1438] An "emotion engine" is a mechanism that analyzes a user's facial expressions and tone of voice to recognize their emotions.
[1439] "Notification" refers to messages or alerts sent to users based on analysis results.
[1440] "Adjustment" refers to the process of appropriately changing the content of notifications based on analysis results and the user's emotional state.
[1441] This invention is a system that, when a user purchases a product in a physical store, uses a smart device (e.g., a smartphone or smart glasses) to capture product information, sends it to a server, which analyzes it and sends appropriate notifications.
[1442] The main flow of the system is as follows:
[1443] Data capture
[1444] The device uses a camera and microphone to capture images of the products the user picks up and audio of the conversation with the store clerk in real time. The images include product labels, expiration dates, prices, etc. The device converts the images to JPEG format and the audio to WAV format and saves them.
[1445] Sending data
[1446] The device sends the captured image and audio data to the server using a secure communication protocol (HTTPS). The data is encrypted during transmission to ensure the safety of the communication.
[1447] Receiving data and preparing for analysis
[1448] The server receives the image and audio data sent from the device and stores them in temporary storage. It checks the integrity of the received data and prepares it for analysis.
[1449] Image data analysis
[1450] The server passes the image data to the generation AI module. The generation AI performs image analysis to extract information such as the product name, price, expiration date, and appearance (damage or not) from the image data. The generation AI evaluates whether the product's expiration date is approaching, whether the product's appearance is intact, and whether the price is appropriate.
[1451] Analysis of audio data
[1452] The server passes the voice data to the generation AI module. The generation AI extracts the clerk's speech and its tone from the voice data and evaluates the possibility of hard selling using natural language processing. Based on the results of the voice data analysis, the generation AI determines whether the clerk is trying to force a sale.
[1453] Emotional Data Analysis
[1454] The server uses an emotion engine to analyze the user's facial expressions and tone of voice based on data acquired from the camera and microphone, which allows the emotion engine to recognize the user's emotions and determine whether the user is expressing discomfort.
[1455] Consolidating and communicating results
[1456] The server combines the results of image analysis, audio analysis, and emotion analysis to evaluate whether the product the user is about to purchase is appropriate. The evaluation results include information such as whether to recommend or not to purchase, and points to note. The server then formats the evaluation results and sends them to the device. A secure communication protocol is used during transmission to ensure data safety. The device then notifies the user of the received evaluation results. The notification is displayed as a pop-up message on the screen and the content is announced via audio.
[1457] Specific examples
[1458] For example, if a user picks up milk that is close to its expiration date, the device captures the label and sends it to the server, which analyzes the expiration date and notifies the user, "This milk is close to its expiration date. We do not recommend purchasing it." If a store clerk tries to pressure the user to buy more than they need, the device captures the conversation and sends it to the server, which analyzes the content and notifies the user, "The store clerk may be trying to pressure the user to buy more than they need. Please be careful."
[1459] Prompt Sentence Examples
[1460] An example of a prompt to be input to the generative AI model is as follows:
[1461] "Binary data of product images," "Binary data of conversational audio"
[1462] This allows users to choose the right product with confidence and enjoy a comfortable shopping experience.The main hardware and software used include a smart device (with camera and microphone), a secure communication protocol (HTTPS), an image analysis module, a voice analysis module, and an emotion engine.
[1463] The flow of the specific processing in the application example 2 will be described with reference to FIG.
[1464] Step 1:
[1465] When a user picks up a product, the device uses its camera to capture an image of the product. The input is the product image, and the output is a JPEG image file. Specifically, the device's camera operates to take a photo of the product label or packaging.
[1466] Step 2:
[1467] The device uses a microphone to capture the audio of the conversation with the store clerk. The input is the audio of the conversation with the store clerk, and the output is a WAV format audio file. Specifically, the device's microphone works and records the content and tone of the clerk's speech.
[1468] Step 3:
[1469] The captured image and audio data are sent to the server via a secure communication protocol (HTTPS). The input is JPEG image data and WAV audio data, and the output is a confirmation of receipt of the sent data. Specifically, the device encrypts the data and sends it to the server.
[1470] Step 4:
[1471] The server temporarily stores the received image data and audio data in storage. The input is the transmitted JPEG image data and WAV audio data, and the output is the data stored in the server's memory area. Specifically, the server checks the integrity of the received data before storing it.
[1472] Step 5:
[1473] The server passes the image data to the generation AI module, which evaluates the product's quality, price, and expiration date. The input is JPEG image data, and the output is the product name, price, expiration date, and quality evaluation results. Specifically, the generation AI analyzes the image and extracts product information.
[1474] Step 6:
[1475] The server passes the voice data to the generation AI module, which analyzes the clerk's remarks to evaluate whether or not there was a hard sell. The input is WAV-format voice data, and the output is the evaluation result of whether or not there was a hard sell. Specifically, the generation AI analyzes the voice data and detects the possibility of a hard sell from the tone and content.
[1476] Step 7:
[1477] The server uses an emotion engine to analyze the user's facial expressions and tone of voice. The input is the user's facial expression data and tone of voice, and the output is the user's emotional assessment result. Specifically, the emotion engine recognizes the user's emotions and determines the degree of discomfort or stress.
[1478] Step 8:
[1479] The server integrates the image analysis results, audio analysis results, and emotion analysis results to generate an integrated evaluation result. The input is the various analysis results, and the output is the integrated evaluation result. Specifically, the server comprehensively evaluates each analysis result and determines the appropriate notification content.
[1480] Step 9:
[1481] The server sends the integrated evaluation results to the user terminal. The input is the integrated evaluation results, and the output is notification data that arrives at the user terminal. Specifically, the server formalizes the evaluation results and sends them to the terminal using a secure communication protocol.
[1482] Step 10:
[1483] The user device then notifies the user of the received evaluation results. The input is notification data, and the output is a pop-up message and a voice notification to the user. Specifically, the device displays the notification content on the screen and communicates it to the user by voice.
[1484] Through the above steps, the present invention allows the user to select appropriate products with confidence and enjoy comfortable shopping.
[1485] The specific processing unit 290 transmits the result of the specific processing to the headset type terminal 314. In the headset type terminal 314, the control unit 46A causes the speaker 240 and the display 343 to output the result of the specific processing. The microphone 238 acquires audio indicating a user input regarding the result of the specific processing. The control unit 46A transmits audio data indicating the user input acquired by the microphone 238 to the data processing device 12. In the data processing device 12, the specific processing unit 290 acquires the audio data.
[1486] The data generation model 58 is a so-called generative AI (Artificial Intelligence). An example of the data generation model 58 is ChatGPT (Internet Search<URL: https: / / openai.com / blog / chatgpt> ), Gemini (Internet search <url: https: gemini.google.com ?hl="ja">) and other generation AIs. The data generation model 58 is obtained by performing deep learning on a neural network. A prompt including an instruction is input to the data generation model 58, and inference data such as voice data indicating voice, text data indicating text, and image data indicating an image is also input. The data generation model 58 performs inference on the input inference data in accordance with the instruction indicated by the prompt, and outputs the inference result in a data format such as voice data and text data. Here, inference refers to, for example, analysis, classification, prediction, and / or summarization.
[1487] In the above embodiment, an example was given in which the specific processing is performed by the data processing device 12, but the technology of the present disclosure is not limited to this, and the specific processing may be performed by the headset type terminal 314.
[1488] [Fourth embodiment]
[1489] FIG. 7 shows an example of the configuration of a data processing system 410 according to the fourth embodiment.
[1490] 7, a data processing system 410 includes a data processing device 12 and a robot 414. An example of the data processing device 12 is a server.
[1491] The data processing device 12 includes a computer 22, a database 24, and a communication I / F 26. The computer 22 is an example of a "computer" according to the technology of the present disclosure. The computer 22 includes a processor 28, a RAM 30, and a storage 32. The processor 28, the RAM 30, and the storage 32 are connected to a bus 34. The database 24 and the communication I / F 26 are also connected to the bus 34. The communication I / F 26 is connected to a network 54. Examples of the network 54 include a WAN (Wide Area Network) and / or a LAN (Local Area Network).
[1492] The robot 414 includes a computer 36, a microphone 238, a speaker 240, a camera 42, a communication I / F 44, and a control target 443. The computer 36 includes a processor 46, a RAM 48, and a storage 50. The processor 46, the RAM 48, and the storage 50 are connected to a bus 52. The microphone 238, the speaker 240, the camera 42, and the control target 443 are also connected to the bus 52.
[1493] The microphone 238 receives instructions and the like from the user 20 by receiving voice uttered by the user 20. The microphone 238 captures the voice uttered by the user 20, converts the captured voice into audio data, and outputs it to the processor 46. The speaker 240 outputs audio in accordance with instructions from the processor 46.
[1494] Camera 42 is a small digital camera equipped with an optical system including a lens, aperture, and shutter, and an imaging element such as a CMOS (Complementary Metal-Oxide-Semiconductor) image sensor or a CCD (Charge Coupled Device) image sensor, and captures images of the surroundings of user 20 (for example, an imaging range defined by an angle of view equivalent to the field of vision of a typical healthy person).
[1495] The communication I / F 44 is connected to a network 54. The communication I / Fs 44 and 26 are responsible for the exchange of various information between the processor 46 and the processor 28 via the network 54. The exchange of various information between the processor 46 and the processor 28 using the communication I / Fs 44 and 26 is carried out in a secure state.
[1496] The control object 443 includes a display device, LEDs in the eyes, and motors for driving the arms, hands, and feet. The posture and gestures of the robot 414 are controlled by controlling the motors of the arms, hands, and feet. Some of the emotions of the robot 414 can be expressed by controlling these motors. In addition, the facial expressions of the robot 414 can also be expressed by controlling the light emission state of the LEDs in the eyes of the robot 414.
[1497] Fig. 8 shows an example of the main functions of the data processing device 12 and the robot 414. As shown in Fig. 8, in the data processing device 12, a specific process is performed by the processor 28. A specific process program 56 is stored in the storage 32.
[1498] The specific processing program 56 is an example of a "program" according to the technology of the present disclosure. The processor 28 reads the specific processing program 56 from the storage 32 and executes the read specific processing program 56 on the RAM 30. The specific processing is realized by the processor 28 operating as a specific processing unit 290 in accordance with the specific processing program 56 executed on the RAM 30.
[1499] The storage 32 stores a data generation model 58 and an emotion identification model 59. The data generation model 58 and the emotion identification model 59 are used by the identification processing unit 290.
[1500] In the robot 414, the processor 46 performs the reception output process. A reception output program 60 is stored in the storage 50. The processor 46 reads the reception output program 60 from the storage 50 and executes the read reception output program 60 on the RAM 48. The reception output process is realized by the processor 46 operating as the control unit 46A in accordance with the reception output program 60 executed on the RAM 48.
[1501] Next, a description will be given of the specific processing performed by the specific processing unit 290 of the data processing device 12. In the following description, the data processing device 12 will be referred to as a "server" and the robot 414 will be referred to as a "terminal."
[1502] System Overview
[1503] This system supports elderly people and children in choosing appropriate products and shopping with peace of mind. Product information captured from a user's device is sent to a server as image data and audio data, which is then analyzed by the server. Based on the analysis results, the system sends appropriate notifications to the user to support their shopping.
[1504] What the program does
[1505] 1. Data capture
[1506] The device uses a camera and microphone to capture images of the products the user picks up and audio of the conversation with the store clerk in real time.
[1507] The images include product labels, expiration dates, prices, etc.
[1508] The device converts images to the appropriate format (e.g., JPEG, PNG) and saves audio in the appropriate format (e.g., WAV, MP3).
[1509] 2. Data transmission
[1510] The device sends the captured image and audio data to the server using a secure communication protocol (e.g., HTTPS), where the data is encrypted to prevent tampering during transmission.
[1511] 3. Receiving data and preparing for analysis
[1512] The server receives the image data and audio data sent from the device, and stores the received data in temporary storage.
[1513] The server checks the integrity of the received data and checks whether it can be parsed.
[1514] 4. Analysis of image data
[1515] The server passes the image data to the Generative AI module, which extracts the following information from the image:
[1516] Product name
[1517] price
[1518] expiration date
[1519] Appearance of the product (whether damaged or not)
[1520] The generative AI will use the extracted data to make the following decisions:
[1521] Is the expiration date approaching?
[1522] Is there any damage to the product's appearance?
[1523] Is the price justified?
[1524] 5. Analysis of audio data
[1525] The server passes the voice data to the generative AI module, which extracts the following information from the voice and performs natural language processing:
[1526] What the store clerk said
[1527] Tone and degree of emphasis
[1528] The generative AI will use the extracted data to make the following decisions:
[1529] Whether the salesperson is pushing unnecessary sales
[1530] 6. Synthesis and evaluation of results
[1531] The server combines the results of image analysis and audio analysis to evaluate whether the product the user is about to purchase is appropriate.
[1532] The evaluation produces the following information:
[1533] Decision to recommend or not recommend purchase
[1534] Things to note when purchasing
[1535] 7. Notification of Results
[1536] The server formats the evaluation results and sends them to the terminal.
[1537] The device will notify the user of the received evaluation results. Notification will be done as follows:
[1538] Displayed as a pop-up message on the screen
[1539] Communicate content through audio
[1540] Specific examples
[1541] Scenario 1: Purchasing an item that is close to its expiration date
[1542] 1. The device captures a milk label that says "Best before: October 10, 2023."
[1543] 2. The device receives the image data and sends it to the server.
[1544] 3. The server receives the image and passes it to the generation AI.
[1545] 4. The generative AI detects the expiration date and recognizes that it is approaching.
[1546] 5. The server formats the results and sends them to the device.
[1547] 6. The device will notify the user that "This milk is nearing its expiration date. It is not recommended that you purchase it."
[1548] 7. The user receives a notification and returns the milk to its original position.
[1549] Scenario 2: The salesperson is trying to push you too hard
[1550] 1. The device captures the store clerk's statement, "Would you like some of this sweets as well?"
[1551] 2. The device receives the voice data and sends it to the server.
[1552] 3. The server receives the audio and passes it to the generation AI.
[1553] 4. Generative AI detects potential hard sales.
[1554] 5. The server formats the results and sends them to the device.
[1555] 6. The device will notify the user, "Please be careful as store clerks may be trying to force you to buy more than you need."
[1556] 7. The user receives a notification and responds to the store clerk, "I don't need it right now."
[1557] Through the above process, the system of the present invention supports elderly people and children in selecting appropriate products and shopping with peace of mind.
[1558] The processing flow will be explained below.
[1559] Step 1:
[1560] The device uses a camera and microphone to capture images of the products the user picks up and audio of the conversation with the store clerk in real time. The images include product labels, expiration dates, prices, etc. The images are converted to JPEG format, and the audio is converted to WAV format.
[1561] Step 2:
[1562] The device sends the captured image and audio data to the server using a secure communication protocol (HTTPS). The data is encrypted during transmission to ensure the safety of the communication.
[1563] Step 3:
[1564] The server receives the image and audio data sent from the device and stores them in temporary storage. It checks the integrity of the received data and prepares it for analysis.
[1565] Step 4:
[1566] The server passes the image data to the generation AI module, which analyzes the image data to extract information such as the product name, price, expiration date, and appearance (damage).
[1567] Step 5:
[1568] Based on the extracted data, the generative AI evaluates whether the product's expiration date is approaching, whether the product's appearance is damaged, and whether the price is appropriate.
[1569] Step 6:
[1570] The server passes the voice data to a generation AI module, which analyzes the data, extracts the content and tone of the salesperson's speech, and uses natural language processing to assess the likelihood of a hard sell.
[1571] Step 7:
[1572] Based on the results of voice data analysis, the generative AI determines whether the store clerk is making unnecessary sales pitches.
[1573] Step 8:
[1574] The server combines the results of image and audio analysis to evaluate whether the product the user is about to purchase is appropriate, and generates information such as whether to recommend or not to purchase it, as well as points to be aware of.
[1575] Step 9:
[1576] The server then formats the evaluation results and sends them to the device, using a secure communication protocol to ensure data safety.
[1577] Step 10:
[1578] The device will then notify the user of the evaluation results, which will be displayed as a pop-up message on the screen and explained to them via audio.
[1579] Step 11:
[1580] The user receives a notification from the device and can decide whether to decline the purchase or respond appropriately to the store clerk. For example, if the product is close to its expiration date, the user can decline the purchase and return it to the store clerk.
[1581] Example 1
[1582] Next, a description will be given of Example 1. In the following description, the data processing device 12 will be referred to as a "server" and the robot 414 will be referred to as a "terminal."
[1583] Elderly people and children often have difficulty choosing the right products when shopping. Specifically, it can be difficult to check the quality, expiration date, and price of a product, and it can be difficult to avoid sales tactics from salespeople. In these situations, a method is needed to help them choose products safely and with peace of mind.
[1584] The specific processing by the specific processing unit 290 of the data processing device 12 in the first embodiment is realized by the following means.
[1585] In this invention, the server includes means for analyzing image data received by the server using a generative AI model to evaluate the product name, price, expiration date, and appearance of the product, means for analyzing voice data received by the server using a generative AI model to detect whether or not a salesperson is trying to pressure a user, and means for sending a notification to the user based on the analysis results, which makes it easier for the user to judge the quality, expiration date, and price of a product and to avoid unnecessary pressure from a salesperson.
[1586] A "user terminal" is a device equipped with a camera and a microphone, and is capable of capturing product information as image data and audio data and transmitting the data to a server.
[1587] "Capture" refers to the process of obtaining information about the product a user picks up as image data or audio data using a camera or microphone.
[1588] "Image data" is information that is digitally recorded using a camera to capture product labels and product appearances.
[1589] "Voice data" refers to information recorded in digital format using a microphone to record what a user or store clerk says.
[1590] A "server" is a device or system that receives image data and audio data sent from a user terminal and performs analysis processing.
[1591] A "generative AI model" is a model that uses machine learning algorithms to analyze image data and audio data and make decisions based on product information and user requests.
[1592] "Analysis" is the act of processing received data and extracting specific information.
[1593] "Product quality" refers to the criteria for evaluating a product's appearance, expiration date, price, etc.
[1594] "Hard selling" is when a salesperson pushes more product on a user than they need.
[1595] "Notifications" are messages or alerts that convey analysis results to users.
[1596] System Overview
[1597] This system supports elderly people and children in choosing appropriate products and shopping with peace of mind. Product information captured from a user's device is sent to a server as image data and audio data, which is then analyzed by the server. Based on the analysis results, the system sends appropriate notifications to the user to support their shopping.
[1598] Hardware and Software Configuration
[1599] This system consists of a user terminal, a server, and a generative AI model.
[1600] User Device
[1601] It is equipped with a camera to capture the product's label and appearance.
[1602] It is equipped with a microphone that records conversations with store staff and the user's voice commands.
[1603] Converting the captured data into an appropriate format (e.g. JPEG, PNG, WAV, MP3).
[1604] server
[1605] Receive data sent from the device via a secure communication protocol (e.g. HTTPS).
[1606] The integrity of the received data is checked and analysis processing is performed.
[1607] Data analysis is performed using generative AI models.
[1608] Data analysis details
[1609] Image data analysis
[1610] The server passes the image data to the generative AI model and extracts the product name, price, expiration date, and product appearance (whether damaged or not).
[1611] Based on the extracted data, the generative AI model determines whether the expiration date is approaching, whether the product is damaged, and whether the price is appropriate.
[1612] Analysis of audio data
[1613] The server passes the audio data to a generative AI model, which analyzes the content and tone of the clerk's speech.
[1614] Generative AI models detect potential hard sales.
[1615] Notification details
[1616] The server formalizes the evaluation results based on the analysis results and notifies the user.
[1617] The device will notify the user of the received evaluation results by displaying a pop-up on the screen or by voice.
[1618] Specific examples
[1619] Scenario 1: Purchasing an item that is close to its expiration date
[1620] 1. A user picks up a bottle of milk that says "Best before: October 10, 2023."
[1621] 2. The device captures the image of the milk label and sends the data to the server.
[1622] 3. The server analyzes the image data, and the generative AI model recognizes that the expiration date is approaching.
[1623] 4. The server formats the results and sends a notification to the device saying, "This milk is nearing its expiration date. It is not recommended for purchase."
[1624] 5. The user receives a notification and returns the milk to its original position.
[1625] Scenario 2: The salesperson is trying to push you too hard
[1626] 1. The user hears the store clerk say, "Would you like some of this candy as well?"
[1627] 2. The device captures this speech and sends the audio data to the server.
[1628] 3. The server analyzes the voice data and a generative AI model detects potential hard sales.
[1629] 4. The server formats the results and sends a notification to the device saying, "The store clerk may be trying to force you to buy something. Please be careful."
[1630] 5. The user receives a notification and replies to the store clerk, "I don't need it right now."
[1631] By using generative AI models for analysis, users can accurately judge the quality of products and whether or not store clerks are trying to push products. This system allows elderly people and children to shop with peace of mind by selecting appropriate products.
[1632] The flow of the identification process in the first embodiment will be described with reference to FIG.
[1633] Step 1: Capture the data
[1634] How it works: Uses the device's camera and microphone to capture information about the product the user picks up.
[1635] Input: Images and audio of the product picked up by the user (e.g., product label, conversation with the store clerk).
[1636] Processing: The device uses its camera to capture the product label (product name, price, expiration date, etc.) and uses a microphone to record the conversation with the store clerk and the user's voice commands. The image data is converted to JPEG or PNG format, and the voice data is converted to WAV or MP3 format.
[1637] Output: Image data and audio data are generated.
[1638] Step 2: Sending data
[1639] Operation: The device sends the captured image and audio data to the server.
[1640] Input: Image and audio data generated in step 1.
[1641] Processing: The device sends encrypted image and audio data to the server using a secure communication protocol such as HTTPS.
[1642] Output: Securely transmitted image and audio data.
[1643] Step 3: Receive data and prepare for analysis
[1644] Operation: The server receives image data and audio data sent from the device and temporarily stores them in storage.
[1645] Input: Image and audio data sent from the device.
[1646] Processing: When the server receives the data, it checks its integrity and whether it can be analyzed. If there are no problems, it temporarily stores it in storage.
[1647] Output: Data ready for analysis.
[1648] Step 4: Analyzing the image data
[1649] How it works: The server passes image data to the generative AI model and extracts product information.
[1650] Input: Image data stored in storage.
[1651] Processing: The generative AI model analyzes the image data and extracts the following information: product name, price, expiration date, and product appearance (damaged or not).
[1652] Output: Extracted product information (e.g. product name, price, expiry date, product appearance).
[1653] Step 5: Analyze the audio data
[1654] How it works: The server passes the audio data to a generative AI model, which analyzes the content and tone of what the clerk is saying.
[1655] Input: Audio data stored in storage.
[1656] Processing: A generative AI model analyzes the audio data and extracts information about the salesperson's speech, including their content, tone, and emphasis. Based on this, it can detect potential pressure.
[1657] Output: Extracted audio information and a judgment on whether there was a hard sell.
[1658] Step 6: Synthesis and evaluation of results
[1659] Operation: The server integrates the image analysis results and the audio analysis results to evaluate whether or not to purchase the item.
[1660] Input: Extracted product information and audio information.
[1661] Processing: The server combines the results of image analysis (e.g., expiration date, product appearance) and voice analysis (e.g., whether there is any hard sell) to evaluate whether the product the user is trying to purchase is appropriate.
[1662] Output: Evaluation result (e.g., recommended / not recommended for purchase, points to note).
[1663] Step 7: Notification of results
[1664] Operation: The server formats the evaluation results and sends them to the terminal.
[1665] Input: Consolidated evaluation results.
[1666] Processing: The server converts the evaluation results into a format that is easy for the user to understand and sends them to the terminal. The terminal receives the evaluation results and notifies the user.
[1667] Output: A notification message that is displayed to the user and / or an audio notification (e.g., "This milk is nearing its expiration date. It is not recommended for purchase.").
[1668] The above is the specific processing flow of this system. By performing appropriate data processing and calculations at each step, users can select products and shop with confidence.
[1669] (Application example 1)
[1670] Next, a description will be given of Application Example 1. In the following description, the data processing device 12 will be referred to as a "server" and the robot 414 will be referred to as a "terminal."
[1671] There is a challenge for certain user groups, such as the elderly and children, to choose appropriate products and shop without anxiety. In particular, there is a risk that they may have difficulty determining the quality, expiration date, price, etc. of a product, or that they may be pressured by store clerks into buying unnecessary products. This creates an environment in which it is difficult to enjoy shopping with peace of mind.
[1672] The specific processing by the specific processing unit 290 of the data processing device 12 in the application example 1 is realized by the following means.
[1673] In this invention, the server includes means for capturing product information using a user terminal, means for transmitting image data and audio data of the captured product information to the server, means for analyzing the received image data in the server to evaluate the quality, price, and expiration date of the product, means for analyzing the received audio data in the server to detect whether the salesperson is trying to pressure the user, means for sending a notification to the user based on the analysis results, and means for capturing and notifying the user of the product information using an application installed on a smartphone. This allows users such as the elderly and children to select appropriate products and shop with peace of mind.
[1674] A "user terminal" is a portable computing device such as a smartphone or tablet.
[1675] "Capture" is the action of acquiring image data or audio data using a camera or microphone.
[1676] "Image data" is visual information captured by a camera or other device expressed in data format.
[1677] "Audio data" is sound information recorded by a microphone or the like expressed in data format.
[1678] A "server" is a computing device that receives data from user terminals over a network and analyzes and processes the data.
[1679] "Analysis" is the process of examining acquired data in detail for evaluation or judgment.
[1680] "Quality" is a property that describes the condition and performance of a product.
[1681] "Price" is the amount paid for a product.
[1682] The "best before" date indicates the period during which a product is safe and delicious to eat.
[1683] "Hard selling" is when a salesperson tries to force you into buying a product.
[1684] "Notifications" are messages or alerts that communicate analysis results to users.
[1685] An "application" is a software program that is installed on a user device, such as a smartphone or tablet, and performs a specific function.
[1686] In order to support specific user groups such as the elderly and children in selecting appropriate products and shopping with peace of mind, the system of the present invention is configured by the following means.
[1687] The system includes a user terminal, a server, and a program for linking them. The user terminal can be a smartphone or tablet. The user terminal is equipped with a camera and microphone to capture product information as image and audio data. The captured data is then sent to the server using a secure protocol.
[1688] The server analyzes the received image and audio data to evaluate the product's quality, price, and expiration date. This evaluation uses image and audio analysis technologies. Image analysis evaluates the product's label, expiration date, price, and appearance (whether damaged or not). Deep learning models and generative AI models are used for this. For example, machine learning frameworks such as TensorFlow and PyTorch can be used.
[1689] Voice data analysis analyzes conversations with store clerks to detect whether or not there is any hard sell. Natural language processing (NLP) technology, such as Google's Dialogflow or OpenAI's GPT model, is used to convert the voice file into text, and the likelihood of hard sell is assessed based on that text.
[1690] The analysis results are integrated and notifications are sent to users based on the evaluation. Notifications are sent through an application installed on the smartphone, and users can check the results via pop-up messages or voice notifications. Based on the analysis results, purchase recommendations or non-recommendations are also notified.
[1691] Specific use cases include the following scenarios:
[1692] Scenario 1: Purchasing an item that is close to its expiration date
[1693] 1. The user device captures a product label that displays "Best before: October 10, 2023."
[1694] 2. The device sends the captured image data to the server.
[1695] 3. The server analyzes the image, and the generative AI model detects the expiration date and recognizes that it is approaching.
[1696] 4. The server sends the results to the device.
[1697] 5. The device will notify the user that "This product's expiration date is approaching. It is not recommended that you purchase it."
[1698] Scenario 2: The salesperson is trying to push you too hard
[1699] 1. The user device captures the store clerk's statement, "Would you like some of this sweets as well?"
[1700] 2. The device sends the captured audio data to the server.
[1701] 3. The server analyzes the audio and a generative AI model detects potential hard sales attempts.
[1702] 4. The server sends the results to the device.
[1703] 5. The device will notify the user, "Please be careful as store clerks may be trying to force you to buy more than you need."
[1704] Prompt Sentence Examples
[1705] Image data: Product label image showing the expiration date "October 10, 2023"
[1706] Audio data: A recording of the store clerk saying, "Would you like some of these sweets as well?"
[1707] Analysis results:
[1708] 1. This product is nearing its expiration date and is not recommended for purchase.
[1709] 2. Be careful, as store clerks may try to pressure you into buying something you don't need.
[1710] In this way, the system of the present invention supports elderly people and children in choosing appropriate products and shopping with peace of mind.
[1711] The flow of the specific processing in the application example 1 will be described with reference to FIG.
[1712] Step 1:
[1713] The user picks up a product, captures the product label with the smartphone camera, and records the conversation with the store clerk using the microphone. The input is the image data captured by the camera and the audio data recorded by the microphone. The output is an image file (e.g., JPEG) and an audio file (e.g., WAV).
[1714] Step 2:
[1715] The device encrypts the captured image and audio data and sends them to the server using a secure communication protocol (e.g., HTTPS) to maintain security. The input is an image file and an audio file, and the output is an encrypted data packet.
[1716] Step 3:
[1717] The server receives the encrypted data sent from the terminal, decrypts it, and temporarily stores it in storage. The input is the encrypted data packet, and the output is the decrypted image data and audio data.
[1718] Step 4:
[1719] The server passes the image data to a generative AI model, which analyzes the product's quality, price, and expiration date. The input is the decoded image data, and the output is the analysis results for the product's quality, price, and expiration date. For example, the generative AI model reads the expiration date label from the image and determines whether the expiration date is approaching.
[1720] Step 5:
[1721] The server passes the audio data to a generative AI model, which analyzes the content and tone of the salesperson's speech. The input is the decoded audio data, and the output is an analysis of the salesperson's likelihood of aggressive sales. For example, the generative AI model converts the audio file into text and detects aggressive sales cues from the text.
[1722] Step 6:
[1723] The server integrates the results of image and audio analysis and evaluates the appropriate notification content for the user. The input is the analysis results of the image and audio data, and the output is the notification content. For example, notification content such as "This product's expiration date is approaching. We do not recommend purchasing it" or "The store clerk may be trying to pressure you into buying something unnecessarily" may be generated.
[1724] Step 7:
[1725] The server then formattes the evaluation results and sends them to the user's device. The input is the consolidated analysis result, and the output is a formatted notification message. For example, the server converts the results into JSON format and sends them to the device using a secure protocol.
[1726] Step 8:
[1727] The device receives the notification from the server and displays it to the user as a pop-up message or a sound notification. The input is the notification message sent from the server, and the output is the notification displayed to the user. For example, "This product is nearing its expiration date. It is not recommended to purchase it."
[1728] By checking this notification, users can choose the right product and shop with peace of mind.
[1729] Furthermore, an emotion engine that estimates the user's emotion may be further combined. That is, the identification processing unit 290 may estimate the user's emotion using the emotion identification model 59, and perform identification processing using the user's emotion.
[1730] System Overview
[1731] This system captures product information from a user's device and sends it to a server as image data or voice data, where the server analyzes the data. It also incorporates an emotion engine that recognizes the user's emotions, and sends appropriate notifications to the user in real time based on the analysis results. This allows elderly people and children to choose appropriate products and shop with peace of mind.
[1732] What the program does
[1733] 1. Data capture
[1734] The device uses a camera and microphone to capture images of the products the user picks up and audio of the conversation with the store associate in real time, including product labels, expiration dates, and prices.
[1735] The device converts and saves images in JPEG format and audio in WAV format.
[1736] 2. Data transmission
[1737] The device sends the captured image and audio data to the server using a secure communication protocol (HTTPS). The data is encrypted during transmission to ensure the safety of the communication.
[1738] 3. Receiving data and preparing for analysis
[1739] The server receives the image and audio data sent from the device and stores them in temporary storage. It checks the integrity of the received data and prepares it for analysis.
[1740] 4. Analysis of image data
[1741] The server passes the image data to the generation AI module, which analyzes the image data to extract information such as the product name, price, expiration date, and appearance (damage).
[1742] The generative AI evaluates whether the product is nearing its expiration date, whether the product's appearance is damaged, and whether the price is appropriate.
[1743] 5. Analysis of audio data
[1744] The server passes the voice data to a generation AI module, which extracts the content and tone of the salesperson's speech from the voice data and uses natural language processing to evaluate the likelihood of hard sales.
[1745] Based on the results of voice data analysis, the generative AI determines whether the store clerk is making unnecessary sales pitches.
[1746] 6. Emotion Data Analysis
[1747] The server uses an emotion engine to analyze the user's facial expressions and tone of voice based on data acquired from the camera and microphone, which allows the emotion engine to recognize the user's emotions and determine whether the user is expressing discomfort.
[1748] 7. Synthesis and evaluation of results
[1749] The server combines the results of image analysis, audio analysis, and emotion analysis to evaluate the suitability of the product the user is about to purchase, and generates information such as whether to recommend or not to purchase it, as well as points to be aware of.
[1750] 8. Notification of Results
[1751] The server then formats the evaluation results and sends them to the device, using a secure communication protocol to ensure data safety.
[1752] The device will then notify the user of the evaluation results, which will be displayed as a pop-up message on the screen and explained to them via audio.
[1753] Specific examples
[1754] Scenario 1: Purchasing an item that is close to its expiration date
[1755] 1. The device captures a milk label that says "Best before: October 10, 2023."
[1756] 2. The device receives the image data and sends it to the server.
[1757] 3. The server receives the image and passes it to the generation AI.
[1758] 4. The generative AI detects the expiration date and recognizes that it is approaching.
[1759] 5. The server formats the results and sends them to the device.
[1760] 6. The device will notify the user that "This milk is nearing its expiration date. It is not recommended that you purchase it."
[1761] 7. The user receives a notification and returns the milk to its original position.
[1762] Scenario 2: The salesperson is trying to push you too hard
[1763] 1. The device captures the store clerk's statement, "Would you like some of this sweets as well?"
[1764] 2. The device receives the voice data and sends it to the server.
[1765] 3. The server receives the audio and passes it to the generation AI.
[1766] 4. Generative AI detects potential hard sales.
[1767] 5. The server formats the results and sends them to the device.
[1768] 6. The device will notify the user, "Please be careful as store clerks may be trying to force you to buy more than you need."
[1769] 7. The user receives a notification and responds to the store clerk, "I don't need it right now."
[1770] Scenario 3: Adjusting notifications based on the user's emotional state
[1771] 1. Your device uses its camera and microphone to capture your facial expressions and tone of voice.
[1772] 2. The device receives the emotion data and sends it to the server.
[1773] 3. The server uses the emotion engine to analyze the user's emotions.
[1774] 4. The emotion engine detects the user's discomfort.
[1775] 5. The server reflects the analysis results in the evaluation and adjusts the notification content.
[1776] 6. The device will then display a tailored notification to the user, saying, "We understand you're feeling annoyed. Choose only what you need."
[1777] Through the above process, the system of the present invention supports elderly people and children in selecting appropriate products and shopping with peace of mind. The introduction of an emotion engine enables more precise support based on the user's emotional state.
[1778] The processing flow will be explained below.
[1779] Step 1:
[1780] The device uses a camera and microphone to capture images of the products the user picks up and audio of the conversation with the store clerk in real time. The images include the product label, price, expiration date, etc. The image data is converted to JPEG format, and the audio data is converted to WAV format.
[1781] Step 2:
[1782] The device sends the captured image and audio data to the server using a secure communication protocol (HTTPS). The data is encrypted during transmission, protecting it from unauthorized access.
[1783] Step 3:
[1784] The server receives the image and audio data sent from the device and stores them in temporary storage. It checks the integrity and completeness of the received data and prepares it for analysis.
[1785] Step 4:
[1786] The server passes the image data to the generation AI module, which analyzes the image and extracts information such as the product name, price, expiration date, and appearance (whether damaged or not).
[1787] Step 5:
[1788] Based on the extracted data, the generative AI evaluates whether the product's expiration date is approaching, whether the product's appearance is damaged, and whether the price is appropriate.
[1789] Step 6:
[1790] The server passes the voice data to the generative AI module, which analyzes the voice and extracts the content and tone of the clerk's speech. Natural language processing (NLP) is also performed on the speech.
[1791] Step 7:
[1792] The AI generator uses voice analysis to determine whether a salesperson is trying to push a customer too hard, and tone analysis is also taken into account.
[1793] Step 8:
[1794] The server passes data captured by the camera and microphone to the emotion engine, which analyzes the user's facial expressions and tone of voice to detect whether the user is feeling uncomfortable or stressed.
[1795] Step 9:
[1796] The server combines the results of image analysis, audio analysis, and emotion analysis to evaluate whether the product the user is about to purchase is appropriate. As a result of the evaluation, it generates information such as whether to recommend or not to purchase the product and points to be careful about. Emotional data is also taken into consideration, and the content of notifications is adjusted as necessary.
[1797] Step 10:
[1798] The server then formats the evaluation results and sends them to the device, using a secure communication protocol to ensure data safety.
[1799] Step 11:
[1800] The device will then notify the user of the evaluation results. The notification will appear as a pop-up message on the screen and will also be announced via audio. Based on the emotional data, a gentle, encouraging message may also be displayed.
[1801] Specific examples
[1802] Scenario 1: Purchasing an item that is close to its expiration date
[1803] 1. The device captures a milk label that says "Best before: October 10, 2023."
[1804] 2. The device receives the image data and sends it to the server.
[1805] 3. The server receives the image and passes it to the generation AI.
[1806] 4. The generative AI detects the expiration date and recognizes that it is approaching.
[1807] 5. The server formats the results and sends them to the device.
[1808] 6. The device will notify the user that "This milk is nearing its expiration date. It is not recommended that you purchase it."
[1809] 7. The user receives a notification and returns the milk to its original position.
[1810] Scenario 2: The salesperson is trying to push you too hard
[1811] 1. The device captures the store clerk's statement, "Would you like some of this sweets as well?"
[1812] 2. The device receives the voice data and sends it to the server.
[1813] 3. The server receives the audio and passes it to the generation AI.
[1814] 4. Generative AI detects potential hard sales.
[1815] 5. The server formats the results and sends them to the device.
[1816] 6. The device will notify the user, "Please be careful as store clerks may be trying to force you to buy more than you need."
[1817] 7. The user receives a notification and responds to the store clerk, "I don't need it right now."
[1818] Scenario 3: Adjusting notifications based on the user's emotional state
[1819] 1. Your device uses its camera and microphone to capture your facial expressions and tone of voice.
[1820] 2. The device receives the emotion data and sends it to the server.
[1821] 3. The server uses the emotion engine to analyze the user's emotions.
[1822] 4. The emotion engine detects the user's discomfort.
[1823] 5. The server reflects the analysis results in the evaluation and adjusts the notification content.
[1824] 6. The device will then display a tailored notification to the user, saying, "We understand you're feeling annoyed. Choose only what you need."
[1825] Example 2
[1826] Next, a description will be given of Example 2. In the following description, the data processing device 12 will be referred to as a "server" and the robot 414 will be referred to as a "terminal."
[1827] When users, such as the elderly and children, purchase products, they face challenges such as difficulty in accurately assessing product information, salesperson pressure, and even their own emotional state. In particular, there is a need for systems that can quickly and accurately evaluate product quality, price, and expiration dates, detect salesperson pressure, and provide real-time notifications that take into account the user's emotional state. Furthermore, there is a lack of a means to comprehensively analyze and evaluate this information in a single system and provide appropriate advice to users.
[1828] The specific processing by the specific processing unit 290 of the data processing device 12 in the second embodiment is realized by the following means.
[1829] In this invention, the server includes means for analyzing received image data and evaluating the quality, price, and expiration date of the product, means for analyzing received voice data and detecting whether the salesperson is trying to pressure you, and means for analyzing the user's facial expression and voice tone using an emotion engine to recognize the user's emotions. This makes it possible to support elderly people and children in choosing appropriate products and shopping with peace of mind.
[1830] "User terminal" means an electronic device used by a user to capture product information and process and transmit that information to a server.
[1831] "Product information" is image data and audio data that includes information about the product's quality, price, expiration date, and so on.
[1832] A "server" is a computer system for receiving, storing, and analyzing data sent from a user terminal.
[1833] "Image data" is electronic data that contains visual information about a product, such as the product label, price, and expiration date.
[1834] "Voice data" refers to electronic data containing the content of a conversation between a salesperson and a user.
[1835] "Analysis" is the process of processing data and extracting specific information or patterns.
[1836] "Quality" is a standard for evaluating a product's condition, performance, reliability, etc.
[1837] "Price" means the amount you are willing to pay for the Goods.
[1838] "Best before date" is information indicating the expiration date of food products and the like.
[1839] "Presence or absence of hard selling" is a state that indicates whether the salesperson is trying to forcefully sell the product.
[1840] An "emotion engine" is software or hardware that analyzes a user's facial expressions and voice to recognize their emotional state.
[1841] "Notification" refers to an informational message sent to the user based on the analysis results.
[1842] This invention is a system that captures product information from a user's device and sends it to a server as image and voice data, where the server analyzes the data. It also incorporates an emotion engine that recognizes the user's emotions, and sends appropriate notifications to the user in real time based on the analysis results. This allows elderly people and children to choose appropriate products and shop with peace of mind.
[1843] System Configuration
[1844] User Device
[1845] A user device is an electronic device equipped with a camera and microphone. This includes smartphones, tablets, electronic devices, etc. A user device:
[1846] Capture of product information (image data and audio data).
[1847] Convert captured data to JPEG and WAV formats.
[1848] Sending data to the server (using the HTTPS protocol).
[1849] server
[1850] The server is a computer system that receives, stores, and analyzes data sent from user terminals. The server performs the following functions:
[1851] Check the integrity of received data and temporarily store it.
[1852] Analysis of image data (using generative AI modules).
[1853] Analysis of voice data (using generative AI modules and natural language processing techniques).
[1854] Parsing sentiment data (using the sentiment engine).
[1855] Integrating analytical results and generating evaluation results.
[1856] Sending the results to the user's device.
[1857] Specific examples of data analysis
[1858] Image data analysis
[1859] The server analyzes the image data using a generative AI module (e.g., TensorFlow or PyTorch). The generative AI extracts the product name, price, expiration date, and appearance (whether damaged or not) from the image data. For example, it analyzes a milk label that reads "Best before: October 10, 2023" and recognizes that the expiration date is approaching. Based on this, it notifies the user that "This milk's expiration date is approaching. It is not recommended that you purchase it."
[1860] Analysis of audio data
[1861] The server analyzes the voice data using automatic speech recognition (ASR) technology (e.g., Google Speech-to-Text API). The generative AI converts the voice data into text and uses natural language processing (NLP) technology (e.g., the BERT model) to analyze the content and tone of the salesperson's speech. For example, it analyzes a salesperson's statement, "Would you like to buy this candy as well?" and evaluates the possibility of a hard sell. Based on this, it notifies the customer, "The salesperson may be trying to force an unnecessary sale. Please be careful."
[1862] Emotional Data Analysis
[1863] The server uses an emotion engine (e.g., OpenCV or DeepFace) to analyze the user's facial expressions and voice tone. This allows the generative AI to recognize the user's emotional state (e.g., displeasure, joy). For example, if the user expresses displeasure, the server notifies them by saying, "It seems you are displeased. Please choose only what you need."
[1864] Prompt Sentence Examples
[1865] Below are some examples of prompt sentences:
[1866] Example prompt to detect products nearing their expiration date:
[1867] "This milk's expiration date is October 10, 2023. It is not recommended for purchase."
[1868] Example prompt to detect salespeople making hard sales pitches:
[1869] "The salesperson asked me if I wanted to buy some sweets while I was there. It might be a case of pushy sales."
[1870] In this way, the system of the present invention combines hardware such as cameras and microphones with software such as generative AI modules and emotion engines to help elderly people and children shop with peace of mind.
[1871] The flow of the identification process in the second embodiment will be described with reference to FIG.
[1872] Program processing flow
[1873] Step 1: Capture the data
[1874] The device uses a camera and microphone to capture images of the products the user picks up and audio of conversations with store staff in real time.
[1875] Input: Product images from the camera, audio data from the microphone
[1876] Output: JPEG image data, WAV audio data
[1877] Specific operation: The camera takes pictures of product labels and prices, and the microphone records conversations. The recorded data is automatically converted to JPEG and WAV format, respectively, and temporarily saved to the device's storage device.
[1878] Step 2: Sending data
[1879] The device sends the captured image and audio data to the server using the HTTPS protocol.
[1880] Input: JPEG image data, WAV audio data
[1881] Output: Encrypted data packet
[1882] Specific operation: Data is encrypted and sent to the server using a secure communication protocol (HTTPS). A hash value is generated during transmission to ensure data integrity.
[1883] Step 3: Receive data and prepare for analysis
[1884] The server receives the image data and audio data sent from the terminal and stores them in temporary storage.
[1885] Input: Encrypted data packet
[1886] Output: Image and audio data with integrity confirmed
[1887] Specific operation: Decrypts encrypted data, compares hash values to verify data integrity, and then stores the data in temporary storage (e.g., a database).
[1888] Step 4: Analyzing the image data
[1889] The server passes the image data to a generative AI module, which evaluates the product's quality, price, and expiration date.
[1890] Input: JPEG format image data
[1891] Output: Product name, price, expiration date, appearance condition evaluation result
[1892] How it works: The generative AI uses image recognition algorithms to extract information from product labels, converts it into text using OCR technology, and then uses quality assessment algorithms to evaluate the product's expiration date and appearance damage.
[1893] Step 5: Analyze the audio data
[1894] The server passes the voice data to a generation AI module to detect whether the salesperson is trying to force a sale.
[1895] Input: WAV format audio data
[1896] Output: Evaluation results of hard selling using natural language processing
[1897] Specific operations: The system converts voice data into text using automatic speech recognition (ASR) technology, and then analyzes the content and tone of the salesperson's speech using natural language processing (NLP) technology. It evaluates the likelihood of a hard sell and stores the results in a database.
[1898] Step 6: Analyze the sentiment data
[1899] The server uses an emotion engine to analyze the user's facial expressions and tone of voice to recognize the user's emotions.
[1900] Input: JPEG image data, WAV audio data
[1901] Output: Emotional state evaluation result
[1902] Specific operation: Analyzes image data from the camera and audio data from the microphone, evaluates the user's facial expressions and tone of voice using an emotion engine, analyzes whether the user is showing signs of discomfort, and records the results.
[1903] Step 7: Synthesis and evaluation of results
[1904] The server combines the results of image analysis, audio analysis, and emotion analysis to evaluate whether the product the user is about to purchase is appropriate.
[1905] Input: Image analysis results, audio analysis results, emotion analysis results
[1906] Output: Evaluation results such as purchase recommendation / non-recommendation and points to note
[1907] Specific operation: The results of each analysis are integrated and a comprehensive evaluation is performed based on a rule-based evaluation model. Information such as purchase recommendations, non-recommendations, and points to be aware of is generated and formalized as evaluation results.
[1908] Step 8: Notification of results
[1909] The server formalizes the evaluation results and sends them to the user terminal.
[1910] Input: Evaluation result
[1911] Output: Notification to user device
[1912] Specific operation: The evaluation results are re-encrypted and sent to the user's device using a secure communication protocol. The device then displays the received evaluation results as a pop-up message and announces the contents via audio.
[1913] In this way, a system can be constructed that processes data and performs data calculations at each step while providing appropriate notifications to the user in real time.
[1914] (Application example 2)
[1915] Next, a description will be given of Application Example 2. In the following description, the data processing device 12 will be referred to as a "server" and the robot 414 will be referred to as a "terminal."
[1916] When purchasing products in physical stores, certain user groups, such as the elderly and children, often find it difficult to evaluate product quality, expiration dates, and prices, and are often annoyed by salespeople's pushy sales tactics. In such situations, it is difficult for them to select the right product and they are unable to shop with peace of mind. Therefore, there is a need for a system that can comprehensively analyze the user's emotional state and detailed product information and provide appropriate notifications.
[1917] The specification process by the specification processing unit 290 of the data processing device 12 in Application Example 2 is realized by the following means. In this invention, the server includes means for capturing product information using a user terminal, means for transmitting image data and audio data of the captured product information to the server, means for analyzing the received image data in the server and evaluating the product quality, price, and expiration date, means for analyzing the received audio data in the server and detecting whether or not the salesperson is trying to pressure the user, means for sending a notification to the user based on the analysis results, and means for analyzing the user's emotions using an emotion engine in the server and adjusting the content of the notification based on the user's emotional state. This allows the user to select appropriate products with confidence and enjoy comfortable shopping.
[1918] "User terminal" refers to an electronic device that has the function of capturing product information and transmitting it to a server.
[1919] "Capture" refers to the act of acquiring image data or audio data using a camera or microphone.
[1920] "Image data" refers to data that represents product photos and labels taken with a camera in digital format.
[1921] "Audio data" refers to data that represents conversations and environmental sounds recorded by a microphone in digital form.
[1922] A "server" refers to a computer system that receives data sent from a user terminal, analyzes it, and returns the results.
[1923] "Analysis" refers to the process of evaluating and judging received image data and audio data using a program.
[1924] "Product quality" refers to the standard by which a product is evaluated based on its condition and appearance.
[1925] "Price" refers to the price of the product.
[1926] "Best before date" refers to the period during which a product such as food will retain its quality.
[1927] "Hard selling" refers to the act of a salesperson forcibly pushing a product on a customer.
[1928] An "emotion engine" is a mechanism that analyzes a user's facial expressions and tone of voice to recognize their emotions.
[1929] "Notification" refers to messages or alerts sent to users based on analysis results.
[1930] "Adjustment" refers to the process of appropriately changing the content of notifications based on analysis results and the user's emotional state.
[1931] This invention is a system that, when a user purchases a product in a physical store, uses a smart device (e.g., a smartphone or smart glasses) to capture product information, sends it to a server, which analyzes it and sends appropriate notifications.
[1932] The main flow of the system is as follows:
[1933] Data capture
[1934] The device uses a camera and microphone to capture images of the products the user picks up and audio of the conversation with the store clerk in real time. The images include product labels, expiration dates, prices, etc. The device converts the images to JPEG format and the audio to WAV format and saves them.
[1935] Sending data
[1936] The device sends the captured image and audio data to the server using a secure communication protocol (HTTPS). The data is encrypted during transmission to ensure the safety of the communication.
[1937] Receiving data and preparing for analysis
[1938] The server receives the image and audio data sent from the device and stores them in temporary storage. It checks the integrity of the received data and prepares it for analysis.
[1939] Image data analysis
[1940] The server passes the image data to the generation AI module. The generation AI performs image analysis to extract information such as the product name, price, expiration date, and appearance (damage or not) from the image data. The generation AI evaluates whether the product's expiration date is approaching, whether the product's appearance is intact, and whether the price is appropriate.
[1941] Analysis of audio data
[1942] The server passes the voice data to the generation AI module. The generation AI extracts the clerk's speech and its tone from the voice data and evaluates the possibility of hard selling using natural language processing. Based on the results of the voice data analysis, the generation AI determines whether the clerk is trying to force a sale.
[1943] Emotional Data Analysis
[1944] The server uses an emotion engine to analyze the user's facial expressions and tone of voice based on data acquired from the camera and microphone, which allows the emotion engine to recognize the user's emotions and determine whether the user is expressing discomfort.
[1945] Consolidating and communicating results
[1946] The server combines the results of image analysis, audio analysis, and emotion analysis to evaluate whether the product the user is about to purchase is appropriate. The evaluation results include information such as whether to recommend or not to purchase, and points to note. The server then formats the evaluation results and sends them to the device. A secure communication protocol is used during transmission to ensure data safety. The device then notifies the user of the received evaluation results. The notification is displayed as a pop-up message on the screen and the content is announced via audio.
[1947] Specific examples
[1948] For example, if a user picks up milk that is close to its expiration date, the device captures the label and sends it to the server, which analyzes the expiration date and notifies the user, "This milk is close to its expiration date. We do not recommend purchasing it." If a store clerk tries to pressure the user to buy more than they need, the device captures the conversation and sends it to the server, which analyzes the content and notifies the user, "The store clerk may be trying to pressure the user to buy more than they need. Please be careful."
[1949] Prompt Sentence Examples
[1950] An example of a prompt to be input to the generative AI model is as follows:
[1951] "Binary data of product images," "Binary data of conversational audio"
[1952] This allows users to choose the right product with confidence and enjoy a comfortable shopping experience.The main hardware and software used include a smart device (with camera and microphone), a secure communication protocol (HTTPS), an image analysis module, a voice analysis module, and an emotion engine.
[1953] The flow of the specific processing in the application example 2 will be described with reference to FIG.
[1954] Step 1:
[1955] When a user picks up a product, the device uses its camera to capture an image of the product. The input is the product image, and the output is a JPEG image file. Specifically, the device's camera operates to take a photo of the product label or packaging.
[1956] Step 2:
[1957] The device uses a microphone to capture the audio of the conversation with the store clerk. The input is the audio of the conversation with the store clerk, and the output is a WAV format audio file. Specifically, the device's microphone works and records the content and tone of the clerk's speech.
[1958] Step 3:
[1959] The captured image and audio data are sent to the server via a secure communication protocol (HTTPS). The input is JPEG image data and WAV audio data, and the output is a confirmation of receipt of the sent data. Specifically, the device encrypts the data and sends it to the server.
[1960] Step 4:
[1961] The server temporarily stores the received image data and audio data in storage. The input is the transmitted JPEG image data and WAV audio data, and the output is the data stored in the server's memory area. Specifically, the server checks the integrity of the received data before storing it.
[1962] Step 5:
[1963] The server passes the image data to the generation AI module, which evaluates the product's quality, price, and expiration date. The input is JPEG image data, and the output is the product name, price, expiration date, and quality evaluation results. Specifically, the generation AI analyzes the image and extracts product information.
[1964] Step 6:
[1965] The server passes the voice data to the generation AI module, which analyzes the clerk's remarks to evaluate whether or not there was a hard sell. The input is WAV-format voice data, and the output is the evaluation result of whether or not there was a hard sell. Specifically, the generation AI analyzes the voice data and detects the possibility of a hard sell from the tone and content.
[1966] Step 7:
[1967] The server uses an emotion engine to analyze the user's facial expressions and tone of voice. The input is the user's facial expression data and tone of voice, and the output is the user's emotional assessment result. Specifically, the emotion engine recognizes the user's emotions and determines the degree of discomfort or stress.
[1968] Step 8:
[1969] The server integrates the image analysis results, audio analysis results, and emotion analysis results to generate an integrated evaluation result. The input is the various analysis results, and the output is the integrated evaluation result. Specifically, the server comprehensively evaluates each analysis result and determines the appropriate notification content.
[1970] Step 9:
[1971] The server sends the integrated evaluation results to the user terminal. The input is the integrated evaluation results, and the output is notification data that arrives at the user terminal. Specifically, the server formalizes the evaluation results and sends them to the terminal using a secure communication protocol.
[1972] Step 10:
[1973] The user device then notifies the user of the received evaluation results. The input is notification data, and the output is a pop-up message and a voice notification to the user. Specifically, the device displays the notification content on the screen and communicates it to the user by voice.
[1974] Through the above steps, the present invention allows the user to select appropriate products with confidence and enjoy comfortable shopping.
[1975] The specific processing unit 290 transmits the result of the specific processing to the robot 414. In the robot 414, the control unit 46A causes the speaker 240 and the control target 443 to output the result of the specific processing. The microphone 238 acquires voice indicating a user input regarding the result of the specific processing. The control unit 46A transmits voice data indicating the user input acquired by the microphone 238 to the data processing device 12. In the data processing device 12, the specific processing unit 290 acquires the voice data.
[1976] The data generation model 58 is a so-called generative AI (Artificial Intelligence). An example of the data generation model 58 is ChatGPT (Internet Search<URL: https: / / openai.com / blog / chatgpt> ), Gemini (Internet search <url: https: gemini.google.com ?hl="ja">) and other generation AIs. The data generation model 58 is obtained by performing deep learning on a neural network. A prompt including an instruction is input to the data generation model 58, and inference data such as voice data indicating voice, text data indicating text, and image data indicating an image is also input. The data generation model 58 performs inference on the input inference data in accordance with the instruction indicated by the prompt, and outputs the inference result in a data format such as voice data and text data. Here, inference refers to, for example, analysis, classification, prediction, and / or summarization.
[1977] In the above embodiment, an example was given in which the specific processing is performed by the data processing device 12, but the technology of the present disclosure is not limited to this, and the specific processing may be performed by the robot 414.
[1978] The emotion identification model 59 as an emotion engine may determine the user's emotion according to a specific mapping. Specifically, the emotion identification model 59 may determine the user's emotion according to an emotion map (see FIG. 9), which is a specific mapping. Similarly, the emotion identification model 59 may determine the robot's emotion, and the identification processing unit 290 may perform identification processing using the robot's emotion.
[1979] FIG. 9 is a diagram illustrating an emotion map 400 on which multiple emotions are mapped. In the emotion map 400, emotions are arranged in concentric circles radiating from the center. Emotions closer to the center of the concentric circles are more primitive. Emotions representing states and actions arising from a state of mind are arranged on the outer edges of the concentric circles. The concept of emotion includes both affect and mental states. Emotions generally generated from reactions occurring in the brain are arranged on the left side of the concentric circles. Emotions generally induced by situational judgment are arranged on the right side of the concentric circles. Emotions generally generated from reactions occurring in the brain and induced by situational judgment are arranged on the upper and lower sides of the concentric circles. Furthermore, the emotion of "pleasure" is arranged on the upper side of the concentric circles, and the emotion of "discomfort" is arranged on the lower side. In this way, in the emotion map 400, multiple emotions are mapped based on the structure by which emotions are generated, and emotions that tend to occur simultaneously are mapped close to each other.
[1980] These emotions are distributed in the 3 o'clock direction on emotion map 400, and typically fluctuate between relief and anxiety. In the right half of emotion map 400, situational awareness dominates over internal sensations, resulting in a sense of calm.
[1981] The inside of emotion map 400 represents what is going on in the mind, and the outside of emotion map 400 represents behavior, so the further you go outside emotion map 400, the more visible the emotions become (the more they are expressed in behavior).
[1982] Human emotions are based on various balances, such as posture and blood sugar levels. When these balances deviate from the ideal, a state of discomfort is indicated, and when they approach the ideal, a state of pleasure is indicated. Emotions can also be created for robots, automobiles, and motorcycles, based on various balances, such as posture and remaining battery life. When these balances deviate from the ideal, a state of discomfort is indicated, and when they approach the ideal, a state of pleasure is indicated. An emotion map can be generated, for example, based on Dr. Mitsuyoshi's emotion map (Research on Voice Emotion Recognition and Emotional Brain Physiological Signal Analysis Systems, Tokushima University, Doctoral Dissertation: https: / / ci.nii.ac.jp / naid / 500000375379). The left half of the emotion map lists emotions belonging to the "reaction" domain, where sensation is dominant. The right half of the emotion map lists emotions belonging to the "situation" domain, where situational awareness is dominant.
[1983] The emotion map defines two emotions that promote learning. One is a negative emotion on the situation side, around the middle of "repentance" or "reflection." In other words, this occurs when the robot experiences negative emotions such as "I never want to feel this way again" or "I don't want to be scolded again." The other is a positive emotion on the response side, around "desire." In other words, this occurs when the robot experiences positive feelings such as "I want more" or "I want to know more."
[1984] The emotion identification model 59 inputs user input into a pre-trained neural network, obtains emotion values indicating each emotion shown in the emotion map 400, and determines the user's emotion. This neural network is pre-trained based on multiple pieces of training data that are combinations of user input and emotion values indicating each emotion shown in the emotion map 400. Furthermore, this neural network is trained so that emotions that are located close to each other have similar values, as in the emotion map 900 shown in FIG. 10. FIG. 10 shows an example in which multiple emotions, "relieved," "calm," and "reassuring," have similar emotion values.
[1985] The system according to the present disclosure has been described above mainly with respect to the functions of the data processing device 12, but the system according to the present disclosure is not necessarily implemented on a server. The system according to the present disclosure may be implemented as a general information processing system. The present disclosure may be implemented, for example, as a software program running on a personal computer or an application running on a smartphone, etc. The method according to the present disclosure may be provided to users in the form of SaaS (Software as a Service).
[1986] In the above embodiment, an example was given in which the specific processing is performed by one computer 22, but the technology of the present disclosure is not limited to this, and the specific processing may be distributed and performed by a plurality of computers including the computer 22. For example, the data generation model 58 may be provided in an external device of the data processing device 12, and data may be generated in the external device in accordance with input data.
[1987] In the above embodiment, an example in which the specific processing program 56 is stored in the storage 32 has been described, but the technology of the present disclosure is not limited to this. For example, the specific processing program 56 may be stored in a portable, computer-readable, non-transitory storage medium such as a USB (Universal Serial Bus) memory. The specific processing program 56 stored in the non-transitory storage medium is installed in the computer 22 of the data processing device 12. The processor 28 executes the specific processing in accordance with the specific processing program 56.
[1988] Alternatively, the specific processing program 56 may be stored in a storage device such as a server connected to the data processing device 12 via the network 54, and the specific processing program 56 may be downloaded and installed on the computer 22 in response to a request from the data processing device 12.
[1989] It is not necessary to store all of the specific processing program 56 in a storage device such as a server connected to the data processing device 12 via the network 54, or to store all of the specific processing program 56 in the storage 32; only a portion of the specific processing program 56 may be stored.
[1990] The hardware resource for executing a specific process can be any of the following processors: An example of a processor is a CPU, which is a general-purpose processor that functions as a hardware resource for executing a specific process by executing software, i.e., a program. Another example of a processor is a dedicated electrical circuit, such as an FPGA (Field-Programmable Gate Array), a PLD (Programmable Logic Device), or an ASIC (Application Specific Integrated Circuit), which is a processor with a circuit configuration designed specifically for executing a specific process. Each processor has built-in or connected memory, and each processor uses the memory to execute the specific process.
[1991] The hardware resource that executes the specific processing may be configured with one of these various processors, or may be configured with a combination of two or more processors of the same or different types (for example, a combination of multiple FPGAs, or a combination of a CPU and an FPGA). Also, the hardware resource that executes the specific processing may be a single processor.
[1992] As an example of a system configured with a single processor, first, one processor is configured by combining one or more CPUs and software, and this processor functions as a hardware resource that executes a specific process. Second, there is a system that uses a processor that realizes the functions of an entire system including multiple hardware resources that execute a specific process on a single IC chip, as typified by SoC (System-on-a-chip). In this way, a specific process is realized using one or more of the above-mentioned various processors as hardware resources.
[1993] Furthermore, the hardware structure of these various processors can be, more specifically, an electric circuit that combines circuit elements such as semiconductor devices. The specific processing described above is merely an example. Therefore, it goes without saying that unnecessary steps may be deleted, new steps may be added, or the processing order may be rearranged, without departing from the spirit of the invention.
[1994] The above-described description and illustrations are a detailed explanation of the parts related to the technology of the present disclosure and are merely an example of the technology of the present disclosure. For example, the above description of the configuration, functions, actions, and effects is an explanation of an example of the configuration, functions, actions, and effects of the parts related to the technology of the present disclosure. Therefore, it goes without saying that unnecessary parts may be deleted, new elements may be added, or replacements may be made to the above-described description and illustrations within the scope of the gist of the technology of the present disclosure. Furthermore, to avoid confusion and facilitate understanding of the parts related to the technology of the present disclosure, the above-described description and illustrations omit explanations of common technical knowledge that do not require particular explanation to enable the implementation of the technology of the present disclosure.
[1995] All publications, patent applications, and technical standards mentioned in this specification are herein incorporated by reference to the same extent as if each individual publication, patent application, or technical standard was specifically and individually indicated to be incorporated by reference.
[1996] The following is further disclosed regarding the above embodiment.
[1997] (Claim 1)
[1998] a means for capturing product information by a user terminal;
[1999] means for transmitting the captured image data and audio data of the product information to a server;
[2000] means for analyzing the image data received in the server and evaluating the quality, price, and expiration date of the product;
[2001] means for analyzing the voice data received in the server and detecting whether or not the salesperson is trying to force a sale;
[2002] means for sending a notification to the user based on the analysis results;
[2003] A system including:
[2004] (Claim 2)
[2005] 2. The system according to claim 1, further comprising means for notifying the user terminal of a recommendation or non-recommendation of a purchase based on the analysis result.
[2006] (Claim 3)
[2007] 2. The system of claim 1, wherein the means for capturing the product information is a user terminal including a camera and a microphone.
[2008] "Example 1"
[2009] (Claim 1)
[2010] a means for capturing product information by a user terminal;
[2011] means for transmitting the captured image data and audio data of the product information to a server;
[2012] A means for analyzing the image data received by the server using a generating AI model to evaluate the product name, price, expiration date, and appearance of the product;
[2013] A means for analyzing the voice data received in the server using a generating AI model to detect whether or not a salesperson is trying to force a sale;
[2014] means for sending a notification to the user based on the analysis results;
[2015] A system including:
[2016] (Claim 2)
[2017] 2. The system according to claim 1, further comprising means for notifying the user terminal of a recommendation or non-recommendation of a purchase based on the analysis result.
[2018] (Claim 3)
[2019] 2. The system of claim 1, wherein the means for capturing the product information is a user terminal including a camera and a microphone.
[2020] "Application Example 1"
[2021] (Claim 1)
[2022] a means for capturing product information by a user terminal;
[2023] means for transmitting the captured image data and audio data of the product information to a server;
[2024] means for analyzing the image data received in the server and evaluating the quality, price, and expiration date of the product;
[2025] means for analyzing the voice data received in the server and detecting whether or not the salesperson is trying to force a sale;
[2026] means for sending a notification to the user based on the analysis results;
[2027] A means for capturing and notifying product information using an application installed on a smartphone;
[2028] A system including:
[2029] (Claim 2)
[2030] 2. The system according to claim 1, further comprising means for notifying the user terminal of a recommendation or non-recommendation of a purchase based on the analysis result.
[2031] (Claim 3)
[2032] 2. The system of claim 1, wherein the means for capturing the product information is a user terminal including a camera and a microphone.
[2033] "Example 2: Combining Emotion Engines"
[2034] (Claim 1)
[2035] a means for capturing product information by a user terminal;
[2036] means for transmitting the captured image data and audio data of the product information to a server;
[2037] means for analyzing the image data received in the server and evaluating the quality, price, and expiration date of the product;
[2038] means for analyzing the voice data received in the server and detecting whether or not the salesperson is trying to force a sale;
[2039] means for analyzing a user's facial expression and tone of voice using an emotion engine in the server to recognize the user's emotion;
[2040] means for sending a notification to the user based on the analysis results;
[2041] A system including:
[2042] (Claim 2)
[2043] 2. The system according to claim 1, further comprising means for notifying the user terminal of a recommendation or non-recommendation of a purchase based on the analysis result.
[2044] (Claim 3)
[2045] 2. The system of claim 1, wherein the means for capturing the product information is a user terminal including a camera and a microphone.
[2046] "Application example 2 when combining emotion engines"
[2047] (Claim 1)
[2048] a means for capturing product information by a user terminal;
[2049] means for transmitting the captured image data and audio data of the product information to a server;
[2050] means for analyzing the image data received in the server and evaluating the quality, price, and expiration date of the product;
[2051] means for analyzing the voice data received in the server and detecting whether or not the salesperson is trying to force a sale;
[2052] means for sending a notification to the user based on the analysis results;
[2053] means for analyzing the user's emotions using an emotion engine in the server and adjusting the notification content based on the user's emotional state;
[2054] A system including:
[2055] (Claim 2)
[2056] 2. The system according to claim 1, further comprising means for notifying the user terminal of a recommendation or non-recommendation of a purchase based on the analysis result.
[2057] (Claim 3)
[2058] 2. The system of claim 1, wherein the means for capturing the product information is a user terminal including a camera and a microphone. [Explanation of symbols]
[2059] 10, 210, 310, 410 Data Processing Systems 12 Data Processing Device 14 Smart Devices 214 Smart Glasses 314 Headset-type terminal 414 Robot< / url:> < / url:> < / url:> < / url:>
Claims
1. a means for capturing product information by a user terminal; means for transmitting the captured image data and audio data of the product information to a server; means for analyzing the image data received in the server and evaluating the quality, price, and expiration date of the product; means for analyzing the voice data received in the server and detecting whether or not the salesperson is trying to force a sale; means for sending a notification to the user based on the analysis results; A system including:
2. The system according to claim 1 , further comprising means for notifying the user terminal of a recommendation or non-recommendation of a purchase based on the analysis result.
3. 2. The system of claim 1, wherein the means for capturing product information is a user terminal including a camera and a microphone.
Citation Information
Patent Citations
Persona chatbot control method and system
JP2022180282A