System
The real-time information providing system uses portable terminals to recognize user emotions and select personalized deals based on store information and user history, addressing the inefficiencies of traditional deal-finding methods and enhancing user convenience and business promotion.
Patent Information
- Application Number
- JP2024182292
- Authority / Receiving Office
- JP · JP
- Patent Type
- Applications
- Current Assignee / Owner
- Priority Date
- 2023-10-23
- Filing Date
- 2024-10-17
- Publication Date
- 2025-05-08
AI Technical Summary
Traditional methods for finding great deals at stores and restaurants require time-consuming individual searches, and existing systems fail to provide personalized and real-time information tailored to users' emotions and preferences.
A real-time information providing system using a portable information terminal that recognizes user emotions, acquires store information through image data and location information, and selects advantageous information based on user attributes and past behavior history, using a server and emotional engine to optimize the information displayed to the user.
Enables users to quickly and easily obtain valuable information at stores and restaurants, improving user convenience and allowing businesses to effectively promote campaigns and discounts, leading to increased customer acquisition and sales.
Smart Images

Figure 2025071789000001_ABST
Abstract
Description
[Technical field]
[0001] The technology of the present disclosure relates to a system. [Background technology]
[0002] Patent document 1 discloses a persona chatbot control method performed by at least one processor, the method including the steps of receiving a user utterance, adding the user utterance to a prompt including a description and related instruction sentence regarding the chatbot character, encoding the prompt, and inputting the encoded prompt into a language model to generate a chatbot utterance in response to the user utterance. [Prior art documents] [Patent documents]
[0003] [Patent Document 1] JP 2022-180282 A Summary of the Invention [Problem to be solved by the invention]
[0004] In the past, when visiting stores or restaurants, you had to search and check individually to find good deals, which was time-consuming and laborious. [Means for solving the problem]
[0005] The present invention provides a real-time information providing system that utilizes a mobile information terminal, thereby enabling a user to easily and quickly obtain advantageous information when visiting a store or restaurant.
[0006] Specifically, the system provides advantageous information to a user when the user carries a mobile information terminal and visits a store, and includes: means for recognizing the user's emotions using an emotion engine that recognizes the user's emotions; means for acquiring information about the store based on image data or location information obtained from the mobile information terminal; means for generating an output result indicating the advantageous information based on a generative AI and a prompt sentence that instructs the user to select the advantageous information from the acquired store information together with the user's attribute data or past behavioral history; and means for outputting the generated output result indicating the advantageous information to the mobile information terminal.
[0007] This allows users to obtain advantageous information on stores and restaurants without the time and effort required for conventional methods, improving convenience. Stores and restaurants can also provide users with effective promotion and campaign information, which leads to increased customer acquisition and sales.
[0008] The store information is obtained by analyzing image data obtained from a mobile information terminal such as smart glasses and recognizing the store's logo or signboard.
[0009] The advantageous information is selected in consideration of the acquired store information, the attributes of the user, and the past behavior history.
[0010] The advantageous information is useful information that the user can obtain at the store, and includes at least one of campaign information, special offer information, and discount information. [Brief description of the drawings]
[0011] [Figure 1] 1 is a conceptual diagram showing an example of the configuration of a data processing system according to a first embodiment. [Diagram 2] 1 is a conceptual diagram showing an example of main functions of a data processing device and a smart device according to a first embodiment. FIG. [Diagram 3]FIG. 11 is a conceptual diagram showing an example of the configuration of a data processing system according to a second embodiment. [Figure 4] FIG. 11 is a conceptual diagram showing an example of main functions of a data processing device and smart glasses according to a second embodiment. [Diagram 5] FIG. 13 is a conceptual diagram showing an example of the configuration of a data processing system according to a third embodiment. [Figure 6] FIG. 11 is a conceptual diagram showing an example of main functions of a data processing device and a headset-type terminal according to a third embodiment. [Figure 7] FIG. 13 is a conceptual diagram showing an example of the configuration of a data processing system according to a fourth embodiment. [Figure 8] FIG. 13 is a conceptual diagram showing an example of main functions of a data processing device and a robot according to a fourth embodiment. [Figure 9] 1 shows an emotion map onto which multiple emotions are mapped. [Figure 10] 1 shows an emotion map onto which multiple emotions are mapped. [Figure 11] 4 is a sequence diagram showing a process flow of the data processing system according to the first embodiment. FIG. [Figure 12] 11 is a sequence diagram showing a process flow of the data processing system in application example 1. FIG. [Figure 13] FIG. 11 is a sequence diagram showing the flow of processing of the data processing system according to the second embodiment when an emotion engine is combined. [Figure 14] FIG. 11 is a sequence diagram showing the flow of processing in the data processing system in application example 2 when combined with an emotion engine. DETAILED DESCRIPTION OF THE PREFERRED EMBODIMENTS
[0012] Hereinafter, an example of an embodiment of a system according to the technology of the present disclosure will be described with reference to the accompanying drawings.
[0013] First, the terms used in the following description will be explained.
[0014] In the following embodiments, a signed processor (hereinafter simply referred to as a "processor") may be one arithmetic device or a combination of multiple arithmetic devices. Furthermore, a processor may be one type of arithmetic device or a combination of multiple types of arithmetic devices. Examples of arithmetic devices include a CPU (Central Processing Unit), a GPU (Graphics Processing Unit), a GPGPU (General-Purpose computing on Graphics Processing Units), an APU (Accelerated Processing Unit), etc.
[0015] In the following embodiments, a signed RAM (Random Access Memory) is a memory in which information is temporarily stored and is used as a working memory by the processor.
[0016] In the following embodiments, the coded storage is one or more non-volatile storage devices that store various programs, various parameters, etc. Examples of non-volatile storage devices include flash memory (SSD (Solid State Drive)), magnetic disks (e.g., hard disks), and magnetic tapes.
[0017] In the following embodiments, a communication I / F (Interface) with a code is an interface including a communication processor and an antenna. The communication I / F controls communication between multiple computers. An example of a communication standard applied to the communication I / F is a wireless communication standard including 5G (5th Generation Mobile Communication System), Wi-Fi (registered trademark), Bluetooth (registered trademark), etc.
[0018] In the following embodiments, "A and / or B" is synonymous with "at least one of A and B." In other words, "A and / or B" means that it may be only A, only B, or a combination of A and B. In addition, in this specification, the same idea as "A and / or B" is also applied when three or more things are expressed by connecting them with "and / or."
[0019] [First embodiment]
[0020] FIG. 1 shows an example of the configuration of a data processing system 10 according to the first embodiment.
[0021] 1, a data processing system 10 includes a data processing device 12 and a smart device 14. An example of the data processing device 12 is a server.
[0022] The data processing device 12 includes a computer 22, a database 24, and a communication I / F 26. The computer 22 is an example of a "computer" according to the technology of the present disclosure.
[0023] The computer 22 includes a processor 28, a RAM 30, and a storage 32. The processor 28, the RAM 30, and the storage 32 are connected to a bus 34. The database 24 and the communication I / F 26 are also connected to the bus 34. The communication I / F 26 is connected to a network 54. Examples of the network 54 include a wide area network (WAN) and / or a local area network (LAN).
[0024] The smart device 14 includes a computer 36, a reception device 38, an output device 40, a camera 42, and a communication I / F 44. The computer 36 includes a processor 46, a RAM 48, and a storage 50. The processor 46, the RAM 48, and the storage 50 are connected to a bus 52. The reception device 38, the output device 40, and the camera 42 are also connected to the bus 52.
[0025] The reception device 38 includes a touch panel 38A and a microphone 38B, and receives user input. The touch panel 38A detects contact with a pointer (e.g., a pen or a finger) to receive user input by the touch of the pointer. The microphone 38B detects the user's voice to receive user input by voice. The control unit 46A transmits data indicating the user input received by the touch panel 38A and the microphone 38B to the data processing device 12. In the data processing device 12, the specific processing unit 290 acquires the data indicating the user input.
[0026] The output device 40 includes a display 40A and a speaker 40B, and presents data to the user 20 by outputting the data in a form of expression that the user 20 can perceive (e.g., voice and / or text). The display 40A displays visible information such as text and images according to instructions from the processor 46. The speaker 40B outputs voice according to instructions from the processor 46. The camera 42 is a compact digital camera equipped with an optical system including a lens, an aperture, and a shutter, and an imaging element such as a Complementary Metal-Oxide-Semiconductor (CMOS) image sensor or a Charge Coupled Device (CCD) image sensor.
[0027] The communication I / F 44 is connected to a network 54. The communication I / Fs 44 and 26 are responsible for transmitting and receiving various types of information between the processor 46 and the processor 28 via the network 54.
[0028] FIG. 2 shows an example of main functions of the data processing device 12 and the smart device 14.
[0029] As shown in Fig. 2, in the data processing device 12, a specific process is performed by the processor 28. A specific process program 56 is stored in the storage 32. The specific process program 56 is an example of a "program" according to the technology of the present disclosure. The processor 28 reads the specific process program 56 from the storage 32, and executes the read specific process program 56 on the RAM 30. The specific process is realized by the processor 28 operating as a specific processing unit 290 in accordance with the specific process program 56 executed on the RAM 30.
[0030] The storage 32 stores a data generation model 58 and an emotion identification model 59. The data generation model 58 and the emotion identification model 59 are used by the identification processing unit 290.
[0031] In the smart device 14, the processor 46 performs the reception output processing. The storage 50 stores a reception output program 60. The reception output program 60 is used together with the specific processing program 56 by the data processing system 10. The processor 46 reads out the reception output program 60 from the storage 50, and executes the read reception output program 60 on the RAM 48. The reception output processing is realized by the processor 46 operating as the control unit 46A in accordance with the reception output program 60 executed on the RAM 48.
[0032] Next, a description will be given of the identification process performed by the identification processing unit 290 of the data processing device 12. In the following description, the data processing device 12 will be referred to as a "server" and the smart device 14 will be referred to as a "terminal."
[0033] The embodiment for carrying out the present invention comprises the following elements.
[0034] 1. Terminal: A device used when a user visits a store or restaurant to acquire images and location information and display special offers.
[0035] 2. Server: A central processing unit that receives images and location information from the terminals, obtains information about stores and restaurants, and processes the information to select deals.
[0036] 3. Database: This is where the server stores data to search for store and restaurant information. This includes store logos, signboard information, campaign information, discount information, etc.
[0037] 4. Program: Software that allows the server to process information obtained from the terminal, select advantageous information, and display it on the terminal. This includes image recognition technology, location information analysis, and consideration of user attributes and past behavioral history.
[0038] (Specific examples)
[0039] A specific example in which a user uses the smart device 14 to visit a cafe will be described.
[0040] 1. A user carries a smart device 14 and enters a cafe.
[0041] 2. The smart device 14 recognizes the cafe's logo or sign and transmits that information to the server.
[0042] 3.The server searches the database for information about the cafe based on the received information and obtains campaign and discount information.
[0043] 4. The server selects deals that are appropriate for the user, taking into account the user's attributes and past behavioral history.
[0044] 5. The server transmits the selected deals to the smart device 14.
[0045] 6. The smart device 14 displays the received deals information, allowing the user to take advantage of campaigns and discounts.
[0046] The above is a specific example of an embodiment of the present invention. By linking the smart device 14 with the server, the user can easily obtain and use information on deals at stores and restaurants.
[0047] The process flow will be explained below.
[0048] Step 1: The server receives images and location information obtained from the smart device 14. Specifically, the smart device 14 recognizes the logo or sign of the cafe and transmits the information to the server.
[0049] Step 2: The server analyzes the received image and location information. Using image recognition technology, it identifies the cafe's logo and sign and obtains the location information.
[0050] Step 3: The server searches the database for information about the cafe. The database contains information about the cafe's campaigns and discounts, and the server retrieves this information.
[0051] Step 4: The server selects the appropriate discount information based on the user's attributes and past behavior history. For example, if the user has visited the same cafe in the past and received a discount at that time, the server selects the discount information based on that information.
[0052] Step 5: The server transmits the selected bargain information to the smart device 14. Specifically, the server generates data for displaying the bargain information and transmits it to the smart device 14.
[0053] Step 6: The smart device 14 displays the received special offer information. The user can check the campaign and discount information at the cafe through the screen of the smart device 14.
[0054] Example 1
[0055] Next, a description will be given of Example 1. In the following description, the data processing device 12 is referred to as a "server" and the smart device 14 is referred to as a "terminal."
[0056] Conventionally, there has been no system that allows users to easily and quickly obtain and use information on deals on the spot when they visit a store. In addition, there is a demand for a method to realize more effective marketing by providing personalized information that takes into account the user's attributes and past behavioral history.
[0057] The specific processing by the specific processing unit 290 of the data processing device 12 in the first embodiment is realized by the following means.
[0058] In this invention, the server includes a means for acquiring store information based on image data and location information acquired from the mobile information terminal when the user visits the store using the mobile information terminal, a means for selecting advantageous information from the acquired store information, and a means for displaying the selected advantageous information on the mobile information terminal, thereby enabling the user to easily acquire and use personalized advantageous information in real time.
[0059] A "portable information terminal" is an electronic device that a user can carry with them, and is a terminal that has the function of taking pictures and acquiring location information.
[0060] A "server" is a central processing unit that receives data from multiple portable information terminals via a network, analyzes the data, and provides related information.
[0061] "Store information" is data related to a specific store, and includes image data of logos and signs, location information, campaign information, discount information, and the like.
[0062] "Bargain information" is promotional information that is beneficial to the user, such as special offers, discounts, and sales information for stores visited by the user.
[0063] "Image data" refers to digital data of an image captured by a user's mobile information terminal, and is the subject of analysis.
[0064] "Location information" refers to geographic coordinate data obtained through the GPS function of a mobile information terminal, for example.
[0065] "User attributes" refers to information relating to a user, such as age, sex, preferences, etc., that is used to provide personalized information.
[0066] "Past behavioral history" refers to a record of stores the user previously visited, products the user purchased, campaigns the user participated in, and the like, and indicates the user's behavioral patterns.
[0067] "Analysis" is the process in which the server extracts and identifies the necessary information based on the image data and location information it receives.
[0068] The present invention relates to a system for providing advantageous information to a user when the user visits a store using a mobile information terminal. Specific embodiments of the system will be described below.
[0069] Hardware and software configuration
[0070] First, the hardware and software required to realize this system will be described.
[0071] 1. Mobile information terminals:
[0072] A mobile information terminal is an electronic device that can be carried by a user and has the functions of taking pictures and acquiring location information. Specifically, a smartphone or a tablet is one such device.
[0073] It has a camera function and can acquire image data.
[0074] It has a GPS function and can obtain location information.
[0075] 2. Server:
[0076] A server is a central processing unit that receives data from multiple portable information terminals via a network, analyzes the data, and provides related information.
[0077] The servers are typically located in the cloud and have the necessary processing power and storage.
[0078] Image recognition libraries such as "OpenCV" are used for image analysis.
[0079] "Google (registered trademark) Maps API" is used for location analysis.
[0080] Machine learning libraries such as "TENSORFLOW (registered trademark)" will be used to analyze user attributes and past behavioral history.
[0081] 3. Database:
[0082] The database stores store information (logos, signs, campaign information, discount information, etc.).
[0083] Program Processing
[0084] The server executes a program for selecting and providing advantageous information suitable for the user based on image data and location information obtained from the mobile information terminal.
[0085] 1. Data Acquisition:
[0086] Image data of a store's logo or signboard photographed by the terminal's camera is acquired.
[0087] The current location information is obtained using the device's GPS function.
[0088] 2. Data transmission:
[0089] The device transmits the acquired image data and location information to the server using Wi-Fi or 4G / 5G networks.
[0090] 3. Image Analysis:
[0091] The server uses OpenCV to analyze the image data and identify store logos and signs, extract feature points within the image, and match them with existing images in a database.
[0092] 4. Information Search:
[0093] Based on the analysis results, the server retrieves detailed information about the relevant store from the database, such as the cafe's menu, opening hours, and location.
[0094] 5. Obtaining Campaign Information:
[0095] The server obtains store campaign and discount information from the database. For example, it obtains information about "new menu item launch commemorative discount."
[0096] 6. Choose the best deals:
[0097] The server selects the most suitable campaign information for a user based on the user's attribute information (e.g., age, gender, preferences) and past behavioral history. To do this, a machine learning algorithm using "TensorFlow" is used.
[0098] 7. Information display:
[0099] The server transmits the selected deals to the terminal, which displays the received deals on its screen, allowing the user to use the information to receive campaigns and discounts.
[0100] Examples
[0101] For example, a specific example in which a user uses a mobile information terminal to visit a cafe is as follows.
[0102] 1. User enters the cafe:
[0103] A user carries a mobile information terminal and enters a cafe.
[0104] 2. The device captures a logo or sign:
[0105] Using the device's camera, the user takes a picture of the cafe's logo or sign.
[0106] 3. The device sends the image and location information to the server:
[0107] The device sends the captured images and location information obtained using the GPS function to the server.
[0108] 4. The server analyzes the image:
[0109] The server analyzes the received image using OpenCV and retrieves information about the corresponding cafe from a database.
[0110] 5. The server gets the campaign information:
[0111] The server retrieves the cafe's latest promotions and discounts from the same database.
[0112] 6. Server selects deals:
[0113] The server selects the most suitable deals by taking into account the user's attributes and past behavioral history.
[0114] 7. The server sends the deals to the device:
[0115] The server sends the selected deals to the terminal.
[0116] 8. Your device will display special offers:
[0117] The device will display the received deals on the screen, for example, "This cafe is now offering 50% off our new coffee menu!"
[0118] 9. User takes advantage of promotions and discounts:
[0119] Users can use the deals displayed on the screen to take advantage of coupons and discounts at the cafe.
[0120] Examples of prompt statements
[0121] "When a user visits a cafe, he or she takes a photo of the cafe's logo with a mobile information terminal and sends it to the server. The server then retrieves information about the cafe from the database, selects the most suitable deals based on the user's attributes and past behavioral history, and sends them to the mobile information terminal. The mobile information terminal displays the received information, and the user takes advantage of the cafe's campaign. Please explain the specific process."
[0122] The above is an embodiment of the present invention. This system allows users to easily obtain and use information about deals at stores.
[0123] The flow of the identification process in the first embodiment will be described with reference to FIG.
[0124] Step 1:
[0125] User enters the cafe
[0126] A user carries a mobile information terminal and enters a cafe.
[0127] Input: Cafe location, user attribute information
[0128] Output: The user is in the store
[0129] Step 2:
[0130] The device photographs logos and signs
[0131] Using the device's camera, the user takes a picture of the cafe's logo or sign.
[0132] Specific actions: The user activates the device's camera, frames the cafe's logo or sign, and presses the capture button.
[0133] Input: Image data of the cafe logo or sign
[0134] Output: Captured image data
[0135] Step 3:
[0136] The device sends the image and location information to the server.
[0137] The terminal transmits the acquired image data and the location information acquired by the GPS function to the server.
[0138] Specific operation: The device uses Wi-Fi or 4G / 5G networks to pack image data and location information into packets and upload them to the server.
[0139] Input: Image data, location information
[0140] Output: Image data and location information sent to the server
[0141] Step 4:
[0142] The server analyzes the image
[0143] The server analyzes the received image data using OpenCV and identifies the cafe's logo and sign.
[0144] Specific operation: Extract feature points in the image and match them with images in an existing database to identify the corresponding store.
[0145] Input: Image data
[0146] Output: Identified store information
[0147] Step 5:
[0148] The server retrieves cafe information from the database.
[0149] The server retrieves detailed information about the corresponding cafe from the database based on the identified store information.
[0150] Specific operation: The server sends a query to the database to obtain information about the relevant cafe, such as its menu, opening hours, and location.
[0151] Input: Identified store information
[0152] Output: Detailed information about the cafe
[0153] Step 6:
[0154] The server retrieves the campaign information.
[0155] The server retrieves the cafe's latest promotions and discounts from the same database.
[0156] Specific operation: The server sends a query to the database to obtain the relevant campaign information.
[0157] Input: Cafe details
[0158] Output: Retrieved campaign information
[0159] Step 7:
[0160] Server selects deals
[0161] The server selects the most suitable deals based on the user's attribute information (age, gender, preferences) and past behavioral history.
[0162] Specific operation: The server inputs the user's attribute data and past behavioral history into a machine learning model and calculates optimized campaign information.
[0163] Input: User attribute information, past behavior history, campaign information
[0164] Output: Selected deals
[0165] Step 8:
[0166] The server sends useful information to the terminal
[0167] The server sends the selected deals to the terminal.
[0168] How it works: The server uses Wi-Fi or 4G / 5G networks to pack deals into packets and upload them to the device.
[0169] Input: Selected Deals
[0170] Output: Deals sent to the terminal
[0171] Step 9:
[0172] The device displays useful information
[0173] The device will display the received deals on the screen.
[0174] Specific operation: A pop-up notification or a dedicated application screen is displayed on the device screen to provide the user with useful information.
[0175] Input: Deals sent to your device
[0176] Output: Deals displayed to the user
[0177] Step 10:
[0178] Users take advantage of promotions and discounts
[0179] Users can use the deals displayed on their devices to take advantage of coupons and discounts at cafes.
[0180] Specific operation: The user shows the coupon screen displayed on the terminal to the cafe cashier to receive the discount.
[0181] Input: Deals displayed to the user
[0182] Output: The user can use the campaign or discount.
[0183] (Application example 1)
[0184] Next, a description will be given of Application Example 1. In the following description, the data processing device 12 will be referred to as a "server" and the smart device 14 will be referred to as a "terminal."
[0185] In the conventional user experience in brick-and-mortar stores, users had to search for store and campaign information themselves, which was not very efficient. In particular, there were limited ways for users to find out what services and discounts were available in stores in real time, so building an information provision system that was effective for users was a challenge.
[0186] The specific processing by the specific processing unit 290 of the data processing device 12 in the application example 1 is realized by the following means.
[0187] In this invention, the server includes a means for acquiring image data and location information, a means for transmitting the acquired image data and location information to the server, a means for the server to search for store information from a database based on the image data and location information and provide the user with selected advantageous information, a means for displaying the advantageous information on the smart glasses, a means for recognizing store logos and signs from the image data, and a means for selecting optimal information in consideration of the user's past behavior history and attributes. This allows the user to acquire and use store information and campaign information in real time.
[0188] "Image data" refers to data that represents visual information acquired by a user using a smart device.
[0189] "Location information" is geographical data such as the latitude and longitude of the user's current location.
[0190] The "server" is a central processing unit that receives image data and location information and searches for store information from a database.
[0191] A "database" is a storage device for storing store information, campaign information, and discount information.
[0192] "Store information" is data including store logos and signs, campaign information, discount information, and the like.
[0193] "Good deals" refers to information about campaigns and discounts that are beneficial to users.
[0194] "Smart glasses" are wearable devices that users wear to display store information and special offers.
[0195] "Image recognition" is a technology that analyzes acquired image data and identifies specific objects or text.
[0196] "Past behavior history" is a record of actions and choices made by a user in the past.
[0197] "User attributes" refers to personal information about the user, such as age, sex, and interests.
[0198] A "generative AI model" is a model generated by artificial intelligence, and is a technology used for data analysis, prediction, and recommendations.
[0199] An embodiment of the present invention will be described below. A system for carrying out the present invention is mainly composed of a terminal for acquiring image data and location information, a server for processing data and providing information, and software for linking these components.
[0200] First, a user visits a physical store using a smart device (such as a smartphone or smart glasses). Image data and location information are acquired using the device's camera and GPS functions. Specific hardware used for this purpose include an on-board webcam, an external camera (e.g., Logitech (registered trademark) C920), a built-in GPS, or an external GPS module (e.g., Holux M-1000C).
[0201] The acquired image data is sent from the terminal to a server. The server receives the image data and location information, and searches a database for store information. Google Cloud Vision API and Amazon Rekognition (registered trademark) are used for image recognition, and Google Maps API is used for location information analysis. The database uses cloud storage such as Amazon RDS (registered trademark).
[0202] The server selects deals suitable for the user based on the acquired store information and the user's past behavioral history and attributes. At this time, a generative AI model (e.g. TensorFlow) is used to analyze user attributes and past behavioral data and recommend the most suitable information.
[0203] The selected deals are then pushed to the user's smart device and displayed to the user, allowing the user to receive real-time deals and improve their shopping experience in-store.
[0204] For example, when a user enters a bookstore, the smart device recognizes the bookstore's logo and location information. The server provides discount information on new books based on the bookstore's information and past purchase history, and displays information about autograph sessions for members only.
[0205] Examples of specific prompts include the following:
[0206] When a user enters a bookstore using their smartphone, we want to display real-time information about discounts on new books and events for members only. Please tell us how to provide the appropriate information using image recognition and GPS information.
[0207] In this way, convenience for users in the store can be greatly improved.
[0208] The flow of the specific process in the application example 1 will be described with reference to FIG.
[0209] Step 1:
[0210] A user visits a store using a smart device. The device's camera takes a picture of the store's logo or sign, and the location information is obtained using the GPS function. The input is camera image data and location information, and the output is data to be sent to the server. The device prepares to send the image data captured by the camera and GPS coordinates to the server.
[0211] Step 2:
[0212] The terminal transmits the acquired image data and location information to the server. The input is the captured image and the acquired location information, and the output is data sent to the server. The terminal encodes the image data and transmits it to the server together with GPS information. At this time, the data is uploaded using an HTTP request.
[0213] Step 3:
[0214] The server processes the received image data and location information. The input is the image data and location information sent from the terminal, and the output is store information retrieved from a database. The server uses image recognition software (e.g. Google Cloud Vision API or Amazon Rekognition) to recognize store logos and signs from the image data. At the same time, it searches the database based on the location information and retrieves the corresponding store information.
[0215] Step 4:
[0216] The server selects deals based on the acquired store information and the user's past behavioral history and attributes. The input is store information, user behavioral history, and user attributes, and the output is the selected deals. The server analyzes the data using a generative AI model (e.g. TensorFlow) and selects the most suitable campaign and discount information.
[0217] Step 5:
[0218] The server sends the selected deals to the terminal. The input is the selected deals, and the output is data transmission to the terminal. The server uses an HTTP request to push the selected information to the user's smart device.
[0219] Step 6:
[0220] The terminal displays the received discount information to the user. The input is the discount information sent from the server, and the output is the information displayed to the user. The terminal uses the push notification function to display the discount information on the display of smart glasses or a smartphone. The user can check campaign information and discounts in real time.
[0221] Furthermore, an emotion engine that estimates the emotion of the user may be combined. That is, the identification processing unit 290 may estimate the emotion of the user using the emotion identification model 59, and perform identification processing using the emotion of the user.
[0222] The embodiment for carrying out the present invention comprises the following elements.
[0223] 1. Terminal: A device used when a user visits a store or restaurant to acquire images and location information and display special offers.
[0224] 2. Server: A central processing unit that receives images and location information from the terminals, obtains information about stores and restaurants, and processes the information to select deals.
[0225] 3. Database: This is where the server stores data to search for store and restaurant information. This includes store logos, signboard information, campaign information, discount information, etc.
[0226] 4. Emotion engine: This engine has the function of recognizing the user's facial expressions and voice characteristics and estimating the user's emotions. The emotion engine is used in cooperation with the server as a means of selecting advantageous information based on the user's emotions.
[0227] 5. Program: Software that allows the server to process information obtained from the terminal, select advantageous information, and display it on the terminal. This includes image recognition technology, location information analysis, emotion engine utilization, and other processes.
[0228] (Specific examples)
[0229] A specific example in which a user uses the smart device 14 to visit a cafe will be described.
[0230] 1. A user carries a smart device 14 and enters a cafe.
[0231] 2. The smart device 14 recognizes the cafe's logo or sign and transmits that information to the server.
[0232] 3.The server searches the database for information about the cafe based on the received information and obtains campaign and discount information.
[0233] 4. The smart glasses simultaneously transmit the user's facial expressions and voice characteristics to the emotion engine.
[0234] 5. The emotion engine estimates the user's emotion and returns the result to the server.
[0235] 6. The server selects deals based on the user's emotions. For example, if the user has a happy expression, the server selects discount deals.
[0236] 7. The server transmits the selected deals to the smart device 14.
[0237] 8. The smart device 14 displays the received discount information based on the user's emotions. For example, the smart device 14 displays a message that matches the user's facial expression together with the discount information.
[0238] The above is a specific example of the embodiment of the present invention. In addition to the cooperation between the smart device 14 and the server, the use of the emotion engine makes it possible to provide advantageous information that matches the user's emotions.
[0239] The process flow will be explained below.
[0240] Step 1: A user visits a cafe using a smart device 14.
[0241] Step 2: The smart device 14 recognizes the logo or sign of the cafe and transmits the information to the server.
[0242] Step 3: The server searches the database for information about the cafe based on the received information.
[0243] Step 4: The smart device 14 simultaneously transmits the user's facial expressions and voice characteristics to the emotion engine.
[0244] Step 5: The emotion engine recognizes the user's facial expressions and voice characteristics and estimates the user's emotions.
[0245] Step 6: The server selects deals based on the user's emotions received from the emotion engine.
[0246] Step 7: The server sends the selected deals to the smart device 14.
[0247] Step 8: The smart device 14 performs display based on the received deals and the user's sentiment.
[0248] Example 2
[0249] Next, a description will be given of Example 2. In the following description, the data processing device 12 is referred to as a "server" and the smart device 14 is referred to as a "terminal."
[0250] Conventional systems provide limited information about the stores that users visit, making it difficult to provide information based on the user's current emotions and individual characteristics. In addition, there is a lack of dynamic information provision to improve the user experience, which means that the functions of the user's smart device cannot be fully utilized.
[0251] The identification process by the identification processing unit 290 of the data processing device 12 in the second embodiment is realized by the following means. In this invention, the server includes a means for a terminal to acquire user's location information, a means for the terminal to acquire image data using a camera and recognize a store's logo or sign based on the image data, a means for the terminal to transmit image data and location information to the server, a means for the server to search a database to acquire store information, a means for analyzing the user's facial expression or voice characteristics using an emotion engine and estimating the user's emotion, a means for the server to select advantageous information based on the user's emotion, a means for the server to transmit the advantageous information to the terminal, and a means for the terminal to display the advantageous information to the user. This makes it possible to acquire information on the store visited by the user from multiple angles and provide advantageous information personalized according to the user's emotion and situation.
[0252] A "terminal" is an electronic device that can be carried by a user and has the functions of acquiring image data, recording location information, and communicating with a server.
[0253] "Location information" is data indicating the geographical location of the user's current location, and is obtained using the GPS function.
[0254] "Image data" is visual information captured using the device's camera, and includes, for example, images of store logos and signs.
[0255] A "server" is a computer system that centrally processes and stores data, and is responsible for receiving and processing data sent from terminals.
[0256] A "database" is a collection of information that stores store logos, signboard information, campaign information, discount information, etc., and is used by the server for search and access.
[0257] An "emotion engine" is software or hardware that analyzes the user's facial expressions and vocal characteristics and estimates the user's emotional state.
[0258] "Bargain information" is information about discounts, campaigns, and the like that is useful to the user, and is selected based on the user's emotions and the stores that the user visits.
[0259] "User facial or vocal features" refers to data such as facial expressions and tone of voice that are used to estimate the user's emotions.
[0260] "Means for transmitting" refers to a communication means for transmitting data from a terminal to a server, and includes, for example, Wi-Fi and mobile data.
[0261] "Display means" refers to a function for visually or audibly presenting the information received by the terminal to the user, and includes screen display and audio guidance.
[0262] The system for implementing the present invention is composed of a user, a terminal, a server, a database, an emotion engine, and a program that connects these. Below, we will specifically explain how these elements work together to realize the invention.
[0263] First, the core of the system is the terminal used by the user. This terminal is designed as a smart device (e.g., smartphone, smart glasses) and has the following functions: It acquires image data using its camera function and collects location information using its GPS function. It also has sensors that detect the user's facial expressions and voice characteristics.
[0264] Next, a server is required to process the data acquired by the terminal. The server performs the following processes.
[0265] 1. Receive image data and location information sent from the device.
[0266] 2. Access the database and search for and retrieve image data of store logos and signs, as well as campaign and discount information.
[0267] 3. Use an emotion engine to estimate the user's emotions from their facial expressions and voice characteristics.
[0268] 4. Select the best deal from the information in the database based on user sentiment.
[0269] 5. Send the selected deals back to your device.
[0270] The database stores detailed data such as store and restaurant logos, signboard information, campaign information, discount information, etc. This database is maintained so that the server can efficiently search and access it.
[0271] The emotion engine estimates emotions by analyzing the user's facial expressions and vocal characteristics. For example, a generative AI model is used to classify emotional states such as "happy," "sad," and "surprised." This model is capable of analyzing the user's real-time emotions with high accuracy.
[0272] As a concrete example, the process when a user visits a cafe is shown below.
[0273] 1. A user enters a cafe with their smart device.
[0274] 2. The device recognizes the cafe's logo or sign, and sends image data and location information to the server.
[0275] 3. The server searches the database for information about the cafe and retrieves campaign and discount information.
[0276] 4. At the same time, the device transmits the user's facial expressions and voice characteristics to the emotion engine.
[0277] 5. The emotion engine estimates the user's emotion and returns the result to the server.
[0278] 6. The server selects deals based on the user's emotions. For example, if the user has a happy expression, the server selects discount deals.
[0279] 7. The server sends the selected deals to the device.
[0280] 8. The device displays information based on the discount information received and the user's emotions. For example, a message that matches the user's facial expression is displayed along with discount information.
[0281] An example of a prompt sentence might be:
[0282] When a user enters a cafe, the smart device scans the cafe's logo or sign and sends the image and location information to the server. The server searches the corresponding cafe information from the database and analyzes the user's facial expressions and voice using an emotion engine. Finally, the server selects deals that match the user's emotions and displays them on the smart device.
[0283] As described above, the present invention provides a system that obtains information about stores visited by a user from multiple angles and provides the user with personalized advantageous information according to the user's emotions and circumstances.
[0284] The flow of the identification process in the second embodiment will be described with reference to FIG.
[0285] Specific flow of system program processing
[0286] Step 1:
[0287] A user enters a store.
[0288] This process starts when a user enters a store with a smart device. For example, assume that the user enters a cafe.
[0289] Step 2:
[0290] The device captures images and location information.
[0291] The device uses its camera to take a picture of the store's logo or sign, and acquires location information using its GPS function. This inputs the image of the cafe's logo and the coordinate data of the current location into the device. The output is the captured image data and location information data.
[0292] Step 3:
[0293] The terminal transmits the information to the server.
[0294] The terminal transmits the acquired image data and location information data to the server. The input is the image data and location information data stored in the terminal, and the data is transmitted by communicating it to the server. The output is the data that has been transferred to the server.
[0295] Step 4:
[0296] The server searches the database.
[0297] The server searches the database based on the received image data and location information. The input is the image data and location information sent to the server, and the output is the corresponding store information (e.g., store name, campaign information, discount information, etc.) retrieved from the database.
[0298] Step 5:
[0299] The server uses an emotion engine to estimate the user's emotion.
[0300] The server inputs the facial expression and voice feature data of the user sent from the terminal into the emotion engine. The emotion engine analyzes this data and estimates the user's emotion. The input is the facial expression and voice feature data of the user, and the output is the user's emotional state (for example, "happy," "sad," "surprised," etc.).
[0301] Step 6:
[0302] The server selects the deals.
[0303] The server selects the optimal deal from the information in the database based on the estimated user emotion. The input is the user's emotional state and store information retrieved from the database, and the output is the deal information (e.g., specific discounts and campaign information) to be provided to the user.
[0304] Step 7:
[0305] The server transmits information to the terminal.
[0306] The server sends the selected deals to the terminal. The input is the deals stored in the server, and the data is sent by communicating it to the terminal. The output is the data that has been transferred to the terminal.
[0307] Step 8:
[0308] The terminal displays the information.
[0309] The terminal displays the received discount information to the user. The input is the discount information sent from the server, and the output is the information that the user can confirm visually or audibly. For example, the message "All items are 10% off today!" is displayed on the terminal screen.
[0310] Through the above processing steps, the present invention makes it possible to obtain information about stores visited by a user from multiple angles and provide personalized deals information according to the user's emotions and circumstances.
[0311] (Application example 2)
[0312] Next, a description will be given of Application Example 2. In the following description, the data processing device 12 will be referred to as a "server" and the smart device 14 will be referred to as a "terminal."
[0313] Conventional store information systems can only provide uniform deals to users, and it is difficult to provide information tailored to the emotions and circumstances of individual users. In addition, since the system does not take into account the emotions of users, the information provided may not always be appropriate, making it difficult to improve user satisfaction.
[0314] The identification process by the identification processing unit 290 of the data processing device 12 in the application example 2 is realized by the following means. In this invention, the server includes a means for recognizing and acquiring store information from image data, a means for acquiring the user's facial expression data together with the image data, a means for selecting advantageous information based on the acquired store information, a means for estimating an emotion based on the user's facial expression data and optimizing the advantageous information based on the emotion, and a means for displaying the optimized advantageous information on the smart glasses. This makes it possible to provide personalized information based on the user's emotions, and is expected to improve user satisfaction.
[0315] The "user-worn device" is a device worn by a user visiting a store, and has the function of acquiring image data and facial expression data.
[0316] The "store information acquisition means" is a function for recognizing store logos and signs from image data obtained from a user-worn device and acquiring that information.
[0317] The "facial expression data acquisition means" is a function for capturing the user's facial expression and acquiring that data.
[0318] The "information selection means" is a function for selecting advantageous information to be presented to the user based on the acquired store information.
[0319] The "emotion estimation means" is a function for estimating the user's emotion based on the acquired facial expression data of the user.
[0320] The "information optimization means" is a function for optimizing selected advantageous information based on the estimated user's emotions.
[0321] The "information display means" is a function for displaying optimized advantageous information on a user-worn device.
[0322] DETAILED DESCRIPTION OF THE PREFERRED EMBODIMENTS
[0323] The system of the present invention consists of smart glasses worn by the user, a server that processes images and location information, and an emotion engine.
[0324] 1. Hardware and Software Configuration
[0325] Smart Glasses
[0326] The smart glasses, worn by the user, have a camera and a display function. The camera captures store logos and signs, and also obtains the user's facial expression data. These data are then sent to a server.
[0327] server
[0328] The server uses the following hardware and software:
[0329] Hardware: High performance processor, storage device
[0330] Software: Python, OpenCV, requests library
[0331] The server analyzes the image data acquired from the smart glasses to recognize store information, and also uses an emotion engine to estimate the user's emotions from their facial expression data.
[0332] Emotion Engine
[0333] The emotion engine is software that analyzes the user's facial expression data and estimates their emotions. The basic emotion estimation algorithm uses a machine learning model.
[0334] 2. Data flow and processing
[0335] The image data acquired by the smart glasses is sent to a server. The server recognizes the store's logo or signboard from the image data and retrieves store information from a database. Next, the server sends the user's facial expression data acquired at the same time to an emotion engine to estimate the user's emotion.
[0336] Store information acquisition method
[0337] The server recognizes logos and signs from the image data and retrieves related store information from the database based on this. Specifically, the server uses the OpenCV library for image recognition.
[0338] Emotion estimation means
[0339] The user's facial expression data is sent to the emotion engine, which uses machine learning to classify the user's emotions into categories such as "happy," "sad," "excited," and "calm."
[0340] Information optimization measures
[0341] The server optimizes the deals based on the user's emotions. For example, if the user has a happy expression, the server will provide discount information.
[0342] Means of selecting and displaying information
[0343] The server sends the optimized information to the smart glasses to display to the user, including messages personalized to the user's emotions.
[0344] 3. Specific Examples
[0345] As a concrete example, consider the case where a user visits a cafe wearing smart glasses.
[0346] 1. When a user enters a cafe, the camera in the smart glasses recognizes the cafe's logo and sign and transmits it to the server.
[0347] 2. The server retrieves information about recognized cafes from the database.
[0348] 3. At the same time, the smart glasses capture the user's facial expression, and if the emotion engine estimates it to be "happy," the server will offer a "Special Happy Hour Discount."
[0349] 4. This information is then displayed on the smart glasses, allowing users to receive real-time deals.
[0350] Example prompts to be input to the generative AI model
[0351] When a smart glasses user enters a store, the store's logo and sign are recognized and the information is sent to the server. Create a program that then analyzes the user's emotions and displays discount and campaign information based on the emotions on the display. Show how to configure it in Python.
[0352] The flow of the specific process in the application example 2 will be described with reference to FIG.
[0353] Step 1:
[0354] Smart glasses capture user actions.
[0355] Input: Image data including the in-store environment and the user's facial expressions captured using the smart glasses camera.
[0356] Output: Image files containing store logos, signage, and user facial expressions.
[0357] How it works: The camera in the smart glasses captures images in real time and stores them in local storage.
[0358] Step 2:
[0359] The terminal transmits the captured image data to the server.
[0360] Input: Image data captured by smart glasses.
[0361] Output: Image data sent to the server.
[0362] Specific operation: The terminal communicates with the server via the Internet using HTTP, and sends image data as a POST request.
[0363] Step 3:
[0364] The server analyzes the image data it receives and recognizes store logos and signs.
[0365] Input: Image data including store logos and signs sent from the device.
[0366] Output: IDs of recognized stores and related information.
[0367] Specific operation: The server processes image data using the OpenCV library and detects and recognizes store logos and signs.
[0368] Step 4:
[0369] The store information recognized by the server is obtained from the database.
[0370] Input: The ID of a recognized store.
[0371] Output: Store details (e.g. store name, campaign information, discount information, etc.).
[0372] Specific operation: The server executes a database query using the store ID as a key to obtain related store information.
[0373] Step 5:
[0374] The terminal captures the user's facial expression data and transmits it to the server.
[0375] Input: Image data containing the user's facial expressions captured by the smart glasses.
[0376] Output: Facial expression data sent to the server.
[0377] Specific operation: The terminal communicates with the server via the Internet via HTTP and sends facial expression data as a POST request.
[0378] Step 6:
[0379] The server analyzes the facial expression data and uses an emotion engine to infer the user's emotions.
[0380] Input: User's facial expression data sent from the device.
[0381] Output: An estimated user emotion classification (e.g. happy, sad, excited, etc.).
[0382] Specific operation: The server uses an emotion engine (machine learning model) to analyze facial expression data and classify emotions.
[0383] Step 7:
[0384] The server selects the best deals based on the user's sentiment.
[0385] Input: Estimated user sentiment and store information.
[0386] Output: Optimized deals based on sentiment.
[0387] Specific operation: Based on the store information, the server selects the deals (discount information and campaign information) that best suit the user's emotions.
[0388] Step 8:
[0389] The server sends the selected information to the terminal and displays it on the smart glasses.
[0390] Enter: Optimized Deals.
[0391] Output: Deals displayed on the smart glasses.
[0392] Specific operation: The server sends the best deals to the device and displays the information on the smart glasses display.
[0393] This allows users to receive real-time, personalized deals when they visit a store.
[0394] The specific processing unit 290 transmits the result of the specific processing to the smart device 14. In the smart device 14, the control unit 46A causes the output device 40 to output the result of the specific processing. The microphone 38B acquires a voice indicating a user input for the result of the specific processing. The control unit 46A transmits the voice data indicating the user input acquired by the microphone 38B to the data processing device 12. In the data processing device 12, the specific processing unit 290 acquires the voice data.
[0395] The data generation model 58 is a so-called generative AI (Artificial Intelligence). An example of the data generation model 58 is ChatGPT (registered trademark) (Internet search<URL: https: / / openai.com / blog / chatgpt> ), Gemini (registered trademark) (Internet search <url: https: gemini.google.com ?hl="ja">) and other generation AIs. The data generation model 58 is obtained by making a neural network perform deep learning. A prompt including an instruction is input to the data generation model 58, and inference data such as voice data indicating a voice, text data indicating a text, and image data indicating an image is input. The data generation model 58 performs inference on the input inference data according to the instruction indicated by the prompt, and outputs the inference result in a data format such as voice data and text data. Here, inference refers to, for example, analysis, classification, prediction, and / or summarization.
[0396] In the above embodiment, an example was given in which the specific process was performed by the data processing device 12, but the technology of the present disclosure is not limited to this, and the specific process may be performed by the smart device 14.
[0397] [Second embodiment]
[0398] FIG. 3 shows an example of the configuration of a data processing system 210 according to the second embodiment.
[0399] 3, the data processing system 210 includes the data processing device 12 and smart glasses 214. An example of the data processing device 12 is a server.
[0400] The data processing device 12 includes a computer 22, a database 24, and a communication I / F 26. The computer 22 is an example of a "computer" according to the technology of the present disclosure. The computer 22 includes a processor 28, a RAM 30, and a storage 32. The processor 28, the RAM 30, and the storage 32 are connected to a bus 34. The database 24 and the communication I / F 26 are also connected to the bus 34. The communication I / F 26 is connected to a network 54. Examples of the network 54 include a wide area network (WAN) and / or a local area network (LAN).
[0401] The smart glasses 214 include a computer 36, a microphone 238, a speaker 240, a camera 42, and a communication I / F 44. The computer 36 includes a processor 46, a RAM 48, and a storage 50. The processor 46, the RAM 48, and the storage 50 are connected to a bus 52. The microphone 238, the speaker 240, and the camera 42 are also connected to the bus 52.
[0402] The microphone 238 receives instructions and the like from the user 20 by receiving voice uttered by the user 20. The microphone 238 captures the voice uttered by the user 20, converts the captured voice into audio data, and outputs it to the processor 46. The speaker 240 outputs the voice according to instructions from the processor 46.
[0403] Camera 42 is a small digital camera equipped with an optical system including a lens, an aperture, and a shutter, and an imaging element such as a CMOS (Complementary Metal-Oxide-Semiconductor) image sensor or a CCD (Charge Coupled Device) image sensor, and captures images of the surroundings of user 20 (e.g., an imaging range defined by an angle of view equivalent to the width of the field of vision of an average healthy person).
[0404] The communication I / F 44 is connected to a network 54. The communication I / Fs 44 and 26 are responsible for transmitting and receiving various types of information between the processor 46 and the processor 28 via the network 54. The transmission and reception of various types of information between the processor 46 and the processor 28 using the communication I / Fs 44 and 26 is performed in a secure state.
[0405] Fig. 4 shows an example of main functions of the data processing device 12 and the smart glasses 214. As shown in Fig. 4, in the data processing device 12, a specific process is performed by the processor 28. A specific process program 56 is stored in the storage 32.
[0406] The specific processing program 56 is an example of a "program" according to the technology of the present disclosure. The processor 28 reads the specific processing program 56 from the storage 32, and executes the read specific processing program 56 on the RAM 30. The specific processing is realized by the processor 28 operating as a specific processing unit 290 in accordance with the specific processing program 56 executed on the RAM 30.
[0407] The storage 32 stores a data generation model 58 and an emotion identification model 59. The data generation model 58 and the emotion identification model 59 are used by the identification processing unit 290.
[0408] In the smart glasses 214, the reception output process is performed by the processor 46. A reception output program 60 is stored in the storage 50. The processor 46 reads the reception output program 60 from the storage 50, and executes the read reception output program 60 on the RAM 48. The reception output process is realized by the processor 46 operating as the control unit 46A in accordance with the reception output program 60 executed on the RAM 48.
[0409] Next, a description will be given of the identification process performed by the identification processing unit 290 of the data processing device 12. In the following description, the data processing device 12 will be referred to as a "server" and the smart glasses 214 will be referred to as a "terminal".
[0410] The embodiment for carrying out the present invention comprises the following elements.
[0411] 1. Terminal: A device used when a user visits a store or restaurant to acquire images and location information and display special offers.
[0412] 2. Server: A central processing unit that receives images and location information from the terminals, obtains information about stores and restaurants, and processes the information to select deals.
[0413] 3. Database: This is where the server stores data to search for store and restaurant information. This includes store logos, signboard information, campaign information, discount information, etc.
[0414] 4. Program: Software that allows the server to process information obtained from the terminal, select advantageous information, and display it on the terminal. This includes image recognition technology, location information analysis, and consideration of user attributes and past behavioral history.
[0415] (Specific examples)
[0416] A specific example is given where a user uses smart glasses 214 to visit a cafe.
[0417] 1. A user puts on the smart glasses 214 and enters a cafe.
[0418] 2. The smart glasses 214 recognize the cafe's logo or sign and send that information to the server.
[0419] 3.The server searches the database for information about the cafe based on the received information and obtains campaign and discount information.
[0420] 4. The server selects deals that are appropriate for the user, taking into account the user's attributes and past behavioral history.
[0421] 5. The server sends the selected deals to the smart glasses 214.
[0422] 6. The smart glasses 214 display the received deals and allow the user to take advantage of promotions and discounts.
[0423] The above is a specific example of an embodiment of the present invention. By linking the smart glasses 214 with a server, the user can easily obtain and use information on deals at stores and restaurants.
[0424] The process flow will be explained below.
[0425] Step 1: The server receives images and location information obtained from the smart glasses 214. Specifically, the smart glasses 214 recognize the logo or sign of the cafe and transmit the information to the server.
[0426] Step 2: The server analyzes the received image and location information. Using image recognition technology, it identifies the cafe's logo and sign and obtains the location information.
[0427] Step 3: The server searches the database for information about the cafe. The database contains information about the cafe's campaigns and discounts, and the server retrieves this information.
[0428] Step 4: The server selects the appropriate discount information based on the user's attributes and past behavior history. For example, if the user has visited the same cafe in the past and received a discount at that time, the server selects the discount information based on that information.
[0429] Step 5: The server transmits the selected deals to the smart glasses 214. Specifically, the server generates data for displaying the deals and transmits it to the smart glasses.
[0430] Step 6: The smart glasses 214 display the received special offers. The user can check the campaign and discount information at the cafe through the screen of the smart glasses 214.
[0431] Example 1
[0432] Next, a description will be given of Example 1. In the following description, the data processing device 12 is referred to as a "server" and the smart glasses 214 are referred to as a "terminal".
[0433] Conventionally, there has been no system that allows users to easily and quickly obtain and use information on deals on the spot when they visit a store. In addition, there is a demand for a method to realize more effective marketing by providing personalized information that takes into account the user's attributes and past behavioral history.
[0434] The specific processing by the specific processing unit 290 of the data processing device 12 in the first embodiment is realized by the following means.
[0435] In this invention, the server includes a means for acquiring store information based on image data and location information acquired from the mobile information terminal when the user visits the store using the mobile information terminal, a means for selecting advantageous information from the acquired store information, and a means for displaying the selected advantageous information on the mobile information terminal, thereby enabling the user to easily acquire and use personalized advantageous information in real time.
[0436] A "portable information terminal" is an electronic device that a user can carry with them, and is a terminal that has the function of taking pictures and acquiring location information.
[0437] A "server" is a central processing unit that receives data from multiple portable information terminals via a network, analyzes the data, and provides related information.
[0438] "Store information" is data related to a specific store, and includes image data of logos and signs, location information, campaign information, discount information, and the like.
[0439] "Bargain information" is promotional information that is beneficial to the user, such as special offers, discounts, and sales information for stores visited by the user.
[0440] "Image data" refers to digital data of an image captured by a user's mobile information terminal, and is the subject of analysis.
[0441] "Location information" refers to geographic coordinate data obtained through the GPS function of a mobile information terminal, for example.
[0442] "User attributes" refers to information relating to a user, such as age, sex, preferences, etc., that is used to provide personalized information.
[0443] "Past behavioral history" refers to a record of stores the user previously visited, products the user purchased, campaigns the user participated in, and the like, and indicates the user's behavioral patterns.
[0444] "Analysis" is the process in which the server extracts and identifies the necessary information based on the image data and location information it receives.
[0445] The present invention relates to a system for providing advantageous information to a user when the user visits a store using a mobile information terminal. Specific embodiments of the system will be described below.
[0446] Hardware and software configuration
[0447] First, the hardware and software required to realize this system will be described.
[0448] 1. Mobile information terminals:
[0449] A mobile information terminal is an electronic device that can be carried by a user and has the functions of taking pictures and acquiring location information. Specifically, a smartphone or a tablet is one such device.
[0450] It has a camera function and can acquire image data.
[0451] It has a GPS function and can obtain location information.
[0452] 2. Server:
[0453] A server is a central processing unit that receives data from multiple portable information terminals via a network, analyzes the data, and provides related information.
[0454] The servers are typically located in the cloud and have the necessary processing power and storage.
[0455] Image recognition libraries such as "OpenCV" are used for image analysis.
[0456] "Google Maps API" is used for location analysis.
[0457] Machine learning libraries such as "TensorFlow" are used to analyze user attributes and past behavioral history.
[0458] 3. Database:
[0459] The database stores store information (logos, signs, campaign information, discount information, etc.).
[0460] Program Processing
[0461] The server executes a program for selecting and providing advantageous information suitable for the user based on image data and location information obtained from the mobile information terminal.
[0462] 1. Data Acquisition:
[0463] Image data of a store's logo or signboard photographed by the terminal's camera is acquired.
[0464] The current location information is obtained using the device's GPS function.
[0465] 2. Data transmission:
[0466] The device transmits the acquired image data and location information to the server using Wi-Fi or 4G / 5G networks.
[0467] 3. Image Analysis:
[0468] The server uses OpenCV to analyze the image data and identify store logos and signs, extract feature points within the image, and match them with existing images in a database.
[0469] 4. Information Search:
[0470] Based on the analysis results, the server retrieves detailed information about the relevant store from the database, such as the cafe's menu, opening hours, and location.
[0471] 5. Obtaining Campaign Information:
[0472] The server obtains store campaign and discount information from the database. For example, it obtains information about "new menu item launch commemorative discount."
[0473] 6. Choose the best deals:
[0474] The server selects the most suitable campaign information for a user based on the user's attribute information (e.g., age, gender, preferences) and past behavioral history. To do this, a machine learning algorithm using "TensorFlow" is used.
[0475] 7. Information display:
[0476] The server transmits the selected deals to the terminal, which displays the received deals on its screen, allowing the user to use the information to receive campaigns and discounts.
[0477] Examples
[0478] For example, a specific example in which a user uses a mobile information terminal to visit a cafe is as follows.
[0479] 1. User enters the cafe:
[0480] A user carries a mobile information terminal and enters a cafe.
[0481] 2. The device captures a logo or sign:
[0482] Using the device's camera, the user takes a picture of the cafe's logo or sign.
[0483] 3. The device sends the image and location information to the server:
[0484] The device sends the captured images and location information obtained using the GPS function to the server.
[0485] 4. The server analyzes the image:
[0486] The server analyzes the received image using OpenCV and retrieves information about the corresponding cafe from a database.
[0487] 5. The server gets the campaign information:
[0488] The server retrieves the cafe's latest promotions and discounts from the same database.
[0489] 6. Server selects deals:
[0490] The server selects the most suitable deals by taking into account the user's attributes and past behavioral history.
[0491] 7. The server sends the deals to the device:
[0492] The server sends the selected deals to the terminal.
[0493] 8. Your device will display special offers:
[0494] The device will display the received deals on the screen, for example, "This cafe is now offering 50% off our new coffee menu!"
[0495] 9. User takes advantage of promotions and discounts:
[0496] Users can use the deals displayed on the screen to take advantage of coupons and discounts at the cafe.
[0497] Examples of prompt statements
[0498] "When a user visits a cafe, he or she takes a photo of the cafe's logo with a mobile information terminal and sends it to the server. The server then retrieves information about the cafe from the database, selects the most suitable deals based on the user's attributes and past behavioral history, and sends them to the mobile information terminal. The mobile information terminal displays the received information, and the user takes advantage of the cafe's campaign. Please explain the specific process."
[0499] The above is an embodiment of the present invention. This system allows users to easily obtain and use information about deals at stores.
[0500] The flow of the identification process in the first embodiment will be described with reference to FIG.
[0501] Step 1:
[0502] User enters the cafe
[0503] A user carries a mobile information terminal and enters a cafe.
[0504] Input: Cafe location, user attribute information
[0505] Output: The user is in the store
[0506] Step 2:
[0507] The device photographs logos and signs
[0508] Using the device's camera, the user takes a picture of the cafe's logo or sign.
[0509] Specific actions: The user activates the device's camera, frames the cafe's logo or sign, and presses the capture button.
[0510] Input: Image data of the cafe logo or sign
[0511] Output: Captured image data
[0512] Step 3:
[0513] The device sends the image and location information to the server.
[0514] The terminal transmits the acquired image data and the location information acquired by the GPS function to the server.
[0515] Specific operation: The device uses Wi-Fi or 4G / 5G networks to pack image data and location information into packets and upload them to the server.
[0516] Input: Image data, location information
[0517] Output: Image data and location information sent to the server
[0518] Step 4:
[0519] The server analyzes the image
[0520] The server analyzes the received image data using OpenCV and identifies the cafe's logo and sign.
[0521] Specific operation: Extract feature points in the image and match them with images in an existing database to identify the corresponding store.
[0522] Input: Image data
[0523] Output: Identified store information
[0524] Step 5:
[0525] The server retrieves cafe information from the database.
[0526] The server retrieves detailed information about the corresponding cafe from the database based on the identified store information.
[0527] Specific operation: The server sends a query to the database to obtain information about the relevant cafe, such as its menu, opening hours, and location.
[0528] Input: Identified store information
[0529] Output: Detailed information about the cafe
[0530] Step 6:
[0531] The server retrieves the campaign information.
[0532] The server retrieves the cafe's latest promotions and discounts from the same database.
[0533] Specific operation: The server sends a query to the database to obtain the relevant campaign information.
[0534] Input: Cafe details
[0535] Output: Retrieved campaign information
[0536] Step 7:
[0537] Server selects deals
[0538] The server selects the most suitable deals based on the user's attribute information (age, gender, preferences) and past behavioral history.
[0539] Specific operation: The server inputs the user's attribute data and past behavioral history into a machine learning model and calculates optimized campaign information.
[0540] Input: User attribute information, past behavior history, campaign information
[0541] Output: Selected deals
[0542] Step 8:
[0543] The server sends useful information to the terminal
[0544] The server sends the selected deals to the terminal.
[0545] How it works: The server uses Wi-Fi or 4G / 5G networks to pack deals into packets and upload them to the device.
[0546] Input: Selected Deals
[0547] Output: Deals sent to the terminal
[0548] Step 9:
[0549] The device displays useful information
[0550] The device will display the received deals on the screen.
[0551] Specific operation: A pop-up notification or a dedicated application screen is displayed on the device screen to provide the user with useful information.
[0552] Input: Deals sent to your device
[0553] Output: Deals displayed to the user
[0554] Step 10:
[0555] Users take advantage of promotions and discounts
[0556] Users can use the deals displayed on their devices to take advantage of coupons and discounts at cafes.
[0557] Specific operation: The user shows the coupon screen displayed on the terminal to the cafe cashier to receive the discount.
[0558] Input: Deals displayed to the user
[0559] Output: The user can use the campaign or discount.
[0560] (Application example 1)
[0561] Next, a description will be given of Application Example 1. In the following description, the data processing device 12 will be referred to as a "server" and the smart glasses 214 will be referred to as a "terminal".
[0562] In the conventional user experience in brick-and-mortar stores, users had to search for store and campaign information themselves, which was not very efficient. In particular, there were limited ways for users to find out what services and discounts were available in stores in real time, so building an information provision system that was effective for users was a challenge.
[0563] The specific processing by the specific processing unit 290 of the data processing device 12 in the application example 1 is realized by the following means.
[0564] In this invention, the server includes a means for acquiring image data and location information, a means for transmitting the acquired image data and location information to the server, a means for the server to search for store information from a database based on the image data and location information and provide the user with selected advantageous information, a means for displaying the advantageous information on the smart glasses, a means for recognizing store logos and signs from the image data, and a means for selecting optimal information in consideration of the user's past behavior history and attributes. This allows the user to acquire and use store information and campaign information in real time.
[0565] "Image data" refers to data that represents visual information acquired by a user using a smart device.
[0566] "Location information" is geographical data such as the latitude and longitude of the user's current location.
[0567] The "server" is a central processing unit that receives image data and location information and searches for store information from a database.
[0568] A "database" is a storage device for storing store information, campaign information, and discount information.
[0569] "Store information" is data including store logos and signs, campaign information, discount information, and the like.
[0570] "Good deals" refers to information about campaigns and discounts that are beneficial to users.
[0571] "Smart glasses" are wearable devices that users wear to display store information and special offers.
[0572] "Image recognition" is a technology that analyzes acquired image data and identifies specific objects or text.
[0573] "Past behavior history" is a record of actions and choices made by a user in the past.
[0574] "User attributes" refers to personal information about the user, such as age, sex, and interests.
[0575] A "generative AI model" is a model generated by artificial intelligence, and is a technology used for data analysis, prediction, and recommendations.
[0576] An embodiment of the present invention will be described below. A system for carrying out the present invention is mainly composed of a terminal for acquiring image data and location information, a server for processing data and providing information, and software for linking these components.
[0577] First, a user visits a physical store using a smart device (such as a smartphone or smart glasses). Image data and location information are acquired using the device's camera and GPS functions. Specific hardware used for this purpose include an on-board webcam, an external camera (e.g., Logitech C920), a built-in GPS, or an external GPS module (e.g., Holux M-1000C).
[0578] The acquired image data is sent from the terminal to a server. The server receives the image data and location information, and searches a database for store information. Google Cloud Vision API and Amazon Rekognition are used for image recognition, and Google Maps API is used for location analysis. The database uses cloud storage such as Amazon RDS.
[0579] The server selects deals suitable for the user based on the acquired store information and the user's past behavioral history and attributes. At this time, a generative AI model (e.g. TensorFlow) is used to analyze user attributes and past behavioral data and recommend the most suitable information.
[0580] The selected deals are then pushed to the user's smart device and displayed to the user, allowing the user to receive real-time deals and improve their shopping experience in-store.
[0581] For example, when a user enters a bookstore, the smart device recognizes the bookstore's logo and location information. The server provides discount information on new books based on the bookstore's information and past purchase history, and displays information about autograph sessions for members only.
[0582] Examples of specific prompts include the following:
[0583] When a user enters a bookstore using their smartphone, we want to display real-time information about discounts on new books and events for members only. Please tell us how to provide the appropriate information using image recognition and GPS information.
[0584] In this way, convenience for users in the store can be greatly improved.
[0585] The flow of the specific process in the application example 1 will be described with reference to FIG.
[0586] Step 1:
[0587] A user visits a store using a smart device. The device's camera takes a picture of the store's logo or sign, and the location information is obtained using the GPS function. The input is camera image data and location information, and the output is data to be sent to the server. The device prepares to send the image data captured by the camera and GPS coordinates to the server.
[0588] Step 2:
[0589] The terminal transmits the acquired image data and location information to the server. The input is the captured image and the acquired location information, and the output is data sent to the server. The terminal encodes the image data and transmits it to the server together with GPS information. At this time, the data is uploaded using an HTTP request.
[0590] Step 3:
[0591] The server processes the received image data and location information. The input is the image data and location information sent from the terminal, and the output is store information retrieved from a database. The server uses image recognition software (e.g. Google Cloud Vision API or Amazon Rekognition) to recognize store logos and signs from the image data. At the same time, it searches the database based on the location information and retrieves the corresponding store information.
[0592] Step 4:
[0593] The server selects deals based on the acquired store information and the user's past behavioral history and attributes. The input is store information, user behavioral history, and user attributes, and the output is the selected deals. The server analyzes the data using a generative AI model (e.g. TensorFlow) and selects the most suitable campaign and discount information.
[0594] Step 5:
[0595] The server sends the selected deals to the terminal. The input is the selected deals, and the output is data transmission to the terminal. The server uses an HTTP request to push the selected information to the user's smart device.
[0596] Step 6:
[0597] The terminal displays the received discount information to the user. The input is the discount information sent from the server, and the output is the information displayed to the user. The terminal uses the push notification function to display the discount information on the display of smart glasses or a smartphone. The user can check campaign information and discounts in real time.
[0598] In addition, an emotion engine that estimates the emotion of the user may be further combined. That is, the identification processing unit 290 may estimate the emotion of the user using the emotion identification model 59, and perform identification processing using the emotion of the user.
[0599] The embodiment for carrying out the present invention comprises the following elements.
[0600] 1. Terminal: A device used when a user visits a store or restaurant to acquire images and location information and display special offers.
[0601] 2. Server: A central processing unit that receives images and location information from the terminals, obtains information about stores and restaurants, and processes the information to select deals.
[0602] 3. Database: This is where the server stores data to search for store and restaurant information. This includes store logos, signboard information, campaign information, discount information, etc.
[0603] 4. Emotion engine: This engine has the function of recognizing the user's facial expressions and voice characteristics and estimating the user's emotions. The emotion engine is used in cooperation with the server as a means of selecting advantageous information based on the user's emotions.
[0604] 5. Program: This is the software that allows the server to process the information obtained from the smart glasses, select deals, and display them on the device. This includes image recognition technology, location information analysis, emotion engine utilization, and other processes.
[0605] (Specific examples)
[0606] A specific example is given where a user visits a cafe using smart glasses.
[0607] 1. A user puts on the smart glasses 214 and enters a cafe.
[0608] 2. The smart glasses 214 recognize the cafe's logo or sign and send that information to the server.
[0609] 3.The server searches the database for information about the cafe based on the received information and obtains campaign and discount information.
[0610] 4. The smart glasses 214 simultaneously transmit the user's facial expressions and voice characteristics to the emotion engine.
[0611] 5. The emotion engine estimates the user's emotion and returns the result to the server.
[0612] 6. The server selects deals based on the user's emotions. For example, if the user has a happy expression, the server selects discount deals.
[0613] 7. The server sends the selected deals to the smart glasses 214.
[0614] 8. The smart glasses 214 display information based on the received discount information and the user's emotions. For example, the smart glasses 214 display a message that matches the user's facial expression along with discount information.
[0615] The above is a specific example of an embodiment of the present invention. In addition to the cooperation between the smart glasses 214 and the server, the use of an emotion engine makes it possible to provide advantageous information that matches the user's emotions.
[0616] The process flow will be explained below.
[0617] Step 1: A user visits a cafe wearing smart glasses 214.
[0618] Step 2: The smart glasses 214 recognize the cafe's logo or sign and send that information to the server.
[0619] Step 3: The server searches the database for information about the cafe based on the received information.
[0620] Step 4: The smart glasses 214 simultaneously transmit the user's facial and vocal characteristics to the emotion engine.
[0621] Step 5: The emotion engine recognizes the user's facial expressions and voice characteristics and estimates the user's emotions.
[0622] Step 6: The server selects deals based on the user's emotions received from the emotion engine.
[0623] Step 7: The server sends the selected deals to the smart glasses 214.
[0624] Step 8: The smart glasses 214 make a display based on the received deals and the user's emotions.
[0625] Example 2
[0626] Next, a description will be given of Example 2. In the following description, the data processing device 12 is referred to as a "server" and the smart glasses 214 are referred to as a "terminal".
[0627] Conventional systems provide limited information about the stores that users visit, making it difficult to provide information based on the user's current emotions and individual characteristics. In addition, there is a lack of dynamic information provision to improve the user experience, which means that the functions of the user's smart device cannot be fully utilized.
[0628] The identification process by the identification processing unit 290 of the data processing device 12 in the second embodiment is realized by the following means. In this invention, the server includes a means for a terminal to acquire user's location information, a means for the terminal to acquire image data using a camera and recognize a store's logo or sign based on the image data, a means for the terminal to transmit image data and location information to the server, a means for the server to search a database to acquire store information, a means for analyzing the user's facial expression or voice characteristics using an emotion engine and estimating the user's emotion, a means for the server to select advantageous information based on the user's emotion, a means for the server to transmit the advantageous information to the terminal, and a means for the terminal to display the advantageous information to the user. This makes it possible to acquire information on the store visited by the user from multiple angles and provide advantageous information personalized according to the user's emotion and situation.
[0629] A "terminal" is an electronic device that can be carried by a user and has the functions of acquiring image data, recording location information, and communicating with a server.
[0630] "Location information" is data indicating the geographical location of the user's current location, and is obtained using the GPS function.
[0631] "Image data" is visual information captured using the device's camera, and includes, for example, images of store logos and signs.
[0632] A "server" is a computer system that centrally processes and stores data, and is responsible for receiving and processing data sent from terminals.
[0633] A "database" is a collection of information that stores store logos, signboard information, campaign information, discount information, etc., and is used by the server for search and access.
[0634] An "emotion engine" is software or hardware that analyzes the user's facial expressions and vocal characteristics and estimates the user's emotional state.
[0635] "Bargain information" is information about discounts, campaigns, and the like that is useful to the user, and is selected based on the user's emotions and the stores that the user visits.
[0636] "User facial or vocal features" refers to data such as facial expressions and tone of voice that are used to estimate the user's emotions.
[0637] "Means for transmitting" refers to a communication means for transmitting data from a terminal to a server, and includes, for example, Wi-Fi and mobile data.
[0638] "Display means" refers to a function for visually or audibly presenting the information received by the terminal to the user, and includes screen display and audio guidance.
[0639] The system for implementing the present invention is composed of a user, a terminal, a server, a database, an emotion engine, and a program that connects these. Below, we will specifically explain how these elements work together to realize the invention.
[0640] First, the core of the system is the terminal used by the user. This terminal is designed as a smart device (e.g., smartphone, smart glasses) and has the following functions: It acquires image data using its camera function and collects location information using its GPS function. It also has sensors that detect the user's facial expressions and voice characteristics.
[0641] Next, a server is required to process the data acquired by the terminal. The server performs the following processes.
[0642] 1. Receive image data and location information sent from the device.
[0643] 2. Access the database and search for and retrieve image data of store logos and signs, as well as campaign and discount information.
[0644] 3. Use an emotion engine to estimate the user's emotions from their facial expressions and voice characteristics.
[0645] 4. Select the best deal from the information in the database based on user sentiment.
[0646] 5. Send the selected deals back to your device.
[0647] The database stores detailed data such as store and restaurant logos, signboard information, campaign information, discount information, etc. This database is maintained so that the server can efficiently search and access it.
[0648] The emotion engine estimates emotions by analyzing the user's facial expressions and vocal characteristics. For example, a generative AI model is used to classify emotional states such as "happy," "sad," and "surprised." This model is capable of analyzing the user's real-time emotions with high accuracy.
[0649] As a concrete example, the process when a user visits a cafe is shown below.
[0650] 1. A user enters a cafe with their smart device.
[0651] 2. The device recognizes the cafe's logo or sign, and sends image data and location information to the server.
[0652] 3. The server searches the database for information about the cafe and retrieves campaign and discount information.
[0653] 4. At the same time, the device transmits the user's facial expressions and voice characteristics to the emotion engine.
[0654] 5. The emotion engine estimates the user's emotion and returns the result to the server.
[0655] 6. The server selects deals based on the user's emotions. For example, if the user has a happy expression, the server selects discount deals.
[0656] 7. The server sends the selected deals to the device.
[0657] 8. The device displays information based on the discount information received and the user's emotions. For example, a message that matches the user's facial expression is displayed along with discount information.
[0658] An example of a prompt sentence might be:
[0659] When a user enters a cafe, the smart device scans the cafe's logo or sign and sends the image and location information to the server. The server searches the corresponding cafe information from the database and analyzes the user's facial expressions and voice using an emotion engine. Finally, the server selects deals that match the user's emotions and displays them on the smart device.
[0660] As described above, the present invention provides a system that obtains information about stores visited by a user from multiple angles and provides the user with personalized advantageous information according to the user's emotions and circumstances.
[0661] The flow of the identification process in the second embodiment will be described with reference to FIG.
[0662] Specific flow of system program processing
[0663] Step 1:
[0664] A user enters a store.
[0665] This process starts when a user enters a store with a smart device. For example, assume that the user enters a cafe.
[0666] Step 2:
[0667] The device captures images and location information.
[0668] The device uses its camera to take a picture of the store's logo or sign, and acquires location information using its GPS function. This inputs the image of the cafe's logo and the coordinate data of the current location into the device. The output is the captured image data and location information data.
[0669] Step 3:
[0670] The terminal transmits the information to the server.
[0671] The terminal transmits the acquired image data and location information data to the server. The input is the image data and location information data stored in the terminal, and the data is transmitted by communicating it to the server. The output is the data that has been transferred to the server.
[0672] Step 4:
[0673] The server searches the database.
[0674] The server searches the database based on the received image data and location information. The input is the image data and location information sent to the server, and the output is the corresponding store information (e.g., store name, campaign information, discount information, etc.) retrieved from the database.
[0675] Step 5:
[0676] The server uses an emotion engine to estimate the user's emotion.
[0677] The server inputs the facial expression and voice feature data of the user sent from the terminal into the emotion engine. The emotion engine analyzes this data and estimates the user's emotion. The input is the facial expression and voice feature data of the user, and the output is the user's emotional state (for example, "happy," "sad," "surprised," etc.).
[0678] Step 6:
[0679] The server selects the deals.
[0680] The server selects the optimal deal from the information in the database based on the estimated user emotion. The input is the user's emotional state and store information retrieved from the database, and the output is the deal information (e.g., specific discounts and campaign information) to be provided to the user.
[0681] Step 7:
[0682] The server transmits information to the terminal.
[0683] The server sends the selected deals to the terminal. The input is the deals stored in the server, and the data is sent by communicating it to the terminal. The output is the data that has been transferred to the terminal.
[0684] Step 8:
[0685] The terminal displays the information.
[0686] The terminal displays the received discount information to the user. The input is the discount information sent from the server, and the output is the information that the user can confirm visually or audibly. For example, the message "All items are 10% off today!" is displayed on the terminal screen.
[0687] Through the above processing steps, the present invention makes it possible to obtain information about stores visited by a user from multiple angles and provide personalized deals information according to the user's emotions and circumstances.
[0688] (Application example 2)
[0689] Next, a description will be given of Application Example 2. In the following description, the data processing device 12 will be referred to as a "server" and the smart glasses 214 will be referred to as a "terminal".
[0690] Conventional store information systems can only provide uniform deals to users, and it is difficult to provide information tailored to the emotions and circumstances of individual users. In addition, since the system does not take into account the emotions of users, the information provided may not always be appropriate, making it difficult to improve user satisfaction.
[0691] The identification process by the identification processing unit 290 of the data processing device 12 in the application example 2 is realized by the following means. In this invention, the server includes a means for recognizing and acquiring store information from image data, a means for acquiring the user's facial expression data together with the image data, a means for selecting advantageous information based on the acquired store information, a means for estimating an emotion based on the user's facial expression data and optimizing the advantageous information based on the emotion, and a means for displaying the optimized advantageous information on the smart glasses. This makes it possible to provide personalized information based on the user's emotions, and is expected to improve user satisfaction.
[0692] The "user-worn device" is a device worn by a user visiting a store, and has the function of acquiring image data and facial expression data.
[0693] The "store information acquisition means" is a function for recognizing store logos and signs from image data obtained from a user-worn device and acquiring that information.
[0694] The "facial expression data acquisition means" is a function for capturing the user's facial expression and acquiring that data.
[0695] The "information selection means" is a function for selecting advantageous information to be presented to the user based on the acquired store information.
[0696] The "emotion estimation means" is a function for estimating the user's emotion based on the acquired facial expression data of the user.
[0697] The "information optimization means" is a function for optimizing selected advantageous information based on the estimated user's emotions.
[0698] The "information display means" is a function for displaying optimized advantageous information on a user-worn device.
[0699] DETAILED DESCRIPTION OF THE PREFERRED EMBODIMENTS
[0700] The system of the present invention consists of smart glasses worn by the user, a server that processes images and location information, and an emotion engine.
[0701] 1. Hardware and Software Configuration
[0702] Smart Glasses
[0703] The smart glasses, worn by the user, have a camera and a display function. The camera captures store logos and signs, and also obtains the user's facial expression data. These data are then sent to a server.
[0704] server
[0705] The server uses the following hardware and software:
[0706] Hardware: High performance processor, storage device
[0707] Software: Python, OpenCV, requests library
[0708] The server analyzes the image data acquired from the smart glasses to recognize store information, and also uses an emotion engine to estimate the user's emotions from their facial expression data.
[0709] Emotion Engine
[0710] The emotion engine is software that analyzes the user's facial expression data and estimates their emotions. The basic emotion estimation algorithm uses a machine learning model.
[0711] 2. Data flow and processing
[0712] The image data acquired by the smart glasses is sent to a server. The server recognizes the store's logo or signboard from the image data and retrieves store information from a database. Next, the server sends the user's facial expression data acquired at the same time to an emotion engine to estimate the user's emotion.
[0713] Store information acquisition method
[0714] The server recognizes logos and signs from the image data and retrieves related store information from the database based on this. Specifically, the server uses the OpenCV library for image recognition.
[0715] Emotion estimation means
[0716] The user's facial expression data is sent to the emotion engine, which uses machine learning to classify the user's emotions into categories such as "happy," "sad," "excited," and "calm."
[0717] Information optimization measures
[0718] The server optimizes the deals based on the user's emotions. For example, if the user has a happy expression, the server will provide discount information.
[0719] Means of selecting and displaying information
[0720] The server sends the optimized information to the smart glasses to display to the user, including messages personalized to the user's emotions.
[0721] 3. Specific Examples
[0722] As a concrete example, consider the case where a user visits a cafe wearing smart glasses.
[0723] 1. When a user enters a cafe, the camera in the smart glasses recognizes the cafe's logo and sign and transmits it to the server.
[0724] 2. The server retrieves information about recognized cafes from the database.
[0725] 3. At the same time, the smart glasses capture the user's facial expression, and if the emotion engine estimates it to be "happy," the server will offer a "Special Happy Hour Discount."
[0726] 4. This information is then displayed on the smart glasses, allowing users to receive real-time deals.
[0727] Example prompts to be input to the generative AI model
[0728] When a smart glasses user enters a store, the store's logo and sign are recognized and the information is sent to the server. Create a program that then analyzes the user's emotions and displays discount and campaign information based on the emotions on the display. Show how to configure it in Python.
[0729] The flow of the specific process in the application example 2 will be described with reference to FIG.
[0730] Step 1:
[0731] Smart glasses capture user actions.
[0732] Input: Image data including the in-store environment and the user's facial expressions captured using the smart glasses camera.
[0733] Output: Image files containing store logos, signage, and user facial expressions.
[0734] How it works: The camera in the smart glasses captures images in real time and stores them in local storage.
[0735] Step 2:
[0736] The terminal transmits the captured image data to the server.
[0737] Input: Image data captured by smart glasses.
[0738] Output: Image data sent to the server.
[0739] Specific operation: The terminal communicates with the server via the Internet using HTTP, and sends image data as a POST request.
[0740] Step 3:
[0741] The server analyzes the image data it receives and recognizes store logos and signs.
[0742] Input: Image data including store logos and signs sent from the device.
[0743] Output: IDs of recognized stores and related information.
[0744] Specific operation: The server processes image data using the OpenCV library and detects and recognizes store logos and signs.
[0745] Step 4:
[0746] The store information recognized by the server is obtained from the database.
[0747] Input: The ID of a recognized store.
[0748] Output: Store details (e.g. store name, campaign information, discount information, etc.).
[0749] Specific operation: The server executes a database query using the store ID as a key to obtain related store information.
[0750] Step 5:
[0751] The terminal captures the user's facial expression data and transmits it to the server.
[0752] Input: Image data containing the user's facial expressions captured by the smart glasses.
[0753] Output: Facial expression data sent to the server.
[0754] Specific operation: The terminal communicates with the server via the Internet via HTTP and sends facial expression data as a POST request.
[0755] Step 6:
[0756] The server analyzes the facial expression data and uses an emotion engine to infer the user's emotions.
[0757] Input: User's facial expression data sent from the device.
[0758] Output: An estimated user emotion classification (e.g. happy, sad, excited, etc.).
[0759] Specific operation: The server uses an emotion engine (machine learning model) to analyze facial expression data and classify emotions.
[0760] Step 7:
[0761] The server selects the best deals based on the user's sentiment.
[0762] Input: Estimated user sentiment and store information.
[0763] Output: Optimized deals based on sentiment.
[0764] Specific operation: Based on the store information, the server selects the deals (discount information and campaign information) that best suit the user's emotions.
[0765] Step 8:
[0766] The server sends the selected information to the terminal and displays it on the smart glasses.
[0767] Enter: Optimized Deals.
[0768] Output: Deals displayed on the smart glasses.
[0769] Specific operation: The server sends the best deals to the device and displays the information on the smart glasses display.
[0770] This allows users to receive real-time, personalized deals when they visit a store.
[0771] The specific processing unit 290 transmits the result of the specific processing to the smart glasses 214. In the smart glasses 214, the control unit 46A causes the speaker 240 to output the result of the specific processing. The microphone 238 acquires a voice indicating a user input for the result of the specific processing. The control unit 46A transmits the voice data indicating the user input acquired by the microphone 238 to the data processing device 12. In the data processing device 12, the specific processing unit 290 acquires the voice data.
[0772] The data generation model 58 is a so-called generative AI (Artificial Intelligence). An example of the data generation model 58 is ChatGPT (Internet search<URL: https: / / openai.com / blog / chatgpt> ), Gemini (Internet search <url: https: gemini.google.com ?hl="ja">) and other generation AIs. The data generation model 58 is obtained by making a neural network perform deep learning. A prompt including an instruction is input to the data generation model 58, and inference data such as voice data indicating a voice, text data indicating a text, and image data indicating an image is input. The data generation model 58 performs inference on the input inference data according to the instruction indicated by the prompt, and outputs the inference result in a data format such as voice data and text data. Here, inference refers to, for example, analysis, classification, prediction, and / or summarization.
[0773] In the above embodiment, an example was given in which the specific processing is performed by the data processing device 12, but the technology of the present disclosure is not limited to this, and the specific processing may be performed by the smart glasses 214.
[0774] [Third embodiment]
[0775] FIG. 5 shows an example of the configuration of a data processing system 310 according to the third embodiment.
[0776] 5, the data processing system 310 includes the data processing device 12 and a headset type terminal 314. An example of the data processing device 12 is a server.
[0777] The data processing device 12 includes a computer 22, a database 24, and a communication I / F 26. The computer 22 is an example of a "computer" according to the technology of the present disclosure. The computer 22 includes a processor 28, a RAM 30, and a storage 32. The processor 28, the RAM 30, and the storage 32 are connected to a bus 34. The database 24 and the communication I / F 26 are also connected to the bus 34. The communication I / F 26 is connected to a network 54. Examples of the network 54 include a wide area network (WAN) and / or a local area network (LAN).
[0778] The headset type terminal 314 includes a computer 36, a microphone 238, a speaker 240, a camera 42, a communication I / F 44, and a display 343. The computer 36 includes a processor 46, a RAM 48, and a storage 50. The processor 46, the RAM 48, and the storage 50 are connected to a bus 52. The microphone 238, the speaker 240, the camera 42, and the display 343 are also connected to the bus 52.
[0779] The microphone 238 receives instructions and the like from the user 20 by receiving voice uttered by the user 20. The microphone 238 captures the voice uttered by the user 20, converts the captured voice into audio data, and outputs it to the processor 46. The speaker 240 outputs the voice according to instructions from the processor 46.
[0780] Camera 42 is a small digital camera equipped with an optical system including a lens, an aperture, and a shutter, and an imaging element such as a CMOS (Complementary Metal-Oxide-Semiconductor) image sensor or a CCD (Charge Coupled Device) image sensor, and captures images of the surroundings of user 20 (e.g., an imaging range defined by an angle of view equivalent to the width of the field of vision of an average healthy person).
[0781] The communication I / F 44 is connected to a network 54. The communication I / Fs 44 and 26 are responsible for transmitting and receiving various types of information between the processor 46 and the processor 28 via the network 54. The transmission and reception of various types of information between the processor 46 and the processor 28 using the communication I / Fs 44 and 26 is performed in a secure state.
[0782] Fig. 6 shows an example of the main functions of the data processing device 12 and the headset type terminal 314. As shown in Fig. 6, in the data processing device 12, a specific process is performed by the processor 28. A specific process program 56 is stored in the storage 32.
[0783] The specific processing program 56 is an example of a "program" according to the technology of the present disclosure. The processor 28 reads the specific processing program 56 from the storage 32, and executes the read specific processing program 56 on the RAM 30. The specific processing is realized by the processor 28 operating as a specific processing unit 290 in accordance with the specific processing program 56 executed on the RAM 30.
[0784] The storage 32 stores a data generation model 58 and an emotion identification model 59. The data generation model 58 and the emotion identification model 59 are used by the identification processing unit 290.
[0785] In the headset type terminal 314, the reception output process is performed by the processor 46. A reception output program 60 is stored in the storage 50. The processor 46 reads the reception output program 60 from the storage 50, and executes the read reception output program 60 on the RAM 48. The reception output process is realized by the processor 46 operating as the control unit 46A in accordance with the reception output program 60 executed on the RAM 48.
[0786] Next, a description will be given of the identification process performed by the identification processing unit 290 of the data processing device 12. In the following description, the data processing device 12 will be referred to as the "server", and the headset type terminal 314 will be referred to as the "terminal".
[0787] The embodiment for carrying out the present invention comprises the following elements.
[0788] 1. Terminal: A device used when a user visits a store or restaurant to acquire images and location information and display special offers.
[0789] 2. Server: A central processing unit that receives images and location information from the terminals, obtains information about stores and restaurants, and processes the information to select deals.
[0790] 3. Database: This is where the server stores data to search for store and restaurant information. This includes store logos, signboard information, campaign information, discount information, etc.
[0791] 4. Program: Software that allows the server to process information obtained from the terminal, select advantageous information, and display it on the terminal. This includes image recognition technology, location information analysis, and consideration of user attributes and past behavioral history.
[0792] (Specific examples)
[0793] A specific example in which a user uses a headset type terminal 314 to visit a cafe will be shown.
[0794] 1. The user calls the headset terminal 314 and enters the cafe.
[0795] 2. The headset terminal 314 recognizes the cafe's logo or sign and transmits that information to the server.
[0796] 3.The server searches the database for information about the cafe based on the received information and obtains campaign and discount information.
[0797] 4. The server selects deals that are appropriate for the user, taking into account the user's attributes and past behavioral history.
[0798] 5. The server transmits the selected deals to the headset terminal 314.
[0799] 6. The headset terminal 314 displays the received deals information, allowing the user to take advantage of campaigns and discounts.
[0800] The above is a specific example of an embodiment of the present invention. By linking the headset terminal 314 with a server, the user can easily obtain and use information about deals at stores and restaurants.
[0801] The process flow will be explained below.
[0802] Step 1: The server receives images and location information obtained from the headset terminal 314. Specifically, the headset terminal 314 recognizes the logo or sign of the cafe, and transmits that information to the server.
[0803] Step 2: The server analyzes the received image and location information. Using image recognition technology, it identifies the cafe's logo and sign and obtains the location information.
[0804] Step 3: The server searches the database for information about the cafe. The database contains information about the cafe's campaigns and discounts, and the server retrieves this information.
[0805] Step 4: The server selects the appropriate discount information based on the user's attributes and past behavior history. For example, if the user has visited the same cafe in the past and received a discount at that time, the server selects the discount information based on that information.
[0806] Step 5: The server transmits the selected advantageous information to the headset type terminal 314. Specifically, data for displaying the advantageous information is generated and transmitted to the headset type terminal 314.
[0807] Step 6: The headset type terminal 314 displays the received special offer information. The user can check the campaign and discount information at the cafe through the screen of the headset type terminal 314.
[0808] Example 1
[0809] Next, a description will be given of Example 1. In the following description, the data processing device 12 will be referred to as a "server" and the headset type terminal 314 will be referred to as a "terminal".
[0810] Conventionally, there has been no system that allows users to easily and quickly obtain and use information on deals on the spot when they visit a store. In addition, there is a demand for a method to realize more effective marketing by providing personalized information that takes into account the user's attributes and past behavioral history.
[0811] The specific processing by the specific processing unit 290 of the data processing device 12 in the first embodiment is realized by the following means.
[0812] In this invention, the server includes a means for acquiring store information based on image data and location information acquired from the mobile information terminal when the user visits the store using the mobile information terminal, a means for selecting advantageous information from the acquired store information, and a means for displaying the selected advantageous information on the mobile information terminal, thereby enabling the user to easily acquire and use personalized advantageous information in real time.
[0813] A "portable information terminal" is an electronic device that a user can carry with them, and is a terminal that has the function of taking pictures and acquiring location information.
[0814] A "server" is a central processing unit that receives data from multiple portable information terminals via a network, analyzes the data, and provides related information.
[0815] "Store information" is data related to a specific store, and includes image data of logos and signs, location information, campaign information, discount information, and the like.
[0816] "Bargain information" is promotional information that is beneficial to the user, such as special offers, discounts, and sales information for stores visited by the user.
[0817] "Image data" refers to digital data of an image captured by a user's mobile information terminal, and is the subject of analysis.
[0818] "Location information" refers to geographic coordinate data obtained through the GPS function of a mobile information terminal, for example.
[0819] "User attributes" refers to information relating to a user, such as age, sex, preferences, etc., that is used to provide personalized information.
[0820] "Past behavioral history" refers to a record of stores the user previously visited, products the user purchased, campaigns the user participated in, and the like, and indicates the user's behavioral patterns.
[0821] "Analysis" is the process in which the server extracts and identifies the necessary information based on the image data and location information it receives.
[0822] The present invention relates to a system for providing advantageous information to a user when the user visits a store using a mobile information terminal. Specific embodiments of the system will be described below.
[0823] Hardware and software configuration
[0824] First, the hardware and software required to realize this system will be described.
[0825] 1. Mobile information terminals:
[0826] A mobile information terminal is an electronic device that can be carried by a user and has the functions of taking pictures and acquiring location information. Specifically, a smartphone or a tablet is one such device.
[0827] It has a camera function and can acquire image data.
[0828] It has a GPS function and can obtain location information.
[0829] 2. Server:
[0830] A server is a central processing unit that receives data from multiple portable information terminals via a network, analyzes the data, and provides related information.
[0831] The servers are typically located in the cloud and have the necessary processing power and storage.
[0832] Image recognition libraries such as "OpenCV" are used for image analysis.
[0833] "Google Maps API" is used for location analysis.
[0834] Machine learning libraries such as "TensorFlow" are used to analyze user attributes and past behavioral history.
[0835] 3. Database:
[0836] The database stores store information (logos, signs, campaign information, discount information, etc.).
[0837] Program Processing
[0838] The server executes a program for selecting and providing advantageous information suitable for the user based on image data and location information obtained from the mobile information terminal.
[0839] 1. Data Acquisition:
[0840] Image data of a store's logo or signboard photographed by the terminal's camera is acquired.
[0841] The current location information is obtained using the device's GPS function.
[0842] 2. Data transmission:
[0843] The device transmits the acquired image data and location information to the server using Wi-Fi or 4G / 5G networks.
[0844] 3. Image Analysis:
[0845] The server uses OpenCV to analyze the image data and identify store logos and signs, extract feature points within the image, and match them with existing images in a database.
[0846] 4. Information Search:
[0847] Based on the analysis results, the server retrieves detailed information about the relevant store from the database, such as the cafe's menu, opening hours, and location.
[0848] 5. Obtaining Campaign Information:
[0849] The server obtains store campaign and discount information from the database. For example, it obtains information about "new menu item launch commemorative discount."
[0850] 6. Choose the best deals:
[0851] The server selects the most suitable campaign information for a user based on the user's attribute information (e.g., age, gender, preferences) and past behavioral history. To do this, a machine learning algorithm using "TensorFlow" is used.
[0852] 7. Information display:
[0853] The server transmits the selected deals to the terminal, which displays the received deals on its screen, allowing the user to use the information to receive campaigns and discounts.
[0854] Examples
[0855] For example, a specific example in which a user uses a mobile information terminal to visit a cafe is as follows.
[0856] 1. User enters the cafe:
[0857] A user carries a mobile information terminal and enters a cafe.
[0858] 2. The device captures a logo or sign:
[0859] Using the device's camera, the user takes a picture of the cafe's logo or sign.
[0860] 3. The device sends the image and location information to the server:
[0861] The device sends the captured images and location information obtained using the GPS function to the server.
[0862] 4. The server analyzes the image:
[0863] The server analyzes the received image using OpenCV and retrieves information about the corresponding cafe from a database.
[0864] 5. The server gets the campaign information:
[0865] The server retrieves the cafe's latest promotions and discounts from the same database.
[0866] 6. Server selects deals:
[0867] The server selects the most suitable deals by taking into account the user's attributes and past behavioral history.
[0868] 7. The server sends the deals to the device:
[0869] The server sends the selected deals to the terminal.
[0870] 8. Your device will display special offers:
[0871] The device will display the received deals on the screen, for example, "This cafe is now offering 50% off our new coffee menu!"
[0872] 9. User takes advantage of promotions and discounts:
[0873] Users can use the deals displayed on the screen to take advantage of coupons and discounts at the cafe.
[0874] Examples of prompt statements
[0875] "When a user visits a cafe, he or she takes a photo of the cafe's logo with a mobile information terminal and sends it to the server. The server then retrieves information about the cafe from the database, selects the most suitable deals based on the user's attributes and past behavioral history, and sends them to the mobile information terminal. The mobile information terminal displays the received information, and the user takes advantage of the cafe's campaign. Please explain the specific process."
[0876] The above is an embodiment of the present invention. This system allows users to easily obtain and use information about deals at stores.
[0877] The flow of the identification process in the first embodiment will be described with reference to FIG.
[0878] Step 1:
[0879] User enters the cafe
[0880] A user carries a mobile information terminal and enters a cafe.
[0881] Input: Cafe location, user attribute information
[0882] Output: The user is in the store
[0883] Step 2:
[0884] The device photographs logos and signs
[0885] Using the device's camera, the user takes a picture of the cafe's logo or sign.
[0886] Specific actions: The user activates the device's camera, frames the cafe's logo or sign, and presses the capture button.
[0887] Input: Image data of the cafe logo or sign
[0888] Output: Captured image data
[0889] Step 3:
[0890] The device sends the image and location information to the server.
[0891] The terminal transmits the acquired image data and the location information acquired by the GPS function to the server.
[0892] Specific operation: The device uses Wi-Fi or 4G / 5G networks to pack image data and location information into packets and upload them to the server.
[0893] Input: Image data, location information
[0894] Output: Image data and location information sent to the server
[0895] Step 4:
[0896] The server analyzes the image
[0897] The server analyzes the received image data using OpenCV and identifies the cafe's logo and sign.
[0898] Specific operation: Extract feature points in the image and match them with images in an existing database to identify the corresponding store.
[0899] Input: Image data
[0900] Output: Identified store information
[0901] Step 5:
[0902] The server retrieves cafe information from the database.
[0903] The server retrieves detailed information about the corresponding cafe from the database based on the identified store information.
[0904] Specific operation: The server sends a query to the database to obtain information about the relevant cafe, such as its menu, opening hours, and location.
[0905] Input: Identified store information
[0906] Output: Detailed information about the cafe
[0907] Step 6:
[0908] The server retrieves the campaign information.
[0909] The server retrieves the cafe's latest promotions and discounts from the same database.
[0910] Specific operation: The server sends a query to the database to obtain the relevant campaign information.
[0911] Input: Cafe details
[0912] Output: Retrieved campaign information
[0913] Step 7:
[0914] Server selects deals
[0915] The server selects the most suitable deals based on the user's attribute information (age, gender, preferences) and past behavioral history.
[0916] Specific operation: The server inputs the user's attribute data and past behavioral history into a machine learning model and calculates optimized campaign information.
[0917] Input: User attribute information, past behavior history, campaign information
[0918] Output: Selected deals
[0919] Step 8:
[0920] The server sends useful information to the terminal
[0921] The server sends the selected deals to the terminal.
[0922] How it works: The server uses Wi-Fi or 4G / 5G networks to pack deals into packets and upload them to the device.
[0923] Input: Selected Deals
[0924] Output: Deals sent to the terminal
[0925] Step 9:
[0926] The device displays useful information
[0927] The device will display the received deals on the screen.
[0928] Specific operation: A pop-up notification or a dedicated application screen is displayed on the device screen to provide the user with useful information.
[0929] Input: Deals sent to your device
[0930] Output: Deals displayed to the user
[0931] Step 10:
[0932] Users take advantage of promotions and discounts
[0933] Users can use the deals displayed on their devices to take advantage of coupons and discounts at cafes.
[0934] Specific operation: The user shows the coupon screen displayed on the terminal to the cafe cashier to receive the discount.
[0935] Input: Deals displayed to the user
[0936] Output: The user can use the campaign or discount.
[0937] (Application example 1)
[0938] Next, a description will be given of Application Example 1. In the following description, the data processing device 12 will be referred to as a "server" and the headset type terminal 314 will be referred to as a "terminal."
[0939] In the conventional user experience in brick-and-mortar stores, users had to search for store and campaign information themselves, which was not very efficient. In particular, there were limited ways for users to find out what services and discounts were available in stores in real time, so building an information provision system that was effective for users was a challenge.
[0940] The specific processing by the specific processing unit 290 of the data processing device 12 in the application example 1 is realized by the following means.
[0941] In this invention, the server includes a means for acquiring image data and location information, a means for transmitting the acquired image data and location information to the server, a means for the server to search for store information from a database based on the image data and location information and provide the user with selected advantageous information, a means for displaying the advantageous information on the smart glasses, a means for recognizing store logos and signs from the image data, and a means for selecting optimal information in consideration of the user's past behavior history and attributes. This allows the user to acquire and use store information and campaign information in real time.
[0942] "Image data" refers to data that represents visual information acquired by a user using a smart device.
[0943] "Location information" is geographical data such as the latitude and longitude of the user's current location.
[0944] The "server" is a central processing unit that receives image data and location information and searches for store information from a database.
[0945] A "database" is a storage device for storing store information, campaign information, and discount information.
[0946] "Store information" is data including store logos and signs, campaign information, discount information, and the like.
[0947] "Good deals" refers to information about campaigns and discounts that are beneficial to users.
[0948] "Smart glasses" are wearable devices that users wear to display store information and special offers.
[0949] "Image recognition" is a technology that analyzes acquired image data and identifies specific objects or text.
[0950] "Past behavior history" is a record of actions and choices made by a user in the past.
[0951] "User attributes" refers to personal information about the user, such as age, sex, and interests.
[0952] A "generative AI model" is a model generated by artificial intelligence, and is a technology used for data analysis, prediction, and recommendations.
[0953] An embodiment of the present invention will be described below. A system for carrying out the present invention is mainly composed of a terminal for acquiring image data and location information, a server for processing data and providing information, and software for linking these components.
[0954] First, a user visits a physical store using a smart device (such as a smartphone or smart glasses). Image data and location information are acquired using the device's camera and GPS functions. Specific hardware used for this purpose include an on-board webcam, an external camera (e.g., Logitech C920), a built-in GPS, or an external GPS module (e.g., Holux M-1000C).
[0955] The acquired image data is sent from the terminal to a server. The server receives the image data and location information, and searches a database for store information. Google Cloud Vision API and Amazon Rekognition are used for image recognition, and Google Maps API is used for location analysis. The database uses cloud storage such as Amazon RDS.
[0956] The server selects deals suitable for the user based on the acquired store information and the user's past behavioral history and attributes. At this time, a generative AI model (e.g. TensorFlow) is used to analyze user attributes and past behavioral data and recommend the most suitable information.
[0957] The selected deals are then pushed to the user's smart device and displayed to the user, allowing the user to receive real-time deals and improve their shopping experience in-store.
[0958] For example, when a user enters a bookstore, the smart device recognizes the bookstore's logo and location information. The server provides discount information on new books based on the bookstore's information and past purchase history, and displays information about autograph sessions for members only.
[0959] Examples of specific prompts include the following:
[0960] When a user enters a bookstore using their smartphone, we want to display real-time information about discounts on new books and events for members only. Please tell us how to provide the appropriate information using image recognition and GPS information.
[0961] In this way, convenience for users in the store can be greatly improved.
[0962] The flow of the specific process in the application example 1 will be described with reference to FIG.
[0963] Step 1:
[0964] A user visits a store using a smart device. The device's camera takes a picture of the store's logo or sign, and the location information is obtained using the GPS function. The input is camera image data and location information, and the output is data to be sent to the server. The device prepares to send the image data captured by the camera and GPS coordinates to the server.
[0965] Step 2:
[0966] The terminal transmits the acquired image data and location information to the server. The input is the captured image and the acquired location information, and the output is data sent to the server. The terminal encodes the image data and transmits it to the server together with GPS information. At this time, the data is uploaded using an HTTP request.
[0967] Step 3:
[0968] The server processes the received image data and location information. The input is the image data and location information sent from the terminal, and the output is store information retrieved from a database. The server uses image recognition software (e.g. Google Cloud Vision API or Amazon Rekognition) to recognize store logos and signs from the image data. At the same time, it searches the database based on the location information and retrieves the corresponding store information.
[0969] Step 4:
[0970] The server selects deals based on the acquired store information and the user's past behavioral history and attributes. The input is store information, user behavioral history, and user attributes, and the output is the selected deals. The server analyzes the data using a generative AI model (e.g. TensorFlow) and selects the most suitable campaign and discount information.
[0971] Step 5:
[0972] The server sends the selected deals to the terminal. The input is the selected deals, and the output is data transmission to the terminal. The server uses an HTTP request to push the selected information to the user's smart device.
[0973] Step 6:
[0974] The terminal displays the received discount information to the user. The input is the discount information sent from the server, and the output is the information displayed to the user. The terminal uses the push notification function to display the discount information on the display of smart glasses or a smartphone. The user can check campaign information and discounts in real time.
[0975] In addition, an emotion engine that estimates the emotion of the user may be further combined. That is, the identification processing unit 290 may estimate the emotion of the user using the emotion identification model 59, and perform identification processing using the emotion of the user.
[0976] The embodiment for carrying out the present invention comprises the following elements.
[0977] 1. Terminal: A device used when a user visits a store or restaurant to acquire images and location information and display special offers.
[0978] 2. Server: A central processing unit that receives images and location information from the terminals, obtains information about stores and restaurants, and processes the information to select deals.
[0979] 3. Database: This is where the server stores data to search for store and restaurant information. This includes store logos, signboard information, campaign information, discount information, etc.
[0980] 4. Emotion engine: This engine has the function of recognizing the user's facial expressions and voice characteristics and estimating the user's emotions. The emotion engine is used in cooperation with the server as a means of selecting advantageous information based on the user's emotions.
[0981] 5. Program: Software that allows the server to process information obtained from the terminal, select advantageous information, and display it on the terminal. This includes image recognition technology, location information analysis, emotion engine utilization, and other processes.
[0982] (Specific examples)
[0983] A specific example in which a user uses a headset type terminal 314 to visit a cafe will be shown.
[0984] 1. The user calls the headset terminal 314 and enters the cafe.
[0985] 2. The headset terminal 314 recognizes the cafe's logo or sign and transmits that information to the server.
[0986] 3.The server searches the database for information about the cafe based on the received information and obtains campaign and discount information.
[0987] 4. The headset type terminal 314 simultaneously transmits the user's facial expressions and voice characteristics to the emotion engine.
[0988] 5. The emotion engine estimates the user's emotion and returns the result to the server.
[0989] 6. The server selects deals based on the user's emotions. For example, if the user has a happy expression, the server selects discount deals.
[0990] 7. The server transmits the selected deals to the headset terminal 314.
[0991] 8. The headset type terminal 314 displays the received discount information based on the user's emotions. For example, the headset type terminal 314 displays a message that matches the user's facial expression together with the discount information.
[0992] The above is a specific example of an embodiment of the present invention. In addition to the cooperation between the headset type terminal 314 and the server, the use of an emotion engine makes it possible to provide advantageous information that matches the user's emotions.
[0993] The process flow will be explained below.
[0994] Step 1: A user visits a cafe using a headset type terminal 314.
[0995] Step 2: The headset terminal 314 recognizes the logo or sign of the cafe and transmits that information to the server.
[0996] Step 3: The server searches the database for information about the cafe based on the received information.
[0997] Step 4: The headset type terminal 314 simultaneously transmits the user's facial expressions and voice characteristics to the emotion engine.
[0998] Step 5: The emotion engine recognizes the user's facial expressions and voice characteristics and estimates the user's emotions.
[0999] Step 6: The server selects deals based on the user's emotions received from the emotion engine.
[1000] Step 7: The server transmits the selected deals to the headset terminal 314.
[1001] Step 8: The headset type terminal 314 performs display based on the received deals and the user's emotions.
[1002] Example 2
[1003] Next, a description will be given of Example 2. In the following description, the data processing device 12 will be referred to as a "server" and the headset type terminal 314 will be referred to as a "terminal".
[1004] Conventional systems provide limited information about the stores that users visit, making it difficult to provide information based on the user's current emotions and individual characteristics. In addition, there is a lack of dynamic information provision to improve the user experience, which means that the functions of the user's smart device cannot be fully utilized.
[1005] The identification process by the identification processing unit 290 of the data processing device 12 in the second embodiment is realized by the following means. In this invention, the server includes a means for a terminal to acquire user's location information, a means for the terminal to acquire image data using a camera and recognize a store's logo or sign based on the image data, a means for the terminal to transmit image data and location information to the server, a means for the server to search a database to acquire store information, a means for analyzing the user's facial expression or voice characteristics using an emotion engine and estimating the user's emotion, a means for the server to select advantageous information based on the user's emotion, a means for the server to transmit the advantageous information to the terminal, and a means for the terminal to display the advantageous information to the user. This makes it possible to acquire information on the store visited by the user from multiple angles and provide advantageous information personalized according to the user's emotion and situation.
[1006] A "terminal" is an electronic device that can be carried by a user and has the functions of acquiring image data, recording location information, and communicating with a server.
[1007] "Location information" is data indicating the geographical location of the user's current location, and is obtained using the GPS function.
[1008] "Image data" is visual information captured using the device's camera, and includes, for example, images of store logos and signs.
[1009] A "server" is a computer system that centrally processes and stores data, and is responsible for receiving and processing data sent from terminals.
[1010] A "database" is a collection of information that stores store logos, signboard information, campaign information, discount information, etc., and is used by the server for search and access.
[1011] An "emotion engine" is software or hardware that analyzes the user's facial expressions and vocal characteristics and estimates the user's emotional state.
[1012] "Bargain information" is information about discounts, campaigns, and the like that is useful to the user, and is selected based on the user's emotions and the stores that the user visits.
[1013] "User facial or vocal features" refers to data such as facial expressions and tone of voice that are used to estimate the user's emotions.
[1014] "Means for transmitting" refers to a communication means for transmitting data from a terminal to a server, and includes, for example, Wi-Fi and mobile data.
[1015] "Display means" refers to a function for visually or audibly presenting the information received by the terminal to the user, and includes screen display and audio guidance.
[1016] The system for implementing the present invention is composed of a user, a terminal, a server, a database, an emotion engine, and a program that connects these. Below, we will specifically explain how these elements work together to realize the invention.
[1017] First, the core of the system is the terminal used by the user. This terminal is designed as a smart device (e.g., smartphone, smart glasses) and has the following functions: It acquires image data using its camera function and collects location information using its GPS function. It also has sensors that detect the user's facial expressions and voice characteristics.
[1018] Next, a server is required to process the data acquired by the terminal. The server performs the following processes.
[1019] 1. Receive image data and location information sent from the device.
[1020] 2. Access the database and search for and retrieve image data of store logos and signs, as well as campaign and discount information.
[1021] 3. Use an emotion engine to estimate the user's emotions from their facial expressions and voice characteristics.
[1022] 4. Select the best deal from the information in the database based on user sentiment.
[1023] 5. Send the selected deals back to your device.
[1024] The database stores detailed data such as store and restaurant logos, signboard information, campaign information, discount information, etc. This database is maintained so that the server can efficiently search and access it.
[1025] The emotion engine estimates emotions by analyzing the user's facial expressions and vocal characteristics. For example, a generative AI model is used to classify emotional states such as "happy," "sad," and "surprised." This model is capable of analyzing the user's real-time emotions with high accuracy.
[1026] As a concrete example, the process when a user visits a cafe is shown below.
[1027] 1. A user enters a cafe with their smart device.
[1028] 2. The device recognizes the cafe's logo or sign, and sends image data and location information to the server.
[1029] 3. The server searches the database for information about the cafe and retrieves campaign and discount information.
[1030] 4. At the same time, the device transmits the user's facial expressions and voice characteristics to the emotion engine.
[1031] 5. The emotion engine estimates the user's emotion and returns the result to the server.
[1032] 6. The server selects deals based on the user's emotions. For example, if the user has a happy expression, the server selects discount deals.
[1033] 7. The server sends the selected deals to the device.
[1034] 8. The device displays information based on the discount information received and the user's emotions. For example, a message that matches the user's facial expression is displayed along with discount information.
[1035] An example of a prompt sentence might be:
[1036] When a user enters a cafe, the smart device scans the cafe's logo or sign and sends the image and location information to the server. The server searches the corresponding cafe information from the database and analyzes the user's facial expressions and voice using an emotion engine. Finally, the server selects deals that match the user's emotions and displays them on the smart device.
[1037] As described above, the present invention provides a system that obtains information about stores visited by a user from multiple angles and provides the user with personalized advantageous information according to the user's emotions and circumstances.
[1038] The flow of the identification process in the second embodiment will be described with reference to FIG.
[1039] Specific flow of system program processing
[1040] Step 1:
[1041] A user enters a store.
[1042] This process starts when a user enters a store with a smart device. For example, assume that the user enters a cafe.
[1043] Step 2:
[1044] The device captures images and location information.
[1045] The device uses its camera to take a picture of the store's logo or sign, and acquires location information using its GPS function. This inputs the image of the cafe's logo and the coordinate data of the current location into the device. The output is the captured image data and location information data.
[1046] Step 3:
[1047] The terminal transmits the information to the server.
[1048] The terminal transmits the acquired image data and location information data to the server. The input is the image data and location information data stored in the terminal, and the data is transmitted by communicating it to the server. The output is the data that has been transferred to the server.
[1049] Step 4:
[1050] The server searches the database.
[1051] The server searches the database based on the received image data and location information. The input is the image data and location information sent to the server, and the output is the corresponding store information (e.g., store name, campaign information, discount information, etc.) retrieved from the database.
[1052] Step 5:
[1053] The server uses an emotion engine to estimate the user's emotion.
[1054] The server inputs the facial expression and voice feature data of the user sent from the terminal into the emotion engine. The emotion engine analyzes this data and estimates the user's emotion. The input is the facial expression and voice feature data of the user, and the output is the user's emotional state (for example, "happy," "sad," "surprised," etc.).
[1055] Step 6:
[1056] The server selects the deals.
[1057] The server selects the optimal deal from the information in the database based on the estimated user emotion. The input is the user's emotional state and store information retrieved from the database, and the output is the deal information (e.g., specific discounts and campaign information) to be provided to the user.
[1058] Step 7:
[1059] The server transmits information to the terminal.
[1060] The server sends the selected deals to the terminal. The input is the deals stored in the server, and the data is sent by communicating it to the terminal. The output is the data that has been transferred to the terminal.
[1061] Step 8:
[1062] The terminal displays the information.
[1063] The terminal displays the received discount information to the user. The input is the discount information sent from the server, and the output is the information that the user can confirm visually or audibly. For example, the message "All items are 10% off today!" is displayed on the terminal screen.
[1064] Through the above processing steps, the present invention makes it possible to obtain information about stores visited by a user from multiple angles and provide personalized deals information according to the user's emotions and circumstances.
[1065] (Application example 2)
[1066] Next, a description will be given of Application Example 2. In the following description, the data processing device 12 will be referred to as a "server", and the headset type terminal 314 will be referred to as a "terminal".
[1067] Conventional store information systems can only provide uniform deals to users, and it is difficult to provide information tailored to the emotions and circumstances of individual users. In addition, since the system does not take into account the emotions of users, the information provided may not always be appropriate, making it difficult to improve user satisfaction.
[1068] The identification process by the identification processing unit 290 of the data processing device 12 in the application example 2 is realized by the following means. In this invention, the server includes a means for recognizing and acquiring store information from image data, a means for acquiring the user's facial expression data together with the image data, a means for selecting advantageous information based on the acquired store information, a means for estimating an emotion based on the user's facial expression data and optimizing the advantageous information based on the emotion, and a means for displaying the optimized advantageous information on the smart glasses. This makes it possible to provide personalized information based on the user's emotions, and is expected to improve user satisfaction.
[1069] The "user-worn device" is a device worn by a user visiting a store, and has the function of acquiring image data and facial expression data.
[1070] The "store information acquisition means" is a function for recognizing store logos and signs from image data obtained from a user-worn device and acquiring that information.
[1071] The "facial expression data acquisition means" is a function for capturing the user's facial expression and acquiring that data.
[1072] The "information selection means" is a function for selecting advantageous information to be presented to the user based on the acquired store information.
[1073] The "emotion estimation means" is a function for estimating the user's emotion based on the acquired facial expression data of the user.
[1074] The "information optimization means" is a function for optimizing selected advantageous information based on the estimated user's emotions.
[1075] The "information display means" is a function for displaying optimized advantageous information on a user-worn device.
[1076] DETAILED DESCRIPTION OF THE PREFERRED EMBODIMENTS
[1077] The system of the present invention consists of smart glasses worn by the user, a server that processes images and location information, and an emotion engine.
[1078] 1. Hardware and Software Configuration
[1079] Smart Glasses
[1080] The smart glasses, worn by the user, have a camera and a display function. The camera captures store logos and signs, and also obtains the user's facial expression data. These data are then sent to a server.
[1081] server
[1082] The server uses the following hardware and software:
[1083] Hardware: High performance processor, storage device
[1084] Software: Python, OpenCV, requests library
[1085] The server analyzes the image data acquired from the smart glasses to recognize store information, and also uses an emotion engine to estimate the user's emotions from their facial expression data.
[1086] Emotion Engine
[1087] The emotion engine is software that analyzes the user's facial expression data and estimates their emotions. The basic emotion estimation algorithm uses a machine learning model.
[1088] 2. Data flow and processing
[1089] The image data acquired by the smart glasses is sent to a server. The server recognizes the store's logo or signboard from the image data and retrieves store information from a database. Next, the server sends the user's facial expression data acquired at the same time to an emotion engine to estimate the user's emotion.
[1090] Store information acquisition method
[1091] The server recognizes logos and signs from the image data and retrieves related store information from the database based on this. Specifically, the server uses the OpenCV library for image recognition.
[1092] Emotion estimation means
[1093] The user's facial expression data is sent to the emotion engine, which uses machine learning to classify the user's emotions into categories such as "happy," "sad," "excited," and "calm."
[1094] Information optimization measures
[1095] The server optimizes the deals based on the user's emotions. For example, if the user has a happy expression, the server will provide discount information.
[1096] Means of selecting and displaying information
[1097] The server sends the optimized information to the smart glasses to display to the user, including messages personalized to the user's emotions.
[1098] 3. Specific Examples
[1099] As a concrete example, consider the case where a user visits a cafe wearing smart glasses.
[1100] 1. When a user enters a cafe, the camera in the smart glasses recognizes the cafe's logo and sign and transmits it to the server.
[1101] 2. The server retrieves information about recognized cafes from the database.
[1102] 3. At the same time, the smart glasses capture the user's facial expression, and if the emotion engine estimates it to be "happy," the server will offer a "Special Happy Hour Discount."
[1103] 4. This information is then displayed on the smart glasses, allowing users to receive real-time deals.
[1104] Example prompts to be input to the generative AI model
[1105] When a smart glasses user enters a store, the store's logo and sign are recognized and the information is sent to the server. Create a program that then analyzes the user's emotions and displays discount and campaign information based on the emotions on the display. Show how to configure it in Python.
[1106] The flow of the specific process in the application example 2 will be described with reference to FIG.
[1107] Step 1:
[1108] Smart glasses capture user actions.
[1109] Input: Image data including the in-store environment and the user's facial expressions captured using the smart glasses camera.
[1110] Output: Image files containing store logos, signage, and user facial expressions.
[1111] How it works: The camera in the smart glasses captures images in real time and stores them in local storage.
[1112] Step 2:
[1113] The terminal transmits the captured image data to the server.
[1114] Input: Image data captured by smart glasses.
[1115] Output: Image data sent to the server.
[1116] Specific operation: The terminal communicates with the server via the Internet using HTTP, and sends image data as a POST request.
[1117] Step 3:
[1118] The server analyzes the image data it receives and recognizes store logos and signs.
[1119] Input: Image data including store logos and signs sent from the device.
[1120] Output: IDs of recognized stores and related information.
[1121] Specific operation: The server processes image data using the OpenCV library and detects and recognizes store logos and signs.
[1122] Step 4:
[1123] The store information recognized by the server is obtained from the database.
[1124] Input: The ID of a recognized store.
[1125] Output: Store details (e.g. store name, campaign information, discount information, etc.).
[1126] Specific operation: The server executes a database query using the store ID as a key to obtain related store information.
[1127] Step 5:
[1128] The terminal captures the user's facial expression data and transmits it to the server.
[1129] Input: Image data containing the user's facial expressions captured by the smart glasses.
[1130] Output: Facial expression data sent to the server.
[1131] Specific operation: The terminal communicates with the server via the Internet via HTTP and sends facial expression data as a POST request.
[1132] Step 6:
[1133] The server analyzes the facial expression data and uses an emotion engine to infer the user's emotions.
[1134] Input: User's facial expression data sent from the device.
[1135] Output: An estimated user emotion classification (e.g. happy, sad, excited, etc.).
[1136] Specific operation: The server uses an emotion engine (machine learning model) to analyze facial expression data and classify emotions.
[1137] Step 7:
[1138] The server selects the best deals based on the user's sentiment.
[1139] Input: Estimated user sentiment and store information.
[1140] Output: Optimized deals based on sentiment.
[1141] Specific operation: Based on the store information, the server selects the deals (discount information and campaign information) that best suit the user's emotions.
[1142] Step 8:
[1143] The server sends the selected information to the terminal and displays it on the smart glasses.
[1144] Enter: Optimized Deals.
[1145] Output: Deals displayed on the smart glasses.
[1146] Specific operation: The server sends the best deals to the device and displays the information on the smart glasses display.
[1147] This allows users to receive real-time, personalized deals when they visit a store.
[1148] The specific processing unit 290 transmits the result of the specific processing to the headset type terminal 314. In the headset type terminal 314, the control unit 46A causes the speaker 240 and the display 343 to output the result of the specific processing. The microphone 238 acquires voice indicating a user input for the result of the specific processing. The control unit 46A transmits voice data indicating the user input acquired by the microphone 238 to the data processing device 12. In the data processing device 12, the specific processing unit 290 acquires the voice data.
[1149] The data generation model 58 is a so-called generative AI (Artificial Intelligence). An example of the data generation model 58 is ChatGPT (Internet search<URL: https: / / openai.com / blog / chatgpt> ), Gemini (Internet search <url: https: gemini.google.com ?hl="ja">) and other generation AIs. The data generation model 58 is obtained by making a neural network perform deep learning. A prompt including an instruction is input to the data generation model 58, and inference data such as voice data indicating a voice, text data indicating a text, and image data indicating an image is input. The data generation model 58 performs inference on the input inference data according to the instruction indicated by the prompt, and outputs the inference result in a data format such as voice data and text data. Here, inference refers to, for example, analysis, classification, prediction, and / or summarization.
[1150] In the above embodiment, an example was given in which the specific processing is performed by the data processing device 12, but the technology of the present disclosure is not limited to this, and the specific processing may be performed by the headset type terminal 314.
[1151] [Fourth embodiment]
[1152] FIG. 7 shows an example of the configuration of a data processing system 410 according to the fourth embodiment.
[1153] 7, a data processing system 410 includes a data processing device 12 and a robot 414. An example of the data processing device 12 is a server.
[1154] The data processing device 12 includes a computer 22, a database 24, and a communication I / F 26. The computer 22 is an example of a "computer" according to the technology of the present disclosure. The computer 22 includes a processor 28, a RAM 30, and a storage 32. The processor 28, the RAM 30, and the storage 32 are connected to a bus 34. The database 24 and the communication I / F 26 are also connected to the bus 34. The communication I / F 26 is connected to a network 54. Examples of the network 54 include a wide area network (WAN) and / or a local area network (LAN).
[1155] The robot 414 includes a computer 36, a microphone 238, a speaker 240, a camera 42, a communication I / F 44, and a control target 443. The computer 36 includes a processor 46, a RAM 48, and a storage 50. The processor 46, the RAM 48, and the storage 50 are connected to a bus 52. In addition, the microphone 238, the speaker 240, the camera 42, and the control target 443 are also connected to the bus 52.
[1156] The microphone 238 receives instructions and the like from the user 20 by receiving voice uttered by the user 20. The microphone 238 captures the voice uttered by the user 20, converts the captured voice into audio data, and outputs it to the processor 46. The speaker 240 outputs the voice according to instructions from the processor 46.
[1157] Camera 42 is a small digital camera equipped with an optical system including a lens, an aperture, and a shutter, and an imaging element such as a CMOS (Complementary Metal-Oxide-Semiconductor) image sensor or a CCD (Charge Coupled Device) image sensor, and captures images of the surroundings of user 20 (e.g., an imaging range defined by an angle of view equivalent to the width of the field of vision of an average healthy person).
[1158] The communication I / F 44 is connected to a network 54. The communication I / Fs 44 and 26 are responsible for transmitting and receiving various types of information between the processor 46 and the processor 28 via the network 54. The transmission and reception of various types of information between the processor 46 and the processor 28 using the communication I / Fs 44 and 26 is performed in a secure state.
[1159] The control target 443 includes a display device, LEDs in the eyes, and motors for driving the arms, hands, legs, etc. The posture and behavior of the robot 414 are controlled by controlling the motors of the arms, hands, legs, etc. Some of the emotions of the robot 414 can be expressed by controlling these motors. In addition, the facial expressions of the robot 414 can also be expressed by controlling the light emission state of the LEDs in the eyes of the robot 414.
[1160] Fig. 8 shows an example of main functions of the data processing device 12 and the robot 414. As shown in Fig. 8, in the data processing device 12, a specific process is performed by the processor 28. A specific process program 56 is stored in the storage 32.
[1161] The specific processing program 56 is an example of a "program" according to the technology of the present disclosure. The processor 28 reads the specific processing program 56 from the storage 32, and executes the read specific processing program 56 on the RAM 30. The specific processing is realized by the processor 28 operating as a specific processing unit 290 in accordance with the specific processing program 56 executed on the RAM 30.
[1162] The storage 32 stores a data generation model 58 and an emotion identification model 59. The data generation model 58 and the emotion identification model 59 are used by the identification processing unit 290.
[1163] In the robot 414, the reception output process is performed by the processor 46. A reception output program 60 is stored in the storage 50. The processor 46 reads the reception output program 60 from the storage 50, and executes the read reception output program 60 on the RAM 48. The reception output process is realized by the processor 46 operating as the control unit 46A in accordance with the reception output program 60 executed on the RAM 48.
[1164] Next, a description will be given of the specific processing by the specific processing unit 290 of the data processing device 12. In the following description, the data processing device 12 will be referred to as a "server" and the robot 414 will be referred to as a "terminal".
[1165] The embodiment for carrying out the present invention comprises the following elements.
[1166] 1. Terminal: A device used when a user visits a store or restaurant to acquire images and location information and display special offers.
[1167] 2. Server: A central processing unit that receives images and location information from the terminals, obtains information about stores and restaurants, and processes the information to select deals.
[1168] 3. Database: This is where the server stores data to search for store and restaurant information. This includes store logos, signboard information, campaign information, discount information, etc.
[1169] 4. Program: This is the software that allows the server to process the information obtained from the terminal, select the deals, and display them on the smart glasses. This includes image recognition technology, location information analysis, and consideration of user attributes and past behavioral history.
[1170] (Specific examples)
[1171] A specific example will be given of a case where a user carries a robot 414 and visits a cafe.
[1172] 1. A user carries a robot 414 and enters a cafe.
[1173] 2. The robot 414 recognizes the cafe's logo or sign and sends that information to the server.
[1174] 3.The server searches the database for information about the cafe based on the received information and obtains campaign and discount information.
[1175] 4. The server selects deals that are appropriate for the user, taking into account the user's attributes and past behavioral history.
[1176] 5. The server sends the selected deals to the robot 414.
[1177] 6. The robot 414 displays the received deals and allows the user to take advantage of promotions and discounts.
[1178] The above is a specific example of an embodiment of the present invention. By linking the robot 414 with the server, the user can easily obtain and use information on deals at stores and restaurants.
[1179] The process flow will be explained below.
[1180] Step 1: The server receives images and location information obtained from the robot 414. Specifically, the robot 414 recognizes the logo and sign of the cafe and transmits that information to the server.
[1181] Step 2: The server analyzes the received image and location information. Using image recognition technology, it identifies the cafe's logo and sign and obtains the location information.
[1182] Step 3: The server searches the database for information about the cafe. The database contains information about the cafe's campaigns and discounts, and the server retrieves this information.
[1183] Step 4: The server selects the appropriate discount information based on the user's attributes and past behavior history. For example, if the user has visited the same cafe in the past and received a discount at that time, the server selects the discount information based on that information.
[1184] Step 5: The server transmits the selected advantageous information to the robot 414. Specifically, the server generates data for displaying the advantageous information and transmits it to the robot 414.
[1185] Step 6: The robot 414 displays the received deals. The user can check the campaigns and discount information at the cafe through the screen of the robot 414.
[1186] Example 1
[1187] Next, a description will be given of Example 1. In the following description, the data processing device 12 is referred to as a "server" and the robot 414 is referred to as a "terminal."
[1188] Conventionally, there has been no system that allows users to easily and quickly obtain and use information on deals on the spot when they visit a store. In addition, there is a demand for a method to realize more effective marketing by providing personalized information that takes into account the user's attributes and past behavioral history.
[1189] The specific processing by the specific processing unit 290 of the data processing device 12 in the first embodiment is realized by the following means.
[1190] In this invention, the server includes a means for acquiring store information based on image data and location information acquired from the mobile information terminal when the user visits the store using the mobile information terminal, a means for selecting advantageous information from the acquired store information, and a means for displaying the selected advantageous information on the mobile information terminal, thereby enabling the user to easily acquire and use personalized advantageous information in real time.
[1191] A "portable information terminal" is an electronic device that a user can carry with them, and is a terminal that has the function of taking pictures and acquiring location information.
[1192] A "server" is a central processing unit that receives data from multiple portable information terminals via a network, analyzes the data, and provides related information.
[1193] "Store information" is data related to a specific store, and includes image data of logos and signs, location information, campaign information, discount information, and the like.
[1194] "Bargain information" is promotional information that is beneficial to the user, such as special offers, discounts, and sales information for stores visited by the user.
[1195] "Image data" refers to digital data of an image captured by a user's mobile information terminal, and is the subject of analysis.
[1196] "Location information" refers to geographic coordinate data obtained through the GPS function of a mobile information terminal, for example.
[1197] "User attributes" refers to information relating to a user, such as age, sex, preferences, etc., that is used to provide personalized information.
[1198] "Past behavioral history" refers to a record of stores the user previously visited, products the user purchased, campaigns the user participated in, and the like, and indicates the user's behavioral patterns.
[1199] "Analysis" is the process in which the server extracts and identifies the necessary information based on the image data and location information it receives.
[1200] The present invention relates to a system for providing advantageous information to a user when the user visits a store using a mobile information terminal. Specific embodiments of the system will be described below.
[1201] Hardware and software configuration
[1202] First, the hardware and software required to realize this system will be described.
[1203] 1. Mobile information terminals:
[1204] A mobile information terminal is an electronic device that can be carried by a user and has the functions of taking pictures and acquiring location information. Specifically, a smartphone or a tablet is one such device.
[1205] It has a camera function and can acquire image data.
[1206] It has a GPS function and can obtain location information.
[1207] 2. Server:
[1208] A server is a central processing unit that receives data from multiple portable information terminals via a network, analyzes the data, and provides related information.
[1209] The servers are typically located in the cloud and have the necessary processing power and storage.
[1210] Image recognition libraries such as "OpenCV" are used for image analysis.
[1211] "Google Maps API" is used for location analysis.
[1212] Machine learning libraries such as "TensorFlow" are used to analyze user attributes and past behavioral history.
[1213] 3. Database:
[1214] The database stores store information (logos, signs, campaign information, discount information, etc.).
[1215] Program Processing
[1216] The server executes a program for selecting and providing advantageous information suitable for the user based on image data and location information obtained from the mobile information terminal.
[1217] 1. Data Acquisition:
[1218] Image data of a store's logo or signboard photographed by the terminal's camera is acquired.
[1219] The current location information is obtained using the device's GPS function.
[1220] 2. Data transmission:
[1221] The device transmits the acquired image data and location information to the server using Wi-Fi or 4G / 5G networks.
[1222] 3. Image Analysis:
[1223] The server uses OpenCV to analyze the image data and identify store logos and signs, extract feature points within the image, and match them with existing images in a database.
[1224] 4. Information Search:
[1225] Based on the analysis results, the server retrieves detailed information about the relevant store from the database, such as the cafe's menu, opening hours, and location.
[1226] 5. Obtaining Campaign Information:
[1227] The server obtains store campaign and discount information from the database. For example, it obtains information about "new menu item launch commemorative discount."
[1228] 6. Choose the best deals:
[1229] The server selects the most suitable campaign information for a user based on the user's attribute information (e.g., age, gender, preferences) and past behavioral history. To do this, a machine learning algorithm using "TensorFlow" is used.
[1230] 7. Information display:
[1231] The server transmits the selected deals to the terminal, which displays the received deals on its screen, allowing the user to use the information to receive campaigns and discounts.
[1232] Examples
[1233] For example, a specific example in which a user uses a mobile information terminal to visit a cafe is as follows.
[1234] 1. User enters the cafe:
[1235] A user carries a mobile information terminal and enters a cafe.
[1236] 2. The device captures a logo or sign:
[1237] Using the device's camera, the user takes a picture of the cafe's logo or sign.
[1238] 3. The device sends the image and location information to the server:
[1239] The device sends the captured images and location information obtained using the GPS function to the server.
[1240] 4. The server analyzes the image:
[1241] The server analyzes the received image using OpenCV and retrieves information about the corresponding cafe from a database.
[1242] 5. The server gets the campaign information:
[1243] The server retrieves the cafe's latest promotions and discounts from the same database.
[1244] 6. Server selects deals:
[1245] The server selects the most suitable deals by taking into account the user's attributes and past behavioral history.
[1246] 7. The server sends the deals to the device:
[1247] The server sends the selected deals to the terminal.
[1248] 8. Your device will display special offers:
[1249] The device will display the received deals on the screen, for example, "This cafe is now offering 50% off our new coffee menu!"
[1250] 9. User takes advantage of promotions and discounts:
[1251] Users can use the deals displayed on the screen to take advantage of coupons and discounts at the cafe.
[1252] Examples of prompt statements
[1253] "When a user visits a cafe, he or she takes a photo of the cafe's logo with a mobile information terminal and sends it to the server. The server then retrieves information about the cafe from the database, selects the most suitable deals based on the user's attributes and past behavioral history, and sends them to the mobile information terminal. The mobile information terminal displays the received information, and the user takes advantage of the cafe's campaign. Please explain the specific process."
[1254] The above is an embodiment of the present invention. This system allows users to easily obtain and use information about deals at stores.
[1255] The flow of the identification process in the first embodiment will be described with reference to FIG.
[1256] Step 1:
[1257] User enters the cafe
[1258] A user carries a mobile information terminal and enters a cafe.
[1259] Input: Cafe location, user attribute information
[1260] Output: The user is in the store
[1261] Step 2:
[1262] The device photographs logos and signs
[1263] Using the device's camera, the user takes a picture of the cafe's logo or sign.
[1264] Specific actions: The user activates the device's camera, frames the cafe's logo or sign, and presses the capture button.
[1265] Input: Image data of the cafe logo or sign
[1266] Output: Captured image data
[1267] Step 3:
[1268] The device sends the image and location information to the server.
[1269] The terminal transmits the acquired image data and the location information acquired by the GPS function to the server.
[1270] Specific operation: The device uses Wi-Fi or 4G / 5G networks to pack image data and location information into packets and upload them to the server.
[1271] Input: Image data, location information
[1272] Output: Image data and location information sent to the server
[1273] Step 4:
[1274] The server analyzes the image
[1275] The server analyzes the received image data using OpenCV and identifies the cafe's logo and sign.
[1276] Specific operation: Extract feature points in the image and match them with images in an existing database to identify the corresponding store.
[1277] Input: Image data
[1278] Output: Identified store information
[1279] Step 5:
[1280] The server retrieves cafe information from the database.
[1281] The server retrieves detailed information about the corresponding cafe from the database based on the identified store information.
[1282] Specific operation: The server sends a query to the database to obtain information about the relevant cafe, such as its menu, opening hours, and location.
[1283] Input: Identified store information
[1284] Output: Detailed information about the cafe
[1285] Step 6:
[1286] The server retrieves the campaign information.
[1287] The server retrieves the cafe's latest promotions and discounts from the same database.
[1288] Specific operation: The server sends a query to the database to obtain the relevant campaign information.
[1289] Input: Cafe details
[1290] Output: Retrieved campaign information
[1291] Step 7:
[1292] Server selects deals
[1293] The server selects the most suitable deals based on the user's attribute information (age, gender, preferences) and past behavioral history.
[1294] Specific operation: The server inputs the user's attribute data and past behavioral history into a machine learning model and calculates optimized campaign information.
[1295] Input: User attribute information, past behavior history, campaign information
[1296] Output: Selected deals
[1297] Step 8:
[1298] The server sends useful information to the terminal
[1299] The server sends the selected deals to the terminal.
[1300] How it works: The server uses Wi-Fi or 4G / 5G networks to pack deals into packets and upload them to the device.
[1301] Input: Selected Deals
[1302] Output: Deals sent to the terminal
[1303] Step 9:
[1304] The device displays useful information
[1305] The device will display the received deals on the screen.
[1306] Specific operation: A pop-up notification or a dedicated application screen is displayed on the device screen to provide the user with useful information.
[1307] Input: Deals sent to your device
[1308] Output: Deals displayed to the user
[1309] Step 10:
[1310] Users take advantage of promotions and discounts
[1311] Users can use the deals displayed on their devices to take advantage of coupons and discounts at cafes.
[1312] Specific operation: The user shows the coupon screen displayed on the terminal to the cafe cashier to receive the discount.
[1313] Input: Deals displayed to the user
[1314] Output: The user can use the campaign or discount.
[1315] (Application example 1)
[1316] Next, a description will be given of Application Example 1. In the following description, the data processing device 12 will be referred to as a "server" and the robot 414 will be referred to as a "terminal."
[1317] In the conventional user experience in brick-and-mortar stores, users had to search for store and campaign information themselves, which was not very efficient. In particular, there were limited ways for users to find out what services and discounts were available in stores in real time, so building an information provision system that was effective for users was a challenge.
[1318] The specific processing by the specific processing unit 290 of the data processing device 12 in the application example 1 is realized by the following means.
[1319] In this invention, the server includes a means for acquiring image data and location information, a means for transmitting the acquired image data and location information to the server, a means for the server to search for store information from a database based on the image data and location information and provide the user with selected advantageous information, a means for displaying the advantageous information on the smart glasses, a means for recognizing store logos and signs from the image data, and a means for selecting optimal information in consideration of the user's past behavior history and attributes. This allows the user to acquire and use store information and campaign information in real time.
[1320] "Image data" refers to data that represents visual information acquired by a user using a smart device.
[1321] "Location information" is geographical data such as the latitude and longitude of the user's current location.
[1322] The "server" is a central processing unit that receives image data and location information and searches for store information from a database.
[1323] A "database" is a storage device for storing store information, campaign information, and discount information.
[1324] "Store information" is data including store logos and signs, campaign information, discount information, and the like.
[1325] "Good deals" refers to information about campaigns and discounts that are beneficial to users.
[1326] "Smart glasses" are wearable devices that users wear to display store information and special offers.
[1327] "Image recognition" is a technology that analyzes acquired image data and identifies specific objects or text.
[1328] "Past behavior history" is a record of actions and choices made by a user in the past.
[1329] "User attributes" refers to personal information about the user, such as age, sex, and interests.
[1330] A "generative AI model" is a model generated by artificial intelligence, and is a technology used for data analysis, prediction, and recommendations.
[1331] An embodiment of the present invention will be described below. A system for carrying out the present invention is mainly composed of a terminal for acquiring image data and location information, a server for processing data and providing information, and software for linking these components.
[1332] First, a user visits a physical store using a smart device (such as a smartphone or smart glasses). Image data and location information are acquired using the device's camera and GPS functions. Specific hardware used for this purpose include an on-board webcam, an external camera (e.g., Logitech C920), a built-in GPS, or an external GPS module (e.g., Holux M-1000C).
[1333] The acquired image data is sent from the terminal to a server. The server receives the image data and location information, and searches a database for store information. Google Cloud Vision API and Amazon Rekognition are used for image recognition, and Google Maps API is used for location analysis. The database uses cloud storage such as Amazon RDS.
[1334] The server selects deals suitable for the user based on the acquired store information and the user's past behavioral history and attributes. At this time, a generative AI model (e.g. TensorFlow) is used to analyze user attributes and past behavioral data and recommend the most suitable information.
[1335] The selected deals are then pushed to the user's smart device and displayed to the user, allowing the user to receive real-time deals and improve their shopping experience in-store.
[1336] For example, when a user enters a bookstore, the smart device recognizes the bookstore's logo and location information. The server provides discount information on new books based on the bookstore's information and past purchase history, and displays information about autograph sessions for members only.
[1337] Examples of specific prompts include the following:
[1338] When a user enters a bookstore using their smartphone, we want to display real-time information about discounts on new books and events for members only. Please tell us how to provide the appropriate information using image recognition and GPS information.
[1339] In this way, convenience for users in the store can be greatly improved.
[1340] The flow of the specific process in the application example 1 will be described with reference to FIG.
[1341] Step 1:
[1342] A user visits a store using a smart device. The device's camera takes a picture of the store's logo or sign, and the location information is obtained using the GPS function. The input is camera image data and location information, and the output is data to be sent to the server. The device prepares to send the image data captured by the camera and GPS coordinates to the server.
[1343] Step 2:
[1344] The terminal transmits the acquired image data and location information to the server. The input is the captured image and the acquired location information, and the output is data sent to the server. The terminal encodes the image data and transmits it to the server together with GPS information. At this time, the data is uploaded using an HTTP request.
[1345] Step 3:
[1346] The server processes the received image data and location information. The input is the image data and location information sent from the terminal, and the output is store information retrieved from a database. The server uses image recognition software (e.g. Google Cloud Vision API or Amazon Rekognition) to recognize store logos and signs from the image data. At the same time, it searches the database based on the location information and retrieves the corresponding store information.
[1347] Step 4:
[1348] The server selects deals based on the acquired store information and the user's past behavioral history and attributes. The input is store information, user behavioral history, and user attributes, and the output is the selected deals. The server analyzes the data using a generative AI model (e.g. TensorFlow) and selects the most suitable campaign and discount information.
[1349] Step 5:
[1350] The server sends the selected deals to the terminal. The input is the selected deals, and the output is data transmission to the terminal. The server uses an HTTP request to push the selected information to the user's smart device.
[1351] Step 6:
[1352] The terminal displays the received discount information to the user. The input is the discount information sent from the server, and the output is the information displayed to the user. The terminal uses the push notification function to display the discount information on the display of smart glasses or a smartphone. The user can check campaign information and discounts in real time.
[1353] In addition, an emotion engine that estimates the emotion of the user may be further combined. That is, the identification processing unit 290 may estimate the emotion of the user using the emotion identification model 59, and perform identification processing using the emotion of the user.
[1354] The embodiment for carrying out the present invention comprises the following elements.
[1355] 1. Terminal: A device used when a user visits a store or restaurant to acquire images and location information and display special offers.
[1356] 2. Server: A central processing unit that receives images and location information from the terminals, obtains information about stores and restaurants, and processes the information to select deals.
[1357] 3. Database: This is where the server stores data to search for store and restaurant information. This includes store logos, signboard information, campaign information, discount information, etc.
[1358] 4. Emotion engine: This engine has the function of recognizing the user's facial expressions and voice characteristics and estimating the user's emotions. The emotion engine is used in cooperation with the server as a means of selecting advantageous information based on the user's emotions.
[1359] 5. Program: Software that allows the server to process information obtained from the terminal, select advantageous information, and display it on the terminal. This includes image recognition technology, location information analysis, emotion engine utilization, and other processes.
[1360] (Specific examples)
[1361] A specific example will be given of a case where a user carries a robot 414 and visits a cafe.
[1362] 1. A user carries a robot 414 and enters a cafe.
[1363] 2. The robot 414 recognizes the cafe's logo or sign and sends that information to the server.
[1364] 3.The server searches the database for information about the cafe based on the received information and obtains campaign and discount information.
[1365] 4. The robot 414 simultaneously transmits the user's facial expressions and voice characteristics to the emotion engine.
[1366] 5. The emotion engine estimates the user's emotion and returns the result to the server.
[1367] 6. The server selects deals based on the user's emotions. For example, if the user has a happy expression, the server selects discount deals.
[1368] 7. The server sends the selected deals to the robot 414.
[1369] 8. The robot 414 displays information based on the received discount information and the user's emotions. For example, the robot 414 displays a message that matches the user's facial expression along with discount information.
[1370] The above is a specific example of the embodiment for carrying out the present invention. In addition to the cooperation between the robot 414 and the server, the use of the emotion engine makes it possible to provide advantageous information that matches the user's emotions.
[1371] The process flow will be explained below.
[1372] Step 1: A user carries a robot 414 and visits a cafe.
[1373] Step 2: The robot 414 recognizes the cafe's logo or sign and sends that information to the server.
[1374] Step 3: The server searches the database for information about the cafe based on the received information.
[1375] Step 4: The robot 414 simultaneously transmits the user's facial expressions and voice characteristics to the emotion engine.
[1376] Step 5: The emotion engine recognizes the user's facial expressions and voice characteristics and estimates the user's emotions.
[1377] Step 6: The server selects deals based on the user's emotions received from the emotion engine.
[1378] Step 7: The server sends the selected deals to the robot 414.
[1379] Step 8: The robot 414 makes a display based on the received deals and the user's emotions.
[1380] Example 2
[1381] Next, a description will be given of Example 2. In the following description, the data processing device 12 is referred to as a "server" and the robot 414 is referred to as a "terminal."
[1382] Conventional systems provide limited information about the stores that users visit, making it difficult to provide information based on the user's current emotions and individual characteristics. In addition, there is a lack of dynamic information provision to improve the user experience, which means that the functions of the user's smart device cannot be fully utilized.
[1383] The identification process by the identification processing unit 290 of the data processing device 12 in the second embodiment is realized by the following means. In this invention, the server includes a means for a terminal to acquire user's location information, a means for the terminal to acquire image data using a camera and recognize a store's logo or sign based on the image data, a means for the terminal to transmit image data and location information to the server, a means for the server to search a database to acquire store information, a means for analyzing the user's facial expression or voice characteristics using an emotion engine and estimating the user's emotion, a means for the server to select advantageous information based on the user's emotion, a means for the server to transmit the advantageous information to the terminal, and a means for the terminal to display the advantageous information to the user. This makes it possible to acquire information on the store visited by the user from multiple angles and provide advantageous information personalized according to the user's emotion and situation.
[1384] A "terminal" is an electronic device that can be carried by a user and has the functions of acquiring image data, recording location information, and communicating with a server.
[1385] "Location information" is data indicating the geographical location of the user's current location, and is obtained using the GPS function.
[1386] "Image data" is visual information captured using the device's camera, and includes, for example, images of store logos and signs.
[1387] A "server" is a computer system that centrally processes and stores data, and is responsible for receiving and processing data sent from terminals.
[1388] A "database" is a collection of information that stores store logos, signboard information, campaign information, discount information, etc., and is used by the server for search and access.
[1389] An "emotion engine" is software or hardware that analyzes the user's facial expressions and vocal characteristics and estimates the user's emotional state.
[1390] "Bargain information" is information about discounts, campaigns, and the like that is useful to the user, and is selected based on the user's emotions and the stores that the user visits.
[1391] "User facial or vocal features" refers to data such as facial expressions and tone of voice that are used to estimate the user's emotions.
[1392] "Means for transmitting" refers to a communication means for transmitting data from a terminal to a server, and includes, for example, Wi-Fi and mobile data.
[1393] "Display means" refers to a function for visually or audibly presenting the information received by the terminal to the user, and includes screen display and audio guidance.
[1394] The system for implementing the present invention is composed of a user, a terminal, a server, a database, an emotion engine, and a program that connects these. Below, we will specifically explain how these elements work together to realize the invention.
[1395] First, the core of the system is the terminal used by the user. This terminal is designed as a smart device (e.g., smartphone, smart glasses) and has the following functions: It acquires image data using its camera function and collects location information using its GPS function. It also has sensors that detect the user's facial expressions and voice characteristics.
[1396] Next, a server is required to process the data acquired by the terminal. The server performs the following processes.
[1397] 1. Receive image data and location information sent from the device.
[1398] 2. Access the database and search for and retrieve image data of store logos and signs, as well as campaign and discount information.
[1399] 3. Use an emotion engine to estimate the user's emotions from their facial expressions and voice characteristics.
[1400] 4. Select the best deal from the information in the database based on user sentiment.
[1401] 5. Send the selected deals back to your device.
[1402] The database stores detailed data such as store and restaurant logos, signboard information, campaign information, discount information, etc. This database is maintained so that the server can efficiently search and access it.
[1403] The emotion engine estimates emotions by analyzing the user's facial expressions and vocal characteristics. For example, a generative AI model is used to classify emotional states such as "happy," "sad," and "surprised." This model is capable of analyzing the user's real-time emotions with high accuracy.
[1404] As a concrete example, the process when a user visits a cafe is shown below.
[1405] 1. A user enters a cafe with their smart device.
[1406] 2. The device recognizes the cafe's logo or sign, and sends image data and location information to the server.
[1407] 3. The server searches the database for information about the cafe and retrieves campaign and discount information.
[1408] 4. At the same time, the device transmits the user's facial expressions and voice characteristics to the emotion engine.
[1409] 5. The emotion engine estimates the user's emotion and returns the result to the server.
[1410] 6. The server selects deals based on the user's emotions. For example, if the user has a happy expression, the server selects discount deals.
[1411] 7. The server sends the selected deals to the device.
[1412] 8. The device displays information based on the discount information received and the user's emotions. For example, a message that matches the user's facial expression is displayed along with discount information.
[1413] An example of a prompt sentence might be:
[1414] When a user enters a cafe, the smart device scans the cafe's logo or sign and sends the image and location information to the server. The server searches the corresponding cafe information from the database and analyzes the user's facial expressions and voice using an emotion engine. Finally, the server selects deals that match the user's emotions and displays them on the smart device.
[1415] As described above, the present invention provides a system that obtains information about stores visited by a user from multiple angles and provides the user with personalized advantageous information according to the user's emotions and circumstances.
[1416] The flow of the identification process in the second embodiment will be described with reference to FIG.
[1417] Specific flow of system program processing
[1418] Step 1:
[1419] A user enters a store.
[1420] This process starts when a user enters a store with a smart device. For example, assume that the user enters a cafe.
[1421] Step 2:
[1422] The device captures images and location information.
[1423] The device uses its camera to take a picture of the store's logo or sign, and acquires location information using its GPS function. This inputs the image of the cafe's logo and the coordinate data of the current location into the device. The output is the captured image data and location information data.
[1424] Step 3:
[1425] The terminal transmits the information to the server.
[1426] The terminal transmits the acquired image data and location information data to the server. The input is the image data and location information data stored in the terminal, and the data is transmitted by communicating it to the server. The output is the data that has been transferred to the server.
[1427] Step 4:
[1428] The server searches the database.
[1429] The server searches the database based on the received image data and location information. The input is the image data and location information sent to the server, and the output is the corresponding store information (e.g., store name, campaign information, discount information, etc.) retrieved from the database.
[1430] Step 5:
[1431] The server uses an emotion engine to estimate the user's emotion.
[1432] The server inputs the facial expression and voice feature data of the user sent from the terminal into the emotion engine. The emotion engine analyzes this data and estimates the user's emotion. The input is the facial expression and voice feature data of the user, and the output is the user's emotional state (for example, "happy," "sad," "surprised," etc.).
[1433] Step 6:
[1434] The server selects the deals.
[1435] The server selects the optimal deal from the information in the database based on the estimated user emotion. The input is the user's emotional state and store information retrieved from the database, and the output is the deal information (e.g., specific discounts and campaign information) to be provided to the user.
[1436] Step 7:
[1437] The server transmits information to the terminal.
[1438] The server sends the selected deals to the terminal. The input is the deals stored in the server, and the data is sent by communicating it to the terminal. The output is the data that has been transferred to the terminal.
[1439] Step 8:
[1440] The terminal displays the information.
[1441] The terminal displays the received discount information to the user. The input is the discount information sent from the server, and the output is the information that the user can confirm visually or audibly. For example, the message "All items are 10% off today!" is displayed on the terminal screen.
[1442] Through the above processing steps, the present invention makes it possible to obtain information about stores visited by a user from multiple angles and provide personalized deals information according to the user's emotions and circumstances.
[1443] (Application example 2)
[1444] Next, a description will be given of Application Example 2. In the following description, the data processing device 12 will be referred to as a "server" and the robot 414 will be referred to as a "terminal".
[1445] Conventional store information systems can only provide uniform deals to users, and it is difficult to provide information tailored to the emotions and circumstances of individual users. In addition, since the system does not take into account the emotions of users, the information provided may not always be appropriate, making it difficult to improve user satisfaction.
[1446] The identification process by the identification processing unit 290 of the data processing device 12 in the application example 2 is realized by the following means. In this invention, the server includes a means for recognizing and acquiring store information from image data, a means for acquiring the user's facial expression data together with the image data, a means for selecting advantageous information based on the acquired store information, a means for estimating an emotion based on the user's facial expression data and optimizing the advantageous information based on the emotion, and a means for displaying the optimized advantageous information on the smart glasses. This makes it possible to provide personalized information based on the user's emotions, and is expected to improve user satisfaction.
[1447] The "user-worn device" is a device worn by a user visiting a store, and has the function of acquiring image data and facial expression data.
[1448] The "store information acquisition means" is a function for recognizing store logos and signs from image data obtained from a user-worn device and acquiring that information.
[1449] The "facial expression data acquisition means" is a function for capturing the user's facial expression and acquiring that data.
[1450] The "information selection means" is a function for selecting advantageous information to be presented to the user based on the acquired store information.
[1451] The "emotion estimation means" is a function for estimating the user's emotion based on the acquired facial expression data of the user.
[1452] The "information optimization means" is a function for optimizing selected advantageous information based on the estimated user's emotions.
[1453] The "information display means" is a function for displaying optimized advantageous information on a user-worn device.
[1454] DETAILED DESCRIPTION OF THE PREFERRED EMBODIMENTS
[1455] The system of the present invention consists of smart glasses worn by the user, a server that processes images and location information, and an emotion engine.
[1456] 1. Hardware and Software Configuration
[1457] Smart Glasses
[1458] The smart glasses, worn by the user, have a camera and a display function. The camera captures store logos and signs, and also obtains the user's facial expression data. These data are then sent to a server.
[1459] server
[1460] The server uses the following hardware and software:
[1461] Hardware: High performance processor, storage device
[1462] Software: Python, OpenCV, requests library
[1463] The server analyzes the image data acquired from the smart glasses to recognize store information, and also uses an emotion engine to estimate the user's emotions from their facial expression data.
[1464] Emotion Engine
[1465] The emotion engine is software that analyzes the user's facial expression data and estimates their emotions. The basic emotion estimation algorithm uses a machine learning model.
[1466] 2. Data flow and processing
[1467] The image data acquired by the smart glasses is sent to a server. The server recognizes the store's logo or signboard from the image data and retrieves store information from a database. Next, the server sends the user's facial expression data acquired at the same time to an emotion engine to estimate the user's emotion.
[1468] Store information acquisition method
[1469] The server recognizes logos and signs from the image data and retrieves related store information from the database based on this. Specifically, the server uses the OpenCV library for image recognition.
[1470] Emotion estimation means
[1471] The user's facial expression data is sent to the emotion engine, which uses machine learning to classify the user's emotions into categories such as "happy," "sad," "excited," and "calm."
[1472] Information optimization measures
[1473] The server optimizes the deals based on the user's emotions. For example, if the user has a happy expression, the server will provide discount information.
[1474] Means of selecting and displaying information
[1475] The server sends the optimized information to the smart glasses to display to the user, including messages personalized to the user's emotions.
[1476] 3. Specific Examples
[1477] As a concrete example, consider the case where a user visits a cafe wearing smart glasses.
[1478] 1. When a user enters a cafe, the camera in the smart glasses recognizes the cafe's logo and sign and transmits it to the server.
[1479] 2. The server retrieves information about recognized cafes from the database.
[1480] 3. At the same time, the smart glasses capture the user's facial expression, and if the emotion engine estimates it to be "happy," the server will offer a "Special Happy Hour Discount."
[1481] 4. This information is then displayed on the smart glasses, allowing users to receive real-time deals.
[1482] Example prompts to be input to the generative AI model
[1483] When a smart glasses user enters a store, the store's logo and sign are recognized and the information is sent to the server. Create a program that then analyzes the user's emotions and displays discount and campaign information based on the emotions on the display. Show how to configure it in Python.
[1484] The flow of the specific process in the application example 2 will be described with reference to FIG.
[1485] Step 1:
[1486] Smart glasses capture user actions.
[1487] Input: Image data including the in-store environment and the user's facial expressions captured using the smart glasses camera.
[1488] Output: Image files containing store logos, signage, and user facial expressions.
[1489] How it works: The camera in the smart glasses captures images in real time and stores them in local storage.
[1490] Step 2:
[1491] The terminal transmits the captured image data to the server.
[1492] Input: Image data captured by smart glasses.
[1493] Output: Image data sent to the server.
[1494] Specific operation: The terminal communicates with the server via the Internet using HTTP, and sends image data as a POST request.
[1495] Step 3:
[1496] The server analyzes the image data it receives and recognizes store logos and signs.
[1497] Input: Image data including store logos and signs sent from the device.
[1498] Output: IDs of recognized stores and related information.
[1499] Specific operation: The server processes image data using the OpenCV library and detects and recognizes store logos and signs.
[1500] Step 4:
[1501] The store information recognized by the server is obtained from the database.
[1502] Input: The ID of a recognized store.
[1503] Output: Store details (e.g. store name, campaign information, discount information, etc.).
[1504] Specific operation: The server executes a database query using the store ID as a key to obtain related store information.
[1505] Step 5:
[1506] The terminal captures the user's facial expression data and transmits it to the server.
[1507] Input: Image data containing the user's facial expressions captured by the smart glasses.
[1508] Output: Facial expression data sent to the server.
[1509] Specific operation: The terminal communicates with the server via the Internet via HTTP and sends facial expression data as a POST request.
[1510] Step 6:
[1511] The server analyzes the facial expression data and uses an emotion engine to infer the user's emotions.
[1512] Input: User's facial expression data sent from the device.
[1513] Output: An estimated user emotion classification (e.g. happy, sad, excited, etc.).
[1514] Specific operation: The server uses an emotion engine (machine learning model) to analyze facial expression data and classify emotions.
[1515] Step 7:
[1516] The server selects the best deals based on the user's sentiment.
[1517] Input: Estimated user sentiment and store information.
[1518] Output: Optimized deals based on sentiment.
[1519] Specific operation: Based on the store information, the server selects the deals (discount information and campaign information) that best suit the user's emotions.
[1520] Step 8:
[1521] The server sends the selected information to the terminal and displays it on the smart glasses.
[1522] Enter: Optimized Deals.
[1523] Output: Deals displayed on the smart glasses.
[1524] Specific operation: The server sends the best deals to the device and displays the information on the smart glasses display.
[1525] This allows users to receive real-time, personalized deals when they visit a store.
[1526] The specific processing unit 290 transmits the result of the specific processing to the robot 414. In the robot 414, the control unit 46A causes the speaker 240 and the control target 443 to output the result of the specific processing. The microphone 238 acquires a voice indicating a user input for the result of the specific processing. The control unit 46A transmits voice data indicating the user input acquired by the microphone 238 to the data processing device 12. In the data processing device 12, the specific processing unit 290 acquires the voice data.
[1527] The data generation model 58 is a so-called generative AI (Artificial Intelligence). An example of the data generation model 58 is ChatGPT (Internet search<URL: https: / / openai.com / blog / chatgpt> ), Gemini (Internet search <url: https: gemini.google.com ?hl="ja">) and other generation AIs. The data generation model 58 is obtained by making a neural network perform deep learning. A prompt including an instruction is input to the data generation model 58, and inference data such as voice data indicating a voice, text data indicating a text, and image data indicating an image is input. The data generation model 58 performs inference on the input inference data according to the instruction indicated by the prompt, and outputs the inference result in a data format such as voice data and text data. Here, inference refers to, for example, analysis, classification, prediction, and / or summarization.
[1528] In the above embodiment, an example was given in which the specific process was performed by the data processing device 12, but the technology of the present disclosure is not limited to this, and the specific process may be performed by the robot 414.
[1529] The emotion identification model 59 as an emotion engine may determine the emotion of the user according to a specific mapping. Specifically, the emotion identification model 59 may determine the emotion of the user according to an emotion map (see FIG. 9), which is a specific mapping. Similarly, the emotion identification model 59 may determine the emotion of the robot, and the identification processing unit 290 may perform identification processing using the emotion of the robot.
[1530] FIG. 9 is a diagram showing an emotion map 400 on which multiple emotions are mapped. In the emotion map 400, emotions are arranged in concentric circles radiating from the center. The closer to the center of the concentric circles, the more primitive emotions are arranged. The more outside the concentric circles, the more emotions that represent states and actions that arise from a state of mind are arranged. Emotions are a concept that includes emotions and mental states. On the left side of the concentric circles, emotions that are generally generated from reactions that occur in the brain are arranged. On the right side of the concentric circles, emotions that are generally induced by situational judgment are arranged. On the upper and lower sides of the concentric circles, emotions that are generally generated from reactions that occur in the brain and are induced by situational judgment are arranged. In addition, on the upper side of the concentric circles, emotions of "pleasure" are arranged, and on the lower side, emotions of "discomfort" are arranged. In this way, in the emotion map 400, multiple emotions are mapped based on the structure in which emotions are generated, and emotions that tend to occur simultaneously are mapped close to each other.
[1531] These emotions are distributed in the three o'clock direction of emotion map 400 and usually fluctuate between relief and anxiety. In the right half of emotion map 400, situational awareness dominates over internal sensations, resulting in a sense of calm.
[1532] The inside of emotion map 400 represents what is going on inside one's mind, and the outside of emotion map 400 represents behavior, so the further out you go on emotion map 400, the more visible the emotions become (the more they are expressed in behavior).
[1533] Here, human emotions are based on various balances such as posture and blood sugar level, and when these balances are far from the ideal, it indicates an unpleasant state, and when they are close to the ideal, it indicates a pleasant state. Emotions can also be created for robots, cars, motorcycles, etc., based on various balances such as posture and battery level, so that when these balances are far from the ideal, it indicates an unpleasant state, and when they are close to the ideal, it indicates a pleasant state. The emotion map may be generated, for example, based on the emotion map of Dr. Mitsuyoshi (Research on speech emotion recognition and emotion brain physiological signal analysis system, Tokushima University, doctoral dissertation: https: / / ci.nii.ac.jp / naid / 500000375379). On the left half of the emotion map, emotions belonging to an area called "reaction" where sensation is dominant are lined up. On the right half of the emotion map, emotions belonging to an area called "situation" where situation recognition is dominant are lined up.
[1534] The emotion map defines two emotions that promote learning. The first is the negative emotion around the middle of "repentance" or "remorse" on the situation side. In other words, this is when the robot experiences negative emotions such as "I never want to feel this way again" or "I don't want to be scolded again." The other is the positive emotion around "desire" on the response side. In other words, this is when the robot has positive feelings such as "I want more" or "I want to know more."
[1535] The emotion identification model 59 inputs the user input to a pre-trained neural network, obtains emotion values indicating each emotion shown in the emotion map 400, and determines the emotion of the user. This neural network is pre-trained based on multiple learning data that are combinations of the user input and emotion values indicating each emotion shown in the emotion map 400. Furthermore, this neural network is trained so that emotions that are located close to each other have similar values, as in the emotion map 900 shown in Fig. 10. Fig. 10 shows an example in which multiple emotions, "relief," "calm," and "encouraging," have similar emotion values.
[1536] In the above embodiment, an example is given in which the specific processing is performed by one computer 22, but the technology of the present disclosure is not limited to this, and the specific processing may be distributed and performed by a plurality of computers including the computer 22. For example, the data generation model 58 may be provided in an external device of the data processing device 12, and the external device may generate data according to input data.
[1537] In the above embodiment, an example in which the specific processing program 56 is stored in the storage 32 has been described, but the technology of the present disclosure is not limited to this. For example, the specific processing program 56 may be stored in a portable computer-readable non-transitory storage medium such as a Universal Serial Bus (USB) memory. The specific processing program 56 stored in the non-transitory storage medium is installed in the computer 22 of the data processing device 12. The processor 28 executes the specific processing in accordance with the specific processing program 56.
[1538] In addition, the specific processing program 56 may be stored in a storage device such as a server connected to the data processing device 12 via the network 54, and the specific processing program 56 may be downloaded and installed on the computer 22 upon request from the data processing device 12.
[1539] It is not necessary to store all of the specific processing program 56 in a storage device such as a server connected to the data processing device 12 via the network 54, or to store all of the specific processing program 56 in the storage 32; only a portion of the specific processing program 56 may be stored.
[1540] As the hardware resource for executing the specific process, various processors as shown below can be used. An example of the processor is a CPU, which is a general-purpose processor that functions as a hardware resource for executing the specific process by executing software, i.e., a program. Another example of the processor is a dedicated electric circuit, which is a processor having a circuit configuration designed exclusively for executing the specific process, such as a Field-Programmable Gate Array (FPGA), a Programmable Logic Device (PLD), or an Application Specific Integrated Circuit (ASIC). Each processor has a built-in or connected memory, and each processor executes the specific process by using the memory.
[1541] The hardware resource that executes the specific process may be one of these various processors, or may be a combination of two or more processors of the same or different types (for example, a combination of multiple FPGAs, or a combination of a CPU and an FPGA). Also, the hardware resource that executes the specific process may be a single processor.
[1542] As an example of a configuration using one processor, first, one processor is configured by combining one or more CPUs and software, and this processor functions as a hardware resource that executes a specific process. Second, there is a configuration using a processor that realizes the functions of the entire system including multiple hardware resources that execute a specific process on a single IC chip, as typified by SoC (System-on-a-chip). In this way, a specific process is realized using one or more of the above-mentioned various processors as hardware resources.
[1543] Furthermore, more specifically, the hardware structure of these various processors can be an electric circuit that combines circuit elements such as semiconductor elements. The specific processes described above are merely examples. It goes without saying that unnecessary steps may be deleted, new steps may be added, or the order of processes may be changed without departing from the spirit of the invention.
[1544] The above description and illustrations are detailed descriptions of the parts related to the technology of the present disclosure, and are merely an example of the technology of the present disclosure. For example, the above description of the configuration, function, action, and effect is an example of the configuration, function, action, and effect of the parts related to the technology of the present disclosure. Therefore, it goes without saying that unnecessary parts may be deleted, new elements may be added, or replacements may be made to the above description and illustrations, within the scope of the gist of the technology of the present disclosure. In addition, in order to avoid confusion and to facilitate understanding of the parts related to the technology of the present disclosure, the above description and illustrations omit explanations of technical common sense that do not require explanation in order to enable the implementation of the technology of the present disclosure.
[1545] In addition, in each of the above embodiments, some of the configurations may be combined in any manner.
[1546] All publications, patent applications, and standards mentioned in this specification are herein incorporated by reference to the same extent as if each individual publication, patent application, or standard was specifically and individually indicated to be incorporated by reference.
[1547] The following is further disclosed regarding the above embodiment.
[1548] (Appendix 1) A system for providing advantageous information to a user wearing smart glasses when the user visits a store, comprising: Acquire information about the store, Select the advantageous information based on the acquired store information, The system displays the selected deals on the smart glasses.
[1549] (Appendix 2) The system according to claim 1, wherein the store information is obtained by recognizing the store's logo or signboard from image data obtained from smart glasses.
[1550] (Appendix 3) 2. The system of claim 1, The system selects the advantageous information based on the attributes or past behavioral history of the user.
[1551] (Appendix 4) 2. The system of claim 1, A system that recognizes the user's emotions using an emotion engine that recognizes the user's emotions, and selects the special offers based on the recognition result.
[1552] (Appendix 5) 5. The system according to claim 4, wherein the emotion engine has a function for recognizing facial expressions and voice characteristics of a user and estimating the user's emotions.
[1553] "Example 1"
[1554] (Claim 1) A system for providing advantageous information to a user when the user visits a store using a mobile information terminal, comprising: A means for acquiring information about the store based on image data and location information obtained from a mobile information terminal; A means for selecting the advantageous information from the acquired store information; a means for displaying the selected advantageous information on the mobile information terminal; A system including:
[1555] (Claim 2) 2. The system according to claim 1, wherein image data obtained from a mobile information terminal is analyzed, and information about the store is acquired by recognizing the logo or signboard of the store.
[1556] (Claim 3) The system according to claim 1, wherein the advantageous information is selected taking into consideration the attributes and past behavioral history of the user together with the acquired store information.
[1557] "Application example 1"
[1558] (Claim 1) A means for acquiring image data and location information; A means for transmitting the acquired image data and position information to a server; A means for the server to search for store information from a database based on the image data and the location information and provide the selected advantageous information to the user; means for displaying the deals on the smart glasses; A means for recognizing store logos and signs from the image data; A means for selecting optimal information in consideration of the user's past behavior history and attributes; A system including:
[1559] (Claim 2) The system according to claim 1, wherein the store information is obtained from a cloud storage database.
[1560] (Claim 3) The system of claim 1 , further comprising a generative AI model for taking into account the user attributes.
[1561] "Example 2 of combining emotion engines"
[1562] (Claim 1)
[1563] A means for the terminal to acquire user location information; A means for the terminal to acquire image data using a camera and recognize a store's logo or signboard based on the image data; A means for transmitting image data and location information from the terminal to a server; A means for the server to search the database and obtain store information; A means for analyzing a facial expression or voice characteristic of a user using an emotion engine to estimate an emotion of the user; A means for the server to select deals based on user sentiment; A means for the server to send advantageous information to the terminal; A means for the terminal to display advantageous information to the user; A system including:
[1564] (Claim 2) 2. The system according to claim 1, wherein a terminal having a GPS function and a camera function is used to obtain the user's location information and image data.
[1565] (Claim 3) 10. The system of claim 1, further comprising an emotion engine for inferring the user's emotion.
[1566] "Application example 2 when combining emotion engines"
[1567] (Claim 1) A system for providing advantageous information to a user when the user wears smart glasses and visits a store, A means for recognizing and acquiring information about the store from image data; means for acquiring facial expression data of a user together with the image data; A means for selecting the advantageous information based on the acquired store information; A means for estimating an emotion based on facial expression data of the user and optimizing the special offer information based on the emotion; means for displaying the optimized deals on the smart glasses; A system including:
[1568] (Claim 2) The system of claim 1, wherein the store information is obtained by recognizing the store's logo or signboard from image data obtained from the smart glasses.
[1569] (Claim 3) The system according to claim 1, wherein the advantageous information is selected based on attributes or past behavioral history of the user. [Explanation of symbols]
[1570] 10, 210, 310, 410 Data Processing Systems 12 Data Processing Device 14 Smart Devices 214 Smart Glasses 314 Headset-type terminal 414 Robot< / url:> < / url:> < / url:> < / url:>
Claims
1. A system for providing advantageous information to a user when the user carries a mobile information terminal and visits a store, comprising: means for recognizing an emotion of the user using an emotion engine for recognizing an emotion of the user; A means for acquiring information about the store based on image data or location information obtained from the mobile information terminal; A means for generating an output result showing the advantageous information based on a prompt sentence instructing the user to select the advantageous information from the acquired store information, attribute data of the user, or a past behavior history, and a generative AI; a means for outputting, to the mobile information terminal, an output result indicating the generated advantageous information; A system including:
2. The system according to claim 1 , wherein image data obtained from the mobile information terminal is analyzed, and information about the store is acquired by recognizing a logo or a signboard of the store.
3. The system according to claim 1 , wherein the advantageous information includes at least one of campaign information, special offer information, and discount information that is useful for the user at the store.
Citation Information
Patent Citations
Persona chatbot control method and system
JP2022180282A