System
A system using facial recognition and real-time data analysis optimizes POP displays in retail environments by personalizing content and updating AI models based on customer behavior, enhancing sales promotion effectiveness.
Patent Information
- Application Number
- JP2024122801
- Authority / Receiving Office
- JP · JP
- Patent Type
- Applications
- Current Assignee / Owner
- Filing Date
- 2024-07-29
- Publication Date
- 2026-02-10
AI Technical Summary
In retail environments like shopping malls and supermarkets, traditional point-of-purchase (POP) displays are ineffective in responding to individual customer needs and providing real-time information, lacking mechanisms to utilize customer purchasing behavior data for personalized and dynamic content generation.
A system incorporating facial recognition, customer attribute analysis, real-time POP content generation, display systems, gesture and voice recognition, personalized information provision, and AI model updating based on collected data to optimize and personalize content.
Enables real-time provision of personalized information and optimal product recommendations, continuously improving the effectiveness of POP content through data-driven updates.
Smart Images

Figure 2026021119000001_ABST
Abstract
Description
[Technical Field]
[0001] The technology of the present disclosure relates to a system. [Background technology]
[0002] Patent document 1 discloses a persona chatbot control method performed by at least one processor, the method including the steps of receiving a user utterance, adding the user utterance to a prompt including an instruction sentence related to a description of the chatbot character, encoding the prompt, and inputting the encoded prompt into a language model to generate a chatbot utterance in response to the user utterance. [Prior art documents] [Patent documents]
[0003] [Patent Document 1] Japanese Patent Publication No. 2022-180282 Summary of the Invention [Problem to be solved by the invention]
[0004] In the retail industry, particularly in shopping malls and supermarkets, effective sales promotions aimed at diverse customer segments are a challenge. Traditional point-of-purchase (POP) creation relies on the craftsmanship of experienced staff, and the results are heavily influenced by human factors. Furthermore, fixed POPs make it difficult to respond to individual customer needs or provide real-time information. Furthermore, in order to continuously create and improve effective POPs, it is desirable to utilize customer purchasing behavior data and make highly effective proposals. [Means for solving the problem]
[0005] The present invention solves these problems with a system that includes a facial recognition system, a customer attribute analysis system, a product information database, a system for generating POP content in real time, a display system for displaying the generated POP content, a system for recognizing customer pointing gestures, a system for recognizing customer questions by voice, a system for providing information based on the voice questions, a system for personalizing information, a system for collecting purchase data, and a system for updating an AI model based on the collected data. By selecting optimal products and services based on customer attributes, optimal proposals can be made that meet individual customer needs. Furthermore, by optimizing the effectiveness of the generated POP content based on actual purchase data, effective POPs that are continuously improved can be generated.
[0006] "Facial recognition means" refers to technology that uses a camera to capture customers' faces and automatically analyzes attribute data such as age, gender, and number of people through AI algorithms.
[0007] "Customer attribute analysis means" refers to a function that uses an AI algorithm to estimate the age, gender, and number of customers from acquired data such as facial images, and updates this data in real time.
[0008] "Product information database" refers to a database that stores and manages information about various products and services, and allows for appropriate searching and reference.
[0009] "Means for generating POP content in real time" refers to technology that automatically selects the most suitable products and services based on customer attribute data and generates POP content dynamically and in a timely manner.
[0010] The "display means for displaying the generated POP content" refers to a display device that receives the POP content sent from the server and displays it to the customer in a multimedia format.
[0011] "Means for recognizing customer pointing gestures" refers to the function of using a camera and gesture recognition technology to detect the specific product that a customer points at among the products displayed on the screen.
[0012] "Means for recognizing customer questions by voice" refers to technology that uses a microphone to capture customer voice questions and converts them into text data through a voice recognition engine.
[0013] "Means for providing information based on voice questions" refers to technology that converts a customer's voice questions into text data, retrieves the relevant information from a server, and provides it.
[0014] "Means for personalizing information" refers to technology that references membership databases and purchase history data to provide customized messages and special offer information to specific customers.
[0015] "Means for collecting purchasing data" refers to a system that collects and stores data on customer purchasing behavior and actual purchased products in real time.
[0016] "Means for updating AI models based on collected data" refers to technology that uses collected purchasing data to continuously train AI algorithms, optimizing the accuracy and effectiveness of POP content generation. [Brief explanation of the drawings]
[0017] [Figure 1] 1 is a conceptual diagram showing an example of the configuration of a data processing system according to a first embodiment. [Figure 2] 1 is a conceptual diagram showing an example of main functions of a data processing device and a smart device according to a first embodiment. [Figure 3] FIG. 10 is a conceptual diagram showing an example of the configuration of a data processing system according to a second embodiment. [Figure 4] FIG. 10 is a conceptual diagram showing an example of main functions of a data processing device and smart glasses according to a second embodiment. [Figure 5]FIG. 10 is a conceptual diagram showing an example of the configuration of a data processing system according to a third embodiment. [Figure 6] FIG. 11 is a conceptual diagram showing an example of main functions of a data processing device and a headset-type terminal according to a third embodiment. [Figure 7] FIG. 10 is a conceptual diagram showing an example of the configuration of a data processing system according to a fourth embodiment. [Figure 8] FIG. 10 is a conceptual diagram showing an example of main functions of a data processing device and a robot according to a fourth embodiment. [Figure 9] 1 shows an emotion map onto which multiple emotions are mapped. [Figure 10] 1 shows an emotion map onto which multiple emotions are mapped. [Figure 11] FIG. 3 is a sequence diagram showing a processing flow of the data processing system according to the first embodiment. [Figure 12] FIG. 10 is a sequence diagram showing the flow of processing in the data processing system in Application Example 1. [Figure 13] FIG. 10 is a sequence diagram showing the flow of processing in the data processing system according to the second embodiment when an emotion engine is combined. [Figure 14] FIG. 10 is a sequence diagram showing the flow of processing in the data processing system in Application Example 2 when an emotion engine is combined. DETAILED DESCRIPTION OF THE INVENTION
[0018] An example of an embodiment of a system according to the technology of the present disclosure will be described below with reference to the accompanying drawings.
[0019] First, the terms used in the following description will be explained.
[0020] In the following embodiments, a coded processor (hereinafter simply referred to as a "processor") may be a single arithmetic device or a combination of multiple arithmetic devices. Furthermore, a processor may be a single type of arithmetic device or a combination of multiple types of arithmetic devices. Examples of arithmetic devices include a CPU (Central Processing Unit), a GPU (Graphics Processing Unit), a GPGPU (General-Purpose computing on Graphics Processing Units), and an APU (Accelerated Processing Unit).
[0021] In the following embodiments, a coded RAM (Random Access Memory) is a memory in which information is temporarily stored and is used as a working memory by a processor.
[0022] In the following embodiments, the coded storage is one or more non-volatile storage devices that store various programs, various parameters, etc. Examples of non-volatile storage devices include flash memory (SSD (Solid State Drive)), magnetic disks (e.g., hard disks), and magnetic tapes.
[0023] In the following embodiments, a communication I / F (Interface) with a symbol is an interface including a communication processor, an antenna, etc. The communication I / F controls communication between multiple computers. Examples of communication standards applied to the communication I / F include wireless communication standards including 5G (5th Generation Mobile Communication System), Wi-Fi (registered trademark), Bluetooth (registered trademark), etc.
[0024] In the following embodiments, "A and / or B" is synonymous with "at least one of A and B." In other words, "A and / or B" means that it may be only A, only B, or a combination of A and B. Furthermore, in this specification, the same concept as "A and / or B" is also applied when three or more things are expressed connected by "and / or."
[0025] [First embodiment]
[0026] FIG. 1 shows an example of the configuration of a data processing system 10 according to the first embodiment.
[0027] 1, a data processing system 10 includes a data processing device 12 and a smart device 14. An example of the data processing device 12 is a server.
[0028] The data processing device 12 includes a computer 22, a database 24, and a communication I / F 26. The computer 22 is an example of a "computer" according to the technology of the present disclosure. The computer 22 includes a processor 28, a RAM 30, and a storage 32. The processor 28, the RAM 30, and the storage 32 are connected to a bus 34. The database 24 and the communication I / F 26 are also connected to the bus 34. The communication I / F 26 is connected to a network 54. Examples of the network 54 include a WAN (Wide Area Network) and / or a LAN (Local Area Network).
[0029] The smart device 14 includes a computer 36, a reception device 38, an output device 40, a camera 42, and a communication I / F 44. The computer 36 includes a processor 46, a RAM 48, and a storage 50. The processor 46, the RAM 48, and the storage 50 are connected to a bus 52. The reception device 38, the output device 40, and the camera 42 are also connected to the bus 52.
[0030] The reception device 38 includes a touch panel 38A, a microphone 38B, and the like, and receives user input. The touch panel 38A detects contact with an indicator (for example, a pen or a finger) to receive user input by the touch of the indicator. The microphone 38B detects the user's voice to receive user input by voice. The control unit 46A transmits data indicating the user input received by the touch panel 38A and the microphone 38B to the data processing device 12. In the data processing device 12, the specific processing unit 290 acquires the data indicating the user input.
[0031] The output device 40 includes a display 40A and a speaker 40B, and presents data to the user 20 by outputting the data in a form of expression that the user 20 can perceive (for example, audio and / or text). The display 40A displays visible information such as text and images in accordance with instructions from the processor 46. The speaker 40B outputs audio in accordance with instructions from the processor 46. The camera 42 is a compact digital camera equipped with an optical system including a lens, aperture, and shutter, and an imaging element such as a CMOS (Complementary Metal-Oxide-Semiconductor) image sensor or a CCD (Charge Coupled Device) image sensor.
[0032] The communication I / F 44 is connected to a network 54. The communication I / Fs 44 and 26 control the exchange of various information between the processor 46 and the processor 28 via the network 54.
[0033] FIG. 2 shows an example of the main functions of the data processing device 12 and the smart device 14.
[0034] 2, in the data processing device 12, a specific process is performed by the processor 28. A specific processing program 56 is stored in the storage 32. The specific processing program 56 is an example of a "program" according to the technology of the present disclosure. The processor 28 reads the specific processing program 56 from the storage 32 and executes the read specific processing program 56 on the RAM 30. The specific process is realized by the processor 28 operating as a specific processing unit 290 in accordance with the specific processing program 56 executed on the RAM 30.
[0035] The storage 32 stores a data generation model 58 and an emotion identification model 59. The data generation model 58 and the emotion identification model 59 are used by the identification processing unit 290.
[0036] In the smart device 14, the processor 46 performs the reception output process. The storage 50 stores a reception output program 60. The reception output program 60 is used in conjunction with the specific processing program 56 by the data processing system 10. The processor 46 reads the reception output program 60 from the storage 50 and executes the read reception output program 60 on the RAM 48. The reception output process is realized by the processor 46 operating as the control unit 46A in accordance with the reception output program 60 executed on the RAM 48.
[0037] Next, a description will be given of the specific processing performed by the specific processing unit 290 of the data processing device 12. In the following description, the data processing device 12 will be referred to as a "server" and the smart device 14 will be referred to as a "terminal."
[0038] This invention relates to a digital signage-type POP system that introduces optimal products and services in real time based on customer attribute information. This system is used in retail stores and shopping malls to provide personalized information to individual customers.
[0039] System Configuration
[0040] This system is broadly composed of a "server," a "terminal," and a "user." Details of each component are explained below.
[0041] server
[0042] The server has the following functions:
[0043] 1. Facial Recognition Methods:
[0044] The server captures customers' faces through a camera and obtains attribute data such as age, gender, and number of people.
[0045] Using high-resolution video data, customer attributes are analyzed using facial recognition algorithms.
[0046] 2. Customer attribute analysis means:
[0047] Based on facial recognition, analyzed attribute data is stored in a database and updated in real time.
[0048] The analysis results are used to provide promotional information to each customer.
[0049] 3. Product information database and personalized message generation:
[0050] The product information database stores detailed information about each product and service, allowing users to search and reference the information they need.
[0051] Customer attribute data is referenced to select the most suitable products and services and dynamically generate POP content.
[0052] Add personalized messages from your membership database and purchase history to generate individual messages for specific customers.
[0053] 4. Learning and optimization:
[0054] The server analyzes collected customer behavioral and purchasing data and updates the AI model to optimize the effectiveness of POP content.
[0055] We continuously collect data and improve our algorithms to generate optimal content.
[0056] Terminal
[0057] The terminal is installed inside the store and serves as an interface with customers.
[0058] 1. Display means:
[0059] Real-time POP content sent from the server is displayed on the display.
[0060] The content can be displayed in the form of images, videos, text, etc.
[0061] 2. Pointing recognition method:
[0062] The camera recognizes the customer's pointing movements in real time.
[0063] A request is made to the server for detailed information about the product pointed to by the customer, and the received information is displayed on the screen.
[0064] 3. Question and Answer Function:
[0065] The customer's voice questions are captured using a microphone and converted into text data using voice recognition technology.
[0066] Information based on the question is obtained from the server and displayed on the screen.
[0067] User
[0068] The user is a customer of the store and performs the following operations.
[0069] 1. Viewing content:
[0070] Users view the POP content displayed on the display and check information about products and services that interest them.
[0071] 2. Questions and Pointing:
[0072] Users can ask for details about products they are interested in by voice or by pointing at specific products on the screen.
[0073] The question and information about the product you point at will be displayed on the screen to help you make a purchasing decision.
[0074] Specific examples
[0075] Example 1: Parent and child visitor scenario
[0076] A parent and child visit a shopping mall. The server captures the faces of the parent and child via camera and determines that the parent is in their late 30s and the child is about 7 years old. Based on this, the device displays information mainly about toys and events for children. When the parent asks, "Where are the toy sales?", the device recognizes the voice question, retrieves the answer from the server, and displays it on the display.
[0077] Example 2: Checking details of individual products
[0078] The customer points at the image of a product displayed on the display. The terminal recognizes the pointing gesture and retrieves detailed information about that product from the server. The product's price, specifications, stock status, and other information are displayed on the display in real time.
[0079] As described above, this system can provide optimal information in real time based on customer attributes and behavior, thereby enhancing sales promotion effectiveness. Furthermore, by continually updating the AI model based on collected data, it is possible to continuously generate effective content.
[0080] The processing flow will be explained below.
[0081] Program processing steps
[0082] Server Processing Steps
[0083] Step 1:
[0084] The server captures the faces of customers in the store through a camera and obtains high-resolution image data.
[0085] Step 2:
[0086] The server analyzes the acquired facial images using an AI algorithm to estimate attribute data such as the customer's age, gender, and number of people.
[0087] Step 3:
[0088] The server updates the analysis results to a database and stores customer attribute data in real time.
[0089] Step 4:
[0090] The server searches for the most suitable products and services from a product information database based on customer attribute data.
[0091] Step 5:
[0092] The server dynamically generates POP content in real time based on the search results.
[0093] Step 6:
[0094] The server references the member database and purchase history and generates personalized messages as needed.
[0095] Step 7:
[0096] The server transmits the generated POP content to the terminal.
[0097] Step 8:
[0098] The server analyzes the collected customer behavior data and purchase data to evaluate the effectiveness of the POP content.
[0099] Step 9:
[0100] Based on the effectiveness evaluation results, the server updates the AI model and improves the accuracy of future content generation algorithms.
[0101] Terminal processing steps
[0102] Step 1:
[0103] The terminal receives the POP content sent from the server.
[0104] Step 2:
[0105] The device displays the received content on the display, which can be in the form of images, videos, text, etc.
[0106] Step 3:
[0107] The device uses a camera to recognize the customer's pointing movements in real time.
[0108] Step 4:
[0109] The terminal requests information about the product pointed at by the customer from the server.
[0110] Step 5:
[0111] The terminal receives the detailed product information sent from the server and displays it on the display.
[0112] Step 6:
[0113] The terminal uses a microphone to capture the customer's voice asking a question.
[0114] Step 7:
[0115] The device uses voice recognition technology to convert the voice into text data and send it to the server.
[0116] Step 8:
[0117] The terminal displays the response information sent from the server on the display.
[0118] User operation steps
[0119] Step 1:
[0120] The user views the POP content displayed on the display.
[0121] Step 2:
[0122] Users can input questions about products or services that interest them by voice.
[0123] Step 3:
[0124] The user points to a particular product on the display to see more information about it.
[0125] Step 4:
[0126] Users make purchasing decisions based on the displayed information.
[0127] Specific examples
[0128] Example 1: Parent-child visit scenario
[0129] Step 1: A parent and child visit a shopping mall. The server captures the faces of the parent and child through a camera and analyzes them.
[0130] Step 2: The server estimates the age and gender of the parents and children and updates the attribute data in the database.
[0131] Step 3: The server selects toys and event information for children based on the parent-child attributes and generates POP content.
[0132] Step 4: The terminal receives the POP content and displays it on the display.
[0133] Step 5: The parent asks, "Where are the toy sales?" The device captures the audio and sends it to the server.
[0134] Step 6: The server generates toy department information based on the query and sends it to the terminal.
[0135] Step 7: The terminal displays the received information on its display.
[0136] Example 2: Checking details of individual products
[0137] Step 1: The customer points to the product image on the display.
[0138] Step 2: The device recognizes the pointing gesture and requests detailed information about the product from the server.
[0139] Step 3: The server obtains the product details and sends them to the terminal.
[0140] Step 4: The terminal displays the received information on its display.
[0141] Step 5: The customer checks the product price, specifications, and stock information and makes a purchasing decision.
[0142] The above is the specific flow of operations at each step. This invention realizes the provision of personalized information in real time based on the attributes and behavior of customers.
[0143] Example 1
[0144] Next, a description will be given of Example 1. In the following description, the data processing device 12 will be referred to as a "server" and the smart device 14 will be referred to as a "terminal."
[0145] Conventional digital signage systems have the problem that the information provided to customers is uniform and not personalized enough for each individual customer. As a result, it is not possible to introduce optimal products and services that meet each customer's interests and needs, and sales promotion effectiveness is limited. Furthermore, there is a lack of a mechanism to effectively collect and analyze customer behavior data and optimize content in real time.
[0146] The specific processing by the specific processing unit 290 of the data processing device 12 in the first embodiment is realized by the following means.
[0147] In this invention, the server includes a means for capturing customer faces and acquiring attribute data, a means for analyzing and saving the acquired attribute data, and a database for accumulating product information. This makes it possible to select and display optimal products and services in real time according to customer attributes. Furthermore, by collecting purchasing data and updating and optimizing the AI model based on this, it is possible to consistently provide effective content.
[0148] "Means for capturing customer faces and acquiring attribute data" refers to technology that uses cameras and sensors to capture images of customers' faces in a store and extract information such as age, gender, and number of customers from that image data.
[0149] "Means for analyzing and storing acquired attribute data" refers to hardware and software for analyzing attribute data obtained by facial recognition technology and storing it in a database.
[0150] A "database that stores product information" is a database system that manages detailed information about each product and service sold in a retail store or shopping mall, and allows for searching and referencing.
[0151] "Means for generating personalized POP content in real time" refers to algorithms and software for dynamically creating and displaying promotional information that is optimal for each customer based on their attribute data.
[0152] The "display means for displaying the generated content" refers to a display device for visually presenting the POP content sent from the server to customers in the form of images, videos, text, etc.
[0153] "Means for recognizing customer pointing movements" refers to technology that uses sensors such as cameras to detect customer pointing movements and analyze those movements.
[0154] "Means for recognizing customer questions via voice" refers to voice recognition technology that uses a microphone to capture the customer's voice and converts the voice data into text.
[0155] "Means for providing information based on voice questions" refers to a system that searches for appropriate information based on the question content obtained by voice recognition and provides that information.
[0156] "Means for personalizing information" refers to algorithms and systems for providing the most appropriate information to individual customers based on their attribute data and purchasing history.
[0157] "Means for collecting purchasing data" refers to a system for collecting customer purchasing history and behavioral data, and storing and managing this in a database.
[0158] "Means for updating AI models based on collected data and optimizing POP content" refers to a system that uses collected customer behavioral and purchasing data to continuously learn and update AI models to generate optimal POP content.
[0159] This invention relates to a digital signage-type POP system that introduces optimal products and services in real time based on customer attribute information. This system is used in retail stores and shopping malls to provide personalized information to individual customers.
[0160] System Configuration
[0161] This system is broadly composed of a "server," a "terminal," and a "user." Each component will be explained in detail below.
[0162] server
[0163] The server has the following functions:
[0164] 1. A means of capturing customer faces and obtaining attribute data
[0165] The server monitors the store in real time through installed cameras, capturing images of customers' faces and acquiring attribute data such as age, gender, number of people, etc. This facial recognition is performed using Python and OpenCV, applying facial recognition algorithms such as YOLO.
[0166] 2. Means for analyzing and storing acquired attribute data
[0167] The server analyzes the acquired customer attribute data and saves it in a database. The attribute data is continuously updated and to manage the information of multiple customers in real time, a database management system such as MySQL or PostgreSQL is used. The data is managed via an API using the Django or Flask framework.
[0168] 3. Database for storing product information
[0169] The product information database stores detailed information about the products it handles, allowing for quick search and reference as needed. This database includes information such as product prices, specifications, and stock status.
[0170] 4. A way to generate personalized POP content in real time
[0171] The server selects the most suitable products and services based on customer attribute data and dynamically generates personalized POP content using a generative AI model (e.g., GPT-3). The generated POP content is converted into HTML or image format as a text message or a list of recommended products and sent to the device.
[0172] 5. Display means for displaying generated content
[0173] The terminal visually displays real-time POP content sent from the server on a display, which can be powered by a Raspberry Pi or Windows PC and runs as a web application in a browser.
[0174] Terminal
[0175] The terminal is installed inside the store and serves as an interface with customers.
[0176] 1. A method for recognizing customer pointing gestures
[0177] The device uses a camera to recognize the customer's pointing gestures in real time, using hand movement recognition algorithms such as TensorFlow and OpenPose, and sends the processing results to a server using a Python script.
[0178] 2. A way to recognize customer questions by voice
[0179] The device uses a microphone to capture the customer's voice and converts it into text using speech recognition technology, using the Google Cloud Speech-to-Text API.
[0180] 3. Means of providing information based on voice queries
[0181] The server searches for appropriate information based on the question obtained through voice recognition and sends that information to the terminal, which then displays the information on its screen.
[0182] User
[0183] The user is a customer of the store and performs the following operations.
[0184] 1. Viewing content
[0185] Users view the POP content displayed on the display and check information about products and services that interest them.
[0186] 2. Questions and Pointing
[0187] Users can ask for details about products they are interested in by voice or by pointing at a specific product on the screen. The question or information about the product they are pointing at is then displayed on the screen to help them make a purchasing decision.
[0188] Specific examples
[0189] Example 1: Parent and child visitor scenario
[0190] A parent and child visit a shopping mall. The server captures the faces of the parent and child via camera and determines that the parent is in their late 30s and the child is about 7 years old. Based on this, the device displays information mainly about toys and events for children. When the parent asks, "Where are the toy sales?", the device recognizes the voice question, retrieves the answer from the server, and displays it on the display.
[0191] Example 2: Checking details of individual products
[0192] The customer points at the image of a product displayed on the display. The terminal recognizes the pointing gesture and retrieves detailed information about that product from the server. The product's price, specifications, stock status, and other information are displayed on the display in real time.
[0193] Prompt Sentence Examples
[0194] "What kind of information should be displayed when a parent and child in their 30s visit?"
[0195] "When a customer points to a specific product, explain how you would retrieve and display information about that product."
[0196] This invention makes it possible to provide optimal information in real time based on customer attributes and behavior, thereby enhancing sales promotion effectiveness. In addition, by continually updating the AI model based on collected data, it is possible to continuously generate effective content.
[0197] The flow of the identification process in the first embodiment will be described with reference to FIG.
[0198] Step 1: Discover customers and obtain attribute data
[0199] The server monitors the store in real time through the installed cameras. It captures video data and uses it as input to extract attribute information such as age, gender, and number of people. This is done using Python and OpenCV, and processing high-resolution video data with facial recognition algorithms such as YOLO.
[0200] Input: Real-time video data from the camera
[0201] Output: Customer attribute data (age, gender, number of people, etc.)
[0202] Step 2: Parse and store customer attribute data
[0203] The server analyzes the acquired customer attribute data and stores it in a database. Database management uses MySQL or PostgreSQL, and data management and updating is performed using the Django or Flask framework.
[0204] Input: Customer attribute data
[0205] Output: Analysis data stored in a database
[0206] Step 3: Generate optimal POP content
[0207] The server selects the most suitable products and services from a product information database based on the accumulated customer attribute data, and dynamically generates personalized POP content using a generative AI model (e.g., GPT-3) and sends it to the device in HTML or image format.
[0208] Input: Customer attribute data, product information
[0209] Output: Personalized POP content
[0210] Step 4: View POP Content
[0211] The terminal displays POP content sent from the server in real time. The displayed content includes images, videos, and text. It runs as a web application in a browser on a Raspberry Pi or Windows PC.
[0212] Input: POP content sent from the server
[0213] Output: Content displayed on the display
[0214] Step 5: Recognizing pointing gestures and providing detailed information
[0215] The device uses a camera to recognize the customer's pointing movements in real time and sends the data to a server. The server then analyzes the movement data and obtains detailed information about the product the customer is pointing at. The movement recognition uses TensorFlow and OpenPose.
[0216] Input: Customer pointing gesture
[0217] Output: Detailed information about the pointed item
[0218] Step 6: Recognize voice questions and provide answers
[0219] The device uses a microphone to capture the customer's voice question and converts it into text using the Google Cloud Speech-to-Text API. The server then analyzes the question from the converted text, searches for and retrieves the appropriate answer, and sends it to the device, which then displays the information on its screen.
[0220] Input: Voice question data
[0221] Output: Answer information for the question
[0222] Step 7: Collect purchasing data and update the AI model
[0223] The server periodically collects customer purchase data and stores it in a database. The collected data is used to train and update the AI model, optimizing the effectiveness of the generated POP content.
[0224] Input: Purchasing data
[0225] Output: Updated AI model, optimized POP content
[0226] (Application example 1)
[0227] Next, a description will be given of Application Example 1. In the following description, the data processing device 12 will be referred to as a "server" and the smart device 14 will be referred to as a "terminal."
[0228] Conventional digital signage-type POP systems were primarily intended to provide information to customers in retail stores and shopping malls. However, in the logistics field, optimizing worker flow and supporting rapid inventory management are important, and there were few systems that met these needs. Therefore, there is a demand for technology that can improve work efficiency and increase the accuracy of inventory management in logistics centers, warehouses, and other on-site locations.
[0229] The specific processing by the specific processing unit 290 of the data processing device 12 in the application example 1 is realized by the following means.
[0230] In this invention, the server includes a face recognition unit, a customer attribute analysis unit, a product information database, a unit for generating POP content in real time, a display unit for displaying the generated POP content, a unit for recognizing customer pointing gestures, a unit for recognizing customer questions by voice, a unit for providing information based on the voice questions, a unit for personalizing information, a unit for collecting purchase data, a unit for updating an AI model based on the collected data, a unit for optimizing worker movement lines, a unit for visually guiding inventory locations, and a unit for linking with smart glasses or a head-mounted display. This makes it possible to optimize worker movement lines and provide real-time visual guidance on inventory locations at sites such as logistics centers and warehouses.
[0231] "Facial recognition means" is a technology that uses a camera to capture the face of a worker and analyzes attribute data such as age and gender based on the image.
[0232] "Customer attribute analysis means" is a technology that analyzes the tendencies and characteristics of specific customers based on attribute data acquired by face recognition means.
[0233] The "product information database" is a database that stores detailed information about each product or service, allowing users to search and reference the information they need.
[0234] "Means for generating POP content in real time" refers to technology for dynamically generating personalized POP content based on data collected in real time.
[0235] "Display means" refers to a device that displays generated POP content and information in real time, and includes display formats such as images, videos, and text.
[0236] The "means for recognizing pointing gestures" is a technology that uses a camera to detect the pointing gestures of workers and customers and analyzes their location information.
[0237] "Voice recognition means" refers to a technology that uses a microphone to capture the voice of a worker or customer and converts it into text data using voice recognition technology.
[0238] "Means for providing information based on voice questions" refers to a technology that acquires appropriate information from a server based on the content of a voice-recognized question and displays it on a display means.
[0239] "Means for personalizing information" refers to technology that generates individually optimized information and messages based on the attributes and behavioral data of customers and workers.
[0240] "Means for collecting purchasing data" refers to technology that collects data on customers' actual purchasing behavior and selected products and stores it in a database.
[0241] "Means for updating AI models based on collected data" refers to technologies for analyzing collected data and continuously training and optimizing AI models.
[0242] "Means to optimize worker movement" refers to technology that analyzes worker location information and movement in real time and proposes optimal movement routes and work procedures.
[0243] "Means for visually guiding inventory location" refers to technology that uses smart glasses or head-mounted displays within logistics centers to visually show workers the exact location of inventory.
[0244] "Means for linking with smart glasses or head-mounted displays" refers to technology that links smart glasses or head-mounted displays with the system to display information in real time and accept input.
[0245] This invention relates to a digital signage system for improving work efficiency in logistics centers and warehouses. This system uses facial recognition to analyze worker attributes, optimize worker movement lines, and visually guide inventory locations. Furthermore, by linking with smart glasses or head-mounted displays, work efficiency can be further improved.
[0246] System Configuration
[0247] This system is broadly composed of a server, terminals, and users. Each component is explained in detail below.
[0248] server
[0249] The server has the following functions:
[0250] 1. Facial Recognition Methods
[0251] The camera captures the worker's face and acquires attribute data such as age and gender, which is then analyzed using a facial recognition algorithm.
[0252] 2. Customer attribute analysis means
[0253] The system analyzes worker characteristics based on attribute data acquired through facial recognition, which is updated in real time and stored in a database.
[0254] 3. Product information database
[0255] This database comprehensively manages inventory information within the distribution center and is used to identify optimal inventory locations and routes.
[0256] 4. A way to generate POP content in real time
[0257] Generate personalized POP content in real time based on worker attribute information and inventory information.
[0258] 5. Optimizing traffic flow
[0259] Analyzes the real-time location information of workers and calculates and suggests the optimal movement route.
[0260] 6. Inventory location visual guide means
[0261] The exact location of inventory is displayed in real time on smart glasses or head-mounted displays.
[0262] 7. AI model update methods
[0263] The collected data is used to train and optimize the AI model, which is an algorithm designed to continuously improve work efficiency.
[0264] Terminal
[0265] The terminals are installed inside the logistics center and serve as an interface with workers.
[0266] 1. Display Means
[0267] Displays real-time POP content and traffic flow information in the form of images, videos, text, etc.
[0268] 2. Pointing Recognition Method
[0269] The camera is used to recognize the worker's pointing gesture, and detailed information is requested from the server based on that gesture.
[0270] 3. Voice Recognition Method
[0271] The worker's voice is captured using a microphone and converted into text data using voice recognition technology.
[0272] 4. Means of providing information on the display
[0273] Information acquired from the server based on questions received by voice is displayed on the display.
[0274] User
[0275] The users are workers in the logistics center and use the system as follows:
[0276] 1. Viewing content
[0277] Workers can visually check the content displayed on the screen and grasp the information necessary for their work.
[0278] 2. Questions and Pointing
[0279] Workers can ask questions by voice or point to specific locations on the screen to obtain information.
[0280] Specific examples
[0281] Example 1: Picking support
[0282] When picking items in a distribution center, the smart glasses optimize worker movement in real time, suggesting the most efficient route, and visually guiding workers to the exact location of inventory, significantly reducing work time.
[0283] Example prompt sentence:
[0284] "Please build an AI model that can suggest optimal movement paths and display inventory locations in real time to help a man in his 30s improve the efficiency of his picking work."
[0285] Example 2: Improving efficiency of warehousing operations
[0286] When receiving new inventory, the smart glasses will instruct workers in real time on where to place the inventory, supporting efficient stocking operations. Facial recognition is used to provide an individually optimized route and placement location.
[0287] Example prompt sentence:
[0288] "Propose an AI model for smart glasses that uses facial recognition to guide workers to the appropriate shelf location when receiving new inventory."
[0289] The flow of the specific processing in the application example 1 will be described with reference to FIG.
[0290] Step 1:
[0291] The server captures the worker's face through a camera. The input data is the image from the camera, and it analyzes attribute data such as age and gender using a facial recognition algorithm. The output is the analyzed worker's attribute data.
[0292] Step 2:
[0293] The server analyzes the characteristics of the workers based on the attribute data acquired in step 1. The input is customer attribute data, which is stored in a database and updated in real time. The output is the updated characteristic data.
[0294] Step 3:
[0295] The server searches and references inventory information from the product information database. The input is the worker's characteristic data and the inventory database, and the output is information on the optimal inventory location and route. This allows data to be collected to optimize worker movement lines.
[0296] Step 4:
[0297] The server generates the most suitable POP content for each worker in real time. The input is the worker's attribute data and inventory information, and the generated content is sent to the terminal. The output is the dynamically generated POP content.
[0298] Step 5:
[0299] The terminal displays the POP content sent from the server. The input is the POP content sent from the server, and the output is the content displayed on the display.
[0300] Step 6:
[0301] The terminal recognizes the worker's pointing gestures in real time. The input is video of the worker's movements captured by the terminal's built-in camera, and the output is the location information of the recognized pointing gesture. Based on this information, a request for detailed information is made to the server.
[0302] Step 7:
[0303] The terminal captures the worker's voice with a microphone and converts it into text data using voice recognition technology. The input is the worker's voice data, and the output is the converted text data.
[0304] Step 8:
[0305] The server acquires the appropriate information based on the question that has been recognized by voice and sends it to the terminal. The input is the text data that has been recognized by voice, and the output is the acquired information. This provides the necessary information to the worker.
[0306] Step 9:
[0307] The server trains and optimizes the AI model based on the collected data. The input is the collected purchase and operation data, and the output is an updated AI model. This updated model is used to guide future flow optimization and inventory location.
[0308] Step 10:
[0309] The terminal displays information sent from the server on smart glasses or a head-mounted display. The input is inventory location data and movement line data sent from the server, and the output is information displayed on the visual device. This allows workers to work efficiently.
[0310] Furthermore, an emotion engine that estimates the user's emotion may be combined. That is, the identification processing unit 290 may estimate the user's emotion using the emotion identification model 59 and perform identification processing using the user's emotion.
[0311] This invention relates to a digital signage-type POP system that introduces optimal products and services in real time based on customer attribute information and emotional data. This system is used in retail stores and shopping malls to provide personalized information to individual customers.
[0312] System Configuration
[0313] This system is broadly composed of a "server," a "terminal," and a "user." Details of each component are explained below.
[0314] server
[0315] The server has the following functions:
[0316] 1. Facial Recognition Methods:
[0317] The server captures the customer's face through a camera and obtains high-resolution image data.
[0318] Facial recognition algorithms are used to obtain customer attribute data such as age, gender, and number of customers.
[0319] 2. Customer attribute analysis tools and sentiment engine:
[0320] Attribute data analyzed from facial recognition data is stored in a database and updated in real time.
[0321] The emotion engine analyzes customer emotions from captured facial images and tone of voice, and generates emotion data.
[0322] 3. Product information database and personalized message generation:
[0323] The product information database stores detailed information about each product and service, allowing users to search and reference the information they need.
[0324] By referencing customer attribute data and sentiment data, the system selects the most suitable products and services and dynamically generates POP content.
[0325] Add personalized messages from your membership database and purchase history to generate individual messages for specific customers.
[0326] 4. Learning and optimization:
[0327] The server analyzes the collected customer behavioral data, purchasing data, and emotional data to evaluate the effectiveness of POP content.
[0328] Based on the results of the effectiveness evaluation, the AI model will be updated to improve the accuracy of future content generation algorithms.
[0329] Terminal
[0330] The terminal is installed inside the store and serves as an interface with customers.
[0331] 1. Display means:
[0332] Real-time POP content sent from the server is displayed on the display.
[0333] The content can be displayed in the form of images, videos, text, etc.
[0334] 2. Pointing recognition method:
[0335] The camera recognizes the customer's pointing movements in real time.
[0336] A request is made to the server for detailed information about the product pointed to by the customer, and the received information is displayed on the screen.
[0337] 3. Question and Answer Function:
[0338] The customer's voice questions are captured using a microphone and converted into text data using voice recognition technology.
[0339] Information based on the question is obtained from the server and displayed on the screen.
[0340] User
[0341] The user is a customer of the store and performs the following operations.
[0342] 1. Viewing content:
[0343] Users view the POP content displayed on the display and check information about products and services that interest them.
[0344] 2. Questions and Pointing:
[0345] Users can ask for details about products they are interested in by voice or by pointing at specific products on the screen.
[0346] The question and information about the product you point at will be displayed on the screen to help you make a purchasing decision.
[0347] Specific examples
[0348] Example 1: Parent-child visit scenario
[0349] A parent and child visit a shopping mall. The server captures the faces of the parent and child via camera and analyzes that the parent is in their late 30s and the child is around 7 years old. When the emotion engine recognizes the parent's facial expression as being interested and having fun, the device displays information mainly about toys and events for children. When the parent asks, "Where are the toy sales?", the device sends the voice question and analysis results to the server and displays the corresponding information on the display.
[0350] Example 2: Checking details of individual products and emotion recognition
[0351] When a customer points at an image of a product displayed on the display, the device recognizes the pointing gesture and requests detailed information about that product from the server. The server retrieves the product information, and the emotion engine also analyzes the reaction. For example, if the customer shows a surprised expression, information about a special campaign will be displayed. The customer then checks the product's price, specifications, and stock status before making a purchasing decision.
[0352] In this way, combining emotion engines enables more personalized and effective information provision to customers. Using emotion data, it is possible to respond to customers' instantaneous reactions and provide a more personalized shopping experience. This configuration makes it possible to maximize sales promotion effects and improve customer satisfaction.
[0353] The processing flow will be explained below.
[0354] Program processing steps
[0355] Server Processing Steps
[0356] Step 1:
[0357] The server captures the faces of customers in the store through a camera and obtains high-resolution image data.
[0358] Step 2:
[0359] The server analyzes the acquired facial images using an AI algorithm to estimate attribute data such as the customer's age, gender, and number of people.
[0360] Step 3:
[0361] The server updates the analysis results to a database in real time and stores customer attribute data.
[0362] Step 4:
[0363] The server uses an emotion engine to analyze facial images, tone of voice, and other factors to generate customer emotion data.
[0364] Step 5:
[0365] The server stores the generated emotion data in a database and updates it in real time.
[0366] Step 6:
[0367] The server searches and selects the most suitable products and services from a product information database based on customer attribute data and emotional data.
[0368] Step 7:
[0369] The server dynamically generates POP content in real time based on the selection results.
[0370] Step 8:
[0371] The server references the member database and purchase history and generates personalized messages as needed.
[0372] Step 9:
[0373] The server transmits the generated POP content to the terminal.
[0374] Step 10:
[0375] The server analyzes the collected customer behavioral data, purchasing data, and emotional data to evaluate the effectiveness of the POP content.
[0376] Step 11:
[0377] Based on the effectiveness evaluation results, the server updates the AI model and improves the accuracy of future content generation algorithms.
[0378] Terminal processing steps
[0379] Step 1:
[0380] The terminal receives the POP content sent from the server.
[0381] Step 2:
[0382] The device displays the received content on the display, which can be in the form of images, videos, text, etc.
[0383] Step 3:
[0384] The device uses a camera to recognize the customer's pointing movements in real time.
[0385] Step 4:
[0386] The terminal requests information about the product pointed at by the customer from the server.
[0387] Step 5:
[0388] The terminal receives the detailed product information sent from the server and displays it on the display.
[0389] Step 6:
[0390] The terminal uses a microphone to capture the customer's voice asking a question.
[0391] Step 7:
[0392] The device uses voice recognition technology to convert the voice into text data and send it to the server.
[0393] Step 8:
[0394] The terminal displays the response information sent from the server on the display.
[0395] User operation steps
[0396] Step 1:
[0397] The user views the POP content displayed on the display.
[0398] Step 2:
[0399] Users can input questions about products or services that interest them by voice.
[0400] Step 3:
[0401] The user points to a particular product on the display to see more information about it.
[0402] Step 4:
[0403] Users make purchasing decisions based on the displayed information.
[0404] Specific examples
[0405] Example 1: Parent-child visit scenario
[0406] Step 1: A parent and child visit a shopping mall. The server captures the faces of the parent and child through a camera and analyzes them.
[0407] Step 2: The server estimates the age and gender of the parents and children and stores the attribute data in a database.
[0408] Step 3: The server analyzes the facial image and tone of voice using an emotion engine to generate emotion data.
[0409] Step 4: The server selects toys and event information for children based on the generated attribute data and emotion data, and generates POP content.
[0410] Step 5: The terminal receives the POP content and displays it on the display.
[0411] Step 6: The parent asks, "Where are the toy sales?" The device captures the audio and sends it to the server.
[0412] Step 7: The server generates toy department information based on the query and sends it to the terminal.
[0413] Step 8: The terminal displays the received information on its display.
[0414] Example 2: Checking details of individual products and emotion recognition
[0415] Step 1: The customer points to the product image on the display.
[0416] Step 2: The device recognizes the pointing gesture and requests detailed information about the product from the server.
[0417] Step 3: The server retrieves the product details and analyzes the customer's emotions using the emotion engine.
[0418] Step 4: The server adds special campaign information based on the product details and customer sentiment data and sends it to the terminal.
[0419] Step 5: The terminal displays the received information on the screen. The customer checks the product price, specifications, and stock information and makes a purchasing decision.
[0420] In this way, combining emotion engines enables effective provision of personalized information to customers. Using emotion data, it is possible to respond to customers' instantaneous reactions and provide a more personalized shopping experience. This configuration makes it possible to maximize sales promotion effects and improve customer satisfaction.
[0421] Example 2
[0422] Next, a description will be given of Example 2. In the following description, the data processing device 12 will be referred to as a "server" and the smart device 14 will be referred to as a "terminal."
[0423] Conventional digital signage and POP systems have had issues with providing personalized information to customers and not being able to dynamically recommend products based on customer emotions. As a result, it has been difficult to provide effective information to customers and increase their desire to purchase. It has also been difficult to optimize generative models using customer behavioral and emotional data.
[0424] The identification process by the identification processing unit 290 of the data processing device 12 in the second embodiment is realized by the following means. In this invention, the server includes a face recognition means, a customer attribute analysis means, and a means for acquiring emotion data. This makes it possible to recommend optimal products and services in real time based on customer attribute information and emotion data. Furthermore, by evaluating the effectiveness of the generated POP content and updating the generation model based on collected data, it becomes possible to continuously improve the accuracy of information provision.
[0425] A "facial recognition means" is an algorithm or device that uses a camera to capture facial images of customers and extract attribute data such as age, gender, and number of people.
[0426] "Customer attribute analysis means" refers to a system or program that analyzes customer attributes based on acquired facial recognition data and stores and updates that data in a database.
[0427] The "product information database" is a data storage system that stores detailed information about products and services in the store and allows for searching and referencing as needed.
[0428] The "means for acquiring emotional data" refers to an algorithm or device that analyzes the customer's facial image and tone of voice to generate emotional data such as joy, interest, or displeasure.
[0429] "Means for generating POP content in real time" refers to a system or program that dynamically generates content that recommends products and services based on customer attribute information and emotional data.
[0430] The "display means" is a display device that visually presents the POP content transmitted from the server to customers in the store.
[0431] The "means for recognizing the customer's pointing behavior" refers to a system or program that uses a sensor such as a camera to detect the customer's pointing behavior in real time and analyzes that information.
[0432] The "means for recognizing customer questions by voice" is a system or program that uses a microphone to capture the voice of a customer's question and converts it into text data using voice recognition technology.
[0433] The "means for providing information based on a voice question" is a system or program that retrieves related information from a server based on a customer's voice question and presents it on a display.
[0434] "Means for personalizing information" refers to a system or program that generates and displays personalized messages optimized for specific customers based on customer attribute data and purchase history.
[0435] A "means for collecting purchasing data" is a system or device for collecting data on customer purchasing behavior and storing it in a database.
[0436] The "means for updating the generative model based on collected data" refers to a system or algorithm that analyzes collected customer behavioral and emotional data and optimizes and updates the generative model.
[0437] MODE FOR CARRYING OUT THE INVENTION
[0438] This invention relates to a digital signage-type POP system that introduces optimal products and services in real time based on customer attribute information and emotional data. By recognizing the customer's face, analyzing their emotions, and recommending products and services based on that, personalized information is provided.
[0439] server
[0440] The server has the following features:
[0441] 1. Facial Recognition Methods:
[0442] The server captures customer faces through a camera, obtains high-resolution image data, and uses a facial recognition algorithm to extract customer attributes such as age, gender, and number of customers.
[0443] 2. Customer attribute analysis means:
[0444] Based on facial recognition data, customer attribute information is analyzed, and the data is stored in a database and updated in real time.
[0445] 3. How to get emotion data:
[0446] The emotion engine analyzes the customer's emotions using facial images, tone of voice, and other data, and generates emotion data. For example, if a customer smiles, the emotion is analyzed as "happiness."
[0447] 4. Product Information Database:
[0448] The product information database stores detailed information about each product and service, which can be searched and referenced as needed. Content is generated that recommends optimal products and services based on customer attribute data and emotional data.
[0449] 5. Personalized message generation:
[0450] Add individual personalized messages from your membership database and purchase history to generate individual messages for specific customers.
[0451] 6. Learning and optimization:
[0452] The server analyzes collected customer behavioral, purchasing, and emotional data to evaluate the effectiveness of the POP content. Based on the results of the effectiveness evaluation, the generative AI model is updated, improving the accuracy of future content generation algorithms.
[0453] Terminal
[0454] The terminals are installed in the store and act as the interface with the customer:
[0455] 1. Display means:
[0456] Real-time POP content sent from the server is displayed on the display. Content can be displayed in the form of images, videos, text, etc.
[0457] 2. Pointing recognition method:
[0458] The device's camera recognizes the customer's pointing gestures in real time, requests detailed information about the product the customer is pointing at from the server, and displays the received information on the display.
[0459] 3. Question and Answer Function:
[0460] The device's microphone captures the customer's voice and converts it into text data using speech recognition technology. Information based on the question is retrieved from the server and displayed on the screen.
[0461] User
[0462] The user is a customer in the store and performs the following actions:
[0463] 1. Viewing content:
[0464] Users view the POP content displayed on the display and check information about products and services that interest them.
[0465] 2. Questions and Pointing:
[0466] Users can ask for details about products they are interested in by voice or by pointing at a specific product on the screen. The content of the question or information about the product they are pointing at is then displayed on the screen.
[0467] Specific examples
[0468] Example 1: Parent-child visit scenario
[0469] A parent and child visit a shopping mall. The server captures the faces of the parent and child via camera and analyzes that the parent is in their late 30s and the child is about 7 years old. When the emotion engine recognizes the parent's facial expression as being interested and having fun, the device displays information mainly about toys and events for children. When the parent asks, "Where are the toy sales?", the device sends the voice question and analysis results to the server and displays the corresponding information on the display.
[0470] Example 2: Checking details of individual products and emotion recognition
[0471] When a customer points at an image of a product displayed on the display, the device recognizes the pointing gesture and requests detailed information about that product from the server. The server retrieves the product information, and the emotion engine also analyzes the reaction. For example, if the customer shows a surprised expression, information about a special campaign will be displayed. The customer then checks the product's price, specifications, and stock status before making a purchasing decision.
[0472] Examples of prompt statements
[0473] 1. "When a customer points to an image of a product on the display, explain how the device displays more information about that product."
[0474] 2. "Please explain specifically how the system will display toys and information about children's events when a parent and child visit the store."
[0475] This invention makes it possible to provide customers with personalized information, which is expected to increase their purchasing motivation. It also responds to customers' instantaneous reactions, providing a more personalized shopping experience.
[0476] The flow of the identification process in the second embodiment will be described with reference to FIG.
[0477] Step 1:
[0478] The server captures the faces of customers in the store using a camera and obtains high-resolution image data. The obtained image data is input into a facial recognition algorithm to extract customer attribute data such as age, gender, and number of customers. The output attribute data is stored in a database.
[0479] Specific behavior:
[0480] When a customer enters a store, the server automatically acquires data from the camera and begins facial recognition. For example, if a woman in her 30s and a man in his 40s visit the store together, their attribute data will be stored in a database.
[0481] Step 2:
[0482] The server inputs the acquired facial image and tone of voice into an emotion engine to generate customer emotion data. The emotion engine analyzes the customer's emotional state, such as joy, interest, or displeasure, and outputs the data. The analyzed emotion data is stored in a database.
[0483] Specific behavior:
[0484] If a customer looks at the display and smiles, the server interprets the smile as "happiness." If the customer's voice is high-pitched and excited, the emotion engine recognizes it as "interested."
[0485] Step 3:
[0486] The server references the product information database and selects the most suitable products and services based on the acquired customer attribute data and emotion data. The selected information is generated as dynamic POP content in real time. The generated POP content is then saved back into the database.
[0487] Specific behavior:
[0488] For example, if a woman in her 30s expresses interest, the server will generate content recommending the latest beauty products, adding special campaign information for related products based on the customer's past purchases.
[0489] Step 4:
[0490] The terminal displays the POP content sent from the server. The content is displayed on the display in the form of images, videos, and text. When a customer points at a specific product on the screen, the camera recognizes the pointing gesture and requests detailed information from the server.
[0491] Specific behavior:
[0492] When a customer points to an image of a toy on the display, the device automatically retrieves and displays detailed information about the product.
[0493] Step 5:
[0494] When a customer asks a question to the display, the device's microphone captures the voice of the question and converts it into text data using voice recognition technology. The converted text data is sent to the server, which returns information based on the question. The acquired information is then displayed on the display.
[0495] Specific behavior:
[0496] When a parent asks, "Where is the toy sale?", the voice question and analysis results are sent to the server, and the corresponding information is displayed on the screen.
[0497] Step 6:
[0498] The server analyzes collected customer behavioral, purchasing, and emotional data to evaluate the effectiveness of POP content. The results of the evaluation are used to update the AI model, improving the accuracy of future content generation algorithms.
[0499] Specific behavior:
[0500] Past data can be used to analyze which content was most effective. For example, data on women in their 30s can be used to learn that a particular beauty product was well-received, and this can be reflected in future recommendations.
[0501] (Application example 2)
[0502] Next, a description will be given of Application Example 2. In the following description, the data processing device 12 will be referred to as a "server" and the smart device 14 will be referred to as a "terminal."
[0503] In today's retail industry, there is a demand for personalized information provision to each individual customer. Furthermore, there is a lack of methods for using emotional data to suggest optimal products and services based on the customer's real-time situation. As a result, customer satisfaction and purchasing motivation have not been sufficiently improved. Therefore, there is a need to provide a system that can recommend optimal products and services in real time based on the customer's attribute information and emotional data.
[0504] The specific processing by the specific processing unit 290 of the data processing device 12 in the application example 2 is realized by the following means.
[0505] In this invention, the server includes a face recognition means, a customer attribute analysis means, a product information database, a means for generating POP content in real time, a display device for displaying the generated POP content, a means for recognizing customer pointing gestures, a means for recognizing customer questions by voice, a means for providing information based on the voice questions, a means for personalizing information, a means for collecting purchase data, a means for updating an AI model based on the collected data, a sentiment analysis means, a means for selecting optimal products and services based on the sentiment data, a means for detecting customer faces in real time using a camera and analyzing their sentiments, and a means for displaying related product information on a display device based on the analysis results. This makes it possible to propose optimal products and services based on the customer's sentiments and attributes in real time, providing a personalized shopping experience and improving customer satisfaction and purchasing motivation.
[0506] A "face recognition means" is a means of capturing a customer's face using a camera and obtaining attribute data such as the customer's age, gender, and number of people from the high-resolution image data.
[0507] The "customer attribute analysis means" is a means for analyzing attribute data acquired by the face recognition means and updating it in real time.
[0508] A "product information database" is a database that stores detailed information about each product or service, allowing users to search and reference the information they need.
[0509] "Means for generating POP content in real time" refers to means for dynamically generating POP content by referencing customer attribute data and emotional data, selecting the most suitable products and services.
[0510] The "display device" is a device for displaying real-time POP content transmitted from the server on a display.
[0511] The "means for recognizing pointing actions" is a means for recognizing a customer's pointing actions in real time through a camera and requesting detailed information about the pointed product from the server.
[0512] The "voice recognition means" is a means of capturing the customer's voice question using a microphone and converting it into text data using voice recognition technology.
[0513] The "means for providing information based on a voice question" is a means for obtaining related information from a server based on the content of a customer's voice question and displaying it on a display.
[0514] The "means for personalizing information" refers to a means for adding personalized messages from a member database or purchase history and generating individual messages for specific customers.
[0515] "Means for collecting purchasing data" refers to the means for collecting and analyzing customer behavioral data, purchasing data, and emotional data.
[0516] "Means for updating the AI model" refers to a means for evaluating the effectiveness of POP content based on collected data and updating the AI model based on the results of that evaluation.
[0517] The "emotion analysis means" is a means for analyzing the customer's emotions from acquired facial images, tone of voice, etc., and generating emotion data.
[0518] "Means for selecting optimal products and services based on emotional data" refers to means for selecting the most appropriate products and services at that time by referring to emotional data.
[0519] "Means for detecting a customer's face in real time using a camera and analyzing emotions" refers to means for capturing a customer's face in real time using a camera and analyzing the customer's emotions from the facial image.
[0520] The "means for displaying related product information on a display device based on the analysis results" refers to means for displaying related product information on a display based on the analyzed customer emotion data and attribute data.
[0521] The present invention relates to a system that introduces optimal products and services in real time based on customer attribute information and emotion data, and is configured as follows.
[0522] server
[0523] The server has the following functions:
[0524] 1. Facial Recognition Methods
[0525] The server captures the customer's face through a camera, obtains high-resolution image data, and uses a facial recognition algorithm to obtain customer attribute data such as age, gender, and number of customers.
[0526] 2. Customer attribute analysis means
[0527] Attribute data acquired through facial recognition is analyzed and updated in real time, ensuring that basic customer information is always kept up to date.
[0528] 3. Product information database
[0529] The product information database stores detailed information about each product and service, allowing you to search and reference the information you need.
[0530] 4. A way to generate POP content in real time
[0531] By referencing customer attribute data and emotional data, the system selects the most suitable products and services and dynamically generates POP content. For example, it uses prompts such as "display the most suitable products for male customers in their 30s who have a surprised expression."
[0532] 5. Display device means
[0533] Real-time POP content sent from the server is displayed on the display. Content can be displayed in the form of images, videos, text, etc.
[0534] 6. Methods for Recognizing Pointing Actions
[0535] The camera recognizes the customer's pointing gestures in real time, requests detailed information about the product the customer is pointing at from the server, and displays the received information on the display.
[0536] 7. Voice recognition
[0537] The customer's voice questions are captured using a microphone and converted into text data using voice recognition technology.
[0538] 8. Means of Providing Information Based on Voice Queries
[0539] Based on the content of the customer's voice question, related information is obtained from the server and displayed on the screen.
[0540] 9. How we personalize your information
[0541] Add personalized messages from your membership database and purchase history to generate individual messages for specific customers.
[0542] 10. Means of collecting purchasing data
[0543] Collect and analyze customer behavioral, purchasing, and emotional data.
[0544] 11. How to update your AI model
[0545] The effectiveness of POP content is evaluated based on the collected data, and the AI model is updated based on the evaluation results.
[0546] 12. Sentiment analysis tool
[0547] The system analyzes customer emotions from acquired facial images and tone of voice, and generates emotional data.
[0548] 13. A way to select the best products and services based on sentiment data
[0549] It is a means of referring to emotional data to select the most appropriate product or service at any given time.
[0550] 14. Real-time facial detection and emotion analysis using cameras
[0551] The camera captures the customer's face in real time and analyzes their emotions from the facial image.
[0552] 15. Means for displaying related product information on a display device based on the analysis results
[0553] Based on the analyzed customer emotional data and attribute data, relevant product information is displayed on the screen.
[0554] Terminal
[0555] The terminal is installed inside the store and serves as an interface with customers.
[0556] Display Means
[0557] Real-time POP content sent from the server is displayed on the display, allowing customers to directly view information about products and services that are optimized for them.
[0558] Pointing recognition method
[0559] The camera recognizes customers' pointing gestures in real time. For example, if a customer points to a product to say, "I'd like to know more about this product," detailed information will be displayed on the screen.
[0560] Voice question answering function
[0561] The microphone is used to capture the customer's voice asking a question. If a customer asks, "Are there any special offers for this product?", the voice is sent to the server and the appropriate information is displayed.
[0562] User
[0563] The user is a customer of the store.
[0564] Viewing content
[0565] Users view the POP content displayed on the display and check information about products and services that interest them.
[0566] Questions and pointing
[0567] Users can ask for details about products they are interested in by voice or by pointing at a specific product on the screen. The question or information about the product they are pointing at is then displayed on the screen to help them make a purchasing decision.
[0568] Specific examples
[0569] By inputting prompts such as, "Please tell me what product information should be provided to a female customer who looks happy," into the generative AI model, the optimal content is displayed.
[0570] In this way, the system of the present invention improves customer satisfaction and maximizes the effectiveness of store sales promotion by providing real-time information based on customer attribute information and emotional data.
[0571] The flow of the specific processing in the application example 2 will be described with reference to FIG.
[0572] Step 1:
[0573] The server captures the customer's face through a camera and obtains high-resolution image data.
[0574] Input: Video data from the camera
[0575] Output: Customer's facial image data
[0576] Specific operation: The camera captures video and sends the video data to the server, which then detects faces in the video and generates facial image data.
[0577] Step 2:
[0578] The server uses a facial recognition algorithm to obtain attribute data such as customer age, gender, and number of customers.
[0579] Input: Customer's facial image data
[0580] Output: Customer attribute data (age, gender, number of people)
[0581] Specific operation: The server runs a facial recognition algorithm, analyzes the facial image data, and extracts attribute information such as age, gender, and number of people.
[0582] Step 3:
[0583] The server stores the acquired attribute data in a database and updates it in real time.
[0584] Input: Customer attribute data
[0585] Output: Updated customer attribute information in the database
[0586] Specific operation: The server stores the attribute data in a database and updates existing customer information as necessary.
[0587] Step 4:
[0588] The server uses an emotion engine to analyze the customer's emotions from acquired facial images and tone of voice, and generates emotion data.
[0589] Input: Facial image data, audio data
[0590] Output: Customer sentiment data
[0591] Specific operation: The emotion engine analyzes facial images and voice data, infers emotions from the customer's facial expressions and tone of voice, and generates that data.
[0592] Step 5:
[0593] The server references customer attribute data and emotion data, selects the most suitable products and services from a product information database, and dynamically generates POP content.
[0594] Input: Customer attribute data, customer sentiment data
[0595] Output: Generated POP content
[0596] Specific operation: The server searches the product information database based on customer attribute information and emotion data, extracts information on related products and services, and generates POP content based on this.
[0597] Step 6:
[0598] The terminal displays the real-time POP content sent from the server on the display.
[0599] Input: Generated POP content
[0600] Output: The content shown on the display
[0601] Specific operation: Receives POP content sent from the server and displays information such as images, videos, and text on the display.
[0602] Step 7:
[0603] The device uses a camera to recognize the customer's pointing movements in real time.
[0604] Input: Video data from the camera
[0605] Output: Pointing motion detection information
[0606] Specific operation: The camera captures the customer's pointing action and sends it to the server for pointing recognition.
[0607] Step 8:
[0608] The server retrieves detailed information about the product pointed to by the customer from the database and displays it on the display.
[0609] Input: Pointing motion detection information
[0610] Output: Product details
[0611] Specific operation: Based on the pointing gesture, the server retrieves the corresponding product information from the database and sends it to the terminal, which then displays the detailed information on the screen.
[0612] Step 9:
[0613] The terminal uses a microphone to capture the customer's voice asking a question and converts it into text data using voice recognition technology.
[0614] Input: Audio data
[0615] Output: Text data of the audio
[0616] How it works: A microphone captures the customer's question, and a speech recognition algorithm on the server converts it into text data.
[0617] Step 10:
[0618] The server acquires information based on the voice question content and displays it on a display.
[0619] Input: Text data of speech
[0620] Output: Information corresponding to the question
[0621] Specific operation: The server analyzes the text data of the voice question, retrieves relevant information from the database, and sends it to the device, which then displays the answer to the question on the display.
[0622] Step 11:
[0623] The server generates personalized messages from a member database and purchase history, and provides individual messages to specific customers.
[0624] Input: Member data, purchase history
[0625] Output: Personalized message
[0626] Specific operation: The server references the member database and purchase history, generates a personalized message tailored to the specific customer, and provides it via the terminal.
[0627] Step 12:
[0628] The server collects and analyzes customer behavioral data, purchasing data, and emotional data.
[0629] Input: Customer behavior data, purchasing data, emotional data
[0630] Output: Analysis data
[0631] Specific operation: The server analyzes the collected data in real time and uses the results to generate the next POP content and respond to customers.
[0632] Step 13:
[0633] The server evaluates the effectiveness of the POP content based on the collected data and updates the AI model based on the evaluation results.
[0634] Input: Customer behavior data, purchasing data, emotional data
[0635] Output: Updated AI model
[0636] Specific operation: The server analyzes the collected data and adjusts the parameters of the AI model to improve the accuracy of subsequent content generation algorithms.
[0637] The specific processing unit 290 transmits the result of the specific processing to the smart device 14. In the smart device 14, the control unit 46A causes the output device 40 to output the result of the specific processing. The microphone 38B acquires audio indicating a user input regarding the result of the specific processing. The control unit 46A transmits audio data indicating the user input acquired by the microphone 38B to the data processing device 12. In the data processing device 12, the specific processing unit 290 acquires the audio data.
[0638] The data generation model 58 is a so-called generative AI (Artificial Intelligence). An example of the data generation model 58 is ChatGPT (Internet Search<URL: https: / / openai.com / blog / chatgpt> ), Gemini (Internet search <url: https: gemini.google.com ?hl="ja">) and other generation AIs. The data generation model 58 is obtained by performing deep learning on a neural network. A prompt including an instruction is input to the data generation model 58, and inference data such as voice data indicating voice, text data indicating text, and image data indicating an image is also input. The data generation model 58 performs inference on the input inference data in accordance with the instruction indicated by the prompt, and outputs the inference result in a data format such as voice data and text data. Here, inference refers to, for example, analysis, classification, prediction, and / or summarization.
[0639] In the above embodiment, an example in which the specific process is performed by the data processing device 12 has been given, but the technology of the present disclosure is not limited to this, and the specific process may be performed by the smart device 14.
[0640] [Second embodiment]
[0641] FIG. 3 shows an example of the configuration of a data processing system 210 according to the second embodiment.
[0642] 3, the data processing system 210 includes the data processing device 12 and smart glasses 214. An example of the data processing device 12 is a server.
[0643] The data processing device 12 includes a computer 22, a database 24, and a communication I / F 26. The computer 22 is an example of a "computer" according to the technology of the present disclosure. The computer 22 includes a processor 28, a RAM 30, and a storage 32. The processor 28, the RAM 30, and the storage 32 are connected to a bus 34. The database 24 and the communication I / F 26 are also connected to the bus 34. The communication I / F 26 is connected to a network 54. Examples of the network 54 include a WAN (Wide Area Network) and / or a LAN (Local Area Network).
[0644] The smart glasses 214 include a computer 36, a microphone 238, a speaker 240, a camera 42, and a communication I / F 44. The computer 36 includes a processor 46, a RAM 48, and a storage 50. The processor 46, the RAM 48, and the storage 50 are connected to a bus 52. The microphone 238, the speaker 240, and the camera 42 are also connected to the bus 52.
[0645] The microphone 238 receives instructions and the like from the user 20 by receiving voice uttered by the user 20. The microphone 238 captures the voice uttered by the user 20, converts the captured voice into audio data, and outputs it to the processor 46. The speaker 240 outputs audio in accordance with instructions from the processor 46.
[0646] Camera 42 is a small digital camera equipped with an optical system including a lens, aperture, and shutter, and an imaging element such as a CMOS (Complementary Metal-Oxide-Semiconductor) image sensor or a CCD (Charge Coupled Device) image sensor, and captures images of the surroundings of user 20 (for example, an imaging range defined by an angle of view equivalent to the field of vision of a typical healthy person).
[0647] The communication I / F 44 is connected to a network 54. The communication I / Fs 44 and 26 are responsible for the exchange of various information between the processor 46 and the processor 28 via the network 54. The exchange of various information between the processor 46 and the processor 28 using the communication I / Fs 44 and 26 is carried out in a secure state.
[0648] Fig. 4 shows an example of the main functions of the data processing device 12 and the smart glasses 214. As shown in Fig. 4, in the data processing device 12, a specific process is performed by the processor 28. A specific process program 56 is stored in the storage 32.
[0649] The specific processing program 56 is an example of a "program" according to the technology of the present disclosure. The processor 28 reads the specific processing program 56 from the storage 32 and executes the read specific processing program 56 on the RAM 30. The specific processing is realized by the processor 28 operating as a specific processing unit 290 in accordance with the specific processing program 56 executed on the RAM 30.
[0650] The storage 32 stores a data generation model 58 and an emotion identification model 59. The data generation model 58 and the emotion identification model 59 are used by the identification processing unit 290.
[0651] In the smart glasses 214, the reception output process is performed by the processor 46. A reception output program 60 is stored in the storage 50. The processor 46 reads the reception output program 60 from the storage 50 and executes the read reception output program 60 on the RAM 48. The reception output process is realized by the processor 46 operating as the control unit 46A in accordance with the reception output program 60 executed on the RAM 48.
[0652] Next, a description will be given of the identification process performed by the identification processing unit 290 of the data processing device 12. In the following description, the data processing device 12 will be referred to as the "server" and the smart glasses 214 will be referred to as the "terminal."
[0653] This invention relates to a digital signage-type POP system that introduces optimal products and services in real time based on customer attribute information. This system is used in retail stores and shopping malls to provide personalized information to individual customers.
[0654] System Configuration
[0655] This system is broadly composed of a "server," a "terminal," and a "user." Details of each component are explained below.
[0656] server
[0657] The server has the following functions:
[0658] 1. Facial Recognition Methods:
[0659] The server captures customers' faces through a camera and obtains attribute data such as age, gender, and number of people.
[0660] Using high-resolution video data, customer attributes are analyzed using facial recognition algorithms.
[0661] 2. Customer attribute analysis means:
[0662] Based on facial recognition, analyzed attribute data is stored in a database and updated in real time.
[0663] The analysis results are used to provide promotional information to each customer.
[0664] 3. Product information database and personalized message generation:
[0665] The product information database stores detailed information about each product and service, allowing users to search and reference the information they need.
[0666] Customer attribute data is referenced to select the most suitable products and services and dynamically generate POP content.
[0667] Add personalized messages from your membership database and purchase history to generate individual messages for specific customers.
[0668] 4. Learning and optimization:
[0669] The server analyzes collected customer behavioral and purchasing data and updates the AI model to optimize the effectiveness of POP content.
[0670] We continuously collect data and improve our algorithms to generate optimal content.
[0671] Terminal
[0672] The terminal is installed inside the store and serves as an interface with customers.
[0673] 1. Display means:
[0674] Real-time POP content sent from the server is displayed on the display.
[0675] The content can be displayed in the form of images, videos, text, etc.
[0676] 2. Pointing recognition method:
[0677] The camera recognizes the customer's pointing movements in real time.
[0678] A request is made to the server for detailed information about the product pointed to by the customer, and the received information is displayed on the screen.
[0679] 3. Question and Answer Function:
[0680] The customer's voice questions are captured using a microphone and converted into text data using voice recognition technology.
[0681] Information based on the question is obtained from the server and displayed on the screen.
[0682] User
[0683] The user is a customer of the store and performs the following operations.
[0684] 1. Viewing content:
[0685] Users view the POP content displayed on the display and check information about products and services that interest them.
[0686] 2. Questions and Pointing:
[0687] Users can ask for details about products they are interested in by voice or by pointing at specific products on the screen.
[0688] The question and information about the product you point at will be displayed on the screen to help you make a purchasing decision.
[0689] Specific examples
[0690] Example 1: Parent and child visitor scenario
[0691] A parent and child visit a shopping mall. The server captures the faces of the parent and child via camera and determines that the parent is in their late 30s and the child is about 7 years old. Based on this, the device displays information mainly about toys and events for children. When the parent asks, "Where are the toy sales?", the device recognizes the voice question, retrieves the answer from the server, and displays it on the display.
[0692] Example 2: Checking details of individual products
[0693] The customer points at the image of a product displayed on the display. The terminal recognizes the pointing gesture and retrieves detailed information about that product from the server. The product's price, specifications, stock status, and other information are displayed on the display in real time.
[0694] As described above, this system can provide optimal information in real time based on customer attributes and behavior, thereby enhancing sales promotion effectiveness. Furthermore, by continually updating the AI model based on collected data, it is possible to continuously generate effective content.
[0695] The processing flow will be explained below.
[0696] Program processing steps
[0697] Server Processing Steps
[0698] Step 1:
[0699] The server captures the faces of customers in the store through a camera and obtains high-resolution image data.
[0700] Step 2:
[0701] The server analyzes the acquired facial images using an AI algorithm to estimate attribute data such as the customer's age, gender, and number of people.
[0702] Step 3:
[0703] The server updates the analysis results to a database and stores customer attribute data in real time.
[0704] Step 4:
[0705] The server searches for the most suitable products and services from a product information database based on customer attribute data.
[0706] Step 5:
[0707] The server dynamically generates POP content in real time based on the search results.
[0708] Step 6:
[0709] The server references the member database and purchase history and generates personalized messages as needed.
[0710] Step 7:
[0711] The server transmits the generated POP content to the terminal.
[0712] Step 8:
[0713] The server analyzes the collected customer behavior data and purchase data to evaluate the effectiveness of the POP content.
[0714] Step 9:
[0715] Based on the effectiveness evaluation results, the server updates the AI model and improves the accuracy of future content generation algorithms.
[0716] Terminal processing steps
[0717] Step 1:
[0718] The terminal receives the POP content sent from the server.
[0719] Step 2:
[0720] The device displays the received content on the display, which can be in the form of images, videos, text, etc.
[0721] Step 3:
[0722] The device uses a camera to recognize the customer's pointing movements in real time.
[0723] Step 4:
[0724] The terminal requests information about the product pointed at by the customer from the server.
[0725] Step 5:
[0726] The terminal receives the detailed product information sent from the server and displays it on the display.
[0727] Step 6:
[0728] The terminal uses a microphone to capture the customer's voice asking a question.
[0729] Step 7:
[0730] The device uses voice recognition technology to convert the voice into text data and send it to the server.
[0731] Step 8:
[0732] The terminal displays the response information sent from the server on the display.
[0733] User operation steps
[0734] Step 1:
[0735] The user views the POP content displayed on the display.
[0736] Step 2:
[0737] Users can input questions about products or services that interest them by voice.
[0738] Step 3:
[0739] The user points to a particular product on the display to see more information about it.
[0740] Step 4:
[0741] Users make purchasing decisions based on the displayed information.
[0742] Specific examples
[0743] Example 1: Parent-child visit scenario
[0744] Step 1: A parent and child visit a shopping mall. The server captures the faces of the parent and child through a camera and analyzes them.
[0745] Step 2: The server estimates the age and gender of the parents and children and updates the attribute data in the database.
[0746] Step 3: The server selects toys and event information for children based on the parent-child attributes and generates POP content.
[0747] Step 4: The terminal receives the POP content and displays it on the display.
[0748] Step 5: The parent asks, "Where are the toy sales?" The device captures the audio and sends it to the server.
[0749] Step 6: The server generates toy department information based on the query and sends it to the terminal.
[0750] Step 7: The terminal displays the received information on its display.
[0751] Example 2: Checking details of individual products
[0752] Step 1: The customer points to the product image on the display.
[0753] Step 2: The device recognizes the pointing gesture and requests detailed information about the product from the server.
[0754] Step 3: The server obtains the product details and sends them to the terminal.
[0755] Step 4: The terminal displays the received information on its display.
[0756] Step 5: The customer checks the product price, specifications, and stock information and makes a purchasing decision.
[0757] The above is the specific flow of operations at each step. This invention realizes the provision of personalized information in real time based on the attributes and behavior of customers.
[0758] Example 1
[0759] Next, a description will be given of Example 1. In the following description, the data processing device 12 will be referred to as a "server" and the smart glasses 214 will be referred to as a "terminal."
[0760] Conventional digital signage systems have the problem that the information provided to customers is uniform and not personalized enough for each individual customer. As a result, it is not possible to introduce optimal products and services that meet each customer's interests and needs, and sales promotion effectiveness is limited. Furthermore, there is a lack of a mechanism to effectively collect and analyze customer behavior data and optimize content in real time.
[0761] The specific processing by the specific processing unit 290 of the data processing device 12 in the first embodiment is realized by the following means.
[0762] In this invention, the server includes a means for capturing customer faces and acquiring attribute data, a means for analyzing and saving the acquired attribute data, and a database for accumulating product information. This makes it possible to select and display optimal products and services in real time according to customer attributes. Furthermore, by collecting purchasing data and updating and optimizing the AI model based on this, it is possible to consistently provide effective content.
[0763] "Means for capturing customer faces and acquiring attribute data" refers to technology that uses cameras and sensors to capture images of customers' faces in a store and extract information such as age, gender, and number of customers from that image data.
[0764] "Means for analyzing and storing acquired attribute data" refers to hardware and software for analyzing attribute data obtained by facial recognition technology and storing it in a database.
[0765] A "database that stores product information" is a database system that manages detailed information about each product and service sold in a retail store or shopping mall, and allows for searching and referencing.
[0766] "Means for generating personalized POP content in real time" refers to algorithms and software for dynamically creating and displaying promotional information that is optimal for each customer based on their attribute data.
[0767] The "display means for displaying the generated content" refers to a display device for visually presenting the POP content sent from the server to customers in the form of images, videos, text, etc.
[0768] "Means for recognizing customer pointing movements" refers to technology that uses sensors such as cameras to detect customer pointing movements and analyze those movements.
[0769] "Means for recognizing customer questions via voice" refers to voice recognition technology that uses a microphone to capture the customer's voice and converts the voice data into text.
[0770] "Means for providing information based on voice questions" refers to a system that searches for appropriate information based on the question content obtained by voice recognition and provides that information.
[0771] "Means for personalizing information" refers to algorithms and systems for providing the most appropriate information to individual customers based on their attribute data and purchasing history.
[0772] "Means for collecting purchasing data" refers to a system for collecting customer purchasing history and behavioral data, and storing and managing this in a database.
[0773] "Means for updating AI models based on collected data and optimizing POP content" refers to a system that uses collected customer behavioral and purchasing data to continuously learn and update AI models to generate optimal POP content.
[0774] This invention relates to a digital signage-type POP system that introduces optimal products and services in real time based on customer attribute information. This system is used in retail stores and shopping malls to provide personalized information to individual customers.
[0775] System Configuration
[0776] This system is broadly composed of a "server," a "terminal," and a "user." Each component will be explained in detail below.
[0777] server
[0778] The server has the following functions:
[0779] 1. A means of capturing customer faces and obtaining attribute data
[0780] The server monitors the store in real time through installed cameras, capturing images of customers' faces and acquiring attribute data such as age, gender, number of people, etc. This facial recognition is performed using Python and OpenCV, applying facial recognition algorithms such as YOLO.
[0781] 2. Means for analyzing and storing acquired attribute data
[0782] The server analyzes the acquired customer attribute data and saves it in a database. The attribute data is continuously updated and to manage the information of multiple customers in real time, a database management system such as MySQL or PostgreSQL is used. The data is managed via an API using the Django or Flask framework.
[0783] 3. Database for storing product information
[0784] The product information database stores detailed information about the products it handles, allowing for quick search and reference as needed. This database includes information such as product prices, specifications, and stock status.
[0785] 4. A way to generate personalized POP content in real time
[0786] The server selects the most suitable products and services based on customer attribute data and dynamically generates personalized POP content using a generative AI model (e.g., GPT-3). The generated POP content is converted into HTML or image format as a text message or a list of recommended products and sent to the device.
[0787] 5. Display means for displaying generated content
[0788] The terminal visually displays real-time POP content sent from the server on a display, which can be powered by a Raspberry Pi or Windows PC and runs as a web application in a browser.
[0789] Terminal
[0790] The terminal is installed inside the store and serves as an interface with customers.
[0791] 1. A method for recognizing customer pointing gestures
[0792] The device uses a camera to recognize the customer's pointing gestures in real time, using hand movement recognition algorithms such as TensorFlow and OpenPose, and sends the processing results to a server using a Python script.
[0793] 2. A way to recognize customer questions by voice
[0794] The device uses a microphone to capture the customer's voice and converts it into text using speech recognition technology, using the Google Cloud Speech-to-Text API.
[0795] 3. Means of providing information based on voice queries
[0796] The server searches for appropriate information based on the question obtained through voice recognition and sends that information to the terminal, which then displays the information on its screen.
[0797] User
[0798] The user is a customer of the store and performs the following operations.
[0799] 1. Viewing content
[0800] Users view the POP content displayed on the display and check information about products and services that interest them.
[0801] 2. Questions and Pointing
[0802] Users can ask for details about products they are interested in by voice or by pointing at a specific product on the screen. The question or information about the product they are pointing at is then displayed on the screen to help them make a purchasing decision.
[0803] Specific examples
[0804] Example 1: Parent and child visitor scenario
[0805] A parent and child visit a shopping mall. The server captures the faces of the parent and child via camera and determines that the parent is in their late 30s and the child is about 7 years old. Based on this, the device displays information mainly about toys and events for children. When the parent asks, "Where are the toy sales?", the device recognizes the voice question, retrieves the answer from the server, and displays it on the display.
[0806] Example 2: Checking details of individual products
[0807] The customer points at the image of a product displayed on the display. The terminal recognizes the pointing gesture and retrieves detailed information about that product from the server. The product's price, specifications, stock status, and other information are displayed on the display in real time.
[0808] Prompt Sentence Examples
[0809] "What kind of information should be displayed when a parent and child in their 30s visit?"
[0810] "When a customer points to a specific product, explain how you would retrieve and display information about that product."
[0811] This invention makes it possible to provide optimal information in real time based on customer attributes and behavior, thereby enhancing sales promotion effectiveness. In addition, by continually updating the AI model based on collected data, it is possible to continuously generate effective content.
[0812] The flow of the identification process in the first embodiment will be described with reference to FIG.
[0813] Step 1: Discover customers and obtain attribute data
[0814] The server monitors the store in real time through the installed cameras. It captures video data and uses it as input to extract attribute information such as age, gender, and number of people. This is done using Python and OpenCV, and processing high-resolution video data with facial recognition algorithms such as YOLO.
[0815] Input: Real-time video data from the camera
[0816] Output: Customer attribute data (age, gender, number of people, etc.)
[0817] Step 2: Parse and store customer attribute data
[0818] The server analyzes the acquired customer attribute data and stores it in a database. Database management uses MySQL or PostgreSQL, and data management and updating is performed using the Django or Flask framework.
[0819] Input: Customer attribute data
[0820] Output: Analysis data stored in a database
[0821] Step 3: Generate optimal POP content
[0822] The server selects the most suitable products and services from a product information database based on the accumulated customer attribute data, and dynamically generates personalized POP content using a generative AI model (e.g., GPT-3) and sends it to the device in HTML or image format.
[0823] Input: Customer attribute data, product information
[0824] Output: Personalized POP content
[0825] Step 4: View POP Content
[0826] The terminal displays POP content sent from the server in real time. The displayed content includes images, videos, and text. It runs as a web application in a browser on a Raspberry Pi or Windows PC.
[0827] Input: POP content sent from the server
[0828] Output: Content displayed on the display
[0829] Step 5: Recognizing pointing gestures and providing detailed information
[0830] The device uses a camera to recognize the customer's pointing movements in real time and sends the data to a server. The server then analyzes the movement data and obtains detailed information about the product the customer is pointing at. The movement recognition uses TensorFlow and OpenPose.
[0831] Input: Customer pointing gesture
[0832] Output: Detailed information about the pointed item
[0833] Step 6: Recognize voice questions and provide answers
[0834] The device uses a microphone to capture the customer's voice question and converts it into text using the Google Cloud Speech-to-Text API. The server then analyzes the question from the converted text, searches for and retrieves the appropriate answer, and sends it to the device, which then displays the information on its screen.
[0835] Input: Voice question data
[0836] Output: Answer information for the question
[0837] Step 7: Collect purchasing data and update the AI model
[0838] The server periodically collects customer purchase data and stores it in a database. The collected data is used to train and update the AI model, optimizing the effectiveness of the generated POP content.
[0839] Input: Purchasing data
[0840] Output: Updated AI model, optimized POP content
[0841] (Application example 1)
[0842] Next, a description will be given of Application Example 1. In the following description, the data processing device 12 will be referred to as a "server" and the smart glasses 214 will be referred to as a "terminal."
[0843] Conventional digital signage-type POP systems were primarily intended to provide information to customers in retail stores and shopping malls. However, in the logistics field, optimizing worker flow and supporting rapid inventory management are important, and there were few systems that met these needs. Therefore, there is a demand for technology that can improve work efficiency and increase the accuracy of inventory management in logistics centers, warehouses, and other on-site locations.
[0844] The specific processing by the specific processing unit 290 of the data processing device 12 in the application example 1 is realized by the following means.
[0845] In this invention, the server includes a face recognition unit, a customer attribute analysis unit, a product information database, a unit for generating POP content in real time, a display unit for displaying the generated POP content, a unit for recognizing customer pointing gestures, a unit for recognizing customer questions by voice, a unit for providing information based on the voice questions, a unit for personalizing information, a unit for collecting purchase data, a unit for updating an AI model based on the collected data, a unit for optimizing worker movement lines, a unit for visually guiding inventory locations, and a unit for linking with smart glasses or a head-mounted display. This makes it possible to optimize worker movement lines and provide real-time visual guidance on inventory locations at sites such as logistics centers and warehouses.
[0846] "Facial recognition means" is a technology that uses a camera to capture the face of a worker and analyzes attribute data such as age and gender based on the image.
[0847] "Customer attribute analysis means" is a technology that analyzes the tendencies and characteristics of specific customers based on attribute data acquired by face recognition means.
[0848] The "product information database" is a database that stores detailed information about each product or service, allowing users to search and reference the information they need.
[0849] "Means for generating POP content in real time" refers to technology for dynamically generating personalized POP content based on data collected in real time.
[0850] "Display means" refers to a device that displays generated POP content and information in real time, and includes display formats such as images, videos, and text.
[0851] The "means for recognizing pointing gestures" is a technology that uses a camera to detect the pointing gestures of workers and customers and analyzes their location information.
[0852] "Voice recognition means" refers to a technology that uses a microphone to capture the voice of a worker or customer and converts it into text data using voice recognition technology.
[0853] "Means for providing information based on voice questions" refers to a technology that acquires appropriate information from a server based on the content of a voice-recognized question and displays it on a display means.
[0854] "Means for personalizing information" refers to technology that generates individually optimized information and messages based on the attributes and behavioral data of customers and workers.
[0855] "Means for collecting purchasing data" refers to technology that collects data on customers' actual purchasing behavior and selected products and stores it in a database.
[0856] "Means for updating AI models based on collected data" refers to technologies for analyzing collected data and continuously training and optimizing AI models.
[0857] "Means to optimize worker movement" refers to technology that analyzes worker location information and movement in real time and proposes optimal movement routes and work procedures.
[0858] "Means for visually guiding inventory location" refers to technology that uses smart glasses or head-mounted displays within logistics centers to visually show workers the exact location of inventory.
[0859] "Means for linking with smart glasses or head-mounted displays" refers to technology that links smart glasses or head-mounted displays with the system to display information in real time and accept input.
[0860] This invention relates to a digital signage system for improving work efficiency in logistics centers and warehouses. This system uses facial recognition to analyze worker attributes, optimize worker movement lines, and visually guide inventory locations. Furthermore, by linking with smart glasses or head-mounted displays, work efficiency can be further improved.
[0861] System Configuration
[0862] This system is broadly composed of a server, terminals, and users. Each component is explained in detail below.
[0863] server
[0864] The server has the following functions:
[0865] 1. Facial Recognition Methods
[0866] The camera captures the worker's face and acquires attribute data such as age and gender, which is then analyzed using a facial recognition algorithm.
[0867] 2. Customer attribute analysis means
[0868] The system analyzes worker characteristics based on attribute data acquired through facial recognition, which is updated in real time and stored in a database.
[0869] 3. Product information database
[0870] This database comprehensively manages inventory information within the distribution center and is used to identify optimal inventory locations and routes.
[0871] 4. A way to generate POP content in real time
[0872] Generate personalized POP content in real time based on worker attribute information and inventory information.
[0873] 5. Optimizing traffic flow
[0874] Analyzes the real-time location information of workers and calculates and suggests the optimal movement route.
[0875] 6. Inventory location visual guide means
[0876] The exact location of inventory is displayed in real time on smart glasses or head-mounted displays.
[0877] 7. AI model update methods
[0878] The collected data is used to train and optimize the AI model, which is an algorithm designed to continuously improve work efficiency.
[0879] Terminal
[0880] The terminals are installed inside the logistics center and serve as an interface with workers.
[0881] 1. Display Means
[0882] Displays real-time POP content and traffic flow information in the form of images, videos, text, etc.
[0883] 2. Pointing Recognition Method
[0884] The camera is used to recognize the worker's pointing gesture, and detailed information is requested from the server based on that gesture.
[0885] 3. Voice Recognition Method
[0886] The worker's voice is captured using a microphone and converted into text data using voice recognition technology.
[0887] 4. Means of providing information on the display
[0888] Information acquired from the server based on questions received by voice is displayed on the display.
[0889] User
[0890] The users are workers in the logistics center and use the system as follows:
[0891] 1. Viewing content
[0892] Workers can visually check the content displayed on the screen and grasp the information necessary for their work.
[0893] 2. Questions and Pointing
[0894] Workers can ask questions by voice or point to specific locations on the screen to obtain information.
[0895] Specific examples
[0896] Example 1: Picking support
[0897] When picking items in a distribution center, the smart glasses optimize worker movement in real time, suggesting the most efficient route, and visually guiding workers to the exact location of inventory, significantly reducing work time.
[0898] Example prompt sentence:
[0899] "Please build an AI model that can suggest optimal movement paths and display inventory locations in real time to help a man in his 30s improve the efficiency of his picking work."
[0900] Example 2: Improving efficiency of warehousing operations
[0901] When receiving new inventory, the smart glasses will instruct workers in real time on where to place the inventory, supporting efficient stocking operations. Facial recognition is used to provide an individually optimized route and placement location.
[0902] Example prompt sentence:
[0903] "Propose an AI model for smart glasses that uses facial recognition to guide workers to the appropriate shelf location when receiving new inventory."
[0904] The flow of the specific processing in the application example 1 will be described with reference to FIG.
[0905] Step 1:
[0906] The server captures the worker's face through a camera. The input data is the image from the camera, and it analyzes attribute data such as age and gender using a facial recognition algorithm. The output is the analyzed worker's attribute data.
[0907] Step 2:
[0908] The server analyzes the characteristics of the workers based on the attribute data acquired in step 1. The input is customer attribute data, which is stored in a database and updated in real time. The output is the updated characteristic data.
[0909] Step 3:
[0910] The server searches and references inventory information from the product information database. The input is the worker's characteristic data and the inventory database, and the output is information on the optimal inventory location and route. This allows data to be collected to optimize worker movement lines.
[0911] Step 4:
[0912] The server generates the most suitable POP content for each worker in real time. The input is the worker's attribute data and inventory information, and the generated content is sent to the terminal. The output is the dynamically generated POP content.
[0913] Step 5:
[0914] The terminal displays the POP content sent from the server. The input is the POP content sent from the server, and the output is the content displayed on the display.
[0915] Step 6:
[0916] The terminal recognizes the worker's pointing gestures in real time. The input is video of the worker's movements captured by the terminal's built-in camera, and the output is the location information of the recognized pointing gesture. Based on this information, a request for detailed information is made to the server.
[0917] Step 7:
[0918] The terminal captures the worker's voice with a microphone and converts it into text data using voice recognition technology. The input is the worker's voice data, and the output is the converted text data.
[0919] Step 8:
[0920] The server acquires the appropriate information based on the question that has been recognized by voice and sends it to the terminal. The input is the text data that has been recognized by voice, and the output is the acquired information. This provides the necessary information to the worker.
[0921] Step 9:
[0922] The server trains and optimizes the AI model based on the collected data. The input is the collected purchase and operation data, and the output is an updated AI model. This updated model is used to guide future flow optimization and inventory location.
[0923] Step 10:
[0924] The terminal displays information sent from the server on smart glasses or a head-mounted display. The input is inventory location data and movement line data sent from the server, and the output is information displayed on the visual device. This allows workers to work efficiently.
[0925] Furthermore, an emotion engine that estimates the user's emotion may be further combined. That is, the identification processing unit 290 may estimate the user's emotion using the emotion identification model 59, and perform identification processing using the user's emotion.
[0926] This invention relates to a digital signage-type POP system that introduces optimal products and services in real time based on customer attribute information and emotional data. This system is used in retail stores and shopping malls to provide personalized information to individual customers.
[0927] System Configuration
[0928] This system is broadly composed of a "server," a "terminal," and a "user." Details of each component are explained below.
[0929] server
[0930] The server has the following functions:
[0931] 1. Facial Recognition Methods:
[0932] The server captures the customer's face through a camera and obtains high-resolution image data.
[0933] Facial recognition algorithms are used to obtain customer attribute data such as age, gender, and number of customers.
[0934] 2. Customer attribute analysis tools and sentiment engine:
[0935] Attribute data analyzed from facial recognition data is stored in a database and updated in real time.
[0936] The emotion engine analyzes customer emotions from captured facial images and tone of voice, and generates emotion data.
[0937] 3. Product information database and personalized message generation:
[0938] The product information database stores detailed information about each product and service, allowing users to search and reference the information they need.
[0939] By referencing customer attribute data and sentiment data, the system selects the most suitable products and services and dynamically generates POP content.
[0940] Add personalized messages from your membership database and purchase history to generate individual messages for specific customers.
[0941] 4. Learning and optimization:
[0942] The server analyzes the collected customer behavioral data, purchasing data, and emotional data to evaluate the effectiveness of POP content.
[0943] Based on the results of the effectiveness evaluation, the AI model will be updated to improve the accuracy of future content generation algorithms.
[0944] Terminal
[0945] The terminal is installed inside the store and serves as an interface with customers.
[0946] 1. Display means:
[0947] Real-time POP content sent from the server is displayed on the display.
[0948] The content can be displayed in the form of images, videos, text, etc.
[0949] 2. Pointing recognition method:
[0950] The camera recognizes the customer's pointing movements in real time.
[0951] A request is made to the server for detailed information about the product pointed to by the customer, and the received information is displayed on the screen.
[0952] 3. Question and Answer Function:
[0953] The customer's voice questions are captured using a microphone and converted into text data using voice recognition technology.
[0954] Information based on the question is obtained from the server and displayed on the screen.
[0955] User
[0956] The user is a customer of the store and performs the following operations.
[0957] 1. Viewing content:
[0958] Users view the POP content displayed on the display and check information about products and services that interest them.
[0959] 2. Questions and Pointing:
[0960] Users can ask for details about products they are interested in by voice or by pointing at specific products on the screen.
[0961] The question and information about the product you point at will be displayed on the screen to help you make a purchasing decision.
[0962] Specific examples
[0963] Example 1: Parent-child visit scenario
[0964] A parent and child visit a shopping mall. The server captures the faces of the parent and child via camera and analyzes that the parent is in their late 30s and the child is around 7 years old. When the emotion engine recognizes the parent's facial expression as being interested and having fun, the device displays information mainly about toys and events for children. When the parent asks, "Where are the toy sales?", the device sends the voice question and analysis results to the server and displays the corresponding information on the display.
[0965] Example 2: Checking details of individual products and emotion recognition
[0966] When a customer points at an image of a product displayed on the display, the device recognizes the pointing gesture and requests detailed information about that product from the server. The server retrieves the product information, and the emotion engine also analyzes the reaction. For example, if the customer shows a surprised expression, information about a special campaign will be displayed. The customer then checks the product's price, specifications, and stock status before making a purchasing decision.
[0967] In this way, combining emotion engines enables more personalized and effective information provision to customers. Using emotion data, it is possible to respond to customers' instantaneous reactions and provide a more personalized shopping experience. This configuration makes it possible to maximize sales promotion effects and improve customer satisfaction.
[0968] The processing flow will be explained below.
[0969] Program processing steps
[0970] Server Processing Steps
[0971] Step 1:
[0972] The server captures the faces of customers in the store through a camera and obtains high-resolution image data.
[0973] Step 2:
[0974] The server analyzes the acquired facial images using an AI algorithm to estimate attribute data such as the customer's age, gender, and number of people.
[0975] Step 3:
[0976] The server updates the analysis results to a database in real time and stores customer attribute data.
[0977] Step 4:
[0978] The server uses an emotion engine to analyze facial images, tone of voice, and other factors to generate customer emotion data.
[0979] Step 5:
[0980] The server stores the generated emotion data in a database and updates it in real time.
[0981] Step 6:
[0982] The server searches and selects the most suitable products and services from a product information database based on customer attribute data and emotional data.
[0983] Step 7:
[0984] The server dynamically generates POP content in real time based on the selection results.
[0985] Step 8:
[0986] The server references the member database and purchase history and generates personalized messages as needed.
[0987] Step 9:
[0988] The server transmits the generated POP content to the terminal.
[0989] Step 10:
[0990] The server analyzes the collected customer behavioral data, purchasing data, and emotional data to evaluate the effectiveness of the POP content.
[0991] Step 11:
[0992] Based on the effectiveness evaluation results, the server updates the AI model and improves the accuracy of future content generation algorithms.
[0993] Terminal processing steps
[0994] Step 1:
[0995] The terminal receives the POP content sent from the server.
[0996] Step 2:
[0997] The device displays the received content on the display, which can be in the form of images, videos, text, etc.
[0998] Step 3:
[0999] The device uses a camera to recognize the customer's pointing movements in real time.
[1000] Step 4:
[1001] The terminal requests information about the product pointed at by the customer from the server.
[1002] Step 5:
[1003] The terminal receives the detailed product information sent from the server and displays it on the display.
[1004] Step 6:
[1005] The terminal uses a microphone to capture the customer's voice asking a question.
[1006] Step 7:
[1007] The device uses voice recognition technology to convert the voice into text data and send it to the server.
[1008] Step 8:
[1009] The terminal displays the response information sent from the server on the display.
[1010] User operation steps
[1011] Step 1:
[1012] The user views the POP content displayed on the display.
[1013] Step 2:
[1014] Users can input questions about products or services that interest them by voice.
[1015] Step 3:
[1016] The user points to a particular product on the display to see more information about it.
[1017] Step 4:
[1018] Users make purchasing decisions based on the displayed information.
[1019] Specific examples
[1020] Example 1: Parent-child visit scenario
[1021] Step 1: A parent and child visit a shopping mall. The server captures the faces of the parent and child through a camera and analyzes them.
[1022] Step 2: The server estimates the age and gender of the parents and children and stores the attribute data in a database.
[1023] Step 3: The server analyzes the facial image and tone of voice using an emotion engine to generate emotion data.
[1024] Step 4: The server selects toys and event information for children based on the generated attribute data and emotion data, and generates POP content.
[1025] Step 5: The terminal receives the POP content and displays it on the display.
[1026] Step 6: The parent asks, "Where are the toy sales?" The device captures the audio and sends it to the server.
[1027] Step 7: The server generates toy department information based on the query and sends it to the terminal.
[1028] Step 8: The terminal displays the received information on its display.
[1029] Example 2: Checking details of individual products and emotion recognition
[1030] Step 1: The customer points to the product image on the display.
[1031] Step 2: The device recognizes the pointing gesture and requests detailed information about the product from the server.
[1032] Step 3: The server retrieves the product details and analyzes the customer's emotions using the emotion engine.
[1033] Step 4: The server adds special campaign information based on the product details and customer sentiment data and sends it to the terminal.
[1034] Step 5: The terminal displays the received information on the screen. The customer checks the product price, specifications, and stock information and makes a purchasing decision.
[1035] In this way, combining emotion engines enables effective provision of personalized information to customers. Using emotion data, it is possible to respond to customers' instantaneous reactions and provide a more personalized shopping experience. This configuration makes it possible to maximize sales promotion effects and improve customer satisfaction.
[1036] Example 2
[1037] Next, a description will be given of Example 2. In the following description, the data processing device 12 will be referred to as a "server" and the smart glasses 214 will be referred to as a "terminal."
[1038] Conventional digital signage and POP systems have had issues with providing personalized information to customers and not being able to dynamically recommend products based on customer emotions. As a result, it has been difficult to provide effective information to customers and increase their desire to purchase. It has also been difficult to optimize generative models using customer behavioral and emotional data.
[1039] The identification process by the identification processing unit 290 of the data processing device 12 in the second embodiment is realized by the following means. In this invention, the server includes a face recognition means, a customer attribute analysis means, and a means for acquiring emotion data. This makes it possible to recommend optimal products and services in real time based on customer attribute information and emotion data. Furthermore, by evaluating the effectiveness of the generated POP content and updating the generation model based on collected data, it becomes possible to continuously improve the accuracy of information provision.
[1040] A "facial recognition means" is an algorithm or device that uses a camera to capture facial images of customers and extract attribute data such as age, gender, and number of people.
[1041] "Customer attribute analysis means" refers to a system or program that analyzes customer attributes based on acquired facial recognition data and stores and updates that data in a database.
[1042] The "product information database" is a data storage system that stores detailed information about products and services in the store and allows for searching and referencing as needed.
[1043] The "means for acquiring emotional data" refers to an algorithm or device that analyzes the customer's facial image and tone of voice to generate emotional data such as joy, interest, or displeasure.
[1044] "Means for generating POP content in real time" refers to a system or program that dynamically generates content that recommends products and services based on customer attribute information and emotional data.
[1045] The "display means" is a display device that visually presents the POP content transmitted from the server to customers in the store.
[1046] The "means for recognizing the customer's pointing behavior" refers to a system or program that uses a sensor such as a camera to detect the customer's pointing behavior in real time and analyzes that information.
[1047] The "means for recognizing customer questions by voice" is a system or program that uses a microphone to capture the voice of a customer's question and converts it into text data using voice recognition technology.
[1048] The "means for providing information based on a voice question" is a system or program that retrieves related information from a server based on a customer's voice question and presents it on a display.
[1049] "Means for personalizing information" refers to a system or program that generates and displays personalized messages optimized for specific customers based on customer attribute data and purchase history.
[1050] A "means for collecting purchasing data" is a system or device for collecting data on customer purchasing behavior and storing it in a database.
[1051] The "means for updating the generative model based on collected data" refers to a system or algorithm that analyzes collected customer behavioral and emotional data and optimizes and updates the generative model.
[1052] MODE FOR CARRYING OUT THE INVENTION
[1053] This invention relates to a digital signage-type POP system that introduces optimal products and services in real time based on customer attribute information and emotional data. By recognizing the customer's face, analyzing their emotions, and recommending products and services based on that, personalized information is provided.
[1054] server
[1055] The server has the following features:
[1056] 1. Facial Recognition Methods:
[1057] The server captures customer faces through a camera, obtains high-resolution image data, and uses a facial recognition algorithm to extract customer attributes such as age, gender, and number of customers.
[1058] 2. Customer attribute analysis means:
[1059] Based on facial recognition data, customer attribute information is analyzed, and the data is stored in a database and updated in real time.
[1060] 3. How to get emotion data:
[1061] The emotion engine analyzes the customer's emotions using facial images, tone of voice, and other data, and generates emotion data. For example, if a customer smiles, the emotion is analyzed as "happiness."
[1062] 4. Product Information Database:
[1063] The product information database stores detailed information about each product and service, which can be searched and referenced as needed. Content is generated that recommends optimal products and services based on customer attribute data and emotional data.
[1064] 5. Personalized message generation:
[1065] Add individual personalized messages from your membership database and purchase history to generate individual messages for specific customers.
[1066] 6. Learning and optimization:
[1067] The server analyzes collected customer behavioral, purchasing, and emotional data to evaluate the effectiveness of the POP content. Based on the results of the effectiveness evaluation, the generative AI model is updated, improving the accuracy of future content generation algorithms.
[1068] Terminal
[1069] The terminals are installed in the store and act as the interface with the customer:
[1070] 1. Display means:
[1071] Real-time POP content sent from the server is displayed on the display. Content can be displayed in the form of images, videos, text, etc.
[1072] 2. Pointing recognition method:
[1073] The device's camera recognizes the customer's pointing gestures in real time, requests detailed information about the product the customer is pointing at from the server, and displays the received information on the display.
[1074] 3. Question and Answer Function:
[1075] The device's microphone captures the customer's voice and converts it into text data using speech recognition technology. Information based on the question is retrieved from the server and displayed on the screen.
[1076] User
[1077] The user is a customer in the store and performs the following actions:
[1078] 1. Viewing content:
[1079] Users view the POP content displayed on the display and check information about products and services that interest them.
[1080] 2. Questions and Pointing:
[1081] Users can ask for details about products they are interested in by voice or by pointing at a specific product on the screen. The content of the question or information about the product they are pointing at is then displayed on the screen.
[1082] Specific examples
[1083] Example 1: Parent-child visit scenario
[1084] A parent and child visit a shopping mall. The server captures the faces of the parent and child via camera and analyzes that the parent is in their late 30s and the child is about 7 years old. When the emotion engine recognizes the parent's facial expression as being interested and having fun, the device displays information mainly about toys and events for children. When the parent asks, "Where are the toy sales?", the device sends the voice question and analysis results to the server and displays the corresponding information on the display.
[1085] Example 2: Checking details of individual products and emotion recognition
[1086] When a customer points at an image of a product displayed on the display, the device recognizes the pointing gesture and requests detailed information about that product from the server. The server retrieves the product information, and the emotion engine also analyzes the reaction. For example, if the customer shows a surprised expression, information about a special campaign will be displayed. The customer then checks the product's price, specifications, and stock status before making a purchasing decision.
[1087] Examples of prompt statements
[1088] 1. "When a customer points to an image of a product on the display, explain how the device displays more information about that product."
[1089] 2. "Please explain specifically how the system will display toys and information about children's events when a parent and child visit the store."
[1090] This invention makes it possible to provide customers with personalized information, which is expected to increase their purchasing motivation. It also responds to customers' instantaneous reactions, providing a more personalized shopping experience.
[1091] The flow of the identification process in the second embodiment will be described with reference to FIG.
[1092] Step 1:
[1093] The server captures the faces of customers in the store using a camera and obtains high-resolution image data. The obtained image data is input into a facial recognition algorithm to extract customer attribute data such as age, gender, and number of customers. The output attribute data is stored in a database.
[1094] Specific behavior:
[1095] When a customer enters a store, the server automatically acquires data from the camera and begins facial recognition. For example, if a woman in her 30s and a man in his 40s visit the store together, their attribute data will be stored in a database.
[1096] Step 2:
[1097] The server inputs the acquired facial image and tone of voice into an emotion engine to generate customer emotion data. The emotion engine analyzes the customer's emotional state, such as joy, interest, or displeasure, and outputs the data. The analyzed emotion data is stored in a database.
[1098] Specific behavior:
[1099] If a customer looks at the display and smiles, the server interprets the smile as "happiness." If the customer's voice is high-pitched and excited, the emotion engine recognizes it as "interested."
[1100] Step 3:
[1101] The server references the product information database and selects the most suitable products and services based on the acquired customer attribute data and emotion data. The selected information is generated as dynamic POP content in real time. The generated POP content is then saved back into the database.
[1102] Specific behavior:
[1103] For example, if a woman in her 30s expresses interest, the server will generate content recommending the latest beauty products, adding special campaign information for related products based on the customer's past purchases.
[1104] Step 4:
[1105] The terminal displays the POP content sent from the server. The content is displayed on the display in the form of images, videos, and text. When a customer points at a specific product on the screen, the camera recognizes the pointing gesture and requests detailed information from the server.
[1106] Specific behavior:
[1107] When a customer points to an image of a toy on the display, the device automatically retrieves and displays detailed information about the product.
[1108] Step 5:
[1109] When a customer asks a question to the display, the device's microphone captures the voice of the question and converts it into text data using voice recognition technology. The converted text data is sent to the server, which returns information based on the question. The acquired information is then displayed on the display.
[1110] Specific behavior:
[1111] When a parent asks, "Where is the toy sale?", the voice question and analysis results are sent to the server, and the corresponding information is displayed on the screen.
[1112] Step 6:
[1113] The server analyzes collected customer behavioral, purchasing, and emotional data to evaluate the effectiveness of POP content. The results of the evaluation are used to update the AI model, improving the accuracy of future content generation algorithms.
[1114] Specific behavior:
[1115] Past data can be used to analyze which content was most effective. For example, data on women in their 30s can be used to learn that a particular beauty product was well-received, and this can be reflected in future recommendations.
[1116] (Application example 2)
[1117] Next, a description will be given of Application Example 2. In the following description, the data processing device 12 will be referred to as a "server" and the smart glasses 214 will be referred to as a "terminal."
[1118] In today's retail industry, there is a demand for personalized information provision to each individual customer. Furthermore, there is a lack of methods for using emotional data to suggest optimal products and services based on the customer's real-time situation. As a result, customer satisfaction and purchasing motivation have not been sufficiently improved. Therefore, there is a need to provide a system that can recommend optimal products and services in real time based on the customer's attribute information and emotional data.
[1119] The specific processing by the specific processing unit 290 of the data processing device 12 in the application example 2 is realized by the following means.
[1120] In this invention, the server includes a face recognition means, a customer attribute analysis means, a product information database, a means for generating POP content in real time, a display device for displaying the generated POP content, a means for recognizing customer pointing gestures, a means for recognizing customer questions by voice, a means for providing information based on the voice questions, a means for personalizing information, a means for collecting purchase data, a means for updating an AI model based on the collected data, a sentiment analysis means, a means for selecting optimal products and services based on the sentiment data, a means for detecting customer faces in real time using a camera and analyzing their sentiments, and a means for displaying related product information on a display device based on the analysis results. This makes it possible to propose optimal products and services based on the customer's sentiments and attributes in real time, providing a personalized shopping experience and improving customer satisfaction and purchasing motivation.
[1121] A "face recognition means" is a means of capturing a customer's face using a camera and obtaining attribute data such as the customer's age, gender, and number of people from the high-resolution image data.
[1122] The "customer attribute analysis means" is a means for analyzing attribute data acquired by the face recognition means and updating it in real time.
[1123] A "product information database" is a database that stores detailed information about each product or service, allowing users to search and reference the information they need.
[1124] "Means for generating POP content in real time" refers to means for dynamically generating POP content by referencing customer attribute data and emotional data, selecting the most suitable products and services.
[1125] The "display device" is a device for displaying real-time POP content transmitted from the server on a display.
[1126] The "means for recognizing pointing actions" is a means for recognizing a customer's pointing actions in real time through a camera and requesting detailed information about the pointed product from the server.
[1127] The "voice recognition means" is a means of capturing the customer's voice question using a microphone and converting it into text data using voice recognition technology.
[1128] The "means for providing information based on a voice question" is a means for obtaining related information from a server based on the content of a customer's voice question and displaying it on a display.
[1129] The "means for personalizing information" refers to a means for adding personalized messages from a member database or purchase history and generating individual messages for specific customers.
[1130] "Means for collecting purchasing data" refers to the means for collecting and analyzing customer behavioral data, purchasing data, and emotional data.
[1131] "Means for updating the AI model" refers to a means for evaluating the effectiveness of POP content based on collected data and updating the AI model based on the results of that evaluation.
[1132] The "emotion analysis means" is a means for analyzing the customer's emotions from acquired facial images, tone of voice, etc., and generating emotion data.
[1133] "Means for selecting optimal products and services based on emotional data" refers to means for selecting the most appropriate products and services at that time by referring to emotional data.
[1134] "Means for detecting a customer's face in real time using a camera and analyzing emotions" refers to means for capturing a customer's face in real time using a camera and analyzing the customer's emotions from the facial image.
[1135] The "means for displaying related product information on a display device based on the analysis results" refers to means for displaying related product information on a display based on the analyzed customer emotion data and attribute data.
[1136] The present invention relates to a system that introduces optimal products and services in real time based on customer attribute information and emotion data, and is configured as follows.
[1137] server
[1138] The server has the following functions:
[1139] 1. Facial Recognition Methods
[1140] The server captures the customer's face through a camera, obtains high-resolution image data, and uses a facial recognition algorithm to obtain customer attribute data such as age, gender, and number of customers.
[1141] 2. Customer attribute analysis means
[1142] Attribute data acquired through facial recognition is analyzed and updated in real time, ensuring that basic customer information is always kept up to date.
[1143] 3. Product information database
[1144] The product information database stores detailed information about each product and service, allowing you to search and reference the information you need.
[1145] 4. A way to generate POP content in real time
[1146] By referencing customer attribute data and emotional data, the system selects the most suitable products and services and dynamically generates POP content. For example, it uses prompts such as "display the most suitable products for male customers in their 30s who have a surprised expression."
[1147] 5. Display device means
[1148] Real-time POP content sent from the server is displayed on the display. Content can be displayed in the form of images, videos, text, etc.
[1149] 6. Methods for Recognizing Pointing Actions
[1150] The camera recognizes the customer's pointing gestures in real time, requests detailed information about the product the customer is pointing at from the server, and displays the received information on the display.
[1151] 7. Voice recognition
[1152] The customer's voice questions are captured using a microphone and converted into text data using voice recognition technology.
[1153] 8. Means of Providing Information Based on Voice Queries
[1154] Based on the content of the customer's voice question, related information is obtained from the server and displayed on the screen.
[1155] 9. How we personalize your information
[1156] Add personalized messages from your membership database and purchase history to generate individual messages for specific customers.
[1157] 10. Means of collecting purchasing data
[1158] Collect and analyze customer behavioral, purchasing, and emotional data.
[1159] 11. How to update your AI model
[1160] The effectiveness of POP content is evaluated based on the collected data, and the AI model is updated based on the evaluation results.
[1161] 12. Sentiment analysis tool
[1162] The system analyzes customer emotions from acquired facial images and tone of voice, and generates emotional data.
[1163] 13. A way to select the best products and services based on sentiment data
[1164] It is a means of referring to emotional data to select the most appropriate product or service at any given time.
[1165] 14. Real-time facial detection and emotion analysis using cameras
[1166] The camera captures the customer's face in real time and analyzes their emotions from the facial image.
[1167] 15. Means for displaying related product information on a display device based on the analysis results
[1168] Based on the analyzed customer emotional data and attribute data, relevant product information is displayed on the screen.
[1169] Terminal
[1170] The terminal is installed inside the store and serves as an interface with customers.
[1171] Display Means
[1172] Real-time POP content sent from the server is displayed on the display, allowing customers to directly view information about products and services that are optimized for them.
[1173] Pointing recognition method
[1174] The camera recognizes customers' pointing gestures in real time. For example, if a customer points to a product to say, "I'd like to know more about this product," detailed information will be displayed on the screen.
[1175] Voice question answering function
[1176] The microphone is used to capture the customer's voice asking a question. If a customer asks, "Are there any special offers for this product?", the voice is sent to the server and the appropriate information is displayed.
[1177] User
[1178] The user is a customer of the store.
[1179] Viewing content
[1180] Users view the POP content displayed on the display and check information about products and services that interest them.
[1181] Questions and pointing
[1182] Users can ask for details about products they are interested in by voice or by pointing at a specific product on the screen. The question or information about the product they are pointing at is then displayed on the screen to help them make a purchasing decision.
[1183] Specific examples
[1184] By inputting prompts such as, "Please tell me what product information should be provided to a female customer who looks happy," into the generative AI model, the optimal content is displayed.
[1185] In this way, the system of the present invention improves customer satisfaction and maximizes the effectiveness of store sales promotion by providing real-time information based on customer attribute information and emotional data.
[1186] The flow of the specific processing in the application example 2 will be described with reference to FIG.
[1187] Step 1:
[1188] The server captures the customer's face through a camera and obtains high-resolution image data.
[1189] Input: Video data from the camera
[1190] Output: Customer's facial image data
[1191] Specific operation: The camera captures video and sends the video data to the server, which then detects faces in the video and generates facial image data.
[1192] Step 2:
[1193] The server uses a facial recognition algorithm to obtain attribute data such as customer age, gender, and number of customers.
[1194] Input: Customer's facial image data
[1195] Output: Customer attribute data (age, gender, number of people)
[1196] Specific operation: The server runs a facial recognition algorithm, analyzes the facial image data, and extracts attribute information such as age, gender, and number of people.
[1197] Step 3:
[1198] The server stores the acquired attribute data in a database and updates it in real time.
[1199] Input: Customer attribute data
[1200] Output: Updated customer attribute information in the database
[1201] Specific operation: The server stores the attribute data in a database and updates existing customer information as necessary.
[1202] Step 4:
[1203] The server uses an emotion engine to analyze the customer's emotions from acquired facial images and tone of voice, and generates emotion data.
[1204] Input: Facial image data, audio data
[1205] Output: Customer sentiment data
[1206] Specific operation: The emotion engine analyzes facial images and voice data, infers emotions from the customer's facial expressions and tone of voice, and generates that data.
[1207] Step 5:
[1208] The server references customer attribute data and emotion data, selects the most suitable products and services from a product information database, and dynamically generates POP content.
[1209] Input: Customer attribute data, customer sentiment data
[1210] Output: Generated POP content
[1211] Specific operation: The server searches the product information database based on customer attribute information and emotion data, extracts information on related products and services, and generates POP content based on this.
[1212] Step 6:
[1213] The terminal displays the real-time POP content sent from the server on the display.
[1214] Input: Generated POP content
[1215] Output: The content shown on the display
[1216] Specific operation: Receives POP content sent from the server and displays information such as images, videos, and text on the display.
[1217] Step 7:
[1218] The device uses a camera to recognize the customer's pointing movements in real time.
[1219] Input: Video data from the camera
[1220] Output: Pointing motion detection information
[1221] Specific operation: The camera captures the customer's pointing action and sends it to the server for pointing recognition.
[1222] Step 8:
[1223] The server retrieves detailed information about the product pointed to by the customer from the database and displays it on the display.
[1224] Input: Pointing motion detection information
[1225] Output: Product details
[1226] Specific operation: Based on the pointing gesture, the server retrieves the corresponding product information from the database and sends it to the terminal, which then displays the detailed information on the screen.
[1227] Step 9:
[1228] The terminal uses a microphone to capture the customer's voice asking a question and converts it into text data using voice recognition technology.
[1229] Input: Audio data
[1230] Output: Text data of the audio
[1231] How it works: A microphone captures the customer's question, and a speech recognition algorithm on the server converts it into text data.
[1232] Step 10:
[1233] The server acquires information based on the voice question content and displays it on a display.
[1234] Input: Text data of speech
[1235] Output: Information corresponding to the question
[1236] Specific operation: The server analyzes the text data of the voice question, retrieves relevant information from the database, and sends it to the device, which then displays the answer to the question on the screen.
[1237] Step 11:
[1238] The server generates personalized messages from a member database and purchase history, and provides individual messages to specific customers.
[1239] Input: Member data, purchase history
[1240] Output: Personalized message
[1241] Specific operation: The server references the member database and purchase history, generates a personalized message tailored to the specific customer, and provides it via the terminal.
[1242] Step 12:
[1243] The server collects and analyzes customer behavioral data, purchasing data, and emotional data.
[1244] Input: Customer behavior data, purchasing data, emotional data
[1245] Output: Analysis data
[1246] Specific operation: The server analyzes the collected data in real time and uses the results to generate the next POP content and respond to customers.
[1247] Step 13:
[1248] The server evaluates the effectiveness of the POP content based on the collected data and updates the AI model based on the evaluation results.
[1249] Input: Customer behavior data, purchasing data, emotional data
[1250] Output: Updated AI model
[1251] Specific operation: The server analyzes the collected data and adjusts the parameters of the AI model to improve the accuracy of subsequent content generation algorithms.
[1252] The specific processing unit 290 transmits the result of the specific processing to the smart glasses 214. In the smart glasses 214, the control unit 46A causes the speaker 240 to output the result of the specific processing. The microphone 238 acquires audio indicating a user input regarding the result of the specific processing. The control unit 46A transmits audio data indicating the user input acquired by the microphone 238 to the data processing device 12. In the data processing device 12, the specific processing unit 290 acquires the audio data.
[1253] The data generation model 58 is a so-called generative AI (Artificial Intelligence). An example of the data generation model 58 is ChatGPT (Internet Search<URL: https: / / openai.com / blog / chatgpt> ), Gemini (Internet search <url: https: gemini.google.com ?hl="ja">) and other generation AIs. The data generation model 58 is obtained by performing deep learning on a neural network. A prompt including an instruction is input to the data generation model 58, and inference data such as voice data indicating voice, text data indicating text, and image data indicating an image is also input. The data generation model 58 performs inference on the input inference data in accordance with the instruction indicated by the prompt, and outputs the inference result in a data format such as voice data and text data. Here, inference refers to, for example, analysis, classification, prediction, and / or summarization.
[1254] In the above embodiment, an example in which the specific processing is performed by the data processing device 12 has been given, but the technology of the present disclosure is not limited to this, and the specific processing may be performed by the smart glasses 214.
[1255] [Third embodiment]
[1256] FIG. 5 shows an example of the configuration of a data processing system 310 according to the third embodiment.
[1257] 5, the data processing system 310 includes the data processing device 12 and a headset type terminal 314. An example of the data processing device 12 is a server.
[1258] The data processing device 12 includes a computer 22, a database 24, and a communication I / F 26. The computer 22 is an example of a "computer" according to the technology of the present disclosure. The computer 22 includes a processor 28, a RAM 30, and a storage 32. The processor 28, the RAM 30, and the storage 32 are connected to a bus 34. The database 24 and the communication I / F 26 are also connected to the bus 34. The communication I / F 26 is connected to a network 54. Examples of the network 54 include a WAN (Wide Area Network) and / or a LAN (Local Area Network).
[1259] The headset type terminal 314 includes a computer 36, a microphone 238, a speaker 240, a camera 42, a communication I / F 44, and a display 343. The computer 36 includes a processor 46, a RAM 48, and a storage 50. The processor 46, the RAM 48, and the storage 50 are connected to a bus 52. The microphone 238, the speaker 240, the camera 42, and the display 343 are also connected to the bus 52.
[1260] The microphone 238 receives instructions and the like from the user 20 by receiving voice uttered by the user 20. The microphone 238 captures the voice uttered by the user 20, converts the captured voice into audio data, and outputs it to the processor 46. The speaker 240 outputs audio in accordance with instructions from the processor 46.
[1261] Camera 42 is a small digital camera equipped with an optical system including a lens, aperture, and shutter, and an imaging element such as a CMOS (Complementary Metal-Oxide-Semiconductor) image sensor or a CCD (Charge Coupled Device) image sensor, and captures images of the surroundings of user 20 (for example, an imaging range defined by an angle of view equivalent to the field of vision of a typical healthy person).
[1262] The communication I / F 44 is connected to a network 54. The communication I / Fs 44 and 26 are responsible for the exchange of various information between the processor 46 and the processor 28 via the network 54. The exchange of various information between the processor 46 and the processor 28 using the communication I / Fs 44 and 26 is carried out in a secure state.
[1263] Fig. 6 shows an example of the main functions of the data processing device 12 and the headset type terminal 314. As shown in Fig. 6, in the data processing device 12, a specific process is performed by the processor 28. A specific process program 56 is stored in the storage 32.
[1264] The specific processing program 56 is an example of a "program" according to the technology of the present disclosure. The processor 28 reads the specific processing program 56 from the storage 32 and executes the read specific processing program 56 on the RAM 30. The specific processing is realized by the processor 28 operating as a specific processing unit 290 in accordance with the specific processing program 56 executed on the RAM 30.
[1265] The storage 32 stores a data generation model 58 and an emotion identification model 59. The data generation model 58 and the emotion identification model 59 are used by the identification processing unit 290.
[1266] In the headset type terminal 314, a reception output process is performed by the processor 46. A reception output program 60 is stored in the storage 50. The processor 46 reads the reception output program 60 from the storage 50 and executes the read reception output program 60 on the RAM 48. The reception output process is realized by the processor 46 operating as the control unit 46A in accordance with the reception output program 60 executed on the RAM 48.
[1267] Next, a description will be given of the identification process performed by the identification processing unit 290 of the data processing device 12. In the following description, the data processing device 12 will be referred to as the "server" and the headset type terminal 314 will be referred to as the "terminal."
[1268] This invention relates to a digital signage-type POP system that introduces optimal products and services in real time based on customer attribute information. This system is used in retail stores and shopping malls to provide personalized information to individual customers.
[1269] System Configuration
[1270] This system is broadly composed of a "server," a "terminal," and a "user." Details of each component are explained below.
[1271] server
[1272] The server has the following functions:
[1273] 1. Facial Recognition Methods:
[1274] The server captures customers' faces through a camera and obtains attribute data such as age, gender, and number of people.
[1275] Using high-resolution video data, customer attributes are analyzed using facial recognition algorithms.
[1276] 2. Customer attribute analysis means:
[1277] Based on facial recognition, analyzed attribute data is stored in a database and updated in real time.
[1278] The analysis results are used to provide promotional information to each customer.
[1279] 3. Product information database and personalized message generation:
[1280] The product information database stores detailed information about each product and service, allowing users to search and reference the information they need.
[1281] Customer attribute data is referenced to select the most suitable products and services and dynamically generate POP content.
[1282] Add personalized messages from your membership database and purchase history to generate individual messages for specific customers.
[1283] 4. Learning and optimization:
[1284] The server analyzes collected customer behavioral and purchasing data and updates the AI model to optimize the effectiveness of POP content.
[1285] We continuously collect data and improve our algorithms to generate optimal content.
[1286] Terminal
[1287] The terminal is installed inside the store and serves as an interface with customers.
[1288] 1. Display means:
[1289] Real-time POP content sent from the server is displayed on the display.
[1290] The content can be displayed in the form of images, videos, text, etc.
[1291] 2. Pointing recognition method:
[1292] The camera recognizes the customer's pointing movements in real time.
[1293] A request is made to the server for detailed information about the product pointed to by the customer, and the received information is displayed on the screen.
[1294] 3. Question and Answer Function:
[1295] The customer's voice questions are captured using a microphone and converted into text data using voice recognition technology.
[1296] Information based on the question is obtained from the server and displayed on the screen.
[1297] User
[1298] The user is a customer of the store and performs the following operations.
[1299] 1. Viewing content:
[1300] Users view the POP content displayed on the display and check information about products and services that interest them.
[1301] 2. Questions and Pointing:
[1302] Users can ask for details about products they are interested in by voice or by pointing at specific products on the screen.
[1303] The question and information about the product you point at will be displayed on the screen to help you make a purchasing decision.
[1304] Specific examples
[1305] Example 1: Parent and child visitor scenario
[1306] A parent and child visit a shopping mall. The server captures the faces of the parent and child via camera and determines that the parent is in their late 30s and the child is about 7 years old. Based on this, the device displays information mainly about toys and events for children. When the parent asks, "Where are the toy sales?", the device recognizes the voice question, retrieves the answer from the server, and displays it on the display.
[1307] Example 2: Checking details of individual products
[1308] The customer points at the image of a product displayed on the display. The terminal recognizes the pointing gesture and retrieves detailed information about that product from the server. The product's price, specifications, stock status, and other information are displayed on the display in real time.
[1309] As described above, this system can provide optimal information in real time based on customer attributes and behavior, thereby enhancing sales promotion effectiveness. Furthermore, by continually updating the AI model based on collected data, it is possible to continuously generate effective content.
[1310] The processing flow will be explained below.
[1311] Program processing steps
[1312] Server Processing Steps
[1313] Step 1:
[1314] The server captures the faces of customers in the store through a camera and obtains high-resolution image data.
[1315] Step 2:
[1316] The server analyzes the acquired facial images using an AI algorithm to estimate attribute data such as the customer's age, gender, and number of people.
[1317] Step 3:
[1318] The server updates the analysis results to a database and stores customer attribute data in real time.
[1319] Step 4:
[1320] The server searches for the most suitable products and services from a product information database based on customer attribute data.
[1321] Step 5:
[1322] The server dynamically generates POP content in real time based on the search results.
[1323] Step 6:
[1324] The server references the member database and purchase history and generates personalized messages as needed.
[1325] Step 7:
[1326] The server transmits the generated POP content to the terminal.
[1327] Step 8:
[1328] The server analyzes the collected customer behavior data and purchase data to evaluate the effectiveness of the POP content.
[1329] Step 9:
[1330] Based on the effectiveness evaluation results, the server updates the AI model and improves the accuracy of future content generation algorithms.
[1331] Terminal processing steps
[1332] Step 1:
[1333] The terminal receives the POP content sent from the server.
[1334] Step 2:
[1335] The device displays the received content on the display, which can be in the form of images, videos, text, etc.
[1336] Step 3:
[1337] The device uses a camera to recognize the customer's pointing movements in real time.
[1338] Step 4:
[1339] The terminal requests information about the product pointed at by the customer from the server.
[1340] Step 5:
[1341] The terminal receives the detailed product information sent from the server and displays it on the display.
[1342] Step 6:
[1343] The terminal uses a microphone to capture the customer's voice asking a question.
[1344] Step 7:
[1345] The device uses voice recognition technology to convert the voice into text data and send it to the server.
[1346] Step 8:
[1347] The terminal displays the response information sent from the server on the display.
[1348] User operation steps
[1349] Step 1:
[1350] The user views the POP content displayed on the display.
[1351] Step 2:
[1352] Users can input questions about products or services that interest them by voice.
[1353] Step 3:
[1354] The user points to a particular product on the display to see more information about it.
[1355] Step 4:
[1356] Users make purchasing decisions based on the displayed information.
[1357] Specific examples
[1358] Example 1: Parent-child visit scenario
[1359] Step 1: A parent and child visit a shopping mall. The server captures the faces of the parent and child through a camera and analyzes them.
[1360] Step 2: The server estimates the age and gender of the parents and children and updates the attribute data in the database.
[1361] Step 3: The server selects toys and event information for children based on the parent-child attributes and generates POP content.
[1362] Step 4: The terminal receives the POP content and displays it on the display.
[1363] Step 5: The parent asks, "Where are the toy sales?" The device captures the audio and sends it to the server.
[1364] Step 6: The server generates toy department information based on the query and sends it to the terminal.
[1365] Step 7: The terminal displays the received information on its display.
[1366] Example 2: Checking details of individual products
[1367] Step 1: The customer points to the product image on the display.
[1368] Step 2: The device recognizes the pointing gesture and requests detailed information about the product from the server.
[1369] Step 3: The server obtains the product details and sends them to the terminal.
[1370] Step 4: The terminal displays the received information on its display.
[1371] Step 5: The customer checks the product price, specifications, and stock information and makes a purchasing decision.
[1372] The above is the specific flow of operations at each step. This invention realizes the provision of personalized information in real time based on the attributes and behavior of customers.
[1373] Example 1
[1374] Next, a description will be given of Example 1. In the following description, the data processing device 12 will be referred to as a "server" and the headset type terminal 314 will be referred to as a "terminal."
[1375] Conventional digital signage systems have the problem that the information provided to customers is uniform and not personalized enough for each individual customer. As a result, it is not possible to introduce optimal products and services that meet each customer's interests and needs, and sales promotion effectiveness is limited. Furthermore, there is a lack of a mechanism to effectively collect and analyze customer behavior data and optimize content in real time.
[1376] The specific processing by the specific processing unit 290 of the data processing device 12 in the first embodiment is realized by the following means.
[1377] In this invention, the server includes a means for capturing customer faces and acquiring attribute data, a means for analyzing and saving the acquired attribute data, and a database for accumulating product information. This makes it possible to select and display optimal products and services in real time according to customer attributes. Furthermore, by collecting purchasing data and updating and optimizing the AI model based on this, it is possible to consistently provide effective content.
[1378] "Means for capturing customer faces and acquiring attribute data" refers to technology that uses cameras and sensors to capture images of customers' faces in a store and extract information such as age, gender, and number of customers from that image data.
[1379] "Means for analyzing and storing acquired attribute data" refers to hardware and software for analyzing attribute data obtained by facial recognition technology and storing it in a database.
[1380] A "database that stores product information" is a database system that manages detailed information about each product and service sold in a retail store or shopping mall, and allows for searching and referencing.
[1381] "Means for generating personalized POP content in real time" refers to algorithms and software for dynamically creating and displaying promotional information that is optimal for each customer based on their attribute data.
[1382] The "display means for displaying the generated content" refers to a display device for visually presenting the POP content sent from the server to customers in the form of images, videos, text, etc.
[1383] "Means for recognizing customer pointing movements" refers to technology that uses sensors such as cameras to detect customer pointing movements and analyze those movements.
[1384] "Means for recognizing customer questions via voice" refers to voice recognition technology that uses a microphone to capture the customer's voice and converts the voice data into text.
[1385] "Means for providing information based on voice questions" refers to a system that searches for appropriate information based on the question content obtained by voice recognition and provides that information.
[1386] "Means for personalizing information" refers to algorithms and systems for providing the most appropriate information to individual customers based on their attribute data and purchasing history.
[1387] "Means for collecting purchasing data" refers to a system for collecting customer purchasing history and behavioral data, and storing and managing this in a database.
[1388] "Means for updating AI models based on collected data and optimizing POP content" refers to a system that uses collected customer behavioral and purchasing data to continuously learn and update AI models to generate optimal POP content.
[1389] This invention relates to a digital signage-type POP system that introduces optimal products and services in real time based on customer attribute information. This system is used in retail stores and shopping malls to provide personalized information to individual customers.
[1390] System Configuration
[1391] This system is broadly composed of a "server," a "terminal," and a "user." Each component will be explained in detail below.
[1392] server
[1393] The server has the following functions:
[1394] 1. A means of capturing customer faces and obtaining attribute data
[1395] The server monitors the store in real time through installed cameras, capturing images of customers' faces and acquiring attribute data such as age, gender, number of people, etc. This facial recognition is performed using Python and OpenCV, applying facial recognition algorithms such as YOLO.
[1396] 2. Means for analyzing and storing acquired attribute data
[1397] The server analyzes the acquired customer attribute data and saves it in a database. The attribute data is continuously updated and to manage the information of multiple customers in real time, a database management system such as MySQL or PostgreSQL is used. The data is managed via an API using the Django or Flask framework.
[1398] 3. Database for storing product information
[1399] The product information database stores detailed information about the products it handles, allowing for quick search and reference as needed. This database includes information such as product prices, specifications, and stock status.
[1400] 4. A way to generate personalized POP content in real time
[1401] The server selects the most suitable products and services based on customer attribute data and dynamically generates personalized POP content using a generative AI model (e.g., GPT-3). The generated POP content is converted into HTML or image format as a text message or a list of recommended products and sent to the device.
[1402] 5. Display means for displaying generated content
[1403] The terminal visually displays real-time POP content sent from the server on a display, which can be powered by a Raspberry Pi or Windows PC and runs as a web application in a browser.
[1404] Terminal
[1405] The terminal is installed inside the store and serves as an interface with customers.
[1406] 1. A method for recognizing customer pointing gestures
[1407] The device uses a camera to recognize the customer's pointing gestures in real time, using hand movement recognition algorithms such as TensorFlow and OpenPose, and sends the processing results to a server using a Python script.
[1408] 2. A way to recognize customer questions by voice
[1409] The device uses a microphone to capture the customer's voice and converts it into text using speech recognition technology, using the Google Cloud Speech-to-Text API.
[1410] 3. Means of providing information based on voice queries
[1411] The server searches for appropriate information based on the question obtained through voice recognition and sends that information to the terminal, which then displays the information on its screen.
[1412] User
[1413] The user is a customer of the store and performs the following operations.
[1414] 1. Viewing content
[1415] Users view the POP content displayed on the display and check information about products and services that interest them.
[1416] 2. Questions and Pointing
[1417] Users can ask for details about products they are interested in by voice or by pointing at a specific product on the screen. The question or information about the product they are pointing at is then displayed on the screen to help them make a purchasing decision.
[1418] Specific examples
[1419] Example 1: Parent and child visitor scenario
[1420] A parent and child visit a shopping mall. The server captures the faces of the parent and child via camera and determines that the parent is in their late 30s and the child is about 7 years old. Based on this, the device displays information mainly about toys and events for children. When the parent asks, "Where are the toy sales?", the device recognizes the voice question, retrieves the answer from the server, and displays it on the display.
[1421] Example 2: Checking details of individual products
[1422] The customer points at the image of a product displayed on the display. The terminal recognizes the pointing gesture and retrieves detailed information about that product from the server. The product's price, specifications, stock status, and other information are displayed on the display in real time.
[1423] Prompt Sentence Examples
[1424] "What kind of information should be displayed when a parent and child in their 30s visit?"
[1425] "When a customer points to a specific product, explain how you would retrieve and display information about that product."
[1426] This invention makes it possible to provide optimal information in real time based on customer attributes and behavior, thereby enhancing sales promotion effectiveness. In addition, by continually updating the AI model based on collected data, it is possible to continuously generate effective content.
[1427] The flow of the identification process in the first embodiment will be described with reference to FIG.
[1428] Step 1: Discover customers and obtain attribute data
[1429] The server monitors the store in real time through the installed cameras. It captures video data and uses it as input to extract attribute information such as age, gender, and number of people. This is done using Python and OpenCV, and processing high-resolution video data with facial recognition algorithms such as YOLO.
[1430] Input: Real-time video data from the camera
[1431] Output: Customer attribute data (age, gender, number of people, etc.)
[1432] Step 2: Parse and store customer attribute data
[1433] The server analyzes the acquired customer attribute data and stores it in a database. Database management uses MySQL or PostgreSQL, and data management and updating is performed using the Django or Flask framework.
[1434] Input: Customer attribute data
[1435] Output: Analysis data stored in a database
[1436] Step 3: Generate optimal POP content
[1437] The server selects the most suitable products and services from a product information database based on the accumulated customer attribute data, and dynamically generates personalized POP content using a generative AI model (e.g., GPT-3) and sends it to the device in HTML or image format.
[1438] Input: Customer attribute data, product information
[1439] Output: Personalized POP content
[1440] Step 4: View POP Content
[1441] The terminal displays POP content sent from the server in real time. The displayed content includes images, videos, and text. It runs as a web application in a browser on a Raspberry Pi or Windows PC.
[1442] Input: POP content sent from the server
[1443] Output: Content displayed on the display
[1444] Step 5: Recognizing pointing gestures and providing detailed information
[1445] The device uses a camera to recognize the customer's pointing movements in real time and sends the data to a server. The server then analyzes the movement data and obtains detailed information about the product the customer is pointing at. The movement recognition uses TensorFlow and OpenPose.
[1446] Input: Customer pointing gesture
[1447] Output: Detailed information about the pointed item
[1448] Step 6: Recognize voice questions and provide answers
[1449] The device uses a microphone to capture the customer's voice question and converts it into text using the Google Cloud Speech-to-Text API. The server then analyzes the question from the converted text, searches for and retrieves the appropriate answer, and sends it to the device, which then displays the information on its screen.
[1450] Input: Voice question data
[1451] Output: Answer information for the question
[1452] Step 7: Collect purchasing data and update the AI model
[1453] The server periodically collects customer purchase data and stores it in a database. The collected data is used to train and update the AI model, optimizing the effectiveness of the generated POP content.
[1454] Input: Purchasing data
[1455] Output: Updated AI model, optimized POP content
[1456] (Application example 1)
[1457] Next, a description will be given of Application Example 1. In the following description, the data processing device 12 will be referred to as a "server" and the headset type terminal 314 will be referred to as a "terminal."
[1458] Conventional digital signage-type POP systems were primarily intended to provide information to customers in retail stores and shopping malls. However, in the logistics field, optimizing worker flow and supporting rapid inventory management are important, and there were few systems that met these needs. Therefore, there is a demand for technology that can improve work efficiency and increase the accuracy of inventory management in logistics centers, warehouses, and other on-site locations.
[1459] The specific processing by the specific processing unit 290 of the data processing device 12 in the application example 1 is realized by the following means.
[1460] In this invention, the server includes a face recognition unit, a customer attribute analysis unit, a product information database, a unit for generating POP content in real time, a display unit for displaying the generated POP content, a unit for recognizing customer pointing gestures, a unit for recognizing customer questions by voice, a unit for providing information based on the voice questions, a unit for personalizing information, a unit for collecting purchase data, a unit for updating an AI model based on the collected data, a unit for optimizing worker movement lines, a unit for visually guiding inventory locations, and a unit for linking with smart glasses or a head-mounted display. This makes it possible to optimize worker movement lines and provide real-time visual guidance on inventory locations at sites such as logistics centers and warehouses.
[1461] "Facial recognition means" is a technology that uses a camera to capture the face of a worker and analyzes attribute data such as age and gender based on the image.
[1462] "Customer attribute analysis means" is a technology that analyzes the tendencies and characteristics of specific customers based on attribute data acquired by face recognition means.
[1463] The "product information database" is a database that stores detailed information about each product or service, allowing users to search and reference the information they need.
[1464] "Means for generating POP content in real time" refers to technology for dynamically generating personalized POP content based on data collected in real time.
[1465] "Display means" refers to a device that displays generated POP content and information in real time, and includes display formats such as images, videos, and text.
[1466] The "means for recognizing pointing gestures" is a technology that uses a camera to detect the pointing gestures of workers and customers and analyzes their location information.
[1467] "Voice recognition means" refers to a technology that uses a microphone to capture the voice of a worker or customer and converts it into text data using voice recognition technology.
[1468] "Means for providing information based on voice questions" refers to a technology that acquires appropriate information from a server based on the content of a voice-recognized question and displays it on a display means.
[1469] "Means for personalizing information" refers to technology that generates individually optimized information and messages based on the attributes and behavioral data of customers and workers.
[1470] "Means for collecting purchasing data" refers to technology that collects data on customers' actual purchasing behavior and selected products and stores it in a database.
[1471] "Means for updating AI models based on collected data" refers to technologies for analyzing collected data and continuously training and optimizing AI models.
[1472] "Means to optimize worker movement" refers to technology that analyzes worker location information and movement in real time and proposes optimal movement routes and work procedures.
[1473] "Means for visually guiding inventory location" refers to technology that uses smart glasses or head-mounted displays within logistics centers to visually show workers the exact location of inventory.
[1474] "Means for linking with smart glasses or head-mounted displays" refers to technology that links smart glasses or head-mounted displays with the system to display information in real time and accept input.
[1475] This invention relates to a digital signage system for improving work efficiency in logistics centers and warehouses. This system uses facial recognition to analyze worker attributes, optimize worker movement lines, and visually guide inventory locations. Furthermore, by linking with smart glasses or head-mounted displays, work efficiency can be further improved.
[1476] System Configuration
[1477] This system is broadly composed of a server, terminals, and users. Each component is explained in detail below.
[1478] server
[1479] The server has the following functions:
[1480] 1. Facial Recognition Methods
[1481] The camera captures the worker's face and acquires attribute data such as age and gender, which is then analyzed using a facial recognition algorithm.
[1482] 2. Customer attribute analysis means
[1483] The system analyzes worker characteristics based on attribute data acquired through facial recognition, which is updated in real time and stored in a database.
[1484] 3. Product information database
[1485] This database comprehensively manages inventory information within the distribution center and is used to identify optimal inventory locations and routes.
[1486] 4. A way to generate POP content in real time
[1487] Generate personalized POP content in real time based on worker attribute information and inventory information.
[1488] 5. Optimizing traffic flow
[1489] Analyzes the real-time location information of workers and calculates and suggests the optimal movement route.
[1490] 6. Inventory location visual guide means
[1491] The exact location of inventory is displayed in real time on smart glasses or head-mounted displays.
[1492] 7. AI model update methods
[1493] The collected data is used to train and optimize the AI model, which is an algorithm designed to continuously improve work efficiency.
[1494] Terminal
[1495] The terminals are installed inside the logistics center and serve as an interface with workers.
[1496] 1. Display Means
[1497] Displays real-time POP content and traffic flow information in the form of images, videos, text, etc.
[1498] 2. Pointing Recognition Method
[1499] The camera is used to recognize the worker's pointing gesture, and detailed information is requested from the server based on that gesture.
[1500] 3. Voice Recognition Method
[1501] The worker's voice is captured using a microphone and converted into text data using voice recognition technology.
[1502] 4. Means of providing information on the display
[1503] Information acquired from the server based on questions received by voice is displayed on the display.
[1504] User
[1505] The users are workers in the logistics center and use the system as follows:
[1506] 1. Viewing content
[1507] Workers can visually check the content displayed on the screen and grasp the information necessary for their work.
[1508] 2. Questions and Pointing
[1509] Workers can ask questions by voice or point to specific locations on the screen to obtain information.
[1510] Specific examples
[1511] Example 1: Picking support
[1512] When picking items in a distribution center, the smart glasses optimize worker movement in real time, suggesting the most efficient route, and visually guiding workers to the exact location of inventory, significantly reducing work time.
[1513] Example prompt sentence:
[1514] "Please build an AI model that can suggest optimal movement paths and display inventory locations in real time to help a man in his 30s improve the efficiency of his picking work."
[1515] Example 2: Improving efficiency of warehousing operations
[1516] When receiving new inventory, the smart glasses will instruct workers in real time on where to place the inventory, supporting efficient stocking operations. Facial recognition is used to provide an individually optimized route and placement location.
[1517] Example prompt sentence:
[1518] "Propose an AI model for smart glasses that uses facial recognition to guide workers to the appropriate shelf location when receiving new inventory."
[1519] The flow of the specific processing in the application example 1 will be described with reference to FIG.
[1520] Step 1:
[1521] The server captures the worker's face through a camera. The input data is the image from the camera, and it analyzes attribute data such as age and gender using a facial recognition algorithm. The output is the analyzed worker's attribute data.
[1522] Step 2:
[1523] The server analyzes the characteristics of the workers based on the attribute data acquired in step 1. The input is customer attribute data, which is stored in a database and updated in real time. The output is the updated characteristic data.
[1524] Step 3:
[1525] The server searches and references inventory information from the product information database. The input is the worker's characteristic data and the inventory database, and the output is information on the optimal inventory location and route. This allows data to be collected to optimize worker movement lines.
[1526] Step 4:
[1527] The server generates the most suitable POP content for each worker in real time. The input is the worker's attribute data and inventory information, and the generated content is sent to the terminal. The output is the dynamically generated POP content.
[1528] Step 5:
[1529] The terminal displays the POP content sent from the server. The input is the POP content sent from the server, and the output is the content displayed on the display.
[1530] Step 6:
[1531] The terminal recognizes the worker's pointing gestures in real time. The input is video of the worker's movements captured by the terminal's built-in camera, and the output is the location information of the recognized pointing gesture. Based on this information, a request for detailed information is made to the server.
[1532] Step 7:
[1533] The terminal captures the worker's voice with a microphone and converts it into text data using voice recognition technology. The input is the worker's voice data, and the output is the converted text data.
[1534] Step 8:
[1535] The server acquires the appropriate information based on the question that has been recognized by voice and sends it to the terminal. The input is the text data that has been recognized by voice, and the output is the acquired information. This provides the necessary information to the worker.
[1536] Step 9:
[1537] The server trains and optimizes the AI model based on the collected data. The input is the collected purchase and operation data, and the output is an updated AI model. This updated model is used to guide future flow optimization and inventory location.
[1538] Step 10:
[1539] The terminal displays information sent from the server on smart glasses or a head-mounted display. The input is inventory location data and movement line data sent from the server, and the output is information displayed on the visual device. This allows workers to work efficiently.
[1540] Furthermore, an emotion engine that estimates the user's emotion may be further combined. That is, the identification processing unit 290 may estimate the user's emotion using the emotion identification model 59, and perform identification processing using the user's emotion.
[1541] This invention relates to a digital signage-type POP system that introduces optimal products and services in real time based on customer attribute information and emotional data. This system is used in retail stores and shopping malls to provide personalized information to individual customers.
[1542] System Configuration
[1543] This system is broadly composed of a "server," a "terminal," and a "user." Details of each component are explained below.
[1544] server
[1545] The server has the following functions:
[1546] 1. Facial Recognition Methods:
[1547] The server captures the customer's face through a camera and obtains high-resolution image data.
[1548] Facial recognition algorithms are used to obtain customer attribute data such as age, gender, and number of customers.
[1549] 2. Customer attribute analysis tools and sentiment engine:
[1550] Attribute data analyzed from facial recognition data is stored in a database and updated in real time.
[1551] The emotion engine analyzes customer emotions from captured facial images and tone of voice, and generates emotion data.
[1552] 3. Product information database and personalized message generation:
[1553] The product information database stores detailed information about each product and service, allowing users to search and reference the information they need.
[1554] By referencing customer attribute data and sentiment data, the system selects the most suitable products and services and dynamically generates POP content.
[1555] Add personalized messages from your membership database and purchase history to generate individual messages for specific customers.
[1556] 4. Learning and optimization:
[1557] The server analyzes the collected customer behavioral data, purchasing data, and emotional data to evaluate the effectiveness of POP content.
[1558] Based on the results of the effectiveness evaluation, the AI model will be updated to improve the accuracy of future content generation algorithms.
[1559] Terminal
[1560] The terminal is installed inside the store and serves as an interface with customers.
[1561] 1. Display means:
[1562] Real-time POP content sent from the server is displayed on the display.
[1563] The content can be displayed in the form of images, videos, text, etc.
[1564] 2. Pointing recognition method:
[1565] The camera recognizes the customer's pointing movements in real time.
[1566] A request is made to the server for detailed information about the product pointed to by the customer, and the received information is displayed on the screen.
[1567] 3. Question and Answer Function:
[1568] The customer's voice questions are captured using a microphone and converted into text data using voice recognition technology.
[1569] Information based on the question is obtained from the server and displayed on the screen.
[1570] User
[1571] The user is a customer of the store and performs the following operations.
[1572] 1. Viewing content:
[1573] Users view the POP content displayed on the display and check information about products and services that interest them.
[1574] 2. Questions and Pointing:
[1575] Users can ask for details about products they are interested in by voice or by pointing at specific products on the screen.
[1576] The question and information about the product you point at will be displayed on the screen to help you make a purchasing decision.
[1577] Specific examples
[1578] Example 1: Parent-child visit scenario
[1579] A parent and child visit a shopping mall. The server captures the faces of the parent and child via camera and analyzes that the parent is in their late 30s and the child is around 7 years old. When the emotion engine recognizes the parent's facial expression as being interested and having fun, the device displays information mainly about toys and events for children. When the parent asks, "Where are the toy sales?", the device sends the voice question and analysis results to the server and displays the corresponding information on the display.
[1580] Example 2: Checking details of individual products and emotion recognition
[1581] When a customer points at an image of a product displayed on the display, the device recognizes the pointing gesture and requests detailed information about that product from the server. The server retrieves the product information, and the emotion engine also analyzes the reaction. For example, if the customer shows a surprised expression, information about a special campaign will be displayed. The customer then checks the product's price, specifications, and stock status before making a purchasing decision.
[1582] In this way, combining emotion engines enables more personalized and effective information provision to customers. Using emotion data, it is possible to respond to customers' instantaneous reactions and provide a more personalized shopping experience. This configuration makes it possible to maximize sales promotion effects and improve customer satisfaction.
[1583] The processing flow will be explained below.
[1584] Program processing steps
[1585] Server Processing Steps
[1586] Step 1:
[1587] The server captures the faces of customers in the store through a camera and obtains high-resolution image data.
[1588] Step 2:
[1589] The server analyzes the acquired facial images using an AI algorithm to estimate attribute data such as the customer's age, gender, and number of people.
[1590] Step 3:
[1591] The server updates the analysis results to a database in real time and stores customer attribute data.
[1592] Step 4:
[1593] The server uses an emotion engine to analyze facial images, tone of voice, and other factors to generate customer emotion data.
[1594] Step 5:
[1595] The server stores the generated emotion data in a database and updates it in real time.
[1596] Step 6:
[1597] The server searches and selects the most suitable products and services from a product information database based on customer attribute data and emotional data.
[1598] Step 7:
[1599] The server dynamically generates POP content in real time based on the selection results.
[1600] Step 8:
[1601] The server references the member database and purchase history and generates personalized messages as needed.
[1602] Step 9:
[1603] The server transmits the generated POP content to the terminal.
[1604] Step 10:
[1605] The server analyzes the collected customer behavioral data, purchasing data, and emotional data to evaluate the effectiveness of the POP content.
[1606] Step 11:
[1607] Based on the effectiveness evaluation results, the server updates the AI model and improves the accuracy of future content generation algorithms.
[1608] Terminal processing steps
[1609] Step 1:
[1610] The terminal receives the POP content sent from the server.
[1611] Step 2:
[1612] The device displays the received content on the display, which can be in the form of images, videos, text, etc.
[1613] Step 3:
[1614] The device uses a camera to recognize the customer's pointing movements in real time.
[1615] Step 4:
[1616] The terminal requests information about the product pointed at by the customer from the server.
[1617] Step 5:
[1618] The terminal receives the detailed product information sent from the server and displays it on the display.
[1619] Step 6:
[1620] The terminal uses a microphone to capture the customer's voice asking a question.
[1621] Step 7:
[1622] The device uses voice recognition technology to convert the voice into text data and send it to the server.
[1623] Step 8:
[1624] The terminal displays the response information sent from the server on the display.
[1625] User operation steps
[1626] Step 1:
[1627] The user views the POP content displayed on the display.
[1628] Step 2:
[1629] Users can input questions about products or services that interest them by voice.
[1630] Step 3:
[1631] The user points to a particular product on the display to see more information about it.
[1632] Step 4:
[1633] Users make purchasing decisions based on the displayed information.
[1634] Specific examples
[1635] Example 1: Parent-child visit scenario
[1636] Step 1: A parent and child visit a shopping mall. The server captures the faces of the parent and child through a camera and analyzes them.
[1637] Step 2: The server estimates the age and gender of the parents and children and stores the attribute data in a database.
[1638] Step 3: The server analyzes the facial image and tone of voice using an emotion engine to generate emotion data.
[1639] Step 4: The server selects toys and event information for children based on the generated attribute data and emotion data, and generates POP content.
[1640] Step 5: The terminal receives the POP content and displays it on the display.
[1641] Step 6: The parent asks, "Where are the toy sales?" The device captures the audio and sends it to the server.
[1642] Step 7: The server generates toy department information based on the query and sends it to the terminal.
[1643] Step 8: The terminal displays the received information on its display.
[1644] Example 2: Checking details of individual products and emotion recognition
[1645] Step 1: The customer points to the product image on the display.
[1646] Step 2: The device recognizes the pointing gesture and requests detailed information about the product from the server.
[1647] Step 3: The server retrieves the product details and analyzes the customer's emotions using the emotion engine.
[1648] Step 4: The server adds special campaign information based on the product details and customer sentiment data and sends it to the terminal.
[1649] Step 5: The terminal displays the received information on the screen. The customer checks the product price, specifications, and stock information and makes a purchasing decision.
[1650] In this way, combining emotion engines enables effective provision of personalized information to customers. Using emotion data, it is possible to respond to customers' instantaneous reactions and provide a more personalized shopping experience. This configuration makes it possible to maximize sales promotion effects and improve customer satisfaction.
[1651] Example 2
[1652] Next, a description will be given of Example 2. In the following description, the data processing device 12 will be referred to as a "server" and the headset type terminal 314 will be referred to as a "terminal."
[1653] Conventional digital signage and POP systems have had issues with providing personalized information to customers and not being able to dynamically recommend products based on customer emotions. As a result, it has been difficult to provide effective information to customers and increase their desire to purchase. It has also been difficult to optimize generative models using customer behavioral and emotional data.
[1654] The identification process by the identification processing unit 290 of the data processing device 12 in the second embodiment is realized by the following means. In this invention, the server includes a face recognition means, a customer attribute analysis means, and a means for acquiring emotion data. This makes it possible to recommend optimal products and services in real time based on customer attribute information and emotion data. Furthermore, by evaluating the effectiveness of the generated POP content and updating the generation model based on collected data, it becomes possible to continuously improve the accuracy of information provision.
[1655] A "facial recognition means" is an algorithm or device that uses a camera to capture facial images of customers and extract attribute data such as age, gender, and number of people.
[1656] "Customer attribute analysis means" refers to a system or program that analyzes customer attributes based on acquired facial recognition data and stores and updates that data in a database.
[1657] The "product information database" is a data storage system that stores detailed information about products and services in the store and allows for searching and referencing as needed.
[1658] The "means for acquiring emotional data" refers to an algorithm or device that analyzes the customer's facial image and tone of voice to generate emotional data such as joy, interest, or displeasure.
[1659] "Means for generating POP content in real time" refers to a system or program that dynamically generates content that recommends products and services based on customer attribute information and emotional data.
[1660] The "display means" is a display device that visually presents the POP content transmitted from the server to customers in the store.
[1661] The "means for recognizing the customer's pointing behavior" refers to a system or program that uses a sensor such as a camera to detect the customer's pointing behavior in real time and analyzes that information.
[1662] The "means for recognizing customer questions by voice" is a system or program that uses a microphone to capture the voice of a customer's question and converts it into text data using voice recognition technology.
[1663] The "means for providing information based on a voice question" is a system or program that retrieves related information from a server based on a customer's voice question and presents it on a display.
[1664] "Means for personalizing information" refers to a system or program that generates and displays personalized messages optimized for specific customers based on customer attribute data and purchase history.
[1665] A "means for collecting purchasing data" is a system or device for collecting data on customer purchasing behavior and storing it in a database.
[1666] The "means for updating the generative model based on collected data" refers to a system or algorithm that analyzes collected customer behavioral and emotional data and optimizes and updates the generative model.
[1667] MODE FOR CARRYING OUT THE INVENTION
[1668] This invention relates to a digital signage-type POP system that introduces optimal products and services in real time based on customer attribute information and emotional data. By recognizing the customer's face, analyzing their emotions, and recommending products and services based on that, personalized information is provided.
[1669] server
[1670] The server has the following features:
[1671] 1. Facial Recognition Methods:
[1672] The server captures customer faces through a camera, obtains high-resolution image data, and uses a facial recognition algorithm to extract customer attributes such as age, gender, and number of customers.
[1673] 2. Customer attribute analysis means:
[1674] Based on facial recognition data, customer attribute information is analyzed, and the data is stored in a database and updated in real time.
[1675] 3. How to get emotion data:
[1676] The emotion engine analyzes the customer's emotions using facial images, tone of voice, and other data, and generates emotion data. For example, if a customer smiles, the emotion is analyzed as "happiness."
[1677] 4. Product Information Database:
[1678] The product information database stores detailed information about each product and service, which can be searched and referenced as needed. Content is generated that recommends optimal products and services based on customer attribute data and emotional data.
[1679] 5. Personalized message generation:
[1680] Add individual personalized messages from your membership database and purchase history to generate individual messages for specific customers.
[1681] 6. Learning and optimization:
[1682] The server analyzes collected customer behavioral, purchasing, and emotional data to evaluate the effectiveness of the POP content. Based on the results of the effectiveness evaluation, the generative AI model is updated, improving the accuracy of future content generation algorithms.
[1683] Terminal
[1684] The terminals are installed in the store and act as the interface with the customer:
[1685] 1. Display means:
[1686] Real-time POP content sent from the server is displayed on the display. Content can be displayed in the form of images, videos, text, etc.
[1687] 2. Pointing recognition method:
[1688] The device's camera recognizes the customer's pointing gestures in real time, requests detailed information about the product the customer is pointing at from the server, and displays the received information on the display.
[1689] 3. Question and Answer Function:
[1690] The device's microphone captures the customer's voice and converts it into text data using speech recognition technology. Information based on the question is retrieved from the server and displayed on the screen.
[1691] User
[1692] The user is a customer in the store and performs the following actions:
[1693] 1. Viewing content:
[1694] Users view the POP content displayed on the display and check information about products and services that interest them.
[1695] 2. Questions and Pointing:
[1696] Users can ask for details about products they are interested in by voice or by pointing at a specific product on the screen. The content of the question or information about the product they are pointing at is then displayed on the screen.
[1697] Specific examples
[1698] Example 1: Parent-child visit scenario
[1699] A parent and child visit a shopping mall. The server captures the faces of the parent and child via camera and analyzes that the parent is in their late 30s and the child is about 7 years old. When the emotion engine recognizes the parent's facial expression as being interested and having fun, the device displays information mainly about toys and events for children. When the parent asks, "Where are the toy sales?", the device sends the voice question and analysis results to the server and displays the corresponding information on the display.
[1700] Example 2: Checking details of individual products and emotion recognition
[1701] When a customer points at an image of a product displayed on the display, the device recognizes the pointing gesture and requests detailed information about that product from the server. The server retrieves the product information, and the emotion engine also analyzes the reaction. For example, if the customer shows a surprised expression, information about a special campaign will be displayed. The customer then checks the product's price, specifications, and stock status before making a purchasing decision.
[1702] Examples of prompt statements
[1703] 1. "When a customer points to an image of a product on the display, explain how the device displays more information about that product."
[1704] 2. "Please explain specifically how the system will display toys and information about children's events when a parent and child visit the store."
[1705] This invention makes it possible to provide customers with personalized information, which is expected to increase their purchasing motivation. It also responds to customers' instantaneous reactions, providing a more personalized shopping experience.
[1706] The flow of the identification process in the second embodiment will be described with reference to FIG.
[1707] Step 1:
[1708] The server captures the faces of customers in the store using a camera and obtains high-resolution image data. The obtained image data is input into a facial recognition algorithm to extract customer attribute data such as age, gender, and number of customers. The output attribute data is stored in a database.
[1709] Specific behavior:
[1710] When a customer enters a store, the server automatically acquires data from the camera and begins facial recognition. For example, if a woman in her 30s and a man in his 40s visit the store together, their attribute data will be stored in a database.
[1711] Step 2:
[1712] The server inputs the acquired facial image and tone of voice into an emotion engine to generate customer emotion data. The emotion engine analyzes the customer's emotional state, such as joy, interest, or displeasure, and outputs the data. The analyzed emotion data is stored in a database.
[1713] Specific behavior:
[1714] If a customer looks at the display and smiles, the server interprets the smile as "happiness." If the customer's voice is high-pitched and excited, the emotion engine recognizes it as "interested."
[1715] Step 3:
[1716] The server references the product information database and selects the most suitable products and services based on the acquired customer attribute data and emotion data. The selected information is generated as dynamic POP content in real time. The generated POP content is then saved back into the database.
[1717] Specific behavior:
[1718] For example, if a woman in her 30s expresses interest, the server will generate content recommending the latest beauty products, adding special campaign information for related products based on the customer's past purchases.
[1719] Step 4:
[1720] The terminal displays the POP content sent from the server. The content is displayed on the display in the form of images, videos, and text. When a customer points at a specific product on the screen, the camera recognizes the pointing gesture and requests detailed information from the server.
[1721] Specific behavior:
[1722] When a customer points to an image of a toy on the display, the device automatically retrieves and displays detailed information about the product.
[1723] Step 5:
[1724] When a customer asks a question to the display, the device's microphone captures the voice of the question and converts it into text data using voice recognition technology. The converted text data is sent to the server, which returns information based on the question. The acquired information is then displayed on the display.
[1725] Specific behavior:
[1726] When a parent asks, "Where is the toy sale?", the voice question and analysis results are sent to the server, and the corresponding information is displayed on the screen.
[1727] Step 6:
[1728] The server analyzes collected customer behavioral, purchasing, and emotional data to evaluate the effectiveness of POP content. The results of the evaluation are used to update the AI model, improving the accuracy of future content generation algorithms.
[1729] Specific behavior:
[1730] Past data can be used to analyze which content was most effective. For example, data on women in their 30s can be used to learn that a particular beauty product was well-received, and this can be reflected in future recommendations.
[1731] (Application example 2)
[1732] Next, a description will be given of Application Example 2. In the following description, the data processing device 12 will be referred to as a "server" and the headset type terminal 314 will be referred to as a "terminal."
[1733] In today's retail industry, there is a demand for personalized information provision to each individual customer. Furthermore, there is a lack of methods for using emotional data to suggest optimal products and services based on the customer's real-time situation. As a result, customer satisfaction and purchasing motivation have not been sufficiently improved. Therefore, there is a need to provide a system that can recommend optimal products and services in real time based on the customer's attribute information and emotional data.
[1734] The specific processing by the specific processing unit 290 of the data processing device 12 in the application example 2 is realized by the following means.
[1735] In this invention, the server includes a face recognition means, a customer attribute analysis means, a product information database, a means for generating POP content in real time, a display device for displaying the generated POP content, a means for recognizing customer pointing gestures, a means for recognizing customer questions by voice, a means for providing information based on the voice questions, a means for personalizing information, a means for collecting purchase data, a means for updating an AI model based on the collected data, a sentiment analysis means, a means for selecting optimal products and services based on the sentiment data, a means for detecting customer faces in real time using a camera and analyzing their sentiments, and a means for displaying related product information on a display device based on the analysis results. This makes it possible to propose optimal products and services based on the customer's sentiments and attributes in real time, providing a personalized shopping experience and improving customer satisfaction and purchasing motivation.
[1736] A "face recognition means" is a means of capturing a customer's face using a camera and obtaining attribute data such as the customer's age, gender, and number of people from the high-resolution image data.
[1737] The "customer attribute analysis means" is a means for analyzing attribute data acquired by the face recognition means and updating it in real time.
[1738] A "product information database" is a database that stores detailed information about each product or service, allowing users to search and reference the information they need.
[1739] "Means for generating POP content in real time" refers to means for dynamically generating POP content by referencing customer attribute data and emotional data, selecting the most suitable products and services.
[1740] The "display device" is a device for displaying real-time POP content transmitted from the server on a display.
[1741] The "means for recognizing pointing actions" is a means for recognizing a customer's pointing actions in real time through a camera and requesting detailed information about the pointed product from the server.
[1742] The "voice recognition means" is a means of capturing the customer's voice question using a microphone and converting it into text data using voice recognition technology.
[1743] The "means for providing information based on a voice question" is a means for obtaining related information from a server based on the content of a customer's voice question and displaying it on a display.
[1744] The "means for personalizing information" refers to a means for adding personalized messages from a member database or purchase history and generating individual messages for specific customers.
[1745] "Means for collecting purchasing data" refers to the means for collecting and analyzing customer behavioral data, purchasing data, and emotional data.
[1746] "Means for updating the AI model" refers to a means for evaluating the effectiveness of POP content based on collected data and updating the AI model based on the results of that evaluation.
[1747] The "emotion analysis means" is a means for analyzing the customer's emotions from acquired facial images, tone of voice, etc., and generating emotion data.
[1748] "Means for selecting optimal products and services based on emotional data" refers to means for selecting the most appropriate products and services at that time by referring to emotional data.
[1749] "Means for detecting a customer's face in real time using a camera and analyzing emotions" refers to means for capturing a customer's face in real time using a camera and analyzing the customer's emotions from the facial image.
[1750] The "means for displaying related product information on a display device based on the analysis results" refers to means for displaying related product information on a display based on the analyzed customer emotion data and attribute data.
[1751] The present invention relates to a system that introduces optimal products and services in real time based on customer attribute information and emotion data, and is configured as follows.
[1752] server
[1753] The server has the following functions:
[1754] 1. Facial Recognition Methods
[1755] The server captures the customer's face through a camera, obtains high-resolution image data, and uses a facial recognition algorithm to obtain customer attribute data such as age, gender, and number of customers.
[1756] 2. Customer attribute analysis means
[1757] Attribute data acquired through facial recognition is analyzed and updated in real time, ensuring that basic customer information is always kept up to date.
[1758] 3. Product information database
[1759] The product information database stores detailed information about each product and service, allowing you to search and reference the information you need.
[1760] 4. A way to generate POP content in real time
[1761] By referencing customer attribute data and emotional data, the system selects the most suitable products and services and dynamically generates POP content. For example, it uses prompts such as "display the most suitable products for male customers in their 30s who have a surprised expression."
[1762] 5. Display device means
[1763] Real-time POP content sent from the server is displayed on the display. Content can be displayed in the form of images, videos, text, etc.
[1764] 6. Methods for Recognizing Pointing Actions
[1765] The camera recognizes the customer's pointing gestures in real time, requests detailed information about the product the customer is pointing at from the server, and displays the received information on the display.
[1766] 7. Voice recognition
[1767] The customer's voice questions are captured using a microphone and converted into text data using voice recognition technology.
[1768] 8. Means of Providing Information Based on Voice Queries
[1769] Based on the content of the customer's voice question, related information is obtained from the server and displayed on the screen.
[1770] 9. How we personalize your information
[1771] Add personalized messages from your membership database and purchase history to generate individual messages for specific customers.
[1772] 10. Means of collecting purchasing data
[1773] Collect and analyze customer behavioral, purchasing, and emotional data.
[1774] 11. How to update your AI model
[1775] The effectiveness of POP content is evaluated based on the collected data, and the AI model is updated based on the evaluation results.
[1776] 12. Sentiment analysis tool
[1777] The system analyzes customer emotions from acquired facial images and tone of voice, and generates emotional data.
[1778] 13. A way to select the best products and services based on sentiment data
[1779] It is a means of referring to emotional data to select the most appropriate product or service at any given time.
[1780] 14. Real-time facial detection and emotion analysis using cameras
[1781] The camera captures the customer's face in real time and analyzes their emotions from the facial image.
[1782] 15. Means for displaying related product information on a display device based on the analysis results
[1783] Based on the analyzed customer emotional data and attribute data, relevant product information is displayed on the screen.
[1784] Terminal
[1785] The terminal is installed inside the store and serves as an interface with customers.
[1786] Display Means
[1787] Real-time POP content sent from the server is displayed on the display, allowing customers to directly view information about products and services that are optimized for them.
[1788] Pointing recognition method
[1789] The camera recognizes customers' pointing gestures in real time. For example, if a customer points to a product to say, "I'd like to know more about this product," detailed information will be displayed on the screen.
[1790] Voice question answering function
[1791] The microphone is used to capture the customer's voice asking a question. If a customer asks, "Are there any special offers for this product?", the content is sent to the server and the appropriate information is displayed.
[1792] User
[1793] The user is a customer of the store.
[1794] Viewing content
[1795] Users view the POP content displayed on the display and check information about products and services that interest them.
[1796] Questions and pointing
[1797] Users can ask for details about products they are interested in by voice or by pointing at a specific product on the screen. The question or information about the product they are pointing at is then displayed on the screen to help them make a purchasing decision.
[1798] Specific examples
[1799] By inputting prompt statements such as, "Please tell me what product information should be provided to a female customer who looks happy," into the generative AI model, the most appropriate content is displayed.
[1800] In this way, the system of the present invention improves customer satisfaction and maximizes the effectiveness of store sales promotion by providing real-time information based on customer attribute information and emotional data.
[1801] The flow of the specific processing in the application example 2 will be described with reference to FIG.
[1802] Step 1:
[1803] The server captures the customer's face through a camera and obtains high-resolution image data.
[1804] Input: Video data from the camera
[1805] Output: Customer's facial image data
[1806] Specific operation: The camera captures video and sends the video data to the server, which then detects faces in the video and generates facial image data.
[1807] Step 2:
[1808] The server uses a facial recognition algorithm to obtain attribute data such as customer age, gender, and number of customers.
[1809] Input: Customer's facial image data
[1810] Output: Customer attribute data (age, gender, number of people)
[1811] Specific operation: The server runs a facial recognition algorithm, analyzes the facial image data, and extracts attribute information such as age, gender, and number of people.
[1812] Step 3:
[1813] The server stores the acquired attribute data in a database and updates it in real time.
[1814] Input: Customer attribute data
[1815] Output: Updated customer attribute information in the database
[1816] Specific operation: The server stores the attribute data in a database and updates existing customer information as necessary.
[1817] Step 4:
[1818] The server uses an emotion engine to analyze the customer's emotions from acquired facial images and tone of voice, and generates emotion data.
[1819] Input: Facial image data, audio data
[1820] Output: Customer sentiment data
[1821] Specific operation: The emotion engine analyzes facial images and voice data, infers emotions from the customer's facial expressions and tone of voice, and generates that data.
[1822] Step 5:
[1823] The server references customer attribute data and emotion data, selects the most suitable products and services from a product information database, and dynamically generates POP content.
[1824] Input: Customer attribute data, customer sentiment data
[1825] Output: Generated POP content
[1826] Specific operation: The server searches the product information database based on customer attribute information and emotion data, extracts information on related products and services, and generates POP content based on this.
[1827] Step 6:
[1828] The terminal displays the real-time POP content sent from the server on the display.
[1829] Input: Generated POP content
[1830] Output: The content shown on the display
[1831] Specific operation: Receives POP content sent from the server and displays information such as images, videos, and text on the display.
[1832] Step 7:
[1833] The device uses a camera to recognize the customer's pointing movements in real time.
[1834] Input: Video data from the camera
[1835] Output: Pointing motion detection information
[1836] Specific operation: The camera captures the customer's pointing action and sends it to the server for pointing recognition.
[1837] Step 8:
[1838] The server retrieves detailed information about the product pointed to by the customer from the database and displays it on the display.
[1839] Input: Pointing motion detection information
[1840] Output: Product details
[1841] Specific operation: Based on the pointing gesture, the server retrieves the corresponding product information from the database and sends it to the terminal, which then displays the detailed information on the screen.
[1842] Step 9:
[1843] The terminal uses a microphone to capture the customer's voice asking a question and converts it into text data using voice recognition technology.
[1844] Input: Audio data
[1845] Output: Text data of the audio
[1846] How it works: A microphone captures the customer's question, and a speech recognition algorithm on the server converts it into text data.
[1847] Step 10:
[1848] The server acquires information based on the voice question content and displays it on a display.
[1849] Input: Text data of speech
[1850] Output: Information corresponding to the question
[1851] Specific operation: The server analyzes the text data of the voice question, retrieves relevant information from the database, and sends it to the device, which then displays the answer to the question on the display.
[1852] Step 11:
[1853] The server generates personalized messages from a member database and purchase history, and provides individual messages to specific customers.
[1854] Input: Member data, purchase history
[1855] Output: Personalized message
[1856] Specific operation: The server references the member database and purchase history, generates a personalized message tailored to the specific customer, and provides it via the terminal.
[1857] Step 12:
[1858] The server collects and analyzes customer behavioral data, purchasing data, and emotional data.
[1859] Input: Customer behavior data, purchasing data, emotional data
[1860] Output: Analysis data
[1861] Specific operation: The server analyzes the collected data in real time and uses the results to generate the next POP content and respond to customers.
[1862] Step 13:
[1863] The server evaluates the effectiveness of the POP content based on the collected data and updates the AI model based on the evaluation results.
[1864] Input: Customer behavior data, purchasing data, emotional data
[1865] Output: Updated AI model
[1866] Specific operation: The server analyzes the collected data and adjusts the parameters of the AI model to improve the accuracy of subsequent content generation algorithms.
[1867] The specific processing unit 290 transmits the result of the specific processing to the headset type terminal 314. In the headset type terminal 314, the control unit 46A causes the speaker 240 and the display 343 to output the result of the specific processing. The microphone 238 acquires audio indicating a user input regarding the result of the specific processing. The control unit 46A transmits audio data indicating the user input acquired by the microphone 238 to the data processing device 12. In the data processing device 12, the specific processing unit 290 acquires the audio data.
[1868] The data generation model 58 is a so-called generative AI (Artificial Intelligence). An example of the data generation model 58 is ChatGPT (Internet Search<URL: https: / / openai.com / blog / chatgpt> ), Gemini (Internet search <url: https: gemini.google.com ?hl="ja">) and other generation AIs. The data generation model 58 is obtained by performing deep learning on a neural network. A prompt including an instruction is input to the data generation model 58, and inference data such as voice data indicating voice, text data indicating text, and image data indicating an image is also input. The data generation model 58 performs inference on the input inference data in accordance with the instruction indicated by the prompt, and outputs the inference result in a data format such as voice data and text data. Here, inference refers to, for example, analysis, classification, prediction, and / or summarization.
[1869] In the above embodiment, an example was given in which the specific processing is performed by the data processing device 12, but the technology of the present disclosure is not limited to this, and the specific processing may be performed by the headset type terminal 314.
[1870] [Fourth embodiment]
[1871] FIG. 7 shows an example of the configuration of a data processing system 410 according to the fourth embodiment.
[1872] 7, a data processing system 410 includes a data processing device 12 and a robot 414. An example of the data processing device 12 is a server.
[1873] The data processing device 12 includes a computer 22, a database 24, and a communication I / F 26. The computer 22 is an example of a "computer" according to the technology of the present disclosure. The computer 22 includes a processor 28, a RAM 30, and a storage 32. The processor 28, the RAM 30, and the storage 32 are connected to a bus 34. The database 24 and the communication I / F 26 are also connected to the bus 34. The communication I / F 26 is connected to a network 54. Examples of the network 54 include a WAN (Wide Area Network) and / or a LAN (Local Area Network).
[1874] The robot 414 includes a computer 36, a microphone 238, a speaker 240, a camera 42, a communication I / F 44, and a control target 443. The computer 36 includes a processor 46, a RAM 48, and a storage 50. The processor 46, the RAM 48, and the storage 50 are connected to a bus 52. The microphone 238, the speaker 240, the camera 42, and the control target 443 are also connected to the bus 52.
[1875] The microphone 238 receives instructions and the like from the user 20 by receiving voice uttered by the user 20. The microphone 238 captures the voice uttered by the user 20, converts the captured voice into audio data, and outputs it to the processor 46. The speaker 240 outputs audio in accordance with instructions from the processor 46.
[1876] Camera 42 is a small digital camera equipped with an optical system including a lens, aperture, and shutter, and an imaging element such as a CMOS (Complementary Metal-Oxide-Semiconductor) image sensor or a CCD (Charge Coupled Device) image sensor, and captures images of the surroundings of user 20 (for example, an imaging range defined by an angle of view equivalent to the field of vision of a typical healthy person).
[1877] The communication I / F 44 is connected to a network 54. The communication I / Fs 44 and 26 are responsible for the exchange of various information between the processor 46 and the processor 28 via the network 54. The exchange of various information between the processor 46 and the processor 28 using the communication I / Fs 44 and 26 is carried out in a secure state.
[1878] The control object 443 includes a display device, LEDs in the eyes, and motors for driving the arms, hands, and feet. The posture and gestures of the robot 414 are controlled by controlling the motors of the arms, hands, and feet. Some of the emotions of the robot 414 can be expressed by controlling these motors. In addition, the facial expressions of the robot 414 can also be expressed by controlling the light emission state of the LEDs in the eyes of the robot 414.
[1879] Fig. 8 shows an example of the main functions of the data processing device 12 and the robot 414. As shown in Fig. 8, in the data processing device 12, a specific process is performed by the processor 28. A specific process program 56 is stored in the storage 32.
[1880] The specific processing program 56 is an example of a "program" according to the technology of the present disclosure. The processor 28 reads the specific processing program 56 from the storage 32 and executes the read specific processing program 56 on the RAM 30. The specific processing is realized by the processor 28 operating as a specific processing unit 290 in accordance with the specific processing program 56 executed on the RAM 30.
[1881] The storage 32 stores a data generation model 58 and an emotion identification model 59. The data generation model 58 and the emotion identification model 59 are used by the identification processing unit 290.
[1882] In the robot 414, the processor 46 performs the reception output process. A reception output program 60 is stored in the storage 50. The processor 46 reads the reception output program 60 from the storage 50 and executes the read reception output program 60 on the RAM 48. The reception output process is realized by the processor 46 operating as the control unit 46A in accordance with the reception output program 60 executed on the RAM 48.
[1883] Next, a description will be given of the specific processing performed by the specific processing unit 290 of the data processing device 12. In the following description, the data processing device 12 will be referred to as a "server" and the robot 414 will be referred to as a "terminal."
[1884] This invention relates to a digital signage-type POP system that introduces optimal products and services in real time based on customer attribute information. This system is used in retail stores and shopping malls to provide personalized information to individual customers.
[1885] System Configuration
[1886] This system is broadly composed of a "server," a "terminal," and a "user." Details of each component are explained below.
[1887] server
[1888] The server has the following functions:
[1889] 1. Facial Recognition Methods:
[1890] The server captures customers' faces through a camera and obtains attribute data such as age, gender, and number of people.
[1891] Using high-resolution video data, customer attributes are analyzed using facial recognition algorithms.
[1892] 2. Customer attribute analysis means:
[1893] Based on facial recognition, analyzed attribute data is stored in a database and updated in real time.
[1894] The analysis results are used to provide promotional information to each customer.
[1895] 3. Product information database and personalized message generation:
[1896] The product information database stores detailed information about each product and service, allowing users to search and reference the information they need.
[1897] Customer attribute data is referenced to select the most suitable products and services and dynamically generate POP content.
[1898] Add personalized messages from your membership database and purchase history to generate individual messages for specific customers.
[1899] 4. Learning and optimization:
[1900] The server analyzes collected customer behavioral and purchasing data and updates the AI model to optimize the effectiveness of POP content.
[1901] We continuously collect data and improve our algorithms to generate optimal content.
[1902] Terminal
[1903] The terminal is installed inside the store and serves as an interface with customers.
[1904] 1. Display means:
[1905] Real-time POP content sent from the server is displayed on the display.
[1906] The content can be displayed in the form of images, videos, text, etc.
[1907] 2. Pointing recognition method:
[1908] The camera recognizes the customer's pointing movements in real time.
[1909] A request is made to the server for detailed information about the product pointed to by the customer, and the received information is displayed on the screen.
[1910] 3. Question and Answer Function:
[1911] The customer's voice questions are captured using a microphone and converted into text data using voice recognition technology.
[1912] Information based on the question is obtained from the server and displayed on the screen.
[1913] User
[1914] The user is a customer of the store and performs the following operations.
[1915] 1. Viewing content:
[1916] Users view the POP content displayed on the display and check information about products and services that interest them.
[1917] 2. Questions and Pointing:
[1918] Users can ask for details about products they are interested in by voice or by pointing at specific products on the screen.
[1919] The question and information about the product you point at will be displayed on the screen to help you make a purchasing decision.
[1920] Specific examples
[1921] Example 1: Parent and child visitor scenario
[1922] A parent and child visit a shopping mall. The server captures the faces of the parent and child via camera and determines that the parent is in their late 30s and the child is about 7 years old. Based on this, the device displays information mainly about toys and events for children. When the parent asks, "Where are the toy sales?", the device recognizes the voice question, retrieves the answer from the server, and displays it on the display.
[1923] Example 2: Checking details of individual products
[1924] The customer points at the image of a product displayed on the display. The terminal recognizes the pointing gesture and retrieves detailed information about that product from the server. The product's price, specifications, stock status, and other information are displayed on the display in real time.
[1925] As described above, this system can provide optimal information in real time based on customer attributes and behavior, thereby enhancing sales promotion effectiveness. Furthermore, by continually updating the AI model based on collected data, it is possible to continuously generate effective content.
[1926] The processing flow will be explained below.
[1927] Program processing steps
[1928] Server Processing Steps
[1929] Step 1:
[1930] The server captures the faces of customers in the store through a camera and obtains high-resolution image data.
[1931] Step 2:
[1932] The server analyzes the acquired facial images using an AI algorithm to estimate attribute data such as the customer's age, gender, and number of people.
[1933] Step 3:
[1934] The server updates the analysis results to a database and stores customer attribute data in real time.
[1935] Step 4:
[1936] The server searches for the most suitable products and services from a product information database based on customer attribute data.
[1937] Step 5:
[1938] The server dynamically generates POP content in real time based on the search results.
[1939] Step 6:
[1940] The server references the member database and purchase history and generates personalized messages as needed.
[1941] Step 7:
[1942] The server transmits the generated POP content to the terminal.
[1943] Step 8:
[1944] The server analyzes the collected customer behavior data and purchase data to evaluate the effectiveness of the POP content.
[1945] Step 9:
[1946] Based on the effectiveness evaluation results, the server updates the AI model and improves the accuracy of future content generation algorithms.
[1947] Terminal processing steps
[1948] Step 1:
[1949] The terminal receives the POP content sent from the server.
[1950] Step 2:
[1951] The device displays the received content on the display, which can be in the form of images, videos, text, etc.
[1952] Step 3:
[1953] The device uses a camera to recognize the customer's pointing movements in real time.
[1954] Step 4:
[1955] The terminal requests information about the product pointed at by the customer from the server.
[1956] Step 5:
[1957] The terminal receives the detailed product information sent from the server and displays it on the display.
[1958] Step 6:
[1959] The terminal uses a microphone to capture the customer's voice asking a question.
[1960] Step 7:
[1961] The device uses voice recognition technology to convert the voice into text data and send it to the server.
[1962] Step 8:
[1963] The terminal displays the response information sent from the server on the display.
[1964] User operation steps
[1965] Step 1:
[1966] The user views the POP content displayed on the display.
[1967] Step 2:
[1968] Users can input questions about products or services that interest them by voice.
[1969] Step 3:
[1970] The user points to a particular product on the display to see more information about it.
[1971] Step 4:
[1972] Users make purchasing decisions based on the displayed information.
[1973] Specific examples
[1974] Example 1: Parent-child visit scenario
[1975] Step 1: A parent and child visit a shopping mall. The server captures the faces of the parent and child through a camera and analyzes them.
[1976] Step 2: The server estimates the age and gender of the parents and children and updates the attribute data in the database.
[1977] Step 3: The server selects toys and event information for children based on the parent-child attributes and generates POP content.
[1978] Step 4: The terminal receives the POP content and displays it on the display.
[1979] Step 5: The parent asks, "Where are the toy sales?" The device captures the audio and sends it to the server.
[1980] Step 6: The server generates toy department information based on the query and sends it to the terminal.
[1981] Step 7: The terminal displays the received information on its display.
[1982] Example 2: Checking details of individual products
[1983] Step 1: The customer points to the product image on the display.
[1984] Step 2: The device recognizes the pointing gesture and requests detailed information about the product from the server.
[1985] Step 3: The server obtains the product details and sends them to the terminal.
[1986] Step 4: The terminal displays the received information on its display.
[1987] Step 5: The customer checks the product price, specifications, and stock information and makes a purchasing decision.
[1988] The above is the specific flow of operations at each step. This invention realizes the provision of personalized information in real time based on the attributes and behavior of customers.
[1989] Example 1
[1990] Next, a description will be given of Example 1. In the following description, the data processing device 12 will be referred to as a "server" and the robot 414 will be referred to as a "terminal."
[1991] Conventional digital signage systems have the problem that the information provided to customers is uniform and not personalized enough for each individual customer. As a result, it is not possible to introduce optimal products and services that meet each customer's interests and needs, and sales promotion effectiveness is limited. Furthermore, there is a lack of a mechanism to effectively collect and analyze customer behavior data and optimize content in real time.
[1992] The specific processing by the specific processing unit 290 of the data processing device 12 in the first embodiment is realized by the following means.
[1993] In this invention, the server includes a means for capturing customer faces and acquiring attribute data, a means for analyzing and saving the acquired attribute data, and a database for accumulating product information. This makes it possible to select and display optimal products and services in real time according to customer attributes. Furthermore, by collecting purchasing data and updating and optimizing the AI model based on this, it is possible to consistently provide effective content.
[1994] "Means for capturing customer faces and acquiring attribute data" refers to technology that uses cameras and sensors to capture images of customers' faces in a store and extract information such as age, gender, and number of customers from that image data.
[1995] "Means for analyzing and storing acquired attribute data" refers to hardware and software for analyzing attribute data obtained by facial recognition technology and storing it in a database.
[1996] A "database that stores product information" is a database system that manages detailed information about each product and service sold in a retail store or shopping mall, and allows for searching and referencing.
[1997] "Means for generating personalized POP content in real time" refers to algorithms and software for dynamically creating and displaying promotional information that is optimal for each customer based on their attribute data.
[1998] The "display means for displaying the generated content" refers to a display device for visually presenting the POP content sent from the server to customers in the form of images, videos, text, etc.
[1999] "Means for recognizing customer pointing movements" refers to technology that uses sensors such as cameras to detect customer pointing movements and analyze those movements.
[2000] "Means for recognizing customer questions via voice" refers to voice recognition technology that uses a microphone to capture the customer's voice and converts the voice data into text.
[2001] "Means for providing information based on voice questions" refers to a system that searches for appropriate information based on the question content obtained by voice recognition and provides that information.
[2002] "Means for personalizing information" refers to algorithms and systems for providing the most appropriate information to individual customers based on their attribute data and purchasing history.
[2003] "Means for collecting purchasing data" refers to a system for collecting customer purchasing history and behavioral data, and storing and managing this in a database.
[2004] "Means for updating AI models based on collected data and optimizing POP content" refers to a system that uses collected customer behavioral and purchasing data to continuously learn and update AI models to generate optimal POP content.
[2005] This invention relates to a digital signage-type POP system that introduces optimal products and services in real time based on customer attribute information. This system is used in retail stores and shopping malls to provide personalized information to individual customers.
[2006] System Configuration
[2007] This system is broadly composed of a "server," a "terminal," and a "user." Each component will be explained in detail below.
[2008] server
[2009] The server has the following functions:
[2010] 1. A means of capturing customer faces and obtaining attribute data
[2011] The server monitors the store in real time through installed cameras, capturing images of customers' faces and acquiring attribute data such as age, gender, number of people, etc. This facial recognition is performed using Python and OpenCV, applying facial recognition algorithms such as YOLO.
[2012] 2. Means for analyzing and storing acquired attribute data
[2013] The server analyzes the acquired customer attribute data and saves it in a database. The attribute data is continuously updated and to manage the information of multiple customers in real time, a database management system such as MySQL or PostgreSQL is used. The data is managed via an API using the Django or Flask framework.
[2014] 3. Database for storing product information
[2015] The product information database stores detailed information about the products it handles, allowing for quick search and reference as needed. This database includes information such as product prices, specifications, and stock status.
[2016] 4. A way to generate personalized POP content in real time
[2017] The server selects the most suitable products and services based on customer attribute data and dynamically generates personalized POP content using a generative AI model (e.g., GPT-3). The generated POP content is converted into HTML or image format as a text message or a list of recommended products and sent to the device.
[2018] 5. Display means for displaying generated content
[2019] The terminal visually displays real-time POP content sent from the server on a display, which can be powered by a Raspberry Pi or Windows PC and runs as a web application in a browser.
[2020] Terminal
[2021] The terminal is installed inside the store and serves as an interface with customers.
[2022] 1. A method for recognizing customer pointing gestures
[2023] The device uses a camera to recognize the customer's pointing gestures in real time, using hand movement recognition algorithms such as TensorFlow and OpenPose, and sends the processing results to a server using a Python script.
[2024] 2. A way to recognize customer questions by voice
[2025] The device uses a microphone to capture the customer's voice and converts it into text using speech recognition technology, using the Google Cloud Speech-to-Text API.
[2026] 3. Means of providing information based on voice queries
[2027] The server searches for appropriate information based on the question obtained through voice recognition and sends that information to the terminal, which then displays the information on its screen.
[2028] User
[2029] The user is a customer of the store and performs the following operations.
[2030] 1. Viewing content
[2031] Users view the POP content displayed on the display and check information about products and services that interest them.
[2032] 2. Questions and Pointing
[2033] Users can ask for details about products they are interested in by voice or by pointing at a specific product on the screen. The question or information about the product they are pointing at is then displayed on the screen to help them make a purchasing decision.
[2034] Specific examples
[2035] Example 1: Parent and child visitor scenario
[2036] A parent and child visit a shopping mall. The server captures the faces of the parent and child via a camera and determines that the parent is in their late 30s and the child is about 7 years old. Based on this, the device displays information mainly about toys and events for children. When the parent asks, "Where are the toy sales?", the device recognizes the voice question, retrieves the answer from the server, and displays it on the display.
[2037] Example 2: Checking details of individual products
[2038] The customer points at the image of a product displayed on the display. The terminal recognizes the pointing gesture and retrieves detailed information about that product from the server. The product's price, specifications, stock status, and other information are displayed on the display in real time.
[2039] Prompt Sentence Examples
[2040] "What kind of information should be displayed when a parent and child in their 30s visit?"
[2041] "When a customer points to a specific product, explain how you would retrieve and display information about that product."
[2042] This invention makes it possible to provide optimal information in real time based on customer attributes and behavior, thereby enhancing sales promotion effectiveness. In addition, by continually updating the AI model based on collected data, it is possible to continuously generate effective content.
[2043] The flow of the identification process in the first embodiment will be described with reference to FIG.
[2044] Step 1: Discover customers and obtain attribute data
[2045] The server monitors the store in real time through the installed cameras. It captures video data and uses it as input to extract attribute information such as age, gender, and number of people. This is done using Python and OpenCV, and processing high-resolution video data with facial recognition algorithms such as YOLO.
[2046] Input: Real-time video data from the camera
[2047] Output: Customer attribute data (age, gender, number of people, etc.)
[2048] Step 2: Parse and store customer attribute data
[2049] The server analyzes the acquired customer attribute data and stores it in a database. Database management uses MySQL or PostgreSQL, and data management and updating is performed using the Django or Flask framework.
[2050] Input: Customer attribute data
[2051] Output: Analysis data stored in a database
[2052] Step 3: Generate optimal POP content
[2053] The server selects the most suitable products and services from a product information database based on the accumulated customer attribute data, and dynamically generates personalized POP content using a generative AI model (e.g., GPT-3) and sends it to the device in HTML or image format.
[2054] Input: Customer attribute data, product information
[2055] Output: Personalized POP content
[2056] Step 4: View POP Content
[2057] The terminal displays POP content sent from the server in real time. The displayed content includes images, videos, and text. It runs as a web application in a browser on a Raspberry Pi or Windows PC.
[2058] Input: POP content sent from the server
[2059] Output: Content displayed on the display
[2060] Step 5: Recognizing pointing gestures and providing detailed information
[2061] The device uses a camera to recognize the customer's pointing movements in real time and sends the data to a server. The server then analyzes the movement data and obtains detailed information about the product the customer is pointing at. The movement recognition uses TensorFlow and OpenPose.
[2062] Input: Customer pointing gesture
[2063] Output: Detailed information about the pointed item
[2064] Step 6: Recognize voice questions and provide answers
[2065] The device uses a microphone to capture the customer's voice question and converts it into text using the Google Cloud Speech-to-Text API. The server then analyzes the question from the converted text, searches for and retrieves the appropriate answer, and sends it to the device, which then displays the information on its screen.
[2066] Input: Voice question data
[2067] Output: Answer information for the question
[2068] Step 7: Collect purchasing data and update the AI model
[2069] The server periodically collects customer purchase data and stores it in a database. The collected data is used to train and update the AI model, optimizing the effectiveness of the generated POP content.
[2070] Input: Purchasing data
[2071] Output: Updated AI model, optimized POP content
[2072] (Application example 1)
[2073] Next, a description will be given of Application Example 1. In the following description, the data processing device 12 will be referred to as a "server" and the robot 414 will be referred to as a "terminal."
[2074] Conventional digital signage-type POP systems were primarily intended to provide information to customers in retail stores and shopping malls. However, in the logistics field, optimizing worker flow and supporting rapid inventory management are important, and there were few systems that met these needs. Therefore, there is a demand for technology that can improve work efficiency and increase the accuracy of inventory management in logistics centers, warehouses, and other on-site locations.
[2075] The specific processing by the specific processing unit 290 of the data processing device 12 in the application example 1 is realized by the following means.
[2076] In this invention, the server includes a face recognition unit, a customer attribute analysis unit, a product information database, a unit for generating POP content in real time, a display unit for displaying the generated POP content, a unit for recognizing customer pointing gestures, a unit for recognizing customer questions by voice, a unit for providing information based on the voice questions, a unit for personalizing information, a unit for collecting purchase data, a unit for updating an AI model based on the collected data, a unit for optimizing worker movement lines, a unit for visually guiding inventory locations, and a unit for linking with smart glasses or a head-mounted display. This makes it possible to optimize worker movement lines and provide real-time visual guidance on inventory locations at sites such as logistics centers and warehouses.
[2077] "Facial recognition means" is a technology that uses a camera to capture the face of a worker and analyzes attribute data such as age and gender based on the image.
[2078] "Customer attribute analysis means" is a technology that analyzes the tendencies and characteristics of specific customers based on attribute data acquired by face recognition means.
[2079] The "product information database" is a database that stores detailed information about each product or service, allowing users to search and reference the information they need.
[2080] "Means for generating POP content in real time" refers to technology for dynamically generating personalized POP content based on data collected in real time.
[2081] "Display means" refers to a device that displays generated POP content and information in real time, and includes display formats such as images, videos, and text.
[2082] The "means for recognizing pointing gestures" is a technology that uses a camera to detect the pointing gestures of workers and customers and analyzes their location information.
[2083] "Voice recognition means" refers to a technology that uses a microphone to capture the voice of a worker or customer and converts it into text data using voice recognition technology.
[2084] "Means for providing information based on voice questions" refers to a technology that acquires appropriate information from a server based on the content of a voice-recognized question and displays it on a display means.
[2085] "Means for personalizing information" refers to technology that generates individually optimized information and messages based on the attributes and behavioral data of customers and workers.
[2086] "Means for collecting purchasing data" refers to technology that collects data on customers' actual purchasing behavior and selected products and stores it in a database.
[2087] "Means for updating AI models based on collected data" refers to technologies for analyzing collected data and continuously training and optimizing AI models.
[2088] "Means to optimize worker movement" refers to technology that analyzes worker location information and movement in real time and proposes optimal movement routes and work procedures.
[2089] "Means for visually guiding inventory location" refers to technology that uses smart glasses or head-mounted displays within logistics centers to visually show workers the exact location of inventory.
[2090] "Means for linking with smart glasses or head-mounted displays" refers to technology that links smart glasses or head-mounted displays with the system to display information in real time and accept input.
[2091] This invention relates to a digital signage system for improving work efficiency in logistics centers and warehouses. This system uses facial recognition to analyze worker attributes, optimize worker movement lines, and visually guide inventory locations. Furthermore, by linking with smart glasses or head-mounted displays, work efficiency can be further improved.
[2092] System Configuration
[2093] This system is broadly composed of a server, terminals, and users. Each component is explained in detail below.
[2094] server
[2095] The server has the following functions:
[2096] 1. Facial Recognition Methods
[2097] The camera captures the worker's face and acquires attribute data such as age and gender, which is then analyzed using a facial recognition algorithm.
[2098] 2. Customer attribute analysis means
[2099] The system analyzes worker characteristics based on attribute data acquired through facial recognition, which is updated in real time and stored in a database.
[2100] 3. Product information database
[2101] This database comprehensively manages inventory information within the distribution center and is used to identify optimal inventory locations and routes.
[2102] 4. A way to generate POP content in real time
[2103] Generate personalized POP content in real time based on worker attribute information and inventory information.
[2104] 5. Optimizing traffic flow
[2105] Analyzes the real-time location information of workers and calculates and suggests the optimal movement route.
[2106] 6. Inventory location visual guide means
[2107] The exact location of inventory is displayed in real time on smart glasses or head-mounted displays.
[2108] 7. AI model update methods
[2109] The collected data is used to train and optimize the AI model, which is an algorithm designed to continuously improve work efficiency.
[2110] Terminal
[2111] The terminals are installed inside the logistics center and serve as an interface with workers.
[2112] 1. Display Means
[2113] Displays real-time POP content and traffic flow information in the form of images, videos, text, etc.
[2114] 2. Pointing Recognition Method
[2115] The camera is used to recognize the worker's pointing gesture, and detailed information is requested from the server based on that gesture.
[2116] 3. Voice Recognition Method
[2117] The worker's voice is captured using a microphone and converted into text data using voice recognition technology.
[2118] 4. Means of providing information on the display
[2119] Information acquired from the server based on questions received by voice is displayed on the display.
[2120] User
[2121] The users are workers in the logistics center and use the system as follows:
[2122] 1. Viewing content
[2123] Workers can visually check the content displayed on the screen and grasp the information necessary for their work.
[2124] 2. Questions and Pointing
[2125] Workers can ask questions by voice or point to specific locations on the screen to obtain information.
[2126] Specific examples
[2127] Example 1: Picking support
[2128] When picking items in a distribution center, the smart glasses optimize worker movement in real time, suggesting the most efficient route, and visually guiding workers to the exact location of inventory, significantly reducing work time.
[2129] Example prompt sentence:
[2130] "Please build an AI model that can suggest optimal movement paths and display inventory locations in real time to help a man in his 30s improve the efficiency of his picking work."
[2131] Example 2: Improving efficiency of warehousing operations
[2132] When receiving new inventory, the smart glasses will instruct workers in real time on where to place the inventory, supporting efficient stocking operations. Facial recognition is used to provide an individually optimized route and placement location.
[2133] Example prompt sentence:
[2134] "Propose an AI model for smart glasses that uses facial recognition to guide workers to the appropriate shelf location when receiving new inventory."
[2135] The flow of the specific processing in the application example 1 will be described with reference to FIG.
[2136] Step 1:
[2137] The server captures the worker's face through a camera. The input data is the image from the camera, and it analyzes attribute data such as age and gender using a facial recognition algorithm. The output is the analyzed worker's attribute data.
[2138] Step 2:
[2139] The server analyzes the characteristics of the workers based on the attribute data acquired in step 1. The input is customer attribute data, which is stored in a database and updated in real time. The output is the updated characteristic data.
[2140] Step 3:
[2141] The server searches and references inventory information from the product information database. The input is the worker's characteristic data and the inventory database, and the output is information on the optimal inventory location and route. This allows data to be collected to optimize worker movement lines.
[2142] Step 4:
[2143] The server generates the most suitable POP content for each worker in real time. The input is the worker's attribute data and inventory information, and the generated content is sent to the terminal. The output is the dynamically generated POP content.
[2144] Step 5:
[2145] The terminal displays the POP content sent from the server. The input is the POP content sent from the server, and the output is the content displayed on the display.
[2146] Step 6:
[2147] The terminal recognizes the worker's pointing gestures in real time. The input is video of the worker's movements captured by the terminal's built-in camera, and the output is the location information of the recognized pointing gesture. Based on this information, a request for detailed information is made to the server.
[2148] Step 7:
[2149] The terminal captures the worker's voice with a microphone and converts it into text data using voice recognition technology. The input is the worker's voice data, and the output is the converted text data.
[2150] Step 8:
[2151] The server acquires the appropriate information based on the question that has been recognized by voice and sends it to the terminal. The input is the text data that has been recognized by voice, and the output is the acquired information. This provides the necessary information to the worker.
[2152] Step 9:
[2153] The server trains and optimizes the AI model based on the collected data. The input is the collected purchase and operation data, and the output is an updated AI model. This updated model is used to guide future flow optimization and inventory location.
[2154] Step 10:
[2155] The terminal displays information sent from the server on smart glasses or a head-mounted display. The input is inventory location data and movement line data sent from the server, and the output is information displayed on the visual device. This allows workers to work efficiently.
[2156] Furthermore, an emotion engine that estimates the user's emotion may be further combined. That is, the identification processing unit 290 may estimate the user's emotion using the emotion identification model 59, and perform identification processing using the user's emotion.
[2157] This invention relates to a digital signage-type POP system that introduces optimal products and services in real time based on customer attribute information and emotional data. This system is used in retail stores and shopping malls to provide personalized information to individual customers.
[2158] System Configuration
[2159] This system is broadly composed of a "server," a "terminal," and a "user." Details of each component are explained below.
[2160] server
[2161] The server has the following functions:
[2162] 1. Facial Recognition Methods:
[2163] The server captures the customer's face through a camera and obtains high-resolution image data.
[2164] Facial recognition algorithms are used to obtain customer attribute data such as age, gender, and number of customers.
[2165] 2. Customer attribute analysis tools and sentiment engine:
[2166] Attribute data analyzed from facial recognition data is stored in a database and updated in real time.
[2167] The emotion engine analyzes customer emotions from captured facial images and tone of voice, and generates emotion data.
[2168] 3. Product information database and personalized message generation:
[2169] The product information database stores detailed information about each product and service, allowing users to search and reference the information they need.
[2170] By referencing customer attribute data and sentiment data, the system selects the most suitable products and services and dynamically generates POP content.
[2171] Add personalized messages from your membership database and purchase history to generate individual messages for specific customers.
[2172] 4. Learning and optimization:
[2173] The server analyzes the collected customer behavioral data, purchasing data, and emotional data to evaluate the effectiveness of POP content.
[2174] Based on the results of the effectiveness evaluation, the AI model will be updated to improve the accuracy of future content generation algorithms.
[2175] Terminal
[2176] The terminal is installed inside the store and serves as an interface with customers.
[2177] 1. Display means:
[2178] Real-time POP content sent from the server is displayed on the display.
[2179] The content can be displayed in the form of images, videos, text, etc.
[2180] 2. Pointing recognition method:
[2181] The camera recognizes the customer's pointing movements in real time.
[2182] A request is made to the server for detailed information about the product pointed to by the customer, and the received information is displayed on the screen.
[2183] 3. Question and Answer Function:
[2184] The customer's voice questions are captured using a microphone and converted into text data using voice recognition technology.
[2185] Information based on the question is obtained from the server and displayed on the screen.
[2186] User
[2187] The user is a customer of the store and performs the following operations.
[2188] 1. Viewing content:
[2189] Users view the POP content displayed on the display and check information about products and services that interest them.
[2190] 2. Questions and Pointing:
[2191] Users can ask for details about products they are interested in by voice or by pointing at specific products on the screen.
[2192] The question and information about the product you point at will be displayed on the screen to help you make a purchasing decision.
[2193] Specific examples
[2194] Example 1: Parent-child visit scenario
[2195] A parent and child visit a shopping mall. The server captures the faces of the parent and child via camera and analyzes that the parent is in their late 30s and the child is around 7 years old. When the emotion engine recognizes the parent's facial expression as being interested and having fun, the device displays information mainly about toys and events for children. When the parent asks, "Where are the toy sales?", the device sends the voice question and analysis results to the server and displays the corresponding information on the display.
[2196] Example 2: Checking details of individual products and emotion recognition
[2197] When a customer points at an image of a product displayed on the display, the device recognizes the pointing gesture and requests detailed information about that product from the server. The server retrieves the product information, and the emotion engine also analyzes the reaction. For example, if the customer shows a surprised expression, information about a special campaign will be displayed. The customer then checks the product's price, specifications, and stock status before making a purchasing decision.
[2198] In this way, combining emotion engines enables more personalized and effective information provision to customers. Using emotion data, it is possible to respond to customers' instantaneous reactions and provide a more personalized shopping experience. This configuration makes it possible to maximize sales promotion effects and improve customer satisfaction.
[2199] The processing flow will be explained below.
[2200] Program processing steps
[2201] Server Processing Steps
[2202] Step 1:
[2203] The server captures the faces of customers in the store through a camera and obtains high-resolution image data.
[2204] Step 2:
[2205] The server analyzes the acquired facial images using an AI algorithm to estimate attribute data such as the customer's age, gender, and number of people.
[2206] Step 3:
[2207] The server updates the analysis results to a database in real time and stores customer attribute data.
[2208] Step 4:
[2209] The server uses an emotion engine to analyze facial images, tone of voice, and other factors to generate customer emotion data.
[2210] Step 5:
[2211] The server stores the generated emotion data in a database and updates it in real time.
[2212] Step 6:
[2213] The server searches and selects the most suitable products and services from a product information database based on customer attribute data and emotional data.
[2214] Step 7:
[2215] The server dynamically generates POP content in real time based on the selection results.
[2216] Step 8:
[2217] The server references the member database and purchase history and generates personalized messages as needed.
[2218] Step 9:
[2219] The server transmits the generated POP content to the terminal.
[2220] Step 10:
[2221] The server analyzes the collected customer behavioral data, purchasing data, and emotional data to evaluate the effectiveness of the POP content.
[2222] Step 11:
[2223] Based on the effectiveness evaluation results, the server updates the AI model and improves the accuracy of future content generation algorithms.
[2224] Terminal processing steps
[2225] Step 1:
[2226] The terminal receives the POP content sent from the server.
[2227] Step 2:
[2228] The device displays the received content on the display, which can be in the form of images, videos, text, etc.
[2229] Step 3:
[2230] The device uses a camera to recognize the customer's pointing movements in real time.
[2231] Step 4:
[2232] The terminal requests information about the product pointed at by the customer from the server.
[2233] Step 5:
[2234] The terminal receives the detailed product information sent from the server and displays it on the display.
[2235] Step 6:
[2236] The terminal uses a microphone to capture the customer's voice asking a question.
[2237] Step 7:
[2238] The device uses voice recognition technology to convert the voice into text data and send it to the server.
[2239] Step 8:
[2240] The terminal displays the response information sent from the server on the display.
[2241] User operation steps
[2242] Step 1:
[2243] The user views the POP content displayed on the display.
[2244] Step 2:
[2245] Users can input questions about products or services that interest them by voice.
[2246] Step 3:
[2247] The user points to a particular product on the display to see more information about it.
[2248] Step 4:
[2249] Users make purchasing decisions based on the displayed information.
[2250] Specific examples
[2251] Example 1: Parent-child visit scenario
[2252] Step 1: A parent and child visit a shopping mall. The server captures the faces of the parent and child through a camera and analyzes them.
[2253] Step 2: The server estimates the age and gender of the parents and children and stores the attribute data in a database.
[2254] Step 3: The server analyzes the facial image and tone of voice using an emotion engine to generate emotion data.
[2255] Step 4: The server selects toys and event information for children based on the generated attribute data and emotion data, and generates POP content. 【...
Claims
1. A facial recognition means; Customer attribute analysis means; A product information database, a means for generating POP content in real time; a display means for displaying the generated POP content; a means for recognizing a customer's pointing gesture; a means for recognizing a customer's question by voice; means for providing information based on a voice query; a means for performing personalization of information; a means for collecting purchasing data; A means to update the AI model based on the collected data; A system including:
2. The system according to claim 1 , further comprising means for selecting optimal products and services based on customer attributes.
3. 2. The system according to claim 1, further comprising means for optimizing the effectiveness of the POP content generated based on actual purchase data.
Citation Information
Patent Citations
Persona chatbot control method and system
JP2022180282A