System
The system uses image recognition and generative AI to optimize product placement and inventory management in retail, addressing inefficiencies and errors, thereby enhancing operational efficiency and sales performance.
Patent Information
- Application Number
- JP2024131459
- Authority / Receiving Office
- JP · JP
- Patent Type
- Applications
- Current Assignee / Owner
- Filing Date
- 2024-08-07
- Publication Date
- 2026-02-20
AI Technical Summary
In retail sales sites, product and inventory management is time-consuming, prone to errors, and difficult to optimize, leading to out-of-stock issues, ordering mistakes, and reduced operational efficiency, with challenges in real-time inventory tracking and product placement affecting sales performance.
A system that uses image recognition technology and generative AI to analyze product locations and quantities, optimize product placement based on sales data, and provide real-time inventory management through visual and voice notifications, allowing users to correct errors via voice commands.
Enhances product and inventory management efficiency, increases sales, and improves decision-making accuracy by streamlining product placement and inventory processes.
Smart Images

Figure 2026028843000001_ABST
Abstract
Description
[Technical Field]
[0001] The technology of the present disclosure relates to a system. [Background technology]
[0002] Patent document 1 discloses a persona chatbot control method performed by at least one processor, the method including the steps of receiving a user utterance, adding the user utterance to a prompt including an instruction sentence related to a description of the chatbot character, encoding the prompt, and inputting the encoded prompt into a language model to generate a chatbot utterance in response to the user utterance. [Prior art documents] [Patent documents]
[0003] [Patent Document 1] Japanese Patent Publication No. 2022-180282 Summary of the Invention [Problem to be solved by the invention]
[0004] In sales sites, particularly in the retail industry, product and inventory management takes a great deal of time, and out-of-stock items and ordering errors lead to losses. Differences in management ability can also result in lost opportunities, and optimizing product placement can be difficult, negatively impacting sales. Furthermore, it can be difficult to accurately grasp product inventory status in real time, making it difficult to make decisions about ordering and stock disposal at the right time. These issues, which lead to reduced operational efficiency and a decline in profits, require solutions. [Means for solving the problem]
[0005] The present invention comprises a system including: means for acquiring image data from product shelves in a store; means for analyzing the acquired image data to identify the product locations and quantities; means for storing the identified information in a database; means for referencing sales data based on the information stored in the database and analyzing the relationship between product placement and sales; means for proposing optimal product placement based on the analysis results; means for visualizing the proposal and notifying the user; and means for reading the notification aloud to the user. Furthermore, the system allows the user to input a voice command in response to the voice-readout content, and the system reflects the correction content in the database, thereby enabling correction of erroneous judgments. Furthermore, the system includes means for generating management advice for ordering and inventory disposal based on inventory results and notifying the user of the management advice, enabling accurate management decisions in real time. This results in more efficient product and inventory management, increased sales, and more accurate management decisions.
[0006] "Image data" refers to image information captured from product shelves in a store, and is used for product recognition and placement analysis.
[0007] "Product location" is information indicating the physical location of each product on a product shelf in a store.
[0008] "Quantity" is information indicating the number of a particular product present in the store.
[0009] A "database" is a digital storage system for storing product information, analysis results, sales data, etc.
[0010] "Sales data" is information indicating the sales performance of each product within a specific period.
[0011] "Analysis" is a computational process for evaluating the relationship between product placement and sales and identifying optimal placement patterns.
[0012] "Suggestion" is information that indicates optimal product placement advice calculated by the generative AI model.
[0013] "Visualization" is the process of presenting analysis results and proposals in a visually easy-to-understand format.
[0014] "Voice reading" is a function that converts text information into voice and provides it to the user audibly.
[0015] A "voice command" is an operation format in which the user gives instructions to the system using voice.
[0016] "Inventory" is the process of checking and recording the real-time inventory status of products in a store.
[0017] "Management advice" is information that suggests optimal actions such as ordering and inventory disposal based on inventory results and sales data. [Brief explanation of the drawings]
[0018] [Figure 1] 1 is a conceptual diagram showing an example of the configuration of a data processing system according to a first embodiment. [Figure 2] 1 is a conceptual diagram showing an example of main functions of a data processing device and a smart device according to a first embodiment. [Figure 3] FIG. 10 is a conceptual diagram showing an example of the configuration of a data processing system according to a second embodiment. [Figure 4] FIG. 10 is a conceptual diagram showing an example of main functions of a data processing device and smart glasses according to a second embodiment. [Figure 5] FIG. 10 is a conceptual diagram showing an example of the configuration of a data processing system according to a third embodiment. [Figure 6] FIG. 11 is a conceptual diagram showing an example of main functions of a data processing device and a headset-type terminal according to a third embodiment. [Figure 7] FIG. 10 is a conceptual diagram showing an example of the configuration of a data processing system according to a fourth embodiment. [Figure 8] FIG. 10 is a conceptual diagram showing an example of main functions of a data processing device and a robot according to a fourth embodiment. [Figure 9] 1 shows an emotion map onto which multiple emotions are mapped. [Figure 10] 1 shows an emotion map onto which multiple emotions are mapped. [Figure 11] FIG. 3 is a sequence diagram showing a processing flow of the data processing system according to the first embodiment. [Figure 12] FIG. 10 is a sequence diagram showing the flow of processing in the data processing system in Application Example 1. [Figure 13] FIG. 10 is a sequence diagram showing the flow of processing in the data processing system according to the second embodiment when an emotion engine is combined. [Figure 14] FIG. 10 is a sequence diagram showing the flow of processing in the data processing system in Application Example 2 when an emotion engine is combined. DETAILED DESCRIPTION OF THE INVENTION
[0019] An example of an embodiment of a system according to the technology of the present disclosure will be described below with reference to the accompanying drawings.
[0020] First, the terms used in the following description will be explained.
[0021] In the following embodiments, a coded processor (hereinafter simply referred to as a "processor") may be a single arithmetic device or a combination of multiple arithmetic devices. Furthermore, a processor may be a single type of arithmetic device or a combination of multiple types of arithmetic devices. Examples of arithmetic devices include a CPU (Central Processing Unit), a GPU (Graphics Processing Unit), a GPGPU (General-Purpose computing on Graphics Processing Units), and an APU (Accelerated Processing Unit).
[0022] In the following embodiments, a coded RAM (Random Access Memory) is a memory in which information is temporarily stored and is used as a working memory by a processor.
[0023] In the following embodiments, the coded storage is one or more non-volatile storage devices that store various programs, various parameters, etc. Examples of non-volatile storage devices include flash memory (SSD (Solid State Drive)), magnetic disks (e.g., hard disks), and magnetic tapes.
[0024] In the following embodiments, a communication I / F (Interface) with a symbol is an interface including a communication processor, an antenna, etc. The communication I / F controls communication between multiple computers. Examples of communication standards applied to the communication I / F include wireless communication standards including 5G (5th Generation Mobile Communication System), Wi-Fi (registered trademark), Bluetooth (registered trademark), etc.
[0025] In the following embodiments, "A and / or B" is synonymous with "at least one of A and B." In other words, "A and / or B" means that it may be only A, only B, or a combination of A and B. Furthermore, in this specification, the same concept as "A and / or B" is also applied when three or more things are expressed connected by "and / or."
[0026] [First embodiment]
[0027] FIG. 1 shows an example of the configuration of a data processing system 10 according to the first embodiment.
[0028] 1, a data processing system 10 includes a data processing device 12 and a smart device 14. An example of the data processing device 12 is a server.
[0029] The data processing device 12 includes a computer 22, a database 24, and a communication I / F 26. The computer 22 is an example of a "computer" according to the technology of the present disclosure. The computer 22 includes a processor 28, a RAM 30, and a storage 32. The processor 28, the RAM 30, and the storage 32 are connected to a bus 34. The database 24 and the communication I / F 26 are also connected to the bus 34. The communication I / F 26 is connected to a network 54. Examples of the network 54 include a WAN (Wide Area Network) and / or a LAN (Local Area Network).
[0030] The smart device 14 includes a computer 36, a reception device 38, an output device 40, a camera 42, and a communication I / F 44. The computer 36 includes a processor 46, a RAM 48, and a storage 50. The processor 46, the RAM 48, and the storage 50 are connected to a bus 52. The reception device 38, the output device 40, and the camera 42 are also connected to the bus 52.
[0031] The reception device 38 includes a touch panel 38A, a microphone 38B, and the like, and receives user input. The touch panel 38A detects contact with an indicator (for example, a pen or a finger) to receive user input by the touch of the indicator. The microphone 38B detects the user's voice to receive user input by voice. The control unit 46A transmits data indicating the user input received by the touch panel 38A and the microphone 38B to the data processing device 12. In the data processing device 12, the specific processing unit 290 acquires the data indicating the user input.
[0032] The output device 40 includes a display 40A and a speaker 40B, and presents data to the user 20 by outputting the data in a form of expression that the user 20 can perceive (for example, audio and / or text). The display 40A displays visible information such as text and images in accordance with instructions from the processor 46. The speaker 40B outputs audio in accordance with instructions from the processor 46. The camera 42 is a compact digital camera equipped with an optical system including a lens, aperture, and shutter, and an imaging element such as a CMOS (Complementary Metal-Oxide-Semiconductor) image sensor or a CCD (Charge Coupled Device) image sensor.
[0033] The communication I / F 44 is connected to a network 54. The communication I / Fs 44 and 26 control the exchange of various information between the processor 46 and the processor 28 via the network 54.
[0034] FIG. 2 shows an example of the main functions of the data processing device 12 and the smart device 14.
[0035] 2, in the data processing device 12, a specific process is performed by the processor 28. A specific processing program 56 is stored in the storage 32. The specific processing program 56 is an example of a "program" according to the technology of the present disclosure. The processor 28 reads the specific processing program 56 from the storage 32 and executes the read specific processing program 56 on the RAM 30. The specific process is realized by the processor 28 operating as a specific processing unit 290 in accordance with the specific processing program 56 executed on the RAM 30.
[0036] The storage 32 stores a data generation model 58 and an emotion identification model 59. The data generation model 58 and the emotion identification model 59 are used by the identification processing unit 290.
[0037] In the smart device 14, the processor 46 performs the reception output process. The storage 50 stores a reception output program 60. The reception output program 60 is used in conjunction with the specific processing program 56 by the data processing system 10. The processor 46 reads the reception output program 60 from the storage 50 and executes the read reception output program 60 on the RAM 48. The reception output process is realized by the processor 46 operating as the control unit 46A in accordance with the reception output program 60 executed on the RAM 48.
[0038] Next, a description will be given of the specific processing performed by the specific processing unit 290 of the data processing device 12. In the following description, the data processing device 12 will be referred to as a "server" and the smart device 14 will be referred to as a "terminal."
[0039] The present invention relates to an optimal product placement support system for the retail industry that uses image recognition technology and generative AI, and aims to analyze the impact of product placement in a store on sales and optimize product placement. Specific implementation methods of the present invention are described below.
[0040] System configuration
[0041] The system consists of the following main components:
[0042] 1. Terminal: Equipped with a camera to take pictures of product shelves in a store and has the function of sending image data to a server. Examples include smartphones and dedicated cameras.
[0043] 2. Server: The central data processing unit that receives and analyzes image data, running the generative AI model.
[0044] 3. User: The person who operates the system and makes management decisions such as optimal product placement, ordering, and inventory disposal.
[0045] Program processing flow
[0046] Image data collection
[0047] The user uses the device to take an image of the store's shelves, focusing on a specific area of the shelves to clearly show the location and quantity of products. The device is configured to automatically send the captured image to the server.
[0048] Image data analysis
[0049] The server analyzes the image data received from the device. Using an image analysis engine, it identifies the location, quantity, and specific product name of each product. After the product information is identified, it is stored in a database.
[0050] Saving to a database
[0051] The identified product information is stored in a database along with attributes such as SKU, location, and inventory quantity, making it possible to immediately grasp the current inventory status.
[0052] Analysis of the relationship between placement and sales
[0053] The server references sales data from a database for a set period and uses a generative AI model to analyze the relationship between product placement and sales performance. This analysis identifies the optimal placement pattern for specific products, maximizing sales.
[0054] Proposal for optimal layout
[0055] Based on the analysis results, the server proposes the optimal product placement. This proposal is visualized, allowing users to compare the current product shelf layout with the proposed optimal placement. The proposal is then sent to the terminal and notified to the user.
[0056] Voice reading and correction
[0057] Users can confirm the notification content from the device by voice. The device will read out product information and stock status using text-to-speech (TTS) technology. If there is an incorrect judgment, users can correct it using voice commands. This correction is reflected in the database again and recalculation is performed.
[0058] Management Advice
[0059] Based on the results of periodic inventory, the server generates management advice such as ordering and discount proposals for inventory clearance. These advices are notified to the user via text and voice, supporting effective management decisions.
[0060] Specific examples
[0061] For example, analysis may reveal that sales data for a particular product category (e.g., snacks) are stagnating if they are placed on the bottom shelf. Based on this, the server may suggest placing snacks in front of the register or at eye level. The suggestion is sent to the device as a concrete visual map, and the user can adjust product placement based on the suggestion. Also, if a voice notification reads, "There are 10 bags of snacks in stock," the user can correct it with a voice command, saying, "There are 15 bags in stock," and the system will immediately update the information.
[0062] These features provide a system that streamlines product and inventory management, increases sales, and enables more accurate business decisions.
[0063] The processing flow will be explained below.
[0064] Step 1:
[0065] Users take a photo of the product shelves in a store using a device such as a smartphone or dedicated camera, and by pressing the capture button, the image is automatically saved on the device.
[0066] Step 2:
[0067] The device automatically sends the captured image data to a server via Wi-Fi or mobile networks.
[0068] Step 3:
[0069] The server analyzes the image data received from the device and uses an image analysis engine to identify products in the image and determine the location and number of each product.
[0070] Step 4:
[0071] The server stores the identified product information in a database, including the SKU, location, and quantity in stock.
[0072] Step 5:
[0073] The server references sales data based on the information stored in the database, extracts past sales data using SQL queries, and analyzes the relationship between product placement and sales.
[0074] Step 6:
[0075] The server uses a generative AI model to analyze product placement and sales data, thereby identifying optimal product placement patterns.
[0076] Step 7:
[0077] The server generates optimal product placement proposals based on the analysis results, which are then visualized.
[0078] Step 8:
[0079] The server sends the generated placement proposal to the terminal, and the user can check the proposal displayed on the terminal.
[0080] Step 9:
[0081] The device reads product information and stock availability to the user aloud, and uses text-to-speech (TTS) technology to audibly output the contents of the visual suggestions.
[0082] Step 10:
[0083] The user can make corrections using voice commands in response to the content of the voice readout. For example, the user can input a correction such as "There are 15 bags in stock."
[0084] Step 11:
[0085] The server receives voice correction commands from the user and reflects them in the database, where the corrected information is recalculated and saved as the latest information.
[0086] Step 12:
[0087] The server generates management advice for ordering and stock disposal based on the results of periodic inventory. This advice is generated based on the amount of stock and sales trends.
[0088] Step 13:
[0089] The server then sends the generated management advice to the terminal, and the user makes decisions about ordering and inventory disposal based on this advice.
[0090] Through the above steps, this system will improve the efficiency of product management and inventory management, thereby increasing sales and improving the accuracy of management decisions.
[0091] Example 1
[0092] Next, a description will be given of Example 1. In the following description, the data processing device 12 will be referred to as a "server" and the smart device 14 will be referred to as a "terminal."
[0093] The retail industry is seeking to optimize product placement and streamline inventory management. Conventional methods rely on manual work to determine product locations and quantities, which takes time and effort and is prone to errors. It is also difficult to analyze the relationship between product placement and sales, requiring significant effort to find optimal product placement. Furthermore, updating inventory information in real time is difficult, making it difficult to make quick management decisions. To solve these issues, a system is needed that can efficiently and accurately manage product placement and inventory.
[0094] The specific processing by the specific processing unit 290 of the data processing device 12 in the first embodiment is realized by the following means.
[0095] In this invention, the server includes a means for analyzing image data and identifying product locations and quantities, a means for proposing optimal product placement using a generative AI model, and a means for referencing sales data based on information stored in a database and analyzing the relationship between product placement and sales. This makes it possible to identify product locations and quantities with high accuracy, update inventory information in real time, and quickly propose optimal product placement to maximize sales.
[0096] "Image data" refers to visual information acquired in the form of photographs or videos of product shelves in a store.
[0097] "Acquisition" refers to the act of obtaining image data of an object using a camera or other photographic equipment.
[0098] "Analysis" is the process of extracting and identifying specific information from acquired image data.
[0099] A "generative AI model" is an artificial intelligence algorithm trained to perform image recognition and data analysis.
[0100] "Location" is information indicating the specific location where a particular product is located.
[0101] "Quantity" is a number that indicates how much of a particular product is on the shelf or in stock.
[0102] A "database" is an electronic recording device or system for systematically storing and managing specific information.
[0103] "Sales data" refers to transaction information when a product is sold, including sales amount and sales quantity.
[0104] "Analysis" is the act of examining and interpreting data to find specific relationships and patterns.
[0105] "Suggestion" refers to showing optimal behavior and placement patterns based on the analysis results.
[0106] "Visualization" refers to displaying data and proposals in a visually easy-to-understand manner.
[0107] "Notification" is the act of informing a user of specific information.
[0108] "Text-to-speech" is a technology that converts text information into speech and conveys it to the user auditorily.
[0109] "Voice command" is a method by which a user gives instructions or inputs using their voice.
[0110] "Inventory" refers to the periodic checking and recording of stock and the quantity of merchandise items.
[0111] "Management advice" refers to proposals based on inventory and sales data to support decision-making such as ordering and inventory disposal.
[0112] This invention is an optimal product placement support system for the retail industry that uses image recognition technology and generative AI models. The system acquires image data of product shelves in stores, analyzes the data to identify product locations and quantities, stores the data in a database, and then compares it with sales data to suggest optimal product placement.
[0113] The system consists of three main components:
[0114] 1. Terminal: A device with a camera function that takes pictures of product shelves in a store. Typical examples are smartphones or dedicated cameras.
[0115] 2. Server: A central data processing unit that receives data, performs image analysis, and optimizes product placement using generative AI models.
[0116] 3. User: The person who operates the system and makes management decisions such as optimal product placement, ordering, and inventory disposal.
[0117] Hardware and Software Configuration
[0118] Terminal
[0119] Users take pictures of product shelves in stores using smartphones or dedicated cameras. These devices have the ability to automatically send the captured image data to a server. The captured images are sent to the server via Wi-Fi or mobile data. The devices are also equipped with text-to-speech (TTS) technology, which allows them to notify users by voice.
[0120] server
[0121] The server receives the image data sent from the device and uses an image analysis engine to identify the location, quantity, and product name of each product. This analysis uses deep learning technology and generative AI models, such as machine learning libraries like TensorFlow and PyTorch. The server then stores the identified information in a database and uses that data to analyze the relationship between sales data and product placement.
[0122] Database
[0123] The product information analyzed by the server is stored in a database, which updates attributes such as SKU (Stock Keeping Unit), location, and inventory quantity in real time. Users can easily check this information.
[0124] Examples of use and specific prompts
[0125] Users take pictures of product shelves in a store and send them to a server using their device. The server then receives the image data, analyzes it, identifies the product locations and quantities, and stores them in a database. Based on the analysis results, the server uses a generative AI model to propose optimal product placement and sends it to the device as a visual map. The user can then confirm the proposal visually and audibly and adjust the product placement.
[0126] Prompt Sentence Examples
[0127] "Please identify the names and quantities of the items shown in this image."
[0128] "Please suggest optimal product placement based on sales data and product placement data."
[0129] "Please read out the inventory."
[0130] · "Use voice commands to correct stock levels."
[0131] This will lead to more efficient product and inventory management, increased sales, and more accurate business decisions.
[0132] The flow of the identification process in the first embodiment will be described with reference to FIG.
[0133] Step 1:
[0134] The user uses a device to take a photo of a product shelf in a store. The captured image is high resolution, and the product name and quantity are clearly visible. For example, a smartphone or dedicated camera is used to capture product images under appropriate lighting. The input is an "image of a product shelf in a store," and the output is "high-resolution image data."
[0135] Step 2:
[0136] The device sends the captured image data to the server. The device has a network connection function and encrypts and sends the image data to the server using Wi-Fi or mobile data. The input is "high-resolution image data" and the output is "a notification of completion of transmission to the server."
[0137] Step 3:
[0138] The server analyzes the image data received from the device. The image analysis engine processes the input image and identifies the location, quantity, and name of the product. A generative AI model is used for this analysis. Specifically, machine learning libraries such as TensorFlow and PyTorch are used. The input is "high-resolution image data," and the output is "the location, quantity, and name of the identified product."
[0139] Step 4:
[0140] The server stores the identified information in a database. The identified product information (SKU, location, and stock quantity) is recorded in the database, and inventory information is updated in real time. The input is the "identified product location, quantity, and name," and the output is the "updated database entry."
[0141] Step 5:
[0142] The server analyzes the relationship between sales data and product placement performance based on the stored information. It uses a generative AI model to find the optimal placement pattern. For example, it compares the placement location of a specific product with sales data to identify the placement that maximizes sales. The inputs are "product information stored in the database" and "sales data," and the output is "identification of the optimal placement pattern."
[0143] Step 6:
[0144] The server proposes optimal product placement and generates a visualized proposal. It provides the user with a layout diagram in an easy-to-understand visual format, allowing them to compare the proposed placements. The proposal is sent to the device. The input is the "identification result of the optimal placement pattern," and the output is the "visualized placement proposal."
[0145] Step 7:
[0146] The terminal notifies the user of the proposal content sent from the server by voice. Using text-to-speech (TTS) technology, the placement proposal content and inventory information are read aloud. The input is a "visualized placement proposal" and the output is a "voice notification."
[0147] Step 8:
[0148] When a user wants to correct a voice notification, they input a voice command. For example, they can correct an inventory quantity by voice, and the correction is sent to the server. The server analyzes the voice command and updates the database. The input is the "user's voice command" and the output is the "updated database entry."
[0149] Step 9:
[0150] Based on the results of periodic inventory, the server generates management advice for ordering and stock disposal. It then notifies the user of this advice via text and voice. For example, the advice might be, "There is an excess of snack food in stock. We suggest selling them at a discount." The inputs are "inventory results" and "database information," and the output is "notification of management advice."
[0151] (Application example 1)
[0152] Next, a description will be given of Application Example 1. In the following description, the data processing device 12 will be referred to as a "server" and the smart device 14 will be referred to as a "terminal."
[0153] It is widely recognized in the retail industry that product placement in a store has a significant impact on sales. However, optimizing shelf placement using traditional methods requires a great deal of effort and time, and effective placement is not always achieved. Furthermore, real-time inventory management and rapid management decisions based on sales data are required, but traditional systems do not adequately support this. Furthermore, store staff are required to manually check and correct product locations and inventory status, which is time-consuming and inefficient.
[0154] The specific processing by the specific processing unit 290 of the data processing device 12 in the application example 1 is realized by the following means.
[0155] In this invention, the server includes means for acquiring image data from product shelves in a store, means for analyzing the acquired image data and identifying product locations and quantities, means for saving the identified information in a database, means for referencing sales data based on the information saved in the database and analyzing the relationship between product placement and sales, means for proposing optimal product placement based on the analysis results, means for visualizing the proposal and notifying the user, means for reading the notification aloud to the user, means for taking pictures of the product shelves using a smartphone, means for providing audio notification of product information and inventory status using text-to-speech technology, means for the user to correct information using voice commands, means for generating optimal placements using a generative AI model that contributes to maximizing sales, and means for inputting prompts to the generative AI model to obtain analysis results. This enables efficient acquisition of product shelf images using a smartphone, optimal placement suggestions using generative AI, and rapid and accurate inventory management and management decisions using voice technology.
[0156] "Image data" refers to information that is acquired as an image of product shelves in a store and is the subject of analysis.
[0157] "Analysis" refers to the act of identifying the location and quantity of products from the acquired image data.
[0158] A "database" is a repository of information that stores identified product information and allows for later reference.
[0159] "Sales data" is information relating to product sales, and is used to analyze the relationship between product placement and sales.
[0160] A "generative AI model" is an artificial intelligence technology used to analyze image data and propose optimal product placement.
[0161] A "prompt statement" is an instruction statement that is input into a generative AI model to obtain analysis results.
[0162] "Visualization" is the act of visually suggesting optimal product placement.
[0163] "Voice notification" refers to the act of using text-to-speech technology to inform users of product information and stock status via voice.
[0164] A "smartphone" is a mobile device used to acquire and send notifications about image data.
[0165] A "voice command" is a spoken input by a user to give instructions to a system using voice recognition technology.
[0166] "Inventory management" is the management activity of determining the quantity of goods and maintaining appropriate stock levels.
[0167] "Management decisions" are decisions made regarding product placement and inventory management.
[0168] MODE FOR CARRYING OUT THE INVENTION
[0169] The present invention relates to a system for optimizing product placement in a retail store using a smartphone. A specific implementation method of this system will be described below.
[0170] System configuration
[0171] The system consists of the following major components:
[0172] 1. Device (smartphone):
[0173] It is equipped with a camera to take pictures of product shelves in the store and transmits the image data to a server.
[0174] It has the ability to provide voice notification of product information and stock status using text-to-speech technology.
[0175] It uses voice recognition technology to receive the user's voice commands and send correction information to the server.
[0176] 2. Server:
[0177] It receives image data and runs a generative AI model that performs the analysis.
[0178] Using an image analysis engine (e.g., YOLOv5, TensorFlow), the location, quantity, and specific product name of each product are identified.
[0179] The identified product information is stored in a database (e.g., Firebase, PostgreSQL).
[0180] Sales data is referenced based on the information stored in the database, and the relationship between product placement and sales is analyzed.
[0181] The system generates the optimal layout that contributes to maximizing sales and notifies the user as a visual map.
[0182] 3. User:
[0183] A smartphone is used to take a photo of the product shelf and send the image data to the server.
[0184] Visually confirm the proposed optimal placement and confirm / correct inventory information notified by voice.
[0185] Make business decisions and place orders or dispose of inventory based on advice provided by the system.
[0186] Program processing
[0187] (Collection of image data)
[0188] Users take pictures of product shelves in a store using the camera on their smartphone, and the smartphone automatically sends the captured image data to the server.
[0189] (Image data analysis)
[0190] The server is equipped with an image analysis engine and generative AI models, such as YOLOv5 and TensorFlow, to analyze the received image data. The analysis identifies the location, quantity, and SKU of each product, and stores this information in a database.
[0191] (Saving to database)
[0192] The identified product information is stored in a database along with attributes such as SKU, location, and stock quantity, making it possible to instantly grasp the current inventory status.
[0193] (Analysis of the relationship between placement and sales)
[0194] The server references sales data for a set period from the database and uses a generative AI model to analyze the relationship between product placement and sales performance. At this time, it inputs prompt statements into the generative AI to obtain the analysis results.
[0195] (Optimal layout proposal)
[0196] A placement pattern that contributes to maximizing sales is generated, and the proposed details are sent to a smartphone as a visual map.
[0197] (audio notification and fix)
[0198] The smartphone uses text-to-speech technology (e.g., Google Cloud Text-to-Speech API) to notify the user of product information and stock availability via voice. When the user provides corrections via voice commands, the smartphone sends the information to the server, which updates the database.
[0199] (Providing management advice)
[0200] Based on the results of periodic inventory, the server generates management advice for ordering and stock disposal, which is communicated to the user via text and voice to support effective management decisions.
[0201] Specific examples
[0202] For example, if a particular snack is placed on the bottom shelf, analysis may reveal that sales data is sluggish. The server then suggests placing the snack in the center of the shelf where it is more visible. This suggestion is sent to the smartphone as a visual map, and the user follows the instructions to rearrange the products. Also, if a voice notification says, "There are 10 bags of snacks in stock," the user can correct it by saying, "There are 15 bags in stock," and the system will immediately update the database.
[0203] (Example of a prompt to input to a generative AI model)
[0204] input:
[0205] "Analyze the shelf image below to identify product locations, SKUs, and stock quantities. SKUs: '12345', '54321', '67890', etc."
[0206] output:
[0207] "Product locations and quantities identified: 4 units of SKU: '12345', 6 units of SKU: '54321', and 10 units of SKU: '67890'. Suggested placement: Place snacks in the center of the shelf."
[0208] The flow of the specific processing in the application example 1 will be described with reference to FIG.
[0209] Program processing steps
[0210] Step 1:
[0211] A user uses the camera on their smartphone to take a picture of a product shelf in a store. The image data taken by the smartphone is the input. The device sends this image data to the server. The output is the image data sent to the server.
[0212] Step 2:
[0213] The server analyzes the received image data. Specifically, it uses an image analysis engine (e.g., YOLOv5, TensorFlow) to identify the product location and quantity. The image data is input, and the product location, quantity, and SKU information are output.
[0214] Step 3:
[0215] The server stores the identified product information in a database. The input is the location, quantity, and SKU information of the identified product, and the product information stored in the database is obtained as the output.
[0216] Step 4:
[0217] The server references sales data based on the information stored in the database and analyzes the relationship between product placement and sales. This analysis is performed using a generative AI model (e.g., TensorFlow). Database information and sales data are input, and the analysis results include information on the relationship between product placement and sales.
[0218] Step 5:
[0219] The server inputs a prompt into the generative AI model to obtain the optimal product placement that will maximize sales. The inputs are the prompt and information from the database, and the output is a proposal for the optimal placement.
[0220] Step 6:
[0221] The server generates a visual map of the optimal layout proposal and sends it to the smartphone. The input is the optimal layout proposal, and the output is the visual map data.
[0222] Step 7:
[0223] The smartphone uses text-to-speech technology (e.g., Google Cloud Text-to-Speech API) to provide the user with product information and stock status via voice. The input is a visual map and product information, and the output is a voice notification.
[0224] Step 8:
[0225] The user listens to the voice notification and, if necessary, issues a voice command to correct the information. The smartphone receives this voice command and converts it into text using voice recognition technology (e.g., Google Speech-to-Text API). The voice command is used as input, and the text of the correction is output.
[0226] Step 9:
[0227] The server receives the user's corrections and updates the database: the text of the corrections is the input, and the database update is the output.
[0228] Step 10:
[0229] Based on the results of periodic inventory, the server generates management advice for ordering and stock disposal and notifies the user. Inventory data is input and management advice is output.
[0230] Combining these steps will improve the efficiency of in-store product placement and inventory management, and will also enable faster management decisions to maximize sales.
[0231] Furthermore, an emotion engine that estimates the user's emotion may be combined. That is, the identification processing unit 290 may estimate the user's emotion using the emotion identification model 59 and perform identification processing using the user's emotion.
[0232] The present invention relates to an optimal product placement support system for the retail industry that uses image recognition technology, generative AI, and an emotion engine. The purpose of the system is to analyze the impact of in-store product placement on sales, optimize product placement, and also take into account the emotional state of users. Specific implementation methods of the present invention are described below.
[0233] System configuration
[0234] The system consists of the following main components:
[0235] 1. Terminal: Equipped with a camera to take pictures of product shelves in a store and has the function of sending image data to a server. Examples include smartphones and dedicated cameras.
[0236] 2. Server: The central data processing unit that receives and analyzes image data, running the generative AI model and emotion engine.
[0237] 3. User: The person who operates the system and makes management decisions such as optimal product placement, ordering, and inventory disposal.
[0238] Program processing flow
[0239] Image data collection
[0240] The user uses the device to take an image of the store's shelves, focusing on a specific area of the shelves to clearly show the location and quantity of products. The device is configured to automatically send the captured image to the server.
[0241] Image data analysis
[0242] The server analyzes the image data received from the device. Using an image analysis engine, it identifies the location, quantity, and specific product name of each product. After the product information is identified, it is stored in a database.
[0243] Saving to a database
[0244] The identified product information is stored in a database along with attributes such as SKU, location, and inventory quantity, making it possible to immediately grasp the current inventory status.
[0245] Analysis of the relationship between placement and sales
[0246] The server references sales data from a database for a set period and uses a generative AI model to analyze the relationship between product placement and sales performance. This analysis identifies the optimal placement pattern for specific products, maximizing sales.
[0247] Proposal for optimal layout
[0248] Based on the analysis results, the server generates an optimal product placement proposal. This proposal is visualized, allowing the user to compare the current product shelf layout with the proposed optimal placement. The proposal is then sent to the terminal and notified to the user.
[0249] Introducing the Emotion Engine
[0250] The device reads product information and inventory status to the user aloud and uses an emotion engine to recognize the user's emotional state by analyzing the user's voice and facial expressions to identify their emotions.
[0251] Adjusting notification content
[0252] Based on the emotional state identified by the emotion engine, the server dynamically adjusts the notification content, for example, providing more concise and positive suggestions if the user is feeling stressed.
[0253] Correcting misjudgments and providing management advice
[0254] The user can make corrections to the voice-readout information by voice command. For example, the user can input a correction such as "There are 15 bags in stock." The server reflects this correction in the database and performs recalculation. Based on the results of periodic inventory, the server also generates management advice for ordering and stock disposal and sends it to the terminal. The user can make effective management decisions based on this advice.
[0255] Specific examples
[0256] For example, analysis may reveal that sales data for a particular product category (e.g., snacks) are sluggish when placed on the bottom shelf. Based on this, the server suggests placing snacks in front of the cash register or at eye level. The suggestions are sent to the device as a concrete visual map, and the user adjusts product placement accordingly. At the same time, an emotion engine analyzes the user's emotional state, providing detailed suggestions if the user has a positive reaction and simple suggestions if the user has a negative reaction. For example, if a voice notification reads, "There are 10 bags of snacks in stock," the user can correct it with a voice command, saying, "There are 15 bags in stock," and the system will immediately update the information.
[0257] These functions provide a system that improves the efficiency of product management and inventory management, increases sales, responds flexibly to user emotions, and enables more accurate management decisions.
[0258] The processing flow will be explained below.
[0259] Step 1:
[0260] Users take a photo of the product shelves in a store using a device such as a smartphone or dedicated camera, and by pressing the capture button, the image is automatically saved on the device.
[0261] Step 2:
[0262] The device automatically sends the captured image data to a server via Wi-Fi or mobile networks.
[0263] Step 3:
[0264] The server analyzes the image data received from the device and uses an image analysis engine to identify products in the image and determine the location and quantity of each product.
[0265] Step 4:
[0266] The server stores the identified product information (SKU, location, stock quantity, etc.) in a database, allowing the current stock status to be immediately grasped.
[0267] Step 5:
[0268] The server references sales data based on the information stored in the database, extracts past sales data using SQL queries, and analyzes the relationship between product placement and sales.
[0269] Step 6:
[0270] The server uses a generative AI model to analyze product placement and sales data, thereby identifying optimal product placement patterns.
[0271] Step 7:
[0272] The server generates optimal product placement proposals based on the analysis results, which are then visualized.
[0273] Step 8:
[0274] The server sends the generated placement proposal to the terminal, and the user can check the proposal displayed on the terminal.
[0275] Step 9:
[0276] The device will read the notification content aloud to the user using text-to-speech (TTS) technology.
[0277] Step 10:
[0278] The device recognizes the user's emotional state using an emotion engine that analyzes the user's voice and facial expressions in real time.
[0279] Step 11:
[0280] Based on the results of the emotion engine, the server dynamically adjusts the notification content, for example, if the user is feeling stressed, it will provide concise, positive suggestions.
[0281] Step 12:
[0282] The user can correct the contents of the voice message by entering a voice command, for example, "There are 15 bags in stock."
[0283] Step 13:
[0284] The server receives the user's voice command and updates the database with the corrected information, which is then recalculated and saved as the latest information.
[0285] Step 14:
[0286] The server generates management advice for ordering and stock disposal based on the results of periodic inventory. This advice is generated based on the amount of stock and sales trends.
[0287] Step 15:
[0288] The server then sends the generated management advice to the terminal, and the user makes decisions about ordering and inventory disposal based on this advice.
[0289] Through these steps, this system will improve the efficiency of product management and inventory management, increase sales, improve the accuracy of management decisions, and even realize flexible responses that take user emotions into consideration. As a specific example, when the stock of a specific product is low, the system will notify the user by saying, "Stock is low. Please consider placing an order," but if the user is feeling emotionally stressed, it will adjust the wording to a softer one, saying, "There is no need to take immediate action. Please check when you have time."
[0290] Example 2
[0291] Next, a description will be given of Example 2. In the following description, the data processing device 12 will be referred to as a "server" and the smart device 14 will be referred to as a "terminal."
[0292] In the retail industry, it is extremely important to accurately analyze the impact that product placement has on sales and propose optimal placements. However, currently, there are limited methods for optimizing shelf placement and proposal systems that take users' emotional states into account. This has prevented efficient inventory management and increased sales from being fully realized. Furthermore, when incorrect product information is identified, correcting it is cumbersome, requiring quick and appropriate management decisions. Given this background, a system is needed that can automatically analyze the relationship between product placement and sales and respond flexibly while taking emotions into consideration.
[0293] The specific processing by the specific processing unit 290 of the data processing device 12 in the second embodiment is realized by the following means.
[0294] In this invention, the server includes means for acquiring image data from product shelves in a store, means for analyzing the acquired image data to identify the product locations and quantities, means for storing the identified information in a database, means for using a generative model to analyze the relationship between product placement and sales by referencing sales data based on the information stored in the database, means for proposing optimal product placement based on the analysis results, means for visualizing the proposal and notifying the terminal, means for reading the notification content aloud, and means for recognizing the user's emotional state from the terminal and dynamically adjusting the notification content based on the user's emotional state. This enables accurate analysis of the relationship between product placement and sales and proposals for optimal placement that take emotions into consideration. It also makes it easier to correct erroneous judgments and generate management advice, enabling quick and accurate management decisions.
[0295] 1. "Image data" refers to digital information, such as still images or videos, that record the state of product shelves in a store.
[0296] 2. "Analysis" refers to the computational process used to identify the location and quantity of products from the acquired image data.
[0297] 3. "Database" means a system for systematically storing and managing specific product information and sales data.
[0298] 4. "Sales Data" means numerical information on the sales of products within a specific period of time.
[0299] 5. A "generative model" is a machine learning algorithm that learns from large amounts of data to derive trends and patterns.
[0300] 6. "Visualization" refers to displaying analysis results and placement proposals as graphs and charts so that users can intuitively understand them.
[0301] 7. "Terminal" means a computer device used by a user to operate the device, in particular a smartphone or tablet device.
[0302] 8. "Notification Content" means the output of information or suggestions provided by the System to the User in the form of text or audio.
[0303] 9. "Emotional state" is state information that represents the user's emotional response and is identified from the user's voice and facial expression.
[0304] 10. "Dynamic adjustment" means changing your response to the situation in real time.
[0305] This invention is a system that acquires image data of product shelves in a store, analyzes the image data to identify the location and quantity of products, and then analyzes sales data based on this to propose optimal product placement.The system uses image recognition technology, generative AI, and an emotion engine, and is able to flexibly respond by taking into account the user's emotional state.
[0306] Equipment and software configuration
[0307] The system consists of the following main components:
[0308] Terminal: A device equipped with a camera for taking pictures of product shelves in a store. Examples include smartphones and dedicated cameras.
[0309] Server: The central data processing unit that receives and analyzes image data. It runs the generative AI model and emotion engine.
[0310] User: The person who operates the system and makes optimal product placement and management decisions.
[0311] Processing steps
[0312] 1. Image data collection
[0313] A user uses a device to take an image of a product shelf in a store. For example, they can use a smartphone camera app to take a picture, focusing on a specific area. The captured image is set to be automatically sent to the server. A dedicated app encodes the image data and sends it to the server using a secure communication protocol (e.g., HTTPS).
[0314] 2. Analysis of image data
[0315] The server analyzes the image data received from the device. Specifically, it uses software such as OpenCV and TensorFlow to identify the location and quantity of products. For example, it uses the YOLO (You Only Look Once) object detection algorithm. The analysis results are stored in a database.
[0316] 3. Analysis of the relationship between placement and sales
[0317] The server references product information and sales data stored in a database and analyzes the relationship between product placement and sales performance using a generative AI model (e.g., H2O.ai's AutoML function). This analysis identifies the optimal placement pattern for a specific product.
[0318] 4. Proposal for optimal layout
[0319] Based on the analysis results, the server generates optimal product placement proposals, which are visualized and displayed in an intuitive format using Tableau or Power BI, and the proposals are sent to the device and notified to the user.
[0320] 5. Introducing the Emotion Engine
[0321] The device reads product information and stock availability to the user aloud, while simultaneously recognizing the user's emotional state using the Microsoft Azure Emotion API and IBM Watson Tone Analyzer. Based on the user's emotional state, the device dynamically adjusts the content of notifications.
[0322] 6. Correcting misjudgments and providing management advice
[0323] The user can make corrections to the information read out loud using voice commands. For example, they can input a correction such as "There are 15 bags in stock." The server reflects this correction in the database and performs recalculation. In addition, based on the results of periodic inventory, the server generates management advice for ordering and stock disposal and sends it to the terminal.
[0324] Specific examples
[0325] For example, if snacks are placed on the bottom shelf, analysis may reveal poor sales data. Based on this, the server suggests placing snacks at eye level. The suggestions are sent to the device as a concrete visual map, allowing the user to adjust product placement. At the same time, an emotion engine analyzes the user's emotional state, providing detailed suggestions if the user has a positive reaction and simple suggestions if the user has a negative reaction. For example, if a voice notification reads, "There are 10 bags of snacks in stock," the user can correct it with a voice command, saying, "There are 15 bags in stock," and the system will immediately update the information.
[0326] Prompt Sentence Examples
[0327] 1. "What image recognition technology do you use to optimize product placement in your stores?"
[0328] 2. "Please show the results of analyzing sales data using a generative AI model."
[0329] 3. "How does the Emotion Engine analyze the user's emotional state and adjust its response?"
[0330] This system will enable more efficient product and inventory management, increased sales, flexible responses that take user feelings into consideration, and more accurate management decisions.
[0331] The flow of the identification process in the second embodiment will be described with reference to FIG.
[0332] Step 1:
[0333] The user uses a device to take a picture of a specific product shelf in a store. The input is a high-resolution image taken using a smartphone camera app, for example. The output is image data with a clear focus on the captured product. Specifically, the image is taken so that the product's position and quantity are visible, and the image data is saved on the device.
[0334] Step 2:
[0335] The device sends the captured image data to the server. The input is the image data acquired in step 1. The output is the image data securely sent to the server. Specifically, a dedicated application encodes the image data into Base64 format and uploads it to the server in real time using the HTTPS protocol.
[0336] Step 3:
[0337] The server analyzes the image data received from the device. The input is the image data sent in step 2. The output is data on the location and quantity of products identified through image analysis. Specifically, it uses TensorFlow and OpenCV to apply object detection algorithms such as YOLO to extract the bounding boxes and labels of each product in the image.
[0338] Step 4:
[0339] The server stores the analysis results in a database. The input is the product location and quantity data identified in step 3. The output is product information stored in the database. Specifically, it is inserted into a MySQL database using an ORM (Object-Relational Mapping) library such as SQLAlchemy.
[0340] Step 5:
[0341] The server retrieves sales data from the database and uses a generative model to analyze the relationship between product placement and sales. The input is the product information and sales data stored in the database. The output is the optimal product placement pattern. Specifically, it uses H2O.ai's AutoML function to derive placement patterns that are expected to increase sales.
[0342] Step 6:
[0343] The server generates optimal product placement proposals based on the analysis results. The input is the optimal product placement pattern obtained in step 5. The output is a visualized placement proposal. Specifically, it is visualized as intuitive graphs and charts using Tableau or Power BI.
[0344] Step 7:
[0345] The server sends the visualized proposal to the terminal and notifies the user. The input is the placement proposal visualized in step 6. The output is the proposal notified to the user. Specifically, it is visually displayed on the terminal via the Internet.
[0346] Step 8:
[0347] The device reads out product information and stock status aloud, while simultaneously recognizing the user's emotional state. The input is the recommendations sent from the server and the user's real-time voice data. The output is the results of recognizing the user's emotional state and the adapted notification content. Specifically, it uses the Microsoft Azure Emotion API and IBM Watson Tone Analyzer to analyze the user's voice and facial expressions.
[0348] Step 9:
[0349] The server dynamically adjusts the notification content based on the emotional state identified by the emotion engine. The input is the user's emotional state sent from the device. The output is the adjusted notification content. For example, if the user is feeling stressed, more concise and positive suggestions are provided.
[0350] Step 10:
[0351] The user inputs a voice command in response to the voice readout, and the server reflects the correction in the database. The input is the user's voice command and the current contents of the database. The output is updated information in the database that reflects the correction. For example, if the correction is made to "There are 15 bags in stock," the information is immediately updated in the database.
[0352] As a result, this system analyzes the relationship between product placement in a store and sales, and realizes flexible and adaptable placement suggestions and inventory management that take into account the user's emotional state.
[0353] (Application example 2)
[0354] Next, a description will be given of Application Example 2. In the following description, the data processing device 12 will be referred to as a "server" and the smart device 14 will be referred to as a "terminal."
[0355] Conventional product placement support systems for the retail industry require a lot of effort to identify product locations and inventory information, and their placement change suggestions do not take into account the user's emotional state, making it difficult to maximize management efficiency and sales. Furthermore, because real-time product placement changes and inventory adjustments are time-consuming, there is a demand for faster management decision-making.
[0356] The identification process by the identification processing unit 290 of the data processing device 12 in Application Example 2 is realized by the following means. In this invention, the server includes means for acquiring image data from product shelves in a store, means for analyzing the acquired image data and identifying the location and quantity of products, means for saving the identified information in a database, means for referencing sales data based on the information saved in the database and analyzing the relationship between product placement and sales, means for proposing an optimal product placement based on the analysis results, means for visualizing the proposal and notifying the user, means for reading the notification content aloud to the user, means for a user to use a smart device to take a real-time photo of the product shelf situation in the store and send the image data to the server, means for analyzing sales data and product information using a generative AI model, means for recognizing the user's emotional state and adjusting the notification content based on the user's emotion, and means for providing real-time information updates and advice in cooperation with a voice assistant. This enables efficient product management and inventory management, maximizing sales, flexible responses based on user emotions, and faster management decisions.
[0357] "Image data" is image information obtained by photographing a product shelf, and is data used to identify the position and quantity of products.
[0358] "In-store shelves" refers to shelves and racks used to display, showcase, and store products in retail stores and brick-and-mortar stores.
[0359] "Means" means a method, technique, or device for achieving a particular purpose.
[0360] "Analysis" refers to the process of analyzing the acquired image data and extracting detailed information such as the location and quantity of the products.
[0361] A "database" is an information system that can efficiently store, manage, and search large amounts of data.
[0362] "Sales data" refers to data containing sales information for a product within a certain period of time, including sales amount, sales quantity, sales date, and the like.
[0363] A "generative AI model" is an algorithm that uses artificial intelligence to analyze data and generate new insights and predictions.
[0364] "Visualization" is the process of presenting data or information in a visually understandable way.
[0365] "User" refers to the person who operates the system and manages product placement and inventory.
[0366] "Notification" is the act of transmitting information from the system to the user.
[0367] A "voice assistant" is a system that interacts with users using voice input to assist them in obtaining information and performing operations.
[0368] "Smart devices" refer to electronic devices equipped with advanced computing power and communication functions, including smartphones and smart glasses.
[0369] "Emotional state" refers to a user's emotional response or state, including happiness, sadness, stress, etc.
[0370] "Real-time" means processing or reacting nearly simultaneously or without delay.
[0371] System Configuration
[0372] This system acquires images of in-store product shelves and optimizes product placement using generative AI models and emotion engines. The system consists of the following main hardware and software components:
[0373] 1. Terminal
[0374] Smart devices include smart glasses and smartphones, which are equipped with cameras and are used to capture images of the product shelves in stores.
[0375] 2. Server
[0376] OpenCV is used for image analysis, which analyzes image data and identifies the location and quantity of products.
[0377] The database used is MySQL, and the analyzed data is saved in the database.
[0378] The generative AI model uses OpenAI's GPT-4 and other models to analyze sales data and product information and generate optimal placement patterns.
[0379] The emotion engine uses Affectiva's SDK to analyze the user's emotional state.
[0380] As a voice assistant, it uses Google Cloud Text-to-Speech to read out the generated notification content aloud.
[0381] 3. Users
[0382] Managers use smart devices to take photos of the shelves and send the data to a server in real time.
[0383] Processing steps
[0384] Image data collection
[0385] The user uses smart glasses or a smartphone to take an image of the store's shelves, focusing on a specific area of the shelves to clearly show the location and quantity of products. The device is set to automatically send the captured image to the server.
[0386] Image data analysis
[0387] The server analyzes the image data received from the device. The image analysis engine uses OpenCV to identify the location, quantity, and specific product name of each product. After the product information is identified, it is saved in a database.
[0388] Saving to a database
[0389] The identified product information is stored in a database along with attributes such as SKU, location, and inventory quantity, making it possible to immediately grasp the current inventory status.
[0390] Analysis of the relationship between placement and sales
[0391] The server references sales data from a database for a set period and uses a generative AI model to analyze the relationship between product placement and sales performance. This analysis identifies the optimal placement pattern for specific products, maximizing sales.
[0392] Proposal for optimal layout
[0393] Based on the analysis results, the server generates an optimal product placement proposal. This proposal is visualized, allowing the user to compare the current product shelf layout with the proposed optimal placement. The proposal is then sent to the terminal and notified to the user.
[0394] Introducing the Emotion Engine
[0395] The device reads product information and inventory status to the user aloud and uses an emotion engine to recognize the user's emotional state by analyzing the user's voice and facial expressions to identify their emotions.
[0396] Adjusting notification content
[0397] Based on the emotional state identified by the emotion engine, the server dynamically adjusts the notification content, for example, providing more concise and positive suggestions if the user is feeling stressed.
[0398] Correcting misjudgments and providing management advice
[0399] The user can make corrections to the voice-read information by voice command. For example, the user can input a correction such as "15 bags in stock." The server reflects this correction in the database and performs recalculation. In addition, based on the results of periodic inventory, the server generates management advice for ordering and stock disposal and sends it to the terminal. The user can make effective management decisions based on this advice.
[0400] Specific examples
[0401] For example, analysis may reveal that sales data for a particular product category (e.g., snacks) are sluggish when placed on the bottom shelf. Based on this, the server suggests placing snacks in front of the register or at eye level. The suggestions are sent to the device as a concrete visual map, and the user adjusts product placement accordingly. The emotion engine also analyzes the user's emotional state, providing detailed suggestions if the user has a positive reaction and simple suggestions if the user has a negative reaction.
[0402] Prompt Sentence Examples
[0403] "Please suggest the optimal product placement in the store. The current placement and sales data are as follows: \[Specific data\]"
[0404] The flow of the specific processing in the application example 2 will be described with reference to FIG.
[0405] Step 1:
[0406] A user uses a smart device (smart glasses or smartphone) to take a picture of the product shelf situation in a store in real time. The captured image data is automatically sent from the device to a server. The input is the image of the product shelf, and the output is the image data sent to the server.
[0407] Step 2:
[0408] The server analyzes the received image data. At this time, an image analysis engine (such as OpenCV) is used to identify the location and quantity of each product. Specifically, object recognition and extraction are performed on the image data. The input is the received image data, and the output is specific information such as the product's location, quantity, and product name.
[0409] Step 3:
[0410] The server stores the identified product information in a database (e.g., MySQL). Here, attributes such as SKU, location, and stock quantity are stored as data. The input is analysis information such as product location and quantity, and the output is product information stored in the database.
[0411] Step 4:
[0412] The server references the information stored in the database and sales data, and uses a generative AI model to analyze the relationship between product placement and sales data. At this time, the generative AI model (such as OpenAI's GPT-4) receives a prompt message: "Please suggest the optimal placement of products in the store. The current placement data and sales data are as follows: \[Specific data\]". The input is the current placement data and sales data, and the output is a proposal for the optimal product placement.
[0413] Step 5:
[0414] The server visualizes the generated optimal placement proposal and sends the results to the terminal. Specifically, a visual map is created. The input is the placement proposal obtained from the generative AI model, and the output is the visualized placement proposal and a notification of it.
[0415] Step 6:
[0416] The device reads out product information and stock availability to the user and analyzes the user's emotional state using an emotion engine (such as Affectiva's SDK). The emotion engine recognizes emotions and sends the results to the server. The input is the user's voice and facial expression, and the output is the identified emotional state.
[0417] Step 7:
[0418] The server dynamically adjusts the notification content based on the emotional state obtained from the emotion engine. If the user is feeling stressed, it provides more concise and positive suggestions, and if the user has a positive reaction, it provides more detailed suggestions. The input is the emotional state obtained from the emotion engine, and the output is the adjusted notification content.
[0419] Step 8:
[0420] The user can make corrections to the voice readout as necessary using voice commands. For example, they can input the corrections, such as "There are 15 bags in stock." The server reflects this correction information in the database and performs recalculation. The input is the user's voice command, and the output is the corrected database information.
[0421] Step 9:
[0422] The server generates management advice for ordering and stock disposal based on the results of periodic inventory and sends it to the terminal. The user makes effective management decisions based on this management advice. The input is the inventory result data, and the output is the generated management advice.
[0423] The specific processing unit 290 transmits the result of the specific processing to the smart device 14. In the smart device 14, the control unit 46A causes the output device 40 to output the result of the specific processing. The microphone 38B acquires audio indicating a user input regarding the result of the specific processing. The control unit 46A transmits audio data indicating the user input acquired by the microphone 38B to the data processing device 12. In the data processing device 12, the specific processing unit 290 acquires the audio data.
[0424] The data generation model 58 is a so-called generative AI (Artificial Intelligence). An example of the data generation model 58 is ChatGPT (Internet Search<URL: https: / / openai.com / blog / chatgpt> ), Gemini (Internet search <url: https: gemini.google.com ?hl="ja">) and other generation AIs. The data generation model 58 is obtained by performing deep learning on a neural network. A prompt including an instruction is input to the data generation model 58, and inference data such as voice data indicating voice, text data indicating text, and image data indicating an image is also input. The data generation model 58 performs inference on the input inference data in accordance with the instruction indicated by the prompt, and outputs the inference result in a data format such as voice data and text data. Here, inference refers to, for example, analysis, classification, prediction, and / or summarization.
[0425] In the above embodiment, an example in which the specific process is performed by the data processing device 12 has been given, but the technology of the present disclosure is not limited to this, and the specific process may be performed by the smart device 14.
[0426] [Second embodiment]
[0427] FIG. 3 shows an example of the configuration of a data processing system 210 according to the second embodiment.
[0428] 3, the data processing system 210 includes the data processing device 12 and smart glasses 214. An example of the data processing device 12 is a server.
[0429] The data processing device 12 includes a computer 22, a database 24, and a communication I / F 26. The computer 22 is an example of a "computer" according to the technology of the present disclosure. The computer 22 includes a processor 28, a RAM 30, and a storage 32. The processor 28, the RAM 30, and the storage 32 are connected to a bus 34. The database 24 and the communication I / F 26 are also connected to the bus 34. The communication I / F 26 is connected to a network 54. Examples of the network 54 include a WAN (Wide Area Network) and / or a LAN (Local Area Network).
[0430] The smart glasses 214 include a computer 36, a microphone 238, a speaker 240, a camera 42, and a communication I / F 44. The computer 36 includes a processor 46, a RAM 48, and a storage 50. The processor 46, the RAM 48, and the storage 50 are connected to a bus 52. The microphone 238, the speaker 240, and the camera 42 are also connected to the bus 52.
[0431] The microphone 238 receives instructions and the like from the user 20 by receiving voice uttered by the user 20. The microphone 238 captures the voice uttered by the user 20, converts the captured voice into audio data, and outputs it to the processor 46. The speaker 240 outputs audio in accordance with instructions from the processor 46.
[0432] Camera 42 is a small digital camera equipped with an optical system including a lens, aperture, and shutter, and an imaging element such as a CMOS (Complementary Metal-Oxide-Semiconductor) image sensor or a CCD (Charge Coupled Device) image sensor, and captures images of the surroundings of user 20 (for example, an imaging range defined by an angle of view equivalent to the field of vision of a typical healthy person).
[0433] The communication I / F 44 is connected to a network 54. The communication I / Fs 44 and 26 are responsible for the exchange of various information between the processor 46 and the processor 28 via the network 54. The exchange of various information between the processor 46 and the processor 28 using the communication I / Fs 44 and 26 is carried out in a secure state.
[0434] Fig. 4 shows an example of the main functions of the data processing device 12 and the smart glasses 214. As shown in Fig. 4, in the data processing device 12, a specific process is performed by the processor 28. A specific process program 56 is stored in the storage 32.
[0435] The specific processing program 56 is an example of a "program" according to the technology of the present disclosure. The processor 28 reads the specific processing program 56 from the storage 32 and executes the read specific processing program 56 on the RAM 30. The specific processing is realized by the processor 28 operating as a specific processing unit 290 in accordance with the specific processing program 56 executed on the RAM 30.
[0436] The storage 32 stores a data generation model 58 and an emotion identification model 59. The data generation model 58 and the emotion identification model 59 are used by the identification processing unit 290.
[0437] In the smart glasses 214, the reception output process is performed by the processor 46. A reception output program 60 is stored in the storage 50. The processor 46 reads the reception output program 60 from the storage 50 and executes the read reception output program 60 on the RAM 48. The reception output process is realized by the processor 46 operating as the control unit 46A in accordance with the reception output program 60 executed on the RAM 48.
[0438] Next, a description will be given of the identification process performed by the identification processing unit 290 of the data processing device 12. In the following description, the data processing device 12 will be referred to as the "server" and the smart glasses 214 will be referred to as the "terminal."
[0439] The present invention relates to an optimal product placement support system for the retail industry that uses image recognition technology and generative AI, and aims to analyze the impact of product placement in a store on sales and optimize product placement. Specific implementation methods of the present invention are described below.
[0440] System configuration
[0441] The system consists of the following main components:
[0442] 1. Terminal: Equipped with a camera to take pictures of product shelves in a store and has the function of sending image data to a server. Examples include smartphones and dedicated cameras.
[0443] 2. Server: The central data processing unit that receives and analyzes image data, running the generative AI model.
[0444] 3. User: The person who operates the system and makes management decisions such as optimal product placement, ordering, and inventory disposal.
[0445] Program processing flow
[0446] Image data collection
[0447] The user uses the device to take an image of the store's shelves, focusing on a specific area of the shelves to clearly show the location and quantity of products. The device is configured to automatically send the captured image to the server.
[0448] Image data analysis
[0449] The server analyzes the image data received from the device. Using an image analysis engine, it identifies the location, quantity, and specific product name of each product. After the product information is identified, it is stored in a database.
[0450] Saving to a database
[0451] The identified product information is stored in a database along with attributes such as SKU, location, and inventory quantity, making it possible to immediately grasp the current inventory status.
[0452] Analysis of the relationship between placement and sales
[0453] The server references sales data from a database for a set period and uses a generative AI model to analyze the relationship between product placement and sales performance. This analysis identifies the optimal placement pattern for specific products, maximizing sales.
[0454] Proposal for optimal layout
[0455] Based on the analysis results, the server proposes the optimal product placement. This proposal is visualized, allowing users to compare the current product shelf layout with the proposed optimal placement. The proposal is then sent to the terminal and notified to the user.
[0456] Voice reading and correction
[0457] Users can confirm the notification content from the device by voice. The device will read out product information and stock status using text-to-speech (TTS) technology. If there is an incorrect judgment, users can correct it using voice commands. This correction is reflected in the database again and recalculation is performed.
[0458] Management Advice
[0459] Based on the results of periodic inventory, the server generates management advice such as ordering and discount proposals for inventory clearance. These advices are notified to the user via text and voice, supporting effective management decisions.
[0460] Specific examples
[0461] For example, analysis may reveal that sales data for a particular product category (e.g., snacks) are stagnating if they are placed on the bottom shelf. Based on this, the server may suggest placing snacks in front of the register or at eye level. The suggestion is sent to the device as a concrete visual map, and the user can adjust product placement based on the suggestion. Also, if a voice notification reads, "There are 10 bags of snacks in stock," the user can correct it with a voice command, saying, "There are 15 bags in stock," and the system will immediately update the information.
[0462] These features provide a system that streamlines product and inventory management, increases sales, and enables more accurate business decisions.
[0463] The processing flow will be explained below.
[0464] Step 1:
[0465] Users take a photo of the product shelves in a store using a device such as a smartphone or dedicated camera, and by pressing the capture button, the image is automatically saved on the device.
[0466] Step 2:
[0467] The device automatically sends the captured image data to a server via Wi-Fi or mobile networks.
[0468] Step 3:
[0469] The server analyzes the image data received from the device and uses an image analysis engine to identify products in the image and determine the location and number of each product.
[0470] Step 4:
[0471] The server stores the identified product information in a database, including the SKU, location, and quantity in stock.
[0472] Step 5:
[0473] The server references sales data based on the information stored in the database, extracts past sales data using SQL queries, and analyzes the relationship between product placement and sales.
[0474] Step 6:
[0475] The server uses a generative AI model to analyze product placement and sales data, thereby identifying optimal product placement patterns.
[0476] Step 7:
[0477] The server generates optimal product placement proposals based on the analysis results, which are then visualized.
[0478] Step 8:
[0479] The server sends the generated placement proposal to the terminal, and the user can check the proposal displayed on the terminal.
[0480] Step 9:
[0481] The device reads product information and stock availability to the user aloud, and uses text-to-speech (TTS) technology to audibly output the contents of the visual suggestions.
[0482] Step 10:
[0483] The user can make corrections using voice commands in response to the content of the voice readout. For example, the user can input a correction such as "There are 15 bags in stock."
[0484] Step 11:
[0485] The server receives voice correction commands from the user and reflects them in the database, where the corrected information is recalculated and saved as the latest information.
[0486] Step 12:
[0487] The server generates management advice for ordering and stock disposal based on the results of periodic inventory. This advice is generated based on the amount of stock and sales trends.
[0488] Step 13:
[0489] The server then sends the generated management advice to the terminal, and the user makes decisions about ordering and inventory disposal based on this advice.
[0490] Through the above steps, this system will improve the efficiency of product management and inventory management, thereby increasing sales and improving the accuracy of management decisions.
[0491] Example 1
[0492] Next, a description will be given of Example 1. In the following description, the data processing device 12 will be referred to as a "server" and the smart glasses 214 will be referred to as a "terminal."
[0493] The retail industry is seeking to optimize product placement and streamline inventory management. Conventional methods rely on manual work to determine product locations and quantities, which takes time and effort and is prone to errors. It is also difficult to analyze the relationship between product placement and sales, requiring significant effort to find optimal product placement. Furthermore, updating inventory information in real time is difficult, making it difficult to make quick management decisions. To solve these issues, a system is needed that can efficiently and accurately manage product placement and inventory.
[0494] The specific processing by the specific processing unit 290 of the data processing device 12 in the first embodiment is realized by the following means.
[0495] In this invention, the server includes a means for analyzing image data and identifying product locations and quantities, a means for proposing optimal product placement using a generative AI model, and a means for referencing sales data based on information stored in a database and analyzing the relationship between product placement and sales. This makes it possible to identify product locations and quantities with high accuracy, update inventory information in real time, and quickly propose optimal product placement to maximize sales.
[0496] "Image data" refers to visual information acquired in the form of photographs or videos of product shelves in a store.
[0497] "Acquisition" refers to the act of obtaining image data of an object using a camera or other photographic equipment.
[0498] "Analysis" is the process of extracting and identifying specific information from acquired image data.
[0499] A "generative AI model" is an artificial intelligence algorithm trained to perform image recognition and data analysis.
[0500] "Location" is information indicating the specific location where a particular product is located.
[0501] "Quantity" is a number that indicates how much of a particular product is on the shelf or in stock.
[0502] A "database" is an electronic recording device or system for systematically storing and managing specific information.
[0503] "Sales data" refers to transaction information when a product is sold, including sales amount and sales quantity.
[0504] "Analysis" is the act of examining and interpreting data to find specific relationships and patterns.
[0505] "Suggestion" refers to showing optimal behavior and placement patterns based on the analysis results.
[0506] "Visualization" refers to displaying data and proposals in a visually easy-to-understand manner.
[0507] "Notification" is the act of informing a user of specific information.
[0508] "Text-to-speech" is a technology that converts text information into speech and conveys it to the user auditorily.
[0509] "Voice command" is a method by which a user gives instructions or inputs using their voice.
[0510] "Inventory" refers to the periodic checking and recording of stock and the quantity of merchandise items.
[0511] "Management advice" refers to proposals based on inventory and sales data to support decision-making such as ordering and inventory disposal.
[0512] This invention is an optimal product placement support system for the retail industry that uses image recognition technology and generative AI models. The system acquires image data of product shelves in stores, analyzes the data to identify product locations and quantities, stores the data in a database, and then compares it with sales data to suggest optimal product placement.
[0513] The system consists of three main components:
[0514] 1. Terminal: A device with a camera function that takes pictures of product shelves in a store. Typical examples are smartphones or dedicated cameras.
[0515] 2. Server: A central data processing unit that receives data, performs image analysis, and optimizes product placement using generative AI models.
[0516] 3. User: The person who operates the system and makes management decisions such as optimal product placement, ordering, and inventory disposal.
[0517] Hardware and Software Configuration
[0518] Terminal
[0519] Users take pictures of product shelves in stores using smartphones or dedicated cameras. These devices have the ability to automatically send the captured image data to a server. The captured images are sent to the server via Wi-Fi or mobile data. The devices are also equipped with text-to-speech (TTS) technology, which allows them to notify users by voice.
[0520] server
[0521] The server receives the image data sent from the device and uses an image analysis engine to identify the location, quantity, and product name of each product. This analysis uses deep learning technology and generative AI models, such as machine learning libraries like TensorFlow and PyTorch. The server then stores the identified information in a database and uses that data to analyze the relationship between sales data and product placement.
[0522] Database
[0523] The product information analyzed by the server is stored in a database, which updates attributes such as SKU (Stock Keeping Unit), location, and inventory quantity in real time. Users can easily check this information.
[0524] Examples of use and specific prompts
[0525] Users take pictures of product shelves in a store and send them to a server using their device. The server then receives the image data, analyzes it, identifies the product locations and quantities, and stores them in a database. Based on the analysis results, the server uses a generative AI model to propose optimal product placement and sends it to the device as a visual map. The user can then confirm the proposal visually and audibly and adjust the product placement.
[0526] Prompt Sentence Examples
[0527] "Please identify the names and quantities of the items shown in this image."
[0528] "Please suggest optimal product placement based on sales data and product placement data."
[0529] "Please read out the inventory."
[0530] · "Use voice commands to correct stock levels."
[0531] This will lead to more efficient product and inventory management, increased sales, and more accurate business decisions.
[0532] The flow of the identification process in the first embodiment will be described with reference to FIG.
[0533] Step 1:
[0534] The user uses a device to take a photo of a product shelf in a store. The captured image is high resolution, and the product name and quantity are clearly visible. For example, a smartphone or dedicated camera is used to capture product images under appropriate lighting. The input is an "image of a product shelf in a store," and the output is "high-resolution image data."
[0535] Step 2:
[0536] The device sends the captured image data to the server. The device has a network connection function and encrypts and sends the image data to the server using Wi-Fi or mobile data. The input is "high-resolution image data" and the output is "a notification of completion of transmission to the server."
[0537] Step 3:
[0538] The server analyzes the image data received from the device. The image analysis engine processes the input image and identifies the location, quantity, and name of the product. A generative AI model is used for this analysis. Specifically, machine learning libraries such as TensorFlow and PyTorch are used. The input is "high-resolution image data," and the output is "the location, quantity, and name of the identified product."
[0539] Step 4:
[0540] The server stores the identified information in a database. The identified product information (SKU, location, and stock quantity) is recorded in the database, and inventory information is updated in real time. The input is the "identified product location, quantity, and name," and the output is the "updated database entry."
[0541] Step 5:
[0542] The server analyzes the relationship between sales data and product placement performance based on the stored information. It uses a generative AI model to find the optimal placement pattern. For example, it compares the placement location of a specific product with sales data to identify the placement that maximizes sales. The inputs are "product information stored in the database" and "sales data," and the output is "identification of the optimal placement pattern."
[0543] Step 6:
[0544] The server proposes optimal product placement and generates a visualized proposal. It provides the user with a layout diagram in an easy-to-understand visual format, allowing them to compare the proposed placements. The proposal is sent to the device. The input is the "identification result of the optimal placement pattern," and the output is the "visualized placement proposal."
[0545] Step 7:
[0546] The terminal notifies the user of the proposal content sent from the server by voice. Using text-to-speech (TTS) technology, the placement proposal content and inventory information are read aloud. The input is a "visualized placement proposal" and the output is a "voice notification."
[0547] Step 8:
[0548] When a user wants to correct a voice notification, they input a voice command. For example, they can correct an inventory quantity by voice, and the correction is sent to the server. The server analyzes the voice command and updates the database. The input is the "user's voice command" and the output is the "updated database entry."
[0549] Step 9:
[0550] Based on the results of periodic inventory, the server generates management advice for ordering and stock disposal. It then notifies the user of this advice via text and voice. For example, the advice might be, "There is an excess of snack food in stock. We suggest selling them at a discount." The inputs are "inventory results" and "database information," and the output is "notification of management advice."
[0551] (Application example 1)
[0552] Next, a description will be given of Application Example 1. In the following description, the data processing device 12 will be referred to as a "server" and the smart glasses 214 will be referred to as a "terminal."
[0553] It is widely recognized in the retail industry that product placement in a store has a significant impact on sales. However, optimizing shelf placement using traditional methods requires a great deal of effort and time, and effective placement is not always achieved. Furthermore, real-time inventory management and rapid management decisions based on sales data are required, but traditional systems do not adequately support this. Furthermore, store staff are required to manually check and correct product locations and inventory status, which is time-consuming and inefficient.
[0554] The specific processing by the specific processing unit 290 of the data processing device 12 in the application example 1 is realized by the following means.
[0555] In this invention, the server includes means for acquiring image data from product shelves in a store, means for analyzing the acquired image data and identifying product locations and quantities, means for saving the identified information in a database, means for referencing sales data based on the information saved in the database and analyzing the relationship between product placement and sales, means for proposing optimal product placement based on the analysis results, means for visualizing the proposal and notifying the user, means for reading the notification aloud to the user, means for taking pictures of the product shelves using a smartphone, means for providing audio notification of product information and inventory status using text-to-speech technology, means for the user to correct information using voice commands, means for generating optimal placements using a generative AI model that contributes to maximizing sales, and means for inputting prompts to the generative AI model to obtain analysis results. This enables efficient acquisition of product shelf images using a smartphone, optimal placement suggestions using generative AI, and rapid and accurate inventory management and management decisions using voice technology.
[0556] "Image data" refers to information that is acquired as an image of product shelves in a store and is the subject of analysis.
[0557] "Analysis" refers to the act of identifying the location and quantity of products from the acquired image data.
[0558] A "database" is a repository of information that stores identified product information and allows for later reference.
[0559] "Sales data" is information relating to product sales, and is used to analyze the relationship between product placement and sales.
[0560] A "generative AI model" is an artificial intelligence technology used to analyze image data and propose optimal product placement.
[0561] A "prompt statement" is an instruction statement that is input into a generative AI model to obtain analysis results.
[0562] "Visualization" is the act of visually suggesting optimal product placement.
[0563] "Voice notification" refers to the act of using text-to-speech technology to inform users of product information and stock status via voice.
[0564] A "smartphone" is a mobile device used to acquire and send notifications about image data.
[0565] A "voice command" is a spoken input by a user to give instructions to a system using voice recognition technology.
[0566] "Inventory management" is the management activity of determining the quantity of goods and maintaining appropriate stock levels.
[0567] "Management decisions" are decisions made regarding product placement and inventory management.
[0568] MODE FOR CARRYING OUT THE INVENTION
[0569] The present invention relates to a system for optimizing product placement in a retail store using a smartphone. A specific implementation method of this system will be described below.
[0570] System configuration
[0571] The system consists of the following major components:
[0572] 1. Device (smartphone):
[0573] It is equipped with a camera to take pictures of product shelves in the store and transmits the image data to a server.
[0574] It has the ability to provide voice notification of product information and stock status using text-to-speech technology.
[0575] It uses voice recognition technology to receive the user's voice commands and send correction information to the server.
[0576] 2. Server:
[0577] It receives image data and runs a generative AI model that performs the analysis.
[0578] Using an image analysis engine (e.g., YOLOv5, TensorFlow), the location, quantity, and specific product name of each product are identified.
[0579] The identified product information is stored in a database (e.g., Firebase, PostgreSQL).
[0580] Sales data is referenced based on the information stored in the database, and the relationship between product placement and sales is analyzed.
[0581] The system generates the optimal layout that contributes to maximizing sales and notifies the user as a visual map.
[0582] 3. User:
[0583] A smartphone is used to take a photo of the product shelf and send the image data to the server.
[0584] Visually confirm the proposed optimal placement and confirm / correct inventory information notified by voice.
[0585] Make business decisions and place orders or dispose of inventory based on advice provided by the system.
[0586] Program processing
[0587] (Collection of image data)
[0588] Users take pictures of product shelves in a store using the camera on their smartphone, and the smartphone automatically sends the captured image data to the server.
[0589] (Image data analysis)
[0590] The server is equipped with an image analysis engine and generative AI models, such as YOLOv5 and TensorFlow, to analyze the received image data. The analysis identifies the location, quantity, and SKU of each product, and stores this information in a database.
[0591] (Saving to database)
[0592] The identified product information is stored in a database along with attributes such as SKU, location, and stock quantity, making it possible to instantly grasp the current inventory status.
[0593] (Analysis of the relationship between placement and sales)
[0594] The server references sales data for a set period from the database and uses a generative AI model to analyze the relationship between product placement and sales performance. At this time, it inputs prompt statements into the generative AI to obtain the analysis results.
[0595] (Optimal layout proposal)
[0596] A placement pattern that contributes to maximizing sales is generated, and the proposed details are sent to a smartphone as a visual map.
[0597] (audio notification and fix)
[0598] The smartphone uses text-to-speech technology (e.g., Google Cloud Text-to-Speech API) to notify the user of product information and stock availability via voice. When the user provides corrections via voice commands, the smartphone sends the information to the server, which updates the database.
[0599] (Providing management advice)
[0600] Based on the results of periodic inventory, the server generates management advice for ordering and stock disposal, which is communicated to the user via text and voice to support effective management decisions.
[0601] Specific examples
[0602] For example, if a particular snack is placed on the bottom shelf, analysis may reveal that sales data is sluggish. The server then suggests placing the snack in the center of the shelf where it is more visible. This suggestion is sent to the smartphone as a visual map, and the user follows the instructions to rearrange the products. Also, if a voice notification says, "There are 10 bags of snacks in stock," the user can correct it by saying, "There are 15 bags in stock," and the system will immediately update the database.
[0603] (Example of a prompt to input to a generative AI model)
[0604] input:
[0605] "Analyze the shelf image below to identify product locations, SKUs, and stock quantities. SKUs: '12345', '54321', '67890', etc."
[0606] output:
[0607] "Product locations and quantities identified: 4 units of SKU: '12345', 6 units of SKU: '54321', and 10 units of SKU: '67890'. Suggested placement: Place snacks in the center of the shelf."
[0608] The flow of the specific processing in the application example 1 will be described with reference to FIG.
[0609] Program processing steps
[0610] Step 1:
[0611] A user uses the camera on their smartphone to take a picture of a product shelf in a store. The image data taken by the smartphone is the input. The device sends this image data to the server. The output is the image data sent to the server.
[0612] Step 2:
[0613] The server analyzes the received image data. Specifically, it uses an image analysis engine (e.g., YOLOv5, TensorFlow) to identify the product location and quantity. The image data is input, and the product location, quantity, and SKU information are output.
[0614] Step 3:
[0615] The server stores the identified product information in a database. The input is the location, quantity, and SKU information of the identified product, and the product information stored in the database is obtained as the output.
[0616] Step 4:
[0617] The server references sales data based on the information stored in the database and analyzes the relationship between product placement and sales. This analysis is performed using a generative AI model (e.g., TensorFlow). Database information and sales data are input, and the analysis results include information on the relationship between product placement and sales.
[0618] Step 5:
[0619] The server inputs a prompt into the generative AI model to obtain the optimal product placement that will maximize sales. The inputs are the prompt and information from the database, and the output is a proposal for the optimal placement.
[0620] Step 6:
[0621] The server generates a visual map of the optimal layout proposal and sends it to the smartphone. The input is the optimal layout proposal, and the output is the visual map data.
[0622] Step 7:
[0623] The smartphone uses text-to-speech technology (e.g., Google Cloud Text-to-Speech API) to provide the user with product information and stock status via voice. The input is a visual map and product information, and the output is a voice notification.
[0624] Step 8:
[0625] The user listens to the voice notification and, if necessary, issues a voice command to correct the information. The smartphone receives this voice command and converts it into text using voice recognition technology (e.g., Google Speech-to-Text API). The voice command is used as input, and the text of the correction is output.
[0626] Step 9:
[0627] The server receives the user's corrections and updates the database: the text of the corrections is the input, and the database update is the output.
[0628] Step 10:
[0629] Based on the results of periodic inventory, the server generates management advice for ordering and stock disposal and notifies the user. Inventory data is input and management advice is output.
[0630] Combining these steps will improve the efficiency of in-store product placement and inventory management, and will also enable faster management decisions to maximize sales.
[0631] Furthermore, an emotion engine that estimates the user's emotion may be further combined. That is, the identification processing unit 290 may estimate the user's emotion using the emotion identification model 59, and perform identification processing using the user's emotion.
[0632] The present invention relates to an optimal product placement support system for the retail industry that uses image recognition technology, generative AI, and an emotion engine. The purpose of the system is to analyze the impact of in-store product placement on sales, optimize product placement, and also take into account the emotional state of users. Specific implementation methods of the present invention are described below.
[0633] System configuration
[0634] The system consists of the following main components:
[0635] 1. Terminal: Equipped with a camera to take pictures of product shelves in a store and has the function of sending image data to a server. Examples include smartphones and dedicated cameras.
[0636] 2. Server: The central data processing unit that receives and analyzes image data, running the generative AI model and emotion engine.
[0637] 3. User: The person who operates the system and makes management decisions such as optimal product placement, ordering, and inventory disposal.
[0638] Program processing flow
[0639] Image data collection
[0640] The user uses the device to take an image of the store's shelves, focusing on a specific area of the shelves to clearly show the location and quantity of products. The device is configured to automatically send the captured image to the server.
[0641] Image data analysis
[0642] The server analyzes the image data received from the device. Using an image analysis engine, it identifies the location, quantity, and specific product name of each product. After the product information is identified, it is stored in a database.
[0643] Saving to a database
[0644] The identified product information is stored in a database along with attributes such as SKU, location, and inventory quantity, making it possible to immediately grasp the current inventory status.
[0645] Analysis of the relationship between placement and sales
[0646] The server references sales data from a database for a set period and uses a generative AI model to analyze the relationship between product placement and sales performance. This analysis identifies the optimal placement pattern for specific products, maximizing sales.
[0647] Proposal for optimal layout
[0648] Based on the analysis results, the server generates an optimal product placement proposal. This proposal is visualized, allowing the user to compare the current product shelf layout with the proposed optimal placement. The proposal is then sent to the terminal and notified to the user.
[0649] Introducing the Emotion Engine
[0650] The device reads product information and inventory status to the user aloud and uses an emotion engine to recognize the user's emotional state by analyzing the user's voice and facial expressions to identify their emotions.
[0651] Adjusting notification content
[0652] Based on the emotional state identified by the emotion engine, the server dynamically adjusts the notification content, for example, providing more concise and positive suggestions if the user is feeling stressed.
[0653] Correcting misjudgments and providing management advice
[0654] The user can make corrections to the voice-readout information by voice command. For example, the user can input a correction such as "There are 15 bags in stock." The server reflects this correction in the database and performs recalculation. Based on the results of periodic inventory, the server also generates management advice for ordering and stock disposal and sends it to the terminal. The user can make effective management decisions based on this advice.
[0655] Specific examples
[0656] For example, analysis may reveal that sales data for a particular product category (e.g., snacks) are sluggish when placed on the bottom shelf. Based on this, the server suggests placing snacks in front of the cash register or at eye level. The suggestions are sent to the device as a concrete visual map, and the user adjusts product placement accordingly. At the same time, an emotion engine analyzes the user's emotional state, providing detailed suggestions if the user has a positive reaction and simple suggestions if the user has a negative reaction. For example, if a voice notification reads, "There are 10 bags of snacks in stock," the user can correct it with a voice command, saying, "There are 15 bags in stock," and the system will immediately update the information.
[0657] These functions provide a system that improves the efficiency of product management and inventory management, increases sales, responds flexibly to user emotions, and enables more accurate management decisions.
[0658] The processing flow will be explained below.
[0659] Step 1:
[0660] Users take a photo of the product shelves in a store using a device such as a smartphone or dedicated camera, and by pressing the capture button, the image is automatically saved on the device.
[0661] Step 2:
[0662] The device automatically sends the captured image data to a server via Wi-Fi or mobile networks.
[0663] Step 3:
[0664] The server analyzes the image data received from the device and uses an image analysis engine to identify products in the image and determine the location and quantity of each product.
[0665] Step 4:
[0666] The server stores the identified product information (SKU, location, stock quantity, etc.) in a database, allowing the current stock status to be immediately grasped.
[0667] Step 5:
[0668] The server references sales data based on the information stored in the database, extracts past sales data using SQL queries, and analyzes the relationship between product placement and sales.
[0669] Step 6:
[0670] The server uses a generative AI model to analyze product placement and sales data, thereby identifying optimal product placement patterns.
[0671] Step 7:
[0672] The server generates optimal product placement proposals based on the analysis results, which are then visualized.
[0673] Step 8:
[0674] The server sends the generated placement proposal to the terminal, and the user can check the proposal displayed on the terminal.
[0675] Step 9:
[0676] The device will read the notification content aloud to the user using text-to-speech (TTS) technology.
[0677] Step 10:
[0678] The device recognizes the user's emotional state using an emotion engine that analyzes the user's voice and facial expressions in real time.
[0679] Step 11:
[0680] Based on the results of the emotion engine, the server dynamically adjusts the notification content, for example, if the user is feeling stressed, it will provide concise, positive suggestions.
[0681] Step 12:
[0682] The user can correct the contents of the voice message by entering a voice command, for example, "There are 15 bags in stock."
[0683] Step 13:
[0684] The server receives the user's voice command and updates the database with the corrected information, which is then recalculated and saved as the latest information.
[0685] Step 14:
[0686] The server generates management advice for ordering and stock disposal based on the results of periodic inventory. This advice is generated based on the amount of stock and sales trends.
[0687] Step 15:
[0688] The server then sends the generated management advice to the terminal, and the user makes decisions about ordering and inventory disposal based on this advice.
[0689] Through these steps, this system will improve the efficiency of product management and inventory management, increase sales, improve the accuracy of management decisions, and even realize flexible responses that take user emotions into consideration. As a specific example, when the stock of a specific product is low, the system will notify the user by saying, "Stock is low. Please consider placing an order," but if the user is feeling emotionally stressed, it will adjust the wording to a softer one, saying, "There is no need to take immediate action. Please check when you have time."
[0690] Example 2
[0691] Next, a description will be given of Example 2. In the following description, the data processing device 12 will be referred to as a "server" and the smart glasses 214 will be referred to as a "terminal."
[0692] In the retail industry, it is extremely important to accurately analyze the impact that product placement has on sales and propose optimal placements. However, currently, there are limited methods for optimizing shelf placement and proposal systems that take users' emotional states into account. This has prevented efficient inventory management and increased sales from being fully realized. Furthermore, when incorrect product information is identified, correcting it is cumbersome, requiring quick and appropriate management decisions. Given this background, a system is needed that can automatically analyze the relationship between product placement and sales and respond flexibly while taking emotions into consideration.
[0693] The specific processing by the specific processing unit 290 of the data processing device 12 in the second embodiment is realized by the following means.
[0694] In this invention, the server includes means for acquiring image data from product shelves in a store, means for analyzing the acquired image data to identify the product locations and quantities, means for storing the identified information in a database, means for using a generative model to analyze the relationship between product placement and sales by referencing sales data based on the information stored in the database, means for proposing optimal product placement based on the analysis results, means for visualizing the proposal and notifying the terminal, means for reading the notification content aloud, and means for recognizing the user's emotional state from the terminal and dynamically adjusting the notification content based on the user's emotional state. This enables accurate analysis of the relationship between product placement and sales and proposals for optimal placement that take emotions into consideration. It also makes it easier to correct erroneous judgments and generate management advice, enabling quick and accurate management decisions.
[0695] 1. "Image data" refers to digital information, such as still images or videos, that record the state of product shelves in a store.
[0696] 2. "Analysis" refers to the computational process used to identify the location and quantity of products from the acquired image data.
[0697] 3. "Database" means a system for systematically storing and managing specific product information and sales data.
[0698] 4. "Sales Data" means numerical information on the sales of products within a specific period of time.
[0699] 5. A "generative model" is a machine learning algorithm that learns from large amounts of data to derive trends and patterns.
[0700] 6. "Visualization" refers to displaying analysis results and placement proposals as graphs and charts so that users can intuitively understand them.
[0701] 7. "Terminal" means a computer device used by a user to operate the device, in particular a smartphone or tablet device.
[0702] 8. "Notification Content" means the output of information or suggestions provided by the System to the User in the form of text or audio.
[0703] 9. "Emotional state" is state information that represents the user's emotional response and is identified from the user's voice and facial expression.
[0704] 10. "Dynamic adjustment" means changing your response to the situation in real time.
[0705] This invention is a system that acquires image data of product shelves in a store, analyzes the image data to identify the location and quantity of products, and then analyzes sales data based on this to propose optimal product placement.The system uses image recognition technology, generative AI, and an emotion engine, and is able to flexibly respond by taking into account the user's emotional state.
[0706] Equipment and software configuration
[0707] The system consists of the following main components:
[0708] Terminal: A device equipped with a camera for taking pictures of product shelves in a store. Examples include smartphones and dedicated cameras.
[0709] Server: The central data processing unit that receives and analyzes image data. It runs the generative AI model and emotion engine.
[0710] User: The person who operates the system and makes optimal product placement and management decisions.
[0711] Processing steps
[0712] 1. Image data collection
[0713] A user uses a device to take an image of a product shelf in a store. For example, they can use a smartphone camera app to take a picture, focusing on a specific area. The captured image is set to be automatically sent to the server. A dedicated app encodes the image data and sends it to the server using a secure communication protocol (e.g., HTTPS).
[0714] 2. Analysis of image data
[0715] The server analyzes the image data received from the device. Specifically, it uses software such as OpenCV and TensorFlow to identify the location and quantity of products. For example, it uses the YOLO (You Only Look Once) object detection algorithm. The analysis results are stored in a database.
[0716] 3. Analysis of the relationship between placement and sales
[0717] The server references product information and sales data stored in a database and analyzes the relationship between product placement and sales performance using a generative AI model (e.g., H2O.ai's AutoML function). This analysis identifies the optimal placement pattern for a specific product.
[0718] 4. Proposal for optimal layout
[0719] Based on the analysis results, the server generates optimal product placement proposals, which are visualized and displayed in an intuitive format using Tableau or Power BI, and the proposals are sent to the device and notified to the user.
[0720] 5. Introducing the Emotion Engine
[0721] The device reads product information and stock availability to the user aloud, while simultaneously recognizing the user's emotional state using the Microsoft Azure Emotion API and IBM Watson Tone Analyzer. Based on the user's emotional state, the device dynamically adjusts the content of notifications.
[0722] 6. Correcting misjudgments and providing management advice
[0723] The user can make corrections to the information read out loud using voice commands. For example, they can input a correction such as "There are 15 bags in stock." The server reflects this correction in the database and performs recalculation. In addition, based on the results of periodic inventory, the server generates management advice for ordering and stock disposal and sends it to the terminal.
[0724] Specific examples
[0725] For example, if snacks are placed on the bottom shelf, analysis may reveal poor sales data. Based on this, the server suggests placing snacks at eye level. The suggestions are sent to the device as a concrete visual map, allowing the user to adjust product placement. At the same time, an emotion engine analyzes the user's emotional state, providing detailed suggestions if the user has a positive reaction and simple suggestions if the user has a negative reaction. For example, if a voice notification reads, "There are 10 bags of snacks in stock," the user can correct it with a voice command, saying, "There are 15 bags in stock," and the system will immediately update the information.
[0726] Prompt Sentence Examples
[0727] 1. "What image recognition technology do you use to optimize product placement in your stores?"
[0728] 2. "Please show the results of analyzing sales data using a generative AI model."
[0729] 3. "How does the Emotion Engine analyze the user's emotional state and adjust its response?"
[0730] This system will enable more efficient product and inventory management, increased sales, flexible responses that take user feelings into consideration, and more accurate management decisions.
[0731] The flow of the identification process in the second embodiment will be described with reference to FIG.
[0732] Step 1:
[0733] The user uses a device to take a picture of a specific product shelf in a store. The input is a high-resolution image taken using a smartphone camera app, for example. The output is image data with a clear focus on the captured product. Specifically, the image is taken so that the product's position and quantity are visible, and the image data is saved on the device.
[0734] Step 2:
[0735] The device sends the captured image data to the server. The input is the image data acquired in step 1. The output is the image data securely sent to the server. Specifically, a dedicated application encodes the image data into Base64 format and uploads it to the server in real time using the HTTPS protocol.
[0736] Step 3:
[0737] The server analyzes the image data received from the device. The input is the image data sent in step 2. The output is data on the location and quantity of products identified through image analysis. Specifically, it uses TensorFlow and OpenCV to apply object detection algorithms such as YOLO to extract the bounding boxes and labels of each product in the image.
[0738] Step 4:
[0739] The server stores the analysis results in a database. The input is the product location and quantity data identified in step 3. The output is product information stored in the database. Specifically, it is inserted into a MySQL database using an ORM (Object-Relational Mapping) library such as SQLAlchemy.
[0740] Step 5:
[0741] The server retrieves sales data from the database and uses a generative model to analyze the relationship between product placement and sales. The input is the product information and sales data stored in the database. The output is the optimal product placement pattern. Specifically, it uses H2O.ai's AutoML function to derive placement patterns that are expected to increase sales.
[0742] Step 6:
[0743] The server generates optimal product placement proposals based on the analysis results. The input is the optimal product placement pattern obtained in step 5. The output is a visualized placement proposal. Specifically, it is visualized as intuitive graphs and charts using Tableau or Power BI.
[0744] Step 7:
[0745] The server sends the visualized proposal to the terminal and notifies the user. The input is the placement proposal visualized in step 6. The output is the proposal notified to the user. Specifically, it is visually displayed on the terminal via the Internet.
[0746] Step 8:
[0747] The device reads out product information and stock status aloud, while simultaneously recognizing the user's emotional state. The input is the recommendations sent from the server and the user's real-time voice data. The output is the results of recognizing the user's emotional state and the adapted notification content. Specifically, it uses the Microsoft Azure Emotion API and IBM Watson Tone Analyzer to analyze the user's voice and facial expressions.
[0748] Step 9:
[0749] The server dynamically adjusts the notification content based on the emotional state identified by the emotion engine. The input is the user's emotional state sent from the device. The output is the adjusted notification content. For example, if the user is feeling stressed, more concise and positive suggestions are provided.
[0750] Step 10:
[0751] The user inputs a voice command in response to the voice readout, and the server reflects the correction in the database. The input is the user's voice command and the current contents of the database. The output is updated information in the database that reflects the correction. For example, if the correction is made to "There are 15 bags in stock," the information is immediately updated in the database.
[0752] As a result, this system analyzes the relationship between product placement in a store and sales, and realizes flexible and adaptable placement suggestions and inventory management that take into account the user's emotional state.
[0753] (Application example 2)
[0754] Next, a description will be given of Application Example 2. In the following description, the data processing device 12 will be referred to as a "server" and the smart glasses 214 will be referred to as a "terminal."
[0755] Conventional product placement support systems for the retail industry require a lot of effort to identify product locations and inventory information, and their placement change suggestions do not take into account the user's emotional state, making it difficult to maximize management efficiency and sales. Furthermore, because real-time product placement changes and inventory adjustments are time-consuming, there is a demand for faster management decision-making.
[0756] The identification process by the identification processing unit 290 of the data processing device 12 in Application Example 2 is realized by the following means. In this invention, the server includes means for acquiring image data from product shelves in a store, means for analyzing the acquired image data and identifying the location and quantity of products, means for saving the identified information in a database, means for referencing sales data based on the information saved in the database and analyzing the relationship between product placement and sales, means for proposing an optimal product placement based on the analysis results, means for visualizing the proposal and notifying the user, means for reading the notification content aloud to the user, means for a user to use a smart device to take a real-time photo of the product shelf situation in the store and send the image data to the server, means for analyzing sales data and product information using a generative AI model, means for recognizing the user's emotional state and adjusting the notification content based on the user's emotion, and means for providing real-time information updates and advice in cooperation with a voice assistant. This enables efficient product management and inventory management, maximizing sales, flexible responses based on user emotions, and faster management decisions.
[0757] "Image data" is image information obtained by photographing a product shelf, and is data used to identify the position and quantity of products.
[0758] "In-store shelves" refers to shelves and racks used to display, showcase, and store products in retail stores and brick-and-mortar stores.
[0759] "Means" means a method, technique, or device for achieving a particular purpose.
[0760] "Analysis" refers to the process of analyzing the acquired image data and extracting detailed information such as the location and quantity of the products.
[0761] A "database" is an information system that can efficiently store, manage, and search large amounts of data.
[0762] "Sales data" refers to data containing sales information for a product within a certain period of time, including sales amount, sales quantity, sales date, and the like.
[0763] A "generative AI model" is an algorithm that uses artificial intelligence to analyze data and generate new insights and predictions.
[0764] "Visualization" is the process of presenting data or information in a visually understandable way.
[0765] "User" refers to the person who operates the system and manages product placement and inventory.
[0766] "Notification" is the act of transmitting information from the system to the user.
[0767] A "voice assistant" is a system that interacts with users using voice input to assist them in obtaining information and performing operations.
[0768] "Smart devices" refer to electronic devices equipped with advanced computing power and communication functions, including smartphones and smart glasses.
[0769] "Emotional state" refers to a user's emotional response or state, including happiness, sadness, stress, etc.
[0770] "Real-time" means processing or reacting nearly simultaneously or without delay.
[0771] System Configuration
[0772] This system acquires images of in-store product shelves and optimizes product placement using generative AI models and emotion engines. The system consists of the following main hardware and software components:
[0773] 1. Terminal
[0774] Smart devices include smart glasses and smartphones, which are equipped with cameras and are used to capture images of the product shelves in stores.
[0775] 2. Server
[0776] OpenCV is used for image analysis, which analyzes image data and identifies the location and quantity of products.
[0777] The database used is MySQL, and the analyzed data is saved in the database.
[0778] The generative AI model uses OpenAI's GPT-4 and other models to analyze sales data and product information and generate optimal placement patterns.
[0779] The emotion engine uses Affectiva's SDK to analyze the user's emotional state.
[0780] As a voice assistant, it uses Google Cloud Text-to-Speech to read out the generated notification content aloud.
[0781] 3. Users
[0782] Managers use smart devices to take photos of the shelves and send the data to a server in real time.
[0783] Processing steps
[0784] Image data collection
[0785] The user uses smart glasses or a smartphone to take an image of the store's shelves, focusing on a specific area of the shelves to clearly show the location and quantity of products. The device is set to automatically send the captured image to the server.
[0786] Image data analysis
[0787] The server analyzes the image data received from the device. The image analysis engine uses OpenCV to identify the location, quantity, and specific product name of each product. After the product information is identified, it is saved in a database.
[0788] Saving to a database
[0789] The identified product information is stored in a database along with attributes such as SKU, location, and inventory quantity, making it possible to immediately grasp the current inventory status.
[0790] Analysis of the relationship between placement and sales
[0791] The server references sales data from a database for a set period and uses a generative AI model to analyze the relationship between product placement and sales performance. This analysis identifies the optimal placement pattern for specific products, maximizing sales.
[0792] Proposal for optimal layout
[0793] Based on the analysis results, the server generates an optimal product placement proposal. This proposal is visualized, allowing the user to compare the current product shelf layout with the proposed optimal placement. The proposal is then sent to the terminal and notified to the user.
[0794] Introducing the Emotion Engine
[0795] The device reads product information and inventory status to the user aloud and uses an emotion engine to recognize the user's emotional state by analyzing the user's voice and facial expressions to identify their emotions.
[0796] Adjusting notification content
[0797] Based on the emotional state identified by the emotion engine, the server dynamically adjusts the notification content, for example, providing more concise and positive suggestions if the user is feeling stressed.
[0798] Correcting misjudgments and providing management advice
[0799] The user can make corrections to the voice-read information by voice command. For example, the user can input a correction such as "15 bags in stock." The server reflects this correction in the database and performs recalculation. In addition, based on the results of periodic inventory, the server generates management advice for ordering and stock disposal and sends it to the terminal. The user can make effective management decisions based on this advice.
[0800] Specific examples
[0801] For example, analysis may reveal that sales data for a particular product category (e.g., snacks) are sluggish when placed on the bottom shelf. Based on this, the server suggests placing snacks in front of the register or at eye level. The suggestions are sent to the device as a concrete visual map, and the user adjusts product placement accordingly. The emotion engine also analyzes the user's emotional state, providing detailed suggestions if the user has a positive reaction and simple suggestions if the user has a negative reaction.
[0802] Prompt Sentence Examples
[0803] "Please suggest the optimal product placement in the store. The current placement and sales data are as follows: \[Specific data\]"
[0804] The flow of the specific processing in the application example 2 will be described with reference to FIG.
[0805] Step 1:
[0806] A user uses a smart device (smart glasses or smartphone) to take a picture of the product shelf situation in a store in real time. The captured image data is automatically sent from the device to a server. The input is the image of the product shelf, and the output is the image data sent to the server.
[0807] Step 2:
[0808] The server analyzes the received image data. At this time, an image analysis engine (such as OpenCV) is used to identify the location and quantity of each product. Specifically, object recognition and extraction are performed on the image data. The input is the received image data, and the output is specific information such as the product's location, quantity, and product name.
[0809] Step 3:
[0810] The server stores the identified product information in a database (e.g., MySQL). Here, attributes such as SKU, location, and stock quantity are stored as data. The input is analysis information such as product location and quantity, and the output is product information stored in the database.
[0811] Step 4:
[0812] The server references the information stored in the database and sales data, and uses a generative AI model to analyze the relationship between product placement and sales data. At this time, the generative AI model (such as OpenAI's GPT-4) receives a prompt message: "Please suggest the optimal placement of products in the store. The current placement data and sales data are as follows: \[Specific data\]". The input is the current placement data and sales data, and the output is a proposal for the optimal product placement.
[0813] Step 5:
[0814] The server visualizes the generated optimal placement proposal and sends the results to the terminal. Specifically, a visual map is created. The input is the placement proposal obtained from the generative AI model, and the output is the visualized placement proposal and a notification of it.
[0815] Step 6:
[0816] The device reads out product information and stock availability to the user and analyzes the user's emotional state using an emotion engine (such as Affectiva's SDK). The emotion engine recognizes emotions and sends the results to the server. The input is the user's voice and facial expression, and the output is the identified emotional state.
[0817] Step 7:
[0818] The server dynamically adjusts the notification content based on the emotional state obtained from the emotion engine. If the user is feeling stressed, it provides more concise and positive suggestions, and if the user has a positive reaction, it provides more detailed suggestions. The input is the emotional state obtained from the emotion engine, and the output is the adjusted notification content.
[0819] Step 8:
[0820] The user can make corrections to the voice readout as necessary using voice commands. For example, they can input the corrections, such as "There are 15 bags in stock." The server reflects this correction information in the database and performs recalculation. The input is the user's voice command, and the output is the corrected database information.
[0821] Step 9:
[0822] The server generates management advice for ordering and stock disposal based on the results of periodic inventory and sends it to the terminal. The user makes effective management decisions based on this management advice. The input is the inventory result data, and the output is the generated management advice.
[0823] The specific processing unit 290 transmits the result of the specific processing to the smart glasses 214. In the smart glasses 214, the control unit 46A causes the speaker 240 to output the result of the specific processing. The microphone 238 acquires audio indicating a user input regarding the result of the specific processing. The control unit 46A transmits audio data indicating the user input acquired by the microphone 238 to the data processing device 12. In the data processing device 12, the specific processing unit 290 acquires the audio data.
[0824] The data generation model 58 is a so-called generative AI (Artificial Intelligence). An example of the data generation model 58 is ChatGPT (Internet Search<URL: https: / / openai.com / blog / chatgpt> ), Gemini (Internet search <url: https: gemini.google.com ?hl="ja">) and other generation AIs. The data generation model 58 is obtained by performing deep learning on a neural network. A prompt including an instruction is input to the data generation model 58, and inference data such as voice data indicating voice, text data indicating text, and image data indicating an image is also input. The data generation model 58 performs inference on the input inference data in accordance with the instruction indicated by the prompt, and outputs the inference result in a data format such as voice data and text data. Here, inference refers to, for example, analysis, classification, prediction, and / or summarization.
[0825] In the above embodiment, an example in which the specific processing is performed by the data processing device 12 has been given, but the technology of the present disclosure is not limited to this, and the specific processing may be performed by the smart glasses 214.
[0826] [Third embodiment]
[0827] FIG. 5 shows an example of the configuration of a data processing system 310 according to the third embodiment.
[0828] 5, the data processing system 310 includes the data processing device 12 and a headset terminal 314. An example of the data processing device 12 is a server.
[0829] The data processing device 12 includes a computer 22, a database 24, and a communication I / F 26. The computer 22 is an example of a "computer" according to the technology of the present disclosure. The computer 22 includes a processor 28, a RAM 30, and a storage 32. The processor 28, the RAM 30, and the storage 32 are connected to a bus 34. The database 24 and the communication I / F 26 are also connected to the bus 34. The communication I / F 26 is connected to a network 54. Examples of the network 54 include a WAN (Wide Area Network) and / or a LAN (Local Area Network).
[0830] The headset type terminal 314 includes a computer 36, a microphone 238, a speaker 240, a camera 42, a communication I / F 44, and a display 343. The computer 36 includes a processor 46, a RAM 48, and a storage 50. The processor 46, the RAM 48, and the storage 50 are connected to a bus 52. The microphone 238, the speaker 240, the camera 42, and the display 343 are also connected to the bus 52.
[0831] The microphone 238 receives instructions and the like from the user 20 by receiving voice uttered by the user 20. The microphone 238 captures the voice uttered by the user 20, converts the captured voice into audio data, and outputs it to the processor 46. The speaker 240 outputs audio in accordance with instructions from the processor 46.
[0832] Camera 42 is a small digital camera equipped with an optical system including a lens, aperture, and shutter, and an imaging element such as a CMOS (Complementary Metal-Oxide-Semiconductor) image sensor or a CCD (Charge Coupled Device) image sensor, and captures images of the surroundings of user 20 (for example, an imaging range defined by an angle of view equivalent to the field of vision of a typical healthy person).
[0833] The communication I / F 44 is connected to a network 54. The communication I / Fs 44 and 26 are responsible for the exchange of various information between the processor 46 and the processor 28 via the network 54. The exchange of various information between the processor 46 and the processor 28 using the communication I / Fs 44 and 26 is carried out in a secure state.
[0834] Fig. 6 shows an example of the main functions of the data processing device 12 and the headset type terminal 314. As shown in Fig. 6, in the data processing device 12, a specific process is performed by the processor 28. A specific process program 56 is stored in the storage 32.
[0835] The specific processing program 56 is an example of a "program" according to the technology of the present disclosure. The processor 28 reads the specific processing program 56 from the storage 32 and executes the read specific processing program 56 on the RAM 30. The specific processing is realized by the processor 28 operating as a specific processing unit 290 in accordance with the specific processing program 56 executed on the RAM 30.
[0836] The storage 32 stores a data generation model 58 and an emotion identification model 59. The data generation model 58 and the emotion identification model 59 are used by the identification processing unit 290.
[0837] In the headset type terminal 314, a reception output process is performed by the processor 46. A reception output program 60 is stored in the storage 50. The processor 46 reads the reception output program 60 from the storage 50 and executes the read reception output program 60 on the RAM 48. The reception output process is realized by the processor 46 operating as the control unit 46A in accordance with the reception output program 60 executed on the RAM 48.
[0838] Next, a description will be given of the identification process performed by the identification processing unit 290 of the data processing device 12. In the following description, the data processing device 12 will be referred to as the "server" and the headset type terminal 314 will be referred to as the "terminal."
[0839] The present invention relates to an optimal product placement support system for the retail industry that uses image recognition technology and generative AI, and aims to analyze the impact of product placement in a store on sales and optimize product placement. Specific implementation methods of the present invention are described below.
[0840] System configuration
[0841] The system consists of the following main components:
[0842] 1. Terminal: Equipped with a camera to take pictures of product shelves in a store and has the function of sending image data to a server. Examples include smartphones and dedicated cameras.
[0843] 2. Server: The central data processing unit that receives and analyzes image data, running the generative AI model.
[0844] 3. User: The person who operates the system and makes management decisions such as optimal product placement, ordering, and inventory disposal.
[0845] Program processing flow
[0846] Image data collection
[0847] The user uses the device to take an image of the store's shelves, focusing on a specific area of the shelves to clearly show the location and quantity of products. The device is configured to automatically send the captured image to the server.
[0848] Image data analysis
[0849] The server analyzes the image data received from the device. Using an image analysis engine, it identifies the location, quantity, and specific product name of each product. After the product information is identified, it is stored in a database.
[0850] Saving to a database
[0851] The identified product information is stored in a database along with attributes such as SKU, location, and inventory quantity, making it possible to immediately grasp the current inventory status.
[0852] Analysis of the relationship between placement and sales
[0853] The server references sales data from a database for a set period and uses a generative AI model to analyze the relationship between product placement and sales performance. This analysis identifies the optimal placement pattern for specific products, maximizing sales.
[0854] Proposal for optimal layout
[0855] Based on the analysis results, the server proposes the optimal product placement. This proposal is visualized, allowing users to compare the current product shelf layout with the proposed optimal placement. The proposal is then sent to the terminal and notified to the user.
[0856] Voice reading and correction
[0857] Users can confirm the notification content from the device by voice. The device will read out product information and stock status using text-to-speech (TTS) technology. If there is an incorrect judgment, users can correct it using voice commands. This correction is reflected in the database again and recalculation is performed.
[0858] Management Advice
[0859] Based on the results of periodic inventory, the server generates management advice such as ordering and discount proposals for inventory clearance. These advices are notified to the user via text and voice, supporting effective management decisions.
[0860] Specific examples
[0861] For example, analysis may reveal that sales data for a particular product category (e.g., snacks) are stagnating if they are placed on the bottom shelf. Based on this, the server may suggest placing snacks in front of the register or at eye level. The suggestion is sent to the device as a concrete visual map, and the user can adjust product placement based on the suggestion. Also, if a voice notification reads, "There are 10 bags of snacks in stock," the user can correct it with a voice command, saying, "There are 15 bags in stock," and the system will immediately update the information.
[0862] These features provide a system that streamlines product and inventory management, increases sales, and enables more accurate business decisions.
[0863] The processing flow will be explained below.
[0864] Step 1:
[0865] Users take a photo of the product shelves in a store using a device such as a smartphone or dedicated camera, and by pressing the capture button, the image is automatically saved on the device.
[0866] Step 2:
[0867] The device automatically sends the captured image data to a server via Wi-Fi or mobile networks.
[0868] Step 3:
[0869] The server analyzes the image data received from the device and uses an image analysis engine to identify products in the image and determine the location and number of each product.
[0870] Step 4:
[0871] The server stores the identified product information in a database, including the SKU, location, and quantity in stock.
[0872] Step 5:
[0873] The server references sales data based on the information stored in the database, extracts past sales data using SQL queries, and analyzes the relationship between product placement and sales.
[0874] Step 6:
[0875] The server uses a generative AI model to analyze product placement and sales data, thereby identifying optimal product placement patterns.
[0876] Step 7:
[0877] The server generates optimal product placement proposals based on the analysis results, which are then visualized.
[0878] Step 8:
[0879] The server sends the generated placement proposal to the terminal, and the user can check the proposal displayed on the terminal.
[0880] Step 9:
[0881] The device reads product information and stock availability to the user aloud, and uses text-to-speech (TTS) technology to audibly output the contents of the visual suggestions.
[0882] Step 10:
[0883] The user can make corrections using voice commands in response to the content of the voice readout. For example, the user can input a correction such as "There are 15 bags in stock."
[0884] Step 11:
[0885] The server receives voice correction commands from the user and reflects them in the database, where the corrected information is recalculated and saved as the latest information.
[0886] Step 12:
[0887] The server generates management advice for ordering and stock disposal based on the results of periodic inventory. This advice is generated based on the amount of stock and sales trends.
[0888] Step 13:
[0889] The server then sends the generated management advice to the terminal, and the user makes decisions about ordering and inventory disposal based on this advice.
[0890] Through the above steps, this system will improve the efficiency of product management and inventory management, thereby increasing sales and improving the accuracy of management decisions.
[0891] Example 1
[0892] Next, a description will be given of Example 1. In the following description, the data processing device 12 will be referred to as a "server" and the headset type terminal 314 will be referred to as a "terminal."
[0893] The retail industry is seeking to optimize product placement and streamline inventory management. Conventional methods rely on manual work to determine product locations and quantities, which takes time and effort and is prone to errors. It is also difficult to analyze the relationship between product placement and sales, requiring significant effort to find optimal product placement. Furthermore, updating inventory information in real time is difficult, making it difficult to make quick management decisions. To solve these issues, a system is needed that can efficiently and accurately manage product placement and inventory.
[0894] The specific processing by the specific processing unit 290 of the data processing device 12 in the first embodiment is realized by the following means.
[0895] In this invention, the server includes a means for analyzing image data and identifying product locations and quantities, a means for proposing optimal product placement using a generative AI model, and a means for referencing sales data based on information stored in a database and analyzing the relationship between product placement and sales. This makes it possible to identify product locations and quantities with high accuracy, update inventory information in real time, and quickly propose optimal product placement to maximize sales.
[0896] "Image data" refers to visual information acquired in the form of photographs or videos of product shelves in a store.
[0897] "Acquisition" refers to the act of obtaining image data of an object using a camera or other photographic equipment.
[0898] "Analysis" is the process of extracting and identifying specific information from acquired image data.
[0899] A "generative AI model" is an artificial intelligence algorithm trained to perform image recognition and data analysis.
[0900] "Location" is information indicating the specific location where a particular product is located.
[0901] "Quantity" is a number that indicates how much of a particular product is on the shelf or in stock.
[0902] A "database" is an electronic recording device or system for systematically storing and managing specific information.
[0903] "Sales data" refers to transaction information when a product is sold, including sales amount and sales quantity.
[0904] "Analysis" is the act of examining and interpreting data to find specific relationships and patterns.
[0905] "Suggestion" refers to showing optimal behavior and placement patterns based on the analysis results.
[0906] "Visualization" refers to displaying data and proposals in a visually easy-to-understand manner.
[0907] "Notification" is the act of informing a user of specific information.
[0908] "Text-to-speech" is a technology that converts text information into speech and conveys it to the user auditorily.
[0909] "Voice command" is a method by which a user gives instructions or inputs using their voice.
[0910] "Inventory" refers to the periodic checking and recording of stock and the quantity of merchandise items.
[0911] "Management advice" refers to proposals based on inventory and sales data to support decision-making such as ordering and inventory disposal.
[0912] This invention is an optimal product placement support system for the retail industry that uses image recognition technology and generative AI models. The system acquires image data of product shelves in stores, analyzes the data to identify product locations and quantities, stores the data in a database, and then compares it with sales data to suggest optimal product placement.
[0913] The system consists of three main components:
[0914] 1. Terminal: A device with a camera function that takes pictures of product shelves in a store. Typical examples are smartphones or dedicated cameras.
[0915] 2. Server: A central data processing unit that receives data, performs image analysis, and optimizes product placement using generative AI models.
[0916] 3. User: The person who operates the system and makes management decisions such as optimal product placement, ordering, and inventory disposal.
[0917] Hardware and Software Configuration
[0918] Terminal
[0919] Users take pictures of product shelves in stores using smartphones or dedicated cameras. These devices have the ability to automatically send the captured image data to a server. The captured images are sent to the server via Wi-Fi or mobile data. The devices are also equipped with text-to-speech (TTS) technology, which allows them to notify users by voice.
[0920] server
[0921] The server receives the image data sent from the device and uses an image analysis engine to identify the location, quantity, and product name of each product. This analysis uses deep learning technology and generative AI models, such as machine learning libraries like TensorFlow and PyTorch. The server then stores the identified information in a database and uses that data to analyze the relationship between sales data and product placement.
[0922] Database
[0923] The product information analyzed by the server is stored in a database, which updates attributes such as SKU (Stock Keeping Unit), location, and inventory quantity in real time. Users can easily check this information.
[0924] Examples of use and specific prompts
[0925] Users take pictures of product shelves in a store and send them to a server using their device. The server then receives the image data, analyzes it, identifies the product locations and quantities, and stores them in a database. Based on the analysis results, the server uses a generative AI model to propose optimal product placement and sends it to the device as a visual map. The user can then confirm the proposal visually and audibly and adjust the product placement.
[0926] Prompt Sentence Examples
[0927] "Please identify the names and quantities of the items shown in this image."
[0928] "Please suggest optimal product placement based on sales data and product placement data."
[0929] "Please read out the inventory."
[0930] · "Use voice commands to correct stock levels."
[0931] This will lead to more efficient product and inventory management, increased sales, and more accurate business decisions.
[0932] The flow of the identification process in the first embodiment will be described with reference to FIG.
[0933] Step 1:
[0934] The user uses a device to take a photo of a product shelf in a store. The captured image is high resolution, and the product name and quantity are clearly visible. For example, a smartphone or dedicated camera is used to capture product images under appropriate lighting. The input is an "image of a product shelf in a store," and the output is "high-resolution image data."
[0935] Step 2:
[0936] The device sends the captured image data to the server. The device has a network connection function and encrypts and sends the image data to the server using Wi-Fi or mobile data. The input is "high-resolution image data" and the output is "a notification of completion of transmission to the server."
[0937] Step 3:
[0938] The server analyzes the image data received from the device. The image analysis engine processes the input image and identifies the location, quantity, and name of the product. A generative AI model is used for this analysis. Specifically, machine learning libraries such as TensorFlow and PyTorch are used. The input is "high-resolution image data," and the output is "the location, quantity, and name of the identified product."
[0939] Step 4:
[0940] The server stores the identified information in a database. The identified product information (SKU, location, and stock quantity) is recorded in the database, and inventory information is updated in real time. The input is the "identified product location, quantity, and name," and the output is the "updated database entry."
[0941] Step 5:
[0942] The server analyzes the relationship between sales data and product placement performance based on the stored information. It uses a generative AI model to find the optimal placement pattern. For example, it compares the placement location of a specific product with sales data to identify the placement that maximizes sales. The inputs are "product information stored in the database" and "sales data," and the output is "identification of the optimal placement pattern."
[0943] Step 6:
[0944] The server proposes optimal product placement and generates a visualized proposal. It provides the user with a layout diagram in an easy-to-understand visual format, allowing them to compare the proposed placements. The proposal is sent to the device. The input is the "identification result of the optimal placement pattern," and the output is the "visualized placement proposal."
[0945] Step 7:
[0946] The terminal notifies the user of the proposal content sent from the server by voice. Using text-to-speech (TTS) technology, the placement proposal content and inventory information are read aloud. The input is a "visualized placement proposal" and the output is a "voice notification."
[0947] Step 8:
[0948] When a user wants to correct a voice notification, they input a voice command. For example, they can correct an inventory quantity by voice, and the correction is sent to the server. The server analyzes the voice command and updates the database. The input is the "user's voice command" and the output is the "updated database entry."
[0949] Step 9:
[0950] Based on the results of periodic inventory, the server generates management advice for ordering and stock disposal. It then notifies the user of this advice via text and voice. For example, the advice might be, "There is an excess of snack food in stock. We suggest selling them at a discount." The inputs are "inventory results" and "database information," and the output is "notification of management advice."
[0951] (Application example 1)
[0952] Next, a description will be given of Application Example 1. In the following description, the data processing device 12 will be referred to as a "server" and the headset type terminal 314 will be referred to as a "terminal."
[0953] It is widely recognized in the retail industry that product placement in a store has a significant impact on sales. However, optimizing shelf placement using traditional methods requires a great deal of effort and time, and effective placement is not always achieved. Furthermore, real-time inventory management and rapid management decisions based on sales data are required, but traditional systems do not adequately support this. Furthermore, store staff are required to manually check and correct product locations and inventory status, which is time-consuming and inefficient.
[0954] The specific processing by the specific processing unit 290 of the data processing device 12 in the application example 1 is realized by the following means.
[0955] In this invention, the server includes means for acquiring image data from product shelves in a store, means for analyzing the acquired image data and identifying product locations and quantities, means for saving the identified information in a database, means for referencing sales data based on the information saved in the database and analyzing the relationship between product placement and sales, means for proposing optimal product placement based on the analysis results, means for visualizing the proposal and notifying the user, means for reading the notification aloud to the user, means for taking pictures of the product shelves using a smartphone, means for providing audio notification of product information and inventory status using text-to-speech technology, means for the user to correct information using voice commands, means for generating optimal placements using a generative AI model that contributes to maximizing sales, and means for inputting prompts to the generative AI model to obtain analysis results. This enables efficient acquisition of product shelf images using a smartphone, optimal placement suggestions using generative AI, and rapid and accurate inventory management and management decisions using voice technology.
[0956] "Image data" refers to information that is acquired as an image of product shelves in a store and is the subject of analysis.
[0957] "Analysis" refers to the act of identifying the location and quantity of products from the acquired image data.
[0958] A "database" is a repository of information that stores identified product information and allows for later reference.
[0959] "Sales data" is information relating to product sales, and is used to analyze the relationship between product placement and sales.
[0960] A "generative AI model" is an artificial intelligence technology used to analyze image data and propose optimal product placement.
[0961] A "prompt statement" is an instruction statement that is input into a generative AI model to obtain analysis results.
[0962] "Visualization" is the act of visually suggesting optimal product placement.
[0963] "Voice notification" refers to the act of using text-to-speech technology to inform users of product information and stock status via voice.
[0964] A "smartphone" is a mobile device used to acquire and send notifications about image data.
[0965] A "voice command" is a spoken input by a user to give instructions to a system using voice recognition technology.
[0966] "Inventory management" is the management activity of determining the quantity of goods and maintaining appropriate stock levels.
[0967] "Management decisions" are decisions made regarding product placement and inventory management.
[0968] MODE FOR CARRYING OUT THE INVENTION
[0969] The present invention relates to a system for optimizing product placement in a retail store using a smartphone. A specific implementation method of this system will be described below.
[0970] System configuration
[0971] The system consists of the following major components:
[0972] 1. Device (smartphone):
[0973] It is equipped with a camera to take pictures of product shelves in the store and transmits the image data to a server.
[0974] It has the ability to provide voice notification of product information and stock status using text-to-speech technology.
[0975] It uses voice recognition technology to receive the user's voice commands and send correction information to the server.
[0976] 2. Server:
[0977] It receives image data and runs a generative AI model that performs the analysis.
[0978] Using an image analysis engine (e.g., YOLOv5, TensorFlow), the location, quantity, and specific product name of each product are identified.
[0979] The identified product information is stored in a database (e.g., Firebase, PostgreSQL).
[0980] Sales data is referenced based on the information stored in the database, and the relationship between product placement and sales is analyzed.
[0981] The system generates the optimal layout that contributes to maximizing sales and notifies the user as a visual map.
[0982] 3. User:
[0983] A smartphone is used to take a photo of the product shelf and send the image data to the server.
[0984] Visually confirm the proposed optimal placement and confirm / correct inventory information notified by voice.
[0985] Make business decisions and place orders or dispose of inventory based on advice provided by the system.
[0986] Program processing
[0987] (Collection of image data)
[0988] Users take pictures of product shelves in a store using the camera on their smartphone, and the smartphone automatically sends the captured image data to the server.
[0989] (Image data analysis)
[0990] The server is equipped with an image analysis engine and generative AI models, such as YOLOv5 and TensorFlow, to analyze the received image data. The analysis identifies the location, quantity, and SKU of each product, and stores this information in a database.
[0991] (Saving to database)
[0992] The identified product information is stored in a database along with attributes such as SKU, location, and stock quantity, making it possible to instantly grasp the current inventory status.
[0993] (Analysis of the relationship between placement and sales)
[0994] The server references sales data for a set period from the database and uses a generative AI model to analyze the relationship between product placement and sales performance. At this time, it inputs prompt statements into the generative AI to obtain the analysis results.
[0995] (Optimal layout proposal)
[0996] A placement pattern that contributes to maximizing sales is generated, and the proposed details are sent to a smartphone as a visual map.
[0997] (audio notification and fix)
[0998] The smartphone uses text-to-speech technology (e.g., Google Cloud Text-to-Speech API) to notify the user of product information and stock availability via voice. When the user provides corrections via voice commands, the smartphone sends the information to the server, which updates the database.
[0999] (Providing management advice)
[1000] Based on the results of periodic inventory, the server generates management advice for ordering and stock disposal, which is communicated to the user via text and voice to support effective management decisions.
[1001] Specific examples
[1002] For example, if a particular snack is placed on the bottom shelf, analysis may reveal that sales data is sluggish. The server then suggests placing the snack in the center of the shelf where it is more visible. This suggestion is sent to the smartphone as a visual map, and the user follows the instructions to rearrange the products. Also, if a voice notification says, "There are 10 bags of snacks in stock," the user can correct it by saying, "There are 15 bags in stock," and the system will immediately update the database.
[1003] (Example of a prompt to input to a generative AI model)
[1004] input:
[1005] "Analyze the shelf image below to identify product locations, SKUs, and stock quantities. SKUs: '12345', '54321', '67890', etc."
[1006] output:
[1007] "Product locations and quantities identified: 4 units of SKU: '12345', 6 units of SKU: '54321', and 10 units of SKU: '67890'. Suggested placement: Place snacks in the center of the shelf."
[1008] The flow of the specific processing in the application example 1 will be described with reference to FIG.
[1009] Program processing steps
[1010] Step 1:
[1011] A user uses the camera on their smartphone to take a picture of a product shelf in a store. The image data taken by the smartphone is the input. The device sends this image data to the server. The output is the image data sent to the server.
[1012] Step 2:
[1013] The server analyzes the received image data. Specifically, it uses an image analysis engine (e.g., YOLOv5, TensorFlow) to identify the product location and quantity. The image data is input, and the product location, quantity, and SKU information are output.
[1014] Step 3:
[1015] The server stores the identified product information in a database. The input is the location, quantity, and SKU information of the identified product, and the product information stored in the database is obtained as the output.
[1016] Step 4:
[1017] The server references sales data based on the information stored in the database and analyzes the relationship between product placement and sales. This analysis is performed using a generative AI model (e.g., TensorFlow). Database information and sales data are input, and the analysis results include information on the relationship between product placement and sales.
[1018] Step 5:
[1019] The server inputs a prompt into the generative AI model to obtain the optimal product placement that will maximize sales. The inputs are the prompt and information from the database, and the output is a proposal for the optimal placement.
[1020] Step 6:
[1021] The server generates a visual map of the optimal layout proposal and sends it to the smartphone. The input is the optimal layout proposal, and the output is the visual map data.
[1022] Step 7:
[1023] The smartphone uses text-to-speech technology (e.g., Google Cloud Text-to-Speech API) to provide the user with product information and stock status via voice. The input is a visual map and product information, and the output is a voice notification.
[1024] Step 8:
[1025] The user listens to the voice notification and, if necessary, issues a voice command to correct the information. The smartphone receives this voice command and converts it into text using voice recognition technology (e.g., Google Speech-to-Text API). The voice command is used as input, and the text of the correction is output.
[1026] Step 9:
[1027] The server receives the user's corrections and updates the database: the text of the corrections is the input, and the database update is the output.
[1028] Step 10:
[1029] Based on the results of periodic inventory, the server generates management advice for ordering and stock disposal and notifies the user. Inventory data is input and management advice is output.
[1030] Combining these steps will improve the efficiency of in-store product placement and inventory management, and will also enable faster management decisions to maximize sales.
[1031] Furthermore, an emotion engine that estimates the user's emotion may be further combined. That is, the identification processing unit 290 may estimate the user's emotion using the emotion identification model 59, and perform identification processing using the user's emotion.
[1032] The present invention relates to an optimal product placement support system for the retail industry that uses image recognition technology, generative AI, and an emotion engine. The purpose of the system is to analyze the impact of in-store product placement on sales, optimize product placement, and also take into account the emotional state of users. Specific implementation methods of the present invention are described below.
[1033] System configuration
[1034] The system consists of the following main components:
[1035] 1. Terminal: Equipped with a camera to take pictures of product shelves in a store and has the function of sending image data to a server. Examples include smartphones and dedicated cameras.
[1036] 2. Server: The central data processing unit that receives and analyzes image data, running the generative AI model and emotion engine.
[1037] 3. User: The person who operates the system and makes management decisions such as optimal product placement, ordering, and inventory disposal.
[1038] Program processing flow
[1039] Image data collection
[1040] The user uses the device to take an image of the store's shelves, focusing on a specific area of the shelves to clearly show the location and quantity of products. The device is configured to automatically send the captured image to the server.
[1041] Image data analysis
[1042] The server analyzes the image data received from the device. Using an image analysis engine, it identifies the location, quantity, and specific product name of each product. After the product information is identified, it is stored in a database.
[1043] Saving to a database
[1044] The identified product information is stored in a database along with attributes such as SKU, location, and inventory quantity, making it possible to immediately grasp the current inventory status.
[1045] Analysis of the relationship between placement and sales
[1046] The server references sales data from a database for a set period and uses a generative AI model to analyze the relationship between product placement and sales performance. This analysis identifies the optimal placement pattern for specific products, maximizing sales.
[1047] Proposal for optimal layout
[1048] Based on the analysis results, the server generates an optimal product placement proposal. This proposal is visualized, allowing the user to compare the current product shelf layout with the proposed optimal placement. The proposal is then sent to the terminal and notified to the user.
[1049] Introducing the Emotion Engine
[1050] The device reads product information and inventory status to the user aloud and uses an emotion engine to recognize the user's emotional state by analyzing the user's voice and facial expressions to identify their emotions.
[1051] Adjusting notification content
[1052] Based on the emotional state identified by the emotion engine, the server dynamically adjusts the notification content, for example, providing more concise and positive suggestions if the user is feeling stressed.
[1053] Correcting misjudgments and providing management advice
[1054] The user can make corrections to the voice-readout information by voice command. For example, the user can input a correction such as "There are 15 bags in stock." The server reflects this correction in the database and performs recalculation. Based on the results of periodic inventory, the server also generates management advice for ordering and stock disposal and sends it to the terminal. The user can make effective management decisions based on this advice.
[1055] Specific examples
[1056] For example, analysis may reveal that sales data for a particular product category (e.g., snacks) are sluggish when placed on the bottom shelf. Based on this, the server suggests placing snacks in front of the cash register or at eye level. The suggestions are sent to the device as a concrete visual map, and the user adjusts product placement accordingly. At the same time, an emotion engine analyzes the user's emotional state, providing detailed suggestions if the user has a positive reaction and simple suggestions if the user has a negative reaction. For example, if a voice notification reads, "There are 10 bags of snacks in stock," the user can correct it with a voice command, saying, "There are 15 bags in stock," and the system will immediately update the information.
[1057] These functions provide a system that improves the efficiency of product management and inventory management, increases sales, responds flexibly to user emotions, and enables more accurate management decisions.
[1058] The processing flow will be explained below.
[1059] Step 1:
[1060] Users take a photo of the product shelves in a store using a device such as a smartphone or dedicated camera, and by pressing the capture button, the image is automatically saved on the device.
[1061] Step 2:
[1062] The device automatically sends the captured image data to a server via Wi-Fi or mobile networks.
[1063] Step 3:
[1064] The server analyzes the image data received from the device and uses an image analysis engine to identify products in the image and determine the location and quantity of each product.
[1065] Step 4:
[1066] The server stores the identified product information (SKU, location, stock quantity, etc.) in a database, allowing the current stock status to be immediately grasped.
[1067] Step 5:
[1068] The server references sales data based on the information stored in the database, extracts past sales data using SQL queries, and analyzes the relationship between product placement and sales.
[1069] Step 6:
[1070] The server uses a generative AI model to analyze product placement and sales data, thereby identifying optimal product placement patterns.
[1071] Step 7:
[1072] The server generates optimal product placement proposals based on the analysis results, which are then visualized.
[1073] Step 8:
[1074] The server sends the generated placement proposal to the terminal, and the user can check the proposal displayed on the terminal.
[1075] Step 9:
[1076] The device will read the notification content aloud to the user using text-to-speech (TTS) technology.
[1077] Step 10:
[1078] The device recognizes the user's emotional state using an emotion engine that analyzes the user's voice and facial expressions in real time.
[1079] Step 11:
[1080] Based on the results of the emotion engine, the server dynamically adjusts the notification content, for example, if the user is feeling stressed, it will provide concise, positive suggestions.
[1081] Step 12:
[1082] The user can correct the contents of the voice message by entering a voice command, for example, "There are 15 bags in stock."
[1083] Step 13:
[1084] The server receives the user's voice command and updates the database with the corrected information, which is then recalculated and saved as the latest information.
[1085] Step 14:
[1086] The server generates management advice for ordering and stock disposal based on the results of periodic inventory. This advice is generated based on the amount of stock and sales trends.
[1087] Step 15:
[1088] The server then sends the generated management advice to the terminal, and the user makes decisions about ordering and inventory disposal based on this advice.
[1089] Through these steps, this system will improve the efficiency of product management and inventory management, increase sales, improve the accuracy of management decisions, and even realize flexible responses that take user emotions into consideration. As a specific example, when the stock of a specific product is low, the system will notify the user by saying, "Stock is low. Please consider placing an order," but if the user is feeling emotionally stressed, it will adjust the wording to a softer one, saying, "There is no need to take immediate action. Please check when you have time."
[1090] Example 2
[1091] Next, a description will be given of Example 2. In the following description, the data processing device 12 will be referred to as a "server" and the headset type terminal 314 will be referred to as a "terminal."
[1092] In the retail industry, it is extremely important to accurately analyze the impact that product placement has on sales and propose optimal placements. However, currently, there are limited methods for optimizing shelf placement and proposal systems that take users' emotional states into account. This has prevented efficient inventory management and increased sales from being fully realized. Furthermore, when incorrect product information is identified, correcting it is cumbersome, requiring quick and appropriate management decisions. Given this background, a system is needed that can automatically analyze the relationship between product placement and sales and respond flexibly while taking emotions into consideration.
[1093] The specific processing by the specific processing unit 290 of the data processing device 12 in the second embodiment is realized by the following means.
[1094] In this invention, the server includes means for acquiring image data from product shelves in a store, means for analyzing the acquired image data to identify the product locations and quantities, means for storing the identified information in a database, means for using a generative model to analyze the relationship between product placement and sales by referencing sales data based on the information stored in the database, means for proposing optimal product placement based on the analysis results, means for visualizing the proposal and notifying the terminal, means for reading the notification content aloud, and means for recognizing the user's emotional state from the terminal and dynamically adjusting the notification content based on the user's emotional state. This enables accurate analysis of the relationship between product placement and sales and proposals for optimal placement that take emotions into consideration. It also makes it easier to correct erroneous judgments and generate management advice, enabling quick and accurate management decisions.
[1095] 1. "Image data" refers to digital information, such as still images or videos, that record the state of product shelves in a store.
[1096] 2. "Analysis" refers to the computational process used to identify the location and quantity of products from the acquired image data.
[1097] 3. "Database" means a system for systematically storing and managing specific product information and sales data.
[1098] 4. "Sales Data" means numerical information on the sales of products within a specific period of time.
[1099] 5. A "generative model" is a machine learning algorithm that learns from large amounts of data to derive trends and patterns.
[1100] 6. "Visualization" refers to displaying analysis results and placement proposals as graphs and charts so that users can intuitively understand them.
[1101] 7. "Terminal" means a computer device used by a user to operate the device, in particular a smartphone or tablet device.
[1102] 8. "Notification Content" means the output of information or suggestions provided by the System to the User in the form of text or audio.
[1103] 9. "Emotional state" is state information that represents the user's emotional response and is identified from the user's voice and facial expression.
[1104] 10. "Dynamic adjustment" means changing your response to the situation in real time.
[1105] This invention is a system that acquires image data of product shelves in a store, analyzes the image data to identify the location and quantity of products, and then analyzes sales data based on this to propose optimal product placement.The system uses image recognition technology, generative AI, and an emotion engine, and is able to flexibly respond by taking into account the user's emotional state.
[1106] Equipment and software configuration
[1107] The system consists of the following main components:
[1108] Terminal: A device equipped with a camera for taking pictures of product shelves in a store. Examples include smartphones and dedicated cameras.
[1109] Server: The central data processing unit that receives and analyzes image data. It runs the generative AI model and emotion engine.
[1110] User: The person who operates the system and makes optimal product placement and management decisions.
[1111] Processing steps
[1112] 1. Image data collection
[1113] A user uses a device to take an image of a product shelf in a store. For example, they can use a smartphone camera app to take a picture, focusing on a specific area. The captured image is set to be automatically sent to the server. A dedicated app encodes the image data and sends it to the server using a secure communication protocol (e.g., HTTPS).
[1114] 2. Analysis of image data
[1115] The server analyzes the image data received from the device. Specifically, it uses software such as OpenCV and TensorFlow to identify the location and quantity of products. For example, it uses the YOLO (You Only Look Once) object detection algorithm. The analysis results are stored in a database.
[1116] 3. Analysis of the relationship between placement and sales
[1117] The server references product information and sales data stored in a database and analyzes the relationship between product placement and sales performance using a generative AI model (e.g., H2O.ai's AutoML function). This analysis identifies the optimal placement pattern for a specific product.
[1118] 4. Proposal for optimal layout
[1119] Based on the analysis results, the server generates optimal product placement proposals, which are visualized and displayed in an intuitive format using Tableau or Power BI, and the proposals are sent to the device and notified to the user.
[1120] 5. Introducing the Emotion Engine
[1121] The device reads product information and stock availability to the user aloud, while simultaneously recognizing the user's emotional state using the Microsoft Azure Emotion API and IBM Watson Tone Analyzer. Based on the user's emotional state, the device dynamically adjusts the content of notifications.
[1122] 6. Correcting misjudgments and providing management advice
[1123] The user can make corrections to the information read out loud using voice commands. For example, they can input a correction such as "There are 15 bags in stock." The server reflects this correction in the database and performs recalculation. In addition, based on the results of periodic inventory, the server generates management advice for ordering and stock disposal and sends it to the terminal.
[1124] Specific examples
[1125] For example, if snacks are placed on the bottom shelf, analysis may reveal poor sales data. Based on this, the server suggests placing snacks at eye level. The suggestions are sent to the device as a concrete visual map, allowing the user to adjust product placement. At the same time, an emotion engine analyzes the user's emotional state, providing detailed suggestions if the user has a positive reaction and simple suggestions if the user has a negative reaction. For example, if a voice notification reads, "There are 10 bags of snacks in stock," the user can correct it with a voice command, saying, "There are 15 bags in stock," and the system will immediately update the information.
[1126] Prompt Sentence Examples
[1127] 1. "What image recognition technology do you use to optimize product placement in your stores?"
[1128] 2. "Please show the results of analyzing sales data using a generative AI model."
[1129] 3. "How does the Emotion Engine analyze the user's emotional state and adjust its response?"
[1130] This system will enable more efficient product and inventory management, increased sales, flexible responses that take user feelings into consideration, and more accurate management decisions.
[1131] The flow of the identification process in the second embodiment will be described with reference to FIG.
[1132] Step 1:
[1133] The user uses a device to take a picture of a specific product shelf in a store. The input is a high-resolution image taken using a smartphone camera app, for example. The output is image data with a clear focus on the captured product. Specifically, the image is taken so that the product's position and quantity are visible, and the image data is saved on the device.
[1134] Step 2:
[1135] The device sends the captured image data to the server. The input is the image data acquired in step 1. The output is the image data securely sent to the server. Specifically, a dedicated application encodes the image data into Base64 format and uploads it to the server in real time using the HTTPS protocol.
[1136] Step 3:
[1137] The server analyzes the image data received from the device. The input is the image data sent in step 2. The output is data on the location and quantity of products identified through image analysis. Specifically, it uses TensorFlow and OpenCV to apply object detection algorithms such as YOLO to extract the bounding boxes and labels of each product in the image.
[1138] Step 4:
[1139] The server stores the analysis results in a database. The input is the product location and quantity data identified in step 3. The output is product information stored in the database. Specifically, it is inserted into a MySQL database using an ORM (Object-Relational Mapping) library such as SQLAlchemy.
[1140] Step 5:
[1141] The server retrieves sales data from the database and uses a generative model to analyze the relationship between product placement and sales. The input is the product information and sales data stored in the database. The output is the optimal product placement pattern. Specifically, it uses H2O.ai's AutoML function to derive placement patterns that are expected to increase sales.
[1142] Step 6:
[1143] The server generates optimal product placement proposals based on the analysis results. The input is the optimal product placement pattern obtained in step 5. The output is a visualized placement proposal. Specifically, it is visualized as intuitive graphs and charts using Tableau or Power BI.
[1144] Step 7:
[1145] The server sends the visualized proposal to the terminal and notifies the user. The input is the placement proposal visualized in step 6. The output is the proposal notified to the user. Specifically, it is visually displayed on the terminal via the Internet.
[1146] Step 8:
[1147] The device reads out product information and stock status aloud, while simultaneously recognizing the user's emotional state. The input is the recommendations sent from the server and the user's real-time voice data. The output is the results of recognizing the user's emotional state and the adapted notification content. Specifically, it uses the Microsoft Azure Emotion API and IBM Watson Tone Analyzer to analyze the user's voice and facial expressions.
[1148] Step 9:
[1149] The server dynamically adjusts the notification content based on the emotional state identified by the emotion engine. The input is the user's emotional state sent from the device. The output is the adjusted notification content. For example, if the user is feeling stressed, more concise and positive suggestions are provided.
[1150] Step 10:
[1151] The user inputs a voice command in response to the voice readout, and the server reflects the correction in the database. The input is the user's voice command and the current contents of the database. The output is updated information in the database that reflects the correction. For example, if the correction is made to "There are 15 bags in stock," the information is immediately updated in the database.
[1152] As a result, this system analyzes the relationship between product placement in a store and sales, and realizes flexible and adaptable placement suggestions and inventory management that take into account the user's emotional state.
[1153] (Application example 2)
[1154] Next, a description will be given of Application Example 2. In the following description, the data processing device 12 will be referred to as a "server" and the headset type terminal 314 will be referred to as a "terminal."
[1155] Conventional product placement support systems for the retail industry require a lot of effort to identify product locations and inventory information, and their placement change suggestions do not take into account the user's emotional state, making it difficult to maximize management efficiency and sales. Furthermore, because real-time product placement changes and inventory adjustments are time-consuming, there is a demand for faster management decision-making.
[1156] The identification process by the identification processing unit 290 of the data processing device 12 in Application Example 2 is realized by the following means. In this invention, the server includes means for acquiring image data from product shelves in a store, means for analyzing the acquired image data and identifying the location and quantity of products, means for saving the identified information in a database, means for referencing sales data based on the information saved in the database and analyzing the relationship between product placement and sales, means for proposing an optimal product placement based on the analysis results, means for visualizing the proposal and notifying the user, means for reading the notification content aloud to the user, means for a user to use a smart device to take a real-time photo of the product shelf situation in the store and send the image data to the server, means for analyzing sales data and product information using a generative AI model, means for recognizing the user's emotional state and adjusting the notification content based on the user's emotion, and means for providing real-time information updates and advice in cooperation with a voice assistant. This enables efficient product management and inventory management, maximizing sales, flexible responses based on user emotions, and faster management decisions.
[1157] "Image data" is image information obtained by photographing a product shelf, and is data used to identify the position and quantity of products.
[1158] "In-store shelves" refers to shelves and racks used to display, showcase, and store products in retail stores and brick-and-mortar stores.
[1159] "Means" means a method, technique, or device for achieving a particular purpose.
[1160] "Analysis" refers to the process of analyzing the acquired image data and extracting detailed information such as the location and quantity of the products.
[1161] A "database" is an information system that can efficiently store, manage, and search large amounts of data.
[1162] "Sales data" refers to data containing sales information for a product within a certain period of time, including sales amount, sales quantity, sales date, and the like.
[1163] A "generative AI model" is an algorithm that uses artificial intelligence to analyze data and generate new insights and predictions.
[1164] "Visualization" is the process of presenting data or information in a visually understandable way.
[1165] "User" refers to the person who operates the system and manages product placement and inventory.
[1166] "Notification" is the act of transmitting information from the system to the user.
[1167] A "voice assistant" is a system that interacts with users using voice input to assist them in obtaining information and performing operations.
[1168] "Smart devices" refer to electronic devices equipped with advanced computing power and communication functions, including smartphones and smart glasses.
[1169] "Emotional state" refers to a user's emotional response or state, including happiness, sadness, stress, etc.
[1170] "Real-time" means processing or reacting nearly simultaneously or without delay.
[1171] System Configuration
[1172] This system acquires images of in-store product shelves and optimizes product placement using generative AI models and emotion engines. The system consists of the following main hardware and software components:
[1173] 1. Terminal
[1174] Smart devices include smart glasses and smartphones, which are equipped with cameras and are used to capture images of the product shelves in stores.
[1175] 2. Server
[1176] OpenCV is used for image analysis, which analyzes image data and identifies the location and quantity of products.
[1177] The database used is MySQL, and the analyzed data is saved in the database.
[1178] The generative AI model uses OpenAI's GPT-4 and other models to analyze sales data and product information and generate optimal placement patterns.
[1179] The emotion engine uses Affectiva's SDK to analyze the user's emotional state.
[1180] As a voice assistant, it uses Google Cloud Text-to-Speech to read out the generated notification content aloud.
[1181] 3. Users
[1182] Managers use smart devices to take photos of the shelves and send the data to a server in real time.
[1183] Processing steps
[1184] Image data collection
[1185] The user uses smart glasses or a smartphone to take an image of the store's shelves, focusing on a specific area of the shelves to clearly show the location and quantity of products. The device is set to automatically send the captured image to the server.
[1186] Image data analysis
[1187] The server analyzes the image data received from the device. The image analysis engine uses OpenCV to identify the location, quantity, and specific product name of each product. After the product information is identified, it is saved in a database.
[1188] Saving to a database
[1189] The identified product information is stored in a database along with attributes such as SKU, location, and inventory quantity, making it possible to immediately grasp the current inventory status.
[1190] Analysis of the relationship between placement and sales
[1191] The server references sales data from a database for a set period and uses a generative AI model to analyze the relationship between product placement and sales performance. This analysis identifies the optimal placement pattern for specific products, maximizing sales.
[1192] Proposal for optimal layout
[1193] Based on the analysis results, the server generates an optimal product placement proposal. This proposal is visualized, allowing the user to compare the current product shelf layout with the proposed optimal placement. The proposal is then sent to the terminal and notified to the user.
[1194] Introducing the Emotion Engine
[1195] The device reads product information and inventory status to the user aloud and uses an emotion engine to recognize the user's emotional state by analyzing the user's voice and facial expressions to identify their emotions.
[1196] Adjusting notification content
[1197] Based on the emotional state identified by the emotion engine, the server dynamically adjusts the notification content, for example, providing more concise and positive suggestions if the user is feeling stressed.
[1198] Correcting misjudgments and providing management advice
[1199] The user can make corrections to the voice-read information by voice command. For example, the user can input a correction such as "15 bags in stock." The server reflects this correction in the database and performs recalculation. In addition, based on the results of periodic inventory, the server generates management advice for ordering and stock disposal and sends it to the terminal. The user can make effective management decisions based on this advice.
[1200] Specific examples
[1201] For example, analysis may reveal that sales data for a particular product category (e.g., snacks) are sluggish when placed on the bottom shelf. Based on this, the server suggests placing snacks in front of the register or at eye level. The suggestions are sent to the device as a concrete visual map, and the user adjusts product placement accordingly. The emotion engine also analyzes the user's emotional state, providing detailed suggestions if the user has a positive reaction and simple suggestions if the user has a negative reaction.
[1202] Prompt Sentence Examples
[1203] "Please suggest the optimal product placement in the store. The current placement and sales data are as follows: \[Specific data\]"
[1204] The flow of the specific processing in the application example 2 will be described with reference to FIG.
[1205] Step 1:
[1206] A user uses a smart device (smart glasses or smartphone) to take a picture of the product shelf situation in a store in real time. The captured image data is automatically sent from the device to a server. The input is the image of the product shelf, and the output is the image data sent to the server.
[1207] Step 2:
[1208] The server analyzes the received image data. At this time, an image analysis engine (such as OpenCV) is used to identify the location and quantity of each product. Specifically, object recognition and extraction are performed on the image data. The input is the received image data, and the output is specific information such as the product's location, quantity, and product name.
[1209] Step 3:
[1210] The server stores the identified product information in a database (e.g., MySQL). Here, attributes such as SKU, location, and stock quantity are stored as data. The input is analysis information such as product location and quantity, and the output is product information stored in the database.
[1211] Step 4:
[1212] The server references the information stored in the database and sales data, and uses a generative AI model to analyze the relationship between product placement and sales data. At this time, the generative AI model (such as OpenAI's GPT-4) receives a prompt message: "Please suggest the optimal placement of products in the store. The current placement data and sales data are as follows: \[Specific data\]". The input is the current placement data and sales data, and the output is a proposal for the optimal product placement.
[1213] Step 5:
[1214] The server visualizes the generated optimal placement proposal and sends the results to the terminal. Specifically, a visual map is created. The input is the placement proposal obtained from the generative AI model, and the output is the visualized placement proposal and a notification of it.
[1215] Step 6:
[1216] The device reads out product information and stock availability to the user and analyzes the user's emotional state using an emotion engine (such as Affectiva's SDK). The emotion engine recognizes emotions and sends the results to the server. The input is the user's voice and facial expression, and the output is the identified emotional state.
[1217] Step 7:
[1218] The server dynamically adjusts the notification content based on the emotional state obtained from the emotion engine. If the user is feeling stressed, it provides more concise and positive suggestions, and if the user has a positive reaction, it provides more detailed suggestions. The input is the emotional state obtained from the emotion engine, and the output is the adjusted notification content.
[1219] Step 8:
[1220] The user can make corrections to the voice readout as necessary using voice commands. For example, they can input the corrections, such as "There are 15 bags in stock." The server reflects this correction information in the database and performs recalculation. The input is the user's voice command, and the output is the corrected database information.
[1221] Step 9:
[1222] The server generates management advice for ordering and stock disposal based on the results of periodic inventory and sends it to the terminal. The user makes effective management decisions based on this management advice. The input is the inventory result data, and the output is the generated management advice.
[1223] The specific processing unit 290 transmits the result of the specific processing to the headset type terminal 314. In the headset type terminal 314, the control unit 46A causes the speaker 240 and the display 343 to output the result of the specific processing. The microphone 238 acquires audio indicating a user input regarding the result of the specific processing. The control unit 46A transmits audio data indicating the user input acquired by the microphone 238 to the data processing device 12. In the data processing device 12, the specific processing unit 290 acquires the audio data.
[1224] The data generation model 58 is a so-called generative AI (Artificial Intelligence). An example of the data generation model 58 is ChatGPT (Internet Search<URL: https: / / openai.com / blog / chatgpt> ), Gemini (Internet search <url: https: gemini.google.com ?hl="ja">) and other generation AIs. The data generation model 58 is obtained by performing deep learning on a neural network. A prompt including an instruction is input to the data generation model 58, and inference data such as voice data indicating voice, text data indicating text, and image data indicating an image is also input. The data generation model 58 performs inference on the input inference data in accordance with the instruction indicated by the prompt, and outputs the inference result in a data format such as voice data and text data. Here, inference refers to, for example, analysis, classification, prediction, and / or summarization.
[1225] In the above embodiment, an example was given in which the specific processing is performed by the data processing device 12, but the technology of the present disclosure is not limited to this, and the specific processing may be performed by the headset type terminal 314.
[1226] [Fourth embodiment]
[1227] FIG. 7 shows an example of the configuration of a data processing system 410 according to the fourth embodiment.
[1228] 7, a data processing system 410 includes a data processing device 12 and a robot 414. An example of the data processing device 12 is a server.
[1229] The data processing device 12 includes a computer 22, a database 24, and a communication I / F 26. The computer 22 is an example of a "computer" according to the technology of the present disclosure. The computer 22 includes a processor 28, a RAM 30, and a storage 32. The processor 28, the RAM 30, and the storage 32 are connected to a bus 34. The database 24 and the communication I / F 26 are also connected to the bus 34. The communication I / F 26 is connected to a network 54. Examples of the network 54 include a WAN (Wide Area Network) and / or a LAN (Local Area Network).
[1230] The robot 414 includes a computer 36, a microphone 238, a speaker 240, a camera 42, a communication I / F 44, and a control target 443. The computer 36 includes a processor 46, a RAM 48, and a storage 50. The processor 46, the RAM 48, and the storage 50 are connected to a bus 52. The microphone 238, the speaker 240, the camera 42, and the control target 443 are also connected to the bus 52.
[1231] The microphone 238 receives instructions and the like from the user 20 by receiving voice uttered by the user 20. The microphone 238 captures the voice uttered by the user 20, converts the captured voice into audio data, and outputs it to the processor 46. The speaker 240 outputs audio in accordance with instructions from the processor 46.
[1232] Camera 42 is a small digital camera equipped with an optical system including a lens, aperture, and shutter, and an imaging element such as a CMOS (Complementary Metal-Oxide-Semiconductor) image sensor or a CCD (Charge Coupled Device) image sensor, and captures images of the surroundings of user 20 (for example, an imaging range defined by an angle of view equivalent to the field of vision of a typical healthy person).
[1233] The communication I / F 44 is connected to a network 54. The communication I / Fs 44 and 26 are responsible for the exchange of various information between the processor 46 and the processor 28 via the network 54. The exchange of various information between the processor 46 and the processor 28 using the communication I / Fs 44 and 26 is carried out in a secure state.
[1234] The control object 443 includes a display device, LEDs in the eyes, and motors for driving the arms, hands, and feet. The posture and gestures of the robot 414 are controlled by controlling the motors of the arms, hands, and feet. Some of the emotions of the robot 414 can be expressed by controlling these motors. In addition, the facial expressions of the robot 414 can also be expressed by controlling the light emission state of the LEDs in the eyes of the robot 414.
[1235] Fig. 8 shows an example of the main functions of the data processing device 12 and the robot 414. As shown in Fig. 8, in the data processing device 12, a specific process is performed by the processor 28. A specific process program 56 is stored in the storage 32.
[1236] The specific processing program 56 is an example of a "program" according to the technology of the present disclosure. The processor 28 reads the specific processing program 56 from the storage 32 and executes the read specific processing program 56 on the RAM 30. The specific processing is realized by the processor 28 operating as a specific processing unit 290 in accordance with the specific processing program 56 executed on the RAM 30.
[1237] The storage 32 stores a data generation model 58 and an emotion identification model 59. The data generation model 58 and the emotion identification model 59 are used by the identification processing unit 290.
[1238] In the robot 414, the processor 46 performs the reception output process. A reception output program 60 is stored in the storage 50. The processor 46 reads the reception output program 60 from the storage 50 and executes the read reception output program 60 on the RAM 48. The reception output process is realized by the processor 46 operating as the control unit 46A in accordance with the reception output program 60 executed on the RAM 48.
[1239] Next, a description will be given of the specific processing performed by the specific processing unit 290 of the data processing device 12. In the following description, the data processing device 12 will be referred to as a "server" and the robot 414 will be referred to as a "terminal."
[1240] The present invention relates to an optimal product placement support system for the retail industry that uses image recognition technology and generative AI, and aims to analyze the impact of product placement in a store on sales and optimize product placement. Specific implementation methods of the present invention are described below.
[1241] System configuration
[1242] The system consists of the following main components:
[1243] 1. Terminal: Equipped with a camera to take pictures of product shelves in a store and has the function of sending image data to a server. Examples include smartphones and dedicated cameras.
[1244] 2. Server: The central data processing unit that receives and analyzes image data, running the generative AI model.
[1245] 3. User: The person who operates the system and makes management decisions such as optimal product placement, ordering, and inventory disposal.
[1246] Program processing flow
[1247] Image data collection
[1248] The user uses the device to take an image of the store's shelves, focusing on a specific area of the shelves to clearly show the location and quantity of products. The device is configured to automatically send the captured image to the server.
[1249] Image data analysis
[1250] The server analyzes the image data received from the device. Using an image analysis engine, it identifies the location, quantity, and specific product name of each product. After the product information is identified, it is stored in a database.
[1251] Saving to a database
[1252] The identified product information is stored in a database along with attributes such as SKU, location, and inventory quantity, making it possible to immediately grasp the current inventory status.
[1253] Analysis of the relationship between placement and sales
[1254] The server references sales data from a database for a set period and uses a generative AI model to analyze the relationship between product placement and sales performance. This analysis identifies the optimal placement pattern for specific products, maximizing sales.
[1255] Proposal for optimal layout
[1256] Based on the analysis results, the server proposes the optimal product placement. This proposal is visualized, allowing users to compare the current product shelf layout with the proposed optimal placement. The proposal is then sent to the terminal and notified to the user.
[1257] Voice reading and correction
[1258] Users can confirm the notification content from the device by voice. The device will read out product information and stock status using text-to-speech (TTS) technology. If there is an incorrect judgment, users can correct it using voice commands. This correction is reflected in the database again and recalculation is performed.
[1259] Management Advice
[1260] Based on the results of periodic inventory, the server generates management advice such as ordering and discount proposals for inventory clearance. These advices are notified to the user via text and voice, supporting effective management decisions.
[1261] Specific examples
[1262] For example, analysis may reveal that sales data for a particular product category (e.g., snacks) are stagnating if they are placed on the bottom shelf. Based on this, the server may suggest placing snacks in front of the register or at eye level. The suggestion is sent to the device as a concrete visual map, and the user can adjust product placement based on the suggestion. Also, if a voice notification reads, "There are 10 bags of snacks in stock," the user can correct it with a voice command, saying, "There are 15 bags in stock," and the system will immediately update the information.
[1263] These features provide a system that streamlines product and inventory management, increases sales, and enables more accurate business decisions.
[1264] The processing flow will be explained below.
[1265] Step 1:
[1266] Users take a photo of the product shelves in a store using a device such as a smartphone or dedicated camera, and by pressing the capture button, the image is automatically saved on the device.
[1267] Step 2:
[1268] The device automatically sends the captured image data to a server via Wi-Fi or mobile networks.
[1269] Step 3:
[1270] The server analyzes the image data received from the device and uses an image analysis engine to identify products in the image and determine the location and number of each product.
[1271] Step 4:
[1272] The server stores the identified product information in a database, including the SKU, location, and quantity in stock.
[1273] Step 5:
[1274] The server references sales data based on the information stored in the database, extracts past sales data using SQL queries, and analyzes the relationship between product placement and sales.
[1275] Step 6:
[1276] The server uses a generative AI model to analyze product placement and sales data, thereby identifying optimal product placement patterns.
[1277] Step 7:
[1278] The server generates optimal product placement proposals based on the analysis results, which are then visualized.
[1279] Step 8:
[1280] The server sends the generated placement proposal to the terminal, and the user can check the proposal displayed on the terminal.
[1281] Step 9:
[1282] The device reads product information and stock availability to the user aloud, and uses text-to-speech (TTS) technology to audibly output the contents of the visual suggestions.
[1283] Step 10:
[1284] The user can make corrections using voice commands in response to the content of the voice readout. For example, the user can input a correction such as "There are 15 bags in stock."
[1285] Step 11:
[1286] The server receives voice correction commands from the user and reflects them in the database, where the corrected information is recalculated and saved as the latest information.
[1287] Step 12:
[1288] The server generates management advice for ordering and stock disposal based on the results of periodic inventory. This advice is generated based on the amount of stock and sales trends.
[1289] Step 13:
[1290] The server then sends the generated management advice to the terminal, and the user makes decisions about ordering and inventory disposal based on this advice.
[1291] Through the above steps, this system will improve the efficiency of product management and inventory management, thereby increasing sales and improving the accuracy of management decisions.
[1292] Example 1
[1293] Next, a description will be given of Example 1. In the following description, the data processing device 12 will be referred to as a "server" and the robot 414 will be referred to as a "terminal."
[1294] The retail industry is seeking to optimize product placement and streamline inventory management. Conventional methods rely on manual work to determine product locations and quantities, which takes time and effort and is prone to errors. It is also difficult to analyze the relationship between product placement and sales, requiring significant effort to find optimal product placement. Furthermore, updating inventory information in real time is difficult, making it difficult to make quick management decisions. To solve these issues, a system is needed that can efficiently and accurately manage product placement and inventory.
[1295] The specific processing by the specific processing unit 290 of the data processing device 12 in the first embodiment is realized by the following means.
[1296] In this invention, the server includes a means for analyzing image data and identifying product locations and quantities, a means for proposing optimal product placement using a generative AI model, and a means for referencing sales data based on information stored in a database and analyzing the relationship between product placement and sales. This makes it possible to identify product locations and quantities with high accuracy, update inventory information in real time, and quickly propose optimal product placement to maximize sales.
[1297] "Image data" refers to visual information acquired in the form of photographs or videos of product shelves in a store.
[1298] "Acquisition" refers to the act of obtaining image data of an object using a camera or other photographic equipment.
[1299] "Analysis" is the process of extracting and identifying specific information from acquired image data.
[1300] A "generative AI model" is an artificial intelligence algorithm trained to perform image recognition and data analysis.
[1301] "Location" is information indicating the specific location where a particular product is located.
[1302] "Quantity" is a number that indicates how much of a particular product is on the shelf or in stock.
[1303] A "database" is an electronic recording device or system for systematically storing and managing specific information.
[1304] "Sales data" refers to transaction information when a product is sold, including sales amount and sales quantity.
[1305] "Analysis" is the act of examining and interpreting data to find specific relationships and patterns.
[1306] "Suggestion" refers to showing optimal behavior and placement patterns based on the analysis results.
[1307] "Visualization" refers to displaying data and proposals in a visually easy-to-understand manner.
[1308] "Notification" is the act of informing a user of specific information.
[1309] "Text-to-speech" is a technology that converts text information into speech and conveys it to the user auditorily.
[1310] "Voice command" is a method by which a user gives instructions or inputs using their voice.
[1311] "Inventory" refers to the periodic checking and recording of stock and the quantity of merchandise items.
[1312] "Management advice" refers to proposals based on inventory and sales data to support decision-making such as ordering and inventory disposal.
[1313] This invention is an optimal product placement support system for the retail industry that uses image recognition technology and generative AI models. The system acquires image data of product shelves in stores, analyzes the data to identify product locations and quantities, stores the data in a database, and then compares it with sales data to suggest optimal product placement.
[1314] The system consists of three main components:
[1315] 1. Terminal: A device with a camera function that takes pictures of product shelves in a store. Typical examples are smartphones or dedicated cameras.
[1316] 2. Server: A central data processing unit that receives data, performs image analysis, and optimizes product placement using generative AI models.
[1317] 3. User: The person who operates the system and makes management decisions such as optimal product placement, ordering, and inventory disposal.
[1318] Hardware and Software Configuration
[1319] Terminal
[1320] Users take pictures of product shelves in stores using smartphones or dedicated cameras. These devices have the ability to automatically send the captured image data to a server. The captured images are sent to the server via Wi-Fi or mobile data. The devices are also equipped with text-to-speech (TTS) technology, which allows them to notify users by voice.
[1321] server
[1322] The server receives the image data sent from the device and uses an image analysis engine to identify the location, quantity, and product name of each product. This analysis uses deep learning technology and generative AI models, such as machine learning libraries like TensorFlow and PyTorch. The server then stores the identified information in a database and uses that data to analyze the relationship between sales data and product placement.
[1323] Database
[1324] The product information analyzed by the server is stored in a database, which updates attributes such as SKU (Stock Keeping Unit), location, and inventory quantity in real time. Users can easily check this information.
[1325] Examples of use and specific prompts
[1326] Users take pictures of product shelves in a store and send them to a server using their device. The server then receives the image data, analyzes it, identifies the product locations and quantities, and stores them in a database. Based on the analysis results, the server uses a generative AI model to propose optimal product placement and sends it to the device as a visual map. The user can then confirm the proposal visually and audibly and adjust the product placement.
[1327] Prompt Sentence Examples
[1328] "Please identify the names and quantities of the items shown in this image."
[1329] "Please suggest optimal product placement based on sales data and product placement data."
[1330] "Please read out the inventory."
[1331] · "Use voice commands to correct stock levels."
[1332] This will lead to more efficient product and inventory management, increased sales, and more accurate business decisions.
[1333] The flow of the identification process in the first embodiment will be described with reference to FIG.
[1334] Step 1:
[1335] The user uses a device to take a photo of a product shelf in a store. The captured image is high resolution, and the product name and quantity are clearly visible. For example, a smartphone or dedicated camera is used to capture product images under appropriate lighting. The input is an "image of a product shelf in a store," and the output is "high-resolution image data."
[1336] Step 2:
[1337] The device sends the captured image data to the server. The device has a network connection function and encrypts and sends the image data to the server using Wi-Fi or mobile data. The input is "high-resolution image data" and the output is "a notification of completion of transmission to the server."
[1338] Step 3:
[1339] The server analyzes the image data received from the device. The image analysis engine processes the input image and identifies the location, quantity, and name of the product. A generative AI model is used for this analysis. Specifically, machine learning libraries such as TensorFlow and PyTorch are used. The input is "high-resolution image data," and the output is "the location, quantity, and name of the identified product."
[1340] Step 4:
[1341] The server stores the identified information in a database. The identified product information (SKU, location, and stock quantity) is recorded in the database, and inventory information is updated in real time. The input is the "identified product location, quantity, and name," and the output is the "updated database entry."
[1342] Step 5:
[1343] The server analyzes the relationship between sales data and product placement performance based on the stored information. It uses a generative AI model to find the optimal placement pattern. For example, it compares the placement location of a specific product with sales data to identify the placement that maximizes sales. The inputs are "product information stored in the database" and "sales data," and the output is "identification of the optimal placement pattern."
[1344] Step 6:
[1345] The server proposes optimal product placement and generates a visualized proposal. It provides the user with a layout diagram in an easy-to-understand visual format, allowing them to compare the proposed placements. The proposal is sent to the device. The input is the "identification result of the optimal placement pattern," and the output is the "visualized placement proposal."
[1346] Step 7:
[1347] The terminal notifies the user of the proposal content sent from the server by voice. Using text-to-speech (TTS) technology, the placement proposal content and inventory information are read aloud. The input is a "visualized placement proposal" and the output is a "voice notification."
[1348] Step 8:
[1349] When a user wants to correct a voice notification, they input a voice command. For example, they can correct an inventory quantity by voice, and the correction is sent to the server. The server analyzes the voice command and updates the database. The input is the "user's voice command" and the output is the "updated database entry."
[1350] Step 9:
[1351] Based on the results of periodic inventory, the server generates management advice for ordering and stock disposal. It then notifies the user of this advice via text and voice. For example, the advice might be, "There is an excess of snack food in stock. We suggest selling them at a discount." The inputs are "inventory results" and "database information," and the output is "notification of management advice."
[1352] (Application example 1)
[1353] Next, a description will be given of Application Example 1. In the following description, the data processing device 12 will be referred to as a "server" and the robot 414 will be referred to as a "terminal."
[1354] It is widely recognized in the retail industry that product placement in a store has a significant impact on sales. However, optimizing shelf placement using traditional methods requires a great deal of effort and time, and effective placement is not always achieved. Furthermore, real-time inventory management and rapid management decisions based on sales data are required, but traditional systems do not adequately support this. Furthermore, store staff are required to manually check and correct product locations and inventory status, which is time-consuming and inefficient.
[1355] The specific processing by the specific processing unit 290 of the data processing device 12 in the application example 1 is realized by the following means.
[1356] In this invention, the server includes means for acquiring image data from product shelves in a store, means for analyzing the acquired image data and identifying product locations and quantities, means for saving the identified information in a database, means for referencing sales data based on the information saved in the database and analyzing the relationship between product placement and sales, means for proposing optimal product placement based on the analysis results, means for visualizing the proposal and notifying the user, means for reading the notification aloud to the user, means for taking pictures of the product shelves using a smartphone, means for providing audio notification of product information and inventory status using text-to-speech technology, means for the user to correct information using voice commands, means for generating optimal placements using a generative AI model that contributes to maximizing sales, and means for inputting prompts to the generative AI model to obtain analysis results. This enables efficient acquisition of product shelf images using a smartphone, optimal placement suggestions using generative AI, and rapid and accurate inventory management and management decisions using voice technology.
[1357] "Image data" refers to information that is acquired as an image of product shelves in a store and is the subject of analysis.
[1358] "Analysis" refers to the act of identifying the location and quantity of products from the acquired image data.
[1359] A "database" is a repository of information that stores identified product information and allows for later reference.
[1360] "Sales data" is information relating to product sales, and is used to analyze the relationship between product placement and sales.
[1361] A "generative AI model" is an artificial intelligence technology used to analyze image data and propose optimal product placement.
[1362] A "prompt statement" is an instruction statement that is input into a generative AI model to obtain analysis results.
[1363] "Visualization" is the act of visually suggesting optimal product placement.
[1364] "Voice notification" refers to the act of using text-to-speech technology to inform users of product information and stock status via voice.
[1365] A "smartphone" is a mobile device used to acquire and send notifications about image data.
[1366] A "voice command" is a spoken input by a user to give instructions to a system using voice recognition technology.
[1367] "Inventory management" is the management activity of determining the quantity of goods and maintaining appropriate stock levels.
[1368] "Management decisions" are decisions made regarding product placement and inventory management.
[1369] MODE FOR CARRYING OUT THE INVENTION
[1370] The present invention relates to a system for optimizing product placement in a retail store using a smartphone. A specific implementation method of this system will be described below.
[1371] System configuration
[1372] The system consists of the following major components:
[1373] 1. Device (smartphone):
[1374] It is equipped with a camera to take pictures of product shelves in the store and transmits the image data to a server.
[1375] It has the ability to provide voice notification of product information and stock status using text-to-speech technology.
[1376] It uses voice recognition technology to receive the user's voice commands and send correction information to the server.
[1377] 2. Server:
[1378] It receives image data and runs a generative AI model that performs the analysis.
[1379] Using an image analysis engine (e.g., YOLOv5, TensorFlow), the location, quantity, and specific product name of each product are identified.
[1380] The identified product information is stored in a database (e.g., Firebase, PostgreSQL).
[1381] Sales data is referenced based on the information stored in the database, and the relationship between product placement and sales is analyzed.
[1382] The system generates the optimal layout that contributes to maximizing sales and notifies the user as a visual map.
[1383] 3. User:
[1384] A smartphone is used to take a photo of the product shelf and send the image data to the server.
[1385] Visually confirm the proposed optimal placement and confirm / correct inventory information notified by voice.
[1386] Make business decisions and place orders or dispose of inventory based on advice provided by the system.
[1387] Program processing
[1388] (Collection of image data)
[1389] Users take pictures of product shelves in a store using the camera on their smartphone, and the smartphone automatically sends the captured image data to the server.
[1390] (Image data analysis)
[1391] The server is equipped with an image analysis engine and generative AI models, such as YOLOv5 and TensorFlow, to analyze the received image data. The analysis identifies the location, quantity, and SKU of each product, and stores this information in a database.
[1392] (Saving to database)
[1393] The identified product information is stored in a database along with attributes such as SKU, location, and stock quantity, making it possible to instantly grasp the current inventory status.
[1394] (Analysis of the relationship between placement and sales)
[1395] The server references sales data for a set period from the database and uses a generative AI model to analyze the relationship between product placement and sales performance. At this time, it inputs prompt statements into the generative AI to obtain the analysis results.
[1396] (Optimal layout proposal)
[1397] A placement pattern that contributes to maximizing sales is generated, and the proposed details are sent to a smartphone as a visual map.
[1398] (audio notification and fix)
[1399] The smartphone uses text-to-speech technology (e.g., Google Cloud Text-to-Speech API) to notify the user of product information and stock availability via voice. When the user provides corrections via voice commands, the smartphone sends the information to the server, which updates the database.
[1400] (Providing management advice)
[1401] Based on the results of periodic inventory, the server generates management advice for ordering and stock disposal, which is communicated to the user via text and voice to support effective management decisions.
[1402] Specific examples
[1403] For example, if a particular snack is placed on the bottom shelf, analysis may reveal that sales data is sluggish. The server then suggests placing the snack in the center of the shelf where it is more visible. This suggestion is sent to the smartphone as a visual map, and the user follows the instructions to rearrange the products. Also, if a voice notification says, "There are 10 bags of snacks in stock," the user can correct it by saying, "There are 15 bags in stock," and the system will immediately update the database.
[1404] (Example of a prompt to input to a generative AI model)
[1405] input:
[1406] "Analyze the shelf image below to identify product locations, SKUs, and stock quantities. SKUs: '12345', '54321', '67890', etc."
[1407] output:
[1408] "Product locations and quantities identified: 4 units of SKU: '12345', 6 units of SKU: '54321', and 10 units of SKU: '67890'. Suggested placement: Place snacks in the center of the shelf."
[1409] The flow of the specific processing in the application example 1 will be described with reference to FIG.
[1410] Program processing steps
[1411] Step 1:
[1412] A user uses the camera on their smartphone to take a picture of a product shelf in a store. The image data taken by the smartphone is the input. The device sends this image data to the server. The output is the image data sent to the server.
[1413] Step 2:
[1414] The server analyzes the received image data. Specifically, it uses an image analysis engine (e.g., YOLOv5, TensorFlow) to identify the product location and quantity. The image data is input, and the product location, quantity, and SKU information are output.
[1415] Step 3:
[1416] The server stores the identified product information in a database. The input is the location, quantity, and SKU information of the identified product, and the product information stored in the database is obtained as the output.
[1417] Step 4:
[1418] The server references sales data based on the information stored in the database and analyzes the relationship between product placement and sales. This analysis is performed using a generative AI model (e.g., TensorFlow). Database information and sales data are input, and the analysis results include information on the relationship between product placement and sales.
[1419] Step 5:
[1420] The server inputs a prompt into the generative AI model to obtain the optimal product placement that will maximize sales. The inputs are the prompt and information from the database, and the output is a proposal for the optimal placement.
[1421] Step 6:
[1422] The server generates a visual map of the optimal layout proposal and sends it to the smartphone. The input is the optimal layout proposal, and the output is the visual map data.
[1423] Step 7:
[1424] The smartphone uses text-to-speech technology (e.g., Google Cloud Text-to-Speech API) to provide the user with product information and stock status via voice. The input is a visual map and product information, and the output is a voice notification.
[1425] Step 8:
[1426] The user listens to the voice notification and, if necessary, issues a voice command to correct the information. The smartphone receives this voice command and converts it into text using voice recognition technology (e.g., Google Speech-to-Text API). The voice command is used as input, and the text of the correction is output.
[1427] Step 9:
[1428] The server receives the user's corrections and updates the database: the text of the corrections is the input, and the database update is the output.
[1429] Step 10:
[1430] Based on the results of periodic inventory, the server generates management advice for ordering and stock disposal and notifies the user. Inventory data is input and management advice is output.
[1431] Combining these steps will improve the efficiency of in-store product placement and inventory management, and will also enable faster management decisions to maximize sales.
[1432] Furthermore, an emotion engine that estimates the user's emotion may be further combined. That is, the identification processing unit 290 may estimate the user's emotion using the emotion identification model 59, and perform identification processing using the user's emotion.
[1433] The present invention relates to an optimal product placement support system for the retail industry that uses image recognition technology, generative AI, and an emotion engine. The purpose of the system is to analyze the impact of in-store product placement on sales, optimize product placement, and also take into account the emotional state of users. Specific implementation methods of the present invention are described below.
[1434] System configuration
[1435] The system consists of the following main components:
[1436] 1. Terminal: Equipped with a camera to take pictures of product shelves in a store and has the function of sending image data to a server. Examples include smartphones and dedicated cameras.
[1437] 2. Server: The central data processing unit that receives and analyzes image data, running the generative AI model and emotion engine.
[1438] 3. User: The person who operates the system and makes management decisions such as optimal product placement, ordering, and inventory disposal.
[1439] Program processing flow
[1440] Image data collection
[1441] The user uses the device to take an image of the store's shelves, focusing on a specific area of the shelves to clearly show the location and quantity of products. The device is configured to automatically send the captured image to the server.
[1442] Image data analysis
[1443] The server analyzes the image data received from the device. Using an image analysis engine, it identifies the location, quantity, and specific product name of each product. After the product information is identified, it is stored in a database.
[1444] Saving to a database
[1445] The identified product information is stored in a database along with attributes such as SKU, location, and inventory quantity, making it possible to immediately grasp the current inventory status.
[1446] Analysis of the relationship between placement and sales
[1447] The server references sales data from a database for a set period and uses a generative AI model to analyze the relationship between product placement and sales performance. This analysis identifies the optimal placement pattern for specific products, maximizing sales.
[1448] Proposal for optimal layout
[1449] Based on the analysis results, the server generates an optimal product placement proposal. This proposal is visualized, allowing the user to compare the current product shelf layout with the proposed optimal placement. The proposal is then sent to the terminal and notified to the user.
[1450] Introducing the Emotion Engine
[1451] The device reads product information and inventory status to the user aloud and uses an emotion engine to recognize the user's emotional state by analyzing the user's voice and facial expressions to identify their emotions.
[1452] Adjusting notification content
[1453] Based on the emotional state identified by the emotion engine, the server dynamically adjusts the notification content, for example, providing more concise and positive suggestions if the user is feeling stressed.
[1454] Correcting misjudgments and providing management advice
[1455] The user can make corrections to the voice-readout information by voice command. For example, the user can input a correction such as "There are 15 bags in stock." The server reflects this correction in the database and performs recalculation. Based on the results of periodic inventory, the server also generates management advice for ordering and stock disposal and sends it to the terminal. The user can make effective management decisions based on this advice.
[1456] Specific examples
[1457] For example, analysis may reveal that sales data for a particular product category (e.g., snacks) are sluggish when placed on the bottom shelf. Based on this, the server suggests placing snacks in front of the cash register or at eye level. The suggestions are sent to the device as a concrete visual map, and the user adjusts product placement accordingly. At the same time, an emotion engine analyzes the user's emotional state, providing detailed suggestions if the user has a positive reaction and simple suggestions if the user has a negative reaction. For example, if a voice notification reads, "There are 10 bags of snacks in stock," the user can correct it with a voice command, saying, "There are 15 bags in stock," and the system will immediately update the information.
[1458] These functions provide a system that improves the efficiency of product management and inventory management, increases sales, responds flexibly to user emotions, and enables more accurate management decisions.
[1459] The processing flow will be explained below.
[1460] Step 1:
[1461] Users take a photo of the product shelves in a store using a device such as a smartphone or dedicated camera, and by pressing the capture button, the image is automatically saved on the device.
[1462] Step 2:
[1463] The device automatically sends the captured image data to a server via Wi-Fi or mobile networks.
[1464] Step 3:
[1465] The server analyzes the image data received from the device and uses an image analysis engine to identify products in the image and determine the location and quantity of each product.
[1466] Step 4:
[1467] The server stores the identified product information (SKU, location, stock quantity, etc.) in a database, allowing the current stock status to be immediately grasped.
[1468] Step 5:
[1469] The server references sales data based on the information stored in the database, extracts past sales data using SQL queries, and analyzes the relationship between product placement and sales.
[1470] Step 6:
[1471] The server uses a generative AI model to analyze product placement and sales data, thereby identifying optimal product placement patterns.
[1472] Step 7:
[1473] The server generates optimal product placement proposals based on the analysis results, which are then visualized.
[1474] Step 8:
[1475] The server sends the generated placement proposal to the terminal, and the user can check the proposal displayed on the terminal.
[1476] Step 9:
[1477] The device will read the notification content aloud to the user using text-to-speech (TTS) technology.
[1478] Step 10:
[1479] The device recognizes the user's emotional state using an emotion engine that analyzes the user's voice and facial expressions in real time.
[1480] Step 11:
[1481] Based on the results of the emotion engine, the server dynamically adjusts the notification content, for example, if the user is feeling stressed, it will provide concise, positive suggestions.
[1482] Step 12:
[1483] The user can correct the contents of the voice message by entering a voice command, for example, "There are 15 bags in stock."
[1484] Step 13:
[1485] The server receives the user's voice command and updates the database with the corrected information, which is then recalculated and saved as the latest information.
[1486] Step 14:
[1487] The server generates management advice for ordering and stock disposal based on the results of periodic inventory. This advice is generated based on the amount of stock and sales trends.
[1488] Step 15:
[1489] The server then sends the generated management advice to the terminal, and the user makes decisions about ordering and inventory disposal based on this advice.
[1490] Through these steps, this system will improve the efficiency of product management and inventory management, increase sales, improve the accuracy of management decisions, and even realize flexible responses that take user emotions into consideration. As a specific example, when the stock of a specific product is low, the system will notify the user by saying, "Stock is low. Please consider placing an order," but if the user is feeling emotionally stressed, it will adjust the wording to a softer one, saying, "There is no need to take immediate action. Please check when you have time."
[1491] Example 2
[1492] Next, a description will be given of Example 2. In the following description, the data processing device 12 will be referred to as a "server" and the robot 414 will be referred to as a "terminal."
[1493] In the retail industry, it is extremely important to accurately analyze the impact that product placement has on sales and propose optimal placements. However, currently, there are limited methods for optimizing shelf placement and proposal systems that take users' emotional states into account. This has prevented efficient inventory management and increased sales from being fully realized. Furthermore, when incorrect product information is identified, correcting it is cumbersome, requiring quick and appropriate management decisions. Given this background, a system is needed that can automatically analyze the relationship between product placement and sales and respond flexibly while taking emotions into consideration.
[1494] The specific processing by the specific processing unit 290 of the data processing device 12 in the second embodiment is realized by the following means.
[1495] In this invention, the server includes means for acquiring image data from product shelves in a store, means for analyzing the acquired image data to identify the product locations and quantities, means for storing the identified information in a database, means for using a generative model to analyze the relationship between product placement and sales by referencing sales data based on the information stored in the database, means for proposing optimal product placement based on the analysis results, means for visualizing the proposal and notifying the terminal, means for reading the notification content aloud, and means for recognizing the user's emotional state from the terminal and dynamically adjusting the notification content based on the user's emotional state. This enables accurate analysis of the relationship between product placement and sales and proposals for optimal placement that take emotions into consideration. It also makes it easier to correct erroneous judgments and generate management advice, enabling quick and accurate management decisions.
[1496] 1. "Image data" refers to digital information, such as still images or videos, that record the state of product shelves in a store.
[1497] 2. "Analysis" refers to the computational process used to identify the location and quantity of products from the acquired image data.
[1498] 3. "Database" means a system for systematically storing and managing specific product information and sales data.
[1499] 4. "Sales Data" means numerical information on the sales of products within a specific period of time.
[1500] 5. A "generative model" is a machine learning algorithm that learns from large amounts of data to derive trends and patterns.
[1501] 6. "Visualization" refers to displaying analysis results and placement proposals as graphs and charts so that users can intuitively understand them.
[1502] 7. "Terminal" means a computer device used by a user to operate the device, in particular a smartphone or tablet device.
[1503] 8. "Notification Content" means the output of information or suggestions provided by the System to the User in the form of text or audio.
[1504] 9. "Emotional state" is state information that represents the user's emotional response and is identified from the user's voice and facial expression.
[1505] 10. "Dynamic adjustment" means changing your response to the situation in real time.
[1506] This invention is a system that acquires image data of product shelves in a store, analyzes the image data to identify the location and quantity of products, and then analyzes sales data based on this to propose optimal product placement.The system uses image recognition technology, generative AI, and an emotion engine, and is able to flexibly respond by taking into account the user's emotional state.
[1507] Equipment and software configuration
[1508] The system consists of the following main components:
[1509] Terminal: A device equipped with a camera for taking pictures of product shelves in a store. Examples include smartphones and dedicated cameras.
[1510] Server: The central data processing unit that receives and analyzes image data. It runs the generative AI model and emotion engine.
[1511] User: The person who operates the system and makes optimal product placement and management decisions.
[1512] Processing steps
[1513] 1. Image data collection
[1514] A user uses a device to take an image of a product shelf in a store. For example, they can use a smartphone camera app to take a picture, focusing on a specific area. The captured image is set to be automatically sent to the server. A dedicated app encodes the image data and sends it to the server using a secure communication protocol (e.g., HTTPS).
[1515] 2. Analysis of image data
[1516] The server analyzes the image data received from the device. Specifically, it uses software such as OpenCV and TensorFlow to identify the location and quantity of products. For example, it uses the YOLO (You Only Look Once) object detection algorithm. The analysis results are stored in a database.
[1517] 3. Analysis of the relationship between placement and sales
[1518] The server references product information and sales data stored in a database and analyzes the relationship between product placement and sales performance using a generative AI model (e.g., H2O.ai's AutoML function). This analysis identifies the optimal placement pattern for a specific product.
[1519] 4. Proposal for optimal layout
[1520] Based on the analysis results, the server generates optimal product placement proposals, which are visualized and displayed in an intuitive format using Tableau or Power BI, and the proposals are sent to the device and notified to the user.
[1521] 5. Introducing the Emotion Engine
[1522] The device reads product information and stock availability to the user aloud, while simultaneously recognizing the user's emotional state using the Microsoft Azure Emotion API and IBM Watson Tone Analyzer. Based on the user's emotional state, the device dynamically adjusts the content of notifications.
[1523] 6. Correcting misjudgments and providing management advice
[1524] The user can make corrections to the information read out loud using voice commands. For example, they can input a correction such as "There are 15 bags in stock." The server reflects this correction in the database and performs recalculation. In addition, based on the results of periodic inventory, the server generates management advice for ordering and stock disposal and sends it to the terminal.
[1525] Specific examples
[1526] For example, if snacks are placed on the bottom shelf, analysis may reveal poor sales data. Based on this, the server suggests placing snacks at eye level. The suggestions are sent to the device as a concrete visual map, allowing the user to adjust product placement. At the same time, an emotion engine analyzes the user's emotional state, providing detailed suggestions if the user has a positive reaction and simple suggestions if the user has a negative reaction. For example, if a voice notification reads, "There are 10 bags of snacks in stock," the user can correct it with a voice command, saying, "There are 15 bags in stock," and the system will immediately update the information.
[1527] Prompt Sentence Examples
[1528] 1. "What image recognition technology do you use to optimize product placement in your stores?"
[1529] 2. "Please show the results of analyzing sales data using a generative AI model."
[1530] 3. "How does the Emotion Engine analyze the user's emotional state and adjust its response?"
[1531] This system will enable more efficient product and inventory management, increased sales, flexible responses that take user feelings into consideration, and more accurate management decisions.
[1532] The flow of the identification process in the second embodiment will be described with reference to FIG.
[1533] Step 1:
[1534] The user uses a device to take a picture of a specific product shelf in a store. The input is a high-resolution image taken using a smartphone camera app, for example. The output is image data with a clear focus on the captured product. Specifically, the image is taken so that the product's position and quantity are visible, and the image data is saved on the device.
[1535] Step 2:
[1536] The device sends the captured image data to the server. The input is the image data acquired in step 1. The output is the image data securely sent to the server. Specifically, a dedicated application encodes the image data into Base64 format and uploads it to the server in real time using the HTTPS protocol.
[1537] Step 3:
[1538] The server analyzes the image data received from the device. The input is the image data sent in step 2. The output is data on the location and quantity of products identified through image analysis. Specifically, it uses TensorFlow and OpenCV to apply object detection algorithms such as YOLO to extract the bounding boxes and labels of each product in the image.
[1539] Step 4:
[1540] The server stores the analysis results in a database. The input is the product location and quantity data identified in step 3. The output is product information stored in the database. Specifically, it is inserted into a MySQL database using an ORM (Object-Relational Mapping) library such as SQLAlchemy.
[1541] Step 5:
[1542] The server retrieves sales data from the database and uses a generative model to analyze the relationship between product placement and sales. The input is the product information and sales data stored in the database. The output is the optimal product placement pattern. Specifically, it uses H2O.ai's AutoML function to derive placement patterns that are expected to increase sales.
[1543] Step 6:
[1544] The server generates optimal product placement proposals based on the analysis results. The input is the optimal product placement pattern obtained in step 5. The output is a visualized placement proposal. Specifically, it is visualized as intuitive graphs and charts using Tableau or Power BI.
[1545] Step 7:
[1546] The server sends the visualized proposal to the terminal and notifies the user. The input is the placement proposal visualized in step 6. The output is the proposal notified to the user. Specifically, it is visually displayed on the terminal via the Internet.
[1547] Step 8:
[1548] The device reads out product information and stock status aloud, while simultaneously recognizing the user's emotional state. The input is the recommendations sent from the server and the user's real-time voice data. The output is the results of recognizing the user's emotional state and the adapted notification content. Specifically, it uses the Microsoft Azure Emotion API and IBM Watson Tone Analyzer to analyze the user's voice and facial expressions.
[1549] Step 9:
[1550] The server dynamically adjusts the notification content based on the emotional state identified by the emotion engine. The input is the user's emotional state sent from the device. The output is the adjusted notification content. For example, if the user is feeling stressed, more concise and positive suggestions are provided.
[1551] Step 10:
[1552] The user inputs a voice command in response to the voice readout, and the server reflects the correction in the database. The input is the user's voice command and the current contents of the database. The output is updated information in the database that reflects the correction. For example, if the correction is made to "There are 15 bags in stock," the information is immediately updated in the database.
[1553] As a result, this system analyzes the relationship between product placement in a store and sales, and realizes flexible and adaptable placement suggestions and inventory management that take into account the user's emotional state.
[1554] (Application example 2)
[1555] Next, a description will be given of Application Example 2. In the following description, the data processing device 12 will be referred to as a "server" and the robot 414 will be referred to as a "terminal."
[1556] Conventional product placement support systems for the retail industry require a lot of effort to identify product locations and inventory information, and their placement change suggestions do not take into account the user's emotional state, making it difficult to maximize management efficiency and sales. Furthermore, because real-time product placement changes and inventory adjustments are time-consuming, there is a demand for faster management decision-making.
[1557] The identification process by the identification processing unit 290 of the data processing device 12 in Application Example 2 is realized by the following means. In this invention, the server includes means for acquiring image data from product shelves in a store, means for analyzing the acquired image data and identifying the location and quantity of products, means for saving the identified information in a database, means for referencing sales data based on the information saved in the database and analyzing the relationship between product placement and sales, means for proposing an optimal product placement based on the analysis results, means for visualizing the proposal and notifying the user, means for reading the notification content aloud to the user, means for a user to use a smart device to take a real-time photo of the product shelf situation in the store and send the image data to the server, means for analyzing sales data and product information using a generative AI model, means for recognizing the user's emotional state and adjusting the notification content based on the user's emotion, and means for providing real-time information updates and advice in cooperation with a voice assistant. This enables efficient product management and inventory management, maximizing sales, flexible responses based on user emotions, and faster management decisions.
[1558] "Image data" is image information obtained by photographing a product shelf, and is data used to identify the position and quantity of products.
[1559] "In-store shelves" refers to shelves and racks used to display, showcase, and store products in retail stores and brick-and-mortar stores.
[1560] "Means" means a method, technique, or device for achieving a particular purpose.
[1561] "Analysis" refers to the process of analyzing the acquired image data and extracting detailed information such as the location and quantity of the products.
[1562] A "database" is an information system that can efficiently store, manage, and search large amounts of data.
[1563] "Sales data" refers to data containing sales information for a product within a certain period of time, including sales amount, sales quantity, sales date, and the like.
[1564] A "generative AI model" is an algorithm that uses artificial intelligence to analyze data and generate new insights and predictions.
[1565] "Visualization" is the process of presenting data or information in a visually understandable way.
[1566] "User" refers to the person who operates the system and manages product placement and inventory.
[1567] "Notification" is the act of transmitting information from the system to the user.
[1568] A "voice assistant" is a system that interacts with users using voice input to assist them in obtaining information and performing operations.
[1569] "Smart devices" refer to electronic devices equipped with advanced computing power and communication functions, including smartphones and smart glasses.
[1570] "Emotional state" refers to a user's emotional response or state, including happiness, sadness, stress, etc.
[1571] "Real-time" means processing or reacting nearly simultaneously or without delay.
[1572] System Configuration
[1573] This system acquires images of in-store product shelves and optimizes product placement using generative AI models and emotion engines. The system consists of the following main hardware and software components:
[1574] 1. Terminal
[1575] Smart devices include smart glasses and smartphones, which are equipped with cameras and are used to capture images of the product shelves in stores.
[1576] 2. Server
[1577] OpenCV is used for image analysis, which analyzes image data and identifies the location and quantity of products.
[1578] The database used is MySQL, and the analyzed data is saved in the database.
[1579] The generative AI model uses OpenAI's GPT-4 and other models to analyze sales data and product information and generate optimal placement patterns.
[1580] The emotion engine uses Affectiva's SDK to analyze the user's emotional state.
[1581] As a voice assistant, it uses Google Cloud Text-to-Speech to read out the generated notification content aloud.
[1582] 3. Users
[1583] Managers use smart devices to take photos of the shelves and send the data to a server in real time.
[1584] Processing steps
[1585] Image data collection
[1586] The user uses smart glasses or a smartphone to take an image of the store's shelves, focusing on a specific area of the shelves to clearly show the location and quantity of products. The device is set to automatically send the captured image to the server.
[1587] Image data analysis
[1588] The server analyzes the image data received from the device. The image analysis engine uses OpenCV to identify the location, quantity, and specific product name of each product. After the product information is identified, it is saved in a database.
[1589] Saving to a database
[1590] The identified product information is stored in a database along with attributes such as SKU, location, and inventory quantity, making it possible to immediately grasp the current inventory status.
[1591] Analysis of the relationship between placement and sales
[1592] The server references sales data from a database for a set period and uses a generative AI model to analyze the relationship between product placement and sales performance. This analysis identifies the optimal placement pattern for specific products, maximizing sales.
[1593] Proposal for optimal layout
[1594] Based on the analysis results, the server generates an optimal product placement proposal. This proposal is visualized, allowing the user to compare the current product shelf layout with the proposed optimal placement. The proposal is then sent to the terminal and notified to the user.
[1595] Introducing the Emotion Engine
[1596] The device reads product information and inventory status to the user aloud and uses an emotion engine to recognize the user's emotional state by analyzing the user's voice and facial expressions to identify their emotions.
[1597] Adjusting notification content
[1598] Based on the emotional state identified by the emotion engine, the server dynamically adjusts the notification content, for example, providing more concise and positive suggestions if the user is feeling stressed.
[1599] Correcting misjudgments and providing management advice
[1600] The user can make corrections to the voice-read information by voice command. For example, the user can input a correction such as "15 bags in stock." The server reflects this correction in the database and performs recalculation. In addition, based on the results of periodic inventory, the server generates management advice for ordering and stock disposal and sends it to the terminal. The user can make effective management decisions based on this advice.
[1601] Specific examples
[1602] For example, analysis may reveal that sales data for a particular product category (e.g., snacks) are sluggish when placed on the bottom shelf. Based on this, the server suggests placing snacks in front of the register or at eye level. The suggestions are sent to the device as a concrete visual map, and the user adjusts product placement accordingly. The emotion engine also analyzes the user's emotional state, providing detailed suggestions if the user has a positive reaction and simple suggestions if the user has a negative reaction.
[1603] Prompt Sentence Examples
[1604] "Please suggest the optimal product placement in the store. The current placement and sales data are as follows: \[Specific data\]"
[1605] The flow of the specific processing in the application example 2 will be described with reference to FIG.
[1606] Step 1:
[1607] A user uses a smart device (smart glasses or smartphone) to take a picture of the product shelf situation in a store in real time. The captured image data is automatically sent from the device to a server. The input is the image of the product shelf, and the output is the image data sent to the server.
[1608] Step 2:
[1609] The server analyzes the received image data. At this time, an image analysis engine (such as OpenCV) is used to identify the location and quantity of each product. Specifically, object recognition and extraction are performed on the image data. The input is the received image data, and the output is specific information such as the product's location, quantity, and product name.
[1610] Step 3:
[1611] The server stores the identified product information in a database (e.g., MySQL). Here, attributes such as SKU, location, and stock quantity are stored as data. The input is analysis information such as product location and quantity, and the output is product information stored in the database.
[1612] Step 4:
[1613] The server references the information stored in the database and sales data, and uses a generative AI model to analyze the relationship between product placement and sales data. At this time, the generative AI model (such as OpenAI's GPT-4) receives a prompt message: "Please suggest the optimal placement of products in the store. The current placement data and sales data are as follows: \[Specific data\]". The input is the current placement data and sales data, and the output is a proposal for the optimal product placement.
[1614] Step 5:
[1615] The server visualizes the generated optimal placement proposal and sends the results to the terminal. Specifically, a visual map is created. The input is the placement proposal obtained from the generative AI model, and the output is the visualized placement proposal and a notification of it.
[1616] Step 6:
[1617] The device reads out product information and stock availability to the user and analyzes the user's emotional state using an emotion engine (such as Affectiva's SDK). The emotion engine recognizes emotions and sends the results to the server. The input is the user's voice and facial expression, and the output is the identified emotional state.
[1618] Step 7:
[1619] The server dynamically adjusts the notification content based on the emotional state obtained from the emotion engine. If the user is feeling stressed, it provides more concise and positive suggestions, and if the user has a positive reaction, it provides more detailed suggestions. The input is the emotional state obtained from the emotion engine, and the output is the adjusted notification content.
[1620] Step 8:
[1621] The user can make corrections to the voice readout as necessary using voice commands. For example, they can input the corrections, such as "There are 15 bags in stock." The server reflects this correction information in the database and performs recalculation. The input is the user's voice command, and the output is the corrected database information.
[1622] Step 9:
[1623] The server generates management advice for ordering and stock disposal based on the results of periodic inventory and sends it to the terminal. The user makes effective management decisions based on this management advice. The input is the inventory result data, and the output is the generated management advice.
[1624] The specific processing unit 290 transmits the result of the specific processing to the robot 414. In the robot 414, the control unit 46A causes the speaker 240 and the control target 443 to output the result of the specific processing. The microphone 238 acquires voice indicating a user input regarding the result of the specific processing. The control unit 46A transmits voice data indicating the user input acquired by the microphone 238 to the data processing device 12. In the data processing device 12, the specific processing unit 290 acquires the voice data.
[1625] The data generation model 58 is a so-called generative AI (Artificial Intelligence). An example of the data generation model 58 is ChatGPT (Internet Search<URL: https: / / openai.com / blog / chatgpt> ), Gemini (Internet search <url: https: gemini.google.com ?hl="ja">) and other generation AIs. The data generation model 58 is obtained by performing deep learning on a neural network. A prompt including an instruction is input to the data generation model 58, and inference data such as voice data indicating voice, text data indicating text, and image data indicating an image is also input. The data generation model 58 performs inference on the input inference data in accordance with the instruction indicated by the prompt, and outputs the inference result in a data format such as voice data and text data. Here, inference refers to, for example, analysis, classification, prediction, and / or summarization.
[1626] In the above embodiment, an example in which the specific processing is performed by the data processing device 12 has been given, but the technology of the present disclosure is not limited to this, and the specific processing may be performed by the robot 414.
[1627] The emotion identification model 59 as an emotion engine may determine the user's emotion according to a specific mapping. Specifically, the emotion identification model 59 may determine the user's emotion according to an emotion map (see FIG. 9), which is a specific mapping. Similarly, the emotion identification model 59 may determine the robot's emotion, and the identification processing unit 290 may perform identification processing using the robot's emotion.
[1628] FIG. 9 is a diagram illustrating an emotion map 400 on which multiple emotions are mapped. In the emotion map 400, emotions are arranged in concentric circles radiating from the center. Emotions closer to the center of the concentric circles are more primitive. Emotions representing states and actions arising from a state of mind are arranged on the outer edges of the concentric circles. The concept of emotion includes both affect and mental states. Emotions generally generated from reactions occurring in the brain are arranged on the left side of the concentric circles. Emotions generally induced by situational judgment are arranged on the right side of the concentric circles. Emotions generally generated from reactions occurring in the brain and induced by situational judgment are arranged on the upper and lower sides of the concentric circles. Furthermore, the emotion of "pleasure" is arranged on the upper side of the concentric circles, and the emotion of "discomfort" is arranged on the lower side. In this way, in the emotion map 400, multiple emotions are mapped based on the structure by which emotions are generated, and emotions that tend to occur simultaneously are mapped close to each other.
[1629] These emotions are distributed in the 3 o'clock direction on emotion map 400, and typically fluctuate between relief and anxiety. In the right half of emotion map 400, situational awareness dominates over internal sensations, resulting in a sense of calm.
[1630] The inside of emotion map 400 represents what is going on in the mind, and the outside of emotion map 400 represents behavior, so the further you go outside emotion map 400, the more visible the emotions become (the more they are expressed in behavior).
[1631] Human emotions are based on various balances, such as posture and blood sugar levels. When these balances deviate from the ideal, a state of discomfort is indicated, and when they approach the ideal, a state of pleasure is indicated. Emotions can also be created for robots, automobiles, and motorcycles, based on various balances, such as posture and remaining battery life. When these balances deviate from the ideal, a state of discomfort is indicated, and when they approach the ideal, a state of pleasure is indicated. An emotion map can be generated, for example, based on Dr. Mitsuyoshi's emotion map (Research on Voice Emotion Recognition and Emotional Brain Physiological Signal Analysis Systems, Tokushima University, Doctoral Dissertation: https: / / ci.nii.ac.jp / naid / 500000375379). The left half of the emotion map lists emotions belonging to the "reaction" domain, where sensation is dominant. The right half of the emotion map lists emotions belonging to the "situation" domain, where situational awareness is dominant.
[1632] The emotion map defines two emotions that promote learning. One is a negative emotion on the situation side, around the middle of "repentance" or "reflection." In other words, this occurs when the robot experiences negative emotions such as "I never want to feel this way again" or "I don't want to be scolded again." The other is a positive emotion on the response side, around "desire." In other words, this occurs when the robot experiences positive feelings such as "I want more" or "I want to know more."
[1633] The emotion identification model 59 inputs user input into a pre-trained neural network, obtains emotion values indicating each emotion shown in the emotion map 400, and determines the user's emotion. This neural network is pre-trained based on multiple pieces of training data that are combinations of user input and emotion values indicating each emotion shown in the emotion map 400. Furthermore, this neural network is trained so that emotions that are located close to each other have similar values, as in the emotion map 900 shown in FIG. 10. FIG. 10 shows an example in which multiple emotions, "relieved," "calm," and "reassuring," have similar emotion values.
[1634] The system according to the present disclosure has been described above mainly with respect to the functions of the data processing device 12, but the system according to the present disclosure is not necessarily implemented on a server. The system according to the present disclosure may be implemented as a general information processing system. The present disclosure may be implemented, for example, as a software program running on a personal computer or an application running on a smartphone, etc. The method according to the present disclosure may be provided to users in the form of SaaS (Software as a Service).
[1635] In the above embodiment, an example was given in which the specific processing is performed by one computer 22, but the technology of the present disclosure is not limited to this, and the specific processing may be distributed and performed by a plurality of computers including the computer 22. For example, the data generation model 58 may be provided in an external device of the data processing device 12, and data may be generated in the external device in accordance with input data.
[1636] In the above embodiment, an example in which the specific processing program 56 is stored in the storage 32 has been described, but the technology of the present disclosure is not limited to this. For example, the specific processing program 56 may be stored in a portable, computer-readable, non-transitory storage medium such as a USB (Universal Serial Bus) memory. The specific processing program 56 stored in the non-transitory storage medium is installed in the computer 22 of the data processing device 12. The processor 28 executes the specific processing in accordance with the specific processing program 56.
[1637] Alternatively, the specific processing program 56 may be stored in a storage device such as a server connected to the data processing device 12 via the network 54, and the specific processing program 56 may be downloaded and installed on the computer 22 in response to a request from the data processing device 12.
[1638] It is not necessary to store all of the specific processing program 56 in a storage device such as a server connected to the data processing device 12 via the network 54, or to store all of the specific processing program 56 in the storage 32; only a portion of the specific processing program 56 may be stored.
[1639] The hardware resource for executing a specific process can be any of the following processors: An example of a processor is a CPU, which is a general-purpose processor that functions as a hardware resource for executing a specific process by executing software, i.e., a program. Another example of a processor is a dedicated electrical circuit, such as an FPGA (Field-Programmable Gate Array), a PLD (Programmable Logic Device), or an ASIC (Application Specific Integrated Circuit), which is a processor with a circuit configuration designed specifically for executing a specific process. Each processor has built-in or connected memory, and each processor uses the memory to execute the specific process.
[1640] The hardware resource that executes the specific processing may be configured with one of these various processors, or may be configured with a combination of two or more processors of the same or different types (for example, a combination of multiple FPGAs, or a combination of a CPU and an FPGA). Also, the hardware resource that executes the specific processing may be a single processor.
[1641] As an example of a system configured with a single processor, first, one processor is configured by combining one or more CPUs and software, and this processor functions as a hardware resource that executes a specific process. Second, there is a system that uses a processor that realizes the functions of an entire system including multiple hardware resources that execute a specific process on a single IC chip, as typified by SoC (System-on-a-chip). In this way, a specific process is realized using one or more of the above-mentioned various processors as hardware resources.
[1642] Furthermore, the hardware structure of these various processors can be, more specifically, an electric circuit that combines circuit elements such as semiconductor devices. The specific processing described above is merely an example. Therefore, it goes without saying that unnecessary steps may be deleted, new steps may be added, or the processing order may be rearranged, without departing from the spirit of the invention.
[1643] The above-described description and illustrations are a detailed explanation of the parts related to the technology of the present disclosure and are merely an example of the technology of the present disclosure. For example, the above description of the configuration, functions, actions, and effects is an explanation of an example of the configuration, functions, actions, and effects of the parts related to the technology of the present disclosure. Therefore, it goes without saying that unnecessary parts may be deleted, new elements may be added, or replacements may be made to the above-described description and illustrations within the scope of the gist of the technology of the present disclosure. Furthermore, to avoid confusion and facilitate understanding of the parts related to the technology of the present disclosure, the above-described description and illustrations omit explanations of common technical knowledge that do not require particular explanation to enable the implementation of the technology of the present disclosure.
[1644] All publications, patent applications, and technical standards mentioned in this specification are herein incorporated by reference to the same extent as if each individual publication, patent application, or technical standard was specifically and individually indicated to be incorporated by reference.
[1645] The following is further disclosed regarding the above embodiment.
[1646] (Claim 1)
[1647] A means for acquiring image data from product shelves in a store;
[1648] A means for analyzing the acquired image data and identifying the position and quantity of the products;
[1649] means for storing the identified information in a database;
[1650] a means for referencing sales data based on the information stored in the database and analyzing the relationship between product placement and sales;
[1651] means for proposing an optimal product placement based on the analysis results;
[1652] a means for visualizing the content of the proposal and notifying the user;
[1653] The system further includes a means for reading the notification content aloud to the user.
[1654] (Claim 2)
[1655] 2. The system according to claim 1, wherein the user inputs a voice command for the voice-read content, and the system is configured to update the database with the corrections.
[1656] (Claim 3)
[1657] 2. The system according to claim 1, further comprising means for generating management advice for ordering or stock disposal based on the results of the inventory and notifying the user of said management advice.
[1658] "Example 1"
[1659] (Claim 1)
[1660] A means for acquiring image data from product shelves in a store;
[1661] A means for analyzing the acquired image data and identifying the position and quantity of the products;
[1662] means for storing the identified information in a database;
[1663] a means for referencing sales data based on the information stored in the database and analyzing the relationship between product placement and sales;
[1664] A means to suggest optimal product placement using a generative AI model including an image analysis engine;
[1665] a means for visualizing the content of the proposal and notifying the user;
[1666] means for reading the notification content aloud to the user;
[1667] A means for a user to take an image of a product shelf and send it to a server using a terminal;
[1668] A system including:
[1669] (Claim 2)
[1670] 2. The system according to claim 1, wherein the user inputs a voice command for the voice-read content, and the system is configured to update the database with the corrections.
[1671] (Claim 3)
[1672] 2. The system according to claim 1, further comprising means for generating management advice for ordering or stock disposal based on the results of the inventory and notifying the user of said management advice.
[1673] "Application Example 1"
[1674] (Claim 1)
[1675] A means for acquiring image data from product shelves in a store;
[1676] A means for analyzing the acquired image data and identifying the position and quantity of the products;
[1677] means for storing the identified information in a database;
[1678] a means for referencing sales data based on the information stored in the database and analyzing the relationship between product placement and sales;
[1679] means for proposing an optimal product placement based on the analysis results;
[1680] a means for visualizing the content of the proposal and notifying the user;
[1681] means for reading the notification content aloud to the user;
[1682] A means for taking an image of a product shelf using a smartphone;
[1683] A means of providing voice notification of product information and stock status using text-to-speech technology;
[1684] a means for the user to correct the information by voice command;
[1685] A means for generating optimal placement that contributes to maximizing sales using a generative AI model;
[1686] A system including a means for inputting a prompt sentence into the generative AI model to obtain an analysis result.
[1687] (Claim 2)
[1688] 2. The system according to claim 1, wherein the user inputs a voice command for the voice-read content, and the system is configured to update the database with the corrections.
[1689] (Claim 3)
[1690] 2. The system according to claim 1, further comprising means for generating management advice for ordering or stock disposal based on the results of the inventory and notifying the user of said management advice.
[1691] "Example 2: Combining Emotion Engines"
[1692] (Claim 1)
[1693] A means for acquiring image data from product shelves in a store;
[1694] A means for analyzing the acquired image data and identifying the position and quantity of the products;
[1695] means for storing the identified information in a database;
[1696] A means using a generative model to refer to sales data based on the information stored in the database and analyze the relationship between product placement and sales;
[1697] means for proposing an optimal product placement based on the analysis results;
[1698] means for visualizing the content of the proposal and notifying the same to a terminal;
[1699] means for reading out the notification content aloud;
[1700] The system includes a means for recognizing a user's emotional state from a terminal and dynamically adjusting notification content based on the emotional state.
[1701] (Claim 2)
[1702] 2. The system according to claim 1, wherein the user inputs a voice command for the voice-read content, and the system is configured to update the...
Claims
1. A means for acquiring image data from product shelves in a store; A means for analyzing the acquired image data and identifying the position and quantity of the products; means for storing the identified information in a database; a means for referencing sales data based on the information stored in the database and analyzing the relationship between product placement and sales; means for proposing an optimal product placement based on the analysis results; a means for visualizing the content of the proposal and notifying the user; The system further includes a means for reading the notification content aloud to the user.
2. 2. The system according to claim 1, wherein the user inputs a voice command for the voice-read content, and the system is configured to reflect the correction content in the database.
3. 2. The system according to claim 1, further comprising means for generating management advice for ordering or stock disposal based on the results of the inventory and notifying the user of said management advice.
Citation Information
Patent Citations
Persona chatbot control method and system
JP2022180282A