Analysis device, analysis system, analysis method, and program
Patent Information
- Application Number
- JP2023500803
- Authority / Receiving Office
- JP · JP
- Patent Type
- Applications
- Priority Date
- 2022-02-10
- Filing Date
- 2022-02-10
- Publication Date
- 2025-11-12
AI Technical Summary
Current methods lack an effective way to analyze and optimize the display state of products in retail environments, such as supermarkets and convenience stores, which affects product visibility and sales floor management.
An analysis system utilizing a machine learning model, including an imaging device and an analysis device with processors, estimates the placement area and display state of products by segmenting sales floor images into regions and determining product placement, quantity, and arrangement, enabling real-time or batch processing to improve product display efficiency.
The system provides accurate estimation of product placement and display state, allowing for timely replenishment and arrangement, enhancing store operations and customer experience by optimizing product visibility and sales floor management.
Abstract
Description
Analysis device, analysis system, analysis method and program
[0001] The present disclosure relates to an analysis device, an analysis system, an analysis method, and a program.
[0002] The use of information technology is progressing in the retail industry, including supermarkets and convenience stores. For example, information technology is now being used in the display of products in stores.
[0003] JP 2018-10372 A International Publication No. 2018 / 179361 JP 2017-182653 A
[0004] An object of the present disclosure is to provide a novel technique for analyzing the display state of products.
[0005] In order to solve the above problem, one aspect of the present disclosure relates to an analysis device having one or more memories and one or more processors, wherein the one or more processors estimate a placement area for a group of products of the same type and estimate the display state of the group of products in the placement area.
[0006] FIG. 1 is a schematic diagram illustrating an analysis system according to an embodiment of the present disclosure. FIG. 2 is a diagram illustrating region segmentation according to an embodiment of the present disclosure. FIG. 3 is a schematic diagram illustrating analysis processing according to an embodiment of the present disclosure. FIG. 4 is a block diagram illustrating a functional configuration of an analysis device according to an embodiment of the present disclosure. FIG. 5 is a diagram illustrating training data according to an embodiment of the present disclosure. FIG. 6 is a diagram illustrating estimation results according to an embodiment of the present disclosure. FIG. 7 is a flowchart illustrating analysis processing according to an embodiment of the present disclosure. FIG. 8 is a block diagram illustrating a hardware configuration of an analysis device according to an embodiment of the present disclosure.
[0007] Hereinafter, embodiments of the present disclosure will be described with reference to the drawings.
[0008] In the following embodiment, an analysis system is disclosed that captures images of a store's sales floor and uses a machine learning model to estimate the display status of products based on the images of the sales floor.
[0009] [Analysis System] First, an analysis system according to an embodiment of the present disclosure will be described with reference to Fig. 1. Fig. 1 is a schematic diagram showing an analysis system according to an embodiment of the present disclosure.
[0010] 1, the analysis system 10 includes an imaging device 20, a user terminal 30, and an analysis device 100. When an image of a sales floor is acquired from the imaging device 20, the analysis device 100 analyzes the acquired image of the sales floor and notifies the user terminal 30 of the display state of the imaged sales floor, instructions for the display state, instructions related to the display state, etc. The display state may be, for example, the product names, number of products, and product alignment degree for each product type displayed in the layout area of the sales floor.
[0011] The imaging device 20 may be, for example, a video camera installed in a store or the like, which captures an image of the sales floor to be imaged and transmits the sales floor video to the analysis device 100. Typically, the imaging device 20 is installed near the sales floor to be imaged and used to observe the sales floor. The imaging device 20 may be fixed to a specific location in the store, or may be a mobile device attached to a robot or cart. This makes it possible to obtain various information. It also makes it possible to reduce the number of imaging devices 20 installed. Furthermore, multiple imaging devices 20 may be installed. This makes it possible to obtain appropriate sales floor video even when blind spots or the like occur.
[0012] User terminal 30 may be, for example, an information processing device such as a personal computer, tablet, or smartphone provided in a store or the like, and acquires information about the display status of various product groups in the sales floor estimated based on the sales floor video from analysis device 100 or a server or the like that stores the analysis results of analysis device 100. For example, user terminal 30 may be equipped with software related to store management and business improvement, such as various software for assisting store clerks in replenishing, replacing, pricing, etc. of products, and may also be equipped with software that allows viewing of the analysis results of analysis device 100. Store clerks may replenishing, replacing, pricing, etc. of products in the sales floor based on data analyzed by various software using the display status and POS (Point of Sales) data acquired from analysis device 100.
[0013] The analysis device 100 may be, for example, an information processing device such as a personal computer installed in a store or a server installed in a location other than the store, such as a headquarters that manages the store or a cloud server. Based on the sales floor video acquired from the image capture device 20, the analysis device 100 estimates the layout areas of each product type displayed on the sales floor and estimates the display status of the product groups arranged in each layout area. The analysis device 100 may acquire the sales floor video acquired from the image capture device 20, or may acquire data obtained by performing predetermined processing on the sales floor video. In such a case, the sales floor video acquired by the image capture device 20 is output to a predetermined processing device, and the processed data is output to the analysis device 100. This facilitates the transmission of information related to the sales floor video via a network and subsequent processing in the analysis device 100. When multiple image capture devices 20 are installed, one processing device may be provided for each of the multiple image capture devices 20.
[0014] Here, the display state refers to the display state of the product group, and may be, for example, the product names, number of products, and product alignment of the product group displayed in each placement area. The analysis device 100 according to this embodiment may use a machine learning model such as a neural network to estimate placement areas for each product type from the sales floor video and estimate the display state of the same type of product group displayed in each placement area. For example, the analysis device 100 performs region segmentation for each product type on frames of the sales floor video such as that shown in FIG. 1 and acquires region-segmented frames such as those shown in FIG. 2. Here, the analysis device 100 may process the sales floor video acquired from the imaging device 20 in real time or in batches.
[0015] In an embodiment described below, as shown in FIG. 3 , the analysis device 100 may input a sales floor video into a machine learning model, divide the sales floor video into regions, and obtain a product area map showing the placement area for each product type, a product center heat map showing the center of each product placed on the sales floor, and a product orientation heat map showing the orientation of each product placed on the sales floor. After obtaining the product area map, product center heat map, and product orientation heat map from the machine learning model, the analysis device 100 can superimpose these data on the sales floor video to check the display status of each product. Note that, in this disclosure, the "center" does not refer to the strict center. Furthermore, the center can be calculated using various methods.
[0016] According to the present disclosure, it is possible to estimate the display status of not only products stored in standard storage containers such as boxes, but also products such as vegetables and fruits stored in irregular bags or wrapping.
[0017] [Analysis Device] Next, an analysis device 100 according to an embodiment of the present disclosure will be described with reference to Fig. 4. Fig. 4 is a block diagram showing the functional configuration of the analysis device 100 according to an embodiment of the present disclosure.
[0018] 4, analysis device 100 has area estimation unit 110 and display state estimation unit 120. Area estimation unit 110 and display state estimation unit 120 are installed in analysis device 100 and are realized by one or more processors executing one or more programs stored in one or more memories.
[0019] The area estimation unit 110 estimates the layout area of a group of products that are to be classified as the same (referred to as a group of products of the same type in this specification) from the sales floor video. Specifically, when the sales floor video is acquired from the imaging device 20, the area estimation unit 110 performs area division for each product type on the frames of the acquired sales floor video, and estimates the layout area for each product type. In a typical sales floor, various products are arranged together on display shelves by product type.
[0020] For example, suppose that in a vegetable section, multiple bags of onions produced in region A (product type 1), multiple bags of onions produced in region B (product type 2), multiple bags of potatoes produced in region C (product type 3), and multiple bags of potatoes produced in region D (product type 4) are grouped by product type and arranged on shelves. When a video of the vegetable section is acquired, the area estimation unit 110 performs preprocessing such as removing moving objects such as people and shopping carts included in the video of the vegetable section and extracting areas of interest, and estimates, in frames of the preprocessed video of the sales section, the layout area where the product group of product type 1 is displayed, the layout area where the product group of product type 2 is displayed, the layout area where the product group of product type 3 is displayed, and the layout area where the product group of product type 4 is displayed.
[0021] In one embodiment, the region estimation unit 110 may use a trained machine learning model to perform region segmentation based on product type on frames of a sales floor video and estimate the placement region for each product type. In other words, the machine learning model may be trained to, when a frame of a sales floor video is input, segment the frame into regions and output a product region map indicating the placement region for each product type. For example, the region estimation unit 110 may input a frame of the sales floor video of the vegetable section shown in FIG. 3 into the trained machine learning model to obtain a product region map, and then superimpose the obtained product region map on the input frame to generate a frame segmented into regions for each product type.
[0022] Here, the machine learning model for area estimation may be realized as, for example, a neural network, and may be trained by supervised learning using pairs of frames of a sales floor video as shown in Fig. 1 and frames annotated with information on the placement area for each product type as shown in Fig. 2 as training data. Specifically, the machine learning model may be an instance segmentation model such as Mask-RCNN (Regional Convolutional Neural Network), and may be trained to predict the bounding box of a detection target and the corresponding segmentation mask for multiple products or product types in a frame.
[0023] Alternatively, the machine learning model may be a convolutional neural network, trained to segment areas by clustering feature vectors in a feature map. That is, areas with similar feature vectors can be considered areas where the same type of product is displayed. Such a convolutional neural network may be trained by tuning a convolutional neural network pre-trained on another large-scale image dataset, such as Imagenet, or it may be trained to assign tentative labels to product areas and predict the label numbers.
[0024] When region segmentation is performed for each frame, there is a possibility that recognition fluctuations may occur, such as one region in one frame being split into two in another frame. For this reason, the region estimation unit 110 may smooth the frames in the time direction to prevent sudden changes from occurring compared to past prediction results. Note that batch processing can utilize not only past prediction results but also future prediction results. On the other hand, when products are replenished or replaced, recognizing a sudden change may be correct. For this reason, the region estimation unit 110 may allow a sudden change without smoothing when the inter-frame difference in the sales floor video is greater than a predetermined threshold.
[0025] The display state estimation unit 120 estimates the display state of the product group in the arrangement area. For example, the display state estimation unit 120 may estimate one or more of the product names, the number of products, and the product alignment of the product group in the arrangement area. Specifically, the display state estimation unit 120 estimates the display state of the product group, such as the product names, the number of products, and the product alignment of the product group, for the product group included in the arrangement area for each product type estimated by the area estimation unit 110, using a trained machine learning model. For example, when the display state estimation unit 120 detects that the number of products in the arrangement area in which the image is captured is low or the product alignment is low, it can identify the product names of the arrangement area and notify a store clerk to replenish the products or rearrange the display of the products. In one embodiment, the display state estimation unit 120 may estimate at least one of the product names, the number of products, and the product alignment of the products in the arrangement area as the display state using a trained machine learning model. When a frame of a sales floor video is input, the machine learning model is trained to output the product name, center position, and / or orientation of a product included in the frame. For example, the product name may be indicated by product identification information, such as a product number pre-assigned to the product name. The center position of each product may be indicated by a symbol (e.g., a circle) indicating the center of each product in the frame, or by a product center heat map such as that shown in FIG. 3 . The orientation of each product may be indicated by a symbol (e.g., a line) indicating the orientation of each product in the frame, or by a product orientation heat map such as that shown in FIG. 3 . When a frame of a sales floor video is input, the machine learning model is trained to output at least one of the product name, product center, and product orientation of the product captured in the frame. Such a machine learning model may be implemented, for example, as a neural network, or may be trained by supervised learning using pairs of a frame of the sales floor video and a frame annotated with information about the product name, center, and / or orientation of each product for each product type in the frame as training data.
[0026] For example, when the display state estimation unit 120 estimates the product names of products displayed in a display area using a trained machine learning model, the machine learning model may identify product identification information, such as product numbers, of the products captured in the input frames. That is, the machine learning model may be implemented as a neural network and trained using supervised learning, using pairs of frames of sales floor video and frames annotated with information about the product identification information of each product in the frames as training data. Once the trained machine learning model is acquired, the display state estimation unit 120 can use the machine learning model to estimate the product names of each product displayed in the frames of the sales floor video. Here, the input frames may be frames segmented by the region estimation unit 110, or may not be segments.
[0027] Alternatively, the machine learning model may be a neural network that determines the feature quantities of products for each product type from frames of the sales floor video. The display state estimation unit 120 may estimate the feature quantities of each product arranged in the frame using the machine learning model, and then identify the product name corresponding to the estimated feature quantities as the product.
[0028] If a product does not fall into any of the existing product categories, the product may be determined as unknown. Furthermore, if external information such as store layout information or POS data is available, that information may be used for estimation. For example, it is possible to narrow down the products to be placed in the sales area being analyzed based on the external information, and obtain a machine learning model for each product category (e.g., vegetables, sweets, etc.) suitable for the products in the sales area being analyzed (e.g., vegetable section, sweets section, etc.), thereby improving estimation accuracy.
[0029] Next, when the display state estimation unit 120 estimates the number of products in a product group in a placement area using a trained machine learning model, the machine learning model may, for example, identify a point within the area of the product captured in the input frame, such as the center of the product. That is, the machine learning model may be implemented as a neural network and trained using supervised learning using pairs of frames of sales floor footage and frames annotated with annotations indicating the center of each product within the frame as training data. For example, FIG. 5 shows an example of annotated frames with the center of each product. In the annotated frames shown, each product is annotated with a circle indicating the center of the packaged product.
[0030] After acquiring the trained machine learning model, the display state estimation unit 120 can use the machine learning model to estimate the center of each product displayed within a frame of the sales floor video. Then, by referencing the frames segmented by the region estimation unit 110, the display state estimation unit 120 can estimate the number of products displayed in each placement area based on the number of estimated center points within each placement area. For example, the display state estimation unit 120 can use a machine learning model that identifies product names and a machine learning model that estimates product center points to generate information indicating the product names and centers of each product for a group of products displayed in each placement area of the segmented frame, such as that shown in FIG. 6 . The display state estimation unit 120 can estimate the product names and the number of products for each product type by counting the number of points (preferably center points) included in each placement area based on the frame. The number of products may also be estimated based on exposed areas of the placement area where products are displayed (e.g., shelves, the bottom of product placement areas on fixtures, etc.), i.e., areas where products are missing.
[0031] If the estimated number of products for a certain placement area estimated in this manner is less than a set value, the display state determination unit may determine that the placement area needs to be replenished with products, and the display state instruction unit may instruct a store clerk or the like to replenish the placement area with products.
[0032] Note that the estimation of the number of products according to the present disclosure is not limited to this. Instead of using dots representing products, a machine learning model may be used that detects bounding boxes indicating the position of each product within the frame. In this case, the display state estimation unit 120 may estimate the number of products by counting the number of bounding boxes included in each placement area. Alternatively, the estimation may be performed using product density. For example, the product point heat map may be regarded as product density, and the number of products may be estimated by integrating a product-centered heat map for each placement area. Alternatively, the display state estimation unit 120 may estimate the number of products in each placement area of the frame using a machine learning model trained to regress the number of products from the feature quantities of the placement area. If the machine learning model is properly trained, the above-described product number estimation based on the regression of product density and product number may also be able to predict the number of hidden products not captured in the frame.
[0033] Next, when the display state estimation unit 120 estimates the degree of alignment of a group of products in a placement area using a trained machine learning model, the machine learning model may, for example, identify the orientation of the products captured in the input frame. That is, the machine learning model may be implemented as, for example, a neural network and trained by supervised learning using pairs of frames of sales floor footage and frames annotated with information about the orientation of each product within the frame as training data. For example, FIG. 5 shows an example of annotated frames annotated with information about the orientation of each product. In the annotated frames shown, each product is annotated with a line indicating the orientation of the packaged product.
[0034] Using the machine learning model trained in this manner, the display state estimation unit 120 estimates the orientation of each product displayed within a frame of the sales floor video, and by referencing the frames divided into regions by the region estimation unit 110, estimates the degree of alignment of the products displayed in each placement region based on the degree of alignment of the estimated orientations within each placement region. For example, the display state estimation unit 120 may use a machine learning model for estimating product centers and a machine learning model for estimating product orientations in combination to first predict the centers of each product within a placement region, predict the orientations of the products relative to each center, and determine the degree of variation in the orientations of each product included within the placement region. Specifically, the display state estimation unit 120 may define the maximum difference in product orientations among products within a placement region as the product alignment degree. If the maximum difference in product orientations for a certain placement region is equal to or greater than a predetermined threshold, the display state determination unit determines that the placement region requires alignment, and the display state instruction unit may instruct a store clerk, etc., via a user terminal, etc., to align the products displayed in the placement region.
[0035] Note that the estimation of the product alignment degree according to the present disclosure is not limited to this and may utilize features. For example, the display state estimation unit 120 may first predict the center of each product within the placement area and determine the degree of variation in the feature vectors of the feature map for each product center within the placement area. Variation in the feature vectors of each product indicates local variations in appearance and misalignment of the product orientation. Alternatively, the display state estimation unit 120 may estimate variations for multiple products placed nearby, such as two products, and aggregate the variations across the entire placement area. Specifically, the display state estimation unit 120 may search for k products nearby each product within the placement area and evaluate the degree of similarity between the orientation of the product and each nearby product using, for example, the dot product of the feature vectors. The display state estimation unit 120 may then determine the product alignment degree for the placement area by calculating the degree of similarity for all pairs of nearby products within the placement area and averaging the results. Alternatively, the display state estimation unit 120 may use a machine learning model trained to calculate the features of the placement area, and use the features calculated for the placement area within the frame as the degree of product alignment.
[0036] [Analysis Processing] Next, an analysis processing according to an embodiment of the present disclosure will be described with reference to Fig. 7. The analysis processing is executed by the above-described analysis device 100, and can be realized, for example, by one or more processors executing a program stored in one or more memories of the analysis device 100. Fig. 7 is a flowchart illustrating the analysis processing according to an embodiment of the present disclosure.
[0037] 7, in step S101, analysis device 100 acquires a sales floor video. Specifically, analysis device 100 acquires the sales floor video from imaging device 20 installed in the sales floor. Here, analysis device 100 may execute the following steps on the acquired sales floor video in real time, or may temporarily store the acquired sales floor video and execute the following steps on the stored sales floor video at an appropriate time. In addition to receiving the sales floor video, analysis device 100 may also acquire data obtained by processing the sales floor video (i.e., data based on the sales floor video).
[0038] In step S102, analysis device 100 estimates the layout areas of a group of products of the same type from the video of the sales floor. For example, analysis device 100 may divide frames of the video of the sales floor into regions and use a machine learning model such as a neural network that has been trained in advance to estimate the layout areas of each type of product displayed in the sales floor, thereby estimating the layout areas of the group of products displayed in the sales floor from the frames of the video of the sales floor.
[0039] In step S103, the analysis device 100 estimates the display state of the product group in the placement area. Specifically, the analysis device 100 uses a machine learning model such as a neural network that has been pre-trained to estimate one or more of the product names, the number of products, and the product alignment of the product group in the placement area for each product type from frames of the sales floor video, and estimates one or more of the product names, the number of products, and the product alignment for each product type displayed in the sales floor from frames of the sales floor video. For example, the analysis device 100 may use a machine learning model such as a neural network that has been trained to estimate the center and / or orientation of each product in each placement area divided from the frames, calculate the number of products based on the number of product centers in the placement area, and determine the product alignment based on the degree of consistency of the orientation of each product in the placement area.
[0040] In step S104, the analysis device 100 determines whether the display state satisfies a predetermined condition. Specifically, the analysis device 100 determines whether the display state requires maintenance by a store clerk or the like. For example, the analysis device 100 may detect whether there is an arrangement area where the estimated number of products is less than a predetermined value. Alternatively, the analysis device 100 may detect whether the estimated product alignment is less than a predetermined value. Alternatively, the analysis device 100 may detect whether the area of the product area is less than a set value. If an arrangement area where the estimated number of products is less than the predetermined value and satisfies the predetermined condition is detected (S104: YES), in step S105, the analysis device 100 notifies a store clerk or the like via a user terminal or the like to replenish or organize products in the detected arrangement area. On the other hand, if an arrangement area that satisfies the predetermined condition is not detected (S104: NO), the analysis device 100 returns to step S101 and repeats the above steps.
[0041] Furthermore, analysis device 100 may store or output to an external storage device a log relating to the satisfaction of a predetermined condition. The log may be used for store management operations, business improvement, etc. Furthermore, for example, such a log may be stored in association with a video of a sales floor that has been determined to satisfy the predetermined condition. This allows the log to be easily used for store management operations, business improvement, etc.
[0042] [Hardware Configuration] Part or all of the analysis device 100 in the above-described embodiments may be configured as hardware, or may be configured as software (program) information processing executed by a CPU (Central Processing Unit) or a GPU (Graphics Processing Unit). When configured as software information processing, software that realizes at least some of the functions of each device in the above-described embodiments may be stored on a non-transitory storage medium (non-transitory computer-readable medium) such as a flexible disk, CD-ROM (Compact Disc-Read Only Memory), or USB (Universal Serial Bus) memory, and the software information processing may be executed by loading the software into a computer. Alternatively, the software may be downloaded via a communication network. Furthermore, the software may be implemented in a circuit such as an ASIC (Application Specific Integrated Circuit) or FPGA (Field Programmable Gate Array), thereby executing the information processing by hardware.
[0043] The type of storage medium that stores the software is not limited. The storage medium is not limited to removable media such as magnetic disks or optical disks, but may be fixed storage media such as hard disks or memory. The storage medium may be provided inside the computer or outside the computer.
[0044] 8 is a block diagram showing an example of the hardware configuration of analysis device 100 in the above-described embodiment. Analysis device 100 may be realized as a computer 7 including, for example, a processor 71, a main storage device 72 (memory), an auxiliary storage device 73 (memory), a network interface 74, and a device interface 75, which are connected via a bus 76.
[0045] Although the computer 7 in FIG. 8 includes one of each component, it may also include multiple of the same component. Also, while FIG. 8 shows one computer 7, the software may be installed on multiple computers, with each of the multiple computers executing the same or different parts of the software. In this case, a distributed computing configuration may be used in which each computer communicates via a network interface 74 or the like to execute processing. In other words, the analysis device 100 in the above-described embodiment may be configured as a system in which one or more computers execute instructions stored in one or more storage devices to achieve its functions. Furthermore, it may be configured such that information transmitted from a terminal is processed by one or more computers provided on a cloud, and the processing results are transmitted to the terminal.
[0046] The various calculations of the analysis device 100 in the above-described embodiment may be executed in parallel using one or more processors, or using multiple computers via a network. Furthermore, the various calculations may be distributed to multiple processing cores within a processor and executed in parallel. Furthermore, some or all of the processes, means, etc. disclosed herein may be executed by at least one of a processor and a storage device provided on a cloud that can communicate with the computer 7 via a network. Thus, the analysis device 100 in the above-described embodiment may be implemented in the form of parallel computing using one or more computers.
[0047] The processor 71 may be an electronic circuit (processing circuit, processing circuitry, CPU, GPU, FPGA, ASIC, or the like) including a computer control device and arithmetic device. The processor 71 may also be a semiconductor device or the like including a dedicated processing circuit. The processor 71 is not limited to an electronic circuit using electronic logic elements, but may also be realized by an optical circuit using optical logic elements. The processor 71 may also include an arithmetic function based on quantum computing.
[0048] The processor 71 performs arithmetic processing based on data and software (programs) input from each device, etc., configured internally of the computer 7, and can output the arithmetic results and control signals to each device, etc. The processor 71 may control each component constituting the computer 7 by executing the OS (Operating System) of the computer 7, applications, etc.
[0049] The analysis device 100 in the above-described embodiment may be realized by one or more processors 71. Here, the processor 71 may refer to one or more electronic circuits arranged on one chip, or may refer to one or more electronic circuits arranged on two or more chips or two or more devices. When multiple electronic circuits are used, the respective electronic circuits may communicate with each other via wire or wirelessly.
[0050] The main memory device 72 is a memory device that stores instructions executed by the processor 71 and various data, and information stored in the main memory device 72 is read by the processor 71. The auxiliary memory device 73 is a memory device other than the main memory device 72. Note that these memory devices refer to any electronic component capable of storing electronic information, and may be semiconductor memory. The semiconductor memory may be either volatile memory or non-volatile memory. The memory device for saving various data in the analysis device 100 in the above-described embodiment may be realized by the main memory device 72 or the auxiliary memory device 73, or may be realized by an internal memory built into the processor 71. For example, the memory unit 72 in the above-described embodiment may be realized by the main memory device 72 or the auxiliary memory device 73.
[0051] Multiple processors may be connected (coupled) to one storage device (memory), or a single processor may be connected. Multiple storage devices (memories) may be connected (coupled) to one processor. When the analysis device 100 in the above-described embodiment is configured with at least one storage device (memory) and multiple processors connected (coupled) to this at least one storage device (memory), it may include a configuration in which at least one of the multiple processors is connected (coupled) to at least one storage device (memory). This configuration may also be realized by storage devices (memories) and processors included in multiple computers. Furthermore, it may include a configuration in which the storage device (memory) is integrated with the processor (for example, a cache memory including an L1 cache and an L2 cache).
[0052] The network interface 74 is an interface for connecting to the communication network 8 wirelessly or via a wire. The network interface 74 may be an appropriate interface, such as one that conforms to an existing communication standard. The network interface 74 may exchange information with an external device 9A connected via the communication network 8. The communication network 8 may be any one of a wide area network (WAN), a local area network (LAN), a personal area network (PAN), etc., or a combination thereof, as long as information is exchanged between the computer 7 and the external device 9A. An example of a WAN is the Internet, an example of a LAN is IEEE 802.11 or Ethernet (registered trademark), and an example of a PAN is Bluetooth (registered trademark) or NFC (Near Field Communication), etc.
[0053] The device interface 75 is an interface such as a USB that directly connects to the external device 9B.
[0054] The external device 9A is a device connected to the computer 7 via a network, and the external device 9B is a device connected directly to the computer 7.
[0055] The external device 9A or the external device 9B may be, for example, an input device. The input device is, for example, a device such as a camera, a microphone, a motion capture device, various sensors, a keyboard, a mouse, or a touch panel, and provides acquired information to the computer 7. Alternatively, the external device 9A or the external device 9B may be a device including an input unit, a memory, and a processor, such as a personal computer, a tablet terminal, or a smartphone.
[0056] Furthermore, the external device 9A or the external device 9B may be, for example, an output device. The output device may be, for example, a display device such as an LCD (Liquid Crystal Display), a CRT (Cathode Ray Tube), a PDP (Plasma Display Panel), or an organic EL (Electro Luminescence) panel, or may be a speaker that outputs sound or the like. Alternatively, the output device may be a device including an output unit, a memory, and a processor, such as a personal computer, a tablet terminal, or a smartphone.
[0057] The external device 9A or the external device 9B may be a storage device (memory). For example, the external device 9A may be a network storage or the like, and the external device 9B may be a storage device such as an HDD.
[0058] Furthermore, the external device 9A or the external device 9B may be a device having some of the functions of the components of each device (the server 100 or the terminal 200) in the above-described embodiment. In other words, the computer 7 may transmit or receive some or all of the processing results of the external device 9A or the external device 9B.
[0059] In this specification (including the claims), when the expression "at least one of a, b, and c" or "at least one of a, b, or c" (including similar expressions) is used, it includes any of a, b, c, ab, ac, bc, or abc. It may also include multiple instances of any element, such as aa, abb, aabbcc, etc. Furthermore, it also includes the addition of elements other than the enumerated elements (a, b, and c), such as having d, as in abcd.
[0060] In this specification (including the claims), when expressions such as "using data as input / based on / according to / in response to" (including similar expressions) are used, unless otherwise specified, this includes cases where various data itself is used as input, or where various data that has been processed in some way (e.g., noise-added, normalized, intermediate representation of various data, etc.) is used as input. Furthermore, when it is stated that a result is obtained "based on / according to / in response to data," this includes cases where the result is obtained based solely on the data in question, as well as cases where the result is obtained in response to other data, factors, conditions, and / or states other than the data in question. Furthermore, when it is stated that "data is output," unless otherwise specified, this includes cases where various data itself is used as output, or where various data that has been processed in some way (e.g., noise-added, normalized, intermediate representation of various data, etc.) is output.
[0061] When the terms "connected" and "coupled" are used in this specification (including the claims), they are intended as open-ended terms that include any of direct connection / coupling, indirect connection / coupling, electrically connection / coupling, communicatively connection / coupling, functionally connection / coupling, and physically connection / coupling. These terms should be interpreted appropriately according to the context in which they are used, but any connection / coupling form that is not intentionally or naturally excluded should be interpreted as being included in these terms without limitation.
[0062] In this specification (including the claims), when the expression "A configured to B" is used, it may include the physical structure of element A having a configuration capable of performing operation B, and the permanent or temporary setting / configuration of element A being configured / set to actually perform operation B. For example, if element A is a general-purpose processor, it is sufficient that the processor has a hardware configuration capable of performing operation B, and is configured to actually perform operation B by setting a permanent or temporary program (instruction). Also, if element A is a dedicated processor or dedicated arithmetic circuit, it is sufficient that the circuit structure of the processor is implemented to actually perform operation B, regardless of whether control instructions and data are actually attached.
[0063] When used in this specification (including the claims), terms implying containing or possessing (e.g., "comprising / including" and "having") are intended to be open-ended terms that include containing or possessing things other than the object designated by the object of the term. When the object of such terms implies no quantity or a singular number (e.g., expressions using the articles "a" or "an"), the expression should be construed as not being limited to a specific number.
[0064] In this specification (including the claims), even if expressions such as "one or more" or "at least one" are used in some places and expressions that do not specify a quantity or that imply a singular number (expressions using the articles "a" or "an") are used in other places, the latter expressions are not intended to mean "one." In general, expressions that do not specify a quantity or that imply a singular number (expressions using the articles "a" or "an") should be interpreted as not necessarily being limited to a specific number.
[0065] In this specification, when a particular advantage / result is described as being obtained with respect to a particular configuration of a certain embodiment, it should be understood that the same advantage / result can also be obtained with one or more other embodiments having the same configuration, unless otherwise stated. However, it should be understood that the presence or absence of the effect generally depends on various factors, conditions, and / or states, etc., and that the effect is not necessarily obtained with the configuration. The effect is merely obtained by the configuration described in the embodiment when various factors, conditions, and / or states, etc. are satisfied, and the effect does not necessarily occur in a claimed invention that defines the same or a similar configuration.
[0066] When used in this specification (including the claims), terms such as "maximize" include finding a global maximum, finding an approximation of a global maximum, finding a local maximum, and finding an approximation of a local maximum, and should be interpreted appropriately according to the context in which the term is used. It also includes finding approximations of these maxima probabilistically or heuristically. Similarly, when used in this specification (including the claims), terms such as "minimize" include finding a global minimum, finding an approximation of a global minimum, finding a local minimum, and finding an approximation of a local minimum, and should be interpreted appropriately according to the context in which the term is used. It also includes finding approximations of these minima probabilistically or heuristically. Similarly, when used in this specification (including the claims), terms such as "optimize" include finding a global optimum, finding an approximation of a global optimum, finding a local optimum, and finding an approximation of a local optimum, and should be interpreted appropriately according to the context in which the term is used. It also includes finding approximations of these optima probabilistically or heuristically.
[0067] In this specification (including claims), when multiple pieces of hardware perform a predetermined process, the pieces of hardware may cooperate to perform the predetermined process, or some of the hardware may perform all of the predetermined process. Furthermore, some of the hardware may perform part of the predetermined process, and other hardware may perform the rest of the predetermined process. In this specification (including claims), when an expression such as "one or more pieces of hardware perform a first process, and the one or more pieces of hardware perform a second process" is used, the hardware performing the first process and the hardware performing the second process may be the same or different. In other words, it is sufficient that the hardware performing the first process and the hardware performing the second process are included in the one or more pieces of hardware. Note that hardware may include an electronic circuit, a device including an electronic circuit, etc.
[0068] In this specification (including the claims), when multiple storage devices (memories) store data, each of the multiple storage devices (memories) may store only a portion of the data, or may store the entire data.
[0069] Although the embodiments of the present disclosure have been described in detail above, the present disclosure is not limited to the individual embodiments described above. Various additions, modifications, substitutions, partial deletions, etc. are possible within the scope of the conceptual idea and spirit of the present invention derived from the content defined in the claims and their equivalents. For example, in all of the above-described embodiments, when numerical values or formulas are used in the explanation, they are shown as examples and are not limited to these. Furthermore, the order of each operation in the embodiments is shown as an example and is not limited to these.
[0070] This application claims priority based on Japanese Patent Application No. 2021-023664, filed on February 17, 2021, the entire contents of which are incorporated herein by reference.
[0071] REFERENCE SIGNS LIST 10 Analysis system 20 Imaging device 30 User terminal 100 Analysis device 110 Area estimation unit 120 Display state estimation unit
Claims
1. one or more memories; one or more processors; and The one or more processors: Estimating a location area of a group of products with the same product name in an image of the sales floor based on the image of the sales floor; notifying a user terminal of information regarding the estimated display state of the product group in the placement area, the information regarding the display state including the product names of the product group in the placement area; Analysis device.
2. The analysis device described in claim 1, wherein the information causes the user terminal to display the placement area and the product name in a display format that associates the placement area and the product name with each other.
3. An analysis device as described in claim 2, wherein the display mode displays boundary information indicating the placement area in association with the product name.
4. The boundary information is displayed to indicate the position of the group of products with the product name on a display shelf on which the group of products is displayed, The analysis device according to claim 3 , wherein the display format indicates a correspondence between the position indicated by the boundary information and the product name.
5. An analysis device as described in claim 3 or 4, wherein the boundary information is displayed superimposed on an image of the group of products in the placement area.
6. An analysis device described in any one of claims 3 to 5, wherein the one or more processors generate data for displaying the boundary information and notify the user terminal of the information including the product name and the data.
7. An analysis device described in any one of claims 3 to 6, wherein the boundary information is a boundary line surrounding the placement area.
8. The analysis device according to claim 1 , wherein the image of the sales floor is an image showing a plurality of types of merchandise displayed in the sales floor.
9. The analysis device according to claim 8 , wherein the one or more processors estimate the placement area for each type of product based on an image of the sales floor showing the multiple types of products displayed.
10. The analysis device according to claim 1 , wherein the one or more processors estimate the display state in the placement area based on the estimated placement area, and notify information about the estimated display state.
11. The analysis device according to claim 1 , wherein the one or more processors estimate the placement region using a first trained machine learning model.
12. The analysis device of claim 11, wherein the first trained machine learning model receives an image of a sales floor showing multiple types of products on display and outputs data indicating multiple placement areas corresponding to each of the multiple types of products.
13. The analysis device according to claim 11 or 12, wherein the first trained machine learning model is trained by supervised learning using training data including images of a sales floor and annotations regarding information on the placement areas of product groups.
14. The analysis device according to claim 1 , wherein an image representing the estimated placement area is superimposed on the image of the sales floor.
15. The analysis device according to claim 1 , wherein the user terminal is a terminal used in a store where the image of the sales floor was captured.
16. The analysis device according to claim 1 , wherein the image of the sales floor is an image of the sales floor from which moving objects have been removed.
17. The analysis device according to claim 1 , wherein the product group arrangement area is an area in which identical products are arranged together.
18. The analysis device according to claim 1 , wherein the one or more processors estimate one or more of product names, product numbers, and product alignment of the group of products in the placement area.
19. 19. The analysis device of claim 1, wherein the one or more processors utilize a second trained machine learning model to estimate one or both of points and orientations within the area of products within the estimated placement area, and estimate one or both of the number of products and product alignment in the group of products based on the estimated one or both of points and orientations.
20. An analysis device as described in any one of claims 1 to 19, wherein when the display state satisfies a predetermined condition, the one or more processors notify information about the display state, including information indicating the estimated placement area, as information about the display state.
21. The one or more processors: An analysis device according to any one of claims 1 to 20, wherein when it is determined that the estimated number of products in the placement area is in a predetermined state, the analysis device notifies the estimated placement area of the product group, the product names of the product group, and the work content for the placement area.
22. The product name included in the notified information is information that identifies an individual product.
22. An analysis device according to any one of claims 1 to 21.
23. The product name included in the notified information is represented by product identification information pre-assigned to the product name.
23. An analysis device according to any one of claims 1 to 22.
24. The product identification information is a product number assigned in advance to the product name.
24. The analysis device of claim 23.
25. A program for causing one or more processors to execute each process in the analysis device described in any one of claims 1 to 24.
26. An analysis device according to any one of claims 1 to 24; an imaging device that captures an image of the sales floor; An analysis system comprising:
27. The analysis system according to claim 26, wherein the imaging device is installed in a store and captures an image of the sales floor.
28. 28. The analysis system according to claim 26, further comprising a user terminal that acquires the information.
29. The analysis system described in Claim 28, wherein the user terminal displays the placement area and the product name in a display format that associates them with each other based on the acquired information.
30. One or more processors Estimating a location area of a group of products with the same product name in an image of the sales floor based on the image of the sales floor; An analysis method that notifies a user terminal of information regarding the display state of the product group in the estimated placement area, the information regarding the display state including the product names of the product group in the placement area.