Information Processing Systems

The information processing system addresses the limitations of existing tracking methods by generating learning data from video analysis, capturing customer interactions and paths, enhancing store optimization through understanding customer behavior and product interactions.

JP7810959B1Active Publication Date: 2026-02-04MARKETVISION CO LTD
View PDF 4 Cites 0 Cited by

Patent Information

Application Number
JP2025013225
Authority / Receiving Office
JP · JP
Patent Type
Patents
Current Assignee / Owner
Filing Date
2025-01-29
Publication Date
2026-02-04
Estimated Expiration
2045-01-29

AI Technical Summary

Technical Problem

Existing methods for tracking customer behavior in stores, such as using IC tags and beacons, are cumbersome and do not effectively capture product interaction data, while camera-based methods only track paths without understanding customer behavior on shelves.

Method used

An information processing system that generates learning data from video data by analyzing customer motions, extracting product interactions, and tagging relevant video segments for machine learning, utilizing a configuration that includes image acquisition, motion analysis, action extraction, and learning data generation units.

Benefits of technology

Enables efficient generation of learning data to understand customer shopping habits, including product interactions and paths, allowing stores to optimize product display based on customer behavior.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure 0007810959000001_ABST
    Figure 0007810959000001_ABST
Patent Text Reader

Abstract

The objective is to provide an information processing system that generates learning data for machine learning from video data. [Solution] An information processing system that generates learning data from video data, the information processing system having an image acquisition processing unit that acquires video data captured by a camera, an action analysis processing unit that analyzes the purchaser's actions from the acquired video data, an action extraction processing unit that extracts video data including trigger actions based on the analyzed actions, and a learning data generation processing unit that generates learning data by tagging all or part of the extracted video data.
Need to check novelty before this filing date? Find Prior Art

Description

[Technical Field]

[0001] The present invention relates to an information processing system that generates learning data for machine learning from video data. [Background technology]

[0002] In convenience stores, supermarkets, and other stores, it is common for products to be sold on display shelves. After entering the store, customers search for the display location of the product they want, place the product in their shopping basket or cart, and then make their purchase.

[0003] If a customer is familiar with the store, they will have some knowledge of the product display locations, and can quickly go to the display location of their destination and purchase the product. This allows the customer to efficiently move around while purchasing products.

[0004] On the other hand, if a customer is not familiar with the store, such as if it is their first time using the store, they will have to search around the store because they do not know where the products are displayed, and they will not be able to move around efficiently.

[0005] Therefore, it is important for stores to understand customer movements in order to improve time efficiency and avoid overcrowding. To do this, a function is required that can record and analyze not only customer movement paths, but also where and when products are picked up within those paths.

[0006] It is also important for stores to know the path that customers take when purchasing products, as this gives them a glimpse into the motivations behind customers' purchases.

[0007] There are various methods for finding out the movement of customers. For example, there are methods using beacons, methods using cameras to take pictures of the inside of the store and track the movement of customers, and methods using the GPS function equipped on smartphones, etc. In addition to these, there is also a system described in Patent Document 1 below. [Prior art documents] [Patent documents]

[0008] [Patent Document 1] Japanese Patent Application Laid-Open No. 2006-209421 Summary of the Invention [Problem to be solved by the invention]

[0009] In the above-mentioned Patent Document 1, IC tags are attached to products and display shelves, and customers wear a mobile terminal that reads the IC tags on their wrists, and by reading the IC tags with the mobile terminal, it is possible to track their behavior and location.

[0010] However, the system in Patent Document 1 is cumbersome because it requires attaching IC tags to each product and display shelf, and customers also need to wear mobile devices. The same applies when using beacons.

[0011] Furthermore, as mentioned above, in the case of a method of installing a camera or other photographic device in a store and tracking the movement of customers from the captured image data to understand their path and movements, this method only allows for understanding the path, which is the route the customer took, but does not allow for understanding how the customer behaved on the shelves or what products they picked up.

[0012] In particular, there are various methods for understanding the behavior of purchasers. For example, when using deep learning, training data is required to learn purchaser behavior. Generating training data requires extracting image data of specific frames from the video data to be used for training, or extracting video data from a certain section into semantic units, and then tagging them. However, performing such work manually is a heavy workload. [Means for solving the problem]

[0013] In view of the above-mentioned problems, the present inventors have invented an information processing system that generates learning data for machine learning from video data.

[0014] A first invention is an information processing system for generating learning data from video data, the information processing system including: an image acquisition processing unit that acquires video data captured by a camera; a motion analysis processing unit that analyzes a purchaser's motion from the acquired video data; and a motion extraction processing unit that extracts video data including a trigger motion based on the analyzed motion. a product analysis processing unit that analyzes products picked up by the purchaser shown in the extracted video data; a learning data generation processing unit that generates learning data by tagging all or part of the extracted video data; and the learning data generation processing unit generates learning data by tagging the product picked up by the consumer to all or a part of the extracted video data. It is an information processing system.

[0015] With the configuration of the present invention, learning data can be extracted from video data captured by a camera. It is a good idea to tag the video data with the products that the customer picks up and the customer's actions that are captured in the video data.

[0016] In the above-mentioned invention, the action extraction processing unit can be configured as an information processing system that extracts video data between the triggering action and a specified action that occurs before or after the triggering action.

[0017] In the above-mentioned invention, the action extraction processing unit can be configured as an information processing system that extracts video data between the triggering action and a predetermined time before or after the triggering action.

[0018] It is preferable to extract video data at intervals determined by these inventions.

[0020] In the above-mentioned invention, the information processing system can be configured as an information processing system in which the learning data generation processing unit generates learning data by tagging all or part of the extracted video data with the movements analyzed by the movement analysis processing unit.

[0021] It is a good idea to tag the video data with the products that the customer picks up and the customer's actions that are captured in the video data.

[0022] In the above-mentioned invention, the information processing system can be configured as an information processing system having a learning model processing unit that generates a learning model using the learning data generated by the learning data generation processing unit, an image input processing unit that inputs video data to be processed into the learning model, and an output processing unit that outputs the output result of the learning model, wherein the image input processing unit inputs video data of the location where a purchaser picked up a product, identified using the movement line in a movement line analysis system and information on the time or duration thereof, into the learning model, and the output processing unit accepts the output of product identification information of the product picked up by the purchaser as the processing result for the video data input into the learning model.

[0023] By linking the information processing system of the present invention with a flow line analysis system, it is possible to determine what products a customer who is the subject of flow line analysis has picked up. As a result, it is possible to determine the route the customer took after entering the store, and the order in which the customers purchased the products. This allows stores to obtain information on the customer's shopping habits, which can be used as a reference when considering where to display products.

[0024] The information processing system of the first invention can be realized by loading and executing the information processing program of the present invention into a computer, that is, the computer comprises an image acquisition processing unit that acquires video data captured by a video capture device, a motion analysis processing unit that analyzes the purchaser's motion from the acquired video data, a motion extraction processing unit that extracts video data including a trigger motion based on the analyzed motion, a product analysis processing unit that analyzes the products picked up by the purchaser shown in the extracted video data; a learning data generation processing unit that generates learning data by tagging all or part of the extracted video data; the learning data generation processing unit generates learning data by tagging the product picked up by the purchaser to all or a part of the extracted video data. It is an information processing program. [Effects of the Invention]

[0025] By using the information processing system of the present invention, learning data for machine learning can be generated from video data. [Brief explanation of the drawings]

[0026] [Figure 1] 1 is a block diagram schematically illustrating an example of a configuration of an information processing system according to the present invention. [Figure 2] FIG. 2 is a block diagram schematically illustrating an example of a hardware configuration of a computer used in the information processing system of the present invention. [Figure 3] 3 is a flowchart showing an example of the overall processing in the information processing system of the present invention. [Figure 4] FIG. 1 is a diagram schematically illustrating a case where a purchaser purchases a product in a store. [Figure 5]10A and 10B are diagrams illustrating an example of a process for extracting image data in an action extraction processing unit. [Figure 6] FIG. 10 is a diagram schematically illustrating an example of learning data in a learning data generation processing unit. [Figure 7] FIG. 10 is a block diagram illustrating an example of the configuration of an information processing system according to a second embodiment. [Figure 8] FIG. 10 is a diagram schematically illustrating an example of a flow line analyzed by a flow line analysis system superimposed on a floor plan of a store. [Figure 9] FIG. 10 is a diagram schematically illustrating the position where a purchaser picks up a product. [Figure 10] 10 is a flowchart illustrating an example of the overall processing in the information processing system according to the second embodiment. DETAILED DESCRIPTION OF THE INVENTION

[0027] FIG. 1 shows a block diagram of an example of the overall processing functions of an information processing system 1 of the present invention. The information processing system 1 uses a management terminal 2 and a photographing device 3. The management terminal 2 is a computer used by an organization such as a company that operates the information processing system 1. The photographing device 3 is a camera that photographs the interior of a retail store or other store, and one or more photographing devices, preferably two or more, are installed in the store. The photographing device 3 may be installed in a fixed location or may be a mobile device that photographs the interior of the store. It may also be a portable communication terminal such as a smartphone. The photographing device 3 may include a photographing device 3 that photographs display shelves and a photographing device 3 that photographs a customer picking up an item from the display shelves. In this case, the photographing device 3 is preferably installed in a position slightly forward and above the display shelves so that the display surface of the display shelves can be photographed so that the customer can photograph which item to pick up from the display shelves.

[0028] The management terminal 2 in the information processing system 1 is realized using a computer. An example of the hardware configuration of a computer is shown in Figure 2. The computer has a calculation device 70 such as a CPU that executes the calculation processing of a program, a storage device 71 such as a RAM or hard disk that stores information, a display device 72 such as a display that displays information, an input device 73 such as a keyboard or mouse that can input information, and a communication device 74 that sends and receives the processing results of the calculation device 70 and the information stored in the storage device 71 via a network such as the Internet or a LAN.

[0029] If the computer is equipped with a touch panel display, the display device 72 may be integrated with the input device 73. Touch panel displays are often used in portable communication terminals such as tablet computers and smartphones, but are not limited to these.

[0030] The touch panel display is a device that integrates the functions of the display device 72 and the input device 73 in that input can be made directly on the display using a predetermined input device (such as a touch panel pen) or a finger.

[0031] The functions of the various means in the present invention are only logically distinct, and may be physically or practically the same area. The order of the processes in the various means of the present invention may be changed as appropriate. In addition, some of the processes may be omitted.

[0032] The management terminal 2 of the information processing system 1 has an image acquisition processing unit 20, a motion analysis processing unit 21, a motion extraction processing unit 22, a commodity analysis processing unit 23, a commodity display information storage unit 24, and a learning data generation processing unit 25.

[0033] The image acquisition processing unit 20 acquires video data (image data) captured by the photographing device 3. The image acquisition processing unit 20 acquires the video data from the photographing device 3 by associating it with identification information that can identify the photographing device 3. It is also advisable to associate the video data with date and time information of the photographing by the photographing device 3. The image acquisition processing unit 20 acquires image data captured by the photographing device 3 of a customer moving around the store, picking up a product from a display shelf, etc.

[0034] The motion analysis processing unit 21 analyzes the motion of the purchaser captured in the image data of each frame of the video data. For example, the image data of each frame of the video data is analyzed to determine whether the purchaser is performing predetermined motions such as positioning themselves in front of a product, stopping in front of a product, extending their hand toward the product, picking up the product (taking up the product), putting the picked-up product into a shopping basket or cart, returning the picked-up product to the display shelf, or returning the product in the shopping basket or cart to the display shelf.

[0035] It is preferable to use deep learning as machine learning for analyzing the movements. In this case, the learning model is a learning model in which the weighting coefficients between neurons in each layer of a neural network consisting of multiple intermediate layers are optimized. A learning model is generated in advance using learning data that associates image data of each frame of video data showing each movement with the consumer movement shown in that image data. Image data of each frame of video data acquired by the image acquisition processing unit 20 is input to the learning model, and information on the corresponding movement is output. Note that if there is no corresponding movement, it may output, for example, "No target movement."

[0036] The motion analysis processing unit 21 stores the video data to be analyzed and the analyzed motion in association with each other.

[0037] The machine learning used by the motion analysis processing unit 21 is not limited to deep learning, and other machine learning methods may also be used, and analysis methods other than machine learning may also be used.

[0038] Based on the actions analyzed by the action analysis processing unit 21, the action extraction processing unit 22 extracts video data a predetermined range before or after a predetermined trigger action. The predetermined range to be extracted may include video data between the trigger action and an action associated with that trigger, or video data a predetermined time before or after the trigger action. This process is schematically shown in FIG. 5. In FIG. 5, when the trigger action is "putting an item into a cart" and the action associated with the trigger action is "stopping in front of the item," upon detecting the trigger action "putting an item into a cart," the action extraction processing unit 22 extracts video data between image data of a frame in the video data associated with the action "putting an item into a cart" and image data of a frame in the video data associated with the earlier action "stopping in front of the item."

[0039] The product analysis processing unit 23 analyzes the behavior of the purchaser captured in the video data extracted by the action extraction processing unit 22. For example, the product analysis processing unit 23 analyzes the product picked up by the purchaser based on video data associated with an action to be analyzed, such as "picking up a product," among the video data extracted by the action extraction processing unit 22, with reference to the product display information storage unit 24 described below.

[0040] As will be described later, the product display information storage unit 24 stores product identification information of products displayed on the display shelves in association with their display locations, and therefore the product analysis processing unit 23 identifies the display location of the product picked up by the customer in image data of a frame of video data associated with the action to be analyzed, and extracts from the product display information storage unit 24 the product identification information of the product placed in that display location to be stored in the product display information storage unit 24. This allows the product analysis processing unit 23 to analyze the product picked up by the customer from the image data of a frame of video data of the action to be analyzed.

[0041] The product display information storage unit 24 stores product identification information of products displayed on the display shelves and their display locations in association with image data of the product's display shelf. Product identification information includes product names as well as codes that can identify products, such as JAN codes. The display locations correspond to position information in the image data of the product's display shelf, or the face positions of the products on the display shelf. The position information in the image data of the product's display shelf includes coordinate information in the image data, such as x-coordinates and y-coordinates. The face positions include information on the shelf level of the product shown in the image data of the product's display shelf and the order of the product on the shelf from the right or left. The face of a product is the side of the product's package that is visible on the display shelf, and the position of that face on the display shelf is stored as the face position.

[0042] The product identification information of the products displayed on the shelves and their display locations are identified using various known techniques and stored in the product display information storage unit 24.

[0043] The learning data generation processing unit 25 uses all or part of the video data extracted by the action extraction processing unit 22 as learning data for deep learning. For example, the learning data generation processing unit 25 annotates the extracted video data and the product identification information of the product analyzed by the product analysis processing unit 23 as tags to generate learning data. The learning data generation processing unit 25 may also annotate the extracted video data and the actions associated with each video data as tags to generate learning data. In this case, the extracted video data may be divided into one or more image data for each action, and the actions may be annotated as tags. This is schematically shown in FIG. 6. [Example]

[0044] Next, an example of processing using the information processing system 1 of the present invention will be described with reference to the flowchart in Fig. 3. Note that the product display information storage unit 24 previously stores product identification information of products displayed on the display shelves in association with their display locations.

[0045] Each camera device 3 installed in the store captures images of the inside of the store (S100). Then, at a predetermined timing, the video data captured by each camera device 3 is acquired and stored by the image acquisition processing unit 20 (S110). At this time, the acquired video data and its shooting date and time information are stored in association with camera device identification information that identifies the camera device 3.

[0046] Then, the motion analysis processing unit 21 analyzes the motion of the purchaser in the image data of each frame of the video data acquired by the image acquisition processing unit 20 (S120). At this time, the image data of all frames of the video data may be analyzed, or the image data of some frames may be analyzed. For example, the analysis may be performed every predetermined number of frames, or every predetermined time period.

[0047] Then, when there is a triggering action among the actions analyzed by the action analysis processing unit 21 and associated with the video data, the action extraction processing unit 22 extracts video data a predetermined range before or after the triggering action (S130). For example, when the action extraction processing unit 22 detects the triggering action "putting an item into a cart", it extracts video data between the action "putting an item into a cart" and the earlier action "stopping in front of an item".

[0048] The product analysis processing unit 23 extracts the behavior of the purchaser captured in the video data extracted by the action extraction processing unit 22 (S140). For example, the product analysis processing unit 23 identifies and extracts product identification information of the product picked up by the purchaser.

[0049] All or part of the video data extracted by the action extraction processing unit 22 is generated as learning data for deep learning (S150).

[0050] Through the above processing, it is possible to generate learning data such as video data when a specific product is acquired, video data when a specific action is performed, etc. Then, by performing deep learning using the generated learning data, it is possible to generate a learning model of the video data when a specific product is acquired, a learning model of the video data when a specific action is performed, etc. [Example]

[0051] Next, a case where the information processing system 1 of the present invention is used in a flow line analysis system 4 (hereinafter referred to as "flow line analysis system 4") that analyzes the trajectory (flow line) of a purchaser's movement within a store will be described. The flow line analysis system 4 can use various known analysis systems that can acquire information on the flow line of a purchaser moving within a store. An example of the configuration in this embodiment is shown in FIG. 7.

[0052] The management terminal 2 in the information processing system 1 includes a learning model processing unit 26, a learning model storage unit 27, an image input processing unit 28, and an output processing unit 29 in addition to the configuration of the first embodiment.

[0053] The learning model processing unit 26 generates a learning model using the learning data generated by the learning data generation processing unit 25. The learning data to be input at this time may be, for example, video data and data tagged with product identification information of products picked up in the video data.

[0054] The learning model storage unit 27 stores the learning model generated by the learning model processing unit 26. An example of the learning model is a learning model tagged with video data and product identification information of a picked-up product.

[0055] The image input processing unit 28 accepts input of video data of the location where a purchaser picks up a product in a store, analyzed by the traffic flow analysis system 4 described below, and inputs the data into the learning model stored in the learning model memory unit 27.

[0056] The output processing unit 29 receives the output of the processing results that are input to the learning model by the image input processing unit 28. Since the image input processing unit 28 receives input of video data of the position where the purchaser picks up the product in the store and inputs this to the learning model, the output processing unit 29 can receive output of product identification information of the product picked up by the purchaser in the video data.

[0057] The flow line analysis system 4 is a system that analyzes the trajectory (flow line) of a customer moving through a store, and various known flow line analysis systems 4 can be used. Flow lines are calculated by recording the customer's locations in chronological order and connecting them with lines to show the trajectory of the customer's movement. Time or period information, such as date and time information at each location, is stored in association with the flow lines. Figure 8 shows an example of flow lines analyzed by the flow line analysis system 4 superimposed on a store floor plan.

[0058] The flow line analysis system 4 can recognize a purchaser from video data captured by a camera 3 installed in a store, for example, and analyze the purchaser's flow line based on the video data. Also, by attaching a beacon or other transmitter to a shopping basket or shopping cart and detecting it, information such as the purchaser's location and time can be analyzed as the purchaser's flow line. Furthermore, by attaching a GPS device to a shopping basket or shopping cart or to a portable communication device such as a smartphone held by the purchaser and detecting its location information, information such as the purchaser's location and time can be analyzed as the purchaser's flow line.

[0059] Various known systems can be used as the flow line analysis system 4, and it is not limited to the above. The flow line analysis system 4 is preferably a system that can associate the location of a purchaser with time or hour information.

[0060] A plurality of camera devices 3 are installed in the store to capture images of the movement of customers when analyzing the movement of customers with the movement line analysis system 4. The camera devices 3 capture images of customers picking up products. The camera devices 3 may be the same as the camera devices 3 that capture video data for generating learning data, or they may be different devices.

[0061] Next, an example of processing by the information processing system 1 in this embodiment will be described with reference to the flowchart of FIG.

[0062] The processing from S200 to S250 in FIG. 10 is the same as the processing from S100 to S150 in FIG. 3, and therefore a description thereof will be omitted.

[0063] The learning model processing unit 26 generates a learning model using the learning data generated in S250, and stores the model in the learning model storage unit 27 (S260).

[0064] The flow line analysis system 4 identifies positions where people stay for a certain period of time or more from data that associates flow lines with information indicating time or duration, such as date and time information. For example, as shown in Figure 9, it identifies positions P1 to P6.

[0065] Then, video data capturing the actions of the purchaser at each of the identified positions P1 to P6 is identified. This identification is performed, for example, by identifying each camera device 3 capturing video of positions P1 to P6 where the purchaser remains for a certain period of time or more, and identifying a predetermined range of video data captured by each camera device 3 that includes the time the purchaser remained. For example, this is performed by extracting from a predetermined storage area video data of each camera device 3 at each of positions P1 to P6 that includes each time. Each camera device 3 capturing video of each of positions P1 to P6 is set so that it captures the product the purchaser is holding in its hand.

[0066] Then, the image input processing unit 28 inputs the above-specified video data captured by each of the image capturing devices 3 at each of the positions P1 to P6 to the learning model storage unit 27 (S270). With this input, deep learning processing is executed using the learning model stored in the learning model storage unit 27, and the output processing unit 29 receives output of product identification information for the products picked up by the purchasers captured in the input video data at each of the positions P1 to P6 (S280).

[0067] By performing the above process, it is possible to identify the location and product that a purchaser picked up in the flow line data, thereby determining when, in what order, and along what flow line the purchaser followed after entering the store to purchase the product. [Industrial Applicability]

[0068] By using the information processing system 1 of the present invention, it is possible to generate learning data for machine learning from video data. [Explanation of symbols]

[0069] 1: Information processing system 2: Management terminal 3: Imaging device 4: Traffic flow analysis system 20: Image acquisition processing unit 21: Motion analysis processing unit 22: Action extraction processing unit 23: Product analysis processing unit 24: Product display information storage section 25: Learning data generation processing unit 26: Learning model processing unit 27: Learning model memory unit 28: Image input processing unit 29: Output processing section 70: Arithmetic device 71:Storage device 72:Display device 73: Input device 74:Communication equipment

Claims

1. An information processing system for generating learning data from video data, The information processing system includes: an image acquisition processing unit that acquires video data captured by the imaging device; a motion analysis processing unit that analyzes the purchaser's motion from the acquired video data; an action extraction processing unit that extracts video data including a trigger action based on the analyzed action; a product analysis processing unit that analyzes products picked up by the purchaser shown in the extracted video data; a learning data generation processing unit that generates learning data by tagging all or part of the extracted video data, The learning data generation processing unit generating learning data by tagging the product picked up by the purchaser to all or a part of the extracted video data; An information processing system comprising:

2. The action extraction processing unit extracting video data between the triggering action and a predetermined action that precedes or follows the triggering action; 2. The information processing system according to claim 1, wherein:

3. The action extraction processing unit extracting video data from the triggering action and a predetermined time before or after the triggering action; 2. The information processing system according to claim 1, wherein:

4. The information processing system includes: The learning data generation processing unit generating learning data by tagging the motion analyzed by the motion analysis processing unit to all or part of the extracted motion data; 4. The information processing system according to claim 1, wherein the information processing system is a data processing system.

5. The information processing system includes: a learning model processing unit that generates a learning model using the learning data generated by the learning data generation processing unit; an image input processing unit that inputs video data to be processed into the learning model; an output processing unit that outputs an output result of the learning model, The image input processing unit Video data of the location where the purchaser picked up the product, which is identified using the flow line and the time or duration information in the flow line analysis system, is input into the learning model; The output processing unit receiving an output of product identification information of the product picked up by the purchaser as a processing result of the video data input to the learning model; 2. The information processing system according to claim 1, wherein:

6. Computer, an image acquisition processing unit that acquires video data captured by the imaging device; a motion analysis processing unit that analyzes the purchaser's motion from the acquired video data; an action extraction processing unit that extracts video data including a trigger action based on the analyzed action; a product analysis processing unit that analyzes the products picked up by the purchaser shown in the extracted video data; a learning data generation processing unit that generates learning data by tagging all or part of the extracted video data; The learning data generation processing unit generating learning data by tagging the product picked up by the purchaser to all or a part of the extracted video data; An information processing program characterized by:

Citation Information

Patent Citations

  • Operation determination program, operation determination method, and operation determination device

    JP2022187215A

  • Information processing program, information processing method, and information processing device

    JP2023098020A

  • Information processing program, method for processing information, and information processor

    JP2023098506A

  • Method and device for recognizing action and commodity managing system

    JP2006209421A