Information processing program, information processing method, and information processing device.
Patent Information
- Authority / Receiving Office
- JP · JP
- Patent Type
- Patents
- Current Assignee / Owner
- FUJITSU LTD
- Filing Date
- 2022-12-14
- Publication Date
- 2026-08-04
AI Technical Summary
【0009】 顧客に相当する第一の物体と、商品に相当する第二の物体との関係性に応じた情報を提供することができる。
Smart Images

Figure 0007899704000001 
Figure 0007899704000002 
Figure 0007899704000003
Abstract
Description
Technical Field
[0001] The present invention relates to an information processing program and the like.
Background Art
[0002] In stores, various efforts are being made to sell more products. For example, information on preset products is displayed on cash registers and the like, and sales staff greet customers. If the sales staff can greet customers appropriately when the customers show interest in a certain product, the customers' willingness to purchase can be increased.
Prior Art Documents
Patent Documents
[0003]
Patent Document 1
Summary of the Invention
Problems to be Solved by the Invention
[0004] As described above, simply displaying product information often fails to increase customers' willingness to purchase. Also, although sales staff greet customers based on their know-how and advice from other sales staff, since it largely depends on the skills of the sales staff themselves, currently, each sales staff is not able to greet customers appropriately.
[0005] Therefore, there is a need to provide information such as product information and information that assists sales staff in greeting customers according to the relationship between the customer and the product.
[0006] For example, it is desirable to present information corresponding to the relationship between a first object corresponding to a customer and a second object corresponding to a product to the sales staff or the customer.
[0007] In one aspect, the present invention aims to provide an information processing program, an information processing method, and an information processing device that can provide information corresponding to the relationship between a first object corresponding to a customer and a second object corresponding to a product. [Means for solving the problem]
[0008] In the first proposal, the computer performs the following processes: The computer acquires an image. By analyzing the acquired image, the computer identifies a first region containing the first object, a second region containing the second object, and the relationships that identify the interaction between the first and second objects. Based on the identified relationships, the computer selects a model from among several models that is related to either the first or second object, and outputs the selected model. [Effects of the Invention]
[0009] It can provide information that corresponds to the relationship between a first object, which represents the customer, and a second object, which represents the product. [Brief explanation of the drawing]
[0010] [Figure 1] Figure 1 shows an example of the system according to this embodiment 1. [Figure 2] Figure 2 is a diagram (1) illustrating the processing of the information processing device according to this embodiment 1. [Figure 3] Figure 3 is a diagram (2) illustrating the processing of the information processing device according to this embodiment 1. [Figure 4] Figure 4 illustrates HOID's machine learning approach. [Figure 5] Figure 5 is a functional block diagram showing the configuration of the information processing device according to this embodiment 1. [Figure 6] Figure 6 shows an example of the data structure of the model table according to this embodiment 1. [Figure 7] Figure 7 shows an example of the data structure of the display device management table according to this embodiment 1. [Figure 8] FIG. 8 is a flowchart showing a processing procedure of the information processing apparatus according to the first embodiment. [Figure 9] FIG. 9 is a diagram showing an example of the system according to the second embodiment. [Figure 10] FIG. 10 is a diagram (1) for explaining the processing of the information processing apparatus according to the second embodiment. [Figure 11] FIG. 11 is a diagram (2) for explaining the processing of the information processing apparatus according to the second embodiment. [Figure 12] FIG. 12 is a functional block diagram showing the configuration of the information processing apparatus according to the second embodiment. [Figure 13] FIG. 13 is a diagram showing an example of the data structure of the model table according to the second embodiment. [Figure 14] FIG. 14 is a flowchart showing a processing procedure of the information processing apparatus according to the second embodiment. [Figure 15] FIG. 15 is a diagram showing an example of the hardware configuration of a computer that realizes the same functions as the information processing apparatus of the embodiment.
MODE FOR CARRYING OUT THE INVENTION
[0011] Hereinafter, embodiments of the information processing program, information processing method, and information processing apparatus disclosed in the present application will be described in detail based on the drawings. Note that the present invention is not limited by this embodiment.
EXAMPLE
[0012] FIG. 1 is a diagram showing an example of the system according to the first embodiment. As shown in FIG. 1, this system includes cameras 10a, 10b, 10c, display devices 15a, 15b, 15c, and an information processing apparatus 100. The cameras 10a to 10c and the information processing apparatus 100 are interconnected via a network. Also, the display devices 15a to 15c and the information processing apparatus 100 are interconnected via a network.
[0013] In FIG. 1, for convenience of explanation, only cameras 10a to 10c and display devices 15a to 15c are shown, but the system according to the first embodiment may have other cameras and other display devices.
[0014] Cameras 10a to 10c are installed at predetermined positions in the store. A plurality of products are arranged in the store. The positions (coordinates) where cameras 10a to 10c are installed are set to be different positions. In the following description, when cameras 10a to 10c are not particularly distinguished, they are referred to as "camera 10".
[0015] Camera 10 captures an image of the store interior and transmits the data of the captured image to the information processing device 100. In the following description, the data of the image that camera 10 transmits to the information processing device 100 is referred to as "image data".
[0016] The image data includes a plurality of image frames in time series. Each image frame is assigned a frame number in ascending order of time series. One image frame is a still image captured by camera 10 at a certain timing. Each image frame may be provided with time data. The image data is set with camera identification information for identifying camera 10 that captured the image data.
[0017] Display devices 15a to 15c are installed at predetermined positions in the store, for example, around the products. The positions (coordinates) where display devices 15a to 15c are installed are set to be different positions. In the following description, when display devices 15a to 15c are not particularly distinguished, they are referred to as "display device 15". Display device 15 displays information about the product output from the information processing device 100.
[0018] The information processing device 100 acquires video data of the store from the camera 10 and analyzes the acquired video data to identify a first area containing customers who are likely to purchase products in the store, a second area containing products, and relationships that identify interactions between customers and products. Based on the identified relationships, the information processing device 100 selects a machine learning model from among multiple machine learning models stored in the memory unit. This allows for the selection of machine learning models related to customers and people, and by utilizing such machine learning models, information can be provided that is tailored to the relationship between customers and people.
[0019] Figures 2 and 3 are diagrams illustrating the processing of the information processing device according to this embodiment 1. First, Figure 2 will be explained. For example, the information processing device 100 analyzes the video data 20 captured by the camera 10 to identify a first area 20a containing the customer, a second area 20b containing the product, and the relationship between the customer and the product. In the example shown in Figure 2, the relationship between the person and the product is described as "grasping". Note that a display device 15 is installed near the product included in the first area 20a.
[0020] In the example shown in Figure 2, the relationship between the first region 20a and the second region 20b was described as "grasping," but the term "relationship" also includes other relationships such as "looking," "touching," and "sitting."
[0021] Let's move on to the explanation of Figure 3. The information processing device 100 has multiple machine learning models. Figure 3 shows machine learning models 30a, 30b, and 30c. For example, machine learning model 30a is a machine learning model specifically for the relationship "seeing". Machine learning model 30b is a machine learning model specifically for the relationship "touching". Machine learning model 30c is a machine learning model specifically for the relationship "grasping". Machine learning models 30a to 30c are Neural Networks (NN), etc.
[0022] The machine learning model 30a is pre-trained with multiple first training data sets corresponding to the relationship "looking". For example, the input to the first training data is image data of a product, and the output (ground truth label) is product information. The product information in the first training data is "product advertising information," etc.
[0023] The machine learning model 30b is pre-trained with multiple second training datasets corresponding to the relationship "touching". For example, the input to the second training dataset is image data of a product, and the output (ground truth label) is product information. The product information in the second training dataset includes "information describing the product's advantages" and "information describing the product's popularity," etc.
[0024] The machine learning model 30c is pre-trained with multiple third-party training data corresponding to the relationship "grasping". For example, the input to the third-party training data corresponding to the relationship "grasping" is image data of a product, and the output (ground truth label) is product information. The product information in the third-party training data is "information explaining the benefits obtained when purchasing the product," etc.
[0025] The information processing device 100 selects a machine learning model from machine learning models 30a to 30c that corresponds to the relationship identified by the process described in Figure 2. For example, if the identified relationship is "holding," the information processing device 100 selects machine learning model 30c shown in Figure 3.
[0026] The information processing device 100 identifies product information for products contained in the second region 20b by inputting image data of the second region 20b, which contains the products, into the selected machine learning model 30c. The information processing device 100 outputs the identified product information to a display device 15 located near the products contained in the second region, allowing the customer to refer to the product information. The product information that the customer is referred to is information output from a machine learning model based on the relationship between the customer and the product, and can increase the customer's willingness to purchase. Note that the product information is an example of "related information" concerning products contained in the second region.
[0027] Incidentally, the information processing device 100 uses HOID (Human Object Interaction Detection) to identify a first domain including customers, a second domain including products, and the relationship between the first and second domains. When the information processing device 100 inputs video data (time-series image frames) into HOID, information on the first domain, the second domain, and the relationships is output.
[0028] Here, we will describe an example of the HOID learning process performed by the information processing device 100. The information processing device 100 uses multiple training data to train a HOID that identifies a first class representing people, a second class representing objects, and the relationships between the first and second classes.
[0029] Each training dataset consists of input image data (image frames) and correct answer information set for that image data.
[0030] The ground truth information includes the classes of the person and object being detected, the class indicating the interaction between the person and the object, and the Bbox (Bounding Box) that represents the region of each class. For example, the ground truth information might include region information for the Something class representing an object, region information for the Human class representing a user, and the relationship indicating the interaction between the Something class and the Human class.
[0031] Furthermore, both the training data and the training data can be configured with multiple classes and multiple interactions, and the trained HOID can recognize multiple classes and multiple interactions.
[0032] Generally, when you create a Something class using standard object recognition, it detects everything unrelated to the task, such as the background, clothing, and accessories. Moreover, since they are all "Something," the image data simply contains a large number of Bboxes, and nothing useful is revealed. In the case of HOID, however, it is clear that there is a specific relationship between a person and an object (such as holding, sitting on, or manipulating it), so this can be used as meaningful information for the task.
[0033] Figure 4 illustrates the machine learning process of HOID. As shown in Figure 4, the information processing device 100 inputs training data into HOID and obtains the output results of HOID. These output results include the human class, object class, and human-object interactions detected by HOID. The information processing device 100 then calculates the error information between the correct information of the training data and the output results of HOID, and performs machine learning of HOID by backpropagation to minimize the error.
[0034] Next, an example of identification processing using HOID will be described. The information processing device 100 inputs each image frame of the video data captured by the camera 10 into HOID and obtains the output result of HOID. The output result of HOID includes the Bbox of a person, the Bbox of an object, the probability value of the interaction between a person and an object (the probability value of each relationship), and the class name. The Bbox of a person corresponds to the first domain described above. The Bbox of an object corresponds to the second domain described above. Based on the output result of HOID, the information processing device 100 identifies the relationship between the first domain and the second domain that has the highest probability value.
[0035] As described above, the information processing device 100 can identify the first region, the second region, and the relationship between them by inputting the video data into the HOID. Alternatively, the information processing device 100 may store machine-learned HOIDs in a memory unit beforehand and use these HOIDs to identify the first region, the second region, and the relationship between them.
[0036] Next, an example configuration of the information processing device 100 that performs the processes shown in Figures 2 and 3 will be described. Figure 5 is a functional block diagram showing the configuration of the information processing device according to this embodiment 1. As shown in Figure 5, this information processing device 100 has a communication unit 110, an input unit 120, a display unit 130, a storage unit 140, and a control unit 150.
[0037] The communication unit 110 performs data communication with the camera 10, display device 15, external devices, etc., via the network. The communication unit 110 is a NIC (Network Interface Card), etc. For example, the communication unit 110 receives video data from the camera 10.
[0038] The input unit 120 is an input device that inputs various types of information to the control unit 150 of the information processing device 100. For example, the input unit 120 can be a keyboard, mouse, touch panel, etc.
[0039] The display unit 130 is a display device that displays information output from the control unit 150.
[0040] The memory unit 140 includes an HOID 141, a video buffer 142, a model table 143, and a display device management table 144. The memory unit 140 is a storage device such as memory.
[0041] HOID141 is the HOID explained in Figure 4, etc. By inputting an image frame of video data into HOID141, the first region, the second region, and the relationship between the first region (objects contained in the first region) and the second region (objects contained in the second region) on the image frame are output.
[0042] The video buffer 142 holds the video data captured by the camera 10. For example, the video buffer 142 stores the video data in association with camera identification information.
[0043] Model table 143 holds information about multiple machine learning models 30a to 30c, as explained in Figure 3. Figure 6 shows an example of the data structure of the model table in this embodiment 1. As shown in Figure 6, this model table 143 associates model identification information, relationships, and machine learning models. Model identification information is information that uniquely identifies a machine learning model. Relationships indicate the relationships corresponding to machine learning models. A machine learning model is a neural network (NN) that takes image data (image frames) as input and outputs product information.
[0044] For example, model identification information "M30a" indicates machine learning model 30a. Machine learning model 30a is a machine learning model that corresponds to the relationship "looking". Model identification information "M30b" indicates machine learning model 30b. Machine learning model 30b is a machine learning model that corresponds to the relationship "touching". Model identification information "M30c" indicates machine learning model 30c. Machine learning model 30c is a machine learning model that corresponds to the relationship "holding".
[0045] The display device management table 144 holds information about the display devices 15 placed in the store. Figure 7 shows an example of the data structure of the display device management table according to this embodiment 1. As shown in Figure 7, this display device management table 144 associates display device identification information with location and camera identification information.
[0046] The display device identification information is information that uniquely identifies the display device 15. For example, the display device identification information for display devices 15a, 15b, and 15c are A15a, A15b, and A15c, respectively. The position indicates the position (coordinates) of the display device 15. The camera identification information is information that identifies the camera 10 closest to the display device 15. For example, the camera identification information C10a, C10b, and C10c correspond to the cameras 10a, 10b, and 10c shown in Figure 1.
[0047] For example, in Figure 7, information is registered indicating that the display device 15a with display device identification information "A15a" is installed at position "(x1, y1)", and the camera 10 closest to the display device 15a is camera 10a with camera identification information "C10a".
[0048] Returning to the explanation of Figure 5, the control unit 150 includes an acquisition unit 151, an analysis unit 152, a identification unit 153, and a learning unit 154. The control unit 150 is a CPU (Central Processing Unit), a GPU (Graphics Processing Unit), etc.
[0049] The acquisition unit 151 acquires video data from the camera 10. As described above, the video data is set with camera identification information of the camera 10 that captured the video data. The acquisition unit 151 stores the video data in the video buffer 142, associating it with the camera identification information.
[0050] The analysis unit 152 acquires video data from the video buffer 142 and analyzes the video data to identify the first region, the second region, and their relationships. These relationships are those between "the first object contained in the first region" and "the second object contained in the second region." For example, the analysis unit 152 inputs the time-series image frames (image data) contained in the video data into the HOID 141 and identifies the first region, the second region, and their relationships on each image frame.
[0051] The analysis unit 152 repeatedly performs the above process to identify the first region, the second region, and relationships for each time-series image frame. When repeatedly identifying the first region, the second region, and relationships, the analysis unit 152 tracks customers included in the first region and products included in the second region. The analysis unit 152 generates information of the analysis results of the above process and outputs it to the identification unit 153. In the following description, the information of the analysis results will be referred to as "analysis result information". For example, the analysis result information associates image data of the first region containing the person being tracked, image data of the second region containing the product being tracked, relationships, and camera identification information of the camera 10 that captured the video data (image frames).
[0052] The identification unit 153 selects a machine learning model corresponding to a relationship from among multiple machine learning models registered in the model table 143, based on the relationships contained in the analysis result information. The identification unit 153 inputs the image data of the second region contained in the analysis result information into the selected machine learning model and obtains the product information output from the selected machine learning model (identifies the product information).
[0053] For example, if the relationship included in the analysis result information is "grasping", the identification unit 153 selects a machine learning model 30c corresponding to "grasping" from the model table 143. The identification unit 153 identifies product information by inputting image data from the second region into the selected machine learning model 30c.
[0054] The identification unit 153 identifies the display device identification information to display the product information based on the camera identification information included in the analysis result information and the display device management table 144. For example, if the camera identification information included in the analysis result information is "C10a", the identification unit 153 identifies the display device identification information "A15a (display device 15a)" to display the product information. The identification unit 153 outputs and displays the product information on the identified display device 15a.
[0055] The identification unit 153 may calculate the customer's position from the coordinates of the second region (customer) on the image frame and the camera parameters of the camera 10 corresponding to the camera identification information, and compare the calculated position with each distance in the display device management table 144. The identification unit 153 identifies the display device identification information that has the shortest distance to the calculated position, and outputs and displays the product information on the display device 15 corresponding to the identified display device identification information.
[0056] The learning unit 154 performs machine learning on HOID141 based on multiple training data sets. The learning unit 154 has acquired multiple training data sets in advance. The learning unit 154 inputs the training data into HOID141, calculates the error information between the correct information of the training data and the output result of HOID141, and performs machine learning on HOID141 by backpropagation to minimize the error. Other explanations regarding machine learning are the same as those explained in Figure 4.
[0057] Furthermore, the learning unit 154 may perform machine learning on the machine learning models 30a, 30b, and 30c included in the model table 143.
[0058] The learning unit 154 trains the machine learning model 30a based on multiple first training data sets. The learning unit 154 takes the first training data as input, calculates error information between the correct labels of the first training data and the output results of the machine learning model 30a, and performs machine learning on the machine learning model 30a by backpropagation to minimize the error.
[0059] The learning unit 154 trains the machine learning model 30b based on multiple second training data sets. The learning unit 154 takes the second training data as input, calculates error information between the correct labels of the second training data and the output results of the machine learning model 30b, and performs machine learning on the machine learning model 30b by backpropagation to minimize the error.
[0060] The learning unit 154 trains the machine learning model 30c based on multiple third training data sets. The learning unit 154 takes the third training data as input, calculates error information between the correct labels of the third training data and the output results of the machine learning model 30b, and performs machine learning on the machine learning model 30c by backpropagation to minimize the error.
[0061] Next, the processing procedure of the information processing device 100 according to this embodiment 1 will be described. Figure 8 is a flowchart of the processing procedure of the information processing device according to this embodiment 1. As shown in Figure 8, the acquisition unit 151 of the information processing device 100 acquires video data from the camera 10 and stores it in the video buffer 142 (step S101).
[0062] The analysis unit 152 of the information processing device 100 inputs each image frame of the video data into the HOID 141 and identifies the first region, the second region, and the relationship between the first and second regions for each image frame (step S102).
[0063] The identification unit 153 of the information processing device 100 selects a machine learning model from the model table 143 based on the identified relationships (step S103). The identification unit 153 identifies product information by inputting image data from a second region into the selected machine learning model (step S104).
[0064] The identification unit 153 identifies the display device 15 based on the display device management table 144 (step S105). The identification unit 153 causes the product information to be displayed on the display device (step S106).
[0065] Next, the effects of the information processing device 100 according to this embodiment 1 will be described. The information processing device 100 acquires video data of the store and analyzes the acquired video data to identify a first area containing customers who are interested in purchasing products in the store, a second area containing products, and relationships that identify the interaction between customers and products. Based on the identified relationships, the information processing device 100 selects a machine learning model from among a plurality of machine learning models stored in the storage unit 140. This allows for the selection of machine learning models related to customers and people, and by utilizing such machine learning models, information can be provided that is appropriate to the relationship between customers and people.
[0066] The information processing device 100 identifies product information by inputting image data from a second domain into a selected machine learning model, and outputs and displays the identified product information on the display device 15. The product information is information output from a machine learning model based on the relationship between the customer and the product, and by showing this product information to the customer, it is possible to increase the customer's willingness to purchase.
[0067] By the way, although the information processing device 100 in this embodiment 1 displays product information on the display device 15, it is not limited to this, and product information may also be displayed on a terminal device used by the customer. The terminal device used by the customer may be a cash register, digital signage, smart card, etc.
[0068] For example, when the analysis unit 152 of the information processing device 100 performs the above processing for each time-series image frame, it tracks the customer included in the first region. Based on the camera parameters of the camera 10 that captured the video data and the coordinates of the first region on the image frame, the analysis unit 152 identifies the customer's location within the store. Based on the identified customer location within the store, the analysis unit 152 identifies the terminal device used by the customer and outputs and displays product information to the identified terminal device. In this way, the information processing device 100 can efficiently show product information to the customer. [Examples]
[0069] Figure 9 shows an example of a system according to this second embodiment. As shown in Figure 9, this system includes cameras 10a, 10b, and 10c, a terminal device 25 held by a sales staff member 26, and an information processing device 200. The cameras 10a to 10c and the information processing device 200 are interconnected via a network. The terminal device 25 and the information processing device 200 are interconnected via a network (wireless).
[0070] For the sake of explanation, Figure 9 shows only cameras 10a to 10c and terminal device 25, but the system according to this embodiment 2 may have other cameras and other terminal devices.
[0071] Cameras 10a to 10c are installed in designated locations within the store. In the following description, unless otherwise specified, cameras 10a to 10c will be referred to as "camera 10". Camera 10 transmits video data to the information processing device 200. Further details regarding camera 10 are the same as those described in Example 1.
[0072] The terminal device 25 is held by the sales staff 26. The terminal device 25 displays customer service information output from the information processing device 200 to assist in customer service.
[0073] The information processing device 200 acquires video data of the store from the camera 10 and analyzes the acquired video data to identify a first area containing customers who are likely to purchase products in the store, a second area containing products, and relationships that identify interactions between customers and products. Based on the identified relationships, the information processing device 200 selects a customer service model from among multiple customer service models stored in the memory unit. This allows the selection of a customer service model related to the customer and the person, and by using such a customer service model, it is possible to provide sales staff 26 and others with information that is appropriate to the relationship between the customer and the person and can assist in customer service (customer service information). The customer service information is an example of "related information" related to products included in the second area. Furthermore, the customer service information is information about the content of customer service related to products included in the second area for customers included in the first area.
[0074] Figures 10 and 11 are diagrams illustrating the processing of the information processing device according to this second embodiment. First, Figure 10 will be explained. For example, the information processing device 200 analyzes the video data 20 captured by the camera 10 to identify a first area 20a including the customer, a second area 20b including the product, and the relationship between the customer and the product. In the example shown in Figure 10, the relationship between the person and the product is "grasped". A sales staff member 26 is assumed to be waiting near the product.
[0075] In the example explained in Figure 10, the relationship between the first region 20a and the second region 20b was described as "grasping," but the relationship also includes other relationships such as "looking," "touching," and "sitting."
[0076] Let's move on to the explanation of Figure 11. The information processing device 200 has multiple customer service models. Figure 11 shows customer service models 40a, 40b, and 40c. For example, customer service model 40a is a machine learning model specifically for the relationship "seeing". Customer service model 40b is a machine learning model specifically for the relationship "touching". Customer service model 40c is a machine learning model specifically for the relationship "grasping". Customer service models 40a to 40c are neural networks, etc.
[0077] The customer service model 40a is pre-trained with multiple fourth-stage training data corresponding to the relationship "looking". For example, the input to the fourth-stage training data corresponding to the relationship "looking" is image data of the product, and the output (correct label) is information about the customer service content to be performed for a customer looking at a product (hereinafter referred to as customer service information). Customer service information for a customer looking at a product may include "show a product in a different color" or "recommend other products".
[0078] The customer service model 40b is pre-trained with multiple fifth-level training data corresponding to the relationship "touching". For example, the input to the fifth-level training data corresponding to the relationship "touching" is image data of the product, and the output (ground truth label) is customer service information to be given to a customer touching the product. Customer service information for a customer touching the product includes "explaining the product's advantages" and "explaining the product's popularity."
[0079] The customer service model 40c is pre-trained with multiple sixth-level training data corresponding to the relationship "holding." For example, the input to the sixth-level training data corresponding to the relationship "holding" is image data of the product, and the output (correct label) is customer service information to be given to a customer who is holding the product. Customer service information for a customer who is holding the product includes "presenting the product's features" and "explaining the advantageous period for purchasing the product."
[0080] The information processing device 200 selects a customer service model from customer service models 40a to 40c that corresponds to the relationship identified by the process described in Figure 10. For example, if the identified relationship is "held," the information processing device 200 selects customer service model 40c.
[0081] The information processing device 200 identifies customer service information for products included in the second area by inputting an image of the second area, which includes the products, into the selected customer service model 40c. The information processing device 200 outputs the identified customer service information to the terminal device 25 used by the sales staff 26 for display. By referring to the customer service information, the sales staff 26 can provide more appropriate customer service to customers included in the first area.
[0082] Next, an example configuration of the information processing device 200 that performs the processes shown in Figures 10 and 11 will be described. Figure 12 is a functional block diagram showing the configuration of the information processing device according to this embodiment 2.
[0083] The description of the communication unit 210, input unit 220, and display unit 230 is the same as the description of the communication unit 110, input unit 120, and display unit 130 described in Figure 5.
[0084] The memory unit 240 includes an HOID 241, a video buffer 242, and a model table 243. The memory unit 240 is a storage device such as memory.
[0085] The description of HOID241 and video buffer 242 is the same as the description of HOID141 and video buffer 142 described in Example 1.
[0086] Model table 243 holds information about multiple customer service models 40a to 40c, as explained in Figure 11. Figure 13 shows an example of the data structure of the model table according to this embodiment 2. As shown in Figure 13, this model table 243 associates model identification information, relationships, and customer service models. Model identification information is information that uniquely identifies a machine learning model. Relationships indicate the relationships corresponding to machine learning models. A customer service model is a neural network that takes image data (image frames) as input and outputs customer service information.
[0087] For example, model identification information "M40a" indicates customer service model 40a. Customer service model 40a is a customer service model corresponding to the relationship "looking". Model identification information "M40b" indicates customer service model 40b. Customer service model 40b is a customer service model corresponding to the relationship "touching". Model identification information "M40c" indicates customer service model 40c. Customer service model 40c is a customer service model corresponding to the relationship "holding".
[0088] Returning to the explanation of Figure 12, the control unit 250 includes an acquisition unit 251, an analysis unit 252, a identification unit 253, and a learning unit 254. The control unit 250 is a CPU, GPU, etc.
[0089] The acquisition unit 251 acquires video data from the camera 10. As described above, the video data is set with camera identification information of the camera 10 that captured the video data. The acquisition unit 251 stores the video data in the video buffer 242, associating it with the camera identification information.
[0090] The analysis unit 252 acquires video data from the video buffer 242 and analyzes the video data to identify the first region, the second region, and their relationships. These relationships are those between "the first object contained in the first region" and "the second object contained in the second region." For example, the analysis unit 252 inputs the time-series image frames (image data) contained in the video data into the HOID 241 and identifies the first region, the second region, and their relationships on each image frame.
[0091] The analysis unit 252 repeatedly performs the above process to identify the first region, the second region, and relationships for each time-series image frame. When repeatedly identifying the first region, the second region, and relationships, the analysis unit 252 tracks customers included in the first region and products included in the second region. The analysis unit 252 generates information of the analysis results of the above process and outputs it to the identification unit 253. In the following description, the information of the analysis results will be referred to as "analysis result information". For example, the analysis result information associates image data of the first region containing the person being tracked, image data of the second region containing the product being tracked, and relationships.
[0092] The identification unit 253 selects a customer service model corresponding to a relationship from multiple customer service models registered in the model table 243, based on the relationships contained in the analysis result information. The identification unit 253 inputs the image data of the second region contained in the analysis result information to the selected customer service model and obtains the customer service information output from the selected customer service model (identifies the customer service information).
[0093] For example, if the relationship included in the analysis result information is "holding," the identification unit 253 selects a customer service model 40c corresponding to "holding" from the model table 243. The identification unit 253 identifies customer service information by inputting image data from the second region into the selected customer service model 40c.
[0094] The specific unit 253 outputs customer service information to the terminal device 25 held by the sales staff 26 and displays it.
[0095] The learning unit 254 performs machine learning on HOID241 based on multiple training data sets. The learning unit 254 has acquired multiple training data sets in advance. The learning unit 254 inputs the training data into HOID241, calculates the error information between the correct information of the training data and the output result of HOID241, and performs machine learning on HOID241 by backpropagation to minimize the error. Other explanations regarding machine learning are the same as those explained in Figure 4.
[0096] Furthermore, the learning unit 254 may perform machine learning on the customer service models 40a, 40b, and 40c included in the model table 243.
[0097] The learning unit 254 trains the customer service model 40a based on multiple fourth training data sets. The learning unit 254 takes the fourth training data as input, calculates error information between the correct labels of the fourth training data and the output results of the customer service model 40a, and performs machine learning on the customer service model 40a by backpropagation to minimize the error.
[0098] The learning unit 254 trains the customer service model 40b based on multiple fifth training data sets. The learning unit 254 takes the fifth training data as input, calculates error information between the correct labels of the fifth training data and the output results of the customer service model 40b, and performs machine learning on the customer service model 40b by backpropagation to minimize the error.
[0099] The learning unit 254 trains the customer service model 40c based on multiple sixth training data sets. The learning unit 254 takes the sixth training data as input, calculates error information between the correct labels of the sixth training data and the output results of the customer service model 40b, and performs machine learning on the customer service model 40c by backpropagation to minimize the error.
[0100] Next, the processing procedure of the information processing device 200 according to this second embodiment will be described. Figure 14 is a flowchart showing the processing procedure of the information processing device according to this second embodiment. As shown in Figure 14, the acquisition unit 251 of the information processing device 200 acquires video data from the camera 10 and stores it in the video buffer 242 (step S201).
[0101] The analysis unit 252 of the information processing device 200 inputs each image frame of the video data into the HOID 241 and identifies the first region, the second region, and the relationship between the first and second regions for each image frame (step S202).
[0102] The identification unit 253 of the information processing device 200 selects a customer service model from the model table 243 based on the identified relationship (step S203). The identification unit 253 identifies customer service information by inputting image data from a second area to the selected customer service model (step S204). The identification unit 253 displays the customer service information on the terminal device (step S205).
[0103] Next, the effects of the information processing device 200 according to this second embodiment will be described. The information processing device 200 acquires video data of the store and analyzes the acquired video data to identify a first area containing customers who are interested in purchasing products in the store, a second area containing products, and relationships that identify the interaction between customers and products. Based on the identified relationships, the information processing device 200 selects a customer service model from among a plurality of customer service models stored in the storage unit 240. The information processing device 200 identifies customer service information by inputting image data of the second area into the selected customer service model, and outputs and displays the identified customer service information on the terminal device 25. The customer service information is information output from a customer service model based on the relationship between the customer and the product, and by showing this customer service information to the sales staff 26, the sales staff 26 can provide customer service that increases the customer's desire to purchase.
[0104] Next, an example of a computer hardware configuration that achieves the same functions as the information processing devices 100 and 200 described above will be explained. Figure 15 shows an example of a computer hardware configuration that achieves the same functions as the information processing devices in the embodiment.
[0105] As shown in Figure 15, the computer 300 includes a CPU 301 that performs various calculations, an input device 302 that receives data input from the user, and a display 303. The computer 300 also includes a communication device 304 and an interface device 305 that exchange data with external devices via a wired or wireless network. Furthermore, the computer 300 includes a RAM 306 for temporarily storing various information and a hard disk drive 307. Each of these devices 301 to 307 is connected to a bus 308.
[0106] The hard disk drive 307 contains an acquisition program 307a, an analysis program 307b, a specific program 307c, and a learning program 307d. The CPU 301 reads each of the programs 307a to 307d and loads them into the RAM 306.
[0107] The acquisition program 307a functions as the acquisition process 306a. The analysis program 307b functions as the analysis process 306b. The identification program 307c functions as the identification process 306c. The learning program 307d functions as the learning process 306d.
[0108] The processing in acquisition process 306a corresponds to the processing in acquisition units 151 and 251. The processing in analysis process 306b corresponds to the processing in analysis units 152 and 252. The processing in identification process 306c corresponds to the processing in identification units 153 and 253. The processing in learning process 306d corresponds to the processing in learning units 154 and 254.
[0109] Furthermore, programs 307a to 307d do not necessarily have to be stored in the hard disk drive 307 from the beginning. For example, each program could be stored in a "portable physical medium" such as a flexible disk (FD), CD-ROM, DVD, magneto-optical disk, or IC card inserted into the computer 300. Then, the computer 300 could read and execute each program 307a to 307d.
[0110] With regard to embodiments including each of the above examples, the following additional information is disclosed.
[0111] (Note 1) Obtain video footage, By analyzing the acquired video, a first region containing the first object, a second region containing the second object, and the relationship between the interaction between the first object and the second object are identified within the video. Based on the identified relationship, select a model from among several models that is related to the first object or the second object. Output the selected model. An information processing program characterized by having a computer perform the processing.
[0112] (Note 2) By analyzing the video, the first object and the second object are tracked, Based on the identified relationship, a machine learning model is selected from among multiple machine learning models to be applied to the second object. By inputting the tracked image of the second object into the selected machine learning model, relevant information about the second object is identified. The relevant information regarding the identified second object is output to the display device associated with the tracked second object. The information processing program described in Appendix 1, characterized in that it causes a computer to perform further processing.
[0113] (Note 3) The first object is a person, The second object is a commodity, The system refers to a memory unit that associates the identified relationship with a plurality of customer service models in which customer service content is defined, and identifies a customer service model from the plurality of customer service models that corresponds to the identified relationship. Based on the identified customer service model, the customer service content related to the object represented by the second object for the person represented by the first object is identified. The identified customer service details are sent to the terminal used by the store employee. The information processing program described in Appendix 1, characterized in that it causes a computer to perform further processing.
[0114] (Note 4) The first object is a person, The second object is a commodity, The identified relationship is associated with a memory unit containing multiple machine learning models that have learned product information, and the machine learning model corresponding to the identified relationship is identified from the multiple machine learning models. By inputting the image of the product represented by the identified second object into the identified machine learning model, product information is identified. The device used by the person indicated by the first object will display the identified product information. The information processing program described in Appendix 1, characterized in that it causes a computer to perform further processing.
[0115] (Note 5) By analyzing the video, the location of people inside the store can be tracked. Based on the tracked location of the person within the store, the device the person is using is identified. The identified terminal will display the identified product information. The information processing program described in Appendix 4, characterized in that it further causes a computer to perform the processing.
[0116] (Note 6) The process for identifying the relationship involves inputting the video into a predetermined model to identify the first region, the second region, and the relationship. The information processing program described in Appendix 1 is characterized in that the predetermined model is a Human Object Interaction Detection (HOID) model in which machine learning has been performed to identify a first class indicating a person purchasing a product and a first region information indicating the region in which the person appears, a second class indicating an object containing a product and a second region information indicating the region in which the object appears, and the interaction between the first class and the second class.
[0117] (Note 7) Obtain the video, By analyzing the acquired video, a first region containing the first object, a second region containing the second object, and the relationship between the interaction between the first object and the second object are identified within the video. Based on the identified relationship, select a model from among several models that is related to the first object or the second object. Output the selected model. An information processing method characterized in that the processing is performed by a computer.
[0118] (Note 8) By analyzing the video, the first object and the second object are tracked, Based on the identified relationship, a machine learning model is selected from among multiple machine learning models to be applied to the second object. By inputting the tracked image of the second object into the selected machine learning model, relevant information about the second object is identified. The relevant information regarding the identified second object is output to the display device associated with the tracked second object. The information processing method described in Appendix 7, characterized in that the processing is further performed by a computer.
[0119] (Note 9) The first object is a person, The second object is a commodity, The system refers to a memory unit that associates the identified relationship with a plurality of customer service models in which customer service content is defined, and identifies a customer service model from the plurality of customer service models that corresponds to the identified relationship. Based on the identified customer service model, the customer service content related to the object represented by the second object for the person represented by the first object is identified. The identified customer service details are sent to the terminal used by the store employee. The information processing method described in Appendix 7, characterized in that the processing is further performed by a computer.
[0120] (Note 10) The first object is a person, The second object is a commodity, The identified relationship is associated with a memory unit containing multiple machine learning models that have learned product information, and the machine learning model corresponding to the identified relationship is identified from the multiple machine learning models. By inputting the image of the product represented by the identified second object into the identified machine learning model, product information is identified. The device used by the person indicated by the first object will display the identified product information. The information processing method described in Appendix 7, characterized in that the processing is further performed by a computer.
[0121] (Note 11) By analyzing the video, the location of people inside the store can be tracked. Based on the tracked location of the person within the store, the device the person is using is identified. The identified terminal will display the identified product information. The information processing method described in Appendix 10, characterized in that the processing is further performed by a computer.
[0122] (Note 12) The process for identifying the relationship involves inputting the video into a predetermined model to identify the first region, the second region, and the relationship. The information processing method according to Appendix 7, characterized in that the predetermined model is a Human Object Interaction Detection (HOID) model in which machine learning has been performed to identify a first class indicating a person purchasing a product and a first region information indicating the region in which the person appears, a second class indicating an object containing a product and a second region information indicating the region in which the object appears, and the interaction between the first class and the second class.
[0123] (Note 13) Obtain the video, By analyzing the acquired video, a first region containing the first object, a second region containing the second object, and the relationship between the interaction between the first object and the second object are identified within the video. Based on the identified relationship, select a model from among several models that is related to the first object or the second object. Output the selected model. An information processing apparatus characterized by having a control unit that performs processing.
[0124] (Note 14) The control unit is, By analyzing the video, the first object and the second object are tracked, Based on the identified relationship, a machine learning model is selected from among multiple machine learning models to be applied to the second object. By inputting the tracked image of the second object into the selected machine learning model, relevant information about the second object is identified. The relevant information regarding the identified second object is output to the display device associated with the tracked second object. The information processing apparatus according to Appendix 13, characterized by further performing processing.
[0125] (Note 15) The control unit is, The first object is a person, The second object is a commodity, The system refers to a memory unit that associates the identified relationship with a plurality of customer service models in which customer service content is defined, and identifies a customer service model from the plurality of customer service models that corresponds to the identified relationship. Based on the identified customer service model, the customer service content related to the object represented by the second object for the person represented by the first object is identified. The identified customer service details are sent to the terminal used by the store employee. The information processing apparatus according to Appendix 13, characterized by further performing processing.
[0126] (Note 16) The control unit is, The first object is a person, The second object is a commodity, The identified relationship is associated with a memory unit containing multiple machine learning models that have learned product information, and the machine learning model corresponding to the identified relationship is identified from the multiple machine learning models. By inputting the image of the product represented by the identified second object into the identified machine learning model, product information is identified. The device used by the person indicated by the first object will display the identified product information. The information processing apparatus according to Appendix 13, characterized by further performing processing.
[0127] (Note 17) The control unit is, By analyzing the video, we can track the location of people within the store. Based on the tracked location of the person within the store, the device the person is using is identified. The identified terminal will display the identified product information. The information processing apparatus according to Appendix 16, characterized by further performing processing.
[0128] (Note 18) The process for identifying the relationship involves inputting the video into a predetermined model to identify the first region, the second region, and the relationship. The information processing device according to Appendix 13, characterized in that the predetermined model is a Human Object Interaction Detection (HOID) model on which machine learning has been performed to identify a first class indicating a person purchasing a product and a first region information indicating the region in which the person appears, a second class indicating an object containing a product and a second region information indicating the region in which the object appears, and the interaction between the first class and the second class. [Explanation of symbols]
[0129] 100,200 Information Processing Devices 110,210 Communications Department 120,220 Input section 130,230 Display section 140,240 storage section 141,241 HOID 142,242 video buffer 143,243 Model Tables 150,250 Control Unit 151,251 Acquisition Department 152,252 Analysis Department 153,253 Specific part 154,254 Learning Department
Claims
1. Acquire video, By analyzing the acquired video footage, a first region containing people, a second region containing products, and the relationships between the interaction between the people and the products are identified within the video footage. Based on the identified relationships, a machine learning model related to the person or product is selected from among multiple machine learning models individually trained for different relationships. Output the selected machine learning model. An information processing program characterized by having a computer perform the processing.
2. By analyzing the video, we can track the person and the product. Based on the identified relationship, a machine learning model to apply to the product is selected from among multiple machine learning models. By inputting the tracked images of the product into the selected machine learning model, relevant information about the product is identified. The relevant information regarding the identified product is output to the display device associated with the tracked product. The information processing program according to claim 1, characterized in that it causes a computer to perform further processing.
3. Referencing a storage unit to which the identified relationship and a plurality of customer service models in which customer service content is defined are associated, and from the plurality of customer service models, identify a customer service model corresponding to the identified relationship, Based on the identified customer service model, the customer service content related to the object represented by the product for the person represented by the person is identified. The identified customer service details are sent to the terminal used by the store employee. The information processing program according to claim 1, characterized in that it causes a computer to perform further processing.
4. Refer to a memory unit that associates the identified relationship with a plurality of machine learning models that have learned product information, and identify a machine learning model from the plurality of machine learning models that corresponds to the identified relationship, By inputting the product image represented by the identified product into the identified machine learning model, product information is identified. The identified product information is displayed on the device used by the aforementioned person. The information processing program according to claim 1, characterized in that it causes a computer to perform further processing.
5. By analyzing the video, we can track the location of people within the store. Based on the tracked location of the person within the store, the device the person is using is identified. The identified terminal will display the identified product information. The information processing program according to claim 4, characterized in that it causes a computer to perform further processing.
6. The process for identifying the relationship involves inputting the video into a predetermined model to identify the first region, the second region, and the relationship. The information processing program according to claim 1, characterized in that the predetermined model is a Human Object Interaction Detection (HOID) model in which machine learning has been performed to identify a first class indicating a person who purchases a product and a first region information indicating a region in which the person appears, a second class indicating an object containing a product and a second region information indicating a region in which the object appears, and the interaction between the first class and the second class.
7. Acquire video, By analyzing the acquired video footage, a first region containing people, a second region containing products, and the relationships between the interaction between the people and the products are identified within the video footage. Based on the identified relationships, a machine learning model related to the person or product is selected from among multiple machine learning models individually trained for different relationships. Output the selected machine learning model. An information processing method characterized in that the processing is performed by a computer.
8. Acquire video, By analyzing the acquired video footage, a first region containing people, a second region containing products, and the relationships between the interaction between the people and the products are identified within the video footage. Based on the identified relationships, a machine learning model related to the person or product is selected from among multiple machine learning models individually trained for different relationships. Output the selected machine learning model. An information processing apparatus characterized by having a control unit that performs processing.