Information processing device and information processing method
The information processing device addresses user skepticism by presenting AI-derived recommendations with accompanying inference basis, enhancing trust and support for purchasing behavior.
Patent Information
- Application Number
- PCT/JP2025/022089
- Authority / Receiving Office
- WO · WO
- Patent Type
- Applications
- Current Assignee / Owner
- Priority Date
- 2024-07-02
- Filing Date
- 2025-06-19
- Publication Date
- 2026-01-08
AI Technical Summary
Users skeptical of AI-based recommendations lack trust in the presented information, leading to insufficient support for their purchasing behavior.
An information processing device that infers user attributes using an AI model and presents both recommended product-related information and the inference basis, enhancing user trust through transparency.
Enhances user trust in AI-based recommendations by providing insight into the reasoning behind product suggestions, thereby improving purchasing behavior support.
Smart Images

Figure JP2025022089_08012026_PF_FP_ABST
Abstract
Description
Information processing device and information processing method
[0001] The present technology relates to an information processing device and method, and in particular to a technology for deriving and presenting recommended information related to commercial products using an AI model.
[0002] There are various technologies for supporting a user's purchase behavior of a product or service (hereinafter referred to as "merchandise"). For example, Patent Literature 1 listed below discloses a technology for deriving products recommended to a user based on an image captured by the user, and presenting images of the derived products to the user as recommended product images.
[0003] Japanese Patent Application Laid-Open No. 2019-207508
[0004] Here, it is conceivable that AI (Artificial Intelligence) technology may be used to derive products recommended to a user. Specifically, an AI model is used to extract user feature information from captured images and infer the user's attributes based on the extracted feature information. Then, using this inference function, purchase history information is constructed in which inferred attribute information for the user is associated with purchase information of the product (identification information for the purchased product and identification information for the sales location). When a product (or its sales location) is recommended, an image of the target user is taken, and the attributes are inferred using the inference function. For example, information on products (or their sales locations) purchased by other users with similar attributes is presented to the target user as recommended information based on the inferred attribute information and the purchase history information in which the attribute information is associated as described above.
[0005] However, some users are skeptical of AI, and for such users, the content of the recommended information may not be trustworthy, which may result in insufficient support for the user's purchasing behavior.
[0006] This technology was developed in consideration of the above-mentioned issues, and aims to improve the reliability of the presented information and the support for user purchasing behavior in systems that use AI models to derive and present recommended information related to products.
[0007] The information processing device according to the present technology includes a feature inference unit that infers feature information of a target user using an AI model based on an image of the target user, an attribute inference unit that infers attributes of the target user using the AI model based on the feature information inferred by the feature inference unit, and a presentation processing unit that inputs recommended product-related information, which is information indicating a product recommended to the target user or a sales location of the product, derived based on purchase history information indicating a product purchase history by a user, the purchase history information being associated with information indicating attributes of the purchasing user, and information indicating the attributes of the target user inferred by the attribute inference unit, and performs processing to present the recommended product-related information and inference basis information indicating an inference basis for the attributes to the target user. According to the above configuration, the target user is presented with not only the recommended product-related information but also inference basis information indicating an inference basis for deriving the recommended product-related information.
[0008] 1 is a diagram illustrating an overview of an information processing system according to an embodiment. FIG. 1 is a front view illustrating an example of the external configuration of an information processing device according to an embodiment. FIG. 2 is a block diagram illustrating an example of the hardware configuration of an information processing device according to an embodiment. FIG. 3 is a block diagram illustrating an example of the configuration of a camera unit included in an information processing device according to an embodiment. FIG. 4 is a block diagram illustrating an example of the hardware configuration of a server device included in an information processing system according to an embodiment. FIG. 5 is a functional block diagram for explaining functions according to an embodiment. FIG. 6 is a diagram illustrating an example of an initial screen. FIG. 7 is a diagram illustrating an example of a recommendation target designation screen. FIG. 8 is a diagram illustrating an example of a feature information designation screen. FIG. 9 is an explanatory diagram of a flow of feature information inference and attribute inference, including setting an AI model. FIG. 10 is a diagram illustrating an example of a derivation result screen. FIG. 11 is a diagram illustrating an example of a display of store guide information. FIG. 12 is a diagram illustrating an example of information presentation in accordance with movement of a movable object. FIG. 13 is a flowchart illustrating an example of a processing procedure for realizing a purchasing support method according to an embodiment. FIG. 14 is a flowchart illustrating an example of a processing procedure for realizing a purchasing support method according to an embodiment. FIG. 15 is a diagram illustrating an example of the configuration of an information processing device according to a modified example corresponding to a case where sensing information from a three-dimensional sensor is used to extract feature information. FIG. 16 is an explanatory diagram of functions of an information processing device according to a modified example.
[0009] Hereinafter, with reference to the accompanying drawings, embodiments of the present technology will be described in the following order: <1. Overview of information processing system> <2. Hardware configuration example of each device> <3. Purchase behavior support method as an embodiment> <4. Processing procedure> <5. Modification example> <6. Summary of embodiment> <7. Present technology>
[0010] 1 is a diagram illustrating an outline of an information processing system including an information processing device 1 according to an embodiment of the present technology. As shown in the figure, the information processing system includes the information processing device 1 and a server device 3.
[0011] In this embodiment, the information processing device 1 is assumed to be a device installed in a merchandise purchasing facility FS, which is a facility where multiple stores are located and where users purchase merchandise, such as a shopping mall or an outlet mall. Here, merchandise is a general term for goods or services, and refers to the object of purchase by users.
[0012] In this example, the information processing device 1 is configured as a digital signage device that is capable of displaying images and presenting various kinds of guidance information related to the facility in response to operations by a user as a customer at the merchandise purchasing facility FS.
[0013] 2 is a front view showing an example of the external configuration of the information processing device 1. The information processing device 1, which serves as a digital signage device, has a display unit 17 on the front side that is capable of displaying images. The information processing device 1 also has a camera unit 23 for capturing an image of a user positioned in front of the device, i.e., a user positioned opposite the screen of the display unit 17. This camera unit 23 is used when providing purchasing behavior support as an embodiment described below.
[0014] In this example, it is assumed that multiple information processing devices 1 as such digital signage devices are installed within the commercial material purchasing facility FS (see FIG. 1 ), but the number of information processing devices 1 installed within the commercial material purchasing facility FS is not limited to multiple and may be one. In other words, it is sufficient for the information processing system to include at least one information processing device 1.
[0015] The server device 3 is a computer device used to manage customer information of the merchandise purchasing facility FS, and is a device that is expected to be used by an administrator of the merchandise purchasing facility FS. The server device 3 is capable of mutual data communication with the information processing device 1 via a network 2, which is a predetermined communication network such as the Internet or a LAN (Local Area Network).
[0016] In the information processing system according to the embodiment, the information processing device 1 has a function of presenting information about recommended products and stores (sales locations) to a user. Specifically, when a user requests information about recommended products and stores (sales locations), the information processing device 1 extracts characteristic information about the user based on an image captured by the camera unit 23 and determines the user's attributes based on the extracted characteristic information. The server device 3 according to the embodiment stores, for each user, purchase history information about users who have used a product purchasing facility FS, in which information about purchased products and their sales locations is associated with information indicating the user's attributes (purchase history information Ip, described below). Upon receiving attribute information about a user who has requested the presentation of recommended information from the information processing device 1, the server device 3 derives information about recommended products or their sales locations, based on the purchase history information associated with the attribute information as described above, as information about recommended products or their sales locations purchased by users with attributes identical or similar to the received attributes. The information processing device 1 presents the information on the product or its sales location derived by the server device 3 to the user via the display unit 17. This realizes purchasing behavior support in the form of presenting information on recommended products and sales locations to a user who has requested the presentation of recommended information.
[0017] In the following description, a user who is the target of receiving information about recommended products and sales locations from the information processing device 1 will be referred to as a "target user."
[0018] 3 is a block diagram showing an example of the hardware configuration of the information processing device 1. As shown in the figure, the information processing device 1 includes a CPU (Central Processing Unit) 11. The CPU 11 executes various processes in accordance with programs stored in a ROM (Read Only Memory) 12 or programs loaded from a storage unit 19 into a RAM (Random Access Memory) 13. The RAM 13 also stores data necessary for the CPU 11 to execute various processes as appropriate.
[0019] The AI processing unit 14 is configured with a programmable arithmetic processing device, such as a CPU, FPGA (Field Programmable Gate Array), or DSP (Digital Signal Processor), and performs inference processing using an AI model. In this example, the AI processing unit 14 performs processing to infer attributes of a target user based on feature information of the target user extracted from images captured by the camera unit 23. Specifically, in this example, a self-organizing map (SOM) is used as the AI model for inferring such attributes. As is well known, an SOM AI model is an AI model that can classify the attributes of a target through unsupervised learning. When information indicating the features of the target is given as input, the model outputs position information indicating the attributes of the target in a map space. Here, the map space is a space in which attributes are defined by coordinate values. Note that the space referred to here is primarily assumed to be a two-dimensional space, but can also include one-dimensional, three-dimensional, or higher-dimensional spaces. In this example, the map space of the SOM is assumed to be a two-dimensional space.
[0020] The AI processing unit 14 is configured to be able to switch the AI model to be used.
[0021] The CPU 11, ROM 12, RAM 13, and AI processing unit 14 are interconnected via a bus BS, to which an input / output interface (I / F) 15 is also connected.
[0022] An input unit 16 consisting of operators and operation devices is connected to the input / output interface 15. For example, the input unit 16 may be various operators and operation devices such as a keyboard, a mouse, keys, a dial, a touch panel, a touch pad, a remote controller, etc. The input unit 16 detects user operations, and the CPU 11 interprets signals corresponding to the input operations.
[0023] The input / output interface 15 is also connected to a display unit 17 such as an LCD (Liquid Crystal Display) or an organic EL (Electro-Luminescence) panel, and an audio output unit 18 such as a speaker.
[0024] The display unit 17 is configured as a display device capable of displaying images. The display unit 17 displays various images on a display screen based on instructions from the CPU 11. For example, the display unit 17 displays various operation menus, icons, messages, etc., i.e., a GUI (Graphical User Interface), based on instructions from the CPU 11.
[0025] The input / output interface 15 is connected to a storage unit 19 configured with an SSD (Solid State Drive) or an HDD (Hard Disk Drive) and a communication unit 20 configured with a modem or the like.
[0026] The storage unit 19 stores various data used by the CPU 11 for processing. In the information processing device 1 of this example, the storage unit 19 stores data of multiple AI models that the AI processing unit 14 can use to infer attributes as an attribute inference model data group MDa. Here, the attribute inference model data group MDa stores AI model data for multiple AI models with different attribute inference processing content as AI model data for constructing an AI model that infers attributes (SOM in this example: hereinafter referred to as "attribute inference model"). For example, if the AI model has a neural network such as a CNN (Convolutional Neural Network), the AI model data here corresponds to parameters indicating the structure of the neural network and parameters as filter coefficients used in convolution processing, etc. Note that the type of AI model data stored in the attribute inference model data group MDa will be described later.
[0027] The communication unit 20 performs communication processing via a transmission path such as the Internet, and communication with various devices via wired / wireless communication, bus communication, and the like.
[0028] A drive 21 is also connected to the input / output interface 15 as required, and a removable recording medium 22 such as a magnetic disk, optical disk, magneto-optical disk, or semiconductor memory is appropriately loaded therein.
[0029] The drive 21 can read data files such as programs used in various processes from the removable recording medium 22. The read data files are stored in the storage unit 19 and used in the processing of the CPU 11. The computer programs read from the removable recording medium 22 are installed in the storage unit 19 as needed.
[0030] Furthermore, the above-mentioned camera unit 23 is connected to the input / output interface 15. Details of the camera unit 23 will be described later with reference to FIG.
[0031] In the information processing device 1 having the hardware configuration described above, for example, software for the processing of this embodiment can be installed via network communication by the communication unit 20 or via the removable recording medium 22. Alternatively, the software may be stored in advance in the ROM 12, the storage unit 19, etc. The CPU 11 performs processing operations based on various programs, thereby executing the information processing and communication processing required by the information processing device 1.
[0032] 4 is a block diagram showing an example of the configuration of the camera unit 23 included in the information processing device 1. As shown in the figure, the camera unit 23 includes an image sensor 24, an imaging optical system 25, an optical system driving unit 26, a camera control unit 27, a memory unit 28, and a communication unit 29. The image sensor 24, the camera control unit 27, the memory unit 28, and the communication unit 29 are connected via a bus 30, and are capable of mutual data communication.
[0033] In this example, the image sensor 24 is configured as a solid-state imaging element such as a CCD (Charge Coupled Device) type or a CMOS (Complementary Metal Oxide Semiconductor) type, and receives light incident through the imaging optical system 25 to obtain an image of the subject.
[0034] The imaging optical system 25 includes lenses such as a cover lens, a zoom lens, and a focus lens, as well as an iris mechanism. Light (incident light) from a subject is guided by the imaging optical system 25 and collected on the light receiving surface of the image sensor 24.
[0035] The optical system driving unit 26 collectively refers to the driving units for the zoom lens, focus lens, and diaphragm mechanism of the imaging optical system 25. Specifically, the optical system driving unit 26 has actuators for driving the zoom lens, focus lens, and diaphragm mechanism, respectively, and driving circuits for the actuators.
[0036] The camera control unit 27 is configured with, for example, a microcomputer having a CPU, ROM, and RAM, and performs overall control of the camera unit 23 by the CPU executing various processes according to programs stored in the ROM or programs loaded into the RAM.
[0037] The camera control unit 27 issues drive instructions to the optical system drive unit 26 to drive the zoom lens, focus lens, diaphragm mechanism, etc. In response to these drive instructions, the optical system drive unit 26 moves the focus lens and zoom lens, opens and closes the diaphragm blades of the diaphragm mechanism, etc.
[0038] The camera control unit 27 also controls the writing and reading of various data to and from the memory unit 28. The memory unit 28 is a non-volatile storage device such as an HDD or a flash memory device, and is used to store data used by the camera control unit 27 to execute various processes. The memory unit 28 can also be used as a storage destination (recording destination) for image data (captured image data) output from the image sensor 24.
[0039] The camera control unit 27 also performs various data communications with external devices via the communication unit 29. The communication unit 29 in this example is configured to be capable of communication via the input / output interface 15 shown in Fig. 3, and to be capable of performing data communications with external devices connected via the input / output interface 15, particularly with the CPU 11 in this example.
[0040] As shown in the figure, the image sensor 24 comprises an imaging unit 41, an image signal processing unit 42, an internal sensor control unit 43, an AI processing unit 44, a memory unit 45, a computer vision processing unit 46, and a communication interface (I / F) 47, each of which is capable of mutual data communication via a bus 48.
[0041] The imaging unit 41 includes a pixel array unit in which pixels having photoelectric conversion elements such as photodiodes are arranged two-dimensionally, and a readout circuit that reads out electrical signals obtained by photoelectric conversion from each pixel included in the pixel array unit. The readout circuit performs, for example, CDS (Correlated Double Sampling) processing, AGC (Automatic Gain Control) processing, etc. on the electrical signals obtained by photoelectric conversion, and further performs A / D (Analog / Digital) conversion processing.
[0042] The image signal processing unit 42 performs preprocessing, synchronization, YC generation, resolution conversion, codec processing, and other processes on the captured image signal as digital data after A / D conversion. Preprocessing includes clamping the R (red), G (green), and B (blue) black levels of the captured image signal to a predetermined level, and correction between the R, G, and B color channels. Preprocessing can also include brightness adjustments such as gamma correction and color adjustments such as white balance adjustments. Synchronization involves color separation, ensuring that each pixel contains all R, G, and B color components. For example, in the case of an image sensor using a Bayer color filter, demosaic processing is performed as the color separation process. YC generation generates (separates) a luminance (Y) signal and a color (C) signal from R, G, and B image data. Resolution conversion involves performing resolution conversion on image data that has undergone various signal processes. In codec processing, the image data that has undergone the various processes described above is subjected to encoding processing and file generation for recording or communication, for example. In codec processing, it is possible to generate moving image file formats such as MPEG-2 (Moving Picture Experts Group) and H.264. It is also possible to generate still image file formats such as JPEG (Joint Photographic Experts Group), TIFF (Tagged Image File Format), and GIF (Graphics Interchange Format).
[0043] The sensor control unit 43 is configured with a microcomputer having, for example, a CPU, ROM, RAM, etc., and comprehensively controls the operation of the image sensor 24. For example, the sensor control unit 43 issues instructions to the imaging unit 41 to control the execution of imaging operations. It also controls the image signal processing unit 42 to execute processing.
[0044] The AI processing unit 44 is configured to have a programmable arithmetic processing device such as a CPU, FPGA, or DSP (Digital Signal Processor), and performs inference processing on the captured image using an AI model. Similar to the AI processing unit 14 described above, the AI processing unit 44 in this example is configured to be able to switch the AI model to be used. The AI processing unit 44 in this example performs processing to infer characteristic information of the target user based on the captured image, and details will be described again later.
[0045] The memory unit 45 is configured with a volatile memory and is used to hold (temporarily store) data necessary for the AI processing unit 44 to perform inference processing. Specifically, the memory unit 45 stores AI model data of multiple AI models that the AI processing unit 44 can use to infer feature information as a feature inference model data group MDc. Details of this feature inference model data group MDc will be explained again later.
[0046] The computer vision processing unit 46 performs rule-based image processing on the captured image data, such as super-resolution processing.
[0047] The communication interface 47 is an interface for communicating with each unit connected via the bus 30, such as the camera control unit 27 and memory unit 28, outside the image sensor 24. Through this communication interface 47, it is possible to output information such as the result of the inference processing by the AI processing unit 44 to the outside of the image sensor 24. Furthermore, the intra-sensor control unit 43 is able to communicate various data with, for example, the camera control unit 27 and devices outside the image sensor 24, through this communication interface 47.
[0048] Figure 5 is a block diagram showing an example of the hardware configuration of the server device 3. The difference from the hardware configuration of the information processing device 1 shown in Figure 3 is that the AI processing unit 14 is omitted. In addition, purchase history information Ip is stored in the memory unit 19 of the server device 3. As mentioned above, the purchase history information Ip is information indicating the purchase history of users who have used the product purchasing facility FS, and is information in which, for each user, information on the purchased products and their sales locations is associated with information indicating the attributes of the user.
[0049] Here, in order to construct the above-described purchase history information Ip, it is necessary to be able to acquire purchase information indicating which products a user, whose attributes have been inferred by the attribute inference process performed by the AI processing unit 14, has purchased at which sales location. Various methods are conceivable for acquiring such purchase information. For example, a surveillance camera installed at the merchandise purchasing facility FS may be used to monitor the behavior of the user, whose attributes have been inferred by the AI processing unit 14, and information on the merchandise purchased by the user and the store (sales location) where the merchandise was purchased may be identified to generate purchase information. Alternatively, a payment app for making payments at the merchandise purchasing facility FS may be installed on the user's terminal (e.g., a smartphone, tablet, etc.). For a user, whose attributes have been inferred by the AI processing unit 14, the information processing device 1 may transmit information on the inferred attributes to the user's terminal. When a payment for the purchase of a merchandise is made, the payment app may upload information on the purchased merchandise and the store where the purchase was made, as well as the user's attribute information, to the server device 3. As such, various methods are conceivable for acquiring user purchase information to construct the purchase history information Ip, and the present invention is not limited to a specific method.
[0050] Here, the CPU included in the server device 3 will be referred to as a “CPU 31 ” to distinguish it from the CPU 11 included in the information processing device 1 .
[0051] The server device 3 is not limited to being configured as a single computer device as shown in Fig. 5, but may be configured as a system of multiple computer devices. The multiple computer devices may be systemized using a LAN or the like, or may be located in a remote location via a VPN (Virtual Private Network) using the Internet or the like. The multiple computer devices may include computer devices as a server group (cloud) available through a cloud computing service.
[0052] <3. Purchasing behavior support method as an embodiment> As can be understood from the above explanation, the information processing system of this embodiment derives and presents recommended information related to merchandise using an AI model based on an image of a user. However, as mentioned above, some users are skeptical of AI, and for such users, it may be difficult to give confidence in the content of the recommended information presented, which may result in insufficient support for the user's purchasing behavior.
[0053] Therefore, in this embodiment, inference basis information showing the basis of the inference process performed during the derivation process of recommended information using an AI model is presented to the user, thereby aiming to improve the trust that the target user has in the recommended information.
[0054] 6 is a functional block diagram for explaining functions of an embodiment of the CPU 11 of the information processing device 1. As shown in the figure, the CPU 11 has functions as a presentation processing unit F1, a model setting processing unit F2, and a relearning processing unit F3.
[0055] The presentation processing unit F1 inputs recommended product-related information, which is information indicating products recommended to the target user or the sales locations of the products, derived based on the purchase history information Ip and information indicating the attributes of the target user estimated by the AI processing unit 14, and performs processing to present the recommended product-related information and inference basis information indicating the basis for inferring the attributes to the target user.
[0056] In order to explain the details of the processing of this presentation processing unit F1, screen transitions and user operations until the target user is presented with recommended product related information will be described with reference to FIGS.
[0057] 7 shows an example of an initial screen G1 displayed on the display unit 17. As shown, the initial screen G1 includes buttons for instructing the presentation of floor guide information for the merchandise purchasing facility FS, buttons for instructing the presentation of discount information, and the like. The initial screen G1 in this example also includes a recommendation function button Br for calling up a recommendation function that derives and presents information related to recommended merchandise. To call up the recommendation function, the user operates the recommendation function button Br.
[0058] In the following description, it is assumed that operations on the screen are performed by touch operations. However, the operations on the screen are not limited to touch operations, and may also be operations using controls for, for example, moving a cursor or performing a selection operation.
[0059] 8 shows an example of a recommendation target specification screen G2 displayed on the display unit 17 in response to the operation of the recommendation function button Br. The recommendation target specification screen G2 is a screen for specifying the type of merchandise to be recommended. In the following description, the merchandise is assumed to be a product. Furthermore, the recommended merchandise-related information is assumed to present information about the product and the store that sells the product.
[0060] As shown in the figure, the recommendation target specification screen G2 displays information indicating the types of products that can be specified as recommendation targets, and displays check boxes cb for each of the product types. The user can specify the product types to be recommended by checking at least one of the check boxes cb.
[0061] The recommendation target designation screen G2 also has a captured image display area Ad for displaying an image captured by the camera unit 23. Specifically, the captured image display area Ad displays an image (a real-time captured image) as a moving image currently being captured by the camera unit 23.
[0062] The recommendation target specification screen G2 displays a back button B1 and an OK button B2. When the back button B1 is operated, the CPU 11 performs a process of returning the display screen of the display unit 17 to the initial screen G1 shown in FIG.
[0063] On the other hand, when the OK button B2 is operated, the CPU 11 displays a characteristic information specification screen G3, as illustrated in FIG. 9 , on the display unit 17. The characteristic information specification screen G3 is a screen for specifying the type of characteristic information the user is permitted to use in the process for deriving recommended product-related information to be presented to the user. For the sake of explanation, three types of characteristic information are exemplified here: "clothing," "gender," and "age." The "clothing" characteristic information may be information indicating the style and color of the clothing worn by the person as the subject, such as "a black long skirt and a white T-shirt" or "a pink dress and a straw hat." The "gender" characteristic information may be information indicating whether the person is male or female. The "age" characteristic information may be information about the person's age or an age group, such as their 20s or 30s.
[0064] The characteristic information specification screen G3 displays information indicating the types of characteristic information that can be used, such as the above-mentioned "clothing," "gender," and "age," as well as check boxes cb for specifying them. The characteristic information specification screen G3 also has a captured image display area Ad, similar to the recommendation target specification screen G2 shown in FIG. 8, as well as a back button B3 and a start button B4.
[0065] The user can specify the characteristic information that is permitted to be used by checking at least one of the check boxes cb.
[0066] When the back button B3 is operated on the characteristic information specification screen G3, the CPU 11 performs processing to return the display screen of the display unit 17 to the recommendation target specification screen G2 shown in FIG.
[0067] In addition, when the start button B4 is operated on the characteristic information specification screen G3, the CPU 11 processes the information to derive recommended product-related information in accordance with the specification of the recommended target on the recommended target specification screen G2 and the specification of the characteristic information on the characteristic information specification screen G3.
[0068] Specifically, the CPU 11 first performs processing to have the AI processing unit 44 infer the characteristic information specified on the characteristic information specification screen G3. That is, for example, if only "clothing" is selected, the CPU 11 performs processing to have the AI processing unit 44 infer the characteristic information of the clothing, and if "clothing" and "gender" are specified, the CPU 11 performs processing to have the AI processing unit 44 infer the characteristic information of the clothing and the characteristic information of the gender.
[0069] In the information processing device 1 of this example, the above-mentioned feature inference model data group MDc stores feature inference model data for each combination of usable feature information as AI model data (hereinafter referred to as "feature inference model data") for realizing an AI model that infers feature information. Specifically, in this example, a total of seven types of feature inference model data are stored: feature inference model data for inferring only feature information of "clothing," feature inference model data for inferring only feature information of "gender," feature inference model data for inferring only feature information of "age," feature inference model data for inferring feature information for each of "clothing" and "gender," feature inference model data for inferring feature information for each of "clothing" and "age," feature inference model data for inferring feature information for each of "gender" and "age," and feature inference model data for inferring feature information for each of "clothing," "gender," and "age."
[0070] The CPU 11 controls the intra-sensor control unit 43 to select feature inference model data corresponding to the combination of feature information specified by the user from the feature inference model data stored as the feature inference model data group MDc, and to execute setting processing of the AI model used by the AI processing unit 44 in accordance with the selected feature inference model data. This allows the AI processing unit 44 to execute inference processing of the feature information specified by the user.
[0071] In addition, in response to having the AI processing unit 44 execute an inference process for the feature information specified by the user, the CPU 11 causes the AI processing unit 14 to execute an inference process for the attribute using the feature information inferred by the AI processing unit 44 as input.
[0072] Here, the SOM used for attribute inference cannot correctly infer attributes unless the combination of types of feature information provided as input is a predetermined combination. That is, for example, an SOM trained by machine learning using feature information of "clothing" and "gender" as input data cannot correctly infer attributes even if feature information of "gender" and "age" is provided as input data. For this reason, in this example, the attribute inference model data group MDa stores multiple attribute inference model data trained by machine learning for each combination of usable feature information. Specifically, a total of seven types of attribute inference model data are stored, each trained by machine learning using each of the seven combinations of feature information exemplified above as input data.
[0073] In response to having the AI processing unit 44 execute an inference process for feature information specified by the user, the CPU 11 selects attribute inference model data corresponding to the combination of feature information specified by the user from the attribute inference model data group MDa and executes a process for setting an AI model as an SOM to be used by the AI processing unit 14 in accordance with the selected attribute inference model data. Then, the CPU 11 causes the AI processing unit 14 to execute an attribute inference process using the feature information inferred by the AI processing unit 44 as input data. As a result, an attribute inference process based on the features of the target user imaged by the camera unit 23 is executed.
[0074] For confirmation, Fig. 10 shows an explanatory diagram of the flow of feature information inference and attribute inference, including the setting of the AI model described above. As described above, when inferring feature information and attributes, an AI model setting process is performed in accordance with user-specified information, and this AI model setting process is performed by the model setting processing unit F2 shown in Fig. 6. In response to the user's specification of feature information to be used, the model setting processing unit F2 causes the intra-sensor control unit 43 to select feature inference model data corresponding to the specified combination of feature information from the feature inference model data group MDc stored in the memory unit 45, and causes the AI processing unit 44 to set a feature inference model in accordance with the selected feature inference model data.
[0075] In addition, the model setting processing unit F2 selects attribute inference model data corresponding to the combination of specified feature information from the attribute inference model data group MDa stored in the memory unit 19, and sets an attribute inference model in the AI processing unit 14 according to the selected attribute inference model data.
[0076] Then, the CPU 11 issues an instruction to the intra-sensor control unit 43 to cause the AI processing unit 44 to execute inference processing of the feature information. The CPU 11 also causes the AI processing unit 44 to execute inference processing of the attribute using the feature information inferred by the AI processing unit 44 as an input.
[0077] When inferring feature information, if multiple pieces of feature information are to be inferred, a configuration can be adopted in which each piece of feature information is inferred using a separate AI model. For example, an AI model for inferring feature information about "clothing," an AI model for inferring feature information about "gender," and an AI model for inferring feature information about "age" can be used separately. In this case, it is conceivable to provide multiple AI processing units 44 for setting each AI model. Alternatively, it is conceivable to adopt a configuration in which a single AI processing unit 44 performs inference of different pieces of feature information in a time-sharing manner.
[0078] Although the above examples of characteristic information include "clothing," "gender," and "age," it is also possible to use other characteristic information, such as characteristic information related to face type, such as round face or square face.
[0079] As described above, in response to operation of the start button B4 on the characteristic information specification screen G3 shown in Figure 9, the CPU 11 causes the AI processing unit 44 to execute an inference process for the characteristic information of the target user, and also causes the AI processing unit 14 to execute an inference process for attributes using the inferred characteristic information as input data.
[0080] In response to the attribute inference results obtained by the AI processing unit 14, the CPU 11 controls the derivation of recommended product-related information based on the inferred attribute information and the above-mentioned purchase history information Ip. Specifically, the CPU 11 transmits the inferred attribute information to the server device 3 and instructs the execution of a process for deriving recommended product-related information based on the transmitted attribute information and the purchase history information Ip. In this example, in order to present recommended product-related information related to products of the type specified on the recommendation target specification screen G2 of FIG. 8 as the recommended product-related information, the CPU 11 transmits information indicating the type of the specified product together with the inferred attribute information to the server device 3.
[0081] Here, the process of deriving recommended product-related information may be performed as a process of deriving information on products purchased by other users whose attributes match or are similar to those of the target user, or information on the sales locations of the purchased products, based on the purchase history information Ip. In this example, the server device 3 (CPU 31) derives recommended product-related information by deriving, as recommended product-related information, products purchased by other users whose attributes match or are similar to those of the target user and belong to a category specified by the target user, as well as information on stores selling the products. In this example, the derivation process selects multiple other users whose attributes match or are similar to those of the target user, and derives, as recommended product-related information, information on products of the corresponding category purchased by each of the other users and the stores selling the products. In other words, multiple pieces of recommended product-related information are derived. Note that deriving multiple pieces of recommended product-related information is merely an example, and a single piece of recommended product-related information may also be derived.
[0082] In the information processing device 1, the presentation processing unit F1 inputs the recommended product-related information derived in this manner in the server device 3, and performs a process of presenting the recommended product-related information and inference basis information Ir that indicates the basis for attribute inference by the AI processing unit 14 to the target user.
[0083] 11 shows an example of a derivation result screen G4 for presenting recommended product-related information and inference basis information Ir to a target user. The derivation result screen G4 includes a recommended information display area Ar for displaying the derived recommended product-related information. As shown in the figure, the recommended information display area Ar in this example displays multiple pieces of recommended product-related information derived by the server device 3.
[0084] Here, the derivation result screen G4 may display store guide information Ig for guiding the location of recommended stores, as well as the recommended product-related information, as exemplified in Fig. 12. Specifically, the store guide information Ig may display, for example, a map of each floor of the product purchasing facility FS and information indicating the location of the recommended stores on the map.
[0085] 11, the derivation result screen G4 displays an attribute map Ma as shown. Specifically, the presentation processing unit F1 performs processing to present the attribute map Ma and information indicating the position of the target user's attributes in the map space. Here, the attribute map Ma is information that visualizes the map space of the SOM used for attribute estimation.
[0086] In this example, information indicating the location of the target user's attributes on the map space is presented by presenting a target user characteristic image Ic, which is an image showing the characteristics of the target user, at a location on the map space that represents the target user's attributes. FIG. 11 shows an example screen corresponding to a case where only "clothing" is specified as the characteristic information to be used on the characteristic information specification screen G3. In this case, the target user characteristic image Ic is an image obtained by visualizing the characteristic information inferred about the target user. For example, if the characteristic information of the target user's clothing is "pink dress," an image showing the "pink dress" is displayed as the target user characteristic image Ic on the attribute map Ma.
[0087] Additionally, on the attribute map Ma, other user characteristic images Ia, which are images showing the characteristic information of other users, are displayed at positions representing the attributes of those other users. The other users referred to here are users other than the target user whose purchase information is stored in the purchase history information Ip. As with the target user characteristic image Ic, these other user characteristic images Ia are displayed as images corresponding to the characteristic information specified on the characteristic information specification screen G3. For example, if characteristic information for "clothing" is specified and the characteristic information of the target other user is "jeans and a white T-shirt," an image showing the "jeans and a white T-shirt" is displayed as the other user characteristic image Ia for the target other user. To realize the display of such other user characteristic images Ia, for example, the purchase history information Ip may be associated with not only attribute information but also characteristic information, so that the presentation processing unit F1 (CPU 11) can acquire the characteristic information of other users when displaying the other user characteristic image Ia.
[0088] As described above, by displaying the target user characteristic image Ic and the other user characteristic image Ia on the attribute map Ma, the target user can intuitively understand his or her own attributes. In particular, by displaying the characteristic information of the target user and other users as images, it is easier to understand what attributes the target user has been classified into.
[0089] Furthermore, in the derivation result screen G4 of this example, the derived recommended product-related information is also displayed as recommendation information Ri on the attribute map Ma. This recommendation information Ri is displayed as the other user who served as the basis for deriving the recommended product-related information, among the other users whose positions (attributes) in the map space are indicated by the other user characteristic image Ia. In other words, if another user "A" is selected as having attributes similar to those of the target user, and information on the products and stores purchased by "A" is derived as recommended product-related information, the recommendation information Ri indicating the recommended product-related information is displayed to indicate the position of "A" on the attribute map Ma. Specifically, in this example, the recommendation information Ri of the recommended product-related information that served as the basis for derivation by "A" is displayed in a manner that points to the other user characteristic image Ia of "A" on the attribute map Ma (the figure shows an example display in which the other user characteristic image Ia is pointed to in the form of a speech bubble).
[0090] Furthermore, in the derivation result screen G4 in this embodiment, inference basis information Ir is presented. Specifically, the presentation processing unit F1 performs processing to present, as the inference basis information Ir, information indicating an imaging target that the AI model focused on during the process of inferring attributes from among imaging targets captured in a captured image. In this example, the inference basis information Ir presented is information indicating an imaging target that the feature inference model that infers feature information focused on during the inference process. Specifically, in this example, information indicating an image area of the imaging target is presented as the information indicating the imaging target. Hereinafter, the image area of the imaging target that the feature inference model focused on during the inference process will be referred to as the "area of interest."
[0091] Here, the region of interest determined by the feature inference model can be said to be the image region of an object that is the detection target in the object detection process that the feature inference model performs to infer feature information about a target event. For example, if the feature inference model infers feature information about "clothing," the object detection process will be performed with the clothing part as the detection target, and therefore the image region of the object detected by the object detection process will be the region of interest. Alternatively, if the feature inference model infers feature information about "gender" or "age," the object detection process will be performed with the face part as the detection target, and therefore the image region of the object detected by the object detection process will be the region of interest.
[0092] The presentation processing unit F1 in this example performs processing to display information indicating such an area of interest as inference basis information Ir in the captured image display area Ad on the derivation result screen G4. The figure shows an example of displaying inference basis information Ir corresponding to the case where "clothing" is specified as the feature information to be used and the area of interest is the skirt part.
[0093] By presenting the inference basis information Ir as described above, it is possible to improve the trust that the target user has in the presented recommended product-related information. In particular, in this example, the inference basis information Ir presents information indicating the image object that the feature inference model focused on during the inference process. For example, if only "clothing" is specified as the feature information to be used, it is indicated that only the clothing part, not the face part, was used as the basis for the inference, thereby improving the trust that the target user has. Furthermore, by presenting information indicating the image area of the image object that was focused on as the inference basis information Ir, the target user can intuitively understand which part of the target user was focused on when making the inference.
[0094] The inference basis information Ir is not limited to information indicating an image area. For example, when "clothing" is specified, text information indicating "clothing" is presented as the inference basis information Ir, and when "clothing" and "gender" are specified, text information indicating "clothing" and "gender" is presented as the inference basis information Ir. The inference basis information Ir is not limited to presentation of information indicating an image area.
[0095] Furthermore, as for the information on the region of interest, instead of information on the image region of the target object detected by the object detection process as described above, it is also possible to use information on the region of interest that is used as information indicating the basis for inference in the field of so-called explainable AI. For example, it is possible to use information on the region of interest detected by CAM (Class Activation Mapping), Grad-CAM, ABN (Attention Branch Network), etc.
[0096] In the derivation result screen G4 of this embodiment, the target user characteristic image Ic on the attribute map Ma, i.e., information indicating the position of the target user's attributes, is presented as a movable object, which is a display object that can be moved on the attribute map Ma. When the target user characteristic image Ic as a movable object is moved by an operation, the presentation processing unit F1 of this embodiment performs processing to present information indicating a product purchased by a user with an attribute corresponding to the moved position or a sales location of the product.
[0097] 13A and 13B show an example of information presentation in response to movement of the target user characteristic image Ic. As shown in FIGS. 13A and 13B, the target user characteristic image Ic can be moved on the attribute map Ma. The reason why the target user characteristic image Ic can be moved on the attribute map Ma in this way is to allow the target user to modify the attributes. In other words, if the target user wants to modify the attributes inferred by the attribute inference model, the target user can simply move the target user characteristic image Ic on the attribute map Ma.
[0098] In this example, as guide information for performing such attribute correction operations, information indicating products purchased by users with attributes corresponding to the destination location or the retail locations of those products is presented, as shown as attribute guide information Gi in the figure. Specifically, in this example, the attribute guide information Gi presents information indicating products purchased by users with attributes matching or similar to the attributes of the destination location and the retail stores selling those products. In this case, the presentation processing unit F1 transmits information indicating the attributes of the destination location to the server device 3 in response to a movement operation of the target user characteristic image Ic, and causes the server device 3 to identify users with attributes matching or similar to the attributes as destination attribute users from among other users managed in the purchase history information Ip. The server device 3 then identifies the products purchased by the destination attribute user and the retail stores selling those products, and performs processing to display information on the identified products and retail stores on the attribute map Ma as attribute guide information Gi. At this time, the attribute guide information Gi is provided in the same manner as the recommendation information Ri described above, by displaying the relevant other users on the attribute map Ma (the figure shows an example of a display in which the relevant other user's other user characteristic image Ia is indicated in the form of a speech bubble, similar to the recommendation information Ri).
[0099] By presenting the attribute guide information Gi as described above, when a target user moves the target user characteristic image Ic on the attribute map Ma to correct his or her own attribute to a different attribute from the inferred attribute, the target user can search for the location where the correction should be made by using information about the product or retail store presented at the destination of the target user characteristic image Ic. This makes it easier to correct the attribute.
[0100] In the example of Figure 13, multiple destination attribute users are identified and attribute guide information Gi is presented for each of these destination attribute users.This increases the amount of guide information that the target user can refer to when performing attribute modification operations, thereby improving the guide performance for attribute modification and making it easier to perform attribute modification operations.
[0101] Although not illustrated, the presentation processing unit F1 performs a process of presenting product-related information corresponding to the attributes of the destination location upon completion of the movement operation of the target user characteristic image Ic. Specifically, the presentation processing unit F1 causes the server device 3 to identify products and stores selling those products that have been purchased by other users with attributes that match or are similar to the attributes of the destination location, and displays information about the identified products and stores in the recommended information display area Ar.
[0102] The completion of the movement operation may be, for example, the timing when the finger is removed from the screen if the movement operation is performed by a touch operation, or when a button such as "movement complete" is operated if the button is provided.
[0103] In the information processing device 1 of this embodiment, assuming that the attribute modification operation described above is possible, a relearning processing unit F3 is provided as a function of the CPU 11. When a movement operation of the target user characteristic image Ic is performed, the relearning processing unit F3 performs a relearning process of the AI model (attribute inference model) used by the AI processing unit 14, using information indicating the attributes of the destination position of the target user characteristic image Ic as training data. Specifically, in response to an operation of the end button B5 on the derivation result screen G4 (i.e., an operation instructing the end of the recommendation process), the relearning processing unit F3 performs a process of saving information indicating the attributes of the destination position (in this example, coordinate information on the map space) as training data. The training data is saved in, for example, the memory unit 19. Then, using the saved training data, the relearning processing unit F3 performs a relearning process on the corresponding attribute inference model (the attribute inference model used in the process of deriving the current recommended product-related information) among the attribute inference models whose attribute inference model data is stored in the attribute inference model data group MDa. This re-learning process is performed as a process of adjusting the parameters for inference so that when the same feature information as this time is input to the corresponding attribute inference model, the same attribute as the correct data is output as the inference result.
[0104] Furthermore, if the target user feature image Ic has not been moved, the relearning processing unit F3 of this example performs a relearning process on the attribute inference model (the AI model used by the AI processing unit 14) using, as training data, information on the attributes inferred by the AI processing unit 14. Specifically, in response to an operation of the end button B5 on the derivation result screen G4, the relearning processing unit F3 performs a process of saving, as training data, information indicating the attributes inferred by the AI processing unit 14 in, for example, the storage unit 19, and also performs a relearning process on the corresponding attribute inference model using the saved training data.
[0105] By performing the above-described re-learning process, the attribute inference model can be updated so that attribute inference can be performed more correctly, thereby improving attribute inference performance.
[0106] 14 and 15 are flowcharts showing an example of a processing procedure for realizing the purchase support method according to the embodiment described above. The processing shown in Fig. 14 and 15 is executed by the CPU 11 of the information processing device 1 based on a program stored in a predetermined storage device such as the ROM 12. It should be noted that when the processing shown in Fig. 14 and 15 is executed, the initial screen G1 is displayed on the display unit 17.
[0107] First, in step S101, the CPU 11 waits for an operation to request use of the recommendation function, specifically, for the operation of the recommendation function button Br.
[0108] If it is determined that the recommendation function button Br has been operated and an operation to request use of the recommendation function has been performed, the CPU 11 proceeds to step S102 and performs processing to display the recommendation target designation screen G2.
[0109] In step S103 following step S102, the CPU 11 waits for the completion of designation of the recommendation target, specifically, for the operation of the above-mentioned OK button B2.
[0110] If it is determined that the OK button B2 has been operated and the specification of the recommendation target has been completed, the CPU 11 proceeds to step S104 and performs processing to display the characteristic information specification screen G3.
[0111] In step S105 following step S104, the CPU 11 waits for the completion of designation of the characteristic information, that is, waits for the operation of the start button B4.
[0112] If it is determined that the start button B4 has been operated and the specification of the feature information has been completed, the CPU 11 proceeds to step S106 and executes a process for setting an AI model according to the combination of the specified feature information. Note that the details of this AI model setting process have already been explained with reference to FIG. 10, so a duplicate explanation will be avoided.
[0113] In step S107 following step S106, the CPU 11 instructs the intra-sensor control unit 43 to execute a feature inference process for the target user. That is, the CPU 11 instructs the intra-sensor control unit 43 to cause the AI processing unit 44 to execute a feature information inference process using the image captured by the imaging unit 41 as input data (i.e., a feature information inference process for the target user).
[0114] In step S108 following step S107, the CPU 11 acquires the characteristic information of the target user, that is, the characteristic information of the target user inferred by the inference process executed in step S107.
[0115] In step S109 following step S108, the CPU 11 instructs the AI processing unit 14 to execute attribute inference processing using the acquired feature information as input data. That is, the CPU 11 instructs the AI processing unit 14 to execute attribute inference processing using the feature information acquired in step S108 as input data.
[0116] In step S110 following step S109, the CPU 11 transmits information about the specified recommendation target and the inferred attribute information to the server device 3, and causes the server device 3 to execute a process for deriving recommended product-related information. That is, the CPU 11 transmits information indicating the type of product specified on the recommendation target specification screen G2 and information indicating the attributes inferred by the inference process executed in step S109 to the server device 3, and causes the server device 3 to execute a process for deriving recommended product-related information based on the purchase history information Ip.
[0117] In step S111 following step S110, the CPU 11 acquires the derived recommended product-related information. Then, in step S112 following step S111, the CPU 11 performs display processing of the derivation result screen G4. As will be understood from the previous explanation, the derived recommended product-related information and the inference basis information Ir are displayed on the derivation result screen G4. Note that the specific display content of the derivation result screen G4 has already been explained with reference to FIG. 11 , so a duplicate explanation will be avoided.
[0118] After executing the process of step S112, the CPU 11 advances the process to step S113 shown in FIG.
[0119] 15 , the CPU 11 determines in step S113 whether the recommended process has ended. Specifically, it determines whether the aforementioned End button B5 has been operated. If it is determined in step S113 that the End button B5 has not been operated and the recommended process has not ended, the CPU 11 proceeds to step S114 to determine whether a position change operation has been performed on the attribute map Ma. That is, it determines whether a movement operation has been performed on the target user characteristic image Ic displayed as a movable object on the attribute map Ma. If it is determined in step S114 that a movement operation has not been performed, the CPU 11 returns to step S113. That is, the CPU 11 waits for either the end of the recommended process or a position change operation on the attribute map Ma by the processes of steps S113 and S114.
[0120] If it is determined in step S114 that a position change operation has been performed, the CPU 11 proceeds to step S115 and changes the correction flag to ON. Here, the correction flag is a flag for managing whether or not an attribute correction operation has been performed by moving the target user characteristic image Ic, and its initial value is OFF (no correction).
[0121] In step S116 following step S115, the CPU 11 performs a display update process in response to the movement of the location. That is, as described above with reference to Fig. 13, the CPU 11 performs a process of displaying, as attribute guide information Gi, information indicating products purchased by users with attributes that match or are similar to the attributes of the location of the movement destination and the stores selling those products.
[0122] In step S117 following step S116, the CPU 11 determines whether the movement is complete, specifically, whether a predetermined condition for completing the movement operation of the target user characteristic image Ic is met, such as the finger that was in contact with the screen for the movement operation being removed from the screen. If the movement is not complete in step S117, the CPU 11 returns to step S116.
[0123] On the other hand, if it is determined in step S117 that the move is complete, the CPU 11 proceeds to step S118, where it performs a process of presenting recommended product-related information according to the attributes of the destination. Specifically, it causes the server device 3 to identify products and stores selling those products that have been purchased by other users with attributes that match or are similar to the attributes of the destination location, and displays information about the identified products and stores in the recommended information display area Ar.
[0124] After executing the process of step S118, the CPU 11 returns to step S113.
[0125] If the CPU 11 determines in step S113 that the end button B5 has been operated and the recommended process has ended, the CPU 11 proceeds to step S119 to determine whether the correction flag is ON. If the CPU 11 determines in step S119 that the correction flag is ON, the CPU 11 proceeds to step S120 to store the corrected map coordinates as corrected data for re-learning, and then proceeds to step S122. As mentioned above, the corrected data may be stored in, for example, the memory unit 19.
[0126] On the other hand, if it is determined in step S119 that the correction flag is not ON, the CPU 11 proceeds to step S121, where it stores the map coordinates obtained as the inference result of the attribute inference process as correct data for re-learning, and then proceeds to step S122.
[0127] In step S122, the CPU 11 resets the correction flag, and in the subsequent step S123, executes a re-learning process for the attribute inference model. Note that the details of this re-learning process have already been explained, so a duplicate explanation will be avoided.
[0128] After executing the process of step S122, the CPU 11 ends the series of processes shown in FIGS.
[0129] Although the above example shows that the re-learning process of the attribute inference model is performed on the information processing device 1 side, this re-learning process can also be performed on the server device 3 side. In this case, the CPU 11 of the information processing device 1 transmits attribute information as supervised data and feature information given to the attribute inference model as input data to the server device 3. Then, the server device 3 performs the re-learning process for the corresponding attribute inference model using the feature information and supervised data.
[0130] 5. Modifications Although the embodiments according to the present technology have been described above, the embodiments are not limited to the specific examples described above, and various modification configurations may be adopted. For example, in the above example, when extracting feature information of a target user, an example is given in which only an image captured by the image sensor 24 of the camera unit 23 is used. However, the feature information may also be extracted using sensing information from a sensor device other than the image sensor 24.
[0131] 16 shows an example of the configuration of an information processing device 1A that corresponds to the case where sensing information by a three-dimensional sensor 50 is used to extract feature information. The differences from the information processing device 1 shown in FIG. 3 are that a three-dimensional sensor 50 is connected to the input / output interface 15 and that a CPU 11A is used instead of the CPU 11.
[0132] The three-dimensional sensor 50 is a sensor device capable of three-dimensional measurement of a subject, and is configured to obtain, as a result of the three-dimensional measurement, an image in which a value indicating the distance to the subject is associated with each pixel, i.e., a range image. Specific examples of the three-dimensional sensor 50 include sensor devices that obtain range images using various distance measurement methods, such as the ToF (Time of Flight) method or the structured light method.
[0133] CPU 11A differs from CPU 11 in that it performs processing to extract feature information based on the three-dimensional measurement results of the subject, based on the three-dimensional measurement results of the subject obtained by three-dimensional sensor 50 and the results of AI image analysis processing of the captured image performed by AI processing unit 44 in camera unit 23. For example, the AI processing unit 44 can execute object detection processing targeting a person to obtain area information (e.g., a bounding box) of the person within the image frame, and by specifying the distance to the person based on the distance image obtained as the three-dimensional measurement result, feature information such as the person's height and physique (e.g., S, M, or L) can be extracted.
[0134] In this case, the AI image analysis processing to be performed by the AI processing unit 44 is not limited to the object detection processing described above, but may also include semantic segmentation processing and key point detection processing (for example, processing to detect the positions of key points in the human body, such as the head, chest, hands, waist, etc.).
[0135] Furthermore, when inferring a user's attributes, information on the user's height and build can be used in addition to the aforementioned characteristic information on "clothing," "gender," and "age."
[0136] In this case, for example, as illustrated in Figure 17, the CPU 11A will have the function of a feature information generation processing unit F11 that generates feature information on height and physique based on object detection information, segment information, and key point information obtained by AI image analysis processing by the AI processing unit 44 of the camera unit 23 and the three-dimensional measurement results by the three-dimensional sensor 50, and the function of an output processing unit F12 that outputs information on clothing, gender, and age obtained by feature information inference processing by the AI processing unit 44 of the camera unit 23 and the height and physique information generated by the feature information generation processing unit F11 as input data for the attribute inference model.
[0137] Note that the height and physique feature information may be inferred using an AI model instead of being generated by software processing by the CPU 11A. For example, an AI model may be used that uses captured images and range images of a person as the subject as input data, performs machine learning using information on the height and physique as training data, and infers the height and physique feature information of the person as the subject from fusion data of the captured images and range images.
[0138] Although the above description is based on the assumption that the target user is a single user, it is also possible to assume that there are multiple target users (i.e., a group of users). In the case where the target user is a group of users, it is conceivable that the combination of the characteristic information of each person constituting the group can be treated as the characteristic information of the group.
[0139] Here, in order to be able to handle cases where the target users are groups, the feature information provided as input data to the attribute inference model differs between when the target users are single and when the target users are groups, so an attribute inference model for single users and an attribute inference model for groups are prepared as attribute inference models. Also, for groups, the feature information provided as input data to the attribute inference model differs depending on the number of members, so an attribute inference model is prepared for each expected number of people in the group. This makes it possible to appropriately infer attributes regardless of the number of target users, as long as it is equal to or less than the expected number.
[0140] It is also possible to distinguish between types of groups, such as couples and families. Figure 18 is an explanatory diagram of a method for determining the group type of family and couple. In the case of a family shown in Figure 18A, it is possible to determine whether at least two adults and one child are recognized in the captured image. In the case of a couple shown in Figure 18B, it is possible to determine whether only two adults are recognized in the captured image.
[0141] In addition, when determining the group type in this way, if information on height and physique is required as characteristic information, a configuration including a three-dimensional sensor 50 and a CPU 11A, like the information processing device 1A, can be considered.
[0142] In the explanation so far, an example has been given in which the feature information of "clothing" is classified by the shape and color of the clothing, but it is also possible to classify the clothing by including the material, pattern, etc. By increasing the number of feature elements that can be taken into account when classifying attributes, it becomes possible to classify attributes in a more diverse and accurate manner, thereby improving the performance of attribute inference.
[0143] 6. Summary of the Embodiment As described above, the information processing device (1 or 1A) according to the embodiment includes a feature inference unit (AI processing unit 44) that infers feature information of the target user using an AI model based on an image of the target user, an attribute inference unit (AI processing unit 14) that infers attributes of the target user using an AI model based on the feature information inferred by the feature inference unit, and a presentation processing unit (F1) that receives recommended product-related information indicating products recommended to the target user or sales locations of the products, derived based on purchase history information indicating a product purchase history by the user, which information is associated with information indicating the attributes of the purchasing user, and information indicating the attributes of the target user inferred by the attribute inference unit, and performs processing to present the recommended product-related information and inference basis information indicating the inference basis for the attributes to the target user. According to the above configuration, the target user is presented with not only the recommended product-related information but also inference basis information indicating the inference basis for deriving the recommended product-related information. By presenting the inference basis information, it is possible to improve the trust that the target user has in the presented recommended product-related information, thereby improving the supportability of the user's purchasing behavior.
[0144] In addition, in the information processing device according to the embodiment, the feature inference unit is configured to be able to infer multiple types of feature information as feature information, and the type of feature information used by the attribute inference unit for attribute inference can be specified by operation. This allows the target user to specify the feature information to be used for attribute inference for multiple types of feature information, such as clothing, gender, age, face type, etc. This prevents the type of feature information that the target user does not want from being used in the process, thereby improving the target user's trust in the process of deriving recommended product-related information.
[0145] Furthermore, the information processing device according to the embodiment includes a model setting processing unit (F2) that selects an attribute inference model corresponding to a combination of feature information specified by an operation from a storage unit (19) that stores a plurality of attribute inference models that have been machine-learned for each combination of feature information that can be specified by an operation, as an attribute inference model that is an AI model used by the attribute inference unit to infer attributes, and sets the selected attribute inference model in the attribute inference unit. This makes it possible to ensure that attributes are appropriately inferred using an attribute inference model that is appropriate for any combination of feature information specified by the target user as feature information to be used for inferring attributes.
[0146] Furthermore, the information processing device according to the embodiment includes an image sensor (24) for capturing an image, and a feature inference unit is provided within the image sensor. This eliminates the need to transfer image data from the image sensor to a downstream processing unit in order to infer feature information. This reduces the amount of data transferred from the image sensor to the downstream processing unit.
[0147] In addition, in the information processing device according to the embodiment, the presentation processing unit performs processing to present, as inference basis information, information indicating an image capture target that the AI model focused on in the process of inferring attributes from among the image capture targets captured in the captured image. By presenting information about the image capture target that the AI model focused on as inference basis information as described above, the target user can recognize the inference basis from the perspective of which part of the captured image is the basis.
[0148] Furthermore, in the information processing device according to the embodiment, the attribute inference unit infers attributes using an AI model that is a self-organizing map model. Because machine learning of the self-organizing map model can be performed by unsupervised learning, it is possible to generate an AI model that infers attributes without preparing teacher data. Therefore, attribute inference can be achieved while reducing the effort required for the learning process.
[0149] Furthermore, in the information processing device according to the embodiment, the presentation processing unit performs processing to present the attribute map and information indicating the position of the target user's attributes in the map space. This allows the target user to intuitively understand what kind of attributes have been inferred.
[0150] Furthermore, in the information processing device according to the embodiment, the presentation processing unit performs processing to present a movable object (target user characteristic image Ic), which is a display object that can be moved on the attribute map, as information indicating the position of the target user's attributes. When the movable object is moved, the presentation processing unit also performs processing to present information indicating a product or a sales location of the product purchased by a user with attributes corresponding to the moved location. This allows the target user to search for the location where the attribute should be corrected, based on the information on the product or sales location presented at the move destination of the movable object. This improves the ease of attribute correction.
[0151] Furthermore, in the information processing device according to the embodiment, the presentation processing unit performs processing to present a movable object, which is a display object that can be moved on an attribute map, as information indicating the position of the attribute of the target user, and also performs processing to present product-related information corresponding to the attribute of the position to which the movable object is moved upon completion of the movement operation of the movable object. This makes it possible to present appropriate information that reflects the attribute modification when presenting recommended product-related information.
[0152] Furthermore, the information processing device according to the embodiment includes a re-learning processing unit (F3) that performs re-learning processing of the AI model used by the attribute inference unit using information indicating the attributes of the destination location as training data. As a result, when the target user modifies the attributes, the attribute inference model is re-learned using information indicating the modified attributes as training data. Therefore, the attribute inference model can be updated to more accurately infer attributes, thereby improving attribute inference performance.
[0153] Furthermore, in the information processing device according to the embodiment, if a moving operation of the movable object is not performed, the re-learning processing unit performs a re-learning process using the attribute information inferred by the attribute inference unit as training data. As a result, if the target user does not modify the attributes inferred by the attribute inference unit, the AI model of the attribute inference unit is re-learned using the inferred attribute information as training data. Therefore, the attribute inference model can be updated so that attributes can be inferred more accurately, thereby improving attribute inference performance.
[0154] An information processing method as an embodiment is an information processing method that uses an AI model to infer characteristic information of the target user based on an image of the target user, infers attributes of the target user based on the inferred characteristic information, inputs recommended product-related information that indicates products recommended to the target user or sales locations of the products, which is derived based on purchase history information that indicates a product purchase history by the user and is associated with information that indicates the attributes of the purchasing user, and information that indicates the estimated attributes of the target user, and presents the recommended product-related information and inference basis information that indicates the basis for inferring the attributes to the target user. Such an information processing method can also achieve the same functions and effects as the information processing device as the above-mentioned embodiment.
[0155] The effects described in this specification are merely examples and are not limiting, and other effects may also be present.
[0156] <7. The Present Technology> The present technology may also have the following configuration: (1) An information processing device comprising: a feature inference unit that infers feature information of a target user using an AI model based on an image of the target user; an attribute inference unit that infers attributes of the target user using an AI model based on the feature information inferred by the feature inference unit; and a presentation processing unit that receives input of recommended product-related information that indicates a product recommended to the target user or a sales location of the product, the recommended product-related information being derived based on purchase history information that indicates a product purchase history by a user and that is associated with information that indicates attributes of the purchasing user, and information that indicates the attributes of the target user inferred by the attribute inference unit, and that performs processing to present the recommended product-related information and inference basis information that indicates a basis for inferring the attributes to the target user. (2) The information processing device described in (1) above, wherein the feature inference unit is configured to be able to infer multiple types of feature information as the feature information, and the type of feature information that the attribute inference unit uses to infer the attributes can be specified by an operation. (3) The information processing device according to (2), further comprising: a model setting processing unit that selects, from a storage unit that stores a plurality of attribute inference models that have been machine-learned for each combination of feature information that can be specified by the operation, an attribute inference model that corresponds to the combination of feature information specified by the operation as an AI model used by the attribute inference unit to infer the attributes, and sets the selected attribute inference model in the attribute inference unit. (4) The information processing device according to any of (1) to (3), further comprising: an image sensor that obtains the captured image, and the feature inference unit is provided within the image sensor. (5) The information processing device according to any of (1) to (4), further comprising: a display processing unit that displays, as the inference basis information, information indicating an image target that an AI model focused on in the process of inferring the attributes, from among image targets captured in the captured image. (6) The information processing device according to any of (1) to (5), wherein the attribute inference unit infers the attributes using an AI model that is a self-organizing map model.(7) The information processing device according to (6), wherein the presentation processing unit performs a process of presenting an attribute map and information indicating the position of the attribute of the target user on a map space. (8) The information processing device according to (7), wherein the presentation processing unit performs a process of presenting a movable object that is a display object that can be moved on the attribute map as the information indicating the position of the attribute of the target user, and, when the movable object is moved by an operation, performs a process of presenting information indicating a product purchased by a user having an attribute corresponding to the position of the moved destination or a sales location of the product. (9) The information processing device according to (7) or (8), wherein the presentation processing unit performs a process of presenting a movable object that is a display object that can be moved on the attribute map as the information indicating the position of the attribute of the target user, and, upon completion of the operation of moving the movable object, performs a process of presenting the product-related information corresponding to the attribute of the position of the moved destination. (10) The information processing device according to (9), further comprising a re-learning processing unit that performs a re-learning process of an AI model used by the attribute inference unit, using the information indicating the attribute of the position of the moved destination as training data. (11) The information processing device according to (10), wherein the relearning processing unit performs the relearning process using information on the attributes inferred by the attribute inference unit as training data when a movement operation of the movable object is not performed. (12) An information processing method comprising: inferring characteristic information of the target user using an AI model based on a captured image of the target user; inferring attributes of the target user using the AI model based on the inferred characteristic information; inputting recommended product-related information that indicates a product recommended to the target user or a sales location of the product, the recommended product-related information being derived based on purchase history information that indicates a product purchase history by a user and that is associated with information that indicates attributes of the purchasing user, and information that indicates the estimated attributes of the target user; and presenting the recommended product-related information and inference basis information that indicates a basis for inferring the attributes to the target user.
[0157] DESCRIPTION OF SYMBOLS 1 Information processing device 2 Network 3 Server device FS Merchandise purchasing facility 11, 11A CPU 14 AI processing unit 17 Display unit 19 Memory unit 24 Image sensor 27 Control unit 41 Imaging unit 43 Sensor internal control unit 44 AI processing unit F1 Presentation processing unit F2 Model setting processing unit F3 Re-learning processing unit G1 Initial screen G2 Recommendation target specification screen G3 Feature information specification screen G4 Derivation result screen Ad Captured image display area Ar Recommended information display area Ir Inference basis information Ig Store guide information Ma Attribute map cb Check box Ic Target user feature image Ia Other user feature image Ri Recommendation information Gi Attribute guide information F11 Feature information generation processing unit F12 Feature information output processing unit
Claims
1. An information processing device comprising: a feature inference unit that infers feature information of a target user using an AI model based on an image of the target user; an attribute inference unit that infers attributes of the target user using an AI model based on the feature information inferred by the feature inference unit; and a presentation processing unit that inputs recommended product related information, which is information indicating a product recommended to the target user or a sales location of the product, derived based on purchase history information indicating a product purchase history by a user, which is associated with information indicating attributes of the purchasing user, and information indicating the attributes of the target user inferred by the attribute inference unit, and performs processing to present the recommended product related information and inference basis information indicating the inference basis for the attributes to the target user.
2. The information processing device according to claim 1, wherein the feature inference unit is configured to be able to infer multiple types of feature information as the feature information, and the type of feature information that the attribute inference unit uses to infer the attribute can be specified by operation.
3. The information processing device according to claim 2, further comprising a model setting processing unit that selects, from a memory unit that stores a plurality of attribute inference models that have been machine-learned for each combination of feature information that can be specified by the operation, an attribute inference model that is an AI model used by the attribute inference unit to infer the attributes, the attribute inference model that corresponds to the combination of feature information specified by the operation, and sets the selected attribute inference model in the attribute inference unit.
4. The information processing device according to claim 1, further comprising an image sensor for obtaining the captured image, wherein the feature inference unit is provided within the image sensor.
5. The information processing device according to claim 1, wherein the presentation processing unit performs processing to present, as the inference basis information, information indicating an object captured in the captured image that the AI model focused on in the process of inferring the attribute.
6. The information processing device according to claim 1, wherein the attribute inference unit infers the attributes using an AI model as a self-organizing map model.
7. The information processing device according to claim 6, wherein the presentation processing unit performs processing to present an attribute map and information indicating the position of the attribute of the target user on a map space.
8. The information processing device according to claim 7, wherein the presentation processing unit performs processing to present a movable object, which is a display object that can be moved on the attribute map, as information indicating the position of the attribute of the target user, and when the movable object is moved by operation, performs processing to present information indicating a product purchased by a user with an attribute corresponding to the position of the moved destination, or a sales location of the product.
9. The information processing device according to claim 7, wherein the presentation processing unit performs processing to present a movable object, which is a display object that can be moved on the attribute map, as information indicating the position of the attribute of the target user, and, upon completion of the movement operation of the movable object, performs processing to present the product-related information corresponding to the attribute of the destination position.
10. The information processing device according to claim 9, further comprising a re-learning processing unit that performs re-learning processing of the AI model used by the attribute inference unit using information indicating the attributes of the destination position as training data.
11. The information processing device according to claim 10, wherein, if the movable object is not moved, the re-learning processing unit performs the re-learning process using the attribute information inferred by the attribute inference unit as training data.
12. An information processing method comprising the steps of: inferring characteristic information of a target user using an AI model based on an image of the target user; inferring attributes of the target user using an AI model based on the inferred characteristic information; inputting recommended product-related information, which is information indicating products recommended to the target user or sales locations of the products, derived based on purchase history information indicating a product purchase history by the user, which is associated with information indicating the attributes of the purchasing user, and information indicating the estimated attributes of the target user; and presenting the recommended product-related information and inference basis information indicating the inference basis for the attributes to the target user.
Citation Information
Patent Citations
Electronic signboard system
JP2017156514A
Digital signage
JP2021051553A
Digital signage system
JP2022163309A
Method and system for dynamically targeting content based on automatic demographics and behavior analysis
US7921036B1
Information processing method and information processing system
WO2023002648A1