Information processing device, information processing method, and information processing program

By introducing special models and data processing mechanisms into the learning unit, the problem of data forgetting in incremental learning is solved, and the knowledge of learning new attributes and retaining old attributes in e-commerce is realized.

JP7676662B2Active Publication Date: 2025-05-14RAKUTEN GROUP INC
View PDF 6 Cites 0 Cited by

Patent Information

Application Number
JP2024519664
Authority / Receiving Office
JP · JP
Patent Type
Patents
Current Assignee / Owner
Filing Date
2023-03-31
Publication Date
2025-05-14
Estimated Expiration
2043-03-31

AI Technical Summary

Technical Problem

When using incremental learning, the existing technology is prone to the problem of data expiration or "forgot" existing data, which is called catastrophic forgetting, and it is difficult to effectively learn new attributes while retaining knowledge of old attributes.

Method used

Using learning units containing the first and second learning units, each attribute is predicted by training a special model, and the attribute prediction model is trained using part of the old data and a full amount of new data to prevent catastrophic forgetting.

Benefits of technology

It realizes learning new attributes from product images in e-commerce, while avoiding the forgetting of old attributes, ensuring the continuous learning and adaptability of the model.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure 0007676662000002
    Figure 0007676662000002
  • Figure 0007676662000003
    Figure 0007676662000003
  • Figure 0007676662000004
    Figure 0007676662000004
Patent Text Reader

Abstract

This information processing device includes a training unit that trains an attribute prediction model for predicting N attributes (N is an integer of 2 or more) of a product from an image including the product. The training unit includes: a first training unit that trains an i-th training model for predicting an i-th attribute (i is an integer of 1 to N-1) of the product by using i-th training data for the i-th attribute; and a second training unit that trains the attribute prediction model by using the i-th training model trained by the first training unit, a part of the i-th training data, and N-th training data for an N-th attribute of the product.
Need to check novelty before this filing date? Find Prior Art

Description

[Technical field]

[0001] The present invention relates to an information processing device, an information processing method, and an information processing program, and more particularly to a technique for predicting attributes of a product from an image including the product. [Background technology]

[0002] In recent years, electronic commerce (EC), the sale of goods over the Internet, has become very common, and many EC sites have been established on the Web to carry out such EC transactions. EC sites are often established in the languages ​​of countries around the world, allowing users (consumers) living in many countries to purchase goods. By accessing EC sites from a personal computer (PC) or a mobile device such as a smartphone, users can select and purchase the goods they want, regardless of time, without having to go to a physical store.

[0003] In order to increase the user's desire to purchase, EC sites may display recommended products that have similar attributes (information specific to products) to products previously purchased by the user on the screen the user is viewing. In addition, when a user wishes to purchase a desired product, the user may search based on the attributes of the product to be purchased. For this reason, identifying product attributes is a common challenge for site operators and product providers in electronic commerce.

[0004] In recent years, a technology has been developed that predicts attributes of a product from an image including the product using a learning model for machine learning. For example, Patent Document 1 discloses a technology that predicts multiple attributes of a product by applying an image including the product to a learning model configured using a neural network. [Prior art documents] [Patent documents]

[0005] [Patent Document 1] JP 2020-71859 A Summary of the Invention [Problem to be solved by the invention]

[0006] According to the technology disclosed in Patent Document 1, by applying an image including a product to a learning model, it becomes possible to automatically predict multiple attributes of the product. Meanwhile, the types and number of products handled in electronic commerce are increasing, and the types of attributes are also increasing accordingly. Therefore, it is desirable to build a learning model that predicts not only existing attributes of a product but also new attributes from an image including the product.

[0007] In order to train a learning model for a new task, a method called incremental learning is known in which the learning model is trained by continuously using additional learning data. According to incremental learning, by training the learning model using additional learning data of new attributes, it is possible to construct a learning model that can predict not only already-trained attributes but also newly-trained attributes.

[0008] However, when training a learning model using successively additional training data, a problem occurs called catastrophic forgetting, in which the previously trained data is "forgotten." For example, the previously trained data may have a large error with respect to the correct answer data, or may be completely forgotten.

[0009] The present invention has been made in consideration of the above-mentioned problems, and aims to provide a technique for learning new attributes of a product from an image containing the product while suppressing catastrophic forgetting. [Means for solving the problem]

[0010] An information processing device according to one embodiment of the present invention has a learning unit that trains an attribute prediction model for predicting N attributes (N is an integer equal to or greater than 2) of a product from an image including the product, and the learning unit has a first learning unit that trains an ith learning model for predicting an ith attribute (i is an integer from 1 to N-1) of the product using ith learning data for the ith attribute, and a second learning unit that trains the attribute prediction model using the ith learning model trained by the first learning unit, a portion of the ith learning data, and Nth learning data for the Nth attribute of the product.

[0011] An information processing method according to one embodiment of the present invention is an information processing method for training an attribute prediction model for predicting N attributes (N is an integer greater than or equal to 2) of a product from an image including the product, and includes training an ith learning model for predicting an ith attribute (i is an integer from 1 to N-1) of the product using ith learning data for the ith attribute, and training the attribute prediction model using the ith learning model trained by the first learning unit, a portion of the ith learning data, and Nth learning data for the Nth attribute of the product.

[0012] An information processing program according to one embodiment of the present invention is an information processing program that causes a computer to train an attribute prediction model for predicting N attributes (N is an integer equal to or greater than 2) of a product from an image including the product, and causes a computer to execute the following steps: training an ith learning model for predicting an ith attribute (i is an integer from 1 to N-1) of the product using ith learning data for the ith attribute; and training the attribute prediction model using the ith learning model trained by the first learning unit, a portion of the ith learning data, and Nth learning data for the Nth attribute of the product. Effect of the Invention

[0013] According to the present invention, it is possible to learn new attributes of a product from an image including the product while suppressing catastrophic forgetting. The above-mentioned objects, aspects, and advantages of the present invention, as well as objects, aspects, and advantages of the present invention not described above, will be understood by those skilled in the art from the following detailed description of the invention by referring to the accompanying drawings and the claims. [Brief description of the drawings]

[0014] [Figure 1] FIG. 1 shows an example of the configuration of a natural language processing system according to an embodiment. [Diagram 2] FIG. 2 illustrates an example of a hardware configuration of an information processing device according to an embodiment. [Figure 3A] FIG. 3A is a diagram for explaining product attributes. [Figure 3B] FIG. 3B shows the relationship between product images and attributes. [Figure 4] FIG. 4 shows a conceptual diagram of attribute learning data. [Figure 5A] FIG. 5A shows the learning procedure of the learning model for the first attribute in the first time step. [Figure 5B] FIG. 5B shows a learning procedure of the learning model for the second attribute from the first attribute in the second time step. [Figure 5C] FIG. 5C shows the learning procedure of the learning model from the first attribute to the third attribute in the third time step. [Figure 5D] FIG. 5D is a diagram for explaining the learning procedure of the learning model for the first attribute to the Nth attribute in the Nth time step. [Figure 6A] FIG. 6A shows a flowchart of a learning process of an attribute prediction model executed by an information processing device according to an embodiment. [Figure 6B] FIG. 6B shows a flowchart of a modified example of the learning process of an attribute prediction model executed by the information processing device according to the embodiment. [Figure 7] FIG. 7 shows a flowchart of an attribute prediction process executed by an information processing device according to the embodiment. DETAILED DESCRIPTION OF THE PREFERRED EMBODIMENTS

[0015] Hereinafter, with reference to the attached drawings, an embodiment for carrying out the present invention will be described in detail. Among the components disclosed below, those having the same functions are given the same reference numerals, and their description will be omitted. Note that the embodiment disclosed below is an example of a means for realizing the present invention, and should be appropriately modified or changed depending on the configuration of the device to which the present invention is applied and various conditions, and the present invention is not limited to the following embodiment. Furthermore, not all of the combinations of features described in the present embodiment are necessarily essential to the solution of the present invention.

[0016] [Functional configuration of information processing device] The information processing device 100 according to this embodiment acquires an image (hereinafter also referred to as a product image) that includes a product (i.e., displays the product) and predicts multiple attributes of the product. A product attribute is information that is specific to the product, such as the visual characteristics of the product. The product attribute can be an indicator when a user purchases the product. Note that, although an example of predicting product attributes is described in this embodiment, this embodiment can also be applied to the case where attributes of any item, including a product, are predicted from an image that includes the item (i.e., displays the item).

[0017] FIG. 1 shows an example of the functional configuration of an information processing device 100 according to this embodiment. The information processing device 100 includes a learning data acquisition unit 101, a learning data management unit 102, a learning unit 103, an image acquisition unit 104, an attribute prediction unit 105, an output unit 106, a learning model storage unit 110, and an attribute data storage unit 120. The learning unit 103 includes a first learning unit 1031 and a second learning unit 1032. The learning model storage unit 110 is configured to store a first attribute specialized model 111-1 to an N-th attribute specialized model 111-N and an attribute prediction model 112. The attribute data storage unit 120 is configured to store a first attribute learning data 121-1 to an N-th attribute learning data 121-N. In the present disclosure, "N" is an integer equal to or greater than 2.

[0018] The learning data acquisition unit 101 acquires learning data (teacher data) for training a learning model. In this embodiment, the learning data includes product images and correct answer data of product attributes. Here, an example of product attributes in this embodiment will be described with reference to Fig. 3A. Fig. 3A is a diagram for explaining product attributes.

[0019] In this embodiment, the products are assumed to be products that can be handled on an EC (Electronic Commerce) site. Since the types and number of products that can be handled on an EC site are enormous, product attributes are set for classified products (product groups). Products can be classified hierarchically. In this embodiment, attributes are assumed for a product category that indicates a higher classification among product classifications, where "clothing" is the product category.

[0020] In FIG. 3A, products in the product category 30 "clothing" are classified into a subcategory 31 indicating a subcategory of the category 30, and a sub-subcategory 32 indicating a further subcategory of the subcategory 31. The subcategory 31 indicates the object on which the "clothing" is worn, and includes "men" and "women" in the example of FIG. 3A. In addition, the subcategory 31 may include "children", "seniors", and "unisex" that does not consider gender. The sub-subcategory 32 indicates the type and shape of the "clothing". When the subcategory 31 is "men", the sub-subcategory 32 includes "T-shirts" and "jeans". In addition, when the subcategory 31 is "men", the sub-subcategory 32 may include "jackets", "coats", and the like.

[0021] In the example of FIG. 3A, attributes are set for products in sub-subcategory 32. When sub-subcategory 32 is "T-shirts", the attributes include "pattern", "sleeve", "neckline", and "color". When sub-subcategory 32 is "jeans", the attributes include "pattern", "fit", and "length". The types of attributes for each product shown in FIG. 3A are merely examples and are not limited to those shown. The types of attributes may be further increased in the future. Although attributes are set for products in sub-subcategory 32 in FIG. 3A, attributes may be set for subcategory 31 or category 30.

[0022] FIG. 3B shows the relationship between the product image and the attribute. FIG. 3B shows data 35 including the attribute 37 of the product 38 included in the product image 36. The product 38 is a product classified in FIG. 3A under the category 30=“Clothing”, subcategory 31=“Men”, and sub-subcategory 32=“T-shirts”. Referring to FIG. 3A, the attribute 37 includes “Pattern”, “Sleeve”, “Neckline”, and “Color”. Furthermore, in the attribute 37, each attribute has a correct feature (feature value) (hereinafter also referred to as a correct feature). The correct features for each attribute of the product 38 in the product image 36 are as shown in FIG. 3B, where “Pattern” is “Border”, “Sleeve” is “Three-quarter sleeve”, “Neckline” is “Round”, and “Color” is “White and black”.

[0023] As described above, the learning data includes a product image and correct answer data of the product attributes. The learning data acquisition unit 101 generates and acquires learning data from, for example, the data 35 shown in FIG. 3B. If the attribute of the product 38 shown in FIG. 3B is "pattern", the features that the "pattern" can have are, for example, "border", "solid", "check", "dot", and "print". In the case of a classification problem, all correct answer data for learning are 1 or 0. That is, there is a choice between 100% (correct answer feature) and 0% (incorrect answer feature). Since the "pattern" of the product 38 included in the product image 36 is "border", in the order of "border", "solid", "check", "dot", and "print", the first feature is the correct answer feature. Therefore, the correct answer data is given in the form of a probability distribution such as {1,0,0,0,0}.

[0024] The learning data management unit 102 stores the learning data acquired by the learning data acquisition unit 101 as attribute learning data for each attribute in the attribute data storage unit 120. Referring to Fig. 3B, the learning data management unit 102 stores a combination of the product image 36 and the correct answer data for the attributes 37 (i.e., "pattern", "sleeve", "neckline", and "color") as attribute learning data for each attribute in the attribute data storage unit 120. Fig. 4 shows a conceptual diagram of the attribute learning data stored in the attribute data storage unit 120.

[0025] As shown in FIG. 1, the attribute data storage unit 120 is configured to store the first attribute learning data 121-1 to the N-th attribute learning data 121-N. For the sake of explanation, it is assumed that the first attribute = "pattern" and the second attribute = "sleeve". For example, based on the data 35 shown in FIG. 3B, the learning data management unit 102 classifies the combination data of the product image 36 and the correct answer data in which "border" 41 is "1" into the first attribute learning data 121-1. In addition, the learning data management unit 102 classifies the combination data of the product image 36 and the correct answer data in which "three-quarter sleeve" 42 is "1" into the second attribute learning data 121- 2In this way, the learning data management unit 102 extracts a product image and one or more attributes from the learning data acquired by the learning data acquisition unit 101, and classifies the combination of the product image and the correct answer data into any one of the first attribute learning data 121-1 to the Nth attribute learning data 121-N for each attribute. As a result, a set of combination data of a product image and correct answer data for the first attribute to the Nth attribute of the product included in the image is stored in each of the first attribute learning data 121-1 to the Nth attribute learning data 121-N.

[0026] In the present embodiment, the learning data management unit 102 is configured to store the learning data acquired by the learning data acquisition unit 101 in the attribute data storage unit 120 for each attribute, but the procedure for storing data in the attribute data storage unit 120 is not limited to this. For example, when the learning data acquisition unit 101 acquires a product image and a correct answer feature of one attribute, the learning data management unit 102 may store combination data of the product image and correct answer data corresponding to the correct answer feature in any one of the first attribute learning data 121-1 to the Nth attribute learning data 121-N according to the attribute. Alternatively, in this case, the learning data acquisition unit 101 may directly store the combination data in any one of the first attribute learning data 121-1 to the Nth attribute learning data 121-N.

[0027] Furthermore, the learning data management unit 102 manages data to be supplied to the learning unit 103 from the first attribute learning data 121-1 to the N-th attribute learning data 121-N stored in the attribute data storage unit 120. Taking the first attribute learning data 121-1 as an example, the learning data management unit 102 may supply a full set of the first attribute learning data 121-1 to the learning unit 103. Alternatively, the learning data management unit 102 may select (sample) a part of the first attribute learning data 121-1 and supply it to the learning unit 103. The part of the data may be selected randomly or according to a predetermined rule. In the present disclosure, the term "full set of data" means a larger amount of data (e.g., including a larger set of learning data) than "part of the data". The learning data management unit 102 may supply the attribute learning data to the learning unit 103 in response to a request from the learning unit 103.

[0028] Returning to the description of FIG. 1, the learning unit 103 learns the first attribute learning data 121-1 to the Nth attribute learning data 121-N and the attribute prediction model 112. These learning models can be configured using a CNN (Convolutional Neural Network). The learning unit 103 includes a first learning unit 1301 and a second learning unit 1302, and is responsible for controlling the overall learning process including controlling the first learning unit 1301 and the second learning unit 1302. The first learning unit 1301 is configured to learn the first attribute dedicated model 111-1 to the Nth attribute dedicated model 111-N. The second learning unit 1302 is configured to learn the attribute prediction model 112 using at least one of the learned first attribute dedicated model 111-1 to the Nth attribute dedicated model 111-N. The learning procedure of the learning unit 103 for the first attribute dedicated model 111-1 to the N-th attribute dedicated model 111-N and the attribute prediction model 112 will be described later.

[0029] The image acquisition unit 104 acquires a product image including a target product (i.e., a product whose attributes are to be predicted). The image acquisition unit 104 may acquire the product image by an input operation by a user (operator) via the input unit 25 (FIG. 2), or may acquire the product image from a storage unit (ROM 22 or RAM 23 in FIG. 2) by a user operation. The image acquisition unit 104 may also acquire a product image received from an external device via the communication I / F 27 (FIG. 2).

[0030] The attribute prediction unit 105 applies the product image acquired by the image acquisition unit 104 to the trained attribute prediction model 112, and predicts a plurality of attributes of the target product included in the product image.

[0031] The output unit 106 outputs information on the attributes predicted by the attribute prediction unit 105 (attribute prediction result). The output unit 106 may output the attribute prediction result in association with, for example, a product image acquired by the image acquisition unit 104. The output unit 106 may display the attribute prediction result on the display unit 26 (FIG. 2). In addition, when a product image is acquired from an external device such as a user device, the output unit 106 may transmit the attribute prediction result to the external device via the communication I / F 27 (FIG. 2) so as to be displayed on the display unit of the external device.

[0032] [Hardware configuration of information processing device] 2 is a block diagram showing an example of a hardware configuration of the information processing device 100 according to this embodiment. The information processing device 100 can be implemented on a single or multiple computers, mobile devices, or any other processing platform. 2, the information processing device 100 is illustrated as being implemented in a single computer, but the information processing device 100 according to the present embodiment may be implemented in a computer system including multiple computers. The multiple computers may be connected to each other via a wired or wireless network so as to be able to communicate with each other.

[0033] 2, the information processing device 100 may include a CPU (Central Processing Unit) 21, a ROM (Read Only Memory) 22, a RAM (Random Access Memory) 23, a HDD (Hard Disk Drive) 24, an input unit 25, a display unit 26, a communication I / F 27, and a system bus 28. The information processing device 100 may also include an external memory. The CPU 21 comprehensively controls the operation of the information processing device 100, and controls each component (22 to 27) via a system bus 28, which is a data transmission path. The CPU 21 is composed of one or more processors. At least one of the one or more processors may be replaced by one or more processors such as an ASIC (Application specific integrated circuit), an FPGA (Field Programmable Gate Array), a DSP (Digital Signal Processor), or a GPU (Graphics Processing Unit).

[0034] The ROM 22 is a non-volatile memory that stores control programs and the like necessary for the CPU 21 to execute processes. Note that the programs may be stored in a non-volatile memory such as the HDD 24 or an SSD (Solid State Drive) or an external memory such as a removable storage medium (not shown). The RAM 23 is a volatile memory and functions as a main memory, a work area, etc. of the CPU 21. That is, the CPU 21 loads necessary programs, etc. from the ROM 22 to the RAM 23 when executing processing, and realizes various functional operations by executing the programs, etc. The learning model storage unit 110 and the attribute data storage unit 120 shown in FIG. 1 can be configured by the RAM 23.

[0035] The HDD 24 stores, for example, various data and various information required when the CPU 21 performs processing using a program. The HDD 24 also stores, for example, various data and various information obtained when the CPU 21 performs processing using a program. The input unit 25 is composed of a keyboard and a pointing device such as a mouse. The display unit 26 is configured with a monitor such as a liquid crystal display (LCD), etc. The display unit 26 may be configured in combination with the input unit 25 to function as a GUI (Graphical User Interface).

[0036] The communication I / F 27 is an interface that controls communication between the information processing device 100 and an external device. The communication I / F 27 provides an interface with a network and executes communication with an external device via the network. Various data, various parameters, etc. are transmitted and received between the external device and the communication I / F 27. In this embodiment, the communication I / F 27 may execute communication via a wired LAN (Local Area Network) or a dedicated line conforming to a communication standard such as Ethernet (registered trademark). However, the network available in this embodiment is not limited to this, and may be configured as a wireless network. This wireless network includes wireless PANs (Personal Area Networks) such as Bluetooth (registered trademark), ZigBee (registered trademark), and UWB (Ultra Wide Band). It also includes wireless LANs (Local Area Networks) such as Wi-Fi (Wireless Fidelity) (registered trademark), and wireless MANs (Metropolitan Area Networks) such as WiMAX (registered trademark). It also includes wireless WANs (Wide Area Networks) such as 4G and 5G defined by 3GPP (Third Generation Partnership Project) (registered trademark). It should be noted that the network is sufficient as long as it can connect the devices to each other so that they can communicate with each other, and the communication standard, scale, and configuration are not limited to those described above.

[0037] At least some of the functions of the information processing device 100 shown in Fig. 1 can be realized by executing a program by the CPU 21. However, at least some of the functions of the information processing device 100 shown in Fig. 1 may be operated as dedicated hardware. In this case, the dedicated hardware may operate under the control of the CPU 21.

[0038] [Learning procedure for learning model] Next, the learning procedure of the learning model by the learning unit 103 will be described with reference to Figs. 5A to 5D. In this embodiment, the learning unit 103 performs incremental learning, which trains the learning model by continuously using additional learning data. Fig. 5A shows the learning procedure of the learning model for the first attribute in the first time step. Fig. 5B shows the learning procedure of the learning model for the first attribute to the second attribute in the second time step. Fig. 5C shows the learning procedure of the learning model for the first attribute to the third attribute in the third time step. Fig. 5D is a diagram for explaining the learning procedure of the learning model for the first attribute to the Nth attribute in the Nth time step. Here, the term "time step" is understood to be a term representing a relative time, and not a term representing a specific time (or time span).

[0039] First, the learning procedure in the first time step will be described with reference to FIG. 5A. In the first time step, the first learning unit 1031 applies the first attribute learning data 121-1 stored in the attribute data storage unit 120 to the first attribute dedicated model 111-1 to learn the first attribute dedicated model 111-1. In order to learn the first attribute dedicated model 111-1, a full set of the first attribute learning data 121-1 is used. Therefore, for example, the first learning unit 1031 may request the full set of the first attribute learning data 121-1 from the learning data management unit 102. In response to this, the learning data management unit 102 may supply the full set of the first attribute learning data 121-1 to the first learning unit 1031.

[0040] The first learning unit 1031 inputs the product images included in the first attribute learning data 121-1 to the first attribute dedicated model 111-1, and obtains a probability distribution of the characteristics that the first attribute can have (hereinafter also referred to as an estimated probability distribution) as the output 501 (estimation result). For example, when the first attribute is "pattern", the characteristics that the "pattern" can have are, for example, "border", "plain", "check", "dot", and "print", as described above. The output 501 is an estimated probability distribution in this order, for example, {0.6, 0.1, 0.1, 0.1, 0.1}. Also, as described above, in the case of the data 35 shown in FIG. 3B, the correct probability distribution, which is the correct answer data, is {1, 0, 0, 0, 0}.

[0041] The first learning unit 1031 acquires an output 501 (estimated probability distribution) and a correct answer probability distribution for all the first attribute learning data 121-1, and calculates an evaluation value (evaluation function) 502 based on the output 501 and the correct answer probability distribution. In this embodiment, cross-entropy (CE) (also referred to as cross-entropy loss) is used as an evaluation value for training an attribute-specific model such as the first attribute-specific model 111-1. The first learning unit 1031 trains the first attribute-specific model 111-1 based on the evaluation value 502 (for example, so as to minimize the evaluation value 502). The first learning unit 1031 stores the trained first attribute-specific model 111-1 in the training model storage unit 110. The trained first attribute-specific model 111-1 is used for the training process in the second time step.

[0042] Next, the learning procedure in the second time step will be described with reference to FIG. 5B. In the second time step, the first learning unit 1031 applies the second attribute learning data 121-2 stored in the attribute data storage unit 120 to the second attribute dedicated model 111-2 to learn the second attribute dedicated model 111-2. As in the procedure described with reference to FIG. 5A, a full set of the second attribute learning data 121-2 is used to learn the second attribute dedicated model 111-2. Therefore, for example, the first learning unit 1031 may request the full set of the second attribute learning data 121-2 from the learning data management unit 102. In response to this, the learning data management unit 102 may supply the full set of the second attribute learning data 121-2 to the first learning unit 1031.

[0043] The first learning unit 1031 inputs the product images included in the second attribute learning data 121-2 to the second attribute dedicated model 111-2 and obtains an estimated probability distribution for the second attribute as the output 503. As in the process in the first time step, the first learning unit 1031 obtains the output 503 (estimated probability distribution) and the correct answer probability distribution for all the second attribute learning data 121-2, and calculates an evaluation value 504 (e.g., cross entropy) based on the output 503 and the correct answer probability distribution. The first learning unit 1031 trains the second attribute dedicated model 111-2 based on the evaluation value 504 (e.g., so as to minimize the evaluation value 504). The first learning unit 1031 stores the trained second attribute dedicated model 111-2 in the learning model storage unit 110. The trained second attribute dedicated model 111-2 is used in the learning process in the third time step.

[0044] Furthermore, in the second time step, the second learning unit 1032 trains the attribute prediction model 112. The attribute prediction model 112 in the second time step is a learning model for predicting the first attribute and the second attribute. The second learning unit 1032 applies the first attribute learning data 121-1 and the second attribute learning data 121-2 stored in the attribute data storage unit 120 to the attribute prediction model 112 to train the attribute prediction model 112. In order to train the attribute prediction model 112, a part of the first attribute learning data 121-1 and a full set of the second attribute learning data 121-2 are used. Therefore, for example, the second learning unit 1032 may request the part of the first attribute learning data 121-1 and the full set of the second attribute learning data 121-2 from the training data management unit 102. In response to this, the learning data management unit 102 may supply a part of the first attribute learning data 121-1 and the full set of the second attribute learning data 121-2 to the second learning unit 1032. Note that in the figures of the present disclosure, a part of the area of ​​the attribute learning data is shaded to indicate that it is a part of the attribute learning data.

[0045] The second learning unit 1032 inputs a product image included in a part of the first attribute learning data 121-1 to the attribute prediction model 112, and obtains an estimated probability distribution for the first attribute as the output 505. The second learning unit 1302 inputs a product image included in a part of the first attribute learning data 121-1 to the first attribute dedicated model 111-1, and obtains an estimated probability distribution for the first attribute as the output 506. The second learning unit 1032 obtains the output 505 (estimated probability distribution) and the output 506 (estimated probability distribution) for all of the part of the first attribute learning data 121-1, and calculates an evaluation value 507 representing the similarity based on the output 505 and the output 506. In this embodiment, KL divergence (Kullback-Leibler divergence) (also referred to as KL divergence loss) is used as the evaluation value representing the similarity. The second learning unit 1032 trains the attribute prediction model 112 so that the output 505 and the output 506 match (i.e., so that the similarity becomes high). When the KL divergence is used as the evaluation value 507, the second learning unit 1032 trains the attribute prediction model 112 so as to minimize the evaluation value 507.

[0046] In this way, the second learning unit 1302 trains the attribute prediction model 112 so that the output 505 from the first attribute dedicated model 111-1 matches the output 506 from the attribute prediction model 112. In other words, the second learning unit 1302 trains a dedicated model trained for a specific attribute (the first attribute dedicated model 111-1 in the example of FIG. 5B) so that the output from the dedicated model matches the output from the learning model to be trained (the attribute prediction model 112 in the example of FIG. 5B), which is so-called knowledge distillation.

[0047] Furthermore, the second learning unit 1032 inputs product images included in the full set of second attribute learning data 121-2 to the attribute prediction model 112, and obtains an estimated probability distribution for the second attribute as output 508. The second learning unit 1032 obtains the output 508 (estimated probability distribution) and a correct answer probability distribution for all of the second attribute learning data 121-2, and calculates an evaluation value 509 (e.g., cross entropy) based on the output 508 and the correct answer probability distribution. The second learning unit 1032 trains the attribute prediction model 112 based on the evaluation value 509 (e.g., so as to minimize the evaluation value 509).

[0048] In this way, the second learning unit 1032 can suppress forgetting of a task for the first attribute that has already been learned by learning the attribute prediction model 112 using knowledge distillation. The second learning unit 1032 stores the learned attribute prediction model 112 in the learning model storage unit 110.

[0049] Next, the learning procedure in the third time step will be described with reference to FIG. 5C. In the third time step, the first learning unit 1031 applies the third attribute learning data 121-3 stored in the attribute data storage unit 120 to the third attribute dedicated model 111-3 to learn the third attribute dedicated model 111-3. As in the procedure described with reference to FIG. 5A, a full set of the third attribute learning data 121-3 is used to learn the third attribute dedicated model 111-3. Therefore, for example, the first learning unit 1031 may request the full set of the third attribute learning data 121-3 from the learning data management unit 102. In response to this, the learning data management unit 102 may supply the full set of the third attribute learning data 121-3 to the first learning unit 1031.

[0050] The first learning unit 1031 inputs the product images included in the third attribute learning data 121-3 to the third attribute dedicated model 111-3 and obtains an estimated probability distribution for the third attribute as an output 510. As in the process in the first time step, the first learning unit 1031 obtains the output 510 (estimated probability distribution) and the correct answer probability distribution for all the third attribute learning data 121-3, and calculates an evaluation value 511 (e.g., cross entropy) based on the output 510 and the correct answer probability distribution. The first learning unit 1031 trains the third attribute dedicated model 111-3 based on the evaluation value 511 (e.g., so as to minimize the evaluation value 511). The first learning unit 1031 stores the trained third attribute dedicated model 111-3 in the learning model storage unit 110. The trained third attribute dedicated model 111-3 is used for the learning process in the fourth time step.

[0051] Furthermore, in the third time step, the second learning unit 1032 trains the attribute prediction model 112. The attribute prediction model 112 in the third time step is a learning model for predicting a third attribute from a first attribute. The second learning unit 1032 applies the first attribute learning data 121-1, the second attribute learning data 121-2, and the third attribute learning data 121-3 stored in the attribute data storage unit 120 to the attribute prediction model 112 to train the attribute prediction model 112. In order to train the attribute prediction model 112, a part of the first attribute learning data 121-1, a part of the second attribute learning data 121-2, and a full set of the third attribute learning data 121-3 are used. Therefore, for example, the second learning unit 1032 may request a part of the first attribute learning data 121-1, a part of the second attribute learning data 121-2, and a full set of the third attribute learning data 121-3 from the training data management unit 102. In response to this, the learning data management unit 102 can supply the second learning unit 1032 with a portion of the first attribute learning data 121-1, a portion of the second attribute learning data 121-2, and the full set of the third attribute learning data 121-3.

[0052] The second learning unit 1032 inputs a product image included in a part of the first attribute learning data 121-1 to the attribute prediction model 112, and obtains an estimated probability distribution for the first attribute as the output 512. The second learning unit 1302 also inputs a product image included in a part of the first attribute learning data 121-1 to the first attribute dedicated model 111-1, and obtains an estimated probability distribution for the first attribute as the output 513. The second learning unit 1032 obtains the output 512 (estimated probability distribution) and the output 513 (estimated probability distribution) for all of the part of the first attribute learning data 121-1, and calculates an evaluation value 514 (for example, KL divergence) based on the output 512 and the output 513. The second learning unit 1032 trains the attribute prediction model 112 based on the evaluation value 514 (for example, to minimize the evaluation value 514). Moreover, the second learning unit 1032 inputs a product image included in a part of the second attribute learning data 121-2 to the attribute prediction model 112, and obtains an estimated probability distribution for the second attribute as the output 515. Moreover, the second learning unit 1302 inputs a product image included in a part of the second attribute learning data 121-2 to the second attribute dedicated model 111-2, and obtains an estimated probability distribution for the second attribute as the output 516. The second learning unit 1032 obtains the output 515 (estimated probability distribution) and the output 516 (estimated probability distribution) for all of the part of the second attribute learning data 121-2, and calculates an evaluation value 517 (for example, KL divergence) based on the output 515 and the output 516. The second learning unit 1032 trains the attribute prediction model 112 based on the evaluation value 517 (for example, to minimize the evaluation value 517).

[0053] Furthermore, the second learning unit 1032 inputs the product images included in the full set of the third attribute learning data 121-3 to the attribute prediction model 112, and obtains an estimated probability distribution for the third attribute as the output 518. 3For all of the above, an output 518 (estimated probability distribution) and a correct answer probability distribution are obtained, and an evaluation value 519 (e.g., cross entropy) is calculated based on the output 518 and the correct answer probability distribution. The second learning unit 1032 trains the attribute prediction model 112 based on the evaluation value 519 (e.g., to minimize the evaluation value 519). Second Learning Department 1031 teeth The trained attribute prediction model 112 is stored in the trained model storage unit 110.

[0054] 5D shows a diagram for explaining the learning procedure at the Nth time step by the first learning unit 1031 and the second learning unit 1302. The learning procedure shown in FIG 5D corresponds to a generalization of the learning procedures at the second time step to the Nth time step, respectively. Similar to the procedure described with reference to Figures 5B and 5C, the first learning unit 1301 inputs the full set of N-th attribute learning data 121-N to the N-th attribute-dedicated model 111-N, and trains the N-th attribute-dedicated model 111-N based on the output 520 and the evaluation value 521. 5B and 5C, the second learning unit 1032 trains the attribute prediction model 112. The attribute prediction model 112 in the N-th time step is a learning model for predicting the N-th attribute from the first attribute. Taking the first attribute, the second attribute, and the (N-1)th attribute as an example, the second learning unit 1032 trains the attribute prediction model 112 by deriving an evaluation value 524 based on outputs 522 and 523 for a portion of the first attribute learning data 121-1, and the second attribute learning data 121-2. 2 The output 525 for a part of the (N-1)th attribute learning data 121- (N-1) 5. The second learning unit 1302 trains the attribute prediction model 112 based on an output 528 for a portion of the Nth attribute learning data 121-N and an evaluation value 530 based on the output 529. Furthermore, the second learning unit 1302 inputs the full set of the Nth attribute learning data 121-N to the attribute prediction model 112, and trains the attribute prediction model 112 based on the output 531 and the evaluation value 532.

[0055] When cross entropy and KL divergence are used as evaluation values, the loss function L used by the second learning unit 1032 to learn the attribute prediction model 112 in the N-th time step can be expressed as in equation (1). TIFF0007676662000001.tif20158Here, CE is the cross entropy, KL is the KL divergence, τ is the temperature, σ is the softmax function, λ is the log-softmax function, and α is the balancing weight. Also, Y o i represents the output from the i-th attribute specialized model (i is an integer equal to or greater than 1). ^ o i Y represents the output for the i-th attribute from the attribute prediction model 112. ^ N Y represents the output for the Nth attribute from the attribute prediction model 112. N represents correct data for the Nth attribute. As can be seen from equation (1), the loss function is based on a value obtained by multiplying the evaluation values ​​(KL divergence) from the first attribute to the (N-1)th attribute by a weight α and adding them together, and a value obtained by multiplying the evaluation value (cross entropy) for the Nth attribute by a weight (1-α). The second learning unit 1032 trains the attribute prediction model 112 so as to optimize the value derived by the loss function (for example, to minimize the value).

[0056] In this way, the learning unit 103 trains the attribute prediction model 112 for predicting the first attribute to the Nth attribute for the first attribute to the (N-1)th attribute according to knowledge distillation. Specifically, for the first attribute to the (N-1)th attribute, the learning unit 103 trains the attribute prediction model 112 using the trained first attribute dedicated model 111-1 to (N-1)th attribute dedicated model 111-(N-1), which are trained using a full set of training data for the attribute. This makes it possible to train the attribute prediction model 112 so as to suppress catastrophic forgetting for the first attribute to the (N-1)th attribute.

[0057] The attribute prediction model 112 trained in each of the second time step to the (N-1)th time step may be stored in the learning model storage unit 110, distinguished from the attribute prediction model 112 trained in the Nth time step. The attribute prediction model 112 trained in the Nth time step is a learning model for predicting the Nth attribute from the first attribute, and the attribute prediction model 112 trained in the nth time step (n is an integer not less than 2 and not more than N-1) is a learning model for predicting the nth attribute from the first attribute.

[0058] [Variations of the learning procedure of the learning model] 5A to 5D, the procedure for each time step has been described, but the second learning unit 1032 may train the attribute prediction model 112 for predicting the Nth attribute from the first attribute using a procedure different from the above. This modification corresponds to a process of training only the attribute prediction model 112 for predicting the Nth attribute from the first attribute.

[0059] In this modification, first, the first learning unit 1301 learns the first attribute dedicated model 111-1 to the (N-1)th attribute dedicated model 111-(N-1), respectively, using the first attribute learning data 121-1 to the (N-1)th attribute learning data 121-(N-1), respectively. The learning procedure corresponds to the learning procedure of the first learning unit 1301 described with reference to Figs. 5A to 5D, for example. Then, the second learning unit 1302 learns the attribute dedicated model 112 according to the procedure described with reference to Fig. 5D. The learning procedure according to this modification also makes it possible to train the attribute prediction model 112 so as to suppress catastrophic forgetting for the first attribute to the (N-1)th attribute.

[0060] [Processing flow] Next, a flow of processing executed by the information processing device 100 according to this embodiment will be described with reference to Figures 6A, 6B, and 7. The processing shown in Figures 6A, 6B, and 7 can be realized by the CPU 21 of the information processing device 100 loading a program stored in the ROM 22 or the like into the RAM 23 and executing it.

[0061] (1-1) Learning process of attribute prediction model 112 FIG. 6A shows a flowchart of a learning process of the attribute prediction model 112 for predicting an N-th attribute from a first attribute, which is executed by the information processing device 100. In S601, the learning unit 103 sets a parameter i to 1. i corresponds to the time step in the learning procedure described with reference to Figs. 5A to 5D. In S602, the learning unit 103 judges whether i is 1 or not. If i is 1 (Yes in S602), the process proceeds to S603, and if not (No in S602), the process proceeds to S605.

[0062] S603 is a learning process in the first time step. In S603, the first learning unit 1031 uses the first attribute learning data 121-1 to learn the first attribute specialized model 111-1. The learning procedure in S603 is as described with reference to FIG. 5A. After learning the first attribute specialized model 111-1, the first learning unit 1031 stores the learned first attribute specialized model 111-1 in the learning model storage unit 110. Then, in S604, the learning unit 103 increments the parameter i. After that, the process proceeds to S602.

[0063] S605 to S606 are learning processes in the i-th time step (i is 2 or more). In S605, the first learning unit 1031 uses the i-th attribute learning data 121-i to learn the i-th attribute specialized model 111-i. After learning the i-th attribute specialized model 111-i, the first learning unit 1031 stores the learned i-th attribute specialized model 111-i in the learning model storage unit 110. Next, in S606, the second learning unit 1032 uses the learned first attribute dedicated model 111-1 to the (i-1)th attribute dedicated model 111-(i-1), a part of each of the first attribute learning data 121-1 to the (i-1)th attribute learning data 121-(i-1), and the i-th attribute learning data 121-i to learn the attribute prediction model 112. The learning procedure in S606 is as described with reference to FIG. 5B to FIG. 5D. After learning the attribute prediction model 112, the second learning unit 1032 may store the attribute prediction model 112 in the learning model storage unit 110. The second learning unit 1302 may store the attribute prediction model 112 trained in the i-th time step in the learning model storage unit 110 so that the model can be identified as a learning model for predicting the i-th attribute from the first attribute. Then, in S607, the learning unit 103 determines whether the parameter i is N or not. If i is N (Yes in S607), the process proceeds to S608, and if not (No in S607), the process proceeds to S604.

[0064] In S608, the second learning unit 1032 stores the learned attribute prediction model 112 in the learning model storage unit 110 as a learning model for predicting the Nth attribute from the first attribute.

[0065] (1-2) Learning process of attribute prediction model 112 (variation example) Next, a modified example of the process of (1-1) described above will be described. Fig. 6B shows a flowchart of a modified example of the learning process of the attribute prediction model 112 executed by the information processing device 100. As described above, this modified example is a process for learning only the attribute prediction model 112 for predicting the Nth attribute from the first attribute.

[0066] In S611, the first learning unit 1031 uses the first attribute learning data 121-1 to the (N-1)th attribute learning data 121-(N-1) to train the first attribute dedicated model 111-1 to the (N-1)th attribute dedicated model 111-(N-1), respectively. The training procedure in S611 is as described with reference to FIG. 5A. After training the first attribute dedicated model 111-1 to the (N-1)th attribute dedicated model 111-(N-1), the first learning unit 1031 stores the trained first attribute dedicated model 111-1 to the (N-1)th attribute dedicated model 111-(N-1) in the training model storage unit 110.

[0067] In S612, the second learning unit 1032 trains the attribute prediction model 112 using the trained first attribute dedicated model 111-1 to the (N-1)th attribute dedicated model 111-(N-1), a portion of each of the first attribute learning data 121-1 to the (N-1)th attribute learning data 121-(N-1), and the Nth attribute learning data 121-N. The learning procedure in S612 is as described with reference to FIG. 5D.

[0068] In S613, the second learning unit 1032 stores the learned attribute prediction model 112 in the learning model storage unit 110 as a learning model for predicting the Nth attribute from the first attribute.

[0069] (2) Prediction of product attributes contained in product images After the above-described process (1-1) or (1-2), the information processing device 100 predicts the attribute of the product included in the product image. Fig. 7 shows a flowchart of the attribute prediction process executed by the information processing device 100. In this process, a trained attribute prediction model 112 for predicting the Nth attribute (N is an integer equal to or greater than 2) from the first attribute is used.

[0070] In S71, the image acquisition unit 104 acquires a product image including a target product (a product for which attributes are predicted). For example, the image acquisition unit 104 acquires the product image by an operator of the information processing device 100 operating the information processing device 100 to access an arbitrary electronic commerce site and then selecting a product image including the target product. The image acquisition unit 104 can also acquire a product image by acquiring a product image or a URL (Uniform Resource Locator) indicating a product image transmitted from an external device such as a user device.

[0071] In S72, the attribute prediction unit 105 inputs the product image acquired by the image acquisition unit 104 to the trained attribute prediction model 112, and predicts the first to second attributes of the target product included in the product image. N The attribute prediction unit 105 predicts N attributes of the attributes. Specifically, the attribute prediction unit 105 extracts the N attribute types of the target product, and predicts the characteristics of each of the N attribute types.

[0072] In S73, the output unit 106 outputs information on the attributes predicted by the attribute prediction unit 105 (attribute prediction result). Specifically, the output unit 106 outputs the N attribute types predicted by the attribute prediction unit 105 and the features (feature values) of each of the N attribute types. The output unit 106 may display the attribute prediction result on the display unit 26 (FIG. 2). In addition, when a product image is acquired from an external device such as a user device, the output unit 106 may transmit the product image to the external device via the communication I / F 27 (FIG. 2) so as to be displayed on the display unit of the external device.

[0073] In this way, when the information processing device 100 according to the present embodiment trains the attribute prediction model 112 for predicting multiple attributes of a product from a product image including the product, the information processing device 100 trains one or more attributes that have already been trained using a dedicated training model trained specifically for the one or more attributes. This makes it possible to train the attribute prediction model 112 with new attributes while suppressing catastrophic forgetting.

[0074] In this embodiment, the learning procedure of the learning model and the attribute prediction procedure for predicting multiple attributes of a product when the product image includes one product have been described. When the product image includes multiple products, the information processing device 100 may be configured to divide the product image into partial images including each of the multiple products by a known image recognition process, and predict multiple attributes for each partial image by the procedure described in this embodiment.

[0075] Although specific embodiments have been described above, the embodiments are merely illustrative and are not intended to limit the scope of the present invention. The apparatus and method described in this specification can be embodied in forms other than those described above. Furthermore, the above-described embodiments can be appropriately omitted, substituted, and modified without departing from the scope of the present invention. Such omitted, substituted, and modified forms are included in the scope of the claims and their equivalents, and belong to the technical scope of the present invention.

[0076] The disclosure of this embodiment includes the following configuration. [1] An information processing device having a learning unit that trains an attribute prediction model for predicting N attributes (N is an integer equal to or greater than 2) of a product from an image including the product, the learning unit having: a first learning unit that trains an i-th learning model for predicting an i-th attribute of the product (i is an integer from 1 to N-1) using i-th learning data for the i-th attribute; and a second learning unit that trains the attribute prediction model using the i-th learning model trained by the first learning unit, a portion of the i-th learning data, and N-th learning data for the N-th attribute of the product.

[0077] [2] The information processing device described in [1], wherein the first learning unit trains the i-th learning model using an estimation result of the i-th attribute obtained by applying the i-th learning data to the i-th learning model.

[0078] [3] The information processing device described in [2], wherein the first learning unit trains the i-th learning model using an evaluation value obtained from the estimation result of the i-th attribute.

[0079] [4] The information processing device according to [3], wherein the evaluation value is cross entropy.

[0080] [5] The information processing device described in any of [1] to [4], wherein the second learning unit trains the attribute prediction model using a first estimation result for each of the first attribute to the (N-1)th attribute obtained by applying a portion of the i-th learning data to the i-th learning model, and a second prediction result for each of the first attribute to the N-th attribute obtained by applying a portion of the i-th learning data and the N-th learning data to the attribute prediction model.

[0081] [6] The information processing device described in [5], wherein the second learning unit trains the attribute prediction model using a first estimation result for each of the first attribute to the (N-1) attributes, (N-1) first evaluation values ​​obtained by a second prediction result for each of the first attribute to the (N-1) attributes, and a second evaluation value obtained by the second prediction result for the N attribute.

[0082] [7] The information processing device according to [6], wherein each of the (N-1) first evaluation values ​​is a KL divergence, and the second evaluation value is a cross entropy.

[0083] [8] The information processing device described in [7], wherein the second learning unit trains the attribute prediction model using a loss function based on a value obtained by multiplying each of the (N-1) first evaluation values ​​by a weight α and adding them up, and a value obtained by multiplying the second evaluation value by a weight (1-α).

[0084] [9] An information processing device described in any of [1] to [8], further comprising: an acquisition unit that acquires an image including a target product; and a prediction unit that inputs the image acquired by the acquisition unit into the attribute prediction model trained by the second learning unit, and predicts the first attribute to the Nth attribute of the target product.

[0085]

[10] An information processing method for training an attribute prediction model for predicting N attributes (N is an integer equal to or greater than 2) of a product from an image including the product, the information processing method comprising: training an ith learning model for predicting an ith attribute (i is an integer from 1 to N-1) of the product using ith learning data for the ith attribute; and training the attribute prediction model using the ith learning model trained by the first learning unit, a portion of the ith learning data, and Nth learning data for the Nth attribute of the product.

[0086]

[11] An information processing program for training an attribute prediction model for predicting N attributes (N is an integer equal to or greater than 2) of a product from an image including the product, the information processing program causing a computer to execute the following steps: training an ith learning model for predicting an ith attribute (i is an integer from 1 to N-1) of the product using ith learning data for the ith attribute; and training the attribute prediction model using the ith learning model trained by the first learning unit, a portion of the ith learning data, and Nth learning data for the Nth attribute of the product. [Explanation of symbols]

[0087] 100: information processing device, 101: learning data acquisition unit, 102: learning data management unit, 103: learning unit, 1301: first learning unit, 1302: second learning unit, 104: image acquisition unit, 105: attribute prediction unit, 106: output unit, 110: learning model storage unit, 111-1: first attribute specialized model, 111-N: Nth attribute dedicated model, 112: attribute prediction model, 120: attribute data storage unit, 121-1: first attribute learning data, 121-N: Nth attribute learning data

Claims

1. a learning unit that learns an attribute prediction model for predicting N attributes (N is an integer equal to or greater than 2) of a product from an image including the product, The learning unit is a first learning unit that learns an i-th learning model for predicting an i-th attribute (i is an integer from 1 to N-1) of the product using i-th learning data for the i-th attribute; a second learning unit that learns the attribute prediction model by using the i-th learning model learned by the first learning unit, a portion of the i-th learning data, and N-th learning data for the N-th attribute of the product; having The second learning unit is a first estimation result for each of the first attribute to the (N−1)th attribute obtained by applying a portion of the i-th training data to the i-th training model; and second estimation results of the first attribute to the N-th attribute obtained by applying a portion of the i-th training data and the N-th training data to the attribute prediction model; and The information processing device trains the attribute prediction model using the above.

2. The first learning unit learns the i-th learning model using an estimation result of the i-th attribute obtained by applying the i-th learning data to the i-th learning model. The information processing device according to claim 1 .

3. The first learning unit learns the i-th learning model by using an evaluation value obtained from the estimation result of the i-th attribute. The information processing device according to claim 2 .

4. The information processing apparatus according to claim 3 , wherein the evaluation value is a cross entropy.

5. the second learning unit learns the attribute prediction model using a first estimation result for each of the first attribute to the (N-1) attributes, (N-1) first evaluation values ​​obtained by a second estimation result for each of the first attribute to the (N-1) attributes, and a second evaluation value obtained by a second estimation result for the N attribute; The information processing device according to claim 1 .

6. Each of the (N-1) first evaluation values ​​is a KL divergence, and each of the second evaluation values ​​is a cross entropy. The information processing device according to claim 5 .

7. The information processing device according to claim 6, wherein the second learning unit trains the attribute prediction model using a loss function based on a value obtained by multiplying each of the (N-1) first evaluation values ​​by a weight α and adding them up, and a value obtained by multiplying the second evaluation value by a weight (1-α).

8. An acquisition unit that acquires an image including a target product; a prediction unit that inputs the image acquired by the acquisition unit into the attribute prediction model trained by the second learning unit to predict a first attribute to an Nth attribute of the target product; Further comprising The information processing device according to claim 1 .

9. An information processing method executed by an information processing device for learning an attribute prediction model for predicting N attributes (N is an integer equal to or greater than 2) of a product from an image including the product, the method comprising: training an i-th learning model for predicting an i-th attribute (i is an integer from 1 to N-1) of the product using i-th learning data for the i-th attribute; training the attribute prediction model using the trained i-th learning model, a portion of the i-th learning data, and N-th learning data for the N-th attribute of the product; Including, Training the attribute prediction model includes: a first estimation result for each of the first attribute to the (N−1)th attribute obtained by applying a portion of the i-th training data to the i-th training model; and second estimation results of the first attribute to the N-th attribute obtained by applying a portion of the i-th training data and the N-th training data to the attribute prediction model; and training the attribute prediction model using Information processing methods.

10. An information processing program for training an attribute prediction model for predicting N attributes (N is an integer equal to or greater than 2) of a product from an image including the product, the information processing program comprising: training an i-th learning model for predicting an i-th attribute (i is an integer from 1 to N-1) of the product using i-th learning data for the i-th attribute; training the attribute prediction model using the trained i-th learning model, a portion of the i-th learning data, and N-th learning data for the N-th attribute of the product; Run the following on your computer: Training the attribute prediction model includes: a first estimation result for each of the first attribute to the (N−1)th attribute obtained by applying a portion of the i-th training data to the i-th training model; and second estimation results of the first attribute to the N-th attribute obtained by applying a portion of the i-th training data and the N-th training data to the attribute prediction model; and The information processing program includes training the attribute prediction model using the above-mentioned method.

Citation Information

Patent Citations

  • Human body attribute recognition model training method, human body attribute recognition method and related device

    CN115082963A

  • Information processing apparatus, information processing method, and information processing program

    JP2018106284A

  • Information processing system, information processing apparatus, server device, program, or method

    JP2020071859A

  • Information processing device and information processing method

    JP2022185558A

  • Learning method and device for specializing an artificial intelligence model for a user organization

    JP2022550094A