Information processing method, program, and information processing device
Patent Information
- Application Number
- JP2023522685
- Authority / Receiving Office
- JP · JP
- Patent Type
- Applications
- Current Assignee / Owner
- Priority Date
- 2022-05-17
- Filing Date
- 2022-05-17
- Publication Date
- 2025-05-14
- Estimated Expiration
- 2042-05-17
AI Technical Summary
Conventional machine learning-based object recognition systems in image processing face inefficiencies in identifying features that lead to inaccurate estimation results, hindering the generation of effective training data.
An information processing method and device that evaluates the estimation results of a learning model by generating and analyzing first evaluation images with specific parameters, identifying features likely to cause incorrect estimations, and using this information to generate targeted learning data to improve model training.
This approach enables efficient identification and rectification of weaknesses in object recognition, allowing for improved learning data generation without the need for extensive training datasets, thus reducing annotation work and enhancing the accuracy of object recognition.
Abstract
Description
Information processing method, program, and information processing device Cross-reference to related applications
[0001] This application claims priority to Japanese Patent Application No. 2021-084954, filed on May 19, 2021, the entire disclosure of which is incorporated herein by reference.
[0002] The present disclosure relates to an information processing method, a program, and an information processing device.
[0003] In recent years, there has been progress in the development of technology that uses machine learning to recognize objects contained in images. Such technology requires a large amount of image data as training data when training a model. Therefore, technology for generating training data has been developed.
[0004] For example, Patent Literature 1 describes generating a plurality of training set images including one or a plurality of products by randomly arranging individual images. Patent Literature 1 also describes that the generated plurality of training set images include training set images in which the individual images at least partially overlap each other.
[0005] For example, Patent Document 2 describes creating a composite image that serves as training data for machine learning by combining a background image and a patch image of a target object based on the probability set in a created target object existence probability map.
[0006] JP 2020-80003 A JP 2018-88223 A
[0007] An information processing method according to one embodiment of the present disclosure includes: obtaining an evaluation result indicating whether the estimation result of a learning model for the first evaluation image is correct or incorrect, based on first evaluation data including data of at least one first evaluation image and correct answer data for the first evaluation image; and performing a specification process, based on the evaluation result, to identify image features that are likely to cause the estimation result of the learning model to be incorrect.
[0008] A program according to one embodiment of the present disclosure causes a computer to: obtain an evaluation result indicating whether the estimation result of a learning model for a first evaluation image is correct or incorrect, based on first evaluation data including data of at least one first evaluation image and correct answer data for the first evaluation image; and perform a specification process based on the evaluation result to identify image features that are likely to cause the estimation result of the learning model to be incorrect.
[0009] An information processing device according to one embodiment of the present disclosure includes a control unit that obtains an evaluation result indicating whether the estimation result of a learning model for a first evaluation image is correct or incorrect based on first evaluation data including data of at least one first evaluation image and correct answer data for the first evaluation image, and the control unit performs a specification process based on the evaluation result to identify image features that are likely to result in an incorrect estimation result of the learning model.
[0010] FIG. 1 is a diagram showing a schematic configuration of a settlement system according to an embodiment of the present disclosure. FIG. 2 is an external view showing the configuration of the information processing system shown in FIG. 1. FIG. 3 is a functional block diagram showing the configuration of the information processing device shown in FIG. 2. FIG. 4 is a diagram showing an example of a first evaluation image according to an embodiment of the present disclosure. FIG. 5 is a diagram showing an example of parameter setting according to an embodiment of the present disclosure. FIG. 6 is a diagram showing another example of parameter setting according to an embodiment of the present disclosure. FIG. 7 is a flowchart showing the operation of a learning support process executed by the information processing system shown in FIG. 1. FIG. 8 is a diagram showing another example of a first evaluation image according to an embodiment of the present disclosure.
[0011] Conventional techniques have room for improvement. For example, if it were possible to identify image features that are likely to cause inaccurate estimation results from a learning model, it would be possible to generate training data more efficiently. The present disclosure provides an improved technology for supporting learning.
[0012] Hereinafter, embodiments of the present disclosure will be described with reference to the drawings. In the following drawings, the same components are denoted by the same reference numerals.
[0013] (System Configuration) The settlement system 1 shown in Fig. 1 is configured as a POS (Point Of Sales) system. The settlement system 1 includes at least one information processing system 3 and a server 4. In this embodiment, the settlement system 1 includes multiple information processing systems 3.
[0014] The information processing system 3 and the server 4 can communicate with each other via a network 2. The network 2 may be any network including the Internet.
[0015] The information processing system 3 may be installed in any store, such as a shop or a restaurant.
[0016] The information processing system 3 is configured as a cash register terminal of a POS system. The information processing system 3 captures an image of a commodity placed on the cash register terminal by a customer. The information processing system 3 performs object recognition on the captured image and estimates which commodity in the store the object contained in the image corresponds to. In this disclosure, "object contained in the image" refers to an object depicted in the image. In this disclosure, the portion depicted as an object in the image, i.e., the portion in the image where the object is depicted, is also referred to as an "object image." By estimating which commodity in the store the placed object corresponds to, the information processing system 3 can calculate the amount to be charged to the customer. The information processing system 3 transmits an estimation result indicating which commodity in the store the placed object corresponds to to the server 4 via the network 2.
[0017] The server 4 receives an estimation result indicating which product in the store the placed object is, from the information processing system 3 via the network 2. Based on the estimation result, the server 4 manages the inventory status of the store in which the information processing system 3 is installed.
[0018] 2 , the information processing system 3 includes an imaging unit 12 and an information processing device 20. The information processing system 3 may further include a mounting table 10, a support column 11, and a display device 13.
[0019] The table 10 includes an upper surface 10s. When paying, a customer places the product they wish to purchase on the upper surface 10s. In this embodiment, the upper surface 10s is substantially rectangular. However, the upper surface 10s may have any shape.
[0020] The support pillar 11 supports the imaging unit 12. The support pillar 11 extends from the side of the mounting table 10 toward above the top surface 10s.
[0021] The imaging unit 12 generates an image signal corresponding to an image by capturing an image. The imaging unit 12 is fixed so as to be able to capture an image of at least a portion of the surface of the mounting table 10. The imaging unit 12 may be fixed so that its optical axis is perpendicular to the upper surface 10s. For example, the imaging unit 12 is fixed to, for example, the tip of the support column 11 so that it can capture an image of the entire upper surface 10s of the mounting table 10 and so that its optical axis is perpendicular to the upper surface 10s. The imaging unit 12 may continuously capture images at any frame rate.
[0022] The display device 13 may be any display. The display device 13 displays an image corresponding to the image signal transmitted from the information processing device 20. The display device 13 may function as a touch screen.
[0023] 3 , the information processing device 20 includes a communication unit 21, an input unit 22, a storage unit 23, and a control unit 24. In this embodiment, the information processing device 20 is configured as a device separate from the imaging unit 12 and the display device 13. However, the information processing device 20 may be configured integrally with at least one of the imaging unit 12, the support column 11, the mounting base 10, and the display device 13, for example.
[0024] The communication unit 21 includes at least one communication module connectable to the network 2. The communication module is, for example, a communication module compatible with standards such as a wired LAN (Local Area Network) or a wireless LAN. The communication unit 21 is connected to the network 2 via the wired LAN or wireless LAN by the communication module.
[0025] The communication unit 21 includes a communication module that can communicate with the image capture unit 12 and the display device 13 via a communication line. The communication module is a communication module that complies with the standard of the communication line. The communication line is configured to include at least one of a wired and a wireless communication line.
[0026] The input unit 22 is capable of receiving input from a user. The input unit 22 includes at least one input interface capable of receiving input from a user. The input interface is, for example, a physical key, a capacitance key, a pointing device, a touch screen integrated with a display, or a microphone. In this embodiment, the input unit 22 is a touch screen integrated with the display device 13.
[0027] The storage unit 23 includes at least one semiconductor memory, at least one magnetic memory, at least one optical memory, or a combination of at least two of these. The semiconductor memory may be, for example, a random access memory (RAM) or a read-only memory (ROM). The RAM may be, for example, a static random access memory (SRAM) or a dynamic random access memory (DRAM). The ROM may be, for example, an electrically erasable programmable read-only memory (EEPROM). The storage unit 23 may function as a main storage device, an auxiliary storage device, or a cache memory. The storage unit 23 stores data used in the operation of the information processing device 20 and data obtained by the operation of the information processing device 20. For example, the storage unit 23 stores system programs, application programs, embedded software, and the like. For example, the storage unit 23 stores a learning model.
[0028] The control unit 24 is configured to include at least one processor, at least one dedicated circuit, or a combination of these. The processor is a general-purpose processor such as a CPU (Central Processing Unit) or a GPU (Graphics Processing Unit), or a dedicated processor specialized for specific processing. The dedicated circuit is, for example, an FPGA (Field-Programmable Gate Array) or an ASIC (Application Specific Integrated Circuit). The control unit 24 controls each unit of the information processing device 20 and executes processing related to the operation of the information processing device 20.
[0029] [Checkout Processing] The control unit 24 receives an image signal from the imaging unit 12 via the communication unit 21. By receiving the image signal, the control unit 24 acquires an image corresponding to the image signal. The control unit 24 acquires an inference result indicating which product in the store an object included in the image corresponds to by object recognition using a learning model. The learning model is generated by machine learning such as deep learning so as to output an inference result when image data is input. The learning model may assign a reliability to the inference result. The reliability is an index indicating the reliability of the inference result. The higher the reliability, the higher the reliability of the inference result. The control unit 24 inputs the image data into the learning model and acquires an inference result output from the learning model. The control unit 24 calculates the amount to be invoiced to the customer based on the acquired inference result. The control unit 24 transmits a signal indicating information about the invoice amount to the display device 13 via the communication unit 21, and causes the display device 13 to display the information about the invoice amount.
[0030] [Identification Process] The control unit 24 identifies weaknesses in the object recognition of the learning model at any timing, such as before or after the operation of the information processing device 20. Identifying weaknesses in the object recognition of the learning model allows the learning model to be efficiently retrained or trained. To identify weaknesses in the object recognition of the learning model, the control unit 24 executes an identification process to identify image features that are likely to cause the estimation result of the learning model to be incorrect. An example of this process will be described below.
[0031] The control unit 24 generates first evaluation data. The first evaluation data includes data of at least one first evaluation image and correct answer data corresponding to the first evaluation image. The correct answer data is, for example, data indicating which product in the store the object in the first evaluation image is. The first evaluation data may include data of multiple first evaluation images and correct answer data corresponding to each of the multiple first evaluation images.
[0032] The control unit 24 sets parameters and generates at least one first evaluation image based on the set parameters. The control unit 24 may generate multiple first evaluation images. The parameters are for setting elements that constitute an image. The parameters correspond to the difficulty of estimating an object in an image. The control unit 24 sets the difficulty level by setting the parameters. The control unit 24 may set the parameters based on the amount of learning data that the learning model has already learned. For example, the control unit 24 sets the parameters so that the difficulty level increases as the amount of learning data that the learning model has already learned increases. When the learning process and the identification process are executed in parallel in curriculum learning, as described below, the control unit 24 may generate the first evaluation image based on the parameters set in the first process.
[0033] Various parameters may be employed. The parameter may be set appropriately depending on the parameter employed. For example, the parameter may be set as a set value such as a numerical value, or may be set as a level. The parameter may include at least one of a parameter set for an object in an image and a parameter set for the environment of the object in the image. A combination of multiple arbitrary parameters may be employed. For example, the following parameters may be employed:
[0034] The parameter may be the number of objects to be detected in the image. The objects to be detected may be, for example, merchandise in a store. The fewer the number of objects to be detected, the lower the difficulty of estimating the objects in the image. The more the number of objects to be detected, the higher the difficulty of estimating the objects in the image.
[0035] The parameter may be the number of objects that are not the detection target. Objects that are not the detection target are, for example, objects other than merchandise in a store. For example, objects that are not the detection target are a customer's fingers, a mobile phone, keys, etc. The fewer the number of objects that are not the detection target, the lower the difficulty of estimating objects in an image. The more the number of objects that are not the detection target, the higher the difficulty of estimating objects in an image.
[0036] The parameter may be the degree of reflected light. Reflected light is one of the parameters set for an object in an image. The smaller the degree of reflected light, the lower the difficulty of estimating the object for which reflected light is set. The greater the degree of reflected light, the higher the difficulty of estimating the object for which reflected light is set.
[0037] The parameter may be an overlap rate. The overlap rate is one of the parameters set for an object in an image. The overlap rate may be set for two or more objects in an image. The overlap rate indicates the degree of overlap between two or more object images in an image. The smaller the overlap rate, the lower the difficulty of estimating the object for which the overlap rate is set. The larger the overlap rate, the higher the difficulty of estimating the object for which the overlap rate is set.
[0038] As an example, the overlap rate may be the ratio of the area of the portion of the lower object image that is overlapped by the upper object image to the area of the lower object image in two or more object images that overlap vertically in an image. In this case, for two overlapping object images 30a and 30b as shown in Figure 4 described below, the overlap rate is the ratio of the area of the portion of the object image 30b that is overlapped by the object image 30a to the area of the object image 30b.
[0039] As another example, the overlap rate may be the area of the portion of the detection frame of the lower object image that overlaps with the detection frame of the upper object image, relative to the area of the detection frame of the lower object image, in two or more object images that overlap vertically in an image. In this case, for two overlapping object images 30a and 30b as shown in FIG. 4 (described later), the overlap rate is the ratio of the area of the portion of detection frame 30b1 that overlaps with detection frame 30a1, relative to the area of detection frame 30b1. Detection frame 30a1 is the detection frame of object image 30a. Detection frame 30b1 is the detection frame of object image 30b.
[0040] The parameter may be the hue of the background image. The hue of the background image is one of the parameters set for the environment of an object in the image. The more the hue of the background image is different from the hue of the object, the lower the difficulty of estimating the object in the image. The more the hue of the background image is closer to the hue of the object, the higher the difficulty of estimating the object in the image.
[0041] The parameter may be the pattern of the background image. The pattern of the background image is one of the parameters set for the environment of the object in the image. The simpler the pattern of the background image, the lower the difficulty of estimating the object in the image. The more complex the pattern of the background image, the higher the difficulty of estimating the object in the image.
[0042] The parameter may be the hue of the illumination light. The hue of the illumination light is one of the parameters set for the environment of the object in the image. The closer the hue of the illumination light is to white or a warm hue, the lower the difficulty of estimating the object in the image. The further the hue of the illumination light is from white and warm hue, the higher the difficulty of estimating the object in the image.
[0043] The parameter may be the illuminance of the illumination light. The illuminance of the illumination light is one of the parameters set for the environment of the object in the image. The further the illuminance of the illumination light deviates from the set range, the more difficult it becomes to estimate the object in the image. The set range may be set appropriately based on the band of electromagnetic waves that the imaging unit 12 can capture, etc.
[0044] An example of a process for generating a first evaluation image based on the set parameters will be described below. The control unit 24 generates the first evaluation image by, for example, adjusting an object image, a background image, etc. based on the set parameters. The control unit 24 may generate the first evaluation image by using a cut-and-paste method on an existing image. The cut-and-paste method is a method for generating an image by cutting an object image from an existing image and pasting it onto a background image, etc. An image including an object image may be used as the existing image. The object image included in the existing image may be labeled with a label indicating the object corresponding to the object image. For example, the control unit 24 cuts out the object image from the existing image according to the label. The control unit 24 generates the first evaluation image by pasting the cut-out object image onto the background image while adjusting the object image, background image, etc. based on the parameters.
[0045] For example, the control unit 24 generates a first evaluation image 30 as shown in Fig. 4. The first evaluation image 30 includes an object image 30a corresponding to a rice ball, an object image 30b corresponding to butter, an object image 30c corresponding to chocolate, and a background image 30d. The rice ball, butter, and chocolate are products in the store, i.e., objects to be detected. In this example, the number of objects to be detected, the overlap rate, and the illuminance of the illumination light are used as parameters.
[0046] 4, the control unit 24 sets the number of objects to be detected as a parameter to 3. The control unit 24 randomly determines each of the three objects to be detected to be rice ball, butter, and chocolate.
[0047] 4, the control unit 24 sets the overlap rate as a parameter to 40%. The control unit 24 randomly determines the two objects to which the overlap rate of 40% is to be set to be the rice ball and the butter. In setting the overlap rate, the control unit 24 determines that, of the object image 30a and the object image 30b, the object image 30b is to be placed on the bottom.
[0048] 4, the control unit 24 sets the illuminance of the illumination light as a parameter to hard. Setting the illuminance of the illumination light to hard indicates that the illuminance of the illumination light is higher than the above-mentioned setting range.
[0049] 4, the control unit 24 generates the first evaluation image 30 by using a cut-and-paste method. For example, the control unit 24 cuts out each of the object images 30a, 30b, and 30c from different existing images or the same existing image. The control unit 24 pastes each of the cut-out object images 30a, 30b, and 30c onto the background image 30d. In this case, the control unit 24 adjusts the overlap rate between the object image 30a and the object image 30b to 40%.
[0050] The control unit 24 obtains an evaluation result of the learning model using the first evaluation data. The evaluation result indicates whether the estimation result of the learning model for the first evaluation image is correct or incorrect. For example, the control unit 24 inputs data of the first evaluation image into the learning model and obtains the estimation result of the learning model. The control unit 24 generates and obtains the evaluation result by comparing the obtained estimation result with the correct data corresponding to the first evaluation image.
[0051] The control unit 24 identifies features of the image that are likely to result in an incorrect estimation result of the learning model based on the evaluation result. In this embodiment, the control unit 24 identifies the features of the image by acquiring feature information indicating the features of the image that are likely to result in an incorrect estimation result of the learning model based on the evaluation result. The control unit 24 may acquire, as feature information, at least any of the parameters set in generating the first evaluation image, including parameters set for the incorrect object and parameters set for the environment of the incorrect object, based on the evaluation result.
[0052] For example, the control unit 24 inputs data of a first evaluation image 30 as shown in FIG. 4 into the learning model. The control unit 24 obtains, as estimation results from the learning model, an estimation result that the object corresponding to object image 30a is bread, an estimation result that the object corresponding to object image 30b is cheese, and an estimation result that the object corresponding to object image 30c is chocolate. In this case, the control unit 24 obtains, as evaluation results, an evaluation result indicating that the estimation result for object image 30a is incorrect, an evaluation result indicating that the estimation result for object image 30b is incorrect, and an evaluation result indicating that the estimation result for object image 30c is correct. The control unit 24 obtains, as feature information, parameters set for the incorrect object, i.e., the 40% overlap rate set for rice balls and butter. The control unit 24 also obtains parameters set for the environment of the incorrect object, i.e., the illuminance setting of the lighting for the hardware. In other words, in an image containing rice balls and butter with an overlap rate set to 40% and the illuminance of the lighting light set to hard, there is a high possibility that the estimation results of the learning model for rice balls and butter will be incorrect.
[0053] When the control unit 24 acquires the feature information, it may generate a signal indicating the feature information. The control unit 24 may transmit the generated signal to the display device 13 via the communication unit 21, causing the display device 13 to display the feature information. The control unit 24 may cause the display device 13 to display the feature information in language. For example, the control unit 24 may cause the display device 13 to display the feature information in language such as, "In an image where the illuminance of the illumination light is high and the overlap rate between the rice ball and the butter is 40% or more, the estimation results of the learning model for the rice ball and the butter are likely to be incorrect." By displaying the feature information on the display device 13, the operator of the information processing device 20 can understand the weaknesses of the learning model in object recognition. By understanding the weaknesses of the learning model in object recognition, the operator can prepare learning data suitable for the learning model.
[0054] Upon acquiring the feature information, the control unit 24 may generate first training data. The first training data is data to be used by the training model to eliminate weaknesses in object recognition of the training model. The first training data includes data of at least one first training image and correct answer data corresponding to the first training image. The correct answer data is, for example, data indicating which product in a store an object in the first training image is. The first training data may include data of a plurality of first training images and correct answer data corresponding to each of the plurality of first training images.
[0055] The control unit 24 may generate at least one first training image of the first training data based on the feature information. The control unit 24 may generate multiple first training images of the first training data. When the control unit 24 acquires parameters as feature information, the control unit 24 may generate the first training image using an object image corresponding to an incorrect object and the acquired parameters. For example, the control unit 24 may acquire, as feature information, a 40% overlap rate and hard illumination light setting for rice balls and butter. In this case, the control unit 24 adjusts the overlap rate of object images 30a and 30b as shown in FIG. 4 to 40% and adjusts the image to hard illumination light, thereby generating the first training image. In the same manner as or similar to generating the first evaluation image, the control unit 24 may generate the first training image by using a cut-and-paste method on an existing image.
[0056] The control unit 24 may train the learning model using the generated first learning data. With this configuration, it is possible to eliminate the weakness of the learning model in object recognition.
[0057] [Learning Process] The control unit 24 may execute the identification process in parallel with a preset learning process. By executing the identification process in parallel with the learning process, it is possible to identify points that the learning model was not able to sufficiently learn in the learning process as weaknesses in the object recognition of the learning model. In the present disclosure, the learning data used in the learning process executed in parallel with the identification process is also referred to as "second learning data."
[0058] In this embodiment, the control unit 24 executes the specific processing in parallel with the learning processing in curriculum learning. Curriculum learning is a learning method in which the difficulty of problems to be learned by the learning model is gradually increased from low to high. The curriculum learning processing according to this embodiment will be described below.
[0059] The control unit 24 repeatedly executes a learning process in the curriculum learning. The repeatedly executed learning process includes a first process, a second process, and a third process.
[0060] The first process is a process for setting parameters corresponding to the above-described difficulty level. In the repeatedly executed learning process, the control unit 24 sets the parameters in the latest first process so that the difficulty level is at least one level higher than the difficulty level corresponding to the parameters set in the previous first process. With this configuration, the difficulty level set as the parameters in the first process increases stepwise as the learning process is repeatedly executed.
[0061] As an example, the number of objects to be detected is adopted as the parameter. The learning process is also assumed to be repeated three times. In this case, in the first process, the control unit 24 sets the number of objects to be detected. For example, in the first process of the first learning process, the control unit 24 sets the number of objects to be detected to one. In the first process of the second learning process, the control unit 24 sets the number of objects to be detected within a range from two to (2 / M), where M is the maximum number of objects to be detected that can be set in one image. In the first process of the third learning process, the control unit 24 sets the number of objects to be detected within a range from (2 / M) to M. In this example, instead of the number of objects to be detected, the number of objects including objects to be detected and objects not to be detected may be adopted as the parameter.
[0062] In the first process, a combination of multiple arbitrary parameters may be employed. That is, in the first process, the control unit 24 may combine multiple arbitrary parameters. In this case, the control unit 24 sets the parameters in the first process so that the overall difficulty level based on the combined multiple parameters increases in stages.
[0063] As an example, as shown in FIG. 5 , a combination of the number of objects to be detected and the illuminance of the illumination light is used as a combination of multiple parameters. In FIG. 5 , the learning process is repeated six times. The control unit 24 sets the number of objects to be detected and the illuminance of the illumination light in the first process so that the overall difficulty level, determined by the number of objects to be detected and the illuminance of the illumination light, gradually increases as the learning process of the learning model progresses from the first to the sixth time. For example, in the first process of the first learning process, the control unit 24 sets the number of objects to be detected to one and the illuminance of the illumination light to normal. The normal setting of the illuminance of the illumination light indicates that the illuminance of the illumination light is within the above-mentioned setting range. In the first process of the second learning process, the control unit 24 sets the number of objects to be detected to one and the hard setting of the illuminance of the illumination light. The hard setting of the illuminance of the illumination light indicates that the illuminance of the illumination light is higher than the above-mentioned setting range, as described above. In the first process of the third learning process, the control unit 24 sets the number of objects to be detected to two and sets the illuminance of the illumination light to normal. In the first process of the fourth learning process, the control unit 24 sets the number of objects to be detected to two and sets the illuminance of the illumination light to hard. In the first process of the fifth learning process, the control unit 24 sets the number of objects to be detected to three and sets the illuminance of the illumination light to normal. In the first process of the sixth learning process, the control unit 24 sets the number of objects to be detected to three and sets the illuminance of the illumination light to hard.
[0064] As another example, as shown in FIG. 6 , a combination of the number of objects to be detected and the overlap rate is adopted as a combination of multiple parameters. In FIG. 6 , the learning process is repeated seven times. In the first process, the control unit 24 sets the number of objects to be detected and the overlap rate so that the overall difficulty level, determined by the number of objects to be detected and the overlap rate, gradually increases as the learning process of the learning model progresses from the first to the seventh time. For example, in the first process of the first learning process, the control unit 24 sets the number of objects to be detected to one and the overlap rate to zero. In the first process of the second learning process, the control unit 24 sets the number of objects to be detected to a range of two to (M / 2) and the overlap rate to zero. As described above, M is the maximum number of objects to be detected that can be set in one image. In the first process of the third learning process, the control unit 24 sets the number of objects to be detected to a range of two to (M / 2) and the overlap rate to a low level. The small level of the overlap rate is the smallest level when the overlap rate is defined at three levels. In the first process of the fourth learning process, the control unit 24 sets the number of objects to be detected within a range of 2 to (M / 2) and sets the overlap rate to a medium level. The medium level of the overlap rate is the intermediate level when the overlap rate is defined at three levels. In the first process of the fifth learning process, the control unit 24 sets the number of objects to be detected within a range of (M / 2) to M and sets the overlap rate to none. In the first process of the sixth learning process, the control unit 24 sets the number of objects to be detected within a range of (M / 2) to M and sets the overlap rate to a small level. In the first process of the seventh learning process, the control unit 24 sets the number of objects to be detected within a range of (M / 2) to M and sets the overlap rate to a medium level. In this example, the number of objects including the objects to be detected and the objects not to be detected may be used as a parameter instead of the number of objects to be detected. That is, in this example, a combination of the number of objects including the objects to be detected and the objects not to be detected and the overlap rate may be used as a combination of multiple parameters.
[0065] The second process is a process of generating second training data based on the parameters set in the first process. The second training data includes data of at least one second training image and correct answer data corresponding to the second training image. The correct answer data is, for example, data indicating which product in a store an object in the second training image is. The second training data may include data of multiple second training images and correct answer data corresponding to each of the multiple second training images. In the second process, the control unit 24 may generate at least one second training image based on the parameters set in the first process, in the same or similar manner as in generating the first evaluation image. The control unit 24 may generate multiple second training images. In the second process, the control unit 24 may generate the second training image by using a cut-and-paste method on an existing image, in the same or similar manner as in generating the first evaluation image.
[0066] In the second process, the control unit 24 may generate second learning images based on parameters newly set in the most recent first process. With this configuration, the difficulty of the questions using the second learning images generated in the second process increases stepwise as the learning process is repeatedly executed.
[0067] The third process is a process of training a learning model using the second training data generated in the second process.
[0068] Here, the control unit 24 may repeatedly execute the learning process until the estimation accuracy of the learning model satisfies the set conditions. The control unit 24 may acquire the estimation accuracy of the learning model used to determine whether the set conditions are satisfied using second evaluation data. The second evaluation data is evaluation data different from the first evaluation data. The second evaluation data includes data of at least one second evaluation image and correct answer data corresponding to the second evaluation image. The correct answer data is, for example, data indicating which product in the store the object in the second evaluation image is. The second evaluation data may include data of multiple second evaluation images and correct answer data corresponding to each of the multiple second evaluation images. The second evaluation image may be an image that is actually captured and generated. By using an image that is actually captured and generated as the second evaluation data, the estimation accuracy of the learning model can be measured more accurately. The control unit 24 may acquire mAP (mean average precision) as the estimation accuracy of the learning model.
[0069] The set condition may be a first condition that the first relevance rate exceeds a first threshold. The first relevance rate may be the estimation accuracy of the learning model calculated based on the estimation result assigned the highest reliability. The first relevance rate may be mAP@1, which will be described later. The first threshold may be set based on the estimation accuracy of the learning model targeted by the information processing system 3.
[0070] The set condition may be a second condition that the second precision exceeds a second threshold. The second precision may be the estimation accuracy of the learning model calculated based on the estimation result assigned the highest reliability and the estimation result assigned the second highest reliability. The second precision may be mAP@2, which will be described later. The second threshold may be set based on the estimation accuracy of the learning model required for operation of the information processing device 20. The second threshold is, for example, 100%.
[0071] For example, the control unit 24 calculates mAP@1 and mAP@2 using the following formula (1). mAP@n is the average value of AP@n. AP@n is the average value of the accuracy rates of the correct estimation results among the estimation results to which reliability has been assigned up to the nth (n is an integer equal to or greater than 1) reliability. If the set of pairs of estimation results and correct answers is q, AP@n is written as "AP(q)@n". The order of reliability is counted from highest to lowest. The control unit 24 calculates AP@n using the following formula (2). In formula (1), the number of questions Q is the number of questions based on the second learning image. In formula (2), the number of questions GTP is the number of questions that were correct in the estimation results to which the reliability levels up to the nth level have been assigned. The accuracy rate P@k is the number of estimation results to which the kth reliability level has been assigned that were correct relative to the number of questions in the estimation results to which the reliability levels up to the kth level have been assigned. The coefficient rel@k is 1 when the estimation result to which the kth reliability level has been assigned is correct. The coefficient rel@k is 0 when the estimation result to which the kth reliability level has been assigned is incorrect.
[0072] The set condition is not limited to the first condition and the second condition, but may be a condition that either the first condition or the second condition is satisfied, or a condition that both the first condition and the second condition are satisfied.
[0073] (System Operation) Fig. 7 is a flowchart showing the operation of the learning support process executed by the information processing system 3 shown in Fig. 1. This operation corresponds to an example of an information processing method according to this embodiment. The control unit 24 may execute the learning support process at any timing. For example, the control unit 24 may execute the learning support process before operation of the information processing device 20, or may execute the learning support process after operation of the information processing device 20.
[0074] The control unit 24 sets the parameters (step S10). The processing of step S10 corresponds to the first processing. If the processing of step S10 is the first time, the control unit 24 sets the parameters to initial values. The initial values of the parameters may be set based on the amount of learning data that the learning model has already learned. If the processing of step S10 has already been executed, the control unit 24 sets the parameters so that the difficulty level is at least one level higher than the difficulty level corresponding to the parameters set in the previous processing of step S10.
[0075] The control unit 24 generates second learning data based on the parameters set in the process of step S10 (step S11). The process of step S11 corresponds to the second process. The control unit 24 generates second learning images based on the parameters set in the process of step S10.
[0076] The control unit 24 trains a learning model using the second learning data generated in the process of step S11 (step S12). The process of step S12 corresponds to the third process. If the process of step S16 described below has already been executed, in the process of step S12, the control unit 24 trains a learning model using the first learning data generated in the process of step S16 and the second learning data generated in the process of step S11.
[0077] The control unit 24 generates first evaluation data based on the parameters set in the process of step S10 (step S13). The control unit 24 generates a first evaluation image based on the parameters set in the process of step S10.
[0078] The control unit 24 acquires the evaluation result of the learning model using the first evaluation data generated in the processing of step S13 (step S14).
[0079] Based on the evaluation result acquired in the process of step S14, the control unit 24 acquires at least one of the parameters set for the object for which the estimation result of the learning model was incorrect and the parameters set for the environment of the object (step S15). That is, the control unit 24 acquires at least one of the parameters set for the object for which the estimation result of the learning model was incorrect and the parameters set for the environment of the object from among the multiple parameters set in the process of step S10.
[0080] The control unit 24 generates first learning data based on the parameters acquired in the process of step S15 (step S16). The control unit 24 generates a first learning image based on the parameters acquired in the process of step S15. The first learning data generated in the process of step S16 is used to train the learning model in the next process of step S12.
[0081] The control unit 24 acquires the estimation accuracy of the learning model used to determine whether the set condition is satisfied from the second evaluation data (step S17).
[0082] The control unit 24 determines whether the estimation accuracy of the learning model acquired in the processing of step S17 satisfies the set conditions (step S18). If the control unit 24 determines that the estimation accuracy of the learning model satisfies the set conditions (step S18: YES), it ends the learning support processing. On the other hand, if the control unit 24 does not determine that the estimation accuracy of the learning model satisfies the set conditions (step S8: NO), it returns to the processing of step S10.
[0083] Here, in the process of step S10, the control unit 24 may set parameters so that the difficulty level is several levels higher than the difficulty level corresponding to the parameters set in the previous process of step S10. In this case, in the process of step S15, the control unit 24 may identify the type of at least one of the parameters set for the object for which the learning model's estimation result was incorrect and the parameters set for the environment of the object. In the process of step S16, the control unit 24 may identify a parameter range for the identified type of parameter, ranging from the set value in the previous process of step S10 to the set value in the latest process of step S10. Furthermore, in the process of step S16, the control unit 24 may generate a first learning image based on at least a portion of the parameter range. For example, it is assumed that an overlap rate is used as a parameter. It is also assumed that the control unit 24 sets the overlap rate to 40% in the previous process of step S10 and to 50% in the latest process of step S10. Furthermore, in the process of step S15, the control unit 24 specifies the overlap rate as the type of parameter set for the object for which the learning model estimation result was incorrect. In this case, in the process of step S16, the control unit 24 specifies the overlap rate range of 40% to 50% as the parameter range. Furthermore, the control unit 26 generates the first learning image based on at least a portion of the overlap rate included in the range of 40% to 50%.
[0084] In this way, in the information processing device 20, the control unit 24 identifies image features that are likely to cause the learning model's estimation result to be incorrect, based on the evaluation results. This configuration makes it possible to identify weaknesses in the learning model's object recognition. By identifying the weaknesses in the learning model's object recognition, it is possible to efficiently generate learning data to eliminate the weaknesses in the learning model's object recognition. Therefore, the learning model can be trained efficiently.
[0085] Here, we consider a case where a weakness in the object recognition of a learning model cannot be identified. In this case, it is necessary to eliminate the weakness in the object recognition of the learning model by training the learning model with training data including a large amount of training image data. However, using a large number of training images may increase the workload of annotation and other tasks.
[0086] In contrast, in the information processing device 20 according to the present embodiment, as described above, the control unit 24 identifies image features that are likely to cause the estimation result of the learning model to be incorrect, thereby identifying weaknesses in the learning model's object recognition. This configuration makes it possible to eliminate weaknesses in the learning model's image recognition without training the learning model with training data that includes a large amount of training image data. Therefore, in this embodiment, the possibility of an increase in the workload of annotations, etc. is reduced.
[0087] Therefore, according to this embodiment, an improved technique for supporting learning can be provided.
[0088] Furthermore, the control unit 24 may generate first training data based on the feature information. With this configuration, it is possible to automatically generate first training data to eliminate weaknesses in the learning model in object recognition. Furthermore, the control unit 24 may train the learning model using the first training data. With this configuration, it is possible to automatically eliminate weaknesses in the learning model in object recognition.
[0089] Furthermore, the control unit 24 may set parameters and generate a first evaluation image based on the set parameters. The control unit 24 may acquire, as feature information, at least one of parameters set for an object for which the learning model's estimation result was incorrect and parameters set for the environment of the object. By setting parameters and generating a first evaluation image based on the set parameters, it becomes possible to appropriately adjust the difficulty of questions using the first evaluation image. By appropriately adjusting the difficulty of questions using the first evaluation image, it is possible to more accurately identify weaknesses in the learning model's object recognition.
[0090] Furthermore, the control unit 24 may execute a preset learning process and an identification process in parallel. In this case, the control unit 24 may generate a first evaluation image based on parameters set in the first process of the learning process. That is, in the information processing method according to this embodiment, generating a first evaluation image may include generating a first evaluation image based on parameters set in the first process. For example, in the process of step S13, the control unit 24 generates a first evaluation image based on parameters set in step S10 as the first process. By generating a first evaluation image based on parameters set in the first process in this manner, a point in the third process where the learning model was not sufficiently trained using the second training data can be identified as a weak point in the object recognition of the learning model. For example, the control unit 24 may acquire a point in the third process where the learning model was not sufficiently trained using the second training data as a parameter in the process of step S15.
[0091] Furthermore, the control unit 24 may repeatedly execute the learning process. In the repeatedly executed learning process, the control unit 24 may set parameters in the latest first process so that the difficulty level is at least one level higher than the difficulty level corresponding to the parameters set in the previous first process. For example, the control unit 24 sets parameters in the latest process of step S10 so that the difficulty level is at least one level higher than the difficulty level corresponding to the parameters set in the previous process of step S10. With this configuration, the difficulty level of the questions using the second learning images increases stepwise as the learning process is repeatedly executed. For example, as the processes of steps S10 to S18 are repeatedly executed, the difficulty level corresponding to the parameters set in the process of step S10 increases stepwise, and the difficulty level of the questions using the second learning images generated in the process of step S11 also increases stepwise.
[0092] Furthermore, the control unit 24 may generate the first evaluation image based on parameters newly set in the most recent first process. That is, in the information processing method according to this embodiment, generating the first evaluation image may include generating the first evaluation image based on parameters newly set in the most recent first process. For example, in the process of step S13, the control unit 24 generates the first evaluation image based on parameters newly set in the process of step S10 as the most recent first process. With this configuration, in the repeatedly executed learning process, the first evaluation image is generated based on the same parameters as the second learning image. For example, the first evaluation image generated in the process of step S13 is generated based on parameters newly set in the most recent process of step S10, just like the second learning image generated in the process of step S11. Because the first evaluation image is generated based on the same parameters as the second learning image, the difficulty of the questions based on the first evaluation image remains the same as the difficulty of the questions based on the second learning image, even if the difficulty of the questions based on the second learning image gradually increases as the learning process is repeatedly executed. With this configuration, the points at which the learning model was not sufficiently trained using the second learning data in the process of step S12 can be acquired more accurately as parameters in the process of step S15.
[0093] Furthermore, the control unit 24 may train the learning model using the first learning data in the repeatedly executed learning process. That is, the information processing method according to this embodiment may include training the learning model using the first learning data in the repeatedly executed learning process. For example, when repeatedly executing the processes of steps S10 to S18, the control unit 24 trains the learning model in the process of step S12 using the first learning data generated in the process of step S16.
[0094] Furthermore, the control unit 24 may acquire, as feature information, at least one of the parameters set for the object for which the learning model estimation result was incorrect and the parameters set for the environment of the object, among the multiple newly set parameters. That is, in the information processing method according to this embodiment, acquiring parameters as feature information may include acquiring, among the multiple newly set parameters, parameters set for the object for which the learning model estimation result was incorrect. Also, in the information processing method according to this embodiment, acquiring parameters as feature information may include acquiring, among the multiple newly set parameters, parameters set for the environment of the object for which the learning model estimation result was incorrect. For example, in the processing of step S15, the control unit 24 acquires, among the multiple parameters set in the latest processing of step S10, parameters set for the object for which the learning model estimation result was incorrect.
[0095] Furthermore, in the repeatedly executed learning process, the control unit 24 may set parameters in the latest first process so that the difficulty level is several levels higher than the difficulty level corresponding to the parameters set in the previous first process. In this case, the control unit 24 may identify the type of at least one of the parameters set for the object for which the learning model's estimation result was incorrect and the parameters set for the environment of the object. The control unit 24 may identify a parameter range for the identified type of parameter, ranging from the set value in the previous first process to the set value in the latest first process. The control unit 24 may generate a first learning image based on at least a portion of the parameters included in the parameter range. For example, as described above, in the latest process of step S10, the control unit 24 may set parameters so that the difficulty level is several levels higher than the difficulty level corresponding to the parameters set in the previous process of step S10. Also, as described above, in the process of step S15, the control unit 24 may identify the type of parameters, etc., set for the object for which the learning model's estimation result was incorrect. As described above, in the process of step S16, the control unit 24 may identify a parameter range and generate the first training image based on at least a portion of the parameter range. By setting the parameters so that the level of difficulty increases by several levels, the time required for curriculum learning can be reduced. Furthermore, even if the parameters are set so that the level of difficulty increases by several levels, the first training image can be generated based on at least a portion of the parameters included in the identified parameter range. With this configuration, the first training data can compensate for any insufficient learning of the learning model during curriculum learning.
[0096] Furthermore, the control unit 24 may repeatedly execute the learning process until the estimation accuracy of the learning model satisfies a set condition. With this configuration, the learning process can be terminated at an appropriate timing.
[0097] Furthermore, the control unit 24 may generate the first evaluation image by using a cut-and-paste method on an existing image. That is, in the information processing method according to this embodiment, generating the first evaluation image may include using a cut-and-paste method on an existing image. By generating the first evaluation image from an existing image, it is possible to reduce the time and cost required to identify weaknesses in the object recognition of the learning model compared to when the first evaluation image is actually captured and generated. Furthermore, by generating the first evaluation image from an existing image, it is possible to easily generate more first evaluation images compared to when the first evaluation image is actually captured and generated. By using more first evaluation images, it is possible to more accurately identify weaknesses in the object recognition of the learning model.
[0098] Furthermore, the control unit 24 may generate the first training image by using a cut-and-paste method on an existing image. That is, in the information processing method according to this embodiment, generating the first training image may include using a cut-and-paste method on an existing image. By generating the first training image from an existing image, the time and cost required to generate the first training data can be reduced. By reducing the time and cost required to generate the first training data, the time and cost required to resolve weaknesses in object recognition of the training model can be reduced.
[0099] Furthermore, the control unit 24 may generate the second training images by using a cut-and-paste method for existing images. That is, in the information processing method according to this embodiment, generating the second training images may include using a cut-and-paste method for existing images. By generating the second training images from existing images, the time and cost required for curriculum learning using the second training data can be reduced compared to when the second training images are actually captured and generated. Furthermore, by generating the second training images from existing images, it is possible to easily generate a larger number of second training images compared to when the second training images are actually captured and generated. Furthermore, even when curriculum learning is performed using second training data including data of second training images generated from existing images, in this embodiment, weaknesses in object recognition of the learning model can be identified.
[0100] Here, the POS system is required to train the learning model to learn about a new product every time a new product is released. Therefore, even after the information processing system 3 is put into operation, the learning model is required to learn about the new product whenever a new product is released. However, depending on the configuration of the information processing system, it may be difficult to train the learning model to learn about a new product after the information processing system is put into operation. Furthermore, in order to train the learning model to learn about a new product, it may be necessary to prepare many training images and evaluation images including the new product.
[0101] In this embodiment, for example, if one image generated by capturing a new product is prepared, the control unit 24 can easily generate a first evaluation image, a first learning image, and a second learning image by using a cut-and-paste method on existing images including the image of the new product. In other words, many first evaluation images, first learning images, and second learning images including the new product can be easily prepared. Furthermore, if the first evaluation images, first learning images, and second learning images including the image of the new product are prepared, the control unit 24 can execute the learning support process at any timing. Therefore, even after the information processing system 3 is put into operation, the learning model can learn about new products.
[0102] Furthermore, the manager may wish to have the learning model learn images that correspond to the usage patterns of the store in which the information processing system 3 is installed. In this case, in this embodiment, the manager can cause the information processing device 20 to generate images that correspond to the usage patterns of the store. Furthermore, in the information processing device 20, the control unit 24 can cause the learning model to learn images that correspond to the usage patterns of the store by the above-described processing.
[0103] For example, suppose that a store frequently experiences cases in which a customer's hand, an object not targeted for detection, is captured in an image. The manager wants the learning model to learn images containing the hand as an image appropriate for the store's usage pattern. In this case, the manager can cause the information processing device 20 to generate a first evaluation image 31 as shown in FIG. 8 . The first evaluation image 31 includes an object image 30a corresponding to a rice ball, an object image 30c corresponding to chocolate, an object image 31a corresponding to a hand, and a background image 30d. To generate the first evaluation image 31, the manager causes the imaging unit 12 to capture and generate an image including the object image 31a. The control unit 24 generates the first evaluation image 31 by using a cut-and-paste method on an image including the object image 31a captured and generated by the imaging unit 12. In the same manner as or similar to generating the first evaluation image 31, the control unit 24 can generate first and second learning images including the object image 31a. The control unit 24 can then perform curriculum learning and identification processing using these images. With this configuration, the learning model can learn images that correspond to the usage pattern of the store in which the information processing system 3 is installed.
[0104] Although the embodiments according to the present disclosure have been described based on the drawings and examples, it should be noted that those skilled in the art would easily be able to make various modifications or alterations based on the present disclosure. Therefore, it should be noted that these modifications and alterations are included within the scope of the present disclosure. For example, the functions included in each component or step can be rearranged so as not to be logically inconsistent, and multiple components or steps can be combined or divided into one.
[0105] For example, in the present embodiment, the control unit 24 has been described as training the learning model in the process of step S12 using the first learning data generated in the process of step S16 and the second learning data generated in the process of step S11. However, the process of training the learning model using the first learning data and the process of training the learning model using the second learning data may be executed as separate processes. For example, immediately after executing the process of step S16, the control unit 24 may train the learning model using the first learning data generated in the process of step S16.
[0106] For example, the information processing method according to the present embodiment has been described as being executed by the information processing device 20. However, the device that executes the information processing method according to the present embodiment is not limited to the information processing device 20. The information processing method according to the present embodiment may be executed by any device. For example, the information processing method according to the present embodiment may be executed by an image generation device, a learning support device, a server 4, or the like. Furthermore, the information processing method according to the present embodiment may be executed as an image generation method or as a learning support method.
[0107] For example, an embodiment is also possible in which a general-purpose computer functions as the information processing device 20 according to this embodiment. Specifically, a program describing the processing content for realizing each function of the information processing device 20 according to this embodiment is stored in the memory of the general-purpose computer, and the program is read and executed by a processor. Therefore, the configuration according to this embodiment can also be realized as a program executable by a processor or a non-transitory computer-readable medium storing the program.
[0108] In the present disclosure, descriptions such as "first" and "second" are identifiers for distinguishing the configuration. In the present disclosure, the configurations distinguished by descriptions such as "first" and "second" can have their numbers exchanged. For example, the first evaluation image can exchange the identifiers "first" and "second" with the second evaluation image. The exchange of identifiers is performed simultaneously. The configurations remain distinguished even after the identifier exchange. The identifiers may be deleted. A configuration from which the identifier has been deleted is distinguished by a symbol. The descriptions of identifiers such as "first" and "second" in the present disclosure should not be used solely to interpret the order of the configurations or to justify the existence of an identifier with a smaller number.
[0109] REFERENCE SIGNS LIST 1 Payment system 2 Network 3 Information processing system 4 Server 10 Placement platform 10s Top surface 11 Support column 12 Imaging unit 13 Display device 20 Information processing device 21 Communication unit 22 Input unit 23 Storage unit 24 Control unit 30 First evaluation image 30a, 30b, 30c, 31a Object image 30d Background image 30a1, 30b1 Detection frame
Claims
1. acquiring an evaluation result indicating whether an estimation result of a learning model for the first evaluation image is correct or incorrect based on first evaluation data including data of at least one first evaluation image and correct answer data for the first evaluation image; and performing a specification process for specifying image features that are likely to result in an incorrect estimation result of the learning model based on the evaluation result.
2. 2. The information processing method according to claim 1, further comprising: generating first learning data including data of at least one first learning image and correct answer data for the first learning image, the first learning image being generated based on feature information indicating features of the image.
3. The information processing method according to claim 2 , further comprising: training the learning model with the first training data.
4. The method further includes setting a parameter corresponding to a degree of difficulty in estimating an object in an image, and generating the first evaluation image based on the set parameter; The information processing method according to claim 3 , wherein the identification process includes a process of acquiring, as the feature information, at least one of the parameters set for an object for which the estimation result of the learning model is incorrect and the parameters set for the environment of the object for which the estimation result is incorrect.
5. The method further includes executing a preset learning process and the specifying process in parallel, the learning process includes a first process of setting the parameters, a second process of generating at least one second learning image based on the set parameters, and a third process of learning the learning model using second learning data including data of the at least one second learning image and answer data for the second learning image; The information processing method according to claim 4 , wherein generating the first evaluation image includes generating the first evaluation image based on the parameters set in the first process.
6. repeatedly executing the learning process; In the repeatedly executed learning process, in the latest first process, the parameter is set so that the difficulty level is at least one level higher than the difficulty level corresponding to the parameter set in the previous first process, The information processing method according to claim 5 , wherein generating the first evaluation image includes generating the first evaluation image based on the parameters newly set in the latest first process.
7. In the repeatedly executed learning process, in the latest first process, the parameters are set so that the difficulty level is higher by several levels than the difficulty level corresponding to the parameters set in the previous first process; Identifying at least one of the types of the parameters set for an object for which the estimation result of the learning model is incorrect and the types of the parameters set for an environment of the object for which the estimation result of the learning model is incorrect; acquiring a parameter range of the parameter whose type has been identified, the parameter range being a range from a setting value in the previous first process to a setting value in the latest first process; The information processing method according to claim 6 , further comprising: generating the first training image based at least in part on the parameter range.
8. The information processing method according to claim 6 , further comprising repeatedly executing the learning process until an estimation accuracy of the learning model satisfies a set condition.
9. 9. The information processing method according to claim 5, wherein generating the first evaluation image, the first learning image, and the second learning image includes using a cut-and-paste method for existing images.
10. On the computer, acquiring an evaluation result indicating whether an estimation result of a learning model for the first evaluation image is correct or incorrect based on first evaluation data including data of at least one first evaluation image and correct answer data for the first evaluation image; A program for executing a process of identifying image features that are likely to result in an incorrect estimation result of the learning model based on the evaluation results.
11. a control unit that acquires an evaluation result indicating whether an estimation result of a learning model for the first evaluation image is correct or incorrect based on first evaluation data including data of at least one first evaluation image and correct answer data for the first evaluation image; The control unit of the information processing device executes a process of identifying image features that are likely to result in an incorrect estimation result of the learning model based on the evaluation result.