Information processing apparatus, information processing method, and program

By generating and evaluating circuits, and utilizing deep learning and generative adversarial networks to generate virtual evaluation image data, the problem of assessing the risk of misidentification in commodity recognition systems is solved, and the recognition accuracy of the system under different lighting environments is improved.

CN117242483BActive Publication Date: 2026-08-04PANASONIC INTELLECTUAL PROPERTY MANAGEMENT CO LTD
View PDF 3 Cites 0 Cited by

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
PANASONIC INTELLECTUAL PROPERTY MANAGEMENT CO LTD
Filing Date
2022-04-08
Publication Date
2026-08-04

AI Technical Summary

Technical Problem

In existing technologies, product recognition systems are prone to misidentification under different lighting conditions, making it difficult to assess the risk of misidentification. In particular, the accuracy of recognizing other items in the environment of misidentified items is difficult to predict.

Method used

By using generation and evaluation circuits, virtual image data of a second item is generated based on the image data of a first item. The image recognition accuracy of the second item is evaluated, and the evaluation result is output. Virtual evaluation image data is generated using deep learning and generative adversarial networks. Attributes such as lighting environment and item orientation are adjusted to assess the risk of misidentification.

Benefits of technology

It enables accurate assessment of misidentified items under different lighting conditions, reduces the risk of misidentification, and improves the accuracy and reliability of the product recognition system.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN117242483B_ABST
    Figure CN117242483B_ABST
Patent Text Reader

Abstract

An information processing apparatus includes: a generation circuit that generates second image data that reproduces an image in which a second article different from a first article is arranged in a first environment, based on attribute information of first image data obtained by photographing the first article in the first environment; an evaluation circuit that evaluates an image recognition accuracy for the second article based on the second image data; and an output circuit that outputs a result of the evaluation of the accuracy.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This disclosure relates to information processing apparatus, information processing methods, and procedures. Background Technology

[0002] Methods for identifying goods through image recognition in retail stores such as supermarkets or convenience stores (e.g., product identification methods) have been studied (e.g., see Patent Document 1).

[0003] Existing technical documents

[0004] Patent documents

[0005] Patent Document 1: Japanese Patent Application Publication No. 2020-061164 Summary of the Invention

[0006] However, there is still room for improvement in methods to reduce the false recognition rate of image recognition in product identification. In particular, methods to prevent the misidentification of other items in the recognition environment of a misidentified item have not been studied in the past.

[0007] The non-limiting embodiments of this disclosure help to provide information processing apparatus, information processing method and program that can evaluate the image recognition accuracy for different items in a specific environment when different items become the objects of image recognition.

[0008] An information processing apparatus according to an embodiment of this disclosure includes: a generation circuit that generates second image data based on attribute information of first image data obtained by photographing a first item in a first environment, the second image data reproducing an image of a second item, which is different from the first item, disposed in the first environment; an evaluation circuit that evaluates the image recognition accuracy for the second item based on the second image data; and an output circuit that outputs the evaluation result of the accuracy.

[0009] It should be noted that these general or specific methods can be implemented by systems, devices, methods, integrated circuits, computer programs, or recording media, or by any combination of systems, devices, methods, integrated circuits, computer programs, and recording media.

[0010] According to one embodiment of this disclosure, when different items in a specific environment become the objects of image recognition, it is possible to evaluate the image recognition accuracy for different items.

[0011] Further advantages and effects of one embodiment of this disclosure will be illustrated by the specification and drawings. These advantages and / or effects are provided by the various embodiments and the features described in the specification and drawings, but not necessarily all of them need to be provided in order to obtain one or more of the same features. Attached Figure Description

[0012] Figure 1 This is a diagram illustrating an example of the structure of a product identification system.

[0013] Figure 2 This is a flowchart illustrating an example of the actions of a product identification system.

[0014] Figure 3 This is a diagram illustrating an example of the risk of misidentification.

[0015] Figure 4 This is a diagram illustrating an example of the risk of misidentification.

[0016] Figure 5 This is a diagram illustrating an example of the risk of misidentification.

[0017] Figure 6 This is a diagram illustrating an example of the risk of misidentification.

[0018] Figure 7 This is a diagram illustrating a derived example of a lighting suggestion.

[0019] Figure 8 This is a diagram illustrating an example of the risk of misidentification.

[0020] Figure 9 This is a diagram illustrating an example of the risk of misidentification.

[0021] Figure 10 This is a diagram illustrating an example of a computer's hardware structure. Detailed Implementation

[0022] Hereinafter, various embodiments of the present disclosure will be described in detail with reference to the accompanying drawings.

[0023] For example, a system for identifying goods using image recognition at the checkout counter (or cashier) in a retail store has been studied (e.g., referred to as a "goods recognition system"). Furthermore, "goods" is an example of "articles." For instance, the object of image recognition may not be limited to goods.

[0024] In image recognition, for example, lighting environments can vary from store to store. For instance, if the lighting environment and the environment in which the product model is learned (e.g., referred to as the "learning environment") differ, misidentification of products (e.g., including cases where products are not identified) may occur.

[0025] When misidentification of a product is caused by a difference between the learning environment and the environment in which the product is being identified (e.g., referred to as the "recognition environment"), other products may also be misidentified in the same recognition environment. More specifically, when a shadow falling on a product in the recognition environment is the cause of misidentification, the likelihood of the same shadow being cast on other products is high. Therefore, when a product is misidentified, it is sometimes desirable to assess (or estimate) the risk (e.g., the "misidentification risk") that would lead to misidentification of other products in the recognition environment where the misidentified product occurred. In other words, it is sometimes desirable to assess whether other products would also be misidentified in an environment where a misidentified product occurred.

[0026] Here, it is sometimes difficult to faithfully reproduce the past identification environment (e.g., lighting environment) in assessing the risk of misidentification of goods that are different from the goods that were misidentified. In addition, it is sometimes difficult to assess the risk of misidentification of other goods at the same time that were misidentified.

[0027] In one embodiment of this disclosure, a method is described, for example, for determining (or assessing) the risk of misidentification of a product that is different from the product that was misidentified in an identification environment where misidentification has occurred.

[0028] (Implementation Method 1)

[0029] [Structure of a Product Recognition System]

[0030] Figure 1 This is a diagram illustrating a structural example of the product identification system 1 according to this embodiment.

[0031] Figure 1 The product recognition system 1 shown may include, for example, an image acquisition unit 11, a product recognition unit 12, a storage unit 13, a misidentification analysis unit 14, and a result display unit 15. The misidentification analysis unit 14 may also be replaced by other names such as "misidentification determination device," "image recognition device," "image processing device," or "information processing device."

[0032] The image acquisition unit 11 acquires, for example, image data (e.g., product image data) of products at a location where image recognition is performed by the product recognition unit 12 (e.g., checkout counter). For example, the image acquisition unit 11 may acquire product image data obtained by photographing the products present at the checkout location using a camera (not shown). For example, the acquired product image data is sent from the image acquisition unit 11 to the product recognition unit 12.

[0033] The product recognition unit 12 can, for example, perform image recognition on products purchased by customers at a store. The product recognition unit 12 can, for example, identify products corresponding to product image data input from the image acquisition unit 11 based on a learned recognition model (hereinafter referred to as a "product recognition model") stored in the storage unit 13. For example, the product recognition unit 12 can output information indicating the product recognition result to the result display unit 15.

[0034] Furthermore, the product recognition unit 12 can, for example, determine whether a product has been misidentified. For instance, the product recognition unit 12 can detect misidentification based on the user's (e.g., a salesperson or customer's) judgment at checkout. More specifically, the product recognition unit 12 can detect misidentification if the user has entered information indicating that misidentification has occurred. Additionally, the product recognition unit 12 can detect misidentification at time points other than checkout. For example, the product recognition unit 12 can detect misidentification if, when summarizing store sales data, the actual inventory and theoretical sales figures do not match. Furthermore, in the event of misidentification, the product recognition unit 12 can, for example, output the product image data of the misidentified product to the misidentification analysis unit 14.

[0035] Storage unit 13 stores, for example, a product recognition model. The product recognition model may include, for example, product image data and product-related information such as the product name.

[0036] The misidentification analysis unit 14, for example, evaluates the accuracy of image recognition for a product. As an example of an accuracy-related indicator, the misidentification analysis unit 14 can determine (or judge or evaluate) an indicator representing the probability of misidentification (hereinafter referred to as "misidentification risk").

[0037] For example, the misidentification analysis unit 14 can generate image data containing other products in the identification environment (e.g., the first environment) based on the attribute information (e.g., also referred to as "attributes") of the image data obtained by the product identification unit 12 from the image data of the product captured in the identification environment (e.g., the first environment), and evaluate the image recognition accuracy for other products based on the generated image data.

[0038] The attributes of image data are, for example, information related to features within the image, and there is a relationship between the attributes and the features such that changing the attribute will change the image's features. For example, as a non-limiting example of the attributes set in the product recognition system 1, attributes such as the direction of illumination, the intensity of illumination, the color temperature of illumination, the presence or absence of blur, the orientation of the object to be recognized (e.g., a product), the intensity of reflection, and individual differences of the object to be recognized can be listed. Furthermore, the types of attributes are not limited to these attributes; for example, other features related to product recognition may also be included.

[0039] For example, the misidentification analysis unit 14 can acquire product image data (hereinafter also referred to as "misidentified image data") of a product that has been determined to be misidentified in the image recognition of the product recognition unit 12. Furthermore, the misidentification analysis unit 14 can, for example, modify the attribute information of product image data (for example, also referred to as "evaluation image data") of a product different from the misidentified product, based on the attribute information of the misidentified image data. Then, the misidentification analysis unit 14 can, for example, evaluate the image recognition accuracy for that product based on the image data with modified attribute information (hereinafter, "virtual evaluation image data").

[0040] The misidentification analysis unit 14 can, for example, output information related to the evaluation results of image recognition accuracy to the result display unit 15.

[0041] The result display unit 15 (or output unit) may, for example, display (or output) the identification result of the product input from the product identification unit 12. Additionally, the result display unit 15 may, for example, display (or output) information related to the misidentification risk determined by the misidentification analysis unit 14. An example of a method for displaying the misidentification risk in the result display unit 15 will be described below.

[0042] Furthermore, at least one of the image acquisition unit 11, the product recognition unit 12, the storage unit 13, and the result display unit 15 may be included in the misidentification analysis unit 14, or may not be included in the misidentification analysis unit 14.

[0043] [Example of the structure of the misidentification analysis unit 14]

[0044] The misidentification analysis unit 14 may include, for example, a performance evaluation image database (DB) 141, an image encoding unit 142, an attribute manipulation unit 143, an image generation unit 144, and an evaluation unit 145. For example, the attribute manipulation unit 143 and the image generation unit 144 may correspond to a generation circuit.

[0045] Image DB 141 for performance evaluation, for example, stores image data of goods that a store may handle (e.g., evaluation image data).

[0046] Image encoding unit 142 encodes, for example, misidentified image data input from product recognition unit 12. Additionally, image encoding unit 142 encodes, for example, at least one evaluation image data stored in performance evaluation image database 141. Image encoding unit 142 may encode both misidentified image data and evaluation image data based on a product image encoding model. Image encoding unit 142 outputs, for example, encoding information (e.g., encoding related to at least one attribute) of the misidentified image and evaluation image to attribute operation unit 143.

[0047] As an example of a product image encoding model (or encoding method), one could cite methods such as deep learning encoders that obtain real vectors of a certain dimension through image convolution, or methods that use an image generation model to generate an image from the encoding (e.g., real vector) and search for an encoding to obtain an image that is closer to the misidentified image.

[0048] The attribute operation unit 143 can, for example, operate on attribute information in the encoded information of the evaluation image data input from the image encoding unit 142. Additionally, the attribute operation unit 143 can, for example, operate on attributes in the encoded information of misidentified image data input from the image encoding unit 142. The attribute operation unit 143 can, for example, perform attribute operations based on a generated image attribute operation model. Furthermore, the term "operation" can be replaced with other terms such as "adjustment" or "change."

[0049] Additionally, the attribute operation unit 143 can, for example, apply (or reflect) the attribute information related to the lighting environment (e.g., lighting conditions) from the attribute information contained in the encoded information of the misidentified image data to the evaluation image data. For example, the attribute operation unit 143 can mix the lighting environment-related attribute information of the misidentified image data and the lighting environment-related attribute information of the evaluation image data, or it can replace the lighting environment-related attribute information of the evaluation image data with the lighting environment-related attribute information of the misidentified image data.

[0050] Additionally, the attribute operation unit 143 may, for example, adjust at least one of the attribute information related to the lighting environment (e.g., a first attribute) and the attribute related to a feature quantity different from the lighting environment (e.g., a second attribute related to the object being identified) for evaluating image data. Furthermore, the second attribute may also be an attribute related to the surrounding environment that is neither different from the lighting nor the object.

[0051] The attribute manipulation unit 143 outputs, for example, the encoded information of the evaluation image after the attribute manipulation to the image generation unit 144.

[0052] Image generation unit 144 generates (in other words, decodes) image data (e.g., virtual evaluation image data) based on the encoding information of the evaluation image input from attribute operation unit 143. Image generation unit 144 may generate images based on a product image generation model. Image generation unit 144 may also use a generator from deep learning to generate virtual evaluation image data. Image generation unit 144 outputs the generated virtual evaluation image data to evaluation unit 145.

[0053] The evaluation unit 145 can evaluate the accuracy of product recognition for virtual evaluation image data input from the image generation unit 144 based on the product recognition model stored in the storage unit 13. For example, the evaluation unit 145 can also determine the risk of misidentification of the product corresponding to the virtual evaluation image data based on the product recognition result. The evaluation unit 145 can, for example, output information related to the evaluation result for the virtual evaluation image data to the result display unit 15.

[0054] [Example of actions in Product Recognition System 1]

[0055] Next, an example of the operation of the above-mentioned commodity identification system 1 will be described.

[0056] Figure 2 This is a flowchart illustrating an example of the operation of the misidentification analysis unit 14 in the product recognition system 1. For example, the misidentification analysis unit 14 can perform operations at a time when the product recognition unit 12 detects a misidentification of a product. Figure 2 The processing shown.

[0057] exist Figure 2 In this process, the misidentification analysis unit 14 encodes, for example, the misidentified image data input from the product identification unit 12 (S101).

[0058] The misidentification analysis unit 14 can, for example, perform the following processing related to recognition accuracy evaluation in steps S102 to S107 for at least one evaluation image (in other words, an image of at least one product). Additionally, it can perform the following processing related to recognition accuracy evaluation in steps S102 to S107 for at least one attribute operation (e.g., the conditions of the attribute operation). For example, the product to be evaluated for recognition accuracy can be predetermined, or the product can be selected by accepting user input. Furthermore, the conditions of the attribute operation applied in the recognition accuracy evaluation can be predetermined, or the conditions of the attribute operation applied in the recognition accuracy evaluation can be selected by accepting user input.

[0059] For example, the misidentification analysis unit 14 encodes the evaluation image data (S102).

[0060] The misidentification analysis unit 14 may, for example, perform at least one attribute operation (e.g., convert or adjust attributes) on the encoded information of the misidentified image data (S103).

[0061] The misidentification analysis unit 14 may, for example, perform at least one attribute operation on the encoded information of the evaluation image data (S104).

[0062] For example, the misidentification analysis unit 14 can perform predetermined attribute operations on at least one of the misidentified images and the evaluation images. These attribute operations may include, for example, operations on attributes such as the direction of illumination, the intensity of illumination, the color temperature of illumination, the intensity of reflection, the orientation of the item being evaluated (e.g., a commodity), or individual differences of the item.

[0063] For example, during the processing in S103, the misidentification analysis unit 14 can also adjust the attributes related to the lighting environment (e.g., the direction, intensity, or color temperature of the lighting) for the encoded information of the misidentified image data. Through this adjustment, for example, it is possible to virtually reproduce an environment that is darker or brighter than the actual environment related to the lighting.

[0064] Additionally, for example, the misidentification analysis unit 14 can adjust the encoded information of the evaluation image data in the processing of S104 to adjust the attributes related to the item being evaluated (e.g., the orientation of the item or individual differences). Through this adjustment, it is possible to reproduce the various states of the goods.

[0065] Furthermore, the misidentification analysis unit 14 may, for example, not perform attribute operations in one or both of S103 and S104.

[0066] The misidentification analysis unit 14, for example, reflects information related to the misidentified image data into information related to the evaluation image data (S105). For instance, the misidentification analysis unit 14 can reflect attribute information related to the lighting environment of the image recognition (or, referred to as "lighting environment information") from the attribute information contained in the encoded information of the misidentified image data into the encoded information (e.g., lighting environment information) of the evaluation image data. Furthermore, the term "reflect" can be replaced with other terms such as "update," "rewrite," "set," "apply," or "mix." For example, the misidentification analysis unit 14 can also perform style mixing or similar processing between the lighting environment information (or, encoding) of the misidentified image data and the lighting environment information (or, encoding) of the evaluation image data.

[0067] The misidentification analysis unit 14 generates virtual evaluation image data, for example, based on the encoding information of the evaluation image data, which reflects the lighting environment information of the misidentified image data (S106). In other words, the evaluation image data is converted into virtual evaluation image data based on the misidentified image data.

[0068] The misidentification analysis unit 14 evaluates the recognition accuracy (e.g., recognition success rate or misidentification risk) for the goods corresponding to the evaluation image data, for example, based on virtual evaluation image data (S107).

[0069] The misidentification analysis unit 14 outputs an evaluation result of the recognition accuracy to the result display unit 15, for example (S108).

[0070] [Methods for Generating Virtual Evaluation Image Data]

[0071] The following is an example illustrating a method for generating virtual evaluation image data.

[0072] <Generation Method 1>

[0073] In generation method 1, virtual evaluation image data can be generated, for example, by performing prescribed image processing on the entire evaluation image data. For example, the misidentification analysis unit 14 can perform image processing such as gamma correction or adding (or removing) Gaussian noise on the evaluation image data to generate virtual evaluation image data.

[0074] For example, the misidentification analysis unit 14 can generate virtual evaluation image data that virtually reproduces the lighting environment corresponding to the misidentified image data by adjusting the gamma correction value or Gaussian noise value in the misidentified image data (e.g., brightness-related attribute information of the misidentified image data).

[0075] In generation method 1, because virtual evaluation image data can be generated through a simple process, the amount of processing (or processing load) required to generate virtual evaluation image data can be suppressed.

[0076] <Generation Method 2>

[0077] In generation method 2, for example, the misidentification analysis unit 14 can use a three-dimensional model to generate virtual evaluation image data. For example, the misidentification analysis unit 14 can generate a three-dimensional model of an item (e.g., a commodity) that is being evaluated, and generate virtual evaluation image data that controls the lighting environment in the three-dimensional model based on the recognition environment corresponding to the misidentification image data (e.g., attribute information related to the recognition environment that was determined to be misidentified).

[0078] For example, the misidentification analysis unit 14 can reproduce the lighting environment, such as the quantity, type, location, or intensity of lighting, and the surrounding environment, such as the location, color, and material of objects (e.g., walls or mirrors) around the place where product identification is performed, in a 3D model based on the recognition environment corresponding to the misidentified image data. Additionally, the misidentification analysis unit 14 can, for example, reproduce the product, which is the object of evaluation for recognition accuracy, in a 3D model based on the evaluation image data.

[0079] In generation method 2, for example, to reproduce the lighting environment, in addition to the lighting state, a 3D model of the surrounding objects at the location where product recognition is performed is generated. Furthermore, at least one of the light source positions and light amounts from multiple lighting sources existing in reality is reproduced. Additionally, to reproduce individual differences among products, even for the same product, a separate 3D model is generated for each recognition object. Therefore, in generation method 2, as long as this data can be collected accurately, high-quality virtual evaluation image data that accurately reflects the surrounding environment, light source, and recognition object can be reproduced.

[0080] <Generation Method 3>

[0081] In generation method 3, for example, the misidentification analysis unit 14 encodes the evaluation image data and generates virtual evaluation image data by reflecting the attributes obtained from encoding the misidentified image data. When reflecting these attributes, the misidentification analysis unit 14 can use a Generative Adversarial Network (GAN) to transform the evaluation image data (in other words, generate virtual evaluation image data). GAN is one of the image generation techniques based on machine learning models using neural networks, such as deep learning. In this embodiment, the misidentification analysis unit 14 uses a GAN, for example, to generate virtual evaluation image data. In other words, in this embodiment, GAN can be used to generate evaluation data, rather than to generate training data for machine learning.

[0082] The misidentification analysis unit 14 can, for example, adjust the set value (e.g., feature quantity) of at least one attribute of the evaluation image data based on the attributes of the misidentified image data, thereby generating virtual evaluation image data. For example, the misidentification analysis unit 14 can virtually change the lighting environment in the evaluation image data by adjusting attributes related to the lighting environment. Additionally, for example, the misidentification analysis unit 14 can virtually change the orientation of the object in the evaluation image data by adjusting attributes related to the orientation of the object. Furthermore, the misidentification analysis unit 14 may not only reflect the attributes of the misidentified image data in the evaluation image data, but also apply arbitrary attributes such as orientation or lighting environment specified by the user to the evaluation image data.

[0083] In generation method 3, for example, because each of the multiple attributes can be changed independently, the lighting environment for evaluating image data can be adjusted more finely compared to generation method 1. Furthermore, in generation method 3, for example, because attributes such as the shooting angle (e.g., the orientation of the product), the direction of the lighting, or individual differences in the identified object can be changed, states that generation method 1 cannot reproduce can be reproduced.

[0084] Furthermore, in generation method 3, virtual evaluation image data can be generated, for example, by manipulating attributes related to the lighting environment (e.g., encoding a portion of the encoded information). Additionally, in generation method 3, virtual evaluation image data can be generated through attribute manipulation without training data. Therefore, in generation method 3, virtual evaluation image data can be generated even without preparing large amounts of data. Thus, compared to generation method 2, generation method 3 can more easily evaluate the image recognition accuracy for items different from those judged as misidentified, thereby reducing the processing amount (or processing load) of generating virtual evaluation image data.

[0085] Furthermore, the machine learning model used to generate virtual evaluation image data in method 3 is not limited to GAN, but can also be other models.

[0086] [Example of displaying risk of misidentification]

[0087] The misidentification analysis unit 14 may output a signal to the result display unit 15 to display the assessment result of the misidentification risk (or an indicator related to recognition accuracy) of the goods corresponding to multiple assessment image data. For example, the signal in the misidentification analysis unit 14 for displaying the assessment result may be output from the assessment unit 145 to the result display unit 15, or it may be output from an output unit not shown to the result display unit 15.

[0088] The results display unit 15 can display the assessment results of the misidentification risk (or indicators related to identification accuracy) of the product based on the signals from the misidentification analysis unit 14. The display method is not limited. As a non-limiting example, a chart display or a list display (overview display) can be used.

[0089] The following is a display example illustrating the risk of misidentification.

[0090] <Example 1>

[0091] Figure 3 , Figure 4 , Figure 5 and Figure 6 This is a diagram illustrating an example of a display screen showing the risk of misidentification in Example 1.

[0092] exist Figure 3 , Figure 4 , Figure 5 and Figure 6 The screen shown may include, for example, an area for selecting conditions related to the recognition performance evaluation (e.g., "Recognition Performance Evaluation Condition Selection" area), an area for displaying the results of the recognition performance evaluation (e.g., "Recognition Performance Evaluation Results" area), and an area for displaying recommended results for the lighting environment (e.g., "Lighting Recommendation Results" area).

[0093] The "Recognition Performance Evaluation Condition Selection" area may include, for example, a region that prompts the user to select a misidentified image whose lighting conditions should be reproduced. Furthermore, a signal (which may be called an "operation signal") corresponding to the user's operation on the screen (e.g., a selection operation such as touch or click) is input to the result display unit 15 (or, alternatively, the misidentification analysis unit 14), thereby executing processing corresponding to the operation signal (e.g., display control).

[0094] Additionally, the "Recognition Performance Evaluation Condition Selection" area may include, for example, an area for selecting the evaluation image from which the recognition performance evaluation results are to be displayed. The evaluation image displaying the recognition performance evaluation results may include, for example, the evaluation image stored in the performance evaluation image DB 141 (e.g., referred to as the "original performance evaluation image"), and an evaluation image obtained by applying conditions of a lighting environment that causes misrecognition (e.g., referred to as the "misrecognition condition") to the evaluation image stored in the performance evaluation image DB 141 (e.g., referred to as the "misrecognition condition application image").

[0095] Additionally, the images used to apply misidentification conditions may also include evaluation image data for which different attribute operations have been applied than those applied to the misidentification conditions. Figures 3-6 In the example shown, as an attribute operation, the following attributes can be selected: the orientation of the object, individual differences of the object, the direction of the lighting, the intensity of the lighting, the color temperature of the lighting, and the intensity of the reflection. Furthermore, the selectable attributes are not limited to these; other attributes may also be selected. For example, when a certain attribute operation is selected, the misidentification analysis unit 14 may apply a predetermined setting value (or, also called "attribute operation intensity") for that attribute operation, or it may determine the setting value by accepting user input (e.g., input from another screen).

[0096] Additionally, the “Identification Performance Evaluation Condition Selection” area may include, for example, a section that prompts the user to select (or decide) whether to implement lighting environment recommendations.

[0097] For example, users can select the misidentified image data that reproduces the misidentification condition on the screen (in...). Figures 3-6 The user selects the image (filename: false_images) and the evaluation image (object for recognition performance evaluation) on the screen and presses the "Execute" button. Alternatively, the user can select to execute a lighting environment suggestion on the screen and press the "Execute" button. The misidentification analysis unit 14 can, for example, start processing related to recognition performance evaluation (e.g., generation of virtual evaluation image data and recognition performance evaluation) after detecting that the user has pressed the execute button.

[0098] like Figures 3-6 As shown, the evaluation result for assessing the recognition performance of an image can be, for example, information representing the probability of successful recognition (e.g., "recognition rate" or "recognition accuracy"), that is, information representing the probability of successfully recognizing the corresponding product in the evaluation image data in product recognition. For example, the lower the recognition rate, the higher the probability of misidentification of the relevant product (in other words, the risk of misidentification or the failure rate of recognition).

[0099] For example, the recognition rate can be calculated based on whether the recognition result matches the correct product, which is obtained when the virtual evaluation image data is recognized using a product recognition model. For example, if the product recognition model determines that the product with the highest similarity to the virtual evaluation image data matches the correct product, the misidentification analysis unit 14 can set a higher recognition rate. Conversely, if the product recognition model determines that the product with the highest similarity to the virtual evaluation image data does not match the correct product, it may be identified as an incorrect product, therefore, the misidentification analysis unit 14 can set a lower recognition rate.

[0100] Furthermore, even if the product assessed as having the highest similarity matches the correct product, if the recognition score for other products is also high, a slight change in the environment could lead to misidentification as another product. Therefore, the misidentification analysis unit 14 can set the misidentification rate when the similarity is high for multiple products to be higher than the misidentification rate when the similarity is high only for the correct product. Additionally, when the similarity is low for all products (e.g., below a predetermined threshold), the product recognition model itself struggles to identify the product, so the misidentification analysis unit 14 can set the low recognition rate to be lower.

[0101] Furthermore, the evaluation results of recognition performance are not limited to the recognition rate; other parameters can also be used. Examples of other parameters include information representing the probability of failure in product recognition (e.g., "false recognition rate") or similarity as mentioned above.

[0102] The misidentification analysis unit 14 can, for example, determine whether the identification rate of each product (e.g., also referred to as a "category") is higher than a predetermined threshold. Figures 3-6(In the example, the recognition rate is 95%). For instance, information related to products with a recognition rate below the threshold (e.g., the number of non-compliant product categories) can also be displayed on the screen. Figures 3-6 In this example, the misidentification analysis unit 14 can output a signal to the results display unit 15 to graphically display the evaluation results (e.g., recognition rate) for each product. Through the graphical display, the user can, for example, comprehensively confirm the recognition performance for multiple products to which the misidentification criteria were applied.

[0103] Additionally, for example, such as Figure 4 and Figure 5 As shown, the misidentification analysis unit 14 can also output a signal to the result display unit 15 to emphasize the charts corresponding to items with a recognition rate below a threshold (e.g., highlighting by coloring or adding texture). This emphasis makes it easier for users to identify items with a high probability of misidentification.

[0104] Additionally, for example, such as Figures 3-6 As shown, you can also select a product's chart by hovering the mouse over it or clicking it in the chart display of recognition rates for each product (e.g., Figures 3-6 In the case of a chart for product A, display product-related information such as product name and recognition rate.

[0105] Furthermore, the display of the recognition rate for each product is not limited to a chart; it can also be a list or other display method. Additionally, the threshold is not limited to 95% and can be other values. For example, the threshold can be determined based on the recognition accuracy required for the use of the product recognition system 1. Furthermore, the threshold can be set in a variable manner.

[0106] Furthermore, the signal output by the misidentification analysis unit 14 to the result display unit 15 is not limited to the signal used for emphasis display. It may also output the following signal, which is used to distinguish between information related to products with recognition rates exceeding a predetermined threshold and information related to products with recognition rates not exceeding the predetermined threshold.

[0107] The following uses Figures 3-6Here are some examples of the display screen. In these examples, the user can input information for "Selecting a misidentified image to reproduce its lighting conditions," "Selecting a recognition performance evaluation image," and "Selecting lighting suggestions." Here, "Selecting a misidentified image to reproduce its lighting conditions" is an input section that accepts the selection of an image used to extract the lighting conditions that should be reflected in the performance evaluation images for each product. "Selecting a recognition performance evaluation image" is an input section that accepts the selection of whether to use the original image (original performance evaluation image) or an image with applied misidentification conditions (misidentification condition application image) when evaluating recognition performance. Furthermore, if "misidentification condition application image" is selected, the attributes to be applied or corrected can be further selected. "Selecting lighting suggestions" is an input section that accepts the selection of whether automatic suggestions are needed; these automatic suggestions are suggestions regarding lighting that is less prone to misidentification.

[0108] Figure 3 The example displayed is a case where the user selected the original performance evaluation image but did not select the misidentification condition application image, and did not select lighting suggestions. Furthermore, because... Figure 3 This is an example of not applying the attributes of misidentified images to the images used for performance evaluation, so in Figure 3 For example, one can choose not to select misidentified images whose lighting conditions are to be reproduced.

[0109] exist Figure 3 In such cases, the misidentification analysis unit 14 may, for example, evaluate the recognition rate of the original performance evaluation image without applying misidentification conditions based on the misidentified image to the original performance evaluation image.

[0110] The product recognition model, for example, uses a model that achieves a high recognition rate when evaluated using the original performance evaluation image. Therefore, the recognition rate for the original performance evaluation image is likely to exceed a predetermined threshold (e.g., 95%). Furthermore, thus, for example, as... Figure 3 As shown, the evaluation results of the recognition performance of the original performance evaluation images show that the recognition rate of all products is higher than the threshold (95%), and there are 0 unqualified product categories.

[0111] Figure 4 The example displayed shows a scenario where the user selected the misidentification condition application image instead of the original performance evaluation image, and also did not select lighting recommendations. Additionally, in Figure 4 For example, misidentified images (filename: false_images) whose lighting conditions were to be reproduced were selected. Furthermore, in... Figure 4 In the middle, because the idea was to directly apply the attributes of misidentified images, each "attribute operation" was not checked individually.

[0112] exist Figure 4 In such cases, the misidentification analysis unit 14 may, for example, apply misidentification conditions (e.g., lighting environment information) based on the selected misidentified image to the original performance evaluation image to generate virtual evaluation image data, and evaluate the performance (e.g., recognition rate) of product recognition for the virtual evaluation image data.

[0113] exist Figure 4 In the example shown, as an evaluation result of the recognition performance of the virtual evaluation image data, it shows that the recognition rate of three products, including product A, is below the threshold (95%), and there are three non-compliant categories. According to... Figure 4 The displayed screen allows the user to assess the recognition performance (e.g., risk of misidentification) across multiple items in the same lighting environment as that for the item corresponding to the misidentified image. Figure 4 In the example shown, the user can confirm that the recognition rate of three products, including product A, is below the threshold, increasing the risk of misidentification.

[0114] Figure 5 The example displayed shows a scenario where the user selected the misidentification condition application image instead of the original performance evaluation image, and also did not select lighting recommendations. Additionally, in Figure 5 In the example, the "Lighting Intensity" attribute is selected. Additionally, in... Figure 5 For example, in this context, misidentified images (filename: false_images) whose lighting conditions are to be reproduced are selected. Thus, in... Figure 5 The text displays the selected attribute from the properties of the reproduced misidentified image. Figure 5 Risk of misidentification under conditions of (lighting intensity in the image).

[0115] exist Figure 5 In such cases, the misidentification analysis unit 14 may, for example, apply attribute operations of lighting intensity to the original performance evaluation image and generate virtual evaluation image data based on the misidentification conditions of the selected misidentified image (e.g., lighting environment information), and evaluate the performance (e.g., recognition rate) of product recognition for the virtual evaluation image data.

[0116] exist Figure 5 In the example shown, as an evaluation result of the recognition performance of the virtual evaluation image data, it shows that the recognition rate of four products, including product A, is below the threshold (95%), and there are four non-compliant categories. According to... Figure 5The displayed screen allows users to not only cross-check the recognition performance (e.g., risk of misidentification) for multiple items in the same lighting environment (misidentification condition), but also to cross-check conditions different from the misidentification condition (in... Figure 5 In this context, the recognition performance (e.g., risk of misidentification) of multiple items under varying lighting intensity is assessed. The aforementioned lighting environment (misidentification condition) refers to the lighting environment for the item corresponding to the misidentified image. Figure 5 In the example shown, the user can identify the risk of misidentification (e.g., potential misidentification risk) when the lighting intensity is further changed for the misidentification conditions.

[0117] In addition, although Figure 5 The example shown is of selecting one attribute operation, but multiple attribute operations can also be selected.

[0118] In addition, although Figure 3 , Figure 4 and Figure 5 The text describes the selection of one image from the original performance evaluation image and the image used to apply the misidentification condition, but it is not limited to this; both the original performance evaluation image and the image used to apply the misidentification condition can also be selected. In this case, for example, it can also display... Figure 3 The chart shown illustrates the case where the original performance evaluation image was selected, and Figure 4 or Figure 5 The chart shown illustrates the application of the misidentification condition to the image. For example, it can be... Figure 3 The results of the recognition performance evaluation shown are, and Figure 4 or Figure 5 The results of the recognition performance evaluation are displayed side by side on the screen, or they can be displayed separately by switching screens, or they can be displayed in a distinguishable superimposed manner.

[0119] As a result, users can recognize the differences in the risk of misidentification caused by the different conditions, namely, whether or not misidentification conditions are applied, and whether or not conditions are added by the application through attribute operations.

[0120] Figure 6 The image shown is an example of a user selecting an image for misidentification conditions instead of the original performance evaluation image, and choosing a lighting suggestion instead. Additionally, in Figure 6 For example, in this context, misidentified images (filename: false_images) whose lighting conditions are to be reproduced are selected. Figure 6 The diagram shows the risk of misidentification in the case of lighting corresponding to the lighting recommendations, after reproducing the lighting conditions based on the misidentified image.

[0121] exist Figure 6 In such cases, the misidentification analysis unit 14 may, for example, apply misidentification conditions (e.g., lighting environment information) based on the selected misidentified image to the original performance evaluation image, generate virtual evaluation image data, and evaluate the performance (e.g., recognition rate) of product recognition for the virtual evaluation image data.

[0122] For example, the evaluation result may show that there are products with a recognition rate below a threshold. The misidentification analysis unit 14, for example, can determine (e.g., search) candidate attribute-related settings that would make the recognition rate of multiple products (e.g., a specified number of products or all products) higher than the threshold by operating on lighting-related attributes (e.g., at least one of lighting orientation, intensity, and color temperature), and output a signal to the results display unit 15 to display lighting suggestion results, which include the determined candidate settings (e.g., suggested values). Figure 6 In the example shown, for the evaluation image in the suggested lighting environment, the recognition performance evaluation results show that the overall recognition rate is above the threshold (95%), and there are 0 non-compliant categories. Additionally, in Figure 6 In the example shown, the direction, intensity, and color temperature of the lighting are displayed as results related to the suggested lighting environment (e.g., attribute settings).

[0123] according to Figure 6 The displayed screen allows the user to confirm lighting adjustment methods (e.g., adding lighting settings, or adjusting the direction, intensity, or color temperature of the lighting) for a lighting environment that is the same as the lighting environment (misidentification condition) for the product corresponding to the misidentified image.

[0124] Furthermore, the conditions for lighting recommendations are not limited to conditions that make the recognition rate of multiple (e.g., all) items higher than a threshold. For example, the condition could also be to reduce the number of items whose recognition rate based on lighting-related attribute operations is below a threshold (e.g., to reach a minimum or below a threshold). Additionally, for example, in the case where there are multiple attribute operations that make the recognition rate based on lighting-related attribute operations higher than a threshold (e.g., a combination of attribute operations), the misidentification analysis unit 14 could also use the attribute operation that makes the average recognition rate of multiple items higher (e.g., maximum) for the lighting recommendation results.

[0125] Furthermore, for example, the misidentification analysis unit 14 can make lighting suggestions to increase the recognition rate of specific products based on user specifications, product price, sales volume, or advertising content. For instance, expensive products have a greater impact on misidentification compared to other products. Additionally, for best-selling products or products featured in advertisements, the increased frequency of their recognition means a higher likelihood of misidentification if the recognition rate is low. Therefore, even if lighting suggestions slightly decrease the recognition rate of other products, they can still improve the recognition rate of those products, thereby reducing the negative impact on the store's overall sales or operations.

[0126] In addition, regarding Figure 6 The lighting recommendations suggested by the misidentification analysis unit 14 (e.g., candidate setting values ​​for attributes) can be determined by the user (in other words, manually) or by the misidentification analysis unit 14 in advance (in other words, automatically).

[0127] Figure 7 (a) indicates a search example of manually searching the lighting environment. Figure 7 (b) is a diagram showing a search example of automatically searching the lighting environment.

[0128] Figure 7 (a) is an example of a display screen where the user can select candidate attributes related to the lighting environment (e.g., the direction of the light, the intensity of the light, and the color temperature of the light). For example, when in Figure 6 In the middle, if the user selects lighting suggestions, it can also transition to Figure 7 The display shown in (a) is as follows.

[0129] exist Figure 7 In (a), the user can specify parameters related to the orientation of the lighting for the product being evaluated (e.g., one of the 360° directions), the lighting intensity (e.g., one of 0 to 100, where 0 represents off lighting and 100 represents maximum lighting intensity), and the color temperature of the lighting (e.g., one of the minimum to maximum values), and press the "Execute" button. The misidentification analysis unit 14, for example, can also, after detecting that the execute button has been pressed, generate virtual evaluation image data based on the lighting conditions selected by the user's input, and output a signal to the result display unit 15 to display the evaluation results of the recognition performance of the virtual evaluation image data (e.g., the number of products with a recognition rate below a threshold (95%)). Thus, the user can, for example, specify more appropriate lighting conditions (e.g., conditions with a higher recognition rate) based on the display of the recognition rate corresponding to the specified lighting conditions.

[0130] For example, it can also be like Figure 6 As shown, based on Figure 7 The recognition rates of multiple products under the lighting conditions determined in (a) are displayed in a graph.

[0131] Figure 7 (b) is a diagram showing examples of combinations of candidate attributes related to the lighting environment, such as the direction of the light, the intensity of the light, and the color temperature of the light. Figure 7 In the example shown in (b), there are 8 candidate orientations for the lighting, 5 candidate intensities for the lighting, and 5 candidate color temperatures for the lighting.

[0132] The misidentification analysis unit 14 can, for example, target a combination of attributes related to the lighting environment ( Figure 7 From some or all of the 200 (=8×5×5) combinations in (b), virtual evaluation image data is generated, and the recognition performance (e.g., recognition rate) is evaluated on the virtual evaluation image data. The misidentification analysis unit 14 may, for example, determine the lighting recommendation result as the combination of fewer items (e.g., below the minimum or specified value) among the multiple combinations of attributes related to the lighting environment that results in a recognition rate below a threshold.

[0133] In this way, by displaying the lighting suggestion results, the misidentification analysis unit 14 can prompt the user to adjust the lighting environment to reduce the risk of misidentification. Thus, for example, the risk of misidentification can be reduced without performing additional learning (or relearning), which is more costly than evaluating product identification.

[0134] Furthermore, the number of lighting recommendations determined by the misidentification analysis unit 14 is not limited to one, but can be multiple. Additionally, the combination (or number of candidates) of attributes related to the lighting environment is not limited to 200, but can be other numbers. For example, it could also be... Figure 7 At least one of the attributes shown in (b) (e.g., the direction, intensity, and color temperature of the lighting) is a fixed value.

[0135] In addition, although Figure 6 The example provided illustrates how to perform operations without selecting an attribute (e.g., with...). Figure 4 (Same), but can also be with Figure 5 Similarly, select one or more attribute operations.

[0136] In addition, although Figure 6 The document illustrates an example of graphically displaying the recognition rate of each product based on lighting recommendations, but is not limited to this. For example, it could also display the recognition rate of each product based on lighting recommendations, as well as the misrecognition rate of each product when misrecognition conditions are applied (e.g., compared to...). Figure 4 or Figure 5(Same). As a result, users can visually confirm the reduced risk of misidentification achieved through lighting suggestions.

[0137] <Example 2>

[0138] In Example 2, for example, the misidentification analysis unit 14 may output a signal to the result display unit 15 to display the evaluation result (e.g., recognition rate, misidentification risk, or recognition score as described later) for new virtual evaluation image data. This new virtual evaluation image data is generated by further adjusting at least one of the attribute information about a feature quantity that is different from the lighting environment, which is the lighting environment for the product corresponding to the misidentified image, for virtual evaluation image data that has been applied with attribute information related to the lighting environment.

[0139] As an example, the misidentification analysis unit 14 can output a signal to the result display unit 15 for performing a summary display, which means that for multiple new virtual evaluation image data generated by making different changes (in other words, adjustments or attribute operations) to the attribute information of the evaluation image data, a summary display is made of the evaluation results (e.g., recognition rate, misidentification risk or recognition score described later) corresponding to each of the multiple virtual evaluation image data.

[0140] Figure 8 This is a diagram illustrating an example of a display screen showing the risk of misidentification in Example 2.

[0141] exist Figure 8 The displayed screen includes areas such as "misidentified image", "evaluation image" (e.g., registered evaluation image), "attribute operation", "attribute operation intensity", "virtual evaluation image", "recognition score", "recognition result" and "added to learning".

[0142] In the "Misidentified Image" area, for example, image data of products that have been determined by the product identification unit 12 to be misidentified can be displayed.

[0143] In the "Registered Evaluation Image" area, for example, an image that is the evaluation object can be displayed from the evaluation images registered in the performance evaluation image DB 141. For example, the product (evaluation image data) that becomes the object of the registered evaluation image can be determined based on the selection made by at least one of the misidentification analysis unit 14 and the user. For example, for products with a recognition rate below a threshold (e.g., 95%) in Display Example 1, as in Display Example 2, a list of recognition scores based on different attribute operations can be displayed. Alternatively, a display can be made to prompt the selection of the object product displayed on the screen of Display Example 2 from among the multiple products whose recognition rates are displayed in Display Example 1.

[0144] In the "Attribute Operations" area, you can, for example, display whether any attributes of the evaluation image can be manipulated, and the types of attributes that can be manipulated. Although in Figure 8 The example shown is an example of displaying one attribute for each evaluation image, but multiple attributes can also be displayed.

[0145] In the "Attribute Operation Intensity" area, the intensity of the attribute operation displayed in the "Attribute Operation" area can be shown, for example. The intensity of the attribute operation can show, for example, the value (level) of one of a number of candidate intensities representing the attribute operation, or the actual value set in each attribute operation (e.g., lighting direction [degrees], lighting intensity [lx], color temperature [K]).

[0146] Furthermore, the attributes of the object being operated on and the intensity of the operation on those attributes can be selected by accepting user input, or can be predetermined by the misidentification analysis unit 14.

[0147] In the "Virtual Evaluation Image" area, for example, a virtual evaluation image can be displayed that has been modified by applying misidentification conditions (e.g., lighting environment-related conditions) to the evaluation image, and the virtual evaluation image obtained by the attribute operations displayed by "Attribute Operations" and "Attribute Operation Intensity". For example, in Figure 8 The system can display virtual evaluation image data obtained by applying lighting conditions corresponding to the misidentified image of a plastic bottle beverage, as well as attribute operations corresponding to "attribute operations" and "attribute operation intensity." For example, in... Figure 8 In the example, when the attribute operation for the evaluation image data of the cup surface is "none", a virtual evaluation image data of the cup surface obtained by applying lighting conditions corresponding to the misidentified image of the plastic bottle beverage can be displayed. Additionally, for example in... Figure 8 In the context of the evaluation image data for the cup surface, when the attribute operation for the cup surface is "orientation change," a virtual evaluation image data of the cup surface can be displayed, showing the cup surface with the lighting conditions corresponding to the misidentified image of the plastic bottle beverage and the orientation changed. Similarly, for example, in... Figure 8 In the process, when the attribute operation for the evaluation image data of the cup surface is "reflection", a virtual evaluation image data of the cup surface can be displayed, which is obtained by applying lighting conditions corresponding to the misidentified image of the plastic bottle beverage and changing the intensity of light reflection.

[0148] In the "Recognition Score" area, for example, the evaluation results for the virtual evaluation image can be displayed. As an example of the evaluation results, a recognition score can be displayed, which is an indicator of the reliability of the recognition of the correct category (correct product) estimated for the virtual evaluation image. The misidentification analysis unit 14 can, for example, determine a recognition score representing the reliability of each category for multiple (e.g., all) categories of product recognition results from the virtual evaluation image. The recognition score can, for example, be a value where the sum of the recognition scores of multiple categories is 1 (in other words, a value normalized to the range of 0 to 1). For example, the closer the recognition score of a category is to 1, the more reliable the recognition result for that category. In other words, the closer the recognition score is to 0, the higher the risk of misidentification.

[0149] For example, in Figure 8 In the process, the misidentification analysis unit 14 can determine the recognition scores of multiple categories (e.g., multiple categories including cup noodles, snacks, groceries, or beverages) in the product recognition results of the evaluation image data of the cup noodles, and display the correct category, i.e., the recognition score of the cup noodles, among the recognition scores of multiple categories on the screen.

[0150] In the "Identification Results" area, for example, one of "Misidentification" and "Correct Identification" can be displayed as the evaluation result of the misidentification analysis unit 14. For example, the misidentification analysis unit 14 can determine "Correct Identification" if the identification score of the correct category is the highest among the identification scores of multiple categories, and determine "Misidentification" if the identification score of a category different from the correct category is the highest.

[0151] In the "Add to Learning" area, for example, a list display may be included, prompting the user to input information that determines whether to use multiple virtual evaluation image data for relearning (or supplementary learning) (e.g., adding to training data). For example, the user can select the checkbox displayed in the "Add to Learning" column to decide to add relevant virtual evaluation image data to the training data. Alternatively, the misidentification analysis unit 14 can also add virtual evaluation image data that is misidentified to the training data (e.g., by selecting the "Add to Learning" checkbox). Alternatively, the misidentification analysis unit 14 can also add virtual evaluation image data with recognition scores below a threshold to the training data (e.g., by selecting the "Add to Learning" checkbox). By adding virtual evaluation image data with recognition results of misidentification or recognition scores below a threshold to the training data as images with a high probability of misidentification, misidentification can be determined more accurately.

[0152] Alternatively, a checkbox for "Add to Learning" that automatically selects virtual evaluation image data meeting predetermined conditions can be provided, or a button indicating automatic selection of the checkbox can be added to reduce the effort required for user-initiated selection. Furthermore, virtual evaluation image data can be automatically added to the training data without user confirmation, thus automating the selection of training data. In this case, the display of the checkbox itself can be omitted. For example, a criterion for automatic selection could be listed as a misidentification result or a recognition score below a threshold.

[0153] For example, if the misidentification analysis unit 14 detects that the "output" button (not shown) has been pressed, it can add the virtual evaluation image data selected in the "Add to Learning" column to the training data.

[0154] Here, the lower the recognition score of the correct category in the virtual evaluation image data, the greater the effect of improving the recognition rate through relearning. Therefore, it is desirable to add this virtual evaluation image data to the training data. For example, if the "recognition score" is below a threshold, the misidentification analysis unit 14 can also output the result display unit 15 with the following parameters: Figure 9 The information of the region where the recognition score is highlighted is shown. Furthermore, the highlighted region is not limited to the recognition score; it may also include other regions of the relevant virtual evaluation image data. By highlighting this region, the misidentification analysis unit 14 can prompt the user to add the relevant virtual evaluation image data to the training data.

[0155] In addition, virtual evaluation image data that has been judged as "correctly identified" can be used as data for correct category supplementation. Generally speaking, the higher the recognition rate, the more difficult it is to improve. Therefore, if the threshold for judging as "correctly identified" is high enough, although the effect will be weakened, it is still possible to expect an improvement in accuracy through such supplementation.

[0156] Thus, according to Example 2, for objects that are likely to be misidentified, candidate images added to the training data are displayed. Therefore, users can easily add images with a higher risk of misidentification (e.g., images with lower recognition scores) to the training data, thereby improving the accuracy of product recognition and reducing the risk of misidentification through supplementary learning.

[0157] Furthermore, for example, the higher the risk of misidentification in virtual evaluation image data, the easier it is to add it to the training data. Conversely, the lower the risk of misidentification in virtual evaluation image data, the less likely it is to be added to the training data. Therefore, for example, it is possible to suppress the addition of new product image data for learning processes that have a higher processing cost than product recognition, thereby improving the efficiency of relearning.

[0158] also, Figure 8 and Figure 9 The displayed content is an example and is not limited to these displayed contents. For example, it could display... Figure 8 and Figure 9 The information shown may include a portion of the misidentified image, registration evaluation image, attribute operation, strength of attribute operation, virtual evaluation image, recognition score, recognition result, and information added to the learning process. Alternatively, other information may be displayed in addition to the information shown in the misidentified image.

[0159] The above illustrates examples of risks of misidentification.

[0160] Thus, the misidentification analysis unit 14, for example, based on the attribute information of misidentified image data obtained by taking pictures of the product in the environment where misidentification occurred, virtually generates image data of other products in the environment where misidentification occurred (e.g., lighting environment) (in other words, recreates images of other products), evaluates the image recognition accuracy (e.g., misidentification risk) for other products, and outputs a signal for displaying (visualizing) the evaluation results of the image recognition accuracy.

[0161] Therefore, the misidentification analysis unit 14 can, for example, determine (or assess) the risk of misidentification of a product that is different from the product that was misidentified in an environment where product misidentification has occurred. Thus, for example, in cases where different items in a specific environment may become the objects of image recognition, the misidentification analysis unit 14 can assess the image recognition accuracy for different items.

[0162] For example, in addition to identifying items misidentified in lighting environments different from the learning environment, users can also identify the risk of misidentification in lighting environments where misidentification occurred for other items, and can take measures such as additional learning and adjustments to the lighting environment based on the risk of misidentification. Therefore, according to this embodiment, it is possible to prevent misidentification of other items in the same recognition environment as the misidentified item, and to reduce the misidentification rate of image recognition.

[0163] Furthermore, the product recognition system 1 can perform product recognition not only indoors (e.g., inside a store), but also outdoors. For example, the misidentification analysis unit 14 can also transform the evaluation image data based on attributes related to the light environment formed by at least one of the lighting environment and sunlight (e.g., misidentification conditions).

[0164] In addition, although Figures 3-6 The document describes how to display the risk of misidentification (e.g., recognition rate) in a chart, but the display method of the risk of misidentification is not limited to a chart and can also be other display methods (e.g., a list).

[0165] In addition, although Figures 3-6 The example shown illustrates how to display the areas of "Recognition Performance Evaluation Condition Selection", "Recognition Performance Evaluation Result" and "Lighting Recommendation Result" on one screen, but it is not limited to this. At least one of the areas of "Recognition Performance Evaluation Condition Selection", "Recognition Performance Evaluation Result" and "Lighting Recommendation Result" can also be displayed on other screens.

[0166] Alternatively, the structural elements included in the product recognition system 1 may be located at the place where product recognition is performed (e.g., a store). Or, for example, the image acquisition unit 11 and the result display unit 15 of the structural elements included in the product recognition system 1 may be located at the place where product recognition is performed (e.g., a store), while other structural elements may be located physically separate from the image acquisition unit 11 and the result display unit 15 (e.g., at least one server). For example, at least one process for misidentification risk assessment may be performed on the server.

[0167] Furthermore, although the example described in the above embodiment of an item as the object of assessment for misidentification risk (or image recognition object) is a product displayed in a store, the item to be assessed for misidentification risk is not limited to products. That is, as long as the misidentification risk is being assessed for an item that is different from the item misidentified in image recognition, the disclosure of this embodiment can also be applied to facilities other than stores or items other than products.

[0168] Alternatively, virtual evaluation image data can be generated using general image processing techniques different from the generation methods 1 to 3 described in the above embodiments. Specifically, examples include geometric transformations of the product's shape or orientation, and chroma or hue transformations. In these cases, it is sufficient to extract the distribution of shape or color from the image data determined to be misidentified and use it as an attribute.

[0169] Furthermore, although the above embodiments describe the application of lighting environment-related attributes of image data deemed misidentified to the evaluation image data, this is not limited to image data deemed misidentified. For example, lighting environment-related attributes of image data not deemed misidentified (e.g., correctly identified image data) can also be applied to the evaluation image data. Therefore, it is possible to reproduce image data of multiple objects in a variety of recognition environments without actually capturing images in the environment, through image processing (e.g., attribute manipulation), thus making it easy to evaluate the image recognition accuracy for multiple objects.

[0170] In addition, the term "...part" in the above embodiments can be replaced with other terms such as "...circuitry", "...component", "device", "unit" or "module".

[0171] Although the embodiments of this disclosure have been described in detail above with reference to the accompanying drawings, the functions of the commodity management system 1 described above can be implemented by a computer program.

[0172] Figure 10 This is a diagram illustrating the hardware structure of a computer that implements the functions of each device through a program. The computer 1100 includes input devices 1101 such as a keyboard or mouse and a touchpad, output devices 1102 such as a monitor or speakers, a CPU (Central Processing Unit) 1103, a GPU (Graphics Processing Unit) 1104, ROM (Read Only Memory) 1105, RAM (Random Access Memory) 1106, a storage device 1107 such as a hard disk drive or SSD (Solid State Drive), a reading device 1108 for reading information from recording media such as DVD-ROM (Digital Versatile Disk Read Only Memory) or USB (Universal Serial Bus) storage, and a transceiver device 1109 for communication via a network. All parts are connected via a bus 1110.

[0173] Furthermore, the reading device 1108 reads the program from a recording medium containing a program for implementing the functions of the aforementioned devices and stores the program in the storage device 1107. Alternatively, the transceiver device 1109 communicates with a server device connected to a network and stores a program downloaded from the server device for implementing the functions of the aforementioned devices in the storage device 1107.

[0174] Next, the CPU 1103 copies the program stored in the storage device 1107 to the RAM 1106, and sequentially reads and executes the commands contained in the program from the RAM 1106, thereby realizing the functions of the above-mentioned devices.

[0175] This disclosure can be implemented by software, hardware, or software in collaboration with hardware.

[0176] The functional blocks used in the above embodiments are implemented partially or wholly as LSIs (Large Scale Integration) of integrated circuits. The processes described in the above embodiments can also be controlled partially or wholly by a single LSI or a combination of LSIs. An LSI can be composed of individual chips, or it can be composed of a single chip containing some or all of the functional blocks. An LSI may also include data input and output. Depending on the degree of integration, an LSI may also be called an "IC (Integrated Circuit)," a "System LSI," a "Super LSI," or an "Ultra LSI."

[0177] The method of integrating LSIs is not limited to LSIs; it can also be implemented using dedicated circuits, general-purpose processors, or special-purpose processors. Alternatively, LSIs can be used to fabricate programmable FPGAs (Field Programmable Gate Arrays), or reconfigurable processors that allow for reconfiguration of the connections or settings of the circuit blocks within the LSI. This disclosure can also be implemented for digital or analog processing.

[0178] Furthermore, if advancements in semiconductor technology or the emergence of other derivative technologies lead to integrated circuit technologies that can replace LSIs, these technologies could also be used to integrate functional blocks. There are also possibilities for applications such as biotechnology.

[0179] This disclosure can be implemented in all kinds of devices, apparatuses, and systems with communication capabilities (collectively referred to as "communication devices"). A communication device may also include a wireless transceiver and processing / control circuitry. The wireless transceiver may also include a receiving unit and a transmitting unit, or perform the functions of these units. The wireless transceiver (transmitting unit, receiving unit) may also include an RF (Radio Frequency) module and one or more antennas. The RF module may also include an amplifier, an RF modulator / demodulator, or similar devices. Non-limiting examples of communication devices include: telephones (mobile phones, smartphones, etc.), tablet computers, personal computers (PCs) (laptops, desktops, laptops, etc.), cameras (digital cameras, digital camcorders, etc.), digital players (digital audio / video players, etc.), wearable devices (wearable cameras, smartwatches, tracking devices, etc.), game consoles, e-book readers, remote health / telemedicine (remote healthcare / medical prescription) devices, vehicles or transportation vehicles with communication capabilities (cars, airplanes, ships, etc.), and combinations of the various devices described above.

[0180] Communication devices are not limited to portable or movable devices, but also include all kinds of devices, equipment, and systems that cannot be carried or fixed. Examples include: smart home devices (home appliances, lighting equipment, smart meters or meters, control panels, etc.), vending machines, and all other "things" that can exist on the IoT (Internet of Things) network.

[0181] In addition, in recent years, a new concept called CPS (Cyber ​​Physical Systems) has gained attention in IoT (Internet of Things) technology, which utilizes information collaboration between physical space and cyberspace to generate new added value. This CPS concept can also be adopted in the above-described implementation.

[0182] That is, as a basic structure of CPS, for example, edge servers configured in physical space and cloud servers configured in information space can be connected via a network, and the processing is distributed by utilizing the processors on both servers. Here, the processing data generated by the edge server or cloud server is preferably generated on a standardized platform. By using such a standardized platform, the efficiency of building systems that include a wide variety of sensor arrays or IoT application software can be improved.

[0183] In the above embodiments, for example, the edge server can be configured in the store to perform product identification processing and product misidentification risk assessment processing. The cloud server can also use data received from the edge server via the network for model learning. Alternatively, for example, the edge server can be configured in the store to perform product identification processing, and the cloud server can use data received from the edge server via the network to perform product misidentification risk assessment processing.

[0184] In addition to data communication via cellular systems, wireless LAN (Local Area Network) systems, and communication satellite systems, communication also includes data communication via a combination of these systems.

[0185] In addition, the communication device also includes devices such as controllers or sensors that are connected or linked to the communication equipment performing the communication functions described in this invention. For example, it includes controllers or sensors that generate control signals or data signals used by the communication equipment performing the communication functions of the communication device.

[0186] In addition, the communication device includes infrastructure equipment that communicates with or controls the various devices described above (not limited to these), such as base stations, access points, and all other devices, equipment, and systems.

[0187] While various embodiments have been described above with reference to the accompanying drawings, this disclosure is not limited to these examples. Those skilled in the art will understand that various modifications and alterations are readily apparent within the scope of the claims, and these also fall within the technical scope of this disclosure. Furthermore, the constituent elements of the above embodiments can be arbitrarily combined without departing from the spirit of the disclosure.

[0188] While specific examples of this disclosure have been described in detail above, these examples are merely illustrative and do not limit the scope of the claims. The technology described in the claims includes various modifications and alterations to the specific examples illustrated above.

[0189] An information processing apparatus according to an embodiment of this disclosure includes: a generation circuit that generates second image data based on attribute information of first image data obtained by photographing a first item in a first environment, the second image data reproducing an image of a second item, which is different from the first item, disposed in the first environment; an evaluation circuit that evaluates the image recognition accuracy for the second item based on the second image data; and an output circuit that outputs the evaluation result of the accuracy.

[0190] In one embodiment of this disclosure, the evaluation result includes information representing the probability of successfully recognizing the second item in the second image data during image recognition; the output circuit outputs a signal for displaying information representing the probability of successfully recognizing the second item in the second image data.

[0191] In one embodiment of this disclosure, the output circuit outputs a signal that is used to distinguish between information about the second item in the second image data where the probability of successfully identifying the second item exceeds a predetermined threshold, and information about the second item where the probability does not exceed the predetermined threshold.

[0192] In one embodiment of this disclosure, the generation circuit generates the second image data by changing the attribute information of the image data obtained by photographing the second item based on the attribute information of the first image data.

[0193] In one embodiment of this disclosure, the generation circuit applies first attribute information related to the lighting environment of the image recognition from the attribute information of the first image data to the image data obtained by photographing the second item, in order to generate the second image data.

[0194] In one embodiment of this disclosure, the evaluation result includes information representing the probability of successfully recognizing the second item in the second image data during image recognition; if the probability is below a threshold, the output circuit outputs a signal for displaying a candidate set value of the first attribute information that would make the probability greater than the threshold.

[0195] In one embodiment of this disclosure, the parameters used to determine the candidate setting value are predetermined or selected by the user.

[0196] In one embodiment of this disclosure, the generation circuit further adjusts at least one of the second attribute information related to feature quantities different from the lighting environment for the second image data to generate new second image data; the output circuit outputs a signal for displaying the evaluation result for the new second image data.

[0197] In one embodiment of this disclosure, the output circuit outputs a signal for displaying the evaluation results corresponding to each of the plurality of new second image data generated by making different adjustments to the second attribute information.

[0198] In one embodiment of this disclosure, the output circuit outputs a signal for displaying in the overview display that prompts for input of information determining whether to relearn using the modified plurality of second image data.

[0199] In one embodiment of this disclosure, the generation circuit uses Generative Adversarial Networks (GANs) to generate the second image data.

[0200] In one embodiment of the information processing method disclosed herein, the information processing apparatus performs the following steps: generating second image data based on attribute information of first image data obtained by photographing a first item in a first environment, wherein the second image data reproduces an image of a second item, which is different from the first item, in the first environment; evaluating the image recognition accuracy for the second item based on the second image data; and outputting the evaluation result of the accuracy.

[0201] An embodiment of the present disclosure describes a process that causes a computer to perform processing to generate second image data based on attribute information of first image data obtained by photographing a first item in a first environment, the second image data reproducing an image of a second item, which is different from the first item, in the first environment; processing to evaluate the image recognition accuracy for the second item based on the second image data; and processing to output the evaluation result of the accuracy.

[0202] The entire contents of the specification, drawings and abstract of the specification contained in Japanese Patent Application No. 2021-077603, filed on April 30, 2021, are incorporated herein by reference.

[0203] Industrial applicability

[0204] One embodiment of this disclosure is useful for a product identification system.

[0205] Explanation of reference numerals in the attached figures

[0206] 1. Product Identification System

[0207] 11 Image Acquisition Unit

[0208] 12. Product Identification Department

[0209] 13 Storage Department

[0210] 14 Misidentification Analysis Department

[0211] 15 Results Display Department

[0212] 141 Image Database for Performance Evaluation

[0213] 142 Image Coding Unit

[0214] 143 Attribute Operation Department

[0215] 144 Image Generation Unit

[0216] 145 Assessment Department

Claims

1. An information processing apparatus, comprising: The generation circuit generates second image data based on the attribute information of the first image data obtained by photographing the first item in the first environment. The second image data reproduces the image of the second item configured in the first environment, which is different from the first item. An evaluation circuit evaluates the image recognition accuracy for the second item based on the second image data. as well as The output circuit outputs the evaluation result of the accuracy. The generation circuit, based on the attribute information of the first image data, modifies the attribute information of the image data obtained by photographing the second item to generate the second image data. The evaluation result includes information indicating the probability of successfully recognizing the second item in the second image data during the image recognition process; When the probability is below a threshold, the output circuit outputs a signal that displays a candidate set value for the attribute information of the first image data that would cause the probability to become greater than the threshold.

2. The information processing apparatus as described in claim 1, wherein, The evaluation result includes information indicating the probability of successfully recognizing the second item in the second image data during the image recognition process; The output circuit outputs a signal to display information indicating the probability of successfully identifying the second item in the second image data.

3. The information processing apparatus as described in claim 2, wherein, The output circuit outputs a signal that is used to distinguish between information about the second item in the second image data where the probability of successfully identifying the second item exceeds a predetermined threshold, and information about the second item where the probability does not exceed the predetermined threshold.

4. The information processing apparatus as described in claim 1, wherein, The generation circuit applies the first attribute information related to the lighting environment of the image recognition from the attribute information of the first image data to the image data obtained by photographing the second item, so as to generate the second image data.

5. The information processing apparatus as described in claim 1, wherein, The parameters used to determine the candidate setting value are predetermined or selected by the user.

6. The information processing apparatus as described in claim 4, wherein, The generation circuit further adjusts at least one of the second attribute information related to feature quantities different from the lighting environment for the second image data to generate new second image data; The output circuit outputs a signal for displaying the evaluation result for the new second image data.

7. The information processing apparatus as claimed in claim 6, wherein, The output circuit outputs a signal that is used to display the evaluation results corresponding to the multiple new second image data generated by making different adjustments to the second attribute information.

8. The information processing apparatus as claimed in claim 7, wherein, The output circuit outputs a signal for displaying the following information in the overview display: the display prompts for input of information that determines whether to relearn using the modified plurality of second image data.

9. The information processing apparatus as claimed in claim 1, wherein, The first image data is the image data that was determined to be misidentified in the image recognition.

10. The information processing apparatus as claimed in claim 1, wherein, The generation circuit uses a generative adversarial network, or GAN, to generate the second image data.

11. An information processing method, wherein, The information processing device performs the following steps: Based on the attribute information of the first image data obtained by photographing the first item in the first environment, second image data is generated, and the second image data reproduces the image of the second item that is configured in the first environment and is different from the first item; Based on the second image data, evaluate the image recognition accuracy for the second item; as well as Output the evaluation results of the accuracy. Specifically, based on the attribute information of the first image data, the attribute information of the image data obtained by photographing the second item is changed to generate the second image data. The evaluation result includes information indicating the probability of successfully recognizing the second item in the second image data during the image recognition process; When the probability is below a threshold, a signal is output to display a candidate setting value for the attribute information of the first image data that would cause the probability to become greater than the threshold.

12. A computer program product comprising a program that causes a computer to perform the following processes: Based on the attribute information of the first image data obtained by photographing the first item in the first environment, the processing generates second image data, which reproduces an image of a second item that is different from the first item in the first environment; Based on the second image data, evaluate the processing of image recognition accuracy for the second item; as well as The processing that outputs the evaluation results of the aforementioned accuracy. Specifically, based on the attribute information of the first image data, the attribute information of the image data obtained by photographing the second item is changed to generate the second image data. The evaluation result includes information indicating the probability of successfully recognizing the second item in the second image data during the image recognition process; When the probability is below a threshold, a signal is output to display a candidate setting value for the attribute information of the first image data that would cause the probability to become greater than the threshold.