Information processing system

The information processing system addresses the labor-intensive and impractical challenges of generating learning data for POPs by detecting and classifying POPs, enhancing accuracy and efficiency in model generation.

JP2025159139AActive Publication Date: 2025-10-17MARKETVISION CO LTD
View PDF 4 Cites 0 Cited by

Patent Information

Application Number
JP2025135642
Authority / Receiving Office
JP · JP
Patent Type
Applications
Current Assignee / Owner
Filing Date
2025-08-18
Publication Date
2025-10-17
Estimated Expiration
2043-09-15

AI Technical Summary

Technical Problem

Existing methods for generating learning data for product promotion objects (POPs) in retail stores are labor-intensive and require pre-labeled image databases, which are impractical for diverse POPs.

Method used

An information processing system that detects POPs from image information, classifies them, and generates a learning model using provisional identification information, reducing workload and eliminating the need for pre-labeled databases.

Benefits of technology

Efficiently generates learning models for POPs, improving classification accuracy by detecting and distinguishing POPs from merchandise and shelf components, and reducing manual labor.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure 2025159139000001_ABST
    Figure 2025159139000001_ABST
Patent Text Reader

Abstract

To provide an information processing system related to the generation of a learning model in machine learning.SOLUTION: An information processing system for performing processing related to a learning model used in machine learning has: a detection processing unit which detects a POP from image information; and a model processing unit which generates a learning model for machine learning by using the detected POP. The model processing unit has: a classification processing unit which classifies image information of the detected POP; a temporary identification information processing unit which associates a classified group with temporary identification information; and a model generation processing unit which performs learning processing in machine learning to generate a learning model by using data for learning containing the image information of the POP included in the group and the temporary identification information.SELECTED DRAWING: Figure 1
Need to check novelty before this filing date? Find Prior Art

Description

[Technical Field]

[0001] The present invention relates to an information processing system used to generate a learning model (network) in machine learning. [Background technology]

[0002] In retail stores and other stores, POPs for product promotion are sometimes placed on display shelves. POPs are attached, mounted, or installed on products, product tags (price cards), shelves, or above or to the sides of shelves.

[0003] There are two types of POPs: manufacturer POPs, which are created by product manufacturers and requested to be installed by retailers, and retail POPs (distribution POPs), which are created and installed by retailers themselves. Since manufacturers request retailers to install manufacturer POPs, manufacturers want to check whether they are installed as requested.

[0004] In the past, in such cases, a manufacturer's representative would visit each store to check the product, but this was a heavy burden. Therefore, it is thought that automatic processing could be implemented.

[0005] An example of automatic processing is machine learning, a method of image analysis using a computer. Machine learning requires that a learning dataset (learning data) be created in advance, and that the computer then loads it into the computer to execute the required learning process.

[0006] However, preparing learning data in advance itself is a heavy workload. To address this issue, systems such as those shown in Patent Documents 1 and 2 below have been disclosed.

[0007] In the invention of Patent Document 1, product images are photographed and codes attached to the products are read, and product images and labels are then created to generate learning data.

[0008] Furthermore, in the invention of Patent Document 2, a classification of an object contained in image data is selected from an image database storing image data including objects that have been labeled with information relating to the classification in advance, and the object is extracted from the image data to generate a template image, thereby generating training data labeled with the template image. [Prior art documents] [Patent documents]

[0009] [Patent Document 1] Japanese Patent Publication No. 2020-095537 [Patent Document 2] Japanese Patent Publication No. 2022-076296 Summary of the Invention [Problem to be solved by the invention]

[0010] In the case of the invention of Patent Document 1, products are prepared, photographed, and the codes attached to the products are read with a code reader to generate learning data, which places a heavy workload on the worker. If there are many products to be read as learning data, the workload becomes enormous. Furthermore, while codes are attached to products, codes are not attached to POPs, so the invention of Patent Document 1 cannot be used to generate learning data for POPs.

[0011] In the case of the invention of Patent Document 2, the processing cannot be performed unless an image database is prepared in which image data including objects labeled with classification information is stored in advance. POPs are diverse, and it is not realistic to prepare an image database with image data labeled for each POP. [Means for solving the problem]

[0012] In view of the above-mentioned problems, the present inventors have invented an information processing system that efficiently generates learning models for use in machine learning. The information processing system of the present invention is particularly effective when used to generate learning models for POPs that are installed on store shelves.

[0013] The first invention is an information processing system that performs processing related to a learning model used in machine learning, and includes a detection processing unit that detects POPs from image information, and a model processing unit that generates a learning model for machine learning using the detected POPs.The model processing unit includes a classification processing unit that classifies image information of the detected POPs, a provisional identification information processing unit that associates the classified groups with provisional identification information, and a model generation processing unit that performs machine learning learning processing using learning data including image information of POPs included in the groups and the provisional identification information to generate a learning model.

[0014] By using this invention, it is possible to efficiently generate a learning model used in machine learning. In particular, image information used as learning data and its label (type) can be detected from the image information and used as a dataset, thereby reducing the workload.

[0015] The inventors of this application noticed that a typical display shelf is equipped with merchandise, such as displayed goods, product tags, shelf components, and POPs. In other words, while POPs themselves are not easy to distinguish because the shapes and designs of the POPs themselves are diverse, it is easy to distinguish merchandise, product tags, and shelf components. Therefore, if an object detected from image data is not a merchandise, product tag, or shelf component, it is determined to be a POP, making it possible to identify the POP in the image data.

[0016] In the above-mentioned invention, the model processing unit can be configured as an information processing system having a cleansing processing unit that performs a cleansing process to exclude image information of POPs that may contain errors from the image information of POPs included in the classified group.

[0017] By performing a cleansing process as in the present invention, it is possible to remove object image information contained in a group that may contain errors, thereby improving the accuracy of classification.

[0018] The first invention can be realized by loading and executing an information processing program of the present invention into a computer. That is, the information processing program causes a computer to function as a detection processing unit that detects POPs from image information and a model processing unit that generates a learning model for machine learning using the detected POPs, wherein the model processing unit includes a classification processing unit that classifies image information of the detected POPs, a provisional identification information processing unit that associates the classified groups with provisional identification information, and a model generation processing unit that executes a machine learning learning process using learning data including image information of POPs included in the groups and the provisional identification information to generate a learning model. [Effects of the Invention]

[0019] By using the information processing system of the present invention, it is possible to efficiently generate a learning model used in machine learning. [Brief explanation of the drawings]

[0020] [Figure 1] FIG. 2 is a block diagram schematically illustrating an example of a processing function of the information processing system of the present invention. [Figure 2] FIG. 2 is a block diagram schematically illustrating an example of processing functions of a detection processing unit of the information processing system of the present invention. [Figure 3] FIG. 2 is a block diagram schematically illustrating an example of a processing function of a model processing unit of the information processing system of the present invention. [Figure 4] FIG. 2 is a block diagram schematically illustrating an example of a hardware configuration of a computer used in the information processing system of the present invention. [Figure 5] FIG. 2 is a diagram illustrating an example of a conceptual diagram of a detection processing unit of the information processing system of the present invention. [Figure 6] FIG. 1 is a diagram showing an example of a conceptual diagram of processing in an information processing system according to the present invention. [Figure 7] FIG. 1 is a diagram showing an example of a conceptual diagram of processing in an information processing system according to the present invention. [Figure 8] FIG. 10 is a diagram showing an example of captured image information. [Figure 9] 3 is a flowchart showing an example of the overall processing process in the information processing system of the present invention. [Figure 10] 10 is a flowchart showing an example of a processing process of a detection processor in the information processing system of the present invention. [Figure 11] 10 is a flowchart showing an example of a processing process of a model processing unit in the information processing system of the present invention. [Figure 12] FIG. 10 is a diagram illustrating an example of a result of executing an object detection process from image information. [Figure 13] 10A and 10B are diagrams illustrating an example of a process for determining the area of ​​an object detected from image information by determining whether the area satisfies a predetermined condition. [Figure 14] 13 is a diagram showing an example of each region after regions that satisfy predetermined conditions have been determined as regions of the target object in the object detection results of FIG. 12. FIG. [Figure 15] FIG. 10 is a diagram showing an example of a state in which object image information is extracted from image information. [Figure 16] 10 is a diagram showing an example of a state in which object image information is classified in a classification processing unit. FIG. [Figure 17] 10 is a diagram showing an example of a state in which temporary identification information is associated in a temporary identification information processing unit. FIG. [Figure 18] FIG. 10 is a diagram illustrating an example of a state in which a learning model is generated using a learning dataset. [Figure 19] FIG. 10 is a diagram showing an example of a state in which sample information is input to a learning model and tentative identification information is output as an output value. [Figure 20] FIG. 10 is a diagram illustrating an example of a state in which temporary identification information and object identification information are linked together. [Figure 21]FIG. 10 is a block diagram illustrating an example of the configuration of a detection processing unit according to a second embodiment. [Figure 22] 10 is a flowchart illustrating an example of a processing process of a detection processing unit in the second embodiment. [Figure 23] FIG. 10 is a diagram illustrating an example of an outline of a process according to a second embodiment. [Figure 24] FIG. 10 is a diagram illustrating an example of an individual determination process in the second embodiment. [Figure 25] FIG. 10 is a diagram schematically illustrating a state in which input of designation of a product display area, a product tag area, and a shelf top area is received for image information. [Figure 26] 11 is a flowchart illustrating an example of a processing process of a detection processing unit in the third embodiment. [Figure 27] FIG. 10 is a block diagram illustrating an example of the configuration of a model processing unit according to a fourth embodiment. DETAILED DESCRIPTION OF THE INVENTION

[0021] An example of the processing function of the information processing system 1 of the present invention is shown in block diagrams in Figures 1 to 3. Figure 1 shows an example of the overall processing function of the information processing system 1 in block diagram form, Figure 2 shows an example of the processing function of the detection processing unit 20 in block diagram form, and Figure 3 shows an example of the processing function of the model processing unit 21 in block diagram form.

[0022] The information processing system 1 is a system that detects objects captured in predetermined image information, and generates a learning model to be used in machine learning using the image information of the detected various objects.

[0023] For example, if the target object is a product displayed on a display shelf, a product tag (price card), a shelf, or a POP attached, mounted, or installed above or to the side of a shelf, the system detects the POP shown in the image information of the display shelf that includes the POP, and uses the POP to generate a learning model to be used in machine learning. Note that the target object will be described using the case of a product displayed on a display shelf, a product tag (price card), a shelf, or a POP attached, mounted, or installed above or to the side of a shelf, but is not limited to this.

[0024] The management terminal 2 and the image information input terminal 3 in the information processing system 1 are realized using a computer. An example of the hardware configuration of a computer is shown in Figure 4. The computer has a calculation device 70 such as a CPU that executes the calculation processing of a program, a storage device 71 such as a RAM or hard disk that stores information, a display device 72 such as a display that displays information, an input device 73 such as a keyboard or mouse that can input information, and a communication device 74 that sends and receives the processing results of the calculation device 70 and the information stored in the storage device 71 via a network such as the Internet or a LAN.

[0025] If the computer is equipped with a touch panel display, the display device 72 may be integrated with the input device 73. Touch panel displays are often used in portable communication terminals such as tablet computers and smartphones, but are not limited to these.

[0026] The touch panel display is a device that integrates the functions of the display device 72 and the input device 73 in that input can be made directly on the display using a predetermined input device (such as a touch panel pen) or a finger.

[0027] In addition to the above devices, the image information input terminal 3 may be equipped with a photographing device such as a camera. A portable communication terminal such as a mobile phone, smartphone, or tablet computer can also be used as the image information input terminal 3. The image information input terminal 3 photographs image information (photographed image information) of a predetermined object and inputs it into the management terminal 2. For example, a display shelf in a store is photographed and the image information is input into the management terminal 2. In this case, products are displayed on the display shelf, and POPs are attached to the display shelf or the products. Therefore, the photographed image information includes images of the displayed products, such as the products displayed on the display shelf, product tags, shelf components, POPs, etc. It is preferable that unnecessary objects such as people are not photographed.

[0028] The functions of the various means in the present invention are only logically distinct, and may be physically or practically the same area. The order of the processes in the various means of the present invention may be changed as appropriate. In addition, some of the processes may be omitted.

[0029] The management terminal 2 includes a detection processing unit 20 and a model processing unit 21.

[0030] The detection processing unit 20 detects image information of an area in which an object is captured from the captured image information (object image information). For example, image information of an area in which a POP is captured (POP image information) is detected as object image information from image information of a display shelf. Figure 5 is an example of a conceptual diagram of processing in the detection processing unit 20. Figure 5(a) is an example of captured image information in which an object is captured, and Figure 5(b) is an example of a state in which an object has been detected from the captured image information. The captured image information in Figure 5(a) is a diagram that schematically illustrates an image of a store display shelf. Figure 5(b) is a diagram that schematically illustrates an example in which a POP is detected as an object. The shaded area in Figure 5(b) is image information of the area of ​​the detected POP.

[0031] The model processing unit 21 generates a learning model to be used in machine learning for the object image information detected by the detection processing unit 20. Examples of conceptual diagrams of the processing in the model processing unit 21 are shown in Figs. 6 and 7. Fig. 6 is a diagram that schematically shows the processing of generating a learning model using the object image information detected by the detection processing unit 20. Fig. 7 is a diagram that schematically shows the processing of linking temporary identification information associated with groups for each classified object image information with identification information of the object (object identification information).

[0032] The detection processing unit 20 includes an image information input reception processing unit 200 , an image information storage unit 201 , an image information alignment processing unit 202 , an object detection processing unit 203 , and an individual discrimination processing unit 204 .

[0033] The image information input reception processing unit 200 accepts input of captured image information of an object captured by the image information input terminal 3 and stores the image information in the image information storage unit 201 (described later). For example, the image information may be input of an image of a store display shelf and stored in the image information storage unit 201. In addition to the captured image information, the image information input terminal 3 may also accept input of accompanying information such as store identification information (e.g., the date and time of capture, the store name), display shelf identification information for identifying the display shelf, and image information identification information for identifying the image information. FIG. 8 shows an example of image information of a display shelf as captured image information. Note that FIG. 8 shows a case where the captured image information shows one display shelf, but multiple display shelves or even a portion of a display shelf may be captured. Furthermore, although not specifically described in the present invention, display shelves and shelves may be long horizontally. Therefore, in this processing, the image may be divided into sections with a certain width and the divided sections may be used as the processing target for writing.

[0034] The image information storage unit 201 stores captured image information received from the image information input terminal 3 in association with accompanying information such as the date and time of capture, store identification information, display shelf identification information, and image information identification information. The captured image information may be any image information that is the subject of processing in the present invention. It also includes image information obtained by combining multiple images of a single display shelf into a single image. The captured image information also includes corrected image information that has been processed by the image information correction processing unit 202 (described later) and image information after distortion correction processing.

[0035] The image information orthogonalization processing unit 202 executes a correction process (orthogonalization process) on the captured image information stored in the image information storage unit 201 to correct the subject so that it is facing the camera in a normal position, thereby generating image information (orthogonalized image information) in which the captured image information has been orthogonalized. Orthogonalization process is a process in which the optical axis of the lens of the image capture device is aligned with the perpendicular direction of the plane of the subject, thereby transforming the image information so that it appears the same as if it were captured from a sufficiently distant location, and examples of this include keystone correction process. Keystone correction process is a correction process performed so that objects shown in the captured image information, such as shelves on a display shelf, are horizontal, and the products displayed there and the product tags (such as price tags) for those products are vertical.

[0036] The image information alignment processor 202 receives input specifying four vertices in the captured image information, and performs keystone correction processing using each of the vertices. The four vertices to be specified may be any four points in the captured image information. For example, if the subject is a display shelf, they may be the four vertices of the shelf levels on the display shelf, or the four vertices of the shelf positions on the display shelf. They may also be the four vertices of a group of two or three shelf levels. Any four points may be specified as the four vertices. The image information alignment processor 202 can use various known processing methods.

[0037] The object detection processing unit 203 executes a process for detecting an object appearing in photographed image information (including orthogonalized photographed image information; the same applies hereinafter in this specification) stored in the image information storage unit 201. The object detection process is preferably performed using so-called deep learning. In this case, the photographed image information is input to a learning model in which the weighting coefficients between neurons in each layer of a neural network consisting of multiple intermediate layers are optimized, and the area of ​​the object appearing in the input photographed image information (or its image information) and the type (class) of the object are output as output values.

[0038] The learning model is trained by using annotation data representing specific objects that may appear in captured image information, and executing a specific learning process using the annotation data. For example, objects that may appear in captured image information of a display shelf, such as products, product tags (price tags), and shelf components, are used as annotation data, and the learning model is trained by executing a specific learning process using the annotation data. For example, data associating product image information with the product's identity or its product name or product code (e.g., JAN code) is used as annotation data, and a learning process for the learning model is executed. Product image information is image information in the annotation data, and is a class indicating that the product is a product, or that the product is a product and its product identification information, such as its product name or product code, is a "product type." Similarly, product tag image information is image information in the annotation data, and is a class indicating that the product is a product tag, or that the product tag and its product name or product identification information are "product tags." Furthermore, shelf component image information is image information in the annotation data, and is a class indicating that the product is a shelf component or the name of a shelf component, such as "shelf component." The types (classes) of objects used as annotation data by the object detection processing unit 203 include items that are generally displayed on display shelves, such as merchandise, product tags, and shelf components, but are not limited to these, and the processing unit may be trained to detect any object.

[0039] The individual discrimination processing unit 204 determines that the object types (classes) detected from the captured image information by the object detection processing unit 203 as types other than a predetermined class are predetermined objects or candidates for predetermined objects. For example, the individual discrimination processing unit 204 determines that the object types detected as types other than merchandise, product tags, or shelf components are POPs or POP candidates. Furthermore, if an object detected as a type (class) of merchandise, product tags, or shelf components satisfies a predetermined condition, the individual discrimination processing unit 204 determines that the object type (class) is a POP or a POP candidate.

[0040] For example, for an object detected as a type of commodity, it is determined whether or not there is a product-specific advertisement (POP) attached to that area, and if it is determined that there is a POP, the type of object in that area is determined to be a POP or a POP candidate. In this case, there are two methods: one is to detect a POP or a POP candidate attached to a displayed product from the area of ​​an object detected as a type of commodity, and the other is to detect a POP or a POP candidate by comparing the image information of the area of ​​an object detected as a type of commodity with the image information of a product without a POP.

[0041] In the former method, a large amount of image information including POPs attached to displayed products is collected in advance, and the POPs are annotated for each image information including the POPs. A learning model that has been trained in advance using image information of normally displayed products without POPs as annotation data is then additionally trained using the image information annotated with the POPs to generate a new learning model. Image information of the area of ​​an object detected as a type of product is input to this new learning model, and during product recognition processing, in addition to outputting product identification information, it can also output whether or not the displayed product in the image information has a POP attached.

[0042] In the latter method, after performing product recognition processing on image information of the area of ​​an object detected as a commodity, standard image information (e.g., specimen information) of the recognized commodity without a POP is compared with the image information of the object area, and the similarity is compared. If the similarity is equal to or greater than a predetermined value, it is determined that the object detected as a commodity does not have a POP, and if the similarity is less than the predetermined value, it is determined that the object detected as a commodity has a POP or a POP candidate.

[0043] In the above-mentioned method, a more accurate determination can be made by comparing image information of an object area detected as a commodity with standard image information of the commodity recognized by the commodity recognition process without a POP, and determining whether or not it is a POP based on the location of the different area. For example, if it is determined that the different area between the two pieces of image information is localized in a specific position, such as near the top, shoulder, or tail of the commodity, it can be determined that it is a POP or a POP candidate.

[0044] Various known processing methods can be used for product recognition processing. For example, a process for identifying product identification information can be performed using deep learning on image information of the object region. In this case, the image information of the object region can be input to a learning model in which the weighting coefficients between neurons in each layer of a neural network consisting of multiple intermediate layers are optimized, and the product identification information can be identified based on the output value. The learning model can be one in which product identification information is assigned as ground truth data to various image information depicting products. In addition, product sample information (image information of the product taken from one or more directions or information such as features based thereon) can be compared with image information of the object region, and the product identification information can be identified using the similarity. Product recognition processing is not limited to these.

[0045] Furthermore, for objects detected as being of a product tag type, the image information of the object's area is determined to meet the requirements of a product tag. For example, OCR recognition processing is performed on the area of ​​an object detected as being of a product tag type, and if it contains characters typically found in product tags, such as a price or product name, it is determined to be a product tag. If not, it is determined to be a POP or a POP candidate. Alternatively, if the results of the OCR recognition processing include predetermined words typically used in POP, such as "member discount," "tester," "sample," or "sample," it may be determined to be a POP or a POP candidate. Furthermore, if handwritten characters are included, it may be determined to be a POP or a POP candidate.

[0046] For objects detected as being classified as shelf components, the system determines whether the image information in the object's area satisfies the requirements for a shelf component. For example, the system performs OCR recognition on the area of ​​an object detected as being classified as a shelf component, and if it contains a specific word, such as a member discount or other word typically used in POPs, it determines the object as a POP or a POP candidate; otherwise, it determines the object as a shelf component.

[0047] When the individual discrimination processing unit 204 determines that the type (class) of image information of an object area is a candidate for a predetermined object, the image information of the object area may be displayed on the display device 72 of the management terminal 2 or the like, and an input as to whether it is an object or not may be accepted, and if the input is an object, the type (class) of the image information of the object area may be set to the object. For example, when the individual discrimination processing unit 204 determines that the type (class) of image information of an object area is a candidate for a POP, the image information of the object area may be displayed on the display device 72 of the management terminal 2 or the like, and an input as to whether it is a POP or not may be accepted, and if the input is a POP, the type (class) of the image information of the object area may be set to the POP.

[0048] The model processing unit 21 has a classification processing unit 210, a tentative identification information processing unit 211, a model generation processing unit 212, a sample information reception processing unit 213, a recognition processing unit 214, and an output processing unit 215. An example of a conceptual diagram of the processing of the model processing unit 21 is shown in FIGS.

[0049] The classification processing unit 210 classifies the object image information of the object identified by the detection processing unit 20 using clustering processing or the like. For example, image information (POP image information) of an area of ​​an object identified as a POP, which is captured in image information of a display shelf, is classified using clustering processing or the like. If the object is a product displayed on a display shelf, a product tag (price card), a shelf, or a POP affixed, attached, or installed above or to the side of a shelf, the object image information is the POP image information. In this case, the image information may be cut out as a rectangle including the POP area, or may be cut out in any shape corresponding to the outer shape of the POP. Furthermore, cutting out the image information may involve physical cutting out, or it may involve identifying the coordinate information of the area and specifying the range.

[0050] Identification information (object identification information) of an object shown in the object image information, for example, if the object is a POP, a code such as information identifying the POP (POP identification information) does not need to be known.

[0051] The classification processing unit 210 performs clustering processing on the plurality of object image information using a known method to classify the image information into a plurality of groups. For example, if the object is a POP, the classification processing unit 210 performs clustering processing on the plurality of POP image information using a known method to classify the POP image information into a plurality of groups.

[0052] Although the classification processing unit 210 classifies the object image information using clustering processing, the object image information may be classified using a method other than clustering processing.

[0053] The temporary identification information processing unit 211 associates temporary identification information (temporary identification information) with each group of object image information classified by the classification processing unit 210. When the object is a POP, temporary identification information (temporary identification information) is associated with each group of POP image information classified by the classification processing unit 210. The temporary identification information may be automatically generated, or input temporary identification information may be associated by a predetermined operation. The temporary identification information associated with the object image information classified into each group is then associated as a label to form a data set and used as learning data.

[0054] The model generation processing unit 212 performs machine learning using the learning data with a known method to generate a learning model. That is, since temporary identification information is associated with each classified group of object image information, the learning model is generated using learning data that uses a dataset in which the object image information and its label are used as temporary identification information. In the case where the object is a POP, since temporary identification information is associated with each group of classified POP image information, the learning model is generated using learning data that uses a dataset in which the POP image information and its label are used as temporary identification information.

[0055] The specimen information reception processing unit 213 receives input of image information to be input to the model generation processing unit 212. The image information received as input here is object image information for which object identification information for identifying the object is known. For example, it is POP image information that may be attached to a display shelf, and is image information for which the corresponding POP identification information (POP identification information) is known. The image information input here will be referred to as specimen information.

[0056] The recognition processing unit 214 inputs the image information (sample information) received by the sample information receiving processing unit 213 into the learning model generated by the model generation processing unit 212, and outputs tentative identification information. In this case, when sample information is input to a learning model in which the weighting coefficients between neurons in each layer of a neural network consisting of multiple intermediate layers are optimized, tentative identification information is output as an output value. When multiple output values ​​are output together with their probabilities, the output value with the highest probabilities is adopted.

[0057] The output processing unit 215 links the tentative identification information, which is the output value output by the recognition processing unit 214, with the object identification information of the sample information used for its input. Then, the output processing unit 215 outputs the object identification information linked to the tentative identification information output by the recognition processing unit 214 as the recognition result of the recognition processing unit 214.

[0058] In addition, instead of outputting object identification information linked to the provisional identification information output by the recognition processing unit 214, the output processing unit 215 may replace the object identification information with the provisional identification information associated with the group output by the classification processing unit 210, and cause the model generation processing unit 212 to perform the machine learning learning process again.

[0059] Furthermore, if the accuracy of the tentative identification information output by the recognition processing unit 214 is equal to or less than a predetermined threshold, the processing of the output processing unit 215 does not need to be executed.

[0060] By performing the above-described processing, a learning model to be used in machine learning can be efficiently generated. [Example]

[0061] Next, an example of a processing process using the information processing system 1 of the present invention will be described with reference to the flowcharts of Figures 9 to 11. In the following explanation, a case will be described in which captured image information of a store's display shelves is input from the image information input terminal 3, a POP area is detected as an object from the captured image information, and a machine learning learning model for identifying POPs is generated using the image information of the detected POP area.

[0062] First, the detection process in the detection processing unit 20 will be described (S100).

[0063] Captured image information of a store's display shelves is input from the image information input terminal 3, and the input is accepted by the image information input acceptance processing unit 200 of the management terminal 2 (S110). An example of the input captured image information is shown in FIG. 5(a). Also accepted are input of the date and time of capture, store identification information, display shelf identification information, and image information identification information of the captured image information. Then, the image information input acceptance processing unit 200 associates the accepted input captured image information, date and time of capture, store identification information, store identification information, and image information identification information of the captured image information, and stores them in the image information storage unit 201.

[0064] When a predetermined operation input is received at the management terminal 2, the image information alignment processing unit 202 extracts the captured image information stored in the image information storage unit 201, receives input of four points for performing alignment processing such as trapezoid correction processing, for example, the four corners of a display shelf, and executes the alignment processing (S120).

[0065] Then, by receiving a predetermined operation input at the management terminal 2, the object detection processing unit 203 executes an object detection process on the normalized image information (S130) and classifies the type of object appearing in the normalized image information (S140).

[0066] FIG. 12 shows a schematic example of the results of executing object detection processing from upright image information. In FIG. 12, "a" represents the type of product, "b" represents product tags, "c" represents shelf components, and "d" represents the type of object detected as anything other than "a" to "c." Note that the type of object detected from the image information only needs to be associated with each region and does not need to be written on each image. Furthermore, the region of the detected object may be cut out from the upright image information, but it does not need to be actually cut out as long as the image information of that region is in a processable state. Note that cutting out image information of a region includes not only actually cutting out image information of that region from the upright image information, but also including the case where the region is specified using coordinate information or the like so that the image information of that region is in a processable state without being cut out.

[0067] The individual discrimination processing unit 204 discriminates and cuts out the areas of objects detected by the object detection processing unit 203, that is, areas of objects of types other than products, product tags, and shelf components (areas of objects "d"), as POP areas (S150, S160).

[0068] In addition, the individual discrimination processing unit 204 determines whether the area of ​​the object detected by the object detection processing unit 203 (S150), which is detected as a type of commodity, satisfies predetermined conditions, and determines whether it is a POP area (S170).

[0069] For example, for an area of ​​an object detected as a type of commodity, it is determined whether or not the area has a product-specific advertisement (POP) attached, and if it is determined that a POP is present, the type of the area is determined to be a POP. For an area of ​​an object detected as a type of commodity, a process is executed to detect POPs attached to displayed products, or a POP is detected by comparing image information of the area of ​​an object detected as a type of commodity with image information of products without POPs (product sample image information).

[0070] Similarly, for an object area detected as a product tag, the system determines whether the image information of the object area satisfies the requirements for a product tag, and for an object detected as a shelf member, the system determines whether the image information of the object area satisfies the requirements for a shelf member.

[0071] As described above, for the area of ​​an object detected as a type of commodity, product tag, or shelf component (S150), it is determined whether the predetermined condition for each object is met (S170), and if the condition is met, all or part of that area is determined to be a POP (S180). An example of this process is shown in Fig. 13. Furthermore, Fig. 14 shows an example of each area after determining that the area of ​​an object detected as a type of commodity, product tag, or shelf component that meets the predetermined condition is determined to be a POP area, based on the object detection results of Fig. 12.

[0072] POP areas in image information can be detected by performing the above-described processing in the detection processing unit 20. This makes it possible to detect image information of various types of POP areas from photographed image information of various display shelves, and to obtain image information that will be used to generate a learning model in the model processing unit 21, which will be described later.

[0073] Next, a process of generating a learning model in the model processing unit 21 using image information of the POP detected by the detection processing unit 20 will be described (S200).

[0074] 15 shows an example of POP image information detected by the detection processing unit 20. At this point, only the POP image information, which is image information of the object, has been detected from the captured image information, so the POP identification information (object identification information) for identifying the POP, which is the object, does not need to be known.

[0075] Then, the classification processing unit 210 performs a known clustering process or the like using the POP image information, which is the object image information, to group similar image information together and classify them into a plurality of groups (S210). This state is shown schematically in FIG. 16.

[0076] The temporary identification information processing unit 211 associates temporary identification information with each group classified by the classification processing unit 210 (S220). For example, arbitrary temporary identification information such as "a123," "b456," "c789," "d012," and "e345" is associated with each group. This state is schematically shown in FIG. 17. In this way, temporary identification information is associated as a label with the object image information of each group. A data set using the POP image information that constitutes this object image information and the labels of the temporary identification information is used as learning data.

[0077] Once the classification processing unit 210 has associated each group with provisional identification information, the model generation processing unit 212 performs machine learning using a known method on the training data (S230) to generate a training model (S240). This is schematically shown in FIG. 18. Note that while FIG. 18 shows a case where one POP image information and provisional identification information are used as the training data set for each group, in practice, it is preferable to perform machine learning by associating each POP image information and provisional identification information in each group as the training data set. This is because, even within the same group, not only are there images that are completely identical, but there are also images that differ to some extent, and by associating these with the provisional identification information for that group and performing machine learning on them, it is possible to improve accuracy.

[0078] Next, the specimen information reception processing unit 213 receives input of specimen information, which is image information for which POP identification information serving as object identification information has been determined, and inputs the received input specimen information to the learning model generated in S240 (S250). The recognition processing unit 214 receives input of specimen information and outputs an output result based on machine learning using the learning model. As this output result, provisional identification information is output (S260). This is schematically shown in FIG. 19.

[0079] The object identification information of the sample information, such as POP identification information, is known in advance. Therefore, the tentative identification information output by the recognition processing unit 214 is considered to correspond to the object identification information of the sample information input to the learning model. Therefore, the output processing unit 215 links the output tentative identification information with the object identification information of the input sample information (S270). This is schematically shown in FIG. 20.

[0080] This linking allows the provisional identification information to be associated with the POP identification information, which is the object identification information. By storing this association, the provisional identification information output by the recognition processing unit 214 can be replaced with the corresponding object identification information and output. When image information to be subjected to identification processing is input into the learning model, the object identification information can be output as the output value of the learning model.

[0081] By using the model processing described above, it is possible to generate a learning model for machine learning of an object. For example, POPs can be detected from image information of a display shelf, and the detected POPs can be used to generate a learning model for machine learning to detect POPs from input image information. [Example]

[0082] Another embodiment of the detection processing in the detection processing unit 20 will be described. The detection processing unit 20 of this embodiment further includes an area detection processing unit 205. An example of the configuration of the detection processing unit 20 in this embodiment is shown in Fig. 21, and an example of the processing process of the detection processing unit 20 is shown in the flowchart of Fig. 22. An example of an outline of the processing is shown in Fig. 23, and an example of individual discrimination processing is shown in Fig. 24.

[0083] The area detection processing unit 205 detects the product display area where products are displayed, the product tag area where product tags are placed, and the upper shelf area which is the area above the shelf, from the image information (including normalized image information) stored in the image information storage unit 201. To detect the product display area, product tag area, and upper shelf area, the operator of the management terminal 2 may manually specify the product display area, product tag area, and upper shelf area, which the area detection processing unit 205 may accept, or the product display area, product tag area, and upper shelf area may be automatically detected from the second time onwards based on the information on the product display area, product tag area, and upper shelf area that was manually input the first time. The area detection processing unit 205 may also detect areas other than these and use them as processing targets.

[0084] Figure 25 shows a schematic diagram of the state in which input specifying the product display area, product tag area, and shelf top area has been received for image information that has been normalized from image information captured of a display shelf on which products are displayed.

[0085] When detecting the product display area, product tag area, and shelf top area, the area detection processing unit 205 may use deep learning to identify the product display area, product tag area, and shelf top area. In this case, the above-mentioned orthogonalized image information may be input to a learning model in which the weighting coefficients between neurons in each layer of a neural network consisting of multiple intermediate layers are optimized, and the product display area, product tag area, and shelf top area may be detected based on the output values. Furthermore, the learning model may use various image information to which the product display area, product tag area, and shelf top area are assigned as correct answer data.

[0086] The area detection processing unit 205 may detect areas using various known methods in addition to using deep learning as described above.

[0087] The area detection processing unit 205 may actually extract the detected area as image information, or may virtually extract the area by specifying the range of the area without actually extracting it as image information. In this case, the range of the area can be specified by coordinates.

[0088] The object detection processing unit 203 in this embodiment may perform object detection processing from the range of the upright image information, as in the detection processing unit 20 in Example 1, or may perform object detection processing for each area detected by the area detection processing unit 205.

[0089] The individual discrimination processing unit 204 in this embodiment discriminates the type of each object detected by the object detection processing unit 203 using the type (class) of the object and the area in which the object is located. The discrimination is performed according to predetermined individual discrimination conditions, such as the discrimination tables shown in Figs. 23 and 24.

[0090] "a1" to "a4", "b1" to "b10", and "d1" to "d5" in FIG. 23 correspond to "a1" to "a4", "b1" to "b10", and "d1" to "d5" in FIG.

[0091] Next, an example of the processing process of the detection processing unit 20 in this embodiment will be described with reference to the flowchart of FIG.

[0092] Captured image information of store display shelves is input from the image information input terminal 3, and the input is accepted by the image information input acceptance processing unit 200 of the management terminal 2 (S300). The image information input acceptance processing unit 200 also accepts input of the image capture date and time, store identification information, display shelf identification information, and image information identification information of the captured image information. The image information input acceptance processing unit 200 then associates the accepted input image information, image capture date and time, store identification information, display shelf identification information, and image information identification information of the captured image information, and stores them in the image information storage unit 201.

[0093] When a predetermined operation input is received at the management terminal 2, the image information alignment processing unit 202 extracts the captured image information stored in the image information storage unit 201, receives input of any four vertices that are used for alignment processing such as trapezoid correction processing, for example, the four corners of a display shelf, and executes the alignment processing (S310).

[0094] Then, by receiving a predetermined operation input at the management terminal 2, the area detection processing unit 205 performs area detection processing on the aligned image information, detects the product display area, product tag area, and shelf top area (S320), and determines the range of each area.

[0095] Furthermore, by receiving a predetermined operation input at the management terminal 2, the object detection processing unit 203 executes an object detection process on the normalized image information (S330) and classifies the type of object appearing in the image information (S340). An example of the result of executing the object detection process on the image information is classified as shown in FIG. 12, similar to the first embodiment.

[0096] Then, for each object in the image information detected by the object detection processing unit 203, the individual discrimination processing unit 204 classifies the detected object in accordance with predetermined discrimination conditions using the type of the object and the area in which the object is located (the area determined in S320).

[0097] Furthermore, the individual discrimination processing unit 204 determines whether a predetermined condition is satisfied (S350) using the type of object in the image information detected by the object detection processing unit 203 and the area where the object is located, and determines whether the type of object is a POP or something other than a POP (such as merchandise, product tags, or shelf components) (S360, S370). The area where each object is located may be determined by any appropriate method, for example, by determining whether the center of gravity of the object's area is included in the area of ​​merchandise, product tags, or the top of the shelf.

[0098] For example, if an object in a product display area is classified as a commodity, the individual discrimination processing unit 204 determines whether or not a commodity-adjacent POP is present in the area of ​​the object. For example, as in the first embodiment, discrimination is made by detecting a POP attached to a displayed commodity from the area of ​​the object detected as a commodity, or by comparing image information of the area of ​​the object detected as a commodity with image information of a commodity-free POP. If a commodity-adjacent POP is detected, the area is discriminated as a POP (a2), and other areas are discriminated as commodities (a1). Furthermore, if a commodity-adjacent POP is not detected, OCR recognition processing is performed on the area to determine whether it contains specific words used in POPs, such as "tester," "sample," or "sample." If it does not, the object in the area is discriminated as a commodity (a1). If it does, the object in the area is discriminated as a POP, such as a tester or sample (a3).

[0099] Furthermore, if the type of an object in the product display area is a product tag, the individual discrimination processing unit 204 determines that the object in that area is a POP (d1).

[0100] Furthermore, if the type of an object in the product display area is a shelf member, the individual discrimination processing unit 204 determines that the object in that area is a shelf member (c).

[0101] Furthermore, if an object is in the product display area and the type of the object is other than products, product tags, or shelf members, it is determined to be a POP (d2).

[0102] If the object is in the product tag area and the type of the object is a product, the individual discrimination processing unit 204 determines that the object is a POP such as a tester or sample (d4).

[0103] Furthermore, if an object in the product tag area is a product tag, the object is subjected to OCR recognition processing. If handwriting is detected in the area, the object is determined to be a POP (retail POP) created by the store (retailer) (b4). If the OCR processing of the object area determines that no handwriting is detected in the area, the object is determined to contain a price and product name indication, and if it contains a word indicating a condition such as "member," the object is determined to be a product tag indicating a set product indication (b2). On the other hand, if the object is determined to contain a price and product name indication but does not contain a word indicating a condition such as "member," the object is determined to be a product tag (b1). If the object is determined not to contain the price and product name indication and contains a point indication, the object is determined to be a "point indication" POP (b3). Furthermore, if there is no point indication, the object is determined to be a POP (manufacturer POP) created by the product manufacturer (b9). Note that while the case where manufacturer POP and retail POP are distinguished and identified here, it may simply be identified as a POP.

[0104] Furthermore, if the type of an object in the product tag area is a shelf member, the individual discrimination processing unit 204 determines that the object in that area is a shelf member (c).

[0105] Furthermore, if an object is in the product tag area and the type of the object is other than products, product tags, or shelf components, OCR recognition processing is performed on that area. It is then determined whether or not it contains specific words used in POP, such as "tester," "sample," or "sample." If it does, it is determined to be a POP such as a tester or sample (d4). On the other hand, if it does not contain specific words used in POP, such as "tester," "sample," or "sample," it is determined to be a POP such as a promotional item (d3).

[0106] If the object is in the upper shelf area and the type of the object is a commodity, the individual discrimination processing unit 204 discriminates that the object is a stock commodity (a4).

[0107] Furthermore, if an object in the upper shelf area is a product tag, OCR recognition processing is performed on the object area. If handwriting is detected in the area, the object is determined to be a retail POP (b8). If the OCR processing of the object area determines that there is no handwriting in the area, and if it is determined that the area contains a price and product name display and a word indicating a condition such as "member," the object is determined to be a product tag indicating a set product display (b6). On the other hand, if it is determined that the object contains a price and product name display but does not contain a word indicating a condition such as "member," the object is determined to be a product tag (b5). If it is determined that the price and product name display are not included and there is a point display, the object is determined to be a "point display" POP (b7). Furthermore, if there is no point display in the above case, the object is determined to be a manufacturer POP (b10).

[0108] Furthermore, if the object is in the upper shelf area and the type of the object is a shelf member, the individual discrimination processing unit 204 determines that the object in that area is a shelf member (c).

[0109] Furthermore, if an object is in the upper shelf area and the type of the object is other than merchandise, merchandise tags, or shelf components, it is determined to be a POP such as a large promotional item (d5).

[0110] As described above, the individual discrimination processing unit 204 can use the area detected by the area detection processing unit 205 and the type of object detected by the object detection processing unit 203 to determine whether the object is a POP or not, and detect the area of ​​the POP in the image information. [Example]

[0111] As a modification of the processing of Examples 1 and 2, the processing shown in the flowchart of FIG. 26 may be performed. In this case, if the type of object detected by the object detection processing unit 203 is other than merchandise, product tags, or shelf components, the individual discrimination processing unit 204 determines that the object is a POP (S460). Then, for other objects, processing similar to the processing of S350 to S370 of FIG. 22 may be performed. That is, using the type of object in the image information detected by the object detection processing unit 203 and the area where the object is located, it is determined whether a predetermined condition is satisfied (S470), and it is determined whether the object is a POP or something other than a POP (such as merchandise, product tags, or shelf components) (S480, S490). Note that the processing from S400 to S440 may be the same as the processing from S300 to S340. [Example]

[0112] In addition to the processing of the detection processing unit 20 in Examples 1 to 3, modified examples of the model processing unit 21 will be described. The classification processing performed by the classification processing unit 210 may be less accurate than classification processing performed by a human. As a result, errors may be mixed in. Therefore, a cleansing processing unit 216 may be provided that performs a cleansing processing on the image information (target image information) of each group classified by the classification processing unit 210, and removes target image information that may contain errors from the image information classified into the groups, thereby reducing the amount of target image information to be used as learning data for the model generation processing unit 212. An example of the configuration of the model processing unit 21 in this case is shown in FIG. 27.

[0113] That is, when the classification processing unit 210 classifies the object image information into groups, index values ​​such as feature amounts that indicate the characteristics of the image information in that group are calculated. Then, in that group, the information distance from a reference value such as the median or average value of the index values ​​for each object image information is calculated (a value calculated using a predetermined formula to calculate the deviation from the reference value to the index value for each image information), and object image information whose information distance deviates by a certain value or a certain ratio or more may be excluded from the group.

[0114] In addition to using information distance for the cleansing process, similarity of image information may also be used. In this case, the similarity between image information in a group is quantified, and if the number of images with a calculated similarity above a certain level is equal to or greater than a predetermined value, the image information is retained as object image information, and if the number is less than the predetermined value, the image information is excluded from the object image information.

[0115] For example, if a group contains five pieces of target image information (image information 1 to image information 5), the similarity between each of the five pieces of image information and the other pieces of image information is calculated. That is, the similarity between image information 1 and image information 5 is calculated, between image information 2 and image information 5, between image information 2 and image information 5, between image information 3 and image information 5, between image information 4 and image information 5, and between image information 4 and image information 5. In this way, the similarity between each piece of image information can be calculated.

[0116] Assume that for image information 1, the similarity with image information 2 is 0.9, the similarity with image information 3 is 0.95, the similarity with image information 4 is 0.3, and the similarity with image information 5 is 0.8; for image information 2, the similarity with image information 3 is 0.8, the similarity with image information 4 is 0.2, and the similarity with image information 5 is 0.6; for image information 3, the similarity with image information 4 is 0.4 and the similarity with image information 5 is 0.9; and for image information 4, the similarity with image information 5 is 0.5.

[0117] Here, if the similarity of the reference image information is 0.75 and the number of pieces of image information below the reference level is three, then for image information 1, only one piece of image information 4 falls below the reference similarity level; for image information 2, two pieces of image information 4 and image information 5 fall below the reference similarity level; for image information 3, only one piece of image information 4 falls below the reference similarity level; for image information 4, four pieces of image information 1, 2, 3, and 5 fall below the reference similarity level; and for image information 5, two pieces of image information 2 and 4 fall below the reference similarity level.

[0118] Therefore, the cleansing processing unit 216 excludes image information 4, which has three or more pieces of image information below the standard similarity, from the group, and executes processing by the provisional identification information processing unit 211 with four pieces of object image information for the group: image information 1, image information 2, image information 3, and image information 5.

[0119] By excluding the deviated image information from the grouped object image information as in this embodiment, it is possible to narrow down the grouped object image information, which leads to an improvement in accuracy. [Example]

[0120] When extracting POP image information as object image information by cutting out a POP placed on a display shelf, etc., multiple similar POPs may be placed adjacent to each other in either the up, down, left, or right direction to attract customers' attention.

[0121] Therefore, after extracting POP image information as object image information from the image information of the display shelf, the classification processing unit 210 compares the similarity of the POP image information as adjacent object image information, and if it satisfies certain conditions, for example, if the similarity is equal to or greater than a predetermined threshold, it may determine that the objects are of the same type and classify the POP image information as object image information of the adjacent POPs into the same group. [Example]

[0122] When targeting POPs of products displayed on shelves, it is preferable to generate a learning model of POPs that may be displayed in advance. However, in this case, the target POPs may number from hundreds to thousands, or even tens of thousands. In this case, if a learning model is created that can recognize thousands to tens of thousands of POPs, the sample information input by the sample information reception processing unit 213 will be in the thousands to tens of thousands. Although the workload is reduced compared to conventional learning models, the workload is still significant.

[0123] On the other hand, it is known in marketing that generally, about 10% of product types tend to account for about 90% of the sales of that category. Therefore, POPs of products with high sales may be used as the specimen information to be input to the specimen information reception processing unit 213. In this case, POPs of products provided by marketing companies or the like may be used.

[0124] This makes it possible to recognize POPs for approximately 90% of products based on sales, which is sufficient for practical purposes as output of object identification information. Note that when a POP for a product with low sales is input to the learning model, the object identification information is not linked to the provisional identification information, and therefore the output processing unit 215 outputs the provisional identification information as is. When the output processing unit 215 outputs provisional identification information that does not associate the provisional identification information with the object identification information, it may display a predetermined message such as "no correspondence information."

[0125] Furthermore, if it is sufficient to recognize only the POPs of products of a specific company or other organization, image information of the company's products may be input as sample information to the sample information reception processing unit 213. This allows object identification information to be output for the POPs of the company's products. On the other hand, when a POP of a product of a company other than the specific company is input, as in the above case, the object identification information and the provisional identification information are not linked, so the output processing unit 215 will output the provisional identification information as is. When the output processing unit 215 outputs provisional identification information that does not associate the provisional identification information with the object identification information, it may display a predetermined message such as "no correspondence information." [Example]

[0126] For the purpose of improving accuracy, the learning model may be reconstructed again for the learning model generated as in Examples 1 to 6. In this case, new sample information may be input, and new object identification information may be linked to tentative identification information to which no object identification information has been linked, or the classification processing unit 210 may perform classification processing using object image information of a new object so that the new object can be recognized. [Example]

[0127] In the above-described first to seventh embodiments, a learning model is generated for identifying POPs displayed on a display shelf from image information of the display shelf, but the present invention can also be applied to other cases. In particular, the present invention is useful for automating the identification of various types of objects from image information.

[0128] As an example, the target may be an animal. For example, this method can be applied to cases where multiple types of animals, such as sea lions and fur seals, live in large numbers in places where people cannot approach, and the types and populations of these animals need to be efficiently identified. In this case, the habitat is photographed from the air using a drone or other device, and object image information for each individual is extracted. The object image information for each individual is then subjected to image classification processing by the classification processing unit 210 and grouped. Then, the temporary identification information processing unit 211 generates learning data by associating temporary identification information with each group, and the model generation processing unit 212 performs machine learning to generate a learning model. The specimen information reception processing unit 213 then accepts input of specimen information by type, and the recognition processing unit 214 inputs the accepted specimen information as input values ​​into the learning model, outputting temporary identification information. This links the temporary identification information to the object identification information (such as the scientific name of the animal) corresponding to the input specimen information, enabling the output processing unit 215 to output the object identification information.

[0129] As in the above, the target object may be a bird or a plant. Even in the case of a bird or a plant, processing can be performed in the same way as for the above-mentioned animal, and processing can be performed by replacing "animal" with "bird" or "plant." [Industrial Applicability]

[0130] By using the information processing system 1 of the present invention, it is possible to efficiently generate a learning model used in machine learning. [Explanation of symbols]

[0131] 1: Information processing system 2: Management terminal 3: Image information input terminal 20: Detection processing unit 21: Model processing section 70: Arithmetic device 71:Storage device 72:Display device 73: Input device 74:Communication equipment 200: Image information input reception processing unit 201: Image information storage unit 202: Image information alignment processing unit 203: Object detection processing unit 204: Individual discrimination processing unit 205: Area detection processing unit 210: Classification processing unit 211: Temporary identification information processing unit 212: Model generation processing unit 213: Specimen information reception processing unit 214: Recognition processing unit 215: Output processing unit 216: Cleansing processing section

Claims

1. An information processing system that performs processing related to a learning model used in machine learning, a detection processing unit that detects POPs from image information; a model processing unit that generates a learning model for machine learning using the detected POP, The model processing unit a classification processing unit that classifies image information of the detected POP; a temporary identification information processing unit that associates the classified groups with temporary identification information; a model generation processing unit that performs machine learning learning processing using learning data including image information of the POPs included in the group and the temporary identification information to generate a learning model; An information processing system comprising:

2. The model processing unit a cleansing processing unit that executes a cleansing process to exclude image information of POPs that may contain errors from the image information of POPs included in the classified groups; 2. The information processing system according to claim 1, further comprising:

3. Computer, a detection processing unit that detects POPs from image information; a model processing unit that generates a learning model for machine learning using the detected POPs; An information processing program that functions as The model processing unit a classification processing unit that classifies image information of the detected POP; a temporary identification information processing unit that associates the classified groups with temporary identification information; a model generation processing unit that performs machine learning learning processing using learning data including image information of the POPs included in the group and the temporary identification information to generate a learning model; An information processing program comprising:

Citation Information

Patent Citations

  • Data processing apparatus and method

    JP2022150552A

  • Information processing device, information processing method, and program

    WO2019064926A1

  • Learning dataset automatic generation system, server, and learning dataset automatic generation program

    JP2020095537A

  • Dataset generation apparatus, generation method, program, system, machine learning apparatus, object recognition apparatus, and picking system

    JP2022076296A