Image recognition device, its control program, and image recognition method

The dual-learning model approach in the image recognition device, utilizing both global and local models, addresses the challenge of recognizing diverse apple varieties by enhancing accuracy through region-specific recognition.

JP7742249B2Active Publication Date: 2025-09-19TOSHIBA TEC KK
View PDF 8 Cites 0 Cited by

Patent Information

Application Number
JP2021101730
Authority / Receiving Office
JP · JP
Patent Type
Patents
Current Assignee / Owner
Priority Date
2020-08-12
Filing Date
2021-06-18
Publication Date
2025-09-19
Estimated Expiration
2041-06-18

AI Technical Summary

Technical Problem

Existing image recognition systems struggle to accurately distinguish between various apple varieties, especially those sold only in specific regions, due to small differences in features captured in images.

Method used

An image recognition device employs a dual-learning model approach, using a global model for nationwide recognition and a region-specific local model to enhance accuracy, with the local model receiving priority in identifying region-exclusive products.

Benefits of technology

The system achieves high-accuracy product recognition by combining nationwide and region-specific models, effectively distinguishing products with small feature variations, thereby improving overall recognition performance.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure 0007742249000001
    Figure 0007742249000001
  • Figure 0007742249000002
    Figure 0007742249000002
  • Figure 0007742249000003
    Figure 0007742249000003
Patent Text Reader

Abstract

To enable accurate recognition even for a commodity having a small difference in a feature amount obtained from a captured image depending on a product type.SOLUTION: An image recognition device recognizes a commodity by deep-layer learning using a common first learning model using a captured image of the commodity as input. The image recognition device recognizes the commodity by deep-layer learning using a specific second learning model using the captured image of the commodity as input. The image recognition device identifies the commodity on the basis of the recognition result by the deep-layer learning using the first learning model and the recognition result by the deep-layer learning using the second learning model.SELECTED DRAWING: Figure 7
Need to check novelty before this filing date? Find Prior Art

Description

[Technical Field]

[0001] FIELD Embodiments of the present invention relate to an image recognition device, a control program for causing a computer to function as the device, and an image recognition method. [Background technology]

[0002] A technology that uses deep learning, such as convolutional neural networks, to recognize products shown in images captured by a camera is already known. Using this technology, an image recognition device has been developed that automatically recognizes products without barcodes, such as fresh produce and fruit. This type of image recognition device recognizes products by inputting images of the products to be recognized into a learning model. The learning model is a machine learning model. The learning model is created by a computer with AI (artificial intelligence) functionality extracting and modeling features from a large amount of image data of various products individually photographed.

[0003] For example, apples are sold nationwide, and there are many different varieties. Some apple varieties are sold nationwide, while others are sold only in specific regions. Furthermore, most apple varieties have small differences in the features obtained from photographed images. For this reason, when performing image recognition of apples using a learning model containing data on various apple varieties, it can be difficult to distinguish the varieties of apples sold only in specific regions. [Prior art documents] [Patent documents]

[0004] [Patent Document 1] Japanese Patent Application Laid-Open No. 2014-032539 Summary of the Invention [Problem to be solved by the invention]

[0005] The problem to be solved by the embodiments of the present invention is to provide an image recognition device that can accurately recognize a product from a captured image. [Means for solving the problem]

[0006] In one embodiment, the image recognition device includes a first recognition means, a second recognition means, and an identification means. The first recognition means receives a photographed image of a product as an input and recognizes the product through deep learning using a common first learning model. The second recognition means receives a photographed image of the product as an input and recognizes the product through deep learning using a unique second learning model. The identification means identifies the product based on the recognition result by the first recognition means and the recognition result by the second recognition means. [Brief explanation of the drawings]

[0007] [Figure 1] 1 is a block diagram showing a schematic configuration of an image recognition system including an image recognition device according to an embodiment. [Figure 2] FIG. 2 is a block diagram showing the main circuit configuration of a center server. [Figure 3] FIG. 2 is a block diagram showing the main circuit configuration of an edge server. [Figure 4] FIG. 2 is a block diagram showing the main circuit configuration of the image recognition device. [Figure 5] FIG. 2 is a block diagram showing the main circuit configuration of the POS terminal. [Figure 6] FIG. 1 is a sequence diagram of main data signals exchanged between a center server, an edge server, and an image recognition device. [Figure 7] 3 is a flowchart showing the main steps of an image recognition process executed by a processor of the image recognition device. [Figure 8] 10 is a flowchart showing the main steps of a process executed by a processor of an edge server. [Figure 9] 10 is a flowchart showing the main steps of a process executed by a processor of the center server. DETAILED DESCRIPTION OF THE INVENTION

[0008] Hereinafter, the embodiments will be described with reference to the drawings. This embodiment relates to an image recognition device used to recognize products such as fresh food and fruit that do not have barcodes attached, in each store of a retailer that operates a chain throughout the country.

[0009] 1 is a block diagram showing a schematic configuration of an image recognition system 100 including an image recognition device 30 according to an embodiment. The image recognition system 100 includes a center server 10, an edge server 20, the image recognition device 30, a POS (Point Of Sales) terminal 40, a first communication network 50, and a second communication network 60.

[0010] In the image recognition system 100, a plurality of edge servers 20 are connected to one center server 10 via a first communication network 50. In the image recognition system 100, a plurality of image recognition devices 30 are connected to each edge server 20 via a second communication network 60. In the image recognition system 100, a POS terminal 40 is connected to each image recognition device 30. Each image recognition device 30 and each POS terminal 40 are connected one-to-one via a wired or wireless communication method.

[0011] The first communication network 50 is a wide-area computer network. The second communication network 60 is a narrower area computer network than the first communication network 50. Both the first communication network 50 and the second communication network 60 can utilize well-known computer networks.

[0012] A set of an image recognition device 30 and a POS terminal 40 is provided for each store. The number of sets of an image recognition device 30 and a POS terminal 40 provided in one store is not particularly limited. It is possible that one store has only one set, or that one store has multiple sets.

[0013] An edge server 20 may be provided for each store, or one for multiple stores in the same region. The number of servers is not limited to one. It is sufficient that the edge server 20 is accessible to multiple stores in the same region. Alternatively, one edge server 20 may be provided for multiple adjacent regions. The number of servers is not limited to one. It is sufficient that the edge server 20 is accessible to stores in multiple adjacent regions. Incidentally, a region refers to, for example, an administrative division such as a city, town, or village, or to a region that groups together adjacent administrative divisions.

[0014] The center server 10 is provided in common to each store in a nationwide chain. The center server 10 may provide computer resources to each edge server 20 as cloud computing.

[0015] The center server 10 is a computer with AI functions. The center server 10 uses the AI ​​functions to create a global model 70 and also updates the global model 70. The global model 70 is a learning model used for image recognition of products such as fresh food and fruit. The global model 70 is a learning model common to each store. The global model 70 is a learning model commonly used for product recognition in each store. It is assumed that the global model 70 will be used in all stores, but it is not denied that there may be exceptional stores that do not use the global model 70. The global model 70 is an example of a first learning model.

[0016] Each edge server 20 is a computer with AI functions. Each edge server 20 uses its AI functions to create and update a local model 80. The local model 80 is a learning model used for image recognition of products such as fresh produce and fruit. The local model 80 is a specific learning model that is not common to a store or the area where the store is located. The local model 80 is used to recognize products in a limited number of stores. The local model 80 is a learning model that is not used in all stores but in a specific range of stores. The local model 80 is an example of a second learning model.

[0017] Each image recognition device 30 is a computer with AI functions. Each image recognition device 30 uses its AI functions to recognize the product depicted in the captured image. Each image recognition device 30 then outputs information about the recognized product to the corresponding POS terminal 40.

[0018] Each POS terminal 40 registers sales data of the product purchased by the consumer based on the product information recognized by the corresponding image recognition device 30. Furthermore, each POS terminal 40 performs processing to settle the commercial transaction with the consumer based on the registered sales data of the product.

[0019] 2 is a block diagram showing the main circuit configuration of the center server 10. The center server 10 includes a processor 11, a main memory 12, an auxiliary storage device 13, an accelerator 14, and a communication interface 15. The center server 10 connects the processor 11 to the main memory 12, the auxiliary storage device 13, the accelerator 14, and the communication interface 15 via a system bus 16. The system bus 16 includes an address bus, a data bus, etc. The center server 10 configures a computer by connecting the processor 11 to the main memory 12 and the auxiliary storage device 13 via the system bus 16.

[0020] The processor 11 corresponds to the central part of the computer. The processor 11 controls each part to realize various functions of the center server 10 in accordance with an operating system or an application program. The processor 11 is, for example, a CPU (Central Processing Unit).

[0021] The main memory 12 corresponds to the main storage portion of the computer. The main memory 12 includes a nonvolatile memory area and a volatile memory area. The main memory 12 stores an operating system or application programs in the nonvolatile memory area. The main memory 12 stores data required for the processor 11 to execute processes for controlling each part in the volatile memory area. The main memory 12 also uses the volatile memory area as a work area where data can be rewritten by the processor 11 as appropriate. The nonvolatile memory area is, for example, ROM (Read Only Memory). The volatile memory area is, for example, RAM (Random Access Memory).

[0022] The auxiliary storage device 13 corresponds to the auxiliary storage portion of the computer. As the auxiliary storage device 13, for example, well-known storage devices such as an EEPROM (Electric Erasable Programmable Read-Only Memory), an HDD (Hard Disc Drive), or an SSD (Solid State Drive) are used singly or in combination. The auxiliary storage device 13 stores data used by the processor 11 when performing various processes, data generated by the processes in the processor 11, etc. The auxiliary storage device 13 may also store application programs.

[0023] The accelerator 14 is a processing unit for recognizing images through deep learning using AI. Deep learning uses, for example, a convolutional neural network. The accelerator 14 can be, for example, a graphics processing unit (GPU) or a field programmable gate array (FPGA).

[0024] The communication interface 15 is responsible for data communication with each edge server 20 connected via the first communication network 50 .

[0025] 3 is a block diagram showing the main circuit configuration of the edge server 20. The edge server 20 includes a processor 21, a main memory 22, an auxiliary storage device 23, an accelerator 24, a first communication interface 25, and a second communication interface 26. The edge server 20 connects the processor 21 to the main memory 22, the auxiliary storage device 23, the accelerator 24, the first communication interface 25, and the second communication interface 26 via a system bus 27. The system bus 27 includes an address bus, a data bus, etc. The edge server 20 configures a computer by connecting the processor 21 to the main memory 22 and the auxiliary storage device 23 via the system bus 27.

[0026] The processor 21, main memory 22, auxiliary storage device 23, and accelerator 24 have the same basic functions as the processor 11, main memory 12, auxiliary storage device 13, and accelerator 14 of the center server 10. Therefore, a description thereof will be omitted here.

[0027] The first communication interface 25 is responsible for data communication with the center server 10 connected via the first communication network 50.

[0028] The second communication interface 26 controls data communication with each image recognition device 30 connected via the second communication network 60 .

[0029] 4 is a block diagram showing the main circuit configuration of an image recognition device 30. The image recognition device 30 includes a processor 31, a main memory 32, an auxiliary storage device 33, an accelerator 34, a device interface 35, a first communication interface 36, and a second communication interface 37. In the image recognition device 30, the processor 31 is connected to the main memory 32, the auxiliary storage device 33, the accelerator 34, the device interface 35, the first communication interface 36, and the second communication interface 37 via a system bus 38. The system bus 38 includes an address bus, a data bus, etc. The image recognition device 30 constitutes a computer by connecting the processor 31 to the main memory 32 and the auxiliary storage device 33 via the system bus 38.

[0030] The processor 31, main memory 32, auxiliary storage device 33, and accelerator 34 have the same basic functions as the processor 11, main memory 12, auxiliary storage device 13, and accelerator 14 of the center server 10. Therefore, a description thereof will be omitted here.

[0031] The device interface 35 is an interface with the imaging device 90. The imaging device 90 is a device that captures an image of a product to be recognized. As the imaging device 90, for example, a CCD camera using a CCD (Charge Coupled Device) as an imaging element is used.

[0032] The first communication interface 36 is responsible for data communication with the edge server 20 connected via the second communication network 60 .

[0033] The second communication interface 37 is responsible for data communication with the POS terminal 40 connected by wire or wirelessly.

[0034] 5 is a block diagram showing the main circuit configuration of the POS terminal 40. The POS terminal 40 includes a processor 41, a main memory 42, an auxiliary storage device 43, a communication interface 44, an input device 45, a display device 46, a printer 47, and a change machine interface 48. The POS terminal 40 connects the processor 41 with the main memory 42, the auxiliary storage device 43, the communication interface 44, the input device 45, the display device 46, the printer 47, and the change machine interface 48 via a system bus 49. The system bus 49 includes an address bus, a data bus, etc. The POS terminal 40 configures a computer by connecting the processor 41 with the main memory 42 and the auxiliary storage device 43 via the system bus 49.

[0035] The processor 41, main memory 42, and auxiliary storage device 43 have the same basic functions as the processor 11, main memory 12, and auxiliary storage device 13 of the center server 10. Therefore, a description thereof will be omitted here.

[0036] The communication interface 44 controls data communication with the image recognition device 30, which is connected via a wired or wireless connection, and also controls data communication with other computer devices, such as a store server (not shown).

[0037] The input device 45 is a device used to input data required for the POS terminal 40. The input device 45 is, for example, a keyboard, a touch panel sensor, a card reader, a code scanner, or the like.

[0038] The display device 46 is a device for displaying information to be presented to the operator or consumer of the POS terminal 40. The display device 46 is, for example, a liquid crystal display, an organic EL (Electro Luminescence) display, or the like.

[0039] The printer 47 is a printer for printing receipts. The change machine interface 48 is responsible for data communication with an automatic change machine (not shown).

[0040] 6 is a sequence diagram of main data signals exchanged between the center server 10, the edge servers 20, and the image recognition device 30. It is assumed that the center server 10 and each edge server 20 already have a learning model of the global model 70 or the local model 80.

[0041] First, the center server 10 distributes the global model 70 to each edge server 20 via the first communication network 50 at an arbitrary timing for distributing a learning model.

[0042] Upon receiving the global model 70 distributed from the center server 10, each edge server 20 stores the global model 70 in the auxiliary storage device 33. Furthermore, each edge server 20 outputs a query command Ca to each image recognition device 30 connected via the second communication network 60 to inquire whether it is ready to receive the learning model. The query command Ca is transmitted to each image recognition device 30 via the second communication network 60.

[0043] The image recognition device 30, which is ready to receive the learning model, outputs an acceptance response command Cb to the edge server 20. The acceptance response command Cb is transmitted to the edge server 20 through the second communication network 60.

[0044] Upon receiving the acceptance response command Cb, the edge server 20 transmits the global model 70 and the local model 80 via the second communication network 60 to the image recognition device 30 that sent the command.

[0045] Upon receiving the global model 70 and the local model 80 transmitted from the edge server 20, the image recognition device 30 stores the global model 70 and the local model 80 in the auxiliary storage device 33. By storing the global model 70 and the local model 80, the image recognition device 30 becomes able to recognize images of products.

[0046] The image recognition device 30, now capable of image recognition, executes the image recognition process described below. Then, the image recognition device 30 outputs training data Da obtained by the image recognition process to the edge server 20. The training data Da is transmitted to the edge server 20 via the second communication network 60.

[0047] The edge server 20 performs additional learning of the local model 80 based on the training data Da transmitted from each image recognition device 30 connected via the second communication network 60. The edge server 20 then outputs learning result data Db, which is the result of the additional learning performed on the local model 80, to the center server 10. The learning result data Db is transmitted to the center server 10 via the first communication network 50.

[0048] The center server 10 updates the global model 70 so as to aggregate the learning result data Db transmitted from each edge server 20 connected via the first communication network 50. Then, at an arbitrary learning model distribution timing, the center server 10 distributes the global model 70 updated by aggregating the learning result data Db to each edge server 20 via the first communication network 50. Thereafter, the center server 10, each edge server 20, and each image recognition device 30 repeat the same operations as above.

[0049] Therefore, the local model 80 possessed by each edge server 20 is updated by additional learning based on the image recognition results processed by one or more image recognition devices 30 connected to that edge server 20. The local model 80 differs for each limited store, that is, for each edge server 20, and each of the multiple local models 80 is updated by additional learning based on the image recognition results processed to recognize products in the limited stores. Meanwhile, the global model 70 possessed by the center server 10 is updated so as to aggregate the learning results of each local model 80 possessed by each edge server 20. The global model 70 is updated so as to aggregate each of the multiple learned local models 80. Such a global model 70 can be used to recognize products that cannot be recognized by each individual local model 80, among products in a store that uses each local model 80.

[0050] Next, the image recognition process executed by the image recognition device 30 will be described. 7 is a flowchart showing the main steps of the image recognition process executed by the processor 31 of the image recognition device 30. The processor 31 executes the image recognition process according to the steps shown in the flowchart of FIG.

[0051] There are no particular limitations on the method for installing the control program into the main memory 32 or the auxiliary storage device 33. The control program can be recorded on a removable recording medium, or can be distributed by communication via a network and installed in the main memory 32 or the auxiliary storage device 33. The recording medium can be in any form, such as a CD-ROM or memory card, as long as it can store the program and is readable by the device.

[0052] When the operator of POS terminal 40 receives a request from the consumer to pay for the purchased items, he / she declares the start of registration by operating input device 45. In response to this declaration, a start signal is output from POS terminal 40 to image recognition device 30. In response to this start signal, processor 31 of image recognition device 30 begins information processing according to the procedure shown in the flowchart of FIG.

[0053] First, the processor 31 activates the imaging device 90 to start an imaging operation in ACT 1. Then, the processor 31 waits for input of photographed image data of an image of a product in ACT 2.

[0054] The operator picks up each item that the consumer will purchase and holds it up to the lens of the imaging device 90. In this way, the imaging device 90 takes a picture of the item.

[0055] Processor 31 performs contour extraction processing and the like on the captured image data input via device interface 35 to determine whether a product has been photographed. If a product has been photographed, processor 31 determines YES in ACT2 and proceeds to ACT3. In ACT3, processor 31 analyzes the captured image data to confirm whether a barcode attached to the product is displayed in the captured image. If a barcode is displayed in the captured image, processor 31 determines YES in ACT3 and proceeds to ACT4. In ACT4, processor 31 executes a well-known barcode recognition process to read a barcoded code from the barcode image. Then, in ACT5, processor 31 outputs a code that is the recognition result of the barcode recognition process to POS terminal 40. Processor 31 then proceeds to ACT16. The processing from ACT16 onwards will be described later.

[0056] On the other hand, if a barcode is not shown in the captured image, the processor 31 determines NO in ACT3 and proceeds to ACT6. The processor 31 then activates the accelerator 34 in ACT6. The processor 31 then commands the accelerator 34 to execute a machine learning algorithm using the global model 70. In response to this command, the accelerator 34 inputs the image data of the captured image into the global model 70 stored in the auxiliary storage device 33, and executes deep learning using, for example, a convolutional neural network to recognize the product shown in the captured image.

[0057] The processor 31 acquires the recognition result A from the accelerator 34 as ACT7. The recognition result A is an item list of products for which the similarity between the feature amounts modeled by the global model 70 and the feature amounts obtained from the captured image is determined to be equal to or greater than a predetermined threshold. The item list is a list that is subdivided down to the product type. That is, the processor 31 stores the results of recognizing the products through deep learning using the global model 70 in the auxiliary storage device 33.

[0058] Next, the processor 31 issues a command (ACT8) to the accelerator 34 to execute a machine learning algorithm using the local model 80. In response to this command, the accelerator 34 inputs the image data of the captured image into the local model 80 stored in the auxiliary storage device 33, and executes deep learning using, for example, a convolutional neural network to recognize the product depicted in the captured image.

[0059] The processor 31 acquires the recognition result B from the accelerator 34 as ACT9. The recognition result B is an item list of products for which the similarity between the feature amount modeled by the local model 80 and the feature amount obtained from the captured image is determined to be equal to or greater than a predetermined threshold. The item list is a list that is subdivided down to the product type. That is, the processor 31 stores the results of recognizing the products through deep learning using the local model 80 in the auxiliary storage device 33.

[0060] In ACT10, the processor 31 makes a final determination between the recognition results A and B. For example, the processor 31 applies a predetermined weight to the similarity to the product item obtained as the recognition result B to increase the similarity. The processor 31 then compares the weighted similarity to the product item obtained as the recognition result B with the similarity to the product item obtained as the recognition result A. The processor 31 then determines, for example, the top three product items in descending order of similarity as candidate products.

[0061] By applying a predetermined weight to the similarity, if the similarity obtained as recognition result A and the similarity obtained as recognition result B are approximately equal, the similarity obtained as recognition result B will be greater due to the weighting process. Therefore, recognition result B will be given priority over recognition result A. Note that weighting may not be necessary. To identify a product, the processor 31 compares the recognition results using the global model 70 and the recognition results using the local model 80, both of which are stored in the auxiliary storage device 33. The processor 31 then identifies the product from either of the recognition results. The list of candidate products may include a mixture of recognition results using the global model 70 and recognition results using the local model 80. Note that if the similarities between the candidate products are approximately equal, either of the recognition results is adopted. By comparing the recognition results using the global model 70 and the recognition results using the local model 80, products that are only sold in a specific region can be recognized using the local model 80. Products that are sold in all stores can be recognized using the global model 70. Note that even if a product that is not normally sold in a specific region is sold as a special offer, it can be recognized using the global model 70.

[0062] In ACT12, the processor 31 outputs the final determination result to the POS terminal 40. As a result, the display device 46 of the POS terminal 40 displays a list of the first to third ranked candidate products obtained as the final determination result.

[0063] The operator of the POS terminal 40 checks whether the product held over the imaging device 90, i.e., the product purchased by the consumer, is included in the list of candidate products. If the product is included in the list, the operator performs an operation to select the product. This operation causes the POS terminal 40 to register sales data of the purchased product. If the purchased product is not included in the list of candidate products, the operator operates the input device 45 to register sales data of the purchased product. The POS terminal 40 transmits data of the product selected by the operator to the image recognition device 30.

[0064] After outputting the final judgment result, the processor 31 checks whether or not the final judgment result needs to be revised in ACT12. For example, the processor 31 determines whether or not a revision needs to be made based on the product data received from the POS terminal 40. If a candidate product ranked second or lower is selected in the POS terminal 40, the final judgment result needs to be revised. The processor 31 determines YES in ACT12 and proceeds to ACT13. The processor 31 corrects the final judgment result in ACT13. Specifically, the processor 31 makes the revision so that the selected candidate product becomes first on the list. The processor 31 then proceeds to ACT14.

[0065] On the other hand, if there is no correction to the final determination result, that is, if the top candidate product is selected, the processor 31 determines NO in ACT12 and proceeds to ACT14.

[0066] The processor 31 generates training data Da as ACT14. The training data Da is data in which a correct answer label is added to a photographed image of a product input via the device interface 35. The correct answer label is information that identifies the candidate product that ranked first in the final judgment result. In other words, if there is no correction to the final judgment result, the correct answer label is information about the product that was set to first place by the final judgment. If there is a correction to the final judgment result, the correct answer label is data about the product that was changed to first place as a result of the correction.

[0067] The processor 31 issues a command to transmit the generated training data Da as ACT 15. In response to this command, the training data Da is transmitted from the image recognition device 30 to the edge server 20, as shown in FIG. Note that the processor 31 does not need to generate training data Da every time. It is assumed that some stores may specially sell regional products that are not normally sold. In such cases, the local model 80 does not contain a corresponding product, and the product is identified from the recognition results of the global model 70. For products that are not normally sold and are unlikely to be sold in the future, it is assumed that the user does not wish to perform additional training on the local model 80. Therefore, for such products, transmission of the training data Da is prohibited. Specifically, when the POS terminal 40 registers sales data of purchased products, training prohibition information indicating that the training data Da should not be transmitted to the edge server 20 is stored in the referenced product file. When a product for which training prohibition information is stored is registered, the processor 31 transmits the training prohibition information to the image recognition device 30. When the transmission means of the image recognition device 30 detects information prohibiting transmission of the training data Da to the edge server 20, i.e., the training prohibition information, from the received data, the transmission means of the image recognition device 30 does not transmit the training data Da.

[0068] After transmitting the training data Da, processor 31 checks whether the end of registration has been declared in ACT 16. If the end of registration has not been declared, processor 31 returns to ACT 2. Processor 31 executes the processes from ACT 2 onward in the same manner as described above.

[0069] Therefore, every time the operator holds the consumer's purchased item over the lens of the imaging device 90, the same processes as those in ACT2 to ACT16 described above are repeatedly executed. Then, when the operator has finished registering all the sales data of the consumer's purchased items, he / she operates the input device 45 to declare the completion of registration.

[0070] When the processor 31 detects that the registration has been closed in the POS terminal 40, the processor 31 determines YES in ACT 16 and proceeds to ACT 17. In ACT 17, the processor 31 stops the imaging operation by the imaging device 90. With this, the processor 31 ends the information processing of the procedure shown in the flowchart of FIG.

[0071] Here, the processor 31 configures a first recognition means by executing the processes of ACT6 and ACT7 in cooperation with the accelerator 34. That is, the processor 31 receives a photographed image of a product as an input and recognizes the product through deep learning using a common global model 70 managed by the center server 10.

[0072] The processor 31 also configures a second recognition means by executing the processes of ACT8 and ACT9 in cooperation with the accelerator 34. That is, the processor 31 similarly receives a photographed image of a product as input and recognizes the product through deep learning using a specific local model 80 managed by the edge server 20.

[0073] Processor 31 then configures identification means by executing the processing of ACT10. That is, processor 31 identifies the product shown in the captured image based on the recognition result obtained by deep learning using global model 70 and the authentication result obtained by deep learning using local model 80. At this time, processor 31 identifies the product by weighting recognition result B obtained by deep learning using local model 80 to give priority to recognition result A obtained by deep learning using global model 70.

[0074] The processor 31 also constitutes a determination means by executing the process of ACT 12. That is, the processor 31 determines whether the product identified by the identification means is correct or not.

[0075] The processor 31 configures a transmitting means by executing the processes of ACT13 to ACT15. That is, the processor 31 transmits to the edge server 20 training data Da in which the product identified by the identifying means is set as the correct label when the determining means determines that the answer is correct, and in which the corrected product is set as the correct label when the answer is determined to be incorrect.

[0076] The processor 31 also constitutes an output means by executing the process of ACT11 in cooperation with the second communication interface 37. That is, the processor 31 outputs information about the product identified by the identification means to the POS terminal 40. Then, the determination means determines whether the product is authentic based on the information obtained from the POS terminal 40 in the process of ACT12.

[0077] The processor 21 of the edge server 20, which receives training data Da from each image recognition device 30 connected via the second communication network 60, is programmed to execute information processing according to the procedure shown in the flowchart of Fig. 8. That is, the processor 21 waits for the training data Da in ACT21. Then, upon receiving the training data Da, the processor 21 determines YES in ACT21 and proceeds to ACT22. The processor 21 stores the training data Da in the auxiliary storage device 23 in ACT22.

[0078] In ACT23, the processor 21 checks whether the number of pieces of training data Da stored in the auxiliary storage device 23 has reached a specified amount. The specified amount is any value greater than "2." The specified amount is, for example, "100." If the number of pieces of training data Da has not reached the specified amount, the processor 21 determines NO in ACT23 and returns to ACT21. The processor 21 waits for the next piece of training data Da.

[0079] If the number of pieces of training data Da reaches a specified amount, the processor 21 determines YES in ACT 23 and proceeds to ACT 24. The processor 21 activates the accelerator 24 in ACT 24. The processor 21 then commands the accelerator 24 to perform additional learning of the local model 80 using the specified amount of training data Da. In response to this command, the accelerator 24 extracts features from the image data of the training data Da, models them as features of products with correct labels, and adds them to the local model 80.

[0080] When the additional learning by the accelerator 24 is completed, the processor 21 outputs the learning result data Db, which is the result of the additional learning, to the center server 10 in ACT25. The learning result data Db is data of the local model 80 updated by the additional learning. After outputting the learning result data Db, the processor 21 deletes the specified amount of training data Da stored in the auxiliary storage device 23 in ACT26. Then, the processor 21 returns to ACT21.

[0081] Thereafter, the processor 21 stores the training data Da received from each image recognition device 30, and whenever the number of data reaches a specified amount, it repeats the processes of additional learning of the local model 80, transmission of the learning result data Db, and deletion of the training data Da.

[0082] Meanwhile, the processor 11 of the center server 10, which receives training data Da from each edge server 20 connected via the first communication network 50, is programmed to execute information processing according to the procedure shown in the flowchart of Figure 9. That is, the processor 11 waits for learning result data Db in ACT31. Then, upon receiving the learning result data Db, the processor 11 determines YES in ACT31 and proceeds to ACT32. The processor 11 stores the learning result data Db in the auxiliary storage device 13 in ACT32.

[0083] In ACT 33, the processor 11 checks whether the number of pieces of learning result data Db stored in the auxiliary storage device 23 has reached a specified amount. The specified amount is any value greater than "2." The specified amount is, for example, "5." If the number of pieces of learning result data Db has not reached the specified amount, the processor 11 determines NO in ACT 33 and returns to ACT 31. The processor 11 waits for the next learning result data Db.

[0084] If the number of pieces of learning result data Db reaches the specified amount, the processor 11 determines YES in ACT 23 and proceeds to ACT 34. The processor 11 starts the accelerator 14 in ACT 34. The processor 11 then commands the accelerator 14 to aggregate the specified amount of learning result data Db into the global model 70. In response to this command, the accelerator 24 updates the global model 70 so that the data of the local model 80, which is the learning result data Db, is aggregated into the global model 70.

[0085] When the accelerator 14 finishes aggregating the learning result data Db, the processor 11 distributes the global model 70 updated by aggregating the learning result data Db to each edge server 20 in ACT35. The processor 11 also deletes a specified amount of the learning result data Db stored in the auxiliary storage device 13 in ACT36. The processor 11 then returns to ACT31.

[0086] In this way, in each edge server 20, additional learning of the local model 80 is performed using the training data Da obtained from the recognition results of each image recognition device 30 connected via the second communication network 60. Each image recognition device 30 connected via the second communication network 60 is installed in the same store or in each store in the same area. Therefore, it can be said that the local model 80 is a learning model specific to the area.

[0087] Meanwhile, the center server 10 aggregates the local models 80 of the edge servers 20 connected via the first communication network 50, and updates the global model 70. Therefore, the global model 70 can be said to be a learning model common throughout the country.

[0088] To each image recognition device 30, a global model 70 managed by the center server 10 and a local model 80 managed by the edge server 20 connected via the second communication network 60 are distributed.

[0089] The image recognition device 30 recognizes products from photographed images of the products input via the device interface 35 through deep learning using the global model 70. The image recognition device 30 also recognizes products from the same photographed images through deep learning using the local model 80. Each image recognition device 30 then identifies the products shown in the photographed images based on the product recognition result A obtained through deep learning using the global model 70 and the product recognition result B obtained through deep learning using the local model 80.

[0090] Therefore, according to this embodiment, it is possible to accurately recognize products from captured images. In particular, in the past, there was a possibility that accurate product recognition could not be performed for products with small variations in feature values. In this embodiment, products are recognized not only by deep learning using a nationwide learning model, i.e., global model 70, but also by deep learning using a region-specific learning model, i.e., local model 80. Therefore, even products with small variations in feature values ​​obtained from captured images depending on the product type can be recognized with high accuracy.

[0091] Moreover, in this embodiment, product recognition result A obtained by deep learning using the global model 70 is weighted against product recognition result B obtained by deep learning using the local model 80, and product identification is given priority over recognition result A. Therefore, even if the difference in feature values ​​between a product sold exclusively in a specific region and a product sold exclusively in other regions is very small, the product sold exclusively in the specific region is identified with priority, thereby further improving accuracy.

[0092] Although the embodiment of the image recognition device has been described above, the embodiment is not limited to this.

[0093] In the above embodiment, image recognition is performed using the global model 70 in ACT6 of Fig. 7, and then image recognition is performed using the local model 80 in ACT8. In this regard, image recognition may be performed first using the local model 80, and then image recognition may be performed using the global model 70.

[0094] In the above embodiment, the image recognition device 30 performs image recognition by deep learning using a convolutional neural network. The image recognition algorithm is not limited to a convolutional neural network. The image recognition device 30 may perform image recognition by using the global model 70 and the local model 80 through deep learning using another image recognition algorithm.

[0095] In the above embodiment, the image recognition device 30 weights the similarity of feature quantities obtained as recognition results, thereby giving priority to recognition results using the local model 80 over recognition results using the global model 70. The weighting is not limited to similarity. Weighting may be applied to an index other than similarity, and priority may be given to recognition results using the local model 80.

[0096] The above embodiment illustrates an image recognition device 30 used to recognize products such as fresh food and fruit that do not have barcodes. However, the use of the image recognition device is not limited to recognition of products such as fresh food and fruit that do not have barcodes. The image recognition device can be generally applied to image recognition devices that recognize products that are sold nationwide and products that are unique to a region from images. The image recognition device is provided with an operation unit, a display unit, etc., and information on the product identified by the identification means is displayed on the display unit. Then, whether the product is correct or incorrect may be input from the operation unit, and the product may be judged to be correct or incorrect from the input. Also, an operation to prohibit transmission of training data may be performed from the operation unit. The first learning model is a global model 70 commonly used in each store for product recognition, and the second learning model is a local model 80 used in a limited number of stores for product recognition. The first learning model may be a higher-level product classification (e.g., apples, pears, etc.), and the second learning model may be a product or variety included in the product classification (one type of apple, two types of apples, one type of pear, two types of pears, etc.). In this case, the region is irrelevant. The higher-level product classification that is common to all stores is used as the first learning model. The subdivided products (or varieties) included in that classification that are specific to that classification are used as the second learning model. Products are recognized by deep learning using the first learning model. That is, the product classification is recognized. Products are recognized by deep learning using the second learning model. That is, specific products or varieties included in the product classification are recognized. Products are identified based on each recognition result, and specific products or varieties included in that product classification are identified. In this case, when selecting products output as candidates on the POS terminal 40, the product classification and variety (or product) are input. Based on the results of this selection, the second learning model is updated by additional learning about the product and variety. The first learning model is updated by additional learning about the product classification. In this way, by separating the first learning model for product classification from the second learning model for specific products or varieties, it is possible to accurately recognize products even if the differences in the features obtained from the captured images are small.

[0097] Although several embodiments of the present invention have been described, these embodiments are presented as examples and are not intended to limit the scope of the invention. These novel embodiments can be embodied in various other forms, and various omissions, substitutions, and modifications can be made without departing from the spirit of the invention. These embodiments and their modifications are included within the scope of the invention and the scope of the inventions and their equivalents as defined in the claims. The inventions described in the original claims of this application are set forth below. [1] An image recognition device comprising: a first recognition means that receives a photographed image of a product as an input and recognizes the product through deep learning using a common first learning model; a second recognition means that receives a photographed image of the product as an input and recognizes the product through deep learning using a specific second learning model; and an identification means that identifies the product based on the recognition result by the first recognition means and the recognition result by the second recognition means. [2] An image recognition device as described in appendix [1], wherein the first learning model is a learning model commonly used in each store for recognizing products, and the second learning model is a learning model used in a limited number of stores for recognizing products. [3] The image recognition device described in appendix [2], wherein the second learning model is different for each limited store and is updated by additional learning based on the image recognition results processed to recognize products in the limited stores, and the first learning model is updated to aggregate the learned second learning model. [4] The image recognition device according to appendix [1], wherein the identification means compares the recognition result of the first recognition means with the recognition result of the second recognition means and identifies the product from either of the recognition results. [5] The image recognition device according to appendix [4], wherein the identification means identifies the product by weighting the recognition result by the second recognition means more heavily than the recognition result by the first recognition means. [6] An image recognition device according to any one of appendices [1] to [4], further comprising: a determination means for determining whether the product identified by the identification means is correct; and a transmission means for transmitting training data to a server that manages the second learning model, the training data having the identified product as the correct label if the determination means determines that the answer is correct, and the corrected product as the correct label if the determination means determines that the answer is incorrect. [7] An image recognition device as described in appendix [1], further comprising an output means for outputting information about the product identified by the identification means to a terminal, and the determination means for determining whether the product is correct or incorrect based on information from the terminal. [8] The image recognition device according to appendix [6], wherein the transmission means does not transmit the training data if it detects information prohibiting the transmission of the training data to the server. [9] A control program for causing a computer of an image recognition device having an interface that receives a photographed image of a product as input to function as a first recognition means that receives a photographed image of the product as input and recognizes the product through deep learning using a common first learning model, a second recognition means that receives a photographed image of the product as input and recognizes the product through deep learning using a specific second learning model, and an identification means that identifies the product based on the recognition results by the first recognition means and the recognition results by the second recognition means.

[10] An image recognition method in which an image recognition device having an interface that receives a photographed image of a product as input recognizes the product using deep learning with a common first learning model as an input, recognizes the product using deep learning with a specific second learning model as an input, and identifies the product based on the product recognition results obtained by deep learning with the first learning model and the product recognition results obtained by deep learning with the second learning model. [Explanation of symbols]

[0098] 10...center server, 20...edge server, 30...image recognition device, 31...processor, 32...main memory, 33...auxiliary storage device, 34...accelerator, 35...device interface, 36, 37...communication interface, 40...POS terminal, 50...first communication network, 60...second communication network, 70...global model, 80...local model, 90...imaging device, 100...image recognition system.

Claims

1. a first recognition means that receives a photographed image of a product as an input, and recognizes the product by categorizing it into varieties through deep learning using a first learning model commonly used to recognize products sold in each store; a second recognition means for inputting a photographed image of the product and recognizing the product by categorizing it into varieties through deep learning using a second learning model that is used to recognize products sold in stores in a specific region, limited to that region; an identification means for identifying the product based on the recognition result by the first recognition means and the recognition result by the second recognition means; An image recognition device comprising:

2. the second learning model is updated by additional learning based on image recognition results processed to recognize products sold in stores in a specific region, The first learning model is updated to aggregate the learned second learning model. The image recognition device according to claim 1.

3. 2. The image recognition device according to claim 1, wherein the identifying means compares the recognition result of the first recognition means with the recognition result of the second recognition means, and identifies the product from either of the recognition results.

4. 4. The image recognition device according to claim 3, wherein the identifying means identifies the product by weighting the recognition result by the second recognition means more heavily than the recognition result by the first recognition means.

5. a determination means for determining whether the product identified by the identification means is genuine; a transmission means for transmitting training data to a server managing the second learning model, the training data having the identified product as a correct label when the determination means determines that the answer is correct, and the corrected product as a correct label when the determination means determines that the answer is incorrect; The image recognition device according to claim 1 , further comprising:

6. an output means for outputting information about the product identified by the identification means to a terminal; Further comprising: The image recognition device according to claim 5 , wherein the determining means determines whether the product is genuine or not based on information from the terminal.

7. 6. The image recognition device according to claim 5, wherein said transmission means does not transmit said training data when said transmission means detects information prohibiting transmission of said training data to said server.

8. A computer of an image recognition device equipped with an interface that inputs a photographed image of a product, a first recognition means for inputting a photographed image of the product and recognizing the product by subdividing it into varieties through deep learning using a first learning model commonly used for recognizing products sold in each store; A second recognition means that receives a photographed image of the product as an input, and recognizes the product by categorizing it into varieties through deep learning using a second learning model that is used to recognize products sold in stores in a specific region, limited to that region; and an identification means for identifying the product based on the recognition result by the first recognition means and the recognition result by the second recognition means; A control program that functions as a

9. An image recognition device equipped with an interface that takes a photographed image of a product as an input, A photographed image of the product is input, and the product is classified into varieties by deep learning using a first learning model commonly used for recognizing products sold in each store; A photographed image of the product is input, and the product is subdivided into varieties by deep learning using a second learning model that is limited to a specific region and is used to recognize products sold in stores in that region; An image recognition method for identifying the product based on the product recognition results obtained by deep learning using the first learning model and the product recognition results obtained by deep learning using the second learning model.

Citation Information

Patent Citations

  • Combined learning method and device based on gradient momentum acceleration

    CN110889509A

  • Recognition dictionary processing device and recognition dictionary processing program

    JP2014021921A

  • Object recognition scanner system, dictionary server, object recognition scanner, dictionary server program and control program

    JP2014032539A

  • Information processing device, and program

    JP2017139019A

  • Model integration device, model integration system, method and program

    JP2018147261A