Learning device, control method thereof, and computer-readable storage medium storing control program
By extracting objects with a high probability of misidentification as learning data, generating and using differential learning data, the recognizer learns that misidentified objects are of different types, thus solving the problem of misidentification among multiple recognizers and improving learning efficiency.
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- TOYOTA JIDOSHA KK
- Filing Date
- 2023-03-13
- Publication Date
- 2026-04-17
AI Technical Summary
In existing technologies, multiple recognizers are prone to misidentification when recognizing multiple types of objects, making efficient learning impossible.
By extracting objects with a high probability of misidentification as learning data, generating and using differential learning data, the recognizer learns to classify misidentified objects as different categories, thus avoiding misidentification.
This approach reduces misidentification and improves learning efficiency when multiple recognizers are used to identify multiple types of objects.
Smart Images

Figure CN116758382B_ABST
Abstract
Description
Background Technology
[0001] This invention relates to a learning device, its control method, and its control program. Technical Field
[0002] In recent years, there has been development of learning devices that enable multiple recognizers to learn in a way that allows for the individual identification of various types of goods (identification objects) sold in stores such as convenience stores.
[0003] For example, Japanese Patent Publication No. 6-54503 discloses a technology related to learning using a recognition device. In Japanese Patent Publication No. 6-54503, a pattern recognition device is disclosed that: an input pattern is identified by comparing it with a recognition dictionary; the true category and the category obtained as a misidentification result are determined based on the recognition result; based on each category, patterns including true categories that cause misidentification and patterns including categories obtained as misidentification results are input as learning patterns; and the recognition dictionary is learned and updated using each pattern. Summary of the Invention
[0004] However, Japanese Patent Publication No. 6-54503 does not disclose or suggest how to enable multiple recognizers to learn efficiently without misidentifying multiple types of objects when they learn by individually recognizing each type of object. Therefore, the structure of Japanese Patent Publication No. 6-54503 presents the problem of not being able to enable multiple recognizers to learn efficiently without misidentifying multiple types of objects.
[0005] The present invention was made in view of the above background, and its object is to provide a learning device, a control method and a control program thereof that enable each of the multiple recognizers of multiple types of objects to learn efficiently in a manner that does not cause misidentification between the multiple types of objects.
[0006] An embodiment of the present invention relates to a learning apparatus that enables a first recognizer to learn by recognizing a first object as a first type of object from multiple types of objects. The learning apparatus includes: an extraction unit that extracts a first misidentified object, which is an object among the multiple types of objects whose probability of being misidentified as the first object by the first recognizer is at least a predetermined ratio; a learning data generation unit that generates first learning data, which includes an image mapped to both the first object and the first misidentified object; and a learning control unit that uses the first learning data to enable the first recognizer to learn that the first misidentified object is an object of a different type from the first object. In the learning process of the first recognizer for recognizing a predetermined type of first object, this learning apparatus enables the first recognizer to learn that the misidentified object is an object of a different type from the first object only when there is an object (misidentified object) of a different type from the first object whose probability of being misidentified as the first object is at least a predetermined ratio. Therefore, this learning device can learn more efficiently than the case where the first recognizer learns that all objects other than the first recognized object are objects of a different kind from the first recognized object. As a result, this learning device enables each recognizer of multiple recognizers that individually recognize multiple kinds of objects to learn efficiently without misrecognizing objects of different kinds.
[0007] Alternatively, if the probability that the first misidentified object is misidentified as the first identified object in the first recognizer is less than the predetermined ratio, the learning control unit may stop learning using the first learning data.
[0008] The system may also include an evaluation unit. The learning data generation unit further generates benchmark learning data, which includes multiple images of each type of recognition object individually mapped onto the multiple types of recognition objects. The learning control unit uses the benchmark learning data to enable the first recognizer to learn in advance in a manner that correctly recognizes the first recognition object as the first recognition object and does not incorrectly recognize recognition objects other than the first recognition object as the first recognition object. The evaluation unit evaluates whether the first recognizer correctly recognizes the first recognition object as the first recognition object and whether it does not incorrectly recognize recognition objects other than the first recognition object as the first recognition object. The extraction unit extracts the first misidentified object from the evaluation result evaluated by the evaluation unit.
[0009] Alternatively, the learning data generation unit can generate the first learning data, which includes multiple images mapped to both the first identified object and the first misidentified object. The learning control unit uses the first learning data to enable the first recognizer to learn that the first misidentified object is an object of a different kind from the first identified object, until the probability of the first misidentified object being misidentified as the first identified object in the first recognizer is less than the predetermined ratio.
[0010] Alternatively, the learning data generation unit may generate the first learning data when it has extracted from the extraction unit a plurality of first misidentified objects that are identified as first identification objects in the first recognizer at a predetermined rate or higher. The first learning data includes a plurality of images in which each of the first misidentified objects is mapped to the first identification object and the plurality of first misidentified objects. The learning control unit uses the first learning data to enable the first recognizer to learn that the plurality of first misidentified objects are objects of a different kind from the first identification object, until the probability of each of the plurality of first misidentified objects being misidentified as the first identification object in the first recognizer is less than the predetermined rate.
[0011] The second recognizer can further learn by recognizing a second object as a second type of object from the multiple types of objects. The extraction unit also extracts a second misidentified object, which is an object among the multiple types of objects that has a probability of being misidentified as the second object by the second recognizer at a predetermined rate or higher. The learning data generation unit also generates second learning data, which includes images of both the second object and the second misidentified object. The learning control unit then uses the second learning data to enable the second recognizer to learn that the second misidentified object is an object of a different type from the second object.
[0012] Alternatively, if the probability that the second misidentified object is misidentified as the second identified object in the second recognizer is less than the predetermined ratio, the learning control unit may stop learning using the second learning data.
[0013] The system may also include an evaluation unit. The learning data generation unit further generates benchmark learning data, which includes multiple images of each of the multiple types of recognition objects individually mapped onto the learning data. The learning control unit uses the benchmark learning data to pre-learn the first recognizer in a way that correctly identifies the first recognition object as the first recognition object and in a way that does not incorrectly identify recognition objects other than the first recognition object as the first recognition object. It also pre-learns the second recognizer in a way that correctly identifies the second recognition object as the second recognition object and in a way that does not incorrectly identify the second recognition object as the first recognition object. The evaluation unit pre-learns how to incorrectly identify objects other than the first identification object as the second identification object. The evaluation unit evaluates whether the first identifier correctly identifies the first identification object as the first identification object and whether it does not incorrectly identify objects other than the first identification object as the first identification object. The evaluation unit also evaluates whether the second identifier correctly identifies the second identification object as the second identification object and whether it does not incorrectly identify objects other than the second identification object as the second identification object. The extraction unit extracts the first misidentified object and the second misidentified object from the evaluation results evaluated by the evaluation unit.
[0014] Alternatively, if a third identification object is added as a new third type of identification object to the plurality of identification objects, the evaluation unit is configured to further evaluate whether the first recognizer did not mistakenly identify the third identification object as the first identification object before learning that the third identification object is not the first identification object through the first recognizer. The extraction unit is configured to extract, from the evaluation results evaluated by the evaluation unit, the first misidentified object among the plurality of identification objects including the third identification object, which has a probability of being misidentified as the first identification object in the first recognizer that is above a predetermined ratio.
[0015] Alternatively, if a third identification object is added as a new third type of identification object to the plurality of identification objects, the evaluation unit is configured to further evaluate whether the first recognizer did not mistakenly identify the third identification object as the first identification object before learning through the first recognizer that the third identification object is not the first identification object, and whether the second recognizer did not mistakenly identify the third identification object as the second identification object before learning through the second recognizer that the third identification object is not the second identification object. The extraction unit is configured to extract from the evaluation results evaluated by the evaluation unit the first misidentified object, which has a predetermined probability of being misidentified as the first identification object in the first recognizer or higher, and the second misidentified object, which has a predetermined probability of being misidentified as the second identification object in the second recognizer or higher, from the plurality of identification objects including the third identification object.
[0016] Alternatively, the learning data generation unit may update the baseline learning data in a manner that also includes an image of the third identification object, which is separately mapped into the plurality of identification objects. If, despite learning using the first learning data through the first recognizer, the probability that the first misidentified object is misidentified as the first identification object in the first recognizer is maintained at or above the predetermined ratio, the learning control unit, after initializing the learning content using the first recognizer, causes the first recognizer to learn using at least the updated baseline learning data.
[0017] Alternatively, the learning data generation unit may update the baseline learning data by including an image of the third recognition object, which is separately mapped from the plurality of recognition objects. If, despite learning using the first learning data through the first recognizer, the probability of the first misidentified object being misidentified as the first recognition object in the first recognizer remains above the predetermined ratio, the learning control unit, after initializing the learning content using the first recognizer, causes the first recognizer to learn using at least the updated baseline learning data. Similarly, if, despite learning using the second learning data through the second recognizer, the probability of the second misidentified object being misidentified as the second recognition object in the second recognizer remains above the predetermined ratio, the second recognizer, after initializing the learning content using the second recognizer, causes the second recognizer to learn using at least the updated baseline learning data.
[0018] The third recognizer can also be further trained by recognizing the third object from the multiple types of objects. The extraction unit also extracts a third misidentified object, which is an object among the multiple types of objects that has a probability of being misidentified as the third object by the third recognizer at a predetermined rate or higher. The learning data generation unit also generates third learning data, which includes images of both the third object and the third misidentified object. The learning control unit also uses the third learning data to enable the third recognizer to learn that the third misidentified object is an object of a different type from the third object.
[0019] Alternatively, if a fourth identification object, which is a fourth type of identification object, is excluded from the multiple types of identification objects, the extraction unit, after initializing the learning content learned by the learning control unit using the first recognition device with the first learning data, extracts a first misidentified object from the multiple types of identification objects other than the fourth identification object. The first misidentified object has a probability of being misidentified as the first identification object by the first recognition device at a predetermined rate or higher. The learning data generation unit generates new first learning data, which includes images of both the first identification object and the newly extracted first misidentified object. In addition to initializing the learning content learned by the first recognition device using the first learning data, the learning control unit causes the first recognition device to learn using the newly generated first learning data.
[0020] Alternatively, if a fourth identification object, which is a fourth type of identification object, has been excluded from the plurality of identification objects, the extraction unit, in a state where the learning content learned using the first learning data with the first recognizer has been initialized by the learning control unit, and in a state where the learning content learned using the second learning data with the second recognizer has been initialized by the learning control unit, may extract from the plurality of identification objects other than the fourth identification object the probability of misidentifying it as the first identification object in the first recognizer being at a predetermined ratio or higher, and the probability of misidentifying it as the second identification object in the second recognizer being at a predetermined ratio or higher. For the second misidentified object with a rate higher than a predetermined rate, the learning data generation unit generates first learning data including images of both the first identified object and the newly extracted first misidentified object, and second learning data including images of both the second identified object and the newly extracted second misidentified object. The learning control unit, in addition to initializing the learning content learned by the first recognizer using the first learning data, causes the first recognizer to learn using the newly generated first learning data, and in addition to initializing the learning content learned by the second recognizer using the second learning data, causes the second recognizer to learn using the newly generated second learning data.
[0021] One embodiment of the present invention relates to a control method for a learning device that enables a first recognizer to learn, at least in a manner that identifies a first recognition object as a first recognition object of a first type from multiple types of recognition objects. The control method of the learning device involves: extracting a first misidentified object, which is a recognition object among the multiple types of recognition objects whose probability of being misidentified as the first recognition object by the first recognizer is greater than or equal to a predetermined ratio; generating first learning data, which includes an image in which both the first recognition object and the first misidentified object are mapped; and using the first learning data, enabling the first recognizer to learn that the first misidentified object is an object of a different type from the first recognition object. In this control method of the learning device, during the learning of the first recognizer for recognizing a first recognition object of a predetermined type, the first recognizer learns that the misidentified object is an object of a different type from the first recognition object only when there is a recognition object (misidentified object) of a different type from the first recognition object whose probability of being misidentified as the first recognition object is greater than or equal to a predetermined ratio. Therefore, the control method of this learning device is more efficient than the case where the first recognizer learns that all objects other than the first recognized object are objects of a different kind from the first recognized object. As a result, the control method of this learning device enables each recognizer of multiple recognizers that individually recognize multiple kinds of objects to learn efficiently without misrecognition between multiple kinds of objects.
[0022] One embodiment of the present invention relates to a control program for a learning device that enables a first recognizer to learn, at least in a manner that identifies a first object as a first type of object from a plurality of types of objects. The control program causes a computer to perform: a process for extracting a first misidentified object, which is an object among the plurality of types of objects whose probability of being misidentified as the first object by the first recognizer is greater than or equal to a predetermined ratio; a process for generating first learning data, which includes an image in which both the first object and the first misidentified object are mapped; and a process for using the first learning data to enable the first recognizer to learn that the first misidentified object is an object of a different type from the first object. In the learning process of the first recognizer for identifying a predetermined type of first object, the control program enables the first recognizer to learn that the misidentified object is an object of a different type from the first object only when there is an object (misidentified object) of a different type from the first object whose probability of being misidentified as the first object is greater than or equal to a predetermined ratio. Therefore, this control program is more efficient at learning than the case where the first recognizer learns that all objects other than the first recognized object are of a different kind from the first recognized object. As a result, this control program enables each of the multiple recognizers of multiple types of objects to learn efficiently without misidentifying objects of different types.
[0023] According to the present invention, a learning apparatus, a control method, and a control program are provided that enable each of the multiple recognizers of multiple types of objects to learn efficiently in a manner that avoids misidentification between the multiple types of objects.
[0024] The above and other objects, features and advantages of this disclosure will be more fully understood from the detailed description given below and the accompanying drawings, which are given by way of illustration only, and should therefore not be considered as limitations on this disclosure. Attached Figure Description
[0025] Figure 1 This is a block diagram illustrating a structural example of the learning system according to Embodiment 1.
[0026] Figure 2 This is a block diagram illustrating a structural example of the learning device according to Embodiment 1.
[0027] Figure 3 It is shown Figure 2 A flowchart illustrating the operation of the learning device shown.
[0028] Figure 4This is a diagram showing an example of an image contained in benchmark learning data.
[0029] Figure 5 This is a diagram showing an example of an image contained in supplemental learning data.
[0030] Figure 6 This is a flowchart illustrating the operation of the learning device according to Embodiment 2.
[0031] Figure 7 This is a diagram showing an example of an image contained in benchmark learning data.
[0032] Figure 8 This is a diagram showing an example of an image contained in supplemental learning data.
[0033] Figure 9 This is a flowchart illustrating the operation of the learning device according to Embodiment 3. Detailed Implementation
[0034] The present invention will now be described according to embodiments thereof, but the invention as defined in the claims is not limited to these embodiments. Furthermore, not all structures described in the embodiments are necessarily necessary means of solving the problem. For clarity of description, the following descriptions and drawings are appropriately omitted and simplified. In the drawings, the same reference numerals are used for the same elements, and repeated descriptions are omitted as necessary.
[0035] <Implementation Method 1>
[0036] Figure 1 This is a block diagram illustrating a structural example of the learning system 1 according to Embodiment 1. The learning system 1 is a system in which multiple recognizers learn by using a learning device 11 to individually recognize each type of object among multiple types of objects. Here, in the learning process of a first recognizer for recognizing a predetermined type of first object, the learning device 11 in the learning system 1 only teaches that a misidentified object (a misidentified object) of a different type from the first object is a different type of object if the probability of misidentification is greater than a predetermined ratio. Therefore, the learning device 11 can learn more efficiently than if the first recognizer learns that all objects other than the first object are different types of objects. As a result, the learning device 11 enables each recognizer of multiple recognizers that individually recognize each type of object among multiple types of objects to learn efficiently without misidentification between multiple types of objects. This will be explained in detail below.
[0037] like Figure 1As shown, the learning system 1 has a learning device 11 and n (n is an integer of 2 or more) recognizers 12_1 to 12_n.
[0038] n identifiers 12_1 to 12_n are devices for individually identifying each of n types of identification objects TG1 to TGn. The n types of identification objects TG1 to TGn are, for example, n types of goods sold in stores such as convenience stores.
[0039] For example, the recognizer (first recognizer) 12_1 is a device for recognizing a first-class object (first recognized object) TG1 from n types of objects. Similarly, the recognizer (second recognizer) 12_2 is a device for recognizing a second-class object (second recognized object) TG2 from n types of objects TG1 to TGn. Likewise, the recognizer 12_i (where i is any integer from 1 to n) is a device for recognizing an ith-class object TTi from n types of objects TG1 to TGn.
[0040] The learning device 11 enables the recognizers 12_1 to 12_n to learn by individually recognizing each of the recognition objects TG1 to TGn. In other words, the learning device 11 trains the recognizers 12_1 to 12_n by individually recognizing each of the recognition objects TG1 to TGn.
[0041] For example, the learning device 11 enables the recognizer 12_1 to learn in a manner that it can recognize the first type of recognition object TG1 from recognition objects TG1 to TGn. Additionally, the learning device 11 enables the recognizer 12_2 to learn in a manner that it can recognize the second type of recognition object TG2 from recognition objects TG1 to TGn. Similarly, the learning device 11 enables the recognizer 12_i to learn in a manner that it can recognize the i-th type of recognition object TTi from recognition objects TG1 to TGn.
[0042] Figure 2 This is a block diagram illustrating an example of the structure of the learning device 11. For example... Figure 2 As shown, the learning device 11 includes at least a learning data generation unit 111, a learning control unit 112, an evaluation unit 113, and an extraction unit 114.
[0043] The learning data generation unit 111 generates learning data used when learning through the recognizers 12_1 to 12_n.
[0044] Specifically, the learning data generation unit 111 generates baseline learning data comprising multiple images of each of the recognition objects TG1 to TGn, individually mapped onto them. In other words, the learning data generation unit 111 generates baseline learning data comprising multiple images of each of the recognition objects TG1 to TGn, individually mapped onto them. Furthermore, the learning data generation unit 111 generates additional learning data. The generation of additional learning data by the learning data generation unit 111 will be described later.
[0045] The learning control unit 112 uses the learning data generated by the learning data generation unit 111 to enable the recognizers 12_1 to 12_n to learn.
[0046] For example, the learning control unit 112 uses baseline learning data to enable the recognizer 12_1 to learn in a way that correctly identifies object TG1 as object TG1 and avoids incorrectly identifying objects other than object TG1 as object TG1. Similarly, the learning control unit 112 uses baseline learning data to enable the recognizer 12_2 to learn in a way that correctly identifies object TG2 as object TG2 and avoids incorrectly identifying objects other than object TG2 as object TG2. Likewise, the learning control unit 112 uses baseline learning data to enable the recognizer 12_i to learn in a way that correctly identifies object TTi as object TTi and avoids incorrectly identifying objects other than object TTi as object TTi. Furthermore, the learning control unit 112 uses additional learning data to enable the recognizers 12_1 to 12_n to learn additionally. The learning control using additional learning data by the learning control unit 112 will be described later.
[0047] The evaluation unit 113 evaluates whether each of the recognizers 12_1 to 12_n correctly recognizes the specified object. Furthermore, during the evaluation process of the recognizers 12_1 to 12_n by the evaluation unit 113, either a physical object can be used for recognition, or an image projected with the object can be used.
[0048] For example, the evaluation unit 113 evaluates whether the identifier 12_1 correctly identifies the identification object TG1 as the identification object TG1 and whether it incorrectly identifies identification objects other than the identification object TG1 as the identification object TG1. Similarly, the evaluation unit 113 evaluates whether the identifier 12_2 correctly identifies the identification object TG2 as the identification object TG2 and whether it incorrectly identifies identification objects other than the identification object TG2 as the identification object TG2. Likewise, the evaluation unit 113 evaluates whether the identifier 12_i correctly identifies the identification object TTi as the identification object TTi and whether it incorrectly identifies identification objects other than the identification object TTi as the identification object TTi.
[0049] In each of the recognizers 12_1 to 12_n, the extraction unit 114 extracts misidentified objects that have a probability of being misidentified as a specified recognition object that is at or above a predetermined ratio.
[0050] For example, the extraction unit 114 extracts the identification objects TG1 to TGn that have a predetermined probability of being misidentified as identification object TG1 in the recognizer 12_1, i.e., misidentified objects (first misidentified objects) FTG1. Similarly, the extraction unit 114 extracts the identification objects TG1 to TGn that have a predetermined probability of being misidentified as identification object TG2 in the recognizer 12_2, i.e., misidentified objects (second misidentified objects) FTG2. Likewise, the extraction unit 114 extracts the identification objects TG1 to TGn that have a predetermined probability of being misidentified as identification object TTi in the recognizer 12_i, i.e., misidentified objects FTGi.
[0051] Here, the learning data generation unit 111 generates not only baseline learning data including multiple images of each recognition object TG1 to TGn individually mapped, but also additional learning data including images of both misidentified objects and recognized objects extracted by the extraction unit 114.
[0052] For example, the learning data generation unit 111 generates additional learning data (first learning data) including images of recognition objects TG1 to TGn, such as those that are misidentified as recognition object TG1 in recognizer 12_1 at a predetermined rate or higher, i.e., images of both misidentified object FTG1 and recognition object TG1. Additionally, the learning data generation unit 111 generates additional learning data (second learning data) including images of recognition objects TG1 to TGn, such as those that are misidentified as recognition object TG2 in recognizer 12_2 at a predetermined rate or higher. Similarly, the learning data generation unit 111 generates additional learning data including images of recognition objects TG1 to TGn, such as those that are misidentified as recognition object TGI in recognizer 12_i at a predetermined rate or higher, i.e., images of both misidentified object FTGi and recognition object TGI.
[0053] Furthermore, the learning control unit 112 uses the additional learning data generated by the learning data generation unit 111 to enable the recognizers 12_1 to 12_n to learn additionally.
[0054] For example, the learning control unit 112 uses additional learning data that includes images mapped to both the misidentified object FTG1 and the identified object TG1, enabling the recognizer 12_1 to learn that the misidentified object FTG1 is a different type of object from the identified object TG1. Similarly, the learning control unit 112 uses additional learning data that includes images mapped to both the misidentified object FTG2 and the identified object TG2, enabling the recognizer 12_2 to learn that the misidentified object FTG2 is a different type of object from the identified object TG2. Likewise, the learning control unit 112 uses additional learning data that includes images mapped to both the misidentified object FTGi and the identified object TGi, enabling the recognizer 12_i to learn that the misidentified object FTGi is a different type of object from the identified object TGi.
[0055] Furthermore, the learning control unit 112 stops learning through the recognizers 12_1 to 12_n when there are no misidentified objects FTG1 to FTGn.
[0056] (The operation of learning device 11)
[0057] Next, use Figures 3-5 This explains the operation of the learning device 11. Figure 3 This is a flowchart illustrating the operation of the learning device 11. Figure 4 This is a diagram showing an example of an image contained in benchmark learning data. Figure 5 This is a diagram showing an example of an image contained in supplemental learning data.
[0058] Furthermore, the following explanation will use the example of the learning device 11 learning to identify the first type of identification object (first identification object) TG1 from the three types of identification objects TG1 to TG3, and the example of the identifier (first identifier) 12_1 learning. Specifically, the explanation will use the example of the learning device 11 learning to identify the product "ABC tea" as the first type of identification object TG1 from the three types of products "ABC tea", "DEF tea", and "GHI tea" that are the three types of identification objects TG1 to TG3.
[0059] First, the learning data generation unit 111 generates reference learning data including multiple images of each recognition object TG1 to TG3 individually mapped (step S101).
[0060] Specifically, the learning data generation department 111, such as Figure 4 As shown in the example, baseline learning data is generated including images of product “ABC Tea” mapped individually as identification object TG1, images of product “DEF Tea” mapped individually as identification object TG2, and images of product “GHI Tea” mapped individually as identification object TG3.
[0061] Then, the learning control unit 112 uses the benchmark learning data to train the recognizer 12_1 (step S102). That is, the learning control unit 112 uses the benchmark learning data to enable the recognizer 12_1 to learn.
[0062] Specifically, the learning control unit 112 uses benchmark learning data to enable the recognizer 12_1 to learn in a way that correctly identifies object TG1 as object TG1 and avoids incorrectly identifying objects TG2 and TG3 as object TG1. In this example, the learning control unit 112 enables the recognizer 12_1 to learn in a way that correctly identifies product "ABC Tea" (object TG1) as product "ABC Tea" and avoids incorrectly identifying products "DEF Tea" (objects TG2) and "GHI Tea" (objects TG3) as product "ABC Tea".
[0063] Next, the evaluation unit 113 evaluates the identifier 12_1 (step S103). Specifically, the evaluation unit 113 evaluates whether the identifier 12_1 correctly identifies the identification object TG1 as the identification object TG1 and whether it incorrectly identifies the identification objects TG2 and TG3 as the identification object TG1. In this example, the evaluation unit 113 evaluates whether the identifier 12_1 correctly identifies the product "ABC Tea" as the product "ABC Tea" and whether it incorrectly identifies the products "DEF Tea" and "GHI Tea" as the products "ABC Tea" and "TG2" and "TG3" as the product "ABC Tea".
[0064] Based on the evaluation results evaluated by the evaluation unit 113, it becomes clear whether there is a misidentified object (first misidentified object) FTG1 among the identified objects TG1 to TG3 that has a probability of being misidentified as identified object TG1 in the recognizer 12_1 at a predetermined rate or higher (step S104).
[0065] Here, if there is no misidentified object FTG1 (No in step S104), it is determined that the recognition performance of the recognizer 12_1 has reached a sufficient level, and the learning of the recognizer 12_1 ends without the need for extraction by the extraction unit 114.
[0066] In contrast, if there is a misidentified object FTG1 ("Yes" in step S104), it is determined that the recognition performance of the recognizer 12_1 has not reached a sufficient level, and the learner continues to learn through the recognizer 12_1.
[0067] Specifically, firstly, the extraction unit 114 extracts the misidentified object FTG1 (No in step S105 → step S106). Then, the learning data generation unit 111 generates additional learning data (first learning data) including the image mapped to both the misidentified object FTG1 extracted by the extraction unit 114 and the identified object TG1 (step S107).
[0068] For example, in the extraction unit 114, the product "GHI Tea" is extracted as the identification object TG3, which is identified as the misidentified object FTG1. In this case, the learning data generation unit 111... Figure 5 As shown in the example, additional learning data is generated that includes images of both the product "ABC Tea" extracted as the identification object TG1 and the product "GHI Tea" extracted as the misidentified object FTG1.
[0069] Then, the learning control unit 112 uses additional learning data to perform additional training on the recognizer 12_1 (step S102). That is, the learning control unit 112 uses additional learning data to make the recognizer 12_1 learn additionally.
[0070] Specifically, the learning control unit 112 uses additional learning data to enable the recognizer 12_1 to learn that the misidentified object FTG1 is a different type of object from the identified object TG1. In this example, the learning control unit 112 uses additional learning data to enable the recognizer 12_1 to learn that the product "GHI Tea" extracted as the misidentified object FTG1 is a different type of object from the product "ABC Tea" as the identified object TG1.
[0071] Next, the evaluation unit 113 evaluates the identifier 12_1 again (step S103). Specifically, the evaluation unit 113 evaluates whether the identifier 12_1 correctly identifies the identification object TG1 as the identification object TG1 and whether it incorrectly identifies the identification objects TG2 and TG3 as the identification object TG1. In this example, the evaluation unit 113 evaluates whether the identifier 12_1 correctly identifies the product "ABC Tea" as the product "ABC Tea" and whether it incorrectly identifies the products "DEF Tea" and "GHI Tea" as the products "ABC Tea" and "TG2" and "TG3" as the product "ABC Tea".
[0072] Based on the evaluation results evaluated by the evaluation unit 113, it becomes clear whether there is an identification object FTG1 among the identification objects TG1 to TG3 that has a probability of being misidentified as identification object TG1 in the identifier 12_1 at a predetermined rate or higher (step S104).
[0073] Here, if there is no misidentified object FTG1 (No in step S104), it is determined that the recognition performance of the recognizer 12_1 has reached a sufficient level, and the learning of the recognizer 12_1 ends without the need for extraction by the extraction unit 114.
[0074] In contrast, if there is a misidentified object FTG1 ("Yes" in step S104), it is determined that the recognition performance of the recognizer 12_1 has not reached a sufficient level, and the learner continues to learn through the recognizer 12_1.
[0075] Then, the process of repeatedly extracting the misidentified object FTG1 (step S105 "No" → step S106), generating additional learning data (step S107), training the recognizer 12_1 using the additional learning data (step S102), and evaluating the recognizer 12_1 (step S103) continues until the misidentified object FTG1 disappears. However, if the number of training iterations using the additional learning data reaches a predetermined number before the misidentified object FTG1 disappears (step S104 "Yes" → step S105 "Yes"), it is determined that the misidentified object FTG1 will not disappear, and this is notified to, for example, a user of this system (step S108). The user who receives the notification can take countermeasures such as changing the learning method.
[0076] Furthermore, the above description of the operation describes the case where the learning device 11 trains the recognizer 12_1, but it is not limited to this. By performing the same processing as for the recognizer 12_1, the learning device 11 can also train recognizers 12_2 and 12_3 other than the recognizer 12_1.
[0077] Furthermore, in the above description of the operation, the example given is that the learning device 11 extracts the product "GHI Tea" (TG3), which is the recognition object, as a false recognition object FTG1 during the training of the recognizer 12_1. However, it is not limited to this. The learning device 11 can extract the product "DEF Tea" (TG2), which is the recognition object, as a false recognition object FTG1, or it can extract both TG2 and TG3 as false recognition objects FTG1.
[0078] Thus, in the learning process of the first recognizer for recognizing a predetermined type of first object, the learning device 11 of this embodiment only teaches the first recognizer to learn that a misidentified object is a different type of object from the first object if the probability of misidentification as the first object is greater than a predetermined ratio. Therefore, the learning device 11 can learn more efficiently than if the first recognizer learns that all objects other than the first object are different types of objects from the first object. As a result, the learning device 11 enables each recognizer of multiple recognizers that individually recognize multiple types of objects to learn efficiently without misidentification between the multiple types of objects.
[0079] <Implementation Method 2>
[0080] In this embodiment, using Figures 6-8 This describes the operation of the learning device 11 when new types of objects to be identified are added. Figure 6 This is a flowchart illustrating the operation of the learning device 11 according to Embodiment 2. Figure 7 This is a diagram showing an example of an image contained in benchmark learning data. Figure 8 This is a diagram showing an example of an image contained in supplemental learning data.
[0081] Furthermore, the following describes the learning device 11, for example, via... Figure 3 The process shown is as follows: after the recognizer (first recognizer) 12_1 learns by recognizing the first type of recognition object (first recognition object) TG1 from three types of recognition objects TG1 to TG3, the recognizer 12_1 learns again when a new type of recognition object TG4 is added.
[0082] First, the identification object TG4 is added (step S201). At this time, the learning data generation unit 111 generates an image separately mapped to the identification object TG4 and includes it in the reference learning data. That is, the learning data generation unit 111 updates the reference learning data to include the reference learning data that separately maps the image of the identification object TG4, in addition to the multiple images of each identification object individually mapped to the identification objects TG1 to TG3 (step S202).
[0083] Specifically, the learning data generation department 111, such as Figure 7As shown in the example, the baseline learning data is updated by including, in addition to individually projecting images of the product "ABC Tea" (which is the identification target TG1), the product "DEF Tea" (which is the identification target TG2), and the product "GHI Tea" (which is the identification target TG3), an image of the product "JKL Tea" (which is the identification target TG4) is also individually projected.
[0084] Furthermore, although the explanation is omitted, for example by Figure 3 The processing shown is performed independently of the training of recognizer 12_1, and recognizer 12_4 is trained to identify object TG4.
[0085] Subsequently, before training the recognizer 12_1 using the updated benchmark learning data, the evaluation unit 113 evaluates the recognizer 12_1 (step S203). Specifically, the evaluation unit 113 evaluates whether the recognizer 12_1 correctly identifies the recognition object TG1 as the recognition object TG1 and whether it incorrectly identifies the recognition objects TG2, TG3, and TG4 as the recognition object TG1. In this example, the evaluation unit 113 evaluates whether the recognizer 12_1 correctly identifies the product "ABC Tea" as the product "ABC Tea" and whether it incorrectly identifies the products "DEF Tea," "GHI Tea," and "JKL Tea" as the product "ABC Tea," which are the recognition objects TG2, TG3, and TG4.
[0086] Based on the evaluation results evaluated by the evaluation unit 113, it becomes clear whether there is a misidentified object (first misidentified object) FTG1 among the identified objects TG1 to TG4 that has a probability of being misidentified as identified object TG1 in the recognizer 12_1 at a predetermined rate or higher (step S204).
[0087] Here, if there is no misidentified object FTG1 (No in step S204), it is determined that the recognition performance of the recognizer 12_1 has reached a sufficient level, and the learning of the recognizer 12_1 ends without the need for extraction by the extraction unit 114.
[0088] In contrast, if there is a misidentified object FTG1 ("Yes" in step S204), it is determined that the recognition performance of the recognizer 12_1 has not reached a sufficient level, and the learner continues to learn through the recognizer 12_1.
[0089] Specifically, firstly, the extraction unit 114 extracts the misidentified object FTG1 (No in step S205 → step S206). Then, the learning data generation unit 111 generates additional learning data (first learning data) including the image mapped to both the misidentified object FTG1 extracted by the extraction unit 114 and the identified object TG1 (step S207).
[0090] For example, in the extraction unit 114, the product "JKL Tea" is extracted as the identification object TG4, which is identified as the misidentified object FTG1. In this case, the learning data generation unit 111... Figure 8 As shown in the example, additional learning data is generated that includes images of both the product "ABC Tea" extracted as the identification object TG1 and the product "JKL Tea" extracted as the misidentified object FTG1.
[0091] Then, the learning control unit 112 uses the additional learning data to perform additional training on the recognizer 12_1 (step S208). That is, the learning control unit 112 uses the additional learning data to make the recognizer 12_1 learn additionally.
[0092] Specifically, the learning control unit 112 uses additional learning data to enable the recognizer 12_1 to learn that the misidentified object FTG1 is a different type of object from the identified object TG1. In this example, the learning control unit 112 uses additional learning data to enable the recognizer 12_1 to learn that the product "JKL Tea" extracted as the misidentified object FTG1 is a different type of object from the product "ABC Tea" as the identified object TG1.
[0093] Next, the evaluation unit 113 evaluates the identifier 12_1 again (step S203). Specifically, the evaluation unit 113 evaluates whether the identifier 12_1 correctly identifies the identification object TG1 as the identification object TG1 and whether it incorrectly identifies the identification objects TG2, TG3, and TG4 as the identification object TG1. In this example, the evaluation unit 113 evaluates whether the identifier 12_1 correctly identifies the product "ABC Tea" as the product "ABC Tea" and whether it incorrectly identifies the products "DEF Tea", "GHI Tea", and "JKL Tea" as the product "ABC Tea" as the products "TG2", "TG3", and "TG4".
[0094] Based on the evaluation results evaluated by the evaluation unit 113, it becomes clear whether there is an identification object FTG1 among the identification objects TG1 to TG4 that has a probability of being misidentified as identification object TG1 in the identifier 12_1 at a predetermined rate or higher (step S204).
[0095] Here, if there is no misidentified object FTG1 (No in step S204), it is determined that the recognition performance of the recognizer 12_1 has reached a sufficient level, and the learning of the recognizer 12_1 ends without the need for extraction by the extraction unit 114.
[0096] In contrast, if there is a misidentified object FTG1 ("Yes" in step S204), it is determined that the recognition performance of the recognizer 12_1 has not reached a sufficient level, and the learner continues to learn through the recognizer 12_1.
[0097] Then, the process of repeatedly extracting the misidentified object FTG1 (step S205 "No" → step S206), generating additional learning data (step S207), training the recognizer 12_1 using the additional learning data (step S208), and evaluating the recognizer 12_1 (step S203) continues until the misidentified object FTG1 disappears.
[0098] However, if the number of training sessions using the additional learning data reaches a predetermined number before the misidentified object FTG1 disappears ("Yes" in step S204 → "Yes" in step S205), the learning content of the recognizer 12_1 is fully initialized only once ("Yes" in step S209 → "Yes" in step S210).
[0099] After initialization, the learning control unit 112 uses the updated benchmark learning data to train the recognizer 12_1 (step S208). That is, the learning control unit 112 retrains the recognizer 12_1 using the updated benchmark learning data, which includes multiple images of each of the recognition objects TG1 to TG4, each individually mapped.
[0100] Next, the evaluation unit 113 evaluates the identifier 12_1 (step S203). Based on the evaluation result obtained by the evaluation unit 113, it becomes clear whether the misidentified object FTG1 exists (step S204).
[0101] Here, if there is no misidentified object FTG1 (No in step S204), it is determined that the recognition performance of the recognizer 12_1 has reached a sufficient level, and the learning of the recognizer 12_1 ends without the need for extraction by the extraction unit 114.
[0102] In contrast, if there is a misidentified object FTG1 ("Yes" in step S204), it is determined that the recognition performance of the recognizer 12_1 has not reached a sufficient level, and the learner continues to learn through the recognizer 12_1.
[0103] Then, the process of repeatedly extracting the misidentified object FTG1 (step S205 "No" → step S206), generating additional learning data (step S207), training the recognizer 12_1 using the additional learning data (step S208), and evaluating the recognizer 12_1 (step S203) continues until the misidentified object FTG1 disappears.
[0104] However, if the number of training iterations using supplementary learning data reaches the predetermined number again before the misidentified object FTG1 disappears (step S204 "Yes" → step S205 "Yes" → step S209 "No"), it is determined that the misidentified object FTG1 will not disappear, and this is notified to, for example, a user of this system (step S211). The user who receives the notification can take countermeasures such as changing the learning method.
[0105] Thus, in the learning process of the first recognizer for recognizing a predetermined type of first object, the learning device 11 of this embodiment only teaches the first recognizer to learn that a misidentified object is a different type of object from the first object if the probability of misidentification as the first object is greater than a predetermined ratio. Therefore, the learning device 11 can learn more efficiently than if the first recognizer learns that all objects other than the first object are different types of objects from the first object. As a result, the learning device 11 enables each recognizer of multiple recognizers that individually recognize multiple types of objects to learn efficiently without misidentification between the multiple types of objects.
[0106] Furthermore, even when a new type of recognition object is added after learning by the first recognizer, the learning device 11 of this embodiment evaluates the first recognizer before relearning it, and only relearns it when necessary. Therefore, the learning device 11 can prevent the first recognizer from needlessly relearning. Moreover, even when relearning the first recognizer, the learning device 11 does so while retaining the already learned content. Thus, the learning device 11 enables the first recognizer to relearn efficiently.
[0107] <Implementation Method 3>
[0108] In this embodiment, using Figure 9 This describes the operation of the learning device 11 when multiple recognition objects are excluded from the case of any recognition object. Figure 9 This is a flowchart illustrating the operation of the learning device 11 according to Embodiment 3.
[0109] Furthermore, the following describes the learning device 11, for example, via... Figure 3 The process shown involves the recognizer (first recognizer) 12_1 learning to recognize the first type of object (first object) TG1 from three types of objects TG1 to TG3, and then, for example, the recognizer 12_1 learning again while excluding the object TG2.
[0110] First, the object to be identified, TG2, is excluded (step S301). At this time, the learning data generation unit 111 can also delete the image containing the object to be identified, TG2, from the multiple images included in the reference learning data. That is, the learning data generation unit 111 can also update the reference learning data to the reference learning data with the image containing the object to be identified, TG2, deleted.
[0111] Next, the learning content using the additional learning data is initialized from recognizer 12_1 (step S302). Therefore, it is unnecessary to determine whether the learning content of recognizer object TG2, which is the same as that of recognizer object TG1, has disappeared, thus improving the performance of recognizer 12_1. Furthermore, at this time, recognizer 12_1 can initialize not only the learning content using the additional learning data but also the learning content using the baseline learning data. That is, all learning content can be initialized from recognizer 12_1.
[0112] Then, before training the recognizer 12_1, the evaluation unit 113 evaluates the recognizer 12_1 (step S303). Specifically, the evaluation unit 113 evaluates whether the recognizer 12_1 correctly recognizes the object TG1 as the object TG1 and whether it incorrectly recognizes the object TG3 as the object TG1.
[0113] Based on the evaluation results evaluated by the evaluation unit 113, it becomes clear whether there is a misidentified object (first misidentified object) FTG1 among the identified objects TG1 and TG3 that has a probability of being misidentified as identified object TG1 in the recognizer 12_1 that is above a predetermined ratio (step S304).
[0114] Here, if there is no misidentified object FTG1 (No in step S304), it is determined that the recognition performance of the recognizer 12_1 has reached a sufficient level, and the learning of the recognizer 12_1 ends without the need for extraction by the extraction unit 114.
[0115] In contrast, if there is a misidentified object FTG1 ("Yes" in step S304), it is determined that the recognition performance of the recognizer 12_1 has not reached a sufficient level, and the learner continues to learn through the recognizer 12_1.
[0116] Specifically, firstly, the extraction unit 114 extracts the misidentified object FTG1 (No in step S305 → step S306). Then, the learning data generation unit 111 generates additional learning data (first learning data) including the image mapped to both the misidentified object FTG1 extracted by the extraction unit 114 and the identified object TG1 (step S307).
[0117] Subsequently, the learning control unit 112 uses additional learning data to perform additional training on the recognizer 12_1 (step S308). That is, the learning control unit 112 uses additional learning data to enable the recognizer 12_1 to learn additionally. Specifically, the learning control unit 112 uses additional learning data to enable the recognizer 12_1 to learn that the misidentified object FTG1 is an object of a different type from the identified object TG1.
[0118] Afterwards, the evaluation unit 113 evaluates the recognizer 12_1 again (step S303). Specifically, the evaluation unit 113 evaluates whether the recognizer 12_1 correctly recognizes the object TG1 as the object TG1 and whether it incorrectly recognizes the object TG3 as the object TG1.
[0119] Based on the evaluation results evaluated by the evaluation unit 113, it becomes clear whether there is an object FTG1 among the objects TG1 and TG3 that is misidentified as object TG1 in the recognizer 12_1 with a probability of more than a predetermined ratio (step S304).
[0120] Here, if there is no misidentified object FTG1 (No in step S304), it is determined that the recognition performance of the recognizer 12_1 has reached a sufficient level, and the learning of the recognizer 12_1 ends without the need for extraction by the extraction unit 114.
[0121] In contrast, if there is a misidentified object FTG1 ("Yes" in step S304), it is determined that the recognition performance of the recognizer 12_1 has not reached a sufficient level, and the learner continues to learn through the recognizer 12_1.
[0122] Then, the process of repeatedly extracting the misidentified object FTG1 (step S305 "No" → step S306), generating additional learning data (step S307), training the recognizer 12_1 using the additional learning data (step S308), and evaluating the recognizer 12_1 (step S303) continues until the misidentified object FTG1 disappears.
[0123] However, if the training time using supplementary learning data reaches a predetermined number of times before the misidentified object FTG1 disappears ("Yes" in step S304 → "Yes" in step S305), it is determined that the misidentified object FTG1 will not disappear, and this is notified to, for example, a user of this system (step S309). The user who receives the notification can take countermeasures such as changing the learning method.
[0124] Thus, in the learning process of the first recognizer for recognizing a predetermined type of first object, the learning device 11 of this embodiment only teaches the first recognizer to learn that a misidentified object is a different type of object from the first object if the probability of misidentification as the first object is greater than a predetermined ratio. Therefore, the learning device 11 can learn more efficiently than if the first recognizer learns that all objects other than the first object are different types of objects from the first object. As a result, the learning device 11 enables each recognizer of multiple recognizers that individually recognize multiple types of objects to learn efficiently without misidentification between the multiple types of objects.
[0125] Furthermore, in this embodiment, when the learning device 11 excludes learned recognition objects that are different from the first recognition object after learning by the first recognition device, it initializes the learning content of the excluded recognition objects from the first recognition device without having to determine whether the learning content of the recognition object of the first recognition object has disappeared, thus improving the performance of the first recognition device.
[0126] Furthermore, the present invention is not limited to the above-described embodiments and can be appropriately modified without departing from the spirit of the invention.
[0127] Furthermore, by having the CPU (Central Processing Unit) execute computer programs, this disclosure enables the implementation of part or all of the control processing in the conveying system.
[0128] The aforementioned program includes a group of instructions (or software code) used, when read by a computer, to cause the computer to perform one or more of the functions described in the embodiments. The program may also be stored on a non-transitory computer-readable medium or a physical storage medium. Not limited thereto, examples of computer-readable media or physical storage media include random-access memory (RAM), read-only memory (ROM), flash memory, solid-state drive (SSD) or other memory technologies, CD-ROM, digital versatile disc (DVD), Blu-ray disc or other optical disc storage devices, magnetic cartridges, magnetic tape, disk storage devices or other magnetic storage devices. The program may also be transmitted via a temporary computer-readable medium or a communication medium. Not limited thereto, examples of temporary computer-readable media or communication media include electrical, optical, acoustic, or other forms of transmission signals.
[0129] It will be apparent from the disclosure described herein that implementation of this disclosure can vary in many ways. Such variations should not be considered as departing from the spirit and scope of this disclosure, and it will be apparent to those skilled in the art that all such modifications are intended to be included within the scope of the appended claims.
Claims
1. A learning device that learns a first recognizer in a manner to recognize a first recognition object that is a first category of recognition objects from a plurality of categories of recognition objects, wherein The learning device includes: The extraction unit extracts a first misidentified object, which is an object among the plurality of identification objects that has a probability of being misidentified as the first identification object in the first identifier that is at least a predetermined ratio. The learning data generation unit generates first learning data, which includes images mapped to both the first identified object and the first misidentified object; as well as The learning control unit uses the first learning data to enable the first recognizer to learn that the first misidentified object is an object of a different kind from the first identified object.
2. The learning device according to claim 1, wherein, If the probability that the first misidentified object is misidentified as the first identified object in the first recognizer is less than the predetermined ratio, the learning control unit stops learning using the first learning data.
3. The learning device according to claim 1 or 2, wherein, It also has an evaluation department. The learning data generation unit also generates benchmark learning data, which includes multiple images of each of the multiple types of recognition objects individually mapped onto the images. The learning control unit uses the baseline learning data to enable the first recognizer to learn in advance in a way that correctly identifies the first recognition object as the first recognition object and in a way that does not incorrectly identify recognition objects other than the first recognition object as the first recognition object. The evaluation unit evaluates whether the first identifier correctly identifies the first identified object as the first identified object and whether it incorrectly identifies an object other than the first identified object as the first identified object. The extraction unit extracts the first misidentified object from the evaluation results evaluated by the evaluation unit.
4. The learning device according to claim 1 or 2, wherein, The learning data generation unit generates the first learning data, which includes multiple images mapped to both the first identified object and the first misidentified object. The learning control unit uses the first learning data to enable the first recognizer to learn that the first misidentified object is an object of a different kind from the first recognized object, until the probability of the first misidentified object being misidentified as the first recognized object in the first recognizer is less than the predetermined ratio.
5. The learning device according to claim 1 or 2, wherein, The learning data generation unit generates first learning data when it extracts from the extraction unit a plurality of first misidentified objects that are identified as first identification objects in the first recognizer at a predetermined rate or higher. The first learning data includes a plurality of images in which each of the first misidentified objects is mapped. The learning control unit uses the first learning data to enable the first recognizer to learn that multiple first misidentified objects are objects of a different kind from the first recognized object, until the probability of each of the multiple first misidentified objects being misidentified as the first recognized object in the first recognizer is less than the predetermined ratio.
6. The learning device according to claim 1 or 2, wherein, The second recognizer learns further by identifying a second object as a second type of object from among the multiple types of objects to be recognized. The extraction unit also extracts a second misidentified object, which is an object among the plurality of identification objects that has a probability of being misidentified as the second identification object in the second identifier that is greater than a predetermined percentage. The learning data generation unit also generates second learning data, which includes images mapped to both the second identified object and the second misidentified object. The learning control unit then uses the second learning data to enable the second recognizer to learn that the second misidentified object is an object of a different kind from the second identified object.
7. The learning device according to claim 6, wherein, If the probability that the second misidentified object is misidentified as the second identified object in the second recognizer is less than the predetermined ratio, the learning control unit stops learning using the second learning data.
8. The learning device according to claim 6, wherein, It also has an evaluation department. The learning data generation unit also generates benchmark learning data, which includes multiple images of each of the multiple types of recognition objects individually mapped onto the images. The learning control unit uses the baseline learning data to pre-learn the first recognizer in a way that correctly identifies the first recognition object as the first recognition object and in a way that does not incorrectly identify recognition objects other than the first recognition object as the first recognition object, and to pre-learn the second recognizer in a way that correctly identifies the second recognition object as the second recognition object and in a way that does not incorrectly identify recognition objects other than the second recognition object as the second recognition object. The evaluation unit evaluates whether the first recognizer correctly identifies the first identified object as the first identified object and whether it does not incorrectly identify an object other than the first identified object as the first identified object. It also evaluates whether the second recognizer correctly identifies the second identified object as the second identified object and whether it does not incorrectly identify an object other than the second identified object as the second identified object. The extraction unit extracts the first misidentified object and the second misidentified object from the evaluation results evaluated by the evaluation unit.
9. The learning device according to claim 7, wherein, It also has an evaluation department. The learning data generation unit also generates benchmark learning data, which includes multiple images of each of the multiple types of recognition objects individually mapped onto the images. The learning control unit uses the baseline learning data to pre-learn the first recognizer in a way that correctly identifies the first recognition object as the first recognition object and in a way that does not incorrectly identify recognition objects other than the first recognition object as the first recognition object, and to pre-learn the second recognizer in a way that correctly identifies the second recognition object as the second recognition object and in a way that does not incorrectly identify recognition objects other than the second recognition object as the second recognition object. The evaluation unit evaluates whether the first recognizer correctly identifies the first identified object as the first identified object and whether it does not incorrectly identify an object other than the first identified object as the first identified object. It also evaluates whether the second recognizer correctly identifies the second identified object as the second identified object and whether it does not incorrectly identify an object other than the second identified object as the second identified object. The extraction unit extracts the first misidentified object and the second misidentified object from the evaluation results evaluated by the evaluation unit.
10. The learning device according to claim 3, wherein, When a third identification object is added as a new third type of identification object to the plurality of identification objects, the evaluation unit is configured to further evaluate whether the first recognizer did not mistakenly identify the third identification object as the first identification object before learning that the third identification object is not the first identification object through the first recognizer. The extraction unit is configured to extract, from the evaluation results evaluated by the evaluation unit, the first misidentified object among the plurality of identification objects including the third identification object, which has a probability of being misidentified as the first identification object in the first identifier at a predetermined rate or higher.
11. The learning device according to claim 8, wherein, When a third identification object is added as a new third type of identification object to the plurality of identification objects, the evaluation unit is configured to further evaluate whether the first recognizer did not incorrectly identify the third identification object as the first identification object before learning through the first recognizer that the third identification object is not the first identification object, and whether the second recognizer did not incorrectly identify the third identification object as the second identification object before learning through the second recognizer that the third identification object is not the second identification object. The extraction unit is configured to extract from the evaluation results evaluated by the evaluation unit the first misidentified object, which is misidentified as the first identified object in the first identifier at a predetermined rate or higher, and the second misidentified object, which is misidentified as the second identified object in the second identifier at a predetermined rate or higher, from the plurality of identified objects including the third identified object.
12. The learning device according to claim 9, wherein, When a third identification object is added as a new third type of identification object to the plurality of identification objects, the evaluation unit is configured to further evaluate whether the first recognizer did not incorrectly identify the third identification object as the first identification object before learning through the first recognizer that the third identification object is not the first identification object, and whether the second recognizer did not incorrectly identify the third identification object as the second identification object before learning through the second recognizer that the third identification object is not the second identification object. The extraction unit is configured to extract from the evaluation results evaluated by the evaluation unit the first misidentified object, which is misidentified as the first identified object in the first identifier at a predetermined rate or higher, and the second misidentified object, which is misidentified as the second identified object in the second identifier at a predetermined rate or higher, from the plurality of identified objects including the third identified object.
13. The learning device according to claim 10, wherein, The learning data generation unit updates the baseline learning data by also including an image of the third recognition object, which is separately mapped from the multiple types of recognition objects. If, despite learning using the first learning data through the first recognizer, the probability of the first misidentified object being misidentified as the first recognized object in the first recognizer is maintained at or above the predetermined ratio, the learning control unit, after initializing the learning content using the first recognizer, causes the first recognizer to learn using at least the updated baseline learning data.
14. The learning device according to claim 11, wherein, The learning data generation unit updates the baseline learning data by also including an image of the third recognition object, which is separately mapped from the multiple types of recognition objects. If, despite learning using the first learning data through the first recognizer, the probability of the first misidentified object being misidentified as the first recognized object in the first recognizer remains above the predetermined ratio, the learning control unit, after initializing the learning content using the first recognizer, causes the first recognizer to learn using at least the updated baseline learning data; and if, despite learning using the second learning data through the second recognizer, the probability of the second misidentified object being misidentified as the second recognized object in the second recognizer remains above the predetermined ratio, the learning control unit, after initializing the learning content using the second recognizer, causes the second recognizer to learn using at least the updated baseline learning data.
15. The learning device according to any one of claims 10 to 14, wherein, The third recognizer learns further by recognizing the third object from the multiple types of recognized objects. The extraction unit also extracts a third misidentified object, which is an object among the plurality of identified objects that has a probability of being misidentified as the third identified object in the third recognizer that is greater than a predetermined percentage. The learning data generation unit also generates third learning data, which includes images mapped to both the third identified object and the third misidentified object. The learning control unit also uses the third learning data to enable the third recognizer to learn that the third misidentified object is an object of a different kind from the third identified object.
16. The learning device according to claim 1 or 2, wherein, When a fourth identification object, which is a fourth type of identification object, is excluded from the plurality of identification objects, the extraction unit, after the learning control unit has initialized the learning content learned using the first learning data with the first recognizer, extracts from the plurality of identification objects, excluding the fourth identification object, a first misidentified object whose probability of being misidentified as the first identification object in the first recognizer is greater than or equal to a predetermined percentage. The learning data generation unit generates the first learning data, which includes images mapped to both the first identified object and the newly extracted first misidentified object. In addition to initializing the learning content learned by the first recognizer using the first learning data, the learning control unit enables the first recognizer to learn using the newly generated first learning data.
17. The learning device according to claim 6, wherein, When a fourth identification object, which is a fourth type of identification object, is excluded from the plurality of identification objects, the extraction unit, in a state where the learning content learned using the first learning data by the first recognizer is initialized by the learning control unit, and in a state where the learning content learned using the second learning data by the second recognizer is initialized by the learning control unit, extracts from the plurality of identification objects other than the fourth identification object the first misidentified object whose probability of being misidentified as the first identification object in the first recognizer is greater than or equal to a predetermined percentage, and extracts from the second misidentified object whose probability of being misidentified as the second identification object in the second recognizer is greater than or equal to a predetermined percentage, the first misidentified object whose probability of being misidentified as the second identification object in the second recognizer. The learning data generation unit generates first learning data, which includes images mapped to both the first identified object and the newly extracted first misidentified object, and second learning data, which includes images mapped to both the second identified object and the newly extracted second misidentified object. In addition to initializing the learning content learned by the first recognizer using the first learning data, the learning control unit enables the first recognizer to learn using the newly generated first learning data, and in addition to initializing the learning content learned by the second recognizer using the second learning data, it enables the second recognizer to learn using the newly generated second learning data.
18. A method for controlling a learning device, wherein the learning device enables a first recognizer to learn in a manner that identifies a first recognition object as a first recognition object of a first category from multiple categories of recognition objects, wherein, The control method of the learning device: Extract the first misidentified object, which is an object among the multiple types of identification objects that has a probability of being misidentified as the first identification object in the first identifier that is greater than a predetermined percentage. Generate first learning data, which includes images mapped to both the first identified object and the first misidentified object. Using the first learning data, the first recognizer learns that the first misidentified object is an object of a different kind from the first identified object.
19. A computer-readable storage medium storing a control program for a learning device, the learning device enabling a first recognizer to learn in a manner that identifies a first recognizer as a first recognizer of a first type from multiple types of recognizer objects, wherein, This control program causes the computer to perform: The process of extracting the first misidentified object, which is an object among the plurality of identification objects that has a probability of being misidentified as the first identification object in the first identifier that is above a predetermined ratio; The process of generating the first learning data includes images mapped to both the first identified object and the first misidentified object; as well as Using the first learning data, the first recognizer learns that the first misidentified object is an object of a different kind than the first identified object.
Citation Information
Patent Citations
pattern recognition device
JP1994054503B2
Object detection improvement based on autonomously selected training samples
US20210326651A1
Training data generation device, training data generation system, training data generation method, and recording medium
WO2022024366A1